drafter: the fast variant for single-shard third-party exports (split from #172) - #181
Conversation
…apture pinned to the V1 runner gptq_lm_head.py assumed the base model's layout: seven model-0000x shards, a .bak from quant_lm_head.py, a quantization_config.json to copy, a model_extra_tensors.safetensors to link. A checkpoint that went through prepare/quant_heads_stream.py has model.safetensors + model-mtp.safetensors with .bak-orig backups and no quantization_config.json, so the script stopped at its first line (no .bak), and past that the model-0000* glob would have left model-mtp.safetensors, which the index points to, out of the variant. It now reads the bf16 lm_head from .bak or .bak-orig, hardlinks every weight shard the source has, copies only the files that exist, and copies model_extra_tensors.safetensors instead of linking it: build_draft_vocab.py rewrites that file in the variant with save_file, and safetensors 0.4.5 through 0.7.0 write it in place, through the hardlink into the source dir (measured; 0.8.0 replaces the file and leaves the source alone). capture.py hooks the V1 runner (vllm.v1.worker.gpu_model_runner). On the pinned 0.28.0 that is the runner this model gets anyway -- a hybrid architecture outside DEFAULT_V2_MODEL_RUNNER_ARCHITECTURES stays on V1 -- so VLLM_USE_V2_MODEL_RUNNER=0 changes nothing today. vLLM 0.29.0 defaults every model to V2, and there the hooks never fire: the capture finishes with rows=0 and the GPTQ that follows calibrates on zeros (KL 0.00000, round-trip error 1.0000, an lm_head of zeros), measured on the 0.29 port (syv-ai#148). No speculation happens in a capture, so V1 is the right runner on both. It also sets FLASHINFER_DISABLE_VERSION_CHECK=1, as both launchers do, since it runs standalone. Measured on Ar4ikov/Qwen3.8-27B-Uncensored-AWQ-W4A16-ASYM after quant_heads_stream.py, 600k teacher-forced UltraChat tokens captured in 10 min on a 3090 (vLLM 0.29.0): RTN int4 lm_head KL 0.00701, GPTQ int4 0.00239 (the base model's published figures: 0.0068 and 0.0029). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y checkpoints The llm-compressor AWQ exports of the uncensored finetune and of the base model, int4 asymmetric g128 with zero points and the vision tower kept, after quant_heads_stream.py and build_draft_vocab.py; the -fast siblings with the int4-GPTQ lm_head (building one from a single-shard export takes the drafter/ fixes in syv-ai#181). The measured rows (vLLM 0.29.0, the syv-ai#148 port with this patch) for MTP, DFlash2, CTX=long, the production line, the fast variant and batch mode, whose shipped 0.972 / 150k did not boot with the tower on (syv-ai#182). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y checkpoints The llm-compressor AWQ exports of the uncensored finetune and of the base model, int4 asymmetric g128 with zero points and the vision tower kept, after quant_heads_stream.py and build_draft_vocab.py; the -fast siblings with the int4-GPTQ lm_head (building one from a single-shard export takes the drafter/ fixes in syv-ai#181). The measured rows (vLLM 0.29.0, the syv-ai#148 port with this patch) for MTP, DFlash2, CTX=long, the production line, the fast variant and batch mode, whose shipped 0.972 / 150k did not boot with the tower on (syv-ai#182). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
build_draft_vocab.py rewrites it with torch.save, which truncates the open inode on every torch, so through the hardlink it rewrote the source's ids while the source's draft head (now copied, not linked) kept the old rows.
|
Merged, with one commit of mine on top — the same bug class you found, in the file next to it. Your two claims, checked independently:
The one I added. And your fix, on its own, would have made that case worse rather than better: before it both files were linked and got rewritten together, so the source ended up with the variant's head and the variant's ids — wrong, but consistent. With Tested the final block behaviourally on the box with a fake checkpoint in the streamed-export layout: shards (including Thanks for splitting this out; it is a better PR on its own, and the V2-runner catch will save #148 a bad day. |
… marlin-int8-asym-zp on the 0.29 fork Applied as-is to cpuchip/vllm qwen38/0.29-hq2 (291980422, author Nikita Davidchuk) and re-exported; PATCHES.md row cut against 0.29.0 and the header names it with -hq2's other rows. Series (42) plus KVarN on v0.29.0 at --fuzz 0 reproduce 291980422's vllm/ tree (0 differing files).
PATCHES.md is the one conflict: the 0.29 table kept, syv-ai#172's marlin-int8-asym-zp row added. The 0.28-cut patch applies to 0.29.0 plus this series at exact context (check_vllm_series.sh: 42 at exact context, 0 offset, 0 fuzz).
Takes upstream's 0.29.0 series wholesale (patches/series, PATCHES.md, the re-cut patches, KVarN 0.29.0), including the four patches syv-ai#148 retired (int4-mq3d-envs, sse-keep-alive, vllm-pr54282-draft-gumbel-salt, xgrammar-spec-terminated). Drops the fork's auth-deny-default.patch from the series and the tree: it is cut against 0.28.0, both of its target files moved in 0.29.0 (serve/utils/server_utils.py -> serve/middleware/authenticate.py, openai/cli_args.py -> launchers/cli_args.py), so it cannot apply at fuzz 0. It returns with the 0.29 port in syv-ai#169. Keeps from the fork: manual-only image builds and the fork's own buildcache ref (docker-image.yml), the guarded resolver source in the bench scripts, the F12 digest-pinned base, and F04 copy-only variant writes in drafter/gptq_lm_head.py (upstream's syv-ai#181 file handling, copy semantics). verify.sh equals upstream: every fork change to it landed via syv-ai#158/syv-ai#171/syv-ai#172.
What
Split out of #172 at review: the
drafter/fixes, which need no GPU to review. Two files.drafter/gptq_lm_head.pybuilds the fast variant from a checkpoint that went throughprepare/quant_heads_stream.py, not only from the base model's layout (sevenmodel-0000xshards, a.bakfromquant_lm_head.py, aquantization_config.json). A streamed export hasmodel.safetensors+model-mtp.safetensorswith.bak-origbackups and noquantization_config.json, so the script stopped at its first line (no.bak), and past that themodel-0000*glob would have leftmodel-mtp.safetensors, which the index points to, out of the variant. Now:lm_headcomes from<shard>.bakor<shard>.bak-orig, with an assert naming both when neither is there;lm_headshard excluded), instead ofmodel-0000*;tokenizer.json,mtp_draft_vocab_ids.ptand the json files are taken only if they exist (the streamed exports also carrypreprocessor_config.json,video_preprocessor_config.jsonandrecipe.yaml, which now come along);model_extra_tensors.safetensorsis copied, not linked.prepare/build_draft_vocab.pyrewrites it in the variant withsave_file(the second step indrafter/README.md). safetensors 0.4.5, 0.5.3, 0.6.2 and 0.7.0 write that file in place, so through the hardlink the rewrite lands in the source dir as well. 0.8.0 replaces the file and leaves the source alone. Measured with a two-line hardlink test on each version. On a venv resolved before 0.8.0, the int8 draft head would have been overwritten with the variant's int4 one.drafter/capture.pysetsVLLM_USE_V2_MODEL_RUNNER=0(setdefault) andFLASHINFER_DISABLE_VERSION_CHECK=1(both launchers export it; the capture runs standalone). The hooks patch the V1 runner (vllm.v1.worker.gpu_model_runner).DEFAULT_V2_MODEL_RUNNER_ARCHITECTURESalready gets V1 (config/vllm.py,use_v2_model_runner).rows=0. The GPTQ that follows then calibrates on zeros: KL 0.00000, round-trip error 1.0000, anlm_headof zeros. It is here so the drafter survives Pin flip to vLLM 0.29.0: the series re-exported at fuzz 0, KVarN on 0.29, the #114 moves, a cold-boot profiling fix, and 0.28 vs 0.29 on every run setting (#106 part two) #148. No speculation happens in a capture, so V1 is the right runner on both.Verification
-fastvariants of both asymmetric AWQ checkpoints on the Hub (uncensored, base) were built with exactly these two files: the box copies differ from this diff in comments only. The run wascapture.py, thengptq_lm_head.py --bits 4 --calib-rows 400000, thenbuild_draft_vocab.pyon the variant.lm_headKL 0.00701, GPTQ int4 0.00239. The base model's published figures are 0.0068 and 0.0029.verify.sh's model section frommainon both variant dirs: 7/7 PASS (lm_head requantized to int4, packed geometry matches, draft head present, three shards, no duplicates across shards). The source dirs keep their int8lm_headand draft head.python -m py_compileon both files.Not done
.bakexists,model-0000*shards match the new filter, and the listed json files exist.setdefaults, the first of which selects the runner 0.28.0 already selects.🤖 Generated with Claude Code