Skip to content

drafter: the fast variant for single-shard third-party exports (split from #172) - #181

Merged
mhenrichsen merged 2 commits into
syv-ai:mainfrom
Ar4ikov:drafter-single-shard
Sep 22, 2026
Merged

mhenrichsen merged 2 commits into
syv-ai:mainfrom
Ar4ikov:drafter-single-shard

Conversation

@Ar4ikov

@Ar4ikov Ar4ikov commented Sep 22, 2026

Copy link
Copy Markdown
Contributor

What

Split out of #172 at review: the drafter/ fixes, which need no GPU to review. Two files.

  1. drafter/gptq_lm_head.py builds the fast variant from a checkpoint that went through prepare/quant_heads_stream.py, not only from the base model's layout (seven model-0000x shards, a .bak from quant_lm_head.py, a quantization_config.json). A streamed export has model.safetensors + model-mtp.safetensors with .bak-orig backups and no quantization_config.json, so the script stopped at its first line (no .bak), and past that the model-0000* glob would have left model-mtp.safetensors, which the index points to, out of the variant. Now:

    • the bf16 lm_head comes from <shard>.bak or <shard>.bak-orig, with an assert naming both when neither is there;
    • every weight shard the source has is hardlinked (backups and the rewritten lm_head shard excluded), instead of model-0000*;
    • tokenizer.json, mtp_draft_vocab_ids.pt and the json files are taken only if they exist (the streamed exports also carry preprocessor_config.json, video_preprocessor_config.json and recipe.yaml, which now come along);
    • model_extra_tensors.safetensors is copied, not linked. prepare/build_draft_vocab.py rewrites it in the variant with save_file (the second step in drafter/README.md). safetensors 0.4.5, 0.5.3, 0.6.2 and 0.7.0 write that file in place, so through the hardlink the rewrite lands in the source dir as well. 0.8.0 replaces the file and leaves the source alone. Measured with a two-line hardlink test on each version. On a venv resolved before 0.8.0, the int8 draft head would have been overwritten with the variant's int4 one.
  2. drafter/capture.py sets VLLM_USE_V2_MODEL_RUNNER=0 (setdefault) and FLASHINFER_DISABLE_VERSION_CHECK=1 (both launchers export it; the capture runs standalone). The hooks patch the V1 runner (vllm.v1.worker.gpu_model_runner).

Verification

  • The -fast variants of both asymmetric AWQ checkpoints on the Hub (uncensored, base) were built with exactly these two files: the box copies differ from this diff in comments only. The run was capture.py, then gptq_lm_head.py --bits 4 --calib-rows 400000, then build_draft_vocab.py on the variant.
  • verify.sh's model section from main on both variant dirs: 7/7 PASS (lm_head requantized to int4, packed geometry matches, draft head present, three shards, no duplicates across shards). The source dirs keep their int8 lm_head and draft head.
  • python -m py_compile on both files.

Not done

  • No run on the base model's own layout after the change. Every path it takes is the old one: .bak exists, model-0000* shards match the new filter, and the listed json files exist.
  • On 0.28.0 the capture was not re-run. The only change there is two setdefaults, the first of which selects the runner 0.28.0 already selects.

🤖 Generated with Claude Code

…apture pinned to the V1 runner

gptq_lm_head.py assumed the base model's layout: seven model-0000x shards, a .bak from
quant_lm_head.py, a quantization_config.json to copy, a model_extra_tensors.safetensors to
link. A checkpoint that went through prepare/quant_heads_stream.py has model.safetensors +
model-mtp.safetensors with .bak-orig backups and no quantization_config.json, so the script
stopped at its first line (no .bak), and past that the model-0000* glob would have left
model-mtp.safetensors, which the index points to, out of the variant. It now reads the bf16
lm_head from .bak or .bak-orig, hardlinks every weight shard the source has, copies only
the files that exist, and copies model_extra_tensors.safetensors instead of linking it:
build_draft_vocab.py rewrites that file in the variant with save_file, and safetensors
0.4.5 through 0.7.0 write it in place, through the hardlink into the source dir (measured;
0.8.0 replaces the file and leaves the source alone).

capture.py hooks the V1 runner (vllm.v1.worker.gpu_model_runner). On the pinned 0.28.0
that is the runner this model gets anyway -- a hybrid architecture outside
DEFAULT_V2_MODEL_RUNNER_ARCHITECTURES stays on V1 -- so VLLM_USE_V2_MODEL_RUNNER=0 changes
nothing today. vLLM 0.29.0 defaults every model to V2, and there the hooks never fire: the
capture finishes with rows=0 and the GPTQ that follows calibrates on zeros (KL 0.00000,
round-trip error 1.0000, an lm_head of zeros), measured on the 0.29 port (syv-ai#148). No
speculation happens in a capture, so V1 is the right runner on both. It also sets
FLASHINFER_DISABLE_VERSION_CHECK=1, as both launchers do, since it runs standalone.

Measured on Ar4ikov/Qwen3.8-27B-Uncensored-AWQ-W4A16-ASYM after quant_heads_stream.py,
600k teacher-forced UltraChat tokens captured in 10 min on a 3090 (vLLM 0.29.0): RTN int4
lm_head KL 0.00701, GPTQ int4 0.00239 (the base model's published figures: 0.0068 and
0.0029).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ar4ikov added a commit to Ar4ikov/HyperQwen that referenced this pull request Sep 22, 2026
…y checkpoints

The llm-compressor AWQ exports of the uncensored finetune and of the base model, int4
asymmetric g128 with zero points and the vision tower kept, after quant_heads_stream.py
and build_draft_vocab.py; the -fast siblings with the int4-GPTQ lm_head (building one from
a single-shard export takes the drafter/ fixes in syv-ai#181). The measured rows (vLLM 0.29.0,
the syv-ai#148 port with this patch) for MTP, DFlash2, CTX=long, the production line, the fast
variant and batch mode, whose shipped 0.972 / 150k did not boot with the tower on (syv-ai#182).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ar4ikov added a commit to Ar4ikov/HyperQwen that referenced this pull request Sep 22, 2026
…y checkpoints

The llm-compressor AWQ exports of the uncensored finetune and of the base model, int4
asymmetric g128 with zero points and the vision tower kept, after quant_heads_stream.py
and build_draft_vocab.py; the -fast siblings with the int4-GPTQ lm_head (building one from
a single-shard export takes the drafter/ fixes in syv-ai#181). The measured rows (vLLM 0.29.0,
the syv-ai#148 port with this patch) for MTP, DFlash2, CTX=long, the production line, the fast
variant and batch mode, whose shipped 0.972 / 150k did not boot with the tower on (syv-ai#182).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
build_draft_vocab.py rewrites it with torch.save, which truncates the open
inode on every torch, so through the hardlink it rewrote the source's ids
while the source's draft head (now copied, not linked) kept the old rows.
@mhenrichsen

Copy link
Copy Markdown
Contributor

Merged, with one commit of mine on top — the same bug class you found, in the file next to it.

Your two claims, checked independently:

  • 0.29 defaults to the V2 runner. 0.28.0's config/vllm.py carries DEFAULT_V2_MODEL_RUNNER_ARCHITECTURES and a hybrid model outside it gets V1; 0.29.0's use_v2_model_runner has no such allowlist (only a ROCm carve-out). So on 0.29 the V1 hooks never fire and the capture calibrates on zeros — a silent all-zeros lm_head. The setdefault is right, and it is exactly the kind of thing Pin flip to vLLM 0.29.0: the series re-exported at fuzz 0, KVarN on 0.29, the #114 moves, a cold-boot profiling fix, and 0.28 vs 0.29 on every run setting (#106 part two) #148 would have broken without anyone noticing.
  • safetensors writes in place through a hardlink on 0.7. Two-line test on the reference box: with 0.7.0, save_file into the variant's linked path left the same inode and the source read back the variant's values; with 0.8.0 (what the box's venv resolves) the file is replaced and the source is untouched. Your version table, confirmed at both ends.

The one I added. mtp_draft_vocab_ids.pt was still hardlinked, and prepare/build_draft_vocab.py:126 rewrites it in the variant with torch.save — which truncates the open inode on every torch version, not just old ones:

torch 2.13.0 | same inode after save: True | SOURCE ids now: [9, 9, 9, 9, 9]

And your fix, on its own, would have made that case worse rather than better: before it both files were linked and got rewritten together, so the source ended up with the variant's head and the variant's ids — wrong, but consistent. With model_extra_tensors now copied and the ids still linked, the source would keep its own draft head and get the variant's id list, so its head rows and its ids would disagree after every variant build, on every version. The commit copies the ids file the same way (it is ~160 KB), and says why in the comment.

Tested the final block behaviourally on the box with a fake checkpoint in the streamed-export layout: shards (including model-mtp.safetensors) and tokenizer.json linked, extras and ids copied, the rewritten lm_head shard and its .bak-orig left behind, recipe.yaml carried — and after both of build_draft_vocab.py's rewrites in the variant, the source reads back unchanged. py_compile clean.

Thanks for splitting this out; it is a better PR on its own, and the V2-runner catch will save #148 a bad day.

@mhenrichsen
mhenrichsen merged commit f98de95 into syv-ai:main Sep 22, 2026
cpuchip added a commit to cpuchip/qwen38-27b-rtx3090 that referenced this pull request Sep 23, 2026
… marlin-int8-asym-zp on the 0.29 fork

Applied as-is to cpuchip/vllm qwen38/0.29-hq2 (291980422, author Nikita Davidchuk) and re-exported; PATCHES.md
row cut against 0.29.0 and the header names it with -hq2's other rows. Series (42) plus KVarN on v0.29.0 at
--fuzz 0 reproduce 291980422's vllm/ tree (0 differing files).
mhenrichsen added a commit to cpuchip/qwen38-27b-rtx3090 that referenced this pull request Sep 23, 2026
PATCHES.md is the one conflict: the 0.29 table kept, syv-ai#172's
marlin-int8-asym-zp row added. The 0.28-cut patch applies to 0.29.0 plus
this series at exact context (check_vllm_series.sh: 42 at exact context,
0 offset, 0 fuzz).
TyroneNel added a commit to TyroneNel/qwen38-27b-rtx3090 that referenced this pull request Sep 23, 2026
Takes upstream's 0.29.0 series wholesale (patches/series, PATCHES.md, the
re-cut patches, KVarN 0.29.0), including the four patches syv-ai#148 retired
(int4-mq3d-envs, sse-keep-alive, vllm-pr54282-draft-gumbel-salt,
xgrammar-spec-terminated).

Drops the fork's auth-deny-default.patch from the series and the tree: it is
cut against 0.28.0, both of its target files moved in 0.29.0
(serve/utils/server_utils.py -> serve/middleware/authenticate.py,
openai/cli_args.py -> launchers/cli_args.py), so it cannot apply at fuzz 0.
It returns with the 0.29 port in syv-ai#169.

Keeps from the fork: manual-only image builds and the fork's own buildcache
ref (docker-image.yml), the guarded resolver source in the bench scripts,
the F12 digest-pinned base, and F04 copy-only variant writes in
drafter/gptq_lm_head.py (upstream's syv-ai#181 file handling, copy semantics).
verify.sh equals upstream: every fork change to it landed via syv-ai#158/syv-ai#171/syv-ai#172.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants