Repository navigation
Conversation
twu3202
requested review from
BBuf,
Edwardf0t1,
Fridge003,
HaiShaw,
Qiaolin-Yu,
Ying1123,
ch-wan,
hnyls2002,
ispobock,
kpham-sgl,
merrymercy and
pyc96
as code owners
September 23, 2026 08:38
3 of 5 tasks
Author
|
@kpham-sgl you merged #35496, which added quantized lm_head support to the DFlash2 selector. This extends the same path to compressed-tensors packed heads, and fixes NEXTN on them too. Could you take a look, and add |
twu3202
force-pushed
the
fix-packed-lm-head-spec
branch
from
September 29, 2026 01:59
cfcd300 to
87bc4a3
Compare
A compressed-tensors pack-quantized lm_head stores its weight as weight_packed / weight_scale / weight_shape and has no .weight. The target serves fine, but NEXTN died in get_embed_and_head and the DFlash2 selector refused the head, because should_apply_lm_head_quant_method rejected any head without .weight (a guard for the ModelOpt dtype reads, from sgl-project#27906). - Admit heads without .weight through the gate; only the three ModelOpt methods still need the tensor for their layout checks. - _compute_lm_head: such heads take the quant_method.apply branch and keep the fp32 activation cast the old fallback gave them; the fallback becomes an explicit error. - qwen3_5 / qwen3_5_text accessors hand out None for a packed head; the eagle worker shares the module through set_lm_head_from_target and fails clearly when it cannot (lm_head_is_packed in spec_utils). Since sgl-project#39643 a pipeline stage runs the same init_lm_head. The multi-layer EAGLE worker, which shares the tensor only, fails clearly on a packed head. - Tests: the DFlash2 selector cases are parametrized over both packed layouts; a new CPU file covers the gate, the logits path, the accessors and init_lm_head.
twu3202
force-pushed
the
fix-packed-lm-head-spec
branch
from
October 3, 2026 13:58
87bc4a3 to
c5a6af2
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
A compressed-tensors checkpoint that quantizes
lm_headin the pack-quantized INT format stores the head asweight_packed/weight_scale/weight_shape; the module has no.weight. The target serves fine on its own, but both speculative paths that borrow the target head refuse it:--speculative-algorithm NEXTNdies at startup inget_embed_and_head:AttributeError: 'ParallelLMHead' object has no attribute 'weight'.--speculative-algorithm DFLASHwith a DFlash2 draft dies on the server's own warm-up request, so the server never becomes ready:DFlash2 selector requires a dense FP16/BF16/FP32 target lm_head or a supported lm_head.quant_method.Both come from
should_apply_lm_head_quant_methodrefusing any head without.weight. That check arrived with #27906 to keep the ModelOptlm_head.weight.dtypereads safe. A packed head is exactly the case wherequant_method.applyis the only way to get logits, and the target's ownLogitsProcessoralready does that through its trailing fallback branch. #35496 added the quant-method path to the DFlash2 selector, but the gate never let a packed head reach it. GGUF heads (qweight, no.weight) sit in the same position.Reproduced on main (525f140) with
Qwen/Qwen3.8-27Bquantized by llm-compressor so thatlm_headis pack-quantized; the recipe is under Accuracy Tests. The branch is now rebased onto a9c97c9; #39643 rewroteinit_lm_headin the spec workers along the way. The CPU tests were re-run on a9c97c9, and the GPU runs below are from 525f140.Modifications
layers/logits_processor.py:should_apply_lm_head_quant_methodno longer requires.weight. Only the three ModelOpt methods, whose branches readlm_head.weight.dtype, are refused when the tensor is missing. In_compute_lm_heada head without a dense weight now takes thequant_method.applybranch and keeps the fp32 activation cast the old fallback gave it under--enable-fp32-lm-head(that flag with a marlin-packed head was and stays unsupported); the old trailing fallback becomes an explicit error, since every head that used to reach it is now admitted by the gate. Only the second, bound definition of the gate is changed; [Fix] Drop the shadowed copies of the lm_head quant-method helpers #39707 removes the shadowed first copy.models/qwen3_5.py,models/qwen3_5_text.py:get_embed_and_headhands outNonefor a packed head instead of reading.weight.EagleDraftWorker.init_lm_headthen shares the whole module throughset_lm_head_from_target, whichQwen3_5ForCausalLMMTPalready implements.speculative/spec_utils.py:lm_head_is_packed(no.weight, and not aPPMissingLayer).speculative/eagle_worker_v2.py: when the target head is packed and the draft cannot take the module (noset_lm_head_from_target, or a hot-token map), fail with a clear error instead of decoding with the draft's own head. Since [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism #39643 a pipeline stage runs the sameinit_lm_head, so under PP a packed head is shared as a module too; I have not run PP, which needs two GPUs.dspark_draft_sampler.py,deepseek_v4_dspark.pyanddomino_utils.pystill readlm_head.weightand keep failing on a packed head as before. DSpark'sproject_through_lm_headgoes through the gate and is fixed along the way.speculative/multi_layer_eagle_worker_v2.py: it shares the head tensor only, so it fails with a clear error on a packed target, before loading anything, instead of handing the drafthead=None. Frozen-KV MTP is left alone: its only draft, Gemma 4, keeps its ownlm_head.qwen3_5_text.set_embed_and_headacceptsNonefor either half, matching what its accessor can now return.test_dflash_logits.pyare parametrized over a head that keeps packed bytes under.weightand one that has no.weightat all. New CPU filetest_spec_packed_target_lm_head.pycovers the gate (admission, and refusal of each ModelOpt method without a dense weight), the fp32 cast, the error path, the three accessors, andinit_lm_headsharing the module, refusing when it cannot, and still accepting a pipeline stage without a head. On main 8 of the 24 cases fail; the rest pin behaviour that must not change.Accuracy Tests
Checkpoints, both built from the official weights with llm-compressor's model-free PTQ. The target has a W4A16 body and a W8A16
lm_head, both pack-quantized; the in-checkpoint MTP layer stays BF16:The reference is the same call without the
lm_headgroup and with"lm_head"added toignore, so its head stays BF16. DFLASH uses the officialz-lab/Qwen3.8-27B-DFlash2draft.One RTX 6000 Ada (48 GB, TP=1),
--disable-overlap-schedule. NEXTN: steps 3, top-k 1, draft tokens 4. DFLASH: block size 8. Accuracy is SGLang's own GSM8K eval, one request at a time so completions can be compared across runs:lm_headAttributeError: 'ParallelLMHead' object has no attribute 'weight'at startupRuntimeError: DFlash2 selector requires a dense ...on the server's warm-up requestAccept length is the mean of the scheduler's
accept lenover every decode-log window of the run, bonus token included. Thesglang:spec_accept_lengthgauge only holds the last window, so it is not used here.Speculative decoding matches non-speculative decoding only up to numerics, so completions are compared against the same checkpoint without speculation. With the packed head, 172/200 NEXTN and 167/200 DFLASH completions match the packed-head target decoding on its own. On main, the BF16-head target matches its own non-speculative run in 167/200 (NEXTN) and 166/200 (DFLASH) completions.
Unit tests (CPU):
test_spec_packed_target_lm_head.py+test_dflash_logits.py23 passed (8 fail on main);test_eagle_draft_extend_logits.py,test_eagle_worker_v2_topk1_fastpath.py,models/test_qwen3_5_packed_weight_loader.py,models/test_qwen3_5_modelopt_fp4.py31 passed; thelm_head_guardcases ofmodel_loader/test_modelopt_loader.py4 passed.Speed Tests and Profiling
No change for dense or ModelOpt heads: the gate returns what it returned before and the same kernel runs. A packed head goes from refusing to start to serving; the draft calls the same
quant_method.applythe target already runs for its own logits. I have no clean timing for it yet, since the GPU these runs used was shared with other jobs.Checklist
CI States
Latest PR Test (Base): ❌ Run #37128008601
Latest PR Test (Extra): ❌ Run #37128008384
Latest PR Test (AMD ROCm 10): ❌ Run #37128008558