Skip to content

[Spec] Let NEXTN and the DFlash2 selector reuse a packed target lm_head - #40883

Open
twu3202 wants to merge 1 commit into
sgl-project:mainfrom
twu3202:fix-packed-lm-head-spec
Open

twu3202 wants to merge 1 commit into
sgl-project:mainfrom
twu3202:fix-packed-lm-head-spec

Conversation

@twu3202

@twu3202 twu3202 commented Sep 23, 2026 •

Copy link
Copy Markdown

Motivation

A compressed-tensors checkpoint that quantizes lm_head in the pack-quantized INT format stores the head as weight_packed / weight_scale / weight_shape; the module has no .weight. The target serves fine on its own, but both speculative paths that borrow the target head refuse it:

  • --speculative-algorithm NEXTN dies at startup in get_embed_and_head: AttributeError: 'ParallelLMHead' object has no attribute 'weight'.
  • --speculative-algorithm DFLASH with a DFlash2 draft dies on the server's own warm-up request, so the server never becomes ready: DFlash2 selector requires a dense FP16/BF16/FP32 target lm_head or a supported lm_head.quant_method.

Both come from should_apply_lm_head_quant_method refusing any head without .weight. That check arrived with #27906 to keep the ModelOpt lm_head.weight.dtype reads safe. A packed head is exactly the case where quant_method.apply is the only way to get logits, and the target's own LogitsProcessor already does that through its trailing fallback branch. #35496 added the quant-method path to the DFlash2 selector, but the gate never let a packed head reach it. GGUF heads (qweight, no .weight) sit in the same position.

Reproduced on main (525f140) with Qwen/Qwen3.8-27B quantized by llm-compressor so that lm_head is pack-quantized; the recipe is under Accuracy Tests. The branch is now rebased onto a9c97c9; #39643 rewrote init_lm_head in the spec workers along the way. The CPU tests were re-run on a9c97c9, and the GPU runs below are from 525f140.

Modifications

  • layers/logits_processor.py: should_apply_lm_head_quant_method no longer requires .weight. Only the three ModelOpt methods, whose branches read lm_head.weight.dtype, are refused when the tensor is missing. In _compute_lm_head a head without a dense weight now takes the quant_method.apply branch and keeps the fp32 activation cast the old fallback gave it under --enable-fp32-lm-head (that flag with a marlin-packed head was and stays unsupported); the old trailing fallback becomes an explicit error, since every head that used to reach it is now admitted by the gate. Only the second, bound definition of the gate is changed; [Fix] Drop the shadowed copies of the lm_head quant-method helpers #39707 removes the shadowed first copy.
  • models/qwen3_5.py, models/qwen3_5_text.py: get_embed_and_head hands out None for a packed head instead of reading .weight. EagleDraftWorker.init_lm_head then shares the whole module through set_lm_head_from_target, which Qwen3_5ForCausalLMMTP already implements.
  • speculative/spec_utils.py: lm_head_is_packed (no .weight, and not a PPMissingLayer). speculative/eagle_worker_v2.py: when the target head is packed and the draft cannot take the module (no set_lm_head_from_target, or a hot-token map), fail with a clear error instead of decoding with the draft's own head. Since [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism #39643 a pipeline stage runs the same init_lm_head, so under PP a packed head is shared as a module too; I have not run PP, which needs two GPUs.
  • Out of scope: dspark_draft_sampler.py, deepseek_v4_dspark.py and domino_utils.py still read lm_head.weight and keep failing on a packed head as before. DSpark's project_through_lm_head goes through the gate and is fixed along the way.
  • speculative/multi_layer_eagle_worker_v2.py: it shares the head tensor only, so it fails with a clear error on a packed target, before loading anything, instead of handing the draft head=None. Frozen-KV MTP is left alone: its only draft, Gemma 4, keeps its own lm_head. qwen3_5_text.set_embed_and_head accepts None for either half, matching what its accessor can now return.
  • Tests: the two DFlash2 selector cases in test_dflash_logits.py are parametrized over a head that keeps packed bytes under .weight and one that has no .weight at all. New CPU file test_spec_packed_target_lm_head.py covers the gate (admission, and refusal of each ModelOpt method without a dense weight), the fp32 cast, the error path, the three accessors, and init_lm_head sharing the module, refusing when it cannot, and still accepting a pipeline stage without a head. On main 8 of the 24 cases fail; the rest pin behaviour that must not change.

Accuracy Tests

Checkpoints, both built from the official weights with llm-compressor's model-free PTQ. The target has a W4A16 body and a W8A16 lm_head, both pack-quantized; the in-checkpoint MTP layer stays BF16:

# llmcompressor==0.14.0, data-free round-to-nearest, runs on CPU
from compressed_tensors.quantization import QuantizationConfig, QuantizationScheme
from compressed_tensors.quantization.quant_scheme import W4A16, W8A16
from llmcompressor import model_free_ptq

model_free_ptq(
    model_stub="Qwen/Qwen3.8-27B",
    save_directory="Qwen3.8-27B-W4A16-lm_head-W8A16",
    config=QuantizationConfig(config_groups={
        "lm_head": QuantizationScheme(targets=["re:.*lm_head$"], **W8A16),  # first match wins
        "body": QuantizationScheme(targets=["Linear"], **W4A16),
    }),
    ignore=["re:.*visual.*", "re:.*mtp.*", "re:.*in_proj_(a|b)$", "re:.*conv1d.*",
            "re:.*norm.*", "re:.*embed_tokens.*"],
)

The reference is the same call without the lm_head group and with "lm_head" added to ignore, so its head stays BF16. DFLASH uses the official z-lab/Qwen3.8-27B-DFlash2 draft.

One RTX 6000 Ada (48 GB, TP=1), --disable-overlap-schedule. NEXTN: steps 3, top-k 1, draft tokens 4. DFLASH: block size 8. Accuracy is SGLang's own GSM8K eval, one request at a time so completions can be compared across runs:

python3 -m sglang.test.run_eval --eval-name gsm8k --num-examples 200 --num-threads 1 --temperature 0 --max-tokens 512
target lm_head branch spec result GSM8K (200) accept length
BF16 main none serves 0.960
BF16 main NEXTN serves 0.970 3.629
BF16 main DFLASH serves 0.970 6.035
W8A16 packed main none serves 0.965
W8A16 packed main NEXTN AttributeError: 'ParallelLMHead' object has no attribute 'weight' at startup
W8A16 packed main DFLASH RuntimeError: DFlash2 selector requires a dense ... on the server's warm-up request
W8A16 packed this PR NEXTN serves 0.965 3.629
W8A16 packed this PR DFLASH serves; the selector is folded into the draft CUDA graph 0.970 6.035

Accept length is the mean of the scheduler's accept len over every decode-log window of the run, bonus token included. The sglang:spec_accept_length gauge only holds the last window, so it is not used here.

Speculative decoding matches non-speculative decoding only up to numerics, so completions are compared against the same checkpoint without speculation. With the packed head, 172/200 NEXTN and 167/200 DFLASH completions match the packed-head target decoding on its own. On main, the BF16-head target matches its own non-speculative run in 167/200 (NEXTN) and 166/200 (DFLASH) completions.

Unit tests (CPU): test_spec_packed_target_lm_head.py + test_dflash_logits.py 23 passed (8 fail on main); test_eagle_draft_extend_logits.py, test_eagle_worker_v2_topk1_fastpath.py, models/test_qwen3_5_packed_weight_loader.py, models/test_qwen3_5_modelopt_fp4.py 31 passed; the lm_head_guard cases of model_loader/test_modelopt_loader.py 4 passed.

Speed Tests and Profiling

No change for dense or ModelOpt heads: the gate returns what it returned before and the same kernel runs. A packed head goes from refusing to start to serving; the draft calls the same quant_method.apply the target already runs for its own logits. I have no clean timing for it yet, since the GPU these runs used was shared with other jobs.

Checklist


CI States

Latest PR Test (Base): ❌ Run #37128008601
Latest PR Test (Extra): ❌ Run #37128008384
Latest PR Test (AMD ROCm 10): ❌ Run #37128008558

@twu3202

twu3202 commented Sep 27, 2026

Copy link
Copy Markdown
Author

@kpham-sgl you merged #35496, which added quantized lm_head support to the DFlash2 selector. This extends the same path to compressed-tensors packed heads, and fixes NEXTN on them too. Could you take a look, and add run-ci if it looks reasonable? Repro and GSM8K numbers are in the description.

@twu3202
twu3202 force-pushed the fix-packed-lm-head-spec branch from cfcd300 to 87bc4a3 Compare September 29, 2026 01:59
A compressed-tensors pack-quantized lm_head stores its weight as
weight_packed / weight_scale / weight_shape and has no .weight. The target
serves fine, but NEXTN died in get_embed_and_head and the DFlash2 selector
refused the head, because should_apply_lm_head_quant_method rejected any
head without .weight (a guard for the ModelOpt dtype reads, from sgl-project#27906).

- Admit heads without .weight through the gate; only the three ModelOpt
  methods still need the tensor for their layout checks.
- _compute_lm_head: such heads take the quant_method.apply branch and keep
  the fp32 activation cast the old fallback gave them; the fallback becomes
  an explicit error.
- qwen3_5 / qwen3_5_text accessors hand out None for a packed head; the
  eagle worker shares the module through set_lm_head_from_target and fails
  clearly when it cannot (lm_head_is_packed in spec_utils). Since sgl-project#39643 a
  pipeline stage runs the same init_lm_head. The multi-layer EAGLE worker,
  which shares the tensor only, fails clearly on a packed head.
- Tests: the DFlash2 selector cases are parametrized over both packed
  layouts; a new CPU file covers the gate, the logits path, the accessors
  and init_lm_head.
@twu3202
twu3202 force-pushed the fix-packed-lm-head-spec branch from 87bc4a3 to c5a6af2 Compare October 3, 2026 13:58

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant