Repository navigation
Conversation
Author
|
Hi — this PR dequantizes W8A8 MTP draft weights for ModelSlim checkpoints on NPU (fixes spec-decode accept rate collapsing to ~0). Could a maintainer / Merge Oncall please add the run-ci label so the CI suite can run? Happy to address any review feedback. Thanks! |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Serving a ModelSlim-quantized Qwen3.5/3.8 checkpoint (main model
W8A8_DYNAMIC) with MTP speculative decoding on NPU produces a numerically broken draft model: every drafted token is rejected (accept len: 1.02, accept rate: 0.01), so the full draft cost is paid for nothing (~15 tok/s vs 56 tok/s on the same 2-card server with a working draft).Root cause.
_mtp_quant_config()forces the MTP draft toquant_config=Nonewhenspeculative_draft_model_quantizationis unset (qwen3_5_mtp.py).mtp.layers.0.self_attn.q_proj.weightisint8(12288, 5120) with a per-rowweight_scale(12288, 1) and zeroweight_offset; only norms/fcareFLOAT(perquant_model_description.json).absmax = 128(the int8 code range) while correctly loadedmtp.fc.weightmatches the checkpoint value exactly. The draft matmuls then run on garbage matrices and the target rejects every proposed token.This is complementary to #34353: that PR covers checkpoints where all
mtp.*entries areFLOAT; here themtp.*projections are quantized, so that branch does not fire. A related loading failure is reported for compressed-tensors checkpoints in #35797.Modifications
python/sglang/srt/models/qwen3_5_mtp.py:_load_mtp_w8a8_dequant_scales(model_path): reads themtp.*.weight_scaleentries listed inquant_model_weights.safetensors.index.json(the ModelSlim index), keyed by the original weight name. Fails soft (warning + empty dict) on any error, so it can never block startup.load_weights, only when the draft runs unquantized (quant_config is None): if a weight's checkpoint name has a scale and the loaded tensor isint8, dequantizew = q.to(f32) * scale.to(f32) -> model dtypebefore name mapping /weight_loader. No-op otherwise (non-int8 dtype, missing scale, or scale/weight shape mismatch).This matches vLLM-ascend behavior (MTP draft runs bf16). The change is effectively NPU-only: on CUDA the draft keeps the target's ModelSlim quant config, so the branch never fires.
Accuracy Tests
NPU verification (Ascend 910, 2 cards, TP2, Qwen3.8-27B-W8A8, EAGLE
num_steps=3 topk=1 num_draft_tokens=4, sgl-kernel-npu 2026.6.1, sglang 0.5.17.dev):test/registered/unit/models/test_qwen3_5_mtp_w8a8_dequant.py, no NPU needed): scale discovery (key naming, non-MTP exclusion, missing index, missing shard) and theload_weightsbranch (exact bf16 value, pass-through with a quant config, shape-mismatch no-op).Speed Tests and Profiling
Same-server before/after comparison in the table above; no separate benchmark.
Checklist
Related
FLOATMTP case + UB-aware mamba scatter (complementary; this PR covers the W8A8 MTP case)CI States
Latest PR Test (Base): ❌ Run #33726017268
Latest PR Test (Extra): ❌ Run #33726017018
Latest PR Test (AMD ROCm 7.2): ❌ Run #33726017107