Skip to content

[NPU] Fix MTP draft loading for ModelSlim W8A8 checkpoints: dequantize int8 MTP weights - #37736

Open
Mr-qiji wants to merge 2 commits into
sgl-project:mainfrom
Mr-qiji:fix/npu-mtp-modelslim-w8a8-dequant
Open

Mr-qiji wants to merge 2 commits into
sgl-project:mainfrom
Mr-qiji:fix/npu-mtp-modelslim-w8a8-dequant

Conversation

@Mr-qiji

@Mr-qiji Mr-qiji commented Sep 3, 2026 •

Copy link
Copy Markdown

Motivation

Serving a ModelSlim-quantized Qwen3.5/3.8 checkpoint (main model W8A8_DYNAMIC) with MTP speculative decoding on NPU produces a numerically broken draft model: every drafted token is rejected (accept len: 1.02, accept rate: 0.01), so the full draft cost is paid for nothing (~15 tok/s vs 56 tok/s on the same 2-card server with a working draft).

Root cause.

  • On NPU, _mtp_quant_config() forces the MTP draft to quant_config=None when speculative_draft_model_quantization is unset (qwen3_5_mtp.py).
  • These checkpoints store the MTP projections as W8A8: mtp.layers.0.self_attn.q_proj.weight is int8 (12288, 5120) with a per-row weight_scale (12288, 1) and zero weight_offset; only norms/fc are FLOAT (per quant_model_description.json).
  • The loader therefore copies the raw int8 codes into bf16 parameters. Verified via a live weight dump: the 6 MTP projection matrices have absmax = 128 (the int8 code range) while correctly loaded mtp.fc.weight matches the checkpoint value exactly. The draft matmuls then run on garbage matrices and the target rejects every proposed token.

This is complementary to #34353: that PR covers checkpoints where all mtp.* entries are FLOAT; here the mtp.* projections are quantized, so that branch does not fire. A related loading failure is reported for compressed-tensors checkpoints in #35797.

Modifications

python/sglang/srt/models/qwen3_5_mtp.py:

  • Add _load_mtp_w8a8_dequant_scales(model_path): reads the mtp.*.weight_scale entries listed in quant_model_weights.safetensors.index.json (the ModelSlim index), keyed by the original weight name. Fails soft (warning + empty dict) on any error, so it can never block startup.
  • In load_weights, only when the draft runs unquantized (quant_config is None): if a weight's checkpoint name has a scale and the loaded tensor is int8, dequantize w = q.to(f32) * scale.to(f32) -> model dtype before name mapping / weight_loader. No-op otherwise (non-int8 dtype, missing scale, or scale/weight shape mismatch).

This matches vLLM-ascend behavior (MTP draft runs bf16). The change is effectively NPU-only: on CUDA the draft keeps the target's ModelSlim quant config, so the branch never fires.

Accuracy Tests

NPU verification (Ascend 910, 2 cards, TP2, Qwen3.8-27B-W8A8, EAGLE num_steps=3 topk=1 num_draft_tokens=4, sgl-kernel-npu 2026.6.1, sglang 0.5.17.dev):

Metric Before After vLLM (NPU, same model) reference
draft projection absmax 128 (raw int8 codes) 0.35–0.96 (correct) —
accept len 1.02 2.98–3.38 2.58–3.53
accept rate 0.01 0.66–0.79 per-position 0.36–0.93
throughput (2 cards) ~15 tok/s 31–56 tok/s accepted 3.7–55 tok/s
  • 512-token generation fully coherent (target model and mamba state commit unaffected by this change).
  • New CPU unit tests (test/registered/unit/models/test_qwen3_5_mtp_w8a8_dequant.py, no NPU needed): scale discovery (key naming, non-MTP exclusion, missing index, missing shard) and the load_weights branch (exact bf16 value, pass-through with a quant config, shape-mismatch no-op).

Speed Tests and Profiling

Same-server before/after comparison in the table above; no separate benchmark.

Checklist

Related


CI States

Latest PR Test (Base): ❌ Run #33726017268
Latest PR Test (Extra): ❌ Run #33726017018
Latest PR Test (AMD ROCm 7.2): ❌ Run #33726017107

@github-actions github-actions Bot added the quant LLM Quantization label Sep 3, 2026
@Mr-qiji

Mr-qiji commented Sep 3, 2026

Copy link
Copy Markdown
Author

Hi — this PR dequantizes W8A8 MTP draft weights for ModelSlim checkpoints on NPU (fixes spec-decode accept rate collapsing to ~0). Could a maintainer / Merge Oncall please add the run-ci label so the CI suite can run? Happy to address any review feedback. Thanks!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

quant LLM Quantization

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant