From bd759e0d39bba4c3f5913d79d1f5165c90100042 Mon Sep 17 00:00:00 2001 From: jiaqiang-dot-liu <295720280+jiaqiang-dot-liu@users.noreply.github.com> Date: Tue, 8 Sep 2026 17:01:12 +0800 Subject: [PATCH] [Quark][MTP] Only drop the draft quant config when the MTP experts are excluded MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Problem `_mtp_quant_config` drops the Quark quant config for the whole MTP draft module if **any** `mtp.*` entry appears in the checkpoint's `exclude` list: ```python if any(isinstance(layer, str) and layer.startswith("mtp.") for layer in exclude_layers): return None ``` That is right for checkpoints that ship the entire MTP module in bf16, which is the case the current comment describes (sgl-project/sglang#23113). It is wrong for checkpoints that exclude only the **dense** sub-modules. The Qwen3.8-2.4T-A95B Quark MXFP4 export excludes `mtp.fc`, `mtp.layers.0.self_attn.*`, `mtp.layers.0.mlp.gate`, `mtp.layers.0.mlp.shared_expert*` — while still shipping MXFP4-**packed** `mtp.layers.*.mlp.experts.*` tensors. For those, returning `None` makes the fused-MoE loader allocate bf16 `[N, hidden]` and then fail copying the packed uint8 `[N, hidden // 2]` checkpoint tensor. Change Only drop the quant config when the MTP **experts** are excluded too: ```python mtp_excluded = [l for l in exclude_layers if isinstance(l, str) and l.startswith("mtp.")] if mtp_excluded and any("mlp.experts" in l for l in mtp_excluded): return None ``` Checkpoints that exclude the whole MTP module still match (their exclude list contains the expert entries), so the #23113 case is unchanged. Verification **Not tested.** No GPU or Python interpreter here, and I do not have either checkpoint to check the exclude lists against. The shape mismatch described above — bf16 `[N, hidden]` versus packed uint8 `[N, hidden // 2]` — is a 2x on the last dimension and matches the failure mode, but I am taking the description of the Qwen3.8 export's exclude list on trust. Before merging: dump `quantization_config.exclude` for both a whole-module-excluded checkpoint and a dense-only-excluded one, and confirm the predicate classifies each correctly. Note on the substring match `"mlp.experts" in layer` is a substring test, so it also matches a hypothetical `mtp.layers.0.mlp.experts_something`. A stricter path-segment match would be tighter; I kept the patch's form since it mirrors how the surrounding code already matches prefixes. Related Five other agent sessions produced variants of this same fix (`_quark_mtp_per_tensor_exclude`, `_mtp_quark_scoped_exclude_gate`, `_qwen3_5_mtp_quark_partial_exclude_gate` x2, `_qwen3_5_mtp_quark_routed_expert_quant_gate`), using regex segment matching or a `should_ignore_layer` probe instead. This is the one that applies cleanly and makes the smallest change; the others are recorded but not submitted. --- *Provenance: originally authored by an automated kernel-optimization agent (hyperloom session `20260816T070000Z`-ish, model `Qwen3.8-2.4T-A95B-Quark-MXFP4`); applies cleanly to current `main`.* --- python/sglang/srt/models/qwen3_5_mtp.py | 21 ++++++++++++++------- 1 file changed, 14 insertions(+), 7 deletions(-) diff --git a/python/sglang/srt/models/qwen3_5_mtp.py b/python/sglang/srt/models/qwen3_5_mtp.py index 31d066a58a2e..717e14e21175 100644 --- a/python/sglang/srt/models/qwen3_5_mtp.py +++ b/python/sglang/srt/models/qwen3_5_mtp.py @@ -70,16 +70,23 @@ def _mtp_quant_config(quant_config): return None if is_npu() and get_spec().speculative_draft_model_quantization is None: return None - # Quark-quantized Qwen3.5 MXFP4 checkpoints ship the MTP module in bf16; - # every `mtp.*` layer appears under the quantization exclude list. Detect - # that and skip quantization here so linear/MoE weight loaders allocate - # bf16 shapes (see sgl-project/sglang#23113). + # Quark-quantized Qwen3.5 MXFP4 checkpoints vary. Some ship the whole MTP + # module in bf16 (every `mtp.*` layer under the quantization exclude list), + # but others (e.g. Qwen3.8-2.4T-A95B-Quark-MXFP4) exclude only the dense + # sub-modules -- mtp.fc, mtp.layers.0.self_attn.*, mtp.layers.0.mlp.gate, + # mtp.layers.0.mlp.shared_expert* -- while still shipping MXFP4-packed + # `mtp.layers.*.mlp.experts.*` tensors. Only drop the quant config when the + # MTP *experts* are excluded too; otherwise the fused-MoE loader allocates + # bf16 [N, hidden] and fails copying the packed uint8 [N, hidden // 2] + # checkpoint tensor (see sgl-project/sglang#23113). if quant_config and quant_config.get_name() == "quark": exclude_layers = getattr(quant_config, "exclude_layers", []) - if any( - isinstance(layer, str) and layer.startswith("mtp.") + mtp_excluded = [ + layer for layer in exclude_layers - ): + if isinstance(layer, str) and layer.startswith("mtp.") + ] + if mtp_excluded and any("mlp.experts" in layer for layer in mtp_excluded): return None return quant_config