Repository navigation
[Quark][MTP] Only drop the draft quant config when the MTP experts are excluded - #18
Closed
jiaqiang-dot-liu wants to merge 1 commit into
Closed
jiaqiang-dot-liu wants to merge 1 commit into
jiaqiang-dot-liu wants to merge 1 commit into
Conversation
…e excluded
Problem
`_mtp_quant_config` drops the Quark quant config for the whole MTP draft module if
**any** `mtp.*` entry appears in the checkpoint's `exclude` list:
```python
if any(isinstance(layer, str) and layer.startswith("mtp.") for layer in exclude_layers):
return None
```
That is right for checkpoints that ship the entire MTP module in bf16, which is the
case the current comment describes (sgl-project#23113).
It is wrong for checkpoints that exclude only the **dense** sub-modules. The
Qwen3.8-2.4T-A95B Quark MXFP4 export excludes `mtp.fc`,
`mtp.layers.0.self_attn.*`, `mtp.layers.0.mlp.gate`,
`mtp.layers.0.mlp.shared_expert*` — while still shipping MXFP4-**packed**
`mtp.layers.*.mlp.experts.*` tensors.
For those, returning `None` makes the fused-MoE loader allocate bf16
`[N, hidden]` and then fail copying the packed uint8 `[N, hidden // 2]`
checkpoint tensor.
Change
Only drop the quant config when the MTP **experts** are excluded too:
```python
mtp_excluded = [l for l in exclude_layers if isinstance(l, str) and l.startswith("mtp.")]
if mtp_excluded and any("mlp.experts" in l for l in mtp_excluded):
return None
```
Checkpoints that exclude the whole MTP module still match (their exclude list
contains the expert entries), so the sgl-project#23113 case is unchanged.
Verification
**Not tested.** No GPU or Python interpreter here, and I do not have either
checkpoint to check the exclude lists against.
The shape mismatch described above — bf16 `[N, hidden]` versus packed uint8
`[N, hidden // 2]` — is a 2x on the last dimension and matches the failure mode,
but I am taking the description of the Qwen3.8 export's exclude list on trust.
Before merging: dump `quantization_config.exclude` for both a whole-module-excluded
checkpoint and a dense-only-excluded one, and confirm the predicate classifies each
correctly.
Note on the substring match
`"mlp.experts" in layer` is a substring test, so it also matches a hypothetical
`mtp.layers.0.mlp.experts_something`. A stricter path-segment match would be
tighter; I kept the patch's form since it mirrors how the surrounding code already
matches prefixes.
Related
Five other agent sessions produced variants of this same fix
(`_quark_mtp_per_tensor_exclude`, `_mtp_quark_scoped_exclude_gate`,
`_qwen3_5_mtp_quark_partial_exclude_gate` x2,
`_qwen3_5_mtp_quark_routed_expert_quant_gate`), using regex segment matching or a
`should_ignore_layer` probe instead. This is the one that applies cleanly and makes
the smallest change; the others are recorded but not submitted.
---
*Provenance: originally authored by an automated kernel-optimization agent
(hyperloom session `20260816T070000Z`-ish, model `Qwen3.8-2.4T-A95B-Quark-MXFP4`);
applies cleanly to current `main`.*
jiaqiang-dot-liu
force-pushed
the
fix/mtp-quark-partial-exclude
branch
from
September 15, 2026 03:10
7200326 to
bd759e0
Compare
Owner
Author
|
Superseded upstream: sgl-project#39064 (0154f72) landed the same predicate (drop the MTP quant config only when mtp.*.mlp.experts are excluded), plus packed_modules_mapping and a unit test. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
_mtp_quant_configdrops the Quark quant config for the whole MTP draft module ifany
mtp.*entry appears in the checkpoint'sexcludelist:That is right for checkpoints that ship the entire MTP module in bf16, which is the
case the current comment describes (sgl-project#23113).
It is wrong for checkpoints that exclude only the dense sub-modules. The
Qwen3.8-2.4T-A95B Quark MXFP4 export excludes
mtp.fc,mtp.layers.0.self_attn.*,mtp.layers.0.mlp.gate,mtp.layers.0.mlp.shared_expert*— while still shipping MXFP4-packedmtp.layers.*.mlp.experts.*tensors.For those, returning
Nonemakes the fused-MoE loader allocate bf16[N, hidden]and then fail copying the packed uint8[N, hidden // 2]checkpoint tensor.
Change
Only drop the quant config when the MTP experts are excluded too:
Checkpoints that exclude the whole MTP module still match (their exclude list
contains the expert entries), so the sgl-project#23113 case is unchanged.
Verification
Not tested. No GPU or Python interpreter here, and I do not have either
checkpoint to check the exclude lists against.
The shape mismatch described above — bf16
[N, hidden]versus packed uint8[N, hidden // 2]— is a 2x on the last dimension and matches the failure mode,but I am taking the description of the Qwen3.8 export's exclude list on trust.
Before merging: dump
quantization_config.excludefor both a whole-module-excludedcheckpoint and a dense-only-excluded one, and confirm the predicate classifies each
correctly.
Note on the substring match
"mlp.experts" in layeris a substring test, so it also matches a hypotheticalmtp.layers.0.mlp.experts_something. A stricter path-segment match would betighter; I kept the patch's form since it mirrors how the surrounding code already
matches prefixes.
Related
Five other agent sessions produced variants of this same fix
(
_quark_mtp_per_tensor_exclude,_mtp_quark_scoped_exclude_gate,_qwen3_5_mtp_quark_partial_exclude_gatex2,_qwen3_5_mtp_quark_routed_expert_quant_gate), using regex segment matching or ashould_ignore_layerprobe instead. This is the one that applies cleanly and makesthe smallest change; the others are recorded but not submitted.
Provenance: originally authored by an automated kernel-optimization agent
(hyperloom session
20260816T070000Z-ish, modelQwen3.8-2.4T-A95B-Quark-MXFP4);applies cleanly to current
main.CI States
Latest PR Test (Base): ❌ Run #34923950761
Latest PR Test (Extra): ❌ Run #34923950432
Latest PR Test (AMD ROCm 10): ❌ Run #34923950926