Skip to content

[Quark][MTP] Only drop the draft quant config when the MTP experts are excluded - #18

Closed
jiaqiang-dot-liu wants to merge 1 commit into
mainfrom
fix/mtp-quark-partial-exclude
Closed

jiaqiang-dot-liu wants to merge 1 commit into
mainfrom
fix/mtp-quark-partial-exclude

Conversation

@jiaqiang-dot-liu

@jiaqiang-dot-liu jiaqiang-dot-liu commented Sep 8, 2026 •

Copy link
Copy Markdown
Owner

Problem

_mtp_quant_config drops the Quark quant config for the whole MTP draft module if
any mtp.* entry appears in the checkpoint's exclude list:

if any(isinstance(layer, str) and layer.startswith("mtp.") for layer in exclude_layers):
    return None

That is right for checkpoints that ship the entire MTP module in bf16, which is the
case the current comment describes (sgl-project#23113).

It is wrong for checkpoints that exclude only the dense sub-modules. The
Qwen3.8-2.4T-A95B Quark MXFP4 export excludes mtp.fc,
mtp.layers.0.self_attn.*, mtp.layers.0.mlp.gate,
mtp.layers.0.mlp.shared_expert* — while still shipping MXFP4-packed
mtp.layers.*.mlp.experts.* tensors.

For those, returning None makes the fused-MoE loader allocate bf16
[N, hidden] and then fail copying the packed uint8 [N, hidden // 2]
checkpoint tensor.

Change

Only drop the quant config when the MTP experts are excluded too:

mtp_excluded = [l for l in exclude_layers if isinstance(l, str) and l.startswith("mtp.")]
if mtp_excluded and any("mlp.experts" in l for l in mtp_excluded):
    return None

Checkpoints that exclude the whole MTP module still match (their exclude list
contains the expert entries), so the sgl-project#23113 case is unchanged.

Verification

Not tested. No GPU or Python interpreter here, and I do not have either
checkpoint to check the exclude lists against.

The shape mismatch described above — bf16 [N, hidden] versus packed uint8
[N, hidden // 2] — is a 2x on the last dimension and matches the failure mode,
but I am taking the description of the Qwen3.8 export's exclude list on trust.

Before merging: dump quantization_config.exclude for both a whole-module-excluded
checkpoint and a dense-only-excluded one, and confirm the predicate classifies each
correctly.

Note on the substring match

"mlp.experts" in layer is a substring test, so it also matches a hypothetical
mtp.layers.0.mlp.experts_something. A stricter path-segment match would be
tighter; I kept the patch's form since it mirrors how the surrounding code already
matches prefixes.

Related

Five other agent sessions produced variants of this same fix
(_quark_mtp_per_tensor_exclude, _mtp_quark_scoped_exclude_gate,
_qwen3_5_mtp_quark_partial_exclude_gate x2,
_qwen3_5_mtp_quark_routed_expert_quant_gate), using regex segment matching or a
should_ignore_layer probe instead. This is the one that applies cleanly and makes
the smallest change; the others are recorded but not submitted.


Provenance: originally authored by an automated kernel-optimization agent
(hyperloom session 20260816T070000Z-ish, model Qwen3.8-2.4T-A95B-Quark-MXFP4);
applies cleanly to current main.


CI States

Latest PR Test (Base): ❌ Run #34923950761
Latest PR Test (Extra): ❌ Run #34923950432
Latest PR Test (AMD ROCm 10): ❌ Run #34923950926

…e excluded

Problem

`_mtp_quant_config` drops the Quark quant config for the whole MTP draft module if
**any** `mtp.*` entry appears in the checkpoint's `exclude` list:

```python
if any(isinstance(layer, str) and layer.startswith("mtp.") for layer in exclude_layers):
    return None
```

That is right for checkpoints that ship the entire MTP module in bf16, which is the
case the current comment describes (sgl-project#23113).

It is wrong for checkpoints that exclude only the **dense** sub-modules. The
Qwen3.8-2.4T-A95B Quark MXFP4 export excludes `mtp.fc`,
`mtp.layers.0.self_attn.*`, `mtp.layers.0.mlp.gate`,
`mtp.layers.0.mlp.shared_expert*` — while still shipping MXFP4-**packed**
`mtp.layers.*.mlp.experts.*` tensors.

For those, returning `None` makes the fused-MoE loader allocate bf16
`[N, hidden]` and then fail copying the packed uint8 `[N, hidden // 2]`
checkpoint tensor.

Change

Only drop the quant config when the MTP **experts** are excluded too:

```python
mtp_excluded = [l for l in exclude_layers if isinstance(l, str) and l.startswith("mtp.")]
if mtp_excluded and any("mlp.experts" in l for l in mtp_excluded):
    return None
```

Checkpoints that exclude the whole MTP module still match (their exclude list
contains the expert entries), so the sgl-project#23113 case is unchanged.

Verification

**Not tested.** No GPU or Python interpreter here, and I do not have either
checkpoint to check the exclude lists against.

The shape mismatch described above — bf16 `[N, hidden]` versus packed uint8
`[N, hidden // 2]` — is a 2x on the last dimension and matches the failure mode,
but I am taking the description of the Qwen3.8 export's exclude list on trust.

Before merging: dump `quantization_config.exclude` for both a whole-module-excluded
checkpoint and a dense-only-excluded one, and confirm the predicate classifies each
correctly.

Note on the substring match

`"mlp.experts" in layer` is a substring test, so it also matches a hypothetical
`mtp.layers.0.mlp.experts_something`. A stricter path-segment match would be
tighter; I kept the patch's form since it mirrors how the surrounding code already
matches prefixes.

Related

Five other agent sessions produced variants of this same fix
(`_quark_mtp_per_tensor_exclude`, `_mtp_quark_scoped_exclude_gate`,
`_qwen3_5_mtp_quark_partial_exclude_gate` x2,
`_qwen3_5_mtp_quark_routed_expert_quant_gate`), using regex segment matching or a
`should_ignore_layer` probe instead. This is the one that applies cleanly and makes
the smallest change; the others are recorded but not submitted.

---

*Provenance: originally authored by an automated kernel-optimization agent
(hyperloom session `20260816T070000Z`-ish, model `Qwen3.8-2.4T-A95B-Quark-MXFP4`);
applies cleanly to current `main`.*
@jiaqiang-dot-liu
jiaqiang-dot-liu force-pushed the fix/mtp-quark-partial-exclude branch from 7200326 to bd759e0 Compare September 15, 2026 03:10
@jiaqiang-dot-liu

Copy link
Copy Markdown
Owner Author

Superseded upstream: sgl-project#39064 (0154f72) landed the same predicate (drop the MTP quant config only when mtp.*.mlp.experts are excluded), plus packed_modules_mapping and a unit test.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant