Skip to content

[Bugfix] Infer the draft path for Qwen3.5 bundled MTP - #39069

Open
omirosh wants to merge 9 commits into
sgl-project:mainfrom
omirosh:fix/qwen3_5-bundled-mtp-draft-path
Open

omirosh wants to merge 9 commits into
sgl-project:mainfrom
omirosh:fix/qwen3_5-bundled-mtp-draft-path

Conversation

@omirosh

@omirosh omirosh commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

Qwen3.5 already bundles its MTP draft in the target checkpoint, so --speculative-draft-model-path is redundant, but omitting it currently fails the load. This PR adds the four Qwen3.5 architectures to the existing list of models that default the draft path to --model-path.

Modifications

Add the four Qwen3.5 architectures next to Qwen4ExpForConditionalGeneration, which is listed for the same reason, and make the warning name the architecture it fired for. It read "DeepSeek MTP does not require setting speculative_draft_model_path", already wrong for the GLM, Bailing, HunYuan and Qwen4-Exp entries.

This cannot regress an existing setup. The new entries only take effect when speculative_draft_model_path is None, which previously could not work for these architectures. Callers that already pass the flag keep the path they passed.

InternS2PreviewForConditionalGeneration, InternS2MobiusForConditionalGeneration and Qwen3MoeForCausalLM have MTP remaps too and are absent for the same reason, one line each if their owners confirm the published checkpoints bundle the draft. Left out because I cannot test them.

Accuracy Tests

None needed: this resolves a launch argument to the same path users pass by hand today. Checked on amd/Qwen3.8-2.4T-A95B-Quark-MXFP4 (Qwen3_5MoeForCausalLM, NEXTN 3/1/4, TP8).

Speed Tests and Profiling

None; argument resolution only, no runtime path changes.

Related: #39064, a separate fix for mixed Quark MXFP4 Qwen3.5 MTP checkpoints.

Checklist

  • Format your code according to the Format code with pre-commit.
  • Add unit tests according to the Run and add unit tests. _handle_eagle_family needs a resolved ModelConfig, so exercising it wants a real checkpoint config rather than a CPU unit test, and the change is four string literals. Happy to add one if you can point me at the preferred fixture.
  • Update documentation according to Write documentations. Nothing to document: this removes a required flag, it does not add one.
  • Provide accuracy and speed benchmark results according to Test the accuracy and Benchmark the speed. Not applicable, as noted above.
  • Follow the SGLang code style guidance.

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #37437644012
Latest PR Test (Extra): ❌ Run #37437643739
Latest PR Test (AMD ROCm 10): ❌ Run #37437644040

Qwen3.5 ships its MTP draft layer inside the target checkpoint. All four
target loaders in `models/qwen3_5.py` skip `mtp` weights on load, and the
draft worker picks them up through the `Qwen3_5ForCausalLMMTP` rewrite in
`ModelConfig._config_draft_model`.

`_handle_eagle_family` does not know this, so none of the four appear in the
list that defaults `speculative_draft_model_path` to `model_path`. Omitting
`--speculative-draft-model-path` therefore leaves it `None`, and it is passed
straight to `ModelConfig.from_server_args` as the draft worker's `model_path`
with no fallback. Unlike DFLASH and UNO, the EAGLE path has no "requires
setting" validation either, so the result is an obscure load failure rather
than a clear message.

Add the four architectures, next to `Qwen4ExpForConditionalGeneration`, which
was added for exactly the same reason. This can only turn a failure into a
working default: callers that already pass the flag keep using what they
passed.

Also make the accompanying warning name the architecture it fired for. It
read "DeepSeek MTP does not require setting speculative_draft_model_path",
which was already inaccurate for the GLM, Bailing, HunYuan and Qwen4-Exp
entries on the list.

Co-authored-by: Cursor <cursoragent@cursor.com>
@omirosh
omirosh force-pushed the fix/qwen3_5-bundled-mtp-draft-path branch from 5a8d8a5 to 0ca32f3 Compare September 11, 2026 09:09
@omirosh
omirosh marked this pull request as ready for review September 15, 2026 14:16
@omirosh

omirosh commented Sep 22, 2026

Copy link
Copy Markdown
Contributor Author

@hnyls2002 @Qiaolin-Yu could you review this? It is argument resolution only: four Qwen3.5 architectures now default --speculative-draft-model-path to --model-path when the flag is omitted, same as Qwen4-Exp. Callers who already pass the flag are unchanged. Happy to add a unit test if you have a preferred fixture.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant