Skip to content

fix(qwen3.5): allow PP+MTP on disaggregated prefill - #39602

Closed
nvpohanh wants to merge 1 commit into
mainfrom
fix/qwen35-pp-mtp-prefill-validation
Closed

nvpohanh wants to merge 1 commit into
mainfrom
fix/qwen35-pp-mtp-prefill-validation

Conversation

@nvpohanh

@nvpohanh nvpohanh commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

[by Codex]

Motivation

Qwen3.5 gained the pipeline-parallel speculative-prefill runtime in #35758, but two older assertions still reject the supported configuration. The published PP4 AgentX recipes therefore have to set PYTHONOPTIMIZE=1 to start.

This removes that workaround for the supported path without opening unsupported PP+spec combinations.

Changes

  • Allow resolved EAGLE/NEXTN with PP only for Qwen3.5 architectures, non-overlap scheduling, and disaggregation-mode=prefill.
  • Keep decode, non-disaggregated, multi-layer EAGLE, other algorithms, and other model architectures rejected.
  • Exempt Qwen3.5 stage-local target layers from the older generic PP/MTP layer-count assertion.
  • Add server-argument and layer-compatibility regression tests.

Related: #39378 handles the corresponding DeepSeek/GLM path independently.

Validation

  • BLACK_NUM_WORKERS=1 SKIP=no-commit-to-branch pre-commit run --all-files --show-diff-on-failure
  • Focused layer compatibility tests: 2 passed.

CI States

Latest PR Test (Base): ❌ Run #34967629389
Latest PR Test (Extra): ❌ Run #34967629188
Latest PR Test (AMD ROCm 10): ❌ Run #34967629495

@nvpohanh

Copy link
Copy Markdown
Collaborator Author

cc @YAMY1234

@YAMY1234 YAMY1234 self-assigned this Sep 15, 2026
@YAMY1234

YAMY1234 commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Thanks for the fix. This overlaps almost entirely with #39378 : applying both yields conflicts in validation_hook.py, layer_setup.py and test_server_args.py, and the layer_setup change becomes dead code once #39378 removes the assertion. Suggest closing this PR and adding the four Qwen3_5* architectures to _PP_EAGLE_SUPPORTED_ARCHITECTURES in #39378?

YAMY1234 added a commit to nvpohanh/sglang that referenced this pull request Sep 18, 2026
Fold the Qwen3.5 allowlist from sgl-project#39602 into the shared
check_pipeline_parallel_compat gate so disaggregated-prefill PP + MTP
works for Qwen3.5 dense/MoE (text and multimodal) alongside DeepSeek/GLM.
Also covers the SGLANG_ENABLE_PP_SPEC aggregate branch that main added
in sgl-project#30775 and that the merge folded into the same function.
@nvpohanh nvpohanh closed this Sep 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants