Repository navigation
Fix disagg PP MTP for GLM-5.2 - #39378
Merged
YAMY1234 merged 11 commits intoSep 19, 2026
Merged
Conversation
nvpohanh
requested review from
ByronHsu,
Duyi-Wang,
HaiShaw,
merrymercy and
sogalin
as code owners
September 14, 2026 06:25
4 of 5 tasks
YAMY1234
approved these changes
Sep 14, 2026
YAMY1234
left a comment
Collaborator
There was a problem hiding this comment.
LGTM. Very clean and clear implementation.
Collaborator
|
/tag-and-rerun-ci |
YAMY1234
requested review from
Fridge003,
ShangmingCai,
Ying1123,
ch-wan,
fzyzcjy,
hnyls2002 and
ispobock
as code owners
September 15, 2026 05:48
10 tasks
Collaborator
This was referenced Sep 15, 2026
# Conflicts: # test/registered/unit/server_args/test_server_args.py
Fold the Qwen3.5 allowlist from sgl-project#39602 into the shared check_pipeline_parallel_compat gate so disaggregated-prefill PP + MTP works for Qwen3.5 dense/MoE (text and multimodal) alongside DeepSeek/GLM. Also covers the SGLANG_ENABLE_PP_SPEC aggregate branch that main added in sgl-project#30775 and that the merge folded into the same function.
Collaborator
This was referenced Sep 20, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Supersedes #39052 while @nvjullin is OOTO. The original two commits are preserved with Julien as their author; this branch rebases the change onto current
mainand addresses every review comment on the original PR.Motivation
GLM-5.2 could not run disaggregated PP prefill with MTP because:
embed_tokens.dsa_topk_indices), which could make decode ranks diverge between eager and CUDA-graph paths.Modifications
embed_tokenson the last DeepSeek/GLM PP stage when speculative decoding is enabled.dsa_topk_indicesthrough PP serialization and reconstruction.Nonecase).Accuracy Tests
Negative control on the exact parent commit (
f539c1fc65e3a0de79b2a0d08805b096fdc85865): launching the same GLM-5.2-NVFP4 PP4 disaggregated-prefill configuration with EAGLE steps 3 on a GB300 failed before server readiness incheck_server_args:This confirms that PP4 prefill with EAGLE did not start before this PR; the failure is the pre-change non-NPU PP + speculative-decoding restriction removed by this change.
GLM-5.2-NVFP4 on GB300 with Dynamo router:
--max-tokens 1638453e0691e21895a3863a606dfd12910c69eba94abThis is within 0.76 percentage points of the original PR's 95.45% steps1/steps3 result. The decode acceptance length is the mean of 143 scheduler samples and is higher than the original PR's reported 3.589.
Tests
BLACK_NUM_WORKERS=1 SKIP=no-commit-to-branch pre-commit run --all-files --show-diff-on-failureChecklist
CI States
Latest PR Test (Base): ✅ Run #35381952475
Latest PR Test (Extra): ❌ Run #35381952439
Latest PR Test (AMD ROCm 10): ❌ Run #35381952427