Skip to content

Fix disagg PP MTP for GLM-5.2 - #39378

Merged
YAMY1234 merged 11 commits into
sgl-project:mainfrom
nvpohanh:codex/fix-disagg-pp-mtp-glm52-review
Sep 19, 2026
Merged

YAMY1234 merged 11 commits into
sgl-project:mainfrom
nvpohanh:codex/fix-disagg-pp-mtp-glm52-review

Conversation

@nvpohanh

@nvpohanh nvpohanh commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator

Supersedes #39052 while @nvjullin is OOTO. The original two commits are preserved with Julien as their author; this branch rebases the change onto current main and addresses every review comment on the original PR.

Motivation

GLM-5.2 could not run disaggregated PP prefill with MTP because:

  1. PP + speculative decoding was rejected outside NPU even though the prefill path already forwards PP proxy tensors from stages that do not host the draft.
  2. The NextN draft is hosted on the last PP stage and borrows the target embedding, but only the first PP stage materialized and loaded embed_tokens.
  3. The PP ring dropped the draft model's DSA IndexShare seed (dsa_topk_indices), which could make decode ranks diverge between eager and CUDA-graph paths.
  4. NIXL needed to map each heterogeneous PP prefill stage to the corresponding layer span in an unpipelined decode worker's MLA pool.

Modifications

  • Allow single-layer EAGLE with PP only for disaggregated prefill on the DeepSeek/GLM model path that supplies the last-stage draft embedding. Decode/null PP + speculative decoding and unsupported model families remain rejected.
  • Materialize and load embed_tokens on the last DeepSeek/GLM PP stage when speculative decoding is enabled.
  • Preserve dsa_topk_indices through PP serialization and reconstruction.
  • Map heterogeneous plain-MLA NIXL transfers through one destination-index dispatcher, while retaining layer-ID/positional pairing for other layouts.
  • Add regression coverage for PP/MLA transfer mapping, PP/EAGLE validation, embedding ownership, and DSA seed round trips (including the None case).

Accuracy Tests

Negative control on the exact parent commit (f539c1fc65e3a0de79b2a0d08805b096fdc85865): launching the same GLM-5.2-NVFP4 PP4 disaggregated-prefill configuration with EAGLE steps 3 on a GB300 failed before server readiness in check_server_args:

AssertionError: Pipeline parallelism is not compatible with overlap schedule, speculative decoding

This confirms that PP4 prefill with EAGLE did not start before this PR; the failure is the pre-change non-NPU PP + speculative-decoding restriction removed by this change.

GLM-5.2-NVFP4 on GB300 with Dynamo router:

  • Prefill: 1 worker, PP4, EAGLE steps 3
  • Decode: 1 worker, TP4, EAGLE steps 3
  • GSM8K: 1,319-example dataset, 20-shot (1,299 scored examples), --max-tokens 16384
  • Model revision: 53e0691e21895a3863a606dfd12910c69eba94ab
  • Hardware: 2 GB300 nodes in one NVL72 segment
configuration GSM8K score decode acceptance length
PP4 prefill + TP4 decode, steps 3/3 94.69% (1230/1299) 3.745 mean (3.00-4.00)

This is within 0.76 percentage points of the original PR's 95.45% steps1/steps3 result. The decode acceptance length is the mean of 143 scheduler samples and is higher than the original PR's reported 3.589.

Tests

  • BLACK_NUM_WORKERS=1 SKIP=no-commit-to-branch pre-commit run --all-files --show-diff-on-failure
  • Targeted PP/MLA, PP/EAGLE, embedding-ownership, and PP DSA round-trip unit tests on the GB300 validation container: 28 passed, 264 deselected, 32 subtests passed.

Checklist

  • Format code with pre-commit.
  • Add unit tests.
  • Update documentation (not applicable; no user-facing interface change).
  • Provide accuracy results.
  • Follow the SGLang code-style guidance.

CI States

Latest PR Test (Base): ✅ Run #35381952475
Latest PR Test (Extra): ❌ Run #35381952439
Latest PR Test (AMD ROCm 10): ❌ Run #35381952427

@YAMY1234 YAMY1234 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Very clean and clear implementation.

@YAMY1234

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 14, 2026
@YAMY1234

Copy link
Copy Markdown
Collaborator

FYI tracking issue for PP x speculative decoding, including the suggestion to fold the Qwen3_5 architectures from #39602 into this allowlist: #39634

# Conflicts:
#	test/registered/unit/server_args/test_server_args.py
nvpohanh and others added 4 commits September 16, 2026 14:32
Fold the Qwen3.5 allowlist from sgl-project#39602 into the shared
check_pipeline_parallel_compat gate so disaggregated-prefill PP + MTP
works for Qwen3.5 dense/MoE (text and multimodal) alongside DeepSeek/GLM.
Also covers the SGLANG_ENABLE_PP_SPEC aggregate branch that main added
in sgl-project#30775 and that the merge folded into the same function.
@YAMY1234

Copy link
Copy Markdown
Collaborator

@YAMY1234
YAMY1234 merged commit 3a64faa into sgl-project:main Sep 19, 2026
303 of 343 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deepseek run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants