Skip to content

[Fix] Step-3.5: build the shared expert without TP under all-to-all MoE backends - #41754

Merged
ch-wan merged 1 commit into
mainfrom
cheng/refactor/ffn-output-parts
Sep 30, 2026
Merged

ch-wan merged 1 commit into
mainfrom
cheng/refactor/ffn-output-parts

Conversation

@ch-wan

@ch-wan ch-wan commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

This PR is part of a stack (oldest at bottom):

Motivation

Step-3.5 adds the outputs of its routed experts and its shared expert, then runs one TP all-reduce on the sum. That is correct when both parts are TP partial sums. Under DeepEP, fuseep and the other all-to-all MoE backends, the combine already returns each token's routed output in full, on the rank that owns the token, but the shared expert stayed TP-sharded:

  • The all-reduce summed the complete routed output again, so it was counted TP times.
  • The shared expert's partial sums were computed on the tokens each rank holds, so the all-reduce added outputs of different tokens.

DeepSeek-V2 and GLM-4-MoE avoid both by building the shared expert with TP size 1 under these backends.

Modifications

  • Build share_expert with tp_size=1 under the backends GLM-4-MoE uses for this: DeepEP, Mooncake, NIXL, MORI, Ascend fuseep, the FlashInfer and FlashInfer MegaMoE all-to-all, and the FlashInfer CUTLASS FP4 all-gather.
  • Step3p5MoEMLP takes the shared expert's output and adds it in place, without allocating a third tensor: into the routed output before reduce_moe_output() on the TP path; on the DeepEP path, where nothing is left to sum, the combined routed output is added into the shared expert's output.
  • The decoder no longer forces the MoE block to skip its reduction and then all-reduces by hand. The FFN exit's flags reach reduce_moe_output() as for the other MoE models.

Accuracy Tests

H200, stepfun-ai/Step-3.5-Flash cut to its first five layers (three dense, two MoE):

  • --tp-size 2, python -m sglang.benchmark.one_batch --correctness-test: identical to the parent commit, and to a second run of the parent.
  • --tp-size 4, greedy sglang.Engine runs with per-step logits: all 99 steps bitwise identical to the parent.
  • --tp-size 4 --ep-size 4 --moe-a2a-backend deepep --moe-runner-backend deep_gemm, against the TP4 run as the reference. The DeepEP forward does not apply moe_router_scaling_factor (3.0 here), which this PR does not change, so the comparison also applies it with a test hook in both trees:
    • This PR: max |Δlogit| 0.17 (logits up to 9.2), all 99 argmax steps agree.
    • Parent: max |Δlogit| 8.09, and the first generated token differs for all three prompts.
    • This PR without the hook: max |Δlogit| 1.36, all 99 argmax steps agree; the remaining difference is the missing scaling factor.
  • 109 affected unit test files (2,697 tests) pass at the top of this stack, the same as on main.

Speed Tests and Profiling

Without an all-to-all backend the same kernels and collectives run. Under the all-to-all backends each rank now runs the whole shared expert on its own tokens instead of a TP shard of it on the same tokens, as in DeepSeek-V2 and GLM-4-MoE; no latency run was made.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): 🚫 Run #36667714896
Latest PR Test (Extra): 🚫 Run #36667714664
Latest PR Test (AMD ROCm 10): ❌ Run #36667714975

@ch-wan

ch-wan commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 29, 2026
@ch-wan
ch-wan force-pushed the cheng/refactor/glm5-next-mhc-pp branch from 1540944 to 91e9577 Compare September 29, 2026 20:29
@ch-wan
ch-wan force-pushed the cheng/refactor/ffn-output-parts branch from 452fd60 to 0b2b00b Compare September 29, 2026 20:29
@ch-wan ch-wan removed the run-ci CI: run the baseline test suite on this PR label Sep 29, 2026
@ch-wan
ch-wan force-pushed the cheng/refactor/glm5-next-mhc-pp branch from 91e9577 to 88e2d70 Compare September 29, 2026 23:48
@ch-wan
ch-wan force-pushed the cheng/refactor/ffn-output-parts branch from 0b2b00b to ace594b Compare September 29, 2026 23:48
@ch-wan
ch-wan force-pushed the cheng/refactor/glm5-next-mhc-pp branch from 88e2d70 to 4e0b589 Compare September 29, 2026 23:56
@ch-wan
ch-wan force-pushed the cheng/refactor/ffn-output-parts branch from ace594b to 78705b2 Compare September 29, 2026 23:56
@ch-wan ch-wan changed the title [Refactor] Complete FFN outputs computed in parts through the exit [Fix] Step-3.5: build the shared expert without TP under all-to-all MoE backends Sep 29, 2026
@ch-wan
ch-wan force-pushed the cheng/refactor/ffn-output-parts branch from 78705b2 to ec1cb47 Compare September 30, 2026 00:17
@ch-wan
ch-wan force-pushed the cheng/refactor/glm5-next-mhc-pp branch from 4e0b589 to e5270f4 Compare September 30, 2026 00:33
@ch-wan
ch-wan force-pushed the cheng/refactor/ffn-output-parts branch from ec1cb47 to 92c25f0 Compare September 30, 2026 00:33
@ch-wan
ch-wan force-pushed the cheng/refactor/glm5-next-mhc-pp branch from e5270f4 to 5e2f6c7 Compare September 30, 2026 01:01
@ch-wan
ch-wan force-pushed the cheng/refactor/ffn-output-parts branch from 92c25f0 to 515f650 Compare September 30, 2026 01:01
Base automatically changed from cheng/refactor/glm5-next-mhc-pp to main September 30, 2026 04:11
@mintlify

mintlify Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated
sglang-doc 🟢 Ready View Preview Sep 30, 2026, 4:14 AM

💡 Tip: Enable Automations to automatically generate PRs for you.

…ends

Step-3.5 adds its routed and shared experts before one TP all-reduce. That
holds when both parts are TP partial sums. Under DeepEP, fuseep and the other
all-to-all backends the combine already returns each token's routed output in
full, on the rank that owns the token, while the shared expert stayed
TP-sharded: the all-reduce then summed the routed output again, and the
shared expert's partial sums were added across ranks that hold different
tokens.

Build the shared expert with TP size 1 under those backends, as DeepSeek-V2
and GLM-4-MoE do, and add it inside the MoE block: before reduce_moe_output()
on the TP path, and to the combined output on the DeepEP path. The decoder no
longer forces the MoE to skip its reduction or all-reduces by hand. Without
an all-to-all backend the same single all-reduce runs on the same sum.
@ch-wan
ch-wan force-pushed the cheng/refactor/ffn-output-parts branch from 515f650 to e543340 Compare September 30, 2026 04:11
@ch-wan
ch-wan merged commit 9ee065e into main Sep 30, 2026
12 of 19 checks passed
@ch-wan
ch-wan deleted the cheng/refactor/ffn-output-parts branch September 30, 2026 04:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) parallel-stages CI: stages dispatch together instead of waiting on each other run-ci-extra CI: also run the extra suite (requires run-ci)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant