Skip to content

[Fix] Reduce the Sarvam dense MLP through its row-parallel projection - #41752

Merged
ch-wan merged 1 commit into
mainfrom
cheng/hot-fix/sarvam-dense-mlp-reduction
Sep 30, 2026
Merged

ch-wan merged 1 commit into
mainfrom
cheng/hot-fix/sarvam-dense-mlp-reduction

Conversation

@ch-wan

@ch-wan ch-wan commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

This PR is part of a stack (oldest at bottom):

Motivation

The dense MLP of SarvamMoEMLADecoderLayer is built with reduce_results=False, and the decoder then all-reduces its output over the TP group whenever attention TP is larger than one and the FFN exit published neither skip flag. That gate tests the wrong group:

  • Under attention DP with attention TP 1, the dense MLP still spans the full TP group. When the exit completes the output in this layer (for example with EAGLE, or a padding mode without reduce-scatter), the output is never summed.
  • With moe_dense_tp_size=1, the dense MLP runs on each rank's own rows and owes no sum, but with attention TP > 1 the decoder sums rows that belong to different tokens.

Modifications

  • Build the dense MLP's projection with the default reduce_results=True, as BailingMoE does. RowParallelLinear reduces over the MLP's own group and skips the reduction exactly when the exit publishes a flag.
  • Remove the decoder's own all-reduce.

Accuracy Tests

B200, sarvamai/sarvam-105b cut to its first four layers (layer 0 dense, layers 1–3 MoE) in bf16, python -m sglang.benchmark.one_batch --correctness-test. Four layers do not produce meaningful text, so logits and tokens are compared. For scale, the TP2 and TP4 references differ by at most 0.09 in the printed logits.

  • --tp-size 2 and --tp-size 4: identical to the parent commit.
  • Attention TP 1 with the sum completed in the layer (--tp-size 4 --dp-size 4 --enable-dp-attention --ep-size 2, eager prefill, triton attention): the parent never reduces the dense MLP output. Against the --tp-size 4 --ep-size 2 reference, the parent differs by up to 4.25 in the logits and generates different tokens; this PR differs by 0.03 and generates the same tokens.
  • --tp-size 4 --dp-size 2 --enable-dp-attention --moe-dense-tp-size 1 (batch size 2): the parent all-reduces the dense MLP output over the four-rank TP group, adding rows of different tokens. This PR runs no such reduction, and its output matches --dp-size 4 with the same dense setting (within 0.05). This configuration still differs from the TP4 reference for another reason that this PR does not change: under --moe-dense-tp-size 1 the MoE layers' shared experts are also built with TP size 1 and then summed over the TP group.
  • --tp-size 2 --dp-size 2 --enable-dp-attention and --tp-size 4 --dp-size 2 --enable-dp-attention: identical to the parent, and close to the reference. There the sum is left to a later step.

97 affected unit test files pass.

Speed Tests and Profiling

Not applicable: in the configurations that were already correct, the same all-reduce runs.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): 🚫 Run #36667620431
Latest PR Test (Extra): 🚫 Run #36667620347
Latest PR Test (AMD ROCm 10): ❌ Run #36667620499

@ch-wan
ch-wan force-pushed the cheng/hot-fix/sarvam-dense-mlp-reduction branch from 565b44a to fa7b7f2 Compare September 29, 2026 20:29
@ch-wan
ch-wan force-pushed the cheng/refactor/final-norm-skip-empty branch from 7916f5b to c1e1041 Compare September 29, 2026 23:47
@ch-wan
ch-wan force-pushed the cheng/hot-fix/sarvam-dense-mlp-reduction branch from fa7b7f2 to e1b4a77 Compare September 29, 2026 23:47
@ch-wan
ch-wan force-pushed the cheng/refactor/final-norm-skip-empty branch from c1e1041 to 8d3cb4e Compare September 29, 2026 23:56
@ch-wan
ch-wan force-pushed the cheng/hot-fix/sarvam-dense-mlp-reduction branch from e1b4a77 to 9c1ad64 Compare September 29, 2026 23:56
@ch-wan
ch-wan force-pushed the cheng/refactor/final-norm-skip-empty branch from 8d3cb4e to 40f0308 Compare September 30, 2026 00:33
@ch-wan
ch-wan force-pushed the cheng/hot-fix/sarvam-dense-mlp-reduction branch from 9c1ad64 to c83bf2c Compare September 30, 2026 00:33
@ch-wan
ch-wan force-pushed the cheng/refactor/final-norm-skip-empty branch from 40f0308 to ba89c53 Compare September 30, 2026 01:01
@ch-wan
ch-wan force-pushed the cheng/hot-fix/sarvam-dense-mlp-reduction branch from c83bf2c to 53e896c Compare September 30, 2026 01:01
The dense MLP of SarvamMoEMLADecoderLayer was built with
reduce_results=False, and the decoder then all-reduced its output over
the TP group whenever attention TP was larger than one and the FFN exit
had published neither skip flag. That gate tests the wrong group:

- under attention DP with attention TP 1 the dense MLP still spans the
  full TP group, so an output the exit completes in this layer (for
  example with EAGLE, or a padding mode without reduce-scatter) was never
  summed;
- with moe_dense_tp_size=1 the dense MLP runs on each rank's own rows and
  owes no sum, but attention TP > 1 summed rows of different tokens.

Build the projection with the default reduce_results=True, as BailingMoE
does: RowParallelLinear reduces over the MLP's own group and skips the
reduction exactly when the exit publishes a flag. The decoder no longer
runs a collective of its own. Both cases are read from the code; no
Sarvam checkpoint was available to reproduce them.
@ch-wan
ch-wan changed the base branch from cheng/refactor/final-norm-skip-empty to main September 30, 2026 04:09
@mintlify

mintlify Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated
sglang-doc 🟢 Ready View Preview Sep 30, 2026, 4:11 AM

💡 Tip: Enable Automations to automatically generate PRs for you.

@ch-wan
ch-wan force-pushed the cheng/hot-fix/sarvam-dense-mlp-reduction branch from 53e896c to cfdddf4 Compare September 30, 2026 04:10
@ch-wan
ch-wan merged commit 843f50f into main Sep 30, 2026
7 of 17 checks passed

This branch was successfully deployed

1 active deployment
staging - docs — cfdddf42 Deployed Sep 30, 2026 by mintlify[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant