Repository navigation
[Refactor] Build seven more decoders from stage boundaries - #41779
Merged
Merged
Conversation
ch-wan
requested review from
BBuf,
Edwardf0t1,
Fridge003,
HaiShaw,
Ying1123,
ispobock and
merrymercy
as code owners
September 29, 2026 23:56
This was referenced Sep 29, 2026
3 of 5 tasks
ch-wan
force-pushed
the
cheng/refactor/layer-boundary-cleanup
branch
from
September 30, 2026 00:17
13d2309 to
d82c3cf
Compare
ch-wan
requested review from
1am9trash,
YAMY1234,
fzyzcjy,
hubertlu-tw,
kkHuang-amd and
rainj-me
as code owners
September 30, 2026 00:17
ch-wan
force-pushed
the
cheng/refactor/plain-stack-boundaries
branch
from
September 30, 2026 00:17
7517a33 to
a62289b
Compare
4 of 5 tasks
ch-wan
force-pushed
the
cheng/refactor/layer-boundary-cleanup
branch
from
September 30, 2026 00:33
d82c3cf to
f0f45a0
Compare
ch-wan
requested review from
Jiminator,
hnyls2002 and
ping1jing2
as code owners
September 30, 2026 00:33
ch-wan
requested review from
JustinTong0323,
iforgetmyname,
sogalin,
whybeyoung,
wisclmy0611 and
zijiexia
as code owners
September 30, 2026 00:33
ch-wan
force-pushed
the
cheng/refactor/plain-stack-boundaries
branch
from
September 30, 2026 00:33
a62289b to
aae865a
Compare
ch-wan
force-pushed
the
cheng/refactor/layer-boundary-cleanup
branch
from
September 30, 2026 01:01
f0f45a0 to
b83b468
Compare
ch-wan
force-pushed
the
cheng/refactor/plain-stack-boundaries
branch
2 times, most recently
from
September 30, 2026 01:51
d62d9dc to
9a59370
Compare
ch-wan
force-pushed
the
cheng/refactor/layer-boundary-cleanup
branch
from
September 30, 2026 04:12
b83b468 to
8a90925
Compare
Base automatically changed from
cheng/refactor/layer-boundary-cleanup
to
main
September 30, 2026 04:16
Contributor
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Automations to automatically generate PRs for you. |
finish_complete_output() moves an FFN output that compute has already reduced, but it ran the move bound for the exit, which on two paths also completes the sum with a reduce-scatter: the residual on attention-TP slices and a CP take-back that sums. The output was then summed twice. Bind the non-summing counterpart of those moves at construction (the slice of the complete output, and the CP take-back) and use it there.
The decoder layer threaded its residual tensor by hand and the model sent both tensors to the next pipeline rank. Declare the attention and MLP stages instead: the attention projection leaves its TP sum to the MLP input, which runs the same all-reduce and fused add + norm as before, and the MLP still completes its own sum and finishes through finish_complete_output(). The model enters and leaves the stack through residual_batch, and EAGLE3 captures come from the boundary.
As for Apertus: the attention projection leaves its sum to the MLP input, which runs the same all-reduce and fused add + norm, the MLP completes its own sum and finishes through finish_complete_output(), and the model enters and leaves the stack through residual_batch. The attention was already built on the attention-TP group; the boundaries now also bring the MLP its rows under attention DP.
The attention projection leaves its TP sum to the MoE input, which runs the same all-reduce and fused add + norm; the MoE block still all-reduces its own output and finishes through finish_complete_output(). Every layer declares a sparse FFN. The model enters and leaves the stack through residual_batch.
As for Apertus: the attention projection leaves its sum to the MLP input, the MLP completes its own sum and finishes through finish_complete_output(), the model enters and leaves the stack through residual_batch, and EAGLE3 captures come from the boundary.
Both attention kinds (KDA and MLA) leave their TP sum to the FFN input, which runs the same all-reduce and fused add + norm; the dense MLP and the MoE block still complete their own sums and finish through finish_complete_output(). Each layer declares whether its FFN is sparse. DSpark captures read the residual stream through residual_batch.snapshot(), the same output plus residual, and still travel to the next pipeline rank next to the stream.
EXAONE 4.0 is post-LN: each sublayer reads the residual as it is and its output is normalized before it is added. Add the matching residual operations, a plain read and a post-norm add, and declare the stages with them. The attention leaves its TP sum to the MLP input, which completes it and then runs the post-norm add; the layer writes the MLP's post-norm add itself, so the stream it hands on is already written. The model now loops over its own layers only. It used to call every layer, including the placeholders of other pipeline ranks, which do not return a pair. The split-prefill path now ends with the same final norm as forward() instead of norming the output added to itself.
As for Apertus: the attention projection leaves its sum to the MLP input, the MLP completes its own sum and finishes through finish_complete_output(), and EAGLE3 / DFlash captures come from the boundary. Between loops the model folds the residual back in, applies the loop norm when the config asks for it, and starts a new stream for the next pass over the layers.
ch-wan
force-pushed
the
cheng/refactor/plain-stack-boundaries
branch
from
September 30, 2026 04:16
9a59370 to
491df15
Compare
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR is part of a stack (oldest at bottom):
Motivation
Apertus, Spark 2.5, Mixtral, Arcee, Kimi-Linear, EXAONE 4.0 and Nanbeige still thread their residual tensor through each decoder layer by hand, and send both tensors to the next pipeline rank themselves. Building them from stage boundaries gives them the same attention and FFN boundaries as the other decoders: the pipeline handoff, the attention-DP rows and the aux hidden-state captures.
Modifications
finish_complete_output()moves an FFN output that compute has already reduced. On two paths the move bound for the exit also completes the sum, with a reduce-scatter onto the attention-TP slices of the residual or a CP reduce-scatter, so such an output was summed twice. The non-summing counterpart of those moves is now bound at construction and used there.residual/post_norm.py: a plain read and a post-norm add (residual + norm(output)), for decoders that read the residual without a norm and normalize each sublayer's output before adding it. The plain read rejects a quantized input format and a post-residual addition, with or without a residual to update.reduce_results=False), which runs the same all-reduce and fused add + norm as before.finish_complete_output(). No sum is left to the next layer.residual_batch(start,from_pp,to_pp,final_norm). EAGLE3 and DFlash captures come from the boundary; Kimi-Linear's DSpark captures read the stream throughresidual_batch.snapshot()and still travel to the next pipeline rank.--pp-size 2fail at launch. Its split-prefill path now ends with the same final norm asforward()instead of norming the output added to itself.--enable-flashinfer-allreduce-fusionor--enable-quant-communications, the attention all-reduce of these models now takes the fused or quantized path. None of these architectures enables the fusion automatically, so the default configuration runs the same kernels as before.Accuracy Tests
Each model on a real checkpoint, the tree before this PR against this PR. TP2 is
python -m sglang.benchmark.one_batch --correctness-test(prefill logits and three generations), run twice on the parent. The server runs greedy-decode four prompts with input and output logprobs and the top-5 logprobs.one_batchswiss-ai/Apertus-8B-2509XHToken/Spark-X2.5-4Bmistralai/Mixtral-8x7B-Instruct-v0.1arcee-ai/AFM-4.5B-Basemoonshotai/Kimi-Linear-48B-A3B-InstructLGAI-EXAONE/EXAONE-4.0-1.2BLGAI-EXAONE/EXAONE-4.0-32BNanbeige/Nanbeige4.2-3B(two loops)main.Not addressed here: EXAONE 4.0 checkpoints with tied embeddings (the 1.2B) still cannot run with PP, because the last pipeline rank has no
lm_head.Not tested: Nanbeige with PP. As before this PR, each pipeline rank runs every loop over its own layers only, so looped checkpoints (
num_loops > 1) are not claimed to work with PP.Speed Tests and Profiling
Not applicable: the same kernels and collectives run in the default configuration.
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): 🚫 Run #36668115383
Latest PR Test (Extra): 🚫 Run #36668115221
Latest PR Test (AMD ROCm 10): ⏳ Run #36668115633