Repository navigation
[Fix] MiniMax-M3: default to the breakable prefill CUDA graph - #41845
Conversation
In-package calls never mark a getter as warned, so every call from a torch.compile'd forward reaches sys._getframe, which Dynamo cannot trace.
UnreducedOutput, DeclaredSum, Contribution, OwedOutput and ExitDecision are built (and Contribution mutated) on every forward inside the decoder stack. Dynamo cannot construct a msgspec.Struct, so a fullgraph torch.compile of that stack, the tc_piecewise prefill CUDA graph, failed. They become __slots__ classes with the same fields, defaults and constructor signatures, snapshot() rebuilds the unreduced output with its constructor, and the no-dataclasses rule records the exception.
MoeDeferredFinalize is built inside the MoE forward when the next layer's input takes over the finalize; like the other per-forward boundary values, DeferredFinalize and it become __slots__ classes Dynamo can construct.
@thanhhao98 Could we try to set default to |
|
@nvpohanh should I fix for tc_piecewise and change to breakable; or just need to change default to breakable? |
The breakable prefill graph called the sparse backend inline, so MiniMax-M3's sparse prefill attention was captured with the capture batch's metadata and replayed against real batches. Wrap the sparse op with eager_on_graph like the dense ops so it reruns against each replay batch.
|
@thanhhao98 tc_piecewose will be removed soon, so let's just change it to breakable if it works. |
Resolve exit.py: main reshaped ExitDecision to (defer_moe_finalize, sum_in_reduce_scatter, complete); keep it a plain __slots__ class with those fields.
|
@nvpohanh I updated this pr to move to breakable only. |
nvpohanh
left a comment
There was a problem hiding this comment.
[by Claude Code] Severity: style | Confidence: High
Issue
The Accuracy Tests and Speed Tests tables give the model, GPU, and parallel configs, but no commands. Nobody can rerun the GSM8K scores or the TTFT and throughput rows without guessing the server flags, the eval settings, and the benchmark settings.
Fix
Add the exact commands for one config, and note what changes for the other:
- the
sglang.launch_servercommand for this PR and for the main baseline, including the flag that turns off the prefill graph (for example--cuda-graph-backend-prefill disabled); - the GSM8K eval command (script, number of shots, parallelism);
- the
sglang.bench_servingcommand (dataset, ISL/OSL, number of prompts, concurrency).
Also give the checkpoint path for MiniMax-M3-NVFP4.
|
/rerun-test test/registered/unit/layers/test_radix_attention.py test/registered/cuda_graph/breakable/test_breakable_cuda_graph.py test/registered/unit/server_args/test_server_args.py |
|
🚀 🚀 |
|
/tag-and-rerun-ci |
|
@ispobock could you review this? Thanks! |
|
/rerun-failed-ci |
|
/rerun-failed-ci |
3 similar comments
|
/rerun-failed-ci |
|
/rerun-failed-ci |
|
/rerun-failed-ci |
Picks up sgl-project#43229, which fixes the unrelated HF Hub 429 CI failures (e.g. https://github.com/sgl-project/sglang/actions/runs/37738591702/job/113443229627) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
[by Claude Code] I've started babysitting the NVIDIA CI of this PR. I'll check it every 30 minutes until all NV pipelines pass. If a relevant job fails, I may rerun it, merge the latest Anyone can stop me at any time: just post a comment asking me to stop (for example, |
|
/rerun-failed-ci |
|
[by Claude Code] All NV pipelines have passed. |
Motivation
MiniMax-M3 defaults to the
tc_piecewiseprefill CUDA graph, and on main that launch dies at boot (reproduced onnightly-dev-cu13-20261005-f70e8c68, GB300 TP4):tc_piecewiseis being removed (#41634), so this PR moves M3 tobreakable. That needed one fix: underbreakable,RadixAttentionran M3's sparse attention inline instead of as an eager break, so the graph replayed the capture batch's metadata and the server died on its first requests withCUDA error: an illegal memory access was encountered.Modifications
breakable, like dense attention and like DSA's indexer and QSA ([QSA] Enable breakable prefill CUDA graphs for text-only Qwen3.8 Flash-Next (capture-safe metadata, MTP side-channel padding) #41729).breakable.Accuracy Tests
End to end on GB300, MiniMax-M3-NVFP4: this PR at its default (
breakable) vs main with the prefill graph disabled, each pair on one node.Across our other runs (two nightlies, with and without DP attention), both settings stay within 0.864-0.877. Image requests replay the graph too: a colored-shape probe (nine images, sent one at a time and all at once) answers 18/18 in every arm. Every arm also passes a bs=1 GSM8K canary and an EOS-termination check.
Speed Tests
Same pairs, ISL 8192 / OSL 1024, median TTFT and output throughput:
With DP attention the graph makes every DP rank replay a shared MAX_LEN bucket, so at concurrency 1 the idle ranks run the busy rank's length. TPOT is unchanged except DP attention at concurrency 64 (19.13 to 18.15 ms).
Reproduce
Container
lmsysorg/sglang:nightly-dev-cu13-20261005-f70e8c68withpython/sglangat this PR's head (baseline: mainefb62ce269), checkpointnvidia/MiniMax-M3-NVFP4, one 4-GPU GB300 node.Checklist
🤖 Generated with Claude Code
CI States
Latest PR Test (Base): ✅ Run #37870659321
Latest PR Test (Extra): ❌ Run #37870659081
Latest PR Test (AMD ROCm 10): ❌ Run #37870659400