Skip to content

[Fix] MiniMax-M3: default to the breakable prefill CUDA graph - #41845

Merged
hnyls2002 merged 14 commits into
sgl-project:mainfrom
thanhhao98:htphan/m3-tc-piecewise-traceable
Oct 9, 2026
Merged

hnyls2002 merged 14 commits into
sgl-project:mainfrom
thanhhao98:htphan/m3-tc-piecewise-traceable

Conversation

@thanhhao98

@thanhhao98 thanhhao98 commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

MiniMax-M3 defaults to the tc_piecewise prefill CUDA graph, and on main that launch dies at boot (reproduced on nightly-dev-cu13-20261005-f70e8c68, GB300 TP4):

Capture prefill CUDA graph failed: Unsupported object.__new__ user-defined class construction
  ... object.__new__(DeclaredSum) is not safe

tc_piecewise is being removed (#41634), so this PR moves M3 to breakable. That needed one fix: under breakable, RadixAttention ran M3's sparse attention inline instead of as an eager break, so the graph replayed the capture batch's metadata and the server died on its first requests with CUDA error: an illegal memory access was encountered.

Modifications

Accuracy Tests

End to end on GB300, MiniMax-M3-NVFP4: this PR at its default (breakable) vs main with the prefill graph disabled, each pair on one node.

config GSM8K 1319 q, this PR GSM8K 1319 q, main without prefill graph
TP4 0.870 0.864
TP4, EP4, DP attention 4 0.864 0.877

Across our other runs (two nightlies, with and without DP attention), both settings stay within 0.864-0.877. Image requests replay the graph too: a colored-shape probe (nine images, sent one at a time and all at once) answers 18/18 in every arm. Every arm also passes a bs=1 GSM8K canary and an EOS-termination check.

Speed Tests

Same pairs, ISL 8192 / OSL 1024, median TTFT and output throughput:

config concurrency TTFT (ms), this PR / main tok/s, this PR / main
TP4 1 181 / 232 205 / 203
TP4 16 1502 / 1505 1441 / 1429
TP4 64 4868 / 4946 2838 / 2812
TP4, EP4, DP attention 4 1 505 / 438 112 / 112
TP4, EP4, DP attention 4 64 4313 / 4670 2854 / 2699

With DP attention the graph makes every DP rank replay a shared MAX_LEN bucket, so at concurrency 1 the idle ranks run the busy rank's length. TPOT is unchanged except DP attention at concurrency 64 (19.13 to 18.15 ms).

Reproduce

Container lmsysorg/sglang:nightly-dev-cu13-20261005-f70e8c68 with python/sglang at this PR's head (baseline: main efb62ce269), checkpoint nvidia/MiniMax-M3-NVFP4, one 4-GPU GB300 node.

# This PR (breakable is the default). Baseline: the same command on main plus --cuda-graph-backend-prefill disabled.
# DP-attention rows add: --ep-size 4 --enable-dp-attention --dp-size 4
SGLANG_DISABLE_MSA=1 python3 -m sglang.launch_server --model-path nvidia/MiniMax-M3-NVFP4 --trust-remote-code \
  --tp 4 --context-length 1048576 --reasoning-parser auto --tool-call-parser auto --port 30000

# GSM8K: 5-shot, temperature 0, 128 in parallel. The bs=1 canary runs --num-questions 40 with --parallel 1 and with --parallel 40.
python3 -m sglang.test.few_shot_gsm8k --num-shots 5 --num-questions 1319 --port 30000

# Speed: --num-prompts 12 / 64 / 256 for --max-concurrency 1 / 16 / 64.
python3 -m sglang.bench_serving --backend sglang --base-url http://127.0.0.1:30000 \
  --dataset-name random --random-input-len 8192 --random-output-len 1024 --random-range-ratio 1.0 \
  --flush-cache --max-concurrency 1 --num-prompts 12

Checklist

🤖 Generated with Claude Code


CI States

Latest PR Test (Base): ✅ Run #37870659321
Latest PR Test (Extra): ❌ Run #37870659081
Latest PR Test (AMD ROCm 10): ❌ Run #37870659400

Hao Phan added 3 commits September 30, 2026 11:34
In-package calls never mark a getter as warned, so every call from a
torch.compile'd forward reaches sys._getframe, which Dynamo cannot trace.
UnreducedOutput, DeclaredSum, Contribution, OwedOutput and ExitDecision are
built (and Contribution mutated) on every forward inside the decoder stack.
Dynamo cannot construct a msgspec.Struct, so a fullgraph torch.compile of
that stack, the tc_piecewise prefill CUDA graph, failed. They become
__slots__ classes with the same fields, defaults and constructor signatures,
snapshot() rebuilds the unreduced output with its constructor, and the
no-dataclasses rule records the exception.
MoeDeferredFinalize is built inside the MoE forward when the next layer's
input takes over the finalize; like the other per-forward boundary values,
DeferredFinalize and it become __slots__ classes Dynamo can construct.
@nvpohanh

nvpohanh commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

MiniMax-M3 defaults to the tc_piecewise prefill CUDA graph

@thanhhao98 Could we try to set default to breakable? I think tc_piecewise is going to be deprecated/removed

@thanhhao98

Copy link
Copy Markdown
Contributor Author

@nvpohanh should I fix for tc_piecewise and change to breakable; or just need to change default to breakable?

Hao Phan added 2 commits October 5, 2026 17:06
The breakable prefill graph called the sparse backend inline, so MiniMax-M3's
sparse prefill attention was captured with the capture batch's metadata and
replayed against real batches. Wrap the sparse op with eager_on_graph like the
dense ops so it reruns against each replay batch.
@thanhhao98 thanhhao98 changed the title [Fix] MiniMax-M3 boots again with its default tc_piecewise prefill graph [Fix] MiniMax-M3: default to the breakable prefill graph and fix both prefill backends Oct 5, 2026
@nvpohanh

nvpohanh commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

@thanhhao98 tc_piecewose will be removed soon, so let's just change it to breakable if it works.
@ch-wan and @Oasis-Git will remove tc_piecewise

Resolve exit.py: main reshaped ExitDecision to (defer_moe_finalize,
sum_in_reduce_scatter, complete); keep it a plain __slots__ class with
those fields.
@thanhhao98 thanhhao98 changed the title [Fix] MiniMax-M3: default to the breakable prefill graph and fix both prefill backends [Fix] MiniMax-M3: default to the breakable prefill CUDA graph Oct 5, 2026
@thanhhao98

Copy link
Copy Markdown
Contributor Author

@nvpohanh I updated this pr to move to breakable only.

@nvpohanh nvpohanh left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[by Claude Code] Severity: style | Confidence: High

Issue

The Accuracy Tests and Speed Tests tables give the model, GPU, and parallel configs, but no commands. Nobody can rerun the GSM8K scores or the TTFT and throughput rows without guessing the server flags, the eval settings, and the benchmark settings.

Fix

Add the exact commands for one config, and note what changes for the other:

  • the sglang.launch_server command for this PR and for the main baseline, including the flag that turns off the prefill graph (for example --cuda-graph-backend-prefill disabled);
  • the GSM8K eval command (script, number of shots, parallelism);
  • the sglang.bench_serving command (dataset, ISL/OSL, number of prompts, concurrency).

Also give the checkpoint path for MiniMax-M3-NVFP4.

@nvpohanh

nvpohanh commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

/rerun-test test/registered/unit/layers/test_radix_attention.py test/registered/cuda_graph/breakable/test_breakable_cuda_graph.py test/registered/unit/server_args/test_server_args.py

@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

🚀 ubuntu-latest (2 tests): ✅ View workflow run

cd test/ && python3 registered/unit/layers/test_radix_attention.py
cd test/ && python3 registered/unit/server_args/test_server_args.py

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/cuda_graph/breakable/test_breakable_cuda_graph.py

@kevin-mii

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@nvpohanh

nvpohanh commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

@ispobock could you review this? Thanks!

@nvpohanh

nvpohanh commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

@nvpohanh

nvpohanh commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

3 similar comments
@nvpohanh

nvpohanh commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

@nvpohanh

nvpohanh commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

@nvpohanh

nvpohanh commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

@nvpohanh

nvpohanh commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

[by Claude Code] I've started babysitting the NVIDIA CI of this PR. I'll check it every 30 minutes until all NV pipelines pass. If a relevant job fails, I may rerun it, merge the latest main into this branch, push a small test-only fix, or post a note here.

Anyone can stop me at any time: just post a comment asking me to stop (for example, stop babysitting). I'll stop at my next check, within 30 minutes, and take no further action on this PR.

@nvpohanh

nvpohanh commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

@nvpohanh

nvpohanh commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

[by Claude Code] All NV pipelines have passed.

@hnyls2002
hnyls2002 merged commit fbf8d73 into sgl-project:main Oct 9, 2026
289 of 346 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants