Repository navigation
[AMD] Add opt-in MiniMax-M3 TP4 indexer context partitioning - #41488
Merged
hnyls2002 merged 7 commits intoOct 2, 2026
Merged
Conversation
ThomasNing
requested review from
BBuf,
DarkSharpness,
Fridge003,
HaiShaw,
HydraQYH,
Qiaolin-Yu,
celve,
hebiao064,
ispobock,
merrymercy and
yuan-luo
as code owners
September 27, 2026 22:36
This was referenced Sep 28, 2026
kevin-mii
pushed a commit
to kevin-mii/sglang
that referenced
this pull request
Oct 1, 2026
…onto M3-opt-0929 Applied the PR diff (head 44632b7) with a 3-way merge; environ.py kept both additions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MiniMaxM3VLConfig keeps the text config's fields on a sub-config, so hf_config.num_key_value_heads raises AttributeError and make_indexer_cp dies before the gate can report a reason. ModelConfig.get_total_num_kv_heads() returns 4 for amd/MiniMax-M3-MXFP4, which is what the gate checks for. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The gate rejected every speculative config, so CP stayed off for the whole EAGLE3 deployment -- which is how MiniMax-M3 is served. It does not have to. main's forward_extend funnels HIP target-verify into forward_decode with per-row seq_lens (prefix + 1..ndt) and the request slot repeated, because a linear EAGLE chain exposes one more KV token to each successive query. Verify rows therefore reach the indexer as ordinary decode queries, each carrying its own slot and causal length, which is exactly what the context-partitioned scorer already handles -- it was never one row per request. Gate on that property and allowlist it, so an algorithm whose verify rows are not independent chain rows (DSPARK's ragged lengths, NGRAM's tree-in-mask, EAGLE with top-k > 1) disables CP rather than reading as a chain and silently mis-scoring. Measured on MI350X with verify rows flowing through CP (TP4, EAGLE3 real acceptance, AgentX agentic traces, 900 s/point, 1M context, index top-k freq 1), total tok/s/GPU: c=24 25,441 -> 28,182 (+10.8%), c=32 32,393 -> 33,952 (+4.8%), c=40 31,571 -> 33,229 (+5.3%); ITL p50 -12 to -28%. GSM8K-500 0.862-0.872. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ming results-gfx950.json is a 629-line snapshot of one machine on one day -- it pins a torch build and a harness hash, nothing reads it, and the README already carries the same numbers as a table. No other committed results file exists under test/ (all five JSON files there are test inputs), so it goes; the harness regenerates the samples. benchmark_cp.py -> bench_cp.py: the repo has 78 bench_*.py under test/ and no benchmark_*.py. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The only thing verifying that CP selects the same blocks as the native TP selector lived in test/manual, which CI never runs -- so a bit-exactness claim (64-bit packed keys, score-descending with ID-ascending ties, the forced init/local ladder) had nothing guarding it. score_local_blocks, select_local_candidates and merge_candidates all take rank as a plain argument, so all four shards run in one process on one GPU; the all-gather a four-rank run adds is the runtime's collective, not this kernel. That makes it a 1-GPU AMD test alongside test_minimax_rocm_verify.py rather than a four-GPU one. Covers bf16 and fp8 at batch 1 and 4; injecting a one-rank shift into the block stride fails it with 15 IDs differing. bench_cp.py stays as the timing harness and keeps its four-rank parity check, which also exercises the collectives. Its docstring carries the run command and the "indexer microbenchmark, not model throughput" caveat, so the README went: the algorithm and the opt-in flag are already documented in the indexer_cp module docstrings and the environ entry, and the dated results table belongs in the PR description, not the tree. test/manual has no other README. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
It matches no pattern in the tree. Of 78 bench_*.py, 73 are registered under test/registered/kernels/benchmark/<group>/ with the shared run_benchmark harness and a reduced ci_range; the 5 under test/manual are short single-process microbenchmarks. This was a 288-line torchrun harness with argparse, JSON output and source hashing, in a directory CI never runs. Its correctness half is now test/registered/amd/test_minimax_indexer_cp.py, which gets the same guarantee in CI. Its timing half cannot honestly move to the registered 1-GPU pattern: a single-process run omits the two all-gathers, and CP must be judged with communication included, so such a benchmark would overstate the win. The end-to-end serving A/B is the measurement that counts and it belongs in the PR description. This leaves the gathers in MiniMaxIndexerCP.__call__ and graph capture of the full chain without a dedicated harness; they are exercised by any four-rank serving run, which is how the numbers in this PR were produced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
kevin-mii
approved these changes
Oct 2, 2026
Collaborator
|
/tag-and-rerun-ci |
hnyls2002
approved these changes
Oct 2, 2026
This was referenced Oct 4, 2026
4 of 5 tasks
4 tasks done
Collaborator
|
Hi @ThomasNing , thanks for your contribution! The test you added And it's still failing as of today https://github.com/sgl-project/sglang/actions/runs/37709475489/job/113091796233#step:6:7682. Could you please take a look? |
This was referenced Oct 8, 2026
2 of 4 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Under TP4, MiniMax-M3's four index-query heads are split across ranks while the index-K cache is replicated. Each rank scans the full index cache for its head. This PR adds an opt-in decode path that partitions the block reads across those four ranks and scores all heads while each K tile is resident.
The measured benefit is in the indexer chain, including communication: at batch 32 / 128K context, latency falls from 109.84 to 41.63 µs (62.1% lower latency, 2.64× speedup). Short contexts and small batches regress, so the feature defaults off. Full-model serving gains have not been measured.
Modifications
SGLANG_MINIMAX_M3_INDEXER_CP=1, default0, for gfx950 TP4 ordinary decode with attention CP/DP size 1, four index heads, four main KV heads, dimension 128, 128-token blocks, max scoring, and top-k 16. Unsupported run configurations log a reason and retain the existing path.rread blocksr, r+4, r+8, ...for all four heads. The cache remains replicated; this reduces read traffic, not cache capacity.test/manual/minimax_m3/indexer_cp/.The initial scope excludes speculation, TBO, HiSparse, FP8 queries, and dense sparse decode. There is no automatic batch/context crossover policy. It does not require a main-attention backend change.
Accuracy Tests
Completed on four MI355X / gfx950 GPUs:
Validation still outstanding: runtime feature-gate/backend dispatch through a loaded server, full-model accuracy, serving TPS/TTFT/TPOT, and measurements with the fallback collective. The harness directly invokes the indexer helper. The shape gate also accepts FP16 queries and other FP8 cache variants, which were not covered by this run.
Speed Tests and Profiling
Four MI355X GPUs, BF16 queries, FP8 E4M3 index cache, AITER custom gather available. Both gathers are included; projections and main attention are excluded. Eight indexer calls per HIP graph, 100 replays per round, seven rounds alternating baseline/CP order. Each sample uses the slowest rank; reported values are the median of seven samples.
Inputs/cache are repeatedly reused and warm, without an L2 flush or rotating-layer protocol. These are microbenchmark results and cannot be interpreted as model throughput gains.
Negative latency change is an improvement. All tested 8K shapes and batches 1–2 at 128K regress. The small 32K/batch-16 gain needs independent repetition. The README contains the complete 17-shape table, runtime versions, source hashes, and the reproduction command;
results-gfx950.jsoncontains every round.Checklist
CI States
Latest PR Test (Base): ✅ Run #36951043442
Latest PR Test (Extra): 🚫 Run #36962666288
Latest PR Test (AMD ROCm 10): ❌ Run #36951042978