Repository navigation
[AMD] MiniMax-M3 indexer CP: packed scoring for EAGLE verify rows - #42614
Merged
hnyls2002 merged 11 commits intoOct 9, 2026
Merged
Conversation
Chain verify presents ndt rows per request, but the CP scorer's 16-row tile holds only WORLD=4 live rows -- one head per lane for a single query -- so it runs 25% occupied and reads every K block once per draft row. Packing a request's draft rows (tile row i = draft row i // WORLD, head i % WORLD) fills the tile and reads each K block once per request. The packed kernel stays separate from _score_shard: folding the plain path in as PACK=1 reproduces it exactly but its vector row addressing measured 3.6% slower on ordinary decode. Rows in a group may straddle a block boundary, so the K-read branch uses the group's longest row and each row's own length restores the per-row masking and forced-block rule. A pack that does not fit the tile falls back to the unpacked kernel. Unused until a caller passes packed_queries > 1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
forward_extend funnels HIP target-verify into forward_decode with req_pool_indices.repeat_interleave(ndt), so a verify batch reaches the indexer as each request's draft rows consecutively -- the layout the packed scorer needs. Pass the draft-token count there and 1 everywhere else; only the CP branch reads it, and CP already serves score-only layers alone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kevin-mii
requested review from
BBuf,
DarkSharpness,
Fridge003,
HaiShaw,
HydraQYH,
Qiaolin-Yu,
celve,
hebiao064,
ispobock,
merrymercy and
yuan-luo
as code owners
October 5, 2026 17:21
_inputs gave request i table row i (slots = arange(batch)), so the slot id equalled the batch row and the packed group index. A scorer that addressed K by row or group index and ignored the slot table still passed both tests. Give each request a slot that is never its own index, out of order, with unused table rows in between. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The PACK=1 rationale and its 3.6% decode cost are a design argument; they belong in the PR body, not above the kernel. The slot-layout comment in the test kept only the fact that makes the slot formula look intentional. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
NDT does not say what it counts without reading its definition. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
It renamed packed_queries for three uses and added nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ines Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator
Author
|
/tag-and-rerun-ci |
4 of 5 tasks
It was registered on stage-b-test-1-gpu-small-amd (MI300, gfx942) with a ROCm-only skip. There the native reference kernel needs 128 KiB of shared memory against a 64 KiB limit and raises OutOfResources. The CP kernels are gfx950-only, so register on stage-b-test-1-gpu-small-amd-mi35x and skip unless gfx95, as the other gfx950 kernel tests do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ores those scores A row with at most top-k blocks keeps all of them in _local_candidates whatever their scores, so zeroing them in the scorer never changed a selection (removing it leaves the parity tests passing, kernel time within +-1%). Also make the score_local_blocks docstring plain text and say which kernel the dot order matches. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
_score_shard checks the slot range before indexing ReqToToken; the packed branch did not, so a padding group longer than top-k blocks would read past the table. The scores were masked out either way, so no selection changed, and an out-of-range read only faults depending on what the allocator mapped there; no registered test can make that fail deterministically, so none is added. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
hnyls2002
approved these changes
Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
With #41488 merged, MiniMax-M3 indexer context partitioning serves EAGLE chain verify: the ROCm verify
funnel hands the indexer each request's
ndtdraft rows as ordinary decode queries. The CP scorer thenscores them one row at a time. Its 16-row tile holds
WORLD=4 live rows per program (one head per lane),so it runs 25% occupied and reads every index-K block once per draft row -- four times per request at
the default 4 draft tokens.
Modifications
kernels/ops/attention/minimax_sparse/decode/indexer_cp.py).score_local_blockstakes
packed_queries; when a request's rows fit the tile it launches_score_shard_packed, which mapstile row
ito (draft rowi // WORLD, headi % WORLD). That fills the tile and reads each K blockonce per request. Rows in a group can straddle a block boundary, so the K-read branch uses the group's
longest row, while masking and the forced init/local blocks use each row's own length. A pack that does
not fit the tile falls back to the existing kernel. Like
_score_shard, it skips the K read for a slotpast the request table (graph padding).
minimax_sparse_decodeandMiniMaxIndexerCPpasspacked_queriesthrough;the backend passes the draft-token count for target-verify batches and 1 otherwise. Only the CP branch
reads it, and CP already serves score-only layers alone, so no other path changes.
The packed kernel stays separate from
_score_shard. Folding the plain path in asPACK=1reproduces itexactly, but the vector row addressing measured 3.6% slower on ordinary decode (30/30 shapes).
Accuracy Tests
Selection is exact, not approximate: the packed kernel must pick the same block IDs the native selector
picks for each row.
test_packed_verify_rows_match_the_native_selector(registered AMD, 1 GPU) builds verify rows the waythe funnel lays them out -- 4 requests x 4 draft rows -- ending at 32767..32770, so one tile mixes
256- and 257-block rows with different local windows. bf16 and fp8: exact ID equality on every head.
Each of these injected bugs fails it: the group's block count in place of each row's, sizing the group
by its shortest row (fails only because the rows straddle a block boundary), dropping the forced local
window, a wrong score-row index, and dropping the fp8-to-bf16 cast (fp8 subtest only).
_inputsnow gives each request a permuted, gapped slot that never equals its row or groupindex. A packed kernel that addresses K by group index instead of through the slot table now fails (2
subtests); before this change it passed.
test_minimax_indexer_cp.pyon gfx950 (MI355X): 2 passed,6 subtests.
than MI300 has, and the CP path is gfx950-only.
0.864 / 0.866 packed vs 0.884 / 0.886 unpacked; this raw few-shot gate moves about +-3.5 points between
server instances of the same build, so read it as no change.
Speed Tests and Profiling
MiniMax-M3 MXFP4 (
amd/MiniMax-M3-MXFP4), TP4 on MI350X, EAGLE3 (Inferact/MiniMax-M3-EAGLE3-GQA,3 steps / 4 draft tokens, real acceptance),
SGLANG_MINIMAX_M3_INDEXER_CP=1,SGLANG_MINIMAX_M3_INDEX_TOPK_FREQ=1. AgentX agentic traces (SemiAnalysis AIPerf, seed 42), 900 s perpoint, both arms side by side on the node's two GPU halves with identical flags. The unpacked arm forces
pack = 1-- the scorer main runs today. Total tok/s per GPU:Measured on a development branch whose CP kernel is identical to this PR and whose verify wiring is
equivalent (packing applies only to CP's score-only layers).
On main. A-B-B-A on one TP4 replica (MI350X, hostcall-capable node,
sgl_kernelbuilt from main),main
44a2558d40vs main + this PR, same flags as above, c=24, 900 s per arm:+4.6% throughput, -8.6% ITL p50; repeated arms agree within 0.4%. A 20-step decode profile of the PR arm
shows only
_score_shard_packed(57 launches per rank per step, one per sparse layer) and no_score_shard.Both arms ran with
--cuda-graph-backend-prefill disabled, because main crashes on its first M3 prefillwith breakable prefill graphs (fixed by #41845); the decode/verify path this PR changes is unaffected.
The gain on main is smaller than on the development branch because main's other per-step costs are larger.
Checklist
🤖 Generated with Claude Code
CI States
Latest PR Test (Base): ❌ Run #37843210552
Latest PR Test (Extra): ❌ Run #37843209355
Latest PR Test (AMD ROCm 10): ❌ Run #37843210373