Skip to content

[AMD] MiniMax-M3 indexer CP: packed scoring for EAGLE verify rows - #42614

Merged
hnyls2002 merged 11 commits into
sgl-project:mainfrom
kevin-mii:feat/m3-indexer-cp-packed-verify
Oct 9, 2026
Merged

hnyls2002 merged 11 commits into
sgl-project:mainfrom
kevin-mii:feat/m3-indexer-cp-packed-verify

Conversation

@kevin-mii

@kevin-mii kevin-mii commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

With #41488 merged, MiniMax-M3 indexer context partitioning serves EAGLE chain verify: the ROCm verify
funnel hands the indexer each request's ndt draft rows as ordinary decode queries. The CP scorer then
scores them one row at a time. Its 16-row tile holds WORLD=4 live rows per program (one head per lane),
so it runs 25% occupied and reads every index-K block once per draft row -- four times per request at
the default 4 draft tokens.

Modifications

  1. Packed scorer (kernels/ops/attention/minimax_sparse/decode/indexer_cp.py). score_local_blocks
    takes packed_queries; when a request's rows fit the tile it launches _score_shard_packed, which maps
    tile row i to (draft row i // WORLD, head i % WORLD). That fills the tile and reads each K block
    once per request. Rows in a group can straddle a block boundary, so the K-read branch uses the group's
    longest row, while masking and the forced init/local blocks use each row's own length. A pack that does
    not fit the tile falls back to the existing kernel. Like _score_shard, it skips the K read for a slot
    past the request table (graph padding).
  2. Wiring (10 lines). minimax_sparse_decode and MiniMaxIndexerCP pass packed_queries through;
    the backend passes the draft-token count for target-verify batches and 1 otherwise. Only the CP branch
    reads it, and CP already serves score-only layers alone, so no other path changes.

The packed kernel stays separate from _score_shard. Folding the plain path in as PACK=1 reproduces it
exactly, but the vector row addressing measured 3.6% slower on ordinary decode (30/30 shapes).

Accuracy Tests

Selection is exact, not approximate: the packed kernel must pick the same block IDs the native selector
picks for each row.

  • test_packed_verify_rows_match_the_native_selector (registered AMD, 1 GPU) builds verify rows the way
    the funnel lays them out -- 4 requests x 4 draft rows -- ending at 32767..32770, so one tile mixes
    256- and 257-block rows with different local windows. bf16 and fp8: exact ID equality on every head.
    Each of these injected bugs fails it: the group's block count in place of each row's, sizing the group
    by its shortest row (fails only because the rows straddle a block boundary), dropping the forced local
    window, a wrong score-row index, and dropping the fp8-to-bf16 cast (fp8 subtest only).
  • The shared _inputs now gives each request a permuted, gapped slot that never equals its row or group
    index. A packed kernel that addresses K by group index instead of through the slot table now fails (2
    subtests); before this change it passed. test_minimax_indexer_cp.py on gfx950 (MI355X): 2 passed,
    6 subtests.
  • The test moves to the MI35x suite and skips off gfx950: its reference kernel needs more shared memory
    than MI300 has, and the CP path is gfx950-only.
  • GSM8K-500, 5-shot: 0.896 packed vs 0.892 unpacked (same A/B as below). On the main A/B below:
    0.864 / 0.866 packed vs 0.884 / 0.886 unpacked; this raw few-shot gate moves about +-3.5 points between
    server instances of the same build, so read it as no change.

Speed Tests and Profiling

MiniMax-M3 MXFP4 (amd/MiniMax-M3-MXFP4), TP4 on MI350X, EAGLE3 (Inferact/MiniMax-M3-EAGLE3-GQA,
3 steps / 4 draft tokens, real acceptance), SGLANG_MINIMAX_M3_INDEXER_CP=1,
SGLANG_MINIMAX_M3_INDEX_TOPK_FREQ=1. AgentX agentic traces (SemiAnalysis AIPerf, seed 42), 900 s per
point, both arms side by side on the node's two GPU halves with identical flags. The unpacked arm forces
pack = 1 -- the scorer main runs today. Total tok/s per GPU:

unpacked (main today) packed (this PR) change ITL p50
c=24 26,576 28,380 +6.8% 12.4 -> 10.2 ms
c=32 32,294 34,282 +6.2% 17.1 -> 12.7 ms

Measured on a development branch whose CP kernel is identical to this PR and whose verify wiring is
equivalent (packing applies only to CP's score-only layers).

On main. A-B-B-A on one TP4 replica (MI350X, hostcall-capable node, sgl_kernel built from main),
main 44a2558d40 vs main + this PR, same flags as above, c=24, 900 s per arm:

arm total tok/s/GPU ITL p50
main 16,912 / 16,983 30.8 / 31.0 ms
main + this PR 17,728 / 17,731 28.5 / 28.0 ms

+4.6% throughput, -8.6% ITL p50; repeated arms agree within 0.4%. A 20-step decode profile of the PR arm
shows only _score_shard_packed (57 launches per rank per step, one per sparse layer) and no _score_shard.
Both arms ran with --cuda-graph-backend-prefill disabled, because main crashes on its first M3 prefill
with breakable prefill graphs (fixed by #41845); the decode/verify path this PR changes is unaffected.
The gain on main is smaller than on the development branch because main's other per-step costs are larger.

Checklist

  • Format your code according to the Format code with pre-commit.
  • Add unit tests according to the Run and add unit tests.
  • Update documentation -- not applicable; no user-facing interface change.
  • Provide accuracy and speed benchmark results.
  • Follow the SGLang code style guidance.

🤖 Generated with Claude Code


CI States

Latest PR Test (Base): ❌ Run #37843210552
Latest PR Test (Extra): ❌ Run #37843209355
Latest PR Test (AMD ROCm 10): ❌ Run #37843210373

kevin-mii and others added 2 commits October 5, 2026 17:21
Chain verify presents ndt rows per request, but the CP scorer's 16-row tile
holds only WORLD=4 live rows -- one head per lane for a single query -- so it
runs 25% occupied and reads every K block once per draft row. Packing a
request's draft rows (tile row i = draft row i // WORLD, head i % WORLD) fills
the tile and reads each K block once per request.

The packed kernel stays separate from _score_shard: folding the plain path in
as PACK=1 reproduces it exactly but its vector row addressing measured 3.6%
slower on ordinary decode. Rows in a group may straddle a block boundary, so the
K-read branch uses the group's longest row and each row's own length restores
the per-row masking and forced-block rule. A pack that does not fit the tile
falls back to the unpacked kernel.

Unused until a caller passes packed_queries > 1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
forward_extend funnels HIP target-verify into forward_decode with
req_pool_indices.repeat_interleave(ndt), so a verify batch reaches the indexer
as each request's draft rows consecutively -- the layout the packed scorer
needs. Pass the draft-token count there and 1 everywhere else; only the CP
branch reads it, and CP already serves score-only layers alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kevin-mii and others added 5 commits October 7, 2026 22:32
_inputs gave request i table row i (slots = arange(batch)), so the slot
id equalled the batch row and the packed group index. A scorer that
addressed K by row or group index and ignored the slot table still
passed both tests. Give each request a slot that is never its own index,
out of order, with unused table rows in between.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The PACK=1 rationale and its 3.6% decode cost are a design argument; they
belong in the PR body, not above the kernel. The slot-layout comment in the
test kept only the fact that makes the slot formula look intentional.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
NDT does not say what it counts without reading its definition.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
It renamed packed_queries for three uses and added nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ines

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kevin-mii

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Oct 7, 2026
kevin-mii and others added 2 commits October 8, 2026 02:16
It was registered on stage-b-test-1-gpu-small-amd (MI300, gfx942) with a
ROCm-only skip. There the native reference kernel needs 128 KiB of shared
memory against a 64 KiB limit and raises OutOfResources. The CP kernels are
gfx950-only, so register on stage-b-test-1-gpu-small-amd-mi35x and skip
unless gfx95, as the other gfx950 kernel tests do.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kevin-mii and others added 2 commits October 8, 2026 20:22
…ores those scores

A row with at most top-k blocks keeps all of them in _local_candidates whatever
their scores, so zeroing them in the scorer never changed a selection (removing
it leaves the parity tests passing, kernel time within +-1%). Also make the
score_local_blocks docstring plain text and say which kernel the dot order matches.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
_score_shard checks the slot range before indexing ReqToToken; the packed branch
did not, so a padding group longer than top-k blocks would read past the table.
The scores were masked out either way, so no selection changed, and an
out-of-range read only faults depending on what the allocator mapped there; no
registered test can make that fail deterministically, so none is added.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@hnyls2002
hnyls2002 merged commit 6d8ddaf into sgl-project:main Oct 9, 2026
407 of 499 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

jit-kernel run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants