Repository navigation
[AMD] Keep the MiniMax sparse decode score tile within gfx942 LDS - #41354
Open
siliangchen-amd wants to merge 1 commit into
Open
siliangchen-amd wants to merge 1 commit into
siliangchen-amd wants to merge 1 commit into
Conversation
siliangchen-amd
requested review from
BBuf,
DarkSharpness,
HaiShaw,
HydraQYH,
celve and
yuan-luo
as code owners
September 26, 2026 14:21
This was referenced Sep 26, 2026
The ROCm heuristic picks a 512-token K tile for batches of at most four. That tile needs 128 KiB of LDS: fine on gfx950, but gfx942 has 64 KiB, so eager decode and CUDA graph capture of small batches fail with Triton OutOfResources. Use the 512 tile on gfx95 only. Co-authored-by: Cursor <cursoragent@cursor.com>
siliangchen-amd
force-pushed
the
gfx942-mxfp8-moe-runner
branch
from
September 27, 2026 01:26
7c308a3 to
5599aa2
Compare
This was referenced Oct 2, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
MiniMaxAI/MiniMax-M3-MXFP8does not start on gfx942 (MI300X / MI325X) on main. #41377 fixed the MXFP8 MoE runner part of this; what remains is the sparse decode score tile.Since #36527 the ROCm heuristic in
minimax_sparse/decode/flash_with_topk_idx.pypicksBLOCK_SIZE_N=512for batches of at most 4. A 512-token K tile needs 128 KiB of LDS. gfx950 has 160 KiB and gfx942 has 64 KiB:This happens in eager decode and when capturing CUDA graphs for bs 1/2/4. The 512 tile is now gfx95-only; gfx942 uses 128, the tile it already used for larger batches. The heuristic only exists on ROCm, so CUDA is unaffected, and gfx950 is unchanged.
End-to-end (MI325X,
SGLANG_USE_AITER=1,lm_evalGSM8K, 200 questions)425a1f8f24, includes #41377)OutOfResourcesabove)TP8 also needs ROCm/aiter#5873. The aiter Triton blockscale GEMM reads past K when split-K rounds the last partition up, and MiniMax-M3's TP8 shared-expert
down_projhas K=384. Without that fix, TP8 decodes token 0 and CUDA-graph capture can hit a memory access fault. TP4 is not affected.Test plan
425a1f8f24.Found by Hyperloom; reviewed and measured by hand.
Hyperloom was optimizing
MiniMaxAI/MiniMax-M3-MXFP8on MI325X (TP8); the server did not start on gfx942.CI States
Latest PR Test (Base): ❌ Run #36285533164
Latest PR Test (Extra): ❌ Run #36285533049
Latest PR Test (AMD ROCm 10): ❌ Run #36285533206