Skip to content

[AMD] Keep the MiniMax sparse decode score tile within gfx942 LDS - #41354

Open
siliangchen-amd wants to merge 1 commit into
sgl-project:mainfrom
siliangchen-amd:gfx942-mxfp8-moe-runner
Open

siliangchen-amd wants to merge 1 commit into
sgl-project:mainfrom
siliangchen-amd:gfx942-mxfp8-moe-runner

Conversation

@siliangchen-amd

@siliangchen-amd siliangchen-amd commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Summary

MiniMaxAI/MiniMax-M3-MXFP8 does not start on gfx942 (MI300X / MI325X) on main. #41377 fixed the MXFP8 MoE runner part of this; what remains is the sparse decode score tile.

Since #36527 the ROCm heuristic in minimax_sparse/decode/flash_with_topk_idx.py picks BLOCK_SIZE_N=512 for batches of at most 4. A 512-token K tile needs 128 KiB of LDS. gfx950 has 160 KiB and gfx942 has 64 KiB:

triton.runtime.errors.OutOfResources: out of resource: shared memory, Required: 131072, Hardware limit: 65536

This happens in eager decode and when capturing CUDA graphs for bs 1/2/4. The 512 tile is now gfx95-only; gfx942 uses 128, the tile it already used for larger batches. The heuristic only exists on ROCm, so CUDA is unaffected, and gfx950 is unchanged.

End-to-end (MI325X, SGLANG_USE_AITER=1, lm_eval GSM8K, 200 questions)

setup main (425a1f8f24, includes #41377) this PR
TP4, CUDA graphs crashes in CUDA-graph capture (OutOfResources above) serves; GSM8K 0.975 / 0.975 (strict / flexible)
TP8, CUDA graphs not run serves; GSM8K 0.96 / 0.96

TP8 also needs ROCm/aiter#5873. The aiter Triton blockscale GEMM reads past K when split-K rounds the last partition up, and MiniMax-M3's TP8 shared-expert down_proj has K=384. Without that fix, TP8 decodes token 0 and CUDA-graph capture can hit a memory access fault. TP4 is not affected.

Test plan

  • End-to-end runs above, on main at 425a1f8f24.

Found by Hyperloom; reviewed and measured by hand.
Hyperloom was optimizing MiniMaxAI/MiniMax-M3-MXFP8 on MI325X (TP8); the server did not start on gfx942.


CI States

Latest PR Test (Base): ❌ Run #36285533164
Latest PR Test (Extra): ❌ Run #36285533049
Latest PR Test (AMD ROCm 10): ❌ Run #36285533206

The ROCm heuristic picks a 512-token K tile for batches of at most four.
That tile needs 128 KiB of LDS: fine on gfx950, but gfx942 has 64 KiB, so
eager decode and CUDA graph capture of small batches fail with Triton
OutOfResources. Use the 512 tile on gfx95 only.

Co-authored-by: Cursor <cursoragent@cursor.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant