Skip to content

[AMD] Add GLM-5.3-Flash MI30x nightly GSM8K accuracy gate - #41475

Draft
michaelzhang-ai wants to merge 2 commits into
mainfrom
amd/glm53-flash-mi30x-nightly-accuracy
Draft

michaelzhang-ai wants to merge 2 commits into
mainfrom
amd/glm53-flash-mi30x-nightly-accuracy

Conversation

@michaelzhang-ai

@michaelzhang-ai michaelzhang-ai commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

What

Adds the gfx942 half of the GLM-5.3-Flash nightly GSM8K gate:

  • Test: test/registered/amd/accuracy/mi30x/test_glm53_flash_eval_mi30x.py
  • Suite: nightly-amd-accuracy-8-gpu-glm53-flash
  • Job: nightly-8-gpu-glm53-flash on linux-mi300-8gpu-sglang, with an 18000 s per-file cap

This is split out of #36903 (merged), which added the MI35x gate. It uses the same recipe, eval and 0.92 threshold.

Depends on

#41136. Without its mHC HIP guard, decode graph capture on gfx942 fails with Unresolved call Op(tl.get_lane_idx) (run 36079282524). This PR stays in draft until #41136 lands.

Why

gfx942 does not get the fast paths gfx950 does. mHC runs on the generic path, because AITER mHC is gfx95-only, and nothing else gives that path nightly coverage for this model.

Notes for the reviewer

  • MoE runner: the test uses the AITER MoE runner rather than the Triton runner that the recipe names for MI300X. On gfx942 the Triton runner produces sequences that never stop and score 0/64. That happened with both DSA top-k backends, while AITER scored 63/64 in the same job (run 36119287477).
  • Budget: the 328 GB checkpoint has taken 2769 to 4143 s to load from this pool's shared cache. The 18000 s cap is deliberately generous so that the full split is kept, which keeps this score comparable to the MI35x one.

Test plan


CI States

Latest PR Test (Base): ❌ Run #37208245343
Latest PR Test (Extra): ❌ Run #37208245229
Latest PR Test (AMD ROCm 10): ❌ Run #37208245463

@github-actions github-actions Bot added the amd label Sep 27, 2026
Base automatically changed from cursor/glm53-flash-amd-nightly-accuracy-748f to main September 28, 2026 05:40
@michaelzhang-ai
michaelzhang-ai force-pushed the amd/glm53-flash-mi30x-nightly-accuracy branch 2 times, most recently from f20a237 to a8be4df Compare September 30, 2026 01:26
michaelzhang-ai added a commit that referenced this pull request Sep 30, 2026
Adds nightly-8-gpu-glm53-flash on linux-mi300-8gpu-sglang, running
test_glm53_flash_eval_mi30x.py (suite nightly-amd-accuracy-8-gpu-glm53-flash)
across rocm10/rocm724/rocm720 with an 18000 s per-file cap. Same recipe,
sgl-eval gsm8k harness and 0.92 threshold as the MI35x gate, except the MoE
runner is AITER rather than Triton on gfx942.

Requires the mHC HIP guard in #41136 to pass on gfx942.
In run 36678520753 the first launch on all three images was still
loading the 328 GB checkpoint at the 5400 s timeout, and only the CI's
online retry brought the server up. Successful loads took 4427-4650 s.
Raise the launch timeout to 9000 s and est_time to 11400 s (launch
budget plus the slowest eval measured, 2374 s). The workflow's 18000 s
per-file cap already covers this.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant