Repository navigation
[AMD] Add GLM-5.3-Flash MI30x nightly GSM8K accuracy gate - #41475
Draft
michaelzhang-ai wants to merge 2 commits into
Draft
michaelzhang-ai wants to merge 2 commits into
michaelzhang-ai wants to merge 2 commits into
Conversation
Base automatically changed from
cursor/glm53-flash-amd-nightly-accuracy-748f
to
main
September 28, 2026 05:40
michaelzhang-ai
force-pushed
the
amd/glm53-flash-mi30x-nightly-accuracy
branch
2 times, most recently
from
September 30, 2026 01:26
f20a237 to
a8be4df
Compare
michaelzhang-ai
added a commit
that referenced
this pull request
Sep 30, 2026
This was referenced Sep 30, 2026
Adds nightly-8-gpu-glm53-flash on linux-mi300-8gpu-sglang, running test_glm53_flash_eval_mi30x.py (suite nightly-amd-accuracy-8-gpu-glm53-flash) across rocm10/rocm724/rocm720 with an 18000 s per-file cap. Same recipe, sgl-eval gsm8k harness and 0.92 threshold as the MI35x gate, except the MoE runner is AITER rather than Triton on gfx942. Requires the mHC HIP guard in #41136 to pass on gfx942.
In run 36678520753 the first launch on all three images was still loading the 328 GB checkpoint at the 5400 s timeout, and only the CI's online retry brought the server up. Successful loads took 4427-4650 s. Raise the launch timeout to 9000 s and est_time to 11400 s (launch budget plus the slowest eval measured, 2374 s). The workflow's 18000 s per-file cap already covers this.
michaelzhang-ai
force-pushed
the
amd/glm53-flash-mi30x-nightly-accuracy
branch
from
October 4, 2026 14:09
3143825 to
d699ba8
Compare
This was referenced Oct 7, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds the gfx942 half of the GLM-5.3-Flash nightly GSM8K gate:
test/registered/amd/accuracy/mi30x/test_glm53_flash_eval_mi30x.pynightly-amd-accuracy-8-gpu-glm53-flashnightly-8-gpu-glm53-flashonlinux-mi300-8gpu-sglang, with an 18000 s per-file capThis is split out of #36903 (merged), which added the MI35x gate. It uses the same recipe, eval and 0.92 threshold.
Depends on
#41136. Without its mHC HIP guard, decode graph capture on gfx942 fails with
Unresolved call Op(tl.get_lane_idx)(run 36079282524). This PR stays in draft until #41136 lands.Why
gfx942 does not get the fast paths gfx950 does. mHC runs on the generic path, because AITER mHC is gfx95-only, and nothing else gives that path nightly coverage for this model.
Notes for the reviewer
Test plan
All three images on the current head with [AMD] [GLM-5.3-Flash] Restore two HIP fallbacks: clamped-SwiGLU Triton MoE and the fused mHC boundary #41136 merged in, full 1319-question split (run 36798759252, branch
amd/glm53-flash-mi30x-validation= this PR's3143825a63+ [AMD] [GLM-5.3-Flash] Restore two HIP fallbacks: clamped-SwiGLU Triton MoE and the fused mHC boundary #41136's2796d33002):rocm10rocm720rocm724All three pass the 0.92 threshold with no request errors, and each server came up on its first launch. Loads were fast in this run, so it does not exercise the slow-cache case the 9000 s timeout is for.
The previous run, on
a8be4dfab8(run 36678520753), also passed on all three images: 0.9735 / 0.9720 / 0.9704 onrocm10/rocm720/rocm724. There, though, the first server launch on every image timed out at the then 5400 sSERVER_LAUNCH_TIMEOUTwhile still loading weights, and only the CI's automatic online retry brought the server up, after loads of 4427-4650 s.3143825a63raisesSERVER_LAUNCH_TIMEOUTto 9000 s andest_timeto 11400 s so a load that slow succeeds on the first attempt; the workflow's 18000 s per-file cap already covers that.Earlier, on the previous harness: full split on
rocm10with [AMD] [GLM-5.3-Flash] Restore two HIP fallbacks: clamped-SwiGLU Triton MoE and the fused mHC boundary #41136 applied scored 0.9750 (1286/1319), 4419 s wall clock including a 2769 s weight load (run 36232707853).The test calls
run_sgl_evaldirectly instead of going throughrun_combined_tests, because #41280 removed thesgl_eval_thinkingparameter the earlier harness relied on. Server arguments and sampling are unchanged.CI States
Latest PR Test (Base): ❌ Run #37208245343
Latest PR Test (Extra): ❌ Run #37208245229
Latest PR Test (AMD ROCm 10): ❌ Run #37208245463