Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
57 changes: 54 additions & 3 deletions .github/workflows/nightly-test-amd.yml
Original file line number Diff line number Diff line change
Expand Up @@ -90,7 +90,8 @@ on:
- nightly-8-gpu-qwen35
- nightly-8-gpu-mi35x-qwen35
- nightly-8-gpu-mi35x-qwen35-triton-dcp
# 8-GPU GLM-5.3-Flash (MI35x)
# 8-GPU GLM-5.3-Flash (MI30x + MI35x)
- nightly-8-gpu-glm53-flash
- nightly-8-gpu-mi35x-glm53-flash
# 8-GPU GLM-5.3 FP8 (MI30x + MI35x)
- nightly-8-gpu-glm53
Expand Down Expand Up @@ -1959,9 +1960,58 @@ jobs:
exit ${TEST_EXIT_CODE:-0}

# ==============================================================================
# 8-GPU GLM-5.3-Flash (MI35x)
# 8-GPU GLM-5.3-Flash (MI30x + MI35x)
#
# Both arches are gated because they run different kernels for the same
# forward pass: gfx950 takes the AITER mHC pre/post, gfx942 the generic mHC
# path. Both use the AITER MoE runner; see the MI30x test for why gfx942
# does not use the cookbook's Triton runner.
# ==============================================================================

nightly-8-gpu-glm53-flash:
name: ${{ format('nightly-8-gpu-glm53-flash ({0}, linux-mi300-8gpu-sglang)', matrix.rocm_version) }}
strategy:
fail-fast: false
matrix:
rocm_version: ${{ fromJson(inputs.rocm_version && inputs.rocm_version != 'all' && format('["{0}"]', inputs.rocm_version) || '["rocm10", "rocm724", "rocm720"]') }}
if: (github.repository == 'sgl-project/sglang' || github.event_name == 'pull_request') && (!(inputs.job_filter || inputs.job_select) || (inputs.job_filter || inputs.job_select) == 'all' || contains(format(',{0},', inputs.job_filter || inputs.job_select), ',nightly-8-gpu-glm53-flash,'))
runs-on: linux-mi300-8gpu-sglang
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
ref: ${{ inputs.ref || github.sha }}

- name: Ensure VRAM is clear
run: bash scripts/ci/amd/ensure_vram_clear.sh rocm

- name: Setup docker (${{ matrix.rocm_version }})
run: |
touch github_summary.md
bash scripts/ci/amd/amd_ci_start_container.sh --rocm-version ${{ matrix.rocm_version }}
env:
GITHUB_WORKSPACE: ${{ github.workspace }}
ENABLE_CACHE_HOST: "1"

- name: Install dependencies
run: |
bash scripts/ci/amd/amd_ci_install_dependency.sh --skip-test-time-deps
bash scripts/ci/amd/amd_ci_exec.sh pip install tabulate

# Measured 4419 s on rocm10, 2769 s of it loading the 328 GB checkpoint;
# this pool's shared cache has taken up to 4650 s for that load alone.
# 18000 s leaves room for a slower image or a cold cache.
- name: Accuracy Test ROCm (8-GPU GLM-5.3-Flash DSA+KDA)
timeout-minutes: 360
run: |
> github_summary.md # Clear summary file
bash scripts/ci/amd/amd_ci_exec.sh -w /sglang-checkout/test \
-e SGLANG_MOE_COPY_WEIGHT_VIEWS_BEFORE_H2D=1 \
-e GITHUB_STEP_SUMMARY="/sglang-checkout/github_summary.md" \
python3 run_suite.py --hw amd --suite nightly-amd-accuracy-8-gpu-glm53-flash --nightly --timeout-per-file 18000 ${{ (github.event_name == 'schedule' || inputs.continue_on_error) && '--continue-on-error' || '' }} || TEST_EXIT_CODE=$?
echo "$(<github_summary.md )" >> $GITHUB_STEP_SUMMARY || true
exit ${TEST_EXIT_CODE:-0}

nightly-8-gpu-mi35x-glm53-flash:
name: ${{ format('nightly-8-gpu-mi35x-glm53-flash ({0}, linux-mi35x-gpu-8)', matrix.rocm_version) }}
strategy:
Expand Down Expand Up @@ -2353,7 +2403,8 @@ jobs:
- nightly-8-gpu-qwen35
- nightly-8-gpu-mi35x-qwen35
- nightly-8-gpu-mi35x-qwen35-triton-dcp
# 8-GPU GLM-5.3-Flash (MI35x)
# 8-GPU GLM-5.3-Flash (MI30x + MI35x)
- nightly-8-gpu-glm53-flash
- nightly-8-gpu-mi35x-glm53-flash
# 8-GPU GLM-5.3 FP8 (MI30x + MI35x)
- nightly-8-gpu-glm53
Expand Down
173 changes: 173 additions & 0 deletions test/registered/amd/accuracy/mi30x/test_glm53_flash_eval_mi30x.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,173 @@
"""MI30x GLM-5.3-Flash GSM8K Accuracy Evaluation Test (8-GPU)

Tests zai-org/GLM-5.3-Flash on MI30x (gfx942) with the AMD FP8 recipe from the
GLM-5.3-Flash cookbook refresh (#36712): TP8 + EP8, BF16 KV cache, TileLang DSA
prefill+decode, Triton linear attention, SGLANG_USE_AITER=1, full decode graphs
at batch sizes 1 and 32. Same eval and threshold as the gfx950 gate in
test_glm53_flash_eval_mi35x.py.

MoE runner: this deviates from the cookbook cell, which names the Triton runner
for MI300X. On current main the Triton MoE runner makes this model generate
without ever stopping on gfx942: every sequence runs to the token cap and
scores zero. Measured in one job that ran three 64-question evals back to back
on the same server image (run 36119287477, 2026-09-26): Triton MoE with the
default fused sgl-kernel top-k scored 0/64 at 2048+ tokens per sequence,
Triton MoE with #36607's portable Torch top-k also scored 0/64 at 2048+ tokens
per sequence, and the AITER MoE runner scored 63/64 at 151 tokens per sequence.
The DSA top-k backend makes no difference, so the fault is the Triton MoE
runner itself, and the cookbook's MI300X MoE recommendation is stale.

gfx942 is not redundant with gfx950 for this model. It runs the generic mHC
path, since AITER mHC is gfx95-only, and nothing else gives that path nightly
coverage for this model. That path reaches gfx942 only with the HIP guard in
#41136: without it, TileLang's HIP codegen cannot lower the tl.get_lane_idx in
the fused mHC post/pre kernel and decode graph capture dies with "Unresolved
call Op(tl.get_lane_idx)" (run 36079282524).

Measured on main plus both HIP fixes in #41136, which this test requires:
0.9750 (1286/1319) on the rocm10 image, with a 2769 s weight load, a 1391 s
eval and 4419 s of wall clock (run 36232707853). That is the same count the
gfx950 gate in test_glm53_flash_eval_mi35x.py scored, so the gfx942 fallback
paths are not costing accuracy relative to the gfx950 fast paths.

Threshold: 0.92 follows this repo's `measured - 0.05` convention for sgl-eval
gsm8k thresholds and matches the gfx950 gate, so the two arches stay directly
comparable.

Runtime: the 328 GB checkpoint has taken 2769-4650 s to load from this pool's
shared cache, and the eval 1333-2374 s on top of that. In run 36678520753 the
first launch on all three images was still loading at 5400 s and only the CI's
online retry brought the server up, so the launch timeout below is 9000 s. The
workflow allows 18000 s. If that ever proves tight, prefer raising it over
trimming the eval: a full-split score is what makes this arch's number
comparable to the gfx950 one.

Eval harness: sgl-eval's gsm8k through run_sgl_eval, rather than the legacy
few-shot scorer that run_combined_tests routes gsm8k to, because
GLM-5.3-Flash thinks by default and the legacy scorer reads the last number
in the response. The parameters below are the accuracy command the cookbook
publishes for this model. See the MI35x file for the longer note.

Registry: nightly-amd-accuracy-8-gpu-glm53-flash suite
"""

import unittest
from types import SimpleNamespace

from sglang.srt.utils import kill_process_tree
from sglang.test.accuracy_test_runner import (
AccuracyTestResult,
write_accuracy_github_summary,
)
from sglang.test.ci.ci_register import register_amd_ci
from sglang.test.sgl_eval_utils import run_sgl_eval
from sglang.test.test_utils import (
DEFAULT_URL_FOR_TEST,
ModelLaunchSettings,
popen_launch_server,
)

# Register for AMD CI - MI30x GLM-5.3-Flash accuracy test. The 9000 s launch
# budget below plus the slowest full-split eval measured (2374 s).
register_amd_ci(
est_time=11400,
suite="nightly-amd-accuracy-8-gpu-glm53-flash",
nightly=True,
)

GLM_53_FLASH_MODEL_PATH = "zai-org/GLM-5.3-Flash"
BASELINE_ACCURACY = 0.92

# Fetching and loading a 328 GB checkpoint against a cold cache is what this
# budget has to cover; the default launch timeout is nowhere near enough.
# Loads have measured up to 4650 s, and a first launch has run past 5400 s.
SERVER_LAUNCH_TIMEOUT = 9000


class TestGLM53FlashEvalMI30x(unittest.TestCase):
"""GLM-5.3-Flash GSM8K Accuracy Evaluation Test for MI30x."""

def test_glm_53_flash(self):
"""Run accuracy test for GLM-5.3-Flash."""
cookbook_args = [
"--ep-size=8",
"--attention-backend=dsa",
"--dsa-prefill-backend=tilelang",
"--dsa-decode-backend=tilelang",
"--linear-attn-backend=triton",
"--kv-cache-dtype=bfloat16",
# Not the cookbook's Triton runner; see the MoE runner note above.
"--moe-runner-backend=aiter",
"--cuda-graph-backend-decode=full",
"--cuda-graph-backend-prefill=disabled",
"--cuda-graph-bs-decode",
"1",
"32",
"--reasoning-parser=glm45",
"--tool-call-parser=glm47",
"--watchdog-timeout=1200",
# Not part of the cookbook cell; purely a load-time win on a
# checkpoint this large, with no effect on numerics.
"--model-loader-extra-config",
'{"enable_multithread_load": true}',
]

model = ModelLaunchSettings(
GLM_53_FLASH_MODEL_PATH,
tp_size=8,
extra_args=cookbook_args,
env={"SGLANG_USE_AITER": "1"},
variant="TP8-EP8",
)

# run_combined_tests routes gsm8k to the legacy scorer, so launch the
# server here and hand the eval to sgl-eval directly.
base_url = DEFAULT_URL_FOR_TEST
process = popen_launch_server(
model.model_path,
base_url,
timeout=SERVER_LAUNCH_TIMEOUT,
other_args=model.extra_args,
env=model.env,
)
try:
metrics = run_sgl_eval(
SimpleNamespace(
base_url=base_url,
model=model.model_path,
eval_name="gsm8k",
num_examples=None,
num_threads=64,
max_tokens=32768,
temperature=1.0,
top_p=0.95,
seed=42,
sgl_eval_thinking=True,
)
)
finally:
kill_process_tree(process.pid)

score = metrics["score"]
passed = score >= BASELINE_ACCURACY
write_accuracy_github_summary(
"GLM-5.3-Flash (MI30x)",
"gsm8k",
[
AccuracyTestResult(
model=model.model_path,
dataset="gsm8k",
passed=passed,
score=score,
baseline_accuracy=BASELINE_ACCURACY,
error=None if passed else "below baseline",
latency=metrics.get("latency"),
variant=model.variant,
)
],
)
self.assertGreaterEqual(score, BASELINE_ACCURACY)


if __name__ == "__main__":
unittest.main()
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,8 @@
second memory pool alongside the paged KV pool, and the mHC pre/post ops sit on
every layer boundary. A single-arch gate would not be enough: gfx950 takes the
AITER mHC pre/post kernels, while gfx942 falls back to the generic mHC path.
This file gates the gfx950 half.
This file gates the gfx950 half; the gfx942 half is
test_glm53_flash_eval_mi30x.py.

Measured on current main: 0.9750 (1286/1319) on the rocm10 image, HF snapshot
eb9eb208eb0d988989d07a6a12d0fdeb5f52574a, with a 312 s weight load, a 473 s
Expand Down
1 change: 1 addition & 0 deletions test/run_suite.py
Original file line number Diff line number Diff line change
Expand Up @@ -160,6 +160,7 @@
"nightly-amd-accuracy-8-gpu-mi35x-kimi-k3",
"nightly-amd-8-gpu-mi35x-qwen38-mxfp4",
"nightly-amd-8-gpu-mi35x-glm52-fp8",
"nightly-amd-accuracy-8-gpu-glm53-flash",
"nightly-amd-8-gpu-mi35x-glm53-flash",
"nightly-amd-accuracy-8-gpu-glm53",
"nightly-amd-8-gpu-mi35x-glm53",
Expand Down
Loading