Skip to content

[Config] [Tune] Add GLM-5.3 Flash GEMM configs for gfx950 - #5602

Merged
yifehuan merged 7 commits into
ROCm:mainfrom
jamesETsmith:feature/glm53-flash-gfx950-tuning
Sep 20, 2026
Merged

yifehuan merged 7 commits into
ROCm:mainfrom
jamesETsmith:feature/glm53-flash-gfx950-tuning

Conversation

@jamesETsmith

@jamesETsmith jamesETsmith commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • add gfx950 A8W8 block-scale GEMM configurations for GLM-5.3 Flash projection shapes
  • add 171 accepted BF16 GEMM configurations selected from 1,826 captured GLM-5.3 Flash shapes
  • preserve the corresponding untuned BF16 shape set for reproducibility and future retuning

Results

  • BF16 tuning completed all tuning on 8 x MI355X (gfx950, 256 CUs)
  • accepted backends: 99 FlyDSL, 35 ASM, 34 Torch, and 3 Opus
  • Still running end to end test to measure the gains and will post those results shortly

Test plan

  • Process all 1,826 captured BF16 shapes
  • Complete all 74 tuning batches
  • Run matched untuned and tuned vLLM sweeps across MTP=0,5
  • Validate all 56 benchmark outputs for request count, failures, and output-token count

Made with Cursor

stefanskiasan and others added 3 commits September 4, 2026 19:55
The GLM-5.3 projection shapes (N=6144 with K=3072/4096/6144, plus
N=2624/K=6144) have no gfx950 entries in a8w8_blockscale_tuned_gemm.csv,
so they fall back to an untuned config. On MI355X that lands at roughly
19.9 us and 1900 GB/s for M<=32.

Tuned with the project's own tuner:

  gemm_a8w8_blockscale_tune.py --libtype all --splitK --compare \
      --update_improved --min_improvement_pct 3 --mp 4

All 24 low-M shapes improved by 13-61 %; (1, 6144, 6144) goes from
20.67 us to 11.00 us and 1900 -> 3458 GB/s. The winning configs use
splitK=3: at NPerBlock=64 and N=6144 only 96 workgroups are launched on
256 CUs, so splitting K is what fills the machine.

A second sweep over M=48..512 is included; there the effect fades as
expected once there is enough work (only 14 of 32 shapes improved by the
3 % threshold, and at N=6144/K=6144 only M=48 and M=64).

End to end in vLLM (TP4, GLM-5.3 MXFP4 + fp8 attention), single-stream
decode went from 171.3 to 184.7 tok/s, GSM8K unchanged at 148/150.

Related: ROCm#5277, which is about splitK not being searched by default.

Signed-off-by: Asan Stefanski <asan.stefanski.claude@gmail.com>
Add tuned block-scale GEMM configurations for TP4 and TP8 workloads,
reducing geometric-mean latency by 39.5% across retained shapes.

💘 Generated with Crush

Assisted-by: Crush:gpt-5.6-sol
Preserve the captured workload shapes and their MI355X winners so AITER can reuse the tuned configurations without rerunning the experiment.

Co-authored-by: Cursor <cursoragent@cursor.com>
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5602 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@jamesETsmith

jamesETsmith commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

E2E update: GLM-5.3-Flash on 8× MI350X (gfx950)

Matched A/B of this PR's AITER configs vs baseline 849709d641853970af8f61973a2242fdf87f059f.

Item Value
Hardware 8× Instinct MI350X (gfx950, 256 CU)
Model zai-org/GLM-5.3-Flash @ 3f1971b7b5f7a528c9c4ef6212c8785298a8c24a
Image vllm/vllm-openai-rocm@sha256:f169e8df365b5112a1e6500543e13d8161374277c22c94d8ed60817324c101e1 (nightly-0bfc7a15d)
Dataset ShareGPT Vicuna unfiltered, OSL=1024, --ignore-eos
Sweep concurrency 1, 2, 3, 4, 8, 16, 32; requests = 10× concurrency
Warmup 16 requests per point

MTP=0 total-token throughput (mean ± 1 SD, n=3)

Conc Baseline tok/s Tuned tok/s Δ throughput
1 105.2 ± 0.1 105.9 ± 0.2 +0.60 ± 0.07%
2 199.4 ± 0.1 202.2 ± 0.2 +1.45 ± 0.09%
3 252.4 ± 0.3 262.0 ± 0.1 +3.81 ± 0.14%
4 332.0 ± 0.4 342.8 ± 2.3 +3.25 ± 0.61%
8 697.8 ± 0.4 708.3 ± 3.9 +1.51 ± 0.51%
16 1228.5 ± 4.5 1247.0 ± 2.0 +1.51 ± 0.28%
32 1976.0 ± 13.1 2002.4 ± 0.8 +1.34 ± 0.70%

@jamesETsmith
jamesETsmith marked this pull request as ready for review September 18, 2026 18:41
@jamesETsmith
jamesETsmith requested a review from a team September 18, 2026 18:41
@github-actions github-actions Bot changed the title [Tune] Add GLM-5.3 Flash GEMM configs for gfx950 [Config] [Tune] Add GLM-5.3 Flash GEMM configs for gfx950 Sep 18, 2026
@jamesETsmith

Copy link
Copy Markdown
Contributor Author

E2E update: SPEED-Bench 8K on 8× MI350X (gfx950)

The ShareGPT benchmark was on really small ISL, so I also ran before and after benchmarks on Speed-Bench 8k. Matched A/B of this PR's AITER configs versus baseline 849709d

Config

Item Value
Hardware 8× Instinct MI350X (gfx950, 256 CU)
Model zai-org/GLM-5.3-Flash @ 3f1971b7b5f7a528c9c4ef6212c8785298a8c24a
Image vllm/vllm-openai-rocm@sha256:f169e8df365b5112a1e6500543e13d8161374277c22c94d8ed60817324c101e1 (nightly-0bfc7a15d)
Dataset NVIDIA SPEED-Bench throughput_8k, low_entropy
Sequence lengths Mean ISL ≈8,013 tokens; OSL=1,024; --ignore-eos
Sweep Concurrency 1, 2, 3, 4, 8, 16, 32; requests = 10× concurrency
Warmup 16 requests per point

MTP=0 output-token throughput

Conc Baseline tok/s Tuned tok/s Δ throughput
1 83.1 83.6 +0.51%
2 149.2 150.7 +1.04%
3 180.2 185.5 +2.90%
4 250.9 258.1 +2.89%
8 479.9 486.4 +1.36%
16 905.1 916.8 +1.29%
32 1,286.8 1,361.0 +5.77%

The tuned configs improved output throughput at all seven tested concurrency points, with a mean improvement of +2.25%. All requests completed successfully and every request generated exactly 1,024 tokens.

Required vLLM nightly compatibility fixes

The pinned nightly required the following ROCm fixes, applied identically to
both A/B legs:

  • vllm#57252 — add the missing record_logical_topk_ready() hook to ROCMAiterMLASparseImpl.
  • vllm#57425 — route the AMD SparseAttnIndexerKpool through forward_hip.
  • Disabled broken JIT kernel warmup as documented in vllm#57248.
  • Set VLLM_USE_BREAKABLE_CUDAGRAPH=0 for the gfx950 graph-capture fault documented in vllm#57227.

These compatibility changes were constant across baseline and tuned runs; the
only A/B variable was the mounted AITER configuration tree.

Keep the faster existing winners so merged config generation succeeds on the first CI pass.

Co-authored-by: Cursor <cursoragent@cursor.com>
@jamesETsmith

Copy link
Copy Markdown
Contributor Author

e2e update: Speed-Bench 8k on 8xMI350 with MTP=5

Used the same setup as outlined above: #5602 (comment)

Improvement

Conc Baseline tok/s Tuned tok/s Δ
1 872.8 908.1 +4.04%
2 2640.7 2585.4 −2.09%
3 3804.9 3921.2 +3.06%
4 4881.7 5309.1 +8.75%
8 7011.2 7156.0 +2.07%
16 11785.4 11534.5 −2.13%
32 14432.4 15129.3 +4.83%

I reran this on a second node and got similar improvement and the conc=2 results switched sign to show an improvement so I don't think the negative change there or with conc=16 are anything to be concerned about.

@zufayu
zufayu requested a review from yifehuan September 20, 2026 01:11

@yifehuan yifehuan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@yifehuan
yifehuan merged commit 9f9cc3a into ROCm:main Sep 20, 2026
56 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants