Repository navigation
[Config] [Tune] Add GLM-5.3 Flash GEMM configs for gfx950 - #5602
Conversation
The GLM-5.3 projection shapes (N=6144 with K=3072/4096/6144, plus
N=2624/K=6144) have no gfx950 entries in a8w8_blockscale_tuned_gemm.csv,
so they fall back to an untuned config. On MI355X that lands at roughly
19.9 us and 1900 GB/s for M<=32.
Tuned with the project's own tuner:
gemm_a8w8_blockscale_tune.py --libtype all --splitK --compare \
--update_improved --min_improvement_pct 3 --mp 4
All 24 low-M shapes improved by 13-61 %; (1, 6144, 6144) goes from
20.67 us to 11.00 us and 1900 -> 3458 GB/s. The winning configs use
splitK=3: at NPerBlock=64 and N=6144 only 96 workgroups are launched on
256 CUs, so splitting K is what fills the machine.
A second sweep over M=48..512 is included; there the effect fades as
expected once there is enough work (only 14 of 32 shapes improved by the
3 % threshold, and at N=6144/K=6144 only M=48 and M=64).
End to end in vLLM (TP4, GLM-5.3 MXFP4 + fp8 attention), single-stream
decode went from 171.3 to 184.7 tok/s, GSM8K unchanged at 148/150.
Related: ROCm#5277, which is about splitK not being searched by default.
Signed-off-by: Asan Stefanski <asan.stefanski.claude@gmail.com>
Add tuned block-scale GEMM configurations for TP4 and TP8 workloads, reducing geometric-mean latency by 39.5% across retained shapes. 💘 Generated with Crush Assisted-by: Crush:gpt-5.6-sol
Preserve the captured workload shapes and their MI355X winners so AITER can reuse the tuned configurations without rerunning the experiment. Co-authored-by: Cursor <cursoragent@cursor.com>
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
E2E update: GLM-5.3-Flash on 8× MI350X (
|
| Item | Value |
|---|---|
| Hardware | 8× Instinct MI350X (gfx950, 256 CU) |
| Model | zai-org/GLM-5.3-Flash @ 3f1971b7b5f7a528c9c4ef6212c8785298a8c24a |
| Image | vllm/vllm-openai-rocm@sha256:f169e8df365b5112a1e6500543e13d8161374277c22c94d8ed60817324c101e1 (nightly-0bfc7a15d) |
| Dataset | ShareGPT Vicuna unfiltered, OSL=1024, --ignore-eos |
| Sweep | concurrency 1, 2, 3, 4, 8, 16, 32; requests = 10× concurrency |
| Warmup | 16 requests per point |
MTP=0 total-token throughput (mean ± 1 SD, n=3)
| Conc | Baseline tok/s | Tuned tok/s | Δ throughput |
|---|---|---|---|
| 1 | 105.2 ± 0.1 | 105.9 ± 0.2 | +0.60 ± 0.07% |
| 2 | 199.4 ± 0.1 | 202.2 ± 0.2 | +1.45 ± 0.09% |
| 3 | 252.4 ± 0.3 | 262.0 ± 0.1 | +3.81 ± 0.14% |
| 4 | 332.0 ± 0.4 | 342.8 ± 2.3 | +3.25 ± 0.61% |
| 8 | 697.8 ± 0.4 | 708.3 ± 3.9 | +1.51 ± 0.51% |
| 16 | 1228.5 ± 4.5 | 1247.0 ± 2.0 | +1.51 ± 0.28% |
| 32 | 1976.0 ± 13.1 | 2002.4 ± 0.8 | +1.34 ± 0.70% |
E2E update: SPEED-Bench 8K on 8× MI350X (
|
| Item | Value |
|---|---|
| Hardware | 8× Instinct MI350X (gfx950, 256 CU) |
| Model | zai-org/GLM-5.3-Flash @ 3f1971b7b5f7a528c9c4ef6212c8785298a8c24a |
| Image | vllm/vllm-openai-rocm@sha256:f169e8df365b5112a1e6500543e13d8161374277c22c94d8ed60817324c101e1 (nightly-0bfc7a15d) |
| Dataset | NVIDIA SPEED-Bench throughput_8k, low_entropy |
| Sequence lengths | Mean ISL ≈8,013 tokens; OSL=1,024; --ignore-eos |
| Sweep | Concurrency 1, 2, 3, 4, 8, 16, 32; requests = 10× concurrency |
| Warmup | 16 requests per point |
MTP=0 output-token throughput
| Conc | Baseline tok/s | Tuned tok/s | Δ throughput |
|---|---|---|---|
| 1 | 83.1 | 83.6 | +0.51% |
| 2 | 149.2 | 150.7 | +1.04% |
| 3 | 180.2 | 185.5 | +2.90% |
| 4 | 250.9 | 258.1 | +2.89% |
| 8 | 479.9 | 486.4 | +1.36% |
| 16 | 905.1 | 916.8 | +1.29% |
| 32 | 1,286.8 | 1,361.0 | +5.77% |
The tuned configs improved output throughput at all seven tested concurrency points, with a mean improvement of +2.25%. All requests completed successfully and every request generated exactly 1,024 tokens.
Required vLLM nightly compatibility fixes
The pinned nightly required the following ROCm fixes, applied identically to
both A/B legs:
- vllm#57252 — add the missing
record_logical_topk_ready()hook toROCMAiterMLASparseImpl. - vllm#57425 — route the AMD
SparseAttnIndexerKpoolthroughforward_hip. - Disabled broken JIT kernel warmup as documented in vllm#57248.
- Set
VLLM_USE_BREAKABLE_CUDAGRAPH=0for the gfx950 graph-capture fault documented in vllm#57227.
These compatibility changes were constant across baseline and tuned runs; the
only A/B variable was the mounted AITER configuration tree.
Keep the faster existing winners so merged config generation succeeds on the first CI pass. Co-authored-by: Cursor <cursoragent@cursor.com>
e2e update: Speed-Bench 8k on 8xMI350 with MTP=5Used the same setup as outlined above: #5602 (comment) Improvement
I reran this on a second node and got similar improvement and the conc=2 results switched sign to show an improvement so I don't think the negative change there or with conc=16 are anything to be concerned about. |
Summary
Results
Test plan
Made with Cursor