Repository navigation
Conversation
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
There was a problem hiding this comment.
🔵 Needs a closer look
Tuning provenance remains unresolved, warranting final human review.
Pull request overview
Adds GLM-5.3-Flash fused-MoE tuning configurations for gfx942.
Changes:
- Adds fifteen untuned token-bucket inputs.
- Adds five tuned single-stage runtime configurations.
Review note (nit): Add benchmark command, ROCm/GPU environment, and checkpoint provenance for reproducibility.
File summaries
| File | Description |
|---|---|
aiter/configs/model_configs/a8w8_blockscale_untuned_fmoe_glm5_3_flash.csv |
Complete tuner input shape ladder. |
aiter/configs/model_configs/a8w8_blockscale_tuned_fmoe_glm5_3_flash.csv |
Tuned gfx942 runtime configurations. |
Review details
Suppressed comments (1)
aiter/configs/model_configs/a8w8_blockscale_tuned_fmoe_glm5_3_flash.csv:1
- [verified] The tuned rows include measured timings, but the PR description does not record the exact benchmark entry point/command, ROCm version, GPU model, or checkpoint. That makes it difficult to reproduce the default-vs-tuned gate when this table is revisited or regenerated. Author must add the command and environment/checkpoint provenance to the PR or a checked-in tuning record.
gfx,cu_num,token,model_dim,inter_dim,expert,topk,act_type,dtype,q_dtype_a,q_dtype_w,q_type,use_g1u1,doweight_stage1,block_m,ksplit,us1,kernelName1,err1,us2,kernelName2,err2,us,run_1stage,xbf16,flat,tflops,bw,_tag
- Files reviewed: 2/2 changed files
- Comments generated: 0
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
Advisory review (static + hand-run; not a merge gate). Validation/Perf ran no GPU stage — reasons are on their lines below. Findings tagged Adds two gfx942 data files for the GLM-5.3-Flash routed fused-MoE shape (model_dim 4096, inter_dim 512, 288 experts, topk 8, fp8 e4m3fnuz per-1x128 blockscale): a five-row tuned table that selects single-stage 64x256 asm kernels at token buckets 1024-16384, plus the fifteen-bucket tuner-input ladder it was tuned from. Review (advisory): |
Thanks @zufayu — provenance was the same ask in each, so it's now in the description: new Environment and Commands sections (exact tuner invocations, container, GPU/driver/ROCm, and the aiter commit for each run), plus a rewritten Verification. On "not independently reproduced" — now confirmed by a different method. A paired vLLM serving sweep rather than the tuner: 131k in / 1k out, 7 concurrencies, 20 requests each, TP4, one image holding both tables with A correction to my own description. I claimed keys are disjoint from all 35 existing On the |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The documented tuning paths do not match committed files, and downstream duplicate tables can cause order-dependent kernel selection.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 1
9d6f3e7 to
b1911bb
Compare
d1e2b2b to
80fa418
Compare
… table The TP2 token-2 row (E=289, topk=9, inter_dim 1024) beat the default dispatch by 3.3% and 4.8% in the two verification runs (72.3 us against 75.4 us), just over the 3% bar. The routed GLM-5.3-Flash table in ROCm#5500 dropped the same shape, with the same kernel, at 2.1% and 2.9%. Leave it to the default here as well, so the routed and fused shared tables cover the same buckets. Signed-off-by: Jin Tao <jintao12@amd.com>
…or gfx942
Tune the routed fused-MoE shape GLM-5.3-Flash dispatches, on gfx942 / cu_num
304. Kernel-time attribution on a TP4 serving run put the routed MoE at 15.1%
of prefill forward time against 2.1% for the a8w8_blockscale GEMMs, and the
MoE was still on default kernels, so it is where the remaining headroom is.
One shape, tuned over aiter's power-of-two token ladder:
model_dim 4096, inter_dim 512, expert 288, topk 8, Silu,
bf16 activations, fp8 e4m3fnuz per_1x128 blockscale, g1u1,
no doweight_stage1
inter_dim is 512 because moe_intermediate_size 2048 is sharded over TP4.
Tuning the bucket ladder covers every token count, since the lookup rounds to
these buckets.
Five of the fifteen buckets are kept. Improvement is end-to-end operator time
from --compare, tuned against the default kernel:
token default(us) tuned(us) speedup
1024 542.18 524.53 1.03x
2048 1002.53 670.24 1.50x
4096 1864.42 868.68 2.15x
8192 2812.08 1524.90 1.84x
16384 4336.83 2764.15 1.57x
Time-weighted speedup over the default across these five is 1.66x (10.56 ms
-> 6.35 ms summed), median per-shape 1.57x. The gain concentrates at large
token counts, so it matters for prefill and large batched decode.
The other ten buckets are deliberately absent, because the default kernel is
equal or faster there and falling back is the better choice:
- tokens 64, 128, 256 and 512 were rejected by the --update_improved gate on
two separate tuning runs, at -0.57%/-0.20%, -9.75%/-5.26%, +1.23%/-0.07%
and -0.80%/+0.33%.
- tokens 1 through 32 were rejected after a dedicated measurement. The tuner
had marked them "no_baseline" and admitted them on the post-run alone,
because the pre-run benchmark returned no timing for them. Benchmarking the
default path directly (--run_config with no CSV) against the tuned kernels
shows the tuned candidates are slower at five of the six, reproduced over
two runs agreeing within 1%:
token default(us) tuned(us) speedup
1 38.70 58.02 0.67x
2 51.55 61.83 0.83x
4 73.88 74.33 0.99x
8 113.85 111.26 1.02x
16 173.23 179.94 0.96x
32 255.15 295.34 0.87x
Only token 8 improves at all, by 2.3%, under the 3% threshold. Shipping
these would have regressed the small-batch decode path by up to 1.5x.
All five rows select block_m 64 and a 64x256 stage-1 kernel:
fmoe_bf16_blockscaleFp8_g1u1_vs_ps_silu_64x256 at tokens 1024, 4096, 8192 and
16384, and fmoe_bf16_blockscaleFp8_g1u1_vs_silu_64x256 at token 2048. Stage 2
is unused; every row runs single-stage.
The untuned sibling carries all fifteen buckets, not the five that were kept,
so it documents the search space the numbers above came from. Re-running the
tuner from it reproduces the same gate decisions.
Keys are disjoint from all 35 existing tuned_fmoe tables, so get_config_file()
merges this file with no duplicate-shape conflict. Verified by loading
AITER_CONFIG_FMOE_FILE with this file in model_configs/ and confirming all five
rows survive the merge. The untuned file is excluded by that glob, which skips
any name containing "untuned".
Co-authored-by: Cursor <cursoragent@cursor.com>
…d-MoE table Extend the routed GLM-5.3-Flash fused-MoE table (E=288, topk=8) from TP4 to TP2 (inter_dim 1024) and TP8 (inter_dim 256) on gfx942 / cu_num 304. Both were tuned over the same asm and cktile families as the TP4 rows, with --compare --update_improved. New rows are kept when the production operator improves by at least 3% in both of two default-against-tuned runs: TP2: tokens 1024-16384, 1.08x-1.86x TP8: tokens 1024-16384, 1.14x-2.37x TP8 also keeps token 1. The default dispatch there returns wrong output (a mismatch in both runs), and the tuned kernel matches an fp32 reference. TP2 tokens 2 and 4 passed the tuner's gate but improved only 1.4-3.0% when re-measured, so they are left out. The five TP4 rows are unchanged and re-measure at 1.02x-2.12x. The table now uses the tuner's current schema: no _tag column and an explicit nt column. The TP4 rows carry nt=0, which is how they already ran, since a missing nt means non-temporal loads off. The untuned table carries the full fifteen-bucket ladder for all three TPs. Signed-off-by: Jin Tao <jintao12@amd.com>
…ed-MoE table The TP8 token-1 row (E=288, topk=8, inter_dim 256) was kept only because the default dispatch returned wrong output at that shape. The cause is a split-K bug in the CK blockscale stage-1 wrapper, fixed separately in "[Bugfix][CK] Keep two K tiles per split in blockscale MoE stage-1 split-K". With that fix the default is correct and takes 32.1 us, against 59.8 us for this row, so the shape is left to the default. Signed-off-by: Jin Tao <jintao12@amd.com>
80fa418 to
355a4c2
Compare
There was a problem hiding this comment.
🔵 Needs a closer look
The changes require final human review because they are too complex or risky for automated approval.
0 open findings
1 resolved since last review
🧠 Review effort: Lite
Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.
* [Config] configs: GLM-5.3 fused shared MoE rows for gfx942 GLM-5.3-Flash with shared-expert fusion dispatches the AITER fused MoE as E=289, topk=9. Add gfx942 / cu_num 304 rows for its TP2, TP4 and TP8 signatures (inter_dim 1024, 512 and 256): 17 tuned one-stage ASM rows, and the 45-row untuned ladder they were tuned from. At TP8 with one token the default dispatch returns wrong output; the tuned row selects a kernel whose output is correct. Signed-off-by: Jin Tao <jintao12@amd.com> * [Config] Drop the TP8 one-token row from the GLM-5.3 fused shared MoE table The TP8 token-1 row (E=289, topk=9, inter_dim 256) was kept only because the default dispatch returned wrong output at that shape. The cause is a split-K bug in the CK blockscale stage-1 wrapper, fixed separately in "[Bugfix][CK] Keep two K tiles per split in blockscale MoE stage-1 split-K". With that fix the default is correct and takes 34.9 us, against 59.8 us for this row, so the shape is left to the default. Signed-off-by: Jin Tao <jintao12@amd.com> * [Config] Drop the TP2 two-token row from the GLM-5.3 fused shared MoE table The TP2 token-2 row (E=289, topk=9, inter_dim 1024) beat the default dispatch by 3.3% and 4.8% in the two verification runs (72.3 us against 75.4 us), just over the 3% bar. The routed GLM-5.3-Flash table in #5500 dropped the same shape, with the same kernel, at 2.1% and 2.9%. Leave it to the default here as well, so the routed and fused shared tables cover the same buckets. Signed-off-by: Jin Tao <jintao12@amd.com> --------- Signed-off-by: Jin Tao <jintao12@amd.com>



Why
GLM-5.3-Flash routed experts dispatch AITER's block-FP8 fused-MoE operator with these gfx942 runtime signatures:
(4096,1024)at TP2,(4096,512)at TP4,(4096,256)at TP8(288,8)QuantType.per_1x128, G1U1,doweight_stage1=0The existing model-config set covers this dispatch only on gfx950 (#5599, TP4). On gfx942 every token bucket uses AITER's fallback kernels.
What this adds
Fifteen gfx942/cu304 rows in a dedicated GLM-5.3 table, all one-stage
64x256withblock_m64:M={1024,2048,4096,8192,16384}The matching untuned table records the complete 15-bucket power-of-two ladder for each TP (45 rows).
All buckets were tuned with the standard fused-MoE tuner over the ASM and CK-tile block-FP8 families. Only rows improving AITER's production operator by at least 3% were retained; the other thirty buckets are intentionally excluded.
At TP8 with one token the fallback currently returns wrong output. That is a split-K bug in the CK stage-1 wrapper, fixed in #5979. With the fix the fallback is correct and faster than the best tuned candidate (32.1 versus 59.8 µs), so this table leaves the shape to it.
#5953 adds the fused shared-expert counterpart (
E=289, topk=9); the dispatch keys are disjoint.Validation
err1=0.0%,run_1stage=1,kernelName2empty, andus2=0.M=1024/2048/4096/8192/16384, mean of two runs:M=1024passed the 3% gate in the original tuning run (1.03x) and is kept as first submitted.M=1-512at every TP. TP2M=2andM=4passed the tuner gate at +5% but re-measured at only +1.4-3.0%. TP4M=1-32measured 0.67-1.02x against the fallback in a dedicated run.test_csv_validation.pyandtest_config_shape_collision.py: 37 passed, 45 subtests passed.tuned_fmoetable on main, or in the open PRs that add fused-MoE rows ([CK] [FlyDSL] [CI] [Bugfix] Explicit gfx in shipped fused-MoE tuning CSVs #5315, [Config] [DSv4] Tuned decode moe kernels for mori-EP backend- #41119 #5821, [Config] [Perf] Retune GLM5 MXFP4 FMoE for gfx950 #5902, [Config] Add tuned a8w4 fused-MoE configs for DeepSeek V4.1-Flash (TP4 + EP, gfx950) #5904, [Config] config: GLM-5.3-Flash fused shared MoE rows for gfx942 #5953)._tag, explicitnt). A downstream table with the same keys, such as thetuned_fmoe_glm53flash_gfx942.csvin our ROCm vLLM image, is therefore flagged by the merge's duplicate check instead of being kept alongside; ours will be retired when this lands.swiglu_limit=10; measurements are in [Config] config: GLM-5.3-Flash fused shared MoE rows for gfx942 #5953.Hardware: MI325X (
gfx942, 304 CUs).Image:
amdsiloai/vllm:vllm-openai-rocm-glm5.3-flash-mi325-28092026-pr57161for the TP2/TP8 tuning and every measurement above; the TP4 rows were tuned inamdsiloai/vllm:vllm-openai-rocm-glm5.3-flash-mi325-08092026.Tuning base: upstream
ROCm/aitermainat2e6209429for TP2/TP8, and this branch atd2bb0945efor TP4.Submission Checklist
main).