Repository navigation
[Config] configs: GLM-5.3 routed MoE rows for gfx950 - #5599
Conversation
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
ROCm/aiter PR #5599 — [Config] configs: GLM-5.3 routed MoE row for gfx950Adds a gfx950 tuned-config row (with its untuned tuner-input table) that routes the GLM-5.3-Flash fused-MoE dispatch key — padded M=16384, model/inter dim 4096/512, 288 experts, top-8, fp8 per-1x128 blockscale, bf16 output — to the 64x256 one-stage ASM kernel instead of the 32x256 fallback. Review (advisory):
|
c0c1fb5 to
a895adb
Compare
Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
a895adb to
52f0d6b
Compare
|
Addressed both actionable findings in head
|
Why
GLM-5.3-Flash routed experts dispatch AITER's one-stage block-FP8 fused-MoE operator with this TP4 runtime signature:
(4096,512)(288,8)QuantType.per_1x128, G1U1,doweight_stage1=0The single kernel performs the gate/up GEMM,
SiLU(gate) * up, down GEMM, and top-8 routed-weight accumulation.The existing model-config set does not cover this dispatch signature. Without exact entries, these shapes use AITER's fallback kernels instead of kernels tuned for the dispatched token buckets.
What this adds
Six gfx950/cu256 rows are added in a dedicated, architecture-specific GLM-5.3 table:
M={4,8}: one-stage16x128M={2048,4096,8192,16384}: one-stage64x256The matching untuned table records the complete 15-shape power-of-two ladder from
M=1throughM=16384.All fifteen buckets were tuned with the standard fused-MoE tuner over the compatible one-stage ASM block-FP8 family. Only rows improving AITER's production operator by at least 3% were retained. Nine candidates were neutral or slower and are intentionally excluded.
The gfx950-specific filenames avoid the add/add path conflict with #5500's shared gfx942/fnuz tables; both PRs can merge independently, and their full dispatch keys remain disjoint.
Validation
err1=0.0%,run_1stage=1,kernelName2empty, andus2=0.M=4improves 7.39-7.98% across two repeated runs.M=8improves 3.08-3.15% across two repeated runs.M=2048/4096/8192/16384improve 35.20% / 50.21% / 32.01% / 21.14% in the tuner comparison.M={1,2,16,32,64,128,256,512,1024}.test_csv_validation.pyandtest_config_shape_collision.py: 32 passed, 37 subtests passed.Hardware: MI355X (
gfx950, 256 CUs).Image:
rocm/sgl-dev:v0.5.19-rocm724-mi35x-20260910.Tuning branch base: upstream
ROCm/aitermainat commitd9f0c0b2efb53c62c5d3f1dc623c08ea8d2656fb.Submission Checklist
main).