Repository navigation
[Config] config: GLM-5.3-Flash fused shared MoE rows for gfx942 - #5953
Conversation
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
One backend per PR: PR title tags & labels: |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The selected one-stage ASM path does not receive the model’s required swiglu_limit=10, so clamped-SiLU semantics are not preserved.
Review effort: Balanced
Findings: 1
Open (1)
What changed in this PR
Adds gfx942 fused shared-expert MoE configurations for GLM-5.3-Flash.
Changes:
- Adds 17 tuned TP2/TP4/TP8 dispatch rows.
- Adds the complete 45-shape tuning input ladder.
| File | Description |
|---|---|
a8w8_blockscale_untuned_fmoe_glm5_3_flash_shared_gfx942.csv |
Defines tuning shapes. |
a8w8_blockscale_tuned_fmoe_glm5_3_flash_shared_gfx942.csv |
Adds selected gfx942 kernels. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
GLM-5.3-Flash with shared-expert fusion dispatches the AITER fused MoE as E=289, topk=9. Add gfx942 / cu_num 304 rows for its TP2, TP4 and TP8 signatures (inter_dim 1024, 512 and 256): 17 tuned one-stage ASM rows, and the 45-row untuned ladder they were tuned from. At TP8 with one token the default dispatch returns wrong output; the tuned row selects a kernel whose output is correct. Signed-off-by: Jin Tao <jintao12@amd.com>
… table The TP8 token-1 row (E=289, topk=9, inter_dim 256) was kept only because the default dispatch returned wrong output at that shape. The cause is a split-K bug in the CK blockscale stage-1 wrapper, fixed separately in "[Bugfix][CK] Keep two K tiles per split in blockscale MoE stage-1 split-K". With that fix the default is correct and takes 34.9 us, against 59.8 us for this row, so the shape is left to the default. Signed-off-by: Jin Tao <jintao12@amd.com>
3175b1d to
2947071
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The PR description documents an unrelated split-K code fix and provides no applicable validation for these configuration rows.
Review effort: Balanced
Findings: 1
Open (1)
Resolved since last review (1)
… table The TP2 token-2 row (E=289, topk=9, inter_dim 1024) beat the default dispatch by 3.3% and 4.8% in the two verification runs (72.3 us against 75.4 us), just over the 3% bar. The routed GLM-5.3-Flash table in ROCm#5500 dropped the same shape, with the same kernel, at 2.1% and 2.9%. Leave it to the default here as well, so the routed and fused shared tables cover the same buckets. Signed-off-by: Jin Tao <jintao12@amd.com>
AITER has fused-MoE configs tuned for GLM-5.3-Flash's fused shared-expert shape (E=289, topk=9) only on gfx950 (ROCm/aiter#5735, in AITER v0.1.23). The gfx942 rows are still in review (ROCm/aiter#5953), so on gfx942 the fused MoE would run untuned fallback kernels. Keep the shared experts as a separate MLP on every other GPU, and warn once when VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS is set there. gfx942 follows once AITER ships its rows. Signed-off-by: Jin Tao <jin.tao@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com>
AITER has fused-MoE configs tuned for GLM-5.3-Flash's fused shared-expert shape (E=289, topk=9) only on gfx950 (ROCm/aiter#5735, in AITER v0.1.23). The gfx942 rows are still in review (ROCm/aiter#5953), so on gfx942 the fused MoE would run untuned fallback kernels. Keep the shared experts as a separate MLP on every other GPU, and warn once when VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS is set there. gfx942 follows once AITER ships its rows. Signed-off-by: Jin Tao <jin.tao@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com>
ROCm/aiter#5953 adds AITER fused-MoE configs tuned for GLM-5.3-Flash's fused shared-expert shape (E=289, topk=9) on gfx942 with 304 CUs (MI300X, MI325X) at TP2, TP4 and TP8. Let VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS=1 fuse the shared experts there too, with the same data, prefill context and expert parallelism limits as on gfx950. Other gfx942 parts and other TP sizes keep the separate shared-expert MLP. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Jin Tao <jin.tao@amd.com>


Why
GLM-5.3-Flash shared-expert fusion changes the routed fused-MoE dispatch from
E=288, topk=8toE=289, topk=9. The ninth expert is always selected with weight 1.0, so one AITER launch runs the routed and shared experts together. The vLLM side is vllm-project/vllm#59221.The gfx942 TP-local signatures are:
(4096,1024),E=289,topk=9(4096,512),E=289,topk=9(4096,256),E=289,topk=9The existing model-config set covers this dispatch only on gfx950 (#5735). On gfx942 every signature uses fallback kernels, up to 2.1x slower than the rows below.
What this adds
Fifteen gfx942/cu304 rows in a dedicated GLM-5.3 fused-shared table, all one-stage
64x256withblock_m64:M={1024,2048,4096,8192,16384}The matching untuned table records the complete 15-shape power-of-two ladder for each TP (45 source rows). These are the same buckets as #5500, the routed (
E=288, topk=8) counterpart on gfx942; the dispatch keys are disjoint.All three signatures were tuned with the standard fused-MoE tuner over the ASM block-FP8 family. Only rows improving AITER's production operator by at least 3% were retained; the other thirty buckets keep fallback dispatch. Two of them are worth noting:
M=1: the fallback returned wrong output until [CK] [Bugfix] Keep two K tiles per split in blockscale MoE stage-1 split-K #5979 (a CK stage-1 split-K bug), which is in this branch's base. With the fix it is correct and faster than the best tuned candidate (34.9 versus 59.8 µs).M=2: the best candidate gained only 3.3% and 4.8% in the two verification runs, about 3 µs. [Config] configs: GLM-5.3-Flash a8w8 blockscale fused-MoE configs for gfx942 #5500 dropped the same shape at 2.1-2.9%, so it stays on the fallback here too.Validation
err1=0.0%,run_1stage=1,kernelName2empty, andus2=0.M=1024/2048/4096/8192/16384, mean of two runs:fused_moe(separate: the routed call plus the shared expert alone) at TP2, TP4 and TP8, from M=2 through M=16384: all outputs finite, minimum cosine 0.999867, maximum mean absolute error 8.99e-3 against mean absolute outputs of 0.27-0.64.test_csv_validation.pyandtest_config_shape_collision.py: 39 passed, 48 subtests passed. All 15 rows survive theAITER_CONFIG_FMOE_FILEmerge.swiglu_limit=10. Withw1scaled so the clamp changes a third of the gate/up values, this table, the two-stage fallback and the one-stage fallback for the unfused shape all match the unclamped fp32 reference (cosine 0.9990-0.9991), not the clamped one (0.84-0.87). Fixing that needs clamp support in the gfx942 kernels, separate from this table.Hardware: MI325X (
gfx942, 304 CUs).Image:
amdsiloai/vllm:vllm-openai-rocm-glm5.3-flash-mi325-28092026-pr57161; the TP2M=2andswiglu_limitre-checks ran in an image built on it.Tuning branch base: upstream
ROCm/aitermainat475cf0f60. The branch is now ona1455f1c4, whose fused-MoE changes do not touch these rows.Submission Checklist
main).