Skip to content

[Config] config: GLM-5.3-Flash fused shared MoE rows for gfx942 - #5953

Merged
zufayu merged 3 commits into
ROCm:mainfrom
jin-amd:glm53-flash-fused-shared-fmoe-gfx942
Oct 10, 2026
Merged

zufayu merged 3 commits into
ROCm:mainfrom
jin-amd:glm53-flash-fused-shared-fmoe-gfx942

Conversation

@jin-amd

@jin-amd jin-amd commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Why

GLM-5.3-Flash shared-expert fusion changes the routed fused-MoE dispatch from E=288, topk=8 to E=289, topk=9. The ninth expert is always selected with weight 1.0, so one AITER launch runs the routed and shared experts together. The vLLM side is vllm-project/vllm#59221.

The gfx942 TP-local signatures are:

  • TP2: model/intermediate (4096,1024), E=289, topk=9
  • TP4: model/intermediate (4096,512), E=289, topk=9
  • TP8: model/intermediate (4096,256), E=289, topk=9

The existing model-config set covers this dispatch only on gfx950 (#5735). On gfx942 every signature uses fallback kernels, up to 2.1x slower than the rows below.

What this adds

Fifteen gfx942/cu304 rows in a dedicated GLM-5.3 fused-shared table, all one-stage 64x256 with block_m 64:

  • TP2, TP4 and TP8: M={1024,2048,4096,8192,16384}

The matching untuned table records the complete 15-shape power-of-two ladder for each TP (45 source rows). These are the same buckets as #5500, the routed (E=288, topk=8) counterpart on gfx942; the dispatch keys are disjoint.

All three signatures were tuned with the standard fused-MoE tuner over the ASM block-FP8 family. Only rows improving AITER's production operator by at least 3% were retained; the other thirty buckets keep fallback dispatch. Two of them are worth noting:

Validation

Hardware: MI325X (gfx942, 304 CUs).
Image: amdsiloai/vllm:vllm-openai-rocm-glm5.3-flash-mi325-28092026-pr57161; the TP2 M=2 and swiglu_limit re-checks ran in an image built on it.
Tuning branch base: upstream ROCm/aiter main at 475cf0f60. The branch is now on a1455f1c4, whose fused-MoE changes do not touch these rows.

Submission Checklist

  • Looked over the ROCm contributing guidelines.
  • Targeting the repository default branch (main).
  • Included successful config-validation and production-operator results.
  • Commit includes the required DCO sign-off.
  • CI green for the updated head.

@jin-amd
jin-amd requested review from a team and a balanced review from Copilot September 29, 2026 18:13
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5953 --add-label <label>

One backend per PR:
A PR changes one kernel backend: [Triton/Gluon] (Triton and Gluon count as one), [HIP], [ASM], [CK], [OPUS] or [FlyDSL]. If the title ends up with two backend tags, split the PR -- as stacked pull requests when one part cannot merge without the other.

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to stop the title rewrites; labels stay in sync either way.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The selected one-stage ASM path does not receive the model’s required swiglu_limit=10, so clamped-SiLU semantics are not preserved.

Review effort: Balanced
Findings: 1 High severity

Open (1)
What changed in this PR

Adds gfx942 fused shared-expert MoE configurations for GLM-5.3-Flash.

Changes:

  • Adds 17 tuned TP2/TP4/TP8 dispatch rows.
  • Adds the complete 45-shape tuning input ladder.
File Description
a8w8_blockscale_untuned_fmoe_glm5_3_flash_shared_gfx942.csv Defines tuning shapes.
a8w8_blockscale_tuned_fmoe_glm5_3_flash_shared_gfx942.csv Adds selected gfx942 kernels.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@zufayu
zufayu self-requested a review September 30, 2026 00:33
@zufayu zufayu self-assigned this Sep 30, 2026
Copilot AI balanced review requested due to automatic review settings September 30, 2026 05:25

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The missing TP8 M=1 tuned row leaves the reported incorrect fallback active.

Review effort: Balanced
Findings: 1 High severity

Open (1)
Resolved since last review (1)

@jin-amd jin-amd changed the title [Config] GLM-5.3-Flash fused shared MoE rows for gfx942 [Config] config: GLM-5.3-Flash fused shared MoE rows for gfx942 Sep 30, 2026
GLM-5.3-Flash with shared-expert fusion dispatches the AITER fused MoE as
E=289, topk=9. Add gfx942 / cu_num 304 rows for its TP2, TP4 and TP8
signatures (inter_dim 1024, 512 and 256): 17 tuned one-stage ASM rows, and
the 45-row untuned ladder they were tuned from.

At TP8 with one token the default dispatch returns wrong output; the tuned
row selects a kernel whose output is correct.

Signed-off-by: Jin Tao <jintao12@amd.com>
… table

The TP8 token-1 row (E=289, topk=9, inter_dim 256) was kept only because
the default dispatch returned wrong output at that shape. The cause is a
split-K bug in the CK blockscale stage-1 wrapper, fixed separately in
"[Bugfix][CK] Keep two K tiles per split in blockscale MoE stage-1
split-K". With that fix the default is correct and takes 34.9 us, against
59.8 us for this row, so the shape is left to the default.

Signed-off-by: Jin Tao <jintao12@amd.com>
Copilot AI balanced review requested due to automatic review settings October 2, 2026 08:36
@jin-amd
jin-amd force-pushed the glm53-flash-fused-shared-fmoe-gfx942 branch from 3175b1d to 2947071 Compare October 2, 2026 08:36

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The PR description documents an unrelated split-K code fix and provides no applicable validation for these configuration rows.

Review effort: Balanced
Findings: 1 Low severity

Open (1)
Resolved since last review (1)

… table

The TP2 token-2 row (E=289, topk=9, inter_dim 1024) beat the default
dispatch by 3.3% and 4.8% in the two verification runs (72.3 us against
75.4 us), just over the 3% bar. The routed GLM-5.3-Flash table in ROCm#5500
dropped the same shape, with the same kernel, at 2.1% and 2.9%. Leave it
to the default here as well, so the routed and fused shared tables cover
the same buckets.

Signed-off-by: Jin Tao <jintao12@amd.com>
Copilot AI balanced review requested due to automatic review settings October 2, 2026 12:22

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The description documents an unrelated split-K implementation rather than the configuration changes and their validation.

Review effort: Balanced
Findings: 1 Low severity

Open (1)

jin-amd added a commit to jin-amd/vllm that referenced this pull request Oct 2, 2026
AITER has fused-MoE configs tuned for GLM-5.3-Flash's fused shared-expert
shape (E=289, topk=9) only on gfx950 (ROCm/aiter#5735, in AITER v0.1.23).
The gfx942 rows are still in review (ROCm/aiter#5953), so on gfx942 the
fused MoE would run untuned fallback kernels.

Keep the shared experts as a separate MLP on every other GPU, and warn
once when VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS is set there. gfx942
follows once AITER ships its rows.

Signed-off-by: Jin Tao <jin.tao@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
jin-amd added a commit to jin-amd/vllm that referenced this pull request Oct 5, 2026
AITER has fused-MoE configs tuned for GLM-5.3-Flash's fused shared-expert
shape (E=289, topk=9) only on gfx950 (ROCm/aiter#5735, in AITER v0.1.23).
The gfx942 rows are still in review (ROCm/aiter#5953), so on gfx942 the
fused MoE would run untuned fallback kernels.

Keep the shared experts as a separate MLP on every other GPU, and warn
once when VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS is set there. gfx942
follows once AITER ships its rows.

Signed-off-by: Jin Tao <jin.tao@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
jin-amd added a commit to jin-amd/vllm that referenced this pull request Oct 8, 2026
ROCm/aiter#5953 adds AITER fused-MoE configs tuned for GLM-5.3-Flash's
fused shared-expert shape (E=289, topk=9) on gfx942 with 304 CUs (MI300X,
MI325X) at TP2, TP4 and TP8. Let VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS=1
fuse the shared experts there too, with the same data, prefill context and
expert parallelism limits as on gfx950. Other gfx942 parts and other TP sizes
keep the separate shared-expert MLP.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Jin Tao <jin.tao@amd.com>

@samremes samremes left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@zufayu
zufayu merged commit 9ec2a5b into ROCm:main Oct 10, 2026
62 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants