Skip to content

[Config] configs: GLM-5.3 routed MoE rows for gfx950 - #5599

Merged
zufayu merged 3 commits into
ROCm:mainfrom
Raiden-Makoto:RM/glm53-fmoe-tuned-config
Sep 21, 2026
Merged

zufayu merged 3 commits into
ROCm:mainfrom
Raiden-Makoto:RM/glm53-fmoe-tuned-config

Conversation

@Raiden-Makoto

@Raiden-Makoto Raiden-Makoto commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Why

GLM-5.3-Flash routed experts dispatch AITER's one-stage block-FP8 fused-MoE operator with this TP4 runtime signature:

  • model/local dimensions: (4096,512)
  • experts/top-k: (288,8)
  • BF16 output with FP8 activations and weights
  • QuantType.per_1x128, G1U1, doweight_stage1=0

The single kernel performs the gate/up GEMM, SiLU(gate) * up, down GEMM, and top-8 routed-weight accumulation.

The existing model-config set does not cover this dispatch signature. Without exact entries, these shapes use AITER's fallback kernels instead of kernels tuned for the dispatched token buckets.

What this adds

Six gfx950/cu256 rows are added in a dedicated, architecture-specific GLM-5.3 table:

  • M={4,8}: one-stage 16x128
  • M={2048,4096,8192,16384}: one-stage 64x256

The matching untuned table records the complete 15-shape power-of-two ladder from M=1 through M=16384.

All fifteen buckets were tuned with the standard fused-MoE tuner over the compatible one-stage ASM block-FP8 family. Only rows improving AITER's production operator by at least 3% were retained. Nine candidates were neutral or slower and are intentionally excluded.

The gfx950-specific filenames avoid the add/add path conflict with #5500's shared gfx942/fnuz tables; both PRs can merge independently, and their full dispatch keys remain disjoint.

Validation

  • All six retained rows have err1=0.0%, run_1stage=1, kernelName2 empty, and us2=0.
  • Production operator validation passed for every retained shape.
  • M=4 improves 7.39-7.98% across two repeated runs.
  • M=8 improves 3.08-3.15% across two repeated runs.
  • M=2048/4096/8192/16384 improve 35.20% / 50.21% / 32.01% / 21.14% in the tuner comparison.
  • Nine rejected buckets remain on their existing fallback: M={1,2,16,32,64,128,256,512,1024}.
  • test_csv_validation.py and test_config_shape_collision.py: 32 passed, 37 subtests passed.
  • Every new dispatch key occurs exactly once across the complete current-main fused-MoE config set.
  • No open AITER PR contains any new dispatch key.
  • No changed file path overlaps open PR [Config] configs: GLM-5.3-Flash a8w8 blockscale fused-MoE configs for gfx942 #5500.

Hardware: MI355X (gfx950, 256 CUs).
Image: rocm/sgl-dev:v0.5.19-rocm724-mi35x-20260910.
Tuning branch base: upstream ROCm/aiter main at commit d9f0c0b2efb53c62c5d3f1dc623c08ea8d2656fb.

Submission Checklist

  • Looked over the ROCm contributing guidelines.
  • Targeting the repository default branch (main).
  • Included successful config-validation and production-operator results.
  • Commit includes the required DCO sign-off.
  • CI green for the updated head.

@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5599 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@Raiden-Makoto
Raiden-Makoto marked this pull request as ready for review September 16, 2026 18:23
@Raiden-Makoto
Raiden-Makoto requested a review from a team September 16, 2026 18:23
@zufayu
zufayu self-requested a review September 17, 2026 01:10
@zufayu

zufayu commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

ROCm/aiter PR #5599 — [Config] configs: GLM-5.3 routed MoE row for gfx950

Adds a gfx950 tuned-config row (with its untuned tuner-input table) that routes the GLM-5.3-Flash fused-MoE dispatch key — padded M=16384, model/inter dim 4096/512, 288 experts, top-8, fp8 per-1x128 blockscale, bf16 output — to the 64x256 one-stage ASM kernel instead of the 32x256 fallback.

Review (advisory): ⚠️ NEEDS WORK
Validation (deterministic): NOT RUN — triage required a target but the PR ships none ("the PR changes runtime code but ships no test target"); the changed paths are a tuned-config table plus its untuned tuner input (data-only — nothing a test target could cover), so the missing target is not a defect on this diff
Perf (advisory): NOT RUN — the PR ships no benchmark entry point (data-only config row; the tuner run that produced it is the perf evidence), and it cannot be re-measured on this box (gfx942 only, the row is gfx950-only)

⚠️ [verified] Open PR #5500 (jin-amd, 2026-09-14, "Add GLM-5.3-Flash a8w8 blockscale fused-MoE configs for gfx942") adds the exact same two file paths — aiter/configs/model_configs/a8w8_blockscale_tuned_fmoe_glm5_3_flash.csv and its untuned sibling — holding gfx942/fnuz rows for the same (4096, 512, 288, 8) shape, so whichever PR merges second hits an add/add conflict and its new-file diff no longer applies; the "no open AITER PR contains the new dispatch key" check passes only because the dispatch key is gfx-scoped (gfx950/fn vs gfx942/fnuz) while the file path is not, and the two untuned tables also disagree (a 16-row token ladder vs this PR's single 16384 row). Author must coordinate with #5500 and rebase onto whichever lands first, merging both archs' rows into one shared table (the 14-column lookup key keeps them distinct — no collision, verified).
📝 [verified] The gfx950 table retains only the padded-16384 bucket while the gfx942 tuning of the same model in open PR #5500 retains tokens 1024-16384, so on gfx950 every prefill token count at or below 8192 keeps the 32x256 fallback, and the PR does not say whether those buckets were tuned and lost or never swept. Reviewer should ask whether the smaller gfx950 buckets were swept by the tuner and lost to the fallback, or never swept at all.

@Raiden-Makoto
Raiden-Makoto force-pushed the RM/glm53-fmoe-tuned-config branch from c0c1fb5 to a895adb Compare September 17, 2026 15:09
Signed-off-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
@Raiden-Makoto
Raiden-Makoto force-pushed the RM/glm53-fmoe-tuned-config branch from a895adb to 52f0d6b Compare September 17, 2026 15:11
@Raiden-Makoto Raiden-Makoto changed the title [Config] configs: GLM-5.3 routed MoE row for gfx950 [Config] configs: GLM-5.3 routed MoE rows for gfx950 Sep 17, 2026
@Raiden-Makoto
Raiden-Makoto marked this pull request as draft September 17, 2026 15:13
@Raiden-Makoto

Raiden-Makoto commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

Addressed both actionable findings in head 52f0d6bb6:

  • Renamed the tuned and untuned tables with a _gfx950 suffix. [Config] configs: GLM-5.3-Flash a8w8 blockscale fused-MoE configs for gfx942 #5500 and this PR now have no overlapping paths and can merge independently; their full dispatch keys remain disjoint.
  • Swept the complete 15-bucket gfx950 power-of-two ladder from M=1 through M=16384 using the same 3% retention gate. Six rows remain: M={4,8,2048,4096,8192,16384}; the other nine are recorded in the untuned table and retain fallback dispatch. M=4 and M=8 were rerun twice and remained above the gate.

@Raiden-Makoto
Raiden-Makoto marked this pull request as ready for review September 17, 2026 15:29
@zufayu
zufayu merged commit 6b35e2b into ROCm:main Sep 21, 2026
91 of 93 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants