Skip to content

[Config] configs: GLM-5.3-Flash a8w8 blockscale fused-MoE configs for gfx942 - #5500

Open
jin-amd wants to merge 3 commits into
ROCm:mainfrom
jin-amd:glm5.3-flash-tuned-fmoe
Open

jin-amd wants to merge 3 commits into
ROCm:mainfrom
jin-amd:glm5.3-flash-tuned-fmoe

Conversation

@jin-amd

@jin-amd jin-amd commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Why

GLM-5.3-Flash routed experts dispatch AITER's block-FP8 fused-MoE operator with these gfx942 runtime signatures:

  • model/local dimensions: (4096,1024) at TP2, (4096,512) at TP4, (4096,256) at TP8
  • experts/top-k: (288,8)
  • BF16 output with FP8 (e4m3fnuz) activations and weights
  • QuantType.per_1x128, G1U1, doweight_stage1=0

The existing model-config set covers this dispatch only on gfx950 (#5599, TP4). On gfx942 every token bucket uses AITER's fallback kernels.

What this adds

Fifteen gfx942/cu304 rows in a dedicated GLM-5.3 table, all one-stage 64x256 with block_m 64:

  • TP2, TP4 and TP8: M={1024,2048,4096,8192,16384}

The matching untuned table records the complete 15-bucket power-of-two ladder for each TP (45 rows).

All buckets were tuned with the standard fused-MoE tuner over the ASM and CK-tile block-FP8 families. Only rows improving AITER's production operator by at least 3% were retained; the other thirty buckets are intentionally excluded.

At TP8 with one token the fallback currently returns wrong output. That is a split-K bug in the CK stage-1 wrapper, fixed in #5979. With the fix the fallback is correct and faster than the best tuned candidate (32.1 versus 59.8 µs), so this table leaves the shape to it.

#5953 adds the fused shared-expert counterpart (E=289, topk=9); the dispatch keys are disjoint.

Validation

Hardware: MI325X (gfx942, 304 CUs).
Image: amdsiloai/vllm:vllm-openai-rocm-glm5.3-flash-mi325-28092026-pr57161 for the TP2/TP8 tuning and every measurement above; the TP4 rows were tuned in amdsiloai/vllm:vllm-openai-rocm-glm5.3-flash-mi325-08092026.
Tuning base: upstream ROCm/aiter main at 2e6209429 for TP2/TP8, and this branch at d2bb0945e for TP4.

Submission Checklist

  • Looked over the ROCm contributing guidelines.
  • Targeting the repository default branch (main).
  • Included successful config-validation and production-operator results.
  • CI green for the updated head.

@jin-amd
jin-amd requested review from a team and a lite review from Copilot September 14, 2026 11:36
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5500 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Tuning provenance remains unresolved, warranting final human review.

Pull request overview

Adds GLM-5.3-Flash fused-MoE tuning configurations for gfx942.

Changes:

  • Adds fifteen untuned token-bucket inputs.
  • Adds five tuned single-stage runtime configurations.

Review note (nit): Add benchmark command, ROCm/GPU environment, and checkpoint provenance for reproducibility.

File summaries
File Description
aiter/configs/model_configs/a8w8_blockscale_untuned_fmoe_glm5_3_flash.csv Complete tuner input shape ladder.
aiter/configs/model_configs/a8w8_blockscale_tuned_fmoe_glm5_3_flash.csv Tuned gfx942 runtime configurations.
Review details

Suppressed comments (1)

aiter/configs/model_configs/a8w8_blockscale_tuned_fmoe_glm5_3_flash.csv:1

  • [verified] The tuned rows include measured timings, but the PR description does not record the exact benchmark entry point/command, ROCm version, GPU model, or checkpoint. That makes it difficult to reproduce the default-vs-tuned gate when this table is revisited or regenerated. Author must add the command and environment/checkpoint provenance to the PR or a checked-in tuning record.
gfx,cu_num,token,model_dim,inter_dim,expert,topk,act_type,dtype,q_dtype_a,q_dtype_w,q_type,use_g1u1,doweight_stage1,block_m,ksplit,us1,kernelName1,err1,us2,kernelName2,err2,us,run_1stage,xbf16,flat,tflops,bw,_tag
  • Files reviewed: 2/2 changed files
  • Comments generated: 0
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@zufayu

zufayu commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Advisory review (static + hand-run; not a merge gate). Validation/Perf ran no GPU stage — reasons are on their lines below. Findings tagged [verified] are traced to code/repro.

Adds two gfx942 data files for the GLM-5.3-Flash routed fused-MoE shape (model_dim 4096, inter_dim 512, 288 experts, topk 8, fp8 e4m3fnuz per-1x128 blockscale): a five-row tuned table that selects single-stage 64x256 asm kernels at token buckets 1024-16384, plus the fifteen-bucket tuner-input ladder it was tuned from.

Review (advisory): ⚠️ NEEDS WORK
Validation (deterministic): NOT RUN — data-only config PR: triage found no test target (target=null; the diff adds only a tuned-config table and a tuner-input CSV under aiter/configs/model_configs/), so there is nothing a test target could have covered
Perf (advisory): NOT RUN — no benchmark target ships with a data-only CSV PR; the PR's own default-vs-tuned µs tables (20 token buckets, two runs agreeing within 1%) are author-measured on one box and not independently reproduced

⚠️ [verified] The tuning provenance behind aiter/configs/model_configs/a8w8_blockscale_tuned_fmoe_glm5_3_flash.csv cannot be reconstructed: the PR names the tool flags (--compare, --update_improved, --run_config) and TP4/gfx942, but ships no exact command lines, no ROCm version, no GPU model, and no aiter commit for the runs that produced the five kept rows and the fifteen rejections. Anyone re-running the tuner later cannot tell whether a changed number is a real kernel regression or a different environment. Author must add the exact tune/compare commands and environment (ROCm version, GPU, driver, aiter commit) to the PR or a checked-in tuning record.

@jin-amd

jin-amd commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

Advisory review (static + hand-run; not a merge gate). Validation/Perf ran no GPU stage — reasons are on their lines below. Findings tagged [verified] are traced to code/repro.

Adds two gfx942 data files for the GLM-5.3-Flash routed fused-MoE shape (model_dim 4096, inter_dim 512, 288 experts, topk 8, fp8 e4m3fnuz per-1x128 blockscale): a five-row tuned table that selects single-stage 64x256 asm kernels at token buckets 1024-16384, plus the fifteen-bucket tuner-input ladder it was tuned from.

Review (advisory): ⚠️ NEEDS WORK Validation (deterministic): NOT RUN — data-only config PR: triage found no test target (target=null; the diff adds only a tuned-config table and a tuner-input CSV under aiter/configs/model_configs/), so there is nothing a test target could have covered Perf (advisory): NOT RUN — no benchmark target ships with a data-only CSV PR; the PR's own default-vs-tuned µs tables (20 token buckets, two runs agreeing within 1%) are author-measured on one box and not independently reproduced

⚠️ [verified] The tuning provenance behind aiter/configs/model_configs/a8w8_blockscale_tuned_fmoe_glm5_3_flash.csv cannot be reconstructed: the PR names the tool flags (--compare, --update_improved, --run_config) and TP4/gfx942, but ships no exact command lines, no ROCm version, no GPU model, and no aiter commit for the runs that produced the five kept rows and the fifteen rejections. Anyone re-running the tuner later cannot tell whether a changed number is a real kernel regression or a different environment. Author must add the exact tune/compare commands and environment (ROCm version, GPU, driver, aiter commit) to the PR or a checked-in tuning record.

Thanks @zufayu — provenance was the same ask in each, so it's now in the description: new Environment and Commands sections (exact tuner invocations, container, GPU/driver/ROCm, and the aiter commit for each run), plus a rewritten Verification.

On "not independently reproduced" — now confirmed by a different method. A paired vLLM serving sweep rather than the tuner: 131k in / 1k out, 7 concurrencies, 20 requests each, TP4, one image holding both tables with AITER_CONFIG_FMOE selecting per arm, control first — and the control arm reproduces the shipping image's kernel selection exactly. Removing the tokens 1–32 rows gives +1.17% output throughput winning 7/7 concurrencies, +3.58% at concurrency 2, +1.31% TPOT, GSM8K 0.9727 in both arms. On "one box": repeat runs of a fixed config here agree to ±0.48% per row over eight repeats, so the low-concurrency points are clear of variance. The gain also sits exactly where the dropped buckets were being selected and decays to ~0 above them — the shape the µs tables predict.

A correction to my own description. I claimed keys are disjoint from all 35 existing tuned_fmoe tables. True in this repo — but the glob takes any *tuned_fmoe*.csv, and our downstream nightly ships tuned_fmoe_glm53flash_gfx942.csv for this same shape, agreeing on all 14 shape keys at all five tokens. I expected the us tie-break to handle that; it doesn't. _tag is part of the dedup key, so a file without the column fills to 0 while this one carries "" — no duplicate is detected and both rows survive, with nothing deciding which the lookup returns (verified over the two tables: 10 rows, 5 shapes, 0 duplicates). So the rule for this shape is one table, not two, and I'll retire the downstream one when this lands.

On the NOT RUN validation stage — agreed that a data-only CSV gives a test target nothing to exercise. If it would be welcome, I'd follow up with a schema check over model_configs/: valid columns against the untuned sibling, no duplicate shape keys within a file, and no cross-file collisions on the keys the merge actually uses. It would have caught the above automatically.

Copilot AI review requested due to automatic review settings September 22, 2026 12:50

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The documented tuning paths do not match committed files, and downstream duplicate tables can cause order-dependent kernel selection.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 Medium severity · 1 Low severity

Open (2)

Comment thread aiter/configs/model_configs/a8w8_blockscale_tuned_fmoe_glm5_3_flash.csv Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Resolve the _tag deduplication conflict and correct the documented tuning paths.

Review effort: Lite
Findings: 1 Medium severity · 1 Low severity

Open (2)

Comment thread aiter/configs/model_configs/a8w8_blockscale_tuned_fmoe_glm5_3_flash.csv Outdated
Copilot AI review requested due to automatic review settings September 25, 2026 11:14
@jin-amd
jin-amd requested a review from samremes September 25, 2026 11:18

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

No unresolved review issues were identified.

Review effort: Lite
Findings: None

Resolved since last review (2)

Copilot AI review requested due to automatic review settings September 28, 2026 10:20
@jin-amd
jin-amd force-pushed the glm5.3-flash-tuned-fmoe branch from 9d6f3e7 to b1911bb Compare September 28, 2026 10:20

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Hardware-specific tuning and configuration-merge behavior warrant final human validation.

Review effort: Lite
Findings: None

Copilot AI lite review requested due to automatic review settings September 29, 2026 22:04

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Remove the regressive token-1 row and document or remove the unvalidated non-512 variants.

Review effort: Lite
Findings: 1 Medium severity

Open (1)

Comment thread aiter/configs/model_configs/a8w8_blockscale_tuned_fmoe_glm5_3_flash.csv Outdated
Copilot AI lite review requested due to automatic review settings September 30, 2026 05:25

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Add the missing TP8 token-1 tuned row to avoid the incorrect default dispatch.

Review effort: Lite
Findings: 1 High severity

Open (1)
Resolved since last review (1)

@jin-amd jin-amd changed the title [Config] [Tune] Add GLM-5.3-Flash a8w8 blockscale fused-MoE configs for gfx942 [Config] configs: GLM-5.3-Flash a8w8 blockscale fused-MoE configs for gfx942 Sep 30, 2026
Copilot AI lite review requested due to automatic review settings September 30, 2026 06:17
@jin-amd
jin-amd force-pushed the glm5.3-flash-tuned-fmoe branch from d1e2b2b to 80fa418 Compare September 30, 2026 06:17

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Add the missing TP8 token-1 tuned configuration row.

Review effort: Lite
Findings: 1 High severity

Open (1)

jin-amd added a commit to jin-amd/aiter that referenced this pull request Oct 2, 2026
… table

The TP2 token-2 row (E=289, topk=9, inter_dim 1024) beat the default
dispatch by 3.3% and 4.8% in the two verification runs (72.3 us against
75.4 us), just over the 3% bar. The routed GLM-5.3-Flash table in ROCm#5500
dropped the same shape, with the same kernel, at 2.1% and 2.9%. Leave it
to the default here as well, so the routed and fused shared tables cover
the same buckets.

Signed-off-by: Jin Tao <jintao12@amd.com>
jin-amd and others added 3 commits October 8, 2026 09:51
…or gfx942

Tune the routed fused-MoE shape GLM-5.3-Flash dispatches, on gfx942 / cu_num
304. Kernel-time attribution on a TP4 serving run put the routed MoE at 15.1%
of prefill forward time against 2.1% for the a8w8_blockscale GEMMs, and the
MoE was still on default kernels, so it is where the remaining headroom is.

One shape, tuned over aiter's power-of-two token ladder:

    model_dim 4096, inter_dim 512, expert 288, topk 8, Silu,
    bf16 activations, fp8 e4m3fnuz per_1x128 blockscale, g1u1,
    no doweight_stage1

inter_dim is 512 because moe_intermediate_size 2048 is sharded over TP4.
Tuning the bucket ladder covers every token count, since the lookup rounds to
these buckets.

Five of the fifteen buckets are kept. Improvement is end-to-end operator time
from --compare, tuned against the default kernel:

    token   default(us)   tuned(us)   speedup
     1024        542.18      524.53     1.03x
     2048       1002.53      670.24     1.50x
     4096       1864.42      868.68     2.15x
     8192       2812.08     1524.90     1.84x
    16384       4336.83     2764.15     1.57x

Time-weighted speedup over the default across these five is 1.66x (10.56 ms
-> 6.35 ms summed), median per-shape 1.57x. The gain concentrates at large
token counts, so it matters for prefill and large batched decode.

The other ten buckets are deliberately absent, because the default kernel is
equal or faster there and falling back is the better choice:

  - tokens 64, 128, 256 and 512 were rejected by the --update_improved gate on
    two separate tuning runs, at -0.57%/-0.20%, -9.75%/-5.26%, +1.23%/-0.07%
    and -0.80%/+0.33%.

  - tokens 1 through 32 were rejected after a dedicated measurement. The tuner
    had marked them "no_baseline" and admitted them on the post-run alone,
    because the pre-run benchmark returned no timing for them. Benchmarking the
    default path directly (--run_config with no CSV) against the tuned kernels
    shows the tuned candidates are slower at five of the six, reproduced over
    two runs agreeing within 1%:

        token   default(us)   tuned(us)   speedup
            1         38.70       58.02     0.67x
            2         51.55       61.83     0.83x
            4         73.88       74.33     0.99x
            8        113.85      111.26     1.02x
           16        173.23      179.94     0.96x
           32        255.15      295.34     0.87x

    Only token 8 improves at all, by 2.3%, under the 3% threshold. Shipping
    these would have regressed the small-batch decode path by up to 1.5x.

All five rows select block_m 64 and a 64x256 stage-1 kernel:
fmoe_bf16_blockscaleFp8_g1u1_vs_ps_silu_64x256 at tokens 1024, 4096, 8192 and
16384, and fmoe_bf16_blockscaleFp8_g1u1_vs_silu_64x256 at token 2048. Stage 2
is unused; every row runs single-stage.

The untuned sibling carries all fifteen buckets, not the five that were kept,
so it documents the search space the numbers above came from. Re-running the
tuner from it reproduces the same gate decisions.

Keys are disjoint from all 35 existing tuned_fmoe tables, so get_config_file()
merges this file with no duplicate-shape conflict. Verified by loading
AITER_CONFIG_FMOE_FILE with this file in model_configs/ and confirming all five
rows survive the merge. The untuned file is excluded by that glob, which skips
any name containing "untuned".

Co-authored-by: Cursor <cursoragent@cursor.com>
…d-MoE table

Extend the routed GLM-5.3-Flash fused-MoE table (E=288, topk=8) from TP4 to
TP2 (inter_dim 1024) and TP8 (inter_dim 256) on gfx942 / cu_num 304. Both
were tuned over the same asm and cktile families as the TP4 rows, with
--compare --update_improved.

New rows are kept when the production operator improves by at least 3% in
both of two default-against-tuned runs:

  TP2: tokens 1024-16384, 1.08x-1.86x
  TP8: tokens 1024-16384, 1.14x-2.37x

TP8 also keeps token 1. The default dispatch there returns wrong output (a
mismatch in both runs), and the tuned kernel matches an fp32 reference.
TP2 tokens 2 and 4 passed the tuner's gate but improved only 1.4-3.0% when
re-measured, so they are left out.

The five TP4 rows are unchanged and re-measure at 1.02x-2.12x.

The table now uses the tuner's current schema: no _tag column and an
explicit nt column. The TP4 rows carry nt=0, which is how they already ran,
since a missing nt means non-temporal loads off.

The untuned table carries the full fifteen-bucket ladder for all three TPs.

Signed-off-by: Jin Tao <jintao12@amd.com>
…ed-MoE table

The TP8 token-1 row (E=288, topk=8, inter_dim 256) was kept only because
the default dispatch returned wrong output at that shape. The cause is a
split-K bug in the CK blockscale stage-1 wrapper, fixed separately in
"[Bugfix][CK] Keep two K tiles per split in blockscale MoE stage-1
split-K". With that fix the default is correct and takes 32.1 us, against
59.8 us for this row, so the shape is left to the default.

Signed-off-by: Jin Tao <jintao12@amd.com>
@jin-amd
jin-amd force-pushed the glm5.3-flash-tuned-fmoe branch from 80fa418 to 355a4c2 Compare October 8, 2026 06:51
Copilot AI lite review requested due to automatic review settings October 8, 2026 06:51

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The changes require final human review because they are too complex or risky for automated approval.

0 open findings

1 resolved since last review

🧠 Review effort: Lite


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

zufayu pushed a commit that referenced this pull request Oct 10, 2026
* [Config] configs: GLM-5.3 fused shared MoE rows for gfx942

GLM-5.3-Flash with shared-expert fusion dispatches the AITER fused MoE as
E=289, topk=9. Add gfx942 / cu_num 304 rows for its TP2, TP4 and TP8
signatures (inter_dim 1024, 512 and 256): 17 tuned one-stage ASM rows, and
the 45-row untuned ladder they were tuned from.

At TP8 with one token the default dispatch returns wrong output; the tuned
row selects a kernel whose output is correct.

Signed-off-by: Jin Tao <jintao12@amd.com>

* [Config] Drop the TP8 one-token row from the GLM-5.3 fused shared MoE table

The TP8 token-1 row (E=289, topk=9, inter_dim 256) was kept only because
the default dispatch returned wrong output at that shape. The cause is a
split-K bug in the CK blockscale stage-1 wrapper, fixed separately in
"[Bugfix][CK] Keep two K tiles per split in blockscale MoE stage-1
split-K". With that fix the default is correct and takes 34.9 us, against
59.8 us for this row, so the shape is left to the default.

Signed-off-by: Jin Tao <jintao12@amd.com>

* [Config] Drop the TP2 two-token row from the GLM-5.3 fused shared MoE table

The TP2 token-2 row (E=289, topk=9, inter_dim 1024) beat the default
dispatch by 3.3% and 4.8% in the two verification runs (72.3 us against
75.4 us), just over the 3% bar. The routed GLM-5.3-Flash table in #5500
dropped the same shape, with the same kernel, at 2.1% and 2.9%. Leave it
to the default here as well, so the routed and fused shared tables cover
the same buckets.

Signed-off-by: Jin Tao <jintao12@amd.com>

---------

Signed-off-by: Jin Tao <jintao12@amd.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants