Skip to content

[Config] [Perf] Retune GLM5 MXFP4 FMoE for gfx950 - #5902

Draft
XiaobingSuper wants to merge 1 commit into
ROCm:mainfrom
XiaobingSuper:config/glm5-fp4-fmoe-retune-gfx950
Draft

XiaobingSuper wants to merge 1 commit into
ROCm:mainfrom
XiaobingSuper:config/glm5-fp4-fmoe-retune-gfx950

Conversation

@XiaobingSuper

@XiaobingSuper XiaobingSuper commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • retune all 64 GLM5 A4W4 FMoE shapes on gfx950 with the merged MXFP4 FlyDSL path
  • update 21 rows that remain at least 2% faster under production Graph replay
  • preserve the other 43 incumbent rows byte-for-byte

Tuning and selection

  • 8x MI355X (gfx950, 256 CU/GPU), one shape per worker/GPU
  • Mxfp4FlydslTuner: 5 warmups and 101 GPU-event iterations per coupled stage1/stage2 candidate
  • production incumbent-vs-candidate replay with a 2% update threshold
  • three balanced-routing HIP Graph rounds over all 64 shapes (300 MoE calls per graph)
  • low-M random/hotspot replay; the route-regressing M=64, I=2048 candidate was rejected

Performance

  • accepted-row geomean speedup: 1.077x
  • projected geomean across all 64 shapes: 1.025x
  • final all-shape Graph sanity replay: 1.016x
  • the M=25 decode bucket (M=32, I=256/512) is unchanged
inter_dim Updated rows Updated-row geomean All-16-shape projected geomean
256 6 1.060x 1.022x
512 4 1.054x 1.013x
1024 3 1.112x 1.020x
2048 8 1.088x 1.043x

Accepted rows

M inter_dim Incumbent us Tuned us Speedup
2 256 35.266 31.590 1.116x
256 256 138.201 132.211 1.045x
1024 256 181.332 173.039 1.048x
2048 256 277.561 263.216 1.054x
4096 256 374.033 361.133 1.036x
32768 256 2292.030 2157.160 1.063x
256 512 235.802 222.697 1.059x
512 512 249.477 243.600 1.024x
1024 512 297.111 282.400 1.052x
4096 512 580.359 536.329 1.082x
4 1024 88.590 74.830 1.184x
1024 1024 497.433 469.473 1.060x
32768 1024 5254.530 4791.180 1.097x
1 2048 47.167 45.819 1.029x
2 2048 96.403 76.191 1.265x
16 2048 476.406 447.295 1.065x
128 2048 829.480 804.122 1.032x
2048 2048 1188.490 1101.310 1.079x
4096 2048 1730.600 1646.990 1.051x
8192 2048 2760.690 2608.290 1.058x
32768 2048 9878.620 8630.950 1.145x

Validation

  • all 64 tuner shapes completed and passed the configured cosine-error gate
  • no NaN or runtime failure in final production Graph replay
  • pytest -q op_tests/tuning_tests/test_csv_validation.py op_tests/tuning_tests/test_config_shape_collision.py: 37 passed, 45 subtests passed
  • unique config keys and git diff --check

Only aiter/configs/model_configs/glm5_fp4_tuned_fmoe.csv is changed.

@XiaobingSuper
XiaobingSuper requested review from a team and a lite review from Copilot September 28, 2026 06:02

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5902 --add-label <label>

One backend per PR:
A PR changes one kernel backend: [Triton/Gluon] (Triton and Gluon count as one), [HIP], [ASM], [CK], [OPUS] or [FlyDSL]. If the title ends up with two backend tags, split the PR -- as stacked pull requests when one part cannot merge without the other.

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to stop the title rewrites; labels stay in sync either way.

@XiaobingSuper
XiaobingSuper marked this pull request as draft September 28, 2026 07:23
@XiaobingSuper
XiaobingSuper force-pushed the config/glm5-fp4-fmoe-retune-gfx950 branch from 7d1c270 to 75336f5 Compare September 28, 2026 07:47
@XiaobingSuper
XiaobingSuper marked this pull request as ready for review September 28, 2026 07:47
Copilot AI review requested due to automatic review settings September 28, 2026 07:47

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@XiaobingSuper
XiaobingSuper marked this pull request as draft September 28, 2026 07:51
@XiaobingSuper
XiaobingSuper force-pushed the config/glm5-fp4-fmoe-retune-gfx950 branch from 75336f5 to 8aebcaf Compare September 28, 2026 07:52
@XiaobingSuper
XiaobingSuper marked this pull request as ready for review September 28, 2026 07:52
Copilot AI review requested due to automatic review settings September 28, 2026 07:52

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@XiaobingSuper
XiaobingSuper marked this pull request as draft September 28, 2026 07:56
Co-authored-by: Cursor <cursoragent@cursor.com>
@XiaobingSuper
XiaobingSuper force-pushed the config/glm5-fp4-fmoe-retune-gfx950 branch from 8aebcaf to fbc4461 Compare September 28, 2026 07:56
@XiaobingSuper
XiaobingSuper marked this pull request as ready for review September 28, 2026 07:56
Copilot AI review requested due to automatic review settings September 28, 2026 07:56

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants