Skip to content

[FlyDSL] Add gfx950 FHMoE tuner on gemm_moe_tune.py - #5751

Draft
a-canadasruiz wants to merge 14 commits into
ROCm:mainfrom
a-canadasruiz:users/acanadas/fhmoe-tuner
Draft

a-canadasruiz wants to merge 14 commits into
ROCm:mainfrom
a-canadasruiz:users/acanadas/fhmoe-tuner

Conversation

@a-canadasruiz

@a-canadasruiz a-canadasruiz commented Sep 22, 2026 •

Copy link
Copy Markdown

Summary

gfx950 fused heterogeneous MoE (FHMoE, DSV4 I384) already serves through fused_moe + aiter/configs/tuned_fhmoe.csv. What was missing is a tuner: that sweeps FlyDSL stage1×stage2 pairs, rejects fast-wrong kernels against torch, and writes a native tuned CSV (block_m, kernelName1, kernelName2, us).

This PR adds that tuner as FhmoeTuner on the existing FMoE script (gemm_moe_tune.py --fhmoe), wires it into op_tune.sh and Tuning Tests like the other families, and does not replace the shipped 12-row tuned_fhmoe.csv.

FHMoE is its own family, not a flag on FMoE tests: different CLI (--fhmoe), different CSVs (untuned_fhmoe.csv / tuned_fhmoe.csv / AITER_CONFIG_FHMOE), gfx950-only, shared-expert + routed FP4 + shared FP8. Serving fused_moe / aiter/fhmoe.py are not in the diff.

Tuner (FhmoeTuner)

Subclass of FmoeTuner, selected with --fhmoe (mutually exclusive with --grouped-gemm and --mxfp4-flydsl on the same __main__).

Catalogue. aiter/configs/untuned_fhmoe.csv: 12 DSV4 I384 rows, tokens 1,2,4,…,2048. Same shape keys as shipped tuned_fhmoe.csv (that file has no us column; the producer adds us).

Arch gate. aiter/fhmoe.py is NotImplementedError off gfx950. gemm_moe_tune.py multiplexes four tuners in one file, so the gate is in __main__: --fhmoe on a non-gfx950 process SystemExits (--fhmoe is only supported on gfx950). Grouped GEMM already does the same for gfx1250. This is not A6W6’s RuntimeError inside tune(): A6W6 is its own script, so a unittest subprocess can still import and skip; here the process dies before FhmoeTuner runs. Changing FHMoE to RuntimeError would not remove the need to skip --fhmoe subprocesses in Tuning Tests, and it would be inconsistent with --grouped-gemm. The skip belongs in the tests (and in op_tune.sh), not by weakening __main__.

How a candidate is timed. Public fused_moe() cannot inject a one-row candidate. supports_dsv4_i384_fhmoe requires every padded-M tier in the config file (e.g. fused_moe(M=16) needs tuned rows for 1, 2, 4, 8, and 16). The worker therefore times _fused_moe_impl with a temporary one-row CSV (kernelName1 / kernelName2 / block_m). That is kernel-pair search, not the serving lookup.

Published token is the serving lookup key. Serving get_2stage_cfgs looks up get_padded_M(token_num), not the raw catalogue token (fused_moe.py; for M < 32768 that is nextPow2, so 3072 → 4096). Candidate timing already wrote that padded key in _write_candidate_csv. post_process used to emit the catalogue token, so a non-power-of-two untuned row would be timed at the padded tier and published under a key serving never reads: the gap-fill job can exit 0, the CSV grows a dead row, and the shape stays uncovered. The tuner now publishes get_padded_M(X) through _lookup_token, the same helper timing uses. Save the padded token only — not X, not both — because serving has a single token column.

The 12 shipped catalogue rows are already 1,2,4,…,2048; pad(X) = X for them and this does not change tuned_fhmoe.csv.

Local check on gfx950, two-pair smoke regex (not the cartesian), --mp 1, ~66s: catalogue token=3 → /tmp/tuned_fhmoe_token3.csv has token=4, us=47.3686, cosine OK (0 failed). BaseTuner.tune_summary still row-matches the -i frame (token=3) against success (token=4), so it prints leftover “untuned” and [Tuning not Finished] / exit 1. That is a key mismatch in the shared summary, not a failed kernel. We do not pad the catalogue on read and we do not rewrite BaseTuner’s message (every family uses it). Serving lookup is the contract this tuner has to match; the 12-row gap-fill job never hits the mismatch.

Search. For each catalogue row, _kernel_pairs builds the legal FlyDSL cartesian (block_m ∈ {16,32,64,128} × stage1 names × stage2 names, after tile/LDS/inter_dim filters). DSV4 I384 INTERLEAVE is on the order of 5632 pairs per token. Winner = min us among pairs with cosine vs torch ≤ errRatio. post_process drops the whole shape if every pair fails. Fast-wrong pairs must not win on latency.

Data. generate_fhmoe_data builds routed MXFP4 + shared FP8 tensors on the host (FMoE’s generate_data stays on CUDA). Dummy expert pad, gate_mode preshuffle. Weights are ~6 GiB per shape and token-independent. run_config regenerates per row sequentially and empty_caches; this is not a 12×6 GiB peak. Left as-is; not a test_run_config bug.

mp_tuner. Same worker pool as FMoE. shape_grouped=True puts every pair of one shape in one task group. Default --timeout is 1800s per group. A full cartesian for one token is hours, so a real tuner run needs --timeout on the order of a day (see Perf). Pipeline smoke is two pairs and stays under 1800s.
--warmup/--iters now go through as run_perftest kwargs on the FHMoE search, not only --run_config. Homogeneous FMoE still uses the 2/101 defaults; that family is out of scope here.

--run_config. Overrides the parent FMoE path. Iterates the tuned CSV, calls public fused_moe with shared-expert args, cosine vs torch, reports e2e_us. --run_config <csv> sets AITER_CONFIG_FHMOE to that file (same env serving uses) and clears FHMoE caches. Tuner us and e2e_us are different clocks (search vs production); do not compare them.

op_tests/test_fhmoe.py. __main__ now runs pytest.main. python3 op_tests/test_fhmoe.py used to exit 0 without collecting tests. That file is the op oracle, not the tuner.

op_tune.sh / Operators Tuning

Job, same gap-fill shape as FMoE:

-i aiter/configs/untuned_fhmoe.csv -o aiter/configs/tuned_fhmoe.csv

No --all. Without --all, FmoeTuner.pre_process drops keys already present in -o (“only kernels that are not in the tuned CSV”). The 12 catalogue rows already match shipped tuned_fhmoe.csv, so on gfx950 this job is a no-op until someone adds a new untuned shape. --all would retune those keys in CI and could overwrite production defaults. This PR is the tuner, not a retune of shipped rows.

Operators Tuning’s runner is linux-aiter-oci-mi300x-1 (gfx942). The script walks the full tune_jobs list, so --fhmoe would SystemExit and fail the workflow. The fhmoe job skips unless get_gfx() == gfx950. get_gfx() follows GPU_ARCHS (same helper as the rest of this script and aiter tests). Live rocminfo / get_gfx_runtime() is the other clock; we did not switch op_tune.sh to it.

Tuning Tests (family tables, not a one-off yaml)

tuning-tests.yaml is unchanged: test_csv_validation, test_tune_pipeline, test_run_config. GPU_ARCHS=gfx950, runner linux-aiter-mi35x-1. Triggers are schedule + workflow_dispatch only — not pull_request. Opening this PR does not launch that workflow.

Grouped GEMM is not a Tuning Tests family (gfx1250; this workflow is gfx950). FHMoE is a family, with an explicit gfx950 skip so --fhmoe SystemExit is not a FAIL if GPU_ARCHS is wrong.

test_csv_validation.py

  • TUNED_CSVS includes tuned_fhmoe.csv.
  • test_fhmoe_no_duplicates uses FHMoE extra keys (act_type, dtypes, shared_expert_id, pads, gate_mode, …), not GEMM M,N,K.
  • Untuned catalogue file must exist.

test_tune_pipeline.py

  • Smoke only: TUNE_MOE_KERNEL_REGEX pins two kn1×kn2 pairs, not the cartesian (timeout_mp1 1200s). Measured ~72s on gfx950.
  • test_fhmoe_mp1 only — no mp_default. One GPU is enough to prove the tuner writes a CSV; all-GPU cartesian smoke is not the family pattern we need.
  • Skip is an if not _is_gfx950(): skipTest inside test_fhmoe_mp1, not a required_gfx dict key. Pipeline methods are one-per-family; there is no loop over TUNER_FAMILIES that would miss a dict field. _is_gfx950() uses get_gfx() (honors GPU_ARCHS), same as CI.

test_run_config.py

  • Family fhmoe: extra_args=["--fhmoe"], config_property=AITER_CONFIG_FHMOE_FILE, required_gfx=gfx950, timeout=1800.
  • Why required_gfx lives on the dict: TestRunConfig loops every family, and TestRunConfigCustom (TUNE_TEST_FAMILY=fhmoe) must skip too. A skip only inside test_fhmoe would still blow up Custom on gfx942. Helper _skip_unless_required_gfx is called from both. Other families do not SystemExit in __main__ on a Tuning Tests gfx950 box, so they do not need this key.
  • Why 1800s: 600s died after token=512 cosine OK. 12 shipped rows finished in 690.9s, all OK. Weights ~6 GiB/row sequential, plus a distinct FlyDSL pair JIT per token.
  • run_config here is correctness + e2e of the shipped CSV, not a tuner search.

test_fhmoe_replay.py (in tree, not on the yaml)

Tuner → serving in a fresh interpreter: tuner times _fused_moe_impl; this checks that public fused_moe(...) with shared-expert args reads only AITER_CONFIG_FHMOE. Empty DSV4 CSV must raise; a CSV FhmoeTuner actually wrote must be the kn1/kn2 that launch. A planted row is not the contract. M=16 needs the padded-M ladder in the file; the test tunes 1,2,4,8,16 and replays 1 and 16.

Other families do not put a sibling replay on tuning-tests.yaml. Listing it there was considered and dropped so FHMoE matches fmoe/a8w8 CI surface. Run it locally if you want the serving lookup proof.

What this PR does not do (on purpose)

  • Does not change serving fused_moe / fhmoe.py.
  • Does not commit a new tuned_fhmoe.csv. A tuner that beats one shipped token is evidence, not a default swap.
  • Does not pass --all on Operators Tuning (would overwrite shipped keys).
  • Does not run the FlyDSL cartesian in CI (token=1 alone is 5632 pairs; default 1800s watchdog would kill the grouped task).
  • Does not add replay / extra tests to tuning-tests.yaml.
  • Does not switch op_tune.sh from get_gfx() to rocminfo.
  • Does not unify pipeline’s _is_gfx950() with run_config’s required_gfx helper. Different dispatch (named method vs family loop + Custom). Both honor get_gfx() / GPU_ARCHS.
  • Does not pad catalogue tokens on read, and does not change BaseTuner.tune_summary. A non-power-of-two -i row still trips [Tuning not Finished] after a successful padded write (token=3 check above). The shipped 1…2048 catalogue does not.
  • Does not pass --warmup/--iters through the homogeneous FMoE search. Only --fhmoe was changed.

Perf evidence (token=1 only — not a CSV update)

Quiet gfx950, HIP_VISIBLE_DEVICES on an idle GPU (device 4; device 1 was 98% and unused). TUNE_MOE_KERNEL_REGEX unset. --timeout 86400 because 5632 pairs are one mp_tuner group. Tune ~1.5 h, then --run_config shipped vs tuner CSV, same public fused_moe, warmup=2, iters=5.

Tuner wrote /tmp/tuned_from_producer.csv (1 row, us>0, 0 failed). Cosine OK on both clocks.

CSV kn1 kn2 e2e_us status
Before shipped tuned_fhmoe.csv token=1 …_t32x64x256_w4_gui_kw4_fp8 …_t32x256x128_atomic 39.58 OK
After tuner-written CSV same kn1 …_t32x128x128_atomic_persist 37.83 OK

Same kn1, different kn2: real cartesian, not the two-pair smoke regex. After is not slower (~4%). One token, 2/5 iters; not a reason to replace the 12-row shipped file in this PR. Tuner us=38.86 on after is a different clock (101 iters in run_perftest); compare e2e to e2e only.

Files

  • csrc/ck_gemm_moe_2stages_codegen/gemm_moe_tune.py — FhmoeTuner, --fhmoe, cosine, run_config, _lookup_token (publish get_padded_M)
  • aiter/configs/untuned_fhmoe.csv — DSV4 I384 catalogue
  • .github/scripts/op_tune.sh — fhmoe job + gfx950 skip
  • op_tests/tuning_tests/test_{csv_validation,tune_pipeline,run_config}.py + README
  • op_tests/tuning_tests/test_fhmoe_replay.py — local tuner replay
  • op_tests/test_fhmoe.py — pytest __main__

Test plan

  • python3 -m unittest op_tests.tuning_tests.test_csv_validation -v
  • python3 -m unittest op_tests.tuning_tests.test_tune_pipeline.TestTunePipeline.test_fhmoe_mp1 -v (gfx950, ~72s, 2 pairs)
  • python3 -m unittest op_tests.tuning_tests.test_run_config.TestRunConfig.test_fhmoe -v (gfx950, 12 shipped rows, 690.9s, all OK)
  • Tuner token=1 cartesian (candidates=5632) + --run_config shipped vs tuner CSV (table above)
  • Catalogue token=3 → tuned CSV token=4 (pad publish), two-pair smoke, gfx950, 66s, us=47.3686. Summary leftover 3 vs 4 / exit 1 is the unpadded -i key, not a cosine fail.
  • tuning-tests.yaml GitHub workflow: not a PR check. Equivalent fhmoe tests already ran on gfx950. Full-suite dispatch (test_tune_pipeline / test_run_config for every family) is optional and not required to land the tuner.

@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5751 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@a-canadasruiz
a-canadasruiz force-pushed the users/acanadas/fhmoe-tuner branch 3 times, most recently from 25a6fb6 to 1dfc0b3 Compare September 23, 2026 11:37
@a-canadasruiz a-canadasruiz changed the title [FlyDSL] Add gfx950 FHMoE producer on gemm_moe_tune.py [FlyDSL] Add gfx950 FHMoE tuner on gemm_moe_tune.py Sep 24, 2026
Fill untuned_fhmoe.csv with DSV4 I384 shapes and add FhmoeTuner (FMoE as
template). generate_fhmoe_data builds routed/shared tensors, pads the dummy
expert, preshuffles from gate_mode, and _run_fhmoe calls fused_moe. Candidate
timing and tuned CSV write are not wired yet.
… quant

DSV4 dense MXFP4 quant peaks at several GiB; keep FmoeTuner.generate_data
on CUDA and build generate_fhmoe_data on CPU so the worker can time fused_moe.
fused_moe() cannot inject a one-row candidate because supports_dsv4_i384_fhmoe requires every padded M tier. Time each FlyDSL s1×s2 pair through _fused_moe_impl and keep the min-us row with block_m/kernelName1/2.
…nerated CSV

python3 op_tests/test_fhmoe.py was exiting 0 without collecting pytest; the tuner times _fused_moe_impl, so a fresh process must prove public fused_moe honors AITER_CONFIG_FHMOE (raise on an empty DSV4 row, launch the generated kernel names otherwise).
Reject fast-wrong FlyDSL pairs, retune shipped keys with --all, replay
the producer CSV through fused_moe, and measure --run_config with shared FP8.
fused_moe(M=16) requires tuned rows for 1, 2, 4, 8, and 16. Tune that
ladder, then replay only 1 and 16 through the public op.
Operators Tuning runs the full job list on gfx942; --fhmoe SystemExit must not fail the workflow.
csv_validation, a two-pair pipeline smoke, and --run_config on the
shipped DSV4 CSV; skip --fhmoe off gfx950 so SystemExit is not a FAIL.
Cosine vs torch already lives in the tuner and Tuning Tests; the 0-GPU mock lock was not a family pattern.
Serving looks up get_padded_M(token), not the catalogue token. post_process
now writes the same key _write_candidate_csv already timed with.
Drop split-k only when it does not divide K, instead of excluding every k_batch!=1.
mp_tuner was getting empty kwargs, so search always used run_perftest 2/101.
A fused-only pair list drops shipped DSV4 winners whose kn1 has no trailing _fp8. Emit the same names as FmoeTuner.s1_variants.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant