Repository navigation
[FlyDSL] Add gfx950 FHMoE tuner on gemm_moe_tune.py - #5751
Draft
a-canadasruiz wants to merge 14 commits into
Draft
a-canadasruiz wants to merge 14 commits into
a-canadasruiz wants to merge 14 commits into
Conversation
Contributor
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
a-canadasruiz
force-pushed
the
users/acanadas/fhmoe-tuner
branch
3 times, most recently
from
September 23, 2026 11:37
25a6fb6 to
1dfc0b3
Compare
1 task done
1 task done
Fill untuned_fhmoe.csv with DSV4 I384 shapes and add FhmoeTuner (FMoE as template). generate_fhmoe_data builds routed/shared tensors, pads the dummy expert, preshuffles from gate_mode, and _run_fhmoe calls fused_moe. Candidate timing and tuned CSV write are not wired yet.
… quant DSV4 dense MXFP4 quant peaks at several GiB; keep FmoeTuner.generate_data on CUDA and build generate_fhmoe_data on CPU so the worker can time fused_moe.
fused_moe() cannot inject a one-row candidate because supports_dsv4_i384_fhmoe requires every padded M tier. Time each FlyDSL s1×s2 pair through _fused_moe_impl and keep the min-us row with block_m/kernelName1/2.
…nerated CSV python3 op_tests/test_fhmoe.py was exiting 0 without collecting pytest; the tuner times _fused_moe_impl, so a fresh process must prove public fused_moe honors AITER_CONFIG_FHMOE (raise on an empty DSV4 row, launch the generated kernel names otherwise).
Reject fast-wrong FlyDSL pairs, retune shipped keys with --all, replay the producer CSV through fused_moe, and measure --run_config with shared FP8.
fused_moe(M=16) requires tuned rows for 1, 2, 4, 8, and 16. Tune that ladder, then replay only 1 and 16 through the public op.
Operators Tuning runs the full job list on gfx942; --fhmoe SystemExit must not fail the workflow.
csv_validation, a two-pair pipeline smoke, and --run_config on the shipped DSV4 CSV; skip --fhmoe off gfx950 so SystemExit is not a FAIL.
Cosine vs torch already lives in the tuner and Tuning Tests; the 0-GPU mock lock was not a family pattern.
Serving looks up get_padded_M(token), not the catalogue token. post_process now writes the same key _write_candidate_csv already timed with.
Drop split-k only when it does not divide K, instead of excluding every k_batch!=1.
mp_tuner was getting empty kwargs, so search always used run_perftest 2/101.
a-canadasruiz
force-pushed
the
users/acanadas/fhmoe-tuner
branch
from
September 25, 2026 11:43
c12acdc to
bdac6e9
Compare
1 task done
A fused-only pair list drops shipped DSV4 winners whose kn1 has no trailing _fp8. Emit the same names as FmoeTuner.s1_variants.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
gfx950 fused heterogeneous MoE (FHMoE, DSV4 I384) already serves through
fused_moe+aiter/configs/tuned_fhmoe.csv. What was missing is a tuner: that sweeps FlyDSL stage1×stage2 pairs, rejects fast-wrong kernels against torch, and writes a native tuned CSV (block_m,kernelName1,kernelName2,us).This PR adds that tuner as
FhmoeTuneron the existing FMoE script (gemm_moe_tune.py --fhmoe), wires it intoop_tune.shand Tuning Tests like the other families, and does not replace the shipped 12-rowtuned_fhmoe.csv.FHMoE is its own family, not a flag on FMoE tests: different CLI (
--fhmoe), different CSVs (untuned_fhmoe.csv/tuned_fhmoe.csv/AITER_CONFIG_FHMOE), gfx950-only, shared-expert + routed FP4 + shared FP8. Servingfused_moe/aiter/fhmoe.pyare not in the diff.Tuner (
FhmoeTuner)Subclass of
FmoeTuner, selected with--fhmoe(mutually exclusive with--grouped-gemmand--mxfp4-flydslon the same__main__).Catalogue.
aiter/configs/untuned_fhmoe.csv: 12 DSV4 I384 rows, tokens 1,2,4,…,2048. Same shape keys as shippedtuned_fhmoe.csv(that file has nouscolumn; the producer addsus).Arch gate.
aiter/fhmoe.pyisNotImplementedErroroff gfx950.gemm_moe_tune.pymultiplexes four tuners in one file, so the gate is in__main__:--fhmoeon a non-gfx950 processSystemExits (--fhmoe is only supported on gfx950). Grouped GEMM already does the same for gfx1250. This is not A6W6’sRuntimeErrorinsidetune(): A6W6 is its own script, so a unittest subprocess can still import and skip; here the process dies beforeFhmoeTunerruns. Changing FHMoE toRuntimeErrorwould not remove the need to skip--fhmoesubprocesses in Tuning Tests, and it would be inconsistent with--grouped-gemm. The skip belongs in the tests (and inop_tune.sh), not by weakening__main__.How a candidate is timed. Public
fused_moe()cannot inject a one-row candidate.supports_dsv4_i384_fhmoerequires every padded-M tier in the config file (e.g.fused_moe(M=16)needs tuned rows for 1, 2, 4, 8, and 16). The worker therefore times_fused_moe_implwith a temporary one-row CSV (kernelName1/kernelName2/block_m). That is kernel-pair search, not the serving lookup.Published token is the serving lookup key. Serving
get_2stage_cfgslooks upget_padded_M(token_num), not the raw catalogue token (fused_moe.py; for M < 32768 that isnextPow2, so 3072 → 4096). Candidate timing already wrote that padded key in_write_candidate_csv.post_processused to emit the catalogue token, so a non-power-of-two untuned row would be timed at the padded tier and published under a key serving never reads: the gap-fill job can exit 0, the CSV grows a dead row, and the shape stays uncovered. The tuner now publishesget_padded_M(X)through_lookup_token, the same helper timing uses. Save the padded token only — notX, not both — because serving has a singletokencolumn.The 12 shipped catalogue rows are already 1,2,4,…,2048;
pad(X) = Xfor them and this does not changetuned_fhmoe.csv.Local check on gfx950, two-pair smoke regex (not the cartesian),
--mp 1, ~66s: catalogue token=3 →/tmp/tuned_fhmoe_token3.csvhastoken=4,us=47.3686, cosine OK (0 failed).BaseTuner.tune_summarystill row-matches the-iframe (token=3) against success (token=4), so it prints leftover “untuned” and[Tuning not Finished]/ exit 1. That is a key mismatch in the shared summary, not a failed kernel. We do not pad the catalogue on read and we do not rewriteBaseTuner’s message (every family uses it). Serving lookup is the contract this tuner has to match; the 12-row gap-fill job never hits the mismatch.Search. For each catalogue row,
_kernel_pairsbuilds the legal FlyDSL cartesian (block_m∈ {16,32,64,128} × stage1 names × stage2 names, after tile/LDS/inter_dimfilters). DSV4 I384 INTERLEAVE is on the order of 5632 pairs per token. Winner = minusamong pairs with cosine vs torch ≤errRatio.post_processdrops the whole shape if every pair fails. Fast-wrong pairs must not win on latency.Data.
generate_fhmoe_databuilds routed MXFP4 + shared FP8 tensors on the host (FMoE’sgenerate_datastays on CUDA). Dummy expert pad,gate_modepreshuffle. Weights are ~6 GiB per shape and token-independent.run_configregenerates per row sequentially andempty_caches; this is not a 12×6 GiB peak. Left as-is; not atest_run_configbug.mp_tuner. Same worker pool as FMoE.
shape_grouped=Trueputs every pair of one shape in one task group. Default--timeoutis 1800s per group. A full cartesian for one token is hours, so a real tuner run needs--timeouton the order of a day (see Perf). Pipeline smoke is two pairs and stays under 1800s.--warmup/--itersnow go through asrun_perftest kwargson the FHMoE search, not only--run_config.Homogeneous FMoE still uses the 2/101 defaults; that family is out of scope here.--run_config. Overrides the parent FMoE path. Iterates the tuned CSV, calls publicfused_moewith shared-expert args, cosine vs torch, reportse2e_us.--run_config <csv>setsAITER_CONFIG_FHMOEto that file (same env serving uses) and clears FHMoE caches. Tunerusande2e_usare different clocks (search vs production); do not compare them.op_tests/test_fhmoe.py.__main__now runspytest.main.python3 op_tests/test_fhmoe.pyused to exit 0 without collecting tests. That file is the op oracle, not the tuner.op_tune.sh / Operators Tuning
Job, same gap-fill shape as FMoE:
No
--all. Without--all,FmoeTuner.pre_processdrops keys already present in-o(“only kernels that are not in the tuned CSV”). The 12 catalogue rows already match shippedtuned_fhmoe.csv, so on gfx950 this job is a no-op until someone adds a new untuned shape.--allwould retune those keys in CI and could overwrite production defaults. This PR is the tuner, not a retune of shipped rows.Operators Tuning’s runner is
linux-aiter-oci-mi300x-1(gfx942). The script walks the fulltune_jobslist, so--fhmoewouldSystemExitand fail the workflow. The fhmoe job skips unlessget_gfx() == gfx950.get_gfx()followsGPU_ARCHS(same helper as the rest of this script and aiter tests). Liverocminfo/get_gfx_runtime()is the other clock; we did not switchop_tune.shto it.Tuning Tests (family tables, not a one-off yaml)
tuning-tests.yamlis unchanged:test_csv_validation,test_tune_pipeline,test_run_config.GPU_ARCHS=gfx950, runnerlinux-aiter-mi35x-1. Triggers are schedule +workflow_dispatchonly — notpull_request. Opening this PR does not launch that workflow.Grouped GEMM is not a Tuning Tests family (gfx1250; this workflow is gfx950). FHMoE is a family, with an explicit gfx950 skip so
--fhmoeSystemExitis not a FAIL ifGPU_ARCHSis wrong.test_csv_validation.pyTUNED_CSVSincludestuned_fhmoe.csv.test_fhmoe_no_duplicatesuses FHMoE extra keys (act_type, dtypes,shared_expert_id, pads,gate_mode, …), not GEMMM,N,K.test_tune_pipeline.pyTUNE_MOE_KERNEL_REGEXpins two kn1×kn2 pairs, not the cartesian (timeout_mp11200s). Measured ~72s on gfx950.test_fhmoe_mp1only — nomp_default. One GPU is enough to prove the tuner writes a CSV; all-GPU cartesian smoke is not the family pattern we need.if not _is_gfx950(): skipTestinsidetest_fhmoe_mp1, not arequired_gfxdict key. Pipeline methods are one-per-family; there is no loop overTUNER_FAMILIESthat would miss a dict field._is_gfx950()usesget_gfx()(honorsGPU_ARCHS), same as CI.test_run_config.pyfhmoe:extra_args=["--fhmoe"],config_property=AITER_CONFIG_FHMOE_FILE,required_gfx=gfx950,timeout=1800.required_gfxlives on the dict:TestRunConfigloops every family, andTestRunConfigCustom(TUNE_TEST_FAMILY=fhmoe) must skip too. A skip only insidetest_fhmoewould still blow up Custom on gfx942. Helper_skip_unless_required_gfxis called from both. Other families do notSystemExitin__main__on a Tuning Tests gfx950 box, so they do not need this key.run_confighere is correctness + e2e of the shipped CSV, not a tuner search.test_fhmoe_replay.py(in tree, not on the yaml)Tuner → serving in a fresh interpreter: tuner times
_fused_moe_impl; this checks that publicfused_moe(...)with shared-expert args reads onlyAITER_CONFIG_FHMOE. Empty DSV4 CSV must raise; a CSVFhmoeTuneractually wrote must be the kn1/kn2 that launch. A planted row is not the contract. M=16 needs the padded-M ladder in the file; the test tunes 1,2,4,8,16 and replays 1 and 16.Other families do not put a sibling replay on
tuning-tests.yaml. Listing it there was considered and dropped so FHMoE matches fmoe/a8w8 CI surface. Run it locally if you want the serving lookup proof.What this PR does not do (on purpose)
fused_moe/fhmoe.py.tuned_fhmoe.csv. A tuner that beats one shipped token is evidence, not a default swap.--allon Operators Tuning (would overwrite shipped keys).tuning-tests.yaml.op_tune.shfromget_gfx()torocminfo._is_gfx950()with run_config’srequired_gfxhelper. Different dispatch (named method vs family loop + Custom). Both honorget_gfx()/GPU_ARCHS.BaseTuner.tune_summary. A non-power-of-two-irow still trips[Tuning not Finished]after a successful padded write (token=3 check above). The shipped 1…2048 catalogue does not.--warmup/--itersthrough the homogeneous FMoE search. Only--fhmoewas changed.Perf evidence (token=1 only — not a CSV update)
Quiet gfx950,
HIP_VISIBLE_DEVICESon an idle GPU (device 4; device 1 was 98% and unused).TUNE_MOE_KERNEL_REGEXunset.--timeout 86400because 5632 pairs are onemp_tunergroup. Tune ~1.5 h, then--run_configshipped vs tuner CSV, same publicfused_moe, warmup=2, iters=5.Tuner wrote
/tmp/tuned_from_producer.csv(1 row,us>0, 0 failed). Cosine OK on both clocks.tuned_fhmoe.csvtoken=1…_t32x64x256_w4_gui_kw4_fp8…_t32x256x128_atomic…_t32x128x128_atomic_persistSame kn1, different kn2: real cartesian, not the two-pair smoke regex. After is not slower (~4%). One token, 2/5 iters; not a reason to replace the 12-row shipped file in this PR. Tuner
us=38.86on after is a different clock (101 iters inrun_perftest); compare e2e to e2e only.Files
csrc/ck_gemm_moe_2stages_codegen/gemm_moe_tune.py—FhmoeTuner,--fhmoe, cosine,run_config,_lookup_token(publishget_padded_M)aiter/configs/untuned_fhmoe.csv— DSV4 I384 catalogue.github/scripts/op_tune.sh— fhmoe job + gfx950 skipop_tests/tuning_tests/test_{csv_validation,tune_pipeline,run_config}.py+ READMEop_tests/tuning_tests/test_fhmoe_replay.py— local tuner replayop_tests/test_fhmoe.py— pytest__main__Test plan
python3 -m unittest op_tests.tuning_tests.test_csv_validation -vpython3 -m unittest op_tests.tuning_tests.test_tune_pipeline.TestTunePipeline.test_fhmoe_mp1 -v(gfx950, ~72s, 2 pairs)python3 -m unittest op_tests.tuning_tests.test_run_config.TestRunConfig.test_fhmoe -v(gfx950, 12 shipped rows, 690.9s, all OK)candidates=5632) +--run_configshipped vs tuner CSV (table above)token=4(padpublish), two-pair smoke, gfx950, 66s,us=47.3686. Summary leftover 3 vs 4 / exit 1 is the unpadded-ikey, not a cosine fail.tuning-tests.yamlGitHub workflow: not a PR check. Equivalent fhmoe tests already ran on gfx950. Full-suite dispatch (test_tune_pipeline/test_run_configfor every family) is optional and not required to land the tuner.