Repository navigation
[Tuning] Share mp_tuner typed statuses and add a central tuning policy - #5846
Draft
amd-bartgips wants to merge 7 commits into
Draft
amd-bartgips wants to merge 7 commits into
amd-bartgips wants to merge 7 commits into
Conversation
Contributor
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
amd-bartgips
added a commit
that referenced
this pull request
Sep 25, 2026
The race, resume journal, finalist rounds, standard-error gate, evidence manifest and re-tune incumbent move to a follow-up on top of #5846, so this PR carries only what a first tuned table needs: - every candidate (or a seeded --candidate-sample) is measured once through the stock mp_tuner, alongside the auto-select incumbent and the shipped tile defaults; - a winner is published only if it beats the incumbent by more than MHA_FWD_INDIFFERENCE_DELTA (2%), and every written row is proven by a fresh-process dispatch probe; - shapes that already have a row for this GPU are skipped unless --all is given; --all --compare --update_improved re-tunes them and replaces a row only if the public operator got faster; - a re-tune that finds auto-select best warns that the existing row still overrides it, since writing no row cannot remove one. Candidates that fail the accuracy check are labelled mismatch rather than crash, against the --errRatio the run was given. Co-authored-by: Cursor <cursoragent@cursor.com>
amd-bartgips
added a commit
that referenced
this pull request
Sep 25, 2026
mp_tuner, base_tuner's measurement_kwargs and _read_csv fix, the FileBaton zombie check and their tests return to #5592's version. They now live in #5846, where the other tuners can use them, and the MHA tuner no longer needs them. base_tuner keeps only the gpu_model filter in run_config that tuned MHA rows rely on. Co-authored-by: Cursor <cursoragent@cursor.com>
This was referenced Sep 25, 2026
amd-bartgips
added a commit
that referenced
this pull request
Sep 28, 2026
The gate against the incumbent used 2%, while --all --compare --update_improved keeps a row only if the public operator got at least --min_improvement_pct (3%) faster. A 2-3% winner could be published by the search and then be refused by the protected re-tune. The gate now uses the same 3% and the same "at least" comparison, as MHA_FWD_MIN_IMPROVEMENT, which is also the value the shared PromotionPolicy in #5846 uses. Co-authored-by: Cursor <cursoragent@cursor.com>
amd-bartgips
added a commit
that referenced
this pull request
Sep 28, 2026
#5846 now has one improvement bar (min_improvement_pct, 3%) for --compare and the gate, and leaves the finalist settings to the search that uses them. This branch adds them back as FinalistPolicy (finalists, rounds, significance_sigma) next to RacePolicy, and changes what the bar means: - gate_against_incumbent takes a noise_pct floor and promotes only at max(min_improvement_pct, noise_pct). Before, the standard-error bar replaced the delta, so enough finalist rounds would publish a real but tiny win that --compare would then refuse. - The floor is two combined standard errors under the exhaustive strategy, and the race's delta under --strategy race. - The MHA tuner reads the bar from --min_improvement_pct, so one flag sets it for both the search and --compare. --delta defaults to it and is only needed to make the race's indifference zone differ. - standard_error_bar and every margin are in percent, as --compare reports them; the evidence manifest records margin_pct, bar_pct and noise_floor_pct per promotion, and the promotion bar per run. Co-authored-by: Cursor <cursoragent@cursor.com>
amd-bartgips
added a commit
that referenced
this pull request
Sep 28, 2026
The race, resume journal, finalist rounds, standard-error gate, evidence manifest and re-tune incumbent move to a follow-up on top of #5846, so this PR carries only what a first tuned table needs: - every candidate (or a seeded --candidate-sample) is measured once through the stock mp_tuner, alongside the auto-select incumbent and the shipped tile defaults; - a winner is published only if it beats the incumbent by more than MHA_FWD_INDIFFERENCE_DELTA (2%), and every written row is proven by a fresh-process dispatch probe; - shapes that already have a row for this GPU are skipped unless --all is given; --all --compare --update_improved re-tunes them and replaces a row only if the public operator got faster; - a re-tune that finds auto-select best warns that the existing row still overrides it, since writing no row cannot remove one. Candidates that fail the accuracy check are labelled mismatch rather than crash, against the --errRatio the run was given.
amd-bartgips
added a commit
that referenced
this pull request
Sep 28, 2026
mp_tuner, base_tuner's measurement_kwargs and _read_csv fix, the FileBaton zombie check and their tests return to #5592's version. They now live in #5846, where the other tuners can use them, and the MHA tuner no longer needs them. base_tuner keeps only the gpu_model filter in run_config that tuned MHA rows rely on.
amd-bartgips
added a commit
that referenced
this pull request
Sep 28, 2026
The gate against the incumbent used 2%, while --all --compare --update_improved keeps a row only if the public operator got at least --min_improvement_pct (3%) faster. A 2-3% winner could be published by the search and then be refused by the protected re-tune. The gate now uses the same 3% and the same "at least" comparison, as MHA_FWD_MIN_IMPROVEMENT, which is also the value the shared PromotionPolicy in #5846 uses.
amd-bartgips
force-pushed
the
users/bartgips/tuning-shared-infra
branch
from
September 28, 2026 14:17
04a2090 to
f36f909
Compare
amd-bartgips
added a commit
that referenced
this pull request
Sep 28, 2026
#5846 now has one improvement bar (min_improvement_pct, 3%) for --compare and the gate, and leaves the finalist settings to the search that uses them. This branch adds them back as FinalistPolicy (finalists, rounds, significance_sigma) next to RacePolicy, and changes what the bar means: - gate_against_incumbent takes a noise_pct floor and promotes only at max(min_improvement_pct, noise_pct). Before, the standard-error bar replaced the delta, so enough finalist rounds would publish a real but tiny win that --compare would then refuse. - The floor is two combined standard errors under the exhaustive strategy, and the race's delta under --strategy race. - The MHA tuner reads the bar from --min_improvement_pct, so one flag sets it for both the search and --compare. --delta defaults to it and is only needed to make the race's indifference zone differ. - standard_error_bar and every margin are in percent, as --compare reports them; the evidence manifest records margin_pct, bar_pct and noise_floor_pct per promotion, and the promotion bar per run.
Moved unchanged from the MHA forward tuner branch (#5762) so that other tuners can use them without waiting for it: - mp_tuner: optional per-candidate status and detail (return_status), a per-candidate progress queue with result_callback, a faulted group keeping the candidates it had already measured, MpTunerTask, broader accelerator-fault detection, and a pool restart on a stale PID map. - TunerCommon.measurement_kwargs, so --warmup and --iters reach the timed call instead of run_perftest's defaults. - _read_csv builds a new frame rather than assigning into a slice. - FileBaton treats a zombie lock holder as dead, so a builder that exited without releasing its lock no longer wedges the module build.
A caller that does not pass return_status picks the fastest finite latency and cannot tell a partial group from a complete one. Keeping the candidates measured before a fault would let it publish the winner of a search that stopped early, so those callers get every candidate of the group back as failed, as before. Callers with return_status keep the measured candidates and see the crash or timeout status on the rest. test_mp_tuner_fault runs one shape group with a faulting candidate through the real pool on one GPU, in both call forms.
aiter/utility/tuning_policy.py holds the thresholds whose justification does not mention a kernel family: - MeasurementPolicy: warmup, iters, errRatio and timeout, which TunerCommon.ARG_DEFAULTS now reads (same values as before). - COMPARE_MIN_IMPROVEMENT_PCT: the --compare --update_improved bar. - PromotionPolicy: the indifference delta a challenger has to clear over the incumbent, and the finalist count and rounds for re-timing. - gate_against_incumbent: the promote-or-retain decision on those numbers, for tuners that measure the incumbent themselves. Standard library only, so the runtime and CPU tests can import it. A family that needs a different threshold overrides it with dataclasses.replace in its own module.
--compare --update_improved and gate_against_incumbent both decide whether a faster configuration is worth replacing what a shape runs today, but read two different numbers (3% and 2%). They now read one: PromotionPolicy.min_improvement_pct, 3.0, named after the --min_improvement_pct flag whose default it already was. - The gate works in percent and promotes at "at least" the bar, as --compare does, and reports margin_pct. - COMPARE_MIN_IMPROVEMENT_PCT and indifference_delta are gone; neither was on main. - finalists and finalist_rounds move to the follow-up that times finalists (#5849), since nothing on main uses them. - batch comes from a new RunPolicy, so no TunerCommon default is a bare literal any more. It sits apart from MeasurementPolicy because it changes how often the tuned CSV is written, not which candidate wins. No default a tuner sees changes.
test_mp_tuner_logic and test_mp_tuner_fault imported Triton before torch, and the start-time test switched its pool from spawn to fork, both for a dynamic-loader abort in spawned workers. Main runs the same test, and the worker-PID test from #5841, under spawn with no Triton import, and the abort does not reproduce on MI355X. The tests now use spawn, the start method mp_tuner forces, and a CPU test no longer needs Triton installed.
Worker death is now caught by #5841's check, which this branch rebases onto. Without it, the immediate pool restart on a stale PID map resubmitted the dead group with a fresh start time, so a candidate that kills its worker restarted the pool forever. - The worker-exited branch drains the progress queue before failing the group, as the other failure branches do, so a return_status caller keeps the candidates the worker measured before it died. The detail names the exited worker. - With return_status, the candidates behind a fault come back as not_run instead of carrying the fault's status. The callback stream is unchanged: it still receives only the candidate that failed. - The mp_tuner docstring describes both result forms, every status and result_callback. - test_mp_tuner_fault adds a candidate that exits its worker with another shape queued behind it, in both call forms. With the dead-worker check disabled it fails at its deadline after repeated pool restarts. The file now runs in the per-PR shards (67 s on one MI355X).
amd-bartgips
force-pushed
the
users/bartgips/tuning-shared-infra
branch
from
September 30, 2026 10:56
f36f909 to
839de6b
Compare
work_group catches an error raised outside worker(), such as a later candidate's gen_data or reference failing, and returns a result list itself. That list failed every candidate of the group, including those already measured, and put all of them on the progress queue. The parent saw a normal return, so its partial-result handling never ran: a return_status caller lost the measured candidates, a checkpointing caller received a second, crash result for each of them, and the candidates behind the abort were checkpointed as tried, so a resume skipped them. work_group now builds those results with _failed_group_results, the same rules the parent applies to a group the pool lost: with return_status, measured candidates keep their results, the one that aborted gets the failure, and the rest are not_run; without it, the whole group fails as before. Only the aborting candidate is published. test_mp_tuner_fault adds a group whose second candidate fails in gen_data (reported by Araceli); the typed case failed before this change.
amd-bartgips
added a commit
that referenced
this pull request
Oct 1, 2026
The race, resume journal, finalist rounds, standard-error gate, evidence manifest and re-tune incumbent move to a follow-up on top of #5846, so this PR carries only what a first tuned table needs: - every candidate (or a seeded --candidate-sample) is measured once through the stock mp_tuner, alongside the auto-select incumbent and the shipped tile defaults; - a winner is published only if it beats the incumbent by more than MHA_FWD_INDIFFERENCE_DELTA (2%), and every written row is proven by a fresh-process dispatch probe; - shapes that already have a row for this GPU are skipped unless --all is given; --all --compare --update_improved re-tunes them and replaces a row only if the public operator got faster; - a re-tune that finds auto-select best warns that the existing row still overrides it, since writing no row cannot remove one. Candidates that fail the accuracy check are labelled mismatch rather than crash, against the --errRatio the run was given.
amd-bartgips
added a commit
that referenced
this pull request
Oct 1, 2026
mp_tuner, base_tuner's measurement_kwargs and _read_csv fix, the FileBaton zombie check and their tests return to #5592's version. They now live in #5846, where the other tuners can use them, and the MHA tuner no longer needs them. base_tuner keeps only the gpu_model filter in run_config that tuned MHA rows rely on.
amd-bartgips
added a commit
that referenced
this pull request
Oct 1, 2026
The gate against the incumbent used 2%, while --all --compare --update_improved keeps a row only if the public operator got at least --min_improvement_pct (3%) faster. A 2-3% winner could be published by the search and then be refused by the protected re-tune. The gate now uses the same 3% and the same "at least" comparison, as MHA_FWD_MIN_IMPROVEMENT, which is also the value the shared PromotionPolicy in #5846 uses.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Tuners built on
mp_tunerandTunerCommonneed four things the shared code does not give them today: to know why a candidate failed, not just that it did; to lose only the faulting candidate when a worker dies; to pass--warmup/--itersthrough to the timed call; and to agree on when a faster candidate is worth publishing.This PR adds those to the shared tuning code, on
main, so every tuner family can use them. The first users are the MHA forward tuner's race and evidence (#5849) and the follow-up to the FHMoE tuner (#5751).What this PR adds
mp_tuner:return_status=Truereturns(info, us, err, status, detail), withstatusone ofok,mismatch,unsupported,crash,timeout,oom_runtime,oom_preflight,not_run. Without it, results keep today's three-tuple shape. The docstring describes every status.result_callback), so a caller can checkpoint as candidates finish.return_status, a faulted shape group keeps the candidates it had already measured. This includes a group whose worker died, which [tuner] Fail a task as soon as its worker process exits #5841's dead-worker check now detects, and a group thatwork_groupaborts itself because a later candidate's inputs or reference failed to build. The candidate that faulted is reported with the fault's status. The candidates behind it never ran; they come back asnot_runand are not passed toresult_callback, so a resume retries them.MpTunerTask, aNamedTuplein the orderwork_groupunpacks, so tasks can be built by name. Plain tuples still work.return_statusgets every candidate of a faulted group back as failed. Without this, the partial-result survival above would let such a caller publish the fastest survivor of a search that never finished. For a worker that dies, this matchesmain. For an in-process accelerator fault (an error mentioning an illegal memory access, a memory access fault, a device-side assert or HIP error 700), it does not:maingave only the faulting candidate −1 and kept the candidates measured before and after it on the same worker, while this PR fails the whole shape, so existing tuners write no row for it. HIP leaves the context unusable after such a fault, so timings taken after it on that worker cannot be trusted.TunerCommon.measurement_kwargs(args): the timing kwargs for anmp_tunertask, so--warmupand--itersreachrun_perftestinstead of its defaults.aiter/utility/tuning_policy.py: thresholds whose justification does not mention a kernel family. Standard library only.TunerCommon.ARG_DEFAULTSnow reads every value below instead of spelling it out; no value changes, and family subclasses keep their own overrides.MeasurementPolicy: warmup 5, iters 101, errRatio 0.05, timeout 1800.RunPolicy:batch100, the number of shapes tuned between writes of the tuned CSV.PromotionPolicy:min_improvement_pct3.0, the default of--min_improvement_pct.gate_against_incumbent(incumbent_us, challenger_us, policy): returnspromotewhen the challenger is at leastmin_improvement_pctfaster, otherwiseretain, plus the margin in percent. With no measured incumbent it promotes and reports no margin._read_csvbuilds a new frame instead of assigning into a slice;FileBatontreats a zombie lock holder as dead, so a builder that exited without releasing its lock no longer wedges every later build of that module.Reading this diff
+1048 / −93 across 11 files; more than half is tests.
mp_tuner.py(+306 / −56): typed statuses, the progress queue, partial-result survival,MpTunerTask, fault detection, and the coverage check in_failed_group_results.tuning_policy.py(+117) andbase_tuner.py(+44 / −17): the new module and whereARG_DEFAULTSreads it.file_baton.py(+8 / −1).No default value changes. For existing callers, results differ only when something goes wrong:
Technical Details
One improvement bar.
--compare --update_improvedbenchmarks each shape twice: once with the kernel it used before tuning, and once with the newly tuned kernel. It keeps the new row only if the shape got at leastmin_improvement_pct(3%) faster.gate_against_incumbentapplies the same test inside the search, for tuners that time the current kernel next to their candidates. Both read the same value, so they cannot disagree on whether a win is big enough.Worker death and the stale PID map. When a candidate kills its worker, the pool starts a replacement whose PID is not in the GPU map, and the next shape it takes raises
KeyError, which restarts the pool. The dead shape must be failed before that restart, or it is resubmitted, kills its worker again, and the pool restarts forever. #5841's check does this: the parent finds the dead worker while it waits on that shape, and the pool hands out shapes in order, so the dead shape is always checked before the one the replacement took. This branch is rebased onto #5841. The dead-worker branch drains the progress queue before failing the shape, as the other failure branches do.test_mp_tuner_faultcovers this path per PR.Unused so far.
gate_against_incumbenthas no caller onmainyet; #5849 and the FHMoE follow-up are its first users. Everything else in the module is read byTunerCommon.Overriding. A family that needs a different bar sets it in its own module with
dataclasses.replace(DEFAULT_PROMOTION, min_improvement_pct=...).No tuner changes. Tuners that restate a base default (
batch: 100in 10 tuners,errRatio: 0.05in 5) keep their copy until their own PR moves them onto the policy.Test Plan
python3 -m unittest op_tests.tuning_tests.test_tuning_policy op_tests.tuning_tests.test_mp_tuner_logic op_tests.tuning_tests.test_tuner_infra op_tests.tuning_tests.test_compare_logic op_tests.tuning_tests.test_csv_validation op_tests.test_jit_cache_transaction(CPU).op_tests/tuning_tests/test_mp_tuner_fault.py(new, one GPU), through the real pool, each case once in the positional call form existing tuners use and once withreturn_status=True:os._exit), with another shape queued behind it, so the replacement worker hits the stale PID map;gen_data, beforeworker()runs, sowork_groupaborts the group itself (case reported by Araceli in review). The typed case also checks thatresult_callbacksees each candidate that ran once and never the one behind the abort.test_tuning_policyandtest_mp_tuner_faultare added to the per-PR test shards (split_tests.sh). In the nightly tuning workflow,test_tuning_policyruns in the level 0+1 job andtest_mp_tuner_faultin the GPU pipeline job.Test Result
ruff checkandblackare clean on the changed files.test_mp_tuner_fault: all 6 tests pass, in 84 s run as a file on one MI355X. With [tuner] Fail a task as soon as its worker process exits #5841's dead-worker check disabled, the worker-death test fails at its deadline after 8 pool restarts, instead of passing or hanging. Before thework_groupfix, the typed pre-launch test failed: the candidate measured before the abort came back ascrash. Before the rebase, with the coverage check removed, the positional-form fault test failed as intended: the working candidate came back at 1.56 us from a group whose other candidate faulted.Related PRs
gate_against_incumbent, and extends the policy with the finalist rounds and their standard-error bar,RacePolicyand the sampling seed. It reads only theresult_callbackstream, whichnot_rundoes not change.mp_tuner.MpTunerTask,measurement_kwargsandgate_against_incumbent.main; this branch is rebased onto it and relies on it for worker death.Owners
aiter/utility/mp_tuner.py,aiter/utility/base_tuner.py: @yzhou103aiter/jit/utils/file_baton.py: @Lingpeng-JinCredit
Part of the
mp_tunerandbase_tunerwork is @amd-yashagar's.Submission Checklist