Skip to content

[Tuning] Share mp_tuner typed statuses and add a central tuning policy - #5846

Draft
amd-bartgips wants to merge 7 commits into
mainfrom
users/bartgips/tuning-shared-infra
Draft

amd-bartgips wants to merge 7 commits into
mainfrom
users/bartgips/tuning-shared-infra

Conversation

@amd-bartgips

@amd-bartgips amd-bartgips commented Sep 25, 2026 •

Copy link
Copy Markdown

Motivation

Tuners built on mp_tuner and TunerCommon need four things the shared code does not give them today: to know why a candidate failed, not just that it did; to lose only the faulting candidate when a worker dies; to pass --warmup/--iters through to the timed call; and to agree on when a faster candidate is worth publishing.

This PR adds those to the shared tuning code, on main, so every tuner family can use them. The first users are the MHA forward tuner's race and evidence (#5849) and the follow-up to the FHMoE tuner (#5751).

What this PR adds

  • mp_tuner:
    • return_status=True returns (info, us, err, status, detail), with status one of ok, mismatch, unsupported, crash, timeout, oom_runtime, oom_preflight, not_run. Without it, results keep today's three-tuple shape. The docstring describes every status.
    • A per-candidate progress queue (result_callback), so a caller can checkpoint as candidates finish.
    • With return_status, a faulted shape group keeps the candidates it had already measured. This includes a group whose worker died, which [tuner] Fail a task as soon as its worker process exits #5841's dead-worker check now detects, and a group that work_group aborts itself because a later candidate's inputs or reference failed to build. The candidate that faulted is reported with the fault's status. The candidates behind it never ran; they come back as not_run and are not passed to result_callback, so a resume retries them.
    • MpTunerTask, a NamedTuple in the order work_group unpacks, so tasks can be built by name. Plain tuples still work.
    • Broader accelerator-fault detection, and a pool restart on a stale PID-to-GPU map.
  • Coverage check: a caller that does not pass return_status gets every candidate of a faulted group back as failed. Without this, the partial-result survival above would let such a caller publish the fastest survivor of a search that never finished. For a worker that dies, this matches main. For an in-process accelerator fault (an error mentioning an illegal memory access, a memory access fault, a device-side assert or HIP error 700), it does not: main gave only the faulting candidate −1 and kept the candidates measured before and after it on the same worker, while this PR fails the whole shape, so existing tuners write no row for it. HIP leaves the context unusable after such a fault, so timings taken after it on that worker cannot be trusted.
  • TunerCommon.measurement_kwargs(args): the timing kwargs for an mp_tuner task, so --warmup and --iters reach run_perftest instead of its defaults.
  • aiter/utility/tuning_policy.py: thresholds whose justification does not mention a kernel family. Standard library only. TunerCommon.ARG_DEFAULTS now reads every value below instead of spelling it out; no value changes, and family subclasses keep their own overrides.
    • MeasurementPolicy: warmup 5, iters 101, errRatio 0.05, timeout 1800.
    • RunPolicy: batch 100, the number of shapes tuned between writes of the tuned CSV.
    • PromotionPolicy: min_improvement_pct 3.0, the default of --min_improvement_pct.
    • gate_against_incumbent(incumbent_us, challenger_us, policy): returns promote when the challenger is at least min_improvement_pct faster, otherwise retain, plus the margin in percent. With no measured incumbent it promotes and reports no margin.
  • Small fixes: _read_csv builds a new frame instead of assigning into a slice; FileBaton treats a zombie lock holder as dead, so a builder that exited without releasing its lock no longer wedges every later build of that module.

Reading this diff

+1048 / −93 across 11 files; more than half is tests.

  • mp_tuner.py (+306 / −56): typed statuses, the progress queue, partial-result survival, MpTunerTask, fault detection, and the coverage check in _failed_group_results.
  • tuning_policy.py (+117) and base_tuner.py (+44 / −17): the new module and where ARG_DEFAULTS reads it.
  • file_baton.py (+8 / −1).

No default value changes. For existing callers, results differ only when something goes wrong:

  • an in-process accelerator fault now fails the whole shape group and restarts the pool, instead of the worker carrying on in a faulted context;
  • a candidate that still times at 0 us after the retries is reported as failed;
  • a stale PID-to-GPU map restarts the pool at once, instead of waiting for another task to time out.

Technical Details

One improvement bar. --compare --update_improved benchmarks each shape twice: once with the kernel it used before tuning, and once with the newly tuned kernel. It keeps the new row only if the shape got at least min_improvement_pct (3%) faster. gate_against_incumbent applies the same test inside the search, for tuners that time the current kernel next to their candidates. Both read the same value, so they cannot disagree on whether a win is big enough.

Worker death and the stale PID map. When a candidate kills its worker, the pool starts a replacement whose PID is not in the GPU map, and the next shape it takes raises KeyError, which restarts the pool. The dead shape must be failed before that restart, or it is resubmitted, kills its worker again, and the pool restarts forever. #5841's check does this: the parent finds the dead worker while it waits on that shape, and the pool hands out shapes in order, so the dead shape is always checked before the one the replacement took. This branch is rebased onto #5841. The dead-worker branch drains the progress queue before failing the shape, as the other failure branches do. test_mp_tuner_fault covers this path per PR.

Unused so far. gate_against_incumbent has no caller on main yet; #5849 and the FHMoE follow-up are its first users. Everything else in the module is read by TunerCommon.

Overriding. A family that needs a different bar sets it in its own module with dataclasses.replace(DEFAULT_PROMOTION, min_improvement_pct=...).

No tuner changes. Tuners that restate a base default (batch: 100 in 10 tuners, errRatio: 0.05 in 5) keep their copy until their own PR moves them onto the policy.

Test Plan

  • python3 -m unittest op_tests.tuning_tests.test_tuning_policy op_tests.tuning_tests.test_mp_tuner_logic op_tests.tuning_tests.test_tuner_infra op_tests.tuning_tests.test_compare_logic op_tests.tuning_tests.test_csv_validation op_tests.test_jit_cache_transaction (CPU).
  • op_tests/tuning_tests/test_mp_tuner_fault.py (new, one GPU), through the real pool, each case once in the positional call form existing tuners use and once with return_status=True:
    • a shape group with a working candidate, a faulting candidate and one behind it;
    • a shape group whose second candidate exits its worker (os._exit), with another shape queued behind it, so the replacement worker hits the stale PID map;
    • a shape group whose second candidate fails in gen_data, before worker() runs, so work_group aborts the group itself (case reported by Araceli in review). The typed case also checks that result_callback sees each candidate that ran once and never the one behind the abort.
  • test_tuning_policy and test_mp_tuner_fault are added to the per-PR test shards (split_tests.sh). In the nightly tuning workflow, test_tuning_policy runs in the level 0+1 job and test_mp_tuner_fault in the GPU pipeline job.

Test Result

  • The 153 CPU tests pass (on an MI355X host); ruff check and black are clean on the changed files.
  • test_mp_tuner_fault: all 6 tests pass, in 84 s run as a file on one MI355X. With [tuner] Fail a task as soon as its worker process exits #5841's dead-worker check disabled, the worker-death test fails at its deadline after 8 pool restarts, instead of passing or hanging. Before the work_group fix, the typed pre-launch test failed: the candidate measured before the abort came back as crash. Before the rebase, with the coverage check removed, the positional-form fault test failed as intended: the working candidate came back at 1.56 us from a group whose other candidate faulted.

Related PRs

Owners

  • aiter/utility/mp_tuner.py, aiter/utility/base_tuner.py: @yzhou103
  • aiter/jit/utils/file_baton.py: @Lingpeng-Jin

Credit

Part of the mp_tuner and base_tuner work is @amd-yashagar's.

Submission Checklist

@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5846 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

amd-bartgips added a commit that referenced this pull request Sep 25, 2026
The race, resume journal, finalist rounds, standard-error gate, evidence
manifest and re-tune incumbent move to a follow-up on top of #5846, so
this PR carries only what a first tuned table needs:

- every candidate (or a seeded --candidate-sample) is measured once
  through the stock mp_tuner, alongside the auto-select incumbent and the
  shipped tile defaults;
- a winner is published only if it beats the incumbent by more than
  MHA_FWD_INDIFFERENCE_DELTA (2%), and every written row is proven by a
  fresh-process dispatch probe;
- shapes that already have a row for this GPU are skipped unless --all
  is given; --all --compare --update_improved re-tunes them and replaces
  a row only if the public operator got faster;
- a re-tune that finds auto-select best warns that the existing row still
  overrides it, since writing no row cannot remove one.

Candidates that fail the accuracy check are labelled mismatch rather than
crash, against the --errRatio the run was given.

Co-authored-by: Cursor <cursoragent@cursor.com>
amd-bartgips added a commit that referenced this pull request Sep 25, 2026
mp_tuner, base_tuner's measurement_kwargs and _read_csv fix, the
FileBaton zombie check and their tests return to #5592's version. They
now live in #5846, where the other tuners can use them, and the MHA
tuner no longer needs them. base_tuner keeps only the gpu_model filter
in run_config that tuned MHA rows rely on.

Co-authored-by: Cursor <cursoragent@cursor.com>
amd-bartgips added a commit that referenced this pull request Sep 28, 2026
The gate against the incumbent used 2%, while --all --compare
--update_improved keeps a row only if the public operator got at least
--min_improvement_pct (3%) faster. A 2-3% winner could be published by
the search and then be refused by the protected re-tune. The gate now
uses the same 3% and the same "at least" comparison, as
MHA_FWD_MIN_IMPROVEMENT, which is also the value the shared
PromotionPolicy in #5846 uses.

Co-authored-by: Cursor <cursoragent@cursor.com>
amd-bartgips added a commit that referenced this pull request Sep 28, 2026
#5846 now has one improvement bar (min_improvement_pct, 3%) for
--compare and the gate, and leaves the finalist settings to the search
that uses them. This branch adds them back as FinalistPolicy (finalists,
rounds, significance_sigma) next to RacePolicy, and changes what the bar
means:

- gate_against_incumbent takes a noise_pct floor and promotes only at
  max(min_improvement_pct, noise_pct). Before, the standard-error bar
  replaced the delta, so enough finalist rounds would publish a real but
  tiny win that --compare would then refuse.
- The floor is two combined standard errors under the exhaustive
  strategy, and the race's delta under --strategy race.
- The MHA tuner reads the bar from --min_improvement_pct, so one flag
  sets it for both the search and --compare. --delta defaults to it and
  is only needed to make the race's indifference zone differ.
- standard_error_bar and every margin are in percent, as --compare
  reports them; the evidence manifest records margin_pct, bar_pct and
  noise_floor_pct per promotion, and the promotion bar per run.

Co-authored-by: Cursor <cursoragent@cursor.com>
amd-bartgips added a commit that referenced this pull request Sep 28, 2026
The race, resume journal, finalist rounds, standard-error gate, evidence
manifest and re-tune incumbent move to a follow-up on top of #5846, so
this PR carries only what a first tuned table needs:

- every candidate (or a seeded --candidate-sample) is measured once
  through the stock mp_tuner, alongside the auto-select incumbent and the
  shipped tile defaults;
- a winner is published only if it beats the incumbent by more than
  MHA_FWD_INDIFFERENCE_DELTA (2%), and every written row is proven by a
  fresh-process dispatch probe;
- shapes that already have a row for this GPU are skipped unless --all
  is given; --all --compare --update_improved re-tunes them and replaces
  a row only if the public operator got faster;
- a re-tune that finds auto-select best warns that the existing row still
  overrides it, since writing no row cannot remove one.

Candidates that fail the accuracy check are labelled mismatch rather than
crash, against the --errRatio the run was given.
amd-bartgips added a commit that referenced this pull request Sep 28, 2026
mp_tuner, base_tuner's measurement_kwargs and _read_csv fix, the
FileBaton zombie check and their tests return to #5592's version. They
now live in #5846, where the other tuners can use them, and the MHA
tuner no longer needs them. base_tuner keeps only the gpu_model filter
in run_config that tuned MHA rows rely on.
amd-bartgips added a commit that referenced this pull request Sep 28, 2026
The gate against the incumbent used 2%, while --all --compare
--update_improved keeps a row only if the public operator got at least
--min_improvement_pct (3%) faster. A 2-3% winner could be published by
the search and then be refused by the protected re-tune. The gate now
uses the same 3% and the same "at least" comparison, as
MHA_FWD_MIN_IMPROVEMENT, which is also the value the shared
PromotionPolicy in #5846 uses.
@amd-bartgips
amd-bartgips force-pushed the users/bartgips/tuning-shared-infra branch from 04a2090 to f36f909 Compare September 28, 2026 14:17
amd-bartgips added a commit that referenced this pull request Sep 28, 2026
#5846 now has one improvement bar (min_improvement_pct, 3%) for
--compare and the gate, and leaves the finalist settings to the search
that uses them. This branch adds them back as FinalistPolicy (finalists,
rounds, significance_sigma) next to RacePolicy, and changes what the bar
means:

- gate_against_incumbent takes a noise_pct floor and promotes only at
  max(min_improvement_pct, noise_pct). Before, the standard-error bar
  replaced the delta, so enough finalist rounds would publish a real but
  tiny win that --compare would then refuse.
- The floor is two combined standard errors under the exhaustive
  strategy, and the race's delta under --strategy race.
- The MHA tuner reads the bar from --min_improvement_pct, so one flag
  sets it for both the search and --compare. --delta defaults to it and
  is only needed to make the race's indifference zone differ.
- standard_error_bar and every margin are in percent, as --compare
  reports them; the evidence manifest records margin_pct, bar_pct and
  noise_floor_pct per promotion, and the promotion bar per run.
Moved unchanged from the MHA forward tuner branch (#5762) so that other
tuners can use them without waiting for it:

- mp_tuner: optional per-candidate status and detail (return_status), a
  per-candidate progress queue with result_callback, a faulted group
  keeping the candidates it had already measured, MpTunerTask, broader
  accelerator-fault detection, and a pool restart on a stale PID map.
- TunerCommon.measurement_kwargs, so --warmup and --iters reach the timed
  call instead of run_perftest's defaults.
- _read_csv builds a new frame rather than assigning into a slice.
- FileBaton treats a zombie lock holder as dead, so a builder that exited
  without releasing its lock no longer wedges the module build.
A caller that does not pass return_status picks the fastest finite
latency and cannot tell a partial group from a complete one. Keeping the
candidates measured before a fault would let it publish the winner of a
search that stopped early, so those callers get every candidate of the
group back as failed, as before. Callers with return_status keep the
measured candidates and see the crash or timeout status on the rest.

test_mp_tuner_fault runs one shape group with a faulting candidate
through the real pool on one GPU, in both call forms.
aiter/utility/tuning_policy.py holds the thresholds whose justification
does not mention a kernel family:

- MeasurementPolicy: warmup, iters, errRatio and timeout, which
  TunerCommon.ARG_DEFAULTS now reads (same values as before).
- COMPARE_MIN_IMPROVEMENT_PCT: the --compare --update_improved bar.
- PromotionPolicy: the indifference delta a challenger has to clear over
  the incumbent, and the finalist count and rounds for re-timing.
- gate_against_incumbent: the promote-or-retain decision on those
  numbers, for tuners that measure the incumbent themselves.

Standard library only, so the runtime and CPU tests can import it. A
family that needs a different threshold overrides it with
dataclasses.replace in its own module.
--compare --update_improved and gate_against_incumbent both decide
whether a faster configuration is worth replacing what a shape runs
today, but read two different numbers (3% and 2%). They now read one:
PromotionPolicy.min_improvement_pct, 3.0, named after the
--min_improvement_pct flag whose default it already was.

- The gate works in percent and promotes at "at least" the bar, as
  --compare does, and reports margin_pct.
- COMPARE_MIN_IMPROVEMENT_PCT and indifference_delta are gone; neither
  was on main.
- finalists and finalist_rounds move to the follow-up that times
  finalists (#5849), since nothing on main uses them.
- batch comes from a new RunPolicy, so no TunerCommon default is a bare
  literal any more. It sits apart from MeasurementPolicy because it
  changes how often the tuned CSV is written, not which candidate wins.

No default a tuner sees changes.
test_mp_tuner_logic and test_mp_tuner_fault imported Triton before torch,
and the start-time test switched its pool from spawn to fork, both for a
dynamic-loader abort in spawned workers. Main runs the same test, and the
worker-PID test from #5841, under spawn with no Triton import, and the
abort does not reproduce on MI355X. The tests now use spawn, the start
method mp_tuner forces, and a CPU test no longer needs Triton installed.
Worker death is now caught by #5841's check, which this branch rebases
onto. Without it, the immediate pool restart on a stale PID map
resubmitted the dead group with a fresh start time, so a candidate that
kills its worker restarted the pool forever.

- The worker-exited branch drains the progress queue before failing the
  group, as the other failure branches do, so a return_status caller keeps
  the candidates the worker measured before it died. The detail names the
  exited worker.
- With return_status, the candidates behind a fault come back as not_run
  instead of carrying the fault's status. The callback stream is unchanged:
  it still receives only the candidate that failed.
- The mp_tuner docstring describes both result forms, every status and
  result_callback.
- test_mp_tuner_fault adds a candidate that exits its worker with another
  shape queued behind it, in both call forms. With the dead-worker check
  disabled it fails at its deadline after repeated pool restarts. The file
  now runs in the per-PR shards (67 s on one MI355X).
@amd-bartgips
amd-bartgips force-pushed the users/bartgips/tuning-shared-infra branch from f36f909 to 839de6b Compare September 30, 2026 10:56
work_group catches an error raised outside worker(), such as a later
candidate's gen_data or reference failing, and returns a result list
itself. That list failed every candidate of the group, including those
already measured, and put all of them on the progress queue. The parent
saw a normal return, so its partial-result handling never ran: a
return_status caller lost the measured candidates, a checkpointing caller
received a second, crash result for each of them, and the candidates
behind the abort were checkpointed as tried, so a resume skipped them.

work_group now builds those results with _failed_group_results, the
same rules the parent applies to a group the pool lost: with
return_status, measured candidates keep their results, the one that
aborted gets the failure, and the rest are not_run; without it, the whole
group fails as before. Only the aborting candidate is published.

test_mp_tuner_fault adds a group whose second candidate fails in gen_data
(reported by Araceli); the typed case failed before this change.
amd-bartgips added a commit that referenced this pull request Oct 1, 2026
The race, resume journal, finalist rounds, standard-error gate, evidence
manifest and re-tune incumbent move to a follow-up on top of #5846, so
this PR carries only what a first tuned table needs:

- every candidate (or a seeded --candidate-sample) is measured once
  through the stock mp_tuner, alongside the auto-select incumbent and the
  shipped tile defaults;
- a winner is published only if it beats the incumbent by more than
  MHA_FWD_INDIFFERENCE_DELTA (2%), and every written row is proven by a
  fresh-process dispatch probe;
- shapes that already have a row for this GPU are skipped unless --all
  is given; --all --compare --update_improved re-tunes them and replaces
  a row only if the public operator got faster;
- a re-tune that finds auto-select best warns that the existing row still
  overrides it, since writing no row cannot remove one.

Candidates that fail the accuracy check are labelled mismatch rather than
crash, against the --errRatio the run was given.
amd-bartgips added a commit that referenced this pull request Oct 1, 2026
mp_tuner, base_tuner's measurement_kwargs and _read_csv fix, the
FileBaton zombie check and their tests return to #5592's version. They
now live in #5846, where the other tuners can use them, and the MHA
tuner no longer needs them. base_tuner keeps only the gpu_model filter
in run_config that tuned MHA rows rely on.
amd-bartgips added a commit that referenced this pull request Oct 1, 2026
The gate against the incumbent used 2%, while --all --compare
--update_improved keeps a row only if the public operator got at least
--min_improvement_pct (3%) faster. A 2-3% winner could be published by
the search and then be refused by the protected re-tune. The gate now
uses the same 3% and the same "at least" comparison, as
MHA_FWD_MIN_IMPROVEMENT, which is also the value the shared
PromotionPolicy in #5846 uses.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant