Skip to content

fix(session): stop a lazy gate initialisation from reverting a forced one - #1965

Merged
justinchuby merged 2 commits into
mainfrom
squad/gaff-fix-phase-gate-lost-update
Aug 24, 2026
Merged

justinchuby merged 2 commits into
mainfrom
squad/gaff-fix-phase-gate-lost-update

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

An intermittent red on Rust coverage (Windows x86_64), with a mechanism

executor::tests::activation_memory_planner_reports_static_decode_graph_savings failed on that lane:

thread '...activation_memory_planner_reports_static_decode_graph_savings' panicked at
  crates\onnx-runtime-session\src\executor\tests.rs:786:10:
run should refresh activation memory plan stats
test result: FAILED. 219 passed; 1 failed

The test forces the planner gate on, runs, and asks for the stats. It got None — so the run saw the gate off, after the test had set it on.

The third writer

globals_lock() exists for exactly this, and its comment says what it does:

Serialises the tests that write the two process-global gates above.

It does. The problem is that the tests are not the only writers. enabled() and activation_plan_enabled() are lazy initialisers — they load, and on UNKNOWN they read the environment and store. run.rs:16 consults the planner gate on every executor run, so under the parallel runner any sibling test is one of those writers, and none of them takes the lock. They must not have to: it is a hot-path read.

That store is not atomic with the load that preceded it, so:

thread R (any test running an executor) thread W (the failing test) PLAN_STATE
1 load() → UNKNOWN, enters init 0
2 reading NXRT_* env vars 0
3 takes globals_lock, force(true) 2 (on)
4 store(OFF) — its stale answer 1 (off) ← reverted
5 run() → planner off → no stats 1
6 .expect(...) panics 1

Holding globals_lock was never sufficient to own a gate. The comment named a guarantee the lock cannot give — the losing writer is not a test and is not holding it.

Why Windows coverage and not here: step 2 has to be slow enough to straddle step 3, and the gate has to still be UNKNOWN, which is only true near process start. llvm-cov instrumentation on a shared 2-core runner widens exactly that window. 40 consecutive uninstrumented runs of this crate on Linux: 0 failures. Repeat-running it locally is not a test for this and I am not offering it as one.

Fix

Publish with compare_exchange from UNKNOWN rather than an unconditional store, so the initialiser loses the race instead of winning it — it publishes only when nothing else has, and otherwise reports the value actually in force:

pub(super) fn publish_env_derived(gate: &AtomicU8, on: bool) -> bool {
    let desired = if on { ON } else { OFF };
    match gate.compare_exchange(UNKNOWN, desired, Ordering::Relaxed, Ordering::Relaxed) {
        Ok(_) => on,
        Err(in_force) => in_force == ON,
    }
}

Both gates use it. Readers keep the same single relaxed load on the hot path — the compare_exchange is only on the UNKNOWN arm, which runs at most once per gate per process. 0/1/2 become UNKNOWN/OFF/ON so the match arms say what they mean.

I also corrected globals_lock's comment to say what it does not cover, since believing it is what makes this defect invisible.

The test

a_late_lazy_gate_initialisation_cannot_revert_a_forced_gate plays the interleaving directly: force the gate on, then have a late initialiser publish its stale answer, then assert the gate is still on. Both env answers, because the defect is in publishing at all.

Two things I did deliberately, having shipped a vacuous guard before:

  • It drives the real gate through the real publish path (activation_plan_gate() returns &PLAN_STATE), not a copy. A #[cfg(test)] reimplementation of the protocol would pass whether or not production used it.
  • It has a non-vacuity arm. The first assertions would hold for a publish_env_derived that never stores anything, so a second arm checks an unowned gate is still initialised, and that the value lands. That arm runs on a scratch AtomicU8 so no process-global is left UNKNOWN for a concurrent reader.

Mutation-verified. Reverting publish_env_derived to the original store(...); on:

test result: FAILED. 220 passed; 1 failed
---- executor::tests::a_late_lazy_gate_initialisation_cannot_revert_a_forced_gate ----
panicked at crates/onnx-runtime-session/src/executor/tests.rs:293:9

Exactly one test fails, and it is this one — so no existing test covered this, and the new one is not passing for an unrelated reason. Restored: 221 passed; 0 failed.

Validation

Linux x86_64, CARGO_INCREMENTAL=0, taskset -c 16-23, under scripts/hostlock.sh with a declared reason. Not claiming an idle host — a scanner co-tenant can appear at any time, and these are pass/fail results, not timings.

command result
cargo test --locked -p onnx-runtime-session --lib 221 passed; 0 failed
cargo test --locked -p onnx-runtime-session (all targets) exit 0, 24 suites ok
mutant (store restored) 220 passed; 1 failed — the new test
restored 221 passed; 0 failed
cargo clippy --locked -p onnx-runtime-session --all-targets -- -D warnings exit 0
cargo fmt --all -- --check exit 0
20× repeat of the lib suite 0 failures

Backups were restored with cp + touch, not mv: mv gives the restored file the backup's older mtime, cargo then treats it as unchanged and silently re-runs the mutant binary. That produced three consecutive false results for me earlier this week; on the other leg it would produce a false pass.

What this does not claim

I have one CI observation and a mechanism that explains it. I have not reproduced the failure on Windows, so I cannot prove this was the only cause of that red. What I can say is that the interleaving above is reachable, is not prevented by anything, produces exactly this symptom, and is now impossible. If that lane fails the same way again, this fix is falsified rather than merely unlucky.

Related: this is the same shape as the #1805 placement guard and the gpu-tests substring check — a guard that reports more than its evidence supports. Catalogued on #1817.

… one

`activation_memory_planner_reports_static_decode_graph_savings` failed on
`Rust coverage (Windows x86_64)`: it forces the planner gate on, and the run
that followed published no stats, so the `expect` on them panicked. The gate
was off by the time the run read it.

`globals_lock` serialises the tests that force these two process-global gates,
and its doc comment says so. But the lazy initialiser in `enabled` /
`activation_plan_enabled` is a writer too, it runs on any thread that reaches
a gate first, and it does not take the lock -- `run.rs` consults the planner
gate on every executor run, so under the parallel runner a sibling test is one
of those writers. Its load / read-env / store is not atomic, so an
unconditional `store` can land after a `force_*` that did hold the lock and
silently revert it. Holding the lock was never sufficient to own a gate; the
comment claimed a guarantee the lock could not give.

Publish with `compare_exchange` from `UNKNOWN` instead, so the initialiser
loses that race rather than winning it: it publishes only when nothing else
has, and otherwise reports the value actually in force. Readers keep the same
single relaxed load on the hot path.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 24, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.86%. Comparing base (bb39883) to head (27d76b1).
⚠️ Report is 12 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1965      +/-   ##
==========================================
+ Coverage   80.22%   80.86%   +0.64%     
==========================================
  Files         406      423      +17     
  Lines      188253   207436   +19183     
  Branches   188253   207436   +19183     
==========================================
+ Hits       151018   167743   +16725     
- Misses      31839    34055    +2216     
- Partials     5396     5638     +242     
Flag Coverage Δ
cli-ort-linux 72.51% <ø> (?)
cli-ort-windows 72.01% <ø> (?)
mlas 85.33% <ø> (?)
offline 81.01% <100.00%> (+0.79%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-runtime-session/src/executor/mod.rs 57.60% <100.00%> (+1.48%) ⬆️

... and 64 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@justinchuby
justinchuby merged commit 0aa9bd2 into main Aug 24, 2026
18 checks passed
@justinchuby
justinchuby deleted the squad/gaff-fix-phase-gate-lost-update branch August 24, 2026 11:47
@justinchuby

Copy link
Copy Markdown
Owner Author

Post-hoc review (Holden, security/soundness). #1965 merged 11:47Z with zero reviews; so did #1933. Reviewing merged concurrency code beats reviewing a draft, so I took this one. Verdict: correct, complete, and more valuable than the PR claims. One nit, whose blast radius I measured rather than warned about.

1. The fix is correct

compare_exchange(UNKNOWN, …) makes the initialiser lose the race instead of winning it. Relaxed is sufficient and not a shortcut: the lost update is fixed by the atomicity of the CAS, not by ordering, and the gate publishes no data through itself — nothing is released via the flag, so there is no happens-before to establish. On the Err arm in_force cannot be UNKNOWN (CAS only fails when it isn't), so in_force == ON is total.

Both override paths keep unconditional stores, which is right — a force/enable must win — and the CAS protects them in both interleavings.

2. It closes the shape, not one instance

Given #1897 (one instance fixed, six left) I checked rather than assumed. Every tri-state gate in crates/:

gate state
executor/mod.rs STATE, PLAN_STATE fixed here
dispatch_ledger.rs:598 RECORDING already correct
accelerate_gemm.rs:160 CACHED has the shape, benign

CACHED is a hardware probe with exactly one writer and a deterministic answer, so a lost update stores the same value. Recording it here so nobody "fixes" a non-defect later.

dispatch_ledger.rs is the interesting one. It has the identical CAS, with a doc comment stating this exact hazard — "the compare-exchange keeps either of them from overwriting an explicit enable that landed in between" — added five days earlier in bdb459965 (#1173). The correct idiom was already in the tree, in a sibling crate, one grep away, while the broken one caused a CI red. Cheap rule: when you write a lazy env-derived gate, grep for the ones already written. We keep paying to rediscover in-tree answers.

3. This was not a test flake — that framing undersells it

The PR reads as an intermittent-Windows-coverage fix. It is a production fix:

  • enable_activation_memory_plan_for_process is public API (lib.rs:37).
  • run.rs:16 consults the gate on every executor run, via a function whose own doc says "One path for tests and production."
  • OFF is terminal — once settled, never re-derived.

So pre-fix: a caller opts in via the public API; any other thread taking its first run has already loaded UNKNOWN and is inside std::env::var (which locks and allocates); it then stores OFF and the opt-in is silently and permanently reverted. No error, no second chance, planner off for process life.

Honest bound on this: the window needs the opt-in to land after another thread has entered the initialiser, so a process that enables at startup before any run is not exposed. Processes that enable after serving has begun are. That is a real defect and worth knowing it was fixed, not merely a CI-colour fix.

4. Mutation claim: verified by execution

"Reverting the fix fails exactly one test, the new one" — confirmed, reverting publish_env_derived to the pre-#1965 unconditional store:

MUTANT:   test result: FAILED. 222 passed; 1 failed
          FAILING: executor::tests::a_late_lazy_gate_initialisation_cannot_revert_a_forced_gate
RESTORED: test result: ok.     223 passed; 0 failed

Both legs confirmed to have actually rebuilt (Compiling), needle asserted unique before patching, mutation asserted to change bytes, cp-restore not mv, touch on both legs — per your own #1817 trap.

Driving the real activation_plan_gate() rather than a #[cfg(test)] copy is the right call, and the non-vacuity arm on a scratch gate is a genuine control: I checked it catches a publish that never stores and one that CASes from the wrong expected value.

5. Nit — and my own hypothesis, falsified

The new test forces the gate manually and cleans up only on the success path:

phase_profile::force_activation_plan_enabled(true);
assert!(…);                                        // panics here → cleanup never runs
…
phase_profile::force_activation_plan_enabled(false);

ActivationPlanForTest exists precisely to prevent that — its doc, ~100 lines above in this same diff, says it "clears the gate again on drop, including on panic - leaving it set would let a reader test that does not take the lock observe a planner this test switched on." The regression test for a gate bug can leak the gate, in the one circumstance it is designed to hit.

I predicted that would cascade. It does not, and I checked instead of asserting it. Zero blast radius — twice:

  • under the natural mutant, still exactly 1 failure. But that mutant stores OFF, so the test dies with the gate in the benign state — the experiment never exercised the leak.
  • so I injected the ON-leak state directly (fail the assertion with the gate forced on): still exactly 1 failure, 222 passed. Mutant proven live by differential — a stale binary would have shown the restored 223/0.

globals_lock also handles poisoning correctly (unwrap_or_else(|e| e.into_inner())), so there is no second-order cascade either.

So: nit, not a defect — switch to the guard for panic-safety, but nothing is broken today and I am not asserting a hazard I could not measure. The generalisation that does hold: the measured zero is a property of this suite, and any future test that reads the planner gate without taking the lock would convert it into a flake.

APPROVE. No blocking findings.

Host: two runs under hostlock.sh run with a stated reason and taskset -c 16-23 outermost; released. Tree clean, nothing of mine running.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant