Skip to content

Correct 45.9's placement resolution: half of it was resolved in the wrong direction - #1811

Merged
justinchuby merged 2 commits into
mainfrom
squad/leon-p45-placement-scope
Aug 23, 2026
Merged

justinchuby merged 2 commits into
mainfrom
squad/leon-p45-placement-scope

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 23, 2026 •

Copy link
Copy Markdown
Owner

Sebastian's #1729 write-up says every t≥8 row either of us published pre-6e8c31ebd was taken on 8 physical cores, so they all need re-taking. My §45.9/45.10 retraction had listed placement as an open confound on my own rows, marked "not yet verified", pending a /proc read queued behind the host owner.

I went to close it. The first draft of this PR closed it in the wrong direction, and Opus review caught that before merge. Both halves are below, because the mistake is the more useful half.

What is actually established

The SPMD decode pool — the one #1729 re-placed — is never entered by a single-node SDPA/MHA graph. From routing code, no timing, no cores:

  • Production SDPA forks on decode_pool_active() (sdpa.rs:1287); only the true arm reaches decode_parallel_output_row_blocks. The false arm goes to task_runtime::chunk_runs_mut.
  • decode_pool_active() is gated on thread-locals IN_SPMD_SCOPE (matmul_nbits.rs:4901) / IN_NUMA_SCOPE (:4891), set only inside with_decode_pool_scope — not process-globals, so no other thread's decode loop can turn it on for the benchmark's engine thread.
  • The only production callers of that scope are under native_decode/ (cpu.rs:445/578/584, proposer.rs:325). bench_generic never mentions native_decode.

That much stands.

What I got wrong

I wrote "the cpus 0-15 mask never governed these measurements". What I had shown was that that pool never governed them. Different statement.

There is a second, independent confinement mechanism, and it did apply:

  • CpuExecutionProvider::initialize (provider.rs:340) calls bound_process_to_decode_budget() (matmul_nbits.rs:4354), which — budget set, no explicit ONNX_GENAI_CPU_DECODE_AFFINITY — performs a real sched_setaffinity. Its doc comment states the reach: "threads spawned afterwards … inherit the mask".
  • bench_generic reaches it: InferenceSession::load → executor/platform.rs:8 → ep.initialize(...). And it sets ONNX_GENAI_CPU_DECODE_THREADS from --native-threads at :670-673, before any session exists. So on every row where ab.py passed a width, the benchmark process confined itself.

It is keyed on the same knob. Sebastian documented this mechanism himself in the #1746 addendum to 2026-08-22-decode-width-scaling.md, including that his int4_decode_loop_ab escapes it by never calling initialize(). Mine is the opposite case. I had read that addendum as being about his harness and never checked mine against it.

The third asymmetry this turned up

Affinity is inherited at thread creation, and the two arms are built on opposite sides of the mask: bench_generic constructs ort_session at :691, then InferenceSession::load at :703 — and it is that load which confines the process. ORT's intra-op pool is spawned before the confinement; the native EP's threads after it. If ORT's pool is built at session construction, the paired runs handed ORT all 32 logical CPUs while the native arm ran on T.

Same direction as the ORT spin-wait tax, and it also grows with T. Recorded as suspected, not observed — settled by reading Cpus_allowed_list per thread from /proc/<pid>/task/*/status on a live paired run, grouped by thread name. Needs a quiet host; queued, not claimed.

Net effect, which is the inverse of the first draft

These rows are not placement-clean and must not be treated as such. The #1729 mechanism specifically is excluded; the self-imposed T-CPU mask is not, and its interaction with the ORT arm is open. They need re-taking regardless, for the spin-tax reason, which was always sufficient on its own.

I have deliberately not re-run anything: host is Sebastian's (qemu aarch64 suite + decode_gap_park, runnable 33).

Also

The file carried two ### 45.9 headings. Since this PR is entirely about which §45.9 claim was wrong, and the cross-references at :4716/:4722 cite it by number, the retraction is renumbered 45.10 and those references updated. 2026-08-22-sdpa-fanout.md cites the other 45.9 and is unaffected.

Rules added

  • Disproving one mechanism is not disproving the class. The second mechanism was one grep for sched_setaffinity away, in a file I had already opened, described in a teammate's document I had already read.
  • The direction that says your data is fine is the one that propagates. An over-stated confound costs a re-measurement; an under-stated one ships a wrong number under somebody else's name.

Scope

Documentation only — one section of the benchmark ledger plus a heading renumber. No code, no measurement, zero cores.

justinchuby and others added 2 commits August 23, 2026 06:04
45.9 left "whether the cpus 0-15 mask reached the single-node MHA graphs" as
not yet verified, pending a /proc read against a live bench process. It did
not need one. It is answerable from the routing code with no timing at all,
and the answer is that it did not.

The production SDPA route forks on decode_pool_active() and only the true arm
reaches the SPMD pool that decode_affinity places. That flag is two
thread-locals, IN_SPMD_SCOPE and IN_NUMA_SCOPE, set only inside
with_decode_pool_scope, whose only production caller is the generation loop in
native_decode/cpu.rs. bench_generic never mentions native_decode -- it drives
graphs through onnx-runtime-session -- so on a single-node graph the flag is
false and the pool is never entered.

The width knob does still bite, because resolve_width honours
ONNX_GENAI_CPU_DECODE_THREADS deliberately, but that pool is capped to
physical cores and never pinned, which is the opposite of the compact mask.
My earlier inference -- same knob, so it cannot be assumed away -- was sound
as caution and wrong as fact: the knob sizes two different pools and only one
of them is placed.

This matters now because decode rows taken before 6e8c31e genuinely do need
re-taking, and without this the 45.7-45.9 SDPA rows would have been swept into
that set. They are not affected by placement. They still need re-taking for
the ORT spin-tax reason, which is sufficient alone.

A confound is a claim like any other and needs the same standard of proof in
both directions. Left recorded as open when it was decidable from three lines
of routing code, it would have sent somebody re-measuring rows that were never
affected.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… section

The first draft of this correction resolved the placement confound in the
wrong direction. It proved the SPMD decode pool (the one #1729 re-placed) is
never entered by a single-node SDPA/MHA graph -- true, and verified from
routing code -- and then concluded "the cpus 0-15 mask never governed these
measurements". Those are not the same statement.

A second, independent confinement mechanism did apply:
bound_process_to_decode_budget (matmul_nbits.rs:4354), called from
CpuExecutionProvider::initialize (provider.rs:340), which bench_generic
reaches through executor/platform.rs:8 on every native session load. It is
keyed on the same ONNX_GENAI_CPU_DECODE_THREADS the harness sets, and applies
a real sched_setaffinity. Sebastian documented it in the #1746 addendum,
noting his own bench escapes it by never calling initialize(); mine does not.

That also surfaced a third harness asymmetry: affinity is inherited at thread
creation, and bench_generic builds ort_session (:691) before the native load
(:703) that applies the mask. ORT's pool is therefore spawned outside the
confinement and the native EP's threads inside it. Recorded as suspected, not
observed, with the /proc experiment that settles it.

Net effect on the re-measurement queue is inverted: these rows are NOT
placement-clean and must not be treated as such.

Also renumbers the second "### 45.9" to 45.10 -- the file carried two
sections with that number, and both this change and the two cross-references
at 4716/4722 cite it by number.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby justinchuby changed the title Resolve 45.9's open placement confound, and narrow it Correct 45.9's placement resolution: half of it was resolved in the wrong direction Aug 23, 2026
@justinchuby
justinchuby merged commit 1bb3afb into main Aug 23, 2026
11 checks passed
@justinchuby
justinchuby deleted the squad/leon-p45-placement-scope branch August 23, 2026 06:47
justinchuby added a commit that referenced this pull request Aug 23, 2026
Follow-up to #1811, which recorded a harness asymmetry as **suspected,
not observed**. It is now observed, and it is the largest of the three.

## The claim #1811 left open

`bench_generic` constructs `ort_session` at `:691` and *then*
`InferenceSession::load` at `:703` — and it is that load which reaches
`bound_process_to_decode_budget()` and applies a process CPU affinity
mask. Affinity is inherited at thread creation, so ORT's intra-op pool
is spawned **outside** the confinement and the native EP's threads
**inside** it.

## Observed

`Cpus_allowed_list` per thread from `/proc/<pid>/task/*/status`, live
paired runs, no outer `taskset`:

| `T` | confined to the budget mask | **unconfined, `0-31`** |
|---|---|---|
| 2 | 3 × `bench_generic`, 1 × `nxrt-task-0` | **1** |
| 4 | 5 × `bench_generic`, 3 × `nxrt-task-N` | **3** |

Every `nxrt-task-N` worker carries the mask. The unconfined threads are
unnamed — so they show the process name — and there are exactly **`T −
1`** of them: ORT's intra-op pool, which reuses the calling thread.

They are not Rayon: `prefill_worker_name` names those `nxgn-prefill-N`,
and the mask is applied *before* `build_global` (`matmul_nbits.rs:4380`
vs `:4396`), so a Rayon worker would be both named and confined. They
are not inter-op (`--ort-inter-threads` defaults to 0). And **being
unconfined is itself the proof of spawn order** — had ORT built its pool
lazily at first `Run`, after the mask, these threads would carry it.

At `T = 16` the mask is measured but the census is not (that launch was
`--native-only`, so it had no ORT arm). Carrying the `T − 1` rule across
gives a native arm on 16 physical cores against ~15 ORT workers roaming
all 32 logical CPUs — including the SMT siblings of the cores the native
arm cannot leave. Marked in the text as inferred, not counted.

This does not replace the spin-tax reason, it sharpens it, and it grows
with `T` — the shape §45.8 read as a native scaling failure.
`--ort-intra-threads T` does **not** equalise it: it equalises thread
*counts*; the CPU *sets* differ by construction.

## Also settles the mask value #1811 declined to assert

| `--native-threads` | mask |
|---|---|
| 2 | `[0, 2]` |
| 4 | `[0, 2, 4, 6]` |
| 16 | `[0, 2, …, 30]` |

`scatter_across_cores` picks physical-core leaders, so the
**two-workers-per-core** pathology of #1729 is excluded on both
mechanisms.

**But not "the mask places well"** — an earlier draft said that, and
§26.2 of this same document refutes it: that spread measured **worse**
(0.133 ms vs 0.079 ms) for a pinned multi-worker decode pool, because
straddling two CCXs costs more than SMT sharing when the working set
fits in one CCX's L3. Both hold. That penalty is paid by a *pool*,
through cross-CCX barrier traffic, and these rows never build one. Now
stated as a trade-off scoped to these rows, pointing at §26.2 for anyone
quoting it for a pool workload.

## Verified mitigation

`ONNX_GENAI_CPU_DECODE_AFFINITY=off` makes
`explicit_decode_affinity_requested()` true (it tests non-empty, not the
value) and the auto-mask stands down; `build_global` still runs, so pool
*size* is unchanged and only the confinement is dropped. Re-probed at `T
= 4`: all 11 threads read `0-31`, symmetric. Re-measurements should set
it, alongside `--native-only` or `--ort-intra-threads 1`.

**Knock-on for #1792**, which reports this knob as *inert* on the
default SPMD path: it is not inert here — it is the only switch that
disables the process self-confinement. The same variable is dead in one
mechanism and load-bearing in the other, which is a worse user-facing
story than either finding alone. Sebastian's call whether that belongs
in the issue body.

## Review

Opus review confirmed the thread identification, the `T−1` arithmetic
against both data points, the spawn-order argument and the mitigation,
and caught two places where this revision repeated — in the safe
direction — the over-claiming pattern of the two revisions it corrects.
Both are fixed in `c195b4fcd`: the `T=16` census is hedged, and the
§26.2 contradiction is reconciled rather than left standing.

## Rule added

Process-wide state applied during setup makes **construction order part
of the experiment design**. "Which arm was built first" silently decided
which arm got the whole machine, and nothing in the harness expresses
that as a decision — it is the order two `let` bindings happen to appear
in.

## Cost

`/proc` reads, not timings, so contention cannot corrupt them — taken on
a busy host (runnable 33) rather than queued behind it. Widest arm was
one `--runs 1` native-only launch. No benchmark, no millisecond claimed.
Documentation only.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
…d a knob that no longer exists (#1822)

Follow-up to #1173, correcting two defects I shipped in it and repairing
the rule they undermined. Docs, one ledger string, one new test, one new
script. No production kernel or routing change.

## 1. The ledger named a route gate that had already been deleted

`PLAN[MatMulF32].shape_gate` said the native `SimdX86` route "gates M=1
on `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` (default off, #1116)".

#1183 shipped that GEMV on by default and removed the probe. `git
merge-base --is-ancestor 5417d04 bdb4599` confirms it landed
**before** #1173 merged — so the ledger was wrong the day it landed.
Today `sgemm_simd` calls `sgemm_simd_variant(a, b, c, m, k, n, true)`
unconditionally and `use_m1_gemv` is a plain parameter that only the
in-process A/B harness passes as `false`. No environment variable
reaches that route.

`docs/performance/CPU_MATMUL_ASSIGNMENT.md:559` already recorded the
correct fact ("It is measured now, and the route is the default. There
is no env probe on the dispatch any more"). Two files in the same
directory disagreed and nothing compared them.

**Now guarded.**
`ledger_prose_only_names_environment_variables_that_still_exist`
requires every `NXRT_*` / `ONNX_GENAI_*` token in the ledger's prose to
still exist as a string literal in the crate's sources. It cannot check
that the description is *right*, only that the knob is *real* — which is
the half that goes stale silently.

Mutation-verified, not just observed green:

```
matmul_f32: ledger prose names environment variable `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV`,
but no source file in this crate contains the literal "ONNX_GENAI_CPU_MM_SIMD_M1_GEMV".
```

## 2. The doc published a toggle A/B that could not have been run

#1173 carried a table captioned **"same binary, same session, toggle the
only difference"**, reporting `decode 1×2048×2048` at 0.146 with
`ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` off against 0.337 with it on, and
called turning it on "the obvious next slice".

Nothing reads that variable. Setting it measures the same route twice;
it cannot produce two different columns. The table is withdrawn and the
retraction kept in the text rather than quietly deleted.

This is the failure mode the document's own graduation rule warns about
— **an arm that was not on the route it was labelled with** — committed
by the document that wrote the rule. It survived review because a
plausible number in a well-formed table is not self-evidently
unmeasured. Readers are pointed at `bench_f32_gemm_ab`, which holds the
route as a function parameter and carries the M≥2 rows as a built-in
control.

## 3. The gap table is re-measured and the ≥5% rule is repaired

The old table was one unguarded invocation per row at an unstated width,
taken before the decode-placement corrections (#1729, #1794, #1811) —
i.e. when the decode pool put 16 workers on 8 physical cores.

New harness: `scripts/bench_native_vs_mlas_width.py`. Arms interleaved
rep by rep so host drift lands on both equally; per-rep `os.wait4`
CPU-efficiency guard adapted from #1809; six reps per arm; two widths.
Raw verdicts, spreads and discards are all reported rather than
summarised away.

**Three findings, all about method rather than kernels.**

| | narrow (6 cores, 1 L3) | wide (32 logical CPUs) |
|---|---|---|
| `matmul_f32 16×512×512` | 1.581, spread 41% | 0.866, spread 134% |
| `matmul_f32 decode 1×2048×2048` | 1.117, spread 21% | 0.934, spread
13% |

- **Two cases change verdict on width alone.** Same binary, same
half-hour, only the CPU mask differs. `x86_sgemm` parallelises over
column strips and MLAS declines to parallelise some shapes, so
interleaving the two *routes* inside one process does not protect the
ratio — it changes both at once.
- **`16×512×512` disagrees with itself on both arms**, alternating
`keep-mlas` / `native-graduates` from a byte-identical binary. **One
more run of the old table could have graduated a route on this row.**
- **The narrow arm is more trustworthy despite having fewer cores** —
spreads 4–42% against 5–134%, and it lost no reps to the guard.
Isolation beat parallelism.

**Softmax now decomposes cleanly**, because no vendored MLAS kernel has
changed since #1173 (the only `mlas-sys` edits are the additive
straggler handshake in `work_stealing_pool.rs`, #828/#1714, which adds
waiting). At matched width the MLAS control arm is stationary to within
4% while native improved **1.24–1.27×** — matching #1416's claim for the
row kernel. The f32 GEMM rows get no such attribution and now say so
explicitly: their control moved **2.0× the wrong way**, so only the
current ratio at a stated width is defensible.

**The rule gains what it lacked**: spread must be smaller than the
claimed win; reps that did not get the CPU are discarded rather than
averaged; a verdict is valid only at a stated width. Under it, `decode
1×2048×2048` — the first f32 GEMM case to show a real native win —
**still does not graduate**: it costs more CPU (cpu_ratio 0.875), does
not hold at 32 threads, and its 21% spread exceeds its 12% win.

## The width claim is verified, not asserted

#1815 landed while this was in progress and observed the neighbouring
`bench_generic` harness spawning its ORT arm *outside* the affinity
confinement it applied to the native arm. That hazard applies to any
`taskset` claim, including mine, so I checked it instead of trusting it
— sampling `Cpus_allowed_list` from `/proc/<pid>/task/*/status` 40×
across a live narrow-arm run:

```
'16,20,22,26,28,30': 478 observations
  native_vs_mlas- 273, mlas-sys-ws-0..4 39 each, nxrt-task-0..4 2 each
'0-31': 1  (the taskset process itself, before exec)
```

Both routes confined identically; no thread escaped. The rule now
requires this check.

## Validation

- `dispatch_ledger` **17/17**, including the new falsifier, after
merging latest `main`.
- `default_artifacts_are_mlas_free` **9/9** — the no-MLAS-in-defaults
invariant is untouched.
- `cargo clippy -p onnx-runtime-ep-cpu --lib --all-targets` clean;
`cargo fmt --check` clean.
- Normal merge of `origin/main` (`aee2b9d11`), no rebase, no conflicts.

## Limitations

- The narrow arm is six cores on one L3 of one x86-64 host. Nothing here
transfers to aarch64 or to a two-socket box.
- The `activations erf 1 Mi` row shows native 13.5% slower at matched
width. The nearest scatter figure is the wide arm's 8% spread, but that
is a spread of *ratios* against a move in a *native time*, so the two
are not strictly commensurable. Its MLAS control also moved 11%.
**Flagged for pinned re-measurement, not reported as a regression.**
- The wide arm was taken with ~4–5 cores of unrelated load present. That
is stated in the doc rather than hidden, and it is why its spreads are
wider; the guard reports which reps were discarded instead of pretending
the host was quiet.
- No production behaviour changes here, so there is no performance claim
to make about the shipped artifact.

Refs #1173, #1183, #1809, #1815, #1416.



## Independent review, and what it changed

An independent adversarial review of the full diff returned **no
blockers** — it confirmed the ancestry argument behind the retraction,
the stationary-control premise for the softmax attribution, and that the
headline case is correctly *refused* by the rule (21% spread against a
12% win). It also found seven real defects, all now fixed in
`f0323f9ed`.

The one that mattered most was in the new test. It only proved the
variable name appeared *somewhere* in the crate, so a variable whose
read site had been deleted but whose name survived in an
`EnvVarGuard::set(...)` line would still have passed — which is the
precise shape of the defect this PR exists to correct. The test now
requires the matching line to be an `env::var(` / `env::var_os(` read or
an `_ENV: &str =` binding.

Verified by mutation in **both** directions:

| mutation | before | after |
|---|---|---|
| reinsert retired `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` into ledger prose |
fails ✅ | fails ✅ |
| retire the two real `NXRT_CPU_GEMM_BACKEND` reads, leaving the literal
only in test guards | **passes ❌** | fails ✅ |

The remaining six were prose defects in the doc: a stated spread range
that contradicted its own table's 82% row, "within 4%" against a table
reading −4.2%, a narrow-arm ratio fused with a wide-arm attribution, a
spread quoted as 7.5% that was 8% *and* compared against an
incommensurable quantity, the CPU-efficiency guard oversold as "what
makes this table measurable at all" (in-process interleaving is what
protects the ratio; the guard catches only *differential* descheduling),
and a one-directional provenance argument standing in for the direct
control measurement that actually carries the softmax attribution.

**Two further defects I found myself while checking the tables against
each other**, neither raised by the review:

- The `ratio` column is a median of per-rep ratios while the `ns/unit`
columns are medians of times. Medians do not distribute over division,
so every row looked internally inconsistent to anyone who tried to
divide it out (`0.0684 / 0.0617 = 1.109` against a stated `1.117`). Now
documented, along with why the per-rep form is the correct one to quote:
it pairs each MLAS invocation with the native invocation it was
interleaved against, which is the entire point of interleaving. The
then→now figures are relabelled as quotients of medians.
- "wider than nine of the twelve wide-arm rows" was eleven of twelve.

## Adopting #1814

`aee2b9d11` (#1814) landed on `main` while this was in review, and it
closes the exact hole the review found in the guard this document
recommends. A differential CPU-efficiency check cannot see contention
that lands evenly on both arms; #1814's confined-set meter reads busy
jiffies on the process's own `Cpus_allowed_list` and subtracts the
process's own CPU, so foreign load shows up directly. The rule now
points at it, and the tables here are explicitly marked as predating it
and guarded by the weaker method.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
… the box (#1885)

Closes Gaff's blocking item on #1806.

## The defect, in his words and then in mine

> `REAP_DIR` is unguarded. … One kill in that window orphans the
directory, after which `reap_if_dead` returns 1 forever and **no stale
lock can ever be reaped again** — the box wedges behind a dead holder
and `acquire` just reports busy, indistinguishable from legitimate
occupancy.

Correct, and the framing is the part worth keeping: **the wedge this
script exists to prevent, relocated into the mechanism that prevents
it.** The lock has an anchor pid, a start time and a reaper. Its own
mutex had none of the three — a bare `mkdir "$REAP_DIR"` released by
`rmdir`.

#1811 bounded it by age (`REAPER_GRACE=60`). That recovers, but it
recovers *by waiting*, and it bought the mirror-image defect: **age is
not evidence of death.** A reaper that is merely slow — loaded box, a
qemu leg, a stalled `stat` — could have its guard cleared while it was
still inside the critical section. Two acquirers then both reap and both
believe they own the host. That is worse than a wedge, because a wedge
is loud and this is silent, and the numbers on both sides are ruined
without either party knowing.

## What it is now

The guard is built the same way as the lock, for the same reasons:

| property | mechanism | why |
|---|---|---|
| published atomically **with** metadata | stage a populated dir, `mv
-T` | `rename(2)` onto a non-empty dir fails `ENOTEMPTY` → test-and-set,
not last-writer-wins. No kill can produce an unattributable guard,
because there is no instant at which one exists |
| owned | `anchor_pid` + `start_time` in `$REAP_DIR/meta` | pids
recycle; "a process with this number exists" is a different question
from "the reaper is still running" |
| dead owner reclaimed **immediately** | `anchor_alive` | no grace, no
waiting, nothing to wedge behind |
| live owner **never** disturbed | `anchor_alive` outranks age | kills
the double-reap that the age rule introduced |
| released only by its owner | pid check in `reaper_release` | a process
whose guard was reclaimed while descheduled must not delete its
successor's guard on the way out |

`anchor_alive` is now **one** predicate, shared with `holder_alive`.
Four call sites each deciding "is this pid alive?" for themselves
produced two of the four defects in #1830; I am not doing that again in
a second module.

`REAPER_GRACE` survives for exactly one residual class — a guard that
exists with no readable anchor. Staging makes that unreachable from this
script; a stray directory at that path from an older version or a
hand-run `mkdir` can still present it, and for that class no liveness
evidence exists, so age is all there is. The existing test for it stays.

## The kill-in-window test is deterministic, and it cannot pass
vacuously

Requested explicitly, and it is the cell I would have written badly:

```sh
HOSTLOCK_REAPER_STALL=30 $HL acquire --owner victim ... &   # seam holds the section open
...wait for $LOCK.reaper/stalled_pid...                     # the pid INSIDE the window
sig "$stalled" 9                                            # SIGKILL lands in the window every run
chk "killing it in the window really does orphan the guard" ...        # ← the orphan EXISTS
chk "and the orphan still names its dead owner"            ...
$HL acquire --owner roy ...
chk "a guard orphaned by SIGKILL does not wedge the next acquirer" "$(st owner)" "roy"
chk "and recovery is immediate, not after REAPER_GRACE"    "$(...)" "clear"
```

Three things this gets right that the obvious version does not:

1. **Deterministic, not opportunistic.** Racing a real reaper's
microsecond window means the kill lands inside it when the scheduler
feels like it. The seam makes the window as wide as I ask.
2. **It asserts the orphan exists first.** Without that line, all three
recovery checks pass equally well when nothing ever leaked — the
vacuous-coverage failure Gaff flagged on Resch's #1805 the same night,
where every assertion sat behind a condition that could silently be
false.
3. **`cleanup` is not called between the kill and the recovery.** The
harness sweeps `$LOCK.reaper`, so a tidy-looking cleanup there would
erase the orphan and the cell would pass **against the defect it exists
to catch.** Cleanup must not mask the defect; here it is a deliberate
omission with a comment saying so.

The seam is production code, so its inertness is asserted too (R8.4):
unset, the reap completes in under 3 s.

## Falsification

Suite **231 → 247**, green. `shellcheck` clean. Nine mutations, all red:

| # | mutation | result |
|---|---|---|
| P1 | dead owner never reclaimed (the original wedge) | 242 passed /
**2 failed** |
| P2 | age-only rule restored, liveness ignored | 240 / **4** |
| P3 | release without the ownership check | 243 / **1** |
| P4 | guard created in place instead of renamed into place | 242 /
**2** |
| P5 | stall seam always fires | 199 / **45** |
| P6 | `anchor_alive` ignores `start_time` (recycled pid) | 242 / **2**
|
| P7 | rename-then-remove keeps the corpse (clear path) | 246 / **1** |
| P8 | rename-then-remove keeps the corpse (release path) | 246 / **1**
|
| P9 | release checks pid but not start time | 246 / **1** |

**P7–P9 came from review, and both findings are the shape this file
keeps producing: the fix was right and the suite proving it could not
tell.**

*P9* — `reaper_release` compared only `$$` while `reaper_clear_if_dead`
compares pid **and** start time. The asymmetry sat exactly where the
consequence is worst: a recycled pid landing on our number makes us
delete a **live successor's** guard, rather than merely failing to tidy
up our own. This box is at ~1.5M pids in four days, so recycling is not
theoretical. R8.3b forges a successor guard carrying the stalled
reaper's own pid with a start time that is not its, and requires it to
survive.

*P7/P8* — dropping the `rm -rf` from either rename-then-remove left all
244 assertions green, because the harness `cleanup()` sweeps those globs
between cells. In production that is unbounded `.dead.*`/`.rel.*` growth
beside the lock. The new assertion runs **before** cleanup,
deliberately: placed after it, it would pass whether or not the code
tidies up at all — the same "cleanup masks the defect" failure the
kill-in-window cell was written to avoid, sitting one cell to the left
of where I was looking for it.

P4 is a structural assertion rather than a behavioural one, and the PR
says so rather than dressing it up: killing between a `mkdir` and a meta
write is a microsecond window, and a seam wide enough to test it would
be wider than the bug. The suite asserts instead that the guard is never
created in place — weaker evidence, honestly labelled.

## The other two items

**`--gate` stays a start admission only.** No change to its semantics.
The README already carries the split — the lock and the gate decide
whether to **start**, `--expect-cores/--min-efficiency` (#1864) decides
whether to **believe** — with Roy's 52% A/A null as the evidence for why
one instrument cannot do both.

**The harness holds the lock across all arms**, and rows carry
`held_by`, `contended`, `runnable_at_acquire` and now `efficiency`.
Unchanged here, restated in the README because "I sampled `ps` and the
host looked free" was a *between-arms* sample of somebody else's
interleaved A/B.

Also documents that this script is **Linux-only by construction**
(`/proc`, `mv -T`, `stat -c`) instead of leaving the question open. A
portability fallback that degraded liveness to `kill -0` would reap live
holders on whichever platform took it — the one error this script must
never make.

No crate code.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant