Repository navigation
Correct 45.9's placement resolution: half of it was resolved in the wrong direction - #1811
Merged
Merged
Conversation
45.9 left "whether the cpus 0-15 mask reached the single-node MHA graphs" as not yet verified, pending a /proc read against a live bench process. It did not need one. It is answerable from the routing code with no timing at all, and the answer is that it did not. The production SDPA route forks on decode_pool_active() and only the true arm reaches the SPMD pool that decode_affinity places. That flag is two thread-locals, IN_SPMD_SCOPE and IN_NUMA_SCOPE, set only inside with_decode_pool_scope, whose only production caller is the generation loop in native_decode/cpu.rs. bench_generic never mentions native_decode -- it drives graphs through onnx-runtime-session -- so on a single-node graph the flag is false and the pool is never entered. The width knob does still bite, because resolve_width honours ONNX_GENAI_CPU_DECODE_THREADS deliberately, but that pool is capped to physical cores and never pinned, which is the opposite of the compact mask. My earlier inference -- same knob, so it cannot be assumed away -- was sound as caution and wrong as fact: the knob sizes two different pools and only one of them is placed. This matters now because decode rows taken before 6e8c31e genuinely do need re-taking, and without this the 45.7-45.9 SDPA rows would have been swept into that set. They are not affected by placement. They still need re-taking for the ORT spin-tax reason, which is sufficient alone. A confound is a claim like any other and needs the same standard of proof in both directions. Left recorded as open when it was decidable from three lines of routing code, it would have sent somebody re-measuring rows that were never affected. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… section The first draft of this correction resolved the placement confound in the wrong direction. It proved the SPMD decode pool (the one #1729 re-placed) is never entered by a single-node SDPA/MHA graph -- true, and verified from routing code -- and then concluded "the cpus 0-15 mask never governed these measurements". Those are not the same statement. A second, independent confinement mechanism did apply: bound_process_to_decode_budget (matmul_nbits.rs:4354), called from CpuExecutionProvider::initialize (provider.rs:340), which bench_generic reaches through executor/platform.rs:8 on every native session load. It is keyed on the same ONNX_GENAI_CPU_DECODE_THREADS the harness sets, and applies a real sched_setaffinity. Sebastian documented it in the #1746 addendum, noting his own bench escapes it by never calling initialize(); mine does not. That also surfaced a third harness asymmetry: affinity is inherited at thread creation, and bench_generic builds ort_session (:691) before the native load (:703) that applies the mask. ORT's pool is therefore spawned outside the confinement and the native EP's threads inside it. Recorded as suspected, not observed, with the /proc experiment that settles it. Net effect on the re-measurement queue is inverted: these rows are NOT placement-clean and must not be treated as such. Also renumbers the second "### 45.9" to 45.10 -- the file carried two sections with that number, and both this change and the two cross-references at 4716/4722 cite it by number. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 23, 2026
Follow-up to #1811, which recorded a harness asymmetry as **suspected, not observed**. It is now observed, and it is the largest of the three. ## The claim #1811 left open `bench_generic` constructs `ort_session` at `:691` and *then* `InferenceSession::load` at `:703` — and it is that load which reaches `bound_process_to_decode_budget()` and applies a process CPU affinity mask. Affinity is inherited at thread creation, so ORT's intra-op pool is spawned **outside** the confinement and the native EP's threads **inside** it. ## Observed `Cpus_allowed_list` per thread from `/proc/<pid>/task/*/status`, live paired runs, no outer `taskset`: | `T` | confined to the budget mask | **unconfined, `0-31`** | |---|---|---| | 2 | 3 × `bench_generic`, 1 × `nxrt-task-0` | **1** | | 4 | 5 × `bench_generic`, 3 × `nxrt-task-N` | **3** | Every `nxrt-task-N` worker carries the mask. The unconfined threads are unnamed — so they show the process name — and there are exactly **`T − 1`** of them: ORT's intra-op pool, which reuses the calling thread. They are not Rayon: `prefill_worker_name` names those `nxgn-prefill-N`, and the mask is applied *before* `build_global` (`matmul_nbits.rs:4380` vs `:4396`), so a Rayon worker would be both named and confined. They are not inter-op (`--ort-inter-threads` defaults to 0). And **being unconfined is itself the proof of spawn order** — had ORT built its pool lazily at first `Run`, after the mask, these threads would carry it. At `T = 16` the mask is measured but the census is not (that launch was `--native-only`, so it had no ORT arm). Carrying the `T − 1` rule across gives a native arm on 16 physical cores against ~15 ORT workers roaming all 32 logical CPUs — including the SMT siblings of the cores the native arm cannot leave. Marked in the text as inferred, not counted. This does not replace the spin-tax reason, it sharpens it, and it grows with `T` — the shape §45.8 read as a native scaling failure. `--ort-intra-threads T` does **not** equalise it: it equalises thread *counts*; the CPU *sets* differ by construction. ## Also settles the mask value #1811 declined to assert | `--native-threads` | mask | |---|---| | 2 | `[0, 2]` | | 4 | `[0, 2, 4, 6]` | | 16 | `[0, 2, …, 30]` | `scatter_across_cores` picks physical-core leaders, so the **two-workers-per-core** pathology of #1729 is excluded on both mechanisms. **But not "the mask places well"** — an earlier draft said that, and §26.2 of this same document refutes it: that spread measured **worse** (0.133 ms vs 0.079 ms) for a pinned multi-worker decode pool, because straddling two CCXs costs more than SMT sharing when the working set fits in one CCX's L3. Both hold. That penalty is paid by a *pool*, through cross-CCX barrier traffic, and these rows never build one. Now stated as a trade-off scoped to these rows, pointing at §26.2 for anyone quoting it for a pool workload. ## Verified mitigation `ONNX_GENAI_CPU_DECODE_AFFINITY=off` makes `explicit_decode_affinity_requested()` true (it tests non-empty, not the value) and the auto-mask stands down; `build_global` still runs, so pool *size* is unchanged and only the confinement is dropped. Re-probed at `T = 4`: all 11 threads read `0-31`, symmetric. Re-measurements should set it, alongside `--native-only` or `--ort-intra-threads 1`. **Knock-on for #1792**, which reports this knob as *inert* on the default SPMD path: it is not inert here — it is the only switch that disables the process self-confinement. The same variable is dead in one mechanism and load-bearing in the other, which is a worse user-facing story than either finding alone. Sebastian's call whether that belongs in the issue body. ## Review Opus review confirmed the thread identification, the `T−1` arithmetic against both data points, the spawn-order argument and the mitigation, and caught two places where this revision repeated — in the safe direction — the over-claiming pattern of the two revisions it corrects. Both are fixed in `c195b4fcd`: the `T=16` census is hedged, and the §26.2 contradiction is reconciled rather than left standing. ## Rule added Process-wide state applied during setup makes **construction order part of the experiment design**. "Which arm was built first" silently decided which arm got the whole machine, and nothing in the harness expresses that as a decision — it is the order two `let` bindings happen to appear in. ## Cost `/proc` reads, not timings, so contention cannot corrupt them — taken on a busy host (runnable 33) rather than queued behind it. Widest arm was one `--runs 1` native-only launch. No benchmark, no millisecond claimed. Documentation only. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 23, 2026
…d a knob that no longer exists (#1822) Follow-up to #1173, correcting two defects I shipped in it and repairing the rule they undermined. Docs, one ledger string, one new test, one new script. No production kernel or routing change. ## 1. The ledger named a route gate that had already been deleted `PLAN[MatMulF32].shape_gate` said the native `SimdX86` route "gates M=1 on `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` (default off, #1116)". #1183 shipped that GEMV on by default and removed the probe. `git merge-base --is-ancestor 5417d04 bdb4599` confirms it landed **before** #1173 merged — so the ledger was wrong the day it landed. Today `sgemm_simd` calls `sgemm_simd_variant(a, b, c, m, k, n, true)` unconditionally and `use_m1_gemv` is a plain parameter that only the in-process A/B harness passes as `false`. No environment variable reaches that route. `docs/performance/CPU_MATMUL_ASSIGNMENT.md:559` already recorded the correct fact ("It is measured now, and the route is the default. There is no env probe on the dispatch any more"). Two files in the same directory disagreed and nothing compared them. **Now guarded.** `ledger_prose_only_names_environment_variables_that_still_exist` requires every `NXRT_*` / `ONNX_GENAI_*` token in the ledger's prose to still exist as a string literal in the crate's sources. It cannot check that the description is *right*, only that the knob is *real* — which is the half that goes stale silently. Mutation-verified, not just observed green: ``` matmul_f32: ledger prose names environment variable `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV`, but no source file in this crate contains the literal "ONNX_GENAI_CPU_MM_SIMD_M1_GEMV". ``` ## 2. The doc published a toggle A/B that could not have been run #1173 carried a table captioned **"same binary, same session, toggle the only difference"**, reporting `decode 1×2048×2048` at 0.146 with `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` off against 0.337 with it on, and called turning it on "the obvious next slice". Nothing reads that variable. Setting it measures the same route twice; it cannot produce two different columns. The table is withdrawn and the retraction kept in the text rather than quietly deleted. This is the failure mode the document's own graduation rule warns about — **an arm that was not on the route it was labelled with** — committed by the document that wrote the rule. It survived review because a plausible number in a well-formed table is not self-evidently unmeasured. Readers are pointed at `bench_f32_gemm_ab`, which holds the route as a function parameter and carries the M≥2 rows as a built-in control. ## 3. The gap table is re-measured and the ≥5% rule is repaired The old table was one unguarded invocation per row at an unstated width, taken before the decode-placement corrections (#1729, #1794, #1811) — i.e. when the decode pool put 16 workers on 8 physical cores. New harness: `scripts/bench_native_vs_mlas_width.py`. Arms interleaved rep by rep so host drift lands on both equally; per-rep `os.wait4` CPU-efficiency guard adapted from #1809; six reps per arm; two widths. Raw verdicts, spreads and discards are all reported rather than summarised away. **Three findings, all about method rather than kernels.** | | narrow (6 cores, 1 L3) | wide (32 logical CPUs) | |---|---|---| | `matmul_f32 16×512×512` | 1.581, spread 41% | 0.866, spread 134% | | `matmul_f32 decode 1×2048×2048` | 1.117, spread 21% | 0.934, spread 13% | - **Two cases change verdict on width alone.** Same binary, same half-hour, only the CPU mask differs. `x86_sgemm` parallelises over column strips and MLAS declines to parallelise some shapes, so interleaving the two *routes* inside one process does not protect the ratio — it changes both at once. - **`16×512×512` disagrees with itself on both arms**, alternating `keep-mlas` / `native-graduates` from a byte-identical binary. **One more run of the old table could have graduated a route on this row.** - **The narrow arm is more trustworthy despite having fewer cores** — spreads 4–42% against 5–134%, and it lost no reps to the guard. Isolation beat parallelism. **Softmax now decomposes cleanly**, because no vendored MLAS kernel has changed since #1173 (the only `mlas-sys` edits are the additive straggler handshake in `work_stealing_pool.rs`, #828/#1714, which adds waiting). At matched width the MLAS control arm is stationary to within 4% while native improved **1.24–1.27×** — matching #1416's claim for the row kernel. The f32 GEMM rows get no such attribution and now say so explicitly: their control moved **2.0× the wrong way**, so only the current ratio at a stated width is defensible. **The rule gains what it lacked**: spread must be smaller than the claimed win; reps that did not get the CPU are discarded rather than averaged; a verdict is valid only at a stated width. Under it, `decode 1×2048×2048` — the first f32 GEMM case to show a real native win — **still does not graduate**: it costs more CPU (cpu_ratio 0.875), does not hold at 32 threads, and its 21% spread exceeds its 12% win. ## The width claim is verified, not asserted #1815 landed while this was in progress and observed the neighbouring `bench_generic` harness spawning its ORT arm *outside* the affinity confinement it applied to the native arm. That hazard applies to any `taskset` claim, including mine, so I checked it instead of trusting it — sampling `Cpus_allowed_list` from `/proc/<pid>/task/*/status` 40× across a live narrow-arm run: ``` '16,20,22,26,28,30': 478 observations native_vs_mlas- 273, mlas-sys-ws-0..4 39 each, nxrt-task-0..4 2 each '0-31': 1 (the taskset process itself, before exec) ``` Both routes confined identically; no thread escaped. The rule now requires this check. ## Validation - `dispatch_ledger` **17/17**, including the new falsifier, after merging latest `main`. - `default_artifacts_are_mlas_free` **9/9** — the no-MLAS-in-defaults invariant is untouched. - `cargo clippy -p onnx-runtime-ep-cpu --lib --all-targets` clean; `cargo fmt --check` clean. - Normal merge of `origin/main` (`aee2b9d11`), no rebase, no conflicts. ## Limitations - The narrow arm is six cores on one L3 of one x86-64 host. Nothing here transfers to aarch64 or to a two-socket box. - The `activations erf 1 Mi` row shows native 13.5% slower at matched width. The nearest scatter figure is the wide arm's 8% spread, but that is a spread of *ratios* against a move in a *native time*, so the two are not strictly commensurable. Its MLAS control also moved 11%. **Flagged for pinned re-measurement, not reported as a regression.** - The wide arm was taken with ~4–5 cores of unrelated load present. That is stated in the doc rather than hidden, and it is why its spreads are wider; the guard reports which reps were discarded instead of pretending the host was quiet. - No production behaviour changes here, so there is no performance claim to make about the shipped artifact. Refs #1173, #1183, #1809, #1815, #1416. ## Independent review, and what it changed An independent adversarial review of the full diff returned **no blockers** — it confirmed the ancestry argument behind the retraction, the stationary-control premise for the softmax attribution, and that the headline case is correctly *refused* by the rule (21% spread against a 12% win). It also found seven real defects, all now fixed in `f0323f9ed`. The one that mattered most was in the new test. It only proved the variable name appeared *somewhere* in the crate, so a variable whose read site had been deleted but whose name survived in an `EnvVarGuard::set(...)` line would still have passed — which is the precise shape of the defect this PR exists to correct. The test now requires the matching line to be an `env::var(` / `env::var_os(` read or an `_ENV: &str =` binding. Verified by mutation in **both** directions: | mutation | before | after | |---|---|---| | reinsert retired `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` into ledger prose | fails ✅ | fails ✅ | | retire the two real `NXRT_CPU_GEMM_BACKEND` reads, leaving the literal only in test guards | **passes ❌** | fails ✅ | The remaining six were prose defects in the doc: a stated spread range that contradicted its own table's 82% row, "within 4%" against a table reading −4.2%, a narrow-arm ratio fused with a wide-arm attribution, a spread quoted as 7.5% that was 8% *and* compared against an incommensurable quantity, the CPU-efficiency guard oversold as "what makes this table measurable at all" (in-process interleaving is what protects the ratio; the guard catches only *differential* descheduling), and a one-directional provenance argument standing in for the direct control measurement that actually carries the softmax attribution. **Two further defects I found myself while checking the tables against each other**, neither raised by the review: - The `ratio` column is a median of per-rep ratios while the `ns/unit` columns are medians of times. Medians do not distribute over division, so every row looked internally inconsistent to anyone who tried to divide it out (`0.0684 / 0.0617 = 1.109` against a stated `1.117`). Now documented, along with why the per-rep form is the correct one to quote: it pairs each MLAS invocation with the native invocation it was interleaved against, which is the entire point of interleaving. The then→now figures are relabelled as quotients of medians. - "wider than nine of the twelve wide-arm rows" was eleven of twelve. ## Adopting #1814 `aee2b9d11` (#1814) landed on `main` while this was in review, and it closes the exact hole the review found in the guard this document recommends. A differential CPU-efficiency check cannot see contention that lands evenly on both arms; #1814's confined-set meter reads busy jiffies on the process's own `Cpus_allowed_list` and subtracts the process's own CPU, so foreign load shows up directly. The rule now points at it, and the tables here are explicitly marked as predating it and guarded by the weaker method. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 23, 2026
… the box (#1885) Closes Gaff's blocking item on #1806. ## The defect, in his words and then in mine > `REAP_DIR` is unguarded. … One kill in that window orphans the directory, after which `reap_if_dead` returns 1 forever and **no stale lock can ever be reaped again** — the box wedges behind a dead holder and `acquire` just reports busy, indistinguishable from legitimate occupancy. Correct, and the framing is the part worth keeping: **the wedge this script exists to prevent, relocated into the mechanism that prevents it.** The lock has an anchor pid, a start time and a reaper. Its own mutex had none of the three — a bare `mkdir "$REAP_DIR"` released by `rmdir`. #1811 bounded it by age (`REAPER_GRACE=60`). That recovers, but it recovers *by waiting*, and it bought the mirror-image defect: **age is not evidence of death.** A reaper that is merely slow — loaded box, a qemu leg, a stalled `stat` — could have its guard cleared while it was still inside the critical section. Two acquirers then both reap and both believe they own the host. That is worse than a wedge, because a wedge is loud and this is silent, and the numbers on both sides are ruined without either party knowing. ## What it is now The guard is built the same way as the lock, for the same reasons: | property | mechanism | why | |---|---|---| | published atomically **with** metadata | stage a populated dir, `mv -T` | `rename(2)` onto a non-empty dir fails `ENOTEMPTY` → test-and-set, not last-writer-wins. No kill can produce an unattributable guard, because there is no instant at which one exists | | owned | `anchor_pid` + `start_time` in `$REAP_DIR/meta` | pids recycle; "a process with this number exists" is a different question from "the reaper is still running" | | dead owner reclaimed **immediately** | `anchor_alive` | no grace, no waiting, nothing to wedge behind | | live owner **never** disturbed | `anchor_alive` outranks age | kills the double-reap that the age rule introduced | | released only by its owner | pid check in `reaper_release` | a process whose guard was reclaimed while descheduled must not delete its successor's guard on the way out | `anchor_alive` is now **one** predicate, shared with `holder_alive`. Four call sites each deciding "is this pid alive?" for themselves produced two of the four defects in #1830; I am not doing that again in a second module. `REAPER_GRACE` survives for exactly one residual class — a guard that exists with no readable anchor. Staging makes that unreachable from this script; a stray directory at that path from an older version or a hand-run `mkdir` can still present it, and for that class no liveness evidence exists, so age is all there is. The existing test for it stays. ## The kill-in-window test is deterministic, and it cannot pass vacuously Requested explicitly, and it is the cell I would have written badly: ```sh HOSTLOCK_REAPER_STALL=30 $HL acquire --owner victim ... & # seam holds the section open ...wait for $LOCK.reaper/stalled_pid... # the pid INSIDE the window sig "$stalled" 9 # SIGKILL lands in the window every run chk "killing it in the window really does orphan the guard" ... # ← the orphan EXISTS chk "and the orphan still names its dead owner" ... $HL acquire --owner roy ... chk "a guard orphaned by SIGKILL does not wedge the next acquirer" "$(st owner)" "roy" chk "and recovery is immediate, not after REAPER_GRACE" "$(...)" "clear" ``` Three things this gets right that the obvious version does not: 1. **Deterministic, not opportunistic.** Racing a real reaper's microsecond window means the kill lands inside it when the scheduler feels like it. The seam makes the window as wide as I ask. 2. **It asserts the orphan exists first.** Without that line, all three recovery checks pass equally well when nothing ever leaked — the vacuous-coverage failure Gaff flagged on Resch's #1805 the same night, where every assertion sat behind a condition that could silently be false. 3. **`cleanup` is not called between the kill and the recovery.** The harness sweeps `$LOCK.reaper`, so a tidy-looking cleanup there would erase the orphan and the cell would pass **against the defect it exists to catch.** Cleanup must not mask the defect; here it is a deliberate omission with a comment saying so. The seam is production code, so its inertness is asserted too (R8.4): unset, the reap completes in under 3 s. ## Falsification Suite **231 → 247**, green. `shellcheck` clean. Nine mutations, all red: | # | mutation | result | |---|---|---| | P1 | dead owner never reclaimed (the original wedge) | 242 passed / **2 failed** | | P2 | age-only rule restored, liveness ignored | 240 / **4** | | P3 | release without the ownership check | 243 / **1** | | P4 | guard created in place instead of renamed into place | 242 / **2** | | P5 | stall seam always fires | 199 / **45** | | P6 | `anchor_alive` ignores `start_time` (recycled pid) | 242 / **2** | | P7 | rename-then-remove keeps the corpse (clear path) | 246 / **1** | | P8 | rename-then-remove keeps the corpse (release path) | 246 / **1** | | P9 | release checks pid but not start time | 246 / **1** | **P7–P9 came from review, and both findings are the shape this file keeps producing: the fix was right and the suite proving it could not tell.** *P9* — `reaper_release` compared only `$$` while `reaper_clear_if_dead` compares pid **and** start time. The asymmetry sat exactly where the consequence is worst: a recycled pid landing on our number makes us delete a **live successor's** guard, rather than merely failing to tidy up our own. This box is at ~1.5M pids in four days, so recycling is not theoretical. R8.3b forges a successor guard carrying the stalled reaper's own pid with a start time that is not its, and requires it to survive. *P7/P8* — dropping the `rm -rf` from either rename-then-remove left all 244 assertions green, because the harness `cleanup()` sweeps those globs between cells. In production that is unbounded `.dead.*`/`.rel.*` growth beside the lock. The new assertion runs **before** cleanup, deliberately: placed after it, it would pass whether or not the code tidies up at all — the same "cleanup masks the defect" failure the kill-in-window cell was written to avoid, sitting one cell to the left of where I was looking for it. P4 is a structural assertion rather than a behavioural one, and the PR says so rather than dressing it up: killing between a `mkdir` and a meta write is a microsecond window, and a seam wide enough to test it would be wider than the bug. The suite asserts instead that the guard is never created in place — weaker evidence, honestly labelled. ## The other two items **`--gate` stays a start admission only.** No change to its semantics. The README already carries the split — the lock and the gate decide whether to **start**, `--expect-cores/--min-efficiency` (#1864) decides whether to **believe** — with Roy's 52% A/A null as the evidence for why one instrument cannot do both. **The harness holds the lock across all arms**, and rows carry `held_by`, `contended`, `runnable_at_acquire` and now `efficiency`. Unchanged here, restated in the README because "I sampled `ps` and the host looked free" was a *between-arms* sample of somebody else's interleaved A/B. Also documents that this script is **Linux-only by construction** (`/proc`, `mv -T`, `stat -c`) instead of leaving the question open. A portability fallback that degraded liveness to `kill -0` would reap live holders on whichever platform took it — the one error this script must never make. No crate code. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Sebastian's #1729 write-up says every
t≥8row either of us published pre-6e8c31ebdwas taken on 8 physical cores, so they all need re-taking. My §45.9/45.10 retraction had listed placement as an open confound on my own rows, marked "not yet verified", pending a/procread queued behind the host owner.I went to close it. The first draft of this PR closed it in the wrong direction, and Opus review caught that before merge. Both halves are below, because the mistake is the more useful half.
What is actually established
The SPMD decode pool — the one
#1729re-placed — is never entered by a single-node SDPA/MHA graph. From routing code, no timing, no cores:decode_pool_active()(sdpa.rs:1287); only the true arm reachesdecode_parallel_output_row_blocks. The false arm goes totask_runtime::chunk_runs_mut.decode_pool_active()is gated on thread-localsIN_SPMD_SCOPE(matmul_nbits.rs:4901) /IN_NUMA_SCOPE(:4891), set only insidewith_decode_pool_scope— not process-globals, so no other thread's decode loop can turn it on for the benchmark's engine thread.native_decode/(cpu.rs:445/578/584,proposer.rs:325).bench_genericnever mentionsnative_decode.That much stands.
What I got wrong
I wrote "the cpus 0-15 mask never governed these measurements". What I had shown was that that pool never governed them. Different statement.
There is a second, independent confinement mechanism, and it did apply:
CpuExecutionProvider::initialize(provider.rs:340) callsbound_process_to_decode_budget()(matmul_nbits.rs:4354), which — budget set, no explicitONNX_GENAI_CPU_DECODE_AFFINITY— performs a realsched_setaffinity. Its doc comment states the reach: "threads spawned afterwards … inherit the mask".bench_genericreaches it:InferenceSession::load→executor/platform.rs:8→ep.initialize(...). And it setsONNX_GENAI_CPU_DECODE_THREADSfrom--native-threadsat:670-673, before any session exists. So on every row whereab.pypassed a width, the benchmark process confined itself.It is keyed on the same knob. Sebastian documented this mechanism himself in the
#1746addendum to2026-08-22-decode-width-scaling.md, including that hisint4_decode_loop_abescapes it by never callinginitialize(). Mine is the opposite case. I had read that addendum as being about his harness and never checked mine against it.The third asymmetry this turned up
Affinity is inherited at thread creation, and the two arms are built on opposite sides of the mask:
bench_genericconstructsort_sessionat:691, thenInferenceSession::loadat:703— and it is that load which confines the process. ORT's intra-op pool is spawned before the confinement; the native EP's threads after it. If ORT's pool is built at session construction, the paired runs handed ORT all 32 logical CPUs while the native arm ran onT.Same direction as the ORT spin-wait tax, and it also grows with
T. Recorded as suspected, not observed — settled by readingCpus_allowed_listper thread from/proc/<pid>/task/*/statuson a live paired run, grouped by thread name. Needs a quiet host; queued, not claimed.Net effect, which is the inverse of the first draft
These rows are not placement-clean and must not be treated as such. The
#1729mechanism specifically is excluded; the self-imposedT-CPU mask is not, and its interaction with the ORT arm is open. They need re-taking regardless, for the spin-tax reason, which was always sufficient on its own.I have deliberately not re-run anything: host is Sebastian's (qemu aarch64 suite +
decode_gap_park, runnable 33).Also
The file carried two
### 45.9headings. Since this PR is entirely about which §45.9 claim was wrong, and the cross-references at:4716/:4722cite it by number, the retraction is renumbered 45.10 and those references updated.2026-08-22-sdpa-fanout.mdcites the other 45.9 and is unaffected.Rules added
grepforsched_setaffinityaway, in a file I had already opened, described in a teammate's document I had already read.Scope
Documentation only — one section of the benchmark ledger plus a heading renumber. No code, no measurement, zero cores.