Skip to content

acc0 width-16: two null hypotheses rejected, and the +23% steal-tiles candidate closed - #1946

Merged
justinchuby merged 4 commits into
mainfrom
squad/roy-w16-null-page-backing
Aug 24, 2026
Merged

justinchuby merged 4 commits into
mainfrom
squad/roy-w16-null-page-backing

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 24, 2026 •

Copy link
Copy Markdown
Owner

The width-16 A/A null: two hypotheses tested, both rejected, target relocated

The width-16 A/A null is the binding constraint on all acc0 CPU work — it is wide enough to refuse the +23% steal-tiles candidate. This PR is the record of narrowing it, and adds the three probes that did the narrowing.

1. Transparent hugepage backing — REJECT

THP is [always] with defrag=madvise on this host, so a 2 MB fault that cannot be served falls back to 4 KB silently and permanently: a per-process startup lottery matching every property of the null. Pre-registered rule (n≥8, Spearman rho ≤ −0.70, thp_frac range ≥ 0.20), 12 trusted launches: range 0.104, rho −0.1888. Both conditions fail.

A 4-launch reconnaissance had returned a perfect rho of −1.0000 over a 0.023 range. The pre-registered range guard is the only reason that is a rejection and not a headline. Third instance of this trap in this workstream.

2. Foreign load on the pinned CPUs — REJECT

The record above stated this falsifier and did not run it. acc0_w16_foreign_load.py reads per-CPU busy jiffies from /proc/stat for exactly the pinned set before and after each launch and subtracts the child's own total CPU (getrusage(RUSAGE_CHILDREN), whole process, not just steady phase). The child is pinned to the same CPUs, so the remainder is CPU time on our cores that is not ours.

Foreign time never exceeds 0.59 CPU-seconds against ~34 — a 1.7% ceiling in every launch, where losing 3.7 of 16 lanes requires 23%. The bound is 13x too small. Spearman rho +0.0210, and the fastest launch carries 1.6x more foreign time than the slow one.

The magnitude bound does not depend on n or on sampling the slow mode, which matters: that run drew only 1 slow launch in 12. Incidence across three clean runs is 5/12, 6/14, 1/12 — not a stable rate, and the doc says not to quote it.

3. What the null actually is

Splitting CPU per token into user and system:

fast mode (11) slow launch
wall ms/token 3.406 – 3.796 5.910 (1.69x)
user s/token 0.0452 – 0.0485 0.0503 (+4.6%)
sys s/token 0.0081 – 0.0113 0.0232 (+170%)

The work is identical. User CPU per token spans 11% across both modes while wall spans 1.69x. The lanes are lost to waiting in worker_wait's yield loop — the sys consumer @sebastian isolated with his blocktime sweep. This is a persistent straggler inside the process; placement (categorical census, 2c3968afb), page backing and foreign load are all now excluded. Remaining candidates: weight-arena placement across the two L3/CCX domains, and per-launch clock/boost state.

What this unblocks

Two consequences for A/B work at width 16, neither requiring a pool fix:

  1. Stratify on an in-launch statistic that cannot see the arm. Effective lanes separates with no overlap and is backed by three runs; sys_frac separates 0.315 vs 0.140–0.200 but on one slow sample, so it is flagged as hypothesis, not result.
  2. Score work-reducing candidates on user CPU per token, where the null is ~15x smaller. Null-immune but narrow: a pure load-balance change like the +23% steal-tiles candidate moves wall and sys while leaving user CPU flat, so this is that candidate's control, not its score.

The +23% candidate remains blocked pending a mode-stratified re-test against its own unmodified rule. Nothing here claims it is real.

Corrections carried rather than deleted

  • The rule in probe 2 tested its range guard before the both-modes-present check, which would have printed REJECT on an uninformative single-mode run. A narrow range in the cause is only evidence once the effect is known to have moved — the opposite reading from probe 1, where a narrow range invalidated an ACCEPT. Fixed before the first launch; the docstring says why.
  • The earlier "the slow mode uses less CPU per token" line generalised from one B launch undercutting one A launch. Mode B's median is above mode A's. Corrected in place in both doc and ledger, because it is the small-n manufacture trap inside the document about that trap.

Contents

  • docs/benchmarks/2026-08-24-acc0-w16-null-page-backing.md (new) — verdict-first record of both rejections, the bimodality, the classifier work and the user/sys split.
  • docs/performance/CPU_MATMUL_ASSIGNMENT.md — ledger updated.
  • benches/acc0_w16_page_backing.py, benches/acc0_w16_mode_split.py, benches/acc0_w16_foreign_load.py (new) — the probes, each carrying its pre-registered rule in its docstring.

Docs and non-compiled scripts only; no Rust changes. Normal auto-merge, no bypass.


Update: the +23% steal-tiles candidate has now been re-tested and is CLOSED

The PR originally ended with "the +23% candidate remains blocked pending a mode-stratified re-test". That re-test is now in this PR.

run n ratio (2 ÷ 1) sign A/A half-width
original, 2026-08-23 8 1.2327 88% 0.2154
this run, unstratified 24 0.9883 38% 0.1478
this run, fast mode only 8 0.9889 38% 0.0323

Stratification does exactly what the null record predicted — the A/A half-width collapses 4.6x, dropping the bar the effect must clear from +44% to +9.7%. The candidate is flat in that sharper instrument.

The original +23% is quantitatively accounted for. Its control arm drew the slow mode in 4 of 8 launches and its test arm in 1 of 8. A slow control deflates the denominator of tps(test)/tps(control) by the published 1.687x mode ratio, so a net 37.5% imbalance can manufacture up to +0.2576 of ratio on its own — more than the +0.2327 observed.

The scorer contains no rule

benches/acc0_w16_mode_stratified.py imports verdict() from acc0_w16_blocktime_ab.py and calls it. Every threshold, the mechanism claim and the width-8 guard are the same code that produced the original REJECT; the only thing it changes is which launches are fed in, and it prints the unmodified verdict beside the stratified one.

The gate I wrote first was blind to the defect

The arm-selectivity gate computes slow-mode rate per configuration, following the harness's rotation. On the original run it reports 33.3% vs 25.0%, ratio 1.33, PASS. It is a correct check for whether filtering is biased and it cannot see this, because the rotation pools the A/A slot and the ratio is formed from control and test specifically. Pooled they look balanced; paired they are 4 against 1. A paired mode imbalance diagnostic is now printed unconditionally and flags the original run UNUSABLE. It costs nothing and would have caught this the day it was taken.

Two things worth carrying past this candidate

  • The mechanism confirmation inverted. The original run's strongest non-verdict evidence was sys_frac falling 0.280 → 0.192 at 88% sign — exactly what removing a straggler wait should look like. At n=24 it is +0.0111 at 46% sign. The slow mode is the high-sys mode, so the mechanism agreed with the effect because both were the same artifact. A directionally-correct mechanism reading corroborates nothing when it is downstream of the confound.
  • The user-CPU control held on both runs (0.04780/0.04801 and 0.04810/0.04836; total spread 1.124x across both configurations and both modes). That is the right shape of control for a load-balance change, and an independent replication of "the modes differ in waiting, not in work" on data taken before that record existed.

Cost, for planning: the stratified arm needs ~3x the launches. Each launch runs three independent width-16 processes and each draws the mode, so at a ~35% slow rate only about a third of launches have all three arms fast.

The 22.2-point straggler wait remains the open target. What is closed is the claim that spare tiles collect it.

…width — page backing rejected

The A/A null at width 16 is the binding constraint on the acc0 work: it is wide
enough to refuse the +23% steal-tiles candidate. This tests a third candidate
mechanism, rejects it, and replaces "unexplained 1.69x spread" with a much
narrower question.

REJECT: transparent-hugepage backing of the weight arena. THP is `[always]`
with `defrag=madvise` here, so a 2 MB fault that cannot be served immediately
falls back to 4 KB silently and permanently -- a per-process lottery that fits
the null's shape exactly. It is not what is happening. Twelve quiet-host
launches: thp_frac 0.823-0.928, range 0.104 against a pre-registered 0.20, and
Spearman rho -0.19 against a required -0.70, while ms_token spans 1.725x. 83-93%
of anonymous memory is already hugepage-backed in every launch, including every
slow one.

The pre-registered range guard is why this is a rejection. A four-launch
reconnaissance returned a *perfect* rho of -1.0000 over a 0.023 range; twelve
launches collapsed it to -0.1888. That is the same small-n manufacture that
produced the 1.1910 dispatcher-pin ACCEPT which did not replicate.

What the runs established is sharper than the rejection:

* The null is two modes, not a spread. Twelve launches split into a slow cluster
  of five agreeing to 1.05%, with the inter-cluster gap 11x the largest
  within-cluster gap. Launch order does not predict membership.
* Both modes reproduce to within 1% in an independent quiet-host run 40 minutes
  later (3.48-3.81 and 5.91-6.03), which makes them a property of the system.
* `park_frac` is REJECTED as a classifier -- a lead from a contaminated run that
  did not survive. Mode B's launch 4 has *fewer* parks than mode A's launch 8.
  `sys_frac` and `cpu_s_per_token` overlap across the modes too.
* Effective lanes -- `cpu_s_per_token / ms_token` -- separates them with no
  overlap: mode A 15.30-16.10 of sixteen, mode B 9.76-12.16, a 3.1-lane gap.
  The slow mode is not burning more CPU. It uses less per token and takes 1.55x
  the wall time. It is running on about ten of its sixteen lanes.

So the question is now "why does a launch that builds 15 workers on 15 verified
distinct physical cores run at two-thirds of its width, for its entire life,
decided before the first token". Leading hypothesis is foreign load on a subset
of the pinned CPUs -- explicitly unproven, with a direct falsifier stated
(per-CPU /proc/stat busy time on the pinned set minus our own workers' CPU
time). If that is it, the block on the +23% candidate becomes a measurement
problem to stratify rather than a question about the pool.

Docs and bench scripts only; no compiled code changes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby enabled auto-merge (squash) August 24, 2026 05:57
…t to straggler wait

The width-16 A/A null record stated a falsifier for its leading hypothesis and
did not run it. This runs it.

acc0_w16_foreign_load.py reads per-CPU busy jiffies from /proc/stat for exactly
the pinned set before and after each launch and subtracts the child's own total
CPU (getrusage(RUSAGE_CHILDREN) deltas, whole process). The child is pinned to
the same CPUs, so the remainder is CPU time on our cores that is not ours.

REJECT. Foreign time never exceeds 0.59 CPU-seconds against ~34 of pinned busy
time -- a 1.7% ceiling in every launch, where losing 3.7 of 16 lanes needs 23%.
The bound is 13x too small, which rejects without depending on n or on sampling
the slow mode. Spearman rho is +0.0210 and the fastest launch in the run carries
1.6x more foreign time than the slow one.

The same run relocates the target. Splitting CPU per token into user and system,
user CPU per token spans 11% across both modes while wall spans 1.69x; the slow
launch is +4.6% user and +170% sys. The work is identical -- the lanes are lost
to waiting in worker_wait's yield loop, i.e. a persistent straggler inside the
process, with placement, page backing and foreign load now all excluded.

Two consequences for A/B work at width 16, neither needing a pool fix: stratify
on an in-launch statistic that cannot see the arm, and score work-reducing
candidates on user CPU per token where the null is 15x smaller (null-immune but
narrow -- it is the steal-tiles candidate's control, not its score). The +23%
candidate stays blocked pending a mode-stratified re-test.

Two rule defects fixed rather than papered over: the range guard was ordered
before the both-modes-present check, which would have printed REJECT on an
uninformative single-mode run; and the pre-existing 'the slow mode uses less
CPU' line generalised from one B launch undercutting one A launch, which is the
small-n manufacture trap inside a document about that trap.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 24, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 81.00%. Comparing base (2f2eb6d) to head (f37c6dd).
⚠️ Report is 13 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1946      +/-   ##
==========================================
+ Coverage   80.42%   81.00%   +0.57%     
==========================================
  Files         420      423       +3     
  Lines      202834   207429    +4595     
  Branches   202834   207429    +4595     
==========================================
+ Hits       163128   168022    +4894     
+ Misses      34143    33768     -375     
- Partials     5563     5639      +76     
Flag Coverage Δ
cli-ort-linux 72.51% <ø> (?)
cli-ort-windows 72.10% <ø> (ø)
mlas 85.33% <ø> (?)
offline 81.15% <ø> (+0.51%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 33 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

…l null

The candidate was refused in August as unprovable: effect +0.2327 against a
required 3x A/A half-width of 0.6462. The honest reading at the time was that
the instrument was too blunt to see a +23% effect. Re-run 24 launches deep with
the same rule, there is no effect to see.

  original (n=8):            ratio 1.2327, sign 88%, A/A half-width 0.2154
  this run (n=24):           ratio 0.9883, sign 38%, A/A half-width 0.1478
  this run, fast mode (n=8): ratio 0.9889, sign 38%, A/A half-width 0.0323

The stratified arm is a 4.6x sharper instrument that would have resolved
anything above +9.7%, and the candidate is flat in it.

The original +23% is quantitatively accounted for. Its control arm drew the slow
mode in 4 of 8 launches and its test arm in 1 of 8; a slow control deflates the
denominator by the published 1.687x mode ratio, so a net 37.5% imbalance can
manufacture up to +0.2576 of ratio on its own -- more than the +0.2327 observed.

acc0_w16_mode_stratified.py contains no rule: it imports verdict() from the A/B
harness and calls it, so every threshold, the mechanism claim and the width-8
guard are the code that produced the original REJECT. It only chooses which
launches are fed in, and prints the unmodified verdict beside the stratified one.

Its first gate -- slow-mode rate per configuration, the check for whether
filtering is arm-selective -- PASSES on the original run at 1.33x and is blind
to the defect, because the rotation pools the A/A slot and the ratio is formed
from control and test specifically. It now also prints a paired mode imbalance
unconditionally, which flags the original run UNUSABLE. That check costs nothing
and would have caught this the day it was taken.

Two things worth carrying beyond this candidate. The original run's mechanism
confirmation (sys_frac falling 0.280 -> 0.192 at 88% sign) inverts at n=24,
because the slow mode is the high-sys mode -- a directionally-correct mechanism
reading corroborates nothing when it is downstream of the same confound. And
user CPU per token stayed flat to 1.12x across both configurations and both
modes in both runs, which is the correct control for a load-balance change and
an independent replication of the A/A finding on data taken before it.

Stratification is validated as a method here and costs ~3x the launches, since
each launch runs three independent width-16 processes that each draw the mode.
The 22.2-point straggler wait remains open; what is closed is spare tiles as its
collector.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby justinchuby changed the title perf(bench): the width-16 A/A null is bimodal and runs at two-thirds width — page backing rejected acc0 width-16: two null hypotheses rejected, and the +23% steal-tiles candidate closed Aug 24, 2026
@justinchuby
justinchuby merged commit eee5da8 into main Aug 24, 2026
14 of 16 checks passed
@justinchuby
justinchuby deleted the squad/roy-w16-null-page-backing branch August 24, 2026 09:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant