Repository navigation
acc0 width-16: two null hypotheses rejected, and the +23% steal-tiles candidate closed - #1946
Merged
Merged
Conversation
…width — page backing rejected The A/A null at width 16 is the binding constraint on the acc0 work: it is wide enough to refuse the +23% steal-tiles candidate. This tests a third candidate mechanism, rejects it, and replaces "unexplained 1.69x spread" with a much narrower question. REJECT: transparent-hugepage backing of the weight arena. THP is `[always]` with `defrag=madvise` here, so a 2 MB fault that cannot be served immediately falls back to 4 KB silently and permanently -- a per-process lottery that fits the null's shape exactly. It is not what is happening. Twelve quiet-host launches: thp_frac 0.823-0.928, range 0.104 against a pre-registered 0.20, and Spearman rho -0.19 against a required -0.70, while ms_token spans 1.725x. 83-93% of anonymous memory is already hugepage-backed in every launch, including every slow one. The pre-registered range guard is why this is a rejection. A four-launch reconnaissance returned a *perfect* rho of -1.0000 over a 0.023 range; twelve launches collapsed it to -0.1888. That is the same small-n manufacture that produced the 1.1910 dispatcher-pin ACCEPT which did not replicate. What the runs established is sharper than the rejection: * The null is two modes, not a spread. Twelve launches split into a slow cluster of five agreeing to 1.05%, with the inter-cluster gap 11x the largest within-cluster gap. Launch order does not predict membership. * Both modes reproduce to within 1% in an independent quiet-host run 40 minutes later (3.48-3.81 and 5.91-6.03), which makes them a property of the system. * `park_frac` is REJECTED as a classifier -- a lead from a contaminated run that did not survive. Mode B's launch 4 has *fewer* parks than mode A's launch 8. `sys_frac` and `cpu_s_per_token` overlap across the modes too. * Effective lanes -- `cpu_s_per_token / ms_token` -- separates them with no overlap: mode A 15.30-16.10 of sixteen, mode B 9.76-12.16, a 3.1-lane gap. The slow mode is not burning more CPU. It uses less per token and takes 1.55x the wall time. It is running on about ten of its sixteen lanes. So the question is now "why does a launch that builds 15 workers on 15 verified distinct physical cores run at two-thirds of its width, for its entire life, decided before the first token". Leading hypothesis is foreign load on a subset of the pinned CPUs -- explicitly unproven, with a direct falsifier stated (per-CPU /proc/stat busy time on the pinned set minus our own workers' CPU time). If that is it, the block on the +23% candidate becomes a measurement problem to stratify rather than a question about the pool. Docs and bench scripts only; no compiled code changes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
enabled auto-merge (squash)
August 24, 2026 05:57
…t to straggler wait The width-16 A/A null record stated a falsifier for its leading hypothesis and did not run it. This runs it. acc0_w16_foreign_load.py reads per-CPU busy jiffies from /proc/stat for exactly the pinned set before and after each launch and subtracts the child's own total CPU (getrusage(RUSAGE_CHILDREN) deltas, whole process). The child is pinned to the same CPUs, so the remainder is CPU time on our cores that is not ours. REJECT. Foreign time never exceeds 0.59 CPU-seconds against ~34 of pinned busy time -- a 1.7% ceiling in every launch, where losing 3.7 of 16 lanes needs 23%. The bound is 13x too small, which rejects without depending on n or on sampling the slow mode. Spearman rho is +0.0210 and the fastest launch in the run carries 1.6x more foreign time than the slow one. The same run relocates the target. Splitting CPU per token into user and system, user CPU per token spans 11% across both modes while wall spans 1.69x; the slow launch is +4.6% user and +170% sys. The work is identical -- the lanes are lost to waiting in worker_wait's yield loop, i.e. a persistent straggler inside the process, with placement, page backing and foreign load now all excluded. Two consequences for A/B work at width 16, neither needing a pool fix: stratify on an in-launch statistic that cannot see the arm, and score work-reducing candidates on user CPU per token where the null is 15x smaller (null-immune but narrow -- it is the steal-tiles candidate's control, not its score). The +23% candidate stays blocked pending a mode-stratified re-test. Two rule defects fixed rather than papered over: the range guard was ordered before the both-modes-present check, which would have printed REJECT on an uninformative single-mode run; and the pre-existing 'the slow mode uses less CPU' line generalised from one B launch undercutting one A launch, which is the small-n manufacture trap inside a document about that trap. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1946 +/- ##
==========================================
+ Coverage 80.42% 81.00% +0.57%
==========================================
Files 420 423 +3
Lines 202834 207429 +4595
Branches 202834 207429 +4595
==========================================
+ Hits 163128 168022 +4894
+ Misses 34143 33768 -375
- Partials 5563 5639 +76
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
…l null The candidate was refused in August as unprovable: effect +0.2327 against a required 3x A/A half-width of 0.6462. The honest reading at the time was that the instrument was too blunt to see a +23% effect. Re-run 24 launches deep with the same rule, there is no effect to see. original (n=8): ratio 1.2327, sign 88%, A/A half-width 0.2154 this run (n=24): ratio 0.9883, sign 38%, A/A half-width 0.1478 this run, fast mode (n=8): ratio 0.9889, sign 38%, A/A half-width 0.0323 The stratified arm is a 4.6x sharper instrument that would have resolved anything above +9.7%, and the candidate is flat in it. The original +23% is quantitatively accounted for. Its control arm drew the slow mode in 4 of 8 launches and its test arm in 1 of 8; a slow control deflates the denominator by the published 1.687x mode ratio, so a net 37.5% imbalance can manufacture up to +0.2576 of ratio on its own -- more than the +0.2327 observed. acc0_w16_mode_stratified.py contains no rule: it imports verdict() from the A/B harness and calls it, so every threshold, the mechanism claim and the width-8 guard are the code that produced the original REJECT. It only chooses which launches are fed in, and prints the unmodified verdict beside the stratified one. Its first gate -- slow-mode rate per configuration, the check for whether filtering is arm-selective -- PASSES on the original run at 1.33x and is blind to the defect, because the rotation pools the A/A slot and the ratio is formed from control and test specifically. It now also prints a paired mode imbalance unconditionally, which flags the original run UNUSABLE. That check costs nothing and would have caught this the day it was taken. Two things worth carrying beyond this candidate. The original run's mechanism confirmation (sys_frac falling 0.280 -> 0.192 at 88% sign) inverts at n=24, because the slow mode is the high-sys mode -- a directionally-correct mechanism reading corroborates nothing when it is downstream of the same confound. And user CPU per token stayed flat to 1.12x across both configurations and both modes in both runs, which is the correct control for a load-balance change and an independent replication of the A/A finding on data taken before it. Stratification is validated as a method here and costs ~3x the launches, since each launch runs three independent width-16 processes that each draw the mode. The 22.2-point straggler wait remains open; what is closed is spare tiles as its collector. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The width-16 A/A null: two hypotheses tested, both rejected, target relocated
The width-16 A/A null is the binding constraint on all acc0 CPU work — it is wide enough to refuse the +23% steal-tiles candidate. This PR is the record of narrowing it, and adds the three probes that did the narrowing.
1. Transparent hugepage backing — REJECT
THP is
[always]withdefrag=madviseon this host, so a 2 MB fault that cannot be served falls back to 4 KB silently and permanently: a per-process startup lottery matching every property of the null. Pre-registered rule (n≥8, Spearman rho ≤ −0.70,thp_fracrange ≥ 0.20), 12 trusted launches: range 0.104, rho −0.1888. Both conditions fail.A 4-launch reconnaissance had returned a perfect rho of −1.0000 over a 0.023 range. The pre-registered range guard is the only reason that is a rejection and not a headline. Third instance of this trap in this workstream.
2. Foreign load on the pinned CPUs — REJECT
The record above stated this falsifier and did not run it.
acc0_w16_foreign_load.pyreads per-CPU busy jiffies from/proc/statfor exactly the pinned set before and after each launch and subtracts the child's own total CPU (getrusage(RUSAGE_CHILDREN), whole process, not just steady phase). The child is pinned to the same CPUs, so the remainder is CPU time on our cores that is not ours.Foreign time never exceeds 0.59 CPU-seconds against ~34 — a 1.7% ceiling in every launch, where losing 3.7 of 16 lanes requires 23%. The bound is 13x too small. Spearman rho +0.0210, and the fastest launch carries 1.6x more foreign time than the slow one.
The magnitude bound does not depend on n or on sampling the slow mode, which matters: that run drew only 1 slow launch in 12. Incidence across three clean runs is 5/12, 6/14, 1/12 — not a stable rate, and the doc says not to quote it.
3. What the null actually is
Splitting CPU per token into user and system:
The work is identical. User CPU per token spans 11% across both modes while wall spans 1.69x. The lanes are lost to waiting in
worker_wait's yield loop — thesysconsumer @sebastian isolated with his blocktime sweep. This is a persistent straggler inside the process; placement (categorical census,2c3968afb), page backing and foreign load are all now excluded. Remaining candidates: weight-arena placement across the two L3/CCX domains, and per-launch clock/boost state.What this unblocks
Two consequences for A/B work at width 16, neither requiring a pool fix:
sys_fracseparates 0.315 vs 0.140–0.200 but on one slow sample, so it is flagged as hypothesis, not result.syswhile leaving user CPU flat, so this is that candidate's control, not its score.The +23% candidate remains blocked pending a mode-stratified re-test against its own unmodified rule. Nothing here claims it is real.
Corrections carried rather than deleted
Contents
docs/benchmarks/2026-08-24-acc0-w16-null-page-backing.md(new) — verdict-first record of both rejections, the bimodality, the classifier work and the user/sys split.docs/performance/CPU_MATMUL_ASSIGNMENT.md— ledger updated.benches/acc0_w16_page_backing.py,benches/acc0_w16_mode_split.py,benches/acc0_w16_foreign_load.py(new) — the probes, each carrying its pre-registered rule in its docstring.Docs and non-compiled scripts only; no Rust changes. Normal auto-merge, no bypass.
Update: the +23% steal-tiles candidate has now been re-tested and is CLOSED
The PR originally ended with "the +23% candidate remains blocked pending a mode-stratified re-test". That re-test is now in this PR.
Stratification does exactly what the null record predicted — the A/A half-width collapses 4.6x, dropping the bar the effect must clear from +44% to +9.7%. The candidate is flat in that sharper instrument.
The original +23% is quantitatively accounted for. Its control arm drew the slow mode in 4 of 8 launches and its test arm in 1 of 8. A slow control deflates the denominator of
tps(test)/tps(control)by the published 1.687x mode ratio, so a net 37.5% imbalance can manufacture up to +0.2576 of ratio on its own — more than the +0.2327 observed.The scorer contains no rule
benches/acc0_w16_mode_stratified.pyimportsverdict()fromacc0_w16_blocktime_ab.pyand calls it. Every threshold, the mechanism claim and the width-8 guard are the same code that produced the original REJECT; the only thing it changes is which launches are fed in, and it prints the unmodified verdict beside the stratified one.The gate I wrote first was blind to the defect
The arm-selectivity gate computes slow-mode rate per configuration, following the harness's rotation. On the original run it reports 33.3% vs 25.0%, ratio 1.33, PASS. It is a correct check for whether filtering is biased and it cannot see this, because the rotation pools the A/A slot and the ratio is formed from control and test specifically. Pooled they look balanced; paired they are 4 against 1. A paired mode imbalance diagnostic is now printed unconditionally and flags the original run
UNUSABLE. It costs nothing and would have caught this the day it was taken.Two things worth carrying past this candidate
sys_fracfalling 0.280 → 0.192 at 88% sign — exactly what removing a straggler wait should look like. At n=24 it is +0.0111 at 46% sign. The slow mode is the high-sysmode, so the mechanism agreed with the effect because both were the same artifact. A directionally-correct mechanism reading corroborates nothing when it is downstream of the confound.Cost, for planning: the stratified arm needs ~3x the launches. Each launch runs three independent width-16 processes and each draws the mode, so at a ~35% slow rate only about a third of launches have all three arms fast.
The 22.2-point straggler wait remains the open target. What is closed is the claim that spare tiles collect it.