Skip to content

docs(perf): the 1.84x acc0 gap is stale — re-measured it is 1.12x at t=1/4/8 - #1852

Merged
justinchuby merged 5 commits into
mainfrom
squad/roy-acc0-gap-remeasure
Aug 23, 2026
Merged

justinchuby merged 5 commits into
mainfrom
squad/roy-acc0-gap-remeasure

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 23, 2026 •

Copy link
Copy Markdown
Owner

The 1.84x acc0 gap is stale. Re-measured, it is ~1.12x.

The published acc0 (accuracy_level = 0, the production default) int4 decode gap
against ORT — 1.84x at t=1 — is what made acc0 the top remaining CPU
MatMulNBits target in the ledger. It dates from e9754e7ef (#1628) and eight
merges have landed since
, three of them direct acc0 kernel work. On current
main the gap measures ~1.12x.

width native tok/s ORT tok/s gap gap range cells (trusted/taken) A/A range
1 27.9 31.2 1.120x 1.112–1.128 3/3, 2 retained 1.025–1.036
4 107.2 122.3 ~1.15x 1.087–1.284 4/6 0.868–1.150
8 211.0 238.0 1.120x 1.089–1.145 3/3 0.997–1.028
16 — — ~1.64x, does not resolve 1.456–1.831 2/3 0.969–1.295

Both arms are tokens_s_total, paired within each launch and then medianed.
Trusted/taken is the harness's verdict; retained is editorial — all three
t=1 cells passed the guard and one was dropped afterwards by me, disclosed
below.

Only t=1 and t=8 resolve. t=4 sits inside its own A/A null (0.868–1.150).
t=16 reads ~1.64x and is the open row — see below; a second revision of
this PR corrects an earlier claim that it was wholly contaminated.

acc0 is no longer the top CPU target — conditional on t=16. At the two
widths that resolve, the remaining ~12% is a kernel efficiency difference that
sits below several other open items. t=16 is the width closest to an
unconfined production process, and a confirmed 1.64x there would reverse that.

Why the movement is kernel — measured, not inferred

The first revision of this PR argued from a control: ORT re-measures at 31.99 ms,
within 4.4% of its published 30.632, therefore the harness is comparable and "the
movement is entirely on our side."
Review objected that this does not follow,
and review was right.
The ORT arm reproducing shows the ORT ruler did not move.
It says nothing about the native ruler, which sits in a different binary and
changed repeatedly over the same window — 81e611c03 (#1722) is literally titled
"make the acc0 native and ORT arms measure one quantity".

So the inference was replaced with a measurement. e9754e7ef's tree is checked
out in a second worktree, its int4_decode_loop_ab rebuilt, and run beside
current main's on the same host, same environment, PROBE_REPS=1 on both so
neither gets a rep loop the other lacks, arms interleaved and the launch order
alternated:

width kernel-only, measured published pair implies verdict
1 1.64x [1.61–1.88], 12 paired cells 1.59x apparent movement is kernel
8 1.82x [1.78–1.89], 6 paired cells 3.08x 3.08x retracted

Retracting the t=8 3.08x

Both old figures reproduce today to within 0.4% — but only unpinned:

published rebuilt e9754e7ef today, unpinned delta
56.307 ms (t=1) 56.519 (56.402 / 56.519 / 56.878) +0.4%
14.091 ms (t=8) 14.115 (14.105 / 14.115 / 14.196) +0.2%

The old bench never called EpFactory::initialize, so it never ran
bound_process_to_decode_budget() and its process was never confined. That
function — physical-core select_budget_cpus included — already existed at
e9754e7ef
, and production always called it; only the bench was missing the
call, which #1766 11cb8e5f3 added. The old t=8 row therefore measured eight
decode workers scattered over 32 logical CPUs onto SMT siblings: a topology no
served session ever ran in.

binary at t=8 8 physical cores 4 cores + SMT siblings unpinned
e9754e7ef 8.430 ms 16.121 ms 14.115 ms
current main 4.664 ms 7.988 ms 4.619 ms

1.67x of the claimed 3.08x was placement, not kernel work. Today's binary is
pin-insensitive (0.99x) because it confines itself. This is exactly the effect
docs/benchmarks/2026-08-21-decode-worker-cpu-placement.md (#1680, ledger §24)
already recorded — landing on a number I quoted two days later.

Two corrections that look right and are not

Recorded so they are not re-applied at this site:

  • The ~11% warmup/spawn handicap of §27 is in tokens_s_total. Both published
    figures are ms_token — the old ort_matmulnbits_baseline.py docstring names
    "the native harness's steady column-2 median" as its comparand, and the
    reproductions above land on it to 0.4%. Deducting 11% from 56.307 yields a
    number no run of either tree produces.
  • The statistic asymmetry that is real points the other way. Old ORT took
    min over reps of a per-Run median; old native was single-shot with no rep
    loop. Best-of-N against single-shot flatters ORT, so it made the old gap look
    worse. Calling the two arms "the same statistic", as the first revision did,
    was wrong.

Headline-table defects fixed in this revision

  • Mixed statistics. The first table printed native median latency beside ORT
    throughput-equivalent and called the ratio a gap, so its columns did not yield
    its own gap figure. Both sides are now tokens_s_total; the mixed variant is
    shown, labelled, and noted to decline (1.113 → 1.098 → 1.085) rather than be flat.
  • False precision at t=4. 1.148x quoted against an A/A null of 0.868–1.150.
    Now "~1.15x, does not resolve".
  • Undisclosed post-hoc discard. "three independent launches per width" was
    false (3 / 6 / 3 / 3 cells across two invocations), and the t=1 1.4% spread depended
    on discarding a cell after seeing it. Retained-cell figures are now published
    beside the headline (1.112x [0.927–1.128], 18.1%, n=3 — the median barely
    moves, the precision does not survive), and the discard rule is stated
    prospectively for next time.
  • Reproducibility. acc0_gap_matrix.py gains --launches, a per-width
    --tokens 1:64,4:192,8:384 map, and a gap column in ORT÷native orientation
    beside ratio, so the Reproduce block names a command that produces the
    published table.

Method preconditions added

  1. The two arms were not getting the same machine. ONNX_GENAI_CPU_DECODE_THREADS=w
    confines the whole native process to w CPUs; the script pinned ORT to all 16
    even CPUs at every width. Measured effect on a quiet host: 1–2% — real, small,
    and now data rather than argument.
  2. The realized width is read back and checked (decode_width requested=4 realized=4 as_requested). Timings cannot detect a vacuous sweep.
  3. LoadWatch samples the runnable count during every arm, refusing above
    width + slack. A pre-check cannot see a competitor that arrives mid-cell — one
    did, and four cells were discarded because of it.
  4. The wide-pin arm turned out to be a contention detector. One t=1 cell passed
    every host-level guard while CPU 0 alone was busy: both matched-pin arms ~2x slow,
    the roaming arm normal. A single-CPU pin is the most fragile cell in any width
    sweep, and it is what every speedup is quoted against.

Still unresolved — and a correction to the first revision

t=16, and it is the row that matters. The first revision of this PR wrote
the width off as "every cell contaminated". That was wrong, and wrong in the
direction that flattered the conclusion.
Two of the three t=16 cells passed
the load guard cleanly (runnable 6, no competitor recorded):

t=16 gap A/A native spread ORT spread
launch 1 1.831 1.295 17.8% 55.4%
launch 2 1.456 0.969 27.7% 19.6%
median 1.643 — — —

So it is ~1.64x from two accepted cells, not "no data". It still does not
resolve, but on the correct ground: the A/A null spans 0.969–1.295 (±30%, against
3.6% at t=1 and 2.8% at t=8) and both arms are unstable at this width. The
cell the guard did refuse reads 1.585 — between the two retained — so the
discard is not load-bearing either way.

This is the one open cell that could reverse the re-ranking, and it needs a
dedicated quiet-host study with launch distributions and a pre-registered A/A
acceptance threshold.

Second-review fixes (this revision)

An adversarial review returned MERGE AFTER FIXES with all six core claims
surviving falsification and seven defects. All are fixed:

  1. t=16 mischaracterised — the headline fix above, propagated to all four
    sites that carried the re-ranking claim.
  2. Cell counts — t=4 is 4 trusted of 6 taken; the published table came
    from two script invocations, not three.
  3. Column semantics — harness-trust and editorial retention were conflated in
    one column; now separated, which makes the t=1 discard visible in the table
    rather than only in prose.
  4. t=1 placement probe samples published, including a 114.94 ms outlier on
    one pinned rep of three (the old binary's bimodality — it is why the t=1 A/B
    range reaches 1.88x). Placement is worth 1.9% at this width, against 1.67x at
    t=8, as expected for a single-threaded process with no SMT sibling to hit.
  5. ORT-side spread disclosed at t=16 (55.4%) — the denominator is no better
    behaved than the numerator there.
  6. Stale checksum constants in int4_decode_loop_ab's module doc corrected,
    with a note that they drift under reduction reassociation (perf(cpu): break the serial f32 reduction chain in the int4 decode GEMV (5.75x t=1) #1667, perf(cpu): delete the integer divisions from the int4 decode zero-point path #1783) and
    that the pattern — block 16 moves under ONNX_GENAI_CPU_MM_INT4_GEBP=0,
    block 32 does not — is the route evidence, not the digits.
  7. --tokens map footgun — a map missing a --threads width died with a bare
    KeyError after the first cell had already waited out the load guard. It now
    refuses at parse time, naming the missing widths.

Validation

Docs plus one benchmark harness; no library code. The harness was stub-validated
end to end (launch loop, per-width token map, gap column, paired per-width summary)
and both binaries were run for the numbers above. Required CI (Fast (Linux x86_64),
Rust quality) must be green before merge — no admin bypass.

…t=1/4/8

The acc0 int4 decode gap against ORT has been the top remaining CPU
MatMulNBits target on the strength of a `1.84x` figure. Re-measured on
`e189244ba` with a matched core budget, a checked realized width and three
independent launches per width, it is **1.120x at t=1, 1.148x at t=4 and
1.120x at t=8** — flat across the measurable range.

The old figure was not mislabelled. It was a correct measurement of a tree
that no longer exists: six merges landed after `e9754e7ef` published it,
three of them direct acc0 kernel work (#1667 broke the serial f32 reduction
chain, 5.75x t=1; #1679 enabled the register-blocked kernel *at
accuracy_level = 0*; #1783 folded the zero-point unpack).

The control is the ORT arm, which reproduces to +4.4% (30.632 -> 31.99 ms) on
the same host, binary, graph and statistic — so the comparison is sound and
the whole movement is on our side: native 56.307 -> 35.36 at t=1 and
14.091 -> 4.57 at t=8. A yesterday's note on this same figure said it was
`t=1`-only and that the production-width gap was unmeasured; both were true,
and both missed that the `t=1` number was itself three kernel merges old.

Three harness defects had to be fixed before the number meant anything, all in
`acc0_gap_matrix.py`:

* **The arms were not getting the same machine.**
  `ONNX_GENAI_CPU_DECODE_THREADS=w` confines the whole native process to `w`
  CPUs; the script pinned ORT to all 16 at every thread count. Measured in
  both configurations, the asymmetry is only 1-2% on a quiet host, which is
  recorded so it is not re-litigated.
* **The realized width was assumed, not checked.** A cell is now refused
  unless the binary reports `as_requested` (width 1's `path=flat` excepted --
  it builds no pool by design).
* **A pre-check cannot see a competitor that arrives mid-cell.** A sibling
  `cargo test` started during this matrix and four cells that had passed
  `wait_quiet` were measured against it. Every arm now runs inside a
  `LoadWatch` sampling the instantaneous runnable count, refused above
  `threads + slack`.

t=16 remains unresolved: every cell at that width was contaminated, and it is
the same width whose launch distribution spans 514% with no known mechanism.

Full record: docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.28%. Comparing base (b4e5797) to head (bd71d63).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1852      +/-   ##
==========================================
- Coverage   80.60%   80.28%   -0.32%     
==========================================
  Files         413      414       +1     
  Lines      199097   203528    +4431     
  Branches   199097   203528    +4431     
==========================================
+ Hits       160476   163410    +2934     
- Misses      33098    34558    +1460     
- Partials     5523     5560      +37     
Flag Coverage Δ
cli-ort-linux 72.51% <ø> (ø)
cli-ort-windows 72.01% <ø> (ø)
mlas 85.20% <ø> (-0.14%) ⬇️
offline 80.41% <ø> (-0.33%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 42 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Roy and others added 2 commits August 23, 2026 16:13
…=8 3.08x

Review of the previous commit made the right objection: the ORT arm
reproducing to +4.4% shows the *ORT* ruler did not move and says nothing
about the native one, which changed repeatedly over the same window
(#1722 is literally titled "make the acc0 native and ORT arms measure one
quantity"). So the inference is replaced with a measurement.

`e9754e7ef`'s bench is rebuilt in a second worktree and run beside current
main's -- same host, same environment, `PROBE_REPS=1` on both so neither
gets a rep loop the other lacks, arms interleaved and the order alternated:

| width | kernel-only, measured | published pair implies | verdict |
|------:|----------------------:|-----------------------:|---------|
| 1 | 1.64x [1.61-1.88], 12 cells | 1.59x | movement is kernel |
| 8 | 1.82x [1.78-1.89], 6 cells | 3.08x | 3.08x RETRACTED |

Both old figures reproduce to within 0.4% -- but only **unpinned**. The old
bench never called `EpFactory::initialize`, so it never ran
`bound_process_to_decode_budget()`; that function, physical-core selection
included, already existed at `e9754e7ef` and production always called it.
The old t=8 row therefore measured eight decode workers scattered over 32
logical CPUs onto SMT siblings -- a topology no served session ever ran in.
#1766 added the call. Pinned to eight physical cores the same old binary
gives 8.430 ms against its unpinned 14.115, and forced onto `0-7`, 16.121.
1.67x of the claimed 3.08x was placement, not kernel work. This is the
effect 2026-08-21-decode-worker-cpu-placement.md (#1680) already recorded,
landing on a number I quoted two days later.

Two corrections that look right and are not, recorded so they are not
re-applied: the ~11% warmup/spawn handicap of §27 is in `tokens_s_total`,
and both published figures are `ms_token` -- the old ORT harness docstring
names "the native harness's `steady` column-2 median" as its comparand and
the reproductions land on it. Deducting 11% yields a number neither tree
produces. The asymmetry that *is* real, old ORT `min` over reps against
old native single-shot, biases in ORT's favour.

The gap conclusion is unchanged and is measured on today's tree, both arms,
matched pins -- but the headline table is rebuilt to fix three defects:

- it mixed statistics (native median latency over ORT throughput-equivalent)
  so its columns did not yield its own gap figure. Both sides are now
  `tokens_s_total`, with the mixed variant shown and labelled;
- `t=4` was quoted as 1.148x when its A/A null spans 0.868-1.150, so the
  gap is inside its own noise floor there. Now "~1.15x, does not resolve";
- "three independent launches per width" was false (3/5/3 cells across two
  invocations) and the `t=1` 1.4% spread depended on an undisclosed post-hoc
  discard. Retained-cell figures are published beside the headline
  (1.112x [0.927-1.128], n=3) and the discard rule is stated prospectively.

`acc0_gap_matrix.py` gains `--launches`, a per-width `--tokens 1:64,4:192`
map and a `gap` column in ORT/native orientation beside `ratio`, so the
Reproduce block names a command that produces the published table.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby marked this pull request as ready for review August 23, 2026 16:15
@justinchuby

Copy link
Copy Markdown
Owner Author

Reviewed as validation, not style. The headline holds — t=1 and t=8 resolve independently and agree at 1.120x, and retracting the 3.08x by rebuilding e9754e7ef beside main rather than arguing from the ORT control is the right move; replacing an inference with a measurement is exactly what the earlier objection deserved. The placement attribution (1.67x of 3.08x) is well supported by the three-way pin table.

My finding is not about the number. It is about why two of your four cells did not resolve, and it has a mechanism your guards are structurally unable to see.

The pin set leaves every SMT sibling unmonitored

PIN = ",".join(str(c) for c in EVEN)     # 0,2,4,...,30

That is one logical CPU per physical core, which is right for the measurement. But it means the 16 odd CPUs — the SMT sibling of every pinned CPU — are outside the monitored set entirely, and wait_quiet/LoadWatch gate on machine-wide loadavg and runnable count, neither of which can be attributed to that shadow set.

Measured just now on this host, under hostlock, two hogs placed on the siblings of two pinned cores:

pinned=[20, 22]   UNMONITORED siblings=[21, 23]

throughput on pinned cpu 20:  solo=8.27  sibling-busy=5.58   ratio=0.674x
CPU-efficiency guard:         solo=1.000 sibling-busy=1.000  (moves 0.000)
loadavg:                      before=4.01  during=3.67       delta=-0.34

A competitor that costs 33% of throughput moves the efficiency guard by 0.000 and moved loadavg the wrong way. The -0.34 is ambient drift swamping two hogs I started myself and knew the PIDs of — which is the point: a 1-minute machine-wide average cannot be attributed to sixteen specific CPUs. (This reproduces 0.67x from an independent pair measured on a different day; it is a stable property of SMT, not a fluke.)

To be precise and not overstate: at this moment ambient load is 3.67, so wait_quiet(threshold=3.0) would block. The problem is not that the threshold is always permissive — it is that the reading is uncorrelated with the harm. On an otherwise quiet host, threshold=3.0 admits ~3 runnable threads, which is enough to seat a hog on the sibling of three pinned cores while the gate reads clean.

Why this predicts exactly your failure pattern

Exposure scales with width: at width w the run occupies w physical cores and is therefore exposed to w unmonitored siblings. t=1 risks one sibling; t=16 risks sixteen. Your table:

width trusted cells A/A range
1 2 of 3 1.025–1.036
4 4 of 5 0.868–1.150
8 3 of 3 0.997–1.028
16 0 of 3 —

t=16 failing to resolve at all is the cell with maximal sibling exposure. I am not claiming that is the cause — I have not reproduced your matrix — but it is a specific, testable mechanism that your instruments cannot currently rule out, and "not resolvable" is presently unexplained.

There is a real asymmetry worth knowing here, and it partly protects you: steady sibling load depresses the native and ORT arms alike, so the paired ratio survives and the A/A null stays tight — which is why a tight null is not evidence of a quiet host. Intermittent sibling load that lands between the native and ORT arms of one launch corrupts the ratio directly, and produces precisely the wide, non-resolving spread seen at t=4. Your absolute columns (27.9 / 107.2 / 211.0 tok/s) get no ratio protection at all.

Two asks

  1. Hold the lock. scripts/hostlock.sh (Add an advisory host lock for benchmark runs #1806) is mandatory for repo-owned saturating runs, and the clarification is explicit that taskset-bounding does not exempt a run. This matrix is a full width sweep and takes none of it. This is also mutual: an unlocked matrix is invisible to everyone else's gate, and unannounced saturating runs cost Roy a full prefill A/B earlier this week.
  2. Gate on the shadow set, not on the machine. Sample per-CPU busy over siblings(PIN) across each cell and discard cells where any sibling was busy. It is a /proc/stat delta, it localizes, and unlike loadavg it cannot be confounded by ambient drift. I have this in a standalone harness and will hand it over rather than have you rebuild it.

This is the second harness with the same blind spot — #1814's foreign-CPU column excludes every SMT sibling by construction (#1824), and provenance still records no scope/cpuset so the shadow set isn't even recoverable after the fact (#1846). I'd rather fix it once in a shared place than three times.

Verdict: REQUEST CHANGES — bounded, on methodology only. The 1.12x headline and the 3.08x retraction I have no objection to; add the lock and a shadow-set gate, and re-run only the cells that did not resolve.

— Pris

justinchuby and others added 2 commits August 23, 2026 17:24
…on them

Second adversarial review found that this record wrote off width 16 as
"0 of 3 cells, all taken against a sibling cargo test". That is false, and
false in the direction that flattered the conclusion: two of the three t=16
cells passed the load guard cleanly (runnable 6, no competitor recorded) and
read 1.831 and 1.456, median 1.643x.

The width still does not resolve, but for the honest reason -- its A/A null
spans 0.969-1.295 (against 3.6% at t=1 and 2.8% at t=8) and both arms show
20-55% intra-run spread -- not because the data was contaminated. The cell the
guard did refuse reads 1.585, between the two retained, so the discard is not
load-bearing either way.

Because t=16 is the width closest to an unconfined production process, the
re-ranking claim ("acc0 is no longer the top CPU MatMulNBits target") is now
stated as conditional on it, in all four places it appears, with the
quiet-host t=16 study named as the thing that settles it.

Also from the review:

- Headline cell counts corrected: t=4 is 4 trusted of 6 taken (was "4 of 5"),
  and the published table came from two script invocations, not three.
- The "trusted cells" column split into harness-trust vs editorial retention.
  All three t=1 cells passed the guard; one was dropped afterwards by me, and
  conflating the two hid that.
- The t=1 placement probe's samples are now published rather than summarised,
  including the 114.94 ms outlier on one pinned rep -- that excursion is the
  old binary's bimodality and is why the t=1 A/B range reaches 1.88x.
- t=16's ORT arm spread (55.4%) disclosed alongside native's; the denominator
  is no better behaved than the numerator at that width.
- The four checksum falsifier constants in int4_decode_loop_ab's module doc
  were stale. Corrected, with a note that they drift with reduction
  reassociation and that the *pattern* -- block 16 moves under
  ONNX_GENAI_CPU_MM_INT4_GEBP=0, block 32 does not -- is the route evidence.
- acc0_gap_matrix.py: a --tokens map missing a --threads width died with a
  bare KeyError after the first cell had already waited out the load guard.
  It now refuses the run at parse time with the missing widths named.
- Reproduce block updated to the two invocations actually run, so the commands
  as written pass the new validation.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby enabled auto-merge (squash) August 23, 2026 17:28
@justinchuby
justinchuby merged commit 4b4dacc into main Aug 23, 2026
15 of 18 checks passed
@justinchuby
justinchuby deleted the squad/roy-acc0-gap-remeasure branch August 23, 2026 17:39
@justinchuby

Copy link
Copy Markdown
Owner Author

Merged as 4b4dacc7e with both required checks green (Fast (Linux x86_64), Rust quality) via normal auto-merge — no admin bypass.

Disposition of the second review's seven defects — all fixed in 9b9c355ee:

  1. t=16 mischaracterised (the one that mattered). I had written the width off as "0 of 3 cells, all contaminated". Two of three passed the load guard cleanly and read 1.831 / 1.456, median 1.643x. It still does not resolve — A/A null 0.969–1.295, both arms 20–55% intra-run spread — but on the correct ground. The re-ranking claim is now stated as conditional on t=16 in all four places it appeared, and the refused cell (1.585, between the two retained) is disclosed so the discard is visibly not load-bearing.
  2. Cell counts: t=4 is 4 trusted of 6; the table came from two invocations, not three.
  3. Harness-trust and editorial retention split into separate column semantics — all three t=1 cells passed the guard; one was dropped by me afterwards.
  4. t=1 placement probe samples published rather than summarised, including the 114.94 ms outlier on one pinned rep of three. Placement is worth 1.9% at t=1 against 1.67x at t=8 — expected, since a single-threaded process has no SMT sibling to collide with.
  5. ORT-side spread at t=16 disclosed (55.4%): the denominator is no better behaved than the numerator there.
  6. All four checksum falsifier constants in int4_decode_loop_ab's module doc were stale. Corrected, with a note that they drift under reduction reassociation (perf(cpu): break the serial f32 reduction chain in the int4 decode GEMV (5.75x t=1) #1667, perf(cpu): delete the integer divisions from the int4 decode zero-point path #1783) and that the pattern — block 16 moves under ONNX_GENAI_CPU_MM_INT4_GEBP=0, block 32 does not — is the route evidence, not the digits.
  7. --tokens map missing a --threads width died with a bare KeyError after the first cell had already waited out the load guard. It now refuses at parse time naming the missing widths; the Reproduce block was rewritten to the two invocations actually run so the published commands pass the new validation.

What this leaves open, and it is the important part. t=16 is the width closest to an unconfined production process and the only one pointing at a large gap. A confirmed ~1.64x there reverses the re-ranking and puts acc0 back at the top. That study — launch distributions, pre-registered A/A acceptance threshold, quiet host — is my next piece of work, and I would rather it be run than have the ~1.12x headline quoted as though it covered every width.

Thanks for catching the t=16 write-off. It was the single claim in the PR that moved the conclusion, and I had made it in the direction that suited the conclusion.

justinchuby added a commit that referenced this pull request Aug 23, 2026
…ranking (#1871)

## The `t=16` condition I merged this morning has failed

`4b4dacc7e` (#1852) re-measured the acc0 (`accuracy_level = 0`,
production
default) int4 decode gap, found the published **1.84x** stale, measured
**1.120x at t=1 and t=8**, and concluded acc0 was no longer the top CPU
MatMulNBits target — **explicitly conditional on `t=16`**, which that
matrix
could not resolve and which reviewers rightly pushed me to stop burying.

This resolves it. **The gap at `t=16` is ~1.78x and acc0 goes back to
the top of
the work list.**

| statistic | gap (ORT ÷ native) | n |
|---|---:|---:|
| run 1, paired median (384 tok/rep) | **1.782** [1.409–1.813] | 6
trusted of 14 |
| run 2, paired median (768 tok/rep) | **1.773** [1.279–2.176] | 11
trusted of 16 |
| pooled | **1.775** [1.279–2.176] | 17 |
| best launch vs best launch | **1.770** (255.4 → 452.1 tok/s) | 17 |
| the half of cells where native ran **fastest** | **1.650** | 9 |
| separate same-session run, different harness | **1.793** [1.375–1.968]
| 10 |

## The pre-registered rule said RANGE ONLY and that verdict is honoured

The acceptance rule is written into the harness **before its first
run**:

```
ACCEPT iff n_trusted >= 8 AND aa_halfwidth <= 0.10 AND |gap-1| >= 3*aa_halfwidth
RANGE ONLY if n is met but the null is too wide
REPORT NOTHING if n is not met
```

Run 1 returned **REPORT NOTHING** (n=6). Run 2 returned **RANGE ONLY**
(A/A half-width 0.377). **Neither run is re-scored under a looser rule
and no
point estimate is claimed.** The range is genuinely wide — 1.28x to
2.18x — but
**the lowest single cell of 30 still exceeds the `t=8` figure of
1.120x.**

Four reasons the central tendency survives the wide null:

1. **The null is symmetric between the arms**, so it is the width, not
either
   implementation. The study alternates which arm is doubled: native A/A
half-width 0.347 (n=15), **ORT's 0.377** (n=15) — ORT's is the wider.
The
original matrix only ever doubled native and could not tell these apart.
2. **Two runs agree to 0.5%** on the median across a doubled token
budget.
3. **Best-vs-best**, the statistic least sensitive to transient
contention,
   gives 1.770x on top of the paired median's 1.775x.
4. **Even native's best mode is 1.650x.** The width-16 bimodality is
real (slow
   half 1.834x) but the gap does not vanish when native runs well.

## Mechanism: the scaling wall is ours, and it is not bandwidth

A second harness runs **both widths inside the same launch**, width
order
rotated. Its rule is also pre-registered and is deliberately a **sign
test on
paired launches**, not a comparison of medians — a ±35% null can move a
median-of-ratios but cannot flip a per-launch sign symmetrically.

| arm | `t=8 → t=16` scaling | range |
|---|---:|---|
| ORT | **1.762x** | 1.485–1.932 |
| native | **1.319x** | 1.131–1.488 |

**Native scaled worse in 10 of 10 paired launches (100%, threshold
80%).** ORT
converts the doubling into 1.76x, near the 2.0x ideal; we get 1.32x and
lose a
third of the added width.

**This rules out the documented bandwidth knee**
(`2026-08-22-decode-width-scaling.md`, 7.52x → 9.26x at t=16). A DRAM
ceiling is
a property of the *host* and would flatten both arms. ORT scales 1.76x
across
the same doubling, on the same host, in the same launch. Whatever binds
us at
`t=16` is in our pool or our kernel.

## Two things recorded rather than quietly fixed

- **Run 1's guard was wrong and its verdict stands.** It refused 8 of 14
cells,
and **none of the refusals was contention** — `competing_load()` was
empty for
all 14 and the pre-check runnable count was 2–4 throughout. The ceiling
came
from a structural estimate (22); the cell's own peak reaches 25. Run 2's
ceiling of 26 is fitted to run 1's peaks, which is **post-hoc with
respect to
  that run**, and is disclosed as such rather than presented as a fresh
  derivation. Run 1's `REPORT NOTHING` is not re-scored.
- **The scaling run's `t=8` gap reads 1.264x, not 1.120x.** The
difference is on
the native side (190 vs 211 tok/s; ORT unchanged at 232–240 vs 238),
most
likely because that harness runs a `t=16` arm beside the `t=8` arm in
one
launch. Published, and **explicitly not offered as a correction** to
1.120x.
It does not touch the scaling result, which is paired within each
launch.

## Also fixed: the pre-check instrument

Both harnesses gate on the **instantaneous runnable count**, not
`getloadavg()`.
This study is the demonstration: right after its own `cargo build` the
load
average sat at **6.86** on a host whose runnable count was **2**, and a
`threshold=3.0` pre-check slept through a fully quiet window without
measuring
anything. Gating the pre-check on the same quantity `LoadWatch` polices
during
the arm makes the two agree.

## Files

- `docs/benchmarks/2026-08-23-acc0-gap-at-width-16.md` — the record
- `crates/onnx-runtime-ep-cpu/benches/acc0_w16_study.py` — the gap study
- `crates/onnx-runtime-ep-cpu/benches/acc0_w8_w16_scaling.py` — the
scaling test
- `2026-08-23-acc0-gap-vs-ort-by-width.md`, `CPU_MATMUL_ASSIGNMENT.md`,
and two
  neighbouring benchmark docs updated where they carried the now-failed
  condition

## Validation

Docs plus two benchmark harnesses; no library code. Both harnesses were
stub-validated across every verdict branch before touching the host, and
both
were run for the numbers above on `0f84888b8`. Required CI (`Fast (Linux
x86_64)`, `Rust quality`) must be green before merge — no admin bypass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant