Skip to content

perf(docs): the acc0 gap at width 16 is ~1.78x — reverses today's re-ranking - #1871

Merged
justinchuby merged 2 commits into
mainfrom
squad/roy-acc0-w16
Aug 23, 2026
Merged

justinchuby merged 2 commits into
mainfrom
squad/roy-acc0-w16

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

The t=16 condition I merged this morning has failed

4b4dacc7e (#1852) re-measured the acc0 (accuracy_level = 0, production
default) int4 decode gap, found the published 1.84x stale, measured
1.120x at t=1 and t=8, and concluded acc0 was no longer the top CPU
MatMulNBits target — explicitly conditional on t=16, which that matrix
could not resolve and which reviewers rightly pushed me to stop burying.

This resolves it. The gap at t=16 is ~1.78x and acc0 goes back to the top of
the work list.

statistic gap (ORT ÷ native) n
run 1, paired median (384 tok/rep) 1.782 [1.409–1.813] 6 trusted of 14
run 2, paired median (768 tok/rep) 1.773 [1.279–2.176] 11 trusted of 16
pooled 1.775 [1.279–2.176] 17
best launch vs best launch 1.770 (255.4 → 452.1 tok/s) 17
the half of cells where native ran fastest 1.650 9
separate same-session run, different harness 1.793 [1.375–1.968] 10

The pre-registered rule said RANGE ONLY and that verdict is honoured

The acceptance rule is written into the harness before its first run:

ACCEPT iff n_trusted >= 8 AND aa_halfwidth <= 0.10 AND |gap-1| >= 3*aa_halfwidth
RANGE ONLY if n is met but the null is too wide
REPORT NOTHING if n is not met

Run 1 returned REPORT NOTHING (n=6). Run 2 returned RANGE ONLY
(A/A half-width 0.377). Neither run is re-scored under a looser rule and no
point estimate is claimed.
The range is genuinely wide — 1.28x to 2.18x — but
the lowest single cell of 30 still exceeds the t=8 figure of 1.120x.

Four reasons the central tendency survives the wide null:

  1. The null is symmetric between the arms, so it is the width, not either
    implementation. The study alternates which arm is doubled: native A/A
    half-width 0.347 (n=15), ORT's 0.377 (n=15) — ORT's is the wider. The
    original matrix only ever doubled native and could not tell these apart.
  2. Two runs agree to 0.5% on the median across a doubled token budget.
  3. Best-vs-best, the statistic least sensitive to transient contention,
    gives 1.770x on top of the paired median's 1.775x.
  4. Even native's best mode is 1.650x. The width-16 bimodality is real (slow
    half 1.834x) but the gap does not vanish when native runs well.

Mechanism: the scaling wall is ours, and it is not bandwidth

A second harness runs both widths inside the same launch, width order
rotated. Its rule is also pre-registered and is deliberately a sign test on
paired launches
, not a comparison of medians — a ±35% null can move a
median-of-ratios but cannot flip a per-launch sign symmetrically.

arm t=8 → t=16 scaling range
ORT 1.762x 1.485–1.932
native 1.319x 1.131–1.488

Native scaled worse in 10 of 10 paired launches (100%, threshold 80%). ORT
converts the doubling into 1.76x, near the 2.0x ideal; we get 1.32x and lose a
third of the added width.

This rules out the documented bandwidth knee
(2026-08-22-decode-width-scaling.md, 7.52x → 9.26x at t=16). A DRAM ceiling is
a property of the host and would flatten both arms. ORT scales 1.76x across
the same doubling, on the same host, in the same launch. Whatever binds us at
t=16 is in our pool or our kernel.

Two things recorded rather than quietly fixed

  • Run 1's guard was wrong and its verdict stands. It refused 8 of 14 cells,
    and none of the refusals was contention — competing_load() was empty for
    all 14 and the pre-check runnable count was 2–4 throughout. The ceiling came
    from a structural estimate (22); the cell's own peak reaches 25. Run 2's
    ceiling of 26 is fitted to run 1's peaks, which is post-hoc with respect to
    that run
    , and is disclosed as such rather than presented as a fresh
    derivation. Run 1's REPORT NOTHING is not re-scored.
  • The scaling run's t=8 gap reads 1.264x, not 1.120x. The difference is on
    the native side (190 vs 211 tok/s; ORT unchanged at 232–240 vs 238), most
    likely because that harness runs a t=16 arm beside the t=8 arm in one
    launch. Published, and explicitly not offered as a correction to 1.120x.
    It does not touch the scaling result, which is paired within each launch.

Also fixed: the pre-check instrument

Both harnesses gate on the instantaneous runnable count, not getloadavg().
This study is the demonstration: right after its own cargo build the load
average sat at 6.86 on a host whose runnable count was 2, and a
threshold=3.0 pre-check slept through a fully quiet window without measuring
anything. Gating the pre-check on the same quantity LoadWatch polices during
the arm makes the two agree.

Files

  • docs/benchmarks/2026-08-23-acc0-gap-at-width-16.md — the record
  • crates/onnx-runtime-ep-cpu/benches/acc0_w16_study.py — the gap study
  • crates/onnx-runtime-ep-cpu/benches/acc0_w8_w16_scaling.py — the scaling test
  • 2026-08-23-acc0-gap-vs-ort-by-width.md, CPU_MATMUL_ASSIGNMENT.md, and two
    neighbouring benchmark docs updated where they carried the now-failed
    condition

Validation

Docs plus two benchmark harnesses; no library code. Both harnesses were
stub-validated across every verdict branch before touching the host, and both
were run for the numbers above on 0f84888b8. Required CI (Fast (Linux x86_64), Rust quality) must be green before merge — no admin bypass.

…ranking

PR #1852 (`4b4dacc7e`) established that the published 1.84x acc0 gap was stale
and measures 1.120x at t=1 and t=8, and concluded acc0 was no longer the top CPU
MatMulNBits target -- explicitly conditional on t=16, which that matrix could not
resolve. This resolves it. The condition fails.

Two runs on `0f84888b8`, 30 launches, acceptance rule written into the harness
before its first run:

  run 1 (384 tok)      1.782x [1.409-1.813]   6 trusted of 14
  run 2 (768 tok)      1.773x [1.279-2.176]  11 trusted of 16
  pooled               1.775x                17
  best-vs-best         1.770x  (255.4 -> 452.1 tok/s)
  native's fast half   1.650x                 9

The pre-registered rule returned REPORT NOTHING on run 1 (n=6) and RANGE ONLY on
run 2 (A/A half-width 0.377 > 0.10). Both verdicts are honoured and neither run
is re-scored: no point estimate is claimed. The direction and order of magnitude
are not in doubt -- the lowest single cell of 30, 1.279x, still exceeds the t=8
figure of 1.120x.

Why it survives a null that wide: the null is symmetric between the arms (native
0.347, ORT 0.377 -- ORT's is the wider), the two runs agree to 0.5% across
different token budgets, a contamination-resistant best-vs-best statistic agrees
at 1.770x, and even the half of cells where native ran fastest gives 1.650x.

The mechanism, measured rather than inferred. A second harness runs both widths
inside the same launch with the width order rotated, and its acceptance rule is
a sign test on paired launches rather than a comparison of medians, because a
+-35% null can move a median-of-ratios but cannot flip a per-launch sign
symmetrically:

  ORT     t=8 -> t=16 scaling  1.762x  [1.485-1.932]
  native  t=8 -> t=16 scaling  1.319x  [1.131-1.488]
  native scaled worse in 10 of 10 paired launches (100%, threshold 80%)

So the width-16 gap is not a new kernel deficiency at that width: it is the t=8
gap plus a scaling failure that is ours alone. That run also reproduces the
width-16 gap at 1.793x from a different harness and token budget -- a fourth
independent estimate.

This rules out the host bandwidth knee documented in
`2026-08-22-decode-width-scaling.md`. A DRAM ceiling is a property of the host
and would flatten both arms; ORT scales 1.76x across the same doubling on the
same host in the same launch.

Two things recorded rather than quietly fixed:

- Run 1's guard refused 8 of 14 cells and none of the refusals was contention --
  `competing_load()` was empty for all 14 and the pre-check runnable count was
  2-4 throughout. The ceiling was set from a structural estimate (22) and the
  cell's own peak reaches 25. Run 2's ceiling of 26 is fitted to run 1's peaks,
  which is post-hoc with respect to that run, and it is disclosed as such; run
  1's REPORT NOTHING stands as issued.
- The scaling run's t=8 gap reads 1.264x against the matrix's 1.120x. The
  difference is on the native side (190 vs 211 tok/s; ORT is unchanged) and is
  most likely this harness running a t=16 arm beside the t=8 arm in the same
  launch. It is published, and explicitly not offered as a correction to 1.120x.

Both harnesses gate their pre-check on the instantaneous runnable count rather
than `getloadavg()`. This study is why: immediately after its own cargo build
the load average sat at 6.86 on a host whose runnable count was 2, and a
threshold=3.0 pre-check slept through a fully quiet window without measuring
anything.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby enabled auto-merge (squash) August 23, 2026 18:46
@codecov

codecov Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.92%. Comparing base (182d1f7) to head (4cf7d7a).
⚠️ Report is 5 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff           @@
##             main    #1871   +/-   ##
=======================================
  Coverage   80.92%   80.92%           
=======================================
  Files         415      415           
  Lines      203912   203912           
  Branches   203912   203912           
=======================================
+ Hits       165021   165024    +3     
+ Misses      33328    33325    -3     
  Partials     5563     5563           
Flag Coverage Δ
cli-ort-linux 72.51% <ø> (ø)
cli-ort-windows 72.01% <ø> (ø)
mlas 85.33% <ø> (ø)
offline 81.08% <ø> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 3 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Picks up the catalog-pin repair (#1870/#1872) that had every PR on the repo
red: #1860 registered pkg.nxrt::KvCacheCapacityAppend without bumping
expanded_registry_catalog_count_is_pinned. Nothing in this branch touches the
shape-inference registry -- the failure was inherited from main, reproduced
locally on a clean checkout, and is not fixed here.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby merged commit 06de4d8 into main Aug 23, 2026
15 of 18 checks passed
@justinchuby
justinchuby deleted the squad/roy-acc0-w16 branch August 23, 2026 20:28
justinchuby added a commit that referenced this pull request Aug 23, 2026
…te the idle (#1887)

## What this answers

[#1871](#1871) established
that the acc0 gap is ~1.78x at width 16 versus 1.12x at width 8, and
left the mechanism open: **when the `t=8 → t=16` doubling returns 1.32x
instead of 2x, are the extra workers idle, or busy and inefficient?**
Wall-clock timing cannot separate those. This adds CPU-seconds
attribution to both arms and answers it.

## The instrument

`process_cpu_time()` in `benches/common` reads `/proc/self/stat`
(thread-group user/sys ticks); `ort_matmulnbits_baseline.py` brackets
`getrusage(RUSAGE_SELF)` over the same window and emits
identically-named fields, so one parser reads both arms. Read directly
rather than through `/usr/bin/time`, whose `Percent of CPU` is
`(user+sys)/wall` — it looks like independent corroboration of a
wall-time result and is actually the same measurement divided by itself.

The decomposition is an **identity**, not a model:

```
tps(16) / tps(8)  ==  2 * R_busy / R_cpu
```

so residual is 0.00% by construction and it doubles as a free per-cell
self-test. Cells failing it are discarded as instrument faults rather
than reported.

**It earned its place before it was used in anger.** The harness's first
run returned `REPORT NOTHING (n_trusted = 1 < 6)`, which was honoured —
nothing was quoted, including the one trusted cell. Diagnosing it found
a defect in *my instrument*, not the host: each quantity was reduced by
its own independent median, and since `tps = tokens/wall` the two sort
in reversed orders, so at an even rep count they selected **different
repetitions**. Identity errors of **4.4%–29.7%** against a quantity that
is algebraically zero. Both producers now emit every CPU field from the
single median-throughput repetition. **No threshold in the
pre-registered rule was changed** — only the instrument feeding it, and
the change was forced by a self-test firing before anything was scored.

## Finding 1 — the CPU inflation is ours, not the host's

13 trusted of 14 launches, identity error 0.00% on every cell:

| arm | speedup 8→16 | `R_cpu` | `R_busy` | busy@8 | busy@16 |
sys_frac@8 | @16 |
|---|---:|---:|---:|---:|---:|---:|---:|
| native | 1.445 | **1.449** | 1.057 | 0.900 | 0.966 | 0.062 | **0.212**
|
| ORT | 1.860 | **1.074** | 0.999 | 1.000 | 0.999 | 0.000 | 0.000 |

ORT ran interleaved, same launches, same 16 cores, same minutes. If the
knee were a memory-bandwidth ceiling, ORT would hit it too. **This kills
the DRAM-plateau attribution.**

## Finding 2 — `busy` is blind to spin-wait, and the shipped wait path
spins

`decode_spmd`'s worker wait spins then `sched_yield`s for up to 500 µs
(`ONNX_GENAI_CPU_DECODE_BLOCKTIME_US`) before parking. A yielding thread
accrues CPU time exactly like a working one.

Dose-response first, because an env var that never reaches the child
produces a beautifully consistent null — w=16, 192 tokens: `sys_s` =
**0.13 / 3.28 / 3.12** at 0 / 500 / 20000 µs while `user_s` = **9.28 /
9.27 / 8.99**. Knob live, ramp saturated by 500 µs, and `user_s`
invariant — so the system time is pure overhead, not work.

Pre-registered A/B, 7 trusted of 10:

| w | ratio (bt0 ÷ bt500) | A/A null | busy@500 | busy@0 | cpu_s/tok
@500 | @0 |
|---:|---:|---:|---:|---:|---:|---:|
| 16 | **0.9960** | 0.9937 | 0.953 | **0.692** | 0.06037 | 0.04838 |
| 8 | 0.9906 | 1.0186 | 0.957 | 0.916 | 0.04078 | 0.03794 |

**Throughput: REJECT** (43% sign consistency against a 5.24% A/A
half-width). Removing the ramp does not make this workload faster, and
the favourable mechanism data does not change that. No regression at
w=8.

## What finding 2 does to finding 1

BURN-DOMINATED keys on `R_busy ≥ 0.90` — "the workers are not idle".
That quantity was masked. Same harness, same unchanged rule,
`--blocktime 0`, 9 trusted of 12, identity 0.00%:

| arm | `R_cpu` | `R_busy` | busy@8 | busy@16 |
|---|---:|---:|---:|---:|
| native | **1.304** | **0.652** | 0.938 | **0.595** |
| ORT | 1.109 | 0.999 | 1.000 | 0.999 |

**VERDICT: MIXED**, 100% sign consistency on *both* ratios.

| configuration | busy@16 | reads as |
|---|---:|---|
| default (500 µs ramp) | 0.966 | pool nearly fully occupied → BURN |
| ramp off | **0.595** | **40% of the pool is not working** |

**The correct attribution is mixed: ~30% more real CPU per token at
width 16 *and* ~40% of the sixteen cores idle.** Both halves are real,
both are ours.

A useful side-effect: over the same 9 cells the wall-derived speedup
swings **0.78x–1.37x** while `R_cpu` holds 1.10–1.34 and `R_busy`
0.52–0.81, both 100% sign-consistent. A competing process steals our
wall clock but does not add to our CPU seconds — which is the whole
reason for the instrument.

## Scope limit, stated up front

This is a **zero-gap** decode loop — exactly the workload where parking
early looks free. **Nothing here licenses changing the shipped blocktime
default**, which exists to protect latency when there *are* gaps; that
needs the gap-aware harness (#1395). What it does establish is that
**any occupancy reading of the decode pool taken at the default
blocktime over-reads by tens of points**, which is now recorded in the
ledger.

This independently corroborates, on a second workload, the
~20%-of-process-CPU figure @sebastian reported for the `worker_wait`
yield ramp — the cost side is confirmed; the latency side remains his.

## Next

1. Localise the **idle** half — `ONNX_GENAI_CPU_DECODE_WORKER_PROFILE` /
#1859's per-worker straggler attribution, to separate load imbalance
from dispatch/wake latency.
2. Localise the **burn** half — per-op attribution at w=8 vs w=16
against the 136.3 MB/token figure.

## Validation

- `cargo build --release -p onnx-runtime-ep-cpu --bench
int4_decode_loop_ab` ✅
- `cargo fmt --check` ✅, `cargo clippy` clean on the touched bench ✅
- `ruff check` clean on all touched `.py` (the one repo-wide F401 is
pre-existing in `ort_baseline.py`, untouched here)
- Harness branch-validators run before use: 13 checks / 0 failures
(cpu-split), 0 failures (blocktime A/B)
- All measurement runs taken under `scripts/hostlock.sh` with
announce-before/after

Both harnesses accept `--replay <json>` to re-score archived data, and
both carry their acceptance rule in the module docstring where it was
written before the first run.


---

# Update: the idle half is now attributed

Follow-up 1 from the "Next" section above is **done in this PR** (commit
`53d3ff076`). `int4_decode_loop_ab` now brackets `SpmdWorkerProfile`
deltas over exactly the window that produces `wall`, and
`acc0_w16_worker_split.py` scores them against a rule written before its
first run. **10 launches, 10 trusted, both fired conditions at 100% sign
consistency.**

**Wake latency is not the problem.** `wake_frac` is 0.006 at w=8 and
**0.051** at w=16; the pre-registered WAKE-BOUND condition did **not**
fire. A perfect wait/wake path recovers at most 5 points.

The mean worker's window at width 16:

| | w=8 | w=16 |
|---|---:|---:|
| useful work | 0.886 | **0.492** |
| straggler wait | 0.031 | **0.222** |
| wake latency | 0.006 | 0.051 |
| dispatcher / serial | 0.077 | 0.235 |

**Not all of that is a defect.** Holding w=8's serial time constant and
halving its parallel time — Amdahl with no defect anywhere — predicts a
0.204 serial share at w=16; observed is 0.235. So:

| component | points | recoverable? |
|---|---:|---|
| straggler wait | 22.2 | **yes** |
| serial in excess of Amdahl | 3.1 | maybe |
| wake latency | 5.1 | partly |
| Amdahl-predicted serial | 20.4 | **no** |

Quoting the 46% residual as recoverable would be wrong by more than 2x.
The same calibration also **rules out pure Amdahl** as the explanation
for the knee: it predicts 0.796 useful work and the pool delivers 0.492.

**The straggler** holds **72%** of last-arrivals against a 6.7% chance
share, does **1.5x** the median work, and almost never parks. Two
candidate mechanisms were tested and are **negative**: not a static
mis-partition (segments divide evenly for every shape here, and the
straggler's identity moves between launches), and not the unpinned
dispatcher colliding with a worker (dispatcher CPU → straggler CPU over
four profiled launches: `30→22`, `6→24`, `6→20`, `{30,20,18}→18`). **No
mechanism is claimed.**

**Nothing can help it today.** `DEFAULT_STEAL_TILES_PER_WORKER = 1`
makes `target == total_workers`, so `work_stealing_segments_aligned`
always falls back to static equal segments — there are no spare tiles to
steal.

## A +23% candidate, rejected

The blocktime A/B harness is generalised to vary any single env knob
(thresholds, A/A arm, rotation and the w=8 guard untouched; replaying
the archived blocktime JSON through the refactored scorer reproduces the
published numbers exactly). Pointed at that default, steal tiles 1 vs 2,
blocktime held at the shipped 500 µs in both arms, 8 trusted:

| w | ratio | sign | A/A | sys_frac | cpu_s/tok |
|---:|---:|---:|---:|---|---|
| 16 | **1.2327** | 88% | 1.1301 | 0.280 → 0.192 | −11.7% |
| 8 | 1.0463 | — | 1.0211 | 0.038 → 0.030 | −2.5% |

**REJECTED and not proposed.** The rule requires the effect to clear 3x
the A/A half-width, and that half-width is **0.2154** in the same run.
Re-running until the null comes in narrow would be choosing the sample
that licenses the conclusion, so it is not done. The mechanism moving in
the predicted direction, and the absence of the expected w=8 regression,
are recorded as observations rather than as a result.

**This makes the width-16 A/A null the binding constraint** — ±21.5%
here, ±35% in the earlier width-16 study — large enough that no
improvement of realistic size can clear a pre-registered bar at this
width. It moves ahead of further kernel work in the ledger.

New record:
`docs/benchmarks/2026-08-23-acc0-width-16-worker-attribution.md`.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant