Repository navigation
perf(docs): the acc0 gap at width 16 is ~1.78x — reverses today's re-ranking - #1871
Merged
Merged
Conversation
…ranking PR #1852 (`4b4dacc7e`) established that the published 1.84x acc0 gap was stale and measures 1.120x at t=1 and t=8, and concluded acc0 was no longer the top CPU MatMulNBits target -- explicitly conditional on t=16, which that matrix could not resolve. This resolves it. The condition fails. Two runs on `0f84888b8`, 30 launches, acceptance rule written into the harness before its first run: run 1 (384 tok) 1.782x [1.409-1.813] 6 trusted of 14 run 2 (768 tok) 1.773x [1.279-2.176] 11 trusted of 16 pooled 1.775x 17 best-vs-best 1.770x (255.4 -> 452.1 tok/s) native's fast half 1.650x 9 The pre-registered rule returned REPORT NOTHING on run 1 (n=6) and RANGE ONLY on run 2 (A/A half-width 0.377 > 0.10). Both verdicts are honoured and neither run is re-scored: no point estimate is claimed. The direction and order of magnitude are not in doubt -- the lowest single cell of 30, 1.279x, still exceeds the t=8 figure of 1.120x. Why it survives a null that wide: the null is symmetric between the arms (native 0.347, ORT 0.377 -- ORT's is the wider), the two runs agree to 0.5% across different token budgets, a contamination-resistant best-vs-best statistic agrees at 1.770x, and even the half of cells where native ran fastest gives 1.650x. The mechanism, measured rather than inferred. A second harness runs both widths inside the same launch with the width order rotated, and its acceptance rule is a sign test on paired launches rather than a comparison of medians, because a +-35% null can move a median-of-ratios but cannot flip a per-launch sign symmetrically: ORT t=8 -> t=16 scaling 1.762x [1.485-1.932] native t=8 -> t=16 scaling 1.319x [1.131-1.488] native scaled worse in 10 of 10 paired launches (100%, threshold 80%) So the width-16 gap is not a new kernel deficiency at that width: it is the t=8 gap plus a scaling failure that is ours alone. That run also reproduces the width-16 gap at 1.793x from a different harness and token budget -- a fourth independent estimate. This rules out the host bandwidth knee documented in `2026-08-22-decode-width-scaling.md`. A DRAM ceiling is a property of the host and would flatten both arms; ORT scales 1.76x across the same doubling on the same host in the same launch. Two things recorded rather than quietly fixed: - Run 1's guard refused 8 of 14 cells and none of the refusals was contention -- `competing_load()` was empty for all 14 and the pre-check runnable count was 2-4 throughout. The ceiling was set from a structural estimate (22) and the cell's own peak reaches 25. Run 2's ceiling of 26 is fitted to run 1's peaks, which is post-hoc with respect to that run, and it is disclosed as such; run 1's REPORT NOTHING stands as issued. - The scaling run's t=8 gap reads 1.264x against the matrix's 1.120x. The difference is on the native side (190 vs 211 tok/s; ORT is unchanged) and is most likely this harness running a t=16 arm beside the t=8 arm in the same launch. It is published, and explicitly not offered as a correction to 1.120x. Both harnesses gate their pre-check on the instantaneous runnable count rather than `getloadavg()`. This study is why: immediately after its own cargo build the load average sat at 6.86 on a host whose runnable count was 2, and a threshold=3.0 pre-check slept through a fully quiet window without measuring anything. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
enabled auto-merge (squash)
August 23, 2026 18:46
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1871 +/- ##
=======================================
Coverage 80.92% 80.92%
=======================================
Files 415 415
Lines 203912 203912
Branches 203912 203912
=======================================
+ Hits 165021 165024 +3
+ Misses 33328 33325 -3
Partials 5563 5563
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
Picks up the catalog-pin repair (#1870/#1872) that had every PR on the repo red: #1860 registered pkg.nxrt::KvCacheCapacityAppend without bumping expanded_registry_catalog_count_is_pinned. Nothing in this branch touches the shape-inference registry -- the failure was inherited from main, reproduced locally on a clean checkout, and is not fixed here. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 23, 2026
…te the idle (#1887) ## What this answers [#1871](#1871) established that the acc0 gap is ~1.78x at width 16 versus 1.12x at width 8, and left the mechanism open: **when the `t=8 → t=16` doubling returns 1.32x instead of 2x, are the extra workers idle, or busy and inefficient?** Wall-clock timing cannot separate those. This adds CPU-seconds attribution to both arms and answers it. ## The instrument `process_cpu_time()` in `benches/common` reads `/proc/self/stat` (thread-group user/sys ticks); `ort_matmulnbits_baseline.py` brackets `getrusage(RUSAGE_SELF)` over the same window and emits identically-named fields, so one parser reads both arms. Read directly rather than through `/usr/bin/time`, whose `Percent of CPU` is `(user+sys)/wall` — it looks like independent corroboration of a wall-time result and is actually the same measurement divided by itself. The decomposition is an **identity**, not a model: ``` tps(16) / tps(8) == 2 * R_busy / R_cpu ``` so residual is 0.00% by construction and it doubles as a free per-cell self-test. Cells failing it are discarded as instrument faults rather than reported. **It earned its place before it was used in anger.** The harness's first run returned `REPORT NOTHING (n_trusted = 1 < 6)`, which was honoured — nothing was quoted, including the one trusted cell. Diagnosing it found a defect in *my instrument*, not the host: each quantity was reduced by its own independent median, and since `tps = tokens/wall` the two sort in reversed orders, so at an even rep count they selected **different repetitions**. Identity errors of **4.4%–29.7%** against a quantity that is algebraically zero. Both producers now emit every CPU field from the single median-throughput repetition. **No threshold in the pre-registered rule was changed** — only the instrument feeding it, and the change was forced by a self-test firing before anything was scored. ## Finding 1 — the CPU inflation is ours, not the host's 13 trusted of 14 launches, identity error 0.00% on every cell: | arm | speedup 8→16 | `R_cpu` | `R_busy` | busy@8 | busy@16 | sys_frac@8 | @16 | |---|---:|---:|---:|---:|---:|---:|---:| | native | 1.445 | **1.449** | 1.057 | 0.900 | 0.966 | 0.062 | **0.212** | | ORT | 1.860 | **1.074** | 0.999 | 1.000 | 0.999 | 0.000 | 0.000 | ORT ran interleaved, same launches, same 16 cores, same minutes. If the knee were a memory-bandwidth ceiling, ORT would hit it too. **This kills the DRAM-plateau attribution.** ## Finding 2 — `busy` is blind to spin-wait, and the shipped wait path spins `decode_spmd`'s worker wait spins then `sched_yield`s for up to 500 µs (`ONNX_GENAI_CPU_DECODE_BLOCKTIME_US`) before parking. A yielding thread accrues CPU time exactly like a working one. Dose-response first, because an env var that never reaches the child produces a beautifully consistent null — w=16, 192 tokens: `sys_s` = **0.13 / 3.28 / 3.12** at 0 / 500 / 20000 µs while `user_s` = **9.28 / 9.27 / 8.99**. Knob live, ramp saturated by 500 µs, and `user_s` invariant — so the system time is pure overhead, not work. Pre-registered A/B, 7 trusted of 10: | w | ratio (bt0 ÷ bt500) | A/A null | busy@500 | busy@0 | cpu_s/tok @500 | @0 | |---:|---:|---:|---:|---:|---:|---:| | 16 | **0.9960** | 0.9937 | 0.953 | **0.692** | 0.06037 | 0.04838 | | 8 | 0.9906 | 1.0186 | 0.957 | 0.916 | 0.04078 | 0.03794 | **Throughput: REJECT** (43% sign consistency against a 5.24% A/A half-width). Removing the ramp does not make this workload faster, and the favourable mechanism data does not change that. No regression at w=8. ## What finding 2 does to finding 1 BURN-DOMINATED keys on `R_busy ≥ 0.90` — "the workers are not idle". That quantity was masked. Same harness, same unchanged rule, `--blocktime 0`, 9 trusted of 12, identity 0.00%: | arm | `R_cpu` | `R_busy` | busy@8 | busy@16 | |---|---:|---:|---:|---:| | native | **1.304** | **0.652** | 0.938 | **0.595** | | ORT | 1.109 | 0.999 | 1.000 | 0.999 | **VERDICT: MIXED**, 100% sign consistency on *both* ratios. | configuration | busy@16 | reads as | |---|---:|---| | default (500 µs ramp) | 0.966 | pool nearly fully occupied → BURN | | ramp off | **0.595** | **40% of the pool is not working** | **The correct attribution is mixed: ~30% more real CPU per token at width 16 *and* ~40% of the sixteen cores idle.** Both halves are real, both are ours. A useful side-effect: over the same 9 cells the wall-derived speedup swings **0.78x–1.37x** while `R_cpu` holds 1.10–1.34 and `R_busy` 0.52–0.81, both 100% sign-consistent. A competing process steals our wall clock but does not add to our CPU seconds — which is the whole reason for the instrument. ## Scope limit, stated up front This is a **zero-gap** decode loop — exactly the workload where parking early looks free. **Nothing here licenses changing the shipped blocktime default**, which exists to protect latency when there *are* gaps; that needs the gap-aware harness (#1395). What it does establish is that **any occupancy reading of the decode pool taken at the default blocktime over-reads by tens of points**, which is now recorded in the ledger. This independently corroborates, on a second workload, the ~20%-of-process-CPU figure @sebastian reported for the `worker_wait` yield ramp — the cost side is confirmed; the latency side remains his. ## Next 1. Localise the **idle** half — `ONNX_GENAI_CPU_DECODE_WORKER_PROFILE` / #1859's per-worker straggler attribution, to separate load imbalance from dispatch/wake latency. 2. Localise the **burn** half — per-op attribution at w=8 vs w=16 against the 136.3 MB/token figure. ## Validation - `cargo build --release -p onnx-runtime-ep-cpu --bench int4_decode_loop_ab` ✅ - `cargo fmt --check` ✅, `cargo clippy` clean on the touched bench ✅ - `ruff check` clean on all touched `.py` (the one repo-wide F401 is pre-existing in `ort_baseline.py`, untouched here) - Harness branch-validators run before use: 13 checks / 0 failures (cpu-split), 0 failures (blocktime A/B) - All measurement runs taken under `scripts/hostlock.sh` with announce-before/after Both harnesses accept `--replay <json>` to re-score archived data, and both carry their acceptance rule in the module docstring where it was written before the first run. --- # Update: the idle half is now attributed Follow-up 1 from the "Next" section above is **done in this PR** (commit `53d3ff076`). `int4_decode_loop_ab` now brackets `SpmdWorkerProfile` deltas over exactly the window that produces `wall`, and `acc0_w16_worker_split.py` scores them against a rule written before its first run. **10 launches, 10 trusted, both fired conditions at 100% sign consistency.** **Wake latency is not the problem.** `wake_frac` is 0.006 at w=8 and **0.051** at w=16; the pre-registered WAKE-BOUND condition did **not** fire. A perfect wait/wake path recovers at most 5 points. The mean worker's window at width 16: | | w=8 | w=16 | |---|---:|---:| | useful work | 0.886 | **0.492** | | straggler wait | 0.031 | **0.222** | | wake latency | 0.006 | 0.051 | | dispatcher / serial | 0.077 | 0.235 | **Not all of that is a defect.** Holding w=8's serial time constant and halving its parallel time — Amdahl with no defect anywhere — predicts a 0.204 serial share at w=16; observed is 0.235. So: | component | points | recoverable? | |---|---:|---| | straggler wait | 22.2 | **yes** | | serial in excess of Amdahl | 3.1 | maybe | | wake latency | 5.1 | partly | | Amdahl-predicted serial | 20.4 | **no** | Quoting the 46% residual as recoverable would be wrong by more than 2x. The same calibration also **rules out pure Amdahl** as the explanation for the knee: it predicts 0.796 useful work and the pool delivers 0.492. **The straggler** holds **72%** of last-arrivals against a 6.7% chance share, does **1.5x** the median work, and almost never parks. Two candidate mechanisms were tested and are **negative**: not a static mis-partition (segments divide evenly for every shape here, and the straggler's identity moves between launches), and not the unpinned dispatcher colliding with a worker (dispatcher CPU → straggler CPU over four profiled launches: `30→22`, `6→24`, `6→20`, `{30,20,18}→18`). **No mechanism is claimed.** **Nothing can help it today.** `DEFAULT_STEAL_TILES_PER_WORKER = 1` makes `target == total_workers`, so `work_stealing_segments_aligned` always falls back to static equal segments — there are no spare tiles to steal. ## A +23% candidate, rejected The blocktime A/B harness is generalised to vary any single env knob (thresholds, A/A arm, rotation and the w=8 guard untouched; replaying the archived blocktime JSON through the refactored scorer reproduces the published numbers exactly). Pointed at that default, steal tiles 1 vs 2, blocktime held at the shipped 500 µs in both arms, 8 trusted: | w | ratio | sign | A/A | sys_frac | cpu_s/tok | |---:|---:|---:|---:|---|---| | 16 | **1.2327** | 88% | 1.1301 | 0.280 → 0.192 | −11.7% | | 8 | 1.0463 | — | 1.0211 | 0.038 → 0.030 | −2.5% | **REJECTED and not proposed.** The rule requires the effect to clear 3x the A/A half-width, and that half-width is **0.2154** in the same run. Re-running until the null comes in narrow would be choosing the sample that licenses the conclusion, so it is not done. The mechanism moving in the predicted direction, and the absence of the expected w=8 regression, are recorded as observations rather than as a result. **This makes the width-16 A/A null the binding constraint** — ±21.5% here, ±35% in the earlier width-16 study — large enough that no improvement of realistic size can clear a pre-registered bar at this width. It moves ahead of further kernel work in the ledger. New record: `docs/benchmarks/2026-08-23-acc0-width-16-worker-attribution.md`. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The
t=16condition I merged this morning has failed4b4dacc7e(#1852) re-measured the acc0 (accuracy_level = 0, productiondefault) int4 decode gap, found the published 1.84x stale, measured
1.120x at t=1 and t=8, and concluded acc0 was no longer the top CPU
MatMulNBits target — explicitly conditional on
t=16, which that matrixcould not resolve and which reviewers rightly pushed me to stop burying.
This resolves it. The gap at
t=16is ~1.78x and acc0 goes back to the top ofthe work list.
The pre-registered rule said RANGE ONLY and that verdict is honoured
The acceptance rule is written into the harness before its first run:
Run 1 returned REPORT NOTHING (n=6). Run 2 returned RANGE ONLY
(A/A half-width 0.377). Neither run is re-scored under a looser rule and no
point estimate is claimed. The range is genuinely wide — 1.28x to 2.18x — but
the lowest single cell of 30 still exceeds the
t=8figure of 1.120x.Four reasons the central tendency survives the wide null:
implementation. The study alternates which arm is doubled: native A/A
half-width 0.347 (n=15), ORT's 0.377 (n=15) — ORT's is the wider. The
original matrix only ever doubled native and could not tell these apart.
gives 1.770x on top of the paired median's 1.775x.
half 1.834x) but the gap does not vanish when native runs well.
Mechanism: the scaling wall is ours, and it is not bandwidth
A second harness runs both widths inside the same launch, width order
rotated. Its rule is also pre-registered and is deliberately a sign test on
paired launches, not a comparison of medians — a ±35% null can move a
median-of-ratios but cannot flip a per-launch sign symmetrically.
t=8 → t=16scalingNative scaled worse in 10 of 10 paired launches (100%, threshold 80%). ORT
converts the doubling into 1.76x, near the 2.0x ideal; we get 1.32x and lose a
third of the added width.
This rules out the documented bandwidth knee
(
2026-08-22-decode-width-scaling.md, 7.52x → 9.26x at t=16). A DRAM ceiling isa property of the host and would flatten both arms. ORT scales 1.76x across
the same doubling, on the same host, in the same launch. Whatever binds us at
t=16is in our pool or our kernel.Two things recorded rather than quietly fixed
and none of the refusals was contention —
competing_load()was empty forall 14 and the pre-check runnable count was 2–4 throughout. The ceiling came
from a structural estimate (22); the cell's own peak reaches 25. Run 2's
ceiling of 26 is fitted to run 1's peaks, which is post-hoc with respect to
that run, and is disclosed as such rather than presented as a fresh
derivation. Run 1's
REPORT NOTHINGis not re-scored.t=8gap reads 1.264x, not 1.120x. The difference is onthe native side (190 vs 211 tok/s; ORT unchanged at 232–240 vs 238), most
likely because that harness runs a
t=16arm beside thet=8arm in onelaunch. Published, and explicitly not offered as a correction to 1.120x.
It does not touch the scaling result, which is paired within each launch.
Also fixed: the pre-check instrument
Both harnesses gate on the instantaneous runnable count, not
getloadavg().This study is the demonstration: right after its own
cargo buildthe loadaverage sat at 6.86 on a host whose runnable count was 2, and a
threshold=3.0pre-check slept through a fully quiet window without measuringanything. Gating the pre-check on the same quantity
LoadWatchpolices duringthe arm makes the two agree.
Files
docs/benchmarks/2026-08-23-acc0-gap-at-width-16.md— the recordcrates/onnx-runtime-ep-cpu/benches/acc0_w16_study.py— the gap studycrates/onnx-runtime-ep-cpu/benches/acc0_w8_w16_scaling.py— the scaling test2026-08-23-acc0-gap-vs-ort-by-width.md,CPU_MATMUL_ASSIGNMENT.md, and twoneighbouring benchmark docs updated where they carried the now-failed
condition
Validation
Docs plus two benchmark harnesses; no library code. Both harnesses were
stub-validated across every verdict branch before touching the host, and both
were run for the numbers above on
0f84888b8. Required CI (Fast (Linux x86_64),Rust quality) must be green before merge — no admin bypass.