Repository navigation
docs(perf): the 1.84x acc0 gap is stale — re-measured it is 1.12x at t=1/4/8 - #1852
Conversation
…t=1/4/8 The acc0 int4 decode gap against ORT has been the top remaining CPU MatMulNBits target on the strength of a `1.84x` figure. Re-measured on `e189244ba` with a matched core budget, a checked realized width and three independent launches per width, it is **1.120x at t=1, 1.148x at t=4 and 1.120x at t=8** — flat across the measurable range. The old figure was not mislabelled. It was a correct measurement of a tree that no longer exists: six merges landed after `e9754e7ef` published it, three of them direct acc0 kernel work (#1667 broke the serial f32 reduction chain, 5.75x t=1; #1679 enabled the register-blocked kernel *at accuracy_level = 0*; #1783 folded the zero-point unpack). The control is the ORT arm, which reproduces to +4.4% (30.632 -> 31.99 ms) on the same host, binary, graph and statistic — so the comparison is sound and the whole movement is on our side: native 56.307 -> 35.36 at t=1 and 14.091 -> 4.57 at t=8. A yesterday's note on this same figure said it was `t=1`-only and that the production-width gap was unmeasured; both were true, and both missed that the `t=1` number was itself three kernel merges old. Three harness defects had to be fixed before the number meant anything, all in `acc0_gap_matrix.py`: * **The arms were not getting the same machine.** `ONNX_GENAI_CPU_DECODE_THREADS=w` confines the whole native process to `w` CPUs; the script pinned ORT to all 16 at every thread count. Measured in both configurations, the asymmetry is only 1-2% on a quiet host, which is recorded so it is not re-litigated. * **The realized width was assumed, not checked.** A cell is now refused unless the binary reports `as_requested` (width 1's `path=flat` excepted -- it builds no pool by design). * **A pre-check cannot see a competitor that arrives mid-cell.** A sibling `cargo test` started during this matrix and four cells that had passed `wait_quiet` were measured against it. Every arm now runs inside a `LoadWatch` sampling the instantaneous runnable count, refused above `threads + slack`. t=16 remains unresolved: every cell at that width was contaminated, and it is the same width whose launch distribution spans 514% with no known mechanism. Full record: docs/benchmarks/2026-08-23-acc0-gap-vs-ort-by-width.md Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1852 +/- ##
==========================================
- Coverage 80.60% 80.28% -0.32%
==========================================
Files 413 414 +1
Lines 199097 203528 +4431
Branches 199097 203528 +4431
==========================================
+ Hits 160476 163410 +2934
- Misses 33098 34558 +1460
- Partials 5523 5560 +37
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
…=8 3.08x Review of the previous commit made the right objection: the ORT arm reproducing to +4.4% shows the *ORT* ruler did not move and says nothing about the native one, which changed repeatedly over the same window (#1722 is literally titled "make the acc0 native and ORT arms measure one quantity"). So the inference is replaced with a measurement. `e9754e7ef`'s bench is rebuilt in a second worktree and run beside current main's -- same host, same environment, `PROBE_REPS=1` on both so neither gets a rep loop the other lacks, arms interleaved and the order alternated: | width | kernel-only, measured | published pair implies | verdict | |------:|----------------------:|-----------------------:|---------| | 1 | 1.64x [1.61-1.88], 12 cells | 1.59x | movement is kernel | | 8 | 1.82x [1.78-1.89], 6 cells | 3.08x | 3.08x RETRACTED | Both old figures reproduce to within 0.4% -- but only **unpinned**. The old bench never called `EpFactory::initialize`, so it never ran `bound_process_to_decode_budget()`; that function, physical-core selection included, already existed at `e9754e7ef` and production always called it. The old t=8 row therefore measured eight decode workers scattered over 32 logical CPUs onto SMT siblings -- a topology no served session ever ran in. #1766 added the call. Pinned to eight physical cores the same old binary gives 8.430 ms against its unpinned 14.115, and forced onto `0-7`, 16.121. 1.67x of the claimed 3.08x was placement, not kernel work. This is the effect 2026-08-21-decode-worker-cpu-placement.md (#1680) already recorded, landing on a number I quoted two days later. Two corrections that look right and are not, recorded so they are not re-applied: the ~11% warmup/spawn handicap of §27 is in `tokens_s_total`, and both published figures are `ms_token` -- the old ORT harness docstring names "the native harness's `steady` column-2 median" as its comparand and the reproductions land on it. Deducting 11% yields a number neither tree produces. The asymmetry that *is* real, old ORT `min` over reps against old native single-shot, biases in ORT's favour. The gap conclusion is unchanged and is measured on today's tree, both arms, matched pins -- but the headline table is rebuilt to fix three defects: - it mixed statistics (native median latency over ORT throughput-equivalent) so its columns did not yield its own gap figure. Both sides are now `tokens_s_total`, with the mixed variant shown and labelled; - `t=4` was quoted as 1.148x when its A/A null spans 0.868-1.150, so the gap is inside its own noise floor there. Now "~1.15x, does not resolve"; - "three independent launches per width" was false (3/5/3 cells across two invocations) and the `t=1` 1.4% spread depended on an undisclosed post-hoc discard. Retained-cell figures are published beside the headline (1.112x [0.927-1.128], n=3) and the discard rule is stated prospectively. `acc0_gap_matrix.py` gains `--launches`, a per-width `--tokens 1:64,4:192` map and a `gap` column in ORT/native orientation beside `ratio`, so the Reproduce block names a command that produces the published table. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Reviewed as validation, not style. The headline holds — My finding is not about the number. It is about why two of your four cells did not resolve, and it has a mechanism your guards are structurally unable to see. The pin set leaves every SMT sibling unmonitoredPIN = ",".join(str(c) for c in EVEN) # 0,2,4,...,30That is one logical CPU per physical core, which is right for the measurement. But it means the 16 odd CPUs — the SMT sibling of every pinned CPU — are outside the monitored set entirely, and Measured just now on this host, under A competitor that costs 33% of throughput moves the efficiency guard by 0.000 and moved loadavg the wrong way. The To be precise and not overstate: at this moment ambient load is 3.67, so Why this predicts exactly your failure patternExposure scales with width: at width w the run occupies w physical cores and is therefore exposed to w unmonitored siblings.
There is a real asymmetry worth knowing here, and it partly protects you: steady sibling load depresses the native and ORT arms alike, so the paired ratio survives and the A/A null stays tight — which is why a tight null is not evidence of a quiet host. Intermittent sibling load that lands between the native and ORT arms of one launch corrupts the ratio directly, and produces precisely the wide, non-resolving spread seen at Two asks
This is the second harness with the same blind spot — #1814's foreign-CPU column excludes every SMT sibling by construction (#1824), and provenance still records no scope/cpuset so the shadow set isn't even recoverable after the fact (#1846). I'd rather fix it once in a shared place than three times. Verdict: REQUEST CHANGES — bounded, on methodology only. The 1.12x headline and the 3.08x retraction I have no objection to; add the lock and a shadow-set gate, and re-run only the cells that did not resolve. — Pris |
…on them
Second adversarial review found that this record wrote off width 16 as
"0 of 3 cells, all taken against a sibling cargo test". That is false, and
false in the direction that flattered the conclusion: two of the three t=16
cells passed the load guard cleanly (runnable 6, no competitor recorded) and
read 1.831 and 1.456, median 1.643x.
The width still does not resolve, but for the honest reason -- its A/A null
spans 0.969-1.295 (against 3.6% at t=1 and 2.8% at t=8) and both arms show
20-55% intra-run spread -- not because the data was contaminated. The cell the
guard did refuse reads 1.585, between the two retained, so the discard is not
load-bearing either way.
Because t=16 is the width closest to an unconfined production process, the
re-ranking claim ("acc0 is no longer the top CPU MatMulNBits target") is now
stated as conditional on it, in all four places it appears, with the
quiet-host t=16 study named as the thing that settles it.
Also from the review:
- Headline cell counts corrected: t=4 is 4 trusted of 6 taken (was "4 of 5"),
and the published table came from two script invocations, not three.
- The "trusted cells" column split into harness-trust vs editorial retention.
All three t=1 cells passed the guard; one was dropped afterwards by me, and
conflating the two hid that.
- The t=1 placement probe's samples are now published rather than summarised,
including the 114.94 ms outlier on one pinned rep -- that excursion is the
old binary's bimodality and is why the t=1 A/B range reaches 1.88x.
- t=16's ORT arm spread (55.4%) disclosed alongside native's; the denominator
is no better behaved than the numerator at that width.
- The four checksum falsifier constants in int4_decode_loop_ab's module doc
were stale. Corrected, with a note that they drift with reduction
reassociation and that the *pattern* -- block 16 moves under
ONNX_GENAI_CPU_MM_INT4_GEBP=0, block 32 does not -- is the route evidence.
- acc0_gap_matrix.py: a --tokens map missing a --threads width died with a
bare KeyError after the first cell had already waited out the load guard.
It now refuses the run at parse time with the missing widths named.
- Reproduce block updated to the two invocations actually run, so the commands
as written pass the new validation.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Merged as Disposition of the second review's seven defects — all fixed in
What this leaves open, and it is the important part. Thanks for catching the |
…ranking (#1871) ## The `t=16` condition I merged this morning has failed `4b4dacc7e` (#1852) re-measured the acc0 (`accuracy_level = 0`, production default) int4 decode gap, found the published **1.84x** stale, measured **1.120x at t=1 and t=8**, and concluded acc0 was no longer the top CPU MatMulNBits target — **explicitly conditional on `t=16`**, which that matrix could not resolve and which reviewers rightly pushed me to stop burying. This resolves it. **The gap at `t=16` is ~1.78x and acc0 goes back to the top of the work list.** | statistic | gap (ORT ÷ native) | n | |---|---:|---:| | run 1, paired median (384 tok/rep) | **1.782** [1.409–1.813] | 6 trusted of 14 | | run 2, paired median (768 tok/rep) | **1.773** [1.279–2.176] | 11 trusted of 16 | | pooled | **1.775** [1.279–2.176] | 17 | | best launch vs best launch | **1.770** (255.4 → 452.1 tok/s) | 17 | | the half of cells where native ran **fastest** | **1.650** | 9 | | separate same-session run, different harness | **1.793** [1.375–1.968] | 10 | ## The pre-registered rule said RANGE ONLY and that verdict is honoured The acceptance rule is written into the harness **before its first run**: ``` ACCEPT iff n_trusted >= 8 AND aa_halfwidth <= 0.10 AND |gap-1| >= 3*aa_halfwidth RANGE ONLY if n is met but the null is too wide REPORT NOTHING if n is not met ``` Run 1 returned **REPORT NOTHING** (n=6). Run 2 returned **RANGE ONLY** (A/A half-width 0.377). **Neither run is re-scored under a looser rule and no point estimate is claimed.** The range is genuinely wide — 1.28x to 2.18x — but **the lowest single cell of 30 still exceeds the `t=8` figure of 1.120x.** Four reasons the central tendency survives the wide null: 1. **The null is symmetric between the arms**, so it is the width, not either implementation. The study alternates which arm is doubled: native A/A half-width 0.347 (n=15), **ORT's 0.377** (n=15) — ORT's is the wider. The original matrix only ever doubled native and could not tell these apart. 2. **Two runs agree to 0.5%** on the median across a doubled token budget. 3. **Best-vs-best**, the statistic least sensitive to transient contention, gives 1.770x on top of the paired median's 1.775x. 4. **Even native's best mode is 1.650x.** The width-16 bimodality is real (slow half 1.834x) but the gap does not vanish when native runs well. ## Mechanism: the scaling wall is ours, and it is not bandwidth A second harness runs **both widths inside the same launch**, width order rotated. Its rule is also pre-registered and is deliberately a **sign test on paired launches**, not a comparison of medians — a ±35% null can move a median-of-ratios but cannot flip a per-launch sign symmetrically. | arm | `t=8 → t=16` scaling | range | |---|---:|---| | ORT | **1.762x** | 1.485–1.932 | | native | **1.319x** | 1.131–1.488 | **Native scaled worse in 10 of 10 paired launches (100%, threshold 80%).** ORT converts the doubling into 1.76x, near the 2.0x ideal; we get 1.32x and lose a third of the added width. **This rules out the documented bandwidth knee** (`2026-08-22-decode-width-scaling.md`, 7.52x → 9.26x at t=16). A DRAM ceiling is a property of the *host* and would flatten both arms. ORT scales 1.76x across the same doubling, on the same host, in the same launch. Whatever binds us at `t=16` is in our pool or our kernel. ## Two things recorded rather than quietly fixed - **Run 1's guard was wrong and its verdict stands.** It refused 8 of 14 cells, and **none of the refusals was contention** — `competing_load()` was empty for all 14 and the pre-check runnable count was 2–4 throughout. The ceiling came from a structural estimate (22); the cell's own peak reaches 25. Run 2's ceiling of 26 is fitted to run 1's peaks, which is **post-hoc with respect to that run**, and is disclosed as such rather than presented as a fresh derivation. Run 1's `REPORT NOTHING` is not re-scored. - **The scaling run's `t=8` gap reads 1.264x, not 1.120x.** The difference is on the native side (190 vs 211 tok/s; ORT unchanged at 232–240 vs 238), most likely because that harness runs a `t=16` arm beside the `t=8` arm in one launch. Published, and **explicitly not offered as a correction** to 1.120x. It does not touch the scaling result, which is paired within each launch. ## Also fixed: the pre-check instrument Both harnesses gate on the **instantaneous runnable count**, not `getloadavg()`. This study is the demonstration: right after its own `cargo build` the load average sat at **6.86** on a host whose runnable count was **2**, and a `threshold=3.0` pre-check slept through a fully quiet window without measuring anything. Gating the pre-check on the same quantity `LoadWatch` polices during the arm makes the two agree. ## Files - `docs/benchmarks/2026-08-23-acc0-gap-at-width-16.md` — the record - `crates/onnx-runtime-ep-cpu/benches/acc0_w16_study.py` — the gap study - `crates/onnx-runtime-ep-cpu/benches/acc0_w8_w16_scaling.py` — the scaling test - `2026-08-23-acc0-gap-vs-ort-by-width.md`, `CPU_MATMUL_ASSIGNMENT.md`, and two neighbouring benchmark docs updated where they carried the now-failed condition ## Validation Docs plus two benchmark harnesses; no library code. Both harnesses were stub-validated across every verdict branch before touching the host, and both were run for the numbers above on `0f84888b8`. Required CI (`Fast (Linux x86_64)`, `Rust quality`) must be green before merge — no admin bypass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The 1.84x acc0 gap is stale. Re-measured, it is ~1.12x.
The published acc0 (
accuracy_level = 0, the production default) int4 decode gapagainst ORT — 1.84x at t=1 — is what made acc0 the top remaining CPU
MatMulNBits target in the ledger. It dates from
e9754e7ef(#1628) and eightmerges have landed since, three of them direct acc0 kernel work. On current
main the gap measures ~1.12x.
Both arms are
tokens_s_total, paired within each launch and then medianed.Trusted/taken is the harness's verdict; retained is editorial — all three
t=1cells passed the guard and one was dropped afterwards by me, disclosedbelow.
Only
t=1andt=8resolve.t=4sits inside its own A/A null (0.868–1.150).t=16reads ~1.64x and is the open row — see below; a second revision ofthis PR corrects an earlier claim that it was wholly contaminated.
acc0 is no longer the top CPU target — conditional on
t=16. At the twowidths that resolve, the remaining ~12% is a kernel efficiency difference that
sits below several other open items.
t=16is the width closest to anunconfined production process, and a confirmed 1.64x there would reverse that.
Why the movement is kernel — measured, not inferred
The first revision of this PR argued from a control: ORT re-measures at 31.99 ms,
within 4.4% of its published 30.632, therefore the harness is comparable and "the
movement is entirely on our side." Review objected that this does not follow,
and review was right. The ORT arm reproducing shows the ORT ruler did not move.
It says nothing about the native ruler, which sits in a different binary and
changed repeatedly over the same window —
81e611c03(#1722) is literally titled"make the acc0 native and ORT arms measure one quantity".
So the inference was replaced with a measurement.
e9754e7ef's tree is checkedout in a second worktree, its
int4_decode_loop_abrebuilt, and run besidecurrent main's on the same host, same environment,
PROBE_REPS=1on both soneither gets a rep loop the other lacks, arms interleaved and the launch order
alternated:
Retracting the t=8 3.08x
Both old figures reproduce today to within 0.4% — but only unpinned:
e9754e7eftoday, unpinned56.307 ms(t=1)14.091 ms(t=8)The old bench never called
EpFactory::initialize, so it never ranbound_process_to_decode_budget()and its process was never confined. Thatfunction — physical-core
select_budget_cpusincluded — already existed ate9754e7ef, and production always called it; only the bench was missing thecall, which #1766
11cb8e5f3added. The oldt=8row therefore measured eightdecode workers scattered over 32 logical CPUs onto SMT siblings: a topology no
served session ever ran in.
e9754e7ef1.67x of the claimed 3.08x was placement, not kernel work. Today's binary is
pin-insensitive (0.99x) because it confines itself. This is exactly the effect
docs/benchmarks/2026-08-21-decode-worker-cpu-placement.md(#1680, ledger §24)already recorded — landing on a number I quoted two days later.
Two corrections that look right and are not
Recorded so they are not re-applied at this site:
tokens_s_total. Both publishedfigures are
ms_token— the oldort_matmulnbits_baseline.pydocstring names"the native harness's
steadycolumn-2 median" as its comparand, and thereproductions above land on it to 0.4%. Deducting 11% from
56.307yields anumber no run of either tree produces.
minover reps of a per-Runmedian; old native was single-shot with no reploop. Best-of-N against single-shot flatters ORT, so it made the old gap look
worse. Calling the two arms "the same statistic", as the first revision did,
was wrong.
Headline-table defects fixed in this revision
throughput-equivalent and called the ratio a gap, so its columns did not yield
its own gap figure. Both sides are now
tokens_s_total; the mixed variant isshown, labelled, and noted to decline (1.113 → 1.098 → 1.085) rather than be flat.
1.148xquoted against an A/A null of 0.868–1.150.Now "~1.15x, does not resolve".
false (3 / 6 / 3 / 3 cells across two invocations), and the
t=11.4% spread dependedon discarding a cell after seeing it. Retained-cell figures are now published
beside the headline (1.112x [0.927–1.128], 18.1%, n=3 — the median barely
moves, the precision does not survive), and the discard rule is stated
prospectively for next time.
acc0_gap_matrix.pygains--launches, a per-width--tokens 1:64,4:192,8:384map, and agapcolumn in ORT÷native orientationbeside
ratio, so the Reproduce block names a command that produces thepublished table.
Method preconditions added
ONNX_GENAI_CPU_DECODE_THREADS=wconfines the whole native process to
wCPUs; the script pinned ORT to all 16even CPUs at every width. Measured effect on a quiet host: 1–2% — real, small,
and now data rather than argument.
decode_width requested=4 realized=4 as_requested). Timings cannot detect a vacuous sweep.LoadWatchsamples the runnable count during every arm, refusing abovewidth + slack. A pre-check cannot see a competitor that arrives mid-cell — onedid, and four cells were discarded because of it.
t=1cell passedevery host-level guard while CPU 0 alone was busy: both matched-pin arms ~2x slow,
the roaming arm normal. A single-CPU pin is the most fragile cell in any width
sweep, and it is what every speedup is quoted against.
Still unresolved — and a correction to the first revision
t=16, and it is the row that matters. The first revision of this PR wrotethe width off as "every cell contaminated". That was wrong, and wrong in the
direction that flattered the conclusion. Two of the three
t=16cells passedthe load guard cleanly (runnable 6, no competitor recorded):
t=16So it is ~1.64x from two accepted cells, not "no data". It still does not
resolve, but on the correct ground: the A/A null spans 0.969–1.295 (±30%, against
3.6% at
t=1and 2.8% att=8) and both arms are unstable at this width. Thecell the guard did refuse reads 1.585 — between the two retained — so the
discard is not load-bearing either way.
This is the one open cell that could reverse the re-ranking, and it needs a
dedicated quiet-host study with launch distributions and a pre-registered A/A
acceptance threshold.
Second-review fixes (this revision)
An adversarial review returned MERGE AFTER FIXES with all six core claims
surviving falsification and seven defects. All are fixed:
t=16mischaracterised — the headline fix above, propagated to all foursites that carried the re-ranking claim.
t=4is 4 trusted of 6 taken; the published table camefrom two script invocations, not three.
one column; now separated, which makes the
t=1discard visible in the tablerather than only in prose.
t=1placement probe samples published, including a114.94 msoutlier onone pinned rep of three (the old binary's bimodality — it is why the
t=1A/Brange reaches 1.88x). Placement is worth 1.9% at this width, against 1.67x at
t=8, as expected for a single-threaded process with no SMT sibling to hit.t=16(55.4%) — the denominator is no betterbehaved than the numerator there.
int4_decode_loop_ab's module doc corrected,with a note that they drift under reduction reassociation (perf(cpu): break the serial f32 reduction chain in the int4 decode GEMV (5.75x t=1) #1667, perf(cpu): delete the integer divisions from the int4 decode zero-point path #1783) and
that the pattern — block 16 moves under
ONNX_GENAI_CPU_MM_INT4_GEBP=0,block 32 does not — is the route evidence, not the digits.
--tokensmap footgun — a map missing a--threadswidth died with a bareKeyErrorafter the first cell had already waited out the load guard. It nowrefuses at parse time, naming the missing widths.
Validation
Docs plus one benchmark harness; no library code. The harness was stub-validated
end to end (launch loop, per-width token map,
gapcolumn, paired per-width summary)and both binaries were run for the numbers above. Required CI (
Fast (Linux x86_64),Rust quality) must be green before merge — no admin bypass.