Skip to content

perf(cpu): enable the register-blocked int4 decode kernel at accuracy_level=0 (1.17-1.69x) - #1679

Merged
justinchuby merged 1 commit into
mainfrom
roy/acc0-route
Aug 21, 2026
Merged

justinchuby merged 1 commit into
mainfrom
roy/acc0-route

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

What this changes

One default flip: borrowed_int4_nblock_enabled() now returns true when unset. Plus the two tests that should have existed to make that flip unnecessary.

The finding

#1104 built borrowed_affine_int4_matmul_nblock — the register-blocked int4 decode kernel — measured 1.46x on a 14B model, proved byte-identical output, added parity tests, and shipped it default off, "until the win is measured, exactly like the toggles that preceded it."

The measurement never happened. The kernel has been sitting in the tree unused ever since, while accuracy_level = 0 — the production default — kept taking the per-column path it was written to replace.

Every parity test passed the whole time, because they all call the kernel directly. Nothing asserted which route production actually took.

Route proof

I did not infer this from reading guards. I instrumented counters from operator entry through the kernel and ran the decode loop:

PROBE_ACCURACY=0  entry_bits4=95 guard_taken=95 gebp=0 prefill_block=0
                  nblock=0  percolumn=95  block_simd=129499136 block_scalar=0

Every call reaches the borrowed guard; every one lands in borrowed_affine_int4_matmul; nblock is never entered. 100% of blocks take the SIMD path, so the per-column path was not falling back to scalar — it was simply the wrong kernel.

Attribution — and a hypothesis of mine that the data killed

Reading the two kernels, the obvious explanation is the per-block horizontal reduction. The per-column path pays a full 8-lane hreduce (extractf128/movehl/shuffle, each dependent on the last) every 32 weights — once per four FMAs — then folds the result into a serial f32 sum. The N-blocked kernel keeps the scale in a vector accumulator and reduces once per column. That looked like the whole story.

I made the column-group width tunable to separate that restructuring from the four-column activation reuse:

route (block 32, min of 6 interleaved) t=1 t=4 vs per-column
per-column (previous default) 56.528 28.518 1.00x
N-blocked, group of 1 55.156 27.861 1.02x
N-blocked, group of 2 47.679 23.885 1.19x
N-blocked, group of 4 38.114 19.238 1.48x

Removing the per-block horizontal reduction is worth 1.02x — nothing measurable. The block loop has enough independent work across blocks for the out-of-order engine to hide that dependency chain entirely. The entire win is the four-column activation reuse (1.45x from group 1 to group 4): at m == 1 each activation vector is loaded once and feeds four columns' FMAs.

A correction to my own first numbers

My first attribution run reported 3.14x for the group-of-1 arm and 4.80x overall. Those were artifacts of my own instrumentation: the route probe incremented an atomic once per 32-weight block, which fired 129M times in the per-column arm and zero times in the N-blocked arm, inflating the baseline ~3x. The table above is from a clean binary with only the group-width knob compiled in. I mention it because the wrong numbers were self-consistent and looked plausible.

Full matrix (accuracy_level = 0, min of 6 interleaved, pinned to physical cores)

threads block off ms/tok on ms/tok speedup
1 32 64.657 38.204 1.69x
1 64 36.275 28.284 1.28x
1 128 29.284 23.720 1.23x
4 32 28.562 19.240 1.48x
4 64 17.811 14.628 1.22x
4 128 14.764 11.994 1.23x
8 32 14.362 9.706 1.48x
8 64 8.975 7.489 1.20x
8 128 7.529 6.244 1.21x
16 32 7.269 4.938 1.47x
16 64 4.577 3.907 1.17x
16 128 3.812 3.239 1.18x

Concurrent sessions, aggregate tok/s (higher better):

threads sessions off on speedup
4 2 39.5 58.5 1.48x
4 4 46.6 75.5 1.62x
8 2 64.8 95.5 1.47x
8 4 79.7 117.6 1.48x

No regression anywhere in the matrix. The win is largest at block 32, which is what llama/qwen ship.

Numerics — the honest version

The existing parity test asserts N-blocked matches per-column within 1e-3. That is not a strong enough contract for a default: it would let a less accurate kernel ship as long as it stayed near the one it replaced. So the new test measures both against an f64 reference.

I expected the replacement to be at least as accurate. It is not. The two differ structurally: per-column forms (dot - actsum*zp) * scale per block and accumulates in one serial f32 chain; N-blocked accumulates sum(dot*scale) and sum(actsum*zp*scale) separately and subtracts once at the end. With unsigned nibbles 0..=15 and midpoint 8 those two running sums are comparable in magnitude and the answer is their difference, so the separated form gives cancellation somewhere to hide.

case per-column N-blocked
k=4096 bs=32 asym 6.739e-6 2.422e-5
k=256 bs=32 asym 1.676e-6 6.197e-6
k=512 bs=32 sym gain=1024 2.899e-5 3.983e-5
k=1024 bs=128 sym gain=256 1.186e-6 1.186e-6

Up to 3.70x worse in relative terms — a real cost of the 1.48x, not something I want buried. It stays at 2.4e-5 worst case, two orders inside the 1e-3 this route's parity tests already accept, and the accuracy_level = 0 contract itself is untouched: both paths are pure f32 FMA over f32 activations and neither is ever diverted through an int8/int16 kernel. The test includes deliberately hostile cases (activation gain up to 1024x, weights pinned near the midpoint) chosen to maximise exactly this cancellation.

Tests

  • acc0_decode_reaches_the_nblocked_kernel_by_default — asserts the route, which is the check whose absence let this sit dormant. Mutation-verified: reverting the default to false fails it. Takes DISPATCH_PROBE_LOCK per the repo's meta-test. Excluded under feature = "mlas", where mlas_sqnbit_owns_fp32_compute legitimately claims acc0 before the borrowed guard; the shipped default artifact links no MLAS, which is the configuration asserted.
  • nblock_holds_the_f32_contract_against_f64 — the differential test above. The 8x ratio guard is a regression bound sized off the measured 3.70x, widened for AVX-512 hosts where the per-column path uses the 512-bit block dot and so has its own error profile.

Validation

roy_validate.sh PASS=21 FAIL=0 SKIP=0, two consecutive full runs. cargo test -p onnx-runtime-ep-cpu 1614 passed, three consecutive runs. --features mlas 160 matmul_nbits tests pass.

Three real failures caught and fixed on the first validation pass, all mine: fmt; the mlas-feature route conflict above; and aarch64 cross-compile, where the new counter was dead code because the N-blocked kernel is x86-only and the cross lane builds tests with -D warnings.

One earlier full run reported PASS=20 FAIL=1 and its log was overwritten before I read it; two subsequent full runs and three consecutive -p onnx-runtime-ep-cpu runs were clean. Flagging rather than dismissing it.

Scope

The env toggle stays, inverted, so the per-column path remains reachable in the same binary for A/B. accuracy_level = 4, aarch64, MLAS and the prefill routes are untouched.

Refs #1676.

…_level=0

#1104 added `borrowed_affine_int4_matmul_nblock` behind a default-*off* env
toggle "until the win is measured". The measurement never happened, so the
production default kept taking the per-column path and the kernel sat unused.

Route counters instrumented from operator entry through the kernel confirm the
dormancy: at accuracy_level=0 every call reached the borrowed guard and every
one went to `borrowed_affine_int4_matmul` (nblock=0, percolumn=95). Every
parity test still passed, because they all call the kernel directly.

Measured 1.17x-1.69x on the decode A/B, best at block 32 (the common
llama/qwen block size), and 1.47x-1.62x aggregate throughput under 2-4
concurrent sessions.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby marked this pull request as ready for review August 21, 2026 19:15
@codecov

codecov Bot commented Aug 21, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.58065% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 81.02%. Comparing base (8aed77a) to head (f25a876).
⚠️ Report is 2 commits behind head on main.

Files with missing lines Patch % Lines
...es/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs 97.58% 3 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1679      +/-   ##
==========================================
+ Coverage   81.00%   81.02%   +0.01%     
==========================================
  Files         384      384              
  Lines      180771   180893     +122     
  Branches   180771   180893     +122     
==========================================
+ Hits       146429   146563     +134     
+ Misses      29396    29384      -12     
  Partials     4946     4946              
Flag Coverage Δ
cli-ort-linux 82.60% <ø> (ø)
cli-ort-windows 82.20% <ø> (+0.09%) ⬆️
mlas 85.05% <ø> (+0.03%) ⬆️
offline 80.89% <97.58%> (+0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...es/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs 79.69% <97.58%> (+0.33%) ⬆️

... and 1 file with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@justinchuby

Copy link
Copy Markdown
Owner Author

Windows ARM64 investigated and cleared before merge.

The run at 96882730088 failed with plugin_ort_e2e exiting 0xc0000005 STATUS_ACCESS_VIOLATION after initializer_chain_still_fuses_into_one_claim. That is a crash, not an assertion, so it needed attribution rather than a reflex rerun.

Attribution: this PR's only functional change is the default of ONNX_GENAI_CPU_MM_INT4_NBLK, and the sole consumer of that gate (matmul_nbits.rs:1566) sits under #[cfg(target_arch = "x86_64")]. On aarch64 the flag is read and never consulted, so the change is a runtime no-op on that lane. The rerun of the same job is now green (27m20s), confirming a flake rather than a regression.

Separately, Rust quality is red on main itself (cargo fmt on onnx-runtime-session, job 96899749480 on 0f40538b2); it passes here only because this branch predates #828. Fixed in #1686.

All 19 required checks green. Merging.

@justinchuby
justinchuby merged commit 99f105d into main Aug 21, 2026
30 of 31 checks passed
@justinchuby
justinchuby deleted the roy/acc0-route branch August 21, 2026 22:20
justinchuby added a commit that referenced this pull request Aug 22, 2026
…f16 layout divergence (#1701)

Three entries in `docs/performance/CPU_MATMUL_ASSIGNMENT.md` for work
that landed or was closed today. Documentation only — no code.

- **§23 — the acc0 route** (`99f105d52`, #1679). #1104's
register-blocked int4 kernel shipped default-off "until the win is
measured" and the measurement never happened, so `accuracy_level = 0`
took the per-column path for the kernel's whole life; route counters
read `nblock = 0`. Records the **corrected** attribution: the first one
was inflated ~3x because the probe's `fetch_add` sat in the timed path.
Clean, the hreduce removal is worth 1.02x — nothing — and the entire
1.48x is four-column activation reuse. The numerics regress up to 3.70x
relatively; disclosed rather than buried.
- **§24 — the t=8 "wash"** (#1680). The premise does not reproduce. The
win is flat through pool width 8 and collapses at 12+; root cause is
`decode_spmd.rs::node_shards` pinning worker *i* to `allowed_cpus()[i]`
in logical order, so 16 workers land on 8 physical cores. Bandwidth (83
GB/s available vs 41 GB/s drawn) and task grain were tried first and
discarded. **No kernel change** — handed to the runtime owner with the
measurement.
- **§25 — the f16/bf16 layout divergence** (`2e1cfb67c`, #1687, closing
#1381). Same math, same bytes, three prices; `[K,N]` crosses a page
every `p`. Software prefetch was tried first and is a recorded negative.
Accuracy moves the same way, so there was no trade to weigh. Also
records the memory-plan coupling that could have gone badly: the #1056
predictor was Apple-only for `MatMul`, so the transpose would have been
invisible to the plan on x86.

§23 and §25 contain the same error in two disguises — an instrument that
changed what it measured, and a numerics test built from
exactly-representable operands (`*0.125`) that could not see
reassociation at all and reported a confident zero. Both are written up
as such, next to §18's version of it.

All seven repo policy scripts pass (`check_publish_order`,
`check_profile_table`, `check_platform_naming`,
`check_dispatch_reachability`, `check_dispatch_manifest`,
`check_feature_gate_coverage`, `verify_documented_env_vars`).

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

The acc0-vs-ORT matrix that #1679 was missing

The §23 record I merged has only route-vs-route numbers (nblock on/off). It
never had the vs-ORT column, which is what the headline "acc0 is 1.84x behind
ORT" was actually about. Here it is — 24 cells, block 32, 24 tokens, 3 reps,
taskset-pinned to even CPUs (distinct physical cores on this SMT host), ratios
are aggregate tokens_s_total, >1.00x means we beat ORT.

model t sessions before/ORT after/ORT after/before
llama 1 1 0.440x 0.709x 1.613x
llama 8 1 0.389x 0.611x 1.568x
llama 8 2 0.747x 1.538x 2.058x
llama 8 4 1.182x 1.722x 1.457x
llama 16 1 0.672x 1.028x 1.529x
llama 16 2 0.895x 1.164x 1.301x
llama 16 4 1.331x 1.828x 1.374x
qwen 8 4 1.013x 1.591x 1.571x
qwen 16 1 0.275x 0.436x 1.584x
qwen 16 2 0.855x 1.231x 1.439x
qwen 16 4 1.638x 2.350x 1.435x

Reading it honestly

The route win is real, uniform, and larger than the ledger says.
after/before is 1.30x–2.06x in every single cell — no reversal anywhere
across two models, four thread counts and three session counts. §23 records
1.48x; the true range is wider and its floor is higher.

But "1.84x behind ORT" does not generalise, and should stop being quoted as a
constant.
It was a single-session, single-op statement. The gap is a
function of concurrency:

  • 1 session — still behind (0.44x–0.71x before, 0.61x–0.71x after).
  • 2 sessions — ahead at 1.16x–1.54x.
  • 4 sessions — ahead at 1.59x–2.35x.

Single-session decode remains the real deficit and is where the next work
belongs. The worst cell is qwen t=16 s=1 at 0.436x (2.3x behind).

One number I am not willing to stand behind yet: the qwen t=16 s=1 ORT
figure is 420.0 tok/s against 101.4 tok/s at s=2. A single session outrunning
two by 4x is not a shape I can explain, so that specific cell is flagged as
suspect and should be re-measured before anyone builds on it. The 0.436x ratio
is reported because suppressing an inconvenient outlier is worse than flagging
it, not because I trust it.

Why nblock=0 kept the kernel dormant

borrowed_int4_nblock_enabled() (matmul_nbits.rs) read
ONNX_GENAI_CPU_MM_INT4_NBLK and ended in .unwrap_or(false). #1104 shipped it
default-off "until the win is measured". The measurement never happened, and —
this is the actual defect — nothing anywhere asserted which route production
took
. Counters at accuracy_level = 0 read:

entry_bits4=95  percolumn=95  nblock=0  block_simd=129,499,136

nblock=0 after 95 entries. The kernel was reachable, tested, benchmarked in
isolation, and never executed by default. A toggle whose expiry condition lives
only in a comment has no expiry condition; the reviewer sees a plausible
sentence and the CI sees nothing at all.

What now selects it

Two changes, and the second is the one that matters:

  1. .unwrap_or(true) — the default flips.
  2. acc0_decode_reaches_the_nblocked_kernel_by_default — reads
    BORROWED_INT4_NBLOCK_TEST_CALLS before and after a real execute at
    accuracy_level = 0 and fails if the count does not move.

The invariant is now "the default route is a checked property" rather than
"the default route is what the comment says". Flipping the default back, or
adding a gate that quietly excludes the decode shape, fails a test instead of
silently costing 1.5x. That is the transferable lesson here — the 1.5x was
never the hard part, noticing it was switched off for months was.

cc #1676

@justinchuby

Copy link
Copy Markdown
Owner Author

Retraction: the acc0 matrix in this thread was measured with two different rulers

The 24-cell acc0 matrix I posted here, the 0.436x headline for qwen t=16 s=1, and the "the gap is concurrency-dependent" reading built on them are withdrawn. PR #1722 fixes the instrument and records the full finding as §27 of the ledger.

What was wrong

The native and ORT arms were not computing the same quantity, and the disagreement was concentrated at exactly sessions = 1:

native (before) ORT (before)
denominator wall included thread spawn + 3 warmup steps warmup ran before t0
session start no barrier; staggered spawn absorbed into wall threading.Barrier
over repetitions single shot min (s=1) / max (s≥2) — the luckiest run
statistic wall-clock aggregate at every s 1000/median_ms at s=1, wall-clock aggregate at s≥2

The last row is the one that manufactures a result. The baseline switched from a best-case statistic to a realistic one at sessions = 2. A baseline that does that will always look strongest at sessions = 1 — which is exactly the shape I reported as "we lose at one session, win at two and four".

At tokens = 24, the native warmup-inside-the-clock defect alone charged 27 steps of work against 24 counted tokens: a flat ~11% handicap the ORT arm never paid.

The corrected reading

Under one definition the same cell reads 0.70x, not 0.436x. Six independent runs per arm show it cannot be quoted more precisely than a range:

native  190.9 195.6 197.8 200.3 201.5 220.6   unimodal, ±8%
ORT     218.3 229.8 246.2 396.9 414.9 427.7   two clusters, 1.79x apart

0.436x was never a measurable quantity.

The gap is still real — here is the part that survives

Both arms eat the same contention, so the within-window comparison is valid. Interleaved native/ORT/native, native A/A partner at 1.018:

arm tok/s achieved GB/s (same 145.7 MB/token footprint)
native acc0 196.0 28.5
ORT 279.5 40.7

ORT sustains 1.43x our bandwidth on the identical footprint in identical conditions. It demonstrates the bandwidth was available, so the deficit is neither the memory system nor the busy host.

Two hypotheses I am explicitly not claiming

  • "ORT is bimodal here." The slow cluster's intra-run spreads were 67.9%/4.7%/12.0% vs the fast cluster's 1.4%/1.1%/0.7%. Elevated spread confined to one mode is the signature of external contention, not internal bimodality.
  • "The kernel is issue-bound at ~14% of AVX2 FMA peak." Probably right and it is my leading hypothesis, but contention only depresses the measurement, so it is a lower bound, and a lower bound cannot establish distance from a ceiling.

One more correction to the record

The MLP-starvation hypothesis is provisionally falsified: aggregate bandwidth is flat at ~22–28 GB/s across s = 1, 2, 4, 8 rather than rising with session count. That has a consequence I should state plainly, because it undercuts something I previously presented as a win: our absolute throughput is flat in session count, so the s=2/s=4 cells were the baseline degrading, not our kernel scaling. No kernel change should be justified by them.

Why this happened, and the environment fix

The host is shared by three agents on 16 physical cores. During this investigation it was concurrently running another agent's llvm-cov on this crate (~2470% CPU) and another agent's run of this same benchmark binary (~1275% CPU); peak load average 31.25. The same cell measured 197.2 tok/s and 22.8 tok/s in different windows — an 8.6x environmental swing.

The trap worth naming: several contaminated runs reported intra-run spread_% under 6%. A tight spread means the contention was steady, not that the host was idle. Intra-run spread is not a contention detector. The new driver refuses to start a cell while another process exceeds 150% CPU and marks the cell UNTRUSTED rather than proceeding silently.

Real ratios will be published once a quiet window is available.

justinchuby added a commit that referenced this pull request Aug 22, 2026
…er-nibble

Sweeping block_size varies how often the per-block epilogue runs while leaving
the nibble count unchanged. Six interleaved, independently launched pairs at
qwen t=16 s=1 acc0:

  block 32   190.7 195.5 204.0 203.9 192.5 199.2   median 197.4
  block 64   307.0 283.3 281.9 298.0 278.0 303.4   median 290.7

1.47x, six out of six, distributions do not overlap, for 1.11x *less* traffic.
A per-nibble cost cannot produce that, which falsifies the unpack/convert
hypothesis recorded earlier in this section.

The code agrees: `chunks = block_size / 32` is exactly 1 at the production
block size, so every four FMAs pay a full epilogue -- a branchy
`BorrowedScales::get` discriminant test, a bounds-checked scale fetch, a
`layout.zero_point` lookup that is a *constant* whenever no zero-point tensor
is supplied (the default path), a broadcast, an FMA, and three scalar ops.

Two honesty notes recorded with it:

* The first single-shot sweep read 2.20x for 32->64. It had drawn a fast
  placement at block 64. Withdrawn in favour of the paired 1.47x -- the same
  trap this section is otherwise about.
* Counting instructions predicts only ~1.14x, not 1.47x, so the decomposition
  is incomplete and is not claimed. Direction established, mechanism not fully
  attributed.

Also records that the host is bistable independently of contention: eight
interleaved ORT/native pairs show both arms landing in a fast or a slow mode
per process launch, with the low-intra-spread runs clustering at each arm's
fast mode. A cpuset mask pins the pool to 16 physical cores across both L3
domains but does not fix which thread lands where, so single-run ratios on
this host are not reproducible whichever arm they favour.

Refs #1676, #1679, #1712

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 22, 2026
…pus review)

Opus review caught me applying a standard to two claims and then not applying
it to a third, which is exactly what the review was asked to look for.

§27 argued that a within-window ratio survives contention because both arms eat
it, and concluded ORT "reaching a bandwidth we do not" was stable even if the
multiplier was not. This document's own bistability table refutes that:

* pair 8 is native 264.9 vs ORT 255.0 -- native out-bandwidths ORT in the same
  shared window, so the direction is not universal;
* native's fast placement, 335.6 tok/s = 48.9 GB/s, is higher than the 40.7
  GB/s ORT figure quoted as decisive.

The A/A of 1.018 does not rescue it: that partner controls native-vs-native
placement across two native launches and says nothing about which L3 placement
the separately launched ORT process drew. Quoting native's slow placement
against ORT's fast one and calling the gap stable is the same lower-bound
objection used to withhold the "ORT is intrinsically bimodal" and "~14% of FMA
peak" claims.

Restated at the strength the evidence carries -- and the conclusion is now
stronger, not weaker: native's own fast mode proves the memory system supplies
at least 48.9 GB/s to our kernel, so our common-case 28.5 GB/s is not a
hardware ceiling but headroom we are leaving. That statement does not depend on
the ORT arm at all.

Two further review fixes:

* Block 16 was mis-attributed to `chunks = block_size / 32` evaluating to zero.
  That line is unreachable for block 16: the dispatch gate at
  matmul_nbits.rs:1568 requires `block_size.is_multiple_of(32)`, so block 16
  never enters the nblock kernel and goes to `borrowed_affine_int4_matmul`.
  Exclusion right, cited path wrong.
* The acc4 regime document still describes the pre-unification ORT protocol and
  an invocation that no longer prints `steady_median_ms=`. Marked superseded
  rather than silently left to mislead.

While there, connected a loose end: that document recorded
ONNX_GENAI_CPU_DECODE_THREADS=2 producing timings identical to =1 and dropped
the row rather than explain it. It reproduces here, and CPU-time accounting
supplies the mechanism -- the =2 run consumes 71% of one core against 98% for
=1 at equal user time, so the second worker is parked rather than computing.

Refs #1676, #1679, #1712

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 22, 2026
…retracts the 0.436x headline) (#1722)

## Summary

The int4 decode A/B was **dividing two numbers that were not the same
statistic**, and the mismatch was concentrated at exactly `sessions = 1`
— the configuration the whole "acc0 single-session gap" conclusion rests
on.

This unifies the two arms, writes the definition down where it cannot
drift again, and **retracts** the 24-cell matrix, the `0.436x` qwen
`t=16 s=1` headline, and the "the gap is concurrency-dependent" reading
posted to #1679 / #1676.

## The four biases

| | native (before) | ORT (before) |
|---|---|---|
| denominator | wall included thread spawn + 3 warmup steps | warmup ran
before `t0` |
| session start | no barrier; staggered spawn absorbed into `wall` |
`threading.Barrier` |
| over repetitions | single shot | `min` (s=1) / `max` (s≥2) — the
luckiest run |
| **statistic** | wall-clock aggregate at every `s` | **`1000/median_ms`
at s=1, wall-clock aggregate at s≥2** |

The last row is the one that manufactures a result: the baseline
switched from a **best-case** statistic to a **realistic** one at
`sessions = 2`. A baseline that does that is guaranteed to look
strongest at `sessions = 1` — which is precisely the shape that was
reported as "we lose at one session and win at two and four".

Separately, at `tokens = 24` the native warmup-inside-the-clock defect
charged 27 steps of work against 24 counted tokens: a flat ~11% handicap
the ORT arm never paid at any session count.

Both sides now use one definition — numerator `sessions * tokens`;
denominator wall from **barrier release** to last join; warmup
**outside** the clock; **median** over repetitions — and both print
`spread_%`.

## What the number actually is

`qwen t=16 s=1 acc=0`, published as **0.436x**, reads **0.70x** under
one definition. Six independent runs per arm show it cannot honestly be
quoted more precisely than a **range**:

```
native  190.9 195.6 197.8 200.3 201.5 220.6   unimodal, ±8%
ORT     218.3 229.8 246.2 396.9 414.9 427.7   two clusters, 1.79x apart
```

**0.436x was never a measurable quantity.** It is `max`-over-reps of
ORT's fast cluster over a single-shot native run carrying an 11%
handicap.

## Two conclusions deliberately *not* drawn

- **"ORT is bimodal, so the anomaly is in the baseline."** Not
supported. The slow cluster's intra-run spreads were 67.9% / 4.7% /
12.0% against the fast cluster's 1.4% / 1.1% / 0.7%. Elevated spread
confined to the slow mode is the signature of **external contention**; a
genuinely bimodal implementation would be stable in *both* modes.
- **"The kernel is issue-bound at ~14% of FMA peak."** Directionally
supported and probably right, but contention only *depresses* the
measurement, so it is a **lower bound** — and a lower bound cannot
establish distance from a ceiling. Recorded as a hypothesis to prove on
a quiet host.

## What does survive

Both arms eat the same contention, so the **within-window** comparison
is valid. Interleaved native / ORT / native, native A/A partner at
**1.018**:

| arm | tok/s | achieved GB/s (same 145.7 MB/token footprint) |
|---|---|---|
| native acc0 | 196.0 | **28.5** |
| ORT | 279.5 | **40.7** |

**ORT sustains 1.43x our bandwidth on the identical footprint in
identical conditions** — it *demonstrates* the bandwidth was available.
So the deficit is real, and it is neither the memory system nor the busy
host.

Also: the **MLP-starvation hypothesis is provisionally falsified** —
aggregate bandwidth is flat at ~22–28 GB/s across `s = 1, 2, 4, 8`
rather than rising. Consequence worth stating plainly: our absolute
throughput is **flat in session count**, so the `s=2`/`s=4` "wins" were
the baseline degrading, not the kernel scaling. No kernel change should
be justified by them.

## Why the measurement environment gets its own section

Mid-investigation the host was found running, concurrently: another
agent's `cargo test`/`llvm-cov` on this crate (~2470% CPU), **another
agent's run of this same benchmark binary** (~1275% CPU), and a stray
`while :; do :; done`. Peak load average **31.25** on 16 physical cores.

The same cell measured **197.2 tok/s** in one window and **22.8 tok/s**
in another — an **8.6x** environmental swing.

The trap: several contaminated runs reported intra-run `spread_%`
**under 6%**. A tight spread means the contention was *steady*, not that
the host was idle. Intra-run spread is not a contention detector.
`acc0_gap_matrix.py` therefore refuses to start a cell while any other
process exceeds 150% CPU, and marks the cell `UNTRUSTED` rather than
silently proceeding — because "give up after a timeout and measure
anyway" is exactly how the bad numbers got made.

## Changes

- `int4_decode_loop_ab.rs` — barrier; warmup outside the clock;
`PROBE_REPS` with median; `spread_%`; the definition as a table in the
module docs.
- `ort_matmulnbits_baseline.py` — deleted `run_one`/`steady_median`
(dead code computing a *different* statistic); all session counts
through `run_concurrent`; `max` → `median`; `spread_pct`.
- `acc0_gap_matrix.py` (new) — matrix driver under the single
definition, per-cell interleaved A/A, tok/s → achieved GB/s, quiet-host
gate with competing-process detection.
- `docs/performance/CPU_MATMUL_ASSIGNMENT.md` — §27.
- Three formatting-only hunks outside `benches/` are `cargo fmt` output
for code that landed unformatted on main in #1715; without them the
repo's fmt gate cannot pass. **No semantic change** — flagging
separately since it means main's fmt gate is not currently enforcing.

## Validation

No shipped code changes (benches are not part of the library). `cargo
fmt --all --check` clean; `cargo clippy --all-targets -p
onnx-runtime-ep-cpu -- -D warnings` clean. Full 21-gate matrix running;
will report before undrafting.

## Follow-ups (not in this PR)

- Re-run the matrix on a quiet host and publish the real ratios.
- Prove or kill the unpack/convert issue-cost hypothesis in
`borrowed_int4_nblock4_avx2` with a mechanism-isolating ablation.
- Dispatch-width evidence handed to the runtime owner (total CPU-seconds
flat across widths; `sys` time up ~20x from `t<=2` to `t>=4`) rather
than tuned around in the kernel.

Refs #1676, #1679, #1712

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant