Repository navigation
perf(cpu): enable the register-blocked int4 decode kernel at accuracy_level=0 (1.17-1.69x) - #1679
Conversation
…_level=0 #1104 added `borrowed_affine_int4_matmul_nblock` behind a default-*off* env toggle "until the win is measured". The measurement never happened, so the production default kept taking the per-column path and the kernel sat unused. Route counters instrumented from operator entry through the kernel confirm the dormancy: at accuracy_level=0 every call reached the borrowed guard and every one went to `borrowed_affine_int4_matmul` (nblock=0, percolumn=95). Every parity test still passed, because they all call the kernel directly. Measured 1.17x-1.69x on the decode A/B, best at block 32 (the common llama/qwen block size), and 1.47x-1.62x aggregate throughput under 2-4 concurrent sessions. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #1679 +/- ##
==========================================
+ Coverage 81.00% 81.02% +0.01%
==========================================
Files 384 384
Lines 180771 180893 +122
Branches 180771 180893 +122
==========================================
+ Hits 146429 146563 +134
+ Misses 29396 29384 -12
Partials 4946 4946
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
|
Windows ARM64 investigated and cleared before merge. The run at Attribution: this PR's only functional change is the default of Separately, All 19 required checks green. Merging. |
…f16 layout divergence (#1701) Three entries in `docs/performance/CPU_MATMUL_ASSIGNMENT.md` for work that landed or was closed today. Documentation only — no code. - **§23 — the acc0 route** (`99f105d52`, #1679). #1104's register-blocked int4 kernel shipped default-off "until the win is measured" and the measurement never happened, so `accuracy_level = 0` took the per-column path for the kernel's whole life; route counters read `nblock = 0`. Records the **corrected** attribution: the first one was inflated ~3x because the probe's `fetch_add` sat in the timed path. Clean, the hreduce removal is worth 1.02x — nothing — and the entire 1.48x is four-column activation reuse. The numerics regress up to 3.70x relatively; disclosed rather than buried. - **§24 — the t=8 "wash"** (#1680). The premise does not reproduce. The win is flat through pool width 8 and collapses at 12+; root cause is `decode_spmd.rs::node_shards` pinning worker *i* to `allowed_cpus()[i]` in logical order, so 16 workers land on 8 physical cores. Bandwidth (83 GB/s available vs 41 GB/s drawn) and task grain were tried first and discarded. **No kernel change** — handed to the runtime owner with the measurement. - **§25 — the f16/bf16 layout divergence** (`2e1cfb67c`, #1687, closing #1381). Same math, same bytes, three prices; `[K,N]` crosses a page every `p`. Software prefetch was tried first and is a recorded negative. Accuracy moves the same way, so there was no trade to weigh. Also records the memory-plan coupling that could have gone badly: the #1056 predictor was Apple-only for `MatMul`, so the transpose would have been invisible to the plan on x86. §23 and §25 contain the same error in two disguises — an instrument that changed what it measured, and a numerics test built from exactly-representable operands (`*0.125`) that could not see reassociation at all and reported a confident zero. Both are written up as such, next to §18's version of it. All seven repo policy scripts pass (`check_publish_order`, `check_profile_table`, `check_platform_naming`, `check_dispatch_reachability`, `check_dispatch_manifest`, `check_feature_gate_coverage`, `verify_documented_env_vars`). --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The acc0-vs-ORT matrix that #1679 was missingThe §23 record I merged has only route-vs-route numbers (nblock on/off). It
Reading it honestlyThe route win is real, uniform, and larger than the ledger says. But "1.84x behind ORT" does not generalise, and should stop being quoted as a
Single-session decode remains the real deficit and is where the next work One number I am not willing to stand behind yet: the qwen t=16 s=1 ORT Why
|
Retraction: the acc0 matrix in this thread was measured with two different rulersThe 24-cell acc0 matrix I posted here, the 0.436x headline for What was wrongThe native and ORT arms were not computing the same quantity, and the disagreement was concentrated at exactly
The last row is the one that manufactures a result. The baseline switched from a best-case statistic to a realistic one at At The corrected readingUnder one definition the same cell reads 0.70x, not 0.436x. Six independent runs per arm show it cannot be quoted more precisely than a range: 0.436x was never a measurable quantity. The gap is still real — here is the part that survivesBoth arms eat the same contention, so the within-window comparison is valid. Interleaved native/ORT/native, native A/A partner at 1.018:
ORT sustains 1.43x our bandwidth on the identical footprint in identical conditions. It demonstrates the bandwidth was available, so the deficit is neither the memory system nor the busy host. Two hypotheses I am explicitly not claiming
One more correction to the recordThe MLP-starvation hypothesis is provisionally falsified: aggregate bandwidth is flat at ~22–28 GB/s across Why this happened, and the environment fixThe host is shared by three agents on 16 physical cores. During this investigation it was concurrently running another agent's The trap worth naming: several contaminated runs reported intra-run Real ratios will be published once a quiet window is available. |
…er-nibble Sweeping block_size varies how often the per-block epilogue runs while leaving the nibble count unchanged. Six interleaved, independently launched pairs at qwen t=16 s=1 acc0: block 32 190.7 195.5 204.0 203.9 192.5 199.2 median 197.4 block 64 307.0 283.3 281.9 298.0 278.0 303.4 median 290.7 1.47x, six out of six, distributions do not overlap, for 1.11x *less* traffic. A per-nibble cost cannot produce that, which falsifies the unpack/convert hypothesis recorded earlier in this section. The code agrees: `chunks = block_size / 32` is exactly 1 at the production block size, so every four FMAs pay a full epilogue -- a branchy `BorrowedScales::get` discriminant test, a bounds-checked scale fetch, a `layout.zero_point` lookup that is a *constant* whenever no zero-point tensor is supplied (the default path), a broadcast, an FMA, and three scalar ops. Two honesty notes recorded with it: * The first single-shot sweep read 2.20x for 32->64. It had drawn a fast placement at block 64. Withdrawn in favour of the paired 1.47x -- the same trap this section is otherwise about. * Counting instructions predicts only ~1.14x, not 1.47x, so the decomposition is incomplete and is not claimed. Direction established, mechanism not fully attributed. Also records that the host is bistable independently of contention: eight interleaved ORT/native pairs show both arms landing in a fast or a slow mode per process launch, with the low-intra-spread runs clustering at each arm's fast mode. A cpuset mask pins the pool to 16 physical cores across both L3 domains but does not fix which thread lands where, so single-run ratios on this host are not reproducible whichever arm they favour. Refs #1676, #1679, #1712 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…pus review) Opus review caught me applying a standard to two claims and then not applying it to a third, which is exactly what the review was asked to look for. §27 argued that a within-window ratio survives contention because both arms eat it, and concluded ORT "reaching a bandwidth we do not" was stable even if the multiplier was not. This document's own bistability table refutes that: * pair 8 is native 264.9 vs ORT 255.0 -- native out-bandwidths ORT in the same shared window, so the direction is not universal; * native's fast placement, 335.6 tok/s = 48.9 GB/s, is higher than the 40.7 GB/s ORT figure quoted as decisive. The A/A of 1.018 does not rescue it: that partner controls native-vs-native placement across two native launches and says nothing about which L3 placement the separately launched ORT process drew. Quoting native's slow placement against ORT's fast one and calling the gap stable is the same lower-bound objection used to withhold the "ORT is intrinsically bimodal" and "~14% of FMA peak" claims. Restated at the strength the evidence carries -- and the conclusion is now stronger, not weaker: native's own fast mode proves the memory system supplies at least 48.9 GB/s to our kernel, so our common-case 28.5 GB/s is not a hardware ceiling but headroom we are leaving. That statement does not depend on the ORT arm at all. Two further review fixes: * Block 16 was mis-attributed to `chunks = block_size / 32` evaluating to zero. That line is unreachable for block 16: the dispatch gate at matmul_nbits.rs:1568 requires `block_size.is_multiple_of(32)`, so block 16 never enters the nblock kernel and goes to `borrowed_affine_int4_matmul`. Exclusion right, cited path wrong. * The acc4 regime document still describes the pre-unification ORT protocol and an invocation that no longer prints `steady_median_ms=`. Marked superseded rather than silently left to mislead. While there, connected a loose end: that document recorded ONNX_GENAI_CPU_DECODE_THREADS=2 producing timings identical to =1 and dropped the row rather than explain it. It reproduces here, and CPU-time accounting supplies the mechanism -- the =2 run consumes 71% of one core against 98% for =1 at equal user time, so the second worker is parked rather than computing. Refs #1676, #1679, #1712 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…retracts the 0.436x headline) (#1722) ## Summary The int4 decode A/B was **dividing two numbers that were not the same statistic**, and the mismatch was concentrated at exactly `sessions = 1` — the configuration the whole "acc0 single-session gap" conclusion rests on. This unifies the two arms, writes the definition down where it cannot drift again, and **retracts** the 24-cell matrix, the `0.436x` qwen `t=16 s=1` headline, and the "the gap is concurrency-dependent" reading posted to #1679 / #1676. ## The four biases | | native (before) | ORT (before) | |---|---|---| | denominator | wall included thread spawn + 3 warmup steps | warmup ran before `t0` | | session start | no barrier; staggered spawn absorbed into `wall` | `threading.Barrier` | | over repetitions | single shot | `min` (s=1) / `max` (s≥2) — the luckiest run | | **statistic** | wall-clock aggregate at every `s` | **`1000/median_ms` at s=1, wall-clock aggregate at s≥2** | The last row is the one that manufactures a result: the baseline switched from a **best-case** statistic to a **realistic** one at `sessions = 2`. A baseline that does that is guaranteed to look strongest at `sessions = 1` — which is precisely the shape that was reported as "we lose at one session and win at two and four". Separately, at `tokens = 24` the native warmup-inside-the-clock defect charged 27 steps of work against 24 counted tokens: a flat ~11% handicap the ORT arm never paid at any session count. Both sides now use one definition — numerator `sessions * tokens`; denominator wall from **barrier release** to last join; warmup **outside** the clock; **median** over repetitions — and both print `spread_%`. ## What the number actually is `qwen t=16 s=1 acc=0`, published as **0.436x**, reads **0.70x** under one definition. Six independent runs per arm show it cannot honestly be quoted more precisely than a **range**: ``` native 190.9 195.6 197.8 200.3 201.5 220.6 unimodal, ±8% ORT 218.3 229.8 246.2 396.9 414.9 427.7 two clusters, 1.79x apart ``` **0.436x was never a measurable quantity.** It is `max`-over-reps of ORT's fast cluster over a single-shot native run carrying an 11% handicap. ## Two conclusions deliberately *not* drawn - **"ORT is bimodal, so the anomaly is in the baseline."** Not supported. The slow cluster's intra-run spreads were 67.9% / 4.7% / 12.0% against the fast cluster's 1.4% / 1.1% / 0.7%. Elevated spread confined to the slow mode is the signature of **external contention**; a genuinely bimodal implementation would be stable in *both* modes. - **"The kernel is issue-bound at ~14% of FMA peak."** Directionally supported and probably right, but contention only *depresses* the measurement, so it is a **lower bound** — and a lower bound cannot establish distance from a ceiling. Recorded as a hypothesis to prove on a quiet host. ## What does survive Both arms eat the same contention, so the **within-window** comparison is valid. Interleaved native / ORT / native, native A/A partner at **1.018**: | arm | tok/s | achieved GB/s (same 145.7 MB/token footprint) | |---|---|---| | native acc0 | 196.0 | **28.5** | | ORT | 279.5 | **40.7** | **ORT sustains 1.43x our bandwidth on the identical footprint in identical conditions** — it *demonstrates* the bandwidth was available. So the deficit is real, and it is neither the memory system nor the busy host. Also: the **MLP-starvation hypothesis is provisionally falsified** — aggregate bandwidth is flat at ~22–28 GB/s across `s = 1, 2, 4, 8` rather than rising. Consequence worth stating plainly: our absolute throughput is **flat in session count**, so the `s=2`/`s=4` "wins" were the baseline degrading, not the kernel scaling. No kernel change should be justified by them. ## Why the measurement environment gets its own section Mid-investigation the host was found running, concurrently: another agent's `cargo test`/`llvm-cov` on this crate (~2470% CPU), **another agent's run of this same benchmark binary** (~1275% CPU), and a stray `while :; do :; done`. Peak load average **31.25** on 16 physical cores. The same cell measured **197.2 tok/s** in one window and **22.8 tok/s** in another — an **8.6x** environmental swing. The trap: several contaminated runs reported intra-run `spread_%` **under 6%**. A tight spread means the contention was *steady*, not that the host was idle. Intra-run spread is not a contention detector. `acc0_gap_matrix.py` therefore refuses to start a cell while any other process exceeds 150% CPU, and marks the cell `UNTRUSTED` rather than silently proceeding — because "give up after a timeout and measure anyway" is exactly how the bad numbers got made. ## Changes - `int4_decode_loop_ab.rs` — barrier; warmup outside the clock; `PROBE_REPS` with median; `spread_%`; the definition as a table in the module docs. - `ort_matmulnbits_baseline.py` — deleted `run_one`/`steady_median` (dead code computing a *different* statistic); all session counts through `run_concurrent`; `max` → `median`; `spread_pct`. - `acc0_gap_matrix.py` (new) — matrix driver under the single definition, per-cell interleaved A/A, tok/s → achieved GB/s, quiet-host gate with competing-process detection. - `docs/performance/CPU_MATMUL_ASSIGNMENT.md` — §27. - Three formatting-only hunks outside `benches/` are `cargo fmt` output for code that landed unformatted on main in #1715; without them the repo's fmt gate cannot pass. **No semantic change** — flagging separately since it means main's fmt gate is not currently enforcing. ## Validation No shipped code changes (benches are not part of the library). `cargo fmt --all --check` clean; `cargo clippy --all-targets -p onnx-runtime-ep-cpu -- -D warnings` clean. Full 21-gate matrix running; will report before undrafting. ## Follow-ups (not in this PR) - Re-run the matrix on a quiet host and publish the real ratios. - Prove or kill the unpack/convert issue-cost hypothesis in `borrowed_int4_nblock4_avx2` with a mechanism-isolating ablation. - Dispatch-width evidence handed to the runtime owner (total CPU-seconds flat across widths; `sys` time up ~20x from `t<=2` to `t>=4`) rather than tuned around in the kernel. Refs #1676, #1679, #1712 --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
What this changes
One default flip:
borrowed_int4_nblock_enabled()now returnstruewhen unset. Plus the two tests that should have existed to make that flip unnecessary.The finding
#1104 built
borrowed_affine_int4_matmul_nblock— the register-blocked int4 decode kernel — measured 1.46x on a 14B model, proved byte-identical output, added parity tests, and shipped it default off, "until the win is measured, exactly like the toggles that preceded it."The measurement never happened. The kernel has been sitting in the tree unused ever since, while
accuracy_level = 0— the production default — kept taking the per-column path it was written to replace.Every parity test passed the whole time, because they all call the kernel directly. Nothing asserted which route production actually took.
Route proof
I did not infer this from reading guards. I instrumented counters from operator entry through the kernel and ran the decode loop:
Every call reaches the borrowed guard; every one lands in
borrowed_affine_int4_matmul;nblockis never entered. 100% of blocks take the SIMD path, so the per-column path was not falling back to scalar — it was simply the wrong kernel.Attribution — and a hypothesis of mine that the data killed
Reading the two kernels, the obvious explanation is the per-block horizontal reduction. The per-column path pays a full 8-lane hreduce (
extractf128/movehl/shuffle, each dependent on the last) every 32 weights — once per four FMAs — then folds the result into a serial f32sum. The N-blocked kernel keeps the scale in a vector accumulator and reduces once per column. That looked like the whole story.I made the column-group width tunable to separate that restructuring from the four-column activation reuse:
Removing the per-block horizontal reduction is worth 1.02x — nothing measurable. The block loop has enough independent work across blocks for the out-of-order engine to hide that dependency chain entirely. The entire win is the four-column activation reuse (1.45x from group 1 to group 4): at
m == 1each activation vector is loaded once and feeds four columns' FMAs.A correction to my own first numbers
My first attribution run reported 3.14x for the group-of-1 arm and 4.80x overall. Those were artifacts of my own instrumentation: the route probe incremented an atomic once per 32-weight block, which fired 129M times in the per-column arm and zero times in the N-blocked arm, inflating the baseline ~3x. The table above is from a clean binary with only the group-width knob compiled in. I mention it because the wrong numbers were self-consistent and looked plausible.
Full matrix (
accuracy_level = 0, min of 6 interleaved, pinned to physical cores)Concurrent sessions, aggregate tok/s (higher better):
No regression anywhere in the matrix. The win is largest at block 32, which is what llama/qwen ship.
Numerics — the honest version
The existing parity test asserts N-blocked matches per-column within 1e-3. That is not a strong enough contract for a default: it would let a less accurate kernel ship as long as it stayed near the one it replaced. So the new test measures both against an f64 reference.
I expected the replacement to be at least as accurate. It is not. The two differ structurally: per-column forms
(dot - actsum*zp) * scaleper block and accumulates in one serial f32 chain; N-blocked accumulatessum(dot*scale)andsum(actsum*zp*scale)separately and subtracts once at the end. With unsigned nibbles0..=15and midpoint 8 those two running sums are comparable in magnitude and the answer is their difference, so the separated form gives cancellation somewhere to hide.Up to 3.70x worse in relative terms — a real cost of the 1.48x, not something I want buried. It stays at 2.4e-5 worst case, two orders inside the 1e-3 this route's parity tests already accept, and the
accuracy_level = 0contract itself is untouched: both paths are pure f32 FMA over f32 activations and neither is ever diverted through an int8/int16 kernel. The test includes deliberately hostile cases (activation gain up to 1024x, weights pinned near the midpoint) chosen to maximise exactly this cancellation.Tests
acc0_decode_reaches_the_nblocked_kernel_by_default— asserts the route, which is the check whose absence let this sit dormant. Mutation-verified: reverting the default tofalsefails it. TakesDISPATCH_PROBE_LOCKper the repo's meta-test. Excluded underfeature = "mlas", wheremlas_sqnbit_owns_fp32_computelegitimately claims acc0 before the borrowed guard; the shipped default artifact links no MLAS, which is the configuration asserted.nblock_holds_the_f32_contract_against_f64— the differential test above. The 8x ratio guard is a regression bound sized off the measured 3.70x, widened for AVX-512 hosts where the per-column path uses the 512-bit block dot and so has its own error profile.Validation
roy_validate.shPASS=21 FAIL=0 SKIP=0, two consecutive full runs.cargo test -p onnx-runtime-ep-cpu1614 passed, three consecutive runs.--features mlas160 matmul_nbits tests pass.Three real failures caught and fixed on the first validation pass, all mine: fmt; the mlas-feature route conflict above; and aarch64 cross-compile, where the new counter was dead code because the N-blocked kernel is x86-only and the cross lane builds tests with
-D warnings.One earlier full run reported
PASS=20 FAIL=1and its log was overwritten before I read it; two subsequent full runs and three consecutive-p onnx-runtime-ep-cpuruns were clean. Flagging rather than dismissing it.Scope
The env toggle stays, inverted, so the per-column path remains reachable in the same binary for A/B.
accuracy_level = 4, aarch64, MLAS and the prefill routes are untouched.Refs #1676.