Skip to content

perf(cpu-ep): stop QLinearMatMul re-allocating its buffers on every call - #1133

Merged
justinchuby merged 2 commits into
mainfrom
squad/roy-qlinear-staging
Aug 17, 2026
Merged

justinchuby merged 2 commits into
mainfrom
squad/roy-qlinear-staging

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 17, 2026 •

Copy link
Copy Markdown
Owner

What this is

QLinearMatMul was 2.27x slower than ORT at K=N=2048, M=128 with a 16-thread pool, but 1.01x — parity — on one thread (plugin-level, interleaved, both sides pinned; 8.281 ms ours vs 8.183 ms ORT). A deficit that only appears when threads are added is not arithmetic, so this started as a decomposition rather than a tuning exercise.

Hardware for every number below: 32-vCPU / 16-core EPYC 9V74 (Zen4, AVX-512), --release --features mlas, both thread pools pinned to the same budget (ONNX_GENAI_MLAS_THREADPOOL_THREADS + RAYON_NUM_THREADS for us, intra-op for ORT). Ratios are written explicitly as ours/ORT, so greater than 1.0 means we are slower.

Where the time was going

The in-tree qlinear_phase_report was extended to time the fused MLAS primitive on its own, so whole - fused isolates everything our wrapper does around it. At K=N=2048, M=128:

threads whole call MLAS fused GEMM+requantize our wrapper
1 8.31 ms 8.17 ms 145 us
4 2.57 ms 2.26 ms 313 us
16 1.70 ms 1.13 ms 576 us (34% of the call)

MLAS scales 7.2x across that range. Our wrapper gets 4x more expensive as threads are added, which is the entire deficit.

The thread pool was ruled out first rather than assumed innocent. An empty parallel_for dispatch measures 0.03 us at T=1, 2.6 us at T=8 and 4.7 us at T=16, and a 100 us tile costs 103.5 us at T=16 — a few microseconds, not a few hundred. It is also insensitive to a 2 ms idle gap, so parking and wake latency are not involved either.

Bisecting the wrapper by making a direct MLAS call allocate the same buffers the kernel allocates found the mechanism:

per-call buffer cost at T=1 cost at T=16
A copy (256 KiB) 13 us 136-168 us
i32 accumulator (1 MiB) ~0 us 301 us
staged output copy ~0 us ~43 us

Per-call buffer churn is cheap on one thread and expensive on sixteen. A fresh multi-page mapping has to be page-faulted in, and MLAS first-touches the accumulator from every worker, so the fault count follows the pool size. Freeing that mapping then forces a TLB shootdown IPI to every core the process is running on. Neither cost exists at T=1, which is exactly why the kernel looked healthy there.

The change

Three buffers, three fixes, none of which changes an output byte:

  • A is borrowed when the view is already dense (dense_bytes returns a Cow), instead of always being copied through to_dense_bytes — which allocates a zeroed Vec and then overwrites all of it, so the copy path was also paying for a pointless memset. A strided or non-host-accessible input still copies exactly as before. The one route that rewrites A (the sign flip on signed-A x unsigned-B) goes through Cow::to_mut, which copies a borrowed operand before touching it, so the caller's input is never written through.
  • The result is requantized straight into the output tensor when that tensor is contiguous and host-accessible. requantize_rows now takes destination: &mut [u8] (exact length, with a length-mismatch guard) rather than appending into a &mut Vec<u8>. A strided or non-host output still stages into a Vec and scatters, via a small OutputSink::{Direct, Staged} enum whose region(base, len) hands out one batch's slice.
  • The i32 accumulator is parked on a thread-local between calls, bounded at 32 MiB so an outsized shape is released rather than retained forever. It is thread-local rather than shared because execute takes &self; a lock would serialise precisely the concurrent calls this is meant to speed up. An early ? return simply drops the buffer, which costs the next call one allocation and cannot affect a result.

This follows an existing in-repo convention rather than inventing one: matmul_nbits.rs has had the same direct_result/owned_result + contiguous_host_slice shape for a while. QLinearMatMul was the outlier.

Result

Kernel level (in-process, our side only — there is no ORT number at this level, so no ratio is quoted here), u8, K=N=2048, M=128:

threads before after change
1 8.31 ms 8.29 ms unchanged
4 2.57 ms 2.35 ms 9% faster
16 1.70 ms 1.17 ms 31% faster

Wrapper cost at 16 threads: 576 us -> 54 us.

Plugin level vs real ORT — this is the apples-to-apples comparison, and the only place a ours/ORT ratio is defensible. Three cdylibs (this branch / main / plain ORT) measured in 11 interleaved rounds of 41 iterations after 3 warmups, T=16 on both sides, medians of the per-round p50s:

K=N=2048 ours before ours after ORT before, ours/ORT after, ours/ORT
u8, M=128 (5 rounds) 1.869 ms 1.178 ms 0.824 ms 2.27x 1.43x
u8, M=1 (11 rounds) 0.0531 ms 0.0530 ms 0.0325 ms 1.63x 1.63x
u8, M=128 at T=1 (3 rounds) 8.281 ms 8.322 ms 8.183 ms 1.01x 1.02x

So the prefill shape goes from 2.27x to 1.43x ours/ORT at 16 threads, and decode is unchanged to three digits — its buffers are a few KiB, so buffer churn was never its problem.

The T=1 row is the control that makes the whole diagnosis falsifiable: at one thread we were already at parity with ORT (1.01x) on the exact shape where we were 2.27x slower at sixteen. A kernel that is level at T=1 and 2.27x behind at T=16 is not losing on arithmetic, and the fix had to be — and was — in what scales with the thread count. That row is also tight enough to trust: p50-to-p90 spread under 1.2% on all three sides, versus 3-4x at T=16.

This is a real gain and it is still not a win. We remain ~1.4x slower than ORT at M=128 and ~1.6x slower at M=1. This PR does not claim otherwise, no dispatch decision rests on it, and both shapes stay on the open list.

Honest note on the tail

Per-round p90s (medians across rounds, u8 M=128): ours before 3.88 ms, ours after 3.29 ms, ORT 0.83 ms. So p90 improves but our dispersion stays much wider than ORT's, whose p90 sits almost on its p50. This box is shared with other agents' builds and each side's rounds run in separate processes at different moments, so I am not claiming the tail gap is a property of the kernel — but I am not hiding it either. It is unexplained and stays on the open list.

The ORT p50 itself moved between 1.11 ms (rounds 1-2) and 0.82 ms (rounds 3-5) on the same box, which is why every ratio above is taken from interleaved rounds and never from numbers measured in different sittings.

What is not improved

  • bench_qlinear_u8_m1 (decode) stays at ~1.63x ours/ORT. Unchanged, still open.
  • T=1 is unchanged, as expected: at one thread the buffers were nearly free.
  • bench_qlinear_i8_m1 was measured in an earlier session at ~0.065 ms ours vs 0.54-0.97 ms ORT — 8-15x in our favour. That number is not re-measured in this PR's runs and nothing here depends on it; it is repeated only to say this change does not touch that path.

Behaviour change worth disclosing

On the new direct path, an error raised part-way through a multi-batch call can leave the output tensor partially written, where previously the output was untouched until the whole result had been staged. This is within ORT's kernel contract (an output is undefined when a kernel returns non-OK) and it is what matmul_nbits.rs already does, but it is a real difference and is called out here rather than buried.

Tests

Three new falsifiers, each of which fails if the corresponding optimisation is wrong rather than merely slow. Each was checked by deliberately breaking the thing it guards:

  • a_contiguous_output_is_written_in_place_and_a_strided_one_is_staged — asserts the route actually taken via thread-local counters. Forcing the direct path off makes it fail, so a silent regression to always-staging is caught.
  • the_sign_flip_route_never_writes_through_to_the_callers_input — asserts the caller's A is byte-identical after a flip-route call, and that the call really borrowed A. Without the second assertion the test would pass vacuously the moment dense_bytes stopped borrowing; forcing dense_bytes to always copy now makes it fail.
  • a_batched_call_lands_every_batch_at_its_own_offset_in_place — catches an off-by-one in the batch * m * n region base. Forcing the base to 0 makes it fail in both feature configurations.

The route counters are thread_local! Cells, not global atomics; the harness runs tests in parallel and a global counter made the assertions flaky.

Review

Independently reviewed by an Opus reviewer, which ran the build, both test configurations, the real-ORT suite, and four falsification experiments of its own. Verdict APPROVE, no MAJOR findings. Both MINOR findings are fixed in the second commit:

  • the flip_a special case (taking an owned copy before the flip) was redundant — Cow::to_mut already copies a borrowed operand — and it re-introduced the very to_dense_bytes zeroed allocation this PR removes. It is gone, and the test that guards the flip now also asserts the borrow happened.
  • the headline ratio previously divided a kernel-level "ours" by a plugin-level "ORT". Every ratio in this body is now plugin-level, same-run, interleaved, and labelled with its level, round count and thread count.

Verification run

  • cargo fmt --all --check — clean
  • cargo clippy --all-targets -p onnx-runtime-ep-cpu, with and without --features mlas — no new warnings (the pre-existing needless_return at matmul_nbits.rs:639 is untouched here and is being fixed separately)
  • cargo test -p onnx-runtime-ep-cpu --lib — debug and release, with --features mlas (1335 pass) and without (1307 pass), 0 failures
  • NXRT_REQUIRE_ORT_TESTS=1 cargo test -p onnx-runtime-ep-cpu-plugin --test plugin_ort_e2e --features mlas --release — 43 pass, 0 fail, against real ONNX Runtime

@codecov

codecov Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 83.47107% with 20 lines in your changes missing coverage. Please review.
✅ Project coverage is 79.95%. Comparing base (1f675d2) to head (8572826).
⚠️ Report is 3 commits behind head on main.

Files with missing lines Patch % Lines
.../onnx-runtime-ep-cpu/src/kernels/qlinear_matmul.rs 83.47% 19 Missing and 1 partial ⚠️
Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff            @@
##             main    #1133    +/-   ##
========================================
  Coverage   79.94%   79.95%            
========================================
  Files         368      368            
  Lines      160794   160899   +105     
  Branches   160794   160899   +105     
========================================
+ Hits       128551   128640    +89     
- Misses      27524    27541    +17     
+ Partials     4719     4718     -1     
Flag Coverage Δ
cli-ort-linux 83.79% <ø> (ø)
cli-ort-windows ?
mlas 85.74% <ø> (+0.06%) ⬆️
offline 79.69% <83.47%> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
.../onnx-runtime-ep-cpu/src/kernels/qlinear_matmul.rs 84.88% <83.47%> (-0.06%) ⬇️

... and 4 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

justinchuby and others added 2 commits August 17, 2026 18:00
At K=N=2048, M=128 the kernel was 2.07x ORT with a 16-thread pool but at
parity on one thread, so the deficit was never arithmetic. Splitting the call
showed MLAS's fused GEMM+requantize scaling 7.2x (8.17ms -> 1.13ms) while
everything our wrapper did around it grew from 145us to 576us -- 34% of the
call -- as threads were added.

The cause is per-call buffer churn, which is cheap on one thread and expensive
on sixteen. Every call allocated `A` (m*k), an `i32` accumulator (m*n*4) and a
staging copy of the result (m*n), zero-filled them, and freed them. Freeing a
fresh multi-page mapping forces a TLB shootdown to every core the process runs
on, and MLAS first-touches the accumulator from every worker, so the cost rises
with the pool size rather than staying a fixed overhead.

Three buffers, three fixes, none of which changes an output byte:

* `A` is borrowed when the view is already dense instead of being copied.
  Only the sign-flip route rewrites `A`, and that still takes a private copy,
  because writing through to the caller's input would be a bug.
* the result is requantized straight into a contiguous output tensor. A
  strided or non-host output still stages and scatters exactly as before.
* the `i32` accumulator is parked on a thread-local between calls, bounded at
  32 MiB so an outsized shape is released rather than retained. It is
  thread-local rather than shared because `execute` takes `&self`, and a lock
  would serialise the calls this is meant to speed up.

Measured on a 32-vCPU EPYC 9V74, K=N=2048 M=128, both pools pinned:

| threads | before | after |
| --- | --- | --- |
| 1 | 8.31 ms | 8.29 ms |
| 4 | 2.57 ms | 2.35 ms |
| 16 | 1.70 ms | 1.17 ms |

Wrapper cost at 16 threads falls from 576us to 54us.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…appened

The review pointed out that the `flip_a` special case -- taking an owned copy
of `A` before the sign flip -- is redundant: the flip goes through
`Cow::to_mut`, which copies a borrowed operand before writing to it, so the
caller's input is safe either way. Removing the special case was confirmed by
experiment to leave the test passing, which means the case was buying nothing
and costing a `to_dense_bytes` -- the zeroed allocation this PR set out to
avoid -- on the signed-A-by-unsigned-B route.

Dropping it leaves that route paying a plain `to_vec` at `to_mut`, and every
other route borrowing.

The test that guards this could pass vacuously: if `dense_bytes` ever stopped
borrowing, the caller's `A` would trivially survive and the assertion would
prove nothing. It now also asserts a borrow counter, and forcing `dense_bytes`
to always copy makes it fail.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby force-pushed the squad/roy-qlinear-staging branch from ef58b51 to 8572826 Compare August 17, 2026 18:12
@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 784.15 µs 1.15 ms +46.8%
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.71 ms 2.29 ms +34.4%
⚠️ matmul/large_generic_f32_threads=8/32x1024x1024 5.51 ms 6.95 ms +26.1%
⚠️ matmul/large_generic_f32_threads=1/32x1024x1024 9.85 ms 12.10 ms +22.8%
⚠️ matmul/medium_generic_f32_threads=1/32x512x512 2.53 ms 3.06 ms +21.0%
⚠️ gather/large_f32_threads=1-internal/131072 34.22 µs 41.18 µs +20.3%
⚠️ matmul/medium_generic_bf16_threads=8/32x512x512 576.91 µs 689.81 µs +19.6%
⚠️ reduce_mean/large_f32_threads=1-internal/262144 992.38 µs 1.16 ms +16.5%
⚠️ sampling_latency/top_p_per_token 390.47 µs 452.65 µs +15.9%
✅ gather/small_f32_threads=1-internal/4096 762.5 ns 847.1 ns +11.1%
✅ matmul/small_generic_f32_threads=1/1x256x256 44.87 µs 48.63 µs +8.4%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.97 ms 2.11 ms +7.4%
✅ gather/medium_bf16_threads=1-internal/32768 2.46 µs 2.58 µs +4.9%
✅ matmul/small_generic_f32_threads=8/1x256x256 55.11 µs 57.71 µs +4.7%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.55 ms 3.71 ms +4.4%
✅ reduce_mean/medium_f32_threads=1-internal/65536 264.22 µs 275.76 µs +4.4%
✅ tokenization/encode_tokens_per_second 393.18 µs 408.24 µs +3.8%
✅ gather/small_bf16_threads=1-internal/4096 521.2 ns 536.6 ns +3.0%
✅ matmul/small_generic_f16_threads=8/1x256x256 40.05 µs 40.97 µs +2.3%
✅ grammar_masking/llguidance_compute_mask/32 77.51 µs 79.06 µs +2.0%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.15 ms 2.18 ms +1.5%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 640.61 µs 649.10 µs +1.3%
✅ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 99.76 µs 100.53 µs +0.8%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 90.71 µs 91.27 µs +0.6%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.80 ms 5.83 ms +0.5%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.72 ms 1.73 ms +0.3%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 531.94 µs 530.60 µs -0.3%
✅ sampling_latency/greedy_per_token 3.25 µs 3.23 µs -0.6%
✅ tokenization/decode_tokens_per_second 6.41 ms 6.31 ms -1.6%
✅ sampling_latency/top_k_per_token 55.30 µs 54.28 µs -1.9%
✅ kv_cache/alloc_dealloc_pages 42.26 µs 41.45 µs -1.9%
✅ qwen3_sampling_processors/top_k_partial_selection 146.17 µs 143.35 µs -1.9%
✅ qwen3_sampling_processors/top_k_top_p_fast 673.73 µs 656.65 µs -2.5%
✅ gather/medium_f32_threads=1-internal/32768 4.67 µs 4.55 µs -2.5%
✅ matmul/medium_generic_f16_threads=8/32x512x512 45.35 µs 44.13 µs -2.7%
✅ logit_processing/seven_processor_chain_per_step 332.46 µs 322.06 µs -3.1%
✅ gather/large_f16_threads=1-internal/131072 17.82 µs 17.17 µs -3.6%
✅ gather/large_bf16_threads=1-internal/131072 22.31 µs 21.40 µs -4.1%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 131.73 µs 125.94 µs -4.4%
✅ add/large_f16_threads=1-internal/4194304 2.18 ms 2.08 ms -4.4%
✅ gather/small_f16_threads=1-internal/4096 613.8 ns 578.8 ns -5.7%
✅ matmul/small_generic_bf16_threads=8/1x256x256 43.52 µs 40.84 µs -6.2%
✅ add/small_f16_threads=1-internal/1024 495.7 ns 464.7 ns -6.2%
✅ gather/medium_f16_threads=1-internal/32768 2.74 µs 2.56 µs -6.5%
✅ matmul/small_generic_f16_threads=1/1x256x256 45.26 µs 42.02 µs -7.2%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 108.48 µs 100.26 µs -7.6%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 591.30 µs 543.94 µs -8.0%
✅ add/small_f32_threads=1-internal/1024 246.8 ns 226.6 ns -8.2%
✅ matmul/medium_generic_f16_threads=1/32x512x512 41.15 µs 37.32 µs -9.3%
✅ add/small_bf16_threads=1-internal/1024 506.9 ns 459.3 ns -9.4%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 200.06 µs 180.25 µs -9.9%
✅ reduce_mean/small_f32_threads=1-internal/4096 18.65 µs 16.68 µs -10.6%
✅ add/large_bf16_threads=1-internal/4194304 1.88 ms 1.67 ms -11.4%
✅ matmul/small_generic_bf16_threads=1/1x256x256 46.09 µs 40.67 µs -11.8%
✅ sampling_latency/min_p_per_token 246.42 µs 213.46 µs -13.4%
✅ add/medium_f32_threads=1-internal/262144 29.14 µs 24.79 µs -15.0%
🟢 add/medium_f16_threads=1-internal/262144 130.32 µs 108.83 µs -16.5%
🟢 add/medium_bf16_threads=1-internal/262144 171.71 µs 115.63 µs -32.7%
🟢 add/large_f32_threads=1-internal/4194304 1.13 ms 743.12 µs -34.5%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 4.13 3.59 6.36 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby
justinchuby marked this pull request as ready for review August 17, 2026 18:54
@justinchuby
justinchuby merged commit c62798b into main Aug 17, 2026
13 of 18 checks passed
@justinchuby
justinchuby deleted the squad/roy-qlinear-staging branch August 17, 2026 21:28
@justinchuby

Copy link
Copy Markdown
Owner Author

Decomposition is right and the gates pass on my box (cargo test -p onnx-runtime-ep-cpu --lib 1306 passed / 0 failed, +2 over main). Ruling the thread pool out by measurement before blaming it, and then bisecting the wrapper by making a direct MLAS call allocate the same buffers, is the right shape — the mechanism is established, not inferred.

But the retention bound is stated per-call and the buffer is retained per-thread.

qlinear_matmul.rs:50:

ust const MAX_RETAINED_ACCUMULATOR_BYTES: usize = 32 << 20;

and :52:

ust thread_local! { static ACCUMULATOR: RefCell<Vec<i32>> = ...

The doc comment reasons entirely in single-buffer terms -- "32 MiB covers every accumulator up to an 8M-element result" -- and never multiplies by the pool width. But ACCUMULATOR is parked on every worker thread that ever runs this kernel, for the life of the process. The real process-wide ceiling is 32 MiB x threads:

  • your 32-vCPU EPYC: 1 GiB
  • my 20-thread box: 640 MiB
  • a 128-vCPU server: 4 GiB

The one bound in the PR is the one number that is not the process's exposure, and it is off by the pool width -- which is exactly the variable this PR is about, since the whole finding is that per-call buffer cost scales with thread count. The fix's ceiling scales with the same variable as the bug's cost.

This is the third instance of one shape, and I want to name it because it keeps costing us a round:

Each time the arithmetic in the comment was right about one copy and silent about N. Under-reporting is worse than reporting zero: zero is obviously blind and gets caught at review, whereas a plausible-looking 32 MiB passes admission and then overruns.

What I need before merging

  1. Measure it, do not reason about it. Poll PeakWorkingSet64 by PID (not name; it reads 0 after exit) at 150-250 ms while running a QLinearMatMul model at RAYON_NUM_THREADS 1 vs 16, and report before/after RSS for both. If retention is invisible at T=1 and visible at T=16, that is the multiplier made observable.
  2. Bound the process, not the thread. Either a process-wide budget shared across threads, or state the ceiling as per-thread x threads and pick the constant so the product is defensible.
  3. Route it through the governor (Every resident weight side-buffer must be in the memory plan before it is allocated #1056). The rule is: declare before allocating, account actual bytes, and be declinable. GovernedWeightCache<T> (governed_weight_cache.rs, 21cd05b3) exists for this and reports the buffer's own length so a wrong prediction is detectable. This kernel's pre-pack already merged ungoverned in ac394fd6; this PR adds a second retained buffer beside it, so it is the natural place to fix both.

Note that MAX_RETAINED_ACCUMULATOR_BYTES releasing anything larger is good and I do not want it removed -- the issue is only that the surviving case is multiplied.

The speed work stands on its own and I expect it to merge once the ceiling is honest.

justinchuby pushed a commit that referenced this pull request Aug 17, 2026
The weight-cache guard failed the test it exists to pass. #1133 parked a
QLinearMatMul i32 accumulator in a thread_local! RefCell<Vec<i32>> bounded by a
per-buffer 32 MiB constant, while the buffer is retained on every worker thread
for the life of the process: 640 MiB on a 20-thread box, 1 GiB on a 32-vCPU one,
4 GiB on a 128-vCPU one. Run against that PR's 459 added lines the existing
regex matched zero, because it only knew OnceLock/OnceCell/LazyLock.

That is the guard committing, one level up, the defect it was written to catch:
a check that is green but structurally incapable of failing on the case that
motivated it.

The new job does not merely also-match thread_local. It asks the question the
first job does not -- what multiplies this buffer -- because the underlying
error has now cost four rounds: #1051 reported 247 MB against 592 MB measured,
#1100's ratio test drove a single instantiation so it could not observe the x2,
and #1133 bounded one copy of an N-per-thread buffer. Each comment was correct
about one copy and silent about N. Under-reporting is worse than reporting zero:
zero is obviously blind and gets caught at review, whereas a plausible 32 MiB
passes admission and then overruns.

The error text also requires that any test for such a buffer drive it from more
than one thread, since a single-instantiation test cannot observe an xN
multiplier -- which is exactly how #1100 shipped with the factor unmeasured.

Falsified against real history rather than assumed to work:

  #1133  (must fail)  matches=1  -> flagged
  #1143  (must pass)  matches=0  -> passes
  #1142  (must pass)  matches=0  -> passes

The pattern is rare in this tree (3 occurrences), so the false-positive cost is
low, and the 'per-thread-bound-reviewed' label records a deliberate judgement
rather than blocking.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
justinchuby added a commit that referenced this pull request Aug 17, 2026
#1140)

## What

Our `f16` `MatMul` kept both operands in 16-bit storage and ran the
portable **blocked half GEMM**. That saves bandwidth while the operands
dominate, but the blocked kernel has no tuned microkernel, so once `M`
grows enough for the GEMM to become compute-bound it loses badly.

**ORT does not do this.** Measured on this host at `M=128, K=N=2048`,
ORT spends **14.27 ms on f16** and **14.39 ms on f32** — the same
number, because it widens `f16` to `f32` and reuses the same tuned
SGEMM. We spent **28.2 ms**.

This declines the blocked half GEMM for `f16` above a measured crossover
and lets every caller fall through to the widened-`f32` SGEMM they
already have. All four `try_matmul_half` call sites already had that
fall-through; none needed a new path.

## Result — plugin-level A/B vs real ORT

`bench_matmul_f16_m128`, `M=128, K=N=2048`, **non-constant B**, pinned
to 16 physical cores, 5 interleaved rounds of before/after/ORT **in one
process**, p50 ms. Ratios are **ours/ORT**, so >1 means we lose. Ranges
span **three independent measurements**, one by the reviewer on a
separate build:

| threads | before | after | gain |
|---|---|---|---|
| 1 | **1.97x–1.99x slower** | **1.05x–1.08x** | **1.85x–1.87x faster**
|
| 16 | **3.27x–3.29x slower** | **2.03x–2.28x** | **1.44x–1.63x faster**
|

`M=1` decode cannot reach the gate (`1 < 16`) and measured unchanged:
**0.356 ms after vs 0.353 ms before** at T=1, inside noise (ORT 0.670,
so decode remains a **0.53x win** for us).

### What is *not* fixed

**T=16 is still ~2x off ORT, and this PR does not claim that range.**
The residual is the per-call widen of a non-constant `B` (`4·K·N` = 16
MiB here): serial in our `to_dense_f32_widen`, parallel in ORT. Evidence
— our f16 costs ~2.2 ms more than our own f32 at T=16 but only ~1.6 ms
more at T=1, i.e. the conversion gets *worse* with threads
(fresh-mapping page faults, the mechanism #1133 documented). **That
range stays on the open work list.**

Constant/initializer `B` — what real LLM weights are — was already
handled by `try_packed_half_prefill` and is untouched. This PR fixes the
**dynamic-B** hole.

## How the thresholds were set

Both constants come from `bench_f16_half_vs_widen` (added here), which
times the two *actual* routes. Pinned, median of 5. Ratio is
`half/widen`, so >1 means widening wins.

**`HALF_WIDEN_MIN_M = 16`** — M sweep at `K=N=2048`, two independent
runs:

| M | T=1 | T=16 |
|---|---|---|
| 2 | 0.67x, 0.61x | 0.93x, 0.83x |
| 8 | 1.01x, 1.05x | 1.26x, 1.38x |
| **16** | **1.33x, 1.30x** | **2.14x, 1.85x** |
| 32 | 1.56x, 1.63x | 3.14x, 3.32x |
| 128 | 1.90x, 2.00x | 3.47x, 3.30x |

`M=16` is the first row that wins repeatably at **both** thread counts.
**`M=8` is a tie at T=1 (1.01x/1.05x) and is deliberately left
unclaimed** despite its T=16 win.

**`HALF_WIDEN_MIN_WEIGHT = 256` elements** — weight sweep at the minimum
claimed `M=16`:

| K×N | elements | T=1 | T=16 |
|---|---|---|---|
| 8×8 | 64 | 0.88x (half wins) | 275x |
| **16×16** | **256** | **1.16x** | **229x** |
| 32×32 | 1024 | 1.21x | 35.9x |
| 128×128 | 16384 | 1.48x | 16.3x |
| 256×256 | 65536 | 1.52x | 7.8x |

The huge T=16 ratios are **not** a widening win — they are the blocked
half GEMM forking a parallel region to multiply an 8×8 matrix (0.27
ms!). Widening sidesteps that; fixing the half path's own small-work
threshold is separate and not attempted here.

## Scope — what deliberately does *not* change

- **`bf16`** keeps the blocked kernel: its crossover was never measured,
and widening bf16 is a different (shift-only) operation, so the f16
measurement does not transfer.
- **Non-MLAS backends** keep today's behaviour.
- **`M < 16`**, including all decode, is bit-for-bit unchanged.
- Nothing is ever handed back to the ORT CPU EP — declining here falls
through to a local widened path, never to an error or an unsupported op.

## Tests

Value-only tests cannot detect this gate being mis-wired, so the tests
assert **which route ran** via a thread-local counter. Every fault
injected — by me and independently by the reviewer — is caught:

| injected fault | caught |
|---|---|
| gate disabled (`&& false`) | ✅ |
| gate always true | ✅ |
| off-by-one (`>` for `>=`) | ✅ |
| inverted (`<` for `>=`) | ✅ |
| `bf16` wrongly included | ✅ |
| weight clause removed | ✅ |
| counter increment deleted | ✅ |
| oracle perturbed 1% | ✅ |

Plus `f16_widened_route_matches_an_f64_oracle_across_the_crossover` at
`M = 15, 16, 19` with odd `K=521, N=517` for tails. The route test also
asserts `auto_detect() == Mlas`, so the gate cannot silently become dead
code.

## Review fixes (all from the independent review)

- **MINOR-1** — the oracle tolerance was `2e-3·√K` = **0.0456** against
a measured max error of **1.8e-6**: ~25,000x too loose, so the route
assertion was carrying the test. Measured the real error (5.4e-7 /
1.2e-6 / 1.8e-6 at M=15/16/19) and set a flat **1e-4**. Verified as a
falsifier: a 1% oracle perturbation now fails, and previously did not.
- **MINOR-2 / NIT-1** — `HALF_WIDEN_MIN_WEIGHT` was an unmeasured guess
of `256*1024`, and its doc called `K·N` a byte count when it is an
element count. Measuring it showed the guess was not merely untuned but
**wrong in the expensive direction**: it excluded exactly the range
where the half path is at its worst (7x–275x slower at T=16). Lowered to
the measured **256 elements**.
- The pre-existing half-dispatch test asserted the half GEMM is *always*
selected; one of its shapes (17×130×11) now crosses the gate, so it is
gate-aware. Its determinism and tolerance checks are route-independent
and still cover both routes.

### Second-round review nits (also fixed)

- The gate-aware assertion **duplicated** the threshold values rather
than referencing them, so a retune would silently desync test from
kernel. It now derives its expectation from the constants. That alone
would make the expectation follow *any* retune — which is exactly why
the reviewer's "change the constant" injections were not caught: nothing
pinned the **values**. The route test now asserts both constants equal
the numbers their doc tables were measured at, and points a future
editor at the benchmark. Added three shapes sitting exactly on and one
step below each threshold. Both previously-missed faults are now caught:

  | injected fault | before | now |
  |---|---|---|
| `HALF_WIDEN_MIN_M` 16→17 | ❌ missed | ✅ *M threshold moved off its
measurement* |
| `HALF_WIDEN_MIN_WEIGHT` 256→512 | ❌ missed | ✅ *weight threshold moved
off its measurement* |

- The two three-digit ratios in the weight table are **noise-dominated
in magnitude** — an independent run put them at 47x and 28x, not 275x
and 229x. The direction is robust across runs; the magnitude is not, and
the doc now says so. The attribution was confirmed in source:
`gemm_impl` splits with `par_chunks_mut` whenever `threads > 1` with
**no** small-work guard, so the production half path really does fork
for trivial problems (this was independently verified by the reviewer,
not just asserted).

## Also

Corrects the `CpuBackend::Mlas` doc, which claimed MLAS was "opt-in (not
auto-selected)". `auto_detect` has returned it by default on x86-64 for
some time; read literally, that comment implies this gate never fires.

## Gate

fmt clean; clippy clean **both** feature configs (only the pre-existing
`needless_return` at `matmul_nbits.rs:639`); `-p onnx-runtime-ep-cpu
--lib` debug **and** release, with mlas (1335 pass) and without (1306
pass); real-ORT `plugin_ort_e2e` **51 pass**.

Both the `no-mlas` build break and its dead-code warnings were found and
fixed by running the second feature config locally.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby pushed a commit that referenced this pull request Aug 18, 2026
Second failure of this guard in one day, and this time the hole was the path
filter rather than the regex.

PR #1132 adds crates/onnx-runtime-ep-cpu/src/dispatch_ledger.rs containing

    static LOG: OnceLock<Mutex<Vec<Observation>>> = OnceLock::new();

an unbounded process-lifetime Vec that grows one entry per dispatch decision
whenever NXRT_CPU_DISPATCH_LEDGER=1. Both jobs reported success against that
PR's 3511 added lines because both looked only at src/kernels/** and the file
sits one directory up. The existing regex would have matched the line on sight.

The question the guard asks -- what multiplies this buffer, and who accounted
for it -- was never specific to kernels; only the original grep was.

Falsified before landing:

  scope src/kernels (before)  #1132 -> 0 matches (green, wrongly)
  scope src        (after)    #1132 -> 1 match, exactly the ledger line
  merged 11043a0 8d4401d c62798b faf489a fdeaf77 1f675d2 under the wider
  scope -> 0 new matches beyond the one #1133 already produced

OnceLock<Instant> and OnceLock<()>, also added by #1132, are correctly not
matched: the patterns require an owned growable payload.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
@justinchuby

Copy link
Copy Markdown
Owner Author

Follow-up on the memory bound (commit c5b3e19, pushed to squad/roy-qlinear-staging). Note #1133 is already merged, so this may need a follow-up PR to land.

Measured RSS multiplier (retention, exact atomic counter; 12 MiB accumulator, m=768/k=64/n=4096):

arm retained max WS by PID
T=1, unbounded (pre-fix) 12 MiB ~18 MiB
T=16, unbounded (pre-fix) 192 MiB (16x) ~242 MiB
T=1, bounded (fix) 12 MiB ~20 MiB
T=16, bounded (fix) 120 MiB (capped <128) transient-dominated

The xN multiplier is observable: retention grows 12 -> 192 MiB (x16) with the per-thread-only bound; the process budget caps it at 120 MiB. (PeakWorkingSet64 read 0 post-exit -- the by-name/post-exit pitfall the skill warns about -- so max-polled-by-PID is the valid figure. Peak WS during execute is transient-dominated: every thread allocates its accumulator during compute regardless of parking, so the retention signal, not peak, is what the fix moves.)

New bound + arithmetic: min(128 MiB, 32 MiB x threads) -- flat 128 MiB once >=4 threads park a full buffer, does not grow with vCPUs. Was: 32 MiB x threads = 640 MiB (20-vCPU) / 1 GiB (32) / 4 GiB (128). Per-thread 32 MiB cap retained (oversized single buffers still released -- surviving-case behaviour kept intact).

Governor wiring (#1056): parking now routes through a declinable GovernedAccumulatorBudget (sibling of GovernedWeightCache: no Default, verdict-required, live_bytes() reports the sum actually parked). The ungoverned constant-B MLAS pre-pack (packed_b, merged in ac394fd) is governed too -- process gate + mlas-side live-byte accounting + graph predictor; declined -> dense path, byte-identical. load.rs folds both predictors into resident_f32_cache_bytes and admits/declines with the f32 weight-cache verdict. Declined -> kernel parks nothing, recomputes per call.

Falsification (the test can actually fail): the_parked_accumulator_is_bounded_process_wide_not_per_thread drives 8 threads (a single-instantiation test cannot observe an xN multiplier -- that is how #1100 slipped) and asserts live <= 128 MiB and live >= 2x buffer. Reverting the fix (disable the process cap) -> FAILS at 262,144,000 bytes (8 x 31.25 MiB) over the 128 MiB cap. Restored -> passes. The atomic counter makes the multiplier a number, not a formula.

Gates: cargo test -p onnx-runtime-ep-cpu --lib = 1311 passed / 0 failed; --features mlas --lib = 1339 passed / 0 failed; clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings = clean.

Speed: parking stays admitted-by-default, so the standalone fast path is unchanged (adds one relaxed CAS per park/take). qlinear_phase_report (release+mlas) still exercises the reuse path; products-alloc (the cost parking eliminates) is a real ~12 ms at m=512. Absolute wall deltas on this 10-thread contended box are within run-to-run spread -- exactly the timing-vs-contention caveat in the skill -- so no measurable regression.

justinchuby added a commit that referenced this pull request Aug 18, 2026
…er-thread (follow-up to #1133) (#1151)

Follow-up to #1133 (merged as c62798b with the per-thread bound
unchanged). Closes #1147 -- I do not restate its arithmetic table;
numbers below are measured against it.

## Measured exposure (this box, 20 logical CPUs)

Harness `examples/qlinear_accumulator_rss.rs` parks the accumulator on
every worker of a rayon pool, polls `WorkingSet64` **by PID** at 150 ms,
and reads the governor's own byte ledger
(`qlinear_accumulator_live_bytes()`). 12 MiB per-thread buffer
(m=768,k=64,n=4096) to keep the demo tractable; the ledger scales
exactly with `per_thread x threads`.

| arm | retained (ledger) | idle WS | peak WS |
| --- | --- | --- | --- |
| T=1,  unbounded (pre-fix) | 12 MiB | 17 MiB | 18 MiB |
| **T=16, unbounded (pre-fix)** | **192 MiB (x16)** | **199 MiB** | 234
MiB |
| T=1,  bounded (fix) | 12 MiB | 17 MiB | 19 MiB |
| **T=16, bounded (fix)** | **120 MiB (capped)** | **128 MiB** | 229 MiB
|

**The multiplier is real and observable.** Retention climbs 12 -> 192
MiB (x16) with the per-thread-only bound; **idle/steady-state RSS tracks
it 17 -> 199 MiB** (the extra ~182 MiB is 15 further threads x 12 MiB).
The fix caps it: idle RSS 199 -> **128 MiB** at T=16.

**Honest caveat, stated plainly:** *peak* WS under active decode is ~230
MiB in **both** arms -- it is transient-dominated (every thread
allocates its accumulator *during* `execute` regardless of parking), so
peak-under-load does not isolate the retention. The exposure the fix
removes is the **steady-state** buffer held between calls (idle column),
and on a full 32 MiB per-thread buffer this box's ceiling is 640 MiB,
not the 12 MiB demo -- exactly #1147's figure. `PeakWorkingSet64` reads
0 after exit (the by-name/post-exit pitfall the profiling skill warns
of); the by-PID poll above is the valid figure.

## New bound + arithmetic

`min(128 MiB, 32 MiB x threads)` -- flat 128 MiB once >=4 threads park a
full buffer, and it does **not** grow with vCPUs: 20-vCPU 640->128,
32-vCPU 1 GiB->128, 128-vCPU 4 GiB->128. Per-thread 32 MiB cap kept, so
an oversized single buffer is still released (surviving-case behaviour,
intact). The doc comment now states the product with the arithmetic, not
a per-buffer figure.

## Governor wiring (#1056)

- Accumulator parks through a declinable `GovernedAccumulatorBudget`
(sibling of `GovernedWeightCache`: no `Default`, verdict-required,
`live_bytes()` reports the sum actually parked across all threads).
Declined -> parks nothing, recomputes per call, byte-identical.
- `packed_b` (the `OnceLock<Option<(QgemmPackKey, QgemmPackedB)>>` at
line 154, ungoverned since ac394fd) is now gated: process gate +
mlas-side live-byte accounting (`mlas_sys`) + graph predictor; declined
-> dense path, byte-identical.
- `load.rs` folds both predictors into `resident_f32_cache_bytes` and
admits/declines with the f32 weight-cache verdict.

## Falsifiable multi-thread test

`the_parked_accumulator_is_bounded_process_wide_not_per_thread` drives
**8 threads** via `pool.broadcast` (a single-instantiation test cannot
observe an xN multiplier -- that is how #1100 shipped) and asserts `live
<= 128 MiB` **and** `live >= 2x buffer` (non-vacuity). **Falsified on
this branch:** reverting the fix (disable the process-cap check in
`try_park`) -> test **FAILS at 262,144,000 bytes** (8 x 31.25 MiB) over
the 128 MiB cap. Restored -> **passes**.

## CI guard (070c098 / 73c3df0)

Both `weight-cache-guard.yml` jobs are **green without weakening the
regex**: this PR *bounds* the existing `thread_local!` buffer rather
than adding a new one, so the `RefCell<Vec>` line is unchanged (not an
added line) and `packed_b` is governed via the mlas ledger without a new
`OnceLock<...>` declaration. I did not touch the guard's patterns. (The
guard cannot retroactively flag main's existing line -- that gap is why
this follow-up is filed against main rather than #1133.)

## Gates (this branch, rebased on origin/main 266a6fe)

- `cargo test -p onnx-runtime-ep-cpu --lib` = **1321 passed / 0 failed**
- `cargo test -p onnx-runtime-ep-cpu --features mlas --lib` = **1351
passed / 0 failed**
- `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings` =
**clean**
- `cargo check -p onnx-genai-engine` = clean

## Speed (no regression)

Parking stays **admitted-by-default**, so the standalone fast path is
unchanged -- one relaxed atomic CAS added per park/take.
`qlinear_phase_report` (release+mlas) still exercises the reuse path;
`products-alloc` (the cost parking eliminates) is ~12 ms at m=512.
Absolute wall deltas on this contended box are within run-to-run spread
(the timing-vs-contention caveat in the profiling skill), so no
measurable regression.

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 24, 2026
…indings" (#1931)

## What

Both jobs in `weight-cache-guard.yml` currently report **"no findings"
when their `git diff` fails**. One line each; no regex, path-filter or
label semantics change.

## The defect

Both jobs end the match pipeline with `|| true`. That is genuinely
required — `grep` exits 1 when nothing matches, and nothing matching is
the success case. But `|| true` cannot distinguish *"the matcher found
nothing"* from *"the producer never ran"*. `git diff` with an
unresolvable `BASE_SHA` exits 128, `set -o pipefail` faithfully reports
it, and `|| true` discards it.

Verified, not reasoned about:

```
$ BASE_SHA=deadbeefdeadbeef bash <the ungoverned-weight-cache run block>
No new ungoverned long-lived buffers.
exit=0
```

A green check, on a guard that saw nothing.

## Why this file specifically

This is the **fourth** instance of one shape here, and the first in the
plumbing rather than the pattern. The header already records three:

- a regex knowing only `OnceLock`/`OnceCell`/`LazyLock`, which matched
**zero** of #1133's 459 added lines — the very PR that motivated it;
- a path filter scoped to `src/kernels/**`, which missed #1132's
`src/dispatch_ledger.rs` one directory up;
- the per-instance/per-thread arithmetic error underneath both.

Every one surfaced as **zero findings, green check**. Each time zero
meant *"I could not see"*, not *"there is nothing there"*. A guard whose
silence is indistinguishable from its success has no negative control —
which is precisely what this file exists to warn about, applied to
itself.

## The fix, and why it is safe

Resolve the diff on its own and let its status stand. The property that
makes this safe:

```
no differences  -> 0
differences     -> 0
bad object      -> 128
```

`git diff` is non-zero **only** on error, so "nothing matched" and
"nothing ran" separate cleanly. The guard flags exactly the lines it
flagged before; the only new failure mode is the one that was previously
invisible.

I deliberately did **not** add an "the diff must be non-empty" control,
though the `paths:` filter would seem to justify it.
`pull_request.base.sha` is the base branch tip, so `git diff BASE HEAD`
is two-dot while GitHub's changed-files filter is three-dot — they
diverge when the base already contains the head's content, and that
control would fire on a legitimate PR. Noting it as
considered-and-rejected rather than missed.

## Verification

Both jobs, four regimes, by extracting the **actual `run:` blocks from
the YAML** rather than retyping an approximation:

| regime | expected | result |
|---|---|---|
| (a) broken producer (bogus `BASE_SHA`) | fail | **1** — was `0`, the
defect |
| (b) healthy diff, nothing matching | pass | `0`, *"No new … buffers."*
|
| (c) real match (`OnceLock<Vec<u8>>` / `thread_local RefCell<Vec<u8>>`)
| fail | `1`, message intact |
| (d) correct review label | pass | `0`, bypass intact |

On (d): my first run reported a false regression on the second job
because I passed `weight-cache-reviewed` — the *first* job's label. The
second uses `per-thread-bound-reviewed`. The test was wrong, not the
code; checked before reporting.

## Provenance

Found while checking a claim from @gaff-1 that a stated failure
mechanism (`QEMU_LD_PREFIX`, exit 0 vs 255) was wrong and that *"anyone
who writes 'check for the error string, because the code can't be
trusted' ships a weaker check"*. That specific concern does not apply to
this repo — qemu appears in no workflow here. But the underlying
question does, so I swept every `run:` step for pipelines whose
producer's failure could be swallowed. Five candidates; `diff-guard.yml`
is safe (`set -euo pipefail`, no `|| true`); these two were not.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant