Skip to content

perf(cpu-ep): give QLinearMatMul a native integer GEMM - #1194

Merged
justinchuby merged 9 commits into
mainfrom
squad/resch-native-qgemm
Aug 19, 2026
Merged

justinchuby merged 9 commits into
mainfrom
squad/resch-native-qgemm

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 18, 2026 •

Copy link
Copy Markdown
Owner

The build we ship had no integer GEMM at all

QLinearMatMul in the default build did this per call:

  1. read_quantized widened operand A to a Vec<i32>, then did it again for operand B. For a
    2048x2048 B that is a 16 MiB allocation and fill on every single call, thrown away at the end
    of it.
  2. A scalar rank-1 update walked A row by row, and each row re-streamed the whole of B. At
    m = 128 that is 512 MiB of traffic for 1 GFLOP of work.

The result was 11.8x ORT at m = 1 and 12.1x at m = 128 — the largest single loss on the
x86-64 CPU EP. The performance doc's QLinearMatMul rows never described this build: they were
taken with --features mlas, which is a research build we do not ship. That is now called out in
the doc.

This adds kernels/qgemm_native.rs, a native byte-operand integer GEMM, and points
qlinear_matmul.rs at it. Nothing defers, and nothing falls back.

Two kernels, chosen by m

shape kernel why
m <= 4 (decode) pack-free fused one pass over A means a packed panel of B is never reused, so packing is pure cost. Accumulators stay in registers across a 256-row k block.
m > 4 (prefill) packed KC/2 pairs of NC columns (KC = 512, NC = 256) is 256 KiB of B, which stays in L2 while every row of A sweeps it.

The inner tile is vpmaddwd over NR = 16 columns and MR = 4 rows, k consumed two rows at a
time.

Why vpmaddwd and not vpmaddubsw

MLAS gets 32 MACs from two instructions using vpmaddubsw, which saturates: it needs a
sign-domain translation of B and its intermediate is only nominally exact. vpmaddwd needs four
instructions for the same 32 MACs, but with centred a in [-255, 255] and raw b in
[-128, 255] a product is at most 65025 and a pair sum at most 130050, so it cannot saturate and
cannot overflow. No sign-domain flip, no reasoning about clamped intermediates.

That instruction-count difference is the whole of the residual gap at m = 1. Closing it means
giving up exact integer arithmetic, which is not a trade I am willing to make for a quantized
kernel whose entire value is that it is exact.

Determinism is structural, not tested-in

The kernel computes sum_k (a - za)(b - zb) as sum_k (a - za) * b - zb * sum_k (a - za), with
every accumulation a wrapping i32 add. Wrapping addition is arithmetic mod 2^32, which is
associative and commutative, so any blocking, tiling, column split, row split or thread count
gives bit-identical output — including on overflow, where the wrap itself is reproducible.
wrapping_overflow_is_reordering_invariant and the_thread_count_cannot_change_the_result assert
exactly that, and the SIMD path is checked bit-for-bit against a portable scalar oracle
(the_simd_kernel_is_bit_identical_to_the_portable_loop, and separately for the fused path).

Numbers

Session A/B against plain ORT, K = N = 2048, u8 x u8, ratio is ours / ORT, lower is better,
p50 of 61 iterations. ORT's own timings moved under 1.5% between the two arms at 1 and 4 threads,
which is the control that makes the comparison mean anything.

M threads before after ours before ours after
1 1 12.20x 2.17x 1.402 ms 0.226 ms
128 1 11.90x 1.20x 99.84 ms 9.99 ms
1 4 37.11x 4.03x 1.379 ms 0.170 ms
128 4 14.44x 1.47x 31.03 ms 3.14 ms
1 16 83.12x 35.63x 2.366 ms 1.336 ms
128 16 42.28x 15.05x 51.05 ms 16.55 ms

i8_m1 goes 0.206 ms to 0.049 ms at one thread.

Kernel-level scaling (bench_qgemm_ab, taskset -c 0-15), with the portable scalar arm as the
control:

shape 1t 2t 4t 8t 16t portable 1t
1x2048x2048 0.229 ms 0.136 0.090 0.098 0.166 4.92 ms (21x)
4x2048x2048 0.565 ms 0.311 0.199 0.237 0.345 4.58 ms (8.1x)
128x2048x2048 8.911 ms 4.773 2.755 2.780 1.991 —
128x5120x5120 53.56 ms 27.06 14.35 8.98 11.51 —

The task grid splits rows as well as columns. Columns alone gave only n / NC tasks — eight for
n = 2048 — so a sixteen-worker pool left half of itself spinning; 128x2048x2048 was 2.69 ms at
sixteen threads against 1.62 ms at eight. Splitting columns further would shrink the panel and
re-walk B; splitting rows duplicates only the pack, about a percent of the GEMM it feeds.

Things I measured and rejected

  • Software prefetch of the next B rows (PREFETCH_ROWS = 8): a consistent 8% regression
    with a stable m = 128 control. The hardware prefetcher already has the sequential stream.
  • Permuting inside the fused inner loop: replaced by accumulators held in the permuted order
    with a single vperm2i128 fixup per k-block flush. Saves eight instructions per 32 MACs.

Left open, deliberately

  • Constant-B packed cache. The pack is repeated per call. Caching it would remove it from
    prefill entirely, but any new weight-derived cache has to go through
    kernels/governed_weight_cache.rs to satisfy the "New weight-derived caches must be governed"
    gate. That is a separate PR with its own eviction story, not a rider on this one.
  • The session-level threading gap. At four threads the session takes 0.170 ms while the kernel
    alone does 0.090 ms, and past eight threads both arms get worse. That is the pre-existing
    oversubscription item — it is present before and after this change, so it is not a regression
    here, and it is the next thing I am working on.

Validation

  • cargo test --release -p onnx-runtime-ep-cpu --lib — 1340 passed, 0 failed.
  • Every onnx-runtime-ep-cpu-plugin suite with NXRT_REQUIRE_ORT_TESTS=1, including the 53-test
    plugin_ort_e2e ORT conformance suite with CPU fallback disabled.
  • cargo clippy -p onnx-runtime-ep-cpu --all-targets clean, cargo fmt --all --check clean.
  • cargo check -p onnx-runtime-ep-cpu --lib --features mlas — the research build still compiles.
  • Reviewed by Claude Opus 4.8 against the memory-safety, lane-semantics, determinism and
    edge-extent claims above; no blockers, two documentation fixes applied.

Refreshed against main (2026-08-18)

The branch was behind main and its red CI wall came from that, not from this
change: crates/onnx-runtime-session/src/executor/mod.rs:175 failed -D dead-code
on current stable, fixed on main by ca32b3adf (#1239) after this branch forked.
origin/main (c55a3fab3) is merged in — no rebase, no force-push.

One conflict, in docs/performance/CPU_MATMUL_ASSIGNMENT.md, resolved as a
union: this branch's #### 3b (the native integer GEMM) and main's
### 4 (the f32 M = 1 GEMV becoming the default, #1091) were both new sections
appended after 3a. Both are kept, in that order. Taking either side would have
silently deleted the other's record.

Revalidated on the merge commit, AVX2/FMA host, no AVX-512:

  • cargo test --release -p onnx-runtime-ep-cpu --lib — 1424 passed, 0 failed,
    18 ignored, including qgemm_i32_matches_the_integer_oracle_for_every_signedness,
    the_simd_kernel_is_bit_identical_to_the_portable_loop,
    wrapping_overflow_is_reordering_invariant and
    the_thread_count_cannot_change_the_result.

The measurements in this PR were taken before the merge; nothing in the merged
range touches qgemm_native.rs, qlinear_matmul.rs, or the CPU threadpool, so
they stand as recorded. The main change that did land in this range (#1091's
f32 M = 1 GEMV default) is on a different kernel family and is documented in
the section-4 text kept above.


Refreshed again against main @ 6a855d5e0, and a real branch bug found

main moved again while this was queued (#1346/#1352/#1361 on the quality lane,
#1154/#1232/#1238 on the CPU side). Merged in normally — no rebase — and
revalidated.

The revalidation caught something the earlier ones had not. Running
-p onnx-runtime-ep-cpu --lib in a debug profile rather than
--release fails:

kernels::qgemm_native::tests::degenerate_extents_do_nothing
  assertion `left == right` failed
  left: 0
 right: 4

degenerate_extents_do_nothing called qgemm with an empty b_zero_points
and n == 4. qgemm opens with debug_assert_eq!(b_zero_points.len(), n),
so that call is not one the function accepts — the test was exercising the
m == 0 early return through an argument list the contract forbids. It passed
every previous run here only because debug_assert compiles out under
--release, which is how I had been validating this branch locally. A debug
test profile fails it, and this is branch-caused: qgemm_native.rs is new in
this PR.

Fixed in 9ca99e538 by sizing the test's zero points to m and n, not by
weakening the assertion — the assertion states the contract the kernel's
indexing depends on, and a caller whose m is zero still has n columns and
still knows their zero points.

-p onnx-runtime-ep-cpu --lib, debug profile: 1440 passed, 0 failed (was
1439 passed, 1 failed).

This is the second time on this stack that the profile a test runs under
decided whether it caught anything. Worth remembering: --release silently
disables every debug_assert in the crate under test, so a local
cargo test --release is not a substitute for what CI runs.


Re-validated on latest main (e0aedd0fa), 2026-08-19

Latest main merged in normally (no rebase). Full re-measurement, 1 thread
pinned, K = N = 2048, 61 iters / 10 warmup, 2 reps, ours_p50 / ort_p50:

case main ours main ratio this PR ours this PR ratio speedup
bench_qlinear_u8_m1 1.418 / 1.435 ms 11.83x / 12.51x 0.121 / 0.123 ms 1.16x / 1.18x 11.7x
bench_qlinear_u8_m128 29.55 / 29.56 ms 3.57x / 3.57x 3.055 / 3.065 ms 0.372x / 0.373x 9.6x
bench_qlinear_i8_m1 1.516 / 1.497 ms 0.215x / 0.212x 0.209 / 0.212 ms 0.030x / 0.030x 7.2x

ORT-side drift between the two arms was 0.7% at m = 128 and 0.0% on i8,
which is the control that makes the comparison mean anything.

At m = 128 we are now 2.7x faster than ORT outright, and m = 1 closes
from 11.8x to 1.16x. These are better than the numbers originally posted above
because the dispatch work in #1077 landed in between.

Review fixes (987aa0c5c)

An independent review found no blockers but two things worth fixing:

  1. aarch64 built with 5 warnings — NR/MR/NC/KC/FUSED_KC are read
    only by the x86 kernels, so every non-x86 target warned on all five. CI
    builds with -D warnings, so this was a branch-caused CI failure waiting to
    happen; the local x86 clippy run could never have caught it. Now #[cfg]-gated
    alongside the code that uses them: 0 warnings on both x86-64 and aarch64.
  2. The fused-parallel path had no end-to-end coverage. Every m <= 4 shape
    in qlinear_matmul_reordered_accumulation_is_bit_identical sat below
    PARALLEL_MIN_WORK, so the pack-free kernel's column split was only ever
    checked at the kernel level, never through requantize_rows. Added
    (4, 1029, 1100), which forks both.

The review independently re-derived the register-shuffle math in numpy
(cvtep*_epi16, permute4x64_epi64(0xD8), unpacklo/hi_epi16, madd_epi16,
permute2x128) against a plain per-column dot product over 2000 tiles with
extreme values — 0 mismatches — and confirmed the vpmaddwd non-saturation
bound for all four operand combos, the wrapping-add determinism claim, and the
absence of out-of-bounds access in every tail path.

Validation on the merged base

  • cargo fmt clean; cargo clippy --all-targets -D warnings clean
  • 1553 onnx-runtime-ep-cpu tests, debug profile (so debug_asserts are live)
  • 55 plugin conformance tests (NXRT_REQUIRE_ORT_TESTS=1, release)
  • every_assigned_node_is_also_executed_by_this_ep and
    every_fixture_loads_with_cpu_fallback_disabled green — nothing defers,
    nothing falls back to the ORT CPU EP
  • aarch64-unknown-linux-gnu cross-check clean, 0 warnings

justinchuby and others added 3 commits August 18, 2026 03:55
The default (non-MLAS) build had no integer GEMM. `execute` widened both
operands to `Vec<i32>` on every call -- 16 MiB for a 2048x2048 weight --
and then ran a scalar rank-1 update that re-streamed all of `B` once per
row of `A`. Section 3 of the matmul doc calls per-call packing "fixed",
but that only ever applied to `--features mlas`; the build we ship was
paying 12x ORT at both decode and prefill, the largest single loss in the
matmul family.

Add `kernels::qgemm_native`: the same arithmetic on the operand bytes,
with an AVX2 kernel behind a portable reference.

`vpmaddwd` is the instruction that fits: it takes eight `i16` pairs and
sums each pair into an `i32`, sixteen multiply-accumulates per
instruction. It is exact here rather than merely close -- a centred `a`
is in [-255, 255] and a raw `b` in [-128, 255], so a pair sum cannot
exceed 130050 -- and unlike `vpmaddubsw` it does not saturate, so no
operand has to be translated into another sign domain first.

Two kernels, chosen by `m`:

  - `m > MR`: `B` is packed into 16-column, k-pair-interleaved tiles in a
    256 KiB L2-resident panel that every row block re-reads. The pack is
    SIMD and costs about 1% of the GEMM it feeds.

  - `m <= MR` (decode): no pack at all. At one row block a panel would be
    written once and read once, and writing `2 * k * n` bytes to serve a
    GEMV that reads `k * n` is most of the call. The interleave happens
    in registers instead, accumulators stay in registers across a `k`
    block, and the column permutation `vpunpcklwd` implies is undone once
    per block rather than once per iteration.

Every accumulation is wrapping `i32`, which is exactly arithmetic mod
2^32 and therefore associative and commutative, so neither the blocking
nor the thread count can change an output bit -- on overflow included.
The tests assert that directly against an integer oracle rather than
assuming it.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
A column block owns a packed panel, so splitting the columns further to
reach every worker shrinks the panel and re-walks `B`. Splitting the rows
instead duplicates only the pack, which is around a percent of the GEMM
it feeds.

Left at one row block, an `n` of 2048 offers eight tasks, so a
sixteen-worker pool leaves half of itself idle and spinning: 128x2048x2048
measured 2.69 ms at sixteen threads against 1.62 ms at eight. With the row
split it is 1.99 ms, and the session A/B at four threads goes from 1.95x
to 1.47x.

Also document the native path in the matmul performance file, including
the fact that the `QLinearMatMul` rows in the ranges table are `--features
mlas` measurements that never described the build we ship.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…D grid

The note still said a task owns whole rows of its column block, which was
true before the row split. Describe the rectangle it actually owns, and
derive the tile width from the block extent so a `block_width` that is
not a multiple of `NR` cannot make the pack and the store disagree.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby marked this pull request as ready for review August 18, 2026 04:17
@justinchuby
justinchuby enabled auto-merge (squash) August 18, 2026 04:17
@github-actions

github-actions Bot commented Aug 18, 2026 •

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 39.84 µs 77.00 µs +93.3%
🔴 block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 55.00 µs 91.79 µs +66.9%
🔴 matmul/large_generic_f32_threads=1/32x1024x1024 9.18 ms 14.59 ms +58.9%
🔴 matmul/large_generic_f32_threads=8/32x1024x1024 3.83 ms 6.02 ms +57.2%
🔴 matmul/medium_generic_bf16_threads=8/32x512x512 486.83 µs 722.59 µs +48.4%
🔴 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 345.54 µs 512.81 µs +48.4%
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 56.36 µs 78.59 µs +39.4%
⚠️ matmul/medium_generic_f16_threads=1/32x512x512 33.89 µs 41.97 µs +23.8%
✅ matmul/small_generic_f32_threads=1/1x256x256 47.59 µs 51.92 µs +9.1%
✅ tokenization/decode_tokens_per_second 5.94 ms 6.47 ms +8.9%
✅ matmul/medium_generic_f16_threads=8/32x512x512 35.35 µs 38.00 µs +7.5%
✅ gather/large_f16_threads=1-internal/131072 16.41 µs 16.96 µs +3.3%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 576.66 µs 595.75 µs +3.3%
✅ sampling_latency/top_k_per_token 48.86 µs 49.87 µs +2.1%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 97.33 µs 97.56 µs +0.2%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.61 ms 1.61 ms +0.2%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 85.40 µs 85.21 µs -0.2%
✅ sampling_latency/top_p_per_token 361.40 µs 358.61 µs -0.8%
✅ tokenization/encode_tokens_per_second 352.59 µs 349.83 µs -0.8%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.13 ms 2.11 ms -1.3%
✅ sampling_latency/greedy_per_token 3.06 µs 3.02 µs -1.6%
✅ matmul/small_generic_bf16_threads=8/1x256x256 39.83 µs 38.98 µs -2.1%
✅ gather/small_bf16_threads=1-internal/4096 525.2 ns 504.7 ns -3.9%
✅ matmul/small_generic_bf16_threads=1/1x256x256 36.43 µs 34.72 µs -4.7%
✅ kv_cache/alloc_dealloc_pages 39.14 µs 37.18 µs -5.0%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.07 ms 1.01 ms -5.7%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.79 ms 1.67 ms -6.7%
✅ gather/medium_bf16_threads=1-internal/32768 2.81 µs 2.61 µs -6.9%
✅ add/small_f16_threads=1-internal/1024 508.9 ns 467.8 ns -8.1%
✅ add/small_bf16_threads=1-internal/1024 503.5 ns 455.7 ns -9.5%
✅ qwen3_sampling_processors/top_k_top_p_fast 692.39 µs 626.04 µs -9.6%
✅ gather/large_bf16_threads=1-internal/131072 18.88 µs 17.06 µs -9.6%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 558.19 µs 489.23 µs -12.4%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.78 ms 2.42 ms -12.8%
✅ sampling_latency/min_p_per_token 224.49 µs 193.63 µs -13.7%
✅ matmul/small_generic_f16_threads=1/1x256x256 41.11 µs 35.42 µs -13.8%
🟢 add/medium_f16_threads=1-internal/262144 131.40 µs 111.59 µs -15.1%
🟢 qwen3_sampling_processors/top_k_top_p_full_sort_baseline 6.17 ms 5.20 ms -15.8%
🟢 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 861.05 µs 723.09 µs -16.0%
🟢 gather/medium_f16_threads=1-internal/32768 3.02 µs 2.52 µs -16.7%
🟢 matmul/small_generic_f16_threads=8/1x256x256 42.59 µs 35.33 µs -17.1%
🟢 add/medium_f32_threads=1-internal/262144 33.03 µs 26.99 µs -18.3%
🟢 add/medium_bf16_threads=1-internal/262144 135.65 µs 110.62 µs -18.4%
🟢 qwen3_sampling_processors/top_k_full_sort_baseline 2.43 ms 1.97 ms -19.1%
🟢 gather/small_f16_threads=1-internal/4096 662.0 ns 523.0 ns -21.0%
🟢 reduce_mean/small_f32_threads=1-internal/4096 19.98 µs 15.77 µs -21.1%
🟢 logit_processing/seven_processor_chain_per_step 373.27 µs 294.42 µs -21.1%
🟢 gather/small_f32_threads=1-internal/4096 918.1 ns 722.6 ns -21.3%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 327.18 µs 252.92 µs -22.7%
🟢 qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 4.27 ms 3.28 ms -23.3%
🟢 gather/medium_f32_threads=1-internal/32768 5.19 µs 3.96 µs -23.7%
🟢 grammar_masking/llguidance_compute_mask/32 97.01 µs 71.98 µs -25.8%
🟢 add/large_bf16_threads=1-internal/4194304 2.43 ms 1.76 ms -27.4%
🟢 add/small_f32_threads=1-internal/1024 294.1 ns 210.0 ns -28.6%
🟢 gather/large_f32_threads=1-internal/131072 54.39 µs 38.00 µs -30.1%
🟢 add/large_f16_threads=1-internal/4194304 2.72 ms 1.78 ms -34.6%
🟢 qwen3_sampling_processors/top_k_partial_selection 211.00 µs 135.02 µs -36.0%
🟢 add/large_f32_threads=1-internal/4194304 1.16 ms 713.05 µs -38.3%
🟢 matmul/small_generic_f32_threads=8/1x256x256 83.27 µs 39.30 µs -52.8%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 4.84 3.63 6.72 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby and others added 3 commits August 18, 2026 23:36
Union-resolved docs/performance/CPU_MATMUL_ASSIGNMENT.md: kept both the
branch's 3b (native QGEMM) and main's section 4 (f32 M=1 GEMV default).
`degenerate_extents_do_nothing` called `qgemm` with an empty
`b_zero_points` and `n == 4`. `qgemm` asserts one zero point per column
on entry, so that call is not one the function accepts. It passed only
because `debug_assert` compiles out under `--release`, which is how this
branch was being validated locally; a debug test profile fails it with
`left: 0, right: 4`.

Fixed the call rather than the assertion. The assertion states the
contract the kernel's indexing depends on, and a caller whose `m` is zero
still has `n` columns and still knows their zero points.

`-p onnx-runtime-ep-cpu --lib` in a **debug** profile: 1440 passed, 0
failed (was 1439 passed, 1 failed).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 19, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 89.58611% with 78 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.87%. Comparing base (4a9f4ec) to head (987aa0c).
⚠️ Report is 81 commits behind head on main.

Files with missing lines Patch % Lines
...es/onnx-runtime-ep-cpu/src/kernels/qgemm_native.rs 89.59% 75 Missing ⚠️
.../onnx-runtime-ep-cpu/src/kernels/qlinear_matmul.rs 89.28% 1 Missing and 2 partials ⚠️
Additional details and impacted files

Impacted file tree graph

@@             Coverage Diff             @@
##             main    #1194       +/-   ##
===========================================
- Coverage   82.10%   80.87%    -1.23%     
===========================================
  Files          12      378      +366     
  Lines        5471   167778   +162307     
  Branches     5471   167778   +162307     
===========================================
+ Hits         4492   135696   +131204     
- Misses        780    27234    +26454     
- Partials      199     4848     +4649     
Flag Coverage Δ
cli-ort-linux 82.60% <ø> (?)
cli-ort-windows 82.10% <ø> (ø)
offline 80.81% <89.58%> (?)

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-runtime-ep-cpu/src/kernels/mod.rs 95.04% <ø> (ø)
.../onnx-runtime-ep-cpu/src/kernels/qlinear_matmul.rs 84.30% <89.28%> (ø)
...es/onnx-runtime-ep-cpu/src/kernels/qgemm_native.rs 89.59% <89.59%> (ø)

... and 363 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

justinchuby and others added 2 commits August 19, 2026 22:58
… to end

The AVX2 tile constants (NR/MR/NC/KC/FUSED_KC) are only read by the x86
kernels, so an aarch64 build warned on all five and would have failed CI's
-D warnings. Gate them to x86 alongside the code that uses them.

Also extend qlinear_matmul_reordered_accumulation_is_bit_identical with a
(4, 1029, 1100) shape. Every existing m <= 4 case sat below PARALLEL_MIN_WORK,
so the pack-free kernel's column split had never been exercised end to end
through requantize_rows -- only at the kernel level.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby merged commit 06b372c into main Aug 19, 2026
6 checks passed
@justinchuby
justinchuby deleted the squad/resch-native-qgemm branch August 19, 2026 23:28
justinchuby added a commit that referenced this pull request Aug 20, 2026
…es mlas`

The guard asserted a *total* borrow count of exactly `before + 1`. Once
QLinearMatMul got its native integer GEMM (#1194), `B` is taken through
`dense_bytes` as well, so the count is 2 and the test fails on `main`
whenever the `mlas` feature is on -- while the property it guards (the
sign-flip route never writes through to the caller's `A`) still holds.

Assert non-vacuity on `A` itself instead: `A` is in the borrow domain, and
the call borrowed something. How many *other* operands a route borrows is
a routing detail this test does not fix.

The failure escapes CI because the mlas lanes are name-filtered.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 20, 2026
…es mlas` (#1523)

## What


`kernels::qlinear_matmul::tests::the_sign_flip_route_never_writes_through_to_the_callers_input`
fails on plain `origin/main` whenever the (non-default) `mlas` feature
is on.

```
$ cargo test --release -p onnx-runtime-ep-cpu --features mlas --lib the_sign_flip_route_never_writes_through
assertion `left == right` failed: the flip route no longer borrows A, so this test proves nothing
  left: 2
 right: 1
```

## Why

The test's non-vacuity guard counted `BORROWED_INPUT_CALLS` across *all*
operands and asserted exactly `before + 1`. When QLinearMatMul got its
native integer GEMM (#1194), `B` started going through `dense_bytes`
too, so the count is now 2.

The property the test actually guards — the MLAS sign-flip route goes
through `Cow::to_mut` and so never writes through to the caller's `A` —
**still holds**. The second assertion (`a.bytes == untouched`) passes.
Only the guard is wrong.

## Fix

Assert non-vacuity on `A` itself:

- `dense_bytes(&a.view())` returns `Cow::Borrowed` — `A` is in the
borrow domain, which is the precondition that makes the property
non-trivial;
- the call borrowed *something* (`> borrows_before`).

How many *other* operands a route borrows is a routing detail this test
does not fix, so it no longer asserts on it.

## Why CI missed it

The `mlas` lanes are name-filtered, so this test is not selected there.
Found while validating #1365, where an unrelated `--features mlas` route
assertion had to be cfg-gated; I said in that PR I would file this
separately.

## Validation

```
cargo test  --release -p onnx-runtime-ep-cpu --features mlas --lib qlinear_matmul   # 32 passed, 0 failed
cargo test  --release -p onnx-runtime-ep-cpu                --lib qlinear_matmul   # 17 passed, 0 failed
cargo clippy --release -p onnx-runtime-ep-cpu --features mlas --lib --tests -- -D warnings   # clean
cargo fmt --all --check
```

Test-only change: no production code touched, no `mlas` symbol reaches
the default artifact.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 20, 2026
`widen16(signed: bool, ..)` -- the u8 -> i16 widen every byte of `B` passes
through in `qgemm_native` -- took signedness as a runtime argument and branched
on it in the innermost loop, once per 16 bytes. `Operand::signed` is fixed for
the whole call, but the `#[target_feature]` boundary stops the compiler
hoisting the test out.

It is now a `const SIGNED: bool` on `widen16`, `fused_strip`,
`accumulate_fused` and `pack_panel`, with the single runtime `match` moved out
to the block dispatcher that already matched on `m`. No arithmetic changed:
accumulation is still wrapping `i32`, so the output is bit-identical, which the
existing both-sign-domain oracle asserts.

Kernel A/B, two prebuilt test binaries alternated over three repetitions with
the harness's `portable` drift control steady to 1.6%: 1.13x at 1x3584x3584,
1.14x at 1x1024x3072 and 1.03x on the packed 128x3584x3584 path, with the two
m=1 ranges non-overlapping between arms.

End-to-end the effect is below what this host can resolve. A first attempt
across two separately built bench binaries read 1.43x and is withdrawn --
a 1.13x kernel cannot yield a 1.43x call. A null control (one binary against
itself) puts the paired end-to-end floor at +/-10%, and re-running each binary
against its own ORT reference gives medians of 2.43x (base) and 2.58x (new).
The change is kept for the mechanism, not for an end-to-end number: the branch
is provably loop-invariant and the output is bit-identical.

This was found by the first `QLinearMatMul` A/B this repository could run.
`scripts/ort_ab/` had no generator for the op, so the ledger's `QLinearMatMul`
rows -- `--features mlas` numbers under a caveat claiming 11.8x-12.0x on the
default build -- had never been re-measured after #1194 landed the native
integer GEMM. `gen_qlinear.py` closes that hole, with cells straddling both of
the kernel's own dispatch gates. The corrected picture: 1.13x at u8 M=128,
0.11x at i8 M=1 (a 9.4x win, because ORT's own i8 path is 35x slower than its
u8 path), and a real loss of ~2.4x-3.2x confined to u8 M=1, which this change
does not close.

Also corrects two stale claims the measurements invalidated: the ledger's
11.8x/12.0x `QLinearMatMul` caveat, and `scripts/ort_ab/README.md`'s
instruction to set `ONNX_GENAI_CPU_MM_HALF_GEBP=0` to reach the half decode
GEMV, which #1613 made unnecessary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 20, 2026
…#1616)

## What

`widen16(signed: bool, ..)` — the `u8 -> i16` widen every byte of `B`
passes through in
`qgemm_native` (the native integer GEMM behind `QLinearMatMul`) — took
signedness as a **runtime
argument** and branched on it in the innermost loop, once per 16 bytes
of `B`. `Operand::signed`
comes from the input dtype and is fixed for the whole call, but the
`#[target_feature]` boundary
stops the compiler hoisting the test out.

It is now a `const SIGNED: bool` on `widen16`, `fused_strip`,
`accumulate_fused` and `pack_panel`,
with the single runtime `match` moved out to the block dispatcher that
already matched on `m`.

**No arithmetic changed.** Accumulation is still wrapping `i32`, so the
output is bit-identical —
asserted by the existing `check(m, k, n, a_signed, b_signed)` oracle
across both sign domains.

## Measurement

Kernel A/B, two prebuilt test binaries alternated over three
repetitions, `bench_qgemm_ab`'s
`portable` arm as a drift control (steady to 1.6% across all six runs):

| shape | base | new | ratio |
| --- | ---: | ---: | ---: |
| 1x3584x3584 | 19.97 GMACS | 22.52 | **1.13x** |
| 1x1024x3072 | 17.52 | 19.89 | **1.14x** |
| 128x3584x3584 | 61.92 | 64.00 | 1.03x |
| *`portable` control* | 3.82 | 3.82 | *1.00x* |

The two `m = 1` ranges do not overlap between arms.

**End to end the effect is below what this host can resolve, and an
earlier claim is retracted.**
A first A/B across two separately built `bench_generic` binaries read
1.43x; that is withdrawn,
because a 1.13x kernel cannot produce a 1.43x call (`1 / (f/R_k + 1 - f)
<= R_k`). A null control
(one binary against itself) puts the paired end-to-end floor at ±10%,
and re-running each binary
against its own ORT reference gives medians of 2.43x (base) and 2.58x
(new). The change is kept for
the mechanism — provably loop-invariant branch, bit-identical output,
clean separation at the level
where it acts — not for an end-to-end number.

## Why this was found now

`scripts/ort_ab/` had **no `QLinearMatMul` generator**, so the ledger's
`QLinearMatMul` rows
(`--features mlas`, under a caveat claiming 11.8x–12.0x on the default
build) had never been
re-measured after #1194 landed the native integer GEMM. `gen_qlinear.py`
closes that hole, with
cells straddling both of the kernel's own dispatch gates
(`PARALLEL_MIN_WORK` and the `m <= MR`
fused/packed split).

The corrected picture on the shipped build, one thread, parity `PASS`
everywhere: **1.13x at
u8 M=128**, **0.11x at i8 M=1** (a 9.4x win — ORT's own i8 path is 35x
slower than its u8 path),
and a real loss of **2.4x–3.2x confined to u8 M=1**, which this change
does not close.

Also localised, for the follow-ups: at `m = 1` a 1 MB L2-resident weight
runs at the same
19.5 GB/s as a 12.85 MB one and every aspect ratio at equal footprint
lands within 17–20 GB/s, so
the kernel is **instruction-bound, not memory-bound**, up to L3.

## Also in here

Two stale claims the measurements invalidated: the ledger's 11.8x/12.0x
`QLinearMatMul` caveat,
and `scripts/ort_ab/README.md`'s instruction to set
`ONNX_GENAI_CPU_MM_HALF_GEBP=0` to reach the
half decode GEMV, which #1613 made unnecessary.

## Still open (recorded, not fixed here)

- The **~2.4x u8 M=1 gap** itself.
- **Parallel scaling at m=1**: 8 threads buy 1.3x where ORT gets 5.8x.
The fused path splits
columns, handing every worker a page-crossing strided walk of `B`; at
m=1 there is no `B` reuse
to protect, so a `k` split with private accumulators is the shape that
streams.
- The single-thread residual is an instruction budget: exact full-range
8-bit needs `vpmaddwd`;
`vpmaddubsw` would halve the uops but saturates unless one operand stays
within ±64, which is
  why ORT's quantizer ships `reduce_range` for non-VNNI AVX2.

Full record:
`docs/benchmarks/2026-08-21-qlinearmatmul-m1-signedness.md`, ledger
section 16.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant