Skip to content

cpu: vectorise x86 borrowed int4 GEMV (AVX2/AVX-512) for #994 - #1021

Merged
justinchuby merged 1 commit into
mainfrom
squad/994-x86-borrowed-int4-simd
Aug 16, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/994-x86-borrowed-int4-simd

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Vectorise the x86 borrowed int4 GEMV (refs #994, #979)

The x86 symmetric/asymmetric int4 borrowed decode path (unified by #979) was a
scalar unpack-convert-FMA loop; only aarch64 had a vectorised route
(borrowed_affine_int4_matmul_m1_neon_dot / affine_int4_block32_dot_neon). At the
14B lm_head shape it measured ~0.78 GB/s, ~1.6% of this box's 49.28 GB/s STREAM
ceiling
(per the #1013 roofline_gemv probe).

What changed

  • New runtime-dispatched f32 SIMD block dot for the borrowed path, reusing the
    existing selected_dot_kernel() / DotKernel seam (no new mechanism, no
    compile-time target-feature gate):
    • borrowed_int4_block_dot_avx2 — AVX2+FMA, serves Avx2 and AvxVnni hosts
      (the borrowed path is f32, so integer VNNI does not apply).
    • borrowed_int4_block_dot_avx512 — AVX-512F, serves Avx512Vnni hosts.
  • Both unpack the 4-bit nibbles into the scalar loop's natural
    w[2i]=lo(byte i), w[2i+1]=hi(byte i) order, widen to f32, and multiply-accumulate
    the activation per 32-lane tile. The per-block scale and the zero-point correction
    stay in the caller, so symmetric (zero_points=None, implicit midpoint 8) and
    asymmetric both work unchanged. Blocks larger than 32 (e.g. block-128) vectorise
    via the 32-lane chunk loop into a single accumulator.
  • Nothing dequantised outlives the call → the Symmetric int4 has no zero-copy CPU kernel, so it pays 8x RAM: adding a constant zero_points=8 cuts footprint 6.1x with byte-identical output #979 zero-copy footprint is preserved.
  • ONNX_GENAI_CPU_DISABLE_INT4_SIMD=1 is a documented A/B escape hatch to the scalar
    reference (same-binary A/B, OnceLock-resolved, mirrors arm64_int4_direct_enabled).
  • aarch64 is untouched except that the shared let _ = dot_kernel; unused-var
    guard is now #[cfg(not(any(aarch64, x86_64)))] (aarch64 already used dot_kernel;
    behaviour there is byte-for-byte the same).

1. Roofline (roofline_gemv, K=5120 N=152064 block_size=32 bits=4, borrowed symmetric int4)

Weight-only GB/s (criterion input). Box: shared i7-13800H (14C/20T, AVX2+AVX-VNNI,
no AVX-512 — so the AVX2 kernel is what runs here; the AVX-512 kernel is compiled
but unexercised on this host). iters=10 repeats=8; two other squad agents active, so I
report median and best (quietest) rather than a single sample.

threads before med / best after med / best after % of 49.28 speedup (best)
1 0.390 / 0.449 1.048 / 1.661 3.4% 3.7×
2 0.604 / 0.700 1.946 / 2.217 4.5% 3.2×
4 0.505 / 0.882 3.682 / 3.944 8.0% 4.5×
8 0.646 / 0.777 4.551 / 5.219 10.6% 6.7×
16 0.768 / 1.214 6.176 / 6.519 13.2% 5.4×
20 1.021 / 1.186 6.711 / 8.133 16.5% 6.9×

Headline (20 threads, quiet): ~8.1 GB/s best / ~6.7 GB/s median, ~16.5% of the
49.28 GB/s ceiling
, up from ~0.78 GB/s. ~6.7–6.9× at the useful thread counts.

Honest-negative note: the scalar loop was a real bottleneck (~7× confirms it), but
we are still ~6× below roofline (16.5%, not ~100%). The residual gap is now
elsewhere — the +25% scales traffic the criterion omits, the strided per-row weight
reads, and the per-execute SPMD barrier — not the nibble unpack. There is more CPU
headroom, but it is no longer in this loop.

2. Byte-identical output

96 greedy tokens, native CPU EP, SIMD on vs ONNX_GENAI_CPU_DISABLE_INT4_SIMD=1:

  • qwen05b-q4 (symmetric, borrowed path): token IDs byte-identical (all 96).
  • qwen2.5-0.5b-q4_0-mobius (asymmetric, same path): token IDs byte-identical (all 96).

This is not a bit-identical f32 transform: the horizontal reduction + FMA reorder
additions vs the scalar loop, so logits differ by a few ULP. That divergence is below
the argmax margin here — tokens are byte-identical on both models — but I am stating it
rather than claiming bit-identity. The unit test locks the block dot to the scalar
reference within |Δ| ≤ |ref|·1e-5 + 1e-4.

3. Footprint unchanged

Peak working set on qwen05b-q4: 383.0 MiB (SIMD) vs 372.2 MiB (scalar); asymmetric
mobius 380.0 vs 379.5 MiB. No f32 expansion (that would be GBs); the borrowed tile does
not outlive the call. Well under the ~434 MB envelope.

4. aarch64

No functional change. The only aarch64-visible edit is the let _ = dot_kernel; cfg
guard (aarch64 already consumed dot_kernel). Cannot be run here; called out explicitly.

Gates (verbatim)

  • cargo fmt -p onnx-runtime-ep-cpu — clean
  • cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings — clean
  • cargo test -p onnx-runtime-ep-cpu:
    • unittests src\lib.rs: test result: ok. 1078 passed; 0 failed; 10 ignored
    • bf16_conformance: ok. 3 passed; 0 failed
    • kernel_numeric_regression: ok. 10 passed; 0 failed
    • mha_ort_parity: ok. 1 passed; 0 failed
    • msft_attention_ort_parity: ok. 1 passed; 0 failed
    • qwen35_ort_parity: ok. 5 passed; 0 failed
    • shared_allocator: ok. 6 passed; 0 failed
  • cargo test -p onnx-genai-engine --lib: test result: ok. 385 passed; 0 failed; 1 ignored

(baseline lib was 1077; +1 is the new borrowed_int4_block_dot_x86_matches_scalar_reference test.)

#994 criterion re-evaluation

Posted on #994. Short version, with the post-fix number:

  • move to GPU: 389 MB / 11.74 GB/s (H2D pinned) = ~33 ms
  • compute on CPU: 389 MB / ~6.7–8.1 GB/s = ~48–58 ms (20 threads, median→best)

GPU H2D still edges CPU-compute in absolute per-token latency, so the placement
argument for lm_head does not come back — but the margin collapsed from ~15×
(the old ~500 ms) to ~1.5×, moving it from "clearly dead" into narrow/ambiguous
territory where the VRAM-savings tradeoff is now defensible. Embedding gather is
unaffected (arithmetic intensity ~0).

The x86 symmetric/asymmetric int4 borrowed decode path (introduced in #979)
was a scalar unpack-convert-FMA loop; only aarch64 had a vectorised route.
At the 14B lm_head shape it hit ~0.78 GB/s, ~1.6% of this box's 49.28 GB/s
STREAM ceiling.

Add a runtime-dispatched f32 SIMD block dot for the borrowed path, reusing
the existing selected_dot_kernel()/DotKernel seam:
- AVX2+FMA kernel (covers Avx2 and AvxVnni hosts; the borrowed path is f32,
  so integer VNNI does not apply).
- AVX-512F kernel for Avx512Vnni hosts.
Both unpack the 4-bit nibbles, widen to f32, and multiply-accumulate the
activation per 32-lane tile; the per-block scale and zero-point correction
stay in the caller, so symmetric (implicit midpoint 8) and asymmetric both
work unchanged. Nothing dequantised outlives the call, preserving the #979
zero-copy footprint.

The horizontal reduction reorders additions vs the scalar loop, so f32
logits differ by a few ULP; it is not a bit-identical f32 transform. Token
IDs stay byte-identical (argmax stable) on qwen05b-q4 (symmetric) and
qwen2.5-0.5b-q4_0-mobius (asymmetric) for 96 greedy tokens.

ONNX_GENAI_CPU_DISABLE_INT4_SIMD=1 is a documented A/B escape hatch to the
scalar reference. aarch64 is untouched.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 15, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 73.48485% with 35 lines in your changes missing coverage. Please review.
✅ Project coverage is 78.79%. Comparing base (6e3dd30) to head (c65e3bd).
⚠️ Report is 4 commits behind head on main.

Files with missing lines Patch % Lines
...es/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs 73.48% 35 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1021      +/-   ##
==========================================
- Coverage   78.79%   78.79%   -0.01%     
==========================================
  Files         364      364              
  Lines      146900   147032     +132     
  Branches   146900   147032     +132     
==========================================
+ Hits       115754   115855     +101     
- Misses      26529    26560      +31     
  Partials     4617     4617              
Flag Coverage Δ
cli-ort-linux 83.52% <ø> (ø)
cli-ort-windows 83.11% <ø> (+0.09%) ⬆️
mlas 80.96% <ø> (+0.23%) ⬆️
offline 78.58% <73.48%> (-0.01%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...es/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs 73.92% <73.48%> (-0.02%) ⬇️

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/small_generic_f32_threads=8/1x256x256 31.93 µs 59.74 µs +87.1%
🔴 gather/large_f16_threads=1-internal/131072 11.99 µs 22.19 µs +85.1%
🔴 matmul/small_generic_f16_threads=8/1x256x256 28.56 µs 48.70 µs +70.5%
🔴 matmul/small_generic_f16_threads=1/1x256x256 28.63 µs 48.04 µs +67.8%
🔴 matmul/large_generic_f32_threads=8/32x1024x1024 3.65 ms 5.06 ms +38.4%
🔴 logit_processing/seven_processor_chain_per_step 296.92 µs 400.47 µs +34.9%
🔴 matmul/small_generic_f32_threads=1/1x256x256 34.62 µs 45.37 µs +31.1%
🔴 sampling_latency/top_p_per_token 352.76 µs 460.06 µs +30.4%
⚠️ matmul/medium_generic_bf16_threads=8/32x512x512 363.87 µs 469.26 µs +29.0%
⚠️ gather/large_f32_threads=1-internal/131072 31.10 µs 40.02 µs +28.7%
⚠️ kv_cache/alloc_dealloc_pages 38.05 µs 48.96 µs +28.7%
⚠️ matmul/medium_generic_f32_threads=8/32x512x512 890.55 µs 1.14 ms +28.4%
⚠️ sampling_latency/top_k_per_token 48.53 µs 61.79 µs +27.3%
⚠️ gather/large_bf16_threads=1-internal/131072 11.61 µs 14.43 µs +24.3%
⚠️ sampling_latency/min_p_per_token 195.06 µs 236.36 µs +21.2%
⚠️ gather/small_bf16_threads=1-internal/4096 492.7 ns 594.3 ns +20.6%
⚠️ matmul/small_generic_bf16_threads=1/1x256x256 29.24 µs 35.15 µs +20.2%
⚠️ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.30 ms 3.95 ms +19.4%
⚠️ add/large_bf16_threads=1-internal/4194304 41.74 ms 48.93 ms +17.2%
⚠️ grammar_masking/llguidance_compute_mask/32 69.77 µs 81.70 µs +17.1%
⚠️ matmul/large_generic_f32_threads=1/32x1024x1024 8.66 ms 10.08 ms +16.3%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 1.97 ms 2.27 ms +15.0%
✅ qwen3_sampling_processors/top_k_partial_selection 134.66 µs 154.23 µs +14.5%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.17 ms 2.48 ms +14.2%
✅ gather/small_f16_threads=1-internal/4096 460.5 ns 525.3 ns +14.1%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 488.02 µs 555.68 µs +13.9%
✅ matmul/small_generic_bf16_threads=8/1x256x256 29.86 µs 33.97 µs +13.8%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 78.05 µs 88.65 µs +13.6%
✅ reduce_mean/small_f32_threads=1-internal/4096 14.81 µs 16.72 µs +12.9%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 494.86 µs 549.02 µs +10.9%
✅ matmul/medium_generic_f16_threads=1/32x512x512 27.52 µs 30.09 µs +9.4%
✅ gather/medium_bf16_threads=1-internal/32768 2.40 µs 2.62 µs +9.0%
✅ qwen3_sampling_processors/top_k_top_p_fast 674.81 µs 728.45 µs +7.9%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.26 ms 5.68 ms +7.9%
✅ sampling_latency/greedy_per_token 3.09 µs 3.32 µs +7.3%
✅ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 40.01 µs 42.88 µs +7.2%
✅ reduce_mean/medium_f32_threads=1-internal/65536 243.91 µs 260.72 µs +6.9%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.02 ms 1.08 ms +6.2%
✅ matmul/medium_generic_f16_threads=8/32x512x512 28.46 µs 30.11 µs +5.8%
✅ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 516.18 µs 540.87 µs +4.8%
✅ gather/medium_f16_threads=1-internal/32768 2.37 µs 2.44 µs +3.1%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.26 ms 1.29 ms +2.4%
✅ gather/small_f32_threads=1-internal/4096 799.6 ns 816.2 ns +2.1%
✅ tokenization/encode_tokens_per_second 351.32 µs 356.31 µs +1.4%
✅ add/large_f16_threads=1-internal/4194304 41.49 ms 42.00 ms +1.2%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 41.40 µs 41.36 µs -0.1%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 73.09 µs 72.55 µs -0.7%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 352.80 µs 348.89 µs -1.1%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.89 ms 1.86 ms -1.5%
✅ tokenization/decode_tokens_per_second 5.90 ms 5.80 ms -1.8%
✅ add/small_bf16_threads=1-internal/1024 13.17 µs 12.44 µs -5.6%
✅ add/small_f32_threads=1-internal/1024 204.3 ns 192.3 ns -5.9%
✅ gather/medium_f32_threads=1-internal/32768 3.88 µs 3.61 µs -7.1%
✅ add/medium_f32_threads=1-internal/262144 2.63 ms 2.37 ms -9.6%
✅ add/medium_f16_threads=1-internal/262144 2.66 ms 2.40 ms -9.9%
✅ add/medium_bf16_threads=1-internal/262144 2.78 ms 2.46 ms -11.7%
✅ add/small_f16_threads=1-internal/1024 13.55 µs 11.95 µs -11.9%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 82.39 µs 71.68 µs -13.0%
🟢 add/large_f32_threads=1-internal/4194304 44.77 ms 38.00 ms -15.1%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.22 3.42 5.19 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby

Copy link
Copy Markdown
Owner Author

Validated and merging.

Host: RTX 4060 laptop, 20 logical CPUs, 68.5 GB RAM, AVX2 + FMA + F16C + AVX-VNNI, no AVX-512 -- so this exercises your borrowed_int4_block_dot_avx2 path, not the AVX-512 one. Same binary throughout, ONNX_GENAI_CPU_DISABLE_INT4_SIMD=0/1, --release --no-default-features --features native-backend, run through --backend native, ONNX_GENAI_CPU_DECODE_THREADS=8.

Measured in CPU time, not wall clock. Wall clock on this box was unusable -- three identical runs of the same configuration gave 39.3 / 25.8 / 16.1 s because another job was running. Process CPU time (TotalProcessorTime) is nearly contention-immune for a fixed thread count, and reproduced to ~2%. Decode isolated by differencing an 88-token run against an 8-token run, qwen05b-symzp:

8 tok 88 tok decode (80 tok) CPU s/token
SIMD disabled (scalar) 55.2 s 390.5 s 335.3 s 4.19
SIMD enabled 25.4 s 172.8 s 147.5 s 1.84

2.27x less CPU work per token, and the fixed cost drops 55.2 → 25.4 s CPU as well, since prefill runs the same kernel.

Correctness: generated text byte-identical across the A/B on qwen05b-symzp (symmetric, no zero_points), qwen05b-q4-zp (asymmetric) and qwen05b-q4. So both the implicit-midpoint-8 symmetric path and the explicit zero-point path are covered.

Footprint preserved (#979): peak working set 378 MB scalar vs 374 MB SIMD on a 367 MB model. Nothing dequantised outlives the call, as claimed.

cargo test -p onnx-runtime-ep-cpu --lib matmul_nbits -- 80 passed, 0 failed, 2 ignored.

Why this matters more than it looks

mlas is not in the CLI's default feature set, so the borrowed path in this PR is what a default CPU build actually runs. Worth putting the two CPU int4 efforts side by side, both measured by me on this host with qwen05b-symzp:

approach decode gain extra memory on qwen14b-symzp
this PR (SIMD the borrowed path) 2.27x none (374 vs 378 MB on the 0.5B)
#1027 (route to MLAS SQNBit) 2.68x +17.4 GB (8.17 → 25.5 GB)

SIMD gets ~85% of the MLAS win at zero footprint cost. That reorders the two: this lands now and becomes the floor everywhere, and #1027 stays blocked until its packed buffer is admitted against the memory budget rather than allocated silently.

@justinchuby
justinchuby merged commit 3cf49d2 into main Aug 16, 2026
9 of 16 checks passed
@justinchuby
justinchuby deleted the squad/994-x86-borrowed-int4-simd branch August 16, 2026 01:51
justinchuby added a commit that referenced this pull request Aug 17, 2026
…nt copy) (#1104)

## Summary

MLAS SQNBit CompFp32 beats our borrowed int4 decode path (#1021) by
~1.25x on CPU (#1027). This PR finds **where** the gap comes from,
proves it is reachable **without** MLAS's session-lifetime resident
packed copy, and ports it.

**The 1.25x is register/N-blocking, not arithmetic and not layout.**
Both kernels are pure f32 FMA (CompFp32 uses no VNNI), so arithmetic is
equal. Reading the MLAS AVX2 M==1 kernel (`sqnbitgemm_kernel_avx2.cpp`)
against our borrowed path (`matmul_nbits.rs`), the concrete differences
are:

| | MLAS SQNBit M==1 | our borrowed path (before) |
| --- | --- | --- |
| output columns per pass | **4** (`NCols4`), 4 live accumulators | 1,
single accumulator |
| activation load | loaded once, **reused across 4 columns** | reloaded
per column |
| horizontal reduction | **once per column** at the end | **once per
block** |
| nibble unpack | shuffle-free (prepacked `\|v0 v16\|…`) |
`unpacklo/unpackhi` interleave |
| resident cost | session-lifetime packed copy (~2x int4 bytes) |
**none** (zero-copy mmap) |

The first three are pure inner-loop restructuring that need **no**
repack. Only the shuffle-free unpack is tied to MLAS's prepacked layout.

## What this implements

`borrowed_affine_int4_matmul_nblock`: reads the **same zero-copy mmap
int4 layout** (no repack, no resident copy) but processes up to four
output columns per pass with four independent f32 accumulators, loads
each activation vector once and reuses it across the group, folds each
block's scale into the running accumulator with a single FMA, carries
the zero-point affine correction in scalar, and does one horizontal
reduction per column. The nibble unpack is byte-for-byte identical to
the per-column AVX2 helper.

Selected for decode on AVX2-capable hosts behind the same-binary A/B
toggle **`ONNX_GENAI_CPU_MM_INT4_NBLK`** (default off; the toggle is a
read-only env probe, production is the only writer — no new
process-global mutable state).

## Where the 1.25x comes from — evidence

All measured on this host (RTX 4060 laptop, 20 logical CPUs,
AVX2+FMA+AVX-VNNI, **no AVX-512, CPU only**), `--backend native`,
`ONNX_GENAI_CPU_DECODE_THREADS=8`, decode isolated as **88-tok −
8-tok**, **process CPU time**, **min of 3**, on `models\qwen05b-symzp`:

| path | short 8 tok | long 88 tok | decode (80 tok) | **CPU s/token** |
**peak RSS** |
| --- | --- | --- | --- | --- | --- |
| ours (borrowed int4) | 6.7 s | 55.5 s | 48.8 s | **0.61** | **389 MB**
|
| **nblk (this PR)** | 4.4 s | 37.9 s | 33.5 s | **0.42** | **389 MB** |
| MLAS SQNBit CompFp32 | 5.6 s | 36.2 s | 30.6 s | **0.38** | **1059
MB** |

**How much is reachable without a resident copy:** nblk recovers
**~83%** of MLAS's decode advantage — `(0.61−0.42)/(0.61−0.38)` — at the
borrowed path's **389 MB**, versus MLAS's **1059 MB**. The residual
nblk↔mlas gap (0.42 vs 0.38) is **within nblk's own run-to-run spread**
(long runs 37.9 / 43.7 / 45.8 s), i.e. below the measurement floor.
Because nblk reads the **normal mmap layout** and still captures
essentially all of the win, **the advantage is inner-loop register
blocking, not layout**; the shuffle-free-unpack contribution is at or
under the noise floor on this host and does not justify a transient
repack.

## The 14B — the model that needed it most

On `models\qwen14b-symzp` MLAS's resident cache is **declined** at the
default ceiling: predicted `resident_f32_cache_bytes = 17,925,488,640`
(16.7 GB) exceeds the residency ceiling,
`f32_weight_cache_admitted=false`, so the MLAS arm peaks at **8189 MB**
— identical to borrowed. **MLAS gives the 14B nothing.** nblk needs no
admission (it is not a resident cache), so it applies here:

| path | decode (40-8 tok) | **CPU s/token** | **peak RSS** |
| --- | --- | --- | --- |
| ours (borrowed int4) | 442.5 s | **13.83** | 8220 MB |
| **nblk (this PR)** | 304.1 s | **9.50** | 8222 MB |

**1.46× faster decode at the same 8.2 GB**, byte-identical — exactly the
model MLAS could not help.

## Byte-identity

Generated text is byte-identical (SHA-256 of `--raw` output) across
**borrowed / nblk / mlas** on:
- `qwen05b-symzp` (symmetric) — `12B539F850DFE618`
- `qwen05b-q4-zp` (asymmetric) — `3EA4EABE60BB60D1`
- `qwen14b-symzp` (40-tok) — `3C517757A47C1395`

## Gates

- `cargo test -p onnx-runtime-ep-cpu --lib`, five consecutive runs:
**1308 / 1308 / 1308 / 1308 / 1308 passed, 0 failed, 11 ignored** each.
- `cargo clippy -p onnx-runtime-ep-cpu --lib -- -D warnings`: clean.
- New parity test `nblock_matches_per_column_borrowed_path` (symmetric +
asymmetric, single/multi-row, non-multiple-of-4 tail group). `cargo test
--features mlas --lib matmul_nbits`: 111 passed, 0 failed, 6 ignored.

## Notes

- Default off; this is an opt-in A/B path, like #1021 and #1027.
Flipping it on by default is a follow-up once it has soaked, but the
measurements above show it is a strict win (faster decode, identical
footprint, byte-identical output).
- Not absorbed: the shuffle-free in-block nibble order. It requires
either MLAS's resident prepack (refused) or a transient repack, and its
isolated contribution is under the measurement floor here — a
well-measured "not worth a repack on this host."

Refs #1021, #1027, #1051, #1056

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
justinchuby added a commit that referenced this pull request Aug 17, 2026
…ed sums)

Replace the row-serial borrowed int4 prefill path's two per-token overheads
with a structural rewrite behind a default-off A/B toggle
(ONNX_GENAI_CPU_MM_INT4_PREFILL), per the method used in #1021/#1027/#1104/#1116:

- One fork-join over disjoint column strips for the whole prefill, instead of
  the row-serial path's m per-row fork-joins.
- The per-block �ctivation_sums Vec is hoisted out of the per-row loop to a
  single m * block_count allocation, independent of the weight size.

Rows are still visited outer-most, so a column's packed bytes are re-read once
per row: weight traffic and the per-element k-reduction order are unchanged, so
output is byte-identical to the row-serial path (both call the shared
�orrowed_int4_output_element). No resident buffer is added (peak RSS
unchanged), satisfying the #1056/#1117 no-session-buffer constraint.

GEMM blocking (reusing a column's bytes across a tile of rows) is deliberately
NOT included here: measured within run-to-run noise on qwen05b and with no
signal on qwen14b (5-rep interleaved), so it is left as an unproven follow-up
to #1117 that can be added or dropped cleanly.

Measured win is model-size dependent: qwen05b prefill slope 1.219 -> 0.842 CPU
s/token (median of 5, ~1.45x); qwen14b shows no measurable change (medians
8.69/8.62 off/on, within a 30-40% within-arm spread) because the fixed per-row
overheads this removes are a negligible fraction of the 14B's larger per-row
compute.

Refs #1117

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
justinchuby added a commit that referenced this pull request Aug 19, 2026
#994) — re-measured, ~3x below roofline, not ~63x (#1013)

## What

Adds `roofline_gemv`, a CPU-only bench probe (behind a lean `gemv-probe`
feature) that measures the **effective memory bandwidth the real CPU
int4 `MatMulNBits` decode kernel achieves** at the 14B `lm_head` shape —
the last unmeasured term in #994's placement criterion `F/C_slow <
W/B_link + F/C_fast`.

It is the compute-side companion to `roofline_bandwidth` (host DRAM
STREAM ceiling) and `roofline_transfer` (host↔device link).

## ⚠️ Retraction: the original "false by ~63×" claim was wrong

**This PR originally reported that the CPU `lm_head` GEMV runs at ~0.78
GB/s — 1.6% of the DRAM ceiling — making CPU compute ~500 ms and
inverting #994's placement criterion by ~63×. That result does not
reproduce and is retracted.**

Two things were wrong with it:

1. **It measured a kernel that no longer exists.** The probe was written
against the #979 borrowed symmetric-int4 path when that path was a
scalar unpack-convert-FMA loop on x86 (only AArch64 had the NEON dot).
#1021 (`3cf49d25`, merged 2026-08-16) vectorised that GEMV for
AVX2/AVX-512, and #1356/#1431 have since reworked the prefill dispatch
around it. The number was real when taken and is meaningless now.
2. **~63× was never the right way to state it even then.** The
comparison mixed a weight-only rate against a total-traffic ceiling and
reported a single best sample from a heavily shared box.

## Re-measured on `origin/main` (`085c49140`), same box, same probe

DRAM read ceiling, `roofline_bandwidth --mib 2048 --seconds 3`:

| threads | GB/s |
|---:|---:|
| 8 | 75.099 |
| 16 | 75.460 |
| 32 | **75.737** |

The int4 `lm_head` GEMV, `roofline_gemv --repeats 7` (K=5120, N=152064,
block_size=32, bits=4, accuracy_level=0; weight 389.3 MB + scales 97.3
MB = 487.2 MB total traffic):

| threads | per-call ms (median) | per-call ms (min) | weight GB/s
(best) | total GB/s (best) |
|---:|---:|---:|---:|---:|
| 8 | 52.560 | 52.330 | 7.439 | 9.311 |
| 16 | 29.349 | 26.388 | 14.752 | 18.464 |
| 32 | 19.785 | 17.936 | 21.705 | **27.166** |

**Corrected result: the kernel reaches 27.2 GB/s of 75.7 GB/s — it is
~2.8× below the roofline, not ~63×.** Per-token `lm_head` compute is
**~18–20 ms**, not ~500 ms.

## What that does to #994's criterion

The criterion no longer inverts. At ~18–20 ms of CPU compute against ~33
ms to move the work to the GPU, CPU `lm_head` is **cheaper**, not 15×
more expensive — the opposite of what this PR originally concluded.
#994's original assumption (that the GEMV runs at DRAM bandwidth, ~7.9
ms) is still optimistic by ~2.5×, but it is the right order of magnitude
and the placement decision it drives stands.

The remaining 2.8× is the honest open question this probe exists to
track: the kernel still widens every nibble to f32 and streams 97 MB of
f32 scales alongside 389 MB of packed weights, so it is not purely
bandwidth-limited even after vectorisation.

## Why this is still worth merging

The conclusion was wrong; the instrument was not. `roofline_gemv` is
what caught the error, and it is the only probe in the tree that
measures the real quantised decode GEMV against the real DRAM ceiling on
the same box. Merging it makes that measurement repeatable instead of a
claim in a PR description.

## How it stays faithful

- Drives the real kernel via the CPU EP's own `get_kernel`/`execute` —
no reimplementation.
- Confirms the borrowed symmetric-int4 dispatch is taken (bits=4,
accuracy_level=0, no prepack/g_idx, symmetric → `None` zero points).
- Runs warmup+timing inside one `with_decode_pool_scope` so each
`execute` dispatches to the **persistent SPMD decode pool** the runtime
actually uses — not a per-call fork/join of the flat pool (which
understated bandwidth ~10×).
- Sweeps thread counts one process per count via
`set_decode_thread_budget`.
- Reports both a weight-only rate (the #994 criterion input) and a
total-traffic rate (honest roofline efficiency, incl. the f32 scales),
and a distribution rather than one sample.

## Scope

**Measurement, not implementation.** No kernels are restructured.
`mlas`/`accuracy_level=4` (MLAS SQNBit) deliberately not measured — the
default artifact carries zero MLAS symbols by design, so it is not the
path the runtime takes.

---------

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Roy <roy@squad.local>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants