Repository navigation
cpu: vectorise x86 borrowed int4 GEMV (AVX2/AVX-512) for #994 - #1021
Conversation
The x86 symmetric/asymmetric int4 borrowed decode path (introduced in #979) was a scalar unpack-convert-FMA loop; only aarch64 had a vectorised route. At the 14B lm_head shape it hit ~0.78 GB/s, ~1.6% of this box's 49.28 GB/s STREAM ceiling. Add a runtime-dispatched f32 SIMD block dot for the borrowed path, reusing the existing selected_dot_kernel()/DotKernel seam: - AVX2+FMA kernel (covers Avx2 and AvxVnni hosts; the borrowed path is f32, so integer VNNI does not apply). - AVX-512F kernel for Avx512Vnni hosts. Both unpack the 4-bit nibbles, widen to f32, and multiply-accumulate the activation per 32-lane tile; the per-block scale and zero-point correction stay in the caller, so symmetric (implicit midpoint 8) and asymmetric both work unchanged. Nothing dequantised outlives the call, preserving the #979 zero-copy footprint. The horizontal reduction reorders additions vs the scalar loop, so f32 logits differ by a few ULP; it is not a bit-identical f32 transform. Token IDs stay byte-identical (argmax stable) on qwen05b-q4 (symmetric) and qwen2.5-0.5b-q4_0-mobius (asymmetric) for 96 greedy tokens. ONNX_GENAI_CPU_DISABLE_INT4_SIMD=1 is a documented A/B escape hatch to the scalar reference. aarch64 is untouched. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #1021 +/- ##
==========================================
- Coverage 78.79% 78.79% -0.01%
==========================================
Files 364 364
Lines 146900 147032 +132
Branches 146900 147032 +132
==========================================
+ Hits 115754 115855 +101
- Misses 26529 26560 +31
Partials 4617 4617
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
|
Validated and merging. Host: RTX 4060 laptop, 20 logical CPUs, 68.5 GB RAM, Measured in CPU time, not wall clock. Wall clock on this box was unusable -- three identical runs of the same configuration gave 39.3 / 25.8 / 16.1 s because another job was running. Process CPU time (
2.27x less CPU work per token, and the fixed cost drops 55.2 → 25.4 s CPU as well, since prefill runs the same kernel. Correctness: generated text byte-identical across the A/B on Footprint preserved (#979): peak working set 378 MB scalar vs 374 MB SIMD on a 367 MB model. Nothing dequantised outlives the call, as claimed.
Why this matters more than it looks
SIMD gets ~85% of the MLAS win at zero footprint cost. That reorders the two: this lands now and becomes the floor everywhere, and #1027 stays blocked until its packed buffer is admitted against the memory budget rather than allocated silently. |
…nt copy) (#1104) ## Summary MLAS SQNBit CompFp32 beats our borrowed int4 decode path (#1021) by ~1.25x on CPU (#1027). This PR finds **where** the gap comes from, proves it is reachable **without** MLAS's session-lifetime resident packed copy, and ports it. **The 1.25x is register/N-blocking, not arithmetic and not layout.** Both kernels are pure f32 FMA (CompFp32 uses no VNNI), so arithmetic is equal. Reading the MLAS AVX2 M==1 kernel (`sqnbitgemm_kernel_avx2.cpp`) against our borrowed path (`matmul_nbits.rs`), the concrete differences are: | | MLAS SQNBit M==1 | our borrowed path (before) | | --- | --- | --- | | output columns per pass | **4** (`NCols4`), 4 live accumulators | 1, single accumulator | | activation load | loaded once, **reused across 4 columns** | reloaded per column | | horizontal reduction | **once per column** at the end | **once per block** | | nibble unpack | shuffle-free (prepacked `\|v0 v16\|…`) | `unpacklo/unpackhi` interleave | | resident cost | session-lifetime packed copy (~2x int4 bytes) | **none** (zero-copy mmap) | The first three are pure inner-loop restructuring that need **no** repack. Only the shuffle-free unpack is tied to MLAS's prepacked layout. ## What this implements `borrowed_affine_int4_matmul_nblock`: reads the **same zero-copy mmap int4 layout** (no repack, no resident copy) but processes up to four output columns per pass with four independent f32 accumulators, loads each activation vector once and reuses it across the group, folds each block's scale into the running accumulator with a single FMA, carries the zero-point affine correction in scalar, and does one horizontal reduction per column. The nibble unpack is byte-for-byte identical to the per-column AVX2 helper. Selected for decode on AVX2-capable hosts behind the same-binary A/B toggle **`ONNX_GENAI_CPU_MM_INT4_NBLK`** (default off; the toggle is a read-only env probe, production is the only writer — no new process-global mutable state). ## Where the 1.25x comes from — evidence All measured on this host (RTX 4060 laptop, 20 logical CPUs, AVX2+FMA+AVX-VNNI, **no AVX-512, CPU only**), `--backend native`, `ONNX_GENAI_CPU_DECODE_THREADS=8`, decode isolated as **88-tok − 8-tok**, **process CPU time**, **min of 3**, on `models\qwen05b-symzp`: | path | short 8 tok | long 88 tok | decode (80 tok) | **CPU s/token** | **peak RSS** | | --- | --- | --- | --- | --- | --- | | ours (borrowed int4) | 6.7 s | 55.5 s | 48.8 s | **0.61** | **389 MB** | | **nblk (this PR)** | 4.4 s | 37.9 s | 33.5 s | **0.42** | **389 MB** | | MLAS SQNBit CompFp32 | 5.6 s | 36.2 s | 30.6 s | **0.38** | **1059 MB** | **How much is reachable without a resident copy:** nblk recovers **~83%** of MLAS's decode advantage — `(0.61−0.42)/(0.61−0.38)` — at the borrowed path's **389 MB**, versus MLAS's **1059 MB**. The residual nblk↔mlas gap (0.42 vs 0.38) is **within nblk's own run-to-run spread** (long runs 37.9 / 43.7 / 45.8 s), i.e. below the measurement floor. Because nblk reads the **normal mmap layout** and still captures essentially all of the win, **the advantage is inner-loop register blocking, not layout**; the shuffle-free-unpack contribution is at or under the noise floor on this host and does not justify a transient repack. ## The 14B — the model that needed it most On `models\qwen14b-symzp` MLAS's resident cache is **declined** at the default ceiling: predicted `resident_f32_cache_bytes = 17,925,488,640` (16.7 GB) exceeds the residency ceiling, `f32_weight_cache_admitted=false`, so the MLAS arm peaks at **8189 MB** — identical to borrowed. **MLAS gives the 14B nothing.** nblk needs no admission (it is not a resident cache), so it applies here: | path | decode (40-8 tok) | **CPU s/token** | **peak RSS** | | --- | --- | --- | --- | | ours (borrowed int4) | 442.5 s | **13.83** | 8220 MB | | **nblk (this PR)** | 304.1 s | **9.50** | 8222 MB | **1.46× faster decode at the same 8.2 GB**, byte-identical — exactly the model MLAS could not help. ## Byte-identity Generated text is byte-identical (SHA-256 of `--raw` output) across **borrowed / nblk / mlas** on: - `qwen05b-symzp` (symmetric) — `12B539F850DFE618` - `qwen05b-q4-zp` (asymmetric) — `3EA4EABE60BB60D1` - `qwen14b-symzp` (40-tok) — `3C517757A47C1395` ## Gates - `cargo test -p onnx-runtime-ep-cpu --lib`, five consecutive runs: **1308 / 1308 / 1308 / 1308 / 1308 passed, 0 failed, 11 ignored** each. - `cargo clippy -p onnx-runtime-ep-cpu --lib -- -D warnings`: clean. - New parity test `nblock_matches_per_column_borrowed_path` (symmetric + asymmetric, single/multi-row, non-multiple-of-4 tail group). `cargo test --features mlas --lib matmul_nbits`: 111 passed, 0 failed, 6 ignored. ## Notes - Default off; this is an opt-in A/B path, like #1021 and #1027. Flipping it on by default is a follow-up once it has soaked, but the measurements above show it is a strict win (faster decode, identical footprint, byte-identical output). - Not absorbed: the shuffle-free in-block nibble order. It requires either MLAS's resident prepack (refused) or a transient repack, and its isolated contribution is under the measurement floor here — a well-measured "not worth a repack on this host." Refs #1021, #1027, #1051, #1056 Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
…ed sums) Replace the row-serial borrowed int4 prefill path's two per-token overheads with a structural rewrite behind a default-off A/B toggle (ONNX_GENAI_CPU_MM_INT4_PREFILL), per the method used in #1021/#1027/#1104/#1116: - One fork-join over disjoint column strips for the whole prefill, instead of the row-serial path's m per-row fork-joins. - The per-block �ctivation_sums Vec is hoisted out of the per-row loop to a single m * block_count allocation, independent of the weight size. Rows are still visited outer-most, so a column's packed bytes are re-read once per row: weight traffic and the per-element k-reduction order are unchanged, so output is byte-identical to the row-serial path (both call the shared �orrowed_int4_output_element). No resident buffer is added (peak RSS unchanged), satisfying the #1056/#1117 no-session-buffer constraint. GEMM blocking (reusing a column's bytes across a tile of rows) is deliberately NOT included here: measured within run-to-run noise on qwen05b and with no signal on qwen14b (5-rep interleaved), so it is left as an unproven follow-up to #1117 that can be added or dropped cleanly. Measured win is model-size dependent: qwen05b prefill slope 1.219 -> 0.842 CPU s/token (median of 5, ~1.45x); qwen14b shows no measurable change (medians 8.69/8.62 off/on, within a 30-40% within-arm spread) because the fixed per-row overheads this removes are a negligible fraction of the 14B's larger per-row compute. Refs #1117 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
#994) — re-measured, ~3x below roofline, not ~63x (#1013) ## What Adds `roofline_gemv`, a CPU-only bench probe (behind a lean `gemv-probe` feature) that measures the **effective memory bandwidth the real CPU int4 `MatMulNBits` decode kernel achieves** at the 14B `lm_head` shape — the last unmeasured term in #994's placement criterion `F/C_slow < W/B_link + F/C_fast`. It is the compute-side companion to `roofline_bandwidth` (host DRAM STREAM ceiling) and `roofline_transfer` (host↔device link). ##⚠️ Retraction: the original "false by ~63×" claim was wrong **This PR originally reported that the CPU `lm_head` GEMV runs at ~0.78 GB/s — 1.6% of the DRAM ceiling — making CPU compute ~500 ms and inverting #994's placement criterion by ~63×. That result does not reproduce and is retracted.** Two things were wrong with it: 1. **It measured a kernel that no longer exists.** The probe was written against the #979 borrowed symmetric-int4 path when that path was a scalar unpack-convert-FMA loop on x86 (only AArch64 had the NEON dot). #1021 (`3cf49d25`, merged 2026-08-16) vectorised that GEMV for AVX2/AVX-512, and #1356/#1431 have since reworked the prefill dispatch around it. The number was real when taken and is meaningless now. 2. **~63× was never the right way to state it even then.** The comparison mixed a weight-only rate against a total-traffic ceiling and reported a single best sample from a heavily shared box. ## Re-measured on `origin/main` (`085c49140`), same box, same probe DRAM read ceiling, `roofline_bandwidth --mib 2048 --seconds 3`: | threads | GB/s | |---:|---:| | 8 | 75.099 | | 16 | 75.460 | | 32 | **75.737** | The int4 `lm_head` GEMV, `roofline_gemv --repeats 7` (K=5120, N=152064, block_size=32, bits=4, accuracy_level=0; weight 389.3 MB + scales 97.3 MB = 487.2 MB total traffic): | threads | per-call ms (median) | per-call ms (min) | weight GB/s (best) | total GB/s (best) | |---:|---:|---:|---:|---:| | 8 | 52.560 | 52.330 | 7.439 | 9.311 | | 16 | 29.349 | 26.388 | 14.752 | 18.464 | | 32 | 19.785 | 17.936 | 21.705 | **27.166** | **Corrected result: the kernel reaches 27.2 GB/s of 75.7 GB/s — it is ~2.8× below the roofline, not ~63×.** Per-token `lm_head` compute is **~18–20 ms**, not ~500 ms. ## What that does to #994's criterion The criterion no longer inverts. At ~18–20 ms of CPU compute against ~33 ms to move the work to the GPU, CPU `lm_head` is **cheaper**, not 15× more expensive — the opposite of what this PR originally concluded. #994's original assumption (that the GEMV runs at DRAM bandwidth, ~7.9 ms) is still optimistic by ~2.5×, but it is the right order of magnitude and the placement decision it drives stands. The remaining 2.8× is the honest open question this probe exists to track: the kernel still widens every nibble to f32 and streams 97 MB of f32 scales alongside 389 MB of packed weights, so it is not purely bandwidth-limited even after vectorisation. ## Why this is still worth merging The conclusion was wrong; the instrument was not. `roofline_gemv` is what caught the error, and it is the only probe in the tree that measures the real quantised decode GEMV against the real DRAM ceiling on the same box. Merging it makes that measurement repeatable instead of a claim in a PR description. ## How it stays faithful - Drives the real kernel via the CPU EP's own `get_kernel`/`execute` — no reimplementation. - Confirms the borrowed symmetric-int4 dispatch is taken (bits=4, accuracy_level=0, no prepack/g_idx, symmetric → `None` zero points). - Runs warmup+timing inside one `with_decode_pool_scope` so each `execute` dispatches to the **persistent SPMD decode pool** the runtime actually uses — not a per-call fork/join of the flat pool (which understated bandwidth ~10×). - Sweeps thread counts one process per count via `set_decode_thread_budget`. - Reports both a weight-only rate (the #994 criterion input) and a total-traffic rate (honest roofline efficiency, incl. the f32 scales), and a distribution rather than one sample. ## Scope **Measurement, not implementation.** No kernels are restructured. `mlas`/`accuracy_level=4` (MLAS SQNBit) deliberately not measured — the default artifact carries zero MLAS symbols by design, so it is not the path the runtime takes. --------- Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com> Co-authored-by: Roy <roy@squad.local>
Vectorise the x86 borrowed int4 GEMV (refs #994, #979)
The x86 symmetric/asymmetric int4 borrowed decode path (unified by #979) was a
scalar unpack-convert-FMA loop; only aarch64 had a vectorised route
(
borrowed_affine_int4_matmul_m1_neon_dot/affine_int4_block32_dot_neon). At the14B
lm_headshape it measured ~0.78 GB/s, ~1.6% of this box's 49.28 GB/s STREAMceiling (per the #1013
roofline_gemvprobe).What changed
existing
selected_dot_kernel()/DotKernelseam (no new mechanism, nocompile-time target-feature gate):
borrowed_int4_block_dot_avx2— AVX2+FMA, servesAvx2andAvxVnnihosts(the borrowed path is f32, so integer VNNI does not apply).
borrowed_int4_block_dot_avx512— AVX-512F, servesAvx512Vnnihosts.w[2i]=lo(byte i), w[2i+1]=hi(byte i)order, widen to f32, and multiply-accumulatethe activation per 32-lane tile. The per-block scale and the zero-point correction
stay in the caller, so symmetric (
zero_points=None, implicit midpoint 8) andasymmetric both work unchanged. Blocks larger than 32 (e.g. block-128) vectorise
via the 32-lane chunk loop into a single accumulator.
ONNX_GENAI_CPU_DISABLE_INT4_SIMD=1is a documented A/B escape hatch to the scalarreference (same-binary A/B,
OnceLock-resolved, mirrorsarm64_int4_direct_enabled).let _ = dot_kernel;unused-varguard is now
#[cfg(not(any(aarch64, x86_64)))](aarch64 already useddot_kernel;behaviour there is byte-for-byte the same).
1. Roofline (
roofline_gemv, K=5120 N=152064 block_size=32 bits=4, borrowed symmetric int4)Weight-only GB/s (criterion input). Box: shared i7-13800H (14C/20T, AVX2+AVX-VNNI,
no AVX-512 — so the AVX2 kernel is what runs here; the AVX-512 kernel is compiled
but unexercised on this host).
iters=10 repeats=8; two other squad agents active, so Ireport median and best (quietest) rather than a single sample.
Headline (20 threads, quiet): ~8.1 GB/s best / ~6.7 GB/s median, ~16.5% of the
49.28 GB/s ceiling, up from ~0.78 GB/s. ~6.7–6.9× at the useful thread counts.
Honest-negative note: the scalar loop was a real bottleneck (~7× confirms it), but
we are still ~6× below roofline (16.5%, not ~100%). The residual gap is now
elsewhere — the +25% scales traffic the criterion omits, the strided per-row weight
reads, and the per-
executeSPMD barrier — not the nibble unpack. There is more CPUheadroom, but it is no longer in this loop.
2. Byte-identical output
96 greedy tokens, native CPU EP, SIMD on vs
ONNX_GENAI_CPU_DISABLE_INT4_SIMD=1:qwen05b-q4(symmetric, borrowed path): token IDs byte-identical (all 96).qwen2.5-0.5b-q4_0-mobius(asymmetric, same path): token IDs byte-identical (all 96).This is not a bit-identical f32 transform: the horizontal reduction + FMA reorder
additions vs the scalar loop, so logits differ by a few ULP. That divergence is below
the argmax margin here — tokens are byte-identical on both models — but I am stating it
rather than claiming bit-identity. The unit test locks the block dot to the scalar
reference within
|Δ| ≤ |ref|·1e-5 + 1e-4.3. Footprint unchanged
Peak working set on
qwen05b-q4: 383.0 MiB (SIMD) vs 372.2 MiB (scalar); asymmetricmobius 380.0 vs 379.5 MiB. No f32 expansion (that would be GBs); the borrowed tile does
not outlive the call. Well under the ~434 MB envelope.
4. aarch64
No functional change. The only aarch64-visible edit is the
let _ = dot_kernel;cfgguard (aarch64 already consumed
dot_kernel). Cannot be run here; called out explicitly.Gates (verbatim)
cargo fmt -p onnx-runtime-ep-cpu— cleancargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings— cleancargo test -p onnx-runtime-ep-cpu:unittests src\lib.rs:test result: ok. 1078 passed; 0 failed; 10 ignoredbf16_conformance:ok. 3 passed; 0 failedkernel_numeric_regression:ok. 10 passed; 0 failedmha_ort_parity:ok. 1 passed; 0 failedmsft_attention_ort_parity:ok. 1 passed; 0 failedqwen35_ort_parity:ok. 5 passed; 0 failedshared_allocator:ok. 6 passed; 0 failedcargo test -p onnx-genai-engine --lib:test result: ok. 385 passed; 0 failed; 1 ignored(baseline lib was 1077; +1 is the new
borrowed_int4_block_dot_x86_matches_scalar_referencetest.)#994 criterion re-evaluation
Posted on #994. Short version, with the post-fix number:
GPU H2D still edges CPU-compute in absolute per-token latency, so the placement
argument for
lm_headdoes not come back — but the margin collapsed from ~15×(the old ~500 ms) to ~1.5×, moving it from "clearly dead" into narrow/ambiguous
territory where the VRAM-savings tradeoff is now defensible. Embedding gather is
unaffected (arithmetic intensity ~0).