Repository navigation
Add roofline_gemv: measure the real CPU int4 lm_head GEMV against DRAM (#994) — re-measured, ~3x below roofline, not ~63x - #1013
Conversation
Adds a CPU-only bench binary ( oofline_gemv, behind a lean gemv-probe feature) that drives the real symmetric-int4 MatMulNBits decode kernel -- the #979 borrowed zero-copy path (accuracy_level=0) -- at the 14B lm_head shape (K=5120, N=152064, block_size=32, bits=4, m=1) and measures the effective memory bandwidth it achieves against the STREAM DRAM ceiling. The probe runs warmup+timing inside a single with_decode_pool_scope so each execute dispatches to the persistent SPMD decode pool the runtime actually uses (not a per-call fork/join of the flat pool), and sweeps thread counts one process per count via set_decode_thread_budget. It reports both a weight-only rate (the #994 criterion input) and a total-traffic rate (honest roofline efficiency including the ~97 MB of f32 scales). Measurement, not implementation: no kernels are restructured. Answers the last unverified term in #994's placement criterion. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1013 +/- ##
===========================================
- Coverage 82.10% 80.66% -1.45%
===========================================
Files 12 377 +365
Lines 5471 166791 +161320
Branches 5471 166791 +161320
===========================================
+ Hits 4492 134542 +130050
- Misses 780 27406 +26626
- Partials 199 4843 +4644
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
## Vectorise the x86 borrowed int4 GEMV (refs #994, #979) The x86 symmetric/asymmetric int4 **borrowed** decode path (unified by #979) was a scalar unpack-convert-FMA loop; only aarch64 had a vectorised route (`borrowed_affine_int4_matmul_m1_neon_dot` / `affine_int4_block32_dot_neon`). At the 14B `lm_head` shape it measured **~0.78 GB/s, ~1.6% of this box's 49.28 GB/s STREAM ceiling** (per the #1013 `roofline_gemv` probe). ### What changed - New runtime-dispatched **f32 SIMD block dot** for the borrowed path, reusing the existing `selected_dot_kernel()` / `DotKernel` seam (no new mechanism, no compile-time target-feature gate): - `borrowed_int4_block_dot_avx2` — AVX2+FMA, serves `Avx2` **and** `AvxVnni` hosts (the borrowed path is f32, so integer VNNI does not apply). - `borrowed_int4_block_dot_avx512` — AVX-512F, serves `Avx512Vnni` hosts. - Both unpack the 4-bit nibbles into the scalar loop's natural `w[2i]=lo(byte i), w[2i+1]=hi(byte i)` order, widen to f32, and multiply-accumulate the activation per 32-lane tile. The per-block scale and the zero-point correction stay in the caller, so **symmetric** (`zero_points=None`, implicit midpoint 8) and **asymmetric** both work unchanged. Blocks larger than 32 (e.g. block-128) vectorise via the 32-lane chunk loop into a single accumulator. - Nothing dequantised outlives the call → the #979 zero-copy footprint is preserved. - `ONNX_GENAI_CPU_DISABLE_INT4_SIMD=1` is a documented A/B escape hatch to the scalar reference (same-binary A/B, `OnceLock`-resolved, mirrors `arm64_int4_direct_enabled`). - **aarch64 is untouched** except that the shared `let _ = dot_kernel;` unused-var guard is now `#[cfg(not(any(aarch64, x86_64)))]` (aarch64 already used `dot_kernel`; behaviour there is byte-for-byte the same). ### 1. Roofline (`roofline_gemv`, K=5120 N=152064 block_size=32 bits=4, borrowed symmetric int4) Weight-only GB/s (criterion input). Box: shared **i7-13800H** (14C/20T, AVX2+AVX-VNNI, **no AVX-512** — so the AVX2 kernel is what runs here; the AVX-512 kernel is compiled but unexercised on this host). `iters=10 repeats=8`; two other squad agents active, so I report median **and** best (quietest) rather than a single sample. | threads | before med / best | after med / best | after % of 49.28 | speedup (best) | |--------:|------------------:|-----------------:|-----------------:|---------------:| | 1 | 0.390 / 0.449 | 1.048 / 1.661 | 3.4% | 3.7× | | 2 | 0.604 / 0.700 | 1.946 / 2.217 | 4.5% | 3.2× | | 4 | 0.505 / 0.882 | 3.682 / 3.944 | 8.0% | 4.5× | | 8 | 0.646 / 0.777 | 4.551 / 5.219 | 10.6% | 6.7× | | 16 | 0.768 / 1.214 | 6.176 / 6.519 | 13.2% | 5.4× | | 20 | 1.021 / 1.186 | 6.711 / 8.133 | 16.5% | 6.9× | Headline (20 threads, quiet): **~8.1 GB/s best / ~6.7 GB/s median, ~16.5% of the 49.28 GB/s ceiling**, up from ~0.78 GB/s. **~6.7–6.9× at the useful thread counts.** **Honest-negative note:** the scalar loop *was* a real bottleneck (~7× confirms it), but we are still **~6× below roofline** (16.5%, not ~100%). The residual gap is now elsewhere — the +25% scales traffic the criterion omits, the strided per-row weight reads, and the per-`execute` SPMD barrier — not the nibble unpack. There is more CPU headroom, but it is no longer in this loop. ### 2. Byte-identical output 96 greedy tokens, native CPU EP, SIMD on vs `ONNX_GENAI_CPU_DISABLE_INT4_SIMD=1`: - `qwen05b-q4` (symmetric, borrowed path): **token IDs byte-identical** (all 96). - `qwen2.5-0.5b-q4_0-mobius` (asymmetric, same path): **token IDs byte-identical** (all 96). This is **not** a bit-identical f32 transform: the horizontal reduction + FMA reorder additions vs the scalar loop, so logits differ by a few ULP. That divergence is below the argmax margin here — tokens are byte-identical on both models — but I am stating it rather than claiming bit-identity. The unit test locks the block dot to the scalar reference within `|Δ| ≤ |ref|·1e-5 + 1e-4`. ### 3. Footprint unchanged Peak working set on `qwen05b-q4`: **383.0 MiB** (SIMD) vs 372.2 MiB (scalar); asymmetric mobius 380.0 vs 379.5 MiB. No f32 expansion (that would be GBs); the borrowed tile does not outlive the call. Well under the ~434 MB envelope. ### 4. aarch64 No functional change. The only aarch64-visible edit is the `let _ = dot_kernel;` cfg guard (aarch64 already consumed `dot_kernel`). Cannot be run here; called out explicitly. ### Gates (verbatim) - `cargo fmt -p onnx-runtime-ep-cpu` — clean - `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings` — clean - `cargo test -p onnx-runtime-ep-cpu`: - `unittests src\lib.rs`: `test result: ok. 1078 passed; 0 failed; 10 ignored` - `bf16_conformance`: `ok. 3 passed; 0 failed` - `kernel_numeric_regression`: `ok. 10 passed; 0 failed` - `mha_ort_parity`: `ok. 1 passed; 0 failed` - `msft_attention_ort_parity`: `ok. 1 passed; 0 failed` - `qwen35_ort_parity`: `ok. 5 passed; 0 failed` - `shared_allocator`: `ok. 6 passed; 0 failed` - `cargo test -p onnx-genai-engine --lib`: `test result: ok. 385 passed; 0 failed; 1 ignored` (baseline lib was 1077; +1 is the new `borrowed_int4_block_dot_x86_matches_scalar_reference` test.) ### #994 criterion re-evaluation Posted on #994. Short version, with the post-fix number: - move to GPU: 389 MB / 11.74 GB/s (H2D pinned) = **~33 ms** - compute on CPU: 389 MB / ~6.7–8.1 GB/s = **~48–58 ms** (20 threads, median→best) GPU H2D still edges CPU-compute in absolute per-token latency, so the placement argument for `lm_head` **does not come back** — but the margin collapsed from ~15× (the old ~500 ms) to ~1.5×, moving it from "clearly dead" into narrow/ambiguous territory where the VRAM-savings tradeoff is now defensible. Embedding gather is unaffected (arithmetic intensity ~0). Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
|
Built and ran the probe on the RTX 4060 laptop box (20 logical CPUs, AVX2 + FMA + F16C + AVX-VNNI, no AVX-512). It builds and runs cleanly, and the CSV shape is right. But I cannot use it yet, for a reason worth fixing in the probe rather than working around. The probe reports wall-clock-derived GB/s, and on a shared box that number is not stable enough to answer the question it exists to answer. Concretely: #1021 has now landed, which vectorises exactly the kernel this probe targets. Measured end-to-end on
A 2.27x kernel change registers as 3-11%. The probe also reports 20 threads as slower than 8 (1042 ms vs 640 ms per call), and 0.6 GB/s where I previously hand-measured 2.46 GB/s at this shape. All three are symptoms of the same cause: another job was on the box (a python process holding 13,246 CPU-seconds), and wall-clock GB/s collapses under contention while the A/B difference compresses toward nothing. I hit this myself an hour earlier and it nearly cost me the #1021 conclusion: three identical runs of one configuration gave 39.3 / 25.8 / 16.1 s wall. Switching to What I would like changed before merging:
The probe itself is the right idea and I want it in the tree -- reproducible measurement infrastructure is worth more than any single measurement. It just has to be honest about when its own numbers are contaminated, which is the same standard we have been holding the kernels to. |
…389 MB is lm_head (#1304) Corrects the placement worked-example in `docs/memory/MEMORY_MANAGEMENT_MODEL_DESIGN.md`. ## Why The section used the embedding gather as its worked example of the placement lever ("moving 389,283,840 B to produce ~10 KB, ~78 ms/token to do no work"). The number is right, the tensor is not — anyone reading it reaches the same conclusion that greenlit a dead build. #1299 records the finding but an issue does not stop a reader of the doc; this fixes the passage itself. ## What - **Tensor identity corrected.** Verified on `qwen14b-zp` (RTX 4060): `model.embed_tokens.qweight` is consumed by `GatherBlockQuantized`, which is **not** a `LazyWeightBoundary`, so it is never paged — it is **resident** and streams nothing per token (of 867 lazy-weight handles, zero are named `embed_tokens`). The 389,283,840 B (`152064 x 5120 x 0.5`, INT4) streaming at key 919 (#945) is `lm_head.weight` via `MatMulNBits` — the vocab projection, which does real work. The gather remains a correct *statement* of the `F ~ 0` principle; it is simply already resident, so there is nothing to move. - **lm_head GEMV result folded in (#1013, x86, info-only reference):** the CPU int4 `MatMulNBits` kernel peaks at ~0.78 GB/s (scalar, no SIMD), so CPU `lm_head` is ~500 ms vs ~33 ms on GPU — the criterion **inverts**; host-placing `lm_head` is a net loss with the current kernel. Remaining gap: the `mlas`/`accuracy_level=4` prepacked path #1013 did not measure. - **Native CUDA-graph capture interaction resolved (#1300):** a per-token device→host→device excursion is compatible with native graph capture **only as an eager seam between captured segments** (token-exact) and **illegal inside** an active capture; seam price ~45–90 µs/token on the 4060. Native path only — **#982 (plugin-EP interspersed-partition hang) is untouched and must not be read as cleared.** Preserves the placement principle and the doc's measured/hardware/model honesty convention; corrects only the tensor identity and the now-known results. Docs-only change. Relates to #1299, #1013, #1300, #994. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
|
For reference, to save possible duplicate work: #1021 merged 2026-08-16 as commit This PR's branch ( |
The doc block quoted one box's DRAM ceiling as if it were the ceiling. Attribute it, add the 32-core figures, and state the rule the probe exists to serve: pair it with roofline_bandwidth on the same box in the same session before a ratio becomes a placement decision. The kernel it drives got ~4x faster between this PR opening and merging, on one host -- the number was never the deliverable, the ability to re-take it is. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Opus review flagged two claims that read wider than the code: the opening sentence and --help said 'the real CPU MatMulNBits decode kernel' when the probe specifically drives the #979 borrowed zero-copy path, and the feature comment said 'no engine' when onnx-genai-engine is a non-optional dep (only its features are off). Both are doc-only, but this PR's whole premise is that a stated number matches what was measured. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…d MatMulNBits prefill (#959, #1091) (#1176) ## Summary On the **default (`mlas` OFF) build** the `dequant-kn` prefill phase now goes to **zero** for every `MatMulNBits` path that reaches the dense fallback: the native CPU EP gets a transposed-B ("NT") SGEMM and non-MLAS prefill routes through it, reusing the already-cached contiguous `Nk` weight instead of materializing a second, transposed `Kn` copy. Refs #959, #1091. **Does not close #959** — the superlinear per-token *decode* cost (~102 s/token at 14B) is a separate open question this does not touch. This closes the *prefill* `Kn`-materialization term. ## The problem (#959) Int4/int8 `MatMulNBits` prefill (`m > 1`) on the default build dequantized the weight to f32 a **second** time in the transposed `Kn` (`[k, n]`) layout — a strided-scatter transpose (each K step written at stride N), **uncached**, then a dense NN GEMM. #959 measured it at **~2.9x** the contiguous `Nk` pass and *degrading with N* (22 s/GB at 0.5B → 38 s/GB at 14B; **266 s of ~357 s** time-to-first-token on qwen2.5-14b). MLAS hosts already avoided this (`try_prefill_mlas_nt`) by feeding the cached `Nk` weight to MLAS's cache-tiled `sgemm` with `trans_b`. The default build had **no** transposed-B GEMM, so its `#[cfg(not(feature = "mlas"))]` sibling returned `false` and fell back to the slow `Kn` path. ## What changed MLAS's advantage here is not magic — it is a packer that reads the B operand row-wise from `[n, k]`. Ported natively: - **`x86_sgemm.rs`** — `sgemm_simd_nt(a, b_nk, c, m, k, n)` computing `C[m,n] = A[m,k] · B_nk[n,k]^T`. It reuses the entire packed GEBP path (`pack_a`, KC/strip blocking, `micro_6x16`) and adds **`pack_b_nt`**, which gathers each output column from its *contiguous* `b_nk` row (unit stride over `k`) into the L1-resident pack tile at stride `NR`. The transpose becomes a **pack-time reshape of an in-cache tile**, not a full-array stride-`N` scatter — exactly why MLAS wins. - **`matmul.rs`** — `nt_gemm_supported(backend)` is the **single** legality predicate (no two `cfg` arms to drift); `gemm_nt_with_backend` dispatches `Mlas → trans_b`, `SimdX86 → sgemm_simd_nt`. - **`matmul_nbits.rs`** — the two cfg-split `try_prefill_mlas_nt` arms collapse into one **`try_prefill_nk_nt`** gated on `nt_gemm_supported`. It dequantizes once into the pre-existing `weight_nk` `OnceLock` (the same slot decode caches into, so a constant weight pays **one** dequant, not two) and calls `gemm_nt_with_backend`. When the NT route runs, the `dequant-kn` profile phase is skipped by construction. - **Host pool (#1143)** — column strips dispatch onto the installed ORT intra-op pool when present, rather than forking a rayon pool beside it; rayon otherwise. Strip decomposition is numerically transparent. ## Correctness — bit-identical **Byte-for-byte identical** to the existing `Kn` dense route, not "within 1e-5". `pack_b_nt` produces the *same* packed panels `pack_b` would from the `[k, n]` transpose (`b_kn[p·n+j] == b_nk[j·k+p]`); with identical A-pack, K-panel order, and microkernel, every output element's f32 accumulation sequence is unchanged. Strip count and pool choice do **not** affect per-element reduction order, so bit-identity holds regardless. Verified (`assert_eq!` on `f32::to_bits`): - Odd/tail shapes: `m ∈ {1, 2, 7, 33}`, `n ∈ {1, 3, 63, 64, 65}`, plus tile-exact and multi-KC. - 64 randomized shapes (differential NN-vs-NT). - End-to-end by the existing 8-bit prefill oracle test (`matmulnbits_8bit_prefill_batched_matches_dequant_f32_oracle`), which reaches the dense fallback and thus the NT route. This matches the standard the MLAS NT route already held itself to (bit-identity to the no-transpose dense GEMM, recorded at `matmul_nbits.rs` ~2048). ## Measurement **`dequant-kn` → 0 is structural.** The `mm_profile::time_prepack("dequant-kn", …)` call lives in the `if !used_fast_nt {…}` branch; whenever the NT route returns `true` that branch is skipped, so the phase is eliminated by construction. `dequant-nk` is unchanged (still one pass, now the only one). **Finding vs the #959 premise:** on current local builds, 4-bit `acc0` prefill takes the borrowed-int4 in-place path (#979/#1117) and 4-bit `acc4` uses the SDOT prepack — *neither reaches the dense fallback*, so no local q4 model emits a `[mm_prepack] phase=dequant-kn` line to drive to zero end-to-end (confirmed empirically with `ONNX_GENAI_PROFILE_MM=1` on qwen05b q4 / q4-acc4 / symzp). The beneficiaries of this change are therefore: **8-bit** weights (`m>1`), **grouped** quantization, **weight_prepacked**, and **4-bit with `accuracy_level != 0`** that falls to the dense fallback. No local 8-bit/grouped model was available for an end-to-end token-to-first-token arm; per the profiling skill I do not report a contended wall-clock figure I cannot defend. **Per-phase microbench** (`nt_prefill_bench`, `#[ignore]`; best-of-7, `--test-threads=1`, release; **contended box — other agents building concurrently**). Bit-identity asserted in the same harness. The `dequant-kn*` arm is a *plain f32 transpose* standing in for the strided `Kn` materialization the NT route removes (the real int4 dequant is ~2.9x heavier per #959): | shape (m=16) | dequant-kn* (transpose) | NN gemm | NT gemm | old (kn*+NN) → new (NT) | |---|---|---|---|---| | k=5120, n=5120 (100 MiB) | 337 ms | 5.7 ms | 5.0 ms | 343 ms → **5.0 ms** | | k=5120, n=13824 (270 MiB) | 641 ms | 15.0 ms | 11.8 ms | 656 ms → **11.8 ms** | | k=13824, n=5120 (270 MiB) | 1307 ms | 15.3 ms | 11.7 ms | 1323 ms → **11.7 ms** | The eliminated transpose term dominates and grows with size (as #959 predicted); the NT GEMM itself is even **slightly faster than NN** here (contiguous per-column B reads pack better). Per #1132, a native-faster-than-MLAS result on some shape is a graduation event for `benches/native_vs_mlas.rs` — noting it, but **not** changing default routing without that gate's measurement. **RSS.** No new long-lived allocation (reuses `weight_nk`; `apack`/`bpack` are per-call local scratch, freed at return). Removing the second full f32 weight materialization cuts the transient f32 footprint of a prefill that hits the dense fallback by one full `[k, n]` copy (e.g. **270 MiB** at k=13824/n=5120). Not measured end-to-end (no local model reaches that path); the reduction is structural. ## Memory rules No new field that outlives a call or scales with weight size — the NT route reuses the pre-existing `weight_nk` `OnceLock`. `apack`/`bpack` are per-call `vec![]` scratch. The added lines are not matched by `weight-cache-guard.yml` (no `OnceLock<…Vec>` / `(RefCell|Cell)<Vec|Box|Arc>` introduced); its regex and path filter are untouched. ## Gates (exact counts) - `cargo test -p onnx-runtime-ep-cpu --lib` → **1324 passed, 0 failed, 17 ignored** - `cargo test -p onnx-runtime-ep-cpu --lib --features mlas` → **1354 passed, 0 failed, 28 ignored** (MLAS route still works and still wins where enabled) - `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings` → **clean** - `cargo fmt --check` → the three changed files are clean (verified with `rustfmt --edition 2024 --check`). Pre-existing diffs remain in three *unrelated* files (`governed_accumulator_budget.rs`, `qlinear_matmul.rs`, `simd_activations.rs`) from a local rustfmt version skew vs CI — left untouched to keep this PR surgical. --- 🤖 Generated with Squad. Flagged **needs review** — please have a squad member review the kernel packing/tail handling before merge. --- ## Update (2026-08-18, Roy) — merged current `main`, plus a production-path A/B ### Merge Merged `main` (`c55a3fab3`), which had since made the #1091 M=1 GEMV the unconditional `SimdX86` route and dropped the `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` toggle this branch still carried. Conflict resolved by keeping **both**: main's default-route test (`the_default_entry_point_routes_m1_to_the_gemv`) and this branch's NT kernel + bit-identity tests. Two follow-on fixes: - **aarch64 cross-arch lane.** `gemm_nt_with_backend` compiled with neither the `mlas` nor the x86 arm has no reader for any parameter, so the `-D warnings` cross-arch pass rejected all six. Bound them in the unsupported arm. (This is what the old `Rust quality` red was: `Cross-target compile check`, nothing else.) - **`mm_profile` gemv phase.** The MLAS route used to time this GEMM on the `gemv` phase; after the two call sites collapsed into one, that timer was lost. Restored — the default build gets a pass-through (`tick()`, the reporter, is MLAS-only), so the shared call site stays `cfg`-free and MLAS profiling is unchanged. ### Production-path A/B (new harness, `benches/matmul_nbits_prefill_ab.rs`) The original body was right that no local model reaches this route, and honest about not reporting a number it could not defend. That gap is now closed the way #1013 closed its own: drive the **real kernel through the EP's own `get_kernel`/`execute`**, at the shapes and inputs that *do* reach the dense fallback — 8-bit prefill, and 4-bit with `g_idx`. The harness uses no symbol this branch introduces, so the identical file runs on `main` and here; both arms were built and run **interleaved**, 3 repetitions each, on the same box. Host: 32-core x86_64, AVX2, default build (**`mlas` off**), release. **Contended** (other agents building; load ~12), so medians of per-arm medians are reported and the ratios — not the absolute ms — are the claim. **Steady state** (weight already resident, per-call prefill cost): | case | k | n | m | main (ms) | this PR (ms) | speedup | |---|---:|---:|---:|---:|---:|---:| | int8 dense fallback | 2048 | 2048 | 8 | 12.114 | **0.548** | **22.1x** | | int8 dense fallback | 2048 | 2048 | 64 | 12.276 | **1.375** | **8.9x** | | int8 dense fallback | 4096 | 11008 | 8 | 53.858 | **4.170** | **12.9x** | | int8 dense fallback | 4096 | 11008 | 64 | 56.560 | **7.595** | **7.4x** | | int4 + `g_idx` fallback | 2048 | 2048 | 8 | 31.583 | **0.563** | **56.1x** | | int4 + `g_idx` fallback | 2048 | 2048 | 64 | 32.405 | **1.419** | **22.8x** | **Cold** (fresh kernel per repetition, so the one-time weight dequant is inside the measurement — the TTFT term #959 attacked): | case | k | n | m | main (ms) | this PR (ms) | speedup | |---|---:|---:|---:|---:|---:|---:| | int8 dense fallback | 2048 | 2048 | 8 | 6.470 | 6.008 | 1.08x | | int8 dense fallback | 2048 | 2048 | 64 | 12.519 | 7.974 | 1.57x | | int8 dense fallback | 4096 | 11008 | 8 | 53.761 | 37.260 | 1.44x | | int8 dense fallback | 4096 | 11008 | 64 | 57.725 | 38.503 | 1.50x | | int4 + `g_idx` fallback | 2048 | 2048 | 8 | 31.636 | 23.200 | 1.36x | | int4 + `g_idx` fallback | 2048 | 2048 | 64 | 32.083 | 25.883 | 1.24x | **Why steady moves 7–56x and cold only ~1.1–1.6x — and why that is the real finding.** The `Kn` route has **no cache**: `dequantize_weight(WeightLayout::Kn)` is called *inside* `execute`, so every prefill call re-materializes the whole transposed f32 weight. The NT route dequantizes into the pre-existing `weight_nk` `OnceLock`, which a constant weight fills **once**. So this change removes not one transpose but *every repeat of it*. Cold (first call) improves by the layout alone — a contiguous `Nk` write instead of the stride-`N` scatter, 1.2–1.6x at these sizes; steady improves by the caching the `Nk` layout makes possible. Confirmed structurally with `ONNX_GENAI_PROFILE_MM=1` over the same harness run: **main emits 72 `phase=dequant-kn` lines, this PR emits 0 — and 24 `phase=dequant-nk`** (one per kernel instance, i.e. the cold arms only; every steady call pays none). **Bit-identity, across builds.** The harness prints an FNV-style digest of the raw output bits. All six rows have the **identical digest on both arms** (`db6ff07f991d431`, `b4c1df9fd8883789`, `7bd7418eacdcb870`, `2eed012609649617`, `d5834afbecf42494`, `cd8fa77262302969`) — bit-identity of the production `execute` result, not just of the kernel driver, verified across two separately compiled builds. ### Scope, restated honestly On the default build, 4-bit `accuracy_level=0` with contiguous (borrowable) inputs takes the zero-copy borrowed int4 path (#979/#1117/#1126) for both decode *and* prefill and returns before the dense fallback. This PR therefore changes: **8-bit prefill**, **4-bit with `g_idx`**, **`weight_prepacked`/non-borrowable** inputs, and **4-bit `accuracy_level != 0`** that falls through. Those are exactly the cases measured above. MLAS builds already had the NT route; their behaviour is unchanged. ### Gates (re-run after the merge) - `cargo test -p onnx-runtime-ep-cpu --lib` → **1420 passed, 0 failed, 18 ignored** - `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings` → clean - `cargo clippy --locked --target aarch64-unknown-linux-gnu --all-targets -p onnx-runtime-ep-cpu -- -D warnings` → clean (the lane that was red) - `cargo fmt --all -- --check` → the files this PR touches are clean; one *inherited* diff remains in `onnx-runtime-ep-cuda/standard_attention.rs` from `main`, fixed separately in #1347. --------- Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com> Co-authored-by: Roy <roy@squad.local>
Review caught that an empty sample slice underflows median()'s mid - 1. Guard it alongside the existing --threads and --block-size checks. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Retraction posted, re-measured on latest main, Opus-reviewed, local gate greenThe PR description has been rewritten. The original "false by ~63×" conclusion is retracted — it measured the scalar x86 borrowed int4 GEMV that #1021 ( Re-measured on
27.2 of 75.7 GB/s — ~2.8× below roofline, not ~63×. Per-token Merging it as an instrument, not for its conclusion. It is the probe that caught the error, and the only one in the tree that measures the real quantised decode GEMV against the real DRAM ceiling on the same box. ReviewOpus review verified the instrument itself rather than the claim, and confirmed by reading the dispatch code that it reaches the path it names: with One non-blocking finding, fixed in Local gate matrix on
|
What
Adds
roofline_gemv, a CPU-only bench probe (behind a leangemv-probefeature) that measures the effective memory bandwidth the real CPU int4MatMulNBitsdecode kernel achieves at the 14Blm_headshape — the last unmeasured term in #994's placement criterionF/C_slow < W/B_link + F/C_fast.It is the compute-side companion to
roofline_bandwidth(host DRAM STREAM ceiling) androofline_transfer(host↔device link).This PR originally reported that the CPU
lm_headGEMV runs at ~0.78 GB/s — 1.6% of the DRAM ceiling — making CPU compute ~500 ms and inverting #994's placement criterion by ~63×. That result does not reproduce and is retracted.Two things were wrong with it:
3cf49d25, merged 2026-08-16) vectorised that GEMV for AVX2/AVX-512, and perf(cpu-ep): fuse the int4 dequant into the prefill GEMM pack step (23.7x) #1356/Decode int4 nibbles once per row tile in MatMulNBits prefill #1431 have since reworked the prefill dispatch around it. The number was real when taken and is meaningless now.Re-measured on
origin/main(085c49140), same box, same probeDRAM read ceiling,
roofline_bandwidth --mib 2048 --seconds 3:The int4
lm_headGEMV,roofline_gemv --repeats 7(K=5120, N=152064, block_size=32, bits=4, accuracy_level=0; weight 389.3 MB + scales 97.3 MB = 487.2 MB total traffic):Corrected result: the kernel reaches 27.2 GB/s of 75.7 GB/s — it is ~2.8× below the roofline, not ~63×. Per-token
lm_headcompute is ~18–20 ms, not ~500 ms.What that does to #994's criterion
The criterion no longer inverts. At ~18–20 ms of CPU compute against ~33 ms to move the work to the GPU, CPU
lm_headis cheaper, not 15× more expensive — the opposite of what this PR originally concluded. #994's original assumption (that the GEMV runs at DRAM bandwidth, ~7.9 ms) is still optimistic by ~2.5×, but it is the right order of magnitude and the placement decision it drives stands.The remaining 2.8× is the honest open question this probe exists to track: the kernel still widens every nibble to f32 and streams 97 MB of f32 scales alongside 389 MB of packed weights, so it is not purely bandwidth-limited even after vectorisation.
Why this is still worth merging
The conclusion was wrong; the instrument was not.
roofline_gemvis what caught the error, and it is the only probe in the tree that measures the real quantised decode GEMV against the real DRAM ceiling on the same box. Merging it makes that measurement repeatable instead of a claim in a PR description.How it stays faithful
get_kernel/execute— no reimplementation.Nonezero points).with_decode_pool_scopeso eachexecutedispatches to the persistent SPMD decode pool the runtime actually uses — not a per-call fork/join of the flat pool (which understated bandwidth ~10×).set_decode_thread_budget.Scope
Measurement, not implementation. No kernels are restructured.
mlas/accuracy_level=4(MLAS SQNBit) deliberately not measured — the default artifact carries zero MLAS symbols by design, so it is not the path the runtime takes.