Skip to content

Add roofline_gemv: measure the real CPU int4 lm_head GEMV against DRAM (#994) — re-measured, ~3x below roofline, not ~63x - #1013

Merged
justinchuby merged 7 commits into
mainfrom
squad/994-cpu-gemv-roofline
Aug 19, 2026
Merged

justinchuby merged 7 commits into
mainfrom
squad/994-cpu-gemv-roofline

Conversation

@justinchuby

@justinchuby justinchuby commented Aug 15, 2026 •

Copy link
Copy Markdown
Owner

What

Adds roofline_gemv, a CPU-only bench probe (behind a lean gemv-probe feature) that measures the effective memory bandwidth the real CPU int4 MatMulNBits decode kernel achieves at the 14B lm_head shape — the last unmeasured term in #994's placement criterion F/C_slow < W/B_link + F/C_fast.

It is the compute-side companion to roofline_bandwidth (host DRAM STREAM ceiling) and roofline_transfer (host↔device link).

⚠️ Retraction: the original "false by ~63×" claim was wrong

This PR originally reported that the CPU lm_head GEMV runs at ~0.78 GB/s — 1.6% of the DRAM ceiling — making CPU compute ~500 ms and inverting #994's placement criterion by ~63×. That result does not reproduce and is retracted.

Two things were wrong with it:

  1. It measured a kernel that no longer exists. The probe was written against the Symmetric int4 has no zero-copy CPU kernel, so it pays 8x RAM: adding a constant zero_points=8 cuts footprint 6.1x with byte-identical output #979 borrowed symmetric-int4 path when that path was a scalar unpack-convert-FMA loop on x86 (only AArch64 had the NEON dot). cpu: vectorise x86 borrowed int4 GEMV (AVX2/AVX-512) for #994 #1021 (3cf49d25, merged 2026-08-16) vectorised that GEMV for AVX2/AVX-512, and perf(cpu-ep): fuse the int4 dequant into the prefill GEMM pack step (23.7x) #1356/Decode int4 nibbles once per row tile in MatMulNBits prefill #1431 have since reworked the prefill dispatch around it. The number was real when taken and is meaningless now.
  2. ~63× was never the right way to state it even then. The comparison mixed a weight-only rate against a total-traffic ceiling and reported a single best sample from a heavily shared box.

Re-measured on origin/main (085c49140), same box, same probe

DRAM read ceiling, roofline_bandwidth --mib 2048 --seconds 3:

threads GB/s
8 75.099
16 75.460
32 75.737

The int4 lm_head GEMV, roofline_gemv --repeats 7 (K=5120, N=152064, block_size=32, bits=4, accuracy_level=0; weight 389.3 MB + scales 97.3 MB = 487.2 MB total traffic):

threads per-call ms (median) per-call ms (min) weight GB/s (best) total GB/s (best)
8 52.560 52.330 7.439 9.311
16 29.349 26.388 14.752 18.464
32 19.785 17.936 21.705 27.166

Corrected result: the kernel reaches 27.2 GB/s of 75.7 GB/s — it is ~2.8× below the roofline, not ~63×. Per-token lm_head compute is ~18–20 ms, not ~500 ms.

What that does to #994's criterion

The criterion no longer inverts. At ~18–20 ms of CPU compute against ~33 ms to move the work to the GPU, CPU lm_head is cheaper, not 15× more expensive — the opposite of what this PR originally concluded. #994's original assumption (that the GEMV runs at DRAM bandwidth, ~7.9 ms) is still optimistic by ~2.5×, but it is the right order of magnitude and the placement decision it drives stands.

The remaining 2.8× is the honest open question this probe exists to track: the kernel still widens every nibble to f32 and streams 97 MB of f32 scales alongside 389 MB of packed weights, so it is not purely bandwidth-limited even after vectorisation.

Why this is still worth merging

The conclusion was wrong; the instrument was not. roofline_gemv is what caught the error, and it is the only probe in the tree that measures the real quantised decode GEMV against the real DRAM ceiling on the same box. Merging it makes that measurement repeatable instead of a claim in a PR description.

How it stays faithful

  • Drives the real kernel via the CPU EP's own get_kernel/execute — no reimplementation.
  • Confirms the borrowed symmetric-int4 dispatch is taken (bits=4, accuracy_level=0, no prepack/g_idx, symmetric → None zero points).
  • Runs warmup+timing inside one with_decode_pool_scope so each execute dispatches to the persistent SPMD decode pool the runtime actually uses — not a per-call fork/join of the flat pool (which understated bandwidth ~10×).
  • Sweeps thread counts one process per count via set_decode_thread_budget.
  • Reports both a weight-only rate (the Placement as a memory lever: run low-arithmetic-intensity ops where their weights already live, instead of moving the weights #994 criterion input) and a total-traffic rate (honest roofline efficiency, incl. the f32 scales), and a distribution rather than one sample.

Scope

Measurement, not implementation. No kernels are restructured. mlas/accuracy_level=4 (MLAS SQNBit) deliberately not measured — the default artifact carries zero MLAS symbols by design, so it is not the path the runtime takes.

Adds a CPU-only bench binary (
oofline_gemv, behind a lean gemv-probe
feature) that drives the real symmetric-int4 MatMulNBits decode kernel -- the
#979 borrowed zero-copy path (accuracy_level=0) -- at the 14B lm_head shape
(K=5120, N=152064, block_size=32, bits=4, m=1) and measures the effective
memory bandwidth it achieves against the STREAM DRAM ceiling.

The probe runs warmup+timing inside a single with_decode_pool_scope so each
execute dispatches to the persistent SPMD decode pool the runtime actually uses
(not a per-call fork/join of the flat pool), and sweeps thread counts one
process per count via set_decode_thread_budget. It reports both a weight-only
rate (the #994 criterion input) and a total-traffic rate (honest roofline
efficiency including the ~97 MB of f32 scales).

Measurement, not implementation: no kernels are restructured. Answers the last
unverified term in #994's placement criterion.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 15, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.66%. Comparing base (4a9f4ec) to head (f3bd8a6).
⚠️ Report is 81 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@             Coverage Diff             @@
##             main    #1013       +/-   ##
===========================================
- Coverage   82.10%   80.66%    -1.45%     
===========================================
  Files          12      377      +365     
  Lines        5471   166791   +161320     
  Branches     5471   166791   +161320     
===========================================
+ Hits         4492   134542   +130050     
- Misses        780    27406    +26626     
- Partials      199     4843     +4644     
Flag Coverage Δ
cli-ort-linux 82.60% <ø> (?)
cli-ort-windows 82.19% <ø> (+0.09%) ⬆️
offline 80.59% <ø> (?)

Flags with carried forward coverage won't be shown. Click here to find out more.
see 365 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actions

github-actions Bot commented Aug 15, 2026 •

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/medium_generic_f16_threads=8/32x512x512 48.36 µs 117.83 µs +143.6%
🔴 matmul/medium_generic_bf16_threads=8/32x512x512 610.04 µs 1.44 ms +136.1%
🔴 matmul/small_generic_f32_threads=8/1x256x256 58.27 µs 114.35 µs +96.3%
🔴 matmul/small_generic_bf16_threads=8/1x256x256 43.26 µs 72.21 µs +66.9%
🔴 matmul/small_generic_f16_threads=8/1x256x256 40.85 µs 62.34 µs +52.6%
🔴 gather/large_bf16_threads=1-internal/131072 14.79 µs 22.51 µs +52.2%
🔴 matmul/medium_generic_bf16_threads=1/32x512x512 637.57 µs 933.52 µs +46.4%
🔴 sampling_latency/greedy_per_token 3.33 µs 4.57 µs +37.1%
⚠️ qwen3_sampling_processors/top_p_fast_after_top_k 529.81 µs 646.69 µs +22.1%
⚠️ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.64 ms 4.28 ms +17.6%
⚠️ matmul/large_generic_f32_threads=1/32x1024x1024 11.40 ms 13.35 ms +17.2%
✅ add/large_f32_threads=1-internal/4194304 646.59 µs 741.16 µs +14.6%
✅ qwen3_sampling_processors/top_k_partial_selection 147.11 µs 167.37 µs +13.8%
✅ kv_cache/alloc_dealloc_pages 38.59 µs 42.44 µs +10.0%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.93 ms 6.47 ms +9.0%
✅ grammar_masking/llguidance_compute_mask/32 77.76 µs 84.77 µs +9.0%
✅ add/small_bf16_threads=1-internal/1024 435.2 ns 471.3 ns +8.3%
✅ gather/medium_bf16_threads=1-internal/32768 2.75 µs 2.96 µs +7.7%
✅ add/medium_f16_threads=1-internal/262144 103.27 µs 111.08 µs +7.6%
✅ logit_processing/seven_processor_chain_per_step 328.46 µs 350.31 µs +6.7%
✅ add/medium_bf16_threads=1-internal/262144 102.37 µs 108.18 µs +5.7%
✅ qwen3_sampling_processors/top_k_top_p_fast 678.91 µs 708.19 µs +4.3%
✅ gather/medium_f16_threads=1-internal/32768 2.81 µs 2.93 µs +4.2%
✅ sampling_latency/min_p_per_token 229.55 µs 236.07 µs +2.8%
✅ sampling_latency/top_p_per_token 424.39 µs 434.89 µs +2.5%
✅ matmul/small_generic_f32_threads=1/1x256x256 47.55 µs 48.54 µs +2.1%
✅ add/small_f16_threads=1-internal/1024 474.0 ns 483.5 ns +2.0%
✅ add/medium_f32_threads=1-internal/262144 25.15 µs 25.52 µs +1.5%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.23 ms 2.25 ms +0.9%
✅ matmul/medium_generic_f16_threads=1/32x512x512 44.05 µs 44.29 µs +0.5%
✅ gather/small_bf16_threads=1-internal/4096 527.1 ns 528.2 ns +0.2%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 111.79 µs 110.99 µs -0.7%
✅ gather/small_f32_threads=1-internal/4096 773.5 ns 764.0 ns -1.2%
✅ gather/large_f16_threads=1-internal/131072 17.11 µs 16.89 µs -1.3%
✅ tokenization/encode_tokens_per_second 435.65 µs 429.13 µs -1.5%
✅ gather/medium_f32_threads=1-internal/32768 4.75 µs 4.67 µs -1.8%
✅ sampling_latency/top_k_per_token 61.71 µs 60.32 µs -2.3%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.13 ms 1.09 ms -2.8%
✅ gather/large_f32_threads=1-internal/131072 40.90 µs 39.03 µs -4.6%
✅ add/small_f32_threads=1-internal/1024 215.7 ns 205.4 ns -4.7%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 609.67 µs 577.00 µs -5.4%
✅ gather/small_f16_threads=1-internal/4096 572.6 ns 538.4 ns -6.0%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 130.80 µs 122.47 µs -6.4%
✅ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 1.19 ms 1.12 ms -6.4%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 102.37 µs 95.05 µs -7.1%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.86 ms 2.63 ms -8.2%
✅ tokenization/decode_tokens_per_second 8.65 ms 7.84 ms -9.4%
✅ reduce_mean/medium_f32_threads=1-internal/65536 306.03 µs 274.16 µs -10.4%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.32 ms 2.07 ms -10.6%
✅ matmul/small_generic_bf16_threads=1/1x256x256 39.90 µs 34.81 µs -12.8%
✅ add/large_bf16_threads=1-internal/4194304 1.96 ms 1.70 ms -13.3%
✅ matmul/small_generic_f16_threads=1/1x256x256 38.99 µs 33.59 µs -13.9%
✅ add/large_f16_threads=1-internal/4194304 2.02 ms 1.73 ms -14.1%
✅ reduce_mean/small_f32_threads=1-internal/4096 18.72 µs 16.06 µs -14.2%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 6.37 ms 5.43 ms -14.8%
🟢 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 105.14 µs 80.66 µs -23.3%
🟢 matmul/large_generic_bf16_threads=8/32x1024x1024 2.30 ms 1.69 ms -26.4%
🟢 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 111.45 µs 81.22 µs -27.1%
🟢 matmul/medium_generic_f32_threads=8/32x512x512 1.66 ms 1.08 ms -35.2%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 4.42 4.45 5.16 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Aug 16, 2026
## Vectorise the x86 borrowed int4 GEMV (refs #994, #979)

The x86 symmetric/asymmetric int4 **borrowed** decode path (unified by
#979) was a
scalar unpack-convert-FMA loop; only aarch64 had a vectorised route
(`borrowed_affine_int4_matmul_m1_neon_dot` /
`affine_int4_block32_dot_neon`). At the
14B `lm_head` shape it measured **~0.78 GB/s, ~1.6% of this box's 49.28
GB/s STREAM
ceiling** (per the #1013 `roofline_gemv` probe).

### What changed
- New runtime-dispatched **f32 SIMD block dot** for the borrowed path,
reusing the
existing `selected_dot_kernel()` / `DotKernel` seam (no new mechanism,
no
  compile-time target-feature gate):
- `borrowed_int4_block_dot_avx2` — AVX2+FMA, serves `Avx2` **and**
`AvxVnni` hosts
    (the borrowed path is f32, so integer VNNI does not apply).
- `borrowed_int4_block_dot_avx512` — AVX-512F, serves `Avx512Vnni`
hosts.
- Both unpack the 4-bit nibbles into the scalar loop's natural
`w[2i]=lo(byte i), w[2i+1]=hi(byte i)` order, widen to f32, and
multiply-accumulate
the activation per 32-lane tile. The per-block scale and the zero-point
correction
stay in the caller, so **symmetric** (`zero_points=None`, implicit
midpoint 8) and
**asymmetric** both work unchanged. Blocks larger than 32 (e.g.
block-128) vectorise
  via the 32-lane chunk loop into a single accumulator.
- Nothing dequantised outlives the call → the #979 zero-copy footprint
is preserved.
- `ONNX_GENAI_CPU_DISABLE_INT4_SIMD=1` is a documented A/B escape hatch
to the scalar
reference (same-binary A/B, `OnceLock`-resolved, mirrors
`arm64_int4_direct_enabled`).
- **aarch64 is untouched** except that the shared `let _ = dot_kernel;`
unused-var
guard is now `#[cfg(not(any(aarch64, x86_64)))]` (aarch64 already used
`dot_kernel`;
  behaviour there is byte-for-byte the same).

### 1. Roofline (`roofline_gemv`, K=5120 N=152064 block_size=32 bits=4,
borrowed symmetric int4)
Weight-only GB/s (criterion input). Box: shared **i7-13800H** (14C/20T,
AVX2+AVX-VNNI,
**no AVX-512** — so the AVX2 kernel is what runs here; the AVX-512
kernel is compiled
but unexercised on this host). `iters=10 repeats=8`; two other squad
agents active, so I
report median **and** best (quietest) rather than a single sample.

| threads | before med / best | after med / best | after % of 49.28 |
speedup (best) |

|--------:|------------------:|-----------------:|-----------------:|---------------:|
| 1  | 0.390 / 0.449 | 1.048 / 1.661 | 3.4% | 3.7× |
| 2  | 0.604 / 0.700 | 1.946 / 2.217 | 4.5% | 3.2× |
| 4  | 0.505 / 0.882 | 3.682 / 3.944 | 8.0% | 4.5× |
| 8  | 0.646 / 0.777 | 4.551 / 5.219 | 10.6% | 6.7× |
| 16 | 0.768 / 1.214 | 6.176 / 6.519 | 13.2% | 5.4× |
| 20 | 1.021 / 1.186 | 6.711 / 8.133 | 16.5% | 6.9× |

Headline (20 threads, quiet): **~8.1 GB/s best / ~6.7 GB/s median,
~16.5% of the
49.28 GB/s ceiling**, up from ~0.78 GB/s. **~6.7–6.9× at the useful
thread counts.**

**Honest-negative note:** the scalar loop *was* a real bottleneck (~7×
confirms it), but
we are still **~6× below roofline** (16.5%, not ~100%). The residual gap
is now
elsewhere — the +25% scales traffic the criterion omits, the strided
per-row weight
reads, and the per-`execute` SPMD barrier — not the nibble unpack. There
is more CPU
headroom, but it is no longer in this loop.

### 2. Byte-identical output
96 greedy tokens, native CPU EP, SIMD on vs
`ONNX_GENAI_CPU_DISABLE_INT4_SIMD=1`:
- `qwen05b-q4` (symmetric, borrowed path): **token IDs byte-identical**
(all 96).
- `qwen2.5-0.5b-q4_0-mobius` (asymmetric, same path): **token IDs
byte-identical** (all 96).

This is **not** a bit-identical f32 transform: the horizontal reduction
+ FMA reorder
additions vs the scalar loop, so logits differ by a few ULP. That
divergence is below
the argmax margin here — tokens are byte-identical on both models — but
I am stating it
rather than claiming bit-identity. The unit test locks the block dot to
the scalar
reference within `|Δ| ≤ |ref|·1e-5 + 1e-4`.

### 3. Footprint unchanged
Peak working set on `qwen05b-q4`: **383.0 MiB** (SIMD) vs 372.2 MiB
(scalar); asymmetric
mobius 380.0 vs 379.5 MiB. No f32 expansion (that would be GBs); the
borrowed tile does
not outlive the call. Well under the ~434 MB envelope.

### 4. aarch64
No functional change. The only aarch64-visible edit is the `let _ =
dot_kernel;` cfg
guard (aarch64 already consumed `dot_kernel`). Cannot be run here;
called out explicitly.

### Gates (verbatim)
- `cargo fmt -p onnx-runtime-ep-cpu` — clean
- `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings` —
clean
- `cargo test -p onnx-runtime-ep-cpu`:
- `unittests src\lib.rs`: `test result: ok. 1078 passed; 0 failed; 10
ignored`
  - `bf16_conformance`: `ok. 3 passed; 0 failed`
  - `kernel_numeric_regression`: `ok. 10 passed; 0 failed`
  - `mha_ort_parity`: `ok. 1 passed; 0 failed`
  - `msft_attention_ort_parity`: `ok. 1 passed; 0 failed`
  - `qwen35_ort_parity`: `ok. 5 passed; 0 failed`
  - `shared_allocator`: `ok. 6 passed; 0 failed`
- `cargo test -p onnx-genai-engine --lib`: `test result: ok. 385 passed;
0 failed; 1 ignored`

(baseline lib was 1077; +1 is the new
`borrowed_int4_block_dot_x86_matches_scalar_reference` test.)

### #994 criterion re-evaluation
Posted on #994. Short version, with the post-fix number:
- move to GPU: 389 MB / 11.74 GB/s (H2D pinned) = **~33 ms**
- compute on CPU: 389 MB / ~6.7–8.1 GB/s = **~48–58 ms** (20 threads,
median→best)

GPU H2D still edges CPU-compute in absolute per-token latency, so the
placement
argument for `lm_head` **does not come back** — but the margin collapsed
from ~15×
(the old ~500 ms) to ~1.5×, moving it from "clearly dead" into
narrow/ambiguous
territory where the VRAM-savings tradeoff is now defensible. Embedding
gather is
unaffected (arithmetic intensity ~0).

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Built and ran the probe on the RTX 4060 laptop box (20 logical CPUs, AVX2 + FMA + F16C + AVX-VNNI, no AVX-512). It builds and runs cleanly, and the CSV shape is right. But I cannot use it yet, for a reason worth fixing in the probe rather than working around.

The probe reports wall-clock-derived GB/s, and on a shared box that number is not stable enough to answer the question it exists to answer.

Concretely: #1021 has now landed, which vectorises exactly the kernel this probe targets. Measured end-to-end on qwen05b-symzp through the real decode path, in process CPU time, SIMD cuts decode from 4.19 to 1.84 CPU-seconds per token -- a 2.27x reduction, reproducing to ~2% across runs. Running this probe across the same A/B (ONNX_GENAI_CPU_DISABLE_INT4_SIMD=1 vs 0) at its default 14B lm_head shape:

threads SIMD off SIMD on ratio
8 0.608 GB/s median 0.625 GB/s 1.03x
20 0.374 GB/s 0.415 GB/s 1.11x

A 2.27x kernel change registers as 3-11%. The probe also reports 20 threads as slower than 8 (1042 ms vs 640 ms per call), and 0.6 GB/s where I previously hand-measured 2.46 GB/s at this shape. All three are symptoms of the same cause: another job was on the box (a python process holding 13,246 CPU-seconds), and wall-clock GB/s collapses under contention while the A/B difference compresses toward nothing.

I hit this myself an hour earlier and it nearly cost me the #1021 conclusion: three identical runs of one configuration gave 39.3 / 25.8 / 16.1 s wall. Switching to TotalProcessorTime made the same comparison reproduce to 2%.

What I would like changed before merging:

  1. Report CPU time alongside wall time -- per-call CPU-seconds and a CPU-time-derived GB/s, next to the existing wall columns. For a fixed-thread compute kernel this is nearly contention-immune, and it is what makes the probe usable on a machine that is not exclusively ours (which is every machine we actually run on).
  2. Report a contention indicator so a reader can tell a bad sample from a good one -- e.g. the ratio of CPU time to (wall x threads). When that drops well below 1 the sample should be marked untrustworthy rather than silently averaged in. The existing --repeats distribution helps, but a spread does not distinguish "this kernel is variable" from "this box was busy".
  3. Please re-take the headline numbers after cpu: vectorise x86 borrowed int4 GEMV (AVX2/AVX-512) for #994 #1021, and state the SIMD state explicitly in the CSV header. The PR's current 0.78 GB/s / ~1.6% of roofline describes a kernel that no longer exists on x86.

The probe itself is the right idea and I want it in the tree -- reproducible measurement infrastructure is worth more than any single measurement. It just has to be honest about when its own numbers are contaminated, which is the same standard we have been holding the kernels to.

justinchuby added a commit that referenced this pull request Aug 18, 2026
…389 MB is lm_head (#1304)

Corrects the placement worked-example in
`docs/memory/MEMORY_MANAGEMENT_MODEL_DESIGN.md`.

## Why
The section used the embedding gather as its worked example of the
placement lever ("moving 389,283,840 B to produce ~10 KB, ~78 ms/token
to do no work"). The number is right, the tensor is not — anyone reading
it reaches the same conclusion that greenlit a dead build. #1299 records
the finding but an issue does not stop a reader of the doc; this fixes
the passage itself.

## What
- **Tensor identity corrected.** Verified on `qwen14b-zp` (RTX 4060):
`model.embed_tokens.qweight` is consumed by `GatherBlockQuantized`,
which is **not** a `LazyWeightBoundary`, so it is never paged — it is
**resident** and streams nothing per token (of 867 lazy-weight handles,
zero are named `embed_tokens`). The 389,283,840 B (`152064 x 5120 x
0.5`, INT4) streaming at key 919 (#945) is `lm_head.weight` via
`MatMulNBits` — the vocab projection, which does real work. The gather
remains a correct *statement* of the `F ~ 0` principle; it is simply
already resident, so there is nothing to move.
- **lm_head GEMV result folded in (#1013, x86, info-only reference):**
the CPU int4 `MatMulNBits` kernel peaks at ~0.78 GB/s (scalar, no SIMD),
so CPU `lm_head` is ~500 ms vs ~33 ms on GPU — the criterion
**inverts**; host-placing `lm_head` is a net loss with the current
kernel. Remaining gap: the `mlas`/`accuracy_level=4` prepacked path
#1013 did not measure.
- **Native CUDA-graph capture interaction resolved (#1300):** a
per-token device→host→device excursion is compatible with native graph
capture **only as an eager seam between captured segments**
(token-exact) and **illegal inside** an active capture; seam price
~45–90 µs/token on the 4060. Native path only — **#982 (plugin-EP
interspersed-partition hang) is untouched and must not be read as
cleared.**

Preserves the placement principle and the doc's measured/hardware/model
honesty convention; corrects only the tensor identity and the now-known
results.

Docs-only change. Relates to #1299, #1013, #1300, #994.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

For reference, to save possible duplicate work: #1021 merged 2026-08-16 as commit 3cf49d25 ("cpu: vectorise x86 borrowed int4 GEMV (AVX2/AVX-512) for #994"). It replaced the scalar unpack-convert-FMA loop in the borrowed symmetric/asymmetric int4 path with a runtime-dispatched AVX2/AVX-512F f32 block-dot (borrowed_int4_block_dot_avx2 / borrowed_int4_block_dot_avx512 in crates/onnx-runtime-ep-cpu/src/kernels/matmul_nbits.rs), on by default with ONNX_GENAI_CPU_DISABLE_INT4_SIMD=1 as a same-binary escape hatch to the scalar reference.

This PR's branch (squad/994-cpu-gemv-roofline) predates that merge, so the ~0.78 GB/s figure here describes the pre-#1021 (scalar) tree. Recording the fact only.

justinchuby and others added 2 commits August 18, 2026 23:54
The doc block quoted one box's DRAM ceiling as if it were the ceiling.
Attribute it, add the 32-core figures, and state the rule the probe
exists to serve: pair it with roofline_bandwidth on the same box in the
same session before a ratio becomes a placement decision. The kernel it
drives got ~4x faster between this PR opening and merging, on one host --
the number was never the deliverable, the ability to re-take it is.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby justinchuby changed the title Measure real CPU int4 lm_head GEMV bandwidth (#994): claim inverts, ~63x below roofline Add roofline_gemv: measure the real CPU int4 lm_head GEMV against DRAM (#994) — re-measured, ~3x below roofline, not ~63x Aug 19, 2026
justinchuby and others added 2 commits August 19, 2026 00:15
Opus review flagged two claims that read wider than the code: the opening
sentence and --help said 'the real CPU MatMulNBits decode kernel' when the
probe specifically drives the #979 borrowed zero-copy path, and the feature
comment said 'no engine' when onnx-genai-engine is a non-optional dep (only
its features are off). Both are doc-only, but this PR's whole premise is
that a stated number matches what was measured.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 19, 2026
…d MatMulNBits prefill (#959, #1091) (#1176)

## Summary

On the **default (`mlas` OFF) build** the `dequant-kn` prefill phase now
goes to **zero** for every `MatMulNBits` path that reaches the dense
fallback: the native CPU EP gets a transposed-B ("NT") SGEMM and
non-MLAS prefill routes through it, reusing the already-cached
contiguous `Nk` weight instead of materializing a second, transposed
`Kn` copy.

Refs #959, #1091. **Does not close #959** — the superlinear per-token
*decode* cost (~102 s/token at 14B) is a separate open question this
does not touch. This closes the *prefill* `Kn`-materialization term.

## The problem (#959)

Int4/int8 `MatMulNBits` prefill (`m > 1`) on the default build
dequantized the weight to f32 a **second** time in the transposed `Kn`
(`[k, n]`) layout — a strided-scatter transpose (each K step written at
stride N), **uncached**, then a dense NN GEMM. #959 measured it at
**~2.9x** the contiguous `Nk` pass and *degrading with N* (22 s/GB at
0.5B → 38 s/GB at 14B; **266 s of ~357 s** time-to-first-token on
qwen2.5-14b).

MLAS hosts already avoided this (`try_prefill_mlas_nt`) by feeding the
cached `Nk` weight to MLAS's cache-tiled `sgemm` with `trans_b`. The
default build had **no** transposed-B GEMM, so its `#[cfg(not(feature =
"mlas"))]` sibling returned `false` and fell back to the slow `Kn` path.

## What changed

MLAS's advantage here is not magic — it is a packer that reads the B
operand row-wise from `[n, k]`. Ported natively:

- **`x86_sgemm.rs`** — `sgemm_simd_nt(a, b_nk, c, m, k, n)` computing
`C[m,n] = A[m,k] · B_nk[n,k]^T`. It reuses the entire packed GEBP path
(`pack_a`, KC/strip blocking, `micro_6x16`) and adds **`pack_b_nt`**,
which gathers each output column from its *contiguous* `b_nk` row (unit
stride over `k`) into the L1-resident pack tile at stride `NR`. The
transpose becomes a **pack-time reshape of an in-cache tile**, not a
full-array stride-`N` scatter — exactly why MLAS wins.
- **`matmul.rs`** — `nt_gemm_supported(backend)` is the **single**
legality predicate (no two `cfg` arms to drift); `gemm_nt_with_backend`
dispatches `Mlas → trans_b`, `SimdX86 → sgemm_simd_nt`.
- **`matmul_nbits.rs`** — the two cfg-split `try_prefill_mlas_nt` arms
collapse into one **`try_prefill_nk_nt`** gated on `nt_gemm_supported`.
It dequantizes once into the pre-existing `weight_nk` `OnceLock` (the
same slot decode caches into, so a constant weight pays **one** dequant,
not two) and calls `gemm_nt_with_backend`. When the NT route runs, the
`dequant-kn` profile phase is skipped by construction.
- **Host pool (#1143)** — column strips dispatch onto the installed ORT
intra-op pool when present, rather than forking a rayon pool beside it;
rayon otherwise. Strip decomposition is numerically transparent.

## Correctness — bit-identical

**Byte-for-byte identical** to the existing `Kn` dense route, not
"within 1e-5". `pack_b_nt` produces the *same* packed panels `pack_b`
would from the `[k, n]` transpose (`b_kn[p·n+j] == b_nk[j·k+p]`); with
identical A-pack, K-panel order, and microkernel, every output element's
f32 accumulation sequence is unchanged. Strip count and pool choice do
**not** affect per-element reduction order, so bit-identity holds
regardless.

Verified (`assert_eq!` on `f32::to_bits`):
- Odd/tail shapes: `m ∈ {1, 2, 7, 33}`, `n ∈ {1, 3, 63, 64, 65}`, plus
tile-exact and multi-KC.
- 64 randomized shapes (differential NN-vs-NT).
- End-to-end by the existing 8-bit prefill oracle test
(`matmulnbits_8bit_prefill_batched_matches_dequant_f32_oracle`), which
reaches the dense fallback and thus the NT route.

This matches the standard the MLAS NT route already held itself to
(bit-identity to the no-transpose dense GEMM, recorded at
`matmul_nbits.rs` ~2048).

## Measurement

**`dequant-kn` → 0 is structural.** The
`mm_profile::time_prepack("dequant-kn", …)` call lives in the `if
!used_fast_nt {…}` branch; whenever the NT route returns `true` that
branch is skipped, so the phase is eliminated by construction.
`dequant-nk` is unchanged (still one pass, now the only one).

**Finding vs the #959 premise:** on current local builds, 4-bit `acc0`
prefill takes the borrowed-int4 in-place path (#979/#1117) and 4-bit
`acc4` uses the SDOT prepack — *neither reaches the dense fallback*, so
no local q4 model emits a `[mm_prepack] phase=dequant-kn` line to drive
to zero end-to-end (confirmed empirically with `ONNX_GENAI_PROFILE_MM=1`
on qwen05b q4 / q4-acc4 / symzp). The beneficiaries of this change are
therefore: **8-bit** weights (`m>1`), **grouped** quantization,
**weight_prepacked**, and **4-bit with `accuracy_level != 0`** that
falls to the dense fallback. No local 8-bit/grouped model was available
for an end-to-end token-to-first-token arm; per the profiling skill I do
not report a contended wall-clock figure I cannot defend.

**Per-phase microbench** (`nt_prefill_bench`, `#[ignore]`; best-of-7,
`--test-threads=1`, release; **contended box — other agents building
concurrently**). Bit-identity asserted in the same harness. The
`dequant-kn*` arm is a *plain f32 transpose* standing in for the strided
`Kn` materialization the NT route removes (the real int4 dequant is
~2.9x heavier per #959):

| shape (m=16) | dequant-kn* (transpose) | NN gemm | NT gemm | old
(kn*+NN) → new (NT) |
|---|---|---|---|---|
| k=5120, n=5120 (100 MiB) | 337 ms | 5.7 ms | 5.0 ms | 343 ms → **5.0
ms** |
| k=5120, n=13824 (270 MiB) | 641 ms | 15.0 ms | 11.8 ms | 656 ms →
**11.8 ms** |
| k=13824, n=5120 (270 MiB) | 1307 ms | 15.3 ms | 11.7 ms | 1323 ms →
**11.7 ms** |

The eliminated transpose term dominates and grows with size (as #959
predicted); the NT GEMM itself is even **slightly faster than NN** here
(contiguous per-column B reads pack better). Per #1132, a
native-faster-than-MLAS result on some shape is a graduation event for
`benches/native_vs_mlas.rs` — noting it, but **not** changing default
routing without that gate's measurement.

**RSS.** No new long-lived allocation (reuses `weight_nk`;
`apack`/`bpack` are per-call local scratch, freed at return). Removing
the second full f32 weight materialization cuts the transient f32
footprint of a prefill that hits the dense fallback by one full `[k, n]`
copy (e.g. **270 MiB** at k=13824/n=5120). Not measured end-to-end (no
local model reaches that path); the reduction is structural.

## Memory rules

No new field that outlives a call or scales with weight size — the NT
route reuses the pre-existing `weight_nk` `OnceLock`. `apack`/`bpack`
are per-call `vec![]` scratch. The added lines are not matched by
`weight-cache-guard.yml` (no `OnceLock<…Vec>` /
`(RefCell|Cell)<Vec|Box|Arc>` introduced); its regex and path filter are
untouched.

## Gates (exact counts)

- `cargo test -p onnx-runtime-ep-cpu --lib` → **1324 passed, 0 failed,
17 ignored**
- `cargo test -p onnx-runtime-ep-cpu --lib --features mlas` → **1354
passed, 0 failed, 28 ignored** (MLAS route still works and still wins
where enabled)
- `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings` →
**clean**
- `cargo fmt --check` → the three changed files are clean (verified with
`rustfmt --edition 2024 --check`). Pre-existing diffs remain in three
*unrelated* files (`governed_accumulator_budget.rs`,
`qlinear_matmul.rs`, `simd_activations.rs`) from a local rustfmt version
skew vs CI — left untouched to keep this PR surgical.

---
🤖 Generated with Squad. Flagged **needs review** — please have a squad
member review the kernel packing/tail handling before merge.


---

## Update (2026-08-18, Roy) — merged current `main`, plus a
production-path A/B

### Merge

Merged `main` (`c55a3fab3`), which had since made the #1091 M=1 GEMV the
unconditional `SimdX86` route and dropped the
`ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` toggle this branch still carried.
Conflict resolved by keeping **both**: main's default-route test
(`the_default_entry_point_routes_m1_to_the_gemv`) and this branch's NT
kernel + bit-identity tests. Two follow-on fixes:

- **aarch64 cross-arch lane.** `gemm_nt_with_backend` compiled with
neither the `mlas` nor the x86 arm has no reader for any parameter, so
the `-D warnings` cross-arch pass rejected all six. Bound them in the
unsupported arm. (This is what the old `Rust quality` red was:
`Cross-target compile check`, nothing else.)
- **`mm_profile` gemv phase.** The MLAS route used to time this GEMM on
the `gemv` phase; after the two call sites collapsed into one, that
timer was lost. Restored — the default build gets a pass-through
(`tick()`, the reporter, is MLAS-only), so the shared call site stays
`cfg`-free and MLAS profiling is unchanged.

### Production-path A/B (new harness,
`benches/matmul_nbits_prefill_ab.rs`)

The original body was right that no local model reaches this route, and
honest about not reporting a number it could not defend. That gap is now
closed the way #1013 closed its own: drive the **real kernel through the
EP's own `get_kernel`/`execute`**, at the shapes and inputs that *do*
reach the dense fallback — 8-bit prefill, and 4-bit with `g_idx`. The
harness uses no symbol this branch introduces, so the identical file
runs on `main` and here; both arms were built and run **interleaved**, 3
repetitions each, on the same box.

Host: 32-core x86_64, AVX2, default build (**`mlas` off**), release.
**Contended** (other agents building; load ~12), so medians of per-arm
medians are reported and the ratios — not the absolute ms — are the
claim.

**Steady state** (weight already resident, per-call prefill cost):

| case | k | n | m | main (ms) | this PR (ms) | speedup |
|---|---:|---:|---:|---:|---:|---:|
| int8 dense fallback | 2048 | 2048 | 8 | 12.114 | **0.548** | **22.1x**
|
| int8 dense fallback | 2048 | 2048 | 64 | 12.276 | **1.375** | **8.9x**
|
| int8 dense fallback | 4096 | 11008 | 8 | 53.858 | **4.170** |
**12.9x** |
| int8 dense fallback | 4096 | 11008 | 64 | 56.560 | **7.595** |
**7.4x** |
| int4 + `g_idx` fallback | 2048 | 2048 | 8 | 31.583 | **0.563** |
**56.1x** |
| int4 + `g_idx` fallback | 2048 | 2048 | 64 | 32.405 | **1.419** |
**22.8x** |

**Cold** (fresh kernel per repetition, so the one-time weight dequant is
inside the measurement — the TTFT term #959 attacked):

| case | k | n | m | main (ms) | this PR (ms) | speedup |
|---|---:|---:|---:|---:|---:|---:|
| int8 dense fallback | 2048 | 2048 | 8 | 6.470 | 6.008 | 1.08x |
| int8 dense fallback | 2048 | 2048 | 64 | 12.519 | 7.974 | 1.57x |
| int8 dense fallback | 4096 | 11008 | 8 | 53.761 | 37.260 | 1.44x |
| int8 dense fallback | 4096 | 11008 | 64 | 57.725 | 38.503 | 1.50x |
| int4 + `g_idx` fallback | 2048 | 2048 | 8 | 31.636 | 23.200 | 1.36x |
| int4 + `g_idx` fallback | 2048 | 2048 | 64 | 32.083 | 25.883 | 1.24x |

**Why steady moves 7–56x and cold only ~1.1–1.6x — and why that is the
real finding.** The `Kn` route has **no cache**:
`dequantize_weight(WeightLayout::Kn)` is called *inside* `execute`, so
every prefill call re-materializes the whole transposed f32 weight. The
NT route dequantizes into the pre-existing `weight_nk` `OnceLock`, which
a constant weight fills **once**. So this change removes not one
transpose but *every repeat of it*. Cold (first call) improves by the
layout alone — a contiguous `Nk` write instead of the stride-`N`
scatter, 1.2–1.6x at these sizes; steady improves by the caching the
`Nk` layout makes possible.

Confirmed structurally with `ONNX_GENAI_PROFILE_MM=1` over the same
harness run: **main emits 72 `phase=dequant-kn` lines, this PR emits 0 —
and 24 `phase=dequant-nk`** (one per kernel instance, i.e. the cold arms
only; every steady call pays none).

**Bit-identity, across builds.** The harness prints an FNV-style digest
of the raw output bits. All six rows have the **identical digest on both
arms** (`db6ff07f991d431`, `b4c1df9fd8883789`, `7bd7418eacdcb870`,
`2eed012609649617`, `d5834afbecf42494`, `cd8fa77262302969`) —
bit-identity of the production `execute` result, not just of the kernel
driver, verified across two separately compiled builds.

### Scope, restated honestly

On the default build, 4-bit `accuracy_level=0` with contiguous
(borrowable) inputs takes the zero-copy borrowed int4 path
(#979/#1117/#1126) for both decode *and* prefill and returns before the
dense fallback. This PR therefore changes: **8-bit prefill**, **4-bit
with `g_idx`**, **`weight_prepacked`/non-borrowable** inputs, and
**4-bit `accuracy_level != 0`** that falls through. Those are exactly
the cases measured above. MLAS builds already had the NT route; their
behaviour is unchanged.

### Gates (re-run after the merge)

- `cargo test -p onnx-runtime-ep-cpu --lib` → **1420 passed, 0 failed,
18 ignored**
- `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings` →
clean
- `cargo clippy --locked --target aarch64-unknown-linux-gnu
--all-targets -p onnx-runtime-ep-cpu -- -D warnings` → clean (the lane
that was red)
- `cargo fmt --all -- --check` → the files this PR touches are clean;
one *inherited* diff remains in
`onnx-runtime-ep-cuda/standard_attention.rs` from `main`, fixed
separately in #1347.

---------

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Roy <roy@squad.local>
justinchuby and others added 2 commits August 19, 2026 22:20
Review caught that an empty sample slice underflows median()'s mid - 1.
Guard it alongside the existing --threads and --block-size checks.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Retraction posted, re-measured on latest main, Opus-reviewed, local gate green

The PR description has been rewritten. The original "false by ~63×" conclusion is retracted — it measured the scalar x86 borrowed int4 GEMV that #1021 (3cf49d25) vectorised on 2026-08-16, so the number was real when taken and is meaningless now. It also mixed a weight-only rate against a total-traffic ceiling and reported one best sample from a shared box.

Re-measured on origin/main 085c49140, same box, same probe:

8 threads 16 threads 32 threads
DRAM ceiling (roofline_bandwidth --mib 2048 --seconds 3) 75.099 GB/s 75.460 GB/s 75.737 GB/s
GEMV per-call, median 52.560 ms 29.349 ms 19.785 ms
GEMV per-call, min 52.330 ms 26.388 ms 17.936 ms
GEMV weight-only, best 7.439 GB/s 14.752 GB/s 21.705 GB/s
GEMV total-traffic, best 9.311 GB/s 18.464 GB/s 27.166 GB/s

27.2 of 75.7 GB/s — ~2.8× below roofline, not ~63×. Per-token lm_head compute is ~18–20 ms, not ~500 ms, so #994's criterion does not invert: CPU lm_head is cheaper than the ~33 ms to move the work to the GPU, the opposite of what this PR originally concluded.

Merging it as an instrument, not for its conclusion. It is the probe that caught the error, and the only one in the tree that measures the real quantised decode GEMV against the real DRAM ceiling on the same box.

Review

Opus review verified the instrument itself rather than the claim, and confirmed by reading the dispatch code that it reaches the path it names: with bits=4, accuracy_level=0, three inputs and no prepack, borrow_optional_packed_zero_points(None) yields the symmetric branch at matmul_nbits.rs:1449–1461, ahead of the resident-f32 route, and mlas_sqnbit_owns_fp32_compute is the cfg(not(feature = "mlas")) stub. Byte counts recomputed independently and exact (weight_bytes=389,283,840, scales_bytes=97,320,960, AI 4.000 FLOP/byte). Thread sweep confirmed faithful — set_decode_thread_budget runs before the first execution and one process serves one thread count, so no pool is reused across counts. cargo tree -i mlas-sys on the gemv-probe feature set returns no match.

One non-blocking finding, fixed in f3bd8a6ba: median() underflows mid - 1 on an empty slice, reachable via --repeats 0, which was unguarded. Now bails like --threads and --block-size.

Local gate matrix on f3bd8a6ba (latest main merged in normally)

A fmt · B build-offline · C test-ep-cpu · D clippy -D warnings · E mlas-feature tests · F clippy-native-be · G clippy aarch64-unknown-linux-gnu -D warnings · I ep-cpu no-default-features · J ep-cpu all-features · K no-MLAS-symbols in the default artifact (nm -D = 0) · L cross-compile script · S × 8 guard scripts — all PASS, 19/20.

The one failure is H, cargo check --target aarch64-pc-windows-msvc: onnx-genai-ort-sys bindgen needs the Windows SDK (fatal error: 'stdlib.h' file not found). Reproduced on plain origin/main in a clean worktree — pre-existing and environmental, not caused by this branch.

Miri is not run for this change: the diff is one bench binary plus its Cargo.toml/Cargo.lock entries, and miri.yml covers onnx-runtime-memory, onnx-runtime-dlpack, onnx-runtime-ep-api and onnx-runtime-ep-cpu — none of which this touches. The probe contains no unsafe.

@justinchuby
justinchuby merged commit 9d1e993 into main Aug 19, 2026
4 checks passed
@justinchuby
justinchuby deleted the squad/994-cpu-gemv-roofline branch August 19, 2026 22:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants