Skip to content

Absorb MLAS's M=1 dense-f32 GEMV into the SimdX86 backend (no resident copy) - #1116

Merged
justinchuby merged 2 commits into
mainfrom
squad/1091-absorb-dense-f32-gemm
Aug 17, 2026
Merged

justinchuby merged 2 commits into
mainfrom
squad/1091-absorb-dense-f32-gemm

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Summary

#1045 won 4.4x on dense f32 MatMul with --features mlas; #1091 asks to make that real for a default (no-mlas) build, by absorbing the mechanism into our own SimdX86 kernel rather than shipping behind MLAS. This PR measures the gap on this host, finds where it comes from, proves it is reachable without MLAS's session-lifetime packed buffer, and ports it. Same shape as #1104.

The 4.4x prefill number does not reproduce on this host. In one binary containing both paths (same-binary A/B via the existing NXRT_CPU_GEMM_BACKEND=mlas|simd toggle), at M=128 prefill our built-in SimdX86 6×16 packed microkernel is already at parity with MLAS (0.87–1.15x, net slightly favoring SimdX86). #1045's 4.4x was an AMD EPYC without AVX-512; here it's gone.

The entire reproducible gap is at M=1 decode GEMV: 2.2–4.6x.

shape (M×K×N) M simd/mlas (before)
1×5120×5120 (o_proj) 1 4.59x
1×5120×7168 (qkv) 1 3.58x
1×5120×13824 (gate/up) 1 3.44x
1×13824×5120 (down) 1 2.49x
1×5120×152064 (lm_head) 1 2.23x
128×5120×5120 128 0.88x (simd faster)
128×5120×13824 128 0.87x (simd faster)
128×13824×5120 128 1.15x

Mechanism — source-cited, both sides

MLAS sgemm.cpp (vendored, MlasSgemmOperation):

// Handle the special case of a small M. The data from matrix B is not
// referenced multiple times, so using a local packed buffer is a wasted
// memory copy.
if (M == 1 && TransA == CblasNoTrans && alpha == 1.0f && ...) {
    SgemmKernelM1Routine(A, B, C, K, N, ldb, beta);   // reads B in place at stride ldb
    return;
}

MLAS routes M==1 to SgemmKernelM1Avx.asm, which streams B in place — K unrolled ×4 (ProcessRowLoop4), N swept contiguously (ProcessColumnLoop) — no pack, no resident buffer.

Ours (x86_sgemm.rs::sgemm_simd) calls pack_b into a bpack scratch unconditionally. At M=1 there is a single A-panel, so each packed B panel is reused zero times — the pack is a wasted full read+write copy of B (K·N f32), ≈3× the memory traffic of a straight GEMV. It is memory traffic, not arithmetic and not layout, and the fix needs no resident buffer.

How much is reachable without a resident copy — all of it

sgemm_simd_m1: for M==1, stream B exactly once (K unrolled ×4, sequential N sweep, C accumulated in cache, Rayon over disjoint column strips). No pack_b, no scratch, no OnceLock, no GovernedWeightCache — it actually removes the bpack allocation at M=1. Exactly #1104's "no resident copy" property; nothing to admit/decline under #1056.

The first attempt (column-major, C in registers) regressed lm_head to 3.72x because it strided B by N (608 KB stride → TLB thrash). Matching MLAS's K-outer / N-inner sequential traversal fixed it — the layout that matters for wide outputs.

A/B result — process CPU time and peak RSS

Same binary, SimdX86 M=1 route toggled by ONNX_GENAI_CPU_MM_SIMD_M1_GEMV (default off, like #1104's ONNX_GENAI_CPU_MM_INT4_NBLK). One arm per process; peak RSS polled by PID every 150 ms; process CPU time (TotalProcessorTime); 5 decode shapes, min-of-30.

arm process CPU time peak RSS
MLAS (SgemmKernelM1) 52.5 s 2978 MB
ours, packed (toggle off) 169.4 s 2982 MB
ours, GEMV (toggle on) 57.3 s 2977 MB

2.96× faster than the packed path (169.4 → 57.3 s), within 1.09× of MLAS, at identical peak RSS. Recovered fraction of the MLAS gap: (169.4 − 57.3)/(169.4 − 52.5) = 95.9%, with zero added footprint. Per-shape simd/mlas after: 5120×5120 1.39x, 5120×7168 1.12x, 5120×13824 1.22x, 13824×5120 1.10x, lm_head 1.11x (all down from 2.2–4.6x).

Numerical output

Not byte-identical to the packed path — the GEMV reassociates the f32 sum (K-unrolled-by-4 running accumulation vs the packed KC-panel order). It matches the naive f64 / Generic reference within the same tolerance the existing SimdX86-vs-reference tests use (1e-3·(1+|e|)), and a new test asserts GEMV-vs-packed agreement within that bound (they differ only by summation order, never in which products are summed). This is reported as a numerical change, not shipped silently: the toggle defaults off.

What could not be ported / caveats

  • No f32 model exercises this path on this host. qwen2.5-14b-f32, qwen2.5-14b-onnx, and every qwen05b variant route their weights through MatMulNBits (int4), which does not touch the dense f32 GEMM. So there is no end-to-end token-identity check here; the A/B is a synthetic in-binary driver (bench_f32_gemm_ab, #[ignore]), reported honestly as such rather than as a model number that never took the path.
  • Prefill (M>1) is unchanged — it is already at parity, so this PR deliberately scopes to M==1, exactly as MLAS special-cases only M==1.
  • Default-off toggle means a default build is not yet faster; recommend flipping it on for SimdX86 in a follow-up once the reassociation is signed off, which is what makes the MLAS-routed speedups do not reach a default build, and the strategy is to absorb them natively #1091 win reach users.

Gates (on this host, not CI)

  • cargo test -p onnx-runtime-ep-cpu --lib ×5: 1321 / 1321 / 1321 / 1321 / 1321 passed, 0 failed, 12 ignored each.
  • cargo clippy -p onnx-runtime-ep-cpu --lib -- -D warnings: clean. Also --tests --features mlas: clean.
  • New unit tests: m1_gemv_shapes, m1_route_matches_packed_within_tolerance.

Refs #1091 #1045. Precedent #1104.

…t copy)

#1045 claimed 4.4x on dense f32 MatMul with `--features mlas`; #1091 asks to
make that real for a default (no-mlas) build. Measured in one binary on this
host (RTX 4060 laptop, 20 logical CPUs, AVX2+FMA, no AVX-512), the 4.4x does
NOT reproduce: at M=128 prefill our built-in `SimdX86` 6x16 packed microkernel
is already at parity with MLAS (0.87-1.15x). The entire reproducible gap is at
M=1 decode GEMV: 2.2-4.6x.

Mechanism (source-cited both sides): MLAS `sgemm.cpp` routes M==1 away from its
packed GEBP -- "The data from matrix B is not referenced multiple times, so
using a local packed buffer is a wasted memory copy" -- to `SgemmKernelM1Avx`,
which streams B in place (K-outer, N-inner, sequential). Our `sgemm_simd` calls
`pack_b` unconditionally, so at M=1 it pays a full read+write copy of B reused
zero times: ~3x the memory traffic. It is memory traffic, not arithmetic and
not layout, and the fix needs no resident buffer.

This ports the mechanism natively: `sgemm_simd_m1` streams B once (K unrolled
x4, sequential N sweep, C accumulated in cache), no pack, no scratch, no
OnceLock/GovernedWeightCache -- it removes the `bpack` allocation at M=1.
Selected for M==1 behind the same-binary A/B toggle
`ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` (default off, like #1104's
`ONNX_GENAI_CPU_MM_INT4_NBLK`).

Measured (process CPU time / peak RSS, 5 decode shapes, min-of-30):
  mlas               52.5 s / 2978 MB
  simd packed (off) 169.4 s / 2982 MB
  simd GEMV  (on)    57.3 s / 2977 MB
2.96x faster than the packed path, within 1.09x of MLAS, at identical RSS --
95.9% of the MLAS gap recovered with zero added footprint.

Not byte-identical to the packed path (f32 summation is reassociated); matches
the naive f64/generic reference within the same tolerance the existing
SimdX86-vs-reference tests use. No int4 model exercises the dense f32 path (they
route through MatMulNBits), so the A/B is a synthetic in-binary driver
(`bench_f32_gemm_ab`), reported as such.

Refs #1091 #1045

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
@codecov

codecov Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 61.08374% with 79 lines in your changes missing coverage. Please review.
✅ Project coverage is 79.83%. Comparing base (9b7a458) to head (1883219).
⚠️ Report is 8 commits behind head on main.

Files with missing lines Patch % Lines
crates/onnx-runtime-ep-cpu/src/kernels/matmul.rs 0.00% 69 Missing ⚠️
...rates/onnx-runtime-ep-cpu/src/kernels/x86_sgemm.rs 92.53% 9 Missing and 1 partial ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1116      +/-   ##
==========================================
- Coverage   80.48%   79.83%   -0.65%     
==========================================
  Files         367      367              
  Lines      157217   157756     +539     
  Branches   157217   157756     +539     
==========================================
- Hits       126531   125948     -583     
- Misses      25974    27084    +1110     
- Partials     4712     4724      +12     
Flag Coverage Δ
cli-ort-linux 83.79% <ø> (ø)
cli-ort-windows 83.40% <ø> (+0.09%) ⬆️
offline 79.69% <61.08%> (-0.67%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...rates/onnx-runtime-ep-cpu/src/kernels/x86_sgemm.rs 96.82% <92.53%> (-2.71%) ⬇️
crates/onnx-runtime-ep-cpu/src/kernels/matmul.rs 77.95% <0.00%> (-8.39%) ⬇️

... and 15 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 gather/medium_f32_threads=1-internal/32768 4.99 µs 7.67 µs +53.7%
⚠️ gather/large_bf16_threads=1-internal/131072 12.83 µs 15.89 µs +23.8%
⚠️ gather/small_bf16_threads=1-internal/4096 536.7 ns 655.8 ns +22.2%
⚠️ matmul/medium_generic_f32_threads=8/32x512x512 944.89 µs 1.12 ms +18.7%
⚠️ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 552.62 µs 648.94 µs +17.4%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 186.24 µs 212.60 µs +14.2%
✅ gather/large_f32_threads=1-internal/131072 37.73 µs 42.89 µs +13.7%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.35 ms 2.64 ms +12.1%
✅ matmul/small_generic_f16_threads=1/1x256x256 34.57 µs 37.89 µs +9.6%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.15 ms 2.36 ms +9.5%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 85.14 µs 90.13 µs +5.9%
✅ tokenization/decode_tokens_per_second 6.30 ms 6.57 ms +4.2%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.28 ms 1.33 ms +4.0%
✅ grammar_masking/llguidance_compute_mask/32 76.02 µs 78.96 µs +3.9%
✅ tokenization/encode_tokens_per_second 382.35 µs 394.81 µs +3.3%
✅ sampling_latency/top_k_per_token 52.76 µs 54.00 µs +2.3%
✅ matmul/medium_generic_f16_threads=1/32x512x512 30.57 µs 31.27 µs +2.3%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 491.93 µs 502.60 µs +2.2%
✅ add/large_f32_threads=1-internal/4194304 666.98 µs 677.31 µs +1.5%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 79.65 µs 80.53 µs +1.1%
✅ sampling_latency/top_p_per_token 382.97 µs 385.68 µs +0.7%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.52 ms 3.54 ms +0.6%
✅ logit_processing/seven_processor_chain_per_step 320.30 µs 322.00 µs +0.5%
✅ qwen3_sampling_processors/top_k_top_p_fast 657.28 µs 660.58 µs +0.5%
✅ kv_cache/alloc_dealloc_pages 40.18 µs 40.36 µs +0.5%
✅ gather/medium_f16_threads=1-internal/32768 2.99 µs 2.99 µs -0.0%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.67 ms 5.67 ms -0.1%
✅ add/large_f16_threads=1-internal/4194304 1.64 ms 1.64 ms -0.1%
✅ sampling_latency/greedy_per_token 3.21 µs 3.20 µs -0.3%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 526.74 µs 524.91 µs -0.3%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 3.86 ms 3.84 ms -0.6%
✅ qwen3_sampling_processors/top_k_partial_selection 145.02 µs 144.06 µs -0.7%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.40 ms 9.31 ms -0.9%
✅ add/large_bf16_threads=1-internal/4194304 1.72 ms 1.69 ms -1.6%
✅ add/medium_bf16_threads=1-internal/262144 105.25 µs 103.53 µs -1.6%
✅ reduce_mean/small_f32_threads=1-internal/4096 15.29 µs 14.98 µs -2.0%
✅ sampling_latency/min_p_per_token 213.64 µs 208.63 µs -2.3%
✅ matmul/medium_generic_f16_threads=8/32x512x512 29.86 µs 29.11 µs -2.5%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.05 ms 1.95 ms -4.8%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 549.15 µs 519.20 µs -5.5%
✅ matmul/small_generic_bf16_threads=8/1x256x256 41.15 µs 38.86 µs -5.5%
✅ add/small_f16_threads=1-internal/1024 502.1 ns 464.8 ns -7.4%
✅ add/medium_f16_threads=1-internal/262144 112.16 µs 103.45 µs -7.8%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 48.92 µs 44.91 µs -8.2%
✅ matmul/small_generic_f16_threads=8/1x256x256 39.27 µs 34.84 µs -11.3%
✅ matmul/small_generic_bf16_threads=1/1x256x256 39.94 µs 35.23 µs -11.8%
✅ add/small_bf16_threads=1-internal/1024 502.6 ns 440.9 ns -12.3%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.28 ms 1.12 ms -12.6%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 445.73 µs 386.24 µs -13.3%
✅ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 56.11 µs 48.26 µs -14.0%
✅ matmul/small_generic_f32_threads=8/1x256x256 48.16 µs 41.31 µs -14.2%
✅ gather/small_f16_threads=1-internal/4096 692.3 ns 592.6 ns -14.4%
✅ matmul/small_generic_f32_threads=1/1x256x256 46.57 µs 39.77 µs -14.6%
🟢 add/medium_f32_threads=1-internal/262144 29.71 µs 25.25 µs -15.0%
🟢 gather/large_f16_threads=1-internal/131072 15.86 µs 13.47 µs -15.1%
🟢 gather/medium_bf16_threads=1-internal/32768 3.27 µs 2.73 µs -16.4%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 318.04 µs 265.63 µs -16.5%
🟢 add/small_f32_threads=1-internal/1024 239.3 ns 194.4 ns -18.8%
🟢 gather/small_f32_threads=1-internal/4096 1.01 µs 704.4 ns -30.3%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.68 3.47 5.64 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby

Copy link
Copy Markdown
Owner Author

Ran your harness myself. The headline finding reproduces and is important; one shape needs explaining before I merge.

Your central correction reproduces, and it matters

With the toggle off, on this host (20 logical CPUs, AVX2+FMA+F16C+AVX-VNNI, no AVX-512), threads=20 iters=30:

shape (M x K x N) simd/mlas
1 x 5120 x 7168 1.71x
1 x 5120 x 5120 2.82x
1 x 5120 x 13824 2.08x
1 x 13824 x 5120 2.40x
1 x 5120 x 152064 2.70x
128 x 5120 x 5120 1.05x
128 x 5120 x 13824 0.72x
128 x 13824 x 5120 0.57x

#1045's 4.4x does not reproduce here, and at prefill shapes we are already at or ahead of MLAS -- 0.57-1.05x at M=128. The entire reproducible gap is M=1. That is a substantive correction to how that PR framed the problem, and it is the kind of thing that only shows up when someone re-measures a claim on different hardware instead of inheriting it.

The mechanism you cite makes it make sense: MLAS routes M==1 to SgemmKernelM1Avx because "the data from matrix B is not referenced multiple times, so using a local packed buffer is a wasted memory copy", while sgemm_simd packs B unconditionally. The fix is doing less work, not more -- which is also why it needs no resident buffer.

What I cannot yet confirm: one shape looks like a regression

With the toggle on, within-run ratios:

shape off on
1 x 5120 x 7168 1.71x 4.52x
1 x 5120 x 5120 2.82x 2.10x
1 x 5120 x 13824 2.08x 1.28x
1 x 13824 x 5120 2.40x 0.95x
1 x 5120 x 152064 2.70x 0.72x

Four of five improve, and the lm_head shape (1 x 5120 x 152064) goes from 2.7x slower than MLAS to 1.4x faster -- that is the shape that matters most for decode, so this is a real win.

But 1 x 5120 x 7168 moves the wrong way, and I cannot tell whether it is real: another agent was building during my "on" run and the MLAS arm itself moved between runs (13.769 → 9.902 ms on that shape, 21.802 → 40.010 ms on another). Absolute times across runs are not comparable under that load; only within-run ratios are, and even those are suspect when both arms are perturbed.

Before merging, please:

  1. Re-run both arms on a quiet machine and report the ratio table for both, ideally interleaved within one process so the arms share conditions.
  2. Explain or fix 1 x 5120 x 7168. If the sequential K-outer/N-inner stream has a shape-dependent weakness -- an N that maps badly onto the cache or the unroll -- that should be stated as a known boundary, and ideally the route should fall back to the packed path where it loses. A kernel that wins four shapes and loses one should say so at the dispatch site, not only in a PR comment.
  3. Confirm the M=128 rows are unchanged with the toggle on (they should be untouched, but the run I did showed movement there too and I could not separate it from load).

On the numerical change

Not being byte-identical to the packed path is acceptable if it stays behind a default-off toggle and the deviation is quantified, which you did against the f64/Generic reference. Keep it default-off until the shape question is settled -- and when it is eventually flipped on, that flip is a separate decision that needs the deviation restated, because "default off, numerically different" and "default on, numerically different" are very different claims to a user.

@justinchuby

Copy link
Copy Markdown
Owner Author

Addendum — my "on" run has a built-in control group, and it says the run was too noisy to judge.

The full table with the toggle on:

shape off on can the toggle affect it?
1 x 5120 x 7168 1.71x 4.52x yes
1 x 5120 x 5120 2.82x 2.10x yes
1 x 5120 x 13824 2.08x 1.28x yes
1 x 13824 x 5120 2.40x 0.95x yes
1 x 5120 x 152064 2.70x 0.72x yes
128 x 5120 x 5120 1.05x 0.70x no
128 x 5120 x 13824 0.72x 0.92x no
128 x 13824 x 5120 0.57x 0.55x no

The M=128 rows cannot be affected by an M=1-only route, so they are a control. They moved by up to 1.5x (1.05 → 0.70) and 1.28x (0.72 → 0.92) anyway. That sets the noise floor for the whole run.

Two consequences:

  1. The 1 x 5120 x 7168 "regression" is not established. 1.71x → 4.52x is larger than the control drift, so it is not explained by noise either -- but with the control moving that much I cannot separate the two. It still needs a quiet re-run before I will call it real or spurious.
  2. The wins I quoted are understated or overstated by an unknown factor, including the lm_head 2.70x → 0.72x. The direction is almost certainly right -- it is far outside the control drift -- but the magnitude is not something to publish.

Please keep these M=128 rows in the harness and report them explicitly as a control in every future run. They cost nothing extra, they are already there, and they turn "the machine was busy" from an excuse into a measurement: if the control rows move, the run is not usable, and you know that before drawing a conclusion rather than after someone questions it.

This generalises beyond this PR. Every A/B we run on this shared box would benefit from an arm the change provably cannot touch. I have been using repeat-spread within one arm for that, which works but costs extra runs; an untouched shape in the same process is cheaper and stronger.

justinchuby pushed a commit that referenced this pull request Aug 17, 2026
This machine is shared with build agents, and the failure mode is not "numbers
are a bit off" -- it is publishing a conclusion that reverses when the box is
quiet. Three habits, learned the expensive way over the past two days.

**CPU time, not wall clock.** Three identical runs of one configuration measured
39.3 / 25.8 / 16.1 s wall while `TotalProcessorTime` reproduced to ~2%. Includes
the RSS-polling snippet, since `PeakWorkingSet64` reads 0 after exit and must be
sampled by PID while the process runs.

**Compare an effect against the spread of its own arm.** A 14B model showed ~20%
within-arm wall spread; anything under ~1.3x there is unmeasured.

**Best: a control arm the change provably cannot touch.** This caught a real case
on #1116: an M=1-only GEMV toggle was under evaluation and the M=128 rows in the
same harness -- which that route cannot reach -- moved 1.05x to 0.70x and 0.72x to
0.92x between runs, setting a ~1.5x noise floor and making an apparent
single-shape regression unadjudicable. Without the control that would have been
argued about; with it the answer was simply "re-run when quiet". A control costs
one extra shape in a harness you already have and is stronger than a distribution.

Also records the differencing method for separating fixed from per-unit cost:
decode per token by differencing two token counts, prefill per token by taking the
slope between two prompt lengths. A single long run cannot separate them and a
single short run is almost all fixed cost; both mistakes have been made here,
including by me.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
Address #1116 review: re-measure both arms interleaved in one process, add a
built-in M=128 control with a usability rule, and adjudicate 1x5120x7168.

- x86_sgemm: split sgemm_simd into an env wrapper + sgemm_simd_variant(...,
  use_m1_gemv) + sgemm_simd_packed, so the A/B harness can drive both variants
  in one process without touching process-global env. For m>=2 both variants
  run identical packed code -- the basis of the control.
- bench_f32_gemm_ab: interleave arms at the iteration level (shared conditions),
  print the M=128 rows as a control every run with a usability verdict
  (gemv/packed must be ~1.0; drift beyond threshold => RUN NOT USABLE).
- Document the measured per-shape dispatch boundary at the M==1 route: the GEMV
  strictly dominates the packed default on all five decode shapes (2.0-2.9x), so
  no packed fall-back (it would be slower); the only residual gap is vs MLAS on
  the two largest shapes, recorded as a known limit.

Clean, control-passing results (process CPU time; wall clock is too noisy for
the 20-thread M=128 rows on this shared box, ~25% floor, so the control is gated
on CPU time which reproduces to ~2-4%):
  CONTROL M=128 gemv/packed = 1.02x, 1.04x (<5%) -> usable; RSS packed==gemv.
  DECODE M=1 gemv/packed per shape: 0.50 0.34 0.43 0.42 0.34 (2.0-2.9x faster).
  1x5120x7168 = 0.50 (2x faster than packed) -- the earlier "wrong way" was noise.
Toggle stays default-off; numerics unchanged (matches f64/Generic ref within
1e-3*(1+|e|)).

Gates: cargo test -p onnx-runtime-ep-cpu --lib x5 = 1321 passed / 0 failed /
12 ignored (all five); cargo clippy -p onnx-runtime-ep-cpu --lib -D warnings clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
@justinchuby

Copy link
Copy Markdown
Owner Author

Re-measurement (control-gated) — pushed b4a99148

Thanks — the control catch was right, and chasing it produced a methodology I think is worth keeping. Addressing all four points.

The core problem your run exposed, generalized

The M=128 rows did move in your run, and they can't. I reproduced that here: on this shared box the 20-thread M=128 GEMMs have a ~25% wall-clock noise floor — min-of-N never clears sustained background load during a whole arm's window. So a wall-clock control keeps (correctly) refusing to certify, run after run.

The fix is your own framing: gate the control on process CPU time, not wall clock. CPU time reproduces to ~2–4% here under the same contention that swings wall clock 25%. So the authoritative numbers below are per-arm isolated-process TotalProcessorTime, and the harness in-binary reports the control every run.

1 + 2. Interleaved re-run, with the M=128 control reported every run

CONTROL (M=128 prefill — identical packed code in every arm, so gemv/packed must be ~1.0):

rep simd_packed (CPU s) simd_gemv (CPU s) mlas (CPU s) gemv/packed drift verdict
1 80.06 81.77 79.75 2.1% OK
2 79.25 82.14 74.98 3.6% OK

Both < 5% ⇒ usable. Peak RSS: packed == gemv == 293 MB, mlas 286 MB. The rule is now built into bench_f32_gemm_ab and I've written it up as the recommended pattern for every A/B harness in the repo: the control rows cost nothing and turn "the box was busy" into a measurement you make before concluding. For m≥2 both simd arms literally call the same function, so their ratio is a pure noise gauge.

4. M=128 untouched — confirmed

Control above: gemv/packed = 1.02–1.04x. Our packed prefill also ≈ MLAS (79–80 vs 75–80 CPU s), consistent with your finding that prefill is already at parity and the entire gap is M=1.

DECODE (M=1) — the actual result

Aggregate over the 5 decode shapes, 2 reps: gemv/packed = 0.33x, 0.36x (≈3x faster than our default) and gemv/mlas = 0.85x, 0.93x (gemv edges ahead of MLAS in aggregate). Peak RSS 2977 vs 2982 MB — identical, bpack removed at M=1.

Per-shape (CPU s):

shape (M=1) mlas packed gemv gemv/packed gemv/mlas
1×5120×7168 (QKV) 5.20 7.08 3.55 0.50 0.68
1×5120×5120 (o_proj) 2.34 6.47 2.17 0.34 0.93
1×5120×13824 (gate/up) 6.95 12.52 5.33 0.43 0.77
1×13824×5120 (down) 4.59 13.31 5.63 0.42 1.22
1×5120×152064 (lm_head) 37.28 130.19 44.03 0.34 1.18

3. 1x5120x7168 — explained

On robust CPU time it is gemv/packed = 0.50 (2× faster than packed), gemv/mlas = 0.68 (1.5× faster than MLAS) — one of the best rows, not a loss. The "1.71x → 4.52x" you saw was exactly the wall-clock noise the control now catches before concluding.

So there is no M=1 shape where GEMV loses to the packed default (0.34–0.50 across all five). A dispatch fall-back to packed would therefore be strictly slower and is deliberately absent. The only residual gap is versus MLAS — not our shipping default — and only on the two largest shapes (down_proj K=13824 → 1.22, lm_head N=152064 → 1.18), where sequential-B streaming trails MLAS's blocked M=1 asm. That MLAS-only boundary is now documented in code at the dispatch site in sgemm_simd_variant, with the per-shape numbers, so the win/loss structure lives at the dispatch, not only here.

Numerics — unchanged stance

Toggle stays default-off. GEMV reassociates the f32 sum vs packed (not byte-identical) but matches the f64/Generic reference within 1e-3·(1+|e|), the same tolerance the existing SimdX86 tests use. "Default off, numerically different"; flipping it on remains a separate decision that must restate the deviation.

Gates

  • cargo test -p onnx-runtime-ep-cpu --lib ×5: 1321 passed / 0 failed / 12 ignored (all five).
  • cargo clippy -p onnx-runtime-ep-cpu --lib -- -D warnings: clean.

@justinchuby

Copy link
Copy Markdown
Owner Author

Discard-and-retry provenance (wall-clock, in-process interleaved) + why the CPU-time number stands

Taking your practical note: I stopped waiting for a quiet box and instead ran the harness repeatedly, keeping only control-passing runs. The interleave you suggested is already in the pushed harness — arms alternate per iteration inside one process, so both simd variants see the same load moment-to-moment (bench_f32_gemm_ab, commit b4a99148).

8 consecutive wall-clock runs, all gated on the M=128 control

run worst |gemv/packed − 1| on M=128 verdict
1 17.2% discarded
2 15.0% discarded
3 25.6% discarded
4 24.0% discarded
5 59.7% discarded
6 28.7% discarded
7 18.6% discarded
8 11.2% discarded (best, still > 5%)

8/8 discarded. prefill-block (#1117) is building on this box and saturates all 20 logical cores, so the M=128 rows — which fully saturate 20 threads themselves — never find a clean floor even with min-of-60 and per-iteration interleave. This is exactly the failure your control is designed to make visible before concluding, and it now does so automatically every run.

The M=1 signal survives the noise anyway

Even across all 8 discarded runs, the decode result never wavered in direction — gemv/packed per shape, min–max over 8 runs:

shape (M=1) gemv/packed range ever loses to packed?
1×5120×7168 0.21–0.52 no
1×5120×5120 0.34–0.62 no
1×5120×13824 0.35–0.52 no
1×13824×5120 0.25–0.62 no
1×5120×152064 0.22–0.27 no

GEMV beats the packed default on every M=1 shape in every one of 8 runs, including the ~60%-noise run. The win is large enough to be unambiguous even when the machine is loud — which is the strongest possible answer to "was 1x5120x7168 real?": yes, it's 2–5× faster than packed in all 8, and the earlier "wrong way" reading was the M=128-style noise this table quantifies.

Why the merge number is the CPU-time one

Wall clock cannot be certified on this box while #1117 builds (table above). Process CPU time can: it is immune to the preemption that inflates wall time, and the CPU-time control passes (gemv/packed = 1.02×, 1.04× on M=128; drift 2.1%, 3.6% < 5%). Those are the numbers in my prior comment and they stand:

  • DECODE M=1, CPU s, gemv/packed: 0.50 0.34 0.43 0.42 0.34 (2.0–2.9× faster than the packed default), gemv/mlas 0.85–0.93× aggregate.
  • CONTROL M=128, CPU s, gemv/packed: 1.02×, 1.04× → usable; peak RSS packed == gemv (293 MB), identical.

On the harness itself

Per-iteration interleave + the M=128 control + usability verdict are in bench_f32_gemm_ab now and match the practice you added to .agents/skills/profiling/SKILL.md. The one gap: the in-binary control is wall-clock, so on a loud box it (correctly) discards rather than certifies. Making it self-certify would mean sampling per-arm TotalProcessorTime via GetProcessTimes FFI inside the harness — a real but non-trivial change. Per your "defensible number this week over perfect harness next week," I've left that as a documented follow-up and used the external CPU-time driver for the authoritative table. Happy to do the FFI version as a separate PR if you want every harness self-certifying.

No code change in this update — it's the measurement provenance you asked for. The kernel, dispatch-site boundary doc, gates (5×1321 passed / clippy clean), and default-off numerics are as in b4a99148.

@justinchuby

Copy link
Copy Markdown
Owner Author

The provenance approach worked, and it answered the question I was blocked on. Merging.

The control did its job automatically

Eight consecutive in-process interleaved runs, every one discarded by the control rule — M=128 drifts of 17.2 / 15.0 / 25.6 / 24.0 / 59.7 / 28.7 / 18.6 / 11.2%, because another agent was saturating all 20 cores. That is exactly what a control is for: the run announced its own unusability before anyone drew a conclusion from it.

Reporting the discards rather than quietly keeping the best-looking run is the part I want to see repeated. "Eight runs discarded, here is each drift, here is the one whose control passes" is far stronger than a single clean-looking table with no provenance.

And the signal survived the noise, which settles my open question

I had flagged 1 x 5120 x 7168 as an apparent regression (1.71x → 4.52x in my own contaminated run) and said I could not adjudicate it. Your eight discarded runs adjudicate it: gemv/packed stayed 0.21–0.62 on every shape in every run, including the one with ~60% control drift. The GEMV path never once loses, on any shape, under any load.

That is a stronger form of evidence than a clean measurement would have been. A single quiet run shows the effect under one condition; eight noisy runs that all agree in direction show it survives conditions. My flagged regression was an artefact of comparing arms measured at different times — the same mistake I made again an hour later on #1126 and had to retract.

Defensible numbers (CPU time, whose control passes at 1.02x/1.04x, drift 2.1%/3.6%): decode gemv/packed 0.34–0.50, i.e. 2–3x faster, identical peak RSS.

What lands

  • The correction that perf(mlas): let MlasGemmBatch use the registered parallel backend (4.4x dense f32 MatMul, reaches ORT parity) #1045's 4.4x does not reproduce here and that prefill (M=128) is already at 0.57–1.05x of MLAS. The entire reproducible gap was M=1. That reframing is worth more than the kernel.
  • The mechanism, cited from MLAS's own source: it declines to pack B at M==1 because "the data from matrix B is not referenced multiple times, so using a local packed buffer is a wasted memory copy". The fix was doing less work, not more — which is why it needs no resident buffer and no Every resident weight side-buffer must be in the memory plan before it is allocated #1056 governance.
  • Default-off, numerically different from the packed path, deviation quantified against the f64/Generic reference. When someone proposes flipping it on, that is a separate decision requiring the deviation to be restated: "default off, numerically different" and "default on, numerically different" are very different claims to a user.

The self-certifying in-binary control via GetProcessTimes is a good follow-up and correctly not blocking. A harness that refuses to report a number it cannot stand behind is the right end state; a defensible number this week beats a perfect harness next week.

@justinchuby
justinchuby merged commit 1f1ce4b into main Aug 17, 2026
8 of 16 checks passed
@justinchuby
justinchuby deleted the squad/1091-absorb-dense-f32-gemm branch August 17, 2026 15:44
justinchuby added a commit that referenced this pull request Aug 17, 2026
`Fast (Linux x86_64)` is red on main for every open PR: #1116 landed
with rustfmt drift.

```
Diff in crates/onnx-runtime-ep-cpu/src/kernels/matmul.rs:4950
Diff in crates/onnx-runtime-ep-cpu/src/kernels/x86_sgemm.rs:349
```

Verified on a pristine detached checkout of `origin/main` @ `1f1ce4b74`,
so it is not introduced by any open branch. This is pure `cargo fmt
--all` output — no semantic change, no test change.

Co-authored-by: Deckard <deckard@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby pushed a commit that referenced this pull request Aug 17, 2026
… the deliverable

Updates the ledger: dense f32 M=1 GEMV is absorbed (#1116), and #1045's headline
4.4x is recorded as **not reproducing** on this host -- `simd/mlas` was already
0.57-1.05x at M=128, so the entire reproducible gap was M=1 decode. Inheriting
that figure would have sent someone optimising prefill, which was not the problem.
The fix was to stop packing B at M==1, matching MLAS's own reasoning that packing
a matrix referenced once is a wasted copy: a win from doing less work.

Also records a pattern that has now held three times in a row. Each brief
predicted a mechanism and the measurement found a different one -- #1104 expected
layout and found register blocking, #1116 expected a 4.4x prefill gap and found
the gap was entirely at M=1, #1126 expected missing GEMM blocking and found
per-row dispatch and allocation overhead. In all three the correction was worth
more than the patch.

The point is not that briefs are unreliable: each hypothesis was specific enough
to direct a measurement that could refute it, which is what a hypothesis is for.
The point is to ask for the mechanism *before* the kernel, because a plan is
cheap to change then and expensive afterwards -- #1104's transient-tile design was
abandoned as unnecessary rather than built and then found unnecessary.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
justinchuby added a commit that referenced this pull request Aug 17, 2026
…ed sums)

Replace the row-serial borrowed int4 prefill path's two per-token overheads
with a structural rewrite behind a default-off A/B toggle
(ONNX_GENAI_CPU_MM_INT4_PREFILL), per the method used in #1021/#1027/#1104/#1116:

- One fork-join over disjoint column strips for the whole prefill, instead of
  the row-serial path's m per-row fork-joins.
- The per-block �ctivation_sums Vec is hoisted out of the per-row loop to a
  single m * block_count allocation, independent of the weight size.

Rows are still visited outer-most, so a column's packed bytes are re-read once
per row: weight traffic and the per-element k-reduction order are unchanged, so
output is byte-identical to the row-serial path (both call the shared
�orrowed_int4_output_element). No resident buffer is added (peak RSS
unchanged), satisfying the #1056/#1117 no-session-buffer constraint.

GEMM blocking (reusing a column's bytes across a tile of rows) is deliberately
NOT included here: measured within run-to-run noise on qwen05b and with no
signal on qwen14b (5-rep interleaved), so it is left as an unproven follow-up
to #1117 that can be added or dropped cleanly.

Measured win is model-size dependent: qwen05b prefill slope 1.219 -> 0.842 CPU
s/token (median of 5, ~1.45x); qwen14b shows no measurable change (medians
8.69/8.62 off/on, within a 30-40% within-arm spread) because the fixed per-row
overheads this removes are a negligible fraction of the 14B's larger per-row
compute.

Refs #1117

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
justinchuby added a commit that referenced this pull request Aug 17, 2026
Merging main brought in #1116's `x86_sgemm` and #1126's int4 prefill work,
which made two things in this branch wrong.

**The bench was measuring the wrong clock.** It reported process CPU time
only, on the reasoning that MLAS might thread where our baselines do not.
That is backwards for a graduation gate: our `x86_sgemm` parallelises over
column strips while MLAS deliberately declines to parallelise some shapes,
so CPU time can show a native route "losing" by 10x precisely when it wins
on latency. Wall time now decides, with CPU time printed beside it as
`cpu_ratio` so the opposite failure — a route that wins by recruiting the
whole machine — is visible too, and flagged as
`native-graduates-but-costs-more-cpu` rather than passing silently. The
bench also names the native backend it measured, because the same table
means different things on an AVX2 host and a host without it.

**`matmul.rs`'s module doc contradicted `backend.rs` after the merge.** It
still said MLAS was opt-in and reached only via `NXRT_CPU_GEMM_BACKEND=mlas`,
which stopped being true on this branch. Same for `KERNEL_PERF.md`, and
`CROSS_PLATFORM.md`'s P1 row claiming an MLAS build would fail on Windows
MSVC and macOS — `mlas-sys` compiles MASM and Apple sources today, so that
row is marked resolved with the real remaining gap named instead.

Re-measured everything on the merged tree. The finding worth keeping: #1116
built the M=1 GEMV absorption (stream B in place rather than packing panels
reused zero times) but left it behind `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV`,
default off, pending a measurement. This bench is that instrument, so the
measurement is now in the migration doc: 2.4x faster native at 1x2048x2048,
one-sided, and it moves nothing but M=1. Still short of MLAS on this host,
so it does not graduate — but it is the shortest path to the first f32 GEMM
graduation. Not flipped here; this PR is infrastructure.

Also recorded that the GEMM numbers came off a busy shared container and
moved up to 2x between runs, and that the bench calls `gemm_with_backend`
directly so neither route gets prepacked weights. Without both caveats the
table reads as a measurement rather than the order-of-magnitude it is.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 23, 2026
…d a knob that no longer exists (#1822)

Follow-up to #1173, correcting two defects I shipped in it and repairing
the rule they undermined. Docs, one ledger string, one new test, one new
script. No production kernel or routing change.

## 1. The ledger named a route gate that had already been deleted

`PLAN[MatMulF32].shape_gate` said the native `SimdX86` route "gates M=1
on `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` (default off, #1116)".

#1183 shipped that GEMV on by default and removed the probe. `git
merge-base --is-ancestor 5417d04 bdb4599` confirms it landed
**before** #1173 merged — so the ledger was wrong the day it landed.
Today `sgemm_simd` calls `sgemm_simd_variant(a, b, c, m, k, n, true)`
unconditionally and `use_m1_gemv` is a plain parameter that only the
in-process A/B harness passes as `false`. No environment variable
reaches that route.

`docs/performance/CPU_MATMUL_ASSIGNMENT.md:559` already recorded the
correct fact ("It is measured now, and the route is the default. There
is no env probe on the dispatch any more"). Two files in the same
directory disagreed and nothing compared them.

**Now guarded.**
`ledger_prose_only_names_environment_variables_that_still_exist`
requires every `NXRT_*` / `ONNX_GENAI_*` token in the ledger's prose to
still exist as a string literal in the crate's sources. It cannot check
that the description is *right*, only that the knob is *real* — which is
the half that goes stale silently.

Mutation-verified, not just observed green:

```
matmul_f32: ledger prose names environment variable `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV`,
but no source file in this crate contains the literal "ONNX_GENAI_CPU_MM_SIMD_M1_GEMV".
```

## 2. The doc published a toggle A/B that could not have been run

#1173 carried a table captioned **"same binary, same session, toggle the
only difference"**, reporting `decode 1×2048×2048` at 0.146 with
`ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` off against 0.337 with it on, and
called turning it on "the obvious next slice".

Nothing reads that variable. Setting it measures the same route twice;
it cannot produce two different columns. The table is withdrawn and the
retraction kept in the text rather than quietly deleted.

This is the failure mode the document's own graduation rule warns about
— **an arm that was not on the route it was labelled with** — committed
by the document that wrote the rule. It survived review because a
plausible number in a well-formed table is not self-evidently
unmeasured. Readers are pointed at `bench_f32_gemm_ab`, which holds the
route as a function parameter and carries the M≥2 rows as a built-in
control.

## 3. The gap table is re-measured and the ≥5% rule is repaired

The old table was one unguarded invocation per row at an unstated width,
taken before the decode-placement corrections (#1729, #1794, #1811) —
i.e. when the decode pool put 16 workers on 8 physical cores.

New harness: `scripts/bench_native_vs_mlas_width.py`. Arms interleaved
rep by rep so host drift lands on both equally; per-rep `os.wait4`
CPU-efficiency guard adapted from #1809; six reps per arm; two widths.
Raw verdicts, spreads and discards are all reported rather than
summarised away.

**Three findings, all about method rather than kernels.**

| | narrow (6 cores, 1 L3) | wide (32 logical CPUs) |
|---|---|---|
| `matmul_f32 16×512×512` | 1.581, spread 41% | 0.866, spread 134% |
| `matmul_f32 decode 1×2048×2048` | 1.117, spread 21% | 0.934, spread
13% |

- **Two cases change verdict on width alone.** Same binary, same
half-hour, only the CPU mask differs. `x86_sgemm` parallelises over
column strips and MLAS declines to parallelise some shapes, so
interleaving the two *routes* inside one process does not protect the
ratio — it changes both at once.
- **`16×512×512` disagrees with itself on both arms**, alternating
`keep-mlas` / `native-graduates` from a byte-identical binary. **One
more run of the old table could have graduated a route on this row.**
- **The narrow arm is more trustworthy despite having fewer cores** —
spreads 4–42% against 5–134%, and it lost no reps to the guard.
Isolation beat parallelism.

**Softmax now decomposes cleanly**, because no vendored MLAS kernel has
changed since #1173 (the only `mlas-sys` edits are the additive
straggler handshake in `work_stealing_pool.rs`, #828/#1714, which adds
waiting). At matched width the MLAS control arm is stationary to within
4% while native improved **1.24–1.27×** — matching #1416's claim for the
row kernel. The f32 GEMM rows get no such attribution and now say so
explicitly: their control moved **2.0× the wrong way**, so only the
current ratio at a stated width is defensible.

**The rule gains what it lacked**: spread must be smaller than the
claimed win; reps that did not get the CPU are discarded rather than
averaged; a verdict is valid only at a stated width. Under it, `decode
1×2048×2048` — the first f32 GEMM case to show a real native win —
**still does not graduate**: it costs more CPU (cpu_ratio 0.875), does
not hold at 32 threads, and its 21% spread exceeds its 12% win.

## The width claim is verified, not asserted

#1815 landed while this was in progress and observed the neighbouring
`bench_generic` harness spawning its ORT arm *outside* the affinity
confinement it applied to the native arm. That hazard applies to any
`taskset` claim, including mine, so I checked it instead of trusting it
— sampling `Cpus_allowed_list` from `/proc/<pid>/task/*/status` 40×
across a live narrow-arm run:

```
'16,20,22,26,28,30': 478 observations
  native_vs_mlas- 273, mlas-sys-ws-0..4 39 each, nxrt-task-0..4 2 each
'0-31': 1  (the taskset process itself, before exec)
```

Both routes confined identically; no thread escaped. The rule now
requires this check.

## Validation

- `dispatch_ledger` **17/17**, including the new falsifier, after
merging latest `main`.
- `default_artifacts_are_mlas_free` **9/9** — the no-MLAS-in-defaults
invariant is untouched.
- `cargo clippy -p onnx-runtime-ep-cpu --lib --all-targets` clean;
`cargo fmt --check` clean.
- Normal merge of `origin/main` (`aee2b9d11`), no rebase, no conflicts.

## Limitations

- The narrow arm is six cores on one L3 of one x86-64 host. Nothing here
transfers to aarch64 or to a two-socket box.
- The `activations erf 1 Mi` row shows native 13.5% slower at matched
width. The nearest scatter figure is the wide arm's 8% spread, but that
is a spread of *ratios* against a move in a *native time*, so the two
are not strictly commensurable. Its MLAS control also moved 11%.
**Flagged for pinned re-measurement, not reported as a regression.**
- The wide arm was taken with ~4–5 cores of unrelated load present. That
is stated in the doc rather than hidden, and it is why its spreads are
wider; the guard reports which reps were discarded instead of pretending
the host was quiet.
- No production behaviour changes here, so there is no performance claim
to make about the shipped artifact.

Refs #1173, #1183, #1809, #1815, #1416.



## Independent review, and what it changed

An independent adversarial review of the full diff returned **no
blockers** — it confirmed the ancestry argument behind the retraction,
the stationary-control premise for the softmax attribution, and that the
headline case is correctly *refused* by the rule (21% spread against a
12% win). It also found seven real defects, all now fixed in
`f0323f9ed`.

The one that mattered most was in the new test. It only proved the
variable name appeared *somewhere* in the crate, so a variable whose
read site had been deleted but whose name survived in an
`EnvVarGuard::set(...)` line would still have passed — which is the
precise shape of the defect this PR exists to correct. The test now
requires the matching line to be an `env::var(` / `env::var_os(` read or
an `_ENV: &str =` binding.

Verified by mutation in **both** directions:

| mutation | before | after |
|---|---|---|
| reinsert retired `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` into ledger prose |
fails ✅ | fails ✅ |
| retire the two real `NXRT_CPU_GEMM_BACKEND` reads, leaving the literal
only in test guards | **passes ❌** | fails ✅ |

The remaining six were prose defects in the doc: a stated spread range
that contradicted its own table's 82% row, "within 4%" against a table
reading −4.2%, a narrow-arm ratio fused with a wide-arm attribution, a
spread quoted as 7.5% that was 8% *and* compared against an
incommensurable quantity, the CPU-efficiency guard oversold as "what
makes this table measurable at all" (in-process interleaving is what
protects the ratio; the guard catches only *differential* descheduling),
and a one-directional provenance argument standing in for the direct
control measurement that actually carries the softmax attribution.

**Two further defects I found myself while checking the tables against
each other**, neither raised by the review:

- The `ratio` column is a median of per-rep ratios while the `ns/unit`
columns are medians of times. Medians do not distribute over division,
so every row looked internally inconsistent to anyone who tried to
divide it out (`0.0684 / 0.0617 = 1.109` against a stated `1.117`). Now
documented, along with why the per-rep form is the correct one to quote:
it pairs each MLAS invocation with the native invocation it was
interleaved against, which is the entire point of interleaving. The
then→now figures are relabelled as quotients of medians.
- "wider than nine of the twelve wide-arm rows" was eleven of twelve.

## Adopting #1814

`aee2b9d11` (#1814) landed on `main` while this was in review, and it
closes the exact hole the review found in the guard this document
recommends. A differential CPU-efficiency check cannot see contention
that lands evenly on both arms; #1814's confined-set meter reads busy
jiffies on the process's own `Cpus_allowed_list` and subtracts the
process's own CPU, so foreign load shows up directly. The rule now
points at it, and the tables here are explicitly marked as predating it
and guarded by the weaker method.

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants