Repository navigation
Absorb MLAS's M=1 dense-f32 GEMV into the SimdX86 backend (no resident copy) - #1116
Conversation
…t copy) #1045 claimed 4.4x on dense f32 MatMul with `--features mlas`; #1091 asks to make that real for a default (no-mlas) build. Measured in one binary on this host (RTX 4060 laptop, 20 logical CPUs, AVX2+FMA, no AVX-512), the 4.4x does NOT reproduce: at M=128 prefill our built-in `SimdX86` 6x16 packed microkernel is already at parity with MLAS (0.87-1.15x). The entire reproducible gap is at M=1 decode GEMV: 2.2-4.6x. Mechanism (source-cited both sides): MLAS `sgemm.cpp` routes M==1 away from its packed GEBP -- "The data from matrix B is not referenced multiple times, so using a local packed buffer is a wasted memory copy" -- to `SgemmKernelM1Avx`, which streams B in place (K-outer, N-inner, sequential). Our `sgemm_simd` calls `pack_b` unconditionally, so at M=1 it pays a full read+write copy of B reused zero times: ~3x the memory traffic. It is memory traffic, not arithmetic and not layout, and the fix needs no resident buffer. This ports the mechanism natively: `sgemm_simd_m1` streams B once (K unrolled x4, sequential N sweep, C accumulated in cache), no pack, no scratch, no OnceLock/GovernedWeightCache -- it removes the `bpack` allocation at M=1. Selected for M==1 behind the same-binary A/B toggle `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` (default off, like #1104's `ONNX_GENAI_CPU_MM_INT4_NBLK`). Measured (process CPU time / peak RSS, 5 decode shapes, min-of-30): mlas 52.5 s / 2978 MB simd packed (off) 169.4 s / 2982 MB simd GEMV (on) 57.3 s / 2977 MB 2.96x faster than the packed path, within 1.09x of MLAS, at identical RSS -- 95.9% of the MLAS gap recovered with zero added footprint. Not byte-identical to the packed path (f32 summation is reassociated); matches the naive f64/generic reference within the same tolerance the existing SimdX86-vs-reference tests use. No int4 model exercises the dense f32 path (they route through MatMulNBits), so the A/B is a synthetic in-binary driver (`bench_f32_gemm_ab`), reported as such. Refs #1091 #1045 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #1116 +/- ##
==========================================
- Coverage 80.48% 79.83% -0.65%
==========================================
Files 367 367
Lines 157217 157756 +539
Branches 157217 157756 +539
==========================================
- Hits 126531 125948 -583
- Misses 25974 27084 +1110
- Partials 4712 4724 +12
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
|
Ran your harness myself. The headline finding reproduces and is important; one shape needs explaining before I merge. Your central correction reproduces, and it mattersWith the toggle off, on this host (20 logical CPUs, AVX2+FMA+F16C+AVX-VNNI, no AVX-512),
#1045's 4.4x does not reproduce here, and at prefill shapes we are already at or ahead of MLAS -- 0.57-1.05x at M=128. The entire reproducible gap is M=1. That is a substantive correction to how that PR framed the problem, and it is the kind of thing that only shows up when someone re-measures a claim on different hardware instead of inheriting it. The mechanism you cite makes it make sense: MLAS routes M==1 to What I cannot yet confirm: one shape looks like a regressionWith the toggle on, within-run ratios:
Four of five improve, and the But Before merging, please:
On the numerical changeNot being byte-identical to the packed path is acceptable if it stays behind a default-off toggle and the deviation is quantified, which you did against the f64/Generic reference. Keep it default-off until the shape question is settled -- and when it is eventually flipped on, that flip is a separate decision that needs the deviation restated, because "default off, numerically different" and "default on, numerically different" are very different claims to a user. |
|
Addendum — my "on" run has a built-in control group, and it says the run was too noisy to judge. The full table with the toggle on:
The M=128 rows cannot be affected by an M=1-only route, so they are a control. They moved by up to 1.5x (1.05 → 0.70) and 1.28x (0.72 → 0.92) anyway. That sets the noise floor for the whole run. Two consequences:
Please keep these M=128 rows in the harness and report them explicitly as a control in every future run. They cost nothing extra, they are already there, and they turn "the machine was busy" from an excuse into a measurement: if the control rows move, the run is not usable, and you know that before drawing a conclusion rather than after someone questions it. This generalises beyond this PR. Every A/B we run on this shared box would benefit from an arm the change provably cannot touch. I have been using repeat-spread within one arm for that, which works but costs extra runs; an untouched shape in the same process is cheaper and stronger. |
This machine is shared with build agents, and the failure mode is not "numbers are a bit off" -- it is publishing a conclusion that reverses when the box is quiet. Three habits, learned the expensive way over the past two days. **CPU time, not wall clock.** Three identical runs of one configuration measured 39.3 / 25.8 / 16.1 s wall while `TotalProcessorTime` reproduced to ~2%. Includes the RSS-polling snippet, since `PeakWorkingSet64` reads 0 after exit and must be sampled by PID while the process runs. **Compare an effect against the spread of its own arm.** A 14B model showed ~20% within-arm wall spread; anything under ~1.3x there is unmeasured. **Best: a control arm the change provably cannot touch.** This caught a real case on #1116: an M=1-only GEMV toggle was under evaluation and the M=128 rows in the same harness -- which that route cannot reach -- moved 1.05x to 0.70x and 0.72x to 0.92x between runs, setting a ~1.5x noise floor and making an apparent single-shape regression unadjudicable. Without the control that would have been argued about; with it the answer was simply "re-run when quiet". A control costs one extra shape in a harness you already have and is stronger than a distribution. Also records the differencing method for separating fixed from per-unit cost: decode per token by differencing two token counts, prefill per token by taking the slope between two prompt lengths. A single long run cannot separate them and a single short run is almost all fixed cost; both mistakes have been made here, including by me. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
Address #1116 review: re-measure both arms interleaved in one process, add a built-in M=128 control with a usability rule, and adjudicate 1x5120x7168. - x86_sgemm: split sgemm_simd into an env wrapper + sgemm_simd_variant(..., use_m1_gemv) + sgemm_simd_packed, so the A/B harness can drive both variants in one process without touching process-global env. For m>=2 both variants run identical packed code -- the basis of the control. - bench_f32_gemm_ab: interleave arms at the iteration level (shared conditions), print the M=128 rows as a control every run with a usability verdict (gemv/packed must be ~1.0; drift beyond threshold => RUN NOT USABLE). - Document the measured per-shape dispatch boundary at the M==1 route: the GEMV strictly dominates the packed default on all five decode shapes (2.0-2.9x), so no packed fall-back (it would be slower); the only residual gap is vs MLAS on the two largest shapes, recorded as a known limit. Clean, control-passing results (process CPU time; wall clock is too noisy for the 20-thread M=128 rows on this shared box, ~25% floor, so the control is gated on CPU time which reproduces to ~2-4%): CONTROL M=128 gemv/packed = 1.02x, 1.04x (<5%) -> usable; RSS packed==gemv. DECODE M=1 gemv/packed per shape: 0.50 0.34 0.43 0.42 0.34 (2.0-2.9x faster). 1x5120x7168 = 0.50 (2x faster than packed) -- the earlier "wrong way" was noise. Toggle stays default-off; numerics unchanged (matches f64/Generic ref within 1e-3*(1+|e|)). Gates: cargo test -p onnx-runtime-ep-cpu --lib x5 = 1321 passed / 0 failed / 12 ignored (all five); cargo clippy -p onnx-runtime-ep-cpu --lib -D warnings clean. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
Re-measurement (control-gated) — pushed
|
| rep | simd_packed (CPU s) | simd_gemv (CPU s) | mlas (CPU s) | gemv/packed drift | verdict |
|---|---|---|---|---|---|
| 1 | 80.06 | 81.77 | 79.75 | 2.1% | OK |
| 2 | 79.25 | 82.14 | 74.98 | 3.6% | OK |
Both < 5% ⇒ usable. Peak RSS: packed == gemv == 293 MB, mlas 286 MB. The rule is now built into bench_f32_gemm_ab and I've written it up as the recommended pattern for every A/B harness in the repo: the control rows cost nothing and turn "the box was busy" into a measurement you make before concluding. For m≥2 both simd arms literally call the same function, so their ratio is a pure noise gauge.
4. M=128 untouched — confirmed
Control above: gemv/packed = 1.02–1.04x. Our packed prefill also ≈ MLAS (79–80 vs 75–80 CPU s), consistent with your finding that prefill is already at parity and the entire gap is M=1.
DECODE (M=1) — the actual result
Aggregate over the 5 decode shapes, 2 reps: gemv/packed = 0.33x, 0.36x (≈3x faster than our default) and gemv/mlas = 0.85x, 0.93x (gemv edges ahead of MLAS in aggregate). Peak RSS 2977 vs 2982 MB — identical, bpack removed at M=1.
Per-shape (CPU s):
| shape (M=1) | mlas | packed | gemv | gemv/packed | gemv/mlas |
|---|---|---|---|---|---|
| 1×5120×7168 (QKV) | 5.20 | 7.08 | 3.55 | 0.50 | 0.68 |
| 1×5120×5120 (o_proj) | 2.34 | 6.47 | 2.17 | 0.34 | 0.93 |
| 1×5120×13824 (gate/up) | 6.95 | 12.52 | 5.33 | 0.43 | 0.77 |
| 1×13824×5120 (down) | 4.59 | 13.31 | 5.63 | 0.42 | 1.22 |
| 1×5120×152064 (lm_head) | 37.28 | 130.19 | 44.03 | 0.34 | 1.18 |
3. 1x5120x7168 — explained
On robust CPU time it is gemv/packed = 0.50 (2× faster than packed), gemv/mlas = 0.68 (1.5× faster than MLAS) — one of the best rows, not a loss. The "1.71x → 4.52x" you saw was exactly the wall-clock noise the control now catches before concluding.
So there is no M=1 shape where GEMV loses to the packed default (0.34–0.50 across all five). A dispatch fall-back to packed would therefore be strictly slower and is deliberately absent. The only residual gap is versus MLAS — not our shipping default — and only on the two largest shapes (down_proj K=13824 → 1.22, lm_head N=152064 → 1.18), where sequential-B streaming trails MLAS's blocked M=1 asm. That MLAS-only boundary is now documented in code at the dispatch site in sgemm_simd_variant, with the per-shape numbers, so the win/loss structure lives at the dispatch, not only here.
Numerics — unchanged stance
Toggle stays default-off. GEMV reassociates the f32 sum vs packed (not byte-identical) but matches the f64/Generic reference within 1e-3·(1+|e|), the same tolerance the existing SimdX86 tests use. "Default off, numerically different"; flipping it on remains a separate decision that must restate the deviation.
Gates
cargo test -p onnx-runtime-ep-cpu --lib×5: 1321 passed / 0 failed / 12 ignored (all five).cargo clippy -p onnx-runtime-ep-cpu --lib -- -D warnings: clean.
Discard-and-retry provenance (wall-clock, in-process interleaved) + why the CPU-time number standsTaking your practical note: I stopped waiting for a quiet box and instead ran the harness repeatedly, keeping only control-passing runs. The interleave you suggested is already in the pushed harness — arms alternate per iteration inside one process, so both simd variants see the same load moment-to-moment ( 8 consecutive wall-clock runs, all gated on the M=128 control
8/8 discarded. The M=1 signal survives the noise anywayEven across all 8 discarded runs, the decode result never wavered in direction —
GEMV beats the packed default on every M=1 shape in every one of 8 runs, including the ~60%-noise run. The win is large enough to be unambiguous even when the machine is loud — which is the strongest possible answer to "was Why the merge number is the CPU-time oneWall clock cannot be certified on this box while #1117 builds (table above). Process CPU time can: it is immune to the preemption that inflates wall time, and the CPU-time control passes (gemv/packed = 1.02×, 1.04× on M=128; drift 2.1%, 3.6% < 5%). Those are the numbers in my prior comment and they stand:
On the harness itselfPer-iteration interleave + the M=128 control + usability verdict are in No code change in this update — it's the measurement provenance you asked for. The kernel, dispatch-site boundary doc, gates (5×1321 passed / clippy clean), and default-off numerics are as in |
|
The provenance approach worked, and it answered the question I was blocked on. Merging. The control did its job automaticallyEight consecutive in-process interleaved runs, every one discarded by the control rule — M=128 drifts of 17.2 / 15.0 / 25.6 / 24.0 / 59.7 / 28.7 / 18.6 / 11.2%, because another agent was saturating all 20 cores. That is exactly what a control is for: the run announced its own unusability before anyone drew a conclusion from it. Reporting the discards rather than quietly keeping the best-looking run is the part I want to see repeated. "Eight runs discarded, here is each drift, here is the one whose control passes" is far stronger than a single clean-looking table with no provenance. And the signal survived the noise, which settles my open questionI had flagged That is a stronger form of evidence than a clean measurement would have been. A single quiet run shows the effect under one condition; eight noisy runs that all agree in direction show it survives conditions. My flagged regression was an artefact of comparing arms measured at different times — the same mistake I made again an hour later on #1126 and had to retract. Defensible numbers (CPU time, whose control passes at 1.02x/1.04x, drift 2.1%/3.6%): decode What lands
The self-certifying in-binary control via |
`Fast (Linux x86_64)` is red on main for every open PR: #1116 landed with rustfmt drift. ``` Diff in crates/onnx-runtime-ep-cpu/src/kernels/matmul.rs:4950 Diff in crates/onnx-runtime-ep-cpu/src/kernels/x86_sgemm.rs:349 ``` Verified on a pristine detached checkout of `origin/main` @ `1f1ce4b74`, so it is not introduced by any open branch. This is pure `cargo fmt --all` output — no semantic change, no test change. Co-authored-by: Deckard <deckard@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… the deliverable Updates the ledger: dense f32 M=1 GEMV is absorbed (#1116), and #1045's headline 4.4x is recorded as **not reproducing** on this host -- `simd/mlas` was already 0.57-1.05x at M=128, so the entire reproducible gap was M=1 decode. Inheriting that figure would have sent someone optimising prefill, which was not the problem. The fix was to stop packing B at M==1, matching MLAS's own reasoning that packing a matrix referenced once is a wasted copy: a win from doing less work. Also records a pattern that has now held three times in a row. Each brief predicted a mechanism and the measurement found a different one -- #1104 expected layout and found register blocking, #1116 expected a 4.4x prefill gap and found the gap was entirely at M=1, #1126 expected missing GEMM blocking and found per-row dispatch and allocation overhead. In all three the correction was worth more than the patch. The point is not that briefs are unreliable: each hypothesis was specific enough to direct a measurement that could refute it, which is what a hypothesis is for. The point is to ask for the mechanism *before* the kernel, because a plan is cheap to change then and expensive afterwards -- #1104's transient-tile design was abandoned as unnecessary rather than built and then found unnecessary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
…ed sums) Replace the row-serial borrowed int4 prefill path's two per-token overheads with a structural rewrite behind a default-off A/B toggle (ONNX_GENAI_CPU_MM_INT4_PREFILL), per the method used in #1021/#1027/#1104/#1116: - One fork-join over disjoint column strips for the whole prefill, instead of the row-serial path's m per-row fork-joins. - The per-block �ctivation_sums Vec is hoisted out of the per-row loop to a single m * block_count allocation, independent of the weight size. Rows are still visited outer-most, so a column's packed bytes are re-read once per row: weight traffic and the per-element k-reduction order are unchanged, so output is byte-identical to the row-serial path (both call the shared �orrowed_int4_output_element). No resident buffer is added (peak RSS unchanged), satisfying the #1056/#1117 no-session-buffer constraint. GEMM blocking (reusing a column's bytes across a tile of rows) is deliberately NOT included here: measured within run-to-run noise on qwen05b and with no signal on qwen14b (5-rep interleaved), so it is left as an unproven follow-up to #1117 that can be added or dropped cleanly. Measured win is model-size dependent: qwen05b prefill slope 1.219 -> 0.842 CPU s/token (median of 5, ~1.45x); qwen14b shows no measurable change (medians 8.69/8.62 off/on, within a 30-40% within-arm spread) because the fixed per-row overheads this removes are a negligible fraction of the 14B's larger per-row compute. Refs #1117 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
Merging main brought in #1116's `x86_sgemm` and #1126's int4 prefill work, which made two things in this branch wrong. **The bench was measuring the wrong clock.** It reported process CPU time only, on the reasoning that MLAS might thread where our baselines do not. That is backwards for a graduation gate: our `x86_sgemm` parallelises over column strips while MLAS deliberately declines to parallelise some shapes, so CPU time can show a native route "losing" by 10x precisely when it wins on latency. Wall time now decides, with CPU time printed beside it as `cpu_ratio` so the opposite failure — a route that wins by recruiting the whole machine — is visible too, and flagged as `native-graduates-but-costs-more-cpu` rather than passing silently. The bench also names the native backend it measured, because the same table means different things on an AVX2 host and a host without it. **`matmul.rs`'s module doc contradicted `backend.rs` after the merge.** It still said MLAS was opt-in and reached only via `NXRT_CPU_GEMM_BACKEND=mlas`, which stopped being true on this branch. Same for `KERNEL_PERF.md`, and `CROSS_PLATFORM.md`'s P1 row claiming an MLAS build would fail on Windows MSVC and macOS — `mlas-sys` compiles MASM and Apple sources today, so that row is marked resolved with the real remaining gap named instead. Re-measured everything on the merged tree. The finding worth keeping: #1116 built the M=1 GEMV absorption (stream B in place rather than packing panels reused zero times) but left it behind `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV`, default off, pending a measurement. This bench is that instrument, so the measurement is now in the migration doc: 2.4x faster native at 1x2048x2048, one-sided, and it moves nothing but M=1. Still short of MLAS on this host, so it does not graduate — but it is the shortest path to the first f32 GEMM graduation. Not flipped here; this PR is infrastructure. Also recorded that the GEMM numbers came off a busy shared container and moved up to 2x between runs, and that the bench calls `gemm_with_backend` directly so neither route gets prepacked weights. Without both caveats the table reads as a measurement rather than the order-of-magnitude it is. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…d a knob that no longer exists (#1822) Follow-up to #1173, correcting two defects I shipped in it and repairing the rule they undermined. Docs, one ledger string, one new test, one new script. No production kernel or routing change. ## 1. The ledger named a route gate that had already been deleted `PLAN[MatMulF32].shape_gate` said the native `SimdX86` route "gates M=1 on `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` (default off, #1116)". #1183 shipped that GEMV on by default and removed the probe. `git merge-base --is-ancestor 5417d04 bdb4599` confirms it landed **before** #1173 merged — so the ledger was wrong the day it landed. Today `sgemm_simd` calls `sgemm_simd_variant(a, b, c, m, k, n, true)` unconditionally and `use_m1_gemv` is a plain parameter that only the in-process A/B harness passes as `false`. No environment variable reaches that route. `docs/performance/CPU_MATMUL_ASSIGNMENT.md:559` already recorded the correct fact ("It is measured now, and the route is the default. There is no env probe on the dispatch any more"). Two files in the same directory disagreed and nothing compared them. **Now guarded.** `ledger_prose_only_names_environment_variables_that_still_exist` requires every `NXRT_*` / `ONNX_GENAI_*` token in the ledger's prose to still exist as a string literal in the crate's sources. It cannot check that the description is *right*, only that the knob is *real* — which is the half that goes stale silently. Mutation-verified, not just observed green: ``` matmul_f32: ledger prose names environment variable `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV`, but no source file in this crate contains the literal "ONNX_GENAI_CPU_MM_SIMD_M1_GEMV". ``` ## 2. The doc published a toggle A/B that could not have been run #1173 carried a table captioned **"same binary, same session, toggle the only difference"**, reporting `decode 1×2048×2048` at 0.146 with `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` off against 0.337 with it on, and called turning it on "the obvious next slice". Nothing reads that variable. Setting it measures the same route twice; it cannot produce two different columns. The table is withdrawn and the retraction kept in the text rather than quietly deleted. This is the failure mode the document's own graduation rule warns about — **an arm that was not on the route it was labelled with** — committed by the document that wrote the rule. It survived review because a plausible number in a well-formed table is not self-evidently unmeasured. Readers are pointed at `bench_f32_gemm_ab`, which holds the route as a function parameter and carries the M≥2 rows as a built-in control. ## 3. The gap table is re-measured and the ≥5% rule is repaired The old table was one unguarded invocation per row at an unstated width, taken before the decode-placement corrections (#1729, #1794, #1811) — i.e. when the decode pool put 16 workers on 8 physical cores. New harness: `scripts/bench_native_vs_mlas_width.py`. Arms interleaved rep by rep so host drift lands on both equally; per-rep `os.wait4` CPU-efficiency guard adapted from #1809; six reps per arm; two widths. Raw verdicts, spreads and discards are all reported rather than summarised away. **Three findings, all about method rather than kernels.** | | narrow (6 cores, 1 L3) | wide (32 logical CPUs) | |---|---|---| | `matmul_f32 16×512×512` | 1.581, spread 41% | 0.866, spread 134% | | `matmul_f32 decode 1×2048×2048` | 1.117, spread 21% | 0.934, spread 13% | - **Two cases change verdict on width alone.** Same binary, same half-hour, only the CPU mask differs. `x86_sgemm` parallelises over column strips and MLAS declines to parallelise some shapes, so interleaving the two *routes* inside one process does not protect the ratio — it changes both at once. - **`16×512×512` disagrees with itself on both arms**, alternating `keep-mlas` / `native-graduates` from a byte-identical binary. **One more run of the old table could have graduated a route on this row.** - **The narrow arm is more trustworthy despite having fewer cores** — spreads 4–42% against 5–134%, and it lost no reps to the guard. Isolation beat parallelism. **Softmax now decomposes cleanly**, because no vendored MLAS kernel has changed since #1173 (the only `mlas-sys` edits are the additive straggler handshake in `work_stealing_pool.rs`, #828/#1714, which adds waiting). At matched width the MLAS control arm is stationary to within 4% while native improved **1.24–1.27×** — matching #1416's claim for the row kernel. The f32 GEMM rows get no such attribution and now say so explicitly: their control moved **2.0× the wrong way**, so only the current ratio at a stated width is defensible. **The rule gains what it lacked**: spread must be smaller than the claimed win; reps that did not get the CPU are discarded rather than averaged; a verdict is valid only at a stated width. Under it, `decode 1×2048×2048` — the first f32 GEMM case to show a real native win — **still does not graduate**: it costs more CPU (cpu_ratio 0.875), does not hold at 32 threads, and its 21% spread exceeds its 12% win. ## The width claim is verified, not asserted #1815 landed while this was in progress and observed the neighbouring `bench_generic` harness spawning its ORT arm *outside* the affinity confinement it applied to the native arm. That hazard applies to any `taskset` claim, including mine, so I checked it instead of trusting it — sampling `Cpus_allowed_list` from `/proc/<pid>/task/*/status` 40× across a live narrow-arm run: ``` '16,20,22,26,28,30': 478 observations native_vs_mlas- 273, mlas-sys-ws-0..4 39 each, nxrt-task-0..4 2 each '0-31': 1 (the taskset process itself, before exec) ``` Both routes confined identically; no thread escaped. The rule now requires this check. ## Validation - `dispatch_ledger` **17/17**, including the new falsifier, after merging latest `main`. - `default_artifacts_are_mlas_free` **9/9** — the no-MLAS-in-defaults invariant is untouched. - `cargo clippy -p onnx-runtime-ep-cpu --lib --all-targets` clean; `cargo fmt --check` clean. - Normal merge of `origin/main` (`aee2b9d11`), no rebase, no conflicts. ## Limitations - The narrow arm is six cores on one L3 of one x86-64 host. Nothing here transfers to aarch64 or to a two-socket box. - The `activations erf 1 Mi` row shows native 13.5% slower at matched width. The nearest scatter figure is the wide arm's 8% spread, but that is a spread of *ratios* against a move in a *native time*, so the two are not strictly commensurable. Its MLAS control also moved 11%. **Flagged for pinned re-measurement, not reported as a regression.** - The wide arm was taken with ~4–5 cores of unrelated load present. That is stated in the doc rather than hidden, and it is why its spreads are wider; the guard reports which reps were discarded instead of pretending the host was quiet. - No production behaviour changes here, so there is no performance claim to make about the shipped artifact. Refs #1173, #1183, #1809, #1815, #1416. ## Independent review, and what it changed An independent adversarial review of the full diff returned **no blockers** — it confirmed the ancestry argument behind the retraction, the stationary-control premise for the softmax attribution, and that the headline case is correctly *refused* by the rule (21% spread against a 12% win). It also found seven real defects, all now fixed in `f0323f9ed`. The one that mattered most was in the new test. It only proved the variable name appeared *somewhere* in the crate, so a variable whose read site had been deleted but whose name survived in an `EnvVarGuard::set(...)` line would still have passed — which is the precise shape of the defect this PR exists to correct. The test now requires the matching line to be an `env::var(` / `env::var_os(` read or an `_ENV: &str =` binding. Verified by mutation in **both** directions: | mutation | before | after | |---|---|---| | reinsert retired `ONNX_GENAI_CPU_MM_SIMD_M1_GEMV` into ledger prose | fails ✅ | fails ✅ | | retire the two real `NXRT_CPU_GEMM_BACKEND` reads, leaving the literal only in test guards | **passes ❌** | fails ✅ | The remaining six were prose defects in the doc: a stated spread range that contradicted its own table's 82% row, "within 4%" against a table reading −4.2%, a narrow-arm ratio fused with a wide-arm attribution, a spread quoted as 7.5% that was 8% *and* compared against an incommensurable quantity, the CPU-efficiency guard oversold as "what makes this table measurable at all" (in-process interleaving is what protects the ratio; the guard catches only *differential* descheduling), and a one-directional provenance argument standing in for the direct control measurement that actually carries the softmax attribution. **Two further defects I found myself while checking the tables against each other**, neither raised by the review: - The `ratio` column is a median of per-rep ratios while the `ns/unit` columns are medians of times. Medians do not distribute over division, so every row looked internally inconsistent to anyone who tried to divide it out (`0.0684 / 0.0617 = 1.109` against a stated `1.117`). Now documented, along with why the per-rep form is the correct one to quote: it pairs each MLAS invocation with the native invocation it was interleaved against, which is the entire point of interleaving. The then→now figures are relabelled as quotients of medians. - "wider than nine of the twelve wide-arm rows" was eleven of twelve. ## Adopting #1814 `aee2b9d11` (#1814) landed on `main` while this was in review, and it closes the exact hole the review found in the guard this document recommends. A differential CPU-efficiency check cannot see contention that lands evenly on both arms; #1814's confined-set meter reads busy jiffies on the process's own `Cpus_allowed_list` and subtracts the process's own CPU, so foreign load shows up directly. The rule now points at it, and the tables here are explicitly marked as predating it and guarded by the weaker method. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Summary
#1045 won 4.4x on dense f32 MatMul with
--features mlas; #1091 asks to make that real for a default (no-mlas) build, by absorbing the mechanism into our ownSimdX86kernel rather than shipping behind MLAS. This PR measures the gap on this host, finds where it comes from, proves it is reachable without MLAS's session-lifetime packed buffer, and ports it. Same shape as #1104.The 4.4x prefill number does not reproduce on this host. In one binary containing both paths (same-binary A/B via the existing
NXRT_CPU_GEMM_BACKEND=mlas|simdtoggle), at M=128 prefill our built-inSimdX866×16 packed microkernel is already at parity with MLAS (0.87–1.15x, net slightly favoringSimdX86). #1045's 4.4x was an AMD EPYC without AVX-512; here it's gone.The entire reproducible gap is at M=1 decode GEMV: 2.2–4.6x.
Mechanism — source-cited, both sides
MLAS
sgemm.cpp(vendored,MlasSgemmOperation):MLAS routes M==1 to
SgemmKernelM1Avx.asm, which streams B in place — K unrolled ×4 (ProcessRowLoop4), N swept contiguously (ProcessColumnLoop) — no pack, no resident buffer.Ours (
x86_sgemm.rs::sgemm_simd) callspack_binto abpackscratch unconditionally. At M=1 there is a single A-panel, so each packed B panel is reused zero times — the pack is a wasted full read+write copy of B (K·N f32), ≈3× the memory traffic of a straight GEMV. It is memory traffic, not arithmetic and not layout, and the fix needs no resident buffer.How much is reachable without a resident copy — all of it
sgemm_simd_m1: for M==1, stream B exactly once (K unrolled ×4, sequential N sweep, C accumulated in cache, Rayon over disjoint column strips). Nopack_b, no scratch, noOnceLock, noGovernedWeightCache— it actually removes thebpackallocation at M=1. Exactly #1104's "no resident copy" property; nothing to admit/decline under #1056.The first attempt (column-major, C in registers) regressed lm_head to 3.72x because it strided B by N (608 KB stride → TLB thrash). Matching MLAS's K-outer / N-inner sequential traversal fixed it — the layout that matters for wide outputs.
A/B result — process CPU time and peak RSS
Same binary,
SimdX86M=1 route toggled byONNX_GENAI_CPU_MM_SIMD_M1_GEMV(default off, like #1104'sONNX_GENAI_CPU_MM_INT4_NBLK). One arm per process; peak RSS polled by PID every 150 ms; process CPU time (TotalProcessorTime); 5 decode shapes, min-of-30.SgemmKernelM1)2.96× faster than the packed path (169.4 → 57.3 s), within 1.09× of MLAS, at identical peak RSS. Recovered fraction of the MLAS gap:
(169.4 − 57.3)/(169.4 − 52.5)= 95.9%, with zero added footprint. Per-shapesimd/mlasafter: 5120×5120 1.39x, 5120×7168 1.12x, 5120×13824 1.22x, 13824×5120 1.10x, lm_head 1.11x (all down from 2.2–4.6x).Numerical output
Not byte-identical to the packed path — the GEMV reassociates the f32 sum (K-unrolled-by-4 running accumulation vs the packed KC-panel order). It matches the naive f64 / Generic reference within the same tolerance the existing
SimdX86-vs-reference tests use (1e-3·(1+|e|)), and a new test asserts GEMV-vs-packed agreement within that bound (they differ only by summation order, never in which products are summed). This is reported as a numerical change, not shipped silently: the toggle defaults off.What could not be ported / caveats
qwen2.5-14b-f32,qwen2.5-14b-onnx, and every qwen05b variant route their weights throughMatMulNBits(int4), which does not touch the dense f32 GEMM. So there is no end-to-end token-identity check here; the A/B is a synthetic in-binary driver (bench_f32_gemm_ab,#[ignore]), reported honestly as such rather than as a model number that never took the path.SimdX86in a follow-up once the reassociation is signed off, which is what makes the MLAS-routed speedups do not reach a default build, and the strategy is to absorb them natively #1091 win reach users.Gates (on this host, not CI)
cargo test -p onnx-runtime-ep-cpu --lib×5: 1321 / 1321 / 1321 / 1321 / 1321 passed, 0 failed, 12 ignored each.cargo clippy -p onnx-runtime-ep-cpu --lib -- -D warnings: clean. Also--tests --features mlas: clean.m1_gemv_shapes,m1_route_matches_packed_within_tolerance.Refs #1091 #1045. Precedent #1104.