Repository navigation
Govern the weight-transpose cache under the memory plan (#1056 item 2) - #1079
Conversation
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #1079 +/- ##
==========================================
+ Coverage 79.70% 80.36% +0.65%
==========================================
Files 369 369
Lines 160559 161491 +932
Branches 160559 161491 +932
==========================================
+ Hits 127975 129778 +1803
+ Misses 27857 26970 -887
- Partials 4727 4743 +16
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
|
| Status | Scenario | Base | PR | Change |
|---|---|---|---|---|
sampling_latency/top_k_per_token |
52.02 µs | 64.25 µs | +23.5% | |
kv_cache/alloc_dealloc_pages |
40.45 µs | 49.54 µs | +22.5% | |
qwen3_sampling_processors/top_k_top_p_full_sort_baseline |
5.66 ms | 6.81 ms | +20.4% | |
grammar_masking/llguidance_compute_mask/32 |
76.67 µs | 91.22 µs | +19.0% | |
sampling_latency/top_p_per_token |
383.84 µs | 449.04 µs | +17.0% | |
tokenization/decode_tokens_per_second |
6.15 ms | 7.14 ms | +16.1% | |
logit_processing/seven_processor_chain_per_step |
327.66 µs | 377.64 µs | +15.3% | |
qwen3_sampling_processors/top_k_full_sort_baseline |
2.14 ms | 2.47 ms | +15.1% | |
| ✅ | sampling_latency/min_p_per_token |
209.56 µs | 240.99 µs | +15.0% |
| ✅ | qwen3_sampling_processors/top_k_partial_selection |
139.69 µs | 159.25 µs | +14.0% |
| ✅ | tokenization/encode_tokens_per_second |
395.73 µs | 447.95 µs | +13.2% |
| ✅ | qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline |
3.52 ms | 3.98 ms | +13.1% |
| ✅ | sampling_latency/greedy_per_token |
3.19 µs | 3.60 µs | +12.6% |
| ✅ | qwen3_sampling_processors/top_k_top_p_fast |
655.96 µs | 723.59 µs | +10.3% |
| ✅ | qwen3_sampling_processors/top_p_fast_after_top_k |
531.29 µs | 581.20 µs | +9.4% |
| ✅ | add/large_f32_threads=1-internal/4194304 |
631.02 µs | 651.24 µs | +3.2% |
| ✅ | gather/small_f16_threads=1-internal/4096 |
478.3 ns | 491.5 ns | +2.8% |
| ✅ | reduce_mean/small_f32_threads=1-internal/4096 |
14.85 µs | 15.14 µs | +1.9% |
| ✅ | add/medium_bf16_threads=1-internal/262144 |
102.12 µs | 103.97 µs | +1.8% |
| ✅ | gather/small_f32_threads=1-internal/4096 |
667.8 ns | 665.7 ns | -0.3% |
| ✅ | add/large_bf16_threads=1-internal/4194304 |
1.67 ms | 1.66 ms | -0.7% |
| ✅ | add/medium_f32_threads=1-internal/262144 |
24.95 µs | 24.73 µs | -0.9% |
| ✅ | add/large_f16_threads=1-internal/4194304 |
1.65 ms | 1.64 ms | -1.1% |
| ✅ | reduce_mean/medium_f32_threads=1-internal/65536 |
245.15 µs | 242.49 µs | -1.1% |
| ✅ | gather/small_bf16_threads=1-internal/4096 |
478.9 ns | 473.3 ns | -1.2% |
| ✅ | add/medium_f16_threads=1-internal/262144 |
105.10 µs | 103.47 µs | -1.5% |
| ✅ | reduce_mean/large_f32_threads=1-internal/262144 |
1.01 ms | 990.15 µs | -1.8% |
| ✅ | add/small_f32_threads=1-internal/1024 |
217.2 ns | 211.5 ns | -2.6% |
| ✅ | matmul/medium_generic_f16_threads=1/32x512x512 |
32.80 µs | 30.63 µs | -6.6% |
| ✅ | block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 |
511.32 µs | 476.31 µs | -6.8% |
| ✅ | add/small_f16_threads=1-internal/1024 |
507.3 ns | 451.6 ns | -11.0% |
| ✅ | block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 |
192.54 µs | 168.52 µs | -12.5% |
| ✅ | add/small_bf16_threads=1-internal/1024 |
512.1 ns | 446.4 ns | -12.8% |
| ✅ | matmul/small_generic_f32_threads=8/1x256x256 |
38.48 µs | 33.23 µs | -13.6% |
| 🟢 | matmul/small_generic_f32_threads=1/1x256x256 |
42.96 µs | 36.50 µs | -15.0% |
| 🟢 | block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 |
60.16 µs | 49.78 µs | -17.2% |
| 🟢 | matmul/small_generic_f16_threads=1/1x256x256 |
38.28 µs | 31.11 µs | -18.7% |
| 🟢 | gather/medium_f16_threads=1-internal/32768 |
2.95 µs | 2.39 µs | -18.9% |
| 🟢 | gather/medium_f32_threads=1-internal/32768 |
4.68 µs | 3.75 µs | -19.9% |
| 🟢 | gather/medium_bf16_threads=1-internal/32768 |
3.02 µs | 2.38 µs | -21.0% |
| 🟢 | matmul/large_generic_f16_threads=1/32x1024x1024 |
100.98 µs | 79.60 µs | -21.2% |
| 🟢 | gather/large_f32_threads=1-internal/131072 |
37.09 µs | 27.41 µs | -26.1% |
| 🟢 | gather/large_bf16_threads=1-internal/131072 |
14.58 µs | 10.56 µs | -27.6% |
| 🟢 | matmul/small_generic_bf16_threads=1/1x256x256 |
44.11 µs | 31.71 µs | -28.1% |
| 🟢 | matmul/large_generic_f32_threads=8/32x1024x1024 |
5.37 ms | 3.82 ms | -28.9% |
| 🟢 | gather/large_f16_threads=1-internal/131072 |
14.96 µs | 10.42 µs | -30.4% |
| 🟢 | matmul/large_generic_bf16_threads=8/32x1024x1024 |
2.09 ms | 1.37 ms | -34.4% |
| 🟢 | block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 |
922.23 µs | 585.69 µs | -36.5% |
| 🟢 | matmul/medium_generic_f32_threads=1/32x512x512 |
3.74 ms | 2.33 ms | -37.6% |
| 🟢 | block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 |
74.39 µs | 46.34 µs | -37.7% |
| 🟢 | matmul/large_generic_bf16_threads=1/32x1024x1024 |
3.16 ms | 1.96 ms | -38.0% |
| 🟢 | matmul/medium_generic_bf16_threads=1/32x512x512 |
876.76 µs | 527.09 µs | -39.9% |
| 🟢 | matmul/medium_generic_f32_threads=8/32x512x512 |
1.59 ms | 912.62 µs | -42.4% |
| 🟢 | matmul/large_generic_f32_threads=1/32x1024x1024 |
16.49 ms | 9.40 ms | -43.0% |
| 🟢 | matmul/medium_generic_f16_threads=8/32x512x512 |
67.78 µs | 30.46 µs | -55.1% |
| 🟢 | matmul/small_generic_f16_threads=8/1x256x256 |
89.99 µs | 31.16 µs | -65.4% |
| 🟢 | matmul/large_generic_f16_threads=8/32x1024x1024 |
273.29 µs | 85.11 µs | -68.9% |
| 🟢 | matmul/small_generic_bf16_threads=8/1x256x256 |
127.22 µs | 31.86 µs | -75.0% |
| 🟢 | matmul/medium_generic_bf16_threads=8/32x512x512 |
1.58 ms | 370.00 µs | -76.6% |
Visual flags:
Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.75 3.45 4.40 }
What this cannot catch
- Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
- Sub-threshold regressions that compound over multiple PRs
- Performance changes that only manifest under GPU execution
- Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)
|
The design is right and the rule you derived is correct -- I verified it independently. But the PR introduces a process-global mutable flag that tests flip while other tests run, and it is already breaking the suite. Requesting changes. The failure
Both pass in isolation:
Passing alone and failing in the suite is the signature of shared process state, not of a wrong assertion. The mechanism is visible in your own diff: Your own gate reported this to you: the summary says "engine 494 passed (2 pre-existing unrelated failures)" -- those two were mine, fixed on main in This is the third time today that process-global state has produced a test that passes alone and fails in company (#983's ORT handle cache, #1033's fp16 GEMV). Worth treating as a known hazard in this repo rather than an accident. What I want
What I verified and agree withThe rule is right. I checked every call site rather than taking the citation on trust:
So on x86/Windows the only populator is The ratio-1.00 evidence is the right shape, and the finding that the cache keys on The over-prediction case is handled correctly: half-typed Fix the test isolation and I will merge. |
…1056) The process-global weight-transpose cache holds one full K x N f32/f16 copy per transposed constant weight for the session. #1056 requires every session-lifetime, weight-scaled allocation to be declared to the memory plan before it is allocated, in the bytes actually allocated, and to be declinable. Reporting (TransposeCache::bytes / weight_transpose_cache_bytes) landed in 1ae696c; this finishes the governing half. - Predictor: weight_transpose_cache_predicted_bytes(&Graph) mirrors the exact kernel decision. On all platforms it counts constant-weight Gemm nodes with transB != 0 (gemm.rs transposed_b, added in #1035) at N*K*4 f32 bytes; the MatMul/FusedMatMulBias transpose call sites are cfg(macos/ios)-only, so the predictor's MatMul arm is likewise cfg-gated -- a binary predicts exactly what its own kernels allocate. The shape-keyed kernel cache instantiates a node once per activation shape, but every instance keys the global cache on (weight address, K, N), so the second instantiation hits the first entry: the transpose is held once per weight, not once per instantiation (no #1051 per-copy multiplier here). Verified by a test that runs the real Gemm kernel at prefill (m=4) and decode (m=1) and asserts predicted == bytes held, 1.00. - Admission gate: set_weight_transpose_cache_enabled(bool). When declined, MatMulPrepack::transposed_b / transposed_b_f16 return None (kernels recompute a transient transpose per call, freed each call) and cached_transpose_f32/f16 compute without inserting -- nothing session-lifetime accrues. Declining is a pure performance tradeoff: the transpose is byte-identical either way. - Wiring: engine load.rs folds the predicted bytes into resident_f32_cache_bytes and calls set_weight_transpose_cache_enabled beside the resident dequant f32 cache (#987) and MLAS SQNBit packed buffer (#1051) gates, so one plan verdict governs all three. Gates: cargo test -p onnx-runtime-ep-cpu --lib (1216 passed, 0 failed, 11 ignored); cargo clippy -p onnx-runtime-ep-cpu --lib -D warnings clean; engine --features native-backend --lib green except two pre-existing native_decode IO failures unrelated to this change (confirmed failing on origin/main). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
…l harness The #1056 admission flag was a bare process-global AtomicBool that the gemm decline test toggled with set_weight_transpose_cache_enabled(false) while the matmul transposed_b tests read it concurrently in the same process, so those tests took the new declined early-return and failed in company (passing alone). This is the same process-global-races-the-harness trap as #983 and #1033. Make the decline decision consult a thread-local override first; only production writes the global. Add a #[cfg(test)] RAII CacheEnabledScope that sets the override on the current thread and restores the previous value on drop (including on panic), so a test's decline can never leak to another thread. Rework the decline test to probe the exact (addr, k, n) key via a test-only f32_cache_contains peek instead of a global byte delta, so a concurrent test caching an unrelated weight cannot mask a leak. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
…s it The decline test probed f32_cache_contains(b_ptr, k, n), but the Gemm kernel installs the entry via transposed_b(&b, n, k), so the resident key is (b_ptr, n, k). Match that order so the admitted-run assertion sees the entry. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
e4d2a27 to
38af02e
Compare
|
Addressed — thanks for the precise diagnosis. The mechanism was exactly as you described: the admission flag was a bare process-global Fix (your preferred option 1 — scope so tests never mutate process state):
No production code path changed; production still reads the global on every worker thread. Rebased on current Unfiltered gate counts, re-run on this machine after the fix:
I also confirmed Carrying your note forward: any process-global mutable state in this repo should assume the parallel harness will race it, and the test access should be thread-local/RAII-scoped from the start. |
|
Update: the isolation fix works -- the two Same commit
Failing assertion is the first one, Before blaming the test I falsified it, because a test that cannot fail is worse than no test. Removing the decline gate in That is a good finding about the design -- the two gates are genuine defence in depth, either one sufficient -- and it establishes the test can detect a broken decline. So something is really populating the global cache for that key under decline, at least sometimes. Two candidate mechanisms, which need different fixes:
The distinction matters more than the test going green: if a worker thread can bypass the thread-local override in tests, something analogous may bypass the global verdict in production -- for instance a transpose performed before the plan's Sent back for root-cause with evidence, and for five consecutive full-suite runs rather than one. A single green run does not establish stability for a test now known to be order- and timing-sensitive; I made exactly that mistake an hour ago, which is why I re-ran. |
The decline/exactness tests peek the process-global cache by (addr, K, N). An address only names a weight while it is live; once freed, the allocator can recycle it for an unrelated buffer of the same dims, leaving a stale entry that makes f32_cache_contains a false positive for a weight that never went through the cache. That is the intermittent 'a declined transpose cache must retain nothing' failure under the parallel harness. Instrumented the gate sites (thread id, override, global, key) and confirmed the production decline is complete: during decline cached_transpose_f32 is never reached and before/after are both false; every insert observed carries override=Some(true), i.e. the admitted phase. So this is a test-only false positive, not a decline gap. Add a #[cfg(test)] TransposeCache::remove and a f32_cache_evict helper, and evict this weight's exact key at the start of both tests so the 'before' state is deterministically empty. Safe under the harness: no live concurrent allocation can share the address, so the only entry evicted is a stale, unused one. Full suite run five times back to back: 1254 passed / 0 failed each time. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
|
Root-caused with instrumentation, not reasoning — thanks for insisting on that, it settled the important question (production vs. test). What I measuredI instrumented both gate sites to print, on every Two facts fall straight out:
So mechanism 1 (a worker thread bypassing the thread-local override) is ruled out, and — the distinction you flagged as mattering most — the production decline is complete: under decline, The real cause: mechanism 2 (address reuse), and it is test-only
Your "remove both gates → fails, either gate alone → passes" falsification is fully consistent with this: the assertion is genuinely capable of catching a broken decline; it is the starting state that was non-deterministic. Fix (deterministic, not a retry or a relaxed assertion)
Production code is unchanged; this commit is test-only. Stability — five consecutive full runs
Plus the two transpose tests looped 120× in isolation: 0 failures. Pushed as |
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
|
Root cause accepted, independently re-verified, merging. Three consecutive full runs on
Together with your five, that is eight clean runs of a test that previously failed roughly every other time. Good enough. The root cause is the right one and the evidence supports it. The instrumentation showing So the defect was a false positive in the probe, not a leak: Worth noting how consistent this is with the rest of the day: Summary of what lands
That closes item 2 of #1056. The remaining item is the audit for a fourth such buffer. |
## What `cargo fmt --all -- --check` currently **fails on `origin/main`** (`bce03cabb`) in five places across `crates/onnx-runtime-ep-cpu/src/kernels/gemm.rs` and `crates/onnx-runtime-ep-cpu/src/kernels/matmul.rs`. This is the `cargo fmt --all` output and nothing else. ## Why it happened Nobody wrote badly-formatted code. #1073, #1079 and #1080 each touched these two files and each was fmt-clean against its own base. Squash-merging them produced a combined text that rustfmt formats differently: - `TRANSPOSE_TEST_LOCK.lock().unwrap_or_else(|e| e.into_inner())` is 61 characters, which exceeds rustfmt's default `chain_width` of 60 once it sits at test-body indentation (two sites). - The `onnx_runtime_ir` import list grew past `max_width` and now wants braces on their own lines. - One `Gemm`/`transB` attribute chain became short enough to fit on a single line after a neighbouring edit. - One stray double blank line at end of a `mod tests`. `main` is unprotected, so no required check re-ran fmt on the merge result and the breakage landed silently. ## How it was found It blocked #1086: that PR's `Fast (Linux x86_64)` and `Rust quality` jobs failed on the PR **merge ref** with diffs in files #1086 does not touch. Reproduced independently by checking out `origin/main` into a clean worktree and running `cargo fmt --all -- --check`. ## Verification - `cargo fmt --all -- --check` -> clean (was: 5 diffs). - Diff is whitespace/line-breaking only; `git diff -w` on the two files is empty apart from the import-brace move. No logic, no behaviour, no test changes. ## Risk None. Mechanical formatter output. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
#1100) ## #1056: bring `MatMulPrepack::dense` under the memory plan Fourth resident, weight-scaled buffer brought under the memory-strategy plan after the resident dequant f32 cache (#987), the MLAS SQNBit packed buffer (#1051), and the weight-transpose cache (#1079). Follows #1056's rule: *any allocation that outlives a single kernel call and scales with weight size must be declared to the plan before it is allocated, in the bytes actually allocated, and must be declinable.* ### The exact condition under which the kernel caches (verified) `MatMulPrepack::dense(index, view)` in `crates/onnx-runtime-ep-cpu/src/kernels/matmul.rs` caches a session-lifetime `Vec<f32>` iff **both**: 1. the operand is a constant initializer (`constant_inputs[index]`), and 2. `to_dense_f32_widen("MatMul", view)` returns `Cow::Owned` — i.e. the operand is **not** already a contiguous f32 view. `to_dense_f32_widen` (`crates/onnx-runtime-ep-cpu/src/dtype.rs:806`) borrows contiguous f32 zero-copy (`Cow::Borrowed`, no cache) and allocates an owned `4*numel` f32 copy for every other float case: f16/bf16/f64, or a strided/column-major f32. So a contiguous f32 constant costs **nothing** (this is why the buffer stayed invisible on the int4 and f32-contiguous models we exercise most), while an f16/bf16/f64 or non-contiguous constant costs a permanent `4*K*N` per kernel instance. ### The shape-keyed instantiation multiplier (×2) — corrected after review The executor's kernel cache is **shape-keyed** (`KernelKey { node, resolved_input_shapes }`, `crates/onnx-runtime-session/src/executor/kernel_cache.rs`). A decoder instantiates each `MatMul` node at **two** activation shapes — prefill (`m>1`) and decode (`m==1`) — as **separate** `MatMulKernel`s, each with its own `MatMulPrepack::dense`. Unlike the weight-transpose cache (process-global, keyed on the weight address, so a second instantiation reuses the first), `dense` is per-instance. The only case that populates `dense` is a constant non-f32 `B` paired with an f32 `A` (a same-half `B`+`A` takes `try_matmul_half`, and both the `m==1` GEMV and the MLAS half-prefill fast paths require `A` to be half, so they are skipped when `A` is f32). In that case **both** the prefill instance and the decode instance take the generic/direct-f32 GEMM that widens `B`, so **each retains its own `4*K*N` copy** for the session. The resident footprint is therefore `2 × 4·K·N`, not `4·K·N`. The first cut of this PR counted **one** copy — an **under**-prediction by 2×, the exact under-reporting defect #1051 corrected for the MLAS packed buffer (it reported 247 MB where steady state grew to ~592 MB, the missing factor being this same prefill/decode shape-keyed doubling). Fixed by adding `MATMUL_DENSE_DECODE_INSTANTIATIONS = 2` (mirroring #1051's `MLAS_PACKED_DECODE_INSTANTIATIONS`) and multiplying the per-node prediction by it. **Measured, not reasoned**: the ratio test below instantiates prefill + decode and counts two live copies. ### What was built - **Predictor** `matmul_dense_cache_predicted_bytes(&Graph)` — mirrors the kernel condition per node (`MatMul` + `FusedMatMulBias`, both operand indices). A graph initializer is contiguous (`WeightRef` carries only dtype + dims, no strides), so from the graph the condition reduces to *"a constant operand whose dtype is a non-f32 float"* → `4*numel`, **× `MATMUL_DENSE_DECODE_INSTANTIATIONS`** for the prefill+decode instances; a contiguous f32 constant → 0. - **`GovernedWeightCache<f32>`** — `dense: [OnceLock<Vec<f32>>; 2]` → `[GovernedWeightCache<f32>; 2]`. Declined → `dense()` widens transiently per call and frees it. The only extension needed was a read-only `filled()` accessor (reuse an existing session copy without ever running a builder); the "array of two" and the `f32` element type were already expressible, so nothing else changed in that type. - **Plan wiring** in `crates/onnx-genai-engine/src/engine/load.rs`, beside `set_resident_dequant_f32_cache_enabled` / `set_mlas_sqnbit_packing_enabled` / `set_weight_transpose_cache_enabled`, folding the prediction into `resident_f32_cache_bytes` and gating on the same `f32_weight_cache_admitted` verdict — one decision governs all four buffers. - **Decline path** — process-global `AtomicBool` written only by production (`set_matmul_dense_cache_enabled`), with a test-only thread-local RAII `DenseCacheEnabledScope` that restores on drop including on panic. No process-global mutable state tests mutate (#983/#1033/#1079). ### Predicted vs actual (ratio 1.00) — now across both instantiations Proven in one run by `predicted_dense_bytes_equal_actual_after_matmul_execution`: it computes the plan's graph prediction, then — mirroring the shape-keyed kernel cache — builds a **separate** `MatMulKernel` for the prefill shape (`m=4`) and the decode shape (`m=1`), executes each with a constant **f16** `B` and an **f32** `A`, asserts each retains a copy, and compares the **summed** `live_bytes()` against the prediction. | case | K×N | instantiations retaining a copy | per-copy | predicted (×2) | summed actual | ratio | |---|---|---|---|---|---|---| | f16 constant B, f32 A | 48×33 | 2 (prefill m=4 + decode m=1) | 6336 B | 12672 B | 12672 B | **1.00** | The test asserts `summed_actual == predicted` and `instances == MATMUL_DENSE_DECODE_INSTANTIATIONS`, so a dropped multiplier would fail it — the single-instantiation version could not. ### What happens on decline `declined_dense_cache_retains_nothing_and_is_byte_identical` runs one kernel instance admitted vs declined: | arm | bytes held after run | output | |---|---|---| | admitted | `40*24*4` = 3840 B (one copy) | reference | | declined | **0 B** | **byte-identical** | Declining is a pure performance tradeoff: the widened f32 is byte-identical whether cached or recomputed. Because the transient widen is freed at the end of each call, no session-lifetime footprint accrues when declined — and with the ×2 multiplier, admitting an f16-weight model now costs the plan the full `2×` it will actually hold. ### Peak RSS before/after — no available model populates this cache Checked every model in `C:\Users\justinchu\dev\models` for plain `MatMul` nodes with a constant operand: - `qwen2.5-14b-f32`, `qwen2.5-14b-onnx`, and all `qwen*`/`qwen05b*` variants: **0 plain MatMul nodes** — every weight routes through `MatMulNBits` (+ `GatherBlockQuantized`). - `gemma-3-27b-onnx`: exactly **1 plain MatMul**, and **both its operands are activations** (there is a `Transpose` feeding it), so `constant_inputs = [false, false]` and it caches nothing. So there is no local model that genuinely populates `dense`, and a CLI A/B would truthfully report predicted = actual = 0 — which proves nothing. Per #1056 I therefore prove predicted==actual and byte-identical-on-decline on a **synthetic graph** (the tests above) rather than reporting a meaningless zero. An f16 peak-RSS A/B awaits an export that uses plain `MatMul` with f16/non-contiguous constant weights. ### Anything not predicted exactly (over-predict, documented) The predictor **over-predicts, never under-predicts**, with documented directions: 1. **Same-half packed path**: when both operands are the same half dtype and contiguous, the node takes `try_matmul_half` and never calls `dense`, so the true cost is 0 while the predictor still counts the constant operand. Whether the *other* operand is half is not a graph-static property of the constant one, so counting is the safe direction. 2. **Prefill-only / multi-shape workloads**: a prefill-only run instantiates one copy, so the ×2 over-estimates there (safe — the gate declines sooner). Conversely a session presenting *more* than two distinct activation shapes to a node (several prompt lengths across `generate` calls) could instantiate more than two; the multiplier follows #1051's convention of bounding the autoregressive decode workload at prefill + decode, and that residual is the same documented class as #1051's. 3. **Residual gap — non-contiguous f32 constant** (a column-major weight): the kernel *would* cache it, but `WeightRef` exposes no strides, so from the graph it is indistinguishable from a contiguous f32 constant and is **not** counted — a potential under-prediction. No such operand occurs in the models this repo exercises: the one documented column-major weight (the lm_head projection, per `contiguous_b_f16`'s doc) is **f16**, so it *is* counted via the dtype rule. Documented in the predictor rather than over-counting every f32 weight (which would inflate the budget on the overwhelmingly common contiguous-f32 case). If f32 column-major MatMul weights ever appear, the loader must expose their layout for the predictor to see them. ### Gates - `cargo test -p onnx-runtime-ep-cpu --lib` — **five consecutive runs**, all `1299 passed; 0 failed; 11 ignored` (46.0 / 49.9 / 52.6 / 52.7 / 49.6 s). Includes the 4 new tests (3 in `matmul`, 1 `filled()` in `governed_weight_cache`). - `cargo test -p onnx-genai-engine --features native-backend --lib` — `496 passed; 0 failed; 1 ignored`. - `cargo clippy -p onnx-runtime-ep-cpu --lib -- -D warnings` — clean. Closes #1056. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624 --------- Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
## Share one MLAS SQNBit packed buffer per weight (#1056) Refs #1027 (MLAS SQNBit route), #1051 (packed-buffer accounting), #1056 (this dedup). ### The problem `MatMulNBits` int4 `accuracy_level=0` nodes route to MLAS SQNBit CompFp32 (#1027). The executor's `KernelCache` is **shape-keyed**, so each node compiles two kernel instances — prefill (`m > 1`) and decode (`m == 1`) — and before this change **each instance packed its own full copy of the same constant weight**. The resident packed footprint was therefore `2x` the single-copy cost, held for the whole session. ### The fix A process-global, weight-identity-keyed store (`MlasPackedCaches`) keyed on `(address, N, K, bits, block_size, has_zero_points, compute_type)`. The first kernel instance to reach a weight packs it once; the sibling instance takes the same `Arc`. The session now holds **one** packed copy per weight. * Keyed on the mmap **address** plus every pack-determining shape/param, so a same-address different-shape weight (allocator recycling a freed address — the #845/#1079 hazard) misses rather than serving the wrong bytes. * `clear_mlas_packed_caches()` runs on `Executor` drop — the **same lifetime boundary** as `weight_transpose::clear_all` — closing the same-address/same-shape/across-lifetimes window. * Accounting is updated **in the same commit**: `MLAS_PACKED_DECODE_INSTANTIATIONS` goes `2 -> 1`, so `resident_dequant_f32_cache_bytes` (the plan's prediction) equals `mlas_sqnbit_packed_live_bytes` (the actual allocation). The existing accounting test additionally asserts **pointer identity** of the shared `Arc` across the prefill and decode instances. * **No new process-global mutable state that tests mutate** (#983/#1033/#1079): production reads only the global store; tests use a `cfg(test)` **thread-local** store (each libtest thread gets a private cache, so no cross-test recycled-address contamination, while a single test's prefill+decode still share). ### Acceptance criteria **1. Predicted bytes == actual bytes, ratio 1.00.** `resident_f32_cache_bytes` (plan prediction) vs `live_total` (profiler actual `SQNBIT_PACKED_LIVE_BYTES`), measured on-model: | model / arm | predicted (bytes) | actual live (bytes) | ratio | |---|---:|---:|---:| | qwen05b after, admitted, multi-token | 316,443,904 | 316,443,904 | **1.00** | | qwen05b before, admitted, multi-token | 632,887,808 | 632,887,808 | 1.00 | | qwen14b after, admitted (18GiB ceiling) | 8,962,744,320 | 8,962,744,320 | **1.00** | The dedup test (`int4_acc0_mlas_packed_accounting_equals_actual_allocated`) that ties the plan's prediction to the profiler's actual bytes stays green and now also asserts the shared `Arc`. **2. Pack count halves.** `ONNX_GENAI_PROFILE_MM=1`, `[mm_prepack] calls=` on `qwen05b-symzp` (169 weight boundaries): | run | before (origin/main) | after (this branch) | |---|---:|---:| | 1-token (single activation shape) | 169 | 169 | | multi-token (prefill + decode shapes) | **338** | **169** | Before, the multi-token run packed twice as many buffers as the 1-token run; after, they pack the **same** count. (The 1-token run already packed 169 before, but the pre-dedup predictor still accounted `2x = 632,887,808` for it — over-report, the safe direction; after, both the pack count and the accounting are single-copy.) **3. Peak RSS + accounted, with ratios, both models, before/after.** Every number measured on this host (Windows, 68,535,443,456 B RAM, CPU-only, AVX2/FMA/F16C/AVX-VNNI). Peak RSS = polled `PeakWorkingSet64` while running; CPU time = `TotalProcessorTime`. **qwen05b-symzp** (weights 366,846,066 B), multi-token autoregressive run: | arm | packs | accounted | live | peak RSS | admitted | |---|---:|---:|---:|---:|:--:| | route OFF (`QNBIT=0`) | – | 0 | – | 485.9 MB | – | | route ON **before** | 338 | 632,887,808 | 632,887,808 | 1135.9 MB | yes | | route ON **after** | 169 | 316,443,904 | 316,443,904 | **818.5 MB** | yes | The packed accounting halved (632,887,808 -> 316,443,904) and peak RSS dropped **317.4 MB** — almost exactly the one deduplicated packed buffer (316,443,904 B). **qwen14b-symzp** (weights 8,549,241,669 B). Default residency ceiling = `0.25 x RAM` = **17,133,860,864 B**. Admission tests the **expanded** footprint (`on-disk weights + packed cache`): | arm | accounted (predicted) | live | expanded footprint | peak RSS | verdict | |---|---:|---:|---:|---:|:--:| | **before**, default ceiling | 17,925,488,640 | 0 (declined) | 26,474,730,309 | 8643.4 MB | **declined** | | **after**, default ceiling | 8,962,744,320 | 0 (declined) | 17,511,985,989 | 8669.3 MB | **declined** | | **after**, ceiling 18 GiB | 8,962,744,320 | 8,962,744,320 | 17,511,985,989 | 17,606.9 MB | **admitted** (ratio 1.00) | **Does the 14B flip declined -> admitted at the default ceiling? No — but only just, and the reason is precise.** The dedup halved the predicted packed cache (17,925,488,640 -> 8,962,744,320) and shrank the expanded footprint from 26,474,730,309 to 17,511,985,989. But admission compares that **expanded** footprint (weights **+** cache), not the cache alone, against the `0.25 x RAM` ceiling of 17,133,860,864 B. After the dedup the expanded footprint is **17,511,985,989 B — still 378,125,125 B (2.2%) over** the default ceiling, so it stays declined and runs the borrowed zero-copy path (peak ~8.6 GB, unchanged from before). What the dedup *does* change is the admission threshold: admitting the 14B previously required a ceiling >= 26,474,730,309 B = **0.386 of RAM**; it now requires >= 17,511,985,989 B = **0.256 of RAM** — i.e. barely above the 0.25 default. `--host-ram-limit 18GiB` (0.263) now admits it at peak 17,606.9 MB, comfortably inside 68.5 GB. So the dedup moves the 14B from "unreachable without allowing a 26.5 GB expansion" to "one notch above the default," but does **not** cross the 0.25 line on its own on this box. (The original prediction that it would flip rested on comparing the ~8.4 GB cache to the ceiling; the gate actually tests the 17.5 GB expanded footprint.) **4. Byte-identical generated text** (SHA-256 of generated text), greedy decode: | prompt | before | after (declined) | after (admitted) | |---|---|---|---| | qwen05b, "…relativity…" 48 tok | `EF7CA14F…` | `EF7CA14F…` | – | | qwen05b, raw "a" 8 tok | `FD4972FF…` | `FD4972FF…` | – | | qwen14b, "…relativity…" 16 tok | `EB88829D…` | `EB88829D…` | `EB88829D…` | Identical across before/after and across admitted/declined (the borrowed zero-copy path and the MLAS packed path produce the same tokens). qwen05b route ON and route OFF also match (`EF7CA14F…`). **5. No new process-global mutable state that tests mutate.** Production writes only the `LazyLock` global; tests use a `cfg(test)` thread-local store restored automatically by each libtest thread ending. No env/RAII toggles were added to production. ### Gates * `cargo test -p onnx-runtime-ep-cpu --features mlas --lib matmul_nbits` — **110 passed, 0 failed, 6 ignored**. * `cargo test -p onnx-runtime-ep-cpu --lib` — **five consecutive runs**: `1269/0/11`, `1269/0/11`, `1269/0/11`, `1269/0/11`, `1269/0/11` (passed / failed / ignored). * `cargo clippy -p onnx-runtime-ep-cpu --lib -- -D warnings` — clean (both default and `--features mlas`). ### Cases not deduplicated None on the constant-weight route. Every MLAS SQNBit route (`weight_prepacked`, static shards, and the `NO_SHARD` A/B) goes through the shared store. The only non-shared fallback is a weight with no stable contiguous host address to key on — which never occurs on the constant-weight (`can_prepack`) route this touches, since the initializer is a contiguous mmap slice. A non-constant weight rebuilds a transient pack per call and retains nothing, so there is nothing to share. --- _Note: `Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624`._ Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
Bring the weight-transpose cache under the memory plan (#1056, item 2)
The process-global weight-transpose cache holds one full
K x Nf32/f16 copy per transposed constant weight for the session (crates/onnx-runtime-ep-cpu/src/kernels/weight_transpose.rs). #1056's rule: any allocation that outlives a single kernel call and scales with weight size must be declared to the plan before allocation, in the bytes actually allocated, and be declinable. Reporting landed in1ae696c9; this PR finishes the governing half, mirroring #1051 (MLAS packed buffer).The exact rule for when the kernels transpose (file:line citations)
I read the
MatMulandGemmkernels to find every call intocached_transpose_f32/f16. The populating call sites are:gemm.rs:119transposed_b(&b, n, k)N*K*4)transB != 0and constantB;Bis first widened to dense f32 byMatMulPrepack::dense. Added by #1035.matmul.rs:1434,1457transposed_b#[cfg(any(macos, ios))]N*K*4)Bmatmul.rs:735transposed_b_f16#[cfg(any(macos, ios))]N*K*2)Bfused_matmul_bias.rs:71transposed_b_f16#[cfg(any(macos, ios))]N*K*2)Key consequence: on x86/Windows (non-Apple, no Accelerate) the only path that populates the cache is
GemmwithtransBand a constantB. EveryMatMultranspose call site is compiled out. The predictor'sMatMul/FusedMatMulBiasarm is therefore#[cfg(any(macos, ios))]-gated too — a binary predicts exactly what its own kernels allocate.Effect (b) from #1051 — measured, does NOT apply here
The shape-keyed
KernelCacheinstantiates a node once for prefill (m>1) and once for decode (m=1). But every instance keys the global cache on(weight address, K, N), so the decode instance hits the prefill entry and allocates nothing extra. The transpose is held once per weight, not once per instantiation — no per-copy multiplier (unlike the MLAS packed buffer, which retains an owned copy per instance). This is asserted directly by the exactness test, which runs the real kernel atm=4thenm=1and observes a single copy.Predicted-vs-actual (ratio 1.00)
Synthetic exactness test
predicted_transpose_bytes_equal_actual_after_gemm_execution(gemm.rs) — runs the realGemmKernelthrough theKerneltrait at prefill (m=4) and decode (m=1) on a constanttransBweight, then asserts the predictor equals the bytes the process-global cache actually holds for that weight:This asserts predictor-vs-kernel, not predictor-vs-formula, so a drift between the two would fail it.
Real models (predicted measured on this host via the loader + predictor):
qwen2.5-0.5b(f32, 1.98 GB data)qwen2.5-0.5b-q4(int4, 873 MB data)Neither exported model contains a single
Gemmnode, and theirMatMultransposes are Apple-only, so the cache is genuinely empty on x86 — predicted0matches actual0. The non-trivial ratio-1.00 proof is the synthetic test above.RSS measurements (this machine, CPU-only, no CUDA runtime)
Peak working set is per-process and reliable; timing is process CPU time (
TotalProcessorTime), not wall clock.qwen2.5-0.5bqwen2.5-0.5b-q4The memory-strategy plan log confirms the wiring end-to-end on the native path:
scope="single_model_native" resident_f32_cache_bytes=0 f32_weight_cache_admitted=true— the transpose predictor (0) is folded intoresident_f32_cache_bytesandset_weight_transpose_cache_enabled(true)is applied from the same verdict.¹ Both exported models failed to load on the native decode backend on this host with
model.io.position_ids_input declares port 'position_ids', but the graph exposes [...]when I first measured (pre-rebase). That was a model/IO-metadata incompatibility related to the same areac5385b2faddressed onorigin/main; I have not re-measured a full generate post-rebase (the CLI release build is ~50 min on this shared host). Regardless, neither exported model contains aGemmnode, so the transpose cache is empty on x86 and the predictor's0is exact for them either way; the authoritative ratio-1.00 proof is the synthetic executor test above.What happens on decline
set_weight_transpose_cache_enabled(false)(called by the plan when the folded resident total does not fit):MatMulPrepack::transposed_b/transposed_b_f16returnNone→ theGemm/MatMulkernels recompute a transient transpose per call (Cow::Owned(transpose_row_major(...)), freed at the end of the call), so no per-kernel-instanceArcis memoized and nothing session-lifetime accrues.cached_transpose_f32/f16compute and return the transpose without inserting it into the global map, soweight_transpose_cache_bytes()stays put.Test
declined_transpose_cache_retains_nothing_and_is_byte_identicalproves both: after a declined run the global byte total is unchanged (retains nothing), and the declined output is byte-identical to the admitted output — declining is a pure performance tradeoff, never a numerical one (both paths call the sametranspose_row_major).Case I could not exercise exactly
The half-typed
Gemmcase: when both operands are f16/bf16 the node takestry_half_gemmand does not transpose, but whetherAis half is not a graph-static property, so the predictor counts every constanttransBGemmatN*K*4. This over-predicts (never under — the #1056-mandated safe direction) for a halfGemm; it is exact for the f32 case. Documented in the predictor's doc comment. No such node exists in the available models, so it did not affect any measurement.Gates (run locally on this machine, unfiltered, after rebasing on
origin/main@8635e4db)cargo test -p onnx-runtime-ep-cpu --lib→ 1254 passed, 0 failed, 11 ignored (includes the 2 new tests)cargo clippy -p onnx-runtime-ep-cpu --lib -- -D warnings→ cleancargo test -p onnx-genai-engine --features native-backend --lib→ 496 passed, 0 failed, 1 ignoredCloses part 2 of #1056.
Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com