Repository navigation
Evict MLAS SQNBit packs by owner, and key them on every operand (#1735) - #1738
Conversation
`MlasPackedCaches` keys the shared SQNBit packs on the operands' addresses and holds no claim on the memory there. An address names a weight only while that weight is live: once freed, the next allocation of matching shape and pack parameters is served the previous weight's packed bytes. The route counter still increments, because the right route did run -- it just multiplied another weight's rows. This is #1726's defect in the packed store rather than the transpose caches, and `clear_mlas_packed_caches()` at executor teardown only closed the cross-model case. Two gaps, closed together: Lifetime. `get_or_build_shards`/`get_or_build_packed` now report whether *this* call installed the entry, and only the installing kernel records ownership. A retiring kernel evicts what it installed, via a small `MlasPackedOwnership` field that owns the `Drop`. Ownership is install-only on purpose: a kernel that merely hit an entry must leave it alone, or a retiring prefill instance would pull the pack out from under a live decode sibling and force a re-pack -- the cost the shared store (#1056) exists to remove. Identity. A CompInt8 pack bakes the scales and zero points into its per-block sums, but the key covered only the quantized weight, so two nodes sharing a `B` initializer under different scales would serve each other packs. The key now carries all three operand addresses, and a weight whose scales have no stable host address is treated as unshareable rather than shared under a partial identity. Also fixes `clear_mlas_packed_caches()`, which drained `MLAS_PACKED_GLOBAL` directly while `cfg(test)` traffic goes to the thread-local store -- a silent no-op in every test that called it. `MlasPackedOwnership` exists so the `Drop` does not land on `MatMulNBitsKernel` itself: a type that implements `Drop` cannot be moved out of, which would make the `..test_kernel(..)` functional-update syntax used throughout the suite illegal. Tests, each verified to fail under its own mutation and no other: - a recycled address must not serve the previous weight's pack (8 alternating rounds, cross-weight reuse required, references taken against a drained store); - a shared pack outlives the sibling that only shared it; - one quantized blob under two live scale tensors must not share a pack; - the store stays consistent under concurrent install/evict; - weight-derived caches are cleared before their buffers are freed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1738 +/- ##
==========================================
+ Coverage 79.75% 80.06% +0.30%
==========================================
Files 408 408
Lines 193197 192932 -265
Branches 193197 192932 -265
==========================================
+ Hits 154088 154468 +380
+ Misses 33759 33109 -650
- Partials 5350 5355 +5
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
Opus review found the recycled-address test weaker than it presented. Two fixes, both to the test: The vacuousness guard tracked only the packed weight's address, but with the extended key a stale hit needs the packed, scales and zero-point addresses to all collide. On an allocator whose size classes do not recycle in lockstep the run could report `recycled=true` while the key never matched -- passing even with per-owner eviction reverted. It now tracks the whole key tuple. More importantly the test no longer depends on recycling at all: each round now asserts the store is empty before it starts, which is what the previous round's retiring kernel guarantees. That holds on every allocator, and it kills a no-op `Drop` deterministically at round 1 rather than waiting for an address to be reused. The final recycling check is consequently no longer an assertion. Address reuse is the allocator's choice, not the code's: under ASan, valgrind, or some musl/macOS configurations no reuse occurs and a correct implementation would have failed. It is reported loudly instead, so a run that skipped that arm can never be read as one that exercised it. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Independent Opus review — dispositionReviewed adversarially against six attack surfaces. Five came back unfounded after the reviewer traced the concrete interleavings; both actionable findings were test-robustness issues and are fixed in Unfounded, with the reasoning worth recording:
Fixed:
Noted, not changed: the env-var save/restore in the two new tests is not panic-safe — an assertion failing between Final validation
Includes Plus: |
|
Merge recordMerging under Justin's standing direct-merge directive (2026-08-19T14:20:27: comprehensive local validation including cross-platform, then merge directly rather than waiting on queued Actions). Local matrix: Required-CI position, stated honestly. One check is red: The remaining checks are queued, not failing. Per the directive I am not waiting on them; if any produces a real code failure afterwards I will fix it on main rather than leave it. |
… it died (#1889) Closes #1745. ## What #1745 actually reports ``` SPMD parity child failed (persistent=true, workers=3, status: exit code: 0xc0000005 (code -1073741819, 0xc0000005)) ``` `0xC0000005` is `STATUS_ACCESS_VIOLATION`. The child **faulted** — it did not fail the parity assertion. `child_status_detail` exists to distinguish those and says fault. Roy did the elimination work on his PR (#1741): the diff didn't touch the test, the converted env vars are inert on that lane, and an identical tree with one empty commit added passed. He also checked that `Rust (Windows ARM64)` is **green on main** — it is not in the #1600 chronic-red set, so this is a real crash rather than a known-red lane. ## The asymmetry this fixes Three test children build the process-wide decode pool. Exactly one of them, `affinity_defer_routing_child`, stops it before exiting — and its comment names *this same* `0xC0000005` on native Windows ARM64 as the reason it was added. The parity child and the realized-width child never got the same treatment, and #1745 is reported against exactly those two. The pool is a module-level `static`, which Rust never `Drop`s. So those two children exit with workers still spinning or parked on `Arc<SharedState>`, and on Windows `ExitProcess` terminates them at an arbitrary instruction — including inside a CRT or heap lock that exit-time cleanup then takes. Applying the existing remedy to the other two is the cheapest way to find out whether it is the same bug. ## What I am *not* claiming I am not claiming this is proven to be the cause. The crash is rare, intermittent, and on a platform this workspace cannot execute; I have no reproduction and no crash dump. This is a hypothesis with in-repo precedent. Which is why the other half of the change exists. ## Saying where it died A #1745 report today says the child faulted and nothing else. Between process start and its single result line the child builds a pool, runs a kernel, encodes bytes and tears the pool down — a fault anywhere in that span produces the same report. "Crashed doing the work" and "crashed at exit after the work" are different bugs with different fixes, and nothing in the log separates them. Breadcrumbs (`NXRT_CHILD_STAGE: site=… stage=…`) fill that in. If the next occurrence shows `body-complete` with no `pools-stopped`, the crash is *inside* teardown and this PR's remedy is where to look. If it shows neither, it is in the work and this PR is a red herring. Either way the next report is worth more than this one. ### One thing I checked rather than assumed The natural suspicion is that the detail was printed and lost — the child writes to a pipe and dies. **That is not what happens.** Rust's stdout is a `LineWriter` on every platform, so a newline-terminated `println!` has already reached the pipe before the fault. Verified rather than reasoned about: a child that writes one newline-terminated line, one unterminated `print!` and one flushed stderr line, then `abort()`s under `Command::output()`, delivers the first and the third and loses only the unterminated fragment. So the gap is not lost output — it is that nothing is emitted between "started" and "produced the result". I mention it because I believed the buffering story first and wrote it into a doc comment before testing it. It was wrong. ### Why stderr, and why the order matters Both readers of these children key on stdout **by content**: `is_environmental_access_violation_crash` treats the result marker's presence as "the child produced its result", and `parity_child_output_mode` scans stdout for the parity payload. Extra stdout lines are a change to two lanes' red/green semantics. Extra stderr lines are not — provided they never contain `panicked at` or `assertion`, the only two strings that classifier reads out of stderr. That is asserted, not assumed. The teardown call is placed **after** the result marker in the two children that never had one, so a crash inside our own teardown keeps the marker on stdout, classifies as non-environmental, and is **reported rather than retried away**. A bug in this code is not a flaky runner. `affinity_defer_routing_child` keeps its existing marker-last order. Flipping it would change the red/green semantics of a currently-green lane, on a platform I cannot execute, to probe a hypothesis that is not yet confirmed. Its breadcrumb answers the same question without touching the verdict. ## `workers_exited`, and a test that flaked `ready` makes "the workers have started" an observable fact that `build` blocks on. Nothing made "the workers have stopped" observable, so the only available evidence of a completed teardown was the absence of thread names in `/proc/self/task`. My first version of the teardown test asserted exactly that — and **it flaked under load**. The futex that unblocks `join` is signalled at `mm_release`, before the kernel unhashes the task, so a joined thread's `/proc` entry can still exist when the next statement runs. That is a race in the *instrument*, and an intermittently-red teardown test would teach precisely the wrong lesson about intermittently-red teardown. So the assertion moved to a counter incremented by the worker before it returns, which is ordered by the join that observes it. The thread counts are still printed — as a report, not an assertion. Same demotion Roy applied to the recycled-address falsifier in #1738. Cost: one `fetch_add` per worker per process, on a `Drop` guard so a panicking worker is counted out too. ## Tests — 3 added, 6 mutations, 6 caught uniquely | injected defect | caught by | over-catch | |---|---|---| | breadcrumb text contains `assertion` | `…invisible_to_the_crash_classifier` | none | | breadcrumb loses its grep prefix | `…invisible_to_the_crash_classifier` | none | | breadcrumb drops `site=` | `…names_both_the_site_and_the_stage` | none | | breadcrumb drops `stage=` | `…names_both_the_site_and_the_stage` | none | | `shutdown` stops joining | `shutdown_pools_is_a_barrier_not_a_request` | none | | worker exit stops being counted | `shutdown_pools_is_a_barrier_not_a_request` | none | Baseline clean. The join falsifier is **probabilistic by nature** — the workers do leave, just not before the next statement — measured at **18/20** on this host, and stated as such in the test's own doc rather than left as an implied certainty. `shutdown_pools_is_a_barrier_not_a_request` runs in a child process because `comm` truncates at 15 bytes, so in a shared test binary another test's pool is indistinguishable from this one's — same reasoning as the existing `a_failed_build_leaves_no_workers_running`. It builds with `blocktime=0` so the workers are **parked, not spinning**: a spinning worker would drift out on the stop flag alone, which would let a shutdown that never woke anyone pass. ## Adjacent, not claimed Roy notes the crash lands at `workers=3` under the persistent pool, i.e. inside the only path that checks realized width, and that this is uncomfortably close to the t=2 dispatch anomaly I own. I have no evidence they are the same bug and am not asserting it. #983 and #1672 carry the same `0xC0000005` signature on other Windows lanes and may share a root cause; #1123 (LoggingManager race) is not yet ruled in or out. ## Validation - `cargo test -p onnx-runtime-ep-cpu --lib -- spmd` — 97 passed, 3 ignored. - `cargo test -p onnx-runtime-ep-cpu --lib -- affinity_defer realized_width parity` — 28 passed. - `cargo clippy -p onnx-runtime-ep-cpu --all-targets -- -D warnings`, default and `--features mlas` — clean. - `cargo fmt --all` — clean. Note `main` is currently red on `Fast (Linux x86_64)` (`the_matrix_holds_for_every_maintained_workflow`, from `7a0cb6c39`), which is unrelated to this change and will need to go green before this can merge through the gate. Waiting, not bypassing. Co-authored-by: Sebastian <seb@example.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Closes #1735.
The defect
MlasPackedCacheskeys the shared SQNBit packs on the operands' addresses and holds no claim on the memory there. An address names a weight only while that weight is live: once freed, the next allocation of matching shape and pack parameters is served the previous weight's packed bytes. The route counter still increments, because the right route did run — it just multiplied another weight's rows. This is #1726's defect in the packed store rather than the transpose caches, andclear_mlas_packed_caches()at executor teardown only closed the cross-model case.Two gaps, closed together
Lifetime.
get_or_build_shards/get_or_build_packednow report whether this call installed the entry, and only the installing kernel records ownership. A retiring kernel evicts what it installed.Ownership is install-only on purpose: a kernel that merely hit an entry must leave it alone, or a retiring prefill instance would pull the pack out from under a live decode sibling and force a re-pack — the cost the shared store (#1056) exists to remove. This is the defect Opus caught in #1732's first draft, so it is asserted directly.
Identity (not in the issue). A CompInt8 pack bakes the scales and zero points into its per-block sums, but the key covered only the quantized weight — two nodes sharing a
Binitializer under different scales would serve each other packs. The key now carries all three operand addresses, and a weight whose scales have no stable host address is treated as unshareable rather than shared under a partial identity. This is the "stable full weight identity where available" the directive asked for, and it also strengthens the recycled-address case, since a colliding address must now match on scales too.Bonus.
clear_mlas_packed_caches()drainedMLAS_PACKED_GLOBALdirectly whilecfg(test)traffic goes to the thread-local store — a silent no-op in every test that called it. Now routed throughwith_mlas_packed_caches.Design notes
MlasPackedOwnershipexists so theDropdoes not land onMatMulNBitsKernelitself: a type that implementsDropcannot be moved out of, which makes the..test_kernel(..)functional-update syntax used throughout the suite illegal (70 compile errors on the first attempt). Moving out of a struct whose fields haveDropis fine.Production teardown ordering is unchanged and now guarded:
clear_weight_transpose_caches()andclear_mlas_packed_caches()run inExecutor::dropbeforeself.buffers.drain().Tests
Every test was verified to fail under its own mutation and no other:
a_recycled_weight_address_must_not_serve_the_previous_weights_mlas_packDropmade a no-op → round 1 serves weight 0's output for weight 1a_shared_mlas_pack_outlives_the_sibling_that_only_shared_itone_quantized_blob_under_two_scale_tensors_must_not_share_a_packscales_addrdropped from the key → second scales tensor gets the first's packthe_packed_store_stays_consistent_under_concurrent_install_and_evictMlasPackedCachesdirectly; the per-test store is thread-local, so kernels on different threads cannot contend)weight_derived_caches_are_cleared_before_their_buffers_are_freedbuffers.drain()The falsifier runs 8 alternating rounds and requires a cross-weight address reuse before it will pass — "some address was reused" is satisfiable by a weight landing on its own former address, which proves nothing. Eight rounds are needed because glibc first serves the ~2 MiB pack by
mmap; freeing one raises the dynamic mmap threshold so later same-size requests come from the brk heap, where addresses are reused. Reference outputs are taken against a drained store, so a stale hit cannot poison the reference and make the comparison a tautology.Routing note: both kernel-level tests drive
try_mlas_sqnbitunderNXRT_CPU_GEMM_BACKEND=mlasat m=4. On x86_64 the native acc4 int4 route wins before MLAS is consulted, and MLAS refuses asymmetric CompInt8 at M=1 on hosts without a correct AVX2 kernel — either would have made the tests silently vacuous.Validation
./roy_validate.sh— PASS=21 FAIL=0 SKIP=0, includingH check-win-arm64(cargo-xwin) andG2 test-aarch64-qemu(the aarch64 suite actually executed, not just type-checked).mlas) build/test/clippy-D warnings,--no-default-features,--all-featuresK no-mlas-artifacts— the default zero-MLAS artifact is untouched; the research-only feature policy is preservedRUST_TEST_THREADS=32× 6 clean runs, 1621 lib tests each-Zmiri-tree-borrows): passes onthe_packed_store_stays_consistent_under_concurrent_install_and_evict. Miri cannot execute the other three — they call into MLAS through FFI. Coverage here is the safe store/ownership logic only, and that limit is stated rather than papered over.The default (no-
mlas) lane caught a real cfg regression during development: an inserted field consumed the#[cfg(feature = "mlas")]that belonged tomlas_shardsat both initializer sites.New weight-derived caches must be governedgate is expected to fire — it matches on diff text, so a modified field declaration reads as a new cache. No new cache is introduced;MlasPackedOwnershipholds keys, not buffers.