Skip to content

Fix plugin-state leaks, callback-table lifetime, and unload gating in the nxmem plugin ABI - #1448

Closed
justinchuby wants to merge 1 commit into
justinchuby-phase-6-plugin-memory-abifrom
justinchuby-memory-phase-6-revision
Closed

justinchuby wants to merge 1 commit into
justinchuby-phase-6-plugin-memory-abifrom
justinchuby-memory-phase-6-revision

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Part of #1186 — Phase 6 only.

Supersedes #1440, which an independent review rejected. #1440 stays open as the record of that attempt; nothing in it was closed, modified or pushed to. This branch starts from its head (03faeeba) rather than from scratch — the review found the crate split, the vtable/version/status design, the C header and the read_prefix clamp sound, and none of those are churned here.

Base is justinchuby-phase-5-process-memory-manager. c65cf036 (Phase 5 tip) is already merged in below.

The two lifetime defects were the same mistake at two scales

Both findings come down to a debt recorded in a comment instead of in a type.

open_allocator documented its own obligation — once the plugin returns Ok it has created state and the host owes it a release, whatever the host then decides about the vtable — and then dropped that obligation on every rejection path that was added after the comment was written. AllocatorCore pinned the plugin's code through an Arc<PluginModule> while freeing the callback context the plugin still held a raw pointer into, because the SAFETY comment reasoned about the synchronous release and the queued case was never in view.

The fix in both places is to make the obligation something you cannot forget rather than something you must remember.

Finding 1 — plugin state stranded on post-Ok rejection

The review listed five unreleased paths. There are six: the missing-required-capabilities check between the mechanism-id check and the first read_capability was not in the list.

That is itself the argument for the guard over per-path calls. A reviewer reading the function carefully still missed one, and a seventh path costs nothing to add and nothing to notice. AbandonOnDrop is armed the instant the plugin returns Ok and disarmed on the last line before the allocator is published, so a path added later is covered before it is written.

A second defect, not in the review: the abandon path could never have released anything. abandon_allocator re-read the vtable through NxmemAllocatorVtable::read_prefix, which calls validate_required internally. The only route into abandon_allocator was a read_prefix failure. The re-read is deterministic over the same bytes, so it failed identically and returned before reaching release. Every post-Ok path leaked, including the two the review credited as correct.

This also corrects the review's note that validate_required "already catches and releases" the mechanism_id == 0 case. It caught it. It did not release it.

read_prefix is now a thin wrapper over read_prefix_unvalidated, which performs every memory-safety check — null, alignment, self-consistency, bounded copy, level clamp — and skips only the required-slot validation. Its doc comment says plainly that it exists for one caller and that release is the only slot safe to invoke on its result.

Finding 2 — callback table freed while queued releases were outstanding

The pinning was at the wrong granularity, exactly as the review said. bridge and callbacks are now one HostCallbackContext behind a single box, and HostBridge carries an outstanding_releases counter that enqueue_release bumps before the ABI call and unwinds on failure.

AllocatorCore::drop now drains, frees, releases, and only then decides the context's fate. If anything is still outstanding it increments LEAKED_CALLBACK_TABLES and mem::forgets the context. Freeing it would be a use-after-free the moment a plugin worker thread reports; leaking it is bounded by exactly the condition that already keeps the module mapped.

The related half the review flagged — drop not draining at all, leaving queued_releases non-zero forever — is fixed by retire_queued_releases_on_drop. It is bounded at 16 passes, not a loop to completion, and stops early when a pass retires nothing. A mechanism is entitled to refuse to drain (sticky does; a real stream-ordered mechanism may simply not be ready) and spinning would hang the caller. When the plugin declines, the count stays non-zero and the table is leaked — the honest outcome, not a hang.

Finding 4 — dropping a MemoryPlugin bypassed the gate

MemoryPlugin now has a Drop that re-runs both halves of the gate and, when it is shut, increments FORCED_MODULE_LEAKS and mem::forgets an Arc clone so dlclose never runs.

The review's point about platform luck is the right one and the docs now say so: glibc unmaps a refcount-zero DSO without DF_1_NODELETE and Rust cdylibs are not marked; macOS merely declining to unmap is not a safety property.

try_unload could no longer destructure self once a Drop existed (E0509). It now clears the factories, clones the Arc, drops self, and unwraps. Drop re-runs the gate on the clean path too, which is one extra plugin call that passes trivially and records no leak — pinned by dropping_a_plugin_with_live_objects_keeps_the_module_mapped.

The test holes shared one root cause

A fixture that merely lies about its size cannot exhibit an out-of-bounds read, because the bytes past the declaration are still valid memory. That is why the short-struct test passed for the wrong reason (Finding 3) and why the bounded read had no coverage at all (Finding 5) — the same fixture served both.

The fixtures are now backed by allocations that really end where they say they do (leak_vtable_bytes, poisoned_buffer), which is what lets a sanitiser see the over-read. Assertions are on NxmemStatusCode and on declared/required sizes, not on substrings of human-readable messages; PluginError::status_code() was added for that.

New poisoned-tail mechanism: an allocation of exactly size_of::<NxmemAllocatorVtable>() filled with 0xAB, with only the minor-0 prefix written over it and release_allocation deliberately populated inside the poisoned region. A host that reads past the declaration or skips the clamp ends up holding it. Asserted through a new publishes_structured_release_slot() accessor, because structured_release_slot() short-circuits on abi_minor < 1 and would hide the difference.

Offsets are pinned as well as sizes. MIN_STRUCT_SIZE_MINOR_0 is derived from offset_of!(Self, pending_release_count), so a field inserted mid-struct keeps the total size plausible while silently moving the boundary every older peer reads up to.

Finding 6 — layout contract on Windows

#[cfg(unix)] dropped from nxmem_c_example_compiles; #[cfg(all(unix, target_pointer_width = "64"))] reduced to #[cfg(target_pointer_width = "64")]. find_cc() now tries cl (probed with /?, since it has no --version) and returns a compiler description so args are built MSVC-style. The loud-failure property is kept: a missing compiler is a panic!, never a skip.

The header carries 49 new NXMEM_LAYOUT_FIELD: offset annotations alongside its 23 sizes, checked two independent ways — against Rust directly via offset_of! (no C compiler needed), and against a C compiler via _Static_assert(offsetof(...)). The rust_offset table is spelled out by hand rather than generated, because a table derived from the definition would agree with it by construction and prove nothing.

I could not verify the MSVC path. This host is macOS/arm64. The cl flags and the /std:c11 requirement for _Static_assert are from documentation, not from a run. Saying otherwise would be the same dishonesty this PR exists to fix.

M3 — plugin-report gate

Covered by a new self-retaining mechanism, which takes a reference to its own allocator state the host never learns about. Every ordinary scenario leaves the host holding something too, so the host's check fires first and the plugin's is never the reason for refusal; a host-zero / plugin-non-zero arrangement is the only way to make that gate load-bearing.

live_views — wired, not removed

Wired. A shared prefix committed into a live allocation is a genuine plugin-side object with a lifetime bounded by the allocation it looks into, and it is one the host cannot count for itself. Block gained a views counter; plugin_commit_shared_prefix increments both.

Views retire at block removal, not at free_block — the deferred-release path removes the block long before the bytes are freed, so retiring on free would leave the axis non-zero across the whole deferred window.

Removing the axis instead would have been an ABI and header change, and would have cost the gate an axis that definitely has a referent once an EP is wired in.

Test binary split

Three gate tests each live in their own binary. Module counters live in the loaded .so and are process-global; sticky never retires and self-retaining never releases, so each permanently poisons them. Two such tests in one binary read each other's residue and force exact assertions to be softened into inequalities — which is precisely the dishonesty that got #1440 rejected. The pre-existing queued_releases >= 1 || ... OR-assertion in the unload-gate test was tightened into two exact assert_eq!s for the same reason.

Mutation evidence

Every fix was mutation-tested: break the production check, confirm red, restore. All mutations are removed from the tree (grep for the markers returns nothing).

# Mutation Result
M1 copy_prefix: if false && declared_size < required RED — 3 unit + a_short_allocator_vtable_is_refused_before_any_slot_is_read. Previously the integration test passed on the status code's name.
M3 try_unload: if false on the report.total() != 0 gate RED — unload_is_refused_when_only_the_plugin_still_owns_something. Was all-green before.
M7 copy_prefix: let readable = size_of::<T>() RED — copy_prefix_reads_exactly_the_declared_prefix_and_zeroes_the_rest. Under Miri: Undefined Behavior: attempting to access 128 bytes, but got alloc357122 which is only 120 bytes from the end of the allocation.
M9 M7 + the effective_minor < 1 nulling removed RED — 4 unit tests and a_poisoned_tail_never_reaches_the_hosts_view_of_the_vtable. Previously all 24 integration tests stayed green.
F1a AbandonOnDrop::drop → no-op RED — 8 tests, including the cascade the review predicted: one stranded allocator disables the gate for every later test in the process.
F1b abandon_allocator back to the validating read_prefix RED — 9 tests. This is the evidence for the dead-code defect.
F2a free the context unconditionally in AllocatorCore::drop RED — dropping_an_allocator_with_a_queued_release_leaks_its_callback_table.
F2b remove retire_queued_releases_on_drop() RED — 5 tests.
F4 Drop for MemoryPlugin → no-op RED — dropping_a_plugin_with_live_objects_keeps_the_module_mapped.
F6 corrupt one header offset annotation RED — both the_header_field_offsets_match_rust and a_c_compiler_agrees_with_the_rust_layouts.

One honest caveat on M7. It is caught deterministically only at unit level, plus non-deterministically under Miri. It cannot be made value-observable at integration level, and this is structural rather than a gap I chose to leave: read_prefix_unvalidated rejects any vtable where declared_size < required_struct_size(declared_minor), and required_struct_size(1) == size_of::<Self>(), so no vtable can be both short and at effective minor ≥ 1. release_allocation is the only field past the minor-0 prefix, and the clamp nulls it independently of the bound. Destroying the bound alone therefore changes no observable value — it is purely an out-of-bounds read, and the only honest way to see it is a sanitiser on an exactly-sized allocation. That is what the Miri result above is.

The exact-size fixture also had to be Vec<u64>-backed rather than Vec<u8>. A Vec<u8> is byte-aligned; the system allocator happens to return 8-aligned blocks but Miri does not, so a byte-backed fixture failed under Miri for the wrong reason — the alignment check, not the bound. Miri caught that too.

Scope: criterion 8 is half done

Deferred release now keeps the plugin module and host callback table pinned, and that half is tested.

The provider/context half is not implemented. No execution provider is wired to this ABI in this phase — there is no onnx-runtime-ep-cuda integration and deferred_release.rs is untouched — so there is nothing to pin. That scope decision was accepted by review; it is stated here and in docs/memory/MEMORY_PLUGIN_ABI.md so it cannot be mistaken for done, and it is tracked for Phase 7.

Validation

macOS/arm64, no CUDA. Nothing in these three crates touches CUDA.

  • cargo test -p onnx-runtime-memory-abi -p onnx-runtime-memory-host -p onnx-runtime-memory-testplugin — 101 passed, 0 failed (abi lib 50, header_contract 4, host lib 9, nxmem_abi 29, callback_pinning 1, plugin_gate 1, unload_gate 1, testplugin lib 6). Baseline at 03faeeba was 86.
  • cargo clippy --all-targets -- -D warnings scoped to the three crates — clean.
  • cargo fmt --check — clean for everything this PR touches.
  • cargo doc --no-deps for the three crates — zero warnings.
  • Miri: cargo +nightly miri test -p onnx-runtime-memory-abi --lib — 50 passed, 0 failed, and it flags M7 as UB. Not run on the host integration tests: they dlopen a real cdylib, which Miri cannot execute.
  • ASan: not run. Miri gave a stronger result on the code that matters.
  • Standalone C, no workspace linking: minimal_plugin.c and a 72-assertion layout probe (23 sizes + 49 offsets) both compile with cc -std=c11 -Wall -Wextra -Werror against the copied header alone.

git diff 03faeeba --stat is 12 files, all inside the three crates plus docs/memory/MEMORY_PLUGIN_ABI.md. No onnx-runtime-memory-governor dependency was added, so the Phase 1–5 lock graph is untouched by construction. No host lock is held across a plugin call. The three accounting axes stay distinct and no quarantine path refunds.

Known pre-existing, not from this PR and not fixed here: completions.rs rustfmt drift (verified byte-identical to base — git diff 03faeeba on that file is empty), matmul_nbits_marlin_numerics, Windows CUDA unnecessary_unwrap, macOS MLAS, governor rustdoc intra-doc links.

CI will sit pending for hours on this repo's backlog. Local validation above is the evidence.

Do not merge, do not enable auto-merge, do not enqueue.

… nxmem

Revision of the Phase 6 plugin memory ABI after review rejection. The
architecture and the ABI design are unchanged; this closes the specific
safety defects and the test holes that let four mutations pass unnoticed.

The two lifetime defects were the same mistake at two scales. `open_allocator`
took on a debt when the plugin returned `Ok` and then dropped it on six
rejection paths, and `AllocatorCore` pinned the plugin's *code* while freeing
the callback context the plugin still held a pointer into. Both are now
carried by the type system rather than by remembering: an `AbandonOnDrop`
guard covers rejection paths that do not exist yet, and the bridge and
callback table live in one boxed unit whose teardown is gated on the
outstanding-release count.

The abandon path could not have worked as written. It re-read the vtable
through `read_prefix`, which validates, and the only way to reach it was a
`read_prefix` failure — so the re-read failed identically and returned before
calling `release`. Splitting out `read_prefix_unvalidated` is what makes the
release reachable at all; the guard alone would have fixed nothing.

`MemoryPlugin` now has a `Drop`. Without one, every early return and unwind
skipped the unload gate silently and left `dlclose` free to unmap a module
with live objects in it. Whether that actually unmaps is a property of the
loader, not of this code, so the drop keeps the module mapped and counts it.

The test holes shared a root cause: a fixture that lies about its size cannot
exhibit an out-of-bounds read, because the bytes past the declaration are
still valid. The short-struct and poisoned-tail fixtures are now backed by
allocations that really end where they say they do, which is what lets Miri
see the over-read. Assertions are on status codes and sizes rather than on
substrings of human-readable messages.

`live_views` is wired rather than removed: a shared prefix committed into a
live allocation is a genuine plugin-side object, and views retire when the
block is removed rather than when its bytes are freed.

The header now pins 49 field offsets as well as 23 sizes, and the layout
tests no longer skip on Windows — MSVC is the toolchain most likely to
disagree about `#[repr(C)]` packing, which is the whole point of the test.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 19, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 80.00000% with 816 lines in your changes missing coverage. Please review.
✅ Project coverage is 79.34%. Comparing base (c65cf03) to head (fb9a72c).

Files with missing lines Patch % Lines
crates/onnx-runtime-memory-host/src/allocator.rs 68.73% 337 Missing and 11 partials ⚠️
crates/onnx-runtime-memory-testplugin/src/lib.rs 74.49% 268 Missing and 20 partials ⚠️
crates/onnx-runtime-memory-host/src/loader.rs 82.50% 59 Missing and 4 partials ⚠️
crates/onnx-runtime-memory-abi/src/vtable.rs 91.50% 49 Missing and 3 partials ⚠️
crates/onnx-runtime-memory-abi/src/version.rs 90.73% 25 Missing and 4 partials ⚠️
crates/onnx-runtime-memory-abi/src/status.rs 87.34% 19 Missing and 1 partial ⚠️
crates/onnx-runtime-memory-abi/src/types.rs 96.92% 9 Missing ⚠️
crates/onnx-runtime-memory-abi/src/lib.rs 85.71% 4 Missing and 2 partials ⚠️
crates/onnx-runtime-memory-host/src/error.rs 98.33% 1 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@                              Coverage Diff                               @@
##           justinchuby-phase-5-process-memory-manager    #1448      +/-   ##
==============================================================================
- Coverage                                       80.01%   79.34%   -0.68%     
==============================================================================
  Files                                             366      373       +7     
  Lines                                          164950   165793     +843     
  Branches                                       164950   165793     +843     
==============================================================================
- Hits                                           131980   131543     -437     
- Misses                                          28150    29407    +1257     
- Partials                                         4820     4843      +23     
Flag Coverage Δ
mlas ?
offline 79.34% <80.00%> (-0.58%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-runtime-memory-host/src/error.rs 98.33% <98.33%> (ø)
crates/onnx-runtime-memory-abi/src/lib.rs 85.71% <85.71%> (ø)
crates/onnx-runtime-memory-abi/src/types.rs 96.92% <96.92%> (ø)
crates/onnx-runtime-memory-abi/src/status.rs 87.34% <87.34%> (ø)
crates/onnx-runtime-memory-abi/src/version.rs 90.73% <90.73%> (ø)
crates/onnx-runtime-memory-abi/src/vtable.rs 91.50% <91.50%> (ø)
crates/onnx-runtime-memory-host/src/loader.rs 82.50% <82.50%> (ø)
crates/onnx-runtime-memory-testplugin/src/lib.rs 74.49% <74.49%> (ø)
crates/onnx-runtime-memory-host/src/allocator.rs 68.73% <68.73%> (ø)

... and 17 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/medium_generic_f16_threads=8/32x512x512 30.08 µs 188.03 µs +525.1%
🔴 matmul/medium_generic_bf16_threads=8/32x512x512 417.14 µs 1.17 ms +179.4%
🔴 matmul/small_generic_bf16_threads=8/1x256x256 33.43 µs 90.79 µs +171.6%
🔴 matmul/medium_generic_f32_threads=8/32x512x512 1.36 ms 2.09 ms +54.5%
🔴 gather/medium_bf16_threads=1-internal/32768 2.68 µs 4.07 µs +51.8%
🔴 matmul/medium_generic_f16_threads=1/32x512x512 30.46 µs 43.04 µs +41.3%
🔴 gather/medium_f32_threads=1-internal/32768 3.95 µs 5.34 µs +35.3%
🔴 reduce_mean/large_f32_threads=1-internal/262144 1.01 ms 1.31 ms +30.4%
⚠️ matmul/medium_generic_bf16_threads=1/32x512x512 607.42 µs 786.10 µs +29.4%
⚠️ gather/large_bf16_threads=1-internal/131072 12.76 µs 16.25 µs +27.4%
⚠️ gather/small_f16_threads=1-internal/4096 494.9 ns 606.1 ns +22.5%
⚠️ reduce_mean/small_f32_threads=1-internal/4096 16.52 µs 20.08 µs +21.5%
⚠️ add/small_bf16_threads=1-internal/1024 454.8 ns 547.7 ns +20.4%
⚠️ matmul/large_generic_f32_threads=8/32x1024x1024 4.04 ms 4.83 ms +19.4%
⚠️ add/large_bf16_threads=1-internal/4194304 1.75 ms 2.07 ms +18.1%
⚠️ reduce_mean/medium_f32_threads=1-internal/65536 291.70 µs 340.29 µs +16.7%
⚠️ matmul/large_generic_f16_threads=8/32x1024x1024 126.65 µs 146.55 µs +15.7%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.56 ms 2.93 ms +14.3%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 10.07 ms 11.50 ms +14.2%
✅ gather/large_f16_threads=1-internal/131072 15.74 µs 17.89 µs +13.7%
✅ qwen3_sampling_processors/top_k_top_p_fast 656.05 µs 720.22 µs +9.8%
✅ gather/medium_f16_threads=1-internal/32768 2.64 µs 2.90 µs +9.8%
✅ gather/small_bf16_threads=1-internal/4096 469.7 ns 509.6 ns +8.5%
✅ grammar_masking/llguidance_compute_mask/32 78.22 µs 83.22 µs +6.4%
✅ gather/small_f32_threads=1-internal/4096 670.5 ns 703.5 ns +4.9%
✅ add/small_f16_threads=1-internal/1024 475.0 ns 490.5 ns +3.3%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 540.30 µs 556.72 µs +3.0%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.86 ms 3.92 ms +1.5%
✅ matmul/small_generic_bf16_threads=1/1x256x256 35.99 µs 36.53 µs +1.5%
✅ logit_processing/seven_processor_chain_per_step 321.04 µs 324.65 µs +1.1%
✅ add/medium_f32_threads=1-internal/262144 26.72 µs 26.94 µs +0.8%
✅ gather/large_f32_threads=1-internal/131072 46.28 µs 46.52 µs +0.5%
✅ add/large_f16_threads=1-internal/4194304 1.82 ms 1.82 ms -0.2%
✅ kv_cache/alloc_dealloc_pages 39.93 µs 39.73 µs -0.5%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 92.47 µs 91.80 µs -0.7%
✅ add/medium_bf16_threads=1-internal/262144 102.46 µs 98.46 µs -3.9%
✅ sampling_latency/top_p_per_token 391.58 µs 376.01 µs -4.0%
✅ add/large_f32_threads=1-internal/4194304 702.06 µs 673.01 µs -4.1%
✅ add/small_f32_threads=1-internal/1024 226.1 ns 214.3 ns -5.2%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 92.24 µs 84.93 µs -7.9%
✅ matmul/small_generic_f32_threads=1/1x256x256 46.55 µs 42.66 µs -8.4%
✅ sampling_latency/greedy_per_token 3.30 µs 2.98 µs -9.8%
✅ add/medium_f16_threads=1-internal/262144 113.76 µs 101.84 µs -10.5%
✅ matmul/small_generic_f32_threads=8/1x256x256 61.47 µs 53.90 µs -12.3%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.90 ms 2.51 ms -13.6%
✅ qwen3_sampling_processors/top_k_partial_selection 171.54 µs 148.00 µs -13.7%
✅ tokenization/encode_tokens_per_second 438.07 µs 377.74 µs -13.8%
✅ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 826.04 µs 706.99 µs -14.4%
✅ sampling_latency/min_p_per_token 237.75 µs 202.53 µs -14.8%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 545.97 µs 464.63 µs -14.9%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 6.44 ms 5.48 ms -14.9%
🟢 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 73.60 µs 60.62 µs -17.6%
🟢 qwen3_sampling_processors/top_k_full_sort_baseline 2.48 ms 2.04 ms -17.7%
🟢 matmul/small_generic_f16_threads=8/1x256x256 48.16 µs 36.31 µs -24.6%
🟢 tokenization/decode_tokens_per_second 8.65 ms 6.02 ms -30.4%
🟢 sampling_latency/top_k_per_token 68.34 µs 46.10 µs -32.5%
🟢 matmul/small_generic_f16_threads=1/1x256x256 52.25 µs 35.15 µs -32.7%
🟢 matmul/large_generic_bf16_threads=8/32x1024x1024 2.67 ms 1.54 ms -42.3%
🟢 block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 97.46 µs 50.19 µs -48.5%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.77 4.41 6.33 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby
justinchuby changed the base branch from justinchuby-phase-5-process-memory-manager to justinchuby-phase-6-plugin-memory-abi August 20, 2026 14:53
justinchuby added a commit that referenced this pull request Aug 22, 2026
> **A100 safety update (head `13c37a7f`):** expected shared-prefix
admission failures are now all-or-private: K/V commit is transactional,
a failed V share rolls K back before enqueue, and only a successful
rollback permits private-KV fallback. Fatal kernel faults remain
process-fatal; the new cross-process VMM probe proves a worker fault
does not affect an actively executing peer, while the current server is
still in-process and does not yet consume physical shared-prefix
metadata. See the [final verified
conclusion](#1579 (comment)).

Collapses phases 1–7 of the #1186 memory architecture rework onto
current `main` as a single merge. The stack forked 270 commits ago (176
of them touching `crates/`), so rebasing it layer by layer would mean
solving the same 17 conflicts fourteen times, with no reviewer ever
looking at the intermediate states.

Closes the stack: #1252 #1263 #1279 #1283 #1301 #1341 #1349 #1426 #1440
#1448 #1454 #1462 #1465 #1468 #1533.

**A review guide is in the first comment.** It is the part worth reading
— this diff is 103 files, but only three decisions in it are ones a
compiler cannot check.

## What was already reviewed, and what wasn't

Every phase was reviewed and approved by a session that was not its
author, under the rejection-lockout rule (a rejected author never writes
the next revision).

| Phase | PR | Verdict |
|---|---|---|
| 1–5 | #1252 → #1426 | approved |
| 6 | #1440, #1448, #1454 | **rejected** ×3 |
| 6 | #1462 | approved |
| 7 | #1465 | **rejected** |
| 7 | #1468 | approved |
| 7 test fixes | #1533 | approved, A100-verified |

The four rejected rounds were not replaced — the later rounds are
stacked **on top of** them. So the tree here is the approved state, but
the history contains the rejected commits. **#1440 / #1448 / #1454 /
#1465 should not be reviewed individually**; they close automatically.

**Not reviewed anywhere:** the conflict resolutions themselves. That is
what this PR is for.

## Verification

Run here, on this merged tree:

- `cargo check --workspace --all-targets` — clean
- `cargo check -p onnx-genai-engine --features cuda,native-backend
--all-targets` — clean. Worth calling out: the default feature set does
**not** compile the `cfg(cuda)` code, which is where the riskiest edits
are. Checking only the default set would have missed a real break (see
the guide).
- `cargo test` over 7 crates — **1105 passed, 2 failed, 88 ignored**
- `cargo clippy --workspace --all-targets`
- `cargo fmt`

**The 2 failures and the 1 clippy error are pre-existing and proven so,
not merge damage:**

-
`platform_capacity::{disk_capacity_is_measured_for_the_working_directory,
an_explicit_byte_limit_is_honored_without_a_device_query}` —
`platform_capacity.rs` is byte-identical to `main` (`git diff
origin/main HEAD -- ` that file is empty). The cause is a macOS-only FFI
layout bug: `fsblkcnt_t` is 4 bytes on macOS (confirmed: `sizeof(struct
statvfs)=64`, `sizeof(fsblkcnt_t)=4`) while the Rust struct declares
`f_blocks`/`f_bfree`/`f_bavail` as `u64`. CI is Linux, where it is 8
bytes. Left alone: it is a real bug but not this PR's.
- `optimizer.rs` clippy `approx_constant` — present on both parents; the
merged file is byte-identical to `main`.
- Three files had `rustfmt` drift already present on `main`
(`matmul_nbits.rs`, `normalization.rs`, `optimizer.rs`). Reverted rather
than swept in, to keep the diff readable.

**Final A100 revalidation (head `13c37a7f`):** full CUDA-memory GPU
suite passed; CUDA EP default-parallel lib suite passed (**488 passed /
17 ignored**); real fp16 GQA covered three requests × two interleaved
decode steps with two shared peers and one transactional private
fallback, byte-identical to independent GPU and CPU references. A
device-started/event-gated worker remained healthy through a peer
process `CUDA_ERROR_ILLEGAL_ADDRESS`, owner exit, replacement worker,
and process restart. Governor tests, targeted Miri, CI Clippy, fmt, and
CUDA honesty also passed.

## Still open after this merges

- `memory-plugin-provider-wiring` — the back half of Phase 6 criterion 8
(provider/context pinning), which fell in the gap between phases.
**#1186 must not be closed until it lands.** This merge makes it more
tractable, not less: see decision 3 in the guide.
- `memory-deferred-invariant-asserts` — two surviving mutants found by
#1462's reviewer, adjudicated NON-BLOCKING and not defects.
- #1533's CUDA honesty guard warns on a legitimately portable anchor.
Should be moved out of the CUDA test binary or allow-listed rather than
left warning.

---------

Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <copilot@github.com>
Copilot-Session: 46c5d75b-8146-489c-b82f-08ee29c27ce4
Copilot-Session: 39ff6824-d35f-4d3b-8f5f-043a7119a100
Copilot-Session: c80f8522-983c-47f7-8241-2155a823aabe
@justinchuby

Copy link
Copy Markdown
Owner Author

Closing: superseded by #1579, which collapsed phases 1–7 into a single merge — landed as a36964280 (squash, 126 files, +43956 −3286).

This PR was rejected under the lockout rule and superseded by the approved revision that was stacked on top of it (not by a replacement of it). Per #1579's own guidance, it was never meant to be reviewed individually.

Containment verified, not assumed:

$ git merge-base --is-ancestor fb9a72c4d ebceab071   # ebceab071 = #1579 head
→ ancestor

Every one of the 15 stack heads (#1252 → #1533) is a literal ancestor of the merged head, and I separately confirmed the merged head's content reached main: of the 126 files #1579 touched, exactly one differs from main — crates/onnx-runtime-session/src/executor/tests.rs, where main carries 120 extra lines from #1703 (Expand-broadcast mask tests). That is the conflict resolution correctly preserving the other side, not content loss.

Why this didn't close itself: #1579's body says Closes the stack: #1252 #1263 …, which GitHub does not parse — the closing keyword must be immediately followed by the reference, and in any case closing keywords only ever close issues, never pull requests. So all 15 stayed open by mechanism, not by intent.

#1186 stays open: #1579 explicitly gates it on memory-plugin-provider-wiring (the back half of Phase 6 criterion 8, provider/context pinning), which has not landed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant