Skip to content

fix(memory): correct Phase 7's unverifiable claims and anchor count_code - #1468

Closed
justinchuby wants to merge 1 commit into
justinchuby-memory-phase-7-vmm-onlyfrom
justinchuby-memory-phase-7-revision
Closed

justinchuby wants to merge 1 commit into
justinchuby-memory-phase-7-vmm-onlyfrom
justinchuby-memory-phase-7-revision

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Part of #1186 — Phase 7 only. Base is justinchuby-sturdy-potato (Phase 6, head 21370921, approved).

Supersedes #1465, which stays open as the review record. Nothing was pushed to it.

Phase 7 removes the built-in CUDA eager allocator (CudaDeviceAllocator, cuMemAlloc) and makes the VMM arena the sole built-in CUDA memory mechanism. #1465's production change was reviewed and no production defect was found. It was rejected on three factual claims that stand in place of verification that cannot be performed on this host, plus one missing test anchor. This revision fixes exactly those. There are no production behaviour changes; the single production string touched is the diagnostic at provider.rs:408.

This host is macOS/arm64 with no CUDA and no NVIDIA driver, so every behavioural CUDA path is GPU-gated and cannot run — 459 tests are ignored. Because so much is unverifiable here, the written claims carry the weight the tests cannot, which is why their accuracy is the acceptance surface.


What changed

1. count_code is anchored against vacuity (blocking)

crates/onnx-runtime-ep-cuda/tests/no_built_in_eager_allocator.rs

the_removed_type_and_flag_are_absent_from_production_code is built entirely on the count_code helper, and every one of its assertions is hits.is_empty(). Replacing the helper's comment filter with .filter(|_line| false) yields a scanner that reads no lines at all, which satisfies every emptiness assertion trivially — the file stayed 6/6 green and the crate pair stayed at 464/0.

I reproduced the reviewer's compound exploit: with the helper blinded, reintroducing pub const CUDA_VMM_ENV: &str = "ONNX_GENAI_CUDA_VMM"; into production vmm_allocator.rs kept the whole file green (M2 below). Without the blinding, that same regression is killed by two tests (M3). So the test works today; the helper it reads through was unpinned.

Correction to #1465's disclosures #1 and #2. They named the_scan_can_observe_an_eager_call_site_that_is_known_to_exist as the protecting anchor. That was wrong: at #1465 that test used count only and never called count_code. The companion the_removal_stays_explained_in_prose_and_the_code_scan_can_tell_the_difference did not anchor it either — it asserts count sees ONNX_GENAI_CUDA_VMM and count_code does not, and both remain true when count_code is blinded (verified: that test stayed green under M1).

The fix extends the anchor test so it now pins both helpers, and its doc comment states why they must be pinned separately. The new assertion is count_code(&crate_src("onnx-runtime-cuda-memory"), "CUDA_PHYSICAL_HANDLE_POOL_BYTES_ENV") being non-empty — a constant declared on a code line that is not going away, since it is the surviving pool-bound variable that the shipped-constraints table documents.

Also added, as a scope note in the module docs: the eager-site allowlist counts malloc_sync/free_sync only. cudnn/mod.rs:621 reaches the device eagerly via .alloc_zeros::<u8>(...) — a third eager device allocation, pre-existing, untouched, and outside the DeviceAllocator seam, so criterion 2 is unaffected. The point is that the allowlist should not read broader than it is.

2. The granularity capability claim is dropped (blocking)

allocation_granularity() (virtual_memory.rs:2079-2093) ends:

if result == cu::CUresult::CUDA_SUCCESS && granularity > 0 { granularity } else { 2 << 20 }

A driver refusal or a reported zero is silently replaced with 2 MiB. Both callers — CudaVirtualBacking::granularity() (:2290) and PhysicalHandlePool (:1378) — go through it, so the granularity == 0 early-return in build() (vmm_allocator.rs:1160) is unreachable from the CUDA provider. Independently verified.

Criterion 3's substance survives. reserve() (virtual_memory.rs:2312) propagates cuMemAddressReserve failures through check() with no fallback, so an unsupported device still fails fatally at construction with the intended diagnostic. The probe simply stands on one leg, not two.

Per the review, I took the smaller in-scope fix — state honestly which call detects the unsupported device — rather than rewriting allocation_granularity(). That would be a production behaviour change to a pre-existing untouched function, and it cannot be exercised on this host.

The brief listed three claim sites. There is a fourth, which I fixed too:

Site Was Now
provider.rs:408 diagnostic "…work on this device with a non-zero allocation granularity" Names cuMemAddressReserve as the detector and says the granularity query is not a capability check
MEMORY_MANAGEMENT_MODEL_DESIGN.md:734 (table) "A reported granularity of zero is treated as 'VMM unsupported'" "This is not a capability probe … an unsupported device is detected by cuMemAddressReserve and not here"
MEMORY_MANAGEMENT_MODEL_DESIGN.md:708 (prose, not in the brief) "cuMemAddressReserve … and the driver's reported allocation granularity, both of which a device without VMM support refuses" Rewritten: cuMemAddressReserve is the single init-time detector; the granularity query is explicitly not a second one
PR-body criterion-3 evidence "build() already calls granularity() (errors on zero)" Corrected in the table below

Leaving :708 would have left the document contradicting its own table.

I also recorded the fact at the two code sites — a doc comment on allocation_granularity and a comment on the now-unreachable granularity == 0 guard. This matters because the error will not self-correct: a CUDA host following the "what a CUDA host must check" item 3 will observe the diagnostic working via reserve() and conclude the documented boundary was accurate.

3. The retained physical-handle pool rows are corrected (blocking)

As shipped in provider.rs:

Path Default pool Source
standalone / plugin, no governor 256 MiB, on Some(DEFAULT_STANDALONE_PHYSICAL_POOL_BYTES) at :841, const 256 << 20 at :277
governed + dynamic lending 256 MiB, on auto_dynamic_lending.then_some(256usize << 20) at :821
governed, non-lending off env var only — the one case documented correctly

The arena constructors do physical_handle_pool_bytes().or(default_pool_bytes) (vmm_allocator.rs:942, 979, 1086, 1116), so ONNX_GENAI_CUDA_PHYSICAL_HANDLE_POOL_BYTES overrides a default that is already present — it does not enable one.

Two documents said the opposite, so this was systematic rather than a typo. Both corrected:

  • docs/memory/MEMORY_MANAGEMENT_MODEL_DESIGN.md:735 — was "Off unless …POOL_BYTES is a positive byte count"
  • wiki/memory/Memory Management for Beginners.md:334 — was "默认关闭;…设为正整数字节数开启". Kept in Chinese, matching the surrounding register.

Criterion 11 requires retained-pool bounds documented as shipped constraints, and the table's own framing is "the first things worth knowing when diagnosing it" — so an operator asking "is device memory being retained?" was getting the wrong answer on the most common path. Both rows now also state that zero or unparseable means "fall back to the path default", never "a pool of zero".

4. Removal-record line count corrected (non-blocking)

device_allocator.rs was 304 lines, not 206. No figure matched 206:

$ git show 21370921:crates/onnx-runtime-cuda-memory/src/device_allocator.rs | wc -l
     304

non-comment 181; non-blank non-comment 166; diffstat agrees at −304. Full record below.

5. Two cheap items (optional, both done)

production_physical_pool_enabled() (vmm_allocator.rs:122) had no coverage at all — mutating its body to true survived everywhere including the engine suite. Its semantics changed in this PR: it no longer requires the removed flag, so ONNX_GENAI_CUDA_PHYSICAL_HANDLE_POOL_BYTES set alone was previously ignored and is now honoured. That change is correct and was disclosed — its sole consumer is engine/load.rs:608 → uses_governed_physical_pool → cuda_weight_startup_reservation, and since the arena now always applies physical_handle_pool_bytes().or(default), leaving the predicate gated on a deleted flag would make the engine mispredict whether an authority-owned pool exists. Behaviour kept; now covered.

The new test computes its expectation from the already-pinned parse_physical_handle_pool_bytes helper rather than restating it, so it asserts the composition — that the predicate asks the environment exactly one question and applies no second condition. Its doc comment states the scope limit plainly: nothing in the workspace calls set_var for this variable, so it is absent when the suite runs and the expectation is false, which is what kills a true body; a developer who has exported the variable will still see it pass, because it checks agreement rather than a fixed answer.

device_allocator_gpu.rs, injection_is_refused_once_the_live_mechanism_has_served_memory. Confirmed: with_memory takes mut self, so the failing call consumes and drops the provider, tearing down the arena while buffer is still outstanding; the test then built a second provider, let _ = provider;, and dropped buffer. The stated intent was not demonstrated, and the drop could hit a teardown assertion or a use-after-free on a real device.

I fixed it, and the fix required a judgement I should flag: with mut self there is no way to hold a provider across a failed with_memory at all, so "the provider is unchanged by the refusal" is not expressible against this API — reordering alone cannot rescue it. The test now releases the buffer through the original provider first (which is what actually demonstrates "the mechanism that served the pointer releases it"), then attempts the injection, which is still refused because ep_allocations is monotonic so served > 0 stays true. Nothing is outstanding when the provider is consumed. The reasoning is written into the test's doc comment. This is GPU-gated so it cannot run here, but it type-checks under --features gpu-tests --all-targets.


Criterion-by-criterion

# Criterion Status Evidence
1 Eager allocator removed from the memory crate ✅ met device_allocator.rs deleted (−304); the_cuda_memory_crate_has_no_eager_allocation_sites
2 No production EP-managed allocation falls back to built-in cuMemAlloc ✅ met Nothing to fall back to; the_removed_type_and_flag_are_absent_from_production_code, now anchored at depth 1. Scope of the allowlist (malloc_sync/free_sync only; cudnn alloc_zeros excluded and why) disclosed in the module docs
3 Unsupported device fails fatally at construction with a diagnostic ⚠️ met on one leg cuMemAddressReserve in CudaVirtualBacking::reserve is the sole init-time detector — failure propagates through check() with no fallback. The granularity probe is not a second leg: allocation_granularity substitutes 2 MiB for a refusal or a zero, making the granularity == 0 guard unreachable from the CUDA provider. Corrected at all four sites; regraded from #1465
4 DeviceAllocator contract unchanged; injection still supported ✅ met Trait untouched; with_memory unchanged and still honoured
5 Selection flag removed, not deprecated ✅ met ONNX_GENAI_CUDA_VMM / CUDA_VMM_ENV deleted; absence pinned in code, presence pinned in prose
6 Exactly one built-in mechanism ✅ met the_memory_crate_provides_exactly_one_built_in_mechanism (exactly 1, and it is vmm_allocator.rs)
7 Injection is authoritative and never silently ignored ✅ met with_memory retires the arena; refused only for foreign device or already-served memory, before the offered allocator is used
8a Structural claims asserted on a host with no GPU ✅ met no_built_in_eager_allocator.rs, 6/6, runs everywhere
8b Behavioural claims asserted ❌ unmet — cannot be met here 459 tests ignored; no CUDA device and no NVIDIA driver on macOS/arm64
9 Removal documented where the reader is ✅ met Design doc, architecture doc, wiki, and in-module prose in vmm_allocator.rs
10 Real-model benchmark showing no regression ⬜ outstanding — no number invented Requires a CUDA host. Deliberately left unticked
11 Retained-pool bounds documented as shipped constraints ✅ met with this PR Both tables corrected: 256 MiB default on standalone/plugin and governed-lending; env-only on governed non-lending; the variable overrides rather than enables. Regraded from #1465
12 Accurate removal record ✅ met Line count corrected 206 → 304; full record below

Criterion 8 is kept split into its met and unmet halves, and criterion 10 stays unticked with no number invented — that handling was explicitly endorsed in review.


Timemachine removal record (criterion 12)

Field Value
Removed file crates/onnx-runtime-cuda-memory/src/device_allocator.rs — 304 lines (181 non-comment; 166 non-blank non-comment)
Removed types CudaDeviceAllocator, QuarantinedCudaAllocation, and impl DeviceAllocator for CudaDeviceAllocator
Removed flags ONNX_GENAI_CUDA_VMM (env var) and its constant CUDA_VMM_ENV
Removed call paths The pub mod device_allocator; export from onnx-runtime-cuda-memory/src/lib.rs; the device_allocator re-export from onnx-runtime-ep-cuda/src/lib.rs; and the CUDA EP's VMM→eager fallback in provider.rs
Last supporting commit 4d2b2cc5 — feat(memory): add stream-ordered owning release, the last commit to touch the file before removal
Last supporting tree 21370921 (Phase 6 head, this PR's base)
Removal commit / PR This PR; originally proposed in #1465
Removal rationale A second built-in mechanism is a second thing every accounting, capture and teardown invariant must hold for, and a fallback that is never measured is a fallback nobody can vouch for. The eager path produced device allocations that were not charged to the ledger and could not be made during CUDA-graph capture, so falling back to it was a silent capability downgrade that surfaced much later as an unexplained accounting or capture failure. Removing the built-in mechanism does not remove the capability: DeviceAllocator is unchanged and a caller who wants eager cuMemAlloc injects it through CudaExecutionProvider::with_memory.
Exact recovery git show 21370921:crates/onnx-runtime-cuda-memory/src/device_allocator.rs

What a CUDA host must still check before merge

Nothing below can be executed on macOS/arm64.

  1. The default path reaches the VMM arena and allocations are charged to the ledger.
  2. with_memory injection is honoured, is authoritative, and is refused in exactly the two intended cases.
  3. An unsupported device fails fatally at construction with the intended diagnostic. Note when checking this: the detector is cuMemAddressReserve and only that. Observing the diagnostic fire does not confirm the granularity probe contributes — it cannot, since a refusal or a zero is replaced with 2 MiB.
  4. Teardown synchronization: in-flight stream work is awaited before physical handles are released.
  5. injection_is_refused_once_the_live_mechanism_has_served_memory passes with its new ordering, and no teardown assertion fires.
  6. Criterion 10: a real-model benchmark showing no regression against Phase 6.
  7. New (from review): a user setting only ONNX_GENAI_CUDA_PHYSICAL_HANDLE_POOL_BYTES on the governed non-lending path newly gets an authority-owned pool — the predicate no longer requires the deleted flag — and therefore newly faces adopt_memory_governor's authority-match check. Confirm that a mismatched authority produces the intended error rather than a surprise at load.

Mutation table

Every mutation used Python exact-string replacement with a count == 1 guard, printed the mutated line, and was re-confirmed with git diff -U1. No sed -i '' '<N>s/...' line-number substitution was used, per the two harness hazards recorded on this task.

# Mutation Result Notes
M1 count_code: .filter(|line| !line.trim_start().starts_with("//")) → .filter(|_line| false) — at ec75da9d, before my fix 🔴 SURVIVED — 6/6 green The defect. Confirms the reviewer's finding. the_removal_stays_explained_… also stayed green, confirming it does not anchor the helper
M2 M1 plus reintroducing pub const CUDA_VMM_ENV: &str = "ONNX_GENAI_CUDA_VMM"; into production vmm_allocator.rs 🔴 SURVIVED — 6/6 green Compound exploit reproduced: a deleted production flag returns undetected
M3 The CUDA_VMM_ENV reintroduction alone, helper healthy ✅ killed by 2 tests the_removed_type_and_flag_are_absent_from_production_code + the_removal_stays_explained_…. Confirms the test works and only the helper was unpinned
M4 M1 repeated after my fix ✅ killed the_scan_can_observe_an_eager_call_site_that_is_known_to_exist fails: "the code scan found no occurrence of a constant that is declared on a code line right now … every count_code(..).is_empty() assertion below is vacuous: {}"
M5 My own new assertion: needle "CUDA_PHYSICAL_HANDLE_POOL_BYTES_ENV" → "…_ENV_NOPE" ✅ killed Proves the new anchor reads the actual scan result rather than passing unconditionally — it is not itself vacuous
M6 production_physical_pool_enabled() body → true ✅ killed by the new test Survived everywhere at ec75da9d, including the engine suite

All mutations restored and each restore re-verified with git diff.

A third harness hazard, hit and caught. After M4 the naive restore needle .filter(|_line| false) occurred twice — once in code and once in my new doc comment, which quotes the mutation to explain it. My harness's count == 1 guard refused the restore rather than silently mutating the wrong line. This is exactly hazard #2 from the brief (the reviewer's count == 1 on true finding 2 and leaving a file mutated), arriving from the opposite direction. I restored using a unique two-line anchor (.lines() + the filter) and verified with git diff. Recording it because the guard is what caught it: a restore step needs the same discipline as the mutation step.


Validation

Baseline established at ec75da9d before any edit.

Suite Baseline After Verdict
memory + CUDA (7 crates: cuda-memory, ep-cuda, memory-abi, memory-governor, memory-host, memory-testplugin, virtual-memory) 696 / 0 / 459 697 / 0 / 459 ✅ +1 = the new production_physical_pool_enabled test
nxmem (memory-abi 55 + memory-host 55 + memory-testplugin 6) 116 / 0 116 / 0 ✅ unchanged
no_built_in_eager_allocator.rs 6 / 0 6 / 0 ✅
cargo fmt -p onnx-runtime-cuda-memory -p onnx-runtime-ep-cuda -- --check clean clean ✅
cargo clippy … --all-targets no diagnostic in touched files no diagnostic in touched files ✅
rustdoc onnx-runtime-cuda-memory (where every new doc comment lives) clean clean ✅
cargo check -p onnx-runtime-ep-cuda --features gpu-tests --all-targets clean clean ✅ GPU-gated edits type-check
onnx-genai-engine (consumer of the changed predicate) — 403 passed, 2 failed ✅ both on the known pre-existing list

Clippy's one error (approximate value of f32::consts::PI) is in crates/onnx-runtime-ep-cuda/src/kernels/standard_attention.rs, which this PR does not touch — pre-existing. Every clippy diagnostic sits in kernels/, optimizer.rs, ep-cpu, or two unrelated GPU test files; none are in the five files changed here. ep-cuda's rustdoc warnings are the pre-existing intra-doc-link set in kernels/ and deferred_release.

Known pre-existing failures were not touched and are not reported.


Corrections to the brief

Stated with evidence rather than complied with silently.

  1. The granularity claim had a fourth site, not three. MEMORY_MANAGEMENT_MODEL_DESIGN.md:708 says the capability is exercised by "cuMemAddressReserve … and the driver's reported allocation granularity, both of which a device without VMM support refuses" — the same wrong claim in prose, a few lines above the table row the brief did list. Fixing only the three listed sites would have left the document contradicting its own corrected table. Fixed.

  2. Edit 5b could not be fixed by reordering. The brief framed it as "fix the ordering if you can do so confidently". Reordering alone cannot work: with_memory(mut self, …) consumes the provider on the error path too, so there is no arrangement in which a provider survives a refused injection. The test's stated intent is not expressible against this API. I rewrote it to assert what is true and safe — release through the original mechanism first, then observe the refusal, which still fires because served is monotonic. Flagging it because it is a slightly larger change than "reorder two lines", though still test-only.

  3. The baseline crate set needed pinning down. -p onnx-runtime-cuda-memory -p onnx-runtime-ep-cuda alone gives 464/0/458, and all eleven memory+CUDA crates give 785/0/459. The brief's 696/0/459 is the seven-crate set named in the table above (40+424+55+95+55+6+21 = 696; ignored 68+390+1 = 459). The brief's numbers were right; I note the composition so the next author reproduces the same set rather than a different one. 464 is separately correct as the two-crate figure quoted in the Edit 1 blinding description.

I found nothing wrong with the substance of Edits 1–4, and I verified each independently rather than taking them on report: the blinding survival, the compound exploit, the 2 MiB substitution and the resulting unreachable guard, all three pool-default paths and the .or(default) override semantics, and the 304-line count.


Do not merge, do not enable auto-merge, do not enqueue. This stays open for human review, like every PR in this stack. CI will sit pending for hours due to the Actions backlog; that says nothing about this change, and local validation is above.

Phase 7's production change was reviewed and found correct. This revises the
claims that stood in place of verification that cannot be done on a host with
no CUDA, and pins the one test helper that was still unpinned.

- Anchor `count_code` in `no_built_in_eager_allocator.rs`. Every assertion
  built on it is an `is_empty()`, so blinding the helper to
  `.filter(|_line| false)` left the file 6/6 green and let a resurrected
  `CUDA_VMM_ENV` back into production undetected. The existing anchor test
  covers `count` only, and the prose/code companion stays green when the
  helper is blinded. Add a positive `count_code` assertion against a constant
  that lives on a code line now.

- Drop the granularity capability claim. `allocation_granularity` substitutes
  2 MiB for a driver refusal or a reported zero, so the arena builder's
  `granularity == 0` guard is unreachable from the CUDA provider and
  `cuMemAddressReserve` is the sole init-time detector. Corrected in the
  provider diagnostic and in both design-doc passages, and recorded at the two
  code sites so it is not re-derived.

- Correct the retained physical-handle pool rows. It is on at 256 MiB by
  default on the standalone/plugin path and on the governed lending path; the
  env var overrides that default rather than enabling a pool. Both the design
  doc and the Chinese wiki table said it was off by default.

- Cover `production_physical_pool_enabled`, which had no coverage at all and
  whose meaning changed in this phase.

- Fix the drop ordering in the GPU-gated late-injection test: `with_memory`
  takes `mut self`, so a refused injection consumed and dropped the provider
  while a buffer was still outstanding.

No production behaviour changes.

Part of #1186 — Phase 7 only.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@justinchuby justinchuby added the timemachine Major removal preserved as an architectural time-machine reference label Aug 19, 2026
@codecov

codecov Bot commented Aug 19, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 79.55%. Comparing base (2137092) to head (31f3a2d).

Additional details and impacted files

Impacted file tree graph

@@                      Coverage Diff                      @@
##           justinchuby-sturdy-potato    #1468      +/-   ##
=============================================================
+ Coverage                      79.38%   79.55%   +0.17%     
=============================================================
  Files                            373      374       +1     
  Lines                         165958   168905    +2947     
  Branches                      165958   168905    +2947     
=============================================================
+ Hits                          131740   134376    +2636     
- Misses                         29372    29669     +297     
- Partials                        4846     4860      +14     
Flag Coverage Δ
mlas 85.09% <ø> (?)
offline 79.45% <100.00%> (+0.07%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...tes/onnx-runtime-cuda-memory/src/virtual_memory.rs 0.00% <ø> (ø)
...ates/onnx-runtime-cuda-memory/src/vmm_allocator.rs 18.19% <100.00%> (+1.64%) ⬆️

... and 4 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 553.93 µs 1.42 ms +156.3%
🔴 matmul/medium_generic_f32_threads=8/32x512x512 900.29 µs 2.22 ms +146.5%
🔴 block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 57.73 µs 138.26 µs +139.5%
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 52.57 µs 107.53 µs +104.5%
🔴 matmul/large_generic_f32_threads=8/32x1024x1024 3.70 ms 7.42 ms +100.5%
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 41.02 µs 82.01 µs +99.9%
🔴 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 347.54 µs 651.53 µs +87.5%
🔴 matmul/medium_generic_f16_threads=8/32x512x512 28.41 µs 49.13 µs +72.9%
🔴 matmul/small_generic_f16_threads=8/1x256x256 27.82 µs 48.01 µs +72.6%
🔴 matmul/small_generic_f32_threads=8/1x256x256 32.29 µs 55.37 µs +71.5%
🔴 matmul/large_generic_f16_threads=8/32x1024x1024 77.63 µs 129.84 µs +67.3%
🔴 gather/large_f32_threads=1-internal/131072 24.87 µs 39.38 µs +58.4%
🔴 gather/large_bf16_threads=1-internal/131072 11.47 µs 17.94 µs +56.3%
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.43 ms 2.23 ms +56.0%
🔴 matmul/medium_generic_f32_threads=1/32x512x512 2.16 ms 3.26 ms +51.2%
🔴 matmul/large_generic_f16_threads=1/32x1024x1024 73.34 µs 110.47 µs +50.6%
🔴 matmul/medium_generic_bf16_threads=8/32x512x512 399.34 µs 588.79 µs +47.4%
🔴 matmul/large_generic_f32_threads=1/32x1024x1024 8.63 ms 12.72 ms +47.3%
🔴 matmul/large_generic_bf16_threads=1/32x1024x1024 1.88 ms 2.78 ms +47.3%
🔴 matmul/small_generic_f32_threads=1/1x256x256 34.15 µs 50.08 µs +46.6%
🔴 matmul/medium_generic_f16_threads=1/32x512x512 31.69 µs 46.33 µs +46.2%
🔴 matmul/small_generic_bf16_threads=8/1x256x256 29.93 µs 43.44 µs +45.1%
🔴 gather/large_f16_threads=1-internal/131072 11.26 µs 16.24 µs +44.2%
🔴 matmul/small_generic_bf16_threads=1/1x256x256 30.26 µs 42.79 µs +41.4%
🔴 matmul/small_generic_f16_threads=1/1x256x256 28.53 µs 40.10 µs +40.6%
🔴 qwen3_sampling_processors/top_k_top_p_fast 666.78 µs 881.02 µs +32.1%
⚠️ gather/medium_f32_threads=1-internal/32768 3.58 µs 4.63 µs +29.5%
⚠️ sampling_latency/min_p_per_token 236.08 µs 301.32 µs +27.6%
⚠️ matmul/medium_generic_bf16_threads=1/32x512x512 532.15 µs 669.02 µs +25.7%
⚠️ gather/small_f32_threads=1-internal/4096 606.8 ns 756.1 ns +24.6%
⚠️ kv_cache/alloc_dealloc_pages 42.75 µs 51.71 µs +20.9%
⚠️ add/large_f32_threads=1-internal/4194304 604.05 µs 723.21 µs +19.7%
⚠️ grammar_masking/llguidance_compute_mask/32 97.32 µs 115.49 µs +18.7%
⚠️ gather/small_f16_threads=1-internal/4096 438.5 ns 519.6 ns +18.5%
⚠️ gather/medium_bf16_threads=1-internal/32768 2.27 µs 2.65 µs +16.6%
⚠️ reduce_mean/large_f32_threads=1-internal/262144 964.41 µs 1.12 ms +16.4%
✅ gather/medium_f16_threads=1-internal/32768 2.24 µs 2.56 µs +14.2%
✅ gather/small_bf16_threads=1-internal/4096 440.1 ns 501.6 ns +14.0%
✅ logit_processing/seven_processor_chain_per_step 364.76 µs 415.41 µs +13.9%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 578.64 µs 644.10 µs +11.3%
✅ tokenization/encode_tokens_per_second 460.79 µs 478.96 µs +3.9%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.57 ms 2.66 ms +3.6%
✅ tokenization/decode_tokens_per_second 7.86 ms 8.04 ms +2.4%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 7.20 ms 7.36 ms +2.3%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 4.60 ms 4.70 ms +2.2%
✅ sampling_latency/top_p_per_token 450.62 µs 456.29 µs +1.3%
✅ add/medium_f16_threads=1-internal/262144 111.71 µs 112.14 µs +0.4%
✅ reduce_mean/small_f32_threads=1-internal/4096 15.88 µs 15.81 µs -0.4%
✅ add/large_bf16_threads=1-internal/4194304 1.69 ms 1.67 ms -1.6%
✅ qwen3_sampling_processors/top_k_partial_selection 182.51 µs 178.24 µs -2.3%
✅ add/large_f16_threads=1-internal/4194304 1.84 ms 1.80 ms -2.4%
✅ reduce_mean/medium_f32_threads=1-internal/65536 244.59 µs 234.47 µs -4.1%
✅ sampling_latency/greedy_per_token 3.92 µs 3.67 µs -6.3%
✅ sampling_latency/top_k_per_token 62.49 µs 58.04 µs -7.1%
✅ add/medium_f32_threads=1-internal/262144 27.54 µs 24.59 µs -10.7%
✅ add/medium_bf16_threads=1-internal/262144 114.56 µs 99.19 µs -13.4%
🟢 add/small_f32_threads=1-internal/1024 245.1 ns 204.2 ns -16.7%
🟢 add/small_f16_threads=1-internal/1024 528.8 ns 440.1 ns -16.8%
🟢 add/small_bf16_threads=1-internal/1024 544.2 ns 440.9 ns -19.0%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 7.99 5.23 5.86 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

@justinchuby
justinchuby changed the base branch from justinchuby-sturdy-potato to justinchuby-memory-phase-7-vmm-only August 20, 2026 14:53
justinchuby added a commit that referenced this pull request Aug 22, 2026
> **A100 safety update (head `13c37a7f`):** expected shared-prefix
admission failures are now all-or-private: K/V commit is transactional,
a failed V share rolls K back before enqueue, and only a successful
rollback permits private-KV fallback. Fatal kernel faults remain
process-fatal; the new cross-process VMM probe proves a worker fault
does not affect an actively executing peer, while the current server is
still in-process and does not yet consume physical shared-prefix
metadata. See the [final verified
conclusion](#1579 (comment)).

Collapses phases 1–7 of the #1186 memory architecture rework onto
current `main` as a single merge. The stack forked 270 commits ago (176
of them touching `crates/`), so rebasing it layer by layer would mean
solving the same 17 conflicts fourteen times, with no reviewer ever
looking at the intermediate states.

Closes the stack: #1252 #1263 #1279 #1283 #1301 #1341 #1349 #1426 #1440
#1448 #1454 #1462 #1465 #1468 #1533.

**A review guide is in the first comment.** It is the part worth reading
— this diff is 103 files, but only three decisions in it are ones a
compiler cannot check.

## What was already reviewed, and what wasn't

Every phase was reviewed and approved by a session that was not its
author, under the rejection-lockout rule (a rejected author never writes
the next revision).

| Phase | PR | Verdict |
|---|---|---|
| 1–5 | #1252 → #1426 | approved |
| 6 | #1440, #1448, #1454 | **rejected** ×3 |
| 6 | #1462 | approved |
| 7 | #1465 | **rejected** |
| 7 | #1468 | approved |
| 7 test fixes | #1533 | approved, A100-verified |

The four rejected rounds were not replaced — the later rounds are
stacked **on top of** them. So the tree here is the approved state, but
the history contains the rejected commits. **#1440 / #1448 / #1454 /
#1465 should not be reviewed individually**; they close automatically.

**Not reviewed anywhere:** the conflict resolutions themselves. That is
what this PR is for.

## Verification

Run here, on this merged tree:

- `cargo check --workspace --all-targets` — clean
- `cargo check -p onnx-genai-engine --features cuda,native-backend
--all-targets` — clean. Worth calling out: the default feature set does
**not** compile the `cfg(cuda)` code, which is where the riskiest edits
are. Checking only the default set would have missed a real break (see
the guide).
- `cargo test` over 7 crates — **1105 passed, 2 failed, 88 ignored**
- `cargo clippy --workspace --all-targets`
- `cargo fmt`

**The 2 failures and the 1 clippy error are pre-existing and proven so,
not merge damage:**

-
`platform_capacity::{disk_capacity_is_measured_for_the_working_directory,
an_explicit_byte_limit_is_honored_without_a_device_query}` —
`platform_capacity.rs` is byte-identical to `main` (`git diff
origin/main HEAD -- ` that file is empty). The cause is a macOS-only FFI
layout bug: `fsblkcnt_t` is 4 bytes on macOS (confirmed: `sizeof(struct
statvfs)=64`, `sizeof(fsblkcnt_t)=4`) while the Rust struct declares
`f_blocks`/`f_bfree`/`f_bavail` as `u64`. CI is Linux, where it is 8
bytes. Left alone: it is a real bug but not this PR's.
- `optimizer.rs` clippy `approx_constant` — present on both parents; the
merged file is byte-identical to `main`.
- Three files had `rustfmt` drift already present on `main`
(`matmul_nbits.rs`, `normalization.rs`, `optimizer.rs`). Reverted rather
than swept in, to keep the diff readable.

**Final A100 revalidation (head `13c37a7f`):** full CUDA-memory GPU
suite passed; CUDA EP default-parallel lib suite passed (**488 passed /
17 ignored**); real fp16 GQA covered three requests × two interleaved
decode steps with two shared peers and one transactional private
fallback, byte-identical to independent GPU and CPU references. A
device-started/event-gated worker remained healthy through a peer
process `CUDA_ERROR_ILLEGAL_ADDRESS`, owner exit, replacement worker,
and process restart. Governor tests, targeted Miri, CI Clippy, fmt, and
CUDA honesty also passed.

## Still open after this merges

- `memory-plugin-provider-wiring` — the back half of Phase 6 criterion 8
(provider/context pinning), which fell in the gap between phases.
**#1186 must not be closed until it lands.** This merge makes it more
tractable, not less: see decision 3 in the guide.
- `memory-deferred-invariant-asserts` — two surviving mutants found by
#1462's reviewer, adjudicated NON-BLOCKING and not defects.
- #1533's CUDA honesty guard warns on a legitimately portable anchor.
Should be moved out of the CUDA test binary or allow-listed rather than
left warning.

---------

Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <copilot@github.com>
Copilot-Session: 46c5d75b-8146-489c-b82f-08ee29c27ce4
Copilot-Session: 39ff6824-d35f-4d3b-8f5f-043a7119a100
Copilot-Session: c80f8522-983c-47f7-8241-2155a823aabe
@justinchuby

Copy link
Copy Markdown
Owner Author

Closing: superseded by #1579, which collapsed phases 1–7 into a single merge — landed as a36964280 (squash, 126 files, +43956 −3286).

This PR was approved by a session that was not its author, and its work is in main.

Containment verified, not assumed:

$ git merge-base --is-ancestor 31f3a2dde ebceab071   # ebceab071 = #1579 head
→ ancestor

Every one of the 15 stack heads (#1252 → #1533) is a literal ancestor of the merged head, and I separately confirmed the merged head's content reached main: of the 126 files #1579 touched, exactly one differs from main — crates/onnx-runtime-session/src/executor/tests.rs, where main carries 120 extra lines from #1703 (Expand-broadcast mask tests). That is the conflict resolution correctly preserving the other side, not content loss.

Why this didn't close itself: #1579's body says Closes the stack: #1252 #1263 …, which GitHub does not parse — the closing keyword must be immediately followed by the reference, and in any case closing keywords only ever close issues, never pull requests. So all 15 stayed open by mechanism, not by intent.

#1186 stays open: #1579 explicitly gates it on memory-plugin-provider-wiring (the back half of Phase 6 criterion 8, provider/context pinning), which has not landed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

timemachine Major removal preserved as an architectural time-machine reference

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant