Skip to content

refactor(memory)!: make VMM the sole built-in CUDA memory mechanism (#1186 Phase 7) - #1465

Closed
justinchuby wants to merge 1 commit into
justinchuby-sturdy-potatofrom
justinchuby-memory-phase-7-vmm-only
Closed

justinchuby wants to merge 1 commit into
justinchuby-sturdy-potatofrom
justinchuby-memory-phase-7-vmm-only

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Part of #1186 — Phase 7 only.

Removes the built-in CUDA eager cuMemAlloc allocator and makes the VMM arena the sole built-in CUDA memory mechanism for EP-managed device allocations. Stacks on Phase 6 (justinchuby-sturdy-potato, 21370921).

What changed

The CUDA EP carried two built-in mechanisms — an eager cuMemAlloc allocator and the VMM arena — with the arena behind an opt-in flag and the eager allocator as the silent fallback when the arena could not be built. A missing capability quietly selected a different mechanism whose bytes were not charged the same way, and the operator's only evidence was a log line.

  • The arena is unconditional. No environment opt-in. ONNX_GENAI_CUDA_VMM is deleted, not deprecated.
  • Unsupported fails at construction. CudaVmmAllocator::build() already exercises the capability (cuMemAddressReserve, and the driver's reported allocation granularity, which a device without VMM support refuses). That failure now returns Err through vmm_unavailable(...), naming the device ordinal, the driver's own message, the entry points that constitute the support boundary, any requested managed limit, and with_memory as the way out.
  • Injection is preserved and becomes authoritative. DeviceAllocator is untouched. with_memory retires the arena and the injected mechanism serves everything afterwards. It is refused — before the offered allocator is used — only for a foreign device, or for a mechanism with memory outstanding. Both return Err; no successful builder call is ignored.
  • One mechanism in the provider. CudaMemory is Vmm | Injected. There is deliberately no third built-in-eager variant, and reintroducing the fallback does not compile (see M8).

Behaviour change worth calling out

production_physical_pool_enabled() was vmm_enabled() && pool_bytes.is_some() and is now pool_bytes.is_some(). A caller who set ONNX_GENAI_CUDA_PHYSICAL_HANDLE_POOL_BYTES without the removed flag previously had it ignored and now has it honoured. This is the intended consequence of the arena becoming unconditional — the pool bound was only ever meaningful for the arena — but it is a real change and is called out rather than buried.

Deliberately not changed

auto_dynamic_lending.then_some(256 MiB) on the governed path is left as-is. It is tempting to give every governed provider a default retained pool now that VMM is universal, but the pool is authority-owned and adopt_memory_governor errors when arena.physical_pool_authority() != governor.authority_id(). Adding a default pool to the governed non-lending path would newly subject it to that check and could introduce adoption failures that cannot be tested on this host. Flagged for the CUDA host.

⚠️ This was developed on macOS/arm64 with no CUDA and no NVIDIA GPU

Not a single CUDA path executed. Read the criteria table for exactly what that leaves unverified. Criterion 10 is outstanding. No benchmark was run, no number is estimated, and the box below is unticked on purpose.

Acceptance criteria

# Criterion Status Evidence
1 Default selects VMM without an opt-in flag met (structurally); met-but-unverifiable-here (at runtime) Flag and selection branch deleted; arena built unconditionally. Runtime proof is the_default_provider_allocates_through_the_built_in_vmm_arena, which is GPU-gated and did not execute.
2 No production EP-managed allocation falls back to built-in cuMemAlloc met the_cuda_memory_crate_has_no_eager_allocation_sites — zero malloc_sync(/free_sync( in onnx-runtime-cuda-memory/src. Ran here, green. Killed by M5.
3 Unsupported VMM fails at init with an actionable diagnostic + documented boundary met (diagnostic); met-but-unverifiable-here (the driver refusal itself) built_in_vmm_failure_is_fatal_and_names_the_support_boundary, a_requested_managed_limit_is_named_in_the_unavailability_diagnostic — both ran here, green, killed by M1. Boundary documented in MEMORY_MANAGEMENT_MODEL_DESIGN.md. No host without VMM support was available to produce a real refusal. See "no new capability probe" below.
4 Injected same-device mechanisms selected authoritatively or rejected; no successful call ignored met (decision logic); met-but-unverifiable-here (the swap) injection_is_refused_for_a_device_this_provider_does_not_serve, replacing_a_mechanism_that_already_served_memory_is_refused_on_both_axes — ran here, green, killed by M2/M3/M4. The authoritative swap itself is an_injected_external_eager_allocator_replaces_the_built_in_arena, GPU-gated, did not execute.
5 VMM supplies the ordinary contract plus virtual-backing and shared-mapping met (unchanged) impl DeviceAllocator for CudaVmmAllocator and its VirtualBacking/shared-mapping surfaces are untouched by this PR; the_memory_crate_provides_exactly_one_built_in_mechanism pins that it is the only one.
6 Provider no longer stores an eager allocator beside a preferred VMM allocator met CudaMemory is Vmm | Injected. VmmInitialization/resolve_vmm_initialization deleted. M8 shows the old shape does not compile.
7 Standalone, governed, native session, plugin, ORT scratch, graph-capture, weight-streaming, KV-growth, shared-prefix paths behaviourally compatible met-but-unverifiable-here All of these compile (--workspace --all-targets, plus onnx-genai-engine --features cuda,native-backend, plus onnx-runtime-ep-cuda --features gpu-tests). Every one of their behavioural tests is GPU-gated and did not execute. Compilation is not behavioural compatibility and is not claimed to be.
8 Targeted tests prove the default reaches VMM and performs no built-in eager cuMemAlloc calls partially met — see below The "no eager calls" half ran here (no_built_in_eager_allocator.rs, 6/6, mutation-killed by M5/M6/M7). The "default reaches VMM" half is GPU-gated and did not execute.
9 Token parity, capture/fallback counters, mapped-byte accounting, underflow counters, unaccounted committed bytes, oversubscription, teardown invariants clean outstanding on a CUDA host; unchanged by inspection This PR does not modify accounting, quarantine or teardown code. The three axes (mapped_bytes/owned_bytes, allocation_bytes/unmapped_bytes) are untouched. No counter was observed at runtime, because none of these paths can run here.
10 Real-model benchmarks show no material regression OUTSTANDING Not run. Requires a CUDA host. No number is estimated and none is presented.
11 Reservation size, 2 MiB granularity, retained-pool bounds, teardown sync, device loss documented as shipped constraints met Constraints table in MEMORY_MANAGEMENT_MODEL_DESIGN.md and in the beginner wiki.
12 timemachine label + removal record met Label applied; record below.
13 Beginner wiki and formal memory docs describe VMM-only built-in behaviour, distinguishing injected mechanisms met docs/memory/MEMORY_MANAGEMENT_MODEL_DESIGN.md, docs/memory/MEMORY_ARCHITECTURE.md, docs/ep-plugin/EP_PLUGIN_EXPORT_INVENTORY.md, wiki/memory/Memory Management for Beginners.md.
  • Criterion 10 — real-model benchmarks. Not run; requires a CUDA host.

Why no new driver capability probe was added (criterion 3)

cudarc 0.19.8 does not expose CU_DEVICE_ATTRIBUTE_VIRTUAL_MEMORY_MANAGEMENT_SUPPORTED. Using the raw attribute integer would add CUDA surface that cannot be exercised on the host this was written on, which is how untested code gets shipped. Instead the existing init-time exercise is made fatal: build() already calls granularity() (errors on zero) and reserve() (cuMemAddressReserve), which is what a device without VMM support refuses. A CUDA host should confirm this produces the intended diagnostic on real unsupported hardware.

Mutation testing

Every mutation was applied, the mutated line printed to confirm it landed, the suite run, and the source restored.

# Mutation Result
M1 Weaken the fatal diagnostic so it no longer says the arena is the only mechanism Killed — built_in_vmm_failure_is_fatal_and_names_the_support_boundary
M2 reject_foreign_device silently accepts a foreign device (false && ...) Killed — injection_is_refused_for_a_device_this_provider_does_not_serve
M3 Drop the committed axis of reject_live_mechanism_replacement Killed — ..._refused_on_both_axes
M4 Drop the served axis of reject_live_mechanism_replacement Killed — ..._refused_on_both_axes
M5 Resurrect a second built-in eager DeviceAllocator in the memory crate Killed — 2 tests: no-eager-sites and exactly-one-mechanism
M6 Add a third eager cuMemAlloc call site in the EP Killed — ..._exactly_the_two_disclosed_ones
M7 Test-infrastructure: blind the source scanner so it reads nothing Killed — 4 tests, including the dedicated anchor the_scan_can_observe_an_eager_call_site_that_is_known_to_exist
M8 Reintroduce the silent eager fallback at the construction site Killed at compile time — two independent errors: crate::device_allocator does not exist, and CudaMemory::Allocator does not exist

Survivors and disclosures

One survivor found, and it was in my own mutation harness, not the code. My first attempt at M3/M4 used a shell loop that split the before/after strings on | — which is inside the expression served > 0 || committed > 0. The "mutation" therefore edited whitespace and nothing else, the suite stayed green, and I recorded a survivor that did not exist. This is Round 1's "fixtures that lied" reproduced one level up, in the tooling that judges the fixtures. It was caught by printing the mutated line; every mutation in the table above was re-run with that verification, and M3/M4 both kill.

Disclosed weaknesses in the surviving assertions:

  1. Under M7, the_cuda_memory_crate_has_no_eager_allocation_sites and the_removed_type_and_flag_are_absent_from_production_code survive individually — a blinded scanner makes "count is zero" trivially true. This is inherent to negative structural assertions and is exactly why the_scan_can_observe_an_eager_call_site_that_is_known_to_exist exists as an anchor. The anchor dies under M7. The two negative tests should never be read without it.
  2. the_removed_type_and_flag_are_absent_from_production_code excludes comment lines, so it cannot see a removed name mentioned in prose. That is deliberate (criterion 12 requires the removal to stay explained in vmm_allocator.rs), and the exclusion is itself pinned by the_removal_stays_explained_in_prose_and_the_code_scan_can_tell_the_difference, which asserts the raw scan sees the mention and the code scan does not — so the two scans provably differ on a live case.
  3. The criterion-8 GPU test's discriminator was chosen to resist the Round-2 failure. commits_on_demand() is a behavioural property of the live mechanism (the arena maps granules on demand; an eager allocator takes physical memory when asked), not a counter the provider sets about itself, and the test contains an explicit premise assertion proving an eager allocator reports false — so "the default reached something else" cannot pass. It still did not execute here, and I am not claiming it did.
  4. Not mutated: any GPU-gated assertion, because a mutation whose test cannot run proves nothing. The GPU tests in this PR are unmutated and unexecuted.

Validation actually performed on this host

Check Result
cargo check --workspace --all-targets --exclude onnx-genai-bench clean
cargo check -p onnx-genai-engine --features cuda,native-backend --all-targets clean
cargo check -p onnx-runtime-ep-cuda --features gpu-tests --all-targets clean (GPU tests type-check)
Tests: cuda-memory, ep-cuda, virtual-memory, memory-governor, memory-abi/host/testplugin 696 passed, 0 failed, 459 ignored
Phase 6 nxmem ABI suite (memory-abi + memory-host + memory-testplugin) 116 passed / 0 failed — still 116/0
no_built_in_eager_allocator.rs (new, non-GPU) 6/6
cargo fmt --check on changed crates clean
cargo clippy --all-targets on changed crates no diagnostic in any file this PR touches
cargo doc --no-deps on changed crates no warning in any file this PR touches

459 ignored is almost entirely GPU-gated tests. That number is the honest size of what could not run.

Not run here: every #[cfg_attr(not(feature = "gpu-tests"), ignore)] test; the full workspace test suite (blocked by the known pre-existing macOS MLAS link failure in mlas-sys); all benchmarks.

Pre-existing and untouched: completions.rs rustfmt drift, matmul_nbits_marlin_numerics, Windows CUDA unnecessary_unwrap, macOS MLAS, onnx-runtime-memory-governor rustdoc links. onnx-runtime-ep-cpu also fails clippy -D warnings at the base commit (verified by stashing) — also not mine.

What a CUDA host must still check before this is mergeable

  1. Criterion 10 — small-allocation, load, prefill and decode benchmarks on a real model. Nothing about performance is claimed here.
  2. Run device_allocator_gpu.rs and every provider::tests GPU test with --features gpu-tests. Four of them are new or rewritten.
  3. Confirm the diagnostic on genuinely VMM-unsupported hardware, since the boundary is inferred from reserve()/granularity() failing rather than from a capability attribute.
  4. Confirm the governed non-lending path is unaffected by removing the eager fallback — the "deliberately not changed" note above is an untested judgement call.
  5. Check nsys/ncu for cuMemAlloc_v2 in the steady-state decode region; only the two disclosed kernel-scratch sites should appear.
  6. Criterion 9's counters and teardown invariants under a real workload.

Criterion 12 — removal record (timemachine)

Removed types

  • CudaDeviceAllocator — crates/onnx-runtime-cuda-memory/src/device_allocator.rs (whole file, 206 lines)
  • QuarantinedCudaAllocation — same file
  • VmmInitialization<T> — crates/onnx-runtime-ep-cuda/src/provider.rs
  • CudaMemory::Allocator variant — renamed to CudaMemory::Injected

Removed flags

  • ONNX_GENAI_CUDA_VMM (constant CUDA_VMM_ENV) — the arena on/off switch

Removed functions / call paths

  • vmm_allocator::vmm_enabled()
  • provider::resolve_vmm_initialization()
  • The eager() construction closure and the if vmm_enabled() || auto_dynamic_lending selection branch in CudaExecutionProvider construction
  • pub use onnx_runtime_cuda_memory::device_allocator re-export from onnx-runtime-ep-cuda
  • GPU tests construction_selected_vmm_rejects_injection_instead_of_ignoring_it, default_allocator_cumemalloc_scales_one_for_one_with_requests; unit tests managed_vmm_failure_is_fatal_before_allocator_fallback, compatibility_vmm_failure_keeps_fallback_available

Last commit that supported them: 21370921 (Phase 6 head, base of this PR)

Why removed: the eager allocator existed only as the fallback for a VMM arena that was not yet the default. Once the arena is the default, keeping it means keeping a path that can be entered silently, whose allocations are not charged the way the arena's are, and whose selection is invisible except in a log line. Criterion 6 names the dual-mechanism provider state directly. The capability is preserved through DeviceAllocator injection, so nothing a caller could do before is now impossible — it just has to be asked for explicitly.

How to recover the implementation:

# The deleted eager allocator, in full:
git show 21370921:crates/onnx-runtime-cuda-memory/src/device_allocator.rs

# The dual-selection state it was chosen by:
git show 21370921:crates/onnx-runtime-ep-cuda/src/provider.rs

# Restore the file onto a working tree:
git checkout 21370921 -- crates/onnx-runtime-cuda-memory/src/device_allocator.rs

# Everything this PR removed, as one diff:
git diff 21370921..HEAD -- crates/onnx-runtime-cuda-memory crates/onnx-runtime-ep-cuda

Note that tests/device_allocator_gpu.rs in this PR contains ExternalEagerAllocator, a working eager cuMemAlloc allocator built from nothing but public API — which is both a test fixture and the migration example for anyone who needs the removed behaviour back.


Do not merge. This stays open for human review, like every PR in the memory stack.

Part of #1186 -- Phase 7 only.

The CUDA EP carried two built-in device memory mechanisms: an eager
`cuMemAlloc` allocator and the VMM arena, with the arena selected by an
opt-in environment flag and the eager allocator serving as the silent
fallback when the arena could not be built. That is the shape the memory
model argues against: a missing capability quietly selected a different
mechanism whose bytes were not charged the same way, and the operator's
only evidence was a log line.

Delete the eager allocator and the dual-selection state. The arena is now
constructed unconditionally, and failing to construct it is fatal at
provider construction with a diagnostic naming the device, the driver's
own message, the driver entry points that constitute the support
boundary, any requested managed limit, and `with_memory` as the supported
way to supply a different mechanism.

Removing the built-in implementation does not remove the capability.
`DeviceAllocator` is unchanged. `with_memory` becomes authoritative
rather than refusing: it retires the arena and the injected mechanism
serves everything afterwards. It is refused, before the offered allocator
is used at all, only for a foreign device or for a mechanism that still
has memory outstanding -- both return `Err`, so no successful builder
call is ignored.

Removed:
- `CudaDeviceAllocator`, `QuarantinedCudaAllocation`
  (crates/onnx-runtime-cuda-memory/src/device_allocator.rs)
- `CUDA_VMM_ENV` / `ONNX_GENAI_CUDA_VMM`, `vmm_enabled()`
- `VmmInitialization`, `resolve_vmm_initialization`
- `CudaMemory::Allocator` (renamed `Injected`)

Behaviour change worth calling out: `production_physical_pool_enabled()`
no longer requires the removed flag, so a caller who set
`ONNX_GENAI_CUDA_PHYSICAL_HANDLE_POOL_BYTES` alone previously had it
ignored and now has it honoured.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 19, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 80.76923% with 5 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.17%. Comparing base (2137092) to head (ec75da9).

Files with missing lines Patch % Lines
...ates/onnx-runtime-cuda-memory/src/vmm_allocator.rs 80.76% 5 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@                      Coverage Diff                      @@
##           justinchuby-sturdy-potato    #1465      +/-   ##
=============================================================
+ Coverage                      79.38%   80.17%   +0.79%     
=============================================================
  Files                            373      374       +1     
  Lines                         165958   168899    +2941     
  Branches                      165958   168899    +2941     
=============================================================
+ Hits                          131740   135413    +3673     
+ Misses                         29372    28622     -750     
- Partials                        4846     4864      +18     
Flag Coverage Δ
mlas 85.22% <ø> (?)
offline 80.08% <80.76%> (+0.69%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...ates/onnx-runtime-cuda-memory/src/vmm_allocator.rs 17.51% <80.76%> (+0.96%) ⬆️

... and 11 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 45.30 µs 119.50 µs +163.8%
🔴 block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 56.54 µs 139.72 µs +147.1%
🔴 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 354.90 µs 619.18 µs +74.5%
⚠️ tokenization/encode_tokens_per_second 348.60 µs 434.93 µs +24.8%
⚠️ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 630.97 µs 773.89 µs +22.7%
⚠️ sampling_latency/greedy_per_token 3.02 µs 3.55 µs +17.5%
✅ tokenization/decode_tokens_per_second 5.85 ms 6.59 ms +12.7%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.25 ms 3.64 ms +12.2%
✅ qwen3_sampling_processors/top_k_partial_selection 130.94 µs 145.07 µs +10.8%
✅ sampling_latency/top_k_per_token 46.54 µs 51.31 µs +10.2%
✅ sampling_latency/top_p_per_token 366.87 µs 403.08 µs +9.9%
✅ qwen3_sampling_processors/top_k_top_p_fast 621.12 µs 660.61 µs +6.4%
✅ sampling_latency/min_p_per_token 192.53 µs 204.75 µs +6.3%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.51 ms 5.57 ms +1.1%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 524.80 µs 530.33 µs +1.1%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.22 ms 2.24 ms +1.0%
✅ grammar_masking/llguidance_compute_mask/32 69.95 µs 70.04 µs +0.1%
✅ logit_processing/seven_processor_chain_per_step 300.88 µs 294.12 µs -2.2%
✅ kv_cache/alloc_dealloc_pages 37.27 µs 36.15 µs -3.0%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.31 ms 1.27 ms -3.0%
✅ gather/large_f16_threads=1-internal/131072 11.96 µs 11.54 µs -3.5%
✅ matmul/medium_generic_f32_threads=8/32x512x512 956.94 µs 923.01 µs -3.5%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 1.96 ms 1.86 ms -5.1%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 541.59 µs 494.17 µs -8.8%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 86.70 µs 78.67 µs -9.3%
✅ matmul/medium_generic_f16_threads=1/32x512x512 30.77 µs 27.51 µs -10.6%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 87.47 µs 77.21 µs -11.7%
✅ gather/small_f32_threads=1-internal/4096 699.5 ns 616.6 ns -11.8%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 4.05 ms 3.54 ms -12.4%
✅ gather/medium_f16_threads=1-internal/32768 2.51 µs 2.20 µs -12.7%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.84 ms 8.55 ms -13.1%
✅ gather/medium_bf16_threads=1-internal/32768 2.54 µs 2.19 µs -13.8%
✅ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 51.02 µs 43.42 µs -14.9%
🟢 matmul/medium_generic_bf16_threads=8/32x512x512 420.62 µs 352.96 µs -16.1%
🟢 matmul/medium_generic_f32_threads=1/32x512x512 2.58 ms 2.14 ms -17.2%
🟢 matmul/medium_generic_f16_threads=8/32x512x512 33.79 µs 27.64 µs -18.2%
🟢 reduce_mean/large_f32_threads=1-internal/262144 1.12 ms 908.14 µs -18.8%
🟢 matmul/small_generic_bf16_threads=1/1x256x256 35.45 µs 28.46 µs -19.7%
🟢 gather/small_f16_threads=1-internal/4096 581.6 ns 436.2 ns -25.0%
🟢 matmul/small_generic_bf16_threads=8/1x256x256 38.81 µs 29.10 µs -25.0%
🟢 add/small_f32_threads=1-internal/1024 246.8 ns 183.6 ns -25.6%
🟢 matmul/small_generic_f16_threads=8/1x256x256 39.16 µs 28.66 µs -26.8%
🟢 add/medium_f32_threads=1-internal/262144 32.31 µs 23.38 µs -27.6%
🟢 gather/small_bf16_threads=1-internal/4096 607.3 ns 435.5 ns -28.3%
🟢 add/small_bf16_threads=1-internal/1024 589.3 ns 411.5 ns -30.2%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 339.05 µs 233.30 µs -31.2%
🟢 reduce_mean/small_f32_threads=1-internal/4096 20.49 µs 13.84 µs -32.4%
🟢 add/small_f16_threads=1-internal/1024 625.2 ns 422.2 ns -32.5%
🟢 add/large_bf16_threads=1-internal/4194304 2.32 ms 1.52 ms -34.5%
🟢 add/large_f16_threads=1-internal/4194304 2.42 ms 1.56 ms -35.6%
🟢 gather/medium_f32_threads=1-internal/32768 5.46 µs 3.42 µs -37.4%
🟢 matmul/small_generic_f32_threads=1/1x256x256 57.94 µs 35.23 µs -39.2%
🟢 add/medium_f16_threads=1-internal/262144 161.34 µs 97.75 µs -39.4%
🟢 gather/large_f32_threads=1-internal/131072 41.22 µs 24.93 µs -39.5%
🟢 gather/large_bf16_threads=1-internal/131072 17.50 µs 10.27 µs -41.3%
🟢 add/medium_bf16_threads=1-internal/262144 197.90 µs 97.51 µs -50.7%
🟢 add/large_f32_threads=1-internal/4194304 1.11 ms 543.90 µs -51.0%
🟢 matmul/small_generic_f32_threads=8/1x256x256 79.35 µs 31.97 µs -59.7%
🟢 matmul/small_generic_f16_threads=1/1x256x256 69.94 µs 27.98 µs -60.0%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.13 3.21 6.37 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Aug 22, 2026
> **A100 safety update (head `13c37a7f`):** expected shared-prefix
admission failures are now all-or-private: K/V commit is transactional,
a failed V share rolls K back before enqueue, and only a successful
rollback permits private-KV fallback. Fatal kernel faults remain
process-fatal; the new cross-process VMM probe proves a worker fault
does not affect an actively executing peer, while the current server is
still in-process and does not yet consume physical shared-prefix
metadata. See the [final verified
conclusion](#1579 (comment)).

Collapses phases 1–7 of the #1186 memory architecture rework onto
current `main` as a single merge. The stack forked 270 commits ago (176
of them touching `crates/`), so rebasing it layer by layer would mean
solving the same 17 conflicts fourteen times, with no reviewer ever
looking at the intermediate states.

Closes the stack: #1252 #1263 #1279 #1283 #1301 #1341 #1349 #1426 #1440
#1448 #1454 #1462 #1465 #1468 #1533.

**A review guide is in the first comment.** It is the part worth reading
— this diff is 103 files, but only three decisions in it are ones a
compiler cannot check.

## What was already reviewed, and what wasn't

Every phase was reviewed and approved by a session that was not its
author, under the rejection-lockout rule (a rejected author never writes
the next revision).

| Phase | PR | Verdict |
|---|---|---|
| 1–5 | #1252 → #1426 | approved |
| 6 | #1440, #1448, #1454 | **rejected** ×3 |
| 6 | #1462 | approved |
| 7 | #1465 | **rejected** |
| 7 | #1468 | approved |
| 7 test fixes | #1533 | approved, A100-verified |

The four rejected rounds were not replaced — the later rounds are
stacked **on top of** them. So the tree here is the approved state, but
the history contains the rejected commits. **#1440 / #1448 / #1454 /
#1465 should not be reviewed individually**; they close automatically.

**Not reviewed anywhere:** the conflict resolutions themselves. That is
what this PR is for.

## Verification

Run here, on this merged tree:

- `cargo check --workspace --all-targets` — clean
- `cargo check -p onnx-genai-engine --features cuda,native-backend
--all-targets` — clean. Worth calling out: the default feature set does
**not** compile the `cfg(cuda)` code, which is where the riskiest edits
are. Checking only the default set would have missed a real break (see
the guide).
- `cargo test` over 7 crates — **1105 passed, 2 failed, 88 ignored**
- `cargo clippy --workspace --all-targets`
- `cargo fmt`

**The 2 failures and the 1 clippy error are pre-existing and proven so,
not merge damage:**

-
`platform_capacity::{disk_capacity_is_measured_for_the_working_directory,
an_explicit_byte_limit_is_honored_without_a_device_query}` —
`platform_capacity.rs` is byte-identical to `main` (`git diff
origin/main HEAD -- ` that file is empty). The cause is a macOS-only FFI
layout bug: `fsblkcnt_t` is 4 bytes on macOS (confirmed: `sizeof(struct
statvfs)=64`, `sizeof(fsblkcnt_t)=4`) while the Rust struct declares
`f_blocks`/`f_bfree`/`f_bavail` as `u64`. CI is Linux, where it is 8
bytes. Left alone: it is a real bug but not this PR's.
- `optimizer.rs` clippy `approx_constant` — present on both parents; the
merged file is byte-identical to `main`.
- Three files had `rustfmt` drift already present on `main`
(`matmul_nbits.rs`, `normalization.rs`, `optimizer.rs`). Reverted rather
than swept in, to keep the diff readable.

**Final A100 revalidation (head `13c37a7f`):** full CUDA-memory GPU
suite passed; CUDA EP default-parallel lib suite passed (**488 passed /
17 ignored**); real fp16 GQA covered three requests × two interleaved
decode steps with two shared peers and one transactional private
fallback, byte-identical to independent GPU and CPU references. A
device-started/event-gated worker remained healthy through a peer
process `CUDA_ERROR_ILLEGAL_ADDRESS`, owner exit, replacement worker,
and process restart. Governor tests, targeted Miri, CI Clippy, fmt, and
CUDA honesty also passed.

## Still open after this merges

- `memory-plugin-provider-wiring` — the back half of Phase 6 criterion 8
(provider/context pinning), which fell in the gap between phases.
**#1186 must not be closed until it lands.** This merge makes it more
tractable, not less: see decision 3 in the guide.
- `memory-deferred-invariant-asserts` — two surviving mutants found by
#1462's reviewer, adjudicated NON-BLOCKING and not defects.
- #1533's CUDA honesty guard warns on a legitimately portable anchor.
Should be moved out of the CUDA test binary or allow-listed rather than
left warning.

---------

Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <copilot@github.com>
Copilot-Session: 46c5d75b-8146-489c-b82f-08ee29c27ce4
Copilot-Session: 39ff6824-d35f-4d3b-8f5f-043a7119a100
Copilot-Session: c80f8522-983c-47f7-8241-2155a823aabe
@justinchuby

Copy link
Copy Markdown
Owner Author

Closing: superseded by #1579, which collapsed phases 1–7 into a single merge — landed as a36964280 (squash, 126 files, +43956 −3286).

This PR was rejected under the lockout rule and superseded by the approved revision that was stacked on top of it (not by a replacement of it). Per #1579's own guidance, it was never meant to be reviewed individually.

Containment verified, not assumed:

$ git merge-base --is-ancestor ec75da9db ebceab071   # ebceab071 = #1579 head
→ ancestor

Every one of the 15 stack heads (#1252 → #1533) is a literal ancestor of the merged head, and I separately confirmed the merged head's content reached main: of the 126 files #1579 touched, exactly one differs from main — crates/onnx-runtime-session/src/executor/tests.rs, where main carries 120 extra lines from #1703 (Expand-broadcast mask tests). That is the conflict resolution correctly preserving the other side, not content loss.

Why this didn't close itself: #1579's body says Closes the stack: #1252 #1263 …, which GitHub does not parse — the closing keyword must be immediately followed by the reference, and in any case closing keywords only ever close issues, never pull requests. So all 15 stayed open by mechanism, not by intent.

#1186 stays open: #1579 explicitly gates it on memory-plugin-provider-wiring (the back half of Phase 6 criterion 8, provider/context pinning), which has not landed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

timemachine Major removal preserved as an architectural time-machine reference

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant