Skip to content

fix(memory): settle prepared releases on CUDA device loss - #1349

Closed
justinchuby wants to merge 1 commit into
justinchuby-1186-memory-deferred-release-phase-4from
justinchuby-fix-phase4-device-loss-settlement
Closed

justinchuby wants to merge 1 commit into
justinchuby-1186-memory-deferred-release-phase-4from
justinchuby-fix-phase4-device-loss-settlement

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Corrective PR for the confirmed high defect in #1341. Targets justinchuby-1186-memory-deferred-release-phase-4 at 4d2b2cc5. Part of #1186.

Do not merge this or any memory PR automatically.

The defect

CudaDeferredReleaseQueue has three device-loss retention sites — already-lost enqueue/poll, the concurrent-loss carry path, and retain_all_pending. All three moved the whole unexecuted action into RetainedOwnership.keep_alive.

For a PreparedReleaseAction that froze a live PreparedAllocationRelease with request: Some(..). It was never executed, never quarantined, and never dropped, so:

  • the binding never recorded the allocation as retained — no quarantined entry at all;
  • queued_releases and active_operations never settled;
  • confirm_context_terminated failed on ContextNotQuiescent, and remove stayed impossible;
  • a permanent Arc cycle formed: queue.retained → prepared request → MemoryBinding → MechanismEntry → ProviderContextEntry → CudaProviderContextPin → queue.

The queue kept itself, its CUDA context, and its streams alive forever after any device loss with work in flight.

The fix

DeferredReleaseAction gains a consuming settle_device_lost hook.

  • The default retains the whole action, unchanged. That stays correct for actions whose ownership is purely physical and that owe no mechanism anything — WeightPageRelease, ReservationTeardownAction, and the synthetic test actions.
  • PreparedReleaseAction overrides it: consumes its request through the existing device-loss settlement path (no allocator call, no refund, no device touch), then hands back only the pinned Arc<dyn DeviceAllocator> — the same residual the normal quarantine path retains, and the one piece that does not pin the provider context. The observer still sees the real terminal outcome; it refunds only the reported unmapped bytes, which are zero, and counts no free.

All three sites now go through one settle_lost_entry helper, so they cannot drift apart again.

PreparedAllocationRelease::quarantine_device_lost (memory-api) exposes the settlement execute already performs behind a device-lost release gate. It is needed because the queue can learn the context is unusable before the mechanism lifecycle is invalidated, and execute would then still be permitted to call the allocator. It is deliberately distinct from quarantine(QuarantineReason::DeviceLost), which records the generic Quarantined state rather than the DeviceLost terminal state that confirmed context termination discharges.

After settlement

  • binding.quarantined holds exactly one record: the exact AllocationIdentity, state: DeviceLost, reason: DeviceLost, exact bytes and retained bytes.
  • mechanism queued_releases == 0 and active_operations == 0.
  • allocator physical release calls == 0.
  • invalidate_device → confirm_context_terminated → remove → remove_provider_context all succeed, per the documented quarantine discharge.
  • Arc::strong_count(&queue) == 1 after that teardown — the cycle is gone.
  • the residual allocator pin is still held, so physical ownership is not reusable.

Tests

Three new production-path tests in crates/onnx-runtime-ep-cuda/tests/deferred_release_queue.rs. They use real enqueue_prepared with a real MemoryBinding over a host-backed allocator and a provider-context resource shaped exactly like CudaProviderContextPin, not CountingRelease or a toy action:

test site covered
device_loss_settles_a_pending_prepared_release_and_frees_the_context retain_all_pending (the reported pending path)
device_loss_settles_a_prepared_release_a_poller_was_holding concurrent-loss carry, deterministic via a blocking fence
concurrent_loss_settles_every_prepared_release_whichever_path_takes_it 32 real requests, 2 pollers racing loss; covers the already-lost-in-loop site

The third site cannot be forced deterministically without adding a test hook to production code: poll takes the execution gate before its per-entry device-loss check, and mark_device_lost needs that same gate, so any in-loop trigger would deadlock. It is covered by the racing test, which asserts the same invariants for every request regardless of which site settled it, and all three sites share one helper.

Controlled revert: restoring the pre-fix behaviour (override replaced by the default) makes all three new tests fail on the binding records exactly one retained allocation: left 0, right 1. With the fix, all 20 tests in the file pass.

Validation

check result
cargo test -p onnx-runtime-ep-cuda 384 lib + all integration, 0 failed
cargo test --test deferred_release_queue 20 passed (17 before)
cargo test -p onnx-runtime-memory-api -p onnx-runtime-memory-governor 123 passed, 0 failed
flake stress 40× queue suite, 15× memory-api, 0 failures; release mode green
cargo fmt (scoped) clean
clippy, changed files 0 lints in deferred_release.rs / deferred_release_queue.rs
rustdoc -D warnings, changed files 0 new issues

Failures compared against the exact #1341 base 4d2b2cc5:

  • cargo clippy -p onnx-runtime-ep-cuda --features cuda -- -D warnings fails at base and at head with the same two errors in the untouched onnx-runtime-ep-cpu (manual RangeInclusive::contains, collapsible_if).
  • RUSTDOCFLAGS=-D warnings cargo doc fails at base and at head on untouched items (shareability.rs KvLayout, weight_paging.rs, cudnn, and the deferred_release.rs line-29 module header this PR does not touch).
  • verify_cuda_test_honesty.py fails identically at base and head for the same three portable files; only deferred_release_queue's count changes 17 → 20.

No new failure class is introduced.

Scope

deferred_release.rs (CUDA), deferred_release_queue.rs (tests), and one additive method in memory-api deferred.rs. No rearchitecture, no Phase-5 work, no change to the CUDA partial-release state machine or its accounting, no unrelated cleanup. cargo fmt --all also wanted to reformat onnx-genai-server/src/routes/completions.rs; that pre-existing deviation was reverted out of this diff.

Head 9ac2b41b, 3 files changed, +502 / −68.

The CUDA deferred release queue's three device-loss retention sites moved
the whole unexecuted action into `RetainedOwnership.keep_alive`. For a
`PreparedReleaseAction` that froze a live `PreparedAllocationRelease`: its
request was never executed, never quarantined, and never dropped, so the
binding never recorded the allocation as retained, `queued_releases` and
`active_operations` never settled, `confirm_context_terminated` and
`remove` stayed impossible, and a permanent cycle formed — queue ->
retained record -> request -> binding -> mechanism -> provider context ->
context pin -> queue.

`DeferredReleaseAction` gains a consuming `settle_device_lost` hook. The
default keeps retaining the whole action, which is right for an action
whose ownership is purely physical (a weight page's allocator/allowance, a
reservation ticket). `PreparedReleaseAction` overrides it: it consumes its
request through the existing device-loss settlement path — no allocator
call, no refund — and hands back only the pinned allocator, the same
residual the normal quarantine path retains and the one piece that does
not pin the provider context. All three sites (already-lost enqueue/poll,
concurrent-loss carry, retain_all_pending) now share one helper.

`PreparedAllocationRelease::quarantine_device_lost` exposes the settlement
`execute` already performs behind a device-lost release gate. It is needed
because the queue can learn the context is unusable before the mechanism
lifecycle is invalidated, and `execute` would then still be permitted to
call the allocator.

The new portable tests use the production path — real `enqueue_prepared`
with a host-backed `MemoryBinding` and a provider-context pin shaped like
the CUDA one — and assert zero allocator releases, device-lost quarantine
against the exact allocation identity, zero queued releases and active
operations, and that the queue's strong count returns to one after the
documented context teardown.

Refs #1341
Part of #1186

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Final independent review: approved; #1341 device-loss rejection resolved. Verified all three loss sites consume through one settlement hook; prepared releases record exact DeviceLost quarantine without allocator/device calls; queued/active pins settle; observer refunds zero; residual allocator ownership breaks the queue→binding→context cycle while retaining physical safety; races/outstanding/fence retention remain balanced; and production enqueue_prepared tests are non-vacuous. Phase 4 remains open/unmerged for human review. Real GPU/Xid/TDR execution remains a residual gate.

@codecov

codecov Bot commented Aug 19, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 0% with 9 lines in your changes missing coverage. Please review.
✅ Project coverage is 79.52%. Comparing base (4d2b2cc) to head (9ac2b41).

Files with missing lines Patch % Lines
crates/onnx-runtime-memory-api/src/deferred.rs 0.00% 9 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@                                 Coverage Diff                                  @@
##           justinchuby-1186-memory-deferred-release-phase-4    #1349      +/-   ##
====================================================================================
- Coverage                                             79.53%   79.52%   -0.01%     
====================================================================================
  Files                                                   365      365              
  Lines                                                162318   162327       +9     
  Branches                                             162318   162327       +9     
====================================================================================
- Hits                                                 129100   129096       -4     
- Misses                                                28501    28513      +12     
- Partials                                               4717     4718       +1     
Flag Coverage Δ
mlas 85.05% <ø> (-0.14%) ⬇️
offline 79.42% <0.00%> (-0.01%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-runtime-memory-api/src/deferred.rs 70.40% <0.00%> (-1.87%) ⬇️

... and 1 file with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 qwen3_sampling_processors/top_k_partial_selection 143.53 µs 189.61 µs +32.1%
⚠️ logit_processing/seven_processor_chain_per_step 328.32 µs 424.61 µs +29.3%
⚠️ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.63 ms 4.68 ms +28.8%
⚠️ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.81 ms 6.83 ms +17.6%
⚠️ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 74.36 µs 86.94 µs +16.9%
⚠️ grammar_masking/llguidance_compute_mask/32 77.55 µs 90.45 µs +16.6%
⚠️ kv_cache/alloc_dealloc_pages 38.83 µs 45.16 µs +16.3%
⚠️ qwen3_sampling_processors/top_k_top_p_fast 675.12 µs 780.17 µs +15.6%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 4.03 ms 4.57 ms +13.3%
✅ gather/large_f32_threads=1-internal/131072 31.21 µs 35.03 µs +12.2%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.29 ms 2.53 ms +10.6%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 1.55 ms 1.70 ms +9.8%
✅ sampling_latency/min_p_per_token 213.23 µs 227.10 µs +6.5%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 559.72 µs 592.19 µs +5.8%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.09 ms 2.18 ms +4.4%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.51 ms 2.57 ms +2.4%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 418.70 µs 428.22 µs +2.3%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 82.81 µs 84.62 µs +2.2%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.69 ms 9.90 ms +2.1%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 438.70 µs 447.57 µs +2.0%
✅ gather/large_f16_threads=1-internal/131072 13.96 µs 14.15 µs +1.3%
✅ sampling_latency/top_p_per_token 389.79 µs 393.64 µs +1.0%
✅ matmul/medium_generic_f16_threads=8/32x512x512 35.54 µs 35.86 µs +0.9%
✅ matmul/medium_generic_f16_threads=1/32x512x512 33.34 µs 33.57 µs +0.7%
✅ sampling_latency/greedy_per_token 3.30 µs 3.30 µs -0.0%
✅ tokenization/encode_tokens_per_second 395.33 µs 394.69 µs -0.2%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 92.05 µs 91.77 µs -0.3%
✅ sampling_latency/top_k_per_token 54.37 µs 53.92 µs -0.8%
✅ tokenization/decode_tokens_per_second 6.55 ms 6.35 ms -2.9%
✅ gather/large_bf16_threads=1-internal/131072 14.23 µs 13.80 µs -3.0%
✅ matmul/small_generic_bf16_threads=8/1x256x256 37.01 µs 35.10 µs -5.2%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 583.43 µs 552.50 µs -5.3%
✅ matmul/small_generic_bf16_threads=1/1x256x256 37.25 µs 34.78 µs -6.6%
✅ gather/medium_bf16_threads=1-internal/32768 2.83 µs 2.58 µs -8.7%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.26 ms 1.13 ms -10.7%
✅ matmul/small_generic_f32_threads=1/1x256x256 45.11 µs 39.99 µs -11.4%
✅ gather/small_f32_threads=1-internal/4096 803.2 ns 685.0 ns -14.7%
✅ matmul/small_generic_f16_threads=8/1x256x256 39.48 µs 33.65 µs -14.8%
🟢 gather/small_f16_threads=1-internal/4096 574.4 ns 488.1 ns -15.0%
🟢 add/small_bf16_threads=1-internal/1024 664.8 ns 547.9 ns -17.6%
🟢 gather/small_bf16_threads=1-internal/4096 594.0 ns 489.3 ns -17.6%
🟢 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 75.65 µs 61.10 µs -19.2%
🟢 add/small_f16_threads=1-internal/1024 573.5 ns 458.6 ns -20.0%
🟢 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 79.10 µs 61.89 µs -21.8%
🟢 reduce_mean/large_f32_threads=1-internal/262144 1.29 ms 998.44 µs -22.6%
🟢 matmul/small_generic_f16_threads=1/1x256x256 45.65 µs 35.03 µs -23.3%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 327.03 µs 249.79 µs -23.6%
🟢 reduce_mean/small_f32_threads=1-internal/4096 20.46 µs 15.32 µs -25.1%
🟢 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 883.06 µs 627.82 µs -28.9%
🟢 gather/medium_f16_threads=1-internal/32768 3.59 µs 2.52 µs -29.8%
🟢 add/medium_f32_threads=1-internal/262144 37.16 µs 25.74 µs -30.7%
🟢 add/large_bf16_threads=1-internal/4194304 2.49 ms 1.68 ms -32.6%
🟢 add/large_f16_threads=1-internal/4194304 2.53 ms 1.69 ms -33.5%
🟢 gather/medium_f32_threads=1-internal/32768 6.28 µs 4.16 µs -33.7%
🟢 add/large_f32_threads=1-internal/4194304 1.04 ms 688.27 µs -34.0%
🟢 add/medium_f16_threads=1-internal/262144 169.31 µs 106.68 µs -37.0%
🟢 add/small_f32_threads=1-internal/1024 316.5 ns 197.9 ns -37.5%
🟢 add/medium_bf16_threads=1-internal/262144 188.48 µs 109.70 µs -41.8%
🟢 matmul/small_generic_f32_threads=8/1x256x256 94.98 µs 39.10 µs -58.8%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 5.08 3.95 6.76 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Aug 22, 2026
> **A100 safety update (head `13c37a7f`):** expected shared-prefix
admission failures are now all-or-private: K/V commit is transactional,
a failed V share rolls K back before enqueue, and only a successful
rollback permits private-KV fallback. Fatal kernel faults remain
process-fatal; the new cross-process VMM probe proves a worker fault
does not affect an actively executing peer, while the current server is
still in-process and does not yet consume physical shared-prefix
metadata. See the [final verified
conclusion](#1579 (comment)).

Collapses phases 1–7 of the #1186 memory architecture rework onto
current `main` as a single merge. The stack forked 270 commits ago (176
of them touching `crates/`), so rebasing it layer by layer would mean
solving the same 17 conflicts fourteen times, with no reviewer ever
looking at the intermediate states.

Closes the stack: #1252 #1263 #1279 #1283 #1301 #1341 #1349 #1426 #1440
#1448 #1454 #1462 #1465 #1468 #1533.

**A review guide is in the first comment.** It is the part worth reading
— this diff is 103 files, but only three decisions in it are ones a
compiler cannot check.

## What was already reviewed, and what wasn't

Every phase was reviewed and approved by a session that was not its
author, under the rejection-lockout rule (a rejected author never writes
the next revision).

| Phase | PR | Verdict |
|---|---|---|
| 1–5 | #1252 → #1426 | approved |
| 6 | #1440, #1448, #1454 | **rejected** ×3 |
| 6 | #1462 | approved |
| 7 | #1465 | **rejected** |
| 7 | #1468 | approved |
| 7 test fixes | #1533 | approved, A100-verified |

The four rejected rounds were not replaced — the later rounds are
stacked **on top of** them. So the tree here is the approved state, but
the history contains the rejected commits. **#1440 / #1448 / #1454 /
#1465 should not be reviewed individually**; they close automatically.

**Not reviewed anywhere:** the conflict resolutions themselves. That is
what this PR is for.

## Verification

Run here, on this merged tree:

- `cargo check --workspace --all-targets` — clean
- `cargo check -p onnx-genai-engine --features cuda,native-backend
--all-targets` — clean. Worth calling out: the default feature set does
**not** compile the `cfg(cuda)` code, which is where the riskiest edits
are. Checking only the default set would have missed a real break (see
the guide).
- `cargo test` over 7 crates — **1105 passed, 2 failed, 88 ignored**
- `cargo clippy --workspace --all-targets`
- `cargo fmt`

**The 2 failures and the 1 clippy error are pre-existing and proven so,
not merge damage:**

-
`platform_capacity::{disk_capacity_is_measured_for_the_working_directory,
an_explicit_byte_limit_is_honored_without_a_device_query}` —
`platform_capacity.rs` is byte-identical to `main` (`git diff
origin/main HEAD -- ` that file is empty). The cause is a macOS-only FFI
layout bug: `fsblkcnt_t` is 4 bytes on macOS (confirmed: `sizeof(struct
statvfs)=64`, `sizeof(fsblkcnt_t)=4`) while the Rust struct declares
`f_blocks`/`f_bfree`/`f_bavail` as `u64`. CI is Linux, where it is 8
bytes. Left alone: it is a real bug but not this PR's.
- `optimizer.rs` clippy `approx_constant` — present on both parents; the
merged file is byte-identical to `main`.
- Three files had `rustfmt` drift already present on `main`
(`matmul_nbits.rs`, `normalization.rs`, `optimizer.rs`). Reverted rather
than swept in, to keep the diff readable.

**Final A100 revalidation (head `13c37a7f`):** full CUDA-memory GPU
suite passed; CUDA EP default-parallel lib suite passed (**488 passed /
17 ignored**); real fp16 GQA covered three requests × two interleaved
decode steps with two shared peers and one transactional private
fallback, byte-identical to independent GPU and CPU references. A
device-started/event-gated worker remained healthy through a peer
process `CUDA_ERROR_ILLEGAL_ADDRESS`, owner exit, replacement worker,
and process restart. Governor tests, targeted Miri, CI Clippy, fmt, and
CUDA honesty also passed.

## Still open after this merges

- `memory-plugin-provider-wiring` — the back half of Phase 6 criterion 8
(provider/context pinning), which fell in the gap between phases.
**#1186 must not be closed until it lands.** This merge makes it more
tractable, not less: see decision 3 in the guide.
- `memory-deferred-invariant-asserts` — two surviving mutants found by
#1462's reviewer, adjudicated NON-BLOCKING and not defects.
- #1533's CUDA honesty guard warns on a legitimately portable anchor.
Should be moved out of the CUDA test binary or allow-listed rather than
left warning.

---------

Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <copilot@github.com>
Copilot-Session: 46c5d75b-8146-489c-b82f-08ee29c27ce4
Copilot-Session: 39ff6824-d35f-4d3b-8f5f-043a7119a100
Copilot-Session: c80f8522-983c-47f7-8241-2155a823aabe
@justinchuby

Copy link
Copy Markdown
Owner Author

Closing: superseded by #1579, which collapsed phases 1–7 into a single merge — landed as a36964280 (squash, 126 files, +43956 −3286).

This PR was approved by a session that was not its author, and its work is in main.

Containment verified, not assumed:

$ git merge-base --is-ancestor 9ac2b41bd ebceab071   # ebceab071 = #1579 head
→ ancestor

Every one of the 15 stack heads (#1252 → #1533) is a literal ancestor of the merged head, and I separately confirmed the merged head's content reached main: of the 126 files #1579 touched, exactly one differs from main — crates/onnx-runtime-session/src/executor/tests.rs, where main carries 120 extra lines from #1703 (Expand-broadcast mask tests). That is the conflict resolution correctly preserving the other side, not content loss.

Why this didn't close itself: #1579's body says Closes the stack: #1252 #1263 …, which GitHub does not parse — the closing keyword must be immediately followed by the reference, and in any case closing keywords only ever close issues, never pull requests. So all 15 stayed open by mechanism, not by intent.

#1186 stays open: #1579 explicitly gates it on memory-plugin-provider-wiring (the back half of Phase 6 criterion 8, provider/context pinning), which has not landed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant