Skip to content

feat(memory): add registry-issued provider bindings - #1279

Closed
justinchuby wants to merge 1 commit into
justinchuby-1186-memory-capabilities-phase2from
justinchuby-1186-memory-binding-phase3
Closed

justinchuby wants to merge 1 commit into
justinchuby-1186-memory-capabilities-phase2from
justinchuby-1186-memory-binding-phase3

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Summary

  • add a narrow BindingRegistry that issues provider-context, authority, mechanism, binding, and allocation identities instead of trusting allocator-reported pointers, TypeId, or tokens
  • pin the selected Arc<dyn DeviceAllocator> plus opaque provider-context and authority resources in every binding/allocation/view/capability handle
  • preserve the existing allocator/capability and with_memory(Arc<dyn DeviceAllocator>) paths as incremental adapters; no memory policy or physical-release behavior changes

Identity source and lifetime graph

The registry is the sole source of ProviderContextIdentity, AuthorityIdentity, MechanismIdentity, BindingId/BindingGeneration, and AllocationGeneration. Each MemoryBinding points to one registered mechanism entry, which owns the allocator and the provider-context/authority lifetime pins. Bound allocation, view, virtual-backing, and shared-mapping metadata clone that same binding.

AllocationGeneration is process-local, opaque, never pointer-derived, and checked together with the full binding identity and allocation metadata. Reusing the same VA therefore cannot validate stale metadata from an earlier allocation.

Switch and invalidation semantics

Changing the selected mechanism affects only future bind(device) calls. Existing bindings remain pinned to their original mechanism, and explicit whole-allocation release still goes through that original DeviceAllocator. Retiring a valid mechanism rejects new work while allowing existing allocations to release.

Device loss removes selection and terminally invalidates affected binding operations, including explicit release. That path performs no allocator callback, physical free, lease release, or delegated-quota refund. After externally observed provider-context/process termination and callback quiescence, confirm_context_terminated retires allocation identity metadata; authority/accounting reconciliation remains outside this registry.

Cross-binding, cross-mechanism, cross-authority, and cross-device metadata is rejected before capability/device callbacks. A split-inner transparent bundle must use the explicit unsafe register_trusted_composite trust boundary; the API does not claim Rust proves hostile compositions coherent.

Lock order and teardown

The registry has two non-nested lock classes:

  1. registry state for registration and current selection
  2. per-mechanism lifecycle/allocation identity state

Lookups clone an entry and release the registry lock before lifecycle inspection. Allocator, capability, validated-view, and deallocation callbacks run with neither lock held. Invalidation never waits; termination confirmation returns ContextNotQuiescent so waiting/polling stays outside registry and governance locks. Provider-context and authority pins can be removed only after their mechanism registrations are quiescent and removed, while outstanding bound handles retain their own Arc pins.

Migration and Phase 3 boundary

The new binding layer lives in onnx-runtime-memory-api and is re-exported from the governor compatibility surface. Existing CPU, CUDA, VMM/shared-prefix, ORT governed allocator, canonical DeviceAllocator release, capability discovery, and erased third-party Arc<dyn DeviceAllocator> behavior remain unchanged.

This PR adds no allocation Drop free, deferred-free queue, event/fence scheduling, synchronization, physical-release completeness claim, partial-unmap recovery, quarantine, residual-handle state machine, pointer-only retry API, or ProcessMemoryManager policy/transaction centralization.

Validation

  • 504 targeted tests passed across memory API/governor (86), CPU allocator compatibility (8), CUDA EP library (382), CUDA memory adapters (8), and ORT governed allocator (20); 22 CUDA hardware tests were ignored by their existing gates
  • concurrency stress passed 20/20 runs, covering 4,800 register/lookup/switch/invalidate/teardown iterations; no loom dependency/infrastructure exists in the workspace
  • scoped cargo clippy -p onnx-runtime-memory-api -p onnx-runtime-memory-governor --all-targets -- -D warnings passed
  • strict RUSTDOCFLAGS='-D warnings' cargo doc -p onnx-runtime-memory-api --no-deps passed
  • git diff --check passed and an independent high-confidence code review found no significant issues

Exact Phase 2 baseline comparison at d120140a:

  • workspace cargo fmt --all -- --check reports only the same pre-existing onnx-genai-server/src/routes/completions.rs formatting diffs at lines 2891 and 2914
  • workspace all-target check reports the same three pre-existing onnx-genai-bench/tests/fused_batch_prefill.rs feature/import errors
  • workspace all-target tests stop on the same pre-existing arm64 MLAS undefined-symbol linker failures
  • strict governor rustdoc reports the same pre-existing private BUDGET_SHARE_DENOMINATOR and unresolved KvLayout links; the changed memory API rustdoc is clean

Part of #1186 — Phase 3 only.

Introduce narrow provider/context, authority, mechanism, binding, and allocation identities with pinned resources, stable allocator-switch behavior, deterministic invalidation, and explicit bound capability adapters.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 18, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 61.60627% with 392 lines in your changes missing coverage. Please review.
✅ Project coverage is 79.87%. Comparing base (d120140) to head (31a8f28).

Files with missing lines Patch % Lines
crates/onnx-runtime-memory-api/src/binding.rs 61.60% 313 Missing and 79 partials ⚠️
Additional details and impacted files

Impacted file tree graph

@@                               Coverage Diff                               @@
##           justinchuby-1186-memory-capabilities-phase2    #1279      +/-   ##
===============================================================================
- Coverage                                        79.99%   79.87%   -0.12%     
===============================================================================
  Files                                              362      363       +1     
  Lines                                           158583   159604    +1021     
  Branches                                        158583   159604    +1021     
===============================================================================
+ Hits                                            126856   127483     +627     
- Misses                                           27111    27426     +315     
- Partials                                          4616     4695      +79     
Flag Coverage Δ
mlas 85.05% <ø> (-0.14%) ⬇️
offline 79.77% <61.60%> (-0.12%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-runtime-memory-api/src/lib.rs 100.00% <ø> (ø)
crates/onnx-runtime-memory-governor/src/lib.rs 83.26% <ø> (ø)
crates/onnx-runtime-memory-api/src/binding.rs 61.60% <61.60%> (ø)

... and 3 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 54.08 µs 127.91 µs +136.5%
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 543.82 µs 1.17 ms +114.9%
🔴 matmul/large_generic_f32_threads=8/32x1024x1024 3.78 ms 6.86 ms +81.4%
🔴 matmul/medium_generic_bf16_threads=8/32x512x512 384.75 µs 657.28 µs +70.8%
🔴 matmul/medium_generic_f32_threads=8/32x512x512 946.76 µs 1.52 ms +60.7%
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 49.32 µs 73.78 µs +49.6%
🔴 matmul/medium_generic_f16_threads=8/32x512x512 31.00 µs 45.20 µs +45.8%
🔴 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 393.52 µs 540.73 µs +37.4%
🔴 matmul/small_generic_f32_threads=8/1x256x256 35.14 µs 46.49 µs +32.3%
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.76 ms 2.32 ms +31.6%
⚠️ qwen3_sampling_processors/top_p_fast_after_top_k 506.51 µs 655.64 µs +29.4%
⚠️ matmul/large_generic_f16_threads=8/32x1024x1024 82.77 µs 105.95 µs +28.0%
⚠️ matmul/small_generic_f16_threads=1/1x256x256 32.74 µs 41.78 µs +27.6%
⚠️ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.68 ms 7.21 ms +26.9%
⚠️ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.46 ms 4.29 ms +24.1%
⚠️ add/small_bf16_threads=1-internal/1024 430.4 ns 529.3 ns +23.0%
⚠️ logit_processing/seven_processor_chain_per_step 312.52 µs 381.88 µs +22.2%
⚠️ matmul/small_generic_f16_threads=8/1x256x256 30.90 µs 36.85 µs +19.2%
⚠️ grammar_masking/llguidance_compute_mask/32 74.65 µs 88.76 µs +18.9%
⚠️ add/small_f16_threads=1-internal/1024 458.4 ns 543.6 ns +18.6%
⚠️ matmul/medium_generic_f16_threads=1/32x512x512 33.80 µs 39.59 µs +17.1%
⚠️ add/medium_f32_threads=1-internal/262144 24.15 µs 28.24 µs +16.9%
⚠️ kv_cache/alloc_dealloc_pages 39.64 µs 46.17 µs +16.5%
⚠️ qwen3_sampling_processors/top_k_top_p_fast 649.48 µs 750.06 µs +15.5%
✅ add/medium_f16_threads=1-internal/262144 107.38 µs 121.01 µs +12.7%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.10 ms 2.34 ms +11.6%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 534.57 µs 594.67 µs +11.2%
✅ qwen3_sampling_processors/top_k_partial_selection 140.20 µs 155.38 µs +10.8%
✅ gather/large_f16_threads=1-internal/131072 16.39 µs 18.00 µs +9.8%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.44 ms 10.23 ms +8.4%
✅ matmul/small_generic_bf16_threads=1/1x256x256 31.28 µs 33.73 µs +7.8%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 80.22 µs 86.33 µs +7.6%
✅ tokenization/encode_tokens_per_second 369.48 µs 394.73 µs +6.8%
✅ sampling_latency/min_p_per_token 204.32 µs 217.81 µs +6.6%
✅ reduce_mean/small_f32_threads=1-internal/4096 16.26 µs 17.27 µs +6.2%
✅ matmul/small_generic_bf16_threads=8/1x256x256 34.61 µs 36.76 µs +6.2%
✅ sampling_latency/top_p_per_token 380.41 µs 402.61 µs +5.8%
✅ sampling_latency/top_k_per_token 50.77 µs 52.71 µs +3.8%
✅ tokenization/decode_tokens_per_second 6.29 ms 6.49 ms +3.2%
✅ add/medium_bf16_threads=1-internal/262144 101.14 µs 104.29 µs +3.1%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.36 ms 2.43 ms +2.9%
✅ sampling_latency/greedy_per_token 3.27 µs 3.27 µs +0.1%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.14 ms 2.14 ms +0.0%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 42.79 µs 42.69 µs -0.2%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.04 ms 1.02 ms -1.6%
✅ reduce_mean/medium_f32_threads=1-internal/65536 267.25 µs 260.15 µs -2.7%
✅ gather/large_bf16_threads=1-internal/131072 16.35 µs 15.46 µs -5.4%
✅ gather/small_f32_threads=1-internal/4096 677.5 ns 628.7 ns -7.2%
✅ add/large_bf16_threads=1-internal/4194304 2.02 ms 1.81 ms -10.6%
✅ gather/small_f16_threads=1-internal/4096 552.1 ns 492.2 ns -10.8%
✅ gather/small_bf16_threads=1-internal/4096 525.5 ns 467.8 ns -11.0%
✅ add/small_f32_threads=1-internal/1024 274.2 ns 241.4 ns -12.0%
✅ gather/medium_f32_threads=1-internal/32768 4.19 µs 3.60 µs -14.0%
🟢 matmul/small_generic_f32_threads=1/1x256x256 42.07 µs 35.76 µs -15.0%
🟢 gather/medium_bf16_threads=1-internal/32768 2.70 µs 2.26 µs -16.5%
🟢 gather/medium_f16_threads=1-internal/32768 2.68 µs 2.24 µs -16.7%
🟢 add/large_f32_threads=1-internal/4194304 876.45 µs 712.90 µs -18.7%
🟢 gather/large_f32_threads=1-internal/131072 50.77 µs 38.96 µs -23.2%
🟢 add/large_f16_threads=1-internal/4194304 2.58 ms 1.66 ms -35.4%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 4.19 3.98 5.67 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Aug 22, 2026
> **A100 safety update (head `13c37a7f`):** expected shared-prefix
admission failures are now all-or-private: K/V commit is transactional,
a failed V share rolls K back before enqueue, and only a successful
rollback permits private-KV fallback. Fatal kernel faults remain
process-fatal; the new cross-process VMM probe proves a worker fault
does not affect an actively executing peer, while the current server is
still in-process and does not yet consume physical shared-prefix
metadata. See the [final verified
conclusion](#1579 (comment)).

Collapses phases 1–7 of the #1186 memory architecture rework onto
current `main` as a single merge. The stack forked 270 commits ago (176
of them touching `crates/`), so rebasing it layer by layer would mean
solving the same 17 conflicts fourteen times, with no reviewer ever
looking at the intermediate states.

Closes the stack: #1252 #1263 #1279 #1283 #1301 #1341 #1349 #1426 #1440
#1448 #1454 #1462 #1465 #1468 #1533.

**A review guide is in the first comment.** It is the part worth reading
— this diff is 103 files, but only three decisions in it are ones a
compiler cannot check.

## What was already reviewed, and what wasn't

Every phase was reviewed and approved by a session that was not its
author, under the rejection-lockout rule (a rejected author never writes
the next revision).

| Phase | PR | Verdict |
|---|---|---|
| 1–5 | #1252 → #1426 | approved |
| 6 | #1440, #1448, #1454 | **rejected** ×3 |
| 6 | #1462 | approved |
| 7 | #1465 | **rejected** |
| 7 | #1468 | approved |
| 7 test fixes | #1533 | approved, A100-verified |

The four rejected rounds were not replaced — the later rounds are
stacked **on top of** them. So the tree here is the approved state, but
the history contains the rejected commits. **#1440 / #1448 / #1454 /
#1465 should not be reviewed individually**; they close automatically.

**Not reviewed anywhere:** the conflict resolutions themselves. That is
what this PR is for.

## Verification

Run here, on this merged tree:

- `cargo check --workspace --all-targets` — clean
- `cargo check -p onnx-genai-engine --features cuda,native-backend
--all-targets` — clean. Worth calling out: the default feature set does
**not** compile the `cfg(cuda)` code, which is where the riskiest edits
are. Checking only the default set would have missed a real break (see
the guide).
- `cargo test` over 7 crates — **1105 passed, 2 failed, 88 ignored**
- `cargo clippy --workspace --all-targets`
- `cargo fmt`

**The 2 failures and the 1 clippy error are pre-existing and proven so,
not merge damage:**

-
`platform_capacity::{disk_capacity_is_measured_for_the_working_directory,
an_explicit_byte_limit_is_honored_without_a_device_query}` —
`platform_capacity.rs` is byte-identical to `main` (`git diff
origin/main HEAD -- ` that file is empty). The cause is a macOS-only FFI
layout bug: `fsblkcnt_t` is 4 bytes on macOS (confirmed: `sizeof(struct
statvfs)=64`, `sizeof(fsblkcnt_t)=4`) while the Rust struct declares
`f_blocks`/`f_bfree`/`f_bavail` as `u64`. CI is Linux, where it is 8
bytes. Left alone: it is a real bug but not this PR's.
- `optimizer.rs` clippy `approx_constant` — present on both parents; the
merged file is byte-identical to `main`.
- Three files had `rustfmt` drift already present on `main`
(`matmul_nbits.rs`, `normalization.rs`, `optimizer.rs`). Reverted rather
than swept in, to keep the diff readable.

**Final A100 revalidation (head `13c37a7f`):** full CUDA-memory GPU
suite passed; CUDA EP default-parallel lib suite passed (**488 passed /
17 ignored**); real fp16 GQA covered three requests × two interleaved
decode steps with two shared peers and one transactional private
fallback, byte-identical to independent GPU and CPU references. A
device-started/event-gated worker remained healthy through a peer
process `CUDA_ERROR_ILLEGAL_ADDRESS`, owner exit, replacement worker,
and process restart. Governor tests, targeted Miri, CI Clippy, fmt, and
CUDA honesty also passed.

## Still open after this merges

- `memory-plugin-provider-wiring` — the back half of Phase 6 criterion 8
(provider/context pinning), which fell in the gap between phases.
**#1186 must not be closed until it lands.** This merge makes it more
tractable, not less: see decision 3 in the guide.
- `memory-deferred-invariant-asserts` — two surviving mutants found by
#1462's reviewer, adjudicated NON-BLOCKING and not defects.
- #1533's CUDA honesty guard warns on a legitimately portable anchor.
Should be moved out of the CUDA test binary or allow-listed rather than
left warning.

---------

Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <copilot@github.com>
Copilot-Session: 46c5d75b-8146-489c-b82f-08ee29c27ce4
Copilot-Session: 39ff6824-d35f-4d3b-8f5f-043a7119a100
Copilot-Session: c80f8522-983c-47f7-8241-2155a823aabe
@justinchuby

Copy link
Copy Markdown
Owner Author

Closing: superseded by #1579, which collapsed phases 1–7 into a single merge — landed as a36964280 (squash, 126 files, +43956 −3286).

This PR was approved by a session that was not its author, and its work is in main.

Containment verified, not assumed:

$ git merge-base --is-ancestor 31a8f288c ebceab071   # ebceab071 = #1579 head
→ ancestor

Every one of the 15 stack heads (#1252 → #1533) is a literal ancestor of the merged head, and I separately confirmed the merged head's content reached main: of the 126 files #1579 touched, exactly one differs from main — crates/onnx-runtime-session/src/executor/tests.rs, where main carries 120 extra lines from #1703 (Expand-broadcast mask tests). That is the conflict resolution correctly preserving the other side, not content loss.

Why this didn't close itself: #1579's body says Closes the stack: #1252 #1263 …, which GitHub does not parse — the closing keyword must be immediately followed by the reference, and in any case closing keywords only ever close issues, never pull requests. So all 15 stayed open by mechanism, not by intent.

#1186 stays open: #1579 explicitly gates it on memory-plugin-provider-wiring (the back half of Phase 6 criterion 8, provider/context pinning), which has not landed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant