Skip to content

feat(memory): extract foundational memory API (Phase 1) - #1252

Closed
justinchuby wants to merge 1 commit into
mainfrom
justinchuby-1186-memory-api-phase1-reland
Closed

justinchuby wants to merge 1 commit into
mainfrom
justinchuby-1186-memory-api-phase1-reland

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Summary

  • Adds dependency-free onnx-runtime-memory-api as the lowest memory mechanism vocabulary layer.
  • Moves only Tier, DeviceKey, AllocationCommitRange, MappedAllocation, SharedDevicePrefix, and SharedPrefixCommitInfo.
  • Preserves the existing onnx-runtime-memory-governor root and allocator-module re-exports, so downstream imports do not need a rename sweep.
  • Updates workspace/default-member/dependency metadata, the dependency-first publish order, CI package lists, and the shipped memory architecture documentation.

Part of #1186, Phase 1 only. This is the safe re-land after #1247; it does not start Phase 2.

Boundary retained in the governor

The following remain in onnx-runtime-memory-governor because they are accounting/governance concerns or are currently coupled to them:

  • DeviceAllocator and HostAllocator: existing signatures use governor-owned MappedPhysicalCapacityToken and MemoryError.
  • MemoryRole and MemoryError: reservation purpose, budget refusal, and governed-capacity outcomes.
  • MemoryAuthorityId, HolderId, ledgers, mapped allowances/tokens, growth authorities/grants, leases, pressure responders, and governor traits.
  • LargeAllocCache and prefix-shareability analysis, which are built over the current governor-owned allocator/admission model.

Dependency direction is one-way: dependency-free onnx-runtime-memory-api <- onnx-runtime-memory-governor <- existing consumers. The API crate has no EP, session, engine, or governor dependency and introduces no cycle.

Behavior/signature preservation

  • The complete DeviceAllocator trait text is byte-identical to current main.
  • CUDA VMM allocator, EP API provider, CUDA provider, and ORT GovernedAllocator source files are unchanged from main.
  • with_memory(Arc<dyn DeviceAllocator>), allocation/release behavior, eager/VMM selection, shared-prefix handling, mapped refunds, synchronization, governance policy, and runtime defaults are unchanged.
  • No capability discovery, mechanism identity, lifecycle/retry/RAII, manager/binding, or Phase 2+ API is introduced.

Validation

  • cargo fmt -p onnx-runtime-memory-api -p onnx-runtime-memory-governor -- --check
  • cargo metadata --locked --no-deps --format-version 1
  • cargo package --locked -p onnx-runtime-memory-api --allow-dirty
  • cargo check --locked --all-targets -p onnx-runtime-memory-api -p onnx-runtime-memory-governor -p onnx-runtime-cuda-memory -p onnx-runtime-ep-api -p onnx-runtime-ep-cuda -p onnx-genai-ort
  • Clippy with -D warnings: all targets for memory-api, memory-governor, CUDA memory, and EP API; library target for ORT
  • Clippy: CUDA EP library (passes with the pre-existing warning noted below)
  • RUSTDOCFLAGS="-D warnings" cargo doc --locked --no-deps -p onnx-runtime-memory-api
  • Rustdoc generation for memory-governor, CUDA memory, EP API, CUDA EP, and ORT
  • Targeted tests: 108 passed, 0 failed, 6 GPU-gated ignored
    • memory-api: 2 passed
    • memory-governor: 64 passed
    • CUDA memory: 8 passed
    • EP API provider: 6 passed
    • CUDA EP provider: 8 passed, 6 ignored
    • ORT governed allocator: 20 passed
  • git diff origin/main...HEAD --check

Pre-existing failures/warnings on current main

  • Workspace-wide cargo fmt --all -- --check reports formatting in crates/onnx-genai-server/src/routes/completions.rs; this branch does not modify that file. The changed Rust packages pass targeted fmt.
  • All-target checks pass but report three existing ORT test/import warnings in unchanged files.
  • Whole affected-graph Clippy with -D warnings reaches existing CPU EP and CUDA EP lints; the changed crates and EP API pass with warnings denied, ORT lib passes with warnings denied, and CUDA EP lib passes with its existing unnecessary_unwrap warning.
  • Strict rustdoc across every affected consumer reaches existing intra-doc-link warnings in unchanged governor, EP API, CUDA EP, and ORT files. The extracted crate passes rustdoc with warnings denied; all affected public docs generate successfully under the repository's current warning policy.

Move only dependency-free memory mechanism primitives into a new onnx-runtime-memory-api crate while preserving the existing governor re-exports, allocator signatures, and runtime behavior. Keep roles, errors, allocator dispatch, capacity accounting, policy, and lifecycle in onnx-runtime-memory-governor.

Part of #1186 Phase 1.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Independent Phase 1 review: approved with no blocking or high-confidence findings.

Verified the full diff is limited to the six stated type moves plus compatibility/workspace/publish/docs wiring; DeviceAllocator/HostAllocator and runtime call sites remain unchanged; old governor root/module paths still resolve; and dependency direction is acyclic (memory-api <- memory-governor <- consumers). Packaging order and strict rustdoc for the new crate are correct.

Residual gate: hosted CUDA compilation remains queued. The re-exported structs/fields are unchanged, so no source-level CUDA compatibility issue was found; the hosted CUDA lanes should still complete before the stack is merged.

@codecov

codecov Bot commented Aug 18, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.62%. Comparing base (5417d04) to head (9e0e20a).
⚠️ Report is 12 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1252      +/-   ##
==========================================
+ Coverage   79.95%   80.62%   +0.67%     
==========================================
  Files         360      360              
  Lines      158518   155718    -2800     
  Branches   158518   155718    -2800     
==========================================
- Hits       126741   125551    -1190     
+ Misses      27160    25560    -1600     
+ Partials     4617     4607      -10     
Flag Coverage Δ
mlas ?
offline 80.62% <100.00%> (+0.77%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-runtime-memory-api/src/allocator.rs 100.00% <100.00%> (ø)
crates/onnx-runtime-memory-api/src/lib.rs 100.00% <100.00%> (ø)
...ates/onnx-runtime-memory-governor/src/allocator.rs 32.59% <ø> (-1.74%) ⬇️
crates/onnx-runtime-memory-governor/src/lib.rs 83.31% <100.00%> (-0.05%) ⬇️

... and 12 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 42.97 µs 84.70 µs +97.1%
🔴 block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 55.46 µs 99.31 µs +79.1%
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 500.83 µs 715.13 µs +42.8%
🔴 matmul/small_generic_bf16_threads=8/1x256x256 34.70 µs 48.42 µs +39.5%
🔴 matmul/small_generic_bf16_threads=1/1x256x256 41.29 µs 54.29 µs +31.5%
⚠️ matmul/large_generic_f16_threads=8/32x1024x1024 92.83 µs 115.78 µs +24.7%
⚠️ matmul/large_generic_bf16_threads=8/32x1024x1024 1.24 ms 1.55 ms +24.6%
⚠️ matmul/large_generic_f16_threads=1/32x1024x1024 86.37 µs 107.47 µs +24.4%
⚠️ matmul/medium_generic_f32_threads=1/32x512x512 2.33 ms 2.88 ms +23.7%
⚠️ matmul/large_generic_f32_threads=8/32x1024x1024 5.49 ms 6.66 ms +21.4%
⚠️ matmul/medium_generic_f16_threads=8/32x512x512 34.94 µs 40.87 µs +17.0%
✅ logit_processing/seven_processor_chain_per_step 298.02 µs 338.74 µs +13.7%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 41.38 µs 46.52 µs +12.4%
✅ kv_cache/alloc_dealloc_pages 35.41 µs 39.62 µs +11.9%
✅ sampling_latency/top_k_per_token 54.44 µs 60.06 µs +10.3%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.84 ms 10.83 ms +10.1%
✅ add/medium_f16_threads=1-internal/262144 107.47 µs 115.08 µs +7.1%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.02 ms 2.16 ms +6.9%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 344.18 µs 365.03 µs +6.1%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 558.09 µs 589.03 µs +5.5%
✅ qwen3_sampling_processors/top_k_top_p_fast 621.63 µs 653.89 µs +5.2%
✅ grammar_masking/llguidance_compute_mask/32 70.33 µs 73.61 µs +4.7%
✅ sampling_latency/greedy_per_token 3.43 µs 3.57 µs +3.8%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.62 ms 3.66 ms +1.2%
✅ tokenization/decode_tokens_per_second 7.07 ms 7.13 ms +0.9%
✅ add/small_f32_threads=1-internal/1024 216.6 ns 218.0 ns +0.7%
✅ reduce_mean/small_f32_threads=1-internal/4096 16.06 µs 16.14 µs +0.5%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.84 ms 5.81 ms -0.6%
✅ add/medium_bf16_threads=1-internal/262144 109.77 µs 107.53 µs -2.0%
✅ sampling_latency/top_p_per_token 422.45 µs 412.20 µs -2.4%
✅ add/large_f16_threads=1-internal/4194304 1.81 ms 1.76 ms -3.0%
✅ gather/medium_bf16_threads=1-internal/32768 2.30 µs 2.23 µs -3.1%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.12 ms 2.06 ms -3.1%
✅ gather/medium_f16_threads=1-internal/32768 2.34 µs 2.27 µs -3.3%
✅ matmul/medium_generic_f16_threads=1/32x512x512 34.89 µs 33.71 µs -3.4%
✅ gather/small_f32_threads=1-internal/4096 645.1 ns 622.2 ns -3.5%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 535.85 µs 514.00 µs -4.1%
✅ gather/small_bf16_threads=1-internal/4096 475.1 ns 451.6 ns -5.0%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 490.61 µs 465.63 µs -5.1%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.30 ms 1.21 ms -7.0%
✅ sampling_latency/min_p_per_token 219.18 µs 203.81 µs -7.0%
✅ gather/small_f16_threads=1-internal/4096 487.8 ns 446.4 ns -8.5%
✅ add/large_bf16_threads=1-internal/4194304 1.81 ms 1.65 ms -8.9%
✅ add/large_f32_threads=1-internal/4194304 719.61 µs 646.26 µs -10.2%
✅ qwen3_sampling_processors/top_k_partial_selection 150.79 µs 135.42 µs -10.2%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.02 ms 913.46 µs -10.5%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 270.61 µs 227.45 µs -15.9%
🟢 gather/medium_f32_threads=1-internal/32768 4.18 µs 3.49 µs -16.6%
🟢 matmul/small_generic_f32_threads=1/1x256x256 41.36 µs 34.45 µs -16.7%
🟢 gather/large_f16_threads=1-internal/131072 13.69 µs 11.36 µs -17.0%
🟢 add/medium_f32_threads=1-internal/262144 32.78 µs 23.80 µs -27.4%
🟢 gather/large_f32_threads=1-internal/131072 30.81 µs 22.20 µs -28.0%
🟢 tokenization/encode_tokens_per_second 594.89 µs 420.56 µs -29.3%
🟢 matmul/small_generic_f16_threads=8/1x256x256 47.62 µs 33.60 µs -29.4%
🟢 matmul/small_generic_f16_threads=1/1x256x256 40.85 µs 28.46 µs -30.3%
🟢 add/small_f16_threads=1-internal/1024 724.9 ns 488.6 ns -32.6%
🟢 gather/large_bf16_threads=1-internal/131072 15.27 µs 10.10 µs -33.9%
🟢 add/small_bf16_threads=1-internal/1024 857.6 ns 436.0 ns -49.2%
🟢 matmul/small_generic_f32_threads=8/1x256x256 70.60 µs 34.63 µs -50.9%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 3.98 3.75 5.01 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

justinchuby added a commit that referenced this pull request Aug 22, 2026
> **A100 safety update (head `13c37a7f`):** expected shared-prefix
admission failures are now all-or-private: K/V commit is transactional,
a failed V share rolls K back before enqueue, and only a successful
rollback permits private-KV fallback. Fatal kernel faults remain
process-fatal; the new cross-process VMM probe proves a worker fault
does not affect an actively executing peer, while the current server is
still in-process and does not yet consume physical shared-prefix
metadata. See the [final verified
conclusion](#1579 (comment)).

Collapses phases 1–7 of the #1186 memory architecture rework onto
current `main` as a single merge. The stack forked 270 commits ago (176
of them touching `crates/`), so rebasing it layer by layer would mean
solving the same 17 conflicts fourteen times, with no reviewer ever
looking at the intermediate states.

Closes the stack: #1252 #1263 #1279 #1283 #1301 #1341 #1349 #1426 #1440
#1448 #1454 #1462 #1465 #1468 #1533.

**A review guide is in the first comment.** It is the part worth reading
— this diff is 103 files, but only three decisions in it are ones a
compiler cannot check.

## What was already reviewed, and what wasn't

Every phase was reviewed and approved by a session that was not its
author, under the rejection-lockout rule (a rejected author never writes
the next revision).

| Phase | PR | Verdict |
|---|---|---|
| 1–5 | #1252 → #1426 | approved |
| 6 | #1440, #1448, #1454 | **rejected** ×3 |
| 6 | #1462 | approved |
| 7 | #1465 | **rejected** |
| 7 | #1468 | approved |
| 7 test fixes | #1533 | approved, A100-verified |

The four rejected rounds were not replaced — the later rounds are
stacked **on top of** them. So the tree here is the approved state, but
the history contains the rejected commits. **#1440 / #1448 / #1454 /
#1465 should not be reviewed individually**; they close automatically.

**Not reviewed anywhere:** the conflict resolutions themselves. That is
what this PR is for.

## Verification

Run here, on this merged tree:

- `cargo check --workspace --all-targets` — clean
- `cargo check -p onnx-genai-engine --features cuda,native-backend
--all-targets` — clean. Worth calling out: the default feature set does
**not** compile the `cfg(cuda)` code, which is where the riskiest edits
are. Checking only the default set would have missed a real break (see
the guide).
- `cargo test` over 7 crates — **1105 passed, 2 failed, 88 ignored**
- `cargo clippy --workspace --all-targets`
- `cargo fmt`

**The 2 failures and the 1 clippy error are pre-existing and proven so,
not merge damage:**

-
`platform_capacity::{disk_capacity_is_measured_for_the_working_directory,
an_explicit_byte_limit_is_honored_without_a_device_query}` —
`platform_capacity.rs` is byte-identical to `main` (`git diff
origin/main HEAD -- ` that file is empty). The cause is a macOS-only FFI
layout bug: `fsblkcnt_t` is 4 bytes on macOS (confirmed: `sizeof(struct
statvfs)=64`, `sizeof(fsblkcnt_t)=4`) while the Rust struct declares
`f_blocks`/`f_bfree`/`f_bavail` as `u64`. CI is Linux, where it is 8
bytes. Left alone: it is a real bug but not this PR's.
- `optimizer.rs` clippy `approx_constant` — present on both parents; the
merged file is byte-identical to `main`.
- Three files had `rustfmt` drift already present on `main`
(`matmul_nbits.rs`, `normalization.rs`, `optimizer.rs`). Reverted rather
than swept in, to keep the diff readable.

**Final A100 revalidation (head `13c37a7f`):** full CUDA-memory GPU
suite passed; CUDA EP default-parallel lib suite passed (**488 passed /
17 ignored**); real fp16 GQA covered three requests × two interleaved
decode steps with two shared peers and one transactional private
fallback, byte-identical to independent GPU and CPU references. A
device-started/event-gated worker remained healthy through a peer
process `CUDA_ERROR_ILLEGAL_ADDRESS`, owner exit, replacement worker,
and process restart. Governor tests, targeted Miri, CI Clippy, fmt, and
CUDA honesty also passed.

## Still open after this merges

- `memory-plugin-provider-wiring` — the back half of Phase 6 criterion 8
(provider/context pinning), which fell in the gap between phases.
**#1186 must not be closed until it lands.** This merge makes it more
tractable, not less: see decision 3 in the guide.
- `memory-deferred-invariant-asserts` — two surviving mutants found by
#1462's reviewer, adjudicated NON-BLOCKING and not defects.
- #1533's CUDA honesty guard warns on a legitimately portable anchor.
Should be moved out of the CUDA test binary or allow-listed rather than
left warning.

---------

Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <copilot@github.com>
Copilot-Session: 46c5d75b-8146-489c-b82f-08ee29c27ce4
Copilot-Session: 39ff6824-d35f-4d3b-8f5f-043a7119a100
Copilot-Session: c80f8522-983c-47f7-8241-2155a823aabe
@justinchuby

Copy link
Copy Markdown
Owner Author

Closing: superseded by #1579, which collapsed phases 1–7 into a single merge — landed as a36964280 (squash, 126 files, +43956 −3286).

This PR was approved by a session that was not its author, and its work is in main.

Containment verified, not assumed:

$ git merge-base --is-ancestor 9e0e20a03 ebceab071   # ebceab071 = #1579 head
→ ancestor

Every one of the 15 stack heads (#1252 → #1533) is a literal ancestor of the merged head, and I separately confirmed the merged head's content reached main: of the 126 files #1579 touched, exactly one differs from main — crates/onnx-runtime-session/src/executor/tests.rs, where main carries 120 extra lines from #1703 (Expand-broadcast mask tests). That is the conflict resolution correctly preserving the other side, not content loss.

Why this didn't close itself: #1579's body says Closes the stack: #1252 #1263 …, which GitHub does not parse — the closing keyword must be immediately followed by the reference, and in any case closing keywords only ever close issues, never pull requests. So all 15 stayed open by mechanism, not by intent.

#1186 stays open: #1579 explicitly gates it on memory-plugin-provider-wiring (the back half of Phase 6 criterion 8, provider/context pinning), which has not landed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant