Skip to content

feat(cuda-ep): boundary-time route-telemetry consumer wiring into #1854 coarse residency (#1810 Slice 7B) - #1971

Merged
justinchuby merged 1 commit into
mainfrom
squad/1810-slice7b-telemetry-residency-consume
Aug 24, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/1810-slice7b-telemetry-residency-consume

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

#1810 Slice 7B — boundary-time route-telemetry consumer

Status: DRAFT — awaiting independent review. Do not merge.

Closes the loop opened by the merged Slice-6/7A producer (expert-route telemetry, PR #1922 e1ec495ee) and the merged Slice-4/5 coarse-boundary residency lifecycle (PR #1854). This is the smallest production seam the Slice-6 design §8 specifies:

expose a boundary-time consumer that produces a per-expert desired-set, and feed that set to the existing Slice 4/5 coarse-boundary plan application (coarse_residency.rs) as its policy input — with no new allocator, no id→slot rewrite, and every mapping change still owned by PMM/VMM.

What it does

New module crates/onnx-runtime-ep-cuda/src/route_residency.rs:

consume_route_window_at_boundary(...):

  1. Gate — default-off via the existing COARSE_RESIDENCY_ENABLE_ENV (coarse_residency_profile_enabled()). When off (shipped default) it returns Disabled before reading the snapshot or touching any allocator.
  2. Safe boundary — re-reads the existing CudaWeightResidency::resize_safe_point and fails closed with RejectedNotSafeBoundary { reason } if a graph is capturing/replaying, an admission is in flight, a deferred release has not settled, execution is multi-device, or a routed-residency guard is live. Coarse-boundary only — never a per-token remap.
  3. Validate — runs the producer's own consume_and_validate on the already-completed window snapshot; fail-closed (→ WholeBank { reason }) on poison / overflow / stale epoch / foreign request / foreign device, or when no in-range experts were recorded.
  4. Plan — turns the routed-expert union into a desired hot set and asks the already-validated StaticProfileResidencyPolicy to shape a ResidencyPlan (the design's "record → desired set" RouteObserverPolicy role — reused, not duplicated, so exactly one validated policy emits PerExpertCandidate).
  5. Apply — hands the plan to the existing CudaWeightResidency::apply_coarse_residency_plan, which remains the sole authority that maps / unmaps / accounts / quarantines / rolls back through PMM/VMM.

RouteWindowConsumeOutcome carries the exact reason on every non-applied path — no silent fallback.

Invariants held by construction

  • Allocates nothing, opens no stream, owns no VA, copies no device bytes — pure host glue between two existing authorities.
  • No remap during capture/replay; coarse-boundary only, never per-token.
  • No new host sync in steady state (only the producer's already-taken snapshot and the existing transition primitive's drain).
  • Default-off & byte-identical when disarmed (two independent default-off switches: telemetry disarmed and this gate off).
  • Same-device fail-closed and PMM/VMM remains the sole mapping/accounting/quarantine/rollback authority.

Tests — tests/route_residency_consume_gpu.rs (6 GPU tests, all passing on an idle A100)

  • disabled_gate_is_structural_no_op — off path is a structural no-op (byte-identical).
  • route_window_hot_set_transitions_cold_experts — routed hot-set stays resident, cold set tiers to host, bytes identical.
  • expert_group_transitions_atomically_from_window — atomic expert-group transition driven from a window.
  • active_capture_and_multi_device_reject_consume — active-capture and multi-device boundaries reject the consume (fail-closed).
  • foreign_identity_and_defective_windows_fail_closed — foreign request/device and poison/overflow/stale windows fail closed to whole-bank.
  • injected_fault_rolls_back_consumer_transition — injected driver fault (fail_nth(Unmap, 3) over an isolated + merged cold range) rolls back range-precisely and quarantines; rollback_count == 1, values_touched == 0, committed_values empty, content bit-identical (mirrors the proven coarse_residency_plan_gpu.rs rollback fixture).

Run:

env -u ONNX_GENAI_WEIGHT_OFFLOAD_COARSE_RESIDENCY_ENABLE CUDA_VISIBLE_DEVICES=<idle> ONNX_GENAI_CUDA_DEVICE=0 \
  cargo test -p onnx-runtime-ep-cuda --features cuda,cuda-13000,gpu-tests --release \
  --test route_residency_consume_gpu -- --ignored --test-threads=1

Honest scope

Like coarse_residency::apply_residency_plan_at_boundary when it shipped (Slice 5), this consumer has no live decode-loop call site yet — wiring it into a running session's request boundary is the next slice. It ships here as the production seam: reachable, default-off, and proven by the GPU tests. cargo fmt applied; cargo clippy on the crate lib is clean for the new file (pre-existing unrelated warnings/errors in optimizer.rs / standard_attention.rs are out of scope).

Constraints honored

Did not touch PagedAttention or IQ1 fusion files. Branch squad/1810-slice7b-telemetry-residency-consume based on latest main (011fbb284, includes #1922 e1ec495ee). #1788/#1800/#1884/#1854 lifecycle regressions preserved (reused, not forked).

Refs #1810.

… coarse residency (#1810 Slice 7B)

Closes the loop opened by the merged Slice-6/7A producer (expert-route
telemetry, PR #1922 `e1ec495ee`) and the merged Slice-4/5 coarse-boundary
residency lifecycle (PR #1854). Adds the smallest production seam that the
Slice-6 design §8 specifies: a boundary-time consumer that turns one already
completed telemetry window into a per-expert desired hot-set and feeds it to
the existing coarse-boundary plan application.

New module `route_residency`:
  * `consume_route_window_at_boundary` — gate (default-off via the existing
    `COARSE_RESIDENCY_ENABLE_ENV`) -> re-read the existing `resize_safe_point`
    and fail closed off any unsafe/in-flight/multi-device boundary -> validate
    the window with the producer's own `consume_and_validate` (fail-closed on
    poison / overflow / stale epoch / foreign request / foreign device) ->
    shape a `ResidencyPlan` via the already-validated
    `StaticProfileResidencyPolicy` (reused, not duplicated) -> apply through
    `CudaWeightResidency::apply_coarse_residency_plan`, which remains the sole
    PMM/VMM mapping/accounting/quarantine/rollback authority.
  * `RouteWindowConsumeOutcome` carries the exact reason on every non-applied
    path — no silent fallback.
  * A `#[cfg(any(test, feature="gpu-tests"))]` fault-injection twin routes
    through `apply_residency_plan_at_boundary_with_phase8_faults` to prove
    range-precise rollback/quarantine end to end.

Invariants held by construction: allocates nothing, opens no stream, owns no
VA, copies no device bytes; no remap during capture/replay; coarse-boundary
only (never per-token); no new host sync in steady state; default-off and
byte-identical when disarmed. No live decode-loop call site yet — ships as the
tested production seam, exactly as `apply_residency_plan_at_boundary` did when
Slice 5 shipped.

Tests (`tests/route_residency_consume_gpu.rs`, 6 GPU tests, all passing):
  * disabled gate is a structural no-op (byte-identical),
  * routed hot-set keeps its experts resident and tiers the cold set,
  * expert group transitions atomically from a window,
  * active-capture and multi-device boundaries reject the consume,
  * foreign identity / defective windows fail closed to whole-bank,
  * injected driver fault rolls back range-precisely and quarantines, content
    bit-identical (mirrors the proven coarse-residency rollback fixture).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby

Copy link
Copy Markdown
Owner Author

Ready for independent review (draft; do not merge). Independent reviewer excluding Deckard and myself, per reviewer-protocol.

Scope: smallest production consumer wiring the completed coarse-boundary route-telemetry window (#1922 e1ec495ee) into the existing #1854 coarse-residency lifecycle. New module route_residency.rs + 6 GPU tests.

Verification just run (idle A100, GPU isolated, --test-threads=1): 6/6 GPU tests pass — disabled no-op byte-identical, routed hot-set resident / cold tiered, atomic group transition, active-capture + multi-device reject, foreign/poison/overflow/stale fail-closed to whole-bank, and injected-fault range-precise rollback (rollback_count==1, values_touched==0, content bit-identical). cargo fmt applied; crate-lib clippy clean for the new file (pre-existing unrelated optimizer.rs/standard_attention.rs findings out of scope).

Invariants: default-off, byte-identical when disarmed; no remap during capture/replay; coarse-boundary only (never per-token); no new allocator/stream/host-sync; PMM/VMM remains the sole map/account/quarantine/rollback authority; same-device fail-closed; no silent fallback.

Honest scope: no live decode-loop call site yet — ships as the reachable, default-off, tested production seam (as apply_residency_plan_at_boundary did at Slice 5). Did not touch PagedAttention or IQ1 fusion files.

Please flag substantive findings; I'll revise on this same branch before requesting final architecture re-review.

@codecov

codecov Bot commented Aug 24, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.72%. Comparing base (011fbb2) to head (94ff7cc).
⚠️ Report is 15 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1971      +/-   ##
==========================================
+ Coverage   80.21%   80.72%   +0.50%     
==========================================
  Files         421      423       +2     
  Lines      203480   207429    +3949     
  Branches   203480   207429    +3949     
==========================================
+ Hits       163231   167450    +4219     
+ Misses      34675    34341     -334     
- Partials     5574     5638      +64     
Flag Coverage Δ
cli-ort-linux 72.51% <ø> (?)
cli-ort-windows 72.10% <ø> (+0.09%) ⬆️
mlas 85.23% <ø> (?)
offline 80.87% <ø> (+0.43%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 30 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@justinchuby

Copy link
Copy Markdown
Owner Author

Independent Review — PR #1971 (Slice 7B route-telemetry consumer)

Verdict: ✅ APPROVE FOR MERGE (draft seam — do not merge until the draft flag is cleared and the base is rebased; see notes)

Scope reviewed: HEAD 94ff7ccde, read-only. 3 files: lib.rs (+1 module registration), src/route_residency.rs (+388, new), tests/route_residency_consume_gpu.rs (+1099, new). Base merge-base 011fbb284 (#1955); #1922 e1ec495ee and #1955 both confirmed in ancestry.

Reachability / dead-code gate (the crux)

consume_route_window_at_boundary has zero non-test production callers (grep-verified: only lib.rs registration + the integration test). Normally that is a blocker. It clears the gate here because it is an honest, default-off, architecturally-required inert seam, not a dead-code overclaim:

⚠️ Merge condition: preserve the honest scope. A future wiring slice must place the call at an existing coarse safe boundary and must not silently flip the enable gate without its own correctness/perf gate — Slice 7A's G1 NO-GO still stands.

Safety/correctness gates — all verified against the actual referenced code

  • Safe-boundary authority ✓ reuses resize_safe_point().blocking_reason(); apply_coarse_residency_plan re-verifies via verify_safe_point. Fail-closed on capture/replay, unsettled deferred releases, admission-in-flight, multi-device, live routed guards — each with a deterministic reason.
  • No capture/replay remap, no hidden sync/alloc ✓ module allocates nothing, opens no stream, owns no VA; only host reads of the producer's already-copied snapshot; coarse-boundary only.
  • Policy vs PMM/VMM authority ✓ StaticProfileResidencyPolicy only shapes the plan; apply_coarse_residency_plan → Slice-4 transition primitive remains the sole map/unmap/account/quarantine/rollback authority. Reused, not duplicated — no file overlap.
  • Same-device fail-closed ✓ check_same_device(...) precheck + multi-device safe-point block.
  • Atomic expert-group / range rollback + quarantine ✓ group decision-shape agreement (Cycle-19 fix) + group-mate precheck suppression + reverse-tier rollback loop.
  • Poison / overflow / stale / foreign ✓ consume_and_validate fails closed to WholeBank on poison, overflow, foreign device, foreign request, stale epoch; empty routed set → WholeBank.
  • Disabled byte-identical ✓ default-off gate returns Disabled before reading snapshot/allocator.
  • No silent WholeBank fallback masking errors ✓ every non-Applied variant carries an explicit reason; Applied carries the full BoundaryApplicationOutcome.

Independent test + lint (run by me)

  • 6/6 GPU tests pass on an idle A100 (CUDA_VISIBLE_DEVICES=1, --test-threads=1, --features cuda,cuda-13000,gpu-tests --release). Non-vacuous: real VMM transitions with committed/reserved accounting + byte-identity readback; disabled no-op byte-identical; hot-set tiering; atomic expert-group; multi-device + active-capture reject (no side effects); foreign/poison/overflow/stale/empty → WholeBank (no side effects); injected fault (fail_nth(Unmap,3)) → range-precise rollback (values_touched=0, rollback_count=1, committed_values=[], rollback_failures=[], quarantined) with content preserved.
  • cargo fmt clean. cargo clippy: the new files produce zero diagnostics. The one blocking clippy error (optimizer.rs:4811, 3.141_592_7 PI-approx literal) and all warnings are in pre-existing files this PR does not touch — out of scope and correctly disclosed. (Note for the pod: that lib-test clippy error is red on the base itself under CI -D warnings, independent of this PR.)

Base / rebase ordering

Do not merge (draft). Approving the seam as reviewed.

— Independent reviewer · requested by Justin Chu · 2026-08-24

@justinchuby
justinchuby marked this pull request as ready for review August 24, 2026 12:23
@justinchuby
justinchuby merged commit 98731d3 into main Aug 24, 2026
17 checks passed
@justinchuby
justinchuby deleted the squad/1810-slice7b-telemetry-residency-consume branch August 24, 2026 12:23
justinchuby added a commit that referenced this pull request Aug 24, 2026
…ndary (default-off, draft) (#2007)

## Slice 7C — production boundary wiring for the route-telemetry
consumer

Wires the merged #1971 boundary-time consumer
(`consume_route_window_at_boundary`) into the **real production
decode/request lifecycle** at the single existing coarse safe boundary,
while remaining **default-off** and **byte-identical when disabled**.

Closes the gap #1971's module docs called out: *"no live decode-loop
call site yet … wiring it into a running session's request boundary is
the next slice."*

### The one boundary
`Executor::finish_device_validation`
(`crates/onnx-runtime-session/src/executor/run.rs`) — the one
request-level host boundary that runs `ep.sync()` after stream/graph
completion, only for top-level (`!nested`) runs, once per decode
step/request, past capture/replay. The consumer is called there and
**nowhere else**. No second boundary mechanism, no model allowlist.

### How
- **ep-api:** new **required**
`ExecutionProvider::consume_route_residency_at_boundary(&self) ->
Result<()>` (no compatibility default). Every in-repo
EP/mock/test-double implements it explicitly — non-residency EPs return
`Ok(())`, the plugin provider forwards to its inner EP, the
planning-only capability gate is `unreachable!()` — so each provider
states its boundary behaviour and the disabled path stays byte-identical
at runtime.
- **session:** the success arm of `finish_device_validation` calls it
after the device-validation latch is confirmed clean.
- **ep-cuda:** `RouteTelemetrySource` (impl for `QMoEKernel`,
compile-time asserted), `RouteResidencyBoundary` (pure binding of
existing authorities — owns no allocator, maps nothing),
`RouteResidencyDiagnostics`, and `run_route_residency_boundary`. The
CUDA EP override:
1. reads the existing default-off gate (`COARSE_RESIDENCY_ENABLE_ENV`)
first — when off (shipped default) it is a single env read: no lock, no
snapshot, no allocation, no CUDA launch, no host sync, no telemetry
reset;
2. looks up one optional installed boundary binding — `None` in
production today (honest reachable seam, exactly like 7A/7B shipped);
3. when bound, drives `resize_safe_point` pre-check →
`route_telemetry_snapshot` → the merged #1971
`consume_route_window_at_boundary` → `reset_route_telemetry_boundary`
**once**; reset + expected-epoch advance fire **only** after a window is
actually consumed (unsafe/disarmed boundaries neither snapshot nor
reset);
  4. records the typed outcome — no silent WholeBank / default-success.

PMM/VMM stays the sole map/unmap/account/quarantine/rollback authority;
coarse cadence/hysteresis stay policy-owned; no per-token remap; no
remap during capture/replay.

### Tests (traverse the production caller, never the raw consumer)
- **CPU** (`executor/tests.rs`, mock EP through `Executor::run`): fires
exactly once per top-level run, after `sync`, not per kernel; **zero**
boundaries for a nested `If`-subgraph run.
- **GPU** (`tests/route_residency_boundary_gpu.rs`, idle A100, through
the EP trait method / its phase-8 fault sibling): disabled structural
no-op; ≥3-replay union → atomic two-member expert-group transition →
next empty window; multi-device + active-capture reject before
consume/reset; foreign-request (multi-request
isolation)/foreign-device/poison/overflow/stale/empty fail closed to
whole-bank; driver-fault rollback preserves content.

`cargo fmt` clean; `cargo clippy --all-targets -D warnings` clean on the
touched crates (ep-api, ep-cuda incl. `gpu-tests`, session); 150 session
executor unit tests pass; the 6 Slice-7B consume GPU tests and the qmoe
route-telemetry GPU test still pass (no regression).

### Scope / honesty
- **Default-off; no performance / tok-s claim.** This is a wiring seam.
Constructing a binding from a running session's expert banks (live
producer registration) is a later slice — production installs no binding
yet.
- Disabled path is byte-identical and adds no CUDA launch, host sync,
allocation, mapping, or telemetry reset.
- Avoided PagedAttention / HCA·CSA / Mobius / BQMoE-fusion files; the
only session caller file touched is `executor/run.rs` (one success-arm
line).

**Draft — stop for independent review; do not merge.**

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 24, 2026
…ate groups (B6) (#2008)

## DeepSeek‑V4 HCA "C1" runtime slice — single accounting authority for
CSA/HCA device state groups (B6)

**Draft. Do not merge — opened for independent review.**

This is the first, GPU‑verified increment of the DeepSeek‑V4 **HCA
(Hierarchical Compressed Attention)** "C1" runtime slice. It closes
design blocker **B6** from `deckard-deepseek-v4-csa-hca-cuda-slice.md`
§7 ("`onnx-genai-kv`/shared accounting is not the authority; a second
*unaccounted* `alloc_raw` path exists for CSA device buffers") by
introducing a **single property‑typed accounting authority** for CSA
(`CompressedSparseAttention`) / HCA device *state groups* in the CUDA
EP.

### What landed
New module `crates/onnx-runtime-ep-cuda/src/kernels/csa_state_group.rs`:

- **`CsaStateGroupDescriptor` — a property gate.** Parsed from the
node's declared attributes + state input edges (`compression_ratio`,
`cache_format`, `num_heads`/`head_dim`/`qk_rope_head_dim`,
`index_head_dim`, `device_count`, presence of the compressed(6)/carry(7)
state edges). **Not a model‑name/shape allowlist.** Carries the exact
reason on refusal via the typed **`CsaStateGroupRefusal`** enum
(UnsupportedRatio, UnknownCacheFormat, Ratio4RequiresFp8Cache,
Ratio4RequiresIndexHeadDim128, Ratio128RejectsFp4, InvalidHeadGeometry,
RopeExceedsHeadDim, MultiDeviceAmbiguity, MissingStateEdge, OutOfMemory,
UnsupportedC1Ratio).
  - `validate()` accepts the op's full ratio‑4/128 range (op‑level).
- `validate_c1_runtime()` is the stricter **C1 admission gate: ratio‑128
(HCA) only** — ratio‑4/MTP are op‑supported but out of C1 scope and are
**typed‑refused, never silently threaded or dense‑fallback'd.**
- **`CsaStateGroupLedger` — the sole, mandatory accountant.** Owns
**bytes, never cursors**; charges per `(request, device)` for isolation;
`try_charge` is a check‑then‑charge under one mutex that **fails closed
BEFORE any physical allocation** (managed limit; `0 = unlimited`, the
steady‑state default); RAII `CsaStateGroupCharge` releases on `Drop` so
teardown returns to baseline. It is **instance state threaded through
`CsaMetrics` (§8) — not a process‑wide global**, and it is **not** an
allocator/cache/page‑manager: it accounts the bytes the existing
physical path (`runtime.alloc_raw`) reserves.

**No optional/no‑ledger escape hatch.** Per the
*no‑backward‑compatibility during development* directive, the ledger is
a **mandatory** `Arc<CsaStateGroupLedger>` on the reservation path (not
an `Option`), so the **type system alone forbids an unaccounted
reservation** — the exact B6 hazard. The earlier `Option<ledger>` seam
was replaced outright rather than wrapped in a compat shim, and every
in‑repo caller/fixture/test was updated atomically. "Disarmed" is
therefore an **unlimited limit (0 sentinel)**, not a structural bypass:
accounting still runs, it simply never refuses, so byte‑identity is
preserved without leaving a second unaccounted path.

Wiring:
- `CsaDeviceBufferManager::reserve` **unconditionally** charges the
ledger **before** `alloc_raw` and holds the RAII charge for the buffer
lifetime (Drop frees buffers, then releases the charge).
- The CUDA CSA factory `create()` builds the descriptor, `validate()`s
it (typed refusal), takes a per‑runner request id, and threads the
(mandatory) ledger + `(request, device)` key.
- `CsaMetrics` gained the `Arc<CsaStateGroupLedger>` plus `pub`
observability accessors (resident/peak/compressed/dense‑ring bytes,
charge failures, active group count, per‑`(request,device)` residency,
get/set limit).

### Invariants honored
- **Default‑off / byte‑identical.** The charge happens **once at runner
construction**, never per token and never on the captured/replay stream;
release is host‑only. With the ledger unlimited (default) the
reservation is byte‑identical to the pre‑accounting path. **No device op
or host sync is added to steady‑state replay.**
- **Fail‑closed.** Over‑limit groups are refused with a typed
`OutOfMemory` **before** any device memory is reserved (no leak).
- **Isolation.** Residency is tracked per `(request, device)`; teardown
returns to baseline.
- **Property‑based refusals** (not model‑name); cursors stay
backend‑owned; PMM/VMM/governor remains the mapping/accounting authority
for physical device memory (this ledger is an accounting seam **in front
of** it, not a competing manager).
- No PagedAttention / Mobius / QMoE / BQMoE / residency / #1971 files
touched.

### Tests & measurement
- **11 CPU‑only** descriptor/ledger tests: accept/refuse matrix,
fail‑closed OOM without mutation, per‑`(request,device)` isolation,
baseline/peak, monotonic ids, C1 ratio gate, limit round‑trip.
- **6 GPU tests** on a real `CsaDeviceBufferManager` (A100‑80GB, GPU 2):
`charge == sum(reserved)`, teardown → baseline,
**fail‑closed‑before‑alloc with no leak**, per‑request isolation,
**stable workspace addresses**, and an **unlimited‑ledger reservation
that accounts without ever refusing** (`charge_failures == 0`, teardown
→ baseline — replacing the old no‑ledger case now that the ledger is
mandatory).
- **Regression:** the existing CSA op‑parity GPU suite — **26 passed / 1
ignored** (the ignored one needs external Mobius artifacts) — stays
green, covering eager + capture/replay + speculative rollback +
observability. This is the byte‑identical‑when‑disarmed proof.
- `cargo fmt` clean; `cargo clippy -D warnings` clean on **lib and
tests**.
- **Overhead:** by construction there is no steady‑state
kernel/host‑sync cost (accounting is one‑time at runner construction,
off the captured stream), so there is no per‑token GPU delta to measure
for this slice; the op‑parity capture/replay suite is the byte‑identity
control.
- Independent code‑review (excluding the original author and myself):
**no substantive findings.**

### Honestly scoped next steps (tracked, NOT in this slice)
1. **Physical‑allocator unification:** route
`CsaStateGroupLedger::try_charge` to delegate to the adopted
`MemoryGovernor` `reserve(Tier::Device, bytes, MemoryRole::Workspace)` /
`onnx-genai-kv` so the *physical* bytes become governor‑visible (final
B6 unification). The ledger's `try_charge` is already shaped as that
adapter seam. This is an architecture migration (not a compat shim),
deliberately deferred.
2. **Native‑decode threading:** the smallest production caller that
threads the ratio‑128 `present_* → past_*` compressed/carry state at the
existing `ensure_capacity` safe boundary and calls
`validate_c1_runtime`, with captured ≥16‑decode **CPU/CUDA bit‑parity**
and `fallbacks == 0`.

CPU op‑level present→past carry/cursor progression across the decode
boundary is already covered by the CPU oracle test
`ratio128_stateful_carry_matches_full_recompute_across_decode_boundary`.

<sub>Co-authored-by: Copilot</sub>

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 24, 2026
… (follow-up to #2008) (#2029)

## B6.2 — Unify CSA/HCA device accounting under the shared
`MemoryGovernor`

**Draft. Do not merge. Independent review required (excluding the
author).**
Follow-up to **#2008** (B6, merged as `7a162e9f4`), now rebased directly
on `main`
(the #2008 dependency has landed, so this is no longer stacked — it is a
clean
single-commit PR on `main`).

### Problem (B6.2)

#2008 introduced `CsaStateGroupLedger` as the single *typed* accountant
for CSA/HCA
device state groups, but it kept its **own byte limit**
(`set_limit`/`limit`) beside
the process-wide `MemoryGovernor`. Physical bytes were reserved through
the EP's
shared `CudaRuntime::alloc_raw`, yet admission was decided by a
**private CSA gate** —
a second set of books. That is exactly the "two allocators / two
ledgers" shape the
B6 design blocker exists to remove.

### Change — one accounting authority, one transaction

Every CSA reservation now admits through the shared governor **before**
any physical
allocation:

```
governor.reserve(Tier::Device, total, MemoryRole::Workspace { step_scoped: false }, holder)
    -> MemoryLease         // held by the RAII CsaStateGroupCharge for the buffers' lifetime
```

- CSA device residency is now counted in
`MemoryGovernor::used(Tier::Device)`, in the
  **same books** as every other device holder (weights, KV, workspaces).
- The per-`(request, device)` / per-class counters
(`compressed`/`carry`/`dense_ring`/
`index`/`scratch`) remain as a **derived attribution mirror** for
observability the
governor deliberately does not itemise. **The mirror never gates
admission**, so it is
  complementary telemetry, not a competing ledger.
- **Single reserve / commit / rollback transaction:** a refused charge
mutates nothing
(fail closed with a typed `OutOfMemory` *before* `alloc_raw`); an alloc
failure
mid-reservation rolls back the physical buffers and releases the
governor lease
**exactly once** via RAII (`Drop` clears the mirror, then the
`MemoryLease` field drops).
- Physical allocation stays on the EP's shared `alloc_raw` inside the
governed
transaction — this removes the *unaccounted* property (the real B6
defect) **without**
inventing pages or touching the later-built
`DeviceAllocator`/`ProcessMemoryManager`
  teardown ordering.

### Invariant / API

- **API (crate-internal):**
- `CsaStateGroupLedger::new(Arc<dyn MemoryGovernor + Send + Sync>)` —
governor-backed.
- `CsaStateGroupLedger::default()` — unlimited reference
`LedgerGovernor` (disarmed,
byte-identical); `#[cfg(test)] with_device_limit(bytes)` for fail-closed
proofs.
- `device_available_bytes()` / `governor_device_used()` replace the
removed
    `limit()` / `set_limit()`.
- `CsaMetrics::with_governor(governor)`;
`csa_state_group_device_available_bytes()` /
`csa_state_group_governor_device_used()` replace the removed `*_limit`
accessors.
  - Provider threads its process governor at registry build
    (`new_with_policy_governor_and_manager`); the default (no-governor)
`build_cuda_registry` path stays on the unlimited reference governor.
- **No backward-compat shim** (per directive): the private-limit API is
deleted and
  every in-repo caller/test updated atomically in this commit.

### Tests (same commit, new semantics)

CPU-only (no GPU):
- `governor_sees_csa_reservation_in_shared_books` — a CSA charge adds to
an unrelated
holder's total on one governor; teardown returns to baseline; the other
holder is
  isolated.
- `charge_reserves_and_releases_governor_exactly_once` — a fake counting
`MemoryGovernor` proves **one reserve, one release** (no double charge,
no second path).
- `ledger_reports_governor_device_ceiling`,
`ledger_fails_closed_over_limit_without_mutating`
— fail closed against the governor's device ceiling, no partial
reservation.

GPU (idle A100, correctness — no perf claim):
- `reserve_charges_ledger_and_releases_on_drop`,
`reservations_are_isolated_per_request`,
  `reserved_workspace_addresses_are_stable`,
`reserve_fails_closed_over_limit_without_leaking` (via
`with_device_limit`),
`unlimited_ledger_accounts_without_refusing` — now assert the
**physical** reservation
is governor-visible (`governor_device_used()`) and returns to baseline
on teardown.
- **CSA op-parity + capture/replay** integration suite unchanged: **26
passed, 1 ignored**
(the ignored one needs external Mobius/MTP artifacts) — this is the
stable-address /
  capture-no-alloc / byte-identity control.

### Validation

- `cargo fmt -- --check`: clean.
- `cargo clippy -p onnx-runtime-ep-cuda --features cuda,cuda-13000 --lib
--tests -- -D warnings`: clean.
- `cargo test -p onnx-runtime-ep-cuda --features cuda,cuda-13000 --lib
csa_` on idle A100:
  **19 passed, 0 failed**.
- `--features cuda,cuda-13000,gpu-tests --release --test
compressed_sparse_attention_gpu`:
  **26 passed, 1 ignored**.

### Scope / honest next steps (NOT in this slice)

- Migrating physical allocation off `alloc_raw` onto a **governed
`DeviceAllocator`
handle** (subject to the strict PMM/VMM teardown ordering) is a further,
separately
  scoped step.
- **`native_decode` present→past** state threading remains the next
stacked slice.
- No PagedAttention / Mobius / QMoE / BQMoE / residency / #1971 files
touched.
- No full-model perf / tok/s claim.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant