Repository navigation
feat(cuda-ep): boundary-time route-telemetry consumer wiring into #1854 coarse residency (#1810 Slice 7B) - #1971
Conversation
… coarse residency (#1810 Slice 7B) Closes the loop opened by the merged Slice-6/7A producer (expert-route telemetry, PR #1922 `e1ec495ee`) and the merged Slice-4/5 coarse-boundary residency lifecycle (PR #1854). Adds the smallest production seam that the Slice-6 design §8 specifies: a boundary-time consumer that turns one already completed telemetry window into a per-expert desired hot-set and feeds it to the existing coarse-boundary plan application. New module `route_residency`: * `consume_route_window_at_boundary` — gate (default-off via the existing `COARSE_RESIDENCY_ENABLE_ENV`) -> re-read the existing `resize_safe_point` and fail closed off any unsafe/in-flight/multi-device boundary -> validate the window with the producer's own `consume_and_validate` (fail-closed on poison / overflow / stale epoch / foreign request / foreign device) -> shape a `ResidencyPlan` via the already-validated `StaticProfileResidencyPolicy` (reused, not duplicated) -> apply through `CudaWeightResidency::apply_coarse_residency_plan`, which remains the sole PMM/VMM mapping/accounting/quarantine/rollback authority. * `RouteWindowConsumeOutcome` carries the exact reason on every non-applied path — no silent fallback. * A `#[cfg(any(test, feature="gpu-tests"))]` fault-injection twin routes through `apply_residency_plan_at_boundary_with_phase8_faults` to prove range-precise rollback/quarantine end to end. Invariants held by construction: allocates nothing, opens no stream, owns no VA, copies no device bytes; no remap during capture/replay; coarse-boundary only (never per-token); no new host sync in steady state; default-off and byte-identical when disarmed. No live decode-loop call site yet — ships as the tested production seam, exactly as `apply_residency_plan_at_boundary` did when Slice 5 shipped. Tests (`tests/route_residency_consume_gpu.rs`, 6 GPU tests, all passing): * disabled gate is a structural no-op (byte-identical), * routed hot-set keeps its experts resident and tiers the cold set, * expert group transitions atomically from a window, * active-capture and multi-device boundaries reject the consume, * foreign identity / defective windows fail closed to whole-bank, * injected driver fault rolls back range-precisely and quarantines, content bit-identical (mirrors the proven coarse-residency rollback fixture). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Ready for independent review (draft; do not merge). Independent reviewer excluding Deckard and myself, per reviewer-protocol. Scope: smallest production consumer wiring the completed coarse-boundary route-telemetry window (#1922 Verification just run (idle A100, GPU isolated, Invariants: default-off, byte-identical when disarmed; no remap during capture/replay; coarse-boundary only (never per-token); no new allocator/stream/host-sync; PMM/VMM remains the sole map/account/quarantine/rollback authority; same-device fail-closed; no silent fallback. Honest scope: no live decode-loop call site yet — ships as the reachable, default-off, tested production seam (as Please flag substantive findings; I'll revise on this same branch before requesting final architecture re-review. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1971 +/- ##
==========================================
+ Coverage 80.21% 80.72% +0.50%
==========================================
Files 421 423 +2
Lines 203480 207429 +3949
Branches 203480 207429 +3949
==========================================
+ Hits 163231 167450 +4219
+ Misses 34675 34341 -334
- Partials 5574 5638 +64
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
Independent Review — PR #1971 (Slice 7B route-telemetry consumer)Verdict: ✅ APPROVE FOR MERGE (draft seam — do not merge until the draft flag is cleared and the base is rebased; see notes) Scope reviewed: HEAD Reachability / dead-code gate (the crux)
Safety/correctness gates — all verified against the actual referenced code
Independent test + lint (run by me)
Base / rebase ordering
Do not merge (draft). Approving the seam as reviewed. — Independent reviewer · requested by Justin Chu · 2026-08-24 |
…ndary (default-off, draft) (#2007) ## Slice 7C — production boundary wiring for the route-telemetry consumer Wires the merged #1971 boundary-time consumer (`consume_route_window_at_boundary`) into the **real production decode/request lifecycle** at the single existing coarse safe boundary, while remaining **default-off** and **byte-identical when disabled**. Closes the gap #1971's module docs called out: *"no live decode-loop call site yet … wiring it into a running session's request boundary is the next slice."* ### The one boundary `Executor::finish_device_validation` (`crates/onnx-runtime-session/src/executor/run.rs`) — the one request-level host boundary that runs `ep.sync()` after stream/graph completion, only for top-level (`!nested`) runs, once per decode step/request, past capture/replay. The consumer is called there and **nowhere else**. No second boundary mechanism, no model allowlist. ### How - **ep-api:** new **required** `ExecutionProvider::consume_route_residency_at_boundary(&self) -> Result<()>` (no compatibility default). Every in-repo EP/mock/test-double implements it explicitly — non-residency EPs return `Ok(())`, the plugin provider forwards to its inner EP, the planning-only capability gate is `unreachable!()` — so each provider states its boundary behaviour and the disabled path stays byte-identical at runtime. - **session:** the success arm of `finish_device_validation` calls it after the device-validation latch is confirmed clean. - **ep-cuda:** `RouteTelemetrySource` (impl for `QMoEKernel`, compile-time asserted), `RouteResidencyBoundary` (pure binding of existing authorities — owns no allocator, maps nothing), `RouteResidencyDiagnostics`, and `run_route_residency_boundary`. The CUDA EP override: 1. reads the existing default-off gate (`COARSE_RESIDENCY_ENABLE_ENV`) first — when off (shipped default) it is a single env read: no lock, no snapshot, no allocation, no CUDA launch, no host sync, no telemetry reset; 2. looks up one optional installed boundary binding — `None` in production today (honest reachable seam, exactly like 7A/7B shipped); 3. when bound, drives `resize_safe_point` pre-check → `route_telemetry_snapshot` → the merged #1971 `consume_route_window_at_boundary` → `reset_route_telemetry_boundary` **once**; reset + expected-epoch advance fire **only** after a window is actually consumed (unsafe/disarmed boundaries neither snapshot nor reset); 4. records the typed outcome — no silent WholeBank / default-success. PMM/VMM stays the sole map/unmap/account/quarantine/rollback authority; coarse cadence/hysteresis stay policy-owned; no per-token remap; no remap during capture/replay. ### Tests (traverse the production caller, never the raw consumer) - **CPU** (`executor/tests.rs`, mock EP through `Executor::run`): fires exactly once per top-level run, after `sync`, not per kernel; **zero** boundaries for a nested `If`-subgraph run. - **GPU** (`tests/route_residency_boundary_gpu.rs`, idle A100, through the EP trait method / its phase-8 fault sibling): disabled structural no-op; ≥3-replay union → atomic two-member expert-group transition → next empty window; multi-device + active-capture reject before consume/reset; foreign-request (multi-request isolation)/foreign-device/poison/overflow/stale/empty fail closed to whole-bank; driver-fault rollback preserves content. `cargo fmt` clean; `cargo clippy --all-targets -D warnings` clean on the touched crates (ep-api, ep-cuda incl. `gpu-tests`, session); 150 session executor unit tests pass; the 6 Slice-7B consume GPU tests and the qmoe route-telemetry GPU test still pass (no regression). ### Scope / honesty - **Default-off; no performance / tok-s claim.** This is a wiring seam. Constructing a binding from a running session's expert banks (live producer registration) is a later slice — production installs no binding yet. - Disabled path is byte-identical and adds no CUDA launch, host sync, allocation, mapping, or telemetry reset. - Avoided PagedAttention / HCA·CSA / Mobius / BQMoE-fusion files; the only session caller file touched is `executor/run.rs` (one success-arm line). **Draft — stop for independent review; do not merge.** --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ate groups (B6) (#2008) ## DeepSeek‑V4 HCA "C1" runtime slice — single accounting authority for CSA/HCA device state groups (B6) **Draft. Do not merge — opened for independent review.** This is the first, GPU‑verified increment of the DeepSeek‑V4 **HCA (Hierarchical Compressed Attention)** "C1" runtime slice. It closes design blocker **B6** from `deckard-deepseek-v4-csa-hca-cuda-slice.md` §7 ("`onnx-genai-kv`/shared accounting is not the authority; a second *unaccounted* `alloc_raw` path exists for CSA device buffers") by introducing a **single property‑typed accounting authority** for CSA (`CompressedSparseAttention`) / HCA device *state groups* in the CUDA EP. ### What landed New module `crates/onnx-runtime-ep-cuda/src/kernels/csa_state_group.rs`: - **`CsaStateGroupDescriptor` — a property gate.** Parsed from the node's declared attributes + state input edges (`compression_ratio`, `cache_format`, `num_heads`/`head_dim`/`qk_rope_head_dim`, `index_head_dim`, `device_count`, presence of the compressed(6)/carry(7) state edges). **Not a model‑name/shape allowlist.** Carries the exact reason on refusal via the typed **`CsaStateGroupRefusal`** enum (UnsupportedRatio, UnknownCacheFormat, Ratio4RequiresFp8Cache, Ratio4RequiresIndexHeadDim128, Ratio128RejectsFp4, InvalidHeadGeometry, RopeExceedsHeadDim, MultiDeviceAmbiguity, MissingStateEdge, OutOfMemory, UnsupportedC1Ratio). - `validate()` accepts the op's full ratio‑4/128 range (op‑level). - `validate_c1_runtime()` is the stricter **C1 admission gate: ratio‑128 (HCA) only** — ratio‑4/MTP are op‑supported but out of C1 scope and are **typed‑refused, never silently threaded or dense‑fallback'd.** - **`CsaStateGroupLedger` — the sole, mandatory accountant.** Owns **bytes, never cursors**; charges per `(request, device)` for isolation; `try_charge` is a check‑then‑charge under one mutex that **fails closed BEFORE any physical allocation** (managed limit; `0 = unlimited`, the steady‑state default); RAII `CsaStateGroupCharge` releases on `Drop` so teardown returns to baseline. It is **instance state threaded through `CsaMetrics` (§8) — not a process‑wide global**, and it is **not** an allocator/cache/page‑manager: it accounts the bytes the existing physical path (`runtime.alloc_raw`) reserves. **No optional/no‑ledger escape hatch.** Per the *no‑backward‑compatibility during development* directive, the ledger is a **mandatory** `Arc<CsaStateGroupLedger>` on the reservation path (not an `Option`), so the **type system alone forbids an unaccounted reservation** — the exact B6 hazard. The earlier `Option<ledger>` seam was replaced outright rather than wrapped in a compat shim, and every in‑repo caller/fixture/test was updated atomically. "Disarmed" is therefore an **unlimited limit (0 sentinel)**, not a structural bypass: accounting still runs, it simply never refuses, so byte‑identity is preserved without leaving a second unaccounted path. Wiring: - `CsaDeviceBufferManager::reserve` **unconditionally** charges the ledger **before** `alloc_raw` and holds the RAII charge for the buffer lifetime (Drop frees buffers, then releases the charge). - The CUDA CSA factory `create()` builds the descriptor, `validate()`s it (typed refusal), takes a per‑runner request id, and threads the (mandatory) ledger + `(request, device)` key. - `CsaMetrics` gained the `Arc<CsaStateGroupLedger>` plus `pub` observability accessors (resident/peak/compressed/dense‑ring bytes, charge failures, active group count, per‑`(request,device)` residency, get/set limit). ### Invariants honored - **Default‑off / byte‑identical.** The charge happens **once at runner construction**, never per token and never on the captured/replay stream; release is host‑only. With the ledger unlimited (default) the reservation is byte‑identical to the pre‑accounting path. **No device op or host sync is added to steady‑state replay.** - **Fail‑closed.** Over‑limit groups are refused with a typed `OutOfMemory` **before** any device memory is reserved (no leak). - **Isolation.** Residency is tracked per `(request, device)`; teardown returns to baseline. - **Property‑based refusals** (not model‑name); cursors stay backend‑owned; PMM/VMM/governor remains the mapping/accounting authority for physical device memory (this ledger is an accounting seam **in front of** it, not a competing manager). - No PagedAttention / Mobius / QMoE / BQMoE / residency / #1971 files touched. ### Tests & measurement - **11 CPU‑only** descriptor/ledger tests: accept/refuse matrix, fail‑closed OOM without mutation, per‑`(request,device)` isolation, baseline/peak, monotonic ids, C1 ratio gate, limit round‑trip. - **6 GPU tests** on a real `CsaDeviceBufferManager` (A100‑80GB, GPU 2): `charge == sum(reserved)`, teardown → baseline, **fail‑closed‑before‑alloc with no leak**, per‑request isolation, **stable workspace addresses**, and an **unlimited‑ledger reservation that accounts without ever refusing** (`charge_failures == 0`, teardown → baseline — replacing the old no‑ledger case now that the ledger is mandatory). - **Regression:** the existing CSA op‑parity GPU suite — **26 passed / 1 ignored** (the ignored one needs external Mobius artifacts) — stays green, covering eager + capture/replay + speculative rollback + observability. This is the byte‑identical‑when‑disarmed proof. - `cargo fmt` clean; `cargo clippy -D warnings` clean on **lib and tests**. - **Overhead:** by construction there is no steady‑state kernel/host‑sync cost (accounting is one‑time at runner construction, off the captured stream), so there is no per‑token GPU delta to measure for this slice; the op‑parity capture/replay suite is the byte‑identity control. - Independent code‑review (excluding the original author and myself): **no substantive findings.** ### Honestly scoped next steps (tracked, NOT in this slice) 1. **Physical‑allocator unification:** route `CsaStateGroupLedger::try_charge` to delegate to the adopted `MemoryGovernor` `reserve(Tier::Device, bytes, MemoryRole::Workspace)` / `onnx-genai-kv` so the *physical* bytes become governor‑visible (final B6 unification). The ledger's `try_charge` is already shaped as that adapter seam. This is an architecture migration (not a compat shim), deliberately deferred. 2. **Native‑decode threading:** the smallest production caller that threads the ratio‑128 `present_* → past_*` compressed/carry state at the existing `ensure_capacity` safe boundary and calls `validate_c1_runtime`, with captured ≥16‑decode **CPU/CUDA bit‑parity** and `fallbacks == 0`. CPU op‑level present→past carry/cursor progression across the decode boundary is already covered by the CPU oracle test `ratio128_stateful_carry_matches_full_recompute_across_decode_boundary`. <sub>Co-authored-by: Copilot</sub> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… (follow-up to #2008) (#2029) ## B6.2 — Unify CSA/HCA device accounting under the shared `MemoryGovernor` **Draft. Do not merge. Independent review required (excluding the author).** Follow-up to **#2008** (B6, merged as `7a162e9f4`), now rebased directly on `main` (the #2008 dependency has landed, so this is no longer stacked — it is a clean single-commit PR on `main`). ### Problem (B6.2) #2008 introduced `CsaStateGroupLedger` as the single *typed* accountant for CSA/HCA device state groups, but it kept its **own byte limit** (`set_limit`/`limit`) beside the process-wide `MemoryGovernor`. Physical bytes were reserved through the EP's shared `CudaRuntime::alloc_raw`, yet admission was decided by a **private CSA gate** — a second set of books. That is exactly the "two allocators / two ledgers" shape the B6 design blocker exists to remove. ### Change — one accounting authority, one transaction Every CSA reservation now admits through the shared governor **before** any physical allocation: ``` governor.reserve(Tier::Device, total, MemoryRole::Workspace { step_scoped: false }, holder) -> MemoryLease // held by the RAII CsaStateGroupCharge for the buffers' lifetime ``` - CSA device residency is now counted in `MemoryGovernor::used(Tier::Device)`, in the **same books** as every other device holder (weights, KV, workspaces). - The per-`(request, device)` / per-class counters (`compressed`/`carry`/`dense_ring`/ `index`/`scratch`) remain as a **derived attribution mirror** for observability the governor deliberately does not itemise. **The mirror never gates admission**, so it is complementary telemetry, not a competing ledger. - **Single reserve / commit / rollback transaction:** a refused charge mutates nothing (fail closed with a typed `OutOfMemory` *before* `alloc_raw`); an alloc failure mid-reservation rolls back the physical buffers and releases the governor lease **exactly once** via RAII (`Drop` clears the mirror, then the `MemoryLease` field drops). - Physical allocation stays on the EP's shared `alloc_raw` inside the governed transaction — this removes the *unaccounted* property (the real B6 defect) **without** inventing pages or touching the later-built `DeviceAllocator`/`ProcessMemoryManager` teardown ordering. ### Invariant / API - **API (crate-internal):** - `CsaStateGroupLedger::new(Arc<dyn MemoryGovernor + Send + Sync>)` — governor-backed. - `CsaStateGroupLedger::default()` — unlimited reference `LedgerGovernor` (disarmed, byte-identical); `#[cfg(test)] with_device_limit(bytes)` for fail-closed proofs. - `device_available_bytes()` / `governor_device_used()` replace the removed `limit()` / `set_limit()`. - `CsaMetrics::with_governor(governor)`; `csa_state_group_device_available_bytes()` / `csa_state_group_governor_device_used()` replace the removed `*_limit` accessors. - Provider threads its process governor at registry build (`new_with_policy_governor_and_manager`); the default (no-governor) `build_cuda_registry` path stays on the unlimited reference governor. - **No backward-compat shim** (per directive): the private-limit API is deleted and every in-repo caller/test updated atomically in this commit. ### Tests (same commit, new semantics) CPU-only (no GPU): - `governor_sees_csa_reservation_in_shared_books` — a CSA charge adds to an unrelated holder's total on one governor; teardown returns to baseline; the other holder is isolated. - `charge_reserves_and_releases_governor_exactly_once` — a fake counting `MemoryGovernor` proves **one reserve, one release** (no double charge, no second path). - `ledger_reports_governor_device_ceiling`, `ledger_fails_closed_over_limit_without_mutating` — fail closed against the governor's device ceiling, no partial reservation. GPU (idle A100, correctness — no perf claim): - `reserve_charges_ledger_and_releases_on_drop`, `reservations_are_isolated_per_request`, `reserved_workspace_addresses_are_stable`, `reserve_fails_closed_over_limit_without_leaking` (via `with_device_limit`), `unlimited_ledger_accounts_without_refusing` — now assert the **physical** reservation is governor-visible (`governor_device_used()`) and returns to baseline on teardown. - **CSA op-parity + capture/replay** integration suite unchanged: **26 passed, 1 ignored** (the ignored one needs external Mobius/MTP artifacts) — this is the stable-address / capture-no-alloc / byte-identity control. ### Validation - `cargo fmt -- --check`: clean. - `cargo clippy -p onnx-runtime-ep-cuda --features cuda,cuda-13000 --lib --tests -- -D warnings`: clean. - `cargo test -p onnx-runtime-ep-cuda --features cuda,cuda-13000 --lib csa_` on idle A100: **19 passed, 0 failed**. - `--features cuda,cuda-13000,gpu-tests --release --test compressed_sparse_attention_gpu`: **26 passed, 1 ignored**. ### Scope / honest next steps (NOT in this slice) - Migrating physical allocation off `alloc_raw` onto a **governed `DeviceAllocator` handle** (subject to the strict PMM/VMM teardown ordering) is a further, separately scoped step. - **`native_decode` present→past** state threading remains the next stacked slice. - No PagedAttention / Mobius / QMoE / BQMoE / residency / #1971 files touched. - No full-model perf / tok/s claim. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
#1810 Slice 7B — boundary-time route-telemetry consumer
Status: DRAFT — awaiting independent review. Do not merge.
Closes the loop opened by the merged Slice-6/7A producer (expert-route telemetry, PR #1922
e1ec495ee) and the merged Slice-4/5 coarse-boundary residency lifecycle (PR #1854). This is the smallest production seam the Slice-6 design §8 specifies:What it does
New module
crates/onnx-runtime-ep-cuda/src/route_residency.rs:consume_route_window_at_boundary(...):COARSE_RESIDENCY_ENABLE_ENV(coarse_residency_profile_enabled()). When off (shipped default) it returnsDisabledbefore reading the snapshot or touching any allocator.CudaWeightResidency::resize_safe_pointand fails closed withRejectedNotSafeBoundary { reason }if a graph is capturing/replaying, an admission is in flight, a deferred release has not settled, execution is multi-device, or a routed-residency guard is live. Coarse-boundary only — never a per-token remap.consume_and_validateon the already-completed window snapshot; fail-closed (→WholeBank { reason }) on poison / overflow / stale epoch / foreign request / foreign device, or when no in-range experts were recorded.StaticProfileResidencyPolicyto shape aResidencyPlan(the design's "record → desired set"RouteObserverPolicyrole — reused, not duplicated, so exactly one validated policy emitsPerExpertCandidate).CudaWeightResidency::apply_coarse_residency_plan, which remains the sole authority that maps / unmaps / accounts / quarantines / rolls back through PMM/VMM.RouteWindowConsumeOutcomecarries the exact reason on every non-applied path — no silent fallback.Invariants held by construction
Tests —
tests/route_residency_consume_gpu.rs(6 GPU tests, all passing on an idle A100)disabled_gate_is_structural_no_op— off path is a structural no-op (byte-identical).route_window_hot_set_transitions_cold_experts— routed hot-set stays resident, cold set tiers to host, bytes identical.expert_group_transitions_atomically_from_window— atomic expert-group transition driven from a window.active_capture_and_multi_device_reject_consume— active-capture and multi-device boundaries reject the consume (fail-closed).foreign_identity_and_defective_windows_fail_closed— foreign request/device and poison/overflow/stale windows fail closed to whole-bank.injected_fault_rolls_back_consumer_transition— injected driver fault (fail_nth(Unmap, 3)over an isolated + merged cold range) rolls back range-precisely and quarantines;rollback_count == 1,values_touched == 0,committed_valuesempty, content bit-identical (mirrors the provencoarse_residency_plan_gpu.rsrollback fixture).Run:
Honest scope
Like
coarse_residency::apply_residency_plan_at_boundarywhen it shipped (Slice 5), this consumer has no live decode-loop call site yet — wiring it into a running session's request boundary is the next slice. It ships here as the production seam: reachable, default-off, and proven by the GPU tests.cargo fmtapplied;cargo clippyon the crate lib is clean for the new file (pre-existing unrelated warnings/errors inoptimizer.rs/standard_attention.rsare out of scope).Constraints honored
Did not touch PagedAttention or IQ1 fusion files. Branch
squad/1810-slice7b-telemetry-residency-consumebased on latestmain(011fbb284, includes #1922e1ec495ee). #1788/#1800/#1884/#1854 lifecycle regressions preserved (reused, not forked).Refs #1810.