Repository navigation
Migrate CUDA static-pin/eviction-class decisions into ResidencyPolicy (#82) - #1787
Merged
Merged
Conversation
…#82) Introduces the first *real* ResidencyPolicy implementation, CudaHotSetResidencyPolicy, that consumes the existing scan-resistant-dense and static-pin (threshold/key-allowlist) env configuration and answers two decisions that used to be computed inline in weight_paging.rs: - eviction_class(boundary): which churn population (LRU vs. StableResident) a boundary's spans join. Previously the free function eviction_for_boundary(); now a thin delegating wrapper around the new policy so every existing self.eviction_for(boundary) call site is unaffected. - should_pin(input): whether one admitted span enters the static hot-set pin. Previously the inline `pin_this` computation in admit_committed_span, reading static_pin_keys()/static_pin_config() plus live inner.pinned/inner.pinned_bytes state directly. The policy is now a pure function of an explicit AdmissionPolicyInput snapshot (key, len_bytes, already_pinned, pinned_bytes_used); admit_committed_span remains the sole executing authority that reads live cache state and calls mark_pinned(). Victim selection (next_evictable_key/smallest_evictable/ evictable_key_by_probe), admission bookkeeping, and the byte-aware/ eviction-order diagnostic probes (#837 item 3 rejected experiment, #888 investigation) deliberately stay in CudaWeightResidency/ResidencyInner — they require live Arc::strong_count/slot-idle state this pure policy must not reach for, and both are off-by-default diagnostics rather than the shipped decision surface this slice migrates. decide() delegates to WholeBankResidentPolicy since this slice does not change per-expert placement; only eviction-class and pin-admission move behind the trait. API changes (crates/onnx-runtime-ep-api/src/weight.rs): - New EvictionClass enum (Lru | StableResident). - New AdmissionPolicyInput struct (key, len_bytes, already_pinned, pinned_bytes_used). - ResidencyPolicy trait gains eviction_class()/should_pin() with LRU/never-pin defaults, so WholeBankResidentPolicy's behavior is unchanged without touching its impl. Tests (crates/onnx-runtime-ep-cuda/src/weight_paging.rs): - hot_set_policy_eviction_class_agrees_with_old_eviction_for_boundary: regression proving old eviction_for_boundary() and the new policy's eviction_class() agree across every (scan_resistant_dense, boundary) combination exercised by this crate. - hot_set_policy_defaults_never_pin_matching_shipped_env: no pin env vars -> never pins (today's shipped default). - hot_set_policy_pin_keys_take_priority_over_threshold: explicit key allow-list wins over the threshold path, and is idempotent (already_pinned suppresses re-pin). - hot_set_policy_threshold_path_respects_budget: below-threshold, at-threshold-within-budget, and at-threshold-over-budget cases. - hot_set_policy_decide_delegates_to_whole_bank_default: decide() output is unchanged from WholeBankResidentPolicy for a pageable catalog. Local validation: onnx-runtime-ep-cuda --lib (517 passed, 31 ignored, 5 CUDA GPUs available), onnx-runtime-ep-api --lib (71 passed), onnx-runtime-session --lib --features cuda (205 passed). clippy -D warnings and fmt clean on touched crates. Independent code-review pass found no substantive issues (behavior confirmed byte-identical to the pre-migration inline logic; no missed call sites). Scope: default no-offload and existing offload behavior are unchanged (byte/handle/H2D/allocation-identical); no cold expert reaches QMoE. Not in scope for this slice: q* CPU/GPU execution, elastic in-place rebuild, wiring plan_double_buffer into production dispatch (still deferred — prefetch_lazy_weights_after's node-lookahead window is the only production prefetch mechanism today and was left untouched), and Resource Governor shrink/grow request semantics (governor budget/lending inputs were traced but no new decision surface was added this slice; deferred to the elastic in-place rebuild slice that follows). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #1787 +/- ##
==========================================
+ Coverage 80.66% 80.79% +0.13%
==========================================
Files 411 411
Lines 199922 199976 +54
Branches 199922 199976 +54
==========================================
+ Hits 161271 161579 +308
+ Misses 33190 32938 -252
+ Partials 5461 5459 -2
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
justinchuby
added a commit
that referenced
this pull request
Aug 23, 2026
…paging authorities (#82) (#1789) Part of #82 (DeepSeek-V4/GLM-5.2 QMoE vertical slice). 4th slice of the pluggable residency-policy line: #1779 (region candidates) → #1784 (policy seam) → #1787 (migrate existing decision logic into `CudaHotSetResidencyPolicy`) → this PR (elastic resize). ## Goal Introduce FreeToken-style elastic expert-residency **resizing** as a policy-issued resize *intent*, executed transactionally by the existing Resource Governor + PMM/VMM/weight-paging authorities. The policy never allocates/frees/copies/relocates a virtual address — it only asks; existing infrastructure executes. ## New contract (`onnx-runtime-ep-api::weight`) - `ResidencyResizeRequest { direction: Grow | Shrink, target_bytes, priority }` - `ResizeSafePoint { capturing, pending_deferred_releases, admission_in_flight, multi_device }` with `is_safe()` / `blocking_reason()` - `ResizeRejection` (`NotSafePoint` / `WouldExposeColdExpert` / `ExecutionFailed` / `NoOp`) - `ResidencyResizePlan::{Accepted(request), Rejected { request, reason }}` produced by the pure `plan_resize(request, safe_point)` — no allocation, copy, or VA touched during planning. - `ResidencyResizeOutcome` (before/after bytes, `accepted_bytes`, `rollback_count`, `safe_point`, `is_success()`). ## Executor side (`onnx-runtime-ep-cuda::weight_paging`) - `CudaWeightResidency::resize_safe_point(device_count)` builds the snapshot from **existing** signals only: `CudaRuntime::is_capturing()`, the provider deferred-release queue's pending count, and the existing in-flight-fill quarantine. No new state. - `CudaWeightResidency::execute_resize(plan, device_count)` is the *only* place a plan becomes bytes actually moved: - **Grow** reuses the existing `MemoryLease::grow` path — the same one `admit_committed_span`'s over-budget handling already uses. - **Shrink** reuses the existing `ReclaimableMappedHolder::reclaim_mapped` (governor-driven LRU eviction + deferred-release-queue-idle wait) when a mapped allowance is installed, or an equivalent plain-lease evict-then-shrink through the same `remove_page` helper otherwise (so process-global resident-byte/eviction telemetry stays in lockstep, not duplicated). - Re-validates the safe point immediately before committing (planning and execution can be separated in time). - Refuses to shrink an *ungoverned* budget (no lease/allowance) — mirrors `set_ungoverned_budget`'s existing refusal to shrink once nothing governs the budget. ## Safe-point invariant A resize may only commit when: no CUDA graph slot in use is capturing/replaying; the provider's deferred-release queue has nothing pending; no page admission is currently in flight; and the call is **not** running under multi-device/TP. No existing barrier/authority in this codebase coordinates a resize across devices, so multi-device fails closed with an explicit reason rather than inventing distributed synchronization (confirmed via search — no TP/multi-device coordination primitive exists yet). ## What can actually be reclaimed this slice Honest scope: only the existing dense/MatMul weight-paging cache's ordinary LRU-evictable, non-pinned resident bytes — the same population `reclaim_mapped` already touches. **No QMoE per-expert reclaim is enabled.** Whole-bank QMoE residency is unchanged; slice 2/3's "no cold expert may reach QMoE" invariant is preserved because nothing here changes per-expert eviction. Per-expert-aware eviction (requiring a routed-residency guarantee) is the next dependency gate. ## Default behavior Unchanged: `budget()`/`adopt_governed_budget()`/`set_ungoverned_budget()` are untouched, and every existing `weight_paging` test still passes without this seam ever being invoked. ## Tests - `crates/onnx-runtime-ep-api/src/weight.rs` — `plan_resize` acceptance/no-op/capture/pending-release/in-flight/multi-device rejection, deterministic blocking-reason priority, outcome success reporting. - `crates/onnx-runtime-ep-cuda/src/weight_paging.rs` — ungoverned cache refuses grow/shrink leaving budget untouched; governed grow moves exactly the requested bytes / fails over-capacity with budget unchanged; zero-byte resize is a rejected no-op; multi-device fails closed; repeated grow/shrink oscillation keeps accounting consistent; rejected-plan outcome preserves the original request's direction. ## Local validation (A100) - `cargo fmt` / `cargo clippy -D warnings` on both touched crates — clean. - `cargo test -p onnx-runtime-ep-cuda --features cuda --lib` — 524 passed, 0 failed, 31 ignored (GPU-only helpers not applicable), including a single-threaded rerun to rule out cross-test device contention. - `cargo test -p onnx-runtime-ep-api --lib` — 79 passed (+ new resize tests, 88 total counting weight.rs additions). - `cargo test -p onnx-runtime-session --lib --features cuda` — 205 passed. ## Independent review A code-review agent inspected the diff before this commit and found two substantive issues, both fixed prior to committing: 1. The plain-lease shrink fallback updated only per-instance accounting and skipped the process-global resident-byte/eviction gauges every other eviction call site updates — now routed through the existing `remove_page` helper so telemetry stays consistent. 2. A plan-time-rejected outcome always reported `direction: Grow` regardless of the original request — `ResidencyResizePlan::Rejected` now carries the original request so telemetry reports the correct direction/bytes even on rejection. ## Explicitly out of scope q* CPU/GPU split, predictor-based prefetch, a new allocator/cache, and enabling per-expert QMoE eviction/reclaim (next dependency gate, once a routed-residency guarantee exists). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This was referenced Aug 23, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #82 — this migrates the existing static-pin/boundary-eviction-class decision logic in
crates/onnx-runtime-ep-cuda/src/weight_paging.rsinto one concreteResidencyPolicyimplementation (CudaHotSetResidencyPolicy), removing duplicated decision authority rather than wrapping it. Issue #82 stays open — this is one slice, not the final state.What moved behind the policy trait
eviction_class(boundary): which churn population (LruvsStableResident) a boundary's spans join. Previously the free functioneviction_for_boundary(scan_resistant_dense, boundary); that function now delegates to the policy (kept as a thin wrapper since it's still called from everyself.eviction_for(boundary)admission site — no call site needed to change).should_pin(input): whether one admitted span enters the static hot-set pin. Previously the inlinepin_thiscomputation insideadmit_committed_span, which read the module-levelstatic_pin_keys()/static_pin_config()caches and reached into liveinner.pinned/inner.pinned_bytescache state directly. It's now a pure function of an explicitAdmissionPolicyInputsnapshot (key,len_bytes,already_pinned,pinned_bytes_used);admit_committed_spanremains the sole place that reads live state and callsmark_pinned().What stays where it was, and why
Per the architecture invariant (PMM/VMM allocation/accounting, VA ownership, and synchronization are non-pluggable authorities), the following are not migrated:
next_evictable_key/smallest_evictable/evictable_key_by_probe) — needs liveArc::strong_count/slot-idle runtime state a pure policy must not hold.mark_pinned— execution/bookkeeping, not a decision.byte_aware(Large-model decode is streaming-bound: 0.24 tok/s on qwen14b, with 57k VMM map/unmap per 16 tokens and ~4.9 GB/s effective H2D #837 item 3, a rejected experiment CUDA provider construction never enables) andevict_order_probe(Does decode correctness depend on eviction ORDER? Byte-aware residency corrupted output (#886) — is that the change, or a latent assumption in the offload path? #888 diagnostic probe) — both off-by-default diagnostics, not the shipped decision surface this slice targets.plan_double_buffer/drive_double_buffer(session crate) — traced this slice; confirmed it is not the production prefetch mechanism (prefetch_lazy_weights_after's node-lookahead window is). Left untouched per "wire only where an existing valid window already exists" — no new predictive wiring was invented.decide()on the new policy delegates toWholeBankResidentPolicy— this slice does not change per-expert placement, only eviction-class and pin-admission decisions move behind the trait.API changes (
crates/onnx-runtime-ep-api/src/weight.rs)EvictionClassenum (Lru | StableResident).AdmissionPolicyInputstruct (key,len_bytes,already_pinned,pinned_bytes_used).ResidencyPolicytrait gainseviction_class()/should_pin()with LRU/never-pin defaults, soWholeBankResidentPolicy's behavior is unchanged without touching itsimpl.Behavior gates preserved
(scan_resistant_dense, boundary)combination, plus tests for pin-keys-over-threshold priority, idempotent re-pin suppression, and budget-respecting threshold admission.Tests added (
crates/onnx-runtime-ep-cuda/src/weight_paging.rs)hot_set_policy_eviction_class_agrees_with_old_eviction_for_boundary— old/new agreement regression across all boundaries × scan-resistant-dense settings.hot_set_policy_defaults_never_pin_matching_shipped_env— no pin env vars set never pins.hot_set_policy_pin_keys_take_priority_over_threshold— explicit key allow-list wins; already-pinned is never re-pinned.hot_set_policy_threshold_path_respects_budget— below-threshold / at-threshold-within-budget / at-threshold-over-budget.hot_set_policy_decide_delegates_to_whole_bank_default—decide()output unchanged from the default policy.Local validation
cargo test -p onnx-runtime-ep-cuda --features cuda --lib: 517 passed, 31 ignored (5× A100-SXM4-80GB available).cargo test -p onnx-runtime-ep-api --lib: 71 passed.cargo test -p onnx-runtime-session --lib --features cuda: 205 passed.cargo clippy -p onnx-runtime-ep-cuda --features cuda --lib -- -D warningsandcargo fmtclean on touched crates.Not in scope for this slice
q* CPU/GPU execution, elastic in-place rebuild,
plan_double_bufferproduction wiring, LRU/static-hot policy replacement (this slice only relocates existing decisions, it doesn't change them), Resource Governor shrink/grow semantics (traced only). These remain for follow-up slices.Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com