Skip to content

Migrate CUDA static-pin/eviction-class decisions into ResidencyPolicy (#82) - #1787

Merged
justinchuby merged 1 commit into
mainfrom
squad/82-qmoe-residency-policy-migration
Aug 22, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/82-qmoe-residency-policy-migration

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Part of #82 — this migrates the existing static-pin/boundary-eviction-class decision logic in crates/onnx-runtime-ep-cuda/src/weight_paging.rs into one concrete ResidencyPolicy implementation (CudaHotSetResidencyPolicy), removing duplicated decision authority rather than wrapping it. Issue #82 stays open — this is one slice, not the final state.

What moved behind the policy trait

  • eviction_class(boundary): which churn population (Lru vs StableResident) a boundary's spans join. Previously the free function eviction_for_boundary(scan_resistant_dense, boundary); that function now delegates to the policy (kept as a thin wrapper since it's still called from every self.eviction_for(boundary) admission site — no call site needed to change).
  • should_pin(input): whether one admitted span enters the static hot-set pin. Previously the inline pin_this computation inside admit_committed_span, which read the module-level static_pin_keys()/static_pin_config() caches and reached into live inner.pinned/inner.pinned_bytes cache state directly. It's now a pure function of an explicit AdmissionPolicyInput snapshot (key, len_bytes, already_pinned, pinned_bytes_used); admit_committed_span remains the sole place that reads live state and calls mark_pinned().

What stays where it was, and why

Per the architecture invariant (PMM/VMM allocation/accounting, VA ownership, and synchronization are non-pluggable authorities), the following are not migrated:

decide() on the new policy delegates to WholeBankResidentPolicy — this slice does not change per-expert placement, only eviction-class and pin-admission decisions move behind the trait.

API changes (crates/onnx-runtime-ep-api/src/weight.rs)

  • EvictionClass enum (Lru | StableResident).
  • AdmissionPolicyInput struct (key, len_bytes, already_pinned, pinned_bytes_used).
  • ResidencyPolicy trait gains eviction_class()/should_pin() with LRU/never-pin defaults, so WholeBankResidentPolicy's behavior is unchanged without touching its impl.

Behavior gates preserved

  • Default no-offload behavior: byte/handle/H2D/allocation-identical (untouched call sites, pure delegation).
  • Existing offload behavior (scan-resistant-dense, static pin threshold/key-allowlist): correctness-identical — proved via a regression test comparing old and new decision outputs across every exercised (scan_resistant_dense, boundary) combination, plus tests for pin-keys-over-threshold priority, idempotent re-pin suppression, and budget-respecting threshold admission.
  • No cold expert reaches QMoE; the whole-bank safety guard is untouched.

Tests added (crates/onnx-runtime-ep-cuda/src/weight_paging.rs)

  • hot_set_policy_eviction_class_agrees_with_old_eviction_for_boundary — old/new agreement regression across all boundaries × scan-resistant-dense settings.
  • hot_set_policy_defaults_never_pin_matching_shipped_env — no pin env vars set never pins.
  • hot_set_policy_pin_keys_take_priority_over_threshold — explicit key allow-list wins; already-pinned is never re-pinned.
  • hot_set_policy_threshold_path_respects_budget — below-threshold / at-threshold-within-budget / at-threshold-over-budget.
  • hot_set_policy_decide_delegates_to_whole_bank_default — decide() output unchanged from the default policy.

Local validation

  • cargo test -p onnx-runtime-ep-cuda --features cuda --lib: 517 passed, 31 ignored (5× A100-SXM4-80GB available).
  • cargo test -p onnx-runtime-ep-api --lib: 71 passed.
  • cargo test -p onnx-runtime-session --lib --features cuda: 205 passed.
  • cargo clippy -p onnx-runtime-ep-cuda --features cuda --lib -- -D warnings and cargo fmt clean on touched crates.
  • Independent code-review pass: no substantive findings (confirmed byte-identical behavior vs. the pre-migration inline logic, no missed call sites).

Not in scope for this slice

q* CPU/GPU execution, elastic in-place rebuild, plan_double_buffer production wiring, LRU/static-hot policy replacement (this slice only relocates existing decisions, it doesn't change them), Resource Governor shrink/grow semantics (traced only). These remain for follow-up slices.

Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com

…#82)

Introduces the first *real* ResidencyPolicy implementation,
CudaHotSetResidencyPolicy, that consumes the existing scan-resistant-dense
and static-pin (threshold/key-allowlist) env configuration and answers two
decisions that used to be computed inline in weight_paging.rs:

- eviction_class(boundary): which churn population (LRU vs. StableResident)
  a boundary's spans join. Previously the free function
  eviction_for_boundary(); now a thin delegating wrapper around the new
  policy so every existing self.eviction_for(boundary) call site is
  unaffected.
- should_pin(input): whether one admitted span enters the static hot-set
  pin. Previously the inline `pin_this` computation in
  admit_committed_span, reading static_pin_keys()/static_pin_config()
  plus live inner.pinned/inner.pinned_bytes state directly. The policy is
  now a pure function of an explicit AdmissionPolicyInput snapshot
  (key, len_bytes, already_pinned, pinned_bytes_used); admit_committed_span
  remains the sole executing authority that reads live cache state and
  calls mark_pinned().

Victim selection (next_evictable_key/smallest_evictable/
evictable_key_by_probe), admission bookkeeping, and the byte-aware/
eviction-order diagnostic probes (#837 item 3 rejected experiment, #888
investigation) deliberately stay in CudaWeightResidency/ResidencyInner —
they require live Arc::strong_count/slot-idle state this pure policy must
not reach for, and both are off-by-default diagnostics rather than the
shipped decision surface this slice migrates.

decide() delegates to WholeBankResidentPolicy since this slice does not
change per-expert placement; only eviction-class and pin-admission move
behind the trait.

API changes (crates/onnx-runtime-ep-api/src/weight.rs):
- New EvictionClass enum (Lru | StableResident).
- New AdmissionPolicyInput struct (key, len_bytes, already_pinned,
  pinned_bytes_used).
- ResidencyPolicy trait gains eviction_class()/should_pin() with
  LRU/never-pin defaults, so WholeBankResidentPolicy's behavior is
  unchanged without touching its impl.

Tests (crates/onnx-runtime-ep-cuda/src/weight_paging.rs):
- hot_set_policy_eviction_class_agrees_with_old_eviction_for_boundary:
  regression proving old eviction_for_boundary() and the new policy's
  eviction_class() agree across every (scan_resistant_dense, boundary)
  combination exercised by this crate.
- hot_set_policy_defaults_never_pin_matching_shipped_env: no pin env vars
  -> never pins (today's shipped default).
- hot_set_policy_pin_keys_take_priority_over_threshold: explicit key
  allow-list wins over the threshold path, and is idempotent
  (already_pinned suppresses re-pin).
- hot_set_policy_threshold_path_respects_budget: below-threshold,
  at-threshold-within-budget, and at-threshold-over-budget cases.
- hot_set_policy_decide_delegates_to_whole_bank_default: decide() output
  is unchanged from WholeBankResidentPolicy for a pageable catalog.

Local validation: onnx-runtime-ep-cuda --lib (517 passed, 31 ignored,
5 CUDA GPUs available), onnx-runtime-ep-api --lib (71 passed),
onnx-runtime-session --lib --features cuda (205 passed). clippy -D
warnings and fmt clean on touched crates. Independent code-review pass
found no substantive issues (behavior confirmed byte-identical to the
pre-migration inline logic; no missed call sites).

Scope: default no-offload and existing offload behavior are unchanged
(byte/handle/H2D/allocation-identical); no cold expert reaches QMoE. Not
in scope for this slice: q* CPU/GPU execution, elastic in-place rebuild,
wiring plan_double_buffer into production dispatch (still deferred —
prefetch_lazy_weights_after's node-lookahead window is the only production
prefetch mechanism today and was left untouched), and Resource Governor
shrink/grow request semantics (governor budget/lending inputs were traced
but no new decision surface was added this slice; deferred to the elastic
in-place rebuild slice that follows).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby merged commit fc6fbfc into main Aug 22, 2026
4 of 5 checks passed
@justinchuby
justinchuby deleted the squad/82-qmoe-residency-policy-migration branch August 22, 2026 23:20
@codecov

codecov Bot commented Aug 22, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 0% with 8 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.79%. Comparing base (ac9c7dd) to head (4dd59bd).
⚠️ Report is 5 commits behind head on main.

Files with missing lines Patch % Lines
crates/onnx-runtime-ep-api/src/weight.rs 0.00% 8 Missing ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1787      +/-   ##
==========================================
+ Coverage   80.66%   80.79%   +0.13%     
==========================================
  Files         411      411              
  Lines      199922   199976      +54     
  Branches   199922   199976      +54     
==========================================
+ Hits       161271   161579     +308     
+ Misses      33190    32938     -252     
+ Partials     5461     5459       -2     
Flag Coverage Δ
cli-ort-linux 72.47% <ø> (ø)
cli-ort-windows 72.06% <ø> (+0.09%) ⬆️
mlas 85.20% <ø> (+0.09%) ⬆️
offline 80.95% <0.00%> (+0.13%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-runtime-ep-api/src/weight.rs 82.67% <0.00%> (-1.16%) ⬇️

... and 6 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

justinchuby added a commit that referenced this pull request Aug 23, 2026
…paging authorities (#82) (#1789)

Part of #82 (DeepSeek-V4/GLM-5.2 QMoE vertical slice). 4th slice of the
pluggable residency-policy line: #1779 (region candidates) → #1784
(policy seam) → #1787 (migrate existing decision logic into
`CudaHotSetResidencyPolicy`) → this PR (elastic resize).

## Goal
Introduce FreeToken-style elastic expert-residency **resizing** as a
policy-issued resize *intent*, executed transactionally by the existing
Resource Governor + PMM/VMM/weight-paging authorities. The policy never
allocates/frees/copies/relocates a virtual address — it only asks;
existing infrastructure executes.

## New contract (`onnx-runtime-ep-api::weight`)
- `ResidencyResizeRequest { direction: Grow | Shrink, target_bytes,
priority }`
- `ResizeSafePoint { capturing, pending_deferred_releases,
admission_in_flight, multi_device }` with `is_safe()` /
`blocking_reason()`
- `ResizeRejection` (`NotSafePoint` / `WouldExposeColdExpert` /
`ExecutionFailed` / `NoOp`)
- `ResidencyResizePlan::{Accepted(request), Rejected { request, reason
}}` produced by the pure `plan_resize(request, safe_point)` — no
allocation, copy, or VA touched during planning.
- `ResidencyResizeOutcome` (before/after bytes, `accepted_bytes`,
`rollback_count`, `safe_point`, `is_success()`).

## Executor side (`onnx-runtime-ep-cuda::weight_paging`)
- `CudaWeightResidency::resize_safe_point(device_count)` builds the
snapshot from **existing** signals only: `CudaRuntime::is_capturing()`,
the provider deferred-release queue's pending count, and the existing
in-flight-fill quarantine. No new state.
- `CudaWeightResidency::execute_resize(plan, device_count)` is the
*only* place a plan becomes bytes actually moved:
- **Grow** reuses the existing `MemoryLease::grow` path — the same one
`admit_committed_span`'s over-budget handling already uses.
- **Shrink** reuses the existing
`ReclaimableMappedHolder::reclaim_mapped` (governor-driven LRU eviction
+ deferred-release-queue-idle wait) when a mapped allowance is
installed, or an equivalent plain-lease evict-then-shrink through the
same `remove_page` helper otherwise (so process-global
resident-byte/eviction telemetry stays in lockstep, not duplicated).
- Re-validates the safe point immediately before committing (planning
and execution can be separated in time).
- Refuses to shrink an *ungoverned* budget (no lease/allowance) —
mirrors `set_ungoverned_budget`'s existing refusal to shrink once
nothing governs the budget.

## Safe-point invariant
A resize may only commit when: no CUDA graph slot in use is
capturing/replaying; the provider's deferred-release queue has nothing
pending; no page admission is currently in flight; and the call is
**not** running under multi-device/TP. No existing barrier/authority in
this codebase coordinates a resize across devices, so multi-device fails
closed with an explicit reason rather than inventing distributed
synchronization (confirmed via search — no TP/multi-device coordination
primitive exists yet).

## What can actually be reclaimed this slice
Honest scope: only the existing dense/MatMul weight-paging cache's
ordinary LRU-evictable, non-pinned resident bytes — the same population
`reclaim_mapped` already touches. **No QMoE per-expert reclaim is
enabled.** Whole-bank QMoE residency is unchanged; slice 2/3's "no cold
expert may reach QMoE" invariant is preserved because nothing here
changes per-expert eviction. Per-expert-aware eviction (requiring a
routed-residency guarantee) is the next dependency gate.

## Default behavior
Unchanged:
`budget()`/`adopt_governed_budget()`/`set_ungoverned_budget()` are
untouched, and every existing `weight_paging` test still passes without
this seam ever being invoked.

## Tests
- `crates/onnx-runtime-ep-api/src/weight.rs` — `plan_resize`
acceptance/no-op/capture/pending-release/in-flight/multi-device
rejection, deterministic blocking-reason priority, outcome success
reporting.
- `crates/onnx-runtime-ep-cuda/src/weight_paging.rs` — ungoverned cache
refuses grow/shrink leaving budget untouched; governed grow moves
exactly the requested bytes / fails over-capacity with budget unchanged;
zero-byte resize is a rejected no-op; multi-device fails closed;
repeated grow/shrink oscillation keeps accounting consistent;
rejected-plan outcome preserves the original request's direction.

## Local validation (A100)
- `cargo fmt` / `cargo clippy -D warnings` on both touched crates —
clean.
- `cargo test -p onnx-runtime-ep-cuda --features cuda --lib` — 524
passed, 0 failed, 31 ignored (GPU-only helpers not applicable),
including a single-threaded rerun to rule out cross-test device
contention.
- `cargo test -p onnx-runtime-ep-api --lib` — 79 passed (+ new resize
tests, 88 total counting weight.rs additions).
- `cargo test -p onnx-runtime-session --lib --features cuda` — 205
passed.

## Independent review
A code-review agent inspected the diff before this commit and found two
substantive issues, both fixed prior to committing:
1. The plain-lease shrink fallback updated only per-instance accounting
and skipped the process-global resident-byte/eviction gauges every other
eviction call site updates — now routed through the existing
`remove_page` helper so telemetry stays consistent.
2. A plan-time-rejected outcome always reported `direction: Grow`
regardless of the original request — `ResidencyResizePlan::Rejected` now
carries the original request so telemetry reports the correct
direction/bytes even on rejection.

## Explicitly out of scope
q* CPU/GPU split, predictor-based prefetch, a new allocator/cache, and
enabling per-expert QMoE eviction/reclaim (next dependency gate, once a
routed-residency guarantee exists).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant