Repository navigation
Add elastic expert-residency resize seam consumed by existing weight-paging authorities (#82) - #1789
Merged
Merged
Conversation
…paging authorities (#82) Adds a policy-issued elastic resize *intent* seam so a ResidencyPolicy can request the expert-bank residency budget grow or shrink, while the existing Resource Governor + PMM/VMM/weight-paging authorities remain the sole owners of allocation, copy, synchronization, and stable VA. New types (onnx-runtime-ep-api::weight): - ResidencyResizeRequest { direction: Grow|Shrink, target_bytes, priority } - ResizeSafePoint { capturing, pending_deferred_releases, admission_in_flight, multi_device } with is_safe()/blocking_reason() - ResizeRejection (NotSafePoint/WouldExposeColdExpert/ExecutionFailed/NoOp) - ResidencyResizePlan::{Accepted(request), Rejected { request, reason }} and the pure plan_resize(request, safe_point) validator -- no allocation, copy, or VA touched by planning. - ResidencyResizeOutcome (before/after bytes, accepted_bytes, rollback_count, safe_point, is_success()). Executor-side (onnx-runtime-ep-cuda::weight_paging): - CudaWeightResidency::resize_safe_point(device_count) builds the safe-point snapshot from the runtime's existing CUDA-graph-capture check, the provider deferred-release queue's pending count, and the existing in-flight-fill quarantine -- no new state. - CudaWeightResidency::execute_resize(plan, device_count) is the one place a resize plan becomes bytes actually moved: Grow reuses the existing MemoryLease::grow path (identical to admit_committed_span's over-budget handling); Shrink reuses the existing ReclaimableMappedHolder::reclaim_mapped when a mapped allowance is installed, or an equivalent plain-lease evict-then-shrink through the same remove_page helper otherwise. Neither mechanism is reimplemented. Re-validates the safe point immediately before committing. Refuses to shrink an ungoverned budget (no lease/allowance), mirroring set_ungoverned_budget's existing refusal. Safe-point invariant: a resize may only commit when no CUDA graph slot in use is capturing/replaying, the provider deferred-release queue has nothing pending, no page admission is in flight, and the call is not running under multi-device/TP (no existing barrier/authority coordinates a resize across devices, so this fails closed with an explicit reason rather than inventing distributed synchronization). What can actually be reclaimed this slice: only the existing dense/MatMul weight-paging cache's ordinary LRU-evictable, non-pinned resident bytes -- the same population reclaim_mapped already touches. No QMoE per-expert reclaim is enabled; whole-bank QMoE residency is unchanged (slice 2/3's no
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #1789 +/- ##
==========================================
- Coverage 80.28% 80.27% -0.01%
==========================================
Files 411 411
Lines 199976 200113 +137
Branches 199976 200113 +137
==========================================
+ Hits 160541 160644 +103
- Misses 33974 34000 +26
- Partials 5461 5469 +8
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
This was referenced Aug 23, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #82 (DeepSeek-V4/GLM-5.2 QMoE vertical slice). 4th slice of the pluggable residency-policy line: #1779 (region candidates) → #1784 (policy seam) → #1787 (migrate existing decision logic into
CudaHotSetResidencyPolicy) → this PR (elastic resize).Goal
Introduce FreeToken-style elastic expert-residency resizing as a policy-issued resize intent, executed transactionally by the existing Resource Governor + PMM/VMM/weight-paging authorities. The policy never allocates/frees/copies/relocates a virtual address — it only asks; existing infrastructure executes.
New contract (
onnx-runtime-ep-api::weight)ResidencyResizeRequest { direction: Grow | Shrink, target_bytes, priority }ResizeSafePoint { capturing, pending_deferred_releases, admission_in_flight, multi_device }withis_safe()/blocking_reason()ResizeRejection(NotSafePoint/WouldExposeColdExpert/ExecutionFailed/NoOp)ResidencyResizePlan::{Accepted(request), Rejected { request, reason }}produced by the pureplan_resize(request, safe_point)— no allocation, copy, or VA touched during planning.ResidencyResizeOutcome(before/after bytes,accepted_bytes,rollback_count,safe_point,is_success()).Executor side (
onnx-runtime-ep-cuda::weight_paging)CudaWeightResidency::resize_safe_point(device_count)builds the snapshot from existing signals only:CudaRuntime::is_capturing(), the provider deferred-release queue's pending count, and the existing in-flight-fill quarantine. No new state.CudaWeightResidency::execute_resize(plan, device_count)is the only place a plan becomes bytes actually moved:MemoryLease::growpath — the same oneadmit_committed_span's over-budget handling already uses.ReclaimableMappedHolder::reclaim_mapped(governor-driven LRU eviction + deferred-release-queue-idle wait) when a mapped allowance is installed, or an equivalent plain-lease evict-then-shrink through the sameremove_pagehelper otherwise (so process-global resident-byte/eviction telemetry stays in lockstep, not duplicated).set_ungoverned_budget's existing refusal to shrink once nothing governs the budget.Safe-point invariant
A resize may only commit when: no CUDA graph slot in use is capturing/replaying; the provider's deferred-release queue has nothing pending; no page admission is currently in flight; and the call is not running under multi-device/TP. No existing barrier/authority in this codebase coordinates a resize across devices, so multi-device fails closed with an explicit reason rather than inventing distributed synchronization (confirmed via search — no TP/multi-device coordination primitive exists yet).
What can actually be reclaimed this slice
Honest scope: only the existing dense/MatMul weight-paging cache's ordinary LRU-evictable, non-pinned resident bytes — the same population
reclaim_mappedalready touches. No QMoE per-expert reclaim is enabled. Whole-bank QMoE residency is unchanged; slice 2/3's "no cold expert may reach QMoE" invariant is preserved because nothing here changes per-expert eviction. Per-expert-aware eviction (requiring a routed-residency guarantee) is the next dependency gate.Default behavior
Unchanged:
budget()/adopt_governed_budget()/set_ungoverned_budget()are untouched, and every existingweight_pagingtest still passes without this seam ever being invoked.Tests
crates/onnx-runtime-ep-api/src/weight.rs—plan_resizeacceptance/no-op/capture/pending-release/in-flight/multi-device rejection, deterministic blocking-reason priority, outcome success reporting.crates/onnx-runtime-ep-cuda/src/weight_paging.rs— ungoverned cache refuses grow/shrink leaving budget untouched; governed grow moves exactly the requested bytes / fails over-capacity with budget unchanged; zero-byte resize is a rejected no-op; multi-device fails closed; repeated grow/shrink oscillation keeps accounting consistent; rejected-plan outcome preserves the original request's direction.Local validation (A100)
cargo fmt/cargo clippy -D warningson both touched crates — clean.cargo test -p onnx-runtime-ep-cuda --features cuda --lib— 524 passed, 0 failed, 31 ignored (GPU-only helpers not applicable), including a single-threaded rerun to rule out cross-test device contention.cargo test -p onnx-runtime-ep-api --lib— 79 passed (+ new resize tests, 88 total counting weight.rs additions).cargo test -p onnx-runtime-session --lib --features cuda— 205 passed.Independent review
A code-review agent inspected the diff before this commit and found two substantive issues, both fixed prior to committing:
remove_pagehelper so telemetry stays consistent.direction: Growregardless of the original request —ResidencyResizePlan::Rejectednow carries the original request so telemetry reports the correct direction/bytes even on rejection.Explicitly out of scope
q* CPU/GPU split, predictor-based prefetch, a new allocator/cache, and enabling per-expert QMoE eviction/reclaim (next dependency gate, once a routed-residency guarantee exists).
Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com