Skip to content

Add elastic expert-residency resize seam consumed by existing weight-paging authorities (#82) - #1789

Merged
justinchuby merged 1 commit into
mainfrom
squad/82-qmoe-residency-elastic-resize
Aug 23, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/82-qmoe-residency-elastic-resize

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Part of #82 (DeepSeek-V4/GLM-5.2 QMoE vertical slice). 4th slice of the pluggable residency-policy line: #1779 (region candidates) → #1784 (policy seam) → #1787 (migrate existing decision logic into CudaHotSetResidencyPolicy) → this PR (elastic resize).

Goal

Introduce FreeToken-style elastic expert-residency resizing as a policy-issued resize intent, executed transactionally by the existing Resource Governor + PMM/VMM/weight-paging authorities. The policy never allocates/frees/copies/relocates a virtual address — it only asks; existing infrastructure executes.

New contract (onnx-runtime-ep-api::weight)

  • ResidencyResizeRequest { direction: Grow | Shrink, target_bytes, priority }
  • ResizeSafePoint { capturing, pending_deferred_releases, admission_in_flight, multi_device } with is_safe() / blocking_reason()
  • ResizeRejection (NotSafePoint / WouldExposeColdExpert / ExecutionFailed / NoOp)
  • ResidencyResizePlan::{Accepted(request), Rejected { request, reason }} produced by the pure plan_resize(request, safe_point) — no allocation, copy, or VA touched during planning.
  • ResidencyResizeOutcome (before/after bytes, accepted_bytes, rollback_count, safe_point, is_success()).

Executor side (onnx-runtime-ep-cuda::weight_paging)

  • CudaWeightResidency::resize_safe_point(device_count) builds the snapshot from existing signals only: CudaRuntime::is_capturing(), the provider deferred-release queue's pending count, and the existing in-flight-fill quarantine. No new state.
  • CudaWeightResidency::execute_resize(plan, device_count) is the only place a plan becomes bytes actually moved:
    • Grow reuses the existing MemoryLease::grow path — the same one admit_committed_span's over-budget handling already uses.
    • Shrink reuses the existing ReclaimableMappedHolder::reclaim_mapped (governor-driven LRU eviction + deferred-release-queue-idle wait) when a mapped allowance is installed, or an equivalent plain-lease evict-then-shrink through the same remove_page helper otherwise (so process-global resident-byte/eviction telemetry stays in lockstep, not duplicated).
    • Re-validates the safe point immediately before committing (planning and execution can be separated in time).
    • Refuses to shrink an ungoverned budget (no lease/allowance) — mirrors set_ungoverned_budget's existing refusal to shrink once nothing governs the budget.

Safe-point invariant

A resize may only commit when: no CUDA graph slot in use is capturing/replaying; the provider's deferred-release queue has nothing pending; no page admission is currently in flight; and the call is not running under multi-device/TP. No existing barrier/authority in this codebase coordinates a resize across devices, so multi-device fails closed with an explicit reason rather than inventing distributed synchronization (confirmed via search — no TP/multi-device coordination primitive exists yet).

What can actually be reclaimed this slice

Honest scope: only the existing dense/MatMul weight-paging cache's ordinary LRU-evictable, non-pinned resident bytes — the same population reclaim_mapped already touches. No QMoE per-expert reclaim is enabled. Whole-bank QMoE residency is unchanged; slice 2/3's "no cold expert may reach QMoE" invariant is preserved because nothing here changes per-expert eviction. Per-expert-aware eviction (requiring a routed-residency guarantee) is the next dependency gate.

Default behavior

Unchanged: budget()/adopt_governed_budget()/set_ungoverned_budget() are untouched, and every existing weight_paging test still passes without this seam ever being invoked.

Tests

  • crates/onnx-runtime-ep-api/src/weight.rs — plan_resize acceptance/no-op/capture/pending-release/in-flight/multi-device rejection, deterministic blocking-reason priority, outcome success reporting.
  • crates/onnx-runtime-ep-cuda/src/weight_paging.rs — ungoverned cache refuses grow/shrink leaving budget untouched; governed grow moves exactly the requested bytes / fails over-capacity with budget unchanged; zero-byte resize is a rejected no-op; multi-device fails closed; repeated grow/shrink oscillation keeps accounting consistent; rejected-plan outcome preserves the original request's direction.

Local validation (A100)

  • cargo fmt / cargo clippy -D warnings on both touched crates — clean.
  • cargo test -p onnx-runtime-ep-cuda --features cuda --lib — 524 passed, 0 failed, 31 ignored (GPU-only helpers not applicable), including a single-threaded rerun to rule out cross-test device contention.
  • cargo test -p onnx-runtime-ep-api --lib — 79 passed (+ new resize tests, 88 total counting weight.rs additions).
  • cargo test -p onnx-runtime-session --lib --features cuda — 205 passed.

Independent review

A code-review agent inspected the diff before this commit and found two substantive issues, both fixed prior to committing:

  1. The plain-lease shrink fallback updated only per-instance accounting and skipped the process-global resident-byte/eviction gauges every other eviction call site updates — now routed through the existing remove_page helper so telemetry stays consistent.
  2. A plan-time-rejected outcome always reported direction: Grow regardless of the original request — ResidencyResizePlan::Rejected now carries the original request so telemetry reports the correct direction/bytes even on rejection.

Explicitly out of scope

q* CPU/GPU split, predictor-based prefetch, a new allocator/cache, and enabling per-expert QMoE eviction/reclaim (next dependency gate, once a routed-residency guarantee exists).

Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com

…paging authorities (#82)

Adds a policy-issued elastic resize *intent* seam so a ResidencyPolicy can
request the expert-bank residency budget grow or shrink, while the existing
Resource Governor + PMM/VMM/weight-paging authorities remain the sole owners
of allocation, copy, synchronization, and stable VA.

New types (onnx-runtime-ep-api::weight):
- ResidencyResizeRequest { direction: Grow|Shrink, target_bytes, priority }
- ResizeSafePoint { capturing, pending_deferred_releases,
  admission_in_flight, multi_device } with is_safe()/blocking_reason()
- ResizeRejection (NotSafePoint/WouldExposeColdExpert/ExecutionFailed/NoOp)
- ResidencyResizePlan::{Accepted(request), Rejected { request, reason }} and
  the pure plan_resize(request, safe_point) validator -- no allocation,
  copy, or VA touched by planning.
- ResidencyResizeOutcome (before/after bytes, accepted_bytes, rollback_count,
  safe_point, is_success()).

Executor-side (onnx-runtime-ep-cuda::weight_paging):
- CudaWeightResidency::resize_safe_point(device_count) builds the safe-point
  snapshot from the runtime's existing CUDA-graph-capture check, the
  provider deferred-release queue's pending count, and the existing
  in-flight-fill quarantine -- no new state.
- CudaWeightResidency::execute_resize(plan, device_count) is the one place a
  resize plan becomes bytes actually moved: Grow reuses the existing
  MemoryLease::grow path (identical to admit_committed_span's over-budget
  handling); Shrink reuses the existing ReclaimableMappedHolder::reclaim_mapped
  when a mapped allowance is installed, or an equivalent plain-lease
  evict-then-shrink through the same remove_page helper otherwise. Neither
  mechanism is reimplemented. Re-validates the safe point immediately before
  committing. Refuses to shrink an ungoverned budget (no lease/allowance),
  mirroring set_ungoverned_budget's existing refusal.

Safe-point invariant: a resize may only commit when no CUDA graph slot in
use is capturing/replaying, the provider deferred-release queue has nothing
pending, no page admission is in flight, and the call is not running under
multi-device/TP (no existing barrier/authority coordinates a resize across
devices, so this fails closed with an explicit reason rather than inventing
distributed synchronization).

What can actually be reclaimed this slice: only the existing dense/MatMul
weight-paging cache's ordinary LRU-evictable, non-pinned resident bytes --
the same population reclaim_mapped already touches. No QMoE per-expert
reclaim is enabled; whole-bank QMoE residency is unchanged (slice 2/3's no
@justinchuby
justinchuby merged commit 10dfcd7 into main Aug 23, 2026
14 of 16 checks passed
@justinchuby
justinchuby deleted the squad/82-qmoe-residency-elastic-resize branch August 23, 2026 00:22
@codecov

codecov Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 94.92754% with 7 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.27%. Comparing base (fc6fbfc) to head (4f92b87).
⚠️ Report is 3 commits behind head on main.

Files with missing lines Patch % Lines
crates/onnx-runtime-ep-api/src/weight.rs 94.92% 4 Missing and 3 partials ⚠️
Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main    #1789      +/-   ##
==========================================
- Coverage   80.28%   80.27%   -0.01%     
==========================================
  Files         411      411              
  Lines      199976   200113     +137     
  Branches   199976   200113     +137     
==========================================
+ Hits       160541   160644     +103     
- Misses      33974    34000      +26     
- Partials     5461     5469       +8     
Flag Coverage Δ
cli-ort-linux ?
cli-ort-windows 72.06% <ø> (+0.09%) ⬆️
mlas 85.10% <ø> (ø)
offline 80.42% <94.92%> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
crates/onnx-runtime-ep-api/src/weight.rs 85.02% <94.92%> (+2.34%) ⬆️

... and 5 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant