Repository navigation
vmm(#1295): usable-VRAM admission cap (cliff->plateau) + vram_free_ms attribution fix - #1325
Merged
Merged
Conversation
… + fix vram_free_ms attribution Issue #1295 (follow-up to #1315). Converts the batch-N streaming cliff from a driver OOM crash into a clean governed refusal by clamping the device-authority ledger ceiling to measured usable VRAM (cuMemGetInfo free x fraction, default 0.90) instead of nominal total. This is a cliff->plateau conversion, NOT a raised N_max: the ledger now refuses the over-subscribing lease before the WDDM driver falls off the fault-in cliff. - governor.rs: clamp_ceiling_to_usable_vram() + usable_mapped_safe_fraction(); DEFAULT_VMM_MAPPED_FRACTION=0.90 (env ONNX_GENAI_VMM_MAPPED_FRACTION), sourced from capacities.vram.free_bytes(); applied at device_ceiling_bytes with tracing::info when it binds. Single shared authority (MoE #1308 reuses it). - Instrumentation fix: release_allocation() timed stream-drain sync inside the GLOBAL_VRAM_FREE_NS bracket, mis-attributing paging-slowed work to "free". Split into GLOBAL_VRAM_FREE_SYNC_NS (surfaced as vram_free_sync_ms). Corrected: sync is ~4% of the bracket; the ~45s free time is genuine unmap/release churn. Headline not substantially retracted. - 3 governor unit tests (fraction validation, usable-sourced clamp, non-vacuous G3 refusal). Benchmark doc with measured/inferred split. Measured: RTX 4060 Laptop 8GB, driver 591.55, CUDA 13.1, WDDM, i7-13800H, single-thread, model qwen14b-zp. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 18, 2026
…ts (not commit-and-subdivide) (#1332) ## #1295 follow-up — MoE VMM sub-granule packing: the trade curve (analysis, no impl) **Reporting before implementing, as the owner asked.** This answers "can VMM commit more at once and subdivide ourselves?" from the already-measured granite router-skew distribution — no packing code was written. ### The question and the trap The 2 MiB CUDA granule wastes **2.7×** on granite int4 experts (~0.75 MiB). The proposed escape — commit a large region (tens of MiB), pack ~85 experts, subdivide ourselves — is favourable on granule *waste* alone. But committing is also the unit of *residency*, so the real question is whether sub-division keeps selectivity. ### Finding (measured, granite-3.0-1b-a400m, 32e/top-8/24L, real IBM router, CPU EP, 192 decode steps; expert 0.75 MiB, granule 2 MiB) **The granularity that minimises resident mapped physical — and the owner's own `waste × transferred` product — is ONE 2 MiB granule holding TWO experts (E=2), not a large region.** | E | waste | resident MiB/layer/step | transfer MiB | sel% | |---|---|---|---|---| | 1 (per-expert) | 2.67× | 16.00 | 6.00 | 100% | | **2 ← optimum** | 1.33× | **13.34** | 10.01 | 167% | | 8 | 1.00× | 19.07 | 19.07 | 318% | | 32 (whole bank) | 1.00× | 24.00 | 24.00 | 400% | Resident physical is **U-shaped with its minimum at E=2 and rises back to whole-bank levels for every larger chunk** — the "tens of MiB" proposal is at the *worst* end. E=2 beats per-expert (16.00→13.34, **−17%**) because the second expert in a granule you already pay for is **free physical**. ### Why large regions are fatal (inferred) Every decode step activates **all 24 layers**, so any chunk spanning layers is hot in 100% of steps. A 64 MiB region (~85 experts) can't fit one layer's 24 MiB bank, so it **must** span layers → always resident → whole-bank pinning. The "near-zero waste" arithmetic is real but irrelevant; the binding objective is resident physical. ### The other answers - **Hotness-sorted packing rescues selectivity, modestly** (~5–14% vs arbitrary at every E; does *not* rescue large chunks). - **Static packing is stable**: decays only **3.4–7.1%** across prose/code/math vs a per-prompt oracle — so packing once is justified. Layer-1/2 always-on experts stay free pins. - **Interaction with #1295/#1325**: span count is **not** the cliff driver, so E=2 does **not** help the cliff — but lower peak mapped physical raises the admission-cap **plateau height** (not `N_max`). Single shared authority preserved. This is not conflated in the doc. ### Decisive framing (which architecture this assumes) - **If we keep our own VMM paging:** the minimal correct change is **E=2** (two experts/granule, hotness-packed, within-layer) — provably the trade-curve optimum, and small (still commits exactly one granule). **Do not build large-region commit-and-subdivide** — dominated on every metric. - **If OS hinting works, sub-division is unnecessary:** 4 KiB OS paging dissolves the granule problem (zero waste + full per-expert selectivity, strictly better than any VMM packing). - **My #1295 finding leans toward OS paging:** the binding cost is WDDM fault-in beyond usable VRAM, not our bookkeeping — so we may be fighting for control we don't benefit from. The condition to **not** build even E=2 is the OS-hint agent confirming WDDM honours residency hints with an adequate prediction window. I did **not** measure hinted-vs-unhinted fault-in (their job) but recorded that per-page fault-in is exactly what a prefetch hint would target. ### Reproduction `scripts/moe_granularity_analysis.py` reuses the exact `moe_router_skew.py` probe, captures per-step selection sets (the committed JSON has only aggregates), and prints the curve for both packings. Deterministic, bit-for-bit. Box: i7-13800H, RTX 4060 Laptop 8GB, driver 591.55, CUDA 13.1, WDDM. Analysis is CPU-EP (the *which experts* question is dtype-/EP-independent); granule is the CUDA device granule. No GPU decode claimed. Doc-only + reproducible script; no source change, nothing to test beyond the script's deterministic output. Follow-up to #1295 (closed by #1325). Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 18, 2026
…don't build it) (#1336) ## #1295 follow-up — re-scope MoE sub-division: the honest "no" The owner narrowed the question after the granularity trade-curve analysis was written: **small experts fit and never page, so granule waste is irrelevant; the real problem is over-subscription with large (≥ granule) experts** (GLM/DeepSeek-class, ~16 MiB = 8× the 2 MiB granule). This PR re-scopes the merged trade-curve doc accordingly and records the conclusion. ### Surviving question and answer *For large ≥-granule experts in the over-subscribed regime, commit large chunks and subdivide, or per-expert commit?* → **Each expert its own commit. Don't build sub-division.** The two arguments that could have survived both fail: - **Call count / churn.** A 16 MiB expert already maps **8 granules**; packing K experts into one chunk cuts the number of `cuMemMap`/`cuMemUnmap` *calls* (1 vs K) but **not granules mapped** (8K either way). The binding cost is **WDDM fault-in** (per-granule), and #1295/#1315 proved **span/commit count is not the cliff driver**. The reducible per-step API overhead is **3.39 ms/step** (16 MiB skewed) vs H2D's **10.9 ms/step** — well under ¼ of per-step paging cost, and **zero** of the fault-in cliff. - **Coherent residency grouping.** At ≥ granule there is no granule-sharing to force, so grouping is a **policy** decision served by per-expert commits + eviction. Packing into a shared commit **removes** independent per-expert eviction (a chunk is resident if *any* member is hot) — strictly worse, for no benefit. Grouping belongs in the offline residency-policy replay, at per-expert granularity. The sub-granule **E=2** result is retained only as arithmetic for the case now set aside. What survives: the **admission cap** (PR #1325, correct under our-VMM *and* OS-paging) and a **per-expert residency policy** (separate agent). If OS 4 KiB hinting is honoured, packing is moot entirely. Doc-only. Box: RTX 4060 Laptop 8GB, driver 591.55, CUDA 13.1, WDDM; model granite-3.0-1b-a400m / GLM-class 16 MiB experts inferred. Follow-up to #1295 (closed by #1325). Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1325 +/- ##
=========================================
+ Coverage 0 80.02% +80.02%
=========================================
Files 0 362 +362
Lines 0 157677 +157677
Branches 0 157677 +157677
=========================================
+ Hits 0 126175 +126175
- Misses 0 26863 +26863
- Partials 0 4639 +4639
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
This was referenced Aug 19, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue #1295 — VMM admission cap (cliff → plateau) +
vram_free_msattribution fixFollow-up to #1315 (the Phase-1 characterisation). Owner-greenlit Phase 2.
What this does
1. Admission cap sourced from usable VRAM. The device-authority ledger ceiling already bounds the sum of all Device-tier mapped physical (weights + KV + activations). It was mis-sourced:
Fraction(0.90)resolved against nominal total (8188 MiB → 7369 MiB), which exceeds measured usable free (~7100–7959 MiB on this shared box). A ceiling above usable lets the ledger permit leases the driver can't satisfy → the WDDM driver OOMs first. This clamps the ceiling tocuMemGetInfofree × fraction (default 0.90), so the ceiling is ≤ usable and the ledger refuses cleanly before the driver falls off the fault-in cliff. No parallel mechanism — the refusal routes through the existingtry_commit_span/reserve_mapped_growthG3 path.This is a cliff→plateau conversion, NOT a raised
N_max. The ceiling does not move; the failure mode changes from a catastrophic driver crash (and 90 svram_freestorms) into a graceful governed refusal. The next reader should not expect more sequences to fit — they should expect over-subscription to be refused instead of thrashed.Derivable, hardware-recorded fraction (not magic — cf #1261).
DEFAULT_VMM_MAPPED_FRACTION = 0.90, overridable viaONNX_GENAI_VMM_MAPPED_FRACTION. The 0.90 protects the ~740 MB margin between usable-free and the fault-in cliff observed in #1315: atfraction = 1.0the governed refusal lands exactly at free with no margin and the run sits on the edge of the fault-in degradation. What would justify changing it: a box where the fault-in cliff sits at a different offset below usable-free (measure it, don't guess).2.
vram_free_msinstrument fix (first-class deliverable).weight_paging.rs::release_allocation(true)synchronised both streams inside theGLOBAL_VRAM_FREE_NSbracket — so "free" time could be waiting on paging-slowed work. That headline (90 s) drove the whole investigation. Split the stream-drain intoGLOBAL_VRAM_FREE_SYNC_NS(surfaced asvram_free_sync_ms); the free bracket now spans only unmap/release/free.Corrected result (2 release samples, qwen14b-zp, RTX 4060 8GB, single-thread):
vram_free_ms= 41.5–45.2 s,vram_free_sync_ms= 1.6–1.9 s → sync is a stable ~4 % of the bracket. The headline is not substantially retracted — the corrected free time is still tens of seconds of genuine unmap/release churn. An honest instrument was still worth building; it refines the number rather than overturning it.Evidence
cuMemSetAccess failed: CUDA_ERROR_OUT_OF_MEMORY(driver crash). After (cap 0.90): cleanMemoryErrorat ceiling6699456921 B=floor(free × 0.90). Atfraction = 1.0: governed refusal at7443841024 B(= free exactly), proving the sourcing-from-free moves refusal driver→ledger.Honest limitation
A live multi-request N=8 throughput plateau could not be captured on this box — the governed batch-N sweep OOMs at load for N≥4 (14B weights ≈ usable VRAM on 8GB). The plateau mechanism is proven by the unit test + the live load-time refusal; the specific 0.80→peak tok/s curve is not measurable here and not claimed.
Cross-slice
Test / build
--features "cuda,native-backend", single-thread. Governor suite: 13 passed / 0 failed / 0 ignored. Pre-existing failures unchanged and unrelated: #1284 (three harness), #1305 (four kernel on this GPU).Closes #1295.