Skip to content

vmm(#1295): usable-VRAM admission cap (cliff->plateau) + vram_free_ms attribution fix - #1325

Merged
justinchuby merged 1 commit into
mainfrom
squad/1295-admission-cap
Aug 18, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/1295-admission-cap

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Issue #1295 — VMM admission cap (cliff → plateau) + vram_free_ms attribution fix

Follow-up to #1315 (the Phase-1 characterisation). Owner-greenlit Phase 2.

What this does

1. Admission cap sourced from usable VRAM. The device-authority ledger ceiling already bounds the sum of all Device-tier mapped physical (weights + KV + activations). It was mis-sourced: Fraction(0.90) resolved against nominal total (8188 MiB → 7369 MiB), which exceeds measured usable free (~7100–7959 MiB on this shared box). A ceiling above usable lets the ledger permit leases the driver can't satisfy → the WDDM driver OOMs first. This clamps the ceiling to cuMemGetInfo free × fraction (default 0.90), so the ceiling is ≤ usable and the ledger refuses cleanly before the driver falls off the fault-in cliff. No parallel mechanism — the refusal routes through the existing try_commit_span / reserve_mapped_growth G3 path.

This is a cliff→plateau conversion, NOT a raised N_max. The ceiling does not move; the failure mode changes from a catastrophic driver crash (and 90 s vram_free storms) into a graceful governed refusal. The next reader should not expect more sequences to fit — they should expect over-subscription to be refused instead of thrashed.

Derivable, hardware-recorded fraction (not magic — cf #1261). DEFAULT_VMM_MAPPED_FRACTION = 0.90, overridable via ONNX_GENAI_VMM_MAPPED_FRACTION. The 0.90 protects the ~740 MB margin between usable-free and the fault-in cliff observed in #1315: at fraction = 1.0 the governed refusal lands exactly at free with no margin and the run sits on the edge of the fault-in degradation. What would justify changing it: a box where the fault-in cliff sits at a different offset below usable-free (measure it, don't guess).

2. vram_free_ms instrument fix (first-class deliverable). weight_paging.rs::release_allocation(true) synchronised both streams inside the GLOBAL_VRAM_FREE_NS bracket — so "free" time could be waiting on paging-slowed work. That headline (90 s) drove the whole investigation. Split the stream-drain into GLOBAL_VRAM_FREE_SYNC_NS (surfaced as vram_free_sync_ms); the free bracket now spans only unmap/release/free.

Corrected result (2 release samples, qwen14b-zp, RTX 4060 8GB, single-thread): vram_free_ms = 41.5–45.2 s, vram_free_sync_ms = 1.6–1.9 s → sync is a stable ~4 % of the bracket. The headline is not substantially retracted — the corrected free time is still tens of seconds of genuine unmap/release churn. An honest instrument was still worth building; it refines the number rather than overturning it.

Evidence

  • Live cliff→plateau: no-limit 14B VMM load — before (nominal-sourced / no headroom): cuMemSetAccess failed: CUDA_ERROR_OUT_OF_MEMORY (driver crash). After (cap 0.90): clean MemoryError at ceiling 6699456921 B = floor(free × 0.90). At fraction = 1.0: governed refusal at 7443841024 B (= free exactly), proving the sourcing-from-free moves refusal driver→ledger.
  • 13 governor unit tests pass (0 failed, 0 ignored), incl. a non-vacuous G3 refusal test.

Honest limitation

A live multi-request N=8 throughput plateau could not be captured on this box — the governed batch-N sweep OOMs at load for N≥4 (14B weights ≈ usable VRAM on 8GB). The plateau mechanism is proven by the unit test + the live load-time refusal; the specific 0.80→peak tok/s curve is not measurable here and not claimed.

Cross-slice

Test / build

--features "cuda,native-backend", single-thread. Governor suite: 13 passed / 0 failed / 0 ignored. Pre-existing failures unchanged and unrelated: #1284 (three harness), #1305 (four kernel on this GPU).

Closes #1295.

… + fix vram_free_ms attribution

Issue #1295 (follow-up to #1315). Converts the batch-N streaming cliff from a
driver OOM crash into a clean governed refusal by clamping the device-authority
ledger ceiling to measured usable VRAM (cuMemGetInfo free x fraction, default
0.90) instead of nominal total. This is a cliff->plateau conversion, NOT a
raised N_max: the ledger now refuses the over-subscribing lease before the WDDM
driver falls off the fault-in cliff.

- governor.rs: clamp_ceiling_to_usable_vram() + usable_mapped_safe_fraction();
  DEFAULT_VMM_MAPPED_FRACTION=0.90 (env ONNX_GENAI_VMM_MAPPED_FRACTION), sourced
  from capacities.vram.free_bytes(); applied at device_ceiling_bytes with
  tracing::info when it binds. Single shared authority (MoE #1308 reuses it).
- Instrumentation fix: release_allocation() timed stream-drain sync inside the
  GLOBAL_VRAM_FREE_NS bracket, mis-attributing paging-slowed work to "free".
  Split into GLOBAL_VRAM_FREE_SYNC_NS (surfaced as vram_free_sync_ms). Corrected:
  sync is ~4% of the bracket; the ~45s free time is genuine unmap/release churn.
  Headline not substantially retracted.
- 3 governor unit tests (fraction validation, usable-sourced clamp, non-vacuous
  G3 refusal). Benchmark doc with measured/inferred split.

Measured: RTX 4060 Laptop 8GB, driver 591.55, CUDA 13.1, WDDM, i7-13800H,
single-thread, model qwen14b-zp.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby merged commit 1014b40 into main Aug 18, 2026
7 of 16 checks passed
@justinchuby
justinchuby deleted the squad/1295-admission-cap branch August 18, 2026 20:31
justinchuby added a commit that referenced this pull request Aug 18, 2026
…ts (not commit-and-subdivide) (#1332)

## #1295 follow-up — MoE VMM sub-granule packing: the trade curve
(analysis, no impl)

**Reporting before implementing, as the owner asked.** This answers "can
VMM commit more at once and subdivide ourselves?" from the
already-measured granite router-skew distribution — no packing code was
written.

### The question and the trap
The 2 MiB CUDA granule wastes **2.7×** on granite int4 experts (~0.75
MiB). The proposed escape — commit a large region (tens of MiB), pack
~85 experts, subdivide ourselves — is favourable on granule *waste*
alone. But committing is also the unit of *residency*, so the real
question is whether sub-division keeps selectivity.

### Finding (measured, granite-3.0-1b-a400m, 32e/top-8/24L, real IBM
router, CPU EP, 192 decode steps; expert 0.75 MiB, granule 2 MiB)
**The granularity that minimises resident mapped physical — and the
owner's own `waste × transferred` product — is ONE 2 MiB granule holding
TWO experts (E=2), not a large region.**

| E | waste | resident MiB/layer/step | transfer MiB | sel% |
|---|---|---|---|---|
| 1 (per-expert) | 2.67× | 16.00 | 6.00 | 100% |
| **2 ← optimum** | 1.33× | **13.34** | 10.01 | 167% |
| 8 | 1.00× | 19.07 | 19.07 | 318% |
| 32 (whole bank) | 1.00× | 24.00 | 24.00 | 400% |

Resident physical is **U-shaped with its minimum at E=2 and rises back
to whole-bank levels for every larger chunk** — the "tens of MiB"
proposal is at the *worst* end. E=2 beats per-expert (16.00→13.34,
**−17%**) because the second expert in a granule you already pay for is
**free physical**.

### Why large regions are fatal (inferred)
Every decode step activates **all 24 layers**, so any chunk spanning
layers is hot in 100% of steps. A 64 MiB region (~85 experts) can't fit
one layer's 24 MiB bank, so it **must** span layers → always resident →
whole-bank pinning. The "near-zero waste" arithmetic is real but
irrelevant; the binding objective is resident physical.

### The other answers
- **Hotness-sorted packing rescues selectivity, modestly** (~5–14% vs
arbitrary at every E; does *not* rescue large chunks).
- **Static packing is stable**: decays only **3.4–7.1%** across
prose/code/math vs a per-prompt oracle — so packing once is justified.
Layer-1/2 always-on experts stay free pins.
- **Interaction with #1295/#1325**: span count is **not** the cliff
driver, so E=2 does **not** help the cliff — but lower peak mapped
physical raises the admission-cap **plateau height** (not `N_max`).
Single shared authority preserved. This is not conflated in the doc.

### Decisive framing (which architecture this assumes)
- **If we keep our own VMM paging:** the minimal correct change is
**E=2** (two experts/granule, hotness-packed, within-layer) — provably
the trade-curve optimum, and small (still commits exactly one granule).
**Do not build large-region commit-and-subdivide** — dominated on every
metric.
- **If OS hinting works, sub-division is unnecessary:** 4 KiB OS paging
dissolves the granule problem (zero waste + full per-expert selectivity,
strictly better than any VMM packing).
- **My #1295 finding leans toward OS paging:** the binding cost is WDDM
fault-in beyond usable VRAM, not our bookkeeping — so we may be fighting
for control we don't benefit from. The condition to **not** build even
E=2 is the OS-hint agent confirming WDDM honours residency hints with an
adequate prediction window. I did **not** measure hinted-vs-unhinted
fault-in (their job) but recorded that per-page fault-in is exactly what
a prefetch hint would target.

### Reproduction
`scripts/moe_granularity_analysis.py` reuses the exact
`moe_router_skew.py` probe, captures per-step selection sets (the
committed JSON has only aggregates), and prints the curve for both
packings. Deterministic, bit-for-bit.

Box: i7-13800H, RTX 4060 Laptop 8GB, driver 591.55, CUDA 13.1, WDDM.
Analysis is CPU-EP (the *which experts* question is
dtype-/EP-independent); granule is the CUDA device granule. No GPU
decode claimed. Doc-only + reproducible script; no source change,
nothing to test beyond the script's deterministic output.

Follow-up to #1295 (closed by #1325).

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
justinchuby added a commit that referenced this pull request Aug 18, 2026
…don't build it) (#1336)

## #1295 follow-up — re-scope MoE sub-division: the honest "no"

The owner narrowed the question after the granularity trade-curve
analysis was written: **small experts fit and never page, so granule
waste is irrelevant; the real problem is over-subscription with large (≥
granule) experts** (GLM/DeepSeek-class, ~16 MiB = 8× the 2 MiB granule).
This PR re-scopes the merged trade-curve doc accordingly and records the
conclusion.

### Surviving question and answer
*For large ≥-granule experts in the over-subscribed regime, commit large
chunks and subdivide, or per-expert commit?* → **Each expert its own
commit. Don't build sub-division.** The two arguments that could have
survived both fail:

- **Call count / churn.** A 16 MiB expert already maps **8 granules**;
packing K experts into one chunk cuts the number of
`cuMemMap`/`cuMemUnmap` *calls* (1 vs K) but **not granules mapped** (8K
either way). The binding cost is **WDDM fault-in** (per-granule), and
#1295/#1315 proved **span/commit count is not the cliff driver**. The
reducible per-step API overhead is **3.39 ms/step** (16 MiB skewed) vs
H2D's **10.9 ms/step** — well under ¼ of per-step paging cost, and
**zero** of the fault-in cliff.
- **Coherent residency grouping.** At ≥ granule there is no
granule-sharing to force, so grouping is a **policy** decision served by
per-expert commits + eviction. Packing into a shared commit **removes**
independent per-expert eviction (a chunk is resident if *any* member is
hot) — strictly worse, for no benefit. Grouping belongs in the offline
residency-policy replay, at per-expert granularity.

The sub-granule **E=2** result is retained only as arithmetic for the
case now set aside. What survives: the **admission cap** (PR #1325,
correct under our-VMM *and* OS-paging) and a **per-expert residency
policy** (separate agent). If OS 4 KiB hinting is honoured, packing is
moot entirely.

Doc-only. Box: RTX 4060 Laptop 8GB, driver 591.55, CUDA 13.1, WDDM;
model granite-3.0-1b-a400m / GLM-class 16 MiB experts inferred.
Follow-up to #1295 (closed by #1325).

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 19, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.02%. Comparing base (e17a695) to head (713cac7).
⚠️ Report is 56 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff            @@
##           main    #1325       +/-   ##
=========================================
+ Coverage      0   80.02%   +80.02%     
=========================================
  Files         0      362      +362     
  Lines         0   157677   +157677     
  Branches      0   157677   +157677     
=========================================
+ Hits          0   126175   +126175     
- Misses        0    26863    +26863     
- Partials      0     4639     +4639     
Flag Coverage Δ
offline 80.02% <ø> (?)

Flags with carried forward coverage won't be shown. Click here to find out more.
see 362 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/medium_generic_f32_threads=8/32x512x512 1.06 ms 1.91 ms +80.9%
🔴 matmul/small_generic_f32_threads=8/1x256x256 33.38 µs 56.19 µs +68.3%
🔴 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 602.59 µs 983.17 µs +63.2%
🔴 gather/large_f32_threads=1-internal/131072 27.10 µs 40.60 µs +49.8%
🔴 matmul/medium_generic_bf16_threads=8/32x512x512 394.91 µs 580.44 µs +47.0%
🔴 gather/large_f16_threads=1-internal/131072 11.37 µs 16.58 µs +45.8%
🔴 gather/large_bf16_threads=1-internal/131072 11.28 µs 16.11 µs +42.8%
🔴 gather/medium_f32_threads=1-internal/32768 3.75 µs 5.14 µs +37.1%
🔴 matmul/small_generic_bf16_threads=8/1x256x256 31.98 µs 43.74 µs +36.8%
🔴 matmul/large_generic_bf16_threads=8/32x1024x1024 1.41 ms 1.91 ms +35.4%
🔴 tokenization/decode_tokens_per_second 7.04 ms 9.30 ms +32.2%
⚠️ matmul/small_generic_f32_threads=1/1x256x256 36.50 µs 47.16 µs +29.2%
⚠️ matmul/small_generic_f16_threads=1/1x256x256 32.37 µs 41.20 µs +27.3%
⚠️ matmul/medium_generic_f16_threads=8/32x512x512 30.71 µs 38.78 µs +26.3%
⚠️ tokenization/encode_tokens_per_second 415.07 µs 519.78 µs +25.2%
⚠️ matmul/medium_generic_bf16_threads=1/32x512x512 533.94 µs 663.90 µs +24.3%
⚠️ gather/small_bf16_threads=1-internal/4096 475.7 ns 590.7 ns +24.2%
⚠️ matmul/medium_generic_f16_threads=1/32x512x512 30.72 µs 37.68 µs +22.7%
⚠️ reduce_mean/medium_f32_threads=1-internal/65536 244.92 µs 297.59 µs +21.5%
⚠️ gather/small_f32_threads=1-internal/4096 669.4 ns 794.3 ns +18.7%
⚠️ add/large_bf16_threads=1-internal/4194304 1.68 ms 1.96 ms +16.7%
⚠️ matmul/small_generic_f16_threads=8/1x256x256 31.62 µs 36.90 µs +16.7%
⚠️ sampling_latency/top_k_per_token 52.16 µs 60.86 µs +16.7%
⚠️ sampling_latency/min_p_per_token 208.42 µs 242.39 µs +16.3%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.65 ms 3.04 ms +15.0%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 83.35 µs 94.75 µs +13.7%
✅ gather/small_f16_threads=1-internal/4096 477.0 ns 542.2 ns +13.7%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 5.73 ms 6.49 ms +13.3%
✅ kv_cache/alloc_dealloc_pages 38.23 µs 42.98 µs +12.4%
✅ gather/medium_f16_threads=1-internal/32768 2.43 µs 2.72 µs +12.3%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 3.91 ms 4.32 ms +10.5%
✅ gather/medium_bf16_threads=1-internal/32768 2.51 µs 2.77 µs +10.4%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 519.07 µs 573.08 µs +10.4%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.61 ms 3.97 ms +10.0%
✅ sampling_latency/greedy_per_token 3.47 µs 3.79 µs +9.4%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 95.25 µs 103.63 µs +8.8%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 2.11 ms 2.28 ms +8.0%
✅ qwen3_sampling_processors/top_k_top_p_fast 657.30 µs 708.49 µs +7.8%
✅ matmul/small_generic_bf16_threads=1/1x256x256 32.14 µs 34.52 µs +7.4%
✅ reduce_mean/small_f32_threads=1-internal/4096 16.36 µs 17.57 µs +7.4%
✅ sampling_latency/top_p_per_token 379.17 µs 403.89 µs +6.5%
✅ qwen3_sampling_processors/top_k_partial_selection 146.65 µs 150.72 µs +2.8%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.06 ms 2.11 ms +2.6%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 9.52 ms 9.76 ms +2.6%
✅ grammar_masking/llguidance_compute_mask/32 85.58 µs 87.43 µs +2.2%
✅ logit_processing/seven_processor_chain_per_step 338.24 µs 339.98 µs +0.5%
✅ add/medium_f16_threads=1-internal/262144 104.83 µs 105.26 µs +0.4%
✅ reduce_mean/large_f32_threads=1-internal/262144 974.12 µs 978.02 µs +0.4%
✅ add/large_f16_threads=1-internal/4194304 1.76 ms 1.75 ms -0.5%
✅ add/small_f16_threads=1-internal/1024 452.6 ns 447.7 ns -1.1%
✅ add/medium_bf16_threads=1-internal/262144 103.72 µs 102.39 µs -1.3%
✅ add/small_bf16_threads=1-internal/1024 448.2 ns 435.0 ns -2.9%
✅ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 66.25 µs 63.63 µs -4.0%
✅ add/medium_f32_threads=1-internal/262144 26.03 µs 24.62 µs -5.4%
✅ add/large_f32_threads=1-internal/4194304 695.25 µs 648.50 µs -6.7%
🟢 add/small_f32_threads=1-internal/1024 236.1 ns 198.2 ns -16.0%
🟢 block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 513.60 µs 387.76 µs -24.5%
🟢 block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 98.50 µs 55.61 µs -43.5%
🟢 block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 115.54 µs 48.11 µs -58.4%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 5.22 3.89 4.88 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Streaming-regime batch-N ceiling is VMM map/unmap churn, not the batching path (N_max~4-5 vs projected ~19 on 8GB)

2 participants