[Model Runner V2] Reserve CUDA graph memory - #53306
WoosukKwon merged 12 commits into
Conversation
Model Runner V2 became the default for dense models in 0.25.1, but its profile_cudagraph_memory() was a placeholder returning 0. Worker. determine_available_memory() subtracts this estimate before sizing the KV cache, so with V2 no headroom was reserved for CUDA graph capture: the KV cache claimed the whole gpu_memory_utilization budget and capture_model() OOMed at startup (e.g. Llama-3.1-70B FP8, TP=8 on L40S). Implement profile_cudagraph_memory() for the V2 runner, reusing its own initialize_kv_cache()/capture_model(): bootstrap a minimal KV cache, capture graphs into a throwaway pool (so their memory is reclaimed and does not pollute the persistent global pool), measure the free-memory delta, then release the profiling state while keeping model weights. Fixes vllm-project#49224 Signed-off-by: Anh Tran <anh.tran.3889@gmail.com>
…ph-memory-profiling
|
/ci run all |
|
✅ Triggered Buildkite CI #85073 for commit |
…ph-memory-profiling
|
/ci run all |
|
✅ Triggered Buildkite CI #85113 for commit |
Signed-off-by: Nick Hill <nickhill123@gmail.com>
…x/v2-runner-cudagraph-memory-profiling
|
/ci run |
|
✅ Triggered Buildkite CI #85217 for commit |
…udagraph-memory-profiling
|
/ci run |
|
✅ Triggered Buildkite CI #85254 for commit |
Upstream vllm-project#53306 filled in the V2 `profile_cudagraph_memory()` stub this branch had also implemented, so the two implementations collided in `vllm/v1/worker/gpu/model_runner.py`. Resolved in favour of upstream's, keeping only the workspace conduit this PR needs on top: - `profile_cudagraph_memory()` now delegates to `cudagraph_utils.profile_cudagraph_memory()` and keeps the `persistent_workspace_profiled` keyword the worker passes to both runners. When the persistent workspace was already profiled, the arena is asserted not to grow across the delegate: the profiling capture re-runs the reservation, which must be a no-op so the measured delta covers the graphs alone and is not double counted against the activation peak that already holds the persistent workspace. - Dropped this branch's `_init_minimal_kv_cache_for_profiling()` and `_cleanup_profiling_kv_cache()` duplicates; `prepare_profiling_workspace()` now uses the upstream module helpers, so both profiling paths share one bootstrap/teardown. - Carried the quantized KV scale-view detach into `_teardown_profiling_state()`. `_k_scale_cache`/`_v_scale_cache` are strided views into the profiling KV cache (triton_attn int8/fp8 per-token-head), so the profiling allocation stays live through KV sizing unless the impl drops them too. V1 already does this. - `initialize_kv_cache()` keeps upstream's unconditional `DraftModelSpeculator.set_attn()`. The `not is_profiling` guard this branch added leaves `attn_cg_support` unset for DFlash speculators, and `init_cudagraph_manager()` below reads it; upstream profiling now also captures speculator graphs through `capture_model()`, so the profiling path needs the full attention state. - Took upstream's `is_profiling` branch for the KV connector, which also installs `NO_OP_KV_CONNECTOR` on the profiling path. Tests follow the same split: the persistent-workspace tests patch the module-level helpers for V2 and the runner methods for V1, the V2 profiling tests cover the delegate and its arena guard, and the `_teardown_profiling_state` test moved to `cudagraph_utils`. Also fixed two stubs that had fallen behind their source: `_get_workspace_routes()` reads `use_dedicated_xqa` (now parametrized, with a non-causal dedicated-XQA decode case), and `capture_model()` reads `model_state.supports_mm_inputs` and `adaptive_verification`. Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
|
Sherlock here, Kevin Luu's CI-monitoring agent. Post-merge main reproduced a Kimi K3 initialization regression twice in the exact H200 Extra Initialization shard:
Both failed The generic Mamba/GDN paths in this PR tolerate a missing per-layer metadata entry during profiling, but Kimi K3's dedicated NVIDIA/AMD KDA implementations still indexed it directly. This PR's exact H200 Extra Initialization shard was blocked in #85254, so the Kimi path was not exercised before merge. Focused fix #53581 passed pre-merge exact #85323 shard 1 92/92 on Resolved on main: fix #53581 merged as |
…)" This reverts commit a4d70be. Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
Profiling captures graphs before the real KV cache is allocated and
discards them at teardown. Any of them captured into the persistent
global graph pool drop its use_count to 0 when released, tripping the
c10 allocator's create_or_incref_pool assert ("use_count > 0 INTERNAL
ASSERT FAILED") when the real capture re-opens that pool.
Point the platform's global graph pool singleton at a throwaway pool
for the whole profiling phase, so objects that bind a pool lazily
during profiling (speculator cudagraph managers created in the
profiling initialize_kv_cache, breakable-CG runners created
mid-capture) land on the throwaway pool by construction. Piecewise
wrappers and the decoder manager bind the pool earlier, so swap those
explicitly (and restore them after). Drop all profiling captures at
teardown.
Fixes the nightly B200 DeepSeek-V4-Flash failure after vllm-project#53306;
reproduced on GB200 with the exact nightly config (DSV4-Flash, TP2+EP,
fp8 KV, deepgemm mega-MoE, MTPx2): crashes before this change, engine
initializes cleanly with it.
Co-authored-by: Kimi Code CLI
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Anh Tran <anh.tran.3889@gmail.com> Signed-off-by: Nick Hill <nickhill123@gmail.com> Co-authored-by: Anh Tran <anh.tran.3889@gmail.com> Co-authored-by: anhtra3889 <anhtra@nvidia.com>
Signed-off-by: Anh Tran <anh.tran.3889@gmail.com> Signed-off-by: Nick Hill <nickhill123@gmail.com> Co-authored-by: Anh Tran <anh.tran.3889@gmail.com> Co-authored-by: anhtra3889 <anhtra@nvidia.com> Signed-off-by: mikeshawcode <michaelwshaw2@gmail.com>
Hybrid KV-cache group sizing for a small sliding-window bucket, and explicit CUDA-graph
memory accounting for the V2 model runner. Both matter for single-user SPEC=dflash2; both are
inert for every other config in this repo (verified: MTP mode's pool is bit-identical).
1. v1/core/kv_cache_utils.py — `_get_kv_cache_groups_uniform_page_size` sets the KV group size
to the SMALLEST bucket of same-type layers. With the DFlash2 drafter that bucket is its 5
sliding-window layers, so the target's 16 full-attention layers get padded to 20 and the 48
GDN layers to 50 ("Add 4 padding layers, may waste at most 25%"): 25% more memory for every
token of context, to pad layers that are not the problem. Sliding-window groups only ever
hold window-many blocks (the manager frees blocks behind the window), so padding THEM is
nearly free. `_prefer_padding_sliding_window_buckets` therefore picks the largest common
divisor of the non-sliding buckets instead, when the smallest bucket is sliding-window-only:
16/48/5 -> group size 8 (no full/GDN padding, 3 padding layers on the 5-layer window group,
which costs ~7 MB per request). Measured on the 3090: 105 -> 78 KB of pool per token, i.e.
45,383 tokens at 40k context -> 69,758 at 64k. MTP mode has no sliding-window bucket, takes
the early return, and keeps its 73,777-token pool.
2. v1/worker/gpu/model_runner.py — the V1 runner profiles its CUDA graphs and the worker
subtracts the estimate from the KV budget; the V2 runner's `profile_cudagraph_memory`
returns 0 ("TBD"), so its ~1.2 GiB of graphs lands ON TOP of gpu_memory_utilization: ask for
0.93 and the process actually peaks near 0.98, which is how the DeltaNet spec path OOMs
mid-request (main README, gotcha 4). Profiling them the way V1 does needs a temporary KV
cache and graph pool; this instead makes the reservation explicit and measurable:
VLLM_V2_CUDAGRAPH_MEM_MIB=<MiB> is subtracted before the KV cache is sized (0 = upstream
behaviour). Read the real figure from the startup line "... and X GiB for CUDAGraph memory".
Note the profiled activation peak itself varies by ~1 GiB between starts on this model, so
for a fixed context length prefer pinning the pool with --kv-cache-memory (what
single-user/start_qwen.sh does for SPEC=dflash2) over tuning gpu_memory_utilization.
Apply from the installed vllm package directory:
patch -p1 -d venv/lib/python3.12/site-packages/vllm < patches/hybrid-kv-groups-v2-cudagraph.patch
Validated against vLLM 0.28.0. Reapply after upgrades.
Retired on the 0.29 port: the VLLM_V2_CUDAGRAPH_MEM_MIB reserve hunk in v1/worker/gpu/model_runner.py.
vLLM 0.29 profiles CUDA graph memory natively (cudagraph_utils.profile_cudagraph_memory, upstream vllm-project#53306), so the
reserve knob is a no-op there; only the kv_cache_utils hunks remain.
Source: syv-ai/qwen38-27b-rtx3090 patches/hybrid-kv-groups-v2-cudagraph.patch (at d4c4cc2). Applied as-is from the patch file; the preamble above is the patch's own.
Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Michael Stufflebeam <cpuchip@gmail.com>
Profile cuda graph memory usage at startup so that it can be factored into the kv cache auto-sizing, as MRV1 already does.
Note we are hoping to replace this with #50779 soon but it will unblock the change to make MRV2 default for all models in the meantime.
This includes the commits from #49233 by @anhtra3889 plus some fixes / additional rework.
Fixes: #49224