Skip to content

[Model Runner V2] Reserve CUDA graph memory - #53306

Merged
WoosukKwon merged 12 commits into
vllm-project:mainfrom
njhill:fix/v2-runner-cudagraph-memory-profiling
Aug 24, 2026
Merged

WoosukKwon merged 12 commits into
vllm-project:mainfrom
njhill:fix/v2-runner-cudagraph-memory-profiling

Conversation

@njhill

@njhill njhill commented Aug 21, 2026 •

Copy link
Copy Markdown
Member

Profile cuda graph memory usage at startup so that it can be factored into the kv cache auto-sizing, as MRV1 already does.

Note we are hoping to replace this with #50779 soon but it will unblock the change to make MRV2 default for all models in the meantime.

This includes the commits from #49233 by @anhtra3889 plus some fixes / additional rework.

Fixes: #49224

anht3889 and others added 7 commits July 21, 2026 04:33
Model Runner V2 became the default for dense models in 0.25.1, but its
profile_cudagraph_memory() was a placeholder returning 0. Worker.
determine_available_memory() subtracts this estimate before sizing the
KV cache, so with V2 no headroom was reserved for CUDA graph capture:
the KV cache claimed the whole gpu_memory_utilization budget and
capture_model() OOMed at startup (e.g. Llama-3.1-70B FP8, TP=8 on L40S).

Implement profile_cudagraph_memory() for the V2 runner, reusing its own
initialize_kv_cache()/capture_model(): bootstrap a minimal KV cache,
capture graphs into a throwaway pool (so their memory is reclaimed and
does not pollute the persistent global pool), measure the free-memory
delta, then release the profiling state while keeping model weights.

Fixes vllm-project#49224

Signed-off-by: Anh Tran <anh.tran.3889@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@njhill

njhill commented Aug 21, 2026

Copy link
Copy Markdown
Member Author

/ci run all

@njhill njhill added the ready ONLY add when PR is ready to merge/full CI is needed label Aug 21, 2026
@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #85073 for commit 36a1131bab0e.

@vllm-project vllm-project deleted a comment from mergify Bot Aug 21, 2026
@njhill

njhill commented Aug 21, 2026

Copy link
Copy Markdown
Member Author

/ci run all

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #85113 for commit c72b87c3d36e.

@github-project-automation github-project-automation Bot moved this to Ready in NVIDIA Aug 22, 2026
njhill added 2 commits August 22, 2026 17:48
Signed-off-by: Nick Hill <nickhill123@gmail.com>
@njhill

njhill commented Aug 23, 2026

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #85217 for commit d3fa7091c309.

@njhill
njhill enabled auto-merge (squash) August 23, 2026 01:08
@mergify mergify Bot added the kimi label Aug 23, 2026
@njhill

njhill commented Aug 23, 2026

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #85254 for commit b2a8c5666c73.

@WoosukKwon
WoosukKwon disabled auto-merge August 24, 2026 06:11
@WoosukKwon
WoosukKwon merged commit a4d70be into vllm-project:main Aug 24, 2026
115 of 117 checks passed
@github-project-automation github-project-automation Bot moved this from Ready to Done in NVIDIA Aug 24, 2026
lesj0610 added a commit to lesj0610/vllm that referenced this pull request Aug 24, 2026
Upstream vllm-project#53306 filled in the V2 `profile_cudagraph_memory()` stub this
branch had also implemented, so the two implementations collided in
`vllm/v1/worker/gpu/model_runner.py`. Resolved in favour of upstream's,
keeping only the workspace conduit this PR needs on top:

- `profile_cudagraph_memory()` now delegates to
  `cudagraph_utils.profile_cudagraph_memory()` and keeps the
  `persistent_workspace_profiled` keyword the worker passes to both
  runners. When the persistent workspace was already profiled, the arena
  is asserted not to grow across the delegate: the profiling capture
  re-runs the reservation, which must be a no-op so the measured delta
  covers the graphs alone and is not double counted against the
  activation peak that already holds the persistent workspace.
- Dropped this branch's `_init_minimal_kv_cache_for_profiling()` and
  `_cleanup_profiling_kv_cache()` duplicates; `prepare_profiling_workspace()`
  now uses the upstream module helpers, so both profiling paths share one
  bootstrap/teardown.
- Carried the quantized KV scale-view detach into
  `_teardown_profiling_state()`. `_k_scale_cache`/`_v_scale_cache` are
  strided views into the profiling KV cache (triton_attn
  int8/fp8 per-token-head), so the profiling allocation stays live
  through KV sizing unless the impl drops them too. V1 already does this.
- `initialize_kv_cache()` keeps upstream's unconditional
  `DraftModelSpeculator.set_attn()`. The `not is_profiling` guard this
  branch added leaves `attn_cg_support` unset for DFlash speculators, and
  `init_cudagraph_manager()` below reads it; upstream profiling now also
  captures speculator graphs through `capture_model()`, so the profiling
  path needs the full attention state.
- Took upstream's `is_profiling` branch for the KV connector, which also
  installs `NO_OP_KV_CONNECTOR` on the profiling path.

Tests follow the same split: the persistent-workspace tests patch the
module-level helpers for V2 and the runner methods for V1, the V2
profiling tests cover the delegate and its arena guard, and the
`_teardown_profiling_state` test moved to `cudagraph_utils`.

Also fixed two stubs that had fallen behind their source:
`_get_workspace_routes()` reads `use_dedicated_xqa` (now parametrized,
with a non-causal dedicated-XQA decode case), and `capture_model()`
reads `model_state.supports_mm_inputs` and `adaptive_verification`.

Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
@khluu

khluu commented Aug 24, 2026 •

Copy link
Copy Markdown
Member

Sherlock here, Kevin Luu's CI-monitoring agent. Post-merge main reproduced a Kimi K3 initialization regression twice in the exact H200 Extra Initialization shard:

Both failed test_can_initialize_large_subset[KimiK3ForConditionalGeneration] during CUDA-graph memory profiling with:

KeyError: 'language_model.model.layers.0.self_attn'

The generic Mamba/GDN paths in this PR tolerate a missing per-layer metadata entry during profiling, but Kimi K3's dedicated NVIDIA/AMD KDA implementations still indexed it directly. This PR's exact H200 Extra Initialization shard was blocked in #85254, so the Kimi path was not exercised before merge.

Focused fix #53581 passed pre-merge exact #85323 shard 1 92/92 on h200-ci-4, then merged as 4c56e62c85ce. Natural merge build #85345 blocked the exact shard, so targeted post-merge main #85350 is now scheduled at the literal merge commit. The incident remains open until #85350 shard 1 passes.

Resolved on main: fix #53581 merged as 4c56e62c85ce, and targeted post-merge main #85350 shard 1 passed 92/92 (3 deselected, 0 failed) on physical h200-ci-4. The exact Kimi K3 case completed CUDA-graph profiling and FlashInfer autotune without the missing-metadata KeyError. No revert of this PR is needed; the omitted Kimi-specific guard is now fixed and confirmed on main.

@njhill
njhill deleted the fix/v2-runner-cudagraph-memory-profiling branch August 24, 2026 14:22
khluu added a commit to khluu/vllm that referenced this pull request Aug 24, 2026
…)"

This reverts commit a4d70be.

Signed-off-by: Kevin Luu <51931015+khluu@users.noreply.github.com>
njhill added a commit to njhill/vllm that referenced this pull request Aug 25, 2026
Profiling captures graphs before the real KV cache is allocated and
discards them at teardown. Any of them captured into the persistent
global graph pool drop its use_count to 0 when released, tripping the
c10 allocator's create_or_incref_pool assert ("use_count > 0 INTERNAL
ASSERT FAILED") when the real capture re-opens that pool.

Point the platform's global graph pool singleton at a throwaway pool
for the whole profiling phase, so objects that bind a pool lazily
during profiling (speculator cudagraph managers created in the
profiling initialize_kv_cache, breakable-CG runners created
mid-capture) land on the throwaway pool by construction. Piecewise
wrappers and the decoder manager bind the pool earlier, so swap those
explicitly (and restore them after). Drop all profiling captures at
teardown.

Fixes the nightly B200 DeepSeek-V4-Flash failure after vllm-project#53306;
reproduced on GB200 with the exact nightly config (DSV4-Flash, TP2+EP,
fp8 KV, deepgemm mega-MoE, MTPx2): crashes before this change, engine
initializes cleanly with it.

Co-authored-by: Kimi Code CLI
Signed-off-by: Nick Hill <nickhill123@gmail.com>
am-cohere pushed a commit to am-cohere/vllm that referenced this pull request Sep 1, 2026
Signed-off-by: Anh Tran <anh.tran.3889@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Anh Tran <anh.tran.3889@gmail.com>
Co-authored-by: anhtra3889 <anhtra@nvidia.com>
mikeshawcode pushed a commit to mikeshawcode/vllm that referenced this pull request Sep 1, 2026
Signed-off-by: Anh Tran <anh.tran.3889@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Anh Tran <anh.tran.3889@gmail.com>
Co-authored-by: anhtra3889 <anhtra@nvidia.com>
Signed-off-by: mikeshawcode <michaelwshaw2@gmail.com>
cpuchip pushed a commit to cpuchip/vllm that referenced this pull request Sep 14, 2026
Hybrid KV-cache group sizing for a small sliding-window bucket, and explicit CUDA-graph
memory accounting for the V2 model runner. Both matter for single-user SPEC=dflash2; both are
inert for every other config in this repo (verified: MTP mode's pool is bit-identical).

1. v1/core/kv_cache_utils.py — `_get_kv_cache_groups_uniform_page_size` sets the KV group size
   to the SMALLEST bucket of same-type layers. With the DFlash2 drafter that bucket is its 5
   sliding-window layers, so the target's 16 full-attention layers get padded to 20 and the 48
   GDN layers to 50 ("Add 4 padding layers, may waste at most 25%"): 25% more memory for every
   token of context, to pad layers that are not the problem. Sliding-window groups only ever
   hold window-many blocks (the manager frees blocks behind the window), so padding THEM is
   nearly free. `_prefer_padding_sliding_window_buckets` therefore picks the largest common
   divisor of the non-sliding buckets instead, when the smallest bucket is sliding-window-only:
   16/48/5 -> group size 8 (no full/GDN padding, 3 padding layers on the 5-layer window group,
   which costs ~7 MB per request). Measured on the 3090: 105 -> 78 KB of pool per token, i.e.
   45,383 tokens at 40k context -> 69,758 at 64k. MTP mode has no sliding-window bucket, takes
   the early return, and keeps its 73,777-token pool.

2. v1/worker/gpu/model_runner.py — the V1 runner profiles its CUDA graphs and the worker
   subtracts the estimate from the KV budget; the V2 runner's `profile_cudagraph_memory`
   returns 0 ("TBD"), so its ~1.2 GiB of graphs lands ON TOP of gpu_memory_utilization: ask for
   0.93 and the process actually peaks near 0.98, which is how the DeltaNet spec path OOMs
   mid-request (main README, gotcha 4). Profiling them the way V1 does needs a temporary KV
   cache and graph pool; this instead makes the reservation explicit and measurable:
   VLLM_V2_CUDAGRAPH_MEM_MIB=<MiB> is subtracted before the KV cache is sized (0 = upstream
   behaviour). Read the real figure from the startup line "... and X GiB for CUDAGraph memory".
   Note the profiled activation peak itself varies by ~1 GiB between starts on this model, so
   for a fixed context length prefer pinning the pool with --kv-cache-memory (what
   single-user/start_qwen.sh does for SPEC=dflash2) over tuning gpu_memory_utilization.

Apply from the installed vllm package directory:
  patch -p1 -d venv/lib/python3.12/site-packages/vllm < patches/hybrid-kv-groups-v2-cudagraph.patch

Validated against vLLM 0.28.0. Reapply after upgrades.

Retired on the 0.29 port: the VLLM_V2_CUDAGRAPH_MEM_MIB reserve hunk in v1/worker/gpu/model_runner.py.
vLLM 0.29 profiles CUDA graph memory natively (cudagraph_utils.profile_cudagraph_memory, upstream vllm-project#53306), so the
reserve knob is a no-op there; only the kv_cache_utils hunks remain.

Source: syv-ai/qwen38-27b-rtx3090 patches/hybrid-kv-groups-v2-cudagraph.patch (at d4c4cc2). Applied as-is from the patch file; the preamble above is the patch's own.

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Michael Stufflebeam <cpuchip@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kimi mrv2 Model Runner V2 specific nvidia ready ONLY add when PR is ready to merge/full CI is needed

Projects

Status: Done

5 participants