[Core] Demo implementation of extensible kv cache memory - #47363
zhuohan123 wants to merge 3 commits into
Conversation
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.
Tip: disable this comment in your organization's Code Review settings.
Port of internal D110967544. The extensible KV cache flow (reserve KV virtual address space up front, capture CUDA graphs first, then size and commit the KV cache from post-capture free memory) previously required a block-major attention backend and rejected Mamba models. This enables it for every backend layout and for Mamba / linear attention: - ExtensibleTensor gains num_segments: the reservation is divided into equal segments that grow in lockstep, with committed bytes forming a prefix of each segment. Physical pages are mapped at allocation-granularity granules and deduped across overlapping ranges, so a granule straddling a segment boundary is mapped exactly once. resize_per_segment_(bytes, zero_new=True) zeroes only the newly committed logical range of each segment. - Each KV cache buffer keeps its layers' physical layout and is committed as one prefix per layout segment. The segment count is derived from the backend's get_kv_cache_shape / get_kv_cache_block_dim / stride order: K/V-split layouts (e.g. FlashAttention) get one prefix per half, block-major layouts (e.g. FlashInfer, MLA) a single prefix. Mamba state pages are block-major per layer, and hybrid-model attention caches are re-strided to block-major, so both use a single segment. - Removed the supports_extensible_kv_cache gate plumbing from EngineCore, Executor, Worker, WorkerBase and GPUModelRunner; a CUDA platform check remains in EngineCore. - enable_extensible_kv_cache is reported as unsupported by the V2 model runner so V2-default models fall back to the V1 runner (which implements the flow); also fixed initialize_kv_cache being called with the extensible kwarg on runners that do not accept it, which broke every default V2-runner boot on this branch. Tested on H100: - tests/utils_/test_extensible_tensor.py (5 passed, incl. new segmented lockstep-grow/zero, granule-dedup and invalid-usage tests) - tests/v1/worker/test_extensible_kv_cache.py (new, 6 passed: segment derivation, split grows both halves, block-major, legacy full commit, Mamba per-layer growth, hybrid attention+Mamba) - E2E Qwen3-0.6B greedy with VLLM_ATTENTION_BACKEND=FLASH_ATTN (a K/V-split backend the old gate rejected): extensible generations byte-identical to the legacy path; log shows reserve then "Extended KV cache to 34663 blocks". V2->V1 auto-fallback path verified as well.
| def _deleter(_managed_ptr: object) -> None: | ||
| _KEEPALIVE.pop(key, None) |
There was a problem hiding this comment.
⚪ Severity: LOW
The callback _deleter removes its own wrapper from _KEEPALIVE while still executing. Since this wrapper is the only reference keeping the CFUNCTYPE object alive, it gets garbage-collected during its execution, causing a Use-After-Free of the C thunk on return.
Helpful? Add 👍 / 👎
💡 Fix Suggestion
Suggestion: Assign the result of _KEEPALIVE.pop(key, None) to a local variable so the CFUNCTYPE object (and the other kept-alive objects) remain referenced until _deleter returns. Without this, the pop discards the only references, causing the CFUNCTYPE's C thunk to be freed while the callback is still on the native call stack (Use-After-Free). A local variable extends the objects' lifetime through the function return.
⚠️ Experimental Feature: This code suggestion is automatically generated. Please review carefully.
| def _deleter(_managed_ptr: object) -> None: | |
| _KEEPALIVE.pop(key, None) | |
| def _deleter(_managed_ptr: object) -> None: | |
| _prevent_gc = _KEEPALIVE.pop(key, None) # noqa: F841 |
|
This pull request has merge conflicts that must be resolved before it can be |
Reserve the KV cache address range with CUDA virtual memory, commit a minimal prefix before CUDA graph capture, measure real post-capture memory usage, then commit the final KV cache size with stable tensor addresses. Opt-in via --enable-extensible-kv-cache. Squashed pick of #47363. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Signed-off-by: Nick Hill <nickhill@us.ibm.com>
…armup memory Adapt #47363's extensible KV cache to the V2 model runner: - V2 allocation (gpu/attn_utils.py): reserve each KV cache tensor's full virtual range with ExtensibleTensor, committing a per-segment block prefix. Segment counts are derived from each backend's physical layout (block dim / stride order), with hybrid attention+Mamba forced block-major to match the re-strided layout. - V2 warmup writes to real block IDs (a contiguous prefix starting at 1), unlike V1's all-zero dummy block tables, so warmup_kernels and run_mixed_prefill_decode_warmup now commit exactly the block prefix they touch via a new ensure_kv_cache_blocks() hook. - Post-warmup measurement: instead of only the CUDA graph pool bytes, the worker measures actual non-KV memory in use after ALL warmup (retained worst-case activation segments, NCCL buffers, CUDA graphs) and reports the excess over the profiling estimate (CompilationTimes.cuda_graph renamed to warmup_memory). The engine's second sizing pass then commits a KV cache that leaves room for the real runtime working set - including the worst-case spec-decode logits all-gather that memory profiling misses today. - Gate extensible mode against KV connectors and sleep mode; drop it from the V2-unsupported feature list. - Extend tests/v1/worker/test_extensible_kv_cache.py with V2 coverage (segment inference, staged prefix commits, hybrid re-stride layout). Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0133xqsNmqLHG9Pyhr5wSp1D Signed-off-by: Nick Hill <nickhill@us.ibm.com>
Reserve the KV cache address range with CUDA virtual memory, commit a minimal prefix before CUDA graph capture, measure real post-capture memory usage, then commit the final KV cache size with stable tensor addresses. Opt-in via --enable-extensible-kv-cache. Squashed pick of #47363. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Signed-off-by: Nick Hill <nickhill@us.ibm.com>
…umbing Ported from the extensible-kv-cache branch (originally #47363): VmmDriver (CUDA cuMem* / ROCm hipMem* bindings with probe), ExtensibleTensor / ExtensibleKVCacheBuffers (grow-only per-segment prefix commits over a stable VA reservation), enable_extensible_kv_cache config/CLI plumbing, and the register_kv_caches views-are-authoritative contract note. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ho2gVXA5r7PhrM2sk8oJUn Signed-off-by: Nick Hill <nickhill123@gmail.com>
…umbing Ported from the extensible-kv-cache branch (originally #47363): VmmDriver (CUDA cuMem* / ROCm hipMem* bindings with probe), ExtensibleTensor / ExtensibleKVCacheBuffers (grow-only per-segment prefix commits over a stable VA reservation), enable_extensible_kv_cache config/CLI plumbing, and the register_kv_caches views-are-authoritative contract note. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ho2gVXA5r7PhrM2sk8oJUn Signed-off-by: Nick Hill <nickhill123@gmail.com>
…umbing Ported from the extensible-kv-cache branch (originally #47363): VmmDriver (CUDA cuMem* / ROCm hipMem* bindings with probe), ExtensibleTensor / ExtensibleKVCacheBuffers (grow-only per-segment prefix commits over a stable VA reservation), enable_extensible_kv_cache config/CLI plumbing, and the register_kv_caches views-are-authoritative contract note. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ho2gVXA5r7PhrM2sk8oJUn Signed-off-by: Nick Hill <nickhill123@gmail.com>
…umbing Ported from the extensible-kv-cache branch (originally #47363): VmmDriver (CUDA cuMem* / ROCm hipMem* bindings with probe), ExtensibleTensor / ExtensibleKVCacheBuffers (grow-only per-segment prefix commits over a stable VA reservation), enable_extensible_kv_cache config/CLI plumbing, and the register_kv_caches views-are-authoritative contract note. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ho2gVXA5r7PhrM2sk8oJUn Signed-off-by: Nick Hill <nickhill123@gmail.com>
…umbing Ported from the extensible-kv-cache branch (originally #47363): VmmDriver (CUDA cuMem* / ROCm hipMem* bindings with probe), ExtensibleTensor / ExtensibleKVCacheBuffers (grow-only per-segment prefix commits over a stable VA reservation), enable_extensible_kv_cache config/CLI plumbing, and the register_kv_caches views-are-authoritative contract note. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ho2gVXA5r7PhrM2sk8oJUn Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Port the library and configuration half of the extensible KV cache demo (#47363). ExtensibleTensor reserves a device virtual address range with CUDA VMM and commits physical pages incrementally, so a buffer can grow without moving its base pointer. The reservation is divided into equal segments that grow in lockstep, which backs KV cache layouts whose block dimension is not outermost, and `segment_view` exposes each committed prefix as a tensor whose storage covers exactly the backed bytes. `--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs and LLM but is not yet consumed; the V2 model runner integration follows. Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Nick Hill <nickhill123@gmail.com>
Purpose
This PR adds an opt-in extensible KV cache path for V1 CUDA workers so automatic KV cache sizing can account for the actual CUDA graph memory footprint instead of relying on a pre-capture estimate.
When enabled, vLLM reserves the upper-bound KV cache address range with CUDA virtual memory, commits only a small initial amount before graph capture, measures the real CUDA graph memory usage, then commits the final KV cache size afterward while keeping KV tensor addresses stable for captured graphs. The change also updates multimodal memory profiling to reserve the full persistent encoder-cache footprint, preventing KV sizing from overestimating available memory for multimodal models.
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.