Skip to content

[Core] Demo implementation of extensible kv cache memory - #47363

Open
zhuohan123 wants to merge 3 commits into
mainfrom
zhuohan/extensible-kv-cache-memory
Open

zhuohan123 wants to merge 3 commits into
mainfrom
zhuohan/extensible-kv-cache-memory

Conversation

@zhuohan123

@zhuohan123 zhuohan123 commented Jul 2, 2026 •

Copy link
Copy Markdown
Member

Purpose

This PR adds an opt-in extensible KV cache path for V1 CUDA workers so automatic KV cache sizing can account for the actual CUDA graph memory footprint instead of relying on a pre-capture estimate.

When enabled, vLLM reserves the upper-bound KV cache address range with CUDA virtual memory, commits only a small initial amount before graph capture, measures the real CUDA graph memory usage, then commits the final KV cache size afterward while keeping KV tensor addresses stable for captured graphs. The change also updates multimodal memory profiling to reserve the full persistent encoder-cache footprint, preventing KV sizing from overestimating available memory for multimodal models.


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

Port of internal D110967544. The extensible KV cache flow (reserve KV
virtual address space up front, capture CUDA graphs first, then size and
commit the KV cache from post-capture free memory) previously required a
block-major attention backend and rejected Mamba models. This enables it
for every backend layout and for Mamba / linear attention:

- ExtensibleTensor gains num_segments: the reservation is divided into
  equal segments that grow in lockstep, with committed bytes forming a
  prefix of each segment. Physical pages are mapped at
  allocation-granularity granules and deduped across overlapping ranges,
  so a granule straddling a segment boundary is mapped exactly once.
  resize_per_segment_(bytes, zero_new=True) zeroes only the newly
  committed logical range of each segment.
- Each KV cache buffer keeps its layers' physical layout and is committed
  as one prefix per layout segment. The segment count is derived from the
  backend's get_kv_cache_shape / get_kv_cache_block_dim / stride order:
  K/V-split layouts (e.g. FlashAttention) get one prefix per half,
  block-major layouts (e.g. FlashInfer, MLA) a single prefix. Mamba state
  pages are block-major per layer, and hybrid-model attention caches are
  re-strided to block-major, so both use a single segment.
- Removed the supports_extensible_kv_cache gate plumbing from EngineCore,
  Executor, Worker, WorkerBase and GPUModelRunner; a CUDA platform check
  remains in EngineCore.
- enable_extensible_kv_cache is reported as unsupported by the V2 model
  runner so V2-default models fall back to the V1 runner (which implements
  the flow); also fixed initialize_kv_cache being called with the
  extensible kwarg on runners that do not accept it, which broke every
  default V2-runner boot on this branch.

Tested on H100:
- tests/utils_/test_extensible_tensor.py (5 passed, incl. new segmented
  lockstep-grow/zero, granule-dedup and invalid-usage tests)
- tests/v1/worker/test_extensible_kv_cache.py (new, 6 passed: segment
  derivation, split grows both halves, block-major, legacy full commit,
  Mamba per-layer growth, hybrid attention+Mamba)
- E2E Qwen3-0.6B greedy with VLLM_ATTENTION_BACKEND=FLASH_ATTN (a K/V-split
  backend the old gate rejected): extensible generations byte-identical to
  the legacy path; log shows reserve then "Extended KV cache to 34663
  blocks". V2->V1 auto-fallback path verified as well.
Comment on lines +361 to +362
def _deleter(_managed_ptr: object) -> None:
_KEEPALIVE.pop(key, None)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚪ Severity: LOW

The callback _deleter removes its own wrapper from _KEEPALIVE while still executing. Since this wrapper is the only reference keeping the CFUNCTYPE object alive, it gets garbage-collected during its execution, causing a Use-After-Free of the C thunk on return.
Helpful? Add 👍 / 👎

💡 Fix Suggestion

Suggestion: Assign the result of _KEEPALIVE.pop(key, None) to a local variable so the CFUNCTYPE object (and the other kept-alive objects) remain referenced until _deleter returns. Without this, the pop discards the only references, causing the CFUNCTYPE's C thunk to be freed while the callback is still on the native call stack (Use-After-Free). A local variable extends the objects' lifetime through the function return.

⚠️ Experimental Feature: This code suggestion is automatically generated. Please review carefully.

Suggested change
def _deleter(_managed_ptr: object) -> None:
_KEEPALIVE.pop(key, None)
def _deleter(_managed_ptr: object) -> None:
_prevent_gc = _KEEPALIVE.pop(key, None) # noqa: F841

@mergify

mergify Bot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @zhuohan123.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Jul 21, 2026
ZJY0516 pushed a commit that referenced this pull request Jul 22, 2026
Reserve the KV cache address range with CUDA virtual memory, commit a
minimal prefix before CUDA graph capture, measure real post-capture
memory usage, then commit the final KV cache size with stable tensor
addresses. Opt-in via --enable-extensible-kv-cache.

Squashed pick of #47363.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Signed-off-by: Nick Hill <nickhill@us.ibm.com>
ZJY0516 pushed a commit that referenced this pull request Jul 22, 2026
…armup memory

Adapt #47363's extensible KV cache to the V2 model runner:

- V2 allocation (gpu/attn_utils.py): reserve each KV cache tensor's full
  virtual range with ExtensibleTensor, committing a per-segment block
  prefix. Segment counts are derived from each backend's physical layout
  (block dim / stride order), with hybrid attention+Mamba forced
  block-major to match the re-strided layout.
- V2 warmup writes to real block IDs (a contiguous prefix starting at 1),
  unlike V1's all-zero dummy block tables, so warmup_kernels and
  run_mixed_prefill_decode_warmup now commit exactly the block prefix
  they touch via a new ensure_kv_cache_blocks() hook.
- Post-warmup measurement: instead of only the CUDA graph pool bytes,
  the worker measures actual non-KV memory in use after ALL warmup
  (retained worst-case activation segments, NCCL buffers, CUDA graphs)
  and reports the excess over the profiling estimate
  (CompilationTimes.cuda_graph renamed to warmup_memory). The engine's
  second sizing pass then commits a KV cache that leaves room for the
  real runtime working set - including the worst-case spec-decode
  logits all-gather that memory profiling misses today.
- Gate extensible mode against KV connectors and sleep mode; drop it
  from the V2-unsupported feature list.
- Extend tests/v1/worker/test_extensible_kv_cache.py with V2 coverage
  (segment inference, staged prefix commits, hybrid re-stride layout).

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0133xqsNmqLHG9Pyhr5wSp1D
Signed-off-by: Nick Hill <nickhill@us.ibm.com>
njhill added a commit that referenced this pull request Jul 31, 2026
Reserve the KV cache address range with CUDA virtual memory, commit a
minimal prefix before CUDA graph capture, measure real post-capture
memory usage, then commit the final KV cache size with stable tensor
addresses. Opt-in via --enable-extensible-kv-cache.

Squashed pick of #47363.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Signed-off-by: Nick Hill <nickhill@us.ibm.com>
njhill added a commit that referenced this pull request Aug 18, 2026
…umbing

Ported from the extensible-kv-cache branch (originally #47363): VmmDriver
(CUDA cuMem* / ROCm hipMem* bindings with probe), ExtensibleTensor /
ExtensibleKVCacheBuffers (grow-only per-segment prefix commits over a
stable VA reservation), enable_extensible_kv_cache config/CLI plumbing,
and the register_kv_caches views-are-authoritative contract note.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ho2gVXA5r7PhrM2sk8oJUn
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Aug 18, 2026
…umbing

Ported from the extensible-kv-cache branch (originally #47363): VmmDriver
(CUDA cuMem* / ROCm hipMem* bindings with probe), ExtensibleTensor /
ExtensibleKVCacheBuffers (grow-only per-segment prefix commits over a
stable VA reservation), enable_extensible_kv_cache config/CLI plumbing,
and the register_kv_caches views-are-authoritative contract note.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ho2gVXA5r7PhrM2sk8oJUn
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Aug 19, 2026
…umbing

Ported from the extensible-kv-cache branch (originally #47363): VmmDriver
(CUDA cuMem* / ROCm hipMem* bindings with probe), ExtensibleTensor /
ExtensibleKVCacheBuffers (grow-only per-segment prefix commits over a
stable VA reservation), enable_extensible_kv_cache config/CLI plumbing,
and the register_kv_caches views-are-authoritative contract note.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ho2gVXA5r7PhrM2sk8oJUn
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Aug 19, 2026
…umbing

Ported from the extensible-kv-cache branch (originally #47363): VmmDriver
(CUDA cuMem* / ROCm hipMem* bindings with probe), ExtensibleTensor /
ExtensibleKVCacheBuffers (grow-only per-segment prefix commits over a
stable VA reservation), enable_extensible_kv_cache config/CLI plumbing,
and the register_kv_caches views-are-authoritative contract note.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ho2gVXA5r7PhrM2sk8oJUn
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Aug 19, 2026
…umbing

Ported from the extensible-kv-cache branch (originally #47363): VmmDriver
(CUDA cuMem* / ROCm hipMem* bindings with probe), ExtensibleTensor /
ExtensibleKVCacheBuffers (grow-only per-segment prefix commits over a
stable VA reservation), enable_extensible_kv_cache config/CLI plumbing,
and the register_kv_caches views-are-authoritative contract note.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ho2gVXA5r7PhrM2sk8oJUn
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 11, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 11, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 11, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 12, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 12, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 13, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 14, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 15, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 15, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 15, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 17, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 17, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 17, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 19, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 19, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 19, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 19, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 25, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 25, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 26, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 28, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill added a commit that referenced this pull request Sep 28, 2026
Port the library and configuration half of the extensible KV cache demo
(#47363). ExtensibleTensor reserves a device virtual address range with
CUDA VMM and commits physical pages incrementally, so a buffer can grow
without moving its base pointer. The reservation is divided into equal
segments that grow in lockstep, which backs KV cache layouts whose block
dimension is not outermost, and `segment_view` exposes each committed
prefix as a tensor whose storage covers exactly the backed bytes.

`--enable-extensible-kv-cache` is plumbed through CacheConfig, EngineArgs
and LLM but is not yet consumed; the V2 model runner integration follows.

Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant