Repository navigation
Conversation
vllm-project#59158 runs model_runner.initialize_kv_cache inside the cuMem "runtime" pool, which also covers the KV connector's register_kv_caches. A connector that allocates and RDMA-registers a GPU buffer there (LMCache PD, pd_buffer_device=cuda) now gets runtime-tagged memory: sleep copies it to CPU and unmaps it, wake maps new pages at the same address, and the NIXL registration still points at the old pages. After a sleep/wake the prefiller sends stale KV and decode outputs are silently wrong. The old pages also stay pinned by the registration, so sleep does not free the buffer and after wake it is held twice. initialize_kv_cache now returns the KV tensors, and the worker hands them to a new runner method, init_kv_connector, after leaving the runtime pool. Block tables and metadata stay in the runtime pool; the KV cache stays in its nested kv_cache pool, and the connector sees the same tensors. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
5 of 20 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
#59158 runs
model_runner.initialize_kv_cacheinside the cuMemruntimepool. That call also covers the KV connector'sregister_kv_caches.A connector that allocates a GPU buffer there and registers it for RDMA now gets runtime-tagged memory. LMCache PD (
LMCacheConnectorV1,pd_buffer_device=cuda) does exactly this:The result:
Fix:
initialize_kv_cache(V1, V2, CPU) returns the KV tensors.init_kv_connector, after leaving the runtime pool.runtime, and the KV cache stays in its nestedkv_cachepool. The connector receives the same tensors.nullcontext, so nothing changes there.Test Plan
test_runtime_state_survives_sleepis reworked intotest_kv_init_placement_survives_sleep(V1/V2 × sleep level 1/2).Worker.initialize_from_configand the realinit_kv_connector, with a fake connector that allocates a GPU tensor inregister_kv_caches.runtime, the KV cache is inkv_cache, the connector gets the same KV tensor, and the connector buffer is outside every cuMem tag. Runtime state and KV survive sleep.register_kv_caches;rc_mlx5, Qwen3-0.6B, 2 GiB PD buffer, sleep level 1 on P./sleep,/wake_up, then 8 new PD requests.test_cumem.pypasses. kv_connector unit tests and worker tests show no new failures (the ones that fail also fail on main in a 1-GPU offline environment). pre-commit including mypy passes.Test Result
runtimeruntimeRe-check on current main (
bdd31c31d), GB300The patch rebases onto
bdd31c31dwith no conflicts, and main still runsregister_kv_cachesinside theruntimepool.LMCache 1P1D, Qwen3-8B, 2 GiB PD buffer on the GPU, NIXL over RDMA. Level-1 sleep on P, 2 cycles; after each wake, 8 new prompts compared byte-for-byte with a run without sleep:
runtime. P holds 2016 MiB more after wake, which is the buffer size. No errors are logged.--attention-backend TRITON_ATTN: LMCache 0.5.5 P/D gives wrong outputs on GB300 with the FlashInfer backend even without sleep.🤖 Generated with Claude Code