Skip to content

[Bugfix] Register KV caches with the connector outside the runtime pool - #59847

Draft
aoshen02 wants to merge 1 commit into
vllm-project:mainfrom
aoshen02:fix/kv-connector-register-outside-runtime-pool
Draft

aoshen02 wants to merge 1 commit into
vllm-project:mainfrom
aoshen02:fix/kv-connector-register-outside-runtime-pool

Conversation

@aoshen02

@aoshen02 aoshen02 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

#59158 runs model_runner.initialize_kv_cache inside the cuMem runtime pool. That call also covers the KV connector's register_kv_caches.

A connector that allocates a GPU buffer there and registers it for RDMA now gets runtime-tagged memory. LMCache PD (LMCacheConnectorV1, pd_buffer_device=cuda) does exactly this:

  • Sleep copies the buffer to CPU and unmaps it.
  • Wake maps new pages at the same address, but the NIXL registration still points at the old pages.

The result:

  • After a sleep/wake, the prefiller sends stale KV and decode outputs are silently wrong.
  • The registration keeps the old pages pinned. Sleep does not free the buffer, and after wake it is held twice.

Fix:

  • initialize_kv_cache (V1, V2, CPU) returns the KV tensors.
  • The worker passes them to a new runner method, init_kv_connector, after leaving the runtime pool.
  • Block tables and metadata stay in runtime, and the KV cache stays in its nested kv_cache pool. The connector receives the same tensors.
  • No connector-specific code. On XPU and without cumem the runtime pool is a nullcontext, so nothing changes there.

Test Plan

  • test_runtime_state_survives_sleep is reworked into test_kv_init_placement_survives_sleep (V1/V2 × sleep level 1/2).
    • It drives the real Worker.initialize_from_config and the real init_kv_connector, with a fake connector that allocates a GPU tensor in register_kv_caches.
    • It checks that block tables are in runtime, the KV cache is in kv_cache, the connector gets the same KV tensor, and the connector buffer is outside every cuMem tag. Runtime state and KV survive sleep.
    • Mutation-checked: each of these mutants fails the test:
      • registration moved back inside the runtime pool (the regression);
      • no runtime pool;
      • V1 skips register_kv_caches;
      • V2 registers copies of the KV tensors.
  • Hardware: H200, LMCache 0.5.5 / NIXL 1.4.1 1P1D over UCX rc_mlx5, Qwen3-0.6B, 2 GiB PD buffer, sleep level 1 on P.
  • Without a connector (Qwen3-8B, sleep level 1): runtime tag, freed memory and greedy output are identical to main.
  • test_cumem.py passes. kv_connector unit tests and worker tests show no new failures (the ones that fail also fail on main in a 1-GPU offline environment). pre-commit including mypy passes.

Test Result

PD buffer tag after-wake outputs == pre-#59158 GPU MiB after wake (pre-#59158)
main, V2 runner runtime 0/8 (no error, no hang) 47277 (45227)
main, V1 runner runtime 0/8 45469 (43421)
this PR, V2 none 8/8 45229
this PR, V1 none 8/8 43421
  • Before sleep, all variants produce byte-identical outputs.
  • On main, the registration is still the same 9 MRs before and after wake.
  • On main, the allocator reports 2 GiB more freed than the device actually frees.

Re-check on current main (bdd31c31d), GB300

The patch rebases onto bdd31c31d with no conflicts, and main still runs register_kv_caches inside the runtime pool.

LMCache 1P1D, Qwen3-8B, 2 GiB PD buffer on the GPU, NIXL over RDMA. Level-1 sleep on P, 2 cycles; after each wake, 8 new prompts compared byte-for-byte with a run without sleep:

after wake 1 after wake 2 P memory after wake
main, Model Runner V2 0/8 0/8 88061 MiB
this PR, Model Runner V2 8/8 8/8 86045 MiB
main, Model Runner V1 0/8 0/8 87501 MiB
this PR, Model Runner V1 8/8 8/8 85485 MiB
#59625 alone 0/8 0/8
#59625 + this PR 8/8 8/8
  • On main, the LMCache PD buffer (2113929216 B) is tagged runtime. P holds 2016 MiB more after wake, which is the buffer size. No errors are logged.
  • Without a connector (Qwen3-8B), memory is identical with and without this PR: 85121 MiB awake, 2077 MiB asleep.
  • These runs use --attention-backend TRITON_ATTN: LMCache 0.5.5 P/D gives wrong outputs on GB300 with the FlashInfer backend even without sleep.

🤖 Generated with Claude Code

vllm-project#59158 runs model_runner.initialize_kv_cache inside the cuMem "runtime"
pool, which also covers the KV connector's register_kv_caches. A connector
that allocates and RDMA-registers a GPU buffer there (LMCache PD,
pd_buffer_device=cuda) now gets runtime-tagged memory: sleep copies it to
CPU and unmaps it, wake maps new pages at the same address, and the NIXL
registration still points at the old pages. After a sleep/wake the
prefiller sends stale KV and decode outputs are silently wrong. The old
pages also stay pinned by the registration, so sleep does not free the
buffer and after wake it is held twice.

initialize_kv_cache now returns the KV tensors, and the worker hands them
to a new runner method, init_kv_connector, after leaving the runtime pool.
Block tables and metadata stay in the runtime pool; the KV cache stays in
its nested kv_cache pool, and the connector sees the same tensors.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mergify mergify Bot added cpu Related to CPU backends mrv2 Model Runner V2 specific bug Something isn't working labels Oct 3, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working cpu Related to CPU backends mrv2 Model Runner V2 specific

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant