Skip to content

[KV Connector] Support sleep mode with MooncakeConnector over RDMA - #59625

Draft
aoshen02 wants to merge 8 commits into
vllm-project:mainfrom
aoshen02:mooncake-sleep-registration
Draft

aoshen02 wants to merge 8 commits into
vllm-project:mainfrom
aoshen02:mooncake-sleep-registration

Conversation

@aoshen02

@aoshen02 aoshen02 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Sleep mode with MooncakeConnector P/D disaggregation is silently wrong on main. Sleep unmaps the KV cache, and wake maps new physical pages at the same addresses. But the RDMA registration still points at the old pages. After a wake, transfers carry stale KV, so outputs are wrong with no error. The old pages also stay pinned.

This PR makes that combination work, including the pause/sleep lifecycle an RL trainer drives on both P and D. It is self-contained and includes the connector API from #59624.

  • Connector API (KVConnectorBase_V1):
    • supports_sleep_mode(kv_transfer_config), a classmethod. It defaults to True and returns False for a connector that hands KV device memory to a NIC or another process and cannot register it again;
    • release_kv_caches() and restore_kv_caches(), which are no-ops by default.
  • Worker:
    • Worker.sleep and Worker.discard(("kv_cache",)) call release_kv_caches() before the KV cache is unmapped;
    • Worker.wake_up calls restore_kv_caches() once the KV cache is mapped again.
  • Config: with enable_sleep_mode, a connector that does not support sleep mode is refused at startup, before any model is loaded. MultiConnector supports it only if every child does.
    • Refused: NIXL and MoRIIOConnector (peers keep the stale registration), MooncakeConnector with a protocol other than RDMA, and MooncakeStoreConnector. On main, the last one gives wrong outputs after a wake and spreads them to other instances through the store; its fix is [KV Connector] Support sleep mode with MooncakeStoreConnector #59934, stacked on this PR.
    • Every other connector behaves as before. LMCacheMPConnector and FlexKVConnectorV1 hand the KV cache to another process; they are still accepted, but not verified with sleep mode.
  • Mooncake sleep:
    • supports sleep mode with mooncake_protocol: rdma only, where a peer drops a stale remote key on the failed access and looks it up again;
    • release waits until no ready block is waiting for D (expired sends excluded) and no pull that writes local blocks is in flight. The wait is bounded by VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT. Release then unregisters the KV cache memory;
    • restore registers the same addresses and sizes again. Both are idempotent.

Behavior change: sleep mode with NIXL, MoRIIO, Mooncake over a protocol other than RDMA, or MooncakeStore now fails at startup instead of silently corrupting KV after a wake. NIXL support is a follow-up (#59635).

Known limitation: if one TP/PP rank's release times out after the other ranks have already slept, the ranks end up in mixed states. This is the same as any partial sleep failure on main.

PD + RL lifecycle (last four commits)

An RL trainer pauses P and D, updates the weights, sleeps or releases KV memory, and resumes. On main plus the commits above, most of these operations fail under load:

  • An abort on P wedges D. P never notified D's waiting pulls, so each waited 480 s while holding one of P's 2 × num_workers fixed sender tasks.
  • Every operation that resets the KV cache fails with "Failed to reset KV cache". Transfers still hold blocks: D's loads, and P's finished requests that D has not pulled yet.

[KV Connector][Mooncake] Answer every pull and free dropped transfers promptly

  • One task per pull. A pull no longer delays others; it holds blocks only once one of num_workers senders is free.
  • Explicit send state. P's send state is WAITING (request running), READY (blocks held for D) or DONE (nothing more to send). _end_send is the only place that ends a send and frees blocks once no write reads them.
  • Every pull gets one prompt answer. This holds whether P aborted the transfer, D released it, it expired, or a send failed.
  • Bounded pulls. Each pull has a deadline (the abort timeout from receipt, 60 s before D's), so no write starts after D could give up.
  • Safe load completion on D. D asks P to drop a transfer whose request finished mid-load, and ends a load only after every producer answered.

[Core][KV Connector] Release transfer-held KV before a cache reset

  • Release, then reset. A pause only stops compute. Once idle, an operation that resets the cache (pause with clear_cache, sleep, release_kv_cache_memory) calls Scheduler.release_transfer_kv() once, waits for the transfers to give back their blocks, then resets.
    • It calls the new KVConnectorBase_V1.abort_transfers() hook. P stops holding finished requests' KV for D, and D asks P to drop loads still in flight. The hook is a no-op by default, MultiConnector fans it out, and Mooncake implements it.
  • Single entry point. This lives in one place, EngineCoreProc._when_idle. A resume fails any such operation still pending.
  • Keep pauses do not wait for exports.
  • The reset recomputes waiting requests that hold KV, including arrived remote loads. reset_prefix_cache(reset_running_requests=True) returns False instead of raising while transfers hold blocks.

[Core][KV Connector] Recompute remote loads cut short by a cache reset. A D request whose KV was still arriving is no longer aborted. Its load is ended early and marked as a failed load with 0 computed tokens, so it recomputes after resume through the existing failed-load recovery, under either kv_load_failure_policy. This is like SGLang's retract, with D recomputing. RL rollouts in flight survive a weight update.

[Examples][Mooncake] Stop the decode stream when its prefill request fails. The example proxy no longer leaves D waiting 480 s for KV that a failed P request never sends.

NIXL keeps the default no-op hook: D reads one-sided and cannot see a revoked lease, so P waits out its 30 s lease. The in-process engine cannot wait for transfers; its cache reset fails while they hold blocks, as on main.

Test Plan

  • Unit:
    • test_mooncake_connector.py: release/restore cycle for each role; waits for blocks ready to send and for a pull in flight, including one still querying the bootstrap server; an expired send or a notify-only pull does not block; timeout with its message, including a wedged transfer loop; idempotence.
    • test_config.py: startup gate for each connector, with and without sleep mode, including the Mooncake protocol and MultiConnector combinations (model-free, CPU).
    • test_sleep_mode_backend.py: the worker releases before unmap and restores after map, only when the KV cache is involved.
    • test_multi_connector.py: release/restore are forwarded to every child.
    • Mutation check: reverting each piece of the change makes at least one test fail.
  • Hardware, 1P1D, Mooncake over RDMA, outputs compared with the awake run (greedy, 30 requests after every wake):
    • 2× GB300 nodes (P and D on different nodes, RoCE 400G), Qwen3-8B, Mooncake 0.3.13.post1 and v0.3.14-rc1;
    • H200 single node, Qwen3-0.6B.

Test Result

PD + RL lifecycle (last four commits):

  • Hardware matrix: GB300 1P1D Mooncake RDMA, Qwen3-8B, 27 cases under load (600 requests, concurrency 96, client disconnects). Cases: pause abort/wait/keep on P, D and both (with and without clear_cache), sleep L0/L1/L2, abort_requests, reset_prefix_cache, release_kv_cache_memory, in-place weight reload, client disconnects.
    • This head on ce49174247: 27/27, before and after the recompute commit.
    • The same commits on main 18f8f96: 27/27 on every revision.
    • Main with the earlier commits: 6/22.
    • NIXL: 19/19 (P pause bounded by its 30 s lease).
  • Unit: scheduler, engine core, Mooncake, MultiConnector and NIXL suites.
    • 561 pass on 18f8f96, including the NIXL suites. The rest fail identically without these commits (single GPU, gated models).
    • 31/31 mutations caught.
  • Review: 6 rounds by three independent reviewers; all approve.

Sleep mode (earlier commits):

  • Unit (GB300 container, main ce49174247 + this PR; config, Mooncake, MultiConnector, worker sleep mode and MoRIIO unit tests): 278 passed, 25 skipped, 2 failed. Both failures also happen without this PR: one needs a gated HF model, the other needs ray. The Mooncake and config tests were repeated 5 times with no flakes (160 passed each time).
  • Mutation check: 19 mutants, all killed.
  • Hardware, GB300 cross-node, both Mooncake versions:
    • Sleep D, sleep P, or sleep both, 3 cycles each; selective wake in both orders; release_kv_cache_memory; P sleeping while D pulls under load: 30/30 after every wake, 0 failed transfers.
    • Level 2 + reload_weights on D and on P: 30/30.
    • TP2 P + TP2 D: 30/30, 0 failed transfers. Main gives 0/30 here with no error.
    • OffloadingConnector, SimpleCPUOffloadConnector, and MultiConnector(Mooncake + Offloading) with sleep mode: start and match the control after every wake.
    • Startup: NIXL, Mooncake over TCP, and Multi(Mooncake + NIXL) are refused before the model loads. Example, Offloading, Multi(Mooncake + Offloading), and NIXL without sleep mode start.
    • GPU memory: 2,077 MiB asleep and 85,101 MiB awake, the same in every cycle.
  • H200: the registrations are deleted on sleep and created again on wake. 30/30 after every wake, including selective wake and release_kv_cache_memory.
  • At scale, this head (d4bd6a6448), GB300 cross-node, Mooncake 0.3.13.post1 over RDMA. Cycles rotate D, P and both, alternating full and selective wake:
    • DeepSeek-V4-Flash, P TP4 + D TP4: 30/30 cycles. 900/900 probes after wake match the awake run, and 4,480/4,480 load requests (64-256 concurrent) succeed. 17,886 transfers, 0 failed. Asleep 8.2 GiB/GPU every cycle, no drift.
    • Qwen3-8B 1P1D: 200/200 cycles. 600/600 probes match, and 29,760/29,760 load requests succeed. 22,278 transfers, 0 failed, no GPU drift.
    • On main, the same V4-Flash setup keeps the KV cache pinned while asleep (198 GiB/GPU still in use), and the first wake runs out of GPU memory.
  • V4-Flash single node, P TP2 + D TP2, Mooncake 0.3.13.post1 and 0.3.14-rc1: sleep D / P / both ×3, selective wake in both orders, release_kv_cache_memory on D / P / both, and P sleep(mode=wait) under 128-concurrent load. 30/30 after every wake (36/36 rounds), 0 failed transfers, no GPU-memory drift.
  • Known: after D wakes, P's first write can hit the stale remote key once. Mooncake then refreshes it and retries, and outputs stay correct. The first request takes up to 3.3 s (TP2, GB300), about 1.5 s (V4-Flash TP4) or about 5 s (H200, Mooncake 0.3.13); on GB300 at TP1 it was not seen.

This PR contains AI-assisted code (Claude Code); commits carry a Co-authored-by trailer.

🤖 Generated with Claude Code

@mergify

mergify Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Documentation preview: https://vllm--59625.org.readthedocs.build/en/59625/

@mergify mergify Bot added documentation Improvements or additions to documentation mrv2 Model Runner V2 specific kv-connector mooncake Mooncake KV-transfer / EC-transfer labels Oct 1, 2026
@aoshen02
aoshen02 force-pushed the mooncake-sleep-registration branch from 61fec5b to 449eff5 Compare October 1, 2026 16:04
@aoshen02
aoshen02 force-pushed the mooncake-sleep-registration branch from 55a5d74 to 5741ac6 Compare October 1, 2026 17:37
@aoshen02
aoshen02 force-pushed the mooncake-sleep-registration branch 6 times, most recently from 8f40fc2 to c5eaeba Compare October 2, 2026 02:11
@mergify

mergify Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @aoshen02.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Oct 2, 2026
@aoshen02
aoshen02 force-pushed the mooncake-sleep-registration branch 2 times, most recently from 1a2bbf3 to 1cd7b97 Compare October 2, 2026 03:35
@aoshen02 aoshen02 changed the title [KV Connector] Mooncake: release the KV cache registration across sleep mode [KV Connector] Mooncake: support sleep mode over RDMA Oct 2, 2026
@mergify mergify Bot removed the needs-rebase label Oct 2, 2026
@aoshen02
aoshen02 force-pushed the mooncake-sleep-registration branch from 1cd7b97 to e0ecf32 Compare October 2, 2026 07:30
@aoshen02
aoshen02 force-pushed the mooncake-sleep-registration branch 6 times, most recently from 89a779c to 424d0b8 Compare October 4, 2026 03:06
@aoshen02 aoshen02 changed the title [KV Connector] Mooncake: support sleep mode over RDMA [KV Connector] Support sleep mode with MooncakeConnector over RDMA Oct 4, 2026
Sleep mode maps new physical pages for the KV cache at the same addresses,
so a transport registration of that memory goes stale. The worker now
releases the KV connector before the KV cache is unmapped and restores it
after it is mapped again; Mooncake implements both over RDMA. A connector
that does not support sleep mode is refused at startup with sleep mode.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@aoshen02
aoshen02 force-pushed the mooncake-sleep-registration branch from 424d0b8 to bc2507d Compare October 4, 2026 04:37
aoshen02 and others added 2 commits October 4, 2026 07:21
- Connectors support sleep mode by default; NIXL and Mooncake over a
  protocol other than RDMA are refused, as they hand the KV cache to a
  peer and cannot re-register it.
- Mooncake release: an expired send no longer holds the release, only
  pulls that write local blocks are waited for, and a wedged transfer
  loop fails at the deadline instead of hanging.
- Drop the EngineCore handshake re-publish; Mooncake publishes none.
- Tests: the expandable-segments sleep test uses a supported connector,
  the config gate gets a model-free test, and the NIXL MNNVL doc no
  longer suggests sleep mode.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MooncakeStoreConnector registers the KV cache with RDMA once. After a
sleep/wake cycle the cache lives in new physical pages, but the store
keeps the old registration, so transfers silently return wrong outputs.
Report supports_sleep_mode() = False so the config check refuses the
combination up front; re-registering on wake is left for a follow-up.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mergify

mergify Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Documentation preview: https://vllm--59625.org.readthedocs.build/en/59625/

- Refuse sleep mode with MoRIIOConnector: like NIXL, its peers keep the
  RDMA registration of the KV cache, which a sleep makes stale.
- A timed-out Mooncake release now always raises the builtin TimeoutError
  with its message, also when the future wait times out (Python 3.10 too).
- Name the configured connector in the sleep mode refusal.
- Tests: cover pull counting through the bootstrap path, pin the timeout
  message, and drop tests duplicated by the config table.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
aoshen02 and others added 3 commits October 5, 2026 08:46
… promptly

A pull for a transfer that P aborted, that D released or that expired
now gets an error reply at once. Before, it waited out
VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT and was then dropped without a
reply. Each waiting pull also held one of 2 x num_workers fixed sender
tasks, so a P-side abort under load wedged every later pull on D.

- P serves each pull in its own task. A pull for a dropped transfer is
  answered at once; a live pull holds blocks only once one of
  num_workers senders is free, and waits for one no longer than its
  deadline (the abort timeout from receipt, before D's), so no write
  outlives D's wait.
- P's send state is explicit: WAITING (request running), READY (blocks
  held for D) and DONE (nothing more to send). One function ends a send
  and frees the blocks once no transfer reads them. A state expires a
  timeout after P is done with the request.
- A failed send ends the transfer at once instead of holding the blocks
  until expiry: D fails the load anyway.
- D asks P to drop the transfer when a request finishes while its KV is
  still arriving, and frees the blocks after P's reply.
- A load ends only once every producer worker has answered, so one
  failure cannot free blocks another worker still writes.
- P no longer raises KeyError when it aborts an unscheduled request.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With prefill/decode disaggregation, a pause that clears the cache (wait
or keep with clear_cache, sleep at level 1 or above) and
release_kv_cache_memory failed with "Failed to reset KV cache": D
requests waiting for remote KV, and P's finished requests whose KV D had
not pulled yet, still held blocks. A pause that keeps the KV instead
waited for those exports, up to VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT
when D was paused too.

A pause only stops compute; a cache reset needs every block back. So:

- Once the engine is idle, an operation that resets the cache first
  calls Scheduler.release_transfer_kv(). It aborts requests whose KV is
  still arriving and asks the connector to stop holding finished
  requests' KV for remote readers (new KVConnectorBase_V1 hook
  abort_pending_sends, fanned out by MultiConnector, implemented for
  Mooncake). The engine waits until the transfers give back their
  blocks, then resets. A resume fails any such operation still pending.
- A paused engine does not wait for KV held for remote readers unless
  it is releasing it, so a pause that keeps the KV returns at once.
- The reset also recomputes waiting requests that hold KV, including
  remote loads that arrived but were not scheduled, sharing the
  preemption path. reset_prefix_cache(reset_running_requests=True)
  returns False while transfers hold blocks, as documented, instead of
  raising.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fails

The Mooncake proxy sends each request to P and D at the same time and
never checked P's result. When the P request failed, for example on a
pooled connection the server had already closed, D waited for KV that
never came, up to VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT. Log the failure
and cancel the D stream instead: D aborts the request on the disconnect
and the client's response is cut short.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
aoshen02 added a commit to aoshen02/vllm that referenced this pull request Oct 5, 2026
Sleep mode keeps the KV cache virtual addresses but maps new physical
pages, so NIXL's memory registration, local prepped dlists, published
handshake payload and every peer's loaded remote agent go stale. On
GB300 (dmabuf, no nvidia_peermem) the stale MR pins the old pages: peers
read old KV without any error and sleep frees nothing.

Implement the vllm-project#59625 hooks for NixlPullConnector:
- release_kv_caches waits for own reads and pending handshakes, releases
  the local dlists, drops the remote engines and deregisters.
- restore_kv_caches registers again and rebuilds the payload with the
  new agent metadata.
- After a wake-up that restored the KV cache, EngineCore hands the
  workers' payloads to the scheduler, which bumps its registration
  epoch and stamps it on every rank's payload; request_finished
  advertises it as remote_registration_epoch. A reader whose loaded
  agent is older drops it and handshakes again, deferring while reads
  through it run.
- Under sleep mode UCX_RCACHE_ENABLE=n and UCX_TLS=^cuda_ipc are set
  process-wide unless already set, and logged once: UCX's registration
  cache and peers' cuda_ipc imports keep released pages pinned.

Only the UCX backend is accepted; NixlPushConnector stays refused.
NIXL_CONNECTOR_VERSION 13 -> 14.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mergify mergify Bot added the scheduler label Oct 5, 2026
A cache reset aborted D requests whose KV was still arriving from P, so
an RL trainer lost those rollouts on every weight update. Keep them
instead, as SGLang's retract does: the load is ended early and the
request recomputes its KV after resume.

- release_transfer_kv() marks each load in flight as a failed load with
  0 computed tokens, the state the recompute policy already handles,
  and asks the connector to end it. Whatever the load's outcome and the
  kv_load_failure_policy, its blocks are freed once the transfer stops
  writing them and the request recomputes; it is never failed.
- The connector hook abort_pending_sends() becomes abort_transfers(): it
  also ends remote loads in flight. Mooncake's D side asks P to drop
  each one with the existing release, and the load ends with P's answer.
- While releasing, the paused engine steps until those loads end.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mergify

mergify Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @aoshen02.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@andakai andakai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Generally LGTM! I have tested and it works well.

One point to discuss about: on main, P waits for D to pull the KV. With this PR, P drops it and D fails those requests under the default kv_load_failure_policy="fail". In a 1P1D test that pauses only P (mode="wait") after its prefills finish and before D pulls the KV, main succeeds on 64/64 requests and this PR fails. Should mode="wait" keep main’s behavior here?

Comment on lines +265 to +267
def release_kv_caches(self) -> None:
for c in self._connectors:
c.release_kv_caches()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This function only unregisters the kvcache, can confuse with release_kv_cache_memory? Maybe all rename to unregister_kv_caches()?

elif event == "short-timeout":
monkeypatch.setenv("VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT", "1")
elif event == "answered":
await asyncio.wait([pull], timeout=3)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

pull may be None when passed to asyncio.wait

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation kv-connector mooncake Mooncake KV-transfer / EC-transfer mrv2 Model Runner V2 specific needs-rebase scheduler

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants