Repository navigation
Conversation
|
Documentation preview: https://vllm--59625.org.readthedocs.build/en/59625/ |
61fec5b to
449eff5
Compare
55a5d74 to
5741ac6
Compare
8f40fc2 to
c5eaeba
Compare
|
This pull request has merge conflicts that must be resolved before it can be |
1a2bbf3 to
1cd7b97
Compare
1cd7b97 to
e0ecf32
Compare
89a779c to
424d0b8
Compare
Sleep mode maps new physical pages for the KV cache at the same addresses, so a transport registration of that memory goes stale. The worker now releases the KV connector before the KV cache is unmapped and restores it after it is mapped again; Mooncake implements both over RDMA. A connector that does not support sleep mode is refused at startup with sleep mode. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
424d0b8 to
bc2507d
Compare
- Connectors support sleep mode by default; NIXL and Mooncake over a protocol other than RDMA are refused, as they hand the KV cache to a peer and cannot re-register it. - Mooncake release: an expired send no longer holds the release, only pulls that write local blocks are waited for, and a wedged transfer loop fails at the deadline instead of hanging. - Drop the EngineCore handshake re-publish; Mooncake publishes none. - Tests: the expandable-segments sleep test uses a supported connector, the config gate gets a model-free test, and the NIXL MNNVL doc no longer suggests sleep mode. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MooncakeStoreConnector registers the KV cache with RDMA once. After a sleep/wake cycle the cache lives in new physical pages, but the store keeps the old registration, so transfers silently return wrong outputs. Report supports_sleep_mode() = False so the config check refuses the combination up front; re-registering on wake is left for a follow-up. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Documentation preview: https://vllm--59625.org.readthedocs.build/en/59625/ |
- Refuse sleep mode with MoRIIOConnector: like NIXL, its peers keep the RDMA registration of the KV cache, which a sleep makes stale. - A timed-out Mooncake release now always raises the builtin TimeoutError with its message, also when the future wait times out (Python 3.10 too). - Name the configured connector in the sleep mode refusal. - Tests: cover pull counting through the bootstrap path, pin the timeout message, and drop tests duplicated by the config table. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… promptly A pull for a transfer that P aborted, that D released or that expired now gets an error reply at once. Before, it waited out VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT and was then dropped without a reply. Each waiting pull also held one of 2 x num_workers fixed sender tasks, so a P-side abort under load wedged every later pull on D. - P serves each pull in its own task. A pull for a dropped transfer is answered at once; a live pull holds blocks only once one of num_workers senders is free, and waits for one no longer than its deadline (the abort timeout from receipt, before D's), so no write outlives D's wait. - P's send state is explicit: WAITING (request running), READY (blocks held for D) and DONE (nothing more to send). One function ends a send and frees the blocks once no transfer reads them. A state expires a timeout after P is done with the request. - A failed send ends the transfer at once instead of holding the blocks until expiry: D fails the load anyway. - D asks P to drop the transfer when a request finishes while its KV is still arriving, and frees the blocks after P's reply. - A load ends only once every producer worker has answered, so one failure cannot free blocks another worker still writes. - P no longer raises KeyError when it aborts an unscheduled request. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With prefill/decode disaggregation, a pause that clears the cache (wait or keep with clear_cache, sleep at level 1 or above) and release_kv_cache_memory failed with "Failed to reset KV cache": D requests waiting for remote KV, and P's finished requests whose KV D had not pulled yet, still held blocks. A pause that keeps the KV instead waited for those exports, up to VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT when D was paused too. A pause only stops compute; a cache reset needs every block back. So: - Once the engine is idle, an operation that resets the cache first calls Scheduler.release_transfer_kv(). It aborts requests whose KV is still arriving and asks the connector to stop holding finished requests' KV for remote readers (new KVConnectorBase_V1 hook abort_pending_sends, fanned out by MultiConnector, implemented for Mooncake). The engine waits until the transfers give back their blocks, then resets. A resume fails any such operation still pending. - A paused engine does not wait for KV held for remote readers unless it is releasing it, so a pause that keeps the KV returns at once. - The reset also recomputes waiting requests that hold KV, including remote loads that arrived but were not scheduled, sharing the preemption path. reset_prefix_cache(reset_running_requests=True) returns False while transfers hold blocks, as documented, instead of raising. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fails The Mooncake proxy sends each request to P and D at the same time and never checked P's result. When the P request failed, for example on a pooled connection the server had already closed, D waited for KV that never came, up to VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT. Log the failure and cancel the D stream instead: D aborts the request on the disconnect and the client's response is cut short. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sleep mode keeps the KV cache virtual addresses but maps new physical pages, so NIXL's memory registration, local prepped dlists, published handshake payload and every peer's loaded remote agent go stale. On GB300 (dmabuf, no nvidia_peermem) the stale MR pins the old pages: peers read old KV without any error and sleep frees nothing. Implement the vllm-project#59625 hooks for NixlPullConnector: - release_kv_caches waits for own reads and pending handshakes, releases the local dlists, drops the remote engines and deregisters. - restore_kv_caches registers again and rebuilds the payload with the new agent metadata. - After a wake-up that restored the KV cache, EngineCore hands the workers' payloads to the scheduler, which bumps its registration epoch and stamps it on every rank's payload; request_finished advertises it as remote_registration_epoch. A reader whose loaded agent is older drops it and handshakes again, deferring while reads through it run. - Under sleep mode UCX_RCACHE_ENABLE=n and UCX_TLS=^cuda_ipc are set process-wide unless already set, and logged once: UCX's registration cache and peers' cuda_ipc imports keep released pages pinned. Only the UCX backend is accepted; NixlPushConnector stays refused. NIXL_CONNECTOR_VERSION 13 -> 14. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A cache reset aborted D requests whose KV was still arriving from P, so an RL trainer lost those rollouts on every weight update. Keep them instead, as SGLang's retract does: the load is ended early and the request recomputes its KV after resume. - release_transfer_kv() marks each load in flight as a failed load with 0 computed tokens, the state the recompute policy already handles, and asks the connector to end it. Whatever the load's outcome and the kv_load_failure_policy, its blocks are freed once the transfer stops writing them and the request recomputes; it is never failed. - The connector hook abort_pending_sends() becomes abort_transfers(): it also ends remote loads in flight. Mooncake's D side asks P to drop each one with the existing release, and the load ends with P's answer. - While releasing, the paused engine steps until those loads end. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
This pull request has merge conflicts that must be resolved before it can be |
andakai
left a comment
There was a problem hiding this comment.
Generally LGTM! I have tested and it works well.
One point to discuss about: on main, P waits for D to pull the KV. With this PR, P drops it and D fails those requests under the default kv_load_failure_policy="fail". In a 1P1D test that pauses only P (mode="wait") after its prefills finish and before D pulls the KV, main succeeds on 64/64 requests and this PR fails. Should mode="wait" keep main’s behavior here?
| def release_kv_caches(self) -> None: | ||
| for c in self._connectors: | ||
| c.release_kv_caches() |
There was a problem hiding this comment.
This function only unregisters the kvcache, can confuse with release_kv_cache_memory? Maybe all rename to unregister_kv_caches()?
| elif event == "short-timeout": | ||
| monkeypatch.setenv("VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT", "1") | ||
| elif event == "answered": | ||
| await asyncio.wait([pull], timeout=3) |
There was a problem hiding this comment.
pull may be None when passed to asyncio.wait
Purpose
Sleep mode with
MooncakeConnectorP/D disaggregation is silently wrong on main. Sleep unmaps the KV cache, and wake maps new physical pages at the same addresses. But the RDMA registration still points at the old pages. After a wake, transfers carry stale KV, so outputs are wrong with no error. The old pages also stay pinned.This PR makes that combination work, including the pause/sleep lifecycle an RL trainer drives on both P and D. It is self-contained and includes the connector API from #59624.
KVConnectorBase_V1):supports_sleep_mode(kv_transfer_config), a classmethod. It defaults toTrueand returnsFalsefor a connector that hands KV device memory to a NIC or another process and cannot register it again;release_kv_caches()andrestore_kv_caches(), which are no-ops by default.Worker.sleepandWorker.discard(("kv_cache",))callrelease_kv_caches()before the KV cache is unmapped;Worker.wake_upcallsrestore_kv_caches()once the KV cache is mapped again.enable_sleep_mode, a connector that does not support sleep mode is refused at startup, before any model is loaded.MultiConnectorsupports it only if every child does.MoRIIOConnector(peers keep the stale registration),MooncakeConnectorwith a protocol other than RDMA, andMooncakeStoreConnector. On main, the last one gives wrong outputs after a wake and spreads them to other instances through the store; its fix is [KV Connector] Support sleep mode with MooncakeStoreConnector #59934, stacked on this PR.LMCacheMPConnectorandFlexKVConnectorV1hand the KV cache to another process; they are still accepted, but not verified with sleep mode.mooncake_protocol: rdmaonly, where a peer drops a stale remote key on the failed access and looks it up again;VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT. Release then unregisters the KV cache memory;Behavior change: sleep mode with NIXL, MoRIIO, Mooncake over a protocol other than RDMA, or MooncakeStore now fails at startup instead of silently corrupting KV after a wake. NIXL support is a follow-up (#59635).
Known limitation: if one TP/PP rank's release times out after the other ranks have already slept, the ranks end up in mixed states. This is the same as any partial sleep failure on main.
PD + RL lifecycle (last four commits)
An RL trainer pauses P and D, updates the weights, sleeps or releases KV memory, and resumes. On main plus the commits above, most of these operations fail under load:
2 × num_workersfixed sender tasks.[KV Connector][Mooncake] Answer every pull and free dropped transfers promptly
num_workerssenders is free.WAITING(request running),READY(blocks held for D) orDONE(nothing more to send)._end_sendis the only place that ends a send and frees blocks once no write reads them.[Core][KV Connector] Release transfer-held KV before a cache reset
clear_cache, sleep,release_kv_cache_memory) callsScheduler.release_transfer_kv()once, waits for the transfers to give back their blocks, then resets.KVConnectorBase_V1.abort_transfers()hook. P stops holding finished requests' KV for D, and D asks P to drop loads still in flight. The hook is a no-op by default, MultiConnector fans it out, and Mooncake implements it.EngineCoreProc._when_idle. A resume fails any such operation still pending.reset_prefix_cache(reset_running_requests=True)returnsFalseinstead of raising while transfers hold blocks.[Core][KV Connector] Recompute remote loads cut short by a cache reset. A D request whose KV was still arriving is no longer aborted. Its load is ended early and marked as a failed load with 0 computed tokens, so it recomputes after resume through the existing failed-load recovery, under either
kv_load_failure_policy. This is like SGLang's retract, with D recomputing. RL rollouts in flight survive a weight update.[Examples][Mooncake] Stop the decode stream when its prefill request fails. The example proxy no longer leaves D waiting 480 s for KV that a failed P request never sends.
NIXL keeps the default no-op hook: D reads one-sided and cannot see a revoked lease, so P waits out its 30 s lease. The in-process engine cannot wait for transfers; its cache reset fails while they hold blocks, as on main.
Test Plan
test_mooncake_connector.py: release/restore cycle for each role; waits for blocks ready to send and for a pull in flight, including one still querying the bootstrap server; an expired send or a notify-only pull does not block; timeout with its message, including a wedged transfer loop; idempotence.test_config.py: startup gate for each connector, with and without sleep mode, including the Mooncake protocol andMultiConnectorcombinations (model-free, CPU).test_sleep_mode_backend.py: the worker releases before unmap and restores after map, only when the KV cache is involved.test_multi_connector.py: release/restore are forwarded to every child.Test Result
PD + RL lifecycle (last four commits):
clear_cache), sleep L0/L1/L2,abort_requests,reset_prefix_cache,release_kv_cache_memory, in-place weight reload, client disconnects.ce49174247: 27/27, before and after the recompute commit.18f8f96: 27/27 on every revision.18f8f96, including the NIXL suites. The rest fail identically without these commits (single GPU, gated models).Sleep mode (earlier commits):
ce49174247+ this PR; config, Mooncake, MultiConnector, worker sleep mode and MoRIIO unit tests): 278 passed, 25 skipped, 2 failed. Both failures also happen without this PR: one needs a gated HF model, the other needsray. The Mooncake and config tests were repeated 5 times with no flakes (160 passed each time).release_kv_cache_memory; P sleeping while D pulls under load: 30/30 after every wake, 0 failed transfers.reload_weightson D and on P: 30/30.OffloadingConnector,SimpleCPUOffloadConnector, andMultiConnector(Mooncake + Offloading)with sleep mode: start and match the control after every wake.release_kv_cache_memory.d4bd6a6448), GB300 cross-node, Mooncake 0.3.13.post1 over RDMA. Cycles rotate D, P and both, alternating full and selective wake:release_kv_cache_memoryon D / P / both, and Psleep(mode=wait)under 128-concurrent load. 30/30 after every wake (36/36 rounds), 0 failed transfers, no GPU-memory drift.This PR contains AI-assisted code (Claude Code); commits carry a
Co-authored-bytrailer.🤖 Generated with Claude Code