Repository navigation
[ROCm][KV Connector] Support sleep mode with MoRIIOConnector over RDMA - #60693
indianspeedster wants to merge 10 commits into
Conversation
|
Documentation preview: https://vllm--60693.org.readthedocs.build/en/60693/ |
|
👋 Hi! Thank you for contributing to the vLLM project. 💬 Join our developer Slack at https://slack.vllm.ai to discuss your PR in PRs do not trigger a full CI run by default. Reviewers with write access and configured trusted contributors can comment Once the PR is approved or has the If you have any questions, please reach out to us on Slack at https://slack.vllm.ai. Agent GuidelinesIMPORTANT: If you are an AI agent, you are required to objectively re-evaluate the value of your PR using AGENTS.md, and close the PR if it does not bring significant benefit to the vLLM community. Failure to do so may result in an immediate ban. 🚀 |
|
This pull request has merge conflicts that must be resolved before it can be |
Sleep mode maps new physical pages for the KV cache at the same addresses, so a transport registration of that memory goes stale. The worker now releases the KV connector before the KV cache is unmapped and restores it after it is mapped again; Mooncake implements both over RDMA. A connector that does not support sleep mode is refused at startup with sleep mode. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- Connectors support sleep mode by default; NIXL and Mooncake over a protocol other than RDMA are refused, as they hand the KV cache to a peer and cannot re-register it. - Mooncake release: an expired send no longer holds the release, only pulls that write local blocks are waited for, and a wedged transfer loop fails at the deadline instead of hanging. - Drop the EngineCore handshake re-publish; Mooncake publishes none. - Tests: the expandable-segments sleep test uses a supported connector, the config gate gets a model-free test, and the NIXL MNNVL doc no longer suggests sleep mode. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MooncakeStoreConnector registers the KV cache with RDMA once. After a sleep/wake cycle the cache lives in new physical pages, but the store keeps the old registration, so transfers silently return wrong outputs. Report supports_sleep_mode() = False so the config check refuses the combination up front; re-registering on wake is left for a follow-up. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Refuse sleep mode with MoRIIOConnector: like NIXL, its peers keep the RDMA registration of the KV cache, which a sleep makes stale. - A timed-out Mooncake release now always raises the builtin TimeoutError with its message, also when the future wait times out (Python 3.10 too). - Name the configured connector in the sleep mode refusal. - Tests: cover pull counting through the bootstrap path, pin the timeout message, and drop tests duplicated by the config table. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… promptly A pull for a transfer that P aborted, that D released or that expired now gets an error reply at once. Before, it waited out VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT and was then dropped without a reply. Each waiting pull also held one of 2 x num_workers fixed sender tasks, so a P-side abort under load wedged every later pull on D. - P serves each pull in its own task. A pull for a dropped transfer is answered at once; a live pull holds blocks only once one of num_workers senders is free, and waits for one no longer than its deadline (the abort timeout from receipt, before D's), so no write outlives D's wait. - P's send state is explicit: WAITING (request running), READY (blocks held for D) and DONE (nothing more to send). One function ends a send and frees the blocks once no transfer reads them. A state expires a timeout after P is done with the request. - A failed send ends the transfer at once instead of holding the blocks until expiry: D fails the load anyway. - D asks P to drop the transfer when a request finishes while its KV is still arriving, and frees the blocks after P's reply. - A load ends only once every producer worker has answered, so one failure cannot free blocks another worker still writes. - P no longer raises KeyError when it aborts an unscheduled request. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With prefill/decode disaggregation, a pause that clears the cache (wait or keep with clear_cache, sleep at level 1 or above) and release_kv_cache_memory failed with "Failed to reset KV cache": D requests waiting for remote KV, and P's finished requests whose KV D had not pulled yet, still held blocks. A pause that keeps the KV instead waited for those exports, up to VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT when D was paused too. A pause only stops compute; a cache reset needs every block back. So: - Once the engine is idle, an operation that resets the cache first calls Scheduler.release_transfer_kv(). It aborts requests whose KV is still arriving and asks the connector to stop holding finished requests' KV for remote readers (new KVConnectorBase_V1 hook abort_pending_sends, fanned out by MultiConnector, implemented for Mooncake). The engine waits until the transfers give back their blocks, then resets. A resume fails any such operation still pending. - A paused engine does not wait for KV held for remote readers unless it is releasing it, so a pause that keeps the KV returns at once. - The reset also recomputes waiting requests that hold KV, including remote loads that arrived but were not scheduled, sharing the preemption path. reset_prefix_cache(reset_running_requests=True) returns False while transfers hold blocks, as documented, instead of raising. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fails The Mooncake proxy sends each request to P and D at the same time and never checked P's result. When the P request failed, for example on a pooled connection the server had already closed, D waited for KV that never came, up to VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT. Log the failure and cancel the D stream instead: D aborts the request on the disconnect and the client's response is cut short. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A cache reset aborted D requests whose KV was still arriving from P, so an RL trainer lost those rollouts on every weight update. Keep them instead, as SGLang's retract does: the load is ended early and the request recomputes its KV after resume. - release_transfer_kv() marks each load in flight as a failed load with 0 computed tokens, the state the recompute policy already handles, and asks the connector to end it. Whatever the load's outcome and the kv_load_failure_policy, its blocks are freed once the transfer stops writing them and the request recomputes; it is never failed. - The connector hook abort_pending_sends() becomes abort_transfers(): it also ends remote loads in flight. Mooncake's D side asks P to drop each one with the existing release, and the load ends with P's answer. - While releasing, the paused engine steps until those loads end. Signed-off-by: aoshen02 <aoshen@inferact.ai> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sleep mode remaps the KV cache, but MoRIIO keeps it registered with the NIC: the registration pins the old pages, so sleep frees nothing and wake runs out of memory, and a peer using the stale registration gets a remote access error that leaves its queue pair unusable. - MoRIIO releases its KV cache registration before sleep: it waits for its transfers, tells every peer that fetched its region metadata to drop it (new INVALIDATE handshake message, acknowledged), then deregisters. After wake it registers again and updates the metadata its handshake listener serves; peers handshake again on next use. - On ROCm, cuMem splits allocations into 256 MB physical chunks, but a dmabuf export covers one allocation, so the KV cache cannot be registered as one region. MoRIIO now requests a single allocation for the KV cache through get_mem_pool_context() (CuMemAllocator.use_memory_pool(single_allocation=True)). - ROCm cuMem allocations are created with the POSIX fd handle type so they can be exported as dmabuf through the allocation handle. - MoRIIO supports sleep mode with the RDMA backend only; xGMI shares memory through hipIpcGetMemHandle, which rejects cuMem memory. Signed-off-by: indianspeedster <cspandey016@gmail.com>
- Use the builtin TimeoutError (ruff UP041; Python >= 3.11 is now the oldest supported version, where asyncio.TimeoutError is an alias). - Assert the pull task exists before waiting on it (mypy). Signed-off-by: indianspeedster <cspandey016@gmail.com>
7753249 to
158d09a
Compare
Overview
Supports sleep mode with
MoRIIOConnectorP/D disaggregation over RDMA on ROCm. Today #59625 refuses this combination at startup because peers keep a stale KV cache registration across sleep; this PR implements #59625'srelease_kv_caches()/restore_kv_caches()hooks for MoRIIO and makes the KV cache registrable under sleep mode.Stacked on #59625 (its commits show in this diff until it merges). This PR adds the last two commits: the MoRIIO change, and a pre-commit fix to #59625's Mooncake code needed after rebasing on
main(builtinTimeoutError, a mypy assert in a test).Needs a MoRI fix to free memory: ROCm/mori#729, merged as ROCm/mori@19e9ea6055 (see Details). It is not in a MoRI release yet; vLLM's
MORI_BRANCH(v1.2.3.post1) needs a bump once one is cut. Without it everything works but sleep frees nothing, as today.Claims
MoRIIOConnector(RDMA, WRITE mode) andenable_sleep_modeserves correct outputs across sleep/wake of P, D or both, level 1 and level 2, full and selective wake, andrelease_kv_cache_memory.std::abort()).backend: rdmaand still refused forbackend: xgmi.Validation
2 nodes, 8x MI355X each, AMD Pollara AINIC (ionic) RoCE, P on node A and D on node B, TP1, Qwen3-0.6B, greedy,
vllm/vllm-openai-rocm:nightly-43b4aaewith this branch (rebased onto that nightly) and the MoRI fix overlaid, default cuMem chunking.enable_sleep_modeRegisterRdmaMemoryRegionAuto failed: ibv_reg_mr and dmabuf registration both failed→ prefill EngineCore abortsreload_weights;release_kv_cache_memoryon D and on P), outputs vs awake baselinewake_upOOMbackend: xgmiRe-validated on 2026-10-09 with an image built from
vllm/vllm-openai-rocm:nightly(vLLM81198e97b), MoRI built from source at ROCm/mori@19e9ea6055 and this branch installed (no overlays; same 2 nodes, model and settings as above):release_kv_cache_memorycycles pause first and resume after waketest_moriio_sleep_mode.py,test_config.py,test_multi_connector.py,test_moriio_connector.pytest_multi_connector.py, which need a Hugging Face download and fail the same way on the unmodified nightlyUnit tests (single node):
tests/v1/kv_connector/unit/test_moriio_*.py,test_config.pyand the newtest_moriio_sleep_mode.py: 250 passed.pre-commitpasses on the changed files.Details
Root causes (each reproduced outside vLLM with MoRI-only and plain-verbs probes):
ibv_reg_mrfails), andhipMemGetHandleForAddressRangeexports only the first chunk's allocation (the dmabuf is 256 MB for an 8 GB range), soibv_reg_dmabuf_mrfails with EINVAL. Fix:CuMemAllocator.use_memory_pool(single_allocation=True)backs each allocation with one physical allocation (a small switch incumem_allocator.cpp); MoRIIO requests it for the KV cache throughget_mem_pool_context()when sleep mode is on, on ROCm. A single 150 GB allocation sleeps and wakes normally.release_kv_caches()waits for its transfers, sendsINVALIDATE <engine_id>to every peer that fetched its region metadata (peers now include their handshake listener address inGET_META) and waits for the ack, then deregisters its regions.restore_kv_caches()registers the same regions again and updates the metadata dict the handshake listener serves in place; peers handshake again on next use.hipMemGetHandleForAddressRangeon cuMem (VMM) memory keeps a reference to the allocation after the fd is closed, so a registered region is never freed on sleep. Exporting the allocation handle withhipMemExportToShareableHandle(POSIX_FD)gives the same kind of dmabuf (the NIC accepts it) without the leak. That needs the allocation to be created with the POSIX fd handle type, which this PR sets for ROCm cuMem allocations, and a change to MoRI'sTryExportDmabufFd: fix(io/rdma): export VMM memory by allocation handle so deregistered regions are freed ROCm/mori#729. fix(io/rdma): export VMM memory by allocation handle so deregistered regions are freed ROCm/mori#729 also handles the retained handle per runtime:hipMemRetainAllocationHandleadds a reference only from HIP 7.12 on (ROCm 7.2.x does not), so MoRI releases it only there; otherwise the allocation would stay held on newer ROCm. It also falls back to the existing export path when the registered range does not fit in one allocation.Limitations
libionicin the image (the image's provider does not match the host driver; MoRI otherwise falls back silently to the management NIC). Unrelated to this change.Pull Request Checklist
I used vLLM's
/pr-checklistskill. (Mandatory for agents, optional for humans).AI assistance was used during the creation of this PR.
Design Fit: Minimizes impact on core components, reuses existing functionality, and justifies added complexity.
Testing and Validation: Validates the change and ensures any added tests are meaningful and reliable, with CI coverage or documented CI resource constraints and validation performed outside CI.
Code Quality and Style: Keeps code and comments clear and concise, and updates relevant documentation and examples.
Pull Request Contents: Includes a brief summary and relevant links, supports claims with evidence, explains root causes and implementation trade-offs, and follows the contributing guide.
Not a duplicate: the only open PR touching MoRIIO and sleep mode is #59625, which refuses the combination; this PR builds on it.
Authored with the help of an AI assistant (Claude Code).