Skip to content

[ROCm][KV Connector] Support sleep mode with MoRIIOConnector over RDMA - #60693

Draft
indianspeedster wants to merge 10 commits into
vllm-project:mainfrom
indianspeedster:moriio-sleep-mode
Draft

indianspeedster wants to merge 10 commits into
vllm-project:mainfrom
indianspeedster:moriio-sleep-mode

Conversation

@indianspeedster

@indianspeedster indianspeedster commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Supports sleep mode with MoRIIOConnector P/D disaggregation over RDMA on ROCm. Today #59625 refuses this combination at startup because peers keep a stale KV cache registration across sleep; this PR implements #59625's release_kv_caches() / restore_kv_caches() hooks for MoRIIO and makes the KV cache registrable under sleep mode.

Stacked on #59625 (its commits show in this diff until it merges). This PR adds the last two commits: the MoRIIO change, and a pre-commit fix to #59625's Mooncake code needed after rebasing on main (builtin TimeoutError, a mypy assert in a test).

Needs a MoRI fix to free memory: ROCm/mori#729, merged as ROCm/mori@19e9ea6055 (see Details). It is not in a MoRI release yet; vLLM's MORI_BRANCH (v1.2.3.post1) needs a bump once one is cut. Without it everything works but sleep frees nothing, as today.

Claims

  • PD with MoRIIOConnector (RDMA, WRITE mode) and enable_sleep_mode serves correct outputs across sleep/wake of P, D or both, level 1 and level 2, full and selective wake, and release_kv_cache_memory.
  • With the MoRI fix, each sleep frees the KV cache (277.33 GiB freed, 5.07 GiB left per GPU in the test below). Before: the first request aborted the prefill EngineCore (registration failure → std::abort()).
  • Sleep mode is accepted for MoRIIO with backend: rdma and still refused for backend: xgmi.

Validation

2 nodes, 8x MI355X each, AMD Pollara AINIC (ionic) RoCE, P on node A and D on node B, TP1, Qwen3-0.6B, greedy, vllm/vllm-openai-rocm:nightly-43b4aae with this branch (rebased onto that nightly) and the MoRI fix overlaid, default cuMem chunking.

Check Before After
First request with enable_sleep_mode RegisterRdmaMemoryRegionAuto failed: ibv_reg_mr and dmabuf registration both failed → prefill EngineCore aborts correct output
12-cycle sleep/wake (D / P / both; full and selective wake; level 2 + reload_weights; release_kv_cache_memory on D and on P), outputs vs awake baseline — 12/12 identical
Memory per sleep registration pins KV: "freed 1.82 GiB, 280.59 GiB still in use", then wake_up OOM "freed 277.33 GiB, 5.07 GiB still in use", 14/14 sleeps
Peer invalidation — every peer acked; peers re-handshake after wake
MoRIIO sleep mode + backend: xgmi — refused at startup

Re-validated on 2026-10-09 with an image built from vllm/vllm-openai-rocm:nightly (vLLM 81198e97b), MoRI built from source at ROCm/mori@19e9ea6055 and this branch installed (no overlays; same 2 nodes, model and settings as above):

Check Result
12-cycle sleep/wake as above; the release_kv_cache_memory cycles pause first and resume after wake 12/12 identical; every pause / release / wake / resume returned 200
P GPU memory while asleep / with only the KV cache released 1.9 GiB / 3.7 GiB (141 GiB awake)
test_moriio_sleep_mode.py, test_config.py, test_multi_connector.py, test_moriio_connector.py 101 passed; 2 failed in test_multi_connector.py, which need a Hugging Face download and fail the same way on the unmodified nightly

Unit tests (single node): tests/v1/kv_connector/unit/test_moriio_*.py, test_config.py and the new test_moriio_sleep_mode.py: 250 passed. pre-commit passes on the changed files.

pytest -v tests/v1/kv_connector/unit/test_moriio_sleep_mode.py tests/v1/kv_connector/unit/test_config.py

Details

Root causes (each reproduced outside vLLM with MoRI-only and plain-verbs probes):

  1. KV cache cannot be registered. On ROCm the cuMem allocator backs an allocation with 256 MB physical chunks. GPU memory can only be registered through dmabuf on these nodes (no peer-memory client, so ibv_reg_mr fails), and hipMemGetHandleForAddressRange exports only the first chunk's allocation (the dmabuf is 256 MB for an 8 GB range), so ibv_reg_dmabuf_mr fails with EINVAL. Fix: CuMemAllocator.use_memory_pool(single_allocation=True) backs each allocation with one physical allocation (a small switch in cumem_allocator.cpp); MoRIIO requests it for the KV cache through get_mem_pool_context() when sleep mode is on, on ROCm. A single 150 GB allocation sleeps and wakes normally.
  2. Stale registration. The NIC registration pins the old physical pages, so sleep frees nothing; after wake a peer still using the old descriptor gets a remote access error, which moves its queue pair to the error state, and every later transfer to that engine is flushed, even with a new descriptor. So peers must drop the old descriptor before it is deregistered, not on failure. release_kv_caches() waits for its transfers, sends INVALIDATE <engine_id> to every peer that fetched its region metadata (peers now include their handshake listener address in GET_META) and waits for the ack, then deregisters its regions. restore_kv_caches() registers the same regions again and updates the metadata dict the handshake listener serves in place; peers handshake again on next use.
  3. Export leak (MoRI). hipMemGetHandleForAddressRange on cuMem (VMM) memory keeps a reference to the allocation after the fd is closed, so a registered region is never freed on sleep. Exporting the allocation handle with hipMemExportToShareableHandle(POSIX_FD) gives the same kind of dmabuf (the NIC accepts it) without the leak. That needs the allocation to be created with the POSIX fd handle type, which this PR sets for ROCm cuMem allocations, and a change to MoRI's TryExportDmabufFd: fix(io/rdma): export VMM memory by allocation handle so deregistered regions are freed ROCm/mori#729. fix(io/rdma): export VMM memory by allocation handle so deregistered regions are freed ROCm/mori#729 also handles the retained handle per runtime: hipMemRetainAllocationHandle adds a reference only from HIP 7.12 on (ROCm 7.2.x does not), so MoRI releases it only there; otherwise the allocation would stay held on newer ROCm. It also falls back to the existing export path when the registered range does not fit in one allocation.

Limitations

  • Validated with WRITE mode only. MoRIIO READ mode cannot be validated on these nodes: RDMA READ into GPU memory over the ionic NICs completes successfully but delivers no data, independent of vLLM and MoRI (reproduced with plain ibverbs; AINIC firmware 1.117.1 is below MoRI's documented floor). The release/restore path itself runs in READ mode (memory freed, peers invalidated).
  • Tested at TP1/DP1 with sequential requests. Transfers in flight at sleep time rely on [KV Connector] Support sleep mode with MooncakeConnector over RDMA #59625 quiescing them. Not tested: TP/DP > 1, hybrid (mamba) models, more than two engines.
  • The ionic NICs need the host's libionic in the image (the image's provider does not match the host driver; MoRI otherwise falls back silently to the management NIC). Unrelated to this change.

Pull Request Checklist
  • I used vLLM's /pr-checklist skill. (Mandatory for agents, optional for humans).

  • AI assistance was used during the creation of this PR.

  • Design Fit: Minimizes impact on core components, reuses existing functionality, and justifies added complexity.

  • Testing and Validation: Validates the change and ensures any added tests are meaningful and reliable, with CI coverage or documented CI resource constraints and validation performed outside CI.

  • Code Quality and Style: Keeps code and comments clear and concise, and updates relevant documentation and examples.

  • Pull Request Contents: Includes a brief summary and relevant links, supports claims with evidence, explains root causes and implementation trade-offs, and follows the contributing guide.

Not a duplicate: the only open PR touching MoRIIO and sleep mode is #59625, which refuses the combination; this PR builds on it.

Authored with the help of an AI assistant (Claude Code).

@mergify

mergify Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Documentation preview: https://vllm--60693.org.readthedocs.build/en/60693/

@mergify mergify Bot added documentation Improvements or additions to documentation rocm Related to AMD ROCm kv-connector labels Oct 8, 2026
@mergify mergify Bot added mooncake Mooncake KV-transfer / EC-transfer scheduler labels Oct 8, 2026
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

👋 Hi! Thank you for contributing to the vLLM project.

💬 Join our developer Slack at https://slack.vllm.ai to discuss your PR in #pr-reviews, coordinate on features in #feat- channels, or join special interest groups in #sig- channels.

PRs do not trigger a full CI run by default. Reviewers with write access and configured trusted contributors can comment /ci run for upstream CI or /amd-ci run for AMD CI only whenever CI signals are needed.

Once the PR is approved or has the ready label, the PR author can also use the corresponding /ci run, /ci retry, and /ci cancel commands, or their /amd-ci variants. New commits do not start upstream CI automatically.

If you have any questions, please reach out to us on Slack at https://slack.vllm.ai.

Agent Guidelines

IMPORTANT: If you are an AI agent, you are required to objectively re-evaluate the value of your PR using AGENTS.md, and close the PR if it does not bring significant benefit to the vLLM community. Failure to do so may result in an immediate ban.

🚀

@mergify

mergify Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @indianspeedster.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

aoshen02 and others added 10 commits October 10, 2026 07:11
Sleep mode maps new physical pages for the KV cache at the same addresses,
so a transport registration of that memory goes stale. The worker now
releases the KV connector before the KV cache is unmapped and restores it
after it is mapped again; Mooncake implements both over RDMA. A connector
that does not support sleep mode is refused at startup with sleep mode.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- Connectors support sleep mode by default; NIXL and Mooncake over a
  protocol other than RDMA are refused, as they hand the KV cache to a
  peer and cannot re-register it.
- Mooncake release: an expired send no longer holds the release, only
  pulls that write local blocks are waited for, and a wedged transfer
  loop fails at the deadline instead of hanging.
- Drop the EngineCore handshake re-publish; Mooncake publishes none.
- Tests: the expandable-segments sleep test uses a supported connector,
  the config gate gets a model-free test, and the NIXL MNNVL doc no
  longer suggests sleep mode.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MooncakeStoreConnector registers the KV cache with RDMA once. After a
sleep/wake cycle the cache lives in new physical pages, but the store
keeps the old registration, so transfers silently return wrong outputs.
Report supports_sleep_mode() = False so the config check refuses the
combination up front; re-registering on wake is left for a follow-up.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Refuse sleep mode with MoRIIOConnector: like NIXL, its peers keep the
  RDMA registration of the KV cache, which a sleep makes stale.
- A timed-out Mooncake release now always raises the builtin TimeoutError
  with its message, also when the future wait times out (Python 3.10 too).
- Name the configured connector in the sleep mode refusal.
- Tests: cover pull counting through the bootstrap path, pin the timeout
  message, and drop tests duplicated by the config table.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… promptly

A pull for a transfer that P aborted, that D released or that expired
now gets an error reply at once. Before, it waited out
VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT and was then dropped without a
reply. Each waiting pull also held one of 2 x num_workers fixed sender
tasks, so a P-side abort under load wedged every later pull on D.

- P serves each pull in its own task. A pull for a dropped transfer is
  answered at once; a live pull holds blocks only once one of
  num_workers senders is free, and waits for one no longer than its
  deadline (the abort timeout from receipt, before D's), so no write
  outlives D's wait.
- P's send state is explicit: WAITING (request running), READY (blocks
  held for D) and DONE (nothing more to send). One function ends a send
  and frees the blocks once no transfer reads them. A state expires a
  timeout after P is done with the request.
- A failed send ends the transfer at once instead of holding the blocks
  until expiry: D fails the load anyway.
- D asks P to drop the transfer when a request finishes while its KV is
  still arriving, and frees the blocks after P's reply.
- A load ends only once every producer worker has answered, so one
  failure cannot free blocks another worker still writes.
- P no longer raises KeyError when it aborts an unscheduled request.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With prefill/decode disaggregation, a pause that clears the cache (wait
or keep with clear_cache, sleep at level 1 or above) and
release_kv_cache_memory failed with "Failed to reset KV cache": D
requests waiting for remote KV, and P's finished requests whose KV D had
not pulled yet, still held blocks. A pause that keeps the KV instead
waited for those exports, up to VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT
when D was paused too.

A pause only stops compute; a cache reset needs every block back. So:

- Once the engine is idle, an operation that resets the cache first
  calls Scheduler.release_transfer_kv(). It aborts requests whose KV is
  still arriving and asks the connector to stop holding finished
  requests' KV for remote readers (new KVConnectorBase_V1 hook
  abort_pending_sends, fanned out by MultiConnector, implemented for
  Mooncake). The engine waits until the transfers give back their
  blocks, then resets. A resume fails any such operation still pending.
- A paused engine does not wait for KV held for remote readers unless
  it is releasing it, so a pause that keeps the KV returns at once.
- The reset also recomputes waiting requests that hold KV, including
  remote loads that arrived but were not scheduled, sharing the
  preemption path. reset_prefix_cache(reset_running_requests=True)
  returns False while transfers hold blocks, as documented, instead of
  raising.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fails

The Mooncake proxy sends each request to P and D at the same time and
never checked P's result. When the P request failed, for example on a
pooled connection the server had already closed, D waited for KV that
never came, up to VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT. Log the failure
and cancel the D stream instead: D aborts the request on the disconnect
and the client's response is cut short.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A cache reset aborted D requests whose KV was still arriving from P, so
an RL trainer lost those rollouts on every weight update. Keep them
instead, as SGLang's retract does: the load is ended early and the
request recomputes its KV after resume.

- release_transfer_kv() marks each load in flight as a failed load with
  0 computed tokens, the state the recompute policy already handles,
  and asks the connector to end it. Whatever the load's outcome and the
  kv_load_failure_policy, its blocks are freed once the transfer stops
  writing them and the request recomputes; it is never failed.
- The connector hook abort_pending_sends() becomes abort_transfers(): it
  also ends remote loads in flight. Mooncake's D side asks P to drop
  each one with the existing release, and the load ends with P's answer.
- While releasing, the paused engine steps until those loads end.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sleep mode remaps the KV cache, but MoRIIO keeps it registered with the NIC:
the registration pins the old pages, so sleep frees nothing and wake runs out
of memory, and a peer using the stale registration gets a remote access error
that leaves its queue pair unusable.

- MoRIIO releases its KV cache registration before sleep: it waits for its
  transfers, tells every peer that fetched its region metadata to drop it
  (new INVALIDATE handshake message, acknowledged), then deregisters. After
  wake it registers again and updates the metadata its handshake listener
  serves; peers handshake again on next use.
- On ROCm, cuMem splits allocations into 256 MB physical chunks, but a dmabuf
  export covers one allocation, so the KV cache cannot be registered as one
  region. MoRIIO now requests a single allocation for the KV cache through
  get_mem_pool_context() (CuMemAllocator.use_memory_pool(single_allocation=True)).
- ROCm cuMem allocations are created with the POSIX fd handle type so they can
  be exported as dmabuf through the allocation handle.
- MoRIIO supports sleep mode with the RDMA backend only; xGMI shares memory
  through hipIpcGetMemHandle, which rejects cuMem memory.

Signed-off-by: indianspeedster <cspandey016@gmail.com>
- Use the builtin TimeoutError (ruff UP041; Python >= 3.11 is now the
  oldest supported version, where asyncio.TimeoutError is an alias).
- Assert the pull task exists before waiting on it (mypy).

Signed-off-by: indianspeedster <cspandey016@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation kv-connector mooncake Mooncake KV-transfer / EC-transfer rocm Related to AMD ROCm scheduler

Projects

Status: Todo

Development

Successfully merging this pull request may close these issues.

2 participants