Repository navigation
[Feature][MRV2][P/D] Support PCP KV transfer - #16080
Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces support for Prefill Context Parallelism (PCP) KV transfer within the MRV2 P/D disaggregation architecture. By moving from a sharding-based approach to a per-request selection of a complete prefill replica, the system now reuses existing TP/group routing and completion tracking to manage KV cache release more effectively. This change enables prefill-side PCP while disabling decode-side PCP, providing a more robust solution for distributed inference. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Ops][Feature] Refactor Mooncake connector to separate PCP and DCP handlingSuggested PR Summary:
### What this PR does / why we need it?
This PR refactors the Mooncake connector's Context Parallelism (CP) handling by separating Prefill Context Parallelism (PCP) and Decode Context Parallelism (DCP) logic. Specifically, PCP now selects one complete P-side KV replica, whereas DCP splits prompt blocks across P workers. It also restricts decode-side PCP size to 1 for consumers and updates the corresponding tests.
I have no feedback to provide as there are no review comments.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Tested with updated and new unit tests in `tests/ut/kv_offload/test_mooncake_connector.py`.19c635c to
c099b60
Compare
9a603f0 to
d9513e8
Compare
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
🟡 Changes recommended
A verified bug in the HMA+DCP group-pull path can miscompute prefill_pp_rank when PCP replicas are used, which risks incorrect routing/reformatting.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
This PR updates MooncakeConnectorV1’s KV-transfer logic for MRV2 P/D disaggregation to replace PCP KV sharding with per-request selection of a complete prefill-side PCP replica, reusing existing TP/group routing and completion tracking to release KV safely.
Changes:
- Update KV block selection, port routing, and group-pull metadata to support PCP-as-replica-selection (with DCP remaining the sharding dimension).
- Enforce a configuration constraint that kv_consumer (decode side) must not enable PCP.
- Refresh and extend unit tests and update PCP feature documentation to reflect P/D disaggregation support.
File summaries
| File | Description |
|---|---|
| vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py | Implements PCP replica selection for KV pulls, adjusts block-id trimming and group-pull construction, and adds a decode-side PCP guard. |
| tests/ut/kv_offload/test_mooncake_connector.py | Updates UTs for the new PCP/DCP semantics and adds coverage for replica routing/completion tracking. |
| docs/source/user_guide/feature_guide/context_parallel.md | Updates the PCP support matrix to indicate P/D disaggregation compatibility. |
Review details
Suppressed comments (1)
vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py:3360
- In the HMA+DCP path,
pp_rankis derived from(port - remote_base_port) // prefill_tp_size, which treats the PCP segment offset as a PP rank whenremote_pcp_size > 1. This can yield out-of-rangeprefill_pp_rankvalues (PP size is 1 when PCP>1) and break group reformatting/routing. Compute PP rank the same way as the non-HMA path by modulo’ing out any replica segment offsets.
for port_idx, port in enumerate(ports):
pulls = []
port_tp = (port - remote_base_port) % prefill_tp_size
pp_rank = (port - remote_base_port) // prefill_tp_size
# Attention uses the leading ports selected for each DCP shard.
- Files reviewed: 3/3 changed files
- Comments generated: 2
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
d9513e8 to
ebd9ef1
Compare
dc42d7e to
d370cea
Compare
d370cea to
e0e3f4d
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
e0e3f4d to
1ba9370
Compare
MRV2 PCP stores a complete KV replica on each prefill PCP rank. Remove PCP from scheduler block counts, group-pull selection and sender tracking so these paths use only the actual KV sharding dimension. Remove the obsolete PCP shard cases while retaining DCP coverage. Signed-off-by: leolee <yihao.li@huawei.com>
Allow MRV2 prefill workers with PCP to serve decode workers without PCP through MooncakeConnectorV1 and MooncakeHybridConnector. Select one complete prefill replica by request ID so all decode TP ranks agree, independently of the existing TP routing. Add PCP-aware worker ports and remote replica sizes to the hybrid connector. Derive peer offsets from worker ranks relative to the prefill instance's base port, preserving the existing handshake metadata and host lookup. Keep TP/PP transfer offsets and block mappings unchanged, including compressed KV, SWA and state/indexer caches. Register requests before dispatching asynchronous reads. Only selected transfer sources wait for decode completion; unused PCP replicas finish locally through the existing request tracker. Keep Hybrid DCP disabled and reject decode-side PCP. Signed-off-by: leolee <yihao.li@huawei.com>
1ba9370 to
fe92635
Compare
|
More information, pls check this : #15254 |
### What this PR does / why we need it? Replace PCP KV sharding with per-request selection of a complete prefill replica. Reuse TP/group routing and existing completion tracking to release KV. ### Does this PR introduce _any_ user-facing change? Enable prefill-side PCP with decode-side PCP disabled for MRV2 P/D disaggregation using MooncakeConnectorV1. ### How was this patch tested? GPQA Diamond first 100 questions, concurrency 100. Qwen3-8B BF16 uses TP2 on both sides; DeepSeek-V4-Flash-w4a8 (DSA) uses TP4 on both sides with MooncakeHybridConnector. | Model | Scenario | P PCP | D PCP | Accuracy | Valid answers | Output cap hits + failed requests (limit) | | --- | --- | ---: | ---: | ---: | ---: | ---: | | Qwen3-8B BF16 | PD baseline | 1 | 1 | 42/100 | 92 | 8 (4096) | | Qwen3-8B BF16 | PCP + PD | 2 | 1 | 44/100 | 95 | 6 (4096) | | DeepSeek-V4-Flash-w4a8 (DSA) | PD baseline | 1 | 1 | 72/100 | 98 | 2 (8192) | | DeepSeek-V4-Flash-w4a8 (DSA) | PCP + PD | 2 | 1 | 81/100 | 99 | 1 (8192) | Both Qwen3 runs passed 18 short/long/batched probes. Remote KV reads and cleanup were verified, with no duplicate DONEs or timeout reclamations. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: leolee <yihao.li@huawei.com>
### What this PR does / why we need it? Replace PCP KV sharding with per-request selection of a complete prefill replica. Reuse TP/group routing and existing completion tracking to release KV. ### Does this PR introduce _any_ user-facing change? Enable prefill-side PCP with decode-side PCP disabled for MRV2 P/D disaggregation using MooncakeConnectorV1. ### How was this patch tested? GPQA Diamond first 100 questions, concurrency 100. Qwen3-8B BF16 uses TP2 on both sides; DeepSeek-V4-Flash-w4a8 (DSA) uses TP4 on both sides with MooncakeHybridConnector. | Model | Scenario | P PCP | D PCP | Accuracy | Valid answers | Output cap hits + failed requests (limit) | | --- | --- | ---: | ---: | ---: | ---: | ---: | | Qwen3-8B BF16 | PD baseline | 1 | 1 | 42/100 | 92 | 8 (4096) | | Qwen3-8B BF16 | PCP + PD | 2 | 1 | 44/100 | 95 | 6 (4096) | | DeepSeek-V4-Flash-w4a8 (DSA) | PD baseline | 1 | 1 | 72/100 | 98 | 2 (8192) | | DeepSeek-V4-Flash-w4a8 (DSA) | PCP + PD | 2 | 1 | 81/100 | 99 | 1 (8192) | Both Qwen3 runs passed 18 short/long/batched probes. Remote KV reads and cleanup were verified, with no duplicate DONEs or timeout reclamations. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: leolee <yihao.li@huawei.com>
### What this PR does / why we need it? Replace PCP KV sharding with per-request selection of a complete prefill replica. Reuse TP/group routing and existing completion tracking to release KV. ### Does this PR introduce _any_ user-facing change? Enable prefill-side PCP with decode-side PCP disabled for MRV2 P/D disaggregation using MooncakeConnectorV1. ### How was this patch tested? GPQA Diamond first 100 questions, concurrency 100. Qwen3-8B BF16 uses TP2 on both sides; DeepSeek-V4-Flash-w4a8 (DSA) uses TP4 on both sides with MooncakeHybridConnector. | Model | Scenario | P PCP | D PCP | Accuracy | Valid answers | Output cap hits + failed requests (limit) | | --- | --- | ---: | ---: | ---: | ---: | ---: | | Qwen3-8B BF16 | PD baseline | 1 | 1 | 42/100 | 92 | 8 (4096) | | Qwen3-8B BF16 | PCP + PD | 2 | 1 | 44/100 | 95 | 6 (4096) | | DeepSeek-V4-Flash-w4a8 (DSA) | PD baseline | 1 | 1 | 72/100 | 98 | 2 (8192) | | DeepSeek-V4-Flash-w4a8 (DSA) | PCP + PD | 2 | 1 | 81/100 | 99 | 1 (8192) | Both Qwen3 runs passed 18 short/long/batched probes. Remote KV reads and cleanup were verified, with no duplicate DONEs or timeout reclamations. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: leolee <yihao.li@huawei.com> Signed-off-by: tianming2009 <13246728590@163.com>
### What this PR does / why we need it? Replace PCP KV sharding with per-request selection of a complete prefill replica. Reuse TP/group routing and existing completion tracking to release KV. ### Does this PR introduce _any_ user-facing change? Enable prefill-side PCP with decode-side PCP disabled for MRV2 P/D disaggregation using MooncakeConnectorV1. ### How was this patch tested? GPQA Diamond first 100 questions, concurrency 100. Qwen3-8B BF16 uses TP2 on both sides; DeepSeek-V4-Flash-w4a8 (DSA) uses TP4 on both sides with MooncakeHybridConnector. | Model | Scenario | P PCP | D PCP | Accuracy | Valid answers | Output cap hits + failed requests (limit) | | --- | --- | ---: | ---: | ---: | ---: | ---: | | Qwen3-8B BF16 | PD baseline | 1 | 1 | 42/100 | 92 | 8 (4096) | | Qwen3-8B BF16 | PCP + PD | 2 | 1 | 44/100 | 95 | 6 (4096) | | DeepSeek-V4-Flash-w4a8 (DSA) | PD baseline | 1 | 1 | 72/100 | 98 | 2 (8192) | | DeepSeek-V4-Flash-w4a8 (DSA) | PCP + PD | 2 | 1 | 81/100 | 99 | 1 (8192) | Both Qwen3 runs passed 18 short/long/batched probes. Remote KV reads and cleanup were verified, with no duplicate DONEs or timeout reclamations. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: leolee <yihao.li@huawei.com> Signed-off-by: like-0517 <ithwlike@126.com>
What this PR does / why we need it?
Replace PCP KV sharding with per-request selection of a complete prefill replica. Reuse TP/group routing and existing completion tracking to release KV.
Does this PR introduce any user-facing change?
Enable prefill-side PCP with decode-side PCP disabled for MRV2 P/D disaggregation using MooncakeConnectorV1.
How was this patch tested?
GPQA Diamond first 100 questions, concurrency 100. Qwen3-8B BF16 uses TP2 on both sides; DeepSeek-V4-Flash-w4a8 (DSA) uses TP4 on both sides with MooncakeHybridConnector.
Both Qwen3 runs passed 18 short/long/batched probes. Remote KV reads and cleanup were verified, with no duplicate DONEs or timeout reclamations.