Skip to content

[Feature][MRV2][P/D] Support PCP KV transfer - #16080

Merged
Tflowers-0129 merged 2 commits into
vllm-project:mainfrom
li1how:feature/pcp-pd
Sep 10, 2026
Merged

Tflowers-0129 merged 2 commits into
vllm-project:mainfrom
li1how:feature/pcp-pd

Conversation

@li1how

@li1how li1how commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Replace PCP KV sharding with per-request selection of a complete prefill replica. Reuse TP/group routing and existing completion tracking to release KV.

Does this PR introduce any user-facing change?

Enable prefill-side PCP with decode-side PCP disabled for MRV2 P/D disaggregation using MooncakeConnectorV1.

How was this patch tested?

GPQA Diamond first 100 questions, concurrency 100. Qwen3-8B BF16 uses TP2 on both sides; DeepSeek-V4-Flash-w4a8 (DSA) uses TP4 on both sides with MooncakeHybridConnector.

Model Scenario P PCP D PCP Accuracy Valid answers Output cap hits + failed requests (limit)
Qwen3-8B BF16 PD baseline 1 1 42/100 92 8 (4096)
Qwen3-8B BF16 PCP + PD 2 1 44/100 95 6 (4096)
DeepSeek-V4-Flash-w4a8 (DSA) PD baseline 1 1 72/100 98 2 (8192)
DeepSeek-V4-Flash-w4a8 (DSA) PCP + PD 2 1 81/100 99 1 (8192)

Both Qwen3 runs passed 18 short/long/batched probes. Remote KV reads and cleanup were verified, with no duplicate DONEs or timeout reclamations.

@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces support for Prefill Context Parallelism (PCP) KV transfer within the MRV2 P/D disaggregation architecture. By moving from a sharding-based approach to a per-request selection of a complete prefill replica, the system now reuses existing TP/group routing and completion tracking to manage KV cache release more effectively. This change enables prefill-side PCP while disabling decode-side PCP, providing a more robust solution for distributed inference.

Highlights

  • PCP KV Transfer Strategy: Replaced the previous PCP KV sharding mechanism with a per-request selection of a complete prefill replica, improving efficiency in disaggregated P/D setups.
  • Logic Refactoring: Updated MooncakeConnectorWorker to handle PCP replica selection and refined the KV transfer block ID calculations to align with the new replica-based architecture.
  • Test Suite Enhancements: Expanded unit tests in test_mooncake_connector.py to cover new request handling scenarios, replica routing, and completion tracking logic.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Refactor Mooncake connector to separate PCP and DCP handling

Suggested PR Summary:

### What this PR does / why we need it?
This PR refactors the Mooncake connector's Context Parallelism (CP) handling by separating Prefill Context Parallelism (PCP) and Decode Context Parallelism (DCP) logic. Specifically, PCP now selects one complete P-side KV replica, whereas DCP splits prompt blocks across P workers. It also restricts decode-side PCP size to 1 for consumers and updates the corresponding tests.

I have no feedback to provide as there are no review comments.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Tested with updated and new unit tests in `tests/ut/kv_offload/test_mooncake_connector.py`.

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 8, 2026
@li1how
li1how force-pushed the feature/pcp-pd branch 2 times, most recently from 9a603f0 to d9513e8 Compare September 9, 2026 01:47
@li1how
li1how marked this pull request as ready for review September 9, 2026 01:47
@li1how
li1how requested review from LCAIZJ and Yikun as code owners September 9, 2026 01:47
Copilot AI lite review requested due to automatic review settings September 9, 2026 01:47
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 9, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
🔒 Security Review ✅ Completed 2026-09-09T01:54:59.631332Z d9513e8 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

A verified bug in the HMA+DCP group-pull path can miscompute prefill_pp_rank when PCP replicas are used, which risks incorrect routing/reformatting.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

This PR updates MooncakeConnectorV1’s KV-transfer logic for MRV2 P/D disaggregation to replace PCP KV sharding with per-request selection of a complete prefill-side PCP replica, reusing existing TP/group routing and completion tracking to release KV safely.

Changes:

  • Update KV block selection, port routing, and group-pull metadata to support PCP-as-replica-selection (with DCP remaining the sharding dimension).
  • Enforce a configuration constraint that kv_consumer (decode side) must not enable PCP.
  • Refresh and extend unit tests and update PCP feature documentation to reflect P/D disaggregation support.
File summaries
File Description
vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py Implements PCP replica selection for KV pulls, adjusts block-id trimming and group-pull construction, and adds a decode-side PCP guard.
tests/ut/kv_offload/test_mooncake_connector.py Updates UTs for the new PCP/DCP semantics and adds coverage for replica routing/completion tracking.
docs/source/user_guide/feature_guide/context_parallel.md Updates the PCP support matrix to indicate P/D disaggregation compatibility.
Review details

Suppressed comments (1)

vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py:3360

  • In the HMA+DCP path, pp_rank is derived from (port - remote_base_port) // prefill_tp_size, which treats the PCP segment offset as a PP rank when remote_pcp_size > 1. This can yield out-of-range prefill_pp_rank values (PP size is 1 when PCP>1) and break group reformatting/routing. Compute PP rank the same way as the non-HMA path by modulo’ing out any replica segment offsets.
            for port_idx, port in enumerate(ports):
                pulls = []
                port_tp = (port - remote_base_port) % prefill_tp_size
                pp_rank = (port - remote_base_port) // prefill_tp_size
                # Attention uses the leading ports selected for each DCP shard.
  • Files reviewed: 3/3 changed files
  • Comments generated: 2
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread docs/source/user_guide/feature_guide/context_parallel.md
Comment thread vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py Outdated
Comment thread vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py
Comment thread vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py
@Ronald1995 Ronald1995 added the ready-precise run selected e2e test for pr label Sep 9, 2026
@li1how
li1how force-pushed the feature/pcp-pd branch 3 times, most recently from dc42d7e to d370cea Compare September 9, 2026 08:25

@Ronald1995 Ronald1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

MRV2 PCP stores a complete KV replica on each prefill PCP rank. Remove
PCP from scheduler block counts, group-pull selection and sender tracking
so these paths use only the actual KV sharding dimension.

Remove the obsolete PCP shard cases while retaining DCP coverage.

Signed-off-by: leolee <yihao.li@huawei.com>
Allow MRV2 prefill workers with PCP to serve decode workers without PCP
through MooncakeConnectorV1 and MooncakeHybridConnector. Select one
complete prefill replica by request ID so all decode TP ranks agree,
independently of the existing TP routing.

Add PCP-aware worker ports and remote replica sizes to the hybrid
connector. Derive peer offsets from worker ranks relative to the prefill
instance's base port, preserving the existing handshake metadata and
host lookup. Keep TP/PP transfer offsets and block mappings unchanged,
including compressed KV, SWA and state/indexer caches.

Register requests before dispatching asynchronous reads. Only selected
transfer sources wait for decode completion; unused PCP replicas finish
locally through the existing request tracker. Keep Hybrid DCP disabled
and reject decode-side PCP.

Signed-off-by: leolee <yihao.li@huawei.com>
Comment thread tests/ut/kv_offload/test_mooncake_connector.py
@Tflowers-0129

Copy link
Copy Markdown
Collaborator

More information, pls check this : #15254

@Tflowers-0129
Tflowers-0129 merged commit 84b7f79 into vllm-project:main Sep 10, 2026
29 checks passed
@li1how
li1how deleted the feature/pcp-pd branch September 10, 2026 03:50
chen-commits pushed a commit to chen-commits/vllm-ascend that referenced this pull request Sep 10, 2026
### What this PR does / why we need it?

Replace PCP KV sharding with per-request selection of a complete prefill
replica. Reuse TP/group routing and existing completion tracking to
release KV.

### Does this PR introduce _any_ user-facing change?

Enable prefill-side PCP with decode-side PCP disabled for MRV2 P/D
disaggregation using MooncakeConnectorV1.

### How was this patch tested?

GPQA Diamond first 100 questions, concurrency 100. Qwen3-8B BF16 uses
TP2 on both sides; DeepSeek-V4-Flash-w4a8 (DSA) uses TP4 on both sides
with MooncakeHybridConnector.

| Model | Scenario | P PCP | D PCP | Accuracy | Valid answers | Output
cap hits + failed requests (limit) |
| --- | --- | ---: | ---: | ---: | ---: | ---: |
| Qwen3-8B BF16 | PD baseline | 1 | 1 | 42/100 | 92 | 8 (4096) |
| Qwen3-8B BF16 | PCP + PD | 2 | 1 | 44/100 | 95 | 6 (4096) |
| DeepSeek-V4-Flash-w4a8 (DSA) | PD baseline | 1 | 1 | 72/100 | 98 | 2
(8192) |
| DeepSeek-V4-Flash-w4a8 (DSA) | PCP + PD | 2 | 1 | 81/100 | 99 | 1
(8192) |

Both Qwen3 runs passed 18 short/long/batched probes. Remote KV reads and
cleanup were verified, with no duplicate DONEs or timeout reclamations.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: leolee <yihao.li@huawei.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

Replace PCP KV sharding with per-request selection of a complete prefill
replica. Reuse TP/group routing and existing completion tracking to
release KV.

### Does this PR introduce _any_ user-facing change?

Enable prefill-side PCP with decode-side PCP disabled for MRV2 P/D
disaggregation using MooncakeConnectorV1.

### How was this patch tested?

GPQA Diamond first 100 questions, concurrency 100. Qwen3-8B BF16 uses
TP2 on both sides; DeepSeek-V4-Flash-w4a8 (DSA) uses TP4 on both sides
with MooncakeHybridConnector.

| Model | Scenario | P PCP | D PCP | Accuracy | Valid answers | Output
cap hits + failed requests (limit) |
| --- | --- | ---: | ---: | ---: | ---: | ---: |
| Qwen3-8B BF16 | PD baseline | 1 | 1 | 42/100 | 92 | 8 (4096) |
| Qwen3-8B BF16 | PCP + PD | 2 | 1 | 44/100 | 95 | 6 (4096) |
| DeepSeek-V4-Flash-w4a8 (DSA) | PD baseline | 1 | 1 | 72/100 | 98 | 2
(8192) |
| DeepSeek-V4-Flash-w4a8 (DSA) | PCP + PD | 2 | 1 | 81/100 | 99 | 1
(8192) |

Both Qwen3 runs passed 18 short/long/batched probes. Remote KV reads and
cleanup were verified, with no duplicate DONEs or timeout reclamations.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: leolee <yihao.li@huawei.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
### What this PR does / why we need it?

Replace PCP KV sharding with per-request selection of a complete prefill
replica. Reuse TP/group routing and existing completion tracking to
release KV.

### Does this PR introduce _any_ user-facing change?

Enable prefill-side PCP with decode-side PCP disabled for MRV2 P/D
disaggregation using MooncakeConnectorV1.

### How was this patch tested?

GPQA Diamond first 100 questions, concurrency 100. Qwen3-8B BF16 uses
TP2 on both sides; DeepSeek-V4-Flash-w4a8 (DSA) uses TP4 on both sides
with MooncakeHybridConnector.

| Model | Scenario | P PCP | D PCP | Accuracy | Valid answers | Output
cap hits + failed requests (limit) |
| --- | --- | ---: | ---: | ---: | ---: | ---: |
| Qwen3-8B BF16 | PD baseline | 1 | 1 | 42/100 | 92 | 8 (4096) |
| Qwen3-8B BF16 | PCP + PD | 2 | 1 | 44/100 | 95 | 6 (4096) |
| DeepSeek-V4-Flash-w4a8 (DSA) | PD baseline | 1 | 1 | 72/100 | 98 | 2
(8192) |
| DeepSeek-V4-Flash-w4a8 (DSA) | PCP + PD | 2 | 1 | 81/100 | 99 | 1
(8192) |

Both Qwen3 runs passed 18 short/long/batched probes. Remote KV reads and
cleanup were verified, with no duplicate DONEs or timeout reclamations.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: leolee <yihao.li@huawei.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
### What this PR does / why we need it?

Replace PCP KV sharding with per-request selection of a complete prefill
replica. Reuse TP/group routing and existing completion tracking to
release KV.

### Does this PR introduce _any_ user-facing change?

Enable prefill-side PCP with decode-side PCP disabled for MRV2 P/D
disaggregation using MooncakeConnectorV1.

### How was this patch tested?

GPQA Diamond first 100 questions, concurrency 100. Qwen3-8B BF16 uses
TP2 on both sides; DeepSeek-V4-Flash-w4a8 (DSA) uses TP4 on both sides
with MooncakeHybridConnector.

| Model | Scenario | P PCP | D PCP | Accuracy | Valid answers | Output
cap hits + failed requests (limit) |
| --- | --- | ---: | ---: | ---: | ---: | ---: |
| Qwen3-8B BF16 | PD baseline | 1 | 1 | 42/100 | 92 | 8 (4096) |
| Qwen3-8B BF16 | PCP + PD | 2 | 1 | 44/100 | 95 | 6 (4096) |
| DeepSeek-V4-Flash-w4a8 (DSA) | PD baseline | 1 | 1 | 72/100 | 98 | 2
(8192) |
| DeepSeek-V4-Flash-w4a8 (DSA) | PCP + PD | 2 | 1 | 81/100 | 99 | 1
(8192) |

Both Qwen3 runs passed 18 short/long/batched probes. Remote KV reads and
cleanup were verified, with no duplicate DONEs or timeout reclamations.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: leolee <yihao.li@huawei.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation module:core module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants