Skip to content

[Feature] [MRV2] Supports the combination of SFA PCP, DCP, and MTP - #16597

Merged
linfeng-yuan merged 7 commits into
vllm-project:mainfrom
weiguihua2:main
Sep 16, 2026
Merged

linfeng-yuan merged 7 commits into
vllm-project:mainfrom
weiguihua2:main

Conversation

@weiguihua2

@weiguihua2 weiguihua2 commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

This PR enables MRV2 speculative decoding with the combined SFA, PCP, DCP, and MTP configuration by keeping KV-cache metadata and draft execution consistent across parallel ranks.

The main changes include:

  • Reuse vLLM's canonical get_dcp_local_seq_lens implementation for interleave-aware, per-rank DCP sequence-length calculation.
  • Build PCP+DCP compact KV gather metadata from the global request block table and the actual allocated block counts. This prevents stale or padded block-table entries from being treated as valid blocks and ensures all DCP ranks use a consistent compact KV layout.
  • Preserve global PCP query metadata and fix replicated slot mapping for padded batches.
  • Align the SFA indexer metadata with replicated KV addresses, PCP-ordered slot mappings, and the main attention block table.
  • Generate speculative DCP metadata for SFA workloads.
  • Preserve the target model's resolved cache layout for separate draft models and normalize the draft parallel configuration.
  • Broadcast replicated draft hidden states from PCP rank 0 so that draft and target execution remain consistent.

Does this PR introduce any user-facing change?

Yes. MRV2 now supports using SFA with PCP, DCP, and MTP together. No new configuration options are introduced.

How was this patch tested?

Added and updated unit tests covering:

  • DCP local sequence lengths for different DCP sizes, interleave sizes, input dtypes, ranks, and empty batches.

  • PCP+DCP compact KV metadata with global request block tables and allocated block counts.

  • Stale and padded block-table entries, including valid block ID 0.

  • Logical DCP collective rank ordering.

  • PCP attention-context block-count snapshots.

  • Replicated SFA indexer metadata and slot-mapping behavior.

  • Existing SFA PCP/DCP metadata-builder paths.

  • vLLM main: vllm-project/vllm@84030bb

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces support for the concurrent use of SFA (Sequence-Parallel Attention), PCP (Prefill Context Parallelism), DCP (Decode Context Parallelism), and MTP (Multi-Token Prediction) within the Ascend backend. By centralizing sequence length calculations and enhancing metadata builders to account for global block table state, the changes ensure consistent KV cache management and synchronization across parallel ranks. These improvements facilitate more robust handling of complex parallel execution modes in speculative decoding scenarios.

Highlights

  • Utility Centralization: Refactored get_dcp_local_seq_lens to vllm.v1.attention.backends.utils to provide a unified implementation across the codebase.
  • Metadata Builder Enhancement: Enhanced AscendSFAPCPDCPMetadataBuilder to support global block table counts, ensuring consistent block ID packing across different PCP ranks.
  • PCP Manager Updates: Updated AscendPCPManager to track and expose global block table counts, enabling more accurate attention context construction.
  • Speculative Decoding Support: Added logic in the Speculator to broadcast hidden states across PCP ranks when replicated PCP is enabled, ensuring consistency during draft generation.
  • Configuration Validation: Updated platform configuration validation to correctly handle speculative decoding scenarios where draft and target models share cache layouts.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Feature] Support PCP and DCP integration for SFA and speculative decoding

Suggested PR Summary:

### What this PR does / why we need it?
This PR refactors context parallel (CP) and speculative decoding integration for Ascend. It imports `get_dcp_local_seq_lens` from vLLM core, updates metadata builders to support PCP+DCP compact KV views, and adds broadcast logic for replicated draft tokens in speculative decoding. It also bypasses parallel config validation for separate draft models.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Tested with new and updated unit tests in `tests/ut/attention/` and `tests/ut/worker/`.

Feedback on Review Comments:
All review comments are valid and point out important robustness improvements, such as using getattr to avoid AttributeErrors, avoiding direct mutation of potentially frozen config objects, and ensuring tensors are contiguous before broadcasting.

Comment thread vllm_ascend/platform.py
Comment thread vllm_ascend/worker/v2/spec_decode/pcp_utils.py
Comment thread vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py Outdated
Comment thread vllm_ascend/attention/context_parallel/sfa_cp.py
Comment thread vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py Outdated

@drslark drslark left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM.

Comment thread vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py

@pisceskkk pisceskkk left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm, leave some nits.

Comment thread vllm_ascend/attention/context_parallel/sfa_cp.py Outdated
Comment thread vllm_ascend/attention/context_parallel/sfa_cp.py
@weiguihua2 weiguihua2 added the ready-all run all e2e test for pr label Sep 15, 2026
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

…ency

Signed-off-by: weiguihua2 <984323595@qq.com>
Signed-off-by: weiguihua2 <984323595@qq.com>
Signed-off-by: weiguihua2 <984323595@qq.com>
Signed-off-by: weiguihua2 <984323595@qq.com>
Signed-off-by: weiguihua2 <984323595@qq.com>
Signed-off-by: weiguihua2 <984323595@qq.com>
@weiguihua2

weiguihua2 commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@linfeng-yuan
linfeng-yuan merged commit 0301ec6 into vllm-project:main Sep 16, 2026
39 of 41 checks passed
lllyys pushed a commit to lllyys/vllm-ascend that referenced this pull request Sep 18, 2026
…t mapping (vllm-project#16597)

Partial backport of vllm-project#16597 to releases/v0.26.0rc: only the SFA DCP
replicated slot-mapping padding fix and its regression test are included.
_build_slot_mapping_replicated_view sized its repeat_interleave output by
num_input_tokens, which can include graph padding, while the repeated query
lengths cover only real tokens; this raises a RuntimeError whenever
num_input_tokens > sum(query_lens). The fix sizes it by the actual query
token count from query_start_loc_cpu.

The SFA-PCP / MTP-PCP feature work in the original commit requires MRV2 PCP
groundwork (vllm-project#15809, vllm-project#14960) that is not on this branch and is intentionally
not included.

(cherry picked from commit 0301ec6)

Signed-off-by: weiguihua2 <984323595@qq.com>
Signed-off-by: lllyys <221728326+lllyys@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants