Skip to content

[Cherry-pick][releases/v0.25.1rc][Feature]dsv4 dsa_cp support dspark (from #13249) - #13316

Closed
vllm-ascend-ci wants to merge 1 commit into
vllm-project:releases/v0.25.1rcfrom
vllm-ascend-ci:cherry-pick/pr-13249-to-releases-v0.25.1rc
Closed

vllm-ascend-ci wants to merge 1 commit into
vllm-project:releases/v0.25.1rcfrom
vllm-ascend-ci:cherry-pick/pr-13249-to-releases-v0.25.1rc

Conversation

@vllm-ascend-ci

@vllm-ascend-ci vllm-ascend-ci commented Aug 1, 2026 •

Copy link
Copy Markdown
Collaborator

Cherry-pick of PR #13249 onto releases/v0.25.1rc.

Original PR: #13249
Original author: @weiguihua2


What this PR does / why we need it?

dsv4 dsa_cp supports the integration of the dspark feature.

  1. Identify DSpark query blocks with causal=False, allowing each CP rank to use the full KV length for valid queries, no longer truncating according to causal rules.
  2. Construct "historical SWA window + complete draft block" ori_sparse_indices for each draft token.
  3. Expand the SWA window to window_size + draft_block_size and pass the sparse indices into the attention operator to achieve non-causal attention within the draft block.

Does this PR introduce any user-facing change?

No

How was this patch tested?

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request integrates the dspark feature into the dsv4 dsa_cp implementation. The changes enable non-causal attention for draft tokens by expanding the sliding window and passing custom sparse indices to the attention operator, ensuring proper handling of KV lengths during speculative decoding.

Highlights

  • DSpark Integration: Enabled dsv4 dsa_cp to support the dspark feature by identifying non-causal query blocks and adjusting attention windowing.
  • Attention Mechanism Updates: Implemented logic to construct historical SWA windows combined with draft blocks, allowing for non-causal attention within draft segments.
  • Testing: Added a new E2E test case for DeepSeek V4 with DSpark and DSA CP enabled to verify acceptance metrics.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Feature] Support DSpark with Context Parallel (DSA CP) for DeepSeek V4

Suggested PR Summary:

### What this PR does / why we need it?
This pull request adds support for DSpark speculative decoding with Context Parallel (DSA CP) for DeepSeek V4. It introduces the necessary metadata fields (`ori_win_left`, `ori_win_right`, and `dspark_swa_indices`) to `AscendDSAReqMetadata` and implements the logic to build and pad `dspark_swa_indices` when non-causal attention is used. It also updates the attention forward pass to utilize these indices and window configurations.

Feedback:
One review comment suggests replacing `F.pad` with `torch.full` and slice assignment to avoid potential compatibility or performance issues with integer tensor padding on Ascend NPUs.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
A new end-to-end test `test_deepseek_v4_dsa_cp_dspark_acceptance_tp4` has been added to verify the acceptance rate of DeepSeek V4 with DSA CP and DSpark speculative decoding.

Comment on lines +509 to +510
if pad_rows:
global_dspark_indices = F.pad(global_dspark_indices, (0, 0, 0, 0, 0, pad_rows), value=-1)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Using F.pad with constant values on integer tensors can sometimes hit compatibility or performance issues on specialized hardware like the Ascend NPU. A more robust and standard approach is to pre-allocate the padded tensor using torch.full and copy the original indices into it.

Suggested change
if pad_rows:
global_dspark_indices = F.pad(global_dspark_indices, (0, 0, 0, 0, 0, pad_rows), value=-1)
if pad_rows:
padded_indices = torch.full(
(num_tokens_pad, *global_dspark_indices.shape[1:]),
-1,
dtype=global_dspark_indices.dtype,
device=global_dspark_indices.device,
)
padded_indices[:global_dspark_indices.shape[0]] = global_dspark_indices
global_dspark_indices = padded_indices

@weiguihua2 weiguihua2 closed this Aug 1, 2026
@weiguihua2 weiguihua2 reopened this Aug 1, 2026
@vllm-ascend-ci
vllm-ascend-ci force-pushed the cherry-pick/pr-13249-to-releases-v0.25.1rc branch from 10960d9 to f8b6658 Compare August 1, 2026 03:21
@weiguihua2

weiguihua2 commented Aug 1, 2026 •

Copy link
Copy Markdown
Collaborator

/rerun
[Bot]: rerun completed.

Rerun:

  • E2E

@weiguihua2 weiguihua2 closed this Aug 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants