Skip to content

[Feature]dsv4 dsa_cp support dspark - #13249

Merged
zzzzwwjj merged 11 commits into
vllm-project:mainfrom
weiguihua2:main_2
Aug 1, 2026
Merged

zzzzwwjj merged 11 commits into
vllm-project:mainfrom
weiguihua2:main_2

Conversation

@weiguihua2

@weiguihua2 weiguihua2 commented Jul 31, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

dsv4 dsa_cp supports the integration of the dspark feature.

  1. Identify DSpark query blocks with causal=False, allowing each CP rank to use the full KV length for valid queries, no longer truncating according to causal rules.
  2. Construct "historical SWA window + complete draft block" ori_sparse_indices for each draft token.
  3. Expand the SWA window to window_size + draft_block_size and pass the sparse indices into the attention operator to achieve non-causal attention within the draft block.

Does this PR introduce any user-facing change?

No

How was this patch tested?

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request integrates DSpark speculative decoding support into the DSA context-parallel framework. By modifying the attention metadata and windowing logic, the changes allow for non-causal attention within draft blocks, improving performance and accuracy for speculative decoding on Ascend hardware. The update ensures that draft tokens can access the full KV length without causal truncation, maintaining consistency across parallel ranks.

Highlights

  • DSpark Integration: Enabled non-causal attention support for DSpark queries within the DSA context-parallel implementation, allowing full KV length usage for draft tokens.
  • Metadata Updates: Extended metadata structures to include sliding window indices and custom window configurations, ensuring proper attention window management during speculative decoding.
  • E2E Testing: Added a new E2E test case to verify DeepSeek-V4 DSpark acceptance rates under context-parallel settings.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Feature] Support DSpark speculative decoding with Context Parallel (DSA CP)

Suggested PR Summary:

### What this PR does / why we need it?
This PR adds support for DSpark speculative decoding with Context Parallel (DSA CP) in the Ascend backend. It introduces non-causal CP drafting metadata building, including computing DSpark sliding window attention (SWA) indices and sparse window parameters, and passing them to the attention forward pass.

Feedback:
In `vllm_ascend/attention/context_parallel/dsa_cp.py`, when `is_noncausal` is `True`, `local_seq_lens` is completely overwritten. Cloning it unconditionally at line 471 is redundant and introduces unnecessary memory allocation overhead. It should only be cloned when `is_noncausal` is `False`.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Added an end-to-end test `test_deepseek_v4_dsa_cp_dspark_acceptance_tp4` in `tests/e2e/pull_request/four_card/spec_decode/test_dspark_deepseekv4.py` to verify the acceptance rate of DeepSeek V4 with DSA CP and DSpark speculative decoding.

@@ -463,21 +469,62 @@ def build_req_metadata_for_drafting(
)
local_query_start_loc = local_query_start_loc.clone()
local_seq_lens = local_seq_lens.clone()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

When is_noncausal is True, local_seq_lens is completely overwritten by torch.where at line 486. Therefore, cloning local_seq_lens at line 471 is redundant and leads to unnecessary memory allocation and copy overhead. We should only clone it when is_noncausal is False.

Suggested change
local_seq_lens = local_seq_lens.clone()
if not is_noncausal:
local_seq_lens = local_seq_lens.clone()

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
cleanup_dist_env_and_memory()


@pytest.mark.parametrize("model_name", MODELS)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it can be merged into test_deepseek_v4_dspark_acceptance_tp4.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should guard both dsa_v1.py and dsa_cp.py. But yes, we can use same function with two different param groups.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Because these are two different and separate processes with different logic, both require monitoring.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've thought about it, and I might be wrong.
If these two test cases are merged, it will become unclear how to skip just one of them when a problem arises.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll think about it over again.

input_positions = common_attn_metadata.positions[:num_input_tokens].long()
self.common_ratio_to_sas_metadata["input_positions"] = input_positions
cos, sin = get_cos_and_sin_dsa(input_positions, use_cache=not has_prefill)
cos, sin = get_cos_and_sin_dsa(input_positions, use_cache=self.num_prefills == 0)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can set has_prefill = self.num_prefills > 0 before this line, and use not has_prefill here.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done

Comment on lines +485 to +486
local_query_lens = local_query_start_loc[1:] - local_query_start_loc[:-1]
local_seq_lens = torch.where(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should we make local_query_lens and local_seq_lens persistent?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done

seq_lens=self.seq_lens_cpu[:num_reqs],
)

if is_noncausal:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe we should move this branch into _build_local_token_metadata

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done

cleanup_dist_env_and_memory()


@pytest.mark.parametrize("model_name", MODELS)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should guard both dsa_v1.py and dsa_cp.py. But yes, we can use same function with two different param groups.

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
Signed-off-by: weiguihua2 <weiguihua2@huawei.com>

@yiz-liu yiz-liu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  1. Use parametrize instead of creating a new case

@yiz-liu yiz-liu added the ready label Jul 31, 2026
Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
@weiguihua2

weiguihua2 commented Aug 1, 2026 •

Copy link
Copy Markdown
Collaborator Author
  1. Use parametrize instead of creating a new case

done

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
@weiguihua2

weiguihua2 commented Aug 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

/cherry-pick releases/v0.25.1rc
[Bot]: cherry-pick branch has been refreshed. Existing PR: #13316

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>

@drslark drslark left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM.
Thanks for your efforts.

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
@weiguihua2

weiguihua2 commented Aug 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

/cherry-pick releases/v0.26.0rc
[Bot]: cherry-pick completed successfully. New PR created: #13322

@zzzzwwjj
zzzzwwjj merged commit fa3b1bd into vllm-project:main Aug 1, 2026
46 checks passed
yiz-liu pushed a commit that referenced this pull request Aug 1, 2026
…13320)

### What this PR does / why we need it?
Cherry-pick of PR #13249
onto releases/v0.25.1rc.

Original PR: #13249

dsv4 dsa_cp supports the integration of the dspark feature.

Identify DSpark query blocks with causal=False, allowing each CP rank to
use the full KV length for valid queries, no longer truncating according
to causal rules.
Construct "historical SWA window + complete draft block"
ori_sparse_indices for each draft token.
Expand the SWA window to window_size + draft_block_size and pass the
sparse indices into the attention operator to achieve non-causal
attention within the draft block.

### Does this PR introduce _any_ user-facing change?
NO

### How was this patch tested?

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@fe784ff

---------

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
yiz-liu pushed a commit that referenced this pull request Aug 3, 2026
…(from #13249) (#13322)

Cherry-pick of PR #13249 onto `releases/v0.26.0rc`.

Original PR: #13249
Original author: @weiguihua2

---
### What this PR does / why we need it?
dsv4 dsa_cp supports the integration of the dspark feature.  
1. Identify DSpark query blocks with causal=False, allowing each CP rank
to use the full KV length for valid queries, no longer truncating
according to causal rules.
2. Construct "historical SWA window + complete draft block"
ori_sparse_indices for each draft token.
3. Expand the SWA window to window_size + draft_block_size and pass the
sparse indices into the attention operator to achieve non-causal
attention within the draft block.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
xiayingqing pushed a commit to xiayingqing/vllm-ascend that referenced this pull request Aug 3, 2026
### What this PR does / why we need it?
dsv4 dsa_cp supports the integration of the dspark feature.  
1. Identify DSpark query blocks with causal=False, allowing each CP rank
to use the full KV length for valid queries, no longer truncating
according to causal rules.
2. Construct "historical SWA window + complete draft block"
ori_sparse_indices for each draft token.
3. Expand the SWA window to window_size + draft_block_size and pass the
sparse indices into the attention operator to achieve non-causal
attention within the draft block.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

---------

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
HMCCMH pushed a commit to hotTea123/vllm-ascend that referenced this pull request Aug 12, 2026
### What this PR does / why we need it?
dsv4 dsa_cp supports the integration of the dspark feature.  
1. Identify DSpark query blocks with causal=False, allowing each CP rank
to use the full KV length for valid queries, no longer truncating
according to causal rules.
2. Construct "historical SWA window + complete draft block"
ori_sparse_indices for each draft token.
3. Expand the SWA window to window_size + draft_block_size and pass the
sparse indices into the attention operator to achieve non-causal
attention within the draft block.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

---------

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
MmMmaru pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 19, 2026
### What this PR does / why we need it?
dsv4 dsa_cp supports the integration of the dspark feature.  
1. Identify DSpark query blocks with causal=False, allowing each CP rank
to use the full KV length for valid queries, no longer truncating
according to causal rules.
2. Construct "historical SWA window + complete draft block"
ori_sparse_indices for each draft token.
3. Expand the SWA window to window_size + draft_block_size and pass the
sparse indices into the attention operator to achieve non-causal
attention within the draft block.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

---------

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
shiqiangA pushed a commit to shiqiangA/vllm-ascend that referenced this pull request Aug 20, 2026
### What this PR does / why we need it?
dsv4 dsa_cp supports the integration of the dspark feature.  
1. Identify DSpark query blocks with causal=False, allowing each CP rank
to use the full KV length for valid queries, no longer truncating
according to causal rules.
2. Construct "historical SWA window + complete draft block"
ori_sparse_indices for each draft token.
3. Expand the SWA window to window_size + draft_block_size and pass the
sparse indices into the attention operator to achieve non-causal
attention within the draft block.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

---------

Signed-off-by: weiguihua2 <weiguihua2@huawei.com>
lijiahang226 pushed a commit to lijiahang226/vllm-ascend that referenced this pull request Aug 21, 2026
…(from vllm-project#13249) (vllm-project#13322)

Cherry-pick of PR vllm-project#13249 onto `releases/v0.26.0rc`.

Original PR: vllm-project#13249
Original author: @weiguihua2

---
### What this PR does / why we need it?
dsv4 dsa_cp supports the integration of the dspark feature.  
1. Identify DSpark query blocks with causal=False, allowing each CP rank
to use the full KV length for valid queries, no longer truncating
according to causal rules.
2. Construct "historical SWA window + complete draft block"
ori_sparse_indices for each draft token.
3. Expand the SWA window to window_size + draft_block_size and pass the
sparse indices into the attention operator to achieve non-causal
attention within the draft block.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
lrf-vm pushed a commit to lijiahang226/vllm-ascend that referenced this pull request Aug 21, 2026
…(from vllm-project#13249) (vllm-project#13322)

Cherry-pick of PR vllm-project#13249 onto `releases/v0.26.0rc`.

Original PR: vllm-project#13249
Original author: @weiguihua2

---
### What this PR does / why we need it?
dsv4 dsa_cp supports the integration of the dspark feature.  
1. Identify DSpark query blocks with causal=False, allowing each CP rank
to use the full KV length for valid queries, no longer truncating
according to causal rules.
2. Construct "historical SWA window + complete draft block"
ori_sparse_indices for each draft token.
3. Expand the SWA window to window_size + draft_block_size and pass the
sparse indices into the attention operator to achieve non-causal
attention within the draft block.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
weijinqian0 pushed a commit that referenced this pull request Sep 8, 2026
### What this PR does / why we need it?

This PR enables MTP=3 speculative decoding in the DSA‑CP path for both
eager and graph modes.
It introduces full ACL graph support for the MTP draft steps, including
pre‑allocated metadata buffers that allow deterministic graph capture
and replay when enable_dsa_cp is true.

### Changes summary

#### MTP > 1 support in DSA‑CP

~Previously build_for_drafting only handled MTP=1 and lacked the
CPU‑side sequence length path, which caused tensor dimension errors for
any MTP > 1.~ Supported by #13249.

This PR updates build_for_drafting to populate CPU sequence lengths, so
that all downstream metadata builders can correctly process multiple
draft steps.

#### Full ACL graph support for MTP draft steps

ACL graph capture and replay require stable tensor addresses across
invocations. To guarantee this, the PR pre‑allocates per‑step buffers
for all draft‑step metadata during builder initialization, and pads
seq_lens_cpu to match the graph‑dispatched batch size, ensuring
deterministic addresses throughout replay.

In addition, a static `update_graph_params` method was added, so that
the graph dispatch can
correctly invoke the DSA-CP attention backend during graph replay.

~#### Depends on PR  #12193.~

~#12193 also fixes metadata mismatch issues in the DSA-CP builder, and
this PR is directly based on the refactored code. Cherry-picking the
commits onto main without #12193 causes runtime errors under concurrent
requests.~

~Only the top commits (after Commits on Jul 21, 2026) belong to this PR.
Please review by focusing on the top commits. Once it is merged, I'll
rebase onto main quickly.`~

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- Added an E2E accuracy test in
  `tests/e2e/pull_request/four_card/test_deepseek_v4.py` with exact
  `expected_token_ids` assertion to guard against regressions in the
  DSA-CP + MTP=3 + full graph path.
- GMS8K accuracy evaluation for MTP+eager and MTP+graph is attached
below.

```
mtp3+eager
+---------+-----------+----------+----------+-------+---------+---------+
| Model   | Dataset   | Metric   | Subset   |   Num |   Score | Cat.0   |
+=========+===========+==========+==========+=======+=========+=========+
| dsv4-f    | gsm8k     | mean_acc | main     |  1319 |  0.9666 | default |
+---------+-----------+----------+----------+-------+---------+---------+ 

mtp3+full graph
+---------+-----------+----------+----------+-------+---------+---------+
| Model   | Dataset   | Metric   | Subset   |   Num |   Score | Cat.0   |
+=========+===========+==========+==========+=======+=========+=========+
| dsv4-f    | gsm8k     | mean_acc | main     |  1319 |  0.9659 | default |
+---------+-----------+----------+----------+-------+---------+---------+ 
```


- vLLM main:
vllm-project/vllm@e6bfe03

---------

Signed-off-by: frankie <wangyongsheng686@gmail.com>
Co-authored-by: frankie <wangyongsheng686@gmail.com>
jiangli221 pushed a commit to jiangli221/vllm-ascend that referenced this pull request Sep 9, 2026
### What this PR does / why we need it?

This PR enables MTP=3 speculative decoding in the DSA‑CP path for both
eager and graph modes.
It introduces full ACL graph support for the MTP draft steps, including
pre‑allocated metadata buffers that allow deterministic graph capture
and replay when enable_dsa_cp is true.

### Changes summary

#### MTP > 1 support in DSA‑CP

~Previously build_for_drafting only handled MTP=1 and lacked the
CPU‑side sequence length path, which caused tensor dimension errors for
any MTP > 1.~ Supported by vllm-project#13249.

This PR updates build_for_drafting to populate CPU sequence lengths, so
that all downstream metadata builders can correctly process multiple
draft steps.

#### Full ACL graph support for MTP draft steps

ACL graph capture and replay require stable tensor addresses across
invocations. To guarantee this, the PR pre‑allocates per‑step buffers
for all draft‑step metadata during builder initialization, and pads
seq_lens_cpu to match the graph‑dispatched batch size, ensuring
deterministic addresses throughout replay.

In addition, a static `update_graph_params` method was added, so that
the graph dispatch can
correctly invoke the DSA-CP attention backend during graph replay.

~#### Depends on PR  vllm-project#12193.~

~vllm-project#12193 also fixes metadata mismatch issues in the DSA-CP builder, and
this PR is directly based on the refactored code. Cherry-picking the
commits onto main without vllm-project#12193 causes runtime errors under concurrent
requests.~

~Only the top commits (after Commits on Jul 21, 2026) belong to this PR.
Please review by focusing on the top commits. Once it is merged, I'll
rebase onto main quickly.`~

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- Added an E2E accuracy test in
  `tests/e2e/pull_request/four_card/test_deepseek_v4.py` with exact
  `expected_token_ids` assertion to guard against regressions in the
  DSA-CP + MTP=3 + full graph path.
- GMS8K accuracy evaluation for MTP+eager and MTP+graph is attached
below.

```
mtp3+eager
+---------+-----------+----------+----------+-------+---------+---------+
| Model   | Dataset   | Metric   | Subset   |   Num |   Score | Cat.0   |
+=========+===========+==========+==========+=======+=========+=========+
| dsv4-f    | gsm8k     | mean_acc | main     |  1319 |  0.9666 | default |
+---------+-----------+----------+----------+-------+---------+---------+ 

mtp3+full graph
+---------+-----------+----------+----------+-------+---------+---------+
| Model   | Dataset   | Metric   | Subset   |   Num |   Score | Cat.0   |
+=========+===========+==========+==========+=======+=========+=========+
| dsv4-f    | gsm8k     | mean_acc | main     |  1319 |  0.9659 | default |
+---------+-----------+----------+----------+-------+---------+---------+ 
```


- vLLM main:
vllm-project/vllm@e6bfe03

---------

Signed-off-by: frankie <wangyongsheng686@gmail.com>
Co-authored-by: frankie <wangyongsheng686@gmail.com>
Leetrytry pushed a commit to Leetrytry/vllm-ascend that referenced this pull request Sep 11, 2026
…(from vllm-project#13249) (vllm-project#13322)

Cherry-pick of PR vllm-project#13249 onto `releases/v0.26.0rc`.

Original PR: vllm-project#13249
Original author: @weiguihua2

---
### What this PR does / why we need it?
dsv4 dsa_cp supports the integration of the dspark feature.  
1. Identify DSpark query blocks with causal=False, allowing each CP rank
to use the full KV length for valid queries, no longer truncating
according to causal rules.
2. Construct "historical SWA window + complete draft block"
ori_sparse_indices for each draft token.
3. Expand the SWA window to window_size + draft_block_size and pass the
sparse indices into the attention operator to achieve non-causal
attention within the draft block.

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

This PR enables MTP=3 speculative decoding in the DSA‑CP path for both
eager and graph modes.
It introduces full ACL graph support for the MTP draft steps, including
pre‑allocated metadata buffers that allow deterministic graph capture
and replay when enable_dsa_cp is true.

### Changes summary

#### MTP > 1 support in DSA‑CP

~Previously build_for_drafting only handled MTP=1 and lacked the
CPU‑side sequence length path, which caused tensor dimension errors for
any MTP > 1.~ Supported by vllm-project#13249.

This PR updates build_for_drafting to populate CPU sequence lengths, so
that all downstream metadata builders can correctly process multiple
draft steps.

#### Full ACL graph support for MTP draft steps

ACL graph capture and replay require stable tensor addresses across
invocations. To guarantee this, the PR pre‑allocates per‑step buffers
for all draft‑step metadata during builder initialization, and pads
seq_lens_cpu to match the graph‑dispatched batch size, ensuring
deterministic addresses throughout replay.

In addition, a static `update_graph_params` method was added, so that
the graph dispatch can
correctly invoke the DSA-CP attention backend during graph replay.

~#### Depends on PR  vllm-project#12193.~

~vllm-project#12193 also fixes metadata mismatch issues in the DSA-CP builder, and
this PR is directly based on the refactored code. Cherry-picking the
commits onto main without vllm-project#12193 causes runtime errors under concurrent
requests.~

~Only the top commits (after Commits on Jul 21, 2026) belong to this PR.
Please review by focusing on the top commits. Once it is merged, I'll
rebase onto main quickly.`~

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- Added an E2E accuracy test in
  `tests/e2e/pull_request/four_card/test_deepseek_v4.py` with exact
  `expected_token_ids` assertion to guard against regressions in the
  DSA-CP + MTP=3 + full graph path.
- GMS8K accuracy evaluation for MTP+eager and MTP+graph is attached
below.

```
mtp3+eager
+---------+-----------+----------+----------+-------+---------+---------+
| Model   | Dataset   | Metric   | Subset   |   Num |   Score | Cat.0   |
+=========+===========+==========+==========+=======+=========+=========+
| dsv4-f    | gsm8k     | mean_acc | main     |  1319 |  0.9666 | default |
+---------+-----------+----------+----------+-------+---------+---------+ 

mtp3+full graph
+---------+-----------+----------+----------+-------+---------+---------+
| Model   | Dataset   | Metric   | Subset   |   Num |   Score | Cat.0   |
+=========+===========+==========+==========+=======+=========+=========+
| dsv4-f    | gsm8k     | mean_acc | main     |  1319 |  0.9659 | default |
+---------+-----------+----------+----------+-------+---------+---------+ 
```


- vLLM main:
vllm-project/vllm@e6bfe03

---------

Signed-off-by: frankie <wangyongsheng686@gmail.com>
Co-authored-by: frankie <wangyongsheng686@gmail.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
### What this PR does / why we need it?

This PR enables MTP=3 speculative decoding in the DSA‑CP path for both
eager and graph modes.
It introduces full ACL graph support for the MTP draft steps, including
pre‑allocated metadata buffers that allow deterministic graph capture
and replay when enable_dsa_cp is true.

### Changes summary

#### MTP > 1 support in DSA‑CP

~Previously build_for_drafting only handled MTP=1 and lacked the
CPU‑side sequence length path, which caused tensor dimension errors for
any MTP > 1.~ Supported by vllm-project#13249.

This PR updates build_for_drafting to populate CPU sequence lengths, so
that all downstream metadata builders can correctly process multiple
draft steps.

#### Full ACL graph support for MTP draft steps

ACL graph capture and replay require stable tensor addresses across
invocations. To guarantee this, the PR pre‑allocates per‑step buffers
for all draft‑step metadata during builder initialization, and pads
seq_lens_cpu to match the graph‑dispatched batch size, ensuring
deterministic addresses throughout replay.

In addition, a static `update_graph_params` method was added, so that
the graph dispatch can
correctly invoke the DSA-CP attention backend during graph replay.

~#### Depends on PR  vllm-project#12193.~

~vllm-project#12193 also fixes metadata mismatch issues in the DSA-CP builder, and
this PR is directly based on the refactored code. Cherry-picking the
commits onto main without vllm-project#12193 causes runtime errors under concurrent
requests.~

~Only the top commits (after Commits on Jul 21, 2026) belong to this PR.
Please review by focusing on the top commits. Once it is merged, I'll
rebase onto main quickly.`~

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- Added an E2E accuracy test in
  `tests/e2e/pull_request/four_card/test_deepseek_v4.py` with exact
  `expected_token_ids` assertion to guard against regressions in the
  DSA-CP + MTP=3 + full graph path.
- GMS8K accuracy evaluation for MTP+eager and MTP+graph is attached
below.

```
mtp3+eager
+---------+-----------+----------+----------+-------+---------+---------+
| Model   | Dataset   | Metric   | Subset   |   Num |   Score | Cat.0   |
+=========+===========+==========+==========+=======+=========+=========+
| dsv4-f    | gsm8k     | mean_acc | main     |  1319 |  0.9666 | default |
+---------+-----------+----------+----------+-------+---------+---------+

mtp3+full graph
+---------+-----------+----------+----------+-------+---------+---------+
| Model   | Dataset   | Metric   | Subset   |   Num |   Score | Cat.0   |
+=========+===========+==========+==========+=======+=========+=========+
| dsv4-f    | gsm8k     | mean_acc | main     |  1319 |  0.9659 | default |
+---------+-----------+----------+----------+-------+---------+---------+
```

- vLLM main:
vllm-project/vllm@e6bfe03

---------

Signed-off-by: frankie <wangyongsheng686@gmail.com>
Co-authored-by: frankie <wangyongsheng686@gmail.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants