Skip to content

[Feature][Model] Fix per-request metadata buffer sizing for DSA-CP under full cudagraph padding - #16168

Merged
yiz-liu merged 2 commits into
vllm-project:mainfrom
frankie-ys:mtp_graph_for_dsa_cp
Sep 18, 2026
Merged

yiz-liu merged 2 commits into
vllm-project:mainfrom
frankie-ys:mtp_graph_for_dsa_cp

Conversation

@frankie-ys

@frankie-ys frankie-ys commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

This PR re-applies the DSA-CP + MTP aclgraph support for DeepSeek V4 that was reverted from main in #16181 (revert of #12599), together with a fix for the crash that motivated scrutiny of the original feature: serving concurrent requests with cudagraph_mode=FULL_DECODE_ONLY failed with aclnnInplaceCopy shape-mismatch errors (EZ1007) or illegal device memory accesses (EZ9999).

Root cause of the crash

In FULL-decode graph mode, the DSA-CP draft path receives num_reqs as the padded request count — the cudagraph capture bucket size plus the FIA dummy request from mixed-batch padding — which can exceed scheduler_config.max_num_seqs. The per-request metadata buffers were sized by max_num_seqs, so:

  1. [:num_reqs] views of these buffers get silently truncated and copy_() into them fails with shape mismatches — e.g. --max-num-seqs 6 with 3 concurrent MTP requests pads 12 tokens to bucket 16, producing Shape [16] vs [6] do not meet the broadcast condition (EZ1007);
  2. triton kernels receiving the whole buffer write past its end, corrupting adjacent device memory (EZ9999: MTE accesses an invalid GM address).

Fix

Compute one graph-mode-aware capacity up front and size all per-request buffers with it:

max_padded_reqs = max(max_num_seqs, max_cudagraph_capture_size) + 1  # +1: FIA dummy request

This covers the step-0 buffers (start_pos_prefill, local_query_start_loc, local_seq_lens), the draft buffers (spec_local_query_start_loc, spec_local_seq_lens, spec_start_pos), and the QLI buffers (qli_seqused_k, qli_cmp_residual_k) with a single capacity source. All usage sites are [:num_reqs] slices or whole-buffer kernel inputs, so the change is a pure capacity gain with negligible (int32-level) memory overhead.

The re-applied feature is fully adapted to current main: it fuses with the sequence-parallel prefill rework (#15549), the dynamic DSA indexer quant_mode (#16224) and the DSA PCP + DSpark support (#15958) already on main.

The feature enables aclgraph capture/replay for DSA-CP with MTP=1 and MTP=3, including stable per-draft-index metadata buffers (spec_sas_metadata, spec_start_pos, spec_local_*, per-draft RoPE cache) so tensor addresses stay fixed across graph capture and replay.

Does this PR introduce any user-facing change?

Yes, it re-enables DSA-CP + MTP with FULL_DECODE_ONLY graphs for DeepSeek V4 (enable_dsa_cp: true + speculative_config + cudagraph_mode: FULL_DECODE_ONLY), which was available before the revert. Users with max_num_seqs smaller than the max cudagraph capture size no longer hit EZ1007 / EZ9999 errors under concurrent load.

The gain show below:
when use dsv4+mtp 3+eager:

c98d2620e0c57e8a2de08fd8281e6903
when use dsv4+ mtp 3 + graph:
74786d0aebc590d420bac3bfbf436f32

GMS8K accuracy evaluation for MTP+eager and MTP+graph is attached below.

dsv4-flash+mtp3+eager
Overall report table: 
+---------+-----------+----------+----------+-------+---------+---------+
| Model   | Dataset   | Metric   | Subset   |   Num |   Score | Cat.0   |
+=========+===========+==========+==========+=======+=========+=========+
| dsv4   | gsm8k     | mean_acc | main     |  1318 |  0.9734 | default |
+---------+-----------+----------+----------+-------+---------+---------+ 
dsv4-flash+mtp3+graph
Overall report table: 
+---------+-----------+----------+----------+-------+---------+---------+
| Model   | Dataset   | Metric   | Subset   |   Num |   Score | Cat.0   |
+=========+===========+==========+==========+=======+=========+=========+
| dsv4    | gsm8k     | mean_acc | main     |  1319 |  0.9742 | default |
+---------+-----------+----------+----------+-------+---------+---------+ 
glm5.2-w4a8c8 + mtp 5 +graph
Overall report table: 
+---------+-----------+----------+----------+-------+---------+---------+
| Model   | Dataset   | Metric   | Subset   |   Num |   Score | Cat.0   |
+=========+===========+==========+==========+=======+=========+=========+
| glm5.2    | gsm8k     | mean_acc | main     |  1319 |   0.978 | default |
+---------+-----------+----------+----------+-------+---------+---------+ 

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a critical bug in the DSA-CP metadata builder where per-request buffers were sized based on the standard maximum sequence count rather than the padded count required during full cudagraph execution. By unifying the sizing logic to account for cudagraph capture sizes and FIA dummy requests, the changes ensure that memory allocations are robust against the larger request counts encountered in specific deployment configurations.

Highlights

  • Buffer Sizing Correction: Updated the allocation logic for per-request metadata buffers to use a unified max_padded_reqs value, ensuring sufficient capacity when running with cudagraph padding.
  • Stability Improvements: Prevented runtime shape mismatch errors during tensor copies and potential memory corruption in Triton kernels by aligning buffer sizes with the actual graph-mode request count.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][BugFix] Size all per-request buffers for graph-mode maximum to prevent memory corruption

Suggested PR Summary:

### What this PR does / why we need it?
This PR sizes all per-request buffers (such as `start_pos_prefill`, `qli_seqused_k`, `qli_cmp_residual_k`, `local_query_start_loc`, and `local_seq_lens`) for the graph-mode maximum (`max_padded_reqs`) instead of `scheduler_config.max_num_seqs`. This prevents issues where views get truncated, `copy_()` operations fail due to shape mismatches, or Triton kernels write past the buffer end and corrupt adjacent device memory.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
No tests were explicitly added, but existing tests should pass with these updated buffer sizes.

I have no further feedback to provide as there are no review comments.

@frankie-ys frankie-ys changed the title [BugFix][Model] Fix per-request metadata buffer sizing for DSA-CP under full cudagraph padding [BugFix][Feature] Fix per-request metadata buffer sizing for DSA-CP under full cudagraph padding Sep 10, 2026
@frankie-ys frankie-ys changed the title [BugFix][Feature] Fix per-request metadata buffer sizing for DSA-CP under full cudagraph padding [BugFix] Fix per-request metadata buffer sizing for DSA-CP under full cudagraph padding Sep 10, 2026
@weiguihua2 weiguihua2 added the ready-precise run selected e2e test for pr label Sep 10, 2026
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

… sizing

main reverted vllm-project#12599 (5d4294f); this re-applies the DSA-CP aclgraph feature on top of current main, keeping the graph-mode-aware per-request buffer sizing (max_padded_reqs) that fixes EZ1007 copy_ shape mismatches and out-of-bounds triton writes when num_reqs is padded beyond max_num_seqs during FULL_DECODE_ONLY aclgraph replay.

Signed-off-by: frankie <wangyongsheng686@gmail.com>
…cp graph

Re-applies the MTP speculative config removed by the vllm-project#12599 revert and adds a concurrent-request accuracy guard: max_num_seqs smaller than the cudagraph capture bucket with 4 concurrent MTP requests exercises the padded draft path.

Signed-off-by: frankie <wangyongsheng686@gmail.com>
@frankie-ys frankie-ys changed the title [BugFix] Fix per-request metadata buffer sizing for DSA-CP under full cudagraph padding [Feature][Model] Fix per-request metadata buffer sizing for DSA-CP under full cudagraph padding Sep 16, 2026
@kunpengW-code

Copy link
Copy Markdown
Collaborator

The PR #16181 was reverted because it caused a nightly test failure. I will trigger the failed test case from that run in the comments, please keep an eye on the execution results.

@kunpengW-code

kunpengW-code commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

/nightly glm-5.2-w4a8c8-sfa-dcp
nightly command triggered.

@frankie-ys

Copy link
Copy Markdown
Contributor Author

/nightly glm-5.2-w4a8c8-sfa-dcp nightly command triggered.

Thanks,I see this test has passed.

@weijinqian0
weijinqian0 self-requested a review September 18, 2026 08:20

@drslark drslark left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM.

Comment thread vllm_ascend/attention/context_parallel/dsa_cp.py
@yiz-liu
yiz-liu merged commit fe167d9 into vllm-project:main Sep 18, 2026
18 checks passed
czydyy added a commit to czydyy/vllm-ascend that referenced this pull request Sep 18, 2026
vllm-project#16544)

Reverts commit 200309d on main_verify.

Clean revert: 41 of 43 surviving files are byte-identical to the
pre-PR state. dsa_cp.py and test_model_runner_v2.py keep later
upstream changes (vllm-project#16168 per-request metadata buffer sizing fix,
vllm-project#15747 spec-pp protocol rename) as intended.

Committed with --no-verify: pre-commit gitleaks/check-logger hooks
fail on this Windows host (/bin/bash path + GBK locale); equivalent
content checks (ruff, codespell, typos, check-logger,
check-long-functions) all pass when run manually.

Signed-off-by: chenzeyu <2978509328@qq.com>
czydyy added a commit to czydyy/vllm-ascend that referenced this pull request Sep 18, 2026
vllm-project#16544)

Reverts commit 200309d on main
(cherry-picked from main_verify 82145c7).

Conflict resolution: ascend_forward_context.py, patch/__init__.py,
worker.py and test_ascend_forward_context.py were entangled with vllm-project#16626
(merged before vllm-project#16544, already reverted on this branch); they are
restored to the pre-vllm-project#16626 state (c7ca0b6~1), the correct composition
of both reverts. dsa_cp.py and test_model_runner_v2.py keep later
upstream changes (vllm-project#16168 metadata buffer sizing fix, vllm-project#15747 spec-pp
protocol rename). All other surviving files match the pre-PR state.

Signed-off-by: chenzeyu <2978509328@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants