Skip to content

[BugFix] Fix speculative decoding graph replay - #16221

Merged
weiguihua2 merged 1 commit into
vllm-project:mainfrom
li1how:fix/speculative-graph-replay
Sep 10, 2026
Merged

weiguihua2 merged 1 commit into
vllm-project:mainfrom
li1how:fix/speculative-graph-replay

Conversation

@li1how

@li1how li1how commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Fix a query-length mismatch during PCP speculative graph replay by rebuilding draft prefill metadata only for DSA and SFA.

Does this PR introduce any user-facing change?

Yes, fixes speculative decoding graph replay failures.

How was this patch tested?

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a bug in the speculative decoding graph replay mechanism. By refining the logic for rebuilding draft prefill metadata to only trigger for DSA-based architectures, the system avoids incorrect metadata states that previously caused query-length mismatches. These changes ensure more reliable speculative decoding performance across different model configurations.

Highlights

  • Speculative Decoding Fix: Resolved a query-length mismatch during PCP speculative graph replay by restricting the draft prefill metadata rebuild process specifically to DSA architectures.
  • Test Suite Expansion: Updated the unit tests to include parameterized test cases covering different attention architectures (DSA vs MLA) and replication states to ensure robust metadata handling.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@li1how
li1how marked this pull request as ready for review September 10, 2026 07:04
Copilot AI lite review requested due to automatic review settings September 10, 2026 07:04
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 10, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
🔒 Security Review ✅ Completed 2026-09-10T07:10:38.640497Z 2a2ca6d Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Feature] Support MLA architecture in draft attention metadata building for speculator

Suggested PR Summary:

### What this PR does / why we need it?
This PR updates `build_draft_attn_metadatas` in `AscendMTPSpeculator` to conditionally prepare replicated prefill attention only when the attention architecture is "DSA". For other architectures like "MLA", it bypasses this preparation and directly returns the original attention metadata. This is necessary to support different attention architectures during draft model prefill.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
The changes were tested by updating the unit test `test_graph_prefill_builds_draft_metadata` in `tests/ut/worker/test_mtp_pcp_speculator_v2.py` to cover different combinations of `attn_architecture`, `replicated_pcp`, and `rebuild_metadata`.

No review comments were provided, so there is no additional feedback to provide.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Warning

Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.

Pull request overview

Fixes PCP speculative decoding graph replay failures caused by query-length mismatches by conditionally rebuilding draft prefill attention metadata only for DSA.

Changes:

  • Gate replicated prefill attention metadata preparation behind attn_architecture == "DSA" during draft-model prefill.
  • Expand the unit test matrix to cover DSA/MLA + replicated/non-replicated PCP combinations.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 2 comments.

File Description
vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py Restricts replicated prefill attention-metadata preparation to DSA to prevent query-length mismatches during graph replay.
tests/ut/worker/test_mtp_pcp_speculator_v2.py Updates UT to parameterize behavior across attention architectures and PCP replication modes.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread tests/ut/worker/test_mtp_pcp_speculator_v2.py
Comment thread vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py Outdated
@li1how
li1how force-pushed the fix/speculative-graph-replay branch from 2a2ca6d to 16120cb Compare September 10, 2026 07:24
Restrict draft prefill metadata rebuilding to DSA and SFA. Reuse existing
metadata for other attention backends to preserve padded query lengths
during speculative graph replay.

Signed-off-by: leolee <yihao.li@huawei.com>
@li1how
li1how force-pushed the fix/speculative-graph-replay branch from 16120cb to 52360b1 Compare September 10, 2026 07:31
@weiguihua2 weiguihua2 added the ready-precise run selected e2e test for pr label Sep 10, 2026
@weiguihua2
weiguihua2 merged commit 00ce2f0 into vllm-project:main Sep 10, 2026
31 of 32 checks passed
@li1how
li1how deleted the fix/speculative-graph-replay branch September 10, 2026 12:45
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

Fix a query-length mismatch during PCP speculative graph replay by
rebuilding draft prefill metadata only for DSA and SFA.

### Does this PR introduce _any_ user-facing change?

Yes, fixes speculative decoding graph replay failures.

### How was this patch tested?


- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: leolee <yihao.li@huawei.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
### What this PR does / why we need it?

Fix a query-length mismatch during PCP speculative graph replay by
rebuilding draft prefill metadata only for DSA and SFA.

### Does this PR introduce _any_ user-facing change?

Yes, fixes speculative decoding graph replay failures.

### How was this patch tested?


- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: leolee <yihao.li@huawei.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
### What this PR does / why we need it?

Fix a query-length mismatch during PCP speculative graph replay by
rebuilding draft prefill metadata only for DSA and SFA.

### Does this PR introduce _any_ user-facing change?

Yes, fixes speculative decoding graph replay failures.

### How was this patch tested?

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: leolee <yihao.li@huawei.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants