Skip to content

[BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings - #16300

Merged
Tflowers-0129 merged 6 commits into
vllm-project:mainfrom
wzx0726:codex/fix-pcp-draft-graph-mappings
Sep 14, 2026
Merged

Tflowers-0129 merged 6 commits into
vllm-project:mainfrom
wzx0726:codex/fix-pcp-draft-graph-mappings

Conversation

@wzx0726

@wzx0726 wzx0726 commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

With replicated PCP, the first draft stage under FULL graph replay must refresh the persistent block-table and slot-mapping buffers captured by the draft model. The MLA/GQA path reuses target attention metadata and skips _prepare_replicated_prefill_attn, leaving those default buffers stale between batches or draft steps.

Extract the existing mapping refresh into _refresh_replicated_prefill_mappings and reuse it from both the preparation path and the MLA/GQA graph path. This preserves MLA/GQA metadata and its padded query layout, while DSA/SFA retain their existing metadata rebuild path with one mapping refresh. No new buffers, flags, or query-padding logic are introduced.

Does this PR introduce any user-facing change?

Fixes stale KV cache mappings for replicated-PCP speculative decoding with MLA/GQA FULL graphs, which can affect draft acceptance. No CLI or configuration changes.

How was this patch tested?

  • Remote isolated CPU unit tests with mocked BlockTables operations: 54 passed across the following files:
    python3 -m pytest tests/ut/worker/test_mtp_pcp_speculator_v2.py tests/ut/worker/test_attn_utils_v2.py -q --tb=short
  • The selected regression tests against the unpatched baseline produced 4 failed, 14 passed, 19 deselected. They detect missing MLA/GQA refreshes and stale slot values.
  • Coverage includes all four attention architectures, request/block changes, padded slots, stable buffer addresses, metadata identity, and non-PCP/dummy/missing-batch bypasses.
  • Ruff lint, Ruff format check on changed files, and git diff --check passed. The full format.sh ci check was not run because Git Bash and pre-commit are unavailable in the local Windows environment.

Validation status: These tests verify the refresh and buffer contracts using CPU tensors and mocks; they do not exercise NPU mapping kernels or graph replay. A service startup failure was reported after deployment and has not yet been diagnosed. NPU startup/capture/replay, end-to-end correctness, and acceptance-rate comparison for this exact patch remain pending.

Per-request KV-length correction is outside the scope of this PR. Existing length construction and update behavior is unchanged.

Signed-off-by: Madilyn537 <zhexuanwu12@gmail.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a bug in replicated PCP speculative decoding where persistent KV cache buffers were not being correctly refreshed for MLA/GQA architectures during FULL graph replay. By refactoring the mapping refresh logic into a shared method, the patch ensures these buffers remain consistent across batches and draft steps without disrupting existing metadata rebuild paths.

Highlights

  • Refactored Mapping Refresh Logic: Extracted the cache mapping refresh logic into a new helper method, _refresh_replicated_prefill_mappings, to centralize buffer updates for replicated PCP.
  • Fixed Stale KV Cache Mappings: Ensured that MLA/GQA paths correctly refresh persistent block-table and slot-mapping buffers during FULL graph replay, preventing stale state between draft steps.
  • Improved Test Coverage: Updated unit tests to cover GQA/MLA architectures and added new regression tests to verify that cache buffers are refreshed appropriately.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Feature] Refresh replicated prefill mappings for draft speculator

Suggested PR Summary:

### What this PR does / why we need it?
This PR refactors the speculator to refresh persistent cache mappings used by replicated drafts during graph prefill. Specifically, it extracts `_refresh_replicated_prefill_mappings` from `_prepare_replicated_prefill_attn` and ensures it is called even when `_prepare_replicated_prefill_attn` is bypassed in `build_draft_attn_metadatas`. It also adds comprehensive unit tests to verify that captured cache buffers are refreshed correctly and that bypass guards preserve metadata.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Tested with new unit tests in `tests/ut/worker/test_mtp_pcp_speculator_v2.py`, including `test_graph_prefill_refreshes_captured_cache_buffers`, `test_prepare_replicated_prefill_preserves_bypass`, and `test_graph_prefill_without_real_batch_preserves_metadata`.

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@wzx0726
wzx0726 marked this pull request as ready for review September 11, 2026 02:04
Signed-off-by: Madilyn537 <zhexuanwu12@gmail.com>
@wzx0726 wzx0726 changed the title [BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings [BugFix][SpecDecode] Fix replicated PCP draft mappings and KV lengths Sep 11, 2026
assert prepared_attn_metadata is not None
attn_metadata = prepared_attn_metadata
else:
self._refresh_replicated_prefill_mappings(num_reqs_padded, num_tokens_padded)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add a TOTO about deleting it when FIA remove its check.

@drslark drslark left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM.

num_tokens_padded=num_tokens_padded,
)

def _prepare_replicated_prefill_attn(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cause prepare_attn in model_runner just modify pcp's variables.

Signed-off-by: Madilyn537 <zhexuanwu12@gmail.com>
@wzx0726 wzx0726 changed the title [BugFix][SpecDecode] Fix replicated PCP draft mappings and KV lengths [BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings Sep 11, 2026
@lijiahang226 lijiahang226 added the ready-precise run selected e2e test for pr label Sep 11, 2026
Signed-off-by: Madilyn537 <zhexuanwu12@gmail.com>
Signed-off-by: Madilyn537 <zhexuanwu12@gmail.com>

@lijiahang226 lijiahang226 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@drslark drslark left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very clear.
LGTM.

@wzx0726

wzx0726 commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun

[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@wzx0726

wzx0726 commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@Tflowers-0129

Copy link
Copy Markdown
Collaborator

LGTM.

@Tflowers-0129
Tflowers-0129 merged commit 6d36d66 into vllm-project:main Sep 14, 2026
46 of 50 checks passed
windshado added a commit to windshado/vllm-ascend that referenced this pull request Sep 14, 2026
…hado/vllm-ascend into main_fix_mrv2_eagle3_mamba

* 'main_fix_mrv2_eagle3_mamba' of https://github.com/windshado/vllm-ascend: (42 commits)
  Update vllm_ascend/worker/v2/model_states/mamba_hybrid.py
  [Feature][Kimi K3 DSPark] Enable TP for context_proj (vllm-project#16344)
  [BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings (vllm-project#16300)
  [Feature][Model] Integrate Triton KeyPool indexing for GLM-5.3-Flash (vllm-project#16253)
  [BugFix][Offloader] Re-bind params to NZ static buffers after npu_format_cast (vllm-project#15415)
  [Feature][Model] Integrate AscendC KDA and causal convolution for GLM-5.3-Flash (vllm-project#16251)
  [Performance][Communicator] Replace per-layer F.pad with cat of a persistent zero block in MoE prepare (vllm-project#16343)
  [Feature][Operator] Add DeepSeek V4.1 sparse attention operators (vllm-project#16422)
  [Doc][Misc] Document batch invariance scheduling limitations (vllm-project#16232)
  [CI][MRV2] Enable mrv2 dspark e2e test (vllm-project#16319)
  [BugFix] Precast MoE gate weight_fp32 to avoid aclop Cast (vllm-project#16189)
  [Feature][MRV2][310P] MRv2 adapting MTP on the 310P for Qwen3.5 (vllm-project#16043)
  [Revert] Revert "[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409)
  [Feature][Ops] Add Triton KeyPool compression and pooled indexing (vllm-project#16243)
  [Feature][Attention] Support NoPE in the shared SFA backend (vllm-project#16252)
  [Performance][Model] Reuse fused mHC operators for GLM-5.3-Flash (vllm-project#16321)
  [Feature][Model] Enable MiniMax-M3 FP8 MSA index score on A5 (vllm-project#15918)
  [Performance][KDA] Reduce preprocessing copies and redundant output masks (vllm-project#16067)
  [Feature][Model][MTP] Support speculative decoding for GLM-5.3-Flash (vllm-project#16214)
  [BugFix][Model] Skip unused hash-router bias when loading DeepSeek-V4 weights (vllm-project#16259)
  ...
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
vllm-project#16300)

### What this PR does / why we need it?

With replicated PCP, the first draft stage under FULL graph replay must
refresh the persistent block-table and slot-mapping buffers captured by
the draft model. The MLA/GQA path reuses target attention metadata and
skips `_prepare_replicated_prefill_attn`, leaving those default buffers
stale between batches or draft steps.

Extract the existing mapping refresh into
`_refresh_replicated_prefill_mappings` and reuse it from both the
preparation path and the MLA/GQA graph path. This preserves MLA/GQA
metadata and its padded query layout, while DSA/SFA retain their
existing metadata rebuild path with one mapping refresh. No new buffers,
flags, or query-padding logic are introduced.

### Does this PR introduce _any_ user-facing change?

Fixes stale KV cache mappings for replicated-PCP speculative decoding
with MLA/GQA FULL graphs, which can affect draft acceptance. No CLI or
configuration changes.

### How was this patch tested?

- Remote isolated CPU unit tests with mocked BlockTables operations:
**54 passed** across the following files:
  ```bash
python3 -m pytest tests/ut/worker/test_mtp_pcp_speculator_v2.py
tests/ut/worker/test_attn_utils_v2.py -q --tb=short
  ```
- The selected regression tests against the unpatched baseline produced
**4 failed, 14 passed, 19 deselected**. They detect missing MLA/GQA
refreshes and stale slot values.
- Coverage includes all four attention architectures, request/block
changes, padded slots, stable buffer addresses, metadata identity, and
non-PCP/dummy/missing-batch bypasses.
- Ruff lint, Ruff format check on changed files, and `git diff --check`
passed. The full `format.sh ci` check was not run because Git Bash and
pre-commit are unavailable in the local Windows environment.

**Validation status:** These tests verify the refresh and buffer
contracts using CPU tensors and mocks; they do not exercise NPU mapping
kernels or graph replay. A service startup failure was reported after
deployment and has not yet been diagnosed. NPU startup/capture/replay,
end-to-end correctness, and acceptance-rate comparison for this exact
patch remain pending.

Per-request KV-length correction is outside the scope of this PR.
Existing length construction and update behavior is unchanged.


- vLLM main:
vllm-project/vllm@a97dacb

---------

Signed-off-by: Madilyn537 <zhexuanwu12@gmail.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
vllm-project#16300)

### What this PR does / why we need it?

With replicated PCP, the first draft stage under FULL graph replay must
refresh the persistent block-table and slot-mapping buffers captured by
the draft model. The MLA/GQA path reuses target attention metadata and
skips `_prepare_replicated_prefill_attn`, leaving those default buffers
stale between batches or draft steps.

Extract the existing mapping refresh into
`_refresh_replicated_prefill_mappings` and reuse it from both the
preparation path and the MLA/GQA graph path. This preserves MLA/GQA
metadata and its padded query layout, while DSA/SFA retain their
existing metadata rebuild path with one mapping refresh. No new buffers,
flags, or query-padding logic are introduced.

### Does this PR introduce _any_ user-facing change?

Fixes stale KV cache mappings for replicated-PCP speculative decoding
with MLA/GQA FULL graphs, which can affect draft acceptance. No CLI or
configuration changes.

### How was this patch tested?

- Remote isolated CPU unit tests with mocked BlockTables operations:
**54 passed** across the following files:
  ```bash
python3 -m pytest tests/ut/worker/test_mtp_pcp_speculator_v2.py
tests/ut/worker/test_attn_utils_v2.py -q --tb=short
  ```
- The selected regression tests against the unpatched baseline produced
**4 failed, 14 passed, 19 deselected**. They detect missing MLA/GQA
refreshes and stale slot values.
- Coverage includes all four attention architectures, request/block
changes, padded slots, stable buffer addresses, metadata identity, and
non-PCP/dummy/missing-batch bypasses.
- Ruff lint, Ruff format check on changed files, and `git diff --check`
passed. The full `format.sh ci` check was not run because Git Bash and
pre-commit are unavailable in the local Windows environment.

**Validation status:** These tests verify the refresh and buffer
contracts using CPU tensors and mocks; they do not exercise NPU mapping
kernels or graph replay. A service startup failure was reported after
deployment and has not yet been diagnosed. NPU startup/capture/replay,
end-to-end correctness, and acceptance-rate comparison for this exact
patch remain pending.

Per-request KV-length correction is outside the scope of this PR.
Existing length construction and update behavior is unchanged.

- vLLM main:
vllm-project/vllm@a97dacb

---------

Signed-off-by: Madilyn537 <zhexuanwu12@gmail.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants