Conversation
…vllm-project#16487) Fix shape mismatches when DCP is enabled with PD-disaggregated recomputation. Short recomputation requests, including last-token recomputation, are now correctly classified as decode requests in the attention, MLA, and SFA metadata builders. Cache DCP state during initialization and use the builder's configuration so metadata construction works outside the current-config context. No API or configuration changes. This fixes failures in DCP-enabled PD-disaggregated recomputation. - Added regression tests covering PD recomputation, mixed query lengths, and behavior when the override does not apply. - Added coverage for metadata construction without an active current-config context. - vLLM main: vllm-project/vllm@a97dacb --------- Signed-off-by: weiguihua2 <weiguihua2@huawei.com> (cherry picked from commit cab3719)
lllyys
requested review from
wangxiyuan,
weijinqian0 and
whx-sjtu
as code owners
September 18, 2026 02:43
Contributor
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [BugFix] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
Contributor
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Attention][Feature] Cache DCP state in metadata builders for PD decode recompute schedulerSuggested PR Summary:
### What this PR does / why we need it?
This PR caches the DCP (Distributed Context Parallel) enabled state during the initialization of metadata builders (`AscendAttentionDCPMetadataBuilder`, `AscendSFADCPMetadataBuilder`, and `AscendMLAMetadataBuilder`). This cached state, along with the builder's own `vllm_config`, is used to determine whether to treat short extends as decodes when splitting decodes and prefills. This avoids relying on `get_current_vllm_config()` during metadata building, which can run outside the active vLLM config context, and ensures correct behavior when the PD decode recompute scheduler is enabled.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Added unit tests in `test_attention_cp.py`, `test_mla_v1.py`, and `test_sfa_cp.py` to verify that the decodes/prefills split correctly uses the builder's config and cached DCP state without requiring an active vLLM config context.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this PR does / why we need it?
Cherry-pick of #16487 to
releases/v0.26.0rc.Fix shape mismatches when DCP is enabled with PD-disaggregated recomputation:
short recomputation requests (including last-token recomputation) are now
correctly classified as decode requests in the attention, MLA, and SFA
metadata builders. DCP state is cached during builder initialization and the
builder's own configuration is used, so metadata construction works outside
the current-config context.
Original PR: #16487 (main commit cab3719)
Backport conflicts and how they were resolved
The cherry-pick conflicted in 4 files (this branch carries 165 backports of
its own;
test_attention_cp.py/test_mla_v1.pylive undertests/ut/attention/a2/here and were applied via rename detection).vllm_ascend/attention/mla_v1.py— 3 hunks:is_pd_decode_recompute_scheduler_enabled(already present inutils.pyfrom an earlier backport).__init__: main's hunk carried the whole PCP block(
pcp_size/pcp_enabled/pcp_rank), which does not exist on thisbranch; added only
self.dcp_enabled = enable_dcp().not (pcp_enabled or dcp_size > 1)to add the DCP + PD-recomputeoverride. This branch's gate is
context_parallel_metadata is None; resolved ascontext_parallel_metadata is None or (self.dcp_enabled and is_pd_decode_recompute_scheduler_enabled(self.vllm_config))— same intent, expressed on this branch's gate. Verified
AscendMlaDCPMetadataBuilder(mla_cp.py) inheritsbuild()fromAscendMLAMetadataBuilder, so the DCP MLA path is covered.attention_cp.py— import hunk only: kept this branch'smemcache_comm_fenceimport path (main moved that symbol elsewhere) andadded the helper import. All functional hunks applied cleanly.
sfa_cp.py— import hunk only: main's side carried its wholeutils/weight_switch/pcp import section, none of which exists here; added
only
from vllm_ascend.utils import is_pd_decode_recompute_scheduler_enabled.Functional hunks applied cleanly.
tests/ut/attention/test_sfa_cp.py— main's version oftest_sfa_dcp_builder_sizes_replicated_view_from_padded_block_tableisparametrized against buffers/patches this branch does not have; applied
only the PR's intent onto this branch's shape of the test (patch
enable_dcp, assert it is called exactly once at init and cached).Adaptations of cleanly-applied test hunks:
patch("vllm.config.get_current_vllm_config_or_none", ...)→patch("vllm.config.get_current_vllm_config", side_effect=AssertionError):this branch's helper fallback calls the latter, so patching the former
would never simulate "no current config context".
patch.object(builder, "_update_parallel_slot_mapping")→_update_dsa_cp_slot_mapping_for_dcp(this branch's method name).pcp_size in (1, 2)— reducedto the single no-DCP case (no PCP in mla_v1 on this branch); it still
asserts the DCP-only override is never evaluated when DCP is off.
Independent review (Codex)
Reviewed pre-submission with Codex against both branches and both commits.
Verdict: APPROVE WITH NITS; both findings addressed before submission:
context_parallel_metadata=None, so theis Nonegate short-circuitedand the new override was never exercised. Fixed: the test now populates
context_parallel_metadata, so classification hinges on the overrideitself, and a recompute-off case asserts short extends stay prefills.
exports both config getters; the patch-target change is needed because
this branch's helper fallback calls
get_current_vllm_config). Corrected.Codex also confirmed: no functional hunks dropped, no other
split_decodes_and_prefillscall sites on this branch need the same fix,and no main-only code was pulled in.
Does this PR introduce any user-facing change?
No. Fixes failures in DCP-enabled PD-disaggregated recomputation.
How was this patch tested?
(see above), in
tests/ut/attention/a2/test_attention_cp.py,tests/ut/attention/a2/test_mla_v1.py,tests/ut/attention/test_sfa_cp.py.