[Feature][MRV2] Support DSA PCP with DSpark - #15958
Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly enhances the capabilities of ModelRunner V2 by integrating DeepSeek-V4 DSpark with Prefill Context Parallelism. The changes enable efficient speculative decoding for DSpark models, ensuring proper resource allocation and data handling within the parallel execution environment. This expands the range of supported models and configurations for high-performance inference. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Ignored Files
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Attention][Feature] Support DSpark speculative decoding with DeepSeek-V4 DSA models under PCPSuggested PR Summary:
### What this PR does / why we need it?
This pull request adds support for DSpark speculative decoding with DeepSeek-V4 DSA models under Prefill Context Parallel (PCP). It updates the PCP manager configuration validation, refactors metadata builders to handle request capacity factors dynamically, and integrates replicated PCP configuration handling for the DSpark speculator. Additionally, documentation is updated to reflect these new capabilities and constraints.
Feedback: An issue was identified in `vllm_ascend/worker/v2/spec_decode/pcp_utils.py` where `replace` is imported from `vllm.config` instead of the standard library `dataclasses`, which is fragile and could lead to runtime import errors.
### Does this PR introduce _any_ user-facing change?
Yes, it adds support for DSpark speculative decoding with DeepSeek-V4 DSA models under PCP. Documentation has been updated to guide users on how to configure this feature.
### How was this patch tested?
The changes were tested using new end-to-end tests in `tests/e2e/pull_request/four_card/context_parallel/test_deepseek_v4.py` and updated unit tests in `tests/ut/worker/test_mtp_pcp_speculator_v2.py` and `tests/ut/worker/test_pcp_manager_v2.py`.There was a problem hiding this comment.
Warning
Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.
Pull request overview
Adds MRV2 PCP support for DSpark (DeepSeek-V4) by replicating draft execution with PCP=1 and adjusts DSA metadata buffer sizing for PCP-local request expansion.
Changes:
- Centralizes “replicated PCP” draft-config preparation and applies it to DSpark speculator initialization.
- Ensures DSpark draft attention is initialized under the correct (PCP=1) config context for graph/eager modes.
- Expands DSA per-request buffer capacity for PCP-local rows and updates unit/e2e/docs to cover DSpark+PCP behavior.
Reviewed changes
Copilot reviewed 11 out of 11 changed files in this pull request and generated 6 comments.
Show a summary per file
| File | Description |
|---|---|
| vllm_ascend/worker/v2/spec_decode/pcp_utils.py | Moves replicated-PCP config preparation into a shared utility and tweaks typing/imports. |
| vllm_ascend/worker/v2/spec_decode/dspark/speculator.py | Initializes DSpark draft with PCP=1 config and sets current config when building attention backends. |
| vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py | Switches to importing shared replicated-PCP config helper. |
| vllm_ascend/worker/v2/pcp_manager.py | Allows DSpark as a supported speculative method under MRV2 PCP. |
| vllm_ascend/attention/dsa_v1.py | Introduces request-capacity factor and uses it for DSA per-request buffers. |
| vllm_ascend/attention/context_parallel/dsa_cp.py | Sets PCP builder request-capacity factor to 2 and removes manual buffer resize. |
| tests/ut/worker/test_pcp_manager_v2.py | Extends validate_config tests to DSpark and adds additional mocking for partitioning path. |
| tests/ut/worker/test_mtp_pcp_speculator_v2.py | Updates patching to account for relocated replace usage in pcp_utils. |
| tests/e2e/pull_request/four_card/context_parallel/test_deepseek_v4.py | Adds DSpark PCP e2e coverage and runtime assertions for FULL_DECODE_ONLY graphs. |
| docs/source/user_guide/feature_guide/context_parallel.md | Documents DSpark as supported within PCP speculative decoding constraints. |
| .github/workflows/scripts/test_config.yaml | Updates estimated runtime for the expanded DeepSeek-V4 PCP e2e test. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
6252630 to
02bb90f
Compare
02bb90f to
bdd4c0b
Compare
a9db232 to
258fed0
Compare
PCP DualChunkSwap can expand each scheduler request into two local rows. QLI buffers sized only for max_num_seqs can overflow on mixed batches when graph padding has not already reserved enough capacity. Use a private request capacity factor to allocate prefill and QLI buffers at construction. Keep the global and draft builders at their original capacity and retain larger graph padding allocations. Signed-off-by: leolee <yihao.li@huawei.com>
Enable DeepSeek-V4 DSpark speculative decoding with PCP by keeping the target model sharded across PCP ranks and replicating the draft model with a logical PCP size of 1 on every rank. Share the replicated PCP configuration helper with autoregressive speculators and initialize draft attention with the draft configuration. Add DSpark to the common PCP speculative decoding validation. Add FULL_DECODE_ONLY graph-mode E2E coverage and compatibility documentation. Update existing PCP unit tests for the shared configuration helper and supported speculative decoding methods. Signed-off-by: leolee <yihao.li@huawei.com>
258fed0 to
762af5f
Compare
### What this PR does / why we need it? Support DeepSeek-V4 DSpark with MRV2 PCP by replicating draft execution with PCP=1. Fix DSA request-buffer capacity for PCP-local rows. ### Does this PR introduce _any_ user-facing change? Enable DSpark with PCP and FULL_DECODE_ONLY graphs. No new configuration options. ### How was this patch tested? Two rounds on Ascend `02bb90f97` and vLLM `e6bfe03ad`: TP4, DSpark 5 tokens, FULL_DECODE_ONLY, GPQA Diamond first 100 questions at concurrency 100, 64K context and 60K output budget. Each run restarted the service; temperature=0 and seed=0. | Round | PCP | Completed requests | Correct answers | Output-limit truncations | Acceptance pos 1 | Acceptance pos 2 | Acceptance pos 3 | Acceptance pos 4 | Acceptance pos 5 | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 1 | 1 | 100/100 | 79/100 | 16 | 85.660% | 68.957% | 53.327% | 39.766% | 28.535% | | 1 | 2 | 100/100 | 74/100 | 19 | 86.147% | 69.985% | 54.531% | 40.717% | 28.770% | | 2 | 1 | 100/100 | 75/100 | 16 | 85.801% | 69.148% | 53.385% | 39.368% | 27.631% | | 2 | 2 | 100/100 | 79/100 | 13 | 85.607% | 68.690% | 52.906% | 39.254% | 28.199% | - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: leolee <yihao.li@huawei.com>
### What this PR does / why we need it? Support DeepSeek-V4 DSpark with MRV2 PCP by replicating draft execution with PCP=1. Fix DSA request-buffer capacity for PCP-local rows. ### Does this PR introduce _any_ user-facing change? Enable DSpark with PCP and FULL_DECODE_ONLY graphs. No new configuration options. ### How was this patch tested? Two rounds on Ascend `02bb90f97` and vLLM `e6bfe03ad`: TP4, DSpark 5 tokens, FULL_DECODE_ONLY, GPQA Diamond first 100 questions at concurrency 100, 64K context and 60K output budget. Each run restarted the service; temperature=0 and seed=0. | Round | PCP | Completed requests | Correct answers | Output-limit truncations | Acceptance pos 1 | Acceptance pos 2 | Acceptance pos 3 | Acceptance pos 4 | Acceptance pos 5 | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 1 | 1 | 100/100 | 79/100 | 16 | 85.660% | 68.957% | 53.327% | 39.766% | 28.535% | | 1 | 2 | 100/100 | 74/100 | 19 | 86.147% | 69.985% | 54.531% | 40.717% | 28.770% | | 2 | 1 | 100/100 | 75/100 | 16 | 85.801% | 69.148% | 53.385% | 39.368% | 27.631% | | 2 | 2 | 100/100 | 79/100 | 13 | 85.607% | 68.690% | 52.906% | 39.254% | 28.199% | - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: leolee <yihao.li@huawei.com>
### What this PR does / why we need it? Support DeepSeek-V4 DSpark with MRV2 PCP by replicating draft execution with PCP=1. Fix DSA request-buffer capacity for PCP-local rows. ### Does this PR introduce _any_ user-facing change? Enable DSpark with PCP and FULL_DECODE_ONLY graphs. No new configuration options. ### How was this patch tested? Two rounds on Ascend `02bb90f97` and vLLM `e6bfe03ad`: TP4, DSpark 5 tokens, FULL_DECODE_ONLY, GPQA Diamond first 100 questions at concurrency 100, 64K context and 60K output budget. Each run restarted the service; temperature=0 and seed=0. | Round | PCP | Completed requests | Correct answers | Output-limit truncations | Acceptance pos 1 | Acceptance pos 2 | Acceptance pos 3 | Acceptance pos 4 | Acceptance pos 5 | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 1 | 1 | 100/100 | 79/100 | 16 | 85.660% | 68.957% | 53.327% | 39.766% | 28.535% | | 1 | 2 | 100/100 | 74/100 | 19 | 86.147% | 69.985% | 54.531% | 40.717% | 28.770% | | 2 | 1 | 100/100 | 75/100 | 16 | 85.801% | 69.148% | 53.385% | 39.368% | 27.631% | | 2 | 2 | 100/100 | 79/100 | 13 | 85.607% | 68.690% | 52.906% | 39.254% | 28.199% | - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: leolee <yihao.li@huawei.com> Signed-off-by: tianming2009 <13246728590@163.com>
### What this PR does / why we need it? Support DeepSeek-V4 DSpark with MRV2 PCP by replicating draft execution with PCP=1. Fix DSA request-buffer capacity for PCP-local rows. ### Does this PR introduce _any_ user-facing change? Enable DSpark with PCP and FULL_DECODE_ONLY graphs. No new configuration options. ### How was this patch tested? Two rounds on Ascend `02bb90f97` and vLLM `e6bfe03ad`: TP4, DSpark 5 tokens, FULL_DECODE_ONLY, GPQA Diamond first 100 questions at concurrency 100, 64K context and 60K output budget. Each run restarted the service; temperature=0 and seed=0. | Round | PCP | Completed requests | Correct answers | Output-limit truncations | Acceptance pos 1 | Acceptance pos 2 | Acceptance pos 3 | Acceptance pos 4 | Acceptance pos 5 | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 1 | 1 | 100/100 | 79/100 | 16 | 85.660% | 68.957% | 53.327% | 39.766% | 28.535% | | 1 | 2 | 100/100 | 74/100 | 19 | 86.147% | 69.985% | 54.531% | 40.717% | 28.770% | | 2 | 1 | 100/100 | 75/100 | 16 | 85.801% | 69.148% | 53.385% | 39.368% | 27.631% | | 2 | 2 | 100/100 | 79/100 | 13 | 85.607% | 68.690% | 52.906% | 39.254% | 28.199% | - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: leolee <yihao.li@huawei.com> Signed-off-by: like-0517 <ithwlike@126.com>
…der full cudagraph padding (#16168) ### What this PR does / why we need it? This PR re-applies the DSA-CP + MTP aclgraph support for DeepSeek V4 that was reverted from main in #16181 (revert of #12599), together with a fix for the crash that motivated scrutiny of the original feature: serving concurrent requests with `cudagraph_mode=FULL_DECODE_ONLY` failed with `aclnnInplaceCopy` shape-mismatch errors (EZ1007) or illegal device memory accesses (EZ9999). **Root cause of the crash** In FULL-decode graph mode, the DSA-CP draft path receives `num_reqs` as the *padded* request count — the cudagraph capture bucket size plus the FIA dummy request from mixed-batch padding — which can exceed `scheduler_config.max_num_seqs`. The per-request metadata buffers were sized by `max_num_seqs`, so: 1. `[:num_reqs]` views of these buffers get silently truncated and `copy_()` into them fails with shape mismatches — e.g. `--max-num-seqs 6` with 3 concurrent MTP requests pads 12 tokens to bucket 16, producing `Shape [16] vs [6] do not meet the broadcast condition (EZ1007)`; 2. triton kernels receiving the whole buffer write past its end, corrupting adjacent device memory (`EZ9999: MTE accesses an invalid GM address`). **Fix** Compute one graph-mode-aware capacity up front and size all per-request buffers with it: max_padded_reqs = max(max_num_seqs, max_cudagraph_capture_size) + 1 # +1: FIA dummy request This covers the step-0 buffers (`start_pos_prefill`, `local_query_start_loc`, `local_seq_lens`), the draft buffers (`spec_local_query_start_loc`, `spec_local_seq_lens`, `spec_start_pos`), and the QLI buffers (`qli_seqused_k`, `qli_cmp_residual_k`) with a single capacity source. All usage sites are `[:num_reqs]` slices or whole-buffer kernel inputs, so the change is a pure capacity gain with negligible (int32-level) memory overhead. The re-applied feature is fully adapted to current main: it fuses with the sequence-parallel prefill rework (#15549), the dynamic DSA indexer quant_mode (#16224) and the DSA PCP + DSpark support (#15958) already on main. The feature enables aclgraph capture/replay for DSA-CP with MTP=1 and MTP=3, including stable per-draft-index metadata buffers (`spec_sas_metadata`, `spec_start_pos`, `spec_local_*`, per-draft RoPE cache) so tensor addresses stay fixed across graph capture and replay. ### Does this PR introduce _any_ user-facing change? Yes, it re-enables DSA-CP + MTP with FULL_DECODE_ONLY graphs for DeepSeek V4 (`enable_dsa_cp: true` + `speculative_config` + `cudagraph_mode: FULL_DECODE_ONLY`), which was available before the revert. Users with `max_num_seqs` smaller than the max cudagraph capture size no longer hit EZ1007 / EZ9999 errors under concurrent load. The gain show below: when use dsv4+mtp 3+eager: <img width="371" height="509" alt="c98d2620e0c57e8a2de08fd8281e6903" src="https://github.com/user-attachments/assets/78d1eeb8-e0ae-4421-b555-84ce5d6b8987" /> when use dsv4+ mtp 3 + graph: <img width="374" height="508" alt="74786d0aebc590d420bac3bfbf436f32" src="https://github.com/user-attachments/assets/7f3eb4d5-7bef-4b83-97f0-e9c825c2486e" /> GMS8K accuracy evaluation for MTP+eager and MTP+graph is attached below. dsv4-flash+mtp3+eager Overall report table: +---------+-----------+----------+----------+-------+---------+---------+ | Model | Dataset | Metric | Subset | Num | Score | Cat.0 | +=========+===========+==========+==========+=======+=========+=========+ | dsv4 | gsm8k | mean_acc | main | 1318 | 0.9734 | default | +---------+-----------+----------+----------+-------+---------+---------+ dsv4-flash+mtp3+graph Overall report table: +---------+-----------+----------+----------+-------+---------+---------+ | Model | Dataset | Metric | Subset | Num | Score | Cat.0 | +=========+===========+==========+==========+=======+=========+=========+ | dsv4 | gsm8k | mean_acc | main | 1319 | 0.9742 | default | +---------+-----------+----------+----------+-------+---------+---------+ glm5.2-w4a8c8 + mtp 5 +graph Overall report table: +---------+-----------+----------+----------+-------+---------+---------+ | Model | Dataset | Metric | Subset | Num | Score | Cat.0 | +=========+===========+==========+==========+=======+=========+=========+ | glm5.2 | gsm8k | mean_acc | main | 1319 | 0.978 | default | +---------+-----------+----------+----------+-------+---------+---------+ --------- Signed-off-by: frankie <wangyongsheng686@gmail.com>
What this PR does / why we need it?
Support DeepSeek-V4 DSpark with MRV2 PCP by replicating draft execution with PCP=1. Fix DSA request-buffer capacity for PCP-local rows.
Does this PR introduce any user-facing change?
Enable DSpark with PCP and FULL_DECODE_ONLY graphs. No new configuration options.
How was this patch tested?
Two rounds on Ascend
02bb90f97and vLLMe6bfe03ad: TP4, DSpark 5 tokens, FULL_DECODE_ONLY, GPQA Diamond first 100 questions at concurrency 100, 64K context and 60K output budget. Each run restarted the service; temperature=0 and seed=0.