[Bugfix][MRV2]Skip D2H copy and synchronize when spec decode is not active - #13382
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request optimizes the NPU model runner by removing redundant synchronization overhead during inference. By recognizing that token counts are deterministic when speculative decoding is not active, the changes bypass unnecessary device-to-host memory transfers, effectively reducing idle time and improving performance for standard workloads. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[v2][ModelRunner][BugFix] Optimize token count synchronization and skip redundant D2H copiesSuggested PR Summary:
### What this PR does / why we need it?
This PR optimizes the model runner by skipping redundant Device-to-Host (D2H) copies of computed tokens when Multi-Token Prediction (MTP) is not used (i.e., `self.speculator` is `None`). Instead, `num_computed_tokens_cpu` is directly synchronized from `num_computed_tokens_np` during `_update_seq_lens_cpu`.
Additionally, feedback has been provided to vectorize the loop that copies these token counts to avoid Python loop overhead and improve performance, especially for larger batch sizes.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Tested with existing model runner unit tests.fdf9215 to
49cdc70
Compare
|
/rerun Failed:
|
|
/rerun |
1 similar comment
|
/rerun |
|
/rerun Failed:
|
|
/rerun Rerun:
|
|
/rerun Failed:
|
2 similar comments
|
/rerun Failed:
|
|
/rerun Failed:
|
|
/retry |
49cdc70 to
6b2af42
Compare
|
/rerun Rerun:
|
Signed-off-by: XYQ <xiayingqing0928@163.com>
6b2af42 to
08d673c
Compare
|
/rerun Failed:
|
2 similar comments
|
/rerun Failed:
|
|
/rerun Failed:
|
…into vllm-new # By shenhui-cli (8) and others # Via GitHub (1) and shenhui-cli (1) * 'vllm-new' of https://github.com/shenhui-cli/vllm-ascend: (34 commits) When the PR modifies any files in the CSRC folder, skip the test case filtering logic and execute all test cases by default. [MRV2][Bugfix] Fix two triton ops errors in model_runner_v2 num_nans_kernel and apply_panalties (vllm-project#13159) [Cherry-pick][main][Doc][BugFix] Update proxy script name in DeepSeek-V3.2 tutorial (from vllm-project#13537) (vllm-project#13575) [Refactor][quantization]remove mxfp_compat compatibility shim and inline dtype references (vllm-project#13447) [MTP][BugFix] Preserve MoE weight loaders during online updates (vllm-project#13337) [Feature][xlite] Support MLA and DSA (`DeepseekV3ForCausalLM`, `DeepseekV32ForCausalLM` and `GlmMoeDsaForCausalLM`) in the xlite adapter (vllm-project#13378) [BugFix][xlite] Fix DP metadata handling in XliteWrapper and pass down the max number of tokens across all DP ranks for collective communications. (vllm-project#11213) [Refactor][Ops] Move expert routing into router classes (vllm-project#13417) [Platform][Refactor] Refactor NPUPlatform for better organization and clarity (vllm-project#13484) [BugFix] fix fiaV2 contiguous err in GQA (vllm-project#13456) [feature][KV Offload] Support Sparse KV Cache Offload (vllm-project#13026) [Bugfix][MRV2]Skip D2H copy and synchronize when spec decode is not active (vllm-project#13382) [Bugfix] Change AscendSFAIndexerCacheSpec to inherit MLAAttentionSpec (vllm-project#12849) [BugFix][FusedMoE] Restore BF16 quant method initialization (vllm-project#13412) [Doc] Fix link errors and add section anchors (vllm-project#13485) [TEST]Revise the A3 case (vllm-project#13495) [CI] modify default cann_version and add build_type support in nightly (vllm-project#13480) [CI]Improve logging, optimize function-level recommendation algorithm, and add early exit for no product code changes (vllm-project#13472) This PR restricts the test case discovery scope to the following directories: tests/e2e/pull_request/ tests/ut/ This PR restricts the test case discovery scope to the following directories: tests/e2e/pull_request/ tests/ut/ ... # Conflicts: # .github/workflows/scripts/test_selector.py
…ctive (vllm-project#13382) ### What this PR does / why we need it? When speculative decoding (MTP) is not active, NPUModelRunner._update_seq_lens_cpu calls synchronize() to wait for an async device-to-host copy before computing seq_lens_cpu. This is unnecessary: without MTP, num_computed_tokens is deterministic (+= num_scheduled_tokens every step, no token rejection), and the parent class GPUModelRunner already maintains num_computed_tokens_np from scheduler values via update_requests each step. This PR: Skips _copy_num_computed_tokens_to_cpu() in postprocess_sampled when speculator is None Uses num_computed_tokens_np (the parent class's CPU-side scheduler snapshot) instead of the D2H pinned buffer in _update_seq_lens_cpu when speculator is None Eliminates an unnecessary synchronize() call and the associated NPU idle bubble for non-MTP workloads The MTP path is unchanged. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? Profiling before/after on NPU without MTP, confirming the synchronize() bubble is eliminated and NPU utilization gap between steps is reduced. - vLLM version: v0.26.0 - vLLM main: vllm-project/vllm@0351e9a Signed-off-by: XYQ <xiayingqing0928@163.com>
…ctive (vllm-project#13382) ### What this PR does / why we need it? When speculative decoding (MTP) is not active, NPUModelRunner._update_seq_lens_cpu calls synchronize() to wait for an async device-to-host copy before computing seq_lens_cpu. This is unnecessary: without MTP, num_computed_tokens is deterministic (+= num_scheduled_tokens every step, no token rejection), and the parent class GPUModelRunner already maintains num_computed_tokens_np from scheduler values via update_requests each step. This PR: Skips _copy_num_computed_tokens_to_cpu() in postprocess_sampled when speculator is None Uses num_computed_tokens_np (the parent class's CPU-side scheduler snapshot) instead of the D2H pinned buffer in _update_seq_lens_cpu when speculator is None Eliminates an unnecessary synchronize() call and the associated NPU idle bubble for non-MTP workloads The MTP path is unchanged. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? Profiling before/after on NPU without MTP, confirming the synchronize() bubble is eliminated and NPU utilization gap between steps is reduced. - vLLM version: v0.26.0 - vLLM main: vllm-project/vllm@0351e9a Signed-off-by: XYQ <xiayingqing0928@163.com>
…ctive (vllm-project#13382) ### What this PR does / why we need it? When speculative decoding (MTP) is not active, NPUModelRunner._update_seq_lens_cpu calls synchronize() to wait for an async device-to-host copy before computing seq_lens_cpu. This is unnecessary: without MTP, num_computed_tokens is deterministic (+= num_scheduled_tokens every step, no token rejection), and the parent class GPUModelRunner already maintains num_computed_tokens_np from scheduler values via update_requests each step. This PR: Skips _copy_num_computed_tokens_to_cpu() in postprocess_sampled when speculator is None Uses num_computed_tokens_np (the parent class's CPU-side scheduler snapshot) instead of the D2H pinned buffer in _update_seq_lens_cpu when speculator is None Eliminates an unnecessary synchronize() call and the associated NPU idle bubble for non-MTP workloads The MTP path is unchanged. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? Profiling before/after on NPU without MTP, confirming the synchronize() bubble is eliminated and NPU utilization gap between steps is reduced. - vLLM version: v0.26.0 - vLLM main: vllm-project/vllm@0351e9a Signed-off-by: XYQ <xiayingqing0928@163.com>
What this PR does / why we need it?
When speculative decoding (MTP) is not active, NPUModelRunner._update_seq_lens_cpu calls synchronize() to wait for an async device-to-host copy before computing seq_lens_cpu. This is unnecessary: without MTP, num_computed_tokens is deterministic (+= num_scheduled_tokens every step, no token rejection), and the parent class GPUModelRunner already maintains num_computed_tokens_np from scheduler values via update_requests each step.
This PR:
Skips _copy_num_computed_tokens_to_cpu() in postprocess_sampled when speculator is None
Uses num_computed_tokens_np (the parent class's CPU-side scheduler snapshot) instead of the D2H pinned buffer in _update_seq_lens_cpu when speculator is None
Eliminates an unnecessary synchronize() call and the associated NPU idle bubble for non-MTP workloads
The MTP path is unchanged.
Does this PR introduce any user-facing change?
No.
How was this patch tested?
Profiling before/after on NPU without MTP, confirming the synchronize() bubble is eliminated and NPU utilization gap between steps is reduced.