Skip to content

[Bugfix][MRV2]Skip D2H copy and synchronize when spec decode is not active - #13382

Merged
linfeng-yuan merged 1 commit into
vllm-project:mainfrom
xiayingqing:perf/skip-d2h-sync-non-mtp
Aug 5, 2026
Merged

linfeng-yuan merged 1 commit into
vllm-project:mainfrom
xiayingqing:perf/skip-d2h-sync-non-mtp

Conversation

@xiayingqing

@xiayingqing xiayingqing commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

When speculative decoding (MTP) is not active, NPUModelRunner._update_seq_lens_cpu calls synchronize() to wait for an async device-to-host copy before computing seq_lens_cpu. This is unnecessary: without MTP, num_computed_tokens is deterministic (+= num_scheduled_tokens every step, no token rejection), and the parent class GPUModelRunner already maintains num_computed_tokens_np from scheduler values via update_requests each step.

This PR:

Skips _copy_num_computed_tokens_to_cpu() in postprocess_sampled when speculator is None
Uses num_computed_tokens_np (the parent class's CPU-side scheduler snapshot) instead of the D2H pinned buffer in _update_seq_lens_cpu when speculator is None
Eliminates an unnecessary synchronize() call and the associated NPU idle bubble for non-MTP workloads
The MTP path is unchanged.

Does this PR introduce any user-facing change?

No.

How was this patch tested?

Profiling before/after on NPU without MTP, confirming the synchronize() bubble is eliminated and NPU utilization gap between steps is reduced.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request optimizes the NPU model runner by removing redundant synchronization overhead during inference. By recognizing that token counts are deterministic when speculative decoding is not active, the changes bypass unnecessary device-to-host memory transfers, effectively reducing idle time and improving performance for standard workloads.

Highlights

  • Optimization of NPU Synchronization: Eliminated an unnecessary device-to-host (D2H) synchronization call in the NPU model runner when speculative decoding (MTP) is disabled.
  • Conditional Logic Update: Updated postprocess_sampled and _update_seq_lens_cpu to conditionally skip D2H copies and rely on the CPU-side scheduler snapshot (num_computed_tokens_np) for non-MTP workloads.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@xiayingqing xiayingqing changed the title Skip D2H copy and synchronize when spec decode is not active [Bugfix][MRV2]Skip D2H copy and synchronize when spec decode is not active Aug 3, 2026
@github-actions

github-actions Bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[v2][ModelRunner][BugFix] Optimize token count synchronization and skip redundant D2H copies

Suggested PR Summary:

### What this PR does / why we need it?
This PR optimizes the model runner by skipping redundant Device-to-Host (D2H) copies of computed tokens when Multi-Token Prediction (MTP) is not used (i.e., `self.speculator` is `None`). Instead, `num_computed_tokens_cpu` is directly synchronized from `num_computed_tokens_np` during `_update_seq_lens_cpu`.

Additionally, feedback has been provided to vectorize the loop that copies these token counts to avoid Python loop overhead and improve performance, especially for larger batch sizes.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Tested with existing model runner unit tests.

Comment thread vllm_ascend/worker/v2/model_runner.py
@xiayingqing
xiayingqing force-pushed the perf/skip-d2h-sync-non-mtp branch from fdf9215 to 49cdc70 Compare August 3, 2026 09:26
@xiayingqing

xiayingqing commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

@xiayingqing

Copy link
Copy Markdown
Contributor Author

/rerun

1 similar comment
@xiayingqing

Copy link
Copy Markdown
Contributor Author

/rerun

@xiayingqing

xiayingqing commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

@xiayingqing

xiayingqing commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun:

  • E2E

@xiayingqing

xiayingqing commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

2 similar comments
@xiayingqing

xiayingqing commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

@xiayingqing

xiayingqing commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

@xiayingqing

Copy link
Copy Markdown
Contributor Author

/retry

@xiayingqing
xiayingqing force-pushed the perf/skip-d2h-sync-non-mtp branch from 49cdc70 to 6b2af42 Compare August 3, 2026 12:47
@xiayingqing

xiayingqing commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun:

  • E2E

Signed-off-by: XYQ <xiayingqing0928@163.com>
@xiayingqing
xiayingqing force-pushed the perf/skip-d2h-sync-non-mtp branch from 6b2af42 to 08d673c Compare August 4, 2026 01:41
@xiayingqing

xiayingqing commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

2 similar comments
@xiayingqing

xiayingqing commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

@xiayingqing

xiayingqing commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

@linfeng-yuan
linfeng-yuan merged commit baed935 into vllm-project:main Aug 5, 2026
48 checks passed
shenhui-cli added a commit to shenhui-cli/vllm-ascend that referenced this pull request Aug 5, 2026
…into vllm-new

# By shenhui-cli (8) and others
# Via GitHub (1) and shenhui-cli (1)
* 'vllm-new' of https://github.com/shenhui-cli/vllm-ascend: (34 commits)
  When the PR modifies any files in the CSRC folder, skip the test case filtering logic and execute all test cases by default.
  [MRV2][Bugfix] Fix two triton ops errors in model_runner_v2 num_nans_kernel and apply_panalties (vllm-project#13159)
  [Cherry-pick][main][Doc][BugFix] Update proxy script name in DeepSeek-V3.2 tutorial (from vllm-project#13537) (vllm-project#13575)
  [Refactor][quantization]remove mxfp_compat compatibility shim and inline dtype references (vllm-project#13447)
  [MTP][BugFix] Preserve MoE weight loaders during online updates (vllm-project#13337)
  [Feature][xlite] Support MLA and DSA (`DeepseekV3ForCausalLM`, `DeepseekV32ForCausalLM` and `GlmMoeDsaForCausalLM`) in the xlite adapter (vllm-project#13378)
  [BugFix][xlite] Fix DP metadata handling in XliteWrapper and pass down the max number of tokens across all DP ranks for collective communications. (vllm-project#11213)
  [Refactor][Ops] Move expert routing into router classes (vllm-project#13417)
  [Platform][Refactor] Refactor NPUPlatform for better organization and clarity (vllm-project#13484)
  [BugFix] fix fiaV2 contiguous err in GQA (vllm-project#13456)
  [feature][KV Offload] Support Sparse KV Cache Offload (vllm-project#13026)
  [Bugfix][MRV2]Skip D2H copy and synchronize when spec decode is not active (vllm-project#13382)
  [Bugfix] Change AscendSFAIndexerCacheSpec to inherit MLAAttentionSpec (vllm-project#12849)
  [BugFix][FusedMoE] Restore BF16 quant method initialization (vllm-project#13412)
  [Doc] Fix link errors and add section anchors (vllm-project#13485)
  [TEST]Revise the A3 case (vllm-project#13495)
  [CI] modify default cann_version and add build_type support in nightly (vllm-project#13480)
  [CI]Improve logging, optimize function-level recommendation algorithm, and add early exit for no product code changes (vllm-project#13472)
  This PR restricts the test case discovery scope to the following directories: tests/e2e/pull_request/ tests/ut/
  This PR restricts the test case discovery scope to the following directories: tests/e2e/pull_request/ tests/ut/
  ...

# Conflicts:
#	.github/workflows/scripts/test_selector.py
HMCCMH pushed a commit to hotTea123/vllm-ascend that referenced this pull request Aug 12, 2026
…ctive (vllm-project#13382)

### What this PR does / why we need it?

When speculative decoding (MTP) is not active,
NPUModelRunner._update_seq_lens_cpu calls synchronize() to wait for an
async device-to-host copy before computing seq_lens_cpu. This is
unnecessary: without MTP, num_computed_tokens is deterministic (+=
num_scheduled_tokens every step, no token rejection), and the parent
class GPUModelRunner already maintains num_computed_tokens_np from
scheduler values via update_requests each step.

This PR:

Skips _copy_num_computed_tokens_to_cpu() in postprocess_sampled when
speculator is None
Uses num_computed_tokens_np (the parent class's CPU-side scheduler
snapshot) instead of the D2H pinned buffer in _update_seq_lens_cpu when
speculator is None
Eliminates an unnecessary synchronize() call and the associated NPU idle
bubble for non-MTP workloads
The MTP path is unchanged.

### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Profiling before/after on NPU without MTP, confirming the synchronize()
bubble is eliminated and NPU utilization gap between steps is reduced.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

Signed-off-by: XYQ <xiayingqing0928@163.com>
MmMmaru pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 19, 2026
…ctive (vllm-project#13382)

### What this PR does / why we need it?

When speculative decoding (MTP) is not active,
NPUModelRunner._update_seq_lens_cpu calls synchronize() to wait for an
async device-to-host copy before computing seq_lens_cpu. This is
unnecessary: without MTP, num_computed_tokens is deterministic (+=
num_scheduled_tokens every step, no token rejection), and the parent
class GPUModelRunner already maintains num_computed_tokens_np from
scheduler values via update_requests each step.

This PR:

Skips _copy_num_computed_tokens_to_cpu() in postprocess_sampled when
speculator is None
Uses num_computed_tokens_np (the parent class's CPU-side scheduler
snapshot) instead of the D2H pinned buffer in _update_seq_lens_cpu when
speculator is None
Eliminates an unnecessary synchronize() call and the associated NPU idle
bubble for non-MTP workloads
The MTP path is unchanged.

### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Profiling before/after on NPU without MTP, confirming the synchronize()
bubble is eliminated and NPU utilization gap between steps is reduced.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

Signed-off-by: XYQ <xiayingqing0928@163.com>
shiqiangA pushed a commit to shiqiangA/vllm-ascend that referenced this pull request Aug 20, 2026
…ctive (vllm-project#13382)

### What this PR does / why we need it?

When speculative decoding (MTP) is not active,
NPUModelRunner._update_seq_lens_cpu calls synchronize() to wait for an
async device-to-host copy before computing seq_lens_cpu. This is
unnecessary: without MTP, num_computed_tokens is deterministic (+=
num_scheduled_tokens every step, no token rejection), and the parent
class GPUModelRunner already maintains num_computed_tokens_np from
scheduler values via update_requests each step.

This PR:

Skips _copy_num_computed_tokens_to_cpu() in postprocess_sampled when
speculator is None
Uses num_computed_tokens_np (the parent class's CPU-side scheduler
snapshot) instead of the D2H pinned buffer in _update_seq_lens_cpu when
speculator is None
Eliminates an unnecessary synchronize() call and the associated NPU idle
bubble for non-MTP workloads
The MTP path is unchanged.

### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Profiling before/after on NPU without MTP, confirming the synchronize()
bubble is eliminated and NPU utilization gap between steps is reduced.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

Signed-off-by: XYQ <xiayingqing0928@163.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants