[Bugfix] Fix the incorrect use of the output parameter in _forward_fia_slidingwindow by acat-rw · Pull Request #6469 · vllm-project/vllm-ascend

acat-rw · 2026-02-01T11:33:47Z

What this PR does / why we need it?

Fix the incorrect use of the output parameter in _forward_fia_slidingwindow:

# Original (incorrect)
output, _ = torch_npu.npu_fused_infer_attention_score(...)
output= output.view(batch_size, self.num_heads, self.head_size)

In the original writing, the output parameter was directly assigned a new value, which is inconsistent with the interface definition, resulting in the inability to directly update output when calling externally.

attn_output, _ = torch_npu.npu_fused_infer_attention_score(...)
attn_output = attn_output.view(batch_size, self.num_heads, self.head_size)
output[:batch_size] = attn_output[:batch_size]

Does this PR introduce any user-facing change?

No change.

Co-authored-by: GoCHuggch59135228@163.com

How was this patch tested?

vLLM ascend version: v0.13.0rc1

github-actions · 2026-02-01T11:33:58Z

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

A PR should do only one thing, smaller PRs enable faster reviews.
Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

gemini-code-assist · 2026-02-01T11:34:02Z

Summary of Changes

Hello @acat-rw, I'm Gemini Code Assist¹! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request resolves a bug in the _forward_fia_slidingwindow function within the vllm_ascend attention module. The core issue was an improper handling of the output parameter, which prevented its intended modification by external callers. The fix ensures that the function correctly updates the output parameter, maintaining the integrity of the function's interface and expected behavior.

Highlights

Parameter Handling Correction: The pull request rectifies an issue in _forward_fia_slidingwindow where the output parameter was incorrectly reassigned, preventing external updates.
Intermediate Variable Introduction: An intermediate variable, attn_output, is now used to capture the result of torch_npu.npu_fused_infer_attention_score.
Correct Output Update: The reshaped attn_output is now explicitly assigned to the output parameter, ensuring proper data flow and adherence to the interface.

🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature	Command	Description
Code Review	`/gemini review`	Performs a code review for the current pull request in its current state.
Pull Request Summary	`/gemini summary`	Provides a summary of the current pull request in its current state.
Comment	@gemini-code-assist	Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help	`/gemini help`	Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

gemini-code-assist

Code Review

This pull request provides a good fix for the incorrect handling of the output parameter in the _forward_fia_slidingwindow function. The change correctly ensures that the output tensor is updated in-place, which aligns with its role as an output parameter.

To align with the repository's style guide, I suggest updating the pull request title. The pull request summary is well-written and follows the prescribed format.

Suggested PR Title:

[Attention][BugFix] Fix incorrect use of output parameter in _forward_fia_slidingwindow

Suggested PR Summary:

### What this PR does / why we need it?
Fix the incorrect use of the `output` parameter in `_forward_fia_slidingwindow`:

Original (incorrect)

output, _ = torch_npu.npu_fused_infer_attention_score(...)
output= output.view(batch_size, self.num_heads, self.head_size)


In the original writing, the `output `parameter was directly assigned a new value, which is inconsistent with the interface definition, resulting in the inability to directly update `output `when calling externally.

attn_output, _ = torch_npu.npu_fused_infer_attention_score(...)
attn_output = attn_output.view(batch_size, self.num_heads, self.head_size)
output[:batch_size] = attn_output[:batch_size]


### Does this PR introduce _any_ user-facing change?
No change.

### How was this patch tested?
vLLM ascend version: v0.13.0rc1

- vLLM version: v0.14.1
- vLLM main: https://github.com/vllm-project/vllm/commit/dc917cceb877dfd13f98c538c4c96158047d98bd

…window Signed-off-by: acat-rw <892882856@qq.com>

weijinqian0 · 2026-02-05T12:58:21Z

Thanks @GoCHug for same fixed pr. #6558

…to qwen3next_rebase * 'main' of https://github.com/vllm-project/vllm-ascend: (59 commits) [Feat.]: 310p support MOE models (vllm-project#6530) [Doc] backport 0.13.0 release note (vllm-project#6584) [CI] Update UT CANN version to 8.5.0 for main branch (vllm-project#6564) [CI] Change A2 runner (vllm-project#6557) [Bugfix] Fix the incorrect use of the output parameter in _forward_fia_slidingwindow (vllm-project#6469) [main2main] upgrade vllm main 0202 (vllm-project#6560) [CI][npugraph_ex]Fix npugraph ex e2e test (vllm-project#6553) [Feature]KV pool supports sparse attention (vllm-project#6339) [bugfix]Fix accuracy issue in PCP/DCP with speculative decoding (vllm-project#6491) perf: adaptive block size selection in linear_persistent kernel (vllm-project#6537) [ModelRunner][Fix] Pads query_start_loc to satisfy FIA/TND constraint (vllm-project#6475) [Bugfix]Fix of Pooling Code and Update of Pooling Usage Guide (vllm-project#6126) [Fusion] Add rmsnorm dynamic quant fusion pass (vllm-project#6274) [Bugfix] Synchronize only the current stream to avoid device sync (vllm-project#6432) [CI] Add long and short prompt tests for DeepSeek-V3.2 (vllm-project#6499) [Refactor] MLP weight prefetch to consistency with MoE Model's prefetching in terms of code and usage (vllm-project#6442) [bugfix][npugraph_ex]duplicate pattern issue (vllm-project#6513) [bugfix][npugraph_ex]add the extra check for allreduce rmsnorm fusion pass (vllm-project#6430) [Quant] GLM4.7-Flash Support W8A8 (vllm-project#6492) [Nightly][BugFix] Remove kv_cache nz test case for test_mla_preprocess_nq.py (vllm-project#6505) ...

…a_slidingwindow (vllm-project#6469) ### What this PR does / why we need it? Fix the incorrect use of the `output` parameter in `_forward_fia_slidingwindow`: ``` # Original (incorrect) output, _ = torch_npu.npu_fused_infer_attention_score(...) output= output.view(batch_size, self.num_heads, self.head_size) ``` In the original writing, the `output `parameter was directly assigned a new value, which is inconsistent with the interface definition, resulting in the inability to directly update `output `when calling externally. ``` attn_output, _ = torch_npu.npu_fused_infer_attention_score(...) attn_output = attn_output.view(batch_size, self.num_heads, self.head_size) output[:batch_size] = attn_output[:batch_size] ``` ### Does this PR introduce _any_ user-facing change? No change. Co-authored-by: GoCHug<gch59135228@163.com> ### How was this patch tested? vLLM ascend version: v0.13.0rc1 Signed-off-by: acat-rw <892882856@qq.com> Signed-off-by: momochenchuw <chenchuw@huawei.com>

…a_slidingwindow (vllm-project#6469) ### What this PR does / why we need it? Fix the incorrect use of the `output` parameter in `_forward_fia_slidingwindow`: ``` # Original (incorrect) output, _ = torch_npu.npu_fused_infer_attention_score(...) output= output.view(batch_size, self.num_heads, self.head_size) ``` In the original writing, the `output `parameter was directly assigned a new value, which is inconsistent with the interface definition, resulting in the inability to directly update `output `when calling externally. ``` attn_output, _ = torch_npu.npu_fused_infer_attention_score(...) attn_output = attn_output.view(batch_size, self.num_heads, self.head_size) output[:batch_size] = attn_output[:batch_size] ``` ### Does this PR introduce _any_ user-facing change? No change. Co-authored-by: GoCHug<gch59135228@163.com> ### How was this patch tested? vLLM ascend version: v0.13.0rc1 Signed-off-by: acat-rw <892882856@qq.com> Signed-off-by: zrj026 <zhangrunjiang026@gmail.com>

…a_slidingwindow (vllm-project#6469) ### What this PR does / why we need it? Fix the incorrect use of the `output` parameter in `_forward_fia_slidingwindow`: ``` # Original (incorrect) output, _ = torch_npu.npu_fused_infer_attention_score(...) output= output.view(batch_size, self.num_heads, self.head_size) ``` In the original writing, the `output `parameter was directly assigned a new value, which is inconsistent with the interface definition, resulting in the inability to directly update `output `when calling externally. ``` attn_output, _ = torch_npu.npu_fused_infer_attention_score(...) attn_output = attn_output.view(batch_size, self.num_heads, self.head_size) output[:batch_size] = attn_output[:batch_size] ``` ### Does this PR introduce _any_ user-facing change? No change. Co-authored-by: GoCHug<gch59135228@163.com> ### How was this patch tested? vLLM ascend version: v0.13.0rc1 Signed-off-by: acat-rw <892882856@qq.com>

…a_slidingwindow (vllm-project#6469) ### What this PR does / why we need it? Fix the incorrect use of the `output` parameter in `_forward_fia_slidingwindow`: ``` # Original (incorrect) output, _ = torch_npu.npu_fused_infer_attention_score(...) output= output.view(batch_size, self.num_heads, self.head_size) ``` In the original writing, the `output `parameter was directly assigned a new value, which is inconsistent with the interface definition, resulting in the inability to directly update `output `when calling externally. ``` attn_output, _ = torch_npu.npu_fused_infer_attention_score(...) attn_output = attn_output.view(batch_size, self.num_heads, self.head_size) output[:batch_size] = attn_output[:batch_size] ``` ### Does this PR introduce _any_ user-facing change? No change. Co-authored-by: GoCHug<gch59135228@163.com> ### How was this patch tested? vLLM ascend version: v0.13.0rc1 Signed-off-by: acat-rw <892882856@qq.com> Signed-off-by: zrj026 <zhangrunjiang026@gmail.com>

…a_slidingwindow (vllm-project#6469) ### What this PR does / why we need it? Fix the incorrect use of the `output` parameter in `_forward_fia_slidingwindow`: ``` # Original (incorrect) output, _ = torch_npu.npu_fused_infer_attention_score(...) output= output.view(batch_size, self.num_heads, self.head_size) ``` In the original writing, the `output `parameter was directly assigned a new value, which is inconsistent with the interface definition, resulting in the inability to directly update `output `when calling externally. ``` attn_output, _ = torch_npu.npu_fused_infer_attention_score(...) attn_output = attn_output.view(batch_size, self.num_heads, self.head_size) output[:batch_size] = attn_output[:batch_size] ``` ### Does this PR introduce _any_ user-facing change? No change. Co-authored-by: GoCHug<gch59135228@163.com> ### How was this patch tested? vLLM ascend version: v0.13.0rc1 Signed-off-by: acat-rw <892882856@qq.com>

acat-rw requested review from weijinqian0 and whx-sjtu as code owners February 1, 2026 11:33

gemini-code-assist bot reviewed Feb 1, 2026

View reviewed changes

weijinqian0 approved these changes Feb 2, 2026

View reviewed changes

whx-sjtu approved these changes Feb 2, 2026

View reviewed changes

Fix the incorrect use of the output parameter in _forward_fia_sliding…

059c89f

…window Signed-off-by: acat-rw <892882856@qq.com>

acat-rw force-pushed the main branch from 349a66d to 059c89f Compare February 2, 2026 07:02

weijinqian0 added ready read for review ready-for-test start test by label for PR labels Feb 2, 2026

Merge branch 'vllm-project:main' into main

6221961

weijinqian0 merged commit 8e66299 into vllm-project:main Feb 5, 2026
17 checks passed

weijinqian0 mentioned this pull request Feb 5, 2026

[BugFix] Fixed an accuracy issue caused by the fact that the attention calculation result was not saved to the output when the FIA used the sliding window method. #6558

Closed

PFloyd0 mentioned this pull request Mar 27, 2026

[Bug]: 自研模型从0.11升级到0.13精度出现问题 #7712

Open

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[Bugfix] Fix the incorrect use of the output parameter in _forward_fia_slidingwindow#6469

[Bugfix] Fix the incorrect use of the output parameter in _forward_fia_slidingwindow#6469
weijinqian0 merged 2 commits intovllm-project:mainfrom
acat-rw:main

acat-rw commented Feb 1, 2026 •

edited by weijinqian0

Loading

Uh oh!

github-actions bot commented Feb 1, 2026

Uh oh!

gemini-code-assist bot commented Feb 1, 2026

Uh oh!

gemini-code-assist bot left a comment

Uh oh!

weijinqian0 commented Feb 5, 2026 •

edited

Loading

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

Conversation

acat-rw commented Feb 1, 2026 • edited by weijinqian0 Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

What this PR does / why we need it?

Does this PR introduce any user-facing change?

How was this patch tested?

Uh oh!

github-actions bot commented Feb 1, 2026

Uh oh!

gemini-code-assist bot commented Feb 1, 2026

Summary of Changes

Highlights

Footnotes

Uh oh!

gemini-code-assist bot left a comment

Choose a reason for hiding this comment

Code Review

Original (incorrect)

Uh oh!

weijinqian0 commented Feb 5, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

acat-rw commented Feb 1, 2026 •

edited by weijinqian0

Loading

weijinqian0 commented Feb 5, 2026 •

edited

Loading