Skip to content

[BugFix] fix fiaV2 contiguous err in GQA - #13456

Merged
zzzzwwjj merged 1 commit into
vllm-project:mainfrom
zzzzzz198:main
Aug 5, 2026
Merged

zzzzwwjj merged 1 commit into
vllm-project:mainfrom
zzzzzz198:main

Conversation

@zzzzzz198

@zzzzzz198 zzzzzz198 commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

When CANN is updated to 9.1.0, the fused_infer_attention operator added a new constraint that key and value should be contiguous when input into this FIAV2 operator. This PR makes the implementation.

Does this PR introduce any user-facing change?

NO.

How was this patch tested?

CI passed.

Signed-off-by: zhaoyhh <zhaoyuhui14@h-partners.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a bug in the GQA (Grouped Query Attention) implementation within the vLLM Ascend backend. By enforcing contiguous memory layouts for key and value tensors, it resolves potential compatibility issues when invoking the NPU-specific fused attention kernel.

Highlights

  • Memory Layout Fix: Ensured that key and value tensors are contiguous before passing them to the NPU fused attention score function to prevent potential memory layout errors during GQA operations.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request ensures that the key and value tensors are contiguous before being passed to the npu_fused_infer_attention_score_v2 function in vllm_ascend/attention/attention_v1.py. The review feedback correctly points out that the PR title and description violate the repository's style guide and provides a compliant suggestion.

Comment on lines +1316 to +1317
key.contiguous(),
value.contiguous(),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The Pull Request title and description do not adhere to the repository style guide. The PR description is currently empty, and the PR title is missing the module prefix.\n\nPlease update the PR title and description to match the following suggested formats:\n\nSuggested PR Title:\n\nmarkdown\n[Attention][BugFix] fix fiaV2 contiguous err in GQA\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\n\nThis PR fixes a contiguous tensor error in Grouped Query Attention (GQA) when using `npu_fused_infer_attention_score_v2`. In GQA, the key and value tensors can be non-contiguous (e.g., due to slicing or striding). Passing non-contiguous tensors to `npu_fused_infer_attention_score_v2` causes runtime errors. This is resolved by calling `.contiguous()` on the `key` and `value` tensors before passing them to the attention score function.\n\n### Does this PR introduce _any_ user-facing change?\n\nNo.\n\n### How was this patch tested?\n\nTested with existing attention tests and GQA workloads on Ascend NPU.\n

References
  1. The Pull Request title and summary must follow the repository style guide format. The title should be in the format [Branch][Module][Action] Pull Request Title, and the summary should contain sections for 'What this PR does / why we need it?', 'Does this PR introduce any user-facing change?', and 'How was this patch tested?'. (link)

@github-actions

github-actions Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@zzzzzz198

zzzzzz198 commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Failed:

  • E2E: still in progress, retry /rerun after completion

@zzzzzz198

zzzzzz198 commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun:

  • E2E

@zzzzwwjj
zzzzwwjj merged commit 0b5dd06 into vllm-project:main Aug 5, 2026
95 of 97 checks passed
shenhui-cli added a commit to shenhui-cli/vllm-ascend that referenced this pull request Aug 5, 2026
…into vllm-new

# By shenhui-cli (8) and others
# Via GitHub (1) and shenhui-cli (1)
* 'vllm-new' of https://github.com/shenhui-cli/vllm-ascend: (34 commits)
  When the PR modifies any files in the CSRC folder, skip the test case filtering logic and execute all test cases by default.
  [MRV2][Bugfix] Fix two triton ops errors in model_runner_v2 num_nans_kernel and apply_panalties (vllm-project#13159)
  [Cherry-pick][main][Doc][BugFix] Update proxy script name in DeepSeek-V3.2 tutorial (from vllm-project#13537) (vllm-project#13575)
  [Refactor][quantization]remove mxfp_compat compatibility shim and inline dtype references (vllm-project#13447)
  [MTP][BugFix] Preserve MoE weight loaders during online updates (vllm-project#13337)
  [Feature][xlite] Support MLA and DSA (`DeepseekV3ForCausalLM`, `DeepseekV32ForCausalLM` and `GlmMoeDsaForCausalLM`) in the xlite adapter (vllm-project#13378)
  [BugFix][xlite] Fix DP metadata handling in XliteWrapper and pass down the max number of tokens across all DP ranks for collective communications. (vllm-project#11213)
  [Refactor][Ops] Move expert routing into router classes (vllm-project#13417)
  [Platform][Refactor] Refactor NPUPlatform for better organization and clarity (vllm-project#13484)
  [BugFix] fix fiaV2 contiguous err in GQA (vllm-project#13456)
  [feature][KV Offload] Support Sparse KV Cache Offload (vllm-project#13026)
  [Bugfix][MRV2]Skip D2H copy and synchronize when spec decode is not active (vllm-project#13382)
  [Bugfix] Change AscendSFAIndexerCacheSpec to inherit MLAAttentionSpec (vllm-project#12849)
  [BugFix][FusedMoE] Restore BF16 quant method initialization (vllm-project#13412)
  [Doc] Fix link errors and add section anchors (vllm-project#13485)
  [TEST]Revise the A3 case (vllm-project#13495)
  [CI] modify default cann_version and add build_type support in nightly (vllm-project#13480)
  [CI]Improve logging, optimize function-level recommendation algorithm, and add early exit for no product code changes (vllm-project#13472)
  This PR restricts the test case discovery scope to the following directories: tests/e2e/pull_request/ tests/ut/
  This PR restricts the test case discovery scope to the following directories: tests/e2e/pull_request/ tests/ut/
  ...

# Conflicts:
#	.github/workflows/scripts/test_selector.py
HMCCMH pushed a commit to hotTea123/vllm-ascend that referenced this pull request Aug 12, 2026
### What this PR does / why we need it?
When CANN is updated to 9.1.0, the fused_infer_attention operator added
a new constraint that key and value should be contiguous when input into
this FIAV2 operator. This PR makes the implementation.

### Does this PR introduce _any_ user-facing change?
NO.

### How was this patch tested?
CI passed.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

Signed-off-by: zhaoyhh <zhaoyuhui14@h-partners.com>
Co-authored-by: zhaoyhh <zhaoyuhui14@h-partners.com>
MmMmaru pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 19, 2026
### What this PR does / why we need it?
When CANN is updated to 9.1.0, the fused_infer_attention operator added
a new constraint that key and value should be contiguous when input into
this FIAV2 operator. This PR makes the implementation.

### Does this PR introduce _any_ user-facing change?
NO.

### How was this patch tested?
CI passed.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

Signed-off-by: zhaoyhh <zhaoyuhui14@h-partners.com>
Co-authored-by: zhaoyhh <zhaoyuhui14@h-partners.com>
shiqiangA pushed a commit to shiqiangA/vllm-ascend that referenced this pull request Aug 20, 2026
### What this PR does / why we need it?
When CANN is updated to 9.1.0, the fused_infer_attention operator added
a new constraint that key and value should be contiguous when input into
this FIAV2 operator. This PR makes the implementation.

### Does this PR introduce _any_ user-facing change?
NO.

### How was this patch tested?
CI passed.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@0351e9a

Signed-off-by: zhaoyhh <zhaoyuhui14@h-partners.com>
Co-authored-by: zhaoyhh <zhaoyuhui14@h-partners.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants