Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions vllm_ascend/attention/attention_v1.py
Original file line number Diff line number Diff line change
Expand Up @@ -1313,8 +1313,8 @@ def forward_fused_infer_attention(
sparse_mode = 3
attn_output, _ = torch_npu.npu_fused_infer_attention_score_v2(
query,
key,
value,
key.contiguous(),
value.contiguous(),
Comment on lines +1316 to +1317

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The Pull Request title and description do not adhere to the repository style guide. The PR description is currently empty, and the PR title is missing the module prefix.\n\nPlease update the PR title and description to match the following suggested formats:\n\nSuggested PR Title:\n\nmarkdown\n[Attention][BugFix] fix fiaV2 contiguous err in GQA\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\n\nThis PR fixes a contiguous tensor error in Grouped Query Attention (GQA) when using `npu_fused_infer_attention_score_v2`. In GQA, the key and value tensors can be non-contiguous (e.g., due to slicing or striding). Passing non-contiguous tensors to `npu_fused_infer_attention_score_v2` causes runtime errors. This is resolved by calling `.contiguous()` on the `key` and `value` tensors before passing them to the attention score function.\n\n### Does this PR introduce _any_ user-facing change?\n\nNo.\n\n### How was this patch tested?\n\nTested with existing attention tests and GQA workloads on Ascend NPU.\n

References
  1. The Pull Request title and summary must follow the repository style guide format. The title should be in the format [Branch][Module][Action] Pull Request Title, and the summary should contain sections for 'What this PR does / why we need it?', 'Does this PR introduce any user-facing change?', and 'How was this patch tested?'. (link)

num_query_heads=self.num_heads,
num_key_value_heads=self.num_kv_heads,
input_layout="TND",
Expand Down
Loading