Skip to content

[BugFix][SFA] Fix padded-index LSE and empty shards for A5 DCP - #16656

Merged
weiguihua2 merged 3 commits into
vllm-project:mainfrom
recky-c:codex/fix-a5-sfa-dcp-lse
Sep 18, 2026
Merged

weiguihua2 merged 3 commits into
vllm-project:mainfrom
recky-c:codex/fix-a5-sfa-dcp-lse

Conversation

@recky-c

@recky-c recky-c commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

A5 SFA DCP decode can select fewer keys on a rank than its local KV length. The sparse index tensor then contains a valid prefix followed by -1 padding. Using the KV length as the softmax extent includes invalid entries and produces incorrect per-rank normalization for the DCP merge.

This patch ports the focused fixes from our standalone SFA operator experiment into the existing custom operator:

  • Bound sparse-mode-0 LSE computation by the actual valid index prefix.
  • Initialize skipped queries to zero output and zero softmax max/sum (reconstructed LSE is negative infinity), including queries with a nonempty cache but no selected local keys.
  • Initialize n2Size before output initialization.
  • Make prefill KV and RoPE views contiguous after the packed DCP gather.
  • Advertise SFA_DCP_REPLICATED_INDEXER on A5 while retaining the platform capability check.

The current custom operator already accepts PA_BSND with return_softmax_lse, so no host-tiling change is needed on this base. No Mooncake changes or bulk operator-source refresh are included.

Does this PR introduce any user-facing change?

Enables the A5 SFA DCP capability and fixes its BF16 KV-cache path. The validated target is full GLM5.2 with W4A8 weights and both SFA C8 and indexer C8 disabled, TP8/DCP8/EP, eager execution, no MTP or PD disaggregation. This PR does not claim C8 support.

The full-model experiment also applied #16325 (tested revision ec6e301d183c5da1023652744b4e9d6395172b86) for indexer metadata. Those changes are not included here; the newer revision of that PR has not been validated as part of this work.

How was this patch tested?

  • Historical experiment: the equivalent numerical fixes in standalone ops-transformer, together with contiguous prefill inputs, the capability bypass and the above [BugFix][Attention] Build SFA indexer DCP metadata independently #16325 revision, passed all six full 78-layer GLM5.2 functional cases on A5 (arithmetic, factual answer, translation, Python generation, long-context retrieval and repeated arithmetic).
  • This is a minimal port to the custom operator on main 4c5ee3320, not the exact operator build used in that experiment. The port has not yet been recompiled or rerun on NPU. Historical results are not evidence of a completed regression for this exact commit.
  • Added an A5 operator regression comparing output and LSE against a PyTorch reference for 0, 1, 127, 128, 129 and 257 selected keys, including a nonempty cache with no selected keys.
  • Added a prefill regression using non-contiguous packed-cache split views; updated the hardware capability expectation.
  • Ruff 0.14.0 check/format, Python syntax parsing and git diff --check passed. The repository-wide bash format.sh ci could not run because pre-commit is not installed in this local environment. New unit/NPU tests have not been executed in this PR preparation.

Performance, graph mode, MTP, C8 and PD combinations remain outside the completed validation scope.

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces critical fixes for the A5 SFA DCP (Data-Parallel Context Parallel) implementation to resolve issues with padded indices and incorrect softmax normalization. By bounding LSE computations to valid prefixes and ensuring proper initialization of skipped queries, the patch improves the reliability of the attention mechanism. Additionally, it optimizes memory layout for prefill operations and formally enables the SFA DCP capability for A5 hardware.

Highlights

  • Sparse Flash Attention (SFA) DCP Fixes: Corrected LSE computation by bounding it to the valid index prefix and ensured proper initialization for skipped queries to prevent incorrect normalization.
  • Memory and Layout Improvements: Ensured prefill KV and RoPE views are contiguous after packed DCP gathering and initialized n2Size to improve operator stability.
  • Platform Capability Updates: Enabled SFA_DCP_REPLICATED_INDEXER on A5 hardware while maintaining existing platform capability checks.
  • Testing and Regression: Added comprehensive regression tests for A5 SFA kernels and prefill contiguous cache handling to ensure numerical correctness.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Feature] Support SFA DCP with LSE and padded indices

Suggested PR Summary:

### What this PR does / why we need it?
This PR enhances the Sparse Flash Attention (SFA) MLA kernel to support returning Softmax LSE when using padded indices under DCP. It introduces a binary search to find the end of valid sparse indices and ensures proper initialization of outputs and LSE values (setting LSE to -inf) for queries with no selected keys. Additionally, it ensures gathered KV cache tensors are contiguous after splitting and registers the `SFA_DCP_REPLICATED_INDEXER` hardware capability.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
- Added `test_sparse_flash_attention_padded_indices_lse` to test SFA with padded indices and LSE.
- Added `test_sfa_dcp_prefill_passes_contiguous_gathered_cache` to verify contiguous gathered cache.

@paddy-admin paddy-admin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the new .contiguous() shouldn't impact perf for A3 since it's overlapping with indexer processing, still some headroom left.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Signed-off-by: chengruiqi (C) <c00913489@china.huawei.com>
@recky-c
recky-c force-pushed the codex/fix-a5-sfa-dcp-lse branch from 830afad to 60c6c32 Compare September 18, 2026 00:33
chengruiqi (C) added 2 commits September 18, 2026 11:48
Signed-off-by: chengruiqi (C) <c00913489@china.huawei.com>
Signed-off-by: chengruiqi (C) <c00913489@china.huawei.com>

@ZT-AIA ZT-AIA left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Make this modification and subsequently add it to the operator description document.

@weiguihua2
weiguihua2 merged commit aff1b74 into vllm-project:main Sep 18, 2026
31 checks passed
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 18, 2026
Merging main kept the old SFA_DCP_REPLICATED_INDEXER name in platform.py
while hardware_profile.py already uses SFA_C8_DCP_REPLICATED_INDEXER from
vllm-project#16656, which made pre-commit mypy fail.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 18, 2026
Rebase the PR onto latest main as a linear history so CI's
`git rebase $BASE_SHA` no longer replays old commits onto the
vllm-project#16832 tree that deleted mrv2_utils.py.

Keep default-V2 selection, the Ascend feature blacklist, and the
SFA C8 DCP hardware capability check from vllm-project#16656.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 19, 2026
Rebase the PR onto latest main as a linear history so CI's
`git rebase $BASE_SHA` no longer replays old commits onto the

Keep default-V2 selection, the Ascend feature blacklist, and the
SFA C8 DCP hardware capability check from vllm-project#16656.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 19, 2026
Rebase the PR onto latest main as a linear history so CI's
`git rebase $BASE_SHA` no longer replays old commits onto the

Keep default-V2 selection, the Ascend feature blacklist, and the
SFA C8 DCP hardware capability check from vllm-project#16656.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 20, 2026
Rebase the PR onto latest main as a linear history so CI's
`git rebase $BASE_SHA` no longer replays old commits onto the

Keep default-V2 selection, the Ascend feature blacklist, and the
SFA C8 DCP hardware capability check from vllm-project#16656.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
yjyang62 added a commit to yjyang62/vllm-ascend that referenced this pull request Sep 20, 2026
Rebase the PR onto latest main as a linear history so CI's
`git rebase $BASE_SHA` no longer replays old commits onto the

Keep default-V2 selection, the Ascend feature blacklist, and the
SFA C8 DCP hardware capability check from vllm-project#16656.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
yjyang62 added a commit to yjyang62/vllm-ascend that referenced this pull request Sep 20, 2026
Rebase the PR onto latest main as a linear history so CI's
`git rebase $BASE_SHA` no longer replays old commits onto the

Keep default-V2 selection, the Ascend feature blacklist, and the
SFA C8 DCP hardware capability check from vllm-project#16656.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
yjyang62 added a commit to yjyang62/vllm-ascend that referenced this pull request Sep 21, 2026
Rebase the PR onto latest main as a linear history so CI's
`git rebase $BASE_SHA` no longer replays old commits onto the

Keep default-V2 selection, the Ascend feature blacklist, and the
SFA C8 DCP hardware capability check from vllm-project#16656.

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants