Skip to content

[Feature][Model] Integrate Triton KeyPool indexing for GLM-5.3-Flash - #16253

Merged
ZT-AIA merged 3 commits into
vllm-project:mainfrom
lijiahang226:glm53-v5-kpool-model
Sep 14, 2026
Merged

ZT-AIA merged 3 commits into
vllm-project:mainfrom
lijiahang226:glm53-v5-kpool-model

Conversation

@lijiahang226

@lijiahang226 lijiahang226 commented Sep 10, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Refs #15665

Connect the GLM-5.3-Flash pooled indexer to IndexerWrapper and the shared NoPE SFA backend. Keep cache metadata and the model-side indexer backend together in vllm_ascend/attention/indexer_kpool.py. Load KPool execution dependencies when the backend is instantiated to avoid circular imports during metadata-only startup.

Keep hardware and PCP/DCP capability checks in backend initialization. Pass complete compressed physical pages and compressor-state metadata to the Triton operators while preserving causal-tail visibility and the cache allocation/grouping contract from #15913.

Dependencies:

Prerequisite commits are combined only for integration validation and are excluded from this branch.

Does this PR introduce any user-facing change?

Yes. GLM-5.3-Flash routes pooled sparse attention through the shared SFA backend and Triton indexer. Unsupported hardware and PCP/DCP configurations are rejected when the backend is initialized.

How was this patch tested?

At PR revision 1ff19ad48, combined with its integration prerequisites:

  • Fresh-process imports of the consolidated metadata module and model backend factory passed. Regression tests cover importing metadata without generic SFA indexing or KPool execution dependencies.
  • 120 unit tests passed, 1 skipped, covering test_glm5next_kpool_model_backend.py, test_glm5next_kv_cache.py, test_glm5next_sfa_routing.py, test_glm5next_indexer_backend.py, test_platform.py, and attention/test_indexer.py.
  • Six native NPU comparisons against the pre-consolidation backend produced identical indices and compressed/state caches: prefill tail, decode tail, pool-boundary decode, prefill top-k selection, FULL-mode padding, and cache-only updates.
  • GitHub CI passed pre-commit, CPU unit tests, selected device tests, and the CI gate.

Earlier full-model validation at PR revision 0bff0eb46, combined with integration prerequisites, passed 320 unit tests (1 skipped) and 8/8 GLM-5.3-Flash-w8a8 inference cases with TP8/EP8, FULL_DECODE_ONLY, and MTP=3. Those results apply to the pre-consolidation revision; the six native comparisons above do not constitute a new full-model or ACLGraph capture/replay run.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request integrates the GLM-5.3-Flash pooled sparse attention mechanism into the vLLM Ascend backend. By leveraging Triton operators for KeyPool compression and routing through the shared SFA infrastructure, the changes enable efficient sparse attention on supported hardware (A2/A3/A5). The implementation simplifies cache management by utilizing complete physical pages and improves routing logic to ensure compatibility with existing serving constraints.

Highlights

  • GLM-5.3-Flash KPool Integration: Integrated the GLM-5.3-Flash pooled indexer with the IndexerWrapper and shared NoPE SFA backend, enabling Triton-based persistent KeyPool compression and pooled-key selection.
  • Indexer Backend Refactoring: Moved model-specific projections and cache binding into the indexer backend, removing the CANN pool-indexer block-splitting path in favor of complete physical page addressing.
  • Routing and Validation: Added explicit validation to reject unsupported PCP/DCP configurations and implemented new regression tests for indexer page addressing and causal tail handling.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Feature] Implement GLM-Next KPool sparse indexer backend and Triton orchestration

Suggested PR Summary:

### What this PR does / why we need it?

This PR implements the model-side backend and Triton orchestration for the GLM-Next pooled sparse indexer (KPool) on Ascend NPU. It replaces the previous placeholder implementation that raised `NotImplementedError` with a fully functional Triton-based indexer.

Key changes include:
- Implementing `Glm5NextKPoolIndexerBackend` to handle the unified indexer API.
- Rewriting `SparseAttnIndexerKpool` to orchestrate Triton kernels (`glm5_next_kpool_state_compress_and_write_cache_triton` and `glm5_next_lightning_indexer_triton`).
- Adding `append_causal_tail` to handle unpooled tokens for CANN SFA compatibility.
- Updating platform routing to support KPool sparse attention on Ascend A2, A3, or A5, while explicitly rejecting context parallelism.
- Removing obsolete helper functions like `select_indexer_block_size` and the unused `fwht128_quant_fp8` module.
- Adding comprehensive unit and end-to-end tests to verify indexer pages, causal tail, and backend routing.

### Does this PR introduce _any_ user-facing change?

No, this is an internal backend implementation for GLM-Next sparse attention indexing on Ascend NPUs.

### How was this patch tested?

Tested with new unit and end-to-end tests:
- `test_glm5next_indexer_pages.py`
- `test_glm5next_kpool_tail.py`
- `test_glm5next_kpool_model_backend.py`
- `test_glm5next_sfa_routing.py`

***

*Note: No review comments were provided for this pull request, so there is no additional feedback to address.*

@lijiahang226
lijiahang226 force-pushed the glm53-v5-kpool-model branch 4 times, most recently from edf69ed to c992386 Compare September 10, 2026 14:42
@lijiahang226
lijiahang226 force-pushed the glm53-v5-kpool-model branch 2 times, most recently from ecab1d6 to c6761f8 Compare September 10, 2026 17:46
@lijiahang226 lijiahang226 added the ready-precise run selected e2e test for pr label Sep 10, 2026 — with ChatGPT Codex Connector
@lijiahang226
lijiahang226 marked this pull request as ready for review September 11, 2026 00:37
Comment thread vllm_ascend/attention/indexer_kpool_backend.py Outdated
Comment thread vllm_ascend/platform.py Outdated
… backend

Connect Triton pooled indexing to the shared SFA path. Keep the indexer backend and its hardware/context-parallel capability checks in attention, preserving physical cache pages, compressor state and causal-tail visibility.

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>

@pisceskkk pisceskkk left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I noticed that the vllm_ascend/attention directory is currently taking on too many responsibilities. The newly added indexer_kpool_backend.py does not seem substantial enough to warrant a separate file, so we can temporarily merge it into indexer_kpool.py.

Keep KPool cache metadata and model execution in one attention module. Update model and test imports and defer KPool operator loading until backend initialization.

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
Use the self-contained KPool execution module without inheriting the generic SFA indexer. This avoids the DeviceOperator import cycle during metadata-only startup. Add regression coverage with execution dependencies unavailable.

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
@lijiahang226

Copy link
Copy Markdown
Collaborator Author

I noticed that the vllm_ascend/attention directory is currently taking on too many responsibilities. The newly added indexer_kpool_backend.py does not seem substantial enough to warrant a separate file, so we can temporarily merge it into indexer_kpool.py.

Sounds good. These two files have been merged.

linfeng-yuan pushed a commit that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

Refs #15665

Extend the shared SFA backend to execute sparse MLA without RoPE. Add
NoPE preprocessing and forward execution, physical-page cache
addressing, graph-stable metadata buffers, and indexer-provided visible
top-k lengths. Keep cache ownership with the indexer and mask unwritten
graph-padding rows before output projections.

Reuse the existing hardware-specific operators:

- A2/A3: `npu_sparse_flash_attention` with NoPE support from #16241, now
merged.
- A5: the DeepSeek-V4 sparse MLA interface in
`vllm_ascend/attention/sparse_flash_mla.py`, backed by the CANN
`cann_ops_transformer` operators.

This PR contains shared attention integration. GLM KeyPool model routing
is handled separately in #16253.

### Does this PR introduce _any_ user-facing change?

Yes. The shared SFA backend supports NoPE attention on A2/A3 and A5 when
the corresponding operators are available. Existing RoPE behavior and
cache allocation/grouping are preserved.

### How was this patch tested?

- The following files passed in a CPU integration run with 319 total
passing tests, covering metadata/addressing, operator dispatch, NoPE
forward behavior, and existing SFA regressions:

  ```bash
  pytest -q \
    tests/ut/attention/test_sfa_nope_metadata.py \
    tests/ut/attention/test_sfa_nope_forward.py \
    tests/ut/attention/test_sfa_v1.py
  ```

- A focused rerun of the same three files on A3 passed 60 tests.
- GitHub CI passed pre-commit, CPU unit tests, selected device tests,
and the CI gate on the current head.

The focused unit tests use mocked operator calls and do not establish
native numerical accuracy. A5 native operator validation remains
outstanding.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
@ZT-AIA
ZT-AIA merged commit ae67c34 into vllm-project:main Sep 14, 2026
30 checks passed
windshado added a commit to windshado/vllm-ascend that referenced this pull request Sep 14, 2026
…hado/vllm-ascend into main_fix_mrv2_eagle3_mamba

* 'main_fix_mrv2_eagle3_mamba' of https://github.com/windshado/vllm-ascend: (42 commits)
  Update vllm_ascend/worker/v2/model_states/mamba_hybrid.py
  [Feature][Kimi K3 DSPark] Enable TP for context_proj (vllm-project#16344)
  [BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings (vllm-project#16300)
  [Feature][Model] Integrate Triton KeyPool indexing for GLM-5.3-Flash (vllm-project#16253)
  [BugFix][Offloader] Re-bind params to NZ static buffers after npu_format_cast (vllm-project#15415)
  [Feature][Model] Integrate AscendC KDA and causal convolution for GLM-5.3-Flash (vllm-project#16251)
  [Performance][Communicator] Replace per-layer F.pad with cat of a persistent zero block in MoE prepare (vllm-project#16343)
  [Feature][Operator] Add DeepSeek V4.1 sparse attention operators (vllm-project#16422)
  [Doc][Misc] Document batch invariance scheduling limitations (vllm-project#16232)
  [CI][MRV2] Enable mrv2 dspark e2e test (vllm-project#16319)
  [BugFix] Precast MoE gate weight_fp32 to avoid aclop Cast (vllm-project#16189)
  [Feature][MRV2][310P] MRv2 adapting MTP on the 310P for Qwen3.5 (vllm-project#16043)
  [Revert] Revert "[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409)
  [Feature][Ops] Add Triton KeyPool compression and pooled indexing (vllm-project#16243)
  [Feature][Attention] Support NoPE in the shared SFA backend (vllm-project#16252)
  [Performance][Model] Reuse fused mHC operators for GLM-5.3-Flash (vllm-project#16321)
  [Feature][Model] Enable MiniMax-M3 FP8 MSA index score on A5 (vllm-project#15918)
  [Performance][KDA] Reduce preprocessing copies and redundant output masks (vllm-project#16067)
  [Feature][Model][MTP] Support speculative decoding for GLM-5.3-Flash (vllm-project#16214)
  [BugFix][Model] Skip unused hash-router bias when loading DeepSeek-V4 weights (vllm-project#16259)
  ...
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
…ject#16252)

### What this PR does / why we need it?

Refs vllm-project#15665

Extend the shared SFA backend to execute sparse MLA without RoPE. Add
NoPE preprocessing and forward execution, physical-page cache
addressing, graph-stable metadata buffers, and indexer-provided visible
top-k lengths. Keep cache ownership with the indexer and mask unwritten
graph-padding rows before output projections.

Reuse the existing hardware-specific operators:

- A2/A3: `npu_sparse_flash_attention` with NoPE support from vllm-project#16241, now
merged.
- A5: the DeepSeek-V4 sparse MLA interface in
`vllm_ascend/attention/sparse_flash_mla.py`, backed by the CANN
`cann_ops_transformer` operators.

This PR contains shared attention integration. GLM KeyPool model routing
is handled separately in vllm-project#16253.

### Does this PR introduce _any_ user-facing change?

Yes. The shared SFA backend supports NoPE attention on A2/A3 and A5 when
the corresponding operators are available. Existing RoPE behavior and
cache allocation/grouping are preserved.

### How was this patch tested?

- The following files passed in a CPU integration run with 319 total
passing tests, covering metadata/addressing, operator dispatch, NoPE
forward behavior, and existing SFA regressions:

  ```bash
  pytest -q \
    tests/ut/attention/test_sfa_nope_metadata.py \
    tests/ut/attention/test_sfa_nope_forward.py \
    tests/ut/attention/test_sfa_v1.py
  ```

- A focused rerun of the same three files on A3 passed 60 tests.
- GitHub CI passed pre-commit, CPU unit tests, selected device tests,
and the CI gate on the current head.

The focused unit tests use mocked operator calls and do not establish
native numerical accuracy. A5 native operator validation remains
outstanding.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
Signed-off-by: tianming2009 <13246728590@163.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
…llm-project#16253)

### What this PR does / why we need it?

Refs vllm-project#15665

Connect the GLM-5.3-Flash pooled indexer to `IndexerWrapper` and the
shared NoPE SFA backend. Keep cache metadata and the model-side indexer
backend together in `vllm_ascend/attention/indexer_kpool.py`. Load KPool
execution dependencies when the backend is instantiated to avoid
circular imports during metadata-only startup.

Keep hardware and PCP/DCP capability checks in backend initialization.
Pass complete compressed physical pages and compressor-state metadata to
the Triton operators while preserving causal-tail visibility and the
cache allocation/grouping contract from vllm-project#15913.

Dependencies:

- vllm-project#16243 provides the Triton operators and indexer backend-selection
hook.
- vllm-project#16252 provides shared NoPE SFA execution and metadata.

Prerequisite commits are combined only for integration validation and
are excluded from this branch.

### Does this PR introduce _any_ user-facing change?

Yes. GLM-5.3-Flash routes pooled sparse attention through the shared SFA
backend and Triton indexer. Unsupported hardware and PCP/DCP
configurations are rejected when the backend is initialized.

### How was this patch tested?

At PR revision `1ff19ad48`, combined with its integration prerequisites:

- Fresh-process imports of the consolidated metadata module and model
backend factory passed. Regression tests cover importing metadata
without generic SFA indexing or KPool execution dependencies.
- 120 unit tests passed, 1 skipped, covering
`test_glm5next_kpool_model_backend.py`, `test_glm5next_kv_cache.py`,
`test_glm5next_sfa_routing.py`, `test_glm5next_indexer_backend.py`,
`test_platform.py`, and `attention/test_indexer.py`.
- Six native NPU comparisons against the pre-consolidation backend
produced identical indices and compressed/state caches: prefill tail,
decode tail, pool-boundary decode, prefill top-k selection, FULL-mode
padding, and cache-only updates.
- GitHub CI passed pre-commit, CPU unit tests, selected device tests,
and the CI gate.

Earlier full-model validation at PR revision `0bff0eb46`, combined with
integration prerequisites, passed 320 unit tests (1 skipped) and 8/8
GLM-5.3-Flash-w8a8 inference cases with TP8/EP8, FULL_DECODE_ONLY, and
MTP=3. Those results apply to the pre-consolidation revision; the six
native comparisons above do not constitute a new full-model or ACLGraph
capture/replay run.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…ject#16252)

### What this PR does / why we need it?

Refs vllm-project#15665

Extend the shared SFA backend to execute sparse MLA without RoPE. Add
NoPE preprocessing and forward execution, physical-page cache
addressing, graph-stable metadata buffers, and indexer-provided visible
top-k lengths. Keep cache ownership with the indexer and mask unwritten
graph-padding rows before output projections.

Reuse the existing hardware-specific operators:

- A2/A3: `npu_sparse_flash_attention` with NoPE support from vllm-project#16241, now
merged.
- A5: the DeepSeek-V4 sparse MLA interface in
`vllm_ascend/attention/sparse_flash_mla.py`, backed by the CANN
`cann_ops_transformer` operators.

This PR contains shared attention integration. GLM KeyPool model routing
is handled separately in vllm-project#16253.

### Does this PR introduce _any_ user-facing change?

Yes. The shared SFA backend supports NoPE attention on A2/A3 and A5 when
the corresponding operators are available. Existing RoPE behavior and
cache allocation/grouping are preserved.

### How was this patch tested?

- The following files passed in a CPU integration run with 319 total
passing tests, covering metadata/addressing, operator dispatch, NoPE
forward behavior, and existing SFA regressions:

  ```bash
  pytest -q \
    tests/ut/attention/test_sfa_nope_metadata.py \
    tests/ut/attention/test_sfa_nope_forward.py \
    tests/ut/attention/test_sfa_v1.py
  ```

- A focused rerun of the same three files on A3 passed 60 tests.
- GitHub CI passed pre-commit, CPU unit tests, selected device tests,
and the CI gate on the current head.

The focused unit tests use mocked operator calls and do not establish
native numerical accuracy. A5 native operator validation remains
outstanding.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
Signed-off-by: like-0517 <ithwlike@126.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…llm-project#16253)

### What this PR does / why we need it?

Refs vllm-project#15665

Connect the GLM-5.3-Flash pooled indexer to `IndexerWrapper` and the
shared NoPE SFA backend. Keep cache metadata and the model-side indexer
backend together in `vllm_ascend/attention/indexer_kpool.py`. Load KPool
execution dependencies when the backend is instantiated to avoid
circular imports during metadata-only startup.

Keep hardware and PCP/DCP capability checks in backend initialization.
Pass complete compressed physical pages and compressor-state metadata to
the Triton operators while preserving causal-tail visibility and the
cache allocation/grouping contract from vllm-project#15913.

Dependencies:

- vllm-project#16243 provides the Triton operators and indexer backend-selection
hook.
- vllm-project#16252 provides shared NoPE SFA execution and metadata.

Prerequisite commits are combined only for integration validation and
are excluded from this branch.

### Does this PR introduce _any_ user-facing change?

Yes. GLM-5.3-Flash routes pooled sparse attention through the shared SFA
backend and Triton indexer. Unsupported hardware and PCP/DCP
configurations are rejected when the backend is initialized.

### How was this patch tested?

At PR revision `1ff19ad48`, combined with its integration prerequisites:

- Fresh-process imports of the consolidated metadata module and model
backend factory passed. Regression tests cover importing metadata
without generic SFA indexing or KPool execution dependencies.
- 120 unit tests passed, 1 skipped, covering
`test_glm5next_kpool_model_backend.py`, `test_glm5next_kv_cache.py`,
`test_glm5next_sfa_routing.py`, `test_glm5next_indexer_backend.py`,
`test_platform.py`, and `attention/test_indexer.py`.
- Six native NPU comparisons against the pre-consolidation backend
produced identical indices and compressed/state caches: prefill tail,
decode tail, pool-boundary decode, prefill top-k selection, FULL-mode
padding, and cache-only updates.
- GitHub CI passed pre-commit, CPU unit tests, selected device tests,
and the CI gate.

Earlier full-model validation at PR revision `0bff0eb46`, combined with
integration prerequisites, passed 320 unit tests (1 skipped) and 8/8
GLM-5.3-Flash-w8a8 inference cases with TP8/EP8, FULL_DECODE_ONLY, and
MTP=3. Those results apply to the pre-consolidation revision; the six
native comparisons above do not constitute a new full-model or ACLGraph
capture/replay run.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants