[Feature][Model] Integrate Triton KeyPool indexing for GLM-5.3-Flash - #16253
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request integrates the GLM-5.3-Flash pooled sparse attention mechanism into the vLLM Ascend backend. By leveraging Triton operators for KeyPool compression and routing through the shared SFA infrastructure, the changes enable efficient sparse attention on supported hardware (A2/A3/A5). The implementation simplifies cache management by utilizing complete physical pages and improves routing logic to ensure compatibility with existing serving constraints. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Attention][Feature] Implement GLM-Next KPool sparse indexer backend and Triton orchestrationSuggested PR Summary:
### What this PR does / why we need it?
This PR implements the model-side backend and Triton orchestration for the GLM-Next pooled sparse indexer (KPool) on Ascend NPU. It replaces the previous placeholder implementation that raised `NotImplementedError` with a fully functional Triton-based indexer.
Key changes include:
- Implementing `Glm5NextKPoolIndexerBackend` to handle the unified indexer API.
- Rewriting `SparseAttnIndexerKpool` to orchestrate Triton kernels (`glm5_next_kpool_state_compress_and_write_cache_triton` and `glm5_next_lightning_indexer_triton`).
- Adding `append_causal_tail` to handle unpooled tokens for CANN SFA compatibility.
- Updating platform routing to support KPool sparse attention on Ascend A2, A3, or A5, while explicitly rejecting context parallelism.
- Removing obsolete helper functions like `select_indexer_block_size` and the unused `fwht128_quant_fp8` module.
- Adding comprehensive unit and end-to-end tests to verify indexer pages, causal tail, and backend routing.
### Does this PR introduce _any_ user-facing change?
No, this is an internal backend implementation for GLM-Next sparse attention indexing on Ascend NPUs.
### How was this patch tested?
Tested with new unit and end-to-end tests:
- `test_glm5next_indexer_pages.py`
- `test_glm5next_kpool_tail.py`
- `test_glm5next_kpool_model_backend.py`
- `test_glm5next_sfa_routing.py`
***
*Note: No review comments were provided for this pull request, so there is no additional feedback to address.*edf69ed to
c992386
Compare
ecab1d6 to
c6761f8
Compare
… backend Connect Triton pooled indexing to the shared SFA path. Keep the indexer backend and its hardware/context-parallel capability checks in attention, preserving physical cache pages, compressor state and causal-tail visibility. Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
c6761f8 to
0bff0eb
Compare
pisceskkk
left a comment
There was a problem hiding this comment.
I noticed that the vllm_ascend/attention directory is currently taking on too many responsibilities. The newly added indexer_kpool_backend.py does not seem substantial enough to warrant a separate file, so we can temporarily merge it into indexer_kpool.py.
Keep KPool cache metadata and model execution in one attention module. Update model and test imports and defer KPool operator loading until backend initialization. Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
Use the self-contained KPool execution module without inheriting the generic SFA indexer. This avoids the DeviceOperator import cycle during metadata-only startup. Add regression coverage with execution dependencies unavailable. Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
Sounds good. These two files have been merged. |
### What this PR does / why we need it? Refs #15665 Extend the shared SFA backend to execute sparse MLA without RoPE. Add NoPE preprocessing and forward execution, physical-page cache addressing, graph-stable metadata buffers, and indexer-provided visible top-k lengths. Keep cache ownership with the indexer and mask unwritten graph-padding rows before output projections. Reuse the existing hardware-specific operators: - A2/A3: `npu_sparse_flash_attention` with NoPE support from #16241, now merged. - A5: the DeepSeek-V4 sparse MLA interface in `vllm_ascend/attention/sparse_flash_mla.py`, backed by the CANN `cann_ops_transformer` operators. This PR contains shared attention integration. GLM KeyPool model routing is handled separately in #16253. ### Does this PR introduce _any_ user-facing change? Yes. The shared SFA backend supports NoPE attention on A2/A3 and A5 when the corresponding operators are available. Existing RoPE behavior and cache allocation/grouping are preserved. ### How was this patch tested? - The following files passed in a CPU integration run with 319 total passing tests, covering metadata/addressing, operator dispatch, NoPE forward behavior, and existing SFA regressions: ```bash pytest -q \ tests/ut/attention/test_sfa_nope_metadata.py \ tests/ut/attention/test_sfa_nope_forward.py \ tests/ut/attention/test_sfa_v1.py ``` - A focused rerun of the same three files on A3 passed 60 tests. - GitHub CI passed pre-commit, CPU unit tests, selected device tests, and the CI gate on the current head. The focused unit tests use mocked operator calls and do not establish native numerical accuracy. A5 native operator validation remains outstanding. - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
…hado/vllm-ascend into main_fix_mrv2_eagle3_mamba * 'main_fix_mrv2_eagle3_mamba' of https://github.com/windshado/vllm-ascend: (42 commits) Update vllm_ascend/worker/v2/model_states/mamba_hybrid.py [Feature][Kimi K3 DSPark] Enable TP for context_proj (vllm-project#16344) [BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings (vllm-project#16300) [Feature][Model] Integrate Triton KeyPool indexing for GLM-5.3-Flash (vllm-project#16253) [BugFix][Offloader] Re-bind params to NZ static buffers after npu_format_cast (vllm-project#15415) [Feature][Model] Integrate AscendC KDA and causal convolution for GLM-5.3-Flash (vllm-project#16251) [Performance][Communicator] Replace per-layer F.pad with cat of a persistent zero block in MoE prepare (vllm-project#16343) [Feature][Operator] Add DeepSeek V4.1 sparse attention operators (vllm-project#16422) [Doc][Misc] Document batch invariance scheduling limitations (vllm-project#16232) [CI][MRV2] Enable mrv2 dspark e2e test (vllm-project#16319) [BugFix] Precast MoE gate weight_fp32 to avoid aclop Cast (vllm-project#16189) [Feature][MRV2][310P] MRv2 adapting MTP on the 310P for Qwen3.5 (vllm-project#16043) [Revert] Revert "[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409) [Feature][Ops] Add Triton KeyPool compression and pooled indexing (vllm-project#16243) [Feature][Attention] Support NoPE in the shared SFA backend (vllm-project#16252) [Performance][Model] Reuse fused mHC operators for GLM-5.3-Flash (vllm-project#16321) [Feature][Model] Enable MiniMax-M3 FP8 MSA index score on A5 (vllm-project#15918) [Performance][KDA] Reduce preprocessing copies and redundant output masks (vllm-project#16067) [Feature][Model][MTP] Support speculative decoding for GLM-5.3-Flash (vllm-project#16214) [BugFix][Model] Skip unused hash-router bias when loading DeepSeek-V4 weights (vllm-project#16259) ...
…ject#16252) ### What this PR does / why we need it? Refs vllm-project#15665 Extend the shared SFA backend to execute sparse MLA without RoPE. Add NoPE preprocessing and forward execution, physical-page cache addressing, graph-stable metadata buffers, and indexer-provided visible top-k lengths. Keep cache ownership with the indexer and mask unwritten graph-padding rows before output projections. Reuse the existing hardware-specific operators: - A2/A3: `npu_sparse_flash_attention` with NoPE support from vllm-project#16241, now merged. - A5: the DeepSeek-V4 sparse MLA interface in `vllm_ascend/attention/sparse_flash_mla.py`, backed by the CANN `cann_ops_transformer` operators. This PR contains shared attention integration. GLM KeyPool model routing is handled separately in vllm-project#16253. ### Does this PR introduce _any_ user-facing change? Yes. The shared SFA backend supports NoPE attention on A2/A3 and A5 when the corresponding operators are available. Existing RoPE behavior and cache allocation/grouping are preserved. ### How was this patch tested? - The following files passed in a CPU integration run with 319 total passing tests, covering metadata/addressing, operator dispatch, NoPE forward behavior, and existing SFA regressions: ```bash pytest -q \ tests/ut/attention/test_sfa_nope_metadata.py \ tests/ut/attention/test_sfa_nope_forward.py \ tests/ut/attention/test_sfa_v1.py ``` - A focused rerun of the same three files on A3 passed 60 tests. - GitHub CI passed pre-commit, CPU unit tests, selected device tests, and the CI gate on the current head. The focused unit tests use mocked operator calls and do not establish native numerical accuracy. A5 native operator validation remains outstanding. - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com> Signed-off-by: tianming2009 <13246728590@163.com>
…llm-project#16253) ### What this PR does / why we need it? Refs vllm-project#15665 Connect the GLM-5.3-Flash pooled indexer to `IndexerWrapper` and the shared NoPE SFA backend. Keep cache metadata and the model-side indexer backend together in `vllm_ascend/attention/indexer_kpool.py`. Load KPool execution dependencies when the backend is instantiated to avoid circular imports during metadata-only startup. Keep hardware and PCP/DCP capability checks in backend initialization. Pass complete compressed physical pages and compressor-state metadata to the Triton operators while preserving causal-tail visibility and the cache allocation/grouping contract from vllm-project#15913. Dependencies: - vllm-project#16243 provides the Triton operators and indexer backend-selection hook. - vllm-project#16252 provides shared NoPE SFA execution and metadata. Prerequisite commits are combined only for integration validation and are excluded from this branch. ### Does this PR introduce _any_ user-facing change? Yes. GLM-5.3-Flash routes pooled sparse attention through the shared SFA backend and Triton indexer. Unsupported hardware and PCP/DCP configurations are rejected when the backend is initialized. ### How was this patch tested? At PR revision `1ff19ad48`, combined with its integration prerequisites: - Fresh-process imports of the consolidated metadata module and model backend factory passed. Regression tests cover importing metadata without generic SFA indexing or KPool execution dependencies. - 120 unit tests passed, 1 skipped, covering `test_glm5next_kpool_model_backend.py`, `test_glm5next_kv_cache.py`, `test_glm5next_sfa_routing.py`, `test_glm5next_indexer_backend.py`, `test_platform.py`, and `attention/test_indexer.py`. - Six native NPU comparisons against the pre-consolidation backend produced identical indices and compressed/state caches: prefill tail, decode tail, pool-boundary decode, prefill top-k selection, FULL-mode padding, and cache-only updates. - GitHub CI passed pre-commit, CPU unit tests, selected device tests, and the CI gate. Earlier full-model validation at PR revision `0bff0eb46`, combined with integration prerequisites, passed 320 unit tests (1 skipped) and 8/8 GLM-5.3-Flash-w8a8 inference cases with TP8/EP8, FULL_DECODE_ONLY, and MTP=3. Those results apply to the pre-consolidation revision; the six native comparisons above do not constitute a new full-model or ACLGraph capture/replay run. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com> Signed-off-by: tianming2009 <13246728590@163.com>
…ject#16252) ### What this PR does / why we need it? Refs vllm-project#15665 Extend the shared SFA backend to execute sparse MLA without RoPE. Add NoPE preprocessing and forward execution, physical-page cache addressing, graph-stable metadata buffers, and indexer-provided visible top-k lengths. Keep cache ownership with the indexer and mask unwritten graph-padding rows before output projections. Reuse the existing hardware-specific operators: - A2/A3: `npu_sparse_flash_attention` with NoPE support from vllm-project#16241, now merged. - A5: the DeepSeek-V4 sparse MLA interface in `vllm_ascend/attention/sparse_flash_mla.py`, backed by the CANN `cann_ops_transformer` operators. This PR contains shared attention integration. GLM KeyPool model routing is handled separately in vllm-project#16253. ### Does this PR introduce _any_ user-facing change? Yes. The shared SFA backend supports NoPE attention on A2/A3 and A5 when the corresponding operators are available. Existing RoPE behavior and cache allocation/grouping are preserved. ### How was this patch tested? - The following files passed in a CPU integration run with 319 total passing tests, covering metadata/addressing, operator dispatch, NoPE forward behavior, and existing SFA regressions: ```bash pytest -q \ tests/ut/attention/test_sfa_nope_metadata.py \ tests/ut/attention/test_sfa_nope_forward.py \ tests/ut/attention/test_sfa_v1.py ``` - A focused rerun of the same three files on A3 passed 60 tests. - GitHub CI passed pre-commit, CPU unit tests, selected device tests, and the CI gate on the current head. The focused unit tests use mocked operator calls and do not establish native numerical accuracy. A5 native operator validation remains outstanding. - vLLM main: vllm-project/vllm@b2f6858 Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com> Signed-off-by: like-0517 <ithwlike@126.com>
…llm-project#16253) ### What this PR does / why we need it? Refs vllm-project#15665 Connect the GLM-5.3-Flash pooled indexer to `IndexerWrapper` and the shared NoPE SFA backend. Keep cache metadata and the model-side indexer backend together in `vllm_ascend/attention/indexer_kpool.py`. Load KPool execution dependencies when the backend is instantiated to avoid circular imports during metadata-only startup. Keep hardware and PCP/DCP capability checks in backend initialization. Pass complete compressed physical pages and compressor-state metadata to the Triton operators while preserving causal-tail visibility and the cache allocation/grouping contract from vllm-project#15913. Dependencies: - vllm-project#16243 provides the Triton operators and indexer backend-selection hook. - vllm-project#16252 provides shared NoPE SFA execution and metadata. Prerequisite commits are combined only for integration validation and are excluded from this branch. ### Does this PR introduce _any_ user-facing change? Yes. GLM-5.3-Flash routes pooled sparse attention through the shared SFA backend and Triton indexer. Unsupported hardware and PCP/DCP configurations are rejected when the backend is initialized. ### How was this patch tested? At PR revision `1ff19ad48`, combined with its integration prerequisites: - Fresh-process imports of the consolidated metadata module and model backend factory passed. Regression tests cover importing metadata without generic SFA indexing or KPool execution dependencies. - 120 unit tests passed, 1 skipped, covering `test_glm5next_kpool_model_backend.py`, `test_glm5next_kv_cache.py`, `test_glm5next_sfa_routing.py`, `test_glm5next_indexer_backend.py`, `test_platform.py`, and `attention/test_indexer.py`. - Six native NPU comparisons against the pre-consolidation backend produced identical indices and compressed/state caches: prefill tail, decode tail, pool-boundary decode, prefill top-k selection, FULL-mode padding, and cache-only updates. - GitHub CI passed pre-commit, CPU unit tests, selected device tests, and the CI gate. Earlier full-model validation at PR revision `0bff0eb46`, combined with integration prerequisites, passed 320 unit tests (1 skipped) and 8/8 GLM-5.3-Flash-w8a8 inference cases with TP8/EP8, FULL_DECODE_ONLY, and MTP=3. Those results apply to the pre-consolidation revision; the six native comparisons above do not constitute a new full-model or ACLGraph capture/replay run. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com> Signed-off-by: like-0517 <ithwlike@126.com>
What this PR does / why we need it?
Refs #15665
Connect the GLM-5.3-Flash pooled indexer to
IndexerWrapperand the shared NoPE SFA backend. Keep cache metadata and the model-side indexer backend together invllm_ascend/attention/indexer_kpool.py. Load KPool execution dependencies when the backend is instantiated to avoid circular imports during metadata-only startup.Keep hardware and PCP/DCP capability checks in backend initialization. Pass complete compressed physical pages and compressor-state metadata to the Triton operators while preserving causal-tail visibility and the cache allocation/grouping contract from #15913.
Dependencies:
Prerequisite commits are combined only for integration validation and are excluded from this branch.
Does this PR introduce any user-facing change?
Yes. GLM-5.3-Flash routes pooled sparse attention through the shared SFA backend and Triton indexer. Unsupported hardware and PCP/DCP configurations are rejected when the backend is initialized.
How was this patch tested?
At PR revision
1ff19ad48, combined with its integration prerequisites:test_glm5next_kpool_model_backend.py,test_glm5next_kv_cache.py,test_glm5next_sfa_routing.py,test_glm5next_indexer_backend.py,test_platform.py, andattention/test_indexer.py.Earlier full-model validation at PR revision
0bff0eb46, combined with integration prerequisites, passed 320 unit tests (1 skipped) and 8/8 GLM-5.3-Flash-w8a8 inference cases with TP8/EP8, FULL_DECODE_ONLY, and MTP=3. Those results apply to the pre-consolidation revision; the six native comparisons above do not constitute a new full-model or ACLGraph capture/replay run.