Repository navigation
[Feature][Model] Integrate AscendC KDA and causal convolution for GLM-5.3-Flash - #16251
Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request integrates AscendC-based KDA and causal convolution operators into the GLM-5.3-Flash model. It enhances cache management by ensuring shared physical-page strides are correctly handled for Mamba cache views, preventing potential data corruption between scheduler blocks. Additionally, it introduces robust testing coverage for operator dispatch and cache isolation to ensure stability. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Ops][Feature] Support GLM5-Next short convolution and KDA operators on AscendSuggested PR Summary:
### What this PR does / why we need it?
This PR integrates the AscendC custom operators (`npu_causal_conv1d_custom`, `recurrent_kda`, and `chunk_kda_fwd`) into the GLM5-Next model implementation for Ascend. It replaces previous Triton/fallback implementations with optimized AscendC operators for prefill, decode, and MTP verification. It also introduces a Triton-based state copy kernel for non-contiguous views and updates the model runner to support shared MLA and Mamba physical pages.
Feedback has been provided to use `tl.where` to clamp the slot index to a safe value (e.g., `0`) when inactive in the Triton state copy kernel (`_copy_conv_state`) to prevent undefined behavior or compilation errors from negative pointer arithmetic on Ascend hardware.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Tested with new end-to-end and unit tests:
- `test_glm5next_causal_conv1d.py`
- `test_glm5next_kda_contracts.py`
- `test_glm5next_kda_dispatch.py`
- `test_glm5next_pooled_cache.py`ba7edcd to
3dfb61b
Compare
37920cd to
4b11596
Compare
4b11596 to
c40b9e5
Compare
|
Current-head applicability update (2026-09-11): #16251 is now at I ran a bounded standalone validation of the exact #16251 GLM-5.3 KDA causal-conv ordinary non-spec prefill path at head Correctness:
Performance protocol: 5 warmups, 20 observations per arm, balanced AB/BA order.
On the negative-slot review comment: the reviewed #16251 wrapper computes So the source-reachability evidence in this ordinary-prefill scope did not establish a negative slot reaching the tested path, and it does not downgrade the R16e valid-row prefill result. It also does not prove the broader wrapper is safe; the suggested clamp remains a legitimate review/hardening item. Coordination note: PR #15885 also covers the ordinary non-spec GLM KDA causal-conv path through a narrower integration surface. I independently validated its actual wrapper as bitwise exact on the same official-weight family, with strong component speedups, while #16251 additionally owns GDN metadata and non-contiguous state staging/writeback. The source overlap is therefore partial rather than a clean duplicate; this note is intended to provide evidence for maintainer coordination, not to recommend which whole PR should win ownership. |
|
The current /rerun |
|
/rerun |
|
/rerun Rerun (failed jobs only):
|
|
/rerun Failed:
|
2a889f6 to
2fc40a1
Compare
|
/rerun |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
1 similar comment
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
2fc40a1 to
565419b
Compare
…olution state Integrate AscendC KDA and convolution for GLM using the same kernel helpers as Kimi K3. Preserve each model's beta processing, state updates and projections. Mask inactive convolution slots before address arithmetic. Retain cache-spec capability detection and physical-page views without changing shared cache allocation or grouping. Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com>
565419b to
6102f76
Compare
|
Current-head hardware validation for
So the current |
ZT-AIA
left a comment
There was a problem hiding this comment.
There are too many duplicate file names. I will restructure them later.
…hado/vllm-ascend into main_fix_mrv2_eagle3_mamba * 'main_fix_mrv2_eagle3_mamba' of https://github.com/windshado/vllm-ascend: (42 commits) Update vllm_ascend/worker/v2/model_states/mamba_hybrid.py [Feature][Kimi K3 DSPark] Enable TP for context_proj (vllm-project#16344) [BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings (vllm-project#16300) [Feature][Model] Integrate Triton KeyPool indexing for GLM-5.3-Flash (vllm-project#16253) [BugFix][Offloader] Re-bind params to NZ static buffers after npu_format_cast (vllm-project#15415) [Feature][Model] Integrate AscendC KDA and causal convolution for GLM-5.3-Flash (vllm-project#16251) [Performance][Communicator] Replace per-layer F.pad with cat of a persistent zero block in MoE prepare (vllm-project#16343) [Feature][Operator] Add DeepSeek V4.1 sparse attention operators (vllm-project#16422) [Doc][Misc] Document batch invariance scheduling limitations (vllm-project#16232) [CI][MRV2] Enable mrv2 dspark e2e test (vllm-project#16319) [BugFix] Precast MoE gate weight_fp32 to avoid aclop Cast (vllm-project#16189) [Feature][MRV2][310P] MRv2 adapting MTP on the 310P for Qwen3.5 (vllm-project#16043) [Revert] Revert "[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409) [Feature][Ops] Add Triton KeyPool compression and pooled indexing (vllm-project#16243) [Feature][Attention] Support NoPE in the shared SFA backend (vllm-project#16252) [Performance][Model] Reuse fused mHC operators for GLM-5.3-Flash (vllm-project#16321) [Feature][Model] Enable MiniMax-M3 FP8 MSA index score on A5 (vllm-project#15918) [Performance][KDA] Reduce preprocessing copies and redundant output masks (vllm-project#16067) [Feature][Model][MTP] Support speculative decoding for GLM-5.3-Flash (vllm-project#16214) [BugFix][Model] Skip unused hash-router bias when loading DeepSeek-V4 weights (vllm-project#16259) ...
…-5.3-Flash (vllm-project#16251) ### What this PR does / why we need it? Refs vllm-project#15665 Integrate the existing AscendC recurrent/chunk KDA and causal-convolution operators into GLM-5.3-Flash. GLM and Kimi K3 share the native KDA call helpers in `vllm_ascend/ops/kda.py`, while retaining their model-specific projections, weight loading, beta preparation, cache updates, and metadata handling. Preserve physical-page state views and cache-spec capability detection without changing the cache allocation, grouping, or shared MLA/Mamba tensor ownership from vllm-project#15913. Mask inactive convolution slots before pointer arithmetic and preserve writeback to noncontiguous state views. The native operators and GDN interfaces are already in main. Full-model validation combines the GLM Flash integration PRs and their prerequisites; prerequisite commits are excluded from this branch. ### Does this PR introduce _any_ user-facing change? Yes. GLM-5.3-Flash uses AscendC KDA and causal convolution. Unsupported operator configurations are rejected during initialization. Kimi K3 retains its existing beta and gate semantics through the shared helpers. ### How was this patch tested? A3 integration validation included PR revision `2fc40a144`: - 320 unit tests passed, 1 skipped, covering GLM/Kimi KDA, cache pages, KeyPool routing, GDN metadata, MTP, ModelSlim, and shared SFA. Updated coverage includes `tests/ut/ops/test_kimi_kda.py` and `tests/ut/models/test_glm5next_kda_contracts.py`. - Four native NPU tests passed, including `tests/e2e/nightly/single_node/ops/singlecard_ops/test_glm5next_conv_state.py` for inactive slots and strided state layouts, plus existing Kimi recurrent/prefill tests. - Twelve comparisons against the prior implementations produced exactly equal outputs and backing-state tensors for GLM/Kimi gate variants, decode, MTP verification, and 65/131-token prefill. - GLM-5.3-Flash-w8a8 passed 8/8 inference cases, including a 5264-token prompt, with TP8/EP8, FULL_DECODE_ONLY, and MTP=3. The current revision rebases on vllm-project#16067 and preserves its direct Q/K/V view handling and removal of redundant Kimi output masks. The shared recurrent helper forwards Q/K/V unchanged; GLM and Kimi regression cases assert that contiguous and separately strided views reach the native operator without Python-side materialization. Eight focused GLM/Kimi contract cases passed in an isolated CPU harness executing the actual function bodies and test cases. Ruff, formatting, `git diff --check`, and GitHub pre-commit passed. The hardware results above apply to the earlier integration; this rebased revision has not had a new hardware run. - vLLM main: vllm-project/vllm@a97dacb Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com> Signed-off-by: tianming2009 <13246728590@163.com>
…-5.3-Flash (vllm-project#16251) ### What this PR does / why we need it? Refs vllm-project#15665 Integrate the existing AscendC recurrent/chunk KDA and causal-convolution operators into GLM-5.3-Flash. GLM and Kimi K3 share the native KDA call helpers in `vllm_ascend/ops/kda.py`, while retaining their model-specific projections, weight loading, beta preparation, cache updates, and metadata handling. Preserve physical-page state views and cache-spec capability detection without changing the cache allocation, grouping, or shared MLA/Mamba tensor ownership from vllm-project#15913. Mask inactive convolution slots before pointer arithmetic and preserve writeback to noncontiguous state views. The native operators and GDN interfaces are already in main. Full-model validation combines the GLM Flash integration PRs and their prerequisites; prerequisite commits are excluded from this branch. ### Does this PR introduce _any_ user-facing change? Yes. GLM-5.3-Flash uses AscendC KDA and causal convolution. Unsupported operator configurations are rejected during initialization. Kimi K3 retains its existing beta and gate semantics through the shared helpers. ### How was this patch tested? A3 integration validation included PR revision `2fc40a144`: - 320 unit tests passed, 1 skipped, covering GLM/Kimi KDA, cache pages, KeyPool routing, GDN metadata, MTP, ModelSlim, and shared SFA. Updated coverage includes `tests/ut/ops/test_kimi_kda.py` and `tests/ut/models/test_glm5next_kda_contracts.py`. - Four native NPU tests passed, including `tests/e2e/nightly/single_node/ops/singlecard_ops/test_glm5next_conv_state.py` for inactive slots and strided state layouts, plus existing Kimi recurrent/prefill tests. - Twelve comparisons against the prior implementations produced exactly equal outputs and backing-state tensors for GLM/Kimi gate variants, decode, MTP verification, and 65/131-token prefill. - GLM-5.3-Flash-w8a8 passed 8/8 inference cases, including a 5264-token prompt, with TP8/EP8, FULL_DECODE_ONLY, and MTP=3. The current revision rebases on vllm-project#16067 and preserves its direct Q/K/V view handling and removal of redundant Kimi output masks. The shared recurrent helper forwards Q/K/V unchanged; GLM and Kimi regression cases assert that contiguous and separately strided views reach the native operator without Python-side materialization. Eight focused GLM/Kimi contract cases passed in an isolated CPU harness executing the actual function bodies and test cases. Ruff, formatting, `git diff --check`, and GitHub pre-commit passed. The hardware results above apply to the earlier integration; this rebased revision has not had a new hardware run. - vLLM main: vllm-project/vllm@a97dacb Signed-off-by: Li Jiahang <216526138+lijiahang226@users.noreply.github.com> Signed-off-by: like-0517 <ithwlike@126.com>
- xr-conv-native: decode 因果卷积原生化(main vllm-project#16251 第 1 步) - pr15-mooncake-connector-glm5.3-flash-a3: mooncake connector 适配 - patch02_topk_fix: lightning indexer 规避 CANN 9.1 TopKV2 宽归约故障(宽度<=4096) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
What this PR does / why we need it?
Refs #15665
Integrate the existing AscendC recurrent/chunk KDA and causal-convolution operators into GLM-5.3-Flash. GLM and Kimi K3 share the native KDA call helpers in
vllm_ascend/ops/kda.py, while retaining their model-specific projections, weight loading, beta preparation, cache updates, and metadata handling.Preserve physical-page state views and cache-spec capability detection without changing the cache allocation, grouping, or shared MLA/Mamba tensor ownership from #15913. Mask inactive convolution slots before pointer arithmetic and preserve writeback to noncontiguous state views.
The native operators and GDN interfaces are already in main. Full-model validation combines the GLM Flash integration PRs and their prerequisites; prerequisite commits are excluded from this branch.
Does this PR introduce any user-facing change?
Yes. GLM-5.3-Flash uses AscendC KDA and causal convolution. Unsupported operator configurations are rejected during initialization. Kimi K3 retains its existing beta and gate semantics through the shared helpers.
How was this patch tested?
A3 integration validation included PR revision
2fc40a144:tests/ut/ops/test_kimi_kda.pyandtests/ut/models/test_glm5next_kda_contracts.py.tests/e2e/nightly/single_node/ops/singlecard_ops/test_glm5next_conv_state.pyfor inactive slots and strided state layouts, plus existing Kimi recurrent/prefill tests.The current revision rebases on #16067 and preserves its direct Q/K/V view handling and removal of redundant Kimi output masks. The shared recurrent helper forwards Q/K/V unchanged; GLM and Kimi regression cases assert that contiguous and separately strided views reach the native operator without Python-side materialization.
Eight focused GLM/Kimi contract cases passed in an isolated CPU harness executing the actual function bodies and test cases. Ruff, formatting,
git diff --check, and GitHub pre-commit passed. The hardware results above apply to the earlier integration; this rebased revision has not had a new hardware run.