Repository navigation
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces the Ascend attention backend integration for the Kimi K3 model. It provides essential support for KDA and MLA mechanisms, including specialized prefill and recurrent execution wrappers, output gating, and optimized A5 dispatch. These changes facilitate the Kimi K3 integration stack by enabling necessary attention-layer functionality while maintaining compatibility with existing vLLM contracts. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Ignored Files
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Attention][Feature] Add Ascend support for Kimi K3 Delta-Attention and enhance MLA featuresSuggested PR Summary:
### What this PR does / why we need it?
This PR introduces Ascend backend support for the Kimi K3 delta-attention (KDA) layer by implementing `AscendKimiK3DeltaAttention` using AscendC prefill and recurrent kernels. It also enhances the MLA (Multi-Head Latent Attention) implementation to support optional RoPE inputs, output gating, and deferred parallel draft reject finalization. Additionally, it updates the GDN attention builder to correctly extract head counts from `linear_attn_config`.
### Does this PR introduce _any_ user-facing change?
Yes, it adds support for Kimi K3 models and enhances MLA capabilities on Ascend devices.
### How was this patch tested?
Tested using newly added unit tests in `tests/ut/attention/a2/test_mla_v1.py` and `tests/ut/ops/test_gdn_attn_builder.py`.
### Review Feedback
1. **MLA RoPE Mode Lookup**: In `mla_v1.py`, the `layer_name` lookup in `static_forward_context` might fail if the keys do not match exactly. It is recommended to fall back to checking `layer_name.rsplit(".", 1)[0]`.
2. **Tensor Padding**: In `mla_v1.py`, using `torch.cat` with a zero tensor is preferred over `F.pad` with 6 elements on a 3D tensor for better compatibility and performance on `torch_npu`.
3. **Safe Namespace Access**: In `kimi_kda.py`, use `getattr(torch.ops, "_C_ascend", None)` to safely check for the custom namespace registration to avoid unhandled `AttributeError`s.
4. **Distributed Check**: In `kimi_kda.py`, avoid calling `get_pcp_group().world_size` unconditionally as it can crash in non-distributed environments; check `self.vllm_config.parallel_config.prefill_context_parallel_size > 1` instead.
5. **GDN Head Count Detection**: In `gdn_attn_builder.py`, support both dictionary and attribute access for `linear_attn_config` to ensure robust head count detection.d75e917 to
7b84ad1
Compare
0f59def to
4365dde
Compare
Implementation walkthroughMetadata construction
The builder compacts empty segments before launching the prefill operator, pads/reset graph inputs for stable ACLGraph replay, and keeps the sequence reorder indices on device. The one-token boundary is state-aware: a one-token request with existing recurrent state is decode, while a genuine one-token prompt without state remains prefill. This prevents the same KDA execution
Backend registration
CoverageThe tests cover prefill/decode classification, one-token boundaries, empty segments, graph padding/reset, causal-convolution metadata, mixed projection loading, chunk prefill, recurrent decode, and padded-output zeroing. |
4365dde to
673d2a0
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
Implement Ascend KDA prefill and recurrent decode with the existing GDN metadata path, AscendC kernels, graph padding, mixed-precision projection loading, and focused boundary tests. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Register the upstream FLA FusedRMSNormGated class with the Ascend backend so K3 output normalization reaches the existing fused Triton kernel instead of decomposed native operations inside kda_attention. Reuse the v0.26 kernel arithmetic and tiling, preserve the loaded parameters and epsilon, and allocate a separate result as in the v0.26 K3 adapter. Cover CustomOp dispatch plus NPU numerics for packed gate strides, FP16/BF16, sigmoid/SiLU, affine and residual/prenorm paths. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Remove stale references to the retired merged gate projection and is_vl_model helper so the KDA test module collects against the current implementation. Keep checkpoint epsilon and prefill/recurrent execution coverage unchanged. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Register the KDA implementation in the existing GDN test module and add the fused norm-gate NPU test to estimated_times. This keeps selective coverage validation complete after rebasing the KDA child PR. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
6d33147 to
cbcea5d
Compare
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | #14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | #14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | #14598 | KDA attention execution and fused RMSNorm gate | | 4 | #14839 | MLA attention and rotary execution | | 5 | #14840 | Attention-residual Triton fusion | | 6 | #14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | #14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | #14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | #14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | #14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
|
Closing this child PR following the merge of parent #14454. |
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate | | 4 | vllm-project#14839 | MLA attention and rotary execution | | 5 | vllm-project#14840 | Attention-residual Triton fusion | | 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: d30086105 <denghaojie1@h-partners.com>
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate | | 4 | vllm-project#14839 | MLA attention and rotary execution | | 5 | vllm-project#14840 | Attention-residual Triton fusion | | 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
What this PR does / why we need it?
FusedRMSNormGatedwith the Ascend CustomOp backend and reuses the existing fused Triton kernel for output normalization and gating. Weight loading, checkpoint epsilon, and kernel arithmetic/tiling are unchanged; the result is allocated separately to preserve the input.Dependencies
How was this patch tested?
git diff --checkandbash format.sh citests/ut/ops/test_gdn_attn_builder.pyandtests/ut/ops/test_kimi_kda.pyDoes this PR introduce any user-facing change?
No model is registered by this PR alone. It provides KDA execution for Kimi K3.