Repository navigation
Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request integrates the Kimi K3 attention-residual fusion into the vLLM Ascend backend. By leveraging Triton, the implementation provides a high-performance path for learned softmax mixtures of residual streams, ensuring compatibility with the current vLLM 0.27 buffer management standards. The changes are strictly internal to the model's operation and do not affect public APIs. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Ops][Feature] Add Triton fusion kernel for Kimi K3 attention-residual mixtureSuggested PR Summary:
### What this PR does / why we need it?
This PR adds a Triton fusion kernel for the Kimi K3 attention-residual mixture along with numerical regression tests. However, the current Triton kernel implementation uses `tl.arange(0, H)` where `H` (7168) is not a power of two, which will cause compilation failures or undefined behavior. The feedback suggests passing a power-of-two block size `BLOCK_H` to the kernel and applying a mask `cols < H` for all loads and stores.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Tested with the newly added nightly single-node test: `tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_kimi_k3_fusions.py`.
Implementation walkthrough
For each token:
The launch uses one program per vector-core partition of the token dimension. Hidden-state columns and the next-power-of-two score vector are compile-time specialized, while the valid block count is passed as the model-level value. The nightly regression compares the Triton result with the FP32 PyTorch formula for:
Both cases use the K3 hidden size and verify BF16 output within the expected tolerance. |
Port the validated attention-residual Triton computation to vLLM 0.27's preallocated contiguous buffer contract and specialize the profile shapes used during graph capture. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
eed7246 to
431fdd1
Compare
Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | #14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | #14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | #14598 | KDA attention execution and fused RMSNorm gate | | 4 | #14839 | MLA attention and rotary execution | | 5 | #14840 | Attention-residual Triton fusion | | 6 | #14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | #14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | #14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | #14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | #14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
|
Closing this child PR following the merge of parent #14454. |
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate | | 4 | vllm-project#14839 | MLA attention and rotary execution | | 5 | vllm-project#14840 | Attention-residual Triton fusion | | 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: d30086105 <denghaojie1@h-partners.com>
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate | | 4 | vllm-project#14839 | MLA attention and rotary execution | | 5 | vllm-project#14840 | Attention-residual Triton fusion | | 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
What this PR does / why we need it?
Dependencies
How was this patch tested?
git diff --checktests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_kimi_k3_fusions.pyDoes this PR introduce any user-facing change?
No direct API change. Kimi K3 uses the fusion internally.