Repository navigation
Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces support for Kimi K3 Multi-Head Latent Attention (MLA) execution on Ascend hardware. The changes include updates to the MLA backend to handle rotary-positioning and weight post-processing, as well as the propagation of causal and non-causal draft metadata. These enhancements enable the vLLM Ascend implementation to support Kimi K3 models effectively. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Attention][Feature] Support MLA without RoPE, output gating, and non-causal speculative decodingSuggested PR Summary:
### What this PR does / why we need it?
This pull request introduces support for Multi-Head Latent Attention (MLA) without Rotary Position Embeddings (RoPE), adds output gating (`g_proj`) support, and enables non-causal (bidirectional) speculative decoding. Specifically, it adds `get_identity_cos_and_sin_mla` to skip rotary cache when RoPE is disabled, handles non-causal sparse modes in Flash Attention, and updates the MLAPO prolog to support native floating-point weights on A5.
Feedback:
Four high-severity issues were identified in the review:
1. Potential `KeyError` in `AscendMLAMetadataBuilder` due to `.attn` suffix mismatch when looking up `static_forward_context`.
2. A similar `KeyError` in `AscendMLAImpl` when `fa_quant_layer` is enabled.
3. A dtype mismatch error in `mla_preprocess_only_decode` when `use_mla_rope` is `False` due to hardcoded `torch.bfloat16`.
4. A potential memory corruption/runtime failure in `_exec_kv_no_rope` due to passing a non-contiguous `k_pe` tensor to `reshape_and_cache`.
### Does this PR introduce _any_ user-facing change?
No, these are internal backend optimizations and feature alignments for speculative decoding and MLA execution on Ascend devices.
### How was this patch tested?
Tested via new unit tests in `tests/ut/attention/a2/test_mla_v1.py` covering draft layer rope mode, target layer nope mode, and mixed rope modes, as well as `tests/ut/ops/test_rotary_embedding.py` for identity cos/sin cache skipping.
Implementation walkthroughMLA metadata and causality
Kimi target MLA uses the normal causal path. Kimi MLA DSpark uses a non-causal multi-token draft block, so its metadata keeps No-RoPE Kimi pathFor layers without MLA RoPE, Output gateWhen the upstream Kimi module provides A5 MLA Prolog V3 weight preparationThe post-loading path prepares the fused Kimi projections for
This uses the standard model post-loading hook and the existing CANN operator interface. Rotary and SP handling
CoverageThe tests cover causal/non-causal metadata, no-RoPE identity buffers, Kimi output gating, A5 native and quantized Prolog V3 preparation, padded heads, and rotary/SP position behavior. |
19f5288 to
aa7c434
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
Preserve per-group RoPE semantics, non-causal multi-token decode metadata, explicit no-RoPE execution, and A5 MLA preprocessing for Kimi K3 target and draft layers. Reuse the current device MLA implementation and PCP compatibility paths. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Keep absent cos/sin metadata intact before calling the existing no-RoPE query and KV paths. The PCP prefill changes introduced unconditional slicing before those paths could bypass rotation. Preserve the distinct actual-query and padded-KV lengths for RoPE layers, and cover pure and mixed prefill with numeric Q/K/V and cache-write assertions. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
aa7c434 to
4d8c0a7
Compare
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | #14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | #14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | #14598 | KDA attention execution and fused RMSNorm gate | | 4 | #14839 | MLA attention and rotary execution | | 5 | #14840 | Attention-residual Triton fusion | | 6 | #14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | #14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | #14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | #14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | #14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
|
Closing this child PR following the merge of parent #14454. |
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate | | 4 | vllm-project#14839 | MLA attention and rotary execution | | 5 | vllm-project#14840 | Attention-residual Triton fusion | | 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: d30086105 <denghaojie1@h-partners.com>
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate | | 4 | vllm-project#14839 | MLA attention and rotary execution | | 5 | vllm-project#14840 | Attention-residual Triton fusion | | 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
What this PR does / why we need it?
Dependencies
How was this patch tested?
tests/ut/attention/a2/test_mla_v1.pyandtests/ut/ops/test_rotary_embedding.py.Does this PR introduce any user-facing change?
Yes. It enables Kimi K3 MLA execution on Ascend.