Repository navigation
Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces support for Kimi K3 DSpark execution within the Ascend speculative decoding runtime. It enables the model to function as a hidden-state drafter, incorporating specialized handling for multimodal embeddings, sequence-parallel sharding, and QuaRot projection alignment. These changes ensure correct integration with the existing speculative decoding infrastructure while maintaining performance through optimized host sequence length management. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[SpecDecode][Feature] Support Kimi K3 DSpark and QuaRot alignmentSuggested PR Summary:
### What this PR does / why we need it?
This PR introduces support for Kimi K3 DSpark (`K3DSparkForCausalLM`) as a hidden state drafter on the Ascend backend. It implements QuaRot target alignment for draft hidden-state projections and shared embedding/lm_head layers. Additionally, it shards parallel draft multimodal embeddings to enter sequence parallelism and publishes exact-at-forward host KV lengths for K3 DSpark. It also adds validation to reject unsupported speculative decoding configurations (e.g., reduce sampling, fine-grained LM-head TP, probabilistic sampling, and batch-size based dynamic speculative decoding).
### Does this PR introduce _any_ user-facing change?
No, this is an internal enhancement to support Kimi K3 DSpark and QuaRot alignment.
### How was this patch tested?
Tested with new unit tests in `tests/ut/spec_decode/test_dspark_proposer.py`, `tests/ut/spec_decode/test_llm_base_proposer.py`, and `tests/ut/spec_decode/test_utils.py`.Review Feedback:
We have identified two high-severity robustness issues in the proposed changes. First, when loading the QuaRot rotation, parsing quant_model_description.json could raise an AttributeError if any intermediate keys are null or not dictionaries; we should catch AttributeError to prevent model loading crashes. Second, directly accessing num_speculative_tokens_per_batch_size on vllm_config.speculative_config can raise an AttributeError if the attribute is missing, so using getattr is recommended.
b3e00fb to
75319a3
Compare
75319a3 to
9b5329b
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
9b5329b to
0756675
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
0756675 to
0133310
Compare
9503528 to
7c85653
Compare
Implementation walkthroughConfiguration normalizationpatch_speculative_config.py maps the supported cropped Qwen3 DSpark descriptor into the architecture and fields expected by the current DSpark runtime. It also derives the DSpark mask/PTD token id during speculative-config initialization. Kimi DSpark registrationllm_base_proposer.py registers the upstream Kimi K3 DSpark class in the hidden-state drafter set and recognizes Kimi multimodal placeholder tokens. The existing proposer lifecycle, embedding sharing, lm_head sharing, and hidden-state combination are reused. Shared-layer preparationThe proposer contains no model-specific QuaRot loader or model-type branch. When embedding or lm_head sharing is selected, it calls an optional prepare_shared_layer hook on the loaded draft model and then calls finish_shared_layer_preparation. Qwen3/Kimi modeling owns rotation loading, its own fc/context projection conversion, creation of draft-owned shared projections, and release of the retained rotation matrix. Models without the hook keep the normal upstream sharing behavior. Per-group attention metadatadspark_proposer.py discovers draft layers inside every KV group. Each AttentionGroup is built with its kernel block size, which may differ from the physical allocation page, and query slot mappings use that logical size. During proposal preparation, device sequence lengths and the model runner canonical CPU mirror are extended by the DSpark query count together. Graph padding extends only the padded tail and does not add a reject-count D2H path. Proposal data flow and coverageTarget auxiliary hidden states are combined by the Kimi draft model, DSpark prepares the anchor/query block, per-group attention metadata is built, and the existing sampling interface returns draft tokens and confidence. Focused tests cover descriptor normalization, Kimi registration, multimodal placeholders, the generic shared-layer hook, per-group block sizes, host/device sequence lengths, graph padding, and proposal-buffer preparation. |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
7c85653 to
a1c80b6
Compare
|
"Which model are you using? Is it RadixArk/Kimi-K3-DSpark?" Please provide it in documentation. |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
c3ddb08 to
3af5820
Compare
047ebc5 to
3015539
Compare
|
Documented the checkpoint explicitly: the public GQA path uses RadixArk/Kimi-K3-DSpark; an MLA run requires a matching Kimi K3 MLA draft checkpoint. This is now stated in the PR description and in the parent deployment guide (#14454). |
Integrate Kimi K3 with the existing Ascend DSpark proposer, including dynamic draft-step configuration, multimodal metadata, corrected host sequence lengths, and model-owned preparation of QuaRot shared embedding and LM-head weights. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
3015539 to
f064a1d
Compare
Keep the generic DSpark sharing path focused on sharing semantics. QuaRot boundary preparation is now owned by the model weight loaders, so remove the optional proposer callbacks and their hook-only tests. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
|
DSpark proposer cleanup update:
Validation: targeted model/spec UT suite passed (46 tests total together with PR #14600), and the full QuaRot K3 + GQA DSpark service matched the established DP0 c2 acceptance baseline: 58.17% current versus 58.35% baseline. |
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | #14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | #14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | #14598 | KDA attention execution and fused RMSNorm gate | | 4 | #14839 | MLA attention and rotary execution | | 5 | #14840 | Attention-residual Triton fusion | | 6 | #14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | #14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | #14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | #14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | #14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
|
Closing this child PR following the merge of parent #14454. |
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate | | 4 | vllm-project#14839 | MLA attention and rotary execution | | 5 | vllm-project#14840 | Attention-residual Triton fusion | | 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: d30086105 <denghaojie1@h-partners.com>
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate | | 4 | vllm-project#14839 | MLA attention and rotary execution | | 5 | vllm-project#14840 | Attention-residual Triton fusion | | 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
What this PR does / why we need it?
Checkpoint
RadixArk/Kimi-K3-DSparkis the public GQA draft checkpoint used by this compatibility path. The MLA path requires a Kimi K3 MLA draft checkpoint matching the target model.Dependencies
Merge after #14600 and #14839. The Kimi hybrid cache layout is provided by #14765.
How was this patch tested?
git diff --checkDoes this PR introduce any user-facing change?
Yes. Kimi K3 can use DSpark speculative decoding on Ascend.