Repository navigation
[Feature][Model] Add Kimi K3 support for vLLM 0.27 - #14454
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request integrates Kimi K3 model support into the vLLM 0.27-based main branch, specifically targeting Ascend hardware. It introduces a suite of custom operators and backend adaptations to enable efficient serving of Kimi K3 variants, including text, multimodal, and MTP architectures, while leveraging native vLLM 0.27 interfaces. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
Suggested PR Title:\n\nmarkdown\n[Attention][Feature] Add RecurrentKda, ChunkKdaFwd, KdaGateCumsum, KdaLayoutSwap12, and MlaPrologV3 operators\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\nThis pull request introduces several new operators for the attention module, including `RecurrentKda`, `ChunkKdaFwd`, `KdaGateCumsum`, `KdaLayoutSwap12`, and `MlaPrologV3`. These operators are implemented with Ascend C kernels, host tiling, and aclnn APIs to support Kimi K3 integration, speculative decoding, and other attention-related features. Additionally, the build script `csrc/build_aclnn.sh` is updated to compile these new operators.\n\nFeedback on the code changes includes:\n- In `aclnn_chunk_kda_fwd.cpp`, the `linearDst` tensor in `KdaFwdCopyMaybeCastAfter` needs to be reshaped to match the flat shape of `linearSrc` to prevent shape validation failures in `KdaLayoutSwap12`.\n- In `recurrent_kda_tiling.cpp` and `kda_gate_cumsum_tiling.cpp`, retrieving the `layout` string attribute using `GetAttrPointer<char>` is incorrect and can cause runtime crashes; `attrs->GetStr` should be used instead.\n\n### Does this PR introduce _any_ user-facing change?\nYes, it introduces new PyTorch and aclnn APIs for the newly added operators: `npu_recurrent_kda`, `npu_chunk_kda_fwd`, `npu_kda_gate_cumsum`, `npu_kda_layout_swap12`, and `npu_mla_prolog_v3`.\n\n### How was this patch tested?\nThe operators can be verified using the provided accuracy tests under `csrc/attention/recurrent_kda/tests/pta/`.\n
c5187b8 to
aeebb5c
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
aeebb5c to
31c4820
Compare
31c4820 to
b7dd3fe
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
b7dd3fe to
7ac9d69
Compare
Remove the Kimi K3 dummy nightly case, its dedicated fixture and workflow entry. Clean up the matching documentation claims while retaining real-checkpoint validation guidance. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Keep SiTU in the existing MoE activation dispatch and evaluate the upstream native formula directly. This avoids constructing a CustomOp after the model-init vLLM config context has exited in the unquantized, antiquant and W4A16 paths. Exercise the existing SiTU regression tests without a forward config context and compare all supported floating dtypes with the upstream native activation. The quantized fused SiTU kernels are unchanged. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Allow uniform handoff batches to use decode graphs once initial state is available, including prompts at N-1 with speculative padding. Keep first-token prefills and nonuniform batches on their existing path without changing DP mode synchronization. Replace the completed-prompt assertion test with CPU behavior checks using the real dispatcher and DP synchronization for DP1 and DP4. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Keep asynchronous accepted-token D2H results outside the mutable InputBatch until they are remapped into current request order, following the ownership fix in vllm-project#13864. Reuse the existing buffers and event, and preserve synchronous scheduling behavior. Trust the per-rank capacity returned by each KV spec instead of multiplying replicated Mamba tables by DCP again. Cover real request replacement and backend reordering, scheduling modes, and final speculative slots for raw and grouped Mamba specs at DCP1/DCP4. Validation: 89 targeted tests passed, 2 upstream-version cases skipped; bash format.sh ci passed. No full-model NPU serving run was performed for this change. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Use the same beta=1.0 default as the existing quantized MoE path before calling native SiTU from the unquantized path. Keep configured beta values unchanged and avoid constructing a runtime CustomOp. Extend the existing upstream parity test to cover an omitted beta across floating dtypes and optional linear clipping. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Give the upstream reference CustomOp only its compilation configuration instead of constructing a full VllmConfig. This avoids platform logging initialization against Ascend mock state left by preceding unit tests. Keep the native upstream numerical comparison and run the MoE call outside the config context. Validate alongside linear, encoder-attention and MoE communication tests. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
|
/rerun Rerun (failed jobs only):
|
Keep absent cos/sin metadata intact before calling the existing no-RoPE query and KV paths. The PCP prefill changes introduced unconditional slicing before those paths could bypass rotation. Preserve the distinct actual-query and padded-KV lengths for RoPE layers, and cover pure and mixed prefill with numeric Q/K/V and cache-write assertions. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Register the upstream FLA FusedRMSNormGated class with the Ascend backend so K3 output normalization reaches the existing fused Triton kernel instead of decomposed native operations inside kda_attention. Reuse the v0.26 kernel arithmetic and tiling, preserve the loaded parameters and epsilon, and allocate a separate result as in the v0.26 K3 adapter. Cover CustomOp dispatch plus NPU numerics for packed gate strides, FP16/BF16, sigmoid/SiLU, affine and residual/prenorm paths. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Add the measured NPU norm-gate test to estimated_times so selective CI coverage validation can schedule the existing test. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Build reduced 16-layer/16-expert K3 configs in one A3 16-device PR test, covering MLA and GQA DSpark, W4A8, DP2/TP8, MTP images, prefix-cache boundaries and P/D transfer with local prefill fallback. Restore DP-wide MC2 padding around local SP shards while preserving rank-local masks and hash token IDs. Match upstream text-only draft handling, use the target language-model head for MTP, and adapt its decoder attention return convention. Exclude K3 MLA from Transformers' inherited MHA divisibility check. Validated all six A3 functional cases on vLLM 0.27.1, MC2 unit tests, targeted mypy for Python 3.10/3.11/3.12, CI routing/coverage and format.sh ci. Dummy results are not accuracy or QuaRot checkpoint validation. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Use four complementary deployment cases: Block5 MLA at TP16, W4A8 GQA at DP2/TP8, legacy MLA at P8/D8, and MTP with an image. Keep production widths and 16 experts in a six-layer target with a reduced attention-residual block. Share local model fixtures and omit redundant fused-expert quantization descriptions. Initialize the text-only multimodal capability in the bare proposer UT fixture to match the base constructor contract. No production runtime changes. Validated all four cases on one A3-16 with vLLM 0.27.1 in 357.92 seconds, 23 proposer unit tests, targeted mypy, CI routing/coverage, and format.sh ci. Set the suite estimate to 360 seconds. Dummy smoke is not checkpoint accuracy or QuaRot validation. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Remove source-text assertions, duplicated constructor and forwarding checks, and redundant mock scaffolding from the Kimi K3 unit tests. Consolidate rotation loading and Mamba copy checks around actual tensor results while retaining cache-capacity, P/D transfer, accepted-token and operator regressions. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
|
/rerun Rerun (failed jobs only):
|
What this PR does / why we need it?
This PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend, including the required Ascend operators, runtime integration, and deployment documentation.
Deployment and validation documentation:
docs/source/tutorials/models/Kimi-K3.mddocs/source/tutorials/models/index.mdThe guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting.
How was this patch tested?
Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and
bash format.sh cipass. A5 performance and full-model serving were not rerun for this change.State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4.
bash format.sh cipasses. Full-model NPU serving was not rerun for the snapshot/capacity changes.Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change.
All changed Python files pass syntax compilation.
Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration.
Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.
Detailed accuracy and performance results remain in the PR comments.
Does this PR introduce any user-facing change?
Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend.