Skip to content

[Feature][Model] Add Kimi K3 support for vLLM 0.27 - #14454

Merged
linfeng-yuan merged 50 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-model-features-main
Aug 28, 2026
Merged

linfeng-yuan merged 50 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-model-features-main

Conversation

@maoxx241

@maoxx241 maoxx241 commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

This PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend, including the required Ascend operators, runtime integration, and deployment documentation.

Deployment and validation documentation:

  • docs/source/tutorials/models/Kimi-K3.md
  • docs/source/tutorials/models/index.md

The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting.

How was this patch tested?

  • Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and bash format.sh ci pass. A5 performance and full-model serving were not rerun for this change.

  • State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. bash format.sh ci passes. Full-model NPU serving was not rerun for the snapshot/capacity changes.

  • Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change.

  • All changed Python files pass syntax compilation.

  • Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration.

  • Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend.

@maoxx241 maoxx241 changed the title [Feat][Model] Add Kimi K3 support for vLLM 0.27 [Feature][Model] Add Kimi K3 support for vLLM 0.27 Aug 17, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request integrates Kimi K3 model support into the vLLM 0.27-based main branch, specifically targeting Ascend hardware. It introduces a suite of custom operators and backend adaptations to enable efficient serving of Kimi K3 variants, including text, multimodal, and MTP architectures, while leveraging native vLLM 0.27 interfaces.

Highlights

  • Kimi K3 Support: Added support for Kimi K3 text, multimodal, MTP, and DSpark architectures on Ascend hardware.
  • Ascend Backend Integration: Composed Kimi K3 models with Ascend KDA and MLA backends, including the addition of the MLA Prolog V3 implementation.
  • New Custom Operators: Introduced multiple custom operators including ChunkKdaFwd, KdaGateCumsum, KdaLayoutSwap12, MlaPrologV3, and RecurrentKda to facilitate efficient model inference.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions github-actions Bot added documentation Improvements or additions to documentation module:tests module:ops labels Aug 17, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@maoxx241 maoxx241 closed this Aug 17, 2026
@maoxx241 maoxx241 reopened this Aug 17, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:\n\nmarkdown\n[Attention][Feature] Add RecurrentKda, ChunkKdaFwd, KdaGateCumsum, KdaLayoutSwap12, and MlaPrologV3 operators\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\nThis pull request introduces several new operators for the attention module, including `RecurrentKda`, `ChunkKdaFwd`, `KdaGateCumsum`, `KdaLayoutSwap12`, and `MlaPrologV3`. These operators are implemented with Ascend C kernels, host tiling, and aclnn APIs to support Kimi K3 integration, speculative decoding, and other attention-related features. Additionally, the build script `csrc/build_aclnn.sh` is updated to compile these new operators.\n\nFeedback on the code changes includes:\n- In `aclnn_chunk_kda_fwd.cpp`, the `linearDst` tensor in `KdaFwdCopyMaybeCastAfter` needs to be reshaped to match the flat shape of `linearSrc` to prevent shape validation failures in `KdaLayoutSwap12`.\n- In `recurrent_kda_tiling.cpp` and `kda_gate_cumsum_tiling.cpp`, retrieving the `layout` string attribute using `GetAttrPointer<char>` is incorrect and can cause runtime crashes; `attrs->GetStr` should be used instead.\n\n### Does this PR introduce _any_ user-facing change?\nYes, it introduces new PyTorch and aclnn APIs for the newly added operators: `npu_recurrent_kda`, `npu_chunk_kda_fwd`, `npu_kda_gate_cumsum`, `npu_kda_layout_swap12`, and `npu_mla_prolog_v3`.\n\n### How was this patch tested?\nThe operators can be verified using the provided accuracy tests under `csrc/attention/recurrent_kda/tests/pta/`.\n

Comment thread csrc/attention/chunk_kda_fwd/op_host/op_api/aclnn_chunk_kda_fwd.cpp
Comment thread csrc/attention/recurrent_kda/op_host/recurrent_kda_tiling.cpp
Comment thread csrc/attention/kda_gate_cumsum/op_host/kda_gate_cumsum_tiling.cpp
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-model-features-main branch 2 times, most recently from c5187b8 to aeebb5c Compare August 17, 2026 13:15
@maoxx241 maoxx241 closed this Aug 17, 2026
@maoxx241 maoxx241 reopened this Aug 17, 2026
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Remove the Kimi K3 dummy nightly case, its dedicated fixture and workflow entry. Clean up the matching documentation claims while retaining real-checkpoint validation guidance.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Keep SiTU in the existing MoE activation dispatch and evaluate the upstream native formula directly. This avoids constructing a CustomOp after the model-init vLLM config context has exited in the unquantized, antiquant and W4A16 paths.

Exercise the existing SiTU regression tests without a forward config context and compare all supported floating dtypes with the upstream native activation. The quantized fused SiTU kernels are unchanged.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Allow uniform handoff batches to use decode graphs once initial state is available, including prompts at N-1 with speculative padding. Keep first-token prefills and nonuniform batches on their existing path without changing DP mode synchronization.

Replace the completed-prompt assertion test with CPU behavior checks using the real dispatcher and DP synchronization for DP1 and DP4.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Keep asynchronous accepted-token D2H results outside the mutable InputBatch until they are remapped into current request order, following the ownership fix in vllm-project#13864. Reuse the existing buffers and event, and preserve synchronous scheduling behavior.

Trust the per-rank capacity returned by each KV spec instead of multiplying replicated Mamba tables by DCP again. Cover real request replacement and backend reordering, scheduling modes, and final speculative slots for raw and grouped Mamba specs at DCP1/DCP4.

Validation: 89 targeted tests passed, 2 upstream-version cases skipped; bash format.sh ci passed. No full-model NPU serving run was performed for this change.
Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Use the same beta=1.0 default as the existing quantized MoE path before calling native SiTU from the unquantized path. Keep configured beta values unchanged and avoid constructing a runtime CustomOp.

Extend the existing upstream parity test to cover an omitted beta across floating dtypes and optional linear clipping.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Give the upstream reference CustomOp only its compilation configuration instead of constructing a full VllmConfig. This avoids platform logging initialization against Ascend mock state left by preceding unit tests.

Keep the native upstream numerical comparison and run the MoE call outside the config context. Validate alongside linear, encoder-attention and MoE communication tests.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@maoxx241

maoxx241 commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

Keep absent cos/sin metadata intact before calling the existing no-RoPE query and KV paths. The PCP prefill changes introduced unconditional slicing before those paths could bypass rotation. Preserve the distinct actual-query and padded-KV lengths for RoPE layers, and cover pure and mixed prefill with numeric Q/K/V and cache-write assertions.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Register the upstream FLA FusedRMSNormGated class with the Ascend backend so K3 output normalization reaches the existing fused Triton kernel instead of decomposed native operations inside kda_attention. Reuse the v0.26 kernel arithmetic and tiling, preserve the loaded parameters and epsilon, and allocate a separate result as in the v0.26 K3 adapter. Cover CustomOp dispatch plus NPU numerics for packed gate strides, FP16/BF16, sigmoid/SiLU, affine and residual/prenorm paths.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Add the measured NPU norm-gate test to estimated_times so selective CI coverage validation can schedule the existing test.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Build reduced 16-layer/16-expert K3 configs in one A3 16-device PR test, covering MLA and GQA DSpark, W4A8, DP2/TP8, MTP images, prefix-cache boundaries and P/D transfer with local prefill fallback.

Restore DP-wide MC2 padding around local SP shards while preserving rank-local masks and hash token IDs. Match upstream text-only draft handling, use the target language-model head for MTP, and adapt its decoder attention return convention. Exclude K3 MLA from Transformers' inherited MHA divisibility check.

Validated all six A3 functional cases on vLLM 0.27.1, MC2 unit tests, targeted mypy for Python 3.10/3.11/3.12, CI routing/coverage and format.sh ci. Dummy results are not accuracy or QuaRot checkpoint validation.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Use four complementary deployment cases: Block5 MLA at TP16, W4A8 GQA at DP2/TP8, legacy MLA at P8/D8, and MTP with an image. Keep production widths and 16 experts in a six-layer target with a reduced attention-residual block. Share local model fixtures and omit redundant fused-expert quantization descriptions.

Initialize the text-only multimodal capability in the bare proposer UT fixture to match the base constructor contract. No production runtime changes.

Validated all four cases on one A3-16 with vLLM 0.27.1 in 357.92 seconds, 23 proposer unit tests, targeted mypy, CI routing/coverage, and format.sh ci. Set the suite estimate to 360 seconds. Dummy smoke is not checkpoint accuracy or QuaRot validation.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Remove source-text assertions, duplicated constructor and forwarding checks, and redundant mock scaffolding from the Kimi K3 unit tests. Consolidate rotation loading and Mamba copy checks around actual tensor results while retaining cache-capacity, P/D transfer, accepted-token and operator regressions.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@maoxx241

maoxx241 commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation module:core module:ops module:quantization module:tests ready-all run all e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants