Skip to content

[Feature][KDA] Add Kimi K3 KDA execution - #14598

Closed
maoxx241 wants to merge 4 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-attention-main
Closed

maoxx241 wants to merge 4 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-attention-main

Conversation

@maoxx241

@maoxx241 maoxx241 commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

  • Adds the Ascend KDA attention builder and runtime.
  • Implements chunk-prefill and recurrent-decode dispatch through the KDA AscendC operators.
  • Handles causal-convolution state packing and mixed W8A8/BF16 projection loading.
  • Registers upstream FusedRMSNormGated with the Ascend CustomOp backend and reuses the existing fused Triton kernel for output normalization and gating. Weight loading, checkpoint epsilon, and kernel arithmetic/tiling are unchanged; the result is allocated separately to preserve the input.
  • Adds focused builder, KDA execution, CustomOp dispatch, and fused norm-gate numerical tests.

Dependencies

How was this patch tested?

  • git diff --check and bash format.sh ci
  • On each of vLLM v0.27.1 and the pinned upstream revision: 40 targeted CPU tests pass (4 existing skips), and 22 real A3 NPU numerical cases pass. Norm-gate coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, affine and residual/prenorm paths, and input preservation.
  • A5 performance and full-model serving were not rerun for the fused norm-gate change.
  • Python bytecode compilation for changed implementation and test files
  • Unit coverage in tests/ut/ops/test_gdn_attn_builder.py and tests/ut/ops/test_kimi_kda.py
  • Integrated KDA prefill/decode coverage with the Kimi K3 checkpoint

Does this PR introduce any user-facing change?

No model is registered by this PR alone. It provides KDA execution for Kimi K3.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces the Ascend attention backend integration for the Kimi K3 model. It provides essential support for KDA and MLA mechanisms, including specialized prefill and recurrent execution wrappers, output gating, and optimized A5 dispatch. These changes facilitate the Kimi K3 integration stack by enabling necessary attention-layer functionality while maintaining compatibility with existing vLLM contracts.

Highlights

  • KDA Integration: Added KDA prefill and recurrent execution wrappers for Kimi K3 model support.
  • MLA Enhancements: Implemented no-RoPE MLA metadata handling, output gating, and attention residual support.
  • A5 Dispatch: Added A5 native MLA Prolog V3 dispatch for BF16 and optional RoPE configurations.
  • Parallel Drafting: Implemented sequence-length finalization for parallel-drafting at the MLA boundary.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/scripts/test_config.yaml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Feature] Add Ascend support for Kimi K3 Delta-Attention and enhance MLA features

Suggested PR Summary:

### What this PR does / why we need it?
This PR introduces Ascend backend support for the Kimi K3 delta-attention (KDA) layer by implementing `AscendKimiK3DeltaAttention` using AscendC prefill and recurrent kernels. It also enhances the MLA (Multi-Head Latent Attention) implementation to support optional RoPE inputs, output gating, and deferred parallel draft reject finalization. Additionally, it updates the GDN attention builder to correctly extract head counts from `linear_attn_config`.

### Does this PR introduce _any_ user-facing change?
Yes, it adds support for Kimi K3 models and enhances MLA capabilities on Ascend devices.

### How was this patch tested?
Tested using newly added unit tests in `tests/ut/attention/a2/test_mla_v1.py` and `tests/ut/ops/test_gdn_attn_builder.py`.

### Review Feedback
1. **MLA RoPE Mode Lookup**: In `mla_v1.py`, the `layer_name` lookup in `static_forward_context` might fail if the keys do not match exactly. It is recommended to fall back to checking `layer_name.rsplit(".", 1)[0]`.
2. **Tensor Padding**: In `mla_v1.py`, using `torch.cat` with a zero tensor is preferred over `F.pad` with 6 elements on a 3D tensor for better compatibility and performance on `torch_npu`.
3. **Safe Namespace Access**: In `kimi_kda.py`, use `getattr(torch.ops, "_C_ascend", None)` to safely check for the custom namespace registration to avoid unhandled `AttributeError`s.
4. **Distributed Check**: In `kimi_kda.py`, avoid calling `get_pcp_group().world_size` unconditionally as it can crash in non-distributed environments; check `self.vllm_config.parallel_config.prefill_context_parallel_size > 1` instead.
5. **GDN Head Count Detection**: In `gdn_attn_builder.py`, support both dictionary and attribute access for `linear_attn_config` to ensure robust head count detection.

Comment thread vllm_ascend/attention/mla_v1.py Outdated
Comment thread vllm_ascend/attention/mla_v1.py Outdated
Comment thread vllm_ascend/ops/kimi_kda.py Outdated
Comment thread vllm_ascend/ops/kimi_kda.py Outdated
Comment thread vllm_ascend/ops/gdn_attn_builder.py
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-attention-main branch 5 times, most recently from d75e917 to 7b84ad1 Compare August 20, 2026 06:28
@maoxx241
maoxx241 marked this pull request as ready for review August 21, 2026 01:31
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-attention-main branch 6 times, most recently from 0f59def to 4365dde Compare August 24, 2026 09:22
@maoxx241 maoxx241 changed the title [Feature][Attention] Add Kimi K3 KDA and MLA integration [Feature][KDA] Add Kimi K3 KDA execution Aug 24, 2026
@maoxx241

Copy link
Copy Markdown
Contributor Author

Implementation walkthrough

Metadata construction

vllm_ascend/ops/gdn_attn_builder.py builds separate metadata for:

  • chunk prefill;
  • recurrent decode;
  • speculative decode;
  • causal-convolution state.

The builder compacts empty segments before launching the prefill operator, pads/reset graph inputs for stable ACLGraph replay, and keeps the sequence reorder indices on device.

The one-token boundary is state-aware: a one-token request with existing recurrent state is decode, while a genuine one-token prompt without state remains prefill. This prevents the same query_len == 1 shape from selecting the wrong KDA path.

KDA execution

AscendKimiK3DeltaAttention in vllm_ascend/ops/kimi_kda.py provides the Ascend implementation used by the upstream Kimi KDA module:

  1. project and pack the Q/K/V, gate, and convolution inputs;
  2. run the custom causal-convolution operator and update its cache;
  3. for prefill, build the cumulative gate and call chunk_kda_fwd;
  4. for decode, normalize the recurrent gate and call recurrent_kda;
  5. zero graph-padding rows before returning the output.

_pack_conv_weights() prepares the convolution weights once in the layout required by the NPU operator. The projection path supports the K3 mixed layout where Q/K/V are quantized while gate-related projections remain floating point.

Backend registration

AscendGDNAttentionBackend returns the new metadata builder, allowing the existing vLLM attention-layer registration to select this implementation without a K3-specific model-runner branch.

Coverage

The tests cover prefill/decode classification, one-token boundaries, empty segments, graph padding/reset, causal-convolution metadata, mixed projection loading, chunk prefill, recurrent decode, and padded-output zeroing.

@maoxx241
maoxx241 force-pushed the codex/kimi-k3-attention-main branch from 4365dde to 673d2a0 Compare August 24, 2026 14:50
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Implement Ascend KDA prefill and recurrent decode with the existing GDN metadata path, AscendC kernels, graph padding, mixed-precision projection loading, and focused boundary tests.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Register the upstream FLA FusedRMSNormGated class with the Ascend backend so K3 output normalization reaches the existing fused Triton kernel instead of decomposed native operations inside kda_attention. Reuse the v0.26 kernel arithmetic and tiling, preserve the loaded parameters and epsilon, and allocate a separate result as in the v0.26 K3 adapter. Cover CustomOp dispatch plus NPU numerics for packed gate strides, FP16/BF16, sigmoid/SiLU, affine and residual/prenorm paths.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Remove stale references to the retired merged gate projection and is_vl_model helper so the KDA test module collects against the current implementation. Keep checkpoint epsilon and prefill/recurrent execution coverage unchanged.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Register the KDA implementation in the existing GDN test module and add the fused norm-gate NPU test to estimated_times. This keeps selective coverage validation complete after rebasing the KDA child PR.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-attention-main branch from 6d33147 to cbcea5d Compare August 27, 2026 16:28
linfeng-yuan pushed a commit that referenced this pull request Aug 28, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | #14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | #14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | #14598 | KDA attention execution and fused RMSNorm gate |
| 4 | #14839 | MLA attention and rotary execution |
| 5 | #14840 | Attention-residual Triton fusion |
| 6 | #14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | #14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | #14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | #14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | #14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
@maoxx241

Copy link
Copy Markdown
Contributor Author

Closing this child PR following the merge of parent #14454.

@maoxx241 maoxx241 closed this Aug 28, 2026
ASH-XING pushed a commit to ASH-XING/vllm-ascend that referenced this pull request Aug 28, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate |
| 4 | vllm-project#14839 | MLA attention and rotary execution |
| 5 | vllm-project#14840 | Attention-residual Triton fusion |
| 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: d30086105 <denghaojie1@h-partners.com>
Lethobenthos20 pushed a commit to Lethobenthos20/vllm-ascend that referenced this pull request Sep 4, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate |
| 4 | vllm-project#14839 | MLA attention and rotary execution |
| 5 | vllm-project#14840 | Attention-residual Triton fusion |
| 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant