Skip to content

[Feature][MoE] Support Kimi K3 SiTU and quantization - #14599

Closed
maoxx241 wants to merge 5 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-situ-moe-main
Closed

maoxx241 wants to merge 5 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-situ-moe-main

Conversation

@maoxx241

@maoxx241 maoxx241 commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

This PR provides the MoE execution and ModelSlim quantization support required by Kimi K3:

  • Propagates SiTU beta and linear-beta parameters through dense and routed experts.
  • Dispatches A2/A3 Dequant-SiTU and A5 MX-SiTU through the existing Ascend fused-MoE path.
  • Extends ModelSlim mappings and layer-skip handling for the Kimi K3 mixed W8A8/BF16 KDA projection layout.
  • Starts shared-expert SP collectives on the auxiliary stream and orders them away from routed dispatch/combine with NPU events.
  • Adds focused fused-MoE, shared-expert, runtime-argument, and ModelSlim tests.

Dependencies

How was this patch tested?

  • Python bytecode compilation for changed implementation and test files
  • Focused fused-MoE, shared-expert, runtime-argument, and ModelSlim coverage
  • A3 BF16 SiTU output comparison against the FP32 reference
  • Integrated TP16/EP serving coverage with text, multimodal, and ACLGraph replay requests
  • MX-SiTU A5 cross-compilation and packaging

Does this PR introduce any user-facing change?

No model is registered by this PR alone. It provides reusable Ascend MoE and ModelSlim quantization support used by the Kimi K3 integration.

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces support for Kimi K3 SiTU activation and its associated quantized MoE execution paths. It enables the propagation of SiTU-specific parameters through dense, routed, and shared-expert layers, while ensuring compatibility with existing LoRA and ModelSlim quantization workflows. The changes provide the necessary infrastructure for Kimi K3 integration, including kernel dispatch logic and expanded regression testing to maintain system stability.

Highlights

  • SiTU Activation Support: Added Kimi K3 SiTU (Sigmoid-Tanh Unit) activation support, including beta and linear-beta parameter propagation.
  • MoE Quantization Integration: Extended MoE layers to support A2/A3 Dequant-SiTU and A5 MX-SiTU execution contracts.
  • LoRA and ModelSlim Compatibility: Updated LoRA and ModelSlim quantization paths to integrate seamlessly with Kimi K3 model structures.
  • Regression Coverage: Added focused unit tests for activation, linear, fused-MoE, LoRA, and ModelSlim components to ensure consistency.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Support Kimi SiTU activation and quantized pipeline

Suggested PR Summary:

### What this PR does / why we need it?

This PR introduces support for Kimi's SiTU gated activation and its quantized pipeline (`GMM1 -> SiTU -> GMM2`) in the Ascend backend. Specifically, it:
- Adds `SituActivationConfig`, `situ_and_mul`, and `AscendSituAndMul` to `vllm_ascend/ops/activation.py`.
- Implements `_w4a8_situ_apply_mlp` in `vllm_ascend/ops/fused_moe/moe_mlp.py` to support the fused W4A8 quantization path for SiTU.
- Updates `_get_column_parallel_op` in `vllm_ascend/ops/linear_op.py` to handle head-wise attention gates (`g_proj`) correctly.
- Integrates `kimi_k3` and `kimi_linear` packed modules mapping in ModelSlim quantization configuration.
- Updates shared experts to use projection input width instead of hidden size for consistency validation.

Feedback:
An issue was identified in `vllm_ascend/ops/linear_op.py` where accessing `model_config.hf_text_config` could raise an `AttributeError` if it is not defined on the `model_config` object. It is recommended to catch `AttributeError` in addition to `AssertionError` to prevent potential runtime crashes.

### Does this PR introduce _any_ user-facing change?

No, this PR only adds backend support and optimizations for Kimi's SiTU activation and ModelSlim quantization configurations.

### How was this patch tested?

The changes are covered by new unit tests added in:
- `tests/ut/lora/test_quant_moe.py`
- `tests/ut/ops/test_fused_moe.py`
- `tests/ut/ops/test_linear.py`
- `tests/ut/quantization/test_modelslim_config.py`

Comment thread vllm_ascend/ops/linear_op.py Outdated
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@maoxx241
maoxx241 force-pushed the codex/kimi-k3-situ-moe-main branch from 2bf616d to 17e16d1 Compare August 20, 2026 04:41
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-situ-moe-main branch 3 times, most recently from a6fed69 to f3a05a9 Compare August 20, 2026 06:34
@maoxx241
maoxx241 marked this pull request as ready for review August 21, 2026 01:31
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-situ-moe-main branch 5 times, most recently from 7ae842b to d8d9080 Compare August 21, 2026 13:17
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@maoxx241
maoxx241 force-pushed the codex/kimi-k3-situ-moe-main branch from d8d9080 to 66b606f Compare August 22, 2026 10:46
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-situ-moe-main branch 3 times, most recently from ceb4f36 to 5b18c8a Compare August 24, 2026 09:25
@maoxx241 maoxx241 changed the title [Feature][MoE] Support Kimi K3 SiTU quantization [Feature][MoE] Support Kimi K3 SiTU execution Aug 24, 2026
@maoxx241

maoxx241 commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor Author

Implementation walkthrough

SiTU propagation

moe_mlp.py, fused_moe.py, and dataclass/moe_mlp.py carry activation_situ_beta and activation_situ_linear_beta through the existing MoE runtime arguments. Dense, routed, shared-expert, floating-point, quantized, and antiquant paths therefore use the same SiTU contract.

Quantized routed-expert path

SiTU is integrated into the existing quant_apply_mlp GMM1 -> activation -> GMM2 flow. After GMM1, A2/A3 uses dequant_situ_quant and A5 uses situ_mx_quant when the following GMM2 requires a quantized activation; the existing scale, grouped-matmul, and QuantType inputs are reused. Antiquant and floating-point execution continue through the common activation flow rather than a parallel K3-only helper.

Shared-expert overlap

For sequence-parallel shared experts, input all-gather, shared projection work, and final reduce-scatter remain on the auxiliary stream. NPU events order the shared collectives away from routed dispatch/combine without host synchronization, while activation/GMM work overlaps the compatible routed stages.

ModelSlim mixed projection

modelslim_config.py recognizes the Kimi K3 packed KDA layout: Q/K/V use the configured quantization method while g_proj, f_a_proj, and b_proj use the existing unquantized linear method. The K3 decision is part of is_layer_skipped_ascend, so it follows the common ModelSlim skip path and remains compatible with the planned per-quant-class refactor.

Coverage

Focused tests cover SiTU argument propagation, floating-point and quantized routed experts, shared-expert event ordering, GMM1/SiTU/GMM2 scales, antiquant execution, and ModelSlim mixed-projection selection.

@maoxx241
maoxx241 force-pushed the codex/kimi-k3-situ-moe-main branch from 5b18c8a to 6bd953c Compare August 24, 2026 14:50
Propagate SiTU parameters through routed and shared experts, add quantized A2/A3 and A5 execution paths, support ModelSlim mixed KDA projections, and order shared-expert sequence-parallel collectives against routed communication.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-situ-moe-main branch from 6bd953c to 327a61e Compare August 24, 2026 15:01
Comment thread vllm_ascend/ops/fused_moe/moe_mlp.py Outdated
Comment thread vllm_ascend/quantization/modelslim_config.py Outdated
Comment thread vllm_ascend/ops/fused_moe/moe_mlp.py Outdated
Comment thread vllm_ascend/quantization/modelslim_config.py
@maoxx241 maoxx241 changed the title [Feature][MoE] Support Kimi K3 SiTU execution [Feature][MoE] Support Kimi K3 SiTU and quantization Aug 25, 2026
Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Keep SiTU in the existing MoE activation dispatch and evaluate the upstream native formula directly. This avoids constructing a CustomOp after the model-init vLLM config context has exited in the unquantized, antiquant and W4A16 paths.

Run the SiTU regression tests without a forward config context and compare supported floating dtypes with the upstream native activation. Also synchronize the existing parent EPLB test forward-context mock. Quantized fused SiTU kernels are unchanged.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Use the same beta=1.0 default as the existing quantized MoE path before calling native SiTU from the unquantized path. Keep configured beta values unchanged and avoid constructing a runtime CustomOp.

Extend the existing upstream parity test to cover an omitted beta across floating dtypes and optional linear clipping.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Give the upstream reference CustomOp only its compilation configuration instead of constructing a full VllmConfig. This avoids platform logging initialization against Ascend mock state left by preceding unit tests.

Keep the native upstream numerical comparison and run the MoE call outside the config context. Validate alongside linear, encoder-attention and MoE communication tests.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
linfeng-yuan pushed a commit that referenced this pull request Aug 28, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | #14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | #14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | #14598 | KDA attention execution and fused RMSNorm gate |
| 4 | #14839 | MLA attention and rotary execution |
| 5 | #14840 | Attention-residual Triton fusion |
| 6 | #14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | #14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | #14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | #14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | #14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
@maoxx241

Copy link
Copy Markdown
Contributor Author

Closing this child PR following the merge of parent #14454.

@maoxx241 maoxx241 closed this Aug 28, 2026
ASH-XING pushed a commit to ASH-XING/vllm-ascend that referenced this pull request Aug 28, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate |
| 4 | vllm-project#14839 | MLA attention and rotary execution |
| 5 | vllm-project#14840 | Attention-residual Triton fusion |
| 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: d30086105 <denghaojie1@h-partners.com>
Lethobenthos20 pushed a commit to Lethobenthos20/vllm-ascend that referenced this pull request Sep 4, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate |
| 4 | vllm-project#14839 | MLA attention and rotary execution |
| 5 | vllm-project#14840 | Attention-residual Triton fusion |
| 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants