Skip to content

[Feature][Spec Decode] Add Kimi K3 DSpark execution - #14601

Closed
maoxx241 wants to merge 2 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-dspark-main
Closed

maoxx241 wants to merge 2 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-dspark-main

Conversation

@maoxx241

@maoxx241 maoxx241 commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

  • Registers the upstream Kimi K3 DSpark model as a hidden-state drafter.
  • Normalizes supported DSpark configuration fields for the proposer.
  • Invokes an optional model-owned shared-layer preparation hook before the generic proposer shares target embedding or LM-head modules; QuaRot loading and matrix conversion remain in modeling.
  • Prepares proposals, verification metadata, multimodal placeholders, and corrected host sequence lengths through the existing Ascend DSpark runtime.
  • Adds focused proposer, configuration, metadata, and utility tests.

Checkpoint

RadixArk/Kimi-K3-DSpark is the public GQA draft checkpoint used by this compatibility path. The MLA path requires a Kimi K3 MLA draft checkpoint matching the target model.

Dependencies

Merge after #14600 and #14839. The Kimi hybrid cache layout is provided by #14765.

How was this patch tested?

  • git diff --check
  • Python syntax compilation for changed implementation and test files
  • Focused proposer, speculative-config, model-owned shared-layer hook, multimodal-placeholder, and metadata tests
  • Integrated GQA/MLA DSpark, QuaRot, ACLGraph, Prefix Cache, C64/C128, and finite non-empty output coverage

Does this PR introduce any user-facing change?

Yes. Kimi K3 can use DSpark speculative decoding on Ascend.

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces support for Kimi K3 DSpark execution within the Ascend speculative decoding runtime. It enables the model to function as a hidden-state drafter, incorporating specialized handling for multimodal embeddings, sequence-parallel sharding, and QuaRot projection alignment. These changes ensure correct integration with the existing speculative decoding infrastructure while maintaining performance through optimized host sequence length management.

Highlights

  • Kimi K3 Integration: Integrated Kimi K3 DSpark as a hidden-state drafter within the Ascend speculative decoding runtime.
  • Multimodal and Parallel Support: Added support for parallel-draft multimodal embeddings and sequence-parallel sharding to improve performance.
  • QuaRot Alignment: Implemented QuaRot draft projection alignment and shared embedding/lm-head boundary handling for model consistency.
  • Optimized Sequence Management: Enhanced host sequence length maintenance and input construction to support K3 DSpark requirements.
  • Testing: Added comprehensive unit tests for the DSpark proposer, metadata utilities, and QuaRot boundary logic.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[SpecDecode][Feature] Support Kimi K3 DSpark and QuaRot alignment

Suggested PR Summary:

### What this PR does / why we need it?
This PR introduces support for Kimi K3 DSpark (`K3DSparkForCausalLM`) as a hidden state drafter on the Ascend backend. It implements QuaRot target alignment for draft hidden-state projections and shared embedding/lm_head layers. Additionally, it shards parallel draft multimodal embeddings to enter sequence parallelism and publishes exact-at-forward host KV lengths for K3 DSpark. It also adds validation to reject unsupported speculative decoding configurations (e.g., reduce sampling, fine-grained LM-head TP, probabilistic sampling, and batch-size based dynamic speculative decoding).

### Does this PR introduce _any_ user-facing change?
No, this is an internal enhancement to support Kimi K3 DSpark and QuaRot alignment.

### How was this patch tested?
Tested with new unit tests in `tests/ut/spec_decode/test_dspark_proposer.py`, `tests/ut/spec_decode/test_llm_base_proposer.py`, and `tests/ut/spec_decode/test_utils.py`.

Review Feedback:
We have identified two high-severity robustness issues in the proposed changes. First, when loading the QuaRot rotation, parsing quant_model_description.json could raise an AttributeError if any intermediate keys are null or not dictionaries; we should catch AttributeError to prevent model loading crashes. Second, directly accessing num_speculative_tokens_per_batch_size on vllm_config.speculative_config can raise an AttributeError if the attribute is missing, so using getattr is recommended.

Comment thread vllm_ascend/spec_decode/llm_base_proposer.py Outdated
Comment thread vllm_ascend/spec_decode/dspark_proposer.py Outdated
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-dspark-main branch 3 times, most recently from b3e00fb to 75319a3 Compare August 20, 2026 06:25
@maoxx241
maoxx241 marked this pull request as ready for review August 21, 2026 01:31
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-dspark-main branch from 75319a3 to 9b5329b Compare August 21, 2026 11:58
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@maoxx241

maoxx241 commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor Author

Implementation walkthrough

Configuration normalization

patch_speculative_config.py maps the supported cropped Qwen3 DSpark descriptor into the architecture and fields expected by the current DSpark runtime. It also derives the DSpark mask/PTD token id during speculative-config initialization.

Kimi DSpark registration

llm_base_proposer.py registers the upstream Kimi K3 DSpark class in the hidden-state drafter set and recognizes Kimi multimodal placeholder tokens. The existing proposer lifecycle, embedding sharing, lm_head sharing, and hidden-state combination are reused.

Shared-layer preparation

The proposer contains no model-specific QuaRot loader or model-type branch. When embedding or lm_head sharing is selected, it calls an optional prepare_shared_layer hook on the loaded draft model and then calls finish_shared_layer_preparation. Qwen3/Kimi modeling owns rotation loading, its own fc/context projection conversion, creation of draft-owned shared projections, and release of the retained rotation matrix. Models without the hook keep the normal upstream sharing behavior.

Per-group attention metadata

dspark_proposer.py discovers draft layers inside every KV group. Each AttentionGroup is built with its kernel block size, which may differ from the physical allocation page, and query slot mappings use that logical size.

During proposal preparation, device sequence lengths and the model runner canonical CPU mirror are extended by the DSpark query count together. Graph padding extends only the padded tail and does not add a reject-count D2H path.

Proposal data flow and coverage

Target auxiliary hidden states are combined by the Kimi draft model, DSpark prepares the anchor/query block, per-group attention metadata is built, and the existing sampling interface returns draft tokens and confidence. Focused tests cover descriptor normalization, Kimi registration, multimodal placeholders, the generic shared-layer hook, per-group block sizes, host/device sequence lengths, graph padding, and proposal-buffer preparation.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@maoxx241
maoxx241 force-pushed the codex/kimi-k3-dspark-main branch from 7c85653 to a1c80b6 Compare August 24, 2026 14:51
@realliujiaxu

Copy link
Copy Markdown
Collaborator

"Which model are you using? Is it RadixArk/Kimi-K3-DSpark?" Please provide it in documentation.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Comment thread vllm_ascend/spec_decode/dspark_proposer.py
Comment thread vllm_ascend/spec_decode/dspark_proposer.py
Comment thread vllm_ascend/spec_decode/llm_base_proposer.py Outdated
Comment thread vllm_ascend/patch/platform/patch_speculative_config.py
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-dspark-main branch from c3ddb08 to 3af5820 Compare August 25, 2026 09:04
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-dspark-main branch 2 times, most recently from 047ebc5 to 3015539 Compare August 25, 2026 10:40
@maoxx241

Copy link
Copy Markdown
Contributor Author

Documented the checkpoint explicitly: the public GQA path uses RadixArk/Kimi-K3-DSpark; an MLA run requires a matching Kimi K3 MLA draft checkpoint. This is now stated in the PR description and in the parent deployment guide (#14454).

Integrate Kimi K3 with the existing Ascend DSpark proposer, including dynamic draft-step configuration, multimodal metadata, corrected host sequence lengths, and model-owned preparation of QuaRot shared embedding and LM-head weights.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-dspark-main branch from 3015539 to f064a1d Compare August 25, 2026 10:50

@drslark drslark left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM.

Comment thread vllm_ascend/spec_decode/llm_base_proposer.py Outdated
Keep the generic DSpark sharing path focused on sharing semantics. QuaRot boundary preparation is now owned by the model weight loaders, so remove the optional proposer callbacks and their hook-only tests.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@maoxx241

Copy link
Copy Markdown
Contributor Author

DSpark proposer cleanup update:

  • Removed the optional prepare_shared_layer and finish_shared_layer_preparation callbacks from the generic proposer.
  • The proposer now only decides whether a draft-owned boundary is kept or a target boundary is shared; model-specific QuaRot loading and rotation live in the modeling layer from PR [Feature][Model] Register Kimi K3 model adapters #14600.
  • Removed the hook-only unit tests; behavior is covered at the model loader boundary instead.

Validation: targeted model/spec UT suite passed (46 tests total together with PR #14600), and the full QuaRot K3 + GQA DSpark service matched the established DP0 c2 acceptance baseline: 58.17% current versus 58.35% baseline.

linfeng-yuan pushed a commit that referenced this pull request Aug 28, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | #14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | #14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | #14598 | KDA attention execution and fused RMSNorm gate |
| 4 | #14839 | MLA attention and rotary execution |
| 5 | #14840 | Attention-residual Triton fusion |
| 6 | #14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | #14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | #14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | #14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | #14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
@maoxx241

Copy link
Copy Markdown
Contributor Author

Closing this child PR following the merge of parent #14454.

@maoxx241 maoxx241 closed this Aug 28, 2026
ASH-XING pushed a commit to ASH-XING/vllm-ascend that referenced this pull request Aug 28, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate |
| 4 | vllm-project#14839 | MLA attention and rotary execution |
| 5 | vllm-project#14840 | Attention-residual Triton fusion |
| 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: d30086105 <denghaojie1@h-partners.com>
Lethobenthos20 pushed a commit to Lethobenthos20/vllm-ascend that referenced this pull request Sep 4, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate |
| 4 | vllm-project#14839 | MLA attention and rotary execution |
| 5 | vllm-project#14840 | Attention-residual Triton fusion |
| 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants