Skip to content

[Performance][Ops] Add Kimi K3 attention residual fusion - #14840

Closed
maoxx241 wants to merge 3 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-attention-residual-main
Closed

maoxx241 wants to merge 3 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-attention-residual-main

Conversation

@maoxx241

@maoxx241 maoxx241 commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

  • Adds the Triton attention-residual fusion used by Kimi K3.
  • Preserves the upstream preallocated residual-buffer contract while specializing the launch shape for Ascend.
  • Covers valid-block counts and residual layouts with a nightly operator test.

Dependencies

  • No child PR is required to review this operator in isolation.
  • The Kimi K3 model adapter PR consumes this fusion.

How was this patch tested?

  • git diff --check
  • Python bytecode compilation for changed implementation and test files
  • Nightly operator coverage in tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_kimi_k3_fusions.py

Does this PR introduce any user-facing change?

No direct API change. Kimi K3 uses the fusion internally.

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request integrates the Kimi K3 attention-residual fusion into the vLLM Ascend backend. By leveraging Triton, the implementation provides a high-performance path for learned softmax mixtures of residual streams, ensuring compatibility with the current vLLM 0.27 buffer management standards. The changes are strictly internal to the model's operation and do not affect public APIs.

Highlights

  • Triton Kernel Implementation: Introduced a new Triton-based attention-residual fusion kernel specifically for the Kimi K3 model architecture, optimized for Ascend hardware.
  • Contract Compliance: Updated the implementation to align with the vLLM 0.27 preallocated residual-buffer contract while maintaining specialized launch shapes.
  • Regression Testing: Added a new nightly operator test to ensure numerical consistency and coverage for various block counts and residual layouts.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Add Triton fusion kernel for Kimi K3 attention-residual mixture

Suggested PR Summary:

### What this PR does / why we need it?
This PR adds a Triton fusion kernel for the Kimi K3 attention-residual mixture along with numerical regression tests. However, the current Triton kernel implementation uses `tl.arange(0, H)` where `H` (7168) is not a power of two, which will cause compilation failures or undefined behavior. The feedback suggests passing a power-of-two block size `BLOCK_H` to the kernel and applying a mask `cols < H` for all loads and stores.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Tested with the newly added nightly single-node test: `tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_kimi_k3_fusions.py`.

Comment thread vllm_ascend/ops/triton/kimi_k3/attention_residual.py
Comment thread vllm_ascend/ops/triton/kimi_k3/attention_residual.py
@maoxx241

Copy link
Copy Markdown
Contributor Author

Implementation walkthrough

apply_attn_res() computes Kimi K3's learned mixture of attention residual streams in one Triton kernel.

For each token:

  1. read the num_valid_blocks residual streams from the preallocated block_residual buffer;
  2. append prefix_sum as the final candidate stream;
  3. RMS-normalize every candidate in FP32;
  4. score each candidate with norm.weight * proj.weight;
  5. apply a softmax across the valid candidates;
  6. form the weighted residual sum and cast it back to the model dtype.

block_residual.shape[1] is treated as allocation capacity, while num_valid_blocks controls the streams participating in the computation. Unused preallocated slots are never scored or accumulated.

The launch uses one program per vector-core partition of the token dimension. Hidden-state columns and the next-power-of-two score vector are compile-time specialized, while the valid block count is passed as the model-level value.

The nightly regression compares the Triton result with the FP32 PyTorch formula for:

  • a partial-capacity buffer, where valid blocks are fewer than allocated slots;
  • the profile-run shape with the full residual capacity.

Both cases use the K3 hidden size and verify BF16 output within the expected tolerance.

Port the validated attention-residual Triton computation to vLLM 0.27's preallocated contiguous buffer contract and specialize the profile shapes used during graph capture.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-attention-residual-main branch from eed7246 to 431fdd1 Compare August 24, 2026 14:50
Comment thread vllm_ascend/ops/triton/kimi_k3/attention_residual.py
ZT-AIA
ZT-AIA previously approved these changes Aug 25, 2026
Comment thread vllm_ascend/ops/triton/kimi_k3/attention_residual.py
Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Aug 25, 2026
linfeng-yuan pushed a commit that referenced this pull request Aug 28, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | #14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | #14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | #14598 | KDA attention execution and fused RMSNorm gate |
| 4 | #14839 | MLA attention and rotary execution |
| 5 | #14840 | Attention-residual Triton fusion |
| 6 | #14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | #14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | #14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | #14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | #14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
@maoxx241

Copy link
Copy Markdown
Contributor Author

Closing this child PR following the merge of parent #14454.

@maoxx241 maoxx241 closed this Aug 28, 2026
ASH-XING pushed a commit to ASH-XING/vllm-ascend that referenced this pull request Aug 28, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate |
| 4 | vllm-project#14839 | MLA attention and rotary execution |
| 5 | vllm-project#14840 | Attention-residual Triton fusion |
| 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: d30086105 <denghaojie1@h-partners.com>
Lethobenthos20 pushed a commit to Lethobenthos20/vllm-ascend that referenced this pull request Sep 4, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate |
| 4 | vllm-project#14839 | MLA attention and rotary execution |
| 5 | vllm-project#14840 | Attention-residual Triton fusion |
| 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation module:ops module:tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants