Skip to content

[Performance][KDA] Compose and overlap gate projections - #14497

Open
Dawn952 wants to merge 2 commits into
vllm-project:releases/v0.26.0rcfrom
Dawn952:perf/kimi-kda-preprocess-fusion
Open

Dawn952 wants to merge 2 commits into
vllm-project:releases/v0.26.0rcfrom
Dawn952:perf/kimi-kda-preprocess-fusion

Conversation

@Dawn952

@Dawn952 Dawn952 commented Aug 18, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Kimi K3 KDA applies B, F, and output-gate projections to the same hidden states. This revision removes the runtime F-A -> F-B projection chain and overlaps the complete float BFG path with the quantized QKV path.

This change:

  • composes checkpoint weights offline as f_proj.weight = f_b_proj.weight @ f_a_proj.weight;
  • keeps the original F-A/F-B source weights as small staging parameters so initial load and later weight reloads can rebuild the composed weight;
  • packs b_proj, the composed f_proj, and full-rank g_proj into one MergedColumnParallelLinear;
  • splits the merged output directly into beta, raw F gate, and output gate;
  • runs the BFG GEMM, split, beta.float().sigmoid(), and reshape vector work on one auxiliary NPU stream;
  • runs QKV MXFP8 dynamic quantization and quantized matmul on the main stream in parallel;
  • closes the dependency chain with main-stream input-ready event -> auxiliary wait -> auxiliary BFG-ready event -> main-stream wait before KDA;
  • leaves the legacy low-rank gate path unchanged when full-rank gating is disabled.

The checkpoint schema and API/configuration surface remain unchanged. The composed F projection changes BF16 rounding order compared with two sequential GEMMs, so model-level accuracy must be checked before merge.

Does this PR introduce any user-facing change?

No API or configuration change. Runtime behavior changes only for Kimi K3 KDA layers using the existing full-rank gate configuration.

How was this patch tested?

  • No new or updated UT was added for this revision, per the requested scope.

  • ruff check, ruff format --check, Python syntax compilation, and git diff --check passed for the changed source files.

  • One-card Ascend 950DT micro-validation on green-52-220 in kimi_k3_test passed using PR head 3b9321605:

    • composed F weight max abs error: 0.0;
    • merged BFG output vs composed-weight reference max abs error: 0.0;
    • MXFP8 dynamic quant + quantized QKV matmul completed on the main stream;
    • BFG cube/vector work completed on a distinct auxiliary stream with the full event wait/record closure;
    • beta/raw-gate/output-gate/QKV outputs had expected shapes and were finite.
  • The same random BF16 microcase measured max abs difference 2.0 between the composed one-GEMM F output and the original two-GEMM output. This is recorded as an expected rounding-order change, not an accuracy conclusion.

  • No vLLM service was started and no inference or benchmark request was sent.

  • vLLM version: v0.26.0

  • vLLM main: vllm-project/vllm@d02df74

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces a performance optimization for the Kimi K3 KDA implementation by fusing multiple float projection operations into a single matmul. By packing the B, F-A, and output-gate weights, the model reduces overhead during the projection phase while preserving existing sharding logic and checkpoint formats. The changes are internal to the model architecture and do not impact public APIs or configuration.

Highlights

  • Performance Optimization: Fused the float B, F-A, and output-gate projections into a single MergedColumnParallelLinear layer to reduce the number of independent matmul launches.
  • Checkpoint Management: Implemented a custom _KDAFusedBFGLinear class to handle sharding and loading of packed weights, ensuring proper handling of replicated F-A projections across TP ranks.
  • Compatibility: Maintained support for the legacy low-rank gate path when full-rank gating is disabled, ensuring no breaking changes for existing configurations.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Fuse KDA's B, F-A, and output-gate projections

Suggested PR Summary:

### What this PR does / why we need it?

This pull request introduces `_KDAFusedBFGLinear` to fuse KDA's float B, F-A, and output-gate projections into a single merged column-parallel linear layer when `use_full_rank_gate` is enabled. This optimization reduces the number of linear projections during the forward pass of `AscendKimiGatedDeltaNetAttention`. Additionally, weight loading logic in `kimi_k3.py` has been updated to support loading the fused BFG projection parameters.

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

The changes were tested using new unit tests added in `tests/ut/models/test_kimi_k3.py` and `tests/ut/ops/test_kimi_kda.py` verifying weight loading and output splitting for the fused BFG projection.

@Dawn952 Dawn952 changed the title [Performance][KDA] Fuse float gate projections [Performance][KDA] Compose and overlap gate projections Aug 21, 2026
Split MXFP dynamic quantization from the fused QKV matmul and overlap it with BFG projections on a dedicated NPU stream.

Add explicit event synchronization and stream-lifetime tracking, harden packed checkpoint routing, and cover the projection loading and staging behavior with unit tests.

Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
@Dawn952
Dawn952 force-pushed the perf/kimi-kda-preprocess-fusion branch from 52129fc to 7e8a99b Compare August 23, 2026 14:21
Compose the F projection from checkpoint weights during loading, pack B/F/G into one float matmul, and make the two overlap stages join through explicit QKV and BFG events. Harden full-rank checkpoint routing and cover global/local shards, reloads, DynamicQuant tuples, and event ordering.

Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant