Skip to content

[Cherry-pick][releases/v0.27.1rc][Performance][KDA] Compose and overlap gate projections on main (from #15416) - #15869

Open
vllm-ascend-ci wants to merge 1 commit into
vllm-project:releases/v0.27.1rcfrom
vllm-ascend-ci:cherry-pick/pr-15416-to-releases-v0.27.1rc
Open

vllm-ascend-ci wants to merge 1 commit into
vllm-project:releases/v0.27.1rcfrom
vllm-ascend-ci:cherry-pick/pr-15416-to-releases-v0.27.1rc

Conversation

@vllm-ascend-ci

@vllm-ascend-ci vllm-ascend-ci commented Sep 6, 2026 •

Copy link
Copy Markdown
Collaborator

Cherry-pick of PR #15416 onto releases/v0.27.1rc.

Original PR: #15416
Original author: @Dawn952


What this PR does / why we need it?

This recreates #15168 without the unrelated mla_v1.py head-padding change and its corresponding test changes.

For the existing mixed-precision Kimi K3 KDA layout, this change:

  • composes f_proj.weight = f_b_proj.weight @ f_a_proj.weight after checkpoint loading and later source-weight reloads;
  • packs beta, the composed F projection, and the output gate into one floating-point BFG projection;
  • runs a two-stage auxiliary-stream schedule:
    1. main-stream MXFP DynamicQuant overlaps the auxiliary BFG GEMM;
    2. main-stream QKV GEMM overlaps auxiliary B/F/G split, beta FP32 sigmoid, and gate reshaping;
  • marks mixed-path beta as preprocessed so the v0.27 dispatch path only slices it, while ordinary upstream raw beta retains its existing FP32 sigmoid path;
  • preserves the configured W8A8 MXFP8 scale_alg when DynamicQuant is split from the linear method;
  • routes global or local F shards through the v0.27 packed loader;
  • leaves
  • vLLM main: vllm-project/vllm@ba07e4a

@github-actions

github-actions Bot commented Sep 6, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request optimizes the performance of the Kimi K3 Delta Attention (KDA) mechanism by introducing a fused BFG projection and a multi-stream execution schedule. By composing F projection weights and overlapping auxiliary stream operations—such as DynamicQuant and QKV GEMM—with gate processing, the implementation reduces latency and improves throughput. These changes also include updates to weight loading and beta preprocessing to support the new fused architecture.

Highlights

  • Performance Optimization: Implemented a two-stage auxiliary-stream schedule to overlap DynamicQuant and QKV GEMM operations with BFG projection processing.
  • Weight Composition: Added logic to compose F projection weights and pack them with beta and output gate into a single fused BFG projection.
  • Beta Preprocessing: Updated _prepare_beta to support preprocessed beta, avoiding redundant sigmoid operations in the auxiliary stream.
  • Model Refactoring: Refactored AscendKimiK3DeltaAttention to utilize the new fused BFG projection and updated checkpoint mapping in AscendKimiLinearModel.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Fuse BFG projections and overlap with QKV dynamic quantization for Kimi K3

Suggested PR Summary:

### What this PR does / why we need it?
This PR refactors the Kimi K3 attention projection by introducing `_KDAFusedBFGLinear` to fuse the B, F, and G projections into a single linear layer. It also implements `_run_overlapped_qkv_bfg` to run the BFG projection on an auxiliary NPU stream, allowing it to overlap with the QKV dynamic quantization on the main stream. This optimizes performance by enabling parallel execution of vector and cube operations.

Feedback: An issue was identified in `_KDAFusedBFGLinear` where `self.tp_rank` is used during weight loading but is never initialized in `__init__`, which will cause an `AttributeError` at runtime.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Added unit tests in `tests/ut/ops/test_kimi_kda.py` and updated existing tests in `tests/ut/models/test_kimi_k3_adapter.py`.

Comment on lines +73 to +74
if self.tp_size != tp_size:
raise ValueError(f"KDA fused BFG TP mismatch: layer={self.tp_size}, attention={tp_size}")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

The self.tp_rank attribute is accessed in _load_f_b_weight (line 124) but is never initialized in _KDAFusedBFGLinear.__init__. Since MergedColumnParallelLinear (and its parent ColumnParallelLinear in vLLM) only sets self.tp_size but not self.tp_rank, this will raise an AttributeError at runtime when loading weights.

We should initialize self.tp_rank in __init__ using get_tensor_model_parallel_rank() from vllm.distributed.

        if self.tp_size != tp_size:
            raise ValueError(f"KDA fused BFG TP mismatch: layer={self.tp_size}, attention={tp_size}")
        from vllm.distributed import get_tensor_model_parallel_rank
        self.tp_rank = get_tensor_model_parallel_rank()

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant