Repository navigation
[Cherry-pick][releases/v0.27.1rc][Performance][KDA] Compose and overlap gate projections on main (from #15416) - #15869
Conversation
Cherry-pick of PR vllm-project#15416 by @Dawn952
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request optimizes the performance of the Kimi K3 Delta Attention (KDA) mechanism by introducing a fused BFG projection and a multi-stream execution schedule. By composing F projection weights and overlapping auxiliary stream operations—such as DynamicQuant and QKV GEMM—with gate processing, the implementation reduces latency and improves throughput. These changes also include updates to weight loading and beta preprocessing to support the new fused architecture. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Ops][Feature] Fuse BFG projections and overlap with QKV dynamic quantization for Kimi K3Suggested PR Summary:
### What this PR does / why we need it?
This PR refactors the Kimi K3 attention projection by introducing `_KDAFusedBFGLinear` to fuse the B, F, and G projections into a single linear layer. It also implements `_run_overlapped_qkv_bfg` to run the BFG projection on an auxiliary NPU stream, allowing it to overlap with the QKV dynamic quantization on the main stream. This optimizes performance by enabling parallel execution of vector and cube operations.
Feedback: An issue was identified in `_KDAFusedBFGLinear` where `self.tp_rank` is used during weight loading but is never initialized in `__init__`, which will cause an `AttributeError` at runtime.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Added unit tests in `tests/ut/ops/test_kimi_kda.py` and updated existing tests in `tests/ut/models/test_kimi_k3_adapter.py`.| if self.tp_size != tp_size: | ||
| raise ValueError(f"KDA fused BFG TP mismatch: layer={self.tp_size}, attention={tp_size}") |
There was a problem hiding this comment.
The self.tp_rank attribute is accessed in _load_f_b_weight (line 124) but is never initialized in _KDAFusedBFGLinear.__init__. Since MergedColumnParallelLinear (and its parent ColumnParallelLinear in vLLM) only sets self.tp_size but not self.tp_rank, this will raise an AttributeError at runtime when loading weights.
We should initialize self.tp_rank in __init__ using get_tensor_model_parallel_rank() from vllm.distributed.
if self.tp_size != tp_size:
raise ValueError(f"KDA fused BFG TP mismatch: layer={self.tp_size}, attention={tp_size}")
from vllm.distributed import get_tensor_model_parallel_rank
self.tp_rank = get_tensor_model_parallel_rank()|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
Cherry-pick of PR #15416 onto
releases/v0.27.1rc.Original PR: #15416
Original author: @Dawn952
What this PR does / why we need it?
This recreates #15168 without the unrelated
mla_v1.pyhead-padding change and its corresponding test changes.For the existing mixed-precision Kimi K3 KDA layout, this change:
f_proj.weight = f_b_proj.weight @ f_a_proj.weightafter checkpoint loading and later source-weight reloads;scale_algwhen DynamicQuant is split from the linear method;