Repository navigation
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces a performance optimization for the Kimi K3 KDA implementation by fusing multiple float projection operations into a single matmul. By packing the B, F-A, and output-gate weights, the model reduces overhead during the projection phase while preserving existing sharding logic and checkpoint formats. The changes are internal to the model architecture and do not impact public APIs or configuration. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Ops][Feature] Fuse KDA's B, F-A, and output-gate projectionsSuggested PR Summary:
### What this PR does / why we need it?
This pull request introduces `_KDAFusedBFGLinear` to fuse KDA's float B, F-A, and output-gate projections into a single merged column-parallel linear layer when `use_full_rank_gate` is enabled. This optimization reduces the number of linear projections during the forward pass of `AscendKimiGatedDeltaNetAttention`. Additionally, weight loading logic in `kimi_k3.py` has been updated to support loading the fused BFG projection parameters.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
The changes were tested using new unit tests added in `tests/ut/models/test_kimi_k3.py` and `tests/ut/ops/test_kimi_kda.py` verifying weight loading and output splitting for the fused BFG projection.Split MXFP dynamic quantization from the fused QKV matmul and overlap it with BFG projections on a dedicated NPU stream. Add explicit event synchronization and stream-lifetime tracking, harden packed checkpoint routing, and cover the projection loading and staging behavior with unit tests. Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
52129fc to
7e8a99b
Compare
Compose the F projection from checkpoint weights during loading, pack B/F/G into one float matmul, and make the two overlap stages join through explicit QKV and BFG events. Harden full-rank checkpoint routing and cover global/local shards, reloads, DynamicQuant tuples, and event ordering. Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
What this PR does / why we need it?
Kimi K3 KDA applies B, F, and output-gate projections to the same hidden states. This revision removes the runtime F-A -> F-B projection chain and overlaps the complete float BFG path with the quantized QKV path.
This change:
f_proj.weight = f_b_proj.weight @ f_a_proj.weight;b_proj, the composedf_proj, and full-rankg_projinto oneMergedColumnParallelLinear;beta.float().sigmoid(), and reshape vector work on one auxiliary NPU stream;The checkpoint schema and API/configuration surface remain unchanged. The composed F projection changes BF16 rounding order compared with two sequential GEMMs, so model-level accuracy must be checked before merge.
Does this PR introduce any user-facing change?
No API or configuration change. Runtime behavior changes only for Kimi K3 KDA layers using the existing full-rank gate configuration.
How was this patch tested?
No new or updated UT was added for this revision, per the requested scope.
ruff check,ruff format --check, Python syntax compilation, andgit diff --checkpassed for the changed source files.One-card Ascend 950DT micro-validation on green-52-220 in
kimi_k3_testpassed using PR head3b9321605:0.0;0.0;The same random BF16 microcase measured max abs difference
2.0between the composed one-GEMM F output and the original two-GEMM output. This is recorded as an expected rounding-order change, not an accuracy conclusion.No vLLM service was started and no inference or benchmark request was sent.
vLLM version: v0.26.0
vLLM main: vllm-project/vllm@d02df74