Skip to content

fix(kimi-k3): reduce attention residual peak memory - #1

Draft
qijiajin wants to merge 6 commits into
maoxx241:codex/kimi-k3-model-adapters-mainfrom
qijiajin:codex/fix-kimi3-attn-res-memory-only
Draft

qijiajin wants to merge 6 commits into
maoxx241:codex/kimi-k3-model-adapters-mainfrom
qijiajin:codex/fix-kimi3-attn-res-memory-only

Conversation

@qijiajin

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Adds one follow-up fix on top of vllm-project#14600.

The current attention-residual scoring path materializes a broadcasted FP32 tensor:

(normalized_without_gamma * score_weight).sum(-1)

At max_num_batched_tokens=24576, the temporary is 5.25 GiB and causes Kimi K3 workers to OOM during determine_available_memory -> profile_run, before KV cache allocation.

Use the mathematically equivalent matrix-vector product to avoid the temporary:

torch.matmul(normalized_without_gamma, score_weight)

Does this PR introduce any user-facing change?

Yes. It prevents the observed Kimi K3 startup OOM without changing APIs or configuration.

How was this patch tested?

  • Added test_ascend_attn_res_avoids_broadcast_score_product.

  • Python syntax compilation passed.

  • git diff --check passed.

  • Real-weight Ascend rerun at 24576 batched tokens is pending.

  • vLLM version: v0.27.1

  • vLLM main: vllm-project/vllm@58d3918

maoxx241 and others added 6 commits August 21, 2026 00:33
Compose the upstream vLLM 0.27 Kimi text, multimodal, DSpark, and MTP contracts with the Ascend MLA, KDA, MoE, and quantization backends.

Use explicit Kimi configuration fields, keep optional projector rotation state typed, and cover model-only adapter behavior without duplicating operator tests.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Route floating-point KDA gate shards into the packed gate projection and retain the Ascend DSpark model's quantization-aware per-layer KV projections through AutoWeightsLoader.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Keep the model child PR type-checkable before its attention and SiTU dependencies merge by documenting the two cross-PR imports with precise mypy error codes. Runtime import behavior is unchanged.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Give the empty capture list an explicit string element type so the focused loader regression passes strict mypy checks.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Exercise the production KDA/MLA geometry with a five-layer, sixteen-expert dummy fixture and verify block-size-plus-one prefix-cache parity, complete outputs, finite log probabilities, and FULL_DECODE_ONLY graph execution on a 16-NPU A3 runner.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Use a matrix-vector product for residual scores so profiling does not materialize a broadcasted FP32 tensor. Add a regression test that verifies the memory-efficient scoring path.

Signed-off-by: q00852295 <qijiajin1@huawei.com>
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-model-adapters-main branch 14 times, most recently from 27d30b4 to 682050c Compare August 25, 2026 10:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants