Repository navigation
Conversation
Compose the upstream vLLM 0.27 Kimi text, multimodal, DSpark, and MTP contracts with the Ascend MLA, KDA, MoE, and quantization backends. Use explicit Kimi configuration fields, keep optional projector rotation state typed, and cover model-only adapter behavior without duplicating operator tests. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Route floating-point KDA gate shards into the packed gate projection and retain the Ascend DSpark model's quantization-aware per-layer KV projections through AutoWeightsLoader. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Keep the model child PR type-checkable before its attention and SiTU dependencies merge by documenting the two cross-PR imports with precise mypy error codes. Runtime import behavior is unchanged. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Give the empty capture list an explicit string element type so the focused loader regression passes strict mypy checks. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Exercise the production KDA/MLA geometry with a five-layer, sixteen-expert dummy fixture and verify block-size-plus-one prefix-cache parity, complete outputs, finite log probabilities, and FULL_DECODE_ONLY graph execution on a 16-NPU A3 runner. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Use a matrix-vector product for residual scores so profiling does not materialize a broadcasted FP32 tensor. Add a regression test that verifies the memory-efficient scoring path. Signed-off-by: q00852295 <qijiajin1@huawei.com>
maoxx241
force-pushed
the
codex/kimi-k3-model-adapters-main
branch
14 times, most recently
from
August 25, 2026 10:55
27d30b4 to
682050c
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this PR does / why we need it?
Adds one follow-up fix on top of vllm-project#14600.
The current attention-residual scoring path materializes a broadcasted FP32 tensor:
At
max_num_batched_tokens=24576, the temporary is 5.25 GiB and causes Kimi K3 workers to OOM duringdetermine_available_memory -> profile_run, before KV cache allocation.Use the mathematically equivalent matrix-vector product to avoid the temporary:
Does this PR introduce any user-facing change?
Yes. It prevents the observed Kimi K3 startup OOM without changing APIs or configuration.
How was this patch tested?
Added
test_ascend_attn_res_avoids_broadcast_score_product.Python syntax compilation passed.
git diff --checkpassed.Real-weight Ascend rerun at 24576 batched tokens is pending.
vLLM version: v0.27.1
vLLM main: vllm-project/vllm@58d3918