Repository navigation
Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request integrates Kimi K3 model support for the Ascend platform. It provides the necessary infrastructure to map Kimi K3 architectures—including multimodal and speculative decoding variants—to Ascend-optimized implementations. By reusing vLLM 0.27's native configuration and processing logic, the changes ensure compatibility while enabling high-performance execution on Ascend hardware. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Ignored Files
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request registers and implements Kimi K3 model adapters for vLLM on Ascend, including text, multimodal, MTP, and DSpark architectures, along with comprehensive unit tests. The review feedback identifies a critical runtime error due to an unsupported return_bias argument in ReplicatedLinear, potential AttributeErrors when directly accessing configuration attributes like activation_situ_beta and rope_parameters, and a violation of the repository's style guide regarding the PR title and summary format.
cc26cc0 to
8eca530
Compare
ff20bf7 to
7bd46ca
Compare
7420dc3 to
a7a97b5
Compare
b0b47b5 to
60ff97f
Compare
4a4bc2f to
8863238
Compare
Current-head GQA DSpark performance resultsTested sources:
Runtime: full W4A8 target plus GQA DSpark draft, DP4 x TP16 x EP64 on four Atlas A3 nodes, 7 draft tokens plus 1 target token, Prefix Cache enabled, All waves satisfying the 30 ms mean-TPOT gateEvery row below passed HTTP 200, exact prompt/output token counts,
The maximum passing per-DP concurrency for 2K/4K/8K/16K/32K/64K/128K/256K is respectively 8/8/8/8/4/2/1/none. Sanitized server-side launch scriptThe script follows the four-node deployment format in #!/usr/bin/env bash
set -euo pipefail
if [[ $# -ne 1 ]]; then
echo "usage: $0 <dp-rank:0..3>" >&2
exit 2
fi
export DP_START_RANK=$1
export MODEL_PATH="<KIMI_K3_FULL_W4A8_PATH>"
export DRAFT_MODEL_PATH="<KIMI_K3_GQA_DSPARK_PATH>"
export LOCAL_IP="<CURRENT_NODE_IP>"
export NODE0_IP="<NODE0_IP>"
export NIC_NAME="<CURRENT_NODE_NIC>"
export SERVICE_PORT="<SERVICE_PORT>"
export RPC_PORT="<DP_RPC_PORT>"
export HCCL_IF_IP=$LOCAL_IP
export GLOO_SOCKET_IFNAME=$NIC_NAME
export TP_SOCKET_IFNAME=$NIC_NAME
export HCCL_SOCKET_IFNAME=$NIC_NAME
export VLLM_HOST_IP=$LOCAL_IP
export VLLM_ENGINE_READY_TIMEOUT_S=7200
export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=7200
export VLLM_USE_V1=1
export VLLM_USE_V2_MODEL_RUNNER=0
export VLLM_WORKER_MULTIPROC_METHOD=spawn
export VLLM_RANDOMIZE_DP_DUMMY_INPUTS=0
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
export HCCL_BUFFSIZE=512
export HCCL_OP_EXPANSION_MODE=AIV
export OMP_PROC_BIND=false
export OMP_NUM_THREADS=1
export OPENBLAS_NUM_THREADS=1
export TASK_QUEUE_ENABLE=1
export HCCL_CONNECT_TIMEOUT=1800
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15
vllm serve "$MODEL_PATH" \
--host 0.0.0.0 \
--port "$SERVICE_PORT" \
--served-model-name kimi-k3-gqa-dspark \
--enable-request-id-headers \
--trust-remote-code \
--tokenizer-mode kimi_k3 \
--enable-auto-tool-choice \
--reasoning-parser kimi_k3 \
--tool-call-parser kimi_k3 \
--data-parallel-size 4 \
--data-parallel-size-local 1 \
--data-parallel-start-rank "$DP_START_RANK" \
--data-parallel-address "$NODE0_IP" \
--data-parallel-rpc-port "$RPC_PORT" \
--tensor-parallel-size 16 \
--enable-expert-parallel \
--all2all-backend allgather_reducescatter \
--dtype bfloat16 \
--quantization ascend \
--safetensors-load-strategy lazy \
--max-model-len 270000 \
--max-num-seqs 16 \
--max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.80 \
--async-scheduling \
--enable-prefix-caching \
--speculative-config \
'{"method":"dspark","model":"'"$DRAFT_MODEL_PATH"'","num_speculative_tokens":7,"draft_tensor_parallel_size":16,"max_model_len":4096,"draft_sample_method":"greedy","enforce_eager":true}' \
--compilation-config \
'{"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[16,32,48,64,80,96,112,128],"cudagraph_num_of_warmups":1}' \
--additional-config \
'{"enable_cpu_binding":false,"enable_shared_expert_dp":false,"multistream_overlap_shared_expert":true,"enable_fused_mc2":0,"enable_reduce_sample":false,"finegrained_tp_config":{"lmhead_tensor_parallel_size":0}}' \
--mm-processor-cache-gb 0 \
--mm-encoder-tp-mode data \
--skip-mm-profiling \
--limit-mm-per-prompt '{"image":1}' |
QuaRot GQA DSpark performance resultsThis is a new QuaRot-only result set. It does not replace the earlier non-rotating W4A8 comment.
The QuaRot checkpoint provides Both result sets use the same four-node DP4 x TP16 x EP64 topology, GQA DSpark with 7 draft tokens plus 1 target token, Prefix Cache, This is not a pure kernel-overhead A/B: the two checkpoints can produce different target-token trajectories, which changes DSpark acceptance and therefore TPOT. The tables retain acceptance metrics so that throughput is not interpreted independently of draft quality. QuaRot waves satisfying the 30 ms mean-TPOT gateThe main sweep completed 26 waves and all 400/400 requests passed HTTP, exact token-count,
QuaRot versus non-rotating maximum passing concurrency
The conservative QuaRot maximum passing per-DP concurrency for 2K/4K/8K/16K/32K/64K/128K/256K is therefore 8/4/none/4/4/4/2/none. Sanitized server-side launch scriptRun once on each node with a unique DP rank. Before launch, the QuaRot checkpoint's quantization descriptor must map #!/usr/bin/env bash
set -euo pipefail
export DP_START_RANK=$1
export MODEL_PATH="<KIMI_K3_FULL_QUAROT_PATH>"
export DRAFT_MODEL_PATH="<KIMI_K3_GQA_DSPARK_PATH>"
export LOCAL_IP="<CURRENT_NODE_IP>"
export NODE0_IP="<NODE0_IP>"
export NIC_NAME="<CURRENT_NODE_NIC>"
export SERVICE_PORT="<SERVICE_PORT>"
export RPC_PORT="<DP_RPC_PORT>"
export HCCL_IF_IP=$LOCAL_IP
export GLOO_SOCKET_IFNAME=$NIC_NAME
export TP_SOCKET_IFNAME=$NIC_NAME
export HCCL_SOCKET_IFNAME=$NIC_NAME
export VLLM_HOST_IP=$LOCAL_IP
export VLLM_ENGINE_READY_TIMEOUT_S=7200
export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=7200
export VLLM_USE_V2_MODEL_RUNNER=0
export VLLM_WORKER_MULTIPROC_METHOD=spawn
export VLLM_RANDOMIZE_DP_DUMMY_INPUTS=0
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
export HCCL_BUFFSIZE=512
export HCCL_OP_EXPANSION_MODE=AIV
export OMP_PROC_BIND=false
export OMP_NUM_THREADS=1
export OPENBLAS_NUM_THREADS=1
export TASK_QUEUE_ENABLE=1
export HCCL_CONNECT_TIMEOUT=1800
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15
vllm serve "$MODEL_PATH" \
--host 0.0.0.0 \
--port "$SERVICE_PORT" \
--served-model-name kimi-k3-full-quarot-gqa-dspark \
--enable-request-id-headers \
--trust-remote-code \
--tokenizer-mode kimi_k3 \
--enable-auto-tool-choice \
--reasoning-parser kimi_k3 \
--tool-call-parser kimi_k3 \
--data-parallel-size 4 \
--data-parallel-size-local 1 \
--data-parallel-start-rank "$DP_START_RANK" \
--data-parallel-address "$NODE0_IP" \
--data-parallel-rpc-port "$RPC_PORT" \
--tensor-parallel-size 16 \
--enable-expert-parallel \
--all2all-backend allgather_reducescatter \
--dtype bfloat16 \
--quantization ascend \
--safetensors-load-strategy lazy \
--max-model-len 270000 \
--max-num-seqs 16 \
--max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.80 \
--async-scheduling \
--enable-prefix-caching \
--speculative-config \
'{"method":"dspark","model":"'"$DRAFT_MODEL_PATH"'","num_speculative_tokens":7,"draft_tensor_parallel_size":16,"max_model_len":4096,"draft_sample_method":"greedy","enforce_eager":true}' \
--compilation-config \
'{"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[16,32,48,64,80,96,112,128],"cudagraph_num_of_warmups":1}' \
--additional-config \
'{"enable_cpu_binding":false,"enable_shared_expert_dp":false,"multistream_overlap_shared_expert":true,"enable_fused_mc2":0,"enable_reduce_sample":false,"finegrained_tp_config":{"lmhead_tensor_parallel_size":0}}' \
--mm-processor-cache-gb 0 \
--mm-encoder-tp-mode data \
--skip-mm-profiling \
--limit-mm-per-prompt '{"image":1}' |
8863238 to
d2ff31e
Compare
d2ff31e to
cb4e9fb
Compare
Implementation walkthroughArchitecture registrationvllm_ascend/models/init.py maps the Kimi K3 text, multimodal, MTP, and DSpark architecture names to the Ascend adapters. Configuration, processor, renderer, parser, and checkpoint naming continue to come from upstream vLLM. Target modelvllm_ascend/models/kimi_k3.py composes each decoder layer from Ascend KDA or MLA attention, the existing fused-MoE path with SiTU parameters, the attention-residual fusion, and the upstream Kimi layer structure. Sequence-parallel sharding occurs before residual allocation, so block_residual remains rank-local and gathers happen only at the defined consumer boundaries. The loader remaps mixed KDA gate projections into the packed runtime parameters and leaves quantization-specific processing to the selected linear methods. DSpark target hidden statesThe target model captures the configured auxiliary layer outputs for DSpark. GQA DSpark consumes the materialized residual form; MLA DSpark consumes the regular hidden-state form. model_runner_v1.py selects this mode from the draft architecture and routes the auxiliary states to the proposer. Draft, QuaRot, and MTP models
Nightly coverageThe reduced five-layer/16-expert fixture exercises TP/EP construction, text execution, graph replay, cold and prefix-hit paths, reset behavior, block-size-plus-one boundaries, complete output, and finite selected-token log probabilities without requiring the full checkpoint. |
75b5f27 to
cf5764e
Compare
72ee680 to
27d30b4
Compare
Compose the upstream Kimi text, multimodal, MTP, and DSpark model structures with Ascend KDA, MLA, MoE, quantization, sequence-parallel, and auxiliary-state contracts. Keep ViT FIA key/value inputs contiguous at the operator boundary, matching the validated v0.26 path. Add a reduced nightly execution fixture and focused model tests. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
27d30b4 to
682050c
Compare
Load missing target embed and LM-head shards during DSpark weight loading, anti-rotate them in the draft basis, and mark the draft-owned boundaries explicitly. This keeps model-specific QuaRot handling out of the proposer while reading only the local TP vocab slice. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
|
QuaRot DSpark boundary loading update:
Validation: targeted model/spec UT suite passed (46 tests total together with PR #14601), and the full QuaRot K3 + GQA DSpark service matched the established DP0 c2 acceptance baseline: 58.17% current versus 58.35% baseline. |
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | #14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | #14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | #14598 | KDA attention execution and fused RMSNorm gate | | 4 | #14839 | MLA attention and rotary execution | | 5 | #14840 | Attention-residual Triton fusion | | 6 | #14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | #14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | #14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | #14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | #14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
|
Closing this child PR following the merge of parent #14454. |
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate | | 4 | vllm-project#14839 | MLA attention and rotary execution | | 5 | vllm-project#14840 | Attention-residual Triton fusion | | 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: d30086105 <denghaojie1@h-partners.com>
### What this PR does / why we need it? This integration PR enables Kimi K3 text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. The implementation is reviewed through the atomic PRs below; this parent owns the Kimi K3 deployment and validation guide. ### Recommended merge order | Order | PR | Responsibility | | ---: | --- | --- | | 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and MX-SiTU operators | | 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token snapshots, and Ascend launch-grid correctness | | 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate | | 4 | vllm-project#14839 | MLA attention and rotary execution | | 5 | vllm-project#14840 | Attention-residual Triton fusion | | 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim quantization adaptation | | 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT FIA contiguous inputs; model-owned QuaRot shared-layer conversion | | 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic shared-layer hook integration | | 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative capacity, and per-rank DCP table sizing | | 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe stateful handoffs | After these PRs merge, the parent-owned change is: - `docs/source/tutorials/models/Kimi-K3.md` - `docs/source/tutorials/models/index.md` The guide covers reduced and full checkpoints, single-node TP16, four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix Cache, server-side validation, GPQA, and performance reporting. ### How was this patch tested? - Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1 and the pinned upstream revision. Coverage includes FP16/BF16, sigmoid/SiLU, packed gate strides, residual/prenorm, and input preservation. The actual upstream CustomOp resolves to the Ascend fused implementation; three ACLGraph replays with fresh inputs match upstream native results. CI mypy and `bash format.sh ci` pass. A5 performance and full-model serving were not rerun for this change. - State and capacity regressions: 94 targeted CPU tests pass, with two post-v0.27.1 coordinator-API cases skipped. Coverage includes real InputBatch replacement/reordering, accepted-token ownership across scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker capacity agreement, and writes to the final speculative Mamba slots at DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not rerun for the snapshot/capacity changes. - Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata tests pass. Distributed NPU end-to-end validation was not rerun for this graph-selection change. - All changed Python files pass syntax compilation. - The parent includes the current child implementations plus the deployment documentation. - Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy, DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff, Mooncake transfer, and model registration. - Full-checkpoint integration coverage includes text, multimodal, tools, streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond. Detailed accuracy and performance results remain in the PR comments. ### Does this PR introduce any user-facing change? Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix Cache, and P/D serving on Ascend. - vLLM version: v0.27.1 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com> Signed-off-by: weinachuan <weinachuan1@huawei.com> Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com> Signed-off-by: yolic66 <747731294@qq.com> Signed-off-by: Dawn952 <zhaojunbo13@huawei.com> Signed-off-by: MQ <maomaoyu870@gmail.com> Co-authored-by: weinachuan <weinachuan1@huawei.com> Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: yolic66 <747731294@qq.com> Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
What this PR does / why we need it?
Dependencies
Merge after #14426, #14597, #14598, #14839, #14840, and #14599.
How was this patch tested?
git diff --checkDoes this PR introduce any user-facing change?
Yes. Kimi K3 checkpoints resolve to the Ascend text, multimodal, MTP, and DSpark model implementations.