Skip to content

[Feature][Model] Register Kimi K3 model adapters - #14600

Closed
maoxx241 wants to merge 2 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-model-adapters-main
Closed

maoxx241 wants to merge 2 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-model-adapters-main

Conversation

@maoxx241

@maoxx241 maoxx241 commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

  • Registers Kimi K3 text, multimodal, MTP, and DSpark model architectures for Ascend.
  • Keeps ViT FIA key and value inputs contiguous at the operator boundary for multimodal execution.
  • Composes the Ascend decoder from the KDA, MLA, SiTU MoE, and attention-residual implementations.
  • Loads merged KDA gates and quantization-aware DSpark context projections through model loading hooks.
  • Loads and retains the QuaRot matrix in K3/Qwen DSpark modeling, then prepares draft-owned unrotated embedding and LM-head weights when those layers must be shared with the target.
  • Selects the GQA or MLA DSpark auxiliary execution stream in the model runner.
  • Adds a five-layer, 16-expert nightly fixture and focused adapter tests.

Dependencies

Merge after #14426, #14597, #14598, #14839, #14840, and #14599.

How was this patch tested?

  • git diff --check
  • Python syntax compilation for changed implementation and test files
  • Focused model-adapter, QuaRot shared-layer, and model-runner tests
  • Nightly execution-parity coverage for TP/EP, decode graphs, prefix hits, reset behavior, block-size boundaries, complete output, and finite token log probabilities
  • Integrated text, multimodal, tools, streaming, QuaRot, Prefix Cache, and ACLGraph serving coverage

Does this PR introduce any user-facing change?

Yes. Kimi K3 checkpoints resolve to the Ascend text, multimodal, MTP, and DSpark model implementations.

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request integrates Kimi K3 model support for the Ascend platform. It provides the necessary infrastructure to map Kimi K3 architectures—including multimodal and speculative decoding variants—to Ascend-optimized implementations. By reusing vLLM 0.27's native configuration and processing logic, the changes ensure compatibility while enabling high-performance execution on Ascend hardware.

Highlights

  • Model Registration: Registered Kimi K3 model architectures, including text, multimodal, MTP, and DSpark variants, into the vLLM Ascend model registry.
  • Architecture Implementation: Implemented Ascend-specific decoder layers, MoE, and MLA attention modules, leveraging native KDA and SiTU backends.
  • Multimodal and Speculative Support: Added support for multimodal projector rotation handling and integrated MTP and DSpark wrappers for speculative decoding.
  • Testing: Added comprehensive unit tests for adapter-level model components and operator contracts to ensure hybrid runtime correctness.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/scripts/test_config.yaml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request registers and implements Kimi K3 model adapters for vLLM on Ascend, including text, multimodal, MTP, and DSpark architectures, along with comprehensive unit tests. The review feedback identifies a critical runtime error due to an unsupported return_bias argument in ReplicatedLinear, potential AttributeErrors when directly accessing configuration attributes like activation_situ_beta and rope_parameters, and a violation of the repository's style guide regarding the PR title and summary format.

Comment thread vllm_ascend/models/kimi_k3_dspark.py
Comment thread vllm_ascend/models/kimi_k3.py
Comment thread vllm_ascend/models/kimi_k3.py
Comment thread vllm_ascend/models/__init__.py
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-model-adapters-main branch 5 times, most recently from cc26cc0 to 8eca530 Compare August 20, 2026 08:19
@maoxx241
maoxx241 marked this pull request as ready for review August 21, 2026 01:31
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-model-adapters-main branch from ff20bf7 to 7bd46ca Compare August 21, 2026 05:33
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-model-adapters-main branch from 7420dc3 to a7a97b5 Compare August 21, 2026 11:58
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-model-adapters-main branch 3 times, most recently from b0b47b5 to 60ff97f Compare August 22, 2026 10:46
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-model-adapters-main branch 3 times, most recently from 4a4bc2f to 8863238 Compare August 23, 2026 07:40
@maoxx241

Copy link
Copy Markdown
Contributor Author

Current-head GQA DSpark performance results

Tested sources:

Runtime: full W4A8 target plus GQA DSpark draft, DP4 x TP16 x EP64 on four Atlas A3 nodes, 7 draft tokens plus 1 target token, Prefix Cache enabled, max_model_len=270000, max_num_seqs=16, and decode graph capture sizes [16,32,48,64,80,96,112,128]. Requests were issued server-side. Outputs were fixed at 1024 tokens and common prefixes were page-aligned to 384-token pages.

All waves satisfying the 30 ms mean-TPOT gate

Every row below passed HTTP 200, exact prompt/output token counts, finish_reason=length, non-empty/non-garbled output, and mean TPOT <= 30 ms. The threshold applies to the mean, while p90 is retained separately. In total, 22 waves and 284/284 requests passed.

Input/Output Prefix Per-DP c (global) Graph TTFT mean TPOT mean / p90 Output tok/s Linearity Mean accepted length Token acceptance Prefix hit Requests
2K/1K 1920 (93.75%) 1 (4) 16 1013.74 ms 22.25 / 34.95 ms 111.4 1.000 5.223 60.33% 75.00% 4/4
2K/1K 1920 (93.75%) 2 (8) 16 29937.14 ms 18.65 / 21.44 ms 139.2 0.625 6.420 77.43% 75.00% 8/8
2K/1K 1920 (93.75%) 4 (16) 32 29187.91 ms 18.94 / 19.58 ms 276.2 0.620 7.183 88.32% 75.00% 16/16
2K/1K 1920 (93.75%) 8 (32) 64 29480.01 ms 21.97 / 22.84 ms 528.0 0.592 7.225 88.93% 75.00% 32/32
4K/1K 3840 (93.75%) 1 (4) 16 1264.36 ms 15.41 / 16.75 ms 222.6 1.000 7.475 92.51% 84.38% 4/4
4K/1K 3840 (93.75%) 2 (8) 16 1393.63 ms 15.78 / 16.33 ms 430.1 0.966 7.529 93.28% 84.38% 8/8
4K/1K 3840 (93.75%) 4 (16) 32 29253.09 ms 17.84 / 19.52 ms 330.4 0.371 7.471 92.44% 84.38% 16/16
4K/1K 3840 (93.75%) 8 (32) 64 29981.23 ms 21.45 / 22.12 ms 474.8 0.267 7.216 88.80% 84.38% 32/32
8K/1K 7680 (93.75%) 1 (4) 16 2101.42 ms 18.52 / 28.56 ms 130.8 1.000 6.335 76.22% 89.06% 4/4
8K/1K 7680 (93.75%) 2 (8) 16 2319.66 ms 17.21 / 17.71 ms 287.4 1.099 6.995 85.64% 89.06% 8/8
8K/1K 7680 (93.75%) 4 (16) 32 2209.14 ms 17.83 / 19.77 ms 729.8 1.395 7.589 94.12% 89.06% 16/16
8K/1K 7680 (93.75%) 8 (32) 64 2934.66 ms 21.14 / 22.13 ms 899.6 0.860 7.412 91.60% 89.06% 32/32
16K/1K 15360 (93.75%) 1 (4) 16 2401.58 ms 15.53 / 16.14 ms 216.4 1.000 7.616 94.51% 91.41% 4/4
16K/1K 15360 (93.75%) 2 (8) 16 2370.83 ms 15.98 / 16.34 ms 418.4 0.967 7.602 94.32% 91.41% 8/8
16K/1K 15360 (93.75%) 4 (16) 32 2604.56 ms 18.21 / 19.06 ms 736.5 0.851 7.467 92.39% 91.41% 16/16
16K/1K 15360 (93.75%) 8 (32) 64 3879.17 ms 21.47 / 22.42 ms 1193.2 0.689 7.524 93.21% 91.41% 32/32
32K/1K 31104 (94.92%) 1 (4) 16 2137.41 ms 23.17 / 30.66 ms 122.2 1.000 5.148 59.26% 93.75% 4/4
32K/1K 31104 (94.92%) 2 (8) 16 2468.94 ms 19.98 / 29.93 ms 231.7 0.948 6.128 73.25% 93.75% 8/8
32K/1K 31104 (94.92%) 4 (16) 32 3373.21 ms 20.05 / 29.96 ms 479.6 0.981 6.941 84.88% 93.75% 16/16
64K/1K 62208 (94.92%) 1 (4) 16 2728.33 ms 25.03 / 28.99 ms 126.4 1.000 4.861 55.16% 94.34% 4/4
64K/1K 62208 (94.92%) 2 (8) 16 3634.54 ms 27.56 / 31.56 ms 216.7 0.857 4.594 51.34% 94.34% 8/8
128K/1K 124416 (94.92%) 1 (4) 16 4003.77 ms 28.56 / 33.82 ms 106.0 1.000 4.334 47.62% 94.63% 4/4

The maximum passing per-DP concurrency for 2K/4K/8K/16K/32K/64K/128K/256K is respectively 8/8/8/8/4/2/1/none.

Sanitized server-side launch script

The script follows the four-node deployment format in docs/source/tutorials/models/Kimi-K3.md. Run it once on each node with a unique DP rank. All four rank-local APIs stay enabled because the benchmark sends requests directly to each DP endpoint.

#!/usr/bin/env bash
set -euo pipefail

if [[ $# -ne 1 ]]; then
  echo "usage: $0 <dp-rank:0..3>" >&2
  exit 2
fi

export DP_START_RANK=$1
export MODEL_PATH="<KIMI_K3_FULL_W4A8_PATH>"
export DRAFT_MODEL_PATH="<KIMI_K3_GQA_DSPARK_PATH>"
export LOCAL_IP="<CURRENT_NODE_IP>"
export NODE0_IP="<NODE0_IP>"
export NIC_NAME="<CURRENT_NODE_NIC>"
export SERVICE_PORT="<SERVICE_PORT>"
export RPC_PORT="<DP_RPC_PORT>"

export HCCL_IF_IP=$LOCAL_IP
export GLOO_SOCKET_IFNAME=$NIC_NAME
export TP_SOCKET_IFNAME=$NIC_NAME
export HCCL_SOCKET_IFNAME=$NIC_NAME
export VLLM_HOST_IP=$LOCAL_IP
export VLLM_ENGINE_READY_TIMEOUT_S=7200
export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=7200
export VLLM_USE_V1=1
export VLLM_USE_V2_MODEL_RUNNER=0
export VLLM_WORKER_MULTIPROC_METHOD=spawn
export VLLM_RANDOMIZE_DP_DUMMY_INPUTS=0
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
export HCCL_BUFFSIZE=512
export HCCL_OP_EXPANSION_MODE=AIV
export OMP_PROC_BIND=false
export OMP_NUM_THREADS=1
export OPENBLAS_NUM_THREADS=1
export TASK_QUEUE_ENABLE=1
export HCCL_CONNECT_TIMEOUT=1800
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15

vllm serve "$MODEL_PATH" \
  --host 0.0.0.0 \
  --port "$SERVICE_PORT" \
  --served-model-name kimi-k3-gqa-dspark \
  --enable-request-id-headers \
  --trust-remote-code \
  --tokenizer-mode kimi_k3 \
  --enable-auto-tool-choice \
  --reasoning-parser kimi_k3 \
  --tool-call-parser kimi_k3 \
  --data-parallel-size 4 \
  --data-parallel-size-local 1 \
  --data-parallel-start-rank "$DP_START_RANK" \
  --data-parallel-address "$NODE0_IP" \
  --data-parallel-rpc-port "$RPC_PORT" \
  --tensor-parallel-size 16 \
  --enable-expert-parallel \
  --all2all-backend allgather_reducescatter \
  --dtype bfloat16 \
  --quantization ascend \
  --safetensors-load-strategy lazy \
  --max-model-len 270000 \
  --max-num-seqs 16 \
  --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.80 \
  --async-scheduling \
  --enable-prefix-caching \
  --speculative-config \
    '{"method":"dspark","model":"'"$DRAFT_MODEL_PATH"'","num_speculative_tokens":7,"draft_tensor_parallel_size":16,"max_model_len":4096,"draft_sample_method":"greedy","enforce_eager":true}' \
  --compilation-config \
    '{"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[16,32,48,64,80,96,112,128],"cudagraph_num_of_warmups":1}' \
  --additional-config \
    '{"enable_cpu_binding":false,"enable_shared_expert_dp":false,"multistream_overlap_shared_expert":true,"enable_fused_mc2":0,"enable_reduce_sample":false,"finegrained_tp_config":{"lmhead_tensor_parallel_size":0}}' \
  --mm-processor-cache-gb 0 \
  --mm-encoder-tp-mode data \
  --skip-mm-profiling \
  --limit-mm-per-prompt '{"image":1}'

@maoxx241

Copy link
Copy Markdown
Contributor Author

QuaRot GQA DSpark performance results

This is a new QuaRot-only result set. It does not replace the earlier non-rotating W4A8 comment.

The QuaRot checkpoint provides global_rotation through its quantization descriptor. Runtime logs confirmed that the DSpark proposer entered the QuaRot path and aligned the shared draft embedding and LM-head weights.

Both result sets use the same four-node DP4 x TP16 x EP64 topology, GQA DSpark with 7 draft tokens plus 1 target token, Prefix Cache, max_model_len=270000, max_num_seqs=16, and decode graph sizes [16,32,48,64,80,96,112,128]. Requests were issued server-side. Outputs were fixed at 1024 tokens and common prefixes were page-aligned to 384-token pages.

This is not a pure kernel-overhead A/B: the two checkpoints can produce different target-token trajectories, which changes DSpark acceptance and therefore TPOT. The tables retain acceptance metrics so that throughput is not interpreted independently of draft quality.

QuaRot waves satisfying the 30 ms mean-TPOT gate

The main sweep completed 26 waves and all 400/400 requests passed HTTP, exact token-count, finish_reason=length, and non-empty/non-garbled-output checks. Of these, 18 waves and 184/184 requests also satisfied mean TPOT <= 30 ms.

Input/Output Prefix Per-DP c (global) Graph TTFT mean TPOT mean / p90 Output tok/s Linearity Mean accepted length Token acceptance Prefix hit Requests
2K/1K 1920 (93.75%) 1 (4) 16 1072.25 ms 21.80 / 29.53 ms 130.9 1.000 5.322 61.74% 75.00% 4/4
2K/1K 1920 (93.75%) 2 (8) 16 1195.12 ms 19.69 / 28.22 ms 244.8 0.935 6.051 72.16% 75.00% 8/8
2K/1K 1920 (93.75%) 4 (16) 32 1233.31 ms 21.50 / 25.69 ms 466.5 0.891 6.225 74.65% 75.00% 16/16
2K/1K 1920 (93.75%) 8 (32) 64 1619.05 ms 26.39 / 36.03 ms 759.9 0.726 5.845 69.21% 75.00% 32/32
4K/1K 3840 (93.75%) 1 (4) 16 1121.31 ms 17.17 / 18.14 ms 208.1 1.000 6.682 81.18% 84.38% 4/4
4K/1K 3840 (93.75%) 2 (8) 16 1308.62 ms 22.88 / 25.78 ms 200.6 0.482 5.202 60.03% 84.38% 8/8
4K/1K 3840 (93.75%) 4 (16) 32 1162.88 ms 27.04 / 47.27 ms 315.4 0.379 4.846 54.94% 84.38% 16/16
16K/1K 15360 (93.75%) 1 (4) 16 2215.82 ms 16.43 / 16.98 ms 209.1 1.000 7.159 87.98% 91.41% 4/4
16K/1K 15360 (93.75%) 2 (8) 16 2328.25 ms 17.30 / 18.00 ms 358.8 0.858 6.969 85.28% 91.41% 8/8
16K/1K 15360 (93.75%) 4 (16) 32 2545.33 ms 20.36 / 21.44 ms 483.6 0.578 6.650 80.71% 91.41% 16/16
32K/1K 31104 (94.92%) 1 (4) 16 2151.76 ms 17.75 / 20.51 ms 176.9 1.000 6.628 80.41% 93.75% 4/4
32K/1K 31104 (94.92%) 2 (8) 16 3306.75 ms 22.48 / 30.55 ms 233.3 0.659 5.480 64.00% 93.75% 8/8
32K/1K 31104 (94.92%) 4 (16) 32 3238.40 ms 23.03 / 32.64 ms 404.4 0.571 5.974 71.06% 93.75% 16/16
64K/1K 62208 (94.92%) 1 (4) 16 2600.37 ms 23.51 / 42.04 ms 89.8 1.000 5.087 58.39% 94.34% 4/4
64K/1K 62208 (94.92%) 2 (8) 16 3628.42 ms 25.81 / 42.02 ms 143.7 0.800 4.865 55.22% 94.34% 8/8
64K/1K 62208 (94.92%) 4 (16) 32 5846.35 ms 24.01 / 30.14 ms 423.1 1.178 6.243 74.91% 94.34% 16/16
128K/1K 124416 (94.92%) 1 (4) 16 4011.85 ms 22.02 / 33.87 ms 105.9 1.000 5.565 65.22% 94.63% 4/4
128K/1K 124416 (94.92%) 2 (8) 16 36096.72 ms 28.51 / 35.15 ms 91.3 0.431 5.021 57.44% 94.63% 8/8

QuaRot versus non-rotating maximum passing concurrency

Input/Output Non-rotating per-DP max QuaRot per-DP max Difference
2K/1K 8 8 same
4K/1K 8 4 QuaRot one tested level lower
8K/1K 8 none stable QuaRot c1 measured 30.31 / 29.91 / 36.18 ms across the main run and two repeats
16K/1K 8 4 QuaRot c8 measured 30.21 / 57.59 / 51.09 ms across the main run and two repeats
32K/1K 4 4 same
64K/1K 2 4 QuaRot one tested level higher, with different DSpark acceptance
128K/1K 1 2 QuaRot one tested level higher, but c2 linearity was 0.431
256K/1K none none both c1 waves exceeded 30 ms

The conservative QuaRot maximum passing per-DP concurrency for 2K/4K/8K/16K/32K/64K/128K/256K is therefore 8/4/none/4/4/4/2/none.

Sanitized server-side launch script

Run once on each node with a unique DP rank. Before launch, the QuaRot checkpoint's quantization descriptor must map global_rotation to its bundled rotation safetensor.

#!/usr/bin/env bash
set -euo pipefail

export DP_START_RANK=$1
export MODEL_PATH="<KIMI_K3_FULL_QUAROT_PATH>"
export DRAFT_MODEL_PATH="<KIMI_K3_GQA_DSPARK_PATH>"
export LOCAL_IP="<CURRENT_NODE_IP>"
export NODE0_IP="<NODE0_IP>"
export NIC_NAME="<CURRENT_NODE_NIC>"
export SERVICE_PORT="<SERVICE_PORT>"
export RPC_PORT="<DP_RPC_PORT>"

export HCCL_IF_IP=$LOCAL_IP
export GLOO_SOCKET_IFNAME=$NIC_NAME
export TP_SOCKET_IFNAME=$NIC_NAME
export HCCL_SOCKET_IFNAME=$NIC_NAME
export VLLM_HOST_IP=$LOCAL_IP
export VLLM_ENGINE_READY_TIMEOUT_S=7200
export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=7200
export VLLM_USE_V2_MODEL_RUNNER=0
export VLLM_WORKER_MULTIPROC_METHOD=spawn
export VLLM_RANDOMIZE_DP_DUMMY_INPUTS=0
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
export HCCL_BUFFSIZE=512
export HCCL_OP_EXPANSION_MODE=AIV
export OMP_PROC_BIND=false
export OMP_NUM_THREADS=1
export OPENBLAS_NUM_THREADS=1
export TASK_QUEUE_ENABLE=1
export HCCL_CONNECT_TIMEOUT=1800
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15

vllm serve "$MODEL_PATH" \
  --host 0.0.0.0 \
  --port "$SERVICE_PORT" \
  --served-model-name kimi-k3-full-quarot-gqa-dspark \
  --enable-request-id-headers \
  --trust-remote-code \
  --tokenizer-mode kimi_k3 \
  --enable-auto-tool-choice \
  --reasoning-parser kimi_k3 \
  --tool-call-parser kimi_k3 \
  --data-parallel-size 4 \
  --data-parallel-size-local 1 \
  --data-parallel-start-rank "$DP_START_RANK" \
  --data-parallel-address "$NODE0_IP" \
  --data-parallel-rpc-port "$RPC_PORT" \
  --tensor-parallel-size 16 \
  --enable-expert-parallel \
  --all2all-backend allgather_reducescatter \
  --dtype bfloat16 \
  --quantization ascend \
  --safetensors-load-strategy lazy \
  --max-model-len 270000 \
  --max-num-seqs 16 \
  --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.80 \
  --async-scheduling \
  --enable-prefix-caching \
  --speculative-config \
    '{"method":"dspark","model":"'"$DRAFT_MODEL_PATH"'","num_speculative_tokens":7,"draft_tensor_parallel_size":16,"max_model_len":4096,"draft_sample_method":"greedy","enforce_eager":true}' \
  --compilation-config \
    '{"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[16,32,48,64,80,96,112,128],"cudagraph_num_of_warmups":1}' \
  --additional-config \
    '{"enable_cpu_binding":false,"enable_shared_expert_dp":false,"multistream_overlap_shared_expert":true,"enable_fused_mc2":0,"enable_reduce_sample":false,"finegrained_tp_config":{"lmhead_tensor_parallel_size":0}}' \
  --mm-processor-cache-gb 0 \
  --mm-encoder-tp-mode data \
  --skip-mm-profiling \
  --limit-mm-per-prompt '{"image":1}'

@maoxx241

maoxx241 commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor Author

Implementation walkthrough

Architecture registration

vllm_ascend/models/init.py maps the Kimi K3 text, multimodal, MTP, and DSpark architecture names to the Ascend adapters. Configuration, processor, renderer, parser, and checkpoint naming continue to come from upstream vLLM.

Target model

vllm_ascend/models/kimi_k3.py composes each decoder layer from Ascend KDA or MLA attention, the existing fused-MoE path with SiTU parameters, the attention-residual fusion, and the upstream Kimi layer structure. Sequence-parallel sharding occurs before residual allocation, so block_residual remains rank-local and gathers happen only at the defined consumer boundaries.

The loader remaps mixed KDA gate projections into the packed runtime parameters and leaves quantization-specific processing to the selected linear methods.

DSpark target hidden states

The target model captures the configured auxiliary layer outputs for DSpark. GQA DSpark consumes the materialized residual form; MLA DSpark consumes the regular hidden-state form. model_runner_v1.py selects this mode from the draft architecture and routes the auxiliary states to the proposer.

Draft, QuaRot, and MTP models

  • kimi_k3_dspark.py builds the MLA draft decoder and loads its context projections through the normal quantization-aware model loader.
  • Qwen3 and Kimi DSpark loaders retain the global QuaRot matrix while transforming their own fc/context projection. The modeling-owned prepare_shared_layer hook creates draft-owned embedding and lm_head projections in the required hidden basis, then releases the retained matrix after sharing.
  • kimi_k3_mtp.py retains the upstream MTP structure and interfaces.
  • The multimodal wrapper reuses the upstream Kimi projector and protocol, and FIA key/value tensors are made contiguous at the operator boundary.

Nightly coverage

The reduced five-layer/16-expert fixture exercises TP/EP construction, text execution, graph replay, cold and prefix-hit paths, reset behavior, block-size-plus-one boundaries, complete output, and finite selected-token log probabilities without requiring the full checkpoint.

@maoxx241
maoxx241 force-pushed the codex/kimi-k3-model-adapters-main branch 3 times, most recently from 75b5f27 to cf5764e Compare August 25, 2026 04:53
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-model-adapters-main branch 3 times, most recently from 72ee680 to 27d30b4 Compare August 25, 2026 10:48
Compose the upstream Kimi text, multimodal, MTP, and DSpark model structures with Ascend KDA, MLA, MoE, quantization, sequence-parallel, and auxiliary-state contracts. Keep ViT FIA key/value inputs contiguous at the operator boundary, matching the validated v0.26 path. Add a reduced nightly execution fixture and focused model tests.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@maoxx241
maoxx241 force-pushed the codex/kimi-k3-model-adapters-main branch from 27d30b4 to 682050c Compare August 25, 2026 10:55
Load missing target embed and LM-head shards during DSpark weight loading, anti-rotate them in the draft basis, and mark the draft-owned boundaries explicitly. This keeps model-specific QuaRot handling out of the proposer while reading only the local TP vocab slice.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@maoxx241

Copy link
Copy Markdown
Contributor Author

QuaRot DSpark boundary loading update:

  • The model loader now owns all QuaRot-specific boundary preparation.
  • When the draft checkpoint omits embed_tokens or lm_head, it locates the corresponding tensor through the target safetensors index, reads only the local TP vocabulary rows, and applies the inverse hidden-basis rotation.
  • Draft checkpoints that already contain either boundary keep their trained weights unchanged.
  • The loader marks synthesized boundaries as draft-owned so the generic proposer does not alias the rotated target modules.
  • Added coverage for real safetensors shard selection, TP row slicing, padding, rotation, and both Qwen3/K3 DSpark loader lifecycles.

Validation: targeted model/spec UT suite passed (46 tests total together with PR #14601), and the full QuaRot K3 + GQA DSpark service matched the established DP0 c2 acceptance baseline: 58.17% current versus 58.35% baseline.

linfeng-yuan pushed a commit that referenced this pull request Aug 28, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | #14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | #14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | #14598 | KDA attention execution and fused RMSNorm gate |
| 4 | #14839 | MLA attention and rotary execution |
| 5 | #14840 | Attention-residual Triton fusion |
| 6 | #14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | #14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | #14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | #14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | #14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
@maoxx241

Copy link
Copy Markdown
Contributor Author

Closing this child PR following the merge of parent #14454.

@maoxx241 maoxx241 closed this Aug 28, 2026
ASH-XING pushed a commit to ASH-XING/vllm-ascend that referenced this pull request Aug 28, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate |
| 4 | vllm-project#14839 | MLA attention and rotary execution |
| 5 | vllm-project#14840 | Attention-residual Triton fusion |
| 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: d30086105 <denghaojie1@h-partners.com>
Lethobenthos20 pushed a commit to Lethobenthos20/vllm-ascend that referenced this pull request Sep 4, 2026
### What this PR does / why we need it?

This integration PR enables Kimi K3 text, multimodal, MTP, DSpark,
Prefix Cache, and P/D serving on Ascend. The implementation is reviewed
through the atomic PRs below; this parent owns the Kimi K3 deployment
and validation guide.

### Recommended merge order

| Order | PR | Responsibility |
| ---: | --- | --- |
| 1 | vllm-project#14426 | AscendC KDA, chunk gated delta rule, Dequant-SiTU, and
MX-SiTU operators |
| 2 | vllm-project#14597 | Hybrid Mamba state-copy, asynchronous accepted-token
snapshots, and Ascend launch-grid correctness |
| 3 | vllm-project#14598 | KDA attention execution and fused RMSNorm gate |
| 4 | vllm-project#14839 | MLA attention and rotary execution |
| 5 | vllm-project#14840 | Attention-residual Triton fusion |
| 6 | vllm-project#14599 | SiTU MoE, shared-expert execution, and K3 ModelSlim
quantization adaptation |
| 7 | vllm-project#14600 | Text, multimodal, MTP, and DSpark model registration; ViT
FIA contiguous inputs; model-owned QuaRot shared-layer conversion |
| 8 | vllm-project#14601 | DSpark speculative-decoding runtime and generic
shared-layer hook integration |
| 9 | vllm-project#14765 | TP8/TP16 GQA/MLA DSpark KV grouping, speculative
capacity, and per-rank DCP table sizing |
| 10 | vllm-project#14602 | Hybrid P/D transfer, proxy retry, and graph-safe
stateful handoffs |

After these PRs merge, the parent-owned change is:

- `docs/source/tutorials/models/Kimi-K3.md`
- `docs/source/tutorials/models/index.md`

The guide covers reduced and full checkpoints, single-node TP16,
four-node DP4/TP16/EP64, GQA/MLA DSpark, two-node P/D, QuaRot, Prefix
Cache, server-side validation, GPQA, and performance reporting.

### How was this patch tested?

- Fused norm-gate dispatch: 39 targeted CPU tests pass (4 existing
skips) and 22 real A3 NPU numerical cases pass on each of vLLM v0.27.1
and the pinned upstream revision. Coverage includes FP16/BF16,
sigmoid/SiLU, packed gate strides, residual/prenorm, and input
preservation. The actual upstream CustomOp resolves to the Ascend fused
implementation; three ACLGraph replays with fresh inputs match upstream
native results. CI mypy and `bash format.sh ci` pass. A5 performance and
full-model serving were not rerun for this change.

- State and capacity regressions: 94 targeted CPU tests pass, with two
post-v0.27.1 coordinator-API cases skipped. Coverage includes real
InputBatch replacement/reordering, accepted-token ownership across
scheduling modes, GDN metadata, Mamba copy ordering, scheduler/worker
capacity agreement, and writes to the final speculative Mamba slots at
DCP1/DCP4. `bash format.sh ci` passes. Full-model NPU serving was not
rerun for the snapshot/capacity changes.

- Handoff graph selection: 12 CPU dispatch cases and 25 GDN metadata
tests pass. Distributed NPU end-to-end validation was not rerun for this
graph-selection change.

- All changed Python files pass syntax compilation.
- The parent includes the current child implementations plus the
deployment documentation.
- Focused coverage includes KDA, MLA, SiTU MoE, Mamba state copy,
DSpark, compressed/hybrid KV cache, Prefix Cache, one-token P/D handoff,
Mooncake transfer, and model registration.
- Full-checkpoint integration coverage includes text, multimodal, tools,
streaming, QuaRot, C64/C128, TP8/TP16, two-node P/D, four-node GQA
DSpark, the known K3 accuracy/hang cases, and GPQA-Diamond.

Detailed accuracy and performance results remain in the PR comments.

### Does this PR introduce any user-facing change?

Yes. Kimi K3 can be deployed with text, multimodal, MTP, DSpark, Prefix
Cache, and P/D serving on Ascend.

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Signed-off-by: weinachuan <weinachuan1@huawei.com>
Signed-off-by: zongersama <48584200+zongersama@users.noreply.github.com>
Signed-off-by: yolic66 <747731294@qq.com>
Signed-off-by: Dawn952 <zhaojunbo13@huawei.com>
Signed-off-by: MQ <maomaoyu870@gmail.com>
Co-authored-by: weinachuan <weinachuan1@huawei.com>
Co-authored-by: zongersama <48584200+zongersama@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: yolic66 <747731294@qq.com>
Co-authored-by: Dawn952 <zhaojunbo13@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant