Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
50 commits
Select commit Hold shift + click to select a range
c4f6030
feat(ops): port Kimi K3 AscendC operators
maoxx241 Aug 21, 2026
0cc801a
refactor(ops): move KDA torch adapters into operator directories
maoxx241 Aug 25, 2026
80282f8
fix(mamba): align hybrid state copy on Ascend
maoxx241 Aug 24, 2026
814c2fd
feat(attention): add Kimi K3 KDA execution
maoxx241 Aug 24, 2026
37cf61c
feat(mla): support Kimi K3 attention on Ascend
maoxx241 Aug 24, 2026
3db03ec
perf(ops): add Kimi K3 attention residual fusion
maoxx241 Aug 24, 2026
88252bc
docs(ops): document K3 attention residual
maoxx241 Aug 25, 2026
7d21d38
fix(ops): enforce attention residual stream capacity
maoxx241 Aug 25, 2026
bdec553
feat(moe): support Kimi K3 SiTU execution
maoxx241 Aug 24, 2026
8fa59cd
refactor(moe): integrate SiTU quantization into common flow
maoxx241 Aug 25, 2026
dae7c13
feat(model): register Kimi K3 adapters
maoxx241 Aug 24, 2026
c2c3af5
feat(spec-decode): add Kimi K3 DSpark execution
maoxx241 Aug 24, 2026
b2aa0ec
feat(kv-cache): group Kimi K3 hybrid caches
maoxx241 Aug 24, 2026
8f2cb31
refactor(kv-cache): address K3 grouping review
maoxx241 Aug 25, 2026
0e83f71
fix(pd): harden Kimi K3 disaggregated serving
maoxx241 Aug 24, 2026
bf40dad
refactor(proxy): inline multimodal retry handling
maoxx241 Aug 25, 2026
60c6cfc
docs(kimi-k3): add deployment and validation guide
maoxx241 Aug 21, 2026
3f43cb8
fix(ops): use relative path for shared KDA adapter header
maoxx241 Aug 25, 2026
d533d5d
fix(model): load QuaRot DSpark boundaries in modeling
maoxx241 Aug 25, 2026
5b5512f
refactor(spec_decode): remove QuaRot model hooks
maoxx241 Aug 25, 2026
6c039a1
feat(ops): update fused chunk KDA forward
Aug 20, 2026
e888390
[Feature] Add mla_prolog_v3 and optional RoPE for MLA prolog (#13355)
zongersama Aug 4, 2026
20bd5a7
[Ops][Feature] Support first-axis non-contiguous kv/kr cache for mla_…
yolic66 Aug 5, 2026
cc00487
[Feature][MLA] Support Kimi K3 no-RoPE MLAPO on A5 (#13507)
Dawn952 Aug 13, 2026
1dc3630
chore(ops): normalize MLA prolog source line endings
Dawn952 Aug 25, 2026
adf50d7
fix(quantization): use forward config for MX scale selection
maoxx241 Aug 26, 2026
da6e071
fix(quantization): carry MX scale algorithm into forward context
maoxx241 Aug 26, 2026
9e8b8f7
fix(ci): cover Kimi KDA and format MX context helper
maoxx241 Aug 26, 2026
8a3fdcd
refactor(kimi-k3): remove operator-integration detours
maoxx241 Aug 26, 2026
ac63c66
fix(quantization): preserve ModelSlim optional metadata
maoxx241 Aug 26, 2026
eb11817
fix(ci): restore 310P unified-core output path
maoxx241 Aug 26, 2026
530884a
fix(kimi-k3): align CI coverage with vLLM versions
maoxx241 Aug 26, 2026
fb5c390
fix(ci): execute SiTU custom op eagerly in CPU tests
maoxx241 Aug 26, 2026
45804ed
fix(mla): preserve defaults for existing attention models
maoxx241 Aug 26, 2026
d7a93b2
fix(mla): default existing models to RoPE
maoxx241 Aug 26, 2026
342ee0c
fix(kv-cache): scope RoPE mode to MLA backends
maoxx241 Aug 26, 2026
2c83f01
refactor(mla): remove no-rope identity metadata
maoxx241 Aug 27, 2026
11bc0aa
fix(kv-cache): preserve grouped Mamba block capacity
maoxx241 Aug 27, 2026
548c0dc
test(kimi-k3): remove dummy execution parity nightly
maoxx241 Aug 27, 2026
4c6924c
fix(moe): compute SiTU without a runtime CustomOp
maoxx241 Aug 27, 2026
6da8cfa
fix(pd): preserve decode graphs for stateful handoffs
maoxx241 Aug 27, 2026
5e327cb
fix(mamba): preserve accepted snapshots and per-rank table capacity
maoxx241 Aug 27, 2026
4a96af9
fix(moe): normalize optional SiTU beta before activation
maoxx241 Aug 27, 2026
17d2491
test(moe): isolate SiTU reference configuration
maoxx241 Aug 27, 2026
b106ea1
fix(mla): preserve absent RoPE metadata during prefill
maoxx241 Aug 27, 2026
5dc5cb3
perf(kda): restore Ascend fused RMSNorm gate dispatch
maoxx241 Aug 27, 2026
3b6c845
fix(ci): register fused norm gate test duration
maoxx241 Aug 27, 2026
8c8deed
test(kimi-k3): cover A3 dummy deployment variants
maoxx241 Aug 27, 2026
b6dce8c
test(kimi-k3): streamline A3 smoke coverage
maoxx241 Aug 27, 2026
ba14b85
test(kimi-k3): prune redundant unit test scaffolding
maoxx241 Aug 27, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .github/workflows/scripts/runner_label.json
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,12 @@
"image_tag": "9.1.0-a3-ubuntu22.04-py3.12",
"csrc_cache_target": "a3-arm64-ubuntu"
},
"linux-aarch64-a3-16": {
"chip": "a3",
"npu_num": 16,
"image_tag": "9.1.0-a3-ubuntu22.04-py3.12",
"csrc_cache_target": "a3-arm64-ubuntu"
},
"linux-aarch64-a3-2-": {
"chip": "a3",
"npu_num": 2,
Expand Down
33 changes: 33 additions & 0 deletions .github/workflows/scripts/test_config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -308,6 +308,7 @@
- tests/ut/models/test_deepseek_v4_compressor.py
- tests/ut/models/test_deepseek_v4_indexer.py
- tests/ut/models/test_deepseek_v4_moe.py
- tests/ut/models/test_kimi_k3_adapter.py
- tests/e2e/pull_request/four_card/test_deepseek_v4.py

- name: models_minimax_m3
Expand All @@ -321,6 +322,30 @@
- tests/e2e/pull_request/one_card/test_minimax_m3_sparse_attn.py
- tests/e2e/pull_request/eight_card/test_minimax_m3.py

- name: models_kimi_k3
optional: true
source_file_dependencies:
- vllm_ascend/models/kimi_k3.py
- vllm_ascend/models/kimi_k3_dspark.py
- vllm_ascend/models/kimi_k3_mtp.py
- vllm_ascend/models/qwen3_dspark.py
- vllm_ascend/worker/model_runner_v1.py
- vllm_ascend/worker/worker.py
- vllm_ascend/attention/mla_v1.py
- vllm_ascend/core/kv_cache_interface.py
- vllm_ascend/ops/kimi_kda.py
- vllm_ascend/ops/fused_moe
- vllm_ascend/ops/triton/kimi_k3
- vllm_ascend/spec_decode/dspark_proposer.py
- vllm_ascend/spec_decode/llm_base_proposer.py
- vllm_ascend/quantization/modelslim_config.py
- vllm_ascend/patch/platform/patch_kv_cache_utils.py
- vllm_ascend/patch/platform/patch_kv_cache_coordinator.py
- vllm_ascend/patch/platform/patch_speculative_config.py
- vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py
tests:
- tests/e2e/pull_request/sixteen_card/test_kimi_k3.py

# Ops
- name: ops_basic
optional: false
Expand Down Expand Up @@ -391,6 +416,7 @@
source_file_dependencies:
- vllm_ascend/ops/gdn.py
- vllm_ascend/ops/gdn_attn_builder.py
- vllm_ascend/ops/kimi_kda.py
- vllm_ascend/ops/triton/gdn_chunk_meta.py
tests:
- tests/ut/ops
Expand Down Expand Up @@ -892,6 +918,7 @@ estimated_times:
tests/e2e/pull_request/four_card/test_qwen3_32b_bf16_ll_performance.py: 1200
tests/e2e/pull_request/two_card/test_gemma4.py: 420
tests/e2e/pull_request/eight_card/test_minimax_m3.py: 630
tests/e2e/pull_request/sixteen_card/test_kimi_k3.py: 360
tests/ut/attention/a2/test_attention_cp.py: 30
tests/ut/attention/a2/test_attention_cp_precision.py: 30
tests/ut/attention/a2/test_attention_v1.py: 30
Expand All @@ -913,6 +940,7 @@ estimated_times:
tests/ut/ops/a2/test_gdn_layerwise_kv.py: 70
tests/ut/ops/a2/test_token_dispatcher.py: 30
tests/ut/ops/a3_2/test_activation.py: 50
tests/ut/ops/a3_2/test_kimi_kda_fused_norm_gate.py: 60
tests/ut/ops/a3_2/test_select_experts.py: 20
tests/ut/quantization/methods/a2/test_w4a16.py: 30
tests/ut/quantization/methods/a2/test_w4a4_flatquant.py: 40
Expand Down Expand Up @@ -971,6 +999,8 @@ runner_mapping:
310p: 310p-4
tests/e2e/pull_request/eight_card:
default: a3-8
tests/e2e/pull_request/sixteen_card:
default: a3-16
tests/ut/.+/a3_2:
default: a3-2
tests/ut/.+/a2:
Expand Down Expand Up @@ -1004,6 +1034,9 @@ partition:
a3-8:
runner_label: linux-aarch64-a3-8-
count: 1
a3-16:
runner_label: linux-aarch64-a3-16
count: 1
cpu-0:
runner_label: linux-amd64-cpu-8-hk
count: 1
3 changes: 2 additions & 1 deletion .gitleaks.toml
Original file line number Diff line number Diff line change
Expand Up @@ -56,7 +56,8 @@ paths = ["^vllm-empty/"]
# Allow tilingKey generic-api-key false positives in the specified file
[[allowlists]]
paths = [
"^csrc/notify_dispatch/op_host/notify_dispatch_tiling.cpp"
"^csrc/notify_dispatch/op_host/notify_dispatch_tiling.cpp",
"^csrc/attention/mla_prolog_v3/op_kernel/mla_prolog_template_tiling_key.h"
]
rules = ["generic-api-key"]

Expand Down
2 changes: 1 addition & 1 deletion .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ repos:
args: [
--toml, pyproject.toml,
'--skip', 'tests/prompts/**,./benchmarks/sonnet.txt,*tests/lora/data/**,build/**,./vllm_ascend.egg-info/**,typos.toml',
'-L', 'CANN,cann,NNAL,nnal,ASCEND,ascend,EnQue,CopyIn,ArchType,AND,ND,tbe,copyin,alog,outter,mata,PARD'
'-L', 'CANN,cann,NNAL,nnal,ASCEND,ascend,EnQue,CopyIn,ArchType,AND,ND,tbe,copyin,alog,outter,mata,PARD,uSeed,LoadIn'
]
additional_dependencies:
- tomli
Expand Down
16 changes: 16 additions & 0 deletions csrc/attention/chunk_kda_fwd/CMakeLists.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# -----------------------------------------------------------------------------------------------------------
# Copyright (c) 2026 Tianjin University, Ltd.
# This program is free software, you can redistribute it and/or modify it under the terms and conditions of
# Please refer to the License for details. You may not use this file except in compliance with the License.
# THIS SOFTWARE IS PROVIDED ON AN "AS IS" BASIS, WITHOUT WARRANTIES OF ANY KIND, EITHER EXPRESS OR IMPLIED,
# INCLUDING BUT NOT LIMITED TO NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE.
# -----------------------------------------------------------------------------------------------------------
file(GLOB CURRENT_DIRS RELATIVE ${CMAKE_CURRENT_SOURCE_DIR} ${CMAKE_CURRENT_SOURCE_DIR}/*)
if(NOT ENABLE_TEST AND NOT BENCHMARK)
list(REMOVE_ITEM CURRENT_DIRS tests)
endif()
foreach(SUB_DIR ${CURRENT_DIRS})
if(EXISTS "${CMAKE_CURRENT_SOURCE_DIR}/${SUB_DIR}/CMakeLists.txt")
add_subdirectory(${SUB_DIR})
endif()
endforeach()
114 changes: 114 additions & 0 deletions csrc/attention/chunk_kda_fwd/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,114 @@
# ChunkKdaFwd

## 功能

`ChunkKdaFwd` 对齐不涉及 CP 切分的 FLA `chunk_kda_fwd` 顶层语义。公共接口接收 raw gate 或已激活的
自然对数 gate;Gate、Prepare、PostWu、FwdH 和 Finalize 均在一个物理 `ChunkKdaFwd` L0 内完成,
L2 不再拼接或依次发射多个阶段 L0。

Shape 符号与布局约定见 [KDA 模型符号表](../README.md#model-shape-symbols)。

## Gate 公式

令 `x = g + dt_bias`。逐 token、逐 K 维的自然对数衰减为:

```text
use_gate_in_kernel = false:
gate = g

use_gate_in_kernel = true, safe_gate = false:
gate = -exp(A_log) * softplus(x)

use_gate_in_kernel = true, safe_gate = true:
gate = lower_bound * sigmoid(exp(A_log) * x)
```

随后在每个 chunk 内计算:

```text
gk_i = cumsum(gate)_i / ln(2)
```

因此后续 `exp2(gk)` 与自然指数 gate 严格绑定,不暴露额外 gate scale。

## 输入

| 名称 | 必选性 | Shape/Dtype | 说明 |
| --- | --- | --- | --- |
| `q/k` | 必选 | 输入 layout 对应 Shape;FP16/BF16 | Query/Key |
| `v` | 必选 | 输入 layout 对应 Shape;与 q 同 dtype | Value |
| `g` | 必选 | 输入 layout 对应 K 维 Shape;FP32/BF16 | raw gate 或已激活自然对数 gate |
| `beta` | 必选 | 去掉 g 的 K 维;FP32/BF16 | Delta 系数 |
| `A_log` | 条件必选 | `[H_v]`,FP32 | `use_gate_in_kernel=true` 时必选 |
| `dt_bias` | 可选 | `[H_v*K]`,FP32 | gate bias |
| `initial_state` | 可选 | `[N,H_v,K,V]` 或 `[N,H_v,V,K]`,FP32 | 由 `state_v_first` 解释 |
| `cu_seqlens` | 可选 | `[N+1]`,INT64 | 变长序列 |
| `chunk_indices` | 可选 | `[2*N_c]`,INT64 | canonical chunk 顺序 |

`layout` 只描述上述输入。BSND/TND 由 L2 使用 `l0op::Transpose` 转为内部 BNSD/NTD。

## 输出

Python 返回顺序为:

```text
(attn_out, final_state, gk, Aqk, Akk, w, u, qg, kg, v_new, h, initial_state)
```

- `attn_out` 固定为 BSND/TND。
- `final_state` 固定按序列排列,末两维服从 `state_v_first`。
- `Aqk/Akk` 始终返回,固定为 head-major。
- `gk/w/u/qg/kg/v_new` 是供反向使用的 head-major 中间量。
- 公开 `h` 固定为 sequence-major;内部 `hCompute` 保持 head-major 供 Finalize 使用。
- 第 12 个返回值是 Python 层对 `initial_state` 的原对象透传,不是 aclnn 输出。

输出保留策略对齐 fla-org
[`chunk_kda_fwd`](https://github.com/fla-org/flash-linear-attention/blob/0f0f0c97af39343855b43bbbaddcedfda5cb9d77/fla/ops/kda/chunk_fwd.py)
提交 `0f0f0c97af39343855b43bbbaddcedfda5cb9d77`:

| 条件 | 返回 |
| --- | --- |
| `output_final_state=true` | 返回 `final_state`,否则为 `None` |
| `use_gate_in_kernel=false` 或 `disable_recompute=true` | 返回 `gk` |
| 始终 | 返回 `Aqk/Akk` |
| `disable_recompute=true` | 返回 `w/u/qg/kg/v_new` |
| `disable_recompute=true` 或 `return_intermediate_states=true` | 返回 `h` |

这是 `fla_npu.ops.ascendc.chunk_kda_fwd` 的低层 12 返回值语义;不涉及 CP。aclnn L2 不接收
`output_final_state/disable_recompute/return_intermediate_states`,每个可选输出是否写出仅由对应
输出指针是否为空决定。`w/u/qg/kg/v_new/h` 的 L0 阶段固定写内部 compute 张量,L2 仅在
对应指针非空时通过 `ViewCopy` 导出;`gkOut` 非空时直接复用为 `gkCompute`,避免目标场景
额外复制整张 FP32 gate。内部 `hCompute` 是 FwdH 到 Finalize 的必需 head-major 阶段结果;
公开 `hOut` 非空时,L2 转为 sequence-major 后导出。`hOut` 为空时仍创建 `hCompute`,但不
作为第 11 个 Python 返回值公开。

## 属性

| 名称 | 默认值 | 支持范围 |
| --- | --- | --- |
| `layout` | `BSND` | `BSND/BNSD/TND/NTD` |
| `scale` | 必传 | 通常为 `K**-0.5` |
| `chunk_size` | `64` | `64/128` |
| `output_final_state` | `false` | bool |
| `safe_gate` | `false` | bool |
| `lower_bound` | `-5.0` | safe raw gate 时 `[-5,0)` |
| `use_gate_in_kernel` | `false` | bool |
| `disable_recompute` | `false` | bool |
| `return_intermediate_states` | `false` | bool |
| `state_v_first` | `false` | bool |

## 支持范围

- A2 (`ascend910b`)、A3 (`ascend910_93`)、A5 (`ascend950`)。
- `K/V` 为 `[16,256]` 内 16 的倍数;交付重点覆盖 K=128、V=128/256。
- `chunk_size` 为 64/128。
- TND/NTD 均支持多 head。
- 变长调用最多 1024 条逻辑序列,rank-4 变长输入要求 B=1。

## 验证

唯一用例规格是 `tests/op_cases/chunk_kda_fwd.json`。数值测试位于
`tests/operators/chunk_kda_fwd/accuracy/`,性能使用 `tests/operators/chunk_kda_fwd/performance/profile.py`
和 `msopprof`。

完整 API 见 [API 文档](docs/api.md),阶段和内存设计见 [设计文档](docs/design.md)。
Loading
Loading