Repository navigation
[NPU] support kimi k3 on A5 and improve performance - #39589
Merged
Merged
Conversation
Signed-off-by: zhaozx-cn <zhaozx2116@163.com>
…tion-TP all-gather and reduce-scatter on the current stream.
Signed-off-by: zhaozx-cn <zhaozx2116@163.com>
Co-authored-by: qyb233 <{"message":"Not Found","documentation_url":"https://docs.github.com/rest/users/emails#list-email-addresses-for-the-authenticated-user","status":"404"}>
* fix(npu): use DeepEP MXFP8 dispatch for W4A8 MXFP MoE * feat(npu): fuse K3 SiTU MXFP8 quantization --------- Co-authored-by: zhaozx-cn <59479021+zhaozx-cn@users.noreply.github.com>
Replace the multi-kernel Ascend KDA extend path (chunk_local_cumsum, chunk_kda_scaled_dot_kkt_fwd, solve_tril_npu, recompute_w_u_fwd_npu, chunk_gated_delta_rule_fwd_h_npu, chunk_gla_fwd_o_gk_npu) with a single torch.ops.npu.chunk_kda_fwd call. The fused op takes one initial state per logical sequence in contiguous [N, H, V, K] layout, while SGLang owns a slot-indexed persistent pool, so the wrapper gathers the active slots before the call and scatters final_state back afterwards. chunk_states comes back in the layout the tracker already expects, which drops the transpose on the intermediate-state path. Forward metadata uses -1 in cache_indices for a padded request. The gather would turn that into a read of the last cache slot, and index_copy_ would then write the padded row's final_state back over that same slot, corrupting the state of whichever live sequence owned it. Clamp the gather indices to 0 so a padded row reads a harmless placeholder, and filter padded rows out of the writeback entirely. Drop the speculative-decoding-only NPU transpose of temporal_state in MambaPool. It presented the pool as a transposed view so the stride-aware verify kernel and state movers saw the canonical [pool, HV, V, K] logical shape; the fused extend path consumes the pool directly in that layout, so every consumer now agrees on one contiguous representation. The adjacent shape comment described the swapped layout and is corrected with it. Also rewrite the conv_states copy for non-tracked entries as a no-op self-copy, so it no longer runs bool-mask indexing (aten::nonzero) or a host numel check on every step.
Unpack PA-NZ latent and RoPE cache pages before prefix projection and attention concatenation. Ordinary MLA writes can use NZ independently of MLAPO, so raw page gathers otherwise mix token and feature dimensions. Add CPU regression coverage for ND/NZ layouts, latent/RoPE dimensions, reordered and repeated pages, empty selections, and partial-page writes. Validated test bodies and forward_extend inputs in an isolated CPU PyTorch harness, including a legacy-path negative control; syntax and targeted Ruff checks passed. A5 GSM8K and NPU execution remain to be validated.
Signed-off-by: zhaozx-cn <zhaozx2116@163.com>
This reverts commit e83efb3.
Signed-off-by: zhaozx-cn <zhaozx2116@163.com>
zhaozx-cn
requested review from
Qiaolin-Yu,
Ying1123,
hanming-lu,
hnyls2002,
merrymercy and
xiezhq-hermann
as code owners
September 15, 2026 09:48
sglang-npu-bot
approved these changes
Sep 16, 2026
Contributor
Author
|
/rerun-failed-ci |
Contributor
Author
|
/rerun-failed-ci |
Contributor
Author
|
/rerun-failed-ci |
1 similar comment
Contributor
Author
|
/rerun-failed-ci |
Contributor
Author
|
/rerun-failed-ci |
Collaborator
|
Before commit 4ab4643, we had already passed all test cases. For details, please refer to...https://github.com/sgl-project/sglang/actions/runs/35219963083/job/105312395665?pr=39589. Commit 4ab4643 only modifies python/sglang/srt/environ.py to resolve conflicts. The evaluation does not affect functionality, so this PR is merged. If there are any issues, please leave a comment under this PR, and we will fix it as soon as possible. Thank you. |
sglang-npu-bot
approved these changes
Sep 18, 2026
gjsheu
added a commit
to gjsheu/sglang
that referenced
this pull request
Sep 20, 2026
…rify The NPU transpose in MambaPool's speculative state was removed in sgl-project#39589, so on NPU the spec Mamba pool now hands out the untransposed `temporal` tensor while the rest of the NPU verify path still works in the [.., HV, K, V] convention: * AscendGDNAttnBackend / AscendHybridLinearAttnBackend reshape the per-draft scratch as (-1, num_value_heads, head_k_dim, head_v_dim); * move_intermediate_cache() (sgl_kernel_npu), which commits the state of the accepted prefix back into the live SSM state after target verify, walks each (heads, K, V) block with a hardcoded element order. Handing out the untransposed tensor flips that order, so the state written back after the first verify step is transposed. Everything generated after it is garbage (word salad), while plain decode is unaffected because it never touches the per-draft scratch or re-commits a state. Because GDN uses head_k_dim == head_v_dim, the transposition changes no shape and no byte count, so nothing crashes or asserts -- it silently produces wrong tokens. Verified on Qwen3.6-35B-A3B (Ascend 910, CANN 9.1.0-a3-B070, TP=2, --speculative-algorithm NEXTN, gamma=4): before: "The szo periódzkod ... " (word salad), accept rate 0.01-0.07 after : "The capital of France is **Paris**." , accept rate 0.14-0.66 KDA keeps the [.., HV, V, K] layout introduced together with it.
gjsheu
added a commit
to gjsheu/sglang
that referenced
this pull request
Sep 20, 2026
…c verify sgl-project#39589 dropped the NPU transpose in MambaPool's speculative state while migrating the Kimi-K3 (KDA) path to the canonical contiguous [.., HV, V, K] layout. KDA is self-consistent there -- ascend_kda_backend.py switches its kernels to that layout and its commit kernel, move_intermediate_cache_kda(), is stride aware -- but the GDN path was left on the [.., HV, K, V] convention: * AscendGDNAttnBackend / AscendHybridLinearAttnBackend reshape the per-draft scratch as (-1, num_value_heads, head_k_dim, head_v_dim); * move_intermediate_cache() (sgl_kernel_npu), which commits the state of the accepted prefix back into the live SSM state after target verify, walks each (heads, K, V) block with a hardcoded element order and is not stride aware. So the state written back after the first verify step is transposed, and every token generated after it is garbage (word salad). Plain decode is unaffected because it never touches the per-draft scratch nor commits a state back. GDN uses head_k_dim == head_v_dim, so the flip changes no shape and no byte count: nothing crashes and no assertion fires. Verified on Qwen3.6-35B-A3B (Ascend 910, CANN 9.1.0-a3-B070, TP=2, --speculative-algorithm NEXTN, gamma=4): before: "The szo periódzkod ... " (word salad), accept rate 0.01-0.07 after : "The capital of France is **Paris**." , accept rate 0.14-0.66 KDA (Kimi-K3) is deliberately left on its untransposed layout; no K3 checkpoint was available on the verification machine, so that path is guarded rather than retested.
TallMessiWu
added a commit
to TallMessiWu/sglang
that referenced
this pull request
Sep 20, 2026
Drop the graph-rebind fix: upstream sgl-project#39589 fixes the same race, and more completely. Both versions order the input rebind before the replay -- ours by doing it on the calling thread, upstream's by blocking on the future -- but upstream reuses one device-bound worker whose executor initializer calls set_device, instead of creating a thread per replay. Take upstream's file whole; nothing of ours is left to carry. The other ten files merged without conflicts.
5 tasks done
5 tasks
6 tasks done
yuychang
added a commit
to yuychang/sglang
that referenced
this pull request
Sep 30, 2026
- Upstream sgl-project#39589 narrowed do_fuse_qkvbfg to quant_config is None and attn_tp == tp - The branch was built when it was attn_tp == tp and (quant_config is None or use_full_rank_gate) - Under the new definition Quark K3 checkpoints turn off the in-proj merge, group-64 and PTPC merge gates - Gate those three ROCm-only paths on a new _attn_tp_is_full_tp (attn_tp == tp) to restore the validated behavior - The loader and the low-rank fused path keep upstream's do_fuse_qkvbfg unchanged
yuychang
added a commit
to yuychang/sglang
that referenced
this pull request
Oct 1, 2026
- Upstream sgl-project#39589 narrowed do_fuse_qkvbfg to quant_config is None and attn_tp == tp - The branch was built when it was attn_tp == tp and (quant_config is None or use_full_rank_gate) - Under the new definition Quark K3 checkpoints turn off the in-proj merge, group-64 and PTPC merge gates - Gate those three ROCm-only paths on a new _attn_tp_is_full_tp (attn_tp == tp) to restore the validated behavior - The loader and the low-rank fused path keep upstream's do_fuse_qkvbfg unchanged
3 of 5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Co-Authored-By: hanwlax
Co-Authored-By: Hexq0210
Co-Authored-By: McZyWu
Co-Authored-By: qybnb
Co-Authored-By: sherdavincl9
Motivation
support kimi k3 on A5 and improve performance.
Modifications
1.add compressed w4a8 mxfp4 moe.
2.shared expert: add fine-grained dual steam and support shared expert specified tp size.
3.add fused qkvg proj.
4.add situ mx quant kernel.
5.add kv nz for k3 mla.
6.add fia v2 for mtp branch.
7.add chunk kda kernel for prefill branch.
8.fix kimi k3 dspark acc pd disaggregation.
Accuracy Tests
Speed Tests and Profiling
128k in 1k out dspark 7 tp32 ep 32
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): 🚫 Run #35314528985
Latest PR Test (Extra): ❌ Run #35314528785
Latest PR Test (AMD ROCm 10): ❌ Run #35314528997