[Kimi-K3] Allow DSPARK verify on cutedsl_mla (fold_sq) - #33650
Merged
hnyls2002 merged 2 commits intoAug 6, 2026
Merged
Conversation
yhyang201
marked this pull request as ready for review
August 5, 2026 08:49
Collaborator
Author
|
A quick summary on picking the decode/verify attention backend for K3 MLA (cute-dsl vs trtllm), based on kernel microbenchmarks on B300: DCP: cute-dsl only — trtllm's decode kernel doesn't implement DCP and errors out; already forced, no choice. Pure TP8 (no DCP):
Why: trtllm slows down ~linearly with (context × q), while cute-dsl stays roughly flat thanks to fold_sq — so the longer the context, the more cute-dsl wins (at 1M the MLA kernel is ~4x faster, e2e decode ~2x). fp8 halves the KV bytes so trtllm holds up longer, pushing the crossover out ~2x. |
sagearc
pushed a commit
to sagearc/sglang
that referenced
this pull request
Aug 13, 2026
…edsl_mla (fold_sq) (sgl-project#33650) (sgl-project#34034) Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com> Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>
saturn-acc
pushed a commit
to saturn-acc/sglang
that referenced
this pull request
Aug 16, 2026
Atituiset
pushed a commit
to Atituiset/sglang
that referenced
this pull request
Sep 10, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
_dspark_verify_on_decode_backendcapscutedsl_mlaatq_len <= 4, from when the cute-dsl MLA decode kernel rejectedq_len >= 5. flashinfer's monolithic MLA decode now foldsseq_len_qinto the head dim (fold_sq, added in flashinfer #3309, gated-fixed in #3664, present in the pinned 0.6.15.post1), soq_len > 4is supported.K3 DSPARK verify uses
q_len = block_size(7) + 1 = 8, which exceeded the cap, so verify silently fell back totrtllm_mla— the fold-less path that re-reads the KV per query row and is slow at long context.Dropping the cap (cute-dsl serves any verify width via fold_sq) routes verify to the fold path. Measured on Kimi-K3, TP8, 8xB300, bf16 kvcache,
bs=1 isl=900000 osl=2048(decode backend the only variable):~2x decode throughput at 900k context, acc_length unchanged.
GPU timeline (nsys, B300, bf16, q=8, 900k) — isolated MLA decode kernel
Isolated flashinfer MLA decode call profiled with nsys (
--trace=cuda), same shapes, decode backend the only variable.trtllm_mla— one fold-lessfmhaSm100f…HQk576…kernel, 1.263 ms:cutedsl_mla—…monolithic mla_decode_fp16 Blackwell…fold_sq kernel + a small split-KV reduction, ~320 µs total (~3.9x faster kernel):The pure MLA attention kernel is ~4x faster; MLA is roughly half of a decode step, so the e2e decode speedup is ~2x.
Which decode backend? (isolated kernel microbenchmark)
Isolated flashinfer MLA decode on B300, flashinfer 0.6.15.post1, µs/call (median). H = query heads per rank (TP1=96, TP4=24, TP8=12, TP16=6; MQA, 1 KV head). q = spec verify tokens.
Typical K3 @ 1M-context configs (per-rank, µs/call, bold = pick)
For K3 DSPARK verify (q>1) at long context, cute-dsl is either the only option (TP1+DCP8) or ~4x faster (TP8 @ 1M) — hence this PR relaxes the guard to allow it.
Decision rule
q = 1 →
trtllm-gen. Always faster, even at 1M ctx (bf16 254 vs 331; fp8 130 vs 164).fold_sqneeds q>1 to help.trtllm-genunavailable →cute-dsl(only option). trtllm-gen fails withcomputeCtaAndClusterConfig: numHeadsQ/numHeadsKv not supportedat TP4 (H=24, all q/ctx) and TP1 (H=96) at long ctx or q≥7. TP8/TP16 always work; cute-dsl works everywhere. Dtype-independent.q > 1 and both available (TP8/TP16): pick by context length — cute-dsl wins above:
Below the threshold, trtllm-gen is ~30% faster wherever it runs.
Why: trtllm-gen re-reads KV per query row → time ≈ O(ctx × q). cute-dsl folds the q verify tokens into the head/MMA tile (F = largest divisor of q with H·F ≤ 128) → flat ~120µs up to ~256k, rising to only ~360µs bf16 / ~170µs fp8 at 1M.
Raw q×ctx sweep, TP8 / H=12 (trtllm | cutedsl µs/call; TP16/H=6 is within noise)
bf16:
fp8:
TODO / follow-ups (draft)
(TP, q, context, kv_dtype)-dependent — where both run and q>1, cutedsl only wins above a context threshold (bf16 q=8 ~100k … q=2 ~400k; fp8 ~2x higher), below which trtllm-gen is ~30% faster. sglang can't switch decode backend after launch (backends are fixed at init and baked into the decode CUDA graphs), so a per-step context-aware choice would need capturing graphs for both backends + a per-batch seq_len dispatcher.cutedsl_mlathe default decode backend for Kimi-K3 (currentlytrtllm_mla) -- needs correctness + no-regression, especially short context / low q where trtllm-gen is faster.CI States
Latest PR Test (Base): ❌ Run #30986145615
Latest PR Test (Extra): ❌ Run #30986145542