Conversation
DeepSeek-V4.1 picks the SM12x attention path on compute capability 12.x
(DeepseekV4FlashInferSM120Attention), but three sites that declare the
sparse-MLA kernel block size still treat SM120 like SM100 and return 128:
* vllm/models/deepseek_v41/sparse_mla.py (DeepseekV4SparseMLABackend)
* vllm/models/deepseek_v41/nvidia/flashinfer_sparse.py
(DeepseekV4FlashInferMLASparseBackend, shared by SM100 and SM120)
* vllm/v1/attention/backends/mla/indexer.py (DeepseekV41IndexerBackend)
The SM120 kernels are built for 64-token pages: FlashInfer's
mla/_sparse_mla_sm120.py has _DECODE_DSV4_PAGE_BLOCK_SIZE = 64, and the
vendored DeepGEMM paged-MQA logits kernel only accepts block_kv in
{32, 64} (csrc/apis/attention.hpp:262). V4.1's indexer feeds
num_states = block_size // compress_ratio to that kernel, and the model
mixes ratio-1 and ratio-2 layers (compress_ratios = 2x0, 18x2, 20x1), so
claiming 128 leaves no working block size at all:
* default -> ratio-1 layers give 128 // 1 = 128 ->
RuntimeError: Assertion error ... block_kv == 32 or block_kv == 64
* --block-size 64 -> ValueError: No common block size for 64
(64 % 128 != 0), because the backends claim 128
Widen the 64 branch to cover SM120 as well. The DeepseekV4SWACache
constructed by DeepseekV4Attention is hardcoded to block_size=32, which
the SM120 sparse-MLA decode dispatch also rejects (page_block_size must
be 64), so make that arch aware too.
Refs: vllm-project#56702, vllm-project#56461
Signed-off-by: zack <51604064+luoyuctl@users.noreply.github.com>
Follow-up to the SM120 kernel block size fix: * Extract the sliding-window cache page size into ``_swa_cache_block_size`` so it can be asserted directly from a test. * Reject a ``--block-size`` that is not a multiple of 64 on SM120 in ``DeepseekV4FlashInferSM120Attention.__init__``. The SM120 sparse-MLA kernels page the cache at 64 tokens, so anything else would otherwise only fail later in the opaque DeepGEMM paged-MQA assert. * Add unit tests that patch the platform capability and assert the declared kernel block sizes for SM90 / SM100 / SM120 / SM121, plus the V4.1 sliding-window page size. Refs: vllm-project#56702, vllm-project#56461 Signed-off-by: zack <51604064+luoyuctl@users.noreply.github.com>
|
@pavanimajety @zyongye This is now ready for review. The scope is limited to the SM120 DeepSeek-V4.1 sparse-MLA/SWA block geometry and related tests. It overlaps with #56509, so I’d especially appreciate guidance on which implementation should be consolidated/kept rather than merging duplicate fixes. I have real SM120 validation from 8× RTX PRO 5000, and the remaining DeepGEMM page32 dependency is tracked separately. |
|
This pull request has merge conflicts that must be resolved before it can be |
Preserve the kv_cache_spec backend interface and FlashInfer 0.7 updates. Remove the premature manager block-size check while retaining SM120 64-token kernel and SWA pages. Co-authored-by: Codex <codex@openai.com> Signed-off-by: Codex <codex@openai.com>
Move the SM120 kernel page-size tests into the existing DeepSeek-V4 indexer block-size tests and assert that the MLA, FlashInfer, and indexer backends resolve a common kernel block size. Simplify the duplicated capability check. Co-authored-by: Claude Signed-off-by: zack <51604064+luoyuctl@users.noreply.github.com>
|
@zyongye @LucasWilkinson Could you take a look when you have time? This declares 64-token kernel pages for the DSv4.1 sparse-MLA, FlashInfer, and indexer backends on SM120, so |
|
@luoyuctl Have you tested latest flashinfer main? Page size is not fixed any more after flashinfer-ai/flashinfer#5197 |
|
@lucifer1004 Thanks, you're right. flashinfer-ai/flashinfer#5197 (merged 2026-09-18, eb5f05b) makes the SM120 sparse-MLA page size a runtime argument, including independent main/extra page sizes for the DSv4.1 dual cache. That removes the kernel-side restriction this PR originally worked around, and it covers the 32-token extra pages I had listed flashinfer-ai/flashinfer#5174 for. The vLLM-side page declarations still need to agree, though. On current main, For reference, this geometry has served on 8× RTX PRO 5000 (SM120), but only on a backport image, not on this PR head:
Prompts from 46 to 10k tokens and 4-way mixed-length concurrency returned correct output. That is a smoke test, not an accuracy or throughput run. Two caveats:
I've updated the description to depend on flashinfer-ai/flashinfer#5197 instead of flashinfer-ai/flashinfer#5174. |
Purpose
Partially addresses #56461; related to #59203.
On SM120/SM121, DeepSeek-V4.1-Flash fails at startup with
ValueError: No common block size for 64. The V4.1 sparse-MLA, FlashInfer sparse-MLA, and indexer backends declare a 128-token kernel page on every non-SM90 GPU. The SM120 FlashInfer decode kernels are only instantiated for 64-token pages (_DECODE_DSV4_PAGE_BLOCK_SIZE = 64), and the MLA and indexer groups share a manager block, so no common kernel block size exists.This PR declares 64-token kernel pages on the SM120 family (same as SM90) for
DeepseekV4SparseMLABackend,DeepseekV4FlashInferMLASparseBackend, andDeepseekV41IndexerBackend. It also sizes the SWA cache at 64 tokens on SM120. SM90 and SM100 behavior is unchanged.Scope and dependencies
This is the vLLM-side geometry only. It makes the MLA, FlashInfer sparse-MLA and indexer backends declare one common kernel page size on SM120, so
select_common_block_sizeno longer fails.e1f418c, which predates it, so the DeepGEMM pin must be bumped.[64, 128]kernel blocks. That geometry is also viable once refactor(sparse-mla): unify SM120 execution, calibration and DSv4.1 support flashinfer-ai/flashinfer#5197 is released. I'm happy to consolidate with it or close this one.Test Plan
pytest tests/v1/attention/test_indexer_deepseek_v4_slot_mapping.py \ -k "deepseek_v41_sparse_mla_kernel_page_size or deepseek_v41_swa_cache_block_size or preserves_deepseek_v41 or shares_uncompressed"The new tests patch the device capability, so they run without an SM120 GPU. They assert that every V4.1 backend supported on the given capability declares the same page size, and that
select_common_block_sizeresolves it (the call that raised in production).Test Result
main, thesm120andsm121cases fail and thesm90/sm100cases pass. So the test reproduces the startup failure and does not change other architectures.pre-commithooks pass.End-to-end SM120 validation (external)
@xzwgit ran this PR end to end on 8× RTX PRO 6000 (SM120), reported in #59203:
0.30.1rc1.dev382+g09c47db1c+ this PR, FlashInfer maind849eca(includes refactor(sparse-mla): unify SM120 execution, calibration and DSv4.1 support flashinfer-ai/flashinfer#5197), DeepGEMM6901431(includes [SM120] Support 32-state pages in FP8 paged-MQA logits DeepGEMM#14),FLASHINFER_DISABLE_VERSION_CHECK=1.--kv-cache-dtype fp8,indexer_kv_dtype=fp8,--language-model-only, no--enforce-eager.-9999.0).vllm bench serve, random dataset,--ignore-eos), all 9 tiers with zero failures:* Cold start; reruns of the same shape measured ~0.2 s.
reasoning_effort: 25, temperature 0): 99/100 (report).This run used a separately installed DeepGEMM at dev HEAD, not a bumped vLLM pin, so the pin bump above is still needed; #59385 proposes it. An earlier backport of this geometry also served correctly on 8× RTX PRO 5000 (SM120, TP8, FP8 KV).
Duplicate check
Checked open PRs for #56461, #59203, and the SM120 DSv4.1 page-size area. #56509 is the only overlapping PR; the differences are described above.
AI assistance (Claude) was used for implementation, tests, and this description. I reviewed every changed line and ran the tests above.
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.