Skip to content

[Bugfix][DSv4.1] Use 64-token sparse-MLA pages on SM120 - #57292

Open
luoyuctl wants to merge 8 commits into
vllm-project:mainfrom
luoyuctl:fix/dsv41-sm120-block-size
Open

luoyuctl wants to merge 8 commits into
vllm-project:mainfrom
luoyuctl:fix/dsv41-sm120-block-size

Conversation

@luoyuctl

@luoyuctl luoyuctl commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Partially addresses #56461; related to #59203.

On SM120/SM121, DeepSeek-V4.1-Flash fails at startup with ValueError: No common block size for 64. The V4.1 sparse-MLA, FlashInfer sparse-MLA, and indexer backends declare a 128-token kernel page on every non-SM90 GPU. The SM120 FlashInfer decode kernels are only instantiated for 64-token pages (_DECODE_DSV4_PAGE_BLOCK_SIZE = 64), and the MLA and indexer groups share a manager block, so no common kernel block size exists.

This PR declares 64-token kernel pages on the SM120 family (same as SM90) for DeepseekV4SparseMLABackend, DeepseekV4FlashInferMLASparseBackend, and DeepseekV41IndexerBackend. It also sizes the SWA cache at 64 tokens on SM120. SM90 and SM100 behavior is unchanged.

Scope and dependencies

This is the vLLM-side geometry only. It makes the MLA, FlashInfer sparse-MLA and indexer backends declare one common kernel page size on SM120, so select_common_block_size no longer fails.

Test Plan

pytest tests/v1/attention/test_indexer_deepseek_v4_slot_mapping.py \
  -k "deepseek_v41_sparse_mla_kernel_page_size or deepseek_v41_swa_cache_block_size or preserves_deepseek_v41 or shares_uncompressed"

The new tests patch the device capability, so they run without an SM120 GPU. They assert that every V4.1 backend supported on the given capability declares the same page size, and that select_common_block_size resolves it (the call that raised in production).

Test Result

  • 9 passed locally (CPU, capability patched).
  • Against the unfixed backends on main, the sm120 and sm121 cases fail and the sm90/sm100 cases pass. So the test reproduces the startup failure and does not change other architectures.
  • pre-commit hooks pass.

End-to-end SM120 validation (external)

@xzwgit ran this PR end to end on 8× RTX PRO 6000 (SM120), reported in #59203:

Input/Output Concurrency Output tput (tok/s) Mean TTFT (s) Mean TPOT (ms)
1K/1K 1 90.0 0.16 11.0
1K/1K 4 301.6 0.19 13.1
4K/4K 1 88.1 1.95* 10.9
4K/4K 4 300.0 0.32 13.3
8K/8K 1 91.3 0.28 10.9
8K/8K 4 299.5 0.41 13.3

* Cold start; reruns of the same shape measured ~0.2 s.

  • gsm8k (first 100 questions, thinking on with reasoning_effort: 25, temperature 0): 99/100 (report).

This run used a separately installed DeepGEMM at dev HEAD, not a bumped vLLM pin, so the pin bump above is still needed; #59385 proposes it. An earlier backport of this geometry also served correctly on 8× RTX PRO 5000 (SM120, TP8, FP8 KV).

Duplicate check

Checked open PRs for #56461, #59203, and the SM120 DSv4.1 page-size area. #56509 is the only overlapping PR; the differences are described above.


AI assistance (Claude) was used for implementation, tests, and this description. I reviewed every changed line and ran the tests above.


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

DeepSeek-V4.1 picks the SM12x attention path on compute capability 12.x
(DeepseekV4FlashInferSM120Attention), but three sites that declare the
sparse-MLA kernel block size still treat SM120 like SM100 and return 128:

* vllm/models/deepseek_v41/sparse_mla.py (DeepseekV4SparseMLABackend)
* vllm/models/deepseek_v41/nvidia/flashinfer_sparse.py
  (DeepseekV4FlashInferMLASparseBackend, shared by SM100 and SM120)
* vllm/v1/attention/backends/mla/indexer.py (DeepseekV41IndexerBackend)

The SM120 kernels are built for 64-token pages: FlashInfer's
mla/_sparse_mla_sm120.py has _DECODE_DSV4_PAGE_BLOCK_SIZE = 64, and the
vendored DeepGEMM paged-MQA logits kernel only accepts block_kv in
{32, 64} (csrc/apis/attention.hpp:262). V4.1's indexer feeds
num_states = block_size // compress_ratio to that kernel, and the model
mixes ratio-1 and ratio-2 layers (compress_ratios = 2x0, 18x2, 20x1), so
claiming 128 leaves no working block size at all:

* default -> ratio-1 layers give 128 // 1 = 128 ->
  RuntimeError: Assertion error ... block_kv == 32 or block_kv == 64
* --block-size 64 -> ValueError: No common block size for 64
  (64 % 128 != 0), because the backends claim 128

Widen the 64 branch to cover SM120 as well. The DeepseekV4SWACache
constructed by DeepseekV4Attention is hardcoded to block_size=32, which
the SM120 sparse-MLA decode dispatch also rejects (page_block_size must
be 64), so make that arch aware too.

Refs: vllm-project#56702, vllm-project#56461

Signed-off-by: zack <51604064+luoyuctl@users.noreply.github.com>
@mergify mergify Bot added deepseek Related to DeepSeek models DSv4.1 Related to DeepSeek-V4.1 models nvidia bug Something isn't working labels Sep 17, 2026
Follow-up to the SM120 kernel block size fix:

* Extract the sliding-window cache page size into ``_swa_cache_block_size``
  so it can be asserted directly from a test.
* Reject a ``--block-size`` that is not a multiple of 64 on SM120 in
  ``DeepseekV4FlashInferSM120Attention.__init__``. The SM120 sparse-MLA
  kernels page the cache at 64 tokens, so anything else would otherwise
  only fail later in the opaque DeepGEMM paged-MQA assert.
* Add unit tests that patch the platform capability and assert the declared
  kernel block sizes for SM90 / SM100 / SM120 / SM121, plus the V4.1
  sliding-window page size.

Refs: vllm-project#56702, vllm-project#56461

Signed-off-by: zack <51604064+luoyuctl@users.noreply.github.com>
@luoyuctl
luoyuctl marked this pull request as ready for review September 23, 2026 02:41

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

Copy link
Copy Markdown
Contributor Author

@pavanimajety @zyongye This is now ready for review.

The scope is limited to the SM120 DeepSeek-V4.1 sparse-MLA/SWA block geometry and related tests. It overlaps with #56509, so I’d especially appreciate guidance on which implementation should be consolidated/kept rather than merging duplicate fixes.

I have real SM120 validation from 8× RTX PRO 5000, and the remaining DeepGEMM page32 dependency is tracked separately.

@mergify

mergify Bot commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @luoyuctl.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 24, 2026
Preserve the kv_cache_spec backend interface and FlashInfer 0.7 updates. Remove the premature manager block-size check while retaining SM120 64-token kernel and SWA pages.

Co-authored-by: Codex <codex@openai.com>
Signed-off-by: Codex <codex@openai.com>
@mergify mergify Bot removed the needs-rebase label Sep 24, 2026
Move the SM120 kernel page-size tests into the existing DeepSeek-V4
indexer block-size tests and assert that the MLA, FlashInfer, and
indexer backends resolve a common kernel block size. Simplify the
duplicated capability check.

Co-authored-by: Claude
Signed-off-by: zack <51604064+luoyuctl@users.noreply.github.com>
@luoyuctl

Copy link
Copy Markdown
Contributor Author

@zyongye @LucasWilkinson Could you take a look when you have time? This declares 64-token kernel pages for the DSv4.1 sparse-MLA, FlashInfer, and indexer backends on SM120, so select_common_block_size no longer fails at startup. #56509 overlaps with a different geometry (ratio-specific [64, 128]); I'd appreciate your call on which approach to keep, and I'm happy to consolidate. cc @lucifer1004 in case you can sanity-check on SM120 hardware.

@lucifer1004

Copy link
Copy Markdown
Contributor

@luoyuctl Have you tested latest flashinfer main? Page size is not fixed any more after flashinfer-ai/flashinfer#5197

@luoyuctl

luoyuctl commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor Author

@lucifer1004 Thanks, you're right. flashinfer-ai/flashinfer#5197 (merged 2026-09-18, eb5f05b) makes the SM120 sparse-MLA page size a runtime argument, including independent main/extra page sizes for the DSv4.1 dual cache. That removes the kernel-side restriction this PR originally worked around, and it covers the 32-token extra pages I had listed flashinfer-ai/flashinfer#5174 for.

The vLLM-side page declarations still need to agree, though. On current main, FLASHINFER_MLA_SPARSE_DSV41 and DeepseekV4SparseMLABackend declare 128-token kernel pages on every non-SM90 GPU, while DeepseekV41IndexerBackend declares 64. The MLA and indexer groups share a manager block, so startup fails in select_common_block_size with "No common block size for 64" (the second failure mode in #59203), regardless of what the kernel accepts. This PR aligns those declarations on SM120, so I think it is still needed.

For reference, this geometry has served on 8× RTX PRO 5000 (SM120), but only on a backport image, not on this PR head:

Prompts from 46 to 10k tokens and 4-way mixed-length concurrency returned correct output. That is a smoke test, not an accuracy or throughput run.

Two caveats:

I've updated the description to depend on flashinfer-ai/flashinfer#5197 instead of flashinfer-ai/flashinfer#5174.

@gitbisector

Copy link
Copy Markdown
Contributor

We run 64-row SM12x pages (64-token SWA pages; 64 states per compressed and indexer page) in production on 4× GB10, TP4 + DCP2, and our count-to-100, structured-output and vision/tool gate passes. The NaN you may hit once pages work is #57156, fixed by #59689.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working deepseek Related to DeepSeek models DSv4.1 Related to DeepSeek-V4.1 models nvidia

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants