Repository navigation
V4.1 on SM120: 64-token SWA pages (overlay + self-test members) and a text-only launch - #12906
Conversation
… and a text-only launch Attempt 13 died on all four ranks in CUDA-graph capture with "SM120 sparse-MLA has no decode kernel for this shape: num_tokens=8, num_heads=16, topk=1152, page_block_size=32". Two independent refusals sit behind it, and neither is a combined window+indexer list (vLLM's SM120 class already passes the window and the indexer picks as separate segments): - page_block_size=32: vllm/models/deepseek_v4_1/attention.py constructs the SWA cache with the literal block_size=32; FlashInfer's SM120 DSv4 decode kernel takes 64-token pages only, and every call with <= 64 query tokens must take it. No engine argument reaches the literal, so it is an overlay member: swa_block_size class attribute, 64 on the SM120 FlashInfer class. - topk=1152: the vision variant widens prefill SWA rows to window + vision_max_n_token. The launch now serves text only: --language-model-only --hf-overrides.vision_n_layers 0. The vision capability is dropped on purpose. The self-test member creates vllm.models.deepseek_v4_1.nvidia.sm120_decode_shape_selftest. Probe run 36857198484 (srv6/7/8, image sha256:0f6039eb...): unpatched FAILS in both modalities, patched+multimodal FAILS at 1152, patched+text-only PASSES all four shapes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ove the claim-skip note out of the body (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rride through it (sole_constructor) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… engine argv before staging Review 73720: the self-test member had no executing consumer, and two of its figures were assumed. - gunbc.spark.v41_group_a_launch v41_sm120_gate runs the image's vllm.models.deepseek_v4_1.nvidia.sm120_decode_shape_selftest on each rank's host with that rank's engine argv (checkpoint mounted read-only at the argv's model path) and stops the launch before the Engram staging on anything short of SELFTEST PASS; the gate's lines land in the receipt. - The self-test now parses that argv with vLLM's own serve parser, and reads the packed bytes per token (FlashInfer _BPT_DSV4), the compressed pools' page (backend kernel block / each compress ratio), the local heads and the decode token ceiling off the installed tree. - Witnesses: every rank's argv is text-only; the gate's argv is the rank's engine argv in its image. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
5297211 to
4c8c10d
Compare
…orted bare get (floor UnimportedBareProvider) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ch typing) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Re review 73720 (REQUEST_CHANGES). Both findings were valid; both are fixed.
Execution evidence: fleet-converge 36865955537 ran on srv5, srv7 and srv8 against the current image, read-only. It used a rank-shaped argv, including
CI is green on c76f902. — sent from proud-bear-339 |
What failed
Attempt 13 (2026-10-01 04:40Z) failed on all four ranks during CUDA-graph capture:
SM120 sparse-MLA has no decode kernel for this shape: num_tokens=8, num_heads=16, topk=1152, d_qk=512, page_block_size=32, model_type=1, extra_topk=0.The traceback runs through
DeepseekV4FlashInferSM120Attention._forward_prefill(flashinfer_sparse.py:895). It was an 8-token prefill chunk in a mixed warmup batch, not decode.The window + extra-topk split the brief proposed would not have fixed this. vLLM's SM120 class already passes the sliding window and the indexer picks as separate segments. The failure has two independent causes instead:
page_block_size=32. FlashInfer's SM120 DSv4 decode kernel only accepts 64-token pages (_DECODE_DSV4_PAGE_BLOCK_SIZE). Every call with 64 or fewer query tokens has to use that kernel, because the paged orchestrator handles prefill only. V4.1'sattention.pybuilds the sliding-window cache with a hardcodedblock_size=32; V4 uses the default of 64. With 32-token pages, every decode step is refused, whatever its topk.topk=1152. This is the prefill sliding-window width,sliding_window (128) + vision_max_n_token (1024). The checkpoint is the vision variant, and vLLM reads that width from the model config.The change
sm120_swa_page_block_size.patch(new overlay member). It turns the hardcoded value into aswa_block_sizeclass attribute with default 32, and sets it to 64 onDeepseekV4FlashInferSM120Attention. Other platform classes keep their current page size. No vLLM launch argument reaches the hardcoded value.sm120_decode_shape_selftest.patch(new overlay member). It createsvllm.models.deepseek_v4_1.nvidia.sm120_decode_shape_selftest. The test calls FlashInfer's DSv4 entry point at the TP4 decode and short-prefill shapes. It reads the page size from the installed tree, and the prefill width from vLLM's ownModelConfigbuilt with the launch's arguments. It applies on its own to an unpatched tree.v41_storage_backed_patch_populationafter the Engram pair, so the runtime key changes. The produced-image row needs a rebuild before the next launch.extdeps.vllm.engine_argsgainsVllmModalityRequestandVllmHfOverride, rendered in the dotted form. The V4.1 measurement profile asks forLanguageModelOnlywithvision_n_layers=0, which renders--language-model-only --hf-overrides.vision_n_layers 0. The request enters the profile key only when it is made, as the Engram request does.The vision capability is dropped on purpose. This launch accepts no images, and its prefill sliding-window width is 128. Serving images again needs an SM120 kernel for the widened width first, and would be a separate product.
Evidence
Probe run https://github.com/gunb-ai/gunbc/actions/runs/36857198484 used the current image
sha256:0f6039eb…, read-only, with the patch applied inside a--rmcontainer. srv6, srv7 and srv8 gave the same results; srv5 printed no output.The checkpoint's
index_topkis 512, so the extra segment is 512 wide, not 1024.Not yet tested by execution: the whole engine starting with these arguments. That is the next launch.
🤖 Generated with Claude Code