Skip to content

V4.1 on SM120: 64-token SWA pages (overlay + self-test members) and a text-only launch - #12906

Merged
briansrls merged 5 commits into
mainfrom
session/proud-bear-339-v41-sm120
Oct 1, 2026
Merged

briansrls merged 5 commits into
mainfrom
session/proud-bear-339-v41-sm120

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

What failed

Attempt 13 (2026-10-01 04:40Z) failed on all four ranks during CUDA-graph capture:
SM120 sparse-MLA has no decode kernel for this shape: num_tokens=8, num_heads=16, topk=1152, d_qk=512, page_block_size=32, model_type=1, extra_topk=0.
The traceback runs through DeepseekV4FlashInferSM120Attention._forward_prefill (flashinfer_sparse.py:895). It was an 8-token prefill chunk in a mixed warmup batch, not decode.

The window + extra-topk split the brief proposed would not have fixed this. vLLM's SM120 class already passes the sliding window and the indexer picks as separate segments. The failure has two independent causes instead:

  1. page_block_size=32. FlashInfer's SM120 DSv4 decode kernel only accepts 64-token pages (_DECODE_DSV4_PAGE_BLOCK_SIZE). Every call with 64 or fewer query tokens has to use that kernel, because the paged orchestrator handles prefill only. V4.1's attention.py builds the sliding-window cache with a hardcoded block_size=32; V4 uses the default of 64. With 32-token pages, every decode step is refused, whatever its topk.
  2. topk=1152. This is the prefill sliding-window width, sliding_window (128) + vision_max_n_token (1024). The checkpoint is the vision variant, and vLLM reads that width from the model config.

The change

  • sm120_swa_page_block_size.patch (new overlay member). It turns the hardcoded value into a swa_block_size class attribute with default 32, and sets it to 64 on DeepseekV4FlashInferSM120Attention. Other platform classes keep their current page size. No vLLM launch argument reaches the hardcoded value.
  • sm120_decode_shape_selftest.patch (new overlay member). It creates vllm.models.deepseek_v4_1.nvidia.sm120_decode_shape_selftest. The test calls FlashInfer's DSv4 entry point at the TP4 decode and short-prefill shapes. It reads the page size from the installed tree, and the prefill width from vLLM's own ModelConfig built with the launch's arguments. It applies on its own to an unpatched tree.
  • Candidate recipe. Both members are keyed into v41_storage_backed_patch_population after the Engram pair, so the runtime key changes. The produced-image row needs a rebuild before the next launch.
  • Launch. extdeps.vllm.engine_args gains VllmModalityRequest and VllmHfOverride, rendered in the dotted form. The V4.1 measurement profile asks for LanguageModelOnly with vision_n_layers=0, which renders --language-model-only --hf-overrides.vision_n_layers 0. The request enters the profile key only when it is made, as the Engram request does.

The vision capability is dropped on purpose. This launch accepts no images, and its prefill sliding-window width is 128. Serving images again needs an SM120 kernel for the widened width first, and would be a separate product.

Evidence

Probe run https://github.com/gunb-ai/gunbc/actions/runs/36857198484 used the current image sha256:0f6039eb…, read-only, with the patch applied inside a --rm container. srv6, srv7 and srv8 gave the same results; srv5 printed no output.

arm result
unpatched, multimodal FAIL (all four shapes refused; reproduces attempt 13)
unpatched, text-only FAIL (page size 32)
patched, multimodal FAIL (decode ran; prefill at 1152 refused)
patched, text-only PASS (all four shapes ran)

The checkpoint's index_topk is 512, so the extra segment is 512 wide, not 1024.

Not yet tested by execution: the whole engine starting with these arguments. That is the next launch.

🤖 Generated with Claude Code

… and a text-only launch

Attempt 13 died on all four ranks in CUDA-graph capture with "SM120 sparse-MLA has no
decode kernel for this shape: num_tokens=8, num_heads=16, topk=1152, page_block_size=32".
Two independent refusals sit behind it, and neither is a combined window+indexer list
(vLLM's SM120 class already passes the window and the indexer picks as separate segments):

- page_block_size=32: vllm/models/deepseek_v4_1/attention.py constructs the SWA cache with
  the literal block_size=32; FlashInfer's SM120 DSv4 decode kernel takes 64-token pages only,
  and every call with <= 64 query tokens must take it. No engine argument reaches the literal,
  so it is an overlay member: swa_block_size class attribute, 64 on the SM120 FlashInfer class.
- topk=1152: the vision variant widens prefill SWA rows to window + vision_max_n_token.
  The launch now serves text only: --language-model-only --hf-overrides.vision_n_layers 0.
  The vision capability is dropped on purpose.

The self-test member creates vllm.models.deepseek_v4_1.nvidia.sm120_decode_shape_selftest.
Probe run 36857198484 (srv6/7/8, image sha256:0f6039eb...): unpatched FAILS in both modalities,
patched+multimodal FAILS at 1152, patched+text-only PASSES all four shapes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Oct 1, 2026
…ove the claim-skip note out of the body (never merge)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
gunbc-ci-auto-heal and others added 2 commits October 1, 2026 12:12
…rride through it (sole_constructor)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… engine argv before staging

Review 73720: the self-test member had no executing consumer, and two of its figures were assumed.
- gunbc.spark.v41_group_a_launch v41_sm120_gate runs the image's
  vllm.models.deepseek_v4_1.nvidia.sm120_decode_shape_selftest on each rank's host with that rank's
  engine argv (checkpoint mounted read-only at the argv's model path) and stops the launch before the
  Engram staging on anything short of SELFTEST PASS; the gate's lines land in the receipt.
- The self-test now parses that argv with vLLM's own serve parser, and reads the packed bytes per token
  (FlashInfer _BPT_DSV4), the compressed pools' page (backend kernel block / each compress ratio), the
  local heads and the decode token ceiling off the installed tree.
- Witnesses: every rank's argv is text-only; the gate's argv is the rank's engine argv in its image.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot force-pushed the session/proud-bear-339-v41-sm120 branch from 5297211 to 4c8c10d Compare October 1, 2026 13:02
gunbc-ci-auto-heal and others added 2 commits October 1, 2026 13:18
…orted bare get (floor UnimportedBareProvider)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ch typing)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Re review 73720 (REQUEST_CHANGES). Both findings were valid; both are fixed.

  1. The self-test had no executing consumer. gunbc.spark.v41_group_a_launch v41_sm120_gate now runs vllm.models.deepseek_v4_1.nvidia.sm120_decode_shape_selftest on each rank's host before the Engram staging.

    • It runs inside the planned image, with that rank's exact engine argv (v41_sm120_selftest_argv). The checkpoint is mounted read-only at the argv's model path.
    • Anything short of SELFTEST PASS stops the launch, and the gate's output goes into the launch receipt.
    • Witnesses: w_the_sm120_gate_runs_each_ranks_engine_argv_in_its_image and w_every_rank_serves_text_only_with_the_vision_layers_overridden_away.
    • The recipe comment now names the gate as the consumer and cites the run where the RED was taken.
  2. Assumed figures. The self-test now parses the rank's argv with vLLM's own serve parser (make_arg_parser) and builds the engine's ModelConfig from it. It reads these values from the installed tree:

    • bytes per token: FlashInfer _BPT_DSV4
    • compressed-pool pages: the backend kernel block over each of the config's compress_ratios
    • padded local heads
    • the decode token ceiling: _DECODE_MAX_TOKENS

    It covers 1, 8 and 64 query tokens.

Execution evidence: fleet-converge 36865955537 ran on srv5, srv7 and srv8 against the current image, read-only. It used a rank-shaped argv, including --headless and the dotted --engram-config flags:

  • unpatched + text-only: FAIL (page 32)
  • patched + multimodal: FAIL (width 1152)
  • patched + text-only: SELFTEST PASS

CI is green on c76f902.

— sent from proud-bear-339

@briansrls
briansrls added this pull request to the merge queue Oct 1, 2026
Merged via the queue into main with commit e9a7ce4 Oct 1, 2026
4 checks passed
@briansrls
briansrls deleted the session/proud-bear-339-v41-sm120 branch October 1, 2026 15:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant