Skip to content

[Bugfix][Model] Restore causal image SWA for DeepSeek V4.1 - #57152

Merged
Isotr0py merged 5 commits into
vllm-project:mainfrom
Juntian777:bugfix/dsv41-vl-swa
Sep 16, 2026
Merged

Isotr0py merged 5 commits into
vllm-project:mainfrom
Juntian777:bugfix/dsv41-vl-swa

Conversation

@Juntian777

@Juntian777 Juntian777 commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Restore causal sliding-window attention for image tokens in DeepSeek V4.1 VL.

V4.1 inherited the image visibility rules used by DeepSeek V4 Vision-Exp: tokens inside an image could attend to earlier image tokens outside the sliding window and to future image tokens. This differs from the official V4.1 reference implementation, where _window_kv uses get_window_topk_idxs and both image and text tokens attend to the same causal window. With a window of 128, query position p should see only [max(0, p - 127), p] through SWA.

The incorrect widening affects full-image prefill as well as continuation chunks. For example, at the beginning of image span [1900, 3948), the inherited rule exposes future image tokens through position 3947; the V4.1 reference permits only positions 1773 through 1900.

The key line in the official get_window_topk_idxs is:

idxs = torch.where(idxs > end, -1, idxs)

The reference masks every index greater than the query position with -1; this is a causal 128-token window, including the current token. Executing the published get_window_topk_idxs at query position 1900 produces exactly 128 indices, from 1773 through 1900, with no future positions.

Changes

  • Disable multimodal prefix-LM/bidirectional-range configuration for V4.1.
  • Make the shared SWA metadata builder respect an explicit is_mm_prefix_lm=False, keeping V4.1 metadata and paged indices causal.
  • Set the V4.1 attention-side image window extension to zero. Image preprocessing, the vision encoder, and vision_max_n_token remain unchanged.
  • Keep the existing causal gather lengths and workspace sizes. The final diff does not change NVIDIA or AMD workspace allocation.
  • Add six regression cases: full prefill and continuation prefill, each with compression ratios 0/1/2. Verify causal key positions in both gathered and paged indices, gather/CPU-plan consistency, and absence of image visibility extensions.

V4 Vision-Exp retains its existing image visibility behavior.

Validation

Run on GB200:

CUDA_VISIBLE_DEVICES=1 .venv/bin/python -m pytest \
  tests/v1/attention/test_deepseek_v4_swa_visible.py \
  tests/kernels/attention/test_flashmla_sparse.py \
  -k 'not test_builder_' -q --tb=short

Result: 30 passed, 3 deselected. The three existing builder tests require an inaccessible gated Hugging Face configuration. The new regression uses a local configuration and requires no model download.

Changed-file pre-commit checks passed. A local FlashMLA numerical sanity check also matched the causal reference, retaining a gather length of 227 and a first-query SWA length of 128 for seq_len=4000, query_len=100, and image span [1900, 3948).

Full-model VL accuracy, generation-length evaluation, and AMD runtime validation were not run. The evidence here establishes attention-window correctness at the metadata/kernel level.

Related work

Open-PR searches for SWA/gather, V4.1/prefill, and V4.1/causal/image found no other matching fix. #56986 concerns cross-layer compressed-KV gather reuse; it does not correct the image attention-window semantics addressed here.

AI assistance

AI assistance was used for analysis, implementation, and validation.

Signed-off-by: Juntian Liu <Juntianl777@gmail.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added deepseek Related to DeepSeek models DSv4.1 Related to DeepSeek-V4.1 models bug Something isn't working labels Sep 16, 2026
@Juntian777
Juntian777 marked this pull request as draft September 16, 2026 10:12
Remove inherited bidirectional image visibility and the gather/workspace expansion. Verify full and continuation prefill against causal SWA indices.

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
@Juntian777 Juntian777 changed the title [Bugfix][Model] Fix incorrect image attention in DeepSeek V4.1 VL [Bugfix][Model] Restore causal image SWA for DeepSeek V4.1 Sep 16, 2026
Comment thread vllm/transformers_utils/configs/deepseek_v41.py Outdated
Comment thread vllm/models/deepseek_v41/attention.py Outdated
@Isotr0py Isotr0py self-assigned this Sep 16, 2026
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
@mergify mergify Bot added the nvidia label Sep 16, 2026
@Isotr0py
Isotr0py marked this pull request as ready for review September 16, 2026 13:18
@Isotr0py Isotr0py added the ready ONLY add when PR is ready to merge/full CI is needed label Sep 16, 2026
@github-project-automation github-project-automation Bot moved this to Ready in NVIDIA Sep 16, 2026
@github-actions

Copy link
Copy Markdown

✅ @Juntian777, CI is now available for this PR.

  • /ci run starts upstream CI; /amd-ci run starts AMD CI only.
  • Your branch must contain every commit currently on its upstream target branch. Merge or rebase onto the latest target branch, then rerun the command. Append --allow-stale to a run command to test an outdated branch at your own risk.
  • /ci retry retries failed jobs in the CI build for the current PR head. If the current head has no CI build, it starts a new CI build for the current head containing only jobs that failed in the latest earlier CI build for this PR.
  • /amd-ci retry retries failed jobs in AMD CI for the current PR head. Use /amd-ci run when the current head has no AMD CI build.
  • /ci cancel cancels scheduled or running CI builds for this PR branch; /amd-ci cancel does the same for AMD CI only.

@Isotr0py

Copy link
Copy Markdown
Member

/ci run

@github-actions

Copy link
Copy Markdown

❌ This PR is 7 commits behind upstream main. Your branch must contain every commit currently on upstream main. No new CI build was started. Merge or rebase onto the latest main, then rerun /ci run. To test this branch at your own risk, use /ci run --allow-stale.

@mergify

mergify Bot commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

Hi @Juntian777, the pre-commit checks have failed. Please run:

uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-files

Then, commit the changes and push to your branch.

For future commits, pre-commit will run automatically on changed files before each commit.

@mergify

mergify Bot commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

Hi @Juntian777, the pre-commit checks have failed. Please run:

uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-files

Then, commit the changes and push to your branch.

For future commits, pre-commit will run automatically on changed files before each commit.

Signed-off-by: Isotr0py <Isotr0py@outlook.com>
@Isotr0py
Isotr0py enabled auto-merge (squash) September 16, 2026 14:56
@Isotr0py

Copy link
Copy Markdown
Member

/ci run --allow-stale

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #89358 for commit e62b3c4abbae.

@Isotr0py
Isotr0py merged commit 9f9e1da into vllm-project:main Sep 16, 2026
156 checks passed
@github-project-automation github-project-automation Bot moved this from Ready to Done in NVIDIA Sep 16, 2026
jagat-primitive-org added a commit to jagat-primitive-org/vllm that referenced this pull request Sep 16, 2026
DeepSeek-V4 vision checkpoints widen prefill sliding-window index rows by
vision_max_n_token so image spans are visible bidirectionally. After vllm-project#57152
the metadata builder gates that on mm_prefix_clamp_sliding_window, but the
attention layer and the warmup-key builders still gate on vision_n_layers,
and none of them look at --language-model-only, so a vision checkpoint served
text-only still sizes every prefill row and warmup key for images that can
never arrive.

Add swa_max_image_tokens(vllm_config): 0 unless the config sets
mm_prefix_clamp_sliding_window, 0 under language_model_only, otherwise
vision_max_n_token. Use it at the four sites that computed the width by hand
(builder, V4 attention layer, both V4 warmup-key builders) so they agree, and
add a builder test for the language-model-only case.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working deepseek Related to DeepSeek models DSv4.1 Related to DeepSeek-V4.1 models nvidia ready ONLY add when PR is ready to merge/full CI is needed

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants