[Bugfix][Model] Restore causal image SWA for DeepSeek V4.1 - #57152
Conversation
Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Remove inherited bidirectional image visibility and the gather/workspace expansion. Verify full and continuation prefill against causal SWA indices. Co-authored-by: Codex <noreply@openai.com> Signed-off-by: Juntian Liu <Juntianl777@gmail.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
|
✅ @Juntian777, CI is now available for this PR.
|
|
/ci run |
|
❌ This PR is 7 commits behind upstream |
|
Hi @Juntian777, the pre-commit checks have failed. Please run: uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-filesThen, commit the changes and push to your branch. For future commits, |
|
Hi @Juntian777, the pre-commit checks have failed. Please run: uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-filesThen, commit the changes and push to your branch. For future commits, |
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
|
/ci run --allow-stale |
|
✅ Triggered Buildkite CI #89358 for commit |
DeepSeek-V4 vision checkpoints widen prefill sliding-window index rows by vision_max_n_token so image spans are visible bidirectionally. After vllm-project#57152 the metadata builder gates that on mm_prefix_clamp_sliding_window, but the attention layer and the warmup-key builders still gate on vision_n_layers, and none of them look at --language-model-only, so a vision checkpoint served text-only still sizes every prefill row and warmup key for images that can never arrive. Add swa_max_image_tokens(vllm_config): 0 unless the config sets mm_prefix_clamp_sliding_window, 0 under language_model_only, otherwise vision_max_n_token. Use it at the four sites that computed the width by hand (builder, V4 attention layer, both V4 warmup-key builders) so they agree, and add a builder test for the language-model-only case. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Summary
Restore causal sliding-window attention for image tokens in DeepSeek V4.1 VL.
V4.1 inherited the image visibility rules used by DeepSeek V4 Vision-Exp: tokens inside an image could attend to earlier image tokens outside the sliding window and to future image tokens. This differs from the official V4.1 reference implementation, where
_window_kvusesget_window_topk_idxsand both image and text tokens attend to the same causal window. With a window of 128, query positionpshould see only[max(0, p - 127), p]through SWA.The incorrect widening affects full-image prefill as well as continuation chunks. For example, at the beginning of image span
[1900, 3948), the inherited rule exposes future image tokens through position 3947; the V4.1 reference permits only positions 1773 through 1900.The key line in the official
get_window_topk_idxsis:The reference masks every index greater than the query position with
-1; this is a causal 128-token window, including the current token. Executing the publishedget_window_topk_idxsat query position 1900 produces exactly 128 indices, from 1773 through 1900, with no future positions.Changes
is_mm_prefix_lm=False, keeping V4.1 metadata and paged indices causal.vision_max_n_tokenremain unchanged.0/1/2. Verify causal key positions in both gathered and paged indices, gather/CPU-plan consistency, and absence of image visibility extensions.V4 Vision-Exp retains its existing image visibility behavior.
Validation
Run on GB200:
CUDA_VISIBLE_DEVICES=1 .venv/bin/python -m pytest \ tests/v1/attention/test_deepseek_v4_swa_visible.py \ tests/kernels/attention/test_flashmla_sparse.py \ -k 'not test_builder_' -q --tb=shortResult: 30 passed, 3 deselected. The three existing builder tests require an inaccessible gated Hugging Face configuration. The new regression uses a local configuration and requires no model download.
Changed-file pre-commit checks passed. A local FlashMLA numerical sanity check also matched the causal reference, retaining a gather length of 227 and a first-query SWA length of 128 for
seq_len=4000,query_len=100, and image span[1900, 3948).Full-model VL accuracy, generation-length evaluation, and AMD runtime validation were not run. The evidence here establishes attention-window correctness at the metadata/kernel level.
Related work
Open-PR searches for SWA/gather, V4.1/prefill, and V4.1/causal/image found no other matching fix. #56986 concerns cross-layer compressed-KV gather reuse; it does not correct the image attention-window semantics addressed here.
AI assistance
AI assistance was used for analysis, implementation, and validation.