[BugFix][Model] Align DeepSeek V4 Vision runtime contracts - #15740
QwertyJack wants to merge 2 commits into
Conversation
Add the DeepSeek V4 multimodal preprocessing and vision runtime, Ascend quantization and MoE routing integration, bidirectional vision attention, and DSpark support. Isolate tokenizer backends per preprocessing thread to make concurrent multimodal requests safe. Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Co-authored-by: MengLong Chen <71744434+dragondream-chen@users.noreply.github.com> Co-authored-by: RenYuKai <184603735+pgzddxx@users.noreply.github.com>
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request aligns the runtime contracts for the DeepSeek-V4 Vision model, focusing on robust metadata handling for multimodal requests. It introduces structural improvements to attention metadata, compressor buffer management, and compilation guards to ensure stable performance across various prefill and decode scenarios. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Attention][Feature] Add support for DeepSeek-V4 vision model on AscendSuggested PR Summary:
### What this PR does / why we need it?
This pull request adds support for the DeepSeek-V4 multimodal vision model (`DeepseekV4-Flash-Vision-Exp`) on Ascend. It introduces the `AscendDeepseekV4ForConditionalGeneration` model wrapper, multimodal preprocessing, bidirectional sliding window attention (SWA) indices for vision, and MoE routing updates to handle vision-specific expert routing bias (`bias_vl`). Additionally, it includes compatibility patches for speculative decoding and quantization prefix mapping.
Feedback has been provided to address performance and robustness issues:
- Avoid CPU-GPU synchronizations in `build_vision_bidirectional_swa_indices` by performing overlap checks on the CPU.
- Avoid redundant tensor allocations in `embed_input_ids` during decode steps by checking `sentinel_mask.any()` before processing sentinel tokens.
- Prevent potential `KeyError` crashes in `_get_block_table_and_slot_mapping` during dummy runs by using `self.requests.get(req_id)` instead of direct indexing.
### Does this PR introduce _any_ user-facing change?
Yes, it adds support for serving DeepSeek-V4 vision models on Ascend.
### How was this patch tested?
Tested with newly added unit tests under `tests/ut/` covering the vision SWA indices, MoE routing, preprocessing, and model loading.| for req_idx, ranges in mm_prefix_ranges.items(): | ||
| if req_idx >= query_lens.shape[0]: | ||
| continue | ||
| for span_start, span_end in ranges: | ||
| if span_end < span_start: | ||
| raise ValueError(f"Invalid image span [{span_start}, {span_end}]") | ||
| if span_end - span_start + 1 > max_image_tokens: | ||
| raise ValueError( | ||
| f"Image span exceeds vision_max_n_token: span=[{span_start}, {span_end}], max={max_image_tokens}" | ||
| ) | ||
| in_span = (req_ids == req_idx) & (positions >= span_start) & (positions <= span_end) | ||
| if bool(in_span.any().item()) and span_end >= int(seq_lens[req_idx].item()): | ||
| raise ValueError( | ||
| "Image spans must be scheduled in a single prefill chunk before bidirectional attention is built" | ||
| ) |
There was a problem hiding this comment.
Calling in_span.any().item() and seq_lens[req_idx].item() inside nested loops over requests and image spans causes multiple CPU-GPU synchronizations per step. Since this is executed in the hot path of every prefill step, these synchronizations can severely degrade performance.
We can completely eliminate these synchronizations by copying seq_lens and query_lens to CPU once before the loop, and performing the overlap check entirely on the CPU using interval overlap logic: has_overlap = (span_start <= seq_len - 1) and (span_end >= seq_len - query_len).
seq_lens_cpu = seq_lens.cpu().tolist()
query_lens_cpu = (query_start_loc[1:] - query_start_loc[:-1]).cpu().tolist()
for req_idx, ranges in mm_prefix_ranges.items():
if req_idx >= len(query_lens_cpu):
continue
seq_len = seq_lens_cpu[req_idx]
query_len = query_lens_cpu[req_idx]
for span_start, span_end in ranges:
if span_end < span_start:
raise ValueError(f"Invalid image span [{span_start}, {span_end}]")
if span_end - span_start + 1 > max_image_tokens:
raise ValueError(
f"Image span exceeds vision_max_n_token: span=[{span_start}, {span_end}], max={max_image_tokens}"
)
in_span = (req_ids == req_idx) & (positions >= span_start) & (positions <= span_end)
has_overlap = (span_start <= seq_len - 1) and (span_end >= seq_len - query_len)
if has_overlap and span_end >= seq_len:
raise ValueError(
"Image spans must be scheduled in a single prefill chunk before bidirectional attention is built"
)| if self.image_start is not None: | ||
| sentinel_mask = image_sentinel_mask(input_ids) | ||
| if is_multimodal is not None: | ||
| sentinel_mask = sentinel_mask & ~is_multimodal.to(input_ids.device) | ||
| table = torch.stack( | ||
| [ | ||
| self.image_start, | ||
| self.image_pad, | ||
| self.image_pad, | ||
| self.image_newline, | ||
| self.image_end, | ||
| ] | ||
| ).to(inputs_embeds.dtype) | ||
| idx = (input_ids - IMAGE_SENTINEL_BASE_ID).clamp(0, 4) | ||
| inputs_embeds = torch.where(sentinel_mask.unsqueeze(-1), table[idx], inputs_embeds) |
There was a problem hiding this comment.
The embed_input_ids method always stacks the sentinel parameters, clamps the input IDs, indexes the table, and performs torch.where on every single execution step, even during decode steps where no sentinel tokens are present in input_ids.
This introduces significant overhead (tensor allocations, stacking, clamping, and indexing) on the critical path of every decode step. Wrapping this logic in if sentinel_mask.any(): completely bypasses these operations and allocations during decode steps, which represent the vast majority of the model's execution steps.
| if self.image_start is not None: | |
| sentinel_mask = image_sentinel_mask(input_ids) | |
| if is_multimodal is not None: | |
| sentinel_mask = sentinel_mask & ~is_multimodal.to(input_ids.device) | |
| table = torch.stack( | |
| [ | |
| self.image_start, | |
| self.image_pad, | |
| self.image_pad, | |
| self.image_newline, | |
| self.image_end, | |
| ] | |
| ).to(inputs_embeds.dtype) | |
| idx = (input_ids - IMAGE_SENTINEL_BASE_ID).clamp(0, 4) | |
| inputs_embeds = torch.where(sentinel_mask.unsqueeze(-1), table[idx], inputs_embeds) | |
| if self.image_start is not None: | |
| sentinel_mask = image_sentinel_mask(input_ids) | |
| if is_multimodal is not None: | |
| sentinel_mask = sentinel_mask & ~is_multimodal.to(input_ids.device) | |
| if sentinel_mask.any(): | |
| table = torch.stack( | |
| [ | |
| self.image_start, | |
| self.image_pad, | |
| self.image_pad, | |
| self.image_newline, | |
| self.image_end, | |
| ] | |
| ).to(inputs_embeds.dtype) | |
| idx = (input_ids - IMAGE_SENTINEL_BASE_ID).clamp(0, 4) | |
| inputs_embeds = torch.where(sentinel_mask.unsqueeze(-1), table[idx], inputs_embeds) |
| req_state = self.requests[req_id] | ||
| for mm_feature in req_state.mm_features or (): |
There was a problem hiding this comment.
Accessing self.requests[req_id] directly can raise a KeyError during dummy runs or profiling steps where self.requests may not be fully populated with the dummy request IDs.
Using self.requests.get(req_id) and checking for None prevents potential engine crashes during startup and initialization.
req_state = self.requests.get(req_id)
mm_features = req_state.mm_features if req_state is not None else None
for mm_feature in mm_features or ():|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [BugFix] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
1 similar comment
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
What this PR does / why we need it?
This is a runtime-only follow-up stacked on #15457. It hardens the
DeepSeek-V4-Flash-Vision-Exp metadata path for graph padding, multimodal prefix
cache hits, and compressed attention without including the SCFA/SAS operator
changes that still require separate review.
Related roadmap: #15462.
What changed
split across prefill chunks.
sequsedbuffer, zero graph padding and stale rows,and pass it to the compressor operator.
outside the compiled backbone while retaining compiled text and decode paths.
IMAGE_SENTINEL_BASE_ID.Does this PR introduce any user-facing change?
Yes. DeepSeek V4 Vision requests retain the correct active-request metadata in
ACLGraph and automatic-prefix-cache paths, and partial image spans fail
explicitly instead of silently constructing invalid bidirectional indices.
How was this patch tested?
Current rebased commit:
Fresh real-weight evidence was collected on 2026-09-04 before rebasing this
patch onto the latest public head of PR #15457. The semantic runtime changes
above are the same, but the service matrix has not been rerun on this rebased
commit yet.
/models/DeepSeek-V4-Flash-Vision-Exp-w8a8, 78 shards;699e220611e859125839a6ca4268c36e1600df1388a5cda2944e77552c781319;49487e0c55d0134bbcefa4566e97e08ecda58325adb200ba71b9178709d5a0a4;ENABLE_DSA_CP=0, ACLGraphFULL_DECODE_ONLY, async scheduling enabled;/v1/models, text, render, single-image, and multi-image smoke: PASS;Full task accuracy, DSpark acceptance, and the complete baseline/DSpark
performance comparison are still running and are not claimed by this PR.
Scope boundaries
No
csrc/or AscendC SCFA/SAS implementation is included. The current SCFAkernel still requires separate review before bidirectional original-KV
attention can be signed off end to end.
DSA context parallel is disabled in the validated topology and is not changed
here.
The independent zero-draft RNG fix remains in [BugFix][Spec Decode] Preserve RNG state for zero-draft requests #13755.
The independent frontend compatibility work remains in [BugFix][Frontend] Align DeepSeek V4 frontend behavior #14632.
The 131,072-token result is a single-request capacity check. Runtime logs
reported approximately 398,789 KV-cache tokens and about 3.0x maximum
concurrency at 133,120 tokens per request; this is not a
128K x bs16claim.vLLM main: vllm-project/vllm@ba07e4a