Skip to content

[BugFix][Model] Align DeepSeek V4 Vision runtime contracts - #15740

Open
QwertyJack wants to merge 2 commits into
vllm-project:mainfrom
QwertyJack:codex/dsv4-vision-runtime-metadata
Open

QwertyJack wants to merge 2 commits into
vllm-project:mainfrom
QwertyJack:codex/dsv4-vision-runtime-metadata

Conversation

@QwertyJack

@QwertyJack QwertyJack commented Sep 4, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

This is a runtime-only follow-up stacked on #15457. It hardens the
DeepSeek-V4-Flash-Vision-Exp metadata path for graph padding, multimodal prefix
cache hits, and compressed attention without including the SCFA/SAS operator
changes that still require separate review.

Related roadmap: #15462.

What changed

  • Preserve active image document ranges when attention metadata is unpadded.
  • Build image ranges only for active image/video requests and reject image spans
    split across prefill chunks.
  • Keep a stable compressor seqused buffer, zero graph padding and stale rows,
    and pass it to the compressor operator.
  • Propagate the actual request count through empty DP dummy runs.
  • Keep encoder and multimodal prefill, including prefix-cache-hit prefill,
    outside the compiled backbone while retaining compiled text and decode paths.
  • Replace the vision router sentinel magic number with
    IMAGE_SENTINEL_BASE_ID.

Does this PR introduce any user-facing change?

Yes. DeepSeek V4 Vision requests retain the correct active-request metadata in
ACLGraph and automatic-prefix-cache paths, and partial image spans fail
explicitly instead of silently constructing invalid bidirectional indices.

How was this patch tested?

Current rebased commit:

pytest -q \
  tests/ut/attention/test_attention_v1.py \
  tests/ut/attention/test_dsa_compressor_seqused.py \
  tests/ut/worker/test_model_runner_v1.py
# 85 passed

ruff format --check <8 changed Python files>
# 8 files already formatted

ruff check <8 changed Python files>
# All checks passed!

git diff --check <PR #15457 head>...HEAD
# passed

Fresh real-weight evidence was collected on 2026-09-04 before rebasing this
patch onto the latest public head of PR #15457. The semantic runtime changes
above are the same, but the service matrix has not been rerun on this rebased
commit yet.

  • checkpoint: /models/DeepSeek-V4-Flash-Vision-Exp-w8a8, 78 shards;
  • checkpoint config SHA256:
    699e220611e859125839a6ca4268c36e1600df1388a5cda2944e77552c781319;
  • checkpoint index SHA256:
    49487e0c55d0134bbcefa4566e97e08ecda58325adb200ba71b9178709d5a0a4;
  • topology: TP8 / DP2 / EP16, ENABLE_DSA_CP=0, ACLGraph
    FULL_DECODE_ONLY, async scheduling enabled;
  • /v1/models, text, render, single-image, and multi-image smoke: PASS;
  • exact 131,072-token prompt plus 8-token completion: HTTP 200, PASS.

Full task accuracy, DSpark acceptance, and the complete baseline/DSpark
performance comparison are still running and are not claimed by this PR.

Scope boundaries

GDzhu01 and others added 2 commits September 4, 2026 10:55
Add the DeepSeek V4 multimodal preprocessing and vision runtime, Ascend quantization and MoE routing integration, bidirectional vision attention, and DSpark support. Isolate tokenizer backends per preprocessing thread to make concurrent multimodal requests safe.

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Co-authored-by: MengLong Chen <71744434+dragondream-chen@users.noreply.github.com>
Co-authored-by: RenYuKai <184603735+pgzddxx@users.noreply.github.com>
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request aligns the runtime contracts for the DeepSeek-V4 Vision model, focusing on robust metadata handling for multimodal requests. It introduces structural improvements to attention metadata, compressor buffer management, and compilation guards to ensure stable performance across various prefill and decode scenarios.

Highlights

  • Vision Metadata Hardening: Added logic to preserve active image document ranges and reject invalid image spans during unpadded attention metadata construction.
  • Compressor Stability: Implemented a stable seqused buffer for the compressor operator to correctly mask graph padding and stale rows.
  • Runtime Compilation: Excluded encoder and multimodal prefill paths from the compiled backbone to ensure compatibility with dynamic input structures.
  • Vision Router: Standardized the vision router sentinel ID using IMAGE_SENTINEL_BASE_ID and enabled bias propagation for vision-specific expert routing.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Attention][Feature] Add support for DeepSeek-V4 vision model on Ascend

Suggested PR Summary:

### What this PR does / why we need it?
This pull request adds support for the DeepSeek-V4 multimodal vision model (`DeepseekV4-Flash-Vision-Exp`) on Ascend. It introduces the `AscendDeepseekV4ForConditionalGeneration` model wrapper, multimodal preprocessing, bidirectional sliding window attention (SWA) indices for vision, and MoE routing updates to handle vision-specific expert routing bias (`bias_vl`). Additionally, it includes compatibility patches for speculative decoding and quantization prefix mapping.

Feedback has been provided to address performance and robustness issues:
- Avoid CPU-GPU synchronizations in `build_vision_bidirectional_swa_indices` by performing overlap checks on the CPU.
- Avoid redundant tensor allocations in `embed_input_ids` during decode steps by checking `sentinel_mask.any()` before processing sentinel tokens.
- Prevent potential `KeyError` crashes in `_get_block_table_and_slot_mapping` during dummy runs by using `self.requests.get(req_id)` instead of direct indexing.

### Does this PR introduce _any_ user-facing change?
Yes, it adds support for serving DeepSeek-V4 vision models on Ascend.

### How was this patch tested?
Tested with newly added unit tests under `tests/ut/` covering the vision SWA indices, MoE routing, preprocessing, and model loading.

Comment on lines +457 to +471
for req_idx, ranges in mm_prefix_ranges.items():
if req_idx >= query_lens.shape[0]:
continue
for span_start, span_end in ranges:
if span_end < span_start:
raise ValueError(f"Invalid image span [{span_start}, {span_end}]")
if span_end - span_start + 1 > max_image_tokens:
raise ValueError(
f"Image span exceeds vision_max_n_token: span=[{span_start}, {span_end}], max={max_image_tokens}"
)
in_span = (req_ids == req_idx) & (positions >= span_start) & (positions <= span_end)
if bool(in_span.any().item()) and span_end >= int(seq_lens[req_idx].item()):
raise ValueError(
"Image spans must be scheduled in a single prefill chunk before bidirectional attention is built"
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Calling in_span.any().item() and seq_lens[req_idx].item() inside nested loops over requests and image spans causes multiple CPU-GPU synchronizations per step. Since this is executed in the hot path of every prefill step, these synchronizations can severely degrade performance.

We can completely eliminate these synchronizations by copying seq_lens and query_lens to CPU once before the loop, and performing the overlap check entirely on the CPU using interval overlap logic: has_overlap = (span_start <= seq_len - 1) and (span_end >= seq_len - query_len).

    seq_lens_cpu = seq_lens.cpu().tolist()
    query_lens_cpu = (query_start_loc[1:] - query_start_loc[:-1]).cpu().tolist()
    for req_idx, ranges in mm_prefix_ranges.items():
        if req_idx >= len(query_lens_cpu):
            continue
        seq_len = seq_lens_cpu[req_idx]
        query_len = query_lens_cpu[req_idx]
        for span_start, span_end in ranges:
            if span_end < span_start:
                raise ValueError(f"Invalid image span [{span_start}, {span_end}]")
            if span_end - span_start + 1 > max_image_tokens:
                raise ValueError(
                    f"Image span exceeds vision_max_n_token: span=[{span_start}, {span_end}], max={max_image_tokens}"
                )
            in_span = (req_ids == req_idx) & (positions >= span_start) & (positions <= span_end)
            has_overlap = (span_start <= seq_len - 1) and (span_end >= seq_len - query_len)
            if has_overlap and span_end >= seq_len:
                raise ValueError(
                    "Image spans must be scheduled in a single prefill chunk before bidirectional attention is built"
                )

Comment on lines +183 to +197
if self.image_start is not None:
sentinel_mask = image_sentinel_mask(input_ids)
if is_multimodal is not None:
sentinel_mask = sentinel_mask & ~is_multimodal.to(input_ids.device)
table = torch.stack(
[
self.image_start,
self.image_pad,
self.image_pad,
self.image_newline,
self.image_end,
]
).to(inputs_embeds.dtype)
idx = (input_ids - IMAGE_SENTINEL_BASE_ID).clamp(0, 4)
inputs_embeds = torch.where(sentinel_mask.unsqueeze(-1), table[idx], inputs_embeds)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The embed_input_ids method always stacks the sentinel parameters, clamps the input IDs, indexes the table, and performs torch.where on every single execution step, even during decode steps where no sentinel tokens are present in input_ids.

This introduces significant overhead (tensor allocations, stacking, clamping, and indexing) on the critical path of every decode step. Wrapping this logic in if sentinel_mask.any(): completely bypasses these operations and allocations during decode steps, which represent the vast majority of the model's execution steps.

Suggested change
if self.image_start is not None:
sentinel_mask = image_sentinel_mask(input_ids)
if is_multimodal is not None:
sentinel_mask = sentinel_mask & ~is_multimodal.to(input_ids.device)
table = torch.stack(
[
self.image_start,
self.image_pad,
self.image_pad,
self.image_newline,
self.image_end,
]
).to(inputs_embeds.dtype)
idx = (input_ids - IMAGE_SENTINEL_BASE_ID).clamp(0, 4)
inputs_embeds = torch.where(sentinel_mask.unsqueeze(-1), table[idx], inputs_embeds)
if self.image_start is not None:
sentinel_mask = image_sentinel_mask(input_ids)
if is_multimodal is not None:
sentinel_mask = sentinel_mask & ~is_multimodal.to(input_ids.device)
if sentinel_mask.any():
table = torch.stack(
[
self.image_start,
self.image_pad,
self.image_pad,
self.image_newline,
self.image_end,
]
).to(inputs_embeds.dtype)
idx = (input_ids - IMAGE_SENTINEL_BASE_ID).clamp(0, 4)
inputs_embeds = torch.where(sentinel_mask.unsqueeze(-1), table[idx], inputs_embeds)

Comment on lines +3265 to +3266
req_state = self.requests[req_id]
for mm_feature in req_state.mm_features or ():

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Accessing self.requests[req_id] directly can raise a KeyError during dummy runs or profiling steps where self.requests may not be fully populated with the dummy request IDs.

Using self.requests.get(req_id) and checking for None prevents potential engine crashes during startup and initialization.

                req_state = self.requests.get(req_id)
                mm_features = req_state.mm_features if req_state is not None else None
                for mm_feature in mm_features or ():

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

1 similar comment
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants