Skip to content

[Feature][Model] Align DeepSeek V4 Vision inputs and cached-image execution - #15783

Open
QwertyJack wants to merge 2 commits into
vllm-project:mainfrom
QwertyJack:codex/dsv4-vision-alignment
Open

QwertyJack wants to merge 2 commits into
vllm-project:mainfrom
QwertyJack:codex/dsv4-vision-alignment

Conversation

@QwertyJack

@QwertyJack QwertyJack commented Sep 4, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

DeepSeek V4 Vision uses the official image-content ordering and two-newline separator during encoding. The supported vLLM parser defaults to front-padding multimodal placeholders and joins interleaved parts with a single newline. This PR applies the behavior only to DeepSeek V4 checkpoints with vision_n_layers > 0.

It adds processor golden coverage derived from the published DeepSeek-V4-Flash-Vision-Exp inference/image_processor.py: square, panorama, portrait, min-pixel, and max-token images; patch tensor SHA256; ViT/LLM grids; sentinel types; perm; and compressor position offsets 0..3.

It also fixes image-prefill execution when encoder outputs are cached. For models requiring raw input tokens, the dummy run compiles a text-only path with inputs_embeds=None. Checking only scheduled_encoder_inputs misses cached images, allowing that compiled path to drop the merged image embeddings. The runner now checks whether scheduled tokens overlap a multimodal placeholder range. Image-prefill chunks use the uncompiled path even on an encoder-cache hit; text-only chunks and decode after the image retain the compiled path. Existing encoder-decoder behavior is preserved.

#15457 is already merged and provides the Ascend model and processor implementation. This PR is independently mergeable on top of that support. #14632 covers general DeepSeek V4 frontend compatibility, while #15317 covers the trailing-system edge case; neither is a dependency of this PR. SCFA operator changes and the separate backport of vllm-project/vllm#54548 are not included here.

Refs #15462

Does this PR introduce any user-facing change?

Yes. DeepSeek V4 Vision requests preserve mixed text/image content order and the official two-newline separator, and cached-image prefill consumes image embeddings correctly. The runner fix applies to raw-token multimodal models; text-only requests retain their existing execution path.

How was this patch tested?

  • python -m pytest -q tests/ut/worker/test_model_runner_v1.py tests/ut/patch/platform/test_deepseek_v4_vision.py tests/ut/models/test_deepseek_v4_vision_preprocess.py tests/ut/models/test_deepseek_v4_vision.py: 60 passed against the local vLLM 0.27.1 integration environment. That environment also contains the separate sparse-placeholder-mask fix.

  • New runner tests cover fresh/cached encoder outputs, placeholder boundaries, chunked prefill, mixed batches, text-only/non-raw-token models, and encoder-decoder behavior.

  • The same runner fix was exercised with real W8A8 weights on the local TP8/DP2/EP16 integration stack, DSpark6, FULL_DECODE_ONLY, async scheduling, and CP disabled. Integrated results were GPQA 178/198 and OCRBench 822/1000 (KIE 181). These results include other model/operator and local numerical-stability patches and are not isolated measurements of this PR or a full numerical/performance sign-off. Those local numerical patches are not included in this PR.

  • Formatting and local checks passed except the offline gitleaks hook, whose downloaded binary cannot execute on this ARM64 host (Exec format error). Remote CI remains the merge gate.

  • vLLM main: vllm-project/vllm@ba07e4a

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request aligns the vLLM-Ascend DeepSeek V4 Vision implementation with the model's official requirements. It introduces a patch to the vLLM multimodal parser to ensure correct content-block ordering and separator usage, and adds extensive unit tests to validate the image processor against official golden outputs.

Highlights

  • Parser Alignment: Patched the vLLM multimodal parser and prompt generator to preserve content-block order and enforce a two-newline separator for DeepSeek V4 Vision models.
  • Processor Validation: Added comprehensive golden tests for the image processor, validating tensor shapes, hashes, and grid properties across various image aspect ratios.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Add compatibility patches and tests for DeepSeek-V4 vision multimodal parser

Suggested PR Summary:

### What this PR does / why we need it?
This PR adds compatibility patches for DeepSeek-V4 vision in vLLM v0.27. It patches `_parse_chat_message_content_parts` and `_get_full_multimodal_text_prompt` in `vllm.entrypoints.chat_utils` to preserve the order of interleaved image placeholders and use a double-newline separator (`\n\n`) for DeepSeek-V4 vision models. It also adds unit tests to verify the image processor outputs against official golden values and to validate the behavior of the patched multimodal parser.

Feedback: The review comments suggest adding defensive programming checks to safely access `tracker.model_config` and to handle potential signature changes in upstream vLLM functions to prevent `AttributeError` or `KeyError`.

### Does this PR introduce _any_ user-facing change?
No, this is an internal compatibility patch for DeepSeek-V4 vision model support on Ascend.

### How was this patch tested?
Tested with new unit tests in `tests/ut/models/test_deepseek_v4_vision_preprocess.py` and `tests/ut/patch/platform/test_deepseek_v4_vision.py`.

Comment thread vllm_ascend/patch/platform/patch_deepseek_v4_vision.py
Comment thread vllm_ascend/patch/platform/patch_deepseek_v4_vision.py Outdated
@QwertyJack
QwertyJack force-pushed the codex/dsv4-vision-alignment branch from c03d51a to 09c835f Compare September 4, 2026 07:43
@QwertyJack
QwertyJack force-pushed the codex/dsv4-vision-alignment branch from 09c835f to f7f6fe2 Compare September 4, 2026 08:56
Preserve multimodal content-part order and the official two-newline separator for DeepSeek V4 vision checkpoints. Add processor golden coverage for image resizing, patch tensors, grids, sentinel layouts, permutations, and position-dependent padding.

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
@QwertyJack
QwertyJack force-pushed the codex/dsv4-vision-alignment branch from f7f6fe2 to cd1b050 Compare September 4, 2026 09:13
@QwertyJack QwertyJack added the ready-precise run selected e2e test for pr label Sep 4, 2026
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
@QwertyJack QwertyJack changed the title [Feature][Model] Align DeepSeek V4 Vision image renderer and processor golden [Feature][Model] Align DeepSeek V4 Vision inputs and cached-image execution Sep 5, 2026
@QwertyJack

QwertyJack commented Sep 5, 2026 •

Copy link
Copy Markdown
Collaborator Author

/e2e tests/e2e/pull_request/four_card/test_pipeline_parallel.py
[Bot]: e2e command triggered. See workflow run for details.
[Bot]: e2e command completed successfully.

@QwertyJack

QwertyJack commented Sep 5, 2026 •

Copy link
Copy Markdown
Collaborator Author

/rerun
[Bot]: rerun completed.

Rerun (failed jobs only):

  • E2E

@github-actions

github-actions Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant