[Feature][Model] Align DeepSeek V4 Vision inputs and cached-image execution - #15783
QwertyJack wants to merge 2 commits into
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request aligns the vLLM-Ascend DeepSeek V4 Vision implementation with the model's official requirements. It introduces a patch to the vLLM multimodal parser to ensure correct content-block ordering and separator usage, and adds extensive unit tests to validate the image processor against official golden outputs. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Ops][Feature] Add compatibility patches and tests for DeepSeek-V4 vision multimodal parserSuggested PR Summary:
### What this PR does / why we need it?
This PR adds compatibility patches for DeepSeek-V4 vision in vLLM v0.27. It patches `_parse_chat_message_content_parts` and `_get_full_multimodal_text_prompt` in `vllm.entrypoints.chat_utils` to preserve the order of interleaved image placeholders and use a double-newline separator (`\n\n`) for DeepSeek-V4 vision models. It also adds unit tests to verify the image processor outputs against official golden values and to validate the behavior of the patched multimodal parser.
Feedback: The review comments suggest adding defensive programming checks to safely access `tracker.model_config` and to handle potential signature changes in upstream vLLM functions to prevent `AttributeError` or `KeyError`.
### Does this PR introduce _any_ user-facing change?
No, this is an internal compatibility patch for DeepSeek-V4 vision model support on Ascend.
### How was this patch tested?
Tested with new unit tests in `tests/ut/models/test_deepseek_v4_vision_preprocess.py` and `tests/ut/patch/platform/test_deepseek_v4_vision.py`.c03d51a to
09c835f
Compare
09c835f to
f7f6fe2
Compare
Preserve multimodal content-part order and the official two-newline separator for DeepSeek V4 vision checkpoints. Add processor golden coverage for image resizing, patch tensors, grids, sentinel layouts, permutations, and position-dependent padding. Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
f7f6fe2 to
cd1b050
Compare
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
|
/e2e tests/e2e/pull_request/four_card/test_pipeline_parallel.py |
|
/rerun Rerun (failed jobs only):
|
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
What this PR does / why we need it?
DeepSeek V4 Vision uses the official image-content ordering and two-newline separator during encoding. The supported vLLM parser defaults to front-padding multimodal placeholders and joins interleaved parts with a single newline. This PR applies the behavior only to DeepSeek V4 checkpoints with
vision_n_layers > 0.It adds processor golden coverage derived from the published DeepSeek-V4-Flash-Vision-Exp
inference/image_processor.py: square, panorama, portrait, min-pixel, and max-token images; patch tensor SHA256; ViT/LLM grids; sentineltypes;perm; and compressor position offsets 0..3.It also fixes image-prefill execution when encoder outputs are cached. For models requiring raw input tokens, the dummy run compiles a text-only path with
inputs_embeds=None. Checking onlyscheduled_encoder_inputsmisses cached images, allowing that compiled path to drop the merged image embeddings. The runner now checks whether scheduled tokens overlap a multimodal placeholder range. Image-prefill chunks use the uncompiled path even on an encoder-cache hit; text-only chunks and decode after the image retain the compiled path. Existing encoder-decoder behavior is preserved.#15457 is already merged and provides the Ascend model and processor implementation. This PR is independently mergeable on top of that support. #14632 covers general DeepSeek V4 frontend compatibility, while #15317 covers the trailing-system edge case; neither is a dependency of this PR. SCFA operator changes and the separate backport of vllm-project/vllm#54548 are not included here.
Refs #15462
Does this PR introduce any user-facing change?
Yes. DeepSeek V4 Vision requests preserve mixed text/image content order and the official two-newline separator, and cached-image prefill consumes image embeddings correctly. The runner fix applies to raw-token multimodal models; text-only requests retain their existing execution path.
How was this patch tested?
python -m pytest -q tests/ut/worker/test_model_runner_v1.py tests/ut/patch/platform/test_deepseek_v4_vision.py tests/ut/models/test_deepseek_v4_vision_preprocess.py tests/ut/models/test_deepseek_v4_vision.py: 60 passed against the local vLLM 0.27.1 integration environment. That environment also contains the separate sparse-placeholder-mask fix.New runner tests cover fresh/cached encoder outputs, placeholder boundaries, chunked prefill, mixed batches, text-only/non-raw-token models, and encoder-decoder behavior.
The same runner fix was exercised with real W8A8 weights on the local TP8/DP2/EP16 integration stack, DSpark6, FULL_DECODE_ONLY, async scheduling, and CP disabled. Integrated results were GPQA 178/198 and OCRBench 822/1000 (KIE 181). These results include other model/operator and local numerical-stability patches and are not isolated measurements of this PR or a full numerical/performance sign-off. Those local numerical patches are not included in this PR.
Formatting and local checks passed except the offline gitleaks hook, whose downloaded binary cannot execute on this ARM64 host (
Exec format error). Remote CI remains the merge gate.vLLM main: vllm-project/vllm@ba07e4a