Skip to content

[Doc][Model] Add DeepSeek V4 Flash Vision deployment guide - #15793

Merged
weijinqian0 merged 10 commits into
vllm-project:mainfrom
GDzhu01:codex/docs-deepseek-v4-vision
Sep 7, 2026
Merged

weijinqian0 merged 10 commits into
vllm-project:mainfrom
GDzhu01:codex/docs-deepseek-v4-vision

Conversation

@GDzhu01

@GDzhu01 GDzhu01 commented Sep 4, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Adds the deployment tutorial and supported-model matrix entry for DeepSeek-V4-Flash-Vision-Exp, following the model support merged in #15457.

The tutorial documents the current experimental W8A8 configurations:

  • one Atlas 800 A3 server (128GB × 8 NPUs), using 16 logical devices with TP4/DP4/EP16;
  • two Atlas 800 A2 servers (64GB × 8 NPU dies each), with two local DP ranks per node and global TP4/DP4/EP16;
  • hardware-specific A3/A2 installation tabs, published images, and the public ModelScope W8A8 QuaRot checkpoint;
  • colocated serving with FULL_DECODE_ONLY ACL Graph and a 130K maximum model length;
  • installation and service verification commands with complete expected output examples;
  • OpenAI-compatible multimodal functional verification;
  • the validated OCRBench V1 result and current limitations.

Prefill-Decode disaggregation is explicitly marked unsupported. The matrix leaves unverified features experimental or unverified rather than claiming support.

Related roadmap: #15462

Does this PR introduce any user-facing change?

No. This is a documentation-only update for an already merged model integration.

How was this patch tested?

  • markdownlint-cli@0.45.0 passes for both changed Markdown files.

  • git diff --check passes.

  • All relative links and referenced anchors in the new tutorial resolve.

  • The Hugging Face and ModelScope model URLs return HTTP 200.

  • Quay API confirms both published image tags are active.

  • The new supported-model matrix row has the same 20 columns as its header.

  • The tutorial is registered exactly once in mkdocs.yml and has matching English and Chinese entries in docs/hooks/nav_titles.py.

  • docs/hooks/nav_titles.py passes Python syntax compilation.

  • Automated scope checks confirm Chapter 9 is unchanged; the FAQ contains only the public FAQ reference requested in review.

  • vLLM main: vllm-project/vllm@e6bfe03

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
@GDzhu01
GDzhu01 requested review from LCAIZJ and Yikun as code owners September 4, 2026 08:21
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request provides documentation for the DeepSeek-V4-Flash-Vision-Exp model on vLLM Ascend. It outlines the validated W8A8 configuration, hardware requirements, and deployment procedures for single-node colocated setups, while explicitly noting current limitations and unsupported features to guide users effectively.

Highlights

  • Documentation Addition: Added a comprehensive deployment tutorial for the DeepSeek-V4-Flash-Vision-Exp model, detailing validated configurations for Ascend A3 hardware.
  • Support Matrix Update: Updated the supported-model matrix to include DeepSeek-V4-Flash-Vision-Exp, reflecting its experimental status and specific deployment constraints.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds a deployment tutorial and updates the supported models matrix for the experimental DeepSeek-V4-Flash-Vision-Exp model on vLLM Ascend. The feedback points out that the PR title and summary do not adhere to the repository's style guide and provides a compliant, structured suggestion.

@@ -0,0 +1,208 @@
# DeepSeek-V4-Flash-Vision-Exp (Experimental)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The Pull Request title and summary do not fully adhere to the repository's PR Summary Style Guide. Specifically, the module [Model] is not one of the standard modules, and the action prefix (e.g., [Feature]) is missing.

Here are the suggested PR Title and PR Summary formatted according to the repository style guide:

Suggested PR Title:

[Doc][Feature] Add DeepSeek V4 Flash Vision deployment guide

Suggested PR Summary:

### What this PR does / why we need it?

Adds the deployment tutorial and supported-model matrix entry for `DeepSeek-V4-Flash-Vision-Exp`, following the model support merged in #15457.

The tutorial documents the currently validated configuration:
- experimental W8A8 serving on one Atlas 800 A3 server with 16 NPUs;
- colocated TP4/DP4/EP16 deployment;
- `FULL_DECODE_ONLY` ACL Graph with a 130K maximum model length;
- OpenAI-compatible multimodal functional verification;
- the validated OCRBench V1 result and current limitations.

Prefill-Decode disaggregation is explicitly marked unsupported. The matrix leaves unverified features experimental or unverified rather than claiming support.

Fixes #15462

### Does this PR introduce _any_ user-facing change?

No. This is a documentation-only update for an already merged model integration.

### How was this patch tested?

- `markdownlint-cli@0.45.0` passes for both changed Markdown files.
- `git diff --check` passes.
- All relative links in the new tutorial resolve to existing repository files.
- The new external Hugging Face model link returns HTTP 200.
- The new supported-model matrix row has the same 20 columns as its header.
References
  1. The PR title and summary must follow the repository style guide format, including specific prefixes for modules and actions, and be presented in markdown code blocks. (link)

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 4, 2026
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


| Model | Download | Hardware requirements |
| --- | --- | --- |
| DeepSeek-V4-Flash-Vision-Exp-w8a8-QuaRot | [ModelScope](https://www.modelscope.cn/models/Eco-Tech/DeepSeek-V4-Flash-Vision-Exp-w8a8-QuaRot) | One Atlas 800 A3 server with 16 visible NPUs |

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One Atlas 800 A3 server with 16 visible NPUs - we always say 8 cards × 128GB/card

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
@zhujianwei-ops

Copy link
Copy Markdown

why deployment guide use dp4 tp4? i think dp2 tp8 is better

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
| Model | Download | Hardware requirements |
| --- | --- | --- |
| DeepSeek-V4-Flash-Vision-Exp-w8a8-QuaRot | [ModelScope](https://www.modelscope.cn/models/Eco-Tech/DeepSeek-V4-Flash-Vision-Exp-w8a8-QuaRot) | One Atlas 800 A3 server (128GB × 8 NPUs) |
| DeepSeek-V4-Flash-Vision-Exp | [Hugging Face](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp) | Original model weights; use the ModelScope W8A8 QuaRot checkpoint for the deployment in this guide |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

模板要求权重链接需要提供ModelScope及Hugging Face,若有请补充


## 4 Installation

### 4.1 Docker Image Installation

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

若仅支持A3,需按模板要求在安装章节开头明确说明,若支持多硬件系列(如 A3/A2 系列),须使用标签页语法展示

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated: restored A2 support and now present A3/A2 container installation with MkDocs tab syntax.

Do not replace either package inside the container independently. Mixing other
vLLM and vLLM Ascend revisions is not supported.

## 4 Installation

@herizhen herizhen Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

需要提供验证命令及预期状态:指导用户通过执行命令(如 docker ps)检查安装结果,说明成功时的状态码或输出特征,并提供完整的回显信息示例。

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated: added a docker ps verification command, the success criterion, and a complete expected output example.

[performance tuning guide](../../developer_guide/performance_and_debug/optimization_and_tuning.md)
for general tuning methods.

## 10 Limitations

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

未提供FAQ章节内容

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated: added Chapter 10 FAQ with the requested reference to the public FAQs; the existing limitations section is retained as Chapter 11.

with the deployment in Section 5.1, and report image resolution, input/output
lengths, request concurrency, TTFT, TPOT, ITL, and throughput.

## 9 Performance Tuning

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

该章节文档结构在模板的要求如下所示:
9 性能调优
├── 9.1 推荐配置
└── 9.2 调优思路
├── 9.2.1 模型特有优化
└── 9.2.2 通用调优参考
请按要求补全相关内容

INFO: Waiting for application startup.
INFO: Application startup complete.
```

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

需提供服务验证方法(如 curl 命令)及预期结果,说明成功特征(如 200 OK),并提供回显示例。

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated: added /health and /v1/models verification commands, HTTP 200 and model-id success criteria, and complete expected output examples for A3 and A2.

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
@weijinqian0
weijinqian0 enabled auto-merge (squash) September 7, 2026 07:34
@weijinqian0
weijinqian0 merged commit 620fd1d into vllm-project:main Sep 7, 2026
19 checks passed
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
…ect#15793)

### What this PR does / why we need it?

Adds the deployment tutorial and supported-model matrix entry for
`DeepSeek-V4-Flash-Vision-Exp`, following the model support merged in
vllm-project#15457.

The tutorial documents the current experimental W8A8 configurations:

- one Atlas 800 A3 server (128GB × 8 NPUs), using 16 logical devices
with TP4/DP4/EP16;
- two Atlas 800 A2 servers (64GB × 8 NPU dies each), with two local DP
ranks per node and global TP4/DP4/EP16;
- hardware-specific A3/A2 installation tabs, published images, and the
public ModelScope W8A8 QuaRot checkpoint;
- colocated serving with `FULL_DECODE_ONLY` ACL Graph and a 130K maximum
model length;
- installation and service verification commands with complete expected
output examples;
- OpenAI-compatible multimodal functional verification;
- the validated OCRBench V1 result and current limitations.

Prefill-Decode disaggregation is explicitly marked unsupported. The
matrix leaves unverified features experimental or unverified rather than
claiming support.

Related roadmap: vllm-project#15462

### Does this PR introduce _any_ user-facing change?

No. This is a documentation-only update for an already merged model
integration.

### How was this patch tested?

- `markdownlint-cli@0.45.0` passes for both changed Markdown files.
- `git diff --check` passes.
- All relative links and referenced anchors in the new tutorial resolve.
- The Hugging Face and ModelScope model URLs return HTTP 200.
- Quay API confirms both published image tags are active.
- The new supported-model matrix row has the same 20 columns as its
header.
- The tutorial is registered exactly once in `mkdocs.yml` and has
matching English and Chinese entries in `docs/hooks/nav_titles.py`.
- `docs/hooks/nav_titles.py` passes Python syntax compilation.
- Automated scope checks confirm Chapter 9 is unchanged; the FAQ
contains only the public FAQ reference requested in review.


- vLLM main:
vllm-project/vllm@e6bfe03

---------

Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants