[Doc][Model] Add DeepSeek V4 Flash Vision deployment guide - #15793
Conversation
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request provides documentation for the DeepSeek-V4-Flash-Vision-Exp model on vLLM Ascend. It outlines the validated W8A8 configuration, hardware requirements, and deployment procedures for single-node colocated setups, while explicitly noting current limitations and unsupported features to guide users effectively. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request adds a deployment tutorial and updates the supported models matrix for the experimental DeepSeek-V4-Flash-Vision-Exp model on vLLM Ascend. The feedback points out that the PR title and summary do not adhere to the repository's style guide and provides a compliant, structured suggestion.
| @@ -0,0 +1,208 @@ | |||
| # DeepSeek-V4-Flash-Vision-Exp (Experimental) | |||
There was a problem hiding this comment.
The Pull Request title and summary do not fully adhere to the repository's PR Summary Style Guide. Specifically, the module [Model] is not one of the standard modules, and the action prefix (e.g., [Feature]) is missing.
Here are the suggested PR Title and PR Summary formatted according to the repository style guide:
Suggested PR Title:
[Doc][Feature] Add DeepSeek V4 Flash Vision deployment guideSuggested PR Summary:
### What this PR does / why we need it?
Adds the deployment tutorial and supported-model matrix entry for `DeepSeek-V4-Flash-Vision-Exp`, following the model support merged in #15457.
The tutorial documents the currently validated configuration:
- experimental W8A8 serving on one Atlas 800 A3 server with 16 NPUs;
- colocated TP4/DP4/EP16 deployment;
- `FULL_DECODE_ONLY` ACL Graph with a 130K maximum model length;
- OpenAI-compatible multimodal functional verification;
- the validated OCRBench V1 result and current limitations.
Prefill-Decode disaggregation is explicitly marked unsupported. The matrix leaves unverified features experimental or unverified rather than claiming support.
Fixes #15462
### Does this PR introduce _any_ user-facing change?
No. This is a documentation-only update for an already merged model integration.
### How was this patch tested?
- `markdownlint-cli@0.45.0` passes for both changed Markdown files.
- `git diff --check` passes.
- All relative links in the new tutorial resolve to existing repository files.
- The new external Hugging Face model link returns HTTP 200.
- The new supported-model matrix row has the same 20 columns as its header.References
- The PR title and summary must follow the repository style guide format, including specific prefixes for modules and actions, and be presented in markdown code blocks. (link)
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
|
|
||
| | Model | Download | Hardware requirements | | ||
| | --- | --- | --- | | ||
| | DeepSeek-V4-Flash-Vision-Exp-w8a8-QuaRot | [ModelScope](https://www.modelscope.cn/models/Eco-Tech/DeepSeek-V4-Flash-Vision-Exp-w8a8-QuaRot) | One Atlas 800 A3 server with 16 visible NPUs | |
There was a problem hiding this comment.
One Atlas 800 A3 server with 16 visible NPUs - we always say 8 cards × 128GB/card
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
|
why deployment guide use dp4 tp4? i think dp2 tp8 is better |
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
| | Model | Download | Hardware requirements | | ||
| | --- | --- | --- | | ||
| | DeepSeek-V4-Flash-Vision-Exp-w8a8-QuaRot | [ModelScope](https://www.modelscope.cn/models/Eco-Tech/DeepSeek-V4-Flash-Vision-Exp-w8a8-QuaRot) | One Atlas 800 A3 server (128GB × 8 NPUs) | | ||
| | DeepSeek-V4-Flash-Vision-Exp | [Hugging Face](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp) | Original model weights; use the ModelScope W8A8 QuaRot checkpoint for the deployment in this guide | |
There was a problem hiding this comment.
模板要求权重链接需要提供ModelScope及Hugging Face,若有请补充
|
|
||
| ## 4 Installation | ||
|
|
||
| ### 4.1 Docker Image Installation |
There was a problem hiding this comment.
若仅支持A3,需按模板要求在安装章节开头明确说明,若支持多硬件系列(如 A3/A2 系列),须使用标签页语法展示
There was a problem hiding this comment.
Updated: restored A2 support and now present A3/A2 container installation with MkDocs tab syntax.
| Do not replace either package inside the container independently. Mixing other | ||
| vLLM and vLLM Ascend revisions is not supported. | ||
|
|
||
| ## 4 Installation |
There was a problem hiding this comment.
需要提供验证命令及预期状态:指导用户通过执行命令(如 docker ps)检查安装结果,说明成功时的状态码或输出特征,并提供完整的回显信息示例。
There was a problem hiding this comment.
Updated: added a docker ps verification command, the success criterion, and a complete expected output example.
| [performance tuning guide](../../developer_guide/performance_and_debug/optimization_and_tuning.md) | ||
| for general tuning methods. | ||
|
|
||
| ## 10 Limitations |
There was a problem hiding this comment.
Updated: added Chapter 10 FAQ with the requested reference to the public FAQs; the existing limitations section is retained as Chapter 11.
| with the deployment in Section 5.1, and report image resolution, input/output | ||
| lengths, request concurrency, TTFT, TPOT, ITL, and throughput. | ||
|
|
||
| ## 9 Performance Tuning |
There was a problem hiding this comment.
该章节文档结构在模板的要求如下所示:
9 性能调优
├── 9.1 推荐配置
└── 9.2 调优思路
├── 9.2.1 模型特有优化
└── 9.2.2 通用调优参考
请按要求补全相关内容
| INFO: Waiting for application startup. | ||
| INFO: Application startup complete. | ||
| ``` | ||
|
|
There was a problem hiding this comment.
需提供服务验证方法(如 curl 命令)及预期结果,说明成功特征(如 200 OK),并提供回显示例。
There was a problem hiding this comment.
Updated: added /health and /v1/models verification commands, HTTP 200 and model-id success criteria, and complete expected output examples for A3 and A2.
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
…ect#15793) ### What this PR does / why we need it? Adds the deployment tutorial and supported-model matrix entry for `DeepSeek-V4-Flash-Vision-Exp`, following the model support merged in vllm-project#15457. The tutorial documents the current experimental W8A8 configurations: - one Atlas 800 A3 server (128GB × 8 NPUs), using 16 logical devices with TP4/DP4/EP16; - two Atlas 800 A2 servers (64GB × 8 NPU dies each), with two local DP ranks per node and global TP4/DP4/EP16; - hardware-specific A3/A2 installation tabs, published images, and the public ModelScope W8A8 QuaRot checkpoint; - colocated serving with `FULL_DECODE_ONLY` ACL Graph and a 130K maximum model length; - installation and service verification commands with complete expected output examples; - OpenAI-compatible multimodal functional verification; - the validated OCRBench V1 result and current limitations. Prefill-Decode disaggregation is explicitly marked unsupported. The matrix leaves unverified features experimental or unverified rather than claiming support. Related roadmap: vllm-project#15462 ### Does this PR introduce _any_ user-facing change? No. This is a documentation-only update for an already merged model integration. ### How was this patch tested? - `markdownlint-cli@0.45.0` passes for both changed Markdown files. - `git diff --check` passes. - All relative links and referenced anchors in the new tutorial resolve. - The Hugging Face and ModelScope model URLs return HTTP 200. - Quay API confirms both published image tags are active. - The new supported-model matrix row has the same 20 columns as its header. - The tutorial is registered exactly once in `mkdocs.yml` and has matching English and Chinese entries in `docs/hooks/nav_titles.py`. - `docs/hooks/nav_titles.py` passes Python syntax compilation. - Automated scope checks confirm Chapter 9 is unchanged; the FAQ contains only the public FAQ reference requested in review. - vLLM main: vllm-project/vllm@e6bfe03 --------- Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
What this PR does / why we need it?
Adds the deployment tutorial and supported-model matrix entry for
DeepSeek-V4-Flash-Vision-Exp, following the model support merged in #15457.The tutorial documents the current experimental W8A8 configurations:
FULL_DECODE_ONLYACL Graph and a 130K maximum model length;Prefill-Decode disaggregation is explicitly marked unsupported. The matrix leaves unverified features experimental or unverified rather than claiming support.
Related roadmap: #15462
Does this PR introduce any user-facing change?
No. This is a documentation-only update for an already merged model integration.
How was this patch tested?
markdownlint-cli@0.45.0passes for both changed Markdown files.git diff --checkpasses.All relative links and referenced anchors in the new tutorial resolve.
The Hugging Face and ModelScope model URLs return HTTP 200.
Quay API confirms both published image tags are active.
The new supported-model matrix row has the same 20 columns as its header.
The tutorial is registered exactly once in
mkdocs.ymland has matching English and Chinese entries indocs/hooks/nav_titles.py.docs/hooks/nav_titles.pypasses Python syntax compilation.Automated scope checks confirm Chapter 9 is unchanged; the FAQ contains only the public FAQ reference requested in review.
vLLM main: vllm-project/vllm@e6bfe03