Skip to content

[v0.25.1rc][BugFix][Frontend] Align DeepSeek V4 system tool rendering - #14035

Merged
wangxiyuan merged 1 commit into
vllm-project:releases/v0.25.1rcfrom
QwertyJack:fix/deepseek-v4-system-tools-renderer-v025
Aug 14, 2026
Merged

wangxiyuan merged 1 commit into
vllm-project:releases/v0.25.1rcfrom
QwertyJack:fix/deepseek-v4-system-tools-renderer-v025

Conversation

@QwertyJack

@QwertyJack QwertyJack commented Aug 11, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

This is a follow-up to #13519. The v0.25.1 release monkey patch inherited an
upstream Python renderer mismatch reported in
vllm-project/vllm#51829.

For DeepSeek V4 requests containing both an existing system message and
top-level tools, the Python wrapper always inserted a synthetic system message
and rendered the tools before the caller's system content. This differs from
both vLLM's Rust renderer and the DeepSeek-V4-Flash-0731 checkpoint reference.

This patch aligns the release Python path with those references:

  • attach request-level tools to a shallow copy of the first existing system
    message;
  • insert a synthetic system message only when no system message exists; and
  • leave caller-owned message dictionaries unchanged.

The reasoning-effort normalization introduced by #13519 is unchanged.

Does this PR introduce any user-facing change?

Yes. DeepSeek V4 requests with both a system message and top-level tools now
render the system content and tool schemas in the same order as the Rust and
checkpoint reference renderers. Requests without a system message keep the
existing synthetic-system behavior.

How was this patch tested?

  • Tested with vLLM v0.25.1 at commit
    752a3a504485790a2e8491cacbb35c137339ad34.
  • VLLM_VERSION=0.25.1 python -m pytest -q tests/ut/patch/platform/test_deepseek_v4_thinking.py
    • Result: 26 passed.
    • Coverage includes existing-system tools, missing-system tools, caller input
      immutability, and default/high/low effort behavior.
  • Tokenizer integration comparison against the checkpoint encoder for
    system/no-system x low/high:
    • all four rendered prompts matched exactly;
    • all four input-ID sequences matched exactly; and
    • caller messages remained unchanged in all four cases.
  • ruff check and ruff format --check on both changed files passed.
  • git diff --check origin/releases/v0.25.1rc...HEAD passed.

vllm-project/vllm@752a3a5

Attach request-level tools to the first existing system message and only synthesize a system message when one is absent. Preserve caller-owned message dictionaries.

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
@QwertyJack
QwertyJack requested a review from wangxiyuan as a code owner August 11, 2026 14:33
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a rendering mismatch in the DeepSeek V4 Python wrapper for vLLM v0.25.1. By aligning the tool attachment logic with the Rust renderer and official checkpoint references, the patch ensures consistent prompt formatting when system messages and tools are both provided. The changes improve reliability for tool-calling workflows while maintaining the integrity of user-provided message structures.

Highlights

  • Tool Rendering Alignment: Updated the DeepSeek V4 Python renderer to correctly attach tools to existing system messages instead of always inserting a new synthetic system message.
  • Synthetic System Message Logic: Ensured that a synthetic system message is only created when no existing system message is present in the request.
  • Immutability Preservation: Implemented shallow copying of system messages before modification to ensure caller-owned message dictionaries remain unchanged.
  • Test Coverage: Added comprehensive unit tests to verify tool attachment behavior, system message handling, and reasoning effort parameter preservation.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][BugFix] Fix DeepSeek V4 tokenizer tool attachment to existing system message

Suggested PR Summary:

### What this PR does / why we need it?
This PR fixes an issue in the DeepSeek V4 tokenizer patch where tools were always inserted as a new system message at the beginning of the message list, even if a system message already existed. The updated logic now searches for an existing system message and attaches the tools to it if found; otherwise, it inserts a new system message at the beginning.

### Does this PR introduce _any_ user-facing change?
Yes, it ensures that tools are correctly attached to the existing system message instead of creating a duplicate system message, which improves compatibility with chat templates that expect a single system message.

### How was this patch tested?
Added unit tests in `tests/ut/patch/platform/test_deepseek_v4_thinking.py` to verify both cases (when a system message is already present and when it is missing).

I have no further feedback to provide as the implementation is correct and well-tested.

@wangxiyuan
wangxiyuan merged commit 99d2ade into vllm-project:releases/v0.25.1rc Aug 14, 2026
10 checks passed
jiaqi-lee pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 15, 2026
…vllm-project#14035)

### What this PR does / why we need it?

This is a follow-up to vllm-project#13519. The v0.25.1 release monkey patch
inherited an
upstream Python renderer mismatch reported in
vllm-project/vllm#51829.

For DeepSeek V4 requests containing both an existing system message and
top-level tools, the Python wrapper always inserted a synthetic system
message
and rendered the tools before the caller's system content. This differs
from
both vLLM's Rust renderer and the DeepSeek-V4-Flash-0731 checkpoint
reference.

This patch aligns the release Python path with those references:

- attach request-level tools to a shallow copy of the first existing
system
  message;
- insert a synthetic system message only when no system message exists;
and
- leave caller-owned message dictionaries unchanged.

The reasoning-effort normalization introduced by vllm-project#13519 is unchanged.

### Does this PR introduce _any_ user-facing change?

Yes. DeepSeek V4 requests with both a system message and top-level tools
now
render the system content and tool schemas in the same order as the Rust
and
checkpoint reference renderers. Requests without a system message keep
the
existing synthetic-system behavior.

### How was this patch tested?

- Tested with vLLM `v0.25.1` at commit
  `752a3a504485790a2e8491cacbb35c137339ad34`.
- `VLLM_VERSION=0.25.1 python -m pytest -q
tests/ut/patch/platform/test_deepseek_v4_thinking.py`
  - Result: `26 passed`.
- Coverage includes existing-system tools, missing-system tools, caller
input
    immutability, and default/high/low effort behavior.
- Tokenizer integration comparison against the checkpoint encoder for
  `system/no-system x low/high`:
  - all four rendered prompts matched exactly;
  - all four input-ID sequences matched exactly; and
  - caller messages remained unchanged in all four cases.
- `ruff check` and `ruff format --check` on both changed files passed.
- `git diff --check origin/releases/v0.25.1rc...HEAD` passed.


vllm-project/vllm@752a3a5

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@fe784ff

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
jiaqi-lee pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 15, 2026
…vllm-project#14035)

### What this PR does / why we need it?

This is a follow-up to vllm-project#13519. The v0.25.1 release monkey patch
inherited an
upstream Python renderer mismatch reported in
vllm-project/vllm#51829.

For DeepSeek V4 requests containing both an existing system message and
top-level tools, the Python wrapper always inserted a synthetic system
message
and rendered the tools before the caller's system content. This differs
from
both vLLM's Rust renderer and the DeepSeek-V4-Flash-0731 checkpoint
reference.

This patch aligns the release Python path with those references:

- attach request-level tools to a shallow copy of the first existing
system
  message;
- insert a synthetic system message only when no system message exists;
and
- leave caller-owned message dictionaries unchanged.

The reasoning-effort normalization introduced by vllm-project#13519 is unchanged.

### Does this PR introduce _any_ user-facing change?

Yes. DeepSeek V4 requests with both a system message and top-level tools
now
render the system content and tool schemas in the same order as the Rust
and
checkpoint reference renderers. Requests without a system message keep
the
existing synthetic-system behavior.

### How was this patch tested?

- Tested with vLLM `v0.25.1` at commit
  `752a3a504485790a2e8491cacbb35c137339ad34`.
- `VLLM_VERSION=0.25.1 python -m pytest -q
tests/ut/patch/platform/test_deepseek_v4_thinking.py`
  - Result: `26 passed`.
- Coverage includes existing-system tools, missing-system tools, caller
input
    immutability, and default/high/low effort behavior.
- Tokenizer integration comparison against the checkpoint encoder for
  `system/no-system x low/high`:
  - all four rendered prompts matched exactly;
  - all four input-ID sequences matched exactly; and
  - caller messages remained unchanged in all four cases.
- `ruff check` and `ruff format --check` on both changed files passed.
- `git diff --check origin/releases/v0.25.1rc...HEAD` passed.

vllm-project/vllm@752a3a5

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@fe784ff

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Signed-off-by: lijiaqi139 <lijiaqi139@huawei.com>
jiaqi-lee pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 17, 2026
…vllm-project#14035)

### What this PR does / why we need it?

This is a follow-up to vllm-project#13519. The v0.25.1 release monkey patch
inherited an
upstream Python renderer mismatch reported in
vllm-project/vllm#51829.

For DeepSeek V4 requests containing both an existing system message and
top-level tools, the Python wrapper always inserted a synthetic system
message
and rendered the tools before the caller's system content. This differs
from
both vLLM's Rust renderer and the DeepSeek-V4-Flash-0731 checkpoint
reference.

This patch aligns the release Python path with those references:

- attach request-level tools to a shallow copy of the first existing
system
  message;
- insert a synthetic system message only when no system message exists;
and
- leave caller-owned message dictionaries unchanged.

The reasoning-effort normalization introduced by vllm-project#13519 is unchanged.

### Does this PR introduce _any_ user-facing change?

Yes. DeepSeek V4 requests with both a system message and top-level tools
now
render the system content and tool schemas in the same order as the Rust
and
checkpoint reference renderers. Requests without a system message keep
the
existing synthetic-system behavior.

### How was this patch tested?

- Tested with vLLM `v0.25.1` at commit
  `752a3a504485790a2e8491cacbb35c137339ad34`.
- `VLLM_VERSION=0.25.1 python -m pytest -q
tests/ut/patch/platform/test_deepseek_v4_thinking.py`
  - Result: `26 passed`.
- Coverage includes existing-system tools, missing-system tools, caller
input
    immutability, and default/high/low effort behavior.
- Tokenizer integration comparison against the checkpoint encoder for
  `system/no-system x low/high`:
  - all four rendered prompts matched exactly;
  - all four input-ID sequences matched exactly; and
  - caller messages remained unchanged in all four cases.
- `ruff check` and `ruff format --check` on both changed files passed.
- `git diff --check origin/releases/v0.25.1rc...HEAD` passed.

vllm-project/vllm@752a3a5

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@fe784ff

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Signed-off-by: lijiaqi139 <lijiaqi139@huawei.com>
jiaqi-lee pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 17, 2026
…vllm-project#14035)

### What this PR does / why we need it?

This is a follow-up to vllm-project#13519. The v0.25.1 release monkey patch
inherited an
upstream Python renderer mismatch reported in
vllm-project/vllm#51829.

For DeepSeek V4 requests containing both an existing system message and
top-level tools, the Python wrapper always inserted a synthetic system
message
and rendered the tools before the caller's system content. This differs
from
both vLLM's Rust renderer and the DeepSeek-V4-Flash-0731 checkpoint
reference.

This patch aligns the release Python path with those references:

- attach request-level tools to a shallow copy of the first existing
system
  message;
- insert a synthetic system message only when no system message exists;
and
- leave caller-owned message dictionaries unchanged.

The reasoning-effort normalization introduced by vllm-project#13519 is unchanged.

### Does this PR introduce _any_ user-facing change?

Yes. DeepSeek V4 requests with both a system message and top-level tools
now
render the system content and tool schemas in the same order as the Rust
and
checkpoint reference renderers. Requests without a system message keep
the
existing synthetic-system behavior.

### How was this patch tested?

- Tested with vLLM `v0.25.1` at commit
  `752a3a504485790a2e8491cacbb35c137339ad34`.
- `VLLM_VERSION=0.25.1 python -m pytest -q
tests/ut/patch/platform/test_deepseek_v4_thinking.py`
  - Result: `26 passed`.
- Coverage includes existing-system tools, missing-system tools, caller
input
    immutability, and default/high/low effort behavior.
- Tokenizer integration comparison against the checkpoint encoder for
  `system/no-system x low/high`:
  - all four rendered prompts matched exactly;
  - all four input-ID sequences matched exactly; and
  - caller messages remained unchanged in all four cases.
- `ruff check` and `ruff format --check` on both changed files passed.
- `git diff --check origin/releases/v0.25.1rc...HEAD` passed.

vllm-project/vllm@752a3a5

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@fe784ff

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Signed-off-by: lijiaqi139 <lijiaqi139@huawei.com>
jiaqi-lee pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 17, 2026
…vllm-project#14035)

### What this PR does / why we need it?

This is a follow-up to vllm-project#13519. The v0.25.1 release monkey patch
inherited an
upstream Python renderer mismatch reported in
vllm-project/vllm#51829.

For DeepSeek V4 requests containing both an existing system message and
top-level tools, the Python wrapper always inserted a synthetic system
message
and rendered the tools before the caller's system content. This differs
from
both vLLM's Rust renderer and the DeepSeek-V4-Flash-0731 checkpoint
reference.

This patch aligns the release Python path with those references:

- attach request-level tools to a shallow copy of the first existing
system
  message;
- insert a synthetic system message only when no system message exists;
and
- leave caller-owned message dictionaries unchanged.

The reasoning-effort normalization introduced by vllm-project#13519 is unchanged.

### Does this PR introduce _any_ user-facing change?

Yes. DeepSeek V4 requests with both a system message and top-level tools
now
render the system content and tool schemas in the same order as the Rust
and
checkpoint reference renderers. Requests without a system message keep
the
existing synthetic-system behavior.

### How was this patch tested?

- Tested with vLLM `v0.25.1` at commit
  `752a3a504485790a2e8491cacbb35c137339ad34`.
- `VLLM_VERSION=0.25.1 python -m pytest -q
tests/ut/patch/platform/test_deepseek_v4_thinking.py`
  - Result: `26 passed`.
- Coverage includes existing-system tools, missing-system tools, caller
input
    immutability, and default/high/low effort behavior.
- Tokenizer integration comparison against the checkpoint encoder for
  `system/no-system x low/high`:
  - all four rendered prompts matched exactly;
  - all four input-ID sequences matched exactly; and
  - caller messages remained unchanged in all four cases.
- `ruff check` and `ruff format --check` on both changed files passed.
- `git diff --check origin/releases/v0.25.1rc...HEAD` passed.

vllm-project/vllm@752a3a5

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@fe784ff

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Signed-off-by: lijiaqi139 <lijiaqi139@huawei.com>
Signed-off-by: jiaqi-lee <15316070896@163.com>
jiaqi-lee pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 19, 2026
…vllm-project#14035)

### What this PR does / why we need it?

This is a follow-up to vllm-project#13519. The v0.25.1 release monkey patch
inherited an
upstream Python renderer mismatch reported in
vllm-project/vllm#51829.

For DeepSeek V4 requests containing both an existing system message and
top-level tools, the Python wrapper always inserted a synthetic system
message
and rendered the tools before the caller's system content. This differs
from
both vLLM's Rust renderer and the DeepSeek-V4-Flash-0731 checkpoint
reference.

This patch aligns the release Python path with those references:

- attach request-level tools to a shallow copy of the first existing
system
  message;
- insert a synthetic system message only when no system message exists;
and
- leave caller-owned message dictionaries unchanged.

The reasoning-effort normalization introduced by vllm-project#13519 is unchanged.

### Does this PR introduce _any_ user-facing change?

Yes. DeepSeek V4 requests with both a system message and top-level tools
now
render the system content and tool schemas in the same order as the Rust
and
checkpoint reference renderers. Requests without a system message keep
the
existing synthetic-system behavior.

### How was this patch tested?

- Tested with vLLM `v0.25.1` at commit
  `752a3a504485790a2e8491cacbb35c137339ad34`.
- `VLLM_VERSION=0.25.1 python -m pytest -q
tests/ut/patch/platform/test_deepseek_v4_thinking.py`
  - Result: `26 passed`.
- Coverage includes existing-system tools, missing-system tools, caller
input
    immutability, and default/high/low effort behavior.
- Tokenizer integration comparison against the checkpoint encoder for
  `system/no-system x low/high`:
  - all four rendered prompts matched exactly;
  - all four input-ID sequences matched exactly; and
  - caller messages remained unchanged in all four cases.
- `ruff check` and `ruff format --check` on both changed files passed.
- `git diff --check origin/releases/v0.25.1rc...HEAD` passed.

vllm-project/vllm@752a3a5

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@fe784ff

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Signed-off-by: lijiaqi139 <lijiaqi139@huawei.com>
Signed-off-by: jiaqi-lee <15316070896@163.com>
jiaqi-lee pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 20, 2026
…vllm-project#14035)

### What this PR does / why we need it?

This is a follow-up to vllm-project#13519. The v0.25.1 release monkey patch
inherited an
upstream Python renderer mismatch reported in
vllm-project/vllm#51829.

For DeepSeek V4 requests containing both an existing system message and
top-level tools, the Python wrapper always inserted a synthetic system
message
and rendered the tools before the caller's system content. This differs
from
both vLLM's Rust renderer and the DeepSeek-V4-Flash-0731 checkpoint
reference.

This patch aligns the release Python path with those references:

- attach request-level tools to a shallow copy of the first existing
system
  message;
- insert a synthetic system message only when no system message exists;
and
- leave caller-owned message dictionaries unchanged.

The reasoning-effort normalization introduced by vllm-project#13519 is unchanged.

### Does this PR introduce _any_ user-facing change?

Yes. DeepSeek V4 requests with both a system message and top-level tools
now
render the system content and tool schemas in the same order as the Rust
and
checkpoint reference renderers. Requests without a system message keep
the
existing synthetic-system behavior.

### How was this patch tested?

- Tested with vLLM `v0.25.1` at commit
  `752a3a504485790a2e8491cacbb35c137339ad34`.
- `VLLM_VERSION=0.25.1 python -m pytest -q
tests/ut/patch/platform/test_deepseek_v4_thinking.py`
  - Result: `26 passed`.
- Coverage includes existing-system tools, missing-system tools, caller
input
    immutability, and default/high/low effort behavior.
- Tokenizer integration comparison against the checkpoint encoder for
  `system/no-system x low/high`:
  - all four rendered prompts matched exactly;
  - all four input-ID sequences matched exactly; and
  - caller messages remained unchanged in all four cases.
- `ruff check` and `ruff format --check` on both changed files passed.
- `git diff --check origin/releases/v0.25.1rc...HEAD` passed.

vllm-project/vllm@752a3a5

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@fe784ff

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Signed-off-by: lijiaqi139 <lijiaqi139@huawei.com>
Signed-off-by: jiaqi-lee <15316070896@163.com>
kunpengW-code pushed a commit that referenced this pull request Aug 24, 2026
…4624)

### What this PR does / why we need it?

Ports the remaining DeepSeek V4 frontend fixes from the v0.25 release
branch to `releases/v0.26.0rc`:

- #14035: attach tools to an existing system message instead of
inserting a second system message;
- #14521: stream long schema-typed string tool arguments incrementally
before the closing parameter tag.

The v0.26 reasoning-effort and thinking-default changes are already
covered by #13993 and #14074, so this PR does not duplicate them.

For long string arguments, the patch remains release-scoped until the
supported vLLM contains vllm-project/vllm#52865.

### Does this PR introduce _any_ user-facing change?

Yes. DeepSeek V4 requests preserve their original system prompt when
tools are present, and long string tool arguments are emitted
incrementally instead of being buffered until the parameter closes.

### How was this patch tested?

Tested after rebasing onto the latest `releases/v0.26.0rc` with its
paired vLLM source:

```bash
PYTHONPATH=$VLLM_SOURCE pytest -q \
  tests/ut/patch/platform/test_deepseek_v4_thinking.py \
  tests/ut/patch/platform/test_deepseek_v4_tool_streaming.py
# 45 passed

ruff check   tests/ut/patch/platform/test_deepseek_v4_thinking.py   tests/ut/patch/platform/test_deepseek_v4_tool_streaming.py   vllm_ascend/patch/__init__.py   vllm_ascend/patch/platform/__init__.py   vllm_ascend/patch/platform/patch_deepseek_v4_thinking.py   vllm_ascend/patch/platform/patch_deepseek_v4_tool_streaming.py
# All checks passed!

ruff format --check   tests/ut/patch/platform/test_deepseek_v4_thinking.py   tests/ut/patch/platform/test_deepseek_v4_tool_streaming.py   vllm_ascend/patch/__init__.py   vllm_ascend/patch/platform/__init__.py   vllm_ascend/patch/platform/patch_deepseek_v4_thinking.py   vllm_ascend/patch/platform/patch_deepseek_v4_tool_streaming.py
# 6 files already formatted

git diff --check upstream/releases/v0.26.0rc...HEAD
```

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
alex7092 pushed a commit to alex7092/vllm-ascend that referenced this pull request Aug 27, 2026
- Fix version number: v0.23.0rc1 -> v0.25.1rc1
- Add missing space after comma
- Fix PR link: vllm-project#14035 now points to correct PR

Signed-off-by: alex7092 <15105576+alex7092@users.noreply.github.com>
Signed-off-by: 刘星 <liuxing35@huawei.com>
Leetrytry pushed a commit to Leetrytry/vllm-ascend that referenced this pull request Sep 11, 2026
…lm-project#14624)

### What this PR does / why we need it?

Ports the remaining DeepSeek V4 frontend fixes from the v0.25 release
branch to `releases/v0.26.0rc`:

- vllm-project#14035: attach tools to an existing system message instead of
inserting a second system message;
- vllm-project#14521: stream long schema-typed string tool arguments incrementally
before the closing parameter tag.

The v0.26 reasoning-effort and thinking-default changes are already
covered by vllm-project#13993 and vllm-project#14074, so this PR does not duplicate them.

For long string arguments, the patch remains release-scoped until the
supported vLLM contains vllm-project/vllm#52865.

### Does this PR introduce _any_ user-facing change?

Yes. DeepSeek V4 requests preserve their original system prompt when
tools are present, and long string tool arguments are emitted
incrementally instead of being buffered until the parameter closes.

### How was this patch tested?

Tested after rebasing onto the latest `releases/v0.26.0rc` with its
paired vLLM source:

```bash
PYTHONPATH=$VLLM_SOURCE pytest -q \
  tests/ut/patch/platform/test_deepseek_v4_thinking.py \
  tests/ut/patch/platform/test_deepseek_v4_tool_streaming.py
# 45 passed

ruff check   tests/ut/patch/platform/test_deepseek_v4_thinking.py   tests/ut/patch/platform/test_deepseek_v4_tool_streaming.py   vllm_ascend/patch/__init__.py   vllm_ascend/patch/platform/__init__.py   vllm_ascend/patch/platform/patch_deepseek_v4_thinking.py   vllm_ascend/patch/platform/patch_deepseek_v4_tool_streaming.py
# All checks passed!

ruff format --check   tests/ut/patch/platform/test_deepseek_v4_thinking.py   tests/ut/patch/platform/test_deepseek_v4_tool_streaming.py   vllm_ascend/patch/__init__.py   vllm_ascend/patch/platform/__init__.py   vllm_ascend/patch/platform/patch_deepseek_v4_thinking.py   vllm_ascend/patch/platform/patch_deepseek_v4_tool_streaming.py
# 6 files already formatted

git diff --check upstream/releases/v0.26.0rc...HEAD
```

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants