Skip to content

[CI] Drop --headless for DeepSeek-V3.2-EP secondary nodes - #16204

Merged
wenjun91 merged 4 commits into
vllm-project:mainfrom
LQDLove:fix/nightly-drop-headless-local1
Sep 10, 2026
Merged

wenjun91 merged 4 commits into
vllm-project:mainfrom
LQDLove:fix/nightly-drop-headless-local1

Conversation

@LQDLove

@LQDLove LQDLove commented Sep 10, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Prefiller/decoder use --data-parallel-start-rank with multi-node DP. On vLLM 0.28, a set start-rank infers hybrid LB; with local=1 that autoswitches to external LB, which forbids --headless on remotes.
  • Keep --data-parallel-start-rank; remove --headless on both prefiller and decoder secondary nodes so remotes expose API servers (PD proxy can target them).

Related failure

Test plan

vLLM autoswitches to data_parallel_external_lb when data_parallel_size_local=1, which forbids --headless on remote engines. Remove --headless from affected nightly configs (DeepSeek-V3.2-EP prefiller, GLM5.1 A2 dual-nodes, Qwen3-235B A2).

Signed-off-by: liaoqidan <1107297340@qq.com>

Co-authored-by: Cursor <cursoragent@cursor.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request updates the nightly CI configuration files to ensure compatibility with vLLM's data parallel external load balancing mode. By removing the --headless flag from specific multi-node configurations where local data parallel size is 1, the changes prevent runtime errors and ensure that secondary nodes can correctly join the proxy target list.

Highlights

  • Configuration Update: Removed the --headless flag from nightly CI configurations for DeepSeek-V3.2-W8A8-EP, GLM5.1-A2, and Qwen3-235B-A22B.
  • Compatibility Fix: Addressed conflicts where --headless is forbidden when data_parallel_size_local is set to 1, which triggers external load balancing in vLLM.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Test][Misc] Remove headless flag from multi-node E2E test configurations

Suggested PR Summary:

### What this PR does / why we need it?
This PR removes the obsolete `--headless` flag from the `vllm serve` deployment commands across several nightly multi-node end-to-end test configuration files, including DeepSeek-V3.2, GLM5, and Qwen3.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Tested via existing nightly E2E CI pipelines.

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

vLLM 0.28 treats data_parallel_start_rank=0 as set (is not None), which infers hybrid LB; with local=1 that autoswitches to external LB and forbids --headless. Prefer classic internal DP: omit start-rank 0 on the primary, keep --headless on the secondary. Restore headless on GLM A2 / Qwen A2 (they never set start-rank 0).

Signed-off-by: liaoqidan <1107297340@qq.com>

Co-authored-by: Cursor <cursoragent@cursor.com>
@LQDLove LQDLove changed the title [CI] Drop --headless for local=1 nightly DP nodes [CI] Avoid local=1 hybrid autoswitch via start-rank Sep 10, 2026
@zhao-stack zhao-stack added the ready-precise run selected e2e test for pr label Sep 10, 2026
LQDLove and others added 2 commits September 10, 2026 20:05
Keep --data-parallel-start-rank 0. Under v0.28 that infers hybrid and autoswitches to external LB when local=1, which forbids --headless on the remote prefiller. Remove only the prefiller secondary --headless so both prefiller ranks expose API servers; decoder headless (local=4) unchanged.

Signed-off-by: liaoqidan <1107297340@qq.com>

Co-authored-by: Cursor <cursoragent@cursor.com>
Mirror prefiller: remove secondary decoder --headless so both decode nodes expose API under start-rank hybrid/external LB.

Signed-off-by: liaoqidan <1107297340@qq.com>

Co-authored-by: Cursor <cursoragent@cursor.com>
@LQDLove LQDLove changed the title [CI] Avoid local=1 hybrid autoswitch via start-rank [CI] Drop --headless for DeepSeek-V3.2-EP secondary nodes Sep 10, 2026
@wenjun91
wenjun91 merged commit 5c4ad3f into vllm-project:main Sep 10, 2026
14 checks passed
weijinqian0 pushed a commit that referenced this pull request Sep 11, 2026
### What this PR does / why we need it?

vLLM #47692 changed the hybrid DP load-balancing inference to treat an
explicit `--data-parallel-start-rank 0` as meaningful. As a result, the
primary node in these multi-node `internal_dp` configurations is now
inferred as hybrid LB, while the remote node is intended to run headless
for internal LB.

PRs #16133, #16186, and #16204 removed `--headless` from remote nodes to
satisfy the resulting handshake, but that changed the topology covered
by the tests. The regular multi-node internal-DP runner sends benchmark
traffic only to the primary node. In disaggregated-prefill cases, the PD
proxy was explicitly designed to exclude headless nodes and target one
API endpoint per DP group. Switching these deployments to
hybrid/external LB therefore changes the intended test coverage.

Restore the intended internal-LB topology in all 17 affected nightly and
weekly configurations:

- remove `--data-parallel-start-rank 0` from the primary node of each DP
group, allowing rank 0 to be inferred without enabling hybrid LB;
- restore `--headless` on the remote node while retaining its non-zero
DP rank offset.

This keeps one API endpoint on the primary node and lets it schedule
requests across all local and remote DP ranks.

### Does this PR introduce _any_ user-facing change?

No. This only updates nightly test deployment configurations.

### How was this patch tested?

- `git diff --check`
- Parsed all 17 modified YAML files with PyYAML and verified their
internal-DP node roles.
- Multi-node NPU nightly and weekly jobs should validate the complete
deployments.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
…ct#16204)

## Summary
- Prefiller/decoder use `--data-parallel-start-rank` with multi-node DP.
On vLLM 0.28, a set start-rank infers hybrid LB; with `local=1` that
autoswitches to external LB, which forbids `--headless` on remotes.
- **Keep** `--data-parallel-start-rank`; **remove** `--headless` on both
prefiller and decoder secondary nodes so remotes expose API servers (PD
proxy can target them).

## Related failure
-
https://github.com/vllm-project/vllm-ascend/actions/runs/34372289942/job/102544223304

## Test plan
- [ ] DeepSeek-V3.2-W8A8-EP starts without `Remote engine must not use
--headless`

- vLLM main:
vllm-project/vllm@b2f6858

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

vLLM #47692 changed the hybrid DP load-balancing inference to treat an
explicit `--data-parallel-start-rank 0` as meaningful. As a result, the
primary node in these multi-node `internal_dp` configurations is now
inferred as hybrid LB, while the remote node is intended to run headless
for internal LB.

PRs vllm-project#16133, vllm-project#16186, and vllm-project#16204 removed `--headless` from remote nodes to
satisfy the resulting handshake, but that changed the topology covered
by the tests. The regular multi-node internal-DP runner sends benchmark
traffic only to the primary node. In disaggregated-prefill cases, the PD
proxy was explicitly designed to exclude headless nodes and target one
API endpoint per DP group. Switching these deployments to
hybrid/external LB therefore changes the intended test coverage.

Restore the intended internal-LB topology in all 17 affected nightly and
weekly configurations:

- remove `--data-parallel-start-rank 0` from the primary node of each DP
group, allowing rank 0 to be inferred without enabling hybrid LB;
- restore `--headless` on the remote node while retaining its non-zero
DP rank offset.

This keeps one API endpoint on the primary node and lets it schedule
requests across all local and remote DP ranks.

### Does this PR introduce _any_ user-facing change?

No. This only updates nightly test deployment configurations.

### How was this patch tested?

- `git diff --check`
- Parsed all 17 modified YAML files with PyYAML and verified their
internal-DP node roles.
- Multi-node NPU nightly and weekly jobs should validate the complete
deployments.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
…ct#16204)

## Summary
- Prefiller/decoder use `--data-parallel-start-rank` with multi-node DP.
On vLLM 0.28, a set start-rank infers hybrid LB; with `local=1` that
autoswitches to external LB, which forbids `--headless` on remotes.
- **Keep** `--data-parallel-start-rank`; **remove** `--headless` on both
prefiller and decoder secondary nodes so remotes expose API servers (PD
proxy can target them).

## Related failure
-
https://github.com/vllm-project/vllm-ascend/actions/runs/34372289942/job/102544223304

## Test plan
- [ ] DeepSeek-V3.2-W8A8-EP starts without `Remote engine must not use
--headless`

- vLLM main:
vllm-project/vllm@b2f6858

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: tianming2009 <13246728590@163.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
### What this PR does / why we need it?

vLLM #47692 changed the hybrid DP load-balancing inference to treat an
explicit `--data-parallel-start-rank 0` as meaningful. As a result, the
primary node in these multi-node `internal_dp` configurations is now
inferred as hybrid LB, while the remote node is intended to run headless
for internal LB.

PRs vllm-project#16133, vllm-project#16186, and vllm-project#16204 removed `--headless` from remote nodes to
satisfy the resulting handshake, but that changed the topology covered
by the tests. The regular multi-node internal-DP runner sends benchmark
traffic only to the primary node. In disaggregated-prefill cases, the PD
proxy was explicitly designed to exclude headless nodes and target one
API endpoint per DP group. Switching these deployments to
hybrid/external LB therefore changes the intended test coverage.

Restore the intended internal-LB topology in all 17 affected nightly and
weekly configurations:

- remove `--data-parallel-start-rank 0` from the primary node of each DP
group, allowing rank 0 to be inferred without enabling hybrid LB;
- restore `--headless` on the remote node while retaining its non-zero
DP rank offset.

This keeps one API endpoint on the primary node and lets it schedule
requests across all local and remote DP ranks.

### Does this PR introduce _any_ user-facing change?

No. This only updates nightly test deployment configurations.

### How was this patch tested?

- `git diff --check`
- Parsed all 17 modified YAML files with PyYAML and verified their
internal-DP node roles.
- Multi-node NPU nightly and weekly jobs should validate the complete
deployments.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…ct#16204)

## Summary
- Prefiller/decoder use `--data-parallel-start-rank` with multi-node DP.
On vLLM 0.28, a set start-rank infers hybrid LB; with `local=1` that
autoswitches to external LB, which forbids `--headless` on remotes.
- **Keep** `--data-parallel-start-rank`; **remove** `--headless` on both
prefiller and decoder secondary nodes so remotes expose API servers (PD
proxy can target them).

## Related failure
-
https://github.com/vllm-project/vllm-ascend/actions/runs/34372289942/job/102544223304

## Test plan
- [ ] DeepSeek-V3.2-W8A8-EP starts without `Remote engine must not use
--headless`

- vLLM main:
vllm-project/vllm@b2f6858

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: like-0517 <ithwlike@126.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
### What this PR does / why we need it?

vLLM #47692 changed the hybrid DP load-balancing inference to treat an
explicit `--data-parallel-start-rank 0` as meaningful. As a result, the
primary node in these multi-node `internal_dp` configurations is now
inferred as hybrid LB, while the remote node is intended to run headless
for internal LB.

PRs vllm-project#16133, vllm-project#16186, and vllm-project#16204 removed `--headless` from remote nodes to
satisfy the resulting handshake, but that changed the topology covered
by the tests. The regular multi-node internal-DP runner sends benchmark
traffic only to the primary node. In disaggregated-prefill cases, the PD
proxy was explicitly designed to exclude headless nodes and target one
API endpoint per DP group. Switching these deployments to
hybrid/external LB therefore changes the intended test coverage.

Restore the intended internal-LB topology in all 17 affected nightly and
weekly configurations:

- remove `--data-parallel-start-rank 0` from the primary node of each DP
group, allowing rank 0 to be inferred without enabling hybrid LB;
- restore `--headless` on the remote node while retaining its non-zero
DP rank offset.

This keeps one API endpoint on the primary node and lets it schedule
requests across all local and remote DP ranks.

### Does this PR introduce _any_ user-facing change?

No. This only updates nightly test deployment configurations.

### How was this patch tested?

- `git diff --check`
- Parsed all 17 modified YAML files with PyYAML and verified their
internal-DP node roles.
- Multi-node NPU nightly and weekly jobs should validate the complete
deployments.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants