Skip to content

[Test] Drop remote --headless for vLLM 0.28 hybrid DP nightlies - #16133

Merged
wenjun91 merged 2 commits into
vllm-project:mainfrom
LQDLove:fix/glm51-hybrid-lb-drop-headless
Sep 9, 2026
Merged

wenjun91 merged 2 commits into
vllm-project:mainfrom
LQDLove:fix/glm51-hybrid-lb-drop-headless

Conversation

@LQDLove

@LQDLove LQDLove commented Sep 9, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

vLLM 0.28 infers hybrid LB whenever --data-parallel-start-rank is set (data_parallel_start_rank is not None). In that mode the DP handshake requires remote engines not to pass --headless:

Remote engine must not use --headless in external or hybrid dp lb mode

Nightly double-node / multi-node jobs still used the old internal-LB remote command (--data-parallel-start-rank + --headless), so startup was rejected during handshake.

Drop --headless on the remote node so both nodes run hybrid LB (each exposes an API and owns its local ranks). Affected nightly YAMLs:

  • tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_198k_function.yaml
  • tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_128k_90_50.yaml
  • tests/e2e/nightly/multi_node/internal_dp/config/GLM5_1-W8A8-A3-dual-nodes.yaml
  • tests/e2e/nightly/multi_node/internal_dp/config/GLM5_2-W8A8-A3-dual-nodes.yaml
  • tests/e2e/nightly/multi_node/internal_dp/config/Kimi-K2_5-W4A8-A2-dual-nodes.yaml

Does this PR introduce any user-facing change?

No.

How was this patch tested?

  • YAML-only change. Inspected vLLM main EngineArgs.create_engine_config and vllm/v1/engine/utils.py handshake (remote_should_be_headless).

  • Nightly CI should re-run the five jobs above.

  • vLLM main: vllm-project/vllm@b2f6858

vLLM 0.28 infers hybrid LB from --data-parallel-start-rank, and that mode rejects --headless on remote nodes during handshake.

Signed-off-by: LQD <1107297340@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request updates the nightly test configurations for GLM-5.1 models to resolve compatibility issues with vLLM 0.28. By removing the --headless flag in hybrid data-parallel mode, the remote engines can now correctly participate in the handshake process, ensuring stable execution in multi-node environments.

Highlights

  • Configuration Update: Removed the --headless flag from GLM-5.1 hybrid data-parallel configurations to comply with vLLM 0.28 requirements.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request removes the --headless flag from the deployment configurations in two GLM-5.1 multi-node test YAML files. The review feedback correctly identifies that the PR Title and Summary do not comply with the repository's style guide and provides suggestions to align them with the required format.

--data-parallel-size 8
--data-parallel-size-local 4
--data-parallel-start-rank 4
--headless

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The Pull Request Title and Summary do not fully adhere to the repository style guide. Specifically, the PR Title is missing the [Action] prefix (e.g., [BugFix]), and the PR Summary headers should use H3 (###) instead of H2 (##).

Please update the PR Title and Summary to match the following suggestions:

Suggested PR Title:

[Test][BugFix] Drop remote --headless for GLM-5.1 hybrid DP

Suggested PR Summary:

### What this PR does / why we need it?

vLLM 0.28 infers hybrid LB whenever `--data-parallel-start-rank` is set. In that mode, the DP handshake rejects remote engines that still pass `--headless` with the error:
`Remote engine must not use --headless in external or hybrid dp lb mode`

This failed the nightly `double-node` test for `GLM-5.1-W8A8C8-A3_198k_function`. This PR drops `--headless` on the remote node so both nodes run hybrid LB (each exposes an API and owns its local ranks). The same leftover flag is also removed from the sibling 128k nightly YAML.

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- Inspected vLLM main `EngineArgs.create_engine_config` (`data_parallel_start_rank is not None` implies hybrid LB) and `vllm/v1/engine/utils.py` handshake (`remote_should_be_headless`).
- YAML-only change; nightly double-node CI should re-run `GLM-5.1-W8A8C8-A3_198k_function`.
References
  1. The PR Title and PR Summary must follow the specified format, with the PR Title using [Branch][Module][Action] format and the PR Summary using specific H3 headers inside markdown code blocks. (link)

@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

Same vLLM 0.28 handshake reject as GLM-5.1 C8: start-rank implies hybrid LB, so GLM-5.1/5.2 A3 dual-node and Kimi-K2.5 A2 remote nodes cannot pass --headless.

Signed-off-by: LQD <1107297340@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@czydyy

czydyy commented Sep 9, 2026 •

Copy link
Copy Markdown
Collaborator

/nightly GLM5_1-W8A8-A3-dual-nodes Kimi-K2_5-W4A8-A2-dual-nodes
nightly command triggered.

@LQDLove LQDLove changed the title [Test] Drop remote --headless for GLM-5.1 hybrid DP [Test] Drop remote --headless for vLLM 0.28 hybrid DP nightlies Sep 9, 2026
@zhao-stack zhao-stack added the ready-precise run selected e2e test for pr label Sep 9, 2026

@Ronald1995 Ronald1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@wenjun91
wenjun91 merged commit 7cb184e into vllm-project:main Sep 9, 2026
28 checks passed
@LQDLove
LQDLove deleted the fix/glm51-hybrid-lb-drop-headless branch September 10, 2026 00:53
chen-commits pushed a commit to chen-commits/vllm-ascend that referenced this pull request Sep 10, 2026
…-project#16133)

### What this PR does / why we need it?

vLLM 0.28 infers hybrid LB whenever `--data-parallel-start-rank` is set
(`data_parallel_start_rank is not None`). In that mode the DP handshake
requires remote engines **not** to pass `--headless`:

```text
Remote engine must not use --headless in external or hybrid dp lb mode
```

Nightly double-node / multi-node jobs still used the old internal-LB
remote command (`--data-parallel-start-rank` + `--headless`), so startup
was rejected during handshake.

Drop `--headless` on the remote node so both nodes run hybrid LB (each
exposes an API and owns its local ranks). Affected nightly YAMLs:

-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_198k_function.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_128k_90_50.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM5_1-W8A8-A3-dual-nodes.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM5_2-W8A8-A3-dual-nodes.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/Kimi-K2_5-W4A8-A2-dual-nodes.yaml`

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- YAML-only change. Inspected vLLM main
`EngineArgs.create_engine_config` and `vllm/v1/engine/utils.py`
handshake (`remote_should_be_headless`).
- Nightly CI should re-run the five jobs above.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: LQD <1107297340@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
weijinqian0 pushed a commit that referenced this pull request Sep 11, 2026
### What this PR does / why we need it?

vLLM #47692 changed the hybrid DP load-balancing inference to treat an
explicit `--data-parallel-start-rank 0` as meaningful. As a result, the
primary node in these multi-node `internal_dp` configurations is now
inferred as hybrid LB, while the remote node is intended to run headless
for internal LB.

PRs #16133, #16186, and #16204 removed `--headless` from remote nodes to
satisfy the resulting handshake, but that changed the topology covered
by the tests. The regular multi-node internal-DP runner sends benchmark
traffic only to the primary node. In disaggregated-prefill cases, the PD
proxy was explicitly designed to exclude headless nodes and target one
API endpoint per DP group. Switching these deployments to
hybrid/external LB therefore changes the intended test coverage.

Restore the intended internal-LB topology in all 17 affected nightly and
weekly configurations:

- remove `--data-parallel-start-rank 0` from the primary node of each DP
group, allowing rank 0 to be inferred without enabling hybrid LB;
- restore `--headless` on the remote node while retaining its non-zero
DP rank offset.

This keeps one API endpoint on the primary node and lets it schedule
requests across all local and remote DP ranks.

### Does this PR introduce _any_ user-facing change?

No. This only updates nightly test deployment configurations.

### How was this patch tested?

- `git diff --check`
- Parsed all 17 modified YAML files with PyYAML and verified their
internal-DP node roles.
- Multi-node NPU nightly and weekly jobs should validate the complete
deployments.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
…-project#16133)

### What this PR does / why we need it?

vLLM 0.28 infers hybrid LB whenever `--data-parallel-start-rank` is set
(`data_parallel_start_rank is not None`). In that mode the DP handshake
requires remote engines **not** to pass `--headless`:

```text
Remote engine must not use --headless in external or hybrid dp lb mode
```

Nightly double-node / multi-node jobs still used the old internal-LB
remote command (`--data-parallel-start-rank` + `--headless`), so startup
was rejected during handshake.

Drop `--headless` on the remote node so both nodes run hybrid LB (each
exposes an API and owns its local ranks). Affected nightly YAMLs:

-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_198k_function.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_128k_90_50.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM5_1-W8A8-A3-dual-nodes.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM5_2-W8A8-A3-dual-nodes.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/Kimi-K2_5-W4A8-A2-dual-nodes.yaml`

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- YAML-only change. Inspected vLLM main
`EngineArgs.create_engine_config` and `vllm/v1/engine/utils.py`
handshake (`remote_should_be_headless`).
- Nightly CI should re-run the five jobs above.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: LQD <1107297340@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

vLLM #47692 changed the hybrid DP load-balancing inference to treat an
explicit `--data-parallel-start-rank 0` as meaningful. As a result, the
primary node in these multi-node `internal_dp` configurations is now
inferred as hybrid LB, while the remote node is intended to run headless
for internal LB.

PRs vllm-project#16133, vllm-project#16186, and vllm-project#16204 removed `--headless` from remote nodes to
satisfy the resulting handshake, but that changed the topology covered
by the tests. The regular multi-node internal-DP runner sends benchmark
traffic only to the primary node. In disaggregated-prefill cases, the PD
proxy was explicitly designed to exclude headless nodes and target one
API endpoint per DP group. Switching these deployments to
hybrid/external LB therefore changes the intended test coverage.

Restore the intended internal-LB topology in all 17 affected nightly and
weekly configurations:

- remove `--data-parallel-start-rank 0` from the primary node of each DP
group, allowing rank 0 to be inferred without enabling hybrid LB;
- restore `--headless` on the remote node while retaining its non-zero
DP rank offset.

This keeps one API endpoint on the primary node and lets it schedule
requests across all local and remote DP ranks.

### Does this PR introduce _any_ user-facing change?

No. This only updates nightly test deployment configurations.

### How was this patch tested?

- `git diff --check`
- Parsed all 17 modified YAML files with PyYAML and verified their
internal-DP node roles.
- Multi-node NPU nightly and weekly jobs should validate the complete
deployments.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
…-project#16133)

### What this PR does / why we need it?

vLLM 0.28 infers hybrid LB whenever `--data-parallel-start-rank` is set
(`data_parallel_start_rank is not None`). In that mode the DP handshake
requires remote engines **not** to pass `--headless`:

```text
Remote engine must not use --headless in external or hybrid dp lb mode
```

Nightly double-node / multi-node jobs still used the old internal-LB
remote command (`--data-parallel-start-rank` + `--headless`), so startup
was rejected during handshake.

Drop `--headless` on the remote node so both nodes run hybrid LB (each
exposes an API and owns its local ranks). Affected nightly YAMLs:

-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_198k_function.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_128k_90_50.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM5_1-W8A8-A3-dual-nodes.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM5_2-W8A8-A3-dual-nodes.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/Kimi-K2_5-W4A8-A2-dual-nodes.yaml`

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- YAML-only change. Inspected vLLM main
`EngineArgs.create_engine_config` and `vllm/v1/engine/utils.py`
handshake (`remote_should_be_headless`).
- Nightly CI should re-run the five jobs above.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: LQD <1107297340@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: tianming2009 <13246728590@163.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
### What this PR does / why we need it?

vLLM #47692 changed the hybrid DP load-balancing inference to treat an
explicit `--data-parallel-start-rank 0` as meaningful. As a result, the
primary node in these multi-node `internal_dp` configurations is now
inferred as hybrid LB, while the remote node is intended to run headless
for internal LB.

PRs vllm-project#16133, vllm-project#16186, and vllm-project#16204 removed `--headless` from remote nodes to
satisfy the resulting handshake, but that changed the topology covered
by the tests. The regular multi-node internal-DP runner sends benchmark
traffic only to the primary node. In disaggregated-prefill cases, the PD
proxy was explicitly designed to exclude headless nodes and target one
API endpoint per DP group. Switching these deployments to
hybrid/external LB therefore changes the intended test coverage.

Restore the intended internal-LB topology in all 17 affected nightly and
weekly configurations:

- remove `--data-parallel-start-rank 0` from the primary node of each DP
group, allowing rank 0 to be inferred without enabling hybrid LB;
- restore `--headless` on the remote node while retaining its non-zero
DP rank offset.

This keeps one API endpoint on the primary node and lets it schedule
requests across all local and remote DP ranks.

### Does this PR introduce _any_ user-facing change?

No. This only updates nightly test deployment configurations.

### How was this patch tested?

- `git diff --check`
- Parsed all 17 modified YAML files with PyYAML and verified their
internal-DP node roles.
- Multi-node NPU nightly and weekly jobs should validate the complete
deployments.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…-project#16133)

### What this PR does / why we need it?

vLLM 0.28 infers hybrid LB whenever `--data-parallel-start-rank` is set
(`data_parallel_start_rank is not None`). In that mode the DP handshake
requires remote engines **not** to pass `--headless`:

```text
Remote engine must not use --headless in external or hybrid dp lb mode
```

Nightly double-node / multi-node jobs still used the old internal-LB
remote command (`--data-parallel-start-rank` + `--headless`), so startup
was rejected during handshake.

Drop `--headless` on the remote node so both nodes run hybrid LB (each
exposes an API and owns its local ranks). Affected nightly YAMLs:

-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_198k_function.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_128k_90_50.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM5_1-W8A8-A3-dual-nodes.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/GLM5_2-W8A8-A3-dual-nodes.yaml`
-
`tests/e2e/nightly/multi_node/internal_dp/config/Kimi-K2_5-W4A8-A2-dual-nodes.yaml`

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

- YAML-only change. Inspected vLLM main
`EngineArgs.create_engine_config` and `vllm/v1/engine/utils.py`
handshake (`remote_should_be_headless`).
- Nightly CI should re-run the five jobs above.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: LQD <1107297340@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: like-0517 <ithwlike@126.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
### What this PR does / why we need it?

vLLM #47692 changed the hybrid DP load-balancing inference to treat an
explicit `--data-parallel-start-rank 0` as meaningful. As a result, the
primary node in these multi-node `internal_dp` configurations is now
inferred as hybrid LB, while the remote node is intended to run headless
for internal LB.

PRs vllm-project#16133, vllm-project#16186, and vllm-project#16204 removed `--headless` from remote nodes to
satisfy the resulting handshake, but that changed the topology covered
by the tests. The regular multi-node internal-DP runner sends benchmark
traffic only to the primary node. In disaggregated-prefill cases, the PD
proxy was explicitly designed to exclude headless nodes and target one
API endpoint per DP group. Switching these deployments to
hybrid/external LB therefore changes the intended test coverage.

Restore the intended internal-LB topology in all 17 affected nightly and
weekly configurations:

- remove `--data-parallel-start-rank 0` from the primary node of each DP
group, allowing rank 0 to be inferred without enabling hybrid LB;
- restore `--headless` on the remote node while retaining its non-zero
DP rank offset.

This keeps one API endpoint on the primary node and lets it schedule
requests across all local and remote DP ranks.

### Does this PR introduce _any_ user-facing change?

No. This only updates nightly test deployment configurations.

### How was this patch tested?

- `git diff --check`
- Parsed all 17 modified YAML files with PyYAML and verified their
internal-DP node roles.
- Multi-node NPU nightly and weekly jobs should validate the complete
deployments.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants