Repository navigation
[Test] Drop remote --headless for vLLM 0.28 hybrid DP nightlies - #16133
Conversation
vLLM 0.28 infers hybrid LB from --data-parallel-start-rank, and that mode rejects --headless on remote nodes during handshake. Signed-off-by: LQD <1107297340@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request updates the nightly test configurations for GLM-5.1 models to resolve compatibility issues with vLLM 0.28. By removing the --headless flag in hybrid data-parallel mode, the remote engines can now correctly participate in the handshake process, ensuring stable execution in multi-node environments. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request removes the --headless flag from the deployment configurations in two GLM-5.1 multi-node test YAML files. The review feedback correctly identifies that the PR Title and Summary do not comply with the repository's style guide and provides suggestions to align them with the required format.
| --data-parallel-size 8 | ||
| --data-parallel-size-local 4 | ||
| --data-parallel-start-rank 4 | ||
| --headless |
There was a problem hiding this comment.
The Pull Request Title and Summary do not fully adhere to the repository style guide. Specifically, the PR Title is missing the [Action] prefix (e.g., [BugFix]), and the PR Summary headers should use H3 (###) instead of H2 (##).
Please update the PR Title and Summary to match the following suggestions:
Suggested PR Title:
[Test][BugFix] Drop remote --headless for GLM-5.1 hybrid DPSuggested PR Summary:
### What this PR does / why we need it?
vLLM 0.28 infers hybrid LB whenever `--data-parallel-start-rank` is set. In that mode, the DP handshake rejects remote engines that still pass `--headless` with the error:
`Remote engine must not use --headless in external or hybrid dp lb mode`
This failed the nightly `double-node` test for `GLM-5.1-W8A8C8-A3_198k_function`. This PR drops `--headless` on the remote node so both nodes run hybrid LB (each exposes an API and owns its local ranks). The same leftover flag is also removed from the sibling 128k nightly YAML.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
- Inspected vLLM main `EngineArgs.create_engine_config` (`data_parallel_start_rank is not None` implies hybrid LB) and `vllm/v1/engine/utils.py` handshake (`remote_should_be_headless`).
- YAML-only change; nightly double-node CI should re-run `GLM-5.1-W8A8C8-A3_198k_function`.References
- The PR Title and PR Summary must follow the specified format, with the PR Title using [Branch][Module][Action] format and the PR Summary using specific H3 headers inside markdown code blocks. (link)
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
Same vLLM 0.28 handshake reject as GLM-5.1 C8: start-rank implies hybrid LB, so GLM-5.1/5.2 A3 dual-node and Kimi-K2.5 A2 remote nodes cannot pass --headless. Signed-off-by: LQD <1107297340@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com>
|
/nightly GLM5_1-W8A8-A3-dual-nodes Kimi-K2_5-W4A8-A2-dual-nodes |
…-project#16133) ### What this PR does / why we need it? vLLM 0.28 infers hybrid LB whenever `--data-parallel-start-rank` is set (`data_parallel_start_rank is not None`). In that mode the DP handshake requires remote engines **not** to pass `--headless`: ```text Remote engine must not use --headless in external or hybrid dp lb mode ``` Nightly double-node / multi-node jobs still used the old internal-LB remote command (`--data-parallel-start-rank` + `--headless`), so startup was rejected during handshake. Drop `--headless` on the remote node so both nodes run hybrid LB (each exposes an API and owns its local ranks). Affected nightly YAMLs: - `tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_198k_function.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_128k_90_50.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/GLM5_1-W8A8-A3-dual-nodes.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/GLM5_2-W8A8-A3-dual-nodes.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/Kimi-K2_5-W4A8-A2-dual-nodes.yaml` ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? - YAML-only change. Inspected vLLM main `EngineArgs.create_engine_config` and `vllm/v1/engine/utils.py` handshake (`remote_should_be_headless`). - Nightly CI should re-run the five jobs above. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: LQD <1107297340@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com>
### What this PR does / why we need it? vLLM #47692 changed the hybrid DP load-balancing inference to treat an explicit `--data-parallel-start-rank 0` as meaningful. As a result, the primary node in these multi-node `internal_dp` configurations is now inferred as hybrid LB, while the remote node is intended to run headless for internal LB. PRs #16133, #16186, and #16204 removed `--headless` from remote nodes to satisfy the resulting handshake, but that changed the topology covered by the tests. The regular multi-node internal-DP runner sends benchmark traffic only to the primary node. In disaggregated-prefill cases, the PD proxy was explicitly designed to exclude headless nodes and target one API endpoint per DP group. Switching these deployments to hybrid/external LB therefore changes the intended test coverage. Restore the intended internal-LB topology in all 17 affected nightly and weekly configurations: - remove `--data-parallel-start-rank 0` from the primary node of each DP group, allowing rank 0 to be inferred without enabling hybrid LB; - restore `--headless` on the remote node while retaining its non-zero DP rank offset. This keeps one API endpoint on the primary node and lets it schedule requests across all local and remote DP ranks. ### Does this PR introduce _any_ user-facing change? No. This only updates nightly test deployment configurations. ### How was this patch tested? - `git diff --check` - Parsed all 17 modified YAML files with PyYAML and verified their internal-DP node roles. - Multi-node NPU nightly and weekly jobs should validate the complete deployments. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com> Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
…-project#16133) ### What this PR does / why we need it? vLLM 0.28 infers hybrid LB whenever `--data-parallel-start-rank` is set (`data_parallel_start_rank is not None`). In that mode the DP handshake requires remote engines **not** to pass `--headless`: ```text Remote engine must not use --headless in external or hybrid dp lb mode ``` Nightly double-node / multi-node jobs still used the old internal-LB remote command (`--data-parallel-start-rank` + `--headless`), so startup was rejected during handshake. Drop `--headless` on the remote node so both nodes run hybrid LB (each exposes an API and owns its local ranks). Affected nightly YAMLs: - `tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_198k_function.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_128k_90_50.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/GLM5_1-W8A8-A3-dual-nodes.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/GLM5_2-W8A8-A3-dual-nodes.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/Kimi-K2_5-W4A8-A2-dual-nodes.yaml` ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? - YAML-only change. Inspected vLLM main `EngineArgs.create_engine_config` and `vllm/v1/engine/utils.py` handshake (`remote_should_be_headless`). - Nightly CI should re-run the five jobs above. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: LQD <1107297340@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com>
### What this PR does / why we need it? vLLM #47692 changed the hybrid DP load-balancing inference to treat an explicit `--data-parallel-start-rank 0` as meaningful. As a result, the primary node in these multi-node `internal_dp` configurations is now inferred as hybrid LB, while the remote node is intended to run headless for internal LB. PRs vllm-project#16133, vllm-project#16186, and vllm-project#16204 removed `--headless` from remote nodes to satisfy the resulting handshake, but that changed the topology covered by the tests. The regular multi-node internal-DP runner sends benchmark traffic only to the primary node. In disaggregated-prefill cases, the PD proxy was explicitly designed to exclude headless nodes and target one API endpoint per DP group. Switching these deployments to hybrid/external LB therefore changes the intended test coverage. Restore the intended internal-LB topology in all 17 affected nightly and weekly configurations: - remove `--data-parallel-start-rank 0` from the primary node of each DP group, allowing rank 0 to be inferred without enabling hybrid LB; - restore `--headless` on the remote node while retaining its non-zero DP rank offset. This keeps one API endpoint on the primary node and lets it schedule requests across all local and remote DP ranks. ### Does this PR introduce _any_ user-facing change? No. This only updates nightly test deployment configurations. ### How was this patch tested? - `git diff --check` - Parsed all 17 modified YAML files with PyYAML and verified their internal-DP node roles. - Multi-node NPU nightly and weekly jobs should validate the complete deployments. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com> Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com>
…-project#16133) ### What this PR does / why we need it? vLLM 0.28 infers hybrid LB whenever `--data-parallel-start-rank` is set (`data_parallel_start_rank is not None`). In that mode the DP handshake requires remote engines **not** to pass `--headless`: ```text Remote engine must not use --headless in external or hybrid dp lb mode ``` Nightly double-node / multi-node jobs still used the old internal-LB remote command (`--data-parallel-start-rank` + `--headless`), so startup was rejected during handshake. Drop `--headless` on the remote node so both nodes run hybrid LB (each exposes an API and owns its local ranks). Affected nightly YAMLs: - `tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_198k_function.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_128k_90_50.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/GLM5_1-W8A8-A3-dual-nodes.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/GLM5_2-W8A8-A3-dual-nodes.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/Kimi-K2_5-W4A8-A2-dual-nodes.yaml` ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? - YAML-only change. Inspected vLLM main `EngineArgs.create_engine_config` and `vllm/v1/engine/utils.py` handshake (`remote_should_be_headless`). - Nightly CI should re-run the five jobs above. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: LQD <1107297340@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: tianming2009 <13246728590@163.com>
### What this PR does / why we need it? vLLM #47692 changed the hybrid DP load-balancing inference to treat an explicit `--data-parallel-start-rank 0` as meaningful. As a result, the primary node in these multi-node `internal_dp` configurations is now inferred as hybrid LB, while the remote node is intended to run headless for internal LB. PRs vllm-project#16133, vllm-project#16186, and vllm-project#16204 removed `--headless` from remote nodes to satisfy the resulting handshake, but that changed the topology covered by the tests. The regular multi-node internal-DP runner sends benchmark traffic only to the primary node. In disaggregated-prefill cases, the PD proxy was explicitly designed to exclude headless nodes and target one API endpoint per DP group. Switching these deployments to hybrid/external LB therefore changes the intended test coverage. Restore the intended internal-LB topology in all 17 affected nightly and weekly configurations: - remove `--data-parallel-start-rank 0` from the primary node of each DP group, allowing rank 0 to be inferred without enabling hybrid LB; - restore `--headless` on the remote node while retaining its non-zero DP rank offset. This keeps one API endpoint on the primary node and lets it schedule requests across all local and remote DP ranks. ### Does this PR introduce _any_ user-facing change? No. This only updates nightly test deployment configurations. ### How was this patch tested? - `git diff --check` - Parsed all 17 modified YAML files with PyYAML and verified their internal-DP node roles. - Multi-node NPU nightly and weekly jobs should validate the complete deployments. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com> Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com> Signed-off-by: tianming2009 <13246728590@163.com>
…-project#16133) ### What this PR does / why we need it? vLLM 0.28 infers hybrid LB whenever `--data-parallel-start-rank` is set (`data_parallel_start_rank is not None`). In that mode the DP handshake requires remote engines **not** to pass `--headless`: ```text Remote engine must not use --headless in external or hybrid dp lb mode ``` Nightly double-node / multi-node jobs still used the old internal-LB remote command (`--data-parallel-start-rank` + `--headless`), so startup was rejected during handshake. Drop `--headless` on the remote node so both nodes run hybrid LB (each exposes an API and owns its local ranks). Affected nightly YAMLs: - `tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_198k_function.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_128k_90_50.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/GLM5_1-W8A8-A3-dual-nodes.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/GLM5_2-W8A8-A3-dual-nodes.yaml` - `tests/e2e/nightly/multi_node/internal_dp/config/Kimi-K2_5-W4A8-A2-dual-nodes.yaml` ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? - YAML-only change. Inspected vLLM main `EngineArgs.create_engine_config` and `vllm/v1/engine/utils.py` handshake (`remote_should_be_headless`). - Nightly CI should re-run the five jobs above. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: LQD <1107297340@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: like-0517 <ithwlike@126.com>
### What this PR does / why we need it? vLLM #47692 changed the hybrid DP load-balancing inference to treat an explicit `--data-parallel-start-rank 0` as meaningful. As a result, the primary node in these multi-node `internal_dp` configurations is now inferred as hybrid LB, while the remote node is intended to run headless for internal LB. PRs vllm-project#16133, vllm-project#16186, and vllm-project#16204 removed `--headless` from remote nodes to satisfy the resulting handshake, but that changed the topology covered by the tests. The regular multi-node internal-DP runner sends benchmark traffic only to the primary node. In disaggregated-prefill cases, the PD proxy was explicitly designed to exclude headless nodes and target one API endpoint per DP group. Switching these deployments to hybrid/external LB therefore changes the intended test coverage. Restore the intended internal-LB topology in all 17 affected nightly and weekly configurations: - remove `--data-parallel-start-rank 0` from the primary node of each DP group, allowing rank 0 to be inferred without enabling hybrid LB; - restore `--headless` on the remote node while retaining its non-zero DP rank offset. This keeps one API endpoint on the primary node and lets it schedule requests across all local and remote DP ranks. ### Does this PR introduce _any_ user-facing change? No. This only updates nightly test deployment configurations. ### How was this patch tested? - `git diff --check` - Parsed all 17 modified YAML files with PyYAML and verified their internal-DP node roles. - Multi-node NPU nightly and weekly jobs should validate the complete deployments. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: jiangkaiqiang <jiangkaiqiang@huawei.com> Co-authored-by: jiangkaiqiang <jiangkaiqiang@huawei.com> Signed-off-by: like-0517 <ithwlike@126.com>
What this PR does / why we need it?
vLLM 0.28 infers hybrid LB whenever
--data-parallel-start-rankis set (data_parallel_start_rank is not None). In that mode the DP handshake requires remote engines not to pass--headless:Nightly double-node / multi-node jobs still used the old internal-LB remote command (
--data-parallel-start-rank+--headless), so startup was rejected during handshake.Drop
--headlesson the remote node so both nodes run hybrid LB (each exposes an API and owns its local ranks). Affected nightly YAMLs:tests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_198k_function.yamltests/e2e/nightly/multi_node/internal_dp/config/GLM-5.1-W8A8C8-A3_128k_90_50.yamltests/e2e/nightly/multi_node/internal_dp/config/GLM5_1-W8A8-A3-dual-nodes.yamltests/e2e/nightly/multi_node/internal_dp/config/GLM5_2-W8A8-A3-dual-nodes.yamltests/e2e/nightly/multi_node/internal_dp/config/Kimi-K2_5-W4A8-A2-dual-nodes.yamlDoes this PR introduce any user-facing change?
No.
How was this patch tested?
YAML-only change. Inspected vLLM main
EngineArgs.create_engine_configandvllm/v1/engine/utils.pyhandshake (remote_should_be_headless).Nightly CI should re-run the five jobs above.
vLLM main: vllm-project/vllm@b2f6858