Skip to content

[Bugfix] Strengthen the check of X-data-parallel-rank in Hybrid LB mode - #32314

Merged
chaunceyjiang merged 3 commits into
vllm-project:mainfrom
openanolis:dtcccc/bugfix
Jan 15, 2026
Merged

chaunceyjiang merged 3 commits into
vllm-project:mainfrom
openanolis:dtcccc/bugfix

Conversation

@dtcccc

@dtcccc dtcccc commented Jan 14, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

When running vllm with hybrid LB mode, e.g.,

CUDA_VISIBLE_DEVICES=0,1 vllm serve  Qwen/Qwen2.5-7B-Instruct --port 8001 --data-parallel-size 4 --data-parallel-size-local 2  --data-parallel-hybrid-lb --enforce-eager
CUDA_VISIBLE_DEVICES=2,3 vllm serve  Qwen/Qwen2.5-7B-Instruct --port 8002 --data-parallel-size 4 --data-parallel-start-rank 2 --data-parallel-size-local 2  --data-parallel-hybrid-lb --enforce-eager

and then

port=8001 # or 8002
curl http://localhost:$port/v1/completions \
  -H "Content-Type: application/json" \
  -H "X-data-parallel-rank: 2" \
  -d '{
    "prompt": "Who are you",
    "max_tokens": 10,
    "temperature": 0
  }'

We will receive 500 Internal Server Error because index 2 is out of range about local dp size. So that the index of self.core_engines in DPLBAsyncMPClient is out of bound. Limit the X-data-parallel-rank of the request to local dp size if using hybrid LB mode to fix this issue.

After this patch, we will receive "data_parallel_rank 2 is out of range [0, 2)." as expected.

Test Plan

Test Result


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.
  • (Optional) Release notes update. If your change is user facing, please update the release notes draft in the Google Doc.

Signed-off-by: Tianchen Ding <dtcccc@linux.alibaba.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a bugfix to properly validate the X-data-parallel-rank in hybrid load balancing mode. The core of the fix is in vllm/v1/engine/input_processor.py, where the validation now correctly uses the local data parallel size as the upper bound for the rank. This prevents an out-of-bounds error and ensures a proper error message is returned to the user. Additionally, a new local_engines_only property has been added to ParallelConfig, which is a good refactoring that simplifies and clarifies the code in multiple files where this condition is checked. The changes are correct and well-implemented.

@njhill njhill left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @dtcccc, very nice changes!

@njhill njhill added the ready ONLY add when PR is ready to merge/full CI is needed label Jan 14, 2026
@njhill
njhill enabled auto-merge (squash) January 14, 2026 16:23
Signed-off-by: Tianchen Ding <dtcccc@linux.alibaba.com>
auto-merge was automatically disabled January 15, 2026 03:14

Head branch was pushed to by a user without write access

@dtcccc

dtcccc commented Jan 15, 2026

Copy link
Copy Markdown
Contributor Author

@njhill Hi, would you please merge it again? The previous commit missed adaptation of a test case.

@chaunceyjiang
chaunceyjiang merged commit 1e58482 into vllm-project:main Jan 15, 2026
52 checks passed
sammysun0711 pushed a commit to sammysun0711/vllm that referenced this pull request Jan 16, 2026
…de (vllm-project#32314)

Signed-off-by: Tianchen Ding <dtcccc@linux.alibaba.com>
akh64bit pushed a commit to akh64bit/vllm that referenced this pull request Jan 16, 2026
…de (vllm-project#32314)

Signed-off-by: Tianchen Ding <dtcccc@linux.alibaba.com>
mystous pushed a commit to mystous/vllm_hybrid that referenced this pull request May 10, 2026
…de (vllm-project#32314)

Signed-off-by: Tianchen Ding <dtcccc@linux.alibaba.com>
my-other-github-account pushed a commit to my-other-github-account/vllm that referenced this pull request May 15, 2026
…de (vllm-project#32314)

Signed-off-by: Tianchen Ding <dtcccc@linux.alibaba.com>
0826joyce pushed a commit to 0826joyce/vllm-serving-optimization that referenced this pull request May 19, 2026
…de (vllm-project#32314)

Signed-off-by: Tianchen Ding <dtcccc@linux.alibaba.com>
plasticchris pushed a commit to plasticchris/vllm that referenced this pull request Jul 20, 2026
…de (vllm-project#32314)

Signed-off-by: Tianchen Ding <dtcccc@linux.alibaba.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working frontend ready ONLY add when PR is ready to merge/full CI is needed v1

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants