Skip to content

fix: properly setup and register vLLM worker for external / hybrid load balancing. Update launch script - #6695

Merged
GuanLuo merged 11 commits into
mainfrom
gluo/1.0.0-dep
Mar 3, 2026
Merged

GuanLuo merged 11 commits into
mainfrom
gluo/1.0.0-dep

Conversation

@GuanLuo

@GuanLuo GuanLuo commented Feb 27, 2026 •

Copy link
Copy Markdown
Contributor

Signed-off-by: Guan Luo 41310872+GuanLuo@users.noreply.github.com

Overview:

Code change:

  • vLLM recently fix the DP rank checking within vLLM engine ([Bugfix] Strengthen the check of X-data-parallel-rank in Hybrid LB mode vllm-project/vllm#32314), so that in the case of external / hybrid load balancing, vLLM AsyncLLM will only accepts request with DP rank within the local range. So on Dynamo routing, we need to convert global DP rank to local DP rank before passing to vLLM engine.
  • To accommodate vLLM's definition of external / hybrid load balancing, instead of only registering "lead" worker (dp_rank=0), vLLM worker should always register itself to Dynamo with the global DP ranks that it manages. In such a way, Dynamo router can route request to corresponding worker and specify the "global" DP rank that the vLLM engine of the worker should internally send to. The worker will convert global rank to local rank as stated in the last bullet point.
  • As now all workers will be registered, instead of "lead" worker set up KV event publisher / StatLogger for all ranks, each worker is responsible for setup of its own ranks.

Script change: We are moving from external load balancing to hybrid load balancing of DPs.

Details:

Example log from KV router with 2 worker each owns 4 DP ranks (hybrid load balancing)

First request

curl localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-30B-A3B",
    "messages": [
    {
        "role": "user",
        "content": "In the heart of Eldoria, an ancient land of boundless magic and mysterious creatures, lies the long-forgotten city of Aeloria. Once a beacon of knowledge and power, Aeloria was buried beneath the shifting sands of time, lost to the world for centuries. You are an intrepid explorer, known for your unparalleled curiosity and courage, who has stumbled upon an ancient map hinting at ests that Aeloria holds a secret so profound that it has the potential to reshape the very fabric of reality. Your journey will take you through treacherous deserts, enchanted forests, and across perilous mountain ranges. Your Task: Character Background: Develop a detailed background for your character. Describe their motivations for seeking out Aeloria, their skills and weaknesses, and any personal connections to the ancient city or its legends. Are they driven by a quest for knowledge, a search for lost familt clue is hidden."
    }
    ],
    "stream": false,
    "max_tokens": 30
  }'
2026-03-03T11:11:49.725734Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709692 dp_rank=0 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:49.725759Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709692 dp_rank=1 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:49.725764Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709692 dp_rank=2 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:49.725770Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709692 dp_rank=3 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:49.725775Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709690 dp_rank=4 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:49.725779Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709690 dp_rank=5 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:49.725783Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709690 dp_rank=6 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:49.725787Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709690 dp_rank=7 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:49.725794Z  INFO dynamo_llm::kv_router::scheduler: Multiple workers tied with same logit, using tree size as tie-breaker
2026-03-03T11:11:49.725802Z  INFO dynamo_llm::kv_router::scheduler: Selected worker: worker_id=7587893241352709690 dp_rank=6, logit: 24.250, cached blocks: 0, tree size: 0, total blocks: 158788

Repeat request

2026-03-03T11:11:58.457304Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709692 dp_rank=0 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:58.457328Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709692 dp_rank=1 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:58.457332Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709692 dp_rank=2 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:58.457337Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709692 dp_rank=3 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:58.457341Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709690 dp_rank=4 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:58.457346Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709690 dp_rank=5 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:58.457351Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709690 dp_rank=6 with 12 cached blocks: 12.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 0.250 + 12.000
2026-03-03T11:11:58.457356Z  INFO dynamo_llm::kv_router::scheduler: Formula for worker_id=7587893241352709690 dp_rank=7 with 0 cached blocks: 24.250 = 1.0 * prefill_blocks + decode_blocks = 1.0 * 12.250 + 12.000
2026-03-03T11:11:58.457361Z  INFO dynamo_llm::kv_router::scheduler: Selected worker: worker_id=7587893241352709690 dp_rank=6, logit: 12.250, cached blocks: 12, tree size: 14, total blocks: 158788

Where should the reviewer start?

Related Issues: (use one of the action keywords Closes / Fixes / Resolves / Relates to)

  • closes GitHub issue: #xxx

Summary by CodeRabbit

  • New Features

    • Support for per-worker data-parallel ranges (start rank + size) so workers handle configurable DP rank windows.
    • Simplified example launches: consolidated single-worker launches and unified endpoints for hybrid data-parallel setups.
  • Chores

    • Metrics and endpoint publishing improved with asynchronous endpoint management and consistent stat publishing across DP ranks.
  • Bug Fixes

    • Monitoring and routing now respect DP start rank when enumerating and checking worker availability.

@GuanLuo
GuanLuo requested a review from a team as a code owner February 27, 2026 20:00
@GuanLuo
GuanLuo requested a review from a team February 27, 2026 20:00
@github-actions github-actions Bot added fix backend::vllm Relates to the vllm backend labels Feb 27, 2026
@coderabbitai

coderabbitai Bot commented Feb 27, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Replaces per-worker data-parallel size semantics with explicit (start_rank, size) ranges across vLLM, kv-router, and runtime config. Iteration and publishing of DP ranks now use worker-specific ranges; non-leader node gating logic was removed and new DP-range helpers were added.

Changes

Cohort / File(s) Summary
vLLM core & worker handlers
components/src/dynamo/vllm/main.py, components/src/dynamo/vllm/handlers.py, components/src/dynamo/vllm/publisher.py
Introduce per-worker DP range usage via get_dp_range_for_worker; iterate dp ranks using start+size; set runtime_config.data_parallel_start_rank and data_parallel_size; remove non-leader gating paths; update stat-logger to accept dp_rank and async endpoint creation.
Launch examples
examples/backends/vllm/launch/dep.sh, examples/backends/vllm/launch/dsr1_dep.sh
Switch from per-GPU/per-rank launches to single consolidated vLLM launch using hybrid DP flags (--data-parallel-hybrid-lb, --data-parallel-start-rank, --data-parallel-size*) and simplified endpoints.
KV-router / multi-worker sequencing
lib/kv-router/src/multi_worker_sequence.rs, lib/llm/src/kv_router/sequence.rs, lib/llm/src/kv_router/scheduler.rs, lib/llm/src/kv_router/queue.rs
Change data structures from dp_size → dp_range (start, size); update constructors, update_workers, and iteration logic to enumerate dp_ranks from start..start+size; propagate new types through scheduler and queue logic.
Discovery & monitoring
lib/llm/src/discovery/worker_monitor.rs, lib/llm/src/kv_router/scheduler.rs
Compute dp_start from runtime_config.data_parallel_start_rank and iterate dp ranks over start..start+size when populating monitoring maps and selector logic.
Runtime config & bindings
lib/llm/src/local_model/runtime_config.rs, lib/bindings/python/rust/llm/local_model.rs
Add data_parallel_start_rank: u32 to ModelRuntimeConfig (with default) and expose a Python setter set_data_parallel_start_rank in bindings.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Poem

🐇 I nibble code and count each rank,

start and size, no more blank,
ranges stitched from end to start,
workers play their rightful part,
hopping logs with joyful spark.

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 inconclusive)

Check name Status Explanation Resolution
Description check ❓ Inconclusive Description covers overview and implementation details but lacks specific reviewer guidance and has unresolved placeholder issue reference. Specify which files require close review and provide the actual related issue number instead of placeholder #xxx.
✅ Passed checks (2 passed)
Check name Status Explanation
Title check ✅ Passed Title accurately summarizes the main changes: fixing vLLM worker registration for external/hybrid load balancing and updating launch scripts.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@GuanLuo GuanLuo changed the title fix: fix vLLM config for DP > 1 to specify external (Dynamo) load balancing fix: drop data_parallel_rank arg in vllm generate() as Dynamo manages DP external load balancing Feb 27, 2026
@GuanLuo
GuanLuo requested a review from a team February 28, 2026 02:33
@GuanLuo
GuanLuo requested a review from a team as a code owner February 28, 2026 02:33
@GuanLuo GuanLuo changed the title fix: drop data_parallel_rank arg in vllm generate() as Dynamo manages DP external load balancing fix: add global DP rank to local DP rank conversion for external / hybrid load balancing. Update launch script Feb 28, 2026
@PeaBrane

Copy link
Copy Markdown
Contributor

Both dep.sh and dsr1_dep.sh still reference $i in --kv-events-config ($((20080 + i))), but the for i in ... loop was removed. $i is now undefined, so the port always evaluates to 20080 — all DP workers will try to bind to the same ZMQ endpoint.

@pull-request-size pull-request-size Bot added size/L and removed size/M labels Mar 3, 2026
@github-actions github-actions Bot added the router Relates to routing, KV-aware routing, etc. label Mar 3, 2026
@GuanLuo GuanLuo changed the title fix: add global DP rank to local DP rank conversion for external / hybrid load balancing. Update launch script fix: properly setup and register vLLM worker for external / hybrid load balancing. Update launch script Mar 3, 2026
@GuanLuo

GuanLuo commented Mar 3, 2026

Copy link
Copy Markdown
Contributor Author

@CodeRabbit full review

@coderabbitai

coderabbitai Bot commented Mar 3, 2026

Copy link
Copy Markdown
Contributor
✅ Actions performed

Full review triggered.

Comment thread components/src/dynamo/vllm/handlers.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🧹 Nitpick comments (1)
lib/kv-router/src/multi_worker_sequence.rs (1)

252-252: Rename new_dp_sizes to new_dp_ranges for semantic clarity.

The tuple payload is now a range (start, size), so the current name is slightly misleading for callers/readers.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@lib/kv-router/src/multi_worker_sequence.rs` at line 252, Rename the parameter
and all internal uses of new_dp_sizes to new_dp_ranges in the update_workers
method to reflect that the tuple is a range (start, size); update the function
signature pub fn update_workers(&self, new_dp_ranges: HashMap<u64, (u32, u32)>)
and replace every occurrence of new_dp_sizes within update_workers, plus update
all call sites that pass or reference that parameter name to use new_dp_ranges
so names remain consistent across the codebase.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@components/src/dynamo/vllm/handlers.py`:
- Around line 491-500: The function _to_local_dp_rank currently converts an
out-of-range explicit dp_rank into None which lets vLLM pick a DP rank; instead,
detect when routing.dp_rank is explicitly provided but falls outside this
worker’s dp_range and raise a hard rejection (e.g., throw a ValueError or a
RequestRoutingError) so the request fails rather than silently falling back.
Update _to_local_dp_rank to raise on out-of-range dp_rank and propagate that
change to any callers that pass data_parallel_rank (ensure callers
handle/propagate the exception); apply the same fix to the other duplicate
checks referenced (around the functions/blocks at the other locations you noted)
so all code paths fail fast on explicit, out-of-range routing.dp_rank.

In `@lib/llm/src/discovery/worker_monitor.rs`:
- Around line 459-465: The cleanup fallback assumes rank 0 which breaks when
data_parallel_start_rank != 0; update both cleanup branches that use vec![0] to
derive the fallback dp ranks from the actual data-parallel range for the worker
(use runtime_config.data_parallel_start_rank and
runtime_config.data_parallel_size to compute dp_start..dp_end) or, if
unavailable, derive them from existing keys in
worker_load_states/known_worker_dp_ranks for that lease_id; locate uses around
known_worker_dp_ranks, dp_start/dp_end, and worker_load_states cleanup branches
and replace the hardcoded vec![0] with a computed set based on those symbols so
stale metric series for non-zero start ranks are correctly cleaned up.

In `@lib/llm/src/kv_router/scheduler.rs`:
- Around line 444-446: The loop end uses unchecked addition of
data_parallel_start_rank and data_parallel_size which can overflow; update the
code around data_parallel_start_rank and the for loop to compute the end with
checked_add (e.g., data_parallel_start_rank.checked_add(data_parallel_size)) and
panic with a clear message if it returns None, then iterate using that computed
end (use the checked end as the upper bound in the for dp_rank in ... loop) so
invalid configs fail fast; refer to the variables data_parallel_start_rank and
data_parallel_size in scheduler.rs when making this change.

---

Nitpick comments:
In `@lib/kv-router/src/multi_worker_sequence.rs`:
- Line 252: Rename the parameter and all internal uses of new_dp_sizes to
new_dp_ranges in the update_workers method to reflect that the tuple is a range
(start, size); update the function signature pub fn update_workers(&self,
new_dp_ranges: HashMap<u64, (u32, u32)>) and replace every occurrence of
new_dp_sizes within update_workers, plus update all call sites that pass or
reference that parameter name to use new_dp_ranges so names remain consistent
across the codebase.

ℹ️ Review info

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 05f10e9 and c6cba54.

📒 Files selected for processing (12)
  • components/src/dynamo/vllm/handlers.py
  • components/src/dynamo/vllm/main.py
  • components/src/dynamo/vllm/publisher.py
  • examples/backends/vllm/launch/dep.sh
  • examples/backends/vllm/launch/dsr1_dep.sh
  • lib/bindings/python/rust/llm/local_model.rs
  • lib/kv-router/src/multi_worker_sequence.rs
  • lib/llm/src/discovery/worker_monitor.rs
  • lib/llm/src/kv_router/queue.rs
  • lib/llm/src/kv_router/scheduler.rs
  • lib/llm/src/kv_router/sequence.rs
  • lib/llm/src/local_model/runtime_config.rs
💤 Files with no reviewable changes (1)
  • components/src/dynamo/vllm/publisher.py

Comment thread components/src/dynamo/vllm/handlers.py
Comment thread lib/llm/src/discovery/worker_monitor.rs
Comment thread lib/llm/src/kv_router/scheduler.rs Outdated
GuanLuo added 4 commits March 3, 2026 12:02
…ancing

Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
… DP rank)

Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
GuanLuo added 3 commits March 3, 2026 12:14
…ting stats logger, KV event publisher for its DP ranks

Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>

@ptarasiewiczNV ptarasiewiczNV left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pls add the explicit logging on the "internal load balancing" code path. Otherwise approve on vllm/python side.

GuanLuo added 2 commits March 3, 2026 12:16
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
Comment thread components/src/dynamo/vllm/handlers.py Outdated
Comment thread components/src/dynamo/vllm/main.py
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
@GuanLuo
GuanLuo merged commit 90d7463 into main Mar 3, 2026
88 of 89 checks passed
@GuanLuo
GuanLuo deleted the gluo/1.0.0-dep branch March 3, 2026 22:00
GuanLuo added a commit that referenced this pull request Mar 3, 2026
…ad balancing. Update launch script (#6695)

Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
saturley-hall pushed a commit that referenced this pull request Mar 3, 2026
…ad balancing. Update launch script (#6695) (#6833)

Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
yao531441 pushed a commit to yao531441/dynamo that referenced this pull request May 13, 2026
…ad balancing. Update launch script (ai-dynamo#6695)

Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>

This branch was previously deployed

1 inactive deployment
GITLAB — e510909f Deployed Mar 3, 2026 by copy-pr-bot[bot] via Trigger CI Pipeline #19024
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend::vllm Relates to the vllm backend fix router Relates to routing, KV-aware routing, etc. size/L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants