Skip to content

fix(e2e): respect workers_config in vLLM/TRT-LLM gRPC multi-worker setup - #525

Merged
slin1237 merged 2 commits into
mainfrom
slin/fix-grpc-multi-worker
Feb 24, 2026
Merged

slin1237 merged 2 commits into
mainfrom
slin/fix-grpc-multi-worker

Conversation

@slin1237

@slin1237 slin1237 commented Feb 24, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Fix _setup_grpc_backend() to actually read workers_config["count"] and launch the requested number of gRPC workers for vLLM and TRT-LLM multi-worker tests
  • Single-worker path (count=1) is unchanged for backward compatibility

What changed

  • e2e_test/fixtures/setup_backend.py: Rewrote _setup_grpc_backend() to support multi-worker mode using the same pattern as _setup_local_backend() — get_workers_by_type() to find existing gRPC REGULAR workers, launch_workers() with WorkerIdentity for missing ones, proper cleanup of all instances

Why

Nightly benchmark run showed HTTP winning every metric by 4-27% aggregate, with vLLM multi-worker gRPC TTFT p99 up to 24× worse than HTTP. Investigation of gateway logs revealed the root cause: vLLM multi-worker tests had 4 HTTP workers but only 1 gRPC worker behind the gateway.

_setup_grpc_backend() accepted workers_config as a parameter but never read it — it always called model_pool.get_grpc_worker() which returns exactly one worker. SGLang was unaffected because SGLang gRPC falls through to _setup_local_backend() which correctly handles multi-worker.

How

  • When count > 1: uses get_workers_by_type(model_id, WorkerType.REGULAR), filters by ConnectionMode.GRPC, launches missing workers via launch_workers(), collects all worker URLs for the gateway
  • When count == 1: preserves existing get_grpc_worker() path (no behavior change)
  • Cleanup paths iterate over all instances instead of a single one
  • Both vLLM and TRT-LLM are fixed since both route through _setup_grpc_backend() (line 125: if is_vllm() or is_trtllm())

Test plan

  • Re-run nightly benchmarks and verify vLLM multi-worker gRPC tests now launch the correct number of workers (check gateway logs for worker count)
  • Verify vLLM single-worker gRPC tests still work (unchanged code path)
  • Verify SGLang tests are unaffected (different code path via _setup_local_backend)

Summary by CodeRabbit

  • Tests
    • Added robust multi-worker support to backend test setup, allowing tests to acquire and reuse multiple worker instances.
    • Improved error handling and cleanup to reliably release acquired workers on failure.
    • Updated test logging and messages to report worker counts and clearly indicate when additional workers are launched.

_setup_grpc_backend() accepted workers_config but never read it — it
always called model_pool.get_grpc_worker() which returns exactly one
worker.  When nightly benchmarks ran multi-worker tests (count=4),
vLLM gRPC got 1 worker while HTTP got 4, causing a 2-24× performance
gap that made protocol comparison results meaningless.

What changed:
- e2e_test/fixtures/setup_backend.py: rewrite _setup_grpc_backend() to
  read workers_config["count"] and, when count > 1, use
  get_workers_by_type() + launch_workers() to acquire/launch the
  requested number of gRPC workers — the same pattern used by
  _setup_local_backend() for SGLang.

Why:
- Nightly benchmark run showed HTTP winning every metric by 4-27%
  aggregate, with vLLM multi-worker gRPC TTFT p99 up to 24× worse.
  Investigation of gateway logs revealed vLLM multi-worker tests had
  4 HTTP workers but only 1 gRPC worker behind the gateway.

How:
- Single-worker path (count=1) is unchanged — still uses
  get_grpc_worker() for backward compatibility.
- Multi-worker path mirrors _setup_local_backend(): finds existing
  GRPC REGULAR workers, releases wrong-mode workers, launches missing
  ones via launch_workers() with WorkerIdentity, and collects all
  worker URLs for the gateway.
- Cleanup paths iterate over all instances instead of a single one.
- Both vLLM and TRT-LLM are fixed since both route through
  _setup_grpc_backend() at line 125.

Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@github-actions github-actions Bot added the tests Test changes label Feb 24, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @slin1237, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a critical issue in the e2e testing infrastructure where vLLM and TRT-LLM gRPC multi-worker configurations were not being correctly provisioned. Previously, only a single gRPC worker was launched regardless of the workers_config setting, leading to inaccurate benchmark results. The changes ensure that the _setup_grpc_backend() function now properly respects the desired worker count, dynamically launching and managing multiple gRPC workers to accurately reflect multi-worker test scenarios.

Highlights

  • Multi-worker gRPC setup fix: The _setup_grpc_backend() function now correctly reads workers_config["count"] to launch the specified number of gRPC workers for vLLM and TRT-LLM multi-worker tests.
  • Worker provisioning logic: Implemented logic to find existing gRPC workers, launch missing ones using launch_workers(), and collect all worker URLs, mirroring the pattern used by _setup_local_backend().
  • Backward compatibility: The single-worker path (when count=1) remains unchanged, preserving existing behavior and compatibility.
  • Enhanced cleanup: Cleanup mechanisms were updated to iterate over and release all worker instances, preventing resource leaks in multi-worker scenarios.
Changelog
  • e2e_test/fixtures/setup_backend.py
    • Refactored _setup_grpc_backend to support multi-worker gRPC configurations for vLLM and TRT-LLM.
    • Introduced logic to acquire existing gRPC workers and launch additional ones as needed based on workers_config["count"].
    • Updated error handling and cleanup routines to correctly release all provisioned worker instances.
    • Maintained the original single-worker setup path for backward compatibility.
Activity
  • An investigation was initiated after nightly benchmark runs showed significant performance discrepancies, with HTTP outperforming gRPC due to incorrect worker provisioning in gRPC multi-worker tests.
  • The root cause was identified as _setup_grpc_backend() not utilizing the workers_config parameter for gRPC worker counts.
  • This pull request was created to implement the necessary fixes to ensure accurate multi-worker gRPC test environments.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Feb 24, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Refactors test fixture backend setup to add multi-worker support across gRPC, local, and PD backends: discovers/releases/reuses/launches worker instances, aggregates worker URLs and model_path, and updates cleanup and logging to handle multiple acquired workers.

Changes

Cohort / File(s) Summary
Multi-Worker Backend Setup
e2e_test/fixtures/setup_backend.py
Refactored _setup_grpc_backend, _setup_local_backend, and _setup_pd_backend_common to support multi-worker flows. Added discovery via get_workers_by_type, filtering by connection mode, reuse or launch of missing workers, aggregation of worker_urls and model_path, enhanced error handling and cleanup to release all acquired instances, and logging that reports worker counts.

Sequence Diagram(s)

sequenceDiagram
    participant Setup as Setup Function
    participant Pool as Model Pool
    participant Workers as Worker Instances
    participant Gateway as Gateway

    rect rgba(100, 150, 200, 0.5)
    note over Setup,Workers: Multi-worker setup (num_workers > 1)

    Setup->>Pool: get_workers_by_type(model_id)
    Pool-->>Workers: list discovered instances
    Setup->>Workers: filter by ConnectionMode / release non-matching
    alt enough existing workers
        Setup->>Workers: reuse required instances
    else need more workers
        Setup->>Workers: launch missing workers
        Workers-->>Pool: register/acquire instances
    end
    Setup->>Setup: collect worker_urls & model_path
    Setup->>Gateway: start(prefill_workers, decode_workers, worker_urls, model_path)
    Gateway-->>Setup: ready
    end

    rect rgba(150, 100, 200, 0.5)
    note over Setup,Workers: Cleanup on failure
    Setup->>Workers: release all acquired instances
    Workers-->>Setup: released
    end
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Poem

🐰 I hopped to find more friends to share the load,
Launched a few, kept others on the road,
Counts replaced a lonely single URL,
Cleanup cleans each pawprint — neat and swell,
This rabbit cheers as workers dance in code.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title directly and clearly summarizes the main objective: fixing the gRPC multi-worker setup to respect workers_config in vLLM/TRT-LLM tests, which matches the core changes in the diff.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch slin/fix-grpc-multi-worker

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

The pull request correctly addresses the issue where _setup_grpc_backend ignored the worker count for vLLM and TRT-LLM gRPC setups. By adopting the multi-worker pattern from _setup_local_backend, it now supports launching and managing multiple gRPC workers. I've suggested improvements to ensure robust resource cleanup in case of failures, more precise validation of the launched worker count, and a review of the timeout mechanism for launching multiple workers. Additionally, I noted that this logic is now duplicated with _setup_local_backend, which presents a refactoring opportunity.

Comment thread e2e_test/fixtures/setup_backend.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@e2e_test/fixtures/setup_backend.py`:
- Around line 509-535: The else branch leaks already-acquired existing_grpc if
model_pool.launch_workers raises and silently proceeds when launch_workers
returns an empty list; fix by (1) building an acquired_workers list (include
existing_grpc) before calling model_pool.launch_workers so cleanup will release
them on exceptions (same pattern as acquired_prefills/decodes), (2) after
calling model_pool.launch_workers, immediately check if new_instances is empty
and call pytest.fail with a clear message if so, and (3) only set instances =
existing_grpc + new_instances after the success checks so the later exception
handler will see the correct acquired list; refer to existing_grpc,
new_instances, instances, and model_pool.launch_workers to locate the changes.
- Around line 579-584: The log currently prints the configured worker target via
num_workers which can be misleading; update the log call that formats "Setup %s
gRPC backend: model=%s, workers=%d, gateway=%s, policy=%s" to pass
len(instances) (the actual allocated worker count) instead of num_workers so the
message reflects real allocation (keep the other fields runtime_label, model_id,
gateway.base_url, gateway_config["policy"] unchanged).

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between c949b32 and 55e56f5.

📒 Files selected for processing (1)
  • e2e_test/fixtures/setup_backend.py

Comment thread e2e_test/fixtures/setup_backend.py
Comment thread e2e_test/fixtures/setup_backend.py
…worker

Address PR review feedback for two bugs:

1. Resource leak: existing_grpc workers (already acquired by
   get_workers_by_type) were not tracked in `instances` until after
   launch_workers completed. If launch_workers raised, those workers
   were never released — blocking GPU allocation for subsequent tests.
   Fix: assign `instances = list(existing_grpc)` immediately and use
   `instances.append()` for new workers.

2. Silent under-provisioning: when launch_workers returned [] due to
   insufficient GPUs, `instances = existing_grpc + []` could be
   non-empty (e.g. 1 of 4 needed), passing the `if not instances`
   check and running the test with fewer workers than configured.
   Fix: fail explicitly when launch_workers returns [] and check
   `len(instances) < num_workers` instead of `not instances`.

Also fix the setup log to report len(instances) (actual) instead of
num_workers (configured target).

Signed-off-by: Simo Lin <linsimo.mark@gmail.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
e2e_test/fixtures/setup_backend.py (2)

704-711: ⚠️ Potential issue | 🟡 Minor

Log still reports configured num_workers, not actual len(instances).

Line 708 logs num_workers (the configured target). The identical issue was fixed in _setup_grpc_backend (line 594 now uses len(instances)). Apply the same fix here for consistency and observability.

🔧 Proposed fix
     logger.info(
         "Setup %s backend: model=%s, workers=%d, gateway=%s, policy=%s",
         backend_name,
         model_id,
-        num_workers,
+        len(instances),
         gateway.base_url,
         gateway_config["policy"],
     )
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@e2e_test/fixtures/setup_backend.py` around lines 704 - 711, The logger call
uses the configured num_workers rather than the actual launched instance count;
update the logger.info invocation that currently prints backend_name, model_id,
num_workers, gateway.base_url, gateway_config["policy"] to instead pass
len(instances) for the workers value (mirror the fix made in
_setup_grpc_backend), so the message reports the actual number of launched
instances.

640-658: ⚠️ Potential issue | 🟠 Major

_setup_local_backend else branch has the same two bugs that were just fixed in _setup_grpc_backend — resource leak + silent under-provisioning.

Bug 1 — existing_for_mode workers leak when launch_workers raises.
instances is [] when launch_workers executes (line 651). If it raises, the except at line 667 iterates over an empty list and all workers already acquired via get_workers_by_type are silently leaked.

Bug 2 — Partial allocation silently proceeds when launch_workers returns [].
There is no if not new_instances: pytest.fail(...) guard. If GPU capacity is insufficient, new_instances = [], so instances = existing_for_mode + []. When existing_for_mode is non-empty (e.g., 1 of 3 needed workers already exists), if not instances: at line 657 is False and the test runs under-provisioned — exactly the bug this PR was written to fix in gRPC.

Apply the same three-part fix used in _setup_grpc_backend (lines 502, 533–537, 543–547):

🐛 Proposed fix mirroring _setup_grpc_backend pattern
         else:
             missing = num_workers - len(existing_for_mode)
             workers_to_launch = [
                 WorkerIdentity(
                     model_id,
                     connection_mode,
                     WorkerType.REGULAR,
                     len(existing_for_mode) + i,
                 )
                 for i in range(missing)
             ]
+            instances = list(existing_for_mode)  # track before launch so cleanup fires on failure
             new_instances = model_pool.launch_workers(workers_to_launch, startup_timeout=300)
+            if not new_instances:
+                pytest.fail(
+                    f"Failed to launch {missing} workers for {model_id}: "
+                    f"GPU allocation failed"
+                )
             # Acquire newly launched instances
             for inst in new_instances:
                 inst.acquire()
-            instances = existing_for_mode + new_instances
+                instances.append(inst)

-        if not instances:
-            pytest.fail(f"Failed to get {num_workers} workers for {model_id}")
+        if len(instances) < num_workers:
+            pytest.fail(
+                f"Only got {len(instances)}/{num_workers} workers for {model_id}"
+            )
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@e2e_test/fixtures/setup_backend.py` around lines 640 - 658, In
_setup_local_backend, fix two issues: (1) avoid leaking already-acquired workers
in existing_for_mode if model_pool.launch_workers raises by wrapping the launch
call in try/except and releasing (calling .release() or the same cleanup used
elsewhere) all entries in existing_for_mode on exception before re-raising or
failing; (2) prevent silent under-provisioning by checking new_instances after
the launch and calling pytest.fail (with a clear message) if new_instances is
empty (mirroring the _setup_grpc_backend pattern); ensure you still acquire
newly launched instances (inst.acquire()) and then set instances =
existing_for_mode + new_instances and finally assert/pytest.fail if instances is
empty. Reference symbols: _setup_local_backend, existing_for_mode,
model_pool.launch_workers, new_instances, inst.acquire(), instances,
pytest.fail.
♻️ Duplicate comments (1)
e2e_test/fixtures/setup_backend.py (1)

492-608: _setup_grpc_backend multi-worker implementation looks correct — fixes from past review are properly applied.

All three issues flagged in the previous review cycle are confirmed resolved:

  • Resource leak: instances = list(existing_grpc) at line 502 pre-seeds cleanup tracking before launch_workers, matching the _setup_pd_backend_common pattern.
  • Silent under-provisioning: if not new_instances: pytest.fail(...) at lines 533–537 explicitly fails fast; the guard at line 543 uses len(instances) < num_workers instead of the weaker if not instances:.
  • Log accuracy: line 594 now emits len(instances) (actual acquired) instead of num_workers (configured target).
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@e2e_test/fixtures/setup_backend.py` around lines 492 - 608, The
implementation of _setup_grpc_backend looks good and the previous issues are
resolved; remove the stray duplicate review tag by deleting the redundant
"[duplicate_comment]" marker from the PR review/comment text so the approval is
unambiguous (no code changes required to functions like _setup_grpc_backend,
instances tracking, or Gateway usage).
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@e2e_test/fixtures/setup_backend.py`:
- Around line 704-711: The logger call uses the configured num_workers rather
than the actual launched instance count; update the logger.info invocation that
currently prints backend_name, model_id, num_workers, gateway.base_url,
gateway_config["policy"] to instead pass len(instances) for the workers value
(mirror the fix made in _setup_grpc_backend), so the message reports the actual
number of launched instances.
- Around line 640-658: In _setup_local_backend, fix two issues: (1) avoid
leaking already-acquired workers in existing_for_mode if
model_pool.launch_workers raises by wrapping the launch call in try/except and
releasing (calling .release() or the same cleanup used elsewhere) all entries in
existing_for_mode on exception before re-raising or failing; (2) prevent silent
under-provisioning by checking new_instances after the launch and calling
pytest.fail (with a clear message) if new_instances is empty (mirroring the
_setup_grpc_backend pattern); ensure you still acquire newly launched instances
(inst.acquire()) and then set instances = existing_for_mode + new_instances and
finally assert/pytest.fail if instances is empty. Reference symbols:
_setup_local_backend, existing_for_mode, model_pool.launch_workers,
new_instances, inst.acquire(), instances, pytest.fail.

---

Duplicate comments:
In `@e2e_test/fixtures/setup_backend.py`:
- Around line 492-608: The implementation of _setup_grpc_backend looks good and
the previous issues are resolved; remove the stray duplicate review tag by
deleting the redundant "[duplicate_comment]" marker from the PR review/comment
text so the approval is unambiguous (no code changes required to functions like
_setup_grpc_backend, instances tracking, or Gateway usage).

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 55e56f5 and ae9c1de.

📒 Files selected for processing (1)
  • e2e_test/fixtures/setup_backend.py

@slin1237
slin1237 merged commit e939bda into main Feb 24, 2026
17 of 20 checks passed
@slin1237
slin1237 deleted the slin/fix-grpc-multi-worker branch February 24, 2026 06:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant