Skip to content

fix(e2e): migrate genai-bench to Docker and fix router pipe hang - #403

Merged
key4ng merged 23 commits into
mainfrom
fix-bench
Feb 11, 2026
Merged

key4ng merged 23 commits into
mainfrom
fix-bench

Conversation

@key4ng

@key4ng key4ng commented Feb 11, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

  1. Router/worker process hangs: The e2e test infrastructure used subprocess.PIPE for stdout/stderr but never read from them. When running with log_level="debug", the 64KB OS pipe buffer filled up, causing the router and worker processes to block indefinitely on write operations.
  2. genai-bench dependency management: Installing genai-bench via pip added complexity to CI setup and required version pinning across multiple workflow files.
  3. Missing worker logs: When benchmarks or tests failed, worker output was lost because it was redirected to PIPE but never captured.

Solution

  1. Fix pipe buffer overflow: Changed subprocess.PIPE to subprocess.DEVNULL when show_output=False in both gateway.py and model_pool.py.
  2. Migrate to Docker: Replaced pip install genai-bench with docker run ghcr.io/moirai-internal/genai-bench:0.0.3. Container is named for docker logs retrieval in CI.
  3. Add log collection: When E2E_LOG_DIR env var is set, worker stdout/stderr are redirected to log files (worker-{model}-{port}.log) and uploaded as artifacts.

Changes

  • e2e_test/infra/gateway.py: Changed subprocess.PIPE to subprocess.DEVNULL for router output
  • e2e_test/infra/model_pool.py: PIPE→DEVNULL fix + log_dir support for worker log file redirection
  • e2e_test/fixtures/pool.py: Read E2E_LOG_DIR env var and pass to ModelPool
  • e2e_test/benchmarks/conftest.py: Rewrote genai-bench invocation to use docker run with named containers and docker logs for CI output
  • nightly-benchmark.yml / pr-test-rust.yml: Replaced pip install with docker pull, added E2E_LOG_DIR and log directory setup
  • test_nightly_perf.py: Added openai/gpt-oss-20b to nightly benchmarks

Test Plan

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • New Features
    • Benchmarks now run inside a standardized container image (configurable via GENAI_BENCH_IMAGE) and honor an optional E2E_LOG_DIR for per-worker and centralized nightly logs.
  • Bug Fixes
    • Improved subprocess output handling to avoid gateway startup hangs when output is hidden.
  • Tests
    • Nightly performance adds a new large-model configuration (openai/gpt-oss-20b).
  • Chores
    • CI pulls the benchmark image earlier and ensures benchmark logs are published to the nightly logs directory.

@coderabbitai

coderabbitai Bot commented Feb 11, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Benchmark execution moved from invoking a local CLI to running genai-bench inside a Docker container; CI workflows were updated to pull and use GENAI_BENCH_IMAGE and expose E2E_LOG_DIR; test infra now passes log_dir into ModelPool and captures per-worker logs. A new nightly model entry for openai/gpt-oss-20b was added.

Changes

Cohort / File(s) Summary
CI workflows
.github/workflows/nightly-benchmark.yml, .github/workflows/pr-test-rust.yml
Added GENAI_BENCH_IMAGE env and E2E_LOG_DIR, added explicit "Pull genai-bench image" steps across matrix jobs, removed/standardized extra_deps usage and switched to a generic dependency install step.
Benchmark runner (containerization)
e2e_test/benchmarks/conftest.py
Rewrote command construction to run genai-bench via docker run (volume mounts, env passthrough including GENAI_BENCH_IMAGE, HF_TOKEN, HF_HOME, E2E_LOG_DIR), added _DEFAULT_IMAGE, and adjusted subprocess/timeout/error handling for container runs.
Nightly test matrix
e2e_test/benchmarks/test_nightly_perf.py
Added ("openai/gpt-oss-20b", "GptOss20b", 1, ["http", "grpc"], {}) to _NIGHTLY_MODELS.
Fixtures: pool, hooks & backend setup
e2e_test/fixtures/pool.py, e2e_test/fixtures/hooks.py, e2e_test/fixtures/setup_backend.py
Passed E2E_LOG_DIR into ModelPool constructor, adjusted pytest evaluation lambda to accept extra kwargs, and added a local type annotation for client in cloud backend setup.
Infra logging & process handling
e2e_test/infra/model_pool.py, e2e_test/infra/gateway.py
ModelPool.__init__ now accepts log_dir; when show_output is False and log_dir provided, per-worker log files are created and stdout/stderr routed to them; log files are closed on eviction/shutdown. gateway._launch redirects stdout/stderr to DEVNULL when show_output is False.

Sequence Diagram

sequenceDiagram
    participant CI as CI Workflow
    participant Test as Test Runner
    participant Docker as Docker Engine
    participant Container as genai-bench Container
    participant Router as Model Router/Server
    participant Logs as E2E Log Dir

    CI->>Docker: Pull GENAI_BENCH_IMAGE
    Test->>Docker: Start container (docker run) with volumes & env (GENAI_BENCH_IMAGE, HF_*, E2E_LOG_DIR)
    Docker->>Container: Launch genai-bench
    Container->>Router: Connect to router_url (http / grpc) and run workload
    Container->>Logs: Write results & per-worker logs to mounted E2E_LOG_DIR
    Container-->>Docker: Exit when finished
    Test->>Logs: Collect and parse results
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • XinyueZhang369
  • slin1237

Poem

🐇 I hop into Docker, nose twitching with glee,
ENV clovers and logs tucked safe where they’ll be,
GPT‑OSS sprouts nightly, a new leafy test,
I nibble the output and file it with zest,
Hopping off home, benchmark carrots for me.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title accurately summarizes the main changes: migrating genai-bench to Docker and fixing router pipe hang issues by replacing PIPE with DEVNULL and adding log directory support.
Docstring Coverage ✅ Passed Docstring coverage is 85.71% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch fix-bench

No actionable comments were generated in the recent review. 🎉


Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions github-actions Bot added ci CI/CD configuration changes tests Test changes labels Feb 11, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @key4ng, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly refactors the end-to-end benchmarking setup by transitioning genai-bench execution to Docker containers. This change aims to standardize the testing environment, making benchmark results more reliable and reproducible across different development and CI/CD environments. Concurrently, a new Mistral model has been added to the nightly performance tests, broadening the scope of models being evaluated.

Highlights

  • Containerized Benchmarking: The genai-bench benchmarks are now executed within a Docker container, enhancing environment consistency and isolation for tests.
  • Docker Command Generation: The _build_command function in conftest.py was updated to construct docker run commands, including necessary volume mounts and environment variable passthrough.
  • New Model Added to Nightly Tests: The Mistral-7B-Instruct-v0.3 model has been integrated into the nightly performance test suite to expand benchmark coverage.
Changelog
  • e2e_test/benchmarks/conftest.py
    • Refactored the _build_command function to generate Docker commands for running genai-bench.
    • Removed the direct dependency on a locally installed genai-bench CLI.
    • Updated error handling to reflect Docker-based execution, checking for Docker availability instead of the genai-bench executable.
  • e2e_test/benchmarks/test_nightly_perf.py
    • Added mistralai/Mistral-7B-Instruct-v0.3 to the list of models included in the nightly performance benchmarks.
Ignored Files
  • Ignored by pattern: .github/workflows/** (2)
    • .github/workflows/nightly-benchmark.yml
    • .github/workflows/pr-test-rust.yml
Activity
  • No specific activity (comments, reviews, or progress updates) has been recorded for this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the genai-bench execution in e2e tests to use Docker, which improves environment consistency. It also adds a new model to the nightly performance tests. My review includes a suggestion to improve the maintainability of the Docker image versioning and points out a side effect of the new model configuration that leads to redundant test runs.

Comment thread e2e_test/benchmarks/conftest.py
Comment thread e2e_test/benchmarks/test_nightly_perf.py Outdated
@key4ng key4ng changed the title test fix(e2e): migrate genai-bench to Docker and fix router pipe hang Feb 11, 2026
@key4ng
key4ng marked this pull request as ready for review February 11, 2026 15:10
@key4ng
key4ng requested a review from CatherineSue as a code owner February 11, 2026 15:10

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
e2e_test/benchmarks/test_nightly_perf.py (1)

98-103: ⚠️ Potential issue | 🔴 Critical

Tensor parallelism mismatch in nightly matrix—update tp from 1 to 2.

The model openai/gpt-oss-20b exists in MODEL_SPECS with tp: 2 (line 90, e2e_test/infra/model_specs.py), but the nightly matrix at line 102 specifies tp: 1. This configuration mismatch will cause runtime failures. Update to use tp: 2 to match the model's definition.

e2e_test/benchmarks/conftest.py (1)

35-95: ⚠️ Potential issue | 🟡 Minor

Mount E2E_LOG_DIR into the container when it's outside the working directory.

If E2E_LOG_DIR is set to an absolute path outside the base directory, the container won't have access to write logs there. The current code passes --log-dir without mounting the directory, causing writes to fail inside the container.

🔧 Suggested fix
-    base_dir = str(Path.cwd())
+    base_dir_path = Path.cwd()
+    base_dir = str(base_dir_path)
@@
-    log_dir = os.environ.get("E2E_LOG_DIR")
-    if log_dir:
-        cmd.extend(["--log-dir", log_dir])
+    log_dir = os.environ.get("E2E_LOG_DIR")
+    if log_dir:
+        log_dir_path = Path(log_dir)
+        if not log_dir_path.is_absolute():
+            log_dir_path = base_dir_path / log_dir_path
+        log_dir_abs = str(log_dir_path.resolve())
+        try:
+            log_dir_path.resolve().relative_to(base_dir_path)
+        except ValueError:
+            cmd.extend(["-v", f"{log_dir_abs}:{log_dir_abs}"])
+        cmd.extend(["--log-dir", log_dir_abs])
🤖 Fix all issues with AI agents
In `@e2e_test/infra/model_pool.py`:
- Around line 327-341: The ModelPool currently accumulates open per-worker log
file handles (tracked only in self._log_files and closed in shutdown()), causing
FD leaks when instances are evicted and recreated; modify ModelPool to track log
file handles per instance key (the same key used in self.instances, e.g.,
"model_id:mode") and ensure the file handle for that key is closed and removed
from the tracking structure whenever an instance is evicted (the eviction path
that removes entries from self.instances), and also update all places that open
worker logs to register the handle under that instance key and to cleanly
close/remove it on eviction or final shutdown (references: ModelPool.__init__,
self._log_files, self.instances, shutdown()).
🧹 Nitpick comments (1)
.github/workflows/nightly-benchmark.yml (1)

152-155: Consider enabling E2E_LOG_DIR for multi-worker/H200 jobs too.

Line 152: single-worker now sets E2E_LOG_DIR; if you want per-worker logs consistently across nightly runs, mirror this env + mkdir in the multi-worker and single-worker-h200 blocks.

Comment thread e2e_test/infra/model_pool.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@e2e_test/infra/model_pool.py`:
- Around line 589-610: If subprocess.Popen can raise after opening a log file,
ensure the opened file handle is closed and removed from self._log_files to
avoid FD leaks: wrap the Popen call in a try/except/finally (or try/except)
around the block where log_file is created (refer to variables log_file and
self._log_files and the Popen invocation that assigns proc), and on exception
close log_file and pop self._log_files[key] (and if created on disk optionally
delete the file); apply the same cleanup logic to the other launch path
referenced around lines 1257-1277 (the HTTP + gRPC launch sections).

Comment thread e2e_test/infra/model_pool.py Outdated
@key4ng

key4ng commented Feb 11, 2026

Copy link
Copy Markdown
Member Author

@key4ng

key4ng commented Feb 11, 2026

Copy link
Copy Markdown
Member Author

stored genai-bench worker logs

Screenshot 2026-02-11 at 10 11 49 AM

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants