Skip to content

fix(benchmarks): increase PD benchmark timeout from 240s to 480s - #681

Merged
slin1237 merged 1 commit into
mainfrom
slin/benchmark-timeout-increase
Mar 9, 2026
Merged

slin1237 merged 1 commit into
mainfrom
slin/benchmark-timeout-increase

Conversation

@slin1237

@slin1237 slin1237 commented Mar 9, 2026 •

Copy link
Copy Markdown
Member

Summary

The PD benchmark (test_pd_perf[pd_http]) consistently times out at the default 240s genai-bench subprocess timeout, causing CI failures with exit code -9.

What changed

  • e2e_test/benchmarks/test_pd_perf.py: Added explicit timeout_sec=480 to the genai_bench_runner call

Why

PD setup spins up 4 SGLang workers (2 prefill + 2 decode) plus a disaggregated gateway, which takes significantly longer to complete 200 requests compared to regular benchmarks. The default 240s timeout (from conftest.py) is insufficient.

Test plan

  • Verified the regular benchmarks (http, grpc) pass fine within 240s
  • Confirmed PD benchmark was consistently failing at exactly 240s (genai-bench timed out after 240s)
  • CI run on this PR should pass the PD benchmark

Summary by CodeRabbit

  • Chores
    • Optimized Docker configuration for benchmarks to cache HuggingFace dependencies, improving test execution efficiency by preventing unnecessary re-downloads across benchmark runs.

@coderabbitai

coderabbitai Bot commented Mar 9, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

The change adds HuggingFace cache directory mounting to the Docker configuration in the benchmarks test setup. The modification locates the HF_HOME environment variable (defaulting to ~/.cache/huggingface), and if present, mounts it as a volume into the container while setting the HF_HOME variable inside the container.

Changes

Cohort / File(s) Summary
Docker Volume Configuration
e2e_test/benchmarks/conftest.py
Adds mounting of HuggingFace cache directory into Docker run command to preserve tokenizer and token caches across test runs, avoiding repeated downloads.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~8 minutes

Suggested reviewers

  • CatherineSue
  • XinyueZhang369
  • key4ng

Poem

🐰 A cache is mounted, oh what delight!
No more downloads through the night,
HuggingFace tokens now persist and stay,
Benchmarks run faster, hip-hooray! 🎉

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Title check ⚠️ Warning The pull request title describes increasing a PD benchmark timeout from 240s to 480s, but the actual changes add HuggingFace cache mounting to the Docker configuration with no timeout modifications. Update the title to reflect the actual changes: 'fix(benchmarks): mount HuggingFace cache in benchmark Docker configuration' or similar description of the cache mounting functionality.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch slin/benchmark-timeout-increase

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a persistent CI issue where the PD benchmark (test_pd_perf[pd_http]) was consistently timing out. The timeout for this specific benchmark has been doubled to accommodate the longer execution time required by its complex setup, which includes multiple SGLang workers and a disaggregated gateway. This adjustment aims to stabilize CI by preventing premature termination of the benchmark.

Highlights

  • PD Benchmark Timeout Increased: The timeout for the test_pd_perf[pd_http] benchmark has been explicitly increased from 240 seconds to 480 seconds.
  • CI Stability Improvement: This change resolves consistent CI failures caused by the genai-bench subprocess timing out during the PD benchmark execution.
Changelog
  • e2e_test/benchmarks/test_pd_perf.py
    • Added timeout_sec=480 to the genai_bench_runner call within the test_pd_perf function.
Activity
  • The author verified that standard HTTP and gRPC benchmarks successfully complete within the default 240-second timeout.
  • The author confirmed that the PD benchmark consistently failed at exactly 240 seconds due to the genai-bench subprocess timing out.
  • The author anticipates that the CI run for this pull request will now successfully pass the PD benchmark.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions github-actions Bot added the tests Test changes label Mar 9, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request increases the timeout for the PD benchmark test to prevent CI failures, which is a necessary change. The implementation is straightforward. I have one suggestion to improve code maintainability by defining the new timeout value as a constant rather than using a magic number.

Comment thread e2e_test/benchmarks/test_pd_perf.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@e2e_test/benchmarks/test_pd_perf.py`:
- Line 22: Add a brief inline comment next to the timeout_sec=480 setting in
test_pd_perf.py explaining why the test needs an extended timeout (e.g.,
"extended to 480s because CI environment can take longer to boot/complete
performance setup or ray cluster provisioning; default is 240s"), similar to the
existing comment for max_requests_per_run so future maintainers understand the
rationale.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 4c880591-aefb-4ad7-8b6f-f25f406c82e0

📥 Commits

Reviewing files that changed from the base of the PR and between 076c8ca and a41f8e4.

📒 Files selected for processing (1)
  • e2e_test/benchmarks/test_pd_perf.py

Comment thread e2e_test/benchmarks/test_pd_perf.py Outdated
genai-bench runs inside an ephemeral Docker container (--rm) that has
no access to the host's HuggingFace cache. Each invocation downloads
the tokenizer fresh from HuggingFace, which intermittently hangs and
causes benchmark timeouts (exit code -9).

Mount the host's HF cache directory (~/.cache/huggingface) into the
container so genai-bench reuses tokenizers already downloaded by the
sglang workers that run on the host.

Also revert the timeout_sec=480 workaround from test_pd_perf since the
root cause is the HF download hang, not insufficient timeout.

Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@slin1237
slin1237 force-pushed the slin/benchmark-timeout-increase branch from a41f8e4 to ae122e0 Compare March 9, 2026 20:25

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
e2e_test/benchmarks/conftest.py (1)

56-66: 🧹 Nitpick | 🔵 Trivial

Minor redundancy with existing HF_HOME pass-through.

When HF_HOME is set in the environment and the directory exists, line 61 explicitly sets -e HF_HOME={hf_home}, but lines 64-66 will also add -e HF_HOME (pass-through). This results in Docker receiving the same variable twice, which is harmless but redundant.

Consider excluding HF_HOME from the loop when it's already handled by the new mount logic:

♻️ Suggested cleanup
     if os.path.isdir(hf_home):
         cmd.extend(["-v", f"{hf_home}:{hf_home}", "-e", f"HF_HOME={hf_home}"])
+        hf_home_handled = True
+    else:
+        hf_home_handled = False

     # Pass through environment variables the container may need
-    for var in ("HF_TOKEN", "HF_HOME"):
+    for var in ("HF_TOKEN",) if hf_home_handled else ("HF_TOKEN", "HF_HOME"):
         if os.environ.get(var):
             cmd.extend(["-e", var])

Alternatively, simply remove "HF_HOME" from the tuple since the new block now handles it explicitly (the mount logic covers both the env-set and default cases).

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@e2e_test/benchmarks/conftest.py` around lines 56 - 66, The code redundantly
adds HF_HOME twice to cmd; update the pass-through loop in conftest.py to skip
HF_HOME when it's already handled by the hf_home mount logic (referencing the
hf_home variable and the cmd.extend calls), e.g. either remove "HF_HOME" from
the tuple and only iterate ("HF_TOKEN",) or add a guard inside the for loop (if
var == "HF_HOME" and os.path.isdir(hf_home): continue) so HF_HOME is not
appended twice.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@e2e_test/benchmarks/conftest.py`:
- Around line 56-66: The code redundantly adds HF_HOME twice to cmd; update the
pass-through loop in conftest.py to skip HF_HOME when it's already handled by
the hf_home mount logic (referencing the hf_home variable and the cmd.extend
calls), e.g. either remove "HF_HOME" from the tuple and only iterate
("HF_TOKEN",) or add a guard inside the for loop (if var == "HF_HOME" and
os.path.isdir(hf_home): continue) so HF_HOME is not appended twice.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: b6f844a4-a172-401a-bb4e-aea164201334

📥 Commits

Reviewing files that changed from the base of the PR and between a41f8e4 and ae122e0.

📒 Files selected for processing (1)
  • e2e_test/benchmarks/conftest.py

@slin1237
slin1237 merged commit c9bbade into main Mar 9, 2026
27 checks passed
@slin1237
slin1237 deleted the slin/benchmark-timeout-increase branch March 9, 2026 21:23
@coderabbitai coderabbitai Bot mentioned this pull request Jun 13, 2026
4 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant