Skip to content

refactor(ci): split nightly benchmark into per-model jobs - #327

Merged
slin1237 merged 2 commits into
mainfrom
nightly-benchmark
Feb 5, 2026
Merged

slin1237 merged 2 commits into
mainfrom
nightly-benchmark

Conversation

@slin1237

@slin1237 slin1237 commented Feb 5, 2026 •

Copy link
Copy Markdown
Member

Summary

Refactors the nightly benchmark workflow to run each model as a separate job instead of running all models in a single pytest session. This improves debuggability and fault isolation.

Closes: N/A (workflow improvement)

What changed

.github/workflows/nightly-benchmark.yml:

  • Replaced monolithic single-sglang and single-vllm jobs with a unified single-worker matrix job
  • Matrix: 9 models × 2 runtimes (sglang, vllm) = 18 separate jobs
  • Added fail-fast: false so one model failure doesn't stop other models
  • Added if: !cancelled() to ensure jobs run even if dependencies fail
  • Updated job dependencies: multi-worker now depends on single-worker
  • Added llama-4-maverick-17b to single-worker matrix (was missing)
  • Temporarily added PR trigger for testing this workflow

e2e_test/benchmarks/test_nightly_perf.py:

  • Enabled _TEST_MODE = True for fast benchmark runs during testing
  • Uses 1 concurrency level, 10 requests, 5min timeout instead of full settings

Why

  • Before: All models ran in a single pytest session. If one model failed, the entire step failed, making it impossible to determine which model caused the failure.
  • After: Each model runs as a separate job with clear pass/fail status in the GitHub UI. Failures are isolated and don't block other models.

How

  • Unified the matrix pattern between single-worker and multi-worker jobs
  • Each matrix entry runs: checkout → install backend → download wheel → run pytest with model-specific -k filter → upload artifacts → cleanup
  • Filter step skips jobs based on models and runtime workflow_dispatch inputs

Test plan

  • PR triggers the workflow (via pull_request trigger on workflow file changes)
  • Each model appears as a separate job in the GitHub Actions UI
  • Jobs complete quickly due to _TEST_MODE = True
  • If a model fails, other models continue running
  • Summary job runs and generates report

Before merging

  • Revert _TEST_MODE = True back to False
  • Uncomment the schedule trigger
  • Remove the pull_request trigger

Summary by CodeRabbit

  • Chores
    • Optimized nightly benchmark workflow to run models and variants in isolated jobs with improved filtering and parallelization controls. Updated job dependencies and artifact naming conventions for better organization and traceability of benchmark results.

What changed:
- .github/workflows/nightly-benchmark.yml: Replace monolithic single-sglang
  and single-vllm jobs with a matrix-based single-worker job (9 models × 2
  runtimes). Each model now runs as a separate job with fail-fast: false,
  allowing failures to be isolated without blocking other models.
- e2e_test/benchmarks/test_nightly_perf.py: Enable _TEST_MODE for faster
  benchmark runs during workflow testing (1 concurrency, 10 requests, 5min
  timeout instead of full benchmark settings).

Why:
- Previous workflow ran all models in a single pytest session, making it
  impossible to debug which model failed and causing the entire workflow
  to fail on any single model failure.
- With per-model jobs, failures are isolated, easier to identify, and don't
  block other models from running.

How:
- Unified single-worker and multi-worker jobs to use the same matrix pattern
- Added `if: !cancelled()` and `fail-fast: false` to ensure all jobs run
- Updated job dependencies: multi-worker depends on single-worker
- Temporarily enabled _TEST_MODE and added PR trigger for testing

Refs: workflow testing
@github-actions github-actions Bot added ci CI/CD configuration changes tests Test changes labels Feb 5, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @slin1237, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly improves the robustness and maintainability of the nightly benchmark CI pipeline. By isolating each model's benchmark run into a separate job, it becomes much easier to identify and debug failures without impacting the overall benchmark suite. This change streamlines the process of evaluating model performance and ensures that issues are pinpointed efficiently.

Highlights

  • CI Workflow Refactoring: The nightly benchmark workflow has been refactored to execute each model as a distinct job, moving away from a single pytest session. This significantly enhances debuggability and fault isolation for benchmark runs.
  • Job Isolation and Resilience: Implemented fail-fast: false to ensure that a failure in one model's benchmark job does not halt the execution of other models. Additionally, if: !cancelled() was added to allow jobs to proceed even if upstream dependencies fail.
  • New Model Inclusion: The llama-4-maverick-17b model has been added to the single-worker matrix, ensuring its inclusion in the nightly benchmark suite.
  • Temporary Test Mode for Benchmarks: The _TEST_MODE flag in e2e_test/benchmarks/test_nightly_perf.py has been temporarily set to True to enable faster, reduced-scale benchmark runs for testing the workflow changes. This will be reverted before merging.
Changelog
  • .github/workflows/nightly-benchmark.yml
    • Replaced monolithic single-sglang and single-vllm jobs with a unified single-worker matrix job, creating 18 separate jobs for 9 models across 2 runtimes (sglang, vllm).
    • Added fail-fast: false to prevent single job failures from stopping the entire workflow.
    • Included if: !cancelled() to ensure jobs run even if dependencies encounter issues.
    • Updated job dependencies, making multi-worker dependent on single-worker.
    • Added llama-4-maverick-17b to the single-worker matrix.
    • Temporarily added a pull_request trigger for testing the workflow changes.
  • e2e_test/benchmarks/test_nightly_perf.py
    • Set _TEST_MODE to True to enable fast benchmark runs during workflow testing, using 1 concurrency level, 10 requests, and a 5-minute timeout.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/nightly-benchmark.yml
Activity
  • The pull request description outlines the intent and changes, and includes a detailed test plan for verifying the workflow modifications.
  • The author has provided a clear 'Before merging' checklist, indicating steps to revert temporary testing configurations.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Feb 5, 2026 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

📝 Walkthrough

Walkthrough

The GitHub Actions nightly benchmark workflow is restructured from separate single-sglang and single-vllm jobs to a unified single-worker job using matrices for models and variants, with conditional filtering and dynamic installation paths based on runtime selection.

Changes

Cohort / File(s) Summary
Nightly Benchmark Workflow Restructuring
.github/workflows/nightly-benchmark.yml
Consolidates separate single-sglang and single-vllm jobs into a matrix-driven single-worker job with per-model and per-variant execution; adds skip-filtering step to gate all downstream steps; updates installation to conditionally select ci_install_sglang.sh or ci_install_vllm.sh based on matrix.variant.runtime; modifies artifact naming with pattern nightly-${{ matrix.model.id }}-single-${{ matrix.variant.id }}-${{ github.run_id }}; updates job dependencies for multi-worker and summarize-benchmarks.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~22 minutes

Possibly related PRs

Suggested reviewers

  • key4ng
  • XinyueZhang369

Poem

🐰 Two jobs once ran their separate race,
Now matrix weaves them into place!
Filters gate each model's flight,
One worker hops with unified might,
Benchmarks bloom in branches bright! 🌙✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: refactoring the nightly benchmark workflow to split jobs by model instead of running all models in one job.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch nightly-benchmark

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the nightly benchmark workflow to run each model in a separate job, which is a great improvement for debuggability and fault isolation. The code change enables a test mode to allow for faster validation of the new workflow. My review includes a suggestion to make this test mode activation more robust by using an environment variable instead of a hardcoded boolean. This will prevent accidentally merging the code with test settings enabled, removing the need for a manual revert step.

Comment on lines +33 to +34
# TODO: Set to False after testing workflow changes
_TEST_MODE = True

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

While the TODO and checklist item in the PR description are good reminders, relying on a manual step to revert _TEST_MODE to False is error-prone. If this change is accidentally merged, the nightly benchmarks would run with incorrect test settings.

A more robust approach is to control this via an environment variable. This way, the production code remains unchanged, and you can enable test mode specifically in your CI workflow for validation runs by setting the environment variable.

Suggested change
# TODO: Set to False after testing workflow changes
_TEST_MODE = True
_TEST_MODE = os.environ.get("E2E_BENCHMARK_TEST_MODE", "False").lower() in ("true", "1")

What changed:
- .github/workflows/nightly-benchmark.yml: Restore schedule trigger, remove
  PR trigger that was added for testing
- e2e_test/benchmarks/test_nightly_perf.py: Set _TEST_MODE back to False
  to run full benchmarks

Why:
- Workflow testing is complete, reverting to production configuration
@github-actions github-actions Bot removed the tests Test changes label Feb 5, 2026
@slin1237
slin1237 merged commit d68a7f9 into main Feb 5, 2026
4 of 5 checks passed
@slin1237
slin1237 deleted the nightly-benchmark branch February 5, 2026 06:27
ppraneth pushed a commit that referenced this pull request Feb 18, 2026
Signed-off-by: ppraneth <pranethparuchuri@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants