Skip to content

e2e: canonical HF model IDs and nightly benchmark runner split - #373

Merged
slin1237 merged 7 commits into
mainfrom
nightly-benchmark-fix
Feb 8, 2026
Merged

slin1237 merged 7 commits into
mainfrom
nightly-benchmark-fix

Conversation

@slin1237

@slin1237 slin1237 commented Feb 8, 2026 •

Copy link
Copy Markdown
Member

Summary

  • standardize E2E model IDs to canonical Hugging Face format (org/model)
  • remove short-ID aliasing in model spec lookup and use canonical IDs across fixtures/tests
  • update nightly benchmark matrix IDs to canonical HF IDs
  • keep DeepSeek on H200 and route other nightly benchmarks to k8s H100 runners
  • add slash-safe handling for nightly outputs/artifacts (model slug + safe folder names)

Additional fixes

  • address pre-commit findings (codespell/ruff/ruff format)
  • preserve --reasoning-parser=gpt-oss where parser type is required

Validation

  • check for broken symlinks............................(no files to check)Skipped
  • detect destroyed symlinks................................................Passed
  • trim trailing whitespace.................................................Passed
  • fix end of files.........................................................Passed
  • check yaml...............................................................Passed
  • check toml...............................................................Passed
  • check for added large files..............................................Passed
  • check for merge conflicts................................................Passed
  • check that scripts with shebangs are executable..........................Passed
  • detect private key.......................................................Passed
  • mixed line ending........................................................Passed
  • don't commit to branch...................................................Passed
  • codespell................................................................Passed
  • rustfmt..................................................................Passed
  • clippy...................................................................Passed
  • ruff (legacy alias)......................................................Passed
  • ruff format..............................................................Passed

Summary by CodeRabbit

  • Tests

    • Updated test model identifiers to standardized path-style model names across suites, changing which models are exercised and test targets.
  • Chores

    • Enhanced nightly benchmark pipeline: improved caching, visibility on cache misses, increased parallelism, dynamic GPU runner selection, per-model backend setup, and slug-based artifact naming.
    • Standardized model specs and added per-model tensor-parallelism overrides via environment configuration.
  • Documentation

    • Minor doc/example and style comment updates to reflect new model identifier formats.

@github-actions github-actions Bot added documentation Improvements or additions to documentation ci CI/CD configuration changes tests Test changes multimodal Multimodal crate changes labels Feb 8, 2026
@coderabbitai

coderabbitai Bot commented Feb 8, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.
📝 Walkthrough

Walkthrough

This PR replaces short model IDs with path-based identifiers across e2e tests and infra, updates MODEL_SPECS and defaults, adds E2E_MODEL_TP_OVERRIDES, and refactors the GitHub Actions nightly-benchmark workflow to a per-model matrix with dynamic runners, caching, and per-model backend setup flags.

Changes

Cohort / File(s) Summary
CI Workflow Refactoring
.github/workflows/nightly-benchmark.yml
Reworked triggers to push/pull_request + workflow_dispatch; added RUSTC_WRAPPER/SCCACHE env; wheel caching with cache-miss conditional build and sccache stats; switched runners to k8s-runner-gpu / dynamic runs-on; expanded model matrix (slug, runs_on, gpu_type, setup_vllm/setup_trtllm, extra_deps); increased parallelism; slug-based artifact naming; per-model backend setup steps.
Model Specs & Defaults
e2e_test/infra/model_specs.py, e2e_test/infra/constants.py
Renamed MODEL_SPECS keys to path-based IDs and updated default constants; added json import and E2E_MODEL_TP_OVERRIDES parsing to apply per-model TP overrides; adjusted/extended some model entries (worker_args, vllm_args, etc.).
Fixtures & Hooks / Test infra
e2e_test/fixtures/hooks.py, e2e_test/fixtures/pool.py, e2e_test/conftest.py, e2e_test/infra/model_pool.py
Docstring/example updates to path-style model IDs; hooks now use get_model_spec(model_id) with KeyError → pytest.UsageError and read tp from spec; examples and docs updated to new IDs.
Test Marker Updates
e2e_test/chat_completions/*, e2e_test/bindings_go/*, e2e_test/responses/*, e2e_test/benchmarks/*, e2e_test/router/test_worker_api.py, e2e_test/responses/*
Replaced many pytest.mark.model strings from legacy short names to path-based IDs (e.g., llama-8b → meta-llama/Llama-3.1-8B-Instruct, qwen-7b → Qwen/Qwen2.5-7B-Instruct, gpt-oss → openai/gpt-oss-20b, etc.). Test target models updated without logic changes.
Nightly test harness
e2e_test/benchmarks/test_nightly_perf.py, e2e_test/benchmarks/nightly_summarize.py
Activated _TEST_MODE, adjusted experiment folder naming to use sanitized safe_model_id, updated nightly model list and metadata, and updated summarize docstrings to match new naming conventions.
Misc docs & scripts
docs/assets/stylesheets/extra.css, multimodal/scripts/generate_vision_golden.py
Minor comment/text fixes in CSS and reformatting of processor/from_pretrained calls; no behavioral changes.
Tests adjusting model usage
e2e_test/fixtures/pool.py, many e2e_test/* test files
Multiple tests and fixtures updated to reference new default model constant and new model identifiers; no logic changes other than marker/default updates.

Sequence Diagram(s)

sequenceDiagram
    participant GitHub as "GitHub Actions"
    participant Runner as "Runner (matrix runs-on)"
    participant Cache as "Wheel Cache / sccache"
    participant Setup as "Backend Setup (vLLM / TRT-LLM / SGLang)"
    participant Job as "Per-model Benchmark Job"
    participant Artifacts as "Artifact Storage"

    GitHub->>Runner: trigger (push / pull_request / workflow_dispatch)
    Runner->>Cache: check wheel cache
    alt cache miss (wheel)
        Runner->>Setup: install Rust, create Python venv, build wheel
        Setup->>Cache: upload wheel cache
        Setup->>Runner: show sccache stats
    end
    Runner->>Setup: conditional per-model backend setup (vLLM/TRT/SGLang) based on matrix flags
    Runner->>Job: run benchmark with MODEL_SPECS, GPU_TYPE, extra_deps, slug
    Job->>Artifacts: upload slug-named artifacts and results
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~50 minutes

Possibly related PRs

Suggested reviewers

  • key4ng
  • XinyueZhang369
  • CatherineSue

Poem

🐰 I hopped through specs and CI’s night,

Slugs and paths now guide each test’s light,
Runners hum and caches keep,
Backends wake for models deep,
A carrot clap — the pipelines bright! 🥕

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely summarizes the main changes: migrating to canonical Hugging Face model IDs and restructuring the nightly benchmark runner assignments.
Docstring Coverage ✅ Passed Docstring coverage is 95.83% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch nightly-benchmark-fix

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @slin1237, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request focuses on a significant infrastructure refinement by standardizing model identification within the End-to-End testing and nightly benchmarking systems. By adopting canonical Hugging Face model IDs, the changes enhance the robustness and maintainability of model referencing and test configurations. The update also includes improvements to nightly benchmark execution, such as optimized runner assignments and safer handling of output directory names, contributing to a more reliable and consistent development environment.

Highlights

  • Standardized Model Identifiers: All End-to-End (E2E) model identifiers across the codebase have been standardized to use canonical Hugging Face model IDs, replacing previous short aliases for improved clarity and consistency.
  • Nightly Benchmark Refinements: Nightly benchmark configurations have been updated to reflect the new canonical Hugging Face model IDs. This includes adjustments to tensor parallelism (tp) values for several models and a revised routing strategy, with DeepSeek remaining on H200 runners and other benchmarks directed to k8s H100 runners.
  • Filesystem-Safe Output Naming: A new mechanism has been implemented to ensure filesystem-safe folder names for nightly benchmark outputs, specifically by replacing slashes in canonical model IDs to prevent path issues.
  • Code Quality Improvements: Various code quality fixes were applied based on pre-commit hook findings, including codespell, ruff, and formatting adjustments. Additionally, logic was ensured to preserve parser types where required.
  • Test Fixture and Configuration Updates: Numerous test files, fixtures, and constants (e.g., DEFAULT_MODEL, MODEL_SPECS) were updated to align with the new canonical Hugging Face model ID naming convention, ensuring all E2E tests correctly reference models.
Changelog
  • docs/assets/stylesheets/extra.css
    • Corrected a minor typo from "re-declared" to "redeclared" in CSS comments.
  • e2e_test/benchmarks/nightly_summarize.py
    • Updated documentation comments for folder name parsing to reflect canonical Hugging Face model IDs.
    • Modified _TEST_MODE to True for testing purposes.
    • Implemented a mechanism to create filesystem-safe folder names for nightly outputs by replacing slashes in model IDs.
    • The _NIGHTLY_MODELS list was updated with canonical Hugging Face model IDs and adjusted tensor parallelism (tp) values for several models.
  • e2e_test/benchmarks/test_go_bindings_perf.py
    • Updated the @pytest.mark.model decorator to use the canonical Hugging Face ID meta-llama/Llama-3.2-1B-Instruct instead of llama-1b.
  • e2e_test/bindings_go/test_go_oai_server.py
    • Replaced all instances of llama-1b with meta-llama/Llama-3.2-1B-Instruct in @pytest.mark.model decorators and associated comments.
  • e2e_test/chat_completions/test_enable_thinking.py
    • Updated the @pytest.mark.model decorator from qwen-30b to Qwen/Qwen3-30B-A3B.
  • e2e_test/chat_completions/test_function_calling.py
    • Updated @pytest.mark.model decorators to use canonical Hugging Face IDs for llama-1b, llama-8b, qwen-7b, and mistral-7b.
  • e2e_test/chat_completions/test_openai_server.py
    • Updated @pytest.mark.model decorators for llama-8b and gpt-oss to their respective canonical Hugging Face IDs.
  • e2e_test/chat_completions/test_reasoning_content.py
    • Updated the @pytest.mark.model decorator from deepseek-7b to deepseek-ai/DeepSeek-R1-Distill-Qwen-7B.
  • e2e_test/chat_completions/test_validation.py
    • Updated @pytest.mark.model decorators from llama-8b to meta-llama/Llama-3.1-8B-Instruct.
  • e2e_test/conftest.py
    • Updated documentation and examples for @pytest.mark.model and E2E_MODELS environment variable to reflect the use of canonical Hugging Face model IDs.
  • e2e_test/fixtures/pool.py
    • Updated documentation and examples for E2E_MODELS, @pytest.mark.model, model_client, and model_base_url to use canonical Hugging Face model IDs.
  • e2e_test/infra/constants.py
    • Changed the DEFAULT_MODEL constant from llama-8b to meta-llama/Llama-3.1-8B-Instruct.
  • e2e_test/infra/model_pool.py
    • Updated comments and examples for WorkerIdentity and ModelInstance keys, and ModelPool usage to consistently refer to models by their canonical Hugging Face IDs.
  • e2e_test/infra/model_specs.py
    • Refactored the MODEL_SPECS dictionary to use canonical Hugging Face model IDs as keys, replacing previous short aliases.
    • All associated DEFAULT_MODEL_PATH constants were also updated.
  • e2e_test/responses/test_builtin_tools.py
    • Updated model_pool.get and @pytest.mark.model for gpt-oss to openai/gpt-oss-20b.
  • e2e_test/responses/test_state_management.py
    • Updated @pytest.mark.model decorators for qwen-14b and gpt-oss to their canonical Hugging Face IDs.
  • e2e_test/responses/test_streaming_events.py
    • Updated @pytest.mark.model decorators for qwen-14b and gpt-oss to their canonical Hugging Face IDs.
  • e2e_test/responses/test_structured_output.py
    • Updated @pytest.mark.model decorators for gpt-oss and qwen-14b to their canonical Hugging Face IDs.
  • e2e_test/responses/test_tools_call.py
    • Updated @pytest.mark.model decorators for gpt-oss and qwen-14b to their canonical Hugging Face IDs.
  • e2e_test/router/test_worker_api.py
    • Updated model_pool.get calls from llama-8b to meta-llama/Llama-3.1-8B-Instruct.
  • multimodal/scripts/generate_vision_golden.py
    • Applied minor formatting changes by removing unnecessary line breaks in from_pretrained calls and parser.add_argument.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/nightly-benchmark.yml
Activity
  • Pre-commit checks for broken symlinks, destroyed symlinks, trailing whitespace, end of files, yaml, toml, added large files, merge conflicts, executable scripts, private keys, mixed line endings, branch commits, codespell, rustfmt, clippy, ruff (legacy alias), and ruff format have all passed.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request effectively standardizes the end-to-end test suite to use canonical Hugging Face model IDs, removing the old short-name aliases. The changes are comprehensive, touching test files, fixtures, and configuration to ensure consistency. I've found two issues: a critical syntax error in e2e_test/infra/model_specs.py due to a misplaced bracket, and a high-severity issue in e2e_test/benchmarks/test_nightly_perf.py where a debug flag _TEST_MODE seems to have been left enabled, which would prevent the nightly benchmarks from running correctly. Once these are addressed, the PR will be in great shape.

Comment thread e2e_test/infra/model_specs.py
Comment thread e2e_test/benchmarks/test_nightly_perf.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Fix all issues with AI agents
In `@e2e_test/benchmarks/nightly_summarize.py`:
- Line 81: The inline comment describing the newest folder format in
nightly_summarize.py uses an invalid slash; update that comment to use the
actual double-underscore separator (`__`) used in folder names. Replace the
example "meta-llama/Llama-3.1-8B-Instruct_grpc_sglang_single" with a correctly
separated example such as "meta-llama__Llama-3.1-8B-Instruct_grpc_sglang_single"
so the comment matches the real `model__protocol_runtime_worker_type`
convention.
- Around line 64-75: The docstring entries in nightly_summarize.py list example
folder names with "/" separators (e.g.,
nightly_meta-llama/Llama-3.1-8B-Instruct_http_sglang_single) but the actual test
naming uses safe_model_id = model_id.replace("/", "__") in test_nightly_perf.py,
so update the docstring examples to use the filesystem-safe "__" separator
(e.g., nightly_meta-llama__Llama-3.1-8B-Instruct_http_sglang_single) or
alternatively adjust the parsing logic to accept both formats; reference the
docstring examples in nightly_summarize.py and the safe_model_id replacement in
test_nightly_perf.py when making the change.

In `@e2e_test/benchmarks/test_nightly_perf.py`:
- Line 32: The test flag _TEST_MODE in test_nightly_perf.py is currently True
which forces the suite into a minimal quick-run (single scenario, single
concurrency, low request count and short timeout); change _TEST_MODE = True to
_TEST_MODE = False so the nightly benchmark uses the full default scenarios,
concurrency list, max request counts and the 3-hour timeout; update the constant
in the file (symbol: _TEST_MODE) and ensure the change is committed before
merging so the comprehensive nightly run is executed.

Comment thread e2e_test/benchmarks/nightly_summarize.py
Comment thread e2e_test/benchmarks/nightly_summarize.py
Comment thread e2e_test/benchmarks/test_nightly_perf.py Outdated
@slin1237
slin1237 force-pushed the nightly-benchmark-fix branch from 4678025 to 9f4c7ea Compare February 8, 2026 08:12

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@e2e_test/benchmarks/test_nightly_perf.py`:
- Around line 57-59: The current slugging replaces "/" with "__" which is lossy
and can cause collisions; change the construction of safe_model_id so it uses a
reversible encoding (e.g., URL-quote or base64) of model_id rather than simple
replacement, then use that encoded string when building experiment_folder while
still storing the original model_id in metadata; update the code that defines
safe_model_id and experiment_folder (referencing safe_model_id, model_id, and
experiment_folder) to perform reversible encoding/decoding to guarantee
uniqueness and avoid overwrites.

Comment on lines +57 to +59
# Keep folder names filesystem-safe while retaining the canonical HF model id in metadata.
safe_model_id = model_id.replace("/", "__")
experiment_folder = f"nightly_{safe_model_id}_{backend}_{runtime}_{worker_type}"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Avoid artifact collisions from lossy model-id slugging.

Replacing “/” with “” is not reversible; model IDs containing “” can collide and overwrite artifacts. Use a reversible encoding (e.g., URL-quote or base64) to preserve uniqueness.

🔧 Proposed fix (reversible encoding)
-    safe_model_id = model_id.replace("/", "__")
+    safe_model_id = quote(model_id, safe="")
+from urllib.parse import quote
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
# Keep folder names filesystem-safe while retaining the canonical HF model id in metadata.
safe_model_id = model_id.replace("/", "__")
experiment_folder = f"nightly_{safe_model_id}_{backend}_{runtime}_{worker_type}"
# Keep folder names filesystem-safe while retaining the canonical HF model id in metadata.
safe_model_id = quote(model_id, safe="")
experiment_folder = f"nightly_{safe_model_id}_{backend}_{runtime}_{worker_type}"
🤖 Prompt for AI Agents
In `@e2e_test/benchmarks/test_nightly_perf.py` around lines 57 - 59, The current
slugging replaces "/" with "__" which is lossy and can cause collisions; change
the construction of safe_model_id so it uses a reversible encoding (e.g.,
URL-quote or base64) of model_id rather than simple replacement, then use that
encoded string when building experiment_folder while still storing the original
model_id in metadata; update the code that defines safe_model_id and
experiment_folder (referencing safe_model_id, model_id, and experiment_folder)
to perform reversible encoding/decoding to guarantee uniqueness and avoid
overwrites.

@slin1237
slin1237 force-pushed the nightly-benchmark-fix branch from 9f4c7ea to d7e302f Compare February 8, 2026 08:57

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In @.github/workflows/nightly-benchmark.yml:
- Line 219: The commented matrix entry for id
"meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8" is missing required fields
and will break the workflow if re-enabled; update the commented line to include
slug, runs_on, and gpu_type to match other entries (e.g., add slug:
meta-llama-Llama-4-Maverick-17B-128E-Instruct-FP8, runs_on: '["8-gpu-h200"]',
gpu_type: H200) while keeping test_class: TestNightlyLlama4MaverickMulti so the
full commented entry mirrors the existing matrix format.
🧹 Nitpick comments (2)
.github/workflows/nightly-benchmark.yml (1)

167-168: Consider extracting duplicated E2E_MODEL_TP_OVERRIDES to workflow-level env.

This JSON string is duplicated in both single-worker (line 168) and multi-worker (line 281) jobs. If the model list changes, both need updating. Move to the top-level env: block to maintain a single source of truth.

♻️ Suggested refactor
 env:
   RUSTC_WRAPPER: sccache
   SCCACHE_GHA_ENABLED: "true"
+  E2E_MODEL_TP_OVERRIDES: '{"meta-llama/Llama-3.1-8B-Instruct":1,"meta-llama/Llama-3.2-1B-Instruct":1,"Qwen/Qwen2.5-7B-Instruct":1,"Qwen/Qwen2.5-14B-Instruct":1,"deepseek-ai/DeepSeek-R1-Distill-Qwen-7B":1,"Qwen/Qwen3-30B-A3B":1,"mistralai/Mistral-7B-Instruct-v0.3":1,"openai/gpt-oss-20b":1}'

Then reference via ${{ env.E2E_MODEL_TP_OVERRIDES }} in both jobs.

e2e_test/infra/model_specs.py (1)

136-138: Consider logging a warning on malformed JSON instead of silent pass.

Silently ignoring JSONDecodeError could hide CI misconfigurations. A warning log would help operators diagnose issues while still falling back safely.

♻️ Suggested improvement
+import logging
+
+_logger = logging.getLogger(__name__)
+
 ...
         except json.JSONDecodeError:
-            # Ignore malformed override config and fall back to canonical specs.
-            pass
+            _logger.warning(
+                "E2E_MODEL_TP_OVERRIDES contains malformed JSON; using canonical specs"
+            )

@slin1237
slin1237 force-pushed the nightly-benchmark-fix branch from d7e302f to 3b5d7f3 Compare February 8, 2026 09:03

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
.github/workflows/nightly-benchmark.yml (1)

114-123: ⚠️ Potential issue | 🟡 Minor

Normalize whitespace in models input before exact-match filtering.

Comma-separated inputs often include spaces (e.g., "a, b"), which will cause the exact grep -Fq to miss matches. Strip whitespace before matching.

Proposed fix
-          MODELS="${{ github.event.inputs.models || 'all' }}"
+          MODELS_RAW="${{ github.event.inputs.models || 'all' }}"
+          MODELS="$(echo "$MODELS_RAW" | tr -d '[:space:]')"
           RUNTIME="${{ github.event.inputs.runtime || 'all' }}"

Also applies to: 224-233

🧹 Nitpick comments (1)
e2e_test/infra/model_specs.py (1)

127-139: Consider logging when TP override JSON is malformed.

The validation logic is thorough, but silently ignoring malformed JSON may make CI debugging harder if someone misconfigures E2E_MODEL_TP_OVERRIDES.

💡 Optional: Add debug logging for invalid overrides
+import logging
+
+logger = logging.getLogger(__name__)
+
 def get_model_spec(model_id: str) -> dict:
     """Get spec for a specific model, raising KeyError if not found."""
     if model_id not in MODEL_SPECS:
         raise KeyError(f"Unknown model: {model_id}. Available: {list(MODEL_SPECS.keys())}")
     spec = dict(MODEL_SPECS[model_id])
     tp_overrides_json = os.environ.get("E2E_MODEL_TP_OVERRIDES")
     if tp_overrides_json:
         try:
             tp_overrides = json.loads(tp_overrides_json)
             if isinstance(tp_overrides, dict):
                 override = tp_overrides.get(model_id)
                 if isinstance(override, int) and override > 0:
                     spec["tp"] = override
+                elif override is not None:
+                    logger.debug("Ignoring invalid TP override for %s: %r", model_id, override)
+            else:
+                logger.debug("E2E_MODEL_TP_OVERRIDES is not a dict: %r", tp_overrides)
         except json.JSONDecodeError:
-            # Ignore malformed override config and fall back to canonical specs.
-            pass
+            logger.debug("Malformed E2E_MODEL_TP_OVERRIDES JSON, using canonical specs")
     return spec

@slin1237
slin1237 force-pushed the nightly-benchmark-fix branch from 3b5d7f3 to c87aa0d Compare February 8, 2026 17:57
@slin1237
slin1237 force-pushed the nightly-benchmark-fix branch from 90d81af to ff6847c Compare February 8, 2026 18:29

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In @.github/workflows/nightly-benchmark.yml:
- Around line 42-44: actionlint flags the custom runner label "k8s-runner-gpu"
used in the build-wheel job as unknown; add that label to the actionlint
configuration (e.g., actionlint.yaml) under allowed runner labels so actionlint
recognizes "k8s-runner-gpu" (and any other custom labels) and stops reporting
false failures for the build-wheel job.

In `@e2e_test/fixtures/hooks.py`:
- Around line 108-116: The calculate_test_gpus function currently swallows
missing model IDs by returning 0 when get_model_spec raises KeyError; change
this to surface a hard failure by re-raising or raising a clear exception (e.g.,
raise KeyError or ValueError with a message like "Unknown model_id: {model_id}")
so mis-typed model markers fail fast; update the except block in
calculate_test_gpus to log/raise the descriptive exception instead of returning
0 so callers immediately see configuration errors.

Comment thread .github/workflows/nightly-benchmark.yml
Comment thread e2e_test/fixtures/hooks.py
- Move DeepSeek-R1-Distill-Qwen-7B from 8-gpu-h200 to k8s-runner-gpu/4-gpu-h100
  for both single and multi worker jobs. DeepSeek has tp=1 and doesn't need H200.
  Only Llama-4-Maverick (tp=8) remains on 8-gpu-h200.
- Switch ci_setup_python_venv.sh from system python3 to uv-managed Python 3.12.
  The 8-gpu-h200 runners have Python 3.10 which doesn't meet smg's >=3.12
  requirement. Using uv to manage Python avoids depending on system packages.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In @.github/workflows/nightly-benchmark.yml:
- Around line 83-85: Update the misleading comment in the nightly-benchmark
workflow to match the matrix: change the note that currently says "DeepSeek runs
on H200; all other models run on k8s H100" to reflect that DeepSeek is on H100
and Llama-4 is on H200 (the same change should be applied to the repeated
comment at the later block around the other occurrence). Locate the comment near
the matrix entries for "DeepSeek" and "Llama-4" and edit the text so it
accurately states "DeepSeek runs on H100; Llama-4 runs on H200; all other models
run on k8s H100" (or equivalent phrasing matching the matrix).

In `@scripts/ci_setup_python_venv.sh`:
- Around line 8-12: The script currently pipes the remote uv installer via curl
| sh (the curl -LsSf https://astral.sh/uv/install.sh invocation and subsequent
export PATH), which is a supply-chain risk; update the call in the uv installer
block so it either downloads a pinned installer URL (e.g., use a specific
release path rather than /install.sh) or downloads the installer to a temporary
file, verifies its checksum/signature, and only then executes it (keep the same
PATH export). Apply the same pattern to the other scripts that call the
installer (the ci_install_vllm.sh and ci_install_sglang.sh invocations) so they
use pinned version URLs or checksum-verified downloads instead of piping an
unverified remote script to sh.

Comment thread .github/workflows/nightly-benchmark.yml Outdated
Comment thread scripts/ci_setup_python_venv.sh Outdated
@slin1237
slin1237 force-pushed the nightly-benchmark-fix branch from be3de81 to 5a1563f Compare February 8, 2026 19:09

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In @.github/workflows/nightly-benchmark.yml:
- Around line 115-121: Normalize whitespace in the MODELS input before doing the
grep match: replace occurrences of surrounding whitespace around commas and trim
leading/trailing spaces so "modelA, modelB" becomes "modelA,modelB" and then use
that normalized variable in the existing SKIP check (the line using MODELS and
the grep with ,${{ matrix.model.id }},). Apply the same normalization in both
places the filter is implemented (the single-worker and multi-worker filter
blocks) and also mirror the change for the analogous block that uses the same
pattern later in the file.

Comment thread .github/workflows/nightly-benchmark.yml
uv venv doesn't seed pip by default, breaking downstream pip install
steps. Revert to system python3 -m venv until H200 runners are updated.
@slin1237
slin1237 force-pushed the nightly-benchmark-fix branch from 5a1563f to 1174dba Compare February 8, 2026 22:13

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In @.github/workflows/nightly-benchmark.yml:
- Line 218: Update the commented matrix entry for id
"meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8" to include the missing
fields so downstream parsing (fromJson(matrix.model.runs_on)) and artifact
naming work: add a slug (e.g.,
meta-llama-Llama-4-Maverick-17B-128E-Instruct-FP8), a runs_on JSON string (e.g.,
'["8-gpu-h200"]'), and gpu_type (e.g., H200) to the commented YAML line so it
becomes a complete matrix entry if re-enabled.
🧹 Nitpick comments (1)
.github/workflows/nightly-benchmark.yml (1)

167-167: Consider extracting E2E_MODEL_TP_OVERRIDES to reduce duplication.

The same JSON string appears in both single-worker (line 167) and multi-worker (line 279). If TP values change, both must be updated. Consider defining this in the workflow-level env section or a reusable composite action.

♻️ Example extraction to workflow env
env:
  RUSTC_WRAPPER: sccache
  SCCACHE_GHA_ENABLED: "true"
  E2E_MODEL_TP_OVERRIDES: '{"meta-llama/Llama-3.1-8B-Instruct":1,...}'

Then reference as ${{ env.E2E_MODEL_TP_OVERRIDES }} in job steps.

Comment thread .github/workflows/nightly-benchmark.yml Outdated
- { id: Qwen/Qwen3-30B-A3B, slug: Qwen-Qwen3-30B-A3B, test_class: TestNightlyQwen30bMulti, runs_on: '["k8s-runner-gpu","4-gpu-h100"]', gpu_type: H100 }
- { id: mistralai/Mistral-7B-Instruct-v0.3, slug: mistralai-Mistral-7B-Instruct-v0.3, test_class: TestNightlyMistral7bMulti, runs_on: '["k8s-runner-gpu","4-gpu-h100"]', gpu_type: H100 }
- { id: openai/gpt-oss-20b, slug: openai-gpt-oss-20b, test_class: TestNightlyGptOssMulti, runs_on: '["k8s-runner-gpu","4-gpu-h100"]', gpu_type: H100 }
# - { id: meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8, test_class: TestNightlyLlama4MaverickMulti } # tp=8, keep disabled for nightly

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Commented matrix entry still incomplete.

The past review flagged that this commented entry is missing slug, runs_on, and gpu_type fields. While marked as addressed, the current code still shows the incomplete format. If re-enabled without these fields, fromJson(matrix.model.runs_on) and artifact naming will fail.

# - { id: meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8, slug: meta-llama-Llama-4-Maverick-17B-128E-Instruct-FP8, test_class: TestNightlyLlama4MaverickMulti, runs_on: '["8-gpu-h200"]', gpu_type: H200 }
🤖 Prompt for AI Agents
In @.github/workflows/nightly-benchmark.yml at line 218, Update the commented
matrix entry for id "meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8" to
include the missing fields so downstream parsing
(fromJson(matrix.model.runs_on)) and artifact naming work: add a slug (e.g.,
meta-llama-Llama-4-Maverick-17B-128E-Instruct-FP8), a runs_on JSON string (e.g.,
'["8-gpu-h200"]'), and gpu_type (e.g., H200) to the commented YAML line so it
becomes a complete matrix entry if re-enabled.

- Extract H200 models (Llama-4-Maverick) into separate single-worker-h200
  job so H200 pool runs independently from H100 max-parallel limits.
  H200 runners were idle waiting behind H100 jobs in the shared queue.
- H100 jobs: runs-on moved to job level, removed per-model runs_on/gpu_type
  from matrix entries since all H100 models share the same runner config.
- H200 jobs: include actions/setup-python@v6 for Python 3.12 since H200
  runners have Python 3.10 which doesn't meet smg's >=3.12 requirement.
- multi-worker-h200 commented out (Maverick multi already disabled).
- Update summarize-benchmarks to depend on single-worker-h200.
- Fix docstring in test_nightly_perf.py to reflect multi-runner setup.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
.github/workflows/nightly-benchmark.yml (1)

113-122: ⚠️ Potential issue | 🟡 Minor

Trim whitespace from the models input before matching.

The filter logic doesn't handle spaces in comma-separated input. If a user enters "modelA, modelB" (with spaces), the grep -Fq match will fail because it looks for , modelB instead of modelB.

🔧 Proposed fix
         run: |
           MODELS="${{ github.event.inputs.models || 'all' }}"
+          MODELS="$(echo "$MODELS" | tr -d '[:space:]')"
           RUNTIME="${{ github.event.inputs.runtime || 'all' }}"
           SKIP="false"
           if [ "$MODELS" != "all" ] && ! echo ",$MODELS," | grep -Fq ",${{ matrix.model.id }},"; then

Apply the same fix to the filter blocks in multi-worker (line 217), single-worker-h200 (line 315).

🧹 Nitpick comments (1)
.github/workflows/nightly-benchmark.yml (1)

158-159: Consider extracting E2E_MODEL_TP_OVERRIDES to a reusable location.

This JSON is duplicated at line 263 (multi-worker job). If model TP requirements change, both locations must be updated manually, risking drift.

Options:

  • Define as a workflow-level environment variable
  • Use a composite action or reusable workflow
  • Store in a JSON file and read it in the step

- Set _TEST_MODE = False to run full benchmarks instead of reduced test runs.
- Enable daily cron schedule at midnight (was disabled for PR testing).
- Remove pull_request trigger since nightly benchmarks shouldn't run on PRs.
@slin1237
slin1237 merged commit f287847 into main Feb 8, 2026
6 checks passed
@slin1237
slin1237 deleted the nightly-benchmark-fix branch February 8, 2026 23:42
ppraneth pushed a commit that referenced this pull request Feb 18, 2026
Signed-off-by: ppraneth <pranethparuchuri@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes documentation Improvements or additions to documentation multimodal Multimodal crate changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant