Skip to content

feat(ci): add minimaxai/minimax-m2 to nightly benchmark - #795

Merged
key4ng merged 23 commits into
smg-project:mainfrom
smfirmin:sfirmin/add-minimax-to-nightly-bench
Mar 20, 2026
Merged

key4ng merged 23 commits into
smg-project:mainfrom
smfirmin:sfirmin/add-minimax-to-nightly-bench

Conversation

@smfirmin

@smfirmin smfirmin commented Mar 17, 2026 •

Copy link
Copy Markdown
Contributor

• ## Description

Problem

minimaxai/minimax-m2 needs to be benchmarked.

Solution

Add minimaxai/minimax-m2 to the nightly benchmark path end to end: define its E2E model spec, register nightly single/multi benchmark test classes, and include it in the nightly benchmark workflow matrix.

Changes

  • Added minimaxai/minimax-m2 to e2e_test/infra/model_specs.py.
  • Configured the model with tp=4 and --trust-remote-code for both worker and vLLM startup.
  • Added ("minimaxai/minimax-m2", "MinimaxM2", 2, ["http", "grpc"], {}) to the nightly benchmark model list in e2e_test/benchmarks/test_nightly_perf.py.
  • Added minimaxai/minimax-m2 to the single-worker matrix in .github/workflows/nightly-benchmark.yml.

Summary by CodeRabbit

  • Tests
    • Nightly benchmarks now include the minimax-m2 model in single- and multi-worker runs (HTTP and gRPC), validated for chat, streaming, function-calling, and reasoning; multi-worker parallelism set to 4.
  • Chores
    • Nightly benchmark workflow triggers only for benchmark test changes and the model matrix simplified to focus on minimax-m2.

Signed-off-by: Sydney Firmin <sydney.firmin@oracle.com>
@github-actions github-actions Bot added ci CI/CD configuration changes tests Test changes labels Mar 17, 2026
@coderabbitai

coderabbitai Bot commented Mar 17, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds minimaxai/minimax-m2 to nightly benchmarks: updates the GitHub Actions matrix and trigger, adds the model to the nightly test matrix, and registers a MODEL_SPECS entry with TP, features, and runtime args. CI now triggers on PRs touching e2e_test/benchmarks/**.

Changes

Cohort / File(s) Summary
CI workflow
.github/workflows/nightly-benchmark.yml
Added a pull_request trigger scoped to e2e_test/benchmarks/**; replaced several model matrix entries with minimaxai/minimax-m2 and updated slug/test_class values for single/multi-worker jobs.
Nightly benchmark tests
e2e_test/benchmarks/test_nightly_perf.py
Appended minimaxai/minimax-m2 to _NIGHTLY_MODELS, causing generation of Single (1 worker) and Multi (4 workers) nightly test classes parameterized for http and grpc.
Model specifications
e2e_test/infra/model_specs.py
Added MODEL_SPECS["minimaxai/minimax-m2"] with resolved model path, tp=4, features ["chat","streaming","function_calling","reasoning"], and worker_args/vllm_args containing --trust-remote-code.

Sequence Diagram(s)

sequenceDiagram
    participant Dev as Developer (PR)
    participant GH as GitHub Actions
    participant Tests as Nightly Tests
    participant Specs as MODEL_SPECS
    participant Runner as Test Runner / Workers

    Dev->>GH: Open PR touching e2e_test/benchmarks/**
    GH->>GH: Trigger nightly-benchmark workflow (pull_request)
    GH->>Tests: Select model matrix (minimaxai/minimax-m2)
    Tests->>Specs: Lookup MODEL_SPECS["minimaxai/minimax-m2"]
    Specs->>Runner: Provide model path, tp=4, worker_args/vllm_args
    Runner->>Tests: Execute http/grpc benchmark jobs (Single/Multi)
    Tests->>GH: Upload artifacts / report results
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested reviewers

  • key4ng
  • slin1237
  • XinyueZhang369

Poem

🐰 Hopping through CI at night,
A new M2 model joins the flight,
Matrices shift and specs take root,
Workers wake and tests compute,
Carrots crunch as pipelines light!

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately and concisely summarizes the main change: adding minimaxai/minimax-m2 to the nightly benchmark system.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request integrates the minimaxai/minimax-m2 model into the existing nightly benchmarking system. This ensures that the performance and stability of this new model are regularly monitored, providing crucial data for its ongoing development and deployment. The changes involve defining the model's operational parameters and registering it within the automated testing framework.

Highlights

  • Model Specification: Defined the minimaxai/minimax-m2 model's end-to-end specification, including tp=4 and --trust-remote-code for both worker and vLLM startup configurations.
  • Nightly Benchmark Inclusion: Added minimaxai/minimax-m2 to the list of models designated for nightly performance benchmarks.
  • Workflow Integration: Included minimaxai/minimax-m2 in the single-worker matrix of the nightly benchmark workflow, ensuring its regular evaluation.
Changelog
  • e2e_test/benchmarks/test_nightly_perf.py
    • Added ("minimaxai/minimax-m2", "MinimaxM2", 2, ["http", "grpc"], {}) to the NIGHTLY_MODELS list.
  • e2e_test/infra/model_specs.py
    • Introduced a new entry for minimaxai/minimax-m2 with its model path, tp=4, features=["chat", "streaming"], and worker_args/vllm_args set to ["--trust-remote-code"].
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/nightly-benchmark.yml
Activity
  • No review comments or other human activity have been recorded for this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds the minimaxai/minimax-m2 model to the nightly benchmark suite. The changes correctly configure the model specification and register it for benchmarking. My review includes one suggestion to update the model's feature list in e2e_test/infra/model_specs.py to more accurately reflect its capabilities, ensuring it's included in all relevant tests. This comment aligns with general best practices and does not contradict any specific rules.

Comment thread e2e_test/infra/model_specs.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/nightly-benchmark.yml:
- Line 124: Remove the spaces immediately inside the curly braces for the YAML
list item containing the minimaxai-minimax-m2 entry; update the line "- { id:
minimaxai/minimax-m2, slug: minimaxai-minimax-m2, test_class:
TestNightlyMinimaxM2Single }" to remove the space after "{" and before "}" (so
the entry becomes "- {id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2,
test_class: TestNightlyMinimaxM2Single}") to satisfy the YAMLlint `braces` rule
while keeping the same keys and values (identifiers: minimaxai/minimax-m2, slug
minimaxai-minimax-m2, test_class TestNightlyMinimaxM2Single).
- Line 124: The multi-worker matrix is missing the generated
TestNightlyMinimaxM2Multi entry, so add an entry with the same repo identifiers
but test_class: TestNightlyMinimaxM2Multi to the multi-worker matrix;
specifically add an item matching id: minimaxai/minimax-m2, slug:
minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Multi so the multi-worker
nightly path runs.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 74e1960f-0ea7-4ee7-956d-df40b7f55e82

📥 Commits

Reviewing files that changed from the base of the PR and between 976f1a6 and dfa8ad1.

📒 Files selected for processing (3)
  • .github/workflows/nightly-benchmark.yml
  • e2e_test/benchmarks/test_nightly_perf.py
  • e2e_test/infra/model_specs.py

Comment thread .github/workflows/nightly-benchmark.yml Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8941f815f2

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread .github/workflows/nightly-benchmark.yml Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 48fd9d607e

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread .github/workflows/nightly-benchmark.yml Outdated
- { id: Qwen/Qwen2.5-7B-Instruct, slug: Qwen-Qwen2.5-7B-Instruct, test_class: TestNightlyQwen7bMulti }
- { id: Qwen/Qwen3-30B-A3B, slug: Qwen-Qwen3-30B-A3B, test_class: TestNightlyQwen30bMulti }
- { id: openai/gpt-oss-20b, slug: openai-gpt-oss-20b, test_class: TestNightlyGptOss20bMulti }
- { id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Multi }

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Move Minimax multi-worker benchmark off 4-GPU runners

In the multi-worker workflow matrix this new entry schedules TestNightlyMinimaxM2Multi on runs-on: 4-gpu-h100, but the generated test class uses workers(count=2) from e2e_test/benchmarks/test_nightly_perf.py and the new model spec sets tp=4 in e2e_test/infra/model_specs.py; start_workers allocates GPUs sequentially by tp, so the second worker is assigned GPUs 4-7 (see e2e_test/infra/worker.py), which exceeds a 4-GPU host and causes that matrix leg to fail consistently instead of producing benchmark data.

Useful? React with 👍 / 👎.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@e2e_test/infra/model_specs.py`:
- Around line 87-88: Update the model mapping key/value to use the canonical
HuggingFace identifier: replace the string "minimaxai/minimax-m2" with
"MiniMaxAI/MiniMax-M2" where the mapping entry currently calls
_resolve_model_path (the dict entry containing "minimaxai/minimax-m2" and the
call to _resolve_model_path should be updated to
_resolve_model_path("MiniMaxAI/MiniMax-M2")). Ensure both the dictionary key (if
used) and the argument passed to _resolve_model_path are corrected to the exact
capitalization shown.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: a49a6bf2-407a-47e5-83e5-df0732ff5b56

📥 Commits

Reviewing files that changed from the base of the PR and between dfa8ad1 and 48fd9d6.

📒 Files selected for processing (2)
  • .github/workflows/nightly-benchmark.yml
  • e2e_test/infra/model_specs.py

Comment thread e2e_test/infra/model_specs.py
@smfirmin smfirmin changed the title add minimaxai/minimax-m2 to nightly benchmark (feat)add minimaxai/minimax-m2 to nightly benchmark Mar 18, 2026
@smfirmin smfirmin changed the title (feat)add minimaxai/minimax-m2 to nightly benchmark feat(ci) add minimaxai/minimax-m2 to nightly benchmark Mar 18, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: eeee7ac1b0

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread e2e_test/infra/model_specs.py
@smfirmin smfirmin changed the title feat(ci) add minimaxai/minimax-m2 to nightly benchmark feat(ci): add minimaxai/minimax-m2 to nightly benchmark Mar 18, 2026
@key4ng key4ng self-assigned this Mar 18, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 23b3a5537f

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread e2e_test/benchmarks/test_nightly_perf.py Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8f9b55b065

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread .github/workflows/nightly-benchmark.yml Outdated
Comment thread .github/workflows/nightly-benchmark.yml Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (1)
.github/workflows/nightly-benchmark.yml (1)

127-127: ⚠️ Potential issue | 🟡 Minor

Fix flow-map brace spacing to satisfy YAMLlint.

The inline maps on Line 127 and Line 241 have spaces immediately inside {} and fail the braces rule.

Proposed fix
-          - { id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Single }
+          - {id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Single}
...
-          - { id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Multi }
+          - {id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Multi}

Also applies to: 241-241

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/nightly-benchmark.yml at line 127, Fix the YAML linter
"braces" rule by removing the spaces inside the inline flow-maps: change
occurrences of "{ id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2,
test_class: TestNightlyMinimaxM2Single }" to use no spaces after "{" or before
"}" (e.g. "{id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2, test_class:
TestNightlyMinimaxM2Single}"), and apply the same removal of inner-brace spacing
for the duplicate inline map instance elsewhere in the file.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/nightly-benchmark.yml:
- Around line 16-18: The workflow has pull_request nested under
workflow_dispatch which is invalid; move the pull_request key out so both
workflow_dispatch and pull_request are siblings at the top level of the
job/triggers block. Edit the YAML so workflow_dispatch remains its own mapping
(only containing inputs if present) and add a separate top-level pull_request
mapping (e.g., pull_request: paths: - 'e2e_test/benchmarks/**') alongside
workflow_dispatch to register the PR trigger correctly.

---

Duplicate comments:
In @.github/workflows/nightly-benchmark.yml:
- Line 127: Fix the YAML linter "braces" rule by removing the spaces inside the
inline flow-maps: change occurrences of "{ id: minimaxai/minimax-m2, slug:
minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Single }" to use no spaces
after "{" or before "}" (e.g. "{id: minimaxai/minimax-m2, slug:
minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Single}"), and apply the
same removal of inner-brace spacing for the duplicate inline map instance
elsewhere in the file.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 1706bc0a-5c17-4679-92cc-cbd5c85b5729

📥 Commits

Reviewing files that changed from the base of the PR and between 48fd9d6 and 8f9b55b.

📒 Files selected for processing (1)
  • .github/workflows/nightly-benchmark.yml

Comment thread .github/workflows/nightly-benchmark.yml Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (1)
.github/workflows/nightly-benchmark.yml (1)

127-127: ⚠️ Potential issue | 🟡 Minor

Fix flow-map brace spacing to satisfy YAMLlint.

Both added matrix entries still use spaces inside {} and are flagged by lint.

Proposed fix
-          - { id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Single }
+          - {id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Single}
...
-          - { id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Multi }
+          - {id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Multi}

Also applies to: 241-241

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/nightly-benchmark.yml at line 127, YAML linter is failing
due to spaces inside flow-map braces; remove the inner spaces for the matrix
entries so they read without spaces inside the braces (e.g. change "{ id:
minimaxai/minimax-m2, slug: minimaxai-minimax-m2, test_class:
TestNightlyMinimaxM2Single }" to "{id: minimaxai/minimax-m2, slug:
minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Single}" and make the same
adjustment for the other similar entry), ensuring all flow-map brace spacing
complies with yamllint.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/nightly-benchmark.yml:
- Around line 16-18: Update the pull_request.paths for the nightly-benchmark
workflow so changes to benchmark wiring and this workflow itself trigger the
job: modify the pull_request.paths entry (pull_request.paths) to include
e2e_test/infra/model_specs.py and the workflow file (for example add
'e2e_test/infra/model_specs.py' and '.github/workflows/nightly-benchmark.yml' or
broaden to 'e2e_test/**') in addition to the existing 'e2e_test/benchmarks/**'
pattern so PRs that change model wiring or the workflow also run the benchmark
validation.

---

Duplicate comments:
In @.github/workflows/nightly-benchmark.yml:
- Line 127: YAML linter is failing due to spaces inside flow-map braces; remove
the inner spaces for the matrix entries so they read without spaces inside the
braces (e.g. change "{ id: minimaxai/minimax-m2, slug: minimaxai-minimax-m2,
test_class: TestNightlyMinimaxM2Single }" to "{id: minimaxai/minimax-m2, slug:
minimaxai-minimax-m2, test_class: TestNightlyMinimaxM2Single}" and make the same
adjustment for the other similar entry), ensuring all flow-map brace spacing
complies with yamllint.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 799e9d2a-ebb5-4f26-8c02-f240a82bb956

📥 Commits

Reviewing files that changed from the base of the PR and between 8f9b55b and fa49cfd.

📒 Files selected for processing (1)
  • .github/workflows/nightly-benchmark.yml

Comment thread .github/workflows/nightly-benchmark.yml Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: fa49cfdf96

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread .github/workflows/nightly-benchmark.yml Outdated
@key4ng

key4ng commented Mar 19, 2026

Copy link
Copy Markdown
Member

@key4ng

key4ng commented Mar 19, 2026

Copy link
Copy Markdown
Member

worker log

torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 576.00 MiB. GPU 2 has a total capacity of 79.18 GiB of which 193.69 MiB is free. Process 5029 has 55.63 GiB memory in use. Including non-PyTorch memory, this process has 23.34 GiB memory in use. Of the allocated memory 21.82 GiB is allocated by PyTorch, and 33.97 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation.  See documentation for Memory Management  (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)

Comment thread e2e_test/benchmarks/test_nightly_perf.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@e2e_test/benchmarks/test_nightly_perf.py`:
- Line 104: The MinimaxM2 benchmark entry is configured with multi_workers=4 for
a model with tp=4 which causes CUDA OOM at startup; locate the tuple for
("minimaxai/minimax-m2", "MinimaxM2", 4, ["http", "grpc"], {}) and reduce its
multi-worker count to a feasible value (e.g., set multi_workers from 4 to 2) or
disable multi-worker mode for "MinimaxM2" temporarily until runner
capacity/placement is adjusted so the nightly multi benchmark can start
reliably.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 19fb5c4b-d6bf-49dc-a1ee-c86d31cf597b

📥 Commits

Reviewing files that changed from the base of the PR and between fa49cfd and 94c3810.

📒 Files selected for processing (1)
  • e2e_test/benchmarks/test_nightly_perf.py

Comment thread e2e_test/benchmarks/test_nightly_perf.py Outdated
smfirmin and others added 7 commits March 19, 2026 15:35
Signed-off-by: Sydney Firmin <sydney.firmin@oracle.com>
Signed-off-by: Sydney Firmin <sydney.firmin@oracle.com>
Signed-off-by: Sydney Firmin <sydney.firmin@oracle.com>
Signed-off-by: Sydney Firmin <sydney.firmin@oracle.com>
Signed-off-by: Sydney Firmin <sydney.firmin@oracle.com>
Signed-off-by: Sydney Firmin <sydney.firmin@oracle.com>
@smfirmin
smfirmin force-pushed the sfirmin/add-minimax-to-nightly-bench branch from 5e59948 to 5ff333e Compare March 19, 2026 22:39

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5ff333e88e

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +145 to +147
_variants = [("Single", 1)]
if _multi_workers > 0:
_variants.append(("Multi", _multi_workers))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep a Minimax multi-worker class in nightly collection

Fresh evidence: this commit now sets Minimax’s worker count to 0, and the new guard here skips generating any TestNightlyMinimaxM2Multi class at collection time. That means nightly benchmarking can only exercise the single-worker path for this model, so routing/load-balancing regressions that appear only with multiple workers will never be measured even before the workflow matrix is considered.

Useful? React with 👍 / 👎.

Comment thread e2e_test/infra/worker.py Outdated
Signed-off-by: Sydney Firmin <sydney.firmin@oracle.com>
@smfirmin
smfirmin force-pushed the sfirmin/add-minimax-to-nightly-bench branch from 5ff333e to 7c2ed96 Compare March 19, 2026 22:48

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b922872bae

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread .github/workflows/nightly-benchmark.yml Outdated
name: "${{ matrix.model.id }} / single-${{ matrix.variant.id }}"
needs: build-wheel
if: ${{ !cancelled() }}
if: false

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Re-enable H100 benchmark jobs

Setting the single-worker job condition to if: false unconditionally skips that entire H100 matrix, and this commit makes the same change for multi-worker as well. As a result, nightly/scheduled/manual runs stop executing the H100 benchmark suites (Llama/Qwen/GPT-OSS paths), so regressions on those model/runtime combinations are no longer detected and their benchmark artifacts disappear.

Useful? React with 👍 / 👎.

@key4ng

key4ng commented Mar 19, 2026

Copy link
Copy Markdown
Member

Signed-off-by: key4ng <rukeyang@gmail.com>
@key4ng
key4ng merged commit 4ecd869 into smg-project:main Mar 20, 2026
63 of 65 checks passed
@smfirmin
smfirmin deleted the sfirmin/add-minimax-to-nightly-bench branch March 20, 2026 16:24
smfirmin added a commit to smfirmin/smg that referenced this pull request Mar 20, 2026
)

Signed-off-by: Sydney Firmin <sydney.firmin@oracle.com>
Signed-off-by: key4ng <rukeyang@gmail.com>
Co-authored-by: key4ng <rukeyang@gmail.com>
smfirmin added a commit to smfirmin/smg that referenced this pull request Mar 20, 2026
)

Signed-off-by: Sydney Firmin <sydney.firmin@oracle.com>
Signed-off-by: key4ng <rukeyang@gmail.com>
Co-authored-by: key4ng <rukeyang@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants