Skip to content

ci(benchmarks): add Devstral 2 to H200 nightly benchmarks - #1068

Closed
smfirmin wants to merge 1 commit into
smg-project:mainfrom
smfirmin:devstral-2-nightly-benchmark
Closed

smfirmin wants to merge 1 commit into
smg-project:mainfrom
smfirmin:devstral-2-nightly-benchmark

Conversation

@smfirmin

@smfirmin smfirmin commented Apr 8, 2026 •

Copy link
Copy Markdown
Contributor

Description

Summary

Add mistralai/Devstral-2-123B-Instruct-2512 to the nightly benchmark coverage on H200 for both sglang and vllm.

Changes

  • add Devstral 2 to MODEL_SPECS with nightly benchmark metadata
  • set Devstral 2 tensor parallelism to 4
  • register nightly benchmark test classes for Devstral 2
  • add Devstral 2 to the single-worker-h200 workflow matrix
  • place the workflow entry immediately after the commented-out minimax-m2 benchmark entry

Notes

  • this enables nightly H200 runs for both sglang and vllm

Summary by CodeRabbit

  • Chores
    • Integrated Devstral-2-123B-Instruct model into the nightly performance benchmarking suite, adding H200 hardware workflow configuration, model specifications, and test infrastructure to evaluate performance across single and multi-worker configurations using HTTP and gRPC protocols.

@github-actions github-actions Bot added ci CI/CD configuration changes tests Test changes labels Apr 8, 2026
@coderabbitai

coderabbitai Bot commented Apr 8, 2026 •

Copy link
Copy Markdown

Caution

Review failed

Pull request was closed or merged during review

📝 Walkthrough

Walkthrough

The PR adds the mistralai/Devstral-2-123B-Instruct-2512 model to the nightly benchmark testing infrastructure by registering it in the CI workflow matrix, test configuration, and model specifications. This enables automated performance testing for this model on H200 single-worker resources.

Changes

Cohort / File(s) Summary
Nightly Benchmark CI Workflow
.github/workflows/nightly-benchmark.yml
Added new model entry to single-worker-h200 job matrix with corresponding slug and test class TestNightlyDevstral2Single.
Benchmark Test Configuration
e2e_test/benchmarks/test_nightly_perf.py
Registered mistralai/Devstral-2-123B-Instruct-2512 in _NIGHTLY_MODELS list with single-worker configuration and http/grpc backend support.
Model Specifications
e2e_test/infra/model_specs.py
Added Devstral model specifications including tensor parallelism (tp=4), features (chat, streaming, function_calling, reasoning), startup timeout, and trust-remote-code flags.

Possibly related PRs

Suggested labels

ci, tests

Suggested reviewers

  • key4ng
  • CatherineSue
  • slin1237
  • XinyueZhang369

Poem

🐰 A model swift, Devstral by name,
Now benchmarks will test its flame,
Through H200 it speeds along,
Added in three places strong,
Performance metrics shall not fail! ⚡


🎯 2 (Simple) | ⏱️ ~10 minutes

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically identifies the main change: adding Devstral 2 model to H200 nightly benchmarks, which aligns with all modifications across three files (workflow config, test file, and model specs).
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@mergify

mergify Bot commented Apr 8, 2026

Copy link
Copy Markdown
Contributor

Hi @smfirmin, the DCO sign-off check has failed. All commits must include a Signed-off-by line.

To fix existing commits:

# Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-lease

To sign off future commits automatically:

  • Use git commit -s every time, or
  • VSCode: enable Git: Always Sign Off in Settings
  • PyCharm: enable Sign-off commit in the Commit tool window

@smfirmin smfirmin closed this Apr 8, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds the mistralai/Devstral-2-123B-Instruct-2512 model to the nightly performance benchmarks and infrastructure specifications. Feedback highlights potential typos in the model ID and date suffix, suggests removing the reasoning feature to ensure consistency with other high-capability models, and notes that the minimax-m2 model was not commented out as expected.

("openai/gpt-oss-20b", "GptOss20b", 1, ["http", "grpc"], {}),
("minimaxai/minimax-m2", "MinimaxM2", 1, ["http", "grpc"], {}),
(
"mistralai/Devstral-2-123B-Instruct-2512",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The model ID mistralai/Devstral-2-123B-Instruct-2512 appears to contain typos. "Devstral" is not a known Mistral model series (likely "Mistral" or "Codestral"), and the date suffix "2512" (December 2025) is likely a typo for "2412" or "2411". An incorrect model ID will cause the benchmark to fail during model download. Additionally, the PR description mentions that minimax-m2 (line 104) should be commented out, but it remains active in the code.

"mistralai/Devstral-2-123B-Instruct-2512": {
"model": _resolve_model_path("mistralai/Devstral-2-123B-Instruct-2512"),
"tp": 4,
"features": ["chat", "streaming", "function_calling", "reasoning"],

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The reasoning feature is included here but is omitted for other similar high-capability models like Llama-3.3-70B-Instruct (line 171). In this codebase, the reasoning feature typically refers to models that support a specific reasoning output field (e.g., DeepSeek-R1 or Harmony-compatible models). Unless this model specifically supports such a field in its API response, this feature should be removed to maintain consistency and avoid triggering incompatible tests.

Suggested change
"features": ["chat", "streaming", "function_calling", "reasoning"],
"features": ["chat", "streaming", "function_calling"],

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant