Repository navigation
Conversation
|
Caution Review failedPull request was closed or merged during review 📝 WalkthroughWalkthroughThe PR adds the mistralai/Devstral-2-123B-Instruct-2512 model to the nightly benchmark testing infrastructure by registering it in the CI workflow matrix, test configuration, and model specifications. This enables automated performance testing for this model on H200 single-worker resources. Changes
Possibly related PRs
Suggested labels
Suggested reviewers
Poem
🎯 2 (Simple) | ⏱️ ~10 minutes 🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Hi @smfirmin, the DCO sign-off check has failed. All commits must include a To fix existing commits: # Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-leaseTo sign off future commits automatically:
|
There was a problem hiding this comment.
Code Review
This pull request adds the mistralai/Devstral-2-123B-Instruct-2512 model to the nightly performance benchmarks and infrastructure specifications. Feedback highlights potential typos in the model ID and date suffix, suggests removing the reasoning feature to ensure consistency with other high-capability models, and notes that the minimax-m2 model was not commented out as expected.
| ("openai/gpt-oss-20b", "GptOss20b", 1, ["http", "grpc"], {}), | ||
| ("minimaxai/minimax-m2", "MinimaxM2", 1, ["http", "grpc"], {}), | ||
| ( | ||
| "mistralai/Devstral-2-123B-Instruct-2512", |
There was a problem hiding this comment.
The model ID mistralai/Devstral-2-123B-Instruct-2512 appears to contain typos. "Devstral" is not a known Mistral model series (likely "Mistral" or "Codestral"), and the date suffix "2512" (December 2025) is likely a typo for "2412" or "2411". An incorrect model ID will cause the benchmark to fail during model download. Additionally, the PR description mentions that minimax-m2 (line 104) should be commented out, but it remains active in the code.
| "mistralai/Devstral-2-123B-Instruct-2512": { | ||
| "model": _resolve_model_path("mistralai/Devstral-2-123B-Instruct-2512"), | ||
| "tp": 4, | ||
| "features": ["chat", "streaming", "function_calling", "reasoning"], |
There was a problem hiding this comment.
The reasoning feature is included here but is omitted for other similar high-capability models like Llama-3.3-70B-Instruct (line 171). In this codebase, the reasoning feature typically refers to models that support a specific reasoning output field (e.g., DeepSeek-R1 or Harmony-compatible models). Unless this model specifically supports such a field in its API response, this feature should be removed to maintain consistency and avoid triggering incompatible tests.
| "features": ["chat", "streaming", "function_calling", "reasoning"], | |
| "features": ["chat", "streaming", "function_calling"], |
Description
Summary
Add
mistralai/Devstral-2-123B-Instruct-2512to the nightly benchmark coverage on H200 for bothsglangandvllm.Changes
MODEL_SPECSwith nightly benchmark metadata4single-worker-h200workflow matrixminimax-m2benchmark entryNotes
sglangandvllmSummary by CodeRabbit