Repository navigation
chore(e2e): include model size in gpt-oss nightly benchmark slug - #384
Conversation
Rename GptOss → GptOss20b so the nightly job names include the parameter count, consistent with every other model (Llama8b, Qwen7b, Deepseek7b, etc.).
Summary of ChangesHello @CatherineSue, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request addresses an ambiguity in nightly benchmark naming by standardizing the slug for the Highlights
Changelog
Ignored Files
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
📝 WalkthroughWalkthroughThe PR renames test class identifiers for the OpenAI GPT-OSS 20B model across nightly benchmark CI workflow and test configuration, standardizing the naming from generic "GptOss" to explicit "GptOss20b" for both single-worker and multi-worker test classes. Changes
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~3 minutes Possibly related PRs
Suggested labels
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing touches
🧪 Generate unit tests (beta)
No actionable comments were generated in the recent review. 🎉 Comment |
There was a problem hiding this comment.
Code Review
The pull request effectively addresses the ambiguity in the gpt-oss model slug by explicitly including the model size, 20b, in its name. This change enhances clarity and consistency across the benchmark suite, aligning with the naming conventions of other models and preventing potential confusion with different gpt-oss variants. The update is straightforward and improves the maintainability of the benchmark configuration.
| ("deepseek-ai/DeepSeek-R1-Distill-Qwen-7B", "Deepseek7b", 4, ["http", "grpc"], {}), | ||
| ("Qwen/Qwen3-30B-A3B", "Qwen30b", 4, ["http", "grpc"], {}), | ||
| ("mistralai/Mistral-7B-Instruct-v0.3", "Mistral7b", 4, ["http", "grpc"], {}), | ||
| ("openai/gpt-oss-20b", "GptOss", 4, ["http", "grpc"], {}), |
There was a problem hiding this comment.
This change correctly updates the slug to include the model size, 20b. This improves clarity and consistency with other model naming conventions, making it easier to distinguish between different gpt-oss variants in job names and test classes.
| ("openai/gpt-oss-20b", "GptOss", 4, ["http", "grpc"], {}), | |
| ("openai/gpt-oss-20b", "GptOss20b", 4, ["http", "grpc"], {}), |
Signed-off-by: ppraneth <pranethparuchuri@gmail.com>
Description
Problem
The nightly benchmark slug for
openai/gpt-oss-20bisGptOss, which omits the model size. Every other model includes its parameter count in the slug (Llama8b,Qwen7b,Deepseek7b, etc.). This makes job names vague and ambiguous—especially since agpt-oss-120bvariant also exists in the codebase.Solution
Rename the slug from
GptOsstoGptOss20bso the generated test class names and CI job names explicitly include the model size.Changes
e2e_test/benchmarks/test_nightly_perf.py:"GptOss"→"GptOss20b"in_NIGHTLY_MODELS.github/workflows/nightly-benchmark.yml: updatetest_classreferences in bothsingle-workerandmulti-workermatrices (TestNightlyGptOssSingle→TestNightlyGptOss20bSingle,TestNightlyGptOssMulti→TestNightlyGptOss20bMulti)Test Plan
No functional change—only renames the slug used for test class generation and CI job naming.
Checklist
cargo +nightly fmtpassescargo clippy --all-targets --all-features -- -D warningspassesSummary by CodeRabbit