[CI] Pin DeepSeek-V4 Flash GPQA nightly to low reasoning effort - #16721
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request updates the AISBench tool to support the Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Benchmark][Feature] Add support for reasoning_effort configuration in aisbenchSuggested PR Summary:
### What this PR does / why we need it?
This PR introduces support for the `reasoning_effort` configuration in the `aisbench` benchmarking tool. This allows specifying the reasoning effort (e.g., "low") for models like DeepSeek-V4-Flash. The changes include:
- Retrieving and propagating `reasoning_effort` in `AisbenchRunner`.
- Updating the request configuration generator to inject `reasoning_effort` into the generated python configuration file.
- Adding a nightly E2E test configuration for `DeepSeek-V4-Flash-W8A8-A3` with `reasoning_effort: low`.
- Adding unit tests to verify the correct generation of the request configuration when `reasoning_effort` is specified.
### Does this PR introduce _any_ user-facing change?
Yes, users can now configure `reasoning_effort` in their `aisbench` configuration files.
### How was this patch tested?
- Added unit test `test_request_config_reasoning_effort` in `tests/ut/tools/test_aisbench.py`.
- Added nightly E2E configuration in `tests/e2e/nightly/single_node/models/configs/DeepSeek-V4-Flash-W8A8-A3.yaml`.I have no feedback to provide as the implementation is correct and well-tested.
Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
900801d to
e7dad98
Compare
Pass optional reasoning_effort through AISBench chat request generation. Set reasoning_effort: low for the DeepSeek-V4 Flash W8A8 A3 nightly GPQA case. Keep thinking: true and the performance case unchanged. Cherry-picked from vllm-project#16721 Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
Pass optional reasoning_effort through AISBench chat request generation. Set reasoning_effort: low for the DeepSeek-V4 Flash W8A8 A3 nightly GPQA case. Keep thinking: true and the performance case unchanged. Cherry-picked from vllm-project#16721 Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com> Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
…-project#16721) ## Summary - Pass the optional `reasoning_effort` benchmark setting through to AISBench chat request generation. - Set `reasoning_effort: low` for the DeepSeek-V4 Flash W8A8 A3 nightly GPQA case. Keep `thinking: true` and the performance case unchanged. - Add a focused test for the generated request config, including unchanged behavior when the setting is absent. ## Context The [failing nightly job](https://github.com/vllm-project/vllm-ascend/actions/runs/34497673764/job/103072819176#logs) evaluated this case without an explicit reasoning effort. The DeepSeek-V4 default thinking strength changed upstream, so this pins the intended GPQA protocol rather than relying on that default. An earlier manual GPQA Diamond run using `reasoning_effort="low"` and `chat_template_kwargs={"thinking": True}` scored 88.38% (175/198) on the e5118d1/vLLM 0.28.0 setup. That is supporting evidence, not a validation of this PR's current-main nightly run. ## Validation - Parsed the modified YAML and Python files; confirmed only the GPQA benchmark selects `low`. - `git diff --check` passed. - The new unit test and full A3 nightly GPQA run are pending CI/runner execution. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Summary
reasoning_effortbenchmark setting through to AISBench chat request generation.reasoning_effort: lowfor the DeepSeek-V4 Flash W8A8 A3 nightly GPQA case. Keepthinking: trueand the performance case unchanged.Context
The failing nightly job evaluated this case without an explicit reasoning effort. The DeepSeek-V4 default thinking strength changed upstream, so this pins the intended GPQA protocol rather than relying on that default.
An earlier manual GPQA Diamond run using
reasoning_effort="low"andchat_template_kwargs={"thinking": True}scored 88.38% (175/198) on the e5118d1/vLLM 0.28.0 setup. That is supporting evidence, not a validation of this PR's current-main nightly run.Validation
Parsed the modified YAML and Python files; confirmed only the GPQA benchmark selects
low.git diff --checkpassed.The new unit test and full A3 nightly GPQA run are pending CI/runner execution.
vLLM main: vllm-project/vllm@84030bb