fix(tests): move gemini reasoning tests off the retired gemini-2.5-flash - #1275
Conversation
The Gemini Developer API now returns 404 for gemini-2.5-flash ("no longer
available to new users"), so both reasoning integration tests fail for the
gemini provider. The rest of the suite already moved to gemini-3-flash-preview
in provider_model_map; only provider_reasoning_model_map was left behind.
No provider change is needed: Gemini 3 keeps thinking_budget for backward
compatibility, and any-llm never sends thinking_level alongside it, which is
the combination the API rejects.
WalkthroughThe reasoning test configuration now uses ChangesGemini reasoning test update
Possibly related issues
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Codecov Report✅ All modified and coverable lines are covered by tests. 🚀 New features to boost your workflow:
|
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/conftest.py`:
- Around line 37-39: Update the comment near the Gemini reasoning model
configuration to replace “the rest of the suite” with wording limited to the
direct Gemini tests, while preserving the explanation about Gemini 2.5
availability and Gemini 3.0 thinking_budget support.
- Around line 40-41: Update the backward-compatibility comment near the provider
configuration to state that thinking_level is the recommended control for newer
Gemini models while thinking_budget remains accepted for compatibility. Remove
the unsupported “goes away for 3.5+” removal timeline without changing the
surrounding configuration.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: de2b148a-2c35-4df6-a7ca-647624b64668
📒 Files selected for processing (1)
tests/conftest.py
| # gemini-2.5-flash is closed to new API keys (404 "no longer available to new users"), so | ||
| # reasoning runs on the same Gemini 3 model the rest of the suite uses. thinking_budget, | ||
| # which is what GoogleProvider sends for reasoning_effort, is still supported on Gemini 3.0 |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
Narrow the “rest of the suite” statement.
LLMProvider.VERTEXAI still uses gemini-2.5-flash at Line 48, and LLMProvider.OPENROUTER still uses a Gemini 2.5 model at Line 60. The comment should state that this change applies to the direct Gemini tests.
Proposed wording
- # reasoning runs on the same Gemini 3 model the rest of the suite uses. thinking_budget,
+ # reasoning runs on the Gemini 3 model used by the direct Gemini tests. thinking_budget,📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| # gemini-2.5-flash is closed to new API keys (404 "no longer available to new users"), so | |
| # reasoning runs on the same Gemini 3 model the rest of the suite uses. thinking_budget, | |
| # which is what GoogleProvider sends for reasoning_effort, is still supported on Gemini 3.0 | |
| # gemini-2.5-flash is closed to new API keys (404 "no longer available to new users"), so | |
| # reasoning runs on the Gemini 3 model used by the direct Gemini tests. thinking_budget, | |
| # which is what GoogleProvider sends for reasoning_effort, is still supported on Gemini 3.0 |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/conftest.py` around lines 37 - 39, Update the comment near the Gemini
reasoning model configuration to replace “the rest of the suite” with wording
limited to the direct Gemini tests, while preserving the explanation about
Gemini 2.5 availability and Gemini 3.0 thinking_budget support.
| # for backward compatibility. Google has signalled it goes away for 3.5+, so the provider | ||
| # will need to send thinking_level before this can point at a newer model. |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/usr/bin/env bash
set -euo pipefail
curl -fsSL 'https://ai.google.dev/gemini-api/docs/generate-content/thinking?hl=en' \
| rg -ni 'thinking.?budget|thinking.?level|backward compatibility'Repository: mozilla-ai/any-llm
Length of output: 7614
🏁 Script executed:
#!/usr/bin/env bash
set -euo pipefail
printf '%s\n' '--- tests/conftest.py ---'
cat -n tests/conftest.py | sed -n '25,55p'
printf '%s\n' '--- Gemini model and reasoning references ---'
rg -n -C 3 'gemini-3-flash-preview|gemini-2\.5-flash|reasoning_effort|thinking_budget|thinking_level|LLMProvider\.(GEMINI|VERTEXAI)' tests src .github 2>/dev/null | head -n 400Repository: mozilla-ai/any-llm
Length of output: 32155
🏁 Script executed:
#!/usr/bin/env bash
set -euo pipefail
printf '%s\n' '--- Gemini provider reasoning conversion ---'
rg -n -C 8 'thinking_budget|thinking_level|reasoning_effort|ThinkingConfig' src/any_llm/providers tests/unit/providers -g '*.py' | head -n 300
printf '%s\n' '--- reasoning integration test ---'
cat -n tests/integration/test_reasoning.py | sed -n '1,135p'Repository: mozilla-ai/any-llm
Length of output: 32681
Correct the thinking_budget lifecycle statement.
Google recommends thinking_level for Gemini 3 models and accepts thinking_budget for backwards compatibility. Replace “goes away for 3.5+” with “thinking_level is the recommended control for newer Gemini models”. Do not state an unsupported removal timeline.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/conftest.py` around lines 40 - 41, Update the backward-compatibility
comment near the provider configuration to state that thinking_level is the
recommended control for newer Gemini models while thinking_budget remains
accepted for compatibility. Remove the unsupported “goes away for 3.5+” removal
timeline without changing the surrounding configuration.
…ng test The non-streaming reasoning test prompt combined "say hello" with "think very briefly" at reasoning_effort=low. Gemini documents that a thought block may carry only a signature and no summary text when the request is simple enough that the model barely reasons, so message.reasoning came back None and the assertion failed on gemini-3-flash-preview. gemini-2.5-flash happened to always emit a summary, which is why this only surfaced with the model change. Use the same prompt the streaming variant already uses, which passes.
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/integration/test_reasoning.py`:
- Around line 35-39: Update the prompt in the reasoning test’s messages fixture
to require a small, multi-step reasoning task rather than a simple greeting,
ensuring both streaming and non-streaming responses produce textual reasoning
summaries. Keep the existing reasoning assertions unchanged.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 9b9e2486-c780-4aaf-8690-7df12d49aee4
📒 Files selected for processing (1)
tests/integration/test_reasoning.py
| # Not "think very briefly": Gemini returns a thought block with a signature and no | ||
| # summary text when a request is simple enough that it barely reasons, which makes the | ||
| # reasoning assertion below fail. See the thought summaries section of | ||
| # https://ai.google.dev/gemini-api/docs/thinking | ||
| messages=[{"role": "user", "content": "Please say hello! Think before you respond."}], |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -eu
printf '%s\n' '--- test file ---'
cat -n tests/integration/test_reasoning.py | sed -n '1,110p'
printf '%s\n' '--- related reasoning assertions and Gemini configuration ---'
rg -n -C 3 'reasoning|thought|Gemini|gemini|think' tests/integration tests/unit pyproject.toml 2>/dev/null | head -240Repository: mozilla-ai/any-llm
Length of output: 24194
🌐 Web query:
Gemini API thought summaries simple prompt thought signature no text summary thinking models documentation
💡 Result:
Gemini models with thinking capabilities reason internally before responding [1][2]. You can access insights into this process through "thought summaries" and maintain context across turns using "thought signatures" [1][3][4]. Thought Summaries Thought summaries provide a text or image-based distillation of the model's internal reasoning [1][3]. By default, these are disabled; you can enable them by configuring thinking_summaries [1][2]. When enabled, the API returns these summaries, which offer insight into the reasoning process while the full "raw" thinking tokens are used for the actual generation [3]. In streaming responses, these are delivered via Server-Sent Events (SSE) with the delta type thought_summary [1][2]. Thought Signatures A thought signature is an encrypted representation of the model's internal reasoning state [1][5]. It is essential for multi-turn conversations and function calling, as it allows the model to maintain reasoning context [3][4]. - Function Calling: If the model generates a function call, a thought signature is typically attached to the first function call part [4]. You must include this signature when sending the conversation history back to the model in the next turn to ensure proper processing and high-quality responses [4]. - API Differences: - In the Interactions API, thought signatures are found in dedicated thought steps [1][6]. - In the generateContent API, there are no dedicated thought blocks; signatures appear as metadata attached to parts, such as within functionCall or final response parts [5][4][6]. Configuration You control the reasoning process using thinkingBudget (or ThinkingConfig in some SDK implementations), which dictates the number of tokens used for reasoning [3][7]. Setting this to 0 disables thinking, while higher values allow for more complex reasoning [7]. Gemini models generate full thoughts to improve response quality, and pricing is based on these full thinking tokens, not just the returned summaries [3].
Citations:
- 1: https://ai.google.dev/gemini-api/docs/interactions/thinking
- 2: https://ai.google.dev/gemini-api/docs/thinking
- 3: https://ai.google.dev/gemini-api/docs/generate-content/thinking
- 4: https://ai.google.dev/gemini-api/docs/generate-content/thought-signatures
- 5: https://ai.google.dev/gemini-api/docs/thought-signatures
- 6: https://ai.google.dev/gemini-api/docs/thinking.md.txt
- 7: https://github.com/google-gemini/cookbook/blob/main/quickstarts/Get_started_thinking.ipynb
🏁 Script executed:
#!/bin/bash
set -eu
printf '%s\n' '--- Gemini provider files ---'
fd -i 'gemini' . --type f | head -80
printf '%s\n' '--- reasoning conversion symbols ---'
rg -n -C 4 'reasoning|thought_summary|thought_signature|thinking_budget|thinkingConfig|thinking_config' . \
-g '*.py' -g '!tests/integration/test_reasoning.py' | head -320
printf '%s\n' '--- reasoning model fixtures ---'
rg -n -C 3 'provider_reasoning_model_map|GEMINI' tests conftest.py .github 2>/dev/null | head -240Repository: mozilla-ai/any-llm
Length of output: 39795
🏁 Script executed:
#!/bin/bash
set -eu
printf '%s\n' '--- Gemini provider outline ---'
ast-grep outline src/any_llm/providers/gemini/gemini.py
printf '%s\n' '--- Gemini provider implementation ---'
cat -n src/any_llm/providers/gemini/gemini.py | sed -n '1,280p'
printf '%s\n' '--- reasoning test streaming section ---'
cat -n tests/integration/test_reasoning.py | sed -n '98,175p'
printf '%s\n' '--- model fixture context ---'
cat -n tests/conftest.py | sed -n '30,50p'Repository: mozilla-ai/any-llm
Length of output: 7101
🏁 Script executed:
#!/bin/bash
set -eu
printf '%s\n' '--- Google provider outline ---'
ast-grep outline src/any_llm/providers/gemini/base.py
printf '%s\n' '--- Google provider reasoning conversion ---'
rg -n -C 12 'reasoning|thought|thinking|GenerateContent|Content|Part|usage' src/any_llm/providers/gemini/base.py
printf '%s\n' '--- Google provider relevant source ---'
cat -n src/any_llm/providers/gemini/base.py | sed -n '1,360p'Repository: mozilla-ai/any-llm
Length of output: 24623
🏁 Script executed:
#!/bin/bash
set -eu
printf '%s\n' '--- Google response helper locations ---'
rg -n -C 8 '_convert_response_to_response_dict|_create_openai_chunk_from_google_chunk|thought_signature|thought_summary|reasoning' src/any_llm/providers/gemini
printf '%s\n' '--- helper source ---'
cat -n src/any_llm/providers/gemini/utils.py | sed -n '1,360p'
printf '%s\n' '--- read-only AST probe of reasoning construction ---'
python3 - <<'PY'
import ast
from pathlib import Path
paths = [
Path("src/any_llm/providers/gemini/base.py"),
Path("src/any_llm/providers/gemini/utils.py"),
]
for path in paths:
tree = ast.parse(path.read_text())
print(f"{path}:")
for node in ast.walk(tree):
if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)):
if node.name in {"_convert_completion_response", "_convert_response_to_response_dict",
"_create_openai_chunk_from_google_chunk"}:
print(f" {node.name} lines {node.lineno}-{node.end_lineno}")
for child in ast.walk(node):
if isinstance(child, ast.Call) and isinstance(child.func, ast.Name):
if child.func.id in {"Reasoning", "ChatCompletionMessage", "ChatCompletionChunk"}:
print(f" constructs {child.func.id} at line {child.lineno}")
if isinstance(child, ast.Call) and isinstance(child.func, ast.Attribute):
if child.func.attr in {"get", "text"}:
print(f" accesses .{child.func.attr} at line {child.lineno}")
PYRepository: mozilla-ai/any-llm
Length of output: 44866
Make the prompt require non-trivial reasoning.
The test requires textual reasoning in both non-streaming and streaming responses. A simple greeting can return a thought signature without summary text, so “Think before you respond” does not guarantee that either assertion passes. Use a small multi-step task, or accept signature-only reasoning if that is the intended contract.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/integration/test_reasoning.py` around lines 35 - 39, Update the prompt
in the reasoning test’s messages fixture to require a small, multi-step
reasoning task rather than a simple greeting, ensuring both streaming and
non-streaming responses produce textual reasoning summaries. Keep the existing
reasoning assertions unchanged.
Note: this PR description was drafted by Claude Opus 5 via back-and-forth with @njbrake. The reasoning and decisions are his; the prose is Claude's.
Description
test_completion_reasoning[gemini]andtest_completion_reasoning_streaming[gemini]have been failing onmainsince at least run 31486350583. Two independent causes, both fixed here.1. Retired model. The Gemini Developer API now 404s
gemini-2.5-flashwith "no longer available to new users", and the CI key is not grandfathered.provider_model_mapalready moved togemini-3-flash-previewand its tests pass, so the key has access; onlyprovider_reasoning_model_mapwas left behind.2. A prompt that suppresses the thought summary. With the model swapped, the non-streaming test still failed on
assert message.reasoning is not None. Its prompt asked the model to "think very briefly" atreasoning_effort="low". Gemini documents that a thought block may carry only a signature and no summary text for "simple requests, where the model didn't reason enough to generate a summary".gemini-2.5-flashhappened to always emit one. The streaming variant already used "Think before you respond" and passed, so this aligns the two prompts.No provider change is needed. Gemini 3.0 keeps
thinking_budgetfor backward compatibility; the 400 people hit comes from sendingthinking_levelandthinking_budgettogether, which any-llm never does. #1276 tracks what happens when this moves to a 3.5+ model.LLMProvider.VERTEXAIstill points atgemini-2.5-flashhere. Vertex is skipped entirely in CI for lack of credentials and deprecates on its own schedule, so changing it would be an unverifiable guess.PR Type
Relevant issues
None filed for this. Evidence is CI run 31578025946. Found while reviewing: #1276 and #1277.
Checklist
run-integration-testslabel rather than locally, since noGEMINI_API_KEYwas available. Run 31623619771: both gemini reasoning tests pass, and the suite has no remaining failures.uv run pytest tests/unit: 2028 passed, 93 skipped.pre-commit run --all-filesclean.AI Usage Information
AI Model used: Claude Opus 5
AI Developer Tool used: Claude Code
Any other info you'd like to share: An earlier draft also rewrote the provider to send
thinking_levelfor Gemini 3, on the mistaken belief that Gemini 3 rejectsthinking_budget. Google's Gemini 3 guide says it is still supported, so that was dropped and filed as [BUG] Gemini provider only sends thinking_budget, which Gemini 3.5+ will reject #1276 instead.I am an AI Agent filling out this form (check box if true)