Skip to content

fix(tests): move gemini reasoning tests off the retired gemini-2.5-flash - #1275

Merged
njbrake merged 3 commits into
mainfrom
fix/gemini-reasoning-model-404
Aug 12, 2026
Merged

njbrake merged 3 commits into
mainfrom
fix/gemini-reasoning-model-404

Conversation

@njbrake

@njbrake njbrake commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Note: this PR description was drafted by Claude Opus 5 via back-and-forth with @njbrake. The reasoning and decisions are his; the prose is Claude's.

Description

test_completion_reasoning[gemini] and test_completion_reasoning_streaming[gemini] have been failing on main since at least run 31486350583. Two independent causes, both fixed here.

1. Retired model. The Gemini Developer API now 404s gemini-2.5-flash with "no longer available to new users", and the CI key is not grandfathered. provider_model_map already moved to gemini-3-flash-preview and its tests pass, so the key has access; only provider_reasoning_model_map was left behind.

2. A prompt that suppresses the thought summary. With the model swapped, the non-streaming test still failed on assert message.reasoning is not None. Its prompt asked the model to "think very briefly" at reasoning_effort="low". Gemini documents that a thought block may carry only a signature and no summary text for "simple requests, where the model didn't reason enough to generate a summary". gemini-2.5-flash happened to always emit one. The streaming variant already used "Think before you respond" and passed, so this aligns the two prompts.

No provider change is needed. Gemini 3.0 keeps thinking_budget for backward compatibility; the 400 people hit comes from sending thinking_level and thinking_budget together, which any-llm never does. #1276 tracks what happens when this moves to a 3.5+ model.

LLMProvider.VERTEXAI still points at gemini-2.5-flash here. Vertex is skipped entirely in CI for lack of credentials and deprecates on its own schedule, so changing it would be an unverifiable guess.

PR Type

  • 🐛 Bug Fix

Relevant issues

None filed for this. Evidence is CI run 31578025946. Found while reviewing: #1276 and #1277.

Checklist

  • I understand the code I am submitting.
  • I have added unit tests that prove my fix/feature works
    • The change is integration test configuration, so the integration tests are the test. No unit test applies.
  • I have run this code locally and verified it fixes the issue.
    • Via the run-integration-tests label rather than locally, since no GEMINI_API_KEY was available. Run 31623619771: both gemini reasoning tests pass, and the suite has no remaining failures.
  • New and existing tests pass locally
    • uv run pytest tests/unit: 2028 passed, 93 skipped. pre-commit run --all-files clean.
  • Documentation was updated where necessary
  • I have read and followed the contribution guidelines
  • AI Usage:
    • No AI was used.
    • AI was used for drafting/refactoring.
    • This is fully AI-generated.

AI Usage Information

  • AI Model used: Claude Opus 5

  • AI Developer Tool used: Claude Code

  • Any other info you'd like to share: An earlier draft also rewrote the provider to send thinking_level for Gemini 3, on the mistaken belief that Gemini 3 rejects thinking_budget. Google's Gemini 3 guide says it is still supported, so that was dropped and filed as [BUG] Gemini provider only sends thinking_budget, which Gemini 3.5+ will reject #1276 instead.

  • I am an AI Agent filling out this form (check box if true)

The Gemini Developer API now returns 404 for gemini-2.5-flash ("no longer
available to new users"), so both reasoning integration tests fail for the
gemini provider. The rest of the suite already moved to gemini-3-flash-preview
in provider_model_map; only provider_reasoning_model_map was left behind.

No provider change is needed: Gemini 3 keeps thinking_budget for backward
compatibility, and any-llm never sends thinking_level alongside it, which is
the combination the API rejects.
@njbrake
njbrake temporarily deployed to integration-tests August 12, 2026 17:08 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Aug 12, 2026 •

Copy link
Copy Markdown

Review Change Stack

Walkthrough

The reasoning test configuration now uses gemini-3-flash-preview. The non-streaming test asks the model to think before responding and documents Gemini thought-block behaviour.

Changes

Gemini reasoning test update

Layer / File(s) Summary
Update Gemini reasoning model
tests/conftest.py
The LLMProvider.GEMINI reasoning model changes to gemini-3-flash-preview. Comments document model availability and reasoning parameter compatibility.
Update reasoning test prompt
tests/integration/test_reasoning.py
The non-streaming test asks the model to think before responding without a time limit. Comments document Gemini thought-block behaviour.

Possibly related issues

  • mozilla-ai/any-llm#1276 — The test changes document thinking_budget compatibility with Gemini 3, which relates to migrating provider logic to thinking_level for newer Gemini models.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the main change: moving reasoning tests from the retired Gemini 2.5 Flash model.
Description check ✅ Passed The description follows the template and clearly explains the two fixes, test evidence, issue status, checklist, and AI usage.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/gemini-reasoning-model-404

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Aug 12, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
see 23 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@njbrake
njbrake temporarily deployed to integration-tests August 12, 2026 17:14 — with GitHub Actions Inactive

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/conftest.py`:
- Around line 37-39: Update the comment near the Gemini reasoning model
configuration to replace “the rest of the suite” with wording limited to the
direct Gemini tests, while preserving the explanation about Gemini 2.5
availability and Gemini 3.0 thinking_budget support.
- Around line 40-41: Update the backward-compatibility comment near the provider
configuration to state that thinking_level is the recommended control for newer
Gemini models while thinking_budget remains accepted for compatibility. Remove
the unsupported “goes away for 3.5+” removal timeline without changing the
surrounding configuration.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: de2b148a-2c35-4df6-a7ca-647624b64668

📥 Commits

Reviewing files that changed from the base of the PR and between 2560796 and 08abe71.

📒 Files selected for processing (1)
  • tests/conftest.py

Comment thread tests/conftest.py
Comment on lines +37 to +39
# gemini-2.5-flash is closed to new API keys (404 "no longer available to new users"), so
# reasoning runs on the same Gemini 3 model the rest of the suite uses. thinking_budget,
# which is what GoogleProvider sends for reasoning_effort, is still supported on Gemini 3.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Narrow the “rest of the suite” statement.

LLMProvider.VERTEXAI still uses gemini-2.5-flash at Line 48, and LLMProvider.OPENROUTER still uses a Gemini 2.5 model at Line 60. The comment should state that this change applies to the direct Gemini tests.

Proposed wording
-        # reasoning runs on the same Gemini 3 model the rest of the suite uses. thinking_budget,
+        # reasoning runs on the Gemini 3 model used by the direct Gemini tests. thinking_budget,
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
# gemini-2.5-flash is closed to new API keys (404 "no longer available to new users"), so
# reasoning runs on the same Gemini 3 model the rest of the suite uses. thinking_budget,
# which is what GoogleProvider sends for reasoning_effort, is still supported on Gemini 3.0
# gemini-2.5-flash is closed to new API keys (404 "no longer available to new users"), so
# reasoning runs on the Gemini 3 model used by the direct Gemini tests. thinking_budget,
# which is what GoogleProvider sends for reasoning_effort, is still supported on Gemini 3.0
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/conftest.py` around lines 37 - 39, Update the comment near the Gemini
reasoning model configuration to replace “the rest of the suite” with wording
limited to the direct Gemini tests, while preserving the explanation about
Gemini 2.5 availability and Gemini 3.0 thinking_budget support.

Comment thread tests/conftest.py
Comment on lines +40 to +41
# for backward compatibility. Google has signalled it goes away for 3.5+, so the provider
# will need to send thinking_level before this can point at a newer model.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

curl -fsSL 'https://ai.google.dev/gemini-api/docs/generate-content/thinking?hl=en' \
  | rg -ni 'thinking.?budget|thinking.?level|backward compatibility'

Repository: mozilla-ai/any-llm

Length of output: 7614


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

printf '%s\n' '--- tests/conftest.py ---'
cat -n tests/conftest.py | sed -n '25,55p'

printf '%s\n' '--- Gemini model and reasoning references ---'
rg -n -C 3 'gemini-3-flash-preview|gemini-2\.5-flash|reasoning_effort|thinking_budget|thinking_level|LLMProvider\.(GEMINI|VERTEXAI)' tests src .github 2>/dev/null | head -n 400

Repository: mozilla-ai/any-llm

Length of output: 32155


🏁 Script executed:

#!/usr/bin/env bash
set -euo pipefail

printf '%s\n' '--- Gemini provider reasoning conversion ---'
rg -n -C 8 'thinking_budget|thinking_level|reasoning_effort|ThinkingConfig' src/any_llm/providers tests/unit/providers -g '*.py' | head -n 300

printf '%s\n' '--- reasoning integration test ---'
cat -n tests/integration/test_reasoning.py | sed -n '1,135p'

Repository: mozilla-ai/any-llm

Length of output: 32681


Correct the thinking_budget lifecycle statement.

Google recommends thinking_level for Gemini 3 models and accepts thinking_budget for backwards compatibility. Replace “goes away for 3.5+” with “thinking_level is the recommended control for newer Gemini models”. Do not state an unsupported removal timeline.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/conftest.py` around lines 40 - 41, Update the backward-compatibility
comment near the provider configuration to state that thinking_level is the
recommended control for newer Gemini models while thinking_budget remains
accepted for compatibility. Remove the unsupported “goes away for 3.5+” removal
timeline without changing the surrounding configuration.

@njbrake njbrake added the run-integration-tests Put this label on a PR to trigger the integration test suite: works with forks label Aug 12, 2026
@github-actions github-actions Bot removed the run-integration-tests Put this label on a PR to trigger the integration test suite: works with forks label Aug 12, 2026
…ng test

The non-streaming reasoning test prompt combined "say hello" with "think very
briefly" at reasoning_effort=low. Gemini documents that a thought block may
carry only a signature and no summary text when the request is simple enough
that the model barely reasons, so message.reasoning came back None and the
assertion failed on gemini-3-flash-preview. gemini-2.5-flash happened to always
emit a summary, which is why this only surfaced with the model change.

Use the same prompt the streaming variant already uses, which passes.
@njbrake njbrake added the run-integration-tests Put this label on a PR to trigger the integration test suite: works with forks label Aug 12, 2026
@njbrake
njbrake temporarily deployed to integration-tests August 12, 2026 17:38 — with GitHub Actions Inactive
@github-actions github-actions Bot removed the run-integration-tests Put this label on a PR to trigger the integration test suite: works with forks label Aug 12, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/integration/test_reasoning.py`:
- Around line 35-39: Update the prompt in the reasoning test’s messages fixture
to require a small, multi-step reasoning task rather than a simple greeting,
ensuring both streaming and non-streaming responses produce textual reasoning
summaries. Keep the existing reasoning assertions unchanged.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9b9e2486-c780-4aaf-8690-7df12d49aee4

📥 Commits

Reviewing files that changed from the base of the PR and between 08abe71 and 408b6f6.

📒 Files selected for processing (1)
  • tests/integration/test_reasoning.py

Comment on lines +35 to +39
# Not "think very briefly": Gemini returns a thought block with a signature and no
# summary text when a request is simple enough that it barely reasons, which makes the
# reasoning assertion below fail. See the thought summaries section of
# https://ai.google.dev/gemini-api/docs/thinking
messages=[{"role": "user", "content": "Please say hello! Think before you respond."}],

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -eu
printf '%s\n' '--- test file ---'
cat -n tests/integration/test_reasoning.py | sed -n '1,110p'
printf '%s\n' '--- related reasoning assertions and Gemini configuration ---'
rg -n -C 3 'reasoning|thought|Gemini|gemini|think' tests/integration tests/unit pyproject.toml 2>/dev/null | head -240

Repository: mozilla-ai/any-llm

Length of output: 24194


🌐 Web query:

Gemini API thought summaries simple prompt thought signature no text summary thinking models documentation

💡 Result:

Gemini models with thinking capabilities reason internally before responding [1][2]. You can access insights into this process through "thought summaries" and maintain context across turns using "thought signatures" [1][3][4]. Thought Summaries Thought summaries provide a text or image-based distillation of the model's internal reasoning [1][3]. By default, these are disabled; you can enable them by configuring thinking_summaries [1][2]. When enabled, the API returns these summaries, which offer insight into the reasoning process while the full "raw" thinking tokens are used for the actual generation [3]. In streaming responses, these are delivered via Server-Sent Events (SSE) with the delta type thought_summary [1][2]. Thought Signatures A thought signature is an encrypted representation of the model's internal reasoning state [1][5]. It is essential for multi-turn conversations and function calling, as it allows the model to maintain reasoning context [3][4]. - Function Calling: If the model generates a function call, a thought signature is typically attached to the first function call part [4]. You must include this signature when sending the conversation history back to the model in the next turn to ensure proper processing and high-quality responses [4]. - API Differences: - In the Interactions API, thought signatures are found in dedicated thought steps [1][6]. - In the generateContent API, there are no dedicated thought blocks; signatures appear as metadata attached to parts, such as within functionCall or final response parts [5][4][6]. Configuration You control the reasoning process using thinkingBudget (or ThinkingConfig in some SDK implementations), which dictates the number of tokens used for reasoning [3][7]. Setting this to 0 disables thinking, while higher values allow for more complex reasoning [7]. Gemini models generate full thoughts to improve response quality, and pricing is based on these full thinking tokens, not just the returned summaries [3].

Citations:


🏁 Script executed:

#!/bin/bash
set -eu
printf '%s\n' '--- Gemini provider files ---'
fd -i 'gemini' . --type f | head -80
printf '%s\n' '--- reasoning conversion symbols ---'
rg -n -C 4 'reasoning|thought_summary|thought_signature|thinking_budget|thinkingConfig|thinking_config' . \
  -g '*.py' -g '!tests/integration/test_reasoning.py' | head -320
printf '%s\n' '--- reasoning model fixtures ---'
rg -n -C 3 'provider_reasoning_model_map|GEMINI' tests conftest.py .github 2>/dev/null | head -240

Repository: mozilla-ai/any-llm

Length of output: 39795


🏁 Script executed:

#!/bin/bash
set -eu
printf '%s\n' '--- Gemini provider outline ---'
ast-grep outline src/any_llm/providers/gemini/gemini.py
printf '%s\n' '--- Gemini provider implementation ---'
cat -n src/any_llm/providers/gemini/gemini.py | sed -n '1,280p'
printf '%s\n' '--- reasoning test streaming section ---'
cat -n tests/integration/test_reasoning.py | sed -n '98,175p'
printf '%s\n' '--- model fixture context ---'
cat -n tests/conftest.py | sed -n '30,50p'

Repository: mozilla-ai/any-llm

Length of output: 7101


🏁 Script executed:

#!/bin/bash
set -eu
printf '%s\n' '--- Google provider outline ---'
ast-grep outline src/any_llm/providers/gemini/base.py
printf '%s\n' '--- Google provider reasoning conversion ---'
rg -n -C 12 'reasoning|thought|thinking|GenerateContent|Content|Part|usage' src/any_llm/providers/gemini/base.py
printf '%s\n' '--- Google provider relevant source ---'
cat -n src/any_llm/providers/gemini/base.py | sed -n '1,360p'

Repository: mozilla-ai/any-llm

Length of output: 24623


🏁 Script executed:

#!/bin/bash
set -eu
printf '%s\n' '--- Google response helper locations ---'
rg -n -C 8 '_convert_response_to_response_dict|_create_openai_chunk_from_google_chunk|thought_signature|thought_summary|reasoning' src/any_llm/providers/gemini
printf '%s\n' '--- helper source ---'
cat -n src/any_llm/providers/gemini/utils.py | sed -n '1,360p'
printf '%s\n' '--- read-only AST probe of reasoning construction ---'
python3 - <<'PY'
import ast
from pathlib import Path

paths = [
    Path("src/any_llm/providers/gemini/base.py"),
    Path("src/any_llm/providers/gemini/utils.py"),
]
for path in paths:
    tree = ast.parse(path.read_text())
    print(f"{path}:")
    for node in ast.walk(tree):
        if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)):
            if node.name in {"_convert_completion_response", "_convert_response_to_response_dict",
                             "_create_openai_chunk_from_google_chunk"}:
                print(f"  {node.name} lines {node.lineno}-{node.end_lineno}")
                for child in ast.walk(node):
                    if isinstance(child, ast.Call) and isinstance(child.func, ast.Name):
                        if child.func.id in {"Reasoning", "ChatCompletionMessage", "ChatCompletionChunk"}:
                            print(f"    constructs {child.func.id} at line {child.lineno}")
                    if isinstance(child, ast.Call) and isinstance(child.func, ast.Attribute):
                        if child.func.attr in {"get", "text"}:
                            print(f"    accesses .{child.func.attr} at line {child.lineno}")
PY

Repository: mozilla-ai/any-llm

Length of output: 44866


Make the prompt require non-trivial reasoning.

The test requires textual reasoning in both non-streaming and streaming responses. A simple greeting can return a thought signature without summary text, so “Think before you respond” does not guarantee that either assertion passes. Use a small multi-step task, or accept signature-only reasoning if that is the intended contract.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/integration/test_reasoning.py` around lines 35 - 39, Update the prompt
in the reasoning test’s messages fixture to require a small, multi-step
reasoning task rather than a simple greeting, ensuring both streaming and
non-streaming responses produce textual reasoning summaries. Keep the existing
reasoning assertions unchanged.

@njbrake
njbrake merged commit 3d5cebd into main Aug 12, 2026
23 of 24 checks passed
@njbrake
njbrake deleted the fix/gemini-reasoning-model-404 branch August 12, 2026 18:17
@github-actions github-actions Bot added the 1.26.0 Included in release 1.26.0 label Aug 17, 2026

This branch was previously deployed

1 inactive deployment
integration-tests — 408b6f60 Deployed Aug 12, 2026 by njbrake via run-docs-tests #2397
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

1.26.0 Included in release 1.26.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant