fix(gemini): reasoning_effort="none" disables thinking instead of hiding it - #5
Merged
Merged
Conversation
Owner
Author
|
Submitted upstream as mozilla-ai#1294 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
reasoning_effort="none"(andNone) currently maps toThinkingConfig(include_thoughts=False), which hides thought summaries instead of turning off thinking."none"already means off everywhere else in this codebase:thinking={"type": "disabled"},think=False,messages_compat.pymaps Anthropic'sdisabledback toreasoning_effort="none".Gemini is the one provider where "none" means "hide but keep billing".
The fix, per path:
thinking_budgetmodels: sendThinkingConfig(include_thoughts=False, thinking_budget=0), so thinking is actually off.thinking_budget=0(e.g.gemini-2.5-pro): Google's own 400 surfaces as the typed error, reacting to the provider's error rather than predicting it with a model list.[gemini] 400 INVALID_ARGUMENT ... "Budget 0 is invalid. This model only works in thinking mode."thinking_levelmodels (3.5+):thinking_levelhas no off value since the lowest tier isMINIMAL."none"maps toThinkingLevel.MINIMALwith alogger.warning, the same clamp-to-nearest convention already used in fix(mistral): clamp reasoning_effort, drop prompt_mode, replay thinking traces mozilla-ai/any-llm#1241 and fix(anthropic): map xhigh to Anthropic xhigh and expose max effort mozilla-ai/any-llm#1108.Why clamp with warning, instead of error?
On 3.5/3.6 the clamp billed zero thought tokens (table below), so
"none"callers get what they asked for. Erroring up front instead would mean guessing from the version string which models those are, the kind of guess mozilla-ai#1189 declined. A model that doesn't supportMINIMALrejects it, and the caller gets the typed 400 (3.7, note below).The warning is because
MINIMALis the least thinking the model allows, not a promise of zero, and mozilla-ai#1107 treats surprise thinking costs as a bug.Nonestays merged with"none"as it has been since mozilla-ai#708;"auto"remains the unset default, so callers who don't set the parameter are untouched.Verified live across every path with
reasoning_effort="none"(gap = thought tokens billed but hidden; before the fix,gemini-2.5-flashbilled 424 hidden thought tokens):InvalidRequestError: "Budget 0 is invalid. This model only works in thinking mode."InvalidRequestError: "Thinking level MINIMAL is not supported", details belowEvery failure surfaces as a typed any-llm exception; no raw
google.genaierrors leak.Why gemini-3.7-flash rejects the clamp
gemini-3.7-flashcannot disable thinking:MINIMALis gone, the new floorLOWthinks for real (220 thought tokens on the test prompt), andthinking_budget=0is silently ignored but still billed. So the 400 is correct, same as the thinking-only pros, and clamping toLOWwould just bring back the bug this PR fixes. The removal also breaks the existing"minimal"mapping (mozilla-ai#1281) on 3.7.For reference,
thinking_budget=0today is honored on 3.5, rejected on 3.6, and ignored-and-billed on 3.7. Three behaviors in three consecutive minor versions is why this PR trusts the SDK enum and provider errors instead of a model matrix.Behavior change
reasoning_effort="none"previously kept thinking active (billed, hidden); it now genuinely disables it. Models that cannot run without thinking now raise Google's 400 instead of silently billing. Callers who relied on "hidden but active" thinking should use"auto"or a graded level. Default ("auto") behavior is unchanged.PR Type
Relevant issues
Fixes #2 (staged on this fork; re-pointed to the upstream issue number at submission)
Checklist
AI Usage Information
AI Model used: Claude (Fable 5)
AI Developer Tool used: Claude Code
Any other info you'd like to share:
I am an AI Agent filling out this form (check box if true)