Repository navigation
fix(bedrock): backport #41870 and the GPT-6 reasoning gate fix to stable/1.100.x for v1.100.2 - #42000
Conversation
…and xAI Grok models on Converse (#41870) Backport of #41870 to stable/1.100.x. Cherry-picked from a6e3a72 (main) with -m 1. converse_transformation.py conflicted because this line has no `import re` and no _is_openai_gpt_reasoning_model helper next to the insertion point. The resolution adds exactly the four hunks #41870 merged: the import, the 16-token constant, _requires_min_max_tokens, and the clamped maxTokens assignment. The test file applied clean.
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
|
|
||
| @staticmethod | ||
| def _requires_min_max_tokens(model: str) -> bool: | ||
| return re.search(r"openai\.gpt-\d|xai\.grok-", model) is not None |
There was a problem hiding this comment.
Hardcoded matching violates the directive to store model capabilities in metadata and misses opaque profiles, so GPT/Grok requests below 16 are rejected upstream
Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)
Knowledge Base Used: Provider adapters and capabilities
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
Backport of #41870, where this was withdrawn. The regex clamps ARNs embedding the model id; opaque application profiles carry none, so get_model_info cannot resolve them
There was a problem hiding this comment.
You're right. In this backport, matching the model string is intentional: it covers inference-profile ARNs that embed openai.gpt-* or xai.grok-*, while opaque application-inference-profile ARNs contain no resolvable model identifier for get_model_info to use. Given that limitation and the behavior established by #41870, this comment does not apply. No change needed.
…oning gate Backports the Converse part of fbc6fb5 from main (PR #31884). The gate only matched openai.gpt-5, so a GPT-6 model fell through to Anthropic's thinking block and Bedrock rejected the first real turn after a Claude Code /model switch with 400 Unknown parameter: 'thinking'. The Nova 2 tool_choice registry keys and the invoke json_mode forwarding in that commit stay on main
1 similar comment
| return re.search(r"openai\.gpt-\d", model) is not None | ||
|
|
There was a problem hiding this comment.
Hardcoded reasoning capability
_is_openai_gpt_reasoning_model decides reasoning support from the openai.gpt-<digit> model name. That result controls both the reasoning.effort mapping and whether thinking is discarded. This violates the repository directive that model-specific capabilities must be stored in model_prices_and_context_window.json and queried through model metadata. The repository requirement must be satisfied before merging.
Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
This mirrors main's merged #31884. A stable backport should not diverge from main, and the family pattern covers new GPT releases without a cost-map entry
There was a problem hiding this comment.
That’s fair. For this PR, matching main’s merged #31884 behavior is the correct backport constraint, and the family-level match is intentionally forward-compatible with future openai.gpt-<digit> Bedrock models without requiring a cost-map update for each release. I’m withdrawing this as a merge blocker; the metadata-vs-family-gate convention can be revisited separately if main changes.
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 5545ca9. Configure here.
TLDR
Problem this solves:
stable/1.100.xline hit the bug fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870 fixed on main: Claude Code's/modelswitch to a Bedrock OpenAI GPT or xAI Grok model fails with a 400, because Claude Code sends amax_tokens=1probe on switch, those Bedrock models require at least 16, and the Bedrock Converse path forwardedmaxTokensverbatimus.openai.gpt-6-astra) still fails with 400Unknown parameter: 'thinking', because the Converse reasoning gate on this line did not recognize GPT-6 models and Claude Code'sthinkingblock was mapped as Anthropic thinking on a model that rejects itstable/1.100.xthe gate existed (from fix(bedrock): map reasoning_effort to reasoning.effort for OpenAI GPT-5.x on Converse #38279) but only matchedopenai.gpt-5How it solves it:
stable/1.100.x:maxTokensis clamped to 16 foropenai.gpt-*andxai.grok-*models on the Bedrock Converse route, other Bedrock models keep pass-through behavioropenai.gpt-<digit>model is recognized, soreasoning_effortmaps toreasoning.effortand an Anthropicthinkingblock is dropped for these models instead of being forwarded to Bedrock1.100.2Backport notes
Cherry-picked from main merge commit a6e3a72 (fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870) with
-m 1, so the branch carries fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870's source and test changes as one commitlitellm/llms/bedrock/chat/converse_transformation.pyconflicted in that cherry-pick because this line never hadimport reor the_is_openai_gpt_reasoning_modelhelper that fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870's hunks sit next to on main; the resolution addsimport re, theBEDROCK_OPENAI_COMPAT_MIN_MAX_TOKENSconstant, the_requires_min_max_tokenshelper, and the clampedmaxTokensassignment exactly as they merged on main, nothing elsetests/test_litellm/llms/bedrock/chat/test_converse_transformation.pyapplied clean, so the ten parametrized cases from fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870 run unchanged on this lineThe reasoning gate change is a hand port of the Converse part of main's fbc6fb5 (fix(pricing): rolling model registry update: Bedrock gpt-6-astra, gpt-image-2.5, Cohere rerank 4, Vertex Grok 4.3/4.6/4.20, Gemini 3.5 audio, OpenAI web search fee, xAI Imagine video, Lyria 3.5, Voyage, ChatGPT GPT-5.5/5.6, Bedrock Mantle, Scaleway dates #31884): the
_is_openai_gpt_reasoning_modelhelper matchingopenai\.gpt-\dand its three call sites (thereasoning_effortmapping,get_supported_openai_params, and thethinkingskip inmap_openai_params). The Nova 2tool_choiceregistry keys and the invokejson_modeforwarding in that commit stay on main. The added parametrized test cases come from the same commit and fail on this line without the gate changeVersion bump to
1.100.2is its own commit, one patch above the latest release on this line (1.100.1)User Flow
Before: a developer using Claude Code through a v1.100.1 proxy cannot use a Bedrock OpenAI GPT or xAI Grok model
ANTHROPIC_BASE_URLpointed at the proxy and type/model us.xai.grok-4.6(or/model us.openai.gpt-6-astra)"max_tokens": 1as a warmup probeAPI error: 400 ... BedrockException ... Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.and the header keeps the model they had (Sonnet 4.5)us.openai.gpt-6-astratakes, and their first message, which Claude Code sends to POST https://litellm-domain/v1/messages with athinkingblock, comes back asAPI Error: 400 ... Unknown parameter: 'thinking'After: the same switch and first message succeed once the proxy runs v1.100.2
ANTHROPIC_BASE_URLpointed at the proxy and type/model us.xai.grok-4.6(or/model us.openai.gpt-6-astra)"max_tokens": 1as a warmup probethinkingblock and is answered by that Bedrock modelRelevant issues
Backport of #41870 and of the Converse reasoning gate fix from #31884 (main). Reported by customers (Pylon #8817 for OpenAI GPT, Pylon #8821 for xAI Grok); one of them runs v1.98.0 today and is moving to the 1.100.x line
Affected release
Linear ticket
Resolves LIT-8154
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Both legs ran the same way and differ only in the commit the proxy was booted from: a no-DB proxy with 2 uvicorn workers on a random port, real Bedrock in us-east-1 (real spend, no mocks), and Claude Code 2.1.278 driven interactively under tmux.
$KEYis the proxy master key and$PORTis the leg's portconfig.yamlProxy, booted from the leg's commit with a Python 3.12 venv
Claude Code, started in an empty folder with a fresh config directory (theme picker and security notes accepted with Enter, folder trust accepted with Down then Enter)
tmux new-session -d -s claude -x 160 -y 45 -c work "env -u ANTHROPIC_API_KEY ANTHROPIC_BASE_URL=http://localhost:$PORT ANTHROPIC_AUTH_TOKEN=$KEY ANTHROPIC_MODEL=us.anthropic.claude-sonnet-4-5-20250929-v1:0 ANTHROPIC_DEFAULT_HAIKU_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0 ANTHROPIC_SMALL_FAST_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0 CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 CLAUDE_CONFIG_DIR=claude-config claude"Before (b7d81e9)
The last curl group (the four reasoning cases) ran in a second proxy boot at this same commit after the floor-only after leg surfaced the
thinking400; every other case below is from the original before legClaude Code, /model us.xai.grok-4.6
In the running Claude Code session, type
/model us.xai.grok-4.6and press Enter. The pane shows (header, then the lines this step added)Type
Reply with the single word okand press Enter. The pane shows (header, then the lines this step added)Claude Code, /model us.openai.gpt-6-astra
In the running Claude Code session, type
/model us.openai.gpt-6-astraand press Enter. The pane shows (header, then the lines this step added)Type
Reply with the single word okand press Enter. The pane shows (header, then the lines this step added)POST /v1/messages, max_tokens 1
Run
Output
Run
Output
POST /v1/chat/completions, max_tokens 1 and max_completion_tokens 1
Run
Output
Run
Output
Run
Output
Run
Output
POST /v1/responses, max_output_tokens 1
Run
Output
Run
Output
Streaming, max_tokens 1
Run
Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body
Run
Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body
Run
Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body
Boundary on us.openai.gpt-6-astra, max_tokens 15 then 16
Run
Output
Run
Output
Controls, a larger max_tokens and the Anthropic models
Run
Output
Run
Output
Run
Output
Reasoning params on us.openai.gpt-6-astra (Claude Code sends
thinkingon its real turns), plus the Sonnet controlRun
Output
Run
Output
Run
Output
Run
Output
After (5545ca9)
Claude Code, /model us.xai.grok-4.6
In the running Claude Code session, type
/model us.xai.grok-4.6and press Enter. The pane shows (header, then the lines this step added)Type
Reply with the single word okand press Enter. The pane shows (header, then the lines this step added)Claude Code, /model us.openai.gpt-6-astra
In the running Claude Code session, type
/model us.openai.gpt-6-astraand press Enter. The pane shows (header, then the lines this step added)Type
Reply with the single word okand press Enter. The pane shows (header, then the lines this step added)POST /v1/messages, max_tokens 1
Run
Output
Run
Output
POST /v1/chat/completions, max_tokens 1 and max_completion_tokens 1
Run
Output
Run
Output
Run
Output
Run
Output
POST /v1/responses, max_output_tokens 1
Run
Output
Run
Output
Streaming, max_tokens 1
Run
Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body
Run
Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body
Run
Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body
Boundary on us.openai.gpt-6-astra, max_tokens 15 then 16
Run
Output
Run
Output
Controls, a larger max_tokens and the Anthropic models
Run
Output
Run
Output
Run
Output
Reasoning params on us.openai.gpt-6-astra (Claude Code sends
thinkingon its real turns), plus the Sonnet controlRun
Output
Run
Output
Run
Output
Run
Output
Observations from the run
Seen on the after leg. The first two follow from the floor this PR adds and do not affect the probe, which only needs the 200; the Haiku line shows Anthropic models are left alone; the rest are Claude Code behavior this PR neither causes nor changes
okturn became a Skill call; declined, re-sent/modelopens aSwitch model?dialogType
🐛 Bug Fix
Caveats (if any)
Low
openai.gpt-*andxai.grok-*models on Converse are clamped; inference-profile ARNs carrying the model id are matched, opaque application-inference-profile ARNs are notbedrock/openai/...(Mantle) andbedrock/invoke/...routes for the same models are untouched and still forward a sub-16 value as is, same as fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870max_tokensbelow 16 on these models now returns up to 16 tokens instead of a 400; no previously working request changes, since Bedrock rejected every value below 16openai.gpt-<digit>model id, so a future GPT-7 model on Converse gets the samereasoning.effortmapping without a further change, same as mainlocal_testing_part1,local_testing_part2,litellm_router_testing, ande2e_openai_endpointswere red the same way on fix(responses): backport mid-stream content_policy_violation fallback routing to stable/1.100.x #41208, the last PR merged intostable/1.100.x(Together AI 400s,cohere.command-r-plus-v1:0no longer served by Bedrock, OpenAI timeout tests, andtest_bad_request_error, where OpenAI now answers an unknown model on the Responses API with a 404 and the test expects a 400), andllm_translation_testingfails only intests/llm_translation/test_together_ai.py(18 Together AI 400s). None of them exercise the Bedrock Converse path;build_docker_database_imageand every other job are greenBlast radius, /live-pr-risk at 5545ca9 against base b7d81e9
Verdict: no dependent path that worked on the base answers worse on the head. Two no-DB proxies (2 workers each) ran 33 scenarios side by side against real Bedrock, each behind a forwarding recorder that logged what LiteLLM sent before re-signing it for AWS. Both sides sent the same header set, and the same number of upstream requests in every scenario but the raw
thinkingcase below, and the only body differences are the two this PR names:inferenceConfig.maxTokensraised to 16, and the GPT-6thinkingfield replaced byreasoning.effortBreaking: none observed
Backward incompatible, all benign, and all the behavior main already ships from #41870 and #31884 (both approved by mateo-berri on main):
max_tokensof0,-1, and JSONtrueonus.openai.gpt-6-astrawere Bedrock 400s on the base and are 200s on the head, sent asmaxTokens: 16.truepasses theisinstance(value, int)guard because a Python bool is an int."1",1.5, and[1]are forwarded unchanged and 400 identically on both sidesreasoning_efforton POST /v1/chat/completions, or athinkingblock on POST /v1/messages, were 400Unknown parameter: 'thinking'on the base and are 200 on the head; the upstream body now carriesadditionalModelRequestFields: {"reasoning": {"effort": "low"}}where the base sent{"thinking": {"type": "enabled", "budget_tokens": 1024}}, andadditionalModelResponseFieldPathsis no longer sent for these modelsthinkingparam on POST /v1/chat/completions for GPT-6 stays a 400, but it is now LiteLLM'sUnsupportedParamsErrorwith no upstream call where the base forwarded it and relayed Bedrock's 400thinkingandoutput_configinsupported_openai_paramsfor the GPT-6 model groups (the plain id, thebedrock/converse/prefix, and an ARN deployment withmodel_info.base_model);reasoning_effortstays. Grok, Haiku, Nova, and the ARN deployment withoutbase_modelare identical on both sides. The dashboard model hub only displays this list, so it shows two fewer params for these modelsRegression risk: spend rows for the floored requests were not read back, because the rig ran without a database; the usage blocks in the head's 200 responses reported 5 output tokens for GPT-6 and 16 for Grok
Dependency graph:
map_openai_paramsmaxTokensbranch via /v1/chat/completions (max_tokens,max_completion_tokens), /v1/messages, /v1/responses,converse-stream, thebedrock/converse/prefix, and a router fallback replay (2 upstream calls on both sides): unit tested and verified live, base 400 and head 200max_tokens16 and 17, Haiku at 1,openai.gpt-oss-20b-1:0at 1, unauthenticated 401 with 0 upstream calls): verified live, identical on both sidesAmazonInvokeNovaConfig, the only subclass ofAmazonConverseConfig: verified live onbedrock/invoke/us.amazon.nova-lite-v1:0,maxTokens: 1on both sidesmodel_info.base_model: unit tested as not floored and verified live, 400 on both sides, so this PR leaves that deployment shape broken the way it already wasreasoning_effort: verified live, still sent as Anthropic-stylethinkingand 200 on both sidespyproject.tomlanduv.lock: version bump only, no resolved package changesNot verified: spend and usage rows in the database, the Admin UI pages that read them, and the user docs in litellm-docs. The legacy
tests/llm_translationandtests/local_testingsuites ran on CircleCI and are covered in the bullet aboveFinal Attestation
Note
Medium Risk
Changes request shaping on a hot Bedrock Converse path (token limits and reasoning/thinking forwarding); scope is narrow to specific model id patterns with tests, but mis-matched model strings could still get wrong clamping or param drops.
Overview
Backports Bedrock Converse fixes for OpenAI-compat and GPT-6 models on
stable/1.100.x(release 1.100.2).On the Converse route,
max_tokens/max_completion_tokensare now raised to at least 16 when the model id matchesopenai.gpt-*orxai.grok-*, so low-token probes (e.g. Claude Code/modelswitches) no longer fail Bedrock’s minimum output-token validation. Other Bedrock models are unchanged.Reasoning handling no longer keys only on
openai.gpt-5: anyopenai.gpt-<digit>model is treated as an OpenAI reasoning model—reasoning_effortmaps toreasoning.effort, Anthropicthinkingis not forwarded, and supported-params reflect that (fixes GPT-6Unknown parameter: 'thinking'). Parametrized unit tests cover the token floor and GPT-6 reasoning paths.Reviewed by Cursor Bugbot for commit 5545ca9. Bugbot is set up for automated code reviews on this repo. Configure here.