Repository navigation
fix(bedrock): backport #41870, the GPT-6 reasoning gate fix, and the Python 3.13 image pin to stable/1.98.x for v1.98.1 - #42002
Conversation
…and xAI Grok models on Converse (#41870) Backport of #41870 to stable/1.98.x. Cherry-picked from a6e3a72 (main) with -m 1. converse_transformation.py conflicted because this line has no `import re` and no _is_openai_gpt_reasoning_model helper next to the insertion point. The resolution adds exactly the four hunks #41870 merged: the import, the 16-token constant, _requires_min_max_tokens, and the clamped maxTokens assignment. The test file applied clean.
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
| if param == "max_tokens" or param == "max_completion_tokens": | ||
| optional_params["maxTokens"] = value | ||
| optional_params["maxTokens"] = ( | ||
| max(value, BEDROCK_OPENAI_COMPAT_MIN_MAX_TOKENS) | ||
| if isinstance(value, int) and self._requires_min_max_tokens(model) | ||
| else value |
There was a problem hiding this comment.
Backport of #41870, where this was withdrawn. The regex clamps ARNs embedding the model id; opaque application profiles carry none, so get_model_info cannot resolve them
There was a problem hiding this comment.
That’s fair. Opaque application-profile ARNs contain no model identifier, and get_model_info cannot resolve the backing model from them, so there is no reliable way to apply this model-specific clamp without risking unrelated profiles. The regex intentionally covers only identifiable model IDs, including ARNs that embed them. I agree this finding should be withdrawn.
| @staticmethod | ||
| def _requires_min_max_tokens(model: str) -> bool: | ||
| return re.search(r"openai\.gpt-\d|xai\.grok-", model) is not None |
There was a problem hiding this comment.
This name-pattern check violates the directive to store model-specific flags in model_prices_and_context_window.json and read them through get_model_info; fix before merging
Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
Same shape as main's merged #41870. The pattern matches model families, not individual models, so new GPT or Grok releases need no cost-map entry
There was a problem hiding this comment.
That’s a fair distinction. I rechecked #41870: _requires_min_max_tokens is not identifying a finite catalog of model capabilities; it encodes Bedrock Converse’s shared minimum-token contract for the openai.gpt-* and xai.grok-* model families. Using get_model_info would require adding or updating every future model ID and would make the fix stale, whereas family matching intentionally covers new releases and mirrors the merged implementation on main. I’m withdrawing this finding; no change is needed.
…oning gate Backports the Converse part of fbc6fb5 from main (PR #31884) together with the OpenAI GPT reasoning gate this line never had (PR #38279: 74e86d3, 9cc276a, 418012a): reasoning_effort maps to reasoning.effort for openai.gpt-<digit> models and Anthropic's thinking block is skipped for them. Without it a GPT-6 model fell through to Anthropic's thinking block and Bedrock rejected the first real turn after a Claude Code /model switch with 400 Unknown parameter: 'thinking'. #38279's cost-map JSON and cross-region test changes, the Nova 2 tool_choice registry keys, and the invoke json_mode forwarding stay on main
| python-3.13 \ | ||
| python-3.13-dev \ |
There was a problem hiding this comment.
Migration image remains unpinned
The primary, backend, gateway, database, and non-root images now install Python 3.13 explicitly, but the separately shipped migrations/Dockerfile remains on the old Wolfi digest and installs the moving python3 metapackage. It can therefore select a different Python version, including an incompatible newer interpreter, causing the migration image to fail during build or startup. Apply the same base-image and Python 3.13 pin to the migration image.
…3 in migrations image (cherry picked from commit 3947374)
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 0c0dda9. Configure here.
TLDR
Problem this solves:
stable/1.98.xline hit the bug fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870 fixed on main: Claude Code's/modelswitch to a Bedrock OpenAI GPT or xAI Grok model fails with a 400, because Claude Code sends amax_tokens=1probe on switch, those Bedrock models require at least 16, and the Bedrock Converse path forwardedmaxTokensverbatimus.openai.gpt-6-astra) still fails with 400Unknown parameter: 'thinking', because the Converse reasoning gate on this line did not recognize GPT-6 models and Claude Code'sthinkingblock was mapped as Anthropic thinking on a model that rejects itstable/1.98.xthe OpenAI GPT reasoning gate did not exist at all (main got it for GPT-5 in fix(bedrock): map reasoning_effort to reasoning.effort for OpenAI GPT-5.x on Converse #38279), so the backport also adds thereasoning.effortmapping and itsBedrockConverseGptReasoningEffortBlocktypepython3, which now resolves to CPython 3.14, whereuvloop0.21.0 no longer builds;build_docker_database_imagewas red on this PR's first tip, so the images move to a pinnedpython-3.13the way main did in 728dec2 and, for the migrations image, 3947374How it solves it:
stable/1.98.x:maxTokensis clamped to 16 foropenai.gpt-*andxai.grok-*models on the Bedrock Converse route, other Bedrock models keep pass-through behavioropenai.gpt-<digit>model is recognized, soreasoning_effortmaps toreasoning.effortand an Anthropicthinkingblock is dropped for these models instead of being forwarded to Bedrock1.98.1Backport notes
-m 1, so the branch carries fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870's source and test changes as one commitlitellm/llms/bedrock/chat/converse_transformation.pyconflicted in that cherry-pick because this line never hadimport reor the_is_openai_gpt_reasoning_modelhelper that fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870's hunks sit next to on main; the resolution addsimport re, theBEDROCK_OPENAI_COMPAT_MIN_MAX_TOKENSconstant, the_requires_min_max_tokenshelper, and the clampedmaxTokensassignment exactly as they merged on main, nothing elsetests/test_litellm/llms/bedrock/chat/test_converse_transformation.pyapplied clean, so the ten parametrized cases from fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870 run unchanged on this line_is_openai_gpt_reasoning_modelhelper matchingopenai\.gpt-\dand its three call sites (thereasoning_effortmapping,get_supported_openai_params, and thethinkingskip inmap_openai_params). The Nova 2tool_choiceregistry keys and the invokejson_modeforwarding in that commit stay on main. The added parametrized test cases come from the same commit and fail on this line without the gate changereasoning_efforttoreasoning.effortbranch and theBedrockConverseGptReasoningEffortBlockTypedDict inlitellm/types/llms/bedrock.py. fix(bedrock): map reasoning_effort to reasoning.effort for OpenAI GPT-5.x on Converse #38279's cost-map JSON and cross-region test changes stay on main. The ported tests drop thelocal_model_cost_mapfixture parameter, which this line's conftest does not define-x) so every Dockerfile, the componentizedmigrations/Dockerfileincluded, installspython-3.13andpython-3.13-devwithUV_PYTHON_DOWNLOADS=0, resolving the base image digest lines to the values those commits set; the only other difference between this line's Dockerfiles and those commits' parents was those digest lines1.98.1is its own commit, one patch above the latest release on this line (1.98.0)User Flow
Before: a developer using Claude Code through a v1.98.0 proxy cannot use a Bedrock OpenAI GPT or xAI Grok model
ANTHROPIC_BASE_URLpointed at the proxy and type/model us.xai.grok-4.6(or/model us.openai.gpt-6-astra)"max_tokens": 1as a warmup probeAPI error: 400 ... BedrockException ... Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.and the header keeps the model they had (Sonnet 4.5)us.openai.gpt-6-astratakes, and their first message, which Claude Code sends to POST https://litellm-domain/v1/messages with athinkingblock, comes back asAPI Error: 400 ... Unknown parameter: 'thinking'After: the same switch and first message succeed once the proxy runs v1.98.1
ANTHROPIC_BASE_URLpointed at the proxy and type/model us.xai.grok-4.6(or/model us.openai.gpt-6-astra)"max_tokens": 1as a warmup probethinkingblock and is answered by that Bedrock modelRelevant issues
Backport of #41870 and of the Converse reasoning gate fix from #31884 (main). Reported by customers (Pylon #8817 for OpenAI GPT, Pylon #8821 for xAI Grok); one of them runs v1.98.0 today and is moving to the 1.100.x line
Affected release
Linear ticket
Resolves LIT-8154
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Both legs ran the same way and differ only in the commit the proxy was booted from: a no-DB proxy with 2 uvicorn workers on a random port, real Bedrock in us-east-1 (real spend, no mocks), and Claude Code 2.1.278 driven interactively under tmux.
$KEYis the proxy master key and$PORTis the leg's portconfig.yamlProxy, booted from the leg's commit with a Python 3.12 venv
Claude Code, started in an empty folder with a fresh config directory (theme picker and security notes accepted with Enter, folder trust accepted with Down then Enter)
tmux new-session -d -s claude -x 160 -y 45 -c work "env -u ANTHROPIC_API_KEY ANTHROPIC_BASE_URL=http://localhost:$PORT ANTHROPIC_AUTH_TOKEN=$KEY ANTHROPIC_MODEL=us.anthropic.claude-sonnet-4-5-20250929-v1:0 ANTHROPIC_DEFAULT_HAIKU_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0 ANTHROPIC_SMALL_FAST_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0 CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 CLAUDE_CONFIG_DIR=claude-config claude"Before (d8f71d7)
The last curl group (the four reasoning cases) ran in a second proxy boot at this same commit after the floor-only after leg surfaced the
thinking400; every other case below is from the original before legClaude Code, /model us.xai.grok-4.6
In the running Claude Code session, type
/model us.xai.grok-4.6and press Enter. The pane shows (header, then the lines this step added)Type
Reply with the single word okand press Enter. The pane shows (header, then the lines this step added)Claude Code, /model us.openai.gpt-6-astra
In the running Claude Code session, type
/model us.openai.gpt-6-astraand press Enter. The pane shows (header, then the lines this step added)Type
Reply with the single word okand press Enter. The pane shows (header, then the lines this step added)POST /v1/messages, max_tokens 1
Run
Output
Run
Output
POST /v1/chat/completions, max_tokens 1 and max_completion_tokens 1
Run
Output
Run
Output
Run
Output
Run
Output
POST /v1/responses, max_output_tokens 1
Run
Output
Run
Output
Streaming, max_tokens 1
Run
Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body
Run
Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body
Run
Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body
Boundary on us.openai.gpt-6-astra, max_tokens 15 then 16
Run
Output
Run
Output
Controls, a larger max_tokens and the Anthropic models
Run
Output
Run
Output
Run
Output
Reasoning params on us.openai.gpt-6-astra (Claude Code sends
thinkingon its real turns), plus the Sonnet controlRun
Output
Run
Output
Run
Output
Run
Output
After (0c0dda9)
Claude Code, /model us.xai.grok-4.6
In the running Claude Code session, type
/model us.xai.grok-4.6and press Enter. The pane shows (header, then the lines this step added)Type
Reply with the single word okand press Enter. The pane shows (header, then the lines this step added)Claude Code, /model us.openai.gpt-6-astra
In the running Claude Code session, type
/model us.openai.gpt-6-astraand press Enter. The pane shows (header, then the lines this step added)Type
Reply with the single word okand press Enter. The pane shows (header, then the lines this step added)POST /v1/messages, max_tokens 1
Run
Output
Run
Output
POST /v1/chat/completions, max_tokens 1 and max_completion_tokens 1
Run
Output
Run
Output
Run
Output
Run
Output
POST /v1/responses, max_output_tokens 1
Run
Output
Run
Output
Streaming, max_tokens 1
Run
Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body
Run
Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body
Run
Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body
Boundary on us.openai.gpt-6-astra, max_tokens 15 then 16
Run
Output
Run
Output
Controls, a larger max_tokens and the Anthropic models
Run
Output
Run
Output
Run
Output
Reasoning params on us.openai.gpt-6-astra (Claude Code sends
thinkingon its real turns), plus the Sonnet controlRun
Output
Run
Output
Run
Output
Run
Output
Observations from the run
Seen on the after leg. The first two follow from the floor this PR adds and do not affect the probe, which only needs the 200; the control line shows Anthropic models are left alone; the rest are Claude Code behavior this PR neither causes nor changes
okturn became a Skill call; declined, re-sent/modelopens aSwitch model?dialogType
🐛 Bug Fix
Caveats (if any)
Low
openai.gpt-*andxai.grok-*models on Converse are clamped; inference-profile ARNs carrying the model id are matched, opaque application-inference-profile ARNs are notbedrock/openai/...(Mantle) andbedrock/invoke/...routes for the same models are untouched and still forward a sub-16 value as is, same as fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870max_tokensbelow 16 on these models now returns up to 16 tokens instead of a 400; no previously working request changes, since Bedrock rejected every value below 16openai.gpt-<digit>model id, so a future GPT-7 model on Converse gets the samereasoning.effortmapping without a further change, same as mainllm_translation_testing(Groq, Together AI, and OpenAI web search cases),local_testing_part1andlocal_testing_part2(Together AI,cohere.command-r-plus-v1:0no longer served by Bedrock, OpenAI timeout tests, a Groq stream chunk case),litellm_router_testing(wildcard routing, timeout, and batch completion against live providers),logging_testing(OpenAI web search cost tracking),image_gen_testing(Google and Vertex image generation), andpass_through_unit_testing(the assistants passthrough case against api.openai.com)not_runwhilebuild_docker_database_imagewas red and only started running once thepython-3.13pin landed:e2e_openai_endpointsfailstest_bad_request_error(OpenAI now answers an unknown model on the Responses API with a 404 where the test expects a 400; the same test is red on fix(bedrock): backport #41870 and the GPT-6 reasoning gate fix to stable/1.100.x for v1.100.2 #42000 and on fix(responses): backport mid-stream content_policy_violation fallback routing to stable/1.100.x #41208, the last PR merged intostable/1.100.x),proxy_pass_through_endpoint_testsfails its two Ruby OpenAI Assistants specs on a 404 from api.openai.com, andbuild_and_teststops attest_proxy_all_modelsbecause Groq no longer servesllama-3.1-8b-instantbuild_docker_database_imagewas red at the first tip (the unpinnedwolfi-baseimage now ships CPython 3.14, whereuvloop0.21.0 does not build) and is green at this tip with thepython-3.13pin, so CI verifies that image.migrations/Dockerfile,backend/Dockerfile,gateway/Dockerfile, the rootDockerfile, anddocker/Dockerfile.non_root(the last two with the dashboard stage stubbed) were also built locally on arm64 at this tip and import cleanly on Python 3.13.15; only the real dashboard build and the amd64 builds are first built by the release workflowBlast radius, /live-pr-risk at 0c0dda9 against base d8f71d7
Verdict: no dependent path that worked on the base answers worse on the head. Two no-DB proxies (2 workers each) ran 31 scenarios side by side against real Bedrock, and six more cases ran through
litellm.completionon each tree, every upstream request passing a forwarding recorder that logged what LiteLLM sent before re-signing it for AWS. Both sides sent the same header set, and the same number of upstream requests in every scenario but the rawthinkingcase below, and the only body differences are the two this PR names:inferenceConfig.maxTokensraised to 16, and the GPTthinkingfield replaced byreasoning.effortBreaking: none observed
Backward incompatible, all benign, and all the behavior main already ships from #41870, #31884, and #38279 (each approved by mateo-berri on main):
max_tokensof0,-1, and JSONtrueonus.openai.gpt-6-astrawere Bedrock 400s on the base and are 200s on the head, sent asmaxTokens: 16.truepasses theisinstance(value, int)guard because a Python bool is an int."1",1.5, and[1]are forwarded unchanged and 400 identically on both sidesreasoning_efforton POST /v1/chat/completions, or athinkingblock on POST /v1/messages, were 400Unknown parameter: 'thinking'on the base and are 200 on the head, andus.openai.gpt-5.6-lunawithreasoning_effortthrough the SDK goes the same way; the upstream body now carriesadditionalModelRequestFields: {"reasoning": {"effort": "low"}}where the base sent{"thinking": {"type": "enabled", "budget_tokens": 1024}}thinkingparam on a chat completion for GPT-6 or GPT-5.6 stays a 400, but it is now LiteLLM'sUnsupportedParamsErrorwith no upstream call where the base forwarded it and relayed Bedrock's 400thinkingandoutput_configinsupported_openai_paramsfor the GPT model groups (the plain id, thebedrock/converse/prefix, and an ARN deployment withmodel_info.base_model);reasoning_effortstays. Grok, Haiku, Nova, gpt-oss, Sonnet, and the ARN deployment withoutbase_modelare identical on both sides. The dashboard model hub only displays this list, so it shows two fewer params for these modelsRegression risk:
wolfi-basedigest and a pinned CPython 3.13. CircleCI built the database image at this tip and ran its proxy suites on it. This run also built the other five locally on arm64 at this tip:migrations/Dockerfile,backend/Dockerfile, andgateway/Dockerfileas is, and the rootDockerfileanddocker/Dockerfile.non_rootwith only their dashboardnpm run buildstage replaced by a stub page. All five report Python 3.13.15, litellm 1.98.1,uvloop0.21.0, and prisma 0.11.0, the backend and gateway images importbackend.mainandgateway.main, and the main image booted with no database and answered GET /health/liveliness with 200. The real dashboard build and every amd64 build are first built by the release workflow. On the base the database image did not build (uvloop0.21.0 on CPython 3.14), and the others install the same unpinnedpython3Dependency graph:
map_openai_paramsmaxTokensbranch via /v1/chat/completions (max_tokens,max_completion_tokens), /v1/messages, /v1/responses,converse-stream, thebedrock/converse/prefix, a router fallback replay (2 upstream calls on both sides), andus.openai.gpt-5.6-lunathrough the SDK: unit tested and verified live, base 400 and head 200_is_openai_gpt_reasoning_modelat its three call sites (reasoning_effortmapping,get_supported_openai_params, thethinkingskip): unit tested and verified live on GPT-6 and GPT-5.6max_tokens16 and 17, Haiku at 1,openai.gpt-oss-20b-1:0at 1 and withreasoning_effort, Sonnet 4.5 withreasoning_effort, unauthenticated 401 with 0 upstream calls): verified live, identical on both sidesAmazonInvokeNovaConfig, the only subclass ofAmazonConverseConfig: verified live onbedrock/invoke/us.amazon.nova-lite-v1:0,maxTokens: 1on both sidesmodel_info.base_model: unit tested as not floored and verified live, 400 on both sides, so this PR leaves that deployment shape broken the way it already wasreasoning_effort: verified live, still sent as Anthropic-stylethinkingand 200 on both sidesBedrockConverseGptReasoningEffortBlockinlitellm/types/llms/bedrock.py: a new type with one user, the mapping abovepyproject.tomlanduv.lock: version bump only, no resolved package changesNot verified: spend and usage rows in the database, the Admin UI pages that read them, the real dashboard build inside the Docker images and any amd64 build outside CircleCI's database image, and the user docs in litellm-docs. The legacy
tests/llm_translationandtests/local_testingsuites ran on CircleCI and are covered in the bullets aboveFinal Attestation