Repository navigation
fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse - #41870
Conversation
…models on Converse Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
| optional_params["maxTokens"] = ( | ||
| self._enforce_min_max_tokens(cast(object, value)) | ||
| if self._is_openai_gpt_reasoning_model(model) |
There was a problem hiding this comment.
Opaque inference-profile ARNs bypass this regex, leaving sub-16 limits unchanged and causing Bedrock 400s. Model-specific constraints must come from metadata via get_model_info
Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)
There was a problem hiding this comment.
Substring regex already clamps ARNs carrying the model id (tests added). Opaque application profiles carry no model id, so get_model_info cannot resolve them either
There was a problem hiding this comment.
You're right. The regex searches for the model identifier anywhere in the input, so inference-profile ARNs that embed openai.gpt-* or xai.grok-* are already covered, as the added tests demonstrate. Opaque application inference-profile ARNs contain no model identifier, so get_model_info cannot reliably resolve their underlying model without an external mapping. This is not an actionable gap; I withdraw the previous finding.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…nverse Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 12120fe. Configure here.
fix(bedrock): backport #41870 and the GPT-6 reasoning gate fix to stable/1.100.x for v1.100.2
fix(bedrock): backport #41870, the GPT-6 reasoning gate fix, and the Python 3.13 image pin to stable/1.98.x for v1.98.1
chore(release): backport #41870 to stable/1.101.x
TLDR
Problem this solves:
/modelswitch to a Bedrock OpenAI GPT or xAI Grok model fails with a 400max_tokens=1probe on switch and those Bedrock models require at least 16maxTokensverbatim, so the probe reached Bedrock as 1How it solves it:
maxTokensto 16 foropenai.gpt-*andxai.grok-*models on the Bedrock Converse routeUser Flow
Before: a developer using Claude Code through the proxy cannot switch to a Bedrock OpenAI GPT or xAI Grok model
ANTHROPIC_BASE_URLpointed at the proxy and type/model us.xai.grok-4.6(or/model us.openai.gpt-6-astra)"max_tokens": 1as a warmup probeAPI error: 400 ... BedrockException ... Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.After: the same switch succeeds
ANTHROPIC_BASE_URLpointed at the proxy and type/model us.xai.grok-4.6(or/model us.openai.gpt-6-astra)"max_tokens": 1as a warmup probeRelevant issues
Reported by customers (Pylon #8817 for OpenAI GPT, Pylon #8821 for xAI Grok)
Affected release
Linear ticket
Resolves LIT-8136
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Both legs boot a LiteLLM proxy from the named commit with 2 uvicorn workers and no DB, and call real Bedrock in us-east-1 with real credentials. The model list maps the bare model ids
us.openai.gpt-6-astra,us.xai.grok-4.6,us.anthropic.claude-sonnet-4-5-20250929-v1:0, andus.anthropic.claude-haiku-4-5-20251001-v1:0tobedrock/<same id>. Claude Code v2.1.277 is driven interactively under tmux withANTHROPIC_BASE_URLpointed at the proxy andANTHROPIC_MODELset to the Sonnet entry, so/modelswitches are the customer's exact flow.KEY=<redacted>is the proxy master key. The grok responses carryredacted_thinkingbase64 payloads over 1 KB each, elided below as…; everything else is verbatimBefore (2edda5a, the merge base; proxy on port 41217)
Claude Code,
/model us.xai.grok-4.6thenReply with the single word ok: the switch shows the Bedrock 400 from themax_tokens: 1probe, the header stays on Sonnet 4.5, and theokis served by SonnetClaude Code,
/model us.openai.gpt-6-astrathenReply with the single word ok: same 400, same silent fallback to SonnetThe proxy access log shows both probes as
POST /v1/messages?beta=true400 (17:21:38 routed to us.xai.grok-4.6, 17:22:47 to us.openai.gpt-6-astra) and bothokreplies served by us.anthropic.claude-sonnet-4-5-20250929-v1:0POST /v1/messages,
max_tokens: 1(Claude Code's probe shape): 400 on both OpenAI-compat models, 200 withoutput_tokens: 1on SonnetPOST /v1/chat/completions,
max_tokens: 1thenmax_completion_tokens: 1: 400 on both models, both param namesPOST /v1/responses,
max_output_tokens: 1: 400 on both modelsStreaming,
max_tokens: 1: the same 400 body, no SSEBoundary on us.openai.gpt-6-astra:
max_tokens: 15is rejected,max_tokens: 16is acceptedControls:
max_tokens: 64on us.xai.grok-4.6 (200,completion_tokens: 64) andmax_tokens: 1on Haiku (200,output_tokens: 1)After (12120fe, the PR tip; proxy on port 43851)
Claude Code,
/model us.xai.grok-4.6thenReply with the single word ok: the switch is accepted, the header shows us.xai.grok-4.6, and theokcomes from grokClaude Code,
/model us.openai.gpt-6-astra(Enter on the "Switch model?" confirmation) thenReply with the single word ok: accepted, served by gpt-6-astraThe proxy access log shows both probes as
POST /v1/messages?beta=true200 (17:19:56 routed to us.xai.grok-4.6, 17:21:37 to us.openai.gpt-6-astra); the whole leg has no 400POST /v1/messages,
max_tokens: 1: 200 on both OpenAI-compat models (gpt-6-astra stops on its own at 13 tokens, grok-4.6 hits the 16-token floor), Sonnet still returns exactlyoutput_tokens: 1POST /v1/chat/completions,
max_tokens: 1thenmax_completion_tokens: 1: 200 on both models, both param namesPOST /v1/responses,
max_output_tokens: 1: 200 on both modelsStreaming,
max_tokens: 1: 200 with a usage chunk of 16 (grok) and 13 (gpt) on chat completions, and amessage_deltawithoutput_tokens: 16on /v1/messagesBoundary on us.openai.gpt-6-astra:
max_tokens: 15now returns 200 (clamped to 16, the model stopped at 11),max_tokens: 16unchangedControls:
max_tokens: 64on us.xai.grok-4.6 still returnscompletion_tokens: 64(the clamp is a floor, not an override) andmax_tokens: 1on Haiku still returnsoutput_tokens: 1Observations from the run, none caused by this PR:
Type
🐛 Bug Fix
Caveats (if any)
Low
openai.gpt-*andxai.grok-*models on Converse are clamped; inference-profile ARNs carrying the model id are matched, opaque application-inference-profile ARNs are not (they carry no model id, so metadata lookup cannot resolve them either)bedrock/openai/...(Mantle) andbedrock/invoke/...routes for the same models are untouched and still forward a sub-16 value as is; the default route for these model ids is Converse, which is what the reports use, and each of those routes would need its own QA legmax_tokensbelow 16 on these models now returns up to 16 tokens instead of a 400. No previously working request changes, since Bedrock rejected every value below 16, and the caller's value still applies at 16 and above (the 64-token control above)test_models_by_providerin litellm_utils_testing (red since feat(proxy): add Amazon Transcribe pass-through with completion-time job pricing #41515), the two retry-count integration tests in integration-management and integration-providers (red since refactor(proxy): make the config file win over the database #41779), andtest_bedrock_invoke_messages_with_all_beta_headersin proxy_e2e_anthropic_messages_tests (LIT-8149, red since fix(anthropic): register thinking-binding-controls-2026-08-01 in beta headers config #41203).proxy-infra / Run testsis not a required check and is red on main tooFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/0eed92dc9c7f470a82a5781cbf03e622
Open in Devin Desktop: https://app.devin.ai/desktop/session/0eed92dc9c7f470a82a5781cbf03e622?variant=devin
Note
Low Risk
Localized parameter mapping in Bedrock Converse; only raises sub-16 limits for specific model ID patterns, with tests covering edge cases.
Overview
Bedrock Converse now enforces a 16-token floor on
maxTokenswhen mapping OpenAImax_tokens/max_completion_tokensfor Bedrock models whose IDs matchopenai.gpt-*orxai.grok-*, including typical inference-profile ARNs that embed those IDs. Values at or above 16 are unchanged; other Bedrock models still pass through as before.This fixes 400 errors when clients (e.g. Claude Code’s
max_tokens=1model-switch probe) hit Bedrock’s minimum output-token requirement on those OpenAI-compat and Grok routes. A new_requires_min_max_tokenshelper andBEDROCK_OPENAI_COMPAT_MIN_MAX_TOKENSconstant drive the clamp; parametrized unit tests cover GPT/Grok vs Claude, both param names, and ARN shapes.Reviewed by Cursor Bugbot for commit 12120fe. Bugbot is set up for automated code reviews on this repo. Configure here.