Skip to content

fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse - #41870

Merged
mateo-berri merged 4 commits into
mainfrom
litellm_bedrock_openai_gpt_min_max_tokens
Sep 19, 2026
Merged

mateo-berri merged 4 commits into
mainfrom
litellm_bedrock_openai_gpt_min_max_tokens

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Claude Code's /model switch to a Bedrock OpenAI GPT or xAI Grok model fails with a 400
  • Claude Code sends a max_tokens=1 probe on switch and those Bedrock models require at least 16
  • The Bedrock Converse path forwarded maxTokens verbatim, so the probe reached Bedrock as 1

How it solves it:

  • Clamp maxTokens to 16 for openai.gpt-* and xai.grok-* models on the Bedrock Converse route
  • Mirrors the clamp the OpenAI Responses path already has (fix(openai/responses): clamp max_output_tokens below API minimum #33098)
  • Chosen over dropping the param (every request would run uncapped) and over failing fast with a clearer 400 (the probe would still fail)
  • Other Bedrock models keep pass-through behavior

User Flow

Before: a developer using Claude Code through the proxy cannot switch to a Bedrock OpenAI GPT or xAI Grok model

  1. They run Claude Code with ANTHROPIC_BASE_URL pointed at the proxy and type /model us.xai.grok-4.6 (or /model us.openai.gpt-6-astra)
  2. Claude Code sends POST https://litellm-domain/v1/messages with "max_tokens": 1 as a warmup probe
  3. They see API error: 400 ... BedrockException ... Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.
  4. The model switch is unusable until they pick a different model

After: the same switch succeeds

  1. They run Claude Code with ANTHROPIC_BASE_URL pointed at the proxy and type /model us.xai.grok-4.6 (or /model us.openai.gpt-6-astra)
  2. Claude Code sends POST https://litellm-domain/v1/messages with "max_tokens": 1 as a warmup probe
  3. The probe returns 200 with a short response
  4. They keep chatting on the Bedrock model

Relevant issues

Reported by customers (Pylon #8817 for OpenAI GPT, Pylon #8821 for xAI Grok)

Affected release

Linear ticket

Resolves LIT-8136

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Both legs boot a LiteLLM proxy from the named commit with 2 uvicorn workers and no DB, and call real Bedrock in us-east-1 with real credentials. The model list maps the bare model ids us.openai.gpt-6-astra, us.xai.grok-4.6, us.anthropic.claude-sonnet-4-5-20250929-v1:0, and us.anthropic.claude-haiku-4-5-20251001-v1:0 to bedrock/<same id>. Claude Code v2.1.277 is driven interactively under tmux with ANTHROPIC_BASE_URL pointed at the proxy and ANTHROPIC_MODEL set to the Sonnet entry, so /model switches are the customer's exact flow. KEY=<redacted> is the proxy master key. The grok responses carry redacted_thinking base64 payloads over 1 KB each, elided below as …; everything else is verbatim

Before (2edda5a, the merge base; proxy on port 41217)

Claude Code, /model us.xai.grok-4.6 then Reply with the single word ok: the switch shows the Bedrock 400 from the max_tokens: 1 probe, the header stays on Sonnet 4.5, and the ok is served by Sonnet

 ▐▛███▛█   Claude Code v2.1.277
▝▜██████▀  Sonnet 4.5 · API Usage Billing
  ▝▝ ▝▝    …/qa-before/work
❯ /model us.xai.grok-4.6
  ⎿  API error: 400 {"type":"error","error":{"type":"invalid_request_error","message":"litellm.BadRequestError: BedrockException - {"message":"The model
     returned the following errors: {\"error\":{\"code\":\"integer_below_min_value\",\"message\":\"Invalid 'max_output_tokens': integer below minimum value.
     Expected a value >= 16, but got 1 instead.\",\"param\":\"max_output_tokens\",\"type\":\"invalid_request_error\"}}"}. Received Model
     Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None"}}
❯ Reply with the single word ok
  Thought for 5s
⏺ ok
✻ Worked for 5s · done 5:22 PM

pr41870-12120fe59b-claude-before-grok.png

Claude Code, /model us.openai.gpt-6-astra then Reply with the single word ok: same 400, same silent fallback to Sonnet

❯ /model us.openai.gpt-6-astra
  ⎿  API error: 400 {"type":"error","error":{"type":"invalid_request_error","message":"litellm.BadRequestError: BedrockException - {"message":"The model
     returned the following errors: {\"error\":{\"code\":\"integer_below_min_value\",\"message\":\"Invalid 'max_output_tokens': integer below minimum value.
     Expected a value >= 16, but got 1 instead.\",\"param\":\"max_output_tokens\",\"type\":\"invalid_request_error\"}}"}. Received Model
     Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None"}}
❯ Reply with the single word ok
  Thought for 3s
⏺ ok
✻ Sautéed for 3s · done 5:23 PM

pr41870-12120fe59b-claude-before-gpt.png

The proxy access log shows both probes as POST /v1/messages?beta=true 400 (17:21:38 routed to us.xai.grok-4.6, 17:22:47 to us.openai.gpt-6-astra) and both ok replies served by us.anthropic.claude-sonnet-4-5-20250929-v1:0

POST /v1/messages, max_tokens: 1 (Claude Code's probe shape): 400 on both OpenAI-compat models, 200 with output_tokens: 1 on Sonnet

KEY=<redacted>
$ curl -s http://localhost:41217/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"type":"error","error":{"type":"invalid_request_error","message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None"}}
HTTP 400

$ curl -s http://localhost:41217/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"type":"error","error":{"type":"invalid_request_error","message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None"}}
HTTP 400

$ curl -s http://localhost:41217/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","id":"msg_bdrk_011CfBqgEZAJenEFNWv42dqs","type":"message","role":"assistant","content":[{"type":"text","text":"Hello"}],"stop_reason":"max_tokens","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":8,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1}}
HTTP 200

POST /v1/chat/completions, max_tokens: 1 then max_completion_tokens: 1: 400 on both models, both param names

KEY=<redacted>
$ curl -s http://localhost:41217/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":null,"code":"400"}}
HTTP 400

$ curl -s http://localhost:41217/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":null,"code":"400"}}
HTTP 400

$ curl -s http://localhost:41217/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_completion_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":null,"code":"400"}}
HTTP 400

$ curl -s http://localhost:41217/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_completion_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":null,"code":"400"}}
HTTP 400

POST /v1/responses, max_output_tokens: 1: 400 on both models

KEY=<redacted>
$ curl -s http://localhost:41217/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_output_tokens":1,"input":"hi"}' -w '\nHTTP %{http_code}\n'
{"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":null,"code":"400"}}
HTTP 400

$ curl -s http://localhost:41217/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_output_tokens":1,"input":"hi"}' -w '\nHTTP %{http_code}\n'
{"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":null,"code":"400"}}
HTTP 400

Streaming, max_tokens: 1: the same 400 body, no SSE

KEY=<redacted>
$ curl -s http://localhost:41217/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":null,"code":"400"}}
HTTP 400

$ curl -s http://localhost:41217/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":null,"code":"400"}}
HTTP 400

$ curl -s http://localhost:41217/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"stream":true,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"type":"error","error":{"type":"invalid_request_error","message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None"}}
HTTP 400

Boundary on us.openai.gpt-6-astra: max_tokens: 15 is rejected, max_tokens: 16 is accepted

KEY=<redacted>
$ curl -s http://localhost:41217/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":15,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 15 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":null,"code":"400"}}
HTTP 400

$ curl -s http://localhost:41217/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":16,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"id":"chatcmpl-ac25e554-a383-4820-a89f-b14aee496d6c","created":1789777242,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help you today?","role":"assistant"}}],"usage":{"completion_tokens":13,"prompt_tokens":7,"total_tokens":20,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":13},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

Controls: max_tokens: 64 on us.xai.grok-4.6 (200, completion_tokens: 64) and max_tokens: 1 on Haiku (200, output_tokens: 1)

KEY=<redacted>
$ curl -s http://localhost:41217/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":64,"messages":[{"role":"user","content":"Write one sentence about the sea."}]}' -w '\nHTTP %{http_code}\n'
{"id":"chatcmpl-89832d5d-d3ae-4417-b76b-6f5aea2ec01d","created":1789777211,"model":"us.xai.grok-4.6","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"","role":"assistant","reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContentBlocks":[{"redactedContent":"…"}]}}}],"usage":{"completion_tokens":64,"prompt_tokens":25,"total_tokens":89,"completion_tokens_details":{},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":25,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

$ curl -s http://localhost:41217/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-haiku-4-5-20251001-v1:0","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'
{"model":"us.anthropic.claude-haiku-4-5-20251001-v1:0","id":"msg_bdrk_lmvxeikxaztsucb23lwpgr4wzcsiltccftryapyzkz6wgml6ipbq","type":"message","role":"assistant","content":[{"type":"text","text":"#"}],"container":null,"stop_reason":"max_tokens","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":8,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard"}}
HTTP 200

After (12120fe, the PR tip; proxy on port 43851)

Claude Code, /model us.xai.grok-4.6 then Reply with the single word ok: the switch is accepted, the header shows us.xai.grok-4.6, and the ok comes from grok

 ▐▛███▛█   Claude Code v2.1.277
▝▜██████▀  us.xai.grok-4.6 · API Usage Billing
  ▝▝ ▝▝    …/qa-after/work
❯ /model us.xai.grok-4.6
  ⎿  Set model to us.xai.grok-4.6 and saved as your default for new sessions
          ANTHROPIC_MODEL is set to Sonnet 4.5 — new sessions use that while it is set
❯ Reply with the single word ok
⏺ ok
✻ Brewed for 5s · done 5:20 PM

pr41870-12120fe59b-claude-after-grok.png

Claude Code, /model us.openai.gpt-6-astra (Enter on the "Switch model?" confirmation) then Reply with the single word ok: accepted, served by gpt-6-astra

❯ /model us.openai.gpt-6-astra
  ⎿  Set model to us.openai.gpt-6-astra and saved as your default for new sessions
          ANTHROPIC_MODEL is set to Sonnet 4.5 — new sessions use that while it is set
❯ Reply with the single word ok
⏺ ok
✻ Worked for 9s · done 5:23 PM

pr41870-12120fe59b-claude-after-gpt.png

The proxy access log shows both probes as POST /v1/messages?beta=true 200 (17:19:56 routed to us.xai.grok-4.6, 17:21:37 to us.openai.gpt-6-astra); the whole leg has no 400

POST /v1/messages, max_tokens: 1: 200 on both OpenAI-compat models (gpt-6-astra stops on its own at 13 tokens, grok-4.6 hits the 16-token floor), Sonnet still returns exactly output_tokens: 1

KEY=<redacted>
$ curl -s http://localhost:43851/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' | jq -c .
{"id":"chatcmpl-4711d281-4c4f-4418-8a05-1106e910702f","type":"message","role":"assistant","model":"us.openai.gpt-6-astra","stop_sequence":null,"usage":{"input_tokens":7,"output_tokens":13},"content":[{"type":"text","text":"Hi! How can I help you today?"}],"stop_reason":"end_turn","stop_details":null}
HTTP 200

$ curl -s http://localhost:43851/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' | jq -c .
{"id":"chatcmpl-157f9bd4-e005-42d8-8617-b32968ae4d54","type":"message","role":"assistant","model":"us.xai.grok-4.6","stop_sequence":null,"usage":{"input_tokens":19,"output_tokens":16},"content":[{"type":"redacted_thinking","data":"…"},{"type":"text","text":""}],"stop_reason":"max_tokens","stop_details":null}
HTTP 200

$ curl -s http://localhost:43851/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' | jq -c .
{"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","id":"msg_bdrk_011CfBqam9kapoKY3tAsXJNb","type":"message","role":"assistant","content":[{"type":"text","text":"Hello"}],"stop_reason":"max_tokens","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":8,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1}}
HTTP 200

POST /v1/chat/completions, max_tokens: 1 then max_completion_tokens: 1: 200 on both models, both param names

KEY=<redacted>
$ curl -s http://localhost:43851/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' | jq -c .
{"id":"chatcmpl-cc14dbb8-efbc-4b9c-8de3-622c89c3499b","created":1789777109,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help you today?","role":"assistant"}}],"usage":{"completion_tokens":13,"prompt_tokens":7,"total_tokens":20,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":13},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

$ curl -s http://localhost:43851/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' | jq -c .
{"id":"chatcmpl-1f94a924-945d-4661-b1f8-d38d4b6c7686","created":1789777110,"model":"us.xai.grok-4.6","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"","role":"assistant","reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContentBlocks":[{"redactedContent":"…"}]}}}],"usage":{"completion_tokens":16,"prompt_tokens":19,"total_tokens":35,"completion_tokens_details":{},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":19,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

$ curl -s http://localhost:43851/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_completion_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' | jq -c .
{"id":"chatcmpl-a37d1a6c-321d-4a29-8593-2fd39adb8fb6","created":1789777119,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help you today?","role":"assistant"}}],"usage":{"completion_tokens":13,"prompt_tokens":7,"total_tokens":20,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":13},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

$ curl -s http://localhost:43851/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_completion_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' | jq -c .
{"id":"chatcmpl-c579e712-778a-49f4-be10-df549f7e4fa2","created":1789777119,"model":"us.xai.grok-4.6","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"","role":"assistant","reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContentBlocks":[{"redactedContent":"…"}]}}}],"usage":{"completion_tokens":16,"prompt_tokens":19,"total_tokens":35,"completion_tokens_details":{},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":19,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

POST /v1/responses, max_output_tokens: 1: 200 on both models

KEY=<redacted>
$ curl -s http://localhost:43851/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_output_tokens":1,"input":"hi"}' -w '\nHTTP %{http_code}\n' | jq -c .
{"id":"resp_-d-WbWy2whYKIXGTyZi1kddYmn7osmfJH0AMhiEH7kWGmj9DMmXPRmAquw6kOfjb4vdJQYl-d3WNjNQK01K2ftFlLZFY7RIN5Ka2XZsWAQIoU0jvaW2qujfHzmQZPiRChnsTPrllFGJjFNA-0prFO6yXzG8psOJPC1AfAmuEeGqrUEOMJLpLDEOBwRfu_9PrcxwQ7Fwpk1LpkEekRt2bJAMvtG9CuzTAJmrCmg09LkXgKxAk8xdCaVHw2RWNP3FCufNN8gcuwpclkYAqvWM3at6lG1USkdk_kGheXOT2E-slSUo_HFcoQVqQrvrW7PoPQ7KGKyEJQNnlsCXk3cirM1kV3PulrHuZWP-aSkoEZMoDsD45eznTe6yrMOPQSeVYuMgFaduAC-_c7_MxSybcAl92rg0DyoqVmfSS5EknZ1O_rjxZ5bLSx1RD8pwNg_kkdYcb5ezW","created_at":1789777129,"error":null,"incomplete_details":null,"instructions":null,"metadata":{},"model":"us.openai.gpt-6-astra","object":"response","output":[{"type":"message","id":"msg_24f5ecc8-5b5c-4886-b779-64314926822f","status":"completed","role":"assistant","content":[{"type":"output_text","text":"Hi! How can I help you today?","annotations":[]}],"phase":null}],"parallel_tool_calls":false,"temperature":0.0,"tool_choice":"auto","tools":[],"top_p":null,"max_output_tokens":null,"previous_response_id":null,"reasoning":null,"status":"completed","text":{},"truncation":null,"usage":{"input_tokens":7,"input_tokens_details":{"audio_tokens":null,"cached_tokens":0,"cached_tokens_details":null,"text_tokens":7,"cache_write_tokens":0},"output_tokens":13,"output_tokens_details":{"audio_tokens":null,"reasoning_tokens":0,"text_tokens":13},"total_tokens":20,"cost":null},"user":null,"store":null}
HTTP 200

$ curl -s http://localhost:43851/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_output_tokens":1,"input":"hi"}' -w '\nHTTP %{http_code}\n' | jq -c .
{"id":"resp_ueFBx_DfXOtZUYs4k-dz-sz4P9BO7VHp9uz_6rUkf55J_bAE1OLt52fsI0-e12O5d8uVL52ZQR8OPBbFdG6H-_P7PAgG-r-ecl0q0pY43mpuvdexoKfdghR-cvfujklERxVdOy_fNRJy90N6QYiqKJCUgiV-zaAXqUVf7reB_7qndsYY1IrAuEtfRPKEnwpZvMVjMVdaw0Lk-kq2q8sOHAt-5Cwi56cDL9E6Qr54ICWgxWJiZ9ptDPScjMXF5PuRkKJ4DoBGps2Eb-Uo5bePd3bHsgav2TcDpyDdxLqjcoepDC4qMnY7yGgpuivSov0IUykbdJKZhcTk1UUCnBqrejOQU6j9SwDJTc-_jd5WZdL92hjI6tbdmf8Gmv8xGIV1G9pz2vN13w7Vmqs-MKg81KwwrNXVmqfo2GDgqoTdr_oPQJ39zRQ7rVJLKXAZOfDNJL3Cy0Rt","created_at":1789777130,"error":null,"incomplete_details":null,"instructions":null,"metadata":{},"model":"us.xai.grok-4.6","object":"response","output":[{"type":"reasoning","id":"rs_e9b2e1f6-6783-4beb-a5b5-2cf35b384fd2","status":"incomplete","role":"assistant","content":[],"phase":null,"encrypted_content":"[{\"type\":\"redacted_thinking\",\"data\":\"…\"}]"},{"type":"message","id":"msg_46a1556c-cd0a-4384-8878-7679eac86299","status":"incomplete","role":"assistant","content":[{"type":"output_text","text":"","annotations":[]}],"phase":null}],"parallel_tool_calls":false,"temperature":0.0,"tool_choice":"auto","tools":[],"top_p":null,"max_output_tokens":null,"previous_response_id":null,"reasoning":null,"status":"incomplete","text":{},"truncation":null,"usage":{"input_tokens":19,"input_tokens_details":{"audio_tokens":null,"cached_tokens":0,"cached_tokens_details":null,"text_tokens":19,"cache_write_tokens":0},"output_tokens":16,"output_tokens_details":{"audio_tokens":null,"reasoning_tokens":0,"text_tokens":null},"total_tokens":35,"cost":null},"user":null,"store":null}
HTTP 200

Streaming, max_tokens: 1: 200 with a usage chunk of 16 (grok) and 13 (gpt) on chat completions, and a message_delta with output_tokens: 16 on /v1/messages

KEY=<redacted>
$ curl -s -N http://localhost:43851/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' | grep -E '"usage":\{[^}]*"completion_tokens"|^HTTP|"error"' | sed -E 's/^data: //' | jq -c 'if type=="object" then {usage: .usage, error: .error, choices_finish: (.choices // [] | map(.finish_reason))} else . end' 2>/dev/null || true
chunks: 4
{"usage":{"completion_tokens":16,"prompt_tokens":19,"total_tokens":35,"completion_tokens_details":{},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":19,"cache_write_tokens":0,"cache_creation_tokens":0},"cost":0.0001474,"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

$ curl -s -N http://localhost:43851/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' | grep -E '"usage":\{[^}]*"completion_tokens"|^HTTP|"error"' | sed -E 's/^data: //' | jq -c 'if type=="object" then {usage: .usage, error: .error, choices_finish: (.choices // [] | map(.finish_reason))} else . end' 2>/dev/null || true
chunks: 12
{"usage":{"completion_tokens":13,"prompt_tokens":7,"total_tokens":20,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":13},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cost":0.0007920000000000001,"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

$ curl -s -N http://localhost:43851/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"stream":true,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'  (raw saved; usage lines and HTTP shown below)
events: 5
1 event: content_block_start
1 event: content_block_stop
1 event: message_delta
1 event: message_start
1 event: message_stop
{"type":"message_start","usage":null,"stop_reason":null}
{"type":"message_delta","usage":{"input_tokens":19,"output_tokens":16},"stop_reason":"max_tokens"}
HTTP 200

Boundary on us.openai.gpt-6-astra: max_tokens: 15 now returns 200 (clamped to 16, the model stopped at 11), max_tokens: 16 unchanged

KEY=<redacted>
$ curl -s http://localhost:43851/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":15,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' | jq -c .
{"id":"chatcmpl-14f11391-4a8f-46fa-847a-376c82cebf6e","created":1789777350,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help?","role":"assistant"}}],"usage":{"completion_tokens":11,"prompt_tokens":7,"total_tokens":18,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":11},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

$ curl -s http://localhost:43851/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":16,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' | jq -c .
{"id":"chatcmpl-5a08b33f-75a9-41c6-a154-dc721abc7ac8","created":1789777359,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help you today?","role":"assistant"}}],"usage":{"completion_tokens":13,"prompt_tokens":7,"total_tokens":20,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":13},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

Controls: max_tokens: 64 on us.xai.grok-4.6 still returns completion_tokens: 64 (the clamp is a floor, not an override) and max_tokens: 1 on Haiku still returns output_tokens: 1

KEY=<redacted>
$ curl -s http://localhost:43851/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":64,"messages":[{"role":"user","content":"Write one sentence about the sea."}]}' -w '\nHTTP %{http_code}\n' | jq -c .
{"id":"chatcmpl-06a8e05b-ff5c-497b-813b-e3743f3ddda3","created":1789777131,"model":"us.xai.grok-4.6","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"","role":"assistant","reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContentBlocks":[{"redactedContent":"…"}]}}}],"usage":{"completion_tokens":64,"prompt_tokens":25,"total_tokens":89,"completion_tokens_details":{},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":25,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

$ curl -s http://localhost:43851/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-haiku-4-5-20251001-v1:0","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' | jq -c .
{"model":"us.anthropic.claude-haiku-4-5-20251001-v1:0","id":"msg_bdrk_u2wimgtk4lrwckdl2l57t7yd2g6vuuyyssj2gdioaqcsdzij5psa","type":"message","role":"assistant","content":[{"type":"text","text":"#"}],"container":null,"stop_reason":"max_tokens","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":8,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard"}}
HTTP 200

Observations from the run, none caused by this PR:

  • grok-4.6 spends the whole 16-token floor on redacted reasoning
  • grok max_tokens 64 control: 64 tokens, all reasoning, empty text
  • Bedrock names the rejected param max_output_tokens, not maxTokens
  • Streaming requests got the same 400 body before, no SSE
  • Failed /model switch left Claude Code silently on Sonnet
  • Second /model switch shows a Switch model? confirmation dialog

Type

🐛 Bug Fix

Caveats (if any)

Low

  • Only openai.gpt-* and xai.grok-* models on Converse are clamped; inference-profile ARNs carrying the model id are matched, opaque application-inference-profile ARNs are not (they carry no model id, so metadata lookup cannot resolve them either)
  • Other Bedrock model families with their own minimums are untouched
  • The explicit bedrock/openai/... (Mantle) and bedrock/invoke/... routes for the same models are untouched and still forward a sub-16 value as is; the default route for these model ids is Converse, which is what the reports use, and each of those routes would need its own QA leg
  • A max_tokens below 16 on these models now returns up to 16 tokens instead of a 400. No previously working request changes, since Bedrock rejected every value below 16, and the caller's value still applies at 16 and above (the 64-token control above)
  • Four CircleCI jobs are red at the tip, each red on main's own latest pipeline (89818) and none touching Bedrock Converse: test_models_by_provider in litellm_utils_testing (red since feat(proxy): add Amazon Transcribe pass-through with completion-time job pricing #41515), the two retry-count integration tests in integration-management and integration-providers (red since refactor(proxy): make the config file win over the database #41779), and test_bedrock_invoke_messages_with_all_beta_headers in proxy_e2e_anthropic_messages_tests (LIT-8149, red since fix(anthropic): register thinking-binding-controls-2026-08-01 in beta headers config #41203). proxy-infra / Run tests is not a required check and is red on main too

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/0eed92dc9c7f470a82a5781cbf03e622
Open in Devin Desktop: https://app.devin.ai/desktop/session/0eed92dc9c7f470a82a5781cbf03e622?variant=devin


Note

Low Risk
Localized parameter mapping in Bedrock Converse; only raises sub-16 limits for specific model ID patterns, with tests covering edge cases.

Overview
Bedrock Converse now enforces a 16-token floor on maxTokens when mapping OpenAI max_tokens / max_completion_tokens for Bedrock models whose IDs match openai.gpt-* or xai.grok-*, including typical inference-profile ARNs that embed those IDs. Values at or above 16 are unchanged; other Bedrock models still pass through as before.

This fixes 400 errors when clients (e.g. Claude Code’s max_tokens=1 model-switch probe) hit Bedrock’s minimum output-token requirement on those OpenAI-compat and Grok routes. A new _requires_min_max_tokens helper and BEDROCK_OPENAI_COMPAT_MIN_MAX_TOKENS constant drive the clamp; parametrized unit tests cover GPT/Grok vs Claude, both param names, and ARN shapes.

Reviewed by Cursor Bugbot for commit 12120fe. Bugbot is set up for automated code reviews on this repo. Configure here.

…models on Converse

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 18, 2026 20:05
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

CLAassistant commented Sep 18, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ mateo-berri
❌ devin-ai-integration[bot]
You have signed the CLA already but the status is still pending? Let us recheck it.

@codspeed

codspeed Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_bedrock_openai_gpt_min_max_tokens (12120fe) with main (500e880)

Open in CodSpeed

@codecov

codecov Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@greptile-apps

greptile-apps Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

The token clamping behavior appears correct, but the unresolved repository-rule violation must be addressed before merging

Findings

  1. P1 Inference ARNs Bypass Clamp ▶

Summary

This PR clamps Bedrock Converse maxTokens values to 16 for OpenAI GPT and xAI Grok model identifiers while preserving existing behavior for other model families. The latest changes simplify the clamp implementation and add coverage for recognizable and opaque inference-profile ARNs

Reviews (3) · Last reviewed commit: "refactor(bedrock): inline maxTokens clam..."

Comment on lines +1009 to +1011
optional_params["maxTokens"] = (
self._enforce_min_max_tokens(cast(object, value))
if self._is_openai_gpt_reasoning_model(model)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Inference ARNs Bypass Clamp

Opaque inference-profile ARNs bypass this regex, leaving sub-16 limits unchanged and causing Bedrock 400s. Model-specific constraints must come from metadata via get_model_info

Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Substring regex already clamps ARNs carrying the model id (tests added). Opaque application profiles carry no model id, so get_model_info cannot resolve them either

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right. The regex searches for the model identifier anywhere in the input, so inference-profile ARNs that embed openai.gpt-* or xai.grok-* are already covered, as the added tests demonstrate. Opaque application inference-profile ARNs contain no model identifier, so get_model_info cannot reliably resolve their underlying model without an external mapping. This is not an actionable gap; I withdraw the previous finding.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…nverse

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT models on Converse fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse Sep 18, 2026
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 12120fe. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit a6e3a72 into main Sep 19, 2026
136 of 142 checks passed
@mateo-berri
mateo-berri deleted the litellm_bedrock_openai_gpt_min_max_tokens branch September 19, 2026 00:44
mateo-berri added a commit that referenced this pull request Sep 19, 2026
fix(bedrock): backport #41870 and the GPT-6 reasoning gate fix to stable/1.100.x for v1.100.2
mateo-berri added a commit that referenced this pull request Sep 19, 2026
fix(bedrock): backport #41870, the GPT-6 reasoning gate fix, and the Python 3.13 image pin to stable/1.98.x for v1.98.1
yuneng-berri added a commit that referenced this pull request Sep 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants