Skip to content

fix(bedrock): backport #41870 and the GPT-6 reasoning gate fix to stable/1.100.x for v1.100.2 - #42000

Merged
mateo-berri merged 3 commits into
stable/1.100.xfrom
litellm_cherrypick_1_100_x
Sep 19, 2026
Merged

mateo-berri merged 3 commits into
stable/1.100.xfrom
litellm_cherrypick_1_100_x

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

Backport notes

  1. Cherry-picked from main merge commit a6e3a72 (fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870) with -m 1, so the branch carries fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870's source and test changes as one commit

  2. litellm/llms/bedrock/chat/converse_transformation.py conflicted in that cherry-pick because this line never had import re or the _is_openai_gpt_reasoning_model helper that fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870's hunks sit next to on main; the resolution adds import re, the BEDROCK_OPENAI_COMPAT_MIN_MAX_TOKENS constant, the _requires_min_max_tokens helper, and the clamped maxTokens assignment exactly as they merged on main, nothing else

  3. tests/test_litellm/llms/bedrock/chat/test_converse_transformation.py applied clean, so the ten parametrized cases from fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870 run unchanged on this line

  4. The reasoning gate change is a hand port of the Converse part of main's fbc6fb5 (fix(pricing): rolling model registry update: Bedrock gpt-6-astra, gpt-image-2.5, Cohere rerank 4, Vertex Grok 4.3/4.6/4.20, Gemini 3.5 audio, OpenAI web search fee, xAI Imagine video, Lyria 3.5, Voyage, ChatGPT GPT-5.5/5.6, Bedrock Mantle, Scaleway dates #31884): the _is_openai_gpt_reasoning_model helper matching openai\.gpt-\d and its three call sites (the reasoning_effort mapping, get_supported_openai_params, and the thinking skip in map_openai_params). The Nova 2 tool_choice registry keys and the invoke json_mode forwarding in that commit stay on main. The added parametrized test cases come from the same commit and fail on this line without the gate change

  5. Version bump to 1.100.2 is its own commit, one patch above the latest release on this line (1.100.1)

User Flow

Before: a developer using Claude Code through a v1.100.1 proxy cannot use a Bedrock OpenAI GPT or xAI Grok model

  1. They run Claude Code with ANTHROPIC_BASE_URL pointed at the proxy and type /model us.xai.grok-4.6 (or /model us.openai.gpt-6-astra)
  2. Claude Code sends POST https://litellm-domain/v1/messages with "max_tokens": 1 as a warmup probe
  3. They see API error: 400 ... BedrockException ... Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead. and the header keeps the model they had (Sonnet 4.5)
  4. Their next message is answered by that old model, so the Bedrock GPT or Grok model never gets used
  5. With the probe fixed on its own (fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870 alone, this PR's first tip), the switch to us.openai.gpt-6-astra takes, and their first message, which Claude Code sends to POST https://litellm-domain/v1/messages with a thinking block, comes back as API Error: 400 ... Unknown parameter: 'thinking'

After: the same switch and first message succeed once the proxy runs v1.100.2

  1. They run Claude Code with ANTHROPIC_BASE_URL pointed at the proxy and type /model us.xai.grok-4.6 (or /model us.openai.gpt-6-astra)
  2. Claude Code sends POST https://litellm-domain/v1/messages with "max_tokens": 1 as a warmup probe
  3. The probe returns 200 and the header shows the new model
  4. Their first message goes to POST https://litellm-domain/v1/messages with the thinking block and is answered by that Bedrock model

Relevant issues

Backport of #41870 and of the Converse reasoning gate fix from #31884 (main). Reported by customers (Pylon #8817 for OpenAI GPT, Pylon #8821 for xAI Grok); one of them runs v1.98.0 today and is moving to the 1.100.x line

Affected release

Linear ticket

Resolves LIT-8154

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Both legs ran the same way and differ only in the commit the proxy was booted from: a no-DB proxy with 2 uvicorn workers on a random port, real Bedrock in us-east-1 (real spend, no mocks), and Claude Code 2.1.278 driven interactively under tmux. $KEY is the proxy master key and $PORT is the leg's port

config.yaml

model_list:
  - model_name: us.openai.gpt-6-astra
    litellm_params:
      model: bedrock/us.openai.gpt-6-astra
      aws_region_name: us-east-1
  - model_name: us.xai.grok-4.6
    litellm_params:
      model: bedrock/us.xai.grok-4.6
      aws_region_name: us-east-1
  - model_name: us.anthropic.claude-sonnet-4-5-20250929-v1:0
    litellm_params:
      model: bedrock/us.anthropic.claude-sonnet-4-5-20250929-v1:0
      aws_region_name: us-east-1
  - model_name: us.anthropic.claude-haiku-4-5-20251001-v1:0
    litellm_params:
      model: bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0
      aws_region_name: us-east-1

general_settings:
  master_key: sk-REDACTED

Proxy, booted from the leg's commit with a Python 3.12 venv

AWS_PROFILE=$BEDROCK_PROFILE AWS_DEFAULT_REGION=us-east-1 .venv/bin/python litellm/proxy/proxy_cli.py --config config.yaml --port $PORT --num_workers 2

Claude Code, started in an empty folder with a fresh config directory (theme picker and security notes accepted with Enter, folder trust accepted with Down then Enter)

tmux new-session -d -s claude -x 160 -y 45 -c work "env -u ANTHROPIC_API_KEY ANTHROPIC_BASE_URL=http://localhost:$PORT ANTHROPIC_AUTH_TOKEN=$KEY ANTHROPIC_MODEL=us.anthropic.claude-sonnet-4-5-20250929-v1:0 ANTHROPIC_DEFAULT_HAIKU_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0 ANTHROPIC_SMALL_FAST_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0 CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 CLAUDE_CONFIG_DIR=claude-config claude"

Before (b7d81e9)

The last curl group (the four reasoning cases) ran in a second proxy boot at this same commit after the floor-only after leg surfaced the thinking 400; every other case below is from the original before leg

Claude Code, /model us.xai.grok-4.6

  1. In the running Claude Code session, type /model us.xai.grok-4.6 and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  Sonnet 4.5 · API Usage Billing
    ❯ /model us.xai.grok-4.6
      ⎿  API error: 400 {"error":{"message":"litellm.BadRequestError: BedrockException - {"message":"The model returned the following errors:
         {\"error\":{\"code\":\"integer_below_min_value\",\"message\":\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1
         instead.\",\"param\":\"max_output_tokens\",\"type\":\"invalid_request_error\"}}"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group
         Fallbacks=None","type":null,"param":null,"code":"400"}}
    

    pr42000-1a14aadd03-claude-1_100_x-before-b7d81e98b6-grok.png

  2. Type Reply with the single word ok and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  Sonnet 4.5 · API Usage Billing
    ❯ Reply with the single word ok
      Thought for 7s
    ⏺ ok
    ✻ Cooked for 7s · done 11:50 AM
    

Claude Code, /model us.openai.gpt-6-astra

  1. In the running Claude Code session, type /model us.openai.gpt-6-astra and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  Sonnet 4.5 · API Usage Billing
    ❯ /model us.openai.gpt-6-astra
      ⎿  API error: 400 {"error":{"message":"litellm.BadRequestError: BedrockException - {"message":"The model returned the following errors:
         {\"error\":{\"code\":\"integer_below_min_value\",\"message\":\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1
         instead.\",\"param\":\"max_output_tokens\",\"type\":\"invalid_request_error\"}}"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group
         Fallbacks=None","type":null,"param":null,"code":"400"}}
    

    pr42000-1a14aadd03-claude-1_100_x-before-b7d81e98b6-gpt.png

  2. Type Reply with the single word ok and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  Sonnet 4.5 · API Usage Billing
    ❯ Reply with the single word ok
      Thought for 4s
    ⏺ ok
    ✻ Worked for 4s · done 11:51 AM
    

POST /v1/messages, max_tokens 1

  1. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  2. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    

POST /v1/chat/completions, max_tokens 1 and max_completion_tokens 1

  1. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  2. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  3. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_completion_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  4. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_completion_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    

POST /v1/responses, max_output_tokens 1

  1. Run

    curl -s http://localhost:$PORT/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_output_tokens":1,"input":"hi"}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  2. Run

    curl -s http://localhost:$PORT/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_output_tokens":1,"input":"hi"}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    

Streaming, max_tokens 1

  1. Run

    curl -s -N http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' > stream-chat-us.openai.gpt-6-astra.raw

    Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body

    HTTP: HTTP 400
    data lines: 0
    event lines: 0
    event types: 
    usage-bearing chunks:
    non-SSE body (if any):
    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    
  2. Run

    curl -s -N http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' > stream-chat-us.xai.grok-4.6.raw

    Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body

    HTTP: HTTP 400
    data lines: 0
    event lines: 0
    event types: 
    usage-bearing chunks:
    non-SSE body (if any):
    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    
  3. Run

    curl -s -N http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"stream":true,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' > stream-messages-grok.raw

    Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body

    HTTP: HTTP 400
    data lines: 0
    event lines: 0
    event types: 
    usage-bearing chunks:
    non-SSE body (if any):
    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    

Boundary on us.openai.gpt-6-astra, max_tokens 15 then 16

  1. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":15,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 15 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  2. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":16,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-f2a4d154-4a14-493b-8884-b3ccb03bb75d","created":1789843671,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help?","role":"assistant"}}],"usage":{"completion_tokens":11,"prompt_tokens":7,"total_tokens":18,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":11},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    

Controls, a larger max_tokens and the Anthropic models

  1. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":64,"messages":[{"role":"user","content":"Write one sentence about the sea."}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-49946c17-50a1-47d6-baa9-3c7a0fb27e66","created":1789843672,"model":"us.xai.grok-4.6","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"","role":"assistant","reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContentBlocks":[{"redactedContent":"…"}]}}}],"usage":{"completion_tokens":64,"prompt_tokens":25,"total_tokens":89,"completion_tokens_details":{},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":25,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","id":"msg_bdrk_011CfDJAs8eoPKva87vycMxj","type":"message","role":"assistant","content":[{"type":"text","text":"Hello"}],"stop_reason":"max_tokens","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":8,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1}}
    HTTP 200
    
  3. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-haiku-4-5-20251001-v1:0","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"model":"us.anthropic.claude-haiku-4-5-20251001-v1:0","id":"msg_bdrk_iujj2rfn5r3gycqbnscz3s6e2iqcakumzavhs5ozn2f6c746e5yq","type":"message","role":"assistant","content":[{"type":"text","text":"Hello"}],"container":null,"stop_reason":"max_tokens","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":8,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard"}}
    HTTP 200
    

Reasoning params on us.openai.gpt-6-astra (Claude Code sends thinking on its real turns), plus the Sonnet control

  1. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":2048,"thinking":{"type":"enabled","budget_tokens":1024},"messages":[{"role":"user","content":"Reply with the single word ok"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"unknown_parameter\\\",\\\"message\\\":\\\"Unknown parameter: 'thinking'.\\\",\\\"param\\\":\\\"thinking\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  2. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":512,"reasoning_effort":"low","messages":[{"role":"user","content":"Reply with the single word ok"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"unknown_parameter\\\",\\\"message\\\":\\\"Unknown parameter: 'thinking'.\\\",\\\"param\\\":\\\"thinking\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  3. Run

    curl -s http://localhost:$PORT/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_output_tokens":512,"reasoning":{"effort":"low"},"input":"Reply with the single word ok"}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"unknown_parameter\\\",\\\"message\\\":\\\"Unknown parameter: 'thinking'.\\\",\\\"param\\\":\\\"thinking\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  4. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","max_tokens":2048,"thinking":{"type":"enabled","budget_tokens":1024},"messages":[{"role":"user","content":"Reply with the single word ok"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","id":"msg_bdrk_011CfDNaKgBbGhvX17m9vxPy","type":"message","role":"assistant","content":[{"type":"thinking","thinking":"The user is asking me to reply with the single word \"ok\". This is a straightforward request that doesn't involve anything harmful or problematic. I should simply respond with \"ok\" as requested.","signature":"…"},{"type":"text","text":"ok"}],"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":42,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":53,"output_tokens_details":{"thinking_tokens":46}}}
    HTTP 200
    

After (5545ca9)

Claude Code, /model us.xai.grok-4.6

  1. In the running Claude Code session, type /model us.xai.grok-4.6 and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  us.xai.grok-4.6 · API Usage Billing
    ❯ /model us.xai.grok-4.6
      ⎿  Set model to us.xai.grok-4.6 and saved as your default for new sessions
              ANTHROPIC_MODEL is set to Sonnet 4.5 — new sessions use that while it is set
                                                                                                                                                  ● high · /effort
    

    pr42000-5545ca9e86-claude-1_100_x-after-5545ca9e86-grok.png

  2. Type Reply with the single word ok and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  us.xai.grok-4.6 · API Usage Billing
    ❯ /model us.xai.grok-4.6
      ⎿  Set model to us.xai.grok-4.6 and saved as your default for new sessions
              ANTHROPIC_MODEL is set to Sonnet 4.5 — new sessions use that while it is set
    ❯ Reply with the single word ok
    ⏺ Skill(claude-api)
      ⎿  Initializing…
      ⎿  Interrupted · What should Claude do instead?
    ✻ Baked for 2s · done 12:50 PM
    ❯ Reply with the single word ok
    ⏺ ok
    ✻ Baked for 1s · done 12:50 PM
    

Claude Code, /model us.openai.gpt-6-astra

  1. In the running Claude Code session, type /model us.openai.gpt-6-astra and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  us.openai.gpt-6-astra · API Usage Billing
    ❯ /model us.openai.gpt-6-astra
      ⎿  Set model to us.openai.gpt-6-astra and saved as your default for new sessions
              ANTHROPIC_MODEL is set to Sonnet 4.5 — new sessions use that while it is set
                                                                                                                                                  ● high · /effort
    

    pr42000-5545ca9e86-claude-1_100_x-after-5545ca9e86-gpt.png

  2. Type Reply with the single word ok and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  us.openai.gpt-6-astra · API Usage Billing
    ❯ /model us.xai.grok-4.6
      ⎿  Set model to us.xai.grok-4.6 and saved as your default for new sessions
              ANTHROPIC_MODEL is set to Sonnet 4.5 — new sessions use that while it is set
    ❯ Reply with the single word ok
    ⏺ Skill(claude-api)
      ⎿  Initializing…
      ⎿  Interrupted · What should Claude do instead?
    ✻ Baked for 2s · done 12:50 PM
    ❯ Reply with the single word ok
    ⏺ ok
    ✻ Baked for 1s · done 12:50 PM
    ❯ /model us.openai.gpt-6-astra
      ⎿  Set model to us.openai.gpt-6-astra and saved as your default for new sessions
              ANTHROPIC_MODEL is set to Sonnet 4.5 — new sessions use that while it is set
    ❯ Reply with the single word ok
    ⏺ ok
    ✻ Sautéed for 1s · done 12:51 PM
    

POST /v1/messages, max_tokens 1

  1. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-9633c4ea-23cc-47a7-a0c6-6af28a6a7575","type":"message","role":"assistant","model":"us.openai.gpt-6-astra","stop_sequence":null,"usage":{"input_tokens":7,"output_tokens":13},"content":[{"type":"text","text":"Hi! How can I help you today?"}],"stop_reason":"end_turn"}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-4f1b6c2b-09b3-4d24-b5b6-884509605dc9","type":"message","role":"assistant","model":"us.xai.grok-4.6","stop_sequence":null,"usage":{"input_tokens":19,"output_tokens":16},"content":[{"type":"redacted_thinking","data":"…"},{"type":"text","text":""}],"stop_reason":"max_tokens"}
    HTTP 200
    

POST /v1/chat/completions, max_tokens 1 and max_completion_tokens 1

  1. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-69c9e016-ecc8-4353-9f31-da329d029094","created":1789847280,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help?","role":"assistant"}}],"usage":{"completion_tokens":11,"prompt_tokens":7,"total_tokens":18,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":11},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-a3aa69ba-5d78-4175-9272-6c62ab063617","created":1789847281,"model":"us.xai.grok-4.6","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"","role":"assistant","reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContentBlocks":[{"redactedContent":"…"}]}}}],"usage":{"completion_tokens":16,"prompt_tokens":19,"total_tokens":35,"completion_tokens_details":{},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":19,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  3. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_completion_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-a8109599-6585-4225-9311-92d228ddc15d","created":1789847282,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help?","role":"assistant"}}],"usage":{"completion_tokens":11,"prompt_tokens":7,"total_tokens":18,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":11},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  4. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_completion_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-9fbee3ae-cfa5-4d0d-bea8-6d3f7ff52f46","created":1789847283,"model":"us.xai.grok-4.6","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"","role":"assistant","reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContentBlocks":[{"redactedContent":"…"}]}}}],"usage":{"completion_tokens":16,"prompt_tokens":19,"total_tokens":35,"completion_tokens_details":{},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":19,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    

POST /v1/responses, max_output_tokens 1

  1. Run

    curl -s http://localhost:$PORT/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_output_tokens":1,"input":"hi"}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"resp_Pd_EBy-jeHKnmWePemdTpHRN3S77qJQ-DgLBBSGaP9WsEm1L4Mep6KYFdz1C8sxhZ18P38mcOLMD9HVaE9zdVcrXn3KqgEwiwleVm0-zicgO8HYCel3zQSs5anKG2IkdvoeOnVO2o7k14sBkhL0nkvO7Uamy9L3YSQKsTAxe6rROqEwpisPTk7l3GlBTL42hwzQspYWDkUp6ZH6sXk5RqyD17X-mScUqVx8I74iWLbKye4bbbEpuDHLUJN9-I6H9zYHy-aIYfD3ZsFi0SuOlTRo65QDzsvL6uCcXLtdAJMgwDpuvTUY1OovWsvA3Tfk6o_3CAPJ7a7Ec-CSV79G1bUPdb3ZtR56LBqDo0hWVRk6tsH5jsopEyx3kvvfZoIB5Ihd1pZ8JCSUZtyn23RvN21yq-IvZeDbr61PrilJHsCUBZXMaL4-_LqAVt1lLPdSSgrjUyi1C","created_at":1789847284,"error":null,"incomplete_details":null,"instructions":null,"metadata":{},"model":"us.openai.gpt-6-astra","object":"response","output":[{"type":"message","id":"msg_2774115b-7c85-4b6c-940e-a6e41547d2fb","stat
    [... 477 characters cut for length ...]
    tokens":11,"output_tokens_details":{"audio_tokens":null,"reasoning_tokens":0,"text_tokens":11},"total_tokens":18,"cost":null},"user":null,"store":null}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_output_tokens":1,"input":"hi"}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"resp_wDGx8HFWBJIN34X6s_UyHKm_AJjYfGA0TTUrvu8yq8bntuQ3Hddtvw5Y6Lh-By-99AeBClmda4YYIUs22RgJ_cabXMMtXaT4wa_rTtpyHLEQzBGk8KLgddfDguTCHzQc4ZfKwerezVWYLjoxad33RMlBxANJ1yGb6Rx49_yMegPk_IT-17NKCIgFYZ7gFpL6DHud7vLulGO-yUiXNqdecSGBEr3GJiUIQS7MdzuPwQyl8aK0_S220eRU8CLFNDlx6JMtR02LilGRRkEfpOrcRbCEbb6TofulvQjthTkED5FrL810A50JbrOmblT4VV_hW_F1ImA0n44zoR4rTNcJRvynRC7fKmauoEY-n1pipTLvkdRZDQalmVkcXFUuKNHHUfIvIZGKSpyNb_FOOU_ROpezUeJ_RsuksTouA226CnxYv6CRiMl51QDw-HaBbQ3gxN_TeviU","created_at":1789847285,"error":null,"incomplete_details":null,"instructions":null,"metadata":{},"model":"us.xai.grok-4.6","object":"response","output":[{"type":"reasoning","id":"rs_1b0db188-0cf6-4f20-90d8-484d95e50db8","status":"
    [... 665 characters cut for length ...]
    kens":16,"output_tokens_details":{"audio_tokens":null,"reasoning_tokens":0,"text_tokens":null},"total_tokens":35,"cost":null},"user":null,"store":null}
    HTTP 200
    

Streaming, max_tokens 1

  1. Run

    curl -s -N http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' > stream-chat-us.openai.gpt-6-astra.raw

    Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body

    HTTP: HTTP 200
    data lines: 12
    event lines: 0
    event types: 
    usage-bearing chunks:
    {"type":null,"usage":{"completion_tokens":13,"prompt_tokens":7,"total_tokens":20,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":13},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0},"delta":null,"message_usage":null,"choices":[{"index":0,"delta":{}}]}
    non-SSE body (if any):
    
  2. Run

    curl -s -N http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' > stream-chat-us.xai.grok-4.6.raw

    Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body

    HTTP: HTTP 200
    data lines: 4
    event lines: 0
    event types: 
    usage-bearing chunks:
    {"type":null,"usage":{"completion_tokens":16,"prompt_tokens":19,"total_tokens":35,"completion_tokens_details":{},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":19,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0},"delta":null,"message_usage":null,"choices":[{"index":0,"delta":{}}]}
    non-SSE body (if any):
    
  3. Run

    curl -s -N http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"stream":true,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' > stream-messages-grok.raw

    Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body

    HTTP: HTTP 200
    data lines: 5
    event lines: 5
    event types:    1 content_block_start;   1 content_block_stop;   1 message_delta;   1 message_start;   1 message_stop;
    usage-bearing chunks:
    {"type":"message_start","usage":null,"delta":null,"message_usage":{"input_tokens":0,"output_tokens":0,"cache_creation_input_tokens":0,"cache_read_input_tokens":0},"choices":null}
    {"type":"message_delta","usage":{"input_tokens":19,"output_tokens":16},"delta":{"stop_reason":"max_tokens"},"message_usage":null,"choices":null}
    non-SSE body (if any):
    

Boundary on us.openai.gpt-6-astra, max_tokens 15 then 16

  1. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":15,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-b2a6e1d9-cd99-4515-b36c-b9f03fdf77b7","created":1789847288,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help?","role":"assistant"}}],"usage":{"completion_tokens":11,"prompt_tokens":7,"total_tokens":18,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":11},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":16,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-a87c7cc4-3a77-4b7b-9314-2f89d3785e61","created":1789847289,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help?","role":"assistant"}}],"usage":{"completion_tokens":11,"prompt_tokens":7,"total_tokens":18,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":11},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    

Controls, a larger max_tokens and the Anthropic models

  1. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":64,"messages":[{"role":"user","content":"Write one sentence about the sea."}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-426fe635-9a50-4239-8f86-761f2a8c0d1a","created":1789847290,"model":"us.xai.grok-4.6","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"","role":"assistant","reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContentBlocks":[{"redactedContent":"…"}]}}}],"usage":{"completion_tokens":64,"prompt_tokens":25,"total_tokens":89,"completion_tokens_details":{},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":25,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","id":"msg_bdrk_011CfDNmQQkfNewQyoTSEjfB","type":"message","role":"assistant","content":[{"type":"text","text":"Hello"}],"stop_reason":"max_tokens","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":8,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1}}
    HTTP 200
    
  3. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-haiku-4-5-20251001-v1:0","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"model":"us.anthropic.claude-haiku-4-5-20251001-v1:0","id":"msg_bdrk_fdlvu3cyzlie2jso3pzaaowegyehb3cfowaw73cdbxe5ht6hjoyq","type":"message","role":"assistant","content":[{"type":"text","text":"#"}],"container":null,"stop_reason":"max_tokens","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":8,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard"}}
    HTTP 200
    

Reasoning params on us.openai.gpt-6-astra (Claude Code sends thinking on its real turns), plus the Sonnet control

  1. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":2048,"thinking":{"type":"enabled","budget_tokens":1024},"messages":[{"role":"user","content":"Reply with the single word ok"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-a0b41843-938d-4552-9609-62bb94940d91","type":"message","role":"assistant","model":"us.openai.gpt-6-astra","stop_sequence":null,"usage":{"input_tokens":12,"output_tokens":5},"content":[{"type":"text","text":"ok"}],"stop_reason":"end_turn"}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":512,"reasoning_effort":"low","messages":[{"role":"user","content":"Reply with the single word ok"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-b76b2829-ca24-49e9-a3c7-168dabdb729e","created":1789847306,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"ok","role":"assistant"}}],"usage":{"completion_tokens":5,"prompt_tokens":12,"total_tokens":17,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":5},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":12,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  3. Run

    curl -s http://localhost:$PORT/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_output_tokens":512,"reasoning":{"effort":"low"},"input":"Reply with the single word ok"}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"resp_zHjOiYx400KDAHy1Lptrm60oQwoGjjD66zFSwP9_M3NU3pJryyoz19v8okwR76U0oRZxVo0BPmUyqBM7pavyYPE87JAz6wniDmcx4klKYl2xzsGO0HoCuj0QdEqU4N5j853JU-bGaEZqN6tzLc6-5iQe6juRMSk64cyqrGl0ZdQyHMKGo6BrCVV9xV_pVGyssCHVBDW-ZZP993fOrPPFEiHlxkKkwYDz5nlyCHMV_u-0nR2-0fDGm4o2sGyK5Db9ZlaVi8_1UHDcduSa7r_C0wxgVTQvCzQpuQzHE5046LUbQekT2uAZAWrbIfyXslPBdjbWxUAQv0JYW6sBcU9pW8XSRb8w6AdBaRhTZNdLaUNt2hm8bd7pdrNId-Dsxm8asVT4WwhBS9dCk2lWwncBwtQjMoR1uS-ExDhmpILiOJfs_BYVz1GdDFWIzEmBMFltdWMfxiLl","created_at":1789847307,"error":null,"incomplete_details":null,"instructions":null,"metadata":{},"model":"us.openai.gpt-6-astra","object":"response","output":[{"type":"message","id":"msg_1af77e0c-bdc8-4484-b76f-a0ed629c8b75","stat
    [... 460 characters cut for length ...]
    t_tokens":5,"output_tokens_details":{"audio_tokens":null,"reasoning_tokens":0,"text_tokens":5},"total_tokens":17,"cost":null},"user":null,"store":null}
    HTTP 200
    
  4. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","max_tokens":2048,"thinking":{"type":"enabled","budget_tokens":1024},"messages":[{"role":"user","content":"Reply with the single word ok"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","id":"msg_bdrk_011CfDNoYTVcK7ago1KsLHPw","type":"message","role":"assistant","content":[{"type":"thinking","thinking":"The user is asking me to reply with the single word \"ok\". This is a straightforward request that doesn't involve anything harmful or problematic. I'll simply respond with \"ok\" as requested.","signature":"…"},{"type":"text","text":"ok"}],"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":42,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":53,"output_tokens_details":{"thinking_tokens":46}}}
    HTTP 200
    

Observations from the run

Seen on the after leg. The first two follow from the floor this PR adds and do not affect the probe, which only needs the 200; the Haiku line shows Anthropic models are left alone; the rest are Claude Code behavior this PR neither causes nor changes

  • Grok spends all 16 floor tokens on redacted reasoning
  • gpt-6-astra answers in 11 to 13 tokens under the floor
  • Haiku control still returns exactly 1 output token
  • Grok's first ok turn became a Skill call; declined, re-sent
  • Second /model opens a Switch model? dialog
  • Zero 400 lines in the proxy log during the run

Type

🐛 Bug Fix

Caveats (if any)

Low

  • Only openai.gpt-* and xai.grok-* models on Converse are clamped; inference-profile ARNs carrying the model id are matched, opaque application-inference-profile ARNs are not
  • The explicit bedrock/openai/... (Mantle) and bedrock/invoke/... routes for the same models are untouched and still forward a sub-16 value as is, same as fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870
  • A max_tokens below 16 on these models now returns up to 16 tokens instead of a 400; no previously working request changes, since Bedrock rejected every value below 16
  • The reasoning gate now matches any openai.gpt-<digit> model id, so a future GPT-7 model on Converse gets the same reasoning.effort mapping without a further change, same as main
  • CircleCI is red on five jobs whose failures are all provider-side and untouched by this diff: local_testing_part1, local_testing_part2, litellm_router_testing, and e2e_openai_endpoints were red the same way on fix(responses): backport mid-stream content_policy_violation fallback routing to stable/1.100.x #41208, the last PR merged into stable/1.100.x (Together AI 400s, cohere.command-r-plus-v1:0 no longer served by Bedrock, OpenAI timeout tests, and test_bad_request_error, where OpenAI now answers an unknown model on the Responses API with a 404 and the test expects a 400), and llm_translation_testing fails only in tests/llm_translation/test_together_ai.py (18 Together AI 400s). None of them exercise the Bedrock Converse path; build_docker_database_image and every other job are green

Blast radius, /live-pr-risk at 5545ca9 against base b7d81e9

Verdict: no dependent path that worked on the base answers worse on the head. Two no-DB proxies (2 workers each) ran 33 scenarios side by side against real Bedrock, each behind a forwarding recorder that logged what LiteLLM sent before re-signing it for AWS. Both sides sent the same header set, and the same number of upstream requests in every scenario but the raw thinking case below, and the only body differences are the two this PR names: inferenceConfig.maxTokens raised to 16, and the GPT-6 thinking field replaced by reasoning.effort

Breaking: none observed

Backward incompatible, all benign, and all the behavior main already ships from #41870 and #31884 (both approved by mateo-berri on main):

  • max_tokens of 0, -1, and JSON true on us.openai.gpt-6-astra were Bedrock 400s on the base and are 200s on the head, sent as maxTokens: 16. true passes the isinstance(value, int) guard because a Python bool is an int. "1", 1.5, and [1] are forwarded unchanged and 400 identically on both sides
  • GPT-6 requests carrying reasoning_effort on POST /v1/chat/completions, or a thinking block on POST /v1/messages, were 400 Unknown parameter: 'thinking' on the base and are 200 on the head; the upstream body now carries additionalModelRequestFields: {"reasoning": {"effort": "low"}} where the base sent {"thinking": {"type": "enabled", "budget_tokens": 1024}}, and additionalModelResponseFieldPaths is no longer sent for these models
  • A raw thinking param on POST /v1/chat/completions for GPT-6 stays a 400, but it is now LiteLLM's UnsupportedParamsError with no upstream call where the base forwarded it and relayed Bedrock's 400
  • GET /utils/supported_openai_params and GET /model_group/info no longer list thinking and output_config in supported_openai_params for the GPT-6 model groups (the plain id, the bedrock/converse/ prefix, and an ARN deployment with model_info.base_model); reasoning_effort stays. Grok, Haiku, Nova, and the ARN deployment without base_model are identical on both sides. The dashboard model hub only displays this list, so it shows two fewer params for these models

Regression risk: spend rows for the floored requests were not read back, because the rig ran without a database; the usage blocks in the head's 200 responses reported 5 output tokens for GPT-6 and 16 for Grok

Dependency graph:

  • map_openai_params maxTokens branch via /v1/chat/completions (max_tokens, max_completion_tokens), /v1/messages, /v1/responses, converse-stream, the bedrock/converse/ prefix, and a router fallback replay (2 upstream calls on both sides): unit tested and verified live, base 400 and head 200
  • Boundary and controls (max_tokens 16 and 17, Haiku at 1, openai.gpt-oss-20b-1:0 at 1, unauthenticated 401 with 0 upstream calls): verified live, identical on both sides
  • AmazonInvokeNovaConfig, the only subclass of AmazonConverseConfig: verified live on bedrock/invoke/us.amazon.nova-lite-v1:0, maxTokens: 1 on both sides
  • Opaque application inference profile ARN for GPT-6, with and without model_info.base_model: unit tested as not floored and verified live, 400 on both sides, so this PR leaves that deployment shape broken the way it already was
  • Grok with reasoning_effort: verified live, still sent as Anthropic-style thinking and 200 on both sides
  • GET /model_group/info and GET /utils/supported_openai_params: verified live through the API, difference listed above; the dashboard model hub page that renders the list was not opened
  • pyproject.toml and uv.lock: version bump only, no resolved package changes

Not verified: spend and usage rows in the database, the Admin UI pages that read them, and the user docs in litellm-docs. The legacy tests/llm_translation and tests/local_testing suites ran on CircleCI and are covered in the bullet above

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • 5545ca9 passes /live-pr-risk

Note

Medium Risk
Changes request shaping on a hot Bedrock Converse path (token limits and reasoning/thinking forwarding); scope is narrow to specific model id patterns with tests, but mis-matched model strings could still get wrong clamping or param drops.

Overview
Backports Bedrock Converse fixes for OpenAI-compat and GPT-6 models on stable/1.100.x (release 1.100.2).

On the Converse route, max_tokens / max_completion_tokens are now raised to at least 16 when the model id matches openai.gpt-* or xai.grok-*, so low-token probes (e.g. Claude Code /model switches) no longer fail Bedrock’s minimum output-token validation. Other Bedrock models are unchanged.

Reasoning handling no longer keys only on openai.gpt-5: any openai.gpt-<digit> model is treated as an OpenAI reasoning model—reasoning_effort maps to reasoning.effort, Anthropic thinking is not forwarded, and supported-params reflect that (fixes GPT-6 Unknown parameter: 'thinking'). Parametrized unit tests cover the token floor and GPT-6 reasoning paths.

Reviewed by Cursor Bugbot for commit 5545ca9. Bugbot is set up for automated code reviews on this repo. Configure here.

…and xAI Grok models on Converse (#41870)

Backport of #41870 to stable/1.100.x. Cherry-picked from a6e3a72 (main) with -m 1.

converse_transformation.py conflicted because this line has no `import re` and no
_is_openai_gpt_reasoning_model helper next to the insertion point. The resolution adds
exactly the four hunks #41870 merged: the import, the 16-token constant,
_requires_min_max_tokens, and the clamped maxTokens assignment. The test file applied clean.
@devin-ai-integration

devin-ai-integration Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

The behavioral backport appears correct, but the outstanding repository requirement against hardcoded model capabilities must be addressed before merging.

Findings

  1. P1 Opaque profiles bypass clamp ▶
  2. P2 Hardcoded reasoning capability ▶

Summary

Backports Bedrock Converse compatibility fixes and bumps the stable release to 1.100.2.

  • Clamps sub-16 token limits for Bedrock OpenAI GPT and xAI Grok models.
  • Extends GPT reasoning handling to GPT-6 and future digit-prefixed GPT generations.
  • Adds regression coverage for token clamping, reasoning-effort mapping, and filtering incompatible thinking parameters.
  • Updates project and lockfile version metadata consistently.

Reviews (4) · Last reviewed commit: "fix(bedrock): match any openai.gpt-<digi..."


@staticmethod
def _requires_min_max_tokens(model: str) -> bool:
return re.search(r"openai\.gpt-\d|xai\.grok-", model) is not None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Opaque profiles bypass clamp

Hardcoded matching violates the directive to store model capabilities in metadata and misses opaque profiles, so GPT/Grok requests below 16 are rejected upstream

Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)

Knowledge Base Used: Provider adapters and capabilities

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Backport of #41870, where this was withdrawn. The regex clamps ARNs embedding the model id; opaque application profiles carry none, so get_model_info cannot resolve them

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right. In this backport, matching the model string is intentional: it covers inference-profile ARNs that embed openai.gpt-* or xai.grok-*, while opaque application-inference-profile ARNs contain no resolvable model identifier for get_model_info to use. Given that limitation and the behavior established by #41870, this comment does not apply. No change needed.

…oning gate

Backports the Converse part of fbc6fb5 from main (PR #31884). The gate only matched
openai.gpt-5, so a GPT-6 model fell through to Anthropic's thinking block and Bedrock
rejected the first real turn after a Claude Code /model switch with 400 Unknown
parameter: 'thinking'. The Nova 2 tool_choice registry keys and the invoke json_mode
forwarding in that commit stay on main
@mateo-berri mateo-berri changed the title fix(bedrock): backport #41870 to stable/1.100.x for v1.100.2 fix(bedrock): backport #41870 and the GPT-6 reasoning gate fix to stable/1.100.x for v1.100.2 Sep 19, 2026
@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 19, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

1 similar comment
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

Comment on lines +303 to +304
return re.search(r"openai\.gpt-\d", model) is not None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Hardcoded reasoning capability

_is_openai_gpt_reasoning_model decides reasoning support from the openai.gpt-<digit> model name. That result controls both the reasoning.effort mapping and whether thinking is discarded. This violates the repository directive that model-specific capabilities must be stored in model_prices_and_context_window.json and queried through model metadata. The repository requirement must be satisfied before merging.

Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This mirrors main's merged #31884. A stable backport should not diverge from main, and the family pattern covers new GPT releases without a cost-map entry

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That’s fair. For this PR, matching main’s merged #31884 behavior is the correct backport constraint, and the family-level match is intentionally forward-compatible with future openai.gpt-<digit> Bedrock models without requiring a cost-map update for each release. I’m withdrawing this as a merge blocker; the metadata-vs-family-gate convention can be revisited separately if main changes.

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 5545ca9. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 44d3c4f into stable/1.100.x Sep 19, 2026
48 of 55 checks passed
@mateo-berri
mateo-berri deleted the litellm_cherrypick_1_100_x branch September 19, 2026 22:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant