Skip to content

fix(bedrock): backport #41870, the GPT-6 reasoning gate fix, and the Python 3.13 image pin to stable/1.98.x for v1.98.1 - #42002

Merged
mateo-berri merged 5 commits into
stable/1.98.xfrom
litellm_cherrypick_1_98_x
Sep 19, 2026
Merged

mateo-berri merged 5 commits into
stable/1.98.xfrom
litellm_cherrypick_1_98_x

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Customers on the stable/1.98.x line hit the bug fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870 fixed on main: Claude Code's /model switch to a Bedrock OpenAI GPT or xAI Grok model fails with a 400, because Claude Code sends a max_tokens=1 probe on switch, those Bedrock models require at least 16, and the Bedrock Converse path forwarded maxTokens verbatim
  • Once the probe passes, the first message on a GPT-6 model (us.openai.gpt-6-astra) still fails with 400 Unknown parameter: 'thinking', because the Converse reasoning gate on this line did not recognize GPT-6 models and Claude Code's thinking block was mapped as Anthropic thinking on a model that rejects it
  • On stable/1.98.x the OpenAI GPT reasoning gate did not exist at all (main got it for GPT-5 in fix(bedrock): map reasoning_effort to reasoning.effort for OpenAI GPT-5.x on Converse #38279), so the backport also adds the reasoning.effort mapping and its BedrockConverseGptReasoningEffortBlock type
  • The line's Docker images install an unpinned python3, which now resolves to CPython 3.14, where uvloop 0.21.0 no longer builds; build_docker_database_image was red on this PR's first tip, so the images move to a pinned python-3.13 the way main did in 728dec2 and, for the migrations image, 3947374

How it solves it:

Backport notes

  1. Cherry-picked from main merge commit a6e3a72 (fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870) with -m 1, so the branch carries fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870's source and test changes as one commit
  2. litellm/llms/bedrock/chat/converse_transformation.py conflicted in that cherry-pick because this line never had import re or the _is_openai_gpt_reasoning_model helper that fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870's hunks sit next to on main; the resolution adds import re, the BEDROCK_OPENAI_COMPAT_MIN_MAX_TOKENS constant, the _requires_min_max_tokens helper, and the clamped maxTokens assignment exactly as they merged on main, nothing else
  3. tests/test_litellm/llms/bedrock/chat/test_converse_transformation.py applied clean, so the ten parametrized cases from fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870 run unchanged on this line
  4. The reasoning gate change is a hand port of the Converse part of main's fbc6fb5 (fix(pricing): rolling model registry update: Bedrock gpt-6-astra, gpt-image-2.5, Cohere rerank 4, Vertex Grok 4.3/4.6/4.20, Gemini 3.5 audio, OpenAI web search fee, xAI Imagine video, Lyria 3.5, Voyage, ChatGPT GPT-5.5/5.6, Bedrock Mantle, Scaleway dates #31884): the _is_openai_gpt_reasoning_model helper matching openai\.gpt-\d and its three call sites (the reasoning_effort mapping, get_supported_openai_params, and the thinking skip in map_openai_params). The Nova 2 tool_choice registry keys and the invoke json_mode forwarding in that commit stay on main. The added parametrized test cases come from the same commit and fail on this line without the gate change
  5. On this line the port also carries the GPT-5 gate from fix(bedrock): map reasoning_effort to reasoning.effort for OpenAI GPT-5.x on Converse #38279 (74e86d3, 9cc276a, 418012a): the reasoning_effort to reasoning.effort branch and the BedrockConverseGptReasoningEffortBlock TypedDict in litellm/types/llms/bedrock.py. fix(bedrock): map reasoning_effort to reasoning.effort for OpenAI GPT-5.x on Converse #38279's cost-map JSON and cross-region test changes stay on main. The ported tests drop the local_model_cost_map fixture parameter, which this line's conftest does not define
  6. Cherry-picked main's 728dec2 and 3947374 (-x) so every Dockerfile, the componentized migrations/Dockerfile included, installs python-3.13 and python-3.13-dev with UV_PYTHON_DOWNLOADS=0, resolving the base image digest lines to the values those commits set; the only other difference between this line's Dockerfiles and those commits' parents was those digest lines
  7. Version bump to 1.98.1 is its own commit, one patch above the latest release on this line (1.98.0)

User Flow

Before: a developer using Claude Code through a v1.98.0 proxy cannot use a Bedrock OpenAI GPT or xAI Grok model

  1. They run Claude Code with ANTHROPIC_BASE_URL pointed at the proxy and type /model us.xai.grok-4.6 (or /model us.openai.gpt-6-astra)
  2. Claude Code sends POST https://litellm-domain/v1/messages with "max_tokens": 1 as a warmup probe
  3. They see API error: 400 ... BedrockException ... Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead. and the header keeps the model they had (Sonnet 4.5)
  4. Their next message is answered by that old model, so the Bedrock GPT or Grok model never gets used
  5. With the probe fixed on its own (fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870 alone, this PR's first tip), the switch to us.openai.gpt-6-astra takes, and their first message, which Claude Code sends to POST https://litellm-domain/v1/messages with a thinking block, comes back as API Error: 400 ... Unknown parameter: 'thinking'

After: the same switch and first message succeed once the proxy runs v1.98.1

  1. They run Claude Code with ANTHROPIC_BASE_URL pointed at the proxy and type /model us.xai.grok-4.6 (or /model us.openai.gpt-6-astra)
  2. Claude Code sends POST https://litellm-domain/v1/messages with "max_tokens": 1 as a warmup probe
  3. The probe returns 200 and the header shows the new model
  4. Their first message goes to POST https://litellm-domain/v1/messages with the thinking block and is answered by that Bedrock model

Relevant issues

Backport of #41870 and of the Converse reasoning gate fix from #31884 (main). Reported by customers (Pylon #8817 for OpenAI GPT, Pylon #8821 for xAI Grok); one of them runs v1.98.0 today and is moving to the 1.100.x line

Affected release

Linear ticket

Resolves LIT-8154

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Both legs ran the same way and differ only in the commit the proxy was booted from: a no-DB proxy with 2 uvicorn workers on a random port, real Bedrock in us-east-1 (real spend, no mocks), and Claude Code 2.1.278 driven interactively under tmux. $KEY is the proxy master key and $PORT is the leg's port

config.yaml

model_list:
  - model_name: us.openai.gpt-6-astra
    litellm_params:
      model: bedrock/us.openai.gpt-6-astra
      aws_region_name: us-east-1
  - model_name: us.xai.grok-4.6
    litellm_params:
      model: bedrock/us.xai.grok-4.6
      aws_region_name: us-east-1
  - model_name: us.anthropic.claude-sonnet-4-5-20250929-v1:0
    litellm_params:
      model: bedrock/us.anthropic.claude-sonnet-4-5-20250929-v1:0
      aws_region_name: us-east-1
  - model_name: us.anthropic.claude-haiku-4-5-20251001-v1:0
    litellm_params:
      model: bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0
      aws_region_name: us-east-1

general_settings:
  master_key: sk-REDACTED

Proxy, booted from the leg's commit with a Python 3.12 venv

AWS_PROFILE=$BEDROCK_PROFILE AWS_DEFAULT_REGION=us-east-1 .venv/bin/python litellm/proxy/proxy_cli.py --config config.yaml --port $PORT --num_workers 2

Claude Code, started in an empty folder with a fresh config directory (theme picker and security notes accepted with Enter, folder trust accepted with Down then Enter)

tmux new-session -d -s claude -x 160 -y 45 -c work "env -u ANTHROPIC_API_KEY ANTHROPIC_BASE_URL=http://localhost:$PORT ANTHROPIC_AUTH_TOKEN=$KEY ANTHROPIC_MODEL=us.anthropic.claude-sonnet-4-5-20250929-v1:0 ANTHROPIC_DEFAULT_HAIKU_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0 ANTHROPIC_SMALL_FAST_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0 CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 CLAUDE_CONFIG_DIR=claude-config claude"

Before (d8f71d7)

The last curl group (the four reasoning cases) ran in a second proxy boot at this same commit after the floor-only after leg surfaced the thinking 400; every other case below is from the original before leg

Claude Code, /model us.xai.grok-4.6

  1. In the running Claude Code session, type /model us.xai.grok-4.6 and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  Sonnet 4.5 · API Usage Billing
    ❯ /model us.xai.grok-4.6
      ⎿  API error: 400 {"error":{"message":"litellm.BadRequestError: BedrockException - {"message":"The model returned the following errors:
         {\"error\":{\"code\":\"integer_below_min_value\",\"message\":\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1
         instead.\",\"param\":\"max_output_tokens\",\"type\":\"invalid_request_error\"}}"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group
         Fallbacks=None","type":null,"param":null,"code":"400"}}
    

    pr42002-c32db1b512-claude-1_98_x-before-d8f71d7bdb-grok.png

  2. Type Reply with the single word ok and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  Sonnet 4.5 · API Usage Billing
    ❯ Reply with the single word ok
      Thought for 5s
    ⏺ ok
    ✻ Churned for 5s · done 11:51 AM
    

Claude Code, /model us.openai.gpt-6-astra

  1. In the running Claude Code session, type /model us.openai.gpt-6-astra and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  Sonnet 4.5 · API Usage Billing
    ❯ /model us.openai.gpt-6-astra
      ⎿  API error: 400 {"error":{"message":"litellm.BadRequestError: BedrockException - {"message":"The model returned the following errors:
         {\"error\":{\"code\":\"integer_below_min_value\",\"message\":\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1
         instead.\",\"param\":\"max_output_tokens\",\"type\":\"invalid_request_error\"}}"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group
         Fallbacks=None","type":null,"param":null,"code":"400"}}
    

    pr42002-c32db1b512-claude-1_98_x-before-d8f71d7bdb-gpt.png

  2. Type Reply with the single word ok and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  Sonnet 4.5 · API Usage Billing
    ❯ Reply with the single word ok
      Thought for 5s
    ⏺ ok
    ✻ Worked for 5s · done 11:52 AM
    

POST /v1/messages, max_tokens 1

  1. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  2. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    

POST /v1/chat/completions, max_tokens 1 and max_completion_tokens 1

  1. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  2. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  3. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_completion_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  4. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_completion_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    

POST /v1/responses, max_output_tokens 1

  1. Run

    curl -s http://localhost:$PORT/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_output_tokens":1,"input":"hi"}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  2. Run

    curl -s http://localhost:$PORT/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_output_tokens":1,"input":"hi"}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    

Streaming, max_tokens 1

  1. Run

    curl -s -N http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' > stream-chat-us.openai.gpt-6-astra.raw

    Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  2. Run

    curl -s -N http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' > stream-chat-us.xai.grok-4.6.raw

    Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  3. Run

    curl -s -N http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"stream":true,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' > stream-messages-grok.raw

    Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.xai.grok-4.6\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    

Boundary on us.openai.gpt-6-astra, max_tokens 15 then 16

  1. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":15,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"integer_below_min_value\\\",\\\"message\\\":\\\"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 15 instead.\\\",\\\"param\\\":\\\"max_output_tokens\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  2. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":16,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-77934863-7c0c-4875-9ff9-61db84f16a7f","created":1789843726,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help?","role":"assistant"}}],"usage":{"completion_tokens":11,"prompt_tokens":7,"total_tokens":18,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":11},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    

Controls, a larger max_tokens and the Anthropic models

  1. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":64,"messages":[{"role":"user","content":"Write one sentence about the sea."}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-56db5e8c-a8d4-4493-80dd-414473d94928","created":1789843742,"model":"us.xai.grok-4.6","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"","role":"assistant","reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContentBlocks":[{"redactedContent":"…"}]}}}],"usage":{"completion_tokens":64,"prompt_tokens":25,"total_tokens":89,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":64},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":25,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","id":"msg_bdrk_011CfDJEq3zdPb6SCcFAyv3j","type":"message","role":"assistant","content":[{"type":"text","text":"Hello"}],"stop_reason":"max_tokens","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":8,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1}}
    HTTP 200
    
  3. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-haiku-4-5-20251001-v1:0","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"model":"us.anthropic.claude-haiku-4-5-20251001-v1:0","id":"msg_bdrk_m5xkyw7kjec75l4yjtwujrca7tgw3d6g7wpmmql5442mcqmkb6la","type":"message","role":"assistant","content":[{"type":"text","text":"#"}],"container":null,"stop_reason":"max_tokens","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":8,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard"}}
    HTTP 200
    

Reasoning params on us.openai.gpt-6-astra (Claude Code sends thinking on its real turns), plus the Sonnet control

  1. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":2048,"thinking":{"type":"enabled","budget_tokens":1024},"messages":[{"role":"user","content":"Reply with the single word ok"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"unknown_parameter\\\",\\\"message\\\":\\\"Unknown parameter: 'thinking'.\\\",\\\"param\\\":\\\"thinking\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  2. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":512,"reasoning_effort":"low","messages":[{"role":"user","content":"Reply with the single word ok"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"unknown_parameter\\\",\\\"message\\\":\\\"Unknown parameter: 'thinking'.\\\",\\\"param\\\":\\\"thinking\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  3. Run

    curl -s http://localhost:$PORT/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_output_tokens":512,"reasoning":{"effort":"low"},"input":"Reply with the single word ok"}' -w '\nHTTP %{http_code}\n'

    Output

    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"The model returned the following errors: {\\\"error\\\":{\\\"code\\\":\\\"unknown_parameter\\\",\\\"message\\\":\\\"Unknown parameter: 'thinking'.\\\",\\\"param\\\":\\\"thinking\\\",\\\"type\\\":\\\"invalid_request_error\\\"}}\"}. Received Model Group=us.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":null,"param":null,"code":"400"}}
    HTTP 400
    
  4. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","max_tokens":2048,"thinking":{"type":"enabled","budget_tokens":1024},"messages":[{"role":"user","content":"Reply with the single word ok"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","id":"msg_bdrk_011CfDPgSdB8sWs5z5iAqPF1","type":"message","role":"assistant","content":[{"type":"thinking","thinking":"The user is asking me to reply with the single word \"ok\". This is a straightforward request that doesn't involve anything harmful or problematic. I'll simply respond with \"ok\" as requested.","signature":"…"},{"type":"text","text":"ok"}],"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":42,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":53,"output_tokens_details":{"thinking_tokens":46}}}
    HTTP 200
    

After (0c0dda9)

Claude Code, /model us.xai.grok-4.6

  1. In the running Claude Code session, type /model us.xai.grok-4.6 and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  us.xai.grok-4.6 · API Usage Billing
    ❯ /model us.xai.grok-4.6
      ⎿  Set model to us.xai.grok-4.6 and saved as your default for new sessions
              ANTHROPIC_MODEL is set to Sonnet 4.5 — new sessions use that while it is set
                                                                                                                                                  ● high · /effort
    

    pr42002-0c0dda9780-claude-1_98_x-after-0c0dda9780-grok.png

  2. Type Reply with the single word ok and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  us.xai.grok-4.6 · API Usage Billing
    ❯ /model us.xai.grok-4.6
      ⎿  Set model to us.xai.grok-4.6 and saved as your default for new sessions
              ANTHROPIC_MODEL is set to Sonnet 4.5 — new sessions use that while it is set
    ❯ Reply with the single word ok
    ⏺ Skill(claude-api)
      ⎿  Initializing…
      ⎿  Interrupted · What should Claude do instead?
    ✻ Crunched for 6s · done 1:16 PM
    ❯ Reply with the single word ok
    ⏺ ok
    ✻ Worked for 1s · done 1:17 PM
    

Claude Code, /model us.openai.gpt-6-astra

  1. In the running Claude Code session, type /model us.openai.gpt-6-astra and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  us.openai.gpt-6-astra · API Usage Billing
    ❯ /model us.openai.gpt-6-astra
      ⎿  Set model to us.openai.gpt-6-astra and saved as your default for new sessions
              ANTHROPIC_MODEL is set to Sonnet 4.5 — new sessions use that while it is set
                                                                                                                                                  ● high · /effort
    

    pr42002-0c0dda9780-claude-1_98_x-after-0c0dda9780-gpt.png

  2. Type Reply with the single word ok and press Enter. The pane shows (header, then the lines this step added)

     ▐▛███▛█   Claude Code v2.1.278
    ▝▜██████▀  us.openai.gpt-6-astra · API Usage Billing
    ❯ /model us.xai.grok-4.6
      ⎿  Set model to us.xai.grok-4.6 and saved as your default for new sessions
              ANTHROPIC_MODEL is set to Sonnet 4.5 — new sessions use that while it is set
    ❯ Reply with the single word ok
    ⏺ Skill(claude-api)
      ⎿  Initializing…
      ⎿  Interrupted · What should Claude do instead?
    ✻ Crunched for 6s · done 1:16 PM
    ❯ Reply with the single word ok
    ⏺ ok
    ✻ Worked for 1s · done 1:17 PM
    ❯ /model us.openai.gpt-6-astra
      ⎿  Set model to us.openai.gpt-6-astra and saved as your default for new sessions
              ANTHROPIC_MODEL is set to Sonnet 4.5 — new sessions use that while it is set
    ❯ Reply with the single word ok
    ⏺ ok
    ✻ Churned for 1s · done 1:18 PM
    

POST /v1/messages, max_tokens 1

  1. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-f9c41461-c9a8-4d1e-8df8-abb887cdd51d","type":"message","role":"assistant","model":"us.openai.gpt-6-astra","stop_sequence":null,"usage":{"input_tokens":7,"output_tokens":13},"content":[{"type":"text","text":"Hi! How can I help you today?"}],"stop_reason":"end_turn"}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-c5e4f4e9-d3b5-422d-8316-1fb5f00274c9","type":"message","role":"assistant","model":"us.xai.grok-4.6","stop_sequence":null,"usage":{"input_tokens":19,"output_tokens":16},"content":[{"type":"redacted_thinking","data":"…"},{"type":"text","text":""}],"stop_reason":"max_tokens"}
    HTTP 200
    

POST /v1/chat/completions, max_tokens 1 and max_completion_tokens 1

  1. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-e85db732-122d-4548-a738-c5147fc714c8","created":1789848844,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help?","role":"assistant"}}],"usage":{"completion_tokens":11,"prompt_tokens":7,"total_tokens":18,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":11},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-e28cc747-274b-48f9-a89f-888d14a6550a","created":1789848845,"model":"us.xai.grok-4.6","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"","role":"assistant","reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContentBlocks":[{"redactedContent":"…"}]}}}],"usage":{"completion_tokens":16,"prompt_tokens":19,"total_tokens":35,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":16},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":19,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  3. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_completion_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-f619834f-b8f6-40d8-98c5-f282468ca537","created":1789848846,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help?","role":"assistant"}}],"usage":{"completion_tokens":11,"prompt_tokens":7,"total_tokens":18,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":11},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  4. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_completion_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-816f04e5-1a48-414c-9049-3052670d0743","created":1789848846,"model":"us.xai.grok-4.6","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"","role":"assistant","reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContentBlocks":[{"redactedContent":"…"}]}}}],"usage":{"completion_tokens":16,"prompt_tokens":19,"total_tokens":35,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":16},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":19,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    

POST /v1/responses, max_output_tokens 1

  1. Run

    curl -s http://localhost:$PORT/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_output_tokens":1,"input":"hi"}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"resp_x0PEAgQgno5ZyrjAOgwGif5NJ847FQ128RwSeRIGdG4PWVNJN8RXnfsduz9AeIrTq-bkWEYO6WcyhPF7HoKszCnNAnI3Rq8aQp9XVB-xv3pKWKNzt7u-jLv4GpxMD-kQUwtZznb7-rUUkb5PgG0lpo0Q54sXBsYye7Sye2hOgsLCgZRJJ2dPDR9jZ8rvigLc_kXoe91mCqFWt5Ynch2w0inG8IAqwtY_QBlzmFNt8Giw1ou-Wd2eEBi4WaOHPA8cUMc0M2ZMpuueFHxyqfSSOE5yDfPtMCOllYzx9JnTr9ZbMGGFcLu_RBwyZpMRb-Z3Pz8U0oKLn2Oluhv3zDnjmk7q2mnlsKvx5FrAO92El6xOl6BYN1lk3YUoji0Orf6OgNwIWvXT14QmgaEVpoY9Ftpnlc9X1YXl3647y1sZrwRSIOP8od0aFDqKLejCVrWGjIG7DNpa","created_at":1789848848,"error":null,"incomplete_details":null,"instructions":null,"metadata":{},"model":"us.openai.gpt-6-astra","object":"response","output":[{"type":"message","id":"chatcmpl-2d77ac25-178e-4408-8ddd-6c9ddeed8bdb",
    [... 472 characters cut for length ...]
    _tokens":0},"output_tokens":13,"output_tokens_details":{"reasoning_tokens":0,"text_tokens":13},"total_tokens":20,"cost":null},"user":null,"store":null}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_output_tokens":1,"input":"hi"}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"resp_pEMTsJp-btyfN7q0B-XowpDoVLFIinGDlzhao418Y9-zbSWsqnPeQc26DB_sqyYJOL7wT2AMNqvSfgQ5CvRswvzHiVpz4JMXjSe0nCOw8bTSvR5dGEsKbnT98vzuiOpEl2NIXeSprCVdio7X_UE--8CrNa4UcISm7byog1MRm0dR9gB5ouThsMBVn_5alMaNp0jZ9LKAlLUK1lxqIjoojthqS4F6YaKp6NYi3gdqp64F_o60ZEpvJrbcDFkXwSyfH6shpMBniE8AnuhAH2BcTL4MrrouHg_si6cSm96_4I_nhf_fG4QlnA42_KMPIUfpUMT7sT0_In1Jyn-yjFcTwP-bhJMQK_7VYICt1J5HJBhiOXiZ1mp8tvtJfvSXxGWfIqXR_Pt4yqrxou5ucIb4EyiLz70OBqWBVOTYsXga_Xp2wPvmbI8YWKIV9QEkwK_3lSolLK_0","created_at":1789848849,"error":null,"incomplete_details":null,"instructions":null,"metadata":{},"model":"us.xai.grok-4.6","object":"response","output":[{"type":"message","id":"chatcmpl-ccf9395a-e8b1-4e96-914a-e595b9ebe3bb","statu
    [... 441 characters cut for length ...]
    _tokens":0},"output_tokens":16,"output_tokens_details":{"reasoning_tokens":0,"text_tokens":16},"total_tokens":35,"cost":null},"user":null,"store":null}
    HTTP 200
    

Streaming, max_tokens 1

  1. Run

    curl -s -N http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":1,"stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' > stream-chat-us.openai.gpt-6-astra.raw

    Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body

    data: {"id":"chatcmpl-b9d2e2e4-921e-40c1-98c0-ca56fd672898","created":1789848850,"model":"us.openai.gpt-6-astra","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hi","role":"assistant"}}],"provider_specific_fields":{}}
    
    data: {"id":"chatcmpl-b9d2e2e4-921e-40c1-98c0-ca56fd672898","created":1789848850,"model":"us.openai.gpt-6-astra","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"!"}}],"provider_specific_fields":{}}
    
    data: {"id":"chatcmpl-b9d2e2e4-921e-40c1-98c0-ca56fd672898","created":1789848850,"model":"us.openai.gpt-6-astra","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" How"}}],"provider_specific_fields":{}}
    
    [... 1745 characters cut for length ...]
    kens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    
    data: [DONE]
    
    
    HTTP 200
    
  2. Run

    curl -s -N http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' > stream-chat-us.xai.grok-4.6.raw

    Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body

    data: {"id":"chatcmpl-dc137159-abc9-40eb-859d-d6e8e9e1a616","created":1789848851,"model":"us.xai.grok-4.6","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContent":{"redactedContent":"…"}},"content":"","role":"assistant"}}],"provider_specific_fields":{}}
    
    data: {"id":"chatcmpl-dc137159-abc9-40eb-859d-d6e8e9e1a616","created":1789848851,"model":"us.xai.grok-4.6","object":"chat.completion.chunk","choices":[{"finish_reason":"length","index":0,"delta":{}}],"provider_specific_fields":{}}
    
    data: {"id":"chatcmpl-dc137159-abc9-40eb-859d-d6e8e9e1a616","created":17898
    [... 511 characters cut for length ...]
    ens":0,"text_tokens":19,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    
    data: [DONE]
    
    
    HTTP 200
    
  3. Run

    curl -s -N http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":1,"stream":true,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n' > stream-messages-grok.raw

    Output, summarized from the saved stream: status, SSE line and event counts, usage chunks, and any non-SSE body

    event: message_start
    data: {"type": "message_start", "message": {"id": "msg_b7d6acb6-7a12-4271-8022-7537d05080df", "type": "message", "role": "assistant", "content": [], "model": "us.xai.grok-4.6", "stop_reason": null, "stop_sequence": null, "usage": {"input_tokens": 0, "output_tokens": 0, "cache_creation_input_tokens": 0, "cache_read_input_tokens": 0}}}
    
    event: content_block_start
    data: {"type": "content_block_start", "index": 0, "content_block": {"type": "text", "text": ""}}
    
    event: content_block_stop
    data: {"type": "content_block_stop", "index": 0}
    
    event: message_delta
    data: {"type": "message_delta", "delta": {"stop_reason": "max_tokens"}, "usage": {"input_tokens": 19, "output_tokens": 16}}
    
    event: message_stop
    data: {"type": "message_stop"}
    
    
    HTTP 200
    

Boundary on us.openai.gpt-6-astra, max_tokens 15 then 16

  1. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":15,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-20e4a49d-68be-4104-bf50-632269b37420","created":1789848853,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! What can I help you with today?","role":"assistant"}}],"usage":{"completion_tokens":14,"prompt_tokens":7,"total_tokens":21,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":14},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":16,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-e3a48706-bbfb-47ff-80ed-ae496d9c70c5","created":1789848854,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi! How can I help?","role":"assistant"}}],"usage":{"completion_tokens":11,"prompt_tokens":7,"total_tokens":18,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":11},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":7,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    

Controls, a larger max_tokens and the Anthropic models

  1. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.xai.grok-4.6","max_tokens":64,"messages":[{"role":"user","content":"Write one sentence about the sea."}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-3f2087c6-9022-4b53-9ff7-8b9f1c449d3d","created":1789848854,"model":"us.xai.grok-4.6","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"","role":"assistant","reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"…"}],"provider_specific_fields":{"reasoningContentBlocks":[{"redactedContent":"…"}]}}}],"usage":{"completion_tokens":64,"prompt_tokens":25,"total_tokens":89,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":64},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":25,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","id":"msg_bdrk_011CfDQkeS4pZ8imFzPA12X7","type":"message","role":"assistant","content":[{"type":"text","text":"Hello"}],"stop_reason":"max_tokens","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":8,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1}}
    HTTP 200
    
  3. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-haiku-4-5-20251001-v1:0","max_tokens":1,"messages":[{"role":"user","content":"hi"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"model":"us.anthropic.claude-haiku-4-5-20251001-v1:0","id":"msg_bdrk_peqz2yireu2q6bgq2paoe54u2o7euspxr3s4femfcsebt36to2na","type":"message","role":"assistant","content":[{"type":"text","text":"Hello"}],"container":null,"stop_reason":"max_tokens","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":8,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard"}}
    HTTP 200
    

Reasoning params on us.openai.gpt-6-astra (Claude Code sends thinking on its real turns), plus the Sonnet control

  1. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":2048,"thinking":{"type":"enabled","budget_tokens":1024},"messages":[{"role":"user","content":"Reply with the single word ok"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-ce35f11b-1a98-4933-a7b1-c98f5eee91b0","type":"message","role":"assistant","model":"us.openai.gpt-6-astra","stop_sequence":null,"usage":{"input_tokens":12,"output_tokens":5},"content":[{"type":"text","text":"ok"}],"stop_reason":"end_turn"}
    HTTP 200
    
  2. Run

    curl -s http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_tokens":512,"reasoning_effort":"low","messages":[{"role":"user","content":"Reply with the single word ok"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"chatcmpl-8a300413-49ae-4770-9b0e-a59560685fc3","created":1789848876,"model":"us.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"ok","role":"assistant"}}],"usage":{"completion_tokens":5,"prompt_tokens":12,"total_tokens":17,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":5},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":12,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
    HTTP 200
    
  3. Run

    curl -s http://localhost:$PORT/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.openai.gpt-6-astra","max_output_tokens":512,"reasoning":{"effort":"low"},"input":"Reply with the single word ok"}' -w '\nHTTP %{http_code}\n'

    Output

    {"id":"resp_jiABxS6Adhe4TF7VckclNPAsuUnr8z4S7Epkqzdx0RcrFJpwgjPdfLxvu7nCv0UewTfz3E4kfEDsOPnKi-KjDlkO6BEPUJwFNykfbeDEQ17x9e0EV_0_506JchtH2mr_Hz7Q71pcrIveGH7ZkkOsJm4jS8oYOWufRv-HoeuDDoS-ItBoi4Qd7rMKzSkNltBmOh5FaBuf0RJWtAzpXsSWy4Q9nBd-Hqq1WG3wi4XJyKhujOSvDdBM93KdSjipyLttKYRbC_twO9zy6Q5uJ11fGeDtV7QyqkSO44oKrz3wh5Paaaa2IC8vk2zxIiDu7MGQf1ZldWy1ZKxH8tcn9SvmXvKtIDpLolpdq3-Hs06HKv7Hk5Q6TVYeW7uFUL5eiOGjGmSoNhEIMUVmjJe6-bqUEd_nY91HuyG0qnHwtoEEzuzA6eFH7fdo_w08Exo1Mekl_UqfAg7Y","created_at":1789848877,"error":null,"incomplete_details":null,"instructions":null,"metadata":{},"model":"us.openai.gpt-6-astra","object":"response","output":[{"type":"message","id":"chatcmpl-f7605042-8690-48c1-8ed8-904fb1d52c17",
    [... 445 characters cut for length ...]
    te_tokens":0},"output_tokens":5,"output_tokens_details":{"reasoning_tokens":0,"text_tokens":5},"total_tokens":17,"cost":null},"user":null,"store":null}
    HTTP 200
    
  4. Run

    curl -s http://localhost:$PORT/v1/messages -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","max_tokens":2048,"thinking":{"type":"enabled","budget_tokens":1024},"messages":[{"role":"user","content":"Reply with the single word ok"}]}' -w '\nHTTP %{http_code}\n'

    Output

    {"model":"us.anthropic.claude-sonnet-4-5-20250929-v1:0","id":"msg_bdrk_011CfDQoHVZi46Y9ftwjMoFn","type":"message","role":"assistant","content":[{"type":"thinking","thinking":"The user is asking me to reply with the single word \"ok\". This is a straightforward request that doesn't involve anything harmful or problematic. I should simply respond with \"ok\" as requested.","signature":"…"},{"type":"text","text":"ok"}],"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":42,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":53,"output_tokens_details":{"thinking_tokens":46}}}
    HTTP 200
    

Observations from the run

Seen on the after leg. The first two follow from the floor this PR adds and do not affect the probe, which only needs the 200; the control line shows Anthropic models are left alone; the rest are Claude Code behavior this PR neither causes nor changes

  • Grok spends all 16 floor tokens on redacted reasoning
  • gpt-6-astra answers in 11 to 14 tokens under the floor
  • Sonnet and Haiku controls still return exactly 1 output token
  • Grok's first ok turn became a Skill call; declined, re-sent
  • Second /model opens a Switch model? dialog
  • Zero 400 lines in the proxy log during the run

Type

🐛 Bug Fix

Caveats (if any)

Low

  • Only openai.gpt-* and xai.grok-* models on Converse are clamped; inference-profile ARNs carrying the model id are matched, opaque application-inference-profile ARNs are not
  • The explicit bedrock/openai/... (Mantle) and bedrock/invoke/... routes for the same models are untouched and still forward a sub-16 value as is, same as fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse #41870
  • A max_tokens below 16 on these models now returns up to 16 tokens instead of a 400; no previously working request changes, since Bedrock rejected every value below 16
  • The reasoning gate now matches any openai.gpt-<digit> model id, so a future GPT-7 model on Converse gets the same reasoning.effort mapping without a further change, same as main
  • CircleCI is red on ten jobs at this tip, every one in a test that calls a live provider outside the Bedrock Converse path. Seven of them fail the identical 51 tests that pipeline 89882 failed at the floor-only tip c32db1b, before the reasoning gate and the Docker pins landed: llm_translation_testing (Groq, Together AI, and OpenAI web search cases), local_testing_part1 and local_testing_part2 (Together AI, cohere.command-r-plus-v1:0 no longer served by Bedrock, OpenAI timeout tests, a Groq stream chunk case), litellm_router_testing (wildcard routing, timeout, and batch completion against live providers), logging_testing (OpenAI web search cost tracking), image_gen_testing (Google and Vertex image generation), and pass_through_unit_testing (the assistants passthrough case against api.openai.com)
  • The other three need the database image, so they were not_run while build_docker_database_image was red and only started running once the python-3.13 pin landed: e2e_openai_endpoints fails test_bad_request_error (OpenAI now answers an unknown model on the Responses API with a 404 where the test expects a 400; the same test is red on fix(bedrock): backport #41870 and the GPT-6 reasoning gate fix to stable/1.100.x for v1.100.2 #42000 and on fix(responses): backport mid-stream content_policy_violation fallback routing to stable/1.100.x #41208, the last PR merged into stable/1.100.x), proxy_pass_through_endpoint_tests fails its two Ruby OpenAI Assistants specs on a 404 from api.openai.com, and build_and_test stops at test_proxy_all_models because Groq no longer serves llama-3.1-8b-instant
  • build_docker_database_image was red at the first tip (the unpinned wolfi-base image now ships CPython 3.14, where uvloop 0.21.0 does not build) and is green at this tip with the python-3.13 pin, so CI verifies that image. migrations/Dockerfile, backend/Dockerfile, gateway/Dockerfile, the root Dockerfile, and docker/Dockerfile.non_root (the last two with the dashboard stage stubbed) were also built locally on arm64 at this tip and import cleanly on Python 3.13.15; only the real dashboard build and the amd64 builds are first built by the release workflow

Blast radius, /live-pr-risk at 0c0dda9 against base d8f71d7

Verdict: no dependent path that worked on the base answers worse on the head. Two no-DB proxies (2 workers each) ran 31 scenarios side by side against real Bedrock, and six more cases ran through litellm.completion on each tree, every upstream request passing a forwarding recorder that logged what LiteLLM sent before re-signing it for AWS. Both sides sent the same header set, and the same number of upstream requests in every scenario but the raw thinking case below, and the only body differences are the two this PR names: inferenceConfig.maxTokens raised to 16, and the GPT thinking field replaced by reasoning.effort

Breaking: none observed

Backward incompatible, all benign, and all the behavior main already ships from #41870, #31884, and #38279 (each approved by mateo-berri on main):

  • max_tokens of 0, -1, and JSON true on us.openai.gpt-6-astra were Bedrock 400s on the base and are 200s on the head, sent as maxTokens: 16. true passes the isinstance(value, int) guard because a Python bool is an int. "1", 1.5, and [1] are forwarded unchanged and 400 identically on both sides
  • GPT-6 requests carrying reasoning_effort on POST /v1/chat/completions, or a thinking block on POST /v1/messages, were 400 Unknown parameter: 'thinking' on the base and are 200 on the head, and us.openai.gpt-5.6-luna with reasoning_effort through the SDK goes the same way; the upstream body now carries additionalModelRequestFields: {"reasoning": {"effort": "low"}} where the base sent {"thinking": {"type": "enabled", "budget_tokens": 1024}}
  • A raw thinking param on a chat completion for GPT-6 or GPT-5.6 stays a 400, but it is now LiteLLM's UnsupportedParamsError with no upstream call where the base forwarded it and relayed Bedrock's 400
  • GET /utils/supported_openai_params and GET /model_group/info no longer list thinking and output_config in supported_openai_params for the GPT model groups (the plain id, the bedrock/converse/ prefix, and an ARN deployment with model_info.base_model); reasoning_effort stays. Grok, Haiku, Nova, gpt-oss, Sonnet, and the ARN deployment without base_model are identical on both sides. The dashboard model hub only displays this list, so it shows two fewer params for these models

Regression risk:

  • Spend rows for the floored requests were not read back, because the rig ran without a database; the usage blocks in the head's 200 responses reported 5 output tokens for GPT-6 and 16 for Grok
  • The Docker changes move every image to a newer wolfi-base digest and a pinned CPython 3.13. CircleCI built the database image at this tip and ran its proxy suites on it. This run also built the other five locally on arm64 at this tip: migrations/Dockerfile, backend/Dockerfile, and gateway/Dockerfile as is, and the root Dockerfile and docker/Dockerfile.non_root with only their dashboard npm run build stage replaced by a stub page. All five report Python 3.13.15, litellm 1.98.1, uvloop 0.21.0, and prisma 0.11.0, the backend and gateway images import backend.main and gateway.main, and the main image booted with no database and answered GET /health/liveliness with 200. The real dashboard build and every amd64 build are first built by the release workflow. On the base the database image did not build (uvloop 0.21.0 on CPython 3.14), and the others install the same unpinned python3

Dependency graph:

  • map_openai_params maxTokens branch via /v1/chat/completions (max_tokens, max_completion_tokens), /v1/messages, /v1/responses, converse-stream, the bedrock/converse/ prefix, a router fallback replay (2 upstream calls on both sides), and us.openai.gpt-5.6-luna through the SDK: unit tested and verified live, base 400 and head 200
  • _is_openai_gpt_reasoning_model at its three call sites (reasoning_effort mapping, get_supported_openai_params, the thinking skip): unit tested and verified live on GPT-6 and GPT-5.6
  • Boundary and controls (max_tokens 16 and 17, Haiku at 1, openai.gpt-oss-20b-1:0 at 1 and with reasoning_effort, Sonnet 4.5 with reasoning_effort, unauthenticated 401 with 0 upstream calls): verified live, identical on both sides
  • AmazonInvokeNovaConfig, the only subclass of AmazonConverseConfig: verified live on bedrock/invoke/us.amazon.nova-lite-v1:0, maxTokens: 1 on both sides
  • Opaque application inference profile ARN for GPT-6, with and without model_info.base_model: unit tested as not floored and verified live, 400 on both sides, so this PR leaves that deployment shape broken the way it already was
  • Grok with reasoning_effort: verified live, still sent as Anthropic-style thinking and 200 on both sides
  • GET /model_group/info and GET /utils/supported_openai_params: verified live through the API, difference listed above; the dashboard model hub page that renders the list was not opened
  • BedrockConverseGptReasoningEffortBlock in litellm/types/llms/bedrock.py: a new type with one user, the mapping above
  • pyproject.toml and uv.lock: version bump only, no resolved package changes

Not verified: spend and usage rows in the database, the Admin UI pages that read them, the real dashboard build inside the Docker images and any amd64 build outside CircleCI's database image, and the user docs in litellm-docs. The legacy tests/llm_translation and tests/local_testing suites ran on CircleCI and are covered in the bullets above

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • 0c0dda9 passes /live-pr-risk

…and xAI Grok models on Converse (#41870)

Backport of #41870 to stable/1.98.x. Cherry-picked from a6e3a72 (main) with -m 1.

converse_transformation.py conflicted because this line has no `import re` and no
_is_openai_gpt_reasoning_model helper next to the insertion point. The resolution adds
exactly the four hunks #41870 merged: the import, the 16-token constant,
_requires_min_max_tokens, and the clamped maxTokens assignment. The test file applied clean.
@devin-ai-integration

devin-ai-integration Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with the prior migration-image inconsistency fixed and no new actionable defects identified.

Findings

  1. P1 Opaque profiles bypass clamp ▶
  2. P1 Migration image remains unpinned ▶
  3. P2 Model capability is hardcoded ▶

Summary

This PR backports Bedrock Converse compatibility fixes and release packaging updates to stable/1.98.x.

  • Clamps sub-16 output-token limits for identifiable Bedrock OpenAI GPT and xAI Grok model families.
  • Maps reasoning effort correctly for Bedrock OpenAI GPT reasoning models and prevents forwarding incompatible Anthropic thinking fields.
  • Adds focused transformation tests for model families, parameter variants, inference-profile identifiers, and unaffected models.
  • Pins the shipped Wolfi images, including the migration image, to Python 3.13 and releases version 1.98.1.

Reviews (3) · Last reviewed commit: "fix(docker): bump wolfi-base for glibc 2..."

Comment on lines 866 to +870
if param == "max_tokens" or param == "max_completion_tokens":
optional_params["maxTokens"] = value
optional_params["maxTokens"] = (
max(value, BEDROCK_OPENAI_COMPAT_MIN_MAX_TOKENS)
if isinstance(value, int) and self._requires_min_max_tokens(model)
else value

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Opaque profiles bypass clamp

Opaque application profile ARNs hide the backing model, so max_tokens=1 remains unclamped and Bedrock rejects compatible-model requests

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Backport of #41870, where this was withdrawn. The regex clamps ARNs embedding the model id; opaque application profiles carry none, so get_model_info cannot resolve them

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That’s fair. Opaque application-profile ARNs contain no model identifier, and get_model_info cannot resolve the backing model from them, so there is no reliable way to apply this model-specific clamp without risking unrelated profiles. The regex intentionally covers only identifiable model IDs, including ARNs that embed them. I agree this finding should be withdrawn.

Comment on lines +285 to +287
@staticmethod
def _requires_min_max_tokens(model: str) -> bool:
return re.search(r"openai\.gpt-\d|xai\.grok-", model) is not None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Model capability is hardcoded

This name-pattern check violates the directive to store model-specific flags in model_prices_and_context_window.json and read them through get_model_info; fix before merging

Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same shape as main's merged #41870. The pattern matches model families, not individual models, so new GPT or Grok releases need no cost-map entry

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That’s a fair distinction. I rechecked #41870: _requires_min_max_tokens is not identifying a finite catalog of model capabilities; it encodes Bedrock Converse’s shared minimum-token contract for the openai.gpt-* and xai.grok-* model families. Using get_model_info would require adding or updating every future model ID and would make the fix stale, whereas family matching intentionally covers new releases and mirrors the merged implementation on main. I’m withdrawing this finding; no change is needed.

…oning gate

Backports the Converse part of fbc6fb5 from main (PR #31884) together with the OpenAI
GPT reasoning gate this line never had (PR #38279: 74e86d3, 9cc276a, 418012a):
reasoning_effort maps to reasoning.effort for openai.gpt-<digit> models and Anthropic's
thinking block is skipped for them. Without it a GPT-6 model fell through to Anthropic's
thinking block and Bedrock rejected the first real turn after a Claude Code /model switch
with 400 Unknown parameter: 'thinking'. #38279's cost-map JSON and cross-region test
changes, the Nova 2 tool_choice registry keys, and the invoke json_mode forwarding stay
on main
(cherry picked from commit 14f392b)
(cherry picked from commit 728dec2)
@mateo-berri mateo-berri changed the title fix(bedrock): backport #41870 to stable/1.98.x for v1.98.1 fix(bedrock): backport #41870, the GPT-6 reasoning gate fix, and the Python 3.13 image pin to stable/1.98.x for v1.98.1 Sep 19, 2026
@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 19, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

Comment thread Dockerfile
Comment on lines +43 to +44
python-3.13 \
python-3.13-dev \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Migration image remains unpinned

The primary, backend, gateway, database, and non-root images now install Python 3.13 explicitly, but the separately shipped migrations/Dockerfile remains on the old Wolfi digest and installs the moving python3 metapackage. It can therefore select a different Python version, including an incompatible newer interpreter, causing the migration image to fail during build or startup. Apply the same base-image and Python 3.13 pin to the migration image.

…3 in migrations image

(cherry picked from commit 3947374)
@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 19, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 0c0dda9. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 86f64b0 into stable/1.98.x Sep 19, 2026
47 of 58 checks passed
@mateo-berri
mateo-berri deleted the litellm_cherrypick_1_98_x branch September 19, 2026 22:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant