Skip to content

feat(bedrock): serve gpt-5.6+ chat completions natively by default, with chat_completions/ opt-in for gpt-oss and grok - #40775

Closed
mateo-berri wants to merge 40 commits into
mainfrom
litellm_bedrock_grok_chat_completions
Closed

mateo-berri wants to merge 40 commits into
mainfrom
litellm_bedrock_grok_chat_completions

Conversation

@mateo-berri

@mateo-berri mateo-berri commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Bedrock GPT 5.6, 6, and 6.1 chat calls were rewritten to Converse and back
  • AWS serves them natively on /openai/v1/chat/completions
  • Each rewrite adds latency and a place for a param to get lost
  • Converse refuses temperature and top_p on these models at every effort

How it solves it:

  • An unprefixed bedrock/<gpt-5.6+ id> goes native when its cost-map row lists /v1/chat/completions
  • Converse still serves guardrailConfig, application inference profile ARNs, the other Converse-only body keys, stop, top_k, and tools with reasoning on
  • Under reasoning_effort: "none", temperature, top_p, the penalties, and logprobs pass through natively
  • With reasoning on they answer a 400 naming them, or are dropped under drop_params
  • A non-string reasoning_effort answers 400 in litellm, or drops under drop_params
  • bedrock/chat_completions/<id> opts gpt-oss and Grok in; bedrock/converse/<id> pins Converse
  • A deployment whose model_id names an application inference profile ARN keeps Converse at the ARN URL, like the bedrock/arn:... form
  • A SigV4 deployment with api_key: "" is signed, never sent as an empty bearer
  • Remote http(s):// images are inlined as data: URLs before the native call
  • A json_schema response_format goes native on GPT 5.6+ and Grok; json_object keeps Converse's handling
  • Reasoning tokens come back in usage, and service_tier as AWS sent it

Intentional product change: unprefixed bedrock/ deployments of the us. and global. GPT 5.6, 6, and 6.1 ids move from Converse to runtime Chat Completions, so response ids, service_tier, and the reasoning fields change shape as the User Flow shows; gpt-oss and Grok stay on Converse unless the deployment is prefixed chat_completions/

User Flow

Before: a developer whose app sends GPT 5.6 chat completions through the gateway gets a Converse-shaped answer, a 400 the moment they send a sampling param, and a silent 200 when they send a malformed reasoning_effort

  1. They send POST https://litellm-domain/v1/chat/completions with "model": "gpt-5.6-sol" (a bedrock/global.openai.gpt-5.6-sol deployment) and a user message
  2. A 200 comes back under a gateway-minted chatcmpl-<uuid> id with no service_tier and no reasoning_tokens in usage
  3. They add "temperature": 0.2, "top_p": 0.9, and "reasoning_effort": "none" and get 400 bedrock does not support parameters: ['temperature', 'top_p']
  4. The same request on gpt-6-sol with drop_params on still gets AWS's 400 This model doesn't support the temperature field
  5. They send "reasoning_effort": 3 (an int, or a list such as ["high"]) on gpt-5.6-sol and get 200 at AWS's default effort, with nothing in the answer saying the value was ignored; the same on a drop_params deployment
  6. A request carrying guardrailConfig, or aimed at an application inference profile ARN deployment, answers 200

After: the same developer gets Bedrock's own Chat Completions answer, with sampling params honored under reasoning_effort: "none", refused by name otherwise, and a malformed reasoning_effort refused by name unless the deployment drops params

  1. They send POST https://litellm-domain/v1/chat/completions with "model": "gpt-5.6-sol" (a bedrock/global.openai.gpt-5.6-sol deployment) and a user message
  2. A 200 comes back under Bedrock's long chatcmpl-... id with "service_tier": "default" and reasoning_tokens in usage
  3. They add "temperature": 0.2, "top_p": 0.9, and "reasoning_effort": "none" and get 200 with both honored; without none the 400 names the params and that way out
  4. The same request on gpt-6-sol with drop_params on gets 200 with the params stripped
  5. They send "reasoning_effort": 3 on gpt-5.6-sol and get 400 with the message "global.openai.gpt-5.6-sol takes reasoning_effort as a string on Bedrock's Chat Completions endpoint, not int. Send one of its named efforts, or set litellm.drop_params = True to drop it" and no call made to AWS; the same request on a drop_params deployment, or with "drop_params": true in the body, gets 200 at AWS's default effort as before
  6. A request carrying guardrailConfig, or aimed at an application inference profile ARN deployment, answers 200 exactly as before, through Converse

Linear ticket

Resolves LIT-8684

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Last updated: 6befa50. QA and /live-pr-risk ran on 6daf0b7, the last commit with a product diff

Before is the merge base c168199 and After is the tip 6daf0b7, each booted from its own worktree with python litellm/proxy/proxy_cli.py --config config.yaml --port <port> --num_workers 2 --use_prisma_db_push on 2026-10-02 (Before on 36349, After on 20753), with DATABASE_URL (one Postgres database per leg) and LITELLM_LICENSE exported explicitly. Both read the config below and the same .env names: LITELLM_MASTER_KEY, LITELLM_LOCAL_MODEL_COST_MAP=True, AWS_REGION_NAME=us-west-2, and AWS_BEARER_TOKEN_BEDROCK. BEDROCK_BASE points at a small forwarding recorder on 127.0.0.1 (one per leg, 37807 and 28723) that logs each request with the status AWS answered and passes it unchanged to https://bedrock-runtime.us-west-2.amazonaws.com, so step 3 of every case is the AWS path and body as they left the proxy (the body digest shows the model, the message count, and the params that matter; a case showing the same path twice is the proxy's retry after AWS answered 500 to the first attempt, which the recorder logged on both legs). The gpt-5.6-sol-sigv4-blank-key deployment has no api_base and signs the Host header, so its calls went straight to AWS and its step 3 reads "none". Every call reached AWS and cost real money. The same 40 cases ran once more on a local merge of 6daf0b7 into main d729f97 (33d3c0f5ff, port 41217): every recorded AWS path and body matched the After leg (on 4 of them AWS answered 500 to a first attempt on one leg and the proxy retried, so only the attempt count differs), and only the model's answer text differed, on 10 of them (qa/threeway-6daf0b7136.txt)

model_list:
  - model_name: gpt-5.6-sol
    litellm_params:
      model: bedrock/global.openai.gpt-5.6-sol
      api_base: os.environ/BEDROCK_BASE
  - model_name: gpt-6-sol
    litellm_params:
      model: bedrock/global.openai.gpt-6-sol
      api_base: os.environ/BEDROCK_BASE
  - model_name: gpt-6.1-sol
    litellm_params:
      model: bedrock/global.openai.gpt-6.1-sol
      api_base: os.environ/BEDROCK_BASE
  - model_name: gpt-5.5
    litellm_params:
      model: bedrock/global.openai.gpt-5.5
      api_base: os.environ/BEDROCK_BASE
  - model_name: gpt-5.6-sol-drop
    litellm_params:
      model: bedrock/global.openai.gpt-5.6-sol
      api_base: os.environ/BEDROCK_BASE
      drop_params: true
  - model_name: gpt-6-sol-drop
    litellm_params:
      model: bedrock/global.openai.gpt-6-sol
      api_base: os.environ/BEDROCK_BASE
      drop_params: true
  - model_name: gpt-5.6-sol-converse
    litellm_params:
      model: bedrock/converse/global.openai.gpt-5.6-sol
      api_base: os.environ/BEDROCK_BASE
  - model_name: gpt-5.6-sol-profile
    litellm_params:
      model: bedrock/arn:aws:bedrock:us-west-2:888602223428:application-inference-profile/6mv9jwwiiudo
      api_base: os.environ/BEDROCK_BASE
  - model_name: gpt-5.6-sol-model-id
    litellm_params:
      model: bedrock/global.openai.gpt-5.6-sol
      model_id: arn:aws:bedrock:us-west-2:888602223428:application-inference-profile/6mv9jwwiiudo
      api_base: os.environ/BEDROCK_BASE
  - model_name: gpt-5.6-sol-sigv4-blank-key
    litellm_params:
      model: bedrock/global.openai.gpt-5.6-sol
      api_key: ""
      aws_profile_name: litellm-bedrock-888
      aws_region_name: us-west-2
  - model_name: gpt-oss-20b
    litellm_params:
      model: bedrock/openai.gpt-oss-20b-1:0
      api_base: os.environ/BEDROCK_BASE
  - model_name: gpt-oss-20b-native
    litellm_params:
      model: bedrock/chat_completions/openai.gpt-oss-20b-1:0
      api_base: os.environ/BEDROCK_BASE
  - model_name: grok-4.6
    litellm_params:
      model: bedrock/us.xai.grok-4.6
      api_base: os.environ/BEDROCK_BASE
  - model_name: grok-4.6-native
    litellm_params:
      model: bedrock/chat_completions/us.xai.grok-4.6
      api_base: os.environ/BEDROCK_BASE
  - model_name: gpt-5.6-sol-runtime-endpoint
    litellm_params:
      model: bedrock/global.openai.gpt-5.6-sol
      api_base: http://127.0.0.1:9
      aws_bedrock_runtime_endpoint: os.environ/BEDROCK_BASE
  - model_name: gpt-5.6-sol-converse-runtime-endpoint
    litellm_params:
      model: bedrock/converse/global.openai.gpt-5.6-sol
      api_base: http://127.0.0.1:9
      aws_bedrock_runtime_endpoint: os.environ/BEDROCK_BASE
  - model_name: gpt-overlong-version
    litellm_params:
      model: bedrock/openai.gpt-777...777
      api_base: os.environ/BEDROCK_BASE
  - model_name: gpt-5.6-sol-env-endpoint
    litellm_params:
      model: bedrock/global.openai.gpt-5.6-sol
      api_base: http://127.0.0.1:9

gpt-overlong-version names bedrock/openai.gpt- followed by 4,301 sevens (cut above for width), the worst case for the GPT version regex. The application inference profile 6mv9jwwiiudo wraps global.openai.gpt-5.6-sol in the same account, and the gpt-5.6-sol-model-id deployment reaches it through model_id instead. $TOOLS is one get_weather tool, $GUARD is a Bedrock guardrail in the same account (49e7k3qphvzq, version DRAFT, trace on), and $FIX is a 96x64 PNG, red on the left and blue on the right, served by GitHub as application/octet-stream behind a 302:

[{"type":"function","function":{"name":"get_weather","description":"Get the current weather for a city","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}}]
{"guardrailIdentifier":"49e7k3qphvzq","guardrailVersion":"DRAFT","trace":"enabled"}

Every curl below also carried -H 'Authorization: Bearer $LITELLM_MASTER_KEY' -H 'Content-Type: application/json', dropped for width. Bodies are trimmed for width too: created and the token detail objects are cut, empty fields are cut, reasoning payloads show as …, a Responses answer keeps its output text, reasoning, and usage, a stream shows its first two and last three data: lines with the total count, and an error shows its first 400 characters

Before (c168199)

chat_gpt56_basic

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-e30ac3fb-be8f-4080-94d1-0819dfc3df11", "model": "gpt-5.6-sol", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 34, "total_tokens": 46}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}; POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt6_basic

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-c93c675a-c01f-467b-8aed-b97375a7b41e", "model": "gpt-6-sol", "finish_reason": "stop", "role": "assistant", "content": "Hi to you.", "usage": {"prompt_tokens": 12, "completion_tokens": 8, "total_tokens": 20, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt61_basic

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6.1-sol","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-76c949a6-15ef-4c63-82cb-a39a460f9c24", "model": "gpt-6.1-sol", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 31, "total_tokens": 43}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-6.1-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt6_json_schema

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Name the capital of France."}],"response_format":{"type":"json_schema","json_schema":{"name":"capital","strict":true,"schema":{"type":"object","properties":{"city":{"type":"string"},"country":{"type":"string"}},"required":["city","country"],"additionalProperties":false}}}}'
  2. Response:
    {"id": "chatcmpl-7625e2bc-3584-44fb-b98f-aee0a6f65ec5", "model": "gpt-6-sol", "finish_reason": "stop", "role": "assistant", "content": "{\"city\": \"Paris\", \"country\": \"France\"}", "usage": {"prompt_tokens": 51, "completion_tokens": 23, "total_tokens": 74, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "toolConfig.tools": ["json_tool_call"]}

chat_gpt56_tools_reasoning_none

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"What is the weather in Paris right now?"}],"tools":'"$TOOLS"',"reasoning_effort":"none"}'
  2. Response:
    {"id": "chatcmpl-05bc9a77-c76a-42a3-9bb1-7271a779933b", "model": "gpt-5.6-sol", "finish_reason": "tool_calls", "role": "assistant", "content": "", "tool_calls": [{"id": "call_4968fbada9885d4bb2f5403658db003c", "name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}], "usage": {"prompt_tokens": 53, "completion_tokens": 18, "total_tokens": 71, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "additionalModelRequestFields": {"reasoning": {"effort": "none"}}, "toolConfig.tools": ["get_weather"]}

chat_gpt6_tools_reasoning_low

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"What is the weather in Paris right now?"}],"tools":'"$TOOLS"',"reasoning_effort":"low"}'
  2. Response:
    {"id": "chatcmpl-4e332db9-52fa-4497-aa52-2abcd5a5cef9", "model": "gpt-6-sol", "finish_reason": "tool_calls", "role": "assistant", "content": "", "tool_calls": [{"id": "call_22de624281db5c06b8c6d06c6774f70a", "name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}], "usage": {"prompt_tokens": 53, "completion_tokens": 18, "total_tokens": 71, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "additionalModelRequestFields": {"reasoning": {"effort": "low"}}, "toolConfig.tools": ["get_weather"]}

chat_gpt61_tools

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6.1-sol","messages":[{"role":"user","content":"What is the weather in Paris right now?"}],"tools":'"$TOOLS"'}'
  2. Response:
    {"id": "chatcmpl-d8e0aec7-a662-4971-940b-f630e590f827", "model": "gpt-6.1-sol", "finish_reason": "tool_calls", "role": "assistant", "content": "", "tool_calls": [{"id": "call_74b66dc8f89d5646a367cbdbfb8aefc4", "name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}], "usage": {"prompt_tokens": 53, "completion_tokens": 18, "total_tokens": 71, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-6.1-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "toolConfig.tools": ["get_weather"]}

chat_gpt56_guardrail

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"guardrailConfig":'"$GUARD"'}'
  2. Response:
    {"id": "chatcmpl-63c7b6ec-4c54-4411-af72-e0f09f38643f", "model": "gpt-5.6-sol", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 39, "total_tokens": 51}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "guardrailConfig": {"guardrailIdentifier": "49e7k3qphvzq", "guardrailVersion": "DRAFT", "trace": "enabled"}}; POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "guardrailConfig": {"guardrailIdentifier": "49e7k3qphvzq", "guardrailVersion": "DRAFT", "trace": "enabled"}}; POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "guardrailConfig": {"guardrailIdentifier": "49e7k3qphvzq", "guardrailVersion": "DRAFT", "trace": "enabled"}}

chat_gpt56_app_profile

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-profile","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-c48e8b10-896b-44ed-82f0-54260b66bc33", "model": "gpt-5.6-sol-profile", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 33, "total_tokens": 45}}
    
  3. Recorded outbound: POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt56_model_id

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-model-id","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-3f901ca7-bd94-454e-a6ec-562271efbe9b", "model": "gpt-5.6-sol-model-id", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 30, "total_tokens": 42}}
    
  3. Recorded outbound: POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt56_sigv4_blank_key

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-sigv4-blank-key","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-91ad8a0e-8f54-43a0-9228-d89079c12ac6", "model": "gpt-5.6-sol-sigv4-blank-key", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 29, "total_tokens": 41}}
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt56_converse_prefix

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-converse","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-0615612d-410e-46b9-9efc-6c1c285e9240", "model": "gpt-5.6-sol-converse", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 31, "total_tokens": 43}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt55_basic

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.5","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-61fd2863-ccf2-4e11-917d-c6ceff919250", "model": "gpt-5.5", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend.", "usage": {"prompt_tokens": 12, "completion_tokens": 9, "total_tokens": 21, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.5/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gptoss_basic

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-oss-20b","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-a4e414a1-a69b-49da-b223-9f2e307bcaf7", "model": "gpt-oss-20b", "finish_reason": "stop", "role": "assistant", "content": "Hi there friend!", "reasoning_content": "…", "usage": {"prompt_tokens": 73, "completion_tokens": 192, "total_tokens": 265, "reasoning_tokens": 184}}
    
  3. Recorded outbound: POST /model/openai.gpt-oss-20b-1%3A0/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gptoss_native_prefix

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-oss-20b-native","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"error": {"message": "litellm.BadRequestError: BedrockException - {\"message\":\"The provided model identifier is invalid.\"}\n\nLiteLLM: model group 'gpt-oss-20b-native' failed with the error above. No fallback was attempted.", "type": "invalid_request_error", "param": null, "code": "400"}}
    
  3. Recorded outbound: POST /model/chat_completions%2Fopenai.gpt-oss-20b-1%3A0/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_grok_basic

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"grok-4.6","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-77f9dae5-2968-4468-a305-cb11d85ab640", "model": "grok-4.6", "finish_reason": "stop", "role": "assistant", "content": "Hello there friend.", "usage": {"prompt_tokens": 24, "completion_tokens": 372, "total_tokens": 396}}
    
  3. Recorded outbound: POST /model/us.xai.grok-4.6/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_grok_native_prefix

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"grok-4.6-native","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"error": {"message": "litellm.NotFoundError: BedrockException - Bedrock Invoke HTTPX: Unknown provider=None, model=chat_completions/us.xai.grok-4.6. Try calling via converse route - `bedrock/converse/<model>`.\n\nLiteLLM: model group 'grok-4.6-native' failed with the error above. No fallback was attempted.", "type": "invalid_request_error", "param": null, "code": "404"}}
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt56_reasoning_low

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"What is 17*23? Answer with the number only."}],"reasoning_effort":"low"}'
  2. Response:
    {"id": "chatcmpl-ed196ed3-06f5-46e7-8dff-3a5eda77b87c", "model": "gpt-5.6-sol", "finish_reason": "stop", "role": "assistant", "content": "391", "usage": {"prompt_tokens": 19, "completion_tokens": 5, "total_tokens": 24, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "additionalModelRequestFields": {"reasoning": {"effort": "low"}}}; POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "additionalModelRequestFields": {"reasoning": {"effort": "low"}}}

chat_gpt56_temperature

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9}'
  2. Response:
    {"error": {"message": "litellm.UnsupportedParamsError: bedrock does not support parameters: ['temperature', 'top_p'], for model=global.openai.gpt-5.6-sol. To drop these, set `litellm.drop_params=True` or for proxy:\n\n`litellm_settings:\n drop_params: true`\n. \n If you want to use these params dynamically send allowed_openai_params=['temperature', 'top_p'] in your request.\n\nLiteLLM: model group
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt56_temperature_drop_params

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9}'
  2. Response:
    {"id": "chatcmpl-b1f1c6ab-9930-4a56-a990-892a19db2fd8", "model": "gpt-5.6-sol-drop", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 34, "total_tokens": 46}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}; POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt56_penalty_drop_params

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"frequency_penalty":0.5,"presence_penalty":0.5}'
  2. Response:
    {"id": "chatcmpl-df0a3ab2-dec8-48f5-b7da-4311c833584b", "model": "gpt-5.6-sol-drop", "finish_reason": "stop", "role": "assistant", "content": "Hello there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 33, "total_tokens": 45}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt6_temperature

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2}'
  2. Response:
    {"error": {"message": "litellm.BadRequestError: BedrockException - {\"message\":\"This model doesn't support the temperature field. Remove temperature and try again.\"}\n\nLiteLLM: model group 'gpt-6-sol' failed with the error above. No fallback was attempted.", "type": "invalid_request_error", "param": null, "code": "400"}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {"temperature": 0.2}}

chat_gpt6_temperature_drop_params

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9}'
  2. Response:
    {"error": {"message": "litellm.BadRequestError: BedrockException - {\"message\":\"This model doesn't support the temperature field. Remove temperature and try again.\"}\n\nLiteLLM: model group 'gpt-6-sol-drop' failed with the error above. No fallback was attempted.", "type": "invalid_request_error", "param": null, "code": "400"}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {"temperature": 0.2, "topP": 0.9}}

chat_gpt6_penalty_drop_params

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"frequency_penalty":0.5,"presence_penalty":0.5}'
  2. Response:
    {"id": "chatcmpl-5d253052-f6cd-4896-8e30-5c38c3d7433e", "model": "gpt-6-sol-drop", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 9, "total_tokens": 21, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt56_low_temperature

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"reasoning_effort":"low"}'
  2. Response:
    {"error": {"message": "litellm.UnsupportedParamsError: bedrock does not support parameters: ['temperature'], for model=global.openai.gpt-5.6-sol. To drop these, set `litellm.drop_params=True` or for proxy:\n\n`litellm_settings:\n drop_params: true`\n. \n If you want to use these params dynamically send allowed_openai_params=['temperature'] in your request.\n\nLiteLLM: model group 'gpt-5.6-sol' fai
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt56_none_temperature

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9,"reasoning_effort":"none"}'
  2. Response:
    {"error": {"message": "litellm.UnsupportedParamsError: bedrock does not support parameters: ['temperature', 'top_p'], for model=global.openai.gpt-5.6-sol. To drop these, set `litellm.drop_params=True` or for proxy:\n\n`litellm_settings:\n drop_params: true`\n. \n If you want to use these params dynamically send allowed_openai_params=['temperature', 'top_p'] in your request.\n\nLiteLLM: model group
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt56_none_penalties_logprobs

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"frequency_penalty":0.5,"presence_penalty":0.5,"logprobs":true,"top_logprobs":2,"reasoning_effort":"none"}'
  2. Response:
    {"error": {"message": "litellm.UnsupportedParamsError: bedrock does not support parameters: ['presence_penalty', 'frequency_penalty', 'logprobs', 'top_logprobs'], for model=global.openai.gpt-5.6-sol. To drop these, set `litellm.drop_params=True` or for proxy:\n\n`litellm_settings:\n drop_params: true`\n. \n If you want to use these params dynamically send allowed_openai_params=['presence_penalty',
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt6_none_temperature

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"reasoning_effort":"none"}'
  2. Response:
    {"error": {"message": "litellm.BadRequestError: BedrockException - {\"message\":\"This model doesn't support the temperature field. Remove temperature and try again.\"}\n\nLiteLLM: model group 'gpt-6-sol' failed with the error above. No fallback was attempted.", "type": "invalid_request_error", "param": null, "code": "400"}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {"temperature": 0.2}, "additionalModelRequestFields": {"reasoning": {"effort": "none"}}}

chat_gpt56_none_temperature_stream

  1. curl -sS -N -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"reasoning_effort":"none","stream":true,"stream_options":{"include_usage":true}}'
  2. Response:
    {"error":{"message":"litellm.UnsupportedParamsError: bedrock does not support parameters: ['temperature'], for model=global.openai.gpt-5.6-sol. To drop these, set `litellm.drop_params=True` or for proxy:\n\n`litellm_settings:\n drop_params: true`\n. \n If you want to use these params dynamically send allowed_openai_params=['temperature'] in your request.\n\nLiteLLM: model group 'gpt-5.6-sol' faile
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt56_stream

  1. curl -sS -N -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"stream":true,"stream_options":{"include_usage":true}}'
  2. Response:
    data: {"id":"chatcmpl-458733e3-26f5-4d25-913c-6fc8165f8660","created":1790906788,"model":"gpt-5.6-sol","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"reasoning_content":"","thinking_blocks":[{"type":"redacted_thinking","data":"cnNuX215VUlMTXU5R3ZIVjVETm02RVdZdVFJRmVMbjFyZkdabVZuZjZ...
    data: {"id":"chatcmpl-458733e3-26f5-4d25-913c-6fc8165f8660","created":1790906788,"model":"gpt-5.6-sol","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hi"}}],"provider_specific_fields":{}}
    ... (9 data: lines total)
    data: {"id":"chatcmpl-458733e3-26f5-4d25-913c-6fc8165f8660","created":1790906788,"model":"gpt-5.6-sol","object":"chat.completion.chunk","choices":[{"finish_reason":"stop","index":0,"delta":{}}],"provider_specific_fields":{}}
    data: {"id":"chatcmpl-458733e3-26f5-4d25-913c-6fc8165f8660","created":1790906788,"model":"gpt-5.6-sol","object":"chat.completion.chunk","choices":[{"index":0,"delta":{}}],"usage":{"completion_tokens":37,"prompt_tokens":12,"total_tokens":49,"completion_tokens_details":{},"prompt_tokens_details":{"cac...
    data: [DONE]
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse-stream with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt56_image_fixture_png

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":[{"type":"text","text":"'"$Q_COLORS"'"},{"type":"image_url","image_url":{"url":"'"$FIX"'"}}]}]}'
  2. Response:
    {"id": "chatcmpl-338ea2f5-d093-45e1-9f55-94b3ae39081e", "model": "gpt-5.6-sol", "finish_reason": "stop", "role": "assistant", "content": "Red and blue", "usage": {"prompt_tokens": 171, "completion_tokens": 7, "total_tokens": 178, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "image_blocks": 1}

messages_gpt56

  1. curl -sS -X POST http://127.0.0.1:36349/v1/messages -d '{"model":"gpt-5.6-sol","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-e0751c10-ea50-4a9c-b21d-0e7a7069fd61", "model": "gpt-5.6-sol", "stop_reason": "end_turn", "content": [{"type": "redacted_thinking"}, {"type": "text", "text": "Hi there, friend!"}], "usage": {"input_tokens": 12, "output_tokens": 38}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {"maxTokens": 200}}

responses_gpt56

  1. curl -sS -X POST http://127.0.0.1:36349/v1/responses -d '{"model":"gpt-5.6-sol","input":"Say hi in three words."}'
  2. Response:
    {"id": "resp_j5Z28JsQGY29e4A3dchwnGmkDuIjnsqqSyHHw8hKJ6rEeRDE9HARMQEiNrsFdAB63u4s9PJvlrVRDtKu4txPPSGEz8XhP4_ss2jfHraxpfAG68Ri0SmQkQaW5YQcI7vhdEX58MIdPtoDAN-y8qUQ0QOnLcIhxtHCYNn_kcJ9Q064lDEGKQdTGik8Krpat0okMwnkGk4R8yNCzPBAy49Ad0hpOrpx1FkatroGDRRLbGja7KTGsTI9XN0f4xXn6Y1jMNjyhi0-N29SNnyW2lk_clEgAhzaFAxtIU4RZYh4BMmO9kluA44YEKnTK3sciUrZ3z22fY_jLuR3FXE1wPecsQeBtgD5nngmEoBUSD-51K8dLpd0fWYCxWYoWqAqb4fmpoH
    
  3. Recorded outbound: POST /openai/v1/responses with {"model": "global.openai.gpt-5.6-sol", "messages": 0}

messages_gpt56_model_id

  1. curl -sS -X POST http://127.0.0.1:36349/v1/messages -d '{"model":"gpt-5.6-sol-model-id","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-3410bf50-ca5a-4a62-bb76-4f396db955e6", "model": "gpt-5.6-sol-model-id", "stop_reason": "end_turn", "content": [{"type": "redacted_thinking"}, {"type": "text", "text": "Hi there, friend!"}], "usage": {"input_tokens": 12, "output_tokens": 37}}
    
  3. Recorded outbound: POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/converse with {"system": false, "messages": 1, "inferenceConfig": {"maxTokens": 200}}

responses_gpt56_model_id

  1. curl -sS -X POST http://127.0.0.1:36349/v1/responses -d '{"model":"gpt-5.6-sol-model-id","input":"Say hi in three words."}'
  2. Response:
    {"id": "resp_KY7vRY3FSGcmw8WTL47Pi0rWaYWx8jq2fe2qP7AQEzwfkQa5SCddOtfhKBzjqpRFwIvRP8Khs41YkmMFEDDDqf5ekAgDZpTjplJAwZLjqoE9FEiRhZRXFQY6DMZ1_6PdIVQRn0wJrUg3EKQkU84DngzDniqW81-rtAnJdzdLszBDn9k0d4kiFgOTBJnjuSFD7p_XnbL5EkdT8lAgPTuDPPb75PZw5VcEmMzAjb9_B1ZvOKbelQ_SKJLHEZYt3GDDRrA4U59giSLTIy8rRa0r6JvcMy7zFoqDWXQmvwlO3LvGJlLFD2Ez6o8OTVRUoD1rwm2D4JWi2WlBdwaKQ5QG5Oft9Y9xPJM8dGTzwMMVarrL_X-EcNRYvUOFqmeqYw8HQfr
    
  3. Recorded outbound: POST /openai/v1/responses with {"model": "global.openai.gpt-5.6-sol", "messages": 0}

responses_gpt56_none_temperature

  1. curl -sS -X POST http://127.0.0.1:36349/v1/responses -d '{"model":"gpt-5.6-sol","input":"Say hi in three words.","temperature":0.2,"reasoning":{"effort":"none"}}'
  2. Response:
    {"error": {"message": "litellm.UnsupportedParamsError: global.openai.gpt-5.6-sol doesn't support temperature=0.2 while reasoning is active. Only temperature=1 is supported unless reasoning.effort resolves to 'none', either set explicitly on the request or declared as the model's default_reasoning_effort. To drop unsupported params set `litellm.drop_params = True`\n\nLiteLLM: model group 'gpt-5.6-s
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt56_runtime_endpoint

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-runtime-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-6d455a42-95bf-4cb6-8bb0-edc40a6b761d", "model": "gpt-5.6-sol-runtime-endpoint", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 32, "total_tokens": 44}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt56_converse_runtime_endpoint

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-converse-runtime-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-95d49e4d-9693-45ff-b943-5c023cbd6ad0", "model": "gpt-5.6-sol-converse-runtime-endpoint", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 31, "total_tokens": 43}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt56_api_base_only_dead_port

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-env-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"error": {"message": "litellm.APIConnectionError: BedrockException - Cannot connect to host REDACTED:9 ssl:<ssl.SSLContext object at 0x110b2a270> [Connect call failed ('REDACTED', 9)]\n\nLiteLLM: model group 'gpt-5.6-sol-env-endpoint' failed with the error above. No fallback was attempted.", "type": "internal_server_error", "param": null, "code": "500"}}
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_overlong_gpt_version

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-overlong-version","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"error": {"message": "litellm.BadRequestError: BedrockException - {\"message\":\"The provided model identifier is invalid.\"}\n\nLiteLLM: model group 'gpt-overlong-version' failed with the error above. No fallback was attempted.", "type": "invalid_request_error", "param": null, "code": "400"}}
    
  3. Recorded outbound: POST /model/openai.gpt-77777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt56_model_id

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-model-id","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-7e88e9db-1f25-415a-8e6c-199b230fa328", "model": "gpt-5.6-sol-model-id", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 32, "total_tokens": 44}}
    
  3. Recorded outbound: POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

messages_gpt56_model_id

  1. curl -sS -X POST http://127.0.0.1:36349/v1/messages -d '{"model":"gpt-5.6-sol-model-id","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-9b62dbe7-0d60-4f8a-9eab-1878e76785c3", "model": "gpt-5.6-sol-model-id", "stop_reason": "end_turn", "content": [{"type": "redacted_thinking"}, {"type": "text", "text": "Hi there, friend!"}], "usage": {"input_tokens": 12, "output_tokens": 31}}
    
  3. Recorded outbound: POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/converse with {"system": false, "messages": 1, "inferenceConfig": {"maxTokens": 200}}

responses_gpt56_model_id

  1. curl -sS -X POST http://127.0.0.1:36349/v1/responses -d '{"model":"gpt-5.6-sol-model-id","input":"Say hi in three words."}'
  2. Response:
    {"id": "resp_kFyF5sWzED2vQBpofNQ-Y8ogtq3mSUkG-rV7f8zg1osjV0YI4E2kKVWMdgq0JZSiXt-iHpw5vDFKXhJ3OzDDHkyQuwsAk5FOqYvYcfDP_Bj_YqczN9gLeT3JdVejIMj5I2VJKc9Vxqa3BtEA8nfcXG-woYf5sGavA9MMuxbJnHR_X-7LhyN5zLU3gX1Tzje8vcWkiQ3va9AB6rVmevGI42_P8f6NA5G0TxX0McgWEn9ZYI-wD5AN5i3xfgitcTCqXotq9cHq2NzJ8KkeXmfbXSdmgzl8gryJ7bNlLWJ4QCxrNPM0sGf6cAywOzbXt0pYssNhjYmNlbMXSi5j3yr5aE0JRTuRu2cLg6k5aQF3llqfmirYXvWrSQXrI_k90yJzkO2
    
  3. Recorded outbound: POST /openai/v1/responses with {"model": "global.openai.gpt-5.6-sol", "messages": 0}

chat_gpt56_sigv4_blank_key

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-sigv4-blank-key","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-6cf35a45-6e7a-4bf8-a8b6-6c0d19637889", "model": "gpt-5.6-sol-sigv4-blank-key", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 33, "total_tokens": 45}}
    
  3. Recorded outbound: none, the proxy answered without calling AWS

The three env_ cases ran with AWS_BEDROCK_RUNTIME_ENDPOINT set to the recorder in the proxy's environment and no aws_bedrock_runtime_endpoint in the deployment (gpt-5.6-sol-env-endpoint, whose api_base is a dead port, and the converse/ pin)

env_chat_gpt56_api_base_dead_port_env_endpoint

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-env-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-524abe81-f89b-45be-bdd8-bbb277133532", "model": "gpt-5.6-sol-env-endpoint", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 32, "total_tokens": 44}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

env_chat_gpt56_converse_prefix

  1. curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-converse","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-47709d57-a0bb-428e-96ce-c469f0261653", "model": "gpt-5.6-sol-converse", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 33, "total_tokens": 45}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

env_messages_gpt56_api_base_dead_port_env_endpoint

  1. curl -sS -X POST http://127.0.0.1:36349/v1/messages -d '{"model":"gpt-5.6-sol-env-endpoint","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-87ed328e-1465-4244-ac39-41b176ea3ff7", "model": "gpt-5.6-sol-env-endpoint", "stop_reason": "end_turn", "content": [{"type": "redacted_thinking"}, {"type": "text", "text": "Hi there, friend!"}], "usage": {"input_tokens": 12, "output_tokens": 36}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {"maxTokens": 200}}

Before fix 2 (6181d76773)

The four cases below also ran at 6181d76773, the previous tip dae29c1 merged with main d7e6154, on port 41217 with the same rig. This is what the audit caught: a model_id override reached the native endpoint and a blank api_key became an empty bearer. 578f26e fixed both (design decisions 20 and 21), and the After leg below shows them at 200 through Converse and SigV4

chat_gpt56_model_id

  1. curl -sS -X POST http://127.0.0.1:41217/v1/chat/completions -d '{"model":"gpt-5.6-sol-model-id","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"error": {"message": "litellm.BadRequestError: BedrockException - {\"error\":{\"message\":\"Unknown parameter: 'model_id'.\",\"type\":\"invalid_request_error\",\"param\":\"model_id\",\"code\":\"unknown_parameter\"}}\n\nLiteLLM: model group 'gpt-5.6-sol-model-id' failed with the error above. No fallback was attempted.", "type": "invalid_request_error", "param": null, "code": "400"}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false}

messages_gpt56_model_id

  1. curl -sS -X POST http://127.0.0.1:41217/v1/messages -d '{"model":"gpt-5.6-sol-model-id","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"type": "error", "error": {"type": "invalid_request_error", "message": "litellm.BadRequestError: BedrockException - {\"error\":{\"message\":\"Unknown parameter: 'model_id'.\",\"type\":\"invalid_request_error\",\"param\":\"model_id\",\"code\":\"unknown_parameter\"}}\n\nLiteLLM: model group 'gpt-5.6-sol-model-id' failed with the error above. No fallback was attempted."}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "max_completion_tokens": 200}

responses_gpt56_model_id

  1. curl -sS -X POST http://127.0.0.1:41217/v1/responses -d '{"model":"gpt-5.6-sol-model-id","input":"Say hi in three words."}'
  2. Response:
    {"id": "resp__bTldW2mywztlZUyRVhSH8zInbKpKYlrVia4vXoUICMMKkXfX17d1-hCMfHFjdalgmFWgY9W3LFEL7CqP1NIpmOKHuuLJRa0zyG-m4O34HJb1VpvMSDUyi0e38XaJKIXsFQfRX76cQz4lvuQCVlJJteVLGtUDCAvECceNkTXasLnyBvgh7Z3SKi9d41BYS5hOauDh4CRIeuygiARW7ig9rfK9XZ3263bc1yRNfzwMP5iVS0RIyXBf1gPBb91JhiSgC7IpnRPiDIh2U7iyURpSdq46FJtCIsh_u8judZPMn6JN8Znm8J_QtiT1RAxsaKX1_58kooUqpeJzx2rKTMmRID5TtA4J6IJpRrR-cZN2eoKeAA2bHgX2szCOBBpW61nHOm
    
  3. Recorded outbound: POST /openai/v1/responses with {"model": "global.openai.gpt-5.6-sol", "messages": 0}

chat_gpt56_sigv4_blank_key

  1. curl -sS -X POST http://127.0.0.1:41217/v1/chat/completions -d '{"model":"gpt-5.6-sol-sigv4-blank-key","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"error": {"message": "litellm.AuthenticationError: BedrockException - {\"error\":{\"message\":\"Authorization header requires 'Credential' parameter. Authorization header requires 'Signature' parameter. Authorization header requires 'SignedHeaders' parameter. (Hashed with SHA-256 and encoded with Base64) Authorization=cQ4NvdQimYUQF6TWcmXs0EbxtO23QVJZkqGmqPqAElg=\",\"type\":\"permission_denied_err
    
  3. Recorded outbound: none, the proxy answered without calling AWS

After (6daf0b7)

chat_gpt56_basic

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-4w6cdodjrc5v6y3a6pzdyl3g2bx5z5tth3e447dm4vb6invddsma", "model": "gpt-5.6-sol", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 31, "total_tokens": 43, "reasoning_tokens": 20}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false}

chat_gpt6_basic

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-ygttb235r5377p4ffedrbv3h62gofnpf3aswx4ml5fecavp74wda", "model": "gpt-6-sol", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 9, "total_tokens": 21, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-6-sol", "messages": 1, "stream": false}

chat_gpt61_basic

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6.1-sol","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-gv4ico67dtzbw53xta6ptditdxawwkbunt4cuw4q4ax7uitxms7a", "model": "gpt-6.1-sol", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 32, "total_tokens": 44, "reasoning_tokens": 21}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-6.1-sol", "messages": 1, "stream": false}

chat_gpt6_json_schema

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Name the capital of France."}],"response_format":{"type":"json_schema","json_schema":{"name":"capital","strict":true,"schema":{"type":"object","properties":{"city":{"type":"string"},"country":{"type":"string"}},"required":["city","country"],"additionalProperties":false}}}}'
  2. Response:
    {"id": "chatcmpl-th6iv6m7vmokwk35f2ryzfibbcubzopeboymbx55cteb6neky34a", "model": "gpt-6-sol", "finish_reason": "stop", "role": "assistant", "content": "{\"city\":\"Paris\",\"country\":\"France\"}", "service_tier": "default", "usage": {"prompt_tokens": 40, "completion_tokens": 29, "total_tokens": 69, "reasoning_tokens": 11}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-6-sol", "messages": 1, "response_format": "json_schema", "stream": false}

chat_gpt56_tools_reasoning_none

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"What is the weather in Paris right now?"}],"tools":'"$TOOLS"',"reasoning_effort":"none"}'
  2. Response:
    {"id": "chatcmpl-w4rh4k6tddkkbm3xqkkeifjiazkcuyh2ingapn4tgorgfaegmxqa", "model": "gpt-5.6-sol", "finish_reason": "tool_calls", "role": "assistant", "tool_calls": [{"id": "call_0", "name": "get_weather", "arguments": "{\"city\":\"Paris\"}"}], "service_tier": "default", "usage": {"prompt_tokens": 53, "completion_tokens": 18, "total_tokens": 71, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "none", "tools": ["get_weather"], "stream": false}

chat_gpt6_tools_reasoning_low

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"What is the weather in Paris right now?"}],"tools":'"$TOOLS"',"reasoning_effort":"low"}'
  2. Response:
    {"id": "chatcmpl-807582b7-eabb-4522-b1b0-8963949cda49", "model": "gpt-6-sol", "finish_reason": "tool_calls", "role": "assistant", "content": "", "tool_calls": [{"id": "call_d7846551837d5ce7881f5464afad2964", "name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}], "usage": {"prompt_tokens": 53, "completion_tokens": 18, "total_tokens": 71, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "additionalModelRequestFields": {"reasoning": {"effort": "low"}}, "toolConfig.tools": ["get_weather"]}

chat_gpt61_tools

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6.1-sol","messages":[{"role":"user","content":"What is the weather in Paris right now?"}],"tools":'"$TOOLS"'}'
  2. Response:
    {"id": "chatcmpl-c06b397a-3854-42f1-a732-143e9c1870d1", "model": "gpt-6.1-sol", "finish_reason": "tool_calls", "role": "assistant", "content": "", "tool_calls": [{"id": "call_35dc9214e0f25a77b520b25013458f63", "name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}], "usage": {"prompt_tokens": 53, "completion_tokens": 18, "total_tokens": 71, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-6.1-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "toolConfig.tools": ["get_weather"]}

chat_gpt56_guardrail

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"guardrailConfig":'"$GUARD"'}'
  2. Response:
    {"id": "chatcmpl-7bb3611e-618a-4ca5-9cd8-3196c6cbe4e6", "model": "gpt-5.6-sol", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 31, "total_tokens": 43}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}, "guardrailConfig": {"guardrailIdentifier": "49e7k3qphvzq", "guardrailVersion": "DRAFT", "trace": "enabled"}}

chat_gpt56_app_profile

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-profile","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-631902a8-c987-4625-92e5-4c828c2601b3", "model": "gpt-5.6-sol-profile", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 44, "total_tokens": 56}}
    
  3. Recorded outbound: POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt56_model_id

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-model-id","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-344ee637-8cec-4a4b-9e1f-c9a03aac6526", "model": "gpt-5.6-sol-model-id", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 33, "total_tokens": 45}}
    
  3. Recorded outbound: POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt56_sigv4_blank_key

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-sigv4-blank-key","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-n6yuc2ty3umabkt3mby577jkt2tzwa6eyzyo22cuebc2nic4oj7a", "model": "gpt-5.6-sol-sigv4-blank-key", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 32, "total_tokens": 44, "reasoning_tokens": 21}}
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt56_converse_prefix

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-converse","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-0a8000a2-1f27-453e-b3ca-26a32585cc60", "model": "gpt-5.6-sol-converse", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 35, "total_tokens": 47}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt55_basic

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.5","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-79c09488-6239-4a7b-a36a-f47012adac4c", "model": "gpt-5.5", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend.", "usage": {"prompt_tokens": 12, "completion_tokens": 9, "total_tokens": 21, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.5/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gptoss_basic

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-oss-20b","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-7d39a111-adfa-4919-a8f7-28b02c9bcf56", "model": "gpt-oss-20b", "finish_reason": "stop", "role": "assistant", "content": "Hey there, friend.", "reasoning_content": "…", "usage": {"prompt_tokens": 73, "completion_tokens": 312, "total_tokens": 385, "reasoning_tokens": 310}}
    
  3. Recorded outbound: POST /model/openai.gpt-oss-20b-1%3A0/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gptoss_native_prefix

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-oss-20b-native","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-ff0f65e5-0a84-4b16-87b2-fea5121afc23", "model": "gpt-oss-20b-native", "finish_reason": "stop", "role": "assistant", "content": "Hey there, friend.", "service_tier": "default", "reasoning_content": "…", "usage": {"prompt_tokens": 73, "completion_tokens": 301, "total_tokens": 374}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "openai.gpt-oss-20b-1:0", "messages": 1, "stream": false}

chat_grok_basic

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"grok-4.6","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-c3c0d2e3-ed7f-4c08-9450-b17e98d9943c", "model": "grok-4.6", "finish_reason": "stop", "role": "assistant", "content": "Hello there friend.", "usage": {"prompt_tokens": 24, "completion_tokens": 295, "total_tokens": 319}}
    
  3. Recorded outbound: POST /model/us.xai.grok-4.6/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_grok_native_prefix

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"grok-4.6-native","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-2adanalomqczswyyyx6qlea4e36wo2f3qwqmxvr6b4pcewk7i24q", "model": "grok-4.6-native", "finish_reason": "stop", "role": "assistant", "content": "Hello there friend.", "service_tier": "default", "usage": {"prompt_tokens": 24, "completion_tokens": 761, "total_tokens": 785, "reasoning_tokens": 750}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "us.xai.grok-4.6", "messages": 1, "stream": false}

chat_gpt56_reasoning_low

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"What is 17*23? Answer with the number only."}],"reasoning_effort":"low"}'
  2. Response:
    {"id": "chatcmpl-qucla657fpruldok4qqiwaeqohxcto3xphciclukorxftx2ox25q", "model": "gpt-5.6-sol", "finish_reason": "stop", "role": "assistant", "content": "391", "service_tier": "default", "usage": {"prompt_tokens": 19, "completion_tokens": 5, "total_tokens": 24, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "low", "stream": false}

chat_gpt56_temperature

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9}'
  2. Response:
    {"error": {"message": "litellm.UnsupportedParamsError: global.openai.gpt-5.6-sol doesn't support ['temperature', 'top_p'] while reasoning is active on Bedrock's Chat Completions endpoint. Set reasoning_effort to 'none' to send them, or set `litellm.drop_params = True` to drop them\n\nLiteLLM: model group 'gpt-5.6-sol' failed with the error above. No fallback was attempted.", "type": "invalid_reque
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt56_temperature_drop_params

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9}'
  2. Response:
    {"id": "chatcmpl-7kquegpw6s53tjcvtfnx6ut7w7y63ujwbcft4zkwqfdcfj43m4kq", "model": "gpt-5.6-sol-drop", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 34, "total_tokens": 46, "reasoning_tokens": 23}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false}

chat_gpt56_penalty_drop_params

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"frequency_penalty":0.5,"presence_penalty":0.5}'
  2. Response:
    {"id": "chatcmpl-lxodlyzssrzcn27mk5ms4id6n6zzxvnggom76lwr3f3k5ofvp2ga", "model": "gpt-5.6-sol-drop", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 31, "total_tokens": 43, "reasoning_tokens": 20}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false}

chat_gpt6_temperature

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2}'
  2. Response:
    {"error": {"message": "litellm.UnsupportedParamsError: global.openai.gpt-6-sol doesn't support ['temperature'] while reasoning is active on Bedrock's Chat Completions endpoint. Set reasoning_effort to 'none' to send them, or set `litellm.drop_params = True` to drop them\n\nLiteLLM: model group 'gpt-6-sol' failed with the error above. No fallback was attempted.", "type": "invalid_request_error", "p
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt6_temperature_drop_params

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9}'
  2. Response:
    {"id": "chatcmpl-uz6udifib26fewvajhyjjjl45fga65njpctxmcemfbdqlb6qesqa", "model": "gpt-6-sol-drop", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 9, "total_tokens": 21, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-6-sol", "messages": 1, "stream": false}

chat_gpt6_penalty_drop_params

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"frequency_penalty":0.5,"presence_penalty":0.5}'
  2. Response:
    {"id": "chatcmpl-h6s2yx5lzejzccmfmomwp27kv3e4vj6yp5ipgfydsopk4a5a2yaq", "model": "gpt-6-sol-drop", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 9, "total_tokens": 21, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-6-sol", "messages": 1, "stream": false}

chat_gpt56_low_temperature

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"reasoning_effort":"low"}'
  2. Response:
    {"error": {"message": "litellm.UnsupportedParamsError: global.openai.gpt-5.6-sol doesn't support ['temperature'] while reasoning is active on Bedrock's Chat Completions endpoint. Set reasoning_effort to 'none' to send them, or set `litellm.drop_params = True` to drop them\n\nLiteLLM: model group 'gpt-5.6-sol' failed with the error above. No fallback was attempted.", "type": "invalid_request_error"
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt56_none_temperature

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9,"reasoning_effort":"none"}'
  2. Response:
    {"id": "chatcmpl-sfdcemwfvbnbtpzjtxrvm4cs5txt5z533ysrub3se2usxanw2gaa", "model": "gpt-5.6-sol", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 9, "total_tokens": 21, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "none", "stream": false, "temperature": 0.2, "top_p": 0.9}; POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "none", "stream": false, "temperature": 0.2, "top_p": 0.9}; POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "none", "stream": false, "temperature": 0.2, "top_p": 0.9}

chat_gpt56_none_penalties_logprobs

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"frequency_penalty":0.5,"presence_penalty":0.5,"logprobs":true,"top_logprobs":2,"reasoning_effort":"none"}'
  2. Response:
    {"id": "chatcmpl-qtoshsnf5wnibrw5r2sujvmi54u7326oap7ukdvkny3qaceihetq", "model": "gpt-5.6-sol", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 9, "total_tokens": 21, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "none", "stream": false, "frequency_penalty": 0.5, "presence_penalty": 0.5}

chat_gpt6_none_temperature

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"reasoning_effort":"none"}'
  2. Response:
    {"id": "chatcmpl-4hdxsc7v5c7rpe3yyivyim6hdx5ebodessncvnxowdgdopmprlya", "model": "gpt-6-sol", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 9, "total_tokens": 21, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-6-sol", "messages": 1, "reasoning_effort": "none", "stream": false, "temperature": 0.2}

chat_gpt56_none_temperature_stream

  1. curl -sS -N -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"reasoning_effort":"none","stream":true,"stream_options":{"include_usage":true}}'
  2. Response:
    data: {"id":"chatcmpl-5mynru2puexv5uelycqg6yqahcw73gdde3bqd7pvkytspbo2dsja","object":"chat.completion.chunk","created":1790907229,"model":"gpt-5.6-sol","service_tier":"default","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]}
    data: {"id":"chatcmpl-5mynru2puexv5uelycqg6yqahcw73gdde3bqd7pvkytspbo2dsja","object":"chat.completion.chunk","created":1790907229,"model":"gpt-5.6-sol","service_tier":"default","choices":[{"index":0,"delta":{"content":"Hi"}}]}
    ... (9 data: lines total)
    data: {"id":"chatcmpl-5mynru2puexv5uelycqg6yqahcw73gdde3bqd7pvkytspbo2dsja","object":"chat.completion.chunk","created":1790907229,"model":"gpt-5.6-sol","service_tier":"default","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
    data: {"id":"chatcmpl-5mynru2puexv5uelycqg6yqahcw73gdde3bqd7pvkytspbo2dsja","created":1790907229,"model":"gpt-5.6-sol","object":"chat.completion.chunk","choices":[{"index":0,"delta":{}}],"usage":{"completion_tokens":9,"prompt_tokens":12,"total_tokens":21,"completion_tokens_details":{"accepted_predic...
    data: [DONE]
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "none", "stream": true, "stream_options": {"include_usage": true}, "temperature": 0.2}

chat_gpt56_stream

  1. curl -sS -N -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"stream":true,"stream_options":{"include_usage":true}}'
  2. Response:
    data: {"id":"chatcmpl-c7cbkullipkkoolklxhb2pidrtaocemr5q2gr5lz6rf2mlooqkva","object":"chat.completion.chunk","created":1790907232,"model":"gpt-5.6-sol","service_tier":"default","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]}
    data: {"id":"chatcmpl-c7cbkullipkkoolklxhb2pidrtaocemr5q2gr5lz6rf2mlooqkva","object":"chat.completion.chunk","created":1790907232,"model":"gpt-5.6-sol","service_tier":"default","choices":[{"index":0,"delta":{"content":"Hi"}}]}
    ... (9 data: lines total)
    data: {"id":"chatcmpl-c7cbkullipkkoolklxhb2pidrtaocemr5q2gr5lz6rf2mlooqkva","object":"chat.completion.chunk","created":1790907232,"model":"gpt-5.6-sol","service_tier":"default","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
    data: {"id":"chatcmpl-c7cbkullipkkoolklxhb2pidrtaocemr5q2gr5lz6rf2mlooqkva","created":1790907232,"model":"gpt-5.6-sol","object":"chat.completion.chunk","choices":[{"index":0,"delta":{}}],"usage":{"completion_tokens":34,"prompt_tokens":12,"total_tokens":46,"completion_tokens_details":{"accepted_predi...
    data: [DONE]
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": true, "stream_options": {"include_usage": true}}

chat_gpt56_image_fixture_png

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":[{"type":"text","text":"'"$Q_COLORS"'"},{"type":"image_url","image_url":{"url":"'"$FIX"'"}}]}]}'
  2. Response:
    {"id": "chatcmpl-ppoblp6qrquixc3w47piazqeoyk7j3r53mxb4dtkpdrumeox4sbq", "model": "gpt-5.6-sol", "finish_reason": "stop", "role": "assistant", "content": "Red and blue", "service_tier": "default", "usage": {"prompt_tokens": 171, "completion_tokens": 7, "total_tokens": 178, "reasoning_tokens": 0}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false, "image_parts": 1, "image_url_prefix": "data:image/png;base64,"}

messages_gpt56

  1. curl -sS -X POST http://127.0.0.1:20753/v1/messages -d '{"model":"gpt-5.6-sol","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-hnppdbuwl7b6jj55bmhb2x3h4bbmb2vnzpt63vaa4hsxrvefwwva", "model": "gpt-5.6-sol", "stop_reason": "end_turn", "content": [{"type": "text", "text": "Hi there, friend!"}], "usage": {"input_tokens": 12, "output_tokens": 39}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "max_completion_tokens": 200}

responses_gpt56

  1. curl -sS -X POST http://127.0.0.1:20753/v1/responses -d '{"model":"gpt-5.6-sol","input":"Say hi in three words."}'
  2. Response:
    {"id": "resp_cXwlyQWP0KVARlbXwqdaOMR1JsVLUOO7FXyv4bWURK05edEnoxVpu4rt_BNaCMtWEI-Zw8mztpOKKYazLuEvb83XvS3WtfN9DxLXbdQ65sY1JBdK7HjEDg4ws6k9WrjK9AWrAgDc9zvTE4dYP3-0xBYJRXQSFbhSOH5TrvfXF66UtXGtL9o3kfVOf6V3i-N5RCS3lh6TdrD3q2f9AoGn59CjUI_FMvsCfB0rhe9fx9IUdN83v1jtt92eFeX7IB77grhiO-WaxdP97LaA9CS-NlZx5eVJUSvOSHzwEspwUr6boCqG7Ku8vS-CLj6tfl2KB3-biYOegZM1VKD51ZqY6tb2ldQi4CMUtwUlHHogSPaNnRGYmH6LLkmXkkkPHLT_MWd
    
  3. Recorded outbound: POST /openai/v1/responses with {"model": "global.openai.gpt-5.6-sol", "messages": 0}; POST /openai/v1/responses with {"model": "global.openai.gpt-5.6-sol", "messages": 0}; POST /openai/v1/responses with {"model": "global.openai.gpt-5.6-sol", "messages": 0}

messages_gpt56_model_id

  1. curl -sS -X POST http://127.0.0.1:20753/v1/messages -d '{"model":"gpt-5.6-sol-model-id","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-246bb253-ef83-4164-8989-d54e0837e383", "model": "gpt-5.6-sol-model-id", "stop_reason": "end_turn", "content": [{"type": "redacted_thinking"}, {"type": "text", "text": "Hi there, friend!"}], "usage": {"input_tokens": 12, "output_tokens": 42}}
    
  3. Recorded outbound: POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/converse with {"system": false, "messages": 1, "inferenceConfig": {"maxTokens": 200}}

responses_gpt56_model_id

  1. curl -sS -X POST http://127.0.0.1:20753/v1/responses -d '{"model":"gpt-5.6-sol-model-id","input":"Say hi in three words."}'
  2. Response:
    {"id": "resp_q0XY_5kxSc-bk97ZBTzULqSxretpsfMh6VsEl6DuR407aPTk_0cPcy6N791ktk5-m4tTv2gEXl_mIPdFwZHYtRMQXLfNzpOJZMgiOrL8K4vQ71eZMH-lIRAslCk6HlAIiIQuH91gQqlLLQibnVDcA2F0UjUeebjaVnNlSObQ7_NGQesVbe14Ioi7v7fBZV7UK4CbCVwfxXGrEgguASxqJOMmrWBebWlNmlSDPBC_3CRTEGb_bECT8BjqFj6SycTppeStzU9aB-OGiMDC_6lGtUYhcInPrkf0AxnDH42ftS1R4vOOPAgC4ccB4JYlWV9M4OZLKC7Vg2ameGSykDn0Qkhd5opGEeAtDIJ92LEevaHF4tphJKdXpCKiyPN-LnOOejc
    
  3. Recorded outbound: POST /openai/v1/responses with {"model": "global.openai.gpt-5.6-sol", "messages": 0}

responses_gpt56_none_temperature

  1. curl -sS -X POST http://127.0.0.1:20753/v1/responses -d '{"model":"gpt-5.6-sol","input":"Say hi in three words.","temperature":0.2,"reasoning":{"effort":"none"}}'
  2. Response:
    {"error": {"message": "litellm.UnsupportedParamsError: global.openai.gpt-5.6-sol doesn't support temperature=0.2 while reasoning is active. Only temperature=1 is supported unless reasoning.effort resolves to 'none', either set explicitly on the request or declared as the model's default_reasoning_effort. To drop unsupported params set `litellm.drop_params = True`\n\nLiteLLM: model group 'gpt-5.6-s
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_gpt56_runtime_endpoint

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-runtime-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-3xir7vpb7ps5prhfyxavxix2x4twbsj2nw7vz7sv4uudw5eaeuiq", "model": "gpt-5.6-sol-runtime-endpoint", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 36, "total_tokens": 48, "reasoning_tokens": 25}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false}

chat_gpt56_converse_runtime_endpoint

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-converse-runtime-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-985fede2-943f-447b-b562-6145a112c24f", "model": "gpt-5.6-sol-converse-runtime-endpoint", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 33, "total_tokens": 45}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

chat_gpt56_api_base_only_dead_port

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-env-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"error": {"message": "litellm.ServiceUnavailableError: BedrockException - Cannot connect to host REDACTED:9 ssl:<ssl.SSLContext object at 0x11235a8f0> [Connect call failed ('REDACTED', 9)]\n\nLiteLLM: model group 'gpt-5.6-sol-env-endpoint' failed with the error above. No fallback was attempted.", "type": "internal_server_error", "param": null, "code": "503"}}
    
  3. Recorded outbound: none, the proxy answered without calling AWS

chat_overlong_gpt_version

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-overlong-version","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"error": {"message": "litellm.BadRequestError: BedrockException - {\"message\":\"The provided model identifier is invalid.\"}\n\nLiteLLM: model group 'gpt-overlong-version' failed with the error above. No fallback was attempted.", "type": "invalid_request_error", "param": null, "code": "400"}}
    
  3. Recorded outbound: POST /model/openai.gpt-77777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

The three env_ cases ran with AWS_BEDROCK_RUNTIME_ENDPOINT set to the recorder in the proxy's environment and no aws_bedrock_runtime_endpoint in the deployment (gpt-5.6-sol-env-endpoint, whose api_base is a dead port, and the converse/ pin)

env_chat_gpt56_api_base_dead_port_env_endpoint

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-env-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-n3c2jufuzydbb2dbyin4jtdx7aapujbyqqenl6gar2hx6nebm6ra", "model": "gpt-5.6-sol-env-endpoint", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "service_tier": "default", "usage": {"prompt_tokens": 12, "completion_tokens": 37, "total_tokens": 49, "reasoning_tokens": 26}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false}

env_chat_gpt56_converse_prefix

  1. curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-converse","messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-aa81ed18-4c45-4c99-ad6c-c5441e3b3f38", "model": "gpt-5.6-sol-converse", "finish_reason": "stop", "role": "assistant", "content": "Hi there, friend!", "usage": {"prompt_tokens": 12, "completion_tokens": 32, "total_tokens": 44}}
    
  3. Recorded outbound: POST /model/global.openai.gpt-5.6-sol/converse with {"system": false, "messages": 1, "inferenceConfig": {}}

env_messages_gpt56_api_base_dead_port_env_endpoint

  1. curl -sS -X POST http://127.0.0.1:20753/v1/messages -d '{"model":"gpt-5.6-sol-env-endpoint","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'
  2. Response:
    {"id": "chatcmpl-6xvvbwwnt6t4odnogohdzeg3r4b5h74nso2joh2ei4gbiamnfyba", "model": "gpt-5.6-sol-env-endpoint", "stop_reason": "end_turn", "content": [{"type": "text", "text": "Hi there, friend!"}], "usage": {"input_tokens": 12, "output_tokens": 34}}
    
  3. Recorded outbound: POST /openai/v1/chat/completions with {"model": "global.openai.gpt-5.6-sol", "messages": 1, "max_completion_tokens": 200}

/live-pr-risk cells

The 30 cells below ran on both legs right after the cases, same proxies and recorders (the last three, the /v1/responses and /v1/messages non-string efforts the graph walk of the new commit added, on a re-boot of the same config and recorder per leg), and once more on the merged tree with identical outcomes but one (a retried first attempt aside): chat_gpt56_runtime_endpoint_stream answered 503 on the After leg because AWS accepted the stream (200 at the recorder) and 12.8 s later sent its own error event (The server had an error while processing your request. Sorry about that!) before any token, which the stream path surfaces as a 503 ServiceUnavailableError (the B2 class among the Low caveats); the identical request answered 200 stream natively on the merged tree and in the earlier run at the previous tip. Status and the AWS path the recorder saw:

Cell Before c168199 After 6daf0b7
model_group_info_supported_params gpt-5.6-sol lists 13 params gpt-5.6-sol lists 31 params
chat_gpt6_guardrail_temperature 400 BadRequestError, Converse 400 UnsupportedParamsError, no AWS call
chat_gpt6drop_guardrail_temperature 400 BadRequestError, Converse 200, Converse
chat_gpt6_guardrail_none_temperature 400 BadRequestError, Converse 400 UnsupportedParamsError, no AWS call
chat_gpt56_profile_temperature 400 BadRequestError, Converse 400 BadRequestError, Converse
chat_gpt56_converse_prefix_temperature 400 UnsupportedParamsError, no AWS call 400 UnsupportedParamsError, no AWS call
cost_chat_gpt56 200, Converse 200, native
cost_messages_gpt56 200, Converse + Converse 200, native + native
cost_responses_gpt56 200, native responses 200, native responses
cost_chat_gpt6 200, Converse 200, native
responses_gptoss_native 400 BadRequestError, Converse 200, native
messages_gpt6_temperature 400 BadRequestError, Converse 400 UnsupportedParamsError, no AWS call
chat_gpt6_stream_usage 200 stream, Converse 200 stream, native
hostile_reasoning_effort_int 200, Converse 400 UnsupportedParamsError, no AWS call
hostile_reasoning_effort_list 200, Converse 400 UnsupportedParamsError, no AWS call
hostile_reasoning_effort_int_drop_params 200, Converse + Converse 200, native
hostile_reasoning_effort_list_drop_params 200, Converse 200, native
hostile_reasoning_effort_int_request_drop_params 200, Converse 200, native
hostile_reasoning_effort_empty 400 BadRequestError, Converse 400 BadRequestError, native
hostile_reasoning_effort_5kb 400 BadRequestError, Converse 400 BadRequestError, native
hostile_temperature_string 400 UnsupportedParamsError, no AWS call 400 UnsupportedParamsError, no AWS call
hostile_temperature_string_none 400 UnsupportedParamsError, no AWS call 400 BadRequestError, native
messages_gpt56_runtime_endpoint 200, Converse 200, native
responses_gpt56_runtime_endpoint 503 ServiceUnavailableError, no AWS call 503 ServiceUnavailableError, no AWS call
chat_gpt56_runtime_endpoint_stream 200 stream, Converse 503 ServiceUnavailableError, native
messages_overlong_gpt_version 400 BadRequestError, Converse 400 BadRequestError, Converse
responses_overlong_gpt_version 400 BadRequestError, Converse 400 BadRequestError, Converse
responses_hostile_reasoning_effort_int 400 BadRequestError, native responses 400 BadRequestError, native responses
responses_hostile_reasoning_effort_int_drop_params 400 BadRequestError, native responses 400 BadRequestError, native responses
messages_hostile_reasoning_effort_int 200, Converse 400 UnsupportedParamsError, no AWS call

Observations from the run

  • chat_completions/ gpt-oss usage carries no reasoning_tokens; PR causes
  • Native GPT 5.6 /v1/messages drops Converse's empty redacted_thinking block; PR causes
  • /v1/responses with reasoning.effort: none plus temperature 400s both legs; unchanged
  • gpt-5.6 reasoning_effort: low reports reasoning_tokens: 0 both legs; unchanged
  • Converse tool answers carry content: "", native omits it; unchanged
  • Non-string reasoning_effort 400s before any AWS call; PR causes
  • Under drop_params the non-string effort is dropped, 200; unchanged
  • /v1/responses non-string reasoning.effort gets AWS's 400 both legs; unchanged
  • Unreachable api_base answers 503, not 500; PR causes
  • /v1/responses ignores model_id and calls the base model; unchanged
  • Application inference profile rows log spend: 0.0; unchanged
  • AWS 500 first attempts retried once, both legs; PR leaves alone
  • AWS in-stream error before any token answers 503; unchanged
  • Just-created deployment bursts 400 on the lagging worker; unchanged

/audit round 2 of 3

Verdict: PASS. Every inventory row ran at the head and matched the outcome written for it before the run; all 68 cells below are PASS or a FAIL row whose decision this run recorded (3 rows, A20a, A20b, B2, each a behavior change named in its row and shipped by this run), and the one decision round 1 left to the author, an int or list reasoning_effort answering AWS's 400 where Converse dropped it and answered 200, is taken in 6daf0b7: litellm answers its own 400 before any AWS call unless drop_params is set, and then drops the value the way Converse did (A24 and the live rows L8 to L8c). Unverified rows: none

Hashes: the product code under every row is 6daf0b7 (the native route, the non-string reasoning_effort rule, and the tests); e518e0e changes no product code: it pins num_retries: 0 on the deployments the A25, B1, B2, and B3 cells use and moves the B1, B2, and B3 bursts onto a proxy the chaos module owns (2 workers, the same shape as B4) whose three deployments come from its config instead of /model/new; merge base c168199. Round 2 ran the whole selection at 6daf0b7 (one proxy per leg from its own worktree and virtualenv, tests/integration/_support/proxy on tests/integration/proxy_config.yaml, 2 workers, --use_prisma_db_push, head on 24778 with database lit8684r2_audit_head, base on 32065 with lit8684r2_audit_base, one Redis on 30859, the scripted OpenAI upstream on 24827 for the unrelated deployments, and tests/integration/_support/bedrock_runtime_peer.py as the Bedrock runtime peer each test module owns; deployments are created per test through /model/new): python tests/integration/run.py providers <the six files> --seed 4106601 --order-seed 0, base audit_r2_base1 (32 failed, 23 passed), head audit_r2_head1 (4 failed, 51 passed in 309.7 s) and audit_r2_head2 (4 failed, 51 passed in 280.23 s), identical selections, no skips, no retries. The four head reds were root-caused to the audit's own tests and the rig, not the product: A25 [429], A25 [500], and B2 assert one wire attempt, but the checked-in config leaves the router's default two retries on with cooldowns disabled, so the scripted 429 and 500 and the killed peer were retried twice (three wire hits, a connect error in place of the disconnect text), where round 1 had passed them with retries set to 0 on the rig out of band (A25 [401] was never retried on either leg); B4, and A20b on the base, timed out at the owned proxy's 70 s readiness deadline in tests/integration/_support/process.py while the box carried other rigs' proxies. A first rerun of those two files with only the retries pinned off (runs superseded_changed_head1, 23 passed, and superseded_changed_head2, B2 red once: State did not converge: 8) exposed a second rig-side cause the retries had masked: /model/new lands on one worker over the keep-alive client and the sibling learns of the fresh deployment through the registry read-through in litellm/proxy/common_utils/registry_read_through.py, where each queued miss on the same key consumes one unit of the resync budget of 20 per 5 s, so B1's 36-call burst spent that budget and 4 of B2's 12 calls on the sibling answered 400 Invalid model name before any wire call (4 resync budget exhausted warnings in the head proxy log); it is pre-existing, outside this PR's diff, and ticketed separately, and a chaos burst on a deployment created seconds earlier is not what the cells measure, so e518e0e gives the three bursts config deployments every worker holds at boot. Per the audit's rerun rule the two files those cells live in (test_bedrock_runtime_chat_completions_sad_wire.py and test_bedrock_runtime_chat_completions_chaos.py, 23 cells) then re-ran whole on a fresh rig from this branch (same shape: head on 24432 with drive40775_audit_head, base on 45636 with drive40775_audit_base, Redis on 36275, upstream on 21619, one proxy up at a time, headroom confirmed first): head changed_head1 (23 passed in 217.99 s) and changed_head2 (23 passed in 210.76 s) with identical collected and passed selections, no skips, no retries, and base changed_base1 (8 failed, 15 passed, every red at a cell the inventory marks red for the merge base). Every other row keeps its 6daf0b7 result, the product code under it being the code at the head; the After column of each table names the commit its rows last ran at. Live rows ran once per leg from live/6daf0b7136/run_live.py against real AWS Bedrock in us-west-2 through the round-2 proxies, with the artifacts (proxy response, llm_provider-x-amzn-requestid, spend row) under live/6daf0b7136/<leg>/<row>.json

tests/integration/providers/test_bedrock_runtime_chat_completions_wire.py

Row Test Before c168199 After 6daf0b7 Verdict
A1 test_openai_sdk_reasoning_request_is_served_by_native_chat_completions red: Converse, litellm-made id native, upstream id kept, reasoning tokens, spend row PASS
A2 test_async_openai_sdk_stream_keeps_the_upstream_id_and_usage red: Converse ids native, upstream id on every chunk, usage last PASS
A3 test_temperature_is_forwarded_natively_when_reasoning_is_off red: 400 UnsupportedParamsError 200 native, temperature forwarded PASS
A4 test_temperature_while_reasoning_is_refused_before_any_wire_request red: 400 with Converse's text 400 "while reasoning is active", no wire call PASS
A5 test_drop_params_deployment_drops_temperature_while_reasoning red: 200 Converse 200 native, temperature dropped PASS
A6 test_guardrail_config_keeps_converse Converse, guardrailConfig in body same PASS
A7 test_converse_prefix_pins_the_model_to_converse Converse Converse PASS
A8 test_application_inference_profile_arn_keeps_converse Converse, ARN target same PASS
A8b test_model_id_application_inference_profile_keeps_converse_at_the_profile_url Converse at the ARN URL same (AWS 400 at dae29c1, fixed in 578f26e) PASS
A9 test_stop_sequences_keep_converse Converse stopSequences same PASS
A10 test_json_object_response_format_keeps_converse Converse Converse PASS
A11 test_json_schema_response_format_is_forwarded_natively red: Converse native, response_format verbatim PASS
A12 test_tools_while_reasoning_keep_converse Converse toolConfig same PASS
A13 test_tools_with_reasoning_off_are_forwarded_natively red: Converse native, tools forwarded PASS
A13b test_empty_tools_list_while_reasoning_stays_native red: Converse native PASS
A14 test_chat_completions_prefix_splits_gpt_oss_reasoning_tag red: no chat_completions/ prefix reasoning_content split from content PASS
A15 test_chat_completions_prefix_splits_gpt_oss_reasoning_tag_across_stream_deltas red: no prefix deltas join to reasoning and answer PASS
A16 test_region_path_model_is_served_natively_without_the_region red: Converse native, region stripped from the body model PASS
A17 test_sigv4_deployment_signs_the_native_request red: Converse native, AWS4-HMAC-SHA256 PASS
A17b test_blank_api_key_on_a_sigv4_deployment_is_signed_not_sent_as_an_empty_bearer red: Converse, signed native, signed, no empty bearer (AWS 403 at dae29c1, fixed in 578f26e) PASS
A18 test_runtime_endpoint_without_api_base_is_used_natively red: Converse at the wire native at the wire PASS
A18b test_runtime_endpoint_wins_over_an_unrelated_api_base red: Converse at the wire native at the wire, nothing at api_base (fixed in dae29c1) PASS
A19 test_api_base_already_naming_the_native_path_is_not_doubled[/openai/v1] red: 404, Converse path appended /openai/v1/chat/completions once PASS
A19 test_api_base_already_naming_the_native_path_is_not_doubled[/openai/v1/chat/completions] red: 404, Converse path appended same PASS

tests/integration/providers/test_bedrock_runtime_chat_completions_sad_wire.py

Row Test Before c168199 After e518e0e Verdict
A20a test_remote_image_url_on_the_shared_proxy_is_rejected_before_any_fetch 500 APIConnectionError, blocked address 400 ImageFetchError, failure spend row, no wire call FAIL, shipped by this run (Backward incompatible, Converse keeps 500, follow-up filed)
A20b test_allowlisted_remote_image_is_inlined_for_the_native_route red: Converse bytes; missing image 500 inlined data:image/png;base64, one GET /image.png; missing image 400 FAIL, shipped by this run (as A20a)
A21 test_response_cache_twin_serves_the_second_request_without_a_second_wire_call one wire call, cache-hit row same PASS
A22 test_model_group_info_lists_the_native_supported_params red: Converse list, 13 params reasoning_effort and logprobs listed, n absent PASS (intended; over-claimed params a Low caveat)
A23 test_thirty_thousand_digit_version_is_classified_quickly_and_served_by_converse 200 Converse, classifier and liveliness fast same (ValueError at 708f5a5, fixed in dae29c1) PASS
A23u test_bad_key_on_the_long_version_model_is_refused_before_any_route 401 fast, no wire call, failure spend row same PASS
A24 test_invalid_reasoning_effort_reaches_the_peer_and_its_400_reaches_the_caller[empty] 400 400, forwarded verbatim PASS
A24 test_invalid_reasoning_effort_reaches_the_peer_and_its_400_reaches_the_caller[five_kb] 400 400, forwarded verbatim PASS
A24 test_non_string_reasoning_effort_is_refused_before_any_wire_request[int] red: 200, Converse drops the int 400 UnsupportedParamsError before any wire call, names drop_params PASS (the author's decision at the read: 400 unless drop_params)
A24 test_non_string_reasoning_effort_is_refused_before_any_wire_request[list] red: 200, Converse drops the list 400 UnsupportedParamsError before any wire call, names drop_params PASS (as the int)
A24 test_drop_params_deployment_drops_a_non_string_reasoning_effort[int] red: 200, Converse 200 native, reasoning_effort absent from the wire body PASS
A24 test_drop_params_deployment_drops_a_non_string_reasoning_effort[list] red: 200, Converse 200 native, reasoning_effort absent from the wire body PASS
A24 test_duplicated_reasoning_effort_key_lets_the_last_value_win 200, last value wins same PASS
A24 test_string_temperature_is_refused_before_any_wire_request 400, no wire call 400, no wire call PASS
A25 test_peer_error_status_reaches_the_caller_and_unrelated_deployments_keep_serving[401] 401 preserved, unrelated deployment 200 same PASS
A25 test_peer_error_status_reaches_the_caller_and_unrelated_deployments_keep_serving[429] 429 preserved, unrelated 200 same PASS (retries pinned off on the cell's deployment in e518e0e, see the hashes note)
A25 test_peer_error_status_reaches_the_caller_and_unrelated_deployments_keep_serving[500] 503 (pre-existing mapping), unrelated 200 same PASS (retries pinned off on the cell's deployment in e518e0e, see the hashes note)
A26 test_absent_reasoning_effort_is_forwarded_as_absent_on_every_repeat[null] 200 x3, 3 wire hits, 3 spend rows same PASS
A26 test_absent_reasoning_effort_is_forwarded_as_absent_on_every_repeat[missing] 200 x3, 3 wire hits, 3 spend rows same PASS

tests/integration/providers/test_bedrock_gpt_responses_native_wire.py and tests/integration/providers/test_bedrock_passthrough_stream_wire.py (existing)

Row Test Before c168199 After 6daf0b7 Verdict
A27 test_openai_sdk_responses_request_is_served_by_the_native_responses_route native /openai/v1/responses same PASS
A27 test_async_openai_sdk_responses_stream_is_served_by_the_native_responses_route native /openai/v1/responses same PASS
A28 test_bedrock_passthrough_converse_stream_response_carries_event_stream_content_type event-stream content type same PASS

tests/integration/providers/test_bedrock_runtime_chat_completions_chaos.py

Row Test Before c168199 After e518e0e Verdict
B1 test_burst_across_every_endpoint_lands_each_response_id_once 36 of 36 200 via Converse, each id once 36 of 36 200 native, each id once PASS (deployment moved into the module's owned proxy config in e518e0e, see the hashes note)
B2 test_peer_killed_mid_burst_fails_only_the_held_calls_and_a_restarted_peer_serves_again red: held non-stream calls 500, stream 503 held calls 503, every success once, recovery 200 FAIL, shipped by this run (Backward incompatible, Converse's stream class); retries pinned off and the deployment moved into the module's owned proxy config in e518e0e, see the hashes note
B3 test_slow_peer_streams_are_forwarded_once_and_terminated 10 forwarded once, [DONE] same PASS (deployment moved into the module's owned proxy config in e518e0e, see the hashes note)
B4 test_worker_sigkill_mid_burst_leaves_the_sibling_serving survivor serves, held calls complete same PASS (round 2's three runs at 6daf0b7 failed this cell at the owned proxy's 70 s readiness deadline with the box carrying other rigs; the rerun below with headroom confirmed came up in time on every run; at a109787, the merge of main, the cell passed every assertion and failed the owned proxy's teardown with Owned proxy required forced cleanup: uvicorn had replaced the killed worker moments before teardown and the replacement was still starting, so it could not honor SIGTERM inside the harness's 30 s, and 6befa50 waits for the replacement worker's startup to complete before leaving the proxy; chaos file reruns at 6befa50: merged_head4 4 passed in 133 s and merged_head5 4 passed in 120 s, no skips, no retries, against a head proxy booted from 6befa50 with two uvicorn workers)

tests/integration/messages_endpoint/providers/bedrock/test_bedrock_messages_gpt_chat_completions_wire.py

Row Test Before c168199 After 6daf0b7 Verdict
C1 test_anthropic_sdk_thinking_budget_reaches_native_chat_completions_as_reasoning_effort red: Converse reasoning.effort medium native reasoning_effort: medium, no thinking PASS
C2 test_anthropic_sdk_stream_with_thinking_budget_is_served_by_native_chat_completions red: converse-stream native, stream consumed PASS
C3 test_raw_thinking_summary_reaches_native_chat_completions_as_the_plain_effort red: Converse, plain effort native reasoning_effort: medium, no summary PASS
C4 test_async_anthropic_sdk_disabled_thinking_reaches_native_chat_completions_as_effort_none red: Converse effort none native reasoning_effort: none PASS
C5 test_identical_messages_requests_reach_the_peer_once_and_log_a_cache_hit_row one wire call, cache-hit row same PASS

Live rows, real AWS Bedrock us-west-2 (never committed), base run 28839dc4 at 02:58Z and head run 06c9bfb2 at 03:11Z on 2026-10-02

Row Request Before c168199 After 6daf0b7 Verdict
L1 chat, reasoning_effort: high 200 Converse 200 native, reasoning_tokens 21 PASS
L2 chat stream, include_usage 200 Converse 200 native, usage in the last chunk PASS
L3 /v1/responses 200 native responses 200 native responses PASS
L4 /v1/messages, thinking budget 2048 200 Converse 200 native PASS
L5 application inference profile ARN 200 Converse at the ARN URL, spend 0.0 same PASS
L6 guardrailConfig 200 Converse 200 Converse PASS
L7 reasoning_effort: none + temperature 400 UnsupportedParamsError, no AWS call 200 native PASS
L8 reasoning_effort: 3 200 Converse, int dropped 400 UnsupportedParamsError, no AWS call PASS (A24's decision)
L8b reasoning_effort: 3, drop_params deployment 200 Converse, int dropped 200 native, effort absent from the wire PASS
L8c reasoning_effort: ["high"], drop_params in the body 200 Converse, list dropped 200 native, effort absent from the wire PASS
L9a chat, model_id ARN 200 Converse at the ARN URL same PASS
L9b /v1/messages, model_id ARN 200 Converse at the ARN URL same PASS
L10 SigV4 aws_profile_name, api_key: "" 200 Converse 200 native, signed PASS

Pass-through reachability: the diff's byte and SSE code (ReasoningTagSplitter, BedrockRuntimeChatCompletionsStreamingHandler) is imported only by litellm/llms/bedrock/chat/chat_completions/transformation.py and its unit test, so A28 is the one pass-through route whose bytes it could touch. The A15 cell sends stream_options from the client because a worker that has not yet loaded a deployment created seconds earlier omits include_usage (pre-existing, Low caveat). B1 matches a non-stream /v1/responses spend row by the upstream id inside its payload because that row can carry the pre-encryption id (pre-existing, Low caveat)

Kept from the 2026-09-26 run

The ten non-Bedrock image cases below ran on 2026-09-26 at the merge base 1f77fa6 and the tip 4900a1a, on the same two-worker rig in us-east-1 with one recorder per provider (Anthropic, Gemini, OpenAI, Vertex AI at vertex_location: global, and Bedrock invoke); the only change to the shared image fetch since then (3e3b9cd) hoists the URL comprehension into a helper, so the cases still describe the tip. Their deployments were:

  - model_name: claude-anthropic
    litellm_params:
      model: anthropic/claude-haiku-4-5
      api_base: os.environ/ANTHROPIC_BASE
  - model_name: claude-invoke
    litellm_params:
      model: bedrock/invoke/us.anthropic.claude-haiku-4-5-20251001-v1:0
      api_base: os.environ/BEDROCK_BASE
  - model_name: gemini-flash
    litellm_params:
      model: gemini/gemini-3.8-flash
      api_base: os.environ/GEMINI_BASE
  - model_name: openai-gpt
    litellm_params:
      model: openai/gpt-5.6
      api_base: os.environ/OPENAI_BASE
  - model_name: vertex-gemini
    litellm_params:
      model: vertex_ai/gemini-3.8-flash
      vertex_project: os.environ/VERTEXAI_PROJECT
      vertex_location: global
      api_base: os.environ/VERTEX_BASE
  - model_name: vertex-claude
    litellm_params:
      model: vertex_ai/claude-haiku-4-5@20251001
      vertex_project: os.environ/VERTEXAI_PROJECT
      vertex_location: global
      api_base: os.environ/VERTEX_BASE

The fixtures live under $ASSETS (https://github.com/mateo-berri/qa-screenshots/releases/download/assets-2; $ASSETS_HTTP is the same path over http://), which GitHub serves as application/octet-stream behind a 302, the case the type inference is for. pr40775-fixture.png and pr40775-pngblob are the same 96x64 PNG, red on the left and blue on the right, one with the .png extension and one without; pr40775-notimage is 160 bytes of text; pr40775-doc.pdf is a one-page PDF reading "The secret word is PELICAN". $Q_COLORS is as above and $Q_PDF is "What is the secret word in this document? Answer with the word only."

Before (1f77fa6)

chat_anthropic_http_fixture_png
  1. curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"claude-anthropic","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS_HTTP/pr40775-fixture.png"}}]}],"max_tokens":200}'
  2. HTTP 400 in 0.6 s:
    {"error":{"message":"litellm.BadRequestError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"invalid_request_error\",\"message\":\"messages.0.content.1.image.source.base64.media_type: Input should be 'image/jpeg', 'image/png', 'image/gif' or 'image/webp'\"},\"request_id\":\"req_011CfSgRX89wVRp3pf6CW92W\"}\n\nLiteLLM: model group 'claude-anthropic' failed with the error above. No fallback was attempted.","type":"invalid_request_error","param":null,"code":"400"}}
    
  3. Recorded outbound: anthropic recorder: POST /v1/messages with max_tokens 200; media application/octet-stream
chat_anthropic_http_notimage
  1. curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"claude-anthropic","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS_HTTP/pr40775-notimage"}}]}],"max_tokens":200}'
  2. HTTP 400 in 0.8 s:
    {"error":{"message":"litellm.BadRequestError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"invalid_request_error\",\"message\":\"messages.0.content.1.image.source.base64.media_type: Input should be 'image/jpeg', 'image/png', 'image/gif' or 'image/webp'\"},\"request_id\":\"req_011CfSgRbUbPiMhmDquaNvV8\"}\n\nLiteLLM: model group 'claude-anthropic' failed with the error above. No fallback was attempted.","type":"invalid_request_error","param":null,"code":"400"}}
    
  3. Recorded outbound: anthropic recorder: POST /v1/messages with max_tokens 200; media application/octet-stream
chat_anthropic_http_pngblob
  1. curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"claude-anthropic","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS_HTTP/pr40775-pngblob"}}]}],"max_tokens":200}'
  2. HTTP 400 in 0.5 s:
    {"error":{"message":"litellm.BadRequestError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"invalid_request_error\",\"message\":\"messages.0.content.1.image.source.base64.media_type: Input should be 'image/jpeg', 'image/png', 'image/gif' or 'image/webp'\"},\"request_id\":\"req_011CfSgRdvSUnthdhTxrCpQd\"}\n\nLiteLLM: model group 'claude-anthropic' failed with the error above. No fallback was attempted.","type":"invalid_request_error","param":null,"code":"400"}}
    
  3. Recorded outbound: anthropic recorder: POST /v1/messages with max_tokens 200; media application/octet-stream
chat_gemini_file_pdf
  1. curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"gemini-flash","messages":[{"role":"user","content":[{"type":"text","text":"$Q_PDF"},{"type":"file","file":{"file_id":"$ASSETS/pr40775-doc.pdf"}}]}],"max_tokens":200}'
  2. HTTP 200 in 1.3 s:
    {"id":"pSq4auaFD4P3jMcPhoDIiQc","model":"gemini-flash","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"PELICAN","role":"assistant","provider_specific_fields":{"thought_signatures":["…"]}},"provider_specific_fields":{"native_finish_reason":"STOP"}}],"usage":{"completion_tokens":77,"prompt_tokens":560,"total_tokens":637},"service_tier":"default"}
    
  3. Recorded outbound: gemini recorder: POST /v1beta/models/gemini-3.8-flash:generateContent with generationConfig {"max_output_tokens":200,"temperature":1.0}; media application/pdf
chat_gemini_fixture_png
  1. curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"gemini-flash","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-fixture.png"}}]}],"max_tokens":200}'
  2. HTTP 200 in 1.7 s:
    {"id":"piq4apDfH_XQjrEP3b6K2Q0","model":"gemini-flash","object":"chat.completion","choices":[{"finish_reason":"length","index":0,"message":{"content":"`x` (0x7","role":"assistant","provider_specific_fields":{"thought_signatures":["…"]}},"provider_specific_fields":{"native_finish_reason":"MAX_TOKENS"}}],"usage":{"completion_tokens":196,"prompt_tokens":0,"total_tokens":224},"service_tier":"default"}
    
  3. Recorded outbound: gemini recorder: POST /v1beta/models/gemini-3.8-flash:generateContent with generationConfig {"max_output_tokens":200,"temperature":1.0}; media application/octet-stream
chat_gemini_pngblob
  1. curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"gemini-flash","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-pngblob"}}]}],"max_tokens":200}'
  2. HTTP 200 in 1.6 s:
    {"id":"qCq4apOFFsrP_uMPtsCz4AE","model":"gemini-flash","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Black and white","role":"assistant","provider_specific_fields":{"thought_signatures":["…"]}},"provider_specific_fields":{"native_finish_reason":"STOP"}}],"usage":{"completion_tokens":192,"prompt_tokens":0,"total_tokens":220},"service_tier":"default"}
    
  3. Recorded outbound: gemini recorder: POST /v1beta/models/gemini-3.8-flash:generateContent with generationConfig {"max_output_tokens":200,"temperature":1.0}; media application/octet-stream
chat_invoke_claude_fixture_png
  1. curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"claude-invoke","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-fixture.png"}}]}],"max_tokens":200}'
  2. HTTP 400 in 0.6 s:
    {"error":{"message":"litellm.BadRequestError: BedrockException - {\"message\":\"messages.0.content.1.image.source.base64.media_type: Input should be 'image/jpeg', 'image/png', 'image/gif' or 'image/webp'\"}\n\nLiteLLM: model group 'claude-invoke' failed with the error above. No fallback was attempted.","type":"invalid_request_error","param":null,"code":"400"}}
    
  3. Recorded outbound: bedrock recorder: POST /model/us.anthropic.claude-haiku-4-5-20251001-v1:0/invoke with max_tokens 200; media application/octet-stream
chat_openai_file_pdf
  1. curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"openai-gpt","messages":[{"role":"user","content":[{"type":"text","text":"$Q_PDF"},{"type":"file","file":{"file_id":"$ASSETS/pr40775-doc.pdf"}}]}],"max_completion_tokens":200}'
  2. HTTP 400 in 0.2 s:
    {"error":{"message":"litellm.BadRequestError: OpenAIException - Invalid file data: 'messages[0].content[1].file.file_data'. Expected a base64-encoded data URL with an application/pdf MIME type (e.g. 'data:application/pdf;base64,SGVsbG8sIFdvcmxkIQ=='), but got unsupported MIME type 'application/octet-stream'.\n\nLiteLLM: model group 'openai-gpt' failed with the error above. No fallback was attempted.","type":"invalid_request_error","param":"messages[0].content[1].file.file_data","code":"400"}}
    
  3. Recorded outbound: openai recorder: POST /v1/chat/completions with max_completion_tokens 200; media data:application/octet-stream;base64,…
chat_vertex_claude_fixture_png
  1. curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"vertex-claude","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-fixture.png"}}]}],"max_tokens":200}'
  2. HTTP 400 in 3.5 s:
    {"error":{"message":"litellm.BadRequestError: Vertex_aiException BadRequestError - {\"type\":\"error\",\"error\":{\"type\":\"invalid_request_error\",\"message\":\"messages.0.content.1.image.source.base64.media_type: Input should be 'image/jpeg', 'image/png', 'image/gif' or 'image/webp'\"},\"request_id\":\"req_vrtx_011CfSiRWuzkZqHKFmgXgB7z\"}\n\nLiteLLM: model group 'vertex-claude' failed with the error above. No fallback was attempted.","type":"invalid_request_error","param":null,"code":"400"}}
    
  3. Recorded outbound: vertex recorder: POST /v1/projects/$VERTEXAI_PROJECT/locations/global/publishers/anthropic/models/claude-haiku-4-5@20251001:rawPredict with max_tokens 200; media application/octet-stream
chat_vertex_gemini_pngblob
  1. curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"vertex-gemini","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-pngblob"}}]}],"max_tokens":400}'
  2. HTTP 400 in 0.6 s:
    {"error":{"message":"litellm.BadRequestError: Vertex_aiException BadRequestError - {\n  \"error\": {\n    \"code\": 400,\n    \"message\": \"Unable to submit request because it has a mimeType parameter with value application/octet-stream, which is not supported. Update the mimeType and try again. Learn more: https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/gemini\",\n    \"status\": \"INVALID_ARGUMENT\"\n  }\n}\n\n\nLiteLLM: model group 'vertex-gemini' failed with the error above. No fallback was attempted.","type":"invalid_request_error","param":null,"code":"400"}}
    
  3. Recorded outbound: vertex recorder: POST /v1/projects/$VERTEXAI_PROJECT/locations/global/publishers/google/models/gemini-3.8-flash:generateContent with generationConfig {"max_output_tokens":400,"temperature":1.0}; media application/octet-stream

After (4900a1a)

chat_anthropic_http_fixture_png
  1. curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"claude-anthropic","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS_HTTP/pr40775-fixture.png"}}]}],"max_tokens":200}'
  2. HTTP 200 in 1.8 s:
    {"id":"chatcmpl-1238f8cd-f386-4f30-8f1f-43426860ff35","model":"claude-anthropic","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Red and blue.","role":"assistant"}}],"usage":{"completion_tokens":7,"prompt_tokens":38,"total_tokens":45,"inference_geo":"not_available","service_tier":"standard"}}
    
  3. Recorded outbound: anthropic recorder: POST /v1/messages with max_tokens 200; media image/png
chat_anthropic_http_notimage
  1. curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"claude-anthropic","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS_HTTP/pr40775-notimage"}}]}],"max_tokens":200}'
  2. HTTP 400 in 0.4 s:
    {"error":{"message":"litellm.BadRequestError: Error: Unable to determine image content type from the server's headers, the URL, or the image bytes. url=https://github.com/mateo-berri/qa-screenshots/releases/download/assets-2/pr40775-notimage\n\nLiteLLM: model group 'claude-anthropic' failed with the error above. No fallback was attempted.","type":"invalid_request_error","param":null,"code":"400"}}
    
  3. Recorded outbound: none, the request never left the proxy
chat_anthropic_http_pngblob
  1. curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"claude-anthropic","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS_HTTP/pr40775-pngblob"}}]}],"max_tokens":200}'
  2. HTTP 200 in 1.0 s:
    {"id":"chatcmpl-694e40f5-7103-4410-93fa-7ef43ab57e7f","model":"claude-anthropic","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Red and blue.","role":"assistant"}}],"usage":{"completion_tokens":7,"prompt_tokens":38,"total_tokens":45,"inference_geo":"not_available","service_tier":"standard"}}
    
  3. Recorded outbound: anthropic recorder: POST /v1/messages with max_tokens 200; media image/png
chat_gemini_file_pdf
  1. curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"gemini-flash","messages":[{"role":"user","content":[{"type":"text","text":"$Q_PDF"},{"type":"file","file":{"file_id":"$ASSETS/pr40775-doc.pdf"}}]}],"max_tokens":200}'
  2. HTTP 200 in 2.4 s:
    {"id":"0zi4avPDPK_f-sAPsobQmAY","model":"gemini-flash","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"PELICAN","role":"assistant","provider_specific_fields":{"thought_signatures":["…"]}},"provider_specific_fields":{"native_finish_reason":"STOP"}}],"usage":{"completion_tokens":75,"prompt_tokens":560,"total_tokens":635},"service_tier":"default"}
    
  3. Recorded outbound: gemini recorder: POST /v1beta/models/gemini-3.8-flash:generateContent with generationConfig {"max_output_tokens":200,"temperature":1.0}; media application/pdf
chat_gemini_fixture_png
  1. curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"gemini-flash","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-fixture.png"}}]}],"max_tokens":200}'
  2. HTTP 200 in 1.4 s:
    {"id":"1ji4avnKCqLTjMcPzKeH2AY","model":"gemini-flash","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Red and blue","role":"assistant","provider_specific_fields":{"thought_signatures":["…"]}},"provider_specific_fields":{"native_finish_reason":"STOP"}}],"usage":{"completion_tokens":60,"prompt_tokens":1097,"total_tokens":1157},"service_tier":"default"}
    
  3. Recorded outbound: gemini recorder: POST /v1beta/models/gemini-3.8-flash:generateContent with generationConfig {"max_output_tokens":200,"temperature":1.0}; media image/png
chat_gemini_pngblob
  1. curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"gemini-flash","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-pngblob"}}]}],"max_tokens":200}'
  2. HTTP 200 in 1.5 s:
    {"id":"1zi4atXEK_n5-8YPusH3UA","model":"gemini-flash","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Red, blue","role":"assistant","provider_specific_fields":{"thought_signatures":["…"]}},"provider_specific_fields":{"native_finish_reason":"STOP"}}],"usage":{"completion_tokens":135,"prompt_tokens":1097,"total_tokens":1232},"service_tier":"default"}
    
  3. Recorded outbound: gemini recorder: POST /v1beta/models/gemini-3.8-flash:generateContent with generationConfig {"max_output_tokens":200,"temperature":1.0}; media image/png
chat_invoke_claude_fixture_png
  1. curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"claude-invoke","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-fixture.png"}}]}],"max_tokens":200}'
  2. HTTP 200 in 1.1 s:
    {"id":"chatcmpl-d618211c-89be-4d66-b2b3-53fd07cb596d","model":"claude-invoke","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Red and blue.","role":"assistant"}}],"usage":{"completion_tokens":7,"prompt_tokens":38,"total_tokens":45,"service_tier":"standard"}}
    
  3. Recorded outbound: bedrock recorder: POST /model/us.anthropic.claude-haiku-4-5-20251001-v1:0/invoke with max_tokens 200; media image/png
chat_openai_file_pdf
  1. curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"openai-gpt","messages":[{"role":"user","content":[{"type":"text","text":"$Q_PDF"},{"type":"file","file":{"file_id":"$ASSETS/pr40775-doc.pdf"}}]}],"max_completion_tokens":200}'
  2. HTTP 200 in 2.7 s:
    {"id":"chatcmpl-ESUENFpHWsVFNqgm7vZdg91B2M34s","model":"openai-gpt","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"PELICAN","role":"assistant","provider_specific_fields":{"refusal":null}}}],"usage":{"completion_tokens":6,"prompt_tokens":243,"total_tokens":249},"service_tier":"default"}
    
  3. Recorded outbound: openai recorder: POST /v1/chat/completions with max_completion_tokens 200; media data:application/pdf;base64,…
chat_vertex_claude_fixture_png
  1. curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"vertex-claude","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-fixture.png"}}]}],"max_tokens":200}'
  2. HTTP 200 in 4.8 s:
    {"id":"chatcmpl-6626aaf7-f6ca-41e9-8d47-f035d076eeaf","model":"vertex-claude","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Red and blue.","role":"assistant"}}],"usage":{"completion_tokens":7,"prompt_tokens":38,"total_tokens":45}}
    
  3. Recorded outbound: vertex recorder: POST /v1/projects/$VERTEXAI_PROJECT/locations/global/publishers/anthropic/models/claude-haiku-4-5@20251001:rawPredict with max_tokens 200; media image/png
chat_vertex_gemini_pngblob
  1. curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"vertex-gemini","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-pngblob"}}]}],"max_tokens":400}'
  2. HTTP 200 in 3.2 s:
    {"id":"-ji4ap3KEL-Vq8YPsPulgAI","model":"vertex-gemini","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Red, blue","role":"assistant","provider_specific_fields":{"thought_signatures":["…"]}},"provider_specific_fields":{"native_finish_reason":"STOP"}}],"usage":{"completion_tokens":72,"prompt_tokens":1096,"total_tokens":1168}}
    
  3. Recorded outbound: vertex recorder: POST /v1/projects/$VERTEXAI_PROJECT/locations/global/publishers/google/models/gemini-3.8-flash:generateContent with generationConfig {"max_output_tokens":400,"temperature":1.0}; media image/png

Type

🆕 New Feature

Design decisions

  1. This PR, opened for Grok, carries gpt-oss and GPT 5.6+ (5.6, 6, 6.1) too, since all of them ride the same endpoint, request config, and cost-map keys. One PR per model family would have reviewed the same routing three times
  2. The route is decided from the model id and its cost-map row. An unprefixed bedrock/<id> goes to runtime Chat Completions only when the id is an OpenAI GPT model at version 5.6 or newer (openai.gpt-5.6-*, openai.gpt-6-*, openai.gpt-6.1-*, read off the id, never gpt-oss) and its row lists /v1/chat/completions in supported_endpoints (the key feat(bedrock): serve the OpenAI models on bedrock-runtime's native Responses API (internal copy of #38489) #42767 already reads for the native /v1/responses route). gpt-oss and Grok rows carry the key too but stay on Converse by default; bedrock/chat_completions/<id> opts a deployment in, and bedrock/converse/<id> pins one to Converse. Application inference profile ARNs take Converse ahead of that check, as before this PR. A region path (bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0) is routing, not a different model, so the route is looked up on the id after the path and the path's region picks the endpoint and the SigV4 scope (an explicit aws_region_name still wins, as on Converse). A vendor-prefix rule (openai., xai., zai.) was raised in review, and AWS's API compatibility table does list Chat Completions on every OpenAI, xAI, and Z.AI row, but the per-model key stays for three reasons. The two companion flags differ within openai. already (GPT 5.6+ refuses tools with reasoning, gpt-oss ignores a response_format schema), so the cost map keeps per-model keys under a prefix rule too. A prefix rule would move every gpt-5.4, gpt-5.5, gpt-oss-safeguard, and GLM deployment to the native route at upgrade with none of their refused params, reasoning shape, tools, or response_format checked there. And AWS's table is per model, not per vendor (Kimi K2 Thinking and DeepSeek-R1 lack Chat Completions while their siblings have it), so a vendor key would be wrong the day it is extended past these three vendors. The 5.6 floor comes from the customer ask (feat(bedrock): route OpenAI models to bedrock-runtime's native Chat Completions API #43264 wanted GPT 5.6+ served natively with an automatic fallback and no new flag): gpt-5.4 and gpt-5.5 keep Converse even if a row of theirs gains the key, and a GPT 5.6+ id with no row, or a row without the key, keeps Converse too. The two capability flags stay, since they say what the native endpoint serves for a model (tools with reasoning, a response_format schema), not which endpoints AWS exposes for it; a daily automation syncing supported_endpoints from AWS's compatibility doc is a follow-up outside this PR
  3. A request carrying guardrailConfig goes to Converse. The alternative was mapping it onto the native endpoint's X-Amzn-Bedrock-GuardrailIdentifier / X-Amzn-Bedrock-GuardrailVersion headers, which keeps the native route but adds exactly the kind of translation this PR removes. The trade-off is that Bedrock Guardrails requests keep paying the Converse translation
  4. The same fallback covers every Converse-only body block (performanceConfig, serviceTier, requestMetadata, outputConfig), the Anthropic-style thinking block on /v1/chat/completions (only Converse forwards it, as additionalModelRequestFields, so Grok's redacted_thinking block still comes back there; /v1/messages maps thinking to reasoning_effort before the route is picked, so it goes native), and operator-owned request metadata (the proxy's Bedrock request-metadata opt-in), since the native endpoint has no field for them. Application inference profile ARNs took Converse before this PR and still do, checked ahead of the flag
  5. GPT 5.6 and GPT 6 function tools go native only with "reasoning_effort": "none", because AWS rejects tools plus reasoning on the native endpoint (checked live on 2026-09-30 on global.openai.gpt-5.6-sol and global.openai.gpt-6-sol). The cost map states the capability positively, supports_bedrock_runtime_chat_completions_tools_with_reasoning: true on gpt-oss and Grok, and a model without it sends a tools request with any other effort, or with no effort at all (AWS applies its default effort there), to Converse. Absent flag means the safer route, so a newly listed model never hits AWS's 400 by default. gpt-6.1 has no none effort on AWS, so its tools requests always take Converse. Legacy functions are gated the same way, since AWS rejects them with the same 400. Serving tools plus reasoning natively through the Responses API is deferred to LIT-8970
  6. max_tokens is sent as max_completion_tokens: every model on this surface takes it and the GPT 5.6+ family rejects max_tokens. An explicit max_completion_tokens wins when both are set. Rejected: renaming only for the GPT family and forwarding max_tokens as sent on gpt-oss and Grok, a second per-family table for a rename every model on this endpoint accepts
  7. n is dropped from the native endpoint's supported params, since AWS serves one choice there, and so are the params each family refuses outright on that endpoint, checked live: both penalties on Grok (a 503 for any non-zero value) and logit_bias on gpt-oss (a 400). On GPT 5.6+ six params are tied to reasoning: AWS takes temperature, top_p, frequency_penalty, presence_penalty, logprobs, and top_logprobs natively under reasoning_effort: "none" and 400s them at every other effort (checked live on 2026-09-30 on gpt-5.6 and gpt-6). So under none the native config forwards all six, and otherwise it answers a 400 naming them and the none way out, or drops them under drop_params. The merge base answered litellm's own bedrock does not support parameters: ['temperature', 'top_p'] 400 for GPT-5.6 at every effort (its rows carry supports_sampling_params: false, which Converse reads) and forwarded them for GPT-6, which AWS 400'd even under drop_params (its rows had no flag). The PR adds supports_sampling_params: false to the gpt-6 and gpt-6.1 rows, so the Converse fallback and a converse/ pin refuse or drop the two before the call the way they already did for GPT-5.6. gpt-6.1 has no none effort, so its six params always 400 or drop. The refused set is one table keyed by model family in the native config. Rejected: a per-model list in the cost map (17 entries for a per-family fact) and composing the OpenAI GPT-5 and xAI configs' lists (they rewrite the body through their own mapping and keep presence_penalty, which AWS's Grok 503s)
  8. gpt-oss returns its reasoning inline as a <reasoning>...</reasoning> prefix on the native endpoint. One splitter serves both the streamed and the non-streamed response, so the two split identically, and the reasoning lands in reasoning_content only. Filling thinking_blocks and provider_specific_fields.reasoningContentBlocks the way Converse did would have kept three copies of the same text
  9. A {"type": "json_schema"} response_format (a pydantic model is converted to that shape) goes native only on models flagged supports_bedrock_runtime_chat_completions_response_format: GPT 5.6+ and Grok, which enforce the schema on AWS's endpoint (checked live, both answer the schema's shape, and Grok honors it on both routes once the reasoning fits the token budget). Every {"type": "json_object"} form keeps Converse's handling on every model, the LiteLLM-specific response_schema variant included (Converse turns that one into a forced json_tool_call and answers the schema as before this PR; the schema-less one is ignored and answers free text as before this PR), because AWS's native endpoint validates by type and answers 400 'messages' must contain the word 'json' for json_object unless the prompt mentions json, which the A/B caught for the response_schema variant at an earlier tip; rejected: sending json_object native only when a message mentions json, which needs the messages plumbed into the route predicate and mirrors AWS's contract heuristically. GPT-OSS accepts response_format on the native endpoint but answers with free text, so it keeps Converse's emulation (a forced json_tool_call tool) and the structured answer it gave before this PR. {"type": "text"} is a no-op on both routes and stays native. The flag is positive so an unflagged model gets the safer route
  10. The route is decided once from the caller's raw params, honoring additional_drop_params, by one helper that both the param mapping and the dispatch call. The first version of this PR derived it again at dispatch from the mapped params, which was wrong for the GPT-OSS response_format fallback: Converse's mapping turns response_format into the json_tool_call tool and drops the field, so a dispatch reading the mapped params picked the native handler and handed it Converse-shaped params. A test pins the dispatch to the Converse URL and body for that case
  11. All three cost map keys carry the supports_ prefix. The cost map guard runs the base branch's schema generator against the PR's cost map and auto-classifies supports_* keys as booleans, so any other name fails the required check until a generator-only PR lands on main first. The prefix also keeps the keys out of the generator's hand-maintained key list. They do not leak into get_model_info, /model/info, the public model hub features, or the UI filters, which all read from a typed model info with explicit fields (verified at runtime)
  12. Grok's native endpoint refuses reasoning_effort: "none" with a 400 (it takes low, medium, high, and xhigh), where Converse never listed reasoning_effort for xai models and dropped it under drop_params, so none answered 200 with AWS's default effort. The native config drops none on Grok before the call (one table of refused effort values keyed by model family, next to the refused params table), so the 200 stays and AWS applies its default effort as before, while the other efforts, which Converse silently dropped, are now forwarded on an opted-in Grok deployment. This also keeps /v1/messages with thinking: {"type": "disabled"} working, since the Anthropic adapter maps that to reasoning_effort: "none". Rejected: routing none to Converse (a whole second code path for a value AWS's default already gives) and raising (a regression from main for a value OpenAI clients send by default). gpt-oss refuses none on both routes and GPT 5.6+ takes it on both, so neither changes
  13. additionalModelRequestFields and top_k join the Converse-only request keys, so a request carrying either keeps main's Converse handling to the byte (the 2026-09-26 merge base leg showed Converse nesting additionalModelRequestFields inside its own block and emitting no topK for these models, both of which are Converse's behavior on main, not this PR's). Rejected: letting them through to the native endpoint, which answers 200 and ignores an unknown key, or leaving top_k to the native drop-or-raise handling, which turns main's 200 into an error for a caller without drop_params. Any other unknown body key reaches the native endpoint as sent, as on every OpenAI-compatible provider
  14. QA: each leg runs from its own worktree with its own virtualenv at the named commit, two uvicorn workers, the us-west-2 deployments behind one forwarding recorder so the AWS path each case hit is read from the wire instead of inferred; the kept image cases ran the same way on 2026-09-26 in us-east-1
  15. stop joins the Converse-only request keys. Converse forwards it as stopSequences, which AWS refuses with a 400 on all three families, and that is what main does. Natively gpt-oss and Grok accept stop but apply it to the hidden reasoning stream too, so stop: ["3"] on "count from 1 to 5" answered 200 with content: "" (gpt-oss) or content: null (Grok) and finish_reason: "stop", a silent empty answer where main failed loudly, and GPT-5.6 raised UnsupportedParamsError from get_optional_params (a different error class and layer) once stop left its native list. Routing stop to Converse keeps main's exact 400 for every caller, so stop is also back out of the GPT refused table since the native config never sees it. Rejected: serving stop natively on gpt-oss and Grok (the empty answers above), or refusing it natively on every family (the error class change)
  16. Remote http(s):// image_url parts are downloaded and inlined as data: URLs before the native call, in transform_request through the shared inline_remote_media and in async_transform_request through async_inline_remote_media (the same helpers the Anthropic and Bedrock invoke transforms use, with inline_remote_image_urls so file parts are left alone). Converse downloaded remote images itself; AWS's native endpoint answers 400 Only inline image data URLs and S3 URLs are supported for a URL, so us.xai.grok-4.6 and the GPT 5.6+ rows (supports_vision: true) would have regressed from 200 to 400 on any vision request. data: and s3:// URLs pass through untouched. The sync inline_remote_media is new in image_handling.py, a sequential mirror of the async one, since the module only had the async walker. The shared fetch kept a generic Content-Type (application/octet-stream) from the image server as the data URL's type, where Converse's BedrockImageProcessor inferred the real type from the extension or the magic bytes through infer_content_type_from_url_and_content; the fetch now runs that same helper, so the native route gets data:image/png for such a server and the Anthropic and Bedrock invoke transforms that already used the fetch get the same inference
  17. aws_bedrock_project_id is not forwarded as an OpenAI-Project header on the native route. AWS documents projects and that header for Bedrock Mantle only, bedrock-runtime has a default project alone and answered 200 to a bogus OpenAI-Project value on 2026-09-26, the native Responses transform sends none, and Converse ignored the param, so a runtime deployment carrying it behaves the same on both routes (checked on 2026-09-26 with the recorder showing no OpenAI-Project header on either leg). The first revision forwarded it the way the Mantle configs do and was dropped, since it sent a header nobody chose to a host that does not read it
  18. A reasoning_effort value that is not a string (an int, a list, an object) answers litellm's own 400 before any call (UnsupportedParamsError, naming the type and the drop_params way out), and is dropped under litellm.drop_params or the request's drop_params so AWS applies its default effort, which is what Converse did on main. That is litellm's convention for an unsupported or malformed param on a provider route (the o-series temperature check, and this PR's own refused-while-reasoning branch), so the native route follows it instead of forwarding the value for AWS's 400 (what a0cef91 did) or always dropping it (which hides a malformed request behind a 200). A string AWS does not know (an empty string, a made-up effort) is still forwarded, since the accepted set is AWS's to define per family
  19. A deployment's model_id override (the per-deployment litellm_params.model_id, used to call an application inference profile ARN while the cost map still reads the base model id) joins the Converse-only request keys, so such a deployment keeps Converse at /model/<ARN>/converse as on main. The native endpoint has no field for a profile ARN: at dae29c1 the request went native with the base model in the URL and model_id inside the body, and AWS answered 400 Unknown parameter: 'model_id' (the audit's A8b row and the chat_gpt56_model_id case). Rejected: putting the ARN in the native body's model field, since AWS documents the native endpoint on model ids and nothing shows it taking an ARN there
  20. The native config's validate_environment resolves the key through bedrock_bearer_token before the inherited OpenAI-compatible one writes the header, so a SigV4 deployment with api_key: "" (what a config template or a UI form leaves behind) is signed instead of sent as Authorization: Bearer with an empty token, which AWS 403s (Authorization header requires 'Credential' parameter, the audit's A17b row and the chat_gpt56_sigv4_blank_key case). Converse already treated an empty key as absent. 4900a1a had removed the earlier override (the OpenAI-Project header); this one only resolves the key and sets no other header
  21. The GPT version regex bounds each version group to three digits (openai\.gpt-(\d{1,3})(?!\d)(?:\.(\d{1,3})(?!\d))?), so a model id carrying thousands of digits in its version is classified in constant time and is not a GPT 5.6+ id: it goes to Converse, which answers AWS's error for the unknown model (the *_overlong_gpt_version cases and the audit's A23 row, which also times /health/liveliness during the request). At 708f5a5 the int conversion of such a version raised ValueError (Python's 4300-digit limit). Rejected: raising that limit, since the id is caller-controlled on a shared proxy
  22. The native URL follows Converse's endpoint precedence: aws_bedrock_runtime_endpoint (or AWS_BEDROCK_RUNTIME_ENDPOINT) wins over api_base, so a deployment setting both keeps sending to the one host Converse used (the *_runtime_endpoint* cases and the audit's A18 and A18b rows); at 708f5a5 the native request went to api_base. An api_base already ending in /openai/v1 or in the full /openai/v1/chat/completions path is not doubled (A19)

Caveats (if any)

Medium

  • An int or list reasoning_effort (3, ["high"]) on a native-routed GPT 5.6+ deployment answers litellm's 400 with no call (takes reasoning_effort as a string ... not int) where Converse dropped the value and answered 200 at the default effort; drop_params, global or on the deployment, keeps the 200 by dropping it (design decision 19, audit rows A24 and L8)
  • Default chat completions on unprefixed bedrock/ deployments of us./global. openai.gpt-5.6-{sol,terra,luna}, openai.gpt-6-{astra,sol,luna}, and openai.gpt-6.1-sol now hit runtime Chat Completions instead of Converse, so response ids, service_tier, and the reasoning fields change shape as the User Flow shows; gpt-oss and Grok stay on Converse unless the deployment is prefixed chat_completions/
  • The 400 for temperature, top_p, the penalties, logprobs, or top_logprobs on GPT 5.6+ without reasoning_effort: "none" now names the params and the none way out, where the merge base answered bedrock does not support parameters: [...], so a client matching that text sees a new message
  • Grok 4.3 stays on the bedrock_mantle/ provider. It has no runtime Chat Completions ID
  • bedrock/converse/<model> is still Converse
  • The reasoning models answer "content": null instead of "" when max_tokens runs out inside the reasoning ("finish_reason": "length") on GPT 5.6+ and an opted-in Grok deployment, since AWS's native endpoint returns null there and Converse returned an empty string
  • /v1/messages on an opted-in Grok deployment answers with a text block alone (an empty content list when max_tokens runs out in the reasoning) where Converse also returned a redacted_thinking block, since the route maps thinking to reasoning_effort and goes native
  • AWS's 400 text inside the BedrockException message changes shape on the native route: Converse's {"message": ...} becomes the native {"error": {"message": ..., "type": "invalid_request_error", "param": ..., "code": ...}}, so a client parsing that text sees the new keys
  • Grok reasoning_effort values other than none reach AWS on an opted-in Grok deployment and change the answer's reasoning budget, where Converse dropped every Grok reasoning_effort and AWS applied its default; none is still dropped so that request answers 200 as before (design decision 13)
  • An unknown body key other than additionalModelRequestFields and top_k reaches AWS's native endpoint as sent and is ignored with a 200, where Converse forwarded it in additionalModelRequestFields
  • Params Converse dropped now reach AWS on the native route and answer 200: store, prompt_cache_key, safety_identifier, seed, modalities, web_search_options, logit_bias on GPT 5.6+ and opted-in Grok, and logprobs and top_logprobs on opted-in gpt-oss and Grok (on GPT 5.6+ those two pass under reasoning_effort: "none" only, design decision 7)
  • get_supported_openai_params and /model_group/info on the native-routed ids return the native endpoint's list (31 params on gpt-5.6-sol where the merge base listed 13): requestMetadata, thinking, and output_config leave it and the OpenAI-style params above join it, so a caller reading that list to pre-filter a request sees a different set
  • stop still answers Converse's 400 on every native-routed model, as on main, although gpt-oss and Grok would accept it natively; serving it natively is a follow-up once the reasoning truncation it causes has an answer (design decision 16)
  • The route flips on the price map row, so with LITELLM_LOCAL_MODEL_COST_MAP unset a hosted cost-map update that lists /v1/chat/completions on a further OpenAI GPT 5.6+ Bedrock row moves that deployment to the native path at its next restart without a code upgrade, the same way /v1/responses opt-ins already work for bedrock_supports_openai_responses; any other id needs the chat_completions/ prefix as well
  • bedrock/<region>/xai.grok-4.6 (a bare id after a region path, no profile prefix) still resolves to invoke because no cost-map row exists for the bare id, as on main; the region-path tests cover the profile-prefixed forms
  • The only per-deployment way to stay on Converse is the bedrock/converse/<model> name; there is no config-level switch to pin a whole deployment to Converse
  • The Bedrock provider docs naming which models go native and the feature trade-off live in litellm-docs, in companion PR docs(bedrock): gpt-5.6 and newer default to the native chat completions route, chat_completions/ opts the rest in litellm-docs#1583, which merges right after this one; until then the model list in the TLDR and design decision 2 are the only description

Low

  • gpt-6.1 has no reasoning_effort: "none" on AWS, so its temperature, top_p, penalties, logprobs, and top_logprobs always 400 (or drop under drop_params) and its function tools always take Converse
  • The bare openai.gpt-6-sol, openai.gpt-6-luna, and openai.gpt-6.1-sol rows (no supported_endpoints), gpt-5.4, gpt-5.5, gpt-oss-safeguard, and the zai. GLM ids stay on Converse although AWS serves them natively; a GPT 5.6+ row joins with a one-line cost-map change, the others also need the chat_completions/ prefix (design decision 2)
  • gpt-6 and gpt-6.1 chat completions carrying temperature or top_p under drop_params now answer 200 with the param stripped on Converse too, where the merge base forwarded them and AWS answered 400 (the rows gain supports_sampling_params: false)
  • /v1/messages on GPT 5.6+ answers a text block alone where Converse added an empty redacted_thinking block
  • /v1/responses still refuses temperature under reasoning.effort: none, a pre-existing Responses check this PR leaves alone
  • A reasoning_effort string AWS does not know (an empty string, a made-up effort) reaches the native endpoint and answers its 400, where Converse dropped it and answered 200; an int or list is refused in litellm or dropped instead (the Medium row above)
    • at 19176f6 a list value crashed to a 500 with no upstream call, a0cef91 forwarded it for AWS's 400, and 6daf0b7 refuses or drops it before the call
  • temperature sent as a string under reasoning_effort: "none" now reaches AWS for its 400 (expected a decimal, but got a string) after one call, where the merge base refused it in litellm with no call
  • The shared image fetch, which the Anthropic, Gemini, Vertex, Snowflake, and Bedrock invoke and Responses transforms also use, now infers the type behind a generic Content-Type (URL extension, then magic bytes), so a png, jpeg, gif, webp, heic, or pdf served as application/octet-stream reaches every one of those providers as its real type; every request this changes was already failing or wrong at the merge base, as the kept image cases show (a 400 from Anthropic, OpenAI, Bedrock invoke, Vertex Gemini, and Vertex Claude, Gemini naming the wrong colors, Responses spending its whole budget on reasoning)
  • A remote image whose type cannot be inferred now answers 400 (ImageFetchError, a BadRequestError) in under a second with no provider call, where the merge base answered a 500 (the Grok case) or forwarded data:application/octet-stream for the provider to reject with its own 400 (the Anthropic case); a format the byte sniff does not know (bmp, tiff, avif, svg) behind a generic Content-Type and no known extension lands in that 400 too, where before it was forwarded as data:application/octet-stream and rejected by the provider
  • At 6befa50 all ten required checks are green (Core checks 3.10 to 3.14, Dashboard build, assert-ci-coverage, cost-map-guard, frontend-lint, lint), and the budget-ratchet and rust-test reds the earlier tips carried are green there too, since the merge of main a206882 brought in main's budget files and its fix for litellm-traces::query named::result_contracts_preserve_public_field_names. Two non-required checks are red at the tip, each main's own and neither touched by this PR's diff against main: Verify schema.d.ts matches the proxy OpenAPI spec (https://github.com/BerriAI/litellm/actions/runs/37084822447/job/111092794281) reports only the Trace* components that main's refactor(traces): type the ClickHouse query help response #44285 (dd86ca5) retyped to the enum "otel_traces" | "agent_traces_by_key" | "spend_logs" without regenerating ui/litellm-dashboard/src/lib/http/schema.d.ts (its last change on main is 688d791), the same check passed at a109787 whose merge ref predates that commit, and this PR changes no dashboard type; osv-scan (https://github.com/BerriAI/litellm/actions/runs/37084822524/job/111092794810) flags braces 3.0.3 in main's own ui/litellm-dashboard/package-lock.json under GHSA-vfj7-8cjw-p6xm, whose entry was modified at 2026-10-02T22:45Z, after main's last green scan at d131c43, and this PR does not touch the lockfile. Net negative to fix here: both fixes are main-side changes (a regenerated schema.d.ts, a lockfile bump) that belong in their own PR against main, not in a Bedrock chat completions change, and the schema one is filed as a follow-up on the private team
  • A gpt-oss answer whose visible text itself starts with a literal <reasoning> prefix is split into reasoning_content like the real one, since AWS frames the reasoning inline with no other discriminator; left as is, because a heuristic to tell the two apart would misfire more often than a literal prefix appears
  • One review thread on transformation.py (reasoning output is not scanned by response guardrails, which read content and streamed delta.content only) is left open for a maintainer. The same gap already exists on Converse's reasoning_content, so this PR does not introduce it, and closing it is a guardrail-translation change, not a Bedrock one
  • Bugbot's High finding on litellm/utils.py (Converse-only kwargs such as guardrailConfig invisible to the route decision) was checked live and is a false positive: pre_process_non_default_params copies those kwargs into passed_params before get_optional_params picks the route, so a guardrailConfig request maps to Converse (the chat_gpt56_guardrail case). Greptile's earlier threads (top_k and additionalModelRequestFields on the Converse fallback, allowed_openai_params restoring none, the refused-effort table living in code) were each answered in thread and withdrawn; its P1 at 22e35c6 (sampling params refused on GPT-5.6 even under reasoning_effort: "none") is fixed in 952adfc, which is what design decision 7 describes
  • A remote image the proxy cannot fetch (a blocked address, a 404, an unknown type) answers 400 ImageFetchError on the native route and still 500 APIConnectionError on Converse (the converse/ pin, the guardrail and ARN fallbacks), as on main for every Converse call (the audit's A20a and A20b rows). Net negative to fix here: aligning Converse means touching BedrockImageProcessor, which every Bedrock invoke and Converse model shares, so it is a follow-up on its own
  • An unreachable api_base or a provider that disconnects mid-request answers 503 ServiceUnavailableError on the native route, where Converse's non-stream path answered 500 APIConnectionError (its stream path already answered 503); the chat_gpt56_api_base_only_dead_port case and the audit's B2 row. Shipped as is: 503 is the class the Converse stream path and the other OpenAI-compatible handlers already use, and a 500 for an unreachable upstream was the odd one out
  • /v1/responses on a deployment carrying model_id still calls the base model natively and ignores the ARN, and prefers api_base over aws_bedrock_runtime_endpoint (503 on the responses_gpt56_runtime_endpoint cell on both legs); both as on main, since the Responses route predates this PR. Net negative to fix here: it is the Responses transform, not the chat route, and gets its own follow-up
  • A stream request that lands on a worker which has not yet loaded a deployment created seconds earlier (/model/new under store_model_in_db) goes out without stream_options.include_usage, since the proxy's usage-tracking step runs before the deployment is resolved; pre-existing on main for every provider (1 of 8 immediate streams in the audit's probe), the audit's A15 cell sends stream_options from the client so its wire literal is deterministic. Net negative to fix here: the fix is in common_request_processing.py, shared by every provider, and gets its own follow-up
  • A non-stream /v1/responses spend row can carry the pre-encryption resp_<base64> id instead of the ciphertext the caller received, because the row id is read before the ResponsesIDSecurity hook rewrites it in place; pre-existing on main for every provider, both audit legs show it, and the chaos cells match such a row by the upstream id inside the payload (the test's TODO names it). Net negative to fix here: it is the spend-log id capture, not the Bedrock route, and gets its own follow-up
  • get_supported_openai_params for the native-routed ids inherits the OpenAI-compatible list, which also names audio, modalities, prediction, store, web_search_options, safety_identifier, prompt_cache_key, and prompt_cache_retention; AWS's runtime endpoint answers 200 and ignores them (the Medium row above), so a caller reading the list to pre-filter sends them for nothing. Net negative to fix: trimming the list means a hand-kept per-endpoint table that drifts from AWS's own as it adds support, for params that already answer 200
  • The application inference profile deployments (gpt-5.6-sol-profile, and model_id ARN rows) log spend: 0.0, since the ARN has no cost-map row and the base model is not read back from it; as on main, both legs of the live audit cells show it. Net negative to fix here: cost attribution for profile ARNs is a cost-map lookup change outside this route
  • The Converse /v1/messages call on a runtime-endpoint deployment posted to AWS twice, 12 s apart, on the merge base leg only (messages_gpt56_runtime_endpoint Before); the native route posts once. Left alone, it is the path this PR moves off
  • The chaos cells (B1 to B3) burst on deployments their own proxy reads from its config, so a burst that starts seconds after /model/new on a 2-worker proxy is no longer a cell; on that shape the worker that did not serve /model/new can answer 400 Invalid model name once the registry read-through's resync budget is spent, pre-existing on main for every provider (the audit's hashes note). Net negative to fix here: it is the read-through in litellm/proxy/common_utils/registry_read_through.py, shared by every route, and gets its own follow-up

Final Attestation

The Before and After legs above ran on 2026-10-02 at the merge base c168199 and the tip 6daf0b7 against live AWS Bedrock in us-west-2 with two uvicorn workers per proxy and a recorder in front of AWS, covering the GPT 5.6, 6, and 6.1 defaults, the Converse fallbacks (guardrails, an application inference profile ARN, a model_id ARN override, the converse/ pin, tools with reasoning), a blank-key SigV4 deployment, the reasoning-tied params under none and under other efforts, a non-string reasoning_effort with and without drop_params, streaming, images, /v1/messages, and /v1/responses, plus the gpt-oss and Grok defaults and their chat_completions/ opt-in, and the same cases ran on a local merge into main. The kept image cases ran on 2026-09-26 at 1f77fa6 and 4900a1a (the image cases' own hashes, kept as recorded); the only change to image_handling.py since then (3e3b9cd) hoists the URL comprehension into a helper, so they still describe the tip. The /audit matrix above ran at c168199 and 6daf0b7 on the scripted runtime peer plus real AWS for the live rows, with the two rerun files at e518e0e, a tests-only commit

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • 61a130c passes /live-pr-risk: its tree is byte-identical to ff8e15d, where the earlier legs kept above ran, since the two commits between them add and then revert schema descriptions
  • daba257 passes /live-pr-risk: the image and stop Before/After leg ran at the merge base and this tip (the 4900a1a run above repeats those cases) with the three earlier native paths re-checked, and CI at this tip is 99 passed, 0 failed, 2 skipped
  • e653217 passes /live-pr-risk: the 39-scenario Before/After leg kept above ran at the merge base and this tip, and CI at this tip is 99 passed, 0 failed, 4 skipped (the skips are the Veria, Bugbot, Bugbot Autofix, and conditional stage-mirror non-gates)
  • c60249f passes /live-pr-risk: the two commits since 61a130c are the merge of main (which only aligned two serde attributes in model_info.rs with main's style) and the move of the transformation tests under tests/unit/, so this PR's diff against main is otherwise unchanged from the one every proof above ran against
  • cbf01c2 passes /live-pr-risk: the delta since c60249f is the shared image fetch (type inference behind a generic Content-Type, ImageFetchError for an image whose type cannot be inferred, the sync inline_remote_media) and the Bedrock chat transform (http(s) images inlined before the native call, stop, top_k, and additionalModelRequestFields kept on Converse, Grok reasoning_effort: "none" dropped), and every dependent of the fetch reachable on this rig ran live at the merge base and the tip with a recorder in front of the provider (Bedrock native chat, Responses, and Messages, Bedrock invoke Claude, Anthropic over http://, Gemini image and PDF, OpenAI PDF, Vertex Gemini, Vertex Claude). Breaking: none. Backward incompatible: an image whose type cannot be inferred is a 400 ImageFetchError instead of a 500 or the provider's own 400, and a png, jpeg, gif, webp, heic, or pdf behind application/octet-stream now reaches the provider typed instead of being rejected. Regression risk: none observed, every Before 200 is an After 200 with the same answer. Not verified: Snowflake, Ollama, and Bedrock Mantle (the same fetch helper, no live deployment on this rig) and GovCloud (no credentials)
  • 4900a1a passes /live-pr-risk: the delta since cbf01c2 removes the native config's validate_environment override, so the route inherits the OpenAI-compatible one (Content-Type only, a bearer only when an api_key is given) while sign_request still signs every call (SigV4, or the AWS_BEARER_TOKEN_BEDROCK bearer the legs used), and aws_bedrock_project_id no longer becomes an OpenAI-Project header. Dependents: every native chat request (the whole After leg re-ran at this tip with the same status codes and outbound paths as at cbf01c2), a deployment carrying aws_bedrock_project_id (the chat_gptoss_project_id case, 200 on both legs with no OpenAI-Project among the recorded request headers on either), the Mantle configs (untouched, they keep sending it), and the native Responses transform (never sent it). Breaking: none. Backward incompatible: a runtime deployment configured with aws_bedrock_project_id under an earlier revision of this PR stops sending a header the runtime host ignored (200 to a bogus value, checked live). Regression risk: none observed
  • a0cef91 passes /live-pr-risk: the delta since 4900a1a is 19176f6 (the reasoning_effort: "none" check in the route decision and the refused-param filter became an equality test, so a list value no longer raises there) and a0cef91 (a type guard in the refused-effort filter, so a non-string reasoning_effort is forwarded instead of crashing). Dependents: the route decision for every Bedrock chat, Messages, and Responses request, get_supported_openai_params and /model_group/info on Bedrock ids, the Converse fallback for guardrailConfig, the converse/ prefix and application inference profile ARNs, cost and usage on all three endpoints, and streaming with include_usage. Every one ran live at the merge base and at this tip with a forwarding recorder in front of Bedrock, one proxy per side with 2 workers (the two sides ran one after the other on the same host at CRITICAL headroom rather than at once), and the same cells re-ran on a local merge of this tip into main 54ae4c5 with identical routes, statuses, wire bodies, and rates. Breaking: a list reasoning_effort answered 500 (unhashable type: 'list', no upstream call) at 19176f6 where the merge base answered 200, fixed in a0cef91 (AWS now answers its own 400 after one call). Backward incompatible: get_supported_openai_params and /model_group/info list the native endpoint's 31 params for GPT 5.6+ ids where the merge base listed 13 or 15 (design decision 7, shipped by the author); a Converse-routed gpt-6 request carrying temperature (a guardrailConfig request) now 400s in litellm with no upstream call where the merge base let AWS 400 it after one call, and drops it under drop_params (the supports_sampling_params Low caveat); an int or list reasoning_effort answers AWS's 400 where Converse dropped it and answered 200 (the Low caveat above); a string temperature under reasoning_effort: "none" reaches AWS for its 400 where litellm refused it with no call; the 400 message for sampling params on GPT 5.6+ names the none way out (the Medium caveat above). Regression risk: none observed beyond those; a hosted cost-map update that lists /v1/chat/completions on a further GPT 5.6+ row moves that deployment to the native route at its next restart (the Medium caveat above). Not verified: Bedrock Mantle and GovCloud (no deployment or credentials on this rig), and a deployment reading the hosted cost map rather than LITELLM_LOCAL_MODEL_COST_MAP
  • 578f26e passes /live-pr-risk: the delta since a0cef91 is dae29c1 (the GPT version regex bounds its digits, so a 4300-digit id no longer raises ValueError, and the native URL is built from aws_bedrock_runtime_endpoint like Converse) and 578f26e (model_id joins the Converse-only keys; the native config resolves the bearer through bedrock_bearer_token). Dependents walked for both: the route decision for every Bedrock chat, Messages, and Responses request (get_bedrock_route, bedrock_route_for_request), get_supported_openai_params and /model_group/info, the Converse fallbacks (guardrailConfig, converse/, the ARN and model_id forms), BaseAWSLLM._sign_request and every caller (Converse, invoke, Responses, the pass-through routes), every reader of aws_bedrock_runtime_endpoint, image inlining, cost and usage on all three endpoints, and streaming with include_usage. Every one ran live at the merge base and at this tip with a forwarding recorder in front of Bedrock, one proxy per side with 2 workers, and the same 43 cases ran on a local merge of this tip into main 4b1d9bf (4f930db3a3) with identical routes, statuses, and wire bodies. Breaking: none (at dae29c1 a model_id deployment answered 400 and a blank-key SigV4 deployment 403, both fixed in 578f26e before any release, shown in the Before fix 2 block). Backward incompatible, each shipped by this run as a decision row in the /audit matrix for the reviewer's read: a remote image the proxy cannot fetch answers 400 ImageFetchError on the native route where Converse and main answer 500; an unreachable api_base or a mid-request disconnect answers 503 where Converse's non-stream path answered 500; get_supported_openai_params lists the native endpoint's 31 params where the merge base listed 13; the sampling-param 400 names the none way out; and one left to the author's decision at the read (the Severe caveat at that tip, decided in 6daf0b7): an int or list reasoning_effort answers AWS's 400 where Converse dropped it and answered 200. Regression risk: a hosted cost-map update listing /v1/chat/completions on a further GPT 5.6+ row moves it to the native route at restart (Medium caveat); the cold-worker stream_options race is pre-existing and provider-independent (Low caveat). Not verified: Bedrock Mantle and GovCloud (no deployment or credentials on this rig), and a deployment reading the hosted cost map rather than LITELLM_LOCAL_MODEL_COST_MAP
  • 6daf0b7 passes /live-pr-risk: the delta since 578f26e is non_string_reasoning_effort in the native transform (an int or list reasoning_effort answers litellm's 400 before any call, or is dropped under drop_params). Dependents walked: map_openai_params on every native-routed chat request, the drop_params resolution (global, deployment, request body), the refused-param filter next to it, and /v1/messages and /v1/responses carrying a non-string effort. Every one ran live at the merge base and this tip with a forwarding recorder in front of Bedrock, one proxy per side with 2 workers (the 30 cells above), and the same cells ran on a local merge of this tip into main d729f97 (33d3c0f5ff) with identical routes, statuses, and wire bodies. Breaking: none. Backward incompatible: an int or list reasoning_effort on a native-routed deployment without drop_params answers litellm's 400 with no AWS call where Converse dropped it and answered 200 (the Medium caveat, design decision 19, shipped by the author), and /v1/messages with a top-level non-string reasoning_effort answers the same 400; under drop_params both answer 200 as before. Regression risk: none observed. Not verified: Bedrock Mantle and GovCloud (no deployment or credentials on this rig), and a deployment reading the hosted cost map rather than LITELLM_LOCAL_MODEL_COST_MAP
  • e518e0e passes /audit: round 2 of 3 at 6daf0b7, the last commit with a product diff, with the sad-path and chaos files rerun at e518e0e (a tests-only commit). Every inventory row ran on both legs and matched its written expectation, with the matrix, hashes, run ids, and live artifacts in the proof section above
  • 6befa50 passes /live-pr-risk: the delta since 6daf0b7 is e518e0e (tests only), the merge of main a206882 (a109787, which changes nothing in this PR's diff against main), and this tests-only commit hardening the B4 chaos cell's teardown; the merge-ref check is the audit rerun on the merged tip named in the /audit line below
  • 6befa50 passes /audit: carried from e518e0e over the merge of main and a tests-only commit (no product diff of this PR's own), with all six audit files rerun at a109787 against a head proxy booted from the merged tip (merged_head1, 54 passed and the B4 teardown failure its row records) and the chaos file rerun at 6befa50 (merged_head4 4 passed in 133 s and merged_head5 4 passed in 120 s, no skips, no retries, against a head proxy booted from 6befa50 with two uvicorn workers)

Note

High Risk
Changes default routing for many Bedrock GPT deployments (Converse to native), altering IDs, reasoning fields, and param validation; incorrect fallback logic could misroute guardrails, tools, or structured output.

Overview
Adds native Bedrock Runtime OpenAI Chat Completions (/openai/v1/chat/completions) for GPT 5.6+ by default and optional bedrock/chat_completions/ opt-in for gpt-oss and Grok, with a new AmazonBedrockRuntimeChatCompletionsConfig handling AWS signing, param mapping (max_completion_tokens, reasoning-effort rules), <reasoning> tag splitting, and remote image inlining.

Routing is now request-aware: bedrock_route_for_request and updated get_bedrock_route pick native chat vs Converse from cost-map flags and Converse-only features (guardrails, model_id ARNs, tools+reasoning, response_format, etc.), keeping param mapping and dispatch aligned. Completion dispatch passes raw request_params into that decision.

Supporting changes: shared inline_remote_media and smarter image MIME inference for native vision; price-map/catalog flags for native response-format and tools-with-reasoning; Responses config strips chat_completions/ from model names for capability checks. Broad integration tests cover wire, sad paths, and chaos.

Reviewed by Cursor Bugbot for commit 6befa50. Bugbot is set up for automated code reviews on this repo. Configure here.

Unspecified bedrock grok was rewritten to Converse. Chat completions now hit bedrock-runtime /openai/v1/chat/completions, and converse/ still uses Converse
@mateo-berri
mateo-berri requested a review from a team September 11, 2026 19:23
@greptile-apps

greptile-apps Bot commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[High risk] Adds new routing path for Bedrock chat completions API.

The PR appears safe to merge; no new actionable issue remains after the latest test-only teardown hardening.

Summary

Adds native Bedrock Runtime Chat Completions support and request-aware routing while preserving Converse for incompatible request features.

  • Routes eligible GPT 5.6+ models natively by default and supports explicit native or Converse route prefixes.
  • Adds native parameter validation, AWS signing, endpoint resolution, streaming response handling, reasoning extraction, image inlining, and response metadata preservation.
  • Extends model capability metadata and adds comprehensive unit, wire, failure-path, and multi-worker integration coverage.
  • The only change since the previous review hardens a chaos test by waiting for the replacement Uvicorn worker to finish startup before proxy teardown.

Reviews (39) · Last reviewed commit: "test(bedrock): wait for the replacement ..."

Comment thread litellm/llms/bedrock/chat/chat_completions/transformation.py Outdated
Comment thread litellm/llms/bedrock/chat/chat_completions/transformation.py
@codspeed

codspeed Bot commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_bedrock_grok_chat_completions (6befa50) with main (dd86ca5)

Open in CodSpeed

@codecov

codecov Bot commented Sep 11, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.63014% with 4 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...ms/bedrock/chat/chat_completions/transformation.py 97.77% 4 Missing ⚠️

📢 Thoughts on this report? Let us know!

@tin-berri tin-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri mateo-berri changed the title feat(bedrock): send grok chat completions through runtime openai path feat(bedrock): serve gpt-oss, gpt-5.6, and grok chat completions on runtime's native openai path Sep 20, 2026
Comment thread litellm/llms/bedrock/chat/chat_completions/transformation.py Dismissed
Comment thread litellm/utils.py Fixed
@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 20, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…at Completions

The cost-map parity tests require every regional variant of a flagged id to carry the same supports_ flags, so the six us-gov gpt-oss entries now carry the native-route flags too. A region path in the model name (bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0) is routing, not a different model: the route is looked up on the id after the path, the path's region picks the endpoint and the SigV4 scope, an explicit aws_region_name still wins, and the body carries the bare id AWS expects
@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 20, 2026
Comment thread tests/integration/providers/test_bedrock_runtime_chat_completions_chaos.py Outdated
Comment thread tests/integration/providers/test_bedrock_gpt_responses_native_wire.py Outdated

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.

Autofix Details

Bugbot Autofix prepared fixes for both issues found in the latest run.

  • ✅ Fixed: Responses spend lookup uses encrypted id
    • The non-stream Responses wire test now looks up spend by the pre-encryption issued id, matching the stream case and the logged request_id.
  • ✅ Fixed: Peer-kill burst test can flake
    • The peer counts answered requests after respond() returns, and the test waits until those six client bodies land before killing so only the held calls fail.

Create PR

Or push these changes by commenting:

@cursor push 593f9f2948
Preview (593f9f2948)
diff --git a/tests/integration/_support/bedrock_runtime_peer.py b/tests/integration/_support/bedrock_runtime_peer.py
--- a/tests/integration/_support/bedrock_runtime_peer.py
+++ b/tests/integration/_support/bedrock_runtime_peer.py
@@ -1,3 +1,4 @@
+import itertools
 import json
 import re
 import threading
@@ -263,14 +264,21 @@
 
 def serve_peer(port: int, received: Synchronized[int], answer_first: int) -> None:
     held: Final = threading.Event()
+    ordinals: Final = itertools.count(1)
+    assign: Final = threading.Lock()
 
     def respond_or_hold(request: Request) -> Reply:
+        with assign:
+            ordinal: Final = next(ordinals)
+        if ordinal > answer_first:
+            with received.get_lock():
+                received.value += 1
+            held.wait()
+            return respond(request)
+        reply: Final = respond(request)
         with received.get_lock():
             received.value += 1
-            ordinal: Final = received.value
-        if ordinal > answer_first:
-            held.wait()
-        return respond(request)
+        return reply
 
     with wire_server(respond_or_hold, port=port):
         threading.Event().wait()

diff --git a/tests/integration/providers/test_bedrock_gpt_responses_native_wire.py b/tests/integration/providers/test_bedrock_gpt_responses_native_wire.py
--- a/tests/integration/providers/test_bedrock_gpt_responses_native_wire.py
+++ b/tests/integration/providers/test_bedrock_gpt_responses_native_wire.py
@@ -85,10 +85,11 @@
         response: Final = raw.parse()
         assert response.output_text == answer(marker), raw.text
         assert response.usage is not None and (response.usage.input_tokens, response.usage.output_tokens) == (30, 5)
-        assert _issued_id(response.id).upstream == f"resp_upstream_{marker}", response.id
+        issued: Final = _issued_id(response.id)
+        assert issued.upstream == f"resp_upstream_{marker}", response.id
         request: Final = _native_request(wire)
         assert _body(request) == {"model": GPT, "input": _prompt(marker)}, request.body
-        assert _spend_row(response.id) == _success_row(model)
+        assert _spend_row(issued.issued) == _success_row(model)
 
 
 async def test_async_openai_sdk_responses_stream_is_served_by_the_native_responses_route(gateway: Gateway) -> None:

diff --git a/tests/integration/providers/test_bedrock_runtime_chat_completions_chaos.py b/tests/integration/providers/test_bedrock_runtime_chat_completions_chaos.py
--- a/tests/integration/providers/test_bedrock_runtime_chat_completions_chaos.py
+++ b/tests/integration/providers/test_bedrock_runtime_chat_completions_chaos.py
@@ -175,7 +175,13 @@
         assert len(owned) == 1, (item.call, owned, success_ids)
 
 
-async def _send(client: httpx.AsyncClient, key: str, model: str, call: _Call) -> _Served:
+async def _send(
+    client: httpx.AsyncClient,
+    key: str,
+    model: str,
+    call: _Call,
+    completed: Synchronized[int] | None = None,
+) -> _Served:
     async with client.stream(
         "POST",
         _path(call.endpoint),
@@ -183,17 +189,27 @@
         headers={"Authorization": f"Bearer {key}", "anthropic-version": "2023-06-01"},
     ) as response:
         raw: Final = await response.aread()
+    if completed is not None:
+        with completed.get_lock():
+            completed.value += 1
     return _Served(
         call=call, status=response.status_code, text=raw.decode(), call_id=response.headers.get("x-litellm-call-id")
     )
 
 
 async def _burst(
-    base_url: str, key: str, model: str, calls: tuple[_Call, ...], *, tolerate_transport_errors: bool = False
+    base_url: str,
+    key: str,
+    model: str,
+    calls: tuple[_Call, ...],
+    *,
+    tolerate_transport_errors: bool = False,
+    completed: Synchronized[int] | None = None,
 ) -> tuple[_Served, ...]:
     async with httpx.AsyncClient(base_url=base_url, timeout=60, trust_env=False) as client:
         results: Final = await asyncio.gather(
-            *(_send(client, key, model, call) for call in calls), return_exceptions=tolerate_transport_errors
+            *(_send(client, key, model, call, completed) for call in calls),
+            return_exceptions=tolerate_transport_errors,
         )
     for result in results:
         assert not isinstance(result, BaseException) or isinstance(result, httpx.TransportError), repr(result)
@@ -260,11 +276,15 @@
     calls: Final = _calls(12, _ENDPOINTS, lambda index: index % 2 == 0)
     recovery: Final = _calls(6, _ENDPOINTS, lambda index: index % 2 == 1)
     port: Final = _free_port()
+    answered: Final = multiprocessing.Value("i", 0)
     with gateway.scenario() as scenario:
         model: Final = _deployment(scenario, f"http://127.0.0.1:{port}")
         with _child_peer(port, answer_first=6) as peer:
-            burst: Final = asyncio.create_task(_burst(str(gateway.client.base_url), gateway.key, model, calls))
+            burst: Final = asyncio.create_task(
+                _burst(str(gateway.client.base_url), gateway.key, model, calls, completed=answered)
+            )
             await asyncio.to_thread(eventually, lambda: peer.received.value, lambda count: count == 12, 60)
+            await asyncio.to_thread(eventually, lambda: answered.value, lambda count: count == 6, 60)
             peer.process.kill()
             peer.process.join(timeout=10)
             served: Final = await burst

You can send follow-ups to the cloud agent here.

Comment thread tests/integration/providers/test_bedrock_gpt_responses_native_wire.py Outdated
Comment thread tests/integration/providers/test_bedrock_runtime_chat_completions_chaos.py Outdated
The chaos peer's shared counter and process now come from the same spawn context, since a fork-context Value handed to a spawn-context process raises on Linux. The peer-kill test waits for the first six answers to reach the client before killing the peer instead of counting accepted requests. The Responses wire tests look the spend row up under both the ciphertext id the caller received and the issued id behind it, matching the chaos file's rule for the pre-encryption row
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

… native chat completions call

A reasoning_effort sent as an int, a list, or an object on a GPT 5.6+ deployment the native
route serves now answers 400 from litellm before any wire request, naming the type and the
drop_params way out, and is dropped under drop_params so AWS applies its default effort, the
way Converse dropped it on main. The tip since a0cef91 forwarded it for AWS to refuse
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri
mateo-berri enabled auto-merge (squash) October 2, 2026 20:24
…chat_completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 6befa50. Configure here.

@mateo-berri
mateo-berri disabled auto-merge October 3, 2026 01:43
@mateo-berri

Copy link
Copy Markdown
Contributor Author

Closing in favor of #44307, the same branch tip opened Devin-authored so a code owner approval can land it without waiting on a reviewer

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants