Repository navigation
feat(bedrock): serve gpt-5.6+ chat completions natively by default, with chat_completions/ opt-in for gpt-oss and grok - #40775
mateo-berri wants to merge 40 commits into
Conversation
Unspecified bedrock grok was rewritten to Converse. Chat completions now hit bedrock-runtime /openai/v1/chat/completions, and converse/ still uses Converse
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
…s native openai path
…e route once from the raw request
|
bugbot run |
…at Completions The cost-map parity tests require every regional variant of a flagged id to carry the same supports_ flags, so the six us-gov gpt-oss entries now carry the native-route flags too. A region path in the model name (bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0) is routing, not a different model: the route is looked up on the id after the path, the path's region picks the endpoint and the SigV4 scope, an explicit aws_region_name still wins, and the body carries the bare id AWS expects
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.
Autofix Details
Bugbot Autofix prepared fixes for both issues found in the latest run.
- ✅ Fixed: Responses spend lookup uses encrypted id
- The non-stream Responses wire test now looks up spend by the pre-encryption issued id, matching the stream case and the logged request_id.
- ✅ Fixed: Peer-kill burst test can flake
- The peer counts answered requests after respond() returns, and the test waits until those six client bodies land before killing so only the held calls fail.
Or push these changes by commenting:
@cursor push 593f9f2948
Preview (593f9f2948)
diff --git a/tests/integration/_support/bedrock_runtime_peer.py b/tests/integration/_support/bedrock_runtime_peer.py
--- a/tests/integration/_support/bedrock_runtime_peer.py
+++ b/tests/integration/_support/bedrock_runtime_peer.py
@@ -1,3 +1,4 @@
+import itertools
import json
import re
import threading
@@ -263,14 +264,21 @@
def serve_peer(port: int, received: Synchronized[int], answer_first: int) -> None:
held: Final = threading.Event()
+ ordinals: Final = itertools.count(1)
+ assign: Final = threading.Lock()
def respond_or_hold(request: Request) -> Reply:
+ with assign:
+ ordinal: Final = next(ordinals)
+ if ordinal > answer_first:
+ with received.get_lock():
+ received.value += 1
+ held.wait()
+ return respond(request)
+ reply: Final = respond(request)
with received.get_lock():
received.value += 1
- ordinal: Final = received.value
- if ordinal > answer_first:
- held.wait()
- return respond(request)
+ return reply
with wire_server(respond_or_hold, port=port):
threading.Event().wait()
diff --git a/tests/integration/providers/test_bedrock_gpt_responses_native_wire.py b/tests/integration/providers/test_bedrock_gpt_responses_native_wire.py
--- a/tests/integration/providers/test_bedrock_gpt_responses_native_wire.py
+++ b/tests/integration/providers/test_bedrock_gpt_responses_native_wire.py
@@ -85,10 +85,11 @@
response: Final = raw.parse()
assert response.output_text == answer(marker), raw.text
assert response.usage is not None and (response.usage.input_tokens, response.usage.output_tokens) == (30, 5)
- assert _issued_id(response.id).upstream == f"resp_upstream_{marker}", response.id
+ issued: Final = _issued_id(response.id)
+ assert issued.upstream == f"resp_upstream_{marker}", response.id
request: Final = _native_request(wire)
assert _body(request) == {"model": GPT, "input": _prompt(marker)}, request.body
- assert _spend_row(response.id) == _success_row(model)
+ assert _spend_row(issued.issued) == _success_row(model)
async def test_async_openai_sdk_responses_stream_is_served_by_the_native_responses_route(gateway: Gateway) -> None:
diff --git a/tests/integration/providers/test_bedrock_runtime_chat_completions_chaos.py b/tests/integration/providers/test_bedrock_runtime_chat_completions_chaos.py
--- a/tests/integration/providers/test_bedrock_runtime_chat_completions_chaos.py
+++ b/tests/integration/providers/test_bedrock_runtime_chat_completions_chaos.py
@@ -175,7 +175,13 @@
assert len(owned) == 1, (item.call, owned, success_ids)
-async def _send(client: httpx.AsyncClient, key: str, model: str, call: _Call) -> _Served:
+async def _send(
+ client: httpx.AsyncClient,
+ key: str,
+ model: str,
+ call: _Call,
+ completed: Synchronized[int] | None = None,
+) -> _Served:
async with client.stream(
"POST",
_path(call.endpoint),
@@ -183,17 +189,27 @@
headers={"Authorization": f"Bearer {key}", "anthropic-version": "2023-06-01"},
) as response:
raw: Final = await response.aread()
+ if completed is not None:
+ with completed.get_lock():
+ completed.value += 1
return _Served(
call=call, status=response.status_code, text=raw.decode(), call_id=response.headers.get("x-litellm-call-id")
)
async def _burst(
- base_url: str, key: str, model: str, calls: tuple[_Call, ...], *, tolerate_transport_errors: bool = False
+ base_url: str,
+ key: str,
+ model: str,
+ calls: tuple[_Call, ...],
+ *,
+ tolerate_transport_errors: bool = False,
+ completed: Synchronized[int] | None = None,
) -> tuple[_Served, ...]:
async with httpx.AsyncClient(base_url=base_url, timeout=60, trust_env=False) as client:
results: Final = await asyncio.gather(
- *(_send(client, key, model, call) for call in calls), return_exceptions=tolerate_transport_errors
+ *(_send(client, key, model, call, completed) for call in calls),
+ return_exceptions=tolerate_transport_errors,
)
for result in results:
assert not isinstance(result, BaseException) or isinstance(result, httpx.TransportError), repr(result)
@@ -260,11 +276,15 @@
calls: Final = _calls(12, _ENDPOINTS, lambda index: index % 2 == 0)
recovery: Final = _calls(6, _ENDPOINTS, lambda index: index % 2 == 1)
port: Final = _free_port()
+ answered: Final = multiprocessing.Value("i", 0)
with gateway.scenario() as scenario:
model: Final = _deployment(scenario, f"http://127.0.0.1:{port}")
with _child_peer(port, answer_first=6) as peer:
- burst: Final = asyncio.create_task(_burst(str(gateway.client.base_url), gateway.key, model, calls))
+ burst: Final = asyncio.create_task(
+ _burst(str(gateway.client.base_url), gateway.key, model, calls, completed=answered)
+ )
await asyncio.to_thread(eventually, lambda: peer.received.value, lambda count: count == 12, 60)
+ await asyncio.to_thread(eventually, lambda: answered.value, lambda count: count == 6, 60)
peer.process.kill()
peer.process.join(timeout=10)
served: Final = await burstYou can send follow-ups to the cloud agent here.
The chaos peer's shared counter and process now come from the same spawn context, since a fork-context Value handed to a spawn-context process raises on Linux. The peer-kill test waits for the first six answers to reach the client before killing the peer instead of counting accepted requests. The Responses wire tests look the spend row up under both the ciphertext id the caller received and the issued id behind it, matching the chaos file's rule for the pre-encryption row
|
bugbot run |
… native chat completions call A reasoning_effort sent as an int, a list, or an object on a GPT 5.6+ deployment the native route serves now answers 400 from litellm before any wire request, naming the type and the drop_params way out, and is dropped under drop_params so AWS applies its default effort, the way Converse dropped it on main. The tip since a0cef91 forwarded it for AWS to refuse
|
bugbot run |
…g deployments on an owned proxy
|
bugbot run |
…chat_completions Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
…e sigkill chaos proxy
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 6befa50. Configure here.
|
Closing in favor of #44307, the same branch tip opened Devin-authored so a code owner approval can land it without waiting on a reviewer |

TLDR
Problem this solves:
/openai/v1/chat/completionstemperatureandtop_pon these models at every effortHow it solves it:
bedrock/<gpt-5.6+ id>goes native when its cost-map row lists/v1/chat/completionsguardrailConfig, application inference profile ARNs, the other Converse-only body keys,stop,top_k, and tools with reasoning onreasoning_effort: "none",temperature,top_p, the penalties, and logprobs pass through nativelydrop_paramsreasoning_effortanswers 400 in litellm, or drops underdrop_paramsbedrock/chat_completions/<id>opts gpt-oss and Grok in;bedrock/converse/<id>pins Conversemodel_idnames an application inference profile ARN keeps Converse at the ARN URL, like thebedrock/arn:...formapi_key: ""is signed, never sent as an empty bearerhttp(s)://images are inlined asdata:URLs before the native calljson_schemaresponse_formatgoes native on GPT 5.6+ and Grok;json_objectkeeps Converse's handlingusage, andservice_tieras AWS sent itIntentional product change: unprefixed
bedrock/deployments of theus.andglobal.GPT 5.6, 6, and 6.1 ids move from Converse to runtime Chat Completions, so response ids,service_tier, and the reasoning fields change shape as the User Flow shows; gpt-oss and Grok stay on Converse unless the deployment is prefixedchat_completions/User Flow
Before: a developer whose app sends GPT 5.6 chat completions through the gateway gets a Converse-shaped answer, a 400 the moment they send a sampling param, and a silent 200 when they send a malformed
reasoning_effort"model": "gpt-5.6-sol"(abedrock/global.openai.gpt-5.6-soldeployment) and a user messagechatcmpl-<uuid>id with noservice_tierand noreasoning_tokensinusage"temperature": 0.2,"top_p": 0.9, and"reasoning_effort": "none"and get 400bedrock does not support parameters: ['temperature', 'top_p']gpt-6-solwithdrop_paramson still gets AWS's 400This model doesn't support the temperature field"reasoning_effort": 3(an int, or a list such as["high"]) ongpt-5.6-soland get 200 at AWS's default effort, with nothing in the answer saying the value was ignored; the same on adrop_paramsdeploymentguardrailConfig, or aimed at an application inference profile ARN deployment, answers 200After: the same developer gets Bedrock's own Chat Completions answer, with sampling params honored under
reasoning_effort: "none", refused by name otherwise, and a malformedreasoning_effortrefused by name unless the deployment drops params"model": "gpt-5.6-sol"(abedrock/global.openai.gpt-5.6-soldeployment) and a user messagechatcmpl-...id with"service_tier": "default"andreasoning_tokensinusage"temperature": 0.2,"top_p": 0.9, and"reasoning_effort": "none"and get 200 with both honored; withoutnonethe 400 names the params and that way outgpt-6-solwithdrop_paramson gets 200 with the params stripped"reasoning_effort": 3ongpt-5.6-soland get 400 with the message "global.openai.gpt-5.6-sol takes reasoning_effort as a string on Bedrock's Chat Completions endpoint, not int. Send one of its named efforts, or setlitellm.drop_params = Trueto drop it" and no call made to AWS; the same request on adrop_paramsdeployment, or with"drop_params": truein the body, gets 200 at AWS's default effort as beforeguardrailConfig, or aimed at an application inference profile ARN deployment, answers 200 exactly as before, through ConverseLinear ticket
Resolves LIT-8684
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Last updated: 6befa50. QA and /live-pr-risk ran on 6daf0b7, the last commit with a product diff
Before is the merge base c168199 and After is the tip 6daf0b7, each booted from its own worktree with
python litellm/proxy/proxy_cli.py --config config.yaml --port <port> --num_workers 2 --use_prisma_db_pushon 2026-10-02 (Before on 36349, After on 20753), withDATABASE_URL(one Postgres database per leg) andLITELLM_LICENSEexported explicitly. Both read the config below and the same.envnames:LITELLM_MASTER_KEY,LITELLM_LOCAL_MODEL_COST_MAP=True,AWS_REGION_NAME=us-west-2, andAWS_BEARER_TOKEN_BEDROCK.BEDROCK_BASEpoints at a small forwarding recorder on127.0.0.1(one per leg, 37807 and 28723) that logs each request with the status AWS answered and passes it unchanged tohttps://bedrock-runtime.us-west-2.amazonaws.com, so step 3 of every case is the AWS path and body as they left the proxy (the body digest shows the model, the message count, and the params that matter; a case showing the same path twice is the proxy's retry after AWS answered 500 to the first attempt, which the recorder logged on both legs). Thegpt-5.6-sol-sigv4-blank-keydeployment has noapi_baseand signs the Host header, so its calls went straight to AWS and its step 3 reads "none". Every call reached AWS and cost real money. The same 40 cases ran once more on a local merge of 6daf0b7 into main d729f97 (33d3c0f5ff, port 41217): every recorded AWS path and body matched the After leg (on 4 of them AWS answered 500 to a first attempt on one leg and the proxy retried, so only the attempt count differs), and only the model's answer text differed, on 10 of them (qa/threeway-6daf0b7136.txt)gpt-overlong-versionnamesbedrock/openai.gpt-followed by 4,301 sevens (cut above for width), the worst case for the GPT version regex. The application inference profile6mv9jwwiiudowrapsglobal.openai.gpt-5.6-solin the same account, and thegpt-5.6-sol-model-iddeployment reaches it throughmodel_idinstead.$TOOLSis oneget_weathertool,$GUARDis a Bedrock guardrail in the same account (49e7k3qphvzq, versionDRAFT, trace on), and$FIXis a 96x64 PNG, red on the left and blue on the right, served by GitHub asapplication/octet-streambehind a 302:[{"type":"function","function":{"name":"get_weather","description":"Get the current weather for a city","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}}]{"guardrailIdentifier":"49e7k3qphvzq","guardrailVersion":"DRAFT","trace":"enabled"}Every curl below also carried
-H 'Authorization: Bearer $LITELLM_MASTER_KEY' -H 'Content-Type: application/json', dropped for width. Bodies are trimmed for width too:createdand the token detail objects are cut, empty fields are cut, reasoning payloads show as…, a Responses answer keeps its output text, reasoning, and usage, a stream shows its first two and last threedata:lines with the total count, and an error shows its first 400 charactersBefore (c168199)
chat_gpt56_basic
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}};POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt6_basic
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt61_basic
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6.1-sol","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-6.1-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt6_json_schema
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Name the capital of France."}],"response_format":{"type":"json_schema","json_schema":{"name":"capital","strict":true,"schema":{"type":"object","properties":{"city":{"type":"string"},"country":{"type":"string"}},"required":["city","country"],"additionalProperties":false}}}}'POST /model/global.openai.gpt-6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "toolConfig.tools": ["json_tool_call"]}chat_gpt56_tools_reasoning_none
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"What is the weather in Paris right now?"}],"tools":'"$TOOLS"',"reasoning_effort":"none"}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "additionalModelRequestFields": {"reasoning": {"effort": "none"}}, "toolConfig.tools": ["get_weather"]}chat_gpt6_tools_reasoning_low
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"What is the weather in Paris right now?"}],"tools":'"$TOOLS"',"reasoning_effort":"low"}'POST /model/global.openai.gpt-6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "additionalModelRequestFields": {"reasoning": {"effort": "low"}}, "toolConfig.tools": ["get_weather"]}chat_gpt61_tools
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6.1-sol","messages":[{"role":"user","content":"What is the weather in Paris right now?"}],"tools":'"$TOOLS"'}'POST /model/global.openai.gpt-6.1-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "toolConfig.tools": ["get_weather"]}chat_gpt56_guardrail
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"guardrailConfig":'"$GUARD"'}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "guardrailConfig": {"guardrailIdentifier": "49e7k3qphvzq", "guardrailVersion": "DRAFT", "trace": "enabled"}};POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "guardrailConfig": {"guardrailIdentifier": "49e7k3qphvzq", "guardrailVersion": "DRAFT", "trace": "enabled"}};POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "guardrailConfig": {"guardrailIdentifier": "49e7k3qphvzq", "guardrailVersion": "DRAFT", "trace": "enabled"}}chat_gpt56_app_profile
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-profile","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt56_model_id
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-model-id","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt56_sigv4_blank_key
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-sigv4-blank-key","messages":[{"role":"user","content":"Say hi in three words."}]}'chat_gpt56_converse_prefix
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-converse","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt55_basic
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.5","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.5/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gptoss_basic
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-oss-20b","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/openai.gpt-oss-20b-1%3A0/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gptoss_native_prefix
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-oss-20b-native","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/chat_completions%2Fopenai.gpt-oss-20b-1%3A0/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_grok_basic
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"grok-4.6","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/us.xai.grok-4.6/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_grok_native_prefix
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"grok-4.6-native","messages":[{"role":"user","content":"Say hi in three words."}]}'chat_gpt56_reasoning_low
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"What is 17*23? Answer with the number only."}],"reasoning_effort":"low"}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "additionalModelRequestFields": {"reasoning": {"effort": "low"}}};POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "additionalModelRequestFields": {"reasoning": {"effort": "low"}}}chat_gpt56_temperature
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9}'chat_gpt56_temperature_drop_params
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}};POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt56_penalty_drop_params
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"frequency_penalty":0.5,"presence_penalty":0.5}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt6_temperature
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2}'POST /model/global.openai.gpt-6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {"temperature": 0.2}}chat_gpt6_temperature_drop_params
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9}'POST /model/global.openai.gpt-6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {"temperature": 0.2, "topP": 0.9}}chat_gpt6_penalty_drop_params
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"frequency_penalty":0.5,"presence_penalty":0.5}'POST /model/global.openai.gpt-6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt56_low_temperature
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"reasoning_effort":"low"}'chat_gpt56_none_temperature
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9,"reasoning_effort":"none"}'chat_gpt56_none_penalties_logprobs
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"frequency_penalty":0.5,"presence_penalty":0.5,"logprobs":true,"top_logprobs":2,"reasoning_effort":"none"}'chat_gpt6_none_temperature
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"reasoning_effort":"none"}'POST /model/global.openai.gpt-6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {"temperature": 0.2}, "additionalModelRequestFields": {"reasoning": {"effort": "none"}}}chat_gpt56_none_temperature_stream
curl -sS -N -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"reasoning_effort":"none","stream":true,"stream_options":{"include_usage":true}}'chat_gpt56_stream
curl -sS -N -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"stream":true,"stream_options":{"include_usage":true}}'POST /model/global.openai.gpt-5.6-sol/converse-streamwith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt56_image_fixture_png
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":[{"type":"text","text":"'"$Q_COLORS"'"},{"type":"image_url","image_url":{"url":"'"$FIX"'"}}]}]}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "image_blocks": 1}messages_gpt56
curl -sS -X POST http://127.0.0.1:36349/v1/messages -d '{"model":"gpt-5.6-sol","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {"maxTokens": 200}}responses_gpt56
curl -sS -X POST http://127.0.0.1:36349/v1/responses -d '{"model":"gpt-5.6-sol","input":"Say hi in three words."}'POST /openai/v1/responseswith{"model": "global.openai.gpt-5.6-sol", "messages": 0}messages_gpt56_model_id
curl -sS -X POST http://127.0.0.1:36349/v1/messages -d '{"model":"gpt-5.6-sol-model-id","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/conversewith{"system": false, "messages": 1, "inferenceConfig": {"maxTokens": 200}}responses_gpt56_model_id
curl -sS -X POST http://127.0.0.1:36349/v1/responses -d '{"model":"gpt-5.6-sol-model-id","input":"Say hi in three words."}'POST /openai/v1/responseswith{"model": "global.openai.gpt-5.6-sol", "messages": 0}responses_gpt56_none_temperature
curl -sS -X POST http://127.0.0.1:36349/v1/responses -d '{"model":"gpt-5.6-sol","input":"Say hi in three words.","temperature":0.2,"reasoning":{"effort":"none"}}'chat_gpt56_runtime_endpoint
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-runtime-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt56_converse_runtime_endpoint
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-converse-runtime-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt56_api_base_only_dead_port
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-env-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'chat_overlong_gpt_version
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-overlong-version","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/openai.gpt-77777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt56_model_id
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-model-id","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}messages_gpt56_model_id
curl -sS -X POST http://127.0.0.1:36349/v1/messages -d '{"model":"gpt-5.6-sol-model-id","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/conversewith{"system": false, "messages": 1, "inferenceConfig": {"maxTokens": 200}}responses_gpt56_model_id
curl -sS -X POST http://127.0.0.1:36349/v1/responses -d '{"model":"gpt-5.6-sol-model-id","input":"Say hi in three words."}'POST /openai/v1/responseswith{"model": "global.openai.gpt-5.6-sol", "messages": 0}chat_gpt56_sigv4_blank_key
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-sigv4-blank-key","messages":[{"role":"user","content":"Say hi in three words."}]}'The three
env_cases ran withAWS_BEDROCK_RUNTIME_ENDPOINTset to the recorder in the proxy's environment and noaws_bedrock_runtime_endpointin the deployment (gpt-5.6-sol-env-endpoint, whoseapi_baseis a dead port, and theconverse/pin)env_chat_gpt56_api_base_dead_port_env_endpoint
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-env-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}env_chat_gpt56_converse_prefix
curl -sS -X POST http://127.0.0.1:36349/v1/chat/completions -d '{"model":"gpt-5.6-sol-converse","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}env_messages_gpt56_api_base_dead_port_env_endpoint
curl -sS -X POST http://127.0.0.1:36349/v1/messages -d '{"model":"gpt-5.6-sol-env-endpoint","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {"maxTokens": 200}}Before fix 2 (6181d76773)
The four cases below also ran at 6181d76773, the previous tip dae29c1 merged with main d7e6154, on port 41217 with the same rig. This is what the audit caught: a
model_idoverride reached the native endpoint and a blankapi_keybecame an empty bearer. 578f26e fixed both (design decisions 20 and 21), and the After leg below shows them at 200 through Converse and SigV4chat_gpt56_model_id
curl -sS -X POST http://127.0.0.1:41217/v1/chat/completions -d '{"model":"gpt-5.6-sol-model-id","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false}messages_gpt56_model_id
curl -sS -X POST http://127.0.0.1:41217/v1/messages -d '{"model":"gpt-5.6-sol-model-id","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "max_completion_tokens": 200}responses_gpt56_model_id
curl -sS -X POST http://127.0.0.1:41217/v1/responses -d '{"model":"gpt-5.6-sol-model-id","input":"Say hi in three words."}'POST /openai/v1/responseswith{"model": "global.openai.gpt-5.6-sol", "messages": 0}chat_gpt56_sigv4_blank_key
curl -sS -X POST http://127.0.0.1:41217/v1/chat/completions -d '{"model":"gpt-5.6-sol-sigv4-blank-key","messages":[{"role":"user","content":"Say hi in three words."}]}'After (6daf0b7)
chat_gpt56_basic
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false}chat_gpt6_basic
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-6-sol", "messages": 1, "stream": false}chat_gpt61_basic
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6.1-sol","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-6.1-sol", "messages": 1, "stream": false}chat_gpt6_json_schema
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Name the capital of France."}],"response_format":{"type":"json_schema","json_schema":{"name":"capital","strict":true,"schema":{"type":"object","properties":{"city":{"type":"string"},"country":{"type":"string"}},"required":["city","country"],"additionalProperties":false}}}}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-6-sol", "messages": 1, "response_format": "json_schema", "stream": false}chat_gpt56_tools_reasoning_none
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"What is the weather in Paris right now?"}],"tools":'"$TOOLS"',"reasoning_effort":"none"}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "none", "tools": ["get_weather"], "stream": false}chat_gpt6_tools_reasoning_low
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"What is the weather in Paris right now?"}],"tools":'"$TOOLS"',"reasoning_effort":"low"}'POST /model/global.openai.gpt-6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "additionalModelRequestFields": {"reasoning": {"effort": "low"}}, "toolConfig.tools": ["get_weather"]}chat_gpt61_tools
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6.1-sol","messages":[{"role":"user","content":"What is the weather in Paris right now?"}],"tools":'"$TOOLS"'}'POST /model/global.openai.gpt-6.1-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "toolConfig.tools": ["get_weather"]}chat_gpt56_guardrail
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"guardrailConfig":'"$GUARD"'}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}, "guardrailConfig": {"guardrailIdentifier": "49e7k3qphvzq", "guardrailVersion": "DRAFT", "trace": "enabled"}}chat_gpt56_app_profile
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-profile","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt56_model_id
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-model-id","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt56_sigv4_blank_key
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-sigv4-blank-key","messages":[{"role":"user","content":"Say hi in three words."}]}'chat_gpt56_converse_prefix
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-converse","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt55_basic
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.5","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.5/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gptoss_basic
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-oss-20b","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/openai.gpt-oss-20b-1%3A0/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gptoss_native_prefix
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-oss-20b-native","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /openai/v1/chat/completionswith{"model": "openai.gpt-oss-20b-1:0", "messages": 1, "stream": false}chat_grok_basic
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"grok-4.6","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/us.xai.grok-4.6/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_grok_native_prefix
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"grok-4.6-native","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /openai/v1/chat/completionswith{"model": "us.xai.grok-4.6", "messages": 1, "stream": false}chat_gpt56_reasoning_low
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"What is 17*23? Answer with the number only."}],"reasoning_effort":"low"}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "low", "stream": false}chat_gpt56_temperature
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9}'chat_gpt56_temperature_drop_params
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false}chat_gpt56_penalty_drop_params
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"frequency_penalty":0.5,"presence_penalty":0.5}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false}chat_gpt6_temperature
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2}'chat_gpt6_temperature_drop_params
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-6-sol", "messages": 1, "stream": false}chat_gpt6_penalty_drop_params
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol-drop","messages":[{"role":"user","content":"Say hi in three words."}],"frequency_penalty":0.5,"presence_penalty":0.5}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-6-sol", "messages": 1, "stream": false}chat_gpt56_low_temperature
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"reasoning_effort":"low"}'chat_gpt56_none_temperature
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"top_p":0.9,"reasoning_effort":"none"}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "none", "stream": false, "temperature": 0.2, "top_p": 0.9};POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "none", "stream": false, "temperature": 0.2, "top_p": 0.9};POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "none", "stream": false, "temperature": 0.2, "top_p": 0.9}chat_gpt56_none_penalties_logprobs
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"frequency_penalty":0.5,"presence_penalty":0.5,"logprobs":true,"top_logprobs":2,"reasoning_effort":"none"}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "none", "stream": false, "frequency_penalty": 0.5, "presence_penalty": 0.5}chat_gpt6_none_temperature
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"reasoning_effort":"none"}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-6-sol", "messages": 1, "reasoning_effort": "none", "stream": false, "temperature": 0.2}chat_gpt56_none_temperature_stream
curl -sS -N -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"temperature":0.2,"reasoning_effort":"none","stream":true,"stream_options":{"include_usage":true}}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "reasoning_effort": "none", "stream": true, "stream_options": {"include_usage": true}, "temperature": 0.2}chat_gpt56_stream
curl -sS -N -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"Say hi in three words."}],"stream":true,"stream_options":{"include_usage":true}}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": true, "stream_options": {"include_usage": true}}chat_gpt56_image_fixture_png
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":[{"type":"text","text":"'"$Q_COLORS"'"},{"type":"image_url","image_url":{"url":"'"$FIX"'"}}]}]}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false, "image_parts": 1, "image_url_prefix": "data:image/png;base64,"}messages_gpt56
curl -sS -X POST http://127.0.0.1:20753/v1/messages -d '{"model":"gpt-5.6-sol","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "max_completion_tokens": 200}responses_gpt56
curl -sS -X POST http://127.0.0.1:20753/v1/responses -d '{"model":"gpt-5.6-sol","input":"Say hi in three words."}'POST /openai/v1/responseswith{"model": "global.openai.gpt-5.6-sol", "messages": 0};POST /openai/v1/responseswith{"model": "global.openai.gpt-5.6-sol", "messages": 0};POST /openai/v1/responseswith{"model": "global.openai.gpt-5.6-sol", "messages": 0}messages_gpt56_model_id
curl -sS -X POST http://127.0.0.1:20753/v1/messages -d '{"model":"gpt-5.6-sol-model-id","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/arn%3Aaws%3Abedrock%3Aus-west-2%3A888602223428%3Aapplication-inference-profile%2F6mv9jwwiiudo/conversewith{"system": false, "messages": 1, "inferenceConfig": {"maxTokens": 200}}responses_gpt56_model_id
curl -sS -X POST http://127.0.0.1:20753/v1/responses -d '{"model":"gpt-5.6-sol-model-id","input":"Say hi in three words."}'POST /openai/v1/responseswith{"model": "global.openai.gpt-5.6-sol", "messages": 0}responses_gpt56_none_temperature
curl -sS -X POST http://127.0.0.1:20753/v1/responses -d '{"model":"gpt-5.6-sol","input":"Say hi in three words.","temperature":0.2,"reasoning":{"effort":"none"}}'chat_gpt56_runtime_endpoint
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-runtime-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false}chat_gpt56_converse_runtime_endpoint
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-converse-runtime-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}chat_gpt56_api_base_only_dead_port
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-env-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'chat_overlong_gpt_version
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-overlong-version","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/openai.gpt-77777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777777/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}The three
env_cases ran withAWS_BEDROCK_RUNTIME_ENDPOINTset to the recorder in the proxy's environment and noaws_bedrock_runtime_endpointin the deployment (gpt-5.6-sol-env-endpoint, whoseapi_baseis a dead port, and theconverse/pin)env_chat_gpt56_api_base_dead_port_env_endpoint
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-env-endpoint","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "stream": false}env_chat_gpt56_converse_prefix
curl -sS -X POST http://127.0.0.1:20753/v1/chat/completions -d '{"model":"gpt-5.6-sol-converse","messages":[{"role":"user","content":"Say hi in three words."}]}'POST /model/global.openai.gpt-5.6-sol/conversewith{"system": false, "messages": 1, "inferenceConfig": {}}env_messages_gpt56_api_base_dead_port_env_endpoint
curl -sS -X POST http://127.0.0.1:20753/v1/messages -d '{"model":"gpt-5.6-sol-env-endpoint","max_tokens":200,"messages":[{"role":"user","content":"Say hi in three words."}]}'POST /openai/v1/chat/completionswith{"model": "global.openai.gpt-5.6-sol", "messages": 1, "max_completion_tokens": 200}/live-pr-risk cells
The 30 cells below ran on both legs right after the cases, same proxies and recorders (the last three, the
/v1/responsesand/v1/messagesnon-string efforts the graph walk of the new commit added, on a re-boot of the same config and recorder per leg), and once more on the merged tree with identical outcomes but one (a retried first attempt aside):chat_gpt56_runtime_endpoint_streamanswered 503 on the After leg because AWS accepted the stream (200 at the recorder) and 12.8 s later sent its own error event (The server had an error while processing your request. Sorry about that!) before any token, which the stream path surfaces as a 503ServiceUnavailableError(the B2 class among the Low caveats); the identical request answered 200 stream natively on the merged tree and in the earlier run at the previous tip. Status and the AWS path the recorder saw:model_group_info_supported_paramschat_gpt6_guardrail_temperaturechat_gpt6drop_guardrail_temperaturechat_gpt6_guardrail_none_temperaturechat_gpt56_profile_temperaturechat_gpt56_converse_prefix_temperaturecost_chat_gpt56cost_messages_gpt56cost_responses_gpt56cost_chat_gpt6responses_gptoss_nativemessages_gpt6_temperaturechat_gpt6_stream_usagehostile_reasoning_effort_inthostile_reasoning_effort_listhostile_reasoning_effort_int_drop_paramshostile_reasoning_effort_list_drop_paramshostile_reasoning_effort_int_request_drop_paramshostile_reasoning_effort_emptyhostile_reasoning_effort_5kbhostile_temperature_stringhostile_temperature_string_nonemessages_gpt56_runtime_endpointresponses_gpt56_runtime_endpointchat_gpt56_runtime_endpoint_streammessages_overlong_gpt_versionresponses_overlong_gpt_versionresponses_hostile_reasoning_effort_intresponses_hostile_reasoning_effort_int_drop_paramsmessages_hostile_reasoning_effort_intObservations from the run
chat_completions/gpt-oss usage carries noreasoning_tokens; PR causes/v1/messagesdrops Converse's emptyredacted_thinkingblock; PR causes/v1/responseswithreasoning.effort: noneplustemperature400s both legs; unchangedreasoning_effort: lowreportsreasoning_tokens: 0both legs; unchangedcontent: "", native omits it; unchangedreasoning_effort400s before any AWS call; PR causesdrop_paramsthe non-string effort is dropped, 200; unchanged/v1/responsesnon-stringreasoning.effortgets AWS's 400 both legs; unchangedapi_baseanswers 503, not 500; PR causes/v1/responsesignoresmodel_idand calls the base model; unchangedspend: 0.0; unchanged/audit round 2 of 3
Verdict: PASS. Every inventory row ran at the head and matched the outcome written for it before the run; all 68 cells below are PASS or a FAIL row whose decision this run recorded (3 rows, A20a, A20b, B2, each a behavior change named in its row and shipped by this run), and the one decision round 1 left to the author, an int or list
reasoning_effortanswering AWS's 400 where Converse dropped it and answered 200, is taken in 6daf0b7: litellm answers its own 400 before any AWS call unlessdrop_paramsis set, and then drops the value the way Converse did (A24 and the live rows L8 to L8c). Unverified rows: noneHashes: the product code under every row is 6daf0b7 (the native route, the non-string
reasoning_effortrule, and the tests); e518e0e changes no product code: it pinsnum_retries: 0on the deployments the A25, B1, B2, and B3 cells use and moves the B1, B2, and B3 bursts onto a proxy the chaos module owns (2 workers, the same shape as B4) whose three deployments come from its config instead of/model/new; merge base c168199. Round 2 ran the whole selection at 6daf0b7 (one proxy per leg from its own worktree and virtualenv,tests/integration/_support/proxyontests/integration/proxy_config.yaml, 2 workers,--use_prisma_db_push, head on 24778 with databaselit8684r2_audit_head, base on 32065 withlit8684r2_audit_base, one Redis on 30859, the scripted OpenAI upstream on 24827 for the unrelated deployments, andtests/integration/_support/bedrock_runtime_peer.pyas the Bedrock runtime peer each test module owns; deployments are created per test through/model/new):python tests/integration/run.py providers <the six files> --seed 4106601 --order-seed 0, baseaudit_r2_base1(32 failed, 23 passed), headaudit_r2_head1(4 failed, 51 passed in 309.7 s) andaudit_r2_head2(4 failed, 51 passed in 280.23 s), identical selections, no skips, no retries. The four head reds were root-caused to the audit's own tests and the rig, not the product: A25 [429], A25 [500], and B2 assert one wire attempt, but the checked-in config leaves the router's default two retries on with cooldowns disabled, so the scripted 429 and 500 and the killed peer were retried twice (three wire hits, a connect error in place of the disconnect text), where round 1 had passed them with retries set to 0 on the rig out of band (A25 [401] was never retried on either leg); B4, and A20b on the base, timed out at the owned proxy's 70 s readiness deadline intests/integration/_support/process.pywhile the box carried other rigs' proxies. A first rerun of those two files with only the retries pinned off (runssuperseded_changed_head1, 23 passed, andsuperseded_changed_head2, B2 red once:State did not converge: 8) exposed a second rig-side cause the retries had masked:/model/newlands on one worker over the keep-alive client and the sibling learns of the fresh deployment through the registry read-through inlitellm/proxy/common_utils/registry_read_through.py, where each queued miss on the same key consumes one unit of the resync budget of 20 per 5 s, so B1's 36-call burst spent that budget and 4 of B2's 12 calls on the sibling answered 400Invalid model namebefore any wire call (4resync budget exhaustedwarnings in the head proxy log); it is pre-existing, outside this PR's diff, and ticketed separately, and a chaos burst on a deployment created seconds earlier is not what the cells measure, so e518e0e gives the three bursts config deployments every worker holds at boot. Per the audit's rerun rule the two files those cells live in (test_bedrock_runtime_chat_completions_sad_wire.pyandtest_bedrock_runtime_chat_completions_chaos.py, 23 cells) then re-ran whole on a fresh rig from this branch (same shape: head on 24432 withdrive40775_audit_head, base on 45636 withdrive40775_audit_base, Redis on 36275, upstream on 21619, one proxy up at a time, headroom confirmed first): headchanged_head1(23 passed in 217.99 s) andchanged_head2(23 passed in 210.76 s) with identical collected and passed selections, no skips, no retries, and basechanged_base1(8 failed, 15 passed, every red at a cell the inventory marks red for the merge base). Every other row keeps its 6daf0b7 result, the product code under it being the code at the head; the After column of each table names the commit its rows last ran at. Live rows ran once per leg fromlive/6daf0b7136/run_live.pyagainst real AWS Bedrock inus-west-2through the round-2 proxies, with the artifacts (proxy response,llm_provider-x-amzn-requestid, spend row) underlive/6daf0b7136/<leg>/<row>.jsontests/integration/providers/test_bedrock_runtime_chat_completions_wire.pytest_openai_sdk_reasoning_request_is_served_by_native_chat_completionstest_async_openai_sdk_stream_keeps_the_upstream_id_and_usagetest_temperature_is_forwarded_natively_when_reasoning_is_offtemperatureforwardedtest_temperature_while_reasoning_is_refused_before_any_wire_requesttest_drop_params_deployment_drops_temperature_while_reasoningtemperaturedroppedtest_guardrail_config_keeps_converseguardrailConfigin bodytest_converse_prefix_pins_the_model_to_conversetest_application_inference_profile_arn_keeps_conversetest_model_id_application_inference_profile_keeps_converse_at_the_profile_urltest_stop_sequences_keep_conversestopSequencestest_json_object_response_format_keeps_conversetest_json_schema_response_format_is_forwarded_nativelyresponse_formatverbatimtest_tools_while_reasoning_keep_conversetoolConfigtest_tools_with_reasoning_off_are_forwarded_nativelytoolsforwardedtest_empty_tools_list_while_reasoning_stays_nativetest_chat_completions_prefix_splits_gpt_oss_reasoning_tagchat_completions/prefixreasoning_contentsplit fromcontenttest_chat_completions_prefix_splits_gpt_oss_reasoning_tag_across_stream_deltastest_region_path_model_is_served_natively_without_the_regiontest_sigv4_deployment_signs_the_native_requestAWS4-HMAC-SHA256test_blank_api_key_on_a_sigv4_deployment_is_signed_not_sent_as_an_empty_bearertest_runtime_endpoint_without_api_base_is_used_nativelytest_runtime_endpoint_wins_over_an_unrelated_api_baseapi_base(fixed in dae29c1)test_api_base_already_naming_the_native_path_is_not_doubled[/openai/v1]/openai/v1/chat/completionsoncetest_api_base_already_naming_the_native_path_is_not_doubled[/openai/v1/chat/completions]tests/integration/providers/test_bedrock_runtime_chat_completions_sad_wire.pytest_remote_image_url_on_the_shared_proxy_is_rejected_before_any_fetchtest_allowlisted_remote_image_is_inlined_for_the_native_routedata:image/png;base64, oneGET /image.png; missing image 400test_response_cache_twin_serves_the_second_request_without_a_second_wire_calltest_model_group_info_lists_the_native_supported_paramsreasoning_effortandlogprobslisted,nabsenttest_thirty_thousand_digit_version_is_classified_quickly_and_served_by_conversetest_bad_key_on_the_long_version_model_is_refused_before_any_routetest_invalid_reasoning_effort_reaches_the_peer_and_its_400_reaches_the_caller[empty]test_invalid_reasoning_effort_reaches_the_peer_and_its_400_reaches_the_caller[five_kb]test_non_string_reasoning_effort_is_refused_before_any_wire_request[int]drop_paramsdrop_params)test_non_string_reasoning_effort_is_refused_before_any_wire_request[list]drop_paramstest_drop_params_deployment_drops_a_non_string_reasoning_effort[int]reasoning_effortabsent from the wire bodytest_drop_params_deployment_drops_a_non_string_reasoning_effort[list]reasoning_effortabsent from the wire bodytest_duplicated_reasoning_effort_key_lets_the_last_value_wintest_string_temperature_is_refused_before_any_wire_requesttest_peer_error_status_reaches_the_caller_and_unrelated_deployments_keep_serving[401]test_peer_error_status_reaches_the_caller_and_unrelated_deployments_keep_serving[429]test_peer_error_status_reaches_the_caller_and_unrelated_deployments_keep_serving[500]test_absent_reasoning_effort_is_forwarded_as_absent_on_every_repeat[null]test_absent_reasoning_effort_is_forwarded_as_absent_on_every_repeat[missing]tests/integration/providers/test_bedrock_gpt_responses_native_wire.pyandtests/integration/providers/test_bedrock_passthrough_stream_wire.py(existing)test_openai_sdk_responses_request_is_served_by_the_native_responses_route/openai/v1/responsestest_async_openai_sdk_responses_stream_is_served_by_the_native_responses_route/openai/v1/responsestest_bedrock_passthrough_converse_stream_response_carries_event_stream_content_typetests/integration/providers/test_bedrock_runtime_chat_completions_chaos.pytest_burst_across_every_endpoint_lands_each_response_id_oncetest_peer_killed_mid_burst_fails_only_the_held_calls_and_a_restarted_peer_serves_againtest_slow_peer_streams_are_forwarded_once_and_terminated[DONE]test_worker_sigkill_mid_burst_leaves_the_sibling_servingOwned proxy required forced cleanup: uvicorn had replaced the killed worker moments before teardown and the replacement was still starting, so it could not honor SIGTERM inside the harness's 30 s, and 6befa50 waits for the replacement worker's startup to complete before leaving the proxy; chaos file reruns at 6befa50: merged_head4 4 passed in 133 s and merged_head5 4 passed in 120 s, no skips, no retries, against a head proxy booted from 6befa50 with two uvicorn workers)tests/integration/messages_endpoint/providers/bedrock/test_bedrock_messages_gpt_chat_completions_wire.pytest_anthropic_sdk_thinking_budget_reaches_native_chat_completions_as_reasoning_effortreasoning.effortmediumreasoning_effort: medium, nothinkingtest_anthropic_sdk_stream_with_thinking_budget_is_served_by_native_chat_completionstest_raw_thinking_summary_reaches_native_chat_completions_as_the_plain_effortreasoning_effort: medium, nosummarytest_async_anthropic_sdk_disabled_thinking_reaches_native_chat_completions_as_effort_nonereasoning_effort: nonetest_identical_messages_requests_reach_the_peer_once_and_log_a_cache_hit_rowLive rows, real AWS Bedrock
us-west-2(never committed), base run28839dc4at 02:58Z and head run06c9bfb2at 03:11Z on 2026-10-02reasoning_effort: highreasoning_tokens21include_usage/v1/responses/v1/messages, thinking budget 2048guardrailConfigreasoning_effort: none+temperaturereasoning_effort: 3reasoning_effort: 3,drop_paramsdeploymentreasoning_effort: ["high"],drop_paramsin the bodymodel_idARN/v1/messages,model_idARNaws_profile_name,api_key: ""Pass-through reachability: the diff's byte and SSE code (
ReasoningTagSplitter,BedrockRuntimeChatCompletionsStreamingHandler) is imported only bylitellm/llms/bedrock/chat/chat_completions/transformation.pyand its unit test, so A28 is the one pass-through route whose bytes it could touch. The A15 cell sendsstream_optionsfrom the client because a worker that has not yet loaded a deployment created seconds earlier omitsinclude_usage(pre-existing, Low caveat). B1 matches a non-stream/v1/responsesspend row by the upstream id inside its payload because that row can carry the pre-encryption id (pre-existing, Low caveat)Kept from the 2026-09-26 run
The ten non-Bedrock image cases below ran on 2026-09-26 at the merge base 1f77fa6 and the tip 4900a1a, on the same two-worker rig in
us-east-1with one recorder per provider (Anthropic, Gemini, OpenAI, Vertex AI atvertex_location: global, and Bedrock invoke); the only change to the shared image fetch since then (3e3b9cd) hoists the URL comprehension into a helper, so the cases still describe the tip. Their deployments were:The fixtures live under
$ASSETS(https://github.com/mateo-berri/qa-screenshots/releases/download/assets-2;$ASSETS_HTTPis the same path overhttp://), which GitHub serves asapplication/octet-streambehind a 302, the case the type inference is for.pr40775-fixture.pngandpr40775-pngblobare the same 96x64 PNG, red on the left and blue on the right, one with the.pngextension and one without;pr40775-notimageis 160 bytes of text;pr40775-doc.pdfis a one-page PDF reading "The secret word is PELICAN".$Q_COLORSis as above and$Q_PDFis "What is the secret word in this document? Answer with the word only."Before (1f77fa6)
chat_anthropic_http_fixture_png
curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"claude-anthropic","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS_HTTP/pr40775-fixture.png"}}]}],"max_tokens":200}'chat_anthropic_http_notimage
curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"claude-anthropic","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS_HTTP/pr40775-notimage"}}]}],"max_tokens":200}'chat_anthropic_http_pngblob
curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"claude-anthropic","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS_HTTP/pr40775-pngblob"}}]}],"max_tokens":200}'chat_gemini_file_pdf
curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"gemini-flash","messages":[{"role":"user","content":[{"type":"text","text":"$Q_PDF"},{"type":"file","file":{"file_id":"$ASSETS/pr40775-doc.pdf"}}]}],"max_tokens":200}'chat_gemini_fixture_png
curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"gemini-flash","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-fixture.png"}}]}],"max_tokens":200}'chat_gemini_pngblob
curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"gemini-flash","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-pngblob"}}]}],"max_tokens":200}'chat_invoke_claude_fixture_png
curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"claude-invoke","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-fixture.png"}}]}],"max_tokens":200}'chat_openai_file_pdf
curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"openai-gpt","messages":[{"role":"user","content":[{"type":"text","text":"$Q_PDF"},{"type":"file","file":{"file_id":"$ASSETS/pr40775-doc.pdf"}}]}],"max_completion_tokens":200}'chat_vertex_claude_fixture_png
curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"vertex-claude","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-fixture.png"}}]}],"max_tokens":200}'chat_vertex_gemini_pngblob
curl -s http://127.0.0.1:36738/v1/chat/completions -d '{"model":"vertex-gemini","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-pngblob"}}]}],"max_tokens":400}'After (4900a1a)
chat_anthropic_http_fixture_png
curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"claude-anthropic","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS_HTTP/pr40775-fixture.png"}}]}],"max_tokens":200}'chat_anthropic_http_notimage
curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"claude-anthropic","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS_HTTP/pr40775-notimage"}}]}],"max_tokens":200}'chat_anthropic_http_pngblob
curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"claude-anthropic","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS_HTTP/pr40775-pngblob"}}]}],"max_tokens":200}'chat_gemini_file_pdf
curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"gemini-flash","messages":[{"role":"user","content":[{"type":"text","text":"$Q_PDF"},{"type":"file","file":{"file_id":"$ASSETS/pr40775-doc.pdf"}}]}],"max_tokens":200}'chat_gemini_fixture_png
curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"gemini-flash","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-fixture.png"}}]}],"max_tokens":200}'chat_gemini_pngblob
curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"gemini-flash","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-pngblob"}}]}],"max_tokens":200}'chat_invoke_claude_fixture_png
curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"claude-invoke","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-fixture.png"}}]}],"max_tokens":200}'chat_openai_file_pdf
curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"openai-gpt","messages":[{"role":"user","content":[{"type":"text","text":"$Q_PDF"},{"type":"file","file":{"file_id":"$ASSETS/pr40775-doc.pdf"}}]}],"max_completion_tokens":200}'chat_vertex_claude_fixture_png
curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"vertex-claude","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-fixture.png"}}]}],"max_tokens":200}'chat_vertex_gemini_pngblob
curl -s http://127.0.0.1:32068/v1/chat/completions -d '{"model":"vertex-gemini","messages":[{"role":"user","content":[{"type":"text","text":"$Q_COLORS"},{"type":"image_url","image_url":{"url":"$ASSETS/pr40775-pngblob"}}]}],"max_tokens":400}'Type
🆕 New Feature
Design decisions
bedrock/<id>goes to runtime Chat Completions only when the id is an OpenAI GPT model at version 5.6 or newer (openai.gpt-5.6-*,openai.gpt-6-*,openai.gpt-6.1-*, read off the id, never gpt-oss) and its row lists/v1/chat/completionsinsupported_endpoints(the key feat(bedrock): serve the OpenAI models on bedrock-runtime's native Responses API (internal copy of #38489) #42767 already reads for the native/v1/responsesroute). gpt-oss and Grok rows carry the key too but stay on Converse by default;bedrock/chat_completions/<id>opts a deployment in, andbedrock/converse/<id>pins one to Converse. Application inference profile ARNs take Converse ahead of that check, as before this PR. A region path (bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0) is routing, not a different model, so the route is looked up on the id after the path and the path's region picks the endpoint and the SigV4 scope (an explicitaws_region_namestill wins, as on Converse). A vendor-prefix rule (openai.,xai.,zai.) was raised in review, and AWS's API compatibility table does list Chat Completions on every OpenAI, xAI, and Z.AI row, but the per-model key stays for three reasons. The two companion flags differ withinopenai.already (GPT 5.6+ refuses tools with reasoning, gpt-oss ignores aresponse_formatschema), so the cost map keeps per-model keys under a prefix rule too. A prefix rule would move every gpt-5.4, gpt-5.5, gpt-oss-safeguard, and GLM deployment to the native route at upgrade with none of their refused params, reasoning shape, tools, orresponse_formatchecked there. And AWS's table is per model, not per vendor (Kimi K2 Thinking and DeepSeek-R1 lack Chat Completions while their siblings have it), so a vendor key would be wrong the day it is extended past these three vendors. The 5.6 floor comes from the customer ask (feat(bedrock): route OpenAI models to bedrock-runtime's native Chat Completions API #43264 wanted GPT 5.6+ served natively with an automatic fallback and no new flag): gpt-5.4 and gpt-5.5 keep Converse even if a row of theirs gains the key, and a GPT 5.6+ id with no row, or a row without the key, keeps Converse too. The two capability flags stay, since they say what the native endpoint serves for a model (tools with reasoning, aresponse_formatschema), not which endpoints AWS exposes for it; a daily automation syncingsupported_endpointsfrom AWS's compatibility doc is a follow-up outside this PRguardrailConfiggoes to Converse. The alternative was mapping it onto the native endpoint'sX-Amzn-Bedrock-GuardrailIdentifier/X-Amzn-Bedrock-GuardrailVersionheaders, which keeps the native route but adds exactly the kind of translation this PR removes. The trade-off is that Bedrock Guardrails requests keep paying the Converse translationperformanceConfig,serviceTier,requestMetadata,outputConfig), the Anthropic-stylethinkingblock on/v1/chat/completions(only Converse forwards it, asadditionalModelRequestFields, so Grok'sredacted_thinkingblock still comes back there;/v1/messagesmapsthinkingtoreasoning_effortbefore the route is picked, so it goes native), and operator-owned request metadata (the proxy's Bedrock request-metadata opt-in), since the native endpoint has no field for them. Application inference profile ARNs took Converse before this PR and still do, checked ahead of the flag"reasoning_effort": "none", because AWS rejects tools plus reasoning on the native endpoint (checked live on 2026-09-30 onglobal.openai.gpt-5.6-solandglobal.openai.gpt-6-sol). The cost map states the capability positively,supports_bedrock_runtime_chat_completions_tools_with_reasoning: trueon gpt-oss and Grok, and a model without it sends a tools request with any other effort, or with no effort at all (AWS applies its default effort there), to Converse. Absent flag means the safer route, so a newly listed model never hits AWS's 400 by default. gpt-6.1 has nononeeffort on AWS, so its tools requests always take Converse. Legacyfunctionsare gated the same way, since AWS rejects them with the same 400. Serving tools plus reasoning natively through the Responses API is deferred to LIT-8970max_tokensis sent asmax_completion_tokens: every model on this surface takes it and the GPT 5.6+ family rejectsmax_tokens. An explicitmax_completion_tokenswins when both are set. Rejected: renaming only for the GPT family and forwardingmax_tokensas sent on gpt-oss and Grok, a second per-family table for a rename every model on this endpoint acceptsnis dropped from the native endpoint's supported params, since AWS serves one choice there, and so are the params each family refuses outright on that endpoint, checked live: both penalties on Grok (a 503 for any non-zero value) andlogit_biason gpt-oss (a 400). On GPT 5.6+ six params are tied to reasoning: AWS takestemperature,top_p,frequency_penalty,presence_penalty,logprobs, andtop_logprobsnatively underreasoning_effort: "none"and 400s them at every other effort (checked live on 2026-09-30 on gpt-5.6 and gpt-6). So undernonethe native config forwards all six, and otherwise it answers a 400 naming them and thenoneway out, or drops them underdrop_params. The merge base answered litellm's ownbedrock does not support parameters: ['temperature', 'top_p']400 for GPT-5.6 at every effort (its rows carrysupports_sampling_params: false, which Converse reads) and forwarded them for GPT-6, which AWS 400'd even underdrop_params(its rows had no flag). The PR addssupports_sampling_params: falseto the gpt-6 and gpt-6.1 rows, so the Converse fallback and aconverse/pin refuse or drop the two before the call the way they already did for GPT-5.6. gpt-6.1 has nononeeffort, so its six params always 400 or drop. The refused set is one table keyed by model family in the native config. Rejected: a per-model list in the cost map (17 entries for a per-family fact) and composing the OpenAI GPT-5 and xAI configs' lists (they rewrite the body through their own mapping and keeppresence_penalty, which AWS's Grok 503s)<reasoning>...</reasoning>prefix on the native endpoint. One splitter serves both the streamed and the non-streamed response, so the two split identically, and the reasoning lands inreasoning_contentonly. Fillingthinking_blocksandprovider_specific_fields.reasoningContentBlocksthe way Converse did would have kept three copies of the same text{"type": "json_schema"}response_format(a pydantic model is converted to that shape) goes native only on models flaggedsupports_bedrock_runtime_chat_completions_response_format: GPT 5.6+ and Grok, which enforce the schema on AWS's endpoint (checked live, both answer the schema's shape, and Grok honors it on both routes once the reasoning fits the token budget). Every{"type": "json_object"}form keeps Converse's handling on every model, the LiteLLM-specificresponse_schemavariant included (Converse turns that one into a forcedjson_tool_calland answers the schema as before this PR; the schema-less one is ignored and answers free text as before this PR), because AWS's native endpoint validates by type and answers 400'messages' must contain the word 'json'forjson_objectunless the prompt mentions json, which the A/B caught for theresponse_schemavariant at an earlier tip; rejected: sendingjson_objectnative only when a message mentions json, which needs the messages plumbed into the route predicate and mirrors AWS's contract heuristically. GPT-OSS acceptsresponse_formaton the native endpoint but answers with free text, so it keeps Converse's emulation (a forcedjson_tool_calltool) and the structured answer it gave before this PR.{"type": "text"}is a no-op on both routes and stays native. The flag is positive so an unflagged model gets the safer routeadditional_drop_params, by one helper that both the param mapping and the dispatch call. The first version of this PR derived it again at dispatch from the mapped params, which was wrong for the GPT-OSSresponse_formatfallback: Converse's mapping turnsresponse_formatinto thejson_tool_calltool and drops the field, so a dispatch reading the mapped params picked the native handler and handed it Converse-shaped params. A test pins the dispatch to the Converse URL and body for that casesupports_prefix. The cost map guard runs the base branch's schema generator against the PR's cost map and auto-classifiessupports_*keys as booleans, so any other name fails the required check until a generator-only PR lands on main first. The prefix also keeps the keys out of the generator's hand-maintained key list. They do not leak intoget_model_info,/model/info, the public model hub features, or the UI filters, which all read from a typed model info with explicit fields (verified at runtime)reasoning_effort: "none"with a 400 (it takeslow,medium,high, andxhigh), where Converse never listedreasoning_effortfor xai models and dropped it underdrop_params, sononeanswered 200 with AWS's default effort. The native config dropsnoneon Grok before the call (one table of refused effort values keyed by model family, next to the refused params table), so the 200 stays and AWS applies its default effort as before, while the other efforts, which Converse silently dropped, are now forwarded on an opted-in Grok deployment. This also keeps/v1/messageswiththinking: {"type": "disabled"}working, since the Anthropic adapter maps that toreasoning_effort: "none". Rejected: routingnoneto Converse (a whole second code path for a value AWS's default already gives) and raising (a regression from main for a value OpenAI clients send by default). gpt-oss refusesnoneon both routes and GPT 5.6+ takes it on both, so neither changesadditionalModelRequestFieldsandtop_kjoin the Converse-only request keys, so a request carrying either keeps main's Converse handling to the byte (the 2026-09-26 merge base leg showed Converse nestingadditionalModelRequestFieldsinside its own block and emitting notopKfor these models, both of which are Converse's behavior on main, not this PR's). Rejected: letting them through to the native endpoint, which answers 200 and ignores an unknown key, or leavingtop_kto the native drop-or-raise handling, which turns main's 200 into an error for a caller withoutdrop_params. Any other unknown body key reaches the native endpoint as sent, as on every OpenAI-compatible providerus-west-2deployments behind one forwarding recorder so the AWS path each case hit is read from the wire instead of inferred; the kept image cases ran the same way on 2026-09-26 inus-east-1stopjoins the Converse-only request keys. Converse forwards it asstopSequences, which AWS refuses with a 400 on all three families, and that is what main does. Natively gpt-oss and Grok acceptstopbut apply it to the hidden reasoning stream too, sostop: ["3"]on "count from 1 to 5" answered 200 withcontent: ""(gpt-oss) orcontent: null(Grok) andfinish_reason: "stop", a silent empty answer where main failed loudly, and GPT-5.6 raisedUnsupportedParamsErrorfromget_optional_params(a different error class and layer) oncestopleft its native list. Routingstopto Converse keeps main's exact 400 for every caller, sostopis also back out of the GPT refused table since the native config never sees it. Rejected: servingstopnatively on gpt-oss and Grok (the empty answers above), or refusing it natively on every family (the error class change)http(s)://image_urlparts are downloaded and inlined asdata:URLs before the native call, intransform_requestthrough the sharedinline_remote_mediaand inasync_transform_requestthroughasync_inline_remote_media(the same helpers the Anthropic and Bedrock invoke transforms use, withinline_remote_image_urlsso file parts are left alone). Converse downloaded remote images itself; AWS's native endpoint answers400 Only inline image data URLs and S3 URLs are supportedfor a URL, sous.xai.grok-4.6and the GPT 5.6+ rows (supports_vision: true) would have regressed from 200 to 400 on any vision request.data:ands3://URLs pass through untouched. The syncinline_remote_mediais new inimage_handling.py, a sequential mirror of the async one, since the module only had the async walker. The shared fetch kept a genericContent-Type(application/octet-stream) from the image server as the data URL's type, where Converse'sBedrockImageProcessorinferred the real type from the extension or the magic bytes throughinfer_content_type_from_url_and_content; the fetch now runs that same helper, so the native route getsdata:image/pngfor such a server and the Anthropic and Bedrock invoke transforms that already used the fetch get the same inferenceaws_bedrock_project_idis not forwarded as anOpenAI-Projectheader on the native route. AWS documents projects and that header for Bedrock Mantle only, bedrock-runtime has a default project alone and answered 200 to a bogusOpenAI-Projectvalue on 2026-09-26, the native Responses transform sends none, and Converse ignored the param, so a runtime deployment carrying it behaves the same on both routes (checked on 2026-09-26 with the recorder showing noOpenAI-Projectheader on either leg). The first revision forwarded it the way the Mantle configs do and was dropped, since it sent a header nobody chose to a host that does not read itreasoning_effortvalue that is not a string (an int, a list, an object) answers litellm's own 400 before any call (UnsupportedParamsError, naming the type and thedrop_paramsway out), and is dropped underlitellm.drop_paramsor the request'sdrop_paramsso AWS applies its default effort, which is what Converse did on main. That is litellm's convention for an unsupported or malformed param on a provider route (the o-seriestemperaturecheck, and this PR's own refused-while-reasoning branch), so the native route follows it instead of forwarding the value for AWS's 400 (what a0cef91 did) or always dropping it (which hides a malformed request behind a 200). A string AWS does not know (an empty string, a made-up effort) is still forwarded, since the accepted set is AWS's to define per familymodel_idoverride (the per-deploymentlitellm_params.model_id, used to call an application inference profile ARN while the cost map still reads the base model id) joins the Converse-only request keys, so such a deployment keeps Converse at/model/<ARN>/converseas on main. The native endpoint has no field for a profile ARN: at dae29c1 the request went native with the base model in the URL andmodel_idinside the body, and AWS answered 400Unknown parameter: 'model_id'(the audit's A8b row and thechat_gpt56_model_idcase). Rejected: putting the ARN in the native body'smodelfield, since AWS documents the native endpoint on model ids and nothing shows it taking an ARN therevalidate_environmentresolves the key throughbedrock_bearer_tokenbefore the inherited OpenAI-compatible one writes the header, so a SigV4 deployment withapi_key: ""(what a config template or a UI form leaves behind) is signed instead of sent asAuthorization: Bearerwith an empty token, which AWS 403s (Authorization header requires 'Credential' parameter, the audit's A17b row and thechat_gpt56_sigv4_blank_keycase). Converse already treated an empty key as absent. 4900a1a had removed the earlier override (theOpenAI-Projectheader); this one only resolves the key and sets no other headeropenai\.gpt-(\d{1,3})(?!\d)(?:\.(\d{1,3})(?!\d))?), so a model id carrying thousands of digits in its version is classified in constant time and is not a GPT 5.6+ id: it goes to Converse, which answers AWS's error for the unknown model (the*_overlong_gpt_versioncases and the audit's A23 row, which also times/health/livelinessduring the request). At 708f5a5 the int conversion of such a version raisedValueError(Python's 4300-digit limit). Rejected: raising that limit, since the id is caller-controlled on a shared proxyaws_bedrock_runtime_endpoint(orAWS_BEDROCK_RUNTIME_ENDPOINT) wins overapi_base, so a deployment setting both keeps sending to the one host Converse used (the*_runtime_endpoint*cases and the audit's A18 and A18b rows); at 708f5a5 the native request went toapi_base. Anapi_basealready ending in/openai/v1or in the full/openai/v1/chat/completionspath is not doubled (A19)Caveats (if any)
Medium
reasoning_effort(3,["high"]) on a native-routed GPT 5.6+ deployment answers litellm's 400 with no call (takes reasoning_effort as a string ... not int) where Converse dropped the value and answered 200 at the default effort;drop_params, global or on the deployment, keeps the 200 by dropping it (design decision 19, audit rows A24 and L8)bedrock/deployments ofus./global.openai.gpt-5.6-{sol,terra,luna},openai.gpt-6-{astra,sol,luna}, andopenai.gpt-6.1-solnow hit runtime Chat Completions instead of Converse, so response ids,service_tier, and the reasoning fields change shape as the User Flow shows; gpt-oss and Grok stay on Converse unless the deployment is prefixedchat_completions/temperature,top_p, the penalties,logprobs, ortop_logprobson GPT 5.6+ withoutreasoning_effort: "none"now names the params and thenoneway out, where the merge base answeredbedrock does not support parameters: [...], so a client matching that text sees a new messagebedrock_mantle/provider. It has no runtime Chat Completions IDbedrock/converse/<model>is still Converse"content": nullinstead of""whenmax_tokensruns out inside the reasoning ("finish_reason": "length") on GPT 5.6+ and an opted-in Grok deployment, since AWS's native endpoint returns null there and Converse returned an empty string/v1/messageson an opted-in Grok deployment answers with atextblock alone (an emptycontentlist whenmax_tokensruns out in the reasoning) where Converse also returned aredacted_thinkingblock, since the route mapsthinkingtoreasoning_effortand goes nativeBedrockExceptionmessage changes shape on the native route: Converse's{"message": ...}becomes the native{"error": {"message": ..., "type": "invalid_request_error", "param": ..., "code": ...}}, so a client parsing that text sees the new keysreasoning_effortvalues other thannonereach AWS on an opted-in Grok deployment and change the answer's reasoning budget, where Converse dropped every Grokreasoning_effortand AWS applied its default;noneis still dropped so that request answers 200 as before (design decision 13)additionalModelRequestFieldsandtop_kreaches AWS's native endpoint as sent and is ignored with a 200, where Converse forwarded it inadditionalModelRequestFieldsstore,prompt_cache_key,safety_identifier,seed,modalities,web_search_options,logit_biason GPT 5.6+ and opted-in Grok, andlogprobsandtop_logprobson opted-in gpt-oss and Grok (on GPT 5.6+ those two pass underreasoning_effort: "none"only, design decision 7)get_supported_openai_paramsand/model_group/infoon the native-routed ids return the native endpoint's list (31 params ongpt-5.6-solwhere the merge base listed 13):requestMetadata,thinking, andoutput_configleave it and the OpenAI-style params above join it, so a caller reading that list to pre-filter a request sees a different setstopstill answers Converse's 400 on every native-routed model, as on main, although gpt-oss and Grok would accept it natively; serving it natively is a follow-up once the reasoning truncation it causes has an answer (design decision 16)LITELLM_LOCAL_MODEL_COST_MAPunset a hosted cost-map update that lists/v1/chat/completionson a further OpenAI GPT 5.6+ Bedrock row moves that deployment to the native path at its next restart without a code upgrade, the same way/v1/responsesopt-ins already work forbedrock_supports_openai_responses; any other id needs thechat_completions/prefix as wellbedrock/<region>/xai.grok-4.6(a bare id after a region path, no profile prefix) still resolves toinvokebecause no cost-map row exists for the bare id, as on main; the region-path tests cover the profile-prefixed formsbedrock/converse/<model>name; there is no config-level switch to pin a whole deployment to ConverseLow
reasoning_effort: "none"on AWS, so itstemperature,top_p, penalties,logprobs, andtop_logprobsalways 400 (or drop underdrop_params) and its function tools always take Converseopenai.gpt-6-sol,openai.gpt-6-luna, andopenai.gpt-6.1-solrows (nosupported_endpoints), gpt-5.4, gpt-5.5, gpt-oss-safeguard, and thezai.GLM ids stay on Converse although AWS serves them natively; a GPT 5.6+ row joins with a one-line cost-map change, the others also need thechat_completions/prefix (design decision 2)temperatureortop_punderdrop_paramsnow answer 200 with the param stripped on Converse too, where the merge base forwarded them and AWS answered 400 (the rows gainsupports_sampling_params: false)/v1/messageson GPT 5.6+ answers atextblock alone where Converse added an emptyredacted_thinkingblock/v1/responsesstill refusestemperatureunderreasoning.effort: none, a pre-existing Responses check this PR leaves alonereasoning_effortstring AWS does not know (an empty string, a made-up effort) reaches the native endpoint and answers its 400, where Converse dropped it and answered 200; an int or list is refused in litellm or dropped instead (the Medium row above)temperaturesent as a string underreasoning_effort: "none"now reaches AWS for its 400 (expected a decimal, but got a string) after one call, where the merge base refused it in litellm with no callContent-Type(URL extension, then magic bytes), so a png, jpeg, gif, webp, heic, or pdf served asapplication/octet-streamreaches every one of those providers as its real type; every request this changes was already failing or wrong at the merge base, as the kept image cases show (a 400 from Anthropic, OpenAI, Bedrock invoke, Vertex Gemini, and Vertex Claude, Gemini naming the wrong colors, Responses spending its whole budget on reasoning)ImageFetchError, aBadRequestError) in under a second with no provider call, where the merge base answered a 500 (the Grok case) or forwardeddata:application/octet-streamfor the provider to reject with its own 400 (the Anthropic case); a format the byte sniff does not know (bmp, tiff, avif, svg) behind a genericContent-Typeand no known extension lands in that 400 too, where before it was forwarded asdata:application/octet-streamand rejected by the providerbudget-ratchetandrust-testreds the earlier tips carried are green there too, since the merge of main a206882 brought in main's budget files and its fix forlitellm-traces::query named::result_contracts_preserve_public_field_names. Two non-required checks are red at the tip, each main's own and neither touched by this PR's diff against main:Verify schema.d.ts matches the proxy OpenAPI spec(https://github.com/BerriAI/litellm/actions/runs/37084822447/job/111092794281) reports only the Trace* components that main's refactor(traces): type the ClickHouse query help response #44285 (dd86ca5) retyped to the enum"otel_traces" | "agent_traces_by_key" | "spend_logs"without regeneratingui/litellm-dashboard/src/lib/http/schema.d.ts(its last change on main is 688d791), the same check passed at a109787 whose merge ref predates that commit, and this PR changes no dashboard type;osv-scan(https://github.com/BerriAI/litellm/actions/runs/37084822524/job/111092794810) flagsbraces3.0.3 in main's ownui/litellm-dashboard/package-lock.jsonunder GHSA-vfj7-8cjw-p6xm, whose entry was modified at 2026-10-02T22:45Z, after main's last green scan at d131c43, and this PR does not touch the lockfile. Net negative to fix here: both fixes are main-side changes (a regenerated schema.d.ts, a lockfile bump) that belong in their own PR against main, not in a Bedrock chat completions change, and the schema one is filed as a follow-up on the private team<reasoning>prefix is split intoreasoning_contentlike the real one, since AWS frames the reasoning inline with no other discriminator; left as is, because a heuristic to tell the two apart would misfire more often than a literal prefix appearstransformation.py(reasoning output is not scanned by response guardrails, which readcontentand streameddelta.contentonly) is left open for a maintainer. The same gap already exists on Converse'sreasoning_content, so this PR does not introduce it, and closing it is a guardrail-translation change, not a Bedrock onelitellm/utils.py(Converse-only kwargs such asguardrailConfiginvisible to the route decision) was checked live and is a false positive:pre_process_non_default_paramscopies those kwargs intopassed_paramsbeforeget_optional_paramspicks the route, so aguardrailConfigrequest maps to Converse (thechat_gpt56_guardrailcase). Greptile's earlier threads (top_kandadditionalModelRequestFieldson the Converse fallback,allowed_openai_paramsrestoringnone, the refused-effort table living in code) were each answered in thread and withdrawn; its P1 at 22e35c6 (sampling params refused on GPT-5.6 even underreasoning_effort: "none") is fixed in 952adfc, which is what design decision 7 describesImageFetchErroron the native route and still 500APIConnectionErroron Converse (theconverse/pin, the guardrail and ARN fallbacks), as on main for every Converse call (the audit's A20a and A20b rows). Net negative to fix here: aligning Converse means touchingBedrockImageProcessor, which every Bedrock invoke and Converse model shares, so it is a follow-up on its ownapi_baseor a provider that disconnects mid-request answers 503ServiceUnavailableErroron the native route, where Converse's non-stream path answered 500APIConnectionError(its stream path already answered 503); thechat_gpt56_api_base_only_dead_portcase and the audit's B2 row. Shipped as is: 503 is the class the Converse stream path and the other OpenAI-compatible handlers already use, and a 500 for an unreachable upstream was the odd one out/v1/responseson a deployment carryingmodel_idstill calls the base model natively and ignores the ARN, and prefersapi_baseoveraws_bedrock_runtime_endpoint(503 on theresponses_gpt56_runtime_endpointcell on both legs); both as on main, since the Responses route predates this PR. Net negative to fix here: it is the Responses transform, not the chat route, and gets its own follow-up/model/newunderstore_model_in_db) goes out withoutstream_options.include_usage, since the proxy's usage-tracking step runs before the deployment is resolved; pre-existing on main for every provider (1 of 8 immediate streams in the audit's probe), the audit's A15 cell sendsstream_optionsfrom the client so its wire literal is deterministic. Net negative to fix here: the fix is incommon_request_processing.py, shared by every provider, and gets its own follow-up/v1/responsesspend row can carry the pre-encryptionresp_<base64>id instead of the ciphertext the caller received, because the row id is read before theResponsesIDSecurityhook rewrites it in place; pre-existing on main for every provider, both audit legs show it, and the chaos cells match such a row by the upstream id inside the payload (the test's TODO names it). Net negative to fix here: it is the spend-log id capture, not the Bedrock route, and gets its own follow-upget_supported_openai_paramsfor the native-routed ids inherits the OpenAI-compatible list, which also namesaudio,modalities,prediction,store,web_search_options,safety_identifier,prompt_cache_key, andprompt_cache_retention; AWS's runtime endpoint answers 200 and ignores them (the Medium row above), so a caller reading the list to pre-filter sends them for nothing. Net negative to fix: trimming the list means a hand-kept per-endpoint table that drifts from AWS's own as it adds support, for params that already answer 200gpt-5.6-sol-profile, andmodel_idARN rows) logspend: 0.0, since the ARN has no cost-map row and the base model is not read back from it; as on main, both legs of the live audit cells show it. Net negative to fix here: cost attribution for profile ARNs is a cost-map lookup change outside this route/v1/messagescall on a runtime-endpoint deployment posted to AWS twice, 12 s apart, on the merge base leg only (messages_gpt56_runtime_endpointBefore); the native route posts once. Left alone, it is the path this PR moves off/model/newon a 2-worker proxy is no longer a cell; on that shape the worker that did not serve/model/newcan answer 400Invalid model nameonce the registry read-through's resync budget is spent, pre-existing on main for every provider (the audit's hashes note). Net negative to fix here: it is the read-through inlitellm/proxy/common_utils/registry_read_through.py, shared by every route, and gets its own follow-upFinal Attestation
The Before and After legs above ran on 2026-10-02 at the merge base c168199 and the tip 6daf0b7 against live AWS Bedrock in
us-west-2with two uvicorn workers per proxy and a recorder in front of AWS, covering the GPT 5.6, 6, and 6.1 defaults, the Converse fallbacks (guardrails, an application inference profile ARN, amodel_idARN override, theconverse/pin, tools with reasoning), a blank-key SigV4 deployment, the reasoning-tied params undernoneand under other efforts, a non-stringreasoning_effortwith and withoutdrop_params, streaming, images,/v1/messages, and/v1/responses, plus the gpt-oss and Grok defaults and theirchat_completions/opt-in, and the same cases ran on a local merge into main. The kept image cases ran on 2026-09-26 at 1f77fa6 and 4900a1a (the image cases' own hashes, kept as recorded); the only change toimage_handling.pysince then (3e3b9cd) hoists the URL comprehension into a helper, so they still describe the tip. The /audit matrix above ran at c168199 and 6daf0b7 on the scripted runtime peer plus real AWS for the live rows, with the two rerun files at e518e0e, a tests-only commitstopBefore/After leg ran at the merge base and this tip (the 4900a1a run above repeats those cases) with the three earlier native paths re-checked, and CI at this tip is 99 passed, 0 failed, 2 skippedmodel_info.rswith main's style) and the move of the transformation tests undertests/unit/, so this PR's diff against main is otherwise unchanged from the one every proof above ran againstContent-Type,ImageFetchErrorfor an image whose type cannot be inferred, the syncinline_remote_media) and the Bedrock chat transform (http(s) images inlined before the native call,stop,top_k, andadditionalModelRequestFieldskept on Converse, Grokreasoning_effort: "none"dropped), and every dependent of the fetch reachable on this rig ran live at the merge base and the tip with a recorder in front of the provider (Bedrock native chat, Responses, and Messages, Bedrock invoke Claude, Anthropic overhttp://, Gemini image and PDF, OpenAI PDF, Vertex Gemini, Vertex Claude). Breaking: none. Backward incompatible: an image whose type cannot be inferred is a 400ImageFetchErrorinstead of a 500 or the provider's own 400, and a png, jpeg, gif, webp, heic, or pdf behindapplication/octet-streamnow reaches the provider typed instead of being rejected. Regression risk: none observed, every Before 200 is an After 200 with the same answer. Not verified: Snowflake, Ollama, and Bedrock Mantle (the same fetch helper, no live deployment on this rig) and GovCloud (no credentials)validate_environmentoverride, so the route inherits the OpenAI-compatible one (Content-Typeonly, a bearer only when anapi_keyis given) whilesign_requeststill signs every call (SigV4, or theAWS_BEARER_TOKEN_BEDROCKbearer the legs used), andaws_bedrock_project_idno longer becomes anOpenAI-Projectheader. Dependents: every native chat request (the whole After leg re-ran at this tip with the same status codes and outbound paths as at cbf01c2), a deployment carryingaws_bedrock_project_id(thechat_gptoss_project_idcase, 200 on both legs with noOpenAI-Projectamong the recorded request headers on either), the Mantle configs (untouched, they keep sending it), and the native Responses transform (never sent it). Breaking: none. Backward incompatible: a runtime deployment configured withaws_bedrock_project_idunder an earlier revision of this PR stops sending a header the runtime host ignored (200 to a bogus value, checked live). Regression risk: none observedreasoning_effort: "none"check in the route decision and the refused-param filter became an equality test, so a list value no longer raises there) and a0cef91 (a type guard in the refused-effort filter, so a non-stringreasoning_effortis forwarded instead of crashing). Dependents: the route decision for every Bedrock chat, Messages, and Responses request,get_supported_openai_paramsand/model_group/infoon Bedrock ids, the Converse fallback forguardrailConfig, theconverse/prefix and application inference profile ARNs, cost and usage on all three endpoints, and streaming withinclude_usage. Every one ran live at the merge base and at this tip with a forwarding recorder in front of Bedrock, one proxy per side with 2 workers (the two sides ran one after the other on the same host at CRITICAL headroom rather than at once), and the same cells re-ran on a local merge of this tip into main 54ae4c5 with identical routes, statuses, wire bodies, and rates. Breaking: a listreasoning_effortanswered 500 (unhashable type: 'list', no upstream call) at 19176f6 where the merge base answered 200, fixed in a0cef91 (AWS now answers its own 400 after one call). Backward incompatible:get_supported_openai_paramsand/model_group/infolist the native endpoint's 31 params for GPT 5.6+ ids where the merge base listed 13 or 15 (design decision 7, shipped by the author); a Converse-routed gpt-6 request carryingtemperature(aguardrailConfigrequest) now 400s in litellm with no upstream call where the merge base let AWS 400 it after one call, and drops it underdrop_params(thesupports_sampling_paramsLow caveat); an int or listreasoning_effortanswers AWS's 400 where Converse dropped it and answered 200 (the Low caveat above); a stringtemperatureunderreasoning_effort: "none"reaches AWS for its 400 where litellm refused it with no call; the 400 message for sampling params on GPT 5.6+ names thenoneway out (the Medium caveat above). Regression risk: none observed beyond those; a hosted cost-map update that lists/v1/chat/completionson a further GPT 5.6+ row moves that deployment to the native route at its next restart (the Medium caveat above). Not verified: Bedrock Mantle and GovCloud (no deployment or credentials on this rig), and a deployment reading the hosted cost map rather thanLITELLM_LOCAL_MODEL_COST_MAPValueError, and the native URL is built fromaws_bedrock_runtime_endpointlike Converse) and 578f26e (model_idjoins the Converse-only keys; the native config resolves the bearer throughbedrock_bearer_token). Dependents walked for both: the route decision for every Bedrock chat, Messages, and Responses request (get_bedrock_route,bedrock_route_for_request),get_supported_openai_paramsand/model_group/info, the Converse fallbacks (guardrailConfig,converse/, the ARN andmodel_idforms),BaseAWSLLM._sign_requestand every caller (Converse, invoke, Responses, the pass-through routes), every reader ofaws_bedrock_runtime_endpoint, image inlining, cost and usage on all three endpoints, and streaming withinclude_usage. Every one ran live at the merge base and at this tip with a forwarding recorder in front of Bedrock, one proxy per side with 2 workers, and the same 43 cases ran on a local merge of this tip into main 4b1d9bf (4f930db3a3) with identical routes, statuses, and wire bodies. Breaking: none (at dae29c1 amodel_iddeployment answered 400 and a blank-key SigV4 deployment 403, both fixed in 578f26e before any release, shown in the Before fix 2 block). Backward incompatible, each shipped by this run as a decision row in the /audit matrix for the reviewer's read: a remote image the proxy cannot fetch answers 400ImageFetchErroron the native route where Converse and main answer 500; an unreachableapi_baseor a mid-request disconnect answers 503 where Converse's non-stream path answered 500;get_supported_openai_paramslists the native endpoint's 31 params where the merge base listed 13; the sampling-param 400 names thenoneway out; and one left to the author's decision at the read (the Severe caveat at that tip, decided in 6daf0b7): an int or listreasoning_effortanswers AWS's 400 where Converse dropped it and answered 200. Regression risk: a hosted cost-map update listing/v1/chat/completionson a further GPT 5.6+ row moves it to the native route at restart (Medium caveat); the cold-workerstream_optionsrace is pre-existing and provider-independent (Low caveat). Not verified: Bedrock Mantle and GovCloud (no deployment or credentials on this rig), and a deployment reading the hosted cost map rather thanLITELLM_LOCAL_MODEL_COST_MAPnon_string_reasoning_effortin the native transform (an int or listreasoning_effortanswers litellm's 400 before any call, or is dropped underdrop_params). Dependents walked:map_openai_paramson every native-routed chat request, thedrop_paramsresolution (global, deployment, request body), the refused-param filter next to it, and/v1/messagesand/v1/responsescarrying a non-string effort. Every one ran live at the merge base and this tip with a forwarding recorder in front of Bedrock, one proxy per side with 2 workers (the 30 cells above), and the same cells ran on a local merge of this tip into main d729f97 (33d3c0f5ff) with identical routes, statuses, and wire bodies. Breaking: none. Backward incompatible: an int or listreasoning_efforton a native-routed deployment withoutdrop_paramsanswers litellm's 400 with no AWS call where Converse dropped it and answered 200 (the Medium caveat, design decision 19, shipped by the author), and/v1/messageswith a top-level non-stringreasoning_effortanswers the same 400; underdrop_paramsboth answer 200 as before. Regression risk: none observed. Not verified: Bedrock Mantle and GovCloud (no deployment or credentials on this rig), and a deployment reading the hosted cost map rather thanLITELLM_LOCAL_MODEL_COST_MAPmerged_head1, 54 passed and the B4 teardown failure its row records) and the chaos file rerun at 6befa50 (merged_head4 4 passed in 133 s and merged_head5 4 passed in 120 s, no skips, no retries, against a head proxy booted from 6befa50 with two uvicorn workers)Note
High Risk
Changes default routing for many Bedrock GPT deployments (Converse to native), altering IDs, reasoning fields, and param validation; incorrect fallback logic could misroute guardrails, tools, or structured output.
Overview
Adds native Bedrock Runtime OpenAI Chat Completions (
/openai/v1/chat/completions) for GPT 5.6+ by default and optionalbedrock/chat_completions/opt-in for gpt-oss and Grok, with a newAmazonBedrockRuntimeChatCompletionsConfighandling AWS signing, param mapping (max_completion_tokens, reasoning-effort rules),<reasoning>tag splitting, and remote image inlining.Routing is now request-aware:
bedrock_route_for_requestand updatedget_bedrock_routepick native chat vs Converse from cost-map flags and Converse-only features (guardrails,model_idARNs, tools+reasoning,response_format, etc.), keeping param mapping and dispatch aligned. Completion dispatch passes rawrequest_paramsinto that decision.Supporting changes: shared
inline_remote_mediaand smarter image MIME inference for native vision; price-map/catalog flags for native response-format and tools-with-reasoning; Responses config stripschat_completions/from model names for capability checks. Broad integration tests cover wire, sad paths, and chaos.Reviewed by Cursor Bugbot for commit 6befa50. Bugbot is set up for automated code reviews on this repo. Configure here.