Skip to content

feat(bedrock): serve the OpenAI models on bedrock-runtime's native Responses API (internal copy of #38489) - #42767

Merged
mateo-berri merged 8 commits into
mainfrom
litellm_bedrock_openai_native_responses
Sep 23, 2026
Merged

mateo-berri merged 8 commits into
mainfrom
litellm_bedrock_openai_native_responses

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Internal copy of #38489 (its two commits cherry-picked with the author preserved) so the work can be driven to merge; the fork PR does not allow maintainer edits

Problem this solves:

  • OpenAI models on Bedrock ran /v1/responses through the Converse bridge
  • Any client sending prompt_cache_key got a 400 from that bridge
  • Codex only worked because the gateway silently dropped that parameter for it
  • Bedrock's native Responses endpoint for these models was never reachable through the gateway

How it solves it:

  • Sends opted-in models to bedrock-runtime's native /openai/v1/responses
  • Opt-in is the price map's supported_endpoints: ["/v1/responses"] on the row
  • Opts in the gpt-5.6 family (from feat(bedrock): serve the OpenAI models on bedrock-runtime's native Responses API #38489) and the gpt-6 family (this copy)
  • Signs with SigV4 or the Bedrock bearer token, host from the region's partition
  • Normalizes Codex-only history items the same way the Mantle route already does
  • Drops tool types bedrock-runtime rejects (Codex's web_search) with a warning
  • Keeps file_search on the emulated path (bugbot finding)
  • Custom endpoints ending in /v1/responses or /responses no longer double the path (bugbot finding)
  • Drops background with a proxy-log warning, since bedrock-runtime rejects it and the bridge never forwarded it
  • Downloads remote http(s) image URLs (input_image blocks in message content and in function_call_output lists, and computer_screenshot outputs of a computer_call_output) and inlines them as data URIs, since bedrock-runtime accepts only data: and s3://; the bridge downloaded the message-content ones and 400'd on both tool-output shapes (bugbot finding for the list shape, the /caveats pass for the screenshot dict)

The gpt-6 opt-in is a decision made here: all six us./global. gpt-6 ids answered 200 on the native endpoint with both auth kinds, and the alternative (leaving them on the bridge) keeps the 400 above for the newest models. Anything not flagged in the price map keeps the bridge exactly as today

Two more decisions made here, both found by probing the native endpoint with every Responses parameter at the first tip: background: true is dropped unconditionally with a proxy-log warning, matching the bridge, which never forwarded it either, and the alternative (raising unless drop_params is set) keeps a 400 on a request the bridge answered 200. Remote http(s) image URLs, input_image blocks in message content and in function_call_output lists and computer_screenshot outputs alike, are fetched and inlined as data URIs through the shared image_handling helpers (the same SSRF policy and size cap every other inlining provider uses), with the sync and async fetchers injected through the config constructor so tests never touch the network; the alternative (documenting the 400) breaks a message-content request the bridge answered 200, and both tool-output shapes were a 400 on the bridge too (Converse rejects an image inside a tool result for this model and never forwards a computer_call_output at all), so inlining them is new capability rather than parity

User Flow

Before: a developer sending prompt_cache_key to bedrock/us.openai.gpt-5.6-sol on /v1/responses gets a 400, and only Codex gets through because the gateway drops the parameter for it

  1. They add bedrock-gpt-5.6-sol (bedrock/us.openai.gpt-5.6-sol) to the gateway config and create a virtual key
  2. From the OpenAI Python SDK they send POST https://litellm-domain/v1/responses with model: bedrock-gpt-5.6-sol, a long instructions string, a short input, and prompt_cache_key: "my-session"
  3. It comes back HTTP 400 litellm.UnsupportedParamsError: bedrock does not support parameters: ['prompt_cache_key'], for model=us.openai.gpt-5.6-sol
  4. They send the same request from Codex CLI (wire_api = "responses"), and it comes back 200 only because the gateway silently dropped prompt_cache_key for that user agent, so the key never reached Bedrock
  5. GET https://litellm-domain/model/info shows the row with supported_endpoints: null

After: the same request comes back 200 from Bedrock's native Responses endpoint, and the repeat call reports the cached tokens

  1. They add bedrock-gpt-5.6-sol (bedrock/us.openai.gpt-5.6-sol) to the gateway config and create a virtual key
  2. From the OpenAI Python SDK they send POST https://litellm-domain/v1/responses with model: bedrock-gpt-5.6-sol, a long instructions string, a short input, and prompt_cache_key: "my-session"
  3. It comes back HTTP 200 with the answer and usage.input_tokens_details.cache_write_tokens above 0
  4. They send the identical request again and it comes back 200 with usage.input_tokens_details.cached_tokens above 0 and a lower x-litellm-response-cost header; the same request from Codex CLI comes back 200 with the parameter forwarded rather than dropped
  5. GET https://litellm-domain/model/info shows the row with supported_endpoints: ["/v1/responses"]

Relevant issues

Copy of #38489. Related: #29818, #36182

Affected release

Linear ticket

Resolves LIT-7790

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Shared setup for both legs: one proxy per leg, each booted with --num_workers 2 on its own random port with its own Postgres database, the same config on both. The config lists bedrock-gpt-5.6-sol (bedrock/us.openai.gpt-5.6-sol, us-east-1, SigV4 keys), bedrock-gpt-5.6-luna-global (bedrock/global.openai.gpt-5.6-luna), bedrock-gpt-5.6-sol-bearer (same model with a Bedrock API key as api_key), bedrock-gpt-6-sol (bedrock/us.openai.gpt-6-sol), bedrock-claude-haiku (bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0, not opted in) and bedrock-mantle-gpt-5.6-sol (bedrock_mantle/openai.gpt-5.6-sol). Every call hits the real Bedrock endpoints. $KEY is a virtual key created on that leg, $MASTER its master key, long_instructions.txt is an 11 KB instructions string (about 2,100 tokens). The Before proxy is on port 49115, the After proxy on port 38534 (each tip's proxy got a fresh port; the port in each After block is the one that run used). Codex CLI 0.155.1 runs interactively in tmux with ~/.codex/config.toml pointing base_url at the proxy, model = "bedrock-gpt-5.6-sol", wire_api = "responses", in a directory holding notes.txt (2 lines) and numbers.txt (5 lines)

Before (e73f949)

Codex CLI session (shell tool round trip)

  1. Prompt typed into the Codex TUI: Run ls in this directory, then tell me the total number of lines across notes.txt and numbers.txt
  2. Codex runs ls and wc -l, then prints Total lines across both files: 7. The session works on the base only because the gateway auto-enables drop_params for Codex's user agent, so the prompt_cache_key Codex sends is thrown away before the Converse bridge sees it

pr42767-e4b816e7ec-codex-before.png

prompt_cache_key on POST /v1/responses, three clients

  1. Same minimal body from three user agents:
BODY='{"model":"bedrock-gpt-5.6-sol","input":"Reply with the single word pong","prompt_cache_key":"lit7790-demo"}'
for ua in "" "codex_cli_rs/0.155.1" "OpenAI/Python 2.33.0"; do
  curl -s -o /tmp/out.json -w "UA=${ua:-<curl default>} HTTP %{http_code}\n" -X POST http://127.0.0.1:49115/v1/responses \
    -H "Authorization: Bearer $KEY" -H 'content-type: application/json' ${ua:+-H "User-Agent: $ua"} -d "$BODY"
  jq -c '{status, text: [.output[]? | select(.type=="message") | .content[]?.text], error: (.error.message // .error | tostring | .[0:120])}' /tmp/out.json
done
  1. Output:
UA=<curl default> HTTP 400 {"status":null,"text":[],"error":"litellm.UnsupportedParamsError: bedrock does not support parameters: ['prompt_cache_key'], for model=us.openai.gpt-5.6-sol. To drop these, set `litellm.drop_params=True` or for proxy:\n\n`litellm_settings:\n drop_params: true`\n. \n If you want to use these params "}
UA=codex_cli_rs/0.155.1 HTTP 200 {"status":"completed","text":["pong"],"error":"null"}
UA=OpenAI/Python 2.33.0 HTTP 400 {"status":null,"text":[],"error":"litellm.UnsupportedParamsError: bedrock does not support parameters: ['prompt_cache_key'], for model=us.openai.gpt-5.6-sol. To drop these, set `litellm.drop_params=True` or for proxy:\n\n`litellm_settings:\n drop_params: true`\n. \n If you want to use these params "}

The same request from the OpenAI Python SDK, twice

  1. openai 2.33.0: OpenAI(base_url="http://127.0.0.1:49115", api_key=KEY).responses.create(model="bedrock-gpt-5.6-sol", instructions=<the 11 KB instructions>, input="Reply with the single word pong", prompt_cache_key="lit7790-sdk"), called twice two seconds apart, printing status, output_text, and usage.input_tokens_details
  2. Output:
before call 1: BadRequestError: Error code: 400 - {'error': {'message': "litellm.UnsupportedParamsError: bedrock does not support parameters: ['prompt_cache_key'], for model=us.openai.gpt-5.6-sol. To drop these, set `litellm.drop_params=True` or for proxy:\n\n`litellm_settings:\n drop_params: true`\n. \n If you want to use these p
before call 2: BadRequestError: Error code: 400 - {'error': {'message': "litellm.UnsupportedParamsError: bedrock does not support parameters: ['prompt_cache_key'], for model=us.openai.gpt-5.6-sol. To drop these, set `litellm.drop_params=True` or for proxy:\n\n`litellm_settings:\n drop_params: true`\n. \n If you want to use these p

prompt_cache_key with a long prompt, twice (native_params.json: 11 KB instructions, prompt_cache_key: "lit7790-qa-session", store: false, include: ["reasoning.encrypted_content"], reasoning: {effort: low, summary: auto})

  1. Command, run twice:
curl -s -D hdr.txt -o out.json -X POST http://127.0.0.1:49115/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d @native_params.json
head -1 hdr.txt; jq -c '{status, input_tokens: .usage.input_tokens, cached: .usage.input_tokens_details.cached_tokens, cache_write: .usage.input_tokens_details.cache_write_tokens, text: [.output[]? | select(.type=="message") | .content[]?.text], error: (.error.message // .error | tostring | .[0:100])}' out.json
  1. Output, both calls identical:
HTTP/1.1 400 Bad Request
{"status":null,"input_tokens":null,"cached":null,"cache_write":null,"text":[],"error":"litellm.UnsupportedParamsError: bedrock does not support parameters: ['prompt_cache_key'], for model=us.openai.gpt-5.6-sol. To"}

Same long prompt without prompt_cache_key, twice (resp_cache.json, Bedrock's implicit caching)

  1. Same command with -d @resp_cache.json, twice, plus grep -i x-litellm-response-cost: hdr.txt
  2. Output:
HTTP/1.1 200 OK
{"status":"completed","input_tokens":2176,"output_tokens":5,"cached":0,"cache_write":2174,"text":["pong"],"error":"null"}
x-litellm-response-cost: 0.0120758
HTTP/1.1 200 OK
{"status":"completed","input_tokens":2176,"output_tokens":5,"cached":2174,"cache_write":0,"text":["pong"],"error":"null"}
x-litellm-response-cost: 0.00107536

Streaming

  1. Command:
curl -s -N -X POST http://127.0.0.1:49115/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' \
  -d '{"model":"bedrock-gpt-5.6-sol","input":"Count from 1 to 5, one number per line","stream":true}' > stream.txt
wc -l < stream.txt; grep -c 'response.output_text.delta' stream.txt
grep 'response.completed' stream.txt | sed 's/^data: //' | jq -c '{status: .response.status, model: .response.model, usage: .response.usage, outer_model: .model}'
  1. Output:
36
9
{"status":"completed","model":"us.openai.gpt-5.6-sol","usage":{"input_tokens":18,"input_tokens_details":{"cached_tokens":0,"text_tokens":18,"cache_write_tokens":0},"output_tokens":13,"output_tokens_details":{"reasoning_tokens":0,"text_tokens":13},"total_tokens":31,"cost":0.0003652},"outer_model":"bedrock-gpt-5.6-sol"}

Tool round trip over raw HTTP (tool_history.json: user turn, a function_call to shell, its function_call_output listing two files, one function tool)

  1. curl -s -o out.json -w 'HTTP %{http_code}\n' -X POST http://127.0.0.1:49115/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d @tool_history.json
  2. Output:
HTTP 200 {"status":"completed","text":["There are 2 files."],"error":"null"}

Other model ids and auth kinds

  1. for m in bedrock-gpt-5.6-luna-global bedrock-gpt-5.6-sol-bearer bedrock-gpt-6-sol; do curl -s -o out.json -w "$m HTTP %{http_code} " -X POST http://127.0.0.1:49115/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d "{\"model\":\"$m\",\"input\":\"Reply with the single word pong\"}"; jq -c '{status, input_tokens: .usage.input_tokens, output_tokens: .usage.output_tokens, text: [.output[]? | select(.type=="message") | .content[]?.text]}' out.json; done
  2. Output:
bedrock-gpt-5.6-luna-global HTTP 200 {"status":"completed","input_tokens":12,"output_tokens":5,"text":["pong"]}
bedrock-gpt-5.6-sol-bearer HTTP 200 {"status":"completed","input_tokens":12,"output_tokens":5,"text":["pong"]}
bedrock-gpt-6-sol HTTP 200 {"status":"completed","input_tokens":12,"output_tokens":5,"text":["pong"]}

Spend rows for the two long-prompt calls

  1. curl -s "http://127.0.0.1:49115/spend/logs?request_id=<x-litellm-call-id>" -H "Authorization: Bearer $MASTER" | jq -c '[.[] | {model, custom_llm_provider, call_type, spend, prompt_tokens, completion_tokens, cache_read_input_tokens, total_tokens}]' for the two x-litellm-call-id headers of the implicit-caching pair (the prompt_cache_key calls were 400s, so they have no rows), then curl -s http://127.0.0.1:49115/key/info -H "Authorization: Bearer $KEY" | jq -c '{alias: .info.key_alias, spend: .info.spend}'
  2. Output:
request_id=eb12c62f-8a22-432e-b791-13f9621aaebd
[{"model":"bedrock/us.openai.gpt-5.6-sol","custom_llm_provider":"bedrock","call_type":"aresponses","spend":0.0120758,"prompt_tokens":2176,"completion_tokens":5,"cache_read_input_tokens":null,"total_tokens":2181}]
request_id=0e073eb2-9e59-4596-887a-43e2daccb660
[{"model":"bedrock/us.openai.gpt-5.6-sol","custom_llm_provider":"bedrock","call_type":"aresponses","spend":0.00107536,"prompt_tokens":2176,"completion_tokens":5,"cache_read_input_tokens":null,"total_tokens":2181}]
{"alias":"lit7790-before2-dev","spend":0.07093042}

Chat completions on the same model

  1. curl -s -o out.json -w 'HTTP %{http_code}\n' -X POST http://127.0.0.1:49115/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"bedrock-gpt-5.6-sol","messages":[{"role":"user","content":"Reply with the single word pong"}]}'; jq -c '{model, text: .choices[0].message.content, usage: {prompt_tokens: .usage.prompt_tokens, completion_tokens: .usage.completion_tokens}}' out.json
  2. Output:
HTTP 200
{"model":"bedrock-gpt-5.6-sol","text":"pong","usage":{"prompt_tokens":12,"completion_tokens":5}}

Neighbouring paths (the /live-pr-risk cases)

  1. A: bedrock-claude-haiku (not opted in) on POST /v1/responses; B: bedrock-mantle-gpt-5.6-sol on POST /v1/responses with the Codex-shaped tool history plus a web_search tool; B2: the mantle model with prompt_cache_key; E: bedrock-gpt-5.6-sol on POST /v1/messages; F: GET /model/info; G: bedrock-gpt-5.6-sol on POST /v1/responses with a file_search tool naming a vector store that does not exist
  2. Output:
=== A unflagged bedrock model on /v1/responses (bedrock-claude-haiku)
HTTP 200 {"model":"bedrock-claude-haiku","status":"completed","input_tokens":14,"output_tokens":5,"text":["pong"],"error":"null"}
=== B mantle /v1/responses with Codex-shaped history plus web_search
HTTP 200 {"model":"bedrock-mantle-gpt-5.6-sol","status":"completed","input_tokens":2992,"output_tokens":20,"text":["There are **2 files**: `notes.txt` and `numbers.txt`."],"error":"null"}
=== B2 mantle /v1/responses plain
HTTP 200 {"model":"bedrock-mantle-gpt-5.6-sol","status":"completed","input_tokens":12,"output_tokens":5,"text":["pong"],"error":"null"}
=== E /v1/messages for a flagged model
HTTP 200 {"model":"bedrock-gpt-5.6-sol","type":"message","stop_reason":"end_turn","text":["pong"],"usage":{"input_tokens":12,"output_tokens":5},"error":"null"}
=== F /model/info
{"litellm_model":"bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0","litellm_provider":"bedrock_converse","mode":"chat","model_name":"bedrock-claude-haiku","supported_endpoints":null,"supports_prompt_caching":true}
{"litellm_model":"bedrock/global.openai.gpt-5.6-luna","litellm_provider":"bedrock_converse","mode":"chat","model_name":"bedrock-gpt-5.6-luna-global","supported_endpoints":null,"supports_prompt_caching":null}
{"litellm_model":"bedrock/us.openai.gpt-5.6-sol","litellm_provider":"bedrock_converse","mode":"chat","model_name":"bedrock-gpt-5.6-sol","supported_endpoints":null,"supports_prompt_caching":null}
{"litellm_model":"bedrock/us.openai.gpt-5.6-sol","litellm_provider":"bedrock_converse","mode":"chat","model_name":"bedrock-gpt-5.6-sol-bearer","supported_endpoints":null,"supports_prompt_caching":null}
{"litellm_model":"bedrock/us.openai.gpt-6-sol","litellm_provider":"bedrock_converse","mode":"chat","model_name":"bedrock-gpt-6-sol","supported_endpoints":null,"supports_prompt_caching":true}
{"litellm_model":"bedrock_mantle/openai.gpt-5.6-sol","litellm_provider":"bedrock_mantle","mode":"responses","model_name":"bedrock-mantle-gpt-5.6-sol","supported_endpoints":["/v1/chat/completions","/v1/responses"],"supports_prompt_caching":true}
=== G file_search tool on a flagged model (emulation path)
HTTP 200 {"model":"bedrock-gpt-5.6-sol","status":"completed","input_tokens":135,"output_tokens":5,"text":["pong"],"error":"null"}

After (b2d53f9)

Codex CLI session (shell tool round trip)

  1. Same prompt in the same Codex TUI, now against the After proxy at this tip
  2. Codex runs ls, then prints notes.txt and numbers.txt contain 7 total lines. (done 3:04 PM, about five seconds after the prompt). The gateway forwards Codex's prompt_cache_key to the native endpoint and drops the web_search tool that endpoint rejects, logging a warning in the proxy log

pr42767-b2d53f9-codex-after.png

prompt_cache_key on POST /v1/responses, three clients

  1. Same loop against port 38534
  2. Output:
UA=<curl default> HTTP 200 {"status":"completed","text":["pong"],"error":"null"}
UA=codex_cli_rs/0.155.1 HTTP 200 {"status":"completed","text":["pong"],"error":"null"}
UA=OpenAI/Python 2.33.0 HTTP 200 {"status":"completed","text":["pong"],"error":"null"}

prompt_cache_key with a long prompt, twice

  1. Same command against port 38534 (prompt_cache_key: "lit7790-qa-session", cold again after the proxy restart at this tip), plus grep -i 'x-litellm-response-cost:' hdr.txt
  2. Output:
HTTP/1.1 200 OK
x-litellm-response-cost: 0.0120758
{"model":"bedrock-gpt-5.6-sol","status":"completed","input_tokens":2176,"output_tokens":5,"cached":0,"cache_write":2174,"text":["pong"],"error":"null"}
HTTP/1.1 200 OK
x-litellm-response-cost: 0.00107536
{"model":"bedrock-gpt-5.6-sol","status":"completed","input_tokens":2176,"output_tokens":5,"cached":2174,"cache_write":0,"text":["pong"],"error":"null"}

The same request from the OpenAI Python SDK, twice

  1. openai 2.33.0: the same two responses.create(...) calls as the Before leg, against port 38534 with prompt_cache_key="lit7790-sdk-b2d53f9"
  2. Output:
after call 1: status=completed text='pong' input_tokens=2176 cached_tokens=0 cache_write=2174
after call 2: status=completed text='pong' input_tokens=2176 cached_tokens=2174 cache_write=0

Same long prompt without prompt_cache_key, twice

  1. Same command with -d @resp_cache.json against port 38534
  2. Output:
HTTP/1.1 200 OK
x-litellm-response-cost: 0.0120758
{"model":"bedrock-gpt-5.6-sol","status":"completed","input_tokens":2176,"output_tokens":5,"cached":0,"cache_write":2174,"text":["pong"],"error":"null"}
HTTP/1.1 200 OK
x-litellm-response-cost: 0.00107536
{"model":"bedrock-gpt-5.6-sol","status":"completed","input_tokens":2176,"output_tokens":5,"cached":2174,"cache_write":0,"text":["pong"],"error":"null"}

Streaming

  1. Same command against port 38534, with sequence_number added to the jq projection
  2. Output:
      36
9
data: {"type":"response.created","response":{"id":"resp_E1W9eaWmN7W8JskGyscYWqkMkru9UcbeXVCIJjhfcO-oGjAHKRGeFXOCP07fvtXKxdLUIliLQBP6hUmrvEiZwt2Wc63QOuP1MS-77kyW

{"status":"completed","model":"us.openai.gpt-5.6-sol","usage":{"input_tokens":18,"input_tokens_details":{"cached_tokens":0,"cache_write_tokens":0},"output_tokens":13,"output_tokens_details":{"reasoning_tokens":0},"total_tokens":31,"cost":0.0003652},"outer_model":"bedrock-gpt-5.6-sol","sequence_number":16}

Tool round trip over raw HTTP

  1. Same command against port 38534
  2. Output:
HTTP 200 {"status":"completed","text":["There are 2 files."],"error":"null"}

Other model ids and auth kinds

  1. Same loop against port 38534
  2. Output:
bedrock-gpt-5.6-luna-global HTTP 200 {"model":"bedrock-gpt-5.6-luna-global","status":"completed","input_tokens":12,"output_tokens":5,"cached":0,"cache_write":0,"text":["pong"],"error":"null"}
bedrock-gpt-5.6-sol-bearer HTTP 200 {"model":"bedrock-gpt-5.6-sol-bearer","status":"completed","input_tokens":12,"output_tokens":5,"cached":0,"cache_write":0,"text":["pong"],"error":"null"}
bedrock-gpt-6-sol HTTP 200 {"model":"bedrock-gpt-6-sol","status":"completed","input_tokens":12,"output_tokens":5,"cached":0,"cache_write":0,"text":["pong"],"error":"null"}

Spend rows for the two prompt_cache_key calls

  1. Same /spend/logs?request_id= and /key/info calls against port 38534 for the two prompt_cache_key call ids
  2. Output:
request_id=8f5429f7-b19d-4def-b160-6d2da0d990c8
[{"model":"bedrock/us.openai.gpt-5.6-sol","custom_llm_provider":"bedrock","call_type":"aresponses","spend":0.0120758,"prompt_tokens":2176,"completion_tokens":5,"cache_read_input_tokens":null,"total_tokens":2181}]
request_id=c2bd2c86-735f-4edd-809d-2b653890952a
[{"model":"bedrock/us.openai.gpt-5.6-sol","custom_llm_provider":"bedrock","call_type":"aresponses","spend":0.00107536,"prompt_tokens":2176,"completion_tokens":5,"cache_read_input_tokens":null,"total_tokens":2181}]
{"alias":"lit7790-after-dev","spend":0.5767394800000002}

Chat completions on the same model

  1. Same command against port 38534
  2. Output:
HTTP 200 {"model":"bedrock-gpt-5.6-sol","text":"pong","usage":{"completion_tokens":5,"prompt_tokens":12,"total_tokens":17,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":5},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":12,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0},"error":null}

Neighbouring paths (the /live-pr-risk cases)

  1. Same six cases against port 38534
  2. Output:
=== A unflagged bedrock model on /v1/responses (bedrock-claude-haiku)
HTTP 200 {"model":"bedrock-claude-haiku","status":"completed","input_tokens":14,"output_tokens":5,"text":["pong"],"error":"null"}
=== B mantle /v1/responses with Codex-shaped history plus web_search
HTTP 200 {"model":"bedrock-mantle-gpt-5.6-sol","status":"completed","input_tokens":2992,"output_tokens":20,"text":["There are **2 files**: `notes.txt` and `numbers.txt`."],"error":"null"}
=== B2 mantle /v1/responses plain
HTTP 200 {"model":"bedrock-mantle-gpt-5.6-sol","status":"completed","input_tokens":12,"output_tokens":5,"text":["pong"],"error":"null"}
=== E /v1/messages for a flagged model
HTTP 200 {"model":"bedrock-gpt-5.6-sol","type":"message","stop_reason":"end_turn","text":["pong"],"usage":{"input_tokens":12,"output_tokens":5},"error":"null"}
=== F /model/info
{"litellm_model":"bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0","litellm_provider":"bedrock_converse","mode":"chat","model_name":"bedrock-claude-haiku","supported_endpoints":null,"supports_prompt_caching":true}
{"litellm_model":"bedrock/global.openai.gpt-5.6-luna","litellm_provider":"bedrock_converse","mode":"chat","model_name":"bedrock-gpt-5.6-luna-global","supported_endpoints":["/v1/responses"],"supports_prompt_caching":null}
{"litellm_model":"bedrock/us.openai.gpt-5.6-sol","litellm_provider":"bedrock_converse","mode":"chat","model_name":"bedrock-gpt-5.6-sol","supported_endpoints":["/v1/responses"],"supports_prompt_caching":null}
{"litellm_model":"bedrock/us.openai.gpt-5.6-sol","litellm_provider":"bedrock_converse","mode":"chat","model_name":"bedrock-gpt-5.6-sol-bearer","supported_endpoints":["/v1/responses"],"supports_prompt_caching":null}
{"litellm_model":"bedrock/us.openai.gpt-6-sol","litellm_provider":"bedrock_converse","mode":"chat","model_name":"bedrock-gpt-6-sol","supported_endpoints":["/v1/responses"],"supports_prompt_caching":true}
{"litellm_model":"bedrock_mantle/openai.gpt-5.6-sol","litellm_provider":"bedrock_mantle","mode":"responses","model_name":"bedrock-mantle-gpt-5.6-sol","supported_endpoints":["/v1/chat/completions","/v1/responses"],"supports_prompt_caching":true}
=== G file_search tool on a flagged model (emulation path)
HTTP 200 {"model":"bedrock-gpt-5.6-sol","status":"completed","input_tokens":135,"output_tokens":5,"text":["pong"],"error":"null"}

Parameters bedrock-runtime rejected at the first tip (background, remote input_image)

  1. Three bodies on POST /v1/responses with model: bedrock-gpt-5.6-sol: {"input":"Reply with the single word pong","background":true}, an input_image whose image_url is https://avatars.githubusercontent.com/u/121462774?s=64&v=4 with the text What shape is this image? One word., and an input_image carrying a 1x1 red PNG as a data: URI with the text What color is this image? One word., each run against the merge-base proxy (port 49115) and this tip (port 38534)
  2. Output at the merge base (the Converse bridge answered all three):
before background HTTP 200 {"status":"completed","text":["pong"],"error":null}
before input_image_https HTTP 200 {"status":"completed","text":["Grapes"],"error":null}
before input_image_data HTTP 200 {"status":"completed","text":["Red"],"error":null}
  1. Output at e4b816e, the first tip, where the native endpoint rejected the first two (from the parameter sweep that found them; the image URL there was a wikimedia one):
after background HTTP 400 {"status":null,"text":"","items":[],"err":"litellm.BadRequestError: BedrockException - {\"error\":{\"message\":\"The background parameter is not supported.\",\"type\":\"invalid_request_error\",\"param\":null,\"code\":"}
after input_image_url HTTP 400 {"status":null,"text":"","items":[],"err":"litellm.BadRequestError: BedrockException - {\"error\":{\"message\":\"unsupported image_url scheme: must be `data:` or `s3://`; got https://upload.wikimedia.org/wiki"}
  1. Output at this tip:
after background HTTP 200 {"status":"completed","text":["pong"],"error":null}
after input_image_https HTTP 200 {"status":"completed","text":["Grapes"],"error":null}
after input_image_data HTTP 200 {"status":"completed","text":["Red"],"error":null}

A remote image inside a function_call_output

  1. One body on POST /v1/responses with model: bedrock-gpt-5.6-sol, a fetch_avatar function tool, and the history user -> function_call -> function_call_output where the tool output is the list [{"type":"input_text","text":"Here is the avatar."},{"type":"input_image","image_url":"https://avatars.githubusercontent.com/u/121462774?s=64&v=4"}] (probe/toolimg/run.sh), run against the merge-base proxy (port 49115), the previous tip, and this tip (port 38534)
  2. Output (the Converse bridge never served this shape either; bugbot raised it at the previous tip):
21:13:34Z before tool_output_image_https HTTP 400 {"status":null,"text":[],"error":"litellm.BadRequestError: BedrockException - {\"message\":\"This model doesn't support the image field for user messages. Remove image and try again.\"}
21:13:34Z 0ed4388 tool_output_image_https HTTP 400 {"status":null,"text":[],"error":"litellm.BadRequestError: BedrockException - {\"error\":{\"message\":\"unsupported image_url scheme: must be `data:` or `s3://`; got https://avatars.githubusercontent.com/u/121462774?s=64&v=4\",\"type\":\"inval"}
21:23:41Z after tool_output_image_https HTTP 200 {"status":"completed","text":["Grapes"],"error":"null"}
22:02:41Z after tool_output_image_https HTTP 200 {"status":"completed","text":["Grapes"],"error":"null"}

Computer use: the bare computer tool and a screenshot round trip

  1. Three bodies on POST /v1/responses with model: bedrock-gpt-5.6-sol and tools: [{"type":"computer"}] (probe/computer/roundtrip.sh): the user turn Take a screenshot of the screen. Use the computer tool., then that turn plus the returned reasoning and computer_call items plus a computer_call_output whose output is {"type":"computer_screenshot","image_url":...}, once with the https avatar URL above and once with the 1x1 red PNG data: URI
  2. Output at the merge base (the Converse bridge drops the computer tool, so the model answers in prose, and it 400s on any computer_call_output):
before tool_only HTTP 200 {"status":"completed","out_types":["reasoning","message"],"actions":null,"text":"I can’t open a browser or capture screenshots in this environment.","error":null}
before roundtrip_https HTTP 400 {"status":null,"out_types":[],"actions":null,"text":"","error":"{'message': 'litellm.BadRequestError: BedrockException - {\"message\":\"Expected toolResult blocks at messages.0.content for the following Ids: call_lit7790_cu\"}\\n\\nLiteLLM:"}
  1. Output at the previous tip, 63868db (the computer tool reaches bedrock-runtime, the data: screenshot round-trips, the https screenshot is the one image shape the inlining missed):
63868db tool_only HTTP 200 {"status":"completed","out_types":["reasoning","computer_call"],"actions":[{"type":"screenshot"}],"text":"","error":null}
63868db roundtrip_https HTTP 400 {"status":null,"out_types":[],"actions":null,"text":"","error":"{'message': 'litellm.BadRequestError: BedrockException - {\"error\":{\"message\":\"unsupported image_url scheme: must be `data:` or `s3://`; got https://avatars.githubusercont"}
63868db roundtrip_data HTTP 200 {"status":"completed","out_types":["reasoning","computer_call"],"actions":[{"type":"screenshot"}],"text":"","error":null}
  1. Output at this tip (port 38534):
after tool_only HTTP 200 {"status":"completed","out_types":["reasoning","computer_call"],"call_id":"call_d2302292e0405c3fa7013ea421cc512e","actions":[{"type":"screenshot"}],"text":"","err":"null"}
after roundtrip_https HTTP 200 {"status":"completed","out_types":["reasoning","computer_call"],"call_id":"call_57170209417d5ba8991711b0db960666","actions":[{"type":"screenshot"}],"text":"","err":"null"}
after roundtrip_data HTTP 200 {"status":"completed","out_types":["reasoning","computer_call"],"call_id":"call_888860a9a1a152028a9b98d0be296af6","actions":[{"type":"screenshot"}],"text":"","err":"null"}

Observations from the run:

  • First call with a new key pays cache-write premium; unchanged
  • Spend rows show cache_read_input_tokens null on both legs; unchanged
  • Stream events name the outer alias, inner model raw; unchanged
  • Native stream drops text_tokens details, adds sequence numbers; caused
  • Codex warns about missing model metadata for the alias; unchanged
  • Codex's web_search tool now dropped with a proxy warning; caused
  • background: true now dropped with a proxy warning; caused
  • Remote input_image URLs now inlined as data URIs; caused
  • Remote input_image URLs inside function_call_output lists inlined too, a 400 on both legs before; caused
  • Remote computer_call_output screenshots inlined too, a 400 at the previous tip; caused
  • Same prompt_cache_key stays warm across proxy restarts; unchanged
  • supports_prompt_caching null on gpt-5.6 price rows; unchanged
  • temperature and top_p under the Codex user agent dropped on both legs, 200 each; unchanged
  • Both legs equally slow under a loaded host (up to 213 s before, 180 s after, 45 s on chat); unchanged

Type

🆕 New Feature

Caveats (if any)

Low

  • Nested additional_tools stay invisible to metadata.allowed_tools on every bridged provider; pre-existing on main, untouched here (from feat(bedrock): serve the OpenAI models on bedrock-runtime's native Responses API #38489's notes)
  • The Codex history normalizer lives in litellm/llms/base_llm/responses/codex_compat.py while the tool hoist lives in litellm/responses/additional_tools.py; consolidating them is a small follow-up
  • The native stream's response.completed usage has no text_tokens sub-fields and the events carry a numeric sequence_number, where the bridge synthesized text_tokens and no sequence numbers; this is the OpenAI wire shape passing through, as it does for openai/ models today
  • Codex's web_search tool is dropped with a proxy-log warning on the native path; the bridge dropped it silently, so nothing a Codex user sees changes
  • A client sending the older computer_use_preview tool type gets it dropped with the same proxy-log warning at this tip (200, the model answers in prose that it cannot take screenshots) where the bridge surfaced bedrock-runtime's 400 listing the supported types; that is the drop-unsupported-tools decision above, the bare computer tool GPT-5.6 and GPT-6 use works end to end (screenshot round trips in the After section), and raising instead would 400 Codex's web_search case too, so left as is
  • supports_prompt_caching is unset on the gpt-5.6 price rows on both legs while the cost headers already price cache reads from the usage block; a price-map follow-up, not this PR's surface
  • Spend rows carry cache_read_input_tokens: null on both legs even when the response reports cached_tokens: 2174; pre-existing spend attribution, unchanged here
  • The first call with a new prompt_cache_key bills the cache-write premium (0.0120758 vs 0.00107536 warm) on both legs, as Bedrock's implicit caching already did
  • include: ["reasoning.encrypted_content"] with reasoning.summary: auto is accepted and returns no reasoning item for these short prompts on both legs
  • A deployment whose model is an application inference profile ARN stays on the bridge unless model_info: {supported_endpoints: ["/v1/responses"]} is set on it, and with that opt-in bedrock-runtime answers 400 Application inference profiles are not supported by this API., so the opt-in is for foundation model ids only; verified on two profiles in the CI account at e4b816e
  • GitHub Actions rust-wheel (LiteLLM Rust, not a required check) failed at 63868db in mypy.stubtest on litellm.rust_bridge._native.embedding, whose stub names its parameters request/args/kwargs while the compiled runtime names them _request/_args/_kwargs; that stub landed on main in feat(embeddings): add native dispatch foundation #42799 (2b3a7f7, 2026-09-23 21:11Z) and main's own run 35920843300 at that commit fails the same way, the PR touches no Rust or stub file, and the job runs here only because the workflow watches litellm/utils.py; the rerun from failed (35921645736) failed the same way
  • CircleCI using_litellm_on_windows failed at 0ed4388 (pipeline 90204, job 2208523) at 63868db (pipeline 90212, job 2208886), and at b2d53f9 (pipeline 90215, job 2209006, 683 s) on the Install Dependencies step (650 s timeout) before any test ran; the same step timed out on main at 2026-09-23 18:11Z (job 2208240) and in 4 of main's last 7 runs, so it is a main flake, not this PR
  • CircleCI local_testing_part1 failed at 63868db (pipeline 90212, job 2208880) on test_get_model_info_bedrock_region alone; the test passes alone and as a whole file at both that tip and the merge base, the failure is the _cached_get_model_info_helper lru_cache seeing another test's entry first under xdist, and main's job 2205275 (2026-09-23 06:09Z) failed on exactly that one test, so it is a main flake, not this PR
  • CircleCI integration-accounting failed at 0ed4388 (pipeline 90204) on tests/integration/spend/test_daily_rollup_retry.py::test_failed_daily_user_rollup_commit_is_retried_so_spend_report_and_daily_activity_agree with the fake upstream receiving a GET where the harness asserts POST; the PR touches no spend or proxy code, the test is unchanged since test(integration): regression tests for August cost tracking and budgeting bugs #42622 added it, and the whole integration workflow passed at 63868db (pipeline 90212), so it was the flake main's last 15 passing runs said it was

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

  • e4b816e passes /live-pr-risk

  • 0ed4388 passes /live-pr-risk

  • 63868db passes /live-pr-risk

  • b2d53f9 passes /live-pr-risk

leonardofreitass and others added 3 commits September 23, 2026 10:47
…sponses API

AWS serves the OpenAI models on bedrock-runtime through an OpenAI-compatible
surface at /openai/v1/responses, alongside Converse. LiteLLM had no Responses
config for the bedrock provider, so /v1/responses fell back to the Chat
Completions bridge and was translated into Converse. A realistic Codex session
does not survive that translation: its function_call / function_call_output
history becomes Converse toolUse / toolResult blocks with no toolConfig, and
Converse rejects the request outright.

Add a Responses config for that surface, opted into per model from the price-map
supported_endpoints so models without the signal keep the bridge exactly as
before. Auth is Bearer when a Bedrock API key is present, SigV4 otherwise.

Both Bedrock endpoints reject the Codex history item types agent_message,
context_compaction and local_shell_call, so the normalization bedrock_mantle
carried privately moves into a shared module and both providers use it. They are
history items, so they only bite from the second turn onward -- a first-turn
smoke test passes and hides the problem. Verified against bedrock-runtime with
global.openai.gpt-5.6-sol: additional_tools is accepted there (unlike on
bedrock-mantle) while those three types are rejected, so the two endpoints do
not share one validator and each provider opts in explicitly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…n suffix

get_complete_url hardcoded amazonaws.com in an f-string, so every non-commercial
partition got the wrong host: cn-north-1 resolved to amazonaws.com instead of
amazonaws.com.cn, and GovCloud/ISO regions were wrong the same way. Defer to
BaseAWSLLM._select_default_endpoint_url, which this config already inherits and
which resolves the suffix per partition.

test_no_fstring_hardcodes_the_commercial_dns_suffix scans the whole tree, so it
caught this even though it is not one of this PR's test files. Register the
config in ENDPOINT_BUILDERS so the cn/GovCloud endpoint sweep covers this
surface from now on rather than only the f-string guard.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@devin-ai-integration

devin-ai-integration Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@codecov

codecov Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.29730% with 6 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/llms/bedrock/common_utils.py 66.66% 4 Missing ⚠️
litellm/llms/bedrock/responses/transformation.py 98.46% 2 Missing ⚠️

📢 Thoughts on this report? Let us know!

@greptile-apps

greptile-apps Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with no outstanding correctness, security, or repository-rule findings

Findings

  1. P2 Provider dispatch outside llms ▶

Summary

This PR routes opted-in OpenAI models on Bedrock through the native Responses API while preserving existing fallback behavior

  • Adds capability-driven selection, Bedrock URL and authentication handling, Codex history normalization, unsupported parameter filtering, and guarded remote-image inlining
  • Reuses the asynchronous Responses transformation hook so remote images do not block async requests
  • Extends model metadata and regression coverage for routing, authentication, custom endpoints, tools, images, streaming, and AWS partitions
  • The latest changes add remote computer_screenshot inlining through the same guarded image-fetch path used for other image blocks

Reviews (7) · Last reviewed commit: "fix(bedrock): inline remote computer scr..."

Comment thread litellm/utils.py
Comment on lines +9077 to +9081
elif litellm.LlmProviders.BEDROCK == provider:
# bedrock-runtime serves the OpenAI models on an OpenAI-compatible surface
# (/openai/v1/responses) alongside Converse. The adapter decides whether a
# given model is on it; None keeps the chat-completions bridge.
return litellm.BedrockOpenAIResponsesConfig.for_model(model)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Provider dispatch outside llms

Repository rules require provider-specific behavior to stay under llms/. This Bedrock dispatch in shared utils.py must move before merging

Rule Used: What: Avoid writing provider-specific code outside of the llms/ directory. Why: This practice ensures better maintainability and reduces complexity over time. Good: ```python # Handle provider-specific logic within llms/vertex_ai/transformation.py ... (source)

Knowledge Base Used: Provider adapters and capabilities

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Every provider's Responses config is picked in this one dispatcher, Mantle and Azure AI included. The Bedrock-specific decision lives in llms/bedrock/responses via for_model

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right. This dispatcher is the established central mapping from provider to Responses config (including Mantle and Azure AI), while the Bedrock-specific capability decision and model handling are correctly encapsulated in llms/bedrock/responses/transformation.py via for_model. The added branch is appropriate, and my “provider dispatch outside llms” comment does not apply here.

@codspeed

codspeed Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_bedrock_openai_native_responses (b2d53f9) with main (170eb7f)

Open in CodSpeed

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

Codex sends a web_search tool on every turn. api.openai.com runs that tool
itself, and the Converse bridge dropped it silently, but bedrock-runtime's
native Responses endpoint rejects the whole request with 400 "web search is
not supported for this request". Filter the request's tools down to the
types bedrock-runtime's own validation error names, logging what was dropped,
through a helper shared with the Mantle route, which already did the same.
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 23, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/llms/bedrock/responses/transformation.py
Comment thread litellm/llms/bedrock/responses/transformation.py
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 23, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/llms/bedrock/responses/transformation.py
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 23, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 23, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit b2d53f9. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants