Repository navigation
feat(bedrock): serve the OpenAI models on bedrock-runtime's native Responses API (internal copy of #38489) - #42767
Conversation
…sponses API AWS serves the OpenAI models on bedrock-runtime through an OpenAI-compatible surface at /openai/v1/responses, alongside Converse. LiteLLM had no Responses config for the bedrock provider, so /v1/responses fell back to the Chat Completions bridge and was translated into Converse. A realistic Codex session does not survive that translation: its function_call / function_call_output history becomes Converse toolUse / toolResult blocks with no toolConfig, and Converse rejects the request outright. Add a Responses config for that surface, opted into per model from the price-map supported_endpoints so models without the signal keep the bridge exactly as before. Auth is Bearer when a Bedrock API key is present, SigV4 otherwise. Both Bedrock endpoints reject the Codex history item types agent_message, context_compaction and local_shell_call, so the normalization bedrock_mantle carried privately moves into a shared module and both providers use it. They are history items, so they only bite from the second turn onward -- a first-turn smoke test passes and hides the problem. Verified against bedrock-runtime with global.openai.gpt-5.6-sol: additional_tools is accepted there (unlike on bedrock-mantle) while those three types are rejected, so the two endpoints do not share one validator and each provider opts in explicitly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…n suffix get_complete_url hardcoded amazonaws.com in an f-string, so every non-commercial partition got the wrong host: cn-north-1 resolved to amazonaws.com instead of amazonaws.com.cn, and GovCloud/ISO regions were wrong the same way. Defer to BaseAWSLLM._select_default_endpoint_url, which this config already inherits and which resolves the suffix per partition. test_no_fstring_hardcodes_the_commercial_dns_suffix scans the whole tree, so it caught this even though it is not one of this PR's test files. Register the config in ENDPOINT_BUILDERS so the cn/GovCloud endpoint sweep covers this surface from now on rather than only the f-string guard. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
|
| elif litellm.LlmProviders.BEDROCK == provider: | ||
| # bedrock-runtime serves the OpenAI models on an OpenAI-compatible surface | ||
| # (/openai/v1/responses) alongside Converse. The adapter decides whether a | ||
| # given model is on it; None keeps the chat-completions bridge. | ||
| return litellm.BedrockOpenAIResponsesConfig.for_model(model) |
There was a problem hiding this comment.
Provider dispatch outside llms
Repository rules require provider-specific behavior to stay under llms/. This Bedrock dispatch in shared utils.py must move before merging
Rule Used: What: Avoid writing provider-specific code outside of the llms/ directory. Why: This practice ensures better maintainability and reduces complexity over time. Good: ```python # Handle provider-specific logic within llms/vertex_ai/transformation.py ... (source)
Knowledge Base Used: Provider adapters and capabilities
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
Every provider's Responses config is picked in this one dispatcher, Mantle and Azure AI included. The Bedrock-specific decision lives in llms/bedrock/responses via for_model
There was a problem hiding this comment.
You're right. This dispatcher is the established central mapping from provider to Responses config (including Mantle and Azure AI), while the Bedrock-specific capability decision and model handling are correctly encapsulated in llms/bedrock/responses/transformation.py via for_model. The added branch is appropriate, and my “provider dispatch outside llms” comment does not apply here.
Codex sends a web_search tool on every turn. api.openai.com runs that tool itself, and the Converse bridge dropped it silently, but bedrock-runtime's native Responses endpoint rejects the whole request with 400 "web search is not supported for this request". Filter the request's tools down to the types bedrock-runtime's own validation error names, logging what was dropped, through a helper shared with the Mantle route, which already did the same.
|
bugbot run |
|
bugbot run |
…native Responses route
|
bugbot run |
|
bugbot run |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit b2d53f9. Configure here.
TLDR
Internal copy of #38489 (its two commits cherry-picked with the author preserved) so the work can be driven to merge; the fork PR does not allow maintainer edits
Problem this solves:
/v1/responsesthrough the Converse bridgeprompt_cache_keygot a 400 from that bridgeHow it solves it:
/openai/v1/responsessupported_endpoints: ["/v1/responses"]on the rowweb_search) with a warningfile_searchon the emulated path (bugbot finding)/v1/responsesor/responsesno longer double the path (bugbot finding)backgroundwith a proxy-log warning, since bedrock-runtime rejects it and the bridge never forwarded ithttp(s)image URLs (input_imageblocks in message content and infunction_call_outputlists, andcomputer_screenshotoutputs of acomputer_call_output) and inlines them as data URIs, since bedrock-runtime accepts onlydata:ands3://; the bridge downloaded the message-content ones and 400'd on both tool-output shapes (bugbot finding for the list shape, the /caveats pass for the screenshot dict)The gpt-6 opt-in is a decision made here: all six
us./global.gpt-6 ids answered 200 on the native endpoint with both auth kinds, and the alternative (leaving them on the bridge) keeps the 400 above for the newest models. Anything not flagged in the price map keeps the bridge exactly as todayTwo more decisions made here, both found by probing the native endpoint with every Responses parameter at the first tip:
background: trueis dropped unconditionally with a proxy-log warning, matching the bridge, which never forwarded it either, and the alternative (raising unlessdrop_paramsis set) keeps a 400 on a request the bridge answered 200. Remotehttp(s)image URLs,input_imageblocks in message content and infunction_call_outputlists andcomputer_screenshotoutputs alike, are fetched and inlined as data URIs through the sharedimage_handlinghelpers (the same SSRF policy and size cap every other inlining provider uses), with the sync and async fetchers injected through the config constructor so tests never touch the network; the alternative (documenting the 400) breaks a message-content request the bridge answered 200, and both tool-output shapes were a 400 on the bridge too (Converse rejects an image inside a tool result for this model and never forwards acomputer_call_outputat all), so inlining them is new capability rather than parityUser Flow
Before: a developer sending
prompt_cache_keytobedrock/us.openai.gpt-5.6-solon/v1/responsesgets a 400, and only Codex gets through because the gateway drops the parameter for itbedrock-gpt-5.6-sol(bedrock/us.openai.gpt-5.6-sol) to the gateway config and create a virtual keymodel: bedrock-gpt-5.6-sol, a longinstructionsstring, a shortinput, andprompt_cache_key: "my-session"litellm.UnsupportedParamsError: bedrock does not support parameters: ['prompt_cache_key'], for model=us.openai.gpt-5.6-solwire_api = "responses"), and it comes back 200 only because the gateway silently droppedprompt_cache_keyfor that user agent, so the key never reached Bedrocksupported_endpoints: nullAfter: the same request comes back 200 from Bedrock's native Responses endpoint, and the repeat call reports the cached tokens
bedrock-gpt-5.6-sol(bedrock/us.openai.gpt-5.6-sol) to the gateway config and create a virtual keymodel: bedrock-gpt-5.6-sol, a longinstructionsstring, a shortinput, andprompt_cache_key: "my-session"usage.input_tokens_details.cache_write_tokensabove 0usage.input_tokens_details.cached_tokensabove 0 and a lowerx-litellm-response-costheader; the same request from Codex CLI comes back 200 with the parameter forwarded rather than droppedsupported_endpoints: ["/v1/responses"]Relevant issues
Copy of #38489. Related: #29818, #36182
Affected release
Linear ticket
Resolves LIT-7790
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Shared setup for both legs: one proxy per leg, each booted with
--num_workers 2on its own random port with its own Postgres database, the same config on both. The config listsbedrock-gpt-5.6-sol(bedrock/us.openai.gpt-5.6-sol, us-east-1, SigV4 keys),bedrock-gpt-5.6-luna-global(bedrock/global.openai.gpt-5.6-luna),bedrock-gpt-5.6-sol-bearer(same model with a Bedrock API key asapi_key),bedrock-gpt-6-sol(bedrock/us.openai.gpt-6-sol),bedrock-claude-haiku(bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0, not opted in) andbedrock-mantle-gpt-5.6-sol(bedrock_mantle/openai.gpt-5.6-sol). Every call hits the real Bedrock endpoints.$KEYis a virtual key created on that leg,$MASTERits master key,long_instructions.txtis an 11 KB instructions string (about 2,100 tokens). The Before proxy is on port 49115, the After proxy on port 38534 (each tip's proxy got a fresh port; the port in each After block is the one that run used). Codex CLI 0.155.1 runs interactively in tmux with~/.codex/config.tomlpointingbase_urlat the proxy,model = "bedrock-gpt-5.6-sol",wire_api = "responses", in a directory holdingnotes.txt(2 lines) andnumbers.txt(5 lines)Before (e73f949)
Codex CLI session (shell tool round trip)
Run ls in this directory, then tell me the total number of lines across notes.txt and numbers.txtlsandwc -l, then printsTotal lines across both files: 7.The session works on the base only because the gateway auto-enablesdrop_paramsfor Codex's user agent, so theprompt_cache_keyCodex sends is thrown away before the Converse bridge sees itprompt_cache_keyon POST /v1/responses, three clientsThe same request from the OpenAI Python SDK, twice
openai2.33.0:OpenAI(base_url="http://127.0.0.1:49115", api_key=KEY).responses.create(model="bedrock-gpt-5.6-sol", instructions=<the 11 KB instructions>, input="Reply with the single word pong", prompt_cache_key="lit7790-sdk"), called twice two seconds apart, printingstatus,output_text, andusage.input_tokens_detailsprompt_cache_keywith a long prompt, twice (native_params.json: 11 KBinstructions,prompt_cache_key: "lit7790-qa-session",store: false,include: ["reasoning.encrypted_content"],reasoning: {effort: low, summary: auto})Same long prompt without
prompt_cache_key, twice (resp_cache.json, Bedrock's implicit caching)-d @resp_cache.json, twice, plusgrep -i x-litellm-response-cost: hdr.txtStreaming
Tool round trip over raw HTTP (
tool_history.json: user turn, afunction_calltoshell, itsfunction_call_outputlisting two files, onefunctiontool)curl -s -o out.json -w 'HTTP %{http_code}\n' -X POST http://127.0.0.1:49115/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d @tool_history.jsonOther model ids and auth kinds
for m in bedrock-gpt-5.6-luna-global bedrock-gpt-5.6-sol-bearer bedrock-gpt-6-sol; do curl -s -o out.json -w "$m HTTP %{http_code} " -X POST http://127.0.0.1:49115/v1/responses -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d "{\"model\":\"$m\",\"input\":\"Reply with the single word pong\"}"; jq -c '{status, input_tokens: .usage.input_tokens, output_tokens: .usage.output_tokens, text: [.output[]? | select(.type=="message") | .content[]?.text]}' out.json; doneSpend rows for the two long-prompt calls
curl -s "http://127.0.0.1:49115/spend/logs?request_id=<x-litellm-call-id>" -H "Authorization: Bearer $MASTER" | jq -c '[.[] | {model, custom_llm_provider, call_type, spend, prompt_tokens, completion_tokens, cache_read_input_tokens, total_tokens}]'for the twox-litellm-call-idheaders of the implicit-caching pair (theprompt_cache_keycalls were 400s, so they have no rows), thencurl -s http://127.0.0.1:49115/key/info -H "Authorization: Bearer $KEY" | jq -c '{alias: .info.key_alias, spend: .info.spend}'Chat completions on the same model
curl -s -o out.json -w 'HTTP %{http_code}\n' -X POST http://127.0.0.1:49115/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{"model":"bedrock-gpt-5.6-sol","messages":[{"role":"user","content":"Reply with the single word pong"}]}'; jq -c '{model, text: .choices[0].message.content, usage: {prompt_tokens: .usage.prompt_tokens, completion_tokens: .usage.completion_tokens}}' out.jsonNeighbouring paths (the /live-pr-risk cases)
bedrock-claude-haiku(not opted in) on POST /v1/responses; B:bedrock-mantle-gpt-5.6-solon POST /v1/responses with the Codex-shaped tool history plus aweb_searchtool; B2: the mantle model withprompt_cache_key; E:bedrock-gpt-5.6-solon POST /v1/messages; F: GET /model/info; G:bedrock-gpt-5.6-solon POST /v1/responses with afile_searchtool naming a vector store that does not existAfter (b2d53f9)
Codex CLI session (shell tool round trip)
ls, then printsnotes.txt and numbers.txt contain 7 total lines.(done 3:04 PM, about five seconds after the prompt). The gateway forwards Codex'sprompt_cache_keyto the native endpoint and drops theweb_searchtool that endpoint rejects, logging a warning in the proxy logprompt_cache_keyon POST /v1/responses, three clientsprompt_cache_keywith a long prompt, twiceprompt_cache_key: "lit7790-qa-session", cold again after the proxy restart at this tip), plusgrep -i 'x-litellm-response-cost:' hdr.txtThe same request from the OpenAI Python SDK, twice
openai2.33.0: the same tworesponses.create(...)calls as the Before leg, against port 38534 withprompt_cache_key="lit7790-sdk-b2d53f9"Same long prompt without
prompt_cache_key, twice-d @resp_cache.jsonagainst port 38534Streaming
sequence_numberadded to the jq projectionTool round trip over raw HTTP
Other model ids and auth kinds
Spend rows for the two
prompt_cache_keycalls/spend/logs?request_id=and/key/infocalls against port 38534 for the twoprompt_cache_keycall idsChat completions on the same model
Neighbouring paths (the /live-pr-risk cases)
Parameters bedrock-runtime rejected at the first tip (
background, remoteinput_image)model: bedrock-gpt-5.6-sol:{"input":"Reply with the single word pong","background":true}, aninput_imagewhoseimage_urlishttps://avatars.githubusercontent.com/u/121462774?s=64&v=4with the textWhat shape is this image? One word., and aninput_imagecarrying a 1x1 red PNG as adata:URI with the textWhat color is this image? One word., each run against the merge-base proxy (port 49115) and this tip (port 38534)A remote image inside a
function_call_outputmodel: bedrock-gpt-5.6-sol, afetch_avatarfunction tool, and the historyuser -> function_call -> function_call_outputwhere the tool output is the list[{"type":"input_text","text":"Here is the avatar."},{"type":"input_image","image_url":"https://avatars.githubusercontent.com/u/121462774?s=64&v=4"}](probe/toolimg/run.sh), run against the merge-base proxy (port 49115), the previous tip, and this tip (port 38534)Computer use: the bare
computertool and a screenshot round tripmodel: bedrock-gpt-5.6-solandtools: [{"type":"computer"}](probe/computer/roundtrip.sh): the user turnTake a screenshot of the screen. Use the computer tool., then that turn plus the returnedreasoningandcomputer_callitems plus acomputer_call_outputwhoseoutputis{"type":"computer_screenshot","image_url":...}, once with the https avatar URL above and once with the 1x1 red PNGdata:URIcomputertool, so the model answers in prose, and it 400s on anycomputer_call_output):computertool reaches bedrock-runtime, thedata:screenshot round-trips, the https screenshot is the one image shape the inlining missed):Observations from the run:
background: truenow dropped with a proxy warning; causedType
🆕 New Feature
Caveats (if any)
Low
additional_toolsstay invisible tometadata.allowed_toolson every bridged provider; pre-existing on main, untouched here (from feat(bedrock): serve the OpenAI models on bedrock-runtime's native Responses API #38489's notes)litellm/llms/base_llm/responses/codex_compat.pywhile the tool hoist lives inlitellm/responses/additional_tools.py; consolidating them is a small follow-upresponse.completedusage has notext_tokenssub-fields and the events carry a numericsequence_number, where the bridge synthesizedtext_tokensand no sequence numbers; this is the OpenAI wire shape passing through, as it does foropenai/models todayweb_searchtool is dropped with a proxy-log warning on the native path; the bridge dropped it silently, so nothing a Codex user sees changescomputer_use_previewtool type gets it dropped with the same proxy-log warning at this tip (200, the model answers in prose that it cannot take screenshots) where the bridge surfaced bedrock-runtime's 400 listing the supported types; that is the drop-unsupported-tools decision above, the barecomputertool GPT-5.6 and GPT-6 use works end to end (screenshot round trips in the After section), and raising instead would 400 Codex'sweb_searchcase too, so left as issupports_prompt_cachingis unset on the gpt-5.6 price rows on both legs while the cost headers already price cache reads from the usage block; a price-map follow-up, not this PR's surfacecache_read_input_tokens: nullon both legs even when the response reportscached_tokens: 2174; pre-existing spend attribution, unchanged hereprompt_cache_keybills the cache-write premium (0.0120758 vs 0.00107536 warm) on both legs, as Bedrock's implicit caching already didinclude: ["reasoning.encrypted_content"]withreasoning.summary: autois accepted and returns no reasoning item for these short prompts on both legsmodelis an application inference profile ARN stays on the bridge unlessmodel_info: {supported_endpoints: ["/v1/responses"]}is set on it, and with that opt-in bedrock-runtime answers 400Application inference profiles are not supported by this API., so the opt-in is for foundation model ids only; verified on two profiles in the CI account at e4b816erust-wheel(LiteLLM Rust, not a required check) failed at 63868db inmypy.stubtestonlitellm.rust_bridge._native.embedding, whose stub names its parametersrequest/args/kwargswhile the compiled runtime names them_request/_args/_kwargs; that stub landed on main in feat(embeddings): add native dispatch foundation #42799 (2b3a7f7, 2026-09-23 21:11Z) and main's own run 35920843300 at that commit fails the same way, the PR touches no Rust or stub file, and the job runs here only because the workflow watcheslitellm/utils.py; the rerun from failed (35921645736) failed the same wayusing_litellm_on_windowsfailed at 0ed4388 (pipeline 90204, job 2208523) at 63868db (pipeline 90212, job 2208886), and at b2d53f9 (pipeline 90215, job 2209006, 683 s) on the Install Dependencies step (650 s timeout) before any test ran; the same step timed out on main at 2026-09-23 18:11Z (job 2208240) and in 4 of main's last 7 runs, so it is a main flake, not this PRlocal_testing_part1failed at 63868db (pipeline 90212, job 2208880) ontest_get_model_info_bedrock_regionalone; the test passes alone and as a whole file at both that tip and the merge base, the failure is the_cached_get_model_info_helperlru_cache seeing another test's entry first under xdist, and main's job 2205275 (2026-09-23 06:09Z) failed on exactly that one test, so it is a main flake, not this PRintegration-accountingfailed at 0ed4388 (pipeline 90204) ontests/integration/spend/test_daily_rollup_retry.py::test_failed_daily_user_rollup_commit_is_retried_so_spend_report_and_daily_activity_agreewith the fake upstream receiving a GET where the harness asserts POST; the PR touches no spend or proxy code, the test is unchanged since test(integration): regression tests for August cost tracking and budgeting bugs #42622 added it, and the whole integration workflow passed at 63868db (pipeline 90212), so it was the flake main's last 15 passing runs said it wasFinal Attestation
The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
e4b816e passes /live-pr-risk
0ed4388 passes /live-pr-risk
63868db passes /live-pr-risk
b2d53f9 passes /live-pr-risk