Repository navigation
feat(bedrock): serve the OpenAI models on bedrock-runtime's native Responses API - #38489
leonardofreitass wants to merge 2 commits into
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
65f31c9 to
485f70d
Compare
Greptile SummaryThis PR adds a native OpenAI Responses adapter for Bedrock Runtime, selects it through model capability metadata, and shares Codex history normalization with Bedrock Mantle.
Confidence Score: 3/5This PR is not safe to merge until the stated authorization dependency is present; the provider-specific selection logic should also be moved under the Bedrock adapter. The native route currently permits nested tool declarations to reach Bedrock without being included in the existing allowed-tool evaluation, while the shared utility layer also takes on Bedrock-specific capability policy. Files Needing Attention: litellm/utils.py and litellm/llms/bedrock/responses/transformation.py
|
| Filename | Overview |
|---|---|
| litellm/llms/bedrock/responses/transformation.py | Introduces the native Bedrock Runtime Responses adapter, endpoint construction, authentication, and Codex history normalization. |
| litellm/llms/base_llm/responses/codex_compat.py | Extracts Codex history-item normalization into a shared pure transformation with focused tests. |
| litellm/llms/bedrock_mantle/responses/transformation.py | Reuses the shared normalizer while preserving Mantle-specific tool hoisting and warning behavior. |
| litellm/utils.py | Selects the native adapter through capability metadata, but exposes the acknowledged authorization dependency and places provider-specific policy in shared code. |
| litellm/llms/bedrock/common_utils.py | Adds metadata-driven detection of Bedrock models supporting the native Responses endpoint. |
| model_prices_and_context_window.json | Opts six regional GPT-5.6 profiles into the native Responses surface. |
Reviews (1): Last reviewed commit: "feat(bedrock): serve the OpenAI models o..." | Re-trigger Greptile
485f70d to
b849db9
Compare
|
@leonardofreitass Thanks for adding the native Bedrock Runtime
Our main question is whether GPT-5.6 Luna prompt caching works end-to-end through this adapter. According to the Amazon Bedrock documentation, GPT-5.6 Luna supports explicit prompt caching with:
Documentation: For a stable prefix of at least 1,024 tokens, could you confirm whether the following request structure is supported and transmitted unchanged to the Bedrock Runtime Responses endpoint? {
"model": "global.openai.gpt-5.6-luna",
"prompt_cache_key": "hindsight:retain:stable-prefix-v1",
"prompt_cache_options": {
"mode": "explicit",
"ttl": "30m"
},
"input": [
{
"type": "message",
"role": "developer",
"content": [
{
"type": "input_text",
"text": "<stable prefix of at least 1,024 tokens>",
"prompt_cache_breakpoint": {
"mode": "explicit"
}
}
]
},
{
"type": "message",
"role": "user",
"content": [
{
"type": "input_text",
"text": "<dynamic retain content>"
}
]
}
]
}In particular, for global.openai.gpt-5.6-luna, is the following behavior supported and tested?
If explicit prompt-cache breakpoint handling, |
b849db9 to
b6f84cb
Compare
|
@hil00137 Thanks for the detailed question. I ran your exact request structure against real Bedrock Runtime rather than answering from the code, so the numbers below are from billed calls.
All nine behaviors are supported. Nothing in this PR needs to change for prompt caching to work.
Measured usage for 5, 6 and 7: I also ran three controls, since a cache hit on its own does not prove the fields did the work:
Two caveats, both pre-existing and outside this PR:
To your closing question: none of this is deferred to a follow-up. It works on this branch as written. |
…sponses API AWS serves the OpenAI models on bedrock-runtime through an OpenAI-compatible surface at /openai/v1/responses, alongside Converse. LiteLLM had no Responses config for the bedrock provider, so /v1/responses fell back to the Chat Completions bridge and was translated into Converse. A realistic Codex session does not survive that translation: its function_call / function_call_output history becomes Converse toolUse / toolResult blocks with no toolConfig, and Converse rejects the request outright. Add a Responses config for that surface, opted into per model from the price-map supported_endpoints so models without the signal keep the bridge exactly as before. Auth is Bearer when a Bedrock API key is present, SigV4 otherwise. Both Bedrock endpoints reject the Codex history item types agent_message, context_compaction and local_shell_call, so the normalization bedrock_mantle carried privately moves into a shared module and both providers use it. They are history items, so they only bite from the second turn onward -- a first-turn smoke test passes and hides the problem. Verified against bedrock-runtime with global.openai.gpt-5.6-sol: additional_tools is accepted there (unlike on bedrock-mantle) while those three types are rejected, so the two endpoints do not share one validator and each provider opts in explicitly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
b6f84cb to
0d17346
Compare
…n suffix get_complete_url hardcoded amazonaws.com in an f-string, so every non-commercial partition got the wrong host: cn-north-1 resolved to amazonaws.com instead of amazonaws.com.cn, and GovCloud/ISO regions were wrong the same way. Defer to BaseAWSLLM._select_default_endpoint_url, which this config already inherits and which resolves the suffix per partition. test_no_fstring_hardcodes_the_commercial_dns_suffix scans the whole tree, so it caught this even though it is not one of this PR's test files. Register the config in ENDPOINT_BUILDERS so the cn/GovCloud endpoint sweep covers this surface from now on rather than only the f-string guard. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
@leonardofreitass The cache write/read results are exactly what we needed. We appreciate |
|
Thank you for this PR — it is the missing piece for Bedrock GPT-5.6 on bedrock-runtime. We hit the AWS validation error in production-shaped traffic:
On current LiteLLM, even Could maintainers please review and approve this as soon as you can? The One follow-up that would unlock most OpenAI-compatible proxies (we only call Happy to test a chat-completions follow-up against real Bedrock if that helps. |
…sponses API (internal copy of #38489) (#42767) * feat(bedrock): serve the OpenAI models on bedrock-runtime's native Responses API AWS serves the OpenAI models on bedrock-runtime through an OpenAI-compatible surface at /openai/v1/responses, alongside Converse. LiteLLM had no Responses config for the bedrock provider, so /v1/responses fell back to the Chat Completions bridge and was translated into Converse. A realistic Codex session does not survive that translation: its function_call / function_call_output history becomes Converse toolUse / toolResult blocks with no toolConfig, and Converse rejects the request outright. Add a Responses config for that surface, opted into per model from the price-map supported_endpoints so models without the signal keep the bridge exactly as before. Auth is Bearer when a Bedrock API key is present, SigV4 otherwise. Both Bedrock endpoints reject the Codex history item types agent_message, context_compaction and local_shell_call, so the normalization bedrock_mantle carried privately moves into a shared module and both providers use it. They are history items, so they only bite from the second turn onward -- a first-turn smoke test passes and hides the problem. Verified against bedrock-runtime with global.openai.gpt-5.6-sol: additional_tools is accepted there (unlike on bedrock-mantle) while those three types are rejected, so the two endpoints do not share one validator and each provider opts in explicitly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(bedrock): build the Responses endpoint from the region's partition suffix get_complete_url hardcoded amazonaws.com in an f-string, so every non-commercial partition got the wrong host: cn-north-1 resolved to amazonaws.com instead of amazonaws.com.cn, and GovCloud/ISO regions were wrong the same way. Defer to BaseAWSLLM._select_default_endpoint_url, which this config already inherits and which resolves the suffix per partition. test_no_fstring_hardcodes_the_commercial_dns_suffix scans the whole tree, so it caught this even though it is not one of this PR's test files. Register the config in ENDPOINT_BUILDERS so the cn/GovCloud endpoint sweep covers this surface from now on rather than only the f-string guard. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(bedrock): opt the gpt-6 family into the native Responses API * fix(bedrock): drop the Responses tool types bedrock-runtime rejects Codex sends a web_search tool on every turn. api.openai.com runs that tool itself, and the Converse bridge dropped it silently, but bedrock-runtime's native Responses endpoint rejects the whole request with 400 "web search is not supported for this request". Filter the request's tools down to the types bedrock-runtime's own validation error names, logging what was dropped, through a helper shared with the Mantle route, which already did the same. * fix(bedrock): emulate file_search and collapse custom Responses paths * fix(bedrock): keep background and remote image inputs working on the native Responses route * fix(bedrock): inline remote images inside tool outputs on the native Responses route * fix(bedrock): inline remote computer screenshots on the native Responses route --------- Co-authored-by: Leonardo Freitas dos Santos <leonardo.freitas.s@outlook.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
TLDR
Problem this solves:
bedrockhas no Responses config, so/v1/responsesrides the Converse bridgeHow it solves it:
/openai/v1/responsessurfacebedrock_mantleUser Flow
Before: a developer running Codex CLI through a LiteLLM gateway on a Bedrock GPT-5.6 model cannot get past the first tool call.
POST https://litellm-domain/v1/responseswith abedrock/global.openai.gpt-5.6-soldeploymentBedrockException - The toolConfig field must be defined when using toolUse and toolResult content blocksAfter: the same session continues normally.
POST https://litellm-domain/v1/responseswith the same deploymentRelevant issues
Related: #29818, #36182
Linear ticket
Pre-Submission checklist
uv run pytest tests/test_litellm/llms/bedrock/responses/ tests/test_litellm/llms/base_llm/responses/ tests/test_litellm/llms/bedrock_mantle/ -v(252 pass; 42 new, both new files at 100% patch coverage). The wider Responses suites pass too:tests/test_litellm/responses/ tests/test_litellm/llms/openai/responses/(1007 pass)upstream/main: format-check-changed, ruff (litellm + tests config), ruff-strict ratchet, type-discipline ratchet, test-quality ratchet, basedpyright (139808, no higher than base), circular-imports, import-safetyfor_model(), and the authorization dependency that prompted the second finding is no longer this PR's (see Caveats → Low).Screenshots / Proof of Fix
End-to-end against real Amazon Bedrock. No mocks. Multi-turn, because the item types at issue are history items — a first-turn request succeeds on both sides and hides the problem entirely.
Shared setup —
bedrock/global.openai.gpt-5.6-solinus-east-1, called vialitellm.aresponses, with a Codex-shaped history:Before (cd63c7e)
bedrock/global.openai.gpt-5.6-solBadRequestError: BedrockException - {"message":"The toolConfig field must be defined when using toolUse and toolResult content blocks."}toolUse/toolResultblocks with notoolConfigAfter (b6f84cb)
HTTP 200, output item types['message'], usagein=84 out=7'BLUEBIRD'— the prior turn survived, normalized into a supported item typeMeasured on
b6f84cb4e0, the commit before the rebase ontomain. The rebase changed onlybedrock_mantle, not the bedrock-runtime path this exercises, so the result stands for the current commit; I have cited the ref actually run rather than the tip.Prompt caching, verified end-to-end
Asked about in this comment. Explicit prompt caching works over this adapter unchanged. Two
litellm.aresponsescalls againstbedrock/global.openai.gpt-5.6-luna, sameprompt_cache_key, identical ~1.2K-token prefix carryingprompt_cache_breakpoint:prompt_cache_key,prompt_cache_optionsandprompt_cache_breakpointreach the wire unmodified — the normalizer returns non-Codex items by identity — and usage comes back onusage.input_tokens_details. Same result ongpt-5.6-sol. No code change was needed for this.Error responses keep the Bedrock request id
get_error_classroutes throughBedrockErrorrather than the OpenAI base, which builds a blank response. Verified against a real 400 from bedrock-runtime:x-amzn-RequestId: 3e832927-48fb-4231-abd2-03fe9ac92d18survives to the caller, where it was previously dropped.Type
🆕 New Feature
Caveats (if any)
Medium
/v1/responsesonlyglobal.openai.gpt-5.6-sol:additional_toolsis accepted there whileagent_message,context_compactionandlocal_shell_callare rejected with400 Invalid 'input': value did not match any expected variant. On bedrock-mantle,additional_toolsis rejected.Low
litellm_internal_stagingontomainand retargeted, sincemainis now the repository's default branch.bedrock_mantlewas the one conflict: fix(responses): hoist Codex additional_tools input items into the chat bridge tools #40989 replaced its private_hoist_codex_additional_toolswith the sharedhoist_additional_tools, which is the same method this PR was refactoring. Resolved by keeping that shared hoist and layering this PR's shared history normalizer on top, so the file now uses both shared helpers and carries neither copy of its own.extract_request_tool_namesto descend intoadditional_tools. That is no longer this PR's dependency. fix(responses): lift additional_tools input items into tools on the chat bridge #38388 was closed in favour of fix(responses): hoist Codex additional_tools input items into the chat bridge tools #40989, which landed the hoist but not the extractor change, so the extractor still reads only top-leveltoolswhile nested tools are hoisted into the live tool list — meaning a key restricted bymetadata.allowed_toolsdoes not see them. That is the state ofmainon the bridge route for every bridged provider, independent of this PR, which touches no file in that path. Probed on this branch: nested tools give[]where an equivalent top-level tool gives its name. Worth a separate issue; happy to open one, or to port the extractor half of fix(responses): lift additional_tools input items into tools on the chat bridge #38388 if a maintainer prefers it here.supported_endpointsis the only signal that selects this surface. An earlier note here referencedtest_bedrock_gpt_5_6_advertises_only_converse_supported_features, which asserted the field was absent; 3023497 removed it from the tree while this PR was open. The opt-in is covered directly by this PR's own suite instead, which reads the shipped map.cache_write_tokensreaches the caller throughextra="allow"onInputTokensDetailsrather than a declared field, so it passes through but is not part of a typed contract. Separately, these price-map entries do not setsupports_prompt_cache_breakpoint, so thecache_control→ breakpoint translation inAnthropicCacheControlHookdoes not apply to them; passingprompt_cache_breakpointdirectly does work, as measured above. Both are pre-existing and out of scope here — happy to set the flag if a maintainer wants it in this PR.litellm/responses/additional_tools.pyas the shared home for Codex tool hoisting, while this PR's history normalizer sits inlitellm/llms/base_llm/responses/codex_compat.py. Those two belong together; consolidating them is a small follow-up either way round.bedrock_mantlekeeps its exact warning wording. The shared normalizer is a pure transform returning the rewritten types, and each caller logs in its own words, so allbedrock_mantletests pass unchanged.Final Attestation
🤖 Generated with Claude Code