Repository navigation
Conversation
|
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 1 · PR risk: 0/10 |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
bugbot run |
…e /responses instead of Converse
170efd4 to
e2966ab
Compare
… bridge flag, honor aws_bedrock_runtime_endpoint on Responses
|
bugbot run |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 30c6687. Configure here.
| model, | ||
| {**passed_params, "_litellm_responses_api_bridge_allowed": _litellm_responses_api_bridge_allowed}, | ||
| additional_drop_params, | ||
| ) |
There was a problem hiding this comment.
Param mapping misses converse-only kwargs
Medium Severity
get_optional_params now decides the new responses vs converse route from passed_params only, so kwargs-only fields such as requestMetadata, guardrailConfig, extra_body.tools, and aws_bedrock_client never reach that decision. Dispatch in completion() still sees the full request and correctly stays on Converse, but mapping uses the chat-completions supported-params list. A tools-plus-reasoning call that also sends requestMetadata now raises UnsupportedParamsError or drops that metadata, even though Converse previously accepted it.
Additional Locations (2)
Reviewed by Cursor Bugbot for commit 30c6687. Configure here.
There was a problem hiding this comment.
pre_process_non_default_params writes every kwarg into passed_params in place (utils.py L4284) before L4592, so requestMetadata and guardrailConfig reach the route decision. Probe matched main


TLDR
Problem this solves:
reasoning_tokensfrom usage, so callers can't see the reasoning they pay forjson_schemaresponse_format when tools are presentHow it solves it:
/openai/v1/responses/v1/responsesin the cost mapthinking, legacyfunctions, deprecatedaws_bedrock_client) still go to Conversefile_search, also viaextra_body) stay on Converselitellm.responses(..., use_chat_completions_api=True)skips the bridge, so those calls keep Converseaws_bedrock_runtime_endpointnow wins overapi_baseon Bedrock /v1/responses, like Converseclient, matching the sync pathUser Flow
Before: a developer calling a Bedrock GPT-5.6+ model with tools and reasoning gets Converse's answer, with no reasoning in it
model: bedrock-gpt-6.1-sol, aget_weathertool, andreasoning_effort: highchat.completionwithtool_calls, an id likechatcmpl-..., andusage.completion_tokens_detailsempty, so they cannot see how many reasoning tokens they paid for"stream": trueandinclude_usageends with the same emptycompletion_tokens_detailsthinkingand tools answers through Converse tooAfter: the same requests are served by the model's native OpenAI Responses API
chat.completionwith the sametool_calls, an id likeresp_...,usage.completion_tokens_details.reasoning_tokensfilled in, and areasoning_itemsentry they can echo back on the next turnreasoning_tokensin the final usage chunkreasoning_effort: "none"still goes to native/openai/v1/chat/completions, and aguardrailConfigrequest still goes to Converse and is still blocked by the guardrailtool_useblock, without Converse'sredacted_thinkingblockRelevant issues
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Live run against real Bedrock (us-east-1, real spend) from a maintainer machine. Two proxies, one per commit, each booted with 2 uvicorn workers and no DB, with
LITELLM_LOCAL_MODEL_COST_MAP=True:python litellm/proxy/proxy_cli.py --config config.yaml --port $PORT --num_workers 2Shared shell variables (
$URLis the base proxy for Before and the head proxy for After, and$GRis a Bedrock guardrail whose topic policy denies coffee):Before (d6db8e8)
/v1/chat/completions, tools + reasoning, gpt-6.1-sol
Command
Output, no reasoning tokens reported
/v1/chat/completions, tools + reasoning, streaming
Command
Output
/v1/chat/completions, tools + reasoning, gpt-5.6-sol
Same as the first case with
"model":"bedrock-gpt-5.6-sol", printing{id, finish, completion_tokens_details}Output
/v1/chat/completions, tool result follow-up turn
Command
Output
Control, reasoning_effort none
Same as the first case with
"model":"bedrock-gpt-5.6-sol","reasoning_effort":"none"and the Paris questionOutput, served by native chat completions
Control, guardrailConfig
Command
Output, blocked by the guardrail on Converse
/v1/responses, tools + reasoning
Command
Output
/v1/messages, thinking + tools
Command
Output
After (30c6687)
/v1/chat/completions, tools + reasoning, gpt-6.1-sol
Same command as Before
Output, reasoning tokens reported and billed, plus an encrypted reasoning item to echo back
/v1/chat/completions, tools + reasoning, streaming
Same command as Before
Output
/v1/chat/completions, tools + reasoning, gpt-5.6-sol
Same command as Before
Output
/v1/chat/completions, tool result follow-up turn
Same command as Before
Output
Control, reasoning_effort none
Same command as Before
Output, still native chat completions
Control, guardrailConfig
Same command as Before
Output, still Converse and still blocked
/v1/responses, tools + reasoning
Same command as Before
Output, unchanged
/v1/messages, thinking + tools
Same command as Before
Output, same tool call, no
redacted_thinkingblockExtra checks on the head proxy, all answered 200 with the expected tool call:
description, and withstrict: trueget weather!(also streamed) stays on Converse and keeps its nameapi_baseandaws_bedrock_runtime_endpointsends chat and /v1/responses to the runtime endpointSurprises:
json_schemaresponse_format with tools; bridge honors itreasoning_itemsskipped re-reasoning on the follow-up turnreasoning_effort: none400s on gpt-6.1-sol, both commits, upstream/live-pr-risk (d6db8e8 vs 30c6687)
No dependent path broke. Each side ran 2 uvicorn workers with Postgres, a recorder in front of real Bedrock and OpenAI, and about 40 request cells, plus 144 integration tests on this branch merged with main
Breaking
None found
Backward incompatible
Each of these follows from Bedrock's Responses API serving the call instead of Converse, and each was observed on both sides
resp_id, and/spend/logs?request_id=<that id>finds no rowauthentication_errorinstead of 403permission_errortool_choice: "none"is now honored, where Converse still called the toolprompt_cache_keynow reaches Bedrock instead of being droppedredacted_thinkingblockaws_bedrock_runtime_endpointnow wins overapi_baseon Bedrock /v1/responses (the base sent the call toapi_base)extra_bodystill answers 400, but the message now carries the shared bridge's stack traceDecisions made in 30c6687, with the options not taken:
[a-zA-Z0-9_-]{1,64}keep Converse, which already renames them for Bedrock and maps them backoutputfor a spaced name)_litellm_responses_api_bridge_allowed, so a caller field namedresponses_api_bridge_allowedcan no longer steer the route or crash it (the earlier head answered 500)api_basefirst, which sent one deployment to two hosts depending on the routeRegression risk
clientfix too, covered by unit tests onlyDependency graph
completion()andget_optional_params, verified live (chat, streaming, fallbacks, guardrails, ARNs, names)clientforwarding, verified live on OpenAI and tested only on Azureget_optional_paramsnew keyword, popped before any provider sees it, verified live on OpenAI and BedrockNot verified
tests/integrationchaos groupType
🆕 New Feature
Caveats (if any)
Low
redacted_thinkingblockclientfix in the shared bridge also applies to OpenAI and Azure bridged callsbedrock/us-west-2/global.openai.gpt-6-sol) stay on Converse, since the Responses config doesn't parse region paths/spend/logs, as with OpenAI's bridge todayextra_body400s with a stack trace in the messageFinal Attestation
Note
Medium Risk
Changes Bedrock request routing and completion dispatch for a common agentic pattern (tools + reasoning); mistakes could send traffic to the wrong API or skip the bridge, but Converse remains the fallback for unsupported cases.
Overview
Bedrock GPT 5.6+ chat requests that combine function tools with reasoning (anything other than
reasoning_effort: "none") no longer default to Converse. When the model’s cost map lists/v1/responses, routing picks a newresponsespath andcompletion()uses the existing chat → Responses API bridge to hit bedrock-runtime’s native/openai/v1/responses, while still returning achat.completion.Bridging is narrow: only OpenAI-valid function tool names, no legacy
functions, noaws_bedrock_client, no tools inextra_body, and no mixed server-side tools (e.g.file_search). Guardrails, Converse-only params, invalid tool names, region-prefixed models, and models without a Responses endpoint still go to Converse;reasoning_effort: "none"stays on native chat completions.skip_responses_api_bridge/_litellm_responses_api_bridge_allowedfeeds Bedrock route selection so bridge behavior stays aligned with other providers. The shared bridge handler now passes the caller’sclientthrough on sync and async paths. Bedrock Responses URL building defers toget_runtime_endpointsoaws_bedrock_runtime_endpointwins overapi_baselike other routes.Reviewed by Cursor Bugbot for commit 30c6687. Bugbot is set up for automated code reviews on this repo. Configure here.