Skip to content

feat(bedrock): bridge gpt-5.6+ function tools with reasoning to native /responses instead of Converse - #45609

Open
6matt wants to merge 2 commits into
BerriAI:mainfrom
6matt:litellm_bedrock_gpt_tools_reasoning_responses_bridge
Open

6matt wants to merge 2 commits into
BerriAI:mainfrom
6matt:litellm_bedrock_gpt_tools_reasoning_responses_bridge

Conversation

@6matt

@6matt 6matt commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Bedrock GPT-5.6+ tool calls with reasoning still fall back to Converse
  • Converse hides reasoning_tokens from usage, so callers can't see the reasoning they pay for
  • Converse also ignores a json_schema response_format when tools are present
  • That's most agentic traffic, since reasoning is on by default

How it solves it:

  • Bridge those requests to bedrock-runtime's native /openai/v1/responses
  • Reuses the chat-to-Responses bridge OpenAI and Azure already use
  • Data-driven: only models listing /v1/responses in the cost map
  • Converse-only fields (guardrails, ARNs, thinking, legacy functions, deprecated aws_bedrock_client) still go to Converse
  • Only plain function tools bridge: server-side tools (e.g. file_search, also via extra_body) stay on Converse
  • Tool names the OpenAI API rejects (spaces, dots, over 64 chars) keep Converse
  • litellm.responses(..., use_chat_completions_api=True) skips the bridge, so those calls keep Converse
  • aws_bedrock_runtime_endpoint now wins over api_base on Bedrock /v1/responses, like Converse
  • Async bridge calls now forward the caller's client, matching the sync path

User Flow

Before: a developer calling a Bedrock GPT-5.6+ model with tools and reasoning gets Converse's answer, with no reasoning in it

  1. They send POST http://localhost:4000/v1/chat/completions with model: bedrock-gpt-6.1-sol, a get_weather tool, and reasoning_effort: high
  2. They get a chat.completion with tool_calls, an id like chatcmpl-..., and usage.completion_tokens_details empty, so they cannot see how many reasoning tokens they paid for
  3. The same request with "stream": true and include_usage ends with the same empty completion_tokens_details
  4. POST http://localhost:4000/v1/messages with thinking and tools answers through Converse too

After: the same requests are served by the model's native OpenAI Responses API

  1. They send the same POST with the same body
  2. They get a chat.completion with the same tool_calls, an id like resp_..., usage.completion_tokens_details.reasoning_tokens filled in, and a reasoning_items entry they can echo back on the next turn
  3. Streaming reports reasoning_tokens in the final usage chunk
  4. reasoning_effort: "none" still goes to native /openai/v1/chat/completions, and a guardrailConfig request still goes to Converse and is still blocked by the guardrail
  5. /v1/responses is unchanged, and /v1/messages returns the same tool_use block, without Converse's redacted_thinking block

Relevant issues

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live run against real Bedrock (us-east-1, real spend) from a maintainer machine. Two proxies, one per commit, each booted with 2 uvicorn workers and no DB, with LITELLM_LOCAL_MODEL_COST_MAP=True:

python litellm/proxy/proxy_cli.py --config config.yaml --port $PORT --num_workers 2
model_list:
  - model_name: bedrock-gpt-6.1-sol
    litellm_params:
      model: bedrock/global.openai.gpt-6.1-sol
      aws_region_name: us-east-1
      aws_access_key_id: os.environ/QA_AWS_ACCESS_KEY_ID
      aws_secret_access_key: os.environ/QA_AWS_SECRET_ACCESS_KEY
  - model_name: bedrock-gpt-5.6-sol
    litellm_params:
      model: bedrock/us.openai.gpt-5.6-sol
      aws_region_name: us-east-1
      aws_access_key_id: os.environ/QA_AWS_ACCESS_KEY_ID
      aws_secret_access_key: os.environ/QA_AWS_SECRET_ACCESS_KEY
litellm_settings:
  drop_params: true
router_settings:
  disable_cooldowns: true

Shared shell variables ($URL is the base proxy for Before and the head proxy for After, and $GR is a Bedrock guardrail whose topic policy denies coffee):

TOOLS='[{"type":"function","function":{"name":"get_weather","description":"Current weather for a city","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}}]'
Q='Cities: Paris, Lima, Oslo, Cairo, Tokyo, Quito. Let n be the number of primes p < 60 for which p^2 + 2 is also prime. Pick the city at position n. Get its weather with the tool.'

Before (d6db8e8)

/v1/chat/completions, tools + reasoning, gpt-6.1-sol

  1. Command

    curl -s -D headers.txt $URL/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
      -d '{"model":"bedrock-gpt-6.1-sol","reasoning_effort":"high","tools":'$TOOLS',"messages":[{"role":"user","content":"'$Q'"}]}' \
      | jq -c '{id: .id[0:9], finish: .choices[0].finish_reason, tool_calls: [.choices[0].message.tool_calls[].function | {name, arguments}], completion_tokens_details: .usage.completion_tokens_details, reasoning_items: (.choices[0].message.reasoning_items | length)}'
    grep -i '^x-litellm-response-cost:' headers.txt
  2. Output, no reasoning tokens reported

    {"id":"chatcmpl-","finish":"tool_calls","tool_calls":[{"name":"get_weather","arguments":"{\"city\": \"Paris\"}"}],"completion_tokens_details":{},"reasoning_items":0}
    x-litellm-response-cost: 0.001276
    

/v1/chat/completions, tools + reasoning, streaming

  1. Command

    curl -sN $URL/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
      -d '{"model":"bedrock-gpt-6.1-sol","reasoning_effort":"high","stream":true,"stream_options":{"include_usage":true},"tools":'$TOOLS',"messages":[{"role":"user","content":"'$Q'"}]}' \
      | grep '"usage":{' | sed 's/^data: //' | jq -c '{completion_tokens: .usage.completion_tokens, completion_tokens_details: .usage.completion_tokens_details}'
  2. Output

    {"completion_tokens":105,"completion_tokens_details":{}}
    

/v1/chat/completions, tools + reasoning, gpt-5.6-sol

  1. Same as the first case with "model":"bedrock-gpt-5.6-sol", printing {id, finish, completion_tokens_details}

  2. Output

    {"id":"chatcmpl-","finish":"tool_calls","completion_tokens_details":{}}
    

/v1/chat/completions, tool result follow-up turn

  1. Command

    curl -s $URL/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
      -d '{"model":"bedrock-gpt-6.1-sol","reasoning_effort":"low","tools":'$TOOLS',"messages":[{"role":"user","content":"What is the weather in Paris? Use the tool, then answer in one sentence."},{"role":"assistant","content":null,"tool_calls":[{"id":"call_qa_001","type":"function","function":{"name":"get_weather","arguments":"{\"city\": \"Paris\"}"}}]},{"role":"tool","tool_call_id":"call_qa_001","content":"{\"temp_c\": 14, \"sky\": \"light rain\"}"}]}' \
      | jq -c '{finish: .choices[0].finish_reason, content: .choices[0].message.content}'
  2. Output

    {"finish":"stop","content":"The weather in Paris is 14°C with light rain."}
    

Control, reasoning_effort none

  1. Same as the first case with "model":"bedrock-gpt-5.6-sol","reasoning_effort":"none" and the Paris question

  2. Output, served by native chat completions

    {"id":"chatcmpl-","finish":"tool_calls","completion_tokens_details":{"accepted_prediction_tokens":0,"audio_tokens":0,"reasoning_tokens":0,"rejected_prediction_tokens":0}}
    

Control, guardrailConfig

  1. Command

    curl -s $URL/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
      -d '{"model":"bedrock-gpt-6.1-sol","reasoning_effort":"low","guardrailConfig":{"guardrailIdentifier":"'$GR'","guardrailVersion":"DRAFT"},"tools":'$TOOLS',"messages":[{"role":"user","content":"What is the best way to brew coffee at home? Also check the weather in Paris."}]}' \
      | jq -c '{id: .id[0:9], finish: .choices[0].finish_reason, content: .choices[0].message.content}'
  2. Output, blocked by the guardrail on Converse

    {"id":"chatcmpl-","finish":"content_filter","content":"Sorry, the model cannot answer this question."}
    

/v1/responses, tools + reasoning

  1. Command

    curl -s $URL/v1/responses -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
      -d '{"model":"bedrock-gpt-6.1-sol","reasoning":{"effort":"high"},"tools":[{"type":"function","name":"get_weather","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}],"input":"'$Q'"}' \
      | jq -c '{status, calls: [.output[] | select(.type=="function_call") | {name, arguments}], output_tokens_details: .usage.output_tokens_details}'
  2. Output

    {"status":"completed","calls":[{"name":"get_weather","arguments":"{\"city\":\"Paris\"}"}],"output_tokens_details":{"audio_tokens":null,"reasoning_tokens":56,"text_tokens":null}}
    

/v1/messages, thinking + tools

  1. Command

    curl -s $URL/v1/messages -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' -H 'anthropic-version: 2023-06-01' \
      -d '{"model":"bedrock-gpt-6.1-sol","max_tokens":4096,"thinking":{"type":"enabled","budget_tokens":2048},"tools":[{"name":"get_weather","description":"Current weather for a city","input_schema":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}],"messages":[{"role":"user","content":"'$Q'"}]}' \
      | jq -c '{stop_reason, blocks: [.content[].type], tool_input: [.content[] | select(.type=="tool_use") | .input]}'
  2. Output

    {"stop_reason":"tool_use","blocks":["redacted_thinking","text","tool_use"],"tool_input":[{"city":"Paris"}]}
    

After (30c6687)

/v1/chat/completions, tools + reasoning, gpt-6.1-sol

  1. Same command as Before

  2. Output, reasoning tokens reported and billed, plus an encrypted reasoning item to echo back

    {"id":"resp_ab72","finish":"tool_calls","tool_calls":[{"name":"get_weather","arguments":"{\"city\":\"Paris\"}"}],"completion_tokens_details":{"reasoning_tokens":83},"reasoning_items":1}
    x-litellm-response-cost: 0.001216
    

/v1/chat/completions, tools + reasoning, streaming

  1. Same command as Before

  2. Output

    {"completion_tokens":104,"completion_tokens_details":{"reasoning_tokens":84}}
    

/v1/chat/completions, tools + reasoning, gpt-5.6-sol

  1. Same command as Before

  2. Output

    {"id":"resp_ylfn","finish":"tool_calls","completion_tokens_details":{"reasoning_tokens":116}}
    

/v1/chat/completions, tool result follow-up turn

  1. Same command as Before

  2. Output

    {"finish":"stop","content":"It’s currently 14°C with light rain in Paris."}
    

Control, reasoning_effort none

  1. Same command as Before

  2. Output, still native chat completions

    {"id":"chatcmpl-","finish":"tool_calls","completion_tokens_details":{"accepted_prediction_tokens":0,"audio_tokens":0,"reasoning_tokens":0,"rejected_prediction_tokens":0}}
    

Control, guardrailConfig

  1. Same command as Before

  2. Output, still Converse and still blocked

    {"id":"chatcmpl-","finish":"content_filter","content":"Sorry, the model cannot answer this question."}
    

/v1/responses, tools + reasoning

  1. Same command as Before

  2. Output, unchanged

    {"status":"completed","calls":[{"name":"get_weather","arguments":"{\"city\":\"Paris\"}"}],"output_tokens_details":{"audio_tokens":null,"reasoning_tokens":64,"text_tokens":null}}
    

/v1/messages, thinking + tools

  1. Same command as Before

  2. Output, same tool call, no redacted_thinking block

    {"stop_reason":"tool_use","blocks":["tool_use"],"tool_input":[{"city":"Paris"}]}
    

Extra checks on the head proxy, all answered 200 with the expected tool call:

  • Tools with no description, and with strict: true
  • An image input plus a tool
  • A two-turn loop that echoes the assistant message back, on head, across base then head, and across head then base
  • A two-turn /v1/messages loop, and /v1/messages streaming
  • A tool named get weather! (also streamed) stays on Converse and keeps its name
  • A deployment with both api_base and aws_bedrock_runtime_endpoint sends chat and /v1/responses to the runtime endpoint
  • The same cases on this branch merged with main at 96ac1ec

Surprises:

  • Converse ignored json_schema response_format with tools; bridge honors it
  • Converse already hit implicit cache on a single system prompt
  • Echoed reasoning_items skipped re-reasoning on the follow-up turn
  • reasoning_effort: none 400s on gpt-6.1-sol, both commits, upstream
  • One bridged call hung 363 s at Bedrock, then 500ed

/live-pr-risk (d6db8e8 vs 30c6687)

No dependent path broke. Each side ran 2 uvicorn workers with Postgres, a recorder in front of real Bedrock and OpenAI, and about 40 request cells, plus 144 integration tests on this branch merged with main

Breaking

None found

Backward incompatible

Each of these follows from Bedrock's Responses API serving the call instead of Converse, and each was observed on both sides

  • Bridged calls return a resp_ id, and /spend/logs?request_id=<that id> finds no row
    • Kept because OpenAI and Azure bridged models do the same today (reproduced on the base with gpt-5.3-codex), so the fix belongs in the shared bridge and is tracked separately
  • Bad AWS credentials answer 401 authentication_error instead of 403 permission_error
    • Kept because that is what Bedrock's Responses API returns, and fallbacks still fire on it
  • tool_choice: "none" is now honored, where Converse still called the tool
  • An unknown top-level request field is dropped instead of reaching Bedrock and failing with 400, same as the OpenAI bridge
  • prompt_cache_key now reaches Bedrock instead of being dropped
  • /v1/messages no longer returns Converse's opaque redacted_thinking block
  • aws_bedrock_runtime_endpoint now wins over api_base on Bedrock /v1/responses (the base sent the call to api_base)
  • A string extra_body still answers 400, but the message now carries the shared bridge's stack trace

Decisions made in 30c6687, with the options not taken:

  • Tool names outside [a-zA-Z0-9_-]{1,64} keep Converse, which already renames them for Bedrock and maps them back
    • Not taken was renaming inside the shared bridge (changes OpenAI and Azure too), or bridging them unchanged (Bedrock answered with an empty output for a spaced name)
  • The internal bridge flag is named _litellm_responses_api_bridge_allowed, so a caller field named responses_api_bridge_allowed can no longer steer the route or crash it (the earlier head answered 500)
    • Not taken was keeping the plain name and stripping it from caller params
  • /v1/responses on Bedrock picks its host with the same helper Converse and native chat use
    • Not taken was keeping api_base first, which sent one deployment to two hosts depending on the route

Regression risk

  • Azure bridged calls get the async client fix too, covered by unit tests only
  • Region-path models stay on Converse, covered by unit tests only

Dependency graph

  • Bedrock route choice, called by completion() and get_optional_params, verified live (chat, streaming, fallbacks, guardrails, ARNs, names)
  • Bedrock Responses URL builder, used by /v1/responses and the bridge, verified live plus a unit test
  • Shared bridge async client forwarding, verified live on OpenAI and tested only on Azure
  • get_optional_params new keyword, popped before any provider sees it, verified live on OpenAI and Bedrock
  • Spend logging, key spend, auth, and health, verified live with the same cost and key spend on both sides

Not verified

  • Azure, since this rig had no Azure deployment
  • tests/integration chaos group
  • Bedrock regions other than us-east-1

Type

🆕 New Feature

Caveats (if any)

Low

  • /v1/messages no longer returns Converse's opaque redacted_thinking block
  • The async client fix in the shared bridge also applies to OpenAI and Azure bridged calls
  • Region-path models (bedrock/us-west-2/global.openai.gpt-6-sol) stay on Converse, since the Responses config doesn't parse region paths
  • GPT-5.4/5.5 still use Converse: they aren't on the native chat route at all
  • Bridged calls can't be found by id in /spend/logs, as with OpenAI's bridge today
    • Net-negative to fix here, since the fix changes the shared bridge for every provider
  • Bad AWS credentials answer 401 instead of 403
    • Net-negative to fix here, since remapping would hide what Bedrock actually said
  • A string extra_body 400s with a stack trace in the message
    • Net-negative to fix here, since it is shared bridge code and the request is invalid either way

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Medium Risk
Changes Bedrock request routing and completion dispatch for a common agentic pattern (tools + reasoning); mistakes could send traffic to the wrong API or skip the bridge, but Converse remains the fallback for unsupported cases.

Overview
Bedrock GPT 5.6+ chat requests that combine function tools with reasoning (anything other than reasoning_effort: "none") no longer default to Converse. When the model’s cost map lists /v1/responses, routing picks a new responses path and completion() uses the existing chat → Responses API bridge to hit bedrock-runtime’s native /openai/v1/responses, while still returning a chat.completion.

Bridging is narrow: only OpenAI-valid function tool names, no legacy functions, no aws_bedrock_client, no tools in extra_body, and no mixed server-side tools (e.g. file_search). Guardrails, Converse-only params, invalid tool names, region-prefixed models, and models without a Responses endpoint still go to Converse; reasoning_effort: "none" stays on native chat completions.

skip_responses_api_bridge / _litellm_responses_api_bridge_allowed feeds Bedrock route selection so bridge behavior stays aligned with other providers. The shared bridge handler now passes the caller’s client through on sync and async paths. Bedrock Responses URL building defers to get_runtime_endpoint so aws_bedrock_runtime_endpoint wins over api_base like other routes.

Reviewed by Cursor Bugbot for commit 30c6687. Bugbot is set up for automated code reviews on this repo. Configure here.

@greptile-apps

greptile-apps Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[High impact] The PR appears safe to merge; no new actionable issue remains.

Summary

This PR sends supported Bedrock function-tool requests with active reasoning through native Responses, then returns the usual chat response.

  • Eligible Bedrock chat requests with reasoning and function tools use native Responses.
  • Async Responses bridge calls use the caller’s client.
  • Bedrock Responses calls use the configured runtime endpoint.

Reviews (6) · Last reviewed commit: "fix(bedrock): keep Converse for tool nam..." · Reviewed by Greptile

Comment thread litellm/main.py
Comment thread litellm/main.py
Comment thread litellm/main.py
@codspeed

codspeed Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing 6matt:litellm_bedrock_gpt_tools_reasoning_responses_bridge (30c6687) with main (96ac1ec)

Open in CodSpeed

Comment thread litellm/main.py
@veria-ai

veria-ai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 1 · PR risk: 0/10

@codecov

codecov Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@6matt

6matt commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/main.py Outdated
@6matt

6matt commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai

@6matt

6matt commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/llms/bedrock/common_utils.py Outdated
Comment thread litellm/main.py Outdated
@6matt

6matt commented Oct 10, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/main.py
@6matt
6matt force-pushed the litellm_bedrock_gpt_tools_reasoning_responses_bridge branch from 170efd4 to e2966ab Compare October 10, 2026 02:10
@6matt

6matt commented Oct 10, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai

… bridge flag, honor aws_bedrock_runtime_endpoint on Responses
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 30c6687. Configure here.

Comment thread litellm/utils.py
model,
{**passed_params, "_litellm_responses_api_bridge_allowed": _litellm_responses_api_bridge_allowed},
additional_drop_params,
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Param mapping misses converse-only kwargs

Medium Severity

get_optional_params now decides the new responses vs converse route from passed_params only, so kwargs-only fields such as requestMetadata, guardrailConfig, extra_body.tools, and aws_bedrock_client never reach that decision. Dispatch in completion() still sees the full request and correctly stays on Converse, but mapping uses the chat-completions supported-params list. A tools-plus-reasoning call that also sends requestMetadata now raises UnsupportedParamsError or drops that metadata, even though Converse previously accepted it.

Additional Locations (2)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 30c6687. Configure here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

pre_process_non_default_params writes every kwarg into passed_params in place (utils.py L4284) before L4592, so requestMetadata and guardrailConfig reach the route decision. Probe matched main

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants