Skip to content

feat(bedrock): route OpenAI models to bedrock-runtime's native Chat Completions API - #43264

Closed
6matt wants to merge 16 commits into
BerriAI:mainfrom
6matt:feat/bedrock-native-openai-chat
Closed

6matt wants to merge 16 commits into
BerriAI:mainfrom
6matt:feat/bedrock-native-openai-chat

Conversation

@6matt

@6matt 6matt commented Sep 26, 2026

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Bedrock OpenAI models (gpt-5.4–6) were forced through Converse for chat
  • Converse 400s on unsupported sampling params instead of dropping them
  • It also loses reasoning tokens and breaks implicit prompt caching

How it solves it:

  • Route these models to bedrock-runtime's native /openai/v1/chat/completions
  • Data-driven from the model's supported_endpoints; no model-name matching
  • Function-tools + reasoning (which the chat surface rejects on gpt-5.6/6) bridge to native /v1/responses, so no Converse regression
  • bedrock/converse/ stays an escape hatch to Converse

User Flow

Before: a developer sends a chat request with a temperature to a Bedrock GPT-6 model and gets a hard 400 instead of an answer.

  1. They send POST https://litellm/v1/chat/completions with model: bedrock/global.openai.gpt-6-luna, temperature: 0.3, reasoning_effort: low
  2. The gateway translates it to Bedrock Converse and calls .../model/global.openai.gpt-6-luna/converse
  3. Bedrock returns 400 This model doesn't support the temperature field. Remove temperature and try again.

After: the same request succeeds; the unsupported sampling param is dropped and the native surface answers.

  1. The developer sends the same POST with the same body
  2. The gateway calls .../openai/v1/chat/completions, having dropped temperature (reasoning model) and kept reasoning_effort
  3. A normal chat completion comes back with content and real usage (incl. reasoning tokens), logged at non-zero spend
  4. A request that also carries function tools is instead sent to .../openai/v1/responses and comes back with normal tool_calls, unchanged from the caller's perspective

Relevant issues

Type

🆕 New Feature

Screenshots / Proof of Fix

Shared setup: AWS profile with Bedrock access in us-east-1; LITELLM_LOCAL_MODEL_COST_MAP=True to load the price-map entries in this PR. model="bedrock/global.openai.gpt-6-luna". drop_params=True. Same request bodies on both sides.

Before (a11a93f)

chat non-stream (temperature + reasoning)

  1. litellm.completion(..., temperature=0.3, reasoning_effort="low", max_completion_tokens=400)
  2. Request hits .../model/global.openai.gpt-6-luna/converse
  3. Raises litellm.BadRequestError: BedrockException - {"message":"This model doesn't support the temperature field. Remove temperature and try again."}

function tools + reasoning

  1. litellm.completion(..., tools=[get_weather], reasoning_effort="low", max_completion_tokens=600)
  2. Request hits .../model/global.openai.gpt-6-luna/converse
  3. Returns a get_weather tool call (Converse served this)

After (5678141)

chat non-stream (temperature + reasoning)

  1. Same call as Before
  2. Request hits .../openai/v1/chat/completions (temperature dropped, reasoning_effort kept)
  3. Returns content: "ok"

chat stream

  1. Same call with stream=True
  2. Request hits .../openai/v1/chat/completions
  3. Streamed deltas assemble to "stream ok"

function tools + reasoning

  1. Same call as Before
  2. Request is bridged to .../openai/v1/responses
  3. Returns finish_reason: tool_calls, object: chat.completion, get_weather({"city":"Paris"}) — unchanged shape for the caller, no Converse

Caveats (if any)

Low

  • Scope: gpt-5.4/5.5/5.6/6 cross-Region inference profiles only; bare on-demand ids excluded (Bedrock rejects them)
  • Function-tools + active reasoning on gpt-5.6/6 is served via the native /v1/responses bridge (chat surface rejects it); gpt-5.4/5.5 serve it on chat directly
  • Version boundary (gpt-5.6+) is measured against bedrock-runtime, not advertised in metadata; a future model shifting it would need the check updated

QA runbook

  • tests/unit/llms/bedrock/chat/test_bedrock_openai_native_chat.py — a model advertising /v1/chat/completions routes to the native config/URL and drops reasoning-model sampling params; function-tools + active reasoning on gpt-5.6/6 bridges to /v1/responses while gpt-5.4/5.5, reasoning=none, and no-tools requests stay on chat; the converse/ prefix stays an escape hatch
    • Register a model with supported_endpoints: ["/v1/chat/completions", "/v1/responses"]
    • Send /v1/chat/completions with temperature=0.5 + reasoning_effort=low; expect temperature dropped, request to .../openai/v1/chat/completions
    • Send the same with function tools + reasoning_effort=low; expect the request to .../openai/v1/responses and normal tool_calls back
    • Send bedrock/converse/<model>; expect it stays on Converse
    • Sanity check: this test makes sense to add and is not hand-wavey or flaky

…ompletions API

bedrock-runtime serves the OpenAI models (gpt-5.4/5.5/5.6/6) on an
OpenAI-compatible surface at /openai/v1/chat/completions, alongside Converse.
Previously the bedrock provider translated every chat request into Converse,
which is lossy for these models: it 400s on unsupported sampling params instead
of dropping them, loses reasoning tokens, and hoists mid-conversation system
messages (breaking implicit prompt caching).

Add BedrockOpenAIChatConfig, the Chat Completions sibling of
BedrockOpenAIResponsesConfig. It inherits OpenAIGPT5Config (the same reasoning
param mapper the openai provider uses) and overrides only the endpoint URL,
Bearer/SigV4 auth, and AWS-param stripping. Routing is data-driven from the
model's supported_endpoints, via a shared bedrock_uses_native_openai_chat
predicate used by the config selector, the completion dispatcher, and
get_optional_params. An explicit bedrock/converse/<model> prefix stays an
escape hatch to the Converse translation.

Enables the surface for the 16 gpt-5.4/5.5/5.6/6 cross-Region inference-profile
entries that advertise /v1/chat/completions.
…sponses API

bedrock-runtime's native chat surface rejects function tools while reasoning is
active on gpt-5.6/6 ("use /v1/responses or set reasoning_effort to 'none'"),
which Converse previously served. Rather than regress that path, extend the
existing responses_api_bridge_check (already used for openai/azure with the same
constraint) to fire for bedrock: a chat request with function tools + active
reasoning on a model that supports the native Responses surface is bridged to
/v1/responses, which handles tools + reasoning and returns proper tool_calls.

The version boundary is measured against bedrock-runtime (gpt-5.6+ reject; 5.4/5.5
serve tools with reasoning natively) via
bedrock_chat_rejects_function_tools_while_reasoning. Custom tools, reasoning_effort
"none", and non-tool requests stay on native chat completions.
@greptile-apps

greptile-apps Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[High risk] Routes OpenAI models through a new Bedrock chat endpoint.

The PR appears safe to merge based on this re-review

Summary

Routes eligible Bedrock OpenAI chat models to the native Chat Completions endpoint, with a Responses bridge for function tools used with reasoning. The latest changes remove a redundant comment and strengthen the signed-body test

Reviews (6) · Last reviewed commit: "test(bedrock): assert exact model field ..."

Comment thread litellm/llms/bedrock/chat/openai_native/transformation.py
Comment thread litellm/llms/bedrock/common_utils.py
Comment thread litellm/main.py
Comment thread litellm/main.py Outdated
Comment thread tests/unit/llms/bedrock/chat/test_bedrock_openai_native_chat.py
Comment thread litellm/llms/bedrock/chat/openai_native/transformation.py Outdated
@veria-ai

veria-ai Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 2 · PR risk: 0/10

…; trim comment

Address review: move the gpt-5.6/6 "native chat rejects function tools while
reasoning" boundary out of a hardcoded version check and into a
bedrock_chat_rejects_function_tools_while_reasoning flag on the price-map entries,
read like the other data-driven Bedrock capabilities, so onboarding a model stays
a JSON change. Add a wire-level test asserting the SigV4-signed streaming body
carries stream:true, and drop a narrating dispatch comment.
@6matt

6matt commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai Pushed 25d0ab3 addressing the review: moved the tools+reasoning boundary to a price-map flag, verified the streaming-body concern is a false positive (added a wire-level test), and trimmed the narrating comment. Please re-review.

@codecov

codecov Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Comment thread litellm/llms/bedrock/chat/openai_native/transformation.py Outdated
…ta on native chat

validate_environment passed caller headers through untouched, so a caller could set
X-Amzn-Bedrock-Request-Metadata and have it signed and recorded in AWS billing /
CloudTrail, forging identity attribution. Reuse the existing ownership merge
(bedrock_request_metadata_headers + merge_bedrock_invoke_headers, as the Invoke
OpenAI path does): the proxy-owned header is dropped from caller input
(case-insensitively) and re-supplied only from trusted litellm_params.
Comment thread litellm/main.py
…ster price-map key

- Add coverage for the get_optional_params branch, the _complete_bedrock native
  dispatch, and predicate edge cases (None model, non-converse route).
- Drop untyped **kwargs from BedrockOpenAIChatConfig.__init__ (BaseAWSLLM takes no
  kwargs; every call site is arg-less) to satisfy the typed-parameter guideline.
- Register bedrock_chat_rejects_function_tools_while_reasoning in the schema
  generator's boolean key table and regenerate the price-map schema.
@6matt

6matt commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai Pushed through 55f819c: fixed the request-attribution forgery (proxy now owns the X-Amzn-Bedrock-Request-Metadata header via the existing ownership merge), typed __init__, added patch coverage for the dispatch/param-mapping/predicate branches, and registered the new price-map boolean key in the schema generator. Please re-review.

… native chat

The model rides in the request body on the native /openai/v1/chat/completions
surface (not the URL like Converse/Invoke), and the generic handler merges
extra_body over the body before SigV4 signing. bedrock isn't in the extra_body
model-stripping path, so a key limited to one model could send
extra_body={"model": "<other>"} and invoke another model on the proxy's AWS
credentials. Pin the authorized model back in sign_request (post-merge, pre-sign),
covering both the SigV4 and Bearer paths.
@6matt

6matt commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai Pushed 582fe71 fixing the extra_body model-authorization bypass on the native chat surface (sign_request now pins the authorized model post-merge/pre-sign). Please re-review.

Comment thread litellm/llms/bedrock/chat/openai_native/transformation.py Outdated
@greptile-apps

greptile-apps Bot commented Sep 26, 2026

Copy link
Copy Markdown
Contributor

Want your agent to iterate on Greptile's feedback? Try greploops.

@codspeed

codspeed Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing 6matt:feat/bedrock-native-openai-chat (ad1f2a4) with main (d08746f)

Open in CodSpeed

…tation

- Add tests for custom_llm_provider, get_error_class, the non-reasoning param
  fallbacks, and the SigV4 sign_request path (brings the adapter to 100% patch
  coverage).
- Move the model-pin suppression inline on the mutated line as
  `# rebind-ok` (LIT011) and drop the misplaced block comment.
@6matt

6matt commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai Pushed 4466a96. Also merged the latest main (cost-map guard was failing on the merge because the schema was stale against the merged cost map — regenerated), brought the adapter to 100% patch coverage, and added the # rebind-ok justification. Please re-review.

Comment thread litellm/llms/bedrock/chat/openai_native/transformation.py Outdated
Comment thread tests/unit/llms/bedrock/chat/test_bedrock_openai_native_chat.py Outdated
@6matt

6matt commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai Pushed c7d2829: dropped the redundant model-pin comment and tightened the signed-body assertion to check the exact "model" field. Please re-review.

The cost-map-guard runs under pull_request_target with the BASE branch's schema
generator (PR code is never executed, only the cost map is read as data), so a
price-map key a PR introduces can never be classified by that generator -- the
guard fails every run with "Unclassified keys". Move the gpt-5.6+ boundary back
into bedrock_chat_rejects_function_tools_while_reasoning as a version check on the
normalized OpenAI model name, matching the openai/azure arms already in the same
bridge, and drop the flag from the 12 cost-map entries, the schema generator, and
the schema. No new cost-map key, so the guard passes.
@krrish-berri-2

Copy link
Copy Markdown
Contributor

seems like a dup of #40775

@6matt

6matt commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

@krrish-berri-2 mostly, but #40775 doesn't include GPT 6 and it routes function-tools + reasoning to converse rather than responses. I can close this one and create a followup after the #40775 lands

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants