Skip to content

feat(bedrock): add openai.gpt-6-astra profiles and route its reasoning_effort like gpt-5.x - #40370

Closed
CrypticDriver wants to merge 1 commit into
BerriAI:litellm_internal_stagingfrom
CrypticDriver:feat/bedrock-runtime-openai-responses
Closed

CrypticDriver wants to merge 1 commit into
BerriAI:litellm_internal_stagingfrom
CrypticDriver:feat/bedrock-runtime-openai-responses

Conversation

@CrypticDriver

@CrypticDriver CrypticDriver commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • bedrock/global.openai.gpt-6-astra (and us.) cannot call tools through LiteLLM
  • The model has no price-map entry, so it falls to the invoke route and tools is rejected or dropped
  • reasoning_effort is mapped to Anthropic thinking because the Converse code only recognises openai.gpt-5

How it solves it:

  • Add us. / global.openai.gpt-6-astra to the price map, same shape as the GPT-5.6 entries
  • Treat every openai.gpt-* (except gpt-oss) as an OpenAI frontier model in the Converse transformation

User Flow

Before: a developer pointing their agent at GPT-6 Astra on Bedrock through the gateway cannot get a single tool call back

  1. They send POST https://litellm-domain/v1/chat/completions with "model": "bedrock/global.openai.gpt-6-astra", a user message and a tools array
  2. With default settings they get 400 bedrock does not support parameters: ['tools'], for model=global.openai.gpt-6-astra
  3. If the proxy admin turns on drop_params, the tools are silently removed, the request goes to .../model/global.openai.gpt-6-astra/invoke, and AWS rejects it because this model does not support Invoke
  4. Either way the agent never receives a tool_calls message

After: the same request comes back with a real tool call

  1. They send the same POST https://litellm-domain/v1/chat/completions with "model": "bedrock/global.openai.gpt-6-astra" and the same tools array
  2. The gateway forwards it to https://bedrock-runtime.us-east-2.amazonaws.com/model/global.openai.gpt-6-astra/converse with a toolConfig
  3. The response is a normal chat completion with finish_reason: "tool_calls" and tool_calls[0].function.name == "get_weather"; stream: true returns the same as SSE deltas
  4. Adding "reasoning_effort": "low" now reaches Bedrock as reasoning.effort instead of failing with Unknown parameter: 'thinking'
  5. https://litellm-domain/ui/?page=logs shows the request with non-zero spend priced from the Bedrock rate card

Relevant issues

Related (not closed by this PR): #40246, #40080

Live verification

Run against a real account in us-east-2 with an IAM role, using the bundled price map (LITELLM_LOCAL_MODEL_COST_MAP=True):

Scenario Result
completion() + tools + reasoning_effort=low, non-stream finish_reason=tool_calls, get_weather {"city": "Beijing"}, cost 0.00142
tool-result round trip final text answer
completion() + tools + reasoning_effort=low, stream=True (us. profile) 8 chunks, finish_reason=tool_calls
Router with a model_list entry tool call returned

Without the Converse change the reasoning_effort case fails with BedrockException {"code":"unknown_parameter","message":"Unknown parameter: 'thinking'."}

Notes

  • Earlier revisions of this PR added a bedrock-runtime /openai/v1/responses backend. That turned out to be unnecessary for tool calling: GPT-6 supports tools on Converse exactly like GPT-5.6, so this revision is the price-map entry plus the one-line model check, the same shape as the GPT-5.6 onboarding. A native Responses path can be a separate PR if a Responses-only feature is needed
  • The us. entry uses the Geo CRIS rate and global. the Global CRIS rate from the model card; both carry the 272K tiers and cache write/read rates like the GPT-5.6 entries
  • GPT-5.6 entries are untouched

Pre-Submission checklist

  • I have added meaningful tests (tests/test_litellm/llms/bedrock/test_cross_region_inference_profile_mapping.py, tests/test_litellm/llms/bedrock/chat/test_converse_transformation.py)
  • The handful of test files covering my change pass locally: those two files plus tests/test_litellm/llms/bedrock/chat, tests/test_litellm/llms/bedrock/test_bedrock_common_utils.py, tests/test_litellm/llms/bedrock_mantle, tests/test_litellm/completion_extras, tests/test_litellm/test_model_prices_schema.py. scripts/ruff_strict_gate.py, scripts/type_discipline_gate.py and scripts/type_check_gate.py pass against the merge-base
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)

@CrypticDriver
CrypticDriver requested a review from a team September 9, 2026 06:00
@codspeed

codspeed Bot commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing CrypticDriver:feat/bedrock-runtime-openai-responses (1505bad) with litellm_internal_staging (47b15ff)

Open in CodSpeed

@codecov

codecov Bot commented Sep 9, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@greptile-apps

greptile-apps Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR adds native Amazon Bedrock Responses API routing for GPT-6 Astra, including SigV4 or bearer authentication, chat-to-Responses parameter shaping, model pricing, and behavioral tests

  • Adds a Bedrock Responses adapter backed by /openai/v1/responses
  • Uses model metadata to select native Responses routing
  • Adds global and US GPT-6 Astra capability and pricing entries
  • Adds URL, authentication, parameter, direct Responses, and non-streaming tool-call tests
  • Requires repository-rule cleanup and streaming regression coverage before merge

Confidence Score: 4/5

The implementation appears behaviorally sound, but explicit repository requirements and streaming regression coverage should be addressed before merging

No concrete runtime or security failure remains, but provider-specific logic is placed outside the required adapter boundary, unnecessary comments violate the source-comment policy, and the critical streaming flow lacks regression coverage

Files Needing Attention: litellm/utils.py, litellm/litellm_core_utils/get_supported_openai_params.py, litellm/llms/bedrock/responses/transformation.py, tests/test_litellm/llms/bedrock/responses/test_bedrock_responses_transformation.py

Important Files Changed

Filename Overview
litellm/llms/bedrock/responses/transformation.py Adds the native Bedrock Responses endpoint, URL construction, authentication, error mapping, and chat parameter configuration
litellm/utils.py Adds Bedrock Responses capability routing, but places provider-specific selection logic outside the required provider directory
litellm/litellm_core_utils/get_supported_openai_params.py Selects Responses-compatible Bedrock parameters through provider-specific logic outside the required provider directory
model_prices_and_context_window.json Adds global and US GPT-6 Astra pricing, limits, endpoint support, and capability metadata
tests/test_litellm/llms/bedrock/responses/test_bedrock_responses_transformation.py Covers URL, authentication, routing, pricing, and complete tool calls, but omits the advertised streaming tool-call flow

Reviews (1): Last reviewed commit: "feat(bedrock): serve openai.gpt-6-astra ..." | Re-trigger Greptile

Comment thread litellm/utils.py Outdated
Comment on lines +8726 to +8738
elif litellm.LlmProviders.BEDROCK == provider:
# bedrock-runtime serves OpenAI frontier models (openai.gpt-6-astra and its
# us./global. inference profiles) on /openai/v1/responses. Data-driven from
# the price map (mode: responses / supported_endpoints); models without
# that signal (Claude, gpt-oss, ...) return None and keep the
# chat-completions emulation exactly as before.
from litellm.llms.bedrock.responses.transformation import (
bedrock_supports_responses,
)

if not bedrock_supports_responses(model, litellm.model_cost):
return None
return litellm.BedrockResponsesAPIConfig()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Provider Logic Outside Adapters

These Bedrock branches violate the directive requiring provider-specific logic to remain under llms/; this requirement must be satisfied before merging

Rule Used: What: Avoid writing provider-specific code outside... (source)

Knowledge Base Used: Provider adapters and capabilities

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reworked in 0927eb1. The chat-side decision now lives in get_bedrock_chat_config under llms/bedrock, which returns BedrockResponsesBridgeChatConfig for price-map mode: responses models. get_optional_params no longer special-cases these models; its Converse branch only applies when the resolved provider config is AmazonConverseConfig, so the bridge config goes through the generic provider_config.map_openai_params fallthrough that already exists. The get_supported_openai_params change is gone entirely. What remains in utils.py is the one-line dispatch in _get_python_responses_api_config, the same shape every other provider uses there

Comment on lines +44 to +45
# Checked longest/most-specific first so a full endpoint URL collapses to host in
# one pass and the appended path never doubles.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Unnecessary Explanatory Comments

These comments restate straightforward suffix and authentication code, violating the directive limiting comments to genuinely complex logic; resolve this before merging

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Trimmed in 0927eb1: the suffix, URL helper and SigV4 comments are gone. What is left is the module docstring, the gate function docstring, one line explaining why validate_environment must not inherit the OPENAI_API_KEY injection, and the reason text the type-discipline gate requires on each mutable-ok suppression

Comment on lines +267 to +281
def test_completion_with_tools_is_bridged_to_responses(
self, local_model_cost_map, gpt6_in_converse_set, fake_aws_env
):
route = respx.post("https://bedrock-runtime.us-east-2.amazonaws.com/openai/v1/responses").mock(
return_value=httpx.Response(200, json=_canned_function_call_response(GPT6_GLOBAL))
)

response = litellm.completion(
model=f"bedrock/{GPT6_GLOBAL}",
messages=[{"role": "user", "content": "weather in Beijing?"}],
tools=[WEATHER_TOOL],
)

assert route.called, "request must go to bedrock-runtime /openai/v1/responses"
sent = json.loads(route.calls.last.request.content)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Streaming Flow Lacks Coverage

The advertised streaming tool-call flow only has complete JSON tests. Add an SSE regression covering tool deltas and finish_reason before merging

Context Used: CLAUDE.md (source)

Knowledge Base Used: Model request execution

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added test_streamed_tool_call_is_reassembled_from_sse in 0927eb1. It mocks a text/event-stream body with response.output_item.added, two function_call_arguments.delta events, output_item.done and response.completed, calls completion(stream=True) on bedrock/global.openai.gpt-6-astra, and asserts the request carried stream: true, the tool name and reassembled arguments come back as tool_call deltas, and the only finish_reason is tool_calls. The same flow was also run live against us-east-2 (8 chunks, arguments assembled, finish_reason tool_calls)

…g_effort like gpt-5.x

GPT-6 Astra on Bedrock (us./global.openai.gpt-6-astra) had no price-map
entry, so bedrock/global.openai.gpt-6-astra fell to the invoke route (which
the model does not support) and `tools` was rejected or dropped before the
request left. The Converse code also only recognised `openai.gpt-5`, so
reasoning_effort on GPT-6 was mapped to Anthropic `thinking` and Bedrock
answered "Unknown parameter: 'thinking'".

- add us./global.openai.gpt-6-astra to the price map with the same shape
  as the GPT-5.6 entries (Converse, tools, reasoning, 272K tiers, cache
  rates from the model card)
- treat every openai.gpt-* except gpt-oss as an OpenAI frontier model in
  the Converse transformation

Verified live in us-east-2: completion with tools and reasoning_effort
(non-stream and stream), tool-result round trip, and Router.

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
@CrypticDriver CrypticDriver changed the title feat(bedrock): serve openai.gpt-6-astra on bedrock-runtime /openai/v1/responses feat(bedrock): add openai.gpt-6-astra profiles and route its reasoning_effort like gpt-5.x Sep 9, 2026
@CrypticDriver
CrypticDriver force-pushed the feat/bedrock-runtime-openai-responses branch from 0927eb1 to 1505bad Compare September 9, 2026 07:17
@CrypticDriver

Copy link
Copy Markdown
Contributor Author

Force-pushed a much smaller revision (1505bad). Testing the price-map entry on its own showed that GPT-6 accepts tools on Converse exactly like GPT-5.6, so the bedrock-runtime /openai/v1/responses backend, the Responses dispatch and the bridge param config are all gone. What remains is the two price-map entries plus one helper in the Converse transformation so openai.gpt-6-* gets the same reasoning_effort handling as openai.gpt-5.*; without that Bedrock rejects the request with Unknown parameter: 'thinking'. Title and description are updated. The earlier Greptile threads refer to files that no longer exist in this PR

@mateo-berri

Copy link
Copy Markdown
Contributor

Closing as superseded: #31884 landed the same Bedrock gpt-6-astra profiles and reasoning_effort routing, verified live on main, shipping from v1.102.0-rc.1. Thanks for the contribution!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants