Repository navigation
feat(bedrock): add openai.gpt-6-astra profiles and route its reasoning_effort like gpt-5.x - #40370
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Greptile SummaryThis PR adds native Amazon Bedrock Responses API routing for GPT-6 Astra, including SigV4 or bearer authentication, chat-to-Responses parameter shaping, model pricing, and behavioral tests
Confidence Score: 4/5The implementation appears behaviorally sound, but explicit repository requirements and streaming regression coverage should be addressed before merging No concrete runtime or security failure remains, but provider-specific logic is placed outside the required adapter boundary, unnecessary comments violate the source-comment policy, and the critical streaming flow lacks regression coverage Files Needing Attention: litellm/utils.py, litellm/litellm_core_utils/get_supported_openai_params.py, litellm/llms/bedrock/responses/transformation.py, tests/test_litellm/llms/bedrock/responses/test_bedrock_responses_transformation.py
|
| Filename | Overview |
|---|---|
| litellm/llms/bedrock/responses/transformation.py | Adds the native Bedrock Responses endpoint, URL construction, authentication, error mapping, and chat parameter configuration |
| litellm/utils.py | Adds Bedrock Responses capability routing, but places provider-specific selection logic outside the required provider directory |
| litellm/litellm_core_utils/get_supported_openai_params.py | Selects Responses-compatible Bedrock parameters through provider-specific logic outside the required provider directory |
| model_prices_and_context_window.json | Adds global and US GPT-6 Astra pricing, limits, endpoint support, and capability metadata |
| tests/test_litellm/llms/bedrock/responses/test_bedrock_responses_transformation.py | Covers URL, authentication, routing, pricing, and complete tool calls, but omits the advertised streaming tool-call flow |
Reviews (1): Last reviewed commit: "feat(bedrock): serve openai.gpt-6-astra ..." | Re-trigger Greptile
| elif litellm.LlmProviders.BEDROCK == provider: | ||
| # bedrock-runtime serves OpenAI frontier models (openai.gpt-6-astra and its | ||
| # us./global. inference profiles) on /openai/v1/responses. Data-driven from | ||
| # the price map (mode: responses / supported_endpoints); models without | ||
| # that signal (Claude, gpt-oss, ...) return None and keep the | ||
| # chat-completions emulation exactly as before. | ||
| from litellm.llms.bedrock.responses.transformation import ( | ||
| bedrock_supports_responses, | ||
| ) | ||
|
|
||
| if not bedrock_supports_responses(model, litellm.model_cost): | ||
| return None | ||
| return litellm.BedrockResponsesAPIConfig() |
There was a problem hiding this comment.
Provider Logic Outside Adapters
These Bedrock branches violate the directive requiring provider-specific logic to remain under llms/; this requirement must be satisfied before merging
Rule Used: What: Avoid writing provider-specific code outside... (source)
Knowledge Base Used: Provider adapters and capabilities
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
Reworked in 0927eb1. The chat-side decision now lives in get_bedrock_chat_config under llms/bedrock, which returns BedrockResponsesBridgeChatConfig for price-map mode: responses models. get_optional_params no longer special-cases these models; its Converse branch only applies when the resolved provider config is AmazonConverseConfig, so the bridge config goes through the generic provider_config.map_openai_params fallthrough that already exists. The get_supported_openai_params change is gone entirely. What remains in utils.py is the one-line dispatch in _get_python_responses_api_config, the same shape every other provider uses there
| # Checked longest/most-specific first so a full endpoint URL collapses to host in | ||
| # one pass and the appended path never doubles. |
There was a problem hiding this comment.
Unnecessary Explanatory Comments
These comments restate straightforward suffix and authentication code, violating the directive limiting comments to genuinely complex logic; resolve this before merging
Context Used: CLAUDE.md (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
Trimmed in 0927eb1: the suffix, URL helper and SigV4 comments are gone. What is left is the module docstring, the gate function docstring, one line explaining why validate_environment must not inherit the OPENAI_API_KEY injection, and the reason text the type-discipline gate requires on each mutable-ok suppression
| def test_completion_with_tools_is_bridged_to_responses( | ||
| self, local_model_cost_map, gpt6_in_converse_set, fake_aws_env | ||
| ): | ||
| route = respx.post("https://bedrock-runtime.us-east-2.amazonaws.com/openai/v1/responses").mock( | ||
| return_value=httpx.Response(200, json=_canned_function_call_response(GPT6_GLOBAL)) | ||
| ) | ||
|
|
||
| response = litellm.completion( | ||
| model=f"bedrock/{GPT6_GLOBAL}", | ||
| messages=[{"role": "user", "content": "weather in Beijing?"}], | ||
| tools=[WEATHER_TOOL], | ||
| ) | ||
|
|
||
| assert route.called, "request must go to bedrock-runtime /openai/v1/responses" | ||
| sent = json.loads(route.calls.last.request.content) |
There was a problem hiding this comment.
The advertised streaming tool-call flow only has complete JSON tests. Add an SSE regression covering tool deltas and finish_reason before merging
Context Used: CLAUDE.md (source)
Knowledge Base Used: Model request execution
There was a problem hiding this comment.
Added test_streamed_tool_call_is_reassembled_from_sse in 0927eb1. It mocks a text/event-stream body with response.output_item.added, two function_call_arguments.delta events, output_item.done and response.completed, calls completion(stream=True) on bedrock/global.openai.gpt-6-astra, and asserts the request carried stream: true, the tool name and reassembled arguments come back as tool_call deltas, and the only finish_reason is tool_calls. The same flow was also run live against us-east-2 (8 chunks, arguments assembled, finish_reason tool_calls)
…g_effort like gpt-5.x GPT-6 Astra on Bedrock (us./global.openai.gpt-6-astra) had no price-map entry, so bedrock/global.openai.gpt-6-astra fell to the invoke route (which the model does not support) and `tools` was rejected or dropped before the request left. The Converse code also only recognised `openai.gpt-5`, so reasoning_effort on GPT-6 was mapped to Anthropic `thinking` and Bedrock answered "Unknown parameter: 'thinking'". - add us./global.openai.gpt-6-astra to the price map with the same shape as the GPT-5.6 entries (Converse, tools, reasoning, 272K tiers, cache rates from the model card) - treat every openai.gpt-* except gpt-oss as an OpenAI frontier model in the Converse transformation Verified live in us-east-2: completion with tools and reasoning_effort (non-stream and stream), tool-result round trip, and Router. Generated with [Claude Code](https://claude.ai/code) via [Happy](https://happy.engineering) Co-Authored-By: Claude <noreply@anthropic.com> Co-Authored-By: Happy <yesreply@happy.engineering>
0927eb1 to
1505bad
Compare
|
Force-pushed a much smaller revision (1505bad). Testing the price-map entry on its own showed that GPT-6 accepts tools on Converse exactly like GPT-5.6, so the bedrock-runtime /openai/v1/responses backend, the Responses dispatch and the bridge param config are all gone. What remains is the two price-map entries plus one helper in the Converse transformation so openai.gpt-6-* gets the same reasoning_effort handling as openai.gpt-5.*; without that Bedrock rejects the request with Unknown parameter: 'thinking'. Title and description are updated. The earlier Greptile threads refer to files that no longer exist in this PR |
|
Closing as superseded: #31884 landed the same Bedrock gpt-6-astra profiles and reasoning_effort routing, verified live on main, shipping from v1.102.0-rc.1. Thanks for the contribution! |
TLDR
Problem this solves:
bedrock/global.openai.gpt-6-astra(andus.) cannot call tools through LiteLLMtoolsis rejected or droppedreasoning_effortis mapped to Anthropicthinkingbecause the Converse code only recognisesopenai.gpt-5How it solves it:
us./global.openai.gpt-6-astrato the price map, same shape as the GPT-5.6 entriesopenai.gpt-*(exceptgpt-oss) as an OpenAI frontier model in the Converse transformationUser Flow
Before: a developer pointing their agent at GPT-6 Astra on Bedrock through the gateway cannot get a single tool call back
"model": "bedrock/global.openai.gpt-6-astra", a user message and atoolsarraybedrock does not support parameters: ['tools'], for model=global.openai.gpt-6-astradrop_params, the tools are silently removed, the request goes to.../model/global.openai.gpt-6-astra/invoke, and AWS rejects it because this model does not support Invoketool_callsmessageAfter: the same request comes back with a real tool call
"model": "bedrock/global.openai.gpt-6-astra"and the sametoolsarrayhttps://bedrock-runtime.us-east-2.amazonaws.com/model/global.openai.gpt-6-astra/conversewith atoolConfigfinish_reason: "tool_calls"andtool_calls[0].function.name == "get_weather";stream: truereturns the same as SSE deltas"reasoning_effort": "low"now reaches Bedrock asreasoning.effortinstead of failing withUnknown parameter: 'thinking'Relevant issues
Related (not closed by this PR): #40246, #40080
Live verification
Run against a real account in
us-east-2with an IAM role, using the bundled price map (LITELLM_LOCAL_MODEL_COST_MAP=True):completion()+ tools +reasoning_effort=low, non-streamfinish_reason=tool_calls,get_weather {"city": "Beijing"}, cost 0.00142completion()+ tools +reasoning_effort=low,stream=True(us.profile)finish_reason=tool_callsRouterwith amodel_listentryWithout the Converse change the
reasoning_effortcase fails withBedrockException {"code":"unknown_parameter","message":"Unknown parameter: 'thinking'."}Notes
/openai/v1/responsesbackend. That turned out to be unnecessary for tool calling: GPT-6 supports tools on Converse exactly like GPT-5.6, so this revision is the price-map entry plus the one-line model check, the same shape as the GPT-5.6 onboarding. A native Responses path can be a separate PR if a Responses-only feature is neededus.entry uses the Geo CRIS rate andglobal.the Global CRIS rate from the model card; both carry the 272K tiers and cache write/read rates like the GPT-5.6 entriesPre-Submission checklist
tests/test_litellm/llms/bedrock/test_cross_region_inference_profile_mapping.py,tests/test_litellm/llms/bedrock/chat/test_converse_transformation.py)tests/test_litellm/llms/bedrock/chat,tests/test_litellm/llms/bedrock/test_bedrock_common_utils.py,tests/test_litellm/llms/bedrock_mantle,tests/test_litellm/completion_extras,tests/test_litellm/test_model_prices_schema.py.scripts/ruff_strict_gate.py,scripts/type_discipline_gate.pyandscripts/type_check_gate.pypass against the merge-base