Repository navigation
feat(bedrock): pass chat through to each model's native openai endpoint on runtime and mantle - #43581
Conversation
…nt on runtime and mantle Co-authored-by: Matthew Lapointe <mlapointe@alpha-sense.com> Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
|
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
1 similar comment
|
bugbot run |
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
…tions' into litellm_bedrock_native_endpoint_passthrough Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> # Conflicts: # litellm/utils.py
|
bugbot run |
…onses bridge Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit c22b3ab. Configure here.
TLDR
Problem this solves:
How it solves it:
supported_endpointsfrom AWS's per-model API compatibility tableIntentional product change: Qwen3, DeepSeek V3.x, Gemma 3, MiniMax M2.x, Mistral Large 3, Kimi K2/K3, and GLM 4.7/5 on Bedrock Runtime now use AWS's native
/openai/v1/chat/completionsinstead of Converse.bedrock/converse/<model>still forces ConverseStacked on #40775. The native Responses bridge for tools with reasoning is adapted from #43264 by Matthew Lapointe, credited as coauthor
User Flow
Before: a developer calling Bedrock GPT-5.6 with function tools and reasoning is silently sent through Converse, and Qwen3 or Kimi requests are translated too
model: bedrock/us.openai.gpt-5.6-sol, a function tool, andreasoning_effort: mediumbedrock/qwen.qwen3-32b-v1:0andbedrock/moonshot.kimi-k2-thinking, and Mantle GPT-5.6 always goes through ResponsesAfter: each request goes to the AWS native OpenAI endpoint the model supports, with the fewest translations
reasoning_effort: medium/openai/v1/responsesand returns a normaltool_callschat response/openai/v1/chat/completions, and plain Mantle GPT-5.6 uses native chatPre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Shared proxy configuration
Both proxies used
/home/ubuntu/bedrock_probe/head_proxy.yamlas the shared model list so all seven model identifiers were available on both sidesBefore: base SHA
4900a1ae85b314e108e6150adfb2d3acef1c0914, port 4001, PID 49065,PYTHONPATH=/home/ubuntu/bedrock_probe/litellm_base,LITELLM_LOCAL_MODEL_COST_MAP=True, imported/home/ubuntu/bedrock_probe/litellm_base/litellm/__init__.py, Qwensupported_endpointswas absent, readiness HTTP 200After: head SHA
bcf0cd93f0b960494c9b6cce6056e3ecf6837499, port 4000, PID 64177,PYTHONPATH=/home/ubuntu/repos/litellm,LITELLM_LOCAL_MODEL_COST_MAP=True, imported/home/ubuntu/repos/litellm/litellm/__init__.py, readiness HTTP 200The head process started at 18:02 UTC after the changed routing sources were modified at 17:58 UTC. The checkout is clean at bcf0cd9 and the process imports from that checkout, so no restart was needed
Cases 1 and 3 through 7 use the
pr_*.request.jsonbodies under/home/ubuntu/bedrock_probe/live_raw, withmax_tokens: 128. Case 2 uses the Kimi body atmax_tokens: 1024. Cases 8 and 9 use thefollowup_*.request.jsonbodies. The same body file was used on Before and After. Commands retain the$LITELLM_MASTER_KEYplaceholder and no key value is stored hereBefore (4900a1a)
1. bedrock/qwen.qwen3-32b-v1:0 tools + reasoning_effort=medium
2. bedrock/moonshot.kimi-k2-thinking plain
3. bedrock/us.openai.gpt-5.6-sol tools + reasoning_effort=medium
4. bedrock/us.openai.gpt-6-sol plain
5. bedrock/us.openai.gpt-6-sol tools + reasoning_effort=medium
6. bedrock_mantle/openai.gpt-5.6-sol plain
7. bedrock_mantle/openai.gpt-5.6-sol tools + reasoning_effort=medium
8. bedrock_mantle/openai.gpt-5.6-sol web_search_options
9. bedrock/us.openai.gpt-5.6-sol legacy functions + reasoning_effort=medium + reasoningSummary=auto
After (bcf0cd9)
1. bedrock/qwen.qwen3-32b-v1:0 tools + reasoning_effort=medium
2. bedrock/moonshot.kimi-k2-thinking plain
3. bedrock/us.openai.gpt-5.6-sol tools + reasoning_effort=medium
4. bedrock/us.openai.gpt-6-sol plain
5. bedrock/us.openai.gpt-6-sol tools + reasoning_effort=medium
6. bedrock_mantle/openai.gpt-5.6-sol plain
7. bedrock_mantle/openai.gpt-5.6-sol tools + reasoning_effort=medium
8. bedrock_mantle/openai.gpt-5.6-sol web_search_options
9. bedrock/us.openai.gpt-5.6-sol legacy functions + reasoning_effort=medium + reasoningSummary=auto
Type
🆕 New Feature
Caveats (if any)
Medium
xai.grok-4.6is not served in us-east-1, so it was not probed thereresponse_formaton Converse until native enforcement is verifiedLow
stream_options.include_usagereturn no usage chunk, same as beforefunctionswith reasoning on GPT-5.6 still 400 on Converse, same as basesupports_none_reasoning_effort: falseon GPT-6 rows bridgesreasoning_effort: none(string or{"effort": "none"}) with tools to ResponsesFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/59ff3c034d1b45b281c922251aec59f1
Open in Devin Desktop: https://app.devin.ai/desktop/session/59ff3c034d1b45b281c922251aec59f1?variant=devin
Requested by: @mateo-berri