Skip to content

feat(bedrock): pass chat through to each model's native openai endpoint on runtime and mantle - #43581

Closed
devin-ai-integration[bot] wants to merge 6 commits into
litellm_bedrock_grok_chat_completionsfrom
litellm_bedrock_native_endpoint_passthrough
Closed

devin-ai-integration[bot] wants to merge 6 commits into
litellm_bedrock_grok_chat_completionsfrom
litellm_bedrock_native_endpoint_passthrough

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Many Bedrock models still go through Converse despite AWS serving native OpenAI endpoints
  • GPT-5.6 and GPT-6 tools with reasoning fall back to Converse
  • Mantle GPT-5.x and GPT-6 always go through the Responses bridge

How it solves it:

  • Backfill supported_endpoints from AWS's per-model API compatibility table
  • Bridge tools with reasoning, or web search, to native Responses when the row lists it
  • Mantle chat-capable GPT rows switch to native chat, bridging only when needed
  • One neutral capability flag covers Runtime and Mantle, read from data

Intentional product change: Qwen3, DeepSeek V3.x, Gemma 3, MiniMax M2.x, Mistral Large 3, Kimi K2/K3, and GLM 4.7/5 on Bedrock Runtime now use AWS's native /openai/v1/chat/completions instead of Converse. bedrock/converse/<model> still forces Converse

Stacked on #40775. The native Responses bridge for tools with reasoning is adapted from #43264 by Matthew Lapointe, credited as coauthor

User Flow

Before: a developer calling Bedrock GPT-5.6 with function tools and reasoning is silently sent through Converse, and Qwen3 or Kimi requests are translated too

  1. They send POST http://localhost:4000/v1/chat/completions with model: bedrock/us.openai.gpt-5.6-sol, a function tool, and reasoning_effort: medium
  2. The proxy rewrites the request into Converse's shape instead of AWS's OpenAI-compatible endpoints
  3. The same happens for bedrock/qwen.qwen3-32b-v1:0 and bedrock/moonshot.kimi-k2-thinking, and Mantle GPT-5.6 always goes through Responses

After: each request goes to the AWS native OpenAI endpoint the model supports, with the fewest translations

  1. They send the same POST http://localhost:4000/v1/chat/completions with the GPT-5.6 function tool and reasoning_effort: medium
  2. The proxy sends it to AWS's native /openai/v1/responses and returns a normal tool_calls chat response
  3. Qwen3 and Kimi go to native /openai/v1/chat/completions, and plain Mantle GPT-5.6 uses native chat

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Shared proxy configuration

Both proxies used /home/ubuntu/bedrock_probe/head_proxy.yaml as the shared model list so all seven model identifiers were available on both sides

Before: base SHA 4900a1ae85b314e108e6150adfb2d3acef1c0914, port 4001, PID 49065, PYTHONPATH=/home/ubuntu/bedrock_probe/litellm_base, LITELLM_LOCAL_MODEL_COST_MAP=True, imported /home/ubuntu/bedrock_probe/litellm_base/litellm/__init__.py, Qwen supported_endpoints was absent, readiness HTTP 200

After: head SHA bcf0cd93f0b960494c9b6cce6056e3ecf6837499, port 4000, PID 64177, PYTHONPATH=/home/ubuntu/repos/litellm, LITELLM_LOCAL_MODEL_COST_MAP=True, imported /home/ubuntu/repos/litellm/litellm/__init__.py, readiness HTTP 200

The head process started at 18:02 UTC after the changed routing sources were modified at 17:58 UTC. The checkout is clean at bcf0cd9 and the process imports from that checkout, so no restart was needed

Cases 1 and 3 through 7 use the pr_*.request.json bodies under /home/ubuntu/bedrock_probe/live_raw, with max_tokens: 128. Case 2 uses the Kimi body at max_tokens: 1024. Cases 8 and 9 use the followup_*.request.json bodies. The same body file was used on Before and After. Commands retain the $LITELLM_MASTER_KEY placeholder and no key value is stored here

Before (4900a1a)

1. bedrock/qwen.qwen3-32b-v1:0 tools + reasoning_effort=medium

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_before_shared_qwen_tools.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_before_shared_qwen_tools.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_qwen_tools.request.json \
  http://127.0.0.1:4001/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: tool_calls
content: empty
tool_call: emit_probe args={"value": "LIVE_qwen_qwen3-32b-v1_0_TOOLS"}
reasoning_content_length: 0
usage: prompt_tokens=185 completion_tokens=36 total_tokens=221
x-litellm-response-cost: 4.935e-05
  1. Upstream: https://bedrock-runtime.us-east-1.amazonaws.com/model/qwen.qwen3-32b-v1%3A0/converse

2. bedrock/moonshot.kimi-k2-thinking plain

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_before_kimi_1024.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_before_kimi_1024.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_kimi_plain_1024.request.json \
  http://127.0.0.1:4001/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: stop
content: LIVE_moonshot_kimi-k2-thinking_PLAIN
reasoning_content_length: 725
usage: prompt_tokens=27 completion_tokens=181 total_tokens=208
x-litellm-response-cost: 0.0004687
  1. Upstream: https://bedrock-runtime.us-east-1.amazonaws.com/model/moonshot.kimi-k2-thinking/converse

3. bedrock/us.openai.gpt-5.6-sol tools + reasoning_effort=medium

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_before_shared_runtime_gpt56_tools.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_before_shared_runtime_gpt56_tools.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_runtime_gpt56_tools.request.json \
  http://127.0.0.1:4001/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: tool_calls
content: empty
tool_call: emit_probe args={"value": "LIVE_us_openai_gpt-5_6-sol_TOOLS"}
reasoning_content_length: 0
usage: prompt_tokens=68 completion_tokens=30 total_tokens=98
x-litellm-response-cost: 0.0009592
  1. Upstream: https://bedrock-runtime.us-east-1.amazonaws.com/model/us.openai.gpt-5.6-sol/converse

4. bedrock/us.openai.gpt-6-sol plain

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_before_shared_runtime_gpt6_plain.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_before_shared_runtime_gpt6_plain.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_runtime_gpt6_plain.request.json \
  http://127.0.0.1:4001/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: stop
content: LIVE_us_openai_gpt-6-sol_PLAIN
reasoning_content_length: 0
usage: prompt_tokens=26 completion_tokens=15 total_tokens=41
x-litellm-response-cost: 0.0002222
  1. Upstream: https://bedrock-runtime.us-east-1.amazonaws.com/model/us.openai.gpt-6-sol/converse

5. bedrock/us.openai.gpt-6-sol tools + reasoning_effort=medium

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_before_shared_runtime_gpt6_tools.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_before_shared_runtime_gpt6_tools.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_runtime_gpt6_tools.request.json \
  http://127.0.0.1:4001/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: tool_calls
content: empty
tool_call: emit_probe args={"value": "LIVE_us_openai_gpt-6-sol_TOOLS"}
reasoning_content_length: 0
usage: prompt_tokens=66 completion_tokens=28 total_tokens=94
x-litellm-response-cost: 0.0004532
  1. Upstream: https://bedrock-runtime.us-east-1.amazonaws.com/model/us.openai.gpt-6-sol/converse

6. bedrock_mantle/openai.gpt-5.6-sol plain

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_before_shared_mantle_gpt56_plain.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_before_shared_mantle_gpt56_plain.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_mantle_gpt56_plain.request.json \
  http://127.0.0.1:4001/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: stop
content: LIVE_openai_gpt-5_6-sol_PLAIN
reasoning_content_length: 0
usage: prompt_tokens=27 completion_tokens=16 total_tokens=43
x-litellm-response-cost: 0.0004708
  1. Upstream: https://bedrock-mantle.us-east-1.api.aws/openai/v1/responses

7. bedrock_mantle/openai.gpt-5.6-sol tools + reasoning_effort=medium

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_before_shared_mantle_gpt56_tools.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_before_shared_mantle_gpt56_tools.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_mantle_gpt56_tools.request.json \
  http://127.0.0.1:4001/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: tool_calls
content: absent
tool_call: emit_probe args={"value":"LIVE_openai_gpt-5_6-sol_TOOLS"}
reasoning_content_length: 0
usage: prompt_tokens=67 completion_tokens=29 total_tokens=96
x-litellm-response-cost: 0.0009328
  1. Upstream: https://bedrock-mantle.us-east-1.api.aws/openai/v1/responses

8. bedrock_mantle/openai.gpt-5.6-sol web_search_options

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_before_4900a1a_mantle_web_search.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_before_4900a1a_mantle_web_search.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/followup_mantle_web_search.request.json \
  http://127.0.0.1:4001/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: stop
content: Seattle, Washington is currently 55°F (13°C), with a feels-like temperature of 55°F. Today's forecast calls for a high near 66°F. Source: FOX 13 Seattle Weather, updated at 10:43 a.m. PDT
reasoning_content_length: 0
usage: prompt_tokens=8719 completion_tokens=174 total_tokens=8893
x-litellm-response-cost: 0.042731359999999996
upstream_tools: [{"type": "web_search"}]
  1. Upstream: https://bedrock-mantle.us-east-1.api.aws/openai/v1/responses

9. bedrock/us.openai.gpt-5.6-sol legacy functions + reasoning_effort=medium + reasoningSummary=auto

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_before_4900a1a_runtime_legacy_functions.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_before_4900a1a_runtime_legacy_functions.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/followup_runtime_legacy_functions.request.json \
  http://127.0.0.1:4001/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 400
error: unknown_parameter
message: Unknown parameter: 'functions'
finish_reason: not available
content and tool_calls: not available
reasoning_content_length: not available
usage: not available
x-litellm-response-cost: 0
  1. Upstream: https://bedrock-runtime.us-east-1.amazonaws.com/model/us.openai.gpt-5.6-sol/converse

After (bcf0cd9)

1. bedrock/qwen.qwen3-32b-v1:0 tools + reasoning_effort=medium

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_qwen_tools.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_qwen_tools.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_qwen_tools.request.json \
  http://127.0.0.1:4000/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: tool_calls
content: absent
tool_call: emit_probe args={"value": "LIVE_qwen_qwen3-32b-v1_0_TOOLS"}
reasoning_content_length: 0
usage: prompt_tokens=190 completion_tokens=36 total_tokens=226
x-litellm-response-cost: 5.01e-05
  1. Upstream: https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1/chat/completions

2. bedrock/moonshot.kimi-k2-thinking plain

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_kimi_1024.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_kimi_1024.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_kimi_plain_1024.request.json \
  http://127.0.0.1:4000/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: stop
content: LIVE_moonshot_kimi-k2-thinking_PLAIN
reasoning_content_length: 317
usage: prompt_tokens=27 completion_tokens=85 total_tokens=112
x-litellm-response-cost: 0.00022870000000000003
  1. Upstream: https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1/chat/completions

3. bedrock/us.openai.gpt-5.6-sol tools + reasoning_effort=medium

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_runtime_gpt56_tools.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_runtime_gpt56_tools.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_runtime_gpt56_tools.request.json \
  http://127.0.0.1:4000/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: tool_calls
content: absent
tool_call: emit_probe args={"value":"LIVE_us_openai_gpt-5_6-sol_TOOLS"}
reasoning_content_length: 0
usage: prompt_tokens=68 completion_tokens=30 total_tokens=98
x-litellm-response-cost: 0.0009592
  1. Upstream: https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1/responses

4. bedrock/us.openai.gpt-6-sol plain

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_runtime_gpt6_plain.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_runtime_gpt6_plain.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_runtime_gpt6_plain.request.json \
  http://127.0.0.1:4000/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: stop
content: LIVE_us_openai_gpt-6-sol_PLAIN
reasoning_content_length: 0
usage: prompt_tokens=26 completion_tokens=25 total_tokens=51
x-litellm-response-cost: 0.00033219999999999994
  1. Upstream: https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1/chat/completions

5. bedrock/us.openai.gpt-6-sol tools + reasoning_effort=medium

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_runtime_gpt6_tools.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_runtime_gpt6_tools.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_runtime_gpt6_tools.request.json \
  http://127.0.0.1:4000/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: tool_calls
content: absent
tool_call: emit_probe args={"value":"LIVE_us_openai_gpt-6-sol_TOOLS"}
reasoning_content_length: 0
usage: prompt_tokens=66 completion_tokens=28 total_tokens=94
x-litellm-response-cost: 0.0004532
  1. Upstream: https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1/responses

6. bedrock_mantle/openai.gpt-5.6-sol plain

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_mantle_gpt56_plain.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_mantle_gpt56_plain.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_mantle_gpt56_plain.request.json \
  http://127.0.0.1:4000/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: stop
content: LIVE_openai_gpt-5_6-sol_PLAIN
reasoning_content_length: 0
usage: prompt_tokens=27 completion_tokens=16 total_tokens=43
x-litellm-response-cost: 0.0004708
  1. Upstream: https://bedrock-mantle.us-east-1.api.aws/openai/v1/chat/completions

7. bedrock_mantle/openai.gpt-5.6-sol tools + reasoning_effort=medium

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_mantle_gpt56_tools.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf_mantle_gpt56_tools.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/pr_mantle_gpt56_tools.request.json \
  http://127.0.0.1:4000/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: tool_calls
content: absent
tool_call: emit_probe args={"value":"LIVE_openai_gpt-5_6-sol_TOOLS"}
reasoning_content_length: 0
usage: prompt_tokens=67 completion_tokens=29 total_tokens=96
x-litellm-response-cost: 0.0009328
  1. Upstream: https://bedrock-mantle.us-east-1.api.aws/openai/v1/responses

8. bedrock_mantle/openai.gpt-5.6-sol web_search_options

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf0cd9_mantle_web_search.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf0cd9_mantle_web_search.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/followup_mantle_web_search.request.json \
  http://127.0.0.1:4000/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 200
finish_reason: stop
content: Seattle's current temperature is 56°F (13°C), with sunny conditions, as of 10:55 a.m. PDT today. Source: Weather Underground
reasoning_content_length: 0
usage: prompt_tokens=8559 completion_tokens=208 total_tokens=8767
x-litellm-response-cost: 0.04277536
upstream_tools: [{"type": "web_search"}]
  1. Upstream: https://bedrock-mantle.us-east-1.api.aws/openai/v1/responses

9. bedrock/us.openai.gpt-5.6-sol legacy functions + reasoning_effort=medium + reasoningSummary=auto

  1. command
curl -sS --max-time 180 \
  -D /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf0cd9_runtime_legacy_functions.headers \
  -o /home/ubuntu/bedrock_probe/live_raw/pr_after_bcf0cd9_runtime_legacy_functions.body \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data-binary @/home/ubuntu/bedrock_probe/live_raw/followup_runtime_legacy_functions.request.json \
  http://127.0.0.1:4000/v1/chat/completions \
  -w '%{http_code}\n'
  1. observed output
HTTP status: 400
error: unknown_parameter
message: Unknown parameter: 'functions'
finish_reason: not available
content and tool_calls: not available
reasoning_content_length: not available
usage: not available
x-litellm-response-cost: 0
  1. Upstream: https://bedrock-runtime.us-east-1.amazonaws.com/model/us.openai.gpt-5.6-sol/converse

Type

🆕 New Feature

Caveats (if any)

Medium

  • Endpoint lists come from AWS docs and live probes on 2026-09-28
    • Region-specific identifiers or provisioned throughput can still reject a listed endpoint
  • Mantle xai.grok-4.6 is not served in us-east-1, so it was not probed there
  • New families keep response_format on Converse until native enforcement is verified

Low

  • Streams without stream_options.include_usage return no usage chunk, same as before
  • Legacy functions with reasoning on GPT-5.6 still 400 on Converse, same as base
  • supports_none_reasoning_effort: false on GPT-6 rows bridges reasoning_effort: none (string or {"effort": "none"}) with tools to Responses
  • Region-path Runtime ids keep tools with reasoning on Converse, since native Responses does not read the path region yet
  • Mantle Responses support now also counts a regional-only row, matching the other capability flags

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/59ff3c034d1b45b281c922251aec59f1
Open in Devin Desktop: https://app.devin.ai/desktop/session/59ff3c034d1b45b281c922251aec59f1?variant=devin
Requested by: @mateo-berri

…nt on runtime and mantle

Co-authored-by: Matthew Lapointe <mlapointe@alpha-sense.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@codecov

codecov Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.08917% with 3 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/litellm_core_utils/responses_api_utils.py 94.87% 2 Missing ⚠️
litellm/llms/bedrock_mantle/common_utils.py 95.65% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[High risk] Refactors Bedrock chat routing and model capability detection.

The PR appears safe to merge based on the changes reviewed.

Summary

The PR expands Bedrock Runtime native-chat coverage and selects native Responses for requests that need it. It also moves eligible Mantle chat calls off the default Responses bridge and updates model capability data and tests.

  • The changes since the previous review refine reasoning-effort handling, reuse Mantle price-map keys, and add regional and dictionary-form reasoning tests.
  • No new actionable issue was established. The previous findings are resolved or were withdrawn.

Reviews (6) · Last reviewed commit: "fix(bedrock): honor regional Mantle rows..."

Comment thread litellm/llms/bedrock/common_utils.py
Comment thread model_prices_and_context_window.json
Comment thread litellm/main.py Outdated
Comment thread litellm/main.py
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

1 similar comment
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor

cursor Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

…tions' into litellm_bedrock_native_endpoint_passthrough

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/utils.py
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

mateo-berri
mateo-berri previously approved these changes Sep 28, 2026

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri
mateo-berri dismissed their stale review September 28, 2026 19:56

Holding until Bugbot reviews

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/llms/bedrock/common_utils.py
Comment thread litellm/llms/bedrock_mantle/common_utils.py
Comment thread litellm/litellm_core_utils/responses_api_utils.py
…onses bridge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit c22b3ab. Configure here.

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

Closing as superseded. #44307 shipped native chat for GPT 5.6+ and #45473 for Grok, while other families stay on Converse unless prefixed chat_completions/

@mateo-berri mateo-berri closed this Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant