From 886f41a5ce86798cc94ac49c60de3bfcd704cd8e Mon Sep 17 00:00:00 2001 From: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Date: Sat, 19 Sep 2026 17:42:49 -0700 Subject: [PATCH 01/10] docs(bedrock): document the native chat completions route for the OpenAI and Grok models --- docs/providers/bedrock.md | 69 ++++++++++++++++++++++++++++++++++----- 1 file changed, 61 insertions(+), 8 deletions(-) diff --git a/docs/providers/bedrock.md b/docs/providers/bedrock.md index c57d08099..63d1d4b64 100644 --- a/docs/providers/bedrock.md +++ b/docs/providers/bedrock.md @@ -7,7 +7,7 @@ ALL Bedrock models (Anthropic, Meta, Deepseek, Mistral, Amazon, etc.) are Suppor | Property | Details | |-------|-------| | Description | Amazon Bedrock is a fully managed service that offers a choice of high-performing foundation models (FMs). | -| Provider Route on LiteLLM | `bedrock/`, [`bedrock/converse/`](#set-converse--invoke-route), [`bedrock/invoke/`](/docs/providers/bedrock#set-converse--invoke-route), [`bedrock/converse_like/`](/docs/providers/bedrock#calling-via-internal-proxy-not-bedrock-url-compatible), `bedrock/llama/`, `bedrock/deepseek_r1/`, `bedrock/qwen3/`, [`bedrock/qwen2/`](./bedrock_imported.md#qwen2-imported-models), [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc), [`bedrock/moonshot`](./bedrock_imported.md#moonshot-kimi-k2-thinking) | +| Provider Route on LiteLLM | `bedrock/` ([native Chat Completions](#native-chat-completions-route) for the OpenAI and Grok models), [`bedrock/converse/`](#set-converse--invoke-route), [`bedrock/invoke/`](/docs/providers/bedrock#set-converse--invoke-route), [`bedrock/converse_like/`](/docs/providers/bedrock#calling-via-internal-proxy-not-bedrock-url-compatible), `bedrock/llama/`, `bedrock/deepseek_r1/`, `bedrock/qwen3/`, [`bedrock/qwen2/`](./bedrock_imported.md#qwen2-imported-models), [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc), [`bedrock/moonshot`](./bedrock_imported.md#moonshot-kimi-k2-thinking) | | Provider Doc | [Amazon Bedrock ↗](https://docs.aws.amazon.com/bedrock/latest/userguide/what-is-bedrock.html) | | Supported OpenAI Endpoints | `/chat/completions`, `/completions`, `/embeddings`, `/images/generations`, `/v1/realtime`| | Rerank Endpoint | `/rerank` | @@ -1644,6 +1644,8 @@ LiteLLM defaults to the `invoke` route. LiteLLM uses the `converse` route for Be To explicitly set the route, do `bedrock/converse/` or `bedrock/invoke/`. +The models AWS serves in the OpenAI format are the exception. See [Native Chat Completions route](#native-chat-completions-route). + E.g. @@ -1669,6 +1671,57 @@ model_list: +## Native Chat Completions route + +AWS serves some Bedrock models on an OpenAI-compatible endpoint, `https://bedrock-runtime.{region}.amazonaws.com/openai/v1/chat/completions`. For those models LiteLLM sends your `/chat/completions` request to that endpoint in the shape it arrived in, instead of translating it to Converse and back. Fewer translations means less latency and fewer places for a parameter to get lost. + +| Model | LiteLLM model name | Default route | +|-------|--------------------|---------------| +| GPT-OSS 20B | `bedrock/openai.gpt-oss-20b-1:0` | Native Chat Completions | +| GPT-OSS 120B | `bedrock/openai.gpt-oss-120b-1:0` | Native Chat Completions | +| GPT-5.6 Sol, Terra, Luna | `bedrock/us.openai.gpt-5.6-sol`, `bedrock/global.openai.gpt-5.6-sol`, and the `terra` / `luna` variants | Native Chat Completions | +| Grok 4.6 | `bedrock/us.xai.grok-4.6`, `bedrock/global.xai.grok-4.6`, `bedrock/us-gov.xai.grok-4.6` | Native Chat Completions | +| Everything else (Claude, Nova, Llama, Mistral, ...) | `bedrock/` | Converse or Invoke, as before | + +A model opts in through `"use_bedrock_runtime_chat_completions": true` on its entry in the [model cost map](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json), so a model AWS lists without Chat Completions support keeps using Converse. Authentication, regions, `aws_bedrock_runtime_endpoint`, and cost tracking work the same on both routes. [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc) is a separate route for imported models and is unchanged. + +The trade-off is that the OpenAI-compatible endpoint has no equivalent for a few Converse features, so LiteLLM falls back to Converse per request when you use one of them: + +| Request | Route used | Why | +|---------|------------|-----| +| `guardrailConfig` in the request body | Converse | AWS takes guardrails on the OpenAI-compatible endpoint as `X-Amzn-Bedrock-Guardrail*` headers and rejects a `guardrailConfig` body field, so LiteLLM keeps those requests on Converse and your guardrail behavior does not change | +| `requestMetadata`, `performanceConfig`, `serviceTier`, or `outputConfig` in the request body | Converse | These Converse body fields are rejected as malformed input on the OpenAI-compatible endpoint | +| `bedrock_request_metadata_fields` set in `litellm_settings` | Converse, for every request | LiteLLM only writes the operator's request metadata onto the Converse body | +| Application inference profile ARN as the model | Converse | LiteLLM cannot tell from the ARN which model it fronts | +| GPT-5.6 with function `tools` and `reasoning_effort` other than `"none"` (or unset) | Converse | AWS only accepts function tools on Chat Completions for GPT-5.6 when `reasoning_effort` is `"none"` | +| `bedrock/converse/` | Converse | You asked for it explicitly | + +What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids look like `call_0` instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, `n` greater than 1 is unsupported (as on Converse), and GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field. + +To keep a model on Converse for every request, set the route explicitly: + + + + +```python +from litellm import completion + +completion(model="bedrock/converse/openai.gpt-oss-20b-1:0", messages=[{"role": "user", "content": "Hello"}]) +``` + + + + +```yaml +model_list: + - model_name: gpt-oss-20b-converse + litellm_params: + model: bedrock/converse/openai.gpt-oss-20b-1:0 +``` + + + + ## Alternate user/assistant messages Use `user_continue_message` to add a default user message, for cases (e.g. Autogen) where the client might not follow alternating user/assistant messages starting and ending with a user message. @@ -1900,7 +1953,7 @@ curl -X POST 'http://0.0.0.0:4000/chat/completions' \ | Property | Details | |----------|---------| -| Provider Route | `bedrock/converse/openai.gpt-oss-20b-1:0`, `bedrock/converse/openai.gpt-oss-120b-1:0` | +| Provider Route | `bedrock/openai.gpt-oss-20b-1:0`, `bedrock/openai.gpt-oss-120b-1:0` ([native Chat Completions](#native-chat-completions-route)); prefix with `bedrock/converse/` to force Converse | | Provider Documentation | [Amazon Bedrock ↗](https://docs.aws.amazon.com/bedrock/latest/userguide/what-is-bedrock.html) | @@ -1917,14 +1970,14 @@ os.environ["AWS_REGION_NAME"] = "us-east-1" # GPT OSS 20B model response = completion( - model="bedrock/converse/openai.gpt-oss-20b-1:0", + model="bedrock/openai.gpt-oss-20b-1:0", messages=[{"role": "user", "content": "Hello, how are you?"}], ) print(response.choices[0].message.content) # GPT OSS 120B model response = completion( - model="bedrock/converse/openai.gpt-oss-120b-1:0", + model="bedrock/openai.gpt-oss-120b-1:0", messages=[{"role": "user", "content": "Explain machine learning in simple terms"}], ) print(response.choices[0].message.content) @@ -1940,14 +1993,14 @@ print(response.choices[0].message.content) model_list: - model_name: gpt-oss-20b litellm_params: - model: bedrock/converse/openai.gpt-oss-20b-1:0 + model: bedrock/openai.gpt-oss-20b-1:0 aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY aws_region_name: os.environ/AWS_REGION_NAME - model_name: gpt-oss-120b litellm_params: - model: bedrock/converse/openai.gpt-oss-120b-1:0 + model: bedrock/openai.gpt-oss-120b-1:0 aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY aws_region_name: os.environ/AWS_REGION_NAME @@ -2140,8 +2193,8 @@ Here's an example of using a bedrock model with LiteLLM. For a complete list, re | Model Name | Command | |----------------------------|------------------------------------------------------------------| -| GPT-OSS 20B | `completion(model='bedrock/converse/openai.gpt-oss-20b-1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']`, `os.environ['AWS_REGION_NAME']` | -| GPT-OSS 120B | `completion(model='bedrock/converse/openai.gpt-oss-120b-1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']`, `os.environ['AWS_REGION_NAME']` | +| GPT-OSS 20B | `completion(model='bedrock/openai.gpt-oss-20b-1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']`, `os.environ['AWS_REGION_NAME']` | +| GPT-OSS 120B | `completion(model='bedrock/openai.gpt-oss-120b-1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']`, `os.environ['AWS_REGION_NAME']` | | Deepseek R1 | `completion(model='bedrock/us.deepseek.r1-v1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']` | | Anthropic Claude Sonnet 4.5 | `completion(model='bedrock/us.anthropic.claude-sonnet-4-5-20250929-v1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']` | | Anthropic Claude-V3.5 Sonnet | `completion(model='bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']` | From 350399323182e76acaa2d79a89f4b6e6f9f154f6 Mon Sep 17 00:00:00 2001 From: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Date: Sat, 19 Sep 2026 18:26:31 -0700 Subject: [PATCH 02/10] docs(bedrock): name the supports_ prefixed cost map flags for the native chat completions route --- docs/providers/bedrock.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/providers/bedrock.md b/docs/providers/bedrock.md index 63d1d4b64..70a75b85a 100644 --- a/docs/providers/bedrock.md +++ b/docs/providers/bedrock.md @@ -1683,7 +1683,7 @@ AWS serves some Bedrock models on an OpenAI-compatible endpoint, `https://bedroc | Grok 4.6 | `bedrock/us.xai.grok-4.6`, `bedrock/global.xai.grok-4.6`, `bedrock/us-gov.xai.grok-4.6` | Native Chat Completions | | Everything else (Claude, Nova, Llama, Mistral, ...) | `bedrock/` | Converse or Invoke, as before | -A model opts in through `"use_bedrock_runtime_chat_completions": true` on its entry in the [model cost map](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json), so a model AWS lists without Chat Completions support keeps using Converse. Authentication, regions, `aws_bedrock_runtime_endpoint`, and cost tracking work the same on both routes. [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc) is a separate route for imported models and is unchanged. +A model opts in through `"supports_bedrock_runtime_chat_completions": true` on its entry in the [model cost map](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json), so a model AWS lists without Chat Completions support keeps using Converse. Authentication, regions, `aws_bedrock_runtime_endpoint`, and cost tracking work the same on both routes. [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc) is a separate route for imported models and is unchanged. The trade-off is that the OpenAI-compatible endpoint has no equivalent for a few Converse features, so LiteLLM falls back to Converse per request when you use one of them: @@ -1693,7 +1693,7 @@ The trade-off is that the OpenAI-compatible endpoint has no equivalent for a few | `requestMetadata`, `performanceConfig`, `serviceTier`, or `outputConfig` in the request body | Converse | These Converse body fields are rejected as malformed input on the OpenAI-compatible endpoint | | `bedrock_request_metadata_fields` set in `litellm_settings` | Converse, for every request | LiteLLM only writes the operator's request metadata onto the Converse body | | Application inference profile ARN as the model | Converse | LiteLLM cannot tell from the ARN which model it fronts | -| GPT-5.6 with function `tools` and `reasoning_effort` other than `"none"` (or unset) | Converse | AWS only accepts function tools on Chat Completions for GPT-5.6 when `reasoning_effort` is `"none"` | +| Function `tools` with `reasoning_effort` other than `"none"` (or unset), on a model without `"supports_bedrock_runtime_chat_completions_tools_with_reasoning": true` in the cost map (the GPT-5.6 family today) | Converse | AWS only accepts function tools on Chat Completions for GPT-5.6 when `reasoning_effort` is `"none"`; GPT-OSS and Grok carry the flag and take tools with any effort | | `bedrock/converse/` | Converse | You asked for it explicitly | What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids look like `call_0` instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, `n` greater than 1 is unsupported (as on Converse), and GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field. From 7b0ad709f077f4a86276191f9be65e66a24aa511 Mon Sep 17 00:00:00 2001 From: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Date: Sat, 19 Sep 2026 18:42:17 -0700 Subject: [PATCH 03/10] docs(bedrock): name the response_format fallback to Converse for gpt-oss on the native route --- docs/providers/bedrock.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/docs/providers/bedrock.md b/docs/providers/bedrock.md index 70a75b85a..3214446ee 100644 --- a/docs/providers/bedrock.md +++ b/docs/providers/bedrock.md @@ -1694,9 +1694,10 @@ The trade-off is that the OpenAI-compatible endpoint has no equivalent for a few | `bedrock_request_metadata_fields` set in `litellm_settings` | Converse, for every request | LiteLLM only writes the operator's request metadata onto the Converse body | | Application inference profile ARN as the model | Converse | LiteLLM cannot tell from the ARN which model it fronts | | Function `tools` with `reasoning_effort` other than `"none"` (or unset), on a model without `"supports_bedrock_runtime_chat_completions_tools_with_reasoning": true` in the cost map (the GPT-5.6 family today) | Converse | AWS only accepts function tools on Chat Completions for GPT-5.6 when `reasoning_effort` is `"none"`; GPT-OSS and Grok carry the flag and take tools with any effort | +| `response_format` other than `{"type": "text"}` (a JSON schema, `json_object`, or a Pydantic model), on a model without `"supports_bedrock_runtime_chat_completions_response_format": true` in the cost map (GPT-OSS today) | Converse | AWS accepts `response_format` for GPT-OSS on Chat Completions but answers with free text anyway, so LiteLLM keeps the Converse emulation (a forced `json_tool_call` tool) that returns the JSON you asked for; GPT-5.6 and Grok carry the flag and enforce the schema natively | | `bedrock/converse/` | Converse | You asked for it explicitly | -What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids look like `call_0` instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, `n` greater than 1 is unsupported (as on Converse), and GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field. +What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids look like `call_0` instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, `response_format` and `service_tier` are sent through as you wrote them, `n` greater than 1 is unsupported (as on Converse), and GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field. To keep a model on Converse for every request, set the route explicitly: From 9da7554e69e95f7d65fbef0c08c42d031f0acae8 Mon Sep 17 00:00:00 2001 From: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Date: Sat, 19 Sep 2026 20:35:38 -0700 Subject: [PATCH 04/10] docs(bedrock): region paths and GovCloud gpt-oss profile ids take the native Chat Completions route --- docs/providers/bedrock.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/providers/bedrock.md b/docs/providers/bedrock.md index 3214446ee..8e9c85e52 100644 --- a/docs/providers/bedrock.md +++ b/docs/providers/bedrock.md @@ -1677,13 +1677,13 @@ AWS serves some Bedrock models on an OpenAI-compatible endpoint, `https://bedroc | Model | LiteLLM model name | Default route | |-------|--------------------|---------------| -| GPT-OSS 20B | `bedrock/openai.gpt-oss-20b-1:0` | Native Chat Completions | -| GPT-OSS 120B | `bedrock/openai.gpt-oss-120b-1:0` | Native Chat Completions | +| GPT-OSS 20B | `bedrock/openai.gpt-oss-20b-1:0`, `bedrock/us-gov.openai.gpt-oss-20b-1:0` | Native Chat Completions | +| GPT-OSS 120B | `bedrock/openai.gpt-oss-120b-1:0`, `bedrock/us-gov.openai.gpt-oss-120b-1:0` | Native Chat Completions | | GPT-5.6 Sol, Terra, Luna | `bedrock/us.openai.gpt-5.6-sol`, `bedrock/global.openai.gpt-5.6-sol`, and the `terra` / `luna` variants | Native Chat Completions | | Grok 4.6 | `bedrock/us.xai.grok-4.6`, `bedrock/global.xai.grok-4.6`, `bedrock/us-gov.xai.grok-4.6` | Native Chat Completions | | Everything else (Claude, Nova, Llama, Mistral, ...) | `bedrock/` | Converse or Invoke, as before | -A model opts in through `"supports_bedrock_runtime_chat_completions": true` on its entry in the [model cost map](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json), so a model AWS lists without Chat Completions support keeps using Converse. Authentication, regions, `aws_bedrock_runtime_endpoint`, and cost tracking work the same on both routes. [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc) is a separate route for imported models and is unchanged. +A model opts in through `"supports_bedrock_runtime_chat_completions": true` on its entry in the [model cost map](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json), so a model AWS lists without Chat Completions support keeps using Converse. Authentication, regions, `aws_bedrock_runtime_endpoint`, and cost tracking work the same on both routes. A region path in the model name (`bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0`) works the same too: the region picks the endpoint and the id after it is what AWS receives, and an explicit `aws_region_name` still wins over the path. [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc) is a separate route for imported models and is unchanged. The trade-off is that the OpenAI-compatible endpoint has no equivalent for a few Converse features, so LiteLLM falls back to Converse per request when you use one of them: From 82c2ea7a105087ceb3f56be7ea93e635fc260964 Mon Sep 17 00:00:00 2001 From: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Date: Sat, 19 Sep 2026 20:36:38 -0700 Subject: [PATCH 05/10] docs(bedrock): name the tool call id shape per model on the native route --- docs/providers/bedrock.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/providers/bedrock.md b/docs/providers/bedrock.md index 8e9c85e52..360118def 100644 --- a/docs/providers/bedrock.md +++ b/docs/providers/bedrock.md @@ -1697,7 +1697,7 @@ The trade-off is that the OpenAI-compatible endpoint has no equivalent for a few | `response_format` other than `{"type": "text"}` (a JSON schema, `json_object`, or a Pydantic model), on a model without `"supports_bedrock_runtime_chat_completions_response_format": true` in the cost map (GPT-OSS today) | Converse | AWS accepts `response_format` for GPT-OSS on Chat Completions but answers with free text anyway, so LiteLLM keeps the Converse emulation (a forced `json_tool_call` tool) that returns the JSON you asked for; GPT-5.6 and Grok carry the flag and enforce the schema natively | | `bedrock/converse/` | Converse | You asked for it explicitly | -What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids look like `call_0` instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, `response_format` and `service_tier` are sent through as you wrote them, `n` greater than 1 is unsupported (as on Converse), and GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field. +What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids are AWS's own (`call_0` for GPT-5.6 and Grok, `chatcmpl-tool-...` for GPT-OSS) instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, `response_format` and `service_tier` are sent through as you wrote them, `n` greater than 1 is unsupported (as on Converse), and GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field. To keep a model on Converse for every request, set the route explicitly: From 21809c9f82650ea9cf5440d901edb797345a5674 Mon Sep 17 00:00:00 2001 From: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Date: Mon, 21 Sep 2026 14:27:56 -0700 Subject: [PATCH 06/10] docs(bedrock): keep every json_object response_format on Converse for the native route --- docs/providers/bedrock.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/docs/providers/bedrock.md b/docs/providers/bedrock.md index 360118def..8db90ec84 100644 --- a/docs/providers/bedrock.md +++ b/docs/providers/bedrock.md @@ -1694,10 +1694,11 @@ The trade-off is that the OpenAI-compatible endpoint has no equivalent for a few | `bedrock_request_metadata_fields` set in `litellm_settings` | Converse, for every request | LiteLLM only writes the operator's request metadata onto the Converse body | | Application inference profile ARN as the model | Converse | LiteLLM cannot tell from the ARN which model it fronts | | Function `tools` with `reasoning_effort` other than `"none"` (or unset), on a model without `"supports_bedrock_runtime_chat_completions_tools_with_reasoning": true` in the cost map (the GPT-5.6 family today) | Converse | AWS only accepts function tools on Chat Completions for GPT-5.6 when `reasoning_effort` is `"none"`; GPT-OSS and Grok carry the flag and take tools with any effort | -| `response_format` other than `{"type": "text"}` (a JSON schema, `json_object`, or a Pydantic model), on a model without `"supports_bedrock_runtime_chat_completions_response_format": true` in the cost map (GPT-OSS today) | Converse | AWS accepts `response_format` for GPT-OSS on Chat Completions but answers with free text anyway, so LiteLLM keeps the Converse emulation (a forced `json_tool_call` tool) that returns the JSON you asked for; GPT-5.6 and Grok carry the flag and enforce the schema natively | +| `response_format` with `"type": "json_object"`, with or without LiteLLM's `response_schema` key, on any model | Converse | AWS's OpenAI-compatible endpoint answers 400 for `json_object` unless a message contains the word "json", so LiteLLM keeps Converse's handling: a `response_schema` becomes a forced `json_tool_call` tool that returns the JSON you asked for, and a schema-less `json_object` behaves as it did on Converse before | +| A JSON schema `response_format` (`{"type": "json_schema", ...}` or a Pydantic model), on a model without `"supports_bedrock_runtime_chat_completions_response_format": true` in the cost map (GPT-OSS today) | Converse | AWS accepts `response_format` for GPT-OSS on Chat Completions but answers with free text anyway, so LiteLLM keeps the Converse emulation (a forced `json_tool_call` tool) that returns the JSON you asked for; GPT-5.6 and Grok carry the flag and enforce the schema natively | | `bedrock/converse/` | Converse | You asked for it explicitly | -What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids are AWS's own (`call_0` for GPT-5.6 and Grok, `chatcmpl-tool-...` for GPT-OSS) instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, `response_format` and `service_tier` are sent through as you wrote them, `n` greater than 1 is unsupported (as on Converse), and GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field. +What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids are AWS's own (`call_0` for GPT-5.6 and Grok, `chatcmpl-tool-...` for GPT-OSS) instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, a JSON schema `response_format` and `service_tier` are sent through as you wrote them, `n` greater than 1 is unsupported (as on Converse), and GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field. To keep a model on Converse for every request, set the route explicitly: From 57938fbe0ade67c057f0f32f068daea027a6fa34 Mon Sep 17 00:00:00 2001 From: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Date: Thu, 24 Sep 2026 15:13:42 -0700 Subject: [PATCH 07/10] docs(bedrock): the native chat completions route opts in through supported_endpoints --- docs/providers/bedrock.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/providers/bedrock.md b/docs/providers/bedrock.md index 8db90ec84..6bd0946b1 100644 --- a/docs/providers/bedrock.md +++ b/docs/providers/bedrock.md @@ -1683,7 +1683,7 @@ AWS serves some Bedrock models on an OpenAI-compatible endpoint, `https://bedroc | Grok 4.6 | `bedrock/us.xai.grok-4.6`, `bedrock/global.xai.grok-4.6`, `bedrock/us-gov.xai.grok-4.6` | Native Chat Completions | | Everything else (Claude, Nova, Llama, Mistral, ...) | `bedrock/` | Converse or Invoke, as before | -A model opts in through `"supports_bedrock_runtime_chat_completions": true` on its entry in the [model cost map](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json), so a model AWS lists without Chat Completions support keeps using Converse. Authentication, regions, `aws_bedrock_runtime_endpoint`, and cost tracking work the same on both routes. A region path in the model name (`bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0`) works the same too: the region picks the endpoint and the id after it is what AWS receives, and an explicit `aws_region_name` still wins over the path. [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc) is a separate route for imported models and is unchanged. +A model opts in through `/v1/chat/completions` in the `supported_endpoints` of its entry in the [model cost map](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json), the same key that opts a model into the native `/v1/responses` route, so a model AWS lists without Chat Completions support keeps using Converse. Authentication, regions, `aws_bedrock_runtime_endpoint`, and cost tracking work the same on both routes. A region path in the model name (`bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0`) works the same too: the region picks the endpoint and the id after it is what AWS receives, and an explicit `aws_region_name` still wins over the path. [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc) is a separate route for imported models and is unchanged. The trade-off is that the OpenAI-compatible endpoint has no equivalent for a few Converse features, so LiteLLM falls back to Converse per request when you use one of them: From d79ac8735a4f93e00b3749f22d7e4f027e21d33b Mon Sep 17 00:00:00 2001 From: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Date: Sat, 26 Sep 2026 14:38:47 -0700 Subject: [PATCH 08/10] docs(bedrock): list the stop, top_k, and thinking fallbacks and the image and Grok effort notes for the native chat route --- docs/providers/bedrock.md | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/docs/providers/bedrock.md b/docs/providers/bedrock.md index 6bd0946b1..d35eefaa6 100644 --- a/docs/providers/bedrock.md +++ b/docs/providers/bedrock.md @@ -1696,9 +1696,12 @@ The trade-off is that the OpenAI-compatible endpoint has no equivalent for a few | Function `tools` with `reasoning_effort` other than `"none"` (or unset), on a model without `"supports_bedrock_runtime_chat_completions_tools_with_reasoning": true` in the cost map (the GPT-5.6 family today) | Converse | AWS only accepts function tools on Chat Completions for GPT-5.6 when `reasoning_effort` is `"none"`; GPT-OSS and Grok carry the flag and take tools with any effort | | `response_format` with `"type": "json_object"`, with or without LiteLLM's `response_schema` key, on any model | Converse | AWS's OpenAI-compatible endpoint answers 400 for `json_object` unless a message contains the word "json", so LiteLLM keeps Converse's handling: a `response_schema` becomes a forced `json_tool_call` tool that returns the JSON you asked for, and a schema-less `json_object` behaves as it did on Converse before | | A JSON schema `response_format` (`{"type": "json_schema", ...}` or a Pydantic model), on a model without `"supports_bedrock_runtime_chat_completions_response_format": true` in the cost map (GPT-OSS today) | Converse | AWS accepts `response_format` for GPT-OSS on Chat Completions but answers with free text anyway, so LiteLLM keeps the Converse emulation (a forced `json_tool_call` tool) that returns the JSON you asked for; GPT-5.6 and Grok carry the flag and enforce the schema natively | +| `stop` sequences | Converse | Converse forwards `stop` as `stopSequences`, which AWS answers with a 400 for these models, the same as before this route existed; sent natively, GPT-OSS and Grok apply `stop` to their hidden reasoning too and answer with empty content, which is worse than the error | +| `top_k` or `additionalModelRequestFields` | Converse | Only Converse forwards these model-specific fields | +| A `thinking` block on `/chat/completions` | Converse | The OpenAI-compatible endpoint has no `thinking` field; on `/v1/messages` LiteLLM maps `thinking` to `reasoning_effort` and the request stays native | | `bedrock/converse/` | Converse | You asked for it explicitly | -What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids are AWS's own (`call_0` for GPT-5.6 and Grok, `chatcmpl-tool-...` for GPT-OSS) instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, a JSON schema `response_format` and `service_tier` are sent through as you wrote them, `n` greater than 1 is unsupported (as on Converse), and GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field. +What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids are AWS's own (`call_0` for GPT-5.6 and Grok, `chatcmpl-tool-...` for GPT-OSS) instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, a JSON schema `response_format` and `service_tier` are sent through as you wrote them, `n` greater than 1 is unsupported (as on Converse), GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field, an `http(s)://` image URL in a message is downloaded by LiteLLM and sent inline as a `data:` URL because the endpoint does not fetch remote images itself, and Grok drops `reasoning_effort: "none"` (it always reasons) while `low`, `medium`, `high`, and `xhigh` are sent through. To keep a model on Converse for every request, set the route explicitly: From 09ca0cbfb50227b6788e8fd5e977273a5c827072 Mon Sep 17 00:00:00 2001 From: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Date: Wed, 30 Sep 2026 15:56:36 -0700 Subject: [PATCH 09/10] docs(bedrock): gpt-5.6 and newer default to the native chat completions route, chat_completions/ opts the rest in --- docs/providers/bedrock.md | 67 +++++++++++++++++++++++++++------------ 1 file changed, 46 insertions(+), 21 deletions(-) diff --git a/docs/providers/bedrock.md b/docs/providers/bedrock.md index 7d6640eda..1fd09ed04 100644 --- a/docs/providers/bedrock.md +++ b/docs/providers/bedrock.md @@ -7,7 +7,7 @@ ALL Bedrock models (Anthropic, Meta, Deepseek, Mistral, Amazon, etc.) are Suppor | Property | Details | |-------|-------| | Description | Amazon Bedrock is a fully managed service that offers a choice of high-performing foundation models (FMs). | -| Provider Route on LiteLLM | `bedrock/` ([native Chat Completions](#native-chat-completions-route) for the OpenAI and Grok models), [`bedrock/converse/`](#set-converse--invoke-route), [`bedrock/invoke/`](/docs/providers/bedrock#set-converse--invoke-route), [`bedrock/converse_like/`](/docs/providers/bedrock#calling-via-internal-proxy-not-bedrock-url-compatible), `bedrock/llama/`, `bedrock/deepseek_r1/`, `bedrock/qwen3/`, [`bedrock/qwen2/`](./bedrock_imported.md#qwen2-imported-models), [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc), [`bedrock/moonshot`](./bedrock_imported.md#moonshot-kimi-k2-thinking) | +| Provider Route on LiteLLM | `bedrock/` ([native Chat Completions](#native-chat-completions-route) for GPT-5.6 and newer), [`bedrock/chat_completions/`](#native-chat-completions-route), [`bedrock/converse/`](#set-converse--invoke-route), [`bedrock/invoke/`](/docs/providers/bedrock#set-converse--invoke-route), [`bedrock/converse_like/`](/docs/providers/bedrock#calling-via-internal-proxy-not-bedrock-url-compatible), `bedrock/llama/`, `bedrock/deepseek_r1/`, `bedrock/qwen3/`, [`bedrock/qwen2/`](./bedrock_imported.md#qwen2-imported-models), [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc), [`bedrock/moonshot`](./bedrock_imported.md#moonshot-kimi-k2-thinking) | | Provider Doc | [Amazon Bedrock ↗](https://docs.aws.amazon.com/bedrock/latest/userguide/what-is-bedrock.html) | | Supported OpenAI Endpoints | `/chat/completions`, `/completions`, `/embeddings`, `/images/generations`, `/v1/realtime`| | Rerank Endpoint | `/rerank` | @@ -1644,7 +1644,7 @@ LiteLLM defaults to the `invoke` route. LiteLLM uses the `converse` route for Be To explicitly set the route, do `bedrock/converse/` or `bedrock/invoke/`. -The models AWS serves in the OpenAI format are the exception. See [Native Chat Completions route](#native-chat-completions-route). +GPT-5.6 and newer are the exception: they default to AWS's OpenAI-compatible endpoint, and `bedrock/chat_completions/` opts any other model AWS serves there in. See [Native Chat Completions route](#native-chat-completions-route). E.g. @@ -1673,17 +1673,18 @@ model_list: ## Native Chat Completions route -AWS serves some Bedrock models on an OpenAI-compatible endpoint, `https://bedrock-runtime.{region}.amazonaws.com/openai/v1/chat/completions`. For those models LiteLLM sends your `/chat/completions` request to that endpoint in the shape it arrived in, instead of translating it to Converse and back. Fewer translations means less latency and fewer places for a parameter to get lost. +AWS serves some Bedrock models on an OpenAI-compatible endpoint, `https://bedrock-runtime.{region}.amazonaws.com/openai/v1/chat/completions`. For those models LiteLLM can send your `/chat/completions` request to that endpoint in the shape it arrived in, instead of translating it to Converse and back. Fewer translations means less latency and fewer places for a parameter to get lost. | Model | LiteLLM model name | Default route | |-------|--------------------|---------------| -| GPT-OSS 20B | `bedrock/openai.gpt-oss-20b-1:0`, `bedrock/us-gov.openai.gpt-oss-20b-1:0` | Native Chat Completions | -| GPT-OSS 120B | `bedrock/openai.gpt-oss-120b-1:0`, `bedrock/us-gov.openai.gpt-oss-120b-1:0` | Native Chat Completions | | GPT-5.6 Sol, Terra, Luna | `bedrock/us.openai.gpt-5.6-sol`, `bedrock/global.openai.gpt-5.6-sol`, and the `terra` / `luna` variants | Native Chat Completions | -| Grok 4.6 | `bedrock/us.xai.grok-4.6`, `bedrock/global.xai.grok-4.6`, `bedrock/us-gov.xai.grok-4.6` | Native Chat Completions | +| GPT-6 Sol, Astra, Luna | `bedrock/us.openai.gpt-6-sol`, `bedrock/global.openai.gpt-6-sol`, and the `astra` / `luna` variants | Native Chat Completions | +| GPT-6.1 Sol | `bedrock/us.openai.gpt-6.1-sol`, `bedrock/global.openai.gpt-6.1-sol` | Native Chat Completions | +| GPT-OSS 20B, 120B | `bedrock/openai.gpt-oss-20b-1:0`, `bedrock/openai.gpt-oss-120b-1:0`, and the `us-gov.` profile ids | Converse; `bedrock/chat_completions/openai.gpt-oss-20b-1:0` for native Chat Completions | +| Grok 4.6 | `bedrock/us.xai.grok-4.6`, `bedrock/global.xai.grok-4.6`, `bedrock/us-gov.xai.grok-4.6` | Converse; `bedrock/chat_completions/us.xai.grok-4.6` for native Chat Completions | | Everything else (Claude, Nova, Llama, Mistral, ...) | `bedrock/` | Converse or Invoke, as before | -A model opts in through `/v1/chat/completions` in the `supported_endpoints` of its entry in the [model cost map](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json), the same key that opts a model into the native `/v1/responses` route, so a model AWS lists without Chat Completions support keeps using Converse. Authentication, regions, `aws_bedrock_runtime_endpoint`, and cost tracking work the same on both routes. A region path in the model name (`bedrock/us-gov-west-1/openai.gpt-oss-20b-1:0`) works the same too: the region picks the endpoint and the id after it is what AWS receives, and an explicit `aws_region_name` still wins over the path. [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc) is a separate route for imported models and is unchanged. +GPT-5.6 and newer (`openai.gpt-5.6-*`, `openai.gpt-6-*`, `openai.gpt-6.1-*`, and later versions) take the native route by default when their entry in the [model cost map](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json) lists `/v1/chat/completions` under `supported_endpoints`, the same key that opts a model into the native `/v1/responses` route. Older GPT ids, GPT-OSS, and Grok keep the Converse route they had before unless you prefix the model with `bedrock/chat_completions/`, and `bedrock/converse/` pins any model to Converse. Authentication, regions, `aws_bedrock_runtime_endpoint`, and cost tracking work the same on both routes. A region path in the model name (`bedrock/chat_completions/us-gov-west-1/openai.gpt-oss-20b-1:0`) works the same too: the region picks the endpoint and the id after it is what AWS receives, and an explicit `aws_region_name` still wins over the path. [`bedrock/openai/`](./bedrock_imported.md#openai-compatible-imported-models-qwen-25-vl-etc) is a separate route for imported models and is unchanged. The trade-off is that the OpenAI-compatible endpoint has no equivalent for a few Converse features, so LiteLLM falls back to Converse per request when you use one of them: @@ -1693,17 +1694,17 @@ The trade-off is that the OpenAI-compatible endpoint has no equivalent for a few | `requestMetadata`, `performanceConfig`, `serviceTier`, or `outputConfig` in the request body | Converse | These Converse body fields are rejected as malformed input on the OpenAI-compatible endpoint | | `bedrock_request_metadata_fields` set in `litellm_settings` | Converse, for every request | LiteLLM only writes the operator's request metadata onto the Converse body | | Application inference profile ARN as the model | Converse | LiteLLM cannot tell from the ARN which model it fronts | -| Function `tools` with `reasoning_effort` other than `"none"` (or unset), on a model without `"supports_bedrock_runtime_chat_completions_tools_with_reasoning": true` in the cost map (the GPT-5.6 family today) | Converse | AWS only accepts function tools on Chat Completions for GPT-5.6 when `reasoning_effort` is `"none"`; GPT-OSS and Grok carry the flag and take tools with any effort | +| Function `tools` with `reasoning_effort` other than `"none"` (or unset), on a model without `"supports_bedrock_runtime_chat_completions_tools_with_reasoning": true` in the cost map (the GPT-5.6, GPT-6, and GPT-6.1 families today) | Converse | AWS only accepts function tools on Chat Completions for these models when `reasoning_effort` is `"none"`, and GPT-6.1 does not accept `"none"` at all, so its tool calls always go through Converse; GPT-OSS and Grok carry the flag and take tools with any effort | | `response_format` with `"type": "json_object"`, with or without LiteLLM's `response_schema` key, on any model | Converse | AWS's OpenAI-compatible endpoint answers 400 for `json_object` unless a message contains the word "json", so LiteLLM keeps Converse's handling: a `response_schema` becomes a forced `json_tool_call` tool that returns the JSON you asked for, and a schema-less `json_object` behaves as it did on Converse before | -| A JSON schema `response_format` (`{"type": "json_schema", ...}` or a Pydantic model), on a model without `"supports_bedrock_runtime_chat_completions_response_format": true` in the cost map (GPT-OSS today) | Converse | AWS accepts `response_format` for GPT-OSS on Chat Completions but answers with free text anyway, so LiteLLM keeps the Converse emulation (a forced `json_tool_call` tool) that returns the JSON you asked for; GPT-5.6 and Grok carry the flag and enforce the schema natively | +| A JSON schema `response_format` (`{"type": "json_schema", ...}` or a Pydantic model), on a model without `"supports_bedrock_runtime_chat_completions_response_format": true` in the cost map (GPT-OSS today) | Converse | AWS accepts `response_format` for GPT-OSS on Chat Completions but answers with free text anyway, so LiteLLM keeps the Converse emulation (a forced `json_tool_call` tool) that returns the JSON you asked for; GPT-5.6 and newer and Grok carry the flag and enforce the schema natively | | `stop` sequences | Converse | Converse forwards `stop` as `stopSequences`, which AWS answers with a 400 for these models, the same as before this route existed; sent natively, GPT-OSS and Grok apply `stop` to their hidden reasoning too and answer with empty content, which is worse than the error | | `top_k` or `additionalModelRequestFields` | Converse | Only Converse forwards these model-specific fields | | A `thinking` block on `/chat/completions` | Converse | The OpenAI-compatible endpoint has no `thinking` field; on `/v1/messages` LiteLLM maps `thinking` to `reasoning_effort` and the request stays native | | `bedrock/converse/` | Converse | You asked for it explicitly | -What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids are AWS's own (`call_0` for GPT-5.6 and Grok, `chatcmpl-tool-...` for GPT-OSS) instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, a JSON schema `response_format` and `service_tier` are sent through as you wrote them, `n` greater than 1 is unsupported (as on Converse), GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field, an `http(s)://` image URL in a message is downloaded by LiteLLM and sent inline as a `data:` URL because the endpoint does not fetch remote images itself, and Grok drops `reasoning_effort: "none"` (it always reasons) while `low`, `medium`, `high`, and `xhigh` are sent through. +What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids are AWS's own (`call_0` for the GPT-5.6 and newer families and Grok, `chatcmpl-tool-...` for GPT-OSS) instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, a JSON schema `response_format` and `service_tier` are sent through as you wrote them, `n` greater than 1 is unsupported (as on Converse), GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field, an `http(s)://` image URL in a message is downloaded by LiteLLM and sent inline as a `data:` URL because the endpoint does not fetch remote images itself, and Grok drops `reasoning_effort: "none"` (it always reasons) while `low`, `medium`, `high`, and `xhigh` are sent through. -To keep a model on Converse for every request, set the route explicitly: +To send GPT-OSS or Grok to the native endpoint, prefix the model with `bedrock/chat_completions/`: @@ -1711,7 +1712,7 @@ To keep a model on Converse for every request, set the route explicitly: ```python from litellm import completion -completion(model="bedrock/converse/openai.gpt-oss-20b-1:0", messages=[{"role": "user", "content": "Hello"}]) +completion(model="bedrock/chat_completions/openai.gpt-oss-20b-1:0", messages=[{"role": "user", "content": "Hello"}]) ``` @@ -1719,9 +1720,33 @@ completion(model="bedrock/converse/openai.gpt-oss-20b-1:0", messages=[{"role": " ```yaml model_list: - - model_name: gpt-oss-20b-converse + - model_name: gpt-oss-20b-native litellm_params: - model: bedrock/converse/openai.gpt-oss-20b-1:0 + model: bedrock/chat_completions/openai.gpt-oss-20b-1:0 +``` + + + + +To keep a GPT-5.6 or newer model on Converse for every request, set the route explicitly: + + + + +```python +from litellm import completion + +completion(model="bedrock/converse/us.openai.gpt-5.6-sol", messages=[{"role": "user", "content": "Hello"}]) +``` + + + + +```yaml +model_list: + - model_name: gpt-5.6-sol-converse + litellm_params: + model: bedrock/converse/us.openai.gpt-5.6-sol ``` @@ -1958,7 +1983,7 @@ curl -X POST 'http://0.0.0.0:4000/chat/completions' \ | Property | Details | |----------|---------| -| Provider Route | `bedrock/openai.gpt-oss-20b-1:0`, `bedrock/openai.gpt-oss-120b-1:0` ([native Chat Completions](#native-chat-completions-route)); prefix with `bedrock/converse/` to force Converse | +| Provider Route | `bedrock/converse/openai.gpt-oss-20b-1:0`, `bedrock/converse/openai.gpt-oss-120b-1:0`; `bedrock/chat_completions/openai.gpt-oss-20b-1:0` for the [native Chat Completions route](#native-chat-completions-route) | | Provider Documentation | [Amazon Bedrock ↗](https://docs.aws.amazon.com/bedrock/latest/userguide/what-is-bedrock.html) | @@ -1975,14 +2000,14 @@ os.environ["AWS_REGION_NAME"] = "us-east-1" # GPT OSS 20B model response = completion( - model="bedrock/openai.gpt-oss-20b-1:0", + model="bedrock/converse/openai.gpt-oss-20b-1:0", messages=[{"role": "user", "content": "Hello, how are you?"}], ) print(response.choices[0].message.content) # GPT OSS 120B model response = completion( - model="bedrock/openai.gpt-oss-120b-1:0", + model="bedrock/converse/openai.gpt-oss-120b-1:0", messages=[{"role": "user", "content": "Explain machine learning in simple terms"}], ) print(response.choices[0].message.content) @@ -1998,14 +2023,14 @@ print(response.choices[0].message.content) model_list: - model_name: gpt-oss-20b litellm_params: - model: bedrock/openai.gpt-oss-20b-1:0 + model: bedrock/converse/openai.gpt-oss-20b-1:0 aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY aws_region_name: os.environ/AWS_REGION_NAME - model_name: gpt-oss-120b litellm_params: - model: bedrock/openai.gpt-oss-120b-1:0 + model: bedrock/converse/openai.gpt-oss-120b-1:0 aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY aws_region_name: os.environ/AWS_REGION_NAME @@ -2263,8 +2288,8 @@ Here's an example of using a bedrock model with LiteLLM. For a complete list, re | Model Name | Command | |----------------------------|------------------------------------------------------------------| -| GPT-OSS 20B | `completion(model='bedrock/openai.gpt-oss-20b-1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']`, `os.environ['AWS_REGION_NAME']` | -| GPT-OSS 120B | `completion(model='bedrock/openai.gpt-oss-120b-1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']`, `os.environ['AWS_REGION_NAME']` | +| GPT-OSS 20B | `completion(model='bedrock/converse/openai.gpt-oss-20b-1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']`, `os.environ['AWS_REGION_NAME']` | +| GPT-OSS 120B | `completion(model='bedrock/converse/openai.gpt-oss-120b-1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']`, `os.environ['AWS_REGION_NAME']` | | Deepseek R1 | `completion(model='bedrock/us.deepseek.r1-v1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']` | | Anthropic Claude Sonnet 4.5 | `completion(model='bedrock/us.anthropic.claude-sonnet-4-5-20250929-v1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']` | | Anthropic Claude-V3.5 Sonnet | `completion(model='bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0', messages=messages)` | `os.environ['AWS_ACCESS_KEY_ID']`, `os.environ['AWS_SECRET_ACCESS_KEY']` | From 2b6b794291d05d5608d4f80aa4f2a2aa46641d95 Mon Sep 17 00:00:00 2001 From: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Date: Wed, 30 Sep 2026 16:53:34 -0700 Subject: [PATCH 10/10] docs(bedrock): sampling and logprob params ride on reasoning_effort none for GPT-5.6 and newer --- docs/providers/bedrock.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/providers/bedrock.md b/docs/providers/bedrock.md index 1fd09ed04..7845a463a 100644 --- a/docs/providers/bedrock.md +++ b/docs/providers/bedrock.md @@ -1704,6 +1704,8 @@ The trade-off is that the OpenAI-compatible endpoint has no equivalent for a few What you will notice on the native route: the response carries AWS's own `id` and `service_tier` fields, tool call ids are AWS's own (`call_0` for the GPT-5.6 and newer families and Grok, `chatcmpl-tool-...` for GPT-OSS) instead of `tooluse_...`, `max_tokens` is sent as `max_completion_tokens`, a JSON schema `response_format` and `service_tier` are sent through as you wrote them, `n` greater than 1 is unsupported (as on Converse), GPT-OSS reasoning comes back in `reasoning_content` (LiteLLM splits it out of the inline `...` prefix AWS returns) without the Converse-only `thinking_blocks` field, an `http(s)://` image URL in a message is downloaded by LiteLLM and sent inline as a `data:` URL because the endpoint does not fetch remote images itself, and Grok drops `reasoning_effort: "none"` (it always reasons) while `low`, `medium`, `high`, and `xhigh` are sent through. +On the GPT-5.6 and newer families, AWS ties `temperature`, `top_p`, `frequency_penalty`, `presence_penalty`, `logprobs`, and `top_logprobs` to reasoning being off: with `reasoning_effort: "none"` LiteLLM sends them through as you wrote them, and with any other effort, or none set, AWS would answer 400, so LiteLLM answers 400 up front naming the parameters, or drops them when `drop_params` is on. GPT-6.1 does not take `"none"` at all, so it never samples on Bedrock. Converse rejects `temperature` and `top_p` for these models at every effort, so `reasoning_effort: "none"` on the native route is the one way to sample them. GPT-OSS always refuses `logit_bias` and Grok always refuses the penalties, at any effort, with the same 400-or-drop handling. + To send GPT-OSS or Grok to the native endpoint, prefix the model with `bedrock/chat_completions/`: