Skip to content

chore(release): backport #41870 to stable/1.101.x - #42332

Merged
yuneng-berri merged 3 commits into
stable/1.101.xfrom
litellm_backport_1_101_x_bp_41870
Sep 21, 2026
Merged

yuneng-berri merged 3 commits into
stable/1.101.xfrom
litellm_backport_1_101_x_bp_41870

Conversation

@yuneng-berri

@yuneng-berri yuneng-berri commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Relevant issues

Backports #41870 onto stable/1.101.x. On Bedrock Converse, OpenAI GPT and xAI Grok models reject a maxTokens below 16, so a request carrying a smaller max_tokens or max_completion_tokens fails upstream instead of being served. The pick clamps the mapped value to that 16-token floor for those two model families only; every other model, including Anthropic and application inference profiles, keeps the caller's value unchanged.

This line's tip is already at 1.101.1 from #41704, and that version has not been released anywhere (no published image, no tag, no release branch), so these changes ride the pending 1.101.1 and there is no version bump in this PR.

The two dependency commits refresh third-party pins that the lock resolves below their current maintenance releases. Both are lock-only regenerations produced with the line's own uv pin (0.11.7); no pyproject range moves, and each moved-set is confined to exactly the one package.

What is included

Adaptation notes

The pick is ADAPTED, in two ways, both forced by this line predating upstream's fbc6fb56ae2:

The upstream diff carries _is_openai_gpt_reasoning_model as surrounding context, and because this line does not have that method the three-way merge presented it as part of the incoming block. It is not part of #41870 and nothing on this line calls it, so it was dropped; only _requires_min_max_tokens, the method #41870 actually adds, was taken. A name-equality check against the source commit confirms the resolved file adds exactly the definitions #41870 adds, and no others.

import re was added to the module-level import block. Upstream gained that import in fbc6fb56ae2, which is not on this line; here re existed only as a function-local import inside _validate_request_metadata. Without it the new helper raises NameError on every request that maps an integer max_tokens, which took all 10 of the pick's own tests and 14 pre-existing ones red. The added line is byte-identical to upstream's, and the pre-existing function-local import was left untouched.

Both adaptations match what this pick already received on the sibling stable lines. #41870 was backported to stable/1.100.x as fff43dfc05 and to stable/1.98.x as 79960073b1, and each of those commits adds import re, the constant and _requires_min_max_tokens, and likewise does not add _is_openai_gpt_reasoning_model. This resolution is the same one, arrived at independently.

Known noise on this line

Two tests fail on stable/1.101.x before any of these commits are applied, and they still fail identically after; they are unrelated to this change:

  • tests/test_litellm/litellm_core_utils/test_get_supported_openai_params.py::test_base_model_adds_capabilities_the_real_model_lacks
  • tests/test_litellm/test_utils.py::test_supports_tool_choice_simple_tests

Screenshots / Proof of Fix

The pick's claim is a pure parameter-mapping behavior, so the proof is the 10 parametrized cases #41870 adds, all of which pass on this line and all of which were red before the import re adaptation:

uv run pytest tests/test_litellm/llms/bedrock/chat/test_converse_transformation.py -q
245 passed, 3 warnings in 0.77s      (baseline on the line tip: 235 passed)

Both directions of the clamp are covered: us.openai.gpt-6-astra and us.xai.grok-4.6 with max_tokens=1 map to 16, the same models with 64 and 32 pass through unchanged, anthropic.claude-sonnet-4-5-20250929-v1:0 with max_tokens=1 stays 1, inference-profile ARNs wrapping a GPT or Grok id clamp, and an application-inference-profile ARN does not.

Resolved helper behavior on the built tree:

BEDROCK_OPENAI_COMPAT_MIN_MAX_TOKENS = 16
_requires_min_max_tokens('us.openai.gpt-6-astra')                    -> True
_requires_min_max_tokens('us.xai.grok-4.6')                          -> True
_requires_min_max_tokens('anthropic.claude-sonnet-4-5-20250929-v1:0') -> False

Regression net, run over tests/test_litellm/llms/bedrock/ plus every test module importing converse_transformation or AmazonConverseConfig, differenced against the same command on the untouched line tip:

baseline   2 failed, 2001 passed, 1 skipped
post-pick  2 failed, 2011 passed, 1 skipped
failure sets identical; 0 new failures, 10 newly passing

Dependency commits, resolved in the built environment:

anyio      4.13.0 -> 4.14.2
soupsieve  2.8.4  -> 2.9

Each lock regeneration moved exactly one package and nothing else, and ruff format --check and ruff check are both clean on the changed file.

Verification

An adversarial review was run over this diff with five independent lenses (correctness, dependents, backward-compatibility, conventions and tests, plus a blind black-box exerciser that was given only the claims and not the code). Every finding that would change the verdict was reproduced or rejected by executing it, not by argument. Nothing was confirmed as a defect this backport introduces, and no existing caller of map_openai_params or AmazonConverseConfig breaks.

The new tests were checked for theater by mutation: five separate mutations of the picked code (helper forced true, helper forced false, floor moved to 15, floor moved to 17, max swapped for min) each take cases red, and every one of the ten cases is the unique or joint-unique kill for at least one mutation. The two negative controls are the only thing that catches a helper forced true, and the two above-floor cases are the only thing that catches the min/max swap.

The subclass that inherits this method, AmazonInvokeNovaConfig, was exercised as well; Nova model ids do not match the pattern, so they pass through unchanged.

One scope limitation is worth stating plainly for reviewers, since it is inherited from upstream rather than introduced here. The pattern is openai\.gpt-\d, which requires a digit after gpt-, so the openai.gpt-oss-* models do not receive the floor:

openai.gpt-oss-120b-1:0        clamped=False
us.openai.gpt-oss-20b-1:0      clamped=False
openai.gpt-5.5                 clamped=True
us.openai.gpt-6-astra          clamped=True
xai.grok-4.3                   clamped=True

Elsewhere this same file gates OpenAI-family behavior on "gpt-oss" in model, so the two are inconsistent. The helper here is byte-identical to the one on the development branch, so this is a faithful carry and not a divergence in the backport; widening it belongs upstream first, then as its own backport.

mateo-berri and others added 3 commits September 21, 2026 14:18
…_max_tokens

fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse

(cherry picked from commit a6e3a72)
@yuneng-berri
yuneng-berri requested a review from a team September 21, 2026 21:26
@greptile-apps

greptile-apps Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

The behavior appears sound, but the explicit model-metadata requirement must be satisfied before merging

Findings

  1. P2 Hardcoded model capability detection ▶

Summary

This PR adjusts Bedrock Converse token mapping for selected OpenAI-compatible model families and refreshes two locked dependencies

  • Clamps maxTokens to 16 for matched OpenAI GPT and xAI Grok identifiers
  • Adds focused parameter-mapping coverage for model IDs and inference-profile ARNs
  • Updates AnyIO to 4.14.2 and Soup Sieve to 2.9
  • The model-family decision should be moved into the repository's model metadata mechanism

Reviews (1) · Last reviewed commit: "chore(deps): bump soupsieve to 2.9.0"

Comment on lines +298 to +300
@staticmethod
def _requires_min_max_tokens(model: str) -> bool:
return re.search(r"openai\.gpt-\d|xai\.grok-", model) is not None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Hardcoded model capability detection

This identifies model families with hardcoded name patterns. Repository rules require model-specific flags to live in model_prices_and_context_window.json and be read through get_model_info, so this must be addressed before merging

Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)

Knowledge Base Used: Provider adapters and capabilities

if param == "max_tokens" or param == "max_completion_tokens":
optional_params["maxTokens"] = value
optional_params["maxTokens"] = (
max(value, BEDROCK_OPENAI_COMPAT_MIN_MAX_TOKENS)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Low: Token rate-limit under-reservation

The rate-limit pre-call hook reserves the caller's original max_tokens value, but this later raises it to 16. An authenticated caller can submit concurrent requests with max_tokens: 1 and receive up to 16 output tokens per one token reserved, exceeding TPM or OTPM limits before post-call reconciliation. Normalize the effective cap before rate-limit hooks run, or make the limiter apply the same model-family minimum.

@veria-ai

veria-ai Bot commented Sep 21, 2026

Copy link
Copy Markdown
Contributor

PR overview

This release backport updates Bedrock Converse request transformation behavior on the stable/1.101.x branch, including minimum output-token handling for affected models.

One security issue remains open in the token rate-limiting path. An authenticated caller can use concurrent low-token requests to under-reserve capacity and temporarily exceed configured token-per-minute limits, though the impact is limited to rate-limit enforcement and later reconciliation.

Open issues (1)

Fixed/addressed: 0 · PR risk: 5/10

@yuneng-berri
yuneng-berri merged commit 1fb7e5a into stable/1.101.x Sep 21, 2026
5 checks passed
@yuneng-berri
yuneng-berri deleted the litellm_backport_1_101_x_bp_41870 branch September 21, 2026 21:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants