Repository navigation
chore(release): backport #41870 to stable/1.101.x - #42332
Conversation
…_max_tokens fix(bedrock): clamp maxTokens to the 16-token minimum for OpenAI GPT and xAI Grok models on Converse (cherry picked from commit a6e3a72)
|
| @staticmethod | ||
| def _requires_min_max_tokens(model: str) -> bool: | ||
| return re.search(r"openai\.gpt-\d|xai\.grok-", model) is not None |
There was a problem hiding this comment.
Hardcoded model capability detection
This identifies model families with hardcoded name patterns. Repository rules require model-specific flags to live in model_prices_and_context_window.json and be read through get_model_info, so this must be addressed before merging
Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)
Knowledge Base Used: Provider adapters and capabilities
| if param == "max_tokens" or param == "max_completion_tokens": | ||
| optional_params["maxTokens"] = value | ||
| optional_params["maxTokens"] = ( | ||
| max(value, BEDROCK_OPENAI_COMPAT_MIN_MAX_TOKENS) |
There was a problem hiding this comment.
Low: Token rate-limit under-reservation
The rate-limit pre-call hook reserves the caller's original max_tokens value, but this later raises it to 16. An authenticated caller can submit concurrent requests with max_tokens: 1 and receive up to 16 output tokens per one token reserved, exceeding TPM or OTPM limits before post-call reconciliation. Normalize the effective cap before rate-limit hooks run, or make the limiter apply the same model-family minimum.
PR overviewThis release backport updates Bedrock Converse request transformation behavior on the stable/1.101.x branch, including minimum output-token handling for affected models. One security issue remains open in the token rate-limiting path. An authenticated caller can use concurrent low-token requests to under-reserve capacity and temporarily exceed configured token-per-minute limits, though the impact is limited to rate-limit enforcement and later reconciliation. Open issues (1)
Fixed/addressed: 0 · PR risk: 5/10 |
Relevant issues
Backports #41870 onto
stable/1.101.x. On Bedrock Converse, OpenAI GPT and xAI Grok models reject amaxTokensbelow 16, so a request carrying a smallermax_tokensormax_completion_tokensfails upstream instead of being served. The pick clamps the mapped value to that 16-token floor for those two model families only; every other model, including Anthropic and application inference profiles, keeps the caller's value unchanged.This line's tip is already at
1.101.1from #41704, and that version has not been released anywhere (no published image, no tag, no release branch), so these changes ride the pending1.101.1and there is no version bump in this PR.The two dependency commits refresh third-party pins that the lock resolves below their current maintenance releases. Both are lock-only regenerations produced with the line's own uv pin (0.11.7); no pyproject range moves, and each moved-set is confined to exactly the one package.
What is included
Adaptation notes
The pick is ADAPTED, in two ways, both forced by this line predating upstream's
fbc6fb56ae2:The upstream diff carries
_is_openai_gpt_reasoning_modelas surrounding context, and because this line does not have that method the three-way merge presented it as part of the incoming block. It is not part of #41870 and nothing on this line calls it, so it was dropped; only_requires_min_max_tokens, the method #41870 actually adds, was taken. A name-equality check against the source commit confirms the resolved file adds exactly the definitions #41870 adds, and no others.import rewas added to the module-level import block. Upstream gained that import infbc6fb56ae2, which is not on this line; herereexisted only as a function-local import inside_validate_request_metadata. Without it the new helper raisesNameErroron every request that maps an integermax_tokens, which took all 10 of the pick's own tests and 14 pre-existing ones red. The added line is byte-identical to upstream's, and the pre-existing function-local import was left untouched.Both adaptations match what this pick already received on the sibling stable lines. #41870 was backported to
stable/1.100.xasfff43dfc05and tostable/1.98.xas79960073b1, and each of those commits addsimport re, the constant and_requires_min_max_tokens, and likewise does not add_is_openai_gpt_reasoning_model. This resolution is the same one, arrived at independently.Known noise on this line
Two tests fail on
stable/1.101.xbefore any of these commits are applied, and they still fail identically after; they are unrelated to this change:tests/test_litellm/litellm_core_utils/test_get_supported_openai_params.py::test_base_model_adds_capabilities_the_real_model_lackstests/test_litellm/test_utils.py::test_supports_tool_choice_simple_testsScreenshots / Proof of Fix
The pick's claim is a pure parameter-mapping behavior, so the proof is the 10 parametrized cases #41870 adds, all of which pass on this line and all of which were red before the
import readaptation:Both directions of the clamp are covered:
us.openai.gpt-6-astraandus.xai.grok-4.6withmax_tokens=1map to 16, the same models with 64 and 32 pass through unchanged,anthropic.claude-sonnet-4-5-20250929-v1:0withmax_tokens=1stays 1, inference-profile ARNs wrapping a GPT or Grok id clamp, and an application-inference-profile ARN does not.Resolved helper behavior on the built tree:
Regression net, run over
tests/test_litellm/llms/bedrock/plus every test module importingconverse_transformationorAmazonConverseConfig, differenced against the same command on the untouched line tip:Dependency commits, resolved in the built environment:
Each lock regeneration moved exactly one package and nothing else, and
ruff format --checkandruff checkare both clean on the changed file.Verification
An adversarial review was run over this diff with five independent lenses (correctness, dependents, backward-compatibility, conventions and tests, plus a blind black-box exerciser that was given only the claims and not the code). Every finding that would change the verdict was reproduced or rejected by executing it, not by argument. Nothing was confirmed as a defect this backport introduces, and no existing caller of
map_openai_paramsorAmazonConverseConfigbreaks.The new tests were checked for theater by mutation: five separate mutations of the picked code (helper forced true, helper forced false, floor moved to 15, floor moved to 17,
maxswapped formin) each take cases red, and every one of the ten cases is the unique or joint-unique kill for at least one mutation. The two negative controls are the only thing that catches a helper forced true, and the two above-floor cases are the only thing that catches the min/max swap.The subclass that inherits this method,
AmazonInvokeNovaConfig, was exercised as well; Nova model ids do not match the pattern, so they pass through unchanged.One scope limitation is worth stating plainly for reviewers, since it is inherited from upstream rather than introduced here. The pattern is
openai\.gpt-\d, which requires a digit aftergpt-, so theopenai.gpt-oss-*models do not receive the floor:Elsewhere this same file gates OpenAI-family behavior on
"gpt-oss" in model, so the two are inconsistent. The helper here is byte-identical to the one on the development branch, so this is a faithful carry and not a divergence in the backport; widening it belongs upstream first, then as its own backport.