Skip to content

chore(release): backport six staged fixes into stable/1.84.x and cut 1.84.5 - #29628

Merged
mateo-berri merged 8 commits into
stable/1.84.xfrom
litellm_cherrypick_1_84_5
Jun 4, 2026
Merged

chore(release): backport six staged fixes into stable/1.84.x and cut 1.84.5#29628
mateo-berri merged 8 commits into
stable/1.84.xfrom
litellm_cherrypick_1_84_5

Conversation

@mateo-berri

@mateo-berri mateo-berri commented Jun 3, 2026

Copy link
Copy Markdown
Contributor

Relevant issues

Backports six fixes that already merged on litellm_internal_staging but never reached the 1.84.x line. Each one is cherry-picked from the squashed commit that landed for its PR, so the set matches what the newer lines already carry. This closes that gap and cuts 1.84.5

Linear ticket

N/A

What is included

Cherry-picked in merge order:

The last two commits are the version bump (1.84.4 to 1.84.5) and the matching uv.lock refresh

#29598 needed a manual conflict resolution. pass_through_endpoints.py is stored with CRLF endings on the stable branches, so the verbatim cherry-pick will not apply, and the new tests sit next to an unrelated test class that this line does not carry. The resolved source change and the added tests are byte-identical to the upstream commit; only the surrounding file context differs, and the dedup guard the fix relies on (has_dispatched_final_stream_success) is present on this line

Pre-Submission checklist

  • The cherry-picked PRs each carry their own tests
  • My PR passes all unit tests on make test-unit
  • Scope is limited to backporting already-merged fixes plus the release bump
  • Greptile review requested

CI (LiteLLM team)

  • Branch creation CI run
    Link:
  • CI run for the last commit
    Link:
  • Merge / cherry-pick CI run
    Links:

Screenshots / Proof of Fix

These are backports of changes already merged, released, and validated on newer lines; each linked PR carries its own proof and review

Type

Bug Fix
Infrastructure

Changes

See the commit list above. No new code beyond the cherry-picks, the version bump, and the lockfile refresh

mateo-berri and others added 8 commits June 3, 2026 22:52
* fix(azure): preserve AD token refresh in v1 OpenAI client path

The /openai/v1/ code path (api_version in {"v1", "latest", "preview"})
constructs a plain OpenAI/AsyncOpenAI client, but only forwarded
`api_key` from `azure_client_params`. When `enable_azure_ad_token_refresh`
is set (or any AD-only auth), `api_key` is None and the client
constructor raised "The api_key client option must be set...", breaking
every Azure call with a v1 api_version.

The OpenAI SDK (>=2.20.0) accepts a callable for `api_key` and re-invokes
it on every request via `_refresh_api_key`, so we now forward
`azure_ad_token_provider` directly — preserving the per-request token
refresh behavior of the regular AzureOpenAI client and avoiding the
expiry hole that resolving the token once at client-creation time would
introduce. Static `azure_ad_token` strings fall through to `api_key`.

For the async path we wrap the sync provider returned by azure-identity
in an async function since AsyncOpenAI expects `Callable[[], Awaitable[str]]`.

Fixes #27945

https://claude.ai/code/session_01UnzrDSFUUgp5T2wRoPMxq5

* fix(azure): offload sync token provider to thread in v1 async wrapper

* fix(azure): include AD credential identity in v1 client cache key

---------

Co-authored-by: Claude <noreply@anthropic.com>
(cherry picked from commit 96a2e8b)
…9264)

* fix(proxy): map stripped batch body.model to proxy alias for auth

replace_model_in_jsonl rewrites JSONL body.model to the provider id before
upload; batch file access checks must resolve that id back to model_name
so keys granted the proxy alias are not rejected with 403.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(proxy): surface resolved proxy alias in batch file 403 detail

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
(cherry picked from commit 70d2748)
* fix(proxy): resolve managed video model ids for auth

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(proxy): cover character_id router model resolution

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
(cherry picked from commit d45e9e4)
…ams (#29310)

* fix(key_generate): allow team members to create keys on org-scoped teams

When a virtual key is created for a team, enterprise logic inherits the
team's organization_id onto the key (add_team_organization_id). Since the
VERIA-55 org-IDOR fix, /key/generate then required the caller to be an
explicit LiteLLM_OrganizationMembership member of that org, returning
403 "Caller is not a member of organization_id=<uuid>". Admins normally
only add users to teams (not orgs), so self-serve key creation regressed
for any user on an org-scoped team (regression since v1.84.0-rc.1).

Skip the org-membership check when organization_id was inherited from the
key's team (organization_id == team_table.organization_id). Team-level
authorization already gates this path, so team membership is sufficient.
The membership check still runs when a caller assigns an organization_id
that did not come from the key's team, preserving the IDOR protection.

Adds regression tests covering both the team-inherited (allowed) and
foreign-org (still blocked) cases.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(key_generate): cover mismatched team org IDOR path on generate

Add test_generate_key_foreign_org_with_mismatched_team_still_enforces_membership
for the case where a team is present but request organization_id differs from
team_table.organization_id. Enterprise inheritance is no-op'd in the test so
the guard is exercised directly; membership validation must still run.

Addresses Greptile review on #29310.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
(cherry picked from commit b11833c)
… reject it (Haiku 4.5) (#29585)

* fix(vertex): strip output_config.effort for models that reject it

Haiku 4.5 on Vertex AI does not support output_config.effort and 400s with
"output_config.effort: Extra inputs are not permitted". PR #27074 emptied
VERTEX_UNSUPPORTED_OUTPUT_CONFIG_KEYS so effort would forward for Opus/Sonnet
4.6+, but that made the strip unconditional across every Vertex Anthropic
model, including ones that don't support it. Claude Code injects effort into
its default Messages payload, so `claude --model claude-haiku-4.5` started
failing.

Make the sanitizer model-aware: drop output_config.effort for models that
don't advertise output_config support (or any reasoning effort level) while
forwarding it for those that do. The fix covers both the chat-completion and
Messages pass-through transformation paths since they share the helper.

* chore(vertex): log at debug when dropping unsupported output_config.effort

Operators pointing an unregistered Vertex Claude alias that does support
effort would otherwise see it stripped with no signal. Debug level keeps it
out of normal logs since Claude Code sends effort on every request.

(cherry picked from commit cc55662)
* fix duplicate cost callbacks for anthropic streaming pass-through

Two bugs caused _PROXY_track_cost_callback to see stream=True +
complete_streaming_response=None on every streaming pass-through request,
making the dedup guard in dispatch_success_handlers permanently inactive:

1. pass_through_endpoints.py created the Logging object with stream=False
   for all requests. _is_assembled_stream_success short-circuits on
   self.stream is not True, so has_dispatched_final_stream_success was
   never set and any second dispatch went through unchecked.
   Fix: set logging_obj.stream = True after stream detection.

2. _create_anthropic_response_logging_payload set complete_streaming_response
   inside the try block after litellm.completion_cost(), so a pricing error
   caused an early return without setting it on model_call_details.
   Fix: set complete_streaming_response before the try block.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix stream

* add stream to logging obj

* test(pass_through): give mock logging object a real model_call_details dict

The anthropic passthrough logging payload now records the assembled
response on model_call_details before cost calculation, which requires
model_call_details to support item assignment. In production it is always
a dict; the existing unit test stubbed the logging object with a bare Mock
whose attribute is not subscriptable, so the new assignment raised
TypeError. Use a real dict to match the production logging object.

* test(pass_through): cover streaming logging-obj stream flag

The streaming branch of pass_through_request that marks the logging object
as streaming (logging_obj.stream and model_call_details["stream"]) had no
unit coverage, so the patch coverage gate flagged it. Add a regression test
that drives a streaming pass-through request through pass_through_request and
asserts the logging object is flagged as a stream before dispatch.

* test(pass_through): cover SSE-response stream flag fallback branch

The auto-detected streaming branch of pass_through_request (when a request
that was not flagged as streaming returns a text/event-stream response) sets
logging_obj.stream and model_call_details["stream"] but had no unit coverage,
so the codecov patch gate failed at 60%. Drive a non-streaming pass-through
request whose upstream response is SSE through pass_through_request and assert
the logging object is flagged as a stream before dispatch.

* fix(pass_through): gate complete_streaming_response on stream flag

perform_redaction only scrubs complete_streaming_response when
model_call_details["stream"] is True. Setting it unconditionally for
non-streaming Anthropic pass-through responses left the assembled
response unredacted in model_call_details, which is handed to logging
callbacks as kwargs when message logging is disabled. Only record it for
actual streaming responses so redaction always applies.

---------

Co-authored-by: mubashir1osmani <mubashir.osmani777@gmail.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
(cherry picked from commit 2bbdbfa)
@mateo-berri
mateo-berri marked this pull request as ready for review June 3, 2026 23:20
@mateo-berri
mateo-berri requested review from a team and yassin-berriai June 3, 2026 23:20
@greptile-apps

greptile-apps Bot commented Jun 3, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR backports six targeted bug fixes from litellm_internal_staging to the stable/1.84.x line and cuts version 1.84.5. Each cherry-pick matches its upstream commit, conflict resolution on pass_through_endpoints.py (CRLF line endings) is explicitly documented, and every fix is accompanied by unit tests.

Confidence Score: 4/5

All six cherry-picks apply cleanly and are well-tested; the only rough edge is an unguarded router call in the batch rate-limiter that differs from the defensive pattern used elsewhere for the same operation.

The backported changes are narrow and well-scoped: AD token refresh, Vertex effort-param gating, batch model mapping, video model routing, team-org key creation, and streaming dedup. Tests cover positive, negative, and edge cases for each fix. The only rough edge is that batch_rate_limiter.py calls resolve_model_name_from_model_id without error handling, while the equivalent call in auth_utils.py wraps it in a try-except; an unexpected router error there would surface as an unhandled exception blocking the batch instead of falling back gracefully.

litellm/proxy/hooks/batch_rate_limiter.py — missing error guard around resolve_model_name_from_model_id call.

Important Files Changed

Filename Overview
litellm/llms/azure/common_utils.py Adds azure_ad_token_provider / azure_ad_token as fallback api_key for the v1 OpenAI client path, with async wrapping via asyncio.to_thread; also enriches the client-cache key with hashed credentials so distinct AD configs get separate clients.
litellm/llms/vertex_ai/vertex_ai_partner_models/anthropic/output_params_utils.py Makes sanitize_vertex_anthropic_output_params model-aware: strips output_config.effort for models that don't accept it (e.g. Haiku 4.5) while forwarding it for Opus/Sonnet 4.6+, using the model-cost table via AnthropicConfig._model_supports_effort_param.
litellm/proxy/auth/auth_utils.py Threads llm_router through the managed-resource-ID resolution chain so decoded video_id / character_id model IDs are resolved to their proxy model names before auth and budget checks.
litellm/proxy/auth/user_api_key_auth.py Passes llm_router to all _get_model_from_request_context call sites so managed video model IDs are resolved before every auth, budget, and access check in the critical request path.
litellm/proxy/hooks/batch_rate_limiter.py Resolves batch body.model provider IDs to proxy model names via llm_router.resolve_model_name_from_model_id before the can_key_call_model check, but unlike the parallel fix in auth_utils.py, this call has no error guard.
litellm/proxy/management_endpoints/key_management_endpoints.py Adds _org_inherited_from_team bypass: when the key's organization_id matches the team's DB-sourced organization_id, the explicit org-membership check is skipped, fixing the v1.84.0-rc.1 regression without relaxing the check for unrelated orgs.
litellm/proxy/pass_through_endpoints/pass_through_endpoints.py Sets logging_obj.stream = True and mirrors it to model_call_details["stream"] on both the explicit-stream path and the SSE-response fallback path, enabling the dedup guard for streaming pass-through requests.
litellm/proxy/pass_through_endpoints/llm_provider_handlers/anthropic_passthrough_logging_handler.py Guards complete_streaming_response population on model_call_details["stream"] is True, preventing unredacted responses from leaking to logging callbacks on non-streaming pass-through calls.

Reviews (1): Last reviewed commit: "chore: update uv.lock for 1.84.5" | Re-trigger Greptile

@mateo-berri
mateo-berri enabled auto-merge June 3, 2026 23:25
Comment on lines +312 to +316
model_to_check = model
if llm_router is not None:
proxy_model_name = llm_router.resolve_model_name_from_model_id(model)
if proxy_model_name is not None:
model_to_check = proxy_model_name

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 The parallel fix in auth_utils.py wraps the same resolve_model_name_from_model_id call in a try-except and falls back to the raw model_id on failure. Here the call is unguarded, so an unexpected exception from the router lookup (e.g. an empty model_list race) would surface as an unhandled error and block the entire batch rather than gracefully falling through to the original model string.

Suggested change
model_to_check = model
if llm_router is not None:
proxy_model_name = llm_router.resolve_model_name_from_model_id(model)
if proxy_model_name is not None:
model_to_check = proxy_model_name
model_to_check = model
if llm_router is not None:
try:
proxy_model_name = llm_router.resolve_model_name_from_model_id(model)
if proxy_model_name is not None:
model_to_check = proxy_model_name
except Exception as e:
verbose_proxy_logger.debug(
"Unable to resolve model_id=%s in batch rate limiter: %s",
model,
str(e),
)

@shivamrawat1 shivamrawat1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maunally qa'd for startup.

@mateo-berri
mateo-berri merged commit db7f25d into stable/1.84.x Jun 4, 2026
61 of 74 checks passed
@mateo-berri
mateo-berri deleted the litellm_cherrypick_1_84_5 branch June 4, 2026 03:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants