Skip to content

feat(mcp): mint gateway-bound envelope at the token endpoint for dcr_bridge oauth_delegate - #32828

Merged
tin-berri merged 12 commits into
litellm_internal_stagingfrom
litellm_lit4338_delegate_token_mint
Jul 13, 2026
Merged

feat(mcp): mint gateway-bound envelope at the token endpoint for dcr_bridge oauth_delegate#32828
tin-berri merged 12 commits into
litellm_internal_stagingfrom
litellm_lit4338_delegate_token_mint

Conversation

@tin-berri

@tin-berri tin-berri commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

Relevant issues

Linear ticket

Resolves LIT-4338 (final PR of the delegate-bridge stack; completes the mint -> admit loop)

Pre-Submission checklist

  • I have added meaningful tests
  • My PR passes all CI/CD checks (e.g., lint, format, unit tests)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

The producer -> consumer loop is proven as a unit contract: the token endpoint mints a gateway-bound envelope, and opening it back through the consumer recovers the sealed key reference and the upstream Authorization exactly, with the raw upstream token never appearing in the bearer the client receives. The full interactive proof (a real client signing in, receiving the envelope, and calling tools that attribute to the litellm identity) exercises this producer together with the admission consumer in the PR below it, and is the recommended combined-stack validation on a live proxy before merge; I did not want to claim run logs this PR does not yet have

pytest tests/test_litellm/proxy/_experimental/mcp_server/test_discoverable_endpoints.py -q
164 passed

Type

🆕 New Feature

Changes

Completes the DCR-bridge oauth_delegate flow: at the OAuth token endpoint, a bridge oauth_delegate server now returns a gateway-bound envelope instead of the raw upstream token, so the client holds one bearer that both admits it (via the consumer PR) and forwards the upstream token, with nothing stored server-side

The mint is modeled as three phases inside exchange_token_with_server, whose failures are values rather than raised exceptions. _prepare_bridge_mint runs before the upstream exchange, fires only when the server is both is_oauth_delegate and is_dcr_bridge, checks every precondition once (master_key set, then a resolvable active litellm key), and returns either a frozen _BridgeMintReady carrying the resolved key hash and the master-key-derived envelope keys, or a _BridgeMintError. Because every precondition lives in prepare and prepare runs before the exchange, a missing master_key or an unresolvable identity fails closed without ever consuming the single-use upstream code or rotating a refresh token, for both grant types, by construction rather than by a guard kept in sync by hand. _finish_bridge_mint runs after the exchange and has no preconditions left that can fail; it seals the upstream grant plus an EnvelopeIdentity bound to that key hash and this server_id via build_bridge_token_response and returns the envelope as the access_token with the envelope's own short expires_in. Its only failure values are properties of the upstream response itself, a missing usable access_token or a token too large to seal

Every resolution step the mint depends on returns a precise tagged value rather than a bare None, because a None that meant several different things is what let the wrong status be reported. Identity resolution returns a _ResolvedKey or one of no_active_key / unavailable / unresolvable, classified the way admission's _reload_admitted_key classifies the same conditions; upstream-lifetime classification returns a positive number of seconds, "unspecified" (absent or unparseable, which the envelope caps), or "expired" (a parseable non-positive value, i.e. an already-dead token); and upstream-grant validation returns a typed grant or one of no_access_token / expired_lifetime. Thin exhaustive mappers (match plus assert_never) lift each vocabulary into one _BridgeMintError taxonomy of eight named failures, and a single _bridge_mint_error_response gives each its truthful RFC 6749 section 5.2 status: 400 invalid_request for a missing litellm credential, 400 unsupported_grant_type for a refresh_token grant, 503 temporarily_unavailable for a transient auth-DB outage, 500 server_error for a gateway that cannot resolve identity or has no master_key, and 502 for an upstream response with no usable token, an already-expired lifetime, or one too large to seal. Because the status is chosen by an exhaustive match over the union, a new failure mode cannot be added without a matching status, and a DB outage can no longer read as a client 400 (mint and admission now agree on the same outage)

Three consequences follow from modelling those outcomes precisely. A database outage during identity resolution is a retryable 503 rather than a 400 that blames the caller. An upstream token whose reported lifetime is already elapsed is rejected with 502 rather than sealed into an hour-long envelope around a dead bearer, while an absent or unparseable lifetime still mints a capped envelope as before. And the refresh_token grant, which a bridge server can never legitimately receive (it seals no upstream refresh_token, so the client holds none to present), is rejected in _prepare_bridge_mint before the exchange, so a stray refresh request can never rotate or consume the client's upstream refresh credential; renewal is re-running authorization_code. Every other server, including true_passthrough bridges and non-bridge oauth_delegate servers, keeps returning the raw upstream token, so flag-off behavior is byte-identical

Sealing the key hash rather than a bare user_id is what lets the consumer reload the live key at admission and enforce the key's current team, org, and tool restrictions plus revocation, instead of admitting under a frozen identity. The hash is a one-way digest, not a usable credential (the edge rejects a bare hash presented as a bearer), and it is the same value get_key_object and the cache/DB layer key the record by, so the sealed reference resolves back to the key at admission. The token endpoint's key resolution is extracted into a shared _resolve_active_litellm_key so the per-user token store (which needs the user_id) and this mint (which needs the key hash) derive from one active-key-gated path

The upstream token response is validated field by field into a typed grant before sealing, so nothing untyped from the provider flows into the envelope, and the raw upstream token never appears in the bearer the client receives. Tests cover the mint-and-open round-trip through the consumer asserting the sealed key hash, the fail-closed path without an active key, both raw-token paths (true_passthrough bridge and non-bridge oauth_delegate), and the shared resolver returning the hash for an active key while rejecting a blocked key


Note

High Risk
Changes OAuth token issuance, API key validation, and master-key-derived credential sealing on a security-sensitive path; misclassification of DB outages or lifetime handling could break clients or admit bad tokens, though behavior is heavily tested and scoped to bridge oauth_delegate.

Overview
For DCR-bridge oauth_delegate MCP servers, the OAuth token endpoint now returns a gateway-bound envelope (llm_env_…) instead of the raw upstream access_token, so one client bearer binds LiteLLM key identity and forwards upstream authorization.

exchange_token_with_server gains a three-phase mint: _prepare_bridge_mint (before upstream POST) checks authorization_code only, master_key, and an active LiteLLM key via _resolve_active_litellm_key; _finish_bridge_mint (after exchange) validates the upstream JSON into a typed grant, seals with build_bridge_token_response, and omits upstream refresh_token from the response. Failures are tagged values mapped through _bridge_mint_error_response to RFC 6749 §5.2 statuses (400 vs 503 vs 500 vs 502), including no upstream call when identity or config fails so single-use codes are not burned.

Key resolution is refactored: _key_is_active (team keys without user_id allowed; bad expiry → inactive), distinct no_active_key / unavailable / unresolvable outcomes, and expires_in handling that rejects already-dead upstream lifetimes while capping unknown ones. Non-bridge and true_passthrough paths still return the raw upstream token.

Extensive tests cover mint/open round-trip, fail-closed pre-exchange paths, grant/lifetime edge cases, and resolver behavior.

Reviewed by Cursor Bugbot for commit 55ff3a2. Bugbot is set up for automated code reviews on this repo. Configure here.

@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_admission branch from eb9e1ef to b2fa98d Compare July 10, 2026 21:37
@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_token_mint branch from 3e0cd5d to 0c2b9ee Compare July 10, 2026 21:38
@greptile-apps

greptile-apps Bot commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

For DCR-bridge oauth_delegate MCP servers the OAuth token endpoint now returns a gateway-bound envelope (llm_env_…) instead of the raw upstream access token, completing the mint → admission loop. Every other server mode (true_passthrough, non-bridge oauth_delegate) keeps returning the raw token unchanged.

  • Three-phase mint in exchange_token_with_server: _prepare_bridge_mint runs before the upstream code exchange, validating grant type, master_key, and litellm identity (via the new _resolve_active_litellm_key) so no single-use code is ever burned on a precondition failure; _finish_bridge_mint runs after, sealing the typed upstream grant into a SealedEnvelope whose expires_in is derived from the JWT's own exp to avoid over-promising lifetime.
  • Tagged failure vocabulary: _KeyResolutionFailure / _BridgeMintError replace bare None returns, and exhaustive match + assert_never mappers route each failure to its RFC 6749 §5.2 status (400 / 503 / 500 / 502), so a DB outage now correctly yields 503 rather than 400.
  • Previous inline comments addressed: the eager access_token = token_response["access_token"] extraction that preceded the bridge check was removed; the non-bridge path extracts it inline after the bridge early-return, so neither path can KeyError on the bridge code path.

Confidence Score: 5/5

Safe to merge: the bridge mint is flag-gated behind both is_oauth_delegate and is_dcr_bridge, so no existing server mode is affected, and the three-phase architecture guarantees single-use codes are never consumed on precondition failures.

The change is well-isolated to the new bridge path and backed by 18 new mock unit tests covering all failure modes, the mint-and-open round-trip, and the flag-off paths. The two inline-comment issues flagged in the previous review iteration (eager access_token extraction before the bridge check, and the companion non-bridge extraction) are both correctly resolved in this diff. No pre-existing tests were weakened. The exhaustive match/assert_never pattern enforces that adding a failure mode requires a matching status, and the _key_is_active malformed-expiry guard prevents an unhandled ValueError from surfacing as a 500.

No files require special attention.

Important Files Changed

Filename Overview
litellm/proxy/_experimental/mcp_server/discoverable_endpoints.py Adds DCR-bridge oauth_delegate mint in a three-phase pipeline (prepare → exchange → finish); introduces tagged failure vocabulary, replaces bare-None key resolution with _ResolvedKey/_KeyResolutionFailure, and removes the pre-existing eager access_token extraction that blocked the bridge path. No new defects identified.
tests/test_litellm/proxy/_experimental/mcp_server/test_discoverable_endpoints.py Adds 18 new mock-only tests covering: mint-and-open round-trip, fail-closed paths (no identity, no master_key, DB outage, unresolvable), refresh_token rejection before exchange, upstream lifetime edge cases (expired, sub-second, unknown, float), too-large token 502, and the resolver's team-key / blocked-key / malformed-expiry behaviours. No existing tests modified.

Reviews (12): Last reviewed commit: "fix(mcp): treat a positive sub-second up..." | Re-trigger Greptile

@codecov

codecov Bot commented Jul 10, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 93.24324% with 10 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
..._experimental/mcp_server/discoverable_endpoints.py 93.24% 10 Missing ⚠️

📢 Thoughts on this report? Let us know!

@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_admission branch from b2fa98d to 9e72416 Compare July 10, 2026 22:05
@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_token_mint branch from 0c2b9ee to 73fa6ad Compare July 10, 2026 22:05
Comment thread litellm/proxy/_experimental/mcp_server/discoverable_endpoints.py Outdated
@veria-ai

veria-ai Bot commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 2 · PR risk: 0/10

@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_admission branch from 9e72416 to 70a4591 Compare July 10, 2026 22:33
@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_token_mint branch from 73fa6ad to 3b04add Compare July 10, 2026 22:34
@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_admission branch from 70a4591 to 5534930 Compare July 10, 2026 23:36
@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_token_mint branch from 3b04add to 09f87c9 Compare July 11, 2026 00:23
Comment thread litellm/proxy/_experimental/mcp_server/discoverable_endpoints.py Outdated
@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_token_mint branch 2 times, most recently from 7198ca7 to 06dde66 Compare July 11, 2026 18:49
@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_admission branch from a81ac2d to ea64ef7 Compare July 11, 2026 19:23
@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_token_mint branch from 06dde66 to 6c4f320 Compare July 11, 2026 19:24
@tin-berri

Copy link
Copy Markdown
Contributor Author

Fixed in d8a927a. The eager access_token = token_response["access_token"] extraction ran before the dcr_bridge branch, so a missing upstream access_token raised an unhandled KeyError and _bridge_grant_from_token_response's nil guard (which maps to a clean 502) was dead code. The extraction now lives only on the non-bridge result path, after the bridge early-return, so the bridge branch reaches its 502 guard; _validate_token_response and the per-user storage step both take the full token_response dict and never used the extracted variable, so moving it changes nothing else. Regression test drives a bridge oauth_delegate exchange whose upstream response omits access_token and asserts 502; without the fix it raises KeyError

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_token_mint branch from d8a927a to cfe3eeb Compare July 11, 2026 21:44
Comment thread litellm/proxy/_experimental/mcp_server/discoverable_endpoints.py Outdated
Base automatically changed from litellm_lit4338_delegate_admission to litellm_internal_staging July 11, 2026 21:59
@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_token_mint branch from cfe3eeb to d95b4b2 Compare July 11, 2026 22:05
@tin-berri

Copy link
Copy Markdown
Contributor Author

Rebased onto the merged #32824 (staging), and addressed the 4/5 finding in d95b4b2. _resolve_active_litellm_key gated on _active_key_user_id, which returns None both for blocked/expired keys and for valid keys that simply have no user_id, so a team-scoped or service-account key was wrongly rejected with invalid_request at the bridge token exchange. Split the active-state gate (_key_is_active: blocked and expiry only) from the user_id extraction: the mint seals the key hash rather than the user, and admission already reloads and enforces the key's live team/blocked/expiry state for a keyless-user key, so those keys can now mint. The per-user token store still resolves no user for such a key, since there is none to key a stored credential by, so that path is unchanged. Regression test caches an active key with user_id=None and asserts it resolves to hash_token(key) while the per-user extractor still returns None; before the fix the resolver returned None and no envelope could be minted

The earlier P1 (eager access_token extraction masking the 502 nil-guard) stays fixed

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@codspeed-hq

codspeed-hq Bot commented Jul 11, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_lit4338_delegate_token_mint (55ff3a2) with litellm_internal_staging (5fc1a3c)1

Open in CodSpeed

Footnotes

  1. No successful run was found on litellm_internal_staging (053e64e) during the generation of this report, so 5fc1a3c was used instead as the comparison base. There might be some changes unrelated to this pull request in this report.

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai rereview

Comment thread litellm/proxy/_experimental/mcp_server/discoverable_endpoints.py Outdated
Comment thread litellm/proxy/_experimental/mcp_server/discoverable_endpoints.py Outdated
@tin-berri

Copy link
Copy Markdown
Contributor Author

Addressed all three of the latest findings together in bd243fa, by changing the resolvers' return types rather than adding a branch per finding. All three were the same defect: a resolution step collapsed distinct outcomes into one None or a silent default, so the error mapper could not assign a truthful status.

Each resolution step now returns a precise tagged value. Identity resolution returns a _ResolvedKey or one of no_active_key / unavailable / unresolvable, classified the way admission's _reload_admitted_key classifies the same conditions. Upstream-lifetime classification returns positive seconds, "unspecified" (absent or unparseable, which the envelope caps), or "expired" (a parseable non-positive value). Upstream-grant validation returns a typed grant or no_access_token / expired_lifetime. Thin exhaustive mappers (match plus assert_never) lift each into one eight-member _BridgeMintError taxonomy, and one _bridge_mint_error_response assigns each its RFC 6749 section 5.2 status.

Finding "infra failures mapped as client errors": a DB outage during identity resolution now returns 503 temporarily_unavailable and a missing prisma_client returns 500, matching how admission statuses the same conditions on the egress side; only a genuinely absent or invalid credential is 400. Mint and admit no longer disagree under one outage.

Finding "zero lifetime becomes one-hour envelope": an explicit non-positive expires_in is now "expired" and rejected with 502 rather than sealed into a 1h envelope around a dead bearer. An absent or unparseable lifetime stays "unspecified" and still mints a capped envelope, since that is the by-design behaviour for an upstream that omits the field.

Finding "refresh grant discards rotated tokens": the refresh_token grant is rejected in _prepare_bridge_mint with unsupported_grant_type before any exchange, so it can never rotate or consume the client's upstream refresh credential. A bridge server seals no upstream refresh_token, so the client never holds one to present; renewal is re-running authorization_code.

Because the status is chosen by an exhaustive match over each union, a new failure mode cannot be added without a matching status, so this class of wrong-status bug cannot recur silently. Each of the three fixes is mutation-checked: reverting it turns its regression test red. Full unit file green at 164 passed, and pre-commit gates (ruff, ruff-strict budget, type-discipline, basedpyright budget, e2e basedpyright, circular imports) are green.

@greptileai

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

Comment thread litellm/proxy/_experimental/mcp_server/discoverable_endpoints.py Outdated
@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai rereview

@tin-berri

Copy link
Copy Markdown
Contributor Author

Fixed in f4298dc. _classify_upstream_lifetime decided expired from int(float(expires_in)), which truncates toward zero, so a positive fractional lifetime in (0, 1) became 0 and read as elapsed. It now decides expired on the parsed numeric value, so only a genuinely non-positive lifetime is expired; a positive sub-second value clamps up to the envelope's 1s floor (whole-second granularity) rather than being rejected, and values >= 1 still truncate toward zero so the envelope never claims more life than the upstream stated. Regression covers the classifier (0.5 and 0.001 clamp to 1, -0.5 stays expired) and the mint (a 0.5s upstream lifetime mints a 200 envelope, not a 502); reverting the fix reddens both.

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit f4298dc. Configure here.

tin-berri added 12 commits July 11, 2026 19:20
The mint bound only user_id/server_id into the envelope, which gave admission
no way to reload the caller's key and enforce its current restrictions. Seal the
hashed authorizing key instead (a one-way digest, not a usable credential), so
admission reloads the live UserAPIKeyAuth by it and the key's team/org/tool
permissions and revocation apply per request.

Extract the token endpoint's key resolution into a shared _resolve_active_litellm_key
so the per-user token store (user_id) and the bridge mint (key hash) derive from one
active-key-gated path, and fail the mint closed with invalid_request when no active
key accompanies the request.
…ks access_token

The eager access_token = token_response["access_token"] extraction ran before
the dcr_bridge branch, so a missing upstream access_token raised an unhandled
KeyError and _bridge_grant_from_token_response's nil guard (which maps to a clean
502) was dead code. Move the extraction onto the non-bridge result path so the
bridge branch reaches its 502 guard.
_resolve_active_litellm_key gated on _active_key_user_id, which returns None both
for blocked/expired keys AND for valid keys with no user_id, so a team-scoped or
service-account key was wrongly rejected with invalid_request at bridge token
exchange. Split the active-state gate (_key_is_active: blocked/expiry only) from
the user_id extraction; the mint seals the key hash, not the user, and admission
already handles a keyless-user key. The per-user token store still gets no user
for such a key, as there is none to key a stored credential by.
…lpers

The keyless-user fix moved these signatures, so their pre-existing Optional[...]
annotations counted against the diff and tripped the UP045 strict-budget gate.
Modernize the four touched return annotations to the X | None form the gate
wants; runtime behavior is unchanged.
Two correctness gaps in the bridge mint. _bridge_grant_from_token_response only
accepted an int expires_in, dropping a float (3600.0) or numeric-string ('3600')
lifetime to None so the envelope fell back to its 1h cap and could outlive a
shorter-lived upstream token; coerce it to a positive int (bool excluded). And
_key_is_active called datetime.fromisoformat on the str|datetime expires outside
the resolver's try, so a malformed stored expiry raised an unhandled 500 instead
of the fail-closed invalid_request; it now fails closed (inactive) on an
unparseable expiry. Regression tests cover int/float/string/bool coercion, the
short-float TTL, and the malformed-expiry fail-closed path.
Findings from a full adversarial review of the mint path across security,
correctness, error-handling, concurrency, and OAuth-protocol dimensions.

- expires_in coercion is now total: int(float(...)) can raise OverflowError on
  Infinity / a giant numeric string, which escaped the ValueError/TypeError catch
  and 500'd the token endpoint. Unified to catch OverflowError too.
- Resolve the litellm identity BEFORE exchanging the single-use upstream code, so
  a missing or transiently-unresolvable identity fails closed with invalid_request
  without burning the code (the mint re-resolves via a cache hit).
- The no-identity failure is now an RFC 6749 5.2-shaped invalid_request
  (JSONResponse, top-level error, no-store) instead of a detail-wrapped
  HTTPException, matching the BYOK OAuth endpoint.
- EnvelopeTooLarge (upstream token too big to seal) surfaces a 502, not a 500.
- The upstream refresh_token is no longer sealed into the envelope: the edge
  never consumes it, so it was dead weight embedding a long-lived upstream
  credential in the client bearer and enlarging the envelope; refresh is a
  follow-up (a dedicated refresh-envelope).

Security review found no exploitable defect (forgery, cross-server/user replay,
leakage, confused-deputy all closed). Regression tests cover the OverflowError,
the code-not-burned path, the RFC-shaped error, the 502, and the dropped refresh.
…te master_key first

Follow-up to the pre-exchange identity gate, which I had only added to the
authorization_code branch and which left the master_key check inside the mint
(after the upstream exchange) - so the very burn-then-fail pattern it was meant to
prevent still applied to refresh_token grants and to a misconfigured gateway.

- Hoist a single pre-exchange gate above the upstream call that covers BOTH grant
  types: it fails closed (invalid_request) on an unresolvable litellm identity and
  500s on an unset master_key BEFORE the single-use code or refresh token is
  exchanged/rotated, so a bad key or a misconfigured gateway never burns the
  upstream credential.
- Report expires_in from the envelope JWT's own second-truncated exp (rounding the
  elapsed portion up) instead of the raw expires_at - now delta, so the client is
  never told the bearer is valid past the ~1s point admission already expires it.

Regression tests assert the upstream exchange is never called on the no-identity
refresh grant and the master_key-unset path, and that the reported expires_in does
not overstate the JWT exp.
…ues pipeline

The dcr_bridge oauth_delegate token mint validated its preconditions in two
places: a pre-exchange guard inside exchange_token_with_server (master_key set,
resolvable litellm identity) and an authoritative re-check inside the post-exchange
_mint_bridge_delegate_token_response. Keeping the two in step by hand is what kept
producing the same class of finding: a precondition guarded on one grant branch but
not the other, master_key checked after the exchange on one path, identity resolved
twice, and each failure raising an ad-hoc HTTPException with its own status and body
shape.

Model the mint as three phases whose failures are values. _prepare_bridge_mint runs
before the exchange, checks every precondition once (master_key, then identity), and
returns either a frozen _BridgeMintReady carrying the resolved key hash and the
master-key-derived envelope keys, or a _BridgeMintError literal. Because every
precondition lives in prepare, and prepare runs before the upstream POST, no failure
can burn the single-use code or rotate a refresh token, for either grant type, by
construction rather than by a guard we have to remember to keep in sync.
_finish_bridge_mint runs after the exchange and has no preconditions left that can
fail; its only failure values are properties of the upstream response itself (no
usable access_token, or a token too large to seal). One mapper,
_bridge_mint_error_response, turns each _BridgeMintError into an RFC 6749 section
5.2-shaped body with a status truthful about where the failure is (400 for the
caller, 500 for gateway config, 502 for the upstream), with an exhaustive match plus
assert_never so a new failure mode cannot be added without a matching status.

Behavior is unchanged for the client. Every failure that previously raised now
returns the same status as an OAuth error body, which is the correct token-endpoint
contract; the three tests that asserted a raised HTTPException now assert the
returned response. _exchange_for_bridge_server additionally asserts the identity
resolver is awaited exactly once for a bridge server and never for a non-bridge one.
…boundary

_finish_bridge_mint floored the reported expires_in at 1. Admission expires the
envelope against the JWT's second-truncated exp, so when the mint lands in the same
second that exp falls on (a sub-second upstream lifetime, for instance), the true
remaining life is 0 and reporting 1 tells the client the bearer lives one second past
the point admission already rejects it. Floor at 0 instead so the reported lifetime
never overstates the exp; the value still cannot go negative.

The regression pins the boundary directly: minting at now=100.25 with a 1s upstream
token seals exp=101, and the reported expires_in is max(0, 101 - ceil(100.25)) = 0.
Under the old floor of 1 it reads 1, so the test fails on that mutation.

Also drops the unused mcp_server parameter from _prepare_bridge_mint; identity and
key derivation there never referenced the server.
…tus is truthful by construction

Three findings landed together, all one defect: a resolution step crushed several distinct outcomes
into a single None or a silent default, so the mint's error mapper could not tell them apart and
assigned the wrong status. Identity resolution mapped a database outage to the same None as a missing
credential, which the mint reported as 400 invalid_request, blaming the caller for a gateway outage
while admission statuses the same outage 503/500. Lifetime coercion mapped an explicit non-positive
expires_in to the same None as an absent one, so an upstream token the IdP reports as already dead was
sealed into an hour-long envelope. And the refresh_token grant was run through the upstream exchange
(which can rotate the client's upstream refresh credential) and its result then discarded, even though
a bridge server seals no refresh_token and the client never holds one to present.

Rather than add a mapping branch per finding, the fix changes the return types so a wrong status is not
representable. Each resolution step now returns a precise tagged value instead of None: identity
resolution returns a _ResolvedKey or one of no_active_key / unavailable / unresolvable, classified the
same way admission's _reload_admitted_key classifies the same conditions; upstream-lifetime
classification returns a positive number of seconds, "unspecified" (absent or unparseable, which the
envelope caps), or "expired" (a parseable non-positive value, an already-dead token); and upstream-grant
validation returns a typed grant or one of no_access_token / expired_lifetime. Thin exhaustive mappers
(match plus assert_never) lift each vocabulary into one bridge-mint taxonomy of eight named failures,
and a single _bridge_mint_error_response gives each its truthful RFC 6749 §5.2 status: 400 for the
caller's missing credential or an unsupported grant, 503 for a transient auth-DB outage, 500 for a
gateway that cannot resolve identity or is not configured, and 502 for an upstream response with no
usable token, an already-expired lifetime, or a token too large to seal. Adding a failure mode now
requires a new literal and a match arm the type checker forces, so the class of wrong-status bug cannot
recur silently.

The refresh_token grant is rejected in _prepare_bridge_mint before the exchange with
unsupported_grant_type, so it can never rotate or consume the client's upstream refresh credential;
renewal is re-running authorization_code, as the sealed refresh_token=None already intends. An absent or
unparseable expires_in still mints a capped envelope (the by-design behaviour for an upstream that omits
the field); only an explicitly-dead lifetime is rejected.

Tests cover the resolver's three failure classes (including a real connection-error outage and a missing
prisma_client), the mint statuses for each (503 before the upstream exchange, 500, 502 on an expired
upstream lifetime, and a capped mint on an unknown one), and the refresh-grant rejection before any
exchange. The three findings are mutation-checked: reverting each fix turns its regression test red.
… expired

_classify_upstream_lifetime decided "expired" from int(float(expires_in)), which truncates toward
zero, so a positive fractional lifetime in (0, 1) became 0 and was misread as already elapsed. That
rejected the mint with 502 in _finish_bridge_mint after the single-use upstream code had already been
consumed, even though the upstream reported a positive remaining lifetime.

Decide expired on the parsed numeric value rather than its truncated int, so only a genuinely
non-positive value is expired. The envelope works in whole seconds and cannot represent a sub-second
lifetime, so a positive value that truncates to 0 clamps up to the 1s floor instead of being rejected.
Values >= 1 still truncate toward zero so the envelope never claims more life than the upstream stated,
and NaN / Infinity / oversized input still read as unparseable ("unspecified").

Regression covers the classifier (0.5 and 0.001 clamp to 1, 1.9 truncates to 1, -0.5 stays expired) and
the mint (a 0.5s upstream lifetime mints a 200 envelope rather than a 502); reverting to the
truncate-then-check reddens both.
@tin-berri
tin-berri force-pushed the litellm_lit4338_delegate_token_mint branch from f4298dc to 55ff3a2 Compare July 12, 2026 02:20
@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai rereview

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 55ff3a2. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM; thanks!

@tin-berri
tin-berri merged commit 3d400b5 into litellm_internal_staging Jul 13, 2026
131 checks passed
@tin-berri
tin-berri deleted the litellm_lit4338_delegate_token_mint branch July 13, 2026 17:36
doonga pushed a commit to greyrock-labs/home-ops that referenced this pull request Jul 29, 2026
…4.0) (#201)

This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.93.0` → `v1.94.0` |

---

### Release Notes

<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>

### [`v1.94.0`](https://github.com/BerriAI/litellm/releases/tag/v1.94.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.94.0...v1.94.0)

##### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.94.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.94.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.94.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

##### What's Changed

- feat(ui): working Test Connection for the complexity auto router by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;32950](https://github.com/BerriAI/litellm/pull/32950)
- fix(xecguard): use StandardLoggingGuardrailInformation in logging hook by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;32911](https://github.com/BerriAI/litellm/pull/32911)
- feat(ui): adopt openapi-react-query ($api) and convert useCustomers by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32949](https://github.com/BerriAI/litellm/pull/32949)
- refactor(ui): colocate the mcp-servers view, keeping the shared mcp\_tools surface by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32968](https://github.com/BerriAI/litellm/pull/32968)
- refactor(ui): convert endpoint usage charts to shadcn/recharts by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32723](https://github.com/BerriAI/litellm/pull/32723)
- fix(proxy-auth): stop unrecognized model namespaces slipping through provider wildcard keys by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32979](https://github.com/BerriAI/litellm/pull/32979)
- feat(router): random-pick multi-model complexity tiers by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32967](https://github.com/BerriAI/litellm/pull/32967)
- fix(xecguard): sanitize scan result before recording it for logging by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;32935](https://github.com/BerriAI/litellm/pull/32935)
- fix(auto\_router): filter embedding models in complexity tab dropdowns, require all tiers, inline validation by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;32978](https://github.com/BerriAI/litellm/pull/32978)
- fix(anthropic): translate raw adaptive thinking for pre-4.6 models on chat completions and Bedrock Converse by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;32944](https://github.com/BerriAI/litellm/pull/32944)
- feat(router): add Router(plugins=\[...]) routing-plugin pipeline by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32972](https://github.com/BerriAI/litellm/pull/32972)
- feat(router): soft-floor adaptive mode for complexity router by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32947](https://github.com/BerriAI/litellm/pull/32947)
- docs(github): add QA runbook section to the PR template by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32965](https://github.com/BerriAI/litellm/pull/32965)
- fix(model\_cost): add supports\_reasoning: false to Gemini image generation models by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32836](https://github.com/BerriAI/litellm/pull/32836)
- build(dev-env): add make bootstrap and unprovisioned-checkout preflight to pre-commit by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32981](https://github.com/BerriAI/litellm/pull/32981)
- ci(ui): report only error-level knip findings in CI by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32971](https://github.com/BerriAI/litellm/pull/32971)
- feat(batches): track cost for unmanaged Bedrock batches, generalize the flag by [@&#8203;Sameerlite](https://github.com/Sameerlite) in [#&#8203;32315](https://github.com/BerriAI/litellm/pull/32315)
- fix(guardrails): walk custom\_tool\_call\_output items in \_content\_utils by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;32969](https://github.com/BerriAI/litellm/pull/32969)
- fix: show and allow editing team model aliases after team creation by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33047](https://github.com/BerriAI/litellm/pull/33047)
- chore(deps): bump pillow to 12.3.0 to resolve osv-scan CVEs by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33093](https://github.com/BerriAI/litellm/pull/33093)
- feat(mcp): mint gateway-bound envelope at the token endpoint for dcr\_bridge oauth\_delegate by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32828](https://github.com/BerriAI/litellm/pull/32828)
- fix(mcp): surface rejected delegate-auth upstream tokens as connect-time 401 by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32741](https://github.com/BerriAI/litellm/pull/32741)
- fix(proxy): track unauthenticated pass-through requests in spend logs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;32410](https://github.com/BerriAI/litellm/pull/32410)
- feat(lasso): send source.type for Used By attribution by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33090](https://github.com/BerriAI/litellm/pull/33090)
- fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@&#8203;thibault-linktree](https://github.com/thibault-linktree) in [#&#8203;33025](https://github.com/BerriAI/litellm/pull/33025)
- fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33099](https://github.com/BerriAI/litellm/pull/33099)
- fix(completion): forward aws credential kwargs into litellm\_params so the responses bridge keeps WIF auth by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32956](https://github.com/BerriAI/litellm/pull/32956)
- fix(ui): respect litellm\_key\_header\_name in BYOK credential save and workflow runs fetches by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33103](https://github.com/BerriAI/litellm/pull/33103)
- refactor(ui): standardize debounce waits behind shared DEBOUNCE\_WAIT\_MS constant by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33040](https://github.com/BerriAI/litellm/pull/33040)
- feat(ui): rebuild the Virtual Keys table on the shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;32991](https://github.com/BerriAI/litellm/pull/32991)
- fix: redact async complete streaming response for custom callbacks by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33106](https://github.com/BerriAI/litellm/pull/33106)
- build(ui): bump [@&#8203;tanstack/react-pacer](https://github.com/tanstack/react-pacer) from 0.2.0 to 0.22.1 by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33041](https://github.com/BerriAI/litellm/pull/33041)
- fix(ui): address Virtual Keys redesign review nits by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33112](https://github.com/BerriAI/litellm/pull/33112)
- fix(openai/responses): clamp max\_output\_tokens below API minimum by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33098](https://github.com/BerriAI/litellm/pull/33098)
- fix(prometheus): read v3 rate limiter remaining values for per-key model gauges by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33119](https://github.com/BerriAI/litellm/pull/33119)
- fix(ui): drop w-full from page-content wrappers to remove 32px horizontal overflow by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33118](https://github.com/BerriAI/litellm/pull/33118)
- refactor(ui): migrate straightforward value debounces to react-pacer by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33042](https://github.com/BerriAI/litellm/pull/33042)
- feat(mcp): interactive SSO sign-in for dcr\_bridge oauth\_delegate DCR clients by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32946](https://github.com/BerriAI/litellm/pull/32946)
- test(proxy): add regression tests for management\_endpoints edge cases by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;32976](https://github.com/BerriAI/litellm/pull/32976)
- fix(auto-router): correct Responses API tool\_choice shape and propagate alias litellm\_params by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32974](https://github.com/BerriAI/litellm/pull/32974)
- fix(ui): render the sidebar scrollbar with shadcn ScrollArea by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33124](https://github.com/BerriAI/litellm/pull/33124)
- refactor(ui): migrate callback debounce sites to react-pacer with regression tests by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33043](https://github.com/BerriAI/litellm/pull/33043)
- chore: add CODEOWNERS for ui and proxy UI build artifacts by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33131](https://github.com/BerriAI/litellm/pull/33131)
- feat(mcp): client-held refresh envelope for the dcr\_bridge oauth\_delegate flow by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32980](https://github.com/BerriAI/litellm/pull/32980)
- feat(ui): rebuild the Teams table on the shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33128](https://github.com/BerriAI/litellm/pull/33128)
- fix(mcp): relay upstream OAuth token and DCR rejections instead of a generic 500 by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33113](https://github.com/BerriAI/litellm/pull/33113)
- fix(keys): persist key\_type so the UI shows correct key scope instead of "All Proxy Models" by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33115](https://github.com/BerriAI/litellm/pull/33115)
- feat(router): opt-in session affinity for complexity router by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;33126](https://github.com/BerriAI/litellm/pull/33126)
- feat(prometheus): expose video duration and image count consumption metrics by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33138](https://github.com/BerriAI/litellm/pull/33138)
- test(e2e): otel trace completeness on /chat/completions by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33132](https://github.com/BerriAI/litellm/pull/33132)
- fix(sso): paginate through all pages when fetching service principal group assignments by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33149](https://github.com/BerriAI/litellm/pull/33149)
- test(e2e): otel trace completeness on /v1/messages by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33133](https://github.com/BerriAI/litellm/pull/33133)
- feat(ui): add adaptive routing settings to Auto-Router v2 by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;33146](https://github.com/BerriAI/litellm/pull/33146)
- refactor(mcp): extract the dcr\_bridge token flow into bridge\_token\_flow\.py by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33141](https://github.com/BerriAI/litellm/pull/33141)
- chore: bump litellm 1.93.0 -> 1.94.0, litellm-enterprise 0.1.49 -> 0.1.50, litellm-proxy-extras 0.4.76 -> 0.4.77 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33229](https://github.com/BerriAI/litellm/pull/33229)
- fix(proxy): route master key to team-scoped models by [@&#8203;kunal2002](https://github.com/kunal2002) in [#&#8203;32926](https://github.com/BerriAI/litellm/pull/32926)
- chore(deps): pin httplib2 and setuptools transitive floors by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33233](https://github.com/BerriAI/litellm/pull/33233)
- feat(ui): left-anchor the Create Key and Create Team CTAs by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33248](https://github.com/BerriAI/litellm/pull/33248)
- fix(anthropic/passthrough): drop incompatible temperature when downgrading adaptive thinking for pre-4.6 models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33244](https://github.com/BerriAI/litellm/pull/33244)
- fix(guardrails): run apply\_guardrail-style model-level pre\_call guardrails at deployment hook by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33136](https://github.com/BerriAI/litellm/pull/33136)
- fix(proxy)!: enforce user budget on team keys (read-time + reservation) with UI opt-out by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;32005](https://github.com/BerriAI/litellm/pull/32005)
- fix(e2e): bound spend-log snapshots to a /spend/logs/v2 window by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33265](https://github.com/BerriAI/litellm/pull/33265)
- test(e2e): cover key rpm/tpm rate limiting, window reset, and pacing headers by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32914](https://github.com/BerriAI/litellm/pull/32914)
- fix(anthropic): use native output capability by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;33235](https://github.com/BerriAI/litellm/pull/33235)
- fix(ci): retry setup-uv installs to survive transient manifest fetch failures by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33279](https://github.com/BerriAI/litellm/pull/33279)
- fix(proxy): never log raw virtual keys in key insertion debug output by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33268](https://github.com/BerriAI/litellm/pull/33268)
- fix(bedrock\_mantle): route xai.grok-4.3 via /openai/v1 frontier path by [@&#8203;marty-sullivan](https://github.com/marty-sullivan) in [#&#8203;33027](https://github.com/BerriAI/litellm/pull/33027)
- feat(pricing): add gemini-omni-flash-preview with video output token pricing by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33274](https://github.com/BerriAI/litellm/pull/33274)
- fix(auth): scope the JWT enterprise gate to actual JWTs by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33296](https://github.com/BerriAI/litellm/pull/33296)
- fix(s3): sanitize slashes in response-id-derived object key file name by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33271](https://github.com/BerriAI/litellm/pull/33271)
- refactor(ui): migrate guardrails table onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33303](https://github.com/BerriAI/litellm/pull/33303)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33308](https://github.com/BerriAI/litellm/pull/33308)
- feat(guardrails): streaming text transformation in generic\_guardrail\_api by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33110](https://github.com/BerriAI/litellm/pull/33110)
- test(e2e): cover model-aware mid-conversation system handling on Bedrock Invoke /v1/messages by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32963](https://github.com/BerriAI/litellm/pull/32963)
- test(claude\_code): move the Claude Code compatibility matrix under tests/e2e by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;32548](https://github.com/BerriAI/litellm/pull/32548)
- chore(ci): sync litellm\_internal\_staging into daily OSS branch by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33337](https://github.com/BerriAI/litellm/pull/33337)
- feat(bedrock guardrails): add resource-less InvokeGuardrailChecks (detect-only) mode by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33299](https://github.com/BerriAI/litellm/pull/33299)
- Revert "chore(ci): sync litellm\_internal\_staging into daily OSS branch" by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33339](https://github.com/BerriAI/litellm/pull/33339)
- fix(websearch): intercept web search on the Responses API by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33129](https://github.com/BerriAI/litellm/pull/33129)
- fix(anthropic-adapter): drop empty content\_block\_delta events by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33315](https://github.com/BerriAI/litellm/pull/33315)
- fix(mcp): persist discovered OAuth endpoints and keep last known good on failed re-discovery by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33286](https://github.com/BerriAI/litellm/pull/33286)
- test(e2e): otel trace completeness on streaming chat, messages, and responses (LIT-3787) by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33234](https://github.com/BerriAI/litellm/pull/33234)
- feat(router): resolve auto-router routing plugins from proxy YAML config by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;33251](https://github.com/BerriAI/litellm/pull/33251)
- test(e2e): failed request error span carries the full untruncated message and status (LIT-4179) by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33304](https://github.com/BerriAI/litellm/pull/33304)
- refactor(ui): migrate tags table onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33314](https://github.com/BerriAI/litellm/pull/33314)
- fix(cli): surface actionable CLI SSO errors when CLI and proxy versions skew by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33309](https://github.com/BerriAI/litellm/pull/33309)
- feat(bedrock\_mantle): add GPT-5.6 sol/terra/luna to model cost map by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33412](https://github.com/BerriAI/litellm/pull/33412)
- chore(codeowners): exempt generated schema.d.ts from UI ownership by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33411](https://github.com/BerriAI/litellm/pull/33411)
- feat(proxy): push-based OTLP billable-request metering for enterprise deployments by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;31592](https://github.com/BerriAI/litellm/pull/31592)
- fix(mcp): cap per-user OAuth token cache TTL at the token's own lifetime by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33346](https://github.com/BerriAI/litellm/pull/33346)
- feat(ui): move Caching out of Experimental into Developer Tools by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33432](https://github.com/BerriAI/litellm/pull/33432)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33425](https://github.com/BerriAI/litellm/pull/33425)
- chore(ci): merge daily internal staging branch by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33335](https://github.com/BerriAI/litellm/pull/33335)
- feat(guardrails): add Compresr guardrail for query-aware context compression by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33295](https://github.com/BerriAI/litellm/pull/33295)
- fix(logging): preserve callback order in get\_combined\_callback\_list by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33005](https://github.com/BerriAI/litellm/pull/33005)
- fix(logging): redact assistant tool call arguments in spend logs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33111](https://github.com/BerriAI/litellm/pull/33111)
- fix(anthropic): honor messages request timeout by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33418](https://github.com/BerriAI/litellm/pull/33418)
- fix(llm\_guard): apply sanitized prompt returned by moderation API to request by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33331](https://github.com/BerriAI/litellm/pull/33331)
- fix(logging): stop pinning large request payloads past request end by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33455](https://github.com/BerriAI/litellm/pull/33455)
- feat(guardrails): forward optional metadata on POST /guardrails/apply\_guardrail by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33067](https://github.com/BerriAI/litellm/pull/33067)
- build: raise requires-python cap to <3.15 so Python 3.14 installs current releases by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33438](https://github.com/BerriAI/litellm/pull/33438)
- feat(ui): add reusable BetaBadge and use it for Projects sidebar item by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33449](https://github.com/BerriAI/litellm/pull/33449)
- fix(mcp): discover missing OAuth scopes and token\_url when authorization\_url is set manually by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33317](https://github.com/BerriAI/litellm/pull/33317)
- feat(ui): show exact license expiration date in usage cards by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33478](https://github.com/BerriAI/litellm/pull/33478)
- test(claude\_code): rename misleading REPO\_ROOT to SUITE\_ROOT in test\_v0\_layout by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33472](https://github.com/BerriAI/litellm/pull/33472)
- build(deps): update ddtrace to the 4.x line by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33484](https://github.com/BerriAI/litellm/pull/33484)
- fix(complexity\_router): return empty dict from \_classifier\_call\_metadata when metadata is absent by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33452](https://github.com/BerriAI/litellm/pull/33452)
- fix(ui/chat): resolve chat routes at render time so navigation works under server\_root\_path by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33446](https://github.com/BerriAI/litellm/pull/33446)
- fix(key management): enforce minimum custom key length and mask short keys in key\_name by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33462](https://github.com/BerriAI/litellm/pull/33462)
- chore(ui): remove unmounted UsageIndicator and the Hide Usage Indicator flag by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33482](https://github.com/BerriAI/litellm/pull/33482)
- test(e2e/claude\_code): add passthrough matrix row for the big-3 clouds and Anthropic API by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33473](https://github.com/BerriAI/litellm/pull/33473)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33491](https://github.com/BerriAI/litellm/pull/33491)
- fix(ui): stop sending the complexity-router pseudo-model to /health/test\_connection by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;33498](https://github.com/BerriAI/litellm/pull/33498)
- feat(cli): add lite up/down to ambiently route Claude Code through the proxy by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;33231](https://github.com/BerriAI/litellm/pull/33231)
- feat(complexity\_router): enable session\_affinity by default by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33500](https://github.com/BerriAI/litellm/pull/33500)
- fix(anthropic): stop 500 on combined thinking+signature streaming chunk by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33505](https://github.com/BerriAI/litellm/pull/33505)
- feat(autoroute): prompt for semantic keywords per tier in configure wizard by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33508](https://github.com/BerriAI/litellm/pull/33508)
- fix(cli/anthropic): unblock lite autoroute proxy deps, adaptive thinking, and thinking+signature streaming by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33507](https://github.com/BerriAI/litellm/pull/33507)
- test(ocr): use mistral-document-ai-2512 in azure\_ai OCR tests by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33489](https://github.com/BerriAI/litellm/pull/33489)
- fix(guardrails): show YAML-defined guardrails in the Guardrail Monitor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;32853](https://github.com/BerriAI/litellm/pull/32853)
- refactor(ui): migrate policies, deleted keys, deleted teams, budgets, and search tools tables onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33357](https://github.com/BerriAI/litellm/pull/33357)
- refactor(ui): migrate vector stores, prompts, and skills tables onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33343](https://github.com/BerriAI/litellm/pull/33343)
- test(e2e): datadog log delivery for successful chat, messages, and responses (LIT-4447) by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33415](https://github.com/BerriAI/litellm/pull/33415)
- fix(cli): make CLI output ASCII-only so it doesn't crash legacy Windows consoles by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33465](https://github.com/BerriAI/litellm/pull/33465)
- fix: remove dead user-cache lookup with None key in spend-update path by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33555](https://github.com/BerriAI/litellm/pull/33555)
- feat(helm): add per-component PodDisruptionBudget and topologySpreadConstraints to componentized chart by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33430](https://github.com/BerriAI/litellm/pull/33430)
- fix(e2e/claude\_code): unblock stage collection, align proxy env names, register compat models by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33433](https://github.com/BerriAI/litellm/pull/33433)
- fix(mcp): index authed request-time tools missing from the semantic filter startup index by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33318](https://github.com/BerriAI/litellm/pull/33318)
- fix(streaming): use provider-reported usage cost for OpenRouter streams by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32255](https://github.com/BerriAI/litellm/pull/32255)
- feat(mcp): issuer-anchored OAuth discovery (RFC 8414 §3.3) to close the authorization-server mix-up by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33450](https://github.com/BerriAI/litellm/pull/33450)
- feat(logging): add user and team level spend and budget to StandardLoggingPayload metadata by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33459](https://github.com/BerriAI/litellm/pull/33459)
- fix(router): cast model\_info cost values to float in \_set\_model\_group\_info by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33556](https://github.com/BerriAI/litellm/pull/33556)
- chore(e2e): establish litellm\_e2e\_staging integration line by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33502](https://github.com/BerriAI/litellm/pull/33502)
- fix(ui): navigate to /ui/login/ with trailing slash via hard navigation by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33561](https://github.com/BerriAI/litellm/pull/33561)
- feat(logging): add structured budget fields to budget rejection failure logs by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33460](https://github.com/BerriAI/litellm/pull/33460)
- fix(streaming): surface upstream connection resets instead of empty 200 streams by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33222](https://github.com/BerriAI/litellm/pull/33222)
- fix(proxy\_cli): reap orphaned prisma query-engine processes when a worker dies by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33424](https://github.com/BerriAI/litellm/pull/33424)
- test(reasoning\_effort\_grid): enable azure fable-5 and opus-4-8 grid cells by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33485](https://github.com/BerriAI/litellm/pull/33485)
- build(deps): bump uvicorn lock to 0.51.0 so worker health-check and jitter flags take effect by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33574](https://github.com/BerriAI/litellm/pull/33574)
- fix(proxy): coerce default\_internal\_user\_params.max\_budget to float on config load by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;32434](https://github.com/BerriAI/litellm/pull/32434)
- fix(router): honor per-request routing\_strategy from key/team router\_settings by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33429](https://github.com/BerriAI/litellm/pull/33429)
- fix(redis): honor ssl value instead of key presence when building async connection pool by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;32590](https://github.com/BerriAI/litellm/pull/32590)
- fix(langfuse\_otel): build per-request OTLP exporter from key and team dynamic Langfuse credentials by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;32437](https://github.com/BerriAI/litellm/pull/32437)
- ci: run zizmor and proxy-db unit tests on PRs targeting litellm\_ branches by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33568](https://github.com/BerriAI/litellm/pull/33568)
- fix(router): apply team/key enable\_tag\_filtering to tag routing by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33436](https://github.com/BerriAI/litellm/pull/33436)
- feat(proxy): add disable\_auto\_add\_proxy\_admin\_to\_teams flag by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33563](https://github.com/BerriAI/litellm/pull/33563)
- fix(proxy): stop stale auth cache re-publish to Redis so key updates propagate across replicas by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33565](https://github.com/BerriAI/litellm/pull/33565)
- feat(e2e): emit structured E2E\_RESULT lines for package status history by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33578](https://github.com/BerriAI/litellm/pull/33578)
- chore: bump litellm-enterprise 0.1.50 -> 0.1.51, litellm-proxy-extras 0.4.77 -> 0.4.78 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33571](https://github.com/BerriAI/litellm/pull/33571)
- fix(docker): restore litellm-proxy-extras source dir in runtime images by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33592](https://github.com/BerriAI/litellm/pull/33592)
- feat(ui): require embedding model for semantic auto router by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33313](https://github.com/BerriAI/litellm/pull/33313)
- feat(scim): ingest and round-trip SCIM entitlements and roles user attributes by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33587](https://github.com/BerriAI/litellm/pull/33587)
- refactor(ui): migrate 5 simple tables onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33548](https://github.com/BerriAI/litellm/pull/33548)
- fix(model\_armor): restore reference attachments via skip\_unscannable\_attachments and remove the attachment count cap by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33554](https://github.com/BerriAI/litellm/pull/33554)
- fix(mcp): keep the MCP reference intact when the semantic filter narrows tools by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33584](https://github.com/BerriAI/litellm/pull/33584)
- fix(sso): stop enforcing UI session budget on CLI login tokens by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33312](https://github.com/BerriAI/litellm/pull/33312)
- test: e2e staging leftovers by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33613](https://github.com/BerriAI/litellm/pull/33613)
- test(e2e/claude\_code): add GPT-5.6 Sol/Terra/Luna columns for OpenAI, Azure OpenAI, and Bedrock Mantle by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33474](https://github.com/BerriAI/litellm/pull/33474)
- test(e2e): otel streaming spans record a real ttft below span duration by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33588](https://github.com/BerriAI/litellm/pull/33588)
- fix(mcp): make the preemptive-401 OAuth challenge decision mode-aware by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33586](https://github.com/BerriAI/litellm/pull/33586)
- fix(ui): show all teams in policy attachment form for admins by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33628](https://github.com/BerriAI/litellm/pull/33628)
- refactor(ui): migrate AI Hub, public hub, and MCP Toolsets tables onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33629](https://github.com/BerriAI/litellm/pull/33629)
- test(e2e): datadog log delivery for streamed routes, read back from the real datadog api by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33566](https://github.com/BerriAI/litellm/pull/33566)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33640](https://github.com/BerriAI/litellm/pull/33640)
- fix(vertex\_ai): surface Gemini grounding toolUsePromptTokenCount in Usage by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33533](https://github.com/BerriAI/litellm/pull/33533)
- test(e2e): harness fixes for stage job green (skips + router/UI/budget) by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33634](https://github.com/BerriAI/litellm/pull/33634)
- fix(router): resolve prompt cache minimum per model instead of a flat 1024 by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33637](https://github.com/BerriAI/litellm/pull/33637)
- fix(logging): classify async anthropic\_messages and generate\_content as async by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33589](https://github.com/BerriAI/litellm/pull/33589)
- fix(ui): remove Chat item from dashboard leftnav by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33647](https://github.com/BerriAI/litellm/pull/33647)
- fix(router): tag-aware pre-routing strategy selection for shared model\_name by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33691](https://github.com/BerriAI/litellm/pull/33691)
- fix(proxy): enforce max\_parallel\_requests as a per-slot concurrency gauge by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;32441](https://github.com/BerriAI/litellm/pull/32441)
- fix(proxy): stop treating upstream model body field as a LiteLLM model on auth-enforced pass-through routes by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33710](https://github.com/BerriAI/litellm/pull/33710)
- fix(mcp): expand toolset grants in shared permission primitives so tools/call honors them by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33612](https://github.com/BerriAI/litellm/pull/33612)
- feat(complexity-router): user-triggered escalation keywords by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33656](https://github.com/BerriAI/litellm/pull/33656)
- fix(fireworks\_ai): bill prompt-cache hits at cache\_read rate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33714](https://github.com/BerriAI/litellm/pull/33714)
- fix(pricing): mark realtime-only gpt-realtime models as mode realtime by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33728](https://github.com/BerriAI/litellm/pull/33728)
- fix(rag): track LLM completion usage and spend for /v1/rag/query by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;32438](https://github.com/BerriAI/litellm/pull/32438)
- feat(anthropic): add enable\_anthropic\_prompt\_caching for automatic cache\_control injection by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33573](https://github.com/BerriAI/litellm/pull/33573)
- fix(anthropic): self-heal on missing thinking-signature errors from Bedrock/Vertex by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33719](https://github.com/BerriAI/litellm/pull/33719)
- fix(proxy): resolve router\_settings.plugins dotted paths and load plugins from installed packages by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33644](https://github.com/BerriAI/litellm/pull/33644)
- test(e2e): budget refusals are 429 for bare keys and team caps block every team key by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33632](https://github.com/BerriAI/litellm/pull/33632)
- feat(router): add router plugin reference catalog by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33746](https://github.com/BerriAI/litellm/pull/33746)
- test(e2e): assert an org budget block is a 429 naming the organization by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33638](https://github.com/BerriAI/litellm/pull/33638)
- fix(proxy): bill partial streamed spend when the client disconnects mid-stream by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33736](https://github.com/BerriAI/litellm/pull/33736)
- test(e2e): delete unreferenced Grafana panel docs by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33743](https://github.com/BerriAI/litellm/pull/33743)
- docs(tests/e2e): align skip-vs-fail docs with the hard-fail contract and scope the no-unit-tests rule by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33755](https://github.com/BerriAI/litellm/pull/33755)
- refactor(e2e): replace bespoke result reporter with standard JUnit report by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33758](https://github.com/BerriAI/litellm/pull/33758)
- test(e2e): user budget across keys and team member budget isolation by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33745](https://github.com/BerriAI/litellm/pull/33745)
- refactor(e2e): remove bob\_the\_builder; drive remediation from a Grafana alert (provisioned outside the repo) by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33749](https://github.com/BerriAI/litellm/pull/33749)
- feat(mcp): per-server outcomes for aggregate tools/list and truthful single-server REST statuses by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33153](https://github.com/BerriAI/litellm/pull/33153)
- test(e2e): mcp suite for key-without-access denial by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33752](https://github.com/BerriAI/litellm/pull/33752)
- chore(ci): merge oss branch by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33784](https://github.com/BerriAI/litellm/pull/33784)
- chore(ci): merge oss branch - July 17th by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33793](https://github.com/BerriAI/litellm/pull/33793)
- fix(ui): migrate tag deletion to shared DeleteResourceModal by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33795](https://github.com/BerriAI/litellm/pull/33795)
- build(rust): raise pyo3 to 0.29 so the native bridge compiles on Python 3.14 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33798](https://github.com/BerriAI/litellm/pull/33798)
- chore(guardrails): remove docstring from singulr module for consistency by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33800](https://github.com/BerriAI/litellm/pull/33800)
- fix(ui): stop credential edit from persisting the masked api key by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33797](https://github.com/BerriAI/litellm/pull/33797)
- build(deps): allow redisvl, pypdf, and openapi-core on Python 3.14 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33801](https://github.com/BerriAI/litellm/pull/33801)
- test(proxy): make streaming-cancel mocks awaitable for the disconnect slot release by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33802](https://github.com/BerriAI/litellm/pull/33802)
- test(e2e): a member's team budget cuts off only that member's key by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33718](https://github.com/BerriAI/litellm/pull/33718)
- test(e2e): a user's max\_budget follows the person across personal and team keys by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33762](https://github.com/BerriAI/litellm/pull/33762)
- test(e2e): skip flaky OpenAI GPT cells; raise multi-window max\_tokens by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33799](https://github.com/BerriAI/litellm/pull/33799)
- chore: remove accidentally committed dist tarball and ignore dist/ by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33805](https://github.com/BerriAI/litellm/pull/33805)
- fix(passthrough): stop classifying plain 'predict'/'search' paths as Vertex by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33658](https://github.com/BerriAI/litellm/pull/33658)
- build(deps): bump mcp lock to 1.28.1 to clear image-scan findings by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33803](https://github.com/BerriAI/litellm/pull/33803)
- test(pricing): pin the realtime mode assertion to the bundled cost map by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33806](https://github.com/BerriAI/litellm/pull/33806)
- fix(proxy): derive session id from Anthropic metadata.user\_id for session affinity by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33723](https://github.com/BerriAI/litellm/pull/33723)
- test(e2e): budget reset diagonal for team, org, user, and [#&#8203;32005](https://github.com/BerriAI/litellm/issues/32005) team-member keys by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33771](https://github.com/BerriAI/litellm/pull/33771)
- fix(proxy): source /v1/models token limits from the cost map instead of Router.get\_model\_group\_info by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33721](https://github.com/BerriAI/litellm/pull/33721)
- fix(fireworks\_ai): correct glm-5p2 prompt-cache read price to $0.14/1M by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33796](https://github.com/BerriAI/litellm/pull/33796)
- feat(proxy): add x-litellm-model-name response header with deployment model string by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33698](https://github.com/BerriAI/litellm/pull/33698)
- feat: add Straiker guardrail integration by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33781](https://github.com/BerriAI/litellm/pull/33781)
- fix(vertex\_ai): exclude Gemini Google Search grounding tokens from input token billing by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33742](https://github.com/BerriAI/litellm/pull/33742)
- feat(fireworks\_ai): map litellm session id to x-session-affinity header for prompt caching by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33717](https://github.com/BerriAI/litellm/pull/33717)
- feat(ui): configure Anthropic automatic prompt caching from the Admin UI by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33581](https://github.com/BerriAI/litellm/pull/33581)
- fix(router): enforce context-window pre-call checks for Responses API input by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33706](https://github.com/BerriAI/litellm/pull/33706)
- fix(otel): restore proxy-level error.\* attributes on v2 failure spans (LIT-4179) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33664](https://github.com/BerriAI/litellm/pull/33664)
- fix(mcp): persist config.yaml DCR clients in a server-scoped store so refresh survives token expiry by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33768](https://github.com/BerriAI/litellm/pull/33768)
- refactor(ui): consolidate Add/Edit credential modals into one CredentialModal by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32572](https://github.com/BerriAI/litellm/pull/32572)
- feat(mcp): add ID-JAG (identity assertion authorization grant) support for MCP egress by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;31516](https://github.com/BerriAI/litellm/pull/31516)
- refactor(ui): migrate policy attachments table onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33827](https://github.com/BerriAI/litellm/pull/33827)
- docs(litellm-rust): add provider coding standards by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33833](https://github.com/BerriAI/litellm/pull/33833)
- test(e2e): rename Gateway to ProxyClient and expose it as a session-scoped pytest fixture by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33750](https://github.com/BerriAI/litellm/pull/33750)
- feat(messages): route Azure Anthropic /messages through Rust behind rust:true by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33616](https://github.com/BerriAI/litellm/pull/33616)
- test(e2e): add Locust throughput load test that runs last by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33748](https://github.com/BerriAI/litellm/pull/33748)
- fix(proxy): resolve team wildcard credentials for vector store files by [@&#8203;shivamrawat1](https://github.com/shivamrawat1) in [#&#8203;33649](https://github.com/BerriAI/litellm/pull/33649)
- refactor(e2e): fold claude\_code HTTP probes onto shared ProxyClient methods by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33760](https://github.com/BerriAI/litellm/pull/33760)
- test(e2e): harden stage flakes for batches, UI, and MCP by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33831](https://github.com/BerriAI/litellm/pull/33831)
- fix(e2e): migrate load suite from e2e\_gateway to ProxyClient by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33839](https://github.com/BerriAI/litellm/pull/33839)
- chore(e2e): remove tests/e2e/docker-compose.yml by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33837](https://github.com/BerriAI/litellm/pull/33837)
- test(e2e): cover /v1/responses openai basic nonstream and stream by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33830](https://github.com/BerriAI/litellm/pull/33830)
- test(e2e): cover /v1/responses openai cost\_logged and tool\_use by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33835](https://github.com/BerriAI/litellm/pull/33835)
- test(e2e): cover /v1/responses OpenAI vision and Anthropic basic by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33838](https://github.com/BerriAI/litellm/pull/33838)
- test(e2e): spendlog cost for streaming /v1/messages via responses bridge by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33753](https://github.com/BerriAI/litellm/pull/33753)
- fix(docker): bake prisma CLI and engines at a fixed path so fresh-DB migrations work for any uid offline by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33853](https://github.com/BerriAI/litellm/pull/33853)
- feat(chat-ui): add personal Logs view scoped to the current user by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33829](https://github.com/BerriAI/litellm/pull/33829)
- chore: bump litellm-proxy-extras 0.4.78 -> 0.4.79 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33855](https://github.com/BerriAI/litellm/pull/33855)
- docs(litellm-rust): require the official Rust Style Guide in agent rules by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33867](https://github.com/BerriAI/litellm/pull/33867)
- fix(router): treat malformed configured token limits as absent on /v1/models by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33864](https://github.com/BerriAI/litellm/pull/33864)
- chore: rebuild admin UI bundle for the rc release by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33857](https://github.com/BerriAI/litellm/pull/33857)
- docs(rust): add provider abstraction standards by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33865](https://github.com/BerriAI/litellm/pull/33865)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33868](https://github.com/BerriAI/litellm/pull/33868)
- chore(release): backport [#&#8203;33929](https://github.com/BerriAI/litellm/issues/33929) to rc/1.94.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34033](https://github.com/BerriAI/litellm/pull/34033)
- chore(release): backport [#&#8203;33810](https://github.com/BerriAI/litellm/issues/33810), [#&#8203;33733](https://github.com/BerriAI/litellm/issues/33733) to rc/1.94.0 and bump litellm-proxy-extras to 0.4.79.post1 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34215](https://github.com/BerriAI/litellm/pull/34215)
- chore(ui): rebuild Next.js bundle on rc/1.94.0 so the Cost Optimization page ships by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34216](https://github.com/BerriAI/litellm/pull/34216)
- chore(release): backport auth, CLI SSO and guardrail fixes to rc/1.94.0 and refresh flagged dependencies by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34640](https://github.com/BerriAI/litellm/pull/34640)
- chore(release): backport [#&#8203;33899](https://github.com/BerriAI/litellm/issues/33899), [#&#8203;33978](https://github.com/BerriAI/litellm/issues/33978), [#&#8203;34582](https://github.com/BerriAI/litellm/issues/34582), [#&#8203;34675](https://github.com/BerriAI/litellm/issues/34675) to rc/1.94.0 and bump litellm-proxy-extras to 0.4.79.post2 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34855](https://github.com/BerriAI/litellm/pull/34855)
- fix(ui): backport cache leakage card layout fix to rc/1.94.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34964](https://github.com/BerriAI/litellm/pull/34964)
- fix(ui): add missing cost-optimization page description on rc/1.94.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34967](https://github.com/BerriAI/litellm/pull/34967)
- feat(ui): mark Cost Optimization as beta in the left nav ([#&#8203;34984](https://github.com/BerriAI/litellm/issues/34984)) by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34987](https://github.com/BerriAI/litellm/pull/34987)
- chore: rebuild Admin UI bundle for v1.94.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34982](https://github.com/BerriAI/litellm/pull/34982)
- fix(cost-optimization): backport the savings chart axis fix and methodology popovers to rc/1.94.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34994](https://github.com/BerriAI/litellm/pull/34994)
- chore: rebuild Admin UI bundle for rc/1.94.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34995](https://github.com/BerriAI/litellm/pull/34995)

**Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.93.0...v1.94.0>

### [`v1.94.0`](https://github.com/BerriAI/litellm/releases/tag/v1.94.0)

##### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.94.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.94.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.94.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

##### What's Changed

- feat(ui): working Test Connection for the complexity auto router by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;32950](https://github.com/BerriAI/litellm/pull/32950)
- fix(xecguard): use StandardLoggingGuardrailInformation in logging hook by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;32911](https://github.com/BerriAI/litellm/pull/32911)
- feat(ui): adopt openapi-react-query ($api) and convert useCustomers by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32949](https://github.com/BerriAI/litellm/pull/32949)
- refactor(ui): colocate the mcp-servers view, keeping the shared mcp\_tools surface by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32968](https://github.com/BerriAI/litellm/pull/32968)
- refactor(ui): convert endpoint usage charts to shadcn/recharts by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32723](https://github.com/BerriAI/litellm/pull/32723)
- fix(proxy-auth): stop unrecognized model namespaces slipping through provider wildcard keys by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32979](https://github.com/BerriAI/litellm/pull/32979)
- feat(router): random-pick multi-model complexity tiers by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32967](https://github.com/BerriAI/litellm/pull/32967)
- fix(xecguard): sanitize scan result before recording it for logging by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;32935](https://github.com/BerriAI/litellm/pull/32935)
- fix(auto\_router): filter embedding models in complexity tab dropdowns, require all tiers, inline validation by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;32978](https://github.com/BerriAI/litellm/pull/32978)
- fix(anthropic): translate raw adaptive thinking for pre-4.6 models on chat completions and Bedrock Converse by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;32944](https://github.com/BerriAI/litellm/pull/32944)
- feat(router): add Router(plugins=\[...]) routing-plugin pipeline by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32972](https://github.com/BerriAI/litellm/pull/32972)
- feat(router): soft-floor adaptive mode for complexity router by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32947](https://github.com/BerriAI/litellm/pull/32947)
- docs(github): add QA runbook section to the PR template by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32965](https://github.com/BerriAI/litellm/pull/32965)
- fix(model\_cost): add supports\_reasoning: false to Gemini image generation models by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32836](https://github.com/BerriAI/litellm/pull/32836)
- build(dev-env): add make bootstrap and unprovisioned-checkout preflight to pre-commit by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32981](https://github.com/BerriAI/litellm/pull/32981)
- ci(ui): report only error-level knip findings in CI by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32971](https://github.com/BerriAI/litellm/pull/32971)
- feat(batches): track cost for unmanaged Bedrock batches, generalize the flag by [@&#8203;Sameerlite](https://github.com/Sameerlite) in [#&#8203;32315](https://github.com/BerriAI/litellm/pull/32315)
- fix(guardrails): walk custom\_tool\_call\_output items in \_content\_utils by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;32969](https://github.com/BerriAI/litellm/pull/32969)
- fix: show and allow editing team model aliases after team creation by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33047](https://github.com/BerriAI/litellm/pull/33047)
- chore(deps): bump pillow to 12.3.0 to resolve osv-scan CVEs by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33093](https://github.com/BerriAI/litellm/pull/33093)
- feat(mcp): mint gateway-bound envelope at the token endpoint for dcr\_bridge oauth\_delegate by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32828](https://github.com/BerriAI/litellm/pull/32828)
- fix(mcp): surface rejected delegate-auth upstream tokens as connect-time 401 by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32741](https://github.com/BerriAI/litellm/pull/32741)
- fix(proxy): track unauthenticated pass-through requests in spend logs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;32410](https://github.com/BerriAI/litellm/pull/32410)
- feat(lasso): send source.type for Used By attribution by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33090](https://github.com/BerriAI/litellm/pull/33090)
- fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@&#8203;thibault-linktree](https://github.com/thibault-linktree) in [#&#8203;33025](https://github.com/BerriAI/litellm/pull/33025)
- fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33099](https://github.com/BerriAI/litellm/pull/33099)
- fix(completion): forward aws credential kwargs into litellm\_params so the responses bridge keeps WIF auth by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32956](https://github.com/BerriAI/litellm/pull/32956)
- fix(ui): respect litellm\_key\_header\_name in BYOK credential save and workflow runs fetches by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33103](https://github.com/BerriAI/litellm/pull/33103)
- refactor(ui): standardize debounce waits behind shared DEBOUNCE\_WAIT\_MS constant by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33040](https://github.com/BerriAI/litellm/pull/33040)
- feat(ui): rebuild the Virtual Keys table on the shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;32991](https://github.com/BerriAI/litellm/pull/32991)
- fix: redact async complete streaming response for custom callbacks by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33106](https://github.com/BerriAI/litellm/pull/33106)
- build(ui): bump [@&#8203;tanstack/react-pacer](https://github.com/tanstack/react-pacer) from 0.2.0 to 0.22.1 by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33041](https://github.com/BerriAI/litellm/pull/33041)
- fix(ui): address Virtual Keys redesign review nits by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33112](https://github.com/BerriAI/litellm/pull/33112)
- fix(openai/responses): clamp max\_output\_tokens below API minimum by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33098](https://github.com/BerriAI/litellm/pull/33098)
- fix(prometheus): read v3 rate limiter remaining values for per-key model gauges by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33119](https://github.com/BerriAI/litellm/pull/33119)
- fix(ui): drop w-full from page-content wrappers to remove 32px horizontal overflow by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33118](https://github.com/BerriAI/litellm/pull/33118)
- refactor(ui): migrate straightforward value debounces to react-pacer by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33042](https://github.com/BerriAI/litellm/pull/33042)
- feat(mcp): interactive SSO sign-in for dcr\_bridge oauth\_delegate DCR clients by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32946](https://github.com/BerriAI/litellm/pull/32946)
- test(proxy): add regression tests for management\_endpoints edge cases by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;32976](https://github.com/BerriAI/litellm/pull/32976)
- fix(auto-router): correct Responses API tool\_choice shape and propagate alias litellm\_params by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32974](https://github.com/BerriAI/litellm/pull/32974)
- fix(ui): render the sidebar scrollbar with shadcn ScrollArea by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33124](https://github.com/BerriAI/litellm/pull/33124)
- refactor(ui): migrate callback debounce sites to react-pacer with regression tests by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33043](https://github.com/BerriAI/litellm/pull/33043)
- chore: add CODEOWNERS for ui and proxy UI build artifacts by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33131](https://github.com/BerriAI/litellm/pull/33131)
- feat(mcp): client-held refresh envelope for the dcr\_bridge oauth\_delegate flow by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32980](https://github.com/BerriAI/litellm/pull/32980)
- feat(ui): rebuild the Teams table on the shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33128](https://github.com/BerriAI/litellm/pull/33128)
- fix(mcp): relay upstream OAuth token and DCR rejections instead of a generic 500 by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33113](https://github.com/BerriAI/litellm/pull/33113)
- fix(keys): persist key\_type so the UI shows correct key scope instead of "All Proxy…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants