Skip to content

chore(ci): promote internal staging to main - #33308

Merged
yuneng-berri merged 143 commits into
mainfrom
litellm_internal_staging
Jul 15, 2026
Merged

chore(ci): promote internal staging to main#33308
yuneng-berri merged 143 commits into
mainfrom
litellm_internal_staging

Conversation

@yuneng-berri

Copy link
Copy Markdown
Contributor

Relevant issues

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • My PR passes all CI/CD checks (e.g., lint, format, unit tests)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Type

🆕 New Feature
🐛 Bug Fix
🧹 Refactoring
📖 Documentation
🚄 Infrastructure
✅ Test

Changes

QA runbook

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

tin-berri and others added 30 commits July 10, 2026 12:11
…ime 401

For MCP servers with auth_type=oauth2 + delegate_auth_to_upstream=true, a
client-supplied upstream token that the upstream rejects was masked: the
upstream 401 raised during tools/list is absorbed by the list handler, so on a
single-server route a rejected token became HTTP 200 with an empty tool list.
Clients showed "0 tools" instead of re-authenticating, and monitoring never saw
an unauthorized signal.

Extend the connect-time preflight _check_passthrough_upstream_auth to probe
delegate-auth servers with the caller's bare Authorization bearer, reusing the
existing _probe_upstream_auth and the RFC 6750 challenge builder, so a rejected
token fails the connect with 401 + WWW-Authenticate error="invalid_token" and a
compliant client re-runs the upstream OAuth flow.

The bare Authorization header is a valid upstream token only when admission took
the delegate bypass, so the delegate target is resolved through
get_mcp_server_by_name (the same resolver admission uses) rather than the wider
allowed-server prefix/access-group matching. A name that reaches a delegate
server only via server_id or an access group is admitted as a real LiteLLM key,
so probing it would leak that key upstream; requiring the admission-resolver
match closes that gap. The probe is gated to single-server routes (matching the
OBO preflight), keyed to the caller's authorized set by server_id, and the
challenge echoes the requested name so aliased routes get the same
resource_metadata URL as the tokenless preemptive challenge. Tokenless requests
keep flowing to the preemptive discovery challenge unchanged.

Resolves LIT-4194
…tion models

vertex_ai/gemini-2.5-flash-image, vertex_ai/gemini-3-pro-image-preview,
vertex_ai/gemini-3.1-flash-image-preview, gemini/gemini-3-pro-image-preview,
and gemini/gemini-3.1-flash-image-preview were missing supports_reasoning
entries; _supports_factory then fell through to the vertex_ai provider-level
config which returns true, causing requests with reasoning_effort to be sent
to an API that rejects them.
The backup file is used by tests; the root model_prices_and_context_window.json
is what gets published to the pricing URL and loaded by the proxy at runtime.
Without this, the proxy would continue resolving supports_reasoning via the
provider-level fallback and returning true for Gemini image generation models.

Also covers vertex_ai/gemini-3-pro-image and vertex_ai/gemini-3.1-flash-image
(non-preview variants) and gemini/gemini-3.1-flash-image which exist only in
the root JSON.
gemini/gemini-3.1-flash-image, vertex_ai/gemini-3-pro-image, and
vertex_ai/gemini-3.1-flash-image existed in the root pricing JSON but not in
litellm/model_prices_and_context_window_backup.json, leaving deployments with
LITELLM_LOCAL_MODEL_COST_MAP=True unprotected. Copies the root entries into
the backup verbatim and extends the regression test to cover all ten gemini
image models, asserting each exists in the local cost map so a missing backup
entry fails the test instead of passing vacuously
…on pre-4.6 models

Clients that pass thinking={"type": "adaptive"} directly (not via the
reasoning_effort alias) on the /chat/completions interface had it forwarded
unmodified to pre-4.6 Anthropic models, which reject the shape. Mirrors the
translation already applied on the native /v1/messages passthrough (#32867):
translate to legacy thinking={type: enabled, budget_tokens}, capped below
max_tokens, dropping thinking when max_tokens can't fit even the minimum
budget. Hoists the shared budget-capping helper onto AnthropicConfig so both
paths use one implementation.
Follow-up to #32867 (native /v1/messages) and the /chat/completions
commit earlier on this branch, extending the same adaptive-thinking
translation to the Bedrock Converse path.

Clients like Claude Code send thinking={type: "adaptive"} on every
request. When routed via Bedrock Converse to pre-4.6 models
(claude-haiku-4-5, claude-sonnet-4-5), this was forwarded as-is and
rejected by the model. Mirrors the translation already applied on the
/chat/completions and /v1/messages paths: map to legacy
thinking={type: enabled, budget_tokens}, capped below max_tokens.

Also fixes the missing custom_llm_provider arg in the chat completions
path's call to AnthropicConfig._map_reasoning_effort.
The rebase onto staging changed _is_adaptive_thinking_model to require
custom_llm_provider (no default), so the one-arg call in the raw adaptive
thinking branch raised TypeError at runtime for any /chat/completions
caller sending thinking={type: adaptive}. Use self._resolved_provider,
matching the reasoning_effort branch just below. Caught by Greptile.
…too small

Adds the regression test for the warning-drop branch in the Converse
adaptive-thinking translation, mirroring the chat completions path's
test_raw_adaptive_thinking_dropped_when_max_tokens_too_small.
The consolidated auto-router tab dropped the Test Connection button because
the shared prepareModelAddRequest helper returns an empty array for an auto
router (it has no model_mappings), so the caller crashed destructuring
result[0].litellmParamsObj. That is the crash in #31590 and the open PR
#31794. #31794 only silenced the crash by pointing the test at
auto_router/complexity_router, which is not a provider model, so the
/health/test_connection health check (a real litellm.ahealth_check
completion) would still error.

Bring the button back and make it meaningful: an auto router dispatches to
saved model groups, so Test Connection now probes those directly. It builds
a deduped target list from the configured tiers (tiers sharing a model group
collapse to one probe) plus the embedding model when semantic keyword
matching is on, then runs a live /health/test_connection against each and
shows per-target pass/fail. This never touches prepareModelAddRequest, so the
original destructure crash cannot recur.

Scope is the recommended complexity router only; the to-be-deprecated
semantic router is untouched. No backend changes.

Supersedes #31794. Resolves #31590.
…test_connection

Live testing showed the first cut was broken: /health/test_connection merges
{...configParams, ...requestParams}, so passing the public model_group name as
the request model overrode the resolved provider model and every tier failed
with "LLM Provider NOT provided". The frontend only has the public group name,
not the underlying litellm_params, so it cannot build the request that endpoint
needs.

Switch to testing each model group the way production actually routes it: send a
minimal request to /v1/chat/completions (or /v1/embeddings for the embedding
model) by public group name through the shared apiClient. The router resolves
the group, credentials, and provider itself, so a green row means the tier is
genuinely reachable. Verified live: voyage embedding returns 200, a tier with a
bad key returns the real provider auth error.

Also address Greptile feedback: rows now update progressively as each probe
settles instead of all at once, and TIER_ORDER is derived through a
`satisfies Record<keyof ComplexityTiers, null>` guard so adding a tier without
listing it is a compile error.
… the responses bridge keeps WIF auth

Chat-completions requests to responses-only Bedrock Mantle models are
bridged to the Responses API, but completion() forwarded only
aws_bedrock_project_id into get_litellm_params, so aws_role_name,
aws_web_identity_token, aws_session_name and the other SigV4 credential
kwargs never reached sign_request and botocore fell back to the default
credential chain ("Bedrock Mantle auth failed: no Bearer token and no
usable AWS credentials"). Forward the whole AWS credential kwarg family,
extracted from the OPTIONAL_KWARGS_KEYS set get_litellm_params already
supports.
max_tokens=1 makes reasoning models (o1/o3/...) return a 400 "max_tokens
reached" because reasoning tokens count against the cap, so a reachable
reasoning tier showed a false failure in Test Connection. Live-verified: o3
400s with the cap and succeeds without it.

Extract the request shape into a pure buildModelGroupTestRequest and cover it
with a test asserting the chat body carries no max_tokens (or
max_completion_tokens), so this regression is caught in unit tests instead of
only against a live reasoning model.
…ction

feat(ui): working Test Connection for the complexity auto router
#32911)

XecGuard's async_logging_hook wrote a bare dict to
standard_logging_object["guardrail_information"] while the typed
contract is Optional[List[StandardLoggingGuardrailInformation]].
Readers that iterated the field walked dict keys, raised on
info.get, or silently dropped the entry from guardrail usage
tracking and spend-log writes

Construct the typed entry and append it to the existing list or
create a new one, matching the shared helper pattern. Record the
configured guardrail name instead of a hardcoded "xecguard" and
pass the GuardrailEventHooks enum for guardrail_mode
…32949)

* feat(ui): adopt openapi-react-query and convert useCustomers to $api

Add openapi-react-query and expose $api = createQueryClient(fetchClient)
alongside fetchClient. Rewrite useCustomers as
$api.useQuery("get", "/customer/list", {}, { enabled, select }), which
derives the query key from method + path (dropping the hand-written
createQueryKeys entry and the manual key) and forwards the request signal
for cancellation. The response type still flows from schema.d.ts as
CustomerResponse[]. Tests assert the path, the admin/token enabled gate,
and the empty-body select fallback.

* test(ui): read the last render's options in useCustomers helper

The lastCallOptions helper was named for the last call but read
mock.calls[0]. Harmless while each test renders once, but it would
silently assert against first-render options if a test ever re-renders.
Read the final call instead.
…ools surface (#32968)

* refactor(ui): colocate the usage view, keeping the shared usage components

Split for the usage (UsagePage) segment. Most of the folder is the usage page's
own view, but four pieces are reused elsewhere and stay in @/components/UsagePage:
TopKeyView (old-usage), KeyModelUsageView and value_formatters (activity_metrics),
and the shared types (activity_metrics, chartUtils). The other 21 files move into
usage/_components, preserving the folder structure.

The external consumers import only the retained files, so they are untouched. The
moved files' imports of the retained files become @/components/UsagePage paths,
other escaping relative imports are absolutized, and lint suppressions are re-keyed
for moved files only. No behavior change.

* refactor(ui): colocate the mcp-servers view, keeping the shared mcp_tools surface
krrish-berri-2 and others added 21 commits July 13, 2026 21:29
…ow.py

discoverable_endpoints.py had grown to 2695 lines mixing FastAPI route handlers with the dcr_bridge token-flow logic, against the no-monster-files convention. This moves the bridge token flow (the litellm-key/user resolution, the SCIM revalidation gate, and the mint/refresh envelope logic with their types and error mappers) into a dedicated bridge_token_flow.py, leaving the route handlers and the shared exchange_token_with_server orchestrator in discoverable_endpoints.py importing from it

Pure relocation, zero behavior change. The moved code is byte-verbatim except one type annotation quoted as a forward reference (_BridgeAuthorizationCode is used only for typing and imported under TYPE_CHECKING to avoid a cycle), and the new module imports nothing from discoverable_endpoints at runtime. 275 tests pass unchanged; the test patch targets for moved internals were repointed to the new module and verified to still apply
refactor(mcp): extract the dcr_bridge token flow into bridge_token_flow.py
Raise the constraint floors for two transitive dependencies so resolution moves them to their latest maintenance releases: httplib2 0.31.2 -> 0.32.0 and setuptools 82.0.1 -> 83.0.0. Both are pulled in only by optional integrations (Google API client, grpc tooling, lunary observability, the nvidia-riva extra), all lower-bound only, so the floors stay inside every requirer's allowed range and a default install is unaffected
Move the Create New Key and Create Team buttons out of the page header's
right-side action slot. On Teams the button now sits in the tab bar's left
slot, separated from the three tabs by a vertical rule, so the CTA and tabs
read as one left-anchored cluster. On Keys, which has no tabs, the button
anchors left on its own row beneath the title.
…ading adaptive thinking for pre-4.6 models (#33244)

* fix(anthropic/passthrough): drop temperature and cap thinking budget when downgrading adaptive thinking for pre-4.6 models

* test(anthropic/passthrough): use sufficient max_tokens for reasoning_effort thinking mapping

* fix(anthropic/passthrough): drop incompatible temperature when downgrading adaptive thinking for pre-4.6 models

Narrow the fix to the temperature reconciliation; the reasoning_effort
budget cap is reverted because the live translation grid relies on
budget_tokens >= max_tokens to reject unsupported effort tiers
(xhigh/max) on budget-mode models, so capping turned those 400s into
200s.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…rails at deployment hook (#33136)

* fix(guardrails): run apply_guardrail-style model-level pre_call guardrails at deployment hook

* fix(guardrails): keep request-body dispatch predicate unchanged

* fix(guardrails): fail closed when proxy extras are missing at deployment hook
…n) with UI opt-out (#32005)

* fix: enforce user budget on team keys

User budget was skipped when the key belonged to a team, letting
users exceed their personal budget by going through a team key.

Remove the team_object guard in _user_max_budget_check so user
budgets are always enforced. Add skip_user_budget_on_team_key
general_settings flag to opt back into the old behavior.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: update test to expect user budget enforcement on team keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): enforce user budget on team keys in reservation path and expose skip flag in UI

Extends the read-time fix so the optimistic budget reservation also reserves the user spend counter for team-scoped keys, register skip_user_budget_on_team_key in ConfigGeneralSettings so /config/field/update accepts it, and surface it as a Boolean toggle on the Admin UI General Settings table via allowed_args in /config/list.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: assert budget_exceeded ProxyException in personal budget test

Tighten the broad pytest.raises(Exception) so the test only passes when
the auth flow rejects with a budget_exceeded ProxyException, and switch
the new ConfigGeneralSettings field to Optional[bool] to match the
surrounding annotation style

* fix: revert to bool | None to stay under UP045 strict budget

---------

Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
The rate-limited batch spend test snapshotted unattributed rows via the
unpaginated /spend/logs whole-table read, which grows with the environment
(58MB on stage) and OOMKilled the e2e runner at its 512Mi limit on every
scheduled run. Gateway.spend_logs_window pages /spend/logs/v2 over an
explicit date window instead, and SpendLogsParams now rejects a filterless
read so the whole-table call cannot come back
…erage

test(e2e): cover key rpm/tpm rate limiting, window reset, and pacing headers
* fix(anthropic): route native structured output

Use model capability metadata so new native structured-output models do not require transformation allowlist changes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(anthropic): pass provider to capability

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(anthropic): cover dotted model IDs

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(anthropic): handle remote capability lag

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
fix(ci): retry setup-uv installs to survive transient manifest fetch failures
…33268)

* fix(proxy): never log raw virtual keys in key insertion debug output

* fix(proxy): tolerate None token in insert_data debug log redaction
…iew_pricing

feat(pricing): add gemini-omni-flash-preview with video output token pricing
With enable_jwt_auth enabled but no enterprise license (premium_user
False), the JWT premium check fired on every request before the token
was inspected, so the master key, sk- virtual keys, and the encrypted
CLI/UI SSO session token that `lite login` issues all 401'd with "JWT
Auth is an enterprise only feature" and were never decoded. That broke
`lite login`, `lite claude`, and the proxy master key on any deployment
that turned JWT auth on without a license.

Move the premium check inside the is_jwt branch so it gates only real
JWTs. Non-JWT credentials fall through to their own auth paths
regardless of license; actual JWTs still require premium, so the
enterprise gate is unchanged for the feature it protects.
@greptile-apps

greptile-apps Bot commented Jul 15, 2026

Copy link
Copy Markdown
Contributor

Too many files changed for review. (323 files found, 100 file limit)

Bypass the limit by tagging @greptile-apps to review.

@CLAassistant

CLAassistant commented Jul 15, 2026

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
6 out of 8 committers have signed the CLA.

✅ tin-berri
✅ yuneng-berri
✅ ryan-crabbe-berri
✅ mubashir1osmani
✅ mateo-berri
✅ yucheng-berri
❌ devin-ai-integration[bot]
❌ krrish-berri-2
You have signed the CLA already but the status is still pending? Let us recheck it.

* feat(ui): migrate guardrails table onto shared DataTable

Move the guardrails list onto the shared DataTable + cell library as the
proof-of-concept for the simple-tables design migration, following the Teams
reference pattern.

Split the table into a thin container (guardrail_table.tsx) and column defs
(guardrailTableColumns.tsx): client-side sort defaulting to created_at desc, a
search + refresh toolbar, IdCell / DateCell / StatusBadge cells, real provider
logos, a rich empty state, and skeleton loading rows. Row actions move into a
per-row overflow menu; deletion stays disabled for config-file guardrails, now
surfaced as a disabled menu item instead of a greyed trash icon. Detail view
and the delete modal remain owned by GuardrailsPanel.

Restyle the "Add New Guardrail" control to the shared Button + dropdown menu.

Update the regression tests for the menu-based actions and drop the now-stale
eslint suppression entry that the rewrite eliminated.

* fix(ui): match guardrails table to the design

Address design-review feedback on the guardrails migration:

- Drop the search + refresh toolbar. The original table had neither and the
  SimpleTable design has no toolbar; the container now just renders the sorted
  table and its empty state.
- Give the Guardrail ID cell the design's hover affordance by rendering it with
  the shared IdentityCell (monospace, chevron on hover) instead of the blue
  IdCell pill.
- Stop pinning the actions column. Pinning added a sticky divider that the
  design and the Teams table don't have; it is now a plain right-aligned menu
  column, matching Teams.

* fix(ui): match loading skeleton row height to loaded rows

The compact skeleton row did not carry the h-8 height that real compact
rows get, so loading rows rendered shorter than loaded ones and the table
height jumped when data arrived. Mirror the same size-based height on the
skeleton row in the shared DataTable so every compact table loads at a
stable height

* test(ui): drop stale onGuardrailUpdated from guardrails table baseProps

The prop was removed from GuardrailTableProps when the toolbar went away;
the test baseProps still listed it. Harmless at the call site since it is
spread rather than an object literal, but dead and worth removing

* fix(ui): remove dead edit_guardrail_form after guardrails migration

The guardrails table migration dropped the last import of EditGuardrailForm,
which knip flags as an unused file. The form was already unreachable before
the migration: the table wired a delete button only, and nothing ever called
handleEditClick to open the modal, so the import was the sole thing keeping
the file referenced. Delete it and prune its now-stale eslint suppression
entry. Guardrail editing is unchanged and lives in the detail view
(GuardrailInfoView)
@yuneng-berri
yuneng-berri merged commit 6a797f9 into main Jul 15, 2026
62 of 64 checks passed
@codspeed-hq

codspeed-hq Bot commented Jul 15, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_internal_staging (f8c49f5) with main (6a797f9)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (f8c49f5) during the generation of this report, so 6a797f9 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report.

doonga pushed a commit to greyrock-labs/home-ops that referenced this pull request Jul 29, 2026
…4.0) (#201)

This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.93.0` → `v1.94.0` |

---

### Release Notes

<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>

### [`v1.94.0`](https://github.com/BerriAI/litellm/releases/tag/v1.94.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.94.0...v1.94.0)

##### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.94.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.94.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.94.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

##### What's Changed

- feat(ui): working Test Connection for the complexity auto router by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;32950](https://github.com/BerriAI/litellm/pull/32950)
- fix(xecguard): use StandardLoggingGuardrailInformation in logging hook by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;32911](https://github.com/BerriAI/litellm/pull/32911)
- feat(ui): adopt openapi-react-query ($api) and convert useCustomers by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32949](https://github.com/BerriAI/litellm/pull/32949)
- refactor(ui): colocate the mcp-servers view, keeping the shared mcp\_tools surface by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32968](https://github.com/BerriAI/litellm/pull/32968)
- refactor(ui): convert endpoint usage charts to shadcn/recharts by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32723](https://github.com/BerriAI/litellm/pull/32723)
- fix(proxy-auth): stop unrecognized model namespaces slipping through provider wildcard keys by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32979](https://github.com/BerriAI/litellm/pull/32979)
- feat(router): random-pick multi-model complexity tiers by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32967](https://github.com/BerriAI/litellm/pull/32967)
- fix(xecguard): sanitize scan result before recording it for logging by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;32935](https://github.com/BerriAI/litellm/pull/32935)
- fix(auto\_router): filter embedding models in complexity tab dropdowns, require all tiers, inline validation by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;32978](https://github.com/BerriAI/litellm/pull/32978)
- fix(anthropic): translate raw adaptive thinking for pre-4.6 models on chat completions and Bedrock Converse by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;32944](https://github.com/BerriAI/litellm/pull/32944)
- feat(router): add Router(plugins=\[...]) routing-plugin pipeline by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32972](https://github.com/BerriAI/litellm/pull/32972)
- feat(router): soft-floor adaptive mode for complexity router by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32947](https://github.com/BerriAI/litellm/pull/32947)
- docs(github): add QA runbook section to the PR template by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32965](https://github.com/BerriAI/litellm/pull/32965)
- fix(model\_cost): add supports\_reasoning: false to Gemini image generation models by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32836](https://github.com/BerriAI/litellm/pull/32836)
- build(dev-env): add make bootstrap and unprovisioned-checkout preflight to pre-commit by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32981](https://github.com/BerriAI/litellm/pull/32981)
- ci(ui): report only error-level knip findings in CI by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32971](https://github.com/BerriAI/litellm/pull/32971)
- feat(batches): track cost for unmanaged Bedrock batches, generalize the flag by [@&#8203;Sameerlite](https://github.com/Sameerlite) in [#&#8203;32315](https://github.com/BerriAI/litellm/pull/32315)
- fix(guardrails): walk custom\_tool\_call\_output items in \_content\_utils by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;32969](https://github.com/BerriAI/litellm/pull/32969)
- fix: show and allow editing team model aliases after team creation by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33047](https://github.com/BerriAI/litellm/pull/33047)
- chore(deps): bump pillow to 12.3.0 to resolve osv-scan CVEs by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33093](https://github.com/BerriAI/litellm/pull/33093)
- feat(mcp): mint gateway-bound envelope at the token endpoint for dcr\_bridge oauth\_delegate by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32828](https://github.com/BerriAI/litellm/pull/32828)
- fix(mcp): surface rejected delegate-auth upstream tokens as connect-time 401 by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32741](https://github.com/BerriAI/litellm/pull/32741)
- fix(proxy): track unauthenticated pass-through requests in spend logs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;32410](https://github.com/BerriAI/litellm/pull/32410)
- feat(lasso): send source.type for Used By attribution by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33090](https://github.com/BerriAI/litellm/pull/33090)
- fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@&#8203;thibault-linktree](https://github.com/thibault-linktree) in [#&#8203;33025](https://github.com/BerriAI/litellm/pull/33025)
- fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33099](https://github.com/BerriAI/litellm/pull/33099)
- fix(completion): forward aws credential kwargs into litellm\_params so the responses bridge keeps WIF auth by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32956](https://github.com/BerriAI/litellm/pull/32956)
- fix(ui): respect litellm\_key\_header\_name in BYOK credential save and workflow runs fetches by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33103](https://github.com/BerriAI/litellm/pull/33103)
- refactor(ui): standardize debounce waits behind shared DEBOUNCE\_WAIT\_MS constant by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33040](https://github.com/BerriAI/litellm/pull/33040)
- feat(ui): rebuild the Virtual Keys table on the shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;32991](https://github.com/BerriAI/litellm/pull/32991)
- fix: redact async complete streaming response for custom callbacks by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33106](https://github.com/BerriAI/litellm/pull/33106)
- build(ui): bump [@&#8203;tanstack/react-pacer](https://github.com/tanstack/react-pacer) from 0.2.0 to 0.22.1 by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33041](https://github.com/BerriAI/litellm/pull/33041)
- fix(ui): address Virtual Keys redesign review nits by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33112](https://github.com/BerriAI/litellm/pull/33112)
- fix(openai/responses): clamp max\_output\_tokens below API minimum by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33098](https://github.com/BerriAI/litellm/pull/33098)
- fix(prometheus): read v3 rate limiter remaining values for per-key model gauges by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33119](https://github.com/BerriAI/litellm/pull/33119)
- fix(ui): drop w-full from page-content wrappers to remove 32px horizontal overflow by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33118](https://github.com/BerriAI/litellm/pull/33118)
- refactor(ui): migrate straightforward value debounces to react-pacer by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33042](https://github.com/BerriAI/litellm/pull/33042)
- feat(mcp): interactive SSO sign-in for dcr\_bridge oauth\_delegate DCR clients by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32946](https://github.com/BerriAI/litellm/pull/32946)
- test(proxy): add regression tests for management\_endpoints edge cases by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;32976](https://github.com/BerriAI/litellm/pull/32976)
- fix(auto-router): correct Responses API tool\_choice shape and propagate alias litellm\_params by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32974](https://github.com/BerriAI/litellm/pull/32974)
- fix(ui): render the sidebar scrollbar with shadcn ScrollArea by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33124](https://github.com/BerriAI/litellm/pull/33124)
- refactor(ui): migrate callback debounce sites to react-pacer with regression tests by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33043](https://github.com/BerriAI/litellm/pull/33043)
- chore: add CODEOWNERS for ui and proxy UI build artifacts by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33131](https://github.com/BerriAI/litellm/pull/33131)
- feat(mcp): client-held refresh envelope for the dcr\_bridge oauth\_delegate flow by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32980](https://github.com/BerriAI/litellm/pull/32980)
- feat(ui): rebuild the Teams table on the shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33128](https://github.com/BerriAI/litellm/pull/33128)
- fix(mcp): relay upstream OAuth token and DCR rejections instead of a generic 500 by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33113](https://github.com/BerriAI/litellm/pull/33113)
- fix(keys): persist key\_type so the UI shows correct key scope instead of "All Proxy Models" by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33115](https://github.com/BerriAI/litellm/pull/33115)
- feat(router): opt-in session affinity for complexity router by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;33126](https://github.com/BerriAI/litellm/pull/33126)
- feat(prometheus): expose video duration and image count consumption metrics by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33138](https://github.com/BerriAI/litellm/pull/33138)
- test(e2e): otel trace completeness on /chat/completions by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33132](https://github.com/BerriAI/litellm/pull/33132)
- fix(sso): paginate through all pages when fetching service principal group assignments by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33149](https://github.com/BerriAI/litellm/pull/33149)
- test(e2e): otel trace completeness on /v1/messages by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33133](https://github.com/BerriAI/litellm/pull/33133)
- feat(ui): add adaptive routing settings to Auto-Router v2 by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;33146](https://github.com/BerriAI/litellm/pull/33146)
- refactor(mcp): extract the dcr\_bridge token flow into bridge\_token\_flow\.py by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33141](https://github.com/BerriAI/litellm/pull/33141)
- chore: bump litellm 1.93.0 -> 1.94.0, litellm-enterprise 0.1.49 -> 0.1.50, litellm-proxy-extras 0.4.76 -> 0.4.77 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33229](https://github.com/BerriAI/litellm/pull/33229)
- fix(proxy): route master key to team-scoped models by [@&#8203;kunal2002](https://github.com/kunal2002) in [#&#8203;32926](https://github.com/BerriAI/litellm/pull/32926)
- chore(deps): pin httplib2 and setuptools transitive floors by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33233](https://github.com/BerriAI/litellm/pull/33233)
- feat(ui): left-anchor the Create Key and Create Team CTAs by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33248](https://github.com/BerriAI/litellm/pull/33248)
- fix(anthropic/passthrough): drop incompatible temperature when downgrading adaptive thinking for pre-4.6 models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33244](https://github.com/BerriAI/litellm/pull/33244)
- fix(guardrails): run apply\_guardrail-style model-level pre\_call guardrails at deployment hook by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33136](https://github.com/BerriAI/litellm/pull/33136)
- fix(proxy)!: enforce user budget on team keys (read-time + reservation) with UI opt-out by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;32005](https://github.com/BerriAI/litellm/pull/32005)
- fix(e2e): bound spend-log snapshots to a /spend/logs/v2 window by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33265](https://github.com/BerriAI/litellm/pull/33265)
- test(e2e): cover key rpm/tpm rate limiting, window reset, and pacing headers by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32914](https://github.com/BerriAI/litellm/pull/32914)
- fix(anthropic): use native output capability by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;33235](https://github.com/BerriAI/litellm/pull/33235)
- fix(ci): retry setup-uv installs to survive transient manifest fetch failures by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33279](https://github.com/BerriAI/litellm/pull/33279)
- fix(proxy): never log raw virtual keys in key insertion debug output by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33268](https://github.com/BerriAI/litellm/pull/33268)
- fix(bedrock\_mantle): route xai.grok-4.3 via /openai/v1 frontier path by [@&#8203;marty-sullivan](https://github.com/marty-sullivan) in [#&#8203;33027](https://github.com/BerriAI/litellm/pull/33027)
- feat(pricing): add gemini-omni-flash-preview with video output token pricing by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33274](https://github.com/BerriAI/litellm/pull/33274)
- fix(auth): scope the JWT enterprise gate to actual JWTs by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33296](https://github.com/BerriAI/litellm/pull/33296)
- fix(s3): sanitize slashes in response-id-derived object key file name by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33271](https://github.com/BerriAI/litellm/pull/33271)
- refactor(ui): migrate guardrails table onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33303](https://github.com/BerriAI/litellm/pull/33303)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33308](https://github.com/BerriAI/litellm/pull/33308)
- feat(guardrails): streaming text transformation in generic\_guardrail\_api by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33110](https://github.com/BerriAI/litellm/pull/33110)
- test(e2e): cover model-aware mid-conversation system handling on Bedrock Invoke /v1/messages by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32963](https://github.com/BerriAI/litellm/pull/32963)
- test(claude\_code): move the Claude Code compatibility matrix under tests/e2e by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;32548](https://github.com/BerriAI/litellm/pull/32548)
- chore(ci): sync litellm\_internal\_staging into daily OSS branch by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33337](https://github.com/BerriAI/litellm/pull/33337)
- feat(bedrock guardrails): add resource-less InvokeGuardrailChecks (detect-only) mode by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33299](https://github.com/BerriAI/litellm/pull/33299)
- Revert "chore(ci): sync litellm\_internal\_staging into daily OSS branch" by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33339](https://github.com/BerriAI/litellm/pull/33339)
- fix(websearch): intercept web search on the Responses API by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33129](https://github.com/BerriAI/litellm/pull/33129)
- fix(anthropic-adapter): drop empty content\_block\_delta events by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33315](https://github.com/BerriAI/litellm/pull/33315)
- fix(mcp): persist discovered OAuth endpoints and keep last known good on failed re-discovery by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33286](https://github.com/BerriAI/litellm/pull/33286)
- test(e2e): otel trace completeness on streaming chat, messages, and responses (LIT-3787) by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33234](https://github.com/BerriAI/litellm/pull/33234)
- feat(router): resolve auto-router routing plugins from proxy YAML config by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;33251](https://github.com/BerriAI/litellm/pull/33251)
- test(e2e): failed request error span carries the full untruncated message and status (LIT-4179) by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33304](https://github.com/BerriAI/litellm/pull/33304)
- refactor(ui): migrate tags table onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33314](https://github.com/BerriAI/litellm/pull/33314)
- fix(cli): surface actionable CLI SSO errors when CLI and proxy versions skew by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33309](https://github.com/BerriAI/litellm/pull/33309)
- feat(bedrock\_mantle): add GPT-5.6 sol/terra/luna to model cost map by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33412](https://github.com/BerriAI/litellm/pull/33412)
- chore(codeowners): exempt generated schema.d.ts from UI ownership by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33411](https://github.com/BerriAI/litellm/pull/33411)
- feat(proxy): push-based OTLP billable-request metering for enterprise deployments by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;31592](https://github.com/BerriAI/litellm/pull/31592)
- fix(mcp): cap per-user OAuth token cache TTL at the token's own lifetime by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33346](https://github.com/BerriAI/litellm/pull/33346)
- feat(ui): move Caching out of Experimental into Developer Tools by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33432](https://github.com/BerriAI/litellm/pull/33432)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33425](https://github.com/BerriAI/litellm/pull/33425)
- chore(ci): merge daily internal staging branch by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33335](https://github.com/BerriAI/litellm/pull/33335)
- feat(guardrails): add Compresr guardrail for query-aware context compression by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33295](https://github.com/BerriAI/litellm/pull/33295)
- fix(logging): preserve callback order in get\_combined\_callback\_list by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33005](https://github.com/BerriAI/litellm/pull/33005)
- fix(logging): redact assistant tool call arguments in spend logs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33111](https://github.com/BerriAI/litellm/pull/33111)
- fix(anthropic): honor messages request timeout by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33418](https://github.com/BerriAI/litellm/pull/33418)
- fix(llm\_guard): apply sanitized prompt returned by moderation API to request by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33331](https://github.com/BerriAI/litellm/pull/33331)
- fix(logging): stop pinning large request payloads past request end by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33455](https://github.com/BerriAI/litellm/pull/33455)
- feat(guardrails): forward optional metadata on POST /guardrails/apply\_guardrail by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33067](https://github.com/BerriAI/litellm/pull/33067)
- build: raise requires-python cap to <3.15 so Python 3.14 installs current releases by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33438](https://github.com/BerriAI/litellm/pull/33438)
- feat(ui): add reusable BetaBadge and use it for Projects sidebar item by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33449](https://github.com/BerriAI/litellm/pull/33449)
- fix(mcp): discover missing OAuth scopes and token\_url when authorization\_url is set manually by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33317](https://github.com/BerriAI/litellm/pull/33317)
- feat(ui): show exact license expiration date in usage cards by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33478](https://github.com/BerriAI/litellm/pull/33478)
- test(claude\_code): rename misleading REPO\_ROOT to SUITE\_ROOT in test\_v0\_layout by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33472](https://github.com/BerriAI/litellm/pull/33472)
- build(deps): update ddtrace to the 4.x line by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33484](https://github.com/BerriAI/litellm/pull/33484)
- fix(complexity\_router): return empty dict from \_classifier\_call\_metadata when metadata is absent by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33452](https://github.com/BerriAI/litellm/pull/33452)
- fix(ui/chat): resolve chat routes at render time so navigation works under server\_root\_path by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33446](https://github.com/BerriAI/litellm/pull/33446)
- fix(key management): enforce minimum custom key length and mask short keys in key\_name by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33462](https://github.com/BerriAI/litellm/pull/33462)
- chore(ui): remove unmounted UsageIndicator and the Hide Usage Indicator flag by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33482](https://github.com/BerriAI/litellm/pull/33482)
- test(e2e/claude\_code): add passthrough matrix row for the big-3 clouds and Anthropic API by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33473](https://github.com/BerriAI/litellm/pull/33473)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33491](https://github.com/BerriAI/litellm/pull/33491)
- fix(ui): stop sending the complexity-router pseudo-model to /health/test\_connection by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;33498](https://github.com/BerriAI/litellm/pull/33498)
- feat(cli): add lite up/down to ambiently route Claude Code through the proxy by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;33231](https://github.com/BerriAI/litellm/pull/33231)
- feat(complexity\_router): enable session\_affinity by default by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33500](https://github.com/BerriAI/litellm/pull/33500)
- fix(anthropic): stop 500 on combined thinking+signature streaming chunk by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33505](https://github.com/BerriAI/litellm/pull/33505)
- feat(autoroute): prompt for semantic keywords per tier in configure wizard by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33508](https://github.com/BerriAI/litellm/pull/33508)
- fix(cli/anthropic): unblock lite autoroute proxy deps, adaptive thinking, and thinking+signature streaming by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33507](https://github.com/BerriAI/litellm/pull/33507)
- test(ocr): use mistral-document-ai-2512 in azure\_ai OCR tests by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33489](https://github.com/BerriAI/litellm/pull/33489)
- fix(guardrails): show YAML-defined guardrails in the Guardrail Monitor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;32853](https://github.com/BerriAI/litellm/pull/32853)
- refactor(ui): migrate policies, deleted keys, deleted teams, budgets, and search tools tables onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33357](https://github.com/BerriAI/litellm/pull/33357)
- refactor(ui): migrate vector stores, prompts, and skills tables onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33343](https://github.com/BerriAI/litellm/pull/33343)
- test(e2e): datadog log delivery for successful chat, messages, and responses (LIT-4447) by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33415](https://github.com/BerriAI/litellm/pull/33415)
- fix(cli): make CLI output ASCII-only so it doesn't crash legacy Windows consoles by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33465](https://github.com/BerriAI/litellm/pull/33465)
- fix: remove dead user-cache lookup with None key in spend-update path by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33555](https://github.com/BerriAI/litellm/pull/33555)
- feat(helm): add per-component PodDisruptionBudget and topologySpreadConstraints to componentized chart by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33430](https://github.com/BerriAI/litellm/pull/33430)
- fix(e2e/claude\_code): unblock stage collection, align proxy env names, register compat models by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33433](https://github.com/BerriAI/litellm/pull/33433)
- fix(mcp): index authed request-time tools missing from the semantic filter startup index by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33318](https://github.com/BerriAI/litellm/pull/33318)
- fix(streaming): use provider-reported usage cost for OpenRouter streams by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32255](https://github.com/BerriAI/litellm/pull/32255)
- feat(mcp): issuer-anchored OAuth discovery (RFC 8414 §3.3) to close the authorization-server mix-up by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33450](https://github.com/BerriAI/litellm/pull/33450)
- feat(logging): add user and team level spend and budget to StandardLoggingPayload metadata by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33459](https://github.com/BerriAI/litellm/pull/33459)
- fix(router): cast model\_info cost values to float in \_set\_model\_group\_info by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33556](https://github.com/BerriAI/litellm/pull/33556)
- chore(e2e): establish litellm\_e2e\_staging integration line by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33502](https://github.com/BerriAI/litellm/pull/33502)
- fix(ui): navigate to /ui/login/ with trailing slash via hard navigation by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33561](https://github.com/BerriAI/litellm/pull/33561)
- feat(logging): add structured budget fields to budget rejection failure logs by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33460](https://github.com/BerriAI/litellm/pull/33460)
- fix(streaming): surface upstream connection resets instead of empty 200 streams by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33222](https://github.com/BerriAI/litellm/pull/33222)
- fix(proxy\_cli): reap orphaned prisma query-engine processes when a worker dies by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33424](https://github.com/BerriAI/litellm/pull/33424)
- test(reasoning\_effort\_grid): enable azure fable-5 and opus-4-8 grid cells by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33485](https://github.com/BerriAI/litellm/pull/33485)
- build(deps): bump uvicorn lock to 0.51.0 so worker health-check and jitter flags take effect by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33574](https://github.com/BerriAI/litellm/pull/33574)
- fix(proxy): coerce default\_internal\_user\_params.max\_budget to float on config load by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;32434](https://github.com/BerriAI/litellm/pull/32434)
- fix(router): honor per-request routing\_strategy from key/team router\_settings by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33429](https://github.com/BerriAI/litellm/pull/33429)
- fix(redis): honor ssl value instead of key presence when building async connection pool by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;32590](https://github.com/BerriAI/litellm/pull/32590)
- fix(langfuse\_otel): build per-request OTLP exporter from key and team dynamic Langfuse credentials by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;32437](https://github.com/BerriAI/litellm/pull/32437)
- ci: run zizmor and proxy-db unit tests on PRs targeting litellm\_ branches by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33568](https://github.com/BerriAI/litellm/pull/33568)
- fix(router): apply team/key enable\_tag\_filtering to tag routing by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33436](https://github.com/BerriAI/litellm/pull/33436)
- feat(proxy): add disable\_auto\_add\_proxy\_admin\_to\_teams flag by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33563](https://github.com/BerriAI/litellm/pull/33563)
- fix(proxy): stop stale auth cache re-publish to Redis so key updates propagate across replicas by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33565](https://github.com/BerriAI/litellm/pull/33565)
- feat(e2e): emit structured E2E\_RESULT lines for package status history by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33578](https://github.com/BerriAI/litellm/pull/33578)
- chore: bump litellm-enterprise 0.1.50 -> 0.1.51, litellm-proxy-extras 0.4.77 -> 0.4.78 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33571](https://github.com/BerriAI/litellm/pull/33571)
- fix(docker): restore litellm-proxy-extras source dir in runtime images by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33592](https://github.com/BerriAI/litellm/pull/33592)
- feat(ui): require embedding model for semantic auto router by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33313](https://github.com/BerriAI/litellm/pull/33313)
- feat(scim): ingest and round-trip SCIM entitlements and roles user attributes by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33587](https://github.com/BerriAI/litellm/pull/33587)
- refactor(ui): migrate 5 simple tables onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33548](https://github.com/BerriAI/litellm/pull/33548)
- fix(model\_armor): restore reference attachments via skip\_unscannable\_attachments and remove the attachment count cap by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33554](https://github.com/BerriAI/litellm/pull/33554)
- fix(mcp): keep the MCP reference intact when the semantic filter narrows tools by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33584](https://github.com/BerriAI/litellm/pull/33584)
- fix(sso): stop enforcing UI session budget on CLI login tokens by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33312](https://github.com/BerriAI/litellm/pull/33312)
- test: e2e staging leftovers by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33613](https://github.com/BerriAI/litellm/pull/33613)
- test(e2e/claude\_code): add GPT-5.6 Sol/Terra/Luna columns for OpenAI, Azure OpenAI, and Bedrock Mantle by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;33474](https://github.com/BerriAI/litellm/pull/33474)
- test(e2e): otel streaming spans record a real ttft below span duration by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33588](https://github.com/BerriAI/litellm/pull/33588)
- fix(mcp): make the preemptive-401 OAuth challenge decision mode-aware by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33586](https://github.com/BerriAI/litellm/pull/33586)
- fix(ui): show all teams in policy attachment form for admins by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33628](https://github.com/BerriAI/litellm/pull/33628)
- refactor(ui): migrate AI Hub, public hub, and MCP Toolsets tables onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33629](https://github.com/BerriAI/litellm/pull/33629)
- test(e2e): datadog log delivery for streamed routes, read back from the real datadog api by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33566](https://github.com/BerriAI/litellm/pull/33566)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33640](https://github.com/BerriAI/litellm/pull/33640)
- fix(vertex\_ai): surface Gemini grounding toolUsePromptTokenCount in Usage by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33533](https://github.com/BerriAI/litellm/pull/33533)
- test(e2e): harness fixes for stage job green (skips + router/UI/budget) by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33634](https://github.com/BerriAI/litellm/pull/33634)
- fix(router): resolve prompt cache minimum per model instead of a flat 1024 by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33637](https://github.com/BerriAI/litellm/pull/33637)
- fix(logging): classify async anthropic\_messages and generate\_content as async by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33589](https://github.com/BerriAI/litellm/pull/33589)
- fix(ui): remove Chat item from dashboard leftnav by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33647](https://github.com/BerriAI/litellm/pull/33647)
- fix(router): tag-aware pre-routing strategy selection for shared model\_name by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33691](https://github.com/BerriAI/litellm/pull/33691)
- fix(proxy): enforce max\_parallel\_requests as a per-slot concurrency gauge by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;32441](https://github.com/BerriAI/litellm/pull/32441)
- fix(proxy): stop treating upstream model body field as a LiteLLM model on auth-enforced pass-through routes by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33710](https://github.com/BerriAI/litellm/pull/33710)
- fix(mcp): expand toolset grants in shared permission primitives so tools/call honors them by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33612](https://github.com/BerriAI/litellm/pull/33612)
- feat(complexity-router): user-triggered escalation keywords by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33656](https://github.com/BerriAI/litellm/pull/33656)
- fix(fireworks\_ai): bill prompt-cache hits at cache\_read rate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33714](https://github.com/BerriAI/litellm/pull/33714)
- fix(pricing): mark realtime-only gpt-realtime models as mode realtime by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33728](https://github.com/BerriAI/litellm/pull/33728)
- fix(rag): track LLM completion usage and spend for /v1/rag/query by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;32438](https://github.com/BerriAI/litellm/pull/32438)
- feat(anthropic): add enable\_anthropic\_prompt\_caching for automatic cache\_control injection by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33573](https://github.com/BerriAI/litellm/pull/33573)
- fix(anthropic): self-heal on missing thinking-signature errors from Bedrock/Vertex by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33719](https://github.com/BerriAI/litellm/pull/33719)
- fix(proxy): resolve router\_settings.plugins dotted paths and load plugins from installed packages by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33644](https://github.com/BerriAI/litellm/pull/33644)
- test(e2e): budget refusals are 429 for bare keys and team caps block every team key by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33632](https://github.com/BerriAI/litellm/pull/33632)
- feat(router): add router plugin reference catalog by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33746](https://github.com/BerriAI/litellm/pull/33746)
- test(e2e): assert an org budget block is a 429 naming the organization by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33638](https://github.com/BerriAI/litellm/pull/33638)
- fix(proxy): bill partial streamed spend when the client disconnects mid-stream by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33736](https://github.com/BerriAI/litellm/pull/33736)
- test(e2e): delete unreferenced Grafana panel docs by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33743](https://github.com/BerriAI/litellm/pull/33743)
- docs(tests/e2e): align skip-vs-fail docs with the hard-fail contract and scope the no-unit-tests rule by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33755](https://github.com/BerriAI/litellm/pull/33755)
- refactor(e2e): replace bespoke result reporter with standard JUnit report by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33758](https://github.com/BerriAI/litellm/pull/33758)
- test(e2e): user budget across keys and team member budget isolation by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33745](https://github.com/BerriAI/litellm/pull/33745)
- refactor(e2e): remove bob\_the\_builder; drive remediation from a Grafana alert (provisioned outside the repo) by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33749](https://github.com/BerriAI/litellm/pull/33749)
- feat(mcp): per-server outcomes for aggregate tools/list and truthful single-server REST statuses by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33153](https://github.com/BerriAI/litellm/pull/33153)
- test(e2e): mcp suite for key-without-access denial by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33752](https://github.com/BerriAI/litellm/pull/33752)
- chore(ci): merge oss branch by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33784](https://github.com/BerriAI/litellm/pull/33784)
- chore(ci): merge oss branch - July 17th by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33793](https://github.com/BerriAI/litellm/pull/33793)
- fix(ui): migrate tag deletion to shared DeleteResourceModal by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33795](https://github.com/BerriAI/litellm/pull/33795)
- build(rust): raise pyo3 to 0.29 so the native bridge compiles on Python 3.14 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33798](https://github.com/BerriAI/litellm/pull/33798)
- chore(guardrails): remove docstring from singulr module for consistency by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33800](https://github.com/BerriAI/litellm/pull/33800)
- fix(ui): stop credential edit from persisting the masked api key by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33797](https://github.com/BerriAI/litellm/pull/33797)
- build(deps): allow redisvl, pypdf, and openapi-core on Python 3.14 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33801](https://github.com/BerriAI/litellm/pull/33801)
- test(proxy): make streaming-cancel mocks awaitable for the disconnect slot release by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33802](https://github.com/BerriAI/litellm/pull/33802)
- test(e2e): a member's team budget cuts off only that member's key by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33718](https://github.com/BerriAI/litellm/pull/33718)
- test(e2e): a user's max\_budget follows the person across personal and team keys by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33762](https://github.com/BerriAI/litellm/pull/33762)
- test(e2e): skip flaky OpenAI GPT cells; raise multi-window max\_tokens by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33799](https://github.com/BerriAI/litellm/pull/33799)
- chore: remove accidentally committed dist tarball and ignore dist/ by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33805](https://github.com/BerriAI/litellm/pull/33805)
- fix(passthrough): stop classifying plain 'predict'/'search' paths as Vertex by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33658](https://github.com/BerriAI/litellm/pull/33658)
- build(deps): bump mcp lock to 1.28.1 to clear image-scan findings by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33803](https://github.com/BerriAI/litellm/pull/33803)
- test(pricing): pin the realtime mode assertion to the bundled cost map by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33806](https://github.com/BerriAI/litellm/pull/33806)
- fix(proxy): derive session id from Anthropic metadata.user\_id for session affinity by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33723](https://github.com/BerriAI/litellm/pull/33723)
- test(e2e): budget reset diagonal for team, org, user, and [#&#8203;32005](https://github.com/BerriAI/litellm/issues/32005) team-member keys by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33771](https://github.com/BerriAI/litellm/pull/33771)
- fix(proxy): source /v1/models token limits from the cost map instead of Router.get\_model\_group\_info by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33721](https://github.com/BerriAI/litellm/pull/33721)
- fix(fireworks\_ai): correct glm-5p2 prompt-cache read price to $0.14/1M by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33796](https://github.com/BerriAI/litellm/pull/33796)
- feat(proxy): add x-litellm-model-name response header with deployment model string by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33698](https://github.com/BerriAI/litellm/pull/33698)
- feat: add Straiker guardrail integration by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33781](https://github.com/BerriAI/litellm/pull/33781)
- fix(vertex\_ai): exclude Gemini Google Search grounding tokens from input token billing by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33742](https://github.com/BerriAI/litellm/pull/33742)
- feat(fireworks\_ai): map litellm session id to x-session-affinity header for prompt caching by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33717](https://github.com/BerriAI/litellm/pull/33717)
- feat(ui): configure Anthropic automatic prompt caching from the Admin UI by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33581](https://github.com/BerriAI/litellm/pull/33581)
- fix(router): enforce context-window pre-call checks for Responses API input by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33706](https://github.com/BerriAI/litellm/pull/33706)
- fix(otel): restore proxy-level error.\* attributes on v2 failure spans (LIT-4179) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33664](https://github.com/BerriAI/litellm/pull/33664)
- fix(mcp): persist config.yaml DCR clients in a server-scoped store so refresh survives token expiry by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33768](https://github.com/BerriAI/litellm/pull/33768)
- refactor(ui): consolidate Add/Edit credential modals into one CredentialModal by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32572](https://github.com/BerriAI/litellm/pull/32572)
- feat(mcp): add ID-JAG (identity assertion authorization grant) support for MCP egress by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;31516](https://github.com/BerriAI/litellm/pull/31516)
- refactor(ui): migrate policy attachments table onto shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33827](https://github.com/BerriAI/litellm/pull/33827)
- docs(litellm-rust): add provider coding standards by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33833](https://github.com/BerriAI/litellm/pull/33833)
- test(e2e): rename Gateway to ProxyClient and expose it as a session-scoped pytest fixture by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33750](https://github.com/BerriAI/litellm/pull/33750)
- feat(messages): route Azure Anthropic /messages through Rust behind rust:true by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33616](https://github.com/BerriAI/litellm/pull/33616)
- test(e2e): add Locust throughput load test that runs last by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33748](https://github.com/BerriAI/litellm/pull/33748)
- fix(proxy): resolve team wildcard credentials for vector store files by [@&#8203;shivamrawat1](https://github.com/shivamrawat1) in [#&#8203;33649](https://github.com/BerriAI/litellm/pull/33649)
- refactor(e2e): fold claude\_code HTTP probes onto shared ProxyClient methods by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33760](https://github.com/BerriAI/litellm/pull/33760)
- test(e2e): harden stage flakes for batches, UI, and MCP by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33831](https://github.com/BerriAI/litellm/pull/33831)
- fix(e2e): migrate load suite from e2e\_gateway to ProxyClient by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;33839](https://github.com/BerriAI/litellm/pull/33839)
- chore(e2e): remove tests/e2e/docker-compose.yml by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33837](https://github.com/BerriAI/litellm/pull/33837)
- test(e2e): cover /v1/responses openai basic nonstream and stream by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33830](https://github.com/BerriAI/litellm/pull/33830)
- test(e2e): cover /v1/responses openai cost\_logged and tool\_use by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33835](https://github.com/BerriAI/litellm/pull/33835)
- test(e2e): cover /v1/responses OpenAI vision and Anthropic basic by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33838](https://github.com/BerriAI/litellm/pull/33838)
- test(e2e): spendlog cost for streaming /v1/messages via responses bridge by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33753](https://github.com/BerriAI/litellm/pull/33753)
- fix(docker): bake prisma CLI and engines at a fixed path so fresh-DB migrations work for any uid offline by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33853](https://github.com/BerriAI/litellm/pull/33853)
- feat(chat-ui): add personal Logs view scoped to the current user by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33829](https://github.com/BerriAI/litellm/pull/33829)
- chore: bump litellm-proxy-extras 0.4.78 -> 0.4.79 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33855](https://github.com/BerriAI/litellm/pull/33855)
- docs(litellm-rust): require the official Rust Style Guide in agent rules by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33867](https://github.com/BerriAI/litellm/pull/33867)
- fix(router): treat malformed configured token limits as absent on /v1/models by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33864](https://github.com/BerriAI/litellm/pull/33864)
- chore: rebuild admin UI bundle for the rc release by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33857](https://github.com/BerriAI/litellm/pull/33857)
- docs(rust): add provider abstraction standards by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33865](https://github.com/BerriAI/litellm/pull/33865)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33868](https://github.com/BerriAI/litellm/pull/33868)
- chore(release): backport [#&#8203;33929](https://github.com/BerriAI/litellm/issues/33929) to rc/1.94.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34033](https://github.com/BerriAI/litellm/pull/34033)
- chore(release): backport [#&#8203;33810](https://github.com/BerriAI/litellm/issues/33810), [#&#8203;33733](https://github.com/BerriAI/litellm/issues/33733) to rc/1.94.0 and bump litellm-proxy-extras to 0.4.79.post1 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34215](https://github.com/BerriAI/litellm/pull/34215)
- chore(ui): rebuild Next.js bundle on rc/1.94.0 so the Cost Optimization page ships by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34216](https://github.com/BerriAI/litellm/pull/34216)
- chore(release): backport auth, CLI SSO and guardrail fixes to rc/1.94.0 and refresh flagged dependencies by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34640](https://github.com/BerriAI/litellm/pull/34640)
- chore(release): backport [#&#8203;33899](https://github.com/BerriAI/litellm/issues/33899), [#&#8203;33978](https://github.com/BerriAI/litellm/issues/33978), [#&#8203;34582](https://github.com/BerriAI/litellm/issues/34582), [#&#8203;34675](https://github.com/BerriAI/litellm/issues/34675) to rc/1.94.0 and bump litellm-proxy-extras to 0.4.79.post2 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34855](https://github.com/BerriAI/litellm/pull/34855)
- fix(ui): backport cache leakage card layout fix to rc/1.94.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34964](https://github.com/BerriAI/litellm/pull/34964)
- fix(ui): add missing cost-optimization page description on rc/1.94.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34967](https://github.com/BerriAI/litellm/pull/34967)
- feat(ui): mark Cost Optimization as beta in the left nav ([#&#8203;34984](https://github.com/BerriAI/litellm/issues/34984)) by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34987](https://github.com/BerriAI/litellm/pull/34987)
- chore: rebuild Admin UI bundle for v1.94.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34982](https://github.com/BerriAI/litellm/pull/34982)
- fix(cost-optimization): backport the savings chart axis fix and methodology popovers to rc/1.94.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34994](https://github.com/BerriAI/litellm/pull/34994)
- chore: rebuild Admin UI bundle for rc/1.94.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34995](https://github.com/BerriAI/litellm/pull/34995)

**Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.93.0...v1.94.0>

### [`v1.94.0`](https://github.com/BerriAI/litellm/releases/tag/v1.94.0)

##### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.94.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.94.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.94.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

##### What's Changed

- feat(ui): working Test Connection for the complexity auto router by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;32950](https://github.com/BerriAI/litellm/pull/32950)
- fix(xecguard): use StandardLoggingGuardrailInformation in logging hook by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;32911](https://github.com/BerriAI/litellm/pull/32911)
- feat(ui): adopt openapi-react-query ($api) and convert useCustomers by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32949](https://github.com/BerriAI/litellm/pull/32949)
- refactor(ui): colocate the mcp-servers view, keeping the shared mcp\_tools surface by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32968](https://github.com/BerriAI/litellm/pull/32968)
- refactor(ui): convert endpoint usage charts to shadcn/recharts by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32723](https://github.com/BerriAI/litellm/pull/32723)
- fix(proxy-auth): stop unrecognized model namespaces slipping through provider wildcard keys by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32979](https://github.com/BerriAI/litellm/pull/32979)
- feat(router): random-pick multi-model complexity tiers by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32967](https://github.com/BerriAI/litellm/pull/32967)
- fix(xecguard): sanitize scan result before recording it for logging by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;32935](https://github.com/BerriAI/litellm/pull/32935)
- fix(auto\_router): filter embedding models in complexity tab dropdowns, require all tiers, inline validation by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;32978](https://github.com/BerriAI/litellm/pull/32978)
- fix(anthropic): translate raw adaptive thinking for pre-4.6 models on chat completions and Bedrock Converse by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;32944](https://github.com/BerriAI/litellm/pull/32944)
- feat(router): add Router(plugins=\[...]) routing-plugin pipeline by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32972](https://github.com/BerriAI/litellm/pull/32972)
- feat(router): soft-floor adaptive mode for complexity router by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32947](https://github.com/BerriAI/litellm/pull/32947)
- docs(github): add QA runbook section to the PR template by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32965](https://github.com/BerriAI/litellm/pull/32965)
- fix(model\_cost): add supports\_reasoning: false to Gemini image generation models by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32836](https://github.com/BerriAI/litellm/pull/32836)
- build(dev-env): add make bootstrap and unprovisioned-checkout preflight to pre-commit by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32981](https://github.com/BerriAI/litellm/pull/32981)
- ci(ui): report only error-level knip findings in CI by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;32971](https://github.com/BerriAI/litellm/pull/32971)
- feat(batches): track cost for unmanaged Bedrock batches, generalize the flag by [@&#8203;Sameerlite](https://github.com/Sameerlite) in [#&#8203;32315](https://github.com/BerriAI/litellm/pull/32315)
- fix(guardrails): walk custom\_tool\_call\_output items in \_content\_utils by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;32969](https://github.com/BerriAI/litellm/pull/32969)
- fix: show and allow editing team model aliases after team creation by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33047](https://github.com/BerriAI/litellm/pull/33047)
- chore(deps): bump pillow to 12.3.0 to resolve osv-scan CVEs by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33093](https://github.com/BerriAI/litellm/pull/33093)
- feat(mcp): mint gateway-bound envelope at the token endpoint for dcr\_bridge oauth\_delegate by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32828](https://github.com/BerriAI/litellm/pull/32828)
- fix(mcp): surface rejected delegate-auth upstream tokens as connect-time 401 by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32741](https://github.com/BerriAI/litellm/pull/32741)
- fix(proxy): track unauthenticated pass-through requests in spend logs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;32410](https://github.com/BerriAI/litellm/pull/32410)
- feat(lasso): send source.type for Used By attribution by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33090](https://github.com/BerriAI/litellm/pull/33090)
- fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@&#8203;thibault-linktree](https://github.com/thibault-linktree) in [#&#8203;33025](https://github.com/BerriAI/litellm/pull/33025)
- fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33099](https://github.com/BerriAI/litellm/pull/33099)
- fix(completion): forward aws credential kwargs into litellm\_params so the responses bridge keeps WIF auth by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;32956](https://github.com/BerriAI/litellm/pull/32956)
- fix(ui): respect litellm\_key\_header\_name in BYOK credential save and workflow runs fetches by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33103](https://github.com/BerriAI/litellm/pull/33103)
- refactor(ui): standardize debounce waits behind shared DEBOUNCE\_WAIT\_MS constant by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33040](https://github.com/BerriAI/litellm/pull/33040)
- feat(ui): rebuild the Virtual Keys table on the shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;32991](https://github.com/BerriAI/litellm/pull/32991)
- fix: redact async complete streaming response for custom callbacks by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33106](https://github.com/BerriAI/litellm/pull/33106)
- build(ui): bump [@&#8203;tanstack/react-pacer](https://github.com/tanstack/react-pacer) from 0.2.0 to 0.22.1 by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33041](https://github.com/BerriAI/litellm/pull/33041)
- fix(ui): address Virtual Keys redesign review nits by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33112](https://github.com/BerriAI/litellm/pull/33112)
- fix(openai/responses): clamp max\_output\_tokens below API minimum by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33098](https://github.com/BerriAI/litellm/pull/33098)
- fix(prometheus): read v3 rate limiter remaining values for per-key model gauges by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;33119](https://github.com/BerriAI/litellm/pull/33119)
- fix(ui): drop w-full from page-content wrappers to remove 32px horizontal overflow by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33118](https://github.com/BerriAI/litellm/pull/33118)
- refactor(ui): migrate straightforward value debounces to react-pacer by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33042](https://github.com/BerriAI/litellm/pull/33042)
- feat(mcp): interactive SSO sign-in for dcr\_bridge oauth\_delegate DCR clients by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32946](https://github.com/BerriAI/litellm/pull/32946)
- test(proxy): add regression tests for management\_endpoints edge cases by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;32976](https://github.com/BerriAI/litellm/pull/32976)
- fix(auto-router): correct Responses API tool\_choice shape and propagate alias litellm\_params by [@&#8203;krrish-berri-2](https://github.com/krrish-berri-2) in [#&#8203;32974](https://github.com/BerriAI/litellm/pull/32974)
- fix(ui): render the sidebar scrollbar with shadcn ScrollArea by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33124](https://github.com/BerriAI/litellm/pull/33124)
- refactor(ui): migrate callback debounce sites to react-pacer with regression tests by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;33043](https://github.com/BerriAI/litellm/pull/33043)
- chore: add CODEOWNERS for ui and proxy UI build artifacts by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33131](https://github.com/BerriAI/litellm/pull/33131)
- feat(mcp): client-held refresh envelope for the dcr\_bridge oauth\_delegate flow by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;32980](https://github.com/BerriAI/litellm/pull/32980)
- feat(ui): rebuild the Teams table on the shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33128](https://github.com/BerriAI/litellm/pull/33128)
- fix(mcp): relay upstream OAuth token and DCR rejections instead of a generic 500 by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;33113](https://github.com/BerriAI/litellm/pull/33113)
- fix(keys): persist key\_type so the UI shows correct key scope instead of "All Proxy…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.