feat(router): resolve auto-router routing plugins from proxy YAML config - #33251
Conversation
Router(plugins=[...]) was Python-SDK constructor only, so proxy/YAML users had no way to configure it, and the merged pipeline narrowed candidates from the outer model alias rather than the auto-router's actual tier pool, making it a no-op for auto_router deployments. Add complexity_router_config.plugins (dotted-path strings resolved via get_instance_fn, the same convention litellm_settings.callbacks uses) and run the resolved plugins against ComplexityRouter's tier pool at every model-pick site, so a policy plugin narrows what get_model_for_tier actually returns instead of the outer alias list. adaptive=True with plugins set now raises at config validation instead of silently ignoring the plugins, since the bandit selector doesn't consume narrowed pools yet. Also fixes a latent bug in Router._generate_model_id: it json.dumps every litellm_params dict value to build a deployment hash id, which crashed once a live plugin object could land inside complexity_router_config.
|
|
Greptile SummaryThis PR wires
Confidence Score: 5/5Safe to merge; all routing paths are covered by tests and the plugin pipeline is fail-closed by design. The change is well-scoped: new behaviour only activates when litellm/router_strategy/complexity_router/complexity_router.py — the misleading comment in the no-user-message branch (lines 993-996) should be corrected before it becomes a maintenance trap.
|
| Filename | Overview |
|---|---|
| litellm/proxy/proxy_server.py | Adds resolve_complexity_router_plugins() to resolve dotted-path plugin strings at proxy startup, with an isinstance + iscoroutinefunction guard that correctly catches both missing run and synchronous run implementations before the first request. |
| litellm/router.py | Adds _json_default_stable_id to produce a stable, memory-address-free serialization fallback for plugin instances in _generate_model_id, preventing deployment-ID churn on hot-reload. |
| litellm/router_strategy/complexity_router/complexity_router.py | Adds _pick_model_for_tier to run the plugin pipeline on the tier's candidate pool; wires it into all three routing paths. A comment in the no-user-message branch incorrectly states that get_model_for_tier still provides a default_model fallback when plugins are active, but _pick_model_for_tier bypasses get_model_for_tier entirely when plugins are configured. |
| litellm/router_strategy/complexity_router/config.py | Adds plugins field to ComplexityRouterConfig and a model validator that hard-rejects plugins + adaptive=True combination since the bandit selector doesn't yet consume narrowed pools. |
| litellm/types/router.py | Adds @runtime_checkable decorator to RoutingPlugin Protocol so that isinstance checks work at config-load time for the startup validation guard. |
| tests/test_litellm/router_strategy/test_complexity_router.py | Adds TestRoutingPlugins with nine scenario tests covering plugin narrowing, zero-candidate raise, metadata propagation, all three routing paths (complexity scorer, keyword override, no-user-message), session-affinity bypass disable, and a backwards-compat regression guard for the no-plugin path. |
| tests/test_litellm/proxy/proxy_server/test_proxy_config.py | Adds four unit tests for resolve_complexity_router_plugins: no-op on missing key, successful dotted-path resolution, rejection of non-RoutingPlugin objects, and rejection of a synchronous run method. |
| tests/test_litellm/router_strategy/test_router_routing_plugins.py | Adds two tests for _json_default_stable_id and _generate_model_id stability when litellm_params contain a plugin instance, ensuring the deployment ID doesn't change between process restarts. |
Reviews (8): Last reviewed commit: "fix(router): address Greptile findings o..." | Re-trigger Greptile
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8d5e033929
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if self.config.default_model: | ||
| return self.config.default_model |
There was a problem hiding this comment.
Fail closed when plugins deny every candidate
When a routing plugin removes every candidate, returning default_model bypasses the plugin's policy; for example, a tenant or budget plugin that intentionally blocks the whole tier will still send the request to whatever default is configured, even if that model was not allowed. The existing Router plugin path treats an empty filtered set as a policy denial and raises, so the complexity-router path should only route to a model that survived the plugin pipeline, or fail closed.
Useful? React with 👍 / 👎.
|
|
||
| if user_message is None: | ||
| verbose_router_logger.debug("ComplexityRouter: No user message found, routing to default model") | ||
| routed_model = self.config.default_model or await self._pick_model_for_tier( |
There was a problem hiding this comment.
Run plugins before choosing the no-user default
When default_model is set, this branch returns it without invoking _pick_model_for_tier, so complexity_router_config.plugins are skipped for requests whose normalized messages contain no role="user" content. In deployments where plugins enforce tenant, provider, or budget policy, a system/assistant-only request can bypass those filters and still hit the default model; run the plugin pipeline against the fallback tier/default before returning.
Useful? React with 👍 / 👎.
| context = RoutingContext( | ||
| raw_messages=raw_messages or [], | ||
| structured_messages=resolved_messages or [], | ||
| candidate_models=list(self._tier_pools().get(tier_key, [])), |
There was a problem hiding this comment.
Seed plugins with deployment models instead of tier aliases
This populates RoutingContext.candidate_models from the tier config entries, which are usually router model-group aliases, while the shared RoutingPlugin contract and existing policies operate on deployment litellm_params.model values such as openai/.... After the hook returns an alias, the downstream router can still select any deployment in that group because these plugin candidates are not carried into the deployment filter, so provider/budget policies cannot reliably remove a specific backend deployment.
Useful? React with 👍 / 👎.
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 2 · PR risk: 0/10 |
… fallback json.dumps(v, default=str) on a litellm_params dict containing a live RoutingPlugin instance fell back to object.__repr__'s default <module.Class object at 0x...>, embedding the instance's memory address. _generate_model_id's hash (and therefore the deployment id) changed on every process restart/hot-reload for any deployment with complexity_router_config.plugins configured, defeating the function's own "consistently generate the same id" contract and orphaning anything keyed on that id across restarts (e.g. Redis-backed per-deployment state). Use the plugin's fully-qualified class name instead, which is stable across restarts.
|
@greptileai review |
…gate router_code_coverage.py's AST scanner requires every router.py function be called by name somewhere in tests/, and flagged the new _json_default_stable_id helper from the previous commit.
…eria AI Session-affinity pin shortcut: async_pre_routing_hook returned a session's first-turn pinned model on every later turn without ever re-running it through the plugin pipeline, so a policy plugin (e.g. a budget cap crossed mid-session) was only enforced on turn one. Now the pin shortcut is disabled whenever plugins are configured, so every turn re-runs _classify_and_route (and therefore the plugins). Plugin resolution validation: get_instance_fn accepts any dotted path and returns whatever object it finds there, so a misconfigured complexity_router_config.plugins entry passed proxy startup silently and only surfaced as a confusing AttributeError on the first request that reached the plugin pipeline. Extracted the resolution logic into resolve_complexity_router_plugins() and added an isinstance(..., RoutingPlugin) check that fails proxy startup immediately with a clear error instead.
…plugin-narrowed tier default_model was never checked against the configured plugins, so it functioned as an unconditional escape hatch around whatever policy a plugin enforces -- a tenant/budget plugin narrowing a tier to zero candidates could still be bypassed by the fallback. Drop the fallback entirely for this path; a plugin narrowing to zero is a policy decision, not something to route around, matching the fail-closed behavior the Router-level plugin pipeline already uses for the same situation. Flagged by Veria AI on PR #33251.
|
@greptileai review |
|
@greptileai review |
…itellm_router_plugins_auto_router_yaml
|
@greptileai review |
…ve_complexity_router_plugins
|
@greptileai review |
…n no-user-message path self.config.default_model or await self._pick_model_for_tier(...) -- Python's `or` short-circuits on a truthy default_model, so _pick_model_for_tier (and therefore the plugin pipeline) never ran at all for the no-user-message path whenever default_model was configured. A tenant/budget plugin's decision was silently bypassable this way even after the other two policy-bypass fixes, since this call site had a different shape from the other three pick sites. Removed the short-circuit; falls through to _pick_model_for_tier -> get_model_for_tier, which already checks the MEDIUM tier before default_model -- the same priority every other call site uses. Flagged by Veria AI on PR #33251.
|
@greptileai review |
Preserve default_model-first priority in the no-user-message path when no plugins are configured, instead of unconditionally flipping to the MEDIUM tier -- the plugin-bypass fix must not silently change model selection for the (much larger) population of users who don't use plugins at all. Gated on self.config.plugins, matching the pattern already used elsewhere in this PR, per CLAUDE.md's guidance against backwards-compat flags when a plain conditional does the job. Also close a gap in the plugin validation added earlier: @runtime_checkable only checks that `run` exists as an attribute, not that it's a coroutine function, so a synchronous `def run(self, context)` passed isinstance(resolved_plugin, RoutingPlugin) at startup and only failed at request time with a confusing TypeError. Added an inspect.iscoroutinefunction check. Both flagged by Greptile on PR #33251.
|
@greptileai review |
…fig (BerriAI#33251) * feat(router): resolve auto-router routing plugins from proxy YAML config Router(plugins=[...]) was Python-SDK constructor only, so proxy/YAML users had no way to configure it, and the merged pipeline narrowed candidates from the outer model alias rather than the auto-router's actual tier pool, making it a no-op for auto_router deployments. Add complexity_router_config.plugins (dotted-path strings resolved via get_instance_fn, the same convention litellm_settings.callbacks uses) and run the resolved plugins against ComplexityRouter's tier pool at every model-pick site, so a policy plugin narrows what get_model_for_tier actually returns instead of the outer alias list. adaptive=True with plugins set now raises at config validation instead of silently ignoring the plugins, since the bandit selector doesn't consume narrowed pools yet. Also fixes a latent bug in Router._generate_model_id: it json.dumps every litellm_params dict value to build a deployment hash id, which crashed once a live plugin object could land inside complexity_router_config. * fix(router): use stable class name, not object repr, in model-id json fallback json.dumps(v, default=str) on a litellm_params dict containing a live RoutingPlugin instance fell back to object.__repr__'s default <module.Class object at 0x...>, embedding the instance's memory address. _generate_model_id's hash (and therefore the deployment id) changed on every process restart/hot-reload for any deployment with complexity_router_config.plugins configured, defeating the function's own "consistently generate the same id" contract and orphaning anything keyed on that id across restarts (e.g. Redis-backed per-deployment state). Use the plugin's fully-qualified class name instead, which is stable across restarts. * test(router): cover _json_default_stable_id for router_code_coverage gate router_code_coverage.py's AST scanner requires every router.py function be called by name somewhere in tests/, and flagged the new _json_default_stable_id helper from the previous commit. * fix(router): close two routing-plugin policy-bypass gaps flagged by Veria AI Session-affinity pin shortcut: async_pre_routing_hook returned a session's first-turn pinned model on every later turn without ever re-running it through the plugin pipeline, so a policy plugin (e.g. a budget cap crossed mid-session) was only enforced on turn one. Now the pin shortcut is disabled whenever plugins are configured, so every turn re-runs _classify_and_route (and therefore the plugins). Plugin resolution validation: get_instance_fn accepts any dotted path and returns whatever object it finds there, so a misconfigured complexity_router_config.plugins entry passed proxy startup silently and only surfaced as a confusing AttributeError on the first request that reached the plugin pipeline. Extracted the resolution logic into resolve_complexity_router_plugins() and added an isinstance(..., RoutingPlugin) check that fails proxy startup immediately with a clear error instead. * fix(router): raise instead of falling back to default_model on empty plugin-narrowed tier default_model was never checked against the configured plugins, so it functioned as an unconditional escape hatch around whatever policy a plugin enforces -- a tenant/budget plugin narrowing a tier to zero candidates could still be bypassed by the fallback. Drop the fallback entirely for this path; a plugin narrowing to zero is a policy decision, not something to route around, matching the fail-closed behavior the Router-level plugin pipeline already uses for the same situation. Flagged by Veria AI on PR BerriAI#33251. * style: ruff format complexity_router.py * style(proxy): use modern str | None instead of Optional[str] in resolve_complexity_router_plugins * fix(router): stop default_model short-circuit from skipping plugins on no-user-message path self.config.default_model or await self._pick_model_for_tier(...) -- Python's `or` short-circuits on a truthy default_model, so _pick_model_for_tier (and therefore the plugin pipeline) never ran at all for the no-user-message path whenever default_model was configured. A tenant/budget plugin's decision was silently bypassable this way even after the other two policy-bypass fixes, since this call site had a different shape from the other three pick sites. Removed the short-circuit; falls through to _pick_model_for_tier -> get_model_for_tier, which already checks the MEDIUM tier before default_model -- the same priority every other call site uses. Flagged by Veria AI on PR BerriAI#33251. * fix(router): address Greptile findings on the plugin-bypass fixes Preserve default_model-first priority in the no-user-message path when no plugins are configured, instead of unconditionally flipping to the MEDIUM tier -- the plugin-bypass fix must not silently change model selection for the (much larger) population of users who don't use plugins at all. Gated on self.config.plugins, matching the pattern already used elsewhere in this PR, per CLAUDE.md's guidance against backwards-compat flags when a plain conditional does the job. Also close a gap in the plugin validation added earlier: @runtime_checkable only checks that `run` exists as an attribute, not that it's a coroutine function, so a synchronous `def run(self, context)` passed isinstance(resolved_plugin, RoutingPlugin) at startup and only failed at request time with a confusing TypeError. Added an inspect.iscoroutinefunction check. Both flagged by Greptile on PR BerriAI#33251.
…4.0) (#201) This PR contains the following updates: | Package | Update | Change | |---|---|---| | [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.93.0` → `v1.94.0` | --- ### Release Notes <details> <summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary> ### [`v1.94.0`](https://github.com/BerriAI/litellm/releases/tag/v1.94.0) [Compare Source](https://github.com/BerriAI/litellm/compare/v1.94.0...v1.94.0) ##### Verify Docker Image Signature All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0). **Verify using the pinned commit hash (recommended):** A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \ ghcr.io/berriai/litellm:v1.94.0 ``` **Verify using the release tag (convenience):** Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/v1.94.0/cosign.pub \ ghcr.io/berriai/litellm:v1.94.0 ``` Expected output: ``` The following checks were performed on each of these signatures: - The cosign claims were validated - The signatures were verified against the specified public key ``` *** ##### What's Changed - feat(ui): working Test Connection for the complexity auto router by [@​akapur99](https://github.com/akapur99) in [#​32950](https://github.com/BerriAI/litellm/pull/32950) - fix(xecguard): use StandardLoggingGuardrailInformation in logging hook by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​32911](https://github.com/BerriAI/litellm/pull/32911) - feat(ui): adopt openapi-react-query ($api) and convert useCustomers by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32949](https://github.com/BerriAI/litellm/pull/32949) - refactor(ui): colocate the mcp-servers view, keeping the shared mcp\_tools surface by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32968](https://github.com/BerriAI/litellm/pull/32968) - refactor(ui): convert endpoint usage charts to shadcn/recharts by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32723](https://github.com/BerriAI/litellm/pull/32723) - fix(proxy-auth): stop unrecognized model namespaces slipping through provider wildcard keys by [@​mateo-berri](https://github.com/mateo-berri) in [#​32979](https://github.com/BerriAI/litellm/pull/32979) - feat(router): random-pick multi-model complexity tiers by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32967](https://github.com/BerriAI/litellm/pull/32967) - fix(xecguard): sanitize scan result before recording it for logging by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​32935](https://github.com/BerriAI/litellm/pull/32935) - fix(auto\_router): filter embedding models in complexity tab dropdowns, require all tiers, inline validation by [@​akapur99](https://github.com/akapur99) in [#​32978](https://github.com/BerriAI/litellm/pull/32978) - fix(anthropic): translate raw adaptive thinking for pre-4.6 models on chat completions and Bedrock Converse by [@​akapur99](https://github.com/akapur99) in [#​32944](https://github.com/BerriAI/litellm/pull/32944) - feat(router): add Router(plugins=\[...]) routing-plugin pipeline by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32972](https://github.com/BerriAI/litellm/pull/32972) - feat(router): soft-floor adaptive mode for complexity router by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32947](https://github.com/BerriAI/litellm/pull/32947) - docs(github): add QA runbook section to the PR template by [@​mateo-berri](https://github.com/mateo-berri) in [#​32965](https://github.com/BerriAI/litellm/pull/32965) - fix(model\_cost): add supports\_reasoning: false to Gemini image generation models by [@​mateo-berri](https://github.com/mateo-berri) in [#​32836](https://github.com/BerriAI/litellm/pull/32836) - build(dev-env): add make bootstrap and unprovisioned-checkout preflight to pre-commit by [@​mateo-berri](https://github.com/mateo-berri) in [#​32981](https://github.com/BerriAI/litellm/pull/32981) - ci(ui): report only error-level knip findings in CI by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32971](https://github.com/BerriAI/litellm/pull/32971) - feat(batches): track cost for unmanaged Bedrock batches, generalize the flag by [@​Sameerlite](https://github.com/Sameerlite) in [#​32315](https://github.com/BerriAI/litellm/pull/32315) - fix(guardrails): walk custom\_tool\_call\_output items in \_content\_utils by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​32969](https://github.com/BerriAI/litellm/pull/32969) - fix: show and allow editing team model aliases after team creation by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33047](https://github.com/BerriAI/litellm/pull/33047) - chore(deps): bump pillow to 12.3.0 to resolve osv-scan CVEs by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33093](https://github.com/BerriAI/litellm/pull/33093) - feat(mcp): mint gateway-bound envelope at the token endpoint for dcr\_bridge oauth\_delegate by [@​tin-berri](https://github.com/tin-berri) in [#​32828](https://github.com/BerriAI/litellm/pull/32828) - fix(mcp): surface rejected delegate-auth upstream tokens as connect-time 401 by [@​tin-berri](https://github.com/tin-berri) in [#​32741](https://github.com/BerriAI/litellm/pull/32741) - fix(proxy): track unauthenticated pass-through requests in spend logs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​32410](https://github.com/BerriAI/litellm/pull/32410) - feat(lasso): send source.type for Used By attribution by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33090](https://github.com/BerriAI/litellm/pull/33090) - fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@​thibault-linktree](https://github.com/thibault-linktree) in [#​33025](https://github.com/BerriAI/litellm/pull/33025) - fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@​tin-berri](https://github.com/tin-berri) in [#​33099](https://github.com/BerriAI/litellm/pull/33099) - fix(completion): forward aws credential kwargs into litellm\_params so the responses bridge keeps WIF auth by [@​mateo-berri](https://github.com/mateo-berri) in [#​32956](https://github.com/BerriAI/litellm/pull/32956) - fix(ui): respect litellm\_key\_header\_name in BYOK credential save and workflow runs fetches by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33103](https://github.com/BerriAI/litellm/pull/33103) - refactor(ui): standardize debounce waits behind shared DEBOUNCE\_WAIT\_MS constant by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33040](https://github.com/BerriAI/litellm/pull/33040) - feat(ui): rebuild the Virtual Keys table on the shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​32991](https://github.com/BerriAI/litellm/pull/32991) - fix: redact async complete streaming response for custom callbacks by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33106](https://github.com/BerriAI/litellm/pull/33106) - build(ui): bump [@​tanstack/react-pacer](https://github.com/tanstack/react-pacer) from 0.2.0 to 0.22.1 by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33041](https://github.com/BerriAI/litellm/pull/33041) - fix(ui): address Virtual Keys redesign review nits by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33112](https://github.com/BerriAI/litellm/pull/33112) - fix(openai/responses): clamp max\_output\_tokens below API minimum by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33098](https://github.com/BerriAI/litellm/pull/33098) - fix(prometheus): read v3 rate limiter remaining values for per-key model gauges by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33119](https://github.com/BerriAI/litellm/pull/33119) - fix(ui): drop w-full from page-content wrappers to remove 32px horizontal overflow by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33118](https://github.com/BerriAI/litellm/pull/33118) - refactor(ui): migrate straightforward value debounces to react-pacer by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33042](https://github.com/BerriAI/litellm/pull/33042) - feat(mcp): interactive SSO sign-in for dcr\_bridge oauth\_delegate DCR clients by [@​tin-berri](https://github.com/tin-berri) in [#​32946](https://github.com/BerriAI/litellm/pull/32946) - test(proxy): add regression tests for management\_endpoints edge cases by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​32976](https://github.com/BerriAI/litellm/pull/32976) - fix(auto-router): correct Responses API tool\_choice shape and propagate alias litellm\_params by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32974](https://github.com/BerriAI/litellm/pull/32974) - fix(ui): render the sidebar scrollbar with shadcn ScrollArea by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33124](https://github.com/BerriAI/litellm/pull/33124) - refactor(ui): migrate callback debounce sites to react-pacer with regression tests by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33043](https://github.com/BerriAI/litellm/pull/33043) - chore: add CODEOWNERS for ui and proxy UI build artifacts by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33131](https://github.com/BerriAI/litellm/pull/33131) - feat(mcp): client-held refresh envelope for the dcr\_bridge oauth\_delegate flow by [@​tin-berri](https://github.com/tin-berri) in [#​32980](https://github.com/BerriAI/litellm/pull/32980) - feat(ui): rebuild the Teams table on the shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33128](https://github.com/BerriAI/litellm/pull/33128) - fix(mcp): relay upstream OAuth token and DCR rejections instead of a generic 500 by [@​tin-berri](https://github.com/tin-berri) in [#​33113](https://github.com/BerriAI/litellm/pull/33113) - fix(keys): persist key\_type so the UI shows correct key scope instead of "All Proxy Models" by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33115](https://github.com/BerriAI/litellm/pull/33115) - feat(router): opt-in session affinity for complexity router by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​33126](https://github.com/BerriAI/litellm/pull/33126) - feat(prometheus): expose video duration and image count consumption metrics by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33138](https://github.com/BerriAI/litellm/pull/33138) - test(e2e): otel trace completeness on /chat/completions by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33132](https://github.com/BerriAI/litellm/pull/33132) - fix(sso): paginate through all pages when fetching service principal group assignments by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33149](https://github.com/BerriAI/litellm/pull/33149) - test(e2e): otel trace completeness on /v1/messages by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33133](https://github.com/BerriAI/litellm/pull/33133) - feat(ui): add adaptive routing settings to Auto-Router v2 by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​33146](https://github.com/BerriAI/litellm/pull/33146) - refactor(mcp): extract the dcr\_bridge token flow into bridge\_token\_flow\.py by [@​tin-berri](https://github.com/tin-berri) in [#​33141](https://github.com/BerriAI/litellm/pull/33141) - chore: bump litellm 1.93.0 -> 1.94.0, litellm-enterprise 0.1.49 -> 0.1.50, litellm-proxy-extras 0.4.76 -> 0.4.77 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33229](https://github.com/BerriAI/litellm/pull/33229) - fix(proxy): route master key to team-scoped models by [@​kunal2002](https://github.com/kunal2002) in [#​32926](https://github.com/BerriAI/litellm/pull/32926) - chore(deps): pin httplib2 and setuptools transitive floors by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33233](https://github.com/BerriAI/litellm/pull/33233) - feat(ui): left-anchor the Create Key and Create Team CTAs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33248](https://github.com/BerriAI/litellm/pull/33248) - fix(anthropic/passthrough): drop incompatible temperature when downgrading adaptive thinking for pre-4.6 models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33244](https://github.com/BerriAI/litellm/pull/33244) - fix(guardrails): run apply\_guardrail-style model-level pre\_call guardrails at deployment hook by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33136](https://github.com/BerriAI/litellm/pull/33136) - fix(proxy)!: enforce user budget on team keys (read-time + reservation) with UI opt-out by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​32005](https://github.com/BerriAI/litellm/pull/32005) - fix(e2e): bound spend-log snapshots to a /spend/logs/v2 window by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33265](https://github.com/BerriAI/litellm/pull/33265) - test(e2e): cover key rpm/tpm rate limiting, window reset, and pacing headers by [@​mateo-berri](https://github.com/mateo-berri) in [#​32914](https://github.com/BerriAI/litellm/pull/32914) - fix(anthropic): use native output capability by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​33235](https://github.com/BerriAI/litellm/pull/33235) - fix(ci): retry setup-uv installs to survive transient manifest fetch failures by [@​mateo-berri](https://github.com/mateo-berri) in [#​33279](https://github.com/BerriAI/litellm/pull/33279) - fix(proxy): never log raw virtual keys in key insertion debug output by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33268](https://github.com/BerriAI/litellm/pull/33268) - fix(bedrock\_mantle): route xai.grok-4.3 via /openai/v1 frontier path by [@​marty-sullivan](https://github.com/marty-sullivan) in [#​33027](https://github.com/BerriAI/litellm/pull/33027) - feat(pricing): add gemini-omni-flash-preview with video output token pricing by [@​mateo-berri](https://github.com/mateo-berri) in [#​33274](https://github.com/BerriAI/litellm/pull/33274) - fix(auth): scope the JWT enterprise gate to actual JWTs by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33296](https://github.com/BerriAI/litellm/pull/33296) - fix(s3): sanitize slashes in response-id-derived object key file name by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33271](https://github.com/BerriAI/litellm/pull/33271) - refactor(ui): migrate guardrails table onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33303](https://github.com/BerriAI/litellm/pull/33303) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33308](https://github.com/BerriAI/litellm/pull/33308) - feat(guardrails): streaming text transformation in generic\_guardrail\_api by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33110](https://github.com/BerriAI/litellm/pull/33110) - test(e2e): cover model-aware mid-conversation system handling on Bedrock Invoke /v1/messages by [@​mateo-berri](https://github.com/mateo-berri) in [#​32963](https://github.com/BerriAI/litellm/pull/32963) - test(claude\_code): move the Claude Code compatibility matrix under tests/e2e by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​32548](https://github.com/BerriAI/litellm/pull/32548) - chore(ci): sync litellm\_internal\_staging into daily OSS branch by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33337](https://github.com/BerriAI/litellm/pull/33337) - feat(bedrock guardrails): add resource-less InvokeGuardrailChecks (detect-only) mode by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33299](https://github.com/BerriAI/litellm/pull/33299) - Revert "chore(ci): sync litellm\_internal\_staging into daily OSS branch" by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33339](https://github.com/BerriAI/litellm/pull/33339) - fix(websearch): intercept web search on the Responses API by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33129](https://github.com/BerriAI/litellm/pull/33129) - fix(anthropic-adapter): drop empty content\_block\_delta events by [@​mateo-berri](https://github.com/mateo-berri) in [#​33315](https://github.com/BerriAI/litellm/pull/33315) - fix(mcp): persist discovered OAuth endpoints and keep last known good on failed re-discovery by [@​tin-berri](https://github.com/tin-berri) in [#​33286](https://github.com/BerriAI/litellm/pull/33286) - test(e2e): otel trace completeness on streaming chat, messages, and responses (LIT-3787) by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33234](https://github.com/BerriAI/litellm/pull/33234) - feat(router): resolve auto-router routing plugins from proxy YAML config by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​33251](https://github.com/BerriAI/litellm/pull/33251) - test(e2e): failed request error span carries the full untruncated message and status (LIT-4179) by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33304](https://github.com/BerriAI/litellm/pull/33304) - refactor(ui): migrate tags table onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33314](https://github.com/BerriAI/litellm/pull/33314) - fix(cli): surface actionable CLI SSO errors when CLI and proxy versions skew by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33309](https://github.com/BerriAI/litellm/pull/33309) - feat(bedrock\_mantle): add GPT-5.6 sol/terra/luna to model cost map by [@​mateo-berri](https://github.com/mateo-berri) in [#​33412](https://github.com/BerriAI/litellm/pull/33412) - chore(codeowners): exempt generated schema.d.ts from UI ownership by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33411](https://github.com/BerriAI/litellm/pull/33411) - feat(proxy): push-based OTLP billable-request metering for enterprise deployments by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​31592](https://github.com/BerriAI/litellm/pull/31592) - fix(mcp): cap per-user OAuth token cache TTL at the token's own lifetime by [@​tin-berri](https://github.com/tin-berri) in [#​33346](https://github.com/BerriAI/litellm/pull/33346) - feat(ui): move Caching out of Experimental into Developer Tools by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33432](https://github.com/BerriAI/litellm/pull/33432) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33425](https://github.com/BerriAI/litellm/pull/33425) - chore(ci): merge daily internal staging branch by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33335](https://github.com/BerriAI/litellm/pull/33335) - feat(guardrails): add Compresr guardrail for query-aware context compression by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33295](https://github.com/BerriAI/litellm/pull/33295) - fix(logging): preserve callback order in get\_combined\_callback\_list by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33005](https://github.com/BerriAI/litellm/pull/33005) - fix(logging): redact assistant tool call arguments in spend logs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33111](https://github.com/BerriAI/litellm/pull/33111) - fix(anthropic): honor messages request timeout by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33418](https://github.com/BerriAI/litellm/pull/33418) - fix(llm\_guard): apply sanitized prompt returned by moderation API to request by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33331](https://github.com/BerriAI/litellm/pull/33331) - fix(logging): stop pinning large request payloads past request end by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33455](https://github.com/BerriAI/litellm/pull/33455) - feat(guardrails): forward optional metadata on POST /guardrails/apply\_guardrail by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33067](https://github.com/BerriAI/litellm/pull/33067) - build: raise requires-python cap to <3.15 so Python 3.14 installs current releases by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33438](https://github.com/BerriAI/litellm/pull/33438) - feat(ui): add reusable BetaBadge and use it for Projects sidebar item by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33449](https://github.com/BerriAI/litellm/pull/33449) - fix(mcp): discover missing OAuth scopes and token\_url when authorization\_url is set manually by [@​tin-berri](https://github.com/tin-berri) in [#​33317](https://github.com/BerriAI/litellm/pull/33317) - feat(ui): show exact license expiration date in usage cards by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33478](https://github.com/BerriAI/litellm/pull/33478) - test(claude\_code): rename misleading REPO\_ROOT to SUITE\_ROOT in test\_v0\_layout by [@​mateo-berri](https://github.com/mateo-berri) in [#​33472](https://github.com/BerriAI/litellm/pull/33472) - build(deps): update ddtrace to the 4.x line by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33484](https://github.com/BerriAI/litellm/pull/33484) - fix(complexity\_router): return empty dict from \_classifier\_call\_metadata when metadata is absent by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33452](https://github.com/BerriAI/litellm/pull/33452) - fix(ui/chat): resolve chat routes at render time so navigation works under server\_root\_path by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33446](https://github.com/BerriAI/litellm/pull/33446) - fix(key management): enforce minimum custom key length and mask short keys in key\_name by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33462](https://github.com/BerriAI/litellm/pull/33462) - chore(ui): remove unmounted UsageIndicator and the Hide Usage Indicator flag by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33482](https://github.com/BerriAI/litellm/pull/33482) - test(e2e/claude\_code): add passthrough matrix row for the big-3 clouds and Anthropic API by [@​mateo-berri](https://github.com/mateo-berri) in [#​33473](https://github.com/BerriAI/litellm/pull/33473) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33491](https://github.com/BerriAI/litellm/pull/33491) - fix(ui): stop sending the complexity-router pseudo-model to /health/test\_connection by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​33498](https://github.com/BerriAI/litellm/pull/33498) - feat(cli): add lite up/down to ambiently route Claude Code through the proxy by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​33231](https://github.com/BerriAI/litellm/pull/33231) - feat(complexity\_router): enable session\_affinity by default by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33500](https://github.com/BerriAI/litellm/pull/33500) - fix(anthropic): stop 500 on combined thinking+signature streaming chunk by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33505](https://github.com/BerriAI/litellm/pull/33505) - feat(autoroute): prompt for semantic keywords per tier in configure wizard by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33508](https://github.com/BerriAI/litellm/pull/33508) - fix(cli/anthropic): unblock lite autoroute proxy deps, adaptive thinking, and thinking+signature streaming by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33507](https://github.com/BerriAI/litellm/pull/33507) - test(ocr): use mistral-document-ai-2512 in azure\_ai OCR tests by [@​mateo-berri](https://github.com/mateo-berri) in [#​33489](https://github.com/BerriAI/litellm/pull/33489) - fix(guardrails): show YAML-defined guardrails in the Guardrail Monitor by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​32853](https://github.com/BerriAI/litellm/pull/32853) - refactor(ui): migrate policies, deleted keys, deleted teams, budgets, and search tools tables onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33357](https://github.com/BerriAI/litellm/pull/33357) - refactor(ui): migrate vector stores, prompts, and skills tables onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33343](https://github.com/BerriAI/litellm/pull/33343) - test(e2e): datadog log delivery for successful chat, messages, and responses (LIT-4447) by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33415](https://github.com/BerriAI/litellm/pull/33415) - fix(cli): make CLI output ASCII-only so it doesn't crash legacy Windows consoles by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33465](https://github.com/BerriAI/litellm/pull/33465) - fix: remove dead user-cache lookup with None key in spend-update path by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33555](https://github.com/BerriAI/litellm/pull/33555) - feat(helm): add per-component PodDisruptionBudget and topologySpreadConstraints to componentized chart by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33430](https://github.com/BerriAI/litellm/pull/33430) - fix(e2e/claude\_code): unblock stage collection, align proxy env names, register compat models by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33433](https://github.com/BerriAI/litellm/pull/33433) - fix(mcp): index authed request-time tools missing from the semantic filter startup index by [@​tin-berri](https://github.com/tin-berri) in [#​33318](https://github.com/BerriAI/litellm/pull/33318) - fix(streaming): use provider-reported usage cost for OpenRouter streams by [@​mateo-berri](https://github.com/mateo-berri) in [#​32255](https://github.com/BerriAI/litellm/pull/32255) - feat(mcp): issuer-anchored OAuth discovery (RFC 8414 §3.3) to close the authorization-server mix-up by [@​tin-berri](https://github.com/tin-berri) in [#​33450](https://github.com/BerriAI/litellm/pull/33450) - feat(logging): add user and team level spend and budget to StandardLoggingPayload metadata by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33459](https://github.com/BerriAI/litellm/pull/33459) - fix(router): cast model\_info cost values to float in \_set\_model\_group\_info by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33556](https://github.com/BerriAI/litellm/pull/33556) - chore(e2e): establish litellm\_e2e\_staging integration line by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33502](https://github.com/BerriAI/litellm/pull/33502) - fix(ui): navigate to /ui/login/ with trailing slash via hard navigation by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33561](https://github.com/BerriAI/litellm/pull/33561) - feat(logging): add structured budget fields to budget rejection failure logs by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33460](https://github.com/BerriAI/litellm/pull/33460) - fix(streaming): surface upstream connection resets instead of empty 200 streams by [@​mateo-berri](https://github.com/mateo-berri) in [#​33222](https://github.com/BerriAI/litellm/pull/33222) - fix(proxy\_cli): reap orphaned prisma query-engine processes when a worker dies by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33424](https://github.com/BerriAI/litellm/pull/33424) - test(reasoning\_effort\_grid): enable azure fable-5 and opus-4-8 grid cells by [@​mateo-berri](https://github.com/mateo-berri) in [#​33485](https://github.com/BerriAI/litellm/pull/33485) - build(deps): bump uvicorn lock to 0.51.0 so worker health-check and jitter flags take effect by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33574](https://github.com/BerriAI/litellm/pull/33574) - fix(proxy): coerce default\_internal\_user\_params.max\_budget to float on config load by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​32434](https://github.com/BerriAI/litellm/pull/32434) - fix(router): honor per-request routing\_strategy from key/team router\_settings by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33429](https://github.com/BerriAI/litellm/pull/33429) - fix(redis): honor ssl value instead of key presence when building async connection pool by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​32590](https://github.com/BerriAI/litellm/pull/32590) - fix(langfuse\_otel): build per-request OTLP exporter from key and team dynamic Langfuse credentials by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​32437](https://github.com/BerriAI/litellm/pull/32437) - ci: run zizmor and proxy-db unit tests on PRs targeting litellm\_ branches by [@​mateo-berri](https://github.com/mateo-berri) in [#​33568](https://github.com/BerriAI/litellm/pull/33568) - fix(router): apply team/key enable\_tag\_filtering to tag routing by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33436](https://github.com/BerriAI/litellm/pull/33436) - feat(proxy): add disable\_auto\_add\_proxy\_admin\_to\_teams flag by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33563](https://github.com/BerriAI/litellm/pull/33563) - fix(proxy): stop stale auth cache re-publish to Redis so key updates propagate across replicas by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33565](https://github.com/BerriAI/litellm/pull/33565) - feat(e2e): emit structured E2E\_RESULT lines for package status history by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33578](https://github.com/BerriAI/litellm/pull/33578) - chore: bump litellm-enterprise 0.1.50 -> 0.1.51, litellm-proxy-extras 0.4.77 -> 0.4.78 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33571](https://github.com/BerriAI/litellm/pull/33571) - fix(docker): restore litellm-proxy-extras source dir in runtime images by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33592](https://github.com/BerriAI/litellm/pull/33592) - feat(ui): require embedding model for semantic auto router by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33313](https://github.com/BerriAI/litellm/pull/33313) - feat(scim): ingest and round-trip SCIM entitlements and roles user attributes by [@​tin-berri](https://github.com/tin-berri) in [#​33587](https://github.com/BerriAI/litellm/pull/33587) - refactor(ui): migrate 5 simple tables onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33548](https://github.com/BerriAI/litellm/pull/33548) - fix(model\_armor): restore reference attachments via skip\_unscannable\_attachments and remove the attachment count cap by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33554](https://github.com/BerriAI/litellm/pull/33554) - fix(mcp): keep the MCP reference intact when the semantic filter narrows tools by [@​tin-berri](https://github.com/tin-berri) in [#​33584](https://github.com/BerriAI/litellm/pull/33584) - fix(sso): stop enforcing UI session budget on CLI login tokens by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33312](https://github.com/BerriAI/litellm/pull/33312) - test: e2e staging leftovers by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33613](https://github.com/BerriAI/litellm/pull/33613) - test(e2e/claude\_code): add GPT-5.6 Sol/Terra/Luna columns for OpenAI, Azure OpenAI, and Bedrock Mantle by [@​mateo-berri](https://github.com/mateo-berri) in [#​33474](https://github.com/BerriAI/litellm/pull/33474) - test(e2e): otel streaming spans record a real ttft below span duration by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33588](https://github.com/BerriAI/litellm/pull/33588) - fix(mcp): make the preemptive-401 OAuth challenge decision mode-aware by [@​tin-berri](https://github.com/tin-berri) in [#​33586](https://github.com/BerriAI/litellm/pull/33586) - fix(ui): show all teams in policy attachment form for admins by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33628](https://github.com/BerriAI/litellm/pull/33628) - refactor(ui): migrate AI Hub, public hub, and MCP Toolsets tables onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33629](https://github.com/BerriAI/litellm/pull/33629) - test(e2e): datadog log delivery for streamed routes, read back from the real datadog api by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33566](https://github.com/BerriAI/litellm/pull/33566) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33640](https://github.com/BerriAI/litellm/pull/33640) - fix(vertex\_ai): surface Gemini grounding toolUsePromptTokenCount in Usage by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33533](https://github.com/BerriAI/litellm/pull/33533) - test(e2e): harness fixes for stage job green (skips + router/UI/budget) by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33634](https://github.com/BerriAI/litellm/pull/33634) - fix(router): resolve prompt cache minimum per model instead of a flat 1024 by [@​tin-berri](https://github.com/tin-berri) in [#​33637](https://github.com/BerriAI/litellm/pull/33637) - fix(logging): classify async anthropic\_messages and generate\_content as async by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33589](https://github.com/BerriAI/litellm/pull/33589) - fix(ui): remove Chat item from dashboard leftnav by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33647](https://github.com/BerriAI/litellm/pull/33647) - fix(router): tag-aware pre-routing strategy selection for shared model\_name by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33691](https://github.com/BerriAI/litellm/pull/33691) - fix(proxy): enforce max\_parallel\_requests as a per-slot concurrency gauge by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​32441](https://github.com/BerriAI/litellm/pull/32441) - fix(proxy): stop treating upstream model body field as a LiteLLM model on auth-enforced pass-through routes by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33710](https://github.com/BerriAI/litellm/pull/33710) - fix(mcp): expand toolset grants in shared permission primitives so tools/call honors them by [@​tin-berri](https://github.com/tin-berri) in [#​33612](https://github.com/BerriAI/litellm/pull/33612) - feat(complexity-router): user-triggered escalation keywords by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33656](https://github.com/BerriAI/litellm/pull/33656) - fix(fireworks\_ai): bill prompt-cache hits at cache\_read rate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33714](https://github.com/BerriAI/litellm/pull/33714) - fix(pricing): mark realtime-only gpt-realtime models as mode realtime by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33728](https://github.com/BerriAI/litellm/pull/33728) - fix(rag): track LLM completion usage and spend for /v1/rag/query by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​32438](https://github.com/BerriAI/litellm/pull/32438) - feat(anthropic): add enable\_anthropic\_prompt\_caching for automatic cache\_control injection by [@​tin-berri](https://github.com/tin-berri) in [#​33573](https://github.com/BerriAI/litellm/pull/33573) - fix(anthropic): self-heal on missing thinking-signature errors from Bedrock/Vertex by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33719](https://github.com/BerriAI/litellm/pull/33719) - fix(proxy): resolve router\_settings.plugins dotted paths and load plugins from installed packages by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33644](https://github.com/BerriAI/litellm/pull/33644) - test(e2e): budget refusals are 429 for bare keys and team caps block every team key by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33632](https://github.com/BerriAI/litellm/pull/33632) - feat(router): add router plugin reference catalog by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33746](https://github.com/BerriAI/litellm/pull/33746) - test(e2e): assert an org budget block is a 429 naming the organization by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33638](https://github.com/BerriAI/litellm/pull/33638) - fix(proxy): bill partial streamed spend when the client disconnects mid-stream by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33736](https://github.com/BerriAI/litellm/pull/33736) - test(e2e): delete unreferenced Grafana panel docs by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33743](https://github.com/BerriAI/litellm/pull/33743) - docs(tests/e2e): align skip-vs-fail docs with the hard-fail contract and scope the no-unit-tests rule by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33755](https://github.com/BerriAI/litellm/pull/33755) - refactor(e2e): replace bespoke result reporter with standard JUnit report by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33758](https://github.com/BerriAI/litellm/pull/33758) - test(e2e): user budget across keys and team member budget isolation by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33745](https://github.com/BerriAI/litellm/pull/33745) - refactor(e2e): remove bob\_the\_builder; drive remediation from a Grafana alert (provisioned outside the repo) by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33749](https://github.com/BerriAI/litellm/pull/33749) - feat(mcp): per-server outcomes for aggregate tools/list and truthful single-server REST statuses by [@​tin-berri](https://github.com/tin-berri) in [#​33153](https://github.com/BerriAI/litellm/pull/33153) - test(e2e): mcp suite for key-without-access denial by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33752](https://github.com/BerriAI/litellm/pull/33752) - chore(ci): merge oss branch by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33784](https://github.com/BerriAI/litellm/pull/33784) - chore(ci): merge oss branch - July 17th by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33793](https://github.com/BerriAI/litellm/pull/33793) - fix(ui): migrate tag deletion to shared DeleteResourceModal by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33795](https://github.com/BerriAI/litellm/pull/33795) - build(rust): raise pyo3 to 0.29 so the native bridge compiles on Python 3.14 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33798](https://github.com/BerriAI/litellm/pull/33798) - chore(guardrails): remove docstring from singulr module for consistency by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33800](https://github.com/BerriAI/litellm/pull/33800) - fix(ui): stop credential edit from persisting the masked api key by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33797](https://github.com/BerriAI/litellm/pull/33797) - build(deps): allow redisvl, pypdf, and openapi-core on Python 3.14 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33801](https://github.com/BerriAI/litellm/pull/33801) - test(proxy): make streaming-cancel mocks awaitable for the disconnect slot release by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33802](https://github.com/BerriAI/litellm/pull/33802) - test(e2e): a member's team budget cuts off only that member's key by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33718](https://github.com/BerriAI/litellm/pull/33718) - test(e2e): a user's max\_budget follows the person across personal and team keys by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33762](https://github.com/BerriAI/litellm/pull/33762) - test(e2e): skip flaky OpenAI GPT cells; raise multi-window max\_tokens by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33799](https://github.com/BerriAI/litellm/pull/33799) - chore: remove accidentally committed dist tarball and ignore dist/ by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33805](https://github.com/BerriAI/litellm/pull/33805) - fix(passthrough): stop classifying plain 'predict'/'search' paths as Vertex by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33658](https://github.com/BerriAI/litellm/pull/33658) - build(deps): bump mcp lock to 1.28.1 to clear image-scan findings by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33803](https://github.com/BerriAI/litellm/pull/33803) - test(pricing): pin the realtime mode assertion to the bundled cost map by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33806](https://github.com/BerriAI/litellm/pull/33806) - fix(proxy): derive session id from Anthropic metadata.user\_id for session affinity by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33723](https://github.com/BerriAI/litellm/pull/33723) - test(e2e): budget reset diagonal for team, org, user, and [#​32005](https://github.com/BerriAI/litellm/issues/32005) team-member keys by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33771](https://github.com/BerriAI/litellm/pull/33771) - fix(proxy): source /v1/models token limits from the cost map instead of Router.get\_model\_group\_info by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33721](https://github.com/BerriAI/litellm/pull/33721) - fix(fireworks\_ai): correct glm-5p2 prompt-cache read price to $0.14/1M by [@​tin-berri](https://github.com/tin-berri) in [#​33796](https://github.com/BerriAI/litellm/pull/33796) - feat(proxy): add x-litellm-model-name response header with deployment model string by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33698](https://github.com/BerriAI/litellm/pull/33698) - feat: add Straiker guardrail integration by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33781](https://github.com/BerriAI/litellm/pull/33781) - fix(vertex\_ai): exclude Gemini Google Search grounding tokens from input token billing by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33742](https://github.com/BerriAI/litellm/pull/33742) - feat(fireworks\_ai): map litellm session id to x-session-affinity header for prompt caching by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33717](https://github.com/BerriAI/litellm/pull/33717) - feat(ui): configure Anthropic automatic prompt caching from the Admin UI by [@​tin-berri](https://github.com/tin-berri) in [#​33581](https://github.com/BerriAI/litellm/pull/33581) - fix(router): enforce context-window pre-call checks for Responses API input by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33706](https://github.com/BerriAI/litellm/pull/33706) - fix(otel): restore proxy-level error.\* attributes on v2 failure spans (LIT-4179) by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33664](https://github.com/BerriAI/litellm/pull/33664) - fix(mcp): persist config.yaml DCR clients in a server-scoped store so refresh survives token expiry by [@​tin-berri](https://github.com/tin-berri) in [#​33768](https://github.com/BerriAI/litellm/pull/33768) - refactor(ui): consolidate Add/Edit credential modals into one CredentialModal by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32572](https://github.com/BerriAI/litellm/pull/32572) - feat(mcp): add ID-JAG (identity assertion authorization grant) support for MCP egress by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​31516](https://github.com/BerriAI/litellm/pull/31516) - refactor(ui): migrate policy attachments table onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33827](https://github.com/BerriAI/litellm/pull/33827) - docs(litellm-rust): add provider coding standards by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33833](https://github.com/BerriAI/litellm/pull/33833) - test(e2e): rename Gateway to ProxyClient and expose it as a session-scoped pytest fixture by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33750](https://github.com/BerriAI/litellm/pull/33750) - feat(messages): route Azure Anthropic /messages through Rust behind rust:true by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33616](https://github.com/BerriAI/litellm/pull/33616) - test(e2e): add Locust throughput load test that runs last by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33748](https://github.com/BerriAI/litellm/pull/33748) - fix(proxy): resolve team wildcard credentials for vector store files by [@​shivamrawat1](https://github.com/shivamrawat1) in [#​33649](https://github.com/BerriAI/litellm/pull/33649) - refactor(e2e): fold claude\_code HTTP probes onto shared ProxyClient methods by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33760](https://github.com/BerriAI/litellm/pull/33760) - test(e2e): harden stage flakes for batches, UI, and MCP by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33831](https://github.com/BerriAI/litellm/pull/33831) - fix(e2e): migrate load suite from e2e\_gateway to ProxyClient by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33839](https://github.com/BerriAI/litellm/pull/33839) - chore(e2e): remove tests/e2e/docker-compose.yml by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33837](https://github.com/BerriAI/litellm/pull/33837) - test(e2e): cover /v1/responses openai basic nonstream and stream by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33830](https://github.com/BerriAI/litellm/pull/33830) - test(e2e): cover /v1/responses openai cost\_logged and tool\_use by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33835](https://github.com/BerriAI/litellm/pull/33835) - test(e2e): cover /v1/responses OpenAI vision and Anthropic basic by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33838](https://github.com/BerriAI/litellm/pull/33838) - test(e2e): spendlog cost for streaming /v1/messages via responses bridge by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33753](https://github.com/BerriAI/litellm/pull/33753) - fix(docker): bake prisma CLI and engines at a fixed path so fresh-DB migrations work for any uid offline by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33853](https://github.com/BerriAI/litellm/pull/33853) - feat(chat-ui): add personal Logs view scoped to the current user by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33829](https://github.com/BerriAI/litellm/pull/33829) - chore: bump litellm-proxy-extras 0.4.78 -> 0.4.79 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33855](https://github.com/BerriAI/litellm/pull/33855) - docs(litellm-rust): require the official Rust Style Guide in agent rules by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33867](https://github.com/BerriAI/litellm/pull/33867) - fix(router): treat malformed configured token limits as absent on /v1/models by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33864](https://github.com/BerriAI/litellm/pull/33864) - chore: rebuild admin UI bundle for the rc release by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33857](https://github.com/BerriAI/litellm/pull/33857) - docs(rust): add provider abstraction standards by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33865](https://github.com/BerriAI/litellm/pull/33865) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33868](https://github.com/BerriAI/litellm/pull/33868) - chore(release): backport [#​33929](https://github.com/BerriAI/litellm/issues/33929) to rc/1.94.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34033](https://github.com/BerriAI/litellm/pull/34033) - chore(release): backport [#​33810](https://github.com/BerriAI/litellm/issues/33810), [#​33733](https://github.com/BerriAI/litellm/issues/33733) to rc/1.94.0 and bump litellm-proxy-extras to 0.4.79.post1 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34215](https://github.com/BerriAI/litellm/pull/34215) - chore(ui): rebuild Next.js bundle on rc/1.94.0 so the Cost Optimization page ships by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34216](https://github.com/BerriAI/litellm/pull/34216) - chore(release): backport auth, CLI SSO and guardrail fixes to rc/1.94.0 and refresh flagged dependencies by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34640](https://github.com/BerriAI/litellm/pull/34640) - chore(release): backport [#​33899](https://github.com/BerriAI/litellm/issues/33899), [#​33978](https://github.com/BerriAI/litellm/issues/33978), [#​34582](https://github.com/BerriAI/litellm/issues/34582), [#​34675](https://github.com/BerriAI/litellm/issues/34675) to rc/1.94.0 and bump litellm-proxy-extras to 0.4.79.post2 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34855](https://github.com/BerriAI/litellm/pull/34855) - fix(ui): backport cache leakage card layout fix to rc/1.94.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34964](https://github.com/BerriAI/litellm/pull/34964) - fix(ui): add missing cost-optimization page description on rc/1.94.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34967](https://github.com/BerriAI/litellm/pull/34967) - feat(ui): mark Cost Optimization as beta in the left nav ([#​34984](https://github.com/BerriAI/litellm/issues/34984)) by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34987](https://github.com/BerriAI/litellm/pull/34987) - chore: rebuild Admin UI bundle for v1.94.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34982](https://github.com/BerriAI/litellm/pull/34982) - fix(cost-optimization): backport the savings chart axis fix and methodology popovers to rc/1.94.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34994](https://github.com/BerriAI/litellm/pull/34994) - chore: rebuild Admin UI bundle for rc/1.94.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34995](https://github.com/BerriAI/litellm/pull/34995) **Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.93.0...v1.94.0> ### [`v1.94.0`](https://github.com/BerriAI/litellm/releases/tag/v1.94.0) ##### Verify Docker Image Signature All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0). **Verify using the pinned commit hash (recommended):** A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \ ghcr.io/berriai/litellm:v1.94.0 ``` **Verify using the release tag (convenience):** Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/v1.94.0/cosign.pub \ ghcr.io/berriai/litellm:v1.94.0 ``` Expected output: ``` The following checks were performed on each of these signatures: - The cosign claims were validated - The signatures were verified against the specified public key ``` *** ##### What's Changed - feat(ui): working Test Connection for the complexity auto router by [@​akapur99](https://github.com/akapur99) in [#​32950](https://github.com/BerriAI/litellm/pull/32950) - fix(xecguard): use StandardLoggingGuardrailInformation in logging hook by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​32911](https://github.com/BerriAI/litellm/pull/32911) - feat(ui): adopt openapi-react-query ($api) and convert useCustomers by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32949](https://github.com/BerriAI/litellm/pull/32949) - refactor(ui): colocate the mcp-servers view, keeping the shared mcp\_tools surface by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32968](https://github.com/BerriAI/litellm/pull/32968) - refactor(ui): convert endpoint usage charts to shadcn/recharts by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32723](https://github.com/BerriAI/litellm/pull/32723) - fix(proxy-auth): stop unrecognized model namespaces slipping through provider wildcard keys by [@​mateo-berri](https://github.com/mateo-berri) in [#​32979](https://github.com/BerriAI/litellm/pull/32979) - feat(router): random-pick multi-model complexity tiers by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32967](https://github.com/BerriAI/litellm/pull/32967) - fix(xecguard): sanitize scan result before recording it for logging by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​32935](https://github.com/BerriAI/litellm/pull/32935) - fix(auto\_router): filter embedding models in complexity tab dropdowns, require all tiers, inline validation by [@​akapur99](https://github.com/akapur99) in [#​32978](https://github.com/BerriAI/litellm/pull/32978) - fix(anthropic): translate raw adaptive thinking for pre-4.6 models on chat completions and Bedrock Converse by [@​akapur99](https://github.com/akapur99) in [#​32944](https://github.com/BerriAI/litellm/pull/32944) - feat(router): add Router(plugins=\[...]) routing-plugin pipeline by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32972](https://github.com/BerriAI/litellm/pull/32972) - feat(router): soft-floor adaptive mode for complexity router by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32947](https://github.com/BerriAI/litellm/pull/32947) - docs(github): add QA runbook section to the PR template by [@​mateo-berri](https://github.com/mateo-berri) in [#​32965](https://github.com/BerriAI/litellm/pull/32965) - fix(model\_cost): add supports\_reasoning: false to Gemini image generation models by [@​mateo-berri](https://github.com/mateo-berri) in [#​32836](https://github.com/BerriAI/litellm/pull/32836) - build(dev-env): add make bootstrap and unprovisioned-checkout preflight to pre-commit by [@​mateo-berri](https://github.com/mateo-berri) in [#​32981](https://github.com/BerriAI/litellm/pull/32981) - ci(ui): report only error-level knip findings in CI by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32971](https://github.com/BerriAI/litellm/pull/32971) - feat(batches): track cost for unmanaged Bedrock batches, generalize the flag by [@​Sameerlite](https://github.com/Sameerlite) in [#​32315](https://github.com/BerriAI/litellm/pull/32315) - fix(guardrails): walk custom\_tool\_call\_output items in \_content\_utils by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​32969](https://github.com/BerriAI/litellm/pull/32969) - fix: show and allow editing team model aliases after team creation by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33047](https://github.com/BerriAI/litellm/pull/33047) - chore(deps): bump pillow to 12.3.0 to resolve osv-scan CVEs by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33093](https://github.com/BerriAI/litellm/pull/33093) - feat(mcp): mint gateway-bound envelope at the token endpoint for dcr\_bridge oauth\_delegate by [@​tin-berri](https://github.com/tin-berri) in [#​32828](https://github.com/BerriAI/litellm/pull/32828) - fix(mcp): surface rejected delegate-auth upstream tokens as connect-time 401 by [@​tin-berri](https://github.com/tin-berri) in [#​32741](https://github.com/BerriAI/litellm/pull/32741) - fix(proxy): track unauthenticated pass-through requests in spend logs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​32410](https://github.com/BerriAI/litellm/pull/32410) - feat(lasso): send source.type for Used By attribution by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33090](https://github.com/BerriAI/litellm/pull/33090) - fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@​thibault-linktree](https://github.com/thibault-linktree) in [#​33025](https://github.com/BerriAI/litellm/pull/33025) - fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@​tin-berri](https://github.com/tin-berri) in [#​33099](https://github.com/BerriAI/litellm/pull/33099) - fix(completion): forward aws credential kwargs into litellm\_params so the responses bridge keeps WIF auth by [@​mateo-berri](https://github.com/mateo-berri) in [#​32956](https://github.com/BerriAI/litellm/pull/32956) - fix(ui): respect litellm\_key\_header\_name in BYOK credential save and workflow runs fetches by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33103](https://github.com/BerriAI/litellm/pull/33103) - refactor(ui): standardize debounce waits behind shared DEBOUNCE\_WAIT\_MS constant by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33040](https://github.com/BerriAI/litellm/pull/33040) - feat(ui): rebuild the Virtual Keys table on the shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​32991](https://github.com/BerriAI/litellm/pull/32991) - fix: redact async complete streaming response for custom callbacks by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33106](https://github.com/BerriAI/litellm/pull/33106) - build(ui): bump [@​tanstack/react-pacer](https://github.com/tanstack/react-pacer) from 0.2.0 to 0.22.1 by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33041](https://github.com/BerriAI/litellm/pull/33041) - fix(ui): address Virtual Keys redesign review nits by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33112](https://github.com/BerriAI/litellm/pull/33112) - fix(openai/responses): clamp max\_output\_tokens below API minimum by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33098](https://github.com/BerriAI/litellm/pull/33098) - fix(prometheus): read v3 rate limiter remaining values for per-key model gauges by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33119](https://github.com/BerriAI/litellm/pull/33119) - fix(ui): drop w-full from page-content wrappers to remove 32px horizontal overflow by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33118](https://github.com/BerriAI/litellm/pull/33118) - refactor(ui): migrate straightforward value debounces to react-pacer by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33042](https://github.com/BerriAI/litellm/pull/33042) - feat(mcp): interactive SSO sign-in for dcr\_bridge oauth\_delegate DCR clients by [@​tin-berri](https://github.com/tin-berri) in [#​32946](https://github.com/BerriAI/litellm/pull/32946) - test(proxy): add regression tests for management\_endpoints edge cases by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​32976](https://github.com/BerriAI/litellm/pull/32976) - fix(auto-router): correct Responses API tool\_choice shape and propagate alias litellm\_params by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32974](https://github.com/BerriAI/litellm/pull/32974) - fix(ui): render the sidebar scrollbar with shadcn ScrollArea by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33124](https://github.com/BerriAI/litellm/pull/33124) - refactor(ui): migrate callback debounce sites to react-pacer with regression tests by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33043](https://github.com/BerriAI/litellm/pull/33043) - chore: add CODEOWNERS for ui and proxy UI build artifacts by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33131](https://github.com/BerriAI/litellm/pull/33131) - feat(mcp): client-held refresh envelope for the dcr\_bridge oauth\_delegate flow by [@​tin-berri](https://github.com/tin-berri) in [#​32980](https://github.com/BerriAI/litellm/pull/32980) - feat(ui): rebuild the Teams table on the shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33128](https://github.com/BerriAI/litellm/pull/33128) - fix(mcp): relay upstream OAuth token and DCR rejections instead of a generic 500 by [@​tin-berri](https://github.com/tin-berri) in [#​33113](https://github.com/BerriAI/litellm/pull/33113) - fix(keys): persist key\_type so the UI shows correct key scope instead of "All Proxy…
Relevant issues
Discussion: #32168
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
@greptileaito re-request a review after pushing changes)make pre-commitcould not complete in the environment this PR was built in:uv syncfails buildinggrpcio==1.78.0(no wheel for macOS x86_64 at that pin) before it reaches basedpyright, and the dashboard node_modules aren't provisioned either, both unrelated to this diff.ruff check .passes clean against the changed files. CI runs in a correctly provisioned environment so its lint job is the authoritative check here.Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Commit under test:
8d5e033929Router(plugins=[...])(merged separately, discussion #32168) only accepted pre-instantiated plugin objects through the Python constructor, and narrowed candidates from the outer model alias rather than the auto-router's real tier pool, so it was unusable through proxy config and a no-op forauto_router/complexity_routerdeployments even from Python. This PR addscomplexity_router_config.plugins, a list of dotted-path strings resolved the same waylitellm_settings.callbacksalready is, and wires the resolved plugins into the tier pool thatget_model_for_tieractually picks from.Config used below (
complexity_router_config.tiers.COMPLEX: [gpt-4o, gpt-4o-mini],default_model: gpt-4o-mini), with aCostCeilingPluginthat drops any candidate whose real per-token input cost exceeds0.000001(gpt-4o-mini=1.5e-07,gpt-4o=2.5e-06, fromlitellm/model_prices_and_context_window_backup.json). All requests below are realopenai/gpt-4oandopenai/gpt-4o-minicalls against the live proxy, billed against a real OpenAI key; the routed model is read offx-litellm-response-costsince gpt-4o-mini's cost (9.45e-06) and gpt-4o's cost (0.0001575) are only reachable by one deployment each here.Before,
config.yaml(nopluginskey incomplexity_router_config):8 identical COMPLEX-tier requests against that config:
gpt-4o (over the would-be ceiling) gets picked 5 of 8 times, as expected from an unfiltered pool.
After,
config.yaml(same config pluscomplexity_router_config.plugins):plugins.cost_ceiling_plugin.cost_ceiling_pluginis a dotted path resolved relative toconfig.yaml's own directory, i.e. a siblingplugins/cost_ceiling_plugin.pyfile.5 identical COMPLEX-tier requests against that config:
gpt-4o never gets picked once the plugin is configured; every request lands on gpt-4o-mini, the only candidate the plugin leaves in the tier pool.
plugins/cost_ceiling_plugin.pyused for both runs:Type
🆕 New Feature
🐛 Bug Fix
Changes
complexity_router_configgains aplugins: list[str] | Nonefield (dotted-path refs, e.g.mypackage.mymodule.my_plugin_instance). The proxy config loader resolves each path through the existingget_instance_fn(same helperlitellm_settings.callbacksand custom guardrails already use) beforeRouter.__init__runs, soComplexityRouterConfig.pluginsends up holding liveRoutingPlugininstances.ComplexityRouter._pick_model_for_tierbuilds aRoutingContextscoped to the classified tier's actual candidate pool (not the outerauto_router/complexity_routeralias), runs the configured plugins in order, and picks from whatever candidates survive. It's wired into all three model-pick sites inside_classify_and_route(the weighted-scoring path, thekeyword_tier_rulesoverride path, and the no-user-message default-tier path) so a policy plugin can't be bypassed by any of those paths. A tier narrowed to zero candidates raises rather than falling back todefault_model, sincedefault_modelwas never checked against the plugins and would otherwise function as an unconditional escape hatch around whatever policy a plugin enforces.adaptive=Truecombined withpluginsraises at config-validation time, since the bandit selector in_soft_floor_pickdoesn't consume plugin-narrowed pools yet.Also fixes
Router._generate_model_id, which builds a deployment id byjson.dumps-ing everylitellm_paramsvalue; it crashed withTypeError: Object ... is not JSON serializablethe moment a resolved plugin instance landed insidecomplexity_router_config. Usesjson.dumps(v, default=...)with a fallback that returns the value's fully-qualified class name. (An earlier revision of this PR useddefault=str, which Greptile correctly flagged:str()on an arbitrary object falls back toobject.__repr__'s<module.Class object at 0x...>, embedding the memory address and making the deployment id change on every hot-reload for any deployment withpluginsconfigured. Fixed in a follow-up commit.)Three policy-bypass gaps flagged by Veria AI's security review, all fixed in follow-up commits: the
session_affinitycache-pin shortcut returned a session's first-turn model on every later turn without ever re-running it through the plugin pipeline, so a plugin's decision (e.g. a budget cap crossed mid-session) was only enforced on turn one -- the pin shortcut is now disabled wheneverpluginsare configured. The no-user-message path usedself.config.default_model or await self._pick_model_for_tier(...), and Python'sorshort-circuits on a truthydefault_model-- the plugin pipeline never ran at all wheneverdefault_modelwas configured; that short-circuit is removed, falling through to the same tier-then-default_model resolution every other call site already uses. Separately,get_instance_fnaccepts any dotted path and returns whatever object it finds there, so a misconfiguredpluginsentry passed proxy startup silently and only surfaced as a confusingAttributeErroron the first request that reached the plugin pipeline -- extracted the resolution logic intoresolve_complexity_router_plugins()with anisinstance(..., RoutingPlugin)check that now fails proxy startup immediately with a clear error instead.Final Attestation