Skip to content

fix(langfuse): stop a collected httpx handler from closing a shared client - #35981

Merged
yassin-berriai merged 1 commit into
litellm_internal_stagingfrom
litellm_langfuse_shared_client_close
Aug 5, 2026
Merged

fix(langfuse): stop a collected httpx handler from closing a shared client#35981
yassin-berriai merged 1 commit into
litellm_internal_stagingfrom
litellm_langfuse_shared_client_close

Conversation

@yassin-berriai

@yassin-berriai yassin-berriai commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Langfuse traces stop silently until restart, no error surfaced
  • A collected HTTP handler closed a client other holders still used
  • Hits per-key and per-team Langfuse credentials

How it solves it:

  • A handler finalizer closes only a client it solely referred to
  • Pooled sockets still released promptly, measured against base
  • close() respects _owns_client, so injected clients survive

Relevant issues

Linear ticket

Resolves LIT-5228

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • My PR passes all CI/CD checks (e.g., lint, format, unit tests)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live proxy on port 15228, real gpt-4o-mini calls billed to a real key, real HTTP ingestion. There are no Langfuse Cloud credentials on this machine, so a local HTTP server on port 15229 stands in for the Langfuse ingest endpoint and prints every event it receives. Everything else is the real path: the real Langfuse SDK and its real background flush thread posting to /api/public/ingestion over the client under test

The bug needs the shared client cache entry to go away, which normally happens on the 1 hour TTL or under the 200 entry size cap. Neither is drivable from curl, so evict_callback.py below is loaded as a proxy callback and forces exactly that event on demand when a request carries "user": "lit5228-evict"

Setup, identical for both runs:

# evict_callback.py
import gc

import litellm
from litellm.caching.llm_caching_handler import LLMClientCache
from litellm.integrations.custom_logger import CustomLogger


class CacheEvictor(CustomLogger):
    async def async_pre_call_hook(self, user_api_key_dict, cache, data, call_type):
        if "lit5228-evict" in str(data.get("user", "")):
            litellm.in_memory_llm_clients_cache = LLMClientCache()
            print(f"[LIT-5228] evicted cache, gc.collect() reclaimed {gc.collect()} objects", flush=True)
        return data


cache_evictor_instance = CacheEvictor()
# config.yaml
model_list:
  - model_name: gpt-4o-mini
    litellm_params:
      model: openai/gpt-4o-mini
      api_key: os.environ/OPENAI_API_KEY

litellm_settings:
  success_callback: ["langfuse"]
  callbacks: evict_callback.cache_evictor_instance

general_settings:
  master_key: sk-lit5228
  allow_client_side_credentials: true
  1. Start the ingest stand-in, then the proxy
python langfuse_sink.py 15229 &
python litellm/proxy/proxy_cli.py --config config.yaml --port 15228 --num_workers 1
  1. Send a request carrying per-tenant Langfuse credentials
curl -s http://127.0.0.1:15228/v1/chat/completions -H 'Authorization: Bearer sk-lit5228' -H 'Content-Type: application/json' \
 -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"say LIT5228-PROXY-BEFORE"}],"max_tokens":10,
      "langfuse_public_key":"pk-lf-lit5228-tenant","langfuse_secret":"sk-lf-lit5228-tenant","langfuse_host":"http://127.0.0.1:15229",
      "metadata":{"generation_name":"LIT5228-PROXY-BEFORE"}}'
  1. Expire the shared client cache entry
curl -s http://127.0.0.1:15228/v1/chat/completions -H 'Authorization: Bearer sk-lit5228' -H 'Content-Type: application/json' \
 -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"evict"}],"max_tokens":5,"user":"lit5228-evict"}'
  1. Send the same tenant request again, with LIT5228-PROXY-AFTER

Before, at 2792887e47d698edaeb8a2d0ad1abfc9576609a6

Steps 2 and 4 both return 200 and the caller sees nothing wrong. An instrumented build shows the tenant logger's ingestion client dying across step 3:

[LIT-5228] before evict: LangFuseLogger client id=4924741584 is_closed=False
[LIT-5228] evicted cache, gc.collect() reclaimed 204 objects
[LIT-5228] after evict:  LangFuseLogger client id=4924741584 is_closed=True

Full ingest log for the run:

SINK RECEIVED trace-create name=litellm-acompletion
SINK RECEIVED generation-create name=LIT5228-PROXY-BEFORE
SINK RECEIVED trace-create name=litellm-acompletion
SINK RECEIVED generation-create name=litellm-acompletion

LIT5228-PROXY-AFTER never arrives. The two middle lines are the eviction request itself, logged through the global callback, which uses a different client

After, at 9478d25b63

Same four steps, same proxy, same real OpenAI calls. Step 3 now reports the client staying open:

[LIT-5228] evicted cache, gc.collect() reclaimed 204 objects
[LIT-5228] after evict: LangFuseLogger client id=4950011152 is_closed=False

Full ingest log for the run:

SINK RECEIVED trace-create name=litellm-acompletion
SINK RECEIVED generation-create name=LIT5228-PROXY-BEFORE
SINK RECEIVED trace-create name=litellm-acompletion
SINK RECEIVED generation-create name=litellm-acompletion
SINK RECEIVED trace-create name=litellm-acompletion
SINK RECEIVED generation-create name=LIT5228-PROXY-AFTER

LIT5228-PROXY-AFTER now lands

Pooled socket reclamation, before and after

Touching a finalizer that closes clients risks trading a silent logging outage for a socket leak, so this is measured rather than argued. The first version of this harness could not have found anything: its test server sent Connection: close, so nothing was ever pooled and both arms reported a flat sockets=1, which is indistinguishable from a clean result. A harness that cannot observe the resource it is measuring reports success. With keep-alive the difference was immediate, and that is why the finalizer is guarded here rather than deleted. 2000 handler create-request-drop cycles against a local keep-alive server, no forced gc.collect() at any point, sockets counted with psutil, live objects counted with gc.get_objects(). Both runs print litellm.__file__ so the tree under test is on the record

[BASE]        start  sockets=1    httpx.Client=0  HTTPHandler=0  ConnectionPool=0
[BASE]   cycle 1000  sockets=1    httpx.Client=0  HTTPHandler=0  ConnectionPool=8
[BASE]   cycle 2000  sockets=1    httpx.Client=0  HTTPHandler=0  ConnectionPool=15

[BRANCH]      start  sockets=1    httpx.Client=0  HTTPHandler=0  ConnectionPool=0
[BRANCH] cycle 1000  sockets=1    httpx.Client=0  HTTPHandler=0  ConnectionPool=15
[BRANCH] cycle 2000  sockets=1    httpx.Client=0  HTTPHandler=0  ConnectionPool=11

Sockets stay pinned at 1, the listening socket, on both. Live pool counts oscillate in the same bounded band on both. An earlier revision of this change dropped the finalizer outright and that showed real drift, branch sockets climbing to 29 while base held at 1, which is why the sole-referrer guard exists instead

Repeated LangFuseLogger() construction, the shape at _health_endpoints.py:291, also tracks base exactly: 40 constructions leave 1 live httpx.Client on both builds

Why the finalizer has to stay at all, rather than trusting refcounting. The handler has no reference cycle and is reclaimed by refcounting, but on the httpcore transport the client's connection pool IS cyclic, so dropping the client does not return its file descriptor. Server in a separate process so only the client side is counted, gc disabled, handler confirmed unreferenced at the point of measurement:

AsyncHTTPHandler, httpcore transport, WITH the guarded finalizer
  baseline                          EST=0  ALL=0  FDS=9
  after request                     EST=1  ALL=1  FDS=10
  after `del h`, gc OFF, no yield   EST=1  ALL=1  FDS=10   pool alive? True
  after a loop turn, gc still OFF   EST=0  ALL=0  FDS=9    <- finalizer released it, no gc involved
  after gc.collect()                EST=0  ALL=0  FDS=9

same handler, same transport, with the guard forced to never fire
  after `del h`, gc OFF, no yield   EST=1  ALL=1  FDS=10   pool alive? True
  after a loop turn, gc still OFF   EST=1  ALL=1  FDS=10   <- still held
  after gc.collect()                EST=0  ALL=0  FDS=9    <- only the cycle collector frees it

That pair is the argument for keeping a guarded finalizer instead of deleting one: with it, the descriptor comes back without any collection pass; without it, the descriptor waits for the cycle collector

This is transport-dependent and the PR should not overstate it. Under litellm's default async aiohttp transport the session is NOT cyclic, dies by refcount, and the descriptor returns on the next loop turn whether or not the finalizer fires. The cyclic case is httpcore, which is every HTTPHandler (there is no sync aiohttp transport) and any AsyncHTTPHandler with aiohttp disabled. HTTPHandler is the one Langfuse uses and the one the 2000-cycle measurement above exercises, so the finalizer is load-bearing exactly where this change touches

These are two different experiments and are not interchangeable: the 2000-cycle keep-alive numbers are about REMOVING the finalizer under handler churn, and the transcripts here are about whether a single dropped handler returns its descriptor. Both matter, neither is evidence for the other

One related change on the async side, and it is an improvement on the previous behaviour rather than a restoration of it: the finalizer now schedules self._client.aclose() rather than self.close(). The old form handed self to a coroutine mid-collection, which resurrected the handler and delayed its reclamation by a loop turn. Closing over the client instead gets both properties, the handler is reclaimed immediately by refcounting and the client it owned is still torn down

Known limitation

_CLIENT_REFCOUNT_WHEN_HANDLER_IS_SOLE_REFERRER is a CPython reference-counting assumption. requires-python allows <3.15, so free-threaded 3.13t and 3.14t are in scope, and sys.getrefcount is not meaningful under deferred reference counting there. The expectation is that it reads high and so fails safe, declining to close and leaving the client to the collector rather than closing one that is still in use, which is the same behaviour as the guard being conservative. That is reasoning, not a measurement, and this is untested on a free-threaded build

The constant is also sensitive to how it is read: binding the client to a local before the call adds a reference and reads 3 instead of 2, which would silently stop the guard from ever firing. Both finalizers are covered by a test that fails in exactly that case, so the drift cannot land unnoticed

Type

🐛 Bug Fix

Changes

_get_httpx_client() and get_async_httpx_client() return a handler cached process-wide in litellm.in_memory_llm_clients_cache, and at the Langfuse call site the handler itself was a local that went out of scope immediately. A consumer that takes handler.client and keeps the raw httpx.Client therefore outlives the handler. Once the cache entry expires on its TTL or is evicted under the 200 entry cap, nothing references the handler, it is collected, and its __del__ called close() on the client the consumer is still using

A finalizer firing proves only that nothing references the handler. It proves nothing about the client, which may be held by anyone the handler gave it to. Both handlers now close during finalization only when they built the client and are still its sole referrer, read from the refcount at the call site. That preserves prompt teardown for the ordinary unshared case, which is what the finalizer was actually buying, while a handed-out client is left for its real owners. LLMClientCache's own docstring already promised this: "This cache intentionally does NOT close clients on eviction ... Clients that are no longer referenced will be garbage-collected normally"

Explicit close() stays, that is a caller deliberately shutting a handler down, but it now honours _owns_client. The setter and the client= constructor argument both record that a caller supplied the client, and a wrapper should not close something it was handed. __aexit__ routes through close() so the two agree

The second half is the Langfuse call site keeping a reference to the handler that owns the client it hands the SDK, so the handler cannot be collected out from under it while the logger lives. It still takes the client from the shared cache, so no additional clients are created per logger. This is defence in depth rather than the load-bearing fix: the sole-referrer guard alone stops the outage, and this makes the Langfuse client's liveness independent of that guard being right

Which half stops the outage: the finalizer guard. Which half is hardening: the Langfuse reference

Reachability

LangFuseLogger has four constructors in-tree and they are not equally exposed. success_callback: ["langfuse"] resolves through custom_logger_registry.py to LangfusePromptManagement, which builds its own standalone client and never touches the cached handler, so that object is not affected. Instrumenting the constructor on a live proxy shows litellm_logging.py:3514 does still fire once per process for that configuration, so a vulnerable logger is constructed on the default setup, but it is not the object that performs the logging and no trace loss follows. It fires once, not per request

The path that both constructs the vulnerable logger and logs through it is LangFuseHandler._create_langfuse_logger_from_credentials, reached with per-key or per-team Langfuse credentials and cached long-lived in the DynamicLoggingCache. That is the configuration the proof above reproduces. _health_endpoints.py:291 constructs one per health check, measured above as no worse than base

Tests

tests/test_litellm/llms/custom_httpx/test_http_handler.py covers both directions: a handler must not close a client it does not own, a consumer holding a client obtained from the cache must survive eviction plus a forced collection, and an exclusively owned client's connection pool must still be closed when its handler is collected, so the guard cannot be quietly widened into a no-op. The eviction tests use a weakref to assert the handler really was collected rather than trusting that it was, and the pool tests assert the pooled connection existed before asserting it went away

tests/test_litellm/integrations/test_langfuse.py covers the client handed to the Langfuse SDK surviving cache eviction, and the loggers continuing to share one cached client

This closes a gap in tests/test_litellm/llms/test_lifecycle_fix.py, which drops a handler and then asserts only that an unrelated client is still open, never asserting anything about the client the dropped handler was holding

Every executable line this PR adds is exercised: 15 of 15, checked with coverage locally per changed line. Both directions are pinned for both handlers, verified by mutation rather than assumed: widening the guard so it fires on a shared client is caught by the survival tests, narrowing it so it never fires is caught by the pool-closure tests, and rebinding the client to a local at either call site is caught by the matching pool-closure test. An earlier revision of the async closure test passed under that last mutant because it allowed the client to have been collected instead of closed, so it was rewritten to assert on the connection pool the same way the sync one does

QA runbook

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR changes HTTP client ownership handling so cache eviction and handler finalization do not close clients still used by Langfuse or other consumers.

  • Retains the cached HTTP handler for each Langfuse logger’s lifetime.
  • Makes explicit close and context-manager cleanup respect client ownership.
  • Adds sync and async lifecycle tests covering shared-client survival and owned-client cleanup.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Important Files Changed

Filename Overview
litellm/integrations/langfuse/langfuse.py Retains the cached HTTP handler alongside the client passed to the Langfuse SDK, preventing premature finalization.
litellm/llms/custom_httpx/http_handler.py Introduces ownership- and refcount-aware client cleanup for synchronous and asynchronous HTTP handlers.
tests/test_litellm/integrations/test_langfuse.py Adds coverage for Langfuse client survival across cache eviction and reuse of the shared cached client.
tests/test_litellm/llms/custom_httpx/test_http_handler.py Adds lifecycle coverage for injected clients, cache eviction, finalization, and connection-pool cleanup.

Reviews (4): Last reviewed commit: "fix(langfuse): stop a collected httpx ha..." | Re-trigger Greptile

@codecov

codecov Bot commented Aug 5, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@yassin-berriai
yassin-berriai force-pushed the litellm_langfuse_shared_client_close branch from 15754d7 to f266024 Compare August 5, 2026 19:49
@yassin-berriai

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review f266024. The design changed since the initial open: the finalizer is kept but guarded, rather than removed.

@yassin-berriai
yassin-berriai force-pushed the litellm_langfuse_shared_client_close branch from f266024 to bebf74c Compare August 5, 2026 20:01
@yassin-berriai

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review bebf74c. Adds regression tests pinning the finalizer guard in both directions for the sync and async handlers.

…lient

A cached HTTPHandler hands its raw httpx.Client out to consumers that keep it
for the process lifetime. When the shared client cache expires the entry on its
TTL or evicts it under the 200-entry cap, nothing references the handler, so it
is collected and its finalizer closed the client those consumers still hold.
Langfuse ingestion then failed silently on the SDK's background flush thread
until the process restarted.

A finalizer running proves only that nothing references the handler; it proves
nothing about the client. Both handlers now close the client during finalization
only when they built it and are still its sole referrer, so an unshared client is
still released promptly and a handed-out one is left alone. That keeps the
pooled-socket reclamation the finalizer was providing, which measures identical
to base over 2000 handler create-and-drop cycles.

Explicit close() stays, now gated on _owns_client so the wrapper never closes a
caller-injected client, and __aexit__ routes through it.

LangFuseLogger also keeps a reference to the handler whose client it hands the
SDK. Previously that handler was a local that went out of scope immediately,
leaving the client reachable only from the SDK. It still shares the cached
client, so no extra clients are created per logger.
@yassin-berriai
yassin-berriai force-pushed the litellm_langfuse_shared_client_close branch from bebf74c to 9478d25 Compare August 5, 2026 20:13
@yassin-berriai

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review 9478d25. Adds an __aexit__ ownership test, bringing every changed line under test.

@codspeed-hq

codspeed-hq Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_langfuse_shared_client_close (9478d25) with litellm_internal_staging (2792887)1

Open in CodSpeed

Footnotes

  1. No successful run was found on litellm_internal_staging (32deaff) during the generation of this report, so 2792887 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report.

@yassin-berriai
yassin-berriai enabled auto-merge (squash) August 5, 2026 21:54
@yassin-berriai
yassin-berriai merged commit 9d58923 into litellm_internal_staging Aug 5, 2026
78 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_langfuse_shared_client_close branch August 5, 2026 21:54
doonga pushed a commit to greyrock-labs/home-ops that referenced this pull request Aug 17, 2026
…7.0) (#336)

This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.96.2` → `v1.97.0` |

---

### Release Notes

<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>

### [`v1.97.0`](https://github.com/BerriAI/litellm/releases/tag/v1.97.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.97.0...v1.97.0)

##### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.97.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.97.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.97.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

##### What's Changed

- feat(proxy): resolve Cursor thinking/fast model-name suffixes on /cursor/chat/completions by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35554](https://github.com/BerriAI/litellm/pull/35554)
- fix(team-callbacks): actually stop logging when disable\_logging is called by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35520](https://github.com/BerriAI/litellm/pull/35520)
- refactor(lint): drop redundant !s f-string conversion flags and fix displaced import-group comments by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35546](https://github.com/BerriAI/litellm/pull/35546)
- fix(proxy): backfill null user\_email on existing users during JWT auth by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34588](https://github.com/BerriAI/litellm/pull/34588)
- feat(playground): add non-streaming response toggle by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35560](https://github.com/BerriAI/litellm/pull/35560)
- feat(teams): apply default organization to new teams from default team settings by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35540](https://github.com/BerriAI/litellm/pull/35540)
- fix(ui): block Playground page for viewer roles on direct URL access by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35676](https://github.com/BerriAI/litellm/pull/35676)
- fix(caching): close evicted LLM clients so their connections are reclaimed by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35492](https://github.com/BerriAI/litellm/pull/35492)
- chore(deps): update brace-expansion, postcss, and gitpython to current patch releases by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35692](https://github.com/BerriAI/litellm/pull/35692)
- refactor(ui): rename the create MCP server component to PascalCase by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35686](https://github.com/BerriAI/litellm/pull/35686)
- fix(openai): drop undefined Union from owns\_wrapped\_http\_client annotation by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35706](https://github.com/BerriAI/litellm/pull/35706)
- fix(openai): drop the undefined Union from owns\_wrapped\_http\_client by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35704](https://github.com/BerriAI/litellm/pull/35704)
- chore(ui): note Google's Agent Platform rename in vector store setup by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;28076](https://github.com/BerriAI/litellm/pull/28076)
- fix(proxy): apply key/team router\_settings.model\_group\_alias by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35486](https://github.com/BerriAI/litellm/pull/35486)
- feat(complexity\_router): default session affinity off and expose it in the UI by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35714](https://github.com/BerriAI/litellm/pull/35714)
- fix(datadog): read team callback dd\_\* params from kwargs instead of blocked dynamic params ([#&#8203;35115](https://github.com/BerriAI/litellm/issues/35115) port) by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35687](https://github.com/BerriAI/litellm/pull/35687)
- refactor(ui): extract the MCP create form's logic and field groups by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35694](https://github.com/BerriAI/litellm/pull/35694)
- test(ui): tier the MCP create tests into unit and integration by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35697](https://github.com/BerriAI/litellm/pull/35697)
- fix(proxy): redact credential headers from request logging copies by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35678](https://github.com/BerriAI/litellm/pull/35678)
- feat(guardrails/rubrik): prompt moderation, response-text blocking, streaming buffer, failure logging by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35722](https://github.com/BerriAI/litellm/pull/35722)
- fix(ui): render Responses API request and response in the logs drawer by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35718](https://github.com/BerriAI/litellm/pull/35718)
- fix(ui): hide guardrail review buttons from non-admin users by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;27535](https://github.com/BerriAI/litellm/pull/27535)
- feat(team): custom metadata validation hook for team create and update by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33353](https://github.com/BerriAI/litellm/pull/33353)
- ci(circleci): install a pinned Rust toolchain on the Linux jobs by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35519](https://github.com/BerriAI/litellm/pull/35519)
- fix(bedrock): stop forwarding no-op toolSpec.strict to Converse by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35688](https://github.com/BerriAI/litellm/pull/35688)
- fix(ui): reject an auto-router keyword rule left empty instead of dropping it by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35705](https://github.com/BerriAI/litellm/pull/35705)
- fix(guardrails/rubrik): attribute blocked requests to the caller that made them by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35734](https://github.com/BerriAI/litellm/pull/35734)
- fix(responses): forward client headers to the provider on /v1/responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34531](https://github.com/BerriAI/litellm/pull/34531)
- feat(spend): add net auto-router savings to the cost-optimization dashboard by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35521](https://github.com/BerriAI/litellm/pull/35521)
- chore(typing): clear basedpyright Any errors in budget reset, access groups, and cache settings by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35719](https://github.com/BerriAI/litellm/pull/35719)
- fix(spend): read what a request cost from the record instead of pricing it again by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35736](https://github.com/BerriAI/litellm/pull/35736)
- perf: install hiredis so redis-py parses replies with its C parser by [@&#8203;Classic298](https://github.com/Classic298) in [#&#8203;35709](https://github.com/BerriAI/litellm/pull/35709)
- feat(ui): show auto-router savings on the cost-optimization dashboard by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35522](https://github.com/BerriAI/litellm/pull/35522)
- perf: build log messages lazily so filtered-out log records cost nothing by [@&#8203;Classic298](https://github.com/Classic298) in [#&#8203;35703](https://github.com/BerriAI/litellm/pull/35703)
- fix(proxy): retry model cost map fetch with Retry-After-aware backoff and keep current map on reload failure by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35739](https://github.com/BerriAI/litellm/pull/35739)
- feat(otel): stamp service tier attributes on inference spans by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35679](https://github.com/BerriAI/litellm/pull/35679)
- fix(proxy): log the model cost map reload failure lazily by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35750](https://github.com/BerriAI/litellm/pull/35750)
- fix(groq): translate web\_search\_options to the browser\_search tool by [@&#8203;hMED22](https://github.com/hMED22) in [#&#8203;34971](https://github.com/BerriAI/litellm/pull/34971)
- feat(ui): add admin-configurable user banner by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35729](https://github.com/BerriAI/litellm/pull/35729)
- fix(e2e): make spend-counter redis connection env-driven for non-cluster deployments by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35732](https://github.com/BerriAI/litellm/pull/35732)
- fix(proxy): make /cursor/chat/completions work with Cursor agent mode by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;34029](https://github.com/BerriAI/litellm/pull/34029)
- fix(proxy): propagate user\_email and bind api\_key on JWT auth attribution paths by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34331](https://github.com/BerriAI/litellm/pull/34331)
- chore(build): move the Admin UI toolchain to Node 24 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35801](https://github.com/BerriAI/litellm/pull/35801)
- test(e2e): vendor API strategy coverage across endpoints by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;34649](https://github.com/BerriAI/litellm/pull/34649)
- chore(deps): upgrade cryptography to 50.0.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35803](https://github.com/BerriAI/litellm/pull/35803)
- test(e2e): cover legacy text /completions endpoint by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;34431](https://github.com/BerriAI/litellm/pull/34431)
- feat(gemini): add gemini-robotics-er-2-preview and gemini-robotics-er-1.6-preview by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35555](https://github.com/BerriAI/litellm/pull/35555)
- test(e2e): move load/perf testing out of the main suite and drop the vllm passthrough test by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35820](https://github.com/BerriAI/litellm/pull/35820)
- feat(lint): enforce Final on locals and freeze function parameters (LIT010, LIT011) by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35807](https://github.com/BerriAI/litellm/pull/35807)
- chore: bump litellm-proxy-extras 0.4.81 -> 0.4.82, litellm 1.96.0 -> 1.97.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35810](https://github.com/BerriAI/litellm/pull/35810)
- fix(bedrock): drop conflicting tool\_choice.type when toolConfig.toolChoice is set by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35738](https://github.com/BerriAI/litellm/pull/35738)
- docs(CLAUDE.md): prefer commas over semicolons when replacing em dashes by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35825](https://github.com/BerriAI/litellm/pull/35825)
- chore(lint): zero out basedpyright headroom for purely local rules by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35828](https://github.com/BerriAI/litellm/pull/35828)
- test(e2e): retry provider-transient statuses at the transport with bounded backoff by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35824](https://github.com/BerriAI/litellm/pull/35824)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35836](https://github.com/BerriAI/litellm/pull/35836)
- refactor(ui): route MCP session tokens through the shared storage helper by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35835](https://github.com/BerriAI/litellm/pull/35835)
- docs(helm): replace the classic chart's 128Mi resource example with the documented 4Gi sizing by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35830](https://github.com/BerriAI/litellm/pull/35830)
- fix(proxy): persist periodic reload schedule state so status survives restarts and fires without store\_model\_in\_db by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35165](https://github.com/BerriAI/litellm/pull/35165)
- fix(router): eagerly fetch Vertex AI deferred stream to surface HTTP errors in \_acompletion fallback path by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;34627](https://github.com/BerriAI/litellm/pull/34627)
- fix(azure\_storage): honor AZURE\_STORAGE\_ENDPOINT\_SUFFIX for sovereign clouds by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35806](https://github.com/BerriAI/litellm/pull/35806)
- fix(proxy): apply key\_alias/key\_hash filters to all /key/list visibility branches by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35840](https://github.com/BerriAI/litellm/pull/35840)
- fix(proxy): enforce per-model budgets against resolved cursor model variants by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35834](https://github.com/BerriAI/litellm/pull/35834)
- feat(ui): reorder Add Auto Router into name + template, with a collapsible detailed config by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35746](https://github.com/BerriAI/litellm/pull/35746)
- test: repair three failing suites on litellm\_internal\_staging by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35845](https://github.com/BerriAI/litellm/pull/35845)
- fix(guardrails): scan model output on the /openai/v1/responses alias by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35818](https://github.com/BerriAI/litellm/pull/35818)
- ci: pin Node on the Playwright UI lanes so npm ci meets the engines floor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35848](https://github.com/BerriAI/litellm/pull/35848)
- fix(pricing): apply OpenAI's gpt-5.6 terra/luna cut to Azure cost map by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;35481](https://github.com/BerriAI/litellm/pull/35481)
- feat(spend): add caller-scoped key/user/team/organization spend report endpoints by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35725](https://github.com/BerriAI/litellm/pull/35725)
- revert: "fix(caching): close evicted LLM clients so their connections are reclaimed ([#&#8203;35492](https://github.com/BerriAI/litellm/issues/35492))" by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35856](https://github.com/BerriAI/litellm/pull/35856)
- refactor(repositories): add prisma protocol seams and a spend-reset unit of work by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35748](https://github.com/BerriAI/litellm/pull/35748)
- perf(streaming): assemble streamed tool-call arguments in linear time by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35826](https://github.com/BerriAI/litellm/pull/35826)
- fix(s3\_v2): sign S3 object URLs with S3SigV4Auth so encoded paths verify by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35726](https://github.com/BerriAI/litellm/pull/35726)
- test(e2e): self-seed the ui suite's password-login users in global setup by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35863](https://github.com/BerriAI/litellm/pull/35863)
- fix(claude-code): create-only skill registration with a PUT update route (LIT-4110) by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;31752](https://github.com/BerriAI/litellm/pull/31752)
- fix(proxy):  fix zguard  httpcode when block input by [@&#8203;jwang-gif](https://github.com/jwang-gif) in [#&#8203;31948](https://github.com/BerriAI/litellm/pull/31948)
- fix(lint): pick the merge-aware base so in-progress merges are not blamed for base drift by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35868](https://github.com/BerriAI/litellm/pull/35868)
- chore: bump litellm-proxy-extras 0.4.82 -> 0.4.83 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35877](https://github.com/BerriAI/litellm/pull/35877)
- feat(ui): add Test Routing to the auto router create form by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35859](https://github.com/BerriAI/litellm/pull/35859)
- fix(ui): derive auto-router preset tests from the bundled preset JSON by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35882](https://github.com/BerriAI/litellm/pull/35882)
- revert: "test(e2e): vendor API strategy coverage across endpoints" ([#&#8203;34649](https://github.com/BerriAI/litellm/issues/34649)) by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35881](https://github.com/BerriAI/litellm/pull/35881)
- chore(deps): bump grpc and golang.org/x modules in the terraform provider by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35844](https://github.com/BerriAI/litellm/pull/35844)
- test(e2e): skip view-backed global spend probes pending LIT-5211 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35875](https://github.com/BerriAI/litellm/pull/35875)
- fix(lint): move the basedpyright heap flag into the type check gate by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35869](https://github.com/BerriAI/litellm/pull/35869)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35876](https://github.com/BerriAI/litellm/pull/35876)
- feat(ui): add role capability gating, migrate Tool Policies route by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35812](https://github.com/BerriAI/litellm/pull/35812)
- refactor(ui): inject the fetch client's base url instead of reading it at import by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35802](https://github.com/BerriAI/litellm/pull/35802)
- chore: remove unused .flake8 config and flake8 dev dependency by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35888](https://github.com/BerriAI/litellm/pull/35888)
- chore: stop advising pre-commit and bootstrap by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35884](https://github.com/BerriAI/litellm/pull/35884)
- fix(auth): name enable\_jwt\_auth when a JWT-shaped key is rejected by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35831](https://github.com/BerriAI/litellm/pull/35831)
- feat(auto-router): make reminder marker pair configurable by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;35874](https://github.com/BerriAI/litellm/pull/35874)
- fix(UI): update anthropic model presets by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35896](https://github.com/BerriAI/litellm/pull/35896)
- fix(bootstrap): switch to the dashboard node floor via nvm or fnm by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35895](https://github.com/BerriAI/litellm/pull/35895)
- perf(pre-commit): run python, dashboard, and gen-api checks concurrently by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35903](https://github.com/BerriAI/litellm/pull/35903)
- feat(spend): derive a default auto-router savings baseline from the hardest tier by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35907](https://github.com/BerriAI/litellm/pull/35907)
- fix(http\_handler): self-heal handler clients closed after cache eviction by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35862](https://github.com/BerriAI/litellm/pull/35862)
- fix(cost\_tracking): keep OpenAI prompt cache token details through usage reassembly by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34812](https://github.com/BerriAI/litellm/pull/34812)
- fix(cost): bill gpt-5.6 prompt cache reads at the cache read rate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34957](https://github.com/BerriAI/litellm/pull/34957)
- fix(batches): account for Responses API usage by [@&#8203;rimysore](https://github.com/rimysore) in [#&#8203;35367](https://github.com/BerriAI/litellm/pull/35367)
- ci: retry Codecov uploads and stop failing jobs on OIDC token flakes by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35251](https://github.com/BerriAI/litellm/pull/35251)
- feat(complexity\_router): let operators rename the four complexity tiers by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;35893](https://github.com/BerriAI/litellm/pull/35893)
- chore(lint): zero stale ruff and LIT headroom and strip inert type: ignore comments by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35928](https://github.com/BerriAI/litellm/pull/35928)
- chore(lint): zero out seven more purely local basedpyright rules by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35927](https://github.com/BerriAI/litellm/pull/35927)
- chore(ui): zero stale headroom on local dashboard eslint budgets by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35929](https://github.com/BerriAI/litellm/pull/35929)
- fix(managed-files): skip rows without file objects by [@&#8203;rimysore](https://github.com/rimysore) in [#&#8203;35365](https://github.com/BerriAI/litellm/pull/35365)
- fix(router): redact fallback tracebacks at the call site and cover the sync deferred stream by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35843](https://github.com/BerriAI/litellm/pull/35843)
- fix(migrations): recover from an interrupted Prisma toolchain install by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35832](https://github.com/BerriAI/litellm/pull/35832)
- fix(lint): bring basedpyright rule counts back under their budget limits by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35962](https://github.com/BerriAI/litellm/pull/35962)
- chore(ui): don't zero out stale headroom except no-console by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35964](https://github.com/BerriAI/litellm/pull/35964)
- fix(proxy): give proxy\_admin\_viewer read parity with proxy\_admin by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35851](https://github.com/BerriAI/litellm/pull/35851)
- refactor(ui): address UI lint budget issues by refactoring UI by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35960](https://github.com/BerriAI/litellm/pull/35960)
- fix(ci): make the env-key doc gate see get\_secret\_bool reads by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35833](https://github.com/BerriAI/litellm/pull/35833)
- fix(caching): re-land evicted LLM client closing ([#&#8203;35492](https://github.com/BerriAI/litellm/issues/35492)) atop self-healing handlers by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35870](https://github.com/BerriAI/litellm/pull/35870)
- fix(proxy): keep the connected DB client when a startup health check fails by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35837](https://github.com/BerriAI/litellm/pull/35837)
- chore(lint): remove litellm/types from the ruff lint exclusion by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35926](https://github.com/BerriAI/litellm/pull/35926)
- feat(sgr): make the gateway middleware the source of truth for successful requests by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35717](https://github.com/BerriAI/litellm/pull/35717)
- feat(auto-router): let operators replace the LLM classifier's system prompt by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;35855](https://github.com/BerriAI/litellm/pull/35855)
- fix(docker): bake the pip image's prisma engines at a world-readable path by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35976](https://github.com/BerriAI/litellm/pull/35976)
- fix(auth): return 403 from the OAuth2 enterprise gate by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35838](https://github.com/BerriAI/litellm/pull/35838)
- fix(router): keep custom model\_info across a price data reload by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35491](https://github.com/BerriAI/litellm/pull/35491)
- fix(proxy): resolve pass-through credentials live from router deployments by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35916](https://github.com/BerriAI/litellm/pull/35916)
- fix(ci): fetch only head and merge-base in lint jobs instead of every branch by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35982](https://github.com/BerriAI/litellm/pull/35982)
- fix(autorouter): match CJK keyword\_tier\_rules that regex word boundaries miss by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;35984](https://github.com/BerriAI/litellm/pull/35984)
- feat(spend): rebuild the auto-router benchmarks backend as a per-session rollup by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35910](https://github.com/BerriAI/litellm/pull/35910)
- refactor(ui): replace hand-rolled query-param routing with nuqs by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35871](https://github.com/BerriAI/litellm/pull/35871)
- fix(docker): bake the componentized prisma engines at /opt/prisma so any uid can start by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35989](https://github.com/BerriAI/litellm/pull/35989)
- fix(migrations): keep the toolchain heal from raising on an unreadable nodeenv cache by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35986](https://github.com/BerriAI/litellm/pull/35986)
- fix(bedrock): sign Bedrock managed-file S3 requests with S3SigV4Auth by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35983](https://github.com/BerriAI/litellm/pull/35983)
- chore(typing): replace Any seams with real types across responses, proxy, and provider adapters by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35809](https://github.com/BerriAI/litellm/pull/35809)
- fix(ai21): resolve the documented AI21\_API\_KEY instead of a misspelled name by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35985](https://github.com/BerriAI/litellm/pull/35985)
- fix(docker): fail the image build when the generated prisma engine paths drift off /opt/prisma by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35979](https://github.com/BerriAI/litellm/pull/35979)
- fix(jina\_ai): resolve the documented JINA\_API\_KEY as a fallback by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35992](https://github.com/BerriAI/litellm/pull/35992)
- fix(proxy): only treat a recoverable database outage as grounds to serve without one by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35864](https://github.com/BerriAI/litellm/pull/35864)
- fix(ci): make every remaining CI checkout shallow by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35997](https://github.com/BerriAI/litellm/pull/35997)
- fix(auto-router): stop the embedding model's context window from failing long requests by [@&#8203;akapur99](https://github.com/akapur99) in [#&#8203;35956](https://github.com/BerriAI/litellm/pull/35956)
- fix(ci): make the env-key doc gate see bare get\_secret and get\_secret\_str reads by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35996](https://github.com/BerriAI/litellm/pull/35996)
- fix(logging): extend secret redaction to records litellm does not emit directly by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35977](https://github.com/BerriAI/litellm/pull/35977)
- test(utils): pin the register\_model replay test to the recorded half by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35994](https://github.com/BerriAI/litellm/pull/35994)
- fix(ci): run every helm test suite, not just the first one per file by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35993](https://github.com/BerriAI/litellm/pull/35993)
- ci: fail the build when a test file or Dockerfile is invoked by no job by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35991](https://github.com/BerriAI/litellm/pull/35991)
- fix(langfuse): stop a collected httpx handler from closing a shared client by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35981](https://github.com/BerriAI/litellm/pull/35981)
- fix(bedrock): grant bedrock:CountTokens in OIDC session policy by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33145](https://github.com/BerriAI/litellm/pull/33145)
- feat(pre-commit): save full lint output to a per-worktree log file by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36004](https://github.com/BerriAI/litellm/pull/36004)
- feat(ui): match auto-router preset models against deployments' underlying model IDs by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35972](https://github.com/BerriAI/litellm/pull/35972)
- fix(core\_helpers): map generic 'error' finish\_reason to 'stop' by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33972](https://github.com/BerriAI/litellm/pull/33972)
- fix(proxy)!: apply request-parameter checks consistently across body, path and form inputs by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36011](https://github.com/BerriAI/litellm/pull/36011)
- fix: rebuild models\_by\_provider in add\_known\_models so cost map reloads reach wildcard expansion by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36010](https://github.com/BerriAI/litellm/pull/36010)
- feat(complexity\_router): report LLM classifier cost per request via routing\_decision and x-litellm-classifier-cost header by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36015](https://github.com/BerriAI/litellm/pull/36015)
- fix(model-prices): correct replicate model key typo by [@&#8203;AkashNaickar](https://github.com/AkashNaickar) in [#&#8203;34800](https://github.com/BerriAI/litellm/pull/34800)
- fix(proxy): register managed batch output files on terminal retrieve by [@&#8203;Souravrajvi0](https://github.com/Souravrajvi0) in [#&#8203;34092](https://github.com/BerriAI/litellm/pull/34092)
- perf(pre-commit): fetch basedpyright base counts from CI artifacts by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35970](https://github.com/BerriAI/litellm/pull/35970)
- fix(ui): sync projects list page index to ?page= so back and reload keep the page by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36003](https://github.com/BerriAI/litellm/pull/36003)
- fix(ui): link project page keys to their virtual key detail by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36002](https://github.com/BerriAI/litellm/pull/36002)
- refactor(ui): drop unreferenced locals from dashboard route components by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35819](https://github.com/BerriAI/litellm/pull/35819)
- fix(ui): opening a project now pushes ?project= so back and deep links work by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36001](https://github.com/BerriAI/litellm/pull/36001)
- refactor(ui): drop unreferenced locals from shared dashboard components by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35821](https://github.com/BerriAI/litellm/pull/35821)
- refactor(ui): drop unreferenced locals from tests and narrow destructures by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36025](https://github.com/BerriAI/litellm/pull/36025)
- fix(guardrails): allow litellm\_content\_filter to run on post\_mcp\_call by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35980](https://github.com/BerriAI/litellm/pull/35980)
- fix(guardrails): scan /v1/messages tool traffic by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35999](https://github.com/BerriAI/litellm/pull/35999)
- refactor(ui): drop dead locals and unused React state across the dashboard by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36026](https://github.com/BerriAI/litellm/pull/36026)
- feat(ui): add the auto-router usage tab to cost optimization by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35995](https://github.com/BerriAI/litellm/pull/35995)
- fix(managed\_files): derive unified output file ids deterministically so concurrent registrations converge by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36019](https://github.com/BerriAI/litellm/pull/36019)
- fix(proxy): send keepalive pings on anthropic messages SSE streams during upstream silence by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36024](https://github.com/BerriAI/litellm/pull/36024)
- fix(managed\_files): return unified ids from unscoped file listing by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36031](https://github.com/BerriAI/litellm/pull/36031)
- fix(arize\_phoenix): lowercase OTLP/gRPC auth metadata key by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34883](https://github.com/BerriAI/litellm/pull/34883)
- fix(auto-router): accept every reminder marker pair a harness emits by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36029](https://github.com/BerriAI/litellm/pull/36029)
- fix(pricing): sync flex/priority tier keys to dated OpenAI snapshot variants by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35923](https://github.com/BerriAI/litellm/pull/35923)
- fix(cost): bill reasoning tokens at the service tier output rate by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35925](https://github.com/BerriAI/litellm/pull/35925)
- fix(proxy): include today's UTC bucket when a daily activity range ends at the caller's current day by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36051](https://github.com/BerriAI/litellm/pull/36051)
- fix: expired-miss share over all measured turns + cost-optimization tab labels by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36037](https://github.com/BerriAI/litellm/pull/36037)
- fix(router): include Bedrock batch/S3 fields and model in deployment credentials by [@&#8203;mpcusack-altos](https://github.com/mpcusack-altos) in [#&#8203;24548](https://github.com/BerriAI/litellm/pull/24548)
- fix(batch): track cost for managed batches with no attributable key/u… by [@&#8203;elinacse](https://github.com/elinacse) in [#&#8203;35468](https://github.com/BerriAI/litellm/pull/35468)
- feat(guardrails): add scan\_only\_tool\_results to scope unified guardrails to tool results by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36014](https://github.com/BerriAI/litellm/pull/36014)
- fix(cost): stop token-pricing the placeholder input on file content calls by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35140](https://github.com/BerriAI/litellm/pull/35140)
- fix(proxy): fetch background responses through the router in CheckResponsesCost by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35137](https://github.com/BerriAI/litellm/pull/35137)
- fix(proxy): yaml store\_prompts\_in\_spend\_logs should take precedence over DB cached value by [@&#8203;Praveena-617](https://github.com/Praveena-617) in [#&#8203;35769](https://github.com/BerriAI/litellm/pull/35769)
- fix(lint): measure the basedpyright budget gate in a gate-owned venv by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36050](https://github.com/BerriAI/litellm/pull/36050)
- docs: cap all GitHub comments at 15-25 words, curb semicolon splices by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36059](https://github.com/BerriAI/litellm/pull/36059)
- chore(lint): name MappingProxyType in the mutable-collection fix messages by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36072](https://github.com/BerriAI/litellm/pull/36072)
- test: roll back runtime model registrations between tests by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36039](https://github.com/BerriAI/litellm/pull/36039)
- refactor(types): cut 653 implicit and explicit Any diagnostics across 11 modules by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36054](https://github.com/BerriAI/litellm/pull/36054)
- fix(proxy): stop resolving the UI session sentinel team on /search\_tools/list by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36061](https://github.com/BerriAI/litellm/pull/36061)
- fix(batches): persist managed file ids for cancelled/failed/expired batches by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36048](https://github.com/BerriAI/litellm/pull/36048)
- fix(batches): register managed output files on batch cancel by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36034](https://github.com/BerriAI/litellm/pull/36034)
- fix(proxy): allow non-admins to reach /user/daily/activity/aggregated by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36062](https://github.com/BerriAI/litellm/pull/36062)
- fix(anthropic): coerce explicit additionalProperties to false in output\_format schema by [@&#8203;dkindlund](https://github.com/dkindlund) in [#&#8203;35811](https://github.com/BerriAI/litellm/pull/35811)
- fix(batches): prevent managed file fallbacks by [@&#8203;rimysore](https://github.com/rimysore) in [#&#8203;35371](https://github.com/BerriAI/litellm/pull/35371)
- chore: ignore the mechanical lint and typing sweeps in git blame by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36076](https://github.com/BerriAI/litellm/pull/36076)
- fix(proxy): warn at startup when max\_budget is set but no database is connected by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36041](https://github.com/BerriAI/litellm/pull/36041)
- fix(proxy): promote caller metadata trace fields into litellm\_metadata by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35866](https://github.com/BerriAI/litellm/pull/35866)
- feat(terraform): sync provider 0.3.0 from the mirror and cut 0.4.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36098](https://github.com/BerriAI/litellm/pull/36098)
- fix(guardrails): honor configured timeout in Zscaler AI Guard by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36110](https://github.com/BerriAI/litellm/pull/36110)
- fix(logging): fall back to litellm\_metadata when metadata is empty by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36105](https://github.com/BerriAI/litellm/pull/36105)
- fix(proxy): re-assert the authenticated identity on passthrough requests by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36121](https://github.com/BerriAI/litellm/pull/36121)
- chore: bump litellm-enterprise 0.1.53 -> 0.1.54, litellm-proxy-extras 0.4.83 -> 0.4.84 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36139](https://github.com/BerriAI/litellm/pull/36139)
- fix(ui): match auto-router preset models against wildcard-expanded model groups by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36111](https://github.com/BerriAI/litellm/pull/36111)
- test(router): assert the auto-router max\_input\_chars kwarg by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36109](https://github.com/BerriAI/litellm/pull/36109)
- fix(ui): allow clearing a key's budget reset from the Edit Key form by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36140](https://github.com/BerriAI/litellm/pull/36140)
- fix(managed\_files): skip unparseable rows when listing managed files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36021](https://github.com/BerriAI/litellm/pull/36021)
- fix(a2a): stop writing per-caller headers onto the shared cached httpx client by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35978](https://github.com/BerriAI/litellm/pull/35978)
- build(deps): bump h2 to 4.4.1 and js-yaml to 4.3.1 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36147](https://github.com/BerriAI/litellm/pull/36147)
- chore: promote staging to main by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36057](https://github.com/BerriAI/litellm/pull/36057)
- fix(azure\_sentinel): respect AZURE\_AUTHORITY\_HOST and derive the Azure Monitor audience per cloud by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36137](https://github.com/BerriAI/litellm/pull/36137)
- fix(bedrock): pass SSE-KMS key through to the batch input-file S3 upload by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35148](https://github.com/BerriAI/litellm/pull/35148)
- fix(anthropic adapter): stop indexing choices\[0] on choiceless streaming chunks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35314](https://github.com/BerriAI/litellm/pull/35314)
- fix(bedrock): normalize /v1/completions and /v1/responses batch records by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35675](https://github.com/BerriAI/litellm/pull/35675)
- fix(proxy): return the real status code when a credential update is rejected by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36166](https://github.com/BerriAI/litellm/pull/36166)
- fix(proxy): improve Headroom /v1/compress HTTP 404 diagnostics by [@&#8203;aayush598](https://github.com/aayush598) in [#&#8203;35952](https://github.com/BerriAI/litellm/pull/35952)
- fix(proxy): invalidate cached project object on project update and delete by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36028](https://github.com/BerriAI/litellm/pull/36028)
- feat(proxy): add apply\_user\_budget\_to\_team\_keys opt-in by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36102](https://github.com/BerriAI/litellm/pull/36102)
- fix(proxy): stop alerting on health probes that lose the planned engine-restart race by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36141](https://github.com/BerriAI/litellm/pull/36141)
- test(docker): gate the componentized gateway and backend images on an arbitrary-uid offline boot by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36136](https://github.com/BerriAI/litellm/pull/36136)
- fix(http): stop pooled clients persisting cookies on the aiohttp jar too by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36149](https://github.com/BerriAI/litellm/pull/36149)
- fix(router): bound fallback-walk work and error-log volume by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36148](https://github.com/BerriAI/litellm/pull/36148)
- ci: wire credential\_endpoints tests into the proxy endpoints job by [@&#8203;cursor](https://github.com/cursor)\[bot] in [#&#8203;36187](https://github.com/BerriAI/litellm/pull/36187)
- docs(keys): document /key/info fields and clarify budget\_reset\_at is the next reset by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36127](https://github.com/BerriAI/litellm/pull/36127)
- fix(azure\_sentinel): add AZURE\_SENTINEL\_AUTHORITY\_HOST as a Sentinel scoped override by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36165](https://github.com/BerriAI/litellm/pull/36165)
- docs(pr-template): add a User Flow section with authoring instructions by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36162](https://github.com/BerriAI/litellm/pull/36162)
- fix(proxy): derive config agent ids from agent\_name so grants survive secret rotation by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36020](https://github.com/BerriAI/litellm/pull/36020)
- chore(ui): regenerate schema.d.ts for the /key/info docstring update by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36210](https://github.com/BerriAI/litellm/pull/36210)
- build(deps): bump gitpython to 3.1.58 to clear osv-scan on staging by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36212](https://github.com/BerriAI/litellm/pull/36212)
- fix(proxy): deny agent access when key and team grants resolve to nothing by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36221](https://github.com/BerriAI/litellm/pull/36221)
- build(deps): defer the second pypdf advisory until the 6.15.0 bump by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36218](https://github.com/BerriAI/litellm/pull/36218)
- fix(a2a): align agent list annotation and test with the tuple return type by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36217](https://github.com/BerriAI/litellm/pull/36217)
- ci: always run the UI API types sync check so it can be required by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36213](https://github.com/BerriAI/litellm/pull/36213)
- build(deps): bump nanoid to 3.3.17 in the dashboard lockfile by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36227](https://github.com/BerriAI/litellm/pull/36227)
- feat(ui): show user email or alias in usage data export by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36232](https://github.com/BerriAI/litellm/pull/36232)
- feat(auto-router): track turns per complexity tier (LIT-5302) by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36209](https://github.com/BerriAI/litellm/pull/36209)
- fix(websearch): restore snippet text in native web\_search\_tool\_result blocks (LIT-5315) by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36228](https://github.com/BerriAI/litellm/pull/36228)
- fix(proxy): resolve entity access groups in the model listing endpoints by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36230](https://github.com/BerriAI/litellm/pull/36230)
- fix(ui): let access groups be a team's only model source, with hover provenance by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36234](https://github.com/BerriAI/litellm/pull/36234)
- fix(managed\_files): return unified output file ids from GET /batches by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36049](https://github.com/BerriAI/litellm/pull/36049)
- test(proxy): compare empty agent list to the tuple get\_agent\_list returns by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36225](https://github.com/BerriAI/litellm/pull/36225)
- fix(otel): name the RPC system and upstream on MCP tool-call spans by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35857](https://github.com/BerriAI/litellm/pull/35857)
- fix(guardrails): chunk oversized Bedrock ApplyGuardrail requests instead of failing by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36119](https://github.com/BerriAI/litellm/pull/36119)
- test(e2e): settle control-plane writes across every replica, not just one by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36247](https://github.com/BerriAI/litellm/pull/36247)
- fix(responses): forward allowed\_openai\_params through the chat completions bridge by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35885](https://github.com/BerriAI/litellm/pull/35885)
- test(proxy): assert the copy \_add\_team\_member\_budget\_table returns by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36244](https://github.com/BerriAI/litellm/pull/36244)
- chore(ui): regenerate dashboard api types for tier\_turns by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36243](https://github.com/BerriAI/litellm/pull/36243)
- refactor(types): declare mirrored pricing fields on ModelInfo by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36215](https://github.com/BerriAI/litellm/pull/36215)
- fix(lint): make strict-gate noqas survive base ruff and flag stale ones by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36257](https://github.com/BerriAI/litellm/pull/36257)
- fix(vertex\_ai): surface real error/status on vertex batch create instead of IndexError 500 by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35141](https://github.com/BerriAI/litellm/pull/35141)
- ci: give the remaining pull\_request workflows a concurrency group by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36252](https://github.com/BerriAI/litellm/pull/36252)
- refactor(lint): graduate zero-violation strict rules and guard the budget ratchet by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36161](https://github.com/BerriAI/litellm/pull/36161)
- fix(proxy): enforce require\_managed\_files on every route that accepts a raw provider id by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35551](https://github.com/BerriAI/litellm/pull/35551)
- chore(typing): clear 1.4k basedpyright Any errors across 21 hotspot files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36282](https://github.com/BerriAI/litellm/pull/36282)
- test: roll back live router replay membership between tests by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36278](https://github.com/BerriAI/litellm/pull/36278)
- chore(ci): sync main into internal staging by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36288](https://github.com/BerriAI/litellm/pull/36288)
- build(lint): rename make pre-commit to make check with a working-tree fallback by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36277](https://github.com/BerriAI/litellm/pull/36277)
- fix(ui): show team BYOK models in team fallback settings by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36241](https://github.com/BerriAI/litellm/pull/36241)
- fix(otel): mark v2 server spans as failed for pre-call errors by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34546](https://github.com/BerriAI/litellm/pull/34546)
- fix(websearch\_interception): bill intercepted searches to the calling key by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35708](https://github.com/BerriAI/litellm/pull/35708)
- chore: remove pre-commit rule by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36295](https://github.com/BerriAI/litellm/pull/36295)
- docs: clarify guideline priority ordering in CLAUDE.md by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36296](https://github.com/BerriAI/litellm/pull/36296)
- feat(router): independent, default-on deployment affinity for the auto-router by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36146](https://github.com/BerriAI/litellm/pull/36146)
- test: repair stale CircleCI contracts by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36293](https://github.com/BerriAI/litellm/pull/36293)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36286](https://github.com/BerriAI/litellm/pull/36286)
- chore: rebuild Admin UI bundle for the 2026-08-08 release by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36297](https://github.com/BerriAI/litellm/pull/36297)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36304](https://github.com/BerriAI/litellm/pull/36304)

##### New Contributors

- [@&#8203;rimysore](https://github.com/rimysore) made their first contribution in [#&#8203;35367](https://github.com/BerriAI/litellm/pull/35367)
- [@&#8203;AkashNaickar](https://github.com/AkashNaickar) made their first contribution in [#&#8203;34800](https://github.com/BerriAI/litellm/pull/34800)
- [@&#8203;Souravrajvi0](https://github.com/Souravrajvi0) made their first contribution in [#&#8203;34092](https://github.com/BerriAI/litellm/pull/34092)
- [@&#8203;elinacse](https://github.com/elinacse) made their first contribution in [#&#8203;35468](https://github.com/BerriAI/litellm/pull/35468)
- [@&#8203;aayush598](https://github.com/aayush598) made their first contribution in [#&#8203;35952](https://github.com/BerriAI/litellm/pull/35952)
- [@&#8203;cursor](https://github.com/cursor)\[bot] made their first contribution in [#&#8203;36187](https://github.com/BerriAI/litellm/pull/36187)

**Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.96.0...v1.97.0>

### [`v1.97.0`](https://github.com/BerriAI/litellm/releases/tag/v1.97.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.96.2...v1.97.0)

##### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.97.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.97.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.97.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

##### What's Changed

- feat(proxy): resolve Cursor thinking/fast model-name suffixes on /cursor/chat/completions by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35554](https://github.com/BerriAI/litellm/pull/35554)
- fix(team-callbacks): actually stop logging when disable\_logging is called by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35520](https://github.com/BerriAI/litellm/pull/35520)
- refactor(lint): drop redundant !s f-string conversion flags and fix displaced import-group comments by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35546](https://github.com/BerriAI/litellm/pull/35546)
- fix(proxy): backfill null user\_email on existing users during JWT auth by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34588](https://github.com/BerriAI/litellm/pull/34588)
- feat(playground): add non-streaming response toggle by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35560](https://github.com/BerriAI/litellm/pull/35560)
- feat(teams): apply default organization to new teams from default team settings by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35540](https://github.com/BerriAI/litellm/pull/35540)
- fix(ui): block Playground page for viewer roles on direct URL access by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35676](https://github.com/BerriAI/litellm/pull/35676)
- fix(caching): close evicted LLM clients so their connections are reclaimed by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35492](https://github.com/BerriAI/litellm/pull/35492)
- chore(deps): update brace-expansion, postcss, and gitpython to current patch releases by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35692](https://github.com/BerriAI/litellm/pull/35692)
- refactor(ui): rename the create MCP server component to PascalCase by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35686](https://github.com/BerriAI/litellm/pull/35686)
- fix(openai): drop undefined Union from owns\_wrapped\_http\_client annotation by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35706](https://github.com/BerriAI/litellm/pull/35706)
- fix(openai): drop the undefined Union from owns\_wrapped\_http\_client by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35704](https://github.com/BerriAI/litellm/pull/35704)
- chore(ui): note Google's Agent Platform rename in vector store setup by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;28076](https://github.com/BerriAI/litellm/pull/28076)
- fix(proxy): apply key/team router\_settings.model\_group\_alias by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;35486](https://github.com/BerriAI/litellm/pull/35486)
- feat(complexity\_router): default session affinity off and expose it in the UI by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35714](https://github.com/BerriAI/litellm/pull/35714)
- fix(datadog): read team callback dd\_\* params from kwargs instead of blocked dynamic params ([#&#8203;35115](https://github.com/BerriAI/litellm/issues/35115) port) by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35687](https://github.com/BerriAI/litellm/pull/35687)
- refactor(ui): extract the MCP create form's logic and field groups by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35694](https://github.com/BerriAI/litellm/pull/35694)
- test(ui): tier the MCP create tests into unit and integration by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35697](https://github.com/BerriAI/litellm/pull/35697)
- fix(proxy): redact credential headers from request logging copies by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35678](https://github.com/BerriAI/litellm/pull/35678)
- feat(guardrails/rubrik): prompt moderation, response-text blocking, streaming buffer, failure logging by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35722](https://github.com/BerriAI/litellm/pull/35722)
- fix(ui): render Responses API request and response in the logs drawer by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35718](https://github.com/BerriAI/litellm/pull/35718)
- fix(ui): hide guardrail review buttons from non-admin users by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;27535](https://github.com/BerriAI/litellm/pull/27535)
- feat(team): custom metadata validation hook for team create and update by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;33353](https://github.com/BerriAI/litellm/pull/33353)
- ci(circleci): install a pinned Rust toolchain on the Linux jobs by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35519](https://github.com/BerriAI/litellm/pull/35519)
- fix(bedrock): stop forwarding no-op toolSpec.strict to Converse by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35688](https://github.com/BerriAI/litellm/pull/35688)
- fix(ui): reject an auto-router keyword rule left empty instead of dropping it by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35705](https://github.com/BerriAI/litellm/pull/35705)
- fix(guardrails/rubrik): attribute blocked requests to the caller that made them by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35734](https://github.com/BerriAI/litellm/pull/35734)
- fix(responses): forward client headers to the provider on /v1/responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34531](https://github.com/BerriAI/litellm/pull/34531)
- feat(spend): add net auto-router savings to the cost-optimization dashboard by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35521](https://github.com/BerriAI/litellm/pull/35521)
- chore(typing): clear basedpyright Any errors in budget reset, access groups, and cache settings by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;35719](https://github.com/BerriAI/litellm/pull/35719)
- fix(spend): read what a request cost from the record instead of pricing it again by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35736](https://github.com/BerriAI/litellm/pull/35736)
- perf: install hiredis so redis-py parses replies with its C parser by [@&#8203;Classic298](https://github.com/Classic298) in [#&#8203;35709](https://github.com/BerriAI/litellm/pull/35709)
- feat(ui): show auto-router savings on the cost-optimization dashboard by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35522](https://github.com/BerriAI/litellm/pull/35522)
- perf: build log messages lazily so filtered-out log records cost nothing by [@&#8203;Classic298](https://github.com/Classic298) in [#&#8203;35703](https://github.com/BerriAI/litellm/pull/35703)
- fix(proxy): retry model cost map fetch with Retry-After-aware backoff and keep current map on reload failure by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;35739](https://github.com/BerriAI/litellm/pull/35739)
- feat(otel): stamp service tier attributes on inference spans by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35679](https://github.com/BerriAI/litellm/pull/35679)
- fix(proxy): log the model cost map reload failure lazily by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35750](https://github.com/BerriAI/litellm/pull/35750)
- fix(groq): translate web\_search\_options to the browser\_search tool by [@&#8203;hMED22](https://github.com/hMED22) in [#&#8203;34971](https://github.com/BerriAI/litellm/pull/34971)
- feat(ui): add admin-configurable user banner by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35729](https://github.com/BerriAI/litellm/pull/35729)
- fix(e2e): make spend-counter redis connection env-driven for non-cluster deployments by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;35732](https://github.com/BerriAI/litellm/pull/35732)
- fix(proxy): make /cursor/chat/completions work with Cursor agent mode by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;34029](https://github.com/BerriAI/litellm/pull/34029)
- fix(proxy): propagate user\_email and bind api\_key on JWT auth attribution paths by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34331](https://github.com/BerriAI/litellm/pull/34331)
- chore(build): move the Admin UI toolchain to Node 24 by [@&#8203;yuneng-berri](https://g…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants