Skip to content

fix(redis): log a timeout streak once per interval instead of one line per cache call - #40817

Merged
yassin-berriai merged 8 commits into
mainfrom
litellm_redis_timeout_log_throttle
Sep 14, 2026
Merged

yassin-berriai merged 8 commits into
mainfrom
litellm_redis_timeout_log_throttle

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • A short Redis latency blip logs one ERROR per cache call until the breaker opens
  • A 15s blip under load produced 385 ERROR lines, all the same timeout
  • The cost tracking callback added its own traceback per request on top

How it solves it:

  • Redis timeout lines are throttled: first one logs, later ones go to DEBUG
  • Every 5s (REDIS_TIMEOUT_LOG_INTERVAL) one line summarizes how many were suppressed
  • Non-timeout Redis errors keep their per-call line, breaker-open refusals stay at DEBUG
  • Spend counter increments treat a Redis timeout like an open breaker: invalidate and return
  • The generic Cache.add_cache failure log only joins the throttle when its backend is Redis

User Flow

Before: an operator whose proxy sits behind a Redis that slows down for a few seconds sees their log volume explode

  1. Their apps keep sending POST http://litellm-domain/v1/chat/completions with {"model":"haiku","messages":[{"role":"user","content":"say hi 1"}],"max_tokens":5} and every one returns 200 with a normal completion
  2. Redis round trips go from sub-millisecond to 70ms for 15 seconds, so each cache call hits the 100ms socket timeout
  3. The proxy log gains about 450 lines in 17 seconds: 385 at ERROR, almost all Got exception from REDIS Timeout reading from ..., plus 27 Error in tracking cost callback - Timeout reading from ... tracebacks
  4. Their log-based alert on ERROR rate fires and the on-call spends time confirming nothing was actually failing

After: the same blip shows up as a handful of lines that say how big it was

  1. Their apps send the same POST http://litellm-domain/v1/chat/completions requests and every one still returns 200
  2. Redis slows down the same way for 15 seconds
  3. The proxy log gains 9 lines: one WARNING for the first timeout, one ERROR summary ... Timeout reading from ... (202 more Redis timeouts since the previous Redis timeout line were logged at DEBUG), one WARNING that the breaker opened, and the rest are unrelated spend queue warnings. No cost tracking tracebacks
  4. If Redis is actually down instead of slow, every failed cache call still logs its own ERROR line, so a real outage looks the same as before

Relevant issues

Docs row for REDIS_TIMEOUT_LOG_INTERVAL: BerriAI/litellm-docs#1436 (merged)

Affected release

Linear ticket

Resolves LIT-7520

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Shared setup for both runs. The proxy talks to Redis through a small TCP relay on 127.0.0.1:31521 that forwards to the real Redis on 31520 and can add a fixed per-direction delay read from a file. Postgres is a real local database. The model is Anthropic claude-haiku-4-5 behind the alias haiku, so every request costs real money

# config.yaml
model_list:
  - model_name: haiku
    litellm_params:
      model: anthropic/claude-haiku-4-5
      api_key: os.environ/ANTHROPIC_API_KEY
litellm_settings:
  cache: true
  cache_params:
    type: redis
    host: 127.0.0.1
    port: 31521
general_settings:
  master_key: sk-lit7520

# env
LITELLM_LOG=WARNING JSON_LOGS=True REDIS_SOCKET_TIMEOUT=0.1 REDIS_CONNECTION_POOL_TIMEOUT=5
REDIS_CIRCUIT_BREAKER_FAILURE_THRESHOLD=5 REDIS_CIRCUIT_BREAKER_RECOVERY_TIMEOUT=60 REDIS_CIRCUIT_BREAKER_TIMEOUT_MIN_DURATION=5.0
PYTHONPATH=<checkout under test>

Both legs print which tree served the proxy and whether the new symbol exists before starting:

python -c "import litellm, litellm.caching.redis_cache as r; print('served-from', litellm.__file__, 'MARKER_present' if hasattr(r, '_RedisTimeoutLogThrottle') else 'MARKER_absent')"

The blip storm is 1200 identical requests over 40 parallel curls, with the relay delay set to 70ms from second 5 to second 20:

seq 1 1200 | xargs -P 40 -I{} sh -c "curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:20520/v1/chat/completions \
  -H 'Authorization: Bearer sk-lit7520' -H 'Content-Type: application/json' \
  -d '{\"model\":\"haiku\",\"messages\":[{\"role\":\"user\",\"content\":\"say hi {}\"}],\"max_tokens\":5}'" > codes.txt &
sleep 5; echo 0.07 > delay.txt; sleep 15; echo 0 > delay.txt; wait
sort codes.txt | uniq -c

Log lines below are the JSON log rendered as time level source message

Before (b27d2cc)

Redis latency blip during 1200 requests

  1. bash start_proxy.sh /home/ubuntu/repos/litellm before.log, output served-from /home/ubuntu/repos/litellm/litellm/__init__.py MARKER_absent
  2. Run the storm above. Output: 1200 200
  3. wc -l before.log grew by 449 lines between the first request and the last one (17 seconds). By level and message:
     299 ERROR   redis_cache.py  LiteLLM Redis Caching: async set() - Got exception from REDIS Timeout reading from 127.0.0.1:31521, Writing value=...
      54 ERROR   redis_cache.py  LiteLLM Redis Caching: async increment_pipeline() - Got exception from REDIS Timeout reading from 127.0.0.1:31521
      27 ERROR   spend_log_error_logger.py  Error in tracking cost callback - Timeout reading from 127.0.0.1:31521   (each with a full traceback)
       5 ERROR   redis_cache.py  Error occurred in async batch get cache - Timeout reading from 127.0.0.1:31521
      28 WARNING redis_cache.py  TTL preservation failed, falling back to regular pipeline: Timeout reading from 127.0.0.1:31521
      27 WARNING redis_cache.py  Redis async_increment_cache_pipeline failed, falling back to in-memory result: Timeout reading from 127.0.0.1:31521
       1 WARNING redis_cache.py  Redis circuit breaker OPENED after 1412 consecutive failures (0 hard connectivity) — fast-failing Redis calls for 60s
       8 WARNING spend_update_queue.py / daily_spend_update_queue.py  Spend update queue is full. Aggregating all entries in queue to concatenate entries.
    
    385 ERROR lines for one 15 second blip, every one of them the same timeout

Redis hard down (non-timeout failure)

  1. Same proxy, docker stop litfix-redis-7520, then three sequential curls of the same body with say hi 1, say hi 2, say hi 3. Output: 200 200 200
  2. New log lines:
    01:28:59 WARNING config_sync_pubsub.py:256 config sync subscriber redis error: Connection closed by server.; reconnecting in 5s
    01:28:59 WARNING auth_cache_invalidation_pubsub.py:175 auth cache invalidation subscriber redis error: Connection closed by server.; reconnecting in 5s
    01:29:01 ERROR redis_cache.py:1577 Error occurred in async batch get cache - Connection closed by server.
    01:29:02 ERROR redis_cache.py:1047 LiteLLM Redis Caching: async set() - Got exception from REDIS Connection closed by server., Writing value={'timestamp': ...
    01:29:02 WARNING redis_cache.py:250 Redis circuit breaker OPENED after 5 consecutive failures (5 hard connectivity) — fast-failing Redis calls for 60s
    01:29:04 WARNING config_sync_pubsub.py:256 config sync subscriber redis error: Connection closed by server.; reconnecting in 10s
    01:29:04 WARNING auth_cache_invalidation_pubsub.py:175 auth cache invalidation subscriber redis error: Connection closed by server.; reconnecting in 10s
    

After (bea22df, head a4e34d6 only rewraps one test line on top of it and leaves every litellm/ file byte-identical)

Redis latency blip during 1200 requests

  1. bash start_proxy.sh /home/ubuntu/repos/litellm_redis_timeout_log_throttle after_v7.log, output served-from /home/ubuntu/repos/litellm_redis_timeout_log_throttle/litellm/__init__.py MARKER_present
  2. Control request, curl ... -d '{"model":"haiku","messages":[{"role":"user","content":"say hi"}],"max_tokens":5}' returns HTTP 200 with "content":"Hi! 👋"
  3. Run the identical storm. Output: 1200 200
  4. wc -l after_v7.log grew from 17 to 26 lines over the same window. All 9 new lines:
    20:00:55 WARNING redis_cache.py TTL preservation failed, falling back to regular pipeline: Timeout reading from 127.0.0.1:31521
    20:00:58 WARNING spend_update_queue.py:47 Spend update queue is full. Aggregating all entries in queue to concatenate entries.
    20:01:00 ERROR redis_cache.py LiteLLM Redis Caching: async set() - Got exception from REDIS: Timeout reading from 127.0.0.1:31521 (202 more Redis timeouts since the previous Redis timeout line were logged at DEBUG)
    20:01:03 WARNING redis_cache.py:252 Redis circuit breaker OPENED after 1515 consecutive failures (0 hard connectivity) — fast-failing Redis calls for 60s
    20:01:04 WARNING spend_update_queue.py:47 Spend update queue is full. Aggregating all entries in queue to concatenate entries.
    20:01:06 WARNING spend_update_queue.py:47 Spend update queue is full. Aggregating all entries in queue to concatenate entries.
    20:01:14 WARNING spend_update_queue.py:47 Spend update queue is full. Aggregating all entries in queue to concatenate entries.
    20:01:15 WARNING spend_update_queue.py:47 Spend update queue is full. Aggregating all entries in queue to concatenate entries.
    20:01:16 WARNING spend_update_queue.py:47 Spend update queue is full. Aggregating all entries in queue to concatenate entries.
    
    Two Redis timeout lines (a single timeout under load 3 seconds before the blip, then the 5 second summary with 202 suppressed) instead of 385 ERROR lines, zero Error in tracking cost callback lines (grep -c returns 0), and nothing new after the breaker opened even though requests kept flowing at 200

Redis hard down (non-timeout failure)

  1. Same proxy, docker stop litfix-redis-7520, then the same three sequential curls. Output: 200 200 200
  2. New log lines, one ERROR per failed cache call exactly as before:
    20:02:30 WARNING config_sync_pubsub.py:256 config sync subscriber redis error: Connection closed by server.; reconnecting in 5s
    20:02:30 WARNING auth_cache_invalidation_pubsub.py:175 auth cache invalidation subscriber redis error: Connection closed by server.; reconnecting in 5s
    20:02:32 ERROR redis_cache.py Error occurred in async batch get cache: Connection closed by server.
    20:02:33 ERROR redis_cache.py LiteLLM Redis Caching: async set() - Got exception from REDIS: Connection closed by server.
    20:02:33 ERROR redis_cache.py Error occurred in async batch get cache: Connection closed by server.
    20:02:33 WARNING redis_cache.py:252 Redis circuit breaker OPENED after 5 consecutive failures (5 hard connectivity) — fast-failing Redis calls for 60s
    20:02:35 WARNING config_sync_pubsub.py:256 config sync subscriber redis error: Connection closed by server.; reconnecting in 10s
    20:02:35 WARNING auth_cache_invalidation_pubsub.py:175 auth cache invalidation subscriber redis error: Connection closed by server.; reconnecting in 10s
    

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • The throttle is process-wide, one window shared by every Redis cache instance and message
    • A proxy with two Redis caches timing out at once gets one summary line, not two
  • The async set() message no longer includes the value being written (it was the bulk of every ERROR line)

Low

  • A Redis timeout in the spend counter increment no longer reaches the cost callback's error path or its alert. The counters are invalidated and reloaded from the database on the next read, same as an open breaker
  • _is_redis_timeout_failure is now public as is_redis_timeout_failure so the proxy can import it without a private-usage type error
  • Cache.add_cache wraps whichever backend it was built with. It checks isinstance(self.cache, RedisCache) before joining the throttle, so a timeout from a non-Redis backend (S3, disk, Qdrant) keeps its per-call ERROR line. The dynamic_cache_object argument to async_add_cache is not inspected; a Redis object passed that way while self.cache is something else logs per call, which is the pre-PR behavior
  • Before was captured at b27d2cce77, the branch's original base, not at the current merge base. git diff b27d2cce77 origin/main -- litellm/caching/redis_cache.py shows main kept every per-call verbose_logger.error this PR replaces and added one more (async_set_cache_pipeline_with_ttls), so the current base can only log more lines per blip, not fewer
  • CodeQL (not a required check) flags 3 py/cyclic-import alerts on the two import lines this PR touches. Both lines already exist on main with the same cycle: proxy_server.py:253 already imports RedisCircuitBreakerOpenError from redis_cache (one of the 3 alerts is about that pre-existing name) and redis_cache.py already imports 5 constants from litellm.constants in the same block. The PR only adds a name to each existing statement. python -c "import litellm.constants, litellm.caching.redis_cache, litellm.proxy.proxy_server" succeeds in each of the 3 import orders
  • Case 1 exercised async batch get, async_increment_pipeline() and, on earlier runs of the same storm, async set(), the rate limiter TTL pipeline and the DualCache increment fallback. Sync get_cache(), the batch-get sync path and the remaining write and list paths (async_set_cache_pipeline, async_set_cache_pipeline_with_ttls, async_set_cache_sadd, async_increment, async_rpush, async_lpop and their pipeline variants) are covered by unit tests, not the live run
    • The storm reuses the master key, so the auth prefetch write async_set_cache_pipeline_with_ttls never fires live. Its timeout path is the async_set_cache_pipeline_with_ttls case of test_write_path_timeouts_inside_the_interval_stay_at_debug, which fails on the previous head

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Low Risk
Changes are mostly logging volume and graceful handling of transient Redis timeouts; spend counters are invalidated on timeout like an open breaker, which could briefly skew in-memory spend until reload.

Overview
Redis cache failures now rate-limit timeout noise instead of emitting one ERROR per call during a slow-Redis blip.

log_redis_failure is the shared path: non-timeout Redis errors still log at the caller’s level every time; timeouts emit one line per REDIS_TIMEOUT_LOG_INTERVAL (default 5s, REDIS_TIMEOUT_LOG_INTERVAL), with repeats at DEBUG and a periodic summary of how many were suppressed. Circuit-breaker-open refusals stay DEBUG. Most RedisCache exception handlers that used raw verbose_logger.error now call this helper (messages no longer embed cache values on some paths).

Cache.add_cache / async variants use _log_add_cache_failure, which only joins the Redis throttle when self.cache is a RedisCache; other backends keep per-call ERROR logging.

is_redis_timeout_failure is public (was _is_redis_timeout_failure) for proxy spend logic: _apply_spend_counter_increments now treats Redis timeouts like an open circuit breaker—invalidate counters and return without raising—reducing cost-callback tracebacks during transient slowness.

Tests cover throttle behavior on Redis ops, add_cache, DualCache fallbacks, and spend-counter timeouts.

Reviewed by Cursor Bugbot for commit a4e34d6. Bugbot is set up for automated code reviews on this repo. Configure here.

Link to Devin session: https://app.devin.ai/sessions/eba56c82452f40a2943f80e119789fb7
Open in Devin Desktop: https://app.devin.ai/desktop/session/eba56c82452f40a2943f80e119789fb7?variant=devin
Requested by: @yassin-berriai


Devin Review

yassin-berriai and others added 2 commits September 12, 2026 01:12
…e per cache call

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…lready-logged cache failure

The cost tracking callback logged its own ERROR with a traceback for every request whose spend counter increment timed out, on top of the cache layer's throttled line. Timeouts now take the same path as breaker-open refusals: invalidate the counters and return. Also exposes is_redis_timeout_failure publicly for that caller and drops the comment on the new constant

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@codspeed

codspeed Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_redis_timeout_log_throttle (a4e34d6) with main (34fe9f7)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR reduces Redis timeout log storms while retaining per-call error visibility for non-timeout Redis failures and non-Redis cache backends.

  • Adds a process-wide, thread-safe timeout throttle driven by a monotonic clock and configurable interval.
  • Routes Redis cache operation failures through centralized logging.
  • Treats spend-counter Redis timeouts like an open circuit breaker after invalidating in-memory counters.
  • Keeps generic cache failures at ERROR unless the configured backend is Redis.
  • Adds coverage for timeout summaries, write paths, fallback behavior, spend counters, and backend-specific logging.

Confidence Score: 5/5

The PR appears safe to merge; the prior throttling and style findings are resolved, and no new actionable failures remain.

The timeout throttle now uses a monotonic clock, authentication-prefetch pipeline failures use the centralized throttle, generic non-Redis cache failures remain at ERROR, and the previously flagged test constructor is wrapped. devin-ai-integration[bot] accepted the process-wide shared-throttle behavior because it intentionally matches the process-wide circuit breaker and documented the tradeoff. The final line-length thread was manually resolved without explanation, and the current line is compliant.

Important Files Changed

Filename Overview
litellm/caching/redis_cache.py Centralizes Redis failure logging and throttles repeated timeout messages while preserving ordinary error handling.
litellm/caching/caching.py Restricts generic add-cache timeout throttling to configured Redis backends.
litellm/proxy/proxy_server.py Handles spend-counter Redis timeouts like breaker-open failures after counter invalidation.
litellm/constants.py Adds the configurable Redis timeout logging interval.
tests/test_litellm/caching/test_redis_cache.py Covers throttling intervals, summaries, non-timeout errors, and Redis write/list operations.
tests/test_litellm/caching/test_caching.py Verifies that only Redis-backed generic cache failures join the timeout throttle.
tests/test_litellm/caching/test_dual_cache.py Verifies timeout throttling when Redis operations fall back to memory and wraps the previously flagged long line.
tests/test_litellm/proxy/proxy_server/test_spend_counters.py Verifies timeout invalidation and propagation behavior for spend-counter updates.

Reviews (6): Last reviewed commit: "test(caching): wrap the DualCache fixtur..." | Re-trigger Greptile

greptile-apps[bot]

This comment was marked as resolved.

…im test docstrings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@codecov

codecov Bot commented Sep 12, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.53731% with 5 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/caching/redis_cache.py 94.23% 3 Missing ⚠️
litellm/caching/caching.py 77.77% 2 Missing ⚠️

📢 Thoughts on this report? Let us know!

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review f681a97: monotonic clock and shorter test docstrings

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

… shared log throttle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review 28f2d1f: adds unit tests for the async write and list timeout paths, no production code change

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@yassin-berriai
yassin-berriai changed the base branch from litellm_internal_staging to main September 14, 2026 19:04
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/caching/redis_cache.py
#	litellm/proxy/proxy_server.py
#	tests/test_litellm/caching/test_redis_cache.py
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review 53c8a82: merge of main, conflict resolution keeps is_redis_timeout_failure and the spend counter timeout return

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

REDIS_CIRCUIT_BREAKER_FAILURE_THRESHOLD,
REDIS_CIRCUIT_BREAKER_RECOVERY_TIMEOUT,
REDIS_CIRCUIT_BREAKER_TIMEOUT_MIN_DURATION,
REDIS_TIMEOUT_LOG_INTERVAL,
from litellm._logging import _redact_string, verbose_proxy_logger, verbose_router_logger
from litellm.caching.caching import DualCache, RedisCache
from litellm.caching.redis_cache import RedisCircuitBreakerOpenError
from litellm.caching.redis_cache import RedisCircuitBreakerOpenError, is_redis_timeout_failure
from litellm._logging import _redact_string, verbose_proxy_logger, verbose_router_logger
from litellm.caching.caching import DualCache, RedisCache
from litellm.caching.redis_cache import RedisCircuitBreakerOpenError
from litellm.caching.redis_cache import RedisCircuitBreakerOpenError, is_redis_timeout_failure

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same static cycle main already has for the neighbouring imports on these lines; importing either module first works at runtime, so leaving as is

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…shared throttle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review c4354c2, which routes async_set_cache_pipeline_with_ttls through log_redis_failure

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@veria-ai please review c4354c2 for the throttled Redis timeout logging and spend counter timeout handling

greptile-apps[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

yassin-berriai and others added 2 commits September 14, 2026 19:59
…kend is Redis

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…imit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review a4e34d6: wraps the flagged test line, generic add_cache stays at ERROR for non-Redis backends

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@veria-ai please review a4e34d6 for the throttled Redis timeout logging and the generic add_cache backend check

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit a4e34d6. Configure here.

@yassin-berriai
yassin-berriai merged commit 856aedd into main Sep 14, 2026
93 of 94 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_redis_timeout_log_throttle branch September 14, 2026 20:34
doonga pushed a commit to greyrock-labs/home-ops that referenced this pull request Sep 28, 2026
…103.0) (#290)

This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.102.1` → `v1.103.0` |

---

### Release Notes

<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>

### [`v1.103.0`](https://github.com/BerriAI/litellm/releases/tag/v1.103.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.102.1...v1.103.0)

#### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.103.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.103.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.103.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

#### What's Changed

- fix(responses): translate the reasoning object into a chat-completion reasoning effort by [@&#8203;joshgarnett](https://github.com/joshgarnett) in [#&#8203;36363](https://github.com/BerriAI/litellm/pull/36363)
- fix(proxy): bound tool and guardrail index create\_many by the spend-log statement budgets by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40561](https://github.com/BerriAI/litellm/pull/40561)
- fix(mcp): require admission for delegated OAuth by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;40923](https://github.com/BerriAI/litellm/pull/40923)
- fix(logging): log one bounded summary for a burst of timed-out LoggingWorker callbacks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40912](https://github.com/BerriAI/litellm/pull/40912)
- fix(fireworks): resolve short model names to long cost map keys by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40929](https://github.com/BerriAI/litellm/pull/40929)
- ci: remove main branch source guard by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;40172](https://github.com/BerriAI/litellm/pull/40172)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;40942](https://github.com/BerriAI/litellm/pull/40942)
- fix(spend\_logs): store litellm\_call\_id and match it in request\_id lookups by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39068](https://github.com/BerriAI/litellm/pull/39068)
- fix(auth): refresh lite login session token grants from the live user and team rows by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;40657](https://github.com/BerriAI/litellm/pull/40657)
- feat(bedrock): support file delete and list for S3-backed managed files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39836](https://github.com/BerriAI/litellm/pull/39836)
- fix(proxy): gate the webhook test alert on proxy admins by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;40814](https://github.com/BerriAI/litellm/pull/40814)
- fix(ui): hide admin write-form tabs on the models page from view-only admins by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;38867](https://github.com/BerriAI/litellm/pull/38867)
- fix(anthropic-adapter): surface mid-stream provider errors as Anthropic error events by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33352](https://github.com/BerriAI/litellm/pull/33352)
- docs(github): add an Affected release section to the PR template by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;40618](https://github.com/BerriAI/litellm/pull/40618)
- docs(e2e): ban unit tests under tests/e2e by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33852](https://github.com/BerriAI/litellm/pull/33852)
- fix(router): preserve Azure Entra ID params in reusable credentials by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40889](https://github.com/BerriAI/litellm/pull/40889)
- docs(user endpoints): remove unsupported soft\_budget param from user docstrings by [@&#8203;shivamrawat1](https://github.com/shivamrawat1) in [#&#8203;36585](https://github.com/BerriAI/litellm/pull/36585)
- feat(friendli): auto-sync Friendli model metadata into price registry by [@&#8203;Lee-Si-Yoon](https://github.com/Lee-Si-Yoon) in [#&#8203;35918](https://github.com/BerriAI/litellm/pull/35918)
- build(deps): bump smol-toml to 1.8.0 to clear GHSA-7w5x-hrqm-74c2 in osv-scan by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40478](https://github.com/BerriAI/litellm/pull/40478)
- fix(bedrock\_mantle): price GovCloud regions from the regional cost row and accept region-prefixed model names by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39846](https://github.com/BerriAI/litellm/pull/39846)
- chore(ci): remerge internal staging by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;40943](https://github.com/BerriAI/litellm/pull/40943)
- test(auth): freeze the cache clock in auth prefetch tests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40996](https://github.com/BerriAI/litellm/pull/40996)
- feat(jwt): allow virtual\_key\_claim\_field per issuer by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40927](https://github.com/BerriAI/litellm/pull/40927)
- fix(cost): bill cached realtime audio tokens at the audio cache-read rate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40627](https://github.com/BerriAI/litellm/pull/40627)
- perf(logging): skip correlation contextvar stamping when request\_correlation\_in\_logs is off by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41054](https://github.com/BerriAI/litellm/pull/41054)
- feat(pricing): add azure gpt-chat-latest rates and drop retired friendliai llama-3.1 entries by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40976](https://github.com/BerriAI/litellm/pull/40976)
- build(deps): re-suppress GHSA-h7x2-h6g9-p789 in osv-scan on main, mlflow still has no fixed release by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41104](https://github.com/BerriAI/litellm/pull/41104)
- fix(otel): cap per-index OpenInference message attributes span-wide by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40562](https://github.com/BerriAI/litellm/pull/40562)
- feat(proxy): add general\_settings.allowed\_file\_extensions for /v1/files uploads by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41106](https://github.com/BerriAI/litellm/pull/41106)
- fix(proxy): forward provider request id headers on mapped error responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40925](https://github.com/BerriAI/litellm/pull/40925)
- fix(router): name the all-deployments-in-cooldown error on 429 responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40995](https://github.com/BerriAI/litellm/pull/40995)
- fix(ui): show the team alias on the model info page and in its raw JSON by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40992](https://github.com/BerriAI/litellm/pull/40992)
- feat(proxy): honor LITELLM\_DISABLE\_ACCESS\_LOG\_PATHS to drop noisy uvicorn access log lines by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41096](https://github.com/BerriAI/litellm/pull/41096)
- fix(prometheus): label pre-call rate limit failures with the resolved api\_provider by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41059](https://github.com/BerriAI/litellm/pull/41059)
- perf(proxy): serialize /model/info listing once with orjson by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41114](https://github.com/BerriAI/litellm/pull/41114)
- fix(utils): stop wrapper\_async submitting the sync success handler twice by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41115](https://github.com/BerriAI/litellm/pull/41115)
- fix(redis): log a timeout streak once per interval instead of one line per cache call by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40817](https://github.com/BerriAI/litellm/pull/40817)
- fix(router): record flat retry attempts and cap retries from attempted\_retries by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40930](https://github.com/BerriAI/litellm/pull/40930)
- refactor(prometheus): source PROXY\_LLM\_PROVIDER\_FALLBACK from litellm.constants by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41118](https://github.com/BerriAI/litellm/pull/41118)
- fix(proxy): hide default credentials login hint when UI\_PASSWORD is set by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41107](https://github.com/BerriAI/litellm/pull/41107)
- fix(cli): show routed models and session stats for LLM API keys by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41116](https://github.com/BerriAI/litellm/pull/41116)
- fix(bedrock/realtime): propagate deferred Nova Sonic stream failures to the router by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41064](https://github.com/BerriAI/litellm/pull/41064)
- fix(proxy): keep org admins' own team memberships in other orgs visible on team list by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41086](https://github.com/BerriAI/litellm/pull/41086)
- feat(model\_info): provider-scoped fill\_missing\_for\_providers backfill from fallback generalization rules by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41093](https://github.com/BerriAI/litellm/pull/41093)
- fix(auth): load team membership once per request and skip prisma on an L1 hit by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41102](https://github.com/BerriAI/litellm/pull/41102)
- refactor(harness): expand independent trace coverage by [@&#8203;yujonglee-berri](https://github.com/yujonglee-berri) in [#&#8203;41120](https://github.com/BerriAI/litellm/pull/41120)
- fix(proxy): release max\_parallel\_requests slot when a realtime session ends without LLM callbacks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41113](https://github.com/BerriAI/litellm/pull/41113)
- fix(router): cool down team deployments on 429 when a sibling serves the same public model by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40991](https://github.com/BerriAI/litellm/pull/40991)
- fix(ui): move tags typed into key metadata JSON into the Tags field by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41023](https://github.com/BerriAI/litellm/pull/41023)
- fix(ui): let team admins grant a team all proxy models by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;40196](https://github.com/BerriAI/litellm/pull/40196)
- chore(lint): graduate 12 rules from the strict-gate ratchet by [@&#8203;HUAHAODIA](https://github.com/HUAHAODIA) in [#&#8203;41048](https://github.com/BerriAI/litellm/pull/41048)
- test: add dedicated CircleCI integration contract foundation by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41066](https://github.com/BerriAI/litellm/pull/41066)
- fix(utils): keep litellm params out of provider request bodies by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41018](https://github.com/BerriAI/litellm/pull/41018)
- fix(openai): keep extra\_headers out of the chat request body on the httpx handler path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41141](https://github.com/BerriAI/litellm/pull/41141)
- test: cover persisted updates and warmed authorization policies by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41070](https://github.com/BerriAI/litellm/pull/41070)
- fix(proxy): resolve x-litellm-call-id from response metadata when routes omit call\_id by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41056](https://github.com/BerriAI/litellm/pull/41056)
- chore(prices): sync Vertex AI prices: 14 models by [@&#8203;berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#&#8203;40955](https://github.com/BerriAI/litellm/pull/40955)
- ci(codeql): exclude noisy Python quality queries by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41142](https://github.com/BerriAI/litellm/pull/41142)
- feat(proxy): predict prompt-cache costs across deployments by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;40877](https://github.com/BerriAI/litellm/pull/40877)
- fix(prompt\_security): keep polling file sanitization through non-terminal statuses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41131](https://github.com/BerriAI/litellm/pull/41131)
- fix(health): skip background health check DB writes when the latest-row read fails by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41145](https://github.com/BerriAI/litellm/pull/41145)
- feat(model\_armor): logging\_only mode scans completed streams after delivery by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40702](https://github.com/BerriAI/litellm/pull/40702)
- fix(bedrock guardrails): derive contextual grounding source and query from plain messages by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41132](https://github.com/BerriAI/litellm/pull/41132)
- fix(cli): drop enum.StrEnum so the CLI imports on Python 3.10 by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41046](https://github.com/BerriAI/litellm/pull/41046)
- fix(responses): route mid-stream error events through exception\_type so content\_policy\_fallbacks fire by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40988](https://github.com/BerriAI/litellm/pull/40988)
- fix(cost): bill gemini-embedding-2 per token and stop double charging audio by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41157](https://github.com/BerriAI/litellm/pull/41157)
- test: bind management E2E callers and isolate JWT actors by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;40892](https://github.com/BerriAI/litellm/pull/40892)
- fix(headroom): protect cache\_control-marked rows anywhere in history by [@&#8203;rad-p44](https://github.com/rad-p44) in [#&#8203;40315](https://github.com/BerriAI/litellm/pull/40315)
- test: add strict stateless provider replay identity by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41149](https://github.com/BerriAI/litellm/pull/41149)
- fix(ci): test checked-out model pricing in unit jobs by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41181](https://github.com/BerriAI/litellm/pull/41181)
- fix(guardrails): write per-message guardrail rewrites back onto Responses input items by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40939](https://github.com/BerriAI/litellm/pull/40939)
- fix(proxy): log the provider usage on deferred /v1/messages calls and price cache writes without a creation rate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41172](https://github.com/BerriAI/litellm/pull/41172)
- fix(guardrails): record not\_run evaluation when scoping leaves nothing to scan by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;39050](https://github.com/BerriAI/litellm/pull/39050)
- fix(responses): preserve provider affinity by [@&#8203;AaronHowell](https://github.com/AaronHowell) in [#&#8203;40228](https://github.com/BerriAI/litellm/pull/40228)
- fix(sdk): keep body and proxy headers on BadRequestError mapped from a litellm\_proxy 400 by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40994](https://github.com/BerriAI/litellm/pull/40994)
- fix(responses): hoist Codex additional\_tools input items into the chat bridge tools by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40989](https://github.com/BerriAI/litellm/pull/40989)
- fix(router): honor team and key provider weights by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41072](https://github.com/BerriAI/litellm/pull/41072)
- test(e2e): verify streamed answers and tool continuation by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41194](https://github.com/BerriAI/litellm/pull/41194)
- fix(cli): label router costs and simplify the routed-model header by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41186](https://github.com/BerriAI/litellm/pull/41186)
- test(spend): reconcile concurrent requests and daily activity by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41188](https://github.com/BerriAI/litellm/pull/41188)
- fix(guardrails): scan the Anthropic top-level system prompt and tool\_use arguments by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40984](https://github.com/BerriAI/litellm/pull/40984)
- fix(router): count num\_retries\_per\_request across fallback hops by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41191](https://github.com/BerriAI/litellm/pull/41191)
- fix(bedrock): grant rerank, retrieve, agent, and agentcore actions in the web identity session policy by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41168](https://github.com/BerriAI/litellm/pull/41168)
- fix(vertex-live): bill Gemini Live sessions end to end (internal copy of [#&#8203;37075](https://github.com/BerriAI/litellm/issues/37075)) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40915](https://github.com/BerriAI/litellm/pull/40915)
- fix(health): resolve litellm\_credential\_name in realtime health checks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41173](https://github.com/BerriAI/litellm/pull/41173)
- feat(proxy): unified custom\_key\_policy hook for key generate, update and regenerate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40921](https://github.com/BerriAI/litellm/pull/40921)
- fix(proxy): enforce custom\_key\_update policy on /key/regenerate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40695](https://github.com/BerriAI/litellm/pull/40695)
- fix(router): preserve session model choice within each complexity tier by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41174](https://github.com/BerriAI/litellm/pull/41174)
- test(pricing): let synced GovCloud Bedrock rows cite the AWS price list by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41263](https://github.com/BerriAI/litellm/pull/41263)
- docs(github): ask for interactive coding-tool proof in the PR template by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41257](https://github.com/BerriAI/litellm/pull/41257)
- feat(proxy): add POST /management/v1/users/bulk for batched user and team membership creation by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41028](https://github.com/BerriAI/litellm/pull/41028)
- fix(credentials): answer 409 on a credential name collision, make Terraform adoption opt-in by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;40917](https://github.com/BerriAI/litellm/pull/40917)
- feat(proxy): add POST /management/v1/users/bulk\_delete and POST /management/v1/teams/{team\_id}/members/bulk\_delete by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41039](https://github.com/BerriAI/litellm/pull/41039)
- fix(proxy): list directly assigned team models in model access errors by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41256](https://github.com/BerriAI/litellm/pull/41256)
- feat(auto-router): allow opted-in team members to manage their routers by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41175](https://github.com/BerriAI/litellm/pull/41175)
- build(rust-bridge): add typed \_native stub and validate it with mypy.stubtest by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41180](https://github.com/BerriAI/litellm/pull/41180)
- feat(guardrails): add new upstream presidio pii entities including german set by [@&#8203;MvdB](https://github.com/MvdB) in [#&#8203;36775](https://github.com/BerriAI/litellm/pull/36775)
- fix(responses): filter bridged kwargs like the native Responses path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41144](https://github.com/BerriAI/litellm/pull/41144)
- test(e2e): cover the reliability retry, cooldown, fallback, and routing-strategy cells by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39857](https://github.com/BerriAI/litellm/pull/39857)
- fix(anthropic): add the per-turn-control beta when a message carries output\_config by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41189](https://github.com/BerriAI/litellm/pull/41189)
- fix(router): bind per-request routing\_strategy override selectors to the request's callbacks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41178](https://github.com/BerriAI/litellm/pull/41178)
- feat(proxy): bind JWT claims to registered agents via agent\_id\_jwt\_field by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40904](https://github.com/BerriAI/litellm/pull/40904)
- fix(proxy): enforce organization budgets when max\_budget is 0 by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;41271](https://github.com/BerriAI/litellm/pull/41271)
- fix(alerting): send llm\_exceptions Slack alert for 5xx HTTPException and ProxyException by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41125](https://github.com/BerriAI/litellm/pull/41125)
- fix(headroom): protect the cached prefix through the last cache\_control breakpoint by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41161](https://github.com/BerriAI/litellm/pull/41161)
- fix(utils): cache custom HuggingFace tokenizers across /utils/token\_counter requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41216](https://github.com/BerriAI/litellm/pull/41216)
- fix(router): keep weighted routing when a deployment id equals a model\_name by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41156](https://github.com/BerriAI/litellm/pull/41156)
- feat(router): add capability classifier as Fuse foundation by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41270](https://github.com/BerriAI/litellm/pull/41270)
- fix(proxy): keep access-group raw SQL writes on the writer while writer\_unavailable is stale by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41283](https://github.com/BerriAI/litellm/pull/41283)
- fix(prometheus): count 401 auth failures in litellm\_proxy\_failed\_requests\_metric by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41170](https://github.com/BerriAI/litellm/pull/41170)
- test: drop tests that pin vendor facts and add the CLAUDE.md rule by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41269](https://github.com/BerriAI/litellm/pull/41269)
- fix(proxy): run the remaining inline token counts off the event loop by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40262](https://github.com/BerriAI/litellm/pull/40262)
- fix(proxy): log blocked streaming guardrail responses as failures, not success by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40191](https://github.com/BerriAI/litellm/pull/40191)
- feat(proxy): add tpd\_limit (tokens per day) for batch submissions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40997](https://github.com/BerriAI/litellm/pull/40997)
- fix(proxy): reconcile budget reservation before enqueuing spend to the DB by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40310](https://github.com/BerriAI/litellm/pull/40310)
- fix(xai): stop sending web\_search\_options to xAI's retired Live Search path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;38278](https://github.com/BerriAI/litellm/pull/38278)
- feat(terraform): add tpm\_limit, rpm\_limit, budget\_duration, allowed\_models to litellm\_team\_member\_add by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;38682](https://github.com/BerriAI/litellm/pull/38682)
- fix(rerank): bill Vertex search\_units from input records and give every rerank response a unique id by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35180](https://github.com/BerriAI/litellm/pull/35180)
- fix(router): stop counting caller-set timeout 408s toward deployment cooldown by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41230](https://github.com/BerriAI/litellm/pull/41230)
- feat(router): add Fuse V2 classifier after capability forecasting by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41272](https://github.com/BerriAI/litellm/pull/41272)
- fix(proxy): keep client User-Agent on auth failure spend logs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41291](https://github.com/BerriAI/litellm/pull/41291)
- fix(proxy): reset budgets by decrementing pre-reset spend instead of zeroing rows by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41279](https://github.com/BerriAI/litellm/pull/41279)
- fix(xai): honor nested web\_search filters on the xAI Responses API by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;38268](https://github.com/BerriAI/litellm/pull/38268)
- fix(router): stop registering a caller-supplied credential as a router deployment by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;41289](https://github.com/BerriAI/litellm/pull/41289)
- fix(router): accept custom\_provider\_map providers before the first completion call by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41300](https://github.com/BerriAI/litellm/pull/41300)
- fix(proxy): return 400 instead of 500 for lone surrogate escapes in request body by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41297](https://github.com/BerriAI/litellm/pull/41297)
- fix(langsmith): keep events appended during an in-flight flush instead of clearing them by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41288](https://github.com/BerriAI/litellm/pull/41288)
- fix(logging): track spend for streams a deployment hook converted to non-streaming by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41171](https://github.com/BerriAI/litellm/pull/41171)
- fix(bedrock): sanitize client tool\_call ids to Bedrock toolUseId constraints by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40872](https://github.com/BerriAI/litellm/pull/40872)
- fix(passthrough): attribute Vertex passthrough successes to the resolved router deployment by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41307](https://github.com/BerriAI/litellm/pull/41307)
- feat(ui): persist Models table search, filters, sort and page in the URL by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41296](https://github.com/BerriAI/litellm/pull/41296)
- fix(jwt-auth): scope JWT key mappings by issuer to prevent cross-issuer collisions by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;41281](https://github.com/BerriAI/litellm/pull/41281)
- feat(openai): add openai\_system\_messages\_first to put system messages first for prompt caching by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41304](https://github.com/BerriAI/litellm/pull/41304)
- feat(ui): add custom request headers to the API Playground by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41309](https://github.com/BerriAI/litellm/pull/41309)
- feat(cli): sync Codex /model picker from proxy /v1/models in lite codex by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40476](https://github.com/BerriAI/litellm/pull/40476)
- chore: bump litellm-enterprise 0.1.67 -> 0.1.68, litellm-proxy-extras 0.4.97 -> 0.4.98, litellm 1.102.0 -> 1.103.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41321](https://github.com/BerriAI/litellm/pull/41321)
- feat: add aihubmix provider pricing entries by [@&#8203;IToSSc](https://github.com/IToSSc) in [#&#8203;41179](https://github.com/BerriAI/litellm/pull/41179)
- feat(auto-router): add per-model Fast mode toggle by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41282](https://github.com/BerriAI/litellm/pull/41282)
- fix(proxy): include litellm\_call\_id in LLM API exception logs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41205](https://github.com/BerriAI/litellm/pull/41205)
- fix(proxy): keep yaml pass-through endpoints visible to auth after db overlay by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41303](https://github.com/BerriAI/litellm/pull/41303)
- fix(proxy): resolve router\_settings.model\_group\_alias before key/team model auth by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41308](https://github.com/BerriAI/litellm/pull/41308)
- fix(ui): block usage export and flag the range when a spend page fails by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41294](https://github.com/BerriAI/litellm/pull/41294)
- fix(vertex\_ai): bill Gemini Omni Interactions usage and Veo sampleCount on passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41322](https://github.com/BerriAI/litellm/pull/41322)
- fix(proxy): honor LITELLM\_LOG for uvicorn and proxy extras loggers by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41306](https://github.com/BerriAI/litellm/pull/41306)
- fix(cost): price native Responses WebSocket turns at their returned service\_tier by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41318](https://github.com/BerriAI/litellm/pull/41318)
- fix(proxy): key model rpm/tpm override takes precedence over team model limit by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41302](https://github.com/BerriAI/litellm/pull/41302)
- fix(proxy): track per-member organization spend by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41255](https://github.com/BerriAI/litellm/pull/41255)
- feat(proxy): add /nvidia\_nim passthrough route for NIM object detection and OCR /v1/infer by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41316](https://github.com/BerriAI/litellm/pull/41316)
- feat(model\_info): add provider-neutral Gemini 2.5+ chat baseline fallback generalization by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41320](https://github.com/BerriAI/litellm/pull/41320)
- fix(spend): sum multi-round session duration in logs UI by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35388](https://github.com/BerriAI/litellm/pull/35388)
- feat(router): limit unlicensed Capability and Fuse v2 routers to one each by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41326](https://github.com/BerriAI/litellm/pull/41326)
- fix(guardrails): resolve caller identity from metadata buckets in custom code guardrail by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41126](https://github.com/BerriAI/litellm/pull/41126)
- fix(e2e): onboard dashboard users through invitations by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41319](https://github.com/BerriAI/litellm/pull/41319)
- feat(guardrails): add Microsoft Agent 365 MCP tool-call guardrail by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;38241](https://github.com/BerriAI/litellm/pull/38241)
- test: drop remaining tests that pin cost-map vendor facts by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41298](https://github.com/BerriAI/litellm/pull/41298)
- feat(ui): show average response time per model in usage model activity by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41313](https://github.com/BerriAI/litellm/pull/41313)
- fix(proxy): preserve Anthropic pricing modifiers in router savings by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41341](https://github.com/BerriAI/litellm/pull/41341)
- feat(guardrails): support pre\_call and during\_call modes for llm\_as\_a\_judge by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41128](https://github.com/BerriAI/litellm/pull/41128)
- fix(gemini): propagate the provider's modelVersion to the response model by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41338](https://github.com/BerriAI/litellm/pull/41338)
- fix(fireworks-ai): bill cache-write, reasoning and audio tokens via the shared cost calculator by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41339](https://github.com/BerriAI/litellm/pull/41339)
- feat(guardrails): singulr v2 API contract with logging\_only, pre\_mcp\_call and post\_mcp\_call by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;41329](https://github.com/BerriAI/litellm/pull/41329)
- ci(image-scan): ignore zlib CVE-2026-85091 until Wolfi ships the fix by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41353](https://github.com/BerriAI/litellm/pull/41353)
- feat(e2e): reuse exact provider responses for 24 hours by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41346](https://github.com/BerriAI/litellm/pull/41346)
- fix(xai): keep 'instructions' on the xAI Responses API so system messages survive web search by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;38254](https://github.com/BerriAI/litellm/pull/38254)
- feat(ui): configure capability and Fuse v2 classifiers by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41315](https://github.com/BerriAI/litellm/pull/41315)
- fix(anthropic): tolerate message\_delta events without usage when streaming by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41336](https://github.com/BerriAI/litellm/pull/41336)
- test(router): ignore deployment-selection logs in the fallback log assertion by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41358](https://github.com/BerriAI/litellm/pull/41358)
- test(proxy): assert budget resets decrement the cleared spend by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41359](https://github.com/BerriAI/litellm/pull/41359)
- fix(e2e): expect models filters to persist after reload by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41348](https://github.com/BerriAI/litellm/pull/41348)
- fix(e2e): record cookie-setting provider responses and keep prompt-caching tests live by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41366](https://github.com/BerriAI/litellm/pull/41366)
- fix(responses): recount tokens when a streamed response completes without usage by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41337](https://github.com/BerriAI/litellm/pull/41337)
- fix(ui): simplify Capability and Fuse advanced routing options by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41371](https://github.com/BerriAI/litellm/pull/41371)
- fix(mcp): authorize JWT OAuth credential persistence by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;41314](https://github.com/BerriAI/litellm/pull/41314)
- feat(router): stream shadow traffic and fan out silent\_model to multiple targets by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41368](https://github.com/BerriAI/litellm/pull/41368)
- perf(content\_filter): scan a bounded window per streamed chunk by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41407](https://github.com/BerriAI/litellm/pull/41407)
- fix(proxy): hide model allowlist from client-facing model access denied errors by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41310](https://github.com/BerriAI/litellm/pull/41310)
- feat(http): opt-in outbound HTTP/2 for httpx clients by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41268](https://github.com/BerriAI/litellm/pull/41268)
- refactor(rust): remove gateway, config, router, realtime, and Rust trace-parity instrumentation by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41432](https://github.com/BerriAI/litellm/pull/41432)
- fix(guardrails): don't add post\_call output scan for MCP-only Presidio modes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40571](https://github.com/BerriAI/litellm/pull/40571)
- chore(prices): sync Azure, Azure AI, Gemini, OpenAI, Bedrock, Together AI, Fireworks and Vertex prices: 278 models, 59 new, 30 deprecated by [@&#8203;berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#&#8203;41154](https://github.com/BerriAI/litellm/pull/41154)
- fix(rag): forward retrieval\_filter from retrieval\_config to vector store search by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34427](https://github.com/BerriAI/litellm/pull/34427)
- refactor(rust): extract auth and cache crates by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41464](https://github.com/BerriAI/litellm/pull/41464)
- fix(proxy): default litellm\_trace\_id to the OTel server span trace id by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41386](https://github.com/BerriAI/litellm/pull/41386)
- chore(codeowners): add ryan and kerry as owners of the cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41333](https://github.com/BerriAI/litellm/pull/41333)
- fix(responses): guard empty-choices chunks in the Responses API streaming bridge by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34455](https://github.com/BerriAI/litellm/pull/34455)
- chore(prices): sync Google Gemini prices: 22 models by [@&#8203;berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#&#8203;41457](https://github.com/BerriAI/litellm/pull/41457)
- fix(bedrock): forward userContext in Knowledge Base Retrieve requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41475](https://github.com/BerriAI/litellm/pull/41475)
- ci(rust): split rust jobs, use nextest and Swatinem/rust-cache by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41480](https://github.com/BerriAI/litellm/pull/41480)
- fix(fireworks\_ai): flatten dict-form reasoning\_effort to its effort string by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41335](https://github.com/BerriAI/litellm/pull/41335)
- fix(proxy): never forward the LiteLLM virtual key to Anthropic on the /anthropic passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41340](https://github.com/BerriAI/litellm/pull/41340)
- fix(proxy): rename AWS Secrets Manager secret when key alias changes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41468](https://github.com/BerriAI/litellm/pull/41468)
- feat(otel): promote nested request metadata keys to litellm.metadata.\* span attributes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41462](https://github.com/BerriAI/litellm/pull/41462)
- fix(proxy): sync AWS Secrets Manager on body-less key regenerate and key alias changes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41458](https://github.com/BerriAI/litellm/pull/41458)
- fix(http\_handler): keep a handler alive while a response it issued is still reading by [@&#8203;max-sixty](https://github.com/max-sixty) in [#&#8203;34829](https://github.com/BerriAI/litellm/pull/34829)
- fix(bedrock): make prompt caching work on the Nova InvokeModel route by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41343](https://github.com/BerriAI/litellm/pull/41343)
- ci(migrations): flag defaulted ADD COLUMN on request-log tables by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41460](https://github.com/BerriAI/litellm/pull/41460)
- feat(prometheus): add customer (end\_user) budget gauges by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41472](https://github.com/BerriAI/litellm/pull/41472)
- fix(otel): drop None metric and event attributes before OTLP export by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36815](https://github.com/BerriAI/litellm/pull/36815)
- fix(anthropic): carry the served model from message\_start onto stream chunks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41446](https://github.com/BerriAI/litellm/pull/41446)
- fix(models): rolling registry audit: Gemini latest aliases, Nova cache pricing, OpenRouter/Together sync, Mistral GLM 5.3, Azure snapshots, Grok caching by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41112](https://github.com/BerriAI/litellm/pull/41112)
- fix(router): count TPM/RPM usage before building rate-limit headers by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41474](https://github.com/BerriAI/litellm/pull/41474)
- feat(guardrails): release buffered stream chunks after each passing scan by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41425](https://github.com/BerriAI/litellm/pull/41425)
- fix!: re-check budget on router fallback targets by [@&#8203;runjivu](https://github.com/runjivu) in [#&#8203;41379](https://github.com/BerriAI/litellm/pull/41379)
- refactor(ocr): move file preparation from the python bridge into litellm-core by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41489](https://github.com/BerriAI/litellm/pull/41489)
- feat(s3): add s3\_log\_prompts\_only option to log prompts without responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41327](https://github.com/BerriAI/litellm/pull/41327)
- feat(team): team-level model\_max\_budget with key-level overrides by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41330](https://github.com/BerriAI/litellm/pull/41330)
- feat(keys): filter /key/list by active, expired, revoked or deleted status and serve deleted keys from /key/info by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41311](https://github.com/BerriAI/litellm/pull/41311)
- feat(proxy): expose lifetime total\_spend on virtual keys by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41403](https://github.com/BerriAI/litellm/pull/41403)
- fix(proxy): release completed max-parallel slots promptly by [@&#8203;elifozdamar](https://github.com/elifozdamar) in [#&#8203;40843](https://github.com/BerriAI/litellm/pull/40843)
- feat(ui): accept ssh clone urls when registering a skill by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35418](https://github.com/BerriAI/litellm/pull/35418)
- fix(prices): dedupe Nova cache\_read\_input\_token\_cost keys left by a text merge by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41496](https://github.com/BerriAI/litellm/pull/41496)
- fix(otel): propagate W3C trace context on HTTP and WebSocket passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40669](https://github.com/BerriAI/litellm/pull/40669)
- fix(proxy): remove duplicate user budget hook that 429'd zero-cost models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41345](https://github.com/BerriAI/litellm/pull/41345)
- test(logging): pick this test's own records out of the shared log batch by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41487](https://github.com/BerriAI/litellm/pull/41487)
- test(together\_ai): move request-shape checks to the mapped file, drop the live ones by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41360](https://github.com/BerriAI/litellm/pull/41360)
- feat(ui): shared URL-state layer for tables and tabs by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41331](https://github.com/BerriAI/litellm/pull/41331)
- feat(e2e): make the provider cache reusable across builds and mount Bedrock behind it by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41402](https://github.com/BerriAI/litellm/pull/41402)
- fix(mcp): fail closed on missing upstream credentials by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;41364](https://github.com/BerriAI/litellm/pull/41364)
- feat(rust): scaffold Redis cache crate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41501](https://github.com/BerriAI/litellm/pull/41501)
- fix(dashscope): forward reasoning\_effort to the provider by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37506](https://github.com/BerriAI/litellm/pull/37506)
- fix(proxy): carry litellm\_call\_id through endpoint specific error logs and failure responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41356](https://github.com/BerriAI/litellm/pull/41356)
- fix(proxy): retry rate-limit fallbacks from a pristine request snapshot by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40596](https://github.com/BerriAI/litellm/pull/40596)
- fix(gemini): map minimal thinking to low for Gemini 3.7 and 3.8 Flash by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41201](https://github.com/BerriAI/litellm/pull/41201)
- fix(proxy): stop forwarding LiteLLM credential headers on Bedrock agent-runtime passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41504](https://github.com/BerriAI/litellm/pull/41504)
- fix(streaming): estimate interrupted Anthropic stream usage from reasoning\_content by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41503](https://github.com/BerriAI/litellm/pull/41503)
- fix(azure\_ai): route Responses API to native /openai/v1/responses for Foundry Models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33856](https://github.com/BerriAI/litellm/pull/33856)
- fix(proxy): show all model groups to proxy admins in /model\_group/info by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41094](https://github.com/BerriAI/litellm/pull/41094)
- feat(proxy): let proxy admins choose which team fields team admins may edit by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;39996](https://github.com/BerriAI/litellm/pull/39996)
- fix(bedrock): neutralize orphaned tool blocks instead of raising or injecting a dummy tool (internal copy of [#&#8203;31400](https://github.com/BerriAI/litellm/issues/31400)) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41513](https://github.com/BerriAI/litellm/pull/41513)
- feat(ui): persist organizations and projects list, detail tab and key table state in the URL by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41445](https://github.com/BerriAI/litellm/pull/41445)
- fix(bedrock\_mantle): accept and forward verbosity on gpt-5.x chat completions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41509](https://github.com/BerriAI/litellm/pull/41509)
- test: cover database transactions and persisted accounting by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41073](https://github.com/BerriAI/litellm/pull/41073)
- test: provider wire contracts, streaming and recovery by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41075](https://github.com/BerriAI/litellm/pull/41075)
- fix(mcp): count admin static headers as api\_key credential slots by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41514](https://github.com/BerriAI/litellm/pull/41514)
- ci: auto-merge provider-info-sync PRs when CI, Greptile and Bugbot are clean by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41494](https://github.com/BerriAI/litellm/pull/41494)
- feat(rust): add standalone framing crate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41500](https://github.com/BerriAI/litellm/pull/41500)
- fix(proxy): enforce tag budgets for tags added by guardrails by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40842](https://github.com/BerriAI/litellm/pull/40842)
- fix(utils): run post-call deployment hook on converted chat streams by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41495](https://github.com/BerriAI/litellm/pull/41495)
- fix(e2e): bind provider-cache recordings to the deployment's test, not the serving process by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41520](https://github.com/BerriAI/litellm/pull/41520)
- test: add extension and browser integration contracts by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41078](https://github.com/BerriAI/litellm/pull/41078)
- fix(logging): scan each log record once and collapse base64 payloads before the secret regex by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40934](https://github.com/BerriAI/litellm/pull/40934)
- fix(spend\_tracking): attribute router-rejected requests to the model group provider by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41507](https://github.com/BerriAI/litellm/pull/41507)
- feat(router): discover token limits for hosted OpenAI-compatible models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41508](https://github.com/BerriAI/litellm/pull/41508)
- feat(proxy): let team admins edit rpm\_limit and max\_budget when enabled by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41525](https://github.com/BerriAI/litellm/pull/41525)
- fix(otel): fit per-index OpenInference messages to the span's remaining attribute budget by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41498](https://github.com/BerriAI/litellm/pull/41498)
- test(aws): verify rotated secret value by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41524](https://github.com/BerriAI/litellm/pull/41524)
- test(e2e): read a deleted key back as deleted, not as a 404 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41551](https://github.com/BerriAI/litellm/pull/41551)
- fix(otel v2): map the caller's Langfuse user, session and tags onto the root and generation spans by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41140](https://github.com/BerriAI/litellm/pull/41140)
- test: fix seven tests left stale by [#&#8203;41311](https://github.com/BerriAI/litellm/issues/41311), [#&#8203;41337](https://github.com/BerriAI/litellm/issues/41337), [#&#8203;39996](https://github.com/BerriAI/litellm/issues/39996), [#&#8203;41310](https://github.com/BerriAI/litellm/issues/41310), [#&#8203;41289](https://github.com/BerriAI/litellm/issues/41289) and [#&#8203;41315](https://github.com/BerriAI/litellm/issues/41315) by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41527](https://github.com/BerriAI/litellm/pull/41527)
- test(budgets): cover management null handling by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41563](https://github.com/BerriAI/litellm/pull/41563)
- test(e2e): drop the auto-router select "opens below" spec by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41568](https://github.com/BerriAI/litellm/pull/41568)
- fix(guardrails): stream Prompt Security post\_call redactions in incremental\_diff mode by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41558](https://github.com/BerriAI/litellm/pull/41558)
- fix(guardrails): give post-call scans the scoped request conversation and tools by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41220](https://github.com/BerriAI/litellm/pull/41220)
- feat(openrouter): add stealth/union-alpha to the model cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41576](https://github.com/BerriAI/litellm/pull/41576)
- test(management): cover project authorization lifecycle by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41573](https://github.com/BerriAI/litellm/pull/41573)
- feat(rust): map Anthropic Messages transformations by [@&#8203;yujonglee-berri](https://github.com/yujonglee-berri) in [#&#8203;41531](https://github.com/BerriAI/litellm/pull/41531)
- fix(e2e): clear the three standing errors in the scheduled Buildkite suite by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41616](https://github.com/BerriAI/litellm/pull/41616)
- refactor(rust\_bridge): declarative route catalog and shared runtime selection by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41479](https://github.com/BerriAI/litellm/pull/41479)
- fix(mock\_completion): keep the resolved provider so router custom pricing resolves for azure\_ai deployments by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41623](https://github.com/BerriAI/litellm/pull/41623)
- fix(mcp): restrict health discovery to virtual key grants by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;41609](https://github.com/BerriAI/litellm/pull/41609)
- fix(mcp): preserve request-selected guardrails during tool execution by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;41619](https://github.com/BerriAI/litellm/pull/41619)
- refactor(ocr): mirror Python provider layout and preserve tests by [@&#8203;yujonglee-berri](https://github.com/yujonglee-berri) in [#&#8203;41550](https://github.com/BerriAI/litellm/pull/41550)
- test(fireworks\_ai): stop pinning vision support on minimax-m3 by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41627](https://github.com/BerriAI/litellm/pull/41627)
- perf(spend\_tracking): index LiteLLM\_SpendLogs by (api\_key, startTime) by [@&#8203;etiennechabert](https://github.com/etiennechabert) in [#&#8203;37983](https://github.com/BerriAI/litellm/pull/37983)
- fix(proxy): reject non-string model with 400 and log its spend as unknown-model by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41633](https://github.com/BerriAI/litellm/pull/41633)
- test(together\_ai): stop pinning successor deprecation status by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41635](https://github.com/BerriAI/litellm/pull/41635)
- chore(prices): sync Together AI prices: 6 models, 6 deprecated \[sync failed: Google Gemini] by [@&#8203;berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#&#8203;41570](https://github.com/BerriAI/litellm/pull/41570)
- fix(budgets): page end-user cache invalidation after a budget reset by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41488](https://github.com/BerriAI/litellm/pull/41488)
- chore: bump litellm-proxy-extras 0.4.98 -> 0.4.99 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41659](https://github.com/BerriAI/litellm/pull/41659)
- fix(tests): resolve the integration support package without run.py's PYTHONPATH by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41373](https://github.com/BerriAI/litellm/pull/41373)
- fix(ui): keep untimed guardrail entries on the request lifecycle by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41374](https://github.com/BerriAI/litellm/pull/41374)
- fix(anthropic-bridge): convert mid-conversation system turns to user turns on /v1/messages to chat completions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41493](https://github.com/BerriAI/litellm/pull/41493)
- fix(bedrock): support aws-sdk-bedrock-runtime 0.10/0.11 in Bedrock Realtime by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41542](https://github.com/BerriAI/litellm/pull/41542)
- feat(cli): deprecate the litellm-proxy entrypoint in favour of lite by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41673](https://github.com/BerriAI/litellm/pull/41673)
- fix(scim): align pagination `count` validation with RFC 7644 by [@&#8203;zachbernstein-sdx](https://github.com/zachbernstein-sdx) in [#&#8203;41444](https://github.com/BerriAI/litellm/pull/41444)
- fix(bedrock): never emit Converse cachePoint for OpenAI-family models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41419](https://github.com/BerriAI/litellm/pull/41419)
- fix(images): stop forwarding the raw image\[] and mask\[] form keys by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39512](https://github.com/BerriAI/litellm/pull/39512)
- feat(management\_v1): bulk update team member budgets by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41632](https://github.com/BerriAI/litellm/pull/41632)
- refactor(rust): extract provider translations by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41690](https://github.com/BerriAI/litellm/pull/41690)
- feat(cli): rename lite autoroute up/down to start/stop, keeping the old names as deprecated aliases by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41672](https://github.com/BerriAI/litellm/pull/41672)
- fix(responses): keep the addressed response id off bridged provider requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41689](https://github.com/BerriAI/litellm/pull/41689)
- fix(license): let a wildcard allowed\_features license grant the auto\_router feature by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41684](https://github.com/BerriAI/litellm/pull/41684)
- fix(ui): list every provider in the cache leakage by-model table by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40875](https://github.com/BerriAI/litellm/pull/40875)
- fix(team): keep a forked member budget's reset window and audit bulk member budget writes by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41686](https://github.com/BerriAI/litellm/pull/41686)
- feat(proxy): add TypeSafe AI Jev evaluate passthrough with registry-priced spend tracking by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41607](https://github.com/BerriAI/litellm/pull/41607)
- test(e2e): cover bedrock batch file upload and create in the us-gov-west-1 partition by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41536](https://github.com/BerriAI/litellm/pull/41536)
- feat(grafana): add all-metrics dashboard and fix stale dashboard\_v2 gauges by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41578](https://github.com/BerriAI/litellm/pull/41578)
- fix(cost): price Azure PTU spillover requests at standard token rates by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41569](https://github.com/BerriAI/litellm/pull/41569)
- build(deps): bump soupsieve to 2.9.2 to clear the osv-scan advisories by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41703](https://github.com/BerriAI/litellm/pull/41703)
- fix(fireworks\_ai): restore supports\_vision on minimax-m3 in the cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41699](https://github.com/BerriAI/litellm/pull/41699)
- feat(policy\_engine): explicit priority for policy attachment execution order by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41571](https://github.com/BerriAI/litellm/pull/41571)
- feat(router): add TypeSafe Jev as a complexity router classifier by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41615](https://github.com…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants