fix(proxy): persist periodic reload schedule state so status survives restarts and fires without store_model_in_db - #35165
Conversation
… restarts and fires without store_model_in_db The model cost map and Anthropic beta headers reload schedules kept their last-run time in a per-pod module global, so GET /schedule/*/status reported last_run null after any restart and the Admin UI showed the reload as never having run. The reload check also only ran from the add_deployment job, which is registered only when store_model_in_db is true, so config-file deployments stored a schedule that never fired. Persist last_run_at and reload_requested_at as dedicated columns on LiteLLM_Config, owned by the reload job and manual reload endpoints, while the schedule endpoints own the param_value JSON (interval_hours); no writer can clobber another's fields. Serve status entirely from the row. Register the check as its own periodic_reload_job outside the store_model_in_db gate. Replace the force_reload boolean with a reload_requested_at timestamp each pod compares against its own in-memory last reload, so a manual reload reaches every pod exactly once instead of being cleared by the first poller. Run the blocking fetches via asyncio.to_thread, and stamp last_run_at with update_many so a schedule cancelled mid-poll is not resurrected.
Greptile SummaryThis PR persists model-cost-map reload state and coordinates fleet-wide reloads through a database-backed revision
Confidence Score: 4/5The PR is not yet safe to merge because database upgrades can discard a pending legacy fleet-wide reload request Existing rows carrying force_reload are initialized with reload_revision 0, while the migration performs no conversion and the new parser intentionally ignores force_reload, so an unconsumed request can disappear during upgrade and leave other pods on stale pricing data Files Needing Attention: litellm-proxy-extras/litellm_proxy_extras/migrations/20260729000000_add_reload_tracking_to_litellm_config/migration.sql
|
| Filename | Overview |
|---|---|
| litellm/proxy/common_utils/periodic_reload_schedule.py | Introduces typed persistence helpers and revision-based due checks for periodic and fleet-wide model-cost-map reloads |
| litellm/proxy/proxy_server.py | Registers the independent reload job, applies revisions only after successful reload bookkeeping, and updates the management endpoints |
| litellm/litellm_core_utils/get_model_cost_map.py | Records the actual initial cost-map load time used to seed per-pod scheduling state |
| litellm-proxy-extras/litellm_proxy_extras/migrations/20260729000000_add_reload_tracking_to_litellm_config/migration.sql | Adds the persisted reload columns but leaves the previously reported legacy-request upgrade gap outstanding |
| schema.prisma | Adds nullable last-run state and a monotonic reload revision to LiteLLM_Config |
| tests/test_litellm/proxy/common_utils/test_periodic_reload_schedule.py | Covers parsing, status derivation, revision checks, cancellation, atomic increments, and field-scoped persistence |
| tests/test_litellm/proxy/test_proxy_server.py | Adds integration-level regression coverage for scheduler registration, retries, startup state, and fleet revision propagation |
Reviews (15): Last reviewed commit: "fix(ui): match the CI-generated user_rol..." | Re-trigger Greptile
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
A pod that had never reloaded kept its in-memory clock at None, and with no interval configured nothing ever set it, so every manual reload request was ignored by every pod except the one serving the click (Greptile P1 on the previous commit). Seed the per-pod timestamp at boot as the time its data was loaded and reload whenever a request or the interval is older than that, which also removes both None special cases from the due predicate. A schedule whose row has no last_run_at fires on the next tick so the first run does not wait a full interval.
|
On the db-push backfill note: acknowledged but intentionally left as-is. The backfill only matters for a manual reload clicked in the seconds before an upgrade and not yet consumed, the legacy boolean only ever reached one other pod anyway, and the remedy is clicking Reload Now again. Handling the legacy JSON flag in code would reintroduce the unclearable-flag loop this PR removes, so migrate-deploy paths get the carry-over and db-push paths degrade to exactly the old behavior's worst case |
|
Acknowledged — that's a reasonable trade-off. The backfill window is narrow (only a pending No further concerns on this point. |
|
@greptileai re review |
1 similar comment
|
@greptileai re review |
…he pod clock from the actual load time Revert the Anthropic beta headers reload path to its previous JSON-flag implementation so this PR only changes the price data reload; the beta headers path keeps working exactly as before and can migrate to the shared module in a follow-up. The unused columns on its config row are inert. Seed model_cost_map_loaded_at from the timestamp get_model_cost_map records at the actual import-time fetch instead of ProxyConfig construction time, closing the startup window where a manual reload request stamped between the fetch and the constructor compared as older than the pod's data and was skipped (Greptile P1 on the previous commit).
|
@greptileai re review |
…d tracking migration The backfill only carried over a manual reload clicked in the seconds before an upgrade, and every upgrade restarts the pods, which re-fetch the cost map at import and so already deliver what that request asked for. Removing it makes the migration schema-only, so prisma db push and prisma migrate deploy leave the database in the same state instead of diverging on a data statement that only one of them runs.
|
Resolved the db-push gap differently in 2887eb3: the legacy force_reload backfill is gone and the migration is now schema-only, so db push and migrate deploy leave the database in the same state and there is no data statement for one path to skip. Dropping it costs nothing in practice; that flag only meant "other pods, re-fetch your prices", and the rolling restart an upgrade performs already does that because every booting pod fetches the cost map at import |
Postgres stores these columns as TIMESTAMP(3) while Python stamps microseconds, so a pod comparing its in-memory clock against the persisted copy of the same instant read as newer and skipped the reload request it had just recorded. Truncate every stamp to milliseconds at the source, and floor the boot seed the same way, so the in-memory value and its persisted copy compare exactly.
…timestamps Comparing a request timestamp against each pod's data age made correctness depend on clock resolution: Postgres stores TIMESTAMP(3) while Python stamps microseconds, and two events inside the same millisecond are indistinguishable no matter how the comparison is written. Replace reload_requested_at with a reload_revision counter the manual reload endpoint increments atomically in the database. Each pod records the revision it last applied and reloads whenever the row's differs, so a request reaches every pod exactly once regardless of clock skew or precision, and concurrent requests publish distinct revisions instead of overwriting one another. A pod adopts the current revision on its first poll, since data it loaded at boot already satisfies any earlier request. Interval reloads still key off the pod's own data age, where hour scale comparisons make precision irrelevant.
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 2 · PR risk: 0/10 |
A pod adopted whatever revision it found on its first poll, so a manual reload published while the pod was starting was marked applied without ever being served and the pod kept the prices it fetched at import. Read the row once at startup instead, right after that fetch, and treat a missing row as revision 0
|
Fixed the first-poll race in 9839b95: the pod now seeds model_cost_map_applied_revision from the row once at startup, right after its boot-time cost map fetch, instead of adopting whatever it finds on the first tick. A missing row seeds 0, so the very first Reload Now still reaches a pod that booted before it, and a database hiccup at boot falls back to the previous adopt-on-first-poll behavior rather than blocking startup |
An earlier ruff format run reflowed the whole file from its 88-column formatting, adding ~1150 lines of churn unrelated to this PR. Replay only the real test changes onto the original formatting
…itellm_lit_4882_persist_reload_schedule # Conflicts: # litellm/litellm_core_utils/get_model_cost_map.py # litellm/proxy/proxy_server.py
Seeding the applied revision at startup left a window: a manual reload published after the import-time cost map fetch but before startup read the row was marked applied without ever being fetched, stranding that pod on stale prices when no interval was configured. A pod now starts unapplied and serves any outstanding request on its first poll, which costs one redundant fetch per boot and removes the window along with the seeding step
|
Closed the startup window in 19b6570 rather than narrowing it again. Seeding the applied revision was still a guess about ordering, so it is gone: a pod starts unapplied and serves any outstanding request on its first poll, then stays quiet. That costs one redundant cost map fetch per pod boot, and only after someone has used Reload Now at least once, which is cheaper than a pod holding stale prices indefinitely when no interval is configured. seed_model_cost_map_revision and its three tests are deleted and the per-pod field is a plain int again Also merged litellm_internal_staging in. Worth flagging one interaction: invalidate_config_param now publishes on the config sync channel, so the reload job stamping last_run_at would have made every pod broadcast a config change and put the whole fleet through add_deployment plus get_credentials on every price reload. The module now calls evict_config_param, which is what this path did before and what test_model_cost_map_reload_does_not_publish_config_change pins |
|
@greptileai re review |
param_value is written with safe_dumps, and a raw row read can return it decoded or as a string depending on the driver. Strict validation rejected the string, so the schedule read as disabled and an admin's configured reloads silently stopped. Mirrors the guard ConfigRepository.get_param already carries for the same column
|
@greptileai re review |
…itellm_lit_4882_persist_reload_schedule # Conflicts: # litellm/litellm_core_utils/get_model_cost_map.py # litellm/proxy/proxy_server.py # tests/test_litellm/litellm_core_utils/test_get_model_cost_map.py # tests/test_litellm/proxy/test_proxy_server.py
prisma rejects a null literal for a Json? column, so update_many writes an interval-less object instead. The fake config table now rejects the same input the database does, which is what the live run caught and the mock did not. Also records the run before adopting the revision, so a failed status write leaves the request unserved for the next poll rather than reporting a run that never landed.
…itellm_lit_4882_persist_reload_schedule # Conflicts: # litellm/proxy/proxy_server.py
Distributed QA: 3 pods, one shared Postgres, mutable cost mapQA only; no source changes. I ran this branch as a fleet to check that the original report in #24930 (a manual "Reload Now" refreshing a single pod, status showing nothing after a restart, and schedules never firing with Fleet setupOne Postgres 16 in docker on 5433, with the branch migration applied through Three pods on 4001, 4002 and 4003, same model_list:
- model_name: qa-sentinel-model
litellm_params:
model: openai/qa-sentinel-model
api_key: sk-fake-qa
api_base: http://127.0.0.1:9/v1
general_settings:
master_key: sk-1234
store_model_in_db: false
proxy_config_reload_interval_seconds: 10To make each pod's reload observable on its own, the cost map source is a local mutable file rather than the GitHub raw URL. A. Manual reload fan-outSentinel changed from One poll interval later every pod serves the new price, and the row carries revision 1 Per-pod logs show the fan-out landing exactly once on each pod that did not serve the request, and no pod repeating it over the following minute of ten-second polls B. Scheduled reload with
|
|
bugbot run |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Want reviews to match your repository better? Bugbot Learning can learn team-specific rules from PR activity. A team admin can enable Learning in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit d4df611. Configure here.
| return False | ||
| if schedule.last_run_at is None: | ||
| verbose_proxy_logger.info("%s reload triggered - schedule has never run", description) | ||
| return True |
There was a problem hiding this comment.
Failed stamp causes reload loop
Medium Severity
When an interval schedule has never persisted last_run_at, pod_reload_is_due keeps returning true on every tick without looking at pod_data_loaded_at. If _check_and_reload_model_cost_map fetches and swaps pricing but record_reload_run errors, the row stays unstamped and the pod refetches on each periodic_reload_job interval.
Additional Locations (1)
Reviewed by Cursor Bugbot for commit d4df611. Configure here.
There was a problem hiding this comment.
I think worst case of refetching is no big deal. Brief DB outages or failovers are not frequent anyways
| return datetime.now(timezone.utc) | ||
|
|
||
|
|
||
| def _parse_interval_hours(param_value: object) -> int | None: |
There was a problem hiding this comment.
Good to have explicit checking (input is object -> output is verified value) instead of force casting. Love it
…itellm_lit_4882_persist_reload_schedule # Conflicts: # litellm/proxy/proxy_server.py
9ea5cfc
into
litellm_internal_staging
…7.0) (#336) This PR contains the following updates: | Package | Update | Change | |---|---|---| | [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.96.2` → `v1.97.0` | --- ### Release Notes <details> <summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary> ### [`v1.97.0`](https://github.com/BerriAI/litellm/releases/tag/v1.97.0) [Compare Source](https://github.com/BerriAI/litellm/compare/v1.97.0...v1.97.0) ##### Verify Docker Image Signature All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0). **Verify using the pinned commit hash (recommended):** A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \ ghcr.io/berriai/litellm:v1.97.0 ``` **Verify using the release tag (convenience):** Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/v1.97.0/cosign.pub \ ghcr.io/berriai/litellm:v1.97.0 ``` Expected output: ``` The following checks were performed on each of these signatures: - The cosign claims were validated - The signatures were verified against the specified public key ``` *** ##### What's Changed - feat(proxy): resolve Cursor thinking/fast model-name suffixes on /cursor/chat/completions by [@​mateo-berri](https://github.com/mateo-berri) in [#​35554](https://github.com/BerriAI/litellm/pull/35554) - fix(team-callbacks): actually stop logging when disable\_logging is called by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35520](https://github.com/BerriAI/litellm/pull/35520) - refactor(lint): drop redundant !s f-string conversion flags and fix displaced import-group comments by [@​mateo-berri](https://github.com/mateo-berri) in [#​35546](https://github.com/BerriAI/litellm/pull/35546) - fix(proxy): backfill null user\_email on existing users during JWT auth by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34588](https://github.com/BerriAI/litellm/pull/34588) - feat(playground): add non-streaming response toggle by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35560](https://github.com/BerriAI/litellm/pull/35560) - feat(teams): apply default organization to new teams from default team settings by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35540](https://github.com/BerriAI/litellm/pull/35540) - fix(ui): block Playground page for viewer roles on direct URL access by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35676](https://github.com/BerriAI/litellm/pull/35676) - fix(caching): close evicted LLM clients so their connections are reclaimed by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35492](https://github.com/BerriAI/litellm/pull/35492) - chore(deps): update brace-expansion, postcss, and gitpython to current patch releases by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35692](https://github.com/BerriAI/litellm/pull/35692) - refactor(ui): rename the create MCP server component to PascalCase by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35686](https://github.com/BerriAI/litellm/pull/35686) - fix(openai): drop undefined Union from owns\_wrapped\_http\_client annotation by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35706](https://github.com/BerriAI/litellm/pull/35706) - fix(openai): drop the undefined Union from owns\_wrapped\_http\_client by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35704](https://github.com/BerriAI/litellm/pull/35704) - chore(ui): note Google's Agent Platform rename in vector store setup by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​28076](https://github.com/BerriAI/litellm/pull/28076) - fix(proxy): apply key/team router\_settings.model\_group\_alias by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35486](https://github.com/BerriAI/litellm/pull/35486) - feat(complexity\_router): default session affinity off and expose it in the UI by [@​tin-berri](https://github.com/tin-berri) in [#​35714](https://github.com/BerriAI/litellm/pull/35714) - fix(datadog): read team callback dd\_\* params from kwargs instead of blocked dynamic params ([#​35115](https://github.com/BerriAI/litellm/issues/35115) port) by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35687](https://github.com/BerriAI/litellm/pull/35687) - refactor(ui): extract the MCP create form's logic and field groups by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35694](https://github.com/BerriAI/litellm/pull/35694) - test(ui): tier the MCP create tests into unit and integration by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35697](https://github.com/BerriAI/litellm/pull/35697) - fix(proxy): redact credential headers from request logging copies by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35678](https://github.com/BerriAI/litellm/pull/35678) - feat(guardrails/rubrik): prompt moderation, response-text blocking, streaming buffer, failure logging by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35722](https://github.com/BerriAI/litellm/pull/35722) - fix(ui): render Responses API request and response in the logs drawer by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35718](https://github.com/BerriAI/litellm/pull/35718) - fix(ui): hide guardrail review buttons from non-admin users by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​27535](https://github.com/BerriAI/litellm/pull/27535) - feat(team): custom metadata validation hook for team create and update by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33353](https://github.com/BerriAI/litellm/pull/33353) - ci(circleci): install a pinned Rust toolchain on the Linux jobs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35519](https://github.com/BerriAI/litellm/pull/35519) - fix(bedrock): stop forwarding no-op toolSpec.strict to Converse by [@​tin-berri](https://github.com/tin-berri) in [#​35688](https://github.com/BerriAI/litellm/pull/35688) - fix(ui): reject an auto-router keyword rule left empty instead of dropping it by [@​tin-berri](https://github.com/tin-berri) in [#​35705](https://github.com/BerriAI/litellm/pull/35705) - fix(guardrails/rubrik): attribute blocked requests to the caller that made them by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35734](https://github.com/BerriAI/litellm/pull/35734) - fix(responses): forward client headers to the provider on /v1/responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34531](https://github.com/BerriAI/litellm/pull/34531) - feat(spend): add net auto-router savings to the cost-optimization dashboard by [@​tin-berri](https://github.com/tin-berri) in [#​35521](https://github.com/BerriAI/litellm/pull/35521) - chore(typing): clear basedpyright Any errors in budget reset, access groups, and cache settings by [@​mateo-berri](https://github.com/mateo-berri) in [#​35719](https://github.com/BerriAI/litellm/pull/35719) - fix(spend): read what a request cost from the record instead of pricing it again by [@​tin-berri](https://github.com/tin-berri) in [#​35736](https://github.com/BerriAI/litellm/pull/35736) - perf: install hiredis so redis-py parses replies with its C parser by [@​Classic298](https://github.com/Classic298) in [#​35709](https://github.com/BerriAI/litellm/pull/35709) - feat(ui): show auto-router savings on the cost-optimization dashboard by [@​tin-berri](https://github.com/tin-berri) in [#​35522](https://github.com/BerriAI/litellm/pull/35522) - perf: build log messages lazily so filtered-out log records cost nothing by [@​Classic298](https://github.com/Classic298) in [#​35703](https://github.com/BerriAI/litellm/pull/35703) - fix(proxy): retry model cost map fetch with Retry-After-aware backoff and keep current map on reload failure by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35739](https://github.com/BerriAI/litellm/pull/35739) - feat(otel): stamp service tier attributes on inference spans by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35679](https://github.com/BerriAI/litellm/pull/35679) - fix(proxy): log the model cost map reload failure lazily by [@​tin-berri](https://github.com/tin-berri) in [#​35750](https://github.com/BerriAI/litellm/pull/35750) - fix(groq): translate web\_search\_options to the browser\_search tool by [@​hMED22](https://github.com/hMED22) in [#​34971](https://github.com/BerriAI/litellm/pull/34971) - feat(ui): add admin-configurable user banner by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35729](https://github.com/BerriAI/litellm/pull/35729) - fix(e2e): make spend-counter redis connection env-driven for non-cluster deployments by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35732](https://github.com/BerriAI/litellm/pull/35732) - fix(proxy): make /cursor/chat/completions work with Cursor agent mode by [@​tin-berri](https://github.com/tin-berri) in [#​34029](https://github.com/BerriAI/litellm/pull/34029) - fix(proxy): propagate user\_email and bind api\_key on JWT auth attribution paths by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34331](https://github.com/BerriAI/litellm/pull/34331) - chore(build): move the Admin UI toolchain to Node 24 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35801](https://github.com/BerriAI/litellm/pull/35801) - test(e2e): vendor API strategy coverage across endpoints by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​34649](https://github.com/BerriAI/litellm/pull/34649) - chore(deps): upgrade cryptography to 50.0.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35803](https://github.com/BerriAI/litellm/pull/35803) - test(e2e): cover legacy text /completions endpoint by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​34431](https://github.com/BerriAI/litellm/pull/34431) - feat(gemini): add gemini-robotics-er-2-preview and gemini-robotics-er-1.6-preview by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35555](https://github.com/BerriAI/litellm/pull/35555) - test(e2e): move load/perf testing out of the main suite and drop the vllm passthrough test by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35820](https://github.com/BerriAI/litellm/pull/35820) - feat(lint): enforce Final on locals and freeze function parameters (LIT010, LIT011) by [@​mateo-berri](https://github.com/mateo-berri) in [#​35807](https://github.com/BerriAI/litellm/pull/35807) - chore: bump litellm-proxy-extras 0.4.81 -> 0.4.82, litellm 1.96.0 -> 1.97.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35810](https://github.com/BerriAI/litellm/pull/35810) - fix(bedrock): drop conflicting tool\_choice.type when toolConfig.toolChoice is set by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35738](https://github.com/BerriAI/litellm/pull/35738) - docs(CLAUDE.md): prefer commas over semicolons when replacing em dashes by [@​mateo-berri](https://github.com/mateo-berri) in [#​35825](https://github.com/BerriAI/litellm/pull/35825) - chore(lint): zero out basedpyright headroom for purely local rules by [@​mateo-berri](https://github.com/mateo-berri) in [#​35828](https://github.com/BerriAI/litellm/pull/35828) - test(e2e): retry provider-transient statuses at the transport with bounded backoff by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35824](https://github.com/BerriAI/litellm/pull/35824) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35836](https://github.com/BerriAI/litellm/pull/35836) - refactor(ui): route MCP session tokens through the shared storage helper by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35835](https://github.com/BerriAI/litellm/pull/35835) - docs(helm): replace the classic chart's 128Mi resource example with the documented 4Gi sizing by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35830](https://github.com/BerriAI/litellm/pull/35830) - fix(proxy): persist periodic reload schedule state so status survives restarts and fires without store\_model\_in\_db by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35165](https://github.com/BerriAI/litellm/pull/35165) - fix(router): eagerly fetch Vertex AI deferred stream to surface HTTP errors in \_acompletion fallback path by [@​deepanshululla](https://github.com/deepanshululla) in [#​34627](https://github.com/BerriAI/litellm/pull/34627) - fix(azure\_storage): honor AZURE\_STORAGE\_ENDPOINT\_SUFFIX for sovereign clouds by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35806](https://github.com/BerriAI/litellm/pull/35806) - fix(proxy): apply key\_alias/key\_hash filters to all /key/list visibility branches by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35840](https://github.com/BerriAI/litellm/pull/35840) - fix(proxy): enforce per-model budgets against resolved cursor model variants by [@​mateo-berri](https://github.com/mateo-berri) in [#​35834](https://github.com/BerriAI/litellm/pull/35834) - feat(ui): reorder Add Auto Router into name + template, with a collapsible detailed config by [@​tin-berri](https://github.com/tin-berri) in [#​35746](https://github.com/BerriAI/litellm/pull/35746) - test: repair three failing suites on litellm\_internal\_staging by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35845](https://github.com/BerriAI/litellm/pull/35845) - fix(guardrails): scan model output on the /openai/v1/responses alias by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35818](https://github.com/BerriAI/litellm/pull/35818) - ci: pin Node on the Playwright UI lanes so npm ci meets the engines floor by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35848](https://github.com/BerriAI/litellm/pull/35848) - fix(pricing): apply OpenAI's gpt-5.6 terra/luna cut to Azure cost map by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​35481](https://github.com/BerriAI/litellm/pull/35481) - feat(spend): add caller-scoped key/user/team/organization spend report endpoints by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35725](https://github.com/BerriAI/litellm/pull/35725) - revert: "fix(caching): close evicted LLM clients so their connections are reclaimed ([#​35492](https://github.com/BerriAI/litellm/issues/35492))" by [@​mateo-berri](https://github.com/mateo-berri) in [#​35856](https://github.com/BerriAI/litellm/pull/35856) - refactor(repositories): add prisma protocol seams and a spend-reset unit of work by [@​mateo-berri](https://github.com/mateo-berri) in [#​35748](https://github.com/BerriAI/litellm/pull/35748) - perf(streaming): assemble streamed tool-call arguments in linear time by [@​mateo-berri](https://github.com/mateo-berri) in [#​35826](https://github.com/BerriAI/litellm/pull/35826) - fix(s3\_v2): sign S3 object URLs with S3SigV4Auth so encoded paths verify by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35726](https://github.com/BerriAI/litellm/pull/35726) - test(e2e): self-seed the ui suite's password-login users in global setup by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35863](https://github.com/BerriAI/litellm/pull/35863) - fix(claude-code): create-only skill registration with a PUT update route (LIT-4110) by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​31752](https://github.com/BerriAI/litellm/pull/31752) - fix(proxy): fix zguard httpcode when block input by [@​jwang-gif](https://github.com/jwang-gif) in [#​31948](https://github.com/BerriAI/litellm/pull/31948) - fix(lint): pick the merge-aware base so in-progress merges are not blamed for base drift by [@​mateo-berri](https://github.com/mateo-berri) in [#​35868](https://github.com/BerriAI/litellm/pull/35868) - chore: bump litellm-proxy-extras 0.4.82 -> 0.4.83 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35877](https://github.com/BerriAI/litellm/pull/35877) - feat(ui): add Test Routing to the auto router create form by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35859](https://github.com/BerriAI/litellm/pull/35859) - fix(ui): derive auto-router preset tests from the bundled preset JSON by [@​tin-berri](https://github.com/tin-berri) in [#​35882](https://github.com/BerriAI/litellm/pull/35882) - revert: "test(e2e): vendor API strategy coverage across endpoints" ([#​34649](https://github.com/BerriAI/litellm/issues/34649)) by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35881](https://github.com/BerriAI/litellm/pull/35881) - chore(deps): bump grpc and golang.org/x modules in the terraform provider by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35844](https://github.com/BerriAI/litellm/pull/35844) - test(e2e): skip view-backed global spend probes pending LIT-5211 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35875](https://github.com/BerriAI/litellm/pull/35875) - fix(lint): move the basedpyright heap flag into the type check gate by [@​mateo-berri](https://github.com/mateo-berri) in [#​35869](https://github.com/BerriAI/litellm/pull/35869) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35876](https://github.com/BerriAI/litellm/pull/35876) - feat(ui): add role capability gating, migrate Tool Policies route by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35812](https://github.com/BerriAI/litellm/pull/35812) - refactor(ui): inject the fetch client's base url instead of reading it at import by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35802](https://github.com/BerriAI/litellm/pull/35802) - chore: remove unused .flake8 config and flake8 dev dependency by [@​mateo-berri](https://github.com/mateo-berri) in [#​35888](https://github.com/BerriAI/litellm/pull/35888) - chore: stop advising pre-commit and bootstrap by [@​mateo-berri](https://github.com/mateo-berri) in [#​35884](https://github.com/BerriAI/litellm/pull/35884) - fix(auth): name enable\_jwt\_auth when a JWT-shaped key is rejected by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35831](https://github.com/BerriAI/litellm/pull/35831) - feat(auto-router): make reminder marker pair configurable by [@​akapur99](https://github.com/akapur99) in [#​35874](https://github.com/BerriAI/litellm/pull/35874) - fix(UI): update anthropic model presets by [@​tin-berri](https://github.com/tin-berri) in [#​35896](https://github.com/BerriAI/litellm/pull/35896) - fix(bootstrap): switch to the dashboard node floor via nvm or fnm by [@​mateo-berri](https://github.com/mateo-berri) in [#​35895](https://github.com/BerriAI/litellm/pull/35895) - perf(pre-commit): run python, dashboard, and gen-api checks concurrently by [@​mateo-berri](https://github.com/mateo-berri) in [#​35903](https://github.com/BerriAI/litellm/pull/35903) - feat(spend): derive a default auto-router savings baseline from the hardest tier by [@​tin-berri](https://github.com/tin-berri) in [#​35907](https://github.com/BerriAI/litellm/pull/35907) - fix(http\_handler): self-heal handler clients closed after cache eviction by [@​mateo-berri](https://github.com/mateo-berri) in [#​35862](https://github.com/BerriAI/litellm/pull/35862) - fix(cost\_tracking): keep OpenAI prompt cache token details through usage reassembly by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34812](https://github.com/BerriAI/litellm/pull/34812) - fix(cost): bill gpt-5.6 prompt cache reads at the cache read rate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34957](https://github.com/BerriAI/litellm/pull/34957) - fix(batches): account for Responses API usage by [@​rimysore](https://github.com/rimysore) in [#​35367](https://github.com/BerriAI/litellm/pull/35367) - ci: retry Codecov uploads and stop failing jobs on OIDC token flakes by [@​mateo-berri](https://github.com/mateo-berri) in [#​35251](https://github.com/BerriAI/litellm/pull/35251) - feat(complexity\_router): let operators rename the four complexity tiers by [@​akapur99](https://github.com/akapur99) in [#​35893](https://github.com/BerriAI/litellm/pull/35893) - chore(lint): zero stale ruff and LIT headroom and strip inert type: ignore comments by [@​mateo-berri](https://github.com/mateo-berri) in [#​35928](https://github.com/BerriAI/litellm/pull/35928) - chore(lint): zero out seven more purely local basedpyright rules by [@​mateo-berri](https://github.com/mateo-berri) in [#​35927](https://github.com/BerriAI/litellm/pull/35927) - chore(ui): zero stale headroom on local dashboard eslint budgets by [@​mateo-berri](https://github.com/mateo-berri) in [#​35929](https://github.com/BerriAI/litellm/pull/35929) - fix(managed-files): skip rows without file objects by [@​rimysore](https://github.com/rimysore) in [#​35365](https://github.com/BerriAI/litellm/pull/35365) - fix(router): redact fallback tracebacks at the call site and cover the sync deferred stream by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35843](https://github.com/BerriAI/litellm/pull/35843) - fix(migrations): recover from an interrupted Prisma toolchain install by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35832](https://github.com/BerriAI/litellm/pull/35832) - fix(lint): bring basedpyright rule counts back under their budget limits by [@​mateo-berri](https://github.com/mateo-berri) in [#​35962](https://github.com/BerriAI/litellm/pull/35962) - chore(ui): don't zero out stale headroom except no-console by [@​mateo-berri](https://github.com/mateo-berri) in [#​35964](https://github.com/BerriAI/litellm/pull/35964) - fix(proxy): give proxy\_admin\_viewer read parity with proxy\_admin by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35851](https://github.com/BerriAI/litellm/pull/35851) - refactor(ui): address UI lint budget issues by refactoring UI by [@​tin-berri](https://github.com/tin-berri) in [#​35960](https://github.com/BerriAI/litellm/pull/35960) - fix(ci): make the env-key doc gate see get\_secret\_bool reads by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35833](https://github.com/BerriAI/litellm/pull/35833) - fix(caching): re-land evicted LLM client closing ([#​35492](https://github.com/BerriAI/litellm/issues/35492)) atop self-healing handlers by [@​mateo-berri](https://github.com/mateo-berri) in [#​35870](https://github.com/BerriAI/litellm/pull/35870) - fix(proxy): keep the connected DB client when a startup health check fails by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35837](https://github.com/BerriAI/litellm/pull/35837) - chore(lint): remove litellm/types from the ruff lint exclusion by [@​mateo-berri](https://github.com/mateo-berri) in [#​35926](https://github.com/BerriAI/litellm/pull/35926) - feat(sgr): make the gateway middleware the source of truth for successful requests by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35717](https://github.com/BerriAI/litellm/pull/35717) - feat(auto-router): let operators replace the LLM classifier's system prompt by [@​akapur99](https://github.com/akapur99) in [#​35855](https://github.com/BerriAI/litellm/pull/35855) - fix(docker): bake the pip image's prisma engines at a world-readable path by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35976](https://github.com/BerriAI/litellm/pull/35976) - fix(auth): return 403 from the OAuth2 enterprise gate by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35838](https://github.com/BerriAI/litellm/pull/35838) - fix(router): keep custom model\_info across a price data reload by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35491](https://github.com/BerriAI/litellm/pull/35491) - fix(proxy): resolve pass-through credentials live from router deployments by [@​mateo-berri](https://github.com/mateo-berri) in [#​35916](https://github.com/BerriAI/litellm/pull/35916) - fix(ci): fetch only head and merge-base in lint jobs instead of every branch by [@​mateo-berri](https://github.com/mateo-berri) in [#​35982](https://github.com/BerriAI/litellm/pull/35982) - fix(autorouter): match CJK keyword\_tier\_rules that regex word boundaries miss by [@​akapur99](https://github.com/akapur99) in [#​35984](https://github.com/BerriAI/litellm/pull/35984) - feat(spend): rebuild the auto-router benchmarks backend as a per-session rollup by [@​tin-berri](https://github.com/tin-berri) in [#​35910](https://github.com/BerriAI/litellm/pull/35910) - refactor(ui): replace hand-rolled query-param routing with nuqs by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35871](https://github.com/BerriAI/litellm/pull/35871) - fix(docker): bake the componentized prisma engines at /opt/prisma so any uid can start by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35989](https://github.com/BerriAI/litellm/pull/35989) - fix(migrations): keep the toolchain heal from raising on an unreadable nodeenv cache by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35986](https://github.com/BerriAI/litellm/pull/35986) - fix(bedrock): sign Bedrock managed-file S3 requests with S3SigV4Auth by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35983](https://github.com/BerriAI/litellm/pull/35983) - chore(typing): replace Any seams with real types across responses, proxy, and provider adapters by [@​mateo-berri](https://github.com/mateo-berri) in [#​35809](https://github.com/BerriAI/litellm/pull/35809) - fix(ai21): resolve the documented AI21\_API\_KEY instead of a misspelled name by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35985](https://github.com/BerriAI/litellm/pull/35985) - fix(docker): fail the image build when the generated prisma engine paths drift off /opt/prisma by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35979](https://github.com/BerriAI/litellm/pull/35979) - fix(jina\_ai): resolve the documented JINA\_API\_KEY as a fallback by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35992](https://github.com/BerriAI/litellm/pull/35992) - fix(proxy): only treat a recoverable database outage as grounds to serve without one by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35864](https://github.com/BerriAI/litellm/pull/35864) - fix(ci): make every remaining CI checkout shallow by [@​mateo-berri](https://github.com/mateo-berri) in [#​35997](https://github.com/BerriAI/litellm/pull/35997) - fix(auto-router): stop the embedding model's context window from failing long requests by [@​akapur99](https://github.com/akapur99) in [#​35956](https://github.com/BerriAI/litellm/pull/35956) - fix(ci): make the env-key doc gate see bare get\_secret and get\_secret\_str reads by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35996](https://github.com/BerriAI/litellm/pull/35996) - fix(logging): extend secret redaction to records litellm does not emit directly by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35977](https://github.com/BerriAI/litellm/pull/35977) - test(utils): pin the register\_model replay test to the recorded half by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35994](https://github.com/BerriAI/litellm/pull/35994) - fix(ci): run every helm test suite, not just the first one per file by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35993](https://github.com/BerriAI/litellm/pull/35993) - ci: fail the build when a test file or Dockerfile is invoked by no job by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35991](https://github.com/BerriAI/litellm/pull/35991) - fix(langfuse): stop a collected httpx handler from closing a shared client by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35981](https://github.com/BerriAI/litellm/pull/35981) - fix(bedrock): grant bedrock:CountTokens in OIDC session policy by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33145](https://github.com/BerriAI/litellm/pull/33145) - feat(pre-commit): save full lint output to a per-worktree log file by [@​mateo-berri](https://github.com/mateo-berri) in [#​36004](https://github.com/BerriAI/litellm/pull/36004) - feat(ui): match auto-router preset models against deployments' underlying model IDs by [@​tin-berri](https://github.com/tin-berri) in [#​35972](https://github.com/BerriAI/litellm/pull/35972) - fix(core\_helpers): map generic 'error' finish\_reason to 'stop' by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33972](https://github.com/BerriAI/litellm/pull/33972) - fix(proxy)!: apply request-parameter checks consistently across body, path and form inputs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36011](https://github.com/BerriAI/litellm/pull/36011) - fix: rebuild models\_by\_provider in add\_known\_models so cost map reloads reach wildcard expansion by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36010](https://github.com/BerriAI/litellm/pull/36010) - feat(complexity\_router): report LLM classifier cost per request via routing\_decision and x-litellm-classifier-cost header by [@​tin-berri](https://github.com/tin-berri) in [#​36015](https://github.com/BerriAI/litellm/pull/36015) - fix(model-prices): correct replicate model key typo by [@​AkashNaickar](https://github.com/AkashNaickar) in [#​34800](https://github.com/BerriAI/litellm/pull/34800) - fix(proxy): register managed batch output files on terminal retrieve by [@​Souravrajvi0](https://github.com/Souravrajvi0) in [#​34092](https://github.com/BerriAI/litellm/pull/34092) - perf(pre-commit): fetch basedpyright base counts from CI artifacts by [@​mateo-berri](https://github.com/mateo-berri) in [#​35970](https://github.com/BerriAI/litellm/pull/35970) - fix(ui): sync projects list page index to ?page= so back and reload keep the page by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36003](https://github.com/BerriAI/litellm/pull/36003) - fix(ui): link project page keys to their virtual key detail by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36002](https://github.com/BerriAI/litellm/pull/36002) - refactor(ui): drop unreferenced locals from dashboard route components by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35819](https://github.com/BerriAI/litellm/pull/35819) - fix(ui): opening a project now pushes ?project= so back and deep links work by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36001](https://github.com/BerriAI/litellm/pull/36001) - refactor(ui): drop unreferenced locals from shared dashboard components by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35821](https://github.com/BerriAI/litellm/pull/35821) - refactor(ui): drop unreferenced locals from tests and narrow destructures by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36025](https://github.com/BerriAI/litellm/pull/36025) - fix(guardrails): allow litellm\_content\_filter to run on post\_mcp\_call by [@​mateo-berri](https://github.com/mateo-berri) in [#​35980](https://github.com/BerriAI/litellm/pull/35980) - fix(guardrails): scan /v1/messages tool traffic by [@​mateo-berri](https://github.com/mateo-berri) in [#​35999](https://github.com/BerriAI/litellm/pull/35999) - refactor(ui): drop dead locals and unused React state across the dashboard by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36026](https://github.com/BerriAI/litellm/pull/36026) - feat(ui): add the auto-router usage tab to cost optimization by [@​tin-berri](https://github.com/tin-berri) in [#​35995](https://github.com/BerriAI/litellm/pull/35995) - fix(managed\_files): derive unified output file ids deterministically so concurrent registrations converge by [@​mateo-berri](https://github.com/mateo-berri) in [#​36019](https://github.com/BerriAI/litellm/pull/36019) - fix(proxy): send keepalive pings on anthropic messages SSE streams during upstream silence by [@​mateo-berri](https://github.com/mateo-berri) in [#​36024](https://github.com/BerriAI/litellm/pull/36024) - fix(managed\_files): return unified ids from unscoped file listing by [@​mateo-berri](https://github.com/mateo-berri) in [#​36031](https://github.com/BerriAI/litellm/pull/36031) - fix(arize\_phoenix): lowercase OTLP/gRPC auth metadata key by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34883](https://github.com/BerriAI/litellm/pull/34883) - fix(auto-router): accept every reminder marker pair a harness emits by [@​tin-berri](https://github.com/tin-berri) in [#​36029](https://github.com/BerriAI/litellm/pull/36029) - fix(pricing): sync flex/priority tier keys to dated OpenAI snapshot variants by [@​mateo-berri](https://github.com/mateo-berri) in [#​35923](https://github.com/BerriAI/litellm/pull/35923) - fix(cost): bill reasoning tokens at the service tier output rate by [@​mateo-berri](https://github.com/mateo-berri) in [#​35925](https://github.com/BerriAI/litellm/pull/35925) - fix(proxy): include today's UTC bucket when a daily activity range ends at the caller's current day by [@​tin-berri](https://github.com/tin-berri) in [#​36051](https://github.com/BerriAI/litellm/pull/36051) - fix: expired-miss share over all measured turns + cost-optimization tab labels by [@​tin-berri](https://github.com/tin-berri) in [#​36037](https://github.com/BerriAI/litellm/pull/36037) - fix(router): include Bedrock batch/S3 fields and model in deployment credentials by [@​mpcusack-altos](https://github.com/mpcusack-altos) in [#​24548](https://github.com/BerriAI/litellm/pull/24548) - fix(batch): track cost for managed batches with no attributable key/u… by [@​elinacse](https://github.com/elinacse) in [#​35468](https://github.com/BerriAI/litellm/pull/35468) - feat(guardrails): add scan\_only\_tool\_results to scope unified guardrails to tool results by [@​mateo-berri](https://github.com/mateo-berri) in [#​36014](https://github.com/BerriAI/litellm/pull/36014) - fix(cost): stop token-pricing the placeholder input on file content calls by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35140](https://github.com/BerriAI/litellm/pull/35140) - fix(proxy): fetch background responses through the router in CheckResponsesCost by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35137](https://github.com/BerriAI/litellm/pull/35137) - fix(proxy): yaml store\_prompts\_in\_spend\_logs should take precedence over DB cached value by [@​Praveena-617](https://github.com/Praveena-617) in [#​35769](https://github.com/BerriAI/litellm/pull/35769) - fix(lint): measure the basedpyright budget gate in a gate-owned venv by [@​mateo-berri](https://github.com/mateo-berri) in [#​36050](https://github.com/BerriAI/litellm/pull/36050) - docs: cap all GitHub comments at 15-25 words, curb semicolon splices by [@​mateo-berri](https://github.com/mateo-berri) in [#​36059](https://github.com/BerriAI/litellm/pull/36059) - chore(lint): name MappingProxyType in the mutable-collection fix messages by [@​mateo-berri](https://github.com/mateo-berri) in [#​36072](https://github.com/BerriAI/litellm/pull/36072) - test: roll back runtime model registrations between tests by [@​mateo-berri](https://github.com/mateo-berri) in [#​36039](https://github.com/BerriAI/litellm/pull/36039) - refactor(types): cut 653 implicit and explicit Any diagnostics across 11 modules by [@​mateo-berri](https://github.com/mateo-berri) in [#​36054](https://github.com/BerriAI/litellm/pull/36054) - fix(proxy): stop resolving the UI session sentinel team on /search\_tools/list by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36061](https://github.com/BerriAI/litellm/pull/36061) - fix(batches): persist managed file ids for cancelled/failed/expired batches by [@​mateo-berri](https://github.com/mateo-berri) in [#​36048](https://github.com/BerriAI/litellm/pull/36048) - fix(batches): register managed output files on batch cancel by [@​mateo-berri](https://github.com/mateo-berri) in [#​36034](https://github.com/BerriAI/litellm/pull/36034) - fix(proxy): allow non-admins to reach /user/daily/activity/aggregated by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36062](https://github.com/BerriAI/litellm/pull/36062) - fix(anthropic): coerce explicit additionalProperties to false in output\_format schema by [@​dkindlund](https://github.com/dkindlund) in [#​35811](https://github.com/BerriAI/litellm/pull/35811) - fix(batches): prevent managed file fallbacks by [@​rimysore](https://github.com/rimysore) in [#​35371](https://github.com/BerriAI/litellm/pull/35371) - chore: ignore the mechanical lint and typing sweeps in git blame by [@​mateo-berri](https://github.com/mateo-berri) in [#​36076](https://github.com/BerriAI/litellm/pull/36076) - fix(proxy): warn at startup when max\_budget is set but no database is connected by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36041](https://github.com/BerriAI/litellm/pull/36041) - fix(proxy): promote caller metadata trace fields into litellm\_metadata by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35866](https://github.com/BerriAI/litellm/pull/35866) - feat(terraform): sync provider 0.3.0 from the mirror and cut 0.4.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36098](https://github.com/BerriAI/litellm/pull/36098) - fix(guardrails): honor configured timeout in Zscaler AI Guard by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36110](https://github.com/BerriAI/litellm/pull/36110) - fix(logging): fall back to litellm\_metadata when metadata is empty by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36105](https://github.com/BerriAI/litellm/pull/36105) - fix(proxy): re-assert the authenticated identity on passthrough requests by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36121](https://github.com/BerriAI/litellm/pull/36121) - chore: bump litellm-enterprise 0.1.53 -> 0.1.54, litellm-proxy-extras 0.4.83 -> 0.4.84 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36139](https://github.com/BerriAI/litellm/pull/36139) - fix(ui): match auto-router preset models against wildcard-expanded model groups by [@​tin-berri](https://github.com/tin-berri) in [#​36111](https://github.com/BerriAI/litellm/pull/36111) - test(router): assert the auto-router max\_input\_chars kwarg by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36109](https://github.com/BerriAI/litellm/pull/36109) - fix(ui): allow clearing a key's budget reset from the Edit Key form by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36140](https://github.com/BerriAI/litellm/pull/36140) - fix(managed\_files): skip unparseable rows when listing managed files by [@​mateo-berri](https://github.com/mateo-berri) in [#​36021](https://github.com/BerriAI/litellm/pull/36021) - fix(a2a): stop writing per-caller headers onto the shared cached httpx client by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35978](https://github.com/BerriAI/litellm/pull/35978) - build(deps): bump h2 to 4.4.1 and js-yaml to 4.3.1 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36147](https://github.com/BerriAI/litellm/pull/36147) - chore: promote staging to main by [@​mateo-berri](https://github.com/mateo-berri) in [#​36057](https://github.com/BerriAI/litellm/pull/36057) - fix(azure\_sentinel): respect AZURE\_AUTHORITY\_HOST and derive the Azure Monitor audience per cloud by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36137](https://github.com/BerriAI/litellm/pull/36137) - fix(bedrock): pass SSE-KMS key through to the batch input-file S3 upload by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35148](https://github.com/BerriAI/litellm/pull/35148) - fix(anthropic adapter): stop indexing choices\[0] on choiceless streaming chunks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35314](https://github.com/BerriAI/litellm/pull/35314) - fix(bedrock): normalize /v1/completions and /v1/responses batch records by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35675](https://github.com/BerriAI/litellm/pull/35675) - fix(proxy): return the real status code when a credential update is rejected by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36166](https://github.com/BerriAI/litellm/pull/36166) - fix(proxy): improve Headroom /v1/compress HTTP 404 diagnostics by [@​aayush598](https://github.com/aayush598) in [#​35952](https://github.com/BerriAI/litellm/pull/35952) - fix(proxy): invalidate cached project object on project update and delete by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36028](https://github.com/BerriAI/litellm/pull/36028) - feat(proxy): add apply\_user\_budget\_to\_team\_keys opt-in by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36102](https://github.com/BerriAI/litellm/pull/36102) - fix(proxy): stop alerting on health probes that lose the planned engine-restart race by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36141](https://github.com/BerriAI/litellm/pull/36141) - test(docker): gate the componentized gateway and backend images on an arbitrary-uid offline boot by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36136](https://github.com/BerriAI/litellm/pull/36136) - fix(http): stop pooled clients persisting cookies on the aiohttp jar too by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36149](https://github.com/BerriAI/litellm/pull/36149) - fix(router): bound fallback-walk work and error-log volume by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36148](https://github.com/BerriAI/litellm/pull/36148) - ci: wire credential\_endpoints tests into the proxy endpoints job by [@​cursor](https://github.com/cursor)\[bot] in [#​36187](https://github.com/BerriAI/litellm/pull/36187) - docs(keys): document /key/info fields and clarify budget\_reset\_at is the next reset by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36127](https://github.com/BerriAI/litellm/pull/36127) - fix(azure\_sentinel): add AZURE\_SENTINEL\_AUTHORITY\_HOST as a Sentinel scoped override by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36165](https://github.com/BerriAI/litellm/pull/36165) - docs(pr-template): add a User Flow section with authoring instructions by [@​mateo-berri](https://github.com/mateo-berri) in [#​36162](https://github.com/BerriAI/litellm/pull/36162) - fix(proxy): derive config agent ids from agent\_name so grants survive secret rotation by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36020](https://github.com/BerriAI/litellm/pull/36020) - chore(ui): regenerate schema.d.ts for the /key/info docstring update by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36210](https://github.com/BerriAI/litellm/pull/36210) - build(deps): bump gitpython to 3.1.58 to clear osv-scan on staging by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36212](https://github.com/BerriAI/litellm/pull/36212) - fix(proxy): deny agent access when key and team grants resolve to nothing by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36221](https://github.com/BerriAI/litellm/pull/36221) - build(deps): defer the second pypdf advisory until the 6.15.0 bump by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36218](https://github.com/BerriAI/litellm/pull/36218) - fix(a2a): align agent list annotation and test with the tuple return type by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36217](https://github.com/BerriAI/litellm/pull/36217) - ci: always run the UI API types sync check so it can be required by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36213](https://github.com/BerriAI/litellm/pull/36213) - build(deps): bump nanoid to 3.3.17 in the dashboard lockfile by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36227](https://github.com/BerriAI/litellm/pull/36227) - feat(ui): show user email or alias in usage data export by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36232](https://github.com/BerriAI/litellm/pull/36232) - feat(auto-router): track turns per complexity tier (LIT-5302) by [@​tin-berri](https://github.com/tin-berri) in [#​36209](https://github.com/BerriAI/litellm/pull/36209) - fix(websearch): restore snippet text in native web\_search\_tool\_result blocks (LIT-5315) by [@​tin-berri](https://github.com/tin-berri) in [#​36228](https://github.com/BerriAI/litellm/pull/36228) - fix(proxy): resolve entity access groups in the model listing endpoints by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36230](https://github.com/BerriAI/litellm/pull/36230) - fix(ui): let access groups be a team's only model source, with hover provenance by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36234](https://github.com/BerriAI/litellm/pull/36234) - fix(managed\_files): return unified output file ids from GET /batches by [@​mateo-berri](https://github.com/mateo-berri) in [#​36049](https://github.com/BerriAI/litellm/pull/36049) - test(proxy): compare empty agent list to the tuple get\_agent\_list returns by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36225](https://github.com/BerriAI/litellm/pull/36225) - fix(otel): name the RPC system and upstream on MCP tool-call spans by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35857](https://github.com/BerriAI/litellm/pull/35857) - fix(guardrails): chunk oversized Bedrock ApplyGuardrail requests instead of failing by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36119](https://github.com/BerriAI/litellm/pull/36119) - test(e2e): settle control-plane writes across every replica, not just one by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36247](https://github.com/BerriAI/litellm/pull/36247) - fix(responses): forward allowed\_openai\_params through the chat completions bridge by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35885](https://github.com/BerriAI/litellm/pull/35885) - test(proxy): assert the copy \_add\_team\_member\_budget\_table returns by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36244](https://github.com/BerriAI/litellm/pull/36244) - chore(ui): regenerate dashboard api types for tier\_turns by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36243](https://github.com/BerriAI/litellm/pull/36243) - refactor(types): declare mirrored pricing fields on ModelInfo by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36215](https://github.com/BerriAI/litellm/pull/36215) - fix(lint): make strict-gate noqas survive base ruff and flag stale ones by [@​mateo-berri](https://github.com/mateo-berri) in [#​36257](https://github.com/BerriAI/litellm/pull/36257) - fix(vertex\_ai): surface real error/status on vertex batch create instead of IndexError 500 by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35141](https://github.com/BerriAI/litellm/pull/35141) - ci: give the remaining pull\_request workflows a concurrency group by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36252](https://github.com/BerriAI/litellm/pull/36252) - refactor(lint): graduate zero-violation strict rules and guard the budget ratchet by [@​mateo-berri](https://github.com/mateo-berri) in [#​36161](https://github.com/BerriAI/litellm/pull/36161) - fix(proxy): enforce require\_managed\_files on every route that accepts a raw provider id by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35551](https://github.com/BerriAI/litellm/pull/35551) - chore(typing): clear 1.4k basedpyright Any errors across 21 hotspot files by [@​mateo-berri](https://github.com/mateo-berri) in [#​36282](https://github.com/BerriAI/litellm/pull/36282) - test: roll back live router replay membership between tests by [@​mateo-berri](https://github.com/mateo-berri) in [#​36278](https://github.com/BerriAI/litellm/pull/36278) - chore(ci): sync main into internal staging by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36288](https://github.com/BerriAI/litellm/pull/36288) - build(lint): rename make pre-commit to make check with a working-tree fallback by [@​mateo-berri](https://github.com/mateo-berri) in [#​36277](https://github.com/BerriAI/litellm/pull/36277) - fix(ui): show team BYOK models in team fallback settings by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36241](https://github.com/BerriAI/litellm/pull/36241) - fix(otel): mark v2 server spans as failed for pre-call errors by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34546](https://github.com/BerriAI/litellm/pull/34546) - fix(websearch\_interception): bill intercepted searches to the calling key by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35708](https://github.com/BerriAI/litellm/pull/35708) - chore: remove pre-commit rule by [@​mateo-berri](https://github.com/mateo-berri) in [#​36295](https://github.com/BerriAI/litellm/pull/36295) - docs: clarify guideline priority ordering in CLAUDE.md by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36296](https://github.com/BerriAI/litellm/pull/36296) - feat(router): independent, default-on deployment affinity for the auto-router by [@​tin-berri](https://github.com/tin-berri) in [#​36146](https://github.com/BerriAI/litellm/pull/36146) - test: repair stale CircleCI contracts by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36293](https://github.com/BerriAI/litellm/pull/36293) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36286](https://github.com/BerriAI/litellm/pull/36286) - chore: rebuild Admin UI bundle for the 2026-08-08 release by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36297](https://github.com/BerriAI/litellm/pull/36297) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36304](https://github.com/BerriAI/litellm/pull/36304) ##### New Contributors - [@​rimysore](https://github.com/rimysore) made their first contribution in [#​35367](https://github.com/BerriAI/litellm/pull/35367) - [@​AkashNaickar](https://github.com/AkashNaickar) made their first contribution in [#​34800](https://github.com/BerriAI/litellm/pull/34800) - [@​Souravrajvi0](https://github.com/Souravrajvi0) made their first contribution in [#​34092](https://github.com/BerriAI/litellm/pull/34092) - [@​elinacse](https://github.com/elinacse) made their first contribution in [#​35468](https://github.com/BerriAI/litellm/pull/35468) - [@​aayush598](https://github.com/aayush598) made their first contribution in [#​35952](https://github.com/BerriAI/litellm/pull/35952) - [@​cursor](https://github.com/cursor)\[bot] made their first contribution in [#​36187](https://github.com/BerriAI/litellm/pull/36187) **Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.96.0...v1.97.0> ### [`v1.97.0`](https://github.com/BerriAI/litellm/releases/tag/v1.97.0) [Compare Source](https://github.com/BerriAI/litellm/compare/v1.96.2...v1.97.0) ##### Verify Docker Image Signature All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0). **Verify using the pinned commit hash (recommended):** A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \ ghcr.io/berriai/litellm:v1.97.0 ``` **Verify using the release tag (convenience):** Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/v1.97.0/cosign.pub \ ghcr.io/berriai/litellm:v1.97.0 ``` Expected output: ``` The following checks were performed on each of these signatures: - The cosign claims were validated - The signatures were verified against the specified public key ``` *** ##### What's Changed - feat(proxy): resolve Cursor thinking/fast model-name suffixes on /cursor/chat/completions by [@​mateo-berri](https://github.com/mateo-berri) in [#​35554](https://github.com/BerriAI/litellm/pull/35554) - fix(team-callbacks): actually stop logging when disable\_logging is called by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35520](https://github.com/BerriAI/litellm/pull/35520) - refactor(lint): drop redundant !s f-string conversion flags and fix displaced import-group comments by [@​mateo-berri](https://github.com/mateo-berri) in [#​35546](https://github.com/BerriAI/litellm/pull/35546) - fix(proxy): backfill null user\_email on existing users during JWT auth by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34588](https://github.com/BerriAI/litellm/pull/34588) - feat(playground): add non-streaming response toggle by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35560](https://github.com/BerriAI/litellm/pull/35560) - feat(teams): apply default organization to new teams from default team settings by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35540](https://github.com/BerriAI/litellm/pull/35540) - fix(ui): block Playground page for viewer roles on direct URL access by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35676](https://github.com/BerriAI/litellm/pull/35676) - fix(caching): close evicted LLM clients so their connections are reclaimed by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35492](https://github.com/BerriAI/litellm/pull/35492) - chore(deps): update brace-expansion, postcss, and gitpython to current patch releases by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35692](https://github.com/BerriAI/litellm/pull/35692) - refactor(ui): rename the create MCP server component to PascalCase by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35686](https://github.com/BerriAI/litellm/pull/35686) - fix(openai): drop undefined Union from owns\_wrapped\_http\_client annotation by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35706](https://github.com/BerriAI/litellm/pull/35706) - fix(openai): drop the undefined Union from owns\_wrapped\_http\_client by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35704](https://github.com/BerriAI/litellm/pull/35704) - chore(ui): note Google's Agent Platform rename in vector store setup by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​28076](https://github.com/BerriAI/litellm/pull/28076) - fix(proxy): apply key/team router\_settings.model\_group\_alias by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35486](https://github.com/BerriAI/litellm/pull/35486) - feat(complexity\_router): default session affinity off and expose it in the UI by [@​tin-berri](https://github.com/tin-berri) in [#​35714](https://github.com/BerriAI/litellm/pull/35714) - fix(datadog): read team callback dd\_\* params from kwargs instead of blocked dynamic params ([#​35115](https://github.com/BerriAI/litellm/issues/35115) port) by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35687](https://github.com/BerriAI/litellm/pull/35687) - refactor(ui): extract the MCP create form's logic and field groups by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35694](https://github.com/BerriAI/litellm/pull/35694) - test(ui): tier the MCP create tests into unit and integration by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35697](https://github.com/BerriAI/litellm/pull/35697) - fix(proxy): redact credential headers from request logging copies by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35678](https://github.com/BerriAI/litellm/pull/35678) - feat(guardrails/rubrik): prompt moderation, response-text blocking, streaming buffer, failure logging by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35722](https://github.com/BerriAI/litellm/pull/35722) - fix(ui): render Responses API request and response in the logs drawer by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35718](https://github.com/BerriAI/litellm/pull/35718) - fix(ui): hide guardrail review buttons from non-admin users by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​27535](https://github.com/BerriAI/litellm/pull/27535) - feat(team): custom metadata validation hook for team create and update by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33353](https://github.com/BerriAI/litellm/pull/33353) - ci(circleci): install a pinned Rust toolchain on the Linux jobs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35519](https://github.com/BerriAI/litellm/pull/35519) - fix(bedrock): stop forwarding no-op toolSpec.strict to Converse by [@​tin-berri](https://github.com/tin-berri) in [#​35688](https://github.com/BerriAI/litellm/pull/35688) - fix(ui): reject an auto-router keyword rule left empty instead of dropping it by [@​tin-berri](https://github.com/tin-berri) in [#​35705](https://github.com/BerriAI/litellm/pull/35705) - fix(guardrails/rubrik): attribute blocked requests to the caller that made them by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35734](https://github.com/BerriAI/litellm/pull/35734) - fix(responses): forward client headers to the provider on /v1/responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34531](https://github.com/BerriAI/litellm/pull/34531) - feat(spend): add net auto-router savings to the cost-optimization dashboard by [@​tin-berri](https://github.com/tin-berri) in [#​35521](https://github.com/BerriAI/litellm/pull/35521) - chore(typing): clear basedpyright Any errors in budget reset, access groups, and cache settings by [@​mateo-berri](https://github.com/mateo-berri) in [#​35719](https://github.com/BerriAI/litellm/pull/35719) - fix(spend): read what a request cost from the record instead of pricing it again by [@​tin-berri](https://github.com/tin-berri) in [#​35736](https://github.com/BerriAI/litellm/pull/35736) - perf: install hiredis so redis-py parses replies with its C parser by [@​Classic298](https://github.com/Classic298) in [#​35709](https://github.com/BerriAI/litellm/pull/35709) - feat(ui): show auto-router savings on the cost-optimization dashboard by [@​tin-berri](https://github.com/tin-berri) in [#​35522](https://github.com/BerriAI/litellm/pull/35522) - perf: build log messages lazily so filtered-out log records cost nothing by [@​Classic298](https://github.com/Classic298) in [#​35703](https://github.com/BerriAI/litellm/pull/35703) - fix(proxy): retry model cost map fetch with Retry-After-aware backoff and keep current map on reload failure by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35739](https://github.com/BerriAI/litellm/pull/35739) - feat(otel): stamp service tier attributes on inference spans by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35679](https://github.com/BerriAI/litellm/pull/35679) - fix(proxy): log the model cost map reload failure lazily by [@​tin-berri](https://github.com/tin-berri) in [#​35750](https://github.com/BerriAI/litellm/pull/35750) - fix(groq): translate web\_search\_options to the browser\_search tool by [@​hMED22](https://github.com/hMED22) in [#​34971](https://github.com/BerriAI/litellm/pull/34971) - feat(ui): add admin-configurable user banner by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35729](https://github.com/BerriAI/litellm/pull/35729) - fix(e2e): make spend-counter redis connection env-driven for non-cluster deployments by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35732](https://github.com/BerriAI/litellm/pull/35732) - fix(proxy): make /cursor/chat/completions work with Cursor agent mode by [@​tin-berri](https://github.com/tin-berri) in [#​34029](https://github.com/BerriAI/litellm/pull/34029) - fix(proxy): propagate user\_email and bind api\_key on JWT auth attribution paths by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34331](https://github.com/BerriAI/litellm/pull/34331) - chore(build): move the Admin UI toolchain to Node 24 by [@​yuneng-berri](https://g…


TLDR
Problem this solves:
How it solves it:
Relevant issues
Fixes #24930
Linear ticket
Resolves LIT-4882
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Reproduction before the fix is documented in #24930 and matches the base branch (b930e2f): schedule a reload from Models + Endpoints -> Price Data Reload, restart the proxy, and GET /schedule/model_cost_map_reload/status returns
"last_run": null, "next_run": nullso the card shows "Never"; withstore_model_in_db: falsethe reload additionally never executes, because the check only ran from the store_model_in_db-gated add_deployment jobEverything below ran at commit d66cfab against a live proxy: fresh Postgres 16,
general_settings.store_model_in_db: false, a config-file model list, migrations applied withprisma migrate deployThe migration adds both columns, and scheduling writes only the admin-owned param_value
One scheduler tick later, with store_model_in_db false, the new periodic_reload_job fires and stamps the run
Two manual reloads each publish a distinct revision, so a second click is never swallowed by the first
Cancelling clears the interval and keeps the row, so the counter carries on. With a delete the re-created row restarted at 1, which is a number pods had already applied, and their next Reload Now would have been skipped everywhere but the pod that served it
Re-schedule, kill the proxy, start it again and query status immediately; this is the exact customer flow that used to report "Never"
Type
🐛 Bug Fix
Changes
The interval an admin configures and the runtime state the reload job produces used to share one JSON blob in LiteLLM_Config, with the last-run time living only in a module-level global. This PR splits them by owner. The schedule endpoints own
param_value({"interval_hours": N}); the reload job and the manual reload endpoint own two new columns,last_run_atandreload_revision, written with field-scoped updates so no writer can clobber another's data. The migration is additive and schema-only, with IF NOT EXISTS guards, soprisma db pushandprisma migrate deployleave the database in the same stateThe check itself moves off the store_model_in_db-gated add_deployment job onto its own
periodic_reload_job, registered whenever a database is configured, so config-file deployments execute their schedules. The status endpoint now deriveslast_runandnext_runpurely from the row, which is what makes the UI survive restarts and multi-pod routing. The logic lives inlitellm/proxy/common_utils/periodic_reload_schedule.py. Scope note: this PR only changes the model cost map path; the Anthropic beta headers reload keeps its previous implementation and can adopt the shared module in a follow-upThe old force_reload boolean is replaced by
reload_revision, a counter the manual reload endpoint increments atomically in the database. Each pod records the revision it last applied and reloads whenever the row's differs, which identifies a request rather than ordering one: there is no clock comparison, so neither skew between pods nor the gap between Postgres TIMESTAMP(3) and Python's microsecond stamps can make a pod miss a reload or serve one twice. Nothing ever clears the counter, so a request reaches every pod exactly once, where the boolean was consumed by the first poller and starved the rest. Concurrent requests publish distinct revisions instead of overwriting each other. A booting pod starts unapplied rather than adopting whatever revision it finds, because it cannot prove that an outstanding request predates the prices it fetched at import; it serves that request on its first poll and is quiet from then on. The cost is one redundant fetch per pod boot, and only once someone has used Reload Now at least once, which is the cheaper side of the trade against a pod sitting on stale prices with no interval configured to rescue it. Interval reloads still key off the pod's own data age, where hour-scale comparisons make precision irrelevantCancelling a schedule clears
param_valueinstead of deleting the row, because the row also carries the counter. Deleting it restarted the count, and a reissued revision matches what pods already applied, so the next Reload Now would reach only the pod that served itA schedule whose row has never recorded a run fires on the next tick rather than one interval later. The job stamps
last_run_atwith update_many so a schedule cancelled mid-poll is not resurrected. The fetch itself goes throughrefetch_model_cost_map, so a failed reload keeps the pod's current pricing and leaves the revision unapplied, which is what makes the next poll retry it rather than record it as servedOne visible output change worth noting: timestamps from these endpoints now carry an explicit +00:00 offset, which fixes the dashboard rendering them shifted by the viewer's UTC offset (the previous naive strings were parsed as local time). Response shapes are unchanged. A follow-up PR ports the dashboard card to react-query and stops it rendering fetch errors as "not scheduled"
Final Attestation
Note
Medium Risk
Touches pricing reload paths and shared LiteLLM_Config schema; behavior changes for multi-pod fleets and deployments without store_model_in_db, though failures keep current pricing and retry.
Overview
Fixes proxy model cost map reload scheduling so last run and status survive restarts, interval jobs run without
store_model_in_db, and manual reload reaches every pod.Database:
LiteLLM_Configgainslast_run_atandreload_revision. Admins still ownparam_value(interval_hours); the reload job owns the new columns via scoped writes. Cancelling a schedule nulls the interval in JSON instead of deleting the row so the revision counter is never reset.Coordination: Replaces the old
force_reloadflag and in-memorylast_model_cost_map_reloadwith a monotonicreload_revisionthat manual reload bumps atomically. Each pod tracksmodel_cost_map_applied_revisionand reloads when it lags the row; interval timing uses per-podmodel_cost_map_loaded_at(fromget_model_cost_map_loaded_at()). Failed fetches leave the revision unapplied for retry.Runtime: New
periodic_reload_schedulemodule and a dedicatedperiodic_reload_job(not gated onstore_model_in_db). Status API reads the DB row; reload logic is centralized in_swap_in_model_cost_map.Reviewed by Cursor Bugbot for commit d4df611. Bugbot is set up for automated code reviews on this repo. Configure here.