Skip to content

fix(proxy): forward resolved provider and deployment pricing in /cost/estimate - #35880

Merged
mateo-berri merged 3 commits into
litellm_internal_stagingfrom
devin_ai_fix_cost_estimate_onprem_provider_35210
Aug 12, 2026
Merged

mateo-berri merged 3 commits into
litellm_internal_stagingfrom
devin_ai_fix_cost_estimate_onprem_provider_35210

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • /cost/estimate 500s for on-prem models like nvidia/zai-org/glm-5.2
  • Configured per-token pricing on such deployments was ignored

How it solves it:

  • Forward the router-resolved provider into the cost calculation
  • Forward the deployment's input/output_cost_per_token as custom pricing, read from litellm_params or model_info (litellm_params wins, matching the router's cost-map registration precedence)
  • Return that configured per-token pricing in the estimate response

User Flow

Before: an admin sizing costs for a self-hosted deployment gets a 500 instead of an estimate

  1. The proxy config lists model_name: nvidia/zai-org/glm-5.2 pointing at a self-hosted OpenAI-compatible server, with input_cost_per_token/output_cost_per_token optionally set
  2. They send POST http://litellm-domain/cost/estimate with {"model": "nvidia/zai-org/glm-5.2", "input_tokens": 1000, "output_tokens": 500}
  3. The response is a 500: "Could not calculate cost for model 'nvidia/zai-org/glm-5.2' (resolved to 'zai-org/GLM-5.2'): litellm.BadRequestError: LLM Provider NOT provided"

After: the same request returns a real estimate that honors the deployment's configured pricing

  1. The proxy config lists the same nvidia/zai-org/glm-5.2 deployment
  2. They send the same POST http://litellm-domain/cost/estimate request
  3. The response is a 200 naming the resolved provider, and when the deployment configures per-token pricing (under litellm_params or model_info) the estimate reflects it (cost_per_request, daily_cost, and the per-token rates)

Relevant issues

Linear ticket

Resolves LIT-5210

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • My PR passes all CI/CD checks (e.g., lint, format, unit tests)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

/cost/estimate is a pure pricing computation; it never dispatches a request to a provider, so there is no real LLM call or spend involved. The realistic end-user action is the curl below, run against a live proxy configured exactly as the ticket describes. Both legs booted a live proxy from the same config: the before leg at the merge base 7e80e094c4, the after leg at this PR's head 0fdbe03c50

model_list:
  - model_name: nvidia/zai-org/glm-5.2
    litellm_params:
      model: zai-org/GLM-5.2
      custom_llm_provider: openai
      api_base: http://localhost:9999/v1
      api_key: sk-fake
  - model_name: nvidia-priced/zai-org/glm-5.2
    litellm_params:
      model: zai-org/GLM-5.2
      custom_llm_provider: openai
      api_base: http://localhost:9999/v1
      api_key: sk-fake
      input_cost_per_token: 0.000001
      output_cost_per_token: 0.000002
  - model_name: nvidia-mi-priced/zai-org/glm-5.2
    litellm_params:
      model: zai-org/GLM-5.2
      custom_llm_provider: openai
      api_base: http://localhost:9999/v1
      api_key: sk-fake
    model_info:
      input_cost_per_token: 0.000003
      output_cost_per_token: 0.000004
  - model_name: gpt-4
    litellm_params:
      model: openai/gpt-4
      api_key: os.environ/OPENAI_API_KEY
general_settings:
  master_key: sk-1234

Before (merge base 7e80e094c4): every on-prem deployment 500s, priced or not, wherever the pricing lives

$ curl -s -X POST http://localhost:43117/cost/estimate -H "Authorization: Bearer sk-1234" \
    -H "Content-Type: application/json" \
    -d '{"model":"nvidia/zai-org/glm-5.2","input_tokens":1000,"output_tokens":500}'
{"detail":{"error":"Could not calculate cost for model 'nvidia/zai-org/glm-5.2' (resolved to 'zai-org/GLM-5.2'): litellm.BadRequestError: LLM Provider NOT provided. Pass in the LLM provider you are trying to call. You passed model=zai-org/GLM-5.2\n Pass model as E.g. For 'Huggingface' inference endpoints pass in `completion(model='huggingface/starcoder',..)` Learn more: https://docs.litellm.ai/docs/providers"}}

$ curl -s -X POST http://localhost:43117/cost/estimate -H "Authorization: Bearer sk-1234" \
    -H "Content-Type: application/json" \
    -d '{"model":"nvidia-priced/zai-org/glm-5.2","input_tokens":1000,"output_tokens":500,"num_requests_per_day":100}'
{"detail":{"error":"Could not calculate cost for model 'nvidia-priced/zai-org/glm-5.2' (resolved to 'zai-org/GLM-5.2'): litellm.BadRequestError: LLM Provider NOT provided. ..."}}

$ curl -s -X POST http://localhost:43117/cost/estimate -H "Authorization: Bearer sk-1234" \
    -H "Content-Type: application/json" \
    -d '{"model":"nvidia-mi-priced/zai-org/glm-5.2","input_tokens":1000,"output_tokens":500}'
{"detail":{"error":"Could not calculate cost for model 'nvidia-mi-priced/zai-org/glm-5.2' (resolved to 'zai-org/GLM-5.2'): litellm.BadRequestError: LLM Provider NOT provided. ..."}}

After (PR head 0fdbe03c50): no configured pricing returns 200 with the resolved provider

$ curl -s -X POST http://localhost:44923/cost/estimate -H "Authorization: Bearer sk-1234" \
    -H "Content-Type: application/json" \
    -d '{"model":"nvidia/zai-org/glm-5.2","input_tokens":1000,"output_tokens":500}'
{"model":"nvidia/zai-org/glm-5.2","input_tokens":1000,"output_tokens":500,"num_requests_per_day":null,"num_requests_per_month":null,"cost_per_request":0.0,"input_cost_per_request":0.0,"output_cost_per_request":0.0,"margin_cost_per_request":0.0,"daily_cost":null,"daily_input_cost":null,"daily_output_cost":null,"daily_margin_cost":null,"monthly_cost":null,"monthly_input_cost":null,"monthly_output_cost":null,"monthly_margin_cost":null,"input_cost_per_token":null,"output_cost_per_token":null,"provider":"openai"}

After (PR head 0fdbe03c50): litellm_params per-token pricing is honored

$ curl -s -X POST http://localhost:44923/cost/estimate -H "Authorization: Bearer sk-1234" \
    -H "Content-Type: application/json" \
    -d '{"model":"nvidia-priced/zai-org/glm-5.2","input_tokens":1000,"output_tokens":500,"num_requests_per_day":100}'
{"model":"nvidia-priced/zai-org/glm-5.2","input_tokens":1000,"output_tokens":500,"num_requests_per_day":100,"num_requests_per_month":null,"cost_per_request":0.002,"input_cost_per_request":0.001,"output_cost_per_request":0.001,"margin_cost_per_request":0.0,"daily_cost":0.2,"daily_input_cost":0.1,"daily_output_cost":0.1,"daily_margin_cost":0.0,"monthly_cost":null,"monthly_input_cost":null,"monthly_output_cost":null,"monthly_margin_cost":null,"input_cost_per_token":1e-6,"output_cost_per_token":2e-6,"provider":"openai"}

After (PR head 0fdbe03c50): model_info per-token pricing (how DB / Admin UI added deployments store it) is honored too

$ curl -s -X POST http://localhost:44923/cost/estimate -H "Authorization: Bearer sk-1234" \
    -H "Content-Type: application/json" \
    -d '{"model":"nvidia-mi-priced/zai-org/glm-5.2","input_tokens":1000,"output_tokens":500,"num_requests_per_day":100}'
{"model":"nvidia-mi-priced/zai-org/glm-5.2","input_tokens":1000,"output_tokens":500,"num_requests_per_day":100,"num_requests_per_month":null,"cost_per_request":0.005,"input_cost_per_request":0.003,"output_cost_per_request":0.002,"margin_cost_per_request":0.0,"daily_cost":0.5,"daily_input_cost":0.3,"daily_output_cost":0.2,"daily_margin_cost":0.0,"monthly_cost":null,"monthly_input_cost":null,"monthly_output_cost":null,"monthly_margin_cost":null,"input_cost_per_token":3e-6,"output_cost_per_token":4e-6,"provider":"openai"}

Public models are unaffected; gpt-4 estimates its mapped cost identically on both legs

$ curl -s -X POST http://localhost:44923/cost/estimate -H "Authorization: Bearer sk-1234" \
    -H "Content-Type: application/json" \
    -d '{"model":"gpt-4","input_tokens":1000,"output_tokens":500}'
{"model":"gpt-4","input_tokens":1000,"output_tokens":500,...,"cost_per_request":0.060000000000000005,"input_cost_per_token":0.00003,"output_cost_per_token":0.00006,"provider":"openai"}

QA observations

  • Unpriced on-prem models now estimate 0.0 silently, by design
  • This PR causes that; configured pricing is the remedy

Type

🐛 Bug Fix

Caveats (if any)

QA runbook

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@CLAassistant

CLAassistant commented Aug 5, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ mateo-berri
❌ devin-ai-integration[bot]
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR forwards router-resolved provider metadata and deployment-specific token pricing into /cost/estimate.

  • Introduces a structured result for resolved model, provider, and custom pricing.
  • Applies deployment pricing from litellm_params or model_info, with litellm_params taking precedence.
  • Adds regression coverage for priced and unpriced self-hosted deployments.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Important Files Changed

Filename Overview
litellm/proxy/management_endpoints/cost_tracking_settings.py Resolves deployment provider and custom token pricing before invoking the existing cost calculator.
tests/test_litellm/proxy/management_endpoints/test_cost_tracking_settings.py Adds coverage for self-hosted model estimation and deployment-pricing precedence.

Reviews (4): Last reviewed commit: "fix(proxy): honor model_info custom pric..." | Re-trigger Greptile

Comment on lines +42 to +45
"""
Pull per-token pricing configured on a deployment so on-prem / self-hosted
models (absent from the public cost map) still estimate a real cost.
"""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 New explanatory comments violate guidance

The new helper and regression tests add explanatory docstrings and comments despite the repository instruction prohibiting new comments unless explicitly requested, adding convention cleanup across both changed files

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@codecov

codecov Bot commented Aug 5, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing devin_ai_fix_cost_estimate_onprem_provider_35210 (0fdbe03) with litellm_internal_staging (7e80e09)

Open in CodSpeed

@devin-ai-integration
devin-ai-integration Bot force-pushed the devin_ai_fix_cost_estimate_onprem_provider_35210 branch from 05735f6 to dcc8480 Compare August 7, 2026 20:39
…/estimate

estimate_cost resolved on-prem aliases (e.g. nvidia/zai-org/glm-5.2) to their
underlying model and custom_llm_provider via the router, then called
completion_cost without either, so provider inference ran on the bare model and
raised "LLM Provider NOT provided"; deployment-configured per-token pricing was
dropped too, so priced on-prem deployments estimated 0. The resolver now returns
a frozen ResolvedCostModel(model, provider, custom_cost_per_token) and
estimate_cost forwards both into completion_cost and surfaces the configured
per-token pricing in the response, deriving that pricing as single Final values.

Resolves LIT-5210
@devin-ai-integration
devin-ai-integration Bot force-pushed the devin_ai_fix_cost_estimate_onprem_provider_35210 branch from dcc8480 to c19ab70 Compare August 7, 2026 20:51
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

Comment thread litellm/proxy/management_endpoints/cost_tracking_settings.py Outdated
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 0fdbe03. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 9bfe593 into litellm_internal_staging Aug 12, 2026
85 checks passed
@mateo-berri
mateo-berri deleted the devin_ai_fix_cost_estimate_onprem_provider_35210 branch August 12, 2026 07:08
doonga pushed a commit to greyrock-labs/home-ops that referenced this pull request Aug 24, 2026
…8.0) (#393)

This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.97.0` → `v1.98.0` |

---

### Release Notes

<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>

### [`v1.98.0`](https://github.com/BerriAI/litellm/releases/tag/v1.98.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.98.0...v1.98.0)

##### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.98.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.98.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.98.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

##### What's Changed

- fix(bedrock): drop toolSpec.strict for Claude Sonnet 5 on Converse by [@&#8203;kr0k](https://github.com/kr0k) in [#&#8203;33196](https://github.com/BerriAI/litellm/pull/33196)
- fix(batches): attribute Vertex passthrough batch cost to key/team/tags by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;34456](https://github.com/BerriAI/litellm/pull/34456)
- docs: rewrite the CLAUDE.md comment rule with explicit exceptions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36301](https://github.com/BerriAI/litellm/pull/36301)
- fix(proxy): scope file list pagination cursors to the caller by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36093](https://github.com/BerriAI/litellm/pull/36093)
- fix(proxy): skip prisma-dependent hooks when no database is attached by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36273](https://github.com/BerriAI/litellm/pull/36273)
- fix(proxy): report has\_more false on caller-scoped file list pages by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36326](https://github.com/BerriAI/litellm/pull/36326)
- fix(proxy): restore management\_v1 query-param validation under fastapi>=0.140.7 by [@&#8203;HuanQian571](https://github.com/HuanQian571) in [#&#8203;35773](https://github.com/BerriAI/litellm/pull/35773)
- fix(proxy): stop /{provider}/v1/files from capturing /openai\_passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36092](https://github.com/BerriAI/litellm/pull/36092)
- chore(typing): remove 914 basedpyright Any errors across 16 hotspot files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36386](https://github.com/BerriAI/litellm/pull/36386)
- fix(router): keep batch fallbacks inside the model group that owns the file by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36181](https://github.com/BerriAI/litellm/pull/36181)
- feat(ptu): configure provisioned-throughput flat cost on a model deployment by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35341](https://github.com/BerriAI/litellm/pull/35341)
- docs: clarify the CLAUDE.md comment exceptions are any-of by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36421](https://github.com/BerriAI/litellm/pull/36421)
- docs: replace the Changes PR template section with Caveats by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36423](https://github.com/BerriAI/litellm/pull/36423)
- fix(bedrock): enable native structured output for GLM 5 and DeepSeek V3.2 by [@&#8203;alexshtf](https://github.com/alexshtf) in [#&#8203;35669](https://github.com/BerriAI/litellm/pull/35669)
- feat(ptu): daily rollup writes per-model PTU flat cost by active hour by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35343](https://github.com/BerriAI/litellm/pull/35343)
- feat(logging): add opt-in session\_id and trace\_id correlation to JSON log records via contextvars by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;34418](https://github.com/BerriAI/litellm/pull/34418)
- feat(ptu): surface PTU flat cost on the daily activity read path by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35391](https://github.com/BerriAI/litellm/pull/35391)
- feat(router): add per-deployment allowed\_fails\_policy and cooldown\_time override support by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;34416](https://github.com/BerriAI/litellm/pull/34416)
- feat(ptu): add PTU inputs to the model form and flat cost to the Usage page by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35393](https://github.com/BerriAI/litellm/pull/35393)
- fix(cost): price dict-shaped image input token details at the image rate by [@&#8203;vairodp](https://github.com/vairodp) in [#&#8203;33490](https://github.com/BerriAI/litellm/pull/33490)
- fix(model\_prices): refresh deprecation dates, correct xAI pricing and add missing provider models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36403](https://github.com/BerriAI/litellm/pull/36403)
- feat(ptu): gate PTU flat-cost attribution behind an opt-in env var by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36138](https://github.com/BerriAI/litellm/pull/36138)
- ci: cache Prisma CLI and engine binaries, split test timeout from setup by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36417](https://github.com/BerriAI/litellm/pull/36417)
- feat(rate limiting): configurable estimated output tokens per key, team and model by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36143](https://github.com/BerriAI/litellm/pull/36143)
- fix(ui): hide admin-only Logs tabs from roles that cannot call their endpoints by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36333](https://github.com/BerriAI/litellm/pull/36333)
- test(proxy): guard management\_v1 against fastapi names removed in supported releases by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36336](https://github.com/BerriAI/litellm/pull/36336)
- fix(ui): gate policy and prompt lookups on an admin capability by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36335](https://github.com/BerriAI/litellm/pull/36335)
- build(deps): bump pypdf to 6.15.0 to clear osv-scan by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36350](https://github.com/BerriAI/litellm/pull/36350)
- fix(proxy): isolate guardrail load failures per row by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36432](https://github.com/BerriAI/litellm/pull/36432)
- fix(ui): gate organization and agent usage views behind capabilities by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36334](https://github.com/BerriAI/litellm/pull/36334)
- fix(reset\_budget\_job): atomic budget cascade with chunked reset scans by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36287](https://github.com/BerriAI/litellm/pull/36287)
- feat(proxy): add GET /v1/indexes to list vector store indexes by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36289](https://github.com/BerriAI/litellm/pull/36289)
- feat(ui): show vector store indexes on the Vector Stores page by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36306](https://github.com/BerriAI/litellm/pull/36306)
- fix(proxy): treat SAML as configured in UI SSO detection by [@&#8203;fancybear-dev](https://github.com/fancybear-dev) in [#&#8203;36196](https://github.com/BerriAI/litellm/pull/36196)
- fix(bedrock): reject Anthropic server-side web\_search tool with actionable error by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36473](https://github.com/BerriAI/litellm/pull/36473)
- fix(ui): open the classifier prompt editor above the edit auto-router form by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36438](https://github.com/BerriAI/litellm/pull/36438)
- fix(arize): trace MCP tool calls instead of crashing on CallToolResult by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36453](https://github.com/BerriAI/litellm/pull/36453)
- refactor(ui): make illegal DataTable prop combinations unrepresentable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36470](https://github.com/BerriAI/litellm/pull/36470)
- fix(ui): scope Virtual Keys and Logs team lists to the caller by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36472](https://github.com/BerriAI/litellm/pull/36472)
- fix(ui): gate the Old Usage page behind a proxy-admin capability by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36469](https://github.com/BerriAI/litellm/pull/36469)
- docs(terraform): describe the provider release as automatic by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36467](https://github.com/BerriAI/litellm/pull/36467)
- feat(proxy): add per-deployment keepalive\_seconds SSE heartbeat to prevent load-balancer timeout on long streams by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;34423](https://github.com/BerriAI/litellm/pull/34423)
- fix(router): cool down failed fallback deployments and correct cooldown TTL after Redis backfill by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;35104](https://github.com/BerriAI/litellm/pull/35104)
- perf(spend): write each daily spend batch in one upsert statement by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36448](https://github.com/BerriAI/litellm/pull/36448)
- fix(ui): gate four sidebar pages on the roles their endpoints allow by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36475](https://github.com/BerriAI/litellm/pull/36475)
- fix(ui): restore the Logs Deleted Teams tab for organization admins by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36478](https://github.com/BerriAI/litellm/pull/36478)
- fix(websearch): stop leaking interception control fields to providers by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36480](https://github.com/BerriAI/litellm/pull/36480)
- test(e2e): cover the Anthropic web\_search server tool on Bedrock by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36443](https://github.com/BerriAI/litellm/pull/36443)
- fix(router): warn when a deployment's credentials contradict its provider by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36486](https://github.com/BerriAI/litellm/pull/36486)
- fix: net prompt-caching savings against the cache-write premium by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36452](https://github.com/BerriAI/litellm/pull/36452)
- feat(ui): deployment affinity toggle for the auto-router by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36302](https://github.com/BerriAI/litellm/pull/36302)
- fix(bedrock): use deployment credentials for AWS requests by [@&#8203;daleselaji-dev](https://github.com/daleselaji-dev) in [#&#8203;36160](https://github.com/BerriAI/litellm/pull/36160)
- fix(anthropic): preserve midturn system corrections by [@&#8203;eugene-yao-zocdoc](https://github.com/eugene-yao-zocdoc) in [#&#8203;34290](https://github.com/BerriAI/litellm/pull/34290)
- fix(email): stop duplicate legacy invitation email and fix its onboarding link by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36455](https://github.com/BerriAI/litellm/pull/36455)
- feat(ui): show models under each tier in routing benchmark chart by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36291](https://github.com/BerriAI/litellm/pull/36291)
- fix(proxy): inject streaming usage cost on openai passthrough streams by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36503](https://github.com/BerriAI/litellm/pull/36503)
- docs: require a user flow and live-proxy proof in bug reports by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36498](https://github.com/BerriAI/litellm/pull/36498)
- fix(proxy): add config\_updated\_at audit timestamp for virtual keys by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36488](https://github.com/BerriAI/litellm/pull/36488)
- docs: require a user flow and a stuck-at proof in feature requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36500](https://github.com/BerriAI/litellm/pull/36500)
- feat(router): add required-AND (&) tag prefix and allow\_fail\_open flag by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;36193](https://github.com/BerriAI/litellm/pull/36193)
- feat(proxy): per-key prompt caching toggle via enable\_prompt\_caching by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36466](https://github.com/BerriAI/litellm/pull/36466)
- fix(bedrock): send tool-search beta header for Haiku 4.5 on Invoke /v1/messages by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36502](https://github.com/BerriAI/litellm/pull/36502)
- fix(bedrock): preserve adaptive thinking effort through the /v1/messages bridge by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36507](https://github.com/BerriAI/litellm/pull/36507)
- ci: retry transient network fetch failures in lint workflow by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36563](https://github.com/BerriAI/litellm/pull/36563)
- fix(ui): stub useIsOrgAdmin in UsageTab tests so useCan needs no QueryClient by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36565](https://github.com/BerriAI/litellm/pull/36565)
- fix(alerting): dedupe scheduled Slack spend reports across pods by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36489](https://github.com/BerriAI/litellm/pull/36489)
- chore(typing): clear 1.6k basedpyright Any errors across 56 files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36543](https://github.com/BerriAI/litellm/pull/36543)
- fix(bedrock): add text block to converse user messages carrying documents by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36499](https://github.com/BerriAI/litellm/pull/36499)
- fix(deps): ship boto3 with the base SDK so bedrock works out of the box by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36568](https://github.com/BerriAI/litellm/pull/36568)
- fix(model\_prices): add provider-announced deprecation dates for Bedrock, Mistral, Cohere and Gemini models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36538](https://github.com/BerriAI/litellm/pull/36538)
- chore: bump litellm-enterprise 0.1.54 -> 0.1.55, litellm-proxy-extras 0.4.84 -> 0.4.85, litellm 1.97.0 -> 1.98.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36577](https://github.com/BerriAI/litellm/pull/36577)
- fix(bedrock\_guardrails): skip ApplyGuardrail when there is no content to scan by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36441](https://github.com/BerriAI/litellm/pull/36441)
- fix(e2e): assert on the gen-AI span that served the stream, not the span count by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36582](https://github.com/BerriAI/litellm/pull/36582)
- test(e2e): harden vendor API coverage by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;34557](https://github.com/BerriAI/litellm/pull/34557)
- test(e2e): add reproducers for passthrough and model budget gaps by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;34657](https://github.com/BerriAI/litellm/pull/34657)
- test(e2e): cover google-native generateContent framing and prometheus queue time by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;34650](https://github.com/BerriAI/litellm/pull/34650)
- chore(ci): promote internal staging to main by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36560](https://github.com/BerriAI/litellm/pull/36560)
- feat(router): make routing groups callable as virtual models and list them in /v1/models by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36519](https://github.com/BerriAI/litellm/pull/36519)
- fix(xai): bill web\_search from server\_side\_tool\_usage\_details by [@&#8203;geraint0923](https://github.com/geraint0923) in [#&#8203;30817](https://github.com/BerriAI/litellm/pull/30817)
- fix(responses): init completed\_response on bridge streaming iterator ([#&#8203;35411](https://github.com/BerriAI/litellm/issues/35411)) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35413](https://github.com/BerriAI/litellm/pull/35413)
- fix(batches): attribute Anthropic passthrough batch cost to the creating key, team and tags by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36468](https://github.com/BerriAI/litellm/pull/36468)
- feat(dashscope): add latest Model Studio models to the cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36496](https://github.com/BerriAI/litellm/pull/36496)
- fix(proxy): track streamed passthrough Responses cost by [@&#8203;william-xue](https://github.com/william-xue) in [#&#8203;36529](https://github.com/BerriAI/litellm/pull/36529)
- fix(model\_prices): advertise native structured output on every Bedrock DeepSeek V3.2 and GLM 5 id by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36597](https://github.com/BerriAI/litellm/pull/36597)
- test(bedrock): repoint live Claude tests off the retired Claude 3 Sonnet by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36600](https://github.com/BerriAI/litellm/pull/36600)
- fix(anthropic): preserve speed=fast in usage for /v1/messages and pass-through by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36447](https://github.com/BerriAI/litellm/pull/36447)
- fix(proxy): forward resolved provider and deployment pricing in /cost/estimate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35880](https://github.com/BerriAI/litellm/pull/35880)
- feat(proxy): global SSE keepalive ping interval for OpenAI-shaped streaming routes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36154](https://github.com/BerriAI/litellm/pull/36154)
- fix(responses): preserve Codex namespace tool calls by [@&#8203;dcadenas](https://github.com/dcadenas) in [#&#8203;32536](https://github.com/BerriAI/litellm/pull/32536)
- fix(nvidia\_nim): preserve image passages and stop sending top\_k to /v1/ranking by [@&#8203;atomic](https://github.com/atomic) in [#&#8203;34177](https://github.com/BerriAI/litellm/pull/34177)
- fix: refactor HTTP handler initialization with client support by [@&#8203;Praveen11558](https://github.com/Praveen11558) in [#&#8203;30952](https://github.com/BerriAI/litellm/pull/30952)
- feat(lint): gate writable TypedDict fields with LIT012 by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36590](https://github.com/BerriAI/litellm/pull/36590)
- perf(proxy): stagger scheduled background jobs across jobs and pods by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36589](https://github.com/BerriAI/litellm/pull/36589)
- test: remove four mirror test files that exercise none of their module by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34635](https://github.com/BerriAI/litellm/pull/34635)
- fix(router): stop re-applying router-selecting request tags to the routed tier's deployments by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36628](https://github.com/BerriAI/litellm/pull/36628)
- test: remove tests that never execute by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36681](https://github.com/BerriAI/litellm/pull/36681)
- fix(ui): align spend and budget columns by [@&#8203;daniel-meismer-zocdoc](https://github.com/daniel-meismer-zocdoc) in [#&#8203;35176](https://github.com/BerriAI/litellm/pull/35176)
- test: rename tests that a later definition shadowed by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36685](https://github.com/BerriAI/litellm/pull/36685)
- fix(passthrough): carry the budget reservation into request metadata by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36592](https://github.com/BerriAI/litellm/pull/36592)
- fix(mcp): bound MCP client requests with a session read timeout by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36675](https://github.com/BerriAI/litellm/pull/36675)
- fix(proxy): log requests rejected for an unparsable body in spend logs by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36673](https://github.com/BerriAI/litellm/pull/36673)
- refactor(ui): migrate cost-optimization to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36629](https://github.com/BerriAI/litellm/pull/36629)
- refactor(ui): migrate cost-tracking to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36631](https://github.com/BerriAI/litellm/pull/36631)
- refactor(ui): migrate admin-panel to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36635](https://github.com/BerriAI/litellm/pull/36635)
- refactor(ui): migrate users dashboard to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36642](https://github.com/BerriAI/litellm/pull/36642)
- refactor(ui): migrate prompts to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36643](https://github.com/BerriAI/litellm/pull/36643)
- refactor(ui): migrate team settings to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36641](https://github.com/BerriAI/litellm/pull/36641)
- refactor(ui): migrate models-and-endpoints to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36648](https://github.com/BerriAI/litellm/pull/36648)
- refactor(ui): migrate policy impact popover to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36653](https://github.com/BerriAI/litellm/pull/36653)
- fix(proxy): expand config-defined model access groups when resolving team models for /v2/model/info by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;34211](https://github.com/BerriAI/litellm/pull/34211)
- fix(batches): strip NUL bytes from passthrough batch tags before the managed object write by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36688](https://github.com/BerriAI/litellm/pull/36688)
- test(e2e-ui): verify UI mutations against the API instead of trusting the toast by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36632](https://github.com/BerriAI/litellm/pull/36632)
- fix(proxy): serialize model reconciles so concurrent model writes stop evicting each other by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36687](https://github.com/BerriAI/litellm/pull/36687)
- chore(e2e): port the compat-matrix cron publisher to tests/e2e/claude\_code by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36465](https://github.com/BerriAI/litellm/pull/36465)
- fix(router): never price a strategy-router alias by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36691](https://github.com/BerriAI/litellm/pull/36691)
- feat(model\_prices): add NVIDIA Nemotron 3.5 Lightning on OpenRouter and DeepInfra by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36696](https://github.com/BerriAI/litellm/pull/36696)
- feat(terraform/aws): make VPC, Aurora, and Redis optional by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36676](https://github.com/BerriAI/litellm/pull/36676)
- feat(ui): warn in the Admin UI when no Redis is configured by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36495](https://github.com/BerriAI/litellm/pull/36495)
- fix(ui): show and edit key-level router settings on a virtual key by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36674](https://github.com/BerriAI/litellm/pull/36674)
- fix(router): forward auto-router alias params from the marker entry, not the first same-name deployment by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36626](https://github.com/BerriAI/litellm/pull/36626)
- fix(bedrock\_mantle): 1M context window and long-context pricing for GPT-5.6 Sol/Terra/Luna by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36698](https://github.com/BerriAI/litellm/pull/36698)
- fix(model\_prices): sync the Groq registry with Groq's docs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36664](https://github.com/BerriAI/litellm/pull/36664)
- fix(router): let untagged requests bypass a tagged pre-routing strategy on shared model names by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36627](https://github.com/BerriAI/litellm/pull/36627)
- fix(spend): stop losing spend log rows when a flush is cancelled by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34826](https://github.com/BerriAI/litellm/pull/34826)
- docs(claude): drop the @&#8203; prefix from the PR template path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36726](https://github.com/BerriAI/litellm/pull/36726)
- fix(langfuse): emit otel trace version and release on the keys langfuse v4 reads by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36702](https://github.com/BerriAI/litellm/pull/36702)
- test(interactions): follow Google spec drift replacing Turn with typed steps by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36730](https://github.com/BerriAI/litellm/pull/36730)
- refactor(ui): migrate team detail controls to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36695](https://github.com/BerriAI/litellm/pull/36695)
- refactor(ui): migrate guardrail and duration controls to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36693](https://github.com/BerriAI/litellm/pull/36693)
- refactor(ui): migrate guardrails-monitor, projects, logs to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34606](https://github.com/BerriAI/litellm/pull/34606)
- refactor(ui): migrate search and user controls to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36694](https://github.com/BerriAI/litellm/pull/36694)
- fix(guardrails): scan and re-emit raw Anthropic SSE streams in the bedrock post-call hook by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36598](https://github.com/BerriAI/litellm/pull/36598)
- fix(helm): render nodeSelector on the migrations job by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36747](https://github.com/BerriAI/litellm/pull/36747)
- fix(langfuse): coerce header-sourced mask and trace-update steering values by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36740](https://github.com/BerriAI/litellm/pull/36740)
- refactor(ui): migrate usage tables to shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36707](https://github.com/BerriAI/litellm/pull/36707)
- refactor(ui): migrate guardrails monitor table to shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36709](https://github.com/BerriAI/litellm/pull/36709)
- refactor(ui): migrate guardrails content tables to shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36708](https://github.com/BerriAI/litellm/pull/36708)
- feat(gemini): day-0 pricing for gemini-3.7-flash by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36792](https://github.com/BerriAI/litellm/pull/36792)
- ci: promote staging to main by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36725](https://github.com/BerriAI/litellm/pull/36725)
- build(deps): bump nanoid to 3.3.18 to clear osv-scan by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36787](https://github.com/BerriAI/litellm/pull/36787)
- fix(router): stop scoring system prompt text for code/technical complexity by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36721](https://github.com/BerriAI/litellm/pull/36721)
- feat(complexity\_router): calibrate the classifier rubric with worked examples, selectable per router by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36578](https://github.com/BerriAI/litellm/pull/36578)
- fix(interactions): map step and turn history to Responses API roles and content types by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36733](https://github.com/BerriAI/litellm/pull/36733)
- fix(ui): restore playground model filtering by endpoint by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36130](https://github.com/BerriAI/litellm/pull/36130)
- fix(proxy/batches): stop forwarding custom\_llm\_provider twice in list and cancel by [@&#8203;anxkhn](https://github.com/anxkhn) in [#&#8203;32813](https://github.com/BerriAI/litellm/pull/32813)
- refactor(ui): migrate TokenFlow and JsonViewer to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36735](https://github.com/BerriAI/litellm/pull/36735)
- feat: pre-adoption shadow eval for the auto-router (blind pairwise judge, derived state) by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36587](https://github.com/BerriAI/litellm/pull/36587)
- refactor(ui): migrate SimpleMessageBlock and SimpleToolCallBlock to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36737](https://github.com/BerriAI/litellm/pull/36737)
- refactor(ui): migrate HistoryTree and CollapsibleMessage to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36738](https://github.com/BerriAI/litellm/pull/36738)
- refactor: replace Any with precise types across responses, proxy, and llms modules by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36763](https://github.com/BerriAI/litellm/pull/36763)
- refactor(ui): migrate TruncatedValue and OutputCard to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36739](https://github.com/BerriAI/litellm/pull/36739)
- refactor(ui): migrate SectionHeader and ToolsSection to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36793](https://github.com/BerriAI/litellm/pull/36793)
- feat(ui): migrate playground chat controls to shadcn by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36129](https://github.com/BerriAI/litellm/pull/36129)
- feat(xai): day-0 pricing for grok-4.6 by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36805](https://github.com/BerriAI/litellm/pull/36805)
- feat(ui): highlight Auto Router in the navbar announcement by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36315](https://github.com/BerriAI/litellm/pull/36315)
- test(e2e): assert the model allow-list permits, not only denies by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36823](https://github.com/BerriAI/litellm/pull/36823)
- fix(proxy): tolerate a concurrent creator when creating spend views by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36824](https://github.com/BerriAI/litellm/pull/36824)
- fix(proxy): honor explicit null budget\_duration on team and key create + clearable UI dropdowns by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36699](https://github.com/BerriAI/litellm/pull/36699)
- feat(model\_prices): add meta/muse-spark-1.2 and its contributor tier by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36717](https://github.com/BerriAI/litellm/pull/36717)
- fix(auth): carry team grants in lite login session tokens by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36826](https://github.com/BerriAI/litellm/pull/36826)
- feat(ui): show provider prompt cache tokens in chat response metrics by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36827](https://github.com/BerriAI/litellm/pull/36827)
- fix(auth): stop the team fallback from widening model access by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36837](https://github.com/BerriAI/litellm/pull/36837)
- fix(proxy/team): resolve member\_delete cleanup by user id, not the addressed email by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36839](https://github.com/BerriAI/litellm/pull/36839)
- fix(cli): launch agents as a child process on Windows by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36822](https://github.com/BerriAI/litellm/pull/36822)
- feat(ui): shadow evals tab beside auto-router usage by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36588](https://github.com/BerriAI/litellm/pull/36588)
- feat(cli): make the hidden `lite` command list configurable by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36816](https://github.com/BerriAI/litellm/pull/36816)
- feat(azure\_ai): add Fireworks FW model pricing on Azure AI Foundry by [@&#8203;emerzon](https://github.com/emerzon) in [#&#8203;35613](https://github.com/BerriAI/litellm/pull/35613)
- fix: enable xhigh reasoning support for gpt-5.4-mini models by [@&#8203;emerzon](https://github.com/emerzon) in [#&#8203;26909](https://github.com/BerriAI/litellm/pull/26909)
- feat(azure-ai): add Grok 4.3 model metadata by [@&#8203;emerzon](https://github.com/emerzon) in [#&#8203;27932](https://github.com/BerriAI/litellm/pull/27932)
- feat(ui): render request metrics on the /ui/chat surface by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36845](https://github.com/BerriAI/litellm/pull/36845)
- fix(ui): stop a deselected MCP server keeping its grant on a virtual key by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36840](https://github.com/BerriAI/litellm/pull/36840)
- fix(team): sweep dangling team references and cache on team delete by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36819](https://github.com/BerriAI/litellm/pull/36819)
- fix(mcp): resolve admin OAuth sessions from any worker via DB-backed drafts by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36844](https://github.com/BerriAI/litellm/pull/36844)
- refactor(ui): migrate usage to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36834](https://github.com/BerriAI/litellm/pull/36834)
- refactor(ui): migrate guardrails-monitor to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36838](https://github.com/BerriAI/litellm/pull/36838)
- refactor(ui): migrate playground to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36847](https://github.com/BerriAI/litellm/pull/36847)
- refactor(ui): migrate guardrails to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36832](https://github.com/BerriAI/litellm/pull/36832)
- fix(batches): stop uncostable batches from starving the cost poll page by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36714](https://github.com/BerriAI/litellm/pull/36714)
- perf(spend-logs): bound retention cleanup so one run cannot saturate the database by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36594](https://github.com/BerriAI/litellm/pull/36594)
- fix(proxy): fail config load when a callbacks entry is not dispatchable by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36858](https://github.com/BerriAI/litellm/pull/36858)
- fix(bedrock): hoist custom.defer\_loading before dropping custom on invoke tools by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36855](https://github.com/BerriAI/litellm/pull/36855)
- fix(access groups): sync assigned\_key\_ids from the key write paths by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36843](https://github.com/BerriAI/litellm/pull/36843)
- fix(mcp): expose client HTTP headers to logging callbacks and hooks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36724](https://github.com/BerriAI/litellm/pull/36724)
- fix(ptu): stop per-token billing on a PTU-configured deployment by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36829](https://github.com/BerriAI/litellm/pull/36829)
- fix(ui): add nvidia riva to the model provider list by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36769](https://github.com/BerriAI/litellm/pull/36769)
- fix(scripts): end make check with a ran/skipped summary and verdict by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36864](https://github.com/BerriAI/litellm/pull/36864)
- fix(proxy): track spend for OpenAI passthrough /v1/embeddings by [@&#8203;lostmartian](https://github.com/lostmartian) in [#&#8203;36660](https://github.com/BerriAI/litellm/pull/36660)
- test(proxy): stop monkeypatch.undo re-planting fixture-mocked prisma\_client by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36872](https://github.com/BerriAI/litellm/pull/36872)
- fix(access groups): sync assigned\_team\_ids from the team write paths by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36825](https://github.com/BerriAI/litellm/pull/36825)
- ci: drop the CircleCI ui\_build and ui\_unit\_tests jobs by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36893](https://github.com/BerriAI/litellm/pull/36893)
- fix(langfuse)!: source the emitted metadata blob from StandardLoggingPayload by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36744](https://github.com/BerriAI/litellm/pull/36744)
- refactor(ui): migrate Navbar off antd to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36902](https://github.com/BerriAI/litellm/pull/36902)
- refactor(ui): migrate log details drawer off antd to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36904](https://github.com/BerriAI/litellm/pull/36904)
- refactor(ui): migrate AI Hub off antd and tremor to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36908](https://github.com/BerriAI/litellm/pull/36908)
- refactor(ui): move the shared dropdowns and selectors onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36924](https://github.com/BerriAI/litellm/pull/36924)
- refactor(ui): move the root-level dashboard components onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36927](https://github.com/BerriAI/litellm/pull/36927)
- refactor(ui): move the settings page and bulk user invite onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36936](https://github.com/BerriAI/litellm/pull/36936)
- refactor(ui): move the cost tracking components onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36955](https://github.com/BerriAI/litellm/pull/36955)
- ci: drop the duplicate proxy\_unit\_tests letter-shard workflow by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36866](https://github.com/BerriAI/litellm/pull/36866)
- refactor(ui): migrate shared common\_components off antd and tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36910](https://github.com/BerriAI/litellm/pull/36910)
- refactor(ui): migrate key info and permissions views off antd and tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36913](https://github.com/BerriAI/litellm/pull/36913)
- feat(proxy): serve Anthropic-native /v1/models for Claude Code gateway discovery by [@&#8203;Ar-maan05](https://github.com/Ar-maan05) in [#&#8203;35455](https://github.com/BerriAI/litellm/pull/35455)
- refactor(ui): migrate router settings and shared badges off antd and tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36915](https://github.com/BerriAI/litellm/pull/36915)
- refactor(ui): move the model hub and model select onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36918](https://github.com/BerriAI/litellm/pull/36918)
- fix(ui): keep the cost tracking removal confirmation open until it settles by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36960](https://github.com/BerriAI/litellm/pull/36960)
- refactor(ui): declare DateRangePickerValue locally instead of importing it from tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36962](https://github.com/BerriAI/litellm/pull/36962)
- fix(main): an explicit provider outranks a known OpenAI model name by [@&#8203;FahimaGold](https://github.com/FahimaGold) in [#&#8203;36800](https://github.com/BerriAI/litellm/pull/36800)
- fix(exception\_mapping): bare 429 in an error body no longer outranks the status code by [@&#8203;FahimaGold](https://github.com/FahimaGold) in [#&#8203;36705](https://github.com/BerriAI/litellm/pull/36705)
- refactor(ui): move MCP permission panels onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36964](https://github.com/BerriAI/litellm/pull/36964)
- refactor(ui): migrate ten small dashboard files off antd and tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36966](https://github.com/BerriAI/litellm/pull/36966)
- fix(proxy): force prisma recreate on postgres cached-plan error by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36428](https://github.com/BerriAI/litellm/pull/36428)
- fix(transcription): stop a zero output rate from zeroing transcription cost by [@&#8203;hMED22](https://github.com/hMED22) in [#&#8203;36914](https://github.com/BerriAI/litellm/pull/36914)
- fix(langfuse): restrict trace steering keys to real langfuse trace fields by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36862](https://github.com/BerriAI/litellm/pull/36862)
- Revert "fix(auth): stop the team fallback from widening model access" ([#&#8203;36837](https://github.com/BerriAI/litellm/issues/36837)) by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36982](https://github.com/BerriAI/litellm/pull/36982)
- fix(ui): show zeroed auto-router usage stats when a window has no sessions by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36868](https://github.com/BerriAI/litellm/pull/36868)
- fix(mcp): keep admin-entered oauth endpoints in management reads by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36888](https://github.com/BerriAI/litellm/pull/36888)
- fix(ui): distinguish hosted and local vLLM in the provider dropdown by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36974](https://github.com/BerriAI/litellm/pull/36974)
- fix(openai,azure): return a length-truncated 200 when the output budget fits no token by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36859](https://github.com/BerriAI/litellm/pull/36859)
- fix(proxy): always emit the Anthropic /v1/models token limits, null when unknown by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36961](https://github.com/BerriAI/litellm/pull/36961)
- feat(helm): add startupProbe and hpa.behavior to the componentized chart by [@&#8203;Louis-Vauterin](https://github.com/Louis-Vauterin) in [#&#8203;36382](https://github.com/BerriAI/litellm/pull/36382)
- fix(proxy): serve aggregate MCP endpoint on bare /mcp instead of 307-redirecting by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;34845](https://github.com/BerriAI/litellm/pull/34845)
- feat(shadow\_eval): add reverse-direction shadow eval jobs by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36865](https://github.com/BerriAI/litellm/pull/36865)
- fix(proxy): requeue Redis spend buffer transactions when the DB commit fails by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33881](https://github.com/BerriAI/litellm/pull/33881)
- feat(search): add Nimble as a search provider by [@&#8203;ilchemla](https://github.com/ilchemla) in [#&#8203;36347](https://github.com/BerriAI/litellm/pull/36347)
- fix(mcp): drop caller host and configured upstream headers from logged metadata by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36901](https://github.com/BerriAI/litellm/pull/36901)
- fix(azure\_ai): recognize real Search doc endpoints so teams can read/write via passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36798](https://github.com/BerriAI/litellm/pull/36798)
- fix(anthropic): aggregate 5m/1h cache-write split across iterations path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34860](https://github.com/BerriAI/litellm/pull/34860)
- fix(anthropic cost): apply regional geo uplift to cached tokens by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34850](https://github.com/BerriAI/litellm/pull/34850)
- fix(ui): match the MCP servers count badge to its sibling permission badges by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36984](https://github.com/BerriAI/litellm/pull/36984)
- fix(batches): mark terminal batch with no output file as processed in CheckBatchCost by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35360](https://github.com/BerriAI/litellm/pull/35360)
- fix(caching): cache anthropic /v1/messages responses, including streaming by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34581](https://github.com/BerriAI/litellm/pull/34581)
- fix(anthropic\_messages): make tool\_result images visible to OpenAI-compatible providers by [@&#8203;hMED22](https://github.com/hMED22) in [#&#8203;34462](https://github.com/BerriAI/litellm/pull/34462)
- feat(fireworks\_ai): translate NIM/vLLM extra params to Fireworks-native args by [@&#8203;milesadkins](https://github.com/milesadkins) in [#&#8203;35969](https://github.com/BerriAI/litellm/pull/35969)
- fix(ui): stop the models tab strip from scrolling vertically by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36993](https://github.com/BerriAI/litellm/pull/36993)
- fix(ui): anchor chips-combobox popups to the field instead of the inner input by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36995](https://github.com/BerriAI/litellm/pull/36995)
- feat(proxy): per-component response cost headers by [@&#8203;erensh27](https://github.com/erensh27) in [#&#8203;36965](https://github.com/BerriAI/litellm/pull/36965)
- fix(cost): track OpenAI/Azure web search tool cost per call by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35286](https://github.com/BerriAI/litellm/pull/35286)
- fix(bedrock): resolve aliases in batch file records by [@&#8203;daleselaji-dev](https://github.com/daleselaji-dev) in [#&#8203;36159](https://github.com/BerriAI/litellm/pull/36159)
- fix: report real token usage on guardrail-blocked /v1/responses replies by [@&#8203;guptaishaan](https://github.com/guptaishaan) in [#&#8203;36907](https://github.com/BerriAI/litellm/pull/36907)
- fix(proxy): requeue spend logs when the DB write fails with a transport error by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36716](https://github.com/BerriAI/litellm/pull/36716)
- fix(cost): tiered pricing supports cache creation cost and is all-or-nothing by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36720](https://github.com/BerriAI/litellm/pull/36720)
- fix(vertex\_ai): translate /v1/embeddings batch rows to the Gemini embedding shape by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35092](https://github.com/BerriAI/litellm/pull/35092)
- docs(claude): require ReadOnly on every TypedDict field (LIT012) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37005](https://github.com/BerriAI/litellm/pull/37005)
- refactor(ui): migrate access group create modal to RHF + zod + shadcn by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37033](https://github.com/BerriAI/litellm/pull/37033)
- refactor(ui): re-sync badge and skeleton onto the base-vega shadcn style by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36991](https://github.com/BerriAI/litellm/pull/36991)
- feat(ui): link user detail team names to team pages by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37022](https://github.com/BerriAI/litellm/pull/37022)
- fix(model\_prices): correct DeepSeek V4 max output tokens by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36925](https://github.com/BerriAI/litellm/pull/36925)
- fix(ui): rename models table Status column to Source by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37021](https://github.com/BerriAI/litellm/pull/37021)
- chore: bump litellm-enterprise 0.1.55 -> 0.1.56, litellm-proxy-extras 0.4.85 -> 0.4.86 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37045](https://github.com/BerriAI/litellm/pull/37045)
- feat(proxy): gate the Global Control Plane worker registry on an enterprise license by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36996](https://github.com/BerriAI/litellm/pull/36996)
- fix(model\_prices): add gemini 3.1 flash tts preview and legacy OpenAI shutdown dates by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36788](https://github.com/BerriAI/litellm/pull/36788)
- fix(panw\_prisma\_airs): surface scan\_id on allowed requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37037](https://github.com/BerriAI/litellm/pull/37037)
- fix(model\_map): flag native structured outputs on Anthropic-direct claude-sonnet-5 and claude-haiku-4-5 by [@&#8203;anmolg1997](https://github.com/anmolg1997) in [#&#8203;35930](https://github.com/BerriAI/litellm/pull/35930)
- fix(router): stop get\_router\_model\_info from wiping cached pricing by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36985](https://github.com/BerriAI/litellm/pull/36985)
- fix(redis): unwrap decorated \_\_init\_\_s when deriving the from\_url kwargs allowlist by [@&#8203;anmolg1997](https://github.com/anmolg1997) in [#&#8203;36654](https://github.com/BerriAI/litellm/pull/36654)
- fix(proxy): reserve the larger declared output budget for TPM limits by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37001](https://github.com/BerriAI/litellm/pull/37001)
- fix(databricks): surface provider usage, including prompt-cache counts, in streaming chunks by [@&#8203;pokepoke81](https://github.com/pokepoke81) in [#&#8203;36943](https://github.com/BerriAI/litellm/pull/36943)
- fix(spend): give a batch's cost row a primary key of its own by [@&#8203;marty-sullivan](https://github.com/marty-sullivan) in [#&#8203;36876](https://github.com/BerriAI/litellm/pull/36876)
- feat: shadow eval samples /v1/messages and /v1/responses traffic by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36830](https://github.com/BerriAI/litellm/pull/36830)
- fix(ptu): stop a PTU deployment billing for grounded search by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37043](https://github.com/BerriAI/litellm/pull/37043)
- fix(fireworks\_ai): support router slugs via routers/ prefix by [@&#8203;heathriel](https://github.com/heathriel) in [#&#8203;34257](https://github.com/BerriAI/litellm/pull/34257)
- fix(bedrock): register managed-batch litellm\_params so they stop leaking to the provider (internal copy of [#&#8203;36633](https://github.com/BerriAI/litellm/issues/36633)) by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37048](https://github.com/BerriAI/litellm/pull/37048)
- fix(bedrock): resolve the managed-batch output bucket on every path that reads it by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37047](https://github.com/BerriAI/litellm/pull/37047)
- fix(bedrock): resolve the managed-batch output bucket on every path that reads it by [@&#8203;marty-sullivan](https://github.com/marty-sullivan) in [#&#8203;36634](https://github.com/BerriAI/litellm/pull/36634)
- feat(scripts): queue heavy gates behind a machine-wide slot lock by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36988](https://github.com/BerriAI/litellm/pull/36988)
- feat(mcp): scope gateway session bearers to the RFC 8707 resource by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35045](https://github.com/BerriAI/litellm/pull/35045)
- feat(ui): direction picker and reverse-mode display for shadow evals by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36994](https://github.com/BerriAI/litellm/pull/36994)
- fix(guardrails): return the full PANW AIRS scan response on blocked requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37036](https://github.com/BerriAI/litellm/pull/37036)
- fix(passthrough): stop forwarding client Accept-Encoding upstream by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37058](https://github.com/BerriAI/litellm/pull/37058)
- fix(batches): account a managed batch's cost exactly once by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37050](https://github.com/BerriAI/litellm/pull/37050)
- fix(panw\_prisma\_airs): scan tool call args as plain text, not a tool\_event by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37038](https://github.com/BerriAI/litellm/pull/37038)
- feat(lint): exempt TypedDict-annotated dict literals from LIT002 by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36869](https://github.com/BerriAI/litellm/pull/36869)
- docs(claude): tell agents to let heavy gates queue for machine-wide slots by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37057](https://github.com/BerriAI/litellm/pull/37057)
- test: unstick the suites CircleCI is failing on by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37059](https://github.com/BerriAI/litellm/pull/37059)
- docs(github): proof-of-fix section shows only the latest run as Before/After with nested cases by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37063](https://github.com/BerriAI/litellm/pull/37063)
- test(e2e): assert provider error shape instead of pinned prose by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37065](https://github.com/BerriAI/litellm/pull/37065)
- fix(ui): de-duplicate the reset budget option and polish shadcn surfaces by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37010](https://github.com/BerriAI/litellm/pull/37010)
- chore: rebuild Admin UI bundle from litellm\_internal\_staging by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37066](https://github.com/BerriAI/litellm/pull/37066)
- test(e2e/ui): assert the log drawer chevrons by their lucide classes by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37069](https://github.com/BerriAI/litellm/pull/37069)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37042](https://github.com/BerriAI/litellm/pull/37042)
- fix(ui): keep completion-mode models in the playground chat dropdown (backport to rc/1.98.0) by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37955](https://github.com/BerriAI/litellm/pull/37955)

##### New Contributors

- [@&#8203;kr0k](https://github.com/kr0k) made their first contribution in [#&#8203;33196](https://github.com/BerriAI/litellm/pull/33196)
- [@&#8203;HuanQian571](https://github.com/HuanQian571) made their first contribution in [#&#8203;35773](https://github.com/BerriAI/litellm/pull/35773)
- [@&#8203;alexshtf](https://github.com/alexshtf) made their first contribution in [#&#8203;35669](https://github.com/BerriAI/litellm/pull/35669)
- [@&#8203;vairodp](https://github.com/vairodp) made their first contribution in [#&#8203;33490](https://github.com/BerriAI/litellm/pull/33490)
- [@&#8203;fancybear-dev](https://github.com/fancybear-dev) made their first contribution in [#&#8203;36196](https://github.com/BerriAI/litellm/pull/36196)
- [@&#8203;daleselaji-dev](https://github.com/daleselaji-dev) made their first contribution in [#&#8203;36160](https://github.com/BerriAI/litellm/pull/36160)
- [@&#8203;eugene-yao-zocdoc](https://github.com/eugene-yao-zocdoc) made their first contribution in [#&#8203;34290](https://github.com/BerriAI/litellm/pull/34290)
- [@&#8203;geraint0923](https://github.com/geraint0923) made their first contribution in [#&#8203;30817](https://github.com/BerriAI/litellm/pull/30817)
- [@&#8203;william-xue](https://github.com/william-xue) made their first contribution in [#&#8203;36529](https://github.com/BerriAI/litellm/pull/36529)
- [@&#8203;dcadenas](https://github.com/dcadenas) made their first contribution in [#&#8203;32536](https://github.com/BerriAI/litellm/pull/32536)
- [@&#8203;atomic](https://github.com/atomic) made their first contribution in [#&#8203;34177](https://github.com/BerriAI/litellm/pull/34177)
- [@&#8203;Praveen11558](https://github.com/Praveen11558) made their first contribution in [#&#8203;30952](https://github.com/BerriAI/litellm/pull/30952)
- [@&#8203;anxkhn](https://github.com/anxkhn) made their first contribution in [#&#8203;32813](https://github.com/BerriAI/litellm/pull/32813)
- [@&#8203;lostmartian](https://github.com/lostmartian) made their first contribution in [#&#8203;36660](https://github.com/BerriAI/litellm/pull/36660)
- [@&#8203;FahimaGold](https://github.com/FahimaGold) made their first contribution in [#&#8203;36800](https://github.com/BerriAI/litellm/pull/36800)
- [@&#8203;Louis-Vauterin](https://github.com/Louis-Vauterin) made their first contribution in [#&#8203;36382](https://github.com/BerriAI/litellm/pull/36382)
- [@&#8203;ilchemla](https://github.com/ilchemla) made their first contribution in [#&#8203;36347](https://github.com/BerriAI/litellm/pull/36347)
- [@&#8203;milesadkins](https://github.com/milesadkins) made their first contribution in [#&#8203;35969](https://github.com/BerriAI/litellm/pull/35969)
- [@&#8203;erensh27](https://github.com/erensh27) made their first contribution in [#&#8203;36965](https://github.com/BerriAI/litellm/pull/36965)
- [@&#8203;guptaishaan](https://github.com/guptaishaan) made their first contribution in [#&#8203;36907](https://github.com/BerriAI/litellm/pull/36907)
- [@&#8203;pokepoke81](https://github.com/pokepoke81) made their first contribution in [#&#8203;36943](https://github.com/BerriAI/litellm/pull/36943)
- [@&#8203;heathriel](https://github.com/heathriel) made their first contribution in [#&#8203;34257](https://github.com/BerriAI/litellm/pull/34257)

**Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.97.0...v1.98.0>

### [`v1.98.0`](https://github.com/BerriAI/litellm/releases/tag/v1.98.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.97.0...v1.98.0)

##### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.98.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.98.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.98.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

##### What's Changed

- fix(bedrock): drop toolSpec.strict for Claude Sonnet 5 on Converse by [@&#8203;kr0k](https://github.com/kr0k) in [#&#8203;33196](https://github.com/BerriAI/litellm/pull/33196)
- fix(batches): attribute Vertex passthrough batch cost to key/team/tags by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;34456](https://github.com/BerriAI/litellm/pull/34456)
- docs: rewrite the CLAUDE.md comment rule with explicit exceptions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36301](https://github.com/BerriAI/litellm/pull/36301)
- fix(proxy): scope file list pagination cursors to the caller by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36093](https://github.com/BerriAI/litellm/pull/36093)
- fix(proxy): skip prisma-dependent hooks when no database is attached by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36273](https://github.com/BerriAI/litellm/pull/36273)
- fix(proxy): report has\_more false on caller-scoped file list pages by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36326](https://github.com/BerriAI/litellm/pull/36326)
- fix(proxy): restore management\_v1 query-param validation under fastapi>=0.140.7 by [@&#8203;HuanQian571](https://github.com/HuanQian571) in [#&#8203;35773](https://github.com/BerriAI/litellm/pull/35773)
- fix(proxy): stop /{provider}/v1/files from capturing /openai\_passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36092](https://github.com/BerriAI/litellm/pull/36092)
- chore(typing): remove 914 basedpyright Any errors across 16 hotspot files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36386](https://github.com/BerriAI/litellm/pull/36386)
- fix(router): keep batch fallbacks inside the model group that owns the file by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36181](https://github.com/BerriAI/litellm/pull/36181)
- feat(ptu): configure provisioned-throughput flat cost on a model deployment by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35341](https://github.com/BerriAI/litellm/pull/35341)
- docs: clarify the CLAUDE.md comment exceptions are any-of by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36421](https://github.com/BerriAI/litellm/pull/36421)
- docs: replace the Changes …
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants