Skip to content

feat(ptu): surface PTU flat cost on the daily activity read path - #35391

Merged
yucheng-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_ptu_read_path
Aug 10, 2026
Merged

feat(ptu): surface PTU flat cost on the daily activity read path#35391
yucheng-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_ptu_read_path

Conversation

@yucheng-berri

@yucheng-berri yucheng-berri commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • The flat cost the rollup writes is not visible anywhere yet
  • Sentinel rollup rows must not pollute the per-key or per-provider breakdowns

How it solves it:

  • Aggregate ptu_flat_cost into SpendMetrics.flat_cost and total_flat_cost
  • Sentinel rows add to parent totals only, never appear as a key

Relevant issues

Linear ticket

Resolves LIT-4077

Pre-Submission checklist

  • I have added meaningful tests
  • My PR passes all CI/CD checks (e.g., lint, format, unit tests)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review

Screenshots / Proof of Fix

Captured live against a proxy on this branch (commit 7ead5b8), real DB, no mocks.

Rebased onto litellm_internal_staging at f6587fa on 2026-08-06, which had moved 441 commits past the original branch point. The whole matrix below was re-run against the rebased head on a live proxy and a real Postgres, so nothing here is carried over from the pre-rebase run

With flat-cost rows and real request rows both present, the response separates them rather than mixing them. Three chat completions went to Gemini 2.5 Flash through a team-scoped key, alongside a 30-day PTU window

GET /team/daily/activity?start_date=2026-07-07&end_date=2026-08-06&page_size=1000
  total_spend        : 0.0056465
  total_flat_cost    : 14880.0
  total_api_requests : 3
  total_tokens       : 2285

  sentinel in api_key breakdown : False
  providers seen                : ['gemini']

The request totals count only the three real calls, so the sentinel rows contribute their flat cost and nothing else. They also stay out of the api_key breakdown, and the provider breakdown holds only the provider that actually served traffic, with no bucket invented for the sentinel's empty provider

The rollup PR wrote a sentinel row of ptu_flat_cost 10 for a team (5 PTU, 2.0/hour, effective from 23:00 so one active hour). With this read path the amount now surfaces on the team daily activity response

GET /team/daily/activity?team_ids=<pr2-proof team>&start_date=2026-07-31&end_date=2026-07-31
-> metadata.total_flat_cost = 10.00

Before this PR the same call returned total_flat_cost 0 because the read path did not read the column. This is the third of a stack; the UI that renders it lands in a follow-up PR

Live run 2026-08-08, commit 8704e38

Ten sentinel rows worth $6,260.00 in LiteLLM_DailyTeamSpend, plus one real Gemini request so request spend and flat cost are both present

GET /team/daily/activity?start_date=2026-08-01&end_date=2026-08-08&page_size=1000

metadata.total_flat_cost   6260.0        matches the DB sum exactly
metadata.total_spend       0.000062      the real request, kept separate

per day   2026-08-01  720.0
          2026-08-05  720.0
          2026-08-06  740.0   = 720 + 20, two deployments
          2026-08-07  1200.0  = 720 + 480
          2026-08-08  720.0

breakdown.models          {"ptu-prod-renamed": 720.0}   display name, not the deployment UUID
breakdown.api_keys        the __ptu_flat_cost__ sentinel does not appear

page_size=1000 matters: the endpoint paginates and the totals are per page, so a bare curl under-reports and reads as a rollup bug. The dashboard passes 1000 and sums pages itself

Type

New Feature

Changes

update_metrics accumulates ptu_flat_cost into SpendMetrics.flat_cost, and _record_to_spend_metrics reads it on the aggregated path; _build_aggregated_sql_query selects SUM(ptu_flat_cost) only for litellm_dailyteamspend and a constant zero otherwise, so every daily table returns the same shape. DailySpendMetadata gains total_flat_cost, populated from the aggregated totals

update_breakdown_metrics guards every api_key sub-breakdown and the top-level api_keys map with the sentinel check, so the ptu_flat_cost row contributes to parent metrics but never becomes a key row; the provider breakdown is skipped entirely for sentinel rows since flat cost is not per-provider request spend. The GROUPING SETS dispatcher applies the same rules: sentinel rows stay out of every api_key sub-breakdown, and because the sentinel has no provider its flat cost is withheld from the provider parent rather than landing under "unknown". The provider bucket itself is still emitted unconditionally, matching the base build: a row predating the api_requests column backfills to all zeroes, and suppressing those would drop a provider the previous build reported

Screens

Flat cost reaching a team's response and rendering alongside request spend, with the sentinel absent from the key breakdown:

Team usage showing flat cost from the read path

Same build, tag usage: the shared response still carries the field at 0.0 and the view renders neither tile, so a non-team consumer sees no change:

Tag usage keeps five tiles and one series

Sentinel rows display their model name

The row keys on the deployment id so a rename cannot move it, and the per-model breakdown
is rendered directly as a label by the Usage page and the daily_with_models export, so the
read path keys a sentinel row on model_group. Two deployments sharing a public name merge
under it, which is the collapse the write path used to do by summing them into one row.
Request rows are untouched. Verified through a four-pod proxy, where the breakdown labels
came back as ptu-a and ptu-renamed rather than UUIDs

Behavior changes

flat_cost and total_flat_cost are added to the shared SpendMetrics and DailySpendMetadata, so every daily activity endpoint returns them, defaulting to 0.0 where no PTU cost exists. This mirrors how compression_saved_tokens and prompt_caching_savings_spend were added to the same models in #33810; the fields are additive and no existing value or status code changes. The per-request api_keys sets used to seed key metadata exclude the sentinel string as well

On the aggregated GROUPING SETS path a team whose only rows for a day are PTU sentinels can now show an unknown provider bucket with flat_cost 0, spend 0 and 0 requests. The sentinel carries no provider, and by the time the provider grouping set is aggregated the api_key column has been grouped away, so a sentinel row cannot be told apart from a genuine row that simply has no provider. Suppressing the bucket by checking for request activity was tried and reverted: rows predating the api_requests column backfill to all zeroes, so that check silently dropped provider buckets the previous build reported. An empty bucket is the lesser cost, and no flat cost is ever credited to a provider on either path. The per-row path skips sentinel rows outright and so shows no such bucket; the two paths differ only in this empty entry

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Open in Devin Review

Note

Medium Risk
Changes billing/analytics response semantics and breakdown rules for sentinel rows; incorrect handling could misreport spend or leak fake API keys in breakdowns on some aggregation paths.

Overview
Daily spend analytics now expose PTU flat cost alongside token spend: SpendMetrics gains flat_cost, and response metadata includes total_flat_cost.

The read path pulls ptu_flat_cost from LiteLLM_DailyTeamSpend (other daily tables return zero so the shape stays consistent). Aggregation maps that into flat_cost on totals and per-date metrics.

Rows keyed with PTU_SENTINEL_API_KEY (__ptu_flat_cost__) still roll their flat cost into parent buckets (e.g. model totals) but are excluded from per-api-key breakdowns, the top-level api_keys map, and provider breakdown; metadata lookups also skip the sentinel string. Dashboard OpenAPI types pick up the new fields.

Reviewed by Cursor Bugbot for commit 6f888e3. Bugbot is set up for automated code reviews on this repo. Configure here.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 2 potential issues.

View 1 additional finding in Devin Review.

Open in Devin Review

Comment thread litellm/proxy/management_endpoints/common_daily_activity.py Outdated
Comment on lines +586 to +587
# Only LiteLLM_DailyTeamSpend carries ptu_flat_cost; other daily tables emit a
# constant zero so the SpendMetrics.flat_cost response shape stays uniform.

@devin-ai-integration devin-ai-integration Bot Jul 31, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Explanatory prose comments added where the repository forbids them

Multi-line narrative comments are added to explain rationale (# A PTU sentinel row keys on the deployment id ... at litellm/proxy/management_endpoints/common_daily_activity.py:221-224) even though the repository's coding guidelines only permit comments that are strictly necessary, tool-directed, or TODO/FIXME, so the code carries prose that must be kept in sync with the logic.
Impact: Future edits can silently leave these explanations stale, which is exactly the maintenance cost the rule exists to prevent.

Where the rule is violated

CLAUDE.md (referenced as mandatory from AGENTS.md) states comments should only exist when "absolutely necessary to explain some very complex business logic (in which case, keep it concise and clear)", as a tool suppression, or as a TODO/FIXME. The PR adds three narrative blocks: litellm/proxy/management_endpoints/common_daily_activity.py:221-224 (why the sentinel keys on model_group), litellm/proxy/management_endpoints/common_daily_activity.py:621-623 (why only the team table sums the column), and litellm/proxy/management_endpoints/common_daily_activity.py:890-895 (why the flat cost is withheld from the provider bucket). The # mutable-ok: pydantic update payload suppression on line 896 is allowed; the surrounding prose is not.

Open in Devin Review

Was this helpful? React with 👍 or 👎 to provide feedback.

@greptile-apps

greptile-apps Bot commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR surfaces PTU flat costs in daily activity responses while keeping sentinel data out of request-oriented breakdowns

  • Adds flat-cost fields to shared spend metrics, response metadata, and generated dashboard types
  • Aggregates team PTU flat cost into daily and grand totals
  • Excludes PTU sentinel keys from API-key metadata and nested key breakdowns
  • Prevents sentinel flat cost from being attributed to provider buckets
  • Adds focused coverage for paginated and grouping-set aggregation behavior

Confidence Score: 5/5

The PR appears safe to merge

No blocking failure remains

Important Files Changed

Filename Overview
litellm/proxy/management_endpoints/common_daily_activity.py Reads and aggregates PTU flat cost while excluding sentinel rows from API-key and provider attribution; the previously reported provider pollution is resolved
litellm/types/proxy/management_endpoints/common_daily_activity.py Adds backward-compatible defaulted flat-cost fields to daily activity response models
tests/test_litellm/proxy/management_endpoints/test_common_daily_activity.py Adds focused tests for flat-cost accumulation, sentinel filtering, provider attribution, and model labels
ui/litellm-dashboard/src/lib/http/schema.d.ts Updates generated dashboard API declarations with the new flat-cost fields

Reviews (18): Last reviewed commit: "feat(ptu): surface PTU flat cost on the ..." | Re-trigger Greptile

@codecov

codecov Bot commented Jul 31, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@yucheng-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 6f888e3. Configure here.

records: Sequence[_GroupingSetsRow],
) -> _AggregatedSpendData:
"""Async wrapper: fetch api_key_metadata, then dispatch on a worker thread."""
api_keys: set[str] = {r.api_key for r in records if r.api_key}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sentinel leaks in grouping-sets path

Medium Severity

This commit wires ptu_flat_cost into the grouping-sets read path and excludes PTU_SENTINEL_API_KEY from key-metadata lookup, but _aggregate_grouping_sets_records_sync still assigns sentinel rows into api_keys, per-model key breakdowns, and provider buckets. Unlike the Python update_breakdown_metrics path, sentinel flat cost can therefore still appear as a fake key and under provider unknown whenever team spend is served through the aggregated query.

Additional Locations (2)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 6f888e3. Configure here.

@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch from 6f888e3 to 62b4fe0 Compare July 31, 2026 20:23
@yucheng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai fixed: the GROUPING SETS dispatcher now excludes the flat-cost sentinel from every api_key breakdown, mirroring the per-row path. Re-review please

@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch 2 times, most recently from 7bfecec to 1fcb0b0 Compare July 31, 2026 20:50
@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch from 1fcb0b0 to ed8b05f Compare July 31, 2026 21:15
@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch from ed8b05f to b93ad00 Compare July 31, 2026 21:18
@yucheng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai rebased onto the updated base; the grouping-sets sentinel guards and read-path coverage are unchanged. Re-review please

@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch from b93ad00 to 2c73e28 Compare July 31, 2026 22:12
@yucheng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai rebased onto the updated base with no content change; the grouping-sets sentinel guards and coverage are unchanged. Re-review please

greptile-apps[bot]

This comment was marked as resolved.

@yucheng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai the grouping-sets provider parent no longer attributes sentinel flat cost to "unknown"; provider buckets stay request-only like the per-row path. Re-review please

@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch from 7c8bf51 to 2d46722 Compare August 1, 2026 00:26
@yucheng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai rebased onto the updated rollup commit; no changes to this PR's own diff. Re-review please

@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch from 2d46722 to cfdbd2d Compare August 1, 2026 00:35
@yucheng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai rebased onto the updated rollup commit; this PR's own diff is unchanged from the 5/5 review. Re-review please

@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch from bdb3bce to e45c82b Compare August 8, 2026 17:02
@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch from e45c82b to 8704e38 Compare August 8, 2026 17:09
@yucheng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai review

@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch from 8704e38 to a082b53 Compare August 8, 2026 18:22
@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch from a082b53 to cfabee0 Compare August 8, 2026 19:16
@yucheng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai review

@yucheng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai review

@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch from 9e29944 to f454d5b Compare August 10, 2026 16:51
Base automatically changed from litellm_ptu_flat_cost_rollup to litellm_internal_staging August 10, 2026 17:20
Aggregate the ptu_flat_cost written by the rollup into SpendMetrics.flat_cost and
DailySpendMetadata.total_flat_cost, so /team/daily/activity returns flat cost
alongside per-request spend. The aggregated SQL path selects ptu_flat_cost only
for LiteLLM_DailyTeamSpend and a constant zero for the other daily tables, keeping
the response shape uniform.

Rows written under the PTU sentinel api_key add their flat cost to every parent
bucket (per-model, per-day, per-team totals) but never appear as an api_key row in
any breakdown, and are excluded from the per-request provider breakdown; the
sentinel string is not a real key alias. Both flat_cost and total_flat_cost default
to zero, so a read of any entity without PTU config is unchanged.

The sentinel row now keys on the deployment id, so the per-model breakdown keys it on
model_group instead. That breakdown key is rendered directly as a label by the Usage page
and the daily_with_models export, and a deployment id there would read as a UUID. Two
deployments sharing a public name merge under it, which is the collapse the write path
used to do by summing them into one row. Request rows are untouched and still key on
model, since their model_group is a routing concept rather than a display name.
@yucheng-berri
yucheng-berri force-pushed the litellm_ptu_read_path branch from f454d5b to 65b0efc Compare August 10, 2026 17:23
@yucheng-berri
yucheng-berri merged commit 457be8f into litellm_internal_staging Aug 10, 2026
81 checks passed
@yucheng-berri
yucheng-berri deleted the litellm_ptu_read_path branch August 10, 2026 17:55
@codspeed-hq

codspeed-hq Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_ptu_read_path (65b0efc) with litellm_internal_staging (f6b9518)1

Open in CodSpeed

Footnotes

  1. No successful run was found on litellm_internal_staging (61218f5) during the generation of this report, so f6b9518 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report.

doonga pushed a commit to greyrock-labs/home-ops that referenced this pull request Aug 24, 2026
…8.0) (#393)

This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.97.0` → `v1.98.0` |

---

### Release Notes

<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>

### [`v1.98.0`](https://github.com/BerriAI/litellm/releases/tag/v1.98.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.98.0...v1.98.0)

##### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.98.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.98.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.98.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

##### What's Changed

- fix(bedrock): drop toolSpec.strict for Claude Sonnet 5 on Converse by [@&#8203;kr0k](https://github.com/kr0k) in [#&#8203;33196](https://github.com/BerriAI/litellm/pull/33196)
- fix(batches): attribute Vertex passthrough batch cost to key/team/tags by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;34456](https://github.com/BerriAI/litellm/pull/34456)
- docs: rewrite the CLAUDE.md comment rule with explicit exceptions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36301](https://github.com/BerriAI/litellm/pull/36301)
- fix(proxy): scope file list pagination cursors to the caller by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36093](https://github.com/BerriAI/litellm/pull/36093)
- fix(proxy): skip prisma-dependent hooks when no database is attached by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36273](https://github.com/BerriAI/litellm/pull/36273)
- fix(proxy): report has\_more false on caller-scoped file list pages by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36326](https://github.com/BerriAI/litellm/pull/36326)
- fix(proxy): restore management\_v1 query-param validation under fastapi>=0.140.7 by [@&#8203;HuanQian571](https://github.com/HuanQian571) in [#&#8203;35773](https://github.com/BerriAI/litellm/pull/35773)
- fix(proxy): stop /{provider}/v1/files from capturing /openai\_passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36092](https://github.com/BerriAI/litellm/pull/36092)
- chore(typing): remove 914 basedpyright Any errors across 16 hotspot files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36386](https://github.com/BerriAI/litellm/pull/36386)
- fix(router): keep batch fallbacks inside the model group that owns the file by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36181](https://github.com/BerriAI/litellm/pull/36181)
- feat(ptu): configure provisioned-throughput flat cost on a model deployment by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35341](https://github.com/BerriAI/litellm/pull/35341)
- docs: clarify the CLAUDE.md comment exceptions are any-of by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36421](https://github.com/BerriAI/litellm/pull/36421)
- docs: replace the Changes PR template section with Caveats by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36423](https://github.com/BerriAI/litellm/pull/36423)
- fix(bedrock): enable native structured output for GLM 5 and DeepSeek V3.2 by [@&#8203;alexshtf](https://github.com/alexshtf) in [#&#8203;35669](https://github.com/BerriAI/litellm/pull/35669)
- feat(ptu): daily rollup writes per-model PTU flat cost by active hour by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35343](https://github.com/BerriAI/litellm/pull/35343)
- feat(logging): add opt-in session\_id and trace\_id correlation to JSON log records via contextvars by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;34418](https://github.com/BerriAI/litellm/pull/34418)
- feat(ptu): surface PTU flat cost on the daily activity read path by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35391](https://github.com/BerriAI/litellm/pull/35391)
- feat(router): add per-deployment allowed\_fails\_policy and cooldown\_time override support by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;34416](https://github.com/BerriAI/litellm/pull/34416)
- feat(ptu): add PTU inputs to the model form and flat cost to the Usage page by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35393](https://github.com/BerriAI/litellm/pull/35393)
- fix(cost): price dict-shaped image input token details at the image rate by [@&#8203;vairodp](https://github.com/vairodp) in [#&#8203;33490](https://github.com/BerriAI/litellm/pull/33490)
- fix(model\_prices): refresh deprecation dates, correct xAI pricing and add missing provider models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36403](https://github.com/BerriAI/litellm/pull/36403)
- feat(ptu): gate PTU flat-cost attribution behind an opt-in env var by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36138](https://github.com/BerriAI/litellm/pull/36138)
- ci: cache Prisma CLI and engine binaries, split test timeout from setup by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36417](https://github.com/BerriAI/litellm/pull/36417)
- feat(rate limiting): configurable estimated output tokens per key, team and model by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36143](https://github.com/BerriAI/litellm/pull/36143)
- fix(ui): hide admin-only Logs tabs from roles that cannot call their endpoints by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36333](https://github.com/BerriAI/litellm/pull/36333)
- test(proxy): guard management\_v1 against fastapi names removed in supported releases by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36336](https://github.com/BerriAI/litellm/pull/36336)
- fix(ui): gate policy and prompt lookups on an admin capability by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36335](https://github.com/BerriAI/litellm/pull/36335)
- build(deps): bump pypdf to 6.15.0 to clear osv-scan by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36350](https://github.com/BerriAI/litellm/pull/36350)
- fix(proxy): isolate guardrail load failures per row by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36432](https://github.com/BerriAI/litellm/pull/36432)
- fix(ui): gate organization and agent usage views behind capabilities by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36334](https://github.com/BerriAI/litellm/pull/36334)
- fix(reset\_budget\_job): atomic budget cascade with chunked reset scans by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36287](https://github.com/BerriAI/litellm/pull/36287)
- feat(proxy): add GET /v1/indexes to list vector store indexes by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36289](https://github.com/BerriAI/litellm/pull/36289)
- feat(ui): show vector store indexes on the Vector Stores page by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36306](https://github.com/BerriAI/litellm/pull/36306)
- fix(proxy): treat SAML as configured in UI SSO detection by [@&#8203;fancybear-dev](https://github.com/fancybear-dev) in [#&#8203;36196](https://github.com/BerriAI/litellm/pull/36196)
- fix(bedrock): reject Anthropic server-side web\_search tool with actionable error by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36473](https://github.com/BerriAI/litellm/pull/36473)
- fix(ui): open the classifier prompt editor above the edit auto-router form by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36438](https://github.com/BerriAI/litellm/pull/36438)
- fix(arize): trace MCP tool calls instead of crashing on CallToolResult by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36453](https://github.com/BerriAI/litellm/pull/36453)
- refactor(ui): make illegal DataTable prop combinations unrepresentable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36470](https://github.com/BerriAI/litellm/pull/36470)
- fix(ui): scope Virtual Keys and Logs team lists to the caller by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36472](https://github.com/BerriAI/litellm/pull/36472)
- fix(ui): gate the Old Usage page behind a proxy-admin capability by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36469](https://github.com/BerriAI/litellm/pull/36469)
- docs(terraform): describe the provider release as automatic by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36467](https://github.com/BerriAI/litellm/pull/36467)
- feat(proxy): add per-deployment keepalive\_seconds SSE heartbeat to prevent load-balancer timeout on long streams by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;34423](https://github.com/BerriAI/litellm/pull/34423)
- fix(router): cool down failed fallback deployments and correct cooldown TTL after Redis backfill by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;35104](https://github.com/BerriAI/litellm/pull/35104)
- perf(spend): write each daily spend batch in one upsert statement by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36448](https://github.com/BerriAI/litellm/pull/36448)
- fix(ui): gate four sidebar pages on the roles their endpoints allow by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36475](https://github.com/BerriAI/litellm/pull/36475)
- fix(ui): restore the Logs Deleted Teams tab for organization admins by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36478](https://github.com/BerriAI/litellm/pull/36478)
- fix(websearch): stop leaking interception control fields to providers by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36480](https://github.com/BerriAI/litellm/pull/36480)
- test(e2e): cover the Anthropic web\_search server tool on Bedrock by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36443](https://github.com/BerriAI/litellm/pull/36443)
- fix(router): warn when a deployment's credentials contradict its provider by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36486](https://github.com/BerriAI/litellm/pull/36486)
- fix: net prompt-caching savings against the cache-write premium by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36452](https://github.com/BerriAI/litellm/pull/36452)
- feat(ui): deployment affinity toggle for the auto-router by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36302](https://github.com/BerriAI/litellm/pull/36302)
- fix(bedrock): use deployment credentials for AWS requests by [@&#8203;daleselaji-dev](https://github.com/daleselaji-dev) in [#&#8203;36160](https://github.com/BerriAI/litellm/pull/36160)
- fix(anthropic): preserve midturn system corrections by [@&#8203;eugene-yao-zocdoc](https://github.com/eugene-yao-zocdoc) in [#&#8203;34290](https://github.com/BerriAI/litellm/pull/34290)
- fix(email): stop duplicate legacy invitation email and fix its onboarding link by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36455](https://github.com/BerriAI/litellm/pull/36455)
- feat(ui): show models under each tier in routing benchmark chart by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36291](https://github.com/BerriAI/litellm/pull/36291)
- fix(proxy): inject streaming usage cost on openai passthrough streams by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36503](https://github.com/BerriAI/litellm/pull/36503)
- docs: require a user flow and live-proxy proof in bug reports by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36498](https://github.com/BerriAI/litellm/pull/36498)
- fix(proxy): add config\_updated\_at audit timestamp for virtual keys by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36488](https://github.com/BerriAI/litellm/pull/36488)
- docs: require a user flow and a stuck-at proof in feature requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36500](https://github.com/BerriAI/litellm/pull/36500)
- feat(router): add required-AND (&) tag prefix and allow\_fail\_open flag by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;36193](https://github.com/BerriAI/litellm/pull/36193)
- feat(proxy): per-key prompt caching toggle via enable\_prompt\_caching by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36466](https://github.com/BerriAI/litellm/pull/36466)
- fix(bedrock): send tool-search beta header for Haiku 4.5 on Invoke /v1/messages by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36502](https://github.com/BerriAI/litellm/pull/36502)
- fix(bedrock): preserve adaptive thinking effort through the /v1/messages bridge by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36507](https://github.com/BerriAI/litellm/pull/36507)
- ci: retry transient network fetch failures in lint workflow by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36563](https://github.com/BerriAI/litellm/pull/36563)
- fix(ui): stub useIsOrgAdmin in UsageTab tests so useCan needs no QueryClient by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36565](https://github.com/BerriAI/litellm/pull/36565)
- fix(alerting): dedupe scheduled Slack spend reports across pods by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36489](https://github.com/BerriAI/litellm/pull/36489)
- chore(typing): clear 1.6k basedpyright Any errors across 56 files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36543](https://github.com/BerriAI/litellm/pull/36543)
- fix(bedrock): add text block to converse user messages carrying documents by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36499](https://github.com/BerriAI/litellm/pull/36499)
- fix(deps): ship boto3 with the base SDK so bedrock works out of the box by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36568](https://github.com/BerriAI/litellm/pull/36568)
- fix(model\_prices): add provider-announced deprecation dates for Bedrock, Mistral, Cohere and Gemini models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36538](https://github.com/BerriAI/litellm/pull/36538)
- chore: bump litellm-enterprise 0.1.54 -> 0.1.55, litellm-proxy-extras 0.4.84 -> 0.4.85, litellm 1.97.0 -> 1.98.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36577](https://github.com/BerriAI/litellm/pull/36577)
- fix(bedrock\_guardrails): skip ApplyGuardrail when there is no content to scan by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36441](https://github.com/BerriAI/litellm/pull/36441)
- fix(e2e): assert on the gen-AI span that served the stream, not the span count by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36582](https://github.com/BerriAI/litellm/pull/36582)
- test(e2e): harden vendor API coverage by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;34557](https://github.com/BerriAI/litellm/pull/34557)
- test(e2e): add reproducers for passthrough and model budget gaps by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;34657](https://github.com/BerriAI/litellm/pull/34657)
- test(e2e): cover google-native generateContent framing and prometheus queue time by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;34650](https://github.com/BerriAI/litellm/pull/34650)
- chore(ci): promote internal staging to main by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36560](https://github.com/BerriAI/litellm/pull/36560)
- feat(router): make routing groups callable as virtual models and list them in /v1/models by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36519](https://github.com/BerriAI/litellm/pull/36519)
- fix(xai): bill web\_search from server\_side\_tool\_usage\_details by [@&#8203;geraint0923](https://github.com/geraint0923) in [#&#8203;30817](https://github.com/BerriAI/litellm/pull/30817)
- fix(responses): init completed\_response on bridge streaming iterator ([#&#8203;35411](https://github.com/BerriAI/litellm/issues/35411)) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35413](https://github.com/BerriAI/litellm/pull/35413)
- fix(batches): attribute Anthropic passthrough batch cost to the creating key, team and tags by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36468](https://github.com/BerriAI/litellm/pull/36468)
- feat(dashscope): add latest Model Studio models to the cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36496](https://github.com/BerriAI/litellm/pull/36496)
- fix(proxy): track streamed passthrough Responses cost by [@&#8203;william-xue](https://github.com/william-xue) in [#&#8203;36529](https://github.com/BerriAI/litellm/pull/36529)
- fix(model\_prices): advertise native structured output on every Bedrock DeepSeek V3.2 and GLM 5 id by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36597](https://github.com/BerriAI/litellm/pull/36597)
- test(bedrock): repoint live Claude tests off the retired Claude 3 Sonnet by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36600](https://github.com/BerriAI/litellm/pull/36600)
- fix(anthropic): preserve speed=fast in usage for /v1/messages and pass-through by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36447](https://github.com/BerriAI/litellm/pull/36447)
- fix(proxy): forward resolved provider and deployment pricing in /cost/estimate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35880](https://github.com/BerriAI/litellm/pull/35880)
- feat(proxy): global SSE keepalive ping interval for OpenAI-shaped streaming routes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36154](https://github.com/BerriAI/litellm/pull/36154)
- fix(responses): preserve Codex namespace tool calls by [@&#8203;dcadenas](https://github.com/dcadenas) in [#&#8203;32536](https://github.com/BerriAI/litellm/pull/32536)
- fix(nvidia\_nim): preserve image passages and stop sending top\_k to /v1/ranking by [@&#8203;atomic](https://github.com/atomic) in [#&#8203;34177](https://github.com/BerriAI/litellm/pull/34177)
- fix: refactor HTTP handler initialization with client support by [@&#8203;Praveen11558](https://github.com/Praveen11558) in [#&#8203;30952](https://github.com/BerriAI/litellm/pull/30952)
- feat(lint): gate writable TypedDict fields with LIT012 by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36590](https://github.com/BerriAI/litellm/pull/36590)
- perf(proxy): stagger scheduled background jobs across jobs and pods by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36589](https://github.com/BerriAI/litellm/pull/36589)
- test: remove four mirror test files that exercise none of their module by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34635](https://github.com/BerriAI/litellm/pull/34635)
- fix(router): stop re-applying router-selecting request tags to the routed tier's deployments by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36628](https://github.com/BerriAI/litellm/pull/36628)
- test: remove tests that never execute by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36681](https://github.com/BerriAI/litellm/pull/36681)
- fix(ui): align spend and budget columns by [@&#8203;daniel-meismer-zocdoc](https://github.com/daniel-meismer-zocdoc) in [#&#8203;35176](https://github.com/BerriAI/litellm/pull/35176)
- test: rename tests that a later definition shadowed by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36685](https://github.com/BerriAI/litellm/pull/36685)
- fix(passthrough): carry the budget reservation into request metadata by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36592](https://github.com/BerriAI/litellm/pull/36592)
- fix(mcp): bound MCP client requests with a session read timeout by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36675](https://github.com/BerriAI/litellm/pull/36675)
- fix(proxy): log requests rejected for an unparsable body in spend logs by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36673](https://github.com/BerriAI/litellm/pull/36673)
- refactor(ui): migrate cost-optimization to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36629](https://github.com/BerriAI/litellm/pull/36629)
- refactor(ui): migrate cost-tracking to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36631](https://github.com/BerriAI/litellm/pull/36631)
- refactor(ui): migrate admin-panel to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36635](https://github.com/BerriAI/litellm/pull/36635)
- refactor(ui): migrate users dashboard to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36642](https://github.com/BerriAI/litellm/pull/36642)
- refactor(ui): migrate prompts to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36643](https://github.com/BerriAI/litellm/pull/36643)
- refactor(ui): migrate team settings to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36641](https://github.com/BerriAI/litellm/pull/36641)
- refactor(ui): migrate models-and-endpoints to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36648](https://github.com/BerriAI/litellm/pull/36648)
- refactor(ui): migrate policy impact popover to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36653](https://github.com/BerriAI/litellm/pull/36653)
- fix(proxy): expand config-defined model access groups when resolving team models for /v2/model/info by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;34211](https://github.com/BerriAI/litellm/pull/34211)
- fix(batches): strip NUL bytes from passthrough batch tags before the managed object write by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36688](https://github.com/BerriAI/litellm/pull/36688)
- test(e2e-ui): verify UI mutations against the API instead of trusting the toast by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36632](https://github.com/BerriAI/litellm/pull/36632)
- fix(proxy): serialize model reconciles so concurrent model writes stop evicting each other by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36687](https://github.com/BerriAI/litellm/pull/36687)
- chore(e2e): port the compat-matrix cron publisher to tests/e2e/claude\_code by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36465](https://github.com/BerriAI/litellm/pull/36465)
- fix(router): never price a strategy-router alias by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36691](https://github.com/BerriAI/litellm/pull/36691)
- feat(model\_prices): add NVIDIA Nemotron 3.5 Lightning on OpenRouter and DeepInfra by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36696](https://github.com/BerriAI/litellm/pull/36696)
- feat(terraform/aws): make VPC, Aurora, and Redis optional by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36676](https://github.com/BerriAI/litellm/pull/36676)
- feat(ui): warn in the Admin UI when no Redis is configured by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36495](https://github.com/BerriAI/litellm/pull/36495)
- fix(ui): show and edit key-level router settings on a virtual key by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36674](https://github.com/BerriAI/litellm/pull/36674)
- fix(router): forward auto-router alias params from the marker entry, not the first same-name deployment by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36626](https://github.com/BerriAI/litellm/pull/36626)
- fix(bedrock\_mantle): 1M context window and long-context pricing for GPT-5.6 Sol/Terra/Luna by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36698](https://github.com/BerriAI/litellm/pull/36698)
- fix(model\_prices): sync the Groq registry with Groq's docs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36664](https://github.com/BerriAI/litellm/pull/36664)
- fix(router): let untagged requests bypass a tagged pre-routing strategy on shared model names by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36627](https://github.com/BerriAI/litellm/pull/36627)
- fix(spend): stop losing spend log rows when a flush is cancelled by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34826](https://github.com/BerriAI/litellm/pull/34826)
- docs(claude): drop the @&#8203; prefix from the PR template path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36726](https://github.com/BerriAI/litellm/pull/36726)
- fix(langfuse): emit otel trace version and release on the keys langfuse v4 reads by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36702](https://github.com/BerriAI/litellm/pull/36702)
- test(interactions): follow Google spec drift replacing Turn with typed steps by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36730](https://github.com/BerriAI/litellm/pull/36730)
- refactor(ui): migrate team detail controls to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36695](https://github.com/BerriAI/litellm/pull/36695)
- refactor(ui): migrate guardrail and duration controls to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36693](https://github.com/BerriAI/litellm/pull/36693)
- refactor(ui): migrate guardrails-monitor, projects, logs to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34606](https://github.com/BerriAI/litellm/pull/34606)
- refactor(ui): migrate search and user controls to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36694](https://github.com/BerriAI/litellm/pull/36694)
- fix(guardrails): scan and re-emit raw Anthropic SSE streams in the bedrock post-call hook by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36598](https://github.com/BerriAI/litellm/pull/36598)
- fix(helm): render nodeSelector on the migrations job by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36747](https://github.com/BerriAI/litellm/pull/36747)
- fix(langfuse): coerce header-sourced mask and trace-update steering values by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36740](https://github.com/BerriAI/litellm/pull/36740)
- refactor(ui): migrate usage tables to shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36707](https://github.com/BerriAI/litellm/pull/36707)
- refactor(ui): migrate guardrails monitor table to shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36709](https://github.com/BerriAI/litellm/pull/36709)
- refactor(ui): migrate guardrails content tables to shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36708](https://github.com/BerriAI/litellm/pull/36708)
- feat(gemini): day-0 pricing for gemini-3.7-flash by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36792](https://github.com/BerriAI/litellm/pull/36792)
- ci: promote staging to main by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36725](https://github.com/BerriAI/litellm/pull/36725)
- build(deps): bump nanoid to 3.3.18 to clear osv-scan by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36787](https://github.com/BerriAI/litellm/pull/36787)
- fix(router): stop scoring system prompt text for code/technical complexity by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36721](https://github.com/BerriAI/litellm/pull/36721)
- feat(complexity\_router): calibrate the classifier rubric with worked examples, selectable per router by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36578](https://github.com/BerriAI/litellm/pull/36578)
- fix(interactions): map step and turn history to Responses API roles and content types by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36733](https://github.com/BerriAI/litellm/pull/36733)
- fix(ui): restore playground model filtering by endpoint by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36130](https://github.com/BerriAI/litellm/pull/36130)
- fix(proxy/batches): stop forwarding custom\_llm\_provider twice in list and cancel by [@&#8203;anxkhn](https://github.com/anxkhn) in [#&#8203;32813](https://github.com/BerriAI/litellm/pull/32813)
- refactor(ui): migrate TokenFlow and JsonViewer to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36735](https://github.com/BerriAI/litellm/pull/36735)
- feat: pre-adoption shadow eval for the auto-router (blind pairwise judge, derived state) by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36587](https://github.com/BerriAI/litellm/pull/36587)
- refactor(ui): migrate SimpleMessageBlock and SimpleToolCallBlock to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36737](https://github.com/BerriAI/litellm/pull/36737)
- refactor(ui): migrate HistoryTree and CollapsibleMessage to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36738](https://github.com/BerriAI/litellm/pull/36738)
- refactor: replace Any with precise types across responses, proxy, and llms modules by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36763](https://github.com/BerriAI/litellm/pull/36763)
- refactor(ui): migrate TruncatedValue and OutputCard to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36739](https://github.com/BerriAI/litellm/pull/36739)
- refactor(ui): migrate SectionHeader and ToolsSection to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36793](https://github.com/BerriAI/litellm/pull/36793)
- feat(ui): migrate playground chat controls to shadcn by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36129](https://github.com/BerriAI/litellm/pull/36129)
- feat(xai): day-0 pricing for grok-4.6 by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36805](https://github.com/BerriAI/litellm/pull/36805)
- feat(ui): highlight Auto Router in the navbar announcement by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36315](https://github.com/BerriAI/litellm/pull/36315)
- test(e2e): assert the model allow-list permits, not only denies by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36823](https://github.com/BerriAI/litellm/pull/36823)
- fix(proxy): tolerate a concurrent creator when creating spend views by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36824](https://github.com/BerriAI/litellm/pull/36824)
- fix(proxy): honor explicit null budget\_duration on team and key create + clearable UI dropdowns by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36699](https://github.com/BerriAI/litellm/pull/36699)
- feat(model\_prices): add meta/muse-spark-1.2 and its contributor tier by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36717](https://github.com/BerriAI/litellm/pull/36717)
- fix(auth): carry team grants in lite login session tokens by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36826](https://github.com/BerriAI/litellm/pull/36826)
- feat(ui): show provider prompt cache tokens in chat response metrics by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36827](https://github.com/BerriAI/litellm/pull/36827)
- fix(auth): stop the team fallback from widening model access by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36837](https://github.com/BerriAI/litellm/pull/36837)
- fix(proxy/team): resolve member\_delete cleanup by user id, not the addressed email by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36839](https://github.com/BerriAI/litellm/pull/36839)
- fix(cli): launch agents as a child process on Windows by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36822](https://github.com/BerriAI/litellm/pull/36822)
- feat(ui): shadow evals tab beside auto-router usage by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36588](https://github.com/BerriAI/litellm/pull/36588)
- feat(cli): make the hidden `lite` command list configurable by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36816](https://github.com/BerriAI/litellm/pull/36816)
- feat(azure\_ai): add Fireworks FW model pricing on Azure AI Foundry by [@&#8203;emerzon](https://github.com/emerzon) in [#&#8203;35613](https://github.com/BerriAI/litellm/pull/35613)
- fix: enable xhigh reasoning support for gpt-5.4-mini models by [@&#8203;emerzon](https://github.com/emerzon) in [#&#8203;26909](https://github.com/BerriAI/litellm/pull/26909)
- feat(azure-ai): add Grok 4.3 model metadata by [@&#8203;emerzon](https://github.com/emerzon) in [#&#8203;27932](https://github.com/BerriAI/litellm/pull/27932)
- feat(ui): render request metrics on the /ui/chat surface by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36845](https://github.com/BerriAI/litellm/pull/36845)
- fix(ui): stop a deselected MCP server keeping its grant on a virtual key by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36840](https://github.com/BerriAI/litellm/pull/36840)
- fix(team): sweep dangling team references and cache on team delete by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36819](https://github.com/BerriAI/litellm/pull/36819)
- fix(mcp): resolve admin OAuth sessions from any worker via DB-backed drafts by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36844](https://github.com/BerriAI/litellm/pull/36844)
- refactor(ui): migrate usage to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36834](https://github.com/BerriAI/litellm/pull/36834)
- refactor(ui): migrate guardrails-monitor to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36838](https://github.com/BerriAI/litellm/pull/36838)
- refactor(ui): migrate playground to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36847](https://github.com/BerriAI/litellm/pull/36847)
- refactor(ui): migrate guardrails to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36832](https://github.com/BerriAI/litellm/pull/36832)
- fix(batches): stop uncostable batches from starving the cost poll page by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36714](https://github.com/BerriAI/litellm/pull/36714)
- perf(spend-logs): bound retention cleanup so one run cannot saturate the database by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36594](https://github.com/BerriAI/litellm/pull/36594)
- fix(proxy): fail config load when a callbacks entry is not dispatchable by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36858](https://github.com/BerriAI/litellm/pull/36858)
- fix(bedrock): hoist custom.defer\_loading before dropping custom on invoke tools by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36855](https://github.com/BerriAI/litellm/pull/36855)
- fix(access groups): sync assigned\_key\_ids from the key write paths by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36843](https://github.com/BerriAI/litellm/pull/36843)
- fix(mcp): expose client HTTP headers to logging callbacks and hooks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36724](https://github.com/BerriAI/litellm/pull/36724)
- fix(ptu): stop per-token billing on a PTU-configured deployment by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36829](https://github.com/BerriAI/litellm/pull/36829)
- fix(ui): add nvidia riva to the model provider list by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36769](https://github.com/BerriAI/litellm/pull/36769)
- fix(scripts): end make check with a ran/skipped summary and verdict by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36864](https://github.com/BerriAI/litellm/pull/36864)
- fix(proxy): track spend for OpenAI passthrough /v1/embeddings by [@&#8203;lostmartian](https://github.com/lostmartian) in [#&#8203;36660](https://github.com/BerriAI/litellm/pull/36660)
- test(proxy): stop monkeypatch.undo re-planting fixture-mocked prisma\_client by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36872](https://github.com/BerriAI/litellm/pull/36872)
- fix(access groups): sync assigned\_team\_ids from the team write paths by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36825](https://github.com/BerriAI/litellm/pull/36825)
- ci: drop the CircleCI ui\_build and ui\_unit\_tests jobs by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36893](https://github.com/BerriAI/litellm/pull/36893)
- fix(langfuse)!: source the emitted metadata blob from StandardLoggingPayload by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36744](https://github.com/BerriAI/litellm/pull/36744)
- refactor(ui): migrate Navbar off antd to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36902](https://github.com/BerriAI/litellm/pull/36902)
- refactor(ui): migrate log details drawer off antd to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36904](https://github.com/BerriAI/litellm/pull/36904)
- refactor(ui): migrate AI Hub off antd and tremor to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36908](https://github.com/BerriAI/litellm/pull/36908)
- refactor(ui): move the shared dropdowns and selectors onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36924](https://github.com/BerriAI/litellm/pull/36924)
- refactor(ui): move the root-level dashboard components onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36927](https://github.com/BerriAI/litellm/pull/36927)
- refactor(ui): move the settings page and bulk user invite onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36936](https://github.com/BerriAI/litellm/pull/36936)
- refactor(ui): move the cost tracking components onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36955](https://github.com/BerriAI/litellm/pull/36955)
- ci: drop the duplicate proxy\_unit\_tests letter-shard workflow by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36866](https://github.com/BerriAI/litellm/pull/36866)
- refactor(ui): migrate shared common\_components off antd and tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36910](https://github.com/BerriAI/litellm/pull/36910)
- refactor(ui): migrate key info and permissions views off antd and tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36913](https://github.com/BerriAI/litellm/pull/36913)
- feat(proxy): serve Anthropic-native /v1/models for Claude Code gateway discovery by [@&#8203;Ar-maan05](https://github.com/Ar-maan05) in [#&#8203;35455](https://github.com/BerriAI/litellm/pull/35455)
- refactor(ui): migrate router settings and shared badges off antd and tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36915](https://github.com/BerriAI/litellm/pull/36915)
- refactor(ui): move the model hub and model select onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36918](https://github.com/BerriAI/litellm/pull/36918)
- fix(ui): keep the cost tracking removal confirmation open until it settles by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36960](https://github.com/BerriAI/litellm/pull/36960)
- refactor(ui): declare DateRangePickerValue locally instead of importing it from tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36962](https://github.com/BerriAI/litellm/pull/36962)
- fix(main): an explicit provider outranks a known OpenAI model name by [@&#8203;FahimaGold](https://github.com/FahimaGold) in [#&#8203;36800](https://github.com/BerriAI/litellm/pull/36800)
- fix(exception\_mapping): bare 429 in an error body no longer outranks the status code by [@&#8203;FahimaGold](https://github.com/FahimaGold) in [#&#8203;36705](https://github.com/BerriAI/litellm/pull/36705)
- refactor(ui): move MCP permission panels onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36964](https://github.com/BerriAI/litellm/pull/36964)
- refactor(ui): migrate ten small dashboard files off antd and tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36966](https://github.com/BerriAI/litellm/pull/36966)
- fix(proxy): force prisma recreate on postgres cached-plan error by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36428](https://github.com/BerriAI/litellm/pull/36428)
- fix(transcription): stop a zero output rate from zeroing transcription cost by [@&#8203;hMED22](https://github.com/hMED22) in [#&#8203;36914](https://github.com/BerriAI/litellm/pull/36914)
- fix(langfuse): restrict trace steering keys to real langfuse trace fields by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36862](https://github.com/BerriAI/litellm/pull/36862)
- Revert "fix(auth): stop the team fallback from widening model access" ([#&#8203;36837](https://github.com/BerriAI/litellm/issues/36837)) by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36982](https://github.com/BerriAI/litellm/pull/36982)
- fix(ui): show zeroed auto-router usage stats when a window has no sessions by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36868](https://github.com/BerriAI/litellm/pull/36868)
- fix(mcp): keep admin-entered oauth endpoints in management reads by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36888](https://github.com/BerriAI/litellm/pull/36888)
- fix(ui): distinguish hosted and local vLLM in the provider dropdown by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36974](https://github.com/BerriAI/litellm/pull/36974)
- fix(openai,azure): return a length-truncated 200 when the output budget fits no token by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36859](https://github.com/BerriAI/litellm/pull/36859)
- fix(proxy): always emit the Anthropic /v1/models token limits, null when unknown by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36961](https://github.com/BerriAI/litellm/pull/36961)
- feat(helm): add startupProbe and hpa.behavior to the componentized chart by [@&#8203;Louis-Vauterin](https://github.com/Louis-Vauterin) in [#&#8203;36382](https://github.com/BerriAI/litellm/pull/36382)
- fix(proxy): serve aggregate MCP endpoint on bare /mcp instead of 307-redirecting by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;34845](https://github.com/BerriAI/litellm/pull/34845)
- feat(shadow\_eval): add reverse-direction shadow eval jobs by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36865](https://github.com/BerriAI/litellm/pull/36865)
- fix(proxy): requeue Redis spend buffer transactions when the DB commit fails by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33881](https://github.com/BerriAI/litellm/pull/33881)
- feat(search): add Nimble as a search provider by [@&#8203;ilchemla](https://github.com/ilchemla) in [#&#8203;36347](https://github.com/BerriAI/litellm/pull/36347)
- fix(mcp): drop caller host and configured upstream headers from logged metadata by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36901](https://github.com/BerriAI/litellm/pull/36901)
- fix(azure\_ai): recognize real Search doc endpoints so teams can read/write via passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36798](https://github.com/BerriAI/litellm/pull/36798)
- fix(anthropic): aggregate 5m/1h cache-write split across iterations path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34860](https://github.com/BerriAI/litellm/pull/34860)
- fix(anthropic cost): apply regional geo uplift to cached tokens by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34850](https://github.com/BerriAI/litellm/pull/34850)
- fix(ui): match the MCP servers count badge to its sibling permission badges by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36984](https://github.com/BerriAI/litellm/pull/36984)
- fix(batches): mark terminal batch with no output file as processed in CheckBatchCost by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35360](https://github.com/BerriAI/litellm/pull/35360)
- fix(caching): cache anthropic /v1/messages responses, including streaming by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34581](https://github.com/BerriAI/litellm/pull/34581)
- fix(anthropic\_messages): make tool\_result images visible to OpenAI-compatible providers by [@&#8203;hMED22](https://github.com/hMED22) in [#&#8203;34462](https://github.com/BerriAI/litellm/pull/34462)
- feat(fireworks\_ai): translate NIM/vLLM extra params to Fireworks-native args by [@&#8203;milesadkins](https://github.com/milesadkins) in [#&#8203;35969](https://github.com/BerriAI/litellm/pull/35969)
- fix(ui): stop the models tab strip from scrolling vertically by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36993](https://github.com/BerriAI/litellm/pull/36993)
- fix(ui): anchor chips-combobox popups to the field instead of the inner input by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36995](https://github.com/BerriAI/litellm/pull/36995)
- feat(proxy): per-component response cost headers by [@&#8203;erensh27](https://github.com/erensh27) in [#&#8203;36965](https://github.com/BerriAI/litellm/pull/36965)
- fix(cost): track OpenAI/Azure web search tool cost per call by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35286](https://github.com/BerriAI/litellm/pull/35286)
- fix(bedrock): resolve aliases in batch file records by [@&#8203;daleselaji-dev](https://github.com/daleselaji-dev) in [#&#8203;36159](https://github.com/BerriAI/litellm/pull/36159)
- fix: report real token usage on guardrail-blocked /v1/responses replies by [@&#8203;guptaishaan](https://github.com/guptaishaan) in [#&#8203;36907](https://github.com/BerriAI/litellm/pull/36907)
- fix(proxy): requeue spend logs when the DB write fails with a transport error by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36716](https://github.com/BerriAI/litellm/pull/36716)
- fix(cost): tiered pricing supports cache creation cost and is all-or-nothing by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36720](https://github.com/BerriAI/litellm/pull/36720)
- fix(vertex\_ai): translate /v1/embeddings batch rows to the Gemini embedding shape by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35092](https://github.com/BerriAI/litellm/pull/35092)
- docs(claude): require ReadOnly on every TypedDict field (LIT012) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37005](https://github.com/BerriAI/litellm/pull/37005)
- refactor(ui): migrate access group create modal to RHF + zod + shadcn by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37033](https://github.com/BerriAI/litellm/pull/37033)
- refactor(ui): re-sync badge and skeleton onto the base-vega shadcn style by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36991](https://github.com/BerriAI/litellm/pull/36991)
- feat(ui): link user detail team names to team pages by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37022](https://github.com/BerriAI/litellm/pull/37022)
- fix(model\_prices): correct DeepSeek V4 max output tokens by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36925](https://github.com/BerriAI/litellm/pull/36925)
- fix(ui): rename models table Status column to Source by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37021](https://github.com/BerriAI/litellm/pull/37021)
- chore: bump litellm-enterprise 0.1.55 -> 0.1.56, litellm-proxy-extras 0.4.85 -> 0.4.86 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37045](https://github.com/BerriAI/litellm/pull/37045)
- feat(proxy): gate the Global Control Plane worker registry on an enterprise license by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36996](https://github.com/BerriAI/litellm/pull/36996)
- fix(model\_prices): add gemini 3.1 flash tts preview and legacy OpenAI shutdown dates by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36788](https://github.com/BerriAI/litellm/pull/36788)
- fix(panw\_prisma\_airs): surface scan\_id on allowed requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37037](https://github.com/BerriAI/litellm/pull/37037)
- fix(model\_map): flag native structured outputs on Anthropic-direct claude-sonnet-5 and claude-haiku-4-5 by [@&#8203;anmolg1997](https://github.com/anmolg1997) in [#&#8203;35930](https://github.com/BerriAI/litellm/pull/35930)
- fix(router): stop get\_router\_model\_info from wiping cached pricing by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36985](https://github.com/BerriAI/litellm/pull/36985)
- fix(redis): unwrap decorated \_\_init\_\_s when deriving the from\_url kwargs allowlist by [@&#8203;anmolg1997](https://github.com/anmolg1997) in [#&#8203;36654](https://github.com/BerriAI/litellm/pull/36654)
- fix(proxy): reserve the larger declared output budget for TPM limits by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37001](https://github.com/BerriAI/litellm/pull/37001)
- fix(databricks): surface provider usage, including prompt-cache counts, in streaming chunks by [@&#8203;pokepoke81](https://github.com/pokepoke81) in [#&#8203;36943](https://github.com/BerriAI/litellm/pull/36943)
- fix(spend): give a batch's cost row a primary key of its own by [@&#8203;marty-sullivan](https://github.com/marty-sullivan) in [#&#8203;36876](https://github.com/BerriAI/litellm/pull/36876)
- feat: shadow eval samples /v1/messages and /v1/responses traffic by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36830](https://github.com/BerriAI/litellm/pull/36830)
- fix(ptu): stop a PTU deployment billing for grounded search by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37043](https://github.com/BerriAI/litellm/pull/37043)
- fix(fireworks\_ai): support router slugs via routers/ prefix by [@&#8203;heathriel](https://github.com/heathriel) in [#&#8203;34257](https://github.com/BerriAI/litellm/pull/34257)
- fix(bedrock): register managed-batch litellm\_params so they stop leaking to the provider (internal copy of [#&#8203;36633](https://github.com/BerriAI/litellm/issues/36633)) by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37048](https://github.com/BerriAI/litellm/pull/37048)
- fix(bedrock): resolve the managed-batch output bucket on every path that reads it by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37047](https://github.com/BerriAI/litellm/pull/37047)
- fix(bedrock): resolve the managed-batch output bucket on every path that reads it by [@&#8203;marty-sullivan](https://github.com/marty-sullivan) in [#&#8203;36634](https://github.com/BerriAI/litellm/pull/36634)
- feat(scripts): queue heavy gates behind a machine-wide slot lock by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36988](https://github.com/BerriAI/litellm/pull/36988)
- feat(mcp): scope gateway session bearers to the RFC 8707 resource by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35045](https://github.com/BerriAI/litellm/pull/35045)
- feat(ui): direction picker and reverse-mode display for shadow evals by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36994](https://github.com/BerriAI/litellm/pull/36994)
- fix(guardrails): return the full PANW AIRS scan response on blocked requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37036](https://github.com/BerriAI/litellm/pull/37036)
- fix(passthrough): stop forwarding client Accept-Encoding upstream by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37058](https://github.com/BerriAI/litellm/pull/37058)
- fix(batches): account a managed batch's cost exactly once by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37050](https://github.com/BerriAI/litellm/pull/37050)
- fix(panw\_prisma\_airs): scan tool call args as plain text, not a tool\_event by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37038](https://github.com/BerriAI/litellm/pull/37038)
- feat(lint): exempt TypedDict-annotated dict literals from LIT002 by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36869](https://github.com/BerriAI/litellm/pull/36869)
- docs(claude): tell agents to let heavy gates queue for machine-wide slots by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37057](https://github.com/BerriAI/litellm/pull/37057)
- test: unstick the suites CircleCI is failing on by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37059](https://github.com/BerriAI/litellm/pull/37059)
- docs(github): proof-of-fix section shows only the latest run as Before/After with nested cases by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37063](https://github.com/BerriAI/litellm/pull/37063)
- test(e2e): assert provider error shape instead of pinned prose by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37065](https://github.com/BerriAI/litellm/pull/37065)
- fix(ui): de-duplicate the reset budget option and polish shadcn surfaces by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37010](https://github.com/BerriAI/litellm/pull/37010)
- chore: rebuild Admin UI bundle from litellm\_internal\_staging by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37066](https://github.com/BerriAI/litellm/pull/37066)
- test(e2e/ui): assert the log drawer chevrons by their lucide classes by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37069](https://github.com/BerriAI/litellm/pull/37069)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37042](https://github.com/BerriAI/litellm/pull/37042)
- fix(ui): keep completion-mode models in the playground chat dropdown (backport to rc/1.98.0) by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37955](https://github.com/BerriAI/litellm/pull/37955)

##### New Contributors

- [@&#8203;kr0k](https://github.com/kr0k) made their first contribution in [#&#8203;33196](https://github.com/BerriAI/litellm/pull/33196)
- [@&#8203;HuanQian571](https://github.com/HuanQian571) made their first contribution in [#&#8203;35773](https://github.com/BerriAI/litellm/pull/35773)
- [@&#8203;alexshtf](https://github.com/alexshtf) made their first contribution in [#&#8203;35669](https://github.com/BerriAI/litellm/pull/35669)
- [@&#8203;vairodp](https://github.com/vairodp) made their first contribution in [#&#8203;33490](https://github.com/BerriAI/litellm/pull/33490)
- [@&#8203;fancybear-dev](https://github.com/fancybear-dev) made their first contribution in [#&#8203;36196](https://github.com/BerriAI/litellm/pull/36196)
- [@&#8203;daleselaji-dev](https://github.com/daleselaji-dev) made their first contribution in [#&#8203;36160](https://github.com/BerriAI/litellm/pull/36160)
- [@&#8203;eugene-yao-zocdoc](https://github.com/eugene-yao-zocdoc) made their first contribution in [#&#8203;34290](https://github.com/BerriAI/litellm/pull/34290)
- [@&#8203;geraint0923](https://github.com/geraint0923) made their first contribution in [#&#8203;30817](https://github.com/BerriAI/litellm/pull/30817)
- [@&#8203;william-xue](https://github.com/william-xue) made their first contribution in [#&#8203;36529](https://github.com/BerriAI/litellm/pull/36529)
- [@&#8203;dcadenas](https://github.com/dcadenas) made their first contribution in [#&#8203;32536](https://github.com/BerriAI/litellm/pull/32536)
- [@&#8203;atomic](https://github.com/atomic) made their first contribution in [#&#8203;34177](https://github.com/BerriAI/litellm/pull/34177)
- [@&#8203;Praveen11558](https://github.com/Praveen11558) made their first contribution in [#&#8203;30952](https://github.com/BerriAI/litellm/pull/30952)
- [@&#8203;anxkhn](https://github.com/anxkhn) made their first contribution in [#&#8203;32813](https://github.com/BerriAI/litellm/pull/32813)
- [@&#8203;lostmartian](https://github.com/lostmartian) made their first contribution in [#&#8203;36660](https://github.com/BerriAI/litellm/pull/36660)
- [@&#8203;FahimaGold](https://github.com/FahimaGold) made their first contribution in [#&#8203;36800](https://github.com/BerriAI/litellm/pull/36800)
- [@&#8203;Louis-Vauterin](https://github.com/Louis-Vauterin) made their first contribution in [#&#8203;36382](https://github.com/BerriAI/litellm/pull/36382)
- [@&#8203;ilchemla](https://github.com/ilchemla) made their first contribution in [#&#8203;36347](https://github.com/BerriAI/litellm/pull/36347)
- [@&#8203;milesadkins](https://github.com/milesadkins) made their first contribution in [#&#8203;35969](https://github.com/BerriAI/litellm/pull/35969)
- [@&#8203;erensh27](https://github.com/erensh27) made their first contribution in [#&#8203;36965](https://github.com/BerriAI/litellm/pull/36965)
- [@&#8203;guptaishaan](https://github.com/guptaishaan) made their first contribution in [#&#8203;36907](https://github.com/BerriAI/litellm/pull/36907)
- [@&#8203;pokepoke81](https://github.com/pokepoke81) made their first contribution in [#&#8203;36943](https://github.com/BerriAI/litellm/pull/36943)
- [@&#8203;heathriel](https://github.com/heathriel) made their first contribution in [#&#8203;34257](https://github.com/BerriAI/litellm/pull/34257)

**Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.97.0...v1.98.0>

### [`v1.98.0`](https://github.com/BerriAI/litellm/releases/tag/v1.98.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.97.0...v1.98.0)

##### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.98.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.98.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.98.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

##### What's Changed

- fix(bedrock): drop toolSpec.strict for Claude Sonnet 5 on Converse by [@&#8203;kr0k](https://github.com/kr0k) in [#&#8203;33196](https://github.com/BerriAI/litellm/pull/33196)
- fix(batches): attribute Vertex passthrough batch cost to key/team/tags by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;34456](https://github.com/BerriAI/litellm/pull/34456)
- docs: rewrite the CLAUDE.md comment rule with explicit exceptions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36301](https://github.com/BerriAI/litellm/pull/36301)
- fix(proxy): scope file list pagination cursors to the caller by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36093](https://github.com/BerriAI/litellm/pull/36093)
- fix(proxy): skip prisma-dependent hooks when no database is attached by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36273](https://github.com/BerriAI/litellm/pull/36273)
- fix(proxy): report has\_more false on caller-scoped file list pages by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36326](https://github.com/BerriAI/litellm/pull/36326)
- fix(proxy): restore management\_v1 query-param validation under fastapi>=0.140.7 by [@&#8203;HuanQian571](https://github.com/HuanQian571) in [#&#8203;35773](https://github.com/BerriAI/litellm/pull/35773)
- fix(proxy): stop /{provider}/v1/files from capturing /openai\_passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36092](https://github.com/BerriAI/litellm/pull/36092)
- chore(typing): remove 914 basedpyright Any errors across 16 hotspot files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36386](https://github.com/BerriAI/litellm/pull/36386)
- fix(router): keep batch fallbacks inside the model group that owns the file by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36181](https://github.com/BerriAI/litellm/pull/36181)
- feat(ptu): configure provisioned-throughput flat cost on a model deployment by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35341](https://github.com/BerriAI/litellm/pull/35341)
- docs: clarify the CLAUDE.md comment exceptions are any-of by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36421](https://github.com/BerriAI/litellm/pull/36421)
- docs: replace the Changes …
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants