Repository navigation
chore(prices): sync Vertex AI prices: 14 models - #40955
Conversation
fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/deepseek-v4-flash-0731: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-vision-exp: fireworks_ai/deepseek-v4-flash-vision-exp: fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/deepseek-v4-pro-0813: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/deepseek-v4p1-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/glm-5p2: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/glm-5p2: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/glm-5p3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/glm-5p3-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/gpt-oss-120b: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/gpt-oss-120b: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/kimi-k2p6: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/kimi-k2p6: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/kimi-k2p7-code: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/kimi-k2p7-code: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/kimi-k3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/kimi-k3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/minimax-m2p7: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/minimax-m2p7: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/minimax-m3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/minimax-m3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/muse-glimmer-30b: fireworks_ai/muse-glimmer-30b: fireworks_ai/accounts/fireworks/models/nemotron-3-ultra-nvfp4: fireworks_ai/nemotron-3-ultra-nvfp4: fireworks_ai/accounts/fireworks/models/qwen3-embedding-8b: fireworks_ai/accounts/fireworks/models/qwen3-reranker-8b: input_cost_per_token fireworks_ai/accounts/fireworks/models/qwen3p7-plus: fireworks_ai/qwen3p7-plus: fireworks_ai/accounts/fireworks/models/qwen3p8-max: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/qwen3p8-max: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/routers/glm-5p2-fast: fireworks_ai/accounts/fireworks/routers/glm-5p3-fast: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost fireworks_ai/accounts/fireworks/routers/kimi-k3-fast: together_ai/arcee-ai/trinity-mini: input_cost_per_token, output_cost_per_token together_ai/arize-ai/qwen-2-1.5b-instruct: babbage-002: input_cost_per_token_batches, output_cost_per_token_batches chat-latest: chatgpt-image-latest: output_cost_per_token, input_cost_per_image_token, output_cost_per_image_token, input_cost_per_token_batches, output_cost_per_token_batches claude-fable-5: claude-fable-5-1: claude-haiku-4-5: claude-mythos-5: claude-mythos-5-1: claude-opus-4-5: claude-opus-4-6: claude-opus-4-7: claude-opus-4-8: claude-opus-5: claude-sonnet-4-5: claude-sonnet-4-6: claude-sonnet-5: davinci-002: input_cost_per_token_batches, output_cost_per_token_batches deep-research-pro-preview-12-2025: cache_read_input_token_cost together_ai/deepseek-ai/deepseek-coder-33b-instruct: input_cost_per_token, output_cost_per_token together_ai/deepseek-ai/DeepSeek-R1-0528:
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Greptile SummaryThis PR synchronizes provider pricing metadata in both copies of LiteLLM’s model-price catalog.
Confidence Score: 5/5The PR appears safe to merge because no outstanding or newly introduced actionable failure remains. The newly added pricing fields match supported metadata keys and existing service-tier accounting behavior, both pricing-map copies remain synchronized, and the only previous root finding was manually resolved without explanation.
|
| Filename | Overview |
|---|---|
| model_prices_and_context_window.json | Updates model pricing fields and source URLs; no new actionable issue was identified in the changes since the previous review. |
| litellm/model_prices_and_context_window_backup.json | Mirrors the primary pricing-map updates and remains byte-for-byte consistent with it. |
Reviews (4): Last reviewed commit: "chore(prices): sync Vertex AI prices: 14..." | Re-trigger Greptile
gemini-2.5-flash: cache_read_input_audio_token_cost gemini-2.5-flash-lite: cache_read_input_audio_token_cost gemini-3-flash-preview: cache_read_input_audio_token_cost gemini-3.1-flash-lite: cache_read_input_audio_token_cost
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…embedding-2 alias Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit e4ebeae. Configure here.
vertex_ai/deep-research-pro-preview-12-2025: cache_read_input_token_cost vertex_ai/gemini-2.5-flash-image: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, input_cost_per_token_priority, output_cost_per_token_batches vertex_ai/gemini-3-flash-preview: input_cost_per_token_flex, output_cost_per_token_flex, input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_flex, cache_read_input_audio_token_cost vertex_ai/gemini-3-pro-image: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_flex, cache_read_input_token_cost_priority, input_cost_per_token_above_200k_tokens, output_cost_per_token_above_200k_tokens, cache_read_input_token_cost_above_200k_tokens, input_cost_per_token_above_200k_tokens_priority, output_cost_per_token_above_200k_tokens_priority, cache_read_input_token_cost_above_200k_tokens_priority vertex_ai/gemini-3.1-flash-image: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost, input_cost_per_token_batches, output_cost_per_token_batches, cache_read_input_token_cost_flex vertex_ai/gemini-3.1-flash-lite: cache_read_input_audio_token_cost vertex_ai/gemini-3.1-flash-lite-image: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex vertex_ai/gemini-3.1-pro-preview: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex vertex_ai/gemini-3.5-flash: vertex_ai/gemini-3.5-flash-lite: cache_read_input_token_cost_priority vertex_ai/gemini-3.6-flash: vertex_ai/gemini-3.7-flash: vertex_ai/gemini-3.8-flash: vertex_ai/gemini-embedding-2:
|
Price check for head c3f8c07: 48 models changed, 9 mismatch against models.dev or OpenRouter, mostly OpenRouter router-level prices. Details in #litellm-provider-info-sync. |
…eads at the published 5.4e-08 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
Price check for head b25b6eb: 48 models changed, 17 mismatch against models.dev or OpenRouter, mostly missing context limits and OpenRouter router-level prices. Details in #litellm-provider-info-sync. |
…103.0) (#290)
This PR contains the following updates:
| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.102.1` → `v1.103.0` |
---
### Release Notes
<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>
### [`v1.103.0`](https://github.com/BerriAI/litellm/releases/tag/v1.103.0)
[Compare Source](https://github.com/BerriAI/litellm/compare/v1.102.1...v1.103.0)
#### Verify Docker Image Signature
All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).
**Verify using the pinned commit hash (recommended):**
A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
ghcr.io/berriai/litellm:v1.103.0
```
**Verify using the release tag (convenience):**
Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/v1.103.0/cosign.pub \
ghcr.io/berriai/litellm:v1.103.0
```
Expected output:
```
The following checks were performed on each of these signatures:
- The cosign claims were validated
- The signatures were verified against the specified public key
```
***
#### What's Changed
- fix(responses): translate the reasoning object into a chat-completion reasoning effort by [@​joshgarnett](https://github.com/joshgarnett) in [#​36363](https://github.com/BerriAI/litellm/pull/36363)
- fix(proxy): bound tool and guardrail index create\_many by the spend-log statement budgets by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40561](https://github.com/BerriAI/litellm/pull/40561)
- fix(mcp): require admission for delegated OAuth by [@​joshua-berri](https://github.com/joshua-berri) in [#​40923](https://github.com/BerriAI/litellm/pull/40923)
- fix(logging): log one bounded summary for a burst of timed-out LoggingWorker callbacks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40912](https://github.com/BerriAI/litellm/pull/40912)
- fix(fireworks): resolve short model names to long cost map keys by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40929](https://github.com/BerriAI/litellm/pull/40929)
- ci: remove main branch source guard by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​40172](https://github.com/BerriAI/litellm/pull/40172)
- chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​40942](https://github.com/BerriAI/litellm/pull/40942)
- fix(spend\_logs): store litellm\_call\_id and match it in request\_id lookups by [@​mateo-berri](https://github.com/mateo-berri) in [#​39068](https://github.com/BerriAI/litellm/pull/39068)
- fix(auth): refresh lite login session token grants from the live user and team rows by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​40657](https://github.com/BerriAI/litellm/pull/40657)
- feat(bedrock): support file delete and list for S3-backed managed files by [@​mateo-berri](https://github.com/mateo-berri) in [#​39836](https://github.com/BerriAI/litellm/pull/39836)
- fix(proxy): gate the webhook test alert on proxy admins by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​40814](https://github.com/BerriAI/litellm/pull/40814)
- fix(ui): hide admin write-form tabs on the models page from view-only admins by [@​mateo-berri](https://github.com/mateo-berri) in [#​38867](https://github.com/BerriAI/litellm/pull/38867)
- fix(anthropic-adapter): surface mid-stream provider errors as Anthropic error events by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33352](https://github.com/BerriAI/litellm/pull/33352)
- docs(github): add an Affected release section to the PR template by [@​mateo-berri](https://github.com/mateo-berri) in [#​40618](https://github.com/BerriAI/litellm/pull/40618)
- docs(e2e): ban unit tests under tests/e2e by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33852](https://github.com/BerriAI/litellm/pull/33852)
- fix(router): preserve Azure Entra ID params in reusable credentials by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40889](https://github.com/BerriAI/litellm/pull/40889)
- docs(user endpoints): remove unsupported soft\_budget param from user docstrings by [@​shivamrawat1](https://github.com/shivamrawat1) in [#​36585](https://github.com/BerriAI/litellm/pull/36585)
- feat(friendli): auto-sync Friendli model metadata into price registry by [@​Lee-Si-Yoon](https://github.com/Lee-Si-Yoon) in [#​35918](https://github.com/BerriAI/litellm/pull/35918)
- build(deps): bump smol-toml to 1.8.0 to clear GHSA-7w5x-hrqm-74c2 in osv-scan by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40478](https://github.com/BerriAI/litellm/pull/40478)
- fix(bedrock\_mantle): price GovCloud regions from the regional cost row and accept region-prefixed model names by [@​mateo-berri](https://github.com/mateo-berri) in [#​39846](https://github.com/BerriAI/litellm/pull/39846)
- chore(ci): remerge internal staging by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​40943](https://github.com/BerriAI/litellm/pull/40943)
- test(auth): freeze the cache clock in auth prefetch tests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40996](https://github.com/BerriAI/litellm/pull/40996)
- feat(jwt): allow virtual\_key\_claim\_field per issuer by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40927](https://github.com/BerriAI/litellm/pull/40927)
- fix(cost): bill cached realtime audio tokens at the audio cache-read rate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40627](https://github.com/BerriAI/litellm/pull/40627)
- perf(logging): skip correlation contextvar stamping when request\_correlation\_in\_logs is off by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41054](https://github.com/BerriAI/litellm/pull/41054)
- feat(pricing): add azure gpt-chat-latest rates and drop retired friendliai llama-3.1 entries by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40976](https://github.com/BerriAI/litellm/pull/40976)
- build(deps): re-suppress GHSA-h7x2-h6g9-p789 in osv-scan on main, mlflow still has no fixed release by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41104](https://github.com/BerriAI/litellm/pull/41104)
- fix(otel): cap per-index OpenInference message attributes span-wide by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40562](https://github.com/BerriAI/litellm/pull/40562)
- feat(proxy): add general\_settings.allowed\_file\_extensions for /v1/files uploads by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41106](https://github.com/BerriAI/litellm/pull/41106)
- fix(proxy): forward provider request id headers on mapped error responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40925](https://github.com/BerriAI/litellm/pull/40925)
- fix(router): name the all-deployments-in-cooldown error on 429 responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40995](https://github.com/BerriAI/litellm/pull/40995)
- fix(ui): show the team alias on the model info page and in its raw JSON by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40992](https://github.com/BerriAI/litellm/pull/40992)
- feat(proxy): honor LITELLM\_DISABLE\_ACCESS\_LOG\_PATHS to drop noisy uvicorn access log lines by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41096](https://github.com/BerriAI/litellm/pull/41096)
- fix(prometheus): label pre-call rate limit failures with the resolved api\_provider by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41059](https://github.com/BerriAI/litellm/pull/41059)
- perf(proxy): serialize /model/info listing once with orjson by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41114](https://github.com/BerriAI/litellm/pull/41114)
- fix(utils): stop wrapper\_async submitting the sync success handler twice by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41115](https://github.com/BerriAI/litellm/pull/41115)
- fix(redis): log a timeout streak once per interval instead of one line per cache call by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40817](https://github.com/BerriAI/litellm/pull/40817)
- fix(router): record flat retry attempts and cap retries from attempted\_retries by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40930](https://github.com/BerriAI/litellm/pull/40930)
- refactor(prometheus): source PROXY\_LLM\_PROVIDER\_FALLBACK from litellm.constants by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41118](https://github.com/BerriAI/litellm/pull/41118)
- fix(proxy): hide default credentials login hint when UI\_PASSWORD is set by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41107](https://github.com/BerriAI/litellm/pull/41107)
- fix(cli): show routed models and session stats for LLM API keys by [@​tin-berri](https://github.com/tin-berri) in [#​41116](https://github.com/BerriAI/litellm/pull/41116)
- fix(bedrock/realtime): propagate deferred Nova Sonic stream failures to the router by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41064](https://github.com/BerriAI/litellm/pull/41064)
- fix(proxy): keep org admins' own team memberships in other orgs visible on team list by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41086](https://github.com/BerriAI/litellm/pull/41086)
- feat(model\_info): provider-scoped fill\_missing\_for\_providers backfill from fallback generalization rules by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41093](https://github.com/BerriAI/litellm/pull/41093)
- fix(auth): load team membership once per request and skip prisma on an L1 hit by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41102](https://github.com/BerriAI/litellm/pull/41102)
- refactor(harness): expand independent trace coverage by [@​yujonglee-berri](https://github.com/yujonglee-berri) in [#​41120](https://github.com/BerriAI/litellm/pull/41120)
- fix(proxy): release max\_parallel\_requests slot when a realtime session ends without LLM callbacks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41113](https://github.com/BerriAI/litellm/pull/41113)
- fix(router): cool down team deployments on 429 when a sibling serves the same public model by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40991](https://github.com/BerriAI/litellm/pull/40991)
- fix(ui): move tags typed into key metadata JSON into the Tags field by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41023](https://github.com/BerriAI/litellm/pull/41023)
- fix(ui): let team admins grant a team all proxy models by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​40196](https://github.com/BerriAI/litellm/pull/40196)
- chore(lint): graduate 12 rules from the strict-gate ratchet by [@​HUAHAODIA](https://github.com/HUAHAODIA) in [#​41048](https://github.com/BerriAI/litellm/pull/41048)
- test: add dedicated CircleCI integration contract foundation by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41066](https://github.com/BerriAI/litellm/pull/41066)
- fix(utils): keep litellm params out of provider request bodies by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41018](https://github.com/BerriAI/litellm/pull/41018)
- fix(openai): keep extra\_headers out of the chat request body on the httpx handler path by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41141](https://github.com/BerriAI/litellm/pull/41141)
- test: cover persisted updates and warmed authorization policies by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41070](https://github.com/BerriAI/litellm/pull/41070)
- fix(proxy): resolve x-litellm-call-id from response metadata when routes omit call\_id by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41056](https://github.com/BerriAI/litellm/pull/41056)
- chore(prices): sync Vertex AI prices: 14 models by [@​berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#​40955](https://github.com/BerriAI/litellm/pull/40955)
- ci(codeql): exclude noisy Python quality queries by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41142](https://github.com/BerriAI/litellm/pull/41142)
- feat(proxy): predict prompt-cache costs across deployments by [@​tin-berri](https://github.com/tin-berri) in [#​40877](https://github.com/BerriAI/litellm/pull/40877)
- fix(prompt\_security): keep polling file sanitization through non-terminal statuses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41131](https://github.com/BerriAI/litellm/pull/41131)
- fix(health): skip background health check DB writes when the latest-row read fails by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41145](https://github.com/BerriAI/litellm/pull/41145)
- feat(model\_armor): logging\_only mode scans completed streams after delivery by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40702](https://github.com/BerriAI/litellm/pull/40702)
- fix(bedrock guardrails): derive contextual grounding source and query from plain messages by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41132](https://github.com/BerriAI/litellm/pull/41132)
- fix(cli): drop enum.StrEnum so the CLI imports on Python 3.10 by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41046](https://github.com/BerriAI/litellm/pull/41046)
- fix(responses): route mid-stream error events through exception\_type so content\_policy\_fallbacks fire by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40988](https://github.com/BerriAI/litellm/pull/40988)
- fix(cost): bill gemini-embedding-2 per token and stop double charging audio by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41157](https://github.com/BerriAI/litellm/pull/41157)
- test: bind management E2E callers and isolate JWT actors by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​40892](https://github.com/BerriAI/litellm/pull/40892)
- fix(headroom): protect cache\_control-marked rows anywhere in history by [@​rad-p44](https://github.com/rad-p44) in [#​40315](https://github.com/BerriAI/litellm/pull/40315)
- test: add strict stateless provider replay identity by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41149](https://github.com/BerriAI/litellm/pull/41149)
- fix(ci): test checked-out model pricing in unit jobs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41181](https://github.com/BerriAI/litellm/pull/41181)
- fix(guardrails): write per-message guardrail rewrites back onto Responses input items by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40939](https://github.com/BerriAI/litellm/pull/40939)
- fix(proxy): log the provider usage on deferred /v1/messages calls and price cache writes without a creation rate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41172](https://github.com/BerriAI/litellm/pull/41172)
- fix(guardrails): record not\_run evaluation when scoping leaves nothing to scan by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​39050](https://github.com/BerriAI/litellm/pull/39050)
- fix(responses): preserve provider affinity by [@​AaronHowell](https://github.com/AaronHowell) in [#​40228](https://github.com/BerriAI/litellm/pull/40228)
- fix(sdk): keep body and proxy headers on BadRequestError mapped from a litellm\_proxy 400 by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40994](https://github.com/BerriAI/litellm/pull/40994)
- fix(responses): hoist Codex additional\_tools input items into the chat bridge tools by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40989](https://github.com/BerriAI/litellm/pull/40989)
- fix(router): honor team and key provider weights by [@​tin-berri](https://github.com/tin-berri) in [#​41072](https://github.com/BerriAI/litellm/pull/41072)
- test(e2e): verify streamed answers and tool continuation by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41194](https://github.com/BerriAI/litellm/pull/41194)
- fix(cli): label router costs and simplify the routed-model header by [@​tin-berri](https://github.com/tin-berri) in [#​41186](https://github.com/BerriAI/litellm/pull/41186)
- test(spend): reconcile concurrent requests and daily activity by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41188](https://github.com/BerriAI/litellm/pull/41188)
- fix(guardrails): scan the Anthropic top-level system prompt and tool\_use arguments by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40984](https://github.com/BerriAI/litellm/pull/40984)
- fix(router): count num\_retries\_per\_request across fallback hops by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41191](https://github.com/BerriAI/litellm/pull/41191)
- fix(bedrock): grant rerank, retrieve, agent, and agentcore actions in the web identity session policy by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41168](https://github.com/BerriAI/litellm/pull/41168)
- fix(vertex-live): bill Gemini Live sessions end to end (internal copy of [#​37075](https://github.com/BerriAI/litellm/issues/37075)) by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40915](https://github.com/BerriAI/litellm/pull/40915)
- fix(health): resolve litellm\_credential\_name in realtime health checks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41173](https://github.com/BerriAI/litellm/pull/41173)
- feat(proxy): unified custom\_key\_policy hook for key generate, update and regenerate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40921](https://github.com/BerriAI/litellm/pull/40921)
- fix(proxy): enforce custom\_key\_update policy on /key/regenerate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40695](https://github.com/BerriAI/litellm/pull/40695)
- fix(router): preserve session model choice within each complexity tier by [@​tin-berri](https://github.com/tin-berri) in [#​41174](https://github.com/BerriAI/litellm/pull/41174)
- test(pricing): let synced GovCloud Bedrock rows cite the AWS price list by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41263](https://github.com/BerriAI/litellm/pull/41263)
- docs(github): ask for interactive coding-tool proof in the PR template by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41257](https://github.com/BerriAI/litellm/pull/41257)
- feat(proxy): add POST /management/v1/users/bulk for batched user and team membership creation by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41028](https://github.com/BerriAI/litellm/pull/41028)
- fix(credentials): answer 409 on a credential name collision, make Terraform adoption opt-in by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​40917](https://github.com/BerriAI/litellm/pull/40917)
- feat(proxy): add POST /management/v1/users/bulk\_delete and POST /management/v1/teams/{team\_id}/members/bulk\_delete by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41039](https://github.com/BerriAI/litellm/pull/41039)
- fix(proxy): list directly assigned team models in model access errors by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41256](https://github.com/BerriAI/litellm/pull/41256)
- feat(auto-router): allow opted-in team members to manage their routers by [@​tin-berri](https://github.com/tin-berri) in [#​41175](https://github.com/BerriAI/litellm/pull/41175)
- build(rust-bridge): add typed \_native stub and validate it with mypy.stubtest by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41180](https://github.com/BerriAI/litellm/pull/41180)
- feat(guardrails): add new upstream presidio pii entities including german set by [@​MvdB](https://github.com/MvdB) in [#​36775](https://github.com/BerriAI/litellm/pull/36775)
- fix(responses): filter bridged kwargs like the native Responses path by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41144](https://github.com/BerriAI/litellm/pull/41144)
- test(e2e): cover the reliability retry, cooldown, fallback, and routing-strategy cells by [@​mateo-berri](https://github.com/mateo-berri) in [#​39857](https://github.com/BerriAI/litellm/pull/39857)
- fix(anthropic): add the per-turn-control beta when a message carries output\_config by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41189](https://github.com/BerriAI/litellm/pull/41189)
- fix(router): bind per-request routing\_strategy override selectors to the request's callbacks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41178](https://github.com/BerriAI/litellm/pull/41178)
- feat(proxy): bind JWT claims to registered agents via agent\_id\_jwt\_field by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40904](https://github.com/BerriAI/litellm/pull/40904)
- fix(proxy): enforce organization budgets when max\_budget is 0 by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​41271](https://github.com/BerriAI/litellm/pull/41271)
- fix(alerting): send llm\_exceptions Slack alert for 5xx HTTPException and ProxyException by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41125](https://github.com/BerriAI/litellm/pull/41125)
- fix(headroom): protect the cached prefix through the last cache\_control breakpoint by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41161](https://github.com/BerriAI/litellm/pull/41161)
- fix(utils): cache custom HuggingFace tokenizers across /utils/token\_counter requests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41216](https://github.com/BerriAI/litellm/pull/41216)
- fix(router): keep weighted routing when a deployment id equals a model\_name by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41156](https://github.com/BerriAI/litellm/pull/41156)
- feat(router): add capability classifier as Fuse foundation by [@​tin-berri](https://github.com/tin-berri) in [#​41270](https://github.com/BerriAI/litellm/pull/41270)
- fix(proxy): keep access-group raw SQL writes on the writer while writer\_unavailable is stale by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41283](https://github.com/BerriAI/litellm/pull/41283)
- fix(prometheus): count 401 auth failures in litellm\_proxy\_failed\_requests\_metric by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41170](https://github.com/BerriAI/litellm/pull/41170)
- test: drop tests that pin vendor facts and add the CLAUDE.md rule by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41269](https://github.com/BerriAI/litellm/pull/41269)
- fix(proxy): run the remaining inline token counts off the event loop by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40262](https://github.com/BerriAI/litellm/pull/40262)
- fix(proxy): log blocked streaming guardrail responses as failures, not success by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40191](https://github.com/BerriAI/litellm/pull/40191)
- feat(proxy): add tpd\_limit (tokens per day) for batch submissions by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40997](https://github.com/BerriAI/litellm/pull/40997)
- fix(proxy): reconcile budget reservation before enqueuing spend to the DB by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40310](https://github.com/BerriAI/litellm/pull/40310)
- fix(xai): stop sending web\_search\_options to xAI's retired Live Search path by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​38278](https://github.com/BerriAI/litellm/pull/38278)
- feat(terraform): add tpm\_limit, rpm\_limit, budget\_duration, allowed\_models to litellm\_team\_member\_add by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​38682](https://github.com/BerriAI/litellm/pull/38682)
- fix(rerank): bill Vertex search\_units from input records and give every rerank response a unique id by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35180](https://github.com/BerriAI/litellm/pull/35180)
- fix(router): stop counting caller-set timeout 408s toward deployment cooldown by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41230](https://github.com/BerriAI/litellm/pull/41230)
- feat(router): add Fuse V2 classifier after capability forecasting by [@​tin-berri](https://github.com/tin-berri) in [#​41272](https://github.com/BerriAI/litellm/pull/41272)
- fix(proxy): keep client User-Agent on auth failure spend logs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41291](https://github.com/BerriAI/litellm/pull/41291)
- fix(proxy): reset budgets by decrementing pre-reset spend instead of zeroing rows by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41279](https://github.com/BerriAI/litellm/pull/41279)
- fix(xai): honor nested web\_search filters on the xAI Responses API by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​38268](https://github.com/BerriAI/litellm/pull/38268)
- fix(router): stop registering a caller-supplied credential as a router deployment by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​41289](https://github.com/BerriAI/litellm/pull/41289)
- fix(router): accept custom\_provider\_map providers before the first completion call by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41300](https://github.com/BerriAI/litellm/pull/41300)
- fix(proxy): return 400 instead of 500 for lone surrogate escapes in request body by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41297](https://github.com/BerriAI/litellm/pull/41297)
- fix(langsmith): keep events appended during an in-flight flush instead of clearing them by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41288](https://github.com/BerriAI/litellm/pull/41288)
- fix(logging): track spend for streams a deployment hook converted to non-streaming by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41171](https://github.com/BerriAI/litellm/pull/41171)
- fix(bedrock): sanitize client tool\_call ids to Bedrock toolUseId constraints by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40872](https://github.com/BerriAI/litellm/pull/40872)
- fix(passthrough): attribute Vertex passthrough successes to the resolved router deployment by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41307](https://github.com/BerriAI/litellm/pull/41307)
- feat(ui): persist Models table search, filters, sort and page in the URL by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41296](https://github.com/BerriAI/litellm/pull/41296)
- fix(jwt-auth): scope JWT key mappings by issuer to prevent cross-issuer collisions by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​41281](https://github.com/BerriAI/litellm/pull/41281)
- feat(openai): add openai\_system\_messages\_first to put system messages first for prompt caching by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41304](https://github.com/BerriAI/litellm/pull/41304)
- feat(ui): add custom request headers to the API Playground by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41309](https://github.com/BerriAI/litellm/pull/41309)
- feat(cli): sync Codex /model picker from proxy /v1/models in lite codex by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40476](https://github.com/BerriAI/litellm/pull/40476)
- chore: bump litellm-enterprise 0.1.67 -> 0.1.68, litellm-proxy-extras 0.4.97 -> 0.4.98, litellm 1.102.0 -> 1.103.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41321](https://github.com/BerriAI/litellm/pull/41321)
- feat: add aihubmix provider pricing entries by [@​IToSSc](https://github.com/IToSSc) in [#​41179](https://github.com/BerriAI/litellm/pull/41179)
- feat(auto-router): add per-model Fast mode toggle by [@​tin-berri](https://github.com/tin-berri) in [#​41282](https://github.com/BerriAI/litellm/pull/41282)
- fix(proxy): include litellm\_call\_id in LLM API exception logs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41205](https://github.com/BerriAI/litellm/pull/41205)
- fix(proxy): keep yaml pass-through endpoints visible to auth after db overlay by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41303](https://github.com/BerriAI/litellm/pull/41303)
- fix(proxy): resolve router\_settings.model\_group\_alias before key/team model auth by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41308](https://github.com/BerriAI/litellm/pull/41308)
- fix(ui): block usage export and flag the range when a spend page fails by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41294](https://github.com/BerriAI/litellm/pull/41294)
- fix(vertex\_ai): bill Gemini Omni Interactions usage and Veo sampleCount on passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41322](https://github.com/BerriAI/litellm/pull/41322)
- fix(proxy): honor LITELLM\_LOG for uvicorn and proxy extras loggers by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41306](https://github.com/BerriAI/litellm/pull/41306)
- fix(cost): price native Responses WebSocket turns at their returned service\_tier by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41318](https://github.com/BerriAI/litellm/pull/41318)
- fix(proxy): key model rpm/tpm override takes precedence over team model limit by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41302](https://github.com/BerriAI/litellm/pull/41302)
- fix(proxy): track per-member organization spend by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41255](https://github.com/BerriAI/litellm/pull/41255)
- feat(proxy): add /nvidia\_nim passthrough route for NIM object detection and OCR /v1/infer by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41316](https://github.com/BerriAI/litellm/pull/41316)
- feat(model\_info): add provider-neutral Gemini 2.5+ chat baseline fallback generalization by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41320](https://github.com/BerriAI/litellm/pull/41320)
- fix(spend): sum multi-round session duration in logs UI by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35388](https://github.com/BerriAI/litellm/pull/35388)
- feat(router): limit unlicensed Capability and Fuse v2 routers to one each by [@​tin-berri](https://github.com/tin-berri) in [#​41326](https://github.com/BerriAI/litellm/pull/41326)
- fix(guardrails): resolve caller identity from metadata buckets in custom code guardrail by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41126](https://github.com/BerriAI/litellm/pull/41126)
- fix(e2e): onboard dashboard users through invitations by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41319](https://github.com/BerriAI/litellm/pull/41319)
- feat(guardrails): add Microsoft Agent 365 MCP tool-call guardrail by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​38241](https://github.com/BerriAI/litellm/pull/38241)
- test: drop remaining tests that pin cost-map vendor facts by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41298](https://github.com/BerriAI/litellm/pull/41298)
- feat(ui): show average response time per model in usage model activity by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41313](https://github.com/BerriAI/litellm/pull/41313)
- fix(proxy): preserve Anthropic pricing modifiers in router savings by [@​tin-berri](https://github.com/tin-berri) in [#​41341](https://github.com/BerriAI/litellm/pull/41341)
- feat(guardrails): support pre\_call and during\_call modes for llm\_as\_a\_judge by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41128](https://github.com/BerriAI/litellm/pull/41128)
- fix(gemini): propagate the provider's modelVersion to the response model by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41338](https://github.com/BerriAI/litellm/pull/41338)
- fix(fireworks-ai): bill cache-write, reasoning and audio tokens via the shared cost calculator by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41339](https://github.com/BerriAI/litellm/pull/41339)
- feat(guardrails): singulr v2 API contract with logging\_only, pre\_mcp\_call and post\_mcp\_call by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​41329](https://github.com/BerriAI/litellm/pull/41329)
- ci(image-scan): ignore zlib CVE-2026-85091 until Wolfi ships the fix by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41353](https://github.com/BerriAI/litellm/pull/41353)
- feat(e2e): reuse exact provider responses for 24 hours by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41346](https://github.com/BerriAI/litellm/pull/41346)
- fix(xai): keep 'instructions' on the xAI Responses API so system messages survive web search by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​38254](https://github.com/BerriAI/litellm/pull/38254)
- feat(ui): configure capability and Fuse v2 classifiers by [@​tin-berri](https://github.com/tin-berri) in [#​41315](https://github.com/BerriAI/litellm/pull/41315)
- fix(anthropic): tolerate message\_delta events without usage when streaming by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41336](https://github.com/BerriAI/litellm/pull/41336)
- test(router): ignore deployment-selection logs in the fallback log assertion by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41358](https://github.com/BerriAI/litellm/pull/41358)
- test(proxy): assert budget resets decrement the cleared spend by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41359](https://github.com/BerriAI/litellm/pull/41359)
- fix(e2e): expect models filters to persist after reload by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41348](https://github.com/BerriAI/litellm/pull/41348)
- fix(e2e): record cookie-setting provider responses and keep prompt-caching tests live by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41366](https://github.com/BerriAI/litellm/pull/41366)
- fix(responses): recount tokens when a streamed response completes without usage by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41337](https://github.com/BerriAI/litellm/pull/41337)
- fix(ui): simplify Capability and Fuse advanced routing options by [@​tin-berri](https://github.com/tin-berri) in [#​41371](https://github.com/BerriAI/litellm/pull/41371)
- fix(mcp): authorize JWT OAuth credential persistence by [@​joshua-berri](https://github.com/joshua-berri) in [#​41314](https://github.com/BerriAI/litellm/pull/41314)
- feat(router): stream shadow traffic and fan out silent\_model to multiple targets by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41368](https://github.com/BerriAI/litellm/pull/41368)
- perf(content\_filter): scan a bounded window per streamed chunk by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41407](https://github.com/BerriAI/litellm/pull/41407)
- fix(proxy): hide model allowlist from client-facing model access denied errors by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41310](https://github.com/BerriAI/litellm/pull/41310)
- feat(http): opt-in outbound HTTP/2 for httpx clients by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41268](https://github.com/BerriAI/litellm/pull/41268)
- refactor(rust): remove gateway, config, router, realtime, and Rust trace-parity instrumentation by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41432](https://github.com/BerriAI/litellm/pull/41432)
- fix(guardrails): don't add post\_call output scan for MCP-only Presidio modes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40571](https://github.com/BerriAI/litellm/pull/40571)
- chore(prices): sync Azure, Azure AI, Gemini, OpenAI, Bedrock, Together AI, Fireworks and Vertex prices: 278 models, 59 new, 30 deprecated by [@​berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#​41154](https://github.com/BerriAI/litellm/pull/41154)
- fix(rag): forward retrieval\_filter from retrieval\_config to vector store search by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34427](https://github.com/BerriAI/litellm/pull/34427)
- refactor(rust): extract auth and cache crates by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41464](https://github.com/BerriAI/litellm/pull/41464)
- fix(proxy): default litellm\_trace\_id to the OTel server span trace id by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41386](https://github.com/BerriAI/litellm/pull/41386)
- chore(codeowners): add ryan and kerry as owners of the cost map by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41333](https://github.com/BerriAI/litellm/pull/41333)
- fix(responses): guard empty-choices chunks in the Responses API streaming bridge by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34455](https://github.com/BerriAI/litellm/pull/34455)
- chore(prices): sync Google Gemini prices: 22 models by [@​berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#​41457](https://github.com/BerriAI/litellm/pull/41457)
- fix(bedrock): forward userContext in Knowledge Base Retrieve requests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41475](https://github.com/BerriAI/litellm/pull/41475)
- ci(rust): split rust jobs, use nextest and Swatinem/rust-cache by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41480](https://github.com/BerriAI/litellm/pull/41480)
- fix(fireworks\_ai): flatten dict-form reasoning\_effort to its effort string by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41335](https://github.com/BerriAI/litellm/pull/41335)
- fix(proxy): never forward the LiteLLM virtual key to Anthropic on the /anthropic passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41340](https://github.com/BerriAI/litellm/pull/41340)
- fix(proxy): rename AWS Secrets Manager secret when key alias changes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41468](https://github.com/BerriAI/litellm/pull/41468)
- feat(otel): promote nested request metadata keys to litellm.metadata.\* span attributes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41462](https://github.com/BerriAI/litellm/pull/41462)
- fix(proxy): sync AWS Secrets Manager on body-less key regenerate and key alias changes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41458](https://github.com/BerriAI/litellm/pull/41458)
- fix(http\_handler): keep a handler alive while a response it issued is still reading by [@​max-sixty](https://github.com/max-sixty) in [#​34829](https://github.com/BerriAI/litellm/pull/34829)
- fix(bedrock): make prompt caching work on the Nova InvokeModel route by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41343](https://github.com/BerriAI/litellm/pull/41343)
- ci(migrations): flag defaulted ADD COLUMN on request-log tables by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41460](https://github.com/BerriAI/litellm/pull/41460)
- feat(prometheus): add customer (end\_user) budget gauges by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41472](https://github.com/BerriAI/litellm/pull/41472)
- fix(otel): drop None metric and event attributes before OTLP export by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36815](https://github.com/BerriAI/litellm/pull/36815)
- fix(anthropic): carry the served model from message\_start onto stream chunks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41446](https://github.com/BerriAI/litellm/pull/41446)
- fix(models): rolling registry audit: Gemini latest aliases, Nova cache pricing, OpenRouter/Together sync, Mistral GLM 5.3, Azure snapshots, Grok caching by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41112](https://github.com/BerriAI/litellm/pull/41112)
- fix(router): count TPM/RPM usage before building rate-limit headers by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41474](https://github.com/BerriAI/litellm/pull/41474)
- feat(guardrails): release buffered stream chunks after each passing scan by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41425](https://github.com/BerriAI/litellm/pull/41425)
- fix!: re-check budget on router fallback targets by [@​runjivu](https://github.com/runjivu) in [#​41379](https://github.com/BerriAI/litellm/pull/41379)
- refactor(ocr): move file preparation from the python bridge into litellm-core by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41489](https://github.com/BerriAI/litellm/pull/41489)
- feat(s3): add s3\_log\_prompts\_only option to log prompts without responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41327](https://github.com/BerriAI/litellm/pull/41327)
- feat(team): team-level model\_max\_budget with key-level overrides by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41330](https://github.com/BerriAI/litellm/pull/41330)
- feat(keys): filter /key/list by active, expired, revoked or deleted status and serve deleted keys from /key/info by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41311](https://github.com/BerriAI/litellm/pull/41311)
- feat(proxy): expose lifetime total\_spend on virtual keys by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41403](https://github.com/BerriAI/litellm/pull/41403)
- fix(proxy): release completed max-parallel slots promptly by [@​elifozdamar](https://github.com/elifozdamar) in [#​40843](https://github.com/BerriAI/litellm/pull/40843)
- feat(ui): accept ssh clone urls when registering a skill by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35418](https://github.com/BerriAI/litellm/pull/35418)
- fix(prices): dedupe Nova cache\_read\_input\_token\_cost keys left by a text merge by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41496](https://github.com/BerriAI/litellm/pull/41496)
- fix(otel): propagate W3C trace context on HTTP and WebSocket passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40669](https://github.com/BerriAI/litellm/pull/40669)
- fix(proxy): remove duplicate user budget hook that 429'd zero-cost models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41345](https://github.com/BerriAI/litellm/pull/41345)
- test(logging): pick this test's own records out of the shared log batch by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41487](https://github.com/BerriAI/litellm/pull/41487)
- test(together\_ai): move request-shape checks to the mapped file, drop the live ones by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41360](https://github.com/BerriAI/litellm/pull/41360)
- feat(ui): shared URL-state layer for tables and tabs by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41331](https://github.com/BerriAI/litellm/pull/41331)
- feat(e2e): make the provider cache reusable across builds and mount Bedrock behind it by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41402](https://github.com/BerriAI/litellm/pull/41402)
- fix(mcp): fail closed on missing upstream credentials by [@​joshua-berri](https://github.com/joshua-berri) in [#​41364](https://github.com/BerriAI/litellm/pull/41364)
- feat(rust): scaffold Redis cache crate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41501](https://github.com/BerriAI/litellm/pull/41501)
- fix(dashscope): forward reasoning\_effort to the provider by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37506](https://github.com/BerriAI/litellm/pull/37506)
- fix(proxy): carry litellm\_call\_id through endpoint specific error logs and failure responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41356](https://github.com/BerriAI/litellm/pull/41356)
- fix(proxy): retry rate-limit fallbacks from a pristine request snapshot by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40596](https://github.com/BerriAI/litellm/pull/40596)
- fix(gemini): map minimal thinking to low for Gemini 3.7 and 3.8 Flash by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41201](https://github.com/BerriAI/litellm/pull/41201)
- fix(proxy): stop forwarding LiteLLM credential headers on Bedrock agent-runtime passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41504](https://github.com/BerriAI/litellm/pull/41504)
- fix(streaming): estimate interrupted Anthropic stream usage from reasoning\_content by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41503](https://github.com/BerriAI/litellm/pull/41503)
- fix(azure\_ai): route Responses API to native /openai/v1/responses for Foundry Models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33856](https://github.com/BerriAI/litellm/pull/33856)
- fix(proxy): show all model groups to proxy admins in /model\_group/info by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41094](https://github.com/BerriAI/litellm/pull/41094)
- feat(proxy): let proxy admins choose which team fields team admins may edit by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​39996](https://github.com/BerriAI/litellm/pull/39996)
- fix(bedrock): neutralize orphaned tool blocks instead of raising or injecting a dummy tool (internal copy of [#​31400](https://github.com/BerriAI/litellm/issues/31400)) by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41513](https://github.com/BerriAI/litellm/pull/41513)
- feat(ui): persist organizations and projects list, detail tab and key table state in the URL by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41445](https://github.com/BerriAI/litellm/pull/41445)
- fix(bedrock\_mantle): accept and forward verbosity on gpt-5.x chat completions by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41509](https://github.com/BerriAI/litellm/pull/41509)
- test: cover database transactions and persisted accounting by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41073](https://github.com/BerriAI/litellm/pull/41073)
- test: provider wire contracts, streaming and recovery by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41075](https://github.com/BerriAI/litellm/pull/41075)
- fix(mcp): count admin static headers as api\_key credential slots by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41514](https://github.com/BerriAI/litellm/pull/41514)
- ci: auto-merge provider-info-sync PRs when CI, Greptile and Bugbot are clean by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41494](https://github.com/BerriAI/litellm/pull/41494)
- feat(rust): add standalone framing crate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41500](https://github.com/BerriAI/litellm/pull/41500)
- fix(proxy): enforce tag budgets for tags added by guardrails by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40842](https://github.com/BerriAI/litellm/pull/40842)
- fix(utils): run post-call deployment hook on converted chat streams by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41495](https://github.com/BerriAI/litellm/pull/41495)
- fix(e2e): bind provider-cache recordings to the deployment's test, not the serving process by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41520](https://github.com/BerriAI/litellm/pull/41520)
- test: add extension and browser integration contracts by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41078](https://github.com/BerriAI/litellm/pull/41078)
- fix(logging): scan each log record once and collapse base64 payloads before the secret regex by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40934](https://github.com/BerriAI/litellm/pull/40934)
- fix(spend\_tracking): attribute router-rejected requests to the model group provider by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41507](https://github.com/BerriAI/litellm/pull/41507)
- feat(router): discover token limits for hosted OpenAI-compatible models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41508](https://github.com/BerriAI/litellm/pull/41508)
- feat(proxy): let team admins edit rpm\_limit and max\_budget when enabled by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41525](https://github.com/BerriAI/litellm/pull/41525)
- fix(otel): fit per-index OpenInference messages to the span's remaining attribute budget by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41498](https://github.com/BerriAI/litellm/pull/41498)
- test(aws): verify rotated secret value by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41524](https://github.com/BerriAI/litellm/pull/41524)
- test(e2e): read a deleted key back as deleted, not as a 404 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41551](https://github.com/BerriAI/litellm/pull/41551)
- fix(otel v2): map the caller's Langfuse user, session and tags onto the root and generation spans by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41140](https://github.com/BerriAI/litellm/pull/41140)
- test: fix seven tests left stale by [#​41311](https://github.com/BerriAI/litellm/issues/41311), [#​41337](https://github.com/BerriAI/litellm/issues/41337), [#​39996](https://github.com/BerriAI/litellm/issues/39996), [#​41310](https://github.com/BerriAI/litellm/issues/41310), [#​41289](https://github.com/BerriAI/litellm/issues/41289) and [#​41315](https://github.com/BerriAI/litellm/issues/41315) by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41527](https://github.com/BerriAI/litellm/pull/41527)
- test(budgets): cover management null handling by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41563](https://github.com/BerriAI/litellm/pull/41563)
- test(e2e): drop the auto-router select "opens below" spec by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41568](https://github.com/BerriAI/litellm/pull/41568)
- fix(guardrails): stream Prompt Security post\_call redactions in incremental\_diff mode by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41558](https://github.com/BerriAI/litellm/pull/41558)
- fix(guardrails): give post-call scans the scoped request conversation and tools by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41220](https://github.com/BerriAI/litellm/pull/41220)
- feat(openrouter): add stealth/union-alpha to the model cost map by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41576](https://github.com/BerriAI/litellm/pull/41576)
- test(management): cover project authorization lifecycle by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41573](https://github.com/BerriAI/litellm/pull/41573)
- feat(rust): map Anthropic Messages transformations by [@​yujonglee-berri](https://github.com/yujonglee-berri) in [#​41531](https://github.com/BerriAI/litellm/pull/41531)
- fix(e2e): clear the three standing errors in the scheduled Buildkite suite by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41616](https://github.com/BerriAI/litellm/pull/41616)
- refactor(rust\_bridge): declarative route catalog and shared runtime selection by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41479](https://github.com/BerriAI/litellm/pull/41479)
- fix(mock\_completion): keep the resolved provider so router custom pricing resolves for azure\_ai deployments by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41623](https://github.com/BerriAI/litellm/pull/41623)
- fix(mcp): restrict health discovery to virtual key grants by [@​joshua-berri](https://github.com/joshua-berri) in [#​41609](https://github.com/BerriAI/litellm/pull/41609)
- fix(mcp): preserve request-selected guardrails during tool execution by [@​joshua-berri](https://github.com/joshua-berri) in [#​41619](https://github.com/BerriAI/litellm/pull/41619)
- refactor(ocr): mirror Python provider layout and preserve tests by [@​yujonglee-berri](https://github.com/yujonglee-berri) in [#​41550](https://github.com/BerriAI/litellm/pull/41550)
- test(fireworks\_ai): stop pinning vision support on minimax-m3 by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41627](https://github.com/BerriAI/litellm/pull/41627)
- perf(spend\_tracking): index LiteLLM\_SpendLogs by (api\_key, startTime) by [@​etiennechabert](https://github.com/etiennechabert) in [#​37983](https://github.com/BerriAI/litellm/pull/37983)
- fix(proxy): reject non-string model with 400 and log its spend as unknown-model by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41633](https://github.com/BerriAI/litellm/pull/41633)
- test(together\_ai): stop pinning successor deprecation status by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41635](https://github.com/BerriAI/litellm/pull/41635)
- chore(prices): sync Together AI prices: 6 models, 6 deprecated \[sync failed: Google Gemini] by [@​berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#​41570](https://github.com/BerriAI/litellm/pull/41570)
- fix(budgets): page end-user cache invalidation after a budget reset by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41488](https://github.com/BerriAI/litellm/pull/41488)
- chore: bump litellm-proxy-extras 0.4.98 -> 0.4.99 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41659](https://github.com/BerriAI/litellm/pull/41659)
- fix(tests): resolve the integration support package without run.py's PYTHONPATH by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41373](https://github.com/BerriAI/litellm/pull/41373)
- fix(ui): keep untimed guardrail entries on the request lifecycle by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​41374](https://github.com/BerriAI/litellm/pull/41374)
- fix(anthropic-bridge): convert mid-conversation system turns to user turns on /v1/messages to chat completions by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41493](https://github.com/BerriAI/litellm/pull/41493)
- fix(bedrock): support aws-sdk-bedrock-runtime 0.10/0.11 in Bedrock Realtime by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41542](https://github.com/BerriAI/litellm/pull/41542)
- feat(cli): deprecate the litellm-proxy entrypoint in favour of lite by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41673](https://github.com/BerriAI/litellm/pull/41673)
- fix(scim): align pagination `count` validation with RFC 7644 by [@​zachbernstein-sdx](https://github.com/zachbernstein-sdx) in [#​41444](https://github.com/BerriAI/litellm/pull/41444)
- fix(bedrock): never emit Converse cachePoint for OpenAI-family models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41419](https://github.com/BerriAI/litellm/pull/41419)
- fix(images): stop forwarding the raw image\[] and mask\[] form keys by [@​mateo-berri](https://github.com/mateo-berri) in [#​39512](https://github.com/BerriAI/litellm/pull/39512)
- feat(management\_v1): bulk update team member budgets by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41632](https://github.com/BerriAI/litellm/pull/41632)
- refactor(rust): extract provider translations by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41690](https://github.com/BerriAI/litellm/pull/41690)
- feat(cli): rename lite autoroute up/down to start/stop, keeping the old names as deprecated aliases by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41672](https://github.com/BerriAI/litellm/pull/41672)
- fix(responses): keep the addressed response id off bridged provider requests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41689](https://github.com/BerriAI/litellm/pull/41689)
- fix(license): let a wildcard allowed\_features license grant the auto\_router feature by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41684](https://github.com/BerriAI/litellm/pull/41684)
- fix(ui): list every provider in the cache leakage by-model table by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​40875](https://github.com/BerriAI/litellm/pull/40875)
- fix(team): keep a forked member budget's reset window and audit bulk member budget writes by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​41686](https://github.com/BerriAI/litellm/pull/41686)
- feat(proxy): add TypeSafe AI Jev evaluate passthrough with registry-priced spend tracking by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41607](https://github.com/BerriAI/litellm/pull/41607)
- test(e2e): cover bedrock batch file upload and create in the us-gov-west-1 partition by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41536](https://github.com/BerriAI/litellm/pull/41536)
- feat(grafana): add all-metrics dashboard and fix stale dashboard\_v2 gauges by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41578](https://github.com/BerriAI/litellm/pull/41578)
- fix(cost): price Azure PTU spillover requests at standard token rates by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41569](https://github.com/BerriAI/litellm/pull/41569)
- build(deps): bump soupsieve to 2.9.2 to clear the osv-scan advisories by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41703](https://github.com/BerriAI/litellm/pull/41703)
- fix(fireworks\_ai): restore supports\_vision on minimax-m3 in the cost map by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41699](https://github.com/BerriAI/litellm/pull/41699)
- feat(policy\_engine): explicit priority for policy attachment execution order by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41571](https://github.com/BerriAI/litellm/pull/41571)
- feat(router): add TypeSafe Jev as a complexity router classifier by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​41615](https://github.com…
Updates 14 models, each read from its provider's own published prices.
OpenAI
No price changes; listed for what its tick could not apply.
Skipped by the parser (7)
gpt-6-astra: long-context prices present but the row states no thresholdgpt-5.6-sol: long-context prices present but the row states no thresholdgpt-5.6-terra: long-context prices present but the row states no thresholdgpt-5.6-luna: long-context prices present but the row states no thresholdsora-2-pro 1024p: size variant has no litellm key stated on the pagesora-2-pro 1080p: size variant has no litellm key stated on the pageft:o4-mini-2025-04-16: data-sharing discount has no litellm keyTogether AI
No price changes; listed for what its tick could not apply.
Skipped by the parser (98)
Prism-ML/Ternary-Bonsai-27B: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/meta/llama-3.2-11b-vision-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/meta/llama-3.2-90b-vision-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/mistralai/mixtral-8x22b-instruct-v01: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/meta/llama-3.3-70b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/nvidia/llama-3.1-nemotron-70b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/meta/llama-3.1-8b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/meta/llama-3.1-70b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/nv-mistralai/mistral-nemo-12b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/mistralai/mixtral-8x7b-instruct-v01: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/nvidia/llama-3.3-nemotron-super-49b-v1: no serverless token price (pricing.input and pricing.output are 0 or absent)google/gemma-3-27b-it: no serverless token price (pricing.input and pricing.output are 0 or absent)meta-llama/Llama-2-7b-chat-hf: no serverless token price (pricing.input and pricing.output are 0 or absent)deepseek-ai/DeepSeek-R1-Distill-Qwen-7B: no serverless token price (pricing.input and pricing.output are 0 or absent)google/gemma-3-1b-it: no serverless token price (pricing.input and pricing.output are 0 or absent)google/gemma-3-4b-it: no serverless token price (pricing.input and pricing.output are 0 or absent)deepcogito/cogito-v1-preview-llama-8B: no serverless token price (pricing.input and pricing.output are 0 or absent)deepcogito/cogito-v1-preview-qwen-32B: no serverless token price (pricing.input and pricing.output are 0 or absent)deepcogito/cogito-v1-preview-qwen-14B: no serverless token price (pricing.input and pricing.output are 0 or absent)deepcogito/cogito-v1-preview-llama-70B: no serverless token price (pricing.input and pricing.output are 0 or absent)deepcogito/cogito-v1-preview-llama-70B-Turbo: no serverless token price (pricing.input and pricing.output are 0 or absent)meta-llama/Llama-3.3-70B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-32B: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-72B: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-3B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-1.5B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-14B: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-7B: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-1.5B: no serverless token price (pricing.input and pricing.output are 0 or absent)meta-llama/Meta-Llama-3.1-70B: no serverless token price (pricing.input and pricing.output are 0 or absent)meta-llama/Llama-3.2-1B: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-7B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-32B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)meta-llama/Llama-3.1-405B: no serverless token price (pricing.input and pricing.output are 0 or absent)agentica-org/DeepCoder-14B-Preview: no serverless token price (pricing.input and pricing.output are 0 or absent)mistralai/Mistral-7B-v0.1: no serverless token price (pricing.input and pricing.output are 0 or absent)mistralai/Devstral-Small-2505: no serverless token price (pricing.input and pricing.output are 0 or absent)mistralai/Mixtral-8x22B-Instruct-v0.1: no serverless token price (pricing.input and pricing.output are 0 or absent)allenai/Molmo-7B-D-0924: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen3-8B: no serverless token price (pricing.input and pricing.output are 0 or absent)Vertex AI
Updates 14 models from the provider's pricing page (page
c7bf48f187b9).Changes
vertex_ai/deep-research-pro-preview-12-2025: cache_read_input_token_cost $0.2/1M, source https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/3-pro-image → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-2.5-flash-image: input_cost_per_token_batches $0.15/1M, input_cost_per_token_flex $0.15/1M, input_cost_per_token_priority $0.54/1M, output_cost_per_token_batches $1.25/1M, output_cost_per_token_flex $1.25/1M, source https://cloud.google.com/vertex-ai/generative-ai/docs/multimodal/image-generation#edit-an-image → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-3-flash-preview: cache_read_input_audio_token_cost $0.1/1M, cache_read_input_token_cost_flex $0.05/1M, input_cost_per_token_batches $0.25/1M, input_cost_per_token_flex $0.25/1M, output_cost_per_token_batches $1.5/1M, output_cost_per_token_flex $1.5/1M, source https://cloud.google.com/vertex-ai/generative-ai/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-3-pro-image: cache_read_input_token_cost $0.2/1M, cache_read_input_token_cost_above_200k_tokens $0.4/1M, cache_read_input_token_cost_above_200k_tokens_priority $0.72/1M, cache_read_input_token_cost_flex $0.1/1M, cache_read_input_token_cost_priority $0.36/1M, input_cost_per_token_above_200k_tokens $4/1M, input_cost_per_token_above_200k_tokens_priority $7.2/1M, input_cost_per_token_flex $1/1M, input_cost_per_token_priority $3.6/1M, output_cost_per_token_above_200k_tokens $18/1M, output_cost_per_token_above_200k_tokens_priority $32.4/1M, output_cost_per_token_flex $6/1M, output_cost_per_token_priority $21.6/1M, source https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/3-pro-image → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-3.1-flash-image: cache_read_input_token_cost $0.05/1M, cache_read_input_token_cost_flex $0.025/1M, input_cost_per_token_batches $0.25/1M, input_cost_per_token_flex $0.25/1M, output_cost_per_token_batches $1.5/1M, output_cost_per_token_flex $1.5/1M, source https://cloud.google.com/vertex-ai/generative-ai/pricing#gemini-models → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-3.1-flash-lite: cache_read_input_audio_token_cost $0.05/1M, source https://cloud.google.com/vertex-ai/generative-ai/pricing#gemini-models → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-3.1-flash-lite-image: cache_read_input_token_cost_flex $0.0125/1M, input_cost_per_token_flex $0.125/1M, output_cost_per_token_flex $0.75/1Mvertex_ai/gemini-3.1-pro-preview: cache_read_input_token_cost_flex $0.2/1M, input_cost_per_token_flex $1/1M, output_cost_per_token_flex $6/1M, source https://cloud.google.com/vertex-ai/generative-ai/pricing#gemini-models → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-3.5-flash: source https://cloud.google.com/vertex-ai/generative-ai/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-3.5-flash-lite: cache_read_input_token_cost_priority $0.05/1M → $0.054/1M, source https://cloud.google.com/vertex-ai/generative-ai/pricing#gemini-models → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-3.6-flash: source https://cloud.google.com/vertex-ai/generative-ai/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-3.7-flash: source https://cloud.google.com/vertex-ai/generative-ai/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-3.8-flash: source https://cloud.google.com/vertex-ai/generative-ai/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-embedding-2: source https://cloud.google.com/vertex-ai/generative-ai/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingSkipped by the parser (3)
gemini-2.5-flash: "Audio Input" above 200K is $0.30, not the base price; litellm has no field for itGemini 2.0 Flash Image Generation: litellm keys only the free experimental model (gemini-2.0-flash-exp-image-generation)Gemini 2.0 Flash Live API: litellm has no vertex_ai/ key for the 2.0 Live APIOpened by the litellm-providers price sync. Every price is read from the provider's own published source and gated before it is applied; the audit trail for each value is in the portal's sync history.