Skip to content

fix(azure): rename max_tokens to max_completion_tokens for gpt-5-chat deployments - #36857

Merged
yassin-berriai merged 1 commit into
litellm_internal_stagingfrom
litellm_lit5549_azure_gpt5_chat_max_completion_tokens
Aug 17, 2026
Merged

yassin-berriai merged 1 commit into
litellm_internal_stagingfrom
litellm_lit5549_azure_gpt5_chat_max_completion_tokens

Conversation

@yassin-berriai

@yassin-berriai yassin-berriai commented Aug 14, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Azure gpt-5-chat deployments 400 on every max_tokens request
  • /health reports those deployments permanently unhealthy
  • The rename that fixes it skipped the whole gpt-5-chat family

How it solves it:

User Flow

Before: a developer whose Azure deployment is named gpt-5-chat gets a 400 on every request that sets max_tokens, and the deployment sits red in /health forever

  1. They send POST https://litellm-domain/v1/chat/completions with {"model": "gpt-5-chat", "messages": [...], "max_tokens": 5}
  2. HTTP 400 comes back: Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead.
  3. Their admin opens GET https://litellm-domain/health?model=gpt-5-chat and sees the deployment under unhealthy_endpoints with that same message, healthy_count 0
  4. The synthetic monitor watching that endpoint pages on-call, and keeps paging

After: the same request succeeds and the same deployment reports healthy

  1. They send the same POST https://litellm-domain/v1/chat/completions with {"model": "gpt-5-chat", "messages": [...], "max_tokens": 5}
  2. HTTP 200 comes back with a normal chat completion
  3. Their admin opens GET https://litellm-domain/health?model=gpt-5-chat and sees the deployment under healthy_endpoints, unhealthy_count 0
  4. The monitor goes quiet

Relevant issues

No upstream issue covers this defect. #13781 is the one this change is careful not to reopen: it reported gpt-5-chat-latest wrongly rejecting temperature, which is why the family was pulled off the GPT-5 reasoning path in the first place. #24779, closed as not planned, claims the same Azure error string for azure/gpt-4o

Linear ticket

Resolves LIT-5549

Pre-Submission checklist

  • I have added meaningful tests
  • My PR passes all CI/CD checks (e.g., lint, format, unit tests)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

There are no Azure credentials on the machine this was captured on, so the run below points a live proxy at a local HTTP server standing in for Azure. That server records the exact JSON body litellm puts on the wire and returns Azure's own 400 whenever max_tokens is present. The wire body is measured, the rejection is simulated, and the claim that Azure really rejects max_tokens for this family rests on a customer's production error plus Azure/azure-sdk-for-net#51844, which names gpt-5-chat among the deployments returning that exact string

Both runs use the same config, four Azure deployments (gpt-5-chat, gpt-5, o3-mini, gpt-4o) all pointed at the stand-in on 127.0.0.1:45550, proxy on 127.0.0.1:45549

Before, at e0f388c (this branch's base):

$ curl -s http://127.0.0.1:45549/health/liveliness
"I'm alive!"

$ curl -s -w '\nHTTP %{http_code}\n' http://127.0.0.1:45549/v1/chat/completions \
    -H 'Authorization: Bearer sk-lit5549' -H 'Content-Type: application/json' \
    -d '{"model":"gpt-5-chat","messages":[{"role":"user","content":"hi"}],"max_tokens":5}'
{"error":{"message":"litellm.BadRequestError: AzureException BadRequestError - Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead.. Received Model Group=gpt-5-chat\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","param":"max_tokens","code":"400"}}
HTTP 400

body litellm put on the wire to Azure:
{"path": "/openai/deployments/gpt-5-chat/chat/completions?api-version=2024-12-01-preview", "body": {"messages": [{"role": "user", "content": "hi"}], "model": "gpt-5-chat", "max_tokens": 5, "stream": false}}

$ curl -s "http://127.0.0.1:45549/health?model=gpt-5-chat" -H 'Authorization: Bearer sk-lit5549'
{"healthy_endpoints":[],"unhealthy_endpoints":[{"api_base":"http://127.0.0.1:45550","model":"azure/gpt-5-chat","max_tokens":5,"error":"litellm.BadRequestError: AzureException BadRequestError - Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead. ...","exception_status":400}],"healthy_count":0,"unhealthy_count":1}

body litellm put on the wire to Azure:
{"path": "/openai/deployments/gpt-5-chat/chat/completions?api-version=2024-12-01-preview", "body": {"messages": [{"role": "user", "content": "What's 1 + 1?"}], "model": "gpt-5-chat", "max_tokens": 5}}

controls, same request shape:
  gpt-5   -> HTTP 200  wire: {"model": "gpt-5", "max_completion_tokens": 5, ...}
  o3-mini -> HTTP 200  wire: {"model": "o3-mini", "max_completion_tokens": 5, ...}
  gpt-4o  -> HTTP 400  wire: {"model": "gpt-4o", "max_tokens": 5, ...}

After, at 1e877f2, since rebased onto e1f3d6e, then 29d8bed, and now 5b0ea66 (this PR). The capture is still the code at the current head: litellm/llms/azure/chat/gpt_transformation.py carries a byte-identical contribution at every one of those commits, and the latest rebase only merged a neighbouring test from #34462 into the test file:

$ curl -s http://127.0.0.1:45549/health/liveliness
"I'm alive!"

$ curl -s -w '\nHTTP %{http_code}\n' http://127.0.0.1:45549/v1/chat/completions \
    -H 'Authorization: Bearer sk-lit5549' -H 'Content-Type: application/json' \
    -d '{"model":"gpt-5-chat","messages":[{"role":"user","content":"hi"}],"max_tokens":5}'
{"id":"chatcmpl-lit5549","created":1,"model":"gpt-5-chat","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"ok","role":"assistant"}}],"usage":{"completion_tokens":1,"prompt_tokens":1,"total_tokens":2}}
HTTP 200

body litellm put on the wire to Azure:
{"path": "/openai/deployments/gpt-5-chat/chat/completions?api-version=2024-12-01-preview", "body": {"messages": [{"role": "user", "content": "hi"}], "model": "gpt-5-chat", "max_completion_tokens": 5, "stream": false}}

$ curl -s "http://127.0.0.1:45549/health?model=gpt-5-chat" -H 'Authorization: Bearer sk-lit5549'
{"healthy_endpoints":[{"api_base":"http://127.0.0.1:45550","model":"azure/gpt-5-chat","max_tokens":5,"model_id":"2f92a5a8f85a34351a0b0e03034c389e3cab56c3101e47e6cbf1cd0f2c3ac4f5"}],"unhealthy_endpoints":[],"healthy_count":1,"unhealthy_count":0}

body litellm put on the wire to Azure:
{"path": "/openai/deployments/gpt-5-chat/chat/completions?api-version=2024-12-01-preview", "body": {"messages": [{"role": "user", "content": "What's 1 + 1?"}], "model": "gpt-5-chat", "max_completion_tokens": 5}}

controls, same request shape:
  gpt-5   -> HTTP 200  wire: {"model": "gpt-5", "max_completion_tokens": 5, ...}
  o3-mini -> HTTP 200  wire: {"model": "o3-mini", "max_completion_tokens": 5, ...}
  gpt-4o  -> HTTP 400  wire: {"model": "gpt-4o", "max_tokens": 5, ...}

The gpt-5 and o3-mini rows are the load-bearing part of the before run: the same harness prints max_completion_tokens whenever litellm emits it, so max_tokens on the gpt-5-chat row means the code picked that key rather than the fixture never carrying the alternative. gpt-4o is unchanged on purpose and its 400 only shows that the stand-in rejects max_tokens from everyone

Type

🐛 Bug Fix

Caveats (if any)

Review notes

Greptile scored this 4/5 on 1e877f2 with two findings it framed as non-blocking follow-ups. I looked into both and am deliberately not changing the code for either

On the predicate being hardcoded rather than read from the model cost map: no such flag exists. I enumerated every support* key present anywhere in model_prices_and_context_window.json, 40 of them across roughly three thousand entries, and none expresses "this deployment rejects the legacy max_tokens key". The nearest neighbours are supports_reasoning, supports_sampling_params and the supports_*_reasoning_effort family, which all answer the reasoning question instead. supports_reasoning is the worst of them to borrow here: azure/gpt-5-chat carries supports_reasoning: true today, so keying the rename off it would rebuild the exact conflation of two independent questions that caused this bug. The map also already spends the name max_tokens on the output-token ceiling, so a capability flag named around it would read as a collision. Model naming is the only source of truth for this capability at present, and all three pre-existing renames in the tree agree: llms/openai/chat/o_series_transformation.py:99 keys off is_o_series_model, and llms/openai/chat/gpt_5_transformation.py:200 and :256 key off is_model_gpt_5_search_model and is_model_gpt_5_model. Introducing a flag would mean populating it correctly across the whole gpt-5 and o-series families before it could be trusted, which is a far larger and riskier change than this defect justifies

On the new branch introducing another parameter mutation: optional_params is a local accumulator here, not caller-owned state. It is created as a fresh {} inside pre_process_optional_params (utils.py:3855), returned into get_optional_params, and threaded through the provider map_openai_params chain with the return value reassigned at every hop, which is the contract BaseConfig.map_openai_params declares. Filling it is what the function is for, and every other branch of the same loop fills it the same way. The new branch is actually the least mutating of the four renames in the tree, since the three cited above each call non_default_params.pop("max_tokens") and reach into the caller's dict, while this one writes only into the accumulator and pops nothing

The automated security review raised a Medium, that a request carrying both max_tokens=1 and max_completion_tokens=10000 reserves one output token while the provider receives ten thousand, and asked for the conflicting fields to be rejected or normalised before the pre-call rate-limit hooks run. The under-reservation is real and worth fixing. It is being fixed in #37001, in the limiter, and deliberately not here. Three measurements, each of which independently rules out doing it in this file

Ordering. The reservation is taken by _estimate_tokens_for_request at proxy/hooks/parallel_request_limiter_v3.py:607, reading data.get("max_tokens") or data.get("max_completion_tokens") straight off the raw request body inside async_pre_call_hook, which fires at proxy/common_request_processing.py:1471 before routing and long before any provider transformation. Driven directly, that body reserves 2 tokens where the same body carrying only max_completion_tokens reserves 10001. Nothing in map_openai_params can move those numbers, because the reservation is already written by the time the transformation runs. Rejecting there would convert a bypass into a 400 after the fact rather than reserve correctly, which is not the same thing

Reach. This is not introduced here. At this branch's merge-base, the same request already drops max_tokens: 1 and emits max_completion_tokens: 10000 for azure/gpt-5, azure/o3-mini, openai/gpt-5 and openai/o3-mini. On a live proxy at base, azure/gpt-5 with both fields returns HTTP 200 with ten thousand on the wire, which is exactly what azure/gpt-5-chat does after this change. Closing it inside AzureOpenAIConfig would shut one of five doors while implying the corridor was secured

Blast radius. Rejecting the combination would break a working configuration. LiteLLM_Params sets extra="allow" and the router merges deployment params underneath client kwargs, so an operator's deployment-level max_tokens default and a client's max_completion_tokens arrive together by design. Run end to end through the router, a deployment carrying max_tokens: 4096 with a client sending max_completion_tokens: 100 returns HTTP 200 and puts {"max_completion_tokens": 100} on the wire, the client value correctly winning. A conflict error would fail every operator in that shape

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Aug 14, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR updates Azure GPT-5 chat parameter translation so legacy max_tokens inputs are emitted as max_completion_tokens without moving chat variants onto the reasoning-model path.

  • Adds the Azure GPT-5-family token-key translation.
  • Adds regression coverage for GPT-5 chat variants, reasoning parameter behavior, and unaffected control models.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Important Files Changed

Filename Overview
litellm/llms/azure/chat/gpt_transformation.py Adds an isolated Azure request-parameter translation for GPT-5 chat deployments while preserving existing reasoning-model classification.
tests/test_litellm/llms/azure/chat/test_azure_chat_gpt_transformation.py Adds focused regression tests covering token-key emission and GPT-5 chat versus reasoning behavior.

Reviews (3): Last reviewed commit: "fix(azure): rename max_tokens to max_com..." | Re-trigger Greptile

name family needs the rename, including the ``gpt-5-chat*`` models that are excluded from
the reasoning path by https://github.com/BerriAI/litellm/issues/13781.
"""
return "gpt-5" in model or "gpt5_series" in model

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Hardcoded model capability predicate

The new predicate infers token-parameter support exclusively from gpt-5 and gpt5_series substrings instead of model metadata. This requires LiteLLM releases for future capability changes and lets naming aliases select the wrong outbound parameter; represent the capability in model metadata and access it through get_model_info.

Rule Used: What: Do not hardcode model-specific flags in the ... (source)

Knowledge Base Used: LLM Provider Adapters

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

for param, value in non_default_params.items():
if param == "tool_choice":
if param == "max_tokens" and renames_max_tokens:
optional_params.setdefault("max_completion_tokens", value)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Parameter mapping mutates input

Calling setdefault directly mutates the caller-provided optional_params mapping, adding another in-place state change that makes ownership and parameter precedence harder to reason about. Construct and return the updated mapping without mutating the function parameter.

Context Used: CLAUDE.md (source)

Comment thread litellm/llms/azure/chat/gpt_transformation.py
@veria-ai

veria-ai Bot commented Aug 14, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 1 · PR risk: 0/10

@codecov

codecov Bot commented Aug 14, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@yassin-berriai

Copy link
Copy Markdown
Contributor Author

@greptileai adjudicated both findings in the PR description under Review notes: no such model-map flag exists, and optional_params is a local accumulator. Head 1e877f2

@greptile-apps

greptile-apps Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Both adjudications hold up.

Model-map flag: Correct that no suitable flag exists. supports_reasoning would be the wrong key to borrow — azure/gpt-5-chat carries it today, so keying the rename off it would recreate the conflation this PR is fixing. Name-based predicates are the existing convention across all three prior renames in the tree, and the new predicate ("gpt-5" in model or "gpt5_series" in model) is consistent with that pattern. No objection.

optional_params mutation: setdefault is actually the most conservative of the four renames — it only writes to the accumulator and leaves non_default_params untouched, while the three existing renames all call non_default_params.pop("max_tokens"). The concern doesn't apply.

The test suite is thorough: it covers the new family (gpt-5-chat, gpt-5-chat-latest, gpt-5-chat-2025-08-07), the positive control for the reasoning path (gpt-5), the regression guard for #13781 (gpt-5-chat* stays off the reasoning path), and an unchanged negative (gpt-4o). Both original findings are resolved by the evidence in the PR description and the diff. The fix looks correct.

@codspeed

codspeed Bot commented Aug 14, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_lit5549_azure_gpt5_chat_max_completion_tokens (5b0ea66) with litellm_internal_staging (973329e)

Open in CodSpeed

@yassin-berriai
yassin-berriai force-pushed the litellm_lit5549_azure_gpt5_chat_max_completion_tokens branch from 1e877f2 to 29d8bed Compare August 14, 2026 17:26
@yassin-berriai

Copy link
Copy Markdown
Contributor Author

@greptileai rebased onto current staging, no code change since the last review. Please re-review head 29d8bed. Rebuttal stands in the PR description under Review notes

@yassin-berriai

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 29d8bed. Configure here.

@yassin-berriai

Copy link
Copy Markdown
Contributor Author

@veria-ai answered in the PR description under Review notes: the reservation predates this change and is fixed in #37001. Head 29d8bed

… deployments

Azure rejects the legacy `max_tokens` key for the whole gpt-5 name family, but
`AzureOpenAIGPT5Config.is_model_gpt_5_model` deliberately excludes `gpt-5-chat*`
so those deployments fall through to `AzureOpenAIConfig`, which sends `max_tokens`
verbatim and gets a 400 back on every request that carries it, `/health` probes
included.

One predicate was answering two independent questions. Split it: the new
`AzureOpenAIConfig.requires_max_completion_tokens` covers the whole gpt-5 name
family and drives only the rename, while `is_model_gpt_5_model` keeps keying
reasoning_effort, the temperature clamp and the dropped penalties off the
reasoning question, so #13781 stays fixed.
@yassin-berriai
yassin-berriai force-pushed the litellm_lit5549_azure_gpt5_chat_max_completion_tokens branch from 29d8bed to 5b0ea66 Compare August 17, 2026 16:29
@yassin-berriai

Copy link
Copy Markdown
Contributor Author

@greptileai please review the current head 5b0ea66. Only change since your 5/5: rebase onto staging, resolving a test-file conflict with #34462

@yassin-berriai

Copy link
Copy Markdown
Contributor Author

Rebased onto staging to clear the conflict. The approval at 29d8bed still holds: identical production diff, test file only gained #34462 neighbour

@yassin-berriai
yassin-berriai merged commit 9b7ed77 into litellm_internal_staging Aug 17, 2026
70 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_lit5549_azure_gpt5_chat_max_completion_tokens branch August 17, 2026 18:27
doonga pushed a commit to greyrock-labs/home-ops that referenced this pull request Sep 5, 2026
…9.1) (#104)

This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.98.0` → `v1.99.1` |

---

### Release Notes

<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>

### [`v1.99.1`](https://github.com/BerriAI/litellm/releases/tag/v1.99.1)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.99.0...v1.99.1)

#### Docker-only release

**This release ships container images only. There is no PyPI package for `1.99.1`.**

`pip install litellm==1.99.1` will not resolve — install the images below, or stay on `1.99.0` on PyPI. The git tag and this release exist so the images are traceable to an exact commit.

| Image                                                                     | Tags                |
| ------------------------------------------------------------------------- | ------------------- |
| `ghcr.io/berriai/litellm` · `docker.io/litellm/litellm`                   | `1.99.1`, `v1.99.1` |
| `ghcr.io/berriai/litellm-database` · `docker.io/litellm/litellm-database` | `1.99.1`, `v1.99.1` |
| `ghcr.io/berriai/litellm-non_root` · `docker.io/litellm/litellm-non_root` | `1.99.1`, `v1.99.1` |

This is the newest stable image, so the rolling `latest` and `main-stable` image tags now point at `1.99.1`.

It carries one fix on top of `1.99.0`: OpenTelemetry v2 spans now emit cache token counts (`gen_ai.usage.cache_creation.input_tokens` and `gen_ai.usage.cache_read.input_tokens`) alongside the cache cost that was already reported. If you compute spend from OTel token counts rather than from LiteLLM's own cost fields, prompt-caching workloads were previously under-counted.

***

#### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.99.1
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.99.1/cosign.pub \
  ghcr.io/berriai/litellm:v1.99.1
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

#### What's Changed

- chore(release): backport [#&#8203;38716](https://github.com/BerriAI/litellm/issues/38716) to stable/1.99.x and cut 1.99.1 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;39179](https://github.com/BerriAI/litellm/pull/39179)

**Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.99.0...v1.99.1>

### [`v1.99.0`](https://github.com/BerriAI/litellm/releases/tag/v1.99.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.98.0...v1.99.0)

#### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.99.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.99.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.99.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

#### What's Changed

- chore(typing): drop 1.3k basedpyright errors across 30 Any hotspot files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37073](https://github.com/BerriAI/litellm/pull/37073)
- fix(proxy): register WebSocket passthrough for OpenAI prefixes by [@&#8203;LHMQ878](https://github.com/LHMQ878) in [#&#8203;36151](https://github.com/BerriAI/litellm/pull/36151)
- fix(bedrock): report uploaded size in the FileObject returned by managed batch uploads by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36392](https://github.com/BerriAI/litellm/pull/36392)
- fix(batches): support AWS Bedrock batch cancellation via `StopModelInvocationJob` by [@&#8203;ArjunPakhan](https://github.com/ArjunPakhan) in [#&#8203;34087](https://github.com/BerriAI/litellm/pull/34087)
- feat: Async Rust OCR Bridge and MCP OAuth UI Restore by [@&#8203;ArjunPakhan](https://github.com/ArjunPakhan) in [#&#8203;31453](https://github.com/BerriAI/litellm/pull/31453)
- fix(batches): don't crash logging when a completed batch has no output file by [@&#8203;MUSE-CODE-SPACE](https://github.com/MUSE-CODE-SPACE) in [#&#8203;34067](https://github.com/BerriAI/litellm/pull/34067)
- fix(UI): add default model pin to complexity router UI by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36615](https://github.com/BerriAI/litellm/pull/36615)
- feat(ui): add Lite mixed-provider auto-router preset by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37068](https://github.com/BerriAI/litellm/pull/37068)
- feat(ui): link key info header to its user, creator, team, and organization by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37187](https://github.com/BerriAI/litellm/pull/37187)
- fix(guardrails): scan text on /guardrails/apply\_guardrail for Azure Content Safety by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36894](https://github.com/BerriAI/litellm/pull/36894)
- feat(bedrock): forward LiteLLM identity and metadata into Bedrock requestMetadata by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36861](https://github.com/BerriAI/litellm/pull/36861)
- fix(azure): rename max\_tokens to max\_completion\_tokens for gpt-5-chat deployments by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36857](https://github.com/BerriAI/litellm/pull/36857)
- fix(bedrock): preserve cache token usage when invocationMetrics replace the usage block by [@&#8203;brian5021](https://github.com/brian5021) in [#&#8203;36878](https://github.com/BerriAI/litellm/pull/36878)
- fix(proxy): registry caches stop per-request tag and end-user Postgres reads in auth by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36801](https://github.com/BerriAI/litellm/pull/36801)
- test(e2e): replay a real tool-search assistant turn back to Bedrock Invoke by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36856](https://github.com/BerriAI/litellm/pull/36856)
- fix(proxy): return 400 naming the missing required param on POST /v1/batches by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37199](https://github.com/BerriAI/litellm/pull/37199)
- fix(ci): bump sqlparse to 0.6.0 to resolve osv-scan CVEs by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37200](https://github.com/BerriAI/litellm/pull/37200)
- fix(ui): stop pairing key spend with the team budget when a key has no budget by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37196](https://github.com/BerriAI/litellm/pull/37196)
- fix(guardrails): record MCP tool guardrail evaluations and blocks in … by [@&#8203;Scott-Wilson-ZocDoc](https://github.com/Scott-Wilson-ZocDoc) in [#&#8203;36978](https://github.com/BerriAI/litellm/pull/36978)
- fix(proxy): return 400 for non-object metadata and litellm\_metadata instead of silent drop or 500 by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37203](https://github.com/BerriAI/litellm/pull/37203)
- fix(anthropic): preserve optional Responses tool properties by [@&#8203;Scott-Wilson-ZocDoc](https://github.com/Scott-Wilson-ZocDoc) in [#&#8203;36979](https://github.com/BerriAI/litellm/pull/36979)
- feat(ui): add user ID request log filter by [@&#8203;daniel-meismer-zocdoc](https://github.com/daniel-meismer-zocdoc) in [#&#8203;36781](https://github.com/BerriAI/litellm/pull/36781)
- fix(anthropic): stop emitting empty thinking blocks on the Responses adapter by [@&#8203;Scott-Wilson-ZocDoc](https://github.com/Scott-Wilson-ZocDoc) in [#&#8203;36033](https://github.com/BerriAI/litellm/pull/36033)
- fix(ui): make per-user usage filter searchable by [@&#8203;daniel-meismer-zocdoc](https://github.com/daniel-meismer-zocdoc) in [#&#8203;36790](https://github.com/BerriAI/litellm/pull/36790)
- refactor(ui): decouple bulk invite from the invite user button by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37061](https://github.com/BerriAI/litellm/pull/37061)
- fix(helm): bound the migrations Job so a blocked migration cannot stall the release by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36975](https://github.com/BerriAI/litellm/pull/36975)
- feat(proxy): let USE\_V2\_MIGRATION\_RESOLVER select the v2 migration resolver by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36258](https://github.com/BerriAI/litellm/pull/36258)
- fix(mcp): scope authorization server issuer by [@&#8203;irosh-colombage-ZocDoc2](https://github.com/irosh-colombage-ZocDoc2) in [#&#8203;36482](https://github.com/BerriAI/litellm/pull/36482)
- fix(responses): unwrap object-form tool\_choice before calling the Responses API by [@&#8203;Scott-Wilson-ZocDoc](https://github.com/Scott-Wilson-ZocDoc) in [#&#8203;36032](https://github.com/BerriAI/litellm/pull/36032)
- test(ui): query antd controls accessibly instead of by internal CSS class by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37014](https://github.com/BerriAI/litellm/pull/37014)
- fix(proxy): bill cancelled and failed batches that still produced an output file by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37205](https://github.com/BerriAI/litellm/pull/37205)
- fix(bedrock): read batch usage by payload shape, not by provider name by [@&#8203;marty-sullivan](https://github.com/marty-sullivan) in [#&#8203;37078](https://github.com/BerriAI/litellm/pull/37078)
- fix(ui): self-contained searchable user filter on the Usage page by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37206](https://github.com/BerriAI/litellm/pull/37206)
- revert: don't fix mcp scope authorization server issuer by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37220](https://github.com/BerriAI/litellm/pull/37220)
- fix(mcp): scope authorization server issuer for named MCP servers by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37204](https://github.com/BerriAI/litellm/pull/37204)
- test(ui): gate dashboard test assertions with testing-library and jest-dom rules by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37018](https://github.com/BerriAI/litellm/pull/37018)
- feat(shadow-eval): name the shadowed key in job responses and the UI headline by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37221](https://github.com/BerriAI/litellm/pull/37221)
- test(ui): assert what collaborators are called with, not merely that they were by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37019](https://github.com/BerriAI/litellm/pull/37019)
- fix(logging): stop deepcopying results redaction cannot redact by [@&#8203;marty-sullivan](https://github.com/marty-sullivan) in [#&#8203;36638](https://github.com/BerriAI/litellm/pull/36638)
- fix(gemini): price gemini 3.6 flash at Google's introductory rates on every service tier by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37197](https://github.com/BerriAI/litellm/pull/37197)
- perf(guardrails): stop sending the conversation twice in the noma v2 payload by [@&#8203;itaimodi](https://github.com/itaimodi) in [#&#8203;36764](https://github.com/BerriAI/litellm/pull/36764)
- fix(streaming): track provider-reported cost when caller omits include\_usage by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35013](https://github.com/BerriAI/litellm/pull/35013)
- fix: stop rust flag from leaking into upstream provider request bodies by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37218](https://github.com/BerriAI/litellm/pull/37218)
- fix(proxy): return 404 instead of 500 for unresolvable batch and file ids on /v1/batches by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37201](https://github.com/BerriAI/litellm/pull/37201)
- fix(bedrock): validate file-content retrieval against the configured output bucket ([#&#8203;26335](https://github.com/BerriAI/litellm/issues/26335)) by [@&#8203;kingdoooo](https://github.com/kingdoooo) in [#&#8203;31435](https://github.com/BerriAI/litellm/pull/31435)
- fix(proxy): reject out-of-range limit on GET /v1/batches with OpenAI-parity 400 by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37198](https://github.com/BerriAI/litellm/pull/37198)
- fix(batches): price a retrieved batch from its deployment's model and rates (internal copy of [#&#8203;37077](https://github.com/BerriAI/litellm/issues/37077)) by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37219](https://github.com/BerriAI/litellm/pull/37219)
- feat(ocr): return Azure Document Intelligence's native payload from /v1/ocr via req\_format=native by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37194](https://github.com/BerriAI/litellm/pull/37194)
- fix(anthropic): fold guardrail-modified leading system rows into top-level system param by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37231](https://github.com/BerriAI/litellm/pull/37231)
- fix(shadow\_eval): copy messages before router call and raise judge output cap by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37232](https://github.com/BerriAI/litellm/pull/37232)
- feat(proxy): add Amazon Comprehend Medical passthrough provider by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37229](https://github.com/BerriAI/litellm/pull/37229)
- test(ui): settle the in-flight search before the loading tests end by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37227](https://github.com/BerriAI/litellm/pull/37227)
- test(cli): use example.com placeholder host in base-url trailing slash test by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37240](https://github.com/BerriAI/litellm/pull/37240)
- feat(complexity\_router): operator-defined tier sets for the LLM classifier by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37226](https://github.com/BerriAI/litellm/pull/37226)
- feat(ui): configure the auto router's heuristic scorer from the Admin UI by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37216](https://github.com/BerriAI/litellm/pull/37216)
- fix(shadow\_eval): drop unused judge reasoning field and salvage truncated verdicts by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37239](https://github.com/BerriAI/litellm/pull/37239)
- feat(proxy): proactive model deprecation alerts and `/model/deprecations` endpoint by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;26900](https://github.com/BerriAI/litellm/pull/26900)
- refactor(ui): move dashboard toasts from antd message/notification onto sonner by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37207](https://github.com/BerriAI/litellm/pull/37207)
- feat(guardrails): track bedrock guardrail usage units per invocation by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37225](https://github.com/BerriAI/litellm/pull/37225)
- fix(proxy): strip callback credentials from the auth object stamped into request metadata by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37233](https://github.com/BerriAI/litellm/pull/37233)
- fix(guardrails): retry usage upserts only on connection errors by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37247](https://github.com/BerriAI/litellm/pull/37247)
- fix(mcp): oauth discovery must not cause outages by [@&#8203;daniel-meismer-zocdoc](https://github.com/daniel-meismer-zocdoc) in [#&#8203;36599](https://github.com/BerriAI/litellm/pull/36599)
- test(ui): await the playground model combobox before clicking it by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36850](https://github.com/BerriAI/litellm/pull/36850)
- refactor(ui): migrate budget and skill forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37262](https://github.com/BerriAI/litellm/pull/37262)
- refactor(ui): migrate tag and memory forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37266](https://github.com/BerriAI/litellm/pull/37266)
- feat(complexity\_router): plan-mode tier floor for coding-agent clients by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37230](https://github.com/BerriAI/litellm/pull/37230)
- refactor(ui): codemod every toast call site onto lib/toast and delete the antd-era facades by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37253](https://github.com/BerriAI/litellm/pull/37253)
- feat(proxy): add /team/daily/activity/aggregated and switch the Usage team tab to it by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36562](https://github.com/BerriAI/litellm/pull/36562)
- refactor(ui): migrate user, logging and policy forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37303](https://github.com/BerriAI/litellm/pull/37303)
- refactor(ui): migrate user, policy, and margin forms to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37305](https://github.com/BerriAI/litellm/pull/37305)
- refactor(ui): migrate the regenerate key and team member forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37300](https://github.com/BerriAI/litellm/pull/37300)
- refactor(ui): migrate CloudZero and cost tracking forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37312](https://github.com/BerriAI/litellm/pull/37312)
- refactor(ui): migrate auto router and credential forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37304](https://github.com/BerriAI/litellm/pull/37304)
- refactor(ui): migrate guardrail and vector store forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37306](https://github.com/BerriAI/litellm/pull/37306)
- refactor(ui): migrate prompt, UI access, plugin and MCP filter forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37297](https://github.com/BerriAI/litellm/pull/37297)
- refactor(ui): drop the unreachable user edit modal by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37327](https://github.com/BerriAI/litellm/pull/37327)
- fix(router): route Responses API input through the auto-router by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37333](https://github.com/BerriAI/litellm/pull/37333)
- refactor(ui): migrate the login, onboarding and search tool forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37334](https://github.com/BerriAI/litellm/pull/37334)
- feat(ui): plan-mode override tier in the auto-router create and edit forms by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37319](https://github.com/BerriAI/litellm/pull/37319)
- refactor(ui): retire the tremor date range picker in favour of the shared advanced picker by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37302](https://github.com/BerriAI/litellm/pull/37302)
- fix(proxy): forward Bedrock event-stream content-type on unbuffered passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33767](https://github.com/BerriAI/litellm/pull/33767)
- fix(azure\_ai): strip non-OpenAI-spec message fields before request by [@&#8203;ayaangazali](https://github.com/ayaangazali) in [#&#8203;34445](https://github.com/BerriAI/litellm/pull/34445)
- fix(proxy): stop leaking the client\_side\_timeout marker to providers by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37346](https://github.com/BerriAI/litellm/pull/37346)
- refactor(ui): migrate the caching, cost tracking, alerting and user detail forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37350](https://github.com/BerriAI/litellm/pull/37350)
- refactor(ui): migrate pass-through, project and access group forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37354](https://github.com/BerriAI/litellm/pull/37354)
- refactor(ui): migrate the vector store creation form to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37353](https://github.com/BerriAI/litellm/pull/37353)
- refactor(ui): move the MCP server forms and detail tabs off tremor by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37329](https://github.com/BerriAI/litellm/pull/37329)
- feat(complexity\_router): custom classifier plugins via classifier\_type 'custom' by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37249](https://github.com/BerriAI/litellm/pull/37249)
- fix(fireworks): skip accounts/ rewrite for FW-\* Foundry deployment ids by [@&#8203;bruno-olivia](https://github.com/bruno-olivia) in [#&#8203;37242](https://github.com/BerriAI/litellm/pull/37242)
- fix(advisor): resolve the advisor sub-call through the proxy router by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36246](https://github.com/BerriAI/litellm/pull/36246)
- fix(router): forward target\_model\_names on file uploads to litellm\_proxy deployments by [@&#8203;shivamrawat1](https://github.com/shivamrawat1) in [#&#8203;36240](https://github.com/BerriAI/litellm/pull/36240)
- refactor(ui): move the internal user detail view off tremor by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37309](https://github.com/BerriAI/litellm/pull/37309)
- fix(responses): strip the responses/ routing prefix on the Responses API path by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37345](https://github.com/BerriAI/litellm/pull/37345)
- fix(main): forward store and prompt\_cache\_key params on chat completions by [@&#8203;Sujithr07](https://github.com/Sujithr07) in [#&#8203;33195](https://github.com/BerriAI/litellm/pull/33195)
- refactor(ui): migrate the model settings and credential rotation modals to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37342](https://github.com/BerriAI/litellm/pull/37342)
- refactor(ui): move the shared key form controls off antd onto shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37348](https://github.com/BerriAI/litellm/pull/37348)
- refactor(ui): move the tag and vector store views off tremor by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37311](https://github.com/BerriAI/litellm/pull/37311)
- refactor(ui): migrate SSO, SCIM and vault forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37347](https://github.com/BerriAI/litellm/pull/37347)
- test(ui): cover the edit project modal's required-field validation by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37363](https://github.com/BerriAI/litellm/pull/37363)
- refactor(ui): migrate the guardrail forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37364](https://github.com/BerriAI/litellm/pull/37364)
- refactor(ui): migrate agent forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37357](https://github.com/BerriAI/litellm/pull/37357)
- refactor(ui): migrate the MCP per-user env vars, toolset and tool arguments forms to react-hook-form and shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37349](https://github.com/BerriAI/litellm/pull/37349)
- fix(anthropic): emit tool\_use content\_block\_start without awaiting the next chunk by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37310](https://github.com/BerriAI/litellm/pull/37310)
- fix(proxy): send SSE keepalives while a slow upstream is still silent by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37322](https://github.com/BerriAI/litellm/pull/37322)
- fix(proxy): let org admins view their organization's usage by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37235](https://github.com/BerriAI/litellm/pull/37235)
- feat(vector\_stores): add Valkey as a managed vector store provider by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37002](https://github.com/BerriAI/litellm/pull/37002)
- refactor(ui): move the virtual key create and edit forms off tremor by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37324](https://github.com/BerriAI/litellm/pull/37324)
- refactor(ui): move the add model and credential forms off tremor by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37325](https://github.com/BerriAI/litellm/pull/37325)
- feat(team-callbacks): add DELETE /team/{team\_id}/callback/{callback\_name} by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37331](https://github.com/BerriAI/litellm/pull/37331)
- refactor(ui): move the teams page and team detail views off tremor by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37317](https://github.com/BerriAI/litellm/pull/37317)
- fix(cost\_calculator): recognize the ultrafast service tier in cost calculation by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37355](https://github.com/BerriAI/litellm/pull/37355)
- refactor(ui): move the cache settings and playground model selector off tremor by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37323](https://github.com/BerriAI/litellm/pull/37323)
- feat(guardrails): count bedrock guardrail cost against spend and budgets by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37362](https://github.com/BerriAI/litellm/pull/37362)
- refactor(ui): move the admin, SSO, SCIM, alerting and fallback forms off tremor by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37315](https://github.com/BerriAI/litellm/pull/37315)
- test(ui): raise vitest test and hook timeouts for CI headroom by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37370](https://github.com/BerriAI/litellm/pull/37370)
- fix(caching): truncate semantic cache embedding input, send extra\_body top-level by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37367](https://github.com/BerriAI/litellm/pull/37367)
- fix(guardrails): cap the date window accepted by /guardrails/usage endpoints by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37380](https://github.com/BerriAI/litellm/pull/37380)
- fix(ui): show select labels on the trigger instead of raw values by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37372](https://github.com/BerriAI/litellm/pull/37372)
- refactor(ui): move the team member search modal off antd Form by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37383](https://github.com/BerriAI/litellm/pull/37383)
- refactor(ui): move the model alias manager onto design tokens and shadcn controls by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37376](https://github.com/BerriAI/litellm/pull/37376)
- refactor(ui): move the MCP tool test form off antd by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37381](https://github.com/BerriAI/litellm/pull/37381)
- refactor(ui): move the model info view and pass-through endpoint forms off tremor by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37308](https://github.com/BerriAI/litellm/pull/37308)
- fix(otel): bound and shut down credential-scoped tracer providers by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36591](https://github.com/BerriAI/litellm/pull/36591)
- fix(proxy): send SSE keepalives on assistants runs and A2A streams by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37368](https://github.com/BerriAI/litellm/pull/37368)
- test(anthropic): pin one content\_block\_stop per tool\_use block on the Responses adapter by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37356](https://github.com/BerriAI/litellm/pull/37356)
- refactor(ui): move the agent, guardrail, prompt, policy and skill forms off tremor by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37320](https://github.com/BerriAI/litellm/pull/37320)
- fix(guardrails): requeue usage rollup rows dropped after retry exhaustion by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37387](https://github.com/BerriAI/litellm/pull/37387)
- feat(proxy): add project-level ITPM and OTPM quotas by [@&#8203;shivijain2323](https://github.com/shivijain2323) in [#&#8203;35110](https://github.com/BerriAI/litellm/pull/35110)
- feat(bedrock): add a config toggle to disable agent-runtime pass-through by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37386](https://github.com/BerriAI/litellm/pull/37386)
- fix(mcp): attach per-user BYOK credential when listing tools for non-oauth2 auth types by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34787](https://github.com/BerriAI/litellm/pull/34787)
- feat(ui): add success, warning and info status tokens by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37393](https://github.com/BerriAI/litellm/pull/37393)
- refactor(ui): drop [@&#8203;tremor/react](https://github.com/tremor/react) and the theming scaffolding it needed by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37394](https://github.com/BerriAI/litellm/pull/37394)
- refactor(ui): move the model info edit form off antd Form by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37392](https://github.com/BerriAI/litellm/pull/37392)
- fix(vector\_stores): stop leaking stored credentials in direct search debug logs by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37373](https://github.com/BerriAI/litellm/pull/37373)
- fix(logging): close three secret-leak paths in verbose logging by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37391](https://github.com/BerriAI/litellm/pull/37391)
- chore: bump litellm-enterprise 0.1.56 -> 0.1.57, litellm-proxy-extras 0.4.86 -> 0.4.87, litellm 1.98.0 -> 1.99.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37395](https://github.com/BerriAI/litellm/pull/37395)
- fix(mcp): bind tool existence check to the selected server by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37388](https://github.com/BerriAI/litellm/pull/37388)
- fix(mcp): serve token-forwarding servers when oauth discovery fails by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37399](https://github.com/BerriAI/litellm/pull/37399)
- fix(databricks): add cost map entries for 14 newer Databricks models by [@&#8203;epistoteles](https://github.com/epistoteles) in [#&#8203;28501](https://github.com/BerriAI/litellm/pull/28501)
- refactor(ui): style the logging settings from semantic tokens by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37385](https://github.com/BerriAI/litellm/pull/37385)
- feat(tinyfish): surface response headers + top-level response extras by [@&#8203;ChenluJi](https://github.com/ChenluJi) in [#&#8203;32448](https://github.com/BerriAI/litellm/pull/32448)
- refactor(ui): codemod the antd Tooltips outside form files onto the shadcn atom by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37402](https://github.com/BerriAI/litellm/pull/37402)
- test(ocr): update Azure DI supported-params assertion for req\_format by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37419](https://github.com/BerriAI/litellm/pull/37419)
- fix(proxy): return no rows when the aggregated activity entity filter is empty by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37414](https://github.com/BerriAI/litellm/pull/37414)
- test: build redaction and batch limiter fixtures the way production does by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37416](https://github.com/BerriAI/litellm/pull/37416)
- test: allow protocol-constrained pass-through routes to declare fewer methods by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37415](https://github.com/BerriAI/litellm/pull/37415)
- test(ui): pin the MCP server edit save payload before the form migration by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37404](https://github.com/BerriAI/litellm/pull/37404)
- test(ui): characterize the create key form payload contract by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37405](https://github.com/BerriAI/litellm/pull/37405)
- refactor(ui): migrate the key edit form off Ant Design onto react-hook-form by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37398](https://github.com/BerriAI/litellm/pull/37398)
- test(ui): repoint the e2e locators at the post-antd form controls by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37421](https://github.com/BerriAI/litellm/pull/37421)
- test: point the live gemini and groq conformance suites at models that still exist by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37422](https://github.com/BerriAI/litellm/pull/37422)
- refactor(ui): extract the create-key payload builder out of create\_key\_button by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37397](https://github.com/BerriAI/litellm/pull/37397)
- fix(types): map nested prompt\_tokens\_details.cache\_creation\_input\_tokens to cache\_write\_tokens by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37377](https://github.com/BerriAI/litellm/pull/37377)
- fix(router): honor key-level tag filtering in pre-routing and pin auto-router e2e regressions by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37366](https://github.com/BerriAI/litellm/pull/37366)
- feat(otel): attribute Prisma database spans to PostgreSQL instead of localhost by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36595](https://github.com/BerriAI/litellm/pull/36595)
- fix(bedrock): degrade gracefully on malformed tool-call arguments by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33842](https://github.com/BerriAI/litellm/pull/33842)
- test: move the remaining live groq call sites off the retired llama models by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37426](https://github.com/BerriAI/litellm/pull/37426)
- fix(ui): highlight the first member search match so Enter picks it by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37429](https://github.com/BerriAI/litellm/pull/37429)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37400](https://github.com/BerriAI/litellm/pull/37400)
- fix(proxy): log spend for OpenAI passthrough embeddings with unmapped models by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37425](https://github.com/BerriAI/litellm/pull/37425)
- fix(router): keep acreate\_file fallbacks inside the requested model group by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37424](https://github.com/BerriAI/litellm/pull/37424)
- fix(proxy): record estimated input tokens in spend logs for failed dispatched requests by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37365](https://github.com/BerriAI/litellm/pull/37365)
- fix: accept bool thinking param instead of crashing with AttributeError by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37423](https://github.com/BerriAI/litellm/pull/37423)
- refactor(ui): migrate the teams form graph off antd Form onto react-hook-form by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37417](https://github.com/BerriAI/litellm/pull/37417)
- fix(ui): restore the cache control Role and Index field hints by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37437](https://github.com/BerriAI/litellm/pull/37437)
- feat(ui): add mounted-field projections for the MCP server form graph by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37440](https://github.com/BerriAI/litellm/pull/37440)
- refactor(ui): extract the MCP server edit save payload into a pure builder by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37436](https://github.com/BerriAI/litellm/pull/37436)
- test: derive vertex batch cost expectation from the cost map by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37444](https://github.com/BerriAI/litellm/pull/37444)
- refactor(ui): port the create key form off antd Form onto react-hook-form by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37442](https://github.com/BerriAI/litellm/pull/37442)
- refactor(ui): port the add model form off antd Form onto react-hook-form by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37446](https://github.com/BerriAI/litellm/pull/37446)
- refactor(ui): host KeyLifecycleSettings tests in react-hook-form instead of antd Form by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37449](https://github.com/BerriAI/litellm/pull/37449)
- fix(ui): rebuild nested and list paths in the mounted-field projection by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37450](https://github.com/BerriAI/litellm/pull/37450)
- fix(ui): gate the pass-through guardrail field inputs when the section is disabled by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37435](https://github.com/BerriAI/litellm/pull/37435)
- fix(tests): keep a host PROXY\_BASE\_URL out of request-derived URL tests by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37451](https://github.com/BerriAI/litellm/pull/37451)
- refactor(ui): port the MCP server forms off antd Form onto react-hook-form by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37483](https://github.com/BerriAI/litellm/pull/37483)
- test(e2e): pin the tag-routing denial to its actual cause by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37432](https://github.com/BerriAI/litellm/pull/37432)
- fix(proxy): read through to the DB on registry misses so just-created models, guardrails, and agents resolve on sibling replicas by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36263](https://github.com/BerriAI/litellm/pull/36263)
- fix(mcp): forward the per-server auth header on OpenAPI tool calls by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37410](https://github.com/BerriAI/litellm/pull/37410)
- chore(typing): drop 1.3k basedpyright errors across 42 Any hotspot files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37439](https://github.com/BerriAI/litellm/pull/37439)
- test(ui): drive fields with change events where the typing is not the behaviour by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37495](https://github.com/BerriAI/litellm/pull/37495)
- fix(ui): restore tab strip styling and panel persistence lost in the shadcn migration by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37403](https://github.com/BerriAI/litellm/pull/37403)
- feat(spend-logs): add lifecycle timestamps by [@&#8203;sytianhe](https://github.com/sytianhe) in [#&#8203;37361](https://github.com/BerriAI/litellm/pull/37361)
- refactor(ui): migrate the antd Button call sites onto the shadcn Button by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37505](https://github.com/BerriAI/litellm/pull/37505)
- test(ui): split the vitest suite into unit, component, integration and type projects by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37488](https://github.com/BerriAI/litellm/pull/37488)
- refactor(ptu): give the rollup a source-agnostic deployment record by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37501](https://github.com/BerriAI/litellm/pull/37501)
- feat(auto-router)!: scope shadow eval jobs to multiple keys by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37251](https://github.com/BerriAI/litellm/pull/37251)
- refactor(ui): migrate the antd Alert call sites onto the shared Alert by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37513](https://github.com/BerriAI/litellm/pull/37513)
- chore(ui): upgrade the dashboard to React 19 by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37411](https://github.com/BerriAI/litellm/pull/37411)
- fix(streaming): accept provider cost objects when propagating usage cost by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36593](https://github.com/BerriAI/litellm/pull/36593)
- fix(complexity-router): gate the reasoning override on a non-SIMPLE score by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37500](https://github.com/BerriAI/litellm/pull/37500)
- fix(mcp): stop reporting failed OpenAPI tool calls as successes by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37496](https://github.com/BerriAI/litellm/pull/37496)
- feat(e2e): add record/replay transport seam and fixture bundle format by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37360](https://github.com/BerriAI/litellm/pull/37360)
- fix(proxy): accept inherited model sentinels in project key limits by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37515](https://github.com/BerriAI/litellm/pull/37515)
- fix(model\_prices): add provider-announced deprecation\_date to 205 registry entries by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37283](https://github.com/BerriAI/litellm/pull/37283)
- fix(model\_prices): correct gemini 3.1 flash image and deepseek v4 pricing, add openai deprecation dates by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37473](https://github.com/BerriAI/litellm/pull/37473)
- fix(model\_prices): set prompt\_cache\_min\_tokens=4096 for Gemini 3.5/3.6/3.7 Flash and 3.1 Pro Preview by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37516](https://github.com/BerriAI/litellm/pull/37516)
- fix(anthropic,bedrock): report provider thinking tokens instead of classifying them as text by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35998](https://github.com/BerriAI/litellm/pull/35998)
- fix(batches): stop one bad output line from zeroing an entire batch's spend by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37457](https://github.com/BerriAI/litellm/pull/37457)
- feat(e2e): canonical content-based match keys for record-and-replay by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37525](https://github.com/BerriAI/litellm/pull/37525)
- feat(cli): add `lite login --config-claude` to wire Claude Code at login by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37507](https://github.com/BerriAI/litellm/pull/37507)
- fix(auth): resolve bare model names against wildcard deployments in model access groups by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37492](https://github.com/BerriAI/litellm/pull/37492)
- docs: run only the tests covering your change, leave suites to CI by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37528](https://github.com/BerriAI/litellm/pull/37528)
- feat(complexity-router): make the reasoning override floor configurable by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37537](https://github.com/BerriAI/litellm/pull/37537)
- fix(ui): drop stale user search answers so Enter commits the current match by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37504](https://github.com/BerriAI/litellm/pull/37504)
- refactor(ui): migrate the remaining dashboard pages off antd by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37524](https://github.com/BerriAI/litellm/pull/37524)
- feat(proxy): auto-suppress the no-Redis banner for confirmed single-worker deployments by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36987](https://github.com/BerriAI/litellm/pull/36987)
- fix(proxy): retry spend updates on Postgres deadlock instead of dropping them by [@&#8203;RayJueWang](https://github.com/RayJueWang) in [#&#8203;34887](https://github.com/BerriAI/litellm/pull/34887)
- feat(search): add Amazon Bedrock AgentCore web search provider by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36331](https://github.com/BerriAI/litellm/pull/36331)
- fix(helm): default litellm-helm to the ghcr.io/berriai/litellm image by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37491](https://github.com/BerriAI/litellm/pull/37491)
- refactor(ui): migrate antd Modal onto the shared shadcn Dialog by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37540](https://github.com/BerriAI/litellm/pull/37540)
- fix(ui): toggle unlimited budget when its text is clicked by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37547](https://github.com/BerriAI/litellm/pull/37547)
- feat(proxy): fast-fail validation for batch input files at /v1/files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37527](https://github.com/BerriAI/litellm/pull/37527)
- chore: gitignore CLAUDE.local.md by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37545](https://github.com/BerriAI/litellm/pull/37545)
- perf(otel): build the credential-scoped tracer Resource once per logger by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37542](https://github.com/BerriAI/litellm/pull/37542)
- fix(ci): gate backend unit tests on the pull request's own file list by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37550](https://github.com/BerriAI/litellm/pull/37550)
- chore(codeowners): require pricing owner approval for the model prices jsons by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37551](https://github.com/BerriAI/litellm/pull/37551)
- fix(proxy): initialize the secret manager before resolving os.environ config references by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37544](https://github.com/BerriAI/litellm/pull/37544)
- refactor(ui): swap [@&#8203;ant-design/icons](https://github.com/ant-design/icons) for lucide-react by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37553](https://github.com/BerriAI/litellm/pull/37553)
- refactor(ui): migrate shared primitives and common components off antd by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37521](https://github.com/BerriAI/litellm/pull/37521)
- fix(spend-logs): backfill created\_at/updated\_at from row endTime instead of migration time by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37554](https://github.com/BerriAI/litellm/pull/37554)
- refactor(ui): migrate the MCP servers pages off antd by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37522](https://github.com/BerriAI/litellm/pull/37522)
- fix(vertex\_ai): apply regional endpoint uplift to cost tracking by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37543](https://github.com/BerriAI/litellm/pull/37543)
- fix(proxy): populate deployment attribution on failed-request spend logs by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37520](https://github.com/BerriAI/litellm/pull/37520)
- refactor(ui): migrate the model and router settings pages off antd by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37523](https://github.com/BerriAI/litellm/pull/37523)
- perf(ci): gate the lint, MCP and dashboard jobs on the pull request's file list by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37559](https://github.com/BerriAI/litellm/pull/37559)
- feat(router): allow per-tier litellm\_params in complexity autorouter config by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37064](https://github.com/BerriAI/litellm/pull/37064)
- fix(anthropic): log partial stream spend when a /v1/messages client disconnects mid-stream by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37558](https://github.com/BerriAI/litellm/pull/37558)
- fix(ui): clear pass-through header rows when the create modal is reopened by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37549](https://github.com/BerriAI/litellm/pull/37549)
- fix(ui): render optional array and object MCP tool parameters as JSON inputs by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37548](https://github.com/BerriAI/litellm/pull/37548)
- feat(proxy)!: default audit logs on for enterprise licenses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37518](https://github.com/BerriAI/litellm/pull/37518)
- feat(ui): standardize the Teams page header by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36897](https://github.com/BerriAI/litellm/pull/36897)
- feat(ptu): accrue flat cost for PTU deployments declared in config.yaml by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37556](https://github.com/BerriAI/litellm/pull/37556)
- refactor(ui): migrate the last antd components off antd onto shadcn by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37569](https://github.com/BerriAI/litellm/pull/37569)
- chore(ui): drop the antd dependency and its leftovers by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37574](https://github.com/BerriAI/litellm/pull/37574)
- refactor(ui): map hardcoded Tailwind palette classes onto semantic tokens by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37576](https://github.com/BerriAI/litellm/pull/37576)
- fix(ptu): hand the prune a plain delete filter the query builder can serialise by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37571](https://github.com/BerriAI/litellm/pull/37571)
- feat(proxy): enqueued-token rate limiting for batches with refund on completion and cancellation by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37539](https://github.com/BerriAI/litellm/pull/37539)
- fix(ui): restore hover feedback and dark-mode variants lost in the token migration by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37579](https://github.com/BerriAI/litellm/pull/37579)
- fix(ci): run the full dashboard suite when a change reaches outside src/ by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37563](https://github.com/BerriAI/litellm/pull/37563)
- fix(ui): keep semantic button colours on hover after the no-op hover cleanup by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37580](https://github.com/BerriAI/litellm/pull/37580)
- fix: add supports\_mid\_conversation\_system to bare first-party Claude cost-map keys by [@&#8203;oneKn8](https://github.com/oneKn8) in [#&#8203;36969](https://github.com/BerriAI/litellm/pull/36969)
- feat: add bedrock grok 4.6 to model cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37517](https://github.com/BerriAI/litellm/pull/37517)
- fix: preserve prompt cache for mid-conversation system on unflagged Claude models by [@&#8203;oneKn8](https://github.com/oneKn8) in [#&#8203;36968](https://github.com/BerriAI/litellm/pull/36968)
- fix(router): routed deployment's own litellm\_params beat forwarded auto\_router marker params by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37615](https://github.com/BerriAI/litellm/pull/37615)
- refactor(ci): fold the nine thin unit-shard callers into one matrix by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37590](https://github.com/BerriAI/litellm/pull/37590)
- chore(ci): close the test-census blind spots and move scripts out of workflows/ by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37586](https://github.com/BerriAI/litellm/pull/37586)
- test: retire tests/old\_proxy\_tests, which holds no tests by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37605](https://github.com/BerriAI/litellm/pull/37605)
- feat(ci): ratchet the test suite's zero-assert, mock-echo and global-state debt by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37588](https://github.com/BerriAI/litellm/pull/37588)
- feat(proxy): native CLI login with OAuth authorization code + PKCE by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37626](https://github.com/BerriAI/litellm/pull/37626)
- feat(prompt-caching): map cache\_control\_injection\_points to OpenAI prompt\_cache\_breakpoint on GPT-5.6+ targets by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37628](https://github.com/BerriAI/litellm/pull/37628)
- fix(realtime): bound Vertex credential resolution and make realtime failures loud by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37604](https://github.com/BerriAI/litellm/pull/37604)
- fix(anthropic): map metadata.user\_id to prompt\_cache\_key on the /v1/messages bridge by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37623](https://github.com/BerriAI/litellm/pull/37623)
- fix(passthrough): resolve vertex live credentials from db model deployments by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37602](https://github.com/BerriAI/litellm/pull/37602)
- fix(prompt\_management): don't route no-prompt\_id requests to prompt managers that can't run them by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37575](https://github.com/BerriAI/litellm/pull/37575)
- test: remove the five test functions a later definition shadows by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37591](https://github.com/BerriAI/litellm/pull/37591)
- feat(ci): guard shard assignment across every sharded test tree by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37593](https://github.com/BerriAI/litellm/pull/37593)
- feat(ui): multi-key shadow eval picker and per-key breakdown by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37389](https://github.com/BerriAI/litellm/pull/37389)
- fix(ui): make dark-mode form controls visible by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37648](https://github.com/BerriAI/litellm/pull/37648)
- fix(ui): give status colours a readable foreground and drop the muted 70% step by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37649](https://github.com/BerriAI/litellm/pull/37649)
- fix(ui): make inline styles and code blocks follow the theme by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37651](https://github.com/BerriAI/litellm/pull/37651)
- fix(ui): move the policy flow builder onto theme tokens by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37654](https://github.com/BerriAI/litellm/pull/37654)
- test: settle three allowlist entries that were open questions by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37598](https://github.com/BerriAI/litellm/pull/37598)
- feat(ci): ratchet tests that skip themselves when a credential is absent by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37612](https://github.com/BerriAI/litellm/pull/37612)
- feat(ci): catch files a -k expression deselects from every job by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37601](https://github.com/BerriAI/litellm/pull/37601)
- test: run the 30 test files stranded in the second mirror by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37595](https://github.com/BerriAI/litellm/pull/37595)
- fix(ui): make hardcoded palette surfaces theme-aware by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37650](https://github.com/BerriAI/litellm/pull/37650)
- feat(cli): store the lite login credential in the OS keychain by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37566](https://github.com/BerriAI/litellm/pull/37566)
- fix(ui): draw one Per Day savings bar per date on Cost Optimization by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37643](https://github.com/BerriAI/litellm/pull/37643)
- feat(mistral): add zai-glm-5-2 and glm-5-2 model pricing by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;37110](https://github.com/BerriAI/litellm/pull/37110)
- feat(complexity\_router): add business classification rubric preset by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37534](https://github.com/BerriAI/litellm/pull/37534)
- feat(ui): serve a dark-mode variant of the LiteLLM logo by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37656](https://github.com/BerriAI/litellm/pull/37656)
- fix(otel): route Phoenix traces to per-key/team projects under otel v2 by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36706](https://github.com/BerriAI/litellm/pull/36706)
- test: replace blind sleeps with deadline waits in callback and caching tests by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37660](https://github.com/BerriAI/litellm/pull/37660)
- fix(cli): keep the --pkce refresh token in the OS keychain, not in token.json by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37665](https://github.com/BerriAI/litellm/pull/37665)
- fix(ui): keep keyword tier rules that target operator-defined tiers when hydrating the edit modal by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37413](https://github.com/BerriAI/litellm/pull/37413)
- feat(proxy): add POST /auto\_router/validate\_complexity\_router\_config to dry-run the complexity-router write gate by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37409](https://github.com/BerriAI/litellm/pull/37409)
- feat(ui): let admins supply a dark-mode variant of their custom logo by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37662](https://github.com/BerriAI/litellm/pull/37662)
- fix(proxy): run pre-call guardrails on batch input file uploads by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37519](https://github.com/BerriAI/litellm/pull/37519)
- feat(ui): add a light/dark/system theme toggle to the top bar by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37669](https://github.com/BerriAI/litellm/pull/37669)
- feat(proxy): redact or drop individual batch records instead of rejecting the file by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37561](https://github.com/BerriAI/litellm/pull/37561)
- refactor(ui): mark dark as beta in the theme menu instead of the toolbar by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37680](https://github.com/BerriAI/litellm/pull/37680)
- ci: lint the test tree for undefined names (F821) and fix all 30 by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37671](https://github.com/BerriAI/litellm/pull/37671)
- feat(e2e): move record/replay to the provider edge (LIT-5745) by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37565](https://github.com/BerriAI/litellm/pull/37565)
- fix(mcp): let a salt-key-orphaned OAuth credential be replaced by re-authorization by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37672](https://github.com/BerriAI/litellm/pull/37672)
- fix(mcp): normalize auth schemes so MCP egress emits exactly one prefix by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37668](https://github.com/BerriAI/litellm/pull/37668)
- test: add six ruff rules that catch tests which cannot fail by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37709](https://github.com/BerriAI/litellm/pull/37709)
- perf(ci): measure unit-shard coverage with the sys.monitoring core by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37589](https://github.com/BerriAI/litellm/pull/37589)
- test: merge three stranded twins into the files that shadow them by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37600](https://github.com/BerriAI/litellm/pull/37600)
- test(ci): reject coverage-allowlist entries that no longer match a file by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37608](https://github.com/BerriAI/litellm/pull/37608)
- feat(ci): assert .github/workflows holds only workflows, correctly named by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37616](https://github.com/BerriAI/litellm/pull/37616)
- feat(ci): freeze the conftest save/restore inventory so it can only shrink by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37621](https://github.com/BerriAI/litellm/pull/37621)
- fix(a2a): accept the whole JSON-RPC id union the spec defines by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37704](https://github.com/BerriAI/litellm/pull/37704)
- fix(ptu): refuse an incomplete config.yaml reservation the way the endpoints do by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37703](https://github.com/BerriAI/litellm/pull/37703)
- chore: bump litellm-enterprise 0.1.57 -> 0.1.58, litellm-proxy-extras 0.4.87 -> 0.4.88 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37717](https://github.com/BerriAI/litellm/pull/37717)
- feat(perplexity): add Agent API third-party models by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;37112](https://github.com/BerriAI/litellm/pull/37112)
- fix(ui): surface the paginated fallback on Cost Optimization by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37659](https://github.com/BerriAI/litellm/pull/37659)
- feat(shadow\_eval)!: gate the per-key budget on dollar spend instead of turns by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37555](https://github.com/BerriAI/litellm/pull/37555)
- feat(ui): per-model reasoning effort in the complexity tier editor by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;37673](https://github.com/Berr…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants