Repository navigation
fix(azure): rename max_tokens to max_completion_tokens for gpt-5-chat deployments - #36857
Conversation
|
|
Greptile SummaryThe PR updates Azure GPT-5 chat parameter translation so legacy
Confidence Score: 5/5The PR appears safe to merge. No blocking failure remains.
|
| Filename | Overview |
|---|---|
| litellm/llms/azure/chat/gpt_transformation.py | Adds an isolated Azure request-parameter translation for GPT-5 chat deployments while preserving existing reasoning-model classification. |
| tests/test_litellm/llms/azure/chat/test_azure_chat_gpt_transformation.py | Adds focused regression tests covering token-key emission and GPT-5 chat versus reasoning behavior. |
Reviews (3): Last reviewed commit: "fix(azure): rename max_tokens to max_com..." | Re-trigger Greptile
| name family needs the rename, including the ``gpt-5-chat*`` models that are excluded from | ||
| the reasoning path by https://github.com/BerriAI/litellm/issues/13781. | ||
| """ | ||
| return "gpt-5" in model or "gpt5_series" in model |
There was a problem hiding this comment.
Hardcoded model capability predicate
The new predicate infers token-parameter support exclusively from gpt-5 and gpt5_series substrings instead of model metadata. This requires LiteLLM releases for future capability changes and lets naming aliases select the wrong outbound parameter; represent the capability in model metadata and access it through get_model_info.
Rule Used: What: Do not hardcode model-specific flags in the ... (source)
Knowledge Base Used: LLM Provider Adapters
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| for param, value in non_default_params.items(): | ||
| if param == "tool_choice": | ||
| if param == "max_tokens" and renames_max_tokens: | ||
| optional_params.setdefault("max_completion_tokens", value) |
There was a problem hiding this comment.
Parameter mapping mutates input
Calling setdefault directly mutates the caller-provided optional_params mapping, adding another in-place state change that makes ownership and parameter precedence harder to reason about. Construct and return the updated mapping without mutating the function parameter.
Context Used: CLAUDE.md (source)
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 1 · PR risk: 0/10 |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
@greptileai adjudicated both findings in the PR description under Review notes: no such model-map flag exists, and optional_params is a local accumulator. Head 1e877f2 |
|
Both adjudications hold up. Model-map flag: Correct that no suitable flag exists.
The test suite is thorough: it covers the new family ( |
1e877f2 to
29d8bed
Compare
|
@greptileai rebased onto current staging, no code change since the last review. Please re-review head 29d8bed. Rebuttal stands in the PR description under Review notes |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 29d8bed. Configure here.
… deployments Azure rejects the legacy `max_tokens` key for the whole gpt-5 name family, but `AzureOpenAIGPT5Config.is_model_gpt_5_model` deliberately excludes `gpt-5-chat*` so those deployments fall through to `AzureOpenAIConfig`, which sends `max_tokens` verbatim and gets a 400 back on every request that carries it, `/health` probes included. One predicate was answering two independent questions. Split it: the new `AzureOpenAIConfig.requires_max_completion_tokens` covers the whole gpt-5 name family and drives only the rename, while `is_model_gpt_5_model` keeps keying reasoning_effort, the temperature clamp and the dropped penalties off the reasoning question, so #13781 stays fixed.
29d8bed to
5b0ea66
Compare
|
@greptileai please review the current head 5b0ea66. Only change since your 5/5: rebase onto staging, resolving a test-file conflict with #34462 |
…9.1) (#104)
This PR contains the following updates:
| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.98.0` → `v1.99.1` |
---
### Release Notes
<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>
### [`v1.99.1`](https://github.com/BerriAI/litellm/releases/tag/v1.99.1)
[Compare Source](https://github.com/BerriAI/litellm/compare/v1.99.0...v1.99.1)
#### Docker-only release
**This release ships container images only. There is no PyPI package for `1.99.1`.**
`pip install litellm==1.99.1` will not resolve — install the images below, or stay on `1.99.0` on PyPI. The git tag and this release exist so the images are traceable to an exact commit.
| Image | Tags |
| ------------------------------------------------------------------------- | ------------------- |
| `ghcr.io/berriai/litellm` · `docker.io/litellm/litellm` | `1.99.1`, `v1.99.1` |
| `ghcr.io/berriai/litellm-database` · `docker.io/litellm/litellm-database` | `1.99.1`, `v1.99.1` |
| `ghcr.io/berriai/litellm-non_root` · `docker.io/litellm/litellm-non_root` | `1.99.1`, `v1.99.1` |
This is the newest stable image, so the rolling `latest` and `main-stable` image tags now point at `1.99.1`.
It carries one fix on top of `1.99.0`: OpenTelemetry v2 spans now emit cache token counts (`gen_ai.usage.cache_creation.input_tokens` and `gen_ai.usage.cache_read.input_tokens`) alongside the cache cost that was already reported. If you compute spend from OTel token counts rather than from LiteLLM's own cost fields, prompt-caching workloads were previously under-counted.
***
#### Verify Docker Image Signature
All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).
**Verify using the pinned commit hash (recommended):**
A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
ghcr.io/berriai/litellm:v1.99.1
```
**Verify using the release tag (convenience):**
Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/v1.99.1/cosign.pub \
ghcr.io/berriai/litellm:v1.99.1
```
Expected output:
```
The following checks were performed on each of these signatures:
- The cosign claims were validated
- The signatures were verified against the specified public key
```
***
#### What's Changed
- chore(release): backport [#​38716](https://github.com/BerriAI/litellm/issues/38716) to stable/1.99.x and cut 1.99.1 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​39179](https://github.com/BerriAI/litellm/pull/39179)
**Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.99.0...v1.99.1>
### [`v1.99.0`](https://github.com/BerriAI/litellm/releases/tag/v1.99.0)
[Compare Source](https://github.com/BerriAI/litellm/compare/v1.98.0...v1.99.0)
#### Verify Docker Image Signature
All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).
**Verify using the pinned commit hash (recommended):**
A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
ghcr.io/berriai/litellm:v1.99.0
```
**Verify using the release tag (convenience):**
Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/v1.99.0/cosign.pub \
ghcr.io/berriai/litellm:v1.99.0
```
Expected output:
```
The following checks were performed on each of these signatures:
- The cosign claims were validated
- The signatures were verified against the specified public key
```
***
#### What's Changed
- chore(typing): drop 1.3k basedpyright errors across 30 Any hotspot files by [@​mateo-berri](https://github.com/mateo-berri) in [#​37073](https://github.com/BerriAI/litellm/pull/37073)
- fix(proxy): register WebSocket passthrough for OpenAI prefixes by [@​LHMQ878](https://github.com/LHMQ878) in [#​36151](https://github.com/BerriAI/litellm/pull/36151)
- fix(bedrock): report uploaded size in the FileObject returned by managed batch uploads by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36392](https://github.com/BerriAI/litellm/pull/36392)
- fix(batches): support AWS Bedrock batch cancellation via `StopModelInvocationJob` by [@​ArjunPakhan](https://github.com/ArjunPakhan) in [#​34087](https://github.com/BerriAI/litellm/pull/34087)
- feat: Async Rust OCR Bridge and MCP OAuth UI Restore by [@​ArjunPakhan](https://github.com/ArjunPakhan) in [#​31453](https://github.com/BerriAI/litellm/pull/31453)
- fix(batches): don't crash logging when a completed batch has no output file by [@​MUSE-CODE-SPACE](https://github.com/MUSE-CODE-SPACE) in [#​34067](https://github.com/BerriAI/litellm/pull/34067)
- fix(UI): add default model pin to complexity router UI by [@​tin-berri](https://github.com/tin-berri) in [#​36615](https://github.com/BerriAI/litellm/pull/36615)
- feat(ui): add Lite mixed-provider auto-router preset by [@​tin-berri](https://github.com/tin-berri) in [#​37068](https://github.com/BerriAI/litellm/pull/37068)
- feat(ui): link key info header to its user, creator, team, and organization by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37187](https://github.com/BerriAI/litellm/pull/37187)
- fix(guardrails): scan text on /guardrails/apply\_guardrail for Azure Content Safety by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36894](https://github.com/BerriAI/litellm/pull/36894)
- feat(bedrock): forward LiteLLM identity and metadata into Bedrock requestMetadata by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36861](https://github.com/BerriAI/litellm/pull/36861)
- fix(azure): rename max\_tokens to max\_completion\_tokens for gpt-5-chat deployments by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36857](https://github.com/BerriAI/litellm/pull/36857)
- fix(bedrock): preserve cache token usage when invocationMetrics replace the usage block by [@​brian5021](https://github.com/brian5021) in [#​36878](https://github.com/BerriAI/litellm/pull/36878)
- fix(proxy): registry caches stop per-request tag and end-user Postgres reads in auth by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36801](https://github.com/BerriAI/litellm/pull/36801)
- test(e2e): replay a real tool-search assistant turn back to Bedrock Invoke by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36856](https://github.com/BerriAI/litellm/pull/36856)
- fix(proxy): return 400 naming the missing required param on POST /v1/batches by [@​mateo-berri](https://github.com/mateo-berri) in [#​37199](https://github.com/BerriAI/litellm/pull/37199)
- fix(ci): bump sqlparse to 0.6.0 to resolve osv-scan CVEs by [@​mateo-berri](https://github.com/mateo-berri) in [#​37200](https://github.com/BerriAI/litellm/pull/37200)
- fix(ui): stop pairing key spend with the team budget when a key has no budget by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37196](https://github.com/BerriAI/litellm/pull/37196)
- fix(guardrails): record MCP tool guardrail evaluations and blocks in … by [@​Scott-Wilson-ZocDoc](https://github.com/Scott-Wilson-ZocDoc) in [#​36978](https://github.com/BerriAI/litellm/pull/36978)
- fix(proxy): return 400 for non-object metadata and litellm\_metadata instead of silent drop or 500 by [@​mateo-berri](https://github.com/mateo-berri) in [#​37203](https://github.com/BerriAI/litellm/pull/37203)
- fix(anthropic): preserve optional Responses tool properties by [@​Scott-Wilson-ZocDoc](https://github.com/Scott-Wilson-ZocDoc) in [#​36979](https://github.com/BerriAI/litellm/pull/36979)
- feat(ui): add user ID request log filter by [@​daniel-meismer-zocdoc](https://github.com/daniel-meismer-zocdoc) in [#​36781](https://github.com/BerriAI/litellm/pull/36781)
- fix(anthropic): stop emitting empty thinking blocks on the Responses adapter by [@​Scott-Wilson-ZocDoc](https://github.com/Scott-Wilson-ZocDoc) in [#​36033](https://github.com/BerriAI/litellm/pull/36033)
- fix(ui): make per-user usage filter searchable by [@​daniel-meismer-zocdoc](https://github.com/daniel-meismer-zocdoc) in [#​36790](https://github.com/BerriAI/litellm/pull/36790)
- refactor(ui): decouple bulk invite from the invite user button by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37061](https://github.com/BerriAI/litellm/pull/37061)
- fix(helm): bound the migrations Job so a blocked migration cannot stall the release by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36975](https://github.com/BerriAI/litellm/pull/36975)
- feat(proxy): let USE\_V2\_MIGRATION\_RESOLVER select the v2 migration resolver by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36258](https://github.com/BerriAI/litellm/pull/36258)
- fix(mcp): scope authorization server issuer by [@​irosh-colombage-ZocDoc2](https://github.com/irosh-colombage-ZocDoc2) in [#​36482](https://github.com/BerriAI/litellm/pull/36482)
- fix(responses): unwrap object-form tool\_choice before calling the Responses API by [@​Scott-Wilson-ZocDoc](https://github.com/Scott-Wilson-ZocDoc) in [#​36032](https://github.com/BerriAI/litellm/pull/36032)
- test(ui): query antd controls accessibly instead of by internal CSS class by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37014](https://github.com/BerriAI/litellm/pull/37014)
- fix(proxy): bill cancelled and failed batches that still produced an output file by [@​mateo-berri](https://github.com/mateo-berri) in [#​37205](https://github.com/BerriAI/litellm/pull/37205)
- fix(bedrock): read batch usage by payload shape, not by provider name by [@​marty-sullivan](https://github.com/marty-sullivan) in [#​37078](https://github.com/BerriAI/litellm/pull/37078)
- fix(ui): self-contained searchable user filter on the Usage page by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37206](https://github.com/BerriAI/litellm/pull/37206)
- revert: don't fix mcp scope authorization server issuer by [@​mateo-berri](https://github.com/mateo-berri) in [#​37220](https://github.com/BerriAI/litellm/pull/37220)
- fix(mcp): scope authorization server issuer for named MCP servers by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​37204](https://github.com/BerriAI/litellm/pull/37204)
- test(ui): gate dashboard test assertions with testing-library and jest-dom rules by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37018](https://github.com/BerriAI/litellm/pull/37018)
- feat(shadow-eval): name the shadowed key in job responses and the UI headline by [@​tin-berri](https://github.com/tin-berri) in [#​37221](https://github.com/BerriAI/litellm/pull/37221)
- test(ui): assert what collaborators are called with, not merely that they were by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37019](https://github.com/BerriAI/litellm/pull/37019)
- fix(logging): stop deepcopying results redaction cannot redact by [@​marty-sullivan](https://github.com/marty-sullivan) in [#​36638](https://github.com/BerriAI/litellm/pull/36638)
- fix(gemini): price gemini 3.6 flash at Google's introductory rates on every service tier by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37197](https://github.com/BerriAI/litellm/pull/37197)
- perf(guardrails): stop sending the conversation twice in the noma v2 payload by [@​itaimodi](https://github.com/itaimodi) in [#​36764](https://github.com/BerriAI/litellm/pull/36764)
- fix(streaming): track provider-reported cost when caller omits include\_usage by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35013](https://github.com/BerriAI/litellm/pull/35013)
- fix: stop rust flag from leaking into upstream provider request bodies by [@​mateo-berri](https://github.com/mateo-berri) in [#​37218](https://github.com/BerriAI/litellm/pull/37218)
- fix(proxy): return 404 instead of 500 for unresolvable batch and file ids on /v1/batches by [@​mateo-berri](https://github.com/mateo-berri) in [#​37201](https://github.com/BerriAI/litellm/pull/37201)
- fix(bedrock): validate file-content retrieval against the configured output bucket ([#​26335](https://github.com/BerriAI/litellm/issues/26335)) by [@​kingdoooo](https://github.com/kingdoooo) in [#​31435](https://github.com/BerriAI/litellm/pull/31435)
- fix(proxy): reject out-of-range limit on GET /v1/batches with OpenAI-parity 400 by [@​mateo-berri](https://github.com/mateo-berri) in [#​37198](https://github.com/BerriAI/litellm/pull/37198)
- fix(batches): price a retrieved batch from its deployment's model and rates (internal copy of [#​37077](https://github.com/BerriAI/litellm/issues/37077)) by [@​mateo-berri](https://github.com/mateo-berri) in [#​37219](https://github.com/BerriAI/litellm/pull/37219)
- feat(ocr): return Azure Document Intelligence's native payload from /v1/ocr via req\_format=native by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37194](https://github.com/BerriAI/litellm/pull/37194)
- fix(anthropic): fold guardrail-modified leading system rows into top-level system param by [@​mateo-berri](https://github.com/mateo-berri) in [#​37231](https://github.com/BerriAI/litellm/pull/37231)
- fix(shadow\_eval): copy messages before router call and raise judge output cap by [@​tin-berri](https://github.com/tin-berri) in [#​37232](https://github.com/BerriAI/litellm/pull/37232)
- feat(proxy): add Amazon Comprehend Medical passthrough provider by [@​mateo-berri](https://github.com/mateo-berri) in [#​37229](https://github.com/BerriAI/litellm/pull/37229)
- test(ui): settle the in-flight search before the loading tests end by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37227](https://github.com/BerriAI/litellm/pull/37227)
- test(cli): use example.com placeholder host in base-url trailing slash test by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37240](https://github.com/BerriAI/litellm/pull/37240)
- feat(complexity\_router): operator-defined tier sets for the LLM classifier by [@​tin-berri](https://github.com/tin-berri) in [#​37226](https://github.com/BerriAI/litellm/pull/37226)
- feat(ui): configure the auto router's heuristic scorer from the Admin UI by [@​tin-berri](https://github.com/tin-berri) in [#​37216](https://github.com/BerriAI/litellm/pull/37216)
- fix(shadow\_eval): drop unused judge reasoning field and salvage truncated verdicts by [@​tin-berri](https://github.com/tin-berri) in [#​37239](https://github.com/BerriAI/litellm/pull/37239)
- feat(proxy): proactive model deprecation alerts and `/model/deprecations` endpoint by [@​mateo-berri](https://github.com/mateo-berri) in [#​26900](https://github.com/BerriAI/litellm/pull/26900)
- refactor(ui): move dashboard toasts from antd message/notification onto sonner by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37207](https://github.com/BerriAI/litellm/pull/37207)
- feat(guardrails): track bedrock guardrail usage units per invocation by [@​mateo-berri](https://github.com/mateo-berri) in [#​37225](https://github.com/BerriAI/litellm/pull/37225)
- fix(proxy): strip callback credentials from the auth object stamped into request metadata by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​37233](https://github.com/BerriAI/litellm/pull/37233)
- fix(guardrails): retry usage upserts only on connection errors by [@​mateo-berri](https://github.com/mateo-berri) in [#​37247](https://github.com/BerriAI/litellm/pull/37247)
- fix(mcp): oauth discovery must not cause outages by [@​daniel-meismer-zocdoc](https://github.com/daniel-meismer-zocdoc) in [#​36599](https://github.com/BerriAI/litellm/pull/36599)
- test(ui): await the playground model combobox before clicking it by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36850](https://github.com/BerriAI/litellm/pull/36850)
- refactor(ui): migrate budget and skill forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37262](https://github.com/BerriAI/litellm/pull/37262)
- refactor(ui): migrate tag and memory forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37266](https://github.com/BerriAI/litellm/pull/37266)
- feat(complexity\_router): plan-mode tier floor for coding-agent clients by [@​tin-berri](https://github.com/tin-berri) in [#​37230](https://github.com/BerriAI/litellm/pull/37230)
- refactor(ui): codemod every toast call site onto lib/toast and delete the antd-era facades by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37253](https://github.com/BerriAI/litellm/pull/37253)
- feat(proxy): add /team/daily/activity/aggregated and switch the Usage team tab to it by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36562](https://github.com/BerriAI/litellm/pull/36562)
- refactor(ui): migrate user, logging and policy forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37303](https://github.com/BerriAI/litellm/pull/37303)
- refactor(ui): migrate user, policy, and margin forms to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37305](https://github.com/BerriAI/litellm/pull/37305)
- refactor(ui): migrate the regenerate key and team member forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37300](https://github.com/BerriAI/litellm/pull/37300)
- refactor(ui): migrate CloudZero and cost tracking forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37312](https://github.com/BerriAI/litellm/pull/37312)
- refactor(ui): migrate auto router and credential forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37304](https://github.com/BerriAI/litellm/pull/37304)
- refactor(ui): migrate guardrail and vector store forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37306](https://github.com/BerriAI/litellm/pull/37306)
- refactor(ui): migrate prompt, UI access, plugin and MCP filter forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37297](https://github.com/BerriAI/litellm/pull/37297)
- refactor(ui): drop the unreachable user edit modal by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37327](https://github.com/BerriAI/litellm/pull/37327)
- fix(router): route Responses API input through the auto-router by [@​mateo-berri](https://github.com/mateo-berri) in [#​37333](https://github.com/BerriAI/litellm/pull/37333)
- refactor(ui): migrate the login, onboarding and search tool forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37334](https://github.com/BerriAI/litellm/pull/37334)
- feat(ui): plan-mode override tier in the auto-router create and edit forms by [@​tin-berri](https://github.com/tin-berri) in [#​37319](https://github.com/BerriAI/litellm/pull/37319)
- refactor(ui): retire the tremor date range picker in favour of the shared advanced picker by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37302](https://github.com/BerriAI/litellm/pull/37302)
- fix(proxy): forward Bedrock event-stream content-type on unbuffered passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33767](https://github.com/BerriAI/litellm/pull/33767)
- fix(azure\_ai): strip non-OpenAI-spec message fields before request by [@​ayaangazali](https://github.com/ayaangazali) in [#​34445](https://github.com/BerriAI/litellm/pull/34445)
- fix(proxy): stop leaking the client\_side\_timeout marker to providers by [@​mateo-berri](https://github.com/mateo-berri) in [#​37346](https://github.com/BerriAI/litellm/pull/37346)
- refactor(ui): migrate the caching, cost tracking, alerting and user detail forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37350](https://github.com/BerriAI/litellm/pull/37350)
- refactor(ui): migrate pass-through, project and access group forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37354](https://github.com/BerriAI/litellm/pull/37354)
- refactor(ui): migrate the vector store creation form to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37353](https://github.com/BerriAI/litellm/pull/37353)
- refactor(ui): move the MCP server forms and detail tabs off tremor by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37329](https://github.com/BerriAI/litellm/pull/37329)
- feat(complexity\_router): custom classifier plugins via classifier\_type 'custom' by [@​tin-berri](https://github.com/tin-berri) in [#​37249](https://github.com/BerriAI/litellm/pull/37249)
- fix(fireworks): skip accounts/ rewrite for FW-\* Foundry deployment ids by [@​bruno-olivia](https://github.com/bruno-olivia) in [#​37242](https://github.com/BerriAI/litellm/pull/37242)
- fix(advisor): resolve the advisor sub-call through the proxy router by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36246](https://github.com/BerriAI/litellm/pull/36246)
- fix(router): forward target\_model\_names on file uploads to litellm\_proxy deployments by [@​shivamrawat1](https://github.com/shivamrawat1) in [#​36240](https://github.com/BerriAI/litellm/pull/36240)
- refactor(ui): move the internal user detail view off tremor by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37309](https://github.com/BerriAI/litellm/pull/37309)
- fix(responses): strip the responses/ routing prefix on the Responses API path by [@​mateo-berri](https://github.com/mateo-berri) in [#​37345](https://github.com/BerriAI/litellm/pull/37345)
- fix(main): forward store and prompt\_cache\_key params on chat completions by [@​Sujithr07](https://github.com/Sujithr07) in [#​33195](https://github.com/BerriAI/litellm/pull/33195)
- refactor(ui): migrate the model settings and credential rotation modals to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37342](https://github.com/BerriAI/litellm/pull/37342)
- refactor(ui): move the shared key form controls off antd onto shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37348](https://github.com/BerriAI/litellm/pull/37348)
- refactor(ui): move the tag and vector store views off tremor by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37311](https://github.com/BerriAI/litellm/pull/37311)
- refactor(ui): migrate SSO, SCIM and vault forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37347](https://github.com/BerriAI/litellm/pull/37347)
- test(ui): cover the edit project modal's required-field validation by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37363](https://github.com/BerriAI/litellm/pull/37363)
- refactor(ui): migrate the guardrail forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37364](https://github.com/BerriAI/litellm/pull/37364)
- refactor(ui): migrate agent forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37357](https://github.com/BerriAI/litellm/pull/37357)
- refactor(ui): migrate the MCP per-user env vars, toolset and tool arguments forms to react-hook-form and shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37349](https://github.com/BerriAI/litellm/pull/37349)
- fix(anthropic): emit tool\_use content\_block\_start without awaiting the next chunk by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​37310](https://github.com/BerriAI/litellm/pull/37310)
- fix(proxy): send SSE keepalives while a slow upstream is still silent by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​37322](https://github.com/BerriAI/litellm/pull/37322)
- fix(proxy): let org admins view their organization's usage by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​37235](https://github.com/BerriAI/litellm/pull/37235)
- feat(vector\_stores): add Valkey as a managed vector store provider by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37002](https://github.com/BerriAI/litellm/pull/37002)
- refactor(ui): move the virtual key create and edit forms off tremor by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37324](https://github.com/BerriAI/litellm/pull/37324)
- refactor(ui): move the add model and credential forms off tremor by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37325](https://github.com/BerriAI/litellm/pull/37325)
- feat(team-callbacks): add DELETE /team/{team\_id}/callback/{callback\_name} by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​37331](https://github.com/BerriAI/litellm/pull/37331)
- refactor(ui): move the teams page and team detail views off tremor by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37317](https://github.com/BerriAI/litellm/pull/37317)
- fix(cost\_calculator): recognize the ultrafast service tier in cost calculation by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37355](https://github.com/BerriAI/litellm/pull/37355)
- refactor(ui): move the cache settings and playground model selector off tremor by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37323](https://github.com/BerriAI/litellm/pull/37323)
- feat(guardrails): count bedrock guardrail cost against spend and budgets by [@​mateo-berri](https://github.com/mateo-berri) in [#​37362](https://github.com/BerriAI/litellm/pull/37362)
- refactor(ui): move the admin, SSO, SCIM, alerting and fallback forms off tremor by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37315](https://github.com/BerriAI/litellm/pull/37315)
- test(ui): raise vitest test and hook timeouts for CI headroom by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37370](https://github.com/BerriAI/litellm/pull/37370)
- fix(caching): truncate semantic cache embedding input, send extra\_body top-level by [@​mateo-berri](https://github.com/mateo-berri) in [#​37367](https://github.com/BerriAI/litellm/pull/37367)
- fix(guardrails): cap the date window accepted by /guardrails/usage endpoints by [@​mateo-berri](https://github.com/mateo-berri) in [#​37380](https://github.com/BerriAI/litellm/pull/37380)
- fix(ui): show select labels on the trigger instead of raw values by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37372](https://github.com/BerriAI/litellm/pull/37372)
- refactor(ui): move the team member search modal off antd Form by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37383](https://github.com/BerriAI/litellm/pull/37383)
- refactor(ui): move the model alias manager onto design tokens and shadcn controls by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37376](https://github.com/BerriAI/litellm/pull/37376)
- refactor(ui): move the MCP tool test form off antd by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37381](https://github.com/BerriAI/litellm/pull/37381)
- refactor(ui): move the model info view and pass-through endpoint forms off tremor by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37308](https://github.com/BerriAI/litellm/pull/37308)
- fix(otel): bound and shut down credential-scoped tracer providers by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36591](https://github.com/BerriAI/litellm/pull/36591)
- fix(proxy): send SSE keepalives on assistants runs and A2A streams by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​37368](https://github.com/BerriAI/litellm/pull/37368)
- test(anthropic): pin one content\_block\_stop per tool\_use block on the Responses adapter by [@​mateo-berri](https://github.com/mateo-berri) in [#​37356](https://github.com/BerriAI/litellm/pull/37356)
- refactor(ui): move the agent, guardrail, prompt, policy and skill forms off tremor by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37320](https://github.com/BerriAI/litellm/pull/37320)
- fix(guardrails): requeue usage rollup rows dropped after retry exhaustion by [@​mateo-berri](https://github.com/mateo-berri) in [#​37387](https://github.com/BerriAI/litellm/pull/37387)
- feat(proxy): add project-level ITPM and OTPM quotas by [@​shivijain2323](https://github.com/shivijain2323) in [#​35110](https://github.com/BerriAI/litellm/pull/35110)
- feat(bedrock): add a config toggle to disable agent-runtime pass-through by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37386](https://github.com/BerriAI/litellm/pull/37386)
- fix(mcp): attach per-user BYOK credential when listing tools for non-oauth2 auth types by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34787](https://github.com/BerriAI/litellm/pull/34787)
- feat(ui): add success, warning and info status tokens by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37393](https://github.com/BerriAI/litellm/pull/37393)
- refactor(ui): drop [@​tremor/react](https://github.com/tremor/react) and the theming scaffolding it needed by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37394](https://github.com/BerriAI/litellm/pull/37394)
- refactor(ui): move the model info edit form off antd Form by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37392](https://github.com/BerriAI/litellm/pull/37392)
- fix(vector\_stores): stop leaking stored credentials in direct search debug logs by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37373](https://github.com/BerriAI/litellm/pull/37373)
- fix(logging): close three secret-leak paths in verbose logging by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37391](https://github.com/BerriAI/litellm/pull/37391)
- chore: bump litellm-enterprise 0.1.56 -> 0.1.57, litellm-proxy-extras 0.4.86 -> 0.4.87, litellm 1.98.0 -> 1.99.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37395](https://github.com/BerriAI/litellm/pull/37395)
- fix(mcp): bind tool existence check to the selected server by [@​mateo-berri](https://github.com/mateo-berri) in [#​37388](https://github.com/BerriAI/litellm/pull/37388)
- fix(mcp): serve token-forwarding servers when oauth discovery fails by [@​tin-berri](https://github.com/tin-berri) in [#​37399](https://github.com/BerriAI/litellm/pull/37399)
- fix(databricks): add cost map entries for 14 newer Databricks models by [@​epistoteles](https://github.com/epistoteles) in [#​28501](https://github.com/BerriAI/litellm/pull/28501)
- refactor(ui): style the logging settings from semantic tokens by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37385](https://github.com/BerriAI/litellm/pull/37385)
- feat(tinyfish): surface response headers + top-level response extras by [@​ChenluJi](https://github.com/ChenluJi) in [#​32448](https://github.com/BerriAI/litellm/pull/32448)
- refactor(ui): codemod the antd Tooltips outside form files onto the shadcn atom by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37402](https://github.com/BerriAI/litellm/pull/37402)
- test(ocr): update Azure DI supported-params assertion for req\_format by [@​mateo-berri](https://github.com/mateo-berri) in [#​37419](https://github.com/BerriAI/litellm/pull/37419)
- fix(proxy): return no rows when the aggregated activity entity filter is empty by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37414](https://github.com/BerriAI/litellm/pull/37414)
- test: build redaction and batch limiter fixtures the way production does by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37416](https://github.com/BerriAI/litellm/pull/37416)
- test: allow protocol-constrained pass-through routes to declare fewer methods by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37415](https://github.com/BerriAI/litellm/pull/37415)
- test(ui): pin the MCP server edit save payload before the form migration by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37404](https://github.com/BerriAI/litellm/pull/37404)
- test(ui): characterize the create key form payload contract by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37405](https://github.com/BerriAI/litellm/pull/37405)
- refactor(ui): migrate the key edit form off Ant Design onto react-hook-form by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37398](https://github.com/BerriAI/litellm/pull/37398)
- test(ui): repoint the e2e locators at the post-antd form controls by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37421](https://github.com/BerriAI/litellm/pull/37421)
- test: point the live gemini and groq conformance suites at models that still exist by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37422](https://github.com/BerriAI/litellm/pull/37422)
- refactor(ui): extract the create-key payload builder out of create\_key\_button by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37397](https://github.com/BerriAI/litellm/pull/37397)
- fix(types): map nested prompt\_tokens\_details.cache\_creation\_input\_tokens to cache\_write\_tokens by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37377](https://github.com/BerriAI/litellm/pull/37377)
- fix(router): honor key-level tag filtering in pre-routing and pin auto-router e2e regressions by [@​mateo-berri](https://github.com/mateo-berri) in [#​37366](https://github.com/BerriAI/litellm/pull/37366)
- feat(otel): attribute Prisma database spans to PostgreSQL instead of localhost by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36595](https://github.com/BerriAI/litellm/pull/36595)
- fix(bedrock): degrade gracefully on malformed tool-call arguments by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33842](https://github.com/BerriAI/litellm/pull/33842)
- test: move the remaining live groq call sites off the retired llama models by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37426](https://github.com/BerriAI/litellm/pull/37426)
- fix(ui): highlight the first member search match so Enter picks it by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37429](https://github.com/BerriAI/litellm/pull/37429)
- chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37400](https://github.com/BerriAI/litellm/pull/37400)
- fix(proxy): log spend for OpenAI passthrough embeddings with unmapped models by [@​mateo-berri](https://github.com/mateo-berri) in [#​37425](https://github.com/BerriAI/litellm/pull/37425)
- fix(router): keep acreate\_file fallbacks inside the requested model group by [@​mateo-berri](https://github.com/mateo-berri) in [#​37424](https://github.com/BerriAI/litellm/pull/37424)
- fix(proxy): record estimated input tokens in spend logs for failed dispatched requests by [@​mateo-berri](https://github.com/mateo-berri) in [#​37365](https://github.com/BerriAI/litellm/pull/37365)
- fix: accept bool thinking param instead of crashing with AttributeError by [@​mateo-berri](https://github.com/mateo-berri) in [#​37423](https://github.com/BerriAI/litellm/pull/37423)
- refactor(ui): migrate the teams form graph off antd Form onto react-hook-form by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37417](https://github.com/BerriAI/litellm/pull/37417)
- fix(ui): restore the cache control Role and Index field hints by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37437](https://github.com/BerriAI/litellm/pull/37437)
- feat(ui): add mounted-field projections for the MCP server form graph by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37440](https://github.com/BerriAI/litellm/pull/37440)
- refactor(ui): extract the MCP server edit save payload into a pure builder by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37436](https://github.com/BerriAI/litellm/pull/37436)
- test: derive vertex batch cost expectation from the cost map by [@​mateo-berri](https://github.com/mateo-berri) in [#​37444](https://github.com/BerriAI/litellm/pull/37444)
- refactor(ui): port the create key form off antd Form onto react-hook-form by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37442](https://github.com/BerriAI/litellm/pull/37442)
- refactor(ui): port the add model form off antd Form onto react-hook-form by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37446](https://github.com/BerriAI/litellm/pull/37446)
- refactor(ui): host KeyLifecycleSettings tests in react-hook-form instead of antd Form by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37449](https://github.com/BerriAI/litellm/pull/37449)
- fix(ui): rebuild nested and list paths in the mounted-field projection by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37450](https://github.com/BerriAI/litellm/pull/37450)
- fix(ui): gate the pass-through guardrail field inputs when the section is disabled by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37435](https://github.com/BerriAI/litellm/pull/37435)
- fix(tests): keep a host PROXY\_BASE\_URL out of request-derived URL tests by [@​mateo-berri](https://github.com/mateo-berri) in [#​37451](https://github.com/BerriAI/litellm/pull/37451)
- refactor(ui): port the MCP server forms off antd Form onto react-hook-form by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37483](https://github.com/BerriAI/litellm/pull/37483)
- test(e2e): pin the tag-routing denial to its actual cause by [@​mateo-berri](https://github.com/mateo-berri) in [#​37432](https://github.com/BerriAI/litellm/pull/37432)
- fix(proxy): read through to the DB on registry misses so just-created models, guardrails, and agents resolve on sibling replicas by [@​mateo-berri](https://github.com/mateo-berri) in [#​36263](https://github.com/BerriAI/litellm/pull/36263)
- fix(mcp): forward the per-server auth header on OpenAPI tool calls by [@​tin-berri](https://github.com/tin-berri) in [#​37410](https://github.com/BerriAI/litellm/pull/37410)
- chore(typing): drop 1.3k basedpyright errors across 42 Any hotspot files by [@​mateo-berri](https://github.com/mateo-berri) in [#​37439](https://github.com/BerriAI/litellm/pull/37439)
- test(ui): drive fields with change events where the typing is not the behaviour by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37495](https://github.com/BerriAI/litellm/pull/37495)
- fix(ui): restore tab strip styling and panel persistence lost in the shadcn migration by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37403](https://github.com/BerriAI/litellm/pull/37403)
- feat(spend-logs): add lifecycle timestamps by [@​sytianhe](https://github.com/sytianhe) in [#​37361](https://github.com/BerriAI/litellm/pull/37361)
- refactor(ui): migrate the antd Button call sites onto the shadcn Button by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37505](https://github.com/BerriAI/litellm/pull/37505)
- test(ui): split the vitest suite into unit, component, integration and type projects by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37488](https://github.com/BerriAI/litellm/pull/37488)
- refactor(ptu): give the rollup a source-agnostic deployment record by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​37501](https://github.com/BerriAI/litellm/pull/37501)
- feat(auto-router)!: scope shadow eval jobs to multiple keys by [@​tin-berri](https://github.com/tin-berri) in [#​37251](https://github.com/BerriAI/litellm/pull/37251)
- refactor(ui): migrate the antd Alert call sites onto the shared Alert by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37513](https://github.com/BerriAI/litellm/pull/37513)
- chore(ui): upgrade the dashboard to React 19 by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37411](https://github.com/BerriAI/litellm/pull/37411)
- fix(streaming): accept provider cost objects when propagating usage cost by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36593](https://github.com/BerriAI/litellm/pull/36593)
- fix(complexity-router): gate the reasoning override on a non-SIMPLE score by [@​tin-berri](https://github.com/tin-berri) in [#​37500](https://github.com/BerriAI/litellm/pull/37500)
- fix(mcp): stop reporting failed OpenAPI tool calls as successes by [@​tin-berri](https://github.com/tin-berri) in [#​37496](https://github.com/BerriAI/litellm/pull/37496)
- feat(e2e): add record/replay transport seam and fixture bundle format by [@​mateo-berri](https://github.com/mateo-berri) in [#​37360](https://github.com/BerriAI/litellm/pull/37360)
- fix(proxy): accept inherited model sentinels in project key limits by [@​mateo-berri](https://github.com/mateo-berri) in [#​37515](https://github.com/BerriAI/litellm/pull/37515)
- fix(model\_prices): add provider-announced deprecation\_date to 205 registry entries by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37283](https://github.com/BerriAI/litellm/pull/37283)
- fix(model\_prices): correct gemini 3.1 flash image and deepseek v4 pricing, add openai deprecation dates by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37473](https://github.com/BerriAI/litellm/pull/37473)
- fix(model\_prices): set prompt\_cache\_min\_tokens=4096 for Gemini 3.5/3.6/3.7 Flash and 3.1 Pro Preview by [@​mateo-berri](https://github.com/mateo-berri) in [#​37516](https://github.com/BerriAI/litellm/pull/37516)
- fix(anthropic,bedrock): report provider thinking tokens instead of classifying them as text by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35998](https://github.com/BerriAI/litellm/pull/35998)
- fix(batches): stop one bad output line from zeroing an entire batch's spend by [@​mateo-berri](https://github.com/mateo-berri) in [#​37457](https://github.com/BerriAI/litellm/pull/37457)
- feat(e2e): canonical content-based match keys for record-and-replay by [@​mateo-berri](https://github.com/mateo-berri) in [#​37525](https://github.com/BerriAI/litellm/pull/37525)
- feat(cli): add `lite login --config-claude` to wire Claude Code at login by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​37507](https://github.com/BerriAI/litellm/pull/37507)
- fix(auth): resolve bare model names against wildcard deployments in model access groups by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37492](https://github.com/BerriAI/litellm/pull/37492)
- docs: run only the tests covering your change, leave suites to CI by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37528](https://github.com/BerriAI/litellm/pull/37528)
- feat(complexity-router): make the reasoning override floor configurable by [@​tin-berri](https://github.com/tin-berri) in [#​37537](https://github.com/BerriAI/litellm/pull/37537)
- fix(ui): drop stale user search answers so Enter commits the current match by [@​mateo-berri](https://github.com/mateo-berri) in [#​37504](https://github.com/BerriAI/litellm/pull/37504)
- refactor(ui): migrate the remaining dashboard pages off antd by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37524](https://github.com/BerriAI/litellm/pull/37524)
- feat(proxy): auto-suppress the no-Redis banner for confirmed single-worker deployments by [@​mateo-berri](https://github.com/mateo-berri) in [#​36987](https://github.com/BerriAI/litellm/pull/36987)
- fix(proxy): retry spend updates on Postgres deadlock instead of dropping them by [@​RayJueWang](https://github.com/RayJueWang) in [#​34887](https://github.com/BerriAI/litellm/pull/34887)
- feat(search): add Amazon Bedrock AgentCore web search provider by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36331](https://github.com/BerriAI/litellm/pull/36331)
- fix(helm): default litellm-helm to the ghcr.io/berriai/litellm image by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​37491](https://github.com/BerriAI/litellm/pull/37491)
- refactor(ui): migrate antd Modal onto the shared shadcn Dialog by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37540](https://github.com/BerriAI/litellm/pull/37540)
- fix(ui): toggle unlimited budget when its text is clicked by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37547](https://github.com/BerriAI/litellm/pull/37547)
- feat(proxy): fast-fail validation for batch input files at /v1/files by [@​mateo-berri](https://github.com/mateo-berri) in [#​37527](https://github.com/BerriAI/litellm/pull/37527)
- chore: gitignore CLAUDE.local.md by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37545](https://github.com/BerriAI/litellm/pull/37545)
- perf(otel): build the credential-scoped tracer Resource once per logger by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​37542](https://github.com/BerriAI/litellm/pull/37542)
- fix(ci): gate backend unit tests on the pull request's own file list by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37550](https://github.com/BerriAI/litellm/pull/37550)
- chore(codeowners): require pricing owner approval for the model prices jsons by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37551](https://github.com/BerriAI/litellm/pull/37551)
- fix(proxy): initialize the secret manager before resolving os.environ config references by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​37544](https://github.com/BerriAI/litellm/pull/37544)
- refactor(ui): swap [@​ant-design/icons](https://github.com/ant-design/icons) for lucide-react by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37553](https://github.com/BerriAI/litellm/pull/37553)
- refactor(ui): migrate shared primitives and common components off antd by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37521](https://github.com/BerriAI/litellm/pull/37521)
- fix(spend-logs): backfill created\_at/updated\_at from row endTime instead of migration time by [@​mateo-berri](https://github.com/mateo-berri) in [#​37554](https://github.com/BerriAI/litellm/pull/37554)
- refactor(ui): migrate the MCP servers pages off antd by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37522](https://github.com/BerriAI/litellm/pull/37522)
- fix(vertex\_ai): apply regional endpoint uplift to cost tracking by [@​mateo-berri](https://github.com/mateo-berri) in [#​37543](https://github.com/BerriAI/litellm/pull/37543)
- fix(proxy): populate deployment attribution on failed-request spend logs by [@​mateo-berri](https://github.com/mateo-berri) in [#​37520](https://github.com/BerriAI/litellm/pull/37520)
- refactor(ui): migrate the model and router settings pages off antd by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37523](https://github.com/BerriAI/litellm/pull/37523)
- perf(ci): gate the lint, MCP and dashboard jobs on the pull request's file list by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37559](https://github.com/BerriAI/litellm/pull/37559)
- feat(router): allow per-tier litellm\_params in complexity autorouter config by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37064](https://github.com/BerriAI/litellm/pull/37064)
- fix(anthropic): log partial stream spend when a /v1/messages client disconnects mid-stream by [@​mateo-berri](https://github.com/mateo-berri) in [#​37558](https://github.com/BerriAI/litellm/pull/37558)
- fix(ui): clear pass-through header rows when the create modal is reopened by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37549](https://github.com/BerriAI/litellm/pull/37549)
- fix(ui): render optional array and object MCP tool parameters as JSON inputs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37548](https://github.com/BerriAI/litellm/pull/37548)
- feat(proxy)!: default audit logs on for enterprise licenses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37518](https://github.com/BerriAI/litellm/pull/37518)
- feat(ui): standardize the Teams page header by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36897](https://github.com/BerriAI/litellm/pull/36897)
- feat(ptu): accrue flat cost for PTU deployments declared in config.yaml by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​37556](https://github.com/BerriAI/litellm/pull/37556)
- refactor(ui): migrate the last antd components off antd onto shadcn by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37569](https://github.com/BerriAI/litellm/pull/37569)
- chore(ui): drop the antd dependency and its leftovers by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37574](https://github.com/BerriAI/litellm/pull/37574)
- refactor(ui): map hardcoded Tailwind palette classes onto semantic tokens by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37576](https://github.com/BerriAI/litellm/pull/37576)
- fix(ptu): hand the prune a plain delete filter the query builder can serialise by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​37571](https://github.com/BerriAI/litellm/pull/37571)
- feat(proxy): enqueued-token rate limiting for batches with refund on completion and cancellation by [@​mateo-berri](https://github.com/mateo-berri) in [#​37539](https://github.com/BerriAI/litellm/pull/37539)
- fix(ui): restore hover feedback and dark-mode variants lost in the token migration by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37579](https://github.com/BerriAI/litellm/pull/37579)
- fix(ci): run the full dashboard suite when a change reaches outside src/ by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37563](https://github.com/BerriAI/litellm/pull/37563)
- fix(ui): keep semantic button colours on hover after the no-op hover cleanup by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37580](https://github.com/BerriAI/litellm/pull/37580)
- fix: add supports\_mid\_conversation\_system to bare first-party Claude cost-map keys by [@​oneKn8](https://github.com/oneKn8) in [#​36969](https://github.com/BerriAI/litellm/pull/36969)
- feat: add bedrock grok 4.6 to model cost map by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37517](https://github.com/BerriAI/litellm/pull/37517)
- fix: preserve prompt cache for mid-conversation system on unflagged Claude models by [@​oneKn8](https://github.com/oneKn8) in [#​36968](https://github.com/BerriAI/litellm/pull/36968)
- fix(router): routed deployment's own litellm\_params beat forwarded auto\_router marker params by [@​mateo-berri](https://github.com/mateo-berri) in [#​37615](https://github.com/BerriAI/litellm/pull/37615)
- refactor(ci): fold the nine thin unit-shard callers into one matrix by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37590](https://github.com/BerriAI/litellm/pull/37590)
- chore(ci): close the test-census blind spots and move scripts out of workflows/ by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37586](https://github.com/BerriAI/litellm/pull/37586)
- test: retire tests/old\_proxy\_tests, which holds no tests by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37605](https://github.com/BerriAI/litellm/pull/37605)
- feat(ci): ratchet the test suite's zero-assert, mock-echo and global-state debt by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37588](https://github.com/BerriAI/litellm/pull/37588)
- feat(proxy): native CLI login with OAuth authorization code + PKCE by [@​mateo-berri](https://github.com/mateo-berri) in [#​37626](https://github.com/BerriAI/litellm/pull/37626)
- feat(prompt-caching): map cache\_control\_injection\_points to OpenAI prompt\_cache\_breakpoint on GPT-5.6+ targets by [@​mateo-berri](https://github.com/mateo-berri) in [#​37628](https://github.com/BerriAI/litellm/pull/37628)
- fix(realtime): bound Vertex credential resolution and make realtime failures loud by [@​mateo-berri](https://github.com/mateo-berri) in [#​37604](https://github.com/BerriAI/litellm/pull/37604)
- fix(anthropic): map metadata.user\_id to prompt\_cache\_key on the /v1/messages bridge by [@​mateo-berri](https://github.com/mateo-berri) in [#​37623](https://github.com/BerriAI/litellm/pull/37623)
- fix(passthrough): resolve vertex live credentials from db model deployments by [@​mateo-berri](https://github.com/mateo-berri) in [#​37602](https://github.com/BerriAI/litellm/pull/37602)
- fix(prompt\_management): don't route no-prompt\_id requests to prompt managers that can't run them by [@​mateo-berri](https://github.com/mateo-berri) in [#​37575](https://github.com/BerriAI/litellm/pull/37575)
- test: remove the five test functions a later definition shadows by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37591](https://github.com/BerriAI/litellm/pull/37591)
- feat(ci): guard shard assignment across every sharded test tree by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37593](https://github.com/BerriAI/litellm/pull/37593)
- feat(ui): multi-key shadow eval picker and per-key breakdown by [@​tin-berri](https://github.com/tin-berri) in [#​37389](https://github.com/BerriAI/litellm/pull/37389)
- fix(ui): make dark-mode form controls visible by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37648](https://github.com/BerriAI/litellm/pull/37648)
- fix(ui): give status colours a readable foreground and drop the muted 70% step by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37649](https://github.com/BerriAI/litellm/pull/37649)
- fix(ui): make inline styles and code blocks follow the theme by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37651](https://github.com/BerriAI/litellm/pull/37651)
- fix(ui): move the policy flow builder onto theme tokens by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37654](https://github.com/BerriAI/litellm/pull/37654)
- test: settle three allowlist entries that were open questions by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37598](https://github.com/BerriAI/litellm/pull/37598)
- feat(ci): ratchet tests that skip themselves when a credential is absent by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37612](https://github.com/BerriAI/litellm/pull/37612)
- feat(ci): catch files a -k expression deselects from every job by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37601](https://github.com/BerriAI/litellm/pull/37601)
- test: run the 30 test files stranded in the second mirror by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37595](https://github.com/BerriAI/litellm/pull/37595)
- fix(ui): make hardcoded palette surfaces theme-aware by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37650](https://github.com/BerriAI/litellm/pull/37650)
- feat(cli): store the lite login credential in the OS keychain by [@​mateo-berri](https://github.com/mateo-berri) in [#​37566](https://github.com/BerriAI/litellm/pull/37566)
- fix(ui): draw one Per Day savings bar per date on Cost Optimization by [@​tin-berri](https://github.com/tin-berri) in [#​37643](https://github.com/BerriAI/litellm/pull/37643)
- feat(mistral): add zai-glm-5-2 and glm-5-2 model pricing by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​37110](https://github.com/BerriAI/litellm/pull/37110)
- feat(complexity\_router): add business classification rubric preset by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37534](https://github.com/BerriAI/litellm/pull/37534)
- feat(ui): serve a dark-mode variant of the LiteLLM logo by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37656](https://github.com/BerriAI/litellm/pull/37656)
- fix(otel): route Phoenix traces to per-key/team projects under otel v2 by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​36706](https://github.com/BerriAI/litellm/pull/36706)
- test: replace blind sleeps with deadline waits in callback and caching tests by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37660](https://github.com/BerriAI/litellm/pull/37660)
- fix(cli): keep the --pkce refresh token in the OS keychain, not in token.json by [@​mateo-berri](https://github.com/mateo-berri) in [#​37665](https://github.com/BerriAI/litellm/pull/37665)
- fix(ui): keep keyword tier rules that target operator-defined tiers when hydrating the edit modal by [@​tin-berri](https://github.com/tin-berri) in [#​37413](https://github.com/BerriAI/litellm/pull/37413)
- feat(proxy): add POST /auto\_router/validate\_complexity\_router\_config to dry-run the complexity-router write gate by [@​tin-berri](https://github.com/tin-berri) in [#​37409](https://github.com/BerriAI/litellm/pull/37409)
- feat(ui): let admins supply a dark-mode variant of their custom logo by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37662](https://github.com/BerriAI/litellm/pull/37662)
- fix(proxy): run pre-call guardrails on batch input file uploads by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​37519](https://github.com/BerriAI/litellm/pull/37519)
- feat(ui): add a light/dark/system theme toggle to the top bar by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37669](https://github.com/BerriAI/litellm/pull/37669)
- feat(proxy): redact or drop individual batch records instead of rejecting the file by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​37561](https://github.com/BerriAI/litellm/pull/37561)
- refactor(ui): mark dark as beta in the theme menu instead of the toolbar by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37680](https://github.com/BerriAI/litellm/pull/37680)
- ci: lint the test tree for undefined names (F821) and fix all 30 by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37671](https://github.com/BerriAI/litellm/pull/37671)
- feat(e2e): move record/replay to the provider edge (LIT-5745) by [@​mateo-berri](https://github.com/mateo-berri) in [#​37565](https://github.com/BerriAI/litellm/pull/37565)
- fix(mcp): let a salt-key-orphaned OAuth credential be replaced by re-authorization by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​37672](https://github.com/BerriAI/litellm/pull/37672)
- fix(mcp): normalize auth schemes so MCP egress emits exactly one prefix by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​37668](https://github.com/BerriAI/litellm/pull/37668)
- test: add six ruff rules that catch tests which cannot fail by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37709](https://github.com/BerriAI/litellm/pull/37709)
- perf(ci): measure unit-shard coverage with the sys.monitoring core by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37589](https://github.com/BerriAI/litellm/pull/37589)
- test: merge three stranded twins into the files that shadow them by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37600](https://github.com/BerriAI/litellm/pull/37600)
- test(ci): reject coverage-allowlist entries that no longer match a file by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37608](https://github.com/BerriAI/litellm/pull/37608)
- feat(ci): assert .github/workflows holds only workflows, correctly named by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37616](https://github.com/BerriAI/litellm/pull/37616)
- feat(ci): freeze the conftest save/restore inventory so it can only shrink by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37621](https://github.com/BerriAI/litellm/pull/37621)
- fix(a2a): accept the whole JSON-RPC id union the spec defines by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​37704](https://github.com/BerriAI/litellm/pull/37704)
- fix(ptu): refuse an incomplete config.yaml reservation the way the endpoints do by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​37703](https://github.com/BerriAI/litellm/pull/37703)
- chore: bump litellm-enterprise 0.1.57 -> 0.1.58, litellm-proxy-extras 0.4.87 -> 0.4.88 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37717](https://github.com/BerriAI/litellm/pull/37717)
- feat(perplexity): add Agent API third-party models by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​37112](https://github.com/BerriAI/litellm/pull/37112)
- fix(ui): surface the paginated fallback on Cost Optimization by [@​tin-berri](https://github.com/tin-berri) in [#​37659](https://github.com/BerriAI/litellm/pull/37659)
- feat(shadow\_eval)!: gate the per-key budget on dollar spend instead of turns by [@​tin-berri](https://github.com/tin-berri) in [#​37555](https://github.com/BerriAI/litellm/pull/37555)
- feat(ui): per-model reasoning effort in the complexity tier editor by [@​tin-berri](https://github.com/tin-berri) in [#​37673](https://github.com/Berr…
TLDR
Problem this solves:
gpt-5-chatdeployments 400 on everymax_tokensrequest/healthreports those deployments permanently unhealthygpt-5-chatfamilyHow it solves it:
requires_max_completion_tokenspredicate covers every gpt-5 namemax_tokensrename keys off it, nothing elseis_model_gpt_5_model, so [Bug]: OpenAI GPT-5 Chat model does not support "temperature" parameter #13781 stays fixedUser Flow
Before: a developer whose Azure deployment is named
gpt-5-chatgets a 400 on every request that setsmax_tokens, and the deployment sits red in/healthforever{"model": "gpt-5-chat", "messages": [...], "max_tokens": 5}Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead.unhealthy_endpointswith that same message,healthy_count0After: the same request succeeds and the same deployment reports healthy
{"model": "gpt-5-chat", "messages": [...], "max_tokens": 5}healthy_endpoints,unhealthy_count0Relevant issues
No upstream issue covers this defect. #13781 is the one this change is careful not to reopen: it reported
gpt-5-chat-latestwrongly rejectingtemperature, which is why the family was pulled off the GPT-5 reasoning path in the first place. #24779, closed as not planned, claims the same Azure error string forazure/gpt-4oLinear ticket
Resolves LIT-5549
Pre-Submission checklist
@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
There are no Azure credentials on the machine this was captured on, so the run below points a live proxy at a local HTTP server standing in for Azure. That server records the exact JSON body litellm puts on the wire and returns Azure's own 400 whenever
max_tokensis present. The wire body is measured, the rejection is simulated, and the claim that Azure really rejectsmax_tokensfor this family rests on a customer's production error plus Azure/azure-sdk-for-net#51844, which namesgpt-5-chatamong the deployments returning that exact stringBoth runs use the same config, four Azure deployments (
gpt-5-chat,gpt-5,o3-mini,gpt-4o) all pointed at the stand-in on 127.0.0.1:45550, proxy on 127.0.0.1:45549Before, at e0f388c (this branch's base):
After, at 1e877f2, since rebased onto e1f3d6e, then 29d8bed, and now 5b0ea66 (this PR). The capture is still the code at the current head:
litellm/llms/azure/chat/gpt_transformation.pycarries a byte-identical contribution at every one of those commits, and the latest rebase only merged a neighbouring test from #34462 into the test file:The
gpt-5ando3-minirows are the load-bearing part of the before run: the same harness printsmax_completion_tokenswhenever litellm emits it, somax_tokenson thegpt-5-chatrow means the code picked that key rather than the fixture never carrying the alternative.gpt-4ois unchanged on purpose and its 400 only shows that the stand-in rejectsmax_tokensfrom everyoneType
🐛 Bug Fix
Caveats (if any)
azure/gpt-4ostill sendsmax_tokens; [Bug]: Azure gpt-4o rejectsmax_tokens— needs translation tomax_completion_tokenslike o-series/gpt-5 #24779 is out of scopeReview notes
Greptile scored this 4/5 on 1e877f2 with two findings it framed as non-blocking follow-ups. I looked into both and am deliberately not changing the code for either
On the predicate being hardcoded rather than read from the model cost map: no such flag exists. I enumerated every
support*key present anywhere inmodel_prices_and_context_window.json, 40 of them across roughly three thousand entries, and none expresses "this deployment rejects the legacymax_tokenskey". The nearest neighbours aresupports_reasoning,supports_sampling_paramsand thesupports_*_reasoning_effortfamily, which all answer the reasoning question instead.supports_reasoningis the worst of them to borrow here:azure/gpt-5-chatcarriessupports_reasoning: truetoday, so keying the rename off it would rebuild the exact conflation of two independent questions that caused this bug. The map also already spends the namemax_tokenson the output-token ceiling, so a capability flag named around it would read as a collision. Model naming is the only source of truth for this capability at present, and all three pre-existing renames in the tree agree:llms/openai/chat/o_series_transformation.py:99keys offis_o_series_model, andllms/openai/chat/gpt_5_transformation.py:200and:256key offis_model_gpt_5_search_modelandis_model_gpt_5_model. Introducing a flag would mean populating it correctly across the whole gpt-5 and o-series families before it could be trusted, which is a far larger and riskier change than this defect justifiesOn the new branch introducing another parameter mutation:
optional_paramsis a local accumulator here, not caller-owned state. It is created as a fresh{}insidepre_process_optional_params(utils.py:3855), returned intoget_optional_params, and threaded through the providermap_openai_paramschain with the return value reassigned at every hop, which is the contractBaseConfig.map_openai_paramsdeclares. Filling it is what the function is for, and every other branch of the same loop fills it the same way. The new branch is actually the least mutating of the four renames in the tree, since the three cited above each callnon_default_params.pop("max_tokens")and reach into the caller's dict, while this one writes only into the accumulator and pops nothingThe automated security review raised a Medium, that a request carrying both
max_tokens=1andmax_completion_tokens=10000reserves one output token while the provider receives ten thousand, and asked for the conflicting fields to be rejected or normalised before the pre-call rate-limit hooks run. The under-reservation is real and worth fixing. It is being fixed in #37001, in the limiter, and deliberately not here. Three measurements, each of which independently rules out doing it in this fileOrdering. The reservation is taken by
_estimate_tokens_for_requestatproxy/hooks/parallel_request_limiter_v3.py:607, readingdata.get("max_tokens") or data.get("max_completion_tokens")straight off the raw request body insideasync_pre_call_hook, which fires atproxy/common_request_processing.py:1471before routing and long before any provider transformation. Driven directly, that body reserves 2 tokens where the same body carrying onlymax_completion_tokensreserves 10001. Nothing inmap_openai_paramscan move those numbers, because the reservation is already written by the time the transformation runs. Rejecting there would convert a bypass into a 400 after the fact rather than reserve correctly, which is not the same thingReach. This is not introduced here. At this branch's merge-base, the same request already drops
max_tokens: 1and emitsmax_completion_tokens: 10000forazure/gpt-5,azure/o3-mini,openai/gpt-5andopenai/o3-mini. On a live proxy at base,azure/gpt-5with both fields returns HTTP 200 with ten thousand on the wire, which is exactly whatazure/gpt-5-chatdoes after this change. Closing it insideAzureOpenAIConfigwould shut one of five doors while implying the corridor was securedBlast radius. Rejecting the combination would break a working configuration.
LiteLLM_Paramssetsextra="allow"and the router merges deployment params underneath client kwargs, so an operator's deployment-levelmax_tokensdefault and a client'smax_completion_tokensarrive together by design. Run end to end through the router, a deployment carryingmax_tokens: 4096with a client sendingmax_completion_tokens: 100returns HTTP 200 and puts{"max_completion_tokens": 100}on the wire, the client value correctly winning. A conflict error would fail every operator in that shapeFinal Attestation