fix(router): eagerly fetch Vertex AI deferred stream to surface HTTP errors in _acompletion fallback path - #34627
Conversation
|
Deepanshu seems not to be a GitHub user. You need a GitHub account to be able to sign the CLA. If you have already a GitHub account, please add the email address used for this commit to your account. You have signed the CLA already but the status is still pending? Let us recheck it. |
|
@greptileai please review |
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 2 · PR risk: 0/10 |
Greptile SummaryThis revision completes the streaming fallback and proxy response-header fixes requested in prior review.
Confidence Score: 5/5The PR appears safe to merge. No blocking failure remains.
|
| Filename | Overview |
|---|---|
| litellm/router.py | Eagerly initializes deferred streams and prevents fallback restarts after recorded client-visible output; the previously reported content-detection gaps are addressed. |
| litellm/proxy/common_request_processing.py | Sanitizes unsafe response headers after hooks and after merging existing ProxyException headers, resolving both previously reported bypass paths. |
| litellm/constants.py | Centralizes the case-insensitive set of framing and browser-facing headers excluded from proxy-generated responses. |
| tests/test_litellm/test_router.py | Adds focused coverage for deferred-stream initialization and text and non-text partial-stream failures. |
| tests/test_litellm/proxy/test_common_request_processing.py | Covers provider, hook-added, mixed response, and pre-existing ProxyException header sanitization. |
| tests/test_litellm/test_redact_string_in_error_paths.py | Confirms the Router logger’s redaction filter sanitizes exception traceback text on the changed fallback-failure path. |
Reviews (14): Last reviewed commit: "chore: retrigger CI (frontend-lint cance..." | Re-trigger Greptile
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
|
Addressed this round's review findings in 9c108db (also fixed the PR base branch to litellm_internal_staging and rewrote the description to match the actual PR template). Fixed, with regression tests:
Investigated and refuted, with evidence (replied inline on each thread):
All 342 tests across the touched files pass locally, |
|
@greptileai please re-review |
|
@veria-ai please re-review |
|
Pushed 2c006cb: added a direct unit test for Two other failing checks are unrelated to this PR's diff and not something I can fix here:
|
|
Pushed 2214139 fixing a real gap Greptile found (confidence went 2/5 -> 3/5 last round): in Also investigated two new CodeQL "module-level cyclic import" alerts on the Also fixed a mistake from my own tooling in the prior round: |
|
@greptileai please re-review |
|
Heads up for a maintainer: no GitHub Actions checks have run for the last 3 pushes (d1cd74c, 2214139, f568959) — only the third-party Veria AI check is firing. The first 3 pushes on this PR ran the full CI matrix (~80 checks) normally, so this looks like the workflow-approval gate for fork PRs needing a fresh 'Approve and run workflows' click rather than anything in the diff. Would appreciate someone approving/re-running the workflows when convenient. |
|
@greptileai please re-review |
|
Pushed a9257f6 addressing the |
|
@greptileai please re-review |
|
Did a live manual verification against real provider APIs (not just unit tests) since Greptile hit 5/5: ran the proxy locally with a Vertex AI deployment pointed at a nonexistent project (forces a real 403) and Anthropic Haiku as the explicit fallback. Confirmed |
a9257f6 to
edd19d2
Compare
|
@veria-ai please re-review |
Manual verification against real provider APIsRan the proxy locally (not mocks) with a Vertex AI deployment pointed at a nonexistent project to force a real 403, and an explicit fallback to Anthropic Haiku. Config: model_list:
- model_name: bad-vertex-model
litellm_params:
model: vertex_ai/gemini-2.0-flash
vertex_project: nonexistent-test-project-verify-fix
vertex_location: us-central1
model_info:
id: bad-vertex
- model_name: anthropic-fallback-model
litellm_params:
model: anthropic/claude-haiku-4-5
api_key: os.environ/ANTHROPIC_API_KEY
router_settings:
num_retries: 0
fallbacks:
- bad-vertex-model:
- anthropic-fallback-modelStart the proxy: CONFIG_FILE_PATH=verify_config.yaml LITELLM_LOG=DEBUG uv run uvicorn litellm.proxy.proxy_server:app --port 4111Request: curl -N -X POST http://localhost:4111/v1/chat/completions \
-H "Authorization: Bearer sk-local-dev-master-key" \
-d '{"model":"bad-vertex-model","messages":[{"role":"user","content":"Say hi in exactly 3 words"}],"stream":true}'Router log (the fix firing): This propagates through Client response (real, live streamed completion from the fallback deployment): Before this fix, the Vertex HTTP error stays hidden inside the generator and fallback never triggers at all. This confirms the eager fetch surfaces it and the fallback delivers a real, working response end to end. |
edd19d2 to
9a2b22f
Compare
|
Rebased onto the latest Also pushed a fix for veria-ai's remaining open finding (arbitrary provider response headers forwarded): added a dedicated |
|
@greptileai please re-review |
|
@veria-ai please re-review |
yassin-berriai
left a comment
There was a problem hiding this comment.
Thanks for the detailed writeup, and for doing a real live-proxy verification rather than leaning on unit tests alone; that part is good evidence. The core mechanic is sound too: fetch_stream() is idempotent (guarded on completion_stream is None) and __anext__ already calls it lazily, so awaiting it eagerly doesn't double-invoke make_call or consume chunks, and the affected paths (Vertex Gemini, Bedrock converse and invoke, Codestral, Predibase, and the MCP chat-completions wrapper) are all genuinely deferred-stream paths where surfacing the HTTP error into the router's except block is the right call
I can't approve it as it stands, though. There are some CI items and one change that needs a maintainer decision rather than a review sign-off
On CI, lint is failing on this diff rather than on staging: the ruff strict gate reports UP006: total 12791 over limit 12789 (this change added 1) pointing at litellm/router.py:295, which is List[ModelResponseStream] in _stream_chunks_have_generated_content. Lowercase list[...] clears it. lint is green on every other currently open PR
On osv-scan, I don't think the "pre-existing, not something I can fix here" read holds. This branch is 67 commits behind litellm_internal_staging, and staging already carries gitpython 3.1.55 while this branch's uv.lock still pins 3.1.54; brace-expansion is the same story. A rebase onto current staging should clear that check without hand-editing any lockfile
The CLA also can't go green in its current shape. The commits are authored under an identity that isn't linked to a GitHub account, which is why the bot reports that the signer "seems not to be a GitHub user". Adding that commit email to your GitHub account and re-running the check should sort it
The bigger item is the mid-stream change. Removing the continuation-resume so that a stream failing after partial content re-raises instead of falling back is a defensible call, but it is a product call, and it isn't what the PR title or #31874 describe. Two open PRs (#30242 and #30743) are currently fixing that exact code path for Anthropic's removal of assistant prefill, so landing this would pull the ground out from under both. The Responses API path in the same file still resumes mid-stream via _build_responses_continuation_input, and tests/router_unit_tests/test_router_aresponses_streaming_fallback.py asserts precisely the semantics being deleted on the chat side, so after this PR /chat/completions and /v1/responses would disagree about what happens when a stream dies after partial content. Relatedly, convert_prefix_message_to_non_prefix_messages loses its router-side caller entirely, and continuation coverage in tests/test_litellm/test_router.py drops from 18 references to 6 because two existing tests get rewritten to assert the new behavior
There is also a second semantic change that the description doesn't call out: the new guard is wider than the branch it replaces, so is_pre_first_chunk=False with an empty generated_content but chunks carrying only tool-call, reasoning, or thinking content used to fall back with the original messages and now re-raises. That one was requested in review so it isn't accidental, but it deserves its own line in the description
Smaller points, in rough priority order. _strip_http_framing_headers in router.py has no production callers since the proxy inlines the dict comprehension instead, so the helper and its three tests are dead weight; tests written to satisfy the router coverage gate rather than to catch a regression aren't really the coverage we want. The eager fetch sits after success_calls += 1, the 200 OK info log, and _track_deployment_metrics, then compensates with a manual success_calls[model_name] -= 1; moving the fetch to immediately after response = await _response removes the need for the decrement and stops emitting a success log for a call that failed. _HTTP_FRAMING_HEADERS is a private HTTP constant living in router.py and imported into the proxy, which is what the two open CodeQL cyclic-import threads are pointing at, and a constants module would resolve both. At tests/test_litellm/test_router.py:4752, mock_fallback.assert_not_called(), "..." is a tuple expression, so the message is inert; the check itself still runs. Finally, the logging change that moves redaction from the call site to global logger configuration is unrelated to the stated fix and would be much easier to reason about on its own
Concretely, what I'd like to see before this can go in:
- rebase onto current
litellm_internal_staging List->listatlitellm/router.py:295- get the CLA signable by linking the commit email to a GitHub account
- drop
_strip_http_framing_headersand its three tests, and move_HTTP_FRAMING_HEADERSout ofrouter.py - move the eager
fetch_stream()above the success bookkeeping and drop the manual decrement - split the mid-stream fallback change and the logging change into separate PRs, coordinating the former with #30242 and #30743
Stripped down to just the deferred-stream fix, this is a clean and useful change, and I'd be happy to see that part land on its own
Removing the continuation-prompt fallback (retrying with the partial response as a prefixed assistant message) so a stream failing after partial content always re-raises instead was a scope decision beyond what this PR's title/issue (BerriAI#31874) describe, and it directly conflicts with BerriAI#30242/BerriAI#30743, which are already fixing the same code path for Anthropic's removal of assistant-message prefill on Sonnet 4.6+/Opus 4.6+. Landing this PR's version first would delete the branch those PRs are patching; landing theirs first would have this PR undo their fix on rebase. Restores the original prefill-based continuation-resume behavior (including the is_pre_first_chunk guard already in litellm_internal_staging) in both _acompletion_streaming_iterator and _completion_streaming_iterator, and removes _stream_chunks_have_generated_content along with the tests that only existed to cover the guard. This PR now only touches the deferred-stream eager-fetch fix and the header-stripping fixes; the non-text-content re-raise idea becomes a follow-up PR built on top of whichever of BerriAI#30242/BerriAI#30743 lands.
…merge _handle_llm_api_exception filtered provider/framing headers once, then merged in post_call_response_headers_hook's return value afterward without re-filtering. The ProxyException branch happened to re-filter after its own header merge, but the HTTPException/httpx.HTTPStatusError/ generic-exception branches passed the post-hook headers straight through unfiltered, so a callback hook (any custom guardrail/logging plugin) returning an unsafe header would bypass the strip entirely for those paths. Filters once, right after the hook merge, so every branch gets the same guarantee.
…his PR" This reverts commit c5ca101.
…m guard Greptile flagged that a structured reasoning-only delta (Delta.reasoning_items, the OpenAI Responses-API-style reasoning item) wasn't recognized as already-streamed content by _stream_chunks_have_generated_content, alongside the existing thinking_blocks/tool_calls checks, so a stream that emitted only reasoning_items before failing could still restart via fallback.
…ence, not list The type_discipline_gate LIT001 check flags mutable-collection parameter annotations. chunks is only iterated, never mutated, so Sequence is the correct read-only annotation and clears the ratcheted budget ceiling.
05a8777 to
ebc3845
Compare
|
@greptileai please re-review |
|
@veria-ai please re-review |
|
bugbot run |
…apper, when mid-stream fallback gives up When content has already streamed and MidStreamFallbackError carries original_exception (e.g. RateLimitError), both the async and sync streaming iterators bare-re-raised the wrapper itself, so the client lost the specific error type/code/provider_specific_fields instead of seeing the real provider error. The fallback-failure path a few lines below already unwraps to original_exception for the same reason; apply the same pattern here. Also extend _stream_chunks_have_generated_content to recognize audio, images, and annotations deltas as generated content, matching is_chunk_non_empty's existing annotations check and Delta's treatment of audio/images as first-class content fields — a stream carrying only one of these before failing was not recognized as already-streamed, so the router could still restart it via fallback after the client had received real content.
|
Round update: fixed both Cursor Bugbot findings in 70e47f4.
Also fixed the earlier Full local suite ( |
|
@greptileai please re-review |
frontend-lint's check-run shows conclusion=cancelled on 70e47f4 with no superseding run, and this PR touches no UI files. Verify schema.d.ts matches the proxy OpenAPI spec is on the previously diagnosed stream_timeout/user_role Union-ordering nondeterminism (e9fc5e5). Empty commit to force a fresh CI run for both rather than a manual rerun, which requires repo admin rights this fork PR doesn't have.
|
Pushed 2f251ac, an empty commit to retrigger CI. `frontend-lint` was showing conclusion=cancelled with no superseding run (this PR touches no UI files), and `Verify schema.d.ts...` is on the previously diagnosed stream_timeout/user_role ordering nondeterminism (e9fc5e5). No code changes in this commit; all 174 local router tests and `make pre-commit` still pass as of 70e47f4. |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 2f251ac. Configure here.
Annotate the three new header constants as Final[frozenset[str]] to match the Final sweep staging landed while this PR was open. Drop the incidental schema.d.ts union-ordering churn: this PR adds no API surface, so the file should match the base byte for byte, and the reordered unions are generator nondeterminism that reddens check-ui-api-types.
|
Merged current There was exactly one textual conflict, in Two small follow-ups on top of the merge, both to match what staging now expects: The three new header constants are annotated The Everything else auto-merged, and |
e2950a8
into
BerriAI:litellm_internal_staging
|
This merged while three follow-ups found during the conflict merge were still being verified, so they are in #35843 rather than here Two are worth flagging to anyone reading this PR later. The #35843 carries a live before/after for the sync half against real provider APIs, with the before leg run on a clean checkout of this merge commit |
…7.0) (#336) This PR contains the following updates: | Package | Update | Change | |---|---|---| | [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.96.2` → `v1.97.0` | --- ### Release Notes <details> <summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary> ### [`v1.97.0`](https://github.com/BerriAI/litellm/releases/tag/v1.97.0) [Compare Source](https://github.com/BerriAI/litellm/compare/v1.97.0...v1.97.0) ##### Verify Docker Image Signature All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0). **Verify using the pinned commit hash (recommended):** A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \ ghcr.io/berriai/litellm:v1.97.0 ``` **Verify using the release tag (convenience):** Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/v1.97.0/cosign.pub \ ghcr.io/berriai/litellm:v1.97.0 ``` Expected output: ``` The following checks were performed on each of these signatures: - The cosign claims were validated - The signatures were verified against the specified public key ``` *** ##### What's Changed - feat(proxy): resolve Cursor thinking/fast model-name suffixes on /cursor/chat/completions by [@​mateo-berri](https://github.com/mateo-berri) in [#​35554](https://github.com/BerriAI/litellm/pull/35554) - fix(team-callbacks): actually stop logging when disable\_logging is called by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35520](https://github.com/BerriAI/litellm/pull/35520) - refactor(lint): drop redundant !s f-string conversion flags and fix displaced import-group comments by [@​mateo-berri](https://github.com/mateo-berri) in [#​35546](https://github.com/BerriAI/litellm/pull/35546) - fix(proxy): backfill null user\_email on existing users during JWT auth by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34588](https://github.com/BerriAI/litellm/pull/34588) - feat(playground): add non-streaming response toggle by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35560](https://github.com/BerriAI/litellm/pull/35560) - feat(teams): apply default organization to new teams from default team settings by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35540](https://github.com/BerriAI/litellm/pull/35540) - fix(ui): block Playground page for viewer roles on direct URL access by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35676](https://github.com/BerriAI/litellm/pull/35676) - fix(caching): close evicted LLM clients so their connections are reclaimed by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35492](https://github.com/BerriAI/litellm/pull/35492) - chore(deps): update brace-expansion, postcss, and gitpython to current patch releases by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35692](https://github.com/BerriAI/litellm/pull/35692) - refactor(ui): rename the create MCP server component to PascalCase by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35686](https://github.com/BerriAI/litellm/pull/35686) - fix(openai): drop undefined Union from owns\_wrapped\_http\_client annotation by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35706](https://github.com/BerriAI/litellm/pull/35706) - fix(openai): drop the undefined Union from owns\_wrapped\_http\_client by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35704](https://github.com/BerriAI/litellm/pull/35704) - chore(ui): note Google's Agent Platform rename in vector store setup by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​28076](https://github.com/BerriAI/litellm/pull/28076) - fix(proxy): apply key/team router\_settings.model\_group\_alias by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35486](https://github.com/BerriAI/litellm/pull/35486) - feat(complexity\_router): default session affinity off and expose it in the UI by [@​tin-berri](https://github.com/tin-berri) in [#​35714](https://github.com/BerriAI/litellm/pull/35714) - fix(datadog): read team callback dd\_\* params from kwargs instead of blocked dynamic params ([#​35115](https://github.com/BerriAI/litellm/issues/35115) port) by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35687](https://github.com/BerriAI/litellm/pull/35687) - refactor(ui): extract the MCP create form's logic and field groups by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35694](https://github.com/BerriAI/litellm/pull/35694) - test(ui): tier the MCP create tests into unit and integration by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35697](https://github.com/BerriAI/litellm/pull/35697) - fix(proxy): redact credential headers from request logging copies by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35678](https://github.com/BerriAI/litellm/pull/35678) - feat(guardrails/rubrik): prompt moderation, response-text blocking, streaming buffer, failure logging by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35722](https://github.com/BerriAI/litellm/pull/35722) - fix(ui): render Responses API request and response in the logs drawer by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35718](https://github.com/BerriAI/litellm/pull/35718) - fix(ui): hide guardrail review buttons from non-admin users by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​27535](https://github.com/BerriAI/litellm/pull/27535) - feat(team): custom metadata validation hook for team create and update by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33353](https://github.com/BerriAI/litellm/pull/33353) - ci(circleci): install a pinned Rust toolchain on the Linux jobs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35519](https://github.com/BerriAI/litellm/pull/35519) - fix(bedrock): stop forwarding no-op toolSpec.strict to Converse by [@​tin-berri](https://github.com/tin-berri) in [#​35688](https://github.com/BerriAI/litellm/pull/35688) - fix(ui): reject an auto-router keyword rule left empty instead of dropping it by [@​tin-berri](https://github.com/tin-berri) in [#​35705](https://github.com/BerriAI/litellm/pull/35705) - fix(guardrails/rubrik): attribute blocked requests to the caller that made them by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35734](https://github.com/BerriAI/litellm/pull/35734) - fix(responses): forward client headers to the provider on /v1/responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34531](https://github.com/BerriAI/litellm/pull/34531) - feat(spend): add net auto-router savings to the cost-optimization dashboard by [@​tin-berri](https://github.com/tin-berri) in [#​35521](https://github.com/BerriAI/litellm/pull/35521) - chore(typing): clear basedpyright Any errors in budget reset, access groups, and cache settings by [@​mateo-berri](https://github.com/mateo-berri) in [#​35719](https://github.com/BerriAI/litellm/pull/35719) - fix(spend): read what a request cost from the record instead of pricing it again by [@​tin-berri](https://github.com/tin-berri) in [#​35736](https://github.com/BerriAI/litellm/pull/35736) - perf: install hiredis so redis-py parses replies with its C parser by [@​Classic298](https://github.com/Classic298) in [#​35709](https://github.com/BerriAI/litellm/pull/35709) - feat(ui): show auto-router savings on the cost-optimization dashboard by [@​tin-berri](https://github.com/tin-berri) in [#​35522](https://github.com/BerriAI/litellm/pull/35522) - perf: build log messages lazily so filtered-out log records cost nothing by [@​Classic298](https://github.com/Classic298) in [#​35703](https://github.com/BerriAI/litellm/pull/35703) - fix(proxy): retry model cost map fetch with Retry-After-aware backoff and keep current map on reload failure by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35739](https://github.com/BerriAI/litellm/pull/35739) - feat(otel): stamp service tier attributes on inference spans by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35679](https://github.com/BerriAI/litellm/pull/35679) - fix(proxy): log the model cost map reload failure lazily by [@​tin-berri](https://github.com/tin-berri) in [#​35750](https://github.com/BerriAI/litellm/pull/35750) - fix(groq): translate web\_search\_options to the browser\_search tool by [@​hMED22](https://github.com/hMED22) in [#​34971](https://github.com/BerriAI/litellm/pull/34971) - feat(ui): add admin-configurable user banner by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35729](https://github.com/BerriAI/litellm/pull/35729) - fix(e2e): make spend-counter redis connection env-driven for non-cluster deployments by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35732](https://github.com/BerriAI/litellm/pull/35732) - fix(proxy): make /cursor/chat/completions work with Cursor agent mode by [@​tin-berri](https://github.com/tin-berri) in [#​34029](https://github.com/BerriAI/litellm/pull/34029) - fix(proxy): propagate user\_email and bind api\_key on JWT auth attribution paths by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34331](https://github.com/BerriAI/litellm/pull/34331) - chore(build): move the Admin UI toolchain to Node 24 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35801](https://github.com/BerriAI/litellm/pull/35801) - test(e2e): vendor API strategy coverage across endpoints by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​34649](https://github.com/BerriAI/litellm/pull/34649) - chore(deps): upgrade cryptography to 50.0.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35803](https://github.com/BerriAI/litellm/pull/35803) - test(e2e): cover legacy text /completions endpoint by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​34431](https://github.com/BerriAI/litellm/pull/34431) - feat(gemini): add gemini-robotics-er-2-preview and gemini-robotics-er-1.6-preview by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35555](https://github.com/BerriAI/litellm/pull/35555) - test(e2e): move load/perf testing out of the main suite and drop the vllm passthrough test by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35820](https://github.com/BerriAI/litellm/pull/35820) - feat(lint): enforce Final on locals and freeze function parameters (LIT010, LIT011) by [@​mateo-berri](https://github.com/mateo-berri) in [#​35807](https://github.com/BerriAI/litellm/pull/35807) - chore: bump litellm-proxy-extras 0.4.81 -> 0.4.82, litellm 1.96.0 -> 1.97.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35810](https://github.com/BerriAI/litellm/pull/35810) - fix(bedrock): drop conflicting tool\_choice.type when toolConfig.toolChoice is set by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35738](https://github.com/BerriAI/litellm/pull/35738) - docs(CLAUDE.md): prefer commas over semicolons when replacing em dashes by [@​mateo-berri](https://github.com/mateo-berri) in [#​35825](https://github.com/BerriAI/litellm/pull/35825) - chore(lint): zero out basedpyright headroom for purely local rules by [@​mateo-berri](https://github.com/mateo-berri) in [#​35828](https://github.com/BerriAI/litellm/pull/35828) - test(e2e): retry provider-transient statuses at the transport with bounded backoff by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35824](https://github.com/BerriAI/litellm/pull/35824) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35836](https://github.com/BerriAI/litellm/pull/35836) - refactor(ui): route MCP session tokens through the shared storage helper by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35835](https://github.com/BerriAI/litellm/pull/35835) - docs(helm): replace the classic chart's 128Mi resource example with the documented 4Gi sizing by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35830](https://github.com/BerriAI/litellm/pull/35830) - fix(proxy): persist periodic reload schedule state so status survives restarts and fires without store\_model\_in\_db by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35165](https://github.com/BerriAI/litellm/pull/35165) - fix(router): eagerly fetch Vertex AI deferred stream to surface HTTP errors in \_acompletion fallback path by [@​deepanshululla](https://github.com/deepanshululla) in [#​34627](https://github.com/BerriAI/litellm/pull/34627) - fix(azure\_storage): honor AZURE\_STORAGE\_ENDPOINT\_SUFFIX for sovereign clouds by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35806](https://github.com/BerriAI/litellm/pull/35806) - fix(proxy): apply key\_alias/key\_hash filters to all /key/list visibility branches by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35840](https://github.com/BerriAI/litellm/pull/35840) - fix(proxy): enforce per-model budgets against resolved cursor model variants by [@​mateo-berri](https://github.com/mateo-berri) in [#​35834](https://github.com/BerriAI/litellm/pull/35834) - feat(ui): reorder Add Auto Router into name + template, with a collapsible detailed config by [@​tin-berri](https://github.com/tin-berri) in [#​35746](https://github.com/BerriAI/litellm/pull/35746) - test: repair three failing suites on litellm\_internal\_staging by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35845](https://github.com/BerriAI/litellm/pull/35845) - fix(guardrails): scan model output on the /openai/v1/responses alias by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35818](https://github.com/BerriAI/litellm/pull/35818) - ci: pin Node on the Playwright UI lanes so npm ci meets the engines floor by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35848](https://github.com/BerriAI/litellm/pull/35848) - fix(pricing): apply OpenAI's gpt-5.6 terra/luna cut to Azure cost map by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​35481](https://github.com/BerriAI/litellm/pull/35481) - feat(spend): add caller-scoped key/user/team/organization spend report endpoints by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35725](https://github.com/BerriAI/litellm/pull/35725) - revert: "fix(caching): close evicted LLM clients so their connections are reclaimed ([#​35492](https://github.com/BerriAI/litellm/issues/35492))" by [@​mateo-berri](https://github.com/mateo-berri) in [#​35856](https://github.com/BerriAI/litellm/pull/35856) - refactor(repositories): add prisma protocol seams and a spend-reset unit of work by [@​mateo-berri](https://github.com/mateo-berri) in [#​35748](https://github.com/BerriAI/litellm/pull/35748) - perf(streaming): assemble streamed tool-call arguments in linear time by [@​mateo-berri](https://github.com/mateo-berri) in [#​35826](https://github.com/BerriAI/litellm/pull/35826) - fix(s3\_v2): sign S3 object URLs with S3SigV4Auth so encoded paths verify by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35726](https://github.com/BerriAI/litellm/pull/35726) - test(e2e): self-seed the ui suite's password-login users in global setup by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35863](https://github.com/BerriAI/litellm/pull/35863) - fix(claude-code): create-only skill registration with a PUT update route (LIT-4110) by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​31752](https://github.com/BerriAI/litellm/pull/31752) - fix(proxy): fix zguard httpcode when block input by [@​jwang-gif](https://github.com/jwang-gif) in [#​31948](https://github.com/BerriAI/litellm/pull/31948) - fix(lint): pick the merge-aware base so in-progress merges are not blamed for base drift by [@​mateo-berri](https://github.com/mateo-berri) in [#​35868](https://github.com/BerriAI/litellm/pull/35868) - chore: bump litellm-proxy-extras 0.4.82 -> 0.4.83 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35877](https://github.com/BerriAI/litellm/pull/35877) - feat(ui): add Test Routing to the auto router create form by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35859](https://github.com/BerriAI/litellm/pull/35859) - fix(ui): derive auto-router preset tests from the bundled preset JSON by [@​tin-berri](https://github.com/tin-berri) in [#​35882](https://github.com/BerriAI/litellm/pull/35882) - revert: "test(e2e): vendor API strategy coverage across endpoints" ([#​34649](https://github.com/BerriAI/litellm/issues/34649)) by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35881](https://github.com/BerriAI/litellm/pull/35881) - chore(deps): bump grpc and golang.org/x modules in the terraform provider by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35844](https://github.com/BerriAI/litellm/pull/35844) - test(e2e): skip view-backed global spend probes pending LIT-5211 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35875](https://github.com/BerriAI/litellm/pull/35875) - fix(lint): move the basedpyright heap flag into the type check gate by [@​mateo-berri](https://github.com/mateo-berri) in [#​35869](https://github.com/BerriAI/litellm/pull/35869) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35876](https://github.com/BerriAI/litellm/pull/35876) - feat(ui): add role capability gating, migrate Tool Policies route by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35812](https://github.com/BerriAI/litellm/pull/35812) - refactor(ui): inject the fetch client's base url instead of reading it at import by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35802](https://github.com/BerriAI/litellm/pull/35802) - chore: remove unused .flake8 config and flake8 dev dependency by [@​mateo-berri](https://github.com/mateo-berri) in [#​35888](https://github.com/BerriAI/litellm/pull/35888) - chore: stop advising pre-commit and bootstrap by [@​mateo-berri](https://github.com/mateo-berri) in [#​35884](https://github.com/BerriAI/litellm/pull/35884) - fix(auth): name enable\_jwt\_auth when a JWT-shaped key is rejected by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35831](https://github.com/BerriAI/litellm/pull/35831) - feat(auto-router): make reminder marker pair configurable by [@​akapur99](https://github.com/akapur99) in [#​35874](https://github.com/BerriAI/litellm/pull/35874) - fix(UI): update anthropic model presets by [@​tin-berri](https://github.com/tin-berri) in [#​35896](https://github.com/BerriAI/litellm/pull/35896) - fix(bootstrap): switch to the dashboard node floor via nvm or fnm by [@​mateo-berri](https://github.com/mateo-berri) in [#​35895](https://github.com/BerriAI/litellm/pull/35895) - perf(pre-commit): run python, dashboard, and gen-api checks concurrently by [@​mateo-berri](https://github.com/mateo-berri) in [#​35903](https://github.com/BerriAI/litellm/pull/35903) - feat(spend): derive a default auto-router savings baseline from the hardest tier by [@​tin-berri](https://github.com/tin-berri) in [#​35907](https://github.com/BerriAI/litellm/pull/35907) - fix(http\_handler): self-heal handler clients closed after cache eviction by [@​mateo-berri](https://github.com/mateo-berri) in [#​35862](https://github.com/BerriAI/litellm/pull/35862) - fix(cost\_tracking): keep OpenAI prompt cache token details through usage reassembly by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34812](https://github.com/BerriAI/litellm/pull/34812) - fix(cost): bill gpt-5.6 prompt cache reads at the cache read rate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34957](https://github.com/BerriAI/litellm/pull/34957) - fix(batches): account for Responses API usage by [@​rimysore](https://github.com/rimysore) in [#​35367](https://github.com/BerriAI/litellm/pull/35367) - ci: retry Codecov uploads and stop failing jobs on OIDC token flakes by [@​mateo-berri](https://github.com/mateo-berri) in [#​35251](https://github.com/BerriAI/litellm/pull/35251) - feat(complexity\_router): let operators rename the four complexity tiers by [@​akapur99](https://github.com/akapur99) in [#​35893](https://github.com/BerriAI/litellm/pull/35893) - chore(lint): zero stale ruff and LIT headroom and strip inert type: ignore comments by [@​mateo-berri](https://github.com/mateo-berri) in [#​35928](https://github.com/BerriAI/litellm/pull/35928) - chore(lint): zero out seven more purely local basedpyright rules by [@​mateo-berri](https://github.com/mateo-berri) in [#​35927](https://github.com/BerriAI/litellm/pull/35927) - chore(ui): zero stale headroom on local dashboard eslint budgets by [@​mateo-berri](https://github.com/mateo-berri) in [#​35929](https://github.com/BerriAI/litellm/pull/35929) - fix(managed-files): skip rows without file objects by [@​rimysore](https://github.com/rimysore) in [#​35365](https://github.com/BerriAI/litellm/pull/35365) - fix(router): redact fallback tracebacks at the call site and cover the sync deferred stream by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35843](https://github.com/BerriAI/litellm/pull/35843) - fix(migrations): recover from an interrupted Prisma toolchain install by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35832](https://github.com/BerriAI/litellm/pull/35832) - fix(lint): bring basedpyright rule counts back under their budget limits by [@​mateo-berri](https://github.com/mateo-berri) in [#​35962](https://github.com/BerriAI/litellm/pull/35962) - chore(ui): don't zero out stale headroom except no-console by [@​mateo-berri](https://github.com/mateo-berri) in [#​35964](https://github.com/BerriAI/litellm/pull/35964) - fix(proxy): give proxy\_admin\_viewer read parity with proxy\_admin by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35851](https://github.com/BerriAI/litellm/pull/35851) - refactor(ui): address UI lint budget issues by refactoring UI by [@​tin-berri](https://github.com/tin-berri) in [#​35960](https://github.com/BerriAI/litellm/pull/35960) - fix(ci): make the env-key doc gate see get\_secret\_bool reads by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35833](https://github.com/BerriAI/litellm/pull/35833) - fix(caching): re-land evicted LLM client closing ([#​35492](https://github.com/BerriAI/litellm/issues/35492)) atop self-healing handlers by [@​mateo-berri](https://github.com/mateo-berri) in [#​35870](https://github.com/BerriAI/litellm/pull/35870) - fix(proxy): keep the connected DB client when a startup health check fails by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35837](https://github.com/BerriAI/litellm/pull/35837) - chore(lint): remove litellm/types from the ruff lint exclusion by [@​mateo-berri](https://github.com/mateo-berri) in [#​35926](https://github.com/BerriAI/litellm/pull/35926) - feat(sgr): make the gateway middleware the source of truth for successful requests by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35717](https://github.com/BerriAI/litellm/pull/35717) - feat(auto-router): let operators replace the LLM classifier's system prompt by [@​akapur99](https://github.com/akapur99) in [#​35855](https://github.com/BerriAI/litellm/pull/35855) - fix(docker): bake the pip image's prisma engines at a world-readable path by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35976](https://github.com/BerriAI/litellm/pull/35976) - fix(auth): return 403 from the OAuth2 enterprise gate by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35838](https://github.com/BerriAI/litellm/pull/35838) - fix(router): keep custom model\_info across a price data reload by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35491](https://github.com/BerriAI/litellm/pull/35491) - fix(proxy): resolve pass-through credentials live from router deployments by [@​mateo-berri](https://github.com/mateo-berri) in [#​35916](https://github.com/BerriAI/litellm/pull/35916) - fix(ci): fetch only head and merge-base in lint jobs instead of every branch by [@​mateo-berri](https://github.com/mateo-berri) in [#​35982](https://github.com/BerriAI/litellm/pull/35982) - fix(autorouter): match CJK keyword\_tier\_rules that regex word boundaries miss by [@​akapur99](https://github.com/akapur99) in [#​35984](https://github.com/BerriAI/litellm/pull/35984) - feat(spend): rebuild the auto-router benchmarks backend as a per-session rollup by [@​tin-berri](https://github.com/tin-berri) in [#​35910](https://github.com/BerriAI/litellm/pull/35910) - refactor(ui): replace hand-rolled query-param routing with nuqs by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35871](https://github.com/BerriAI/litellm/pull/35871) - fix(docker): bake the componentized prisma engines at /opt/prisma so any uid can start by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35989](https://github.com/BerriAI/litellm/pull/35989) - fix(migrations): keep the toolchain heal from raising on an unreadable nodeenv cache by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35986](https://github.com/BerriAI/litellm/pull/35986) - fix(bedrock): sign Bedrock managed-file S3 requests with S3SigV4Auth by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35983](https://github.com/BerriAI/litellm/pull/35983) - chore(typing): replace Any seams with real types across responses, proxy, and provider adapters by [@​mateo-berri](https://github.com/mateo-berri) in [#​35809](https://github.com/BerriAI/litellm/pull/35809) - fix(ai21): resolve the documented AI21\_API\_KEY instead of a misspelled name by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35985](https://github.com/BerriAI/litellm/pull/35985) - fix(docker): fail the image build when the generated prisma engine paths drift off /opt/prisma by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35979](https://github.com/BerriAI/litellm/pull/35979) - fix(jina\_ai): resolve the documented JINA\_API\_KEY as a fallback by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35992](https://github.com/BerriAI/litellm/pull/35992) - fix(proxy): only treat a recoverable database outage as grounds to serve without one by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35864](https://github.com/BerriAI/litellm/pull/35864) - fix(ci): make every remaining CI checkout shallow by [@​mateo-berri](https://github.com/mateo-berri) in [#​35997](https://github.com/BerriAI/litellm/pull/35997) - fix(auto-router): stop the embedding model's context window from failing long requests by [@​akapur99](https://github.com/akapur99) in [#​35956](https://github.com/BerriAI/litellm/pull/35956) - fix(ci): make the env-key doc gate see bare get\_secret and get\_secret\_str reads by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35996](https://github.com/BerriAI/litellm/pull/35996) - fix(logging): extend secret redaction to records litellm does not emit directly by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35977](https://github.com/BerriAI/litellm/pull/35977) - test(utils): pin the register\_model replay test to the recorded half by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35994](https://github.com/BerriAI/litellm/pull/35994) - fix(ci): run every helm test suite, not just the first one per file by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35993](https://github.com/BerriAI/litellm/pull/35993) - ci: fail the build when a test file or Dockerfile is invoked by no job by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35991](https://github.com/BerriAI/litellm/pull/35991) - fix(langfuse): stop a collected httpx handler from closing a shared client by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35981](https://github.com/BerriAI/litellm/pull/35981) - fix(bedrock): grant bedrock:CountTokens in OIDC session policy by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33145](https://github.com/BerriAI/litellm/pull/33145) - feat(pre-commit): save full lint output to a per-worktree log file by [@​mateo-berri](https://github.com/mateo-berri) in [#​36004](https://github.com/BerriAI/litellm/pull/36004) - feat(ui): match auto-router preset models against deployments' underlying model IDs by [@​tin-berri](https://github.com/tin-berri) in [#​35972](https://github.com/BerriAI/litellm/pull/35972) - fix(core\_helpers): map generic 'error' finish\_reason to 'stop' by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33972](https://github.com/BerriAI/litellm/pull/33972) - fix(proxy)!: apply request-parameter checks consistently across body, path and form inputs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36011](https://github.com/BerriAI/litellm/pull/36011) - fix: rebuild models\_by\_provider in add\_known\_models so cost map reloads reach wildcard expansion by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36010](https://github.com/BerriAI/litellm/pull/36010) - feat(complexity\_router): report LLM classifier cost per request via routing\_decision and x-litellm-classifier-cost header by [@​tin-berri](https://github.com/tin-berri) in [#​36015](https://github.com/BerriAI/litellm/pull/36015) - fix(model-prices): correct replicate model key typo by [@​AkashNaickar](https://github.com/AkashNaickar) in [#​34800](https://github.com/BerriAI/litellm/pull/34800) - fix(proxy): register managed batch output files on terminal retrieve by [@​Souravrajvi0](https://github.com/Souravrajvi0) in [#​34092](https://github.com/BerriAI/litellm/pull/34092) - perf(pre-commit): fetch basedpyright base counts from CI artifacts by [@​mateo-berri](https://github.com/mateo-berri) in [#​35970](https://github.com/BerriAI/litellm/pull/35970) - fix(ui): sync projects list page index to ?page= so back and reload keep the page by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36003](https://github.com/BerriAI/litellm/pull/36003) - fix(ui): link project page keys to their virtual key detail by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36002](https://github.com/BerriAI/litellm/pull/36002) - refactor(ui): drop unreferenced locals from dashboard route components by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35819](https://github.com/BerriAI/litellm/pull/35819) - fix(ui): opening a project now pushes ?project= so back and deep links work by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36001](https://github.com/BerriAI/litellm/pull/36001) - refactor(ui): drop unreferenced locals from shared dashboard components by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35821](https://github.com/BerriAI/litellm/pull/35821) - refactor(ui): drop unreferenced locals from tests and narrow destructures by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36025](https://github.com/BerriAI/litellm/pull/36025) - fix(guardrails): allow litellm\_content\_filter to run on post\_mcp\_call by [@​mateo-berri](https://github.com/mateo-berri) in [#​35980](https://github.com/BerriAI/litellm/pull/35980) - fix(guardrails): scan /v1/messages tool traffic by [@​mateo-berri](https://github.com/mateo-berri) in [#​35999](https://github.com/BerriAI/litellm/pull/35999) - refactor(ui): drop dead locals and unused React state across the dashboard by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36026](https://github.com/BerriAI/litellm/pull/36026) - feat(ui): add the auto-router usage tab to cost optimization by [@​tin-berri](https://github.com/tin-berri) in [#​35995](https://github.com/BerriAI/litellm/pull/35995) - fix(managed\_files): derive unified output file ids deterministically so concurrent registrations converge by [@​mateo-berri](https://github.com/mateo-berri) in [#​36019](https://github.com/BerriAI/litellm/pull/36019) - fix(proxy): send keepalive pings on anthropic messages SSE streams during upstream silence by [@​mateo-berri](https://github.com/mateo-berri) in [#​36024](https://github.com/BerriAI/litellm/pull/36024) - fix(managed\_files): return unified ids from unscoped file listing by [@​mateo-berri](https://github.com/mateo-berri) in [#​36031](https://github.com/BerriAI/litellm/pull/36031) - fix(arize\_phoenix): lowercase OTLP/gRPC auth metadata key by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34883](https://github.com/BerriAI/litellm/pull/34883) - fix(auto-router): accept every reminder marker pair a harness emits by [@​tin-berri](https://github.com/tin-berri) in [#​36029](https://github.com/BerriAI/litellm/pull/36029) - fix(pricing): sync flex/priority tier keys to dated OpenAI snapshot variants by [@​mateo-berri](https://github.com/mateo-berri) in [#​35923](https://github.com/BerriAI/litellm/pull/35923) - fix(cost): bill reasoning tokens at the service tier output rate by [@​mateo-berri](https://github.com/mateo-berri) in [#​35925](https://github.com/BerriAI/litellm/pull/35925) - fix(proxy): include today's UTC bucket when a daily activity range ends at the caller's current day by [@​tin-berri](https://github.com/tin-berri) in [#​36051](https://github.com/BerriAI/litellm/pull/36051) - fix: expired-miss share over all measured turns + cost-optimization tab labels by [@​tin-berri](https://github.com/tin-berri) in [#​36037](https://github.com/BerriAI/litellm/pull/36037) - fix(router): include Bedrock batch/S3 fields and model in deployment credentials by [@​mpcusack-altos](https://github.com/mpcusack-altos) in [#​24548](https://github.com/BerriAI/litellm/pull/24548) - fix(batch): track cost for managed batches with no attributable key/u… by [@​elinacse](https://github.com/elinacse) in [#​35468](https://github.com/BerriAI/litellm/pull/35468) - feat(guardrails): add scan\_only\_tool\_results to scope unified guardrails to tool results by [@​mateo-berri](https://github.com/mateo-berri) in [#​36014](https://github.com/BerriAI/litellm/pull/36014) - fix(cost): stop token-pricing the placeholder input on file content calls by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35140](https://github.com/BerriAI/litellm/pull/35140) - fix(proxy): fetch background responses through the router in CheckResponsesCost by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35137](https://github.com/BerriAI/litellm/pull/35137) - fix(proxy): yaml store\_prompts\_in\_spend\_logs should take precedence over DB cached value by [@​Praveena-617](https://github.com/Praveena-617) in [#​35769](https://github.com/BerriAI/litellm/pull/35769) - fix(lint): measure the basedpyright budget gate in a gate-owned venv by [@​mateo-berri](https://github.com/mateo-berri) in [#​36050](https://github.com/BerriAI/litellm/pull/36050) - docs: cap all GitHub comments at 15-25 words, curb semicolon splices by [@​mateo-berri](https://github.com/mateo-berri) in [#​36059](https://github.com/BerriAI/litellm/pull/36059) - chore(lint): name MappingProxyType in the mutable-collection fix messages by [@​mateo-berri](https://github.com/mateo-berri) in [#​36072](https://github.com/BerriAI/litellm/pull/36072) - test: roll back runtime model registrations between tests by [@​mateo-berri](https://github.com/mateo-berri) in [#​36039](https://github.com/BerriAI/litellm/pull/36039) - refactor(types): cut 653 implicit and explicit Any diagnostics across 11 modules by [@​mateo-berri](https://github.com/mateo-berri) in [#​36054](https://github.com/BerriAI/litellm/pull/36054) - fix(proxy): stop resolving the UI session sentinel team on /search\_tools/list by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36061](https://github.com/BerriAI/litellm/pull/36061) - fix(batches): persist managed file ids for cancelled/failed/expired batches by [@​mateo-berri](https://github.com/mateo-berri) in [#​36048](https://github.com/BerriAI/litellm/pull/36048) - fix(batches): register managed output files on batch cancel by [@​mateo-berri](https://github.com/mateo-berri) in [#​36034](https://github.com/BerriAI/litellm/pull/36034) - fix(proxy): allow non-admins to reach /user/daily/activity/aggregated by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36062](https://github.com/BerriAI/litellm/pull/36062) - fix(anthropic): coerce explicit additionalProperties to false in output\_format schema by [@​dkindlund](https://github.com/dkindlund) in [#​35811](https://github.com/BerriAI/litellm/pull/35811) - fix(batches): prevent managed file fallbacks by [@​rimysore](https://github.com/rimysore) in [#​35371](https://github.com/BerriAI/litellm/pull/35371) - chore: ignore the mechanical lint and typing sweeps in git blame by [@​mateo-berri](https://github.com/mateo-berri) in [#​36076](https://github.com/BerriAI/litellm/pull/36076) - fix(proxy): warn at startup when max\_budget is set but no database is connected by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36041](https://github.com/BerriAI/litellm/pull/36041) - fix(proxy): promote caller metadata trace fields into litellm\_metadata by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35866](https://github.com/BerriAI/litellm/pull/35866) - feat(terraform): sync provider 0.3.0 from the mirror and cut 0.4.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36098](https://github.com/BerriAI/litellm/pull/36098) - fix(guardrails): honor configured timeout in Zscaler AI Guard by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36110](https://github.com/BerriAI/litellm/pull/36110) - fix(logging): fall back to litellm\_metadata when metadata is empty by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36105](https://github.com/BerriAI/litellm/pull/36105) - fix(proxy): re-assert the authenticated identity on passthrough requests by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36121](https://github.com/BerriAI/litellm/pull/36121) - chore: bump litellm-enterprise 0.1.53 -> 0.1.54, litellm-proxy-extras 0.4.83 -> 0.4.84 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36139](https://github.com/BerriAI/litellm/pull/36139) - fix(ui): match auto-router preset models against wildcard-expanded model groups by [@​tin-berri](https://github.com/tin-berri) in [#​36111](https://github.com/BerriAI/litellm/pull/36111) - test(router): assert the auto-router max\_input\_chars kwarg by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36109](https://github.com/BerriAI/litellm/pull/36109) - fix(ui): allow clearing a key's budget reset from the Edit Key form by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36140](https://github.com/BerriAI/litellm/pull/36140) - fix(managed\_files): skip unparseable rows when listing managed files by [@​mateo-berri](https://github.com/mateo-berri) in [#​36021](https://github.com/BerriAI/litellm/pull/36021) - fix(a2a): stop writing per-caller headers onto the shared cached httpx client by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35978](https://github.com/BerriAI/litellm/pull/35978) - build(deps): bump h2 to 4.4.1 and js-yaml to 4.3.1 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36147](https://github.com/BerriAI/litellm/pull/36147) - chore: promote staging to main by [@​mateo-berri](https://github.com/mateo-berri) in [#​36057](https://github.com/BerriAI/litellm/pull/36057) - fix(azure\_sentinel): respect AZURE\_AUTHORITY\_HOST and derive the Azure Monitor audience per cloud by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36137](https://github.com/BerriAI/litellm/pull/36137) - fix(bedrock): pass SSE-KMS key through to the batch input-file S3 upload by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35148](https://github.com/BerriAI/litellm/pull/35148) - fix(anthropic adapter): stop indexing choices\[0] on choiceless streaming chunks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35314](https://github.com/BerriAI/litellm/pull/35314) - fix(bedrock): normalize /v1/completions and /v1/responses batch records by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35675](https://github.com/BerriAI/litellm/pull/35675) - fix(proxy): return the real status code when a credential update is rejected by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36166](https://github.com/BerriAI/litellm/pull/36166) - fix(proxy): improve Headroom /v1/compress HTTP 404 diagnostics by [@​aayush598](https://github.com/aayush598) in [#​35952](https://github.com/BerriAI/litellm/pull/35952) - fix(proxy): invalidate cached project object on project update and delete by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36028](https://github.com/BerriAI/litellm/pull/36028) - feat(proxy): add apply\_user\_budget\_to\_team\_keys opt-in by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36102](https://github.com/BerriAI/litellm/pull/36102) - fix(proxy): stop alerting on health probes that lose the planned engine-restart race by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36141](https://github.com/BerriAI/litellm/pull/36141) - test(docker): gate the componentized gateway and backend images on an arbitrary-uid offline boot by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36136](https://github.com/BerriAI/litellm/pull/36136) - fix(http): stop pooled clients persisting cookies on the aiohttp jar too by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36149](https://github.com/BerriAI/litellm/pull/36149) - fix(router): bound fallback-walk work and error-log volume by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36148](https://github.com/BerriAI/litellm/pull/36148) - ci: wire credential\_endpoints tests into the proxy endpoints job by [@​cursor](https://github.com/cursor)\[bot] in [#​36187](https://github.com/BerriAI/litellm/pull/36187) - docs(keys): document /key/info fields and clarify budget\_reset\_at is the next reset by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36127](https://github.com/BerriAI/litellm/pull/36127) - fix(azure\_sentinel): add AZURE\_SENTINEL\_AUTHORITY\_HOST as a Sentinel scoped override by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36165](https://github.com/BerriAI/litellm/pull/36165) - docs(pr-template): add a User Flow section with authoring instructions by [@​mateo-berri](https://github.com/mateo-berri) in [#​36162](https://github.com/BerriAI/litellm/pull/36162) - fix(proxy): derive config agent ids from agent\_name so grants survive secret rotation by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36020](https://github.com/BerriAI/litellm/pull/36020) - chore(ui): regenerate schema.d.ts for the /key/info docstring update by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36210](https://github.com/BerriAI/litellm/pull/36210) - build(deps): bump gitpython to 3.1.58 to clear osv-scan on staging by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36212](https://github.com/BerriAI/litellm/pull/36212) - fix(proxy): deny agent access when key and team grants resolve to nothing by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36221](https://github.com/BerriAI/litellm/pull/36221) - build(deps): defer the second pypdf advisory until the 6.15.0 bump by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36218](https://github.com/BerriAI/litellm/pull/36218) - fix(a2a): align agent list annotation and test with the tuple return type by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36217](https://github.com/BerriAI/litellm/pull/36217) - ci: always run the UI API types sync check so it can be required by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36213](https://github.com/BerriAI/litellm/pull/36213) - build(deps): bump nanoid to 3.3.17 in the dashboard lockfile by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36227](https://github.com/BerriAI/litellm/pull/36227) - feat(ui): show user email or alias in usage data export by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36232](https://github.com/BerriAI/litellm/pull/36232) - feat(auto-router): track turns per complexity tier (LIT-5302) by [@​tin-berri](https://github.com/tin-berri) in [#​36209](https://github.com/BerriAI/litellm/pull/36209) - fix(websearch): restore snippet text in native web\_search\_tool\_result blocks (LIT-5315) by [@​tin-berri](https://github.com/tin-berri) in [#​36228](https://github.com/BerriAI/litellm/pull/36228) - fix(proxy): resolve entity access groups in the model listing endpoints by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36230](https://github.com/BerriAI/litellm/pull/36230) - fix(ui): let access groups be a team's only model source, with hover provenance by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36234](https://github.com/BerriAI/litellm/pull/36234) - fix(managed\_files): return unified output file ids from GET /batches by [@​mateo-berri](https://github.com/mateo-berri) in [#​36049](https://github.com/BerriAI/litellm/pull/36049) - test(proxy): compare empty agent list to the tuple get\_agent\_list returns by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36225](https://github.com/BerriAI/litellm/pull/36225) - fix(otel): name the RPC system and upstream on MCP tool-call spans by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35857](https://github.com/BerriAI/litellm/pull/35857) - fix(guardrails): chunk oversized Bedrock ApplyGuardrail requests instead of failing by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36119](https://github.com/BerriAI/litellm/pull/36119) - test(e2e): settle control-plane writes across every replica, not just one by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36247](https://github.com/BerriAI/litellm/pull/36247) - fix(responses): forward allowed\_openai\_params through the chat completions bridge by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35885](https://github.com/BerriAI/litellm/pull/35885) - test(proxy): assert the copy \_add\_team\_member\_budget\_table returns by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36244](https://github.com/BerriAI/litellm/pull/36244) - chore(ui): regenerate dashboard api types for tier\_turns by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36243](https://github.com/BerriAI/litellm/pull/36243) - refactor(types): declare mirrored pricing fields on ModelInfo by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36215](https://github.com/BerriAI/litellm/pull/36215) - fix(lint): make strict-gate noqas survive base ruff and flag stale ones by [@​mateo-berri](https://github.com/mateo-berri) in [#​36257](https://github.com/BerriAI/litellm/pull/36257) - fix(vertex\_ai): surface real error/status on vertex batch create instead of IndexError 500 by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35141](https://github.com/BerriAI/litellm/pull/35141) - ci: give the remaining pull\_request workflows a concurrency group by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36252](https://github.com/BerriAI/litellm/pull/36252) - refactor(lint): graduate zero-violation strict rules and guard the budget ratchet by [@​mateo-berri](https://github.com/mateo-berri) in [#​36161](https://github.com/BerriAI/litellm/pull/36161) - fix(proxy): enforce require\_managed\_files on every route that accepts a raw provider id by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35551](https://github.com/BerriAI/litellm/pull/35551) - chore(typing): clear 1.4k basedpyright Any errors across 21 hotspot files by [@​mateo-berri](https://github.com/mateo-berri) in [#​36282](https://github.com/BerriAI/litellm/pull/36282) - test: roll back live router replay membership between tests by [@​mateo-berri](https://github.com/mateo-berri) in [#​36278](https://github.com/BerriAI/litellm/pull/36278) - chore(ci): sync main into internal staging by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36288](https://github.com/BerriAI/litellm/pull/36288) - build(lint): rename make pre-commit to make check with a working-tree fallback by [@​mateo-berri](https://github.com/mateo-berri) in [#​36277](https://github.com/BerriAI/litellm/pull/36277) - fix(ui): show team BYOK models in team fallback settings by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36241](https://github.com/BerriAI/litellm/pull/36241) - fix(otel): mark v2 server spans as failed for pre-call errors by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34546](https://github.com/BerriAI/litellm/pull/34546) - fix(websearch\_interception): bill intercepted searches to the calling key by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35708](https://github.com/BerriAI/litellm/pull/35708) - chore: remove pre-commit rule by [@​mateo-berri](https://github.com/mateo-berri) in [#​36295](https://github.com/BerriAI/litellm/pull/36295) - docs: clarify guideline priority ordering in CLAUDE.md by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36296](https://github.com/BerriAI/litellm/pull/36296) - feat(router): independent, default-on deployment affinity for the auto-router by [@​tin-berri](https://github.com/tin-berri) in [#​36146](https://github.com/BerriAI/litellm/pull/36146) - test: repair stale CircleCI contracts by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36293](https://github.com/BerriAI/litellm/pull/36293) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36286](https://github.com/BerriAI/litellm/pull/36286) - chore: rebuild Admin UI bundle for the 2026-08-08 release by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36297](https://github.com/BerriAI/litellm/pull/36297) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36304](https://github.com/BerriAI/litellm/pull/36304) ##### New Contributors - [@​rimysore](https://github.com/rimysore) made their first contribution in [#​35367](https://github.com/BerriAI/litellm/pull/35367) - [@​AkashNaickar](https://github.com/AkashNaickar) made their first contribution in [#​34800](https://github.com/BerriAI/litellm/pull/34800) - [@​Souravrajvi0](https://github.com/Souravrajvi0) made their first contribution in [#​34092](https://github.com/BerriAI/litellm/pull/34092) - [@​elinacse](https://github.com/elinacse) made their first contribution in [#​35468](https://github.com/BerriAI/litellm/pull/35468) - [@​aayush598](https://github.com/aayush598) made their first contribution in [#​35952](https://github.com/BerriAI/litellm/pull/35952) - [@​cursor](https://github.com/cursor)\[bot] made their first contribution in [#​36187](https://github.com/BerriAI/litellm/pull/36187) **Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.96.0...v1.97.0> ### [`v1.97.0`](https://github.com/BerriAI/litellm/releases/tag/v1.97.0) [Compare Source](https://github.com/BerriAI/litellm/compare/v1.96.2...v1.97.0) ##### Verify Docker Image Signature All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0). **Verify using the pinned commit hash (recommended):** A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \ ghcr.io/berriai/litellm:v1.97.0 ``` **Verify using the release tag (convenience):** Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/v1.97.0/cosign.pub \ ghcr.io/berriai/litellm:v1.97.0 ``` Expected output: ``` The following checks were performed on each of these signatures: - The cosign claims were validated - The signatures were verified against the specified public key ``` *** ##### What's Changed - feat(proxy): resolve Cursor thinking/fast model-name suffixes on /cursor/chat/completions by [@​mateo-berri](https://github.com/mateo-berri) in [#​35554](https://github.com/BerriAI/litellm/pull/35554) - fix(team-callbacks): actually stop logging when disable\_logging is called by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35520](https://github.com/BerriAI/litellm/pull/35520) - refactor(lint): drop redundant !s f-string conversion flags and fix displaced import-group comments by [@​mateo-berri](https://github.com/mateo-berri) in [#​35546](https://github.com/BerriAI/litellm/pull/35546) - fix(proxy): backfill null user\_email on existing users during JWT auth by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34588](https://github.com/BerriAI/litellm/pull/34588) - feat(playground): add non-streaming response toggle by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35560](https://github.com/BerriAI/litellm/pull/35560) - feat(teams): apply default organization to new teams from default team settings by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35540](https://github.com/BerriAI/litellm/pull/35540) - fix(ui): block Playground page for viewer roles on direct URL access by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35676](https://github.com/BerriAI/litellm/pull/35676) - fix(caching): close evicted LLM clients so their connections are reclaimed by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35492](https://github.com/BerriAI/litellm/pull/35492) - chore(deps): update brace-expansion, postcss, and gitpython to current patch releases by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35692](https://github.com/BerriAI/litellm/pull/35692) - refactor(ui): rename the create MCP server component to PascalCase by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35686](https://github.com/BerriAI/litellm/pull/35686) - fix(openai): drop undefined Union from owns\_wrapped\_http\_client annotation by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35706](https://github.com/BerriAI/litellm/pull/35706) - fix(openai): drop the undefined Union from owns\_wrapped\_http\_client by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35704](https://github.com/BerriAI/litellm/pull/35704) - chore(ui): note Google's Agent Platform rename in vector store setup by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​28076](https://github.com/BerriAI/litellm/pull/28076) - fix(proxy): apply key/team router\_settings.model\_group\_alias by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​35486](https://github.com/BerriAI/litellm/pull/35486) - feat(complexity\_router): default session affinity off and expose it in the UI by [@​tin-berri](https://github.com/tin-berri) in [#​35714](https://github.com/BerriAI/litellm/pull/35714) - fix(datadog): read team callback dd\_\* params from kwargs instead of blocked dynamic params ([#​35115](https://github.com/BerriAI/litellm/issues/35115) port) by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35687](https://github.com/BerriAI/litellm/pull/35687) - refactor(ui): extract the MCP create form's logic and field groups by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35694](https://github.com/BerriAI/litellm/pull/35694) - test(ui): tier the MCP create tests into unit and integration by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35697](https://github.com/BerriAI/litellm/pull/35697) - fix(proxy): redact credential headers from request logging copies by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35678](https://github.com/BerriAI/litellm/pull/35678) - feat(guardrails/rubrik): prompt moderation, response-text blocking, streaming buffer, failure logging by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35722](https://github.com/BerriAI/litellm/pull/35722) - fix(ui): render Responses API request and response in the logs drawer by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35718](https://github.com/BerriAI/litellm/pull/35718) - fix(ui): hide guardrail review buttons from non-admin users by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​27535](https://github.com/BerriAI/litellm/pull/27535) - feat(team): custom metadata validation hook for team create and update by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33353](https://github.com/BerriAI/litellm/pull/33353) - ci(circleci): install a pinned Rust toolchain on the Linux jobs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35519](https://github.com/BerriAI/litellm/pull/35519) - fix(bedrock): stop forwarding no-op toolSpec.strict to Converse by [@​tin-berri](https://github.com/tin-berri) in [#​35688](https://github.com/BerriAI/litellm/pull/35688) - fix(ui): reject an auto-router keyword rule left empty instead of dropping it by [@​tin-berri](https://github.com/tin-berri) in [#​35705](https://github.com/BerriAI/litellm/pull/35705) - fix(guardrails/rubrik): attribute blocked requests to the caller that made them by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35734](https://github.com/BerriAI/litellm/pull/35734) - fix(responses): forward client headers to the provider on /v1/responses by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34531](https://github.com/BerriAI/litellm/pull/34531) - feat(spend): add net auto-router savings to the cost-optimization dashboard by [@​tin-berri](https://github.com/tin-berri) in [#​35521](https://github.com/BerriAI/litellm/pull/35521) - chore(typing): clear basedpyright Any errors in budget reset, access groups, and cache settings by [@​mateo-berri](https://github.com/mateo-berri) in [#​35719](https://github.com/BerriAI/litellm/pull/35719) - fix(spend): read what a request cost from the record instead of pricing it again by [@​tin-berri](https://github.com/tin-berri) in [#​35736](https://github.com/BerriAI/litellm/pull/35736) - perf: install hiredis so redis-py parses replies with its C parser by [@​Classic298](https://github.com/Classic298) in [#​35709](https://github.com/BerriAI/litellm/pull/35709) - feat(ui): show auto-router savings on the cost-optimization dashboard by [@​tin-berri](https://github.com/tin-berri) in [#​35522](https://github.com/BerriAI/litellm/pull/35522) - perf: build log messages lazily so filtered-out log records cost nothing by [@​Classic298](https://github.com/Classic298) in [#​35703](https://github.com/BerriAI/litellm/pull/35703) - fix(proxy): retry model cost map fetch with Retry-After-aware backoff and keep current map on reload failure by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​35739](https://github.com/BerriAI/litellm/pull/35739) - feat(otel): stamp service tier attributes on inference spans by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35679](https://github.com/BerriAI/litellm/pull/35679) - fix(proxy): log the model cost map reload failure lazily by [@​tin-berri](https://github.com/tin-berri) in [#​35750](https://github.com/BerriAI/litellm/pull/35750) - fix(groq): translate web\_search\_options to the browser\_search tool by [@​hMED22](https://github.com/hMED22) in [#​34971](https://github.com/BerriAI/litellm/pull/34971) - feat(ui): add admin-configurable user banner by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35729](https://github.com/BerriAI/litellm/pull/35729) - fix(e2e): make spend-counter redis connection env-driven for non-cluster deployments by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​35732](https://github.com/BerriAI/litellm/pull/35732) - fix(proxy): make /cursor/chat/completions work with Cursor agent mode by [@​tin-berri](https://github.com/tin-berri) in [#​34029](https://github.com/BerriAI/litellm/pull/34029) - fix(proxy): propagate user\_email and bind api\_key on JWT auth attribution paths by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34331](https://github.com/BerriAI/litellm/pull/34331) - chore(build): move the Admin UI toolchain to Node 24 by [@​yuneng-berri](https://g…
TLDR
Problem this solves:
How it solves it:
Relevant issues
Closes #31874
Linear ticket
Pre-Submission checklist
@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Ran the proxy locally against real provider APIs with this config (a Vertex AI deployment pointed at a nonexistent project to force a real 403, and an explicit fallback to Anthropic Haiku):
The router log shows the fix firing:
router.pyraisesVertexAIError: ... PERMISSION_DENIED ...from inside the eager fetch, propagating through_acompletion's except block into the normal retry/fallback path instead of staying hidden inside the stream generator. The client-facing response comes back as a real, live streamed completion from the fallback deployment:Before the fix, the Vertex HTTP error stays hidden inside the generator and fallback never triggers; this run confirms the eager fetch surfaces it and the fallback delivers a real, working response.
Also did a live adversarial test of the header-stripping fix: drove a request through the real ASGI app with
litellm.acompletionpatched to raise an exception carrying attacker-controlled headers (x-frame-options: ALLOWALL,content-security-policy: default-src *,access-control-allow-origin: https://evil.example.com). With the fix, those get stripped and litellm's existingSecurityHeadersMiddlewaredefaults (DENY,frame-ancestors 'none') win; with the fix reverted, the malicious values leak straight through to the client. Legitimate headers (x-request-id) pass through unaffected in both cases. Full trace and before/after in the PR comments.The mid-stream re-raise guard and its
thinking_blocks/tool-call broadening are covered by unit tests rather than a fresh live run, since forcing a real provider to fail mid-stream after partial content (rather than immediately) isn't reliably reproducible against a live API.Type
🐛 Bug Fix
Changes
litellm/router.pyadds an eagerfetch_stream()call in_acompletionwhencompletion_stream is Noneandmake_call is not None, the deferred-HTTP pattern Vertex AI and Bedrock use. Any HTTP error fromfetch_stream()now propagates through_acompletion's existing except block, which records the failure, triggers cooldown, and runs the normal fallback chain._acompletion_streaming_iteratorand_completion_streaming_iterator(sync) re-raise aMidStreamFallbackErrorinstead of silently retrying once any content already streamed to the caller, whether that's text (generated_content), a tool-call/function-call/reasoning-only chunk, or a thinking-only chunk (Delta.thinking_blocks), all detected via the wrapper's rawchunks; retrying in that case would otherwise send a second, unrelated response after content the client already received.Note on scope: this touches the same mid-stream continuation-resume code path that #30242/#30743 are patching for a different reason (Anthropic's removal of assistant-message prefill on Sonnet 4.6+/Opus 4.6+). An earlier revision of this PR pulled the re-raise change out entirely to avoid that conflict, at a maintainer's request. It's back in because Greptile's automated review treats a stream that silently retries after partial content as a P1 finding and won't clear this PR without it. I'll rebase this specific piece onto whichever of #30242/#30743 lands, once one does, so
/chat/completionspicks up their capability-aware continuation-building instead of reverting to the plain prefill-resume this restores.litellm/proxy/common_request_processing.pystrips HTTP-framing headers (content-length,transfer-encoding,content-encoding,content-type,set-cookie,cookie,proxy-authenticate,proxy-authorization) and browser-facing security headers (access-control-allow-origin,content-security-policy,clear-site-data,strict-transport-security,x-frame-options,cross-origin-*-policy) from provider exception headers before they reach the client. This is applied once after the response-headers hook merge (so a callback hook returning an unsafe header can't bypass the strip) and again on the pre-existing-ProxyExceptionbranch (where a naive dict merge could otherwise let them through unstripped). This lives in the proxy layer rather thanRouter._acompletion, sinceRouteris also used directly as an SDK and stripping headers there would drop legitimate provider metadata for callers who never go through the proxy's response construction. The header constants live inlitellm/constants.pyrather thanrouter.pyto avoid a router-proxy import cycle.tests/test_litellm/test_router.pyadds tests for the eager-fetch behavior, the framing-header handling, and the mid-stream re-raise guard, including the tool-call/reasoning-only and thinking-block cases.tests/test_litellm/proxy/test_common_request_processing.pyadds tests for the proxy-layer header stripping, including the pre-existing-ProxyExceptionbranch, the browser-security-header case, and the response-headers-hook case.tests/test_litellm/test_redact_string_in_error_paths.pyadds a test confirming the router's fallback-failure log doesn't leak secrets throughexc_info, since the codebase'sSecretRedactionFilteralready redactsexc_info-rendered tracebacks regardless of whether the call site formats them explicitly.Final Attestation
Note
High Risk
Changes core router streaming/fallback behavior and proxy error response headers; mis-handling could break fallbacks or alter client-visible errors and security headers.
Overview
Fixes Vertex/Bedrock deferred streaming so HTTP failures happen inside
_acompletion’s error path (eagerfetch_stream()), enabling cooldown,fail_calls, and router fallbacks instead of errors surfacing only on first chunk read.Mid-stream streaming fallbacks no longer retry with continuation prompts once any output was already sent. Sync/async streaming iterators re-raise
MidStreamFallbackError(ororiginal_exceptionwhen set) when partial content exists, including non-text deltas detected via_stream_chunks_have_generated_content; pre-first-chunk failures still fall back with the original messages.The proxy strips provider exception headers in
UNSAFE_PROXY_RESPONSE_HEADERS(HTTP framing + browser security/CORS) before building error responses, including after the response-headers hook and on mergedProxyExceptionheaders. Header lists live inconstants.pyso the router/SDK keeps full provider metadata.Router fallback-failure logging switches to
exc_info=Trueinstead of embedding a redacted traceback string. Tests cover header stripping, deferred-stream behavior, mid-stream re-raise, and secret redaction in fallback logs.Reviewed by Cursor Bugbot for commit 2f251ac. Bugbot is set up for automated code reviews on this repo. Configure here.