fix(xai): bill web_search from server_side_tool_usage_details - #30817
Conversation
|
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
Greptile SummaryThe PR updates xAI web-search billing to use server-side tool call counts while preserving the Responses API usage schema
Confidence Score: 5/5The PR appears safe to merge No blocking failure remains
|
| Filename | Overview |
|---|---|
| litellm/llms/xai/cost_calculator.py | Bills reported web-search calls using model pricing when configured and the current xAI fallback otherwise |
| litellm/llms/xai/chat/transformation.py | Preserves xAI server-side tool usage details on normalized chat usage for downstream billing |
| litellm/llms/xai/responses/transformation.py | Removes the schema-breaking usage replacement so Responses clients retain input/output token fields |
| litellm/responses/utils.py | Generically carries provider usage extras into a logging-only chat Usage object without mutating client-visible responses |
| litellm/litellm_core_utils/llm_cost_calc/tool_call_cost_tracking.py | Detects reported server-side web-search calls and recognizes dictionary-shaped Responses output items |
Reviews (5): Last reviewed commit: "fix(cost-tracking): bill web searches re..." | Re-trigger Greptile
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 2 · PR risk: 0/10 |
403eb38 to
8f0c4a8
Compare
4b3d028 to
f0a9a2e
Compare
Use usage.server_side_tool_usage_details.web_search_calls at $5/1k calls instead of legacy num_sources_used/web_search_requests. Preserve tool usage details through Responses usage transform for accurate response cost.
Avoid hard-coding provider-specific usage keys in shared Responses utilities; forward any non-standard usage attributes onto chat Usage for provider cost tracking (e.g. server_side_tool_usage_details).
Revert shared responses/utils.py extras forwarding. Attach server_side_tool_usage_details on chat Usage inside XAIResponsesAPIConfig so cost calc keeps web_search_calls without provider logic in shared utils.
Also inspect model_extra when attaching server_side_tool_usage_details for Responses cost tracking.
Treat positive web_search_calls as a web-search signal in built-in tool cost gating, and mirror counts onto prompt_tokens_details.web_search_requests when attaching xAI tool usage details so charges are not skipped.
Use search_context_cost_per_query from the model cost map (with $5/1k fallback) so web search billing can change via pricing JSON updates.
Apply server_side_tool_usage_details on completed/incomplete/failed streaming events so stream=true web_search is billed like non-stream.
Add unit tests for server_side_tool_usage_details extraction/attach and streaming completed-event pass-through in XAIResponsesAPIConfig.
Add unit tests for apply_server_side_tool_usage_details_to_usage edge cases and model_info-driven web_search per-call pricing fallbacks.
Gate web search like OpenAI (output/annotations/web_search_requests). xAI uses server_side_tool_usage_details only for per-call cost math, with web_search_requests mirrored in llms/xai for existing gate compatibility.
xAI already converts Responses usage to chat Usage so web_search_calls survive cost tracking. The chat completions bridge then re-ran the Responses usage transform and crashed on missing input_tokens. Pass through already-chat Usage and chat-shaped dumps instead
f0a9a2e to
749a8b0
Compare
|
|
||
| chat_usage: Final = ResponseAPILoggingUtils._transform_response_api_usage_to_chat_usage(response.usage) | ||
| apply_server_side_tool_usage_details_to_usage(chat_usage, details) | ||
| response.usage = chat_usage # pyright: ignore[reportAttributeAccessIssue] # extra # rebind-ok: chat Usage |
There was a problem hiding this comment.
Responses usage schema is replaced
When an xAI Responses API result includes server_side_tool_usage_details, this assignment replaces its ResponseAPIUsage with a chat Usage object, causing SDK and proxy clients to receive prompt_tokens and completion_tokens instead of the contractually expected input_tokens and output_tokens.
Knowledge Base Used: LLM Provider Adapters
Drop the transform overrides that swapped response.usage to the chat shape, which broke the /v1/responses client contract. Provider extras like server_side_tool_usage_details already survive validation via ResponseAPIUsage extra fields, so the shared usage bridge now carries them onto the bridged chat Usage generically. The web_search_call output gate also reads dict output items, since items that fail SDK validation stay plain dicts, and the chat path gains billing tests.
…itellm_fix_xai_web_search_cost_billing
…ridge Gemini image usage carries prompt_tokens and friends as extra fields on ResponseAPIUsage, which collided with the bridge's explicit kwargs and raised TypeError. Exclude keys the bridge already sets explicitly.
|
bugbot run |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit d96f76c. Configure here.
mateo-berri
left a comment
There was a problem hiding this comment.
LGTM. Thank you for the contribution!
b4f5e46
into
BerriAI:litellm_internal_staging
…8.0) (#393)
This PR contains the following updates:
| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.97.0` → `v1.98.0` |
---
### Release Notes
<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>
### [`v1.98.0`](https://github.com/BerriAI/litellm/releases/tag/v1.98.0)
[Compare Source](https://github.com/BerriAI/litellm/compare/v1.98.0...v1.98.0)
##### Verify Docker Image Signature
All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).
**Verify using the pinned commit hash (recommended):**
A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
ghcr.io/berriai/litellm:v1.98.0
```
**Verify using the release tag (convenience):**
Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/v1.98.0/cosign.pub \
ghcr.io/berriai/litellm:v1.98.0
```
Expected output:
```
The following checks were performed on each of these signatures:
- The cosign claims were validated
- The signatures were verified against the specified public key
```
***
##### What's Changed
- fix(bedrock): drop toolSpec.strict for Claude Sonnet 5 on Converse by [@​kr0k](https://github.com/kr0k) in [#​33196](https://github.com/BerriAI/litellm/pull/33196)
- fix(batches): attribute Vertex passthrough batch cost to key/team/tags by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​34456](https://github.com/BerriAI/litellm/pull/34456)
- docs: rewrite the CLAUDE.md comment rule with explicit exceptions by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36301](https://github.com/BerriAI/litellm/pull/36301)
- fix(proxy): scope file list pagination cursors to the caller by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36093](https://github.com/BerriAI/litellm/pull/36093)
- fix(proxy): skip prisma-dependent hooks when no database is attached by [@​mateo-berri](https://github.com/mateo-berri) in [#​36273](https://github.com/BerriAI/litellm/pull/36273)
- fix(proxy): report has\_more false on caller-scoped file list pages by [@​mateo-berri](https://github.com/mateo-berri) in [#​36326](https://github.com/BerriAI/litellm/pull/36326)
- fix(proxy): restore management\_v1 query-param validation under fastapi>=0.140.7 by [@​HuanQian571](https://github.com/HuanQian571) in [#​35773](https://github.com/BerriAI/litellm/pull/35773)
- fix(proxy): stop /{provider}/v1/files from capturing /openai\_passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36092](https://github.com/BerriAI/litellm/pull/36092)
- chore(typing): remove 914 basedpyright Any errors across 16 hotspot files by [@​mateo-berri](https://github.com/mateo-berri) in [#​36386](https://github.com/BerriAI/litellm/pull/36386)
- fix(router): keep batch fallbacks inside the model group that owns the file by [@​mateo-berri](https://github.com/mateo-berri) in [#​36181](https://github.com/BerriAI/litellm/pull/36181)
- feat(ptu): configure provisioned-throughput flat cost on a model deployment by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35341](https://github.com/BerriAI/litellm/pull/35341)
- docs: clarify the CLAUDE.md comment exceptions are any-of by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36421](https://github.com/BerriAI/litellm/pull/36421)
- docs: replace the Changes PR template section with Caveats by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36423](https://github.com/BerriAI/litellm/pull/36423)
- fix(bedrock): enable native structured output for GLM 5 and DeepSeek V3.2 by [@​alexshtf](https://github.com/alexshtf) in [#​35669](https://github.com/BerriAI/litellm/pull/35669)
- feat(ptu): daily rollup writes per-model PTU flat cost by active hour by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35343](https://github.com/BerriAI/litellm/pull/35343)
- feat(logging): add opt-in session\_id and trace\_id correlation to JSON log records via contextvars by [@​deepanshululla](https://github.com/deepanshululla) in [#​34418](https://github.com/BerriAI/litellm/pull/34418)
- feat(ptu): surface PTU flat cost on the daily activity read path by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35391](https://github.com/BerriAI/litellm/pull/35391)
- feat(router): add per-deployment allowed\_fails\_policy and cooldown\_time override support by [@​deepanshululla](https://github.com/deepanshululla) in [#​34416](https://github.com/BerriAI/litellm/pull/34416)
- feat(ptu): add PTU inputs to the model form and flat cost to the Usage page by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35393](https://github.com/BerriAI/litellm/pull/35393)
- fix(cost): price dict-shaped image input token details at the image rate by [@​vairodp](https://github.com/vairodp) in [#​33490](https://github.com/BerriAI/litellm/pull/33490)
- fix(model\_prices): refresh deprecation dates, correct xAI pricing and add missing provider models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36403](https://github.com/BerriAI/litellm/pull/36403)
- feat(ptu): gate PTU flat-cost attribution behind an opt-in env var by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36138](https://github.com/BerriAI/litellm/pull/36138)
- ci: cache Prisma CLI and engine binaries, split test timeout from setup by [@​mateo-berri](https://github.com/mateo-berri) in [#​36417](https://github.com/BerriAI/litellm/pull/36417)
- feat(rate limiting): configurable estimated output tokens per key, team and model by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36143](https://github.com/BerriAI/litellm/pull/36143)
- fix(ui): hide admin-only Logs tabs from roles that cannot call their endpoints by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36333](https://github.com/BerriAI/litellm/pull/36333)
- test(proxy): guard management\_v1 against fastapi names removed in supported releases by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36336](https://github.com/BerriAI/litellm/pull/36336)
- fix(ui): gate policy and prompt lookups on an admin capability by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36335](https://github.com/BerriAI/litellm/pull/36335)
- build(deps): bump pypdf to 6.15.0 to clear osv-scan by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36350](https://github.com/BerriAI/litellm/pull/36350)
- fix(proxy): isolate guardrail load failures per row by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36432](https://github.com/BerriAI/litellm/pull/36432)
- fix(ui): gate organization and agent usage views behind capabilities by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36334](https://github.com/BerriAI/litellm/pull/36334)
- fix(reset\_budget\_job): atomic budget cascade with chunked reset scans by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36287](https://github.com/BerriAI/litellm/pull/36287)
- feat(proxy): add GET /v1/indexes to list vector store indexes by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36289](https://github.com/BerriAI/litellm/pull/36289)
- feat(ui): show vector store indexes on the Vector Stores page by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36306](https://github.com/BerriAI/litellm/pull/36306)
- fix(proxy): treat SAML as configured in UI SSO detection by [@​fancybear-dev](https://github.com/fancybear-dev) in [#​36196](https://github.com/BerriAI/litellm/pull/36196)
- fix(bedrock): reject Anthropic server-side web\_search tool with actionable error by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36473](https://github.com/BerriAI/litellm/pull/36473)
- fix(ui): open the classifier prompt editor above the edit auto-router form by [@​tin-berri](https://github.com/tin-berri) in [#​36438](https://github.com/BerriAI/litellm/pull/36438)
- fix(arize): trace MCP tool calls instead of crashing on CallToolResult by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36453](https://github.com/BerriAI/litellm/pull/36453)
- refactor(ui): make illegal DataTable prop combinations unrepresentable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36470](https://github.com/BerriAI/litellm/pull/36470)
- fix(ui): scope Virtual Keys and Logs team lists to the caller by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36472](https://github.com/BerriAI/litellm/pull/36472)
- fix(ui): gate the Old Usage page behind a proxy-admin capability by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36469](https://github.com/BerriAI/litellm/pull/36469)
- docs(terraform): describe the provider release as automatic by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36467](https://github.com/BerriAI/litellm/pull/36467)
- feat(proxy): add per-deployment keepalive\_seconds SSE heartbeat to prevent load-balancer timeout on long streams by [@​deepanshululla](https://github.com/deepanshululla) in [#​34423](https://github.com/BerriAI/litellm/pull/34423)
- fix(router): cool down failed fallback deployments and correct cooldown TTL after Redis backfill by [@​deepanshululla](https://github.com/deepanshululla) in [#​35104](https://github.com/BerriAI/litellm/pull/35104)
- perf(spend): write each daily spend batch in one upsert statement by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36448](https://github.com/BerriAI/litellm/pull/36448)
- fix(ui): gate four sidebar pages on the roles their endpoints allow by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36475](https://github.com/BerriAI/litellm/pull/36475)
- fix(ui): restore the Logs Deleted Teams tab for organization admins by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36478](https://github.com/BerriAI/litellm/pull/36478)
- fix(websearch): stop leaking interception control fields to providers by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36480](https://github.com/BerriAI/litellm/pull/36480)
- test(e2e): cover the Anthropic web\_search server tool on Bedrock by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36443](https://github.com/BerriAI/litellm/pull/36443)
- fix(router): warn when a deployment's credentials contradict its provider by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36486](https://github.com/BerriAI/litellm/pull/36486)
- fix: net prompt-caching savings against the cache-write premium by [@​tin-berri](https://github.com/tin-berri) in [#​36452](https://github.com/BerriAI/litellm/pull/36452)
- feat(ui): deployment affinity toggle for the auto-router by [@​tin-berri](https://github.com/tin-berri) in [#​36302](https://github.com/BerriAI/litellm/pull/36302)
- fix(bedrock): use deployment credentials for AWS requests by [@​daleselaji-dev](https://github.com/daleselaji-dev) in [#​36160](https://github.com/BerriAI/litellm/pull/36160)
- fix(anthropic): preserve midturn system corrections by [@​eugene-yao-zocdoc](https://github.com/eugene-yao-zocdoc) in [#​34290](https://github.com/BerriAI/litellm/pull/34290)
- fix(email): stop duplicate legacy invitation email and fix its onboarding link by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​36455](https://github.com/BerriAI/litellm/pull/36455)
- feat(ui): show models under each tier in routing benchmark chart by [@​tin-berri](https://github.com/tin-berri) in [#​36291](https://github.com/BerriAI/litellm/pull/36291)
- fix(proxy): inject streaming usage cost on openai passthrough streams by [@​mateo-berri](https://github.com/mateo-berri) in [#​36503](https://github.com/BerriAI/litellm/pull/36503)
- docs: require a user flow and live-proxy proof in bug reports by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36498](https://github.com/BerriAI/litellm/pull/36498)
- fix(proxy): add config\_updated\_at audit timestamp for virtual keys by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36488](https://github.com/BerriAI/litellm/pull/36488)
- docs: require a user flow and a stuck-at proof in feature requests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36500](https://github.com/BerriAI/litellm/pull/36500)
- feat(router): add required-AND (&) tag prefix and allow\_fail\_open flag by [@​deepanshululla](https://github.com/deepanshululla) in [#​36193](https://github.com/BerriAI/litellm/pull/36193)
- feat(proxy): per-key prompt caching toggle via enable\_prompt\_caching by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36466](https://github.com/BerriAI/litellm/pull/36466)
- fix(bedrock): send tool-search beta header for Haiku 4.5 on Invoke /v1/messages by [@​mateo-berri](https://github.com/mateo-berri) in [#​36502](https://github.com/BerriAI/litellm/pull/36502)
- fix(bedrock): preserve adaptive thinking effort through the /v1/messages bridge by [@​mateo-berri](https://github.com/mateo-berri) in [#​36507](https://github.com/BerriAI/litellm/pull/36507)
- ci: retry transient network fetch failures in lint workflow by [@​mateo-berri](https://github.com/mateo-berri) in [#​36563](https://github.com/BerriAI/litellm/pull/36563)
- fix(ui): stub useIsOrgAdmin in UsageTab tests so useCan needs no QueryClient by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36565](https://github.com/BerriAI/litellm/pull/36565)
- fix(alerting): dedupe scheduled Slack spend reports across pods by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36489](https://github.com/BerriAI/litellm/pull/36489)
- chore(typing): clear 1.6k basedpyright Any errors across 56 files by [@​mateo-berri](https://github.com/mateo-berri) in [#​36543](https://github.com/BerriAI/litellm/pull/36543)
- fix(bedrock): add text block to converse user messages carrying documents by [@​mateo-berri](https://github.com/mateo-berri) in [#​36499](https://github.com/BerriAI/litellm/pull/36499)
- fix(deps): ship boto3 with the base SDK so bedrock works out of the box by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​36568](https://github.com/BerriAI/litellm/pull/36568)
- fix(model\_prices): add provider-announced deprecation dates for Bedrock, Mistral, Cohere and Gemini models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36538](https://github.com/BerriAI/litellm/pull/36538)
- chore: bump litellm-enterprise 0.1.54 -> 0.1.55, litellm-proxy-extras 0.4.84 -> 0.4.85, litellm 1.97.0 -> 1.98.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36577](https://github.com/BerriAI/litellm/pull/36577)
- fix(bedrock\_guardrails): skip ApplyGuardrail when there is no content to scan by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36441](https://github.com/BerriAI/litellm/pull/36441)
- fix(e2e): assert on the gen-AI span that served the stream, not the span count by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36582](https://github.com/BerriAI/litellm/pull/36582)
- test(e2e): harden vendor API coverage by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​34557](https://github.com/BerriAI/litellm/pull/34557)
- test(e2e): add reproducers for passthrough and model budget gaps by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​34657](https://github.com/BerriAI/litellm/pull/34657)
- test(e2e): cover google-native generateContent framing and prometheus queue time by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​34650](https://github.com/BerriAI/litellm/pull/34650)
- chore(ci): promote internal staging to main by [@​tin-berri](https://github.com/tin-berri) in [#​36560](https://github.com/BerriAI/litellm/pull/36560)
- feat(router): make routing groups callable as virtual models and list them in /v1/models by [@​tin-berri](https://github.com/tin-berri) in [#​36519](https://github.com/BerriAI/litellm/pull/36519)
- fix(xai): bill web\_search from server\_side\_tool\_usage\_details by [@​geraint0923](https://github.com/geraint0923) in [#​30817](https://github.com/BerriAI/litellm/pull/30817)
- fix(responses): init completed\_response on bridge streaming iterator ([#​35411](https://github.com/BerriAI/litellm/issues/35411)) by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35413](https://github.com/BerriAI/litellm/pull/35413)
- fix(batches): attribute Anthropic passthrough batch cost to the creating key, team and tags by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36468](https://github.com/BerriAI/litellm/pull/36468)
- feat(dashscope): add latest Model Studio models to the cost map by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36496](https://github.com/BerriAI/litellm/pull/36496)
- fix(proxy): track streamed passthrough Responses cost by [@​william-xue](https://github.com/william-xue) in [#​36529](https://github.com/BerriAI/litellm/pull/36529)
- fix(model\_prices): advertise native structured output on every Bedrock DeepSeek V3.2 and GLM 5 id by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36597](https://github.com/BerriAI/litellm/pull/36597)
- test(bedrock): repoint live Claude tests off the retired Claude 3 Sonnet by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36600](https://github.com/BerriAI/litellm/pull/36600)
- fix(anthropic): preserve speed=fast in usage for /v1/messages and pass-through by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36447](https://github.com/BerriAI/litellm/pull/36447)
- fix(proxy): forward resolved provider and deployment pricing in /cost/estimate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35880](https://github.com/BerriAI/litellm/pull/35880)
- feat(proxy): global SSE keepalive ping interval for OpenAI-shaped streaming routes by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36154](https://github.com/BerriAI/litellm/pull/36154)
- fix(responses): preserve Codex namespace tool calls by [@​dcadenas](https://github.com/dcadenas) in [#​32536](https://github.com/BerriAI/litellm/pull/32536)
- fix(nvidia\_nim): preserve image passages and stop sending top\_k to /v1/ranking by [@​atomic](https://github.com/atomic) in [#​34177](https://github.com/BerriAI/litellm/pull/34177)
- fix: refactor HTTP handler initialization with client support by [@​Praveen11558](https://github.com/Praveen11558) in [#​30952](https://github.com/BerriAI/litellm/pull/30952)
- feat(lint): gate writable TypedDict fields with LIT012 by [@​mateo-berri](https://github.com/mateo-berri) in [#​36590](https://github.com/BerriAI/litellm/pull/36590)
- perf(proxy): stagger scheduled background jobs across jobs and pods by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36589](https://github.com/BerriAI/litellm/pull/36589)
- test: remove four mirror test files that exercise none of their module by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34635](https://github.com/BerriAI/litellm/pull/34635)
- fix(router): stop re-applying router-selecting request tags to the routed tier's deployments by [@​mateo-berri](https://github.com/mateo-berri) in [#​36628](https://github.com/BerriAI/litellm/pull/36628)
- test: remove tests that never execute by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36681](https://github.com/BerriAI/litellm/pull/36681)
- fix(ui): align spend and budget columns by [@​daniel-meismer-zocdoc](https://github.com/daniel-meismer-zocdoc) in [#​35176](https://github.com/BerriAI/litellm/pull/35176)
- test: rename tests that a later definition shadowed by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36685](https://github.com/BerriAI/litellm/pull/36685)
- fix(passthrough): carry the budget reservation into request metadata by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36592](https://github.com/BerriAI/litellm/pull/36592)
- fix(mcp): bound MCP client requests with a session read timeout by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36675](https://github.com/BerriAI/litellm/pull/36675)
- fix(proxy): log requests rejected for an unparsable body in spend logs by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36673](https://github.com/BerriAI/litellm/pull/36673)
- refactor(ui): migrate cost-optimization to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36629](https://github.com/BerriAI/litellm/pull/36629)
- refactor(ui): migrate cost-tracking to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36631](https://github.com/BerriAI/litellm/pull/36631)
- refactor(ui): migrate admin-panel to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36635](https://github.com/BerriAI/litellm/pull/36635)
- refactor(ui): migrate users dashboard to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36642](https://github.com/BerriAI/litellm/pull/36642)
- refactor(ui): migrate prompts to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36643](https://github.com/BerriAI/litellm/pull/36643)
- refactor(ui): migrate team settings to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36641](https://github.com/BerriAI/litellm/pull/36641)
- refactor(ui): migrate models-and-endpoints to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36648](https://github.com/BerriAI/litellm/pull/36648)
- refactor(ui): migrate policy impact popover to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36653](https://github.com/BerriAI/litellm/pull/36653)
- fix(proxy): expand config-defined model access groups when resolving team models for /v2/model/info by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​34211](https://github.com/BerriAI/litellm/pull/34211)
- fix(batches): strip NUL bytes from passthrough batch tags before the managed object write by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36688](https://github.com/BerriAI/litellm/pull/36688)
- test(e2e-ui): verify UI mutations against the API instead of trusting the toast by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36632](https://github.com/BerriAI/litellm/pull/36632)
- fix(proxy): serialize model reconciles so concurrent model writes stop evicting each other by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36687](https://github.com/BerriAI/litellm/pull/36687)
- chore(e2e): port the compat-matrix cron publisher to tests/e2e/claude\_code by [@​mateo-berri](https://github.com/mateo-berri) in [#​36465](https://github.com/BerriAI/litellm/pull/36465)
- fix(router): never price a strategy-router alias by [@​tin-berri](https://github.com/tin-berri) in [#​36691](https://github.com/BerriAI/litellm/pull/36691)
- feat(model\_prices): add NVIDIA Nemotron 3.5 Lightning on OpenRouter and DeepInfra by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36696](https://github.com/BerriAI/litellm/pull/36696)
- feat(terraform/aws): make VPC, Aurora, and Redis optional by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36676](https://github.com/BerriAI/litellm/pull/36676)
- feat(ui): warn in the Admin UI when no Redis is configured by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36495](https://github.com/BerriAI/litellm/pull/36495)
- fix(ui): show and edit key-level router settings on a virtual key by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36674](https://github.com/BerriAI/litellm/pull/36674)
- fix(router): forward auto-router alias params from the marker entry, not the first same-name deployment by [@​mateo-berri](https://github.com/mateo-berri) in [#​36626](https://github.com/BerriAI/litellm/pull/36626)
- fix(bedrock\_mantle): 1M context window and long-context pricing for GPT-5.6 Sol/Terra/Luna by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36698](https://github.com/BerriAI/litellm/pull/36698)
- fix(model\_prices): sync the Groq registry with Groq's docs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36664](https://github.com/BerriAI/litellm/pull/36664)
- fix(router): let untagged requests bypass a tagged pre-routing strategy on shared model names by [@​mateo-berri](https://github.com/mateo-berri) in [#​36627](https://github.com/BerriAI/litellm/pull/36627)
- fix(spend): stop losing spend log rows when a flush is cancelled by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34826](https://github.com/BerriAI/litellm/pull/34826)
- docs(claude): drop the @​ prefix from the PR template path by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36726](https://github.com/BerriAI/litellm/pull/36726)
- fix(langfuse): emit otel trace version and release on the keys langfuse v4 reads by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36702](https://github.com/BerriAI/litellm/pull/36702)
- test(interactions): follow Google spec drift replacing Turn with typed steps by [@​mateo-berri](https://github.com/mateo-berri) in [#​36730](https://github.com/BerriAI/litellm/pull/36730)
- refactor(ui): migrate team detail controls to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36695](https://github.com/BerriAI/litellm/pull/36695)
- refactor(ui): migrate guardrail and duration controls to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36693](https://github.com/BerriAI/litellm/pull/36693)
- refactor(ui): migrate guardrails-monitor, projects, logs to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34606](https://github.com/BerriAI/litellm/pull/34606)
- refactor(ui): migrate search and user controls to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36694](https://github.com/BerriAI/litellm/pull/36694)
- fix(guardrails): scan and re-emit raw Anthropic SSE streams in the bedrock post-call hook by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36598](https://github.com/BerriAI/litellm/pull/36598)
- fix(helm): render nodeSelector on the migrations job by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36747](https://github.com/BerriAI/litellm/pull/36747)
- fix(langfuse): coerce header-sourced mask and trace-update steering values by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36740](https://github.com/BerriAI/litellm/pull/36740)
- refactor(ui): migrate usage tables to shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36707](https://github.com/BerriAI/litellm/pull/36707)
- refactor(ui): migrate guardrails monitor table to shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36709](https://github.com/BerriAI/litellm/pull/36709)
- refactor(ui): migrate guardrails content tables to shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36708](https://github.com/BerriAI/litellm/pull/36708)
- feat(gemini): day-0 pricing for gemini-3.7-flash by [@​mateo-berri](https://github.com/mateo-berri) in [#​36792](https://github.com/BerriAI/litellm/pull/36792)
- ci: promote staging to main by [@​mateo-berri](https://github.com/mateo-berri) in [#​36725](https://github.com/BerriAI/litellm/pull/36725)
- build(deps): bump nanoid to 3.3.18 to clear osv-scan by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36787](https://github.com/BerriAI/litellm/pull/36787)
- fix(router): stop scoring system prompt text for code/technical complexity by [@​tin-berri](https://github.com/tin-berri) in [#​36721](https://github.com/BerriAI/litellm/pull/36721)
- feat(complexity\_router): calibrate the classifier rubric with worked examples, selectable per router by [@​tin-berri](https://github.com/tin-berri) in [#​36578](https://github.com/BerriAI/litellm/pull/36578)
- fix(interactions): map step and turn history to Responses API roles and content types by [@​mateo-berri](https://github.com/mateo-berri) in [#​36733](https://github.com/BerriAI/litellm/pull/36733)
- fix(ui): restore playground model filtering by endpoint by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​36130](https://github.com/BerriAI/litellm/pull/36130)
- fix(proxy/batches): stop forwarding custom\_llm\_provider twice in list and cancel by [@​anxkhn](https://github.com/anxkhn) in [#​32813](https://github.com/BerriAI/litellm/pull/32813)
- refactor(ui): migrate TokenFlow and JsonViewer to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36735](https://github.com/BerriAI/litellm/pull/36735)
- feat: pre-adoption shadow eval for the auto-router (blind pairwise judge, derived state) by [@​tin-berri](https://github.com/tin-berri) in [#​36587](https://github.com/BerriAI/litellm/pull/36587)
- refactor(ui): migrate SimpleMessageBlock and SimpleToolCallBlock to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36737](https://github.com/BerriAI/litellm/pull/36737)
- refactor(ui): migrate HistoryTree and CollapsibleMessage to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36738](https://github.com/BerriAI/litellm/pull/36738)
- refactor: replace Any with precise types across responses, proxy, and llms modules by [@​mateo-berri](https://github.com/mateo-berri) in [#​36763](https://github.com/BerriAI/litellm/pull/36763)
- refactor(ui): migrate TruncatedValue and OutputCard to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36739](https://github.com/BerriAI/litellm/pull/36739)
- refactor(ui): migrate SectionHeader and ToolsSection to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36793](https://github.com/BerriAI/litellm/pull/36793)
- feat(ui): migrate playground chat controls to shadcn by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​36129](https://github.com/BerriAI/litellm/pull/36129)
- feat(xai): day-0 pricing for grok-4.6 by [@​mateo-berri](https://github.com/mateo-berri) in [#​36805](https://github.com/BerriAI/litellm/pull/36805)
- feat(ui): highlight Auto Router in the navbar announcement by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36315](https://github.com/BerriAI/litellm/pull/36315)
- test(e2e): assert the model allow-list permits, not only denies by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36823](https://github.com/BerriAI/litellm/pull/36823)
- fix(proxy): tolerate a concurrent creator when creating spend views by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36824](https://github.com/BerriAI/litellm/pull/36824)
- fix(proxy): honor explicit null budget\_duration on team and key create + clearable UI dropdowns by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​36699](https://github.com/BerriAI/litellm/pull/36699)
- feat(model\_prices): add meta/muse-spark-1.2 and its contributor tier by [@​mateo-berri](https://github.com/mateo-berri) in [#​36717](https://github.com/BerriAI/litellm/pull/36717)
- fix(auth): carry team grants in lite login session tokens by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36826](https://github.com/BerriAI/litellm/pull/36826)
- feat(ui): show provider prompt cache tokens in chat response metrics by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36827](https://github.com/BerriAI/litellm/pull/36827)
- fix(auth): stop the team fallback from widening model access by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36837](https://github.com/BerriAI/litellm/pull/36837)
- fix(proxy/team): resolve member\_delete cleanup by user id, not the addressed email by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36839](https://github.com/BerriAI/litellm/pull/36839)
- fix(cli): launch agents as a child process on Windows by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36822](https://github.com/BerriAI/litellm/pull/36822)
- feat(ui): shadow evals tab beside auto-router usage by [@​tin-berri](https://github.com/tin-berri) in [#​36588](https://github.com/BerriAI/litellm/pull/36588)
- feat(cli): make the hidden `lite` command list configurable by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36816](https://github.com/BerriAI/litellm/pull/36816)
- feat(azure\_ai): add Fireworks FW model pricing on Azure AI Foundry by [@​emerzon](https://github.com/emerzon) in [#​35613](https://github.com/BerriAI/litellm/pull/35613)
- fix: enable xhigh reasoning support for gpt-5.4-mini models by [@​emerzon](https://github.com/emerzon) in [#​26909](https://github.com/BerriAI/litellm/pull/26909)
- feat(azure-ai): add Grok 4.3 model metadata by [@​emerzon](https://github.com/emerzon) in [#​27932](https://github.com/BerriAI/litellm/pull/27932)
- feat(ui): render request metrics on the /ui/chat surface by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36845](https://github.com/BerriAI/litellm/pull/36845)
- fix(ui): stop a deselected MCP server keeping its grant on a virtual key by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36840](https://github.com/BerriAI/litellm/pull/36840)
- fix(team): sweep dangling team references and cache on team delete by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36819](https://github.com/BerriAI/litellm/pull/36819)
- fix(mcp): resolve admin OAuth sessions from any worker via DB-backed drafts by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36844](https://github.com/BerriAI/litellm/pull/36844)
- refactor(ui): migrate usage to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36834](https://github.com/BerriAI/litellm/pull/36834)
- refactor(ui): migrate guardrails-monitor to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36838](https://github.com/BerriAI/litellm/pull/36838)
- refactor(ui): migrate playground to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36847](https://github.com/BerriAI/litellm/pull/36847)
- refactor(ui): migrate guardrails to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36832](https://github.com/BerriAI/litellm/pull/36832)
- fix(batches): stop uncostable batches from starving the cost poll page by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36714](https://github.com/BerriAI/litellm/pull/36714)
- perf(spend-logs): bound retention cleanup so one run cannot saturate the database by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36594](https://github.com/BerriAI/litellm/pull/36594)
- fix(proxy): fail config load when a callbacks entry is not dispatchable by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36858](https://github.com/BerriAI/litellm/pull/36858)
- fix(bedrock): hoist custom.defer\_loading before dropping custom on invoke tools by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36855](https://github.com/BerriAI/litellm/pull/36855)
- fix(access groups): sync assigned\_key\_ids from the key write paths by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36843](https://github.com/BerriAI/litellm/pull/36843)
- fix(mcp): expose client HTTP headers to logging callbacks and hooks by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36724](https://github.com/BerriAI/litellm/pull/36724)
- fix(ptu): stop per-token billing on a PTU-configured deployment by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36829](https://github.com/BerriAI/litellm/pull/36829)
- fix(ui): add nvidia riva to the model provider list by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36769](https://github.com/BerriAI/litellm/pull/36769)
- fix(scripts): end make check with a ran/skipped summary and verdict by [@​mateo-berri](https://github.com/mateo-berri) in [#​36864](https://github.com/BerriAI/litellm/pull/36864)
- fix(proxy): track spend for OpenAI passthrough /v1/embeddings by [@​lostmartian](https://github.com/lostmartian) in [#​36660](https://github.com/BerriAI/litellm/pull/36660)
- test(proxy): stop monkeypatch.undo re-planting fixture-mocked prisma\_client by [@​mateo-berri](https://github.com/mateo-berri) in [#​36872](https://github.com/BerriAI/litellm/pull/36872)
- fix(access groups): sync assigned\_team\_ids from the team write paths by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36825](https://github.com/BerriAI/litellm/pull/36825)
- ci: drop the CircleCI ui\_build and ui\_unit\_tests jobs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36893](https://github.com/BerriAI/litellm/pull/36893)
- fix(langfuse)!: source the emitted metadata blob from StandardLoggingPayload by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36744](https://github.com/BerriAI/litellm/pull/36744)
- refactor(ui): migrate Navbar off antd to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36902](https://github.com/BerriAI/litellm/pull/36902)
- refactor(ui): migrate log details drawer off antd to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36904](https://github.com/BerriAI/litellm/pull/36904)
- refactor(ui): migrate AI Hub off antd and tremor to shadcn by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36908](https://github.com/BerriAI/litellm/pull/36908)
- refactor(ui): move the shared dropdowns and selectors onto shadcn primitives by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36924](https://github.com/BerriAI/litellm/pull/36924)
- refactor(ui): move the root-level dashboard components onto shadcn primitives by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36927](https://github.com/BerriAI/litellm/pull/36927)
- refactor(ui): move the settings page and bulk user invite onto shadcn primitives by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36936](https://github.com/BerriAI/litellm/pull/36936)
- refactor(ui): move the cost tracking components onto shadcn primitives by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36955](https://github.com/BerriAI/litellm/pull/36955)
- ci: drop the duplicate proxy\_unit\_tests letter-shard workflow by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36866](https://github.com/BerriAI/litellm/pull/36866)
- refactor(ui): migrate shared common\_components off antd and tremor by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36910](https://github.com/BerriAI/litellm/pull/36910)
- refactor(ui): migrate key info and permissions views off antd and tremor by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36913](https://github.com/BerriAI/litellm/pull/36913)
- feat(proxy): serve Anthropic-native /v1/models for Claude Code gateway discovery by [@​Ar-maan05](https://github.com/Ar-maan05) in [#​35455](https://github.com/BerriAI/litellm/pull/35455)
- refactor(ui): migrate router settings and shared badges off antd and tremor by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36915](https://github.com/BerriAI/litellm/pull/36915)
- refactor(ui): move the model hub and model select onto shadcn primitives by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36918](https://github.com/BerriAI/litellm/pull/36918)
- fix(ui): keep the cost tracking removal confirmation open until it settles by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36960](https://github.com/BerriAI/litellm/pull/36960)
- refactor(ui): declare DateRangePickerValue locally instead of importing it from tremor by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36962](https://github.com/BerriAI/litellm/pull/36962)
- fix(main): an explicit provider outranks a known OpenAI model name by [@​FahimaGold](https://github.com/FahimaGold) in [#​36800](https://github.com/BerriAI/litellm/pull/36800)
- fix(exception\_mapping): bare 429 in an error body no longer outranks the status code by [@​FahimaGold](https://github.com/FahimaGold) in [#​36705](https://github.com/BerriAI/litellm/pull/36705)
- refactor(ui): move MCP permission panels onto shadcn primitives by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36964](https://github.com/BerriAI/litellm/pull/36964)
- refactor(ui): migrate ten small dashboard files off antd and tremor by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36966](https://github.com/BerriAI/litellm/pull/36966)
- fix(proxy): force prisma recreate on postgres cached-plan error by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36428](https://github.com/BerriAI/litellm/pull/36428)
- fix(transcription): stop a zero output rate from zeroing transcription cost by [@​hMED22](https://github.com/hMED22) in [#​36914](https://github.com/BerriAI/litellm/pull/36914)
- fix(langfuse): restrict trace steering keys to real langfuse trace fields by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36862](https://github.com/BerriAI/litellm/pull/36862)
- Revert "fix(auth): stop the team fallback from widening model access" ([#​36837](https://github.com/BerriAI/litellm/issues/36837)) by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36982](https://github.com/BerriAI/litellm/pull/36982)
- fix(ui): show zeroed auto-router usage stats when a window has no sessions by [@​tin-berri](https://github.com/tin-berri) in [#​36868](https://github.com/BerriAI/litellm/pull/36868)
- fix(mcp): keep admin-entered oauth endpoints in management reads by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36888](https://github.com/BerriAI/litellm/pull/36888)
- fix(ui): distinguish hosted and local vLLM in the provider dropdown by [@​mateo-berri](https://github.com/mateo-berri) in [#​36974](https://github.com/BerriAI/litellm/pull/36974)
- fix(openai,azure): return a length-truncated 200 when the output budget fits no token by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36859](https://github.com/BerriAI/litellm/pull/36859)
- fix(proxy): always emit the Anthropic /v1/models token limits, null when unknown by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36961](https://github.com/BerriAI/litellm/pull/36961)
- feat(helm): add startupProbe and hpa.behavior to the componentized chart by [@​Louis-Vauterin](https://github.com/Louis-Vauterin) in [#​36382](https://github.com/BerriAI/litellm/pull/36382)
- fix(proxy): serve aggregate MCP endpoint on bare /mcp instead of 307-redirecting by [@​tin-berri](https://github.com/tin-berri) in [#​34845](https://github.com/BerriAI/litellm/pull/34845)
- feat(shadow\_eval): add reverse-direction shadow eval jobs by [@​tin-berri](https://github.com/tin-berri) in [#​36865](https://github.com/BerriAI/litellm/pull/36865)
- fix(proxy): requeue Redis spend buffer transactions when the DB commit fails by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33881](https://github.com/BerriAI/litellm/pull/33881)
- feat(search): add Nimble as a search provider by [@​ilchemla](https://github.com/ilchemla) in [#​36347](https://github.com/BerriAI/litellm/pull/36347)
- fix(mcp): drop caller host and configured upstream headers from logged metadata by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​36901](https://github.com/BerriAI/litellm/pull/36901)
- fix(azure\_ai): recognize real Search doc endpoints so teams can read/write via passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36798](https://github.com/BerriAI/litellm/pull/36798)
- fix(anthropic): aggregate 5m/1h cache-write split across iterations path by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34860](https://github.com/BerriAI/litellm/pull/34860)
- fix(anthropic cost): apply regional geo uplift to cached tokens by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34850](https://github.com/BerriAI/litellm/pull/34850)
- fix(ui): match the MCP servers count badge to its sibling permission badges by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36984](https://github.com/BerriAI/litellm/pull/36984)
- fix(batches): mark terminal batch with no output file as processed in CheckBatchCost by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35360](https://github.com/BerriAI/litellm/pull/35360)
- fix(caching): cache anthropic /v1/messages responses, including streaming by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​34581](https://github.com/BerriAI/litellm/pull/34581)
- fix(anthropic\_messages): make tool\_result images visible to OpenAI-compatible providers by [@​hMED22](https://github.com/hMED22) in [#​34462](https://github.com/BerriAI/litellm/pull/34462)
- feat(fireworks\_ai): translate NIM/vLLM extra params to Fireworks-native args by [@​milesadkins](https://github.com/milesadkins) in [#​35969](https://github.com/BerriAI/litellm/pull/35969)
- fix(ui): stop the models tab strip from scrolling vertically by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36993](https://github.com/BerriAI/litellm/pull/36993)
- fix(ui): anchor chips-combobox popups to the field instead of the inner input by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36995](https://github.com/BerriAI/litellm/pull/36995)
- feat(proxy): per-component response cost headers by [@​erensh27](https://github.com/erensh27) in [#​36965](https://github.com/BerriAI/litellm/pull/36965)
- fix(cost): track OpenAI/Azure web search tool cost per call by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35286](https://github.com/BerriAI/litellm/pull/35286)
- fix(bedrock): resolve aliases in batch file records by [@​daleselaji-dev](https://github.com/daleselaji-dev) in [#​36159](https://github.com/BerriAI/litellm/pull/36159)
- fix: report real token usage on guardrail-blocked /v1/responses replies by [@​guptaishaan](https://github.com/guptaishaan) in [#​36907](https://github.com/BerriAI/litellm/pull/36907)
- fix(proxy): requeue spend logs when the DB write fails with a transport error by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36716](https://github.com/BerriAI/litellm/pull/36716)
- fix(cost): tiered pricing supports cache creation cost and is all-or-nothing by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36720](https://github.com/BerriAI/litellm/pull/36720)
- fix(vertex\_ai): translate /v1/embeddings batch rows to the Gemini embedding shape by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​35092](https://github.com/BerriAI/litellm/pull/35092)
- docs(claude): require ReadOnly on every TypedDict field (LIT012) by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37005](https://github.com/BerriAI/litellm/pull/37005)
- refactor(ui): migrate access group create modal to RHF + zod + shadcn by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37033](https://github.com/BerriAI/litellm/pull/37033)
- refactor(ui): re-sync badge and skeleton onto the base-vega shadcn style by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​36991](https://github.com/BerriAI/litellm/pull/36991)
- feat(ui): link user detail team names to team pages by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37022](https://github.com/BerriAI/litellm/pull/37022)
- fix(model\_prices): correct DeepSeek V4 max output tokens by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36925](https://github.com/BerriAI/litellm/pull/36925)
- fix(ui): rename models table Status column to Source by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​37021](https://github.com/BerriAI/litellm/pull/37021)
- chore: bump litellm-enterprise 0.1.55 -> 0.1.56, litellm-proxy-extras 0.4.85 -> 0.4.86 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37045](https://github.com/BerriAI/litellm/pull/37045)
- feat(proxy): gate the Global Control Plane worker registry on an enterprise license by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​36996](https://github.com/BerriAI/litellm/pull/36996)
- fix(model\_prices): add gemini 3.1 flash tts preview and legacy OpenAI shutdown dates by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36788](https://github.com/BerriAI/litellm/pull/36788)
- fix(panw\_prisma\_airs): surface scan\_id on allowed requests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37037](https://github.com/BerriAI/litellm/pull/37037)
- fix(model\_map): flag native structured outputs on Anthropic-direct claude-sonnet-5 and claude-haiku-4-5 by [@​anmolg1997](https://github.com/anmolg1997) in [#​35930](https://github.com/BerriAI/litellm/pull/35930)
- fix(router): stop get\_router\_model\_info from wiping cached pricing by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36985](https://github.com/BerriAI/litellm/pull/36985)
- fix(redis): unwrap decorated \_\_init\_\_s when deriving the from\_url kwargs allowlist by [@​anmolg1997](https://github.com/anmolg1997) in [#​36654](https://github.com/BerriAI/litellm/pull/36654)
- fix(proxy): reserve the larger declared output budget for TPM limits by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​37001](https://github.com/BerriAI/litellm/pull/37001)
- fix(databricks): surface provider usage, including prompt-cache counts, in streaming chunks by [@​pokepoke81](https://github.com/pokepoke81) in [#​36943](https://github.com/BerriAI/litellm/pull/36943)
- fix(spend): give a batch's cost row a primary key of its own by [@​marty-sullivan](https://github.com/marty-sullivan) in [#​36876](https://github.com/BerriAI/litellm/pull/36876)
- feat: shadow eval samples /v1/messages and /v1/responses traffic by [@​tin-berri](https://github.com/tin-berri) in [#​36830](https://github.com/BerriAI/litellm/pull/36830)
- fix(ptu): stop a PTU deployment billing for grounded search by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​37043](https://github.com/BerriAI/litellm/pull/37043)
- fix(fireworks\_ai): support router slugs via routers/ prefix by [@​heathriel](https://github.com/heathriel) in [#​34257](https://github.com/BerriAI/litellm/pull/34257)
- fix(bedrock): register managed-batch litellm\_params so they stop leaking to the provider (internal copy of [#​36633](https://github.com/BerriAI/litellm/issues/36633)) by [@​mateo-berri](https://github.com/mateo-berri) in [#​37048](https://github.com/BerriAI/litellm/pull/37048)
- fix(bedrock): resolve the managed-batch output bucket on every path that reads it by [@​mateo-berri](https://github.com/mateo-berri) in [#​37047](https://github.com/BerriAI/litellm/pull/37047)
- fix(bedrock): resolve the managed-batch output bucket on every path that reads it by [@​marty-sullivan](https://github.com/marty-sullivan) in [#​36634](https://github.com/BerriAI/litellm/pull/36634)
- feat(scripts): queue heavy gates behind a machine-wide slot lock by [@​mateo-berri](https://github.com/mateo-berri) in [#​36988](https://github.com/BerriAI/litellm/pull/36988)
- feat(mcp): scope gateway session bearers to the RFC 8707 resource by [@​tin-berri](https://github.com/tin-berri) in [#​35045](https://github.com/BerriAI/litellm/pull/35045)
- feat(ui): direction picker and reverse-mode display for shadow evals by [@​tin-berri](https://github.com/tin-berri) in [#​36994](https://github.com/BerriAI/litellm/pull/36994)
- fix(guardrails): return the full PANW AIRS scan response on blocked requests by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37036](https://github.com/BerriAI/litellm/pull/37036)
- fix(passthrough): stop forwarding client Accept-Encoding upstream by [@​mateo-berri](https://github.com/mateo-berri) in [#​37058](https://github.com/BerriAI/litellm/pull/37058)
- fix(batches): account a managed batch's cost exactly once by [@​mateo-berri](https://github.com/mateo-berri) in [#​37050](https://github.com/BerriAI/litellm/pull/37050)
- fix(panw\_prisma\_airs): scan tool call args as plain text, not a tool\_event by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37038](https://github.com/BerriAI/litellm/pull/37038)
- feat(lint): exempt TypedDict-annotated dict literals from LIT002 by [@​mateo-berri](https://github.com/mateo-berri) in [#​36869](https://github.com/BerriAI/litellm/pull/36869)
- docs(claude): tell agents to let heavy gates queue for machine-wide slots by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​37057](https://github.com/BerriAI/litellm/pull/37057)
- test: unstick the suites CircleCI is failing on by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37059](https://github.com/BerriAI/litellm/pull/37059)
- docs(github): proof-of-fix section shows only the latest run as Before/After with nested cases by [@​mateo-berri](https://github.com/mateo-berri) in [#​37063](https://github.com/BerriAI/litellm/pull/37063)
- test(e2e): assert provider error shape instead of pinned prose by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37065](https://github.com/BerriAI/litellm/pull/37065)
- fix(ui): de-duplicate the reset budget option and polish shadcn surfaces by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37010](https://github.com/BerriAI/litellm/pull/37010)
- chore: rebuild Admin UI bundle from litellm\_internal\_staging by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37066](https://github.com/BerriAI/litellm/pull/37066)
- test(e2e/ui): assert the log drawer chevrons by their lucide classes by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37069](https://github.com/BerriAI/litellm/pull/37069)
- chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37042](https://github.com/BerriAI/litellm/pull/37042)
- fix(ui): keep completion-mode models in the playground chat dropdown (backport to rc/1.98.0) by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​37955](https://github.com/BerriAI/litellm/pull/37955)
##### New Contributors
- [@​kr0k](https://github.com/kr0k) made their first contribution in [#​33196](https://github.com/BerriAI/litellm/pull/33196)
- [@​HuanQian571](https://github.com/HuanQian571) made their first contribution in [#​35773](https://github.com/BerriAI/litellm/pull/35773)
- [@​alexshtf](https://github.com/alexshtf) made their first contribution in [#​35669](https://github.com/BerriAI/litellm/pull/35669)
- [@​vairodp](https://github.com/vairodp) made their first contribution in [#​33490](https://github.com/BerriAI/litellm/pull/33490)
- [@​fancybear-dev](https://github.com/fancybear-dev) made their first contribution in [#​36196](https://github.com/BerriAI/litellm/pull/36196)
- [@​daleselaji-dev](https://github.com/daleselaji-dev) made their first contribution in [#​36160](https://github.com/BerriAI/litellm/pull/36160)
- [@​eugene-yao-zocdoc](https://github.com/eugene-yao-zocdoc) made their first contribution in [#​34290](https://github.com/BerriAI/litellm/pull/34290)
- [@​geraint0923](https://github.com/geraint0923) made their first contribution in [#​30817](https://github.com/BerriAI/litellm/pull/30817)
- [@​william-xue](https://github.com/william-xue) made their first contribution in [#​36529](https://github.com/BerriAI/litellm/pull/36529)
- [@​dcadenas](https://github.com/dcadenas) made their first contribution in [#​32536](https://github.com/BerriAI/litellm/pull/32536)
- [@​atomic](https://github.com/atomic) made their first contribution in [#​34177](https://github.com/BerriAI/litellm/pull/34177)
- [@​Praveen11558](https://github.com/Praveen11558) made their first contribution in [#​30952](https://github.com/BerriAI/litellm/pull/30952)
- [@​anxkhn](https://github.com/anxkhn) made their first contribution in [#​32813](https://github.com/BerriAI/litellm/pull/32813)
- [@​lostmartian](https://github.com/lostmartian) made their first contribution in [#​36660](https://github.com/BerriAI/litellm/pull/36660)
- [@​FahimaGold](https://github.com/FahimaGold) made their first contribution in [#​36800](https://github.com/BerriAI/litellm/pull/36800)
- [@​Louis-Vauterin](https://github.com/Louis-Vauterin) made their first contribution in [#​36382](https://github.com/BerriAI/litellm/pull/36382)
- [@​ilchemla](https://github.com/ilchemla) made their first contribution in [#​36347](https://github.com/BerriAI/litellm/pull/36347)
- [@​milesadkins](https://github.com/milesadkins) made their first contribution in [#​35969](https://github.com/BerriAI/litellm/pull/35969)
- [@​erensh27](https://github.com/erensh27) made their first contribution in [#​36965](https://github.com/BerriAI/litellm/pull/36965)
- [@​guptaishaan](https://github.com/guptaishaan) made their first contribution in [#​36907](https://github.com/BerriAI/litellm/pull/36907)
- [@​pokepoke81](https://github.com/pokepoke81) made their first contribution in [#​36943](https://github.com/BerriAI/litellm/pull/36943)
- [@​heathriel](https://github.com/heathriel) made their first contribution in [#​34257](https://github.com/BerriAI/litellm/pull/34257)
**Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.97.0...v1.98.0>
### [`v1.98.0`](https://github.com/BerriAI/litellm/releases/tag/v1.98.0)
[Compare Source](https://github.com/BerriAI/litellm/compare/v1.97.0...v1.98.0)
##### Verify Docker Image Signature
All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).
**Verify using the pinned commit hash (recommended):**
A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
ghcr.io/berriai/litellm:v1.98.0
```
**Verify using the release tag (convenience):**
Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:
```bash
cosign verify \
--key https://raw.githubusercontent.com/BerriAI/litellm/v1.98.0/cosign.pub \
ghcr.io/berriai/litellm:v1.98.0
```
Expected output:
```
The following checks were performed on each of these signatures:
- The cosign claims were validated
- The signatures were verified against the specified public key
```
***
##### What's Changed
- fix(bedrock): drop toolSpec.strict for Claude Sonnet 5 on Converse by [@​kr0k](https://github.com/kr0k) in [#​33196](https://github.com/BerriAI/litellm/pull/33196)
- fix(batches): attribute Vertex passthrough batch cost to key/team/tags by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​34456](https://github.com/BerriAI/litellm/pull/34456)
- docs: rewrite the CLAUDE.md comment rule with explicit exceptions by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36301](https://github.com/BerriAI/litellm/pull/36301)
- fix(proxy): scope file list pagination cursors to the caller by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36093](https://github.com/BerriAI/litellm/pull/36093)
- fix(proxy): skip prisma-dependent hooks when no database is attached by [@​mateo-berri](https://github.com/mateo-berri) in [#​36273](https://github.com/BerriAI/litellm/pull/36273)
- fix(proxy): report has\_more false on caller-scoped file list pages by [@​mateo-berri](https://github.com/mateo-berri) in [#​36326](https://github.com/BerriAI/litellm/pull/36326)
- fix(proxy): restore management\_v1 query-param validation under fastapi>=0.140.7 by [@​HuanQian571](https://github.com/HuanQian571) in [#​35773](https://github.com/BerriAI/litellm/pull/35773)
- fix(proxy): stop /{provider}/v1/files from capturing /openai\_passthrough by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36092](https://github.com/BerriAI/litellm/pull/36092)
- chore(typing): remove 914 basedpyright Any errors across 16 hotspot files by [@​mateo-berri](https://github.com/mateo-berri) in [#​36386](https://github.com/BerriAI/litellm/pull/36386)
- fix(router): keep batch fallbacks inside the model group that owns the file by [@​mateo-berri](https://github.com/mateo-berri) in [#​36181](https://github.com/BerriAI/litellm/pull/36181)
- feat(ptu): configure provisioned-throughput flat cost on a model deployment by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​35341](https://github.com/BerriAI/litellm/pull/35341)
- docs: clarify the CLAUDE.md comment exceptions are any-of by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​36421](https://github.com/BerriAI/litellm/pull/36421)
- docs: replace the Changes …
TLDR
Problem this solves:
num_sources_usedat $25/1k sourcesserver_side_tool_usage_details.web_search_callsinsteadHow it solves it:
web_search_callsat xAI tools pricing, $5/1k callssearch_context_cost_per_queryfrom the model map when set/v1/responsesusage schema untouched: tool details already survive as extra fieldsUser Flow
Before: a developer who gives Grok the web_search tool gets searches and cited answers back, but the gateway bills every search at $0, so tracked spend under-reports the real xAI invoice
"model": "xai/grok-4.5","tools": [{"type": "web_search"}], and a question about today's newsusageblock shows"server_side_tool_usage_details": {"web_search_calls": 2}x-litellm-response-costheader and https://litellm-domain/ui/?page=logs show token cost only, with nothing added for the 2 searches"stream": trueends the same way: the finalresponse.completedevent carries the same usage and the logged spend again omits the searchesAfter: the same request bills each search call at xAI's tools rate, so tracked spend matches the xAI invoice
"model": "xai/grok-4.5","tools": [{"type": "web_search"}], and a question about today's newsusageblock shows"server_side_tool_usage_details": {"web_search_calls": 2}in the same OpenAI Responses shape as beforex-litellm-response-costheader and https://litellm-domain/ui/?page=logs now show token cost plus $0.005 per search call, $0.01 for the 2 calls"stream": truebills identically off the finalresponse.completedeventsearch_context_cost_per_queryon the modelRelevant issues
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Live before/after QA against real xAI grok-4.5, one proxy booted per commit with its own fresh DB: the before leg at the merge base bea3187, the after leg at this PR's head d96f76c. On every call below, xAI's own
usage.cost_in_usd_ticksequals token cost plus $0.005 perweb_search_calls, so the provider itself corroborates the billed rate to the last digit/v1/responses, non-streaming
Before at bea3187: usage reported
"server_side_tool_usage_details": {"web_search_calls": 4}yetx-litellm-response-cost: 0.040014and the spend log matched token-only cost exactly, while xAI's ticks said the call really cost $0.060014, so all 4 searches went unbilledAfter at d96f76c: usage reported
web_search_calls: 1, andx-litellm-response-cost: 0.022532= $0.017532 tokens + 1 x $0.005, spend log identical and equal to xAI's own ticks, with the usage block still in the OpenAI Responses schema (input_tokens/output_tokens, details carried alongside)/v1/responses, streaming
Same request body plus
"stream": true, billed off the finalresponse.completedevent (no cost header exists on a stream, so the spend log is the evidence)Before at bea3187: final event usage showed
web_search_calls: 4, spend log 0.0483732 = token-only exactly, versus $0.0683732 per xAI's ticksAfter at d96f76c: final event usage showed
web_search_calls: 1in the unchanged Responses schema, spend log 0.022794 = $0.017794 tokens + 1 x $0.005, equal to xAI's own ticks/v1/chat/completions
Before at bea3187: the client-facing chat usage carried no web-search fields at all even though the debug log showed xAI reporting
web_search_calls: 1upstream, andx-litellm-response-cost: 0.0167504was token-only, versus $0.0217504 per xAI's ticksAfter at d96f76c: chat usage carried
web_search_calls: 7andx-litellm-response-cost: 0.095734= $0.060734 tokens + 7 x $0.005, spend log identical and equal to xAI's own ticks/v1/messages
Before at bea3187: the searches ran (xAI reported
web_search_calls: 7upstream) and bothx-litellm-response-cost: 0.0831932and the spend log matched token-only cost exactly, while xAI's ticks said the call really cost $0.1181932, so all 7 searches went unbilledAfter at d96f76c: upstream reported
web_search_calls: 2and the spend log billed 0.033998 = $0.023998 tokens + 2 x $0.005, matching xAI's own ticks to the last digit, though this endpoint'sx-litellm-response-costheader still shows the token-only $0.023998 (see observations and caveats)Observations from the QA runs
num_sources_used: 0, the retired billing fieldmax_usescap, ran extra searches, pre-existingType
🐛 Bug Fix
Caveats (if any)
Final Attestation