test(e2e): otel trace completeness on streaming chat, messages, and responses (LIT-3787) - #33234
Conversation
Greptile SummaryAdds streaming OTEL trace completeness coverage alongside the existing non-streaming tests, targeting the stream-consumption span lifecycle (duplicate/orphaned gen-AI spans historically caused by #27757 and the #30590 startup race).
Confidence Score: 5/5Safe to merge; all changes are confined to tests/e2e and have no effect on production runtime behavior. The changes are test-only and well-documented with evidence of passing runs against the current proxy and failing runs against the foil commit. The only noted imperfection is that Anthropic two-line SSE errors store only the tests/e2e/e2e_http.py — the
|
| Filename | Overview |
|---|---|
| tests/e2e/e2e_http.py | Adds stream_error field to StreamingResponse and updates _streaming_outcome to capture the first SSE error event; Anthropic two-line errors lose detail because the control event: error line is captured before the data: line. |
| tests/e2e/logging/test_otel_trace_e2e.py | Adds three well-structured streaming OTEL trace completeness tests (chat, messages, responses); require_cost_span flag correctly relaxes the responses cost-span assertion per LIT-4428; _tag helper and _settled_names changes are clean. |
| tests/e2e/logging/logging_client.py | Adds stream parameter to messages_raw and responses_raw; stream=None for the non-streaming path is correctly excluded by exclude_none=True in the transport, preserving backward compatibility. |
| tests/e2e/models.py | Adds optional `stream: bool |
| tests/e2e/coverage_registry/logging.yaml | Registers new logging.otel.stream.exports_metric P0 coverage row correctly referencing all three streaming surfaces. |
Reviews (6): Last reviewed commit: "test(e2e): assert streamed responses spe..." | Re-trigger Greptile
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
d862a91 to
22ceb38
Compare
Adds the logging.otel.stream.exports_metric registry row (streaming closes the gen-AI span from the stream-consumption path, a lifecycle historically prone to duplicate or orphaned spans) and its first covering test: one streamed chat call must export ONE complete trace with the same tree contract as the non-streaming tests, plus stream-specific assertions - the response actually streamed (event-stream content type, at least one chunk), exactly one gen-AI span exists for the call, and the span records litellm.request.streaming=true
…7) (#33247) * test(e2e): OTEL trace completeness on streaming /v1/messages Covers the messages surface of logging.otel.stream.exports_metric: same tree contract as the non-streaming tests plus the stream-specific assertions (the response actually streamed, exactly one gen-AI span, and the span records litellm.request.streaming=true). Adds a stream field to the shared AnthropicMessagesBody and a stream mode to messages_raw * test(e2e): make the stream field coercion explicit per review * test(e2e): otel trace completeness on streaming /v1/responses (LIT-3787) (#33262) * test(e2e): OTEL trace completeness on streaming /v1/responses Covers the responses surface of logging.otel.stream.exports_metric: same tree contract plus the stream-side assertions (event-stream content type, chunks consumed, exactly one gen-AI span). Two assertions are knowingly relaxed on this surface, both verified against live traces and tracked in LIT-4428: the responses route does not stamp litellm.request.streaming on the gen-AI span, and the spend write for a streamed responses call records spend correctly but emits no batch_write_to_db cost span. Adds a stream mode to responses_raw * test(e2e): parenthesize the settle-names ternary and make stream coercion explicit per review
c77bd78 to
3cccd36
Compare
|
bugbot run |
… per review A streamed call commits its HTTP 200 before the upstream completes, so upstream failures arrive as SSE error events inside an otherwise-successful response; the consumed body is elided, so nothing could detect them before the trace assertions timed out with a misleading message. StreamingResponse now captures the first in-stream error event and the three streamed tests assert it is absent, so an upstream failure dies in seconds carrying the provider's actual error text. Kill-proof: with the quota-dead default model the responses test now fails immediately at this assertion quoting insufficient_quota
mateo-berri
left a comment
There was a problem hiding this comment.
The one finding is a diagnostic quality issue in the stream-error detection pattern that doesn't affect whether the OTEL assertions themselves pass or fail correctly
Legit or no? If legit, worth addressing?
The detection patterns only matched OpenAI's unspaced error JSON; Anthropic streams an 'event: error' control line with spaced JSON, which slipped past the guard. Covers both vendors' formats
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit ad61aae. Configure here.
… docstrings The responses-stream test now polls /spend/logs and requires a positive-spend aresponses row for the test key: the cost-write SPAN is knowingly absent on this surface (LIT-4428), so without this the test carried no billing witness at all, unlike its siblings whose required cost-write span is the instrumented spend write itself. Docstrings for the three streamed tests are author-written
817582e
into
litellm_internal_staging
…esponses (LIT-3787) (BerriAI#33234)
…4.0) (#201) This PR contains the following updates: | Package | Update | Change | |---|---|---| | [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.93.0` → `v1.94.0` | --- ### Release Notes <details> <summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary> ### [`v1.94.0`](https://github.com/BerriAI/litellm/releases/tag/v1.94.0) [Compare Source](https://github.com/BerriAI/litellm/compare/v1.94.0...v1.94.0) ##### Verify Docker Image Signature All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0). **Verify using the pinned commit hash (recommended):** A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \ ghcr.io/berriai/litellm:v1.94.0 ``` **Verify using the release tag (convenience):** Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/v1.94.0/cosign.pub \ ghcr.io/berriai/litellm:v1.94.0 ``` Expected output: ``` The following checks were performed on each of these signatures: - The cosign claims were validated - The signatures were verified against the specified public key ``` *** ##### What's Changed - feat(ui): working Test Connection for the complexity auto router by [@​akapur99](https://github.com/akapur99) in [#​32950](https://github.com/BerriAI/litellm/pull/32950) - fix(xecguard): use StandardLoggingGuardrailInformation in logging hook by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​32911](https://github.com/BerriAI/litellm/pull/32911) - feat(ui): adopt openapi-react-query ($api) and convert useCustomers by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32949](https://github.com/BerriAI/litellm/pull/32949) - refactor(ui): colocate the mcp-servers view, keeping the shared mcp\_tools surface by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32968](https://github.com/BerriAI/litellm/pull/32968) - refactor(ui): convert endpoint usage charts to shadcn/recharts by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32723](https://github.com/BerriAI/litellm/pull/32723) - fix(proxy-auth): stop unrecognized model namespaces slipping through provider wildcard keys by [@​mateo-berri](https://github.com/mateo-berri) in [#​32979](https://github.com/BerriAI/litellm/pull/32979) - feat(router): random-pick multi-model complexity tiers by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32967](https://github.com/BerriAI/litellm/pull/32967) - fix(xecguard): sanitize scan result before recording it for logging by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​32935](https://github.com/BerriAI/litellm/pull/32935) - fix(auto\_router): filter embedding models in complexity tab dropdowns, require all tiers, inline validation by [@​akapur99](https://github.com/akapur99) in [#​32978](https://github.com/BerriAI/litellm/pull/32978) - fix(anthropic): translate raw adaptive thinking for pre-4.6 models on chat completions and Bedrock Converse by [@​akapur99](https://github.com/akapur99) in [#​32944](https://github.com/BerriAI/litellm/pull/32944) - feat(router): add Router(plugins=\[...]) routing-plugin pipeline by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32972](https://github.com/BerriAI/litellm/pull/32972) - feat(router): soft-floor adaptive mode for complexity router by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32947](https://github.com/BerriAI/litellm/pull/32947) - docs(github): add QA runbook section to the PR template by [@​mateo-berri](https://github.com/mateo-berri) in [#​32965](https://github.com/BerriAI/litellm/pull/32965) - fix(model\_cost): add supports\_reasoning: false to Gemini image generation models by [@​mateo-berri](https://github.com/mateo-berri) in [#​32836](https://github.com/BerriAI/litellm/pull/32836) - build(dev-env): add make bootstrap and unprovisioned-checkout preflight to pre-commit by [@​mateo-berri](https://github.com/mateo-berri) in [#​32981](https://github.com/BerriAI/litellm/pull/32981) - ci(ui): report only error-level knip findings in CI by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32971](https://github.com/BerriAI/litellm/pull/32971) - feat(batches): track cost for unmanaged Bedrock batches, generalize the flag by [@​Sameerlite](https://github.com/Sameerlite) in [#​32315](https://github.com/BerriAI/litellm/pull/32315) - fix(guardrails): walk custom\_tool\_call\_output items in \_content\_utils by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​32969](https://github.com/BerriAI/litellm/pull/32969) - fix: show and allow editing team model aliases after team creation by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33047](https://github.com/BerriAI/litellm/pull/33047) - chore(deps): bump pillow to 12.3.0 to resolve osv-scan CVEs by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33093](https://github.com/BerriAI/litellm/pull/33093) - feat(mcp): mint gateway-bound envelope at the token endpoint for dcr\_bridge oauth\_delegate by [@​tin-berri](https://github.com/tin-berri) in [#​32828](https://github.com/BerriAI/litellm/pull/32828) - fix(mcp): surface rejected delegate-auth upstream tokens as connect-time 401 by [@​tin-berri](https://github.com/tin-berri) in [#​32741](https://github.com/BerriAI/litellm/pull/32741) - fix(proxy): track unauthenticated pass-through requests in spend logs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​32410](https://github.com/BerriAI/litellm/pull/32410) - feat(lasso): send source.type for Used By attribution by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33090](https://github.com/BerriAI/litellm/pull/33090) - fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@​thibault-linktree](https://github.com/thibault-linktree) in [#​33025](https://github.com/BerriAI/litellm/pull/33025) - fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@​tin-berri](https://github.com/tin-berri) in [#​33099](https://github.com/BerriAI/litellm/pull/33099) - fix(completion): forward aws credential kwargs into litellm\_params so the responses bridge keeps WIF auth by [@​mateo-berri](https://github.com/mateo-berri) in [#​32956](https://github.com/BerriAI/litellm/pull/32956) - fix(ui): respect litellm\_key\_header\_name in BYOK credential save and workflow runs fetches by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33103](https://github.com/BerriAI/litellm/pull/33103) - refactor(ui): standardize debounce waits behind shared DEBOUNCE\_WAIT\_MS constant by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33040](https://github.com/BerriAI/litellm/pull/33040) - feat(ui): rebuild the Virtual Keys table on the shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​32991](https://github.com/BerriAI/litellm/pull/32991) - fix: redact async complete streaming response for custom callbacks by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33106](https://github.com/BerriAI/litellm/pull/33106) - build(ui): bump [@​tanstack/react-pacer](https://github.com/tanstack/react-pacer) from 0.2.0 to 0.22.1 by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33041](https://github.com/BerriAI/litellm/pull/33041) - fix(ui): address Virtual Keys redesign review nits by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33112](https://github.com/BerriAI/litellm/pull/33112) - fix(openai/responses): clamp max\_output\_tokens below API minimum by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33098](https://github.com/BerriAI/litellm/pull/33098) - fix(prometheus): read v3 rate limiter remaining values for per-key model gauges by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33119](https://github.com/BerriAI/litellm/pull/33119) - fix(ui): drop w-full from page-content wrappers to remove 32px horizontal overflow by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33118](https://github.com/BerriAI/litellm/pull/33118) - refactor(ui): migrate straightforward value debounces to react-pacer by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33042](https://github.com/BerriAI/litellm/pull/33042) - feat(mcp): interactive SSO sign-in for dcr\_bridge oauth\_delegate DCR clients by [@​tin-berri](https://github.com/tin-berri) in [#​32946](https://github.com/BerriAI/litellm/pull/32946) - test(proxy): add regression tests for management\_endpoints edge cases by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​32976](https://github.com/BerriAI/litellm/pull/32976) - fix(auto-router): correct Responses API tool\_choice shape and propagate alias litellm\_params by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32974](https://github.com/BerriAI/litellm/pull/32974) - fix(ui): render the sidebar scrollbar with shadcn ScrollArea by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33124](https://github.com/BerriAI/litellm/pull/33124) - refactor(ui): migrate callback debounce sites to react-pacer with regression tests by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33043](https://github.com/BerriAI/litellm/pull/33043) - chore: add CODEOWNERS for ui and proxy UI build artifacts by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33131](https://github.com/BerriAI/litellm/pull/33131) - feat(mcp): client-held refresh envelope for the dcr\_bridge oauth\_delegate flow by [@​tin-berri](https://github.com/tin-berri) in [#​32980](https://github.com/BerriAI/litellm/pull/32980) - feat(ui): rebuild the Teams table on the shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33128](https://github.com/BerriAI/litellm/pull/33128) - fix(mcp): relay upstream OAuth token and DCR rejections instead of a generic 500 by [@​tin-berri](https://github.com/tin-berri) in [#​33113](https://github.com/BerriAI/litellm/pull/33113) - fix(keys): persist key\_type so the UI shows correct key scope instead of "All Proxy Models" by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33115](https://github.com/BerriAI/litellm/pull/33115) - feat(router): opt-in session affinity for complexity router by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​33126](https://github.com/BerriAI/litellm/pull/33126) - feat(prometheus): expose video duration and image count consumption metrics by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33138](https://github.com/BerriAI/litellm/pull/33138) - test(e2e): otel trace completeness on /chat/completions by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33132](https://github.com/BerriAI/litellm/pull/33132) - fix(sso): paginate through all pages when fetching service principal group assignments by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33149](https://github.com/BerriAI/litellm/pull/33149) - test(e2e): otel trace completeness on /v1/messages by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33133](https://github.com/BerriAI/litellm/pull/33133) - feat(ui): add adaptive routing settings to Auto-Router v2 by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​33146](https://github.com/BerriAI/litellm/pull/33146) - refactor(mcp): extract the dcr\_bridge token flow into bridge\_token\_flow\.py by [@​tin-berri](https://github.com/tin-berri) in [#​33141](https://github.com/BerriAI/litellm/pull/33141) - chore: bump litellm 1.93.0 -> 1.94.0, litellm-enterprise 0.1.49 -> 0.1.50, litellm-proxy-extras 0.4.76 -> 0.4.77 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33229](https://github.com/BerriAI/litellm/pull/33229) - fix(proxy): route master key to team-scoped models by [@​kunal2002](https://github.com/kunal2002) in [#​32926](https://github.com/BerriAI/litellm/pull/32926) - chore(deps): pin httplib2 and setuptools transitive floors by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33233](https://github.com/BerriAI/litellm/pull/33233) - feat(ui): left-anchor the Create Key and Create Team CTAs by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33248](https://github.com/BerriAI/litellm/pull/33248) - fix(anthropic/passthrough): drop incompatible temperature when downgrading adaptive thinking for pre-4.6 models by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33244](https://github.com/BerriAI/litellm/pull/33244) - fix(guardrails): run apply\_guardrail-style model-level pre\_call guardrails at deployment hook by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33136](https://github.com/BerriAI/litellm/pull/33136) - fix(proxy)!: enforce user budget on team keys (read-time + reservation) with UI opt-out by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​32005](https://github.com/BerriAI/litellm/pull/32005) - fix(e2e): bound spend-log snapshots to a /spend/logs/v2 window by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33265](https://github.com/BerriAI/litellm/pull/33265) - test(e2e): cover key rpm/tpm rate limiting, window reset, and pacing headers by [@​mateo-berri](https://github.com/mateo-berri) in [#​32914](https://github.com/BerriAI/litellm/pull/32914) - fix(anthropic): use native output capability by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​33235](https://github.com/BerriAI/litellm/pull/33235) - fix(ci): retry setup-uv installs to survive transient manifest fetch failures by [@​mateo-berri](https://github.com/mateo-berri) in [#​33279](https://github.com/BerriAI/litellm/pull/33279) - fix(proxy): never log raw virtual keys in key insertion debug output by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33268](https://github.com/BerriAI/litellm/pull/33268) - fix(bedrock\_mantle): route xai.grok-4.3 via /openai/v1 frontier path by [@​marty-sullivan](https://github.com/marty-sullivan) in [#​33027](https://github.com/BerriAI/litellm/pull/33027) - feat(pricing): add gemini-omni-flash-preview with video output token pricing by [@​mateo-berri](https://github.com/mateo-berri) in [#​33274](https://github.com/BerriAI/litellm/pull/33274) - fix(auth): scope the JWT enterprise gate to actual JWTs by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33296](https://github.com/BerriAI/litellm/pull/33296) - fix(s3): sanitize slashes in response-id-derived object key file name by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33271](https://github.com/BerriAI/litellm/pull/33271) - refactor(ui): migrate guardrails table onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33303](https://github.com/BerriAI/litellm/pull/33303) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33308](https://github.com/BerriAI/litellm/pull/33308) - feat(guardrails): streaming text transformation in generic\_guardrail\_api by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33110](https://github.com/BerriAI/litellm/pull/33110) - test(e2e): cover model-aware mid-conversation system handling on Bedrock Invoke /v1/messages by [@​mateo-berri](https://github.com/mateo-berri) in [#​32963](https://github.com/BerriAI/litellm/pull/32963) - test(claude\_code): move the Claude Code compatibility matrix under tests/e2e by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​32548](https://github.com/BerriAI/litellm/pull/32548) - chore(ci): sync litellm\_internal\_staging into daily OSS branch by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33337](https://github.com/BerriAI/litellm/pull/33337) - feat(bedrock guardrails): add resource-less InvokeGuardrailChecks (detect-only) mode by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33299](https://github.com/BerriAI/litellm/pull/33299) - Revert "chore(ci): sync litellm\_internal\_staging into daily OSS branch" by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33339](https://github.com/BerriAI/litellm/pull/33339) - fix(websearch): intercept web search on the Responses API by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33129](https://github.com/BerriAI/litellm/pull/33129) - fix(anthropic-adapter): drop empty content\_block\_delta events by [@​mateo-berri](https://github.com/mateo-berri) in [#​33315](https://github.com/BerriAI/litellm/pull/33315) - fix(mcp): persist discovered OAuth endpoints and keep last known good on failed re-discovery by [@​tin-berri](https://github.com/tin-berri) in [#​33286](https://github.com/BerriAI/litellm/pull/33286) - test(e2e): otel trace completeness on streaming chat, messages, and responses (LIT-3787) by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33234](https://github.com/BerriAI/litellm/pull/33234) - feat(router): resolve auto-router routing plugins from proxy YAML config by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​33251](https://github.com/BerriAI/litellm/pull/33251) - test(e2e): failed request error span carries the full untruncated message and status (LIT-4179) by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33304](https://github.com/BerriAI/litellm/pull/33304) - refactor(ui): migrate tags table onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33314](https://github.com/BerriAI/litellm/pull/33314) - fix(cli): surface actionable CLI SSO errors when CLI and proxy versions skew by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33309](https://github.com/BerriAI/litellm/pull/33309) - feat(bedrock\_mantle): add GPT-5.6 sol/terra/luna to model cost map by [@​mateo-berri](https://github.com/mateo-berri) in [#​33412](https://github.com/BerriAI/litellm/pull/33412) - chore(codeowners): exempt generated schema.d.ts from UI ownership by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33411](https://github.com/BerriAI/litellm/pull/33411) - feat(proxy): push-based OTLP billable-request metering for enterprise deployments by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​31592](https://github.com/BerriAI/litellm/pull/31592) - fix(mcp): cap per-user OAuth token cache TTL at the token's own lifetime by [@​tin-berri](https://github.com/tin-berri) in [#​33346](https://github.com/BerriAI/litellm/pull/33346) - feat(ui): move Caching out of Experimental into Developer Tools by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33432](https://github.com/BerriAI/litellm/pull/33432) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33425](https://github.com/BerriAI/litellm/pull/33425) - chore(ci): merge daily internal staging branch by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33335](https://github.com/BerriAI/litellm/pull/33335) - feat(guardrails): add Compresr guardrail for query-aware context compression by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33295](https://github.com/BerriAI/litellm/pull/33295) - fix(logging): preserve callback order in get\_combined\_callback\_list by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33005](https://github.com/BerriAI/litellm/pull/33005) - fix(logging): redact assistant tool call arguments in spend logs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33111](https://github.com/BerriAI/litellm/pull/33111) - fix(anthropic): honor messages request timeout by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33418](https://github.com/BerriAI/litellm/pull/33418) - fix(llm\_guard): apply sanitized prompt returned by moderation API to request by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33331](https://github.com/BerriAI/litellm/pull/33331) - fix(logging): stop pinning large request payloads past request end by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33455](https://github.com/BerriAI/litellm/pull/33455) - feat(guardrails): forward optional metadata on POST /guardrails/apply\_guardrail by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33067](https://github.com/BerriAI/litellm/pull/33067) - build: raise requires-python cap to <3.15 so Python 3.14 installs current releases by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33438](https://github.com/BerriAI/litellm/pull/33438) - feat(ui): add reusable BetaBadge and use it for Projects sidebar item by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33449](https://github.com/BerriAI/litellm/pull/33449) - fix(mcp): discover missing OAuth scopes and token\_url when authorization\_url is set manually by [@​tin-berri](https://github.com/tin-berri) in [#​33317](https://github.com/BerriAI/litellm/pull/33317) - feat(ui): show exact license expiration date in usage cards by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33478](https://github.com/BerriAI/litellm/pull/33478) - test(claude\_code): rename misleading REPO\_ROOT to SUITE\_ROOT in test\_v0\_layout by [@​mateo-berri](https://github.com/mateo-berri) in [#​33472](https://github.com/BerriAI/litellm/pull/33472) - build(deps): update ddtrace to the 4.x line by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33484](https://github.com/BerriAI/litellm/pull/33484) - fix(complexity\_router): return empty dict from \_classifier\_call\_metadata when metadata is absent by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33452](https://github.com/BerriAI/litellm/pull/33452) - fix(ui/chat): resolve chat routes at render time so navigation works under server\_root\_path by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33446](https://github.com/BerriAI/litellm/pull/33446) - fix(key management): enforce minimum custom key length and mask short keys in key\_name by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33462](https://github.com/BerriAI/litellm/pull/33462) - chore(ui): remove unmounted UsageIndicator and the Hide Usage Indicator flag by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33482](https://github.com/BerriAI/litellm/pull/33482) - test(e2e/claude\_code): add passthrough matrix row for the big-3 clouds and Anthropic API by [@​mateo-berri](https://github.com/mateo-berri) in [#​33473](https://github.com/BerriAI/litellm/pull/33473) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33491](https://github.com/BerriAI/litellm/pull/33491) - fix(ui): stop sending the complexity-router pseudo-model to /health/test\_connection by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​33498](https://github.com/BerriAI/litellm/pull/33498) - feat(cli): add lite up/down to ambiently route Claude Code through the proxy by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​33231](https://github.com/BerriAI/litellm/pull/33231) - feat(complexity\_router): enable session\_affinity by default by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33500](https://github.com/BerriAI/litellm/pull/33500) - fix(anthropic): stop 500 on combined thinking+signature streaming chunk by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33505](https://github.com/BerriAI/litellm/pull/33505) - feat(autoroute): prompt for semantic keywords per tier in configure wizard by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33508](https://github.com/BerriAI/litellm/pull/33508) - fix(cli/anthropic): unblock lite autoroute proxy deps, adaptive thinking, and thinking+signature streaming by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33507](https://github.com/BerriAI/litellm/pull/33507) - test(ocr): use mistral-document-ai-2512 in azure\_ai OCR tests by [@​mateo-berri](https://github.com/mateo-berri) in [#​33489](https://github.com/BerriAI/litellm/pull/33489) - fix(guardrails): show YAML-defined guardrails in the Guardrail Monitor by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​32853](https://github.com/BerriAI/litellm/pull/32853) - refactor(ui): migrate policies, deleted keys, deleted teams, budgets, and search tools tables onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33357](https://github.com/BerriAI/litellm/pull/33357) - refactor(ui): migrate vector stores, prompts, and skills tables onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33343](https://github.com/BerriAI/litellm/pull/33343) - test(e2e): datadog log delivery for successful chat, messages, and responses (LIT-4447) by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33415](https://github.com/BerriAI/litellm/pull/33415) - fix(cli): make CLI output ASCII-only so it doesn't crash legacy Windows consoles by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33465](https://github.com/BerriAI/litellm/pull/33465) - fix: remove dead user-cache lookup with None key in spend-update path by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33555](https://github.com/BerriAI/litellm/pull/33555) - feat(helm): add per-component PodDisruptionBudget and topologySpreadConstraints to componentized chart by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33430](https://github.com/BerriAI/litellm/pull/33430) - fix(e2e/claude\_code): unblock stage collection, align proxy env names, register compat models by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33433](https://github.com/BerriAI/litellm/pull/33433) - fix(mcp): index authed request-time tools missing from the semantic filter startup index by [@​tin-berri](https://github.com/tin-berri) in [#​33318](https://github.com/BerriAI/litellm/pull/33318) - fix(streaming): use provider-reported usage cost for OpenRouter streams by [@​mateo-berri](https://github.com/mateo-berri) in [#​32255](https://github.com/BerriAI/litellm/pull/32255) - feat(mcp): issuer-anchored OAuth discovery (RFC 8414 §3.3) to close the authorization-server mix-up by [@​tin-berri](https://github.com/tin-berri) in [#​33450](https://github.com/BerriAI/litellm/pull/33450) - feat(logging): add user and team level spend and budget to StandardLoggingPayload metadata by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33459](https://github.com/BerriAI/litellm/pull/33459) - fix(router): cast model\_info cost values to float in \_set\_model\_group\_info by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33556](https://github.com/BerriAI/litellm/pull/33556) - chore(e2e): establish litellm\_e2e\_staging integration line by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33502](https://github.com/BerriAI/litellm/pull/33502) - fix(ui): navigate to /ui/login/ with trailing slash via hard navigation by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33561](https://github.com/BerriAI/litellm/pull/33561) - feat(logging): add structured budget fields to budget rejection failure logs by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33460](https://github.com/BerriAI/litellm/pull/33460) - fix(streaming): surface upstream connection resets instead of empty 200 streams by [@​mateo-berri](https://github.com/mateo-berri) in [#​33222](https://github.com/BerriAI/litellm/pull/33222) - fix(proxy\_cli): reap orphaned prisma query-engine processes when a worker dies by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33424](https://github.com/BerriAI/litellm/pull/33424) - test(reasoning\_effort\_grid): enable azure fable-5 and opus-4-8 grid cells by [@​mateo-berri](https://github.com/mateo-berri) in [#​33485](https://github.com/BerriAI/litellm/pull/33485) - build(deps): bump uvicorn lock to 0.51.0 so worker health-check and jitter flags take effect by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33574](https://github.com/BerriAI/litellm/pull/33574) - fix(proxy): coerce default\_internal\_user\_params.max\_budget to float on config load by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​32434](https://github.com/BerriAI/litellm/pull/32434) - fix(router): honor per-request routing\_strategy from key/team router\_settings by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33429](https://github.com/BerriAI/litellm/pull/33429) - fix(redis): honor ssl value instead of key presence when building async connection pool by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​32590](https://github.com/BerriAI/litellm/pull/32590) - fix(langfuse\_otel): build per-request OTLP exporter from key and team dynamic Langfuse credentials by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​32437](https://github.com/BerriAI/litellm/pull/32437) - ci: run zizmor and proxy-db unit tests on PRs targeting litellm\_ branches by [@​mateo-berri](https://github.com/mateo-berri) in [#​33568](https://github.com/BerriAI/litellm/pull/33568) - fix(router): apply team/key enable\_tag\_filtering to tag routing by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33436](https://github.com/BerriAI/litellm/pull/33436) - feat(proxy): add disable\_auto\_add\_proxy\_admin\_to\_teams flag by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33563](https://github.com/BerriAI/litellm/pull/33563) - fix(proxy): stop stale auth cache re-publish to Redis so key updates propagate across replicas by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33565](https://github.com/BerriAI/litellm/pull/33565) - feat(e2e): emit structured E2E\_RESULT lines for package status history by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33578](https://github.com/BerriAI/litellm/pull/33578) - chore: bump litellm-enterprise 0.1.50 -> 0.1.51, litellm-proxy-extras 0.4.77 -> 0.4.78 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33571](https://github.com/BerriAI/litellm/pull/33571) - fix(docker): restore litellm-proxy-extras source dir in runtime images by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33592](https://github.com/BerriAI/litellm/pull/33592) - feat(ui): require embedding model for semantic auto router by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33313](https://github.com/BerriAI/litellm/pull/33313) - feat(scim): ingest and round-trip SCIM entitlements and roles user attributes by [@​tin-berri](https://github.com/tin-berri) in [#​33587](https://github.com/BerriAI/litellm/pull/33587) - refactor(ui): migrate 5 simple tables onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33548](https://github.com/BerriAI/litellm/pull/33548) - fix(model\_armor): restore reference attachments via skip\_unscannable\_attachments and remove the attachment count cap by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33554](https://github.com/BerriAI/litellm/pull/33554) - fix(mcp): keep the MCP reference intact when the semantic filter narrows tools by [@​tin-berri](https://github.com/tin-berri) in [#​33584](https://github.com/BerriAI/litellm/pull/33584) - fix(sso): stop enforcing UI session budget on CLI login tokens by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33312](https://github.com/BerriAI/litellm/pull/33312) - test: e2e staging leftovers by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33613](https://github.com/BerriAI/litellm/pull/33613) - test(e2e/claude\_code): add GPT-5.6 Sol/Terra/Luna columns for OpenAI, Azure OpenAI, and Bedrock Mantle by [@​mateo-berri](https://github.com/mateo-berri) in [#​33474](https://github.com/BerriAI/litellm/pull/33474) - test(e2e): otel streaming spans record a real ttft below span duration by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33588](https://github.com/BerriAI/litellm/pull/33588) - fix(mcp): make the preemptive-401 OAuth challenge decision mode-aware by [@​tin-berri](https://github.com/tin-berri) in [#​33586](https://github.com/BerriAI/litellm/pull/33586) - fix(ui): show all teams in policy attachment form for admins by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33628](https://github.com/BerriAI/litellm/pull/33628) - refactor(ui): migrate AI Hub, public hub, and MCP Toolsets tables onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33629](https://github.com/BerriAI/litellm/pull/33629) - test(e2e): datadog log delivery for streamed routes, read back from the real datadog api by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33566](https://github.com/BerriAI/litellm/pull/33566) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33640](https://github.com/BerriAI/litellm/pull/33640) - fix(vertex\_ai): surface Gemini grounding toolUsePromptTokenCount in Usage by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33533](https://github.com/BerriAI/litellm/pull/33533) - test(e2e): harness fixes for stage job green (skips + router/UI/budget) by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33634](https://github.com/BerriAI/litellm/pull/33634) - fix(router): resolve prompt cache minimum per model instead of a flat 1024 by [@​tin-berri](https://github.com/tin-berri) in [#​33637](https://github.com/BerriAI/litellm/pull/33637) - fix(logging): classify async anthropic\_messages and generate\_content as async by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33589](https://github.com/BerriAI/litellm/pull/33589) - fix(ui): remove Chat item from dashboard leftnav by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33647](https://github.com/BerriAI/litellm/pull/33647) - fix(router): tag-aware pre-routing strategy selection for shared model\_name by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33691](https://github.com/BerriAI/litellm/pull/33691) - fix(proxy): enforce max\_parallel\_requests as a per-slot concurrency gauge by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​32441](https://github.com/BerriAI/litellm/pull/32441) - fix(proxy): stop treating upstream model body field as a LiteLLM model on auth-enforced pass-through routes by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33710](https://github.com/BerriAI/litellm/pull/33710) - fix(mcp): expand toolset grants in shared permission primitives so tools/call honors them by [@​tin-berri](https://github.com/tin-berri) in [#​33612](https://github.com/BerriAI/litellm/pull/33612) - feat(complexity-router): user-triggered escalation keywords by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33656](https://github.com/BerriAI/litellm/pull/33656) - fix(fireworks\_ai): bill prompt-cache hits at cache\_read rate by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33714](https://github.com/BerriAI/litellm/pull/33714) - fix(pricing): mark realtime-only gpt-realtime models as mode realtime by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33728](https://github.com/BerriAI/litellm/pull/33728) - fix(rag): track LLM completion usage and spend for /v1/rag/query by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​32438](https://github.com/BerriAI/litellm/pull/32438) - feat(anthropic): add enable\_anthropic\_prompt\_caching for automatic cache\_control injection by [@​tin-berri](https://github.com/tin-berri) in [#​33573](https://github.com/BerriAI/litellm/pull/33573) - fix(anthropic): self-heal on missing thinking-signature errors from Bedrock/Vertex by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33719](https://github.com/BerriAI/litellm/pull/33719) - fix(proxy): resolve router\_settings.plugins dotted paths and load plugins from installed packages by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33644](https://github.com/BerriAI/litellm/pull/33644) - test(e2e): budget refusals are 429 for bare keys and team caps block every team key by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33632](https://github.com/BerriAI/litellm/pull/33632) - feat(router): add router plugin reference catalog by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33746](https://github.com/BerriAI/litellm/pull/33746) - test(e2e): assert an org budget block is a 429 naming the organization by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33638](https://github.com/BerriAI/litellm/pull/33638) - fix(proxy): bill partial streamed spend when the client disconnects mid-stream by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33736](https://github.com/BerriAI/litellm/pull/33736) - test(e2e): delete unreferenced Grafana panel docs by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33743](https://github.com/BerriAI/litellm/pull/33743) - docs(tests/e2e): align skip-vs-fail docs with the hard-fail contract and scope the no-unit-tests rule by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33755](https://github.com/BerriAI/litellm/pull/33755) - refactor(e2e): replace bespoke result reporter with standard JUnit report by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33758](https://github.com/BerriAI/litellm/pull/33758) - test(e2e): user budget across keys and team member budget isolation by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33745](https://github.com/BerriAI/litellm/pull/33745) - refactor(e2e): remove bob\_the\_builder; drive remediation from a Grafana alert (provisioned outside the repo) by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33749](https://github.com/BerriAI/litellm/pull/33749) - feat(mcp): per-server outcomes for aggregate tools/list and truthful single-server REST statuses by [@​tin-berri](https://github.com/tin-berri) in [#​33153](https://github.com/BerriAI/litellm/pull/33153) - test(e2e): mcp suite for key-without-access denial by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33752](https://github.com/BerriAI/litellm/pull/33752) - chore(ci): merge oss branch by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33784](https://github.com/BerriAI/litellm/pull/33784) - chore(ci): merge oss branch - July 17th by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33793](https://github.com/BerriAI/litellm/pull/33793) - fix(ui): migrate tag deletion to shared DeleteResourceModal by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33795](https://github.com/BerriAI/litellm/pull/33795) - build(rust): raise pyo3 to 0.29 so the native bridge compiles on Python 3.14 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33798](https://github.com/BerriAI/litellm/pull/33798) - chore(guardrails): remove docstring from singulr module for consistency by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33800](https://github.com/BerriAI/litellm/pull/33800) - fix(ui): stop credential edit from persisting the masked api key by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33797](https://github.com/BerriAI/litellm/pull/33797) - build(deps): allow redisvl, pypdf, and openapi-core on Python 3.14 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33801](https://github.com/BerriAI/litellm/pull/33801) - test(proxy): make streaming-cancel mocks awaitable for the disconnect slot release by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33802](https://github.com/BerriAI/litellm/pull/33802) - test(e2e): a member's team budget cuts off only that member's key by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33718](https://github.com/BerriAI/litellm/pull/33718) - test(e2e): a user's max\_budget follows the person across personal and team keys by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33762](https://github.com/BerriAI/litellm/pull/33762) - test(e2e): skip flaky OpenAI GPT cells; raise multi-window max\_tokens by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33799](https://github.com/BerriAI/litellm/pull/33799) - chore: remove accidentally committed dist tarball and ignore dist/ by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33805](https://github.com/BerriAI/litellm/pull/33805) - fix(passthrough): stop classifying plain 'predict'/'search' paths as Vertex by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33658](https://github.com/BerriAI/litellm/pull/33658) - build(deps): bump mcp lock to 1.28.1 to clear image-scan findings by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33803](https://github.com/BerriAI/litellm/pull/33803) - test(pricing): pin the realtime mode assertion to the bundled cost map by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33806](https://github.com/BerriAI/litellm/pull/33806) - fix(proxy): derive session id from Anthropic metadata.user\_id for session affinity by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33723](https://github.com/BerriAI/litellm/pull/33723) - test(e2e): budget reset diagonal for team, org, user, and [#​32005](https://github.com/BerriAI/litellm/issues/32005) team-member keys by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33771](https://github.com/BerriAI/litellm/pull/33771) - fix(proxy): source /v1/models token limits from the cost map instead of Router.get\_model\_group\_info by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33721](https://github.com/BerriAI/litellm/pull/33721) - fix(fireworks\_ai): correct glm-5p2 prompt-cache read price to $0.14/1M by [@​tin-berri](https://github.com/tin-berri) in [#​33796](https://github.com/BerriAI/litellm/pull/33796) - feat(proxy): add x-litellm-model-name response header with deployment model string by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33698](https://github.com/BerriAI/litellm/pull/33698) - feat: add Straiker guardrail integration by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33781](https://github.com/BerriAI/litellm/pull/33781) - fix(vertex\_ai): exclude Gemini Google Search grounding tokens from input token billing by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33742](https://github.com/BerriAI/litellm/pull/33742) - feat(fireworks\_ai): map litellm session id to x-session-affinity header for prompt caching by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33717](https://github.com/BerriAI/litellm/pull/33717) - feat(ui): configure Anthropic automatic prompt caching from the Admin UI by [@​tin-berri](https://github.com/tin-berri) in [#​33581](https://github.com/BerriAI/litellm/pull/33581) - fix(router): enforce context-window pre-call checks for Responses API input by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33706](https://github.com/BerriAI/litellm/pull/33706) - fix(otel): restore proxy-level error.\* attributes on v2 failure spans (LIT-4179) by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33664](https://github.com/BerriAI/litellm/pull/33664) - fix(mcp): persist config.yaml DCR clients in a server-scoped store so refresh survives token expiry by [@​tin-berri](https://github.com/tin-berri) in [#​33768](https://github.com/BerriAI/litellm/pull/33768) - refactor(ui): consolidate Add/Edit credential modals into one CredentialModal by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32572](https://github.com/BerriAI/litellm/pull/32572) - feat(mcp): add ID-JAG (identity assertion authorization grant) support for MCP egress by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​31516](https://github.com/BerriAI/litellm/pull/31516) - refactor(ui): migrate policy attachments table onto shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33827](https://github.com/BerriAI/litellm/pull/33827) - docs(litellm-rust): add provider coding standards by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33833](https://github.com/BerriAI/litellm/pull/33833) - test(e2e): rename Gateway to ProxyClient and expose it as a session-scoped pytest fixture by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33750](https://github.com/BerriAI/litellm/pull/33750) - feat(messages): route Azure Anthropic /messages through Rust behind rust:true by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33616](https://github.com/BerriAI/litellm/pull/33616) - test(e2e): add Locust throughput load test that runs last by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33748](https://github.com/BerriAI/litellm/pull/33748) - fix(proxy): resolve team wildcard credentials for vector store files by [@​shivamrawat1](https://github.com/shivamrawat1) in [#​33649](https://github.com/BerriAI/litellm/pull/33649) - refactor(e2e): fold claude\_code HTTP probes onto shared ProxyClient methods by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33760](https://github.com/BerriAI/litellm/pull/33760) - test(e2e): harden stage flakes for batches, UI, and MCP by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33831](https://github.com/BerriAI/litellm/pull/33831) - fix(e2e): migrate load suite from e2e\_gateway to ProxyClient by [@​mubashir1osmani](https://github.com/mubashir1osmani) in [#​33839](https://github.com/BerriAI/litellm/pull/33839) - chore(e2e): remove tests/e2e/docker-compose.yml by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33837](https://github.com/BerriAI/litellm/pull/33837) - test(e2e): cover /v1/responses openai basic nonstream and stream by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33830](https://github.com/BerriAI/litellm/pull/33830) - test(e2e): cover /v1/responses openai cost\_logged and tool\_use by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33835](https://github.com/BerriAI/litellm/pull/33835) - test(e2e): cover /v1/responses OpenAI vision and Anthropic basic by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33838](https://github.com/BerriAI/litellm/pull/33838) - test(e2e): spendlog cost for streaming /v1/messages via responses bridge by [@​yassin-berriai](https://github.com/yassin-berriai) in [#​33753](https://github.com/BerriAI/litellm/pull/33753) - fix(docker): bake prisma CLI and engines at a fixed path so fresh-DB migrations work for any uid offline by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33853](https://github.com/BerriAI/litellm/pull/33853) - feat(chat-ui): add personal Logs view scoped to the current user by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33829](https://github.com/BerriAI/litellm/pull/33829) - chore: bump litellm-proxy-extras 0.4.78 -> 0.4.79 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33855](https://github.com/BerriAI/litellm/pull/33855) - docs(litellm-rust): require the official Rust Style Guide in agent rules by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33867](https://github.com/BerriAI/litellm/pull/33867) - fix(router): treat malformed configured token limits as absent on /v1/models by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33864](https://github.com/BerriAI/litellm/pull/33864) - chore: rebuild admin UI bundle for the rc release by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33857](https://github.com/BerriAI/litellm/pull/33857) - docs(rust): add provider abstraction standards by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33865](https://github.com/BerriAI/litellm/pull/33865) - chore(ci): promote internal staging to main by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33868](https://github.com/BerriAI/litellm/pull/33868) - chore(release): backport [#​33929](https://github.com/BerriAI/litellm/issues/33929) to rc/1.94.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34033](https://github.com/BerriAI/litellm/pull/34033) - chore(release): backport [#​33810](https://github.com/BerriAI/litellm/issues/33810), [#​33733](https://github.com/BerriAI/litellm/issues/33733) to rc/1.94.0 and bump litellm-proxy-extras to 0.4.79.post1 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34215](https://github.com/BerriAI/litellm/pull/34215) - chore(ui): rebuild Next.js bundle on rc/1.94.0 so the Cost Optimization page ships by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34216](https://github.com/BerriAI/litellm/pull/34216) - chore(release): backport auth, CLI SSO and guardrail fixes to rc/1.94.0 and refresh flagged dependencies by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34640](https://github.com/BerriAI/litellm/pull/34640) - chore(release): backport [#​33899](https://github.com/BerriAI/litellm/issues/33899), [#​33978](https://github.com/BerriAI/litellm/issues/33978), [#​34582](https://github.com/BerriAI/litellm/issues/34582), [#​34675](https://github.com/BerriAI/litellm/issues/34675) to rc/1.94.0 and bump litellm-proxy-extras to 0.4.79.post2 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34855](https://github.com/BerriAI/litellm/pull/34855) - fix(ui): backport cache leakage card layout fix to rc/1.94.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34964](https://github.com/BerriAI/litellm/pull/34964) - fix(ui): add missing cost-optimization page description on rc/1.94.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34967](https://github.com/BerriAI/litellm/pull/34967) - feat(ui): mark Cost Optimization as beta in the left nav ([#​34984](https://github.com/BerriAI/litellm/issues/34984)) by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34987](https://github.com/BerriAI/litellm/pull/34987) - chore: rebuild Admin UI bundle for v1.94.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34982](https://github.com/BerriAI/litellm/pull/34982) - fix(cost-optimization): backport the savings chart axis fix and methodology popovers to rc/1.94.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34994](https://github.com/BerriAI/litellm/pull/34994) - chore: rebuild Admin UI bundle for rc/1.94.0 by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​34995](https://github.com/BerriAI/litellm/pull/34995) **Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.93.0...v1.94.0> ### [`v1.94.0`](https://github.com/BerriAI/litellm/releases/tag/v1.94.0) ##### Verify Docker Image Signature All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0). **Verify using the pinned commit hash (recommended):** A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \ ghcr.io/berriai/litellm:v1.94.0 ``` **Verify using the release tag (convenience):** Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules: ```bash cosign verify \ --key https://raw.githubusercontent.com/BerriAI/litellm/v1.94.0/cosign.pub \ ghcr.io/berriai/litellm:v1.94.0 ``` Expected output: ``` The following checks were performed on each of these signatures: - The cosign claims were validated - The signatures were verified against the specified public key ``` *** ##### What's Changed - feat(ui): working Test Connection for the complexity auto router by [@​akapur99](https://github.com/akapur99) in [#​32950](https://github.com/BerriAI/litellm/pull/32950) - fix(xecguard): use StandardLoggingGuardrailInformation in logging hook by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​32911](https://github.com/BerriAI/litellm/pull/32911) - feat(ui): adopt openapi-react-query ($api) and convert useCustomers by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32949](https://github.com/BerriAI/litellm/pull/32949) - refactor(ui): colocate the mcp-servers view, keeping the shared mcp\_tools surface by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32968](https://github.com/BerriAI/litellm/pull/32968) - refactor(ui): convert endpoint usage charts to shadcn/recharts by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32723](https://github.com/BerriAI/litellm/pull/32723) - fix(proxy-auth): stop unrecognized model namespaces slipping through provider wildcard keys by [@​mateo-berri](https://github.com/mateo-berri) in [#​32979](https://github.com/BerriAI/litellm/pull/32979) - feat(router): random-pick multi-model complexity tiers by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32967](https://github.com/BerriAI/litellm/pull/32967) - fix(xecguard): sanitize scan result before recording it for logging by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​32935](https://github.com/BerriAI/litellm/pull/32935) - fix(auto\_router): filter embedding models in complexity tab dropdowns, require all tiers, inline validation by [@​akapur99](https://github.com/akapur99) in [#​32978](https://github.com/BerriAI/litellm/pull/32978) - fix(anthropic): translate raw adaptive thinking for pre-4.6 models on chat completions and Bedrock Converse by [@​akapur99](https://github.com/akapur99) in [#​32944](https://github.com/BerriAI/litellm/pull/32944) - feat(router): add Router(plugins=\[...]) routing-plugin pipeline by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32972](https://github.com/BerriAI/litellm/pull/32972) - feat(router): soft-floor adaptive mode for complexity router by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32947](https://github.com/BerriAI/litellm/pull/32947) - docs(github): add QA runbook section to the PR template by [@​mateo-berri](https://github.com/mateo-berri) in [#​32965](https://github.com/BerriAI/litellm/pull/32965) - fix(model\_cost): add supports\_reasoning: false to Gemini image generation models by [@​mateo-berri](https://github.com/mateo-berri) in [#​32836](https://github.com/BerriAI/litellm/pull/32836) - build(dev-env): add make bootstrap and unprovisioned-checkout preflight to pre-commit by [@​mateo-berri](https://github.com/mateo-berri) in [#​32981](https://github.com/BerriAI/litellm/pull/32981) - ci(ui): report only error-level knip findings in CI by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​32971](https://github.com/BerriAI/litellm/pull/32971) - feat(batches): track cost for unmanaged Bedrock batches, generalize the flag by [@​Sameerlite](https://github.com/Sameerlite) in [#​32315](https://github.com/BerriAI/litellm/pull/32315) - fix(guardrails): walk custom\_tool\_call\_output items in \_content\_utils by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​32969](https://github.com/BerriAI/litellm/pull/32969) - fix: show and allow editing team model aliases after team creation by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33047](https://github.com/BerriAI/litellm/pull/33047) - chore(deps): bump pillow to 12.3.0 to resolve osv-scan CVEs by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33093](https://github.com/BerriAI/litellm/pull/33093) - feat(mcp): mint gateway-bound envelope at the token endpoint for dcr\_bridge oauth\_delegate by [@​tin-berri](https://github.com/tin-berri) in [#​32828](https://github.com/BerriAI/litellm/pull/32828) - fix(mcp): surface rejected delegate-auth upstream tokens as connect-time 401 by [@​tin-berri](https://github.com/tin-berri) in [#​32741](https://github.com/BerriAI/litellm/pull/32741) - fix(proxy): track unauthenticated pass-through requests in spend logs by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​32410](https://github.com/BerriAI/litellm/pull/32410) - feat(lasso): send source.type for Used By attribution by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33090](https://github.com/BerriAI/litellm/pull/33090) - fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@​thibault-linktree](https://github.com/thibault-linktree) in [#​33025](https://github.com/BerriAI/litellm/pull/33025) - fix(responses): continue MCP gateway tool turns from the final response and surface failures by [@​tin-berri](https://github.com/tin-berri) in [#​33099](https://github.com/BerriAI/litellm/pull/33099) - fix(completion): forward aws credential kwargs into litellm\_params so the responses bridge keeps WIF auth by [@​mateo-berri](https://github.com/mateo-berri) in [#​32956](https://github.com/BerriAI/litellm/pull/32956) - fix(ui): respect litellm\_key\_header\_name in BYOK credential save and workflow runs fetches by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33103](https://github.com/BerriAI/litellm/pull/33103) - refactor(ui): standardize debounce waits behind shared DEBOUNCE\_WAIT\_MS constant by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33040](https://github.com/BerriAI/litellm/pull/33040) - feat(ui): rebuild the Virtual Keys table on the shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​32991](https://github.com/BerriAI/litellm/pull/32991) - fix: redact async complete streaming response for custom callbacks by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33106](https://github.com/BerriAI/litellm/pull/33106) - build(ui): bump [@​tanstack/react-pacer](https://github.com/tanstack/react-pacer) from 0.2.0 to 0.22.1 by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33041](https://github.com/BerriAI/litellm/pull/33041) - fix(ui): address Virtual Keys redesign review nits by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33112](https://github.com/BerriAI/litellm/pull/33112) - fix(openai/responses): clamp max\_output\_tokens below API minimum by [@​devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#​33098](https://github.com/BerriAI/litellm/pull/33098) - fix(prometheus): read v3 rate limiter remaining values for per-key model gauges by [@​yucheng-berri](https://github.com/yucheng-berri) in [#​33119](https://github.com/BerriAI/litellm/pull/33119) - fix(ui): drop w-full from page-content wrappers to remove 32px horizontal overflow by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33118](https://github.com/BerriAI/litellm/pull/33118) - refactor(ui): migrate straightforward value debounces to react-pacer by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33042](https://github.com/BerriAI/litellm/pull/33042) - feat(mcp): interactive SSO sign-in for dcr\_bridge oauth\_delegate DCR clients by [@​tin-berri](https://github.com/tin-berri) in [#​32946](https://github.com/BerriAI/litellm/pull/32946) - test(proxy): add regression tests for management\_endpoints edge cases by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​32976](https://github.com/BerriAI/litellm/pull/32976) - fix(auto-router): correct Responses API tool\_choice shape and propagate alias litellm\_params by [@​krrish-berri-2](https://github.com/krrish-berri-2) in [#​32974](https://github.com/BerriAI/litellm/pull/32974) - fix(ui): render the sidebar scrollbar with shadcn ScrollArea by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33124](https://github.com/BerriAI/litellm/pull/33124) - refactor(ui): migrate callback debounce sites to react-pacer with regression tests by [@​ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#​33043](https://github.com/BerriAI/litellm/pull/33043) - chore: add CODEOWNERS for ui and proxy UI build artifacts by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33131](https://github.com/BerriAI/litellm/pull/33131) - feat(mcp): client-held refresh envelope for the dcr\_bridge oauth\_delegate flow by [@​tin-berri](https://github.com/tin-berri) in [#​32980](https://github.com/BerriAI/litellm/pull/32980) - feat(ui): rebuild the Teams table on the shared DataTable by [@​yuneng-berri](https://github.com/yuneng-berri) in [#​33128](https://github.com/BerriAI/litellm/pull/33128) - fix(mcp): relay upstream OAuth token and DCR rejections instead of a generic 500 by [@​tin-berri](https://github.com/tin-berri) in [#​33113](https://github.com/BerriAI/litellm/pull/33113) - fix(keys): persist key\_type so the UI shows correct key scope instead of "All Proxy…
Relevant issues
Related: LIT-4428 (product gap found while writing the responses test)
Linear ticket
https://linear.app/litellm-ai/issue/LIT-3787 (streaming follow-up of the merged non-streaming scenarios)
https://linear.app/litellm-ai/issue/LIT-4428 (tracks the two assertions the responses test knowingly relaxes)
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
Screenshots / Proof of Fix
This PR carries the full streaming trace-completeness suite. It was reviewed as a 3-PR stack (this PR, #33247 messages, #33262 responses) and the two stacked PRs were merged down into this branch; content is byte-identical to the verified stack tip. Uses the jaeger destination and read-back infrastructure merged in #33132; no new infra. Verification run on 2026-07-14, proxies from source at exact refs (no docker buildx on this machine); outputs verbatim.
Pass on current code, DEFAULT models with no override; the whole file (three non-streaming + three streaming tests), including streamed responses on gpt-5.5:
Evidence-first probes shaped the assertions before they were written. Streamed chat and messages spans carry the correlation tag, streaming fingerprints, and appear exactly once:
The responses probe found a product gap, filed as LIT-4428: a successful streamed /v1/responses call records spend correctly but its trace never receives the cost-write span, and its gen-AI span carries no litellm.request.streaming tag:
The responses test therefore relaxes exactly those two assertions (documented in its docstring with the ticket id); the chat and messages tests keep the strict versions. Tighten when LIT-4428 lands. Historical note: the earliest responses runs happened during an OpenAI org quota outage and used the env-overridable gemini model, which is why some probe outputs above are gemini-labeled; the outage is also how the in-stream error guard below earned its place (a streamed call commits its HTTP 200 before the upstream fails).
In-stream failures cannot masquerade as success (added per Cursor review): StreamingResponse captures the first SSE error event (OpenAI and Anthropic formats) and every streamed test asserts none arrived. Kill-proof, run while the org key was still quota-dead:
Fail-before-fix at 1bd603d (parent of fix(otel): one v2 logger owns the global provider; scope tenant OTLP creds per exporter #30590, the same foil the merged non-streaming stack was proven against):
On that ref a streamed responses call exports nothing tagged to the destination at all, an even earlier break than the orphan. Re-run 2026-07-15 with the restored OpenAI key on the DEFAULT models: all three fail again, this boot with the orphan on every route including gpt-5.5 (which failure mode appears per route depends on how the #30590 startup race lands on a given boot; the test fails on either):
A second candidate foil was evaluated and rejected with cause: 8f25942 (parent of #27757, the duplicate-streaming-spans fix) predates the OTEL v2 architecture entirely, so under this suite's v2 destination config it fails trivially at the readiness check rather than meaningfully at the duplicate assertion; the duplicate-span regression class is instead pinned by the exactly-one-gen-AI-span assertion in all three tests.
make lint-e2e-basedpyright-> 0 errors;coverage_registry.collector --strictpasses (Logging & Guardrails 5/51 -> 6/52 with the new stream row).Type
✅ Test
Changes
Streaming closes the gen-AI span from the stream-consumption path: the span can only end after the last chunk arrives and usage is aggregated, a lifecycle that historically produced duplicate spans (#27757) and is equally exposed to the orphaning #30590 fixed (verified at its parent commit). None of the merged non-streaming tests exercise this path.
coverage_registry/logging.yaml: adds thelogging.otel.stream.exports_metricP0 row (event: stream, surfaces chat_completions/messages/responses, matching how langfuse declares separate stream cells); each test declares its surface viaexercised_on.logging/test_otel_trace_e2e.py: three streamed tests with the shared completeness contract (recorded state via/health/readiness/details, then exactly ONE trace for the call'slitellm.call_id, single root SERVER span, auth/db children, gen-AI span whose CHILD_OF chain reaches the root) plus stream-specific assertions: the response actually streamed, no in-stream error event arrived, exactly ONE gen-AI span exists for the call, and on the chat/messages surfaces the span recordslitellm.request.streaming=true._assert_complete_trace/_settled_namesgain arequire_cost_spanflag so the responses surface can skip the cost-span requirement (LIT-4428) while the siblings keep it strict. A typed_taghelper reads span tags.e2e_http.py(shared harness):StreamingResponsegains astream_errorfield capturing the first in-stream SSE error event, covering both OpenAI (unspaced JSON) and Anthropic (event: errorcontrol line, spaced JSON) formats.logging/logging_client.py:messages_rawandresponses_rawgain astreammode that consumes the SSE body and counts events, mirroringchat_raw;ResponsesRequestBodygains an optionalstreamfield.models.py:AnthropicMessagesBodygains an optionalstreamfield (mirrorsChatBody).No new cluster infra: the same jaeger deployment #33132 requires covers all three tests.
Behavior changes
None; test-only changes under tests/e2e.
QA runbook
tests/e2e/logging/test_otel_trace_e2e.py::TestOtelTraceCompleteness::test_chat_completions_stream_exports_complete_trace - one successful streamed chat call shows up at the OTEL destination as one connected trace tree with a single streaming-marked model-call span
OpenTelemetryV2appears insuccess_callbacks{"models":["claude-haiku-4-5"],"key_alias":"otel-stream-chat-<uniq>"}and save the returned keycurl -NPOST /chat/completions with"stream": true,max_tokens: 16, and a unique phrase; retry the first call on 401 for a few seconds; confirmcontent-type: text/event-stream, thatdata:chunks actually arrive with no in-stream"type":"error"event, and note thex-litellm-call-idresponse headerlitellm), search by taglitellm.call_id=<that id>and wait for the trace (spans flush in batches; the cost write lands last)POST /chat/completions(kind server) and no span referencing a parent missing from the traceauth /chat/completions, apostgres ...span, andbatch_write_to_db _PROXY_track_cost_callbackchat claude-haiku-4-5span (kind client), that its parent chain reaches the root, and that its tags includelitellm.request.streaming: truetests/e2e/logging/test_otel_trace_e2e.py::TestOtelTraceCompleteness::test_messages_stream_exports_complete_trace - one successful streamed Anthropic-native call shows up at the OTEL destination as one connected trace tree with a single streaming-marked model-call span
OpenTelemetryV2appears insuccess_callbacks{"models":["claude-haiku-4-5"],"key_alias":"otel-stream-messages-<uniq>"}and save the returned keycurl -NPOST /v1/messages with"stream": true,max_tokens: 16, and a unique phrase; retry the first call on 401 for a few seconds; confirmcontent-type: text/event-stream, thatdata:events actually arrive with no in-stream error event, and note thex-litellm-call-idresponse headerlitellm.call_id=<that id>and wait for the tracePOST /v1/messages(kind server) and no span referencing a parent missing from the traceauth /v1/messages, apostgres ...span, andbatch_write_to_db _PROXY_track_cost_callbackchat claude-haiku-4-5span (kind client), that its parent chain reaches the root, and that its tags includelitellm.request.streaming: truetests/e2e/logging/test_otel_trace_e2e.py::TestOtelTraceCompleteness::test_responses_stream_exports_complete_trace - one successful streamed Responses API call shows up at the OTEL destination as one connected trace tree with a single model-call span
OpenTelemetryV2appears insuccess_callbacks{"models":["gpt-5.5"],"key_alias":"otel-stream-responses-<uniq>"}and save the returned keycurl -NPOST /v1/responses with"stream": true,max_output_tokens: 64, and a unique phrase; retry the first call on 401 for a few seconds; confirmcontent-type: text/event-stream, that events ending inresponse.completedarrive (an in-stream"type":"error"event means the upstream failed despite the 200), and note thex-litellm-call-idheaderlitellm.call_id=<that id>and wait for the tracePOST /v1/responses(kind server) and no span referencing a parent missing from the traceauth /v1/responsesand apostgres ...span (the cost-write span is knowingly absent on this surface; see LIT-4428)chat gpt-5.5span (kind client) and that its parent chain reaches the rootFinal attestation: "I agree this test is testing what the writer wanted to test."
Note
Low Risk
Test-only changes under tests/e2e; no production runtime behavior is modified.
Overview
Adds streaming OTEL trace completeness coverage alongside the existing non-streaming tests, targeting the stream-consumption span lifecycle (duplicate/orphaned gen-AI spans).
Registers
logging.otel.stream.exports_metricin the e2e coverage registry and adds three Jaeger-backed tests for streamed/chat/completions,/v1/messages, and/v1/responses. Shared helpers gainrequire_cost_span(skipped for streamed responses per LIT-4428) and a_taghelper for span assertions. Chat/messages streams assertlitellm.request.streaming=true, a single gen-AI span, real SSE delivery, and no in-stream errors; the responses stream test documents relaxed cost-span and streaming-tag expectations until LIT-4428.E2E harness:
StreamingResponse.stream_errorcaptures the first SSE error on HTTP 200 streams;messages_raw/responses_rawand request models gain optionalstreamflags for consumed SSE bodies.Reviewed by Cursor Bugbot for commit ad61aae. Bugbot is set up for automated code reviews on this repo. Configure here.