Repository navigation
fix(streaming): keep the served service_tier on streamed chunks and spend rows - #42870
Conversation
…hunks and spend rows Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
bugbot run |
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
…ge chunk Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
|
|
bugbot run |
…rtial spend Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… paths Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…s bill partial spend Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…tream wrapper Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…_service_tier Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> # Conflicts: # tests/test_litellm/litellm_core_utils/test_litellm_logging.py
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ved tiers in the stream billing integration test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…bill it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…wargs dict Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit aaf6c94. Configure here.
Follow-up check on whether this is OpenAI not sending us this stuff, or it's just that we were getting them from the provider's chunks, but just not forwarding it to the client or using it would be good @devin-ai-integration can you do the follow-up and post a GitHub comment on what's happening? |
|
OpenAI does send it: the Responses stream carries |
Resolves the conflict at the end of tests/unit/litellm_core_utils/test_streaming_handler.py, where BerriAI#42870 appended a test after the same context: both blocks kept, main's first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…served at default #42870 added both the rule that a served default or standard tier bills at base pricing and records no service_tier, and streamed tests expecting the row to record 'default'. They have failed on every scheduled litellm-e2e run since. The tests now map the served tier to the pricing basis the bill must record and check input is billed at that basis's rate; the messages case registers custom rates so the rate check has something to compare against
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden #43063 stamps used_client_oauth_token into spend-log metadata, so test_async_gcs_pub_sub_v1 failed on main with an extra metadata key * test(ui): give the auto-router threshold save wait room for the availability debounce #42625 keeps Save disabled while a 300ms-debounced availability check runs. This test waits for Save right after the change, so the whole debounce lands inside waitFor's 1s default and it times out under CI load. It is the recurring UI Unit Tests failure on main since #42625 landed * test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default #42870 added both the rule that a served default or standard tier bills at base pricing and records no service_tier, and streamed tests expecting the row to record 'default'. They have failed on every scheduled litellm-e2e run since. The tests now map the served tier to the pricing basis the bill must record and check input is billed at that basis's rate; the messages case registers custom rates so the rate check has something to compare against * test(e2e-ui): wait for the call-id search before hovering the logs row The row the spec hovers is already on the unfiltered first page, so it was found before the search request returned. The search response then re-rendered the table under the mouse, and the Base UI tooltip never opened. Reproduced with Playwright against a local proxy: hovering right after the fill never shows the tooltip, hovering after the search response shows the call id every time * test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off The case picked the cheapest Together row flagged supports_response_schema. DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model the reasoning_effort=none case already exercises, and Together lists it with structured output support * test(integration): read the agent 365 guardrail status by its own name in spend logs The MCP shard runs under xdist against one database, and a sibling file creates a default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so that filter's 'success' entry could land first in guardrail_information and the test read it instead of the agent 365 verdict * test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check gc.collect() inside the caplog window can collect a pending task an earlier test left on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this test's records. The check still counts every LiteLLM logger, and unretrieved task exceptions on this loop still go through the asserted exception handler * test(e2e-ui): fill the create-tag fields inside the dialog #42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description') match two elements and Playwright's strict mode fails the create step * test(integration): run integration proxies with the CI license Multi-worker proxies start each uvicorn worker in a fresh process, so every worker reads the license from its environment. Forward LITELLM_LICENSE into the proxy and test runner environments * ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5 Every pull request saved its own uv, maturin, Rust and Prisma caches, about 4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries within minutes. Pull request jobs then missed every cache, downloaded all dependencies from PyPI and hit the install step timeouts. Pull requests now restore only, and main keeps the caches warm for them. test-linting and check-ui-api-types run only on pull requests and keep saving codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity keybase account, so every upload failed signature verification. 5.5.5 reads it from codecovsecops; the key ID matches the one signing the current CLI * test(unit): join the session-minting thread before collecting the handler asyncio.to_thread resumes the test as soon as the worker sets its result, while the pool thread can still hold the work item and through it the handler. gc.collect() then cannot finalize the handler and the session stays open. A pool that shuts down before the test continues drops that reference * test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early owned_proxy_process released its reserved port and the proxy bound it only after full startup, so another xdist worker or an outgoing connection could take it first and the proxy exited with 'address already in use'. The launch now retries on a fresh port when that happens and stops every failed attempt. uvicorn closes idle keep-alive connections after 5 seconds and httpx expired them at the same 5 seconds, so a request sent right at that mark could reuse a socket the server was closing and get 'Connection reset by peer'. Gateway clients now drop idle connections after 2 seconds * ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build The release profile builds with fat LTO and one codegen unit, so the final link of litellm-cache-s3 runs silently for minutes. Successful builds take 711 to 749 seconds, right at the default 10 minute no-output limit, and about 30% of recent runs were killed there * test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections The proxy retries the database about every 30 seconds and each retry opens roughly one connection, so a 5-connection outage took 3 to 4 retries to clear and recovery landed between 60 and 90 seconds, straddling the test's 80 second reset window. A fixed 10 second outage still refuses the immediate reconnect and recovers on the next retry * ci: move the unit-test uv cache split into a composite action check_workflow_startup_safety sums every setup step's timeout, so the save and restore variants each counted 5 minutes although only one runs. One composite step keeps the setup ceiling at 35 minutes * test(unit): point tiktoken at the bundled cache for every unit test The rust_bridge tokenizer tests loaded o200k_base before any test in their xdist worker had imported default_encoding, so tiktoken fell back to the temp cache and tried to download under pytest-socket. Move the session fixture from litellm_core_utils/conftest.py to the root unit conftest. * test(integration): answer model discovery probes in the hosted_vllm wire tests The router's periodic upstream model info refresh sends GET /v1/models to hosted_vllm deployments, so a wire server that is live during a refresh sees an extra request. Answer the probe with an empty model list and leave it out of the provider-call assertions, matching the responses bridge tests.
TLDR
Problem this solves:
service_tierOpenAI stamped on the responseservice_tierHow it solves it:
response.createdand stamps it on every translated chunkservice_tieron every chunkOpenAIChatCompletionStreamingHandler) keepsservice_tieron the parsed chunk instead of dropping it, so those deployments can bill*_priorityrates tooflex/balanced/priority/fast/ultrafastwins, a serveddefault/standard(or GeminiON_DEMAND) forces base rates, andauto/scale/unknown/absent echoes fall back to the requested tierIntentional product change: a request for
prioritythat the provider downgrades todefaultis now billed at base rates instead of priority rates, matching the provider's own invoiceUser Flow
Before: a developer streams a chat completion and can not tell from the logs which tier OpenAI actually served, so tier pricing is unverifiable
"stream": true,stream_options.include_usageand noservice_tier"service_tier": "default", but the final usage chunk comes back withoutservice_tier(on gpt-5.x models via the Responses bridge, no chunk has it)metadata.cost_breakdown.service_tierisnullAfter: the same request shows the served tier on every chunk and on the spend row
"stream": true,stream_options.include_usageand noservice_tier"service_tier": "default"metadata.cost_breakdown.service_tieras"default"and the row is billed at that tier's ratesSecond flow, the caller asks for a tier and the provider serves a different one
Before: they send the same POST with
"service_tier": "priority", OpenAI is short on priority capacity and serves it at"service_tier": "default"(visible on every chunk), and the spend row bills priority rates anyway because cost calc read the requested tier firstAfter: the same request bills base rates and the spend row shows
"default", matching the provider invoice. A request forprioritythat comes backpriority,autoorscale, or with no tier at all, still bills priority, so callers on providers that do not echo the tier see no changeIntegration coverage across the related tickets
tests/integration/spend/test_service_tier_stream_billing.pyruns a real proxy, Postgres and a scripted upstream that stamps the served tier on the stream. Each test bills at distinct default and tiered rates, so a wrong tier cannot produce the expected spend. Run on origin/main (b248b1c) with the same test file, then on this branch:Mutation: revert
_get_service_tier_from_chunksto return None and every priority row above bills at default rates; drop theservice_tierkwarg fromDatabricksChatResponseIterator.chunk_parserand only the databricks row goes red; make_resolve_billable_service_tierreturn the requested tier first and the downgrade row bills priority againLinear ticket
Resolves LIT-8514
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Setup: proxy on localhost:4000 with a Postgres spend log DB and a deployment
nano-tierregistered asopenai/gpt-4.1-nano(real OpenAI key). Same request both times:Before (a3d791f)
Streamed chunks carry the served tier
service_tier:Spend row records the served tier
/spend/logsfor that request id{ "request_id": "chatcmpl-ERUcP5GEoz6ZvN0Eh7wxoej1n0HxV", "model": "openai/gpt-4.1-nano", "spend": 1.7e-06, "service_tier": null, "input_cost": 9e-07, "output_cost": 8e-07 }After (e6722a9)
Streamed chunks carry the served tier
Spend row records the served tier
/spend/logsfor that request id{ "request_id": "chatcmpl-ERUarhGzcTOQV7GhLS6wZDLTnkUUV", "model": "openai/gpt-4.1-nano", "spend": 1.7e-06, "service_tier": "default", "input_cost": 9e-07, "output_cost": 8e-07 }Responses bridge (gpt-5-nano, deployment
nano5-tier), same curlBefore (a3d791f):
After (69fe084):
The last two frames on the tip:
/v1/messages disconnect (gpt-4.1-nano, chat-adapter route)
Recipe on both legs: register
openai/gpt-4.1-nanoviaPOST /model/newwithinput_cost_per_token4e-5,output_cost_per_token8e-5 and*_priority6e-5 / 1.6e-4, then abandon a stream and read the key's spend rows:Before (69fe084): the client saw
message_start,content_block_startand text deltas, the proxy loggedRecorded streaming client disconnect with error_code=499, and/spend/logsreturned[]after the flush (writer loggedSpend Logs transactions: 0)After (047b506): same stream, proxy logged
Billing partial streamed spend for 3 chunks after client disconnect, and the row landed:{ "request_id": "msg_b9dda6dd-6415-4413-b199-b15e81de63be", "call_type": "anthropic_messages", "spend": 0.00124, "prompt_tokens": 27, "completion_tokens": 2, "model": "openai/gpt-4.1-nano", "metadata": {"cost_breakdown": {"input_cost": 0.00108, "output_cost": 0.00016, "total_cost": 0.00124, "service_tier": "default"}} }The two new e2e tests (
-k "responses_stream or messages_stream") pass against that proxy: 2 passed in 120s, live OpenAIValidation on this branch:
tests/e2e/quota_management/spend_tracking/test_service_tier_pricing_e2e.pystreamed cases failed on the merge base and pass on the tip (2 passed, 119s, live OpenAI). Unit regression tests intest_streaming_chunk_builder_utils.py,test_streaming_handler.py,test_streaming_helpers.pyand the Responses transformation test file. Mutations reverting the chunk builder tier scan, the reverse scan order, the Responses terminal tier copy, the Responses stream-scoped tier memo (test_every_bridged_chunk_after_response_created_carries_the_served_service_tier), the fast serializer field and the final response stamping each fail at least one of those tests.python -m coverage_registry.collector --strictpasses with the two new registry cells.scripts/type_discipline_gate.py,scripts/ruff_strict_gate.py,scripts/type_check_gate.py --base origin/mainand the test-quality gate all passIntegration matrix (scripted upstream, real proxy, Postgres and Redis cache on)
tests/integration/spend/test_service_tier_stream_billing.py: the upstream stampsservice_tier: "priority"on every chunk and the deployment registers 10x priority rates, so a bill at the wrong tier can not match. 4 passed in 19.67s on the tipBefore the last two commits the completed /v1/messages row billed 0.11 at default rates with a null tier (the OpenAI-compatible chunk parser dropped
service_tier), and the disconnected one wrote no row at all (withlitellm.cacheon,AnthropicMessagesStreamCacheWritersat between the router wrapper and the chat stream and hid.chunks)Endpoint x outcome coverage
test_streamed_call_records_and_bills_the_served_tier)TestStreamingClientDisconnectBillingsuite, chunk builder tier tests)response.completed(unittest_responses_completed_event_bills_the_served_service_tier, e2etest_responses_stream_records_the_served_tier)response.completed)test_messages_stream_records_the_served_tier); the Anthropic wire format has no tier fieldLITELLM_USE_CHAT_COMPLETIONS_URL_FOR_ANTHROPIC_MESSAGES, or any non OpenAI/Azure backend) fixed here: partial spend billed with the tier (unittest_disconnect_bills_partial_spend_for_anthropic_adapter_stream, fails before the fix, integration matrix above, live proof below). The default OpenAI route goes through the Responses adapter and bills nothing on disconnect, tracked with the native Responses gap in LIT-8603Type
🐛 Bug Fix
Caveats (if any)
Low
response.completed, so it still bills nothing on disconnect (LIT-8603)autoand OpenAI can echoscale; both count as no signal, so the requested tier still decides thereservice_tierargument passed straight intocompletion_cost(the Responses WS partitioner) still overrides both request and responsehosted_vllm/deployments for /v1/messages because OpenAI and Azure backends take the Responses adapter by default; the OpenAI chat-adapter route is covered by the unit test and the live proofQA runbook
tests/e2e/quota_management/spend_tracking/test_service_tier_pricing_e2e.py::TestServiceTierPricing::test_streamed_call_records_and_bills_the_served_tier - a streamed call with no tier requested records the tier OpenAI served on its spend row and bills at that tier's rates
openai/gpt-4.1-nanowith distinctinput_cost_per_token,output_cost_per_token,input_cost_per_token_priorityandoutput_cost_per_token_priority(needsOPENAI_API_KEYandSTORE_MODEL_IN_DB=True)"stream": true,stream_options.include_usageand noservice_tier; note theservice_tieron the chunksmetadata.cost_breakdown.service_tierto equal the tier seen on the chunks andspendto equal prompt and completion tokens times that tier's ratestests/e2e/quota_management/spend_tracking/test_service_tier_pricing_e2e.py::TestServiceTierPricing::test_every_streamed_chunk_carries_the_served_tier - every relayed chunk of a streamed OpenAI call carries the served
service_tier"stream": trueandstream_options.include_usagedata:chunk, including the final usage-only chunk, to contain the same non-emptyservice_tiertests/e2e/quota_management/spend_tracking/test_service_tier_pricing_e2e.py::TestServiceTierPricing::test_responses_stream_records_the_served_tier - a streamed /v1/responses call records the tier carried on
response.completedon its spend row"stream": true; noteresponse.service_tieron theresponse.completedeventmetadata.cost_breakdown.service_tierto equal that tiertests/e2e/quota_management/spend_tracking/test_service_tier_pricing_e2e.py::TestServiceTierPricing::test_messages_stream_records_the_served_tier - a streamed /v1/messages call on an OpenAI-backed deployment records the served tier on its spend row
"stream": trueand a fresh keymetadata.cost_breakdown.service_tierto be a non-null tier that has custom rates registeredFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/53a7da17cdb047a196848100235a5421
Open in Devin Desktop: https://app.devin.ai/desktop/session/53a7da17cdb047a196848100235a5421?variant=devin
Requested by: @kerry-berri