Skip to content

fix(streaming): keep the served service_tier on streamed chunks and spend rows - #42870

Merged
kerry-berri merged 26 commits into
mainfrom
litellm_stream_served_service_tier
Sep 29, 2026
Merged

kerry-berri merged 26 commits into
mainfrom
litellm_stream_served_service_tier

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Streamed OpenAI calls lost the service_tier OpenAI stamped on the response
  • Spend rows for streamed calls recorded no served tier
  • Some relayed chunks (all of them on the Responses bridge) dropped service_tier

How it solves it:

  • Stream reassembly keeps the last tier the provider stamped on a chunk
  • The Responses-to-chat bridge remembers the tier from response.created and stamps it on every translated chunk
  • Final and usage-only chunks carry the tier through to the caller and spend logging
  • The proxy fast serializer includes service_tier on every chunk
  • The /v1/messages chat-adapter stream (and the router's fallback wrapper and the response-cache writer around it) exposes the chat chunks it wraps, so a client disconnect mid-stream bills the partial spend and its tier the way /v1/chat/completions already does
  • The generic OpenAI-compatible SSE chunk parser (hosted_vllm, openai_like and every provider inheriting OpenAIChatCompletionStreamingHandler) keeps service_tier on the parsed chunk instead of dropping it, so those deployments can bill *_priority rates too
  • Databricks streams keep the served tier on parsed chunks and the Databricks cost path bills it (LIT-8121)
  • Cost calc now bills the tier the provider served over the tier the caller requested: a served flex/balanced/priority/fast/ultrafast wins, a served default/standard (or Gemini ON_DEMAND) forces base rates, and auto/scale/unknown/absent echoes fall back to the requested tier

Intentional product change: a request for priority that the provider downgrades to default is now billed at base rates instead of priority rates, matching the provider's own invoice

User Flow

Before: a developer streams a chat completion and can not tell from the logs which tier OpenAI actually served, so tier pricing is unverifiable

  1. They send POST http://localhost:4000/v1/chat/completions with "stream": true, stream_options.include_usage and no service_tier
  2. OpenAI serves the call at "service_tier": "default", but the final usage chunk comes back without service_tier (on gpt-5.x models via the Responses bridge, no chunk has it)
  3. They GET http://localhost:4000/spend/logs?request_id=... and the row's metadata.cost_breakdown.service_tier is null

After: the same request shows the served tier on every chunk and on the spend row

  1. They send the same POST http://localhost:4000/v1/chat/completions with "stream": true, stream_options.include_usage and no service_tier
  2. Every chunk, including the usage chunk, comes back with "service_tier": "default"
  3. GET http://localhost:4000/spend/logs?request_id=... shows metadata.cost_breakdown.service_tier as "default" and the row is billed at that tier's rates

Second flow, the caller asks for a tier and the provider serves a different one

Before: they send the same POST with "service_tier": "priority", OpenAI is short on priority capacity and serves it at "service_tier": "default" (visible on every chunk), and the spend row bills priority rates anyway because cost calc read the requested tier first

After: the same request bills base rates and the spend row shows "default", matching the provider invoice. A request for priority that comes back priority, auto or scale, or with no tier at all, still bills priority, so callers on providers that do not echo the tier see no change

Integration coverage across the related tickets

tests/integration/spend/test_service_tier_stream_billing.py runs a real proxy, Postgres and a scripted upstream that stamps the served tier on the stream. Each test bills at distinct default and tiered rates, so a wrong tier cannot produce the expected spend. Run on origin/main (b248b1c) with the same test file, then on this branch:

test ticket main this PR
completed chat stream LIT-8514 fail (chunks carry no tier) pass
disconnected chat stream LIT-8514 fail (billed 0.016 at default, expected 0.16) pass
completed /v1/messages stream LIT-8514 fail (billed 0.11, expected 1.1) pass
disconnected /v1/messages stream LIT-8514 fail (no spend row written) pass
azure chat stream LIT-2850 fail (chunks carry no tier) pass
databricks chat stream LIT-8121 fail (chunks carry no tier) pass
responses bridge stream LIT-8514 fail (first chunks carry no tier) pass
gemini flex via trafficType LIT-6287, LIT-6292 pass pass
requested priority, served default served-over-requested fail (billed priority) pass
requested priority, served auto served-over-requested pass pass

Mutation: revert _get_service_tier_from_chunks to return None and every priority row above bills at default rates; drop the service_tier kwarg from DatabricksChatResponseIterator.chunk_parser and only the databricks row goes red; make _resolve_billable_service_tier return the requested tier first and the downgrade row bills priority again

Linear ticket

Resolves LIT-8514

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Setup: proxy on localhost:4000 with a Postgres spend log DB and a deployment nano-tier registered as openai/gpt-4.1-nano (real OpenAI key). Same request both times:

curl -sN http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H "Content-Type: application/json" \
  -d '{"model":"nano-tier","messages":[{"role":"user","content":"Say hi"}],"stream":true,"stream_options":{"include_usage":true}}' > stream.txt
echo "chunks: $(grep -c '^data: {' stream.txt)  with service_tier: $(grep -c '"service_tier":"default"' stream.txt)"
RID=$(grep -o 'chatcmpl-[A-Za-z0-9]*' stream.txt | head -1); sleep 15
curl -s "http://localhost:4000/spend/logs?request_id=$RID" -H "Authorization: Bearer $LITELLM_MASTER_KEY" | python3 -c "
import json,sys; r=json.load(sys.stdin)[0]; cb=r['metadata']['cost_breakdown']
print(json.dumps({'request_id':r['request_id'],'model':r['model'],'spend':r['spend'],'service_tier':cb['service_tier'],'input_cost':cb['input_cost'],'output_cost':cb['output_cost']},indent=1))"

Before (a3d791f)

Streamed chunks carry the served tier

  1. Ran the curl above
  2. Observed:
    chunks: 5  with service_tier: 4
    
    The final usage chunk has no service_tier:
    data: {"id":"chatcmpl-ERUcP5GEoz6ZvN0Eh7wxoej1n0HxV","created":1790221261,"model":"nano-tier","object":"chat.completion.chunk","system_fingerprint":"fp_a8d11c9b56","choices":[{"index":0,"delta":{}}],"usage":{"completion_tokens":2,"prompt_tokens":9,"total_tokens":11,...,"cost":1.6999999999999998e-6}}
    

Spend row records the served tier

  1. Queried /spend/logs for that request id
  2. Observed:
    {
     "request_id": "chatcmpl-ERUcP5GEoz6ZvN0Eh7wxoej1n0HxV",
     "model": "openai/gpt-4.1-nano",
     "spend": 1.7e-06,
     "service_tier": null,
     "input_cost": 9e-07,
     "output_cost": 8e-07
    }

After (e6722a9)

Streamed chunks carry the served tier

  1. Ran the curl above
  2. Observed:
    chunks: 5  with service_tier: 5
    
    The final usage chunk now carries it:
    data: {"id":"chatcmpl-ERUarhGzcTOQV7GhLS6wZDLTnkUUV","created":1790221165,"model":"nano-tier","object":"chat.completion.chunk","system_fingerprint":"fp_a8d11c9b56","choices":[{"index":0,"delta":{}}],"usage":{"completion_tokens":2,"prompt_tokens":9,"total_tokens":11,...,"cost":1.6999999999999998e-6},"service_tier":"default"}
    

Spend row records the served tier

  1. Queried /spend/logs for that request id
  2. Observed:
    {
     "request_id": "chatcmpl-ERUarhGzcTOQV7GhLS6wZDLTnkUUV",
     "model": "openai/gpt-4.1-nano",
     "spend": 1.7e-06,
     "service_tier": "default",
     "input_cost": 9e-07,
     "output_cost": 8e-07
    }

Responses bridge (gpt-5-nano, deployment nano5-tier), same curl

Before (a3d791f):

chunks: 13  with service_tier: 0
{"request_id": "chatcmpl-ERV6rbGcRRj72um9gYk7Her31VoR2", "model": "openai/gpt-5-nano", "spend": 8.48e-05, "service_tier": null, ...}

After (69fe084):

chunks: 13  with service_tier: 13
{"request_id": "chatcmpl-ERV4d5Mj7NGGMbp2nQYWGxQSJygFH", "model": "openai/gpt-5-nano", "spend": 5.96e-05, "service_tier": "default", ...}

The last two frames on the tip:

data: {"id":"chatcmpl-ERV4d5Mj7NGGMbp2nQYWGxQSJygFH","object":"chat.completion.chunk","created":1790223012,"model":"nano5-tier","service_tier":"default","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-ERV4d5Mj7NGGMbp2nQYWGxQSJygFH",...,"choices":[{"index":0,"delta":{}}],"usage":{"completion_tokens":148,"prompt_tokens":8,"total_tokens":156,...},"service_tier":"default"}

/v1/messages disconnect (gpt-4.1-nano, chat-adapter route)

Recipe on both legs: register openai/gpt-4.1-nano via POST /model/new with input_cost_per_token 4e-5, output_cost_per_token 8e-5 and *_priority 6e-5 / 1.6e-4, then abandon a stream and read the key's spend rows:

timeout 3s curl -sN -X POST http://localhost:4000/v1/messages \
  -H "Authorization: Bearer $VK" -H "Content-Type: application/json" \
  -d '{"model":"tier-discount-chat","max_tokens":400,"stream":true,
       "messages":[{"role":"user","content":"Tell me a very long detailed story about a dragon and a lighthouse keeper, at least 500 words"}]}'
sleep 90
curl -s "http://localhost:4000/spend/logs?api_key=$VK" -H "Authorization: Bearer $MK"

Before (69fe084): the client saw message_start, content_block_start and text deltas, the proxy logged Recorded streaming client disconnect with error_code=499, and /spend/logs returned [] after the flush (writer logged Spend Logs transactions: 0)

After (047b506): same stream, proxy logged Billing partial streamed spend for 3 chunks after client disconnect, and the row landed:

{
    "request_id": "msg_b9dda6dd-6415-4413-b199-b15e81de63be",
    "call_type": "anthropic_messages",
    "spend": 0.00124,
    "prompt_tokens": 27,
    "completion_tokens": 2,
    "model": "openai/gpt-4.1-nano",
    "metadata": {"cost_breakdown": {"input_cost": 0.00108, "output_cost": 0.00016, "total_cost": 0.00124, "service_tier": "default"}}
}

The two new e2e tests (-k "responses_stream or messages_stream") pass against that proxy: 2 passed in 120s, live OpenAI

Validation on this branch: tests/e2e/quota_management/spend_tracking/test_service_tier_pricing_e2e.py streamed cases failed on the merge base and pass on the tip (2 passed, 119s, live OpenAI). Unit regression tests in test_streaming_chunk_builder_utils.py, test_streaming_handler.py, test_streaming_helpers.py and the Responses transformation test file. Mutations reverting the chunk builder tier scan, the reverse scan order, the Responses terminal tier copy, the Responses stream-scoped tier memo (test_every_bridged_chunk_after_response_created_carries_the_served_service_tier), the fast serializer field and the final response stamping each fail at least one of those tests. python -m coverage_registry.collector --strict passes with the two new registry cells. scripts/type_discipline_gate.py, scripts/ruff_strict_gate.py, scripts/type_check_gate.py --base origin/main and the test-quality gate all pass

Integration matrix (scripted upstream, real proxy, Postgres and Redis cache on)

tests/integration/spend/test_service_tier_stream_billing.py: the upstream stamps service_tier: "priority" on every chunk and the deployment registers 10x priority rates, so a bill at the wrong tier can not match. 4 passed in 19.67s on the tip

test spend tokens cost_breakdown.service_tier
completed /v1/chat/completions 1.1 30/40 priority
disconnected /v1/chat/completions 0.16 14/1 priority
completed /v1/messages (chat adapter) 1.1 30/40 priority
disconnected /v1/messages (chat adapter) 0.18 16/1 priority

Before the last two commits the completed /v1/messages row billed 0.11 at default rates with a null tier (the OpenAI-compatible chunk parser dropped service_tier), and the disconnected one wrote no row at all (with litellm.cache on, AnthropicMessagesStreamCacheWriter sat between the router wrapper and the chat stream and hid .chunks)

Endpoint x outcome coverage

complete stream client disconnect mid-stream
/v1/chat/completions billed at served tier (e2e test_streamed_call_records_and_bills_the_served_tier) partial spend billed with the tier (existing TestStreamingClientDisconnectBilling suite, chunk builder tier tests)
/v1/chat/completions via the Responses bridge billed at served tier (bridge unit test, live proof above) same partial-spend path as chat, chunks stamped by the bridge
/v1/responses (native) already worked: tier is read off response.completed (unit test_responses_completed_event_bills_the_served_service_tier, e2e test_responses_stream_records_the_served_tier) not billed at all, tracked separately in LIT-8603 (the Responses iterator keeps no events and usage only arrives on response.completed)
/v1/messages (OpenAI backend) billed at served tier by the inner stream (e2e test_messages_stream_records_the_served_tier); the Anthropic wire format has no tier field chat-adapter path (LITELLM_USE_CHAT_COMPLETIONS_URL_FOR_ANTHROPIC_MESSAGES, or any non OpenAI/Azure backend) fixed here: partial spend billed with the tier (unit test_disconnect_bills_partial_spend_for_anthropic_adapter_stream, fails before the fix, integration matrix above, live proof below). The default OpenAI route goes through the Responses adapter and bills nothing on disconnect, tracked with the native Responses gap in LIT-8603

Type

🐛 Bug Fix

Caveats (if any)

Low

  • Partial billing on disconnect for /v1/messages only reaches the chat-adapter route. The default OpenAI route uses the Responses adapter, whose iterator keeps no events and gets usage only on response.completed, so it still bills nothing on disconnect (LIT-8603)
  • Groq rewrites a missing echo to auto and OpenAI can echo scale; both count as no signal, so the requested tier still decides there
  • An explicit service_tier argument passed straight into completion_cost (the Responses WS partitioner) still overrides both request and response
  • The integration matrix uses hosted_vllm/ deployments for /v1/messages because OpenAI and Azure backends take the Responses adapter by default; the OpenAI chat-adapter route is covered by the unit test and the live proof
  • The usage-only unit fixture stamps the tier on the source chunk, so the usage-chunk copy path is covered by the live e2e test rather than the unit test

QA runbook

  • tests/e2e/quota_management/spend_tracking/test_service_tier_pricing_e2e.py::TestServiceTierPricing::test_streamed_call_records_and_bills_the_served_tier - a streamed call with no tier requested records the tier OpenAI served on its spend row and bills at that tier's rates

    • POST /model/new with the master key registering openai/gpt-4.1-nano with distinct input_cost_per_token, output_cost_per_token, input_cost_per_token_priority and output_cost_per_token_priority (needs OPENAI_API_KEY and STORE_MODEL_IN_DB=True)
    • POST /v1/chat/completions to that model with "stream": true, stream_options.include_usage and no service_tier; note the service_tier on the chunks
    • GET /spend/logs?request_id= and expect metadata.cost_breakdown.service_tier to equal the tier seen on the chunks and spend to equal prompt and completion tokens times that tier's rates
    • Sanity check: this test makes sense to add and is not hand-wavey (e.g., assert actual expected spend instead of just spend > 0) or potentially flaky
  • tests/e2e/quota_management/spend_tracking/test_service_tier_pricing_e2e.py::TestServiceTierPricing::test_every_streamed_chunk_carries_the_served_tier - every relayed chunk of a streamed OpenAI call carries the served service_tier

    • POST /v1/chat/completions to the same model with "stream": true and stream_options.include_usage
    • Expect every data: chunk, including the final usage-only chunk, to contain the same non-empty service_tier
    • Sanity check: this test makes sense to add and is not hand-wavey (e.g., assert actual expected spend instead of just spend > 0) or potentially flaky
  • tests/e2e/quota_management/spend_tracking/test_service_tier_pricing_e2e.py::TestServiceTierPricing::test_responses_stream_records_the_served_tier - a streamed /v1/responses call records the tier carried on response.completed on its spend row

    • POST /v1/responses to the same model with "stream": true; note response.service_tier on the response.completed event
    • GET /spend/logs?request_id= and expect metadata.cost_breakdown.service_tier to equal that tier
    • Sanity check: this test makes sense to add and is not hand-wavey (e.g., assert actual expected spend instead of just spend > 0) or potentially flaky
  • tests/e2e/quota_management/spend_tracking/test_service_tier_pricing_e2e.py::TestServiceTierPricing::test_messages_stream_records_the_served_tier - a streamed /v1/messages call on an OpenAI-backed deployment records the served tier on its spend row

    • POST /v1/messages to the same model with "stream": true and a fresh key
    • GET /spend/logs?api_key= and expect the row's metadata.cost_breakdown.service_tier to be a non-null tier that has custom rates registered
    • Sanity check: this test makes sense to add and is not hand-wavey (e.g., assert actual expected spend instead of just spend > 0) or potentially flaky

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/53a7da17cdb047a196848100235a5421
Open in Devin Desktop: https://app.devin.ai/desktop/session/53a7da17cdb047a196848100235a5421?variant=devin
Requested by: @kerry-berri

…hunks and spend rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 24, 2026 03:42
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Adds service tier tracking to streaming responses.

The PR appears safe to merge; no outstanding or new actionable findings remain.

Summary

The PR carries provider-served tiers through streamed chat chunks, stream reassembly, and spend logging, and prices calls using the served tier when available.

  • Exposes the inner chat stream to /v1/messages disconnect billing through its adapter, cache, and router wrappers.
  • Adds integration and unit coverage for served-tier billing across chat, Messages, Responses bridging, Azure, Databricks, and Gemini streams.

Reviews (13) · Last reviewed commit: "refactor(cost): drop explanatory comment..."

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@codspeed

codspeed Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_stream_served_service_tier (aaf6c94) with main (d2cbc94)

Open in CodSpeed

@codecov

codecov Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.41379% with 3 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/router.py 84.61% 2 Missing ⚠️
litellm/proxy/proxy_server.py 50.00% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

…ge chunk

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@CLAassistant

CLAassistant commented Sep 24, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
0 out of 2 committers have signed the CLA.

❌ devin-ai-integration[bot]
❌ kerry-berri
You have signed the CLA already but the status is still pending? Let us recheck it.

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

kerry and others added 2 commits September 25, 2026 00:43
…rtial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Comment thread litellm/llms/anthropic/experimental_pass_through/adapters/streaming_iterator.py Outdated

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/llms/anthropic/experimental_pass_through/adapters/streaming_iterator.py Outdated
kerry and others added 3 commits September 25, 2026 01:28
…s bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…tream wrapper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…_service_tier

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	tests/test_litellm/litellm_core_utils/test_litellm_logging.py
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

kerry and others added 2 commits September 27, 2026 00:51
…ved tiers in the stream billing integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…bill it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/llms/databricks/chat/transformation.py Outdated
…wargs dict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@cursor

cursor Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit aaf6c94. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks!

@kerry-berri
kerry-berri merged commit e814532 into main Sep 29, 2026
104 of 108 checks passed
@mateo-berri

Copy link
Copy Markdown
Contributor

but the final usage chunk comes back without service_tier (on gpt-5.x models via the Responses bridge, no chunk has it)

Follow-up check on whether this is OpenAI not sending us this stuff, or it's just that we were getting them from the provider's chunks, but just not forwarding it to the client or using it would be good

@devin-ai-integration can you do the follow-up and post a GitHub comment on what's happening?

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

OpenAI does send it: the Responses stream carries service_tier on response.created/in_progress/completed (full Response object), never on delta events. Before this PR the bridge only stamped the terminal chunk and reassembly dropped it, so the flow wording "no chunk has it" overstated it

lets-order-some-fries added a commit to lets-order-some-fries/litellm that referenced this pull request Sep 30, 2026
Resolves the conflict at the end of
tests/unit/litellm_core_utils/test_streaming_handler.py, where BerriAI#42870
appended a test after the same context: both blocks kept, main's first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
yuneng-berri added a commit that referenced this pull request Oct 1, 2026
…served at default

#42870 added both the rule that a served default or standard tier bills at
base pricing and records no service_tier, and streamed tests expecting the
row to record 'default'. They have failed on every scheduled litellm-e2e run
since. The tests now map the served tier to the pricing basis the bill must
record and check input is billed at that basis's rate; the messages case
registers custom rates so the rate check has something to compare against
yuneng-berri added a commit that referenced this pull request Oct 1, 2026
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden

#43063 stamps used_client_oauth_token into spend-log metadata, so
test_async_gcs_pub_sub_v1 failed on main with an extra metadata key

* test(ui): give the auto-router threshold save wait room for the availability debounce

#42625 keeps Save disabled while a 300ms-debounced availability check runs.
This test waits for Save right after the change, so the whole debounce lands
inside waitFor's 1s default and it times out under CI load. It is the
recurring UI Unit Tests failure on main since #42625 landed

* test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default

#42870 added both the rule that a served default or standard tier bills at
base pricing and records no service_tier, and streamed tests expecting the
row to record 'default'. They have failed on every scheduled litellm-e2e run
since. The tests now map the served tier to the pricing basis the bill must
record and check input is billed at that basis's rate; the messages case
registers custom rates so the rate check has something to compare against

* test(e2e-ui): wait for the call-id search before hovering the logs row

The row the spec hovers is already on the unfiltered first page, so it was
found before the search request returned. The search response then
re-rendered the table under the mouse, and the Base UI tooltip never opened.
Reproduced with Playwright against a local proxy: hovering right after the
fill never shows the tooltip, hovering after the search response shows the
call id every time

* test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off

The case picked the cheapest Together row flagged supports_response_schema.
DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the
pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token
budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model
the reasoning_effort=none case already exercises, and Together lists it with
structured output support

* test(integration): read the agent 365 guardrail status by its own name in spend logs

The MCP shard runs under xdist against one database, and a sibling file creates a
default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so
that filter's 'success' entry could land first in guardrail_information and the test
read it instead of the agent 365 verdict

* test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check

gc.collect() inside the caplog window can collect a pending task an earlier test left
on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this
test's records. The check still counts every LiteLLM logger, and unretrieved task
exceptions on this loop still go through the asserted exception handler

* test(e2e-ui): fill the create-tag fields inside the dialog

#42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag
Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description')
match two elements and Playwright's strict mode fails the create step

* test(integration): run integration proxies with the CI license

Multi-worker proxies start each uvicorn worker in a fresh process, so every
worker reads the license from its environment. Forward LITELLM_LICENSE into the
proxy and test runner environments

* ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5

Every pull request saved its own uv, maturin, Rust and Prisma caches, about
4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries
within minutes. Pull request jobs then missed every cache, downloaded all
dependencies from PyPI and hit the install step timeouts. Pull requests now
restore only, and main keeps the caches warm for them. test-linting and
check-ui-api-types run only on pull requests and keep saving

codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity
keybase account, so every upload failed signature verification. 5.5.5 reads it
from codecovsecops; the key ID matches the one signing the current CLI

* test(unit): join the session-minting thread before collecting the handler

asyncio.to_thread resumes the test as soon as the worker sets its result,
while the pool thread can still hold the work item and through it the
handler. gc.collect() then cannot finalize the handler and the session stays
open. A pool that shuts down before the test continues drops that reference

* test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early

owned_proxy_process released its reserved port and the proxy bound it only
after full startup, so another xdist worker or an outgoing connection could
take it first and the proxy exited with 'address already in use'. The launch
now retries on a fresh port when that happens and stops every failed attempt.

uvicorn closes idle keep-alive connections after 5 seconds and httpx expired
them at the same 5 seconds, so a request sent right at that mark could reuse a
socket the server was closing and get 'Connection reset by peer'. Gateway
clients now drop idle connections after 2 seconds

* ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build

The release profile builds with fat LTO and one codegen unit, so the final
link of litellm-cache-s3 runs silently for minutes. Successful builds take
711 to 749 seconds, right at the default 10 minute no-output limit, and about
30% of recent runs were killed there

* test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections

The proxy retries the database about every 30 seconds and each retry opens
roughly one connection, so a 5-connection outage took 3 to 4 retries to clear
and recovery landed between 60 and 90 seconds, straddling the test's 80 second
reset window. A fixed 10 second outage still refuses the immediate reconnect
and recovers on the next retry

* ci: move the unit-test uv cache split into a composite action

check_workflow_startup_safety sums every setup step's timeout, so the save and
restore variants each counted 5 minutes although only one runs. One composite
step keeps the setup ceiling at 35 minutes

* test(unit): point tiktoken at the bundled cache for every unit test

The rust_bridge tokenizer tests loaded o200k_base before any test in their
xdist worker had imported default_encoding, so tiktoken fell back to the
temp cache and tried to download under pytest-socket. Move the session
fixture from litellm_core_utils/conftest.py to the root unit conftest.

* test(integration): answer model discovery probes in the hosted_vllm wire tests

The router's periodic upstream model info refresh sends GET /v1/models to
hosted_vllm deployments, so a wire server that is live during a refresh
sees an extra request. Answer the probe with an empty model list and leave
it out of the provider-call assertions, matching the responses bridge
tests.

This branch is waiting to be deployed

1 waiting deployment
e2e-changed — aaf6c947 Waiting Sep 29, 2026 by devin-ai-integration[bot] via Run changed e2e tests against the stage-mirror stack #14462
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants