Skip to content

fix(streaming): backport text-completion usage fix and e2e provider-flake tolerance to rc/1.103.0 - #43400

Merged
yuneng-berri merged 2 commits into
rc/1.103.0from
litellm_rc_1_103_0_backport_42628_43047
Sep 27, 2026
Merged

yuneng-berri merged 2 commits into
rc/1.103.0from
litellm_rc_1_103_0_backport_42628_43047

Conversation

@yuneng-berri

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Streaming /v1/completions with include_usage can end in a MockValSer error instead of usage
  • Two rc/1.103.0 e2e cells fail on provider-side behavior, not LiteLLM code

How it solves it:

User Flow

Before: a developer streaming a legacy text completion with usage turned on gets an error where the usage should be

  1. They send POST http://localhost:4000/v1/completions with "stream": true and "stream_options": {"include_usage": true}
  2. The text chunks arrive normally
  3. The last event is an error mentioning 'MockValSer' object is not an instance of 'SchemaSerializer', and no usage arrives

After: the same request ends with real token counts

  1. They send the same POST http://localhost:4000/v1/completions
  2. The text chunks arrive normally
  3. The last chunk carries usage with prompt, completion and total tokens

Relevant issues

Backport of #42628 and #43047

Affected release

regression in v1.94.0-rc.1

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

The MockValSer bug came in with #32255 (commit 8a4942340e, first in v1.94.0-rc.1), which copies the OpenAI SDK's CompletionUsage onto litellm's streamed response as is. The SDK defers building its pydantic serializers, so in a process where nothing has built that class yet, dumping the chunk fails. Run alone against the real OpenAI API, tests/local_testing/test_streaming.py::test_openai_stream_options_call_text_completion passes at the parent of 8a4942340e, fails at 8a4942340e, fails on the rc/1.103.0 tip, and passes on this branch. In CircleCI it only fails when no earlier test on the same worker has built the class, which is why it looks flaky there (5 failures in the last 59 main runs)

test_streaming_handler.py passes on this branch (138 tests), and #43047's two new cases fail with its source change reverted. The e2e files touched by #42628 collect, and test_e2e_http.py plus test_batch_cleanup.py pass (71 tests)

Adaptations from the main versions:

Type

🐛 Bug Fix
✅ Test

Caveats (if any)

Medium

devin-ai-integration Bot and others added 2 commits September 26, 2026 18:55
…2628)

* test(e2e): tolerate provider-side flakes on five full-suite cells

Mistral OCR retries a provider-relayed 429 with backoff, the Vertex vision
probe turns reasoning off so its 32 tokens go to the answer, the Vertex
cache cell spaces eight never-seen prefixes 15s apart around Google's
nondeterministic minimum-token rejection and prices the cached tokens
instead of prompt_tokens, and the Azure content-policy cell resends the
jailbreak prompt while Azure skips its filter

* test(e2e): shorten the new helper docstrings

* test(e2e): accept a relayed provider 429 on the rust OCR cells

The gateway already retries a provider 429 three times per call and the
Mistral key is shared across pipelines, so a throttle can hold across all
four attempts of the OCR cell. After the bounded retries the cell now
accepts the gateway's faithful relay of the provider's 429 (throttling_error,
code 429) as its second expected outcome; the gateway's own 429 and any
other error still fail the cell at once.

* test(e2e): drop the harness unit tests, the live cells cover the helpers

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
(cherry picked from commit 41ca465)
…43047)

* fix(streaming): keep litellm Usage on text-completion usage chunks

* fix(streaming): convert provider usage to litellm Usage instead of dropping it

(cherry picked from commit be35b22)
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ yuneng-berri
❌ devin-ai-integration[bot]
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Fixes text-completion usage handling and adds provider rate-limit tolerance to e2e tests.

The PR should not merge until the OCR test requires a successful document and the explicit variable-annotation requirement is satisfied

Findings

  1. P1 OCR can pass without OCR ▶
  2. P2 New locals lack Final ▶

Summary

The PR converts SDK usage into LiteLLM usage for streamed text completions and adjusts live e2e tests for provider variability

  • The streaming regression test checks token counts and usage details for OpenAI and Azure text-completion paths
  • The OCR change permits persistent provider 429s to pass without an OCR result, weakening that test’s signal

Reviews (1) · Last reviewed commit: "fix(streaming): keep litellm Usage on te..."

Comment on lines +187 to +188
case RateLimitedError() as outcome:
_assert_provider_rate_limit_relayed(model, outcome)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 OCR can pass without OCR If all four attempts return a matching provider 429, this case passes without checking a document. All four provider cases could pass without successful OCR, violating the repository rule against weakening existing tests to mask regressions.

Rule Used: What: Flag any modifications to existing tests and verify they don't weaken test coverage or mask regressions. Why: Developers may alter tests to make failing code pass rather than fix the actual bug, hiding regressions. Good: ``` // Test updated t... (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment thread tests/e2e/e2e_http.py
for attempt in range(1, attempts):
match issue():
case RateLimitedError(body=body, retry_after_seconds=retry_after) if PROVIDER_RATE_LIMIT_MARKER in body:
delay = retry_after or PROVIDER_RATE_LIMIT_BACKOFF_SECONDS * (1 << (attempt - 1))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 New locals lack Final The new delay local lacks : Final, as do candidate and resp elsewhere in this change. The repository requires every variable to have that annotation; satisfy this requirement before merging.

Context Used: AGENTS.md (source)

@yuneng-berri
yuneng-berri merged commit cc111d1 into rc/1.103.0 Sep 27, 2026
5 of 6 checks passed
@yuneng-berri
yuneng-berri deleted the litellm_rc_1_103_0_backport_42628_43047 branch September 27, 2026 02:12

This branch is waiting to be deployed

1 waiting deployment
e2e-changed — 9decce63 Waiting Sep 27, 2026 by yuneng-berri via Run changed e2e tests against the stage-mirror stack #13232
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants