Skip to content

fix(router): stream /v1/messages lifecycle frames live when no fallback can take over - #43600

Merged
yassin-berriai merged 5 commits into
mainfrom
litellm_messages_stream_live_lifecycle
Sep 29, 2026
Merged

yassin-berriai merged 5 commits into
mainfrom
litellm_messages_stream_live_lifecycle

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • /v1/messages streams hold message_start until the first content delta
  • Adaptive thinking on Bedrock and Vertex can run for minutes before the first delta
  • Clients see only pings for about two minutes, and message_start shows up with the first delta

How it solves it:

  • Hold lifecycle frames back only when a fallback can actually take over
  • That check mirrors the fallback dispatcher: order levels, weighted failover, content-policy and generic fallbacks, request overrides
  • With no fallback, or disable_fallbacks, every frame streams live
  • Pings are always forwarded live before content, so the connection stays warm
  • Only a whole ping frame goes live: a chunk counts as a ping when it starts with event: ping and ends on a frame boundary, so a ping the transport splits across two reads stays in the buffer, in order, instead of its head jumping ahead of message_start

Intentional product change: when a fallback is configured, a ping can now reach the client before message_start. Anthropic's SDKs ignore ping events, and leading pings were already forwarded this way

Mutation check on the regression tests: dropping the startswith(b"event: ping") guard fails test_is_anthropic_ping_chunk_only_matches_whole_ping_frames; dropping the frame boundary guard also fails test_anthropic_messages_split_ping_stays_in_order_behind_buffered_lifecycle_frame; restoring the pre-fix classifier fails those two plus one more. All green with the fix in place

Supersedes #39566, which gated the buffering on fallbacks alone and was written by @Swigler, credited as co-author on this branch. The Bedrock half of the issue (1024 byte aiter_bytes chunks splitting frames) already landed on main in #42607

User Flow

Before: a developer streaming Claude with adaptive thinking through the proxy gets no message_start until thinking ends

  1. They send POST https://litellm-domain/v1/messages with "stream": true and adaptive thinking, to a model group with no fallbacks
  2. The upstream sends message_start right away, but the client only receives pings
  3. message_start, content_block_start and the first delta arrive together once thinking ends, about two minutes later

After: the stream starts immediately

  1. They send the same POST https://litellm-domain/v1/messages
  2. message_start and content_block_start reach the client as soon as the upstream sends them
  3. Deltas follow as they arrive, and pings keep the connection alive in between

Relevant issues

Fixes #39431

Affected release

regression in v1.100.0

Linear ticket

Resolves LIT-8908

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Real Bedrock Claude Opus 4.7 and 4.8 through a local proxy, no mocks. Proxy config: opus47 -> bedrock/us.anthropic.claude-opus-4-7, opus48 -> bedrock/us.anthropic.claude-opus-4-8, num_retries: 0. The no fallback case has no fallbacks; the fallback case adds router_settings.fallbacks: [{"opus47": ["opus48"]}]

Request req.json: model: opus47, "stream": true, "thinking": {"type": "adaptive"}, max_tokens: 16000, a long multiplication puzzle that keeps the model thinking before the first delta. Each event: line is stamped with seconds since the request started:

start=$(date +%s.%N)
curl -sN http://localhost:$PORT/v1/messages -H "Authorization: Bearer sk-1234" -H "Content-Type: application/json" -d @req.json \
| while IFS= read -r line; do case "$line" in event:*) printf '%.3f %s\n' "$(echo "$(date +%s.%N) - $start" | bc)" "$line";; esac; done

Before (6c34c6e)

No fallback configured

  1. Send the request to opus47 three times
  2. message_start arrives at 105.5s, 118.0s and 117.5s, in the same instant as content_block_start and the first delta; only pings arrive before that
  3. Same against opus48: message_start at 153.2s, 129.8s and 147.8s

Fallback configured

  1. Send the request to opus47 once
  2. message_start arrives only with the first content, 3.5s after the last ping:
16.491 event: ping
19.978 event: message_start
19.980 event: content_block_start
19.983 event: content_block_delta
...
20.565 event: message_delta
20.568 event: message_stop

After (af71a09, same tree as eb1c353 where this run was captured; af71a09 is an empty co-author commit)

No fallback configured

  1. Send the request to opus47 once
  2. message_start and content_block_start arrive within 1.7s, 30s ahead of the first delta, with pings in between:
1.570 event: message_start
1.712 event: content_block_start
16.711 event: ping
31.711 event: ping
31.911 event: content_block_delta
31.913 event: content_block_delta
31.917 event: content_block_stop
31.919 event: content_block_start
31.921 event: content_block_delta
...
32.747 event: message_delta
32.749 event: message_stop

Fallback configured

  1. Send the request to opus47 once
  2. Pings reach the client live every 15s; message_start is still held until content so a mid-stream fallback can start a clean message:
17.525 event: ping
32.526 event: ping
47.526 event: ping
49.590 event: message_start
49.592 event: content_block_start
49.594 event: content_block_delta
...
53.748 event: message_delta
53.750 event: message_stop

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • With a fallback configured, message_start is still held until content
    • That is what keeps a mid-stream fallback from producing two message lifecycles
    • Pings now flow live in that case, which covers idle read timeouts

Low

  • With no fallback, a retriable provider error frame is now forwarded verbatim instead of raised
  • A ping split across two transport reads is buffered instead of forwarded live while lifecycle frames are held, so that one keepalive reaches the client with the first content rather than on arrival
  • The misc / Run tests shard fails on main too (tests/unit/interactions/test_openapi_compliance.py, same four tests on b9ba36c, 6d4ccf7 and 9fd78ff), unrelated to this change

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/a83062f73f974849912463d9465dba14
Open in Devin Desktop: https://app.devin.ai/desktop/session/a83062f73f974849912463d9465dba14?variant=devin
Requested by: @yassin-berriai

…fallback can take over

The /v1/messages streaming wrapper buffered message_start and
content_block_start until the first content_block_delta and dropped
pings behind buffered frames unconditionally, even for requests no
fallback could ever recover. With adaptive thinking on Bedrock or
Vertex the client saw no bytes for the whole thinking pass and hit
read timeouts.

Buffering now applies only while a fallback can still take over
(generic or refusal chain resolving), and a ping is always forwarded
live since it carries no lifecycle and keeps the connection alive.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

devin-ai-integration Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Changes how the router handles Anthropic message streaming and ping frames.

The PR appears safe to merge based on the changes reviewed.

Summary

The PR forwards Anthropic Messages lifecycle frames immediately when no fallback can take over, while retaining lifecycle buffering when fallback remains possible. It also forwards complete ping frames live and adds unit and wire-level coverage.

Reviews (5) · Last reviewed commit: "fix(router): keep a transport-split ping..."

Comment thread litellm/router.py
Comment thread litellm/router.py Outdated
@greptile-apps

This comment has been minimized.

@codecov

codecov Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_messages_stream_live_lifecycle (af71a09) with main (6d4ccf7)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (317430d) during the generation of this report, so 6d4ccf7 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

…tream gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/router.py Outdated
…gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/router.py
…ames instead of forwarding its head live

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ mateo-berri
❌ yassin-berriai
You have signed the CLA already but the status is still pending? Let us recheck it.

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

Co-authored-by: Radu Swigler <radu.porumba@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit af71a09. Configure here.

@yassin-berriai
yassin-berriai merged commit d463049 into main Sep 29, 2026
100 of 101 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_messages_stream_live_lifecycle branch September 29, 2026 16:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: /v1/messages streaming delays message_start until the model's thinking pass finishes, even with no fallbacks configured

3 participants