Repository navigation
fix(router): stream /v1/messages lifecycle frames live when no fallback can take over - #43600
Conversation
…fallback can take over The /v1/messages streaming wrapper buffered message_start and content_block_start until the first content_block_delta and dropped pings behind buffered frames unconditionally, even for requests no fallback could ever recover. With adaptive thinking on Bedrock or Vertex the client saw no bytes for the whole thinking pass and hit read timeouts. Buffering now applies only while a fallback can still take over (generic or refusal chain resolving), and a ping is always forwarded live since it carries no lifecycle and keeps the connection alive. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
This comment has been minimized.
This comment has been minimized.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…tream gate Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…gate Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
…ames instead of forwarding its head live Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
|
|
bugbot run |
Co-authored-by: Radu Swigler <radu.porumba@gmail.com> Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit af71a09. Configure here.
TLDR
Problem this solves:
/v1/messagesstreams holdmessage_startuntil the first content deltamessage_startshows up with the first deltaHow it solves it:
disable_fallbacks, every frame streams livepingframe goes live: a chunk counts as a ping when it starts withevent: pingand ends on a frame boundary, so a ping the transport splits across two reads stays in the buffer, in order, instead of its head jumping ahead ofmessage_startIntentional product change: when a fallback is configured, a
pingcan now reach the client beforemessage_start. Anthropic's SDKs ignore ping events, and leading pings were already forwarded this wayMutation check on the regression tests: dropping the
startswith(b"event: ping")guard failstest_is_anthropic_ping_chunk_only_matches_whole_ping_frames; dropping the frame boundary guard also failstest_anthropic_messages_split_ping_stays_in_order_behind_buffered_lifecycle_frame; restoring the pre-fix classifier fails those two plus one more. All green with the fix in placeSupersedes #39566, which gated the buffering on
fallbacksalone and was written by @Swigler, credited as co-author on this branch. The Bedrock half of the issue (1024 byteaiter_byteschunks splitting frames) already landed onmainin #42607User Flow
Before: a developer streaming Claude with adaptive thinking through the proxy gets no
message_startuntil thinking ends"stream": trueand adaptive thinking, to a model group with no fallbacksmessage_startright away, but the client only receives pingsmessage_start,content_block_startand the first delta arrive together once thinking ends, about two minutes laterAfter: the stream starts immediately
message_startandcontent_block_startreach the client as soon as the upstream sends themRelevant issues
Fixes #39431
Affected release
regression in v1.100.0
Linear ticket
Resolves LIT-8908
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Real Bedrock Claude Opus 4.7 and 4.8 through a local proxy, no mocks. Proxy config:
opus47->bedrock/us.anthropic.claude-opus-4-7,opus48->bedrock/us.anthropic.claude-opus-4-8,num_retries: 0. The no fallback case has nofallbacks; the fallback case addsrouter_settings.fallbacks: [{"opus47": ["opus48"]}]Request
req.json:model: opus47,"stream": true,"thinking": {"type": "adaptive"},max_tokens: 16000, a long multiplication puzzle that keeps the model thinking before the first delta. Eachevent:line is stamped with seconds since the request started:Before (6c34c6e)
No fallback configured
opus47three timesmessage_startarrives at 105.5s, 118.0s and 117.5s, in the same instant ascontent_block_startand the first delta; only pings arrive before thatopus48:message_startat 153.2s, 129.8s and 147.8sFallback configured
opus47oncemessage_startarrives only with the first content, 3.5s after the last ping:After (af71a09, same tree as eb1c353 where this run was captured; af71a09 is an empty co-author commit)
No fallback configured
opus47oncemessage_startandcontent_block_startarrive within 1.7s, 30s ahead of the first delta, with pings in between:Fallback configured
opus47oncemessage_startis still held until content so a mid-stream fallback can start a clean message:Type
🐛 Bug Fix
Caveats (if any)
Medium
message_startis still held until contentLow
misc / Run testsshard fails onmaintoo (tests/unit/interactions/test_openapi_compliance.py, same four tests on b9ba36c, 6d4ccf7 and 9fd78ff), unrelated to this changeFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/a83062f73f974849912463d9465dba14
Open in Devin Desktop: https://app.devin.ai/desktop/session/a83062f73f974849912463d9465dba14?variant=devin
Requested by: @yassin-berriai