Repository navigation
fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker - #42607
Conversation
…ding them in a 1024-byte chunker Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…gression test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
|
bugbot run |
…ssages_stream_passthrough
|
bugbot run |
|
bugbot run |
…m_bedrock_v1_messages_stream_passthrough
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 0d2546e. Configure here.
fix(bedrock): backport the /v1/messages Invoke streaming pass-through to rc/1.103.0 (#42607)
…da70a chore(release): backport #42607 to stable/1.101.x and cut 1.101.2
…ding them in a 1024-byte chunker (BerriAI#42607) * fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(bedrock): apply ruff format to invoke messages stream passthrough Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(bedrock): drop drive-by reformat of existing invoke messages tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(bedrock): collect streamed chunks into a tuple in passthrough regression test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(bedrock): give the passthrough regression test a 10s first-chunk budget * test(bedrock): type the eventstream frame helper's payload as Mapping[str, object] * test(bedrock): take the gated byte stream's chunks as an immutable Sequence --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> (cherry picked from commit 975bd28)
TLDR
Problem this solves:
/v1/messageson Bedrock Invoke held upstream bytes in a 1024-byte chunkerHow it solves it:
DEFAULT_CHUNK_SIZE = 1024on the Invoke messages decoderUser Flow
Before: a Claude Code user on a
bedrock/invoke/Claude deployment sees a dead stream while the model writes a large file"stream": truemessage_startandcontent_block_startright away, but they add up to under 1024 bytes, so nothing reaches the clientAfter: the same request streams each event the moment Bedrock sends it
"stream": truemessage_startandcontent_block_startarrive at the client within milliseconds of Bedrock sending themRelevant issues
Affected release
Linear ticket
Resolves LIT-8266
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Live proxy A/B against real Bedrock, no mocks. One proxy per commit, each from its own worktree, both started with the same one-model config on a random free port. Before is a worktree at the merge base 238f434 on port 48630, After is a worktree at the tip 0d2546e on port 21465:
The streaming probe is the customer's request shape, an immediate tool call whose input is large, with no text before it. Each
event:line is stamped with the wall clock and the seconds since the request was sent:Two things to know when reading the timelines. The
pingevents every 15 s are the proxy's own SSE keepalive; Bedrock sends no ping frames on this path. And Bedrock itself goes quiet for 30 to 40 s aftercontent_block_startwhile the model generates the tool input. That upstream pause is the same on both legs, so the thing to compare is when the events Bedrock sent before the pause reach the client.What Bedrock sends before that pause, captured with the same request sent straight to Bedrock and the frames sized as botocore decodes them:
message_start687 B andcontent_block_start319 B at +1.4 s, then a firstcontent_block_deltaof 259 B in that capture, then nothing until +41.9 s. Bedrock pads every event by a random amount, and whether a first delta lands before the pause varies per request, so the pre-pause total sits right around 1024 bytes. When it stays under, as in the Before leg below, the chunker holds every event until Bedrock resumes.Claude Code is the customer's client, so it was driven on both legs too: the interactive TUI (v2.1.280) under tmux, pointed at each proxy with
ANTHROPIC_BASE_URL=http://localhost:<port> ANTHROPIC_AUTH_TOKEN=sk-1234 ANTHROPIC_MODEL=bedrock-invoke-sonnet-4-6 ANTHROPIC_SMALL_FAST_MODEL=bedrock-invoke-sonnet-4-6 ANTHROPIC_DEFAULT_HAIKU_MODEL=bedrock-invoke-sonnet-4-6 CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 claude --dangerously-skip-permissions, given/clearand then the prompt "Immediately use the Write tool, with no text before or after the tool call, to create nginx-guide-.md: a detailed 1500 word guide to configuring nginx as a reverse proxy, with sections, config examples, and troubleshooting.", with the pane captured 12 s, 25 s, and 40 s after the prompt was sent. Six attempts were run in total, three at the tip; attempts 5 and 6 are shown.Before (238f434)
Large tool input over /v1/messages, streaming
message_startreaches the client at +45.6 s, only once Bedrock resumed and the buffer crossed 1024 bytes. 1773 events in total, the largest gaps 15.0 s before a ping, 13.2 s before acontent_block_delta, and 8.8 s beforemessage_startNon-streaming /v1/messages parity
curl -s http://localhost:48630/v1/messages -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' -d '{"model":"bedrock-invoke-sonnet-4-6","max_tokens":60,"messages":[{"role":"user","content":"Reply with exactly the word ok."}]}'typemessage,roleassistant,stop_reasonend_turn, onetextblock; top-level keyscontent, id, model, role, stop_details, stop_reason, stop_sequence, type, usage; usage keyscache_creation, cache_creation_input_tokens, cache_read_input_tokens, input_tokens, output_tokens, all counts positiveText-only streaming parity
curl -sN http://localhost:48630/v1/messages -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' -d '{"model":"bedrock-invoke-sonnet-4-6","max_tokens":120,"stream":true,"messages":[{"role":"user","content":"Write two short sentences about the sea."}]}'message_start content_block_start content_block_delta x23 content_block_stop message_delta message_stop, themessage_deltausage carrying the same five keys with a positiveoutput_tokensSDK entrypoint,
litellm.anthropic.messages.acreate(stream=True)await litellm.anthropic.messages.acreate(model="bedrock/invoke/us.anthropic.claude-sonnet-4-6", max_tokens=120, stream=True, messages=[{"role": "user", "content": "Write two short sentences about the sea."}])and iterate the streammessage_start content_block_start content_block_delta x17 content_block_stop message_delta message_stopClaude Code, interactive
↓ 125 tokensat 12 s, 25 s, and still at 40 s; attempt 5 sits on↓ 114 tokensat 12 s and 25 s and showsWrote 262 lines to nginx-guide-5.mdby 40 s. No "Waiting for API response" or "check your network" line appeared on any of the six attemptsAfter (0d2546e)
Large tool input over /v1/messages, streaming
message_start,content_block_start, and the firstcontent_block_deltareach the client at +2.2 s, the moment Bedrock sent them. The keepalives then cover Bedrock's own pause and the deltas resume at +40.4 s. 1747 events in total, the largest gaps 15.0 s and 15.0 s before pings and 8.2 s before acontent_block_delta; no event was held behind the pauseNon-streaming /v1/messages parity
typemessage,roleassistant,stop_reasonend_turn, onetextblock, the same nine top-level keys and the same five usage keys, all counts positiveText-only streaming parity
message_start content_block_start content_block_delta x19 content_block_stop message_delta message_stop, themessage_deltausage carrying the same five keys with a positiveoutput_tokensSDK entrypoint,
litellm.anthropic.messages.acreate(stream=True)message_start content_block_start content_block_delta x18 content_block_stop message_delta message_stopClaude Code, interactive
↓ 94 tokensat 12 s and 25 s and showsWrote 279 lines to nginx-guide-6.mdby 40 s; attempt 5 sits on↓ 106 tokensat 13 s, 27 s, and 41 s. The on-screen stall is the same length on both legs (it is Bedrock's own tool-input pause, and the proxy keepalives keep the connection visibly alive on both), so this leg does not separate the two commits; the wire timelines above doType
🐛 Bug Fix
Caveats (if any)
Low
chunk_size=1024makes the new test time outmisc / Run testswas red on main run 35807522965 (fix(utils): isolate callback errors in async_post_call_success_deployment_hook #42535) and got cured by test(utils): raise the post-success hook error from a guardrail in the failure-hook regression #42646, merged into this branch at 0d2546e; the rust-wheel red on main b0ac23d (feat(logger): dispatch Python logging through the Rust diagnostics processor #42616) is not required and not this PR'sBlast radius at 0d2546e:
aiter_bytes()now yields per network read, and botocore'sEventStreamBufferalready reassembles frames across arbitrary byte boundaries, so partial frames are handled exactly as before__init__; the class-levelDEFAULT_CHUNK_SIZEhad no other reader. Reached by/v1/messagesstreaming onbedrock/invoke/<claude model>and bylitellm.anthropic.messages.acreate(stream=True), both driven above. Converse (bedrock/<model>) and/v1/chat/completionsnever used this decoderFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/5dbfd4f448f044159872e6dfa698dfbc
Open in Devin Desktop: https://app.devin.ai/desktop/session/5dbfd4f448f044159872e6dfa698dfbc?variant=devin
Requested by: @mateo-berri
Note
Low Risk
Narrow change to Bedrock Invoke messages async streaming read sizing; aligns with pass-through behavior elsewhere and is covered by a new streaming regression test.
Overview
Bedrock Invoke
/v1/messagesstreaming no longer buffers upstream bytes in 1024-byte chunks before decoding.get_async_streaming_response_iteratornow feedshttpx_response.aiter_bytes()straight into the AWS event-stream decoder, and the messages-specificDEFAULT_CHUNK_SIZEoverride onAmazonAnthropicClaudeMessagesStreamDecoderis removed.That fixes cases where early SSE events (e.g.
message_start) stayed below 1 KB and never reached the client during long upstream pauses—clients like Claude Code could show a dead connection until the buffer filled.A regression test builds a minimal AWS event-stream frame and a gated
httpxbyte stream to assert the firstmessage_startSSE is emitted before the upstream resumes.Reviewed by Cursor Bugbot for commit 0d2546e. Bugbot is set up for automated code reviews on this repo. Configure here.