Skip to content

fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker - #42607

Merged
mateo-berri merged 9 commits into
mainfrom
litellm_bedrock_v1_messages_stream_passthrough
Sep 23, 2026
Merged

mateo-berri merged 9 commits into
mainfrom
litellm_bedrock_v1_messages_stream_passthrough

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • /v1/messages on Bedrock Invoke held upstream bytes in a 1024-byte chunker
  • Complete SSE events sat in that buffer while Bedrock paused mid-stream
  • Claude Code saw a dead stream while the model wrote a large file

How it solves it:

  • Pass Bedrock bytes straight through, like the chat/completions Bedrock paths do
  • Drop the hardcoded DEFAULT_CHUNK_SIZE = 1024 on the Invoke messages decoder

User Flow

Before: a Claude Code user on a bedrock/invoke/ Claude deployment sees a dead stream while the model writes a large file

  1. They ask Claude Code to write a large file, which sends POST http://localhost:4000/v1/messages with "stream": true
  2. Bedrock emits message_start and content_block_start right away, but they add up to under 1024 bytes, so nothing reaches the client
  3. Bedrock then goes quiet for the whole time the model generates the tool input (over two minutes on large files)
  4. Claude Code sees zero events since the request started and shows "Waiting for API response... check your network"
  5. Only when Bedrock resumes and the buffer crosses 1024 bytes do all held events land at once

After: the same request streams each event the moment Bedrock sends it

  1. They ask Claude Code to write the same large file, sending the same POST http://localhost:4000/v1/messages with "stream": true
  2. message_start and content_block_start arrive at the client within milliseconds of Bedrock sending them
  3. Bedrock goes quiet for the same stretch while the model generates the tool input
  4. Claude Code has already received the first events, so the pause is the same one it sees on the Anthropic API directly
  5. Later events stream through one by one as Bedrock emits them, not in 1 KB batches

Relevant issues

Affected release

Linear ticket

Resolves LIT-8266

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live proxy A/B against real Bedrock, no mocks. One proxy per commit, each from its own worktree, both started with the same one-model config on a random free port. Before is a worktree at the merge base 238f434 on port 48630, After is a worktree at the tip 0d2546e on port 21465:

# lit8266-qa.yaml
model_list:
  - model_name: bedrock-invoke-sonnet-4-6
    litellm_params:
      model: bedrock/invoke/us.anthropic.claude-sonnet-4-6
      aws_region_name: us-east-1

general_settings:
  master_key: sk-1234
AWS_BEARER_TOKEN_BEDROCK=... AWS_REGION_NAME=us-east-1 uv run python -m litellm.proxy.proxy_cli --config lit8266-qa.yaml --port <port>

The streaming probe is the customer's request shape, an immediate tool call whose input is large, with no text before it. Each event: line is stamped with the wall clock and the seconds since the request was sent:

BODY='{"model":"bedrock-invoke-sonnet-4-6","max_tokens":6000,"stream":true,"system":"Call the write_file tool immediately. No text before the tool call.","tools":[{"name":"write_file","description":"Write docs/reverse-proxy.md to disk with the given content","input_schema":{"type":"object","properties":{"content":{"type":"string"}},"required":["content"]}}],"messages":[{"role":"user","content":"Use write_file to create a detailed 1500 word guide to configuring nginx as a reverse proxy, with sections, config examples, and troubleshooting."}]}'
START=$(gdate +%s%3N); echo "request sent $(gdate +%T.%3N)"
curl -sN "http://localhost:<port>/v1/messages" -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' -d "$BODY" \
  | perl -MTime::HiRes=time -ne "BEGIN{\$|=1; \$s=$START/1000} if(/^event:/){my \$t=time; my @l=localtime(\$t); printf \"%02d:%02d:%02d.%03d  +%6.3fs  %s\", \$l[2],\$l[1],\$l[0],(\$t-int(\$t))*1000, \$t-\$s, \$_}"

Two things to know when reading the timelines. The ping events every 15 s are the proxy's own SSE keepalive; Bedrock sends no ping frames on this path. And Bedrock itself goes quiet for 30 to 40 s after content_block_start while the model generates the tool input. That upstream pause is the same on both legs, so the thing to compare is when the events Bedrock sent before the pause reach the client.

What Bedrock sends before that pause, captured with the same request sent straight to Bedrock and the frames sized as botocore decodes them: message_start 687 B and content_block_start 319 B at +1.4 s, then a first content_block_delta of 259 B in that capture, then nothing until +41.9 s. Bedrock pads every event by a random amount, and whether a first delta lands before the pause varies per request, so the pre-pause total sits right around 1024 bytes. When it stays under, as in the Before leg below, the chunker holds every event until Bedrock resumes.

Claude Code is the customer's client, so it was driven on both legs too: the interactive TUI (v2.1.280) under tmux, pointed at each proxy with ANTHROPIC_BASE_URL=http://localhost:<port> ANTHROPIC_AUTH_TOKEN=sk-1234 ANTHROPIC_MODEL=bedrock-invoke-sonnet-4-6 ANTHROPIC_SMALL_FAST_MODEL=bedrock-invoke-sonnet-4-6 ANTHROPIC_DEFAULT_HAIKU_MODEL=bedrock-invoke-sonnet-4-6 CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 claude --dangerously-skip-permissions, given /clear and then the prompt "Immediately use the Write tool, with no text before or after the tool call, to create nginx-guide-.md: a detailed 1500 word guide to configuring nginx as a reverse proxy, with sections, config examples, and troubleshooting.", with the pane captured 12 s, 25 s, and 40 s after the prompt was sent. Six attempts were run in total, three at the tip; attempts 5 and 6 are shown.

Before (238f434)

Large tool input over /v1/messages, streaming

  1. Run the probe against port 48630
  2. Observed: the first event of any kind is the proxy's keepalive at +21.8 s, and message_start reaches the client at +45.6 s, only once Bedrock resumed and the buffer crossed 1024 bytes. 1773 events in total, the largest gaps 15.0 s before a ping, 13.2 s before a content_block_delta, and 8.8 s before message_start
request sent 18:43:46.395
18:44:08.167  +21.804s  event: ping
18:44:23.164  +36.801s  event: ping
18:44:31.941  +45.577s  event: message_start
18:44:31.942  +45.579s  event: content_block_start
18:44:31.942  +45.579s  event: content_block_delta
18:44:31.942  +45.579s  event: content_block_delta
18:44:31.942  +45.579s  event: content_block_delta
18:44:31.942  +45.579s  event: content_block_delta
...
18:44:50.330  +63.967s  event: content_block_stop
18:44:50.330  +63.967s  event: message_delta
18:44:50.331  +63.967s  event: message_stop

Non-streaming /v1/messages parity

  1. curl -s http://localhost:48630/v1/messages -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' -d '{"model":"bedrock-invoke-sonnet-4-6","max_tokens":60,"messages":[{"role":"user","content":"Reply with exactly the word ok."}]}'
  2. Observed: type message, role assistant, stop_reason end_turn, one text block; top-level keys content, id, model, role, stop_details, stop_reason, stop_sequence, type, usage; usage keys cache_creation, cache_creation_input_tokens, cache_read_input_tokens, input_tokens, output_tokens, all counts positive

Text-only streaming parity

  1. curl -sN http://localhost:48630/v1/messages -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' -d '{"model":"bedrock-invoke-sonnet-4-6","max_tokens":120,"stream":true,"messages":[{"role":"user","content":"Write two short sentences about the sea."}]}'
  2. Observed: 28 events, message_start content_block_start content_block_delta x23 content_block_stop message_delta message_stop, the message_delta usage carrying the same five keys with a positive output_tokens

SDK entrypoint, litellm.anthropic.messages.acreate(stream=True)

  1. From the worktree at this commit, await litellm.anthropic.messages.acreate(model="bedrock/invoke/us.anthropic.claude-sonnet-4-6", max_tokens=120, stream=True, messages=[{"role": "user", "content": "Write two short sentences about the sea."}]) and iterate the stream
  2. Observed: first event at 1.419 s, 22 events, message_start content_block_start content_block_delta x17 content_block_stop message_delta message_stop

Claude Code, interactive

  1. Send the prompt to the Claude Code session pointed at port 48630 and capture the pane at 12 s, 25 s, and 40 s
  2. Observed, attempt 6: the spinner sits on ↓ 125 tokens at 12 s, 25 s, and still at 40 s; attempt 5 sits on ↓ 114 tokens at 12 s and 25 s and shows Wrote 262 lines to nginx-guide-5.md by 40 s. No "Waiting for API response" or "check your network" line appeared on any of the six attempts

pr42607-0d2546e74e-lit8266-cc2-before-a6-t25.png
pr42607-0d2546e74e-lit8266-cc2-before-a6-t40.png
pr42607-0d2546e74e-lit8266-cc2-before-a5-t25.png

After (0d2546e)

Large tool input over /v1/messages, streaming

  1. Run the probe against port 21465
  2. Observed: message_start, content_block_start, and the first content_block_delta reach the client at +2.2 s, the moment Bedrock sent them. The keepalives then cover Bedrock's own pause and the deltas resume at +40.4 s. 1747 events in total, the largest gaps 15.0 s and 15.0 s before pings and 8.2 s before a content_block_delta; no event was held behind the pause
request sent 19:29:13.849
19:29:16.007  + 2.167s  event: message_start
19:29:16.010  + 2.169s  event: content_block_start
19:29:16.010  + 2.169s  event: content_block_delta
19:29:31.018  +17.178s  event: ping
19:29:46.052  +32.211s  event: ping
19:29:54.239  +40.399s  event: content_block_delta
19:29:54.240  +40.399s  event: content_block_delta
19:29:54.240  +40.399s  event: content_block_delta
...
19:29:55.326  +41.485s  event: content_block_stop
19:29:55.393  +41.553s  event: message_delta
19:29:55.413  +41.572s  event: message_stop

Non-streaming /v1/messages parity

  1. Same curl against port 21465
  2. Observed: identical to Before: type message, role assistant, stop_reason end_turn, one text block, the same nine top-level keys and the same five usage keys, all counts positive

Text-only streaming parity

  1. Same curl against port 21465
  2. Observed: 24 events, message_start content_block_start content_block_delta x19 content_block_stop message_delta message_stop, the message_delta usage carrying the same five keys with a positive output_tokens

SDK entrypoint, litellm.anthropic.messages.acreate(stream=True)

  1. Same call from the worktree at this commit
  2. Observed: first event at 1.330 s, 23 events, message_start content_block_start content_block_delta x18 content_block_stop message_delta message_stop

Claude Code, interactive

  1. Send the prompt to the Claude Code session pointed at port 21465 and capture the pane at 12 s, 25 s, and 40 s
  2. Observed, attempt 6: the spinner sits on ↓ 94 tokens at 12 s and 25 s and shows Wrote 279 lines to nginx-guide-6.md by 40 s; attempt 5 sits on ↓ 106 tokens at 13 s, 27 s, and 41 s. The on-screen stall is the same length on both legs (it is Bedrock's own tool-input pause, and the proxy keepalives keep the connection visibly alive on both), so this leg does not separate the two commits; the wire timelines above do

pr42607-0d2546e74e-lit8266-cc2-after-a6-t25.png
pr42607-0d2546e74e-lit8266-cc2-after-a6-t40.png
pr42607-0d2546e74e-lit8266-cc2-after-a5-t25.png

Type

🐛 Bug Fix

Caveats (if any)

Low

Blast radius at 0d2546e:

  • Breaking: none. No signature, config key, or response shape changes
  • Backward incompatible: none. Event order, event count, and usage keys match the base on every leg above
  • Regression risk: delivery cadence only. aiter_bytes() now yields per network read, and botocore's EventStreamBuffer already reassembles frames across arbitrary byte boundaries, so partial frames are handled exactly as before
  • Dependency graph: one call site and one subclass __init__; the class-level DEFAULT_CHUNK_SIZE had no other reader. Reached by /v1/messages streaming on bedrock/invoke/<claude model> and by litellm.anthropic.messages.acreate(stream=True), both driven above. Converse (bedrock/<model>) and /v1/chat/completions never used this decoder
  • Not verified: nothing; there is no sync streaming iterator on this path to check

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • 0d2546e passes /live-pr-risk

Link to Devin session: https://app.devin.ai/sessions/5dbfd4f448f044159872e6dfa698dfbc
Open in Devin Desktop: https://app.devin.ai/desktop/session/5dbfd4f448f044159872e6dfa698dfbc?variant=devin
Requested by: @mateo-berri


Note

Low Risk
Narrow change to Bedrock Invoke messages async streaming read sizing; aligns with pass-through behavior elsewhere and is covered by a new streaming regression test.

Overview
Bedrock Invoke /v1/messages streaming no longer buffers upstream bytes in 1024-byte chunks before decoding. get_async_streaming_response_iterator now feeds httpx_response.aiter_bytes() straight into the AWS event-stream decoder, and the messages-specific DEFAULT_CHUNK_SIZE override on AmazonAnthropicClaudeMessagesStreamDecoder is removed.

That fixes cases where early SSE events (e.g. message_start) stayed below 1 KB and never reached the client during long upstream pauses—clients like Claude Code could show a dead connection until the buffer filled.

A regression test builds a minimal AWS event-stream frame and a gated httpx byte stream to assert the first message_start SSE is emitted before the upstream resumes.

Reviewed by Cursor Bugbot for commit 0d2546e. Bugbot is set up for automated code reviews on this repo. Configure here.

…ding them in a 1024-byte chunker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge because the streaming change is narrowly scoped and the regression test directly covers the delayed-first-event failure.

Summary

Removes the fixed 1024-byte read size from Bedrock Invoke /v1/messages streaming so AWS event-stream frames can be decoded and forwarded as soon as upstream bytes arrive.

  • Passes the default httpx byte iterator directly to the Bedrock stream decoder.
  • Removes the messages decoder’s custom chunk-size override.
  • Adds a gated asynchronous regression test proving a small first frame is emitted before the upstream resumes.

Reviews (7) · Last reviewed commit: "Merge branch 'main' of https://github.co..."

mateo-berri and others added 2 commits September 22, 2026 22:54
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@codecov

codecov Bot commented Sep 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_bedrock_v1_messages_stream_passthrough (0d2546e) with main (b395bfe)

Open in CodSpeed

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

…gression test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 0d2546e. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 975bd28 into main Sep 23, 2026
94 checks passed
@mateo-berri
mateo-berri deleted the litellm_bedrock_v1_messages_stream_passthrough branch September 23, 2026 03:01
mateo-berri added a commit that referenced this pull request Sep 23, 2026
fix(bedrock): backport the /v1/messages Invoke streaming pass-through to rc/1.103.0 (#42607)
yuneng-berri added a commit that referenced this pull request Sep 23, 2026
…da70a

chore(release): backport #42607 to stable/1.101.x and cut 1.101.2
leikaiwei pushed a commit to leikaiwei/litellm that referenced this pull request Sep 30, 2026
…ding them in a 1024-byte chunker (BerriAI#42607)

* fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(bedrock): apply ruff format to invoke messages stream passthrough

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(bedrock): drop drive-by reformat of existing invoke messages tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): collect streamed chunks into a tuple in passthrough regression test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): give the passthrough regression test a 10s first-chunk budget

* test(bedrock): type the eventstream frame helper's payload as Mapping[str, object]

* test(bedrock): take the gated byte stream's chunks as an immutable Sequence

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
(cherry picked from commit 975bd28)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant