Skip to content

fix(bedrock): backport the /v1/messages Invoke streaming pass-through to rc/1.103.0 (#42607) - #42658

Merged
mateo-berri merged 1 commit into
rc/1.103.0from
litellm_backport_42607_rc_1_103_0
Sep 23, 2026
Merged

mateo-berri merged 1 commit into
rc/1.103.0from
litellm_backport_42607_rc_1_103_0

Conversation

@mateo-berri

@mateo-berri mateo-berri commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

User Flow

Before: a Claude Code user on a bedrock/invoke/ Claude deployment sees a dead stream while the model writes a large file

  1. They ask Claude Code to write a large file, which sends POST http://localhost:4000/v1/messages with "stream": true
  2. Bedrock emits message_start and content_block_start right away, but they add up to under 1024 bytes, so nothing reaches the client
  3. Bedrock then goes quiet for the whole time the model generates the tool input (over two minutes on large files)
  4. Claude Code sees zero events since the request started and shows "Waiting for API response... check your network"
  5. Only when Bedrock resumes and the buffer crosses 1024 bytes do all held events land at once

After: the same request streams each event the moment Bedrock sends it

  1. They ask Claude Code to write the same large file, sending the same POST http://localhost:4000/v1/messages with "stream": true
  2. message_start and content_block_start arrive at the client within milliseconds of Bedrock sending them
  3. Bedrock goes quiet for the same stretch while the model generates the tool input
  4. Claude Code has already received the first events, so the pause is the same one it sees on the Anthropic API directly
  5. Later events stream through one by one as Bedrock emits them, not in 1 KB batches

Backported PR

PR Fix Regressed in Commit here Pick
#42607 /v1/messages Invoke streaming held events behind a 1024-byte chunker not a regression, there since #10710 (May 2025) 4a00468 clean

Cherry-picked with git cherry-pick -x 975bd28549, no conflicts, so the diff here is main's byte for byte: the decoder's aiter_bytes() no longer asks httpx for 1024-byte chunks, and the subclass override of DEFAULT_CHUNK_SIZE is gone. #42607 carries no backport-stable label because the chunker is not a regression: it has been on every line since #10710, which is why the pick targets the line the customer runs instead of the newest stable. stable/* lines are untouched

Relevant issues

Affected release

not a regression: present in every release since #10710, including v1.103.0-rc.1 (tagged at ccf5e8c, the current tip of rc/1.103.0)

Linear ticket

Resolves LIT-8392

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests (fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker #42607's test_get_async_streaming_response_iterator_yields_small_frame_before_upstream_pauses comes along unchanged; the whole file is 123 passed on this line)
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.): no required checks on this line; make check passes at 4a00468, and the tip's CI is pipeline 90133, its three reds classified under Caveats
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live proxy A/B on this line against real Bedrock, no mocks. One proxy per commit, each from its own worktree, both started with the same one-model config on a random free port. Before is a detached checkout of the rc/1.103.0 tip ccf5e8c (the commit v1.103.0-rc.1 is tagged at) on port 50413, After is this branch at 4a00468 on port 42507:

# lit8266-qa.yaml
model_list:
  - model_name: bedrock-invoke-sonnet-4-6
    litellm_params:
      model: bedrock/invoke/us.anthropic.claude-sonnet-4-6
      aws_region_name: us-east-1

general_settings:
  master_key: sk-1234
AWS_BEARER_TOKEN_BEDROCK=... AWS_REGION_NAME=us-east-1 uv run python -m litellm.proxy.proxy_cli --config lit8266-qa.yaml --port <port>

The streaming probe is the customer's request shape, an immediate tool call whose input is large, with no text before it. Each event: line is stamped with the wall clock and the seconds since the request was sent:

BODY='{"model":"bedrock-invoke-sonnet-4-6","max_tokens":6000,"stream":true,"system":"Call the write_file tool immediately. No text before the tool call.","tools":[{"name":"write_file","description":"Write docs/reverse-proxy.md to disk with the given content","input_schema":{"type":"object","properties":{"content":{"type":"string"}},"required":["content"]}}],"messages":[{"role":"user","content":"Use write_file to create a detailed 1500 word guide to configuring nginx as a reverse proxy, with sections, config examples, and troubleshooting."}]}'
START=$(gdate +%s%3N); echo "request sent $(gdate +%T.%3N)"
curl -sN "http://localhost:<port>/v1/messages" -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' -d "$BODY" \
  | perl -MTime::HiRes=time -ne "BEGIN{\$|=1; \$s=$START/1000} if(/^event:/){my \$t=time; my @l=localtime(\$t); printf \"%02d:%02d:%02d.%03d  +%6.3fs  %s\", \$l[2],\$l[1],\$l[0],(\$t-int(\$t))*1000, \$t-\$s, \$_}"

Two things to know when reading the timelines. The ping events every 15 s are the proxy's own SSE keepalive; Bedrock sends no ping frames on this path. And Bedrock itself goes quiet for 30 to 40 s after content_block_start while the model generates the tool input. That upstream pause is the same on both legs, so the thing to compare is when the events Bedrock sent before the pause reach the client. The two probes below were sent at the same instant, one per proxy

Claude Code is the customer's client, so it was driven on both legs too: the interactive TUI (v2.1.280) under tmux, pointed at each proxy with ANTHROPIC_BASE_URL=http://localhost:<port> ANTHROPIC_AUTH_TOKEN=sk-1234 ANTHROPIC_MODEL=bedrock-invoke-sonnet-4-6 ANTHROPIC_SMALL_FAST_MODEL=bedrock-invoke-sonnet-4-6 ANTHROPIC_DEFAULT_HAIKU_MODEL=bedrock-invoke-sonnet-4-6 CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 claude --dangerously-skip-permissions, given /clear and then the prompt "Immediately use the Write tool, with no text before or after the tool call, to create nginx-guide-1.md: a detailed 1500 word guide to configuring nginx as a reverse proxy, with sections, config examples, and troubleshooting.", with the pane captured 12 s, 25 s, and 40 s after the prompt was sent

Before (ccf5e8c)

Large tool input over /v1/messages, streaming

  1. Run the probe against port 50413
  2. Observed: the first event of any kind is the proxy's keepalive at +17.8 s, and message_start reaches the client at +45.5 s, only once Bedrock resumed and the buffer crossed 1024 bytes. 1868 events in total, the largest gaps 15.0 s before a ping, 12.7 s before message_start, and 0.4 s before a content_block_delta
request sent 20:10:50.721
20:11:08.555  +17.849s  event: ping
20:11:23.555  +32.849s  event: ping
20:11:36.234  +45.527s  event: message_start
20:11:36.234  +45.528s  event: content_block_start
20:11:36.234  +45.528s  event: content_block_delta
20:11:36.234  +45.528s  event: content_block_delta
20:11:36.234  +45.528s  event: content_block_delta
20:11:36.234  +45.528s  event: content_block_delta
...
20:11:38.338  +47.632s  event: content_block_stop
20:11:38.339  +47.632s  event: message_delta
20:11:38.339  +47.632s  event: message_stop

Non-streaming /v1/messages parity

  1. curl -s http://localhost:50413/v1/messages -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' -d '{"model":"bedrock-invoke-sonnet-4-6","max_tokens":60,"messages":[{"role":"user","content":"Reply with exactly the word ok."}]}'
  2. Observed: type message, role assistant, stop_reason end_turn, one text block; top-level keys content, id, model, role, stop_details, stop_reason, stop_sequence, type, usage; usage keys cache_creation, cache_creation_input_tokens, cache_read_input_tokens, input_tokens, output_tokens, all counts positive

Text-only streaming parity

  1. curl -sN http://localhost:50413/v1/messages -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' -d '{"model":"bedrock-invoke-sonnet-4-6","max_tokens":120,"stream":true,"messages":[{"role":"user","content":"Write two short sentences about the sea."}]}'
  2. Observed: 20 events, message_start content_block_start content_block_delta x15 content_block_stop message_delta message_stop, the message_delta usage carrying the same five keys with a positive output_tokens

SDK entrypoint, litellm.anthropic.messages.acreate(stream=True)

  1. From the worktree at this commit, await litellm.anthropic.messages.acreate(model="bedrock/invoke/us.anthropic.claude-sonnet-4-6", max_tokens=120, stream=True, messages=[{"role": "user", "content": "Write two short sentences about the sea."}]) and iterate the stream
  2. Observed: first event at 1.744 s, 19 events, message_start content_block_start content_block_delta x14 content_block_stop message_delta message_stop

Claude Code, interactive

  1. Send the prompt to the Claude Code session pointed at port 50413 and capture the pane at 12 s, 25 s, and 40 s
  2. Observed: the spinner sits on Processing… ↓ 108 tokens at 12 s, 25 s, and still at 40 s. No "Waiting for API response" or "check your network" line appeared

pr42658-4a00468f38-lit8266-rc-cc-before-a1-t25.png
pr42658-4a00468f38-lit8266-rc-cc-before-a1-t40.png

After (4a00468)

Large tool input over /v1/messages, streaming

  1. Run the probe against port 42507
  2. Observed: message_start, content_block_start, and the first content_block_delta reach the client at +3.3 s, the moment Bedrock sent them. The keepalives then cover Bedrock's own pause and the deltas resume at +41.2 s. 1768 events in total, the largest gaps 15.0 s and 15.0 s before pings and 8.0 s before a content_block_delta; no event was held behind the pause
request sent 20:10:50.724
20:10:53.966  + 3.259s  event: message_start
20:10:53.967  + 3.261s  event: content_block_start
20:10:53.967  + 3.261s  event: content_block_delta
20:11:08.971  +18.264s  event: ping
20:11:23.972  +33.266s  event: ping
20:11:31.950  +41.243s  event: content_block_delta
20:11:31.950  +41.243s  event: content_block_delta
20:11:31.950  +41.243s  event: content_block_delta
...
20:11:32.704  +41.998s  event: content_block_stop
20:11:32.726  +42.020s  event: message_delta
20:11:32.733  +42.027s  event: message_stop

Non-streaming /v1/messages parity

  1. Same curl against port 42507
  2. Observed: identical to Before: type message, role assistant, stop_reason end_turn, one text block, the same nine top-level keys and the same five usage keys, all counts positive

Text-only streaming parity

  1. Same curl against port 42507
  2. Observed: 21 events, message_start content_block_start content_block_delta x16 content_block_stop message_delta message_stop, the message_delta usage carrying the same five keys with a positive output_tokens

SDK entrypoint, litellm.anthropic.messages.acreate(stream=True)

  1. Same call from the worktree at this commit
  2. Observed: first event at 1.377 s, 21 events, message_start content_block_start content_block_delta x16 content_block_stop message_delta message_stop

Claude Code, interactive

  1. Send the prompt to the Claude Code session pointed at port 42507 and capture the pane at 12 s, 25 s, and 40 s
  2. Observed: the spinner sits on Metamorphosing… ↓ 111 tokens at 12 s, 25 s, and 40 s. The on-screen stall is the same length on both legs (it is Bedrock's own tool-input pause, and the proxy keepalives keep the connection visibly alive on both), so this leg does not separate the two commits; the wire timelines above do

pr42658-4a00468f38-lit8266-rc-cc-after-a1-t25.png
pr42658-4a00468f38-lit8266-rc-cc-after-a1-t40.png

Type

🐛 Bug Fix

Caveats (if any)

Low

  • The remaining silent stretch is Bedrock's: it sends nothing while it generates a large tool input, and this PR only stops holding the events it already sent
  • The Claude Code leg does not separate the commits: the same stall shows on both, and no "check your network" line appeared on this line either
    • The proxy's 15 s keepalive pings reach Claude Code on both legs, so what pushed the customer's client over into the banner is not established by this run
  • The stall only reproduces when the bytes before Bedrock's pause stay under 1024, an immediate tool call with no text first, which is the Write flow the customer hit
  • Regression test hand-rolls one AWS eventstream frame; botocore ships no encoder
  • Mutation check on this line: passing chunk_size=1024 back into aiter_bytes() makes the picked test fail
  • CI at 4a00468 (pipeline 90133) has three reds, none on this diff's path, each rerun once from failed (build_and_test rerun, integration rerun); nothing is ported into this PR for them
    • integration-extensions: test_tool_error_remains_error_and_healthy_sibling_returns_value, the same single failure the base's own run on cc44919 shows (job 2193617), and the rerun (job 2208161) fails it again
    • e2e_ui_testing: the proxy-admin/keyBudget spec is red on main's scheduled runs 90101, 90114, and 90144 too; the mcp/mcpTools spec waits on a live upstream MCP server's tool list; the rerun (job 2208159) fails the same two specs again
    • local_testing_part2: test_openai_stream_options_call_text_completion fails inside the OpenAI text-completion client with pydantic's 'MockValSer' object is not an instance of 'SchemaSerializer'; green on main's 90101 and 90114 runs, which fail other tests; green on the rerun (job 2208158)

Blast radius at 4a00468:

  • Breaking: none. No signature, config key, or response shape changes
  • Backward incompatible: none. Event order, event count, and usage keys match the base on every leg above
  • Regression risk: delivery cadence only. aiter_bytes() now yields per network read, and botocore's EventStreamBuffer already reassembles frames across arbitrary byte boundaries, so partial frames are handled exactly as before. httpx is 0.28.1 on this line, same as main
  • Dependency graph: one call site and one subclass __init__; the class-level DEFAULT_CHUNK_SIZE had no other reader on this line (the DEFAULT_CHUNK_SIZE in litellm/constants.py is the RAG text splitter's, unrelated). Reached by /v1/messages streaming on bedrock/invoke/<claude model> and by litellm.anthropic.messages.acreate(stream=True), both driven above. The mantle path calls the Anthropic base iterator, and Converse (bedrock/<model>) and /v1/chat/completions never used this decoder
  • Not verified: nothing; there is no sync streaming iterator on this path to check

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • 4a00468 passes /live-pr-risk

Note

Low Risk
Scoped to async Bedrock Invoke Anthropic messages streaming delivery timing; event framing and response shapes are unchanged, with regression coverage for small early frames.

Overview
Bedrock Invoke /v1/messages streaming no longer buffers early SSE behind a 1024-byte read chunk, so clients (e.g. Claude Code) get message_start and tool events as soon as Bedrock sends them instead of waiting for the buffer to fill during long upstream pauses.

get_async_streaming_response_iterator now feeds the AWS event-stream decoder from httpx_response.aiter_bytes() with httpx’s default chunking, and the messages-specific AmazonAnthropicClaudeMessagesStreamDecoder override that forced DEFAULT_CHUNK_SIZE = 1024 is removed. A regression test simulates a gated upstream stream and asserts the first message_start SSE is yielded before the second frame arrives.

Reviewed by Cursor Bugbot for commit 4a00468. Bugbot is set up for automated code reviews on this repo. Configure here.

…ding them in a 1024-byte chunker (#42607)

* fix(bedrock): stream /v1/messages Invoke bytes through instead of holding them in a 1024-byte chunker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(bedrock): apply ruff format to invoke messages stream passthrough

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(bedrock): drop drive-by reformat of existing invoke messages tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): collect streamed chunks into a tuple in passthrough regression test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): give the passthrough regression test a 10s first-chunk budget

* test(bedrock): type the eventstream frame helper's payload as Mapping[str, object]

* test(bedrock): take the gated byte stream's chunks as an immutable Sequence

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
(cherry picked from commit 975bd28)
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with the streaming latency fix covered by a focused regression test and no correctness or security issues identified

Summary

This backport removes the 1024-byte HTTPX chunking threshold from Bedrock Invoke Anthropic Messages streaming so complete AWS event-stream frames reach the decoder as soon as the transport yields them

  • Removes the decoder subclass's unused fixed chunk-size override
  • Adds an in-memory regression test that holds the upstream after a sub-1024-byte frame and verifies immediate message_start delivery
  • Confirms the stream resumes and emits message_stop after the upstream gate opens

Reviews (1) · Last reviewed commit: "fix(bedrock): stream /v1/messages Invoke..."

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 4a00468. Configure here.

@codecov

codecov Bot commented Sep 23, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@mateo-berri
mateo-berri merged commit ed9ae4f into rc/1.103.0 Sep 23, 2026
63 of 67 checks passed
@mateo-berri
mateo-berri deleted the litellm_backport_42607_rc_1_103_0 branch September 23, 2026 18:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants