Skip to content

fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking - #42871

Closed
clonylu wants to merge 1 commit into
BerriAI:mainfrom
clonylu:fix/responses-stream-signature-only-reasoning
Closed

clonylu wants to merge 1 commit into
BerriAI:mainfrom
clonylu:fix/responses-stream-signature-only-reasoning

Conversation

@clonylu

@clonylu clonylu commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Streaming /v1/responses on Claude drops signature-only thinking: no reasoning item
  • Default shape for Claude Fable 5.1, Opus 5.5 and Bedrock adaptive

How it solves it:

  • Open the reasoning item on a signed or redacted thinking block
  • Keep signed blocks with empty text through stream assembly

User Flow

Before:

  1. A client sends a streaming /v1/responses request with include: ["reasoning.encrypted_content"] to a Claude deployment on Bedrock, or to Claude Fable 5.1 or Opus 5.5 on any provider
  2. The stream opens only a message item and response.completed output is [message], while reasoning_tokens is non-zero
  3. On the next turn there is no reasoning item to send back, so the model continues without its earlier reasoning

After:

  1. Same request
  2. The stream opens reasoning then message, and response.completed output is [reasoning, message] with encrypted_content set
  3. The client replays the reasoning item and the model keeps its earlier reasoning

Relevant issues

Fixes #42869

Related: #41362 keeps the same signed empty thinking block on /v1/messages requests

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Tests: one regression test in each mapped file, test_streaming_iterator_transformation.py (sync and async) and test_streaming_chunk_builder_utils.py. All three fail on unpatched main and pass here. The rest of tests/test_litellm/responses/litellm_completion_transformation/ and the test_streaming_chunk_builder_* files give the same pass/fail set with and without this change. scripts/check_type_discipline.py counts are unchanged in the two library files

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Shared setup: a litellm --config proxy run from source, with one deployment pointed at a stand-in upstream that streams an Anthropic message with one thinking block and then text, so the thinking block's shape can be set per case

model_list:
  - model_name: anthropic-fable
    litellm_params:
      model: anthropic/claude-fable-5-1
      api_base: http://127.0.0.1:9913
      api_key: fake
curl -sN http://localhost:4000/v1/responses -H "Authorization: Bearer sk-1234" -H "content-type: application/json" \
  -d '{"model":"anthropic-fable","stream":true,"input":"What is 2+2?","max_output_tokens":200,"include":["reasoning.encrypted_content"]}'

Before (6dbd65b)

signature-only thinking block (thinking: "", signature set)

  1. Run the curl above
  2. response.output_item.added events: message only
  3. response.completed output: [message], and the reasoning encrypted_content is missing

thinking block with text (thinking: "Two plus two is four.", signature set)

  1. Run the curl above
  2. response.output_item.added events: reasoning, message
  3. response.completed output: [reasoning, message], and the reasoning encrypted_content is set

After (bc9b6f8)

signature-only thinking block (thinking: "", signature set)

  1. Run the curl above
  2. response.output_item.added events: reasoning, message
  3. response.completed output: [reasoning, message], and the reasoning encrypted_content is set

thinking block with text (thinking: "Two plus two is four.", signature set)

  1. Run the curl above
  2. response.output_item.added events: reasoning, message
  3. response.completed output: [reasoning, message], and the reasoning encrypted_content is set

The stand-in fixes the block shape per case. With real providers, through a proxy on v1.101.0-rc.1 carrying the equivalent change, Claude Fable 5.1 on Vertex AI went from no reasoning item to one the provider accepted on replay in the next turn. Unpatched, real providers show the bug on Bedrock (Claude Opus 4.8, Fable 5.1, Opus 5.5) with or without a reasoning effort, and on Vertex AI whenever no thinking text comes back. The full matrix is in #42869

Type

🐛 Bug Fix

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@clonylu
clonylu requested a review from a team September 24, 2026 03:44
@greptile-apps

greptile-apps Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

The behavioral fix appears sound, but the explicit source-comment and mapped-test placement requirements must be satisfied before merging

Findings

  1. P2 Unnecessary source comments ▶
  2. P2 Incorrect regression test placement ▶

Summary

This PR preserves signature-only thinking blocks during stream assembly and opens a Responses API reasoning item when signed or redacted reasoning arrives without text

  • Signature-only blocks now survive completed-response assembly
  • Streaming reasoning detection now considers signature and redacted-data fields
  • Regression tests verify reasoning-item emission and replayable encrypted content
  • The implementation and test placement need cleanup to satisfy the repository’s explicit contribution guide

Reviews (1) · Last reviewed commit: "fix(responses): emit the reasoning item ..."

def _flush_thinking_block() -> None:
nonlocal current_thinking_text_parts, current_signature
if len(current_thinking_text_parts) > 0 and current_signature:
# A signed block with empty thinking text still carries replayable reasoning; keep it.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Unnecessary source comments This restates the signature check, violating the repository’s source-comment policy. Remove similar comments in streaming_iterator.py before merging.

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@@ -0,0 +1,128 @@
"""Streaming /v1/responses must surface a reasoning item for a SIGNATURE-ONLY thinking block.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Incorrect regression test placement This bug fix creates a separate test file, violating the requirement to extend the existing mapped test file before merging.

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@codecov

codecov Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing clonylu:fix/responses-stream-signature-only-reasoning (bc9b6f8) with main (09ebb28)

Open in CodSpeed

…r signature-only thinking

Anthropic models return thinking blocks with empty text and the reasoning carried in the
signature: Claude Fable 5.1 and Claude Opus 5.5 by default, and Bedrock adaptive thinking
with or without an effort. On streaming /v1/responses the chat->Responses bridge opened a
reasoning output item only on reasoning_content text
(LiteLLMCompletionStreamingIterator._ensure_output_item_for_chunk), and
ChunkProcessor.get_combined_thinking_content kept an assembled thinking block only when it
had thinking text. Such a response emitted no reasoning item mid-stream and none in
response.completed, so a streaming Responses client could not replay the reasoning even
though the reasoning tokens were billed. Non-streaming /v1/responses was unaffected.

Open the reasoning item when the delta carries a signed or redacted thinking block, and
keep a signed block through stream assembly even when its thinking text is empty.
Unsigned text-only fragments are still dropped. The reasoning-text path is unchanged.
@clonylu
clonylu force-pushed the fix/responses-stream-signature-only-reasoning branch from d9118d9 to bc9b6f8 Compare September 24, 2026 07:52

@lets-order-some-fries lets-order-some-fries left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I hit the same behaviour independently while looking at #42869 and had a reproduction at main 571ada0b0f; I re-ran it against this PR's head bc9b6f8a5c (2026-09-25).

Before, on 571ada0b0f, a signature-only thinking stream gave:

"signature_only/async": {"output_item.added": ["message"], "completed.output": ["message"], "encrypted_content": null}

At this PR's head, same script, sync and async:

"signature_only/async": {"output_item.added": ["reasoning", "message"], "completed.output": ["reasoning", "message"],
 "encrypted_content": "[{\"type\":\"thinking\",\"thinking\":\"\",\"signature\":\"SIG_abc123\"}]"}

It looks like a root-cause fix rather than a patch over the symptom: the new _delta_has_signed_thinking_block (litellm/responses/litellm_completion_transformation/streaming_iterator.py:76-79) widens the predicate at :944, and _flush_thinking_block (litellm/litellm_core_utils/streaming_chunk_builder_utils.py:688) no longer requires non-empty thinking text — both sites the bug needed.

Two questions from things I ran:

  1. The latch at streaming_iterator.py:934 means only the first chunk with choices picks the first item. For a plain Anthropic stream that is fine — offline, CustomStreamWrapper drops the empty message_start/content_block_start chunks, so the signature chunk really is first. But for [text, signature-only thinking, text, stop] I get announced output_item.added = [("message", 0)] while response.completed output is [("reasoning", 0), ("message", 1)]; on main both are [("message", 0)], so the index disagreement is new here. Does interleaved thinking need handling too, or is a client rebuilding the response by output_index out of scope?

  2. The redacted branch (b.get("data"), streaming_iterator.py:78) has no test although the body says "signed or redacted". I checked it works — redacted-only at head gives [reasoning, message] with the redacted blob in encrypted_content. Would one more parametrization of the new test be worth it?

Separately, and pre-existing rather than yours: the reasoning-close block exists only in __anext__ (:1019-1064); __next__ (:1084) has no counterpart, so sync signature-only streams get output_item.added(reasoning) with no reasoning_summary_text.done / reasoning_summary_part.done / output_item.done. The new test parametrizes sync_mode but asserts only added_item_types[0] and the completed output, so it passes either way. Same on main for the reasoning_content path.

Tests I ran (own worktree, PYTHONPATH confirmed resolving litellm to the worktree):

  • The two touched test files: 93 passed.
  • tests/test_litellm/responses/ + the builder test file: head 20 failed, 747 passed, main 20 failed, 744 passed, identical failure set (test_responses_websocket_all_providers.py URL tests) — the +3 are this PR's new tests. Four modules deselected for missing fastapi/mcp/websockets locally.
  • tests/test_litellm/litellm_core_utils/: head 109 failed, 3431 passed vs main 40 failed, 3500 passed; all 69 extra are test_tokenizer.py ModuleNotFoundError: No module named 'litellm.rust_bridge._native', i.e. the compiled extension missing from my worktree, not the PR.

I did not test Bedrock's converse streaming shape, and I could not reproduce the end-to-end proxy claim here (no network), so I checked the equivalent path offline through ModelResponseIterator.chunk_parser.

@devin-ai-integration

Copy link
Copy Markdown
Contributor

Thanks for the fix. Tests moved to tests/unit on main, so this is rebased and superseded by #43414 with your commit kept. Closing

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants