Skip to content

fix: recover streamed responses completed output - #30934

Closed
emsi wants to merge 2 commits into
BerriAI:litellm_internal_stagingfrom
emsi:litellm_fix_responses_completed_output_clean
Closed

fix: recover streamed responses completed output#30934
emsi wants to merge 2 commits into
BerriAI:litellm_internal_stagingfrom
emsi:litellm_fix_responses_completed_output_clean

Conversation

@emsi

@emsi emsi commented Jun 21, 2026

Copy link
Copy Markdown
Contributor

Relevant issues

Related to #25429 and #26179

Pre-Submission checklist

  • I have added meaningful tests
  • My PR passes all unit tests on make test-unit
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have requested a Greptile review by commenting @greptileai and received a Confidence Score of 4/5

Screenshots / Proof of Fix

I reproduced the customer-visible issue through a local LiteLLM proxy, then reran the same request after this patch

Proxy command:

uv run --extra proxy litellm --config /tmp/litellm-local-auth/litellm.config.yaml --debug

Request:

curl -N http://localhost:4000/v1/responses \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer test-key' \
  -d '{"model":"gpt-5.5","input":"Reply exactly: OK my lord","stream":true}'

Before the patch, the stream emitted assistant text through response.output_item.done, but the terminal response.completed.response.output was empty

After the patch, the terminal response.completed.response.output includes the recovered assistant output, including OK my lord

Type

Bug Fix
Test

Changes

This updates the Responses streaming iterator to remember completed streamed output items for the lifetime of a request. When the provider sends a terminal response.completed event with an empty response.output, LiteLLM backfills that output from the already streamed response.output_item.done and response.output_text.done events. If the provider sends a non-empty terminal output, LiteLLM preserves it as the authoritative payload

The regression tests cover the empty terminal output recovery path, preservation of authoritative terminal output, and isolation of recovered output item dictionaries from later mutation

Checks run:

uv run --no-sync python scripts/ruff_strict_gate.py --base origin/litellm_internal_staging
uv run --extra proxy pytest tests/test_litellm/llms/chatgpt/responses/test_chatgpt_responses_transformation.py tests/test_litellm/test_responses_streaming_container_ownership.py -q
uv run --extra proxy pytest tests/test_litellm/completion_extras/litellm_responses_transformation/test_completion_extras_litellm_responses_transformation_transformation.py -q
uv run --extra proxy ruff check litellm/responses/streaming_iterator.py tests/test_litellm/llms/chatgpt/responses/test_chatgpt_responses_transformation.py
uv run --extra proxy ruff format --check --line-length 88 litellm/responses/streaming_iterator.py tests/test_litellm/llms/chatgpt/responses/test_chatgpt_responses_transformation.py
git diff --check

@codecov

codecov Bot commented Jun 21, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@emsi
emsi force-pushed the litellm_fix_responses_completed_output_clean branch from 5ca1378 to f585101 Compare June 21, 2026 20:00
@emsi

emsi commented Jun 21, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai

@greptile-apps

greptile-apps Bot commented Jun 21, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR fixes a bug where the terminal response.completed SSE event could arrive with an empty output array even though the assistant text had already been delivered via intermediate response.output_item.done / response.output_text.done events. LiteLLM now remembers those intermediate items and backfills the completed response when the provider omits the output.

  • BaseResponsesAPIStreamingIterator gains two tracking dicts (_streamed_output_items, _streamed_text_only_output_items) that are populated as chunks arrive and cleared on response.created for multi-turn safety; _recovered_streamed_output_items deep-copies on retrieval to prevent downstream mutation from corrupting internal state.
  • Five new mock-only unit tests validate the recovery path, authoritative-output preservation, copy isolation, state reset, and invalid-payload edge cases.

Confidence Score: 5/5

Safe to merge; the backfill is a purely additive no-op when the provider sends a non-empty output, and the reset on response.created guards against state leaking across turns.

The change is well scoped: it only activates when response.completed.response.output is falsy, it is gated by an early-return on every other event type, and it deep-copies the recovered items so later mutations cannot corrupt the stored state. The five new tests cover the core paths (recovery, authoritative preservation, copy isolation, reset, and invalid payloads), and the existing test suite for the transformation layer is unchanged.

No files require special attention.

Important Files Changed

Filename Overview
litellm/responses/streaming_iterator.py Adds _record_streamed_output_chunk and _backfill_completed_response_output to BaseResponsesAPIStreamingIterator to recover empty response.completed output from previously streamed output_item.done/output_text.done events; state is reset on response.created for multi-turn safety; deep-copy isolation is applied on retrieval.
tests/test_litellm/llms/chatgpt/responses/test_chatgpt_responses_transformation.py Adds five focused unit tests covering: recovery from output_item.done, preservation of authoritative output, deep-copy isolation, response.created reset, and invalid-payload guard clauses; all tests use mocks and no real network calls.

Reviews (3): Last reviewed commit: "test: isolate recovered responses output" | Re-trigger Greptile

Comment thread litellm/responses/streaming_iterator.py Outdated
@emsi

emsi commented Jun 21, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai

@Sameerlite

Copy link
Copy Markdown
Contributor

I can see it coming:

data: {"type":"response.content_part.done","item_id":"msg_0c7ddb558aa467bd006a392a3e1d908197bb795b7929636a31","output_index":1,"content_index":0,"part":{"type":"output_text","text":"OK my lord","annotations":[],"logprobs":[]},"sequence_number":10,"model":"gpt-5.5"}

data: {"type":"response.output_item.done","output_index":1,"sequence_number":11,"item":{"id":"msg_0c7ddb558aa467bd006a392a3e1d908197bb795b7929636a31","type":"message","status":"completed","content":[{"type":"output_text","annotations":[],"logprobs":[],"text":"OK my lord"}],"phase":"final_answer","role":"assistant"},"model":"gpt-5.5"}

data: {"type":"response.completed","response":{"id":"resp_2G960yVbeBkrfqoqAooAtnOVC-YbYks4IfGoSl4z7rjLPsceG_sNN8A7-TYdWoB4RSv0hzntd01vncNJ7shYYoHrR1rmzEZl21oV4nOam6VUIJtE0K4A9ROZZm2aS3IiHPgMM-0jl6p9M3NWPmgCcwQ6q3L9PKoK81j60g8vU1B3ayfeMtlbmZJOup_Tpd38_oqd971jo-RN3WqVGJ6CQFvZsVAW4NeDS0qpyncGnJYfoIqiqCcoijKeBeiEypAX_ECo2nXBB7mG1tHjiFy-r1LS3xuLDTcBIlc8sjmBANrURmcvf7S0O1xfGD8MFjxAF2FpifdKNFktz4FO0HoD82a5kq1RNdEBXOSk5jQZFXS-bWUD_LpuhatKE245MsiK-5yk7VHekUGvki3izXnz-WTPW6uU8im9X_XRi8MfoxTwBg3F_OcxL8OqZw2lnzgPnp3rp0-9fMT2LuAOviXuFXUZ","created_at":1782131260,"metadata":{},"model":"gpt-5.5-2026-04-23","object":"response","output":[{"id":"rs_0c7ddb558aa467bd006a392a3ddd408197aacec0eef39af1c2","summary":[],"type":"reasoning","content":[]},{"id":"msg_0c7ddb558aa467bd006a392a3e1d908197bb795b7929636a31","content":[{"annotations":[],"text":"OK my lord","type":"output_text","logprobs":[]}],"role":"assistant","status":"completed","type":"message","phase":"final_answer"}],"parallel_tool_calls":true,"temperature":1.0,"tool_choice":"auto","tools":[],"top_p":0.98,"reasoning":{"context":"current_turn","effort":"medium","summary":null},"status":"completed","text":{"format":{"type":"text"},"verbosity":"medium"},"truncation":"disabled","usage":{"input_tokens":12,"input_tokens_details":{"cached_tokens":0},"output_tokens":18,"output_tokens_details":{"reasoning_tokens":9},"total_tokens":30},"store":true,"background":false,"completed_at":1782131262,"frequency_penalty":0.0,"presence_penalty":0.0,"prompt_cache_retention":"24h","service_tier":"default","top_logprobs":0},"sequence_number":12,"model":"gpt-5.5"}

Closing this as it doesn't have enough evidence to show what it is fixing. Please create a issue if it is not related to the issues you mentioned and then create a PR

@Sameerlite Sameerlite closed this Jun 22, 2026
@emsi

emsi commented Jun 22, 2026

Copy link
Copy Markdown
Contributor Author

I can see it coming:

data: {"type":"response.content_part.done","item_id":"msg_0c7ddb558aa467bd006a392a3e1d908197bb795b7929636a31","output_index":1,"content_index":0,"part":{"type":"output_text","text":"OK my lord","annotations":[],"logprobs":[]},"sequence_number":10,"model":"gpt-5.5"}

data: {"type":"response.output_item.done","output_index":1,"sequence_number":11,"item":{"id":"msg_0c7ddb558aa467bd006a392a3e1d908197bb795b7929636a31","type":"message","status":"completed","content":[{"type":"output_text","annotations":[],"logprobs":[],"text":"OK my lord"}],"phase":"final_answer","role":"assistant"},"model":"gpt-5.5"}

data: {"type":"response.completed","response":{"id":"resp_2G960yVbeBkrfqoqAooAtnOVC-YbYks4IfGoSl4z7rjLPsceG_sNN8A7-TYdWoB4RSv0hzntd01vncNJ7shYYoHrR1rmzEZl21oV4nOam6VUIJtE0K4A9ROZZm2aS3IiHPgMM-0jl6p9M3NWPmgCcwQ6q3L9PKoK81j60g8vU1B3ayfeMtlbmZJOup_Tpd38_oqd971jo-RN3WqVGJ6CQFvZsVAW4NeDS0qpyncGnJYfoIqiqCcoijKeBeiEypAX_ECo2nXBB7mG1tHjiFy-r1LS3xuLDTcBIlc8sjmBANrURmcvf7S0O1xfGD8MFjxAF2FpifdKNFktz4FO0HoD82a5kq1RNdEBXOSk5jQZFXS-bWUD_LpuhatKE245MsiK-5yk7VHekUGvki3izXnz-WTPW6uU8im9X_XRi8MfoxTwBg3F_OcxL8OqZw2lnzgPnp3rp0-9fMT2LuAOviXuFXUZ","created_at":1782131260,"metadata":{},"model":"gpt-5.5-2026-04-23","object":"response","output":[{"id":"rs_0c7ddb558aa467bd006a392a3ddd408197aacec0eef39af1c2","summary":[],"type":"reasoning","content":[]},{"id":"msg_0c7ddb558aa467bd006a392a3e1d908197bb795b7929636a31","content":[{"annotations":[],"text":"OK my lord","type":"output_text","logprobs":[]}],"role":"assistant","status":"completed","type":"message","phase":"final_answer"}],"parallel_tool_calls":true,"temperature":1.0,"tool_choice":"auto","tools":[],"top_p":0.98,"reasoning":{"context":"current_turn","effort":"medium","summary":null},"status":"completed","text":{"format":{"type":"text"},"verbosity":"medium"},"truncation":"disabled","usage":{"input_tokens":12,"input_tokens_details":{"cached_tokens":0},"output_tokens":18,"output_tokens_details":{"reasoning_tokens":9},"total_tokens":30},"store":true,"background":false,"completed_at":1782131262,"frequency_penalty":0.0,"presence_penalty":0.0,"prompt_cache_retention":"24h","service_tier":"default","top_logprobs":0},"sequence_number":12,"model":"gpt-5.5"}

Closing this as it doesn't have enough evidence to show what it is fixing. Please create a issue if it is not related to the issues you mentioned and then create a PR

You should have read #25429, there are specific conditions in which the problem happens. The fix is here for everyone to use.

crognlie added a commit to crognlie/litellm that referenced this pull request Jun 25, 2026
…one events

chatgpt.com's Codex backend streams assistant content via
response.output_item.done and sends a terminal response.completed with
an empty output array. BaseResponsesAPIStreamingIterator._process_chunk
had no accumulation logic, so completed_response.response.output was
always [] and the chat-completions bridge raised "Unknown items in
responses API response: []".

Accumulate output_item.done payloads by index in _streamed_output_items
and backfill them into response.completed before storing
completed_response, only when the provider sends empty output.
Non-empty authoritative output is left untouched so other providers
are unaffected. A secondary _streamed_text_only_items dict handles
providers that emit output_text.done without a preceding
output_item.done. Items are serialized to plain dicts via model_dump()
so the downstream _handle_raw_dict_response_item callback in the
transformation layer can process them.

Seven regression tests in
tests/test_litellm/responses/test_streaming_iterator_output_recovery.py
cover: empty-output baseline, core backfill, authoritative output
preservation, multi-item index ordering, output_text.done fallback,
output_item.done precedence, response.incomplete backfill, and the
dict-type contract required by the transformation layer.

Fixes BerriAI#25429. Supersedes BerriAI#30934
crognlie added a commit to crognlie/litellm that referenced this pull request Jul 15, 2026
…one events

chatgpt.com's Codex backend streams assistant content via
response.output_item.done and sends a terminal response.completed with
an empty output array. BaseResponsesAPIStreamingIterator._process_chunk
had no accumulation logic, so completed_response.response.output was
always [] and the chat-completions bridge raised "Unknown items in
responses API response: []".

Accumulate output_item.done payloads by index in _streamed_output_items
and backfill them into response.completed before storing
completed_response, only when the provider sends empty output.
Non-empty authoritative output is left untouched so other providers
are unaffected. A secondary _streamed_text_only_items dict handles
providers that emit output_text.done without a preceding
output_item.done. Items are serialized to plain dicts via model_dump()
so the downstream _handle_raw_dict_response_item callback in the
transformation layer can process them.

Seven regression tests in
tests/test_litellm/responses/test_streaming_iterator_output_recovery.py
cover: empty-output baseline, core backfill, authoritative output
preservation, multi-item index ordering, output_text.done fallback,
output_item.done precedence, response.incomplete backfill, and the
dict-type contract required by the transformation layer.

Fixes BerriAI#25429. Supersedes BerriAI#30934
crognlie added a commit to crognlie/litellm that referenced this pull request Jul 22, 2026
…one events

chatgpt.com's Codex backend streams assistant content via
response.output_item.done and sends a terminal response.completed with
an empty output array. BaseResponsesAPIStreamingIterator._process_chunk
had no accumulation logic, so completed_response.response.output was
always [] and the chat-completions bridge raised "Unknown items in
responses API response: []".

Accumulate output_item.done payloads by index in _streamed_output_items
and backfill them into response.completed before storing
completed_response, only when the provider sends empty output.
Non-empty authoritative output is left untouched so other providers
are unaffected. A secondary _streamed_text_only_items dict handles
providers that emit output_text.done without a preceding
output_item.done. Items are serialized to plain dicts via model_dump()
so the downstream _handle_raw_dict_response_item callback in the
transformation layer can process them.

Seven regression tests in
tests/test_litellm/responses/test_streaming_iterator_output_recovery.py
cover: empty-output baseline, core backfill, authoritative output
preservation, multi-item index ordering, output_text.done fallback,
output_item.done precedence, response.incomplete backfill, and the
dict-type contract required by the transformation layer.

Fixes BerriAI#25429. Supersedes BerriAI#30934
crognlie added a commit to crognlie/litellm that referenced this pull request Jul 30, 2026
…one events

chatgpt.com's Codex backend streams assistant content via
response.output_item.done and sends a terminal response.completed with
an empty output array. BaseResponsesAPIStreamingIterator._process_chunk
had no accumulation logic, so completed_response.response.output was
always [] and the chat-completions bridge raised "Unknown items in
responses API response: []".

Accumulate output_item.done payloads by index in _streamed_output_items
and backfill them into response.completed before storing
completed_response, only when the provider sends empty output.
Non-empty authoritative output is left untouched so other providers
are unaffected. A secondary _streamed_text_only_items dict handles
providers that emit output_text.done without a preceding
output_item.done. Items are serialized to plain dicts via model_dump()
so the downstream _handle_raw_dict_response_item callback in the
transformation layer can process them.

Seven regression tests in
tests/test_litellm/responses/test_streaming_iterator_output_recovery.py
cover: empty-output baseline, core backfill, authoritative output
preservation, multi-item index ordering, output_text.done fallback,
output_item.done precedence, response.incomplete backfill, and the
dict-type contract required by the transformation layer.

Fixes BerriAI#25429. Supersedes BerriAI#30934
crognlie added a commit to crognlie/litellm that referenced this pull request Jul 31, 2026
…one events

chatgpt.com's Codex backend streams assistant content via
response.output_item.done and sends a terminal response.completed with
an empty output array. BaseResponsesAPIStreamingIterator._process_chunk
had no accumulation logic, so completed_response.response.output was
always [] and the chat-completions bridge raised "Unknown items in
responses API response: []".

Accumulate output_item.done payloads by index in _streamed_output_items
and backfill them into response.completed before storing
completed_response, only when the provider sends empty output.
Non-empty authoritative output is left untouched so other providers
are unaffected. A secondary _streamed_text_only_items dict handles
providers that emit output_text.done without a preceding
output_item.done. Items are serialized to plain dicts via model_dump()
so the downstream _handle_raw_dict_response_item callback in the
transformation layer can process them.

Seven regression tests in
tests/test_litellm/responses/test_streaming_iterator_output_recovery.py
cover: empty-output baseline, core backfill, authoritative output
preservation, multi-item index ordering, output_text.done fallback,
output_item.done precedence, response.incomplete backfill, and the
dict-type contract required by the transformation layer.

Fixes BerriAI#25429. Supersedes BerriAI#30934
crognlie added a commit to crognlie/litellm that referenced this pull request Aug 5, 2026
…one events

chatgpt.com's Codex backend streams assistant content via
response.output_item.done and sends a terminal response.completed with
an empty output array. BaseResponsesAPIStreamingIterator._process_chunk
had no accumulation logic, so completed_response.response.output was
always [] and the chat-completions bridge raised "Unknown items in
responses API response: []".

Accumulate output_item.done payloads by index in _streamed_output_items
and backfill them into response.completed before storing
completed_response, only when the provider sends empty output.
Non-empty authoritative output is left untouched so other providers
are unaffected. A secondary _streamed_text_only_items dict handles
providers that emit output_text.done without a preceding
output_item.done. Items are serialized to plain dicts via model_dump()
so the downstream _handle_raw_dict_response_item callback in the
transformation layer can process them.

Seven regression tests in
tests/test_litellm/responses/test_streaming_iterator_output_recovery.py
cover: empty-output baseline, core backfill, authoritative output
preservation, multi-item index ordering, output_text.done fallback,
output_item.done precedence, response.incomplete backfill, and the
dict-type contract required by the transformation layer.

Fixes BerriAI#25429. Supersedes BerriAI#30934
crognlie added a commit to crognlie/litellm that referenced this pull request Aug 15, 2026
…one events

chatgpt.com's Codex backend streams assistant content via
response.output_item.done and sends a terminal response.completed with
an empty output array. BaseResponsesAPIStreamingIterator._process_chunk
had no accumulation logic, so completed_response.response.output was
always [] and the chat-completions bridge raised "Unknown items in
responses API response: []".

Accumulate output_item.done payloads by index in _streamed_output_items
and backfill them into response.completed before storing
completed_response, only when the provider sends empty output.
Non-empty authoritative output is left untouched so other providers
are unaffected. A secondary _streamed_text_only_items dict handles
providers that emit output_text.done without a preceding
output_item.done. Items are serialized to plain dicts via model_dump()
so the downstream _handle_raw_dict_response_item callback in the
transformation layer can process them.

Seven regression tests in
tests/test_litellm/responses/test_streaming_iterator_output_recovery.py
cover: empty-output baseline, core backfill, authoritative output
preservation, multi-item index ordering, output_text.done fallback,
output_item.done precedence, response.incomplete backfill, and the
dict-type contract required by the transformation layer.

Fixes BerriAI#25429. Supersedes BerriAI#30934
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants