Skip to content

fix(responses): emit typed streaming failure events - #40243

Merged
mateo-berri merged 8 commits into
BerriAI:mainfrom
zoroyihan7:fix-responses-stream-error-events
Sep 19, 2026
Merged

mateo-berri merged 8 commits into
BerriAI:mainfrom
zoroyihan7:fix-responses-stream-error-events

Conversation

@zoroyihan7

@zoroyihan7 zoroyihan7 commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • A mid-stream provider error on /v1/responses arrived as a Chat-style data: {"error": ...} frame
  • No event: response.failed, no sequence_number, no data: [DONE], so clients saw a dropped connection
  • The provider's error code and message were replaced by "code": "500" and a litellm.InternalServerError: prefix

How it solves it:

  • Emit event: response.failed with the response id, the next sequence_number, and the provider's code and message
  • End the failed stream with data: [DONE], the same close a completed stream gets
  • Background polling reads the data: line of that named frame, so GET /v1/responses/{id} on a background: true request reports failed with the provider's code and message instead of completed with no error
  • Chat Completions and Cursor streams keep their existing error frames

User Flow

Before: a developer running Codex CLI through the gateway sees "stream disconnected before completion" with no reason when the provider rejects a request mid-stream

  1. They point Codex CLI at https://litellm-domain/v1 with wire_api = "responses" and send a prompt
  2. Codex sends POST https://litellm-domain/v1/responses with "stream": true
  3. The proxy answers HTTP 200 and streams response.created, response.in_progress, and the first output events
  4. The provider rejects the request mid-stream; the stream ends with data: {"error": {"message": "litellm.InternalServerError: InternalServerError: OpenAIException - Your request was flagged for possible cybersecurity risk and was not completed", "type": null, "param": null, "code": "500"}} and closes with no response.failed event and no data: [DONE]
  5. Codex shows stream disconnected before completion: stream closed before response.completed, and the real reason is only visible in the proxy's pod logs

After: the same rejection arrives as a response.failed event and Codex shows the provider's reason

  1. They point Codex CLI at https://litellm-domain/v1 with wire_api = "responses" and send a prompt
  2. Codex sends POST https://litellm-domain/v1/responses with "stream": true
  3. The proxy answers HTTP 200 and streams response.created, response.in_progress, and the first output events
  4. The provider rejects the request mid-stream; the stream ends with event: response.failed whose data carries the same response id, the next sequence_number, and "error": {"code": "cyber_policy", "message": "Your request was flagged for possible cybersecurity risk and was not completed"}, then data: [DONE]
  5. Codex shows the provider's block notice ("This content can't be shown") instead of a disconnect

Relevant issues

Supersedes #39748, which kept the Chat-style error frame on Responses streams and only appended [DONE]

Affected release

Linear ticket

Resolves LIT-5620

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.). The one red job, proxy-infra / Run tests, is not required and fails test_login_throttle_settings_are_not_hot_applied_from_the_database on every recent main run too; this PR does not touch login throttling
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Real gpt-6-astra calls through two DB-less proxies, each started from a checkout of the named commit on a free random port with --num_workers 2 (2 uvicorn workers, one proxy process per port; the changed surface keeps no per-process state, so there is no second instance to alternate across), with the same OPENAI_API_KEY. Proxy A (port 30790) fronts OpenAI and carries a callback that lets the first text delta through and then ends the stream with the in-band rejection OpenAI sends for a cybersecurity block (HTTP 200 first, then an error whose code is cyber_policy); that block cannot be triggered on demand, so the callback replays its shape. Proxy B is the proxy under test (port 32466 for Before, 56341 for After): gpt-6-astra goes straight to OpenAI and gpt-6-astra-chained goes through proxy A, with a callback that rejects any stream whose text contains "banana", no retries, and no fallbacks

Proxy A, config_a.yaml and hook_cyber.py:

model_list:
  - model_name: gpt-6-astra
    litellm_params:
      model: openai/gpt-6-astra
      api_key: os.environ/OPENAI_API_KEY
      num_retries: 0
      max_retries: 0
router_settings:
  num_retries: 0
  fallbacks: []
litellm_settings:
  callbacks: [hook_cyber.policy]
from typing import Final

import litellm
from litellm.integrations.custom_logger import CustomLogger

_MESSAGE: Final = "Your request was flagged for possible cybersecurity risk and was not completed"
_BODY: Final = {"code": "cyber_policy", "message": _MESSAGE, "type": None}


class UpstreamCyberPolicy(CustomLogger):
    async def async_post_call_streaming_iterator_hook(self, user_api_key_dict, response, request_data):
        async for chunk in response:
            yield chunk
            if getattr(chunk, "type", None) == "response.output_text.delta":
                raise litellm.InternalServerError(message=_MESSAGE, llm_provider="openai", model="gpt-6-astra", body=_BODY)


policy = UpstreamCyberPolicy()

Proxy B, config_b.yaml and hook_banana.py:

model_list:
  - model_name: gpt-6-astra
    litellm_params:
      model: openai/gpt-6-astra
      api_key: os.environ/OPENAI_API_KEY
      num_retries: 0
      max_retries: 0
  - model_name: gpt-6-astra-chained
    litellm_params:
      model: openai/gpt-6-astra
      api_base: http://127.0.0.1:30790/v1
      api_key: sk-upstream-stand-in
      num_retries: 0
      max_retries: 0
router_settings:
  num_retries: 0
  fallbacks: []
general_settings:
  coordination_redis:
    host: 127.0.0.1
    port: 48685
litellm_settings:
  callbacks: [hook_banana.policy]
  responses:
    background_mode:
      polling_via_cache: "all"

The Redis on port 48685 is a plain redis-server started for this run; it serves proxy B's background: true polling state, which the polling legs read back through GET /v1/responses/{id}

from fastapi import HTTPException
from litellm.integrations.custom_logger import CustomLogger


class OutputPolicy(CustomLogger):
    async def async_post_call_streaming_iterator_hook(self, user_api_key_dict, response, request_data):
        text = ""
        async for chunk in response:
            if getattr(chunk, "type", None) == "response.output_text.delta":
                text += chunk.delta
                if "banana" in text.lower():
                    raise HTTPException(status_code=400, detail="Output contains blocked keyword: banana")
            yield chunk


policy = OutputPolicy()

The curl legs use leg <port> <model> <prompt> <out>, which saves the raw stream and prints the status, the event: lines, the frame counts, and the last two data frames:

leg() {
  curl --silent --show-error --no-buffer --max-time 90 -D "$4.headers" -o "$4.sse" -w "HTTP_STATUS:%{http_code}\n" \
    "http://127.0.0.1:$1/v1/responses" -H 'Content-Type: application/json' \
    --data-binary "{\"model\":\"$2\",\"input\":\"$3\",\"stream\":true,\"max_output_tokens\":256,\"reasoning\":{\"effort\":\"low\"}}"
  echo "curl exit: $?"
  echo "events: $(grep '^event:' "$4.sse" | sed 's/^event: //' | tr '\n' ' ')"
  echo "data lines: $(grep -c '^data:' "$4.sse"), DONE lines: $(grep -c 'data: \[DONE\]' "$4.sse")"
  grep '^data:' "$4.sse" | tail -2 | cut -c1-400
}

The polling legs use poll <port> <model> <prompt> <out>, which creates the response with "background": true, prints the id and status it got back, then polls GET /v1/responses/{id} every 2 seconds until the status is terminal and prints that state's id, status, error, and joined output text:

poll() {
  curl --silent --show-error --max-time 60 -o "$4.create.json" -w "HTTP_STATUS:%{http_code}\n" \
    "http://127.0.0.1:$1/v1/responses" -H 'Content-Type: application/json' \
    --data-binary "{\"model\":\"$2\",\"input\":\"$3\",\"background\":true,\"max_output_tokens\":256,\"reasoning\":{\"effort\":\"low\"}}"
  id=$(jq -r .id "$4.create.json")
  echo "create: $(jq -c '{id, status, error}' "$4.create.json")"
  for i in $(seq 1 30); do
    sleep 2
    curl --silent --max-time 20 -o "$4.poll.json" "http://127.0.0.1:$1/v1/responses/$id"
    case "$(jq -r .status "$4.poll.json")" in completed|failed|incomplete|cancelled) echo "terminal after $((i*2))s"; break;; esac
  done
  echo "GET: $(jq -c '{id, status, error, output_text: ([.output[]?.content[]?.text?] | join(""))}' "$4.poll.json")"
}

The Codex CLI legs run Codex CLI 0.155.1 in a fresh CODEX_HOME with this config.toml (the port is proxy B's):

model = "gpt-6-astra"
model_provider = "litellm"
model_reasoning_effort = "low"
model_max_output_tokens = 256

[model_providers.litellm]
name = "Local LiteLLM QA"
base_url = "http://127.0.0.1:<proxy B port>/v1"
env_key = "LITELLM_QA_KEY"
wire_api = "responses"
requires_openai_auth = false
supports_websockets = false
request_max_retries = 0
stream_max_retries = 0

Before (6e750e3)

Output policy rejects a generated word

  1. leg 32466 gpt-6-astra "Reply with exactly banana. Do not use tools." banana_before
  2. Observed:
curl: (18) transfer closed with outstanding read data remaining
HTTP_STATUS:200
curl exit: 18
events:
data lines: 4, DONE lines: 0
data: {"type": "response.output_item.added", "output_index": 0, "item": {"id": "msg_0b60fbbeef174f3f006aadbe44ebbc87d0b129805b771bda0d", "type": "message", "status": "in_progress", "content": [], "phase": "final_answer", "role": "assistant"}, "sequence_number": 2, "model": "gpt-6-astra"}
data: {"type": "response.content_part.added", "item_id": "msg_0b60fbbeef174f3f006aadbe44ebbc87d0b129805b771bda0d", "output_index": 0, "content_index": 0, "part": {"type": "output_text", "annotations": [], "logprobs": [], "text": ""}, "sequence_number": 3, "model": "gpt-6-astra"}
  1. The connection is cut after response.content_part.added: no failure event, no [DONE], and curl reports the transfer closed with data outstanding

Upstream provider rejects mid-stream

  1. leg 32466 gpt-6-astra-chained "Say hello in one word." cyber_before
  2. Observed:
HTTP_STATUS:200
curl exit: 0
events:
data lines: 6, DONE lines: 0
data: {"type": "response.output_text.delta", "item_id": "msg_099bc2c47110a8a5006aadbe56319087d09699f8ea7965d1a7", "output_index": 0, "content_index": 0, "delta": "Hello", "logprobs": [], "obfuscation": "WaONNEKyNsS", "sequence_number": 4, "model": "gpt-6-astra-chained"}
data: {"error": {"message": "litellm.InternalServerError: InternalServerError: OpenAIException - Your request was flagged for possible cybersecurity risk and was not completed", "type": null, "param": null, "code": "500"}}
  1. The stream ends on a Chat-style error frame: no event: line, no sequence_number, the provider's cyber_policy code replaced by "500", and no [DONE]

Codex CLI 0.155.1

  1. With base_url pointing at port 32466, run codex interactively in a tmux pane and send Reply with exactly banana. Do not use tools.
  2. Codex shows stream disconnected before completion: Transport error: network error: error decoding response body

pr40243-c181c927d0-codex_before.png

  1. With model = "gpt-6-astra-chained", send Say hello in one word.
  2. Codex prints Hello and then stream disconnected before completion: stream closed before response.completed

pr40243-c181c927d0-codex_cyber-before.png

Completed stream unchanged

  1. leg 32466 gpt-6-astra "Say hello in one word." happy_before
  2. Observed:
HTTP_STATUS:200
curl exit: 0
events:
data lines: 10, DONE lines: 1
data: {"type": "response.completed", "sequence_number": 8, "response": {"id": "resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3ZpZGVyOm9wZW5haTttb2RlbF9pZDozMDhmOTA4OWJmMGQ1ZTMyNmI0NDdiZmU4MjQ0YzhmM2U0NGIxNzA0OThhMzVmMTAzYjI2Y2RlMDA0OWE5MDQ3O3Jlc3BvbnNlX2lkOnJlc3BfMDc2OTkxYjYzMzJlMWZiMTAwNmFhZGJlNWM4MmRjODdkMGEwNWUzN2E2Mzc4Mzg2Yzc=", "status": "completed", "error": null, ...}}
data: [DONE]

Background polling

  1. poll 32466 gpt-6-astra "Reply with exactly banana. Do not use tools." poll_banana_before
  2. Observed:
HTTP_STATUS:200
create: {"id":"litellm_poll_14d73481-f69b-48a6-85ca-3b30a811db1b","status":"queued","error":null}
terminal after 2s
GET: {"id":"litellm_poll_14d73481-f69b-48a6-85ca-3b30a811db1b","status":"failed","error":{"type":"internal_error","message":"400: Output contains blocked keyword: banana","code":"background_streaming_error"},"output_text":""}
  1. poll 32466 gpt-6-astra-chained "Say hello in one word." poll_cyber_before
  2. Observed:
HTTP_STATUS:200
create: {"id":"litellm_poll_9ccf2c6f-e340-4672-8510-3a629ec3ab1f","status":"queued","error":null}
terminal after 2s
GET: {"id":"litellm_poll_9ccf2c6f-e340-4672-8510-3a629ec3ab1f","status":"completed","error":null,"output_text":"Hello"}
  1. The policy rejection surfaces only through the background task's own exception handler, with the generic background_streaming_error code; the upstream rejection is polled back as completed with no error at all, because the Chat-style error frame carries no type and the stream ends without a terminal event
  2. poll 32466 gpt-6-astra "Say hello in one word." poll_happy_before
  3. Observed:
HTTP_STATUS:200
create: {"id":"litellm_poll_5dd9d52f-e016-4a83-8b97-b723e4981e6a","status":"queued","error":null}
terminal after 2s
GET: {"id":"litellm_poll_5dd9d52f-e016-4a83-8b97-b723e4981e6a","status":"completed","error":null,"output_text":"Hello!"}

After (c181c92)

Output policy rejects a generated word

  1. leg 56341 gpt-6-astra "Reply with exactly banana. Do not use tools." banana_after
  2. Observed:
HTTP_STATUS:200
curl exit: 0
events: response.failed
data lines: 6, DONE lines: 1
data: {"type": "response.failed", "sequence_number": 4, "response": {"id": "resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3ZpZGVyOm9wZW5haTttb2RlbF9pZDozMDhmOTA4OWJmMGQ1ZTMyNmI0NDdiZmU4MjQ0YzhmM2U0NGIxNzA0OThhMzVmMTAzYjI2Y2RlMDA0OWE5MDQ3O3Jlc3BvbnNlX2lkOnJlc3BfMDU3YTM0ZTUzYWVlMmY4NjAwNmFhZGMyNTViZWM4ODdkMDk4YjAxZmNkMzc0M2E1OGQ=", "status": "failed", "error": {"code": "invalid_request_error", "message": "400: Output contains blocked keyword: banana"}, ...}}
data: [DONE]
  1. The stream ends with event: response.failed carrying the same response id, sequence_number 4 right after the last streamed event, the policy's reason, and then data: [DONE]; curl exits 0

Upstream provider rejects mid-stream

  1. leg 56341 gpt-6-astra-chained "Say hello in one word." cyber_after
  2. Observed:
HTTP_STATUS:200
curl exit: 0
events: response.failed
data lines: 7, DONE lines: 1
data: {"type": "response.failed", "sequence_number": 5, "response": {"id": "resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3ZpZGVyOm9wZW5haTttb2RlbF9pZDozMDhmOTA4OWJmMGQ1ZTMyNmI0NDdiZmU4MjQ0YzhmM2U0NGIxNzA0OThhMzVmMTAzYjI2Y2RlMDA0OWE5MDQ3O3Jlc3BvbnNlX2lkOnJlc3BfMDg2MzU4MWU1YThmNGFiNTAwNmFhZGMyNWQwMzVjODdkMDk0NTIzY2E5ZWI1NzZjMDQ=", "status": "failed", "error": {"code": "cyber_policy", "message": "Your request was flagged for possible cybersecurity risk and was not completed"}, ...}}
data: [DONE]
  1. The provider's own cyber_policy code and message reach the client unchanged, after the Hello delta, followed by data: [DONE]

Codex CLI 0.155.1

  1. With base_url pointing at port 56341, run codex interactively in a tmux pane and send Reply with exactly banana. Do not use tools.
  2. Codex shows stream disconnected before completion: 400: Output contains blocked keyword: banana, the policy's reason instead of a transport error

pr40243-c181c927d0-codex_after.png

  1. With model = "gpt-6-astra-chained", send Say hello in one word.
  2. Codex shows OpenAI's own block notice, This content can't be shown / We take extra care with some cybersecurity requests, instead of a disconnect

pr40243-c181c927d0-codex_cyber-after.png

Completed stream unchanged

  1. leg 56341 gpt-6-astra "Say hello in one word." happy_after
  2. Observed:
HTTP_STATUS:200
curl exit: 0
events:
data lines: 11, DONE lines: 1
data: {"type": "response.completed", "sequence_number": 9, "response": {"id": "resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3ZpZGVyOm9wZW5haTttb2RlbF9pZDozMDhmOTA4OWJmMGQ1ZTMyNmI0NDdiZmU4MjQ0YzhmM2U0NGIxNzA0OThhMzVmMTAzYjI2Y2RlMDA0OWE5MDQ3O3Jlc3BvbnNlX2lkOnJlc3BfMDBiODMzZDcxOTQ3MGY0MTAwNmFhZGMyNjFmZTgwODdkMGE5MDgxNDc2OTBkMGU5MTM=", "status": "completed", "error": null, ...}}
data: [DONE]
  1. Same event sequence ending in response.completed and one [DONE]; this run streamed one more response.output_text.delta than the Before run (11 data lines against 10), which is OpenAI's chunking, not the proxy

Background polling

  1. poll 56341 gpt-6-astra "Reply with exactly banana. Do not use tools." poll_banana_after
  2. Observed:
HTTP_STATUS:200
create: {"id":"litellm_poll_cd007380-faba-4712-985b-4e62ae66db45","status":"queued","error":null}
terminal after 2s
GET: {"id":"litellm_poll_cd007380-faba-4712-985b-4e62ae66db45","status":"failed","error":{"code":"invalid_request_error","message":"400: Output contains blocked keyword: banana"},"output_text":""}
  1. poll 56341 gpt-6-astra-chained "Say hello in one word." poll_cyber_after
  2. Observed:
HTTP_STATUS:200
create: {"id":"litellm_poll_2e26d579-b008-47d6-9713-9441ae0d994a","status":"queued","error":null}
terminal after 2s
GET: {"id":"litellm_poll_2e26d579-b008-47d6-9713-9441ae0d994a","status":"failed","error":{"code":"cyber_policy","message":"Your request was flagged for possible cybersecurity risk and was not completed"},"output_text":"Hello"}
  1. Both rejections now poll back as failed with the policy's or the provider's own code and message, read from the response.failed frame the background task recorded; the Hello the provider streamed before rejecting stays on the response
  2. poll 56341 gpt-6-astra "Say hello in one word." poll_happy_after
  3. Observed:
HTTP_STATUS:200
create: {"id":"litellm_poll_a00e30a3-5e89-4898-9c2d-e0731e42cc3a","status":"queued","error":null}
terminal after 2s
GET: {"id":"litellm_poll_a00e30a3-5e89-4898-9c2d-e0731e42cc3a","status":"completed","error":null,"output_text":"Hello!"}

Type

🐛 Bug Fix

Caveats (if any)

Low

  • A failure after response.completed already went out gets only data: [DONE], never a second terminal event
  • Failed Responses streams now end with data: [DONE] like the proxy's completed streams; OpenAI itself sends none
  • A request the provider rejected still fails; it now says why instead of dropping the connection
  • The provider rejection in the proof is replayed by a callback on an upstream proxy, since OpenAI's block cannot be triggered on demand
  • The body relay that gives response.failed the provider's code and message also reaches error responses on the other endpoints: an InternalServerError or APIError mapped from an OpenAI-compatible upstream now reports that upstream body's type (server_error for an OpenAI 500) where it used to report LiteLLM's status-derived internal_server_error, or null on a Chat Completions stream error frame. Observed on a chained 500 through /v1/chat/completions: Before "type":"internal_server_error", After "type":"server_error", same HTTP status, same "code":"500", same message. A 429 keeps throttling_error
  • The response.failed frame carries output: [] even when deltas already streamed before the failure (the chained leg had printed Hello). Chosen over accumulating output items in the proxy's error state because the client already holds every delta from the same connection and GET /v1/responses/{id} on a background: true request keeps the accumulated text (output_text Hello in the polling leg); a client that reads response.output off the failed event to recover partial text gets an empty list

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • c181c92 passes /live-pr-risk

@zoroyihan7
zoroyihan7 requested a review from a team September 8, 2026 11:25
@CLAassistant

CLAassistant commented Sep 8, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@codspeed

codspeed Bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing zoroyihan7:fix-responses-stream-error-events (c181c92) with main (b0887b6)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (bc6b540) during the generation of this report, so b0887b6 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@greptile-apps

greptile-apps Bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge; the latest polling adjustment correctly records named Responses failure frames and no outstanding defects remain.

Summary

This PR makes mid-stream Responses API failures conform to the Responses SSE protocol while preserving existing Chat Completions behavior.

  • Tracks emitted response identity and sequence state to construct a typed response.failed terminal event.
  • Preserves provider error codes, messages, and bodies through exception mapping.
  • Terminates handled failed streams with [DONE].
  • Updates background polling to recognize named SSE failure frames.
  • Adds coverage for upstream failures, serialization failures, protocol isolation, numeric status codes, and polling state.

Reviews (3) · Last reviewed commit: "fix(proxy): record response.failed frame..."

Comment thread litellm/proxy/proxy_server.py
@codecov

codecov Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.72131% with 4 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...ellm/proxy/common_utils/responses_stream_errors.py 96.93% 3 Missing ⚠️
litellm/proxy/proxy_server.py 91.66% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@zoroyihan7

Copy link
Copy Markdown
Contributor Author

@yucheng-berri Could you review this Responses streaming fix? CI and CLA pass, and the accounting discussion is resolved

@yuneng-berri
yuneng-berri deleted the branch BerriAI:main September 13, 2026 04:50
@yuneng-berri yuneng-berri reopened this Sep 13, 2026
@mateo-berri
mateo-berri changed the base branch from litellm_internal_staging to main September 17, 2026 00:44
RateLimitError and InternalServerError now carry the provider body, so
the OpenAI exception mapper keeps upstream codes like cyber_policy and
the upstream message instead of a generic mapped one

The proxy's response.failed event prefers the upstream body's code,
message, and type over the mapped exception's, and numeric error codes
in an error event map to their own HTTP status
Comment thread litellm/proxy/common_utils/responses_stream_errors.py Fixed
Emit data: [DONE] after event: response.failed, and after a late failure
when a terminal event already went out, so OpenAI SDK clients see the
same stream end as a completed response. Restore the lazy OpenAPI
snapshot to its Python 3.12 rendering, which is what CI regenerates.
Ends _response_error_code in an unconditional return so CodeQL stops flagging mixed explicit and implicit returns. No behavior change: every status maps as before.
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/common_utils/responses_stream_errors.py
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit c181c92. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks for the contribution!

@mateo-berri
mateo-berri merged commit cda022c into BerriAI:main Sep 19, 2026
90 of 91 checks passed
pull Bot pushed a commit to foreign-sub/litellm that referenced this pull request Sep 19, 2026
…upstream 500

PR BerriAI#40243 started carrying the upstream error body on InternalServerError so the Responses response.failed event can report the provider's code and message, and openai's APIError.__init__ took the body's type along with it. The proxy then answered an OpenAI-compatible upstream 500 with type server_error while a 502 and a 503 kept internal_server_error, and the integration contract in test_observed_routing.py went red. Pin the type the way RateLimitError pins throttling_error, keeping the body.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants