Repository navigation
fix(responses): emit typed streaming failure events - #40243
Conversation
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
|
@yucheng-berri Could you review this Responses streaming fix? CI and CLA pass, and the accounting discussion is resolved |
RateLimitError and InternalServerError now carry the provider body, so the OpenAI exception mapper keeps upstream codes like cyber_policy and the upstream message instead of a generic mapped one The proxy's response.failed event prefers the upstream body's code, message, and type over the mapped exception's, and numeric error codes in an error event map to their own HTTP status
Emit data: [DONE] after event: response.failed, and after a late failure when a terminal event already went out, so OpenAI SDK clients see the same stream end as a completed response. Restore the lazy OpenAPI snapshot to its Python 3.12 rendering, which is what CI regenerates.
Ends _response_error_code in an unconditional return so CodeQL stops flagging mixed explicit and implicit returns. No behavior change: every status maps as before.
|
bugbot run |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit c181c92. Configure here.
mateo-berri
left a comment
There was a problem hiding this comment.
LGTM. Thanks for the contribution!
…upstream 500 PR BerriAI#40243 started carrying the upstream error body on InternalServerError so the Responses response.failed event can report the provider's code and message, and openai's APIError.__init__ took the body's type along with it. The proxy then answered an OpenAI-compatible upstream 500 with type server_error while a 502 and a 503 kept internal_server_error, and the integration contract in test_observed_routing.py went red. Pin the type the way RateLimitError pins throttling_error, keeping the body.
TLDR
Problem this solves:
/v1/responsesarrived as a Chat-styledata: {"error": ...}frameevent: response.failed, nosequence_number, nodata: [DONE], so clients saw a dropped connection"code": "500"and alitellm.InternalServerError:prefixHow it solves it:
event: response.failedwith the response id, the nextsequence_number, and the provider's code and messagedata: [DONE], the same close a completed stream getsdata:line of that named frame, soGET /v1/responses/{id}on abackground: truerequest reportsfailedwith the provider's code and message instead ofcompletedwith no errorUser Flow
Before: a developer running Codex CLI through the gateway sees "stream disconnected before completion" with no reason when the provider rejects a request mid-stream
https://litellm-domain/v1withwire_api = "responses"and send a prompt"stream": trueresponse.created,response.in_progress, and the first output eventsdata: {"error": {"message": "litellm.InternalServerError: InternalServerError: OpenAIException - Your request was flagged for possible cybersecurity risk and was not completed", "type": null, "param": null, "code": "500"}}and closes with noresponse.failedevent and nodata: [DONE]stream disconnected before completion: stream closed before response.completed, and the real reason is only visible in the proxy's pod logsAfter: the same rejection arrives as a
response.failedevent and Codex shows the provider's reasonhttps://litellm-domain/v1withwire_api = "responses"and send a prompt"stream": trueresponse.created,response.in_progress, and the first output eventsevent: response.failedwhose data carries the same response id, the nextsequence_number, and"error": {"code": "cyber_policy", "message": "Your request was flagged for possible cybersecurity risk and was not completed"}, thendata: [DONE]Relevant issues
Supersedes #39748, which kept the Chat-style error frame on Responses streams and only appended
[DONE]Affected release
Linear ticket
Resolves LIT-5620
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or moreproxy-infra / Run tests, is not required and failstest_login_throttle_settings_are_not_hot_applied_from_the_databaseon every recentmainrun too; this PR does not touch login throttling@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Real
gpt-6-astracalls through two DB-less proxies, each started from a checkout of the named commit on a free random port with--num_workers 2(2 uvicorn workers, one proxy process per port; the changed surface keeps no per-process state, so there is no second instance to alternate across), with the sameOPENAI_API_KEY. Proxy A (port 30790) fronts OpenAI and carries a callback that lets the first text delta through and then ends the stream with the in-band rejection OpenAI sends for a cybersecurity block (HTTP 200 first, then an error whose code iscyber_policy); that block cannot be triggered on demand, so the callback replays its shape. Proxy B is the proxy under test (port 32466 for Before, 56341 for After):gpt-6-astragoes straight to OpenAI andgpt-6-astra-chainedgoes through proxy A, with a callback that rejects any stream whose text contains "banana", no retries, and no fallbacksProxy A,
config_a.yamlandhook_cyber.py:Proxy B,
config_b.yamlandhook_banana.py:The Redis on port 48685 is a plain
redis-serverstarted for this run; it serves proxy B'sbackground: truepolling state, which the polling legs read back throughGET /v1/responses/{id}The curl legs use
leg <port> <model> <prompt> <out>, which saves the raw stream and prints the status, theevent:lines, the frame counts, and the last two data frames:The polling legs use
poll <port> <model> <prompt> <out>, which creates the response with"background": true, prints the id and status it got back, then pollsGET /v1/responses/{id}every 2 seconds until the status is terminal and prints that state's id, status, error, and joined output text:The Codex CLI legs run Codex CLI 0.155.1 in a fresh
CODEX_HOMEwith thisconfig.toml(the port is proxy B's):Before (6e750e3)
Output policy rejects a generated word
leg 32466 gpt-6-astra "Reply with exactly banana. Do not use tools." banana_beforeresponse.content_part.added: no failure event, no[DONE], and curl reports the transfer closed with data outstandingUpstream provider rejects mid-stream
leg 32466 gpt-6-astra-chained "Say hello in one word." cyber_beforeevent:line, nosequence_number, the provider'scyber_policycode replaced by"500", and no[DONE]Codex CLI 0.155.1
base_urlpointing at port 32466, runcodexinteractively in a tmux pane and sendReply with exactly banana. Do not use tools.stream disconnected before completion: Transport error: network error: error decoding response bodymodel = "gpt-6-astra-chained", sendSay hello in one word.Helloand thenstream disconnected before completion: stream closed before response.completedCompleted stream unchanged
leg 32466 gpt-6-astra "Say hello in one word." happy_beforeBackground polling
poll 32466 gpt-6-astra "Reply with exactly banana. Do not use tools." poll_banana_beforepoll 32466 gpt-6-astra-chained "Say hello in one word." poll_cyber_beforebackground_streaming_errorcode; the upstream rejection is polled back ascompletedwith no error at all, because the Chat-style error frame carries notypeand the stream ends without a terminal eventpoll 32466 gpt-6-astra "Say hello in one word." poll_happy_beforeAfter (c181c92)
Output policy rejects a generated word
leg 56341 gpt-6-astra "Reply with exactly banana. Do not use tools." banana_afterevent: response.failedcarrying the same response id,sequence_number4 right after the last streamed event, the policy's reason, and thendata: [DONE]; curl exits 0Upstream provider rejects mid-stream
leg 56341 gpt-6-astra-chained "Say hello in one word." cyber_aftercyber_policycode and message reach the client unchanged, after theHellodelta, followed bydata: [DONE]Codex CLI 0.155.1
base_urlpointing at port 56341, runcodexinteractively in a tmux pane and sendReply with exactly banana. Do not use tools.stream disconnected before completion: 400: Output contains blocked keyword: banana, the policy's reason instead of a transport errormodel = "gpt-6-astra-chained", sendSay hello in one word.This content can't be shown / We take extra care with some cybersecurity requests, instead of a disconnectCompleted stream unchanged
leg 56341 gpt-6-astra "Say hello in one word." happy_afterresponse.completedand one[DONE]; this run streamed one moreresponse.output_text.deltathan the Before run (11 data lines against 10), which is OpenAI's chunking, not the proxyBackground polling
poll 56341 gpt-6-astra "Reply with exactly banana. Do not use tools." poll_banana_afterpoll 56341 gpt-6-astra-chained "Say hello in one word." poll_cyber_afterfailedwith the policy's or the provider's own code and message, read from theresponse.failedframe the background task recorded; theHellothe provider streamed before rejecting stays on the responsepoll 56341 gpt-6-astra "Say hello in one word." poll_happy_afterType
🐛 Bug Fix
Caveats (if any)
Low
response.completedalready went out gets onlydata: [DONE], never a second terminal eventdata: [DONE]like the proxy's completed streams; OpenAI itself sends nonebodyrelay that givesresponse.failedthe provider's code and message also reaches error responses on the other endpoints: anInternalServerErrororAPIErrormapped from an OpenAI-compatible upstream now reports that upstream body'stype(server_errorfor an OpenAI 500) where it used to report LiteLLM's status-derivedinternal_server_error, ornullon a Chat Completions stream error frame. Observed on a chained 500 through/v1/chat/completions: Before"type":"internal_server_error", After"type":"server_error", same HTTP status, same"code":"500", same message. A 429 keepsthrottling_errorresponse.failedframe carriesoutput: []even when deltas already streamed before the failure (the chained leg had printedHello). Chosen over accumulating output items in the proxy's error state because the client already holds every delta from the same connection andGET /v1/responses/{id}on abackground: truerequest keeps the accumulated text (output_textHelloin the polling leg); a client that readsresponse.outputoff the failed event to recover partial text gets an empty listFinal Attestation