Skip to content

fix(utils): isolate callback errors in async_post_call_success_deployment_hook - #42535

Merged
yucheng-berri merged 5 commits into
mainfrom
litellm_isolate_success_deployment_hook_errors
Sep 23, 2026
Merged

yucheng-berri merged 5 commits into
mainfrom
litellm_isolate_success_deployment_hook_errors

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

  • Each callback's async_post_call_success_deployment_hook await is wrapped in try/except
  • A raising hook is logged with its class and call_type, then skipped
  • The response stays as it was and later hooks still run
  • CustomGuardrail callbacks re-raise, since raising is how a post-call guardrail blocks

User Flow

Before: an operator with a chat-oriented custom callback gets a 400 back for every video job, even though the provider accepted it

  1. They send POST https://litellm-domain/v1/videos with {"model": "sora-2", "prompt": "a cat", "seconds": "4", "size": "720x1280"}
  2. They get back 400 {"error": {"message": "Invalid request format: 'VideoObject' object has no attribute 'choices'", ...}}
  3. The provider dashboard shows three queued video jobs for that one request

After: the same request returns the queued video job

  1. They send the same POST https://litellm-domain/v1/videos with the same body
  2. They get back 200 {"id": "video_...", "object": "video", "status": "queued", ...}
  3. The proxy log carries one line naming the callback class and call type that raised, and the provider shows one job

Relevant issues

Regression introduced by #42354 (merged as 12f7930)

Affected release

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.): 90 green, 3 red, each red one adjudicated in Caveats below
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live proxy on each side (python litellm/proxy/proxy_cli.py --config config.yaml --port <port>), real OpenAI key, same config on both sides:

model_list:
  - model_name: sora-2
    litellm_params:
      model: openai/sora-2
      api_key: os.environ/OPENAI_API_KEY
  - model_name: gpt-5.6
    litellm_params:
      model: openai/gpt-5.6
      api_key: os.environ/OPENAI_API_KEY
litellm_settings:
  callbacks: ["hooks.chat_shaped", "hooks.recording"]
guardrails:
  - guardrail_name: risk-blocker
    litellm_params:
      guardrail: hooks.BlockingGuardrail
      mode: post_call
      default_on: false

hooks.py on PYTHONPATH: chat_shaped is a CustomLogger whose async_post_call_success_deployment_hook reads response.choices[0].message.content and returns None, recording is a CustomLogger that logs the response type and returns None, BlockingGuardrail is a CustomGuardrail whose async_post_call_success_hook raises HTTPException(400, {"error": "blocked by risk guardrail"})

Before (9d299d0)

Video job with the chat-shaped hook

  1. curl -s http://localhost:4302/v1/videos -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' -d '{"model":"sora-2","prompt":"a cat","seconds":"4","size":"720x1280"}' -w '\nHTTP %{http_code}\n'
  2. {"error":{"message":"Invalid request format: 'VideoObject' object has no attribute 'choices'","type":"invalid_request_error","param":null,"code":"400"}} HTTP 400
  3. Proxy log: RISK chat_shaped saw VideoObject call_type=CallTypes.avideo_generation three times (three upstream jobs), recording never ran

Chat completion with the same hooks

  1. curl -s http://localhost:4302/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' -d '{"model":"gpt-5.6","messages":[{"role":"user","content":"say hi"}]}' -w '\nHTTP %{http_code}\n'
  2. {"id":"chatcmpl-...","model":"gpt-5.6","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi!","role":"assistant",...}}],...} HTTP 200, both hooks logged

Chat completion with a per-request post_call guardrail that raises

  1. Same curl with "guardrails":["risk-blocker"] added to the body
  2. {"error":{"message":"blocked by risk guardrail","type":"invalid_request_error","param":null,"code":"400","provider_specific_fields":{"error":"blocked by risk guardrail","guardrail_name":"risk-blocker","guardrail_mode":"post_call"}}} HTTP 400

After (c8e9756, production diff identical to 6c03b71; the last commit only moves the integration cases into test_callback_delivery.py)

Video job with the chat-shaped hook

  1. curl -s http://localhost:4301/v1/videos -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' -d '{"model":"sora-2","prompt":"a cat","seconds":"4","size":"720x1280"}' -w '\nHTTP %{http_code}\n'
  2. {"id":"video_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3ZpZGVyOm9wZW5haTttb2RlbF9pZDpzb3JhLTI7dmlkZW9faWQ6dmlkZW9fNmFiMmY2N2E0NWM4ODE5MTk3MzI4YWE3NjdkNjRhNDIwN2M2Yzg0YTlhMTY2NTQw","object":"video","status":"queued",...,"seconds":"4","size":"720x1280","model":"sora-2","usage":{"duration_seconds":4.0}} HTTP 200
  3. Proxy log: RISK chat_shaped saw VideoObject once, then LiteLLM:ERROR: utils.py:1452 - async_post_call_success_deployment_hook error in ChatShapedHook for call_type=CallTypes.avideo_generation, then RISK recording saw VideoObject call_type=CallTypes.avideo_generation

Chat completion with the same hooks

  1. Same curl against port 4301
  2. Same 200 response shape, both hooks logged

Chat completion with a per-request post_call guardrail that raises

  1. Same curl with "guardrails":["risk-blocker"] against port 4301
  2. Same 400 blocked by risk guardrail body, so guardrail blocking still works

Admin UI, same request on both sides

The screen recording of this walkthrough (base 9d299d0 on port 4102, an earlier tip addc9bb on port 4101, real sora-2 call on each; the production change is unchanged since) was handed to the requester with the session report. To reproduce: start each proxy with the config above and its own DATABASE_URL, send the video curl, open http://127.0.0.1:4101/ui/?page=logs and http://localhost:4301/ui/?page=logs (different origins, so the two dashboards keep separate sessions), log in with admin and the master key, search the Logs page by the returned request id and open the row. Base shows Failure with Message: 'VideoObject' object has no attribute 'choices' in the detail panel. The PR tip shows Success, call type avideo_generation, cost $0.40 and Retries: None

Live audit matrix (head 6c03b71 vs base 9d299d0, production code identical to c8e9756)

Two proxies, two uvicorn workers each, own Postgres per side, real OpenAI and Anthropic keys, no mocks. Callbacks registered: a chat-shaped hook that reads response.choices, a hook that always raises ValueError, and a recording hook that returns the response, plus a post_call CustomGuardrail (default_on: false) that raises 400 when named in guardrails

Rerun at the tip 6c03b71: head returns 200 on every non-streaming call type with two of three hooks raising: /v1/chat/completions (curl chatcmpl-ER2YXFB4..., openai sync, openai async), /v1/responses (openai sync and async), /v1/messages (anthropic sync msg_011CfKCzGUEk... and async), /v1/embeddings, and a real sora-2 /v1/videos job (video_bGl0ZWxs...MTY2NTQw, queued). The proxy log carries one async_post_call_success_deployment_hook error in <Hook> for call_type=CallTypes.<type> per raising hook and one recording saw ... line per request, so the third hook ran every time. Base on the same cells returns 500 audit: always raising hook for chat, 500 'dict' object has no attribute 'choices' for messages, 400 'ResponsesAPIResponse' object has no attribute 'choices' for responses, 400 'EmbeddingResponse' ... for embeddings and 400 'VideoObject' ... for videos. Streaming chat, responses and messages return 200 on both sides because streaming provider responses do not pass through this hook

guardrails: ["audit-blocker"] returns 400 blocked by audit guardrail on head for chat, videos and embeddings; on base the ordinary hook exception fired first (500 on chat, 400 choices error on the other two) and masked the guardrail. Bad model name (400), invalid provider key (401), no api key (401), /key/health and /health/readiness (200) are identical on both sides. guardrails as an integer gives 500 'int' object is not iterable and guardrails as a bare string is iterated per character on both sides: pre-existing, not touched here

Three identical chat requests produce three distinct spend rows with one row each. A 32-request async burst across chat, responses, embeddings and streaming chat lands 24/24 spend rows, no duplicates. Killing one worker with SIGKILL mid burst drops only the 3 in-flight connections on that worker, a replacement worker spawns, all 27 completed ids have exactly one success row, and the next 8 requests are 200. SIGHUP restarts the workers and drops the requests in flight on both head and base

Full call-type matrix (head c8e9756 vs base 9d299d0)

Same two proxies, one ordinary CustomLogger that always raises plus a recording hook, real keys for OpenAI, Anthropic, Gemini, Cohere, Mistral, Tavily and Azure OpenAI, no mocks. One real request per call type that reaches the hook. Head returns 200 and logs the isolation line and the recording line, base returns 500 audit: always raising hook, on all of: chat, text completions, embeddings, moderations, anthropic messages, responses, responses input items, gemini generateContent, rerank, search, vector store search, ocr, azure pass-through, image generation, image edit, speech, transcription, file create/retrieve/list/delete, batch create/retrieve, fine-tuning job create/retrieve/cancel, container create/list/retrieve/delete, skill create, rag ingest and query, and a completed sora-2 video with list, content, remix, extension, edit, create character and get character. MCP tools/call over the real /mcp streamable-http transport returns the tool result on head and a tool error carrying the hook message on base. Through the SDK, file content, container file upload and list, vector store search, gemini, azure pass-through and A2A asend_message (completion bridge) all return on head and raise the hook's ValueError on base

Not entering the hook on either side, unchanged by this PR: the proxy handlers for file content, container file upload and list, and the /mcp-rest/tools/call facade (they bypass the @client wrapper), every streaming call (gemini streamGenerateContent plus streaming chat, responses and messages), arealtime and aresponses_websocket, and functions absent from CallTypes (video status, list batches, skills list/get/delete, vector store CRUD, evals, interactions)

Not driven: the sandbox family (acreate_sandbox, arun_code, adelete_sandbox, acode_interpreter_tool), whose only providers are E2B (no key available) and a self-hosted OpenSandbox server. Same wrapper and dispatcher. The mock_response converted stream path and the async pre-call hook are also not exercised (out of scope, see Caveats). Per-cell commands, status codes, response ids and log lines are in the matrix report handed to the requester (matrix_report.md)

Type

🐛 Bug Fix

Caveats (if any)

Severe

  • Error semantics change for every call type routed through the async success deployment hook, not only video
    • a non-guardrail hook that raises no longer fails the request; it is logged and the original response is returned
    • operators who relied on a raising hook to reject responses must make it a CustomGuardrail
    • approved in the task brief that requested this PR

Medium

  • There is no sync post_call_success_deployment_hook dispatcher; the only other caller is responses/streaming_iterator.py, which already wraps the hook in a bare try/except and drops the exception without logging. Unchanged here
  • async_pre_call_deployment_hook still propagates hook exceptions; unchanged here, and pre-call hooks may legitimately reject a request
  • Guardrail exemption is by isinstance(callback, CustomGuardrail), so a guardrail built on plain CustomLogger would now be swallowed

Low

  • CI: rust-wheel fails the same seven tests/test_litellm_rust tests on main 05d7fb2 (run 35772779792), proxy-infra / Run tests fails on the merge base 9d299d0 too, and codecov/patch (37.5%) is below target because the misc unit shard that owns tests/test_litellm/test_utils.py collected [0 items] in CI and the workflow treats pytest exit 5 as a pass, so the new branches in litellm/utils.py get no upload even though the mapped regression is green locally. None of the three touch the changed lines
  • tests/test_litellm/test_utils.py has two failures unrelated to this PR that also fail on main (test_generate_azure_ad_redis_token, test_aaamodel_prices_and_context_window_json_is_valid)

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Tests

tests/test_litellm/test_utils.py::test_success_deployment_hook_raising_keeps_response_and_runs_later_hooks is parametrized over VideoObject, EmbeddingResponse, ResponsesAPIResponse, ImageResponse, RerankResponse, TranscriptionResponse and ModelResponse (chat and anthropic messages call types) and asserts the response is returned unchanged, the later hook still runs and exactly one verbose_logger.exception record names the callback class and call type. A second test checks a hook that raises after an earlier hook rewrote the response keeps the rewrite. The guardrail test checks a raising CustomGuardrail still aborts. Mutation: replacing the catch with a bare raise fails 9 of the 13 mapped tests

tests/integration/observability/test_callback_delivery.py::test_response_survives_raising_success_deployment_hook (added to the existing callback file, four parametrized cases) starts a real proxy with an always-raising CustomLogger callback and a scripted upstream, and sends /v1/chat/completions, /v1/embeddings, /v1/responses and /v1/videos. All four are 500 hook rejected <Type> for CallTypes.<type> on base 9d299d0 and 200 with the upstream payload intact at the tip

ran /live-pr-risk at 6c03b71 (production diff unchanged since the first run and unchanged by the test-only move to c8e9756, live matrix rerun on that code) and found no regressions/backward incompatible risks beyond the approved one above

Link to Devin session: https://app.devin.ai/sessions/d954ca781d894d068319a115b73cf73b
Open in Devin Desktop: https://app.devin.ai/desktop/session/d954ca781d894d068319a115b73cf73b?variant=devin
Requested by: @yucheng-berri

yucheng-berri and others added 2 commits September 22, 2026 19:04
…ment_hook

A callback that raises inside async_post_call_success_deployment_hook no longer
fails the completed request. The exception is logged with the callback class and
call_type, the response stays as it was, and later callbacks still run. Guardrail
callbacks are exempt because raising is how a post-call guardrail blocks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@greptile-apps

greptile-apps Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with no outstanding correctness or repository-policy findings.

Findings

  1. P2 New test file violates policy ▶

Summary

This PR isolates failures from ordinary post-call success callbacks so completed requests retain their responses and later callbacks still run, while preserving exception propagation for CustomGuardrail callbacks.

Changes include:

  • Catching and logging ordinary callback exceptions per callback.
  • Preserving earlier callback response rewrites after a later callback fails.
  • Adding unit coverage across multiple response and call types.
  • Moving the proxy-level regression test into the existing callback-delivery integration suite and updating its contract mappings.

Reviews (4) · Last reviewed commit: "test(integration): move raising success ..."

Comment thread litellm/utils.py
Comment thread tests/test_litellm/test_utils.py Outdated
@codspeed

codspeed Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_isolate_success_deployment_hook_errors (c8e9756) with main (d7c27cd)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (0d6ee3d) during the generation of this report, so d7c27cd was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@codecov

codecov Bot commented Sep 22, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 37.50000% with 5 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/utils.py 37.50% 5 Missing ⚠️

📢 Thoughts on this report? Let us know!

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…sing hook

Parametrize the unit regression over video, embedding, responses, image, rerank,
transcription, chat and anthropic messages responses and assert the failure log
names the callback and call type. Run the integration test through a real proxy
for /v1/chat/completions, /v1/embeddings, /v1/responses and /v1/videos

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@@ -0,0 +1,171 @@
import base64

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 New test file violates policy

This bug fix creates a separate integration test file instead of extending the existing mapped callback test file. Repository policy requires bug fixes to extend the existing mapped test file, so this requirement must be satisfied before merging.

Context Used: AGENTS.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

…callback delivery file

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit c8e9756. Configure here.

@yucheng-berri
yucheng-berri merged commit 944f44d into main Sep 23, 2026
96 of 98 checks passed
@yucheng-berri
yucheng-berri deleted the litellm_isolate_success_deployment_hook_errors branch September 23, 2026 00:59
yuneng-berri pushed a commit that referenced this pull request Sep 23, 2026
…e failure-hook regression (#42646)

* test(utils): raise the post-success hook error from a guardrail in the failure-hook regression

Since #42535 a plain logger raising inside async_post_call_success_deployment_hook
is logged and the completed request returns, so the regression added by #36657 for
"a post-success error never reaches async_post_call_failure_deployment_hook" failed
with DID NOT RAISE on every main run once #42603 revived the misc unit shard. The
raising callback is now a CustomGuardrail, the one kind of callback whose
post-success raise still propagates, which keeps the original assertions intact

* test(utils): type the guardrail's success-hook request_data as a Mapping

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants