Skip to content

fix(utils): stop wrapper_async submitting the sync success handler twice - #41058

Merged
yassin-berriai merged 1 commit into
litellm_internal_stagingfrom
litellm_fix_wrapper_async_double_sync_success_handler
Sep 14, 2026
Merged

yassin-berriai merged 1 commit into
litellm_internal_stagingfrom
litellm_fix_wrapper_async_double_sync_success_handler

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Async requests ran the sync success pipeline twice per call
  • Two logging threads mutated the same logging object at once
  • Legacy success_callback functions fired twice per request

How it solves it:

  • The async logging task no longer re-submits the sync success handler
  • The direct dispatch at request completion is the one remaining submission

User Flow

Before: an operator with a plain success_callback function configured sees it fire twice for every request

  1. They add litellm_settings.success_callback: ["my_module.my_callback"] to the proxy config and start the proxy with --detailed_debug
  2. They send POST http://localhost:4000/v1/chat/completions with {"model":"gpt-5.4-mini","messages":[{"role":"user","content":"Reply with the single word: pong"}],"max_tokens":16}
  3. The response is 200 with a chatcmpl-... object and "content":"pong"
  4. The proxy log shows two Logging Details LiteLLM-Success Call lines for that one request, and their callback receives the same request twice, from two different threads at the same time

After: the same request drives the callback exactly once

  1. They add litellm_settings.success_callback: ["my_module.my_callback"] to the proxy config and start the proxy with --detailed_debug
  2. They send POST http://localhost:4000/v1/chat/completions with {"model":"gpt-5.4-mini","messages":[{"role":"user","content":"Reply with the single word: pong"}],"max_tokens":16}
  3. The response is 200 with a chatcmpl-... object and "content":"pong"
  4. The proxy log shows one Logging Details LiteLLM-Success Call line for that request, and their callback receives it once

Relevant issues

Affected release

Linear ticket

Resolves LIT-6931

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Root cause: wrapper_async ends every successful async call in _dispatch_success_logging, which submits logging_obj.success_handler to the logging thread pool directly and also schedules _client_async_logging_helper, and that helper submitted the very same success_handler a second time. Both executor threads then ran _success_handler_body against one logging object. The fix deletes the second submission from the helper; the helper now only enqueues async_success_handler. The diff is a 9 line removal in litellm/utils.py plus one regression test.

Scope: this PR covers the non-streaming async path (acompletion, aresponses, aembedding, everything that goes through wrapper_async). Streaming completions finish their logging in CustomStreamWrapper through dispatch_success_handlers, a separate code path this PR does not touch. The sync wrapper submits success_handler once and is also unchanged.

Shared setup (both arms):

# ~/lit6931/config.yaml
model_list:
  - model_name: gpt-5.4-mini
    litellm_params:
      model: openai/gpt-5.4-mini
      api_key: os.environ/OPENAI_API_KEY

litellm_settings:
  success_callback: ["lit6931_callbacks.customer_success_callback"]

general_settings:
  master_key: sk-lit6931
# ~/lit6931/lit6931_callbacks.py
def customer_success_callback(kwargs, completion_response, start_time, end_time):
    pass
sudo service postgresql start
(cd litellm/proxy && uv run --no-sync prisma db push --accept-data-loss --skip-generate)
PYTHONPATH=/home/ubuntu/repos/litellm uv run --no-sync litellm --config ~/lit6931/config.yaml --detailed_debug --port 46931 > ~/lit6931/proxy_<arm>.log 2>&1

The same request is sent three times in each arm, then the debug log is counted. Each Logging Details LiteLLM-Success Call line is printed once per run of the sync success_handler, so the expected count is 3 (one per request).

for i in 1 2 3; do
  curl -s http://127.0.0.1:46931/v1/chat/completions \
    -H "Authorization: Bearer sk-lit6931" \
    -H "Content-Type: application/json" \
    -d '{"model":"gpt-5.4-mini","messages":[{"role":"user","content":"Reply with the single word: pong"}],"max_tokens":16}'
  echo; sleep 2
done
grep -c "Logging Details LiteLLM-Success Call" ~/lit6931/proxy_<arm>.log
grep -c "Logging Details LiteLLM-Async Success Call" ~/lit6931/proxy_<arm>.log
grep -c "Writing spend log to db" ~/lit6931/proxy_<arm>.log

Before (b3882d8)

  1. sha256sum litellm/utils.py on the served checkout matched git show b3882d8e43:litellm/utils.py (0ee6366f...), and litellm.__file__ in the proxy environment resolved to /home/ubuntu/repos/litellm/litellm/__init__.py
  2. The three curls returned 200 with "content":"pong" (chatcmpl-ENx1YQI0OCMFHJUiVyvWONatDOy1Z, chatcmpl-ENx1bRIzJ09xiZRQ8cJ0zEZx8PHf2, chatcmpl-ENx1eJbI8R5drem6ltvT8JQGqOGNF)
  3. Counts from proxy_before.log:
$ grep -c "Logging Details LiteLLM-Success Call" proxy_before.log
6
$ grep -c "Logging Details LiteLLM-Async Success Call" proxy_before.log
3
$ grep -c "Writing spend log to db" proxy_before.log
3
$ grep -n "Logging Details LiteLLM-Success Call" proxy_before.log
405:09:12:21 - LiteLLM:DEBUG: litellm_logging.py:2498 - Logging Details LiteLLM-Success Call: Cache_hit=None
435:09:12:21 - LiteLLM:DEBUG: litellm_logging.py:2498 - Logging Details LiteLLM-Success Call: Cache_hit=None
592:09:12:24 - LiteLLM:DEBUG: litellm_logging.py:2498 - Logging Details LiteLLM-Success Call: Cache_hit=None
594:09:12:24 - LiteLLM:DEBUG: litellm_logging.py:2498 - Logging Details LiteLLM-Success Call: Cache_hit=None
803:09:12:26 - LiteLLM:DEBUG: litellm_logging.py:2498 - Logging Details LiteLLM-Success Call: Cache_hit=None
804:09:12:26 - LiteLLM:DEBUG: litellm_logging.py:2498 - Logging Details LiteLLM-Success Call: Cache_hit=None

Two sync success-handler runs per request (lines 592/594 and 803/804 are back-to-back entries for a single request). The sync callback fires twice per request.

After (1e0bc3d)

  1. Proxy restarted on the PR tip: sha256sum litellm/utils.py gave f6b32231..., matching git show 1e0bc3d38e:litellm/utils.py
  2. The identical three curls returned 200 with "content":"pong":
{"id":"chatcmpl-ENx74jGoy4Sb0QRicF208LTo6V3Wr","created":1789377482,"model":"gpt-5.4-mini","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"pong","role":"assistant","provider_specific_fields":{"refusal":null},"annotations":[]},"provider_specific_fields":{}}],"usage":{"completion_tokens":4,"prompt_tokens":13,"total_tokens":17,...},"service_tier":"default"}
{"id":"chatcmpl-ENx77nV0prdUlYI7xrwWncIPCxhFN","created":1789377485,"model":"gpt-5.4-mini","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"pong","role":"assistant",...}}],"usage":{"completion_tokens":4,"prompt_tokens":13,"total_tokens":17,...},"service_tier":"default"}
{"id":"chatcmpl-ENx7Ahh7YSjo8xtnxNg53PTabL168","created":1789377488,"model":"gpt-5.4-mini","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"pong","role":"assistant",...}}],"usage":{"completion_tokens":4,"prompt_tokens":13,"total_tokens":17,...},"service_tier":"default"}
  1. Counts from proxy_after.log:
$ grep -c "Logging Details LiteLLM-Success Call" proxy_after.log
3
$ grep -c "Logging Details LiteLLM-Async Success Call" proxy_after.log
3
$ grep -c "Writing spend log to db" proxy_after.log
3
$ grep -n "Logging Details LiteLLM-Success Call" proxy_after.log
405:09:18:03 - LiteLLM:DEBUG: litellm_logging.py:2498 - Logging Details LiteLLM-Success Call: Cache_hit=None
677:09:18:06 - LiteLLM:DEBUG: litellm_logging.py:2498 - Logging Details LiteLLM-Success Call: Cache_hit=None
886:09:18:08 - LiteLLM:DEBUG: litellm_logging.py:2498 - Logging Details LiteLLM-Success Call: Cache_hit=None

One sync success-handler run per request. The async handler count and the spend log count are unchanged at 3, so the fix removed only the duplicate.

Regression test: tests/test_litellm/test_utils.py::test_acompletion_runs_a_custom_logger_sync_logging_hook_exactly_once registers a CustomLogger whose logging_hook records the response id and then blocks on a gate, runs one acompletion through the real logging thread pool, releases the gate, waits for every submitted logging future and asserts the hook saw the response id exactly once. The gate matters because success_handler sets has_run_logging only after the hooks run, so a sequential re-run would be skipped while two concurrent threads both get through, which is the same race the customer hit. On the merge base it fails 3/3 with Left contains one more item: 'chatcmpl-...', on this branch it passes 3/3. tests/test_litellm/test_utils.py, tests/test_litellm/proxy/guardrails/test_deferred_guardrail_logging.py, tests/test_litellm/responses/test_no_duplicate_spend_logs.py and tests/test_litellm/litellm_core_utils/test_litellm_logging.py pass together (731 passed).

Taxonomy audit (A-BB) of the diff: B6 is the defect fixed (the same handler released twice). C1-C7 not applicable: every caller of _dispatch_success_logging still gets exactly one sync submission, including the is_litellm_internal_call, is_completion_with_fallbacks and deferred-guardrail branches. F3: streaming is a separate path (dispatch_success_handlers in streaming_handler.py) and is written into the scope above rather than changed. T4: the test wraps executor.submit only to collect the futures it needs to wait on; the real thread pool and the real success_handler run, and the assertion is on the callback the user configured. T5: success_callback is set through monkeypatch and restored. H2: no comments added.

Type

🐛 Bug Fix

Caveats (if any)

Low

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Medium Risk
Touches the async success logging path used by all non-streaming async API calls; incorrect deduplication could skip sync callbacks, but the change removes redundancy while keeping the existing _dispatch_success_logging submission.

Overview
Fixes duplicate sync success logging on non-streaming async calls (acompletion, etc.) by removing a second handle_sync_success_callbacks_for_async_calls invocation from _client_async_logging_helper in litellm/utils.py.

Async completions still enqueue async_success_handler via GLOBAL_LOGGING_WORKER; sync success_callback / CustomLogger hooks are now submitted only from _dispatch_success_logging, so legacy callbacks and logging threads no longer run twice or race on the same logging object.

Adds test_acompletion_runs_a_custom_logger_sync_logging_hook_exactly_once, which gates a sync logging_hook and tracks logging executor futures to assert the hook sees one response id per acompletion.

Reviewed by Cursor Bugbot for commit b86de17. Bugbot is set up for automated code reviews on this repo. Configure here.

Link to Devin session: https://app.devin.ai/sessions/c0c9636e652649e0865582ddb107409d
Open in Devin Desktop: https://app.devin.ai/desktop/session/c0c9636e652649e0865582ddb107409d?variant=devin
Requested by: @yassin-berriai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@codspeed

codspeed Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_fix_wrapper_async_double_sync_success_handler (b86de17) with litellm_internal_staging (b3882d8)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR prevents successful non-streaming async requests from submitting the synchronous success pipeline twice.

  • Removes the redundant sync callback dispatch from _client_async_logging_helper.
  • Retains async callback delivery and the existing direct sync dispatch.
  • Adds a behavioral regression test that runs the real logging hook and verifies exactly one invocation.
  • Fully types the newly added test helpers and queues.

Confidence Score: 5/5

The PR appears safe to merge; the duplicate sync callback path is removed and no outstanding findings remain.

The current dispatch flow retains one synchronous success submission and one asynchronous callback enqueue per successful non-streaming async request. The regression test exercises the real hook through the real thread pool, and the annotations added since the previous review satisfy the repository requirement. Both previous threads are resolved.

Important Files Changed

Filename Overview
litellm/utils.py Removes the duplicate synchronous success-callback submission while preserving async logging delivery.
tests/test_litellm/test_utils.py Adds a concurrency-sensitive behavioral regression test and fully types its callback, executor wrapper, and queues.

Reviews (4): Last reviewed commit: "fix(utils): stop wrapper_async submittin..." | Re-trigger Greptile

Comment thread tests/test_litellm/test_utils.py Outdated
@codecov

codecov Bot commented Sep 14, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_fix_wrapper_async_double_sync_success_handler branch from 1e0bc3d to 8d0324d Compare September 14, 2026 09:43
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review 8d0324d: the regression test now runs the real logging thread pool and asserts on the configured callback

@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_fix_wrapper_async_double_sync_success_handler branch from 8d0324d to f379671 Compare September 14, 2026 09:49
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review f379671: same test, plus a test-quality-ok reason on the submit wrapper so the lint gate passes

Comment thread tests/test_litellm/test_utils.py Outdated
_client_async_logging_helper re-submitted logging_obj.success_handler to the
executor after _dispatch_success_logging had already done so, running the same
success pipeline twice per async request and racing on shared logging state.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_fix_wrapper_async_double_sync_success_handler branch from f379671 to b86de17 Compare September 14, 2026 10:02
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review b86de17: the test helper parameters and queues are now fully typed

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit b86de17. Configure here.

@yassin-berriai
yassin-berriai merged commit 1f8bae7 into litellm_internal_staging Sep 14, 2026
93 of 94 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_fix_wrapper_async_double_sync_success_handler branch September 14, 2026 19:43
@yassin-berriai
yassin-berriai restored the litellm_fix_wrapper_async_double_sync_success_handler branch September 14, 2026 19:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants