Repository navigation
fix(logging_worker): make flush() survive an event loop change - #42355
Conversation
flush() awaited join() on whatever queue the worker held, even one bound to an event loop that has since closed. Its unfinished counter is never decremented on the new loop, so the first flush() after a loop change hung until pytest-timeout killed it and every later one raised "is bound to a different event loop" from the queue's Event. The CircleCI unit job has been red on every branch since the first tests that flush without enqueueing landed, and an SDK script that flushes from a second asyncio.run() hangs the same way. flush() now goes through start() first, which carries the tasks stranded on the previous loop onto the current one and guarantees a worker there to drain them, the same loop-change handling every other entry point already had.
…ter a loop change
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
…r_flush_loop_change
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 9388602. Configure here.
TLDR
Problem this solves:
unitjob red on every branch since feat(auto-router): add JEV classifier alongside LLM classifier #41886test_jev_classifiercases: one 90s hang, 11 "bound to a different event loop"GLOBAL_LOGGING_WORKER.flush()joins a queue left on a closed event loopasyncio.run()hang the same wayHow it solves it:
flush()callsstart()first, so it joins the current loop's queueUser Flow
Before: a contributor pushes any branch and the required CircleCI
unitjob comes back red on 12 tests their change never touchedunitjob runs about 6,400 tests and finishes red: 12 failed, all intests/unit/router_strategy/complexity_router/test_jev_classifier.pyFailed: Timeout (>90.0s) from pytest-timeout.and the other 11 readRuntimeError: <asyncio.locks.Event ...> is bound to a different event loopuv run pytest tests/unit/router_strategy/complexity_router/test_jev_classifier.pylocally and get 43 passed, so nothing points at their changeAfter: the same push comes back green and the job finishes about 90 seconds sooner
unitjob runs the same tests and finishes green, the 12 jev cases includeduv run pytest tests/unit/router_strategy/complexity_router/test_jev_classifier.pylocally and get 43 passedThe same bug reaches SDK users who flush the logging worker themselves
Before: a script that runs
litellm.acompletionin oneasyncio.run()and thenasyncio.run(GLOBAL_LOGGING_WORKER.flush())never returns from the flushchatcmpl-...idAfter: the same script returns from the flush right away
chatcmpl-...idRelevant issues
Affected release
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Shared setup: a worktree of this repo after
make bootstrap(Python 3.14.3, pytest-asyncio 1.3.0, pytest-xdist 3.8.0, pytest-timeout 2.4.0), every command run from the worktree root. Case 3 needsOPENAI_API_KEYin the environment and uses this script, saved assdk_flush.pyoutside the repo:Before (3d26a29)
The exact CircleCI
unitjob command over the wholetests/unittreeCommand
Observed output
The 12 are exactly the cases CircleCI job 2197604 on
mainfails (4xtest_jev_http_errors_do_not_dispatch_successful_usage, 8xtest_jev_invalid_usage_never_reaches_spend_callbacks). The other 184 only fail on this laptop (missingvertexai, secret-detection scanner cases) and are not in the CircleCI resultsTwo modules on one worker, the smallest sequence that reproduces
Command
Observed output
The RuntimeError traceback ends in
logging_worker.py:490: in flush -> await self._queue.join()SDK script: one real OpenAI call, then flush from a second
asyncio.run()Command
Observed output
After (e86ba8b)
The tip is now 9388602: a merge of
main(37f1670) that resolves the branch's stale lint script, then a test-only commit that swaps the two new tests' list trackers forAsyncMock. The merge brings in commits touching neitherlogging_worker.pynor the two test modules above, and the PR'slogging_worker.pydiff against its base is byte-identical before and after both commits, so the runs below still describe the tip; the last item is CircleCI's ownunitjob run at 9388602 itselfThe exact CircleCI
unitjob command over the wholetests/unittreeSame command
Observed output
The failure sets differ by exactly the 12 jev cases (12 only before, 0 only after); the 90-second hang is gone (an earlier run of the same command at 212ab63, the fix commit alone, finished in 38.77s on an idle machine with the same 184 failures)
Two modules on one worker, the smallest sequence that reproduces
Same command
Observed output
SDK script: one real OpenAI call, then flush from a second
asyncio.run()Same command
Observed output
CircleCI's own
unitjob at the tip (9388602)Job: https://app.circleci.com/pipelines/github/BerriAI/litellm/90019/workflows/c658d518-d040-4418-b6d9-c876fb3c885f/jobs/2198653 (pipeline 90019, started 2026-09-21 23:22Z)
Observed result, read back from the CircleCI tests API for job 2198653
Same-hour control on branches without this fix: pipeline 90018 (23:01Z, job 2198616) and pipeline 90015 (22:41Z, job 2198465) both still fail the 12 jev cases
Type
🐛 Bug Fix
Caveats (if any)
Low
flush()now starts a worker on the calling loop when none runs therelitellm/orenterprise/callsflush(), only tests and SDK users_flush_on_exit) never goes throughflush(), so it is untouchedflush()also runs callbacks earlier tests left on the previous loopensure_initialized_and_enqueuealready didflush()then returns with the leftover parked until the next loop change or atexitlogging_worker.py; the GitHub required checks (the ruleset's list, all GitHub Actions jobs) are greenunit: only the wall-clock-bound secret detection case above, also red on main pipeline 89979llm_translation_testing(10 Fireworks document-inlining cases, one hosted_vllmunhashable type: 'list', one Bedrock Moonshot JSON stream) andlocal_testing_part1(two Bedrockget_model_infocases): identical failure lists on pipelines 90015 and 90018 of other branchesintegration-costandintegration-providers: red on every pipeline from 90009 (21:54Z) through 90019 across seven branches, green on 90008 and 90010 whose tips predate the newest main commitsFinal Attestation
The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
e86ba8b passes /live-pr-risk