Skip to content

test: deflake redis semantic cache sys.modules leak, LangSmith init loop patch, wall-clock stagger assertion, and zombie grandchild check in the fake prisma cli - #39895

Merged
mateo-berri merged 13 commits into
litellm_internal_stagingfrom
litellm_deflake_20260905
Sep 12, 2026
Merged

mateo-berri merged 13 commits into
litellm_internal_stagingfrom
litellm_deflake_20260905

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 5, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Four unit tests failed once and passed on CI rerun, no code change between
  • Redis semantic cache tests dropped litellm.proxy.proxy_server from sys.modules
  • LangSmith init test patched asyncio.get_running_loop globally, GC finalizers hit the mock
  • Scheduled job stagger test read cron fire times off the wall clock, broke at 00:15 UTC
  • Fake prisma CLI counted a zombie grandchild as still alive, so process-tree asserts timed out

How it solves it:

  • Fake only the two redisvl modules with MonkeyPatch.setitem, not all of sys.modules
  • Run the LangSmith test under a real event loop and wait for the scheduled task to flush a queued event
  • Compare the staggered and unstaggered cron triggers from a fixed start instead of datetime.now()
  • Read /proc/<pid>/stat and treat state Z as gone, since a zombie has already been killed
  • The pgbouncer readiness budget fix from the 2026-09-11 run was dropped here because test(pgbouncer): stop the never-listens replacement test flaking under CI load #40830 landed the same change on staging

User Flow

Not applicable. This PR only changes unit tests; proxy behavior is unchanged

Relevant issues

Linear ticket

Flakes covered, by run date

2026-09-12

tests/test_litellm/proxy/db/test_pgbouncer.py::TestPgBouncerProcess::test_a_replacement_that_never_listens_is_replaced_again

Flaked 20 times in the window (2026-09-11 09:20 to 2026-09-12 09:20 UTC, 16,181 workflow runs scanned, 652 failed job logs parsed), always in Unit Tests / proxy-infra / Run tests with the same assert pooler.start() is None reporting PgBouncerError(reason='127.0.0.1:<port> is served by another process, not the pgbouncer that was started'). Same-sha evidence: https://github.com/BerriAI/litellm/actions/runs/34659643270 (attempt 1 failed, attempt 2 passed), plus failures on litellm_internal_staging itself at https://github.com/BerriAI/litellm/actions/runs/34652853654 and https://github.com/BerriAI/litellm/actions/runs/34648862502 where the next staging run passed. Same mechanism as the 2026-09-11 entry below

Every one of the 20 failures ran before #40830 merged into staging at 02:32 UTC on 2026-09-12 with ready_timeout_seconds=2.0 and the log assertion moved after the port file rewrite. That is the same fix this branch carried since 2026-09-11 at 3.0s, so the branch change was dropped in favor of the staging version when staging was merged in (441f863). The diff of this PR no longer touches test_pgbouncer.py

Verification of the upstream fix on the merged branch: the test ran 20 times idle and 20 times under 16 busy-loop processes on 8 cores (load average about 17) with 0 failures, and the whole test_pgbouncer.py file passes both serially and under -n 4

One other same-sha flip was not a staging flake: tests/test_litellm_rust/ocr/test_callbacks.py::test_native_azure_ocr_releases_token_provider_after_terminal_outcome[cancellation] failed on https://github.com/BerriAI/litellm/actions/runs/34638756281 and passed on https://github.com/BerriAI/litellm/actions/runs/34638764815 at sha 7b9d0fe, but that sha is on a feature branch that rewrites the Azure credential provider the test holds a weakref to, so it is that branch's own work in progress. Real breakages seen on every attempt, confined to feature branches and left for their authors: tests/test_litellm_rust/test_ocr.py ('RecordedOCRRequest' object is not subscriptable) on litellm_ocr_azure_mistral and litellm_ocr_mistral_native, the retained-callbacks suite on litellm_rust_retained_callbacks, test_rust_bridge.py (no attribute '_run_rust_ocr') on litellm_ocr_completion_boundary, and test_aaamodel_prices_and_context_window_json_is_valid on the price-sync branches (output_cost_per_video_per_second_batches not in the schema). Infra, not test debt: 15 ai-gateway release image jobs failing with No such container: ai-gateway then passing on retry, 14 e2e-changed-tests approval-gate flips, one proxy-infra runner abort mid-run (https://github.com/BerriAI/litellm/actions/runs/34646984186), and one osv-scan flip

2026-09-11

tests/test_litellm/proxy/db/test_pgbouncer.py::TestPgBouncerProcess::test_a_replacement_that_never_listens_is_replaced_again

Flaked 8 times in the window (2026-09-10 09:15 to 2026-09-11 09:15 UTC, 15,430 workflow runs scanned), every time in Unit Tests / proxy-infra / Run tests on attempt 1, where pytest's own --reruns 2 ran it three times in a row and it failed all three with assert pooler.start() is None reporting PgBouncerError(reason='127.0.0.1:<port> is served by another process, not the pgbouncer that was started'). Seven of the eight runs passed on attempt 2 of the same sha, the eighth has no attempt 2:

Mechanism: wall-clock dependence on interpreter start time. The test builds a PgBouncerProcess with ready_timeout_seconds=0.3 so the later replacement that binds the wrong port gets rejected quickly, but the same 0.3s budget also applies to the first, healthy start(). The fake pooler is a Python script, and in this shard the child interpreter also loads the pytest-cov subprocess hook (COVERAGE_CORE=sysmon, --cov) before it binds anything: 0.19s on an idle box here, more on a 4-worker xdist runner sharing 2 vCPUs with other subprocess-spawning db tests. _wait_ready polls every 0.1s for TCP and Unix socket together, so when the child gets its TCP bind in between the last in-budget poll and the post-deadline check, or has bound TCP but not yet the Unix socket, the supervisor reports the port as held by a stranger and start() returns an error. Reproduced locally without edits by running the test with --cov under 64 busy-loop processes: 2 failures in 10 runs (did not start listening within 0s, the other branch of the same deadline). Idle, 0 failures in 20

Fix (since superseded by #40830, see 2026-09-12): ready_timeout_seconds=3.0, matching test_a_listener_that_grabs_the_port_after_the_spawn_is_not_taken_for_the_pooler in the same class, and a 10s ceiling on the one _wait_until that now has to outlast that budget plus the restart delay before the "did not start listening" log line appears. Every other step still waits on a condition (port listening, port closed, log record present), not a duration. The test takes about 3.5s instead of 0.5s

Mutation check: with _retry_restart no longer starting the supervisor thread, the test fails at _wait_until(lambda: _listening(port)) because nothing replaces the rejected pooler. Reverted

Not touched, reported as a real breakage instead: tests/test_litellm/proxy/anthropic_endpoints/test_claude_code_marketplace.py::test_archive_source_registers_and_is_served_verbatim_in_marketplace fails on every attempt with TypeError: get_marketplace() missing 1 required positional argument: 'request', and the Terraform Provider job keeps failing on every attempt as noted under 2026-09-10. Fail-then-pass on Documentation validation, Helm, CodSpeed, Conventional PR Title and OSV Scan were network, cancellation or advisory timing, not test flakes

2026-09-10

No new flake in the window (2026-09-09 09:20 to 2026-09-10 09:20 UTC, 17,493 workflow runs scanned). The only same-sha fail-then-pass on a test was ui-unit-tests on PR #40367 sha c59b321: add_auto_router_tab.test.tsx > carries a preset's per-tier reasoning effort through to the create payload failed on https://github.com/BerriAI/litellm/actions/runs/34328740942 (attempt 1, 08:28 UTC) and passed on https://github.com/BerriAI/litellm/actions/runs/34430807521 (attempt 1, 02:52 UTC the next day). That is not a flake. Both runs are pull_request events, so the job checks out the merge of the PR head into staging, and staging moved between them: #40341 repointed the Anthropic REASONING preset at claude-fable-5-1 while the test still hardcoded claude-opus-5, then #40456 (2c836f4, 21:14 UTC) made the assertion read the preset. The second run merged in that fix. Nothing to change here

Every other fail-then-pass was infra: nine e2e-changed-tests runs whose first attempt was concurrency-cancelled mid detection (DETECT_RESULT: cancelled), one cancelled Conventional PR Title run, and an osv-scan that started failing when GHSA-7w5x-hrqm-74c2 (smol-toml) was published between two runs of the same sha. One real breakage, reported in Slack rather than touched here: the Terraform Provider job fails TestResourceKeyUpdateFailureKeepsPriorState (resource_key_test.go:356: failed update persisted budget_duration="bad" into state) on every attempt of every PR into staging since #40512 added it, for example https://github.com/BerriAI/litellm/actions/runs/34434802755 and https://github.com/BerriAI/litellm/actions/runs/34459965452

None of the four tests this PR fixes failed anywhere in the window. Branch re-merged with litellm_internal_staging (c956c24)

2026-09-09

tests/test_litellm/proxy/db/test_check_migration.py::test_migrate_diff_stops_at_its_budget_and_takes_its_process_tree_with_it

Flaked 8 times in the window, every time in Unit Tests on attempt 1 with assert fake_prisma_cli.grandchild_is_gone(within_seconds=5) failing as assert False, then PASSED on attempt 2 of the same sha:

Mechanism: the test helper depended on who reaps orphans. run_prisma does kill the whole process group on timeout (SIGKILL via killpg), so the fake CLI's grandchild dies. But the grandchild's parent (the fake CLI) died in the same kill, so the dead grandchild is reparented to the nearest subreaper, and until that ancestor calls wait on it, the process stays a zombie. FakePrismaCli.grandchild_is_gone checked liveness with os.waitpid (fails with ChildProcessError, the pid is not our child) and os.kill(pid, 0), which succeeds on a zombie. So the helper reported "still alive" for as long as the ancestor took to reap, and whether that fits in 5 seconds depends on the runner's process tree, not on the code under test. Not reproducible in a plain shell where init reaps at once; reproduced 100% locally by running pytest under a parent that sets PR_SET_CHILD_SUBREAPER and never reaps, which is the shape of a CI runner's job wrapper. The same run also failed tests/test_litellm/proxy/db/test_prisma_client.py::test_db_push_timeout_takes_its_process_tree_with_it, which shares the helper. Production code is correct, the kill happens; only the observation was wrong

Fix: grandchild_is_gone also reads /proc/<pid>/stat and returns true when the state field is Z. A zombie has already been killed, which is what the assertion is about. Nothing else in the helper or the tests changed, and the 5 second window and the assertion stay as they were

Mutation check: with _kill_process_group changed to process.kill() (kill only the CLI, not the group), both process-tree tests fail because the grandchild keeps sleeping and is neither gone nor a zombie. Mutation reverted

Branch re-merged with litellm_internal_staging (db4dee5). Gates on fdd6f60: both process-tree tests pass under the non-reaping subreaper wrapper, each ran 20x in a loop at 0 failures, test_check_migration.py, test_prisma_client.py and the whole tests/test_litellm/proxy/db/ directory pass, make lint and make check pass

Not touched this run: a same-sha fail-then-pass cluster in tests/test_litellm/proxy/_experimental/mcp_server/test_mcp_proxy_mode.py (authorization tests seeing an empty tool set, e.g. assert set() == {'beta-add', 'alpha-add'}). It was not reproduced locally and the mechanism was not pinned from the logs, so it is reported as unreproduced rather than guessed at

2026-09-08

tests/test_litellm/proxy/common_utils/test_scheduled_job_stagger.py::test_explicit_offset_overrides_the_derived_one_and_zero_pins_a_job

Flaked once in the window, in Unit Tests / proxy-infra / Run tests, RERUN twice then FAILED on attempt 1, PASSED on attempt 2 of the same sha:

Mechanism: wall-clock dependence. The test built two schedulers and read job.next_run_time after scheduler.start(paused=True), which evaluates the PTU rollup cron (00:15 UTC daily) against datetime.now(). Between 00:15:00 and 00:15:07 UTC the unstaggered cron has already rolled to tomorrow, while the +7s offset trigger evaluates the base cron at now - 7s, which is still today, and yields today 00:15:07. The assertion then saw 2026-09-08 00:15:07 - 2026-09-09 00:15 instead of 7 seconds. Reproduced locally under faketime '2026-09-09 00:14:55' with the identical assertion shape; the production code behaves correctly, only the test's reference point moved

Fix: the test now takes the PTU trigger off each scheduler and compares _fire_times(trigger, start, 1) from a fixed start, the same way test_default_cron_is_staggered_and_keeps_its_offset_on_every_later_fire already did. A small _trigger_of(scheduler, job_id) helper replaces the duplicated next(job.trigger for ...) in both tests. The test no longer awaits anything, so it is a plain def

Mutation check: _OffsetTrigger.get_next_fire_time returning the base fire time without + self.offset fails the test (00:15 - 00:15 == 7s); _clamped_override returning 0 fails it (assert 0 == 7). Both mutations reverted

Branch re-merged with litellm_internal_staging (1af7a40). Gates re-run on 61e6640: the fixed test 20x at 0 failures, the owning file 22 passed, the redis semantic cache and LangSmith init files each 20x at 0 failures, make lint and make check pass

2026-09-07

No new flake in the window. Every failed CI test in the last 24 hours either failed on every attempt (real breakage, reported in Slack) or was infra. The one same-SHA fail-then-pass, tests/documentation_tests/test_env_keys.py on https://github.com/BerriAI/litellm/actions/runs/34026362388 (attempt 1 failed, attempt 2 passed), was the docs job checking out the head of BerriAI/litellm-docs: the CLF AI Gateway docs page landed there (litellm-docs#1242, 10:05 UTC) between the two attempts. That is the cross-repo dependency working as designed, not a flaky test, so nothing was changed for it

Branch re-merged with litellm_internal_staging (9a79691). The test-quality-budget.json ratchet from the first run was reverted in ae2f1ae: staging now says the scheduled automation owns that file and PR branches must not touch it (#39937). Neither fixed test flaked in the window, and both passed without rerun on the previous tip 7d93d9d

2026-09-06

tests/test_litellm/responses/litellm_completion_transformation/test_session_handler.py::test_message_history_looks_up_the_decoded_chat_completion_id

Same flake, one more occurrence, in responses-caching-types / Run tests (Python 3.10) (--reruns 2), RERUN then PASSED inside attempt 1. The run's head (sha 80b1d4b, branch litellm_ocr_core_ownership) does not contain this PR's fix, and its gw0 worker ran the whole of test_redis_semantic_cache.py before the session handler test, which is the leak order described below:

No new fix; covered by the redis semantic cache change in this PR. Branch re-merged with litellm_internal_staging (7d93d9d) and the gates re-run: 20x loops of both fixed tests at 0 failures, owning files green, make lint and make check pass

2026-09-05

tests/test_litellm/responses/litellm_completion_transformation/test_session_handler.py::test_message_history_looks_up_the_decoded_chat_completion_id

Flaked 3 times in the window, each time in responses-caching-types / Run tests (Python 3.10) (--reruns 2), RERUN then PASSED inside attempt 1 of the run:

Mechanism: shared module state leaking across tests, ordering dependent. tests/test_litellm/caching/test_redis_semantic_cache.py wrapped the first import of litellm.caching.redis_semantic_cache in patch.dict("sys.modules", {...redisvl fakes...}). patch.dict snapshots the whole dict on enter and restores it on exit, so every module first imported inside the block (including litellm.proxy.proxy_server, pulled in transitively) was removed from sys.modules while staying cached as the proxy_server attribute on the already imported litellm.proxy package. The session handler test then did patch("litellm.proxy.proxy_server.prisma_client", fake), which resolves through the stale package attribute, while ResponsesSessionHandler did from litellm.proxy.proxy_server import prisma_client inside the function, which re-imported a fresh module from sys.modules with prisma_client = None. The fake never saw the query, so assert fake_prisma_client.db.calls == [(request_id,)] failed with [] == [...]. The re-import also re-synced the package attribute, which is why the rerun passed. Reproduced locally by running the redis file's tests before the session handler test in the same worker order as CI (gw1): 1 failed before the fix, 379 passed after

Fix: _fake_redisvl_modules uses pytest.MonkeyPatch.context() with setitem on exactly the two redisvl keys, and the remaining redisvl.query.filter case uses the test's monkeypatch fixture. Nothing else in sys.modules is touched or restored

Mutation check: with response_id lookup changed to use the raw encoded id, the session handler test fails; with the redis cache filter_expression set to None, 7 redis tests fail. Both mutations reverted

tests/test_litellm/integrations/test_langsmith_init.py::TestLangsmithLoggerInit::test_langsmith_init_starts_periodic_flush_with_running_loop

Flaked once in the window, in integrations / Run tests, RERUN then PASSED inside attempt 1:

Mechanism: a global patch racing the cyclic GC. The test patched asyncio.get_running_loop for the whole process while constructing LangsmithLogger. AsyncHTTPHandler.__del__ calls asyncio.get_running_loop() and, when it gets a loop, loop.create_task(client.aclose()). When the cyclic collector finalized an orphaned handler left behind by an earlier test inside the patched window, that create_task landed on the same mock_loop, and mock_loop.create_task.assert_called_once() failed with Called 2 times (the extra call is AsyncClient.aclose). Reproduced locally by forcing a GC of an orphaned handler during LangsmithLogger.__init__: exact same assertion message as CI

Fix: the test is now async, so a real running loop exists. It builds the logger with flush_interval=0.01, queues one event, replaces async_send_batch on that instance with an AsyncMock that sets an asyncio.Event, and waits on that event (5s ceiling) to prove the task init scheduled is the periodic flusher doing its job. Then it cancels and awaits the task. No asyncio patching, so unrelated finalizers cannot interfere

Mutation check: _start_periodic_flush_task returning None fails the isinstance assert; scheduling asyncio.sleep(0) instead of periodic_flush times out waiting for the batch send. Both mutations reverted

Gates

Re-run on 441f863 (2026-09-12): test_redis_semantic_cache.py 20x, the LangSmith periodic flush test 20x plus its file, test_scheduled_job_stagger.py 20x, the eight files that use FakePrismaCli 20x (575 tests per pass), the whole tests/test_litellm/proxy/db/ directory once (1013 tests) and tests/test_litellm/caching/ once, all with 0 failures. The pgbouncer test ran 20x idle and 20x under CPU load with 0 failures on the upstream fix. make lint passes, make check passes (test-only scope)

Review loop

Greptile's one P2 (exact __qualname__ assertion is structural) was fixed in 8a7dc64 by asserting on the flush itself. Greptile re-reviewed the tip at 5/5 and resolved its thread. Bugbot reviewed 8a7dc64 and found no issues

Rolling PR bookkeeping

No other open litellm_deflake_ PRs existed at run time, so nothing was absorbed or closed. The 2026-09-06 through 2026-09-12 runs found this PR to be the only open litellm_deflake_ PR again and rolled their evidence in here. Dropped on 2026-09-12: the test_pgbouncer.py readiness budget change, because #40830 landed the same fix on staging first

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Test-only change with no proxy behavior to demonstrate end to end, so there is no User Flow for /qa to drive at the merge base vs the tip: booting a proxy at either commit answers every request identically because no file under litellm/ changed. The evidence is the CI run links above (first attempt RERUN, then PASSED on the same sha) plus the local reproductions described per flake

/live-pr-risk: the diff touches four test files, so the changed-symbol dependency graph is empty (no callers, overrides, registry lookups, response fields, config keys, or schema). Nothing to A/B live

Type

✅ Test

Caveats (if any)

Low

  • _wait_ready still says "served by another process" when its own child bound TCP but not yet the Unix socket by the deadline; the retry behavior is the same either way, only the log line misleads. Not changed here since this PR no longer touches pgbouncer
  • _is_zombie reads /proc, so on macOS a zombie still counts as alive; the fake CLI is a POSIX shebang script and CI is Linux
  • 24 other test files still use whole-dict patch.dict("sys.modules", ...); same leak class, not touched here
  • test_start_periodic_flush_task_returns_none_without_running_loop still patches asyncio.get_running_loop globally; did not flake in the window, left alone
  • test_operator_supplied_cron_keeps_its_exact_schedule and test_disabling_the_stagger_leaves_every_schedule_untouched still compare two datetime.now() based schedules; the window is microseconds around 03:00 and 00:15 UTC, no flake seen, left alone

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/1d1da9d617b942b6862ed9f10a682847
Open in Devin Desktop: https://app.devin.ai/desktop/session/1d1da9d617b942b6862ed9f10a682847?variant=devin
Requested by: @mateo-berri


Note

Low Risk
Only test harness and assertion timing change; no runtime proxy, auth, or data-path code is modified.

Overview
Test-only deflake PR — no production code changes.

Redis semantic cache tests stop using whole-dict patch.dict("sys.modules", …) and instead inject fakes for only the two redisvl modules via a shared _fake_redisvl_modules helper (MonkeyPatch.setitem), plus monkeypatch for redisvl.query.filter, so imports like litellm.proxy.proxy_server are not dropped from sys.modules between tests.

The LangSmith periodic-flush test runs under a real asyncio loop: it uses a short flush_interval, stubs async_send_batch to signal an Event, and waits for an actual flush instead of globally patching asyncio.get_running_loop.

Scheduled job stagger’s explicit-offset test is now synchronous and compares staggered vs unstaggered PTU cron triggers with _fire_times from a fixed UTC start (and a small _trigger_of helper), avoiding wall-clock sensitivity around the 00:15 cron boundary.

Proxy DB test conftest treats zombie grandchildren as gone by reading /proc/<pid>/stat when os.kill(pid, 0) still succeeds, fixing flaky grandchild_is_gone timeouts under subreapers that haven’t reaped yet.

Reviewed by Cursor Bugbot for commit 441f863. Bugbot is set up for automated code reviews on this repo. Configure here.

Link to Devin session: https://app.devin.ai/sessions/5e0925839b2d4a9b94790b03156e3036
Open in Devin Desktop: https://app.devin.ai/desktop/session/5e0925839b2d4a9b94790b03156e3036?variant=devin

…ed on CI rerun

The Redis semantic cache tests wrapped the first import of litellm.caching.redis_semantic_cache in patch.dict("sys.modules", ...), which snapshots and restores all of sys.modules on exit. Every module first imported inside the block, including litellm.proxy.proxy_server, was dropped from sys.modules while staying cached as an attribute on the litellm.proxy package. The next test that patched litellm.proxy.proxy_server.<attr> hit the stale attribute while production code re-imported a fresh module, so the patch never reached it. Replace the whole-dict patch with MonkeyPatch.setitem on the two redisvl keys only

The LangSmith init test globally patched asyncio.get_running_loop while constructing the logger. Any orphaned AsyncHTTPHandler finalized by the cyclic GC during that window also called loop.create_task on the mock, tripping assert_called_once. Run the test under a real event loop and assert on the real task instead of patching asyncio

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 5, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR deflakes four test areas without changing production behavior:

  • Replaces whole-dictionary sys.modules patching with targeted RedisVL module stubs
  • Exercises LangSmith periodic flushing on a real event loop
  • Compares staggered cron triggers from a fixed timestamp
  • Treats killed but unreaped Linux zombie grandchildren as gone

Confidence Score: 5/5

The PR appears safe to merge, with no outstanding correctness, security, or repository-rule issues

The test changes preserve behavioral regression coverage while removing shared-module, event-loop, wall-clock, and zombie-reaping sources of nondeterminism. The previous structural coroutine assertion was fixed and its thread was resolved

Important Files Changed

Filename Overview
tests/test_litellm/caching/test_redis_semantic_cache.py Limits RedisVL stubbing to specific module entries so teardown does not remove unrelated imports
tests/test_litellm/integrations/test_langsmith_init.py Verifies periodic flushing through observable batch delivery on a real event loop
tests/test_litellm/proxy/common_utils/test_scheduled_job_stagger.py Removes wall-clock dependence by comparing cron trigger results from a fixed UTC start
tests/test_litellm/proxy/db/conftest.py Recognizes Linux zombie grandchildren as already terminated in process-tree assertions

Reviews (6): Last reviewed commit: "Merge remote-tracking branch 'origin/lit..." | Re-trigger Greptile

Comment thread tests/test_litellm/integrations/test_langsmith_init.py Outdated
Replace the coroutine __qualname__ check with a functional check: queue one event, run with a short flush interval, and wait for async_send_batch to be awaited by the task init scheduled

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@codecov

codecov Bot commented Sep 5, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

mateo-berri and others added 3 commits September 6, 2026 09:18
…mation owns it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

mateo-berri and others added 2 commits September 8, 2026 09:19
…all clock

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title test: deflake redis semantic cache sys.modules leak and LangSmith init loop patch test: deflake redis semantic cache sys.modules leak, LangSmith init loop patch, and wall-clock stagger assertion Sep 8, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

mateo-berri and others added 2 commits September 9, 2026 09:17
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title test: deflake redis semantic cache sys.modules leak, LangSmith init loop patch, and wall-clock stagger assertion test: deflake redis semantic cache sys.modules leak, LangSmith init loop patch, wall-clock stagger assertion, and zombie grandchild check in the fake prisma cli Sep 9, 2026
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

1 similar comment
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

mateo-berri and others added 2 commits September 11, 2026 09:18
…CI worker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title test: deflake redis semantic cache sys.modules leak, LangSmith init loop patch, wall-clock stagger assertion, and zombie grandchild check in the fake prisma cli test: deflake redis semantic cache sys.modules leak, LangSmith init loop patch, wall-clock stagger assertion, zombie grandchild check in the fake prisma cli, and pgbouncer fake pooler readiness budget Sep 11, 2026
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…itellm_deflake_20260905

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	tests/test_litellm/proxy/db/test_pgbouncer.py
@devin-ai-integration devin-ai-integration Bot changed the title test: deflake redis semantic cache sys.modules leak, LangSmith init loop patch, wall-clock stagger assertion, zombie grandchild check in the fake prisma cli, and pgbouncer fake pooler readiness budget test: deflake redis semantic cache sys.modules leak, LangSmith init loop patch, wall-clock stagger assertion, and zombie grandchild check in the fake prisma cli Sep 12, 2026
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 441f863. Configure here.

@mateo-berri mateo-berri left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit d72ae3b into litellm_internal_staging Sep 12, 2026
85 checks passed
@mateo-berri
mateo-berri deleted the litellm_deflake_20260905 branch September 12, 2026 22:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant