Repository navigation
test: deflake redis semantic cache sys.modules leak, LangSmith init loop patch, wall-clock stagger assertion, and zombie grandchild check in the fake prisma cli - #39895
Conversation
…ed on CI rerun
The Redis semantic cache tests wrapped the first import of litellm.caching.redis_semantic_cache in patch.dict("sys.modules", ...), which snapshots and restores all of sys.modules on exit. Every module first imported inside the block, including litellm.proxy.proxy_server, was dropped from sys.modules while staying cached as an attribute on the litellm.proxy package. The next test that patched litellm.proxy.proxy_server.<attr> hit the stale attribute while production code re-imported a fresh module, so the patch never reached it. Replace the whole-dict patch with MonkeyPatch.setitem on the two redisvl keys only
The LangSmith init test globally patched asyncio.get_running_loop while constructing the logger. Any orphaned AsyncHTTPHandler finalized by the cyclic GC during that window also called loop.create_task on the mock, tripping assert_called_once. Run the test under a real event loop and assert on the real task instead of patching asyncio
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
Greptile SummaryThis PR deflakes four test areas without changing production behavior:
Confidence Score: 5/5The PR appears safe to merge, with no outstanding correctness, security, or repository-rule issues The test changes preserve behavioral regression coverage while removing shared-module, event-loop, wall-clock, and zombie-reaping sources of nondeterminism. The previous structural coroutine assertion was fixed and its thread was resolved
|
| Filename | Overview |
|---|---|
| tests/test_litellm/caching/test_redis_semantic_cache.py | Limits RedisVL stubbing to specific module entries so teardown does not remove unrelated imports |
| tests/test_litellm/integrations/test_langsmith_init.py | Verifies periodic flushing through observable batch delivery on a real event loop |
| tests/test_litellm/proxy/common_utils/test_scheduled_job_stagger.py | Removes wall-clock dependence by comparing cron trigger results from a fixed UTC start |
| tests/test_litellm/proxy/db/conftest.py | Recognizes Linux zombie grandchildren as already terminated in process-tree assertions |
Reviews (6): Last reviewed commit: "Merge remote-tracking branch 'origin/lit..." | Re-trigger Greptile
Replace the coroutine __qualname__ check with a functional check: queue one event, run with a short flush interval, and wait for async_send_batch to be awaited by the task init scheduled Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
bugbot run |
…itellm_deflake_20260905
…itellm_deflake_20260905
…mation owns it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
…itellm_deflake_20260905
…all clock Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
…itellm_deflake_20260905
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
1 similar comment
|
bugbot run |
…itellm_deflake_20260905
…itellm_deflake_20260905
…CI worker Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
…itellm_deflake_20260905 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> # Conflicts: # tests/test_litellm/proxy/db/test_pgbouncer.py
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 441f863. Configure here.
TLDR
Problem this solves:
litellm.proxy.proxy_serverfromsys.modulesasyncio.get_running_loopglobally, GC finalizers hit the mockHow it solves it:
redisvlmodules withMonkeyPatch.setitem, not all ofsys.modulesdatetime.now()/proc/<pid>/statand treat stateZas gone, since a zombie has already been killedUser Flow
Not applicable. This PR only changes unit tests; proxy behavior is unchanged
Relevant issues
Linear ticket
Flakes covered, by run date
2026-09-12
tests/test_litellm/proxy/db/test_pgbouncer.py::TestPgBouncerProcess::test_a_replacement_that_never_listens_is_replaced_again
Flaked 20 times in the window (2026-09-11 09:20 to 2026-09-12 09:20 UTC, 16,181 workflow runs scanned, 652 failed job logs parsed), always in
Unit Tests / proxy-infra / Run testswith the sameassert pooler.start() is NonereportingPgBouncerError(reason='127.0.0.1:<port> is served by another process, not the pgbouncer that was started'). Same-sha evidence: https://github.com/BerriAI/litellm/actions/runs/34659643270 (attempt 1 failed, attempt 2 passed), plus failures onlitellm_internal_stagingitself at https://github.com/BerriAI/litellm/actions/runs/34652853654 and https://github.com/BerriAI/litellm/actions/runs/34648862502 where the next staging run passed. Same mechanism as the 2026-09-11 entry belowEvery one of the 20 failures ran before #40830 merged into staging at 02:32 UTC on 2026-09-12 with
ready_timeout_seconds=2.0and the log assertion moved after the port file rewrite. That is the same fix this branch carried since 2026-09-11 at 3.0s, so the branch change was dropped in favor of the staging version when staging was merged in (441f863). The diff of this PR no longer touchestest_pgbouncer.pyVerification of the upstream fix on the merged branch: the test ran 20 times idle and 20 times under 16 busy-loop processes on 8 cores (load average about 17) with 0 failures, and the whole
test_pgbouncer.pyfile passes both serially and under-n 4One other same-sha flip was not a staging flake:
tests/test_litellm_rust/ocr/test_callbacks.py::test_native_azure_ocr_releases_token_provider_after_terminal_outcome[cancellation]failed on https://github.com/BerriAI/litellm/actions/runs/34638756281 and passed on https://github.com/BerriAI/litellm/actions/runs/34638764815 at sha 7b9d0fe, but that sha is on a feature branch that rewrites the Azure credential provider the test holds a weakref to, so it is that branch's own work in progress. Real breakages seen on every attempt, confined to feature branches and left for their authors:tests/test_litellm_rust/test_ocr.py('RecordedOCRRequest' object is not subscriptable) onlitellm_ocr_azure_mistralandlitellm_ocr_mistral_native, the retained-callbacks suite onlitellm_rust_retained_callbacks,test_rust_bridge.py(no attribute '_run_rust_ocr') onlitellm_ocr_completion_boundary, andtest_aaamodel_prices_and_context_window_json_is_validon the price-sync branches (output_cost_per_video_per_second_batchesnot in the schema). Infra, not test debt: 15ai-gateway release imagejobs failing withNo such container: ai-gatewaythen passing on retry, 14e2e-changed-testsapproval-gate flips, oneproxy-infrarunner abort mid-run (https://github.com/BerriAI/litellm/actions/runs/34646984186), and one osv-scan flip2026-09-11
tests/test_litellm/proxy/db/test_pgbouncer.py::TestPgBouncerProcess::test_a_replacement_that_never_listens_is_replaced_again
Flaked 8 times in the window (2026-09-10 09:15 to 2026-09-11 09:15 UTC, 15,430 workflow runs scanned), every time in
Unit Tests / proxy-infra / Run testson attempt 1, where pytest's own--reruns 2ran it three times in a row and it failed all three withassert pooler.start() is NonereportingPgBouncerError(reason='127.0.0.1:<port> is served by another process, not the pgbouncer that was started'). Seven of the eight runs passed on attempt 2 of the same sha, the eighth has no attempt 2:Mechanism: wall-clock dependence on interpreter start time. The test builds a
PgBouncerProcesswithready_timeout_seconds=0.3so the later replacement that binds the wrong port gets rejected quickly, but the same 0.3s budget also applies to the first, healthystart(). The fake pooler is a Python script, and in this shard the child interpreter also loads the pytest-cov subprocess hook (COVERAGE_CORE=sysmon,--cov) before it binds anything: 0.19s on an idle box here, more on a 4-worker xdist runner sharing 2 vCPUs with other subprocess-spawning db tests._wait_readypolls every 0.1s for TCP and Unix socket together, so when the child gets its TCP bind in between the last in-budget poll and the post-deadline check, or has bound TCP but not yet the Unix socket, the supervisor reports the port as held by a stranger andstart()returns an error. Reproduced locally without edits by running the test with--covunder 64 busy-loop processes: 2 failures in 10 runs (did not start listening within 0s, the other branch of the same deadline). Idle, 0 failures in 20Fix (since superseded by #40830, see 2026-09-12):
ready_timeout_seconds=3.0, matchingtest_a_listener_that_grabs_the_port_after_the_spawn_is_not_taken_for_the_poolerin the same class, and a 10s ceiling on the one_wait_untilthat now has to outlast that budget plus the restart delay before the "did not start listening" log line appears. Every other step still waits on a condition (port listening, port closed, log record present), not a duration. The test takes about 3.5s instead of 0.5sMutation check: with
_retry_restartno longer starting the supervisor thread, the test fails at_wait_until(lambda: _listening(port))because nothing replaces the rejected pooler. RevertedNot touched, reported as a real breakage instead:
tests/test_litellm/proxy/anthropic_endpoints/test_claude_code_marketplace.py::test_archive_source_registers_and_is_served_verbatim_in_marketplacefails on every attempt withTypeError: get_marketplace() missing 1 required positional argument: 'request', and the Terraform Provider job keeps failing on every attempt as noted under 2026-09-10. Fail-then-pass on Documentation validation, Helm, CodSpeed, Conventional PR Title and OSV Scan were network, cancellation or advisory timing, not test flakes2026-09-10
No new flake in the window (2026-09-09 09:20 to 2026-09-10 09:20 UTC, 17,493 workflow runs scanned). The only same-sha fail-then-pass on a test was
ui-unit-testson PR #40367 sha c59b321:add_auto_router_tab.test.tsx > carries a preset's per-tier reasoning effort through to the create payloadfailed on https://github.com/BerriAI/litellm/actions/runs/34328740942 (attempt 1, 08:28 UTC) and passed on https://github.com/BerriAI/litellm/actions/runs/34430807521 (attempt 1, 02:52 UTC the next day). That is not a flake. Both runs arepull_requestevents, so the job checks out the merge of the PR head into staging, and staging moved between them: #40341 repointed the Anthropic REASONING preset atclaude-fable-5-1while the test still hardcodedclaude-opus-5, then #40456 (2c836f4, 21:14 UTC) made the assertion read the preset. The second run merged in that fix. Nothing to change hereEvery other fail-then-pass was infra: nine
e2e-changed-testsruns whose first attempt was concurrency-cancelled mid detection (DETECT_RESULT: cancelled), one cancelledConventional PR Titlerun, and an osv-scan that started failing when GHSA-7w5x-hrqm-74c2 (smol-toml) was published between two runs of the same sha. One real breakage, reported in Slack rather than touched here: the Terraform Provider job failsTestResourceKeyUpdateFailureKeepsPriorState(resource_key_test.go:356: failed update persisted budget_duration="bad" into state) on every attempt of every PR into staging since #40512 added it, for example https://github.com/BerriAI/litellm/actions/runs/34434802755 and https://github.com/BerriAI/litellm/actions/runs/34459965452None of the four tests this PR fixes failed anywhere in the window. Branch re-merged with
litellm_internal_staging(c956c24)2026-09-09
tests/test_litellm/proxy/db/test_check_migration.py::test_migrate_diff_stops_at_its_budget_and_takes_its_process_tree_with_it
Flaked 8 times in the window, every time in
Unit Testson attempt 1 withassert fake_prisma_cli.grandchild_is_gone(within_seconds=5)failing asassert False, then PASSED on attempt 2 of the same sha:Mechanism: the test helper depended on who reaps orphans.
run_prismadoes kill the whole process group on timeout (SIGKILL viakillpg), so the fake CLI's grandchild dies. But the grandchild's parent (the fake CLI) died in the same kill, so the dead grandchild is reparented to the nearest subreaper, and until that ancestor callswaiton it, the process stays a zombie.FakePrismaCli.grandchild_is_gonechecked liveness withos.waitpid(fails withChildProcessError, the pid is not our child) andos.kill(pid, 0), which succeeds on a zombie. So the helper reported "still alive" for as long as the ancestor took to reap, and whether that fits in 5 seconds depends on the runner's process tree, not on the code under test. Not reproducible in a plain shell where init reaps at once; reproduced 100% locally by running pytest under a parent that setsPR_SET_CHILD_SUBREAPERand never reaps, which is the shape of a CI runner's job wrapper. The same run also failedtests/test_litellm/proxy/db/test_prisma_client.py::test_db_push_timeout_takes_its_process_tree_with_it, which shares the helper. Production code is correct, the kill happens; only the observation was wrongFix:
grandchild_is_gonealso reads/proc/<pid>/statand returns true when the state field isZ. A zombie has already been killed, which is what the assertion is about. Nothing else in the helper or the tests changed, and the 5 second window and the assertion stay as they wereMutation check: with
_kill_process_groupchanged toprocess.kill()(kill only the CLI, not the group), both process-tree tests fail because the grandchild keeps sleeping and is neither gone nor a zombie. Mutation revertedBranch re-merged with
litellm_internal_staging(db4dee5). Gates on fdd6f60: both process-tree tests pass under the non-reaping subreaper wrapper, each ran 20x in a loop at 0 failures,test_check_migration.py,test_prisma_client.pyand the wholetests/test_litellm/proxy/db/directory pass,make lintandmake checkpassNot touched this run: a same-sha fail-then-pass cluster in
tests/test_litellm/proxy/_experimental/mcp_server/test_mcp_proxy_mode.py(authorization tests seeing an empty tool set, e.g.assert set() == {'beta-add', 'alpha-add'}). It was not reproduced locally and the mechanism was not pinned from the logs, so it is reported as unreproduced rather than guessed at2026-09-08
tests/test_litellm/proxy/common_utils/test_scheduled_job_stagger.py::test_explicit_offset_overrides_the_derived_one_and_zero_pins_a_job
Flaked once in the window, in
Unit Tests / proxy-infra / Run tests, RERUN twice then FAILED on attempt 1, PASSED on attempt 2 of the same sha:Mechanism: wall-clock dependence. The test built two schedulers and read
job.next_run_timeafterscheduler.start(paused=True), which evaluates the PTU rollup cron (00:15 UTC daily) againstdatetime.now(). Between 00:15:00 and 00:15:07 UTC the unstaggered cron has already rolled to tomorrow, while the +7s offset trigger evaluates the base cron atnow - 7s, which is still today, and yields today 00:15:07. The assertion then saw2026-09-08 00:15:07 - 2026-09-09 00:15instead of 7 seconds. Reproduced locally underfaketime '2026-09-09 00:14:55'with the identical assertion shape; the production code behaves correctly, only the test's reference point movedFix: the test now takes the PTU trigger off each scheduler and compares
_fire_times(trigger, start, 1)from a fixedstart, the same waytest_default_cron_is_staggered_and_keeps_its_offset_on_every_later_firealready did. A small_trigger_of(scheduler, job_id)helper replaces the duplicatednext(job.trigger for ...)in both tests. The test no longer awaits anything, so it is a plaindefMutation check:
_OffsetTrigger.get_next_fire_timereturning the base fire time without+ self.offsetfails the test (00:15 - 00:15 == 7s);_clamped_overridereturning 0 fails it (assert 0 == 7). Both mutations revertedBranch re-merged with
litellm_internal_staging(1af7a40). Gates re-run on 61e6640: the fixed test 20x at 0 failures, the owning file 22 passed, the redis semantic cache and LangSmith init files each 20x at 0 failures,make lintandmake checkpass2026-09-07
No new flake in the window. Every failed CI test in the last 24 hours either failed on every attempt (real breakage, reported in Slack) or was infra. The one same-SHA fail-then-pass,
tests/documentation_tests/test_env_keys.pyon https://github.com/BerriAI/litellm/actions/runs/34026362388 (attempt 1 failed, attempt 2 passed), was the docs job checking out the head of BerriAI/litellm-docs: the CLF AI Gateway docs page landed there (litellm-docs#1242, 10:05 UTC) between the two attempts. That is the cross-repo dependency working as designed, not a flaky test, so nothing was changed for itBranch re-merged with
litellm_internal_staging(9a79691). Thetest-quality-budget.jsonratchet from the first run was reverted in ae2f1ae: staging now says the scheduled automation owns that file and PR branches must not touch it (#39937). Neither fixed test flaked in the window, and both passed without rerun on the previous tip 7d93d9d2026-09-06
tests/test_litellm/responses/litellm_completion_transformation/test_session_handler.py::test_message_history_looks_up_the_decoded_chat_completion_id
Same flake, one more occurrence, in
responses-caching-types / Run tests (Python 3.10)(--reruns 2), RERUN then PASSED inside attempt 1. The run's head (sha 80b1d4b, branch litellm_ocr_core_ownership) does not contain this PR's fix, and its gw0 worker ran the whole oftest_redis_semantic_cache.pybefore the session handler test, which is the leak order described below:No new fix; covered by the redis semantic cache change in this PR. Branch re-merged with
litellm_internal_staging(7d93d9d) and the gates re-run: 20x loops of both fixed tests at 0 failures, owning files green,make lintandmake checkpass2026-09-05
tests/test_litellm/responses/litellm_completion_transformation/test_session_handler.py::test_message_history_looks_up_the_decoded_chat_completion_id
Flaked 3 times in the window, each time in
responses-caching-types / Run tests (Python 3.10)(--reruns 2), RERUN then PASSED inside attempt 1 of the run:Mechanism: shared module state leaking across tests, ordering dependent.
tests/test_litellm/caching/test_redis_semantic_cache.pywrapped the first import oflitellm.caching.redis_semantic_cacheinpatch.dict("sys.modules", {...redisvl fakes...}).patch.dictsnapshots the whole dict on enter and restores it on exit, so every module first imported inside the block (includinglitellm.proxy.proxy_server, pulled in transitively) was removed fromsys.moduleswhile staying cached as theproxy_serverattribute on the already importedlitellm.proxypackage. The session handler test then didpatch("litellm.proxy.proxy_server.prisma_client", fake), which resolves through the stale package attribute, whileResponsesSessionHandlerdidfrom litellm.proxy.proxy_server import prisma_clientinside the function, which re-imported a fresh module fromsys.moduleswithprisma_client = None. The fake never saw the query, soassert fake_prisma_client.db.calls == [(request_id,)]failed with[] == [...]. The re-import also re-synced the package attribute, which is why the rerun passed. Reproduced locally by running the redis file's tests before the session handler test in the same worker order as CI (gw1): 1 failed before the fix, 379 passed afterFix:
_fake_redisvl_modulesusespytest.MonkeyPatch.context()withsetitemon exactly the tworedisvlkeys, and the remainingredisvl.query.filtercase uses the test'smonkeypatchfixture. Nothing else insys.modulesis touched or restoredMutation check: with
response_idlookup changed to use the raw encoded id, the session handler test fails; with the redis cachefilter_expressionset toNone, 7 redis tests fail. Both mutations revertedtests/test_litellm/integrations/test_langsmith_init.py::TestLangsmithLoggerInit::test_langsmith_init_starts_periodic_flush_with_running_loop
Flaked once in the window, in
integrations / Run tests, RERUN then PASSED inside attempt 1:Mechanism: a global patch racing the cyclic GC. The test patched
asyncio.get_running_loopfor the whole process while constructingLangsmithLogger.AsyncHTTPHandler.__del__callsasyncio.get_running_loop()and, when it gets a loop,loop.create_task(client.aclose()). When the cyclic collector finalized an orphaned handler left behind by an earlier test inside the patched window, thatcreate_tasklanded on the samemock_loop, andmock_loop.create_task.assert_called_once()failed withCalled 2 times(the extra call isAsyncClient.aclose). Reproduced locally by forcing a GC of an orphaned handler duringLangsmithLogger.__init__: exact same assertion message as CIFix: the test is now
async, so a real running loop exists. It builds the logger withflush_interval=0.01, queues one event, replacesasync_send_batchon that instance with anAsyncMockthat sets anasyncio.Event, and waits on that event (5s ceiling) to prove the task init scheduled is the periodic flusher doing its job. Then it cancels and awaits the task. Noasynciopatching, so unrelated finalizers cannot interfereMutation check:
_start_periodic_flush_taskreturningNonefails the isinstance assert; schedulingasyncio.sleep(0)instead ofperiodic_flushtimes out waiting for the batch send. Both mutations revertedGates
Re-run on 441f863 (2026-09-12):
test_redis_semantic_cache.py20x, the LangSmith periodic flush test 20x plus its file,test_scheduled_job_stagger.py20x, the eight files that useFakePrismaCli20x (575 tests per pass), the wholetests/test_litellm/proxy/db/directory once (1013 tests) andtests/test_litellm/caching/once, all with 0 failures. The pgbouncer test ran 20x idle and 20x under CPU load with 0 failures on the upstream fix.make lintpasses,make checkpasses (test-only scope)Review loop
Greptile's one P2 (exact
__qualname__assertion is structural) was fixed in 8a7dc64 by asserting on the flush itself. Greptile re-reviewed the tip at 5/5 and resolved its thread. Bugbot reviewed 8a7dc64 and found no issuesRolling PR bookkeeping
No other open
litellm_deflake_PRs existed at run time, so nothing was absorbed or closed. The 2026-09-06 through 2026-09-12 runs found this PR to be the only openlitellm_deflake_PR again and rolled their evidence in here. Dropped on 2026-09-12: thetest_pgbouncer.pyreadiness budget change, because #40830 landed the same fix on staging firstPre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Test-only change with no proxy behavior to demonstrate end to end, so there is no User Flow for /qa to drive at the merge base vs the tip: booting a proxy at either commit answers every request identically because no file under
litellm/changed. The evidence is the CI run links above (first attempt RERUN, then PASSED on the same sha) plus the local reproductions described per flake/live-pr-risk: the diff touches four test files, so the changed-symbol dependency graph is empty (no callers, overrides, registry lookups, response fields, config keys, or schema). Nothing to A/B live
Type
✅ Test
Caveats (if any)
Low
_wait_readystill says "served by another process" when its own child bound TCP but not yet the Unix socket by the deadline; the retry behavior is the same either way, only the log line misleads. Not changed here since this PR no longer touches pgbouncer_is_zombiereads/proc, so on macOS a zombie still counts as alive; the fake CLI is a POSIX shebang script and CI is Linuxpatch.dict("sys.modules", ...); same leak class, not touched heretest_start_periodic_flush_task_returns_none_without_running_loopstill patchesasyncio.get_running_loopglobally; did not flake in the window, left alonetest_operator_supplied_cron_keeps_its_exact_scheduleandtest_disabling_the_stagger_leaves_every_schedule_untouchedstill compare twodatetime.now()based schedules; the window is microseconds around 03:00 and 00:15 UTC, no flake seen, left aloneFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/1d1da9d617b942b6862ed9f10a682847
Open in Devin Desktop: https://app.devin.ai/desktop/session/1d1da9d617b942b6862ed9f10a682847?variant=devin
Requested by: @mateo-berri
Note
Low Risk
Only test harness and assertion timing change; no runtime proxy, auth, or data-path code is modified.
Overview
Test-only deflake PR — no production code changes.
Redis semantic cache tests stop using whole-dict
patch.dict("sys.modules", …)and instead inject fakes for only the tworedisvlmodules via a shared_fake_redisvl_moduleshelper (MonkeyPatch.setitem), plusmonkeypatchforredisvl.query.filter, so imports likelitellm.proxy.proxy_serverare not dropped fromsys.modulesbetween tests.The LangSmith periodic-flush test runs under a real asyncio loop: it uses a short
flush_interval, stubsasync_send_batchto signal anEvent, and waits for an actual flush instead of globally patchingasyncio.get_running_loop.Scheduled job stagger’s explicit-offset test is now synchronous and compares staggered vs unstaggered PTU cron triggers with
_fire_timesfrom a fixed UTC start (and a small_trigger_ofhelper), avoiding wall-clock sensitivity around the 00:15 cron boundary.Proxy DB test
conftesttreats zombie grandchildren as gone by reading/proc/<pid>/statwhenos.kill(pid, 0)still succeeds, fixing flakygrandchild_is_gonetimeouts under subreapers that haven’t reaped yet.Reviewed by Cursor Bugbot for commit 441f863. Bugbot is set up for automated code reviews on this repo. Configure here.
Link to Devin session: https://app.devin.ai/sessions/5e0925839b2d4a9b94790b03156e3036
Open in Devin Desktop: https://app.devin.ai/desktop/session/5e0925839b2d4a9b94790b03156e3036?variant=devin