fix(web): join dashboard eager-reconcile thread on shutdown - #113265
JoaoMarcos44 wants to merge 1 commit into
Conversation
Keep the statedb-eager-reconcile handle and join it in lifespan teardown before hosted-room stop so TestClient portals cannot leave sqlite workers alive into interpreter finalization.
|
Summary What changed
Strengths
Findings
Verdict Reviewed using Hermes-Agent |
|
Thanks @JoaoMarcos44 — this was the right diagnosis and the right mechanism (lifespan owns the The same fix landed on Closing as superseded; the runner-side half of #113186 (reporting the segfault as a crash rather than "0 failed, no tests ran") is #115587. Apologies that the credit didn't route through a cherry-pick here — the priority is noted on the issue. |
Summary
Dashboard lifespan now keeps the
statedb-eager-reconcilethread handle and joins it on shutdown, before hosted-room teardown. A TestClient/lifespan portal can no longer leave native sqlite workers running into pytest/interpreter finalization.Fixes #113186
Observed problem
tests/hermes_cli/test_web_profiles_off_loop.py(and any other file that starts the dashboard lifespan repeatedly) intermittently dies withFatal Python error: Segmentation faultin sqlite while CI prints0 failedplus1 file where no tests ran. The faulthandler dump shows manystatedb-eager-reconcilethreads in_open_session_db_at_pathwhile the main thread is already inpytest_sessionfinish.statedb-eager-reconcilethread remains, so later tests and interpreter teardown cannot racesqlite3_step/ connection init.origin/main(fdb7b216ecat publication prep; worktree then rebased onto it):_lifespanstartsthreading.Thread(..., daemon=True, name="statedb-eager-reconcile")without retaining the object, andfinallyjoinshosted-room-startupbut never the reconcile worker.Root cause
The dashboard is a view layer and must not delay the socket, so schema reconcile is a daemon thread started before the lifespan yield (
#79531/#80037, GH-73083). That is still correct for startup.The missing owner is shutdown.
TestClientas a context manager runs lifespanfinallyat portal exit. Hosted-room recovery is cancelled and joined there; the reconcile worker is not. Each test therefore leaks a daemon thread that can still be insidecheck_same_thread=Falsesqlite while:state.db, andpytest_sessionfinish).That is the segfault shape in the issue, not a test-local assertion failure. A per-store lock around read-only
SessionDBinit does not stop the worker from outliving the app.Impact and blast radius
TestClient(app)repeatedly (test_web_profiles_off_loop.py,test_web_server.py, and siblings). Production clean shutdown ofhermes dashboard/ Desktopservealso waits up to 5s for reconcile to finish before hosted-room stop, which avoids a sqlite race withstop_hosted_room_service.#112075). Fresh-home hosted-room boot/SIGBUS (#104828)._STATEDB_EAGER_RECONCILE_JOIN_TIMEOUT_S(5.0s) if the worker is still in sqlite.Solution
In
hermes_cli/web_server.py::_lifespan:statedb_eager_reconcile_threadinstead ofThread(...).start().finallywith a 5.0s bound before hosted-room cancel/stop/join, so two sqlite users are not tearing down at once.daemon=Trueso a wedged store still cannot pin Desktop bind (Desktop startup fails with GIL stall on Windows — _warm_gateway_module() import blocks event loop 15-22s #73083).Alternatives considered and rejected:
#113235): does not join the worker; adds a second lock registry besidehermes_state_registry; leaves the interpreter-finalization race if the thread outlives the portal.#104828's hosted-room Event ordering /connect_trackedSIGBUS work: different bug class; this PR only closes the outliving reconcile worker._eager_reconcile_own_session_dbor making the thread non-daemon: would either delay bind or change heal semantics.Related work (not duplicated)
#113235(KoNit-K) — lock around read-only SessionDB init. Incomplete for this issue: the reporter's required owner is shutdown join. This PR does not use that lock or that test.#104828(Alish3r) — fresh-homestate.dbboot/SIGBUS with hosted-room tracking. Complementary. This PR does not take its hosted-room or sqlite-tracking files.#112075(teknium1) — test leak sweep must notSessionDB.close()a live probe thread. Complementary; different files. Join-on-shutdown does not replace a safe sweep.No issue/PR comments were posted on those threads.
Compatibility and residual risks
joinreturns and a daemon worker can still exist. That bound matches "do not hang shutdown"; a wedged sqlite lock remains a residual process-lifetime issue.local-runtime-bootis still fire-and-forget. It is not the sqlite path in the dump.Segmentation faultinrun_tests_parallel.pyis intentionally out of scope. The real fix is that the worker must not outlive the app.Tests and verification
All commands used
scripts/run_tests.shwithHERMES_PYTHON=C:/Users/Nitro/hermes-agent/.venv/Scripts/python.exefrom an isolated worktree. Do not treat the full suite as run.scripts/run_tests.sh tests/hermes_cli/test_web_server_reconcile_shutdown.py -qon unfixedweb_server.py(HEAD production file restored)statedb-eager-reconcilethread(s) after TestClient exit; sequential case had two live daemonsassert _alive_reconcile_threads()inside the portal (worker already finished during slow lifespan startup); retry passed. Tests were then rewritten so the in-portal assertion holds the worker on an Event. Not used as final green evidence.scripts/run_tests.sh tests/hermes_cli/test_web_server_reconcile_shutdown.py tests/hermes_cli/test_web_server_boot_handshake.py tests/hermes_cli/test_web_profiles_off_loop.py -qon final rebased head52d82ad3eescripts/run_tests.sh tests/hermes_cli/test_web_server.py -k 'startup_eager_reconcile' -qruff check hermes_cli/web_server.py tests/hermes_cli/test_web_server_reconcile_shutdown.pygit diff --checkscripts/run_tests.shgraphify updategraphify-out/graph.jsonabsentFiles and documentation
hermes_cli/web_server.py— retain and joinstatedb-eager-reconcileon lifespan shutdown.tests/hermes_cli/test_web_server_reconcile_shutdown.py— behavior contracts: worker is alive during the portal; sequential portals do not leave named threads after shutdown.No changelog/docs change; this is an internal lifecycle join.
Publication identity
fix/113186-join-eager-reconcile-on-shutdown52d82ad3eec6a9fdbefe6b4abea84078cde3fa90origin/mainfdb7b216ecdde7ebe8409c2b81d74685784afa62C:/Users/Nitro/hermes-agentcheckout left onmainwith its pre-existing untracked files untouched