Skip to content

web dashboard: state.db reconcile worker no longer segfaults the test interpreter (joined at lifespan shutdown) - #113552

Merged
teknium1 merged 1 commit into
mainfrom
fix/web-profiles-off-loop-exit-segfault
Sep 17, 2026
Merged

teknium1 merged 1 commit into
mainfrom
fix/web-profiles-off-loop-exit-segfault

Conversation

@teknium1

Copy link
Copy Markdown
Collaborator

The dashboard's startup state.db reconcile worker is now joined by the lifespan, so it can no longer have its sqlite connection closed from another thread — the test_web_profiles_off_loop.py segfault is gone.

Symptom

tests/hermes_cli/test_web_profiles_off_loop.py ..............Fatal Python error: Segmentation fault on CI (PR #113430, run 35153362037 attempt 1) and ~1/4 locally, after all 18 tests had passed.

Root cause

_lifespan started _eager_reconcile_own_session_db on a daemon thread and never joined it; the CI faulthandler dump shows the worker inside conn.execute(...) in web_server_sessions._open_probed while the main thread's autouse _close_leaked_session_dbs teardown called SessionDB.close() on that same connection — a cross-thread sqlite3_close on a statement being stepped. Ten earlier tests' workers were still queued on _session_db_bootstrap_lock (each TestClient context bootstraps a fresh tmp store; slow on the runner), so the leak compounded across the file.

Change

  • hermes_cli/web_server.py::_lifespan — the statedb-eager-reconcile worker is a regular (non-daemon) thread held by the lifespan and join()ed in the shutdown finally. Startup is unchanged (the open still runs off the ready-probe path, Desktop startup fails with GIL stall on Windows — _warm_gateway_module() import blocks event loop 15-22s #73083); the join is bounded by SessionDB's write patience, so shutdown cannot hang on it.
  • tests/hermes_cli/test_web_server_boot_handshake.py::test_lifespan_shutdown_joins_statedb_reconcile_worker — invariant: the worker is off the startup path AND is finished (no statedb-eager-reconcile thread alive) once the TestClient context exits. Red on origin/main, green here.

Twin sweep (hermes_cli/web_server_*.py, tui_gateway/*.py): the hosted-room start thread is already stopped+joined; the Desktop cron ticker is a stop-event-driven long-lived loop whose join could block on a running job — different class, left alone.

Validation

origin/main this PR
Stress probe (reconcile made slow, conftest-style sweep during probe), 12 runs 9/12 SIGSEGV, 12/12 closed a live conn cross-thread 0/12 SIGSEGV, 0 live conns after lifespan
scripts/run_tests.sh tests/hermes_cli/test_web_profiles_off_loop.py ×10 10/10 green locally (race needs a slow runner to fire) 10/10 green
New invariant test FAIL (finished.is_set() False) PASS
tests/hermes_cli/test_web_server.py + off-loop/boot-handshake/eventloop files — 220 passed

Live repro: before — probe /tmp/batchbots/flakefix/probe_exit_segv.py on origin/main: Current thread … web_server_sessions.py line 143 in _open_probed + Fatal Python error: Segmentation fault (9/12), identical to the CI trace; after — same probe 0/12, registry live conns: 0.

The probe inflates _session_db_read_probe_statements with a heavy recursive CTE so the worker is deterministically mid-step when the sweep runs; it is a diagnostic, not a committed test.

Infographic

Reconcile worker joined at shutdown

The dashboard lifespan started `_eager_reconcile_own_session_db` on a
daemon thread and never joined it. Under pytest each TestClient context
spawned one; the fresh tmp HERMES_HOME store makes every worker take the
bootstrap path, so on a slow CI runner ten of them were still queued on
`_session_db_bootstrap_lock` when the test's autouse leaked-DB sweep ran
`SessionDB.close()` on the connection the live worker was stepping in
`_open_probed` -> cross-thread `sqlite3_close` on an active statement ->
`Fatal Python error: Segmentation fault` after every test had passed
(PR #113430 run 35153362037, attempt 1; ~1/4 locally).

The worker is now a regular (non-daemon) thread that the lifespan
`finally` joins, so its connection is only ever closed by the thread that
opened it and it cannot outlive the server or the interpreter. Startup is
unchanged (the open still happens off the ready-probe path); the join is
bounded by SessionDB's write patience, so shutdown cannot hang on it.
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 2838c97 — fix(web): join the state.db eager-reconcile worker at lifesp

debug info

CI timings

CI timings · View report · View job

Wall time 6m7s vs 6m (+1.9%). 8 job(s) slower, 3 faster, 1 unchanged.

  • OS-specific tests / Windows-only tests: +26.0s
  • Python tests / Run tests: +9.0s
  • OS-specific tests / macOS-only tests: +5.0s
  • Python tests / e2e: +4.0s
  • Check no case-colliding filenames / check-case-collisions: +4.0s

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/dashboard Web dashboard / control panel UI (dashboard/, landing) area/sessions Session lifecycle, resume, persistence, history labels Sep 16, 2026
@alt-glitch

Copy link
Copy Markdown

This was generated by AI during triage.

Supersedes #113265 (same mechanism: join the statedb-eager-reconcile worker at lifespan shutdown). Related: #113186 (the flake), #112075 and #113235 (test-leak / lock-side approaches to the same segfault).

@kyssta-exe

Copy link
Copy Markdown

Summary

Fixes the test_web_profiles_off_loop.py interpreter segfault by joining the state.db eager-reconcile worker at lifespan shutdown instead of leaving it as a fire-and-forget daemon thread. The worker still starts off the ready-probe path, so startup latency is unchanged. A new invariant test pins the behavior.

What changed

  • hermes_cli/web_server.py: statedb-eager-reconcile thread is now non-daemon, held in eager_reconcile_thread, and join()ed on the shutdown path.
  • tests/hermes_cli/test_web_server_boot_handshake.py: new test_lifespan_shutdown_joins_statedb_reconcile_worker asserts startup returns before the worker finishes and that no such thread survives TestClient exit.
  • Checks: full CI green on this branch (Windows/macOS, ruff, lints).

Strengths

  • Root cause is precisely diagnosed (cross-thread sqlite3_close while stepping, compounded by queued bootstrap workers) and validated with a stress probe (9/12 SIGSEGV before, 0/12 after).
  • Minimal, surgical diff (+39/−5); comment explains why non-daemon + join is safe (bounded by SessionDB write patience).
  • Regression test asserts both halves of the invariant (off startup path AND joined at shutdown).

Findings

  • No blocking issues. Two non-blocking polish tips: (1) consider a defensive join(timeout=...) plus log rather than an unbounded join, so a future unbounded reconcile path can never stall shutdown; (2) the join reads naturally in the shutdown finally, but confirm it runs even when startup itself raises before the thread handle is assigned (guard with locals()/None check if that path exists).

Verdict

Looks good to merge
Reviewed using Hermes-Agent

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/sessions Session lifecycle, resume, persistence, history comp/dashboard Web dashboard / control panel UI (dashboard/, landing) P2 Medium — degraded but workaround exists type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants