Skip to content

fix(state): dashboard and hermes sessions list no longer fail with 'database is locked' in DELETE journal mode - #120386

Merged
teknium1 merged 1 commit into
mainfrom
fix/delete-mode-read-busy
Sep 23, 2026
Merged

teknium1 merged 1 commit into
mainfrom
fix/delete-mode-read-busy

Conversation

@teknium1

Copy link
Copy Markdown
Collaborator

State.db readers in DELETE (rollback-journal) mode now wait out another process's commit instead of failing with database is locked. This affected the dashboard's /api/sessions, hermes sessions list, and reads on a gateway's or CLI's own handle.

Changes

  • hermes_state.py: new _READ_BUSY_TIMEOUT_S = 5.0. This is the busy budget the WAL read pool already used; the pool now takes it from the constant too.
    • Read-only SessionDB handles (dashboard, hermes sessions list/stats/pinned, cross-profile readers) are opened with it. Before, they got the 1 s write timeout.
    • A DELETE-mode read on the writer connection (_read_ctx's locked path) raises busy_timeout for that read and restores the previous value afterward. Writes keep their 1 s timeout and the jittered application-level retry in _execute_write.
  • tests/hermes_state/test_delete_mode_read_busy_wait.py: a separate process holds BEGIN EXCLUSIVE for 2 s on a DELETE-mode DB. The test checks that a read-only handle and a writer handle both finish list_sessions_rich and session_count, and that the writer connection's busy_timeout is back to 1000 afterward. It fails on origin/main (2 failed, database is locked from _fts_table_probe and _read_all) and passes with the fix.

Root cause: in DELETE mode, a reader needs a SHARED lock, and every commit from another process blocks that lock through its journal and db fsyncs. Reads ran with a 1 s busy timeout that was sized for writes, which retry on their own.

Live repro (multi-process, real surfaces)

Setup: temp HOME/HERMES_HOME with database.journal_mode: delete (the production path; each child asserts _wal_active is False and PRAGMA journal_mode == delete). N writer processes run real SessionDB.append_message with a 20 ms pace. Alongside them:

  • the real dashboard app (hermes_cli.web_server.app on uvicorn) polled over HTTP at /api/sessions?order=recent;
  • real hermes sessions list subprocesses, one in flight at a time;
  • one process that reads through its own writer SessionDB (the gateway shape): list_sessions_rich, session_count, get_messages.
run (30 s) dashboard /api/sessions hermes sessions list writer-handle reads writer errors
base, 3 writers 87×200, 8×503, 1×500 12 ok, 1 failed: "Could not open your session history database. Run: hermes sessions repair" — 0
base, 3 writers (rerun) 87×200, 2×503 12 ok, 1 raw database is locked traceback 160 ok, 3 locked 0
base, 4 writers 83×200, 2×503 17 ok, 1 failed 115 ok, 3 locked 0
fix, 3 writers 144×200 10/10 ok 108 ok, 0 locked 0
fix, 3 writers (rerun) 207×200 10/10 ok 123 ok, 0 locked 0
fix, 4 writers 62×200 16/16 ok 108 ok, 0 locked 0

The base 500 is the same lock showing up as vtable constructor failed: messages_fts during the read-only open probe. The 503 and the CLI failure told users to repair a store that was healthy. With the fix, the slowest dashboard request under load was 1.8–3.0 s; on base the timed-out requests failed at about 1.5 s.

Validation

  • scripts/run_tests.sh tests/hermes_state: 1108 passed, 0 failed. test_hermes_state.py hit the per-file timeout under host load in the directory run, so it was run alone: 275 passed.
  • tests/hermes_cli/test_sessions*.py + test_web_server_sessions*: 41 passed. Every tests/hermes_cli file that exercises /api/sessions: 352 passed.
  • ruff, check_compat_pointers, check-windows-footguns, check_no_tmp_literals, git diff --check: clean.
  • Not done: a browser (SPA) pass. browser_exec refuses loopback addresses, so the dashboard was measured over HTTP against the real app.

Related: #102621 (@ialmeida-jera) retries the session-list page/pinned/count statements 2× after the busy timeout expires. This PR fixes the cause at the connection layer, and that covers every read on read-only handles, including the open probe that #102621 does not reach. Please assess #102621 for supersession after this lands.

Infographic

infographic

Under rollback-journal (DELETE) mode a reader needs a SHARED lock, and
every commit from another process blocks it across its journal+db fsyncs.
Read-only SessionDB handles were opened with a 1 s busy timeout, and
writer handles serve DELETE-mode reads on the writer connection, whose
1 s timeout exists for writes (they retry at application level). Under a
busy gateway, readers gave up after 1 s:

- dashboard GET /api/sessions -> 503 "Session store is busy", or 500 when
  the lock surfaced through the FTS5 vtable constructor during the
  read-only open probe;
- `hermes sessions list` -> "Could not open your session history
  database. Run: hermes sessions repair" (a healthy store) or a raw
  `database is locked` traceback;
- in-process reads on a gateway/CLI writer handle -> `database is locked`.

Reads now get the same 5 s SQLite busy budget the WAL read pool already
used (_READ_BUSY_TIMEOUT_S): read-only handles are opened with it, and a
DELETE-mode read on the writer connection raises busy_timeout for the
read and restores it after, so writes keep their short timeout and the
jittered application-level retry.
@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 0a044d1 — fix(state): DELETE-journal readers wait out another process'

debug info

CI timings

CI timings · View report · View job

Wall time 6m37s vs 8m (-17.3%). 6 job(s) slower, 6 faster,

  • Python lints / Windows footguns (blocking): +88.0s
  • OS-specific tests / Windows-only tests: -56.0s
  • Check no committed infographics / check-no-committed-infographics: -47.0s
  • Profile artifact check / Reject profile archives: -45.0s
  • Python tests / Run tests: +43.0s

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/cli CLI entry point, hermes_cli/, setup wizard comp/dashboard Web dashboard / control panel UI (dashboard/, landing) comp/gateway Gateway runner, session dispatch, delivery area/sessions Session lifecycle, resume, persistence, history sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades labels Sep 23, 2026
@teknium1 teknium1 added the ci-reviewed applied to manually approve dangerous changes label Sep 23, 2026
@teknium1
teknium1 merged commit c871622 into main Sep 23, 2026
38 checks passed
@teknium1
teknium1 deleted the fix/delete-mode-read-busy branch September 23, 2026 17:25
teknium1 added a commit that referenced this pull request Sep 23, 2026
After #120386 raised the read-only busy timeout to 5 s, retrying a lock
inside _open_read_only multiplied the wait to ~20 s on blocking callers
(TUI profile loop, exit epilogue, hermes status). The connection already
waited the read budget; only transient disk-I/O errors are retried now.

Probe (30 s exclusive DELETE-mode lock): 20.17 s -> 5.0 s, still
classified as a transient lock.
teknium1 added a commit that referenced this pull request Sep 23, 2026
After #120386 raised the read-only busy timeout to 5 s, retrying a lock
inside _open_read_only multiplied the wait to ~20 s on blocking callers
(TUI profile loop, exit epilogue, hermes status). The connection already
waited the read budget; only transient disk-I/O errors are retried now.

Probe (30 s exclusive DELETE-mode lock): 20.17 s -> 5.0 s, still
classified as a transient lock.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/sessions Session lifecycle, resume, persistence, history ci-reviewed applied to manually approve dangerous changes comp/cli CLI entry point, hermes_cli/, setup wizard comp/dashboard Web dashboard / control panel UI (dashboard/, landing) comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants