fix(state): dashboard and hermes sessions list no longer fail with 'database is locked' in DELETE journal mode - #120386
Merged
Merged
Conversation
Under rollback-journal (DELETE) mode a reader needs a SHARED lock, and every commit from another process blocks it across its journal+db fsyncs. Read-only SessionDB handles were opened with a 1 s busy timeout, and writer handles serve DELETE-mode reads on the writer connection, whose 1 s timeout exists for writes (they retry at application level). Under a busy gateway, readers gave up after 1 s: - dashboard GET /api/sessions -> 503 "Session store is busy", or 500 when the lock surfaced through the FTS5 vtable constructor during the read-only open probe; - `hermes sessions list` -> "Could not open your session history database. Run: hermes sessions repair" (a healthy store) or a raw `database is locked` traceback; - in-process reads on a gateway/CLI writer handle -> `database is locked`. Reads now get the same 5 s SQLite busy budget the WAL read pool already used (_READ_BUSY_TIMEOUT_S): read-only handles are opened with it, and a DELETE-mode read on the writer connection raises busy_timeout for the read and restores it after, so writes keep their short timeout and the jittered application-level retry.
૮ >ﻌ< ა ci reviewran on 0a044d1 — fix(state): DELETE-journal readers wait out another process' debug infoCI timingsCI timings · View report · View jobWall time 6m37s vs 8m (-17.3%). 6 job(s) slower, 6 faster,
|
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
After #120386 raised the read-only busy timeout to 5 s, retrying a lock inside _open_read_only multiplied the wait to ~20 s on blocking callers (TUI profile loop, exit epilogue, hermes status). The connection already waited the read budget; only transient disk-I/O errors are retried now. Probe (30 s exclusive DELETE-mode lock): 20.17 s -> 5.0 s, still classified as a transient lock.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
After #120386 raised the read-only busy timeout to 5 s, retrying a lock inside _open_read_only multiplied the wait to ~20 s on blocking callers (TUI profile loop, exit epilogue, hermes status). The connection already waited the read budget; only transient disk-I/O errors are retried now. Probe (30 s exclusive DELETE-mode lock): 20.17 s -> 5.0 s, still classified as a transient lock.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
State.db readers in DELETE (rollback-journal) mode now wait out another process's commit instead of failing with
database is locked. This affected the dashboard's/api/sessions,hermes sessions list, and reads on a gateway's or CLI's own handle.Changes
hermes_state.py: new_READ_BUSY_TIMEOUT_S = 5.0. This is the busy budget the WAL read pool already used; the pool now takes it from the constant too.SessionDBhandles (dashboard,hermes sessions list/stats/pinned, cross-profile readers) are opened with it. Before, they got the 1 s write timeout._read_ctx's locked path) raisesbusy_timeoutfor that read and restores the previous value afterward. Writes keep their 1 s timeout and the jittered application-level retry in_execute_write.tests/hermes_state/test_delete_mode_read_busy_wait.py: a separate process holdsBEGIN EXCLUSIVEfor 2 s on a DELETE-mode DB. The test checks that a read-only handle and a writer handle both finishlist_sessions_richandsession_count, and that the writer connection'sbusy_timeoutis back to 1000 afterward. It fails onorigin/main(2 failed,database is lockedfrom_fts_table_probeand_read_all) and passes with the fix.Root cause: in DELETE mode, a reader needs a SHARED lock, and every commit from another process blocks that lock through its journal and db fsyncs. Reads ran with a 1 s busy timeout that was sized for writes, which retry on their own.
Live repro (multi-process, real surfaces)
Setup: temp
HOME/HERMES_HOMEwithdatabase.journal_mode: delete(the production path; each child asserts_wal_active is FalseandPRAGMA journal_mode == delete). N writer processes run realSessionDB.append_messagewith a 20 ms pace. Alongside them:hermes_cli.web_server.appon uvicorn) polled over HTTP at/api/sessions?order=recent;hermes sessions listsubprocesses, one in flight at a time;SessionDB(the gateway shape):list_sessions_rich,session_count,get_messages./api/sessionshermes sessions listdatabase is lockedtracebackThe base 500 is the same lock showing up as
vtable constructor failed: messages_ftsduring the read-only open probe. The 503 and the CLI failure told users to repair a store that was healthy. With the fix, the slowest dashboard request under load was 1.8–3.0 s; on base the timed-out requests failed at about 1.5 s.Validation
scripts/run_tests.sh tests/hermes_state: 1108 passed, 0 failed.test_hermes_state.pyhit the per-file timeout under host load in the directory run, so it was run alone: 275 passed.tests/hermes_cli/test_sessions*.py+test_web_server_sessions*: 41 passed. Everytests/hermes_clifile that exercises/api/sessions: 352 passed.check_compat_pointers,check-windows-footguns,check_no_tmp_literals,git diff --check: clean.browser_execrefuses loopback addresses, so the dashboard was measured over HTTP against the real app.Related: #102621 (@ialmeida-jera) retries the session-list page/pinned/count statements 2× after the busy timeout expires. This PR fixes the cause at the connection layer, and that covers every read on read-only handles, including the open probe that #102621 does not reach. Please assess #102621 for supersession after this lands.
Infographic