fix(state): SessionDB open waits out a DELETE-mode lock reported as 'vtable constructor failed' instead of failing - #120488
Merged
Merged
Conversation
૮ >ﻌ< ა ci reviewran on dc01ffe — fix(state): SessionDB open waits out a lock lost inside the debug infoCI timingsCI timings · View report · View jobWall time 5m46s vs 5m57s (-3.1%). 5 job(s) slower, 7 faster, 1 unchanged.
|
…ructor In rollback-journal (DELETE) mode a sibling process can take the write lock between schema load and the messages_fts probe. FTS5's xConnect then fails its %_config read and SQLite reports SQLITE_BUSY with the text "vtable constructor failed: messages_fts". Every state.db lock classifier matched on the words "locked"/"busy", so: - a writable SessionDB() failed after 1s instead of waiting out the lock with _WRITE_PATIENCE_S, and callers disabled persistence for the run; - a read-only open (dashboard, `hermes sessions list`, cross-profile readers) failed on the first busy timeout with no retry at all; - the error read as not transient (dashboard 500, not 503) and as persistence cause "unknown" instead of "locked". Add hermes_state_errors.is_sqlite_lock_error: SQLITE_BUSY/SQLITE_LOCKED by result code when SQLite supplies one, text only when it does not (our own re-raised messages, RPC-wrapped strings). Route the writer open patience loop, the _execute_write retry, the reconcile re-raise, the WAL->DELETE flip, the maintenance holder probe, is_transient_sqlite_error and classify_persistence_error through it. The read-only open retries a lock inside its existing bounded retry budget, next to the transient IOERR case.
After #120386 raised the read-only busy timeout to 5 s, retrying a lock inside _open_read_only multiplied the wait to ~20 s on blocking callers (TUI profile loop, exit epilogue, hermes status). The connection already waited the read budget; only transient disk-I/O errors are retried now. Probe (30 s exclusive DELETE-mode lock): 20.17 s -> 5.0 s, still classified as a transient lock.
teknium1
force-pushed
the
fix/sessiondb-open-busy-vtable
branch
from
September 23, 2026 18:04
dc01ffe to
d1ae72d
Compare
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
The FTS fail-open detach now waits up to the caller's write budget (20 s / 60 s) for the write lock, so the one-time quarantine check before the loop left a long window: a sibling that quarantined the file meanwhile still got its triggers dropped and the stale breadcrumb committed on the quarantined handle. Re-check the handle flag and the process-wide storage latch at the top of every attempt, via the same _raise_if_db_corrupt(storage=True) that _execute_write runs per attempt. Classify the retryable lock error with is_sqlite_lock_error (result code first) instead of a locked/busy substring match, matching #120488.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
The FTS fail-open detach now waits up to the caller's write budget (20 s / 60 s) for the write lock, so the one-time quarantine check before the loop left a long window: a sibling that quarantined the file meanwhile still got its triggers dropped and the stale breadcrumb committed on the quarantined handle. Re-check the handle flag and the process-wide storage latch at the top of every attempt, via the same _raise_if_db_corrupt(storage=True) that _execute_write runs per attempt. Classify the retryable lock error with is_sqlite_lock_error (result code first) instead of a locked/busy substring match, matching #120488.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A
SessionDBopen, writable or read-only, now waits out another process's lock in DELETE journal mode even when SQLite words the lock asvtable constructor failed: messages_fts. Before this change the open failed after 1 s, persistence was disabled for the run, and the dashboard answered 500.Live repro: before:
SessionDB()raisedOperationalError: vtable constructor failed: messages_fts(SQLITE_BUSY) after 1.06 s. The read-only open raised after 1.02 s. The error counted as not transient and its cause asunknown. After: both opens wait for the 2.5 s lock, then succeed (2.66 s / 2.55 s) withfts_enabled=Trueand the seeded search hit. If a 7 s lock outlasts the bounded read-only budget, the open raises after 4.24 s and the error counts as transient (503) with causelocked.Root cause
When FTS5's table constructor loses the lock while reading
messages_fts_config, SQLite keeps result codeSQLITE_BUSYbut replaces the message. Every state.db lock classifier looked for the words "locked" or "busy" in the message.Changes
hermes_state_errors.py: newis_sqlite_lock_error. It checks forSQLITE_BUSY/SQLITE_LOCKEDby result code when SQLite supplies one. Only when there is no code (our own re-raised messages, RPC-wrapped strings) does it fall back to the text.is_transient_sqlite_error(dashboard 503 vs 500) andclassify_persistence_error(lockedbucket) use the result code too.hermes_state.py:_execute_writeretry now use the helper._open_read_onlynow retries a lock inside its existing bounded retry budget, alongside the transient IOERR case. Each retry waits the busy timeout again. Before, a lock on a read-only open was never retried._reconcile_columnsre-raise (hermes_state_schema.py)hermes_state_wal.py)hermes_state_holders.py)developer-guide/session-storage.mdnow says lock contention is classified by result code and describes open patience for both open kinds.SessionDB()SessionDB(read_only=True)unknownlockedTests
New tests in
tests/hermes_state/test_write_lock_patience.py. They are red onorigin/main(3 failed) and green with the fix:test_open_waits_out_lock_lost_inside_fts_constructor[writer|read_only]: real DELETE-modeSessionDB. A second connection takesBEGIN EXCLUSIVEfor 2.5 s at the moment the open reaches themessages_ftsprobe. The open must succeed with FTS enabled and the transcript readable.test_lock_lost_inside_fts_constructor_classifies_as_busy: a realvtable constructor failederror from SQLite must be transient and classify aslocked.No existing test was changed. Ran
tests/hermes_state/,tests/conformance/persistence/,tests/hermes_cli/test_web_server.pyand the other test files that use the touched predicates: 1251 passed. Two timing tests failed under host load (~155) and pass when re-run alone. Also clean: ruff,check_no_tmp_literals,check-windows-footguns --all,git diff --check.Related
hermes sessions listno longer fail with 'database is locked' in DELETE journal mode #120386 (open) raises the busy timeout for DELETE-mode reads. This PR is complementary: it covers thevtable constructor failedwording and the app-level retry. Swept for competing PRs (fix(state): converge concurrent WAL initialization #101384, fix(state): retry contended session list queries #102621, fix(state): enforce write patience during lock acquisition #79140); none classify lock errors by result code on the open path.Infographic