Skip to content

fix(state): SessionDB open waits out a DELETE-mode lock reported as 'vtable constructor failed' instead of failing - #120488

Merged
teknium1 merged 2 commits into
mainfrom
fix/sessiondb-open-busy-vtable
Sep 23, 2026
Merged

teknium1 merged 2 commits into
mainfrom
fix/sessiondb-open-busy-vtable

Conversation

@teknium1

Copy link
Copy Markdown
Collaborator

A SessionDB open, writable or read-only, now waits out another process's lock in DELETE journal mode even when SQLite words the lock as vtable constructor failed: messages_fts. Before this change the open failed after 1 s, persistence was disabled for the run, and the dashboard answered 500.

Live repro: before: SessionDB() raised OperationalError: vtable constructor failed: messages_fts (SQLITE_BUSY) after 1.06 s. The read-only open raised after 1.02 s. The error counted as not transient and its cause as unknown. After: both opens wait for the 2.5 s lock, then succeed (2.66 s / 2.55 s) with fts_enabled=True and the seeded search hit. If a 7 s lock outlasts the bounded read-only budget, the open raises after 4.24 s and the error counts as transient (503) with cause locked.

Root cause

When FTS5's table constructor loses the lock while reading messages_fts_config, SQLite keeps result code SQLITE_BUSY but replaces the message. Every state.db lock classifier looked for the words "locked" or "busy" in the message.

Changes

  • hermes_state_errors.py: new is_sqlite_lock_error. It checks for SQLITE_BUSY / SQLITE_LOCKED by result code when SQLite supplies one. Only when there is no code (our own re-raised messages, RPC-wrapped strings) does it fall back to the text. is_transient_sqlite_error (dashboard 503 vs 500) and classify_persistence_error (locked bucket) use the result code too.
  • hermes_state.py:
    • The writer open's lock-patience loop and the _execute_write retry now use the helper.
    • _open_read_only now retries a lock inside its existing bounded retry budget, alongside the transient IOERR case. Each retry waits the busy timeout again. Before, a lock on a read-only open was never retried.
  • The rest of the class moved to the same helper:
    • the _reconcile_columns re-raise (hermes_state_schema.py)
    • the WAL→DELETE flip (hermes_state_wal.py)
    • the maintenance lock probe (hermes_state_holders.py)
  • Docs: developer-guide/session-storage.md now says lock contention is classified by result code and describes open patience for both open kinds.
Open under a DELETE-mode lock (2.5 s hold) before after
writable SessionDB() raises after 1.06 s opens after 2.66 s
SessionDB(read_only=True) raises after 1.02 s opens after 2.55 s
error once the budget is exhausted not transient (500), cause unknown transient (503), cause locked

Tests

New tests in tests/hermes_state/test_write_lock_patience.py. They are red on origin/main (3 failed) and green with the fix:

  • test_open_waits_out_lock_lost_inside_fts_constructor[writer|read_only]: real DELETE-mode SessionDB. A second connection takes BEGIN EXCLUSIVE for 2.5 s at the moment the open reaches the messages_fts probe. The open must succeed with FTS enabled and the transcript readable.
  • test_lock_lost_inside_fts_constructor_classifies_as_busy: a real vtable constructor failed error from SQLite must be transient and classify as locked.

No existing test was changed. Ran tests/hermes_state/, tests/conformance/persistence/, tests/hermes_cli/test_web_server.py and the other test files that use the touched predicates: 1251 passed. Two timing tests failed under host load (~155) and pass when re-run alone. Also clean: ruff, check_no_tmp_literals, check-windows-footguns --all, git diff --check.

Related

Infographic

infographic

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on dc01ffe — fix(state): SessionDB open waits out a lock lost inside the

debug info

CI timings

CI timings · View report · View job

Wall time 5m46s vs 5m57s (-3.1%). 5 job(s) slower, 7 faster, 1 unchanged.

  • Check contributors / check-attribution: +216.0s
  • Python lints / Windows footguns (blocking): +29.0s
  • Check no case-colliding filenames / check-case-collisions: -19.0s
  • OS-specific tests / Windows-only tests: +16.0s
  • Check no committed infographics / check-no-committed-infographics: +10.0s

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint area/sessions Session lifecycle, resume, persistence, history sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state labels Sep 23, 2026
@teknium1 teknium1 added the ci-reviewed applied to manually approve dangerous changes label Sep 23, 2026
…ructor

In rollback-journal (DELETE) mode a sibling process can take the write lock
between schema load and the messages_fts probe. FTS5's xConnect then fails its
%_config read and SQLite reports SQLITE_BUSY with the text "vtable constructor
failed: messages_fts". Every state.db lock classifier matched on the words
"locked"/"busy", so:

- a writable SessionDB() failed after 1s instead of waiting out the lock with
  _WRITE_PATIENCE_S, and callers disabled persistence for the run;
- a read-only open (dashboard, `hermes sessions list`, cross-profile readers)
  failed on the first busy timeout with no retry at all;
- the error read as not transient (dashboard 500, not 503) and as persistence
  cause "unknown" instead of "locked".

Add hermes_state_errors.is_sqlite_lock_error: SQLITE_BUSY/SQLITE_LOCKED by
result code when SQLite supplies one, text only when it does not (our own
re-raised messages, RPC-wrapped strings). Route the writer open patience loop,
the _execute_write retry, the reconcile re-raise, the WAL->DELETE flip, the
maintenance holder probe, is_transient_sqlite_error and
classify_persistence_error through it. The read-only open retries a lock
inside its existing bounded retry budget, next to the transient IOERR case.
After #120386 raised the read-only busy timeout to 5 s, retrying a lock
inside _open_read_only multiplied the wait to ~20 s on blocking callers
(TUI profile loop, exit epilogue, hermes status). The connection already
waited the read budget; only transient disk-I/O errors are retried now.

Probe (30 s exclusive DELETE-mode lock): 20.17 s -> 5.0 s, still
classified as a transient lock.
@teknium1
teknium1 force-pushed the fix/sessiondb-open-busy-vtable branch from dc01ffe to d1ae72d Compare September 23, 2026 18:04
@teknium1
teknium1 merged commit 16fe260 into main Sep 23, 2026
34 checks passed
@teknium1
teknium1 deleted the fix/sessiondb-open-busy-vtable branch September 23, 2026 18:35
teknium1 added a commit that referenced this pull request Sep 23, 2026
The FTS fail-open detach now waits up to the caller's write budget (20 s /
60 s) for the write lock, so the one-time quarantine check before the loop
left a long window: a sibling that quarantined the file meanwhile still got
its triggers dropped and the stale breadcrumb committed on the quarantined
handle. Re-check the handle flag and the process-wide storage latch at the
top of every attempt, via the same _raise_if_db_corrupt(storage=True) that
_execute_write runs per attempt.

Classify the retryable lock error with is_sqlite_lock_error (result code
first) instead of a locked/busy substring match, matching #120488.
teknium1 added a commit that referenced this pull request Sep 23, 2026
The FTS fail-open detach now waits up to the caller's write budget (20 s /
60 s) for the write lock, so the one-time quarantine check before the loop
left a long window: a sibling that quarantined the file meanwhile still got
its triggers dropped and the stale breadcrumb committed on the quarantined
handle. Re-check the handle flag and the process-wide storage latch at the
top of every attempt, via the same _raise_if_db_corrupt(storage=True) that
_execute_write runs per attempt.

Classify the retryable lock error with is_sqlite_lock_error (result code
first) instead of a locked/busy substring match, matching #120488.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/sessions Session lifecycle, resume, persistence, history ci-reviewed applied to manually approve dangerous changes comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P2 Medium — degraded but workaround exists sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants