Skip to content

fix(state): a busy write lock no longer loses the message that hit a corrupt search index - #120534

Merged
teknium1 merged 2 commits into
mainfrom
fix/fts-fail-open-lock-patience
Sep 23, 2026
Merged

teknium1 merged 2 commits into
mainfrom
fix/fts-fail-open-lock-patience

Conversation

@teknium1

@teknium1 teknium1 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

A message written while the session search index is corrupt is no longer lost when another Hermes process (gateway + TUI, gateway + CLI) holds the state.db write lock at the moment the index gets detached.

Changes

  • hermes_state_fts.py::_enter_fts_fail_open: the detach (stale breadcrumb + FTS trigger drop) now waits out database is locked / busy with the same jittered retry as _execute_write, on the caller's write budget. The default is _WRITE_PATIENCE_S, used by the search fail-open callers.
  • hermes_state.py::_execute_write passes its own deadline / patience_s into the detach, so a transcript write keeps its 60 s budget.
  • Every retry iteration re-checks this handle's quarantine flag and the process-wide storage latch (_raise_if_db_corrupt(storage=True), the same check _execute_write runs per attempt). A sibling that quarantines the file while the detach waits for the lock now stops it: no trigger drop and no stale breadcrumb are committed on a quarantined handle, and StateDbCorruptError surfaces.
  • Lock errors are classified with hermes_state_errors.is_sqlite_lock_error (result code first, text only as fallback), consistent with fix(state): SessionDB open waits out a DELETE-mode lock reported as 'vtable constructor failed' instead of failing #120488, instead of a locked/busy substring match.
  • Behaviour change for session search: the two session_search fail-open call sites in hermes_state_search.py (_match_rows and the messages_fts MATCH arm of _search_messages_impl) pass no budget, so their detach now inherits the default _WRITE_PATIENCE_S (20 s) of write-lock patience instead of giving up after the 1 s busy timeout. A search that hits a corrupt index while a sibling holds the write lock can therefore block up to 20 s before falling back to canonical LIKE, rather than raising.
  • tests/hermes_state/test_fts_index_fail_open.py: two new invariant tests, plus the existing _enter_fts_fail_open stub updated for the new keyword arguments.

Root cause: the detach ran a single BEGIN IMMEDIATE on the writer connection, whose busy timeout is only 1 s, and gave up on database is locked. The canonical write then escaped as database disk image is malformed. The usual lock holder is a sibling writer detaching the same corrupt index, so under load the second writer lost its turn.

Validation

Check Base (origin/main) This PR
Live probe: a second process takes BEGIN IMMEDIATE right as the corruption surfaces and holds it 2.5 s 3/3 lost the write: DatabaseError: fts5: corrupt structure record after 1.02 s 3/3 row landed after the holder released, FTS detached
E2E torture chamber fts_corruption_fail_open (#120171) with a 3 s pre-detach window and a 1.5 s detach hold injected via a scratch copy, WAL + DELETE arms, 2 runs each 4/4 FAILED: Could not detach corrupt FTS indexes …: database is locked → DatabaseError('database disk image is malformed'), same as the union-run flake 4/4 passed; the second writer detached ~1.5–7 s later instead of giving up
New test test_detach_waits_out_a_sibling_holding_the_write_lock FAILED passed
New test test_quarantine_while_detach_waits_commits_nothing (review follow-up) FAILED on the previous PR head 01d2550 (assert 0 > 0: triggers dropped on the quarantined file) passed
tests/hermes_state/ (133 files) — 1111 passed, 14 skipped. test_hermes_state.py hit the 300 s per-file cap under host load, then passed alone (275 passed). Its slowest test also takes ~15 s on base.
Full tests/e2e/core/sqlite/ (#120171 branch + this fix, no injection) — 26 passed

Live repro: before, a sibling holding the lock for 2.5 s cost the write every time (3/3). After, the write lands every time (3/3), and the torture-chamber episode goes from 4/4 red to 4/4 green under the same injection.

Found while triaging the E2E sqlite torture-chamber flake on #120171. The follow-on kill9_everything "roles still running" failure was a cascade from that episode leaving its gateway writer alive. The tests/core-e2e head already fixes that part: f601e22 stops a failed episode's stragglers in a finally.

Infographic

infographic

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on f0d730c — fix(state): stop a waiting FTS detach once the file is quara

debug info

CI timings

CI timings · View report · View job

Wall time 4m53s vs 9m13s (-47.0%). 5 job(s) slower, 6 faster, 1 unchanged.

  • Detect affected areas: -45.0s
  • Python tests / Run tests: +22.0s
  • Python tests / e2e: +18.0s
  • OS-specific tests / Windows-only tests: -13.0s
  • OS-specific tests / macOS-only tests: -10.0s

@arkheioncorp arkheioncorp left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

APPROVE

Fixes a real problem: a busy write lock on a corrupt FTS index was causing the writer to give up after its 1-second busy timeout, losing the canonical write. Now the detach retry loop waits out the lock holder on the write budget (_WRITE_PATIENCE_S), so the canonical row lands instead of escaping as "database disk image is malformed".

What is solid:

  • The new deadline/patience_s parameters propagate from _execute_write, so the wait budget is the same one the writer already budgeted for its retry loop. Consistent.
  • The sibling-holding-the-lock test is well-constructed: a thread takes BEGIN IMMEDIATE after the corruption check, holds it past the 1s busy timeout, and the writer waits it out. The message lands, _fts_stale is True, _db_corrupt is False. Good coverage of the exact scenario.
  • The check_then_contend monkeypatch is careful to only trigger the sibling on the first hit (guard if hit and not holder), avoiding duplicate threads.
  • The pytest.skip for SQLite builds that defer FTS shadow corruption past the insert trigger is the right call — this is a build-dependent behavior and the test should not fail on those builds.

One minor note:

  • The test uses threading.Timer(1.6, release.set) to release the sibling after 1.6s. With the writer's _WRITE_PATIENCE_S default and the sibling holding for "10" seconds (the release.wait(10)), this seems to rely on the Timer firing before the sibling's wait times out. The 1.6s is well within the writer's patience budget, but if _WRITE_PATIENCE_S were shorter than 1.6s this test could flake. Confirm the default patience is comfortably above 1.6s, or make the Timer interval a function of the patience setting.

No security issues. The lock-wait is a correctness fix, not an exposure.

@alt-glitch alt-glitch added type/bug Something isn't working P1 High — major feature broken, no workaround comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint area/sessions Session lifecycle, resume, persistence, history sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state labels Sep 23, 2026
@teknium1 teknium1 added the ci-reviewed applied to manually approve dangerous changes label Sep 23, 2026
…rrupt FTS index

When a canonical write trips a corrupt FTS index, SessionDB detaches the derived
indexes (breadcrumb + trigger drop) and retries the write. The detach ran one
BEGIN IMMEDIATE on the writer connection, whose busy timeout is only 1 s, and gave
up on "database is locked" — so the canonical write escaped as "database disk
image is malformed". The usual lock holder is a sibling writer (gateway + TUI)
detaching the same index, so under load the second writer's turn was lost.

The detach now waits out lock contention on the caller's write budget with the
same jittered retry as _execute_write (default _WRITE_PATIENCE_S for the search
fail-open callers).

Repro: a second process takes BEGIN IMMEDIATE the instant the corruption error
surfaces and holds it 2.5 s. Base: append raises after 1.02 s (3/3). Fixed: the
row lands after the holder releases, FTS detached (3/3). Found by the E2E sqlite
torture chamber (fts_corruption_fail_open) at load ~200.
The FTS fail-open detach now waits up to the caller's write budget (20 s /
60 s) for the write lock, so the one-time quarantine check before the loop
left a long window: a sibling that quarantined the file meanwhile still got
its triggers dropped and the stale breadcrumb committed on the quarantined
handle. Re-check the handle flag and the process-wide storage latch at the
top of every attempt, via the same _raise_if_db_corrupt(storage=True) that
_execute_write runs per attempt.

Classify the retryable lock error with is_sqlite_lock_error (result code
first) instead of a locked/busy substring match, matching #120488.
@teknium1
teknium1 force-pushed the fix/fts-fail-open-lock-patience branch from 01d2550 to f0d730c Compare September 23, 2026 23:47
@teknium1
teknium1 merged commit 82c5afb into main Sep 23, 2026
34 checks passed
@teknium1
teknium1 deleted the fix/fts-fail-open-lock-patience branch September 23, 2026 23:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/sessions Session lifecycle, resume, persistence, history ci-reviewed applied to manually approve dangerous changes comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P1 High — major feature broken, no workaround sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants