Skip to content

fix: retry WAL conversion on lock contention at connection setup - #361

Merged
stephenschoettler merged 1 commit into
stephenschoettler:mainfrom
ai-ag2026:pr/harden-concurrent-migration-test
Jul 27, 2026
Merged

stephenschoettler merged 1 commit into
stephenschoettler:mainfrom
ai-ag2026:pr/harden-concurrent-migration-test

Conversation

@ai-ag2026

Copy link
Copy Markdown
Contributor

Summary

Follow-up to #346 (concurrent-startup migration race), one phase earlier in the same failure class. configure_connection runs PRAGMA journal_mode=WAL as its first statement. Converting a rollback-journal database to WAL needs the exclusive lock, and SQLite can return SQLITE_BUSY for that upgrade without consulting the busy handler while sibling connections are mid-setup on the same file. Concurrent process startup (gateway + CLI + sub-agents — the documented topology) on a not-yet-WAL database therefore crashed sporadically with sqlite3.OperationalError: database is locked during store construction. Steady state is immune: once the database is WAL, the pragma is a plain read.

Two changes:

  • db_bootstrap.configure_connection: set busy_timeout first, then run the WAL conversion through _execute_wal_conversion_with_lock_retry — a bounded exponential-backoff retry (5ms→250ms steps, total budget SQLITE_BUSY_TIMEOUT_MS) that only swallows lock-contention errors and re-raises everything else or on budget exhaustion.
  • tests/test_crash_safe_wal.py::TestConcurrentStartupMigration: the regression test that exposed this had a second bug of its own — barrier.wait() without timeout. The thread that crashed pre-barrier left the seven survivors parked forever, so instead of a red test the whole pytest run deadlocked (observed twice as a validate_release --full hang at 4%, 8 threads in futex wait, zero CPU). The barrier now uses a timeout and abort() on failure, join is bounded with a liveness assert, and the final assert reports the root-cause exception instead of hiding it.

Why

The hang variant is worse than the crash variant: it silently stalls any CI/validation run that executes the suite, with no error to act on. And the underlying crash is the same first-boot scenario #346 fixed for column DDL — a fresh install or an upgrade from a pre-WAL database with several Hermes processes starting together.

Validation

ruff check db_bootstrap.py tests/test_crash_safe_wal.py     # All checks passed!
python -m pytest tests/ -q                                   # 1590 passed, 12 xfailed

Reproduction & proof (Linux, 8-thread barrier race on a rollback-journal seed DB):

  • Unfixed: standalone repro script hit database is locked from configure_connection (PRAGMA journal_mode=WAL) within 15 attempts; the unhardened test hung 2/10 runs (killed only by an external timeout).
  • Fixed: 60/60 repro attempts clean; hardened test file looped 30× with zero failures or hangs.

Notes

  • Retry scope is deliberately the WAL conversion only — the other pragmas are per-connection settings that do not take the exclusive-lock upgrade path.
  • The hardened test now degrades into an actionable assertion (with the recorded root-cause exception) if any future setup-phase error appears, instead of deadlocking the suite.

Refs

@ai-ag2026
ai-ag2026 force-pushed the pr/harden-concurrent-migration-test branch from 7e90eec to 88a235f Compare July 10, 2026 05:44
@ai-ag2026

Copy link
Copy Markdown
Contributor Author

@codex review current head 88a235f. Rebasing refresh only: stable patch-id preserved, direct parent/merge-base current main d2618c7. Validation: focused crash-safe WAL suite 14 passed independently; full 1605 passed, 12 xfailed; ruff/static/shell/diff checks passed.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. What shall we delve into next?

Reviewed commit: 88a235f653

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@ai-ag2026
ai-ag2026 force-pushed the pr/harden-concurrent-migration-test branch from 88a235f to c2d7709 Compare July 12, 2026 01:00
@ai-ag2026

Copy link
Copy Markdown
Contributor Author

@codex review current head c2d7709. Rebased independently onto current main dd9a5a7 with no conflicts; fork and PR refs read back at this exact SHA. Local verification: focused pytest 14 passed; full suite 1605 passed, 12 xfailed; ruff, compileall/py_compile, shell syntax, and diff checks passed.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Swish!

Reviewed commit: c2d770901b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Converting a rollback-journal database to WAL needs the exclusive lock,
and SQLite can return SQLITE_BUSY for that upgrade without consulting
the busy handler while sibling connections are mid-setup on the same
file. Concurrent process startup (gateway + CLI + sub-agents) on a
not-yet-WAL database therefore crashed sporadically with 'database is
locked' from configure_connection.

Wrap the journal_mode pragma in a bounded exponential-backoff retry
(budget = SQLITE_BUSY_TIMEOUT_MS) and set busy_timeout first. Once the
database is in WAL mode the pragma is a plain read, so steady state is
unaffected.

Also harden the concurrent-migration regression test that exposed this:
its barrier had no timeout, so the pre-barrier failure parked the seven
surviving threads forever and deadlocked the whole test run instead of
reporting the error. The barrier now times out and aborts on failure,
join is bounded, and the assert surfaces the root-cause exception.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@ai-ag2026
ai-ag2026 force-pushed the pr/harden-concurrent-migration-test branch from c2d7709 to 09a2307 Compare July 26, 2026 13:15
@ai-ag2026

Copy link
Copy Markdown
Contributor Author

Merge-ready on 09a2307: rebased onto current main (424a6b2), CI 6/6 green on that head, no open review threads.

Self-contained — one retry path in db_bootstrap.py plus its test; no dependency on any other PR from this fork. Kept open as part of the reduced queue (#416).

@stephenschoettler
stephenschoettler requested a review from Tosko4 July 27, 2026 00:30

@Tosko4 Tosko4 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed 09a2307 against current main.

I reproduced the startup failure on the exact base with a rollback-journal database held under an exclusive lock: the base failed immediately with database is locked, while this head waited for release and completed in WAL mode.

Validation on the exact head:

  • tests/test_crash_safe_wal.py: 14 passed
  • concurrent migration regression: 30 consecutive passes
  • full suite: 2297 passed, 12 xfailed
  • full release validation: passed

The retry remains scoped to lock-related OperationalErrors, non-lock failures still propagate, and the hardened barrier can no longer hide a setup failure behind a deadlocked test run. I found no blocking issue in this change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants