Skip to content

fix: serialize concurrent FTS bootstrap repair - #486

Closed
100yenadmin wants to merge 1 commit into
stephenschoettler:mainfrom
100yenadmin:research/multiprocess-sqlite-safety
Closed

100yenadmin wants to merge 1 commit into
stephenschoettler:mainfrom
100yenadmin:research/multiprocess-sqlite-safety

Conversation

@100yenadmin

@100yenadmin 100yenadmin commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • serialize constructor-time external-content FTS repair with one short BEGIN IMMEDIATE ownership transaction;
  • re-inspect FTS table, shadow-table, and trigger state after ownership is acquired so losing processes accept the winner's valid schema instead of recreating it;
  • preserve SQLite lock provenance rather than converting contention into corruption or destructive repair;
  • retain caller-owned transaction compatibility through a savepoint path;
  • ensure a trigger disappearance observed on the healthy fast path re-enters ownership before any trigger installation;
  • add deterministic trigger-race and six-process spawn regressions with message conservation, token-level search, FTS5 integrity, shadow-table parity, clean reopen, and bounded lock-contention coverage.

Why

Two independent processes could both observe an incomplete fresh-database FTS schema and then race through unconditional virtual-table creation. The loser failed startup with:

sqlite3.OperationalError: table "messages_fts" already exists

The existing WAL conversion retry happens earlier and does not own the later FTS repair boundary.

repair_external_content_fts() now separates cheap preflight detection from the destructive repair boundary:

  1. detect structural or deep repair need;
  2. acquire BEGIN IMMEDIATE when the connection does not already own a transaction;
  3. re-check complete FTS/shadow/trigger state under ownership;
  4. let only the owner drop, create, rebuild, and install triggers;
  5. commit atomically;
  6. let concurrent losers acquire the transaction after the winner, re-inspect, and accept the complete winner state.

If SQLite reports lock/busy contention on the repair-needed path, the exception remains an availability failure. It is not classified as FTS corruption and does not authorize destructive repair. Classification uses SQLite BUSY/LOCKED result codes plus genuine lock-message fallbacks; unrelated error text containing timeout or busy is not treated as lock provenance.

Validation

  • Focused validation: python -m pytest -q -p no:cacheprovider tests/test_db_bootstrap_fts.py -> 35 passed
  • Default validation:
    • pytest tests/test_lcm_core.py tests/test_lcm_engine.py tests/test_packaging_install.py -q — exact grouped command not run; tests/test_lcm_core.py passed 292, and the release validator's broader focused gate passed.
    • pytest -q — run; 2311 passed, 1 skipped, 12 xfailed, 4 failed on the known macOS /var versus /private/var externalization-path baseline.
    • bash -lc 'ulimit -n 1024 && pytest -q' — exact template command not run; the release validator's low-FD gate produced the same 2311 passed, 1 skipped, 12 xfailed, 4 failed baseline-only result.
    • python -m compileall -q .
    • python -m py_compile scripts/import_lossless_claw.py
    • bash -n scripts/install.sh scripts/update.sh
    • git diff --check
  • Release validation: scripts/validate_release.sh --full --keep-going --output /tmp/hermes-lcm-release-validation-475-lock-classifier-final-20260803 -> all diff, compilation, shell, focused, benchmark, and stress gates passed; aggregate remained red only for the same ordinary/low-FD baseline failures above.
  • Static validation: python -m ruff check db_bootstrap.py tests/test_db_bootstrap_fts.py -> passed.
  • GitHub CI on f0fd360804afbe6aeddb2bc12f6c86d0a2b2f165: workflow lint, lint, and Python 3.11–3.14 passed.
  • Workflow validation, if workflows changed: actionlint — not applicable; this PR changes no workflow files.

Spawned-process acceptance

The candidate regression uses six independent spawn workers. Each child first observes the structurally incomplete fresh FTS state, then a second barrier releases all six from that same stale observation. This deterministically drives the pre-fix implementation into competing virtual-table creation while allowing the candidate's post-BEGIN IMMEDIATE recheck to select one owner.

unmodified pinned base + test-only diff:
1 failed with sqlite3.OperationalError: table "messages_fts" already exists

candidate, same forced schedule:
5/5 consecutive root rounds passed
10/10 consecutive final adversarial-review rounds passed

The regression checks:

  • 24/24 unique messages persisted;
  • four token-level FTS matches per worker marker;
  • messages_fts_docsize parity;
  • FTS5 integrity command;
  • PRAGMA integrity_check and foreign_key_check;
  • clean MessageStore reopen.

The focused suite separately checks bounded contention returning OperationalError('database is locked') without partial FTS artifacts.

Trigger-disappearance ownership acceptance

A second real SQLite connection drops one canonical product trigger after the initial healthy/deep precheck. SQLite tracing on the repair connection records whether each observed CREATE TRIGGER statement executes inside a transaction.

unmodified pinned base + test-only diff:
1 failed; CREATE TRIGGER transaction states were [False, False, False]

candidate, same forced schedule:
10/10 consecutive rounds passed; every CREATE TRIGGER state was True

The healthy fast path now returns without executing trigger DDL only when the second trigger check remains complete. An observed disappearance falls through to BEGIN IMMEDIATE, complete-state reinspection, and owner-only trigger recreation.

A separate preserved 12-process/fork reconnaissance probe also passed constructor startup, but it is not cited as the six-process regression and is not included in this PR.

Notes

  • The no-repair fast path retains the existing metadata cleanup/commit behavior and the connection's configured busy timeout. The temporary SQLITE_BUSY_TIMEOUT_MS ownership budget applies only after repair-worthy state is observed.
  • Connections already inside a caller-owned transaction retain historical commit behavior and use a savepoint. A savepoint does not add cross-process ownership to an already-active deferred transaction; the serialized guarantee is for the fresh constructor path.
  • Winner acceptance rechecks table/shadow presence, FTS declaration, indexed column, content-to-docsize parity, and canonical trigger names; it does not compare existing trigger bodies.
  • Stale same-name trigger definitions remain PR fix: repair stale FTS triggers left by schema migrations #445's repair contract. If both changes land, stale-trigger detection, recheck, drop, and canonical recreation must occur inside this ownership transaction. The branches must not be combined mechanically.
  • This PR does not claim to solve ingest idempotency, compaction/publication ownership, lifecycle generation/CAS, cleanup liveness, or durable maintenance debt.
  • The four ordinary/low-FD failures are the pre-existing macOS path-alias baseline and do not touch the modified FTS paths.

Closes #475.

Refs #475

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f0fd360804

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread db_bootstrap.py
owner_rebuild_needed = (
_fts_needs_rebuild_structural(conn, spec) if winner_state_needs_repair else False
)
if not owner_rebuild_needed and not structural_repair_needed and deep_repair_needed:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Recheck deep corruption when triggers are missing

When an FTS trigger is missing and updates have also caused same-row-count index drift, structural_repair_needed is true solely because of the trigger, so this condition prevents the owner from running _fts_needs_rebuild. The repair recreates the trigger and reports success with rebuilt=False, but the existing index remains corrupt and searches continue returning stale or missing results; explicit doctor repair should still perform the deep check after fixing trigger-only structural state.

Useful? React with 👍 / 👎.

@100yenadmin

Copy link
Copy Markdown
Contributor Author

Closing in favour of #505, the maintenance roll-up — this is commit 1 of that stack.

Nothing is dropped — the commits are carried across unchanged, so the review history here stays meaningful and the work is not rewritten. The consolidation is packaging: a system-level change reads better as one ordered stack than as several PRs that have to be merged in the right sequence to make sense.

Apologies for the churn on your queue.

@100yenadmin 100yenadmin closed this Aug 5, 2026
@100yenadmin
100yenadmin deleted the research/multiprocess-sqlite-safety branch August 5, 2026 08:39
stephenschoettler added a commit that referenced this pull request Sep 3, 2026
* fix: serialize concurrent FTS bootstrap repair

Adopt PR #486 commit f0fd360 onto current main.

Include only the FTS deep-repair correction from PR #505 commit 5af57d4 so missing triggers do not mask same-row-count corruption. Add transaction-boundary regressions required by the local acceptance packet.

Refs #475

* fix: preserve due FTS checks during trigger repair
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: concurrent fresh-database startup can race FTS repair and fail initialization

1 participant