Make schema-column migrations idempotent under concurrent startup - #346
Merged
stephenschoettler merged 1 commit intoJul 8, 2026
Merged
stephenschoettler merged 1 commit into
stephenschoettler merged 1 commit into
Conversation
In the multi-agent deployment (gateway + CLI sessions + sub-agents) every process opens its own connection to the same lcm.db and runs the startup migrations concurrently. Each column migration did a check-`PRAGMA table_info`-then-`ALTER TABLE ADD COLUMN` in autocommit with no guard, so two processes could both observe a column as absent and both issue the ALTER; the loser raised `sqlite3.OperationalError: duplicate column name`, which propagated out of `_init_db` and crashed store construction. This bites exactly at an upgrade boundary, when many processes restart together. Add `add_column_if_missing(conn, existing_columns, column, alter_sql)` in db_bootstrap.py: it keeps the existing "skip if already present" check and additionally swallows exactly the `duplicate column name` OperationalError (any other error still propagates), making the ALTER idempotent under concurrency. Route all eleven column migrations through it (db_bootstrap `ensure_lifecycle_state_columns` / `ensure_message_origin_columns`, `store._ensure_source_column` / `_ensure_conversation_id_column`, `dag._ensure_source_window_columns`). Behaviour is unchanged on the non-racing path (same check, same ALTERs); only the concurrent loser now skips instead of crashing. Rollback: revert this commit; the migrations return to the prior unguarded check-then-ALTER. Validation: - New regression test `TestConcurrentStartupMigration` races `ensure_message_origin_columns` from 8 threads (each its own connection to one seeded pre-v5 DB, FTS-free to isolate the column DDL). Without the guard 7/8 threads raise `duplicate column name: conversation_id`; with it every thread migrates and the column exists exactly once. - ruff clean; full suite passes. Note (out of scope, follow-up): the FTS structural rebuild (`repair_external_content_fts`) has a separate concurrency window on a different trigger (missing shadow tables / corruption, not column upgrades); a column migration does not change row counts so it does not enter that path. Left for a focused follow-up.
stephenschoettler
approved these changes
Jul 8, 2026
stephenschoettler
left a comment
Owner
There was a problem hiding this comment.
Approved via Hermes Kanban reviewer gate t_d925191a. Evidence: live CI green 6/6, no unresolved review threads, and local validation passed on the reviewed head.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
add_column_if_missing(conn, existing_columns, column, alter_sql)indb_bootstrap.pyand route all column migrations through it. It keeps the existing "skip if already present" check and additionally swallows exactly theduplicate column nameOperationalError(any other error still propagates), making eachALTER TABLE ... ADD COLUMNidempotent under concurrency.ensure_lifecycle_state_columns/ensure_message_origin_columns(db_bootstrap.py),MessageStore._ensure_source_column/_ensure_conversation_id_column(store.py),SummaryDAG._ensure_source_window_columns(dag.py).Why
In the multi-agent deployment (gateway + CLI sessions + sub-agents) every process opens its own connection to the same
lcm.dband runs the startup migrations concurrently. Each column migration did check-PRAGMA table_info-then-ALTERin autocommit with no guard, so two processes could both observe a column as absent and both issue the ALTER. The loser raisedsqlite3.OperationalError: duplicate column name, which propagated out of_init_dband crashed store construction. This bites precisely at an upgrade boundary, when many processes restart together.Behaviour is unchanged on the non-racing path (same check, same ALTERs); only the concurrent loser now skips instead of crashing.
Validation
TestConcurrentStartupMigration(tests/test_crash_safe_wal.py) racesensure_message_origin_columnsfrom 8 threads, each with its own connection to one seeded pre-v5 DB (FTS-free, to isolate the column DDL). Without the guard 7/8 threads raiseduplicate column name: conversation_id; with it every thread migrates and the column exists exactly once.ruff check .-> clean.PYTHONPATH=<hermes-agent> python -m pytest -q -o addopts=->1576 passed, 12 xfailed.scripts/validate_release.sh --full->PASS: release validation full(all 12 gates, incl. low-fd).Notes
repair_external_content_fts) has a separate concurrency window on a different trigger (missing shadow tables / corruption, not column upgrades — a column migration does not change row counts, so it does not enter that path). Left for a focused follow-up so this PR stays scoped to the verified crash.Refs
Surfaced by a comparison against the more mature lossless-claw, whose
runLcmMigrationsserializes migrations underBEGIN EXCLUSIVE.