Skip to content

fix(sessions): storage maintenance refuses while a writer holds state.db; retired-WAL guard tells users what to do (#110054) - #117687

Merged
teknium1 merged 5 commits into
mainfrom
fix/b466-statedb-recovery
Sep 21, 2026
Merged

teknium1 merged 5 commits into
mainfrom
fix/b466-statedb-recovery

Conversation

@teknium1

@teknium1 teknium1 commented Sep 20, 2026 •

Copy link
Copy Markdown
Collaborator

hermes sessions optimize / optimize-storage / prune now refuse while another Hermes process holds state.db, naming each holder as PID N (command), and the retired-WAL guard tells a person what happened and the one thing to do, with a recovery guide to link to.

  • Held-store gate (hermes_cli/sessions_cmd.py::cmd_sessions via _HELD_STORE_ACTIONS, text in hermes_state_holders.py::held_store_refusal): reuses the same fail-closed foreign_state_db_holders scan doctor/repair use (an incomplete scan refuses too), exits 1, --force overrides, --dry-run previews are never gated. The Desktop console's sessions optimize (hermes_cli/console_engine.py::_sessions_optimize) gets the same refusal.
  • Two-layer guard text (hermes_state_errors._DELETED_WAL_GENERATION_MSG): first sentence for humans — nothing is lost, quit every Hermes process on the profile, hermes doctor names the holders, never doctor --fix / never delete files while they run, docs link — then the operator detail. Classifier fingerprint deleted state.db-wal or state.db-shm unchanged; storage_replaced code unchanged.
  • Same first steps on every surface: hermes_state_user_copy (CLI banner, TUI/Desktop RPC error via _db_unavailable_error, gateway home-channel notice) and agent/turn_explainers.py (chat bubble). The gateway notice no longer hardcodes doctor --fix + gateway restart for every non-corrupt cause — it now uses the cause table's action.
  • Docs: new website/docs/user-guide/session-storage-recovery.md (three steps, do-nots, why maintenance refuses, the files beside state.db: retired-wal-*/manifest.json, pre-update-emergency-*.bak, corrupt backups, snapshots), registered in sidebars.ts, linked from the guard text, developer-guide/state-db-recovery.md and user-guide/sessions.md.

Live repro (temp HERMES_HOME + fake HOME; a second real process opens SessionDB on the store and stays live):

base 31d237cab9d3 this PR
hermes sessions optimize with holder live proceeds: Optimized 2 FTS index(es), VACUUM + TRUNCATE checkpoint, rc=0 Refusing … PID 2220316 (python …/probe): state.db, state.db-shm, state.db-wal … Override with --force … guide URL, rc=1
optimize-storage --yes, prune --yes with holder live proceed refuse, rc=1
optimize --force with holder live — proceeds (Optimized 2 FTS index(es)), rc=0
prune --dry-run with holder live preview preview (not gated)
holder exited, optimize proceeds proceeds, rc=0
hermes doctor with holder live names 1 process(es) holding the DB open, no bare --fix nudge (#116297 landed) same

Honest note: on this host (SQLite 3.53 + the #110544 OFD-lock fix) the holder's next write after the base-side VACUUM still succeeded — the unlink is prevented at the WAL level — so the field-reported "every agent refuses" cascade did not reproduce here; the gate is about never running the rewrite underneath a live writer in the first place, which is the maintainer's 09-20 ask.

Root cause: the storage rewrite commands had no admission check at all, so optimize-storage under a live fleet was a producer of the very state whose only recovery text was an operator runbook.

Review round 2 (09-20): the bug class is now closed on both producers.

  • hermes_state_maintenance.py::maybe_auto_prune_and_vacuum ran the SAME rewrite (VACUUM + TRUNCATE checkpoint) automatically from CLI startup and the gateway constructor with no holder scan. It now runs the same foreign_state_db_holders admission before the VACUUM branch and SKIPS (debug log + vacuum_skipped_holders) — automatic maintenance never refuses a turn, it defers the rewrite.
  • Console (hermes_cli/console_engine.py) parses --force instead of rejecting every argument, so the override the refusal advertises is reachable from the surface that printed it.
  • The gateway home-channel OPERATOR notice appends then `hermes gateway restart` for every store-level cause again (session-scoped causes — compression, compression_closed, turn_lease — keep the user-phrased action).
  • Gate path: cmd_sessions resolves it via _default_db_path(); no hermes sessions subcommand accepts a --db/alternate-store option, so that IS the operated-on store — asserted in the test.
  • Docs: doctor refuses the checkpoint only while it can see a process holding the retired log.

Tests: tests/hermes_cli/test_sessions_held_store_gate.py (parametrized over optimize / optimize-storage / prune through cmd_sessions with a real subprocess holder: refuse+name PID until --force; dry-run preview passes, delete waits for a quiet store — red with only the call-site wiring reverted). Plus tests/hermes_state/test_auto_vacuum_holder_gate.py (auto-VACUUM skips under a real holder, VACUUMs once it exits) and a console --force reachability test; all three are red with their fix reverted. Runs: tests/hermes_cli + tests/hermes_state + tests/gateway = 2343 files, 23606 passed; the 9 failures are pre-existing host/load flakes (dashboard auth-gate port, update venv repair, sidebar-cache concurrency, hygiene timing) that pass in isolation.

Part of #110054 — closes the maintainer's 09-20 producer atom, the two-layer guard text and the user-facing recovery page; the Desktop one-click "stop holders" surface and whether Hermes may kill foreign holders (#110073) remain maintainer decisions.

Dropped hunks

None cherry-picked. #110073 (@JoaoMarcos44, +1033/-35: Desktop maintenance surface + POST /api/ops/... kill path) and #110179 (@ngpestelos, +952/-40: _import_db_member publish-under-holders + diverted-transcript replay) are sized REVIEW_ONLY; #110179's ~10-LOC publish atom is described in the lane report as a credited follow-up carve.

Infographic

Maintenance refuses while a writer is live

….db; human-first retired-WAL guard text + recovery guide

`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).

The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.

New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
@alt-glitch alt-glitch added type/bug Something isn't working P1 High — major feature broken, no workaround comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/cli CLI entry point, hermes_cli/, setup wizard comp/gateway Gateway runner, session dispatch, delivery area/sessions Session lifecycle, resume, persistence, history sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state labels Sep 20, 2026
@github-actions

github-actions Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 6b83a81 — fix(gateway): keep the operator restart tail on the home-cha

debug info

CI timings

CI timings · View report · View job

Wall time 6m39s vs 7m37s (-12.7%). 10 job(s) slower, 2 faster, 1 unchanged.

  • OS-specific tests / macOS-only tests: +31.0s
  • Python lints / Windows footguns (blocking): +20.0s
  • Docs Site / docs-site-checks: +16.0s
  • Check contributors / check-attribution: +14.0s
  • Python tests / Run tests: -6.0s

…er, not the db object

The admission gate read `db.db_path`, which made the refusal depend on whatever object
`SessionDB` resolved to; the CLI tests substitute a lightweight double and CI went red with
AttributeError: 'FakeDB' object has no attribute 'db_path'. The path now comes from
`_default_db_path()` — the exact resolver `SessionDB()` itself uses two lines above — so the
scan targets the same file in production and stays reachable regardless of the db object.
@kyssta-exe

Copy link
Copy Markdown

Summary

Adds a fail-closed admission gate so sessions optimize / optimize-storage / prune refuse while another Hermes process holds state.db, and rewrites the retired-WAL guard text so users get a human remedy instead of an operator runbook. The root-cause framing is right: maintenance rewrites were a producer of the very retired-WAL state the old copy couldn't recover from.

What changed

  • hermes_state_holders.py: new held_store_refusal() — fail-closed foreign_state_db_holders scan, names each holder PID N (command), includes --force override hint and recovery-guide URL.
  • hermes_cli/sessions_cmd.py + console_engine.py: gate wired into cmd_sessions (optimize/optimize-storage/prune, dry-run exempt) and Desktop console sessions optimize.
  • hermes_state_errors.py, hermes_state_user_copy.py, turn_explainers.py, run_notifications.py: two-layer guard copy; gateway notice now uses the cause table's action instead of hardcoding doctor --fix + restart.
  • New website/docs/user-guide/session-storage-recovery.md (+ sidebar/docs links) and tests/hermes_cli/test_sessions_held_store_gate.py with a real subprocess holder.

Strengths

  • Reuses the doctor/repair holder scan with the same fail-closed semantics (incomplete scan refuses, never reads as all-clear) — no second implementation to drift.
  • --dry-run previews exempt, --force override preserved, and the live-repro table in the PR body covers both paths with exit codes.
  • Tests drive the production entry point with a real second process, not a mocked scan.

Findings

  • console_engine.py: _sessions_optimize rejects all args (_expect_no_args) yet the refusal text cites a --force hint — Desktop console users have no override path. Either drop the hint there or accept it as intentional; worth one line in the refusal call.
  • sessions_cmd.py:1026: the gate scans _default_db_path() rather than the path of the already-opened db. If a profile override ever makes those diverge, the gate checks the wrong file — consider passing the opened db's path through.
  • Non-blocking: the --force help text is long and repeated verbatim 3x in subcommands/sessions.py; a shared constant would prevent drift.

Verdict

Looks good to merge — the two findings are minor and safe to follow up.

Reviewed using Hermes-Agent

…ess holds state.db

maybe_auto_prune_and_vacuum() runs the same store rewrite as `hermes sessions optimize`
(VACUUM + TRUNCATE checkpoint) from CLI startup and the gateway constructor, with no holder
scan — so the manual command was gated while the automatic producer of the same #110054
failure was not. The VACUUM branch now runs the same foreign_state_db_holders admission and
SKIPS (debug log + a vacuum_skipped_holders count in the result) when a sibling writer holds
the store or a WAL sidecar. Housekeeping never refuses a turn; it only defers the rewrite to
the next run.
…rtises

The Desktop/dashboard console printed `hermes sessions optimize --force` as the override, but
_sessions_optimize rejected every argument — with a gateway running the command could only
ever refuse. It now parses --force itself (and the hint names the console form).
…age notice

The cause-table action is user-phrased ("Send your message again once compression finishes"),
so the OPERATOR notice lost "then `hermes gateway restart`" for store-level failures that stay
broken until the gateway is restarted. The tail is appended for every cause except the
session-scoped ones that clear on their own (compression, compression_closed, turn_lease).

Also: the held-store refusal test is parametrized over optimize / optimize-storage / prune
(optimize-storage, the command the issue names as the field producer, was uncovered) and
asserts the refusal names the same store SessionDB opened — no `hermes sessions` subcommand
can point the command at another database. Docs: doctor refuses the checkpoint only while it
can see a process holding the RETIRED log.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/sessions Session lifecycle, resume, persistence, history comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/cli CLI entry point, hermes_cli/, setup wizard comp/gateway Gateway runner, session dispatch, delivery P1 High — major feature broken, no workaround sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants