fix(sessions): storage maintenance refuses while a writer holds state.db; retired-WAL guard tells users what to do (#110054) - #117687
Merged
Conversation
….db; human-first retired-WAL guard text + recovery guide `hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed holder scan doctor and repair use before rewriting the store. While a gateway, Desktop, dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as `PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning, `--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and every agent answered every turn with the retired-WAL refusal until all writers were stopped by hand (#110054, maintainer follow-up 09-20). The DeletedWalGenerationError text is now two layers: a first sentence for the person reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete files while they run, docs link), then the operator detail. The classifier fingerprint "deleted state.db-wal or state.db-shm" is unchanged. The cause table (`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway home-channel notice) and the chat explainer carry the same first steps; the gateway notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause, which for a held retired generation is the second-writer trap. New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the guard text, the developer state-db-recovery page and the sessions guide): the three steps, the do-nots, why maintenance refuses, and what the files beside state.db are (retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups, snapshots).
૮ >ﻌ< ა ci reviewran on 6b83a81 — fix(gateway): keep the operator restart tail on the home-cha debug infoCI timingsCI timings · View report · View jobWall time 6m39s vs 7m37s (-12.7%). 10 job(s) slower, 2 faster, 1 unchanged.
|
…er, not the db object The admission gate read `db.db_path`, which made the refusal depend on whatever object `SessionDB` resolved to; the CLI tests substitute a lightweight double and CI went red with AttributeError: 'FakeDB' object has no attribute 'db_path'. The path now comes from `_default_db_path()` — the exact resolver `SessionDB()` itself uses two lines above — so the scan targets the same file in production and stays reachable regardless of the db object.
SummaryAdds a fail-closed admission gate so What changed
Strengths
Findings
VerdictLooks good to merge — the two findings are minor and safe to follow up. Reviewed using Hermes-Agent |
…ess holds state.db maybe_auto_prune_and_vacuum() runs the same store rewrite as `hermes sessions optimize` (VACUUM + TRUNCATE checkpoint) from CLI startup and the gateway constructor, with no holder scan — so the manual command was gated while the automatic producer of the same #110054 failure was not. The VACUUM branch now runs the same foreign_state_db_holders admission and SKIPS (debug log + a vacuum_skipped_holders count in the result) when a sibling writer holds the store or a WAL sidecar. Housekeeping never refuses a turn; it only defers the rewrite to the next run.
…rtises The Desktop/dashboard console printed `hermes sessions optimize --force` as the override, but _sessions_optimize rejected every argument — with a gateway running the command could only ever refuse. It now parses --force itself (and the hint names the console form).
…age notice
The cause-table action is user-phrased ("Send your message again once compression finishes"),
so the OPERATOR notice lost "then `hermes gateway restart`" for store-level failures that stay
broken until the gateway is restarted. The tail is appended for every cause except the
session-scoped ones that clear on their own (compression, compression_closed, turn_lease).
Also: the held-store refusal test is parametrized over optimize / optimize-storage / prune
(optimize-storage, the command the issue names as the field producer, was uncovered) and
asserts the refusal names the same store SessionDB opened — no `hermes sessions` subcommand
can point the command at another database. Docs: doctor refuses the checkpoint only while it
can see a process holding the RETIRED log.
This was referenced Sep 21, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
hermes sessions optimize/optimize-storage/prunenow refuse while another Hermes process holdsstate.db, naming each holder asPID N (command), and the retired-WAL guard tells a person what happened and the one thing to do, with a recovery guide to link to.hermes_cli/sessions_cmd.py::cmd_sessionsvia_HELD_STORE_ACTIONS, text inhermes_state_holders.py::held_store_refusal): reuses the same fail-closedforeign_state_db_holdersscan doctor/repair use (an incomplete scan refuses too), exits 1,--forceoverrides,--dry-runpreviews are never gated. The Desktop console'ssessions optimize(hermes_cli/console_engine.py::_sessions_optimize) gets the same refusal.hermes_state_errors._DELETED_WAL_GENERATION_MSG): first sentence for humans — nothing is lost, quit every Hermes process on the profile,hermes doctornames the holders, neverdoctor --fix/ never delete files while they run, docs link — then the operator detail. Classifier fingerprintdeleted state.db-wal or state.db-shmunchanged;storage_replacedcode unchanged.hermes_state_user_copy(CLI banner, TUI/Desktop RPC error via_db_unavailable_error, gateway home-channel notice) andagent/turn_explainers.py(chat bubble). The gateway notice no longer hardcodesdoctor --fix+gateway restartfor every non-corrupt cause — it now uses the cause table's action.website/docs/user-guide/session-storage-recovery.md(three steps, do-nots, why maintenance refuses, the files besidestate.db:retired-wal-*/manifest.json,pre-update-emergency-*.bak, corrupt backups, snapshots), registered insidebars.ts, linked from the guard text,developer-guide/state-db-recovery.mdanduser-guide/sessions.md.Live repro (temp HERMES_HOME + fake HOME; a second real process opens
SessionDBon the store and stays live):31d237cab9d3hermes sessions optimizewith holder liveOptimized 2 FTS index(es), VACUUM + TRUNCATE checkpoint, rc=0Refusing … PID 2220316 (python …/probe): state.db, state.db-shm, state.db-wal…Override with --force… guide URL, rc=1optimize-storage --yes,prune --yeswith holder liveoptimize --forcewith holder liveOptimized 2 FTS index(es)), rc=0prune --dry-runwith holder liveoptimizehermes doctorwith holder live1 process(es) holding the DB open, no bare--fixnudge (#116297 landed)Honest note: on this host (SQLite 3.53 + the #110544 OFD-lock fix) the holder's next write after the base-side VACUUM still succeeded — the unlink is prevented at the WAL level — so the field-reported "every agent refuses" cascade did not reproduce here; the gate is about never running the rewrite underneath a live writer in the first place, which is the maintainer's 09-20 ask.
Root cause: the storage rewrite commands had no admission check at all, so
optimize-storageunder a live fleet was a producer of the very state whose only recovery text was an operator runbook.Review round 2 (09-20): the bug class is now closed on both producers.
hermes_state_maintenance.py::maybe_auto_prune_and_vacuumran the SAME rewrite (VACUUM + TRUNCATE checkpoint) automatically from CLI startup and the gateway constructor with no holder scan. It now runs the sameforeign_state_db_holdersadmission before the VACUUM branch and SKIPS (debug log +vacuum_skipped_holders) — automatic maintenance never refuses a turn, it defers the rewrite.hermes_cli/console_engine.py) parses--forceinstead of rejecting every argument, so the override the refusal advertises is reachable from the surface that printed it.then `hermes gateway restart`for every store-level cause again (session-scoped causes — compression, compression_closed, turn_lease — keep the user-phrased action).cmd_sessionsresolves it via_default_db_path(); nohermes sessionssubcommand accepts a--db/alternate-store option, so that IS the operated-on store — asserted in the test.Tests:
tests/hermes_cli/test_sessions_held_store_gate.py(parametrized over optimize / optimize-storage / prune throughcmd_sessionswith a real subprocess holder: refuse+name PID until--force; dry-run preview passes, delete waits for a quiet store — red with only the call-site wiring reverted). Plustests/hermes_state/test_auto_vacuum_holder_gate.py(auto-VACUUM skips under a real holder, VACUUMs once it exits) and a console--forcereachability test; all three are red with their fix reverted. Runs:tests/hermes_cli+tests/hermes_state+tests/gateway= 2343 files, 23606 passed; the 9 failures are pre-existing host/load flakes (dashboard auth-gate port, update venv repair, sidebar-cache concurrency, hygiene timing) that pass in isolation.Part of #110054 — closes the maintainer's 09-20 producer atom, the two-layer guard text and the user-facing recovery page; the Desktop one-click "stop holders" surface and whether Hermes may kill foreign holders (#110073) remain maintainer decisions.
Dropped hunks
None cherry-picked. #110073 (@JoaoMarcos44, +1033/-35: Desktop maintenance surface +
POST /api/ops/...kill path) and #110179 (@ngpestelos, +952/-40:_import_db_memberpublish-under-holders + diverted-transcript replay) are sized REVIEW_ONLY; #110179's ~10-LOC publish atom is described in the lane report as a credited follow-up carve.Infographic