Repository navigation
fix(gateway): replay routed profiles' pending-message spools at boot - #123585
Totoro-qaq wants to merge 1 commit into
Conversation
A routed turn on a multiplexed gateway runs inside its profile's HERMES_HOME, and _get_flush_dir follows that home, so a transcript backlog spooled while the profile's store is unwritable lands in profiles/<name>/pending_messages/. Boot recovery only scanned the launch home, and the runtime drain keys on an in-memory set that a restart empties, so after a restart nothing read those files again: the messages never reached state.db and the files stayed on disk. Recover the launch home first as before, then each served profile that has a pending_messages directory, inside that profile's HERMES_HOME so the replay also lands in its own state.db. A failure in one profile is logged and does not stop the others.
|
🤖 Hermes Agent automated review: PR reviewed. Diff analyzed (128 lines changed). Check CI status and manual review recommended. |
|
Mirrored this exact tip on my fork and re-ran the boot-recovery witness. Base behavior: 1 failure with the routed profile spool left stranded; patched tip: 1 pass; restoring base reproduces the failure. The useful invariant for me is that recovery has to enter the routed profile's HERMES_HOME before calling recover_pending_to_db(), otherwise the replay can target the wrong state.db. No additional gap found. |
|
Re the cluster note in #123658: this PR doesn't touch It composes with #117310: that PR decides which store a cap-dropped message replays into, and this one makes boot scan the routed profiles' spool dirs at all. Today |
…3584, salvage #123585) Tighten the boot replay helper and prove the class in one test: two homes under set_multiplex_active(True), a message spooled under each while its store was unwritable, both replayed into their own state.db, no spool file left, and the launch scope restored afterwards (A->B->A). Import gateway.run at module level so the suite's real-home I/O guard does not trip on the lazy bootstrap import.
…3584, salvage #123585) Tighten the boot replay helper and prove the class in one test: two homes under set_multiplex_active(True), a message spooled under each while its store was unwritable, both replayed into their own state.db, no spool file left, and the launch scope restored afterwards (A->B->A). Import gateway.run at module level so the suite's real-home I/O guard does not trip on the lazy bootstrap import.
…3584, salvage #123585) Tighten the boot replay helper and prove the class in one test: two homes under set_multiplex_active(True), a message spooled under each while its store was unwritable, both replayed into their own state.db, no spool file left, and the launch scope restored afterwards (A->B->A). Import gateway.run at module level so the suite's real-home I/O guard does not trip on the lazy bootstrap import.
…3584, salvage #123585) Tighten the boot replay helper and prove the class in one test: two homes under set_multiplex_active(True), a message spooled under each while its store was unwritable, both replayed into their own state.db, no spool file left, and the launch scope restored afterwards (A->B->A). Import gateway.run at module level so the suite's real-home I/O guard does not trip on the lazy bootstrap import.
…3584, salvage #123585) Tighten the boot replay helper and prove the class in one test: two homes under set_multiplex_active(True), a message spooled under each while its store was unwritable, both replayed into their own state.db, no spool file left, and the launch scope restored afterwards (A->B->A). Import gateway.run at module level so the suite's real-home I/O guard does not trip on the lazy bootstrap import.
…3584, salvage #123585) Tighten the boot replay helper and prove the class in one test: two homes under set_multiplex_active(True), a message spooled under each while its store was unwritable, both replayed into their own state.db, no spool file left, and the launch scope restored afterwards (A->B->A). Import gateway.run at module level so the suite's real-home I/O guard does not trip on the lazy bootstrap import.
What does this PR do?
On a multiplexed gateway, a routed turn runs inside its profile's HERMES_HOME, and
_get_flush_dir()follows that home, so a transcript backlog spooled while the profile's store is unwritable lands inprofiles/<name>/pending_messages/. Boot recovery only scanned the launch home, and the runtime drain keys on the in-memory_spooled_drop_sessionsset that a restart empties. After a restart nothing read those files: the messages never reachedstate.dband the files stayed on disk.Boot recovery now replays the launch home first, as before, then each served profile that has a
pending_messagesdirectory, inside that profile's HERMES_HOME sorecover_pending_to_db()opens (and appends to) that profile's store. A failure in one profile is logged and does not stop the others. Single-profile gateways are unchanged.Related Issue
Fixes #123584
Type of Change
Changes Made
gateway/run.py: the boot_recover_pendingclosure now calls a module-level_recover_pending_flushes(runner), which adds the per-profile pass.tests/gateway/test_profile_spool_recovery.py: spools a message under a routed profile's home, runs boot recovery, and checks the message is in the profile'sstate.dband the spool file is gone.How to Test
scripts/run_tests.sh tests/gateway/test_profile_spool_recovery.py: fails with the pre-fix behaviour (assert 0 == 1, the file stays inprofiles/work/pending_messages/) and passes with this change.scripts/run_tests.sh tests/gateway/test_profile_spool_recovery.py tests/gateway/test_shutdown_flush.py tests/gateway/test_pending_queue_spool.py: 18 passed.scripts/run_tests.sh tests/gateway/: 8848 passed, 77 skipped. The known failures intest_media_download_retry.py(Slack ×3) andtest_update_streaming.py::test_gateway_flag_enables_gateway_prompt_for_stashfail the same way onmain.test_failure_writer_ownership.pyfailed once under the parallel full run and passes 3/3 when run alone, with or without this change; it does not touch recovery.ruff check(0.15.10) on the changed files: clean.Checklist
Code
Documentation & Housekeeping
cli-config.yaml.example— N/ACONTRIBUTING.md/AGENTS.md— N/A