Skip to content

fix(gateway): replay routed profiles' pending-message spools at boot - #123585

Closed
Totoro-qaq wants to merge 1 commit into
NousResearch:mainfrom
Totoro-qaq:fix/profile-spool-recovery
Closed

Totoro-qaq wants to merge 1 commit into
NousResearch:mainfrom
Totoro-qaq:fix/profile-spool-recovery

Conversation

@Totoro-qaq

Copy link
Copy Markdown

What does this PR do?

On a multiplexed gateway, a routed turn runs inside its profile's HERMES_HOME, and _get_flush_dir() follows that home, so a transcript backlog spooled while the profile's store is unwritable lands in profiles/<name>/pending_messages/. Boot recovery only scanned the launch home, and the runtime drain keys on the in-memory _spooled_drop_sessions set that a restart empties. After a restart nothing read those files: the messages never reached state.db and the files stayed on disk.

Boot recovery now replays the launch home first, as before, then each served profile that has a pending_messages directory, inside that profile's HERMES_HOME so recover_pending_to_db() opens (and appends to) that profile's store. A failure in one profile is logged and does not stop the others. Single-profile gateways are unchanged.

Related Issue

Fixes #123584

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)

Changes Made

  • gateway/run.py: the boot _recover_pending closure now calls a module-level _recover_pending_flushes(runner), which adds the per-profile pass.
  • tests/gateway/test_profile_spool_recovery.py: spools a message under a routed profile's home, runs boot recovery, and checks the message is in the profile's state.db and the spool file is gone.

How to Test

  1. scripts/run_tests.sh tests/gateway/test_profile_spool_recovery.py: fails with the pre-fix behaviour (assert 0 == 1, the file stays in profiles/work/pending_messages/) and passes with this change.
  2. scripts/run_tests.sh tests/gateway/test_profile_spool_recovery.py tests/gateway/test_shutdown_flush.py tests/gateway/test_pending_queue_spool.py: 18 passed.
  3. scripts/run_tests.sh tests/gateway/: 8848 passed, 77 skipped. The known failures in test_media_download_retry.py (Slack ×3) and test_update_streaming.py::test_gateway_flag_enables_gateway_prompt_for_stash fail the same way on main. test_failure_writer_ownership.py failed once under the parallel full run and passes 3/3 when run alone, with or without this change; it does not touch recovery.
  4. ruff check (0.15.10) on the changed files: clean.

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits
  • I searched for existing PRs to make sure this isn't a duplicate (fix(gateway): route cap-dropped transcript recovery to the owning store #117310 is related but scans only the launch directory)
  • My PR contains only changes related to this fix
  • I've run the affected tests
  • I've added tests for my changes
  • I've tested on my platform: macOS 26 (arm64), Python 3.14

Documentation & Housekeeping

  • Documentation: docstring — N/A otherwise
  • cli-config.yaml.example — N/A
  • CONTRIBUTING.md / AGENTS.md — N/A
  • Cross-platform impact — N/A (path resolution only)
  • Tool descriptions/schemas — N/A

A routed turn on a multiplexed gateway runs inside its profile's
HERMES_HOME, and _get_flush_dir follows that home, so a transcript
backlog spooled while the profile's store is unwritable lands in
profiles/<name>/pending_messages/. Boot recovery only scanned the launch
home, and the runtime drain keys on an in-memory set that a restart
empties, so after a restart nothing read those files again: the
messages never reached state.db and the files stayed on disk.

Recover the launch home first as before, then each served profile that
has a pending_messages directory, inside that profile's HERMES_HOME so
the replay also lands in its own state.db. A failure in one profile is
logged and does not stop the others.
@kiramakes

Copy link
Copy Markdown

🤖 Hermes Agent automated review: PR reviewed. Diff analyzed (128 lines changed). Check CI status and manual review recommended.

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/gateway Gateway runner, session dispatch, delivery area/profiles Multi-profile isolation, HERMES_HOME scoping area/sessions Session lifecycle, resume, persistence, history sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Sep 26, 2026

kvnloo commented Sep 26, 2026

Copy link
Copy Markdown

Mirrored this exact tip on my fork and re-ran the boot-recovery witness. Base behavior: 1 failure with the routed profile spool left stranded; patched tip: 1 pass; restoring base reproduces the failure. The useful invariant for me is that recovery has to enter the routed profile's HERMES_HOME before calling recover_pending_to_db(), otherwise the replay can target the wrong state.db. No additional gap found.

@Totoro-qaq

Copy link
Copy Markdown
Author

Re the cluster note in #123658: this PR doesn't touch _recover_one_payload or shutdown_flush.py. It only changes the boot loop in gateway/run.py, so recovery also enters each routed profile's HERMES_HOME and drains its pending_messages/. git merge-tree against the current heads of #117310 and #122718 is clean for both, so no rebase is needed whichever lands first.

It composes with #117310: that PR decides which store a cap-dropped message replays into, and this one makes boot scan the routed profiles' spool dirs at all. Today recover_pending_to_db only globs the launch home's.

teknium1 added a commit that referenced this pull request Sep 28, 2026
…3584, salvage #123585)

Tighten the boot replay helper and prove the class in one test: two homes under
set_multiplex_active(True), a message spooled under each while its store was
unwritable, both replayed into their own state.db, no spool file left, and the
launch scope restored afterwards (A->B->A). Import gateway.run at module level so
the suite's real-home I/O guard does not trip on the lazy bootstrap import.
teknium1 added a commit that referenced this pull request Sep 28, 2026
…3584, salvage #123585)

Tighten the boot replay helper and prove the class in one test: two homes under
set_multiplex_active(True), a message spooled under each while its store was
unwritable, both replayed into their own state.db, no spool file left, and the
launch scope restored afterwards (A->B->A). Import gateway.run at module level so
the suite's real-home I/O guard does not trip on the lazy bootstrap import.
teknium1 added a commit that referenced this pull request Sep 28, 2026
…3584, salvage #123585)

Tighten the boot replay helper and prove the class in one test: two homes under
set_multiplex_active(True), a message spooled under each while its store was
unwritable, both replayed into their own state.db, no spool file left, and the
launch scope restored afterwards (A->B->A). Import gateway.run at module level so
the suite's real-home I/O guard does not trip on the lazy bootstrap import.
teknium1 added a commit that referenced this pull request Sep 28, 2026
…3584, salvage #123585)

Tighten the boot replay helper and prove the class in one test: two homes under
set_multiplex_active(True), a message spooled under each while its store was
unwritable, both replayed into their own state.db, no spool file left, and the
launch scope restored afterwards (A->B->A). Import gateway.run at module level so
the suite's real-home I/O guard does not trip on the lazy bootstrap import.
teknium1 added a commit that referenced this pull request Sep 28, 2026
…3584, salvage #123585)

Tighten the boot replay helper and prove the class in one test: two homes under
set_multiplex_active(True), a message spooled under each while its store was
unwritable, both replayed into their own state.db, no spool file left, and the
launch scope restored afterwards (A->B->A). Import gateway.run at module level so
the suite's real-home I/O guard does not trip on the lazy bootstrap import.
teknium1 added a commit that referenced this pull request Sep 28, 2026
…3584, salvage #123585)

Tighten the boot replay helper and prove the class in one test: two homes under
set_multiplex_active(True), a message spooled under each while its store was
unwritable, both replayed into their own state.db, no spool file left, and the
launch scope restored afterwards (A->B->A). Import gateway.run at module level so
the suite's real-home I/O guard does not trip on the lazy bootstrap import.
@Totoro-qaq

Copy link
Copy Markdown
Author

Landed via #126157 as 195bf48. Closing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/profiles Multi-profile isolation, HERMES_HOME scoping area/sessions Session lifecycle, resume, persistence, history comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Multiplexed gateway: a routed profile's spooled transcript backlog is never replayed after a restart

4 participants