Skip to content

fix(gateway): warn when sessions are stranded under an unclaimed profile namespace - #113884

Closed
PragmaticMeliorist wants to merge 1 commit into
NousResearch:mainfrom
PragmaticMeliorist:fix/session-orphan-guard
Closed

PragmaticMeliorist wants to merge 1 commit into
NousResearch:mainfrom
PragmaticMeliorist:fix/session-orphan-guard

Conversation

@PragmaticMeliorist

Copy link
Copy Markdown

What does this PR do?

Adds a boot-time warning when sessions sit under a session-key namespace that no currently-served profile claims.

Session keys are namespaced agent:<profile>:... (agent:main for the default profile, agent:<name> for a named one). rekey_profile_state() already migrates this namespace correctly, but only when a profile is explicitly renamed. On one of our boxes, gateway.multiplex_profiles flipped to default-on during an update — a configuration default change, not a rename — and every existing session key started resolving under a different namespace. 246 sessions carrying 9,764 messages stopped resolving. Nothing warned. Each chat simply opened empty on its next message, and "please continue" had nothing to continue from. We only found it by noticing chat history was gone.

This adds detection, not remediation: at boot, for each served profile's state.db, check whether any stored session key's namespace maps to a profile the live gateway doesn't currently serve, and log one actionable warning per orphaned namespace naming the existing hermes profiles migrate-identity remedy. It never rewrites the database and can't block boot — a missing/unreadable DB or any exception during the check returns {}/logs at DEBUG and boot proceeds. Choosing a migration target is deliberately left to the operator, since guessing wrong would merge two profiles' histories.

Related Issue

None filed; root-caused from a live incident on our own deployment. Related to the profile-rename rekey work (#111926/#111927) — this covers the un-announced case that rekey-on-rename doesn't reach.

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)
  • ✅ Tests (adding or improving test coverage)

Changes Made

  • New gateway/session_orphan_guard.py: find_orphaned_namespaces(db_path, live_profiles) reads sessions.session_key read-only and maps orphaned namespace -> count; warn_orphaned_namespaces() logs one warning per orphan and never raises.
  • GatewayProfileReconcileMixin: calls the guard once per boot for each served profile's state.db, after profile signatures are computed.
  • New test file: detection on an orphaned namespace, silence on the healthy case, and safety on a missing/corrupt database.

How to Test

scripts/run_tests.sh tests/gateway/test_session_orphan_guard.py -q

3 passed, 0 failed. Also ran the adjacent multiplex-serve and profile-signature suites to confirm no interaction with the boot path:

scripts/run_tests.sh tests/gateway/test_multiplex_hot_serve.py tests/gateway/test_profile_serve_signature.py tests/gateway/test_session_orphan_guard.py -q

7 passed, 0 failed.

Checklist

Code

  • I've read the Contributing Guide and repository agent instructions
  • My commit messages follow Conventional Commits
  • I searched for existing PRs to make sure this isn't a duplicate (searched session-namespace/multiplex_profiles/orphaned-session PRs; fix(gateway): parse named-profile namespaces in session keys (#105931) #105942 fixes named-profile namespace parsing in routing/sibling-match code, a different bug from an un-announced namespace change stranding existing sessions with no detection)
  • My PR contains only changes related to this fix
  • I've run the complete test suite and all tests pass (focused suite above)
  • I've added tests for my changes

Documentation & Housekeeping

  • Documentation update — N/A; this is a log-only detection, no user-facing config or workflow change
  • cli-config.yaml.example — N/A; no config keys changed
  • CONTRIBUTING.md / AGENTS.md — N/A
  • Cross-platform impact considered; sqlite3 read-only URI open works identically cross-platform
  • Tool descriptions/schemas — N/A

…ile namespace

Session keys carry an agent:<profile>: namespace, so a namespace change moves every
future key. rekey_profile_state() already migrates this correctly, but only runs on an
EXPLICIT profile rename — an un-announced change (a config default flip such as
gateway.multiplex_profiles defaulting on) silently strands every prior session. On one
box that was 246 sessions / 9,764 messages: each chat quietly opened empty instead of
continuing, with nothing logged.

Detect it at boot and log one actionable warning per orphaned namespace, naming the
migrate-identity remedy. Detection only: choosing a migration target is the operator's
call, since a wrong guess would merge two profiles' histories. Best-effort throughout —
a missing/unreadable DB returns {} and the check can never block boot.
@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/gateway Gateway runner, session dispatch, delivery area/profiles Multi-profile isolation, HERMES_HOME scoping area/sessions Session lifecycle, resume, persistence, history sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades labels Sep 17, 2026
teknium1 added a commit that referenced this pull request Sep 19, 2026
…file durable state

The per-profile store model (#88734), the parent-inheritance fence (#88381),
profile-stamped topic rows (#76423) and profile-prefixed voice keys (#75198)
are all forward-only: they put NEW state under the right profile and refuse to
widen existing damage, but nothing walks the stores and settles what earlier
releases left crossed. #113884 found 246 sessions stranded that way and could
only warn.

`hermes sessions repair-profiles` scans every profile's state.db plus the
gateway's voice-mode and sessions.json files and names six kinds of crossing:

1. `profile_name` disagreeing with the row's own session key -> relabel;
2. rows physically in another profile's store -> move (all message
   generations, usage rows, system prompt) to the owning store, parents before
   children so lineage survives, copy-then-delete so a crash leaves a duplicate
   the next run settles;
3. `parent_session_id` crossing namespaces -> sever (own identity kept);
4. routing rows outside the default store under multiplexing -> move (an
   existing row wins); routing rows for a profile that no longer exists -> drop;
5. Telegram topic bindings and voice-mode entries missing their bot's profile
   -> relabel from the sessions that hold the chat (ambiguous chats reported);
6. sessions.json mirror entries for an unclaimed namespace -> drop (the legacy
   import re-injects them into routing every boot).

Report-only by default. `--apply` refuses while a gateway owns any store, takes
a quick snapshot of every store first, and is idempotent. Two cases are
reported but never guessed: rows keyed to a profile that does not exist, and
`agent:main` rows inside a named profile's store (`--legacy-main rekey|move`
says which of the two histories they are).

Storage side lives in `hermes_state_profile_repair.py` (SessionDB mixin);
orchestration across stores in `hermes_cli/sessions_repair_profiles.py`; the
CLI face in `hermes_cli/sessions_cmd_repair_profiles.py` (pre-DB handler: it
opens every store itself).

Part of #88715 (PR-6). Closes the remediation gap #113884 only warns about.
@teknium1

Copy link
Copy Markdown
Collaborator

The stranded-history case is now repairable: hermes sessions repair-profiles --legacy-main rekey (#115689, 67757285f6) adopts agent:main: rows inside a named store as that profile's own history; dry-run first, snapshots before apply.

@teknium1 teknium1 closed this Sep 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/profiles Multi-profile isolation, HERMES_HOME scoping area/sessions Session lifecycle, resume, persistence, history comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants