Skip to content

fix(gateway): multiplex secondary profiles show as unreachable in per-profile status checks - #101487

Open
RicksCleaners wants to merge 2 commits into
NousResearch:mainfrom
RicksCleaners:fix/multiplex-secondary-profile-cron-and-status-sync
Open

RicksCleaners wants to merge 2 commits into
NousResearch:mainfrom
RicksCleaners:fix/multiplex-secondary-profile-cron-and-status-sync

Conversation

@RicksCleaners

Copy link
Copy Markdown

What

Fixes a multiplex-gateway bug where a healthy SECONDARY profile (e.g. a
non-default bot identity served alongside the primary under
gateway.multiplex_profiles: true) shows as unreachable/not-running in
any UI or tool that scopes its liveness/status check to that specific
profile — even though its bot and cron jobs are working fine under the
shared process.

Reported symptom: user could not open a secondary profile's chat in the
Desktop app, despite that profile's Slack bot responding normally and its
cron jobs running successfully.

Root cause (two layered bugs, both fixed here)

A multiplex gateway is a single shared process serving several profiles
from one bare gateway run command with no per-profile -p flag.

  1. gateway/run.py — _start_secondary_profile_adapters() only ever
    refreshed the active/default profile's own gateway_state.json after
    startup. Each secondary profile's own file just rotted at whatever it
    last said — often a dead PID left over from before multiplexing took
    over, or from before that profile was ever run standalone. Any caller
    that scopes its liveness check to a specific profile's home directory
    (e.g. a desktop app's per-profile connectivity check, or
    resolve_gateway_liveness(profile_dir=...)) would therefore see a
    stale/dead PID and report the profile as not running.

  2. gateway/status.py — even after refreshing the top-level record,
    _record_matches_live_gateway_pid() validates a live PID's command
    line against the profile's -p <name>/--profile <name> flag — which
    a multiplex gateway's single SHARED process never carries (it serves
    every profile from one bare command line naming none of them
    specifically). This check can never pass for a multiplex secondary
    profile even given a freshly-written, genuinely-live record.

  3. A third, subtler layer found while verifying the first two fixes live:
    each PLATFORM entry inside gateway_state.json (e.g. platforms.slack)
    carries its own writer_pid/writer_start_time fingerprint, separate
    from the top-level pid. hermes_cli/web_server.py's cross-profile
    /api/status aggregation (_owned_profile_platforms) only includes a
    platform entry when that fingerprint EXACTLY matches the profile's live
    gateway process. Fixing only the top-level record (bugs 1+2) left each
    platform entry still stamped with a stale writer identity, so the
    profile correctly reported running while showing ZERO connected
    platforms to any caller reading the aggregation — reproducing the exact
    "can't reach this profile's chat" symptom through a different code path.
    _start_secondary_profile_adapters() now also re-stamps every
    currently-connected platform for that profile via
    write_runtime_status(platform=..., path=...) right after the
    top-level refresh.

Fix

  • gateway/status.py:

    • _record_matches_live_gateway_pid() accepts a new multiplex_secondary
      marker on a record to skip the per-profile -p <name> command-line
      check for that record. The live-cmdline "looks like a gateway" check
      plus the (pid, start_time) PID-reuse guard still apply — this marker
      is only ever gateway-written, never user-controllable input.
    • write_runtime_status() gains two new optional kwargs:
      • multiplex_secondary: bool — stamps the marker above.
      • path: Path — writes to an explicit path instead of the resolved
        _get_runtime_status_path(). Needed because
        _get_process_hermes_home() deliberately ignores the HERMES_HOME
        contextvar override (a documented anti-leak fix from a past issue —
        gateway identity files must not follow an active per-session
        profile-dispatch override into the wrong directory), so wrapping the
        write in the existing _profile_runtime_scope() context manager
        silently no-ops and rewrites the ACTIVE profile's own file again. An
        explicit path is the only way to deliberately target another
        profile's own gateway_state.json from the active profile's process.
        Both new kwargs are additive/keyword-only — fully backward
        compatible with all existing call sites.
  • gateway/run.py:

    • _start_secondary_profile_adapters() now, for every served secondary
      profile, (a) writes that profile's own gateway_state.json with
      gateway_state="running" + multiplex_secondary=True via the new
      path= kwarg, and (b) re-stamps each of that profile's currently
      connected adapters (from self._profile_adapters[profile_name]) via
      write_runtime_status(platform=..., platform_state="connected", path=...) so every platform entry's writer identity matches the live
      process too.

Tests

tests/gateway/test_multiplex_secondary_gateway_state.py (new, 6 tests):

  • write_runtime_status correctly stamps/omits the multiplex_secondary
    marker.
  • _record_matches_live_gateway_pid skips the per-profile cmdline check
    when the marker is present, and (sibling/contrast test) still requires
    it when absent — preserving the existing PID-reuse protection for the
    "one dedicated process per profile" deployment that check was built for.
  • End-to-end: after a refresh, resolve_gateway_liveness() (the exact
    function status/dashboard surfaces call) reports the profile running.
  • Platform entries are correctly re-stamped to the live process's own
    (writer_pid, writer_start_time), matching what
    _owned_profile_platforms requires to include them in the aggregation.

All new tests pass; existing tests/gateway/test_status.py (74 tests)
and related cron/multiplex suites pass unchanged — no regressions.

Verification

Verified live in production on a real multiplex deployment serving 5
profiles: before the fix, all 4 secondary profiles' gateway_state.json
files carried dead PIDs from a prior pre-multiplex process generation,
and resolve_gateway_liveness() scoped to each profile's home reported
running=False. After the fix + a graceful gateway restart, all 4 report
running=True with the live PID, and /api/status?profile=<name> shows
each profile's platforms as connected with the correct live-process
writer identity.

Hermes VPS Agent added 2 commits September 2, 2026 17:03
…r multiplex

A multiplex gateway (gateway.multiplex_profiles: true) is a single shared
process serving several profiles from one bare `gateway run` command with
no per-profile -p flag. Post-startup, only the active/default profile's
own gateway_state.json ever got refreshed -- each secondary profile's file
just rotted at whatever it last said (often a dead PID left over from
before multiplexing took over, or before that profile was ever run
standalone).

Any UI/tool that scopes its liveness check to a specific profile's own
gateway_state.json (e.g. a desktop app's per-profile connectivity check)
would therefore report a genuinely healthy secondary profile as
not-running, even though its bot/cron jobs work fine under the shared
process.

Two-part fix:

1. gateway/run.py: _start_secondary_profile_adapters now writes each
   secondary profile's own gateway_state.json after adapters connect,
   stamping multiplex_secondary=true. Uses write_runtime_status's new
   explicit path= kwarg rather than _profile_runtime_scope, since
   _get_process_hermes_home() deliberately ignores the HERMES_HOME
   contextvar override (NousResearch#56986) -- the scope-manager approach would
   silently no-op and rewrite the active profile's own file again.

2. gateway/status.py: _record_matches_live_gateway_pid now accepts a
   record's multiplex_secondary marker to skip its normal per-profile
   `-p <name>` command-line check, which can never match a multiplex
   process's bare shared argv. The (pid, start_time) match plus the
   'looks like a gateway' cmdline check still guard against PID reuse;
   the marker is only ever gateway-written, never user-controllable.

Regression tests in tests/gateway/test_multiplex_secondary_gateway_state.py
cover both layers plus an end-to-end resolve_gateway_liveness() check.

Verified live in production: after this fix + a graceful gateway restart,
all 4 secondary profiles (previously showing dead PIDs from a stale
pre-multiplex snapshot) resolve running=True via resolve_gateway_liveness.
…ve process

The gateway_state.json top-level refresh (previous commit) fixed a
secondary profile's own liveness check, but each per-platform entry
(gateway_state.json's platforms.slack etc.) carries its OWN writer_pid/
writer_start_time fingerprint, separate from the top-level pid.

hermes_cli/web_server.py's cross-profile /api/status aggregation
(_owned_profile_platforms) only includes a platform entry when that
fingerprint EXACTLY matches the profile's live gateway process. Left
unstamped, a secondary profile's gateway_state.json can correctly say
"running" while every platform entry still carries a stale/dead writer
identity (e.g. a pre-multiplex standalone PID) -- so the profile reports
zero connected platforms to any caller reading the aggregation. That
reproduces "can't reach this profile's chat in the desktop app" even
after the top-level fix, since Desktop's connectivity view goes through
this aggregation.

_start_secondary_profile_adapters now also calls
write_runtime_status(platform=..., path=...) for each of that profile's
currently-connected adapters right after the top-level refresh, so every
platform entry gets re-stamped with the live process's own
(writer_pid, writer_start_time).

New regression test asserts a re-stamped platform entry's writer identity
exactly matches the live process, mirroring _owned_profile_platforms's
ownership check.

Found live in production while investigating a recurrence of the same
'Pete can't open Summer's chat in Desktop' report: Summer's
gateway_state.json top-level fields were already correctly refreshed
(pid=292100, running), but her platforms.slack entry was still stamped
writer_pid=193239 (a dead pre-multiplex PID) -- confirmed root cause via
inspection before this fix.
@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/gateway Gateway runner, session dispatch, delivery area/profiles Multi-profile isolation, HERMES_HOME scoping sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Sep 2, 2026
@alt-glitch

Copy link
Copy Markdown

This was generated by AI during triage.

Related: multiplex profile-state family #88047, #97120, and #86946. This PR specifically repairs per-profile liveness identity for the shared process.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/profiles Multi-profile isolation, HERMES_HOME scoping comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants