fix(gateway): multiplex secondary profiles show as unreachable in per-profile status checks - #101487
Open
RicksCleaners wants to merge 2 commits into
Conversation
added 2 commits
September 2, 2026 17:03
…r multiplex A multiplex gateway (gateway.multiplex_profiles: true) is a single shared process serving several profiles from one bare `gateway run` command with no per-profile -p flag. Post-startup, only the active/default profile's own gateway_state.json ever got refreshed -- each secondary profile's file just rotted at whatever it last said (often a dead PID left over from before multiplexing took over, or before that profile was ever run standalone). Any UI/tool that scopes its liveness check to a specific profile's own gateway_state.json (e.g. a desktop app's per-profile connectivity check) would therefore report a genuinely healthy secondary profile as not-running, even though its bot/cron jobs work fine under the shared process. Two-part fix: 1. gateway/run.py: _start_secondary_profile_adapters now writes each secondary profile's own gateway_state.json after adapters connect, stamping multiplex_secondary=true. Uses write_runtime_status's new explicit path= kwarg rather than _profile_runtime_scope, since _get_process_hermes_home() deliberately ignores the HERMES_HOME contextvar override (NousResearch#56986) -- the scope-manager approach would silently no-op and rewrite the active profile's own file again. 2. gateway/status.py: _record_matches_live_gateway_pid now accepts a record's multiplex_secondary marker to skip its normal per-profile `-p <name>` command-line check, which can never match a multiplex process's bare shared argv. The (pid, start_time) match plus the 'looks like a gateway' cmdline check still guard against PID reuse; the marker is only ever gateway-written, never user-controllable. Regression tests in tests/gateway/test_multiplex_secondary_gateway_state.py cover both layers plus an end-to-end resolve_gateway_liveness() check. Verified live in production: after this fix + a graceful gateway restart, all 4 secondary profiles (previously showing dead PIDs from a stale pre-multiplex snapshot) resolve running=True via resolve_gateway_liveness.
…ve process The gateway_state.json top-level refresh (previous commit) fixed a secondary profile's own liveness check, but each per-platform entry (gateway_state.json's platforms.slack etc.) carries its OWN writer_pid/ writer_start_time fingerprint, separate from the top-level pid. hermes_cli/web_server.py's cross-profile /api/status aggregation (_owned_profile_platforms) only includes a platform entry when that fingerprint EXACTLY matches the profile's live gateway process. Left unstamped, a secondary profile's gateway_state.json can correctly say "running" while every platform entry still carries a stale/dead writer identity (e.g. a pre-multiplex standalone PID) -- so the profile reports zero connected platforms to any caller reading the aggregation. That reproduces "can't reach this profile's chat in the desktop app" even after the top-level fix, since Desktop's connectivity view goes through this aggregation. _start_secondary_profile_adapters now also calls write_runtime_status(platform=..., path=...) for each of that profile's currently-connected adapters right after the top-level refresh, so every platform entry gets re-stamped with the live process's own (writer_pid, writer_start_time). New regression test asserts a re-stamped platform entry's writer identity exactly matches the live process, mirroring _owned_profile_platforms's ownership check. Found live in production while investigating a recurrence of the same 'Pete can't open Summer's chat in Desktop' report: Summer's gateway_state.json top-level fields were already correctly refreshed (pid=292100, running), but her platforms.slack entry was still stamped writer_pid=193239 (a dead pre-multiplex PID) -- confirmed root cause via inspection before this fix.
10 tasks done
12 of 13 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Fixes a multiplex-gateway bug where a healthy SECONDARY profile (e.g. a
non-default bot identity served alongside the primary under
gateway.multiplex_profiles: true) shows as unreachable/not-running inany UI or tool that scopes its liveness/status check to that specific
profile — even though its bot and cron jobs are working fine under the
shared process.
Reported symptom: user could not open a secondary profile's chat in the
Desktop app, despite that profile's Slack bot responding normally and its
cron jobs running successfully.
Root cause (two layered bugs, both fixed here)
A multiplex gateway is a single shared process serving several profiles
from one bare
gateway runcommand with no per-profile-pflag.gateway/run.py—_start_secondary_profile_adapters()only everrefreshed the active/default profile's own
gateway_state.jsonafterstartup. Each secondary profile's own file just rotted at whatever it
last said — often a dead PID left over from before multiplexing took
over, or from before that profile was ever run standalone. Any caller
that scopes its liveness check to a specific profile's home directory
(e.g. a desktop app's per-profile connectivity check, or
resolve_gateway_liveness(profile_dir=...)) would therefore see astale/dead PID and report the profile as not running.
gateway/status.py— even after refreshing the top-level record,_record_matches_live_gateway_pid()validates a live PID's commandline against the profile's
-p <name>/--profile <name>flag — whicha multiplex gateway's single SHARED process never carries (it serves
every profile from one bare command line naming none of them
specifically). This check can never pass for a multiplex secondary
profile even given a freshly-written, genuinely-live record.
A third, subtler layer found while verifying the first two fixes live:
each PLATFORM entry inside
gateway_state.json(e.g.platforms.slack)carries its own
writer_pid/writer_start_timefingerprint, separatefrom the top-level
pid.hermes_cli/web_server.py's cross-profile/api/statusaggregation (_owned_profile_platforms) only includes aplatform entry when that fingerprint EXACTLY matches the profile's live
gateway process. Fixing only the top-level record (bugs 1+2) left each
platform entry still stamped with a stale writer identity, so the
profile correctly reported
runningwhile showing ZERO connectedplatforms to any caller reading the aggregation — reproducing the exact
"can't reach this profile's chat" symptom through a different code path.
_start_secondary_profile_adapters()now also re-stamps everycurrently-connected platform for that profile via
write_runtime_status(platform=..., path=...)right after thetop-level refresh.
Fix
gateway/status.py:_record_matches_live_gateway_pid()accepts a newmultiplex_secondarymarker on a record to skip the per-profile
-p <name>command-linecheck for that record. The live-cmdline "looks like a gateway" check
plus the
(pid, start_time)PID-reuse guard still apply — this markeris only ever gateway-written, never user-controllable input.
write_runtime_status()gains two new optional kwargs:multiplex_secondary: bool— stamps the marker above.path: Path— writes to an explicit path instead of the resolved_get_runtime_status_path(). Needed because_get_process_hermes_home()deliberately ignores theHERMES_HOMEcontextvar override (a documented anti-leak fix from a past issue —
gateway identity files must not follow an active per-session
profile-dispatch override into the wrong directory), so wrapping the
write in the existing
_profile_runtime_scope()context managersilently no-ops and rewrites the ACTIVE profile's own file again. An
explicit
pathis the only way to deliberately target anotherprofile's own
gateway_state.jsonfrom the active profile's process.Both new kwargs are additive/keyword-only — fully backward
compatible with all existing call sites.
gateway/run.py:_start_secondary_profile_adapters()now, for every served secondaryprofile, (a) writes that profile's own
gateway_state.jsonwithgateway_state="running"+multiplex_secondary=Truevia the newpath=kwarg, and (b) re-stamps each of that profile's currentlyconnected adapters (from
self._profile_adapters[profile_name]) viawrite_runtime_status(platform=..., platform_state="connected", path=...)so every platform entry's writer identity matches the liveprocess too.
Tests
tests/gateway/test_multiplex_secondary_gateway_state.py(new, 6 tests):write_runtime_statuscorrectly stamps/omits themultiplex_secondarymarker.
_record_matches_live_gateway_pidskips the per-profile cmdline checkwhen the marker is present, and (sibling/contrast test) still requires
it when absent — preserving the existing PID-reuse protection for the
"one dedicated process per profile" deployment that check was built for.
resolve_gateway_liveness()(the exactfunction status/dashboard surfaces call) reports the profile running.
(writer_pid, writer_start_time), matching what_owned_profile_platformsrequires to include them in the aggregation.All new tests pass; existing
tests/gateway/test_status.py(74 tests)and related cron/multiplex suites pass unchanged — no regressions.
Verification
Verified live in production on a real multiplex deployment serving 5
profiles: before the fix, all 4 secondary profiles'
gateway_state.jsonfiles carried dead PIDs from a prior pre-multiplex process generation,
and
resolve_gateway_liveness()scoped to each profile's home reportedrunning=False. After the fix + a graceful gateway restart, all 4 reportrunning=Truewith the live PID, and/api/status?profile=<name>showseach profile's platforms as
connectedwith the correct live-processwriter identity.