fix(cli): poll launchd for fresh PID to avoid false DOWN on macOS gateway respawn - #109250
Closed
salch-cred wants to merge 1 commit into
Closed
salch-cred wants to merge 1 commit into
salch-cred wants to merge 1 commit into
Conversation
10 tasks done
|
Closing: the patch is behaviourally identical to main. The new launchd probe is gated on |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #94743
Root Cause
After
hermes updatecallslaunchd kickstart(or a SIGUSR1 graceful restart), launchd acknowledges the restart and KeepAlive respawns the gateway nearly instantly. However,_collect_fleet_snapshotprobes fleet health usingcollect_fleet_versions, which readsgateway_state.jsonfrom disk. The new gateway process takes a few seconds to write that file.During that brief window,
collect_fleet_versionssees:gateway_state.jsonstill containing the old PID (now dead)pre_restart_pids→ triggersstate="down"The existing 30-second poll loop retries on
downrows, which should eventually resolve once the new process writes its state file. However, if startup is slow (e.g. heavy Python import graph, large model cache warm-up) and the 30s deadline passes before the file is written,hermes updateexits 1 even though launchd has already successfully respawned the service.Fix
Added
_launchd_fresh_pids_for_profiles(), which querieslaunchctl printfor each profile that produced adownrow. If launchd already supervises a live PID for that service (i.e. KeepAlive already respawned it), we know the gateway is settling — it's alive but hasn't yet writtengateway_state.json. In that case, the poll loopcontinues rather than accepting the staledownverdict, even if the 30-second deadline has passed.This correctly distinguishes:
The fix is macOS-only (
sys.platform == 'darwin'gate) and wrapped in broad exception handling so it cannot break updates on any other platform or in any failure mode.