Skip to content

hermes gateway stop no longer exits 1 (and gets revived) when the stop watcher beats the CLI's SIGTERM - #120152

Merged
teknium1 merged 3 commits into
mainfrom
fix/gateway-stop-stays-planned
Sep 23, 2026
Merged

teknium1 merged 3 commits into
mainfrom
fix/gateway-stop-stays-planned

Conversation

@teknium1

Copy link
Copy Markdown
Collaborator

hermes gateway stop now always exits the gateway 0, so a Restart=on-failure (or any exit-code-driven) supervisor no longer revives a gateway the operator just stopped.

  • Root cause: the CLI writes the planned-stop marker, THEN sends SIGTERM. The gateway's planned-stop watcher (0.5 s poll) can fire in between; its shutdown call consumes the marker as planned, so the CLI's trailing SIGTERM finds no marker, is classified as an external kill, and the gateway exits 1.
  • Fix (gateway/run.py::_start_gateway_make_shutdown_signal_handler, +8 lines): the handler remembers that it already accepted a planned stop, and the trailing SIGTERM of that same stop stays planned. A --replace takeover still takes its own branch, and a bare SIGTERM with no planned stop before it is still an external kill (exit 1, supervisor revives).
  • Same fix covers every marker-then-signal path: hermes gateway stop, the orphan reaper, the launchd fallback, and the systemd ExecStop= marker from fix(gateway): direct systemctl restart/stop exits 0 instead of logging 'Failed with result exit-code' (#116551, salvage #116589) #116730, which also writes the marker before systemd delivers SIGTERM and so hit the same race.
  • Test: one invariant in tests/gateway/test_planned_stop_watcher.py drives the real handler with a real self-targeting marker file. The watcher's call consumes the marker, then the CLI's SIGTERM arrives. Control: a bare SIGTERM on a fresh handler is still flagged. Red on base (assert True is False), green with the fix.
  • The E2E that found this (core parity matrix, graceful_exit cell on gateway (fake adapter) + api_server, forces the worst-case interleaving) lands separately in the E2E-suite PR.

Live repro (real gateway processes, driven by the parity E2E driver: marker written for the gateway PID, SIGTERM sent only after the watcher consumed it):

gateway (fake adapter) api_server
before (origin/main 07646a7) parity cells red: ['graceful_exit'], host exit code: 1 parity cells red: ['graceful_exit'], host exit code: 1
after (this branch) green, exit 0 green, exit 0

Under load the natural interleaving hit this 1 in 24 stops. The driver now forces it every time.

Related, not superseded: #107230 (@lawcheck) and earlier #41690 / #41642 / #24351 target a different root cause: a systemd-initiated stop with no marker at all, which main already handles via the ExecStop= marker (#116730). None of them touches the watcher-consumes-marker race. #107230's parent-is-systemd heuristic would mask this race only under systemd, not for launchd, s6, a bare hermes gateway run under another supervisor, or the orphan reaper. No open or closed PR fixes this race (swept: planned stop marker, planned-stop watcher, gateway stop exit 1, consume_planned_stop_marker_for_self, shutdown_signal_handler, planned stop SIGTERM race).

Validation: scripts/run_tests.sh tests/gateway/ (8660 passed, 1 failed, 31 skipped. The one failure, test_session_hygiene_turnhold_adoption.py::test_turn_hold_keeps_admission_and_adopts_watermark_fenced_summary ('denied' == 'timeout'), is a load-dependent timing flake that also fails 2 of 4 standalone runs on unmodified origin/main at host load ~90. It doesn't touch the signal handler.); ruff check clean; check_no_tmp_literals.py gateway tests/gateway clean; git diff --check clean.

Infographic

gateway stop stays stopped

… CLI's SIGTERM

`hermes gateway stop` writes the planned-stop marker, then sends SIGTERM.
The gateway's planned-stop watcher (0.5 s poll) can fire in between: its
shutdown call consumes the marker as planned, and the CLI's SIGTERM that
follows finds no marker, is classified as an external kill and the gateway
exits 1 — which Restart=on-failure supervisors answer by reviving a gateway
the operator just stopped. Remember an accepted planned stop in the handler
so the trailing SIGTERM of the same stop is treated as planned.

Found by the core parity E2E matrix (gateway/api_server graceful_exit
cell), reproduced deterministically by signalling after the watcher has
consumed the marker.
…marker stays planned

Invariant for the planned-stop race: the watcher consumes the marker, the
CLI trailing SIGTERM must not read as an external kill; a bare SIGTERM with
no planned stop before it still does.
@teknium1 teknium1 added the ci-reviewed applied to manually approve dangerous changes label Sep 23, 2026
@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/gateway Gateway runner, session dispatch, delivery sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Sep 23, 2026
@teknium1
teknium1 merged commit 671c53c into main Sep 23, 2026
36 checks passed
@teknium1
teknium1 deleted the fix/gateway-stop-stays-planned branch September 23, 2026 13:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-reviewed applied to manually approve dangerous changes comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants