Skip to content

fix(gateway): keep a strong reference to the SIGTERM/SIGINT shutdown task - #83864

Open
briandevans wants to merge 7 commits into
NousResearch:mainfrom
briandevans:fix/gateway-shutdown-task-anchor-27768
Open

briandevans wants to merge 7 commits into
NousResearch:mainfrom
briandevans:fix/gateway-shutdown-task-anchor-27768

Conversation

@briandevans

@briandevans briandevans commented Aug 11, 2026 •

Copy link
Copy Markdown

What does this PR do?

shutdown_signal_handler in gateway/run.py is installed for both SIGINT and SIGTERM, and it ended by scheduling the shutdown and throwing the handle away:

asyncio.create_task(runner.stop())

This repo already argues that this is a bug — on the other signal path. request_restart() (the SIGUSR1 restart handler) creates a task of exactly the same shape, deliberately keeps a reference to it, and writes out why:

We still hold a strong reference in self._restart_task: a bare
asyncio.create_task() keeps only a weak reference, so the event
loop may garbage-collect a still-pending task mid-flight. The
cancel loop in _stop_impl explicitly skips _restart_task for the
same reason it skips _stop_task.

One signal path was hardened against that hazard and the comment explaining why is sitting in the file. The other — the path every systemctl stop, every hermes gateway stop and every interactive Ctrl+C takes — was left bare. This PR completes that hardening by mirroring _restart_task.

The failure window, stated precisely

stop() is two-stage and only the inner stage is self-anchoring:

  • the inner task is created as self._stop_task = asyncio.create_task(_stop_impl()) and awaited, so it is strongly referenced;
  • the outer task — the runner.stop() coroutine the signal handler wraps — had no reference at all.

So the damaging window is narrow and specific: if the outer task is collected before stop() reaches self._stop_task = asyncio.create_task(_stop_impl()), _stop_impl is never created and the gateway never tears down. It keeps running until systemd's TimeoutStopSec escalates to SIGKILL — in-flight turns are never drained, sessions are never finalized, and the operator sees "I sent stop and it hung, then it got killed and I lost the turn."

A collection after that point is harmless: the inner task survives on self._stop_task and shutdown completes normally. I am not claiming that every collection loses the shutdown, and this is a latent-hazard fix in the same sense the _restart_task one was — the asyncio docs and this file's own comment both say the handle must be kept.

Why the obvious fix is the wrong one

The naive version — parking the task in self._background_tasks — is self-cancelling. _stop_impl's teardown loop cancels every member of that set and skips exactly two handles, _stop_task and _restart_task. Putting the shutdown handle there would have _stop_impl cancel the very task that is awaiting it, trading a garbage-collection bug for a cancellation bug. That is why this uses a dedicated _shutdown_task and adds the matching exemption instead.

Why the assignment is guarded

shutdown_signal_handler is not a once-only path — a second Ctrl+C, a SIGINT followed by the service manager's SIGTERM, or _run_planned_stop_watcher driving the same callable from its polling thread via loop.call_soon_threadsafe(shutdown_handler, None) all re-enter it. (That watcher's own docstring notes that on POSIX "the signal handler always races us to consuming the marker file", and its _draining gate does not close until _stop_impl is already running.) An unconditional assignment would let the second call overwrite the reference to the task performing the live teardown, putting us straight back in the window above. Skipping the duplicate is behaviour-preserving: stop() short-circuits on self._stop_task, so the second task only ever existed to await the first one's work.

Scope

Deliberately fenced to the signal-handler statement only. There are two other open PRs anchoring create_task sites in this file — #17966 and #45372, both on the watcher sites (_run_process_watcher, _session_expiry_watcher, _platform_reconnect_watcher and friends inside start()). Neither touches the signal handler, and this PR does not touch any watcher site, so all three apply in any order. The regression tests live in a new file for the same reason: #41642 and #41690 are both adding their own test files in this area, and a shared test file would be a needless conflict.

Related Issue

No filed issue — found by auditing the two signal paths in gateway/run.py against each other, after the SIGUSR1 path's comment made the invariant explicit.

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)
  • ✨ New feature (non-breaking change that adds functionality)
  • 🔒 Security fix
  • 📝 Documentation update
  • ✅ Tests (adding or improving test coverage)
  • ♻️ Refactor (no behavior change)
  • 🎯 New skill (bundled or hub)

Changes Made

Six atomic commits:

  • gateway/run.py — anchor the task the SIGINT/SIGTERM handler creates in runner._shutdown_task, and declare the attribute at both existing sites (the class annotation beside _stop_task / _restart_task, and __init__) so bare runners built via object.__new__ inherit the same None default.
  • gateway/run.py — only re-point that anchor when the previous shutdown task is absent or done(), so a repeat signal cannot drop the handle to the live teardown.
  • gateway/run.py — add _shutdown_task to _stop_impl's cancel-sweep skip list, beside _stop_task and _restart_task. Like _restart_task, the handle is deliberately kept out of _background_tasks, so this branch does not fire on any current path; it is recorded there because that loop is the single place enforcing "never cancel a task that is awaiting _stop_task", and the handles satisfying that description should not be treated differently by it.
  • tests/gateway/test_shutdown_task_anchor.py — regression tests.
  • gateway/run.py — extract the anchoring and the guard out of the handler closure into GatewayRunner._schedule_shutdown_task(), no behaviour change. The handler is built inside the gateway start path and cannot be imported, which had pushed the tests into asserting the shape of the source; AGENTS.md bans that ("Never read source code in tests"), so the behaviour is made importable instead.
  • tests/gateway/test_shutdown_task_anchor.py — replace the four source-parsing assertions with behavioural tests against that method.

How to Test

uv run --with pytest --with pytest-asyncio python3 -m pytest tests/gateway/test_shutdown_task_anchor.py -v

Five tests, all behavioural — no source text is read. Every production hunk is mutation-covered, i.e. each one is independently load-bearing:

mutation applied to gateway/run.py result
return the task without storing it in _shutdown_task 3 tests fail
drop the "previous shutdown still in flight" guard 1 test fails (..._repeat_signal_does_not_replace_the_in_flight_shutdown_task)
delete the _shutdown_task branch from _stop_impl's cancel sweep 1 test fails (test_stop_does_not_cancel_the_anchored_shutdown_task)
none 5 / 5 pass

What they assert:

  • scheduling a shutdown leaves the task reachable from the runner and not only from the event loop — that reachability is the whole fix;
  • a repeat signal reuses the in-flight task and does not start a second stop();
  • once a shutdown has finished, a later signal does schedule again, so the guard cannot wedge the gateway;
  • _stop_impl's cancel sweep leaves _shutdown_task alone while still cancelling an ordinary background task — driven against a real runner.stop() on the restart_test_helpers bare-runner harness.

Adjacent suites run green on this branch: test_gateway_shutdown.py, test_planned_stop_watcher.py, test_restart_drain.py, test_clean_shutdown_marker.py, test_shutdown_cache_cleanup.py, test_startup_restart_race.py, test_restart_resume_pending.py, test_background_command.py, test_api_server_active_work_drain.py, test_cron_active_work_drain.py, test_session_race_guard.py, test_max_concurrent_sessions.py, test_gateway_process_exit.py, test_external_drain_control.py — 121 passed, 2 skipped.

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix/feature (no unrelated commits)
  • I've run pytest tests/ -q and all tests pass — ran the focused + adjacent suites listed above, not the full suite
  • I've added tests for my changes (required for bug fixes, strongly encouraged for features)
  • I've tested on my platform: macOS (Darwin 25.4), Python 3.11

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings) — or N/A (inline comments only; no user-facing docs affected)
  • I've updated cli-config.yaml.example if I added/changed config keys — or N/A
  • I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows — or N/A
  • I've considered cross-platform impact (Windows, macOS) per the compatibility guide — the handler is unreachable on Windows (add_signal_handler raises NotImplementedError there), but _run_planned_stop_watcher invokes the same callable on every platform, so the anchor and the repeat-call guard apply to the Windows fallback path too. Verified on macOS only.
  • I've updated tool descriptions/schemas if I changed tool behavior — or N/A

`shutdown_signal_handler` ended with a bare
`asyncio.create_task(runner.stop())` and discarded the handle. The event
loop keeps only a weak reference to a task, so a still-pending task can
be garbage-collected mid-flight.

This file already states that hazard, and already fixes it — on the other
signal path. `request_restart()` (SIGUSR1) creates a task with the same
shape and deliberately keeps it in `self._restart_task`, with the reason
written out inline: "a bare asyncio.create_task() keeps only a weak
reference, so the event loop may garbage-collect a still-pending task
mid-flight." The path that every `systemctl stop`, `hermes gateway stop`
and interactive Ctrl+C takes was left bare.

The damaging window is precise: `stop()` only becomes self-anchoring once
it reaches `self._stop_task = asyncio.create_task(_stop_impl())`. A
collection before that point means `_stop_impl` is never created at all,
so the gateway does not tear down — it keeps running until systemd's
TimeoutStopSec escalates to SIGKILL, with in-flight turns undrained and
sessions never finalized. A collection after that point is harmless,
because the inner task is strongly referenced by `self._stop_task`.

Adds `_shutdown_task` alongside `_stop_task` / `_restart_task` at both
existing declaration sites (class annotation and `__init__`), so bare
runners built via `object.__new__` in the shutdown-path tests get the
same `None` default the other two have.
`shutdown_signal_handler` is not a once-only path, so assigning
`runner._shutdown_task` unconditionally would let a second invocation
overwrite the reference to the task that is actually performing the
teardown, re-opening the collection window the anchor exists to close.

Three ordinary ways it fires twice:

- a second Ctrl+C while the first drain is still running;
- an interactive SIGINT followed by the service manager's SIGTERM;
- `_run_planned_stop_watcher`, which drives the same callable from its
  polling thread via `loop.call_soon_threadsafe(shutdown_handler, None)`.
  Its `_draining` gate does not close until `_stop_impl` is already
  running, and its own docstring records that on POSIX "the signal
  handler always races us to consuming the marker file".

Re-point the anchor only when the previous task is absent or done. This
is behaviour-preserving for the duplicate call: `stop()` short-circuits
on `self._stop_task` and awaits the first teardown, so the second task
never did work of its own — it only existed to await the first.
`_stop_impl` cancels every entry in `self._background_tasks` and skips
exactly two handles: `_stop_task`, and `_restart_task` because it "is
awaiting _stop_task right now; cancelling it would propagate
CancelledError into this _stop_impl and skip _shutdown_event.set() /
_exit_code = 75 (NousResearch#12875)".

`_shutdown_task` has that same relationship — it holds the outer
`runner.stop()` coroutine, which is parked in `await self._stop_task`
while this loop runs. It is the reason the naive fix for the missing
anchor is wrong: parking the handle in `_background_tasks` would trade a
garbage-collection bug for a cancellation bug, with `_stop_impl`
cancelling the very task waiting on it.

Like `_restart_task`, `_shutdown_task` is deliberately kept out of
`_background_tasks`, so this branch does not fire on any current path.
It is recorded here because this loop is the single place that enforces
the "never cancel a task awaiting `_stop_task`" invariant, and the two
handles that satisfy that description should not be treated differently
by it.
…mption

Five regression tests, all red on the unfixed file and green after,
each pinned to one of the three production hunks:

- `_shutdown_task` exists as a class-level default beside `_stop_task`
  and `_restart_task`, so the bare runners the shutdown-path tests build
  via `object.__new__` inherit `None`;
- `shutdown_signal_handler` contains no `asyncio.create_task(...)` whose
  result is discarded;
- the task it creates is anchored on the runner as `_shutdown_task`;
- that assignment is guarded by a `done()` check, so a repeat signal
  cannot replace the handle to a still-running teardown;
- `_stop_impl`'s cancel sweep leaves `_shutdown_task` alone while still
  cancelling an ordinary background task.

The last one runs against a real `stop()` on the bare-runner harness in
`restart_test_helpers`. The first four are asserted by parsing
`gateway/run.py` with `ast`, because `shutdown_signal_handler` is a
closure built inside the gateway start path and cannot be imported —
the same technique `test_adapter_connect_is_reconnect_contract.py`
already uses for a contract that is likewise unreachable by import.

Lives in its own file rather than extending `test_gateway_shutdown.py`
so it does not collide with the other open changes in this area.
Copilot AI lite review requested due to automatic review settings August 11, 2026 11:33

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR hardens the gateway shutdown path by keeping a strong reference to the runner.stop() task created by the SIGINT/SIGTERM shutdown handler, mirroring the existing _restart_task pattern used for SIGUSR1 restart. This prevents a narrow but real failure window where the outer stop() task could be garbage-collected before it becomes self-anchoring, leading to hung shutdowns until forced SIGKILL.

Changes:

  • Add GatewayRunner._shutdown_task and anchor the SIGINT/SIGTERM shutdown task to it (with a done()-guard to avoid overwriting a live shutdown).
  • Ensure _stop_impl’s background-task cancel sweep skips _shutdown_task (like _stop_task / _restart_task).
  • Add a regression test file covering the anchoring and cancel-sweep behavior.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

File Description
gateway/run.py Anchors SIGINT/SIGTERM shutdown task on the runner and exempts it from the shutdown cancel sweep.
tests/gateway/test_shutdown_task_anchor.py Adds regression coverage for shutdown-task anchoring and cancel-sweep behavior (includes source-structure assertions).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +42 to +45
def _shutdown_signal_handler_node() -> ast.FunctionDef:
"""The single ``shutdown_signal_handler`` definition in ``gateway/run.py``."""
tree = ast.parse(RUN_PY.read_text(encoding="utf-8"))
matches = [

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, and fixed — I checked AGENTS.md and "Never read source code in tests" is a hard ban, so this was a real defect in the PR rather than a style preference. Reading the source was a workaround for shutdown_signal_handler being a closure inside the gateway start path; the right move is to make the behaviour importable, not to assert on text.

Done in 101874c80aa (refactor) + 4a24ede20f4 (tests), current head 4a24ede20f4:

  • GatewayRunner._schedule_shutdown_task() now owns the create-and-anchor and the "don't re-point a live shutdown" check, and returns the task that owns the shutdown. The handler is a one-line call to it. No behaviour change.
  • All four source-parsing assertions are gone. The replacements call the method: test_scheduling_shutdown_anchors_the_task_on_the_runner, test_a_repeat_signal_does_not_replace_the_in_flight_shutdown_task, test_a_later_signal_schedules_again_once_the_previous_one_finished, plus the already-behavioural test_stop_does_not_cancel_the_anchored_shutdown_task — all in tests/gateway/test_shutdown_task_anchor.py.

Each production hunk is mutation-covered so none of these can pass against a broken implementation: returning the task without storing it fails 3 tests, dropping the in-flight guard fails the repeat-signal test, and deleting the _shutdown_task branch from _stop_impl's cancel sweep fails the sweep test.

@briandevans

Copy link
Copy Markdown
Author

CI audit — the single failure on this branch is a pre-existing baseline on clean origin/main (c0106e50e). Zero failures are in touched code.

Check Symptom Root cause on main
Python tests / Run tests slice 5/12 FAILED tests/gateway/test_multiplex_busy_input_mode.py::test_profile_route_and_nonmultiplexed_resolution_preserve_boundaries Same test, same slice, fails identically on origin/main itself (c0106e50e, job 93715438449 — 232 files, 2672 passed, 1 failed). Nothing in this PR touches test_multiplex_busy_input_mode.py or the multiplex profile-routing path.

The slice that actually runs this PR's code is 8/12, and it is green: ✓ tests/gateway/test_shutdown_task_anchor.py (5✓), slice total 3272 tests passed, 0 failed.

…ndler

Move the anchoring and the repeat-signal guard out of the
`shutdown_signal_handler` closure and onto `GatewayRunner`, so the
behaviour is reachable from a test.

The handler is built inside the gateway start path and cannot be
imported, which had pushed its regression tests into asserting the shape
of `gateway/run.py`'s source. AGENTS.md bans that outright ("Never read
source code in tests" — it passes when the implementation is subtly
broken and fails on correct refactors), so the fix is to make the
behaviour importable rather than to assert on text. The following commit
replaces those tests with behavioural ones against this method.

No behaviour change: the handler now calls
`runner._schedule_shutdown_task()`, which performs the same
create-and-anchor and the same "don't re-point a live shutdown" check,
and additionally returns the task that owns the shutdown so callers can
await or inspect it.
The previous revision of this file asserted structural properties by
reading and `ast.parse`-ing `gateway/run.py`, because the behaviour lived
in a closure that could not be imported. AGENTS.md bans that outright
("Never read source code in tests"): such a test passes when the
implementation is subtly broken and fails on a correct refactor.

The preceding commit made the behaviour importable as
`GatewayRunner._schedule_shutdown_task`, so all four structural
assertions are replaced by tests that call it:

- scheduling a shutdown leaves the task reachable from the runner, not
  only from the event loop, which is what stops it being collected while
  still pending;
- a repeat signal reuses the in-flight task and does not start a second
  `stop()`;
- once the previous shutdown has finished, a later signal does schedule
  again, so the guard cannot wedge the gateway;
- `_stop_impl`'s cancel sweep leaves `_shutdown_task` alone while still
  cancelling an ordinary background task (unchanged, already behavioural).

Each production hunk is mutation-covered: dropping the anchor fails 3
tests, dropping the repeat-signal guard fails the repeat test, and
removing the cancel-sweep exemption fails the sweep test.
@alt-glitch alt-glitch added type/bug Something isn't working comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Aug 11, 2026
@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference, author can ignore or act on any point.

fix(gateway): keep a strong reference to the SIGTERM/SIGINT shutdown task

  1. The anchored _shutdown_task is never awaited and gets no done-callback: if stop() raises, the exception sits un-retrieved on the task ("Task exception was never retrieved" noise, and a teardown failure that is easy to miss). Consider task.add_done_callback(...) that logs the exception.
  2. After a completed shutdown, _shutdown_task still holds the finished task; a late repeat signal creates a NEW task that calls stop(), which short-circuits on self._stop_task. Worth confirming that short-circuit path returns promptly rather than hanging when the loop is already winding down (a create_task on a stopping loop can also raise).
  3. Confirmed: _run_planned_stop_watcher fires via loop.call_soon_threadsafe(shutdown_handler, None) (gateway/run.py:27269), so the handler always runs on the loop thread — no cross-thread create_task hazard. The re-entrancy guard (returning the in-flight task instead of re-pointing the anchor) is the right call, and the test coverage (strong ref, no re-point on repeat signal, cancel-sweep exemption) is thorough.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants