Skip to content

fix(gateway): preserve /update completion across restart - #80172

Open
boudyai wants to merge 1 commit into
NousResearch:mainfrom
boudyai:fix/gateway-update-restart-notification
Open

boudyai wants to merge 1 commit into
NousResearch:mainfrom
boudyai:fix/gateway-update-restart-notification

Conversation

@boudyai

@boudyai boudyai commented Aug 6, 2026 •

Copy link
Copy Markdown

What does this PR do?

Makes /update completion notifications survive the gateway restart they report.

The updater writes .update_exit_code before asking systemd (or another platform supervisor) to restart the gateway. The old gateway watcher can therefore observe success while shutdown is already in progress. It currently attempts the final send and then unconditionally deletes .update_pending.json, .update_output.txt, and .update_exit_code, including when delivery fails. The new gateway starts with no durable target left and cannot notify the initiating chat.

This PR turns success notification into a two-phase lifecycle:

  1. The initiating gateway records its PID in the pending marker and may announce that restart is beginning.
  2. It preserves the completion markers instead of consuming them.
  3. The freshly started gateway sends the definitive “update finished / gateway online” notification and cleans up only after confirmed delivery.

Both raised exceptions and SendResult(success=False) preserve and restore the marker for retry. Legacy markers without an initiating PID retain their existing behavior.

Related Issue

N/A — no matching open or closed issue/PR was found after searching update, restart, completion-notification, and marker-related terms.

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)
  • ✨ New feature (non-breaking change that adds functionality)
  • 🔒 Security fix
  • 📝 Documentation update
  • ✅ Tests (adding or improving test coverage)
  • ♻️ Refactor (no behavior change)
  • 🎯 New skill (bundled or hub)

Changes Made

  • gateway/slash_commands.py: persist the initiating gateway PID in the durable update target.
  • gateway/run.py: defer success to the restarted gateway, preserve markers on exceptions and soft send failures, and explicitly report that the gateway is back online.
  • tests/gateway/test_update_command.py: cover PID handoff, successful post-restart delivery, cleanup, exception retry, and SendResult(success=False) retry.
  • tests/gateway/test_update_streaming.py: reproduce the old-watcher race and verify soft-failure retries.

How to Test

  1. Start a messaging gateway under a restart-capable supervisor and invoke /update.

  2. Confirm the initiating process can send Update installed / Restarting gateway without deleting the durable markers.

  3. Confirm the new gateway PID sends Update finished / Gateway restarted and is back online, then removes the markers.

  4. Make adapter.send() raise or return SendResult(success=False) and confirm the pending marker remains retryable.

  5. Run:

    scripts/run_tests.sh \
      tests/gateway/test_update_command.py \
      tests/gateway/test_update_streaming.py \
      tests/gateway/test_startup_restart_race.py -q

    Result: 27 passed.

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix/feature (no unrelated commits)
  • I've run pytest tests/ -q and all tests pass
  • I've added tests for my changes (required for bug fixes, strongly encouraged for features)
  • I've tested on my platform: Linux/systemd reproduction; macOS 26.5.1 targeted test suite

The full repository suite was not run locally. The three directly related test files pass via the canonical runner, and Ruff passes for every changed Python file.

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings) — N/A; no user-facing configuration or syntax changed
  • I've updated cli-config.yaml.example if I added/changed config keys — N/A; no config keys changed
  • I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows — N/A; the existing marker/watcher architecture is retained
  • I've considered cross-platform impact (Windows, macOS) per the compatibility guide — os.getpid() and durable marker handoff are platform-neutral
  • I've updated tool descriptions/schemas if I changed tool behavior — N/A; no model tool changed

Screenshots / Logs

Before:

.update_exit_code appears before supervisor restart
→ old gateway attempts adapter.send()
→ shutdown/network race fails delivery
→ completion markers are deleted anyway
→ new gateway has no notification target

Validation:

Summary: 3 files, 27 tests passed, 0 failed
Ruff: All checks passed

Make gateway updates use a two-phase notification lifecycle. Persist the initiating gateway PID with the durable update target, let the old process announce that restart is beginning without consuming completion state, and let only the freshly started gateway send the definitive online notification.

Preserve pending, output, and exit-code markers when platform delivery raises or returns SendResult(success=False), so reconnect/startup watchers can retry instead of silently losing the result. Clean markers only after confirmed delivery.

This closes the race where the updater writes its success marker before the supervisor stops the old gateway: the old watcher previously attempted a final send during shutdown and unconditionally deleted all markers even when delivery failed.

Validation: 27 gateway update, streaming, and startup-race tests passed; Ruff and diff checks passed.
@boudyai
boudyai force-pushed the fix/gateway-update-restart-notification branch from f73068d to 401227e Compare August 6, 2026 08:21
@alt-glitch alt-glitch added type/bug Something isn't working comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades labels Aug 6, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants