Skip to content

fix(gateway): give up on an update notice whose platform never connects - #113756

Closed
gkd2323c wants to merge 1 commit into
NousResearch:mainfrom
gkd2323c:fix/gateway-update-notification-stale-marker
Closed

gkd2323c wants to merge 1 commit into
NousResearch:mainfrom
gkd2323c:fix/gateway-update-notification-stale-marker

Conversation

@gkd2323c

Copy link
Copy Markdown

Bug

_send_update_notification defers with no upper bound when the marker's platform has no adapter. That defer is right for the case it was written for: /update restarts the gateway, so the requesting platform's adapter is briefly absent when the new process checks, and dropping the markers there would silently lose the "update finished" notice.

But when no adapter will ever appear — the marker names a platform this install does not run — the defer never resolves. _send_update_notification returns False, so run_startup sees the markers still on disk and calls _schedule_update_notification_watch(); the watcher polls every 2s and re-logs

Update notification deferred: telegram adapter not connected yet

Nothing breaks that cycle, because the markers can never be delivered and so are never cleared.

Concretely: one install had a marker written 2026-07-08 still doing this today, every 2 seconds, across every gateway restart — about 19,000 lines of gateway.log. The marker named a platform with no configured credentials, so no adapter was ever going to appear.

Root cause

gateway/run_notifications.py, in _send_update_notification:

if chat_id and not adapter:
    # Target platform not reconnected yet (common right after the update's restart): keep the
    # markers for a later retry instead of silently losing the notification.
    return _defer("Update notification deferred: %s adapter not connected yet", platform_str)

The comment reads the absence as temporary ("not reconnected yet"). Nothing checks that assumption, and _defer renames claimed back to pending each time, so the markers stay on disk and the next boot re-arms the watcher.

Fix

Bound the wait using the timestamp the marker already carries. slash_commands.py stamps datetime.now().isoformat() when it writes the marker, and that value survives the pending ↔ claimed rename — the file mtime does not, which rules the mtime out as an age source. Once the marker is older than the cap the notice is undeliverable by any retry, so it is logged once at WARNING, the markers are cleared, and the method returns True, the definitive answer run_startup keys off to stop rescheduling. Markers with no parseable timestamp (written before the field existed) keep the previous retry behavior rather than have the code guess an age.

Changed in gateway/run_notifications.py: the if chat_id and not adapter: branch, a new _marker_age_seconds() helper beside _marker_profile(), and a new _UPDATE_NOTIFY_MAX_ADAPTER_WAIT_SECONDS constant.

Tests

tests/gateway/test_update_command.py gains three cases in TestSendUpdateNotification:

  • a marker past the cap is dropped, every marker file is removed, and a WARNING is emitted
  • a marker inside the cap is still held — the bound must not swallow the notice it exists to protect
  • a marker with no timestamp still retries

Reverting the source change alone makes the first case fail with assert False is True, which is the reported behavior. The file's other 17 cases pass, as do test_platform_reconnect.py, test_update_streaming.py, test_startup_restart_race.py and test_heartbeat_watch_restore.py (69 total).

Risk

Low. The new branch is reachable only when the adapter is already missing (today's defer) and the marker is older than an hour, so no currently-deliverable notification changes path. The threshold is a plain module constant with no config surface, and the helper returns None on anything it cannot parse.

@whyyagswhy

Copy link
Copy Markdown

Independent verification on the PR head (26da902): tests/gateway/test_update_command.py passes locally, 20/20 on Linux (canonical runner).

Bounded wait with a sensible default: a notice naming a never-configured platform stops re-logging forever after an hour. No findings.

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/gateway Gateway runner, session dispatch, delivery area/install-update Installer, updater, packaging, wheels, doctor sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Sep 17, 2026
@alt-glitch

Copy link
Copy Markdown

This was generated by AI during triage.

Related: #87091 also stops the permanent retry loop for unreachable update-notification targets. #87091 checks whether the platform is disabled; this PR uses an age cap, so these are competing approaches rather than duplicates.

An /update run leaves a marker naming the chat to notify. When that platform's
adapter is not connected as the update finishes, _send_update_notification
defers and keeps the markers for a later retry — correct for the restart that
/update itself triggers, where the adapter reconnects seconds later.

Nothing bounds that wait. If the platform is not configured at all, no adapter
will ever appear and the defer never resolves. The startup path reschedules the
watcher whenever the markers are still on disk, so the marker outlives every
restart: it re-logs "adapter not connected yet" each poll_interval (2s), in
every process, indefinitely. A marker written months ago was still doing this
on one install, filling gateway.log with ~19k duplicate lines.

Bound the wait using the timestamp the marker already carries. Once the marker
is older than the cap the notice is undeliverable by any retry, so log it once
at WARNING, clear the markers, and return True — the definitive answer that
stops the caller rescheduling. Markers with no timestamp (written before that
field existed) keep the previous retry behavior rather than guess an age.

Tests cover the three behaviors: a stale marker is dropped and clears every
marker file, a recent marker is still held for the reconnecting adapter, and a
marker without a timestamp still retries.
@teknium1

Copy link
Copy Markdown
Collaborator

Landed via #118312 (rebase-merge, your authorship kept on the commit) — merged at 09f847d. A post-update notice whose platform never connects now stops retrying after 1 h and removes its markers with a single WARNING. Thanks @gkd2323c.

@teknium1 teknium1 closed this Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/install-update Installer, updater, packaging, wheels, doctor comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants