Repository navigation
fix(gateway): stop retrying /update notification for platforms disabled in this profile - #87091
Open
RichardHojunJang wants to merge 1 commit into
Conversation
A stale `.update_pending.json` naming a platform this profile does not run pins a permanent retry loop for the life of the gateway. `_send_update_notification` treats "no adapter for the target platform" as a single case and always defers, preserving the markers so a reconnecting adapter can still be notified (the behaviour NousResearch#39091 added, which is correct for a platform that IS enabled and merely still connecting). But `config.platforms` is pre-seeded with disabled placeholders for the whole platform catalog, so a marker can name a platform with `enabled=False`. No adapter will ever appear for it. The watcher then polls every 2s, and each poll re-reads the marker, logs, and rewrites it -- forever. Observed on a Slack-only profile carrying an orphan marker that named telegram: a steady 0.5 lines/sec, ~43k lines/day, which grew to 78% of the gateway log (9,972 of 12,699 lines) before it was found. Split the two cases. When the target platform is provably not enabled for this profile, the target is unreachable by construction: consume the markers and log once at WARNING with the reason. When it is enabled, keep deferring exactly as before. The enabled-check fails open. A missing or unreadable config returns True, so an unexpected config shape keeps the existing retry behaviour and can never discard a deliverable notification; only a positive, readable `enabled=False` ends the retry. Tests: - disabled platform: markers consumed, no retry, nothing misdelivered - enabled but reconnecting: markers preserved (guards the discard path from over-reaching) - unreadable config: fails open to the old deferral Verified the first test fails on the unpatched tree (`assert False is True`) and that the other 17 tests in the class still pass there, so it pins this bug rather than the surrounding behaviour.
RichardHojunJang
force-pushed
the
snow/pr-update-notify-unconfigured-platform-20260815
branch
from
August 15, 2026 15:43
60fd7e9 to
898fca1
Compare
fix(gateway): stop retrying /update notification for platforms disabled in this profile
|
Collaborator
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A stale
.update_pending.jsonnaming a platform the profile does not run pins a permanent retry loop for the life of the gateway._send_update_notificationtreats "no adapter for the target platform" as one case and always defers, preserving the markers so a reconnecting adapter can still be notified. That is correct — and deliberate — for a platform that is enabled and merely still connecting (the behaviour #39091 added).But
config.platformsis pre-seeded with disabled placeholders for the whole platform catalog, so a marker can name a platform withenabled=False. No adapter will ever appear for it. The watcher polls every 2s, and each poll re-reads the marker, logs, and rewrites it — forever.This PR splits the two cases.
Reproduction
Found in production on a Slack-only profile carrying an orphan marker that named telegram:
The marker was left by an interrupted update:
.update_exit_codewas124(timeout), so the update was long finished — only the unreachable target remained.Runtime probe of the affected profile confirms the shape:
telegramis present inconfig.platformsbut disabled — soself.adapters.get(platform)returnsNoneon every single poll, permanently.The fix
Where
mainhas one branch, there are two distinct situations:WARNINGwith the reason.The new
_update_target_platform_is_enabledhelper fails open: a missing or unreadable config returnsTrue, so an unexpected config shape keeps the existing retry behaviour and can never discard a deliverable notification. Only a positive, readableenabled=Falseends the retry.Mere key presence is explicitly not trusted — the
enabledflag is the load-bearing part. This is the same pre-seeded-placeholder trap already documented on_scale_to_zero_active_messaging_platforms("the F25 bug"), and the helper's docstring points at it so the next reader does not re-learn it.Test plan
Three new tests:
test_disabled_platform_marker_is_discarded_not_retriedtest_enabled_platform_still_defers_while_reconnectingtest_unreadable_platform_config_fails_open_to_deferralHeld-in verification — the regression was confirmed to actually pin this bug, not the surrounding behaviour. Reverting only
gateway/run.pywhile keeping the new tests:The new test fails on the unpatched tree with the exact expected assertion, and the other 17 tests in the class still pass — so it is not an over-broad change detector.
ruff checkpasses on both files. (ruff format --checkreports these files as unformatted both before and after this change, so it is pre-existing and untouched here.)Relationship to other open PRs
Checked before writing this; all three touch the same markers but none cover this case:
SendResult(success=False)soft-failurePlatform(platform_str)raising. A valid platform that is merely disabled passes that check and still defers (its own line 359-361)This PR is additive to all three: they make delivery more reliable when a target is reachable; this one ends the retry when the target provably is not.
Sensitive data audit
gateway/run.pyandtests/gateway/test_update_command.pychanged.111/222), matching the existing convention in this file.