fix(discord): stop typing indicator getting stuck on rate-limited channels - #34
Merged
Merged
Conversation
…nnels The Discord typing-indicator loop (_typing_loop) could leave a permanent "<bot> is typing…" bubble after a response was already sent, on channels that are sustained-429 rate-limited (e.g. multiple bots sharing one channel). Two defects: 1. On a 429 the loop slept retry_after then continued with NO cap, so a persistently rate-limited channel spun forever re-POSTing ~every 1s (observed in gateway.log as endless 'Typing indicator rate-limited … retrying in 1.0s'). stop_typing() relied solely on task.cancel(), which races a cancel landing during that back-off sleep. 2. No way to interrupt the back-off sleep promptly. Fix: - Add a per-channel asyncio.Event stop signal. stop_typing() sets it FIRST (then cancels), and the loop checks it before every POST and sleeps via an interruptible wait — so a stop during a 429 back-off ends the loop at once. - Cap consecutive failed POSTs (_DISCORD_TYPING_MAX_CONSECUTIVE_FAILURES=5); a healthy POST resets the counter. A wedged channel gives up instead of spinning forever. Adds tests/gateway/test_discord_typing_stop.py — the two regression tests fail without this fix (stop_typing times out mid-backoff; loop never terminates) and pass with it. No regression in test_keep_typing_timeout.py or test_discord_connect.py.
🔎 Lint report:
|
| Rule | Count |
|---|---|
invalid-assignment |
2 |
unresolved-import |
1 |
First entries
tests/gateway/test_discord_typing_stop.py:33: [invalid-assignment] invalid-assignment: Object of type `SimpleNamespace` is not assignable to attribute `http` on type `ModuleType | MagicMock`
tests/gateway/test_discord_typing_stop.py:34: [invalid-assignment] invalid-assignment: Object of type `Literal[True]` is not assignable to attribute `_typing_test_ready` on type `ModuleType | MagicMock`
tests/gateway/test_discord_typing_stop.py:21: [unresolved-import] unresolved-import: Cannot resolve imported module `pytest`
✅ Fixed issues: none
Unchanged: 5118 pre-existing issues carried over.
Diagnostics are surfaced as warnings — this check never fails the build.
Kyzcreig
added a commit
that referenced
this pull request
Jun 21, 2026
A Discord "is typing…" bubble could stay lit indefinitely after a turn was fully delivered, with NO rate-limiting involved (distinct from the 429 stuck- bubble fixed in #34). Field report 2026-06-20: a bubble stayed lit ~26 min on a channel with zero 429s and zero turns since the last response. Root cause: a check-then-act race across teardown. Under concurrent same- channel turns, a late send_typing() — from BasePlatformAdapter._keep_typing() refresh tick, or the gateway follow-up restart (gateway/run.py:16006) — can RECREATE the per-channel _typing_loop AFTER the owning turn's stop_typing() already popped+cancelled it. The duplicate-guard (chat_id in _typing_tasks) is empty at that moment, so a new orphaned loop arms and re-POSTs the typing indicator (~every 12s) forever with no live owner to cancel it. base.py's own _keep_typing finally already documented this recreate path. Fix: a monotonic per-chat typing OWNERSHIP token, issued by the base adapter at turn start (_issue_typing_token) and threaded through _keep_typing -> send_typing -> _typing_loop. A late/stale send_typing carries its own now-superseded token and refuses to arm; the loop re-checks ownership each iteration and self-terminates if superseded. stop_typing never mutates the owner dict (single incrementer = _issue_typing_token; single popper = _stop_typing_refresh on owner-match), so the owner-match teardown compare can't be raced. The Discord adapter co-locates task+stop_event+token in three dicts popped together and only clears shared state when the loop is still the registered owner (no clobber of a newer turn's loop). token=None preserves the legacy unconditional behavior for all existing call sites and other adapters, so this is opt-in and behavior-preserving everywhere else. The base-level primitive covers the whole bug class; the Discord adapter is the per-adapter consumer in this PR. Other adapters with the same _typing_tasks/stop_typing shape (signal/yuanbao/slack/bluebubbles/weixin) can opt in with a one-line loop check as a follow-up. tests/gateway/test_discord_typing_recreate_race.py: - test_orphaned_loop_outlives_stop drives the REAL unmodified adapter through the field interleaving and was captured RED on main (orphan survived + kept POSTing) before this fix; GREEN after. - 3 structural guards for the token primitive (late-arm refusal, no-false-stop of a superseding turn, bounded owner dict). All existing typing tests (test_discord_typing_stop.py #34 429 cases, test_keep_typing_timeout.py) stay green.
Kyzcreig
added a commit
that referenced
this pull request
Jun 21, 2026
…#69) * fix(discord): stop "is typing…" orphaned by typing-loop recreate race A Discord "is typing…" bubble could stay lit indefinitely after a turn was fully delivered, with NO rate-limiting involved (distinct from the 429 stuck- bubble fixed in #34). Field report 2026-06-20: a bubble stayed lit ~26 min on a channel with zero 429s and zero turns since the last response. Root cause: a check-then-act race across teardown. Under concurrent same- channel turns, a late send_typing() — from BasePlatformAdapter._keep_typing() refresh tick, or the gateway follow-up restart (gateway/run.py:16006) — can RECREATE the per-channel _typing_loop AFTER the owning turn's stop_typing() already popped+cancelled it. The duplicate-guard (chat_id in _typing_tasks) is empty at that moment, so a new orphaned loop arms and re-POSTs the typing indicator (~every 12s) forever with no live owner to cancel it. base.py's own _keep_typing finally already documented this recreate path. Fix: a monotonic per-chat typing OWNERSHIP token, issued by the base adapter at turn start (_issue_typing_token) and threaded through _keep_typing -> send_typing -> _typing_loop. A late/stale send_typing carries its own now-superseded token and refuses to arm; the loop re-checks ownership each iteration and self-terminates if superseded. stop_typing never mutates the owner dict (single incrementer = _issue_typing_token; single popper = _stop_typing_refresh on owner-match), so the owner-match teardown compare can't be raced. The Discord adapter co-locates task+stop_event+token in three dicts popped together and only clears shared state when the loop is still the registered owner (no clobber of a newer turn's loop). token=None preserves the legacy unconditional behavior for all existing call sites and other adapters, so this is opt-in and behavior-preserving everywhere else. The base-level primitive covers the whole bug class; the Discord adapter is the per-adapter consumer in this PR. Other adapters with the same _typing_tasks/stop_typing shape (signal/yuanbao/slack/bluebubbles/weixin) can opt in with a one-line loop check as a follow-up. tests/gateway/test_discord_typing_recreate_race.py: - test_orphaned_loop_outlives_stop drives the REAL unmodified adapter through the field interleaving and was captured RED on main (orphan survived + kept POSTing) before this fix; GREEN after. - 3 structural guards for the token primitive (late-arm refusal, no-false-stop of a superseding turn, bounded owner dict). All existing typing tests (test_discord_typing_stop.py #34 429 cases, test_keep_typing_timeout.py) stay green. * fix(typing): guard _typing_owner access for partially-initialized adapters CI caught a real regression: test adapters that build via object.__new__() and skip __init__ (e.g. tests/gateway/test_active_session_text_merge.py's _DummyAdapter) do not have _typing_owner, so _process_message_background -> _issue_typing_token raised AttributeError. Harden the three token accessors and the _stop_typing_refresh pop site to create/read _typing_owner defensively via getattr, mirroring how other partially-initialized-adapter state is tolerated. No behavior change for real adapters (which init the dict). --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The Discord typing indicator (
DiscordAdapter._typing_loop) can leave a permanent "<bot> is typing…" bubble in a channel after the response was already sent, on channels that get sustained 429 rate-limits (commonly: multiple bots sharing one channel, so typing events collide).Observed in
gateway.logas an endless stream:Root cause
Two defects in the typing loop lifecycle:
retry_afterandcontinued unconditionally. A persistently rate-limited channel spun forever, re-POSTing roughly every second.stop_typingraces the back-off sleep. It relied solely ontask.cancel(). A cancel that lands while the loop is parked inawait asyncio.sleep(retry_after)can be swallowed/raced, so the loop keeps the bubble alive past the sent response.Fix
asyncio.Eventstop signal.stop_typing()sets it first (then cancels), and the loop checks it before every POST and sleeps via an interruptible wait — so a stop during a 429 back-off ends the loop immediately._DISCORD_TYPING_MAX_CONSECUTIVE_FAILURES = 5); a healthy POST resets the counter. A wedged channel gives up instead of spinning forever.Both
_typing_tasksand the new_typing_stop_eventsare cleaned up in the loop'sfinallyand instop_typing.Tests
Adds
tests/gateway/test_discord_typing_stop.py(3 tests):test_stop_typing_interrupts_429_backoff— stop during a 30s back-off must end the loop in <2s.test_typing_loop_gives_up_after_persistent_rate_limit— a permanently-429'd channel self-terminates at the cap.test_successful_post_resets_failure_counter— an intermittent 429→success must not trip the cap.Gold-standard verified: with the impl change stashed, the two regression tests fail (stop times out mid-backoff; loop never terminates — reproducing the exact
retrying in …log). With the fix they pass. No regression intest_keep_typing_timeout.pyortest_discord_connect.py(25 passed).