Skip to content

fix(discord): stop typing indicator getting stuck on rate-limited channels - #34

Merged
Kyzcreig merged 1 commit into
mainfrom
fix/discord-typing-stuck-bubble
Jun 12, 2026
Merged

Kyzcreig merged 1 commit into
mainfrom
fix/discord-typing-stuck-bubble

Conversation

@Kyzcreig

Copy link
Copy Markdown
Collaborator

Problem

The Discord typing indicator (DiscordAdapter._typing_loop) can leave a permanent "<bot> is typing…" bubble in a channel after the response was already sent, on channels that get sustained 429 rate-limits (commonly: multiple bots sharing one channel, so typing events collide).

Observed in gateway.log as an endless stream:

WARNING ...discord...adapter: Typing indicator rate-limited for <channel>; retrying in 1.0s
WARNING ...discord...adapter: Typing indicator rate-limited for <channel>; retrying in 1.0s
... (forever)

Root cause

Two defects in the typing loop lifecycle:

  1. No failure cap. On a 429 the loop slept retry_after and continued unconditionally. A persistently rate-limited channel spun forever, re-POSTing roughly every second.
  2. stop_typing races the back-off sleep. It relied solely on task.cancel(). A cancel that lands while the loop is parked in await asyncio.sleep(retry_after) can be swallowed/raced, so the loop keeps the bubble alive past the sent response.

Fix

  • Add a per-channel asyncio.Event stop signal. stop_typing() sets it first (then cancels), and the loop checks it before every POST and sleeps via an interruptible wait — so a stop during a 429 back-off ends the loop immediately.
  • Cap consecutive failed POSTs (_DISCORD_TYPING_MAX_CONSECUTIVE_FAILURES = 5); a healthy POST resets the counter. A wedged channel gives up instead of spinning forever.

Both _typing_tasks and the new _typing_stop_events are cleaned up in the loop's finally and in stop_typing.

Tests

Adds tests/gateway/test_discord_typing_stop.py (3 tests):

  • test_stop_typing_interrupts_429_backoff — stop during a 30s back-off must end the loop in <2s.
  • test_typing_loop_gives_up_after_persistent_rate_limit — a permanently-429'd channel self-terminates at the cap.
  • test_successful_post_resets_failure_counter — an intermittent 429→success must not trip the cap.

Gold-standard verified: with the impl change stashed, the two regression tests fail (stop times out mid-backoff; loop never terminates — reproducing the exact retrying in … log). With the fix they pass. No regression in test_keep_typing_timeout.py or test_discord_connect.py (25 passed).

…nnels

The Discord typing-indicator loop (_typing_loop) could leave a permanent
"<bot> is typing…" bubble after a response was already sent, on channels
that are sustained-429 rate-limited (e.g. multiple bots sharing one channel).

Two defects:
1. On a 429 the loop slept retry_after then continued with NO cap, so a
   persistently rate-limited channel spun forever re-POSTing ~every 1s
   (observed in gateway.log as endless 'Typing indicator rate-limited …
   retrying in 1.0s'). stop_typing() relied solely on task.cancel(), which
   races a cancel landing during that back-off sleep.
2. No way to interrupt the back-off sleep promptly.

Fix:
- Add a per-channel asyncio.Event stop signal. stop_typing() sets it FIRST
  (then cancels), and the loop checks it before every POST and sleeps via an
  interruptible wait — so a stop during a 429 back-off ends the loop at once.
- Cap consecutive failed POSTs (_DISCORD_TYPING_MAX_CONSECUTIVE_FAILURES=5);
  a healthy POST resets the counter. A wedged channel gives up instead of
  spinning forever.

Adds tests/gateway/test_discord_typing_stop.py — the two regression tests
fail without this fix (stop_typing times out mid-backoff; loop never
terminates) and pass with it. No regression in test_keep_typing_timeout.py
or test_discord_connect.py.
@github-actions

Copy link
Copy Markdown

🔎 Lint report: fix/discord-typing-stuck-bubble vs origin/main

ruff

Total: 0 on HEAD, 0 on base (➖ 0)

🆕 New issues: none

✅ Fixed issues: none

Unchanged: 0 pre-existing issues carried over.

ty (type checker)

Total: 9906 on HEAD, 9903 on base (🆕 +3)

🆕 New issues (3):

Rule Count
invalid-assignment 2
unresolved-import 1
First entries
tests/gateway/test_discord_typing_stop.py:33: [invalid-assignment] invalid-assignment: Object of type `SimpleNamespace` is not assignable to attribute `http` on type `ModuleType | MagicMock`
tests/gateway/test_discord_typing_stop.py:34: [invalid-assignment] invalid-assignment: Object of type `Literal[True]` is not assignable to attribute `_typing_test_ready` on type `ModuleType | MagicMock`
tests/gateway/test_discord_typing_stop.py:21: [unresolved-import] unresolved-import: Cannot resolve imported module `pytest`

✅ Fixed issues: none

Unchanged: 5118 pre-existing issues carried over.

Diagnostics are surfaced as warnings — this check never fails the build.

@Kyzcreig
Kyzcreig merged commit d401577 into main Jun 12, 2026
23 checks passed
@Kyzcreig
Kyzcreig deleted the fix/discord-typing-stuck-bubble branch June 12, 2026 02:37
Kyzcreig added a commit that referenced this pull request Jun 21, 2026
A Discord "is typing…" bubble could stay lit indefinitely after a turn was
fully delivered, with NO rate-limiting involved (distinct from the 429 stuck-
bubble fixed in #34). Field report 2026-06-20: a bubble stayed lit ~26 min on
a channel with zero 429s and zero turns since the last response.

Root cause: a check-then-act race across teardown. Under concurrent same-
channel turns, a late send_typing() — from BasePlatformAdapter._keep_typing()
refresh tick, or the gateway follow-up restart (gateway/run.py:16006) — can
RECREATE the per-channel _typing_loop AFTER the owning turn's stop_typing()
already popped+cancelled it. The duplicate-guard (chat_id in _typing_tasks)
is empty at that moment, so a new orphaned loop arms and re-POSTs the typing
indicator (~every 12s) forever with no live owner to cancel it. base.py's own
_keep_typing finally already documented this recreate path.

Fix: a monotonic per-chat typing OWNERSHIP token, issued by the base adapter
at turn start (_issue_typing_token) and threaded through
_keep_typing -> send_typing -> _typing_loop. A late/stale send_typing carries
its own now-superseded token and refuses to arm; the loop re-checks ownership
each iteration and self-terminates if superseded. stop_typing never mutates
the owner dict (single incrementer = _issue_typing_token; single popper =
_stop_typing_refresh on owner-match), so the owner-match teardown compare
can't be raced. The Discord adapter co-locates task+stop_event+token in three
dicts popped together and only clears shared state when the loop is still the
registered owner (no clobber of a newer turn's loop). token=None preserves
the legacy unconditional behavior for all existing call sites and other
adapters, so this is opt-in and behavior-preserving everywhere else.

The base-level primitive covers the whole bug class; the Discord adapter is
the per-adapter consumer in this PR. Other adapters with the same
_typing_tasks/stop_typing shape (signal/yuanbao/slack/bluebubbles/weixin) can
opt in with a one-line loop check as a follow-up.

tests/gateway/test_discord_typing_recreate_race.py:
- test_orphaned_loop_outlives_stop drives the REAL unmodified adapter through
  the field interleaving and was captured RED on main (orphan survived +
  kept POSTing) before this fix; GREEN after.
- 3 structural guards for the token primitive (late-arm refusal, no-false-stop
  of a superseding turn, bounded owner dict).
All existing typing tests (test_discord_typing_stop.py #34 429 cases,
test_keep_typing_timeout.py) stay green.
Kyzcreig added a commit that referenced this pull request Jun 21, 2026
…#69)

* fix(discord): stop "is typing…" orphaned by typing-loop recreate race

A Discord "is typing…" bubble could stay lit indefinitely after a turn was
fully delivered, with NO rate-limiting involved (distinct from the 429 stuck-
bubble fixed in #34). Field report 2026-06-20: a bubble stayed lit ~26 min on
a channel with zero 429s and zero turns since the last response.

Root cause: a check-then-act race across teardown. Under concurrent same-
channel turns, a late send_typing() — from BasePlatformAdapter._keep_typing()
refresh tick, or the gateway follow-up restart (gateway/run.py:16006) — can
RECREATE the per-channel _typing_loop AFTER the owning turn's stop_typing()
already popped+cancelled it. The duplicate-guard (chat_id in _typing_tasks)
is empty at that moment, so a new orphaned loop arms and re-POSTs the typing
indicator (~every 12s) forever with no live owner to cancel it. base.py's own
_keep_typing finally already documented this recreate path.

Fix: a monotonic per-chat typing OWNERSHIP token, issued by the base adapter
at turn start (_issue_typing_token) and threaded through
_keep_typing -> send_typing -> _typing_loop. A late/stale send_typing carries
its own now-superseded token and refuses to arm; the loop re-checks ownership
each iteration and self-terminates if superseded. stop_typing never mutates
the owner dict (single incrementer = _issue_typing_token; single popper =
_stop_typing_refresh on owner-match), so the owner-match teardown compare
can't be raced. The Discord adapter co-locates task+stop_event+token in three
dicts popped together and only clears shared state when the loop is still the
registered owner (no clobber of a newer turn's loop). token=None preserves
the legacy unconditional behavior for all existing call sites and other
adapters, so this is opt-in and behavior-preserving everywhere else.

The base-level primitive covers the whole bug class; the Discord adapter is
the per-adapter consumer in this PR. Other adapters with the same
_typing_tasks/stop_typing shape (signal/yuanbao/slack/bluebubbles/weixin) can
opt in with a one-line loop check as a follow-up.

tests/gateway/test_discord_typing_recreate_race.py:
- test_orphaned_loop_outlives_stop drives the REAL unmodified adapter through
  the field interleaving and was captured RED on main (orphan survived +
  kept POSTing) before this fix; GREEN after.
- 3 structural guards for the token primitive (late-arm refusal, no-false-stop
  of a superseding turn, bounded owner dict).
All existing typing tests (test_discord_typing_stop.py #34 429 cases,
test_keep_typing_timeout.py) stay green.

* fix(typing): guard _typing_owner access for partially-initialized adapters

CI caught a real regression: test adapters that build via object.__new__()
and skip __init__ (e.g. tests/gateway/test_active_session_text_merge.py's
_DummyAdapter) do not have _typing_owner, so _process_message_background ->
_issue_typing_token raised AttributeError. Harden the three token accessors
and the _stop_typing_refresh pop site to create/read _typing_owner defensively
via getattr, mirroring how other partially-initialized-adapter state is
tolerated. No behavior change for real adapters (which init the dict).

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
@Kyzcreig
Kyzcreig restored the fix/discord-typing-stuck-bubble branch September 21, 2026 10:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant