Skip to content

fix(telegram): stop flood control from rewriting a finalized reply as plain text - #100446

Open
AlexxRussell wants to merge 1 commit into
NousResearch:mainfrom
AlexxRussell:fix/telegram-edit-fallback-only-on-parse-errors
Open

AlexxRussell wants to merge 1 commit into
NousResearch:mainfrom
AlexxRussell:fix/telegram-edit-fallback-only-on-parse-errors

Conversation

@AlexxRussell

@AlexxRussell AlexxRussell commented Sep 1, 2026 •

Copy link
Copy Markdown

What does this PR do?

TelegramAdapter.edit_message sends a finalized reply with
parse_mode=MarkdownV2. When that call raises, an inner handler rewrites the
message as stripped plain text. That rescue is right for a MarkdownV2 parse
failure, where the markup is what Telegram rejected. It was applied to flood
control too, because the handler caught bare Exception.

A RetryAfter refusal says nothing about the markup, so a reply whose
MarkdownV2 was perfectly valid arrived with its syntax showing: headings as a
literal ##, links as a literal [text](url).

Observed on a live gateway, twice in three days:

2026-08-30 13:01:17 ERROR   [Telegram] Fallback send also failed: flood_control:45.0
2026-08-31 10:46:01 WARNING [Telegram] Telegram flood control, waiting 3.0s
2026-08-31 10:46:04 WARNING [Telegram] MarkdownV2 edit failed, falling back to plain text: Flood control exceeded. Retry in 269 seconds
2026-08-31 10:46:04 WARNING [Telegram] Telegram flood control, waiting 269.0s

The outer handler in the same method was already written for exactly this, and
its own comment states the intent: short flood waits are retried inline, and an
over-cap wait "return[s] a failure immediately so streaming can fall back to a
normal final send instead of leaving a truncated partial". It rarely ran,
because the inner handler swallowed the flood exception first (the plain-text
rescue could itself be refused by flood control and reach the outer handler,
but only after the downgrade had already been attempted).

The scope is deliberately limited to flood control, which is explained under
Changes Made below.

Related Issue

No existing issue. Reproduction and log evidence are above.

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)

Changes Made

plugins/platforms/telegram/adapter.py

  • Two shared helpers. _looks_like_flood_error recognises a refusal from
    PTB's retry_after or from Telegram's own Flood control exceeded. Retry in Ns, which the outer handler's "retry after" test never matched.
    _flood_wait_seconds normalises the delay, including reading it out of that
    message when no attribute is present, so a text-matched refusal still fails
    closed on a long wait instead of retrying blind inside the window.

  • The inner finalize fallback re-raises on flood control so it reaches the
    outer handler. The not-modified shortcut is unchanged. The bounded retry
    after a short flood wait keeps its own plain-text rescue for a parse
    rejection the refused first attempt never reached, and does not wait on a
    second flood refusal.

  • The same rule now applies inside _edit_overflow_split, for both the
    first-chunk edit and the continuation sends. Those carried their own
    catch-all downgrades, so any reply past the 4,096 unit cap kept the old
    behaviour regardless of the guard above.

  • The inline flood retry re-sends the MarkdownV2 render rather than the raw
    content. It previously re-edited with the unformatted source, which showed
    the same raw syntax whenever a short flood wait hit a finalized reply. When
    a genuine parse failure has already degraded the payload, the retry
    re-sends that plain text instead, rather than reinstating markup Telegram
    just rejected.

  • retry_after is normalised before it is compared against the inline wait
    cap. python-telegram-bot exposes it as a datetime.timedelta under
    PTB_TIMEDELTA=1, its supported migration mode for the 23.x default, and the
    comparison raised TypeError from inside the exception handler, so
    edit_message escaped without retrying or returning a SendResult.

Why only flood control. Re-raising a timeout would return
retryable=True, and GatewayStreamConsumer does not consult retryable on
an edit failure: it sets _fallback_final_send and sends the missing tail. A
timeout can mean Telegram applied the edit and only the response was lost (the
send path notes the same at its own retry), in which case the tail would be
sent again on top of the complete answer. Losing formatting is the smaller
cost, so timeouts keep the existing rescue and
test_transient_network_error_keeps_the_plain_text_rescue pins that as a
decision. Teaching the consumer to honour retryable on edits looks like the
real fix and is a larger change than this one; happy to open it separately if
that is wanted.

tests/gateway/test_telegram_flood_keeps_markdown.py (new)

29 cases (16 test functions, several parametrized) drive the real adapter
through a bot that replays a scripted failure sequence, and test the two
helpers directly:

  • an over-cap flood wait fails closed and issues no plain-text edit
  • an under-cap flood wait retries inline with the same MarkdownV2 payload
  • both retry_after shapes, float and timedelta
  • a genuine parse failure is still rescued as plain text (3 message shapes)
  • [parse error, RetryAfter(1), success] keeps the degraded payload on the
    retry instead of reinstating the rejected markup
  • a parse rejection that only surfaces on the retry after a flood wait still
    gets the plain-text rescue
  • a refusal carrying no retry_after at all, recognised from the message,
    with the encoded 269s surviving into the capped result
  • the delay that actually reaches asyncio.sleep
  • an overflowing reply: first-chunk edit and continuation sends under flood
    control fail closed instead of downgrading
  • a timeout keeps the plain-text rescue, per the reasoning above
  • not-modified is still a no-op success
  • the helpers' classification and wait extraction, case by case

How to Test

$ pytest tests/gateway/test_telegram_flood_keeps_markdown.py -q
29 passed

The test module imports _looks_like_flood_error and _flood_wait_seconds, the
two helpers this PR adds, so against an unmodified main it fails at collection
(the helpers do not exist there) rather than as individual assertion failures.
It is a fix-pinning suite: with the adapter change applied it is green, and each
behaviour it locks was mutation-checked. The captured payload that the
finalized-reply test asserts against is the bug verbatim (a finalized reply
rewritten as plain text under flood control):

'text': '## Best souvlaki choice\n\nGo to [Athinaiko Souvlaki](https://...) at ...'

Each mechanism was mutation-checked rather than assumed. Dropping the
"flood control" text clause, removing the message-parsed wait, reverting the
timedelta normalisation, deleting the line that carries the degraded payload
into the retry, removing the overflow guards and removing the retry's own
parse rescue each fail at least one test.

No regressions across the Telegram surface plus the streaming consumer:

$ pytest tests/gateway/test_telegram_*.py tests/gateway/test_stream_consumer.py -q

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix
  • I've run the affected suites and all tests pass
  • I've added tests for my changes
  • I've tested on my platform: macOS 15 (Darwin 25.6) and Ubuntu 22.04 (gateway host)

Documentation & Housekeeping

  • I've updated relevant documentation (docstrings and inline rationale) or N/A
  • I've updated cli-config.yaml.example if I added/changed config keys or N/A (no config keys)
  • I've updated CONTRIBUTING.md or AGENTS.md or N/A
  • I've considered cross-platform impact (no platform-specific code paths)
  • I've updated tool descriptions/schemas or N/A

@alt-glitch alt-glitch added type/bug Something isn't working comp/plugins Plugin system and bundled plugins platform/telegram Telegram bot adapter P3 Low — cosmetic, nice to have sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Sep 1, 2026
@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference; please use your judgment.

A carefully reasoned and well-tested fix. The bug is real: the finalize-edit fallback caught bare Exception, so a flood-control refusal (which says nothing about markup) triggered the "strip MarkdownV2 and resend as plain text" rescue, showing users literal ##/[text](url) syntax for perfectly valid formatting. The fix re-raises flood errors from the inner handler so the already-correct outer flood logic runs (adapter.py:5923-5925), and it hardens detection with _looks_like_flood_error (covering both PTB's retry_after and Telegram's own "Flood control ... Retry in Ns" wording) plus _flood_wait_seconds (normalizing the timedelta shape under PTB_TIMEDELTA=1, adapter.py:262-292). Tracking _retry_text/_retry_parse_mode through each branch ensures the inline flood retry resends the right payload (raw for previews, MarkdownV2 for finalize, plain after a parse fallback) — all pinned by tests, including the markdown→plain→flood sequence. The deliberate narrowness (timeouts keep the plain-text rescue because the consumer ignores retryable) is documented and test-pinned, which is exactly the right call.

Points worth attention:

  • Overflow edge: an under-cap flood raised from _edit_overflow_split reaches the outer inline retry, which re-sends _retry_text — the full content (adapter.py:6011-6016). For a reply that overflowed the 4096 cap in the first place, that single edit will be rejected as too long, so the inline retry is wasted (it fails closed rather than looping, so it's safe, just suboptimal). Consider re-entering the split path on the flood retry, or noting that overflow + flood intentionally fails closed.
  • _flood_wait_seconds returning default (1.0) when the message text carries no parseable delay is a reasonable floor, but combined with _looks_like_flood_error matching "flood control" text with no numeric wait, a vague refusal would retry blind after 1s. The comment frames this as the intended fallback; just flagging that the regex covers the standard Telegram wording, and anything else falls back to 1.0.

Verdict: LGTM

@AlexxRussell
AlexxRussell force-pushed the fix/telegram-edit-fallback-only-on-parse-errors branch from cca8195 to e86cb5f Compare September 5, 2026 12:22
@AlexxRussell

Copy link
Copy Markdown
Author

Rebased onto current main (2e24e06) and re-implemented against the refactored adapter: the finalize edit fallback now lives in the shared _edit_markdown_or_plain helper, so the guard moved there along with a small retry_state hand-off so the inline flood retry re-sends the payload that was actually in flight.

Same fix and the same deliberate narrowness as before. New in this revision:

  • the inline flood retry keeps the helper's plain-text rescue and the not-modified no-op when the MarkdownV2 retry itself is rejected (the refused first attempt never reached that rejection, so the retry can), and a second flood refusal fails closed;
  • the overflow paths no longer downgrade a flood-refused chunk: the first-chunk edit fails closed with the shared flood result and a continuation reports the failure instead of resending unformatted.

Tests moved to tests/gateway/test_telegram_flood_keeps_markdown.py (29 tests; each guard was mutation-tested). The Telegram and stream consumer suites pass locally on main.

@AlexxRussell
AlexxRussell force-pushed the fix/telegram-edit-fallback-only-on-parse-errors branch from e86cb5f to 0eeac51 Compare September 5, 2026 13:04
@AlexxRussell

Copy link
Copy Markdown
Author

Cosmetic follow-up (22dc0dd): wrapped the new comments and calls to 100 columns. Body corrected: the test file is tests/gateway/test_telegram_flood_keeps_markdown.py with 29 cases.

@AlexxRussell

Copy link
Copy Markdown
Author

Docs follow-up (6b1202e): fixed a self-contradiction in the test docstring ("could never run" then "even when it did run") and the fix-pinning note in the body (the test module imports helpers not on main, so against unmodified main it fails at collection).

@AlexxRussell

Copy link
Copy Markdown
Author

Rebased onto current main (fc8d15d77). No conflicts, and the change content is unchanged.

Worth saying why it needed one. #104370 reworked _send_with_retry in this same send path: per-attempt rate-limit classification, a 60s inline retry_after cap, and a typed failure returned inside a flood penalty. I re-checked the guard against that base rather than treating the clean merge as proof it still held. The finalize edit fallback still degrades to plain text only on parse errors, and the retry_state hand-off still re-sends the payload that was actually in flight. A flood-capped send now returns the typed failure before the fallback is reachable, which is the behaviour this PR was arguing for.

tests/gateway/test_telegram_*.py is 688 passed, 2 skipped.

@AlexxRussell
AlexxRussell force-pushed the fix/telegram-edit-fallback-only-on-parse-errors branch from 0a7504f to 90b839f Compare September 19, 2026 23:19
@AlexxRussell

Copy link
Copy Markdown
Author

Rebased onto current main (8a92051f20), replayed as one commit
(90b839f51e).

One conflict, in edit_message, and the resolution drops a line this branch
carried, so it is worth stating rather than leaving in the diff.

The branch had the flood log before the inline-cap check:

if _looks_like_flood_error(e):
    wait = _flood_wait_seconds(e)
    logger.warning("[%s] Telegram flood control, waiting %.1fs", self.name, wait)
    if wait > _FLOOD_INLINE_WAIT_CAP_SECS:
        return _flood_cap_result(wait)

That ordering came from this branch's base, not from this change. Main has since
moved the log to after the cap check, with a comment giving the reason
("waiting 33.0s" followed by no wait misled an investigation). Carrying our copy
through the rebase would have reinstated the misleading log and printed it
twice, inside a PR about something else. Resolved in favour of main, so only the
detection is this PR's:

if _looks_like_flood_error(e):
    wait = _flood_wait_seconds(e)

The contribution is unchanged. A flood refusal is re-raised out of
_edit_markdown_or_plain instead of being answered with a plain-text downgrade,
the inline retry replays _retry_state rather than re-sending stripped text, and
_edit_overflow_split follows the same rule on the first chunk and the
continuations.

tests/gateway/test_telegram_flood_keeps_markdown.py is 29 passed. The Telegram
sweep (tests/gateway -k telegram) is 900 passed with 5 skipped, against 871
passed with 5 skipped on clean main at the same commit, which is the 29 new tests
and nothing else moved.

One weakness I would rather flag than leave for you to find: the module imports
_looks_like_flood_error and _flood_wait_seconds at module level, so against
unmodified main it fails collection rather than failing behaviourally, which
proves the symbols exist rather than proving anything about behaviour. Measured
properly by adding only the two helpers to main and none of the behaviour
change: 11 of the 29 still fail, so that is the load-bearing number. Happy to
restructure the two helper unit tests behind a module attribute lookup if you
would rather the whole file stayed behaviour-diffable.

… plain text

The finalize edit helper _edit_markdown_or_plain caught bare Exception and
answered every failure with a _strip_mdv2 plain-text rewrite. That rescue is
right for a MarkdownV2 parse error, where the markup is what Telegram
rejected. It also fired on flood control, which says nothing about the
markup, so a reply whose MarkdownV2 was perfectly valid arrived with its
syntax showing: headings as a literal ## and links as a literal [text](url).
The flood branch in edit_message could never run while the helper swallowed
the exception first, and when it did run its inline retry re-sent the raw
content, which showed the same syntax.

Fix, ported to the refactored adapter:

- Shared helpers _looks_like_flood_error (retry_after attribute, or
  Telegram's own "Flood control exceeded. Retry in Ns" text, which the old
  "retry after" test did not match) and _flood_wait_seconds (float, or
  timedelta under PTB_TIMEDELTA=1, or the delay parsed from the message).
  Comparing a timedelta against the inline cap raised TypeError from inside
  the handler.
- _edit_markdown_or_plain re-raises flood errors and, through an optional
  retry_state dict, records the payload in flight (the MarkdownV2 render, or
  the stripped text after a genuine parse fallback).
- edit_message's flood branch uses the helpers and retries with that
  recorded payload instead of raw content. Because the retry now carries
  MarkdownV2 it can meet a parse rejection the refused first attempt never
  reached, so it keeps the same plain-text rescue and not-modified no-op the
  helper gives a first attempt; a second flood refusal fails closed.
- _edit_overflow_split fails closed with the shared flood result on a
  flood-refused first chunk; _send_overflow_continuation reports the
  failure instead of resending the chunk unformatted.

Deliberately narrow: parse errors, timeouts and "message to edit not found"
keep the plain-text rescue. GatewayStreamConsumer does not consult
retryable on an edit failure, so re-raising a timeout would send the tail
again on top of an edit Telegram may already have applied.

Tests: 29 in tests/gateway/test_telegram_flood_keeps_markdown.py, covering
the over-cap and inline paths, timedelta and text-only refusals, the parse
fallback that must survive (before and after a flood wait), the overflow
first chunk and continuation, and the helper classification. Each guard was
mutation-tested.
@AlexxRussell
AlexxRussell force-pushed the fix/telegram-edit-fallback-only-on-parse-errors branch from 90b839f to a7db31c Compare September 19, 2026 23:32
@AlexxRussell

Copy link
Copy Markdown
Author

One more fix on this branch (a7db31c44b), found by checking what the rebase
INHERITED rather than only what it changed.

Upstream added a second-refusal block to edit_message after this branch's
base: when the retry that follows the inline wait is refused again, it fails
closed with _flood_cap_result(float(retry_wait) if retry_wait is not None else wait). This PR exists because that method meets retry_after in more than one
shape, and that inherited block does its own raw conversion, so two refusals in
a row behave worse than one:

  • retry_after as a datetime.timedelta (PTB's PTB_TIMEDELTA=1 migration
    mode for the 23.x default) raises TypeError out of float(), from inside
    the exception handler, so edit_message raises instead of returning a
    SendResult at all;
  • a refusal that states its delay only in Telegram's wording ("Flood control
    exceeded. Retry in Ns") is not matched by the "retry after" test, so it
    returns a plain failure that the ledger reads as ordinary and never
    redelivers.

Those are the exact two gaps _looks_like_flood_error and _flood_wait_seconds
were written to close, so the block now uses them and the local retry_wait
lookup is gone.

Neither is a regression from this branch: on its old base that path ended at a
plain failure, and the TypeError is reachable on main today under
PTB_TIMEDELTA=1. It is in scope here because this is the PR that gives the
method one way to read a refusal.

New test test_the_second_flood_refusal_is_normalised_whatever_shape_it_has
covers all three shapes. Reverting only the adapter fails its timedelta and
text-only cases, so it is a behaviour test rather than a symbol test.
tests/gateway/test_telegram_flood_keeps_markdown.py is now 32 passed.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/plugins Plugin system and bundled plugins P3 Low — cosmetic, nice to have platform/telegram Telegram bot adapter sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants