Skip to content

feat(telegram): single-message streaming + Bot API 10.3 extras (stop button, effects, ephemerals), all opt-in (#110564) - #110611

Open
Maziyang2 wants to merge 5 commits into
NousResearch:mainfrom
Maziyang2:feat/telegram-single-message-streaming
Open

Maziyang2 wants to merge 5 commits into
NousResearch:mainfrom
Maziyang2:feat/telegram-single-message-streaming

Conversation

@Maziyang2

@Maziyang2 Maziyang2 commented Sep 14, 2026 •

Copy link
Copy Markdown

Why this matters

One smooth message instead of a fragmented stream — plus four Bot API 10.3 capabilities, each behind its own opt-in switch.

Long Telegram answers that cross tool calls currently arrive as a chain of separate messages (a new send at every boundary). This PR keeps the entire turn in one editable message with a compact live activity overlay, and adds the modern Bot API 10.3 extras — everything opt-in, defaults off:

  • Single-message streaming (streaming_single_message) — one message per turn; thinking/tool activity as an expandable overlay; per-switch controls for activity, thinking, and 4096-split behavior; composes with quiet tool-progress modes (off/log).
  • Stop button (stop_button) — Bot API 10.3 dismissible stop control on draft previews; a press routes to /stop with the standard semantics.
  • Message effects (message_effects) — a completion effect (🎉 …) on the turn-final reply (private chats), with a minimum-turn-age gate.
  • Ephemeral messages (ephemeral_messages) — busy-ack whispered to the sender in groups; never falls back to a group-wide send.
  • Checklist rendering — native sendChecklist where the account type supports it (the business-account limitation is documented).

Companion PR — a two-part messaging upgrade

This is the first half of a coherent messaging/Telegram upgrade; both PRs are designed to land and be used together:

PR What it adds
1 this PR single-message streaming + Bot API 10.3 extras (stop button, effects, ephemerals) — each with an opt-in switch
2 #111631 cross-surface continuity — turns served from the Desktop app / TUI are mirrored to the platform chat, so the record stays complete

They are split into two PRs only for compatibility / independent reviewability — neither depends on the other's code. The mirror PR reuses this PR's long-text parameter semantics (streaming_single_message_4096_split) by design, so the two behave consistently when enabled together.

What does this PR do?

Adds an opt-in Telegram streaming mode that keeps one editable preview for the entire turn — text emitted on both sides of tool calls keeps updating the same message instead of sealing a new one at every tool boundary.

Verified behavior (local suites + a mock-adapter scenario bench):

  • base: interim narration + 2 tool boundaries + final → 3 separate sends (the fragmentation)
  • opted in: 1 send + in-place edits — final edit = the authoritative final text (interim replaced at adoption, no duplicated prefix)

Relation to #110571

Builds on the approach prototyped in #110571 by @KoNit-K — same config key (display.platforms.telegram.streaming_single_message) and the same one-line boundary skip — so maintainers can prefer either PR or fold them together without config churn. Extensions in this PR:

  1. Quiet-progress gate accepts off and log — tool_progress: log emits no chat bubbles either (the gateway's own tool_progress_enabled treats {"off", "log"} as quiet), so both quiet modes can host the single message.
  2. Commentary continuity — mid-turn commentary still posts its own message, but no longer resets the evolving preview (the resets are suppressed in the same way cumulative transports already suppress them).
  3. Docs — website/docs/user-guide/messaging/index.md + cli-config.yaml.example.
  4. Tests — boundary skip across two tool boundaries, commentary continuity, quiet-mode gating (off/log/new + non-Telegram).

Related Issue

Implements #110564 (companion to #110571).

Type of Change

  • ✨ New feature (non-breaking change that adds functionality)

Changes Made

  • gateway/display_config.py — option streaming_single_message (default false) + normalizer.
  • gateway/run_turn.py — resolves the opt-in (Telegram-only, quiet tool-progress off/log) into the consumer config.
  • gateway/stream_consumer.py — single_message_per_turn: skip the segment reset at tool boundaries and after commentary; docstrings updated.
  • tests/gateway/test_stream_consumer.py — 3 new tests.
  • website/docs/user-guide/messaging/index.md, cli-config.yaml.example — docs.

How to Test

  1. scripts/run_tests.sh tests/gateway/test_stream_consumer.py → base 3 failed, 71 passed → with the change 74 passed, 1 skipped, 0 failed.
  2. Adjacent suites — 10 files, 98 passed, 0 failed: display_config 23 · streaming_tts_consumer 17 · tool_progress 15 · draft 12 · fresh_final 10 · nested_config 9 · abandon-on-death 4 · thread_routing 4 · silence 3 · per_platform_defaults 1.
  3. Scenario bench (mock adapter; interim + 2 boundaries + a final distinct from the accumulated prefix): off → 3 sends; on → 1 send + edits.

Full pytest tests/ -q not run locally (suite too large); scoped suites above + CI for the rest. Tested on macOS (Darwin 27), Python 3.11.

Checklist

Code

  • I've read the Contributing Guide
  • Commit messages follow Conventional Commits
  • Searched for existing PRs — not a duplicate; companion to feat(gateway): add Telegram single-message streaming #110571 (same config key; maintainers may take either or consolidate)
  • Only changes related to this feature
  • pytest tests/ -q — ran the scoped suites instead (see How to Test); CI covers the full matrix
  • Tests added for the changes
  • Tested on my platform: macOS (Darwin 27)

Documentation & Housekeeping

  • Docs updated (messaging guide + cli-config.yaml.example)
  • CONTRIBUTING.md / AGENTS.md — N/A
  • Cross-platform impact considered — additive, Telegram-gated, default off; non-Telegram behavior unchanged
  • Tool descriptions/schemas — N/A

@alt-glitch alt-glitch added type/feature New feature or request P3 Low — cosmetic, nice to have comp/gateway Gateway runner, session dispatch, delivery platform/telegram Telegram bot adapter area/streaming Streaming responses: gateway delivery, provider wire sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Sep 14, 2026
@alt-glitch

Copy link
Copy Markdown

This was generated by AI during triage.

Related: #110571 (earlier minimal version by @KoNit-K, same streaming_single_message config key). This PR extends it with quiet-mode (off/log) gating, commentary continuity, docs and tests. Flagging the pair so a maintainer can pick one or fold them.

@Maziyang2

Copy link
Copy Markdown
Author

Hardening pass update — docs + live-API verification & activity overlay

Since the first revision we exercised this feature directly against the real Telegram Bot API (current docs, 10.3) and the live bot, and folded the findings back in. Pushed commits:

Verified against Bot API docs + live probes (macOS host, system proxy):

  • editMessageText on messages sent by the bot has no time window (the documented 48h limit is only for business messages not sent by the bot); deleteMessage requires <48h — all cleanup paths here run within seconds.
  • 4096-char cap confirmed live: 4096 ✓ / 4097 → MESSAGE_TOO_LONG; and the cap counts code points (2049 emoji, i.e. 4098 UTF-16 units, was accepted) — the consumer's utf16_len accounting is therefore conservative (can only split early, never fail).
  • Rapid edits on one message: 6 consecutive edits @ ~0.6s, all 200s; no flood penalties.
  • sendMessageDraft is an ephemeral 30-second preview (private chats only) and must be followed by a real sendMessage to persist — so the single-message mode composes correctly on both transports: drafts where available (preview animates, then one persistent message), edit path elsewhere (one evolving message).

Hardening fixes in this revision:

  • The gate now asks for the effective tool-progress mode via a shared _effective_tool_progress_mode() helper (same one the display path uses, env bridge + legacy overrides included) — previously HERMES_TOOL_PROGRESS_MODE could leave bubbles on while the gate believed progress was quiet (test proven red → green).
  • Robust against MagicMock-style probes (is True guards on the new capability reads) so existing progress suites are unaffected.

New: transient activity overlay (opt-in, inside the same message):

  • tool starts (default on; streaming_single_message_activity) and thinking snippets (opt-in; streaming_single_message_thinking) render under a --- rule while the turn runs, and are replaced by real text / never in the final edit.
  • Plumbed from the runner: tool.started + _thinking events route into the consumer's existing overlay channel (the same one native streams use) even though progress bubbles are quiet.

Tests: 74→(now) green in test_stream_consumer.py incl. the new overlay-switch gate test; test_stream_consumer_tool_progress.py +7 (gates + drain behavior incl. draft-lane integration); 10 surrounding suites green; adversarial scenario bench (overflow paging, edit-fail/raise fallbacks, no-id sentinel, cancel, fresh-final) shows zero lost/duplicated text.

Follow-ups noted (not in this PR): draft can_stop stop-button wiring (Bot API 10.3) and a group-chat fallback (drafts are private-chat-only; the edit path already handles groups).

@Maziyang2

Copy link
Copy Markdown
Author

4096-split policy switch (deferred vs eager pagination)

Added display.platforms.<p>.streaming_single_message_4096_split (default false) per the owner's rule for over-limit single-message turns:

  • false (default) — deferred pagination: the single preview is not cut mid-turn; the live phase stays one message. If the completed text overflows, pagination arrives at completion via the adapter's overflow split (content complete — nothing is lost). Rich-eligible content first rides the platform's 32,768-char cap before any split.
  • true — eager ≤4096 sealing mid-stream: the reply arrives as several long (≤4096) messages as it fills.

Tests: deferred keeps sends == 1 with the over-limit final edit carrying the whole payload; eager seals the filled head into its own message; both paths replayed for zero lost/duplicated characters; config resolution (default / override / master-gate collapse). Suite: 101 passed in the two streaming files. Docs updated (cli-config + messaging guide).

@Maziyang2

Copy link
Copy Markdown
Author

Update: Telegram platform features (Bot API 10.3), each behind its own opt-in switch — commit cf88086d0a.

  • Stop button (telegram.extra.stop_button): DM draft previews carry can_stop via the raw API; a stopped_message_generation update re-dispatches as a synthetic /stop into the standard interrupt path. Disabled or unauthorized presses are ignored (fail-closed).
  • Checklists (telegram.extra.checklist_emoji + ephemeral_for-style user scope): native sendChecklist when the constraint is satisfiable; business-accounts-only per the Bot API, so other chats keep the MarkdownV2 fallback.
  • Message effects (display.platforms.telegram.message_effects + message_effect_min_seconds, default 60 s): an animated effect on the final reply of long turns. Emoji→effect-id map validated live against the Bot API (👍 👎 🔥 🎉 🎊 all echo effect_id; ❤ requires paid sends → excluded); numeric strings pass through as explicit ids. Private chats only, per the API.
  • Ephemeral (telegram.extra.ephemeral_messages): group busy-acks are whispered to the sender via ephemeral_message_parameters; send failures never fall back to a group-wide send.

Also in this push: stateful multi-line expandable-blockquote conversion (the activity-overlay framing), and a Telegram adapter-extras test file plus busy-ack ephemeral tests. Docs: website/docs/user-guide/messaging/index.md + cli-config.yaml.example.

Verification: 12-suite gateway regression 307 passed / 0 failed (format, adapter extras, consumer, display-config, busy family, polling-progress).

@Maziyang2 Maziyang2 changed the title feat(gateway): Telegram single-message streaming with quiet-progress gate + docs (#110564) feat(telegram): single-message streaming + Bot API 10.3 extras (stop button, effects, ephemerals), all opt-in (#110564) Sep 15, 2026
@Maziyang2
Maziyang2 force-pushed the feat/telegram-single-message-streaming branch from a225bb5 to 7d4c4e7 Compare September 16, 2026 13:20
@Maziyang2

Copy link
Copy Markdown
Author

Refresh: rebased onto current main — MERGEABLE again. The one conflict (gateway/run_turn.py) came from the tool-progress refactor; the single-message gate now resolves the effective mode via the shared resolve_tool_progress() instead of our earlier private helper, so the quiet-progress gate and the mode the turn actually runs with stay in lockstep by construction. No behavior changes beyond that alignment.

Local verification: every test file this PR touches is green (stream consumer, tool-progress, telegram adapter/format/rich-messages, display config, busy ack — 250 tests).

Unchanged context: this ships alongside #111631 (mirroring local-surface turns) — both fully opt-in behind independent switches, no code dependency between the two, so either can land first. Happy to adjust scope or split further if that helps review.

…gate + docs (NousResearch#110564)

Keeps one editable streaming preview alive for a whole agent turn — text on
both sides of tool calls keeps updating the same message instead of sealing a
new one at every boundary. Opt-in via
display.platforms.telegram.streaming_single_message; Telegram-only; active
while text tool-progress is quiet (off/log).

Approach prototyped in NousResearch#110571 (@KoNit-K) — same config key and boundary
skip — extended with:

- quiet-progress gate also accepts tool_progress: log (no chat bubbles either;
  matches the gateway's own {"off", "log"} quiet definition);
- mid-turn commentary no longer resets the evolving preview;
- docs (messaging guide + cli-config.yaml.example) and boundary tests
  (2 tool boundaries, commentary continuity, quiet-mode gating).

Tests: scripts/run_tests.sh tests/gateway/test_stream_consumer.py
       base: 3 failed → with change: 74 passed, 1 skipped;
       adjacent stream/display suites: 98 passed, 0 failed.
…vity overlay, docs

Follow-up hardening on the single-message streaming opt-in (NousResearch#110564), found and
fixed by exercising the real Telegram Bot API (10.3 docs + live probes) against
the feature:

- Gate on the EFFECTIVE tool-progress mode via the shared
  `resolve_tool_progress()` resolver — the same winning source the display
  path runs with (YAML wins, null inherits to the env bridge, tier defaults
  last) — so the single-message gate can never disagree with the mode the
  turn actually runs with. Previously the env bridge could leave tool
  bubbles ON while the gate treated progress as quiet. Red → green test
  included.
- Transient activity overlay (new, opt-in): while the single preview runs, tool
  starts and thinking snippets render under a `---` rule inside the same
  message and are replaced the moment real text arrives — the final edit never
  carries them. Switches: display.platforms.<p>.streaming_single_message_activity
  (default true; quiet tool_progress would otherwise show nothing during tool
  runs) and streaming_single_message_thinking (opt-in). Works on both Telegram
  transports: the ephemeral draft preview (sendMessageDraft) and the edit path;
  tool lines reuse the consumer's existing native-overlay machinery.
- Wire the runner: tool.started events and `_thinking` (scratch/reasoning)
  events route into the overlay even though progress bubbles are quiet — before
  the progress-queue guard, since single mode runs without a queue.

Docs: cli-config.yaml.example + messaging guide updated; verified against the
Bot API reference (edit window: unlimited for bot-sent messages; drafts: 30s
ephemeral preview, private chats only; 4096-char cap enforced server-side with
codepoint counting — the consumer's utf16 accounting stays conservative).
Per the requested rule for the single-message streaming mode:

- `streaming_single_message_4096_split: true` — over-limit content is cut into
  several ≤4096 messages as the preview fills (classic eager sealing).
- Default `false` — deferred pagination: the preview is NOT cut while the turn
  runs; an over-limit turn-final is paged by the adapter's overflow split
  instead, so the live phase stays one message and pages only appear when the
  text genuinely overflows (rich-eligible content first rides the platform's
  32,768-character cap). Nothing is ever lost — pagination just arrives later.

Consumer gate `_eager_overflow_split()`; config plumbing in run_turn +
display_config; docs in cli-config.yaml.example and the messaging guide.

Tests: deferred keeps `sends == 1` with an over-limit final edit; eager seals
into multiple messages; both replayed for zero lost/duplicated characters;
config resolution incl. master-gate collapse. 101 passed in the two suites.
…sts, message effects, ephemeral scope

- stop_button extra: DM draft previews carry can_stop via the raw API; a
  stopped_message_generation update re-dispatches as a synthetic /stop into
  the standard interrupt path; disabled or unauthorized presses are ignored
- checklist_emoji extra + sender scope → native sendChecklist (business
  accounts only; MarkdownV2 fallback elsewhere)
- message_effects switch: validated effect ids (👍👎🔥🎉🎊, live-verified;
  ❤ requires paid sends and is excluded), numeric pass-through, private chats
  only; attach on turn-final sends after message_effect_min_seconds (60s)
- ephemeral_messages extra: group busy-acks whisper to the sender via
  extra_metadata; no group-wide leak on failure
- formatter: stateful multi-line expandable-quote conversion (style A)\n\nTests: telegram format/adapter-extras/consumer/busy-ack suites green; 12-suite
gateway regression 307 passed.
.docs: messaging/index.md + cli-config extras section
- _try_send_rich attaches message_effect_id from metadata under the same
  gate as the legacy path (turn-final + DM-only + verified id); a completed
  reply rendered via sendRichMessage no longer silently drops the effect
- tests: rich send carries the field when requested, omits it otherwise

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/streaming Streaming responses: gateway delivery, provider wire comp/gateway Gateway runner, session dispatch, delivery P3 Low — cosmetic, nice to have platform/telegram Telegram bot adapter sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/feature New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants