Skip to content

fix(slack): harden Assistant status lifecycle - #62248

Closed
KCAYAAI wants to merge 3 commits into
NousResearch:mainfrom
KCAYAAI:fix/clarify-resume-typing-after-reply
Closed

KCAYAAI wants to merge 3 commits into
NousResearch:mainfrom
KCAYAAI:fix/clarify-resume-typing-after-reply

Conversation

@KCAYAAI

@KCAYAAI KCAYAAI commented Jul 10, 2026 •

Copy link
Copy Markdown

Summary

Harden the Slack Assistant status lifecycle in both directions:

  • Resume the existing run's typing/status indicator after a clarify reply resolves.
  • Clear the exact Slack thread's status when a turn finishes, including concurrent-thread, transient-clear-failure, streaming-finalize, and send-error paths.

Together these changes prevent both kinds of misleading Slack UI state: an active run that looks silent after clarification, and a completed run that remains stuck on is thinking... / Evaluating....

Problems

Clarification replies do not resume status

The clarify callback deliberately pauses typing while it waits so Slack's Assistant API does not disable the compose box. When the user answers, the gateway intercepts the reply and resumes the existing agent run, but the interception path returned an empty acknowledgment without re-enabling typing. The run continued silently until the independent long-running heartbeat fired.

Completed turns can leave stale status

Slack Assistant status is server-side state keyed by channel and thread. The adapter previously tracked only one thread_ts per chat_id, so overlapping threads could overwrite each other. Cleanup also removed tracking before Slack confirmed assistant.threads.setStatus(status=""); a transient failure therefore left a visible status with no local record available for retry. Failed final sends could leave the same stale state.

Fix

  • Resume the active adapter's typing state after a pending clarification resolves.
  • Track active Slack Assistant statuses per channel and thread.
  • Clear only the completed turn's exact thread after final post/edit delivery.
  • Preserve other concurrent threads in the same DM/channel.
  • Remove a tracked thread only after Slack confirms the empty-status clear.
  • Keep failed clears tracked so later cleanup attempts can retry.
  • Bound per-conversation status tracking at 128 entries so permanent
    API/scope failures cannot create an unbounded local leak.
  • Resolve reply_to/metadata once and attempt exact-thread cleanup when a
    final Slack send fails.
  • Keep all status operations best-effort so platform failures do not break the agent turn.

Reproduction

Clarify resume

  1. Have an agent invoke the clarify tool in Slack.
  2. Answer the clarification prompt.
  3. Before this fix, the agent continues processing but Slack does not restore is thinking....

Stale status after completion

  1. Start one or more agent turns in Slack Assistant threads in the same DM/channel.
  2. Let a turn send its final response, or let its final send fail.
  3. Before this fix, Slack can continue showing is thinking... / Evaluating... even though the turn has finished.
  4. If the empty-status API call fails once, later cleanup cannot retry because the thread was already removed from local tracking.

Verification

  • pytest -q -o 'addopts=' tests/gateway/test_clarify_active_session_bypass.py tests/gateway/test_typing_indicator_toggle.py tests/gateway/test_keep_typing_timeout.py tests/gateway/test_slack.py
    • 232 passed
  • ruff check gateway/run.py plugins/platforms/slack/adapter.py tests/gateway/test_clarify_active_session_bypass.py tests/gateway/test_slack.py
    • passed
  • git diff --check
    • passed

Regression coverage includes:

  • clarification reply resumes paused typing;
  • concurrent Slack threads remain independently tracked;
  • exact-thread final post/edit cleanup;
  • failed clear remains tracked for retry;
  • mixed clear success/failure retains only failed threads;
  • permanent status-set failures remain bounded;
  • failed final sends target the exact thread for both metadata and
    reply_to-only callers.

Live verification

  • Linux
  • Slack Socket Mode / Assistant thread status
  • Two gateway profiles sharing the same runtime source
  • Reproduced a stale post-response status, applied the patch, restarted both profiles, and confirmed in Slack that the indicator now appears during work and clears after the final response.

Related: #24117, #32295, #8387

@alt-glitch alt-glitch added type/bug Something isn't working comp/gateway Gateway runner, session dispatch, delivery platform/slack Slack app adapter sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages P2 Medium — degraded but workaround exists labels Jul 10, 2026
@teknium1 teknium1 added the sweeper:blast-moderate Sweeper blast radius: moderate — a subsystem or single platform label Jul 11, 2026
@KCAYAAI KCAYAAI changed the title fix(gateway): resume typing after clarify reply fix(slack): harden Assistant status lifecycle Jul 12, 2026
@teknium1

Copy link
Copy Markdown
Collaborator

Partially merged via #64621 — your clarify-resume commit was cherry-picked onto current main with your authorship preserved in git log (e9d564c09). That was the surviving half of this PR: the typing indicator now resumes immediately after a user answers a clarify prompt, matching what /approve already did.

The status-tracking half (per-thread set, send-error clearing, edit_message metadata) was superseded by #63709, which landed the same fixes with workspace-scoped (team, channel, thread) keys for Slack Connect setups.

Thanks @KCAYAAI!

@teknium1 teknium1 closed this Jul 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists platform/slack Slack app adapter sweeper:blast-moderate Sweeper blast radius: moderate — a subsystem or single platform sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants