Skip to content

Fix/stream disconnect thread hang - #2207

Closed
gesila1073 wants to merge 2 commits into
HolmesGPT:masterfrom
gesila1073:fix/stream-disconnect-thread-hang
Closed

gesila1073 wants to merge 2 commits into
HolmesGPT:masterfrom
gesila1073:fix/stream-disconnect-thread-hang

Conversation

@gesila1073

@gesila1073 gesila1073 commented Jun 19, 2026 •

Copy link
Copy Markdown

Summary by CodeRabbit

Release Notes

  • Improvements
    • Users can now interrupt ongoing chat requests with proper cancellation signaling throughout the system
    • Improved handling of client disconnections during streaming responses with enhanced cleanup procedures
    • Better error handling for interrupted chat operations

@linux-foundation-easycla

Copy link
Copy Markdown

CLA Missing ID

@coderabbitai

coderabbitai Bot commented Jun 19, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Introduces a threading.Event-based cancellation mechanism for streaming /api/chat responses. A new LLMInterruptedError exception is defined in holmes/core/exceptions.py and re-exported from tool_calling_llm.py. The server adds an async disconnect watcher that sets the event on HTTP client disconnect, propagates it through the streaming pipeline, and the stream formatter silently swallows the interruption instead of emitting an SSE error.

Changes

HTTP Disconnect Cancellation for Streaming Chat

Layer / File(s) Summary
LLMInterruptedError definition and call_stream cancellation
holmes/core/exceptions.py, holmes/core/tool_calling_llm.py
Defines LLMInterruptedError in exceptions.py, re-exports it from tool_calling_llm.py, adds log statements before raising, and switches the ThreadPoolExecutor to an explicit try/finally that calls shutdown(wait=False, cancel_futures=True) on both cancellation and normal exit.
stream_chat_formatter silent interruption handling
holmes/utils/stream.py
Imports LLMInterruptedError and adds a dedicated except clause that logs the interruption and returns early, bypassing the existing generic SSE error emission path.
Server disconnect watcher, _stream_with_trace_cleanup, and /api/chat wiring
server.py
Adds asyncio/anyio.from_thread imports; introduces _schedule_disconnect_watcher to poll request.is_disconnected() and set a threading.Event; reworks _stream_with_trace_cleanup to accept and set an optional cancel_event in its finally block; wires stream_cancel_event into the /api/chat SSE response wrapper and usage recorder.
Conversation worker cancel_event propagation
holmes/core/conversations_worker/worker.py
_run_chat_and_publish creates a per-execution threading.Event and passes it as cancel_event into request_ai.call_stream.
Streaming cancellation and cleanup tests
tests/test_server_endpoints.py
Adds three tests: stream_chat_formatter yields nothing on LLMInterruptedError; _stream_with_trace_cleanup sets the cancel event and ends the trace span on normal completion and when the upstream stream raises an exception.

Sequence Diagram

sequenceDiagram
    participant HTTPClient
    participant api_chat as /api/chat handler
    participant DisconnectWatcher as _schedule_disconnect_watcher
    participant CancelEvent as stream_cancel_event
    participant call_stream as ToolCallingLLM.call_stream
    participant Formatter as stream_chat_formatter
    participant Cleanup as _stream_with_trace_cleanup

    api_chat->>CancelEvent: create threading.Event()
    api_chat->>DisconnectWatcher: schedule with CancelEvent
    api_chat->>Cleanup: wrap SSE stream with cancel_event
    loop polling
        DisconnectWatcher->>HTTPClient: request.is_disconnected()
    end
    HTTPClient-->>DisconnectWatcher: disconnected=True
    DisconnectWatcher->>CancelEvent: set()
    CancelEvent-->>call_stream: cancel_event.is_set() → raise LLMInterruptedError
    Formatter->>Formatter: catch LLMInterruptedError → return (no SSE error)
    Cleanup->>Cleanup: finally → end OTel span, set CancelEvent
    CancelEvent-->>DisconnectWatcher: watcher exits loop
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • HolmesGPT/holmesgpt#759: Directly touches ToolCallingLLM.call_stream, the /api/chat streaming path, and stream_chat_formatter — the same three components this PR extends with cancel_event/LLMInterruptedError handling.
  • HolmesGPT/holmesgpt#854: Modifies the same except block in stream_chat_formatter (adding RateLimitError SSE emission) that this PR modifies to add the LLMInterruptedError early-return path.
  • HolmesGPT/holmesgpt#935: Also modifies stream_chat_formatter's exception-handling logic, overlapping directly with this PR's new LLMInterruptedError handler at the same function level.

Suggested reviewers

  • nherment
  • moshemorad
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 43.75% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Fix/stream disconnect thread hang' clearly and specifically describes the main objective of this changeset: fixing a thread hang issue related to stream disconnection.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Jun 19, 2026

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit a6fb763
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a3595bba7076400087f1e47
😎 Deploy Preview https://deploy-preview-2207--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (2)
tests/test_server_endpoints.py (1)

597-621: 💤 Low value

Consider moving import threading to the top of the file.

The import threading on lines 604 and 632 is repeated inside each test function. Moving it to the file's import section would be more consistent with the coding guidelines.

Also applies to: 624-649

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/test_server_endpoints.py` around lines 597 - 621, The import statement
for threading is repeated inside multiple test functions including
test_stream_with_trace_cleanup_sets_cancel_event_on_normal_completion. Remove
the import threading statement from inside all test function bodies and add it
once at the top of the file with the other import statements to follow standard
Python conventions and avoid code duplication.

Source: Coding guidelines

holmes/core/conversations_worker/worker.py (1)

862-863: 💤 Low value

Cancel event plumbed but currently unused in worker context.

The cancel_event is created and passed to call_stream, but unlike the server path (which has _schedule_disconnect_watcher), nothing sets this event in the worker context. This is fine — it provides infrastructure for future stop-conversation support — but currently cancellation won't trigger for worker-driven chats.

Also applies to: 947-947

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@holmes/core/conversations_worker/worker.py` around lines 862 - 863, The
cancel_event created in the worker context is prepared and passed to call_stream
but there is no mechanism to actually trigger it, unlike the server path which
uses _schedule_disconnect_watcher. To address this, add a clarifying comment
near where cancel_event is instantiated documenting that while the event
provides infrastructure for future stop-conversation support, it is currently
intentionally not being set in the worker context and should be implemented when
that feature is added in the future.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@holmes/core/conversations_worker/worker.py`:
- Around line 862-863: The cancel_event created in the worker context is
prepared and passed to call_stream but there is no mechanism to actually trigger
it, unlike the server path which uses _schedule_disconnect_watcher. To address
this, add a clarifying comment near where cancel_event is instantiated
documenting that while the event provides infrastructure for future
stop-conversation support, it is currently intentionally not being set in the
worker context and should be implemented when that feature is added in the
future.

In `@tests/test_server_endpoints.py`:
- Around line 597-621: The import statement for threading is repeated inside
multiple test functions including
test_stream_with_trace_cleanup_sets_cancel_event_on_normal_completion. Remove
the import threading statement from inside all test function bodies and add it
once at the top of the file with the other import statements to follow standard
Python conventions and avoid code duplication.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: c088394d-07c9-4ef4-bb5b-17d0157c299d

📥 Commits

Reviewing files that changed from the base of the PR and between 5b9f943 and a6fb763.

📒 Files selected for processing (7)
  • holmes/core/conversations_worker/worker.py
  • holmes/core/exceptions.py
  • holmes/core/models.py
  • holmes/core/tool_calling_llm.py
  • holmes/utils/stream.py
  • server.py
  • tests/test_server_endpoints.py

@gesila1073 gesila1073 closed this Jun 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants