Skip to content

ROB-1689 call_stream support for api chat - #759

Merged
moshemorad merged 16 commits into
masterfrom
ROB-1689-stream-ask-prework
Aug 7, 2025
Merged

moshemorad merged 16 commits into
masterfrom
ROB-1689-stream-ask-prework

Conversation

@RoiGlinik

Copy link
Copy Markdown
Collaborator

improve and simplify stream_call function
make stream message more flexible for other return formats.

@coderabbitai

coderabbitai Bot commented Jul 30, 2025 •

Copy link
Copy Markdown
Contributor

Walkthrough

A new boolean stream field was added to the chat request model. The tool-calling LLM's streaming logic was refactored to yield structured event objects instead of SSE strings, removing direct HTTP streaming and SSE helpers. A new module was created for SSE formatting utilities and streaming message models. Server endpoints were updated to support and utilize the new streaming formatters and enable streaming responses.

Changes

Cohort / File(s) Change Summary
Model Update
holmes/core/models.py
Added a stream: bool = Field(default=False) field to the ChatRequestBaseModel Pydantic model.
LLM Streaming Refactor
holmes/core/tool_calling_llm.py
Refactored call_stream to remove direct HTTP streaming and SSE string generation; now yields structured StreamMessage objects. Updated method signature to accept msgs instead of runbooks and removed the stream parameter. Simplified error handling and tool call retry logic. Removed create_sse_message helper function.
Streaming Utilities
holmes/utils/stream.py
Added a new module defining StreamEvents enum, StreamMessage Pydantic model, SSE formatting function, and generator functions to format investigation and chat streams into SSE-compliant strings.
Server Streaming Integration
server.py
Updated /api/stream/investigate endpoint to use stream_investigate_formatter with runbooks and removed commented-out streaming code. Added streaming support to /api/chat endpoint using stream_chat_formatter when stream flag is true, returning a StreamingResponse.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Note

⚡️ Unit Test Generation is now available in beta!

Learn more here, or try it out under "Finishing Touches" below.

✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch ROB-1689-stream-ask-prework

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share
🪧 Tips

Chat

There are 3 ways to chat with CodeRabbit:

  • Review comments: Directly reply to a review comment made by CodeRabbit. Example:
    • I pushed a fix in commit <commit_id>, please review it.
    • Explain this complex logic.
    • Open a follow-up GitHub issue for this discussion.
  • Files and specific lines of code (under the "Files changed" tab): Tag @coderabbitai in a new review comment at the desired location with your query. Examples:
    • @coderabbitai explain this code block.
  • PR comments: Tag @coderabbitai in a new PR comment to ask questions about the PR branch. For the best results, please provide a very specific query, as very limited context is provided in this mode. Examples:
    • @coderabbitai gather interesting stats about this repository and render them as a table. Additionally, render a pie chart showing the language distribution in the codebase.
    • @coderabbitai read src/utils.ts and explain its main purpose.
    • @coderabbitai read the files in the src/scheduler package and generate a class diagram using mermaid and a README in the markdown format.

Support

Need help? Create a ticket on our support page for assistance with any issues or questions.

CodeRabbit Commands (Invoked using PR comments)

  • @coderabbitai pause to pause the reviews on a PR.
  • @coderabbitai resume to resume the paused reviews.
  • @coderabbitai review to trigger an incremental review. This is useful when automatic reviews are disabled for the repository.
  • @coderabbitai full review to do a full review from scratch and review all the files again.
  • @coderabbitai summary to regenerate the summary of the PR.
  • @coderabbitai generate docstrings to generate docstrings for this PR.
  • @coderabbitai generate sequence diagram to generate a sequence diagram of the changes in this PR.
  • @coderabbitai generate unit tests to generate unit tests for this PR.
  • @coderabbitai resolve resolve all the CodeRabbit review comments.
  • @coderabbitai configuration to show the current CodeRabbit configuration for the repository.
  • @coderabbitai help to get help.

Other keywords and placeholders

  • Add @coderabbitai ignore anywhere in the PR description to prevent this PR from being reviewed.
  • Add @coderabbitai summary to generate the high-level summary at a specific location in the PR description.
  • Add @coderabbitai anywhere in the PR title to generate the title automatically.

CodeRabbit Configuration File (.coderabbit.yaml)

  • You can programmatically configure CodeRabbit by adding a .coderabbit.yaml file to the root of your repository.
  • Please see the configuration documentation for more information.
  • If your editor has YAML language server enabled, you can add the path at the top of this file to enable auto-completion and validation: # yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json

Documentation and Community

  • Visit our Documentation for detailed information on how to use CodeRabbit.
  • Join our Discord Community to get help, request features, and share feedback.
  • Follow us on X/Twitter for updates and announcements.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 3acb328 and 0620000.

📒 Files selected for processing (4)
  • holmes/core/models.py (2 hunks)
  • holmes/core/tool_calling_llm.py (5 hunks)
  • holmes/utils/stream.py (1 hunks)
  • server.py (3 hunks)
🧰 Additional context used
📓 Path-based instructions (1)
**/*.py

📄 CodeRabbit Inference Engine (CLAUDE.md)

**/*.py: Use Ruff for formatting and linting (configured in pyproject.toml)
Type hints required (mypy configuration in pyproject.toml)

Files:

  • server.py
  • holmes/core/models.py
  • holmes/utils/stream.py
  • holmes/core/tool_calling_llm.py
🧠 Learnings (2)
server.py (1)

Learnt from: nherment
PR: #610
File: .github/workflows/llm-evaluation.yaml:39-42
Timestamp: 2025-07-08T08:45:41.069Z
Learning: The robusta-dev/holmesgpt codebase has comprehensive existing validation for Azure environment variables (AZURE_API_BASE, AZURE_API_KEY, AZURE_API_VERSION) and MODEL in tests/llm/utils/classifiers.py, tests/llm/conftest.py, and holmes/core/llm.py. Don't suggest adding redundant validation logic.

holmes/core/models.py (1)

Learnt from: nherment
PR: #436
File: tests/llm/utils/mock_utils.py:240-249
Timestamp: 2025-06-05T12:23:27.634Z
Learning: The holmesgpt project uses Python >= 3.10 and prefers modern type hint syntax like list[str], dict[str, int] over importing equivalent types from the typing module like List[str], Dict[str, int].

🧬 Code Graph Analysis (1)
holmes/utils/stream.py (1)
holmes/core/tool_calling_llm.py (1)
  • call_stream (544-676)
🪛 Ruff (0.12.2)
holmes/utils/stream.py

19-19: Do not use mutable data structures for argument defaults

Replace with None; initialize within function

(B006)

holmes/core/tool_calling_llm.py

606-608: Within an except clause, raise exceptions with raise ... from err or raise ... from None to distinguish them from errors in exception handling

(B904)

⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (3)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.12)
🔇 Additional comments (10)
holmes/core/models.py (1)

92-92: LGTM! Clean addition of streaming support.

The stream field addition to ChatRequestBaseModel is well-implemented with appropriate typing and default value. This enables clients to opt into streaming responses, which aligns perfectly with the PR objectives.

server.py (2)

26-26: LGTM! Import of streaming utilities.

The import of streaming formatters enables the server endpoints to use the new structured streaming approach.


162-174: LGTM! Clean refactoring to use structured streaming.

The streaming investigation endpoint now uses the new stream_investigate_formatter which properly handles the structured StreamMessage objects from ai.call_stream. The runbooks parameter is correctly passed for inclusion in the output.

holmes/utils/stream.py (3)

8-17: LGTM! Well-designed streaming message structure.

The StreamEvents enum and StreamMessage model provide a clean, structured approach to handling streaming events. This separation of concerns improves maintainability compared to the previous direct SSE string generation.


23-42: LGTM! Clean investigation stream formatting.

The stream_investigate_formatter properly processes ANSWER_END events by extracting structured sections and analysis using the existing process_response_into_sections function. The runbooks parameter is correctly included in the output.


44-60: LGTM! Clean chat stream formatting.

The stream_chat_formatter properly handles chat-specific streaming by including analysis, conversation history, and follow-up actions in the response. The optional followups parameter provides good flexibility.

holmes/core/tool_calling_llm.py (4)

38-38: LGTM! Import of streaming utilities.

The import enables the method to use the new structured streaming events.


544-563: LGTM! Clean method signature and documentation.

The method signature is well-updated to use msgs parameter instead of the previous runbooks parameter, and the documentation clearly explains that this doesn't use stream=True but streams Holmes iterations instead.


637-642: LGTM! Clean structured event generation.

The method now yields structured StreamMessage objects with ANSWER_END events containing both content and messages. This is much cleaner than the previous SSE string generation approach.


656-672: LGTM! Well-structured tool event generation.

The method properly yields START_TOOL and TOOL_RESULT events during tool execution, providing good visibility into the AI's progress. The concurrent execution is preserved while adding streaming capabilities.

Comment thread holmes/core/tool_calling_llm.py Outdated
Comment thread holmes/core/tool_calling_llm.py
Comment thread holmes/utils/stream.py Outdated
Comment thread server.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

♻️ Duplicate comments (1)
holmes/core/tool_calling_llm.py (1)

579-579: Tool disabling logic still uses wrong variable.

This is the same issue flagged in previous reviews - the condition should use the local max_steps variable instead of self.max_steps to respect any adjustments made during retry logic.

Apply this fix:

-            tools = None if i == max_steps else tools
+            tools = None if i == max_steps else tools

Wait, I notice the code already shows i == max_steps on line 579, but the past review comment mentioned i == self.max_steps - 1. Let me check if this has actually been fixed.

🧹 Nitpick comments (1)
holmes/core/tool_calling_llm.py (1)

679-679: Consider using local max_steps variable in error message.

For consistency and accuracy, consider using the local max_steps variable instead of self.max_steps in the error message, especially since max_steps can be incremented during retry logic.

-        raise Exception(
-            f"Too many LLM calls - exceeded max_steps: {i}/{self.max_steps}"
-        )
+        raise Exception(
+            f"Too many LLM calls - exceeded max_steps: {i}/{max_steps}"
+        )
📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 90b7243 and 897cc41.

📒 Files selected for processing (1)
  • holmes/core/tool_calling_llm.py (5 hunks)
🧰 Additional context used
📓 Path-based instructions (1)
**/*.py

📄 CodeRabbit Inference Engine (CLAUDE.md)

**/*.py: Use Ruff for formatting and linting (configured in pyproject.toml)
Type hints required (mypy configuration in pyproject.toml)

Files:

  • holmes/core/tool_calling_llm.py
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (3)
  • GitHub Check: llm_evals
  • GitHub Check: Pre-commit checks
  • GitHub Check: Pre-commit checks
🔇 Additional comments (8)
holmes/core/tool_calling_llm.py (8)

41-41: LGTM - Import supports new streaming architecture.

The import of StreamEvents and StreamMessage correctly supports the refactored streaming implementation that yields structured objects instead of raw SSE strings.


550-554: LGTM - Method signature aligns with new streaming architecture.

The signature changes are well-designed:

  • Removing stream parameter since this method always streams structured events
  • Adding msgs parameter for flexible message construction
  • Making system_prompt optional supports different use cases
  • Removing runbooks shifts responsibility appropriately to callers

556-566: LGTM - Flexible message construction.

The new message construction logic cleanly handles different input scenarios - system prompt, user prompt, and pre-constructed messages. This provides good flexibility for various use cases.


596-603: LGTM - Simplified LLM completion call.

The completion call is now simplified to always use stream=False, which aligns with the new architecture of streaming at the iteration level rather than at the LLM response level.


612-614: LGTM - Exception handling preserves context.

The exception handling has been correctly updated to use from e to preserve the original exception context, addressing the previous review feedback.


641-645: LGTM - Clean streaming event generation.

The method now yields structured StreamMessage objects with ANSWER_END events containing both content and message history. This is a clean design that separates concerns between event generation and SSE formatting.


660-663: LGTM - Tool start event generation.

Yielding START_TOOL events when tools are submitted provides good visibility into tool execution progress for streaming consumers.


673-676: LGTM - Tool result event generation.

Yielding TOOL_RESULT events with structured tool results provides comprehensive information about tool execution outcomes to streaming consumers.

@RoiGlinik RoiGlinik changed the title DRAFT ROB-1689 call_stream support for api chat ROB-1689 call_stream support for api chat Jul 31, 2025
@moshemorad
moshemorad self-requested a review August 3, 2025 11:30

@moshemorad moshemorad left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

🧹 Nitpick comments (1)
holmes/core/tool_calling_llm.py (1)

681-683: Verify exception message consistency.

The exception message uses both i and self.max_steps which could be confusing. Consider using the local max_steps variable for consistency with the rest of the method.

        raise Exception(
-            f"Too many LLM calls - exceeded max_steps: {i}/{self.max_steps}"
+            f"Too many LLM calls - exceeded max_steps: {i}/{max_steps}"
        )
📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between b46d64a and 8866e9b.

📒 Files selected for processing (2)
  • holmes/core/models.py (2 hunks)
  • holmes/core/tool_calling_llm.py (6 hunks)
🚧 Files skipped from review as they are similar to previous changes (1)
  • holmes/core/models.py
🧰 Additional context used
📓 Path-based instructions (1)
**/*.py

📄 CodeRabbit Inference Engine (CLAUDE.md)

**/*.py: ALWAYS place Python imports at the top of the file, not inside functions or methods
Use Ruff for formatting and linting (configured in pyproject.toml)
Type hints required (mypy configuration in pyproject.toml)

Files:

  • holmes/core/tool_calling_llm.py
🧠 Learnings (1)
📚 Learning: the robusta-dev/holmesgpt codebase has comprehensive existing validation for azure environment varia...
Learnt from: nherment
PR: robusta-dev/holmesgpt#610
File: .github/workflows/llm-evaluation.yaml:39-42
Timestamp: 2025-07-08T08:45:41.069Z
Learning: The robusta-dev/holmesgpt codebase has comprehensive existing validation for Azure environment variables (AZURE_API_BASE, AZURE_API_KEY, AZURE_API_VERSION) and MODEL in tests/llm/utils/classifiers.py, tests/llm/conftest.py, and holmes/core/llm.py. Don't suggest adding redundant validation logic.

Applied to files:

  • holmes/core/tool_calling_llm.py
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (3)
  • GitHub Check: Pre-commit checks
  • GitHub Check: llm_evals
  • GitHub Check: Pre-commit checks
🔇 Additional comments (10)
holmes/core/tool_calling_llm.py (10)

15-16: LGTM - Import organization follows coding guidelines.

The imports are properly placed at the top of the file as required by the coding guidelines.


41-41: LGTM - New streaming utilities import.

The import of StreamEvents and StreamMessage aligns with the refactoring to use structured streaming objects instead of raw SSE strings.


552-556: Method signature change improves flexibility.

The removal of the stream parameter and replacement of runbooks with msgs makes the method more generic and allows for different message structures to be passed in.


558-568: Clear documentation and improved message construction.

The docstring clearly explains that this doesn't use llm.completion(stream=true) but streams iterations instead. The conditional message construction logic is well-structured and handles different input scenarios properly.


570-575: Variable initialization aligns with refactoring.

The use of tool_calls: list[dict] = [] with explicit type annotation and max_steps = self.max_steps local variable assignment is consistent with the streaming refactoring approach.


581-582: Tool disabling logic is correct.

The condition tools = None if i == max_steps else tools correctly disables tools on the final iteration, which addresses the previous review concern about max_steps consistency. The logic now matches the non-streaming version.


598-617: Exception handling properly preserves context.

The addition of from e in the exception chain at line 615 correctly addresses the previous review feedback about preserving exception context. The non-streaming LLM completion call is appropriate for this iteration-based streaming approach.


619-648: Structured streaming implementation is well-designed.

The refactoring to yield StreamMessage objects with appropriate event types (ANSWER_END) instead of raw SSE strings provides better abstraction and flexibility. The early return when no tool calls are present is efficient and correct.


663-666: Tool start event streaming is well-implemented.

Yielding StreamMessage with START_TOOL event immediately when submitting tool execution provides real-time feedback to the client about tool execution progress.


676-679: Tool result streaming completes the event flow.

The TOOL_RESULT event with as_streaming_tool_result_response() data provides comprehensive tool execution feedback and maintains consistency with the streaming architecture.

@github-actions

github-actions Bot commented Aug 7, 2025

Copy link
Copy Markdown
Contributor

Results of HolmesGPT evals

  • ask_holmes: 25/42 test cases were successful, 0 regressions, 1 skipped, 16 mock failures
Test suite Test case Status
ask 01_how_many_pods ✅
ask 02_what_is_wrong_with_pod 🔧
ask 03_what_is_the_command_to_port_forward 🔧
ask 04_related_k8s_events ↪️
ask 05_image_version 🔧
ask 09_crashpod ✅
ask 10_image_pull_backoff 🔧
ask 11_init_containers ✅
ask 14_pending_resources ✅
ask 15_failed_readiness_probe ✅
ask 17_oom_kill ✅
ask 18_crash_looping_v2 ✅
ask 19_detect_missing_app_details 🔧
ask 24_misconfigured_pvc 🔧
ask 28_permissions_error ✅
ask 29_events_from_alert_manager 🔧
ask 39_failed_toolset 🔧
ask 41_setup_argo ✅
ask 42_dns_issues_steps_new_tools 🔧
ask 43_current_datetime_from_prompt ✅
ask 45_fetch_deployment_logs_simple ✅
ask 51_logs_summarize_errors 🔧
ask 53_logs_find_term ✅
ask 54_not_truncated_when_getting_pods 🔧
ask 59_label_based_counting ✅
ask 60_count_less_than 🔧
ask 61_exact_match_counting ✅
ask 63_fetch_error_logs_no_errors ✅
ask 77_liveness_probe_misconfiguration 🔧
ask 79_configmap_mount_issue 🔧
ask 83_secret_not_found 🔧
ask 86_configmap_like_but_secret 🔧
ask 88_affinity_like_but_taints 🔧
ask 89_runbook_missing_cloudwatch 🔧
ask 90_runbook_basic_selection 🔧
ask 93_calling_datadog ✅
ask 93_calling_datadog ✅
ask 93_calling_datadog ✅
ask 97_logs_clarification_needed ✅
ask 100_historical_logs 🔧
ask 24a_misconfigured_pvc_basic 🔧
ask 13a_pending_node_selector_basic 🔧

Legend

  • ✅ the test was successful
  • ↪️ the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🔧 the test failed due to mock data issues (not a code regression)
  • ❌ the test failed and should be fixed before merging the PR

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants