Skip to content

feat(runbook): llm select the correct runbook from catalog - #533

Closed
mainred wants to merge 5 commits into
masterfrom
runbook-selection
Closed

mainred wants to merge 5 commits into
masterfrom
runbook-selection

Conversation

@mainred

@mainred mainred commented Jun 17, 2025 •

Copy link
Copy Markdown
Collaborator

fix: #473

Changes in this PR:

  • allo user to specify customized runbook under holmes ask
  • introduce a RunbookCatalogManager, which returns runbook per user question by LLM
  • RunbookCatalogManager will return custom runbook if any
  • Add a runbook catalog, which contains internal runbooks written in markdown

impacted function

only holmes ask CLI is expected to be impacted

Why use LLM to pick the runbook

The original thought is to introduce a runbook tool to fetch runbook from the runbook catalog folder, but considering we'll have more runbooks to be introduced and selecting the runbook by the either label or tag is limited, introducing LLM to pick the matched one should be the final solution. Also, it's best to select the runbook after customer specify their question, and the following actions should be based on runbook. With all these in mind, I decided to use LLM to pick the runbook during the early phase of the investigation. This is similar to we have achieved in holmes investigate

Tests

➜  holmesgpt git:(runbook-selection) ✗ ./dist/holmes/holmes ask "detect why the k8s pod client under namespace test-ns cannot resolve dns" --max-steps 20
Failed to fetch holmes info                                                                                                          robusta_client.py:23
loading toolset status from cache                                                                                                  toolset_manager.py:216
User: detect why the k8s pod client under namespace test-ns cannot resolve dns
Found runbook for question 'detect why the k8s pod client under namespace test-ns cannot resolve dns':                            tool_calling_llm.py:181
networking/dns_troubleshooting_instructions.md
Running tool kubectl_find_resource: kubectl get -A --show-labels -o wide pod | grep client                                                   tools.py:125
Running tool kubectl_describe: kubectl describe pod client -n test-ns                                                                        tools.py:125
Running tool kubectl_logs_all_containers: kubectl logs client -n test-ns --all-containers                                                    tools.py:125
Token limit exceeded. Truncating tool responses.                                                                                  tool_calling_llm.py:201
Running tool kubectl_describe: kubectl describe pod coredns- -n kube-system                                                                  tools.py:125
Running tool kubectl_describe: kubectl describe service kube-dns -n kube-system                                                              tools.py:125
Running tool kubectl_get_yaml: kubectl get -o yaml pod client -n test-ns                                                                     tools.py:125
Running tool trace_dns_gadget: fetch DNS queries and responses of pod test-ns/client for 30 seconds                                          tools.py:125
Running tool kubectl_get_by_kind_in_namespace: kubectl get --show-labels -o wide networkpolicy -n test-ns                                    tools.py:125
Running tool kubectl_get_yaml: kubectl get -o yaml networkpolicy default-deny-egress -n test-ns                                              tools.py:125
tool_calling_llm.call - completed in 9 iterations - 74078ms                                                                      performance_timing.py:41
AI: Pod client in namespace test-ns cannot resolve DNS because the default-deny-egress NetworkPolicy blocks all outbound traffic, including DNS requests to
CoreDNS (10.0.0.10). No egress rules are defined, so DNS queries never reach the DNS server.

To fix: Add an egress rule to allow UDP/TCP traffic to 10.0.0.10 on port 53 in the default-deny-egress NetworkPolicy.

@coderabbitai

coderabbitai Bot commented Jun 17, 2025 •

Copy link
Copy Markdown
Contributor

Walkthrough

This update introduces a runbook catalog system for HolmesGPT, enabling dynamic selection and retrieval of operational runbooks based on user queries. It adds a catalog JSON, new runbook management classes, CLI/configuration support for custom runbooks, and documentation. The tool-calling LLM now integrates runbook guidance into its workflow.

Changes

File(s) Change Summary
holmes/config.py, holmes/core/tool_calling_llm.py Updated imports for RunbookCatalogManager; integrated runbook catalog manager into LLM creation and tool-calling workflow; added logic to retrieve and append relevant runbook content to user prompts.
holmes/core/runbooks.py Added RunbookCatalogManager class to manage and retrieve runbooks based on user questions using an LLM and runbook catalog; includes error handling and logging.
holmes/main.py Extended CLI and configuration to support passing custom runbooks; updated ask command to accept and propagate custom runbooks.
holmes/plugins/runbooks/init.py Introduced runbook catalog models (RunbookCatalogEntry, RunbookCatalog), functions to load the catalog JSON, and a helper for the runbook directory path.
holmes/plugins/runbooks/catalog.json New JSON file listing runbook metadata, currently including a DNS troubleshooting runbook.
holmes/plugins/runbooks/networking/dns_troubleshooting_instructions.md Added a markdown runbook with step-by-step DNS troubleshooting guidelines for Kubernetes environments.
holmes/plugins/runbooks/README.md New documentation describing the runbook system, catalog structure, and usage within HolmesGPT.

Sequence Diagram(s)

sequenceDiagram
    participant User
    participant CLI/Config
    participant ToolCallingLLM
    participant RunbookCatalogManager
    participant LLM

    User->>CLI/Config: Provide query (optionally with custom runbooks)
    CLI/Config->>ToolCallingLLM: Initialize with runbook catalog manager
    ToolCallingLLM->>RunbookCatalogManager: get_runbook_by_question(user_question)
    alt Custom runbooks provided
        RunbookCatalogManager-->>ToolCallingLLM: Return combined custom runbooks
    else No custom runbooks
        RunbookCatalogManager->>LLM: Query for matching runbook link
        LLM-->>RunbookCatalogManager: Return runbook link
        RunbookCatalogManager->>RunbookCatalogManager: Load runbook file content
        RunbookCatalogManager-->>ToolCallingLLM: Return runbook content and link
    end
    ToolCallingLLM->>ToolCallingLLM: Append runbook to user prompt
    ToolCallingLLM->>User: Return LLM response with runbook guidance
Loading

Assessment against linked issues

Objective Addressed Explanation
Implement a pre-planning phase to pick the right runbooks based on user prompt (#473) ✅

Assessment against linked issues: Out-of-scope changes

No out-of-scope changes detected.

Suggested reviewers

  • arikalon1
✨ Finishing Touches
  • 📝 Generate Docstrings

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share
🪧 Tips

Chat

There are 3 ways to chat with CodeRabbit:

  • Review comments: Directly reply to a review comment made by CodeRabbit. Example:
    • I pushed a fix in commit <commit_id>, please review it.
    • Explain this complex logic.
    • Open a follow-up GitHub issue for this discussion.
  • Files and specific lines of code (under the "Files changed" tab): Tag @coderabbitai in a new review comment at the desired location with your query. Examples:
    • @coderabbitai explain this code block.
    • @coderabbitai modularize this function.
  • PR comments: Tag @coderabbitai in a new PR comment to ask questions about the PR branch. For the best results, please provide a very specific query, as very limited context is provided in this mode. Examples:
    • @coderabbitai gather interesting stats about this repository and render them as a table. Additionally, render a pie chart showing the language distribution in the codebase.
    • @coderabbitai read src/utils.ts and explain its main purpose.
    • @coderabbitai read the files in the src/scheduler package and generate a class diagram using mermaid and a README in the markdown format.
    • @coderabbitai help me debug CodeRabbit configuration file.

Support

Need help? Create a ticket on our support page for assistance with any issues or questions.

Note: Be mindful of the bot's finite context window. It's strongly recommended to break down tasks such as reading entire modules into smaller chunks. For a focused discussion, use review comments to chat about specific files and their changes, instead of using the PR comments.

CodeRabbit Commands (Invoked using PR comments)

  • @coderabbitai pause to pause the reviews on a PR.
  • @coderabbitai resume to resume the paused reviews.
  • @coderabbitai review to trigger an incremental review. This is useful when automatic reviews are disabled for the repository.
  • @coderabbitai full review to do a full review from scratch and review all the files again.
  • @coderabbitai summary to regenerate the summary of the PR.
  • @coderabbitai generate docstrings to generate docstrings for this PR.
  • @coderabbitai generate sequence diagram to generate a sequence diagram of the changes in this PR.
  • @coderabbitai resolve resolve all the CodeRabbit review comments.
  • @coderabbitai configuration to show the current CodeRabbit configuration for the repository.
  • @coderabbitai help to get help.

Other keywords and placeholders

  • Add @coderabbitai ignore anywhere in the PR description to prevent this PR from being reviewed.
  • Add @coderabbitai summary to generate the high-level summary at a specific location in the PR description.
  • Add @coderabbitai anywhere in the PR title to generate the title automatically.

CodeRabbit Configuration File (.coderabbit.yaml)

  • You can programmatically configure CodeRabbit by adding a .coderabbit.yaml file to the root of your repository.
  • Please see the configuration documentation for more information.
  • If your editor has YAML language server enabled, you can add the path at the top of this file to enable auto-completion and validation: # yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json

Documentation and Community

  • Visit our Documentation for detailed information on how to use CodeRabbit.
  • Join our Discord Community to get help, request features, and share feedback.
  • Follow us on X/Twitter for updates and announcements.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🧹 Nitpick comments (11)
holmes/plugins/runbooks/README.md (2)

1-3: Minor grammar – add missing article

“The Runbooks folder …” reads better with the definite article:

-Runbooks folder contains operational runbooks …
+The runbooks folder contains operational runbooks …

21-23: Proper noun & sentence clarity

  1. “markdown” → “Markdown” (proper noun).
  2. Slight wording tweak for clarity:
-Catalog specified in [catalog.json](catalog.json) contains a collection of runbooks written in markdown.
+The catalog defined in [catalog.json](catalog.json) contains Markdown runbooks.
holmes/plugins/runbooks/__init__.py (1)

78-88: Clean-up: shadowing built-ins and variable naming

  1. catalogPath → catalog_path for PEP-8.
  2. The loop variable file shadows the built-in file.
-    catalogPath = os.path.join(dir_path, CATALOG_FILE)
-    if not os.path.isfile(catalogPath):
+    catalog_path = os.path.join(dir_path, CATALOG_FILE)
+    if not os.path.isfile(catalog_path):
         return None
-
-    with open(catalogPath) as file:
-        catalog_dict = json.load(file)
+    with open(catalog_path, "r") as fh:
+        catalog_dict = json.load(fh)
holmes/main.py (1)

155-160: Option help text: plural vs singular

The help string says “Path to a custom runbooks”.
Grammar is off and might confuse users:

-help="Path to a custom runbooks (can specify -r multiple times to add multiple runbooks)",
+help="Path(s) to custom runbook files (use -r multiple times)",
holmes/plugins/runbooks/networking/dns_troubleshooting_instructions.md (2)

8-10: Typo & wording

“follow the troubleshoot guide” → “follow the troubleshooting guide”.

-*   Instead of provide next steps to the user, you need to follow the troubleshoot guide to execute the steps.
+*   Instead of providing next steps to the user, follow the troubleshooting guide to execute the steps.

56-67: Markdown lint – list indentation & bare URLs

Indentation levels break MD007 and bare URLs break MD034.
Example fix for one block (apply similarly to the rest):

-*   **CRITICAL:** ALWAYS refer to the official Kubernetes DNS debugging guide for detailed troubleshooting and solutions:
-    *   Main guide: https://kubernetes.io/docs/tasks/administer-cluster/dns-debugging-resolution/
-    *   CoreDNS specific: https://kubernetes.io/docs/tasks/administer-cluster/dns-custom-nameservers/ (for CoreDNS customization which might be relevant)
+* **CRITICAL:** ALWAYS refer to the official Kubernetes DNS debugging guide:
+  * Main guide: <https://kubernetes.io/docs/tasks/administer-cluster/dns-debugging-resolution/>
+  * CoreDNS specific: <https://kubernetes.io/docs/tasks/administer-cluster/dns-custom-nameservers/> (for CoreDNS customization)

This resolves both MD007 (indent) and MD034 (wrap URLs in <>).

holmes/config.py (1)

269-276: Reuse the same LLM instance to avoid double initialisation

self._get_llm() is called twice – once for the catalog manager and again for ToolCallingLLM.
Instantiate once and reuse; this avoids duplicate network/session setup and makes mocking easier in tests.

holmes/core/runbooks.py (2)

56-60: Inefficient string build and missing separator

combined_runbooks is built via += inside a loop.
Use "\n".join(...) – faster and clearer:

-            combined_runbooks = ""
-            for runbook_str in self.runbooks:
-                combined_runbooks += f"* {runbook_str}\n"
+            combined_runbooks = "\n".join(f"* {rb}" for rb in self.runbooks)

82-86: Unnecessary else after early-return

After returning for the empty-string case, the else: block is redundant – de-indent its body for cleaner flow.

holmes/core/tool_calling_llm.py (2)

118-129: Constructor: propagate runbook_catalog_manager to subclasses

IssueInvestigator calls super().__init__(tool_executor, max_steps, llm) and therefore loses the catalog manager reference. Consider giving it a default of None and letting derived classes pass it when relevant, or document the limitation explicitly.


745-747: Parameter naming – this builds the prompt but reads oddly

add_runbook_to_user_prompt(user_prompt, runbook) actually expects the question then the runbook. Rename first parameter to question or swap the arguments for clarity.

📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 59107a7 and cdade5e.

📒 Files selected for processing (8)
  • holmes/config.py (2 hunks)
  • holmes/core/runbooks.py (2 hunks)
  • holmes/core/tool_calling_llm.py (6 hunks)
  • holmes/main.py (4 hunks)
  • holmes/plugins/runbooks/README.md (1 hunks)
  • holmes/plugins/runbooks/__init__.py (3 hunks)
  • holmes/plugins/runbooks/catalog.json (1 hunks)
  • holmes/plugins/runbooks/networking/dns_troubleshooting_instructions.md (1 hunks)
🧰 Additional context used
🧬 Code Graph Analysis (2)
holmes/main.py (2)
holmes/core/prompt.py (1)
  • append_file_to_user_prompt (4-8)
holmes/config.py (1)
  • create_console_toolcalling_llm (259-276)
holmes/core/runbooks.py (3)
holmes/core/issue.py (1)
  • Issue (13-54)
holmes/core/llm.py (1)
  • LLM (30-58)
holmes/plugins/runbooks/__init__.py (3)
  • Runbook (23-33)
  • get_runbook_folder (91-92)
  • load_catalog (78-88)
🪛 GitHub Actions: Build and test HolmesGPT
holmes/plugins/runbooks/catalog.json

[error] 11-15: End-of-file-fixer hook fixed missing newline at end of file.

holmes/core/tool_calling_llm.py

[error] 218-226: Ruff-format hook reformatted code due to line break and parentheses style changes.


[error] 579-587: Ruff-format hook reformatted code due to line break and parentheses style changes.

🪛 LanguageTool
holmes/plugins/runbooks/README.md

[uncategorized] ~2-~2: You might be missing the article “the” here.
Context: # Runbooks Runbooks folder contains operational runbooks fo...

(AI_EN_LECTOR_MISSING_DETERMINER_THE)


[uncategorized] ~8-~8: Possible missing preposition found.
Context: ... - Standardize operational processes - Enable quick onboarding for new team members -...

(AI_HYDRA_LEO_MISSING_TO)


[grammar] ~21-~21: Did you mean the formatting language “Markdown” (= proper noun)?
Context: ...ins a collection of runbooks written in markdown. During runtime, LLM will compare the r...

(MARKDOWN_NNP)

holmes/plugins/runbooks/networking/dns_troubleshooting_instructions.md

[grammar] ~8-~8: The word ‘troubleshoot’ is a verb. Did you mean the noun “troubleshooting” or “troubleshooting guide”?
Context: ...eps to the user, you need to follow the troubleshoot guide to execute the steps. * When ge...

(PREPOSITION_VERB)


[style] ~64-~64: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: ...od dnsPolicy and dnsConfig. * If NetworkPolicies are suspected, suggest ...

(ENGLISH_WORD_REPEAT_BEGINNING_RULE)


[style] ~65-~65: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: ...olicy definitions to allow DNS. * If CoreDNS configuration seems problematic...

(ENGLISH_WORD_REPEAT_BEGINNING_RULE)


[style] ~66-~66: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: ...rnetes guide on customizing it. * If upstream DNS resolution is failing, sug...

(ENGLISH_WORD_REPEAT_BEGINNING_RULE)

🪛 markdownlint-cli2 (0.17.2)
holmes/plugins/runbooks/networking/dns_troubleshooting_instructions.md

58-58: Unordered list indentation
Expected: 2; Actual: 4

(MD007, ul-indent)


58-58: Bare URL used
null

(MD034, no-bare-urls)


59-59: Unordered list indentation
Expected: 2; Actual: 4

(MD007, ul-indent)


59-59: Bare URL used
null

(MD034, no-bare-urls)


62-62: Unordered list indentation
Expected: 2; Actual: 4

(MD007, ul-indent)


63-63: Unordered list indentation
Expected: 2; Actual: 4

(MD007, ul-indent)


64-64: Unordered list indentation
Expected: 2; Actual: 4

(MD007, ul-indent)


65-65: Unordered list indentation
Expected: 2; Actual: 4

(MD007, ul-indent)


66-66: Unordered list indentation
Expected: 2; Actual: 4

(MD007, ul-indent)

🪛 Pylint (3.3.7)
holmes/plugins/runbooks/__init__.py

[refactor] 57-57: Too few public methods (0/2)

(R0903)


[refactor] 69-69: Too few public methods (0/2)

(R0903)

holmes/core/tool_calling_llm.py

[error] 178-178: Possibly using variable 'user_question' before assignment

(E0606)


[refactor] 221-229: Unnecessary "else" after "raise", remove the "else" and de-indent the code inside it

(R1720)

holmes/core/runbooks.py

[refactor] 58-58: Consider using str.join(sequence) for concatenating strings from an iterable

(R1713)


[refactor] 82-86: Unnecessary "else" after "return", remove the "else" and de-indent the code inside it

(R1705)


[refactor] 45-45: Too many return statements (7/6)

(R0911)


[refactor] 33-33: Too few public methods (1/2)

(R0903)

⏰ Context from checks skipped due to timeout of 90000ms (1)
  • GitHub Check: build (3.12)
🔇 Additional comments (1)
holmes/main.py (1)

336-337: Pass-through argument name drift

custom_runbooks from CLI is forwarded as custom_runbooks (snake_case) – good.
Verify Config.load_from_file indeed expects the same kwarg; mismatched names silently drop data.
If it expects custom_runbooks_from_cli, adapt here.

Comment thread holmes/plugins/runbooks/catalog.json Outdated
Comment thread holmes/plugins/runbooks/__init__.py
Comment thread holmes/config.py
Comment thread holmes/core/tool_calling_llm.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

♻️ Duplicate comments (1)
holmes/core/tool_calling_llm.py (1)

172-183: Runbook never reaches the LLM + possible NameError

  1. user_question is not initialised before the loop – if the messages list happens to contain no "user" role message, the NameError raised at get_runbook_by_question(user_question) will crash the call path.

  2. Even when a runbook is found, only the local user_prompt variable is updated; the original messages list that is sent to the LLM remains unchanged. Consequently, the LLM never receives the runbook.

-        for message in messages:
-            if message.get("role") == "user":
-                user_question = message.get("content", "")
+        # Grab the *first* user message and guard against missing ones
+        user_question = ""
+        for message in messages:
+            if message.get("role") == "user" and not user_question:
+                user_question = message.get("content", "")
 ...
-                user_prompt = add_runbook_to_user_prompt(user_question, runbook)  # type: ignore
+                user_prompt = add_runbook_to_user_prompt(user_question, runbook)
+                # Inject the augmented prompt back into the message list
+                for msg in messages:
+                    if msg.get("role") == "user":
+                        msg["content"] = user_prompt
+                        break

This diff both prevents the potential NameError and ensures the runbook content is actually forwarded to the model.

🧹 Nitpick comments (1)
holmes/core/tool_calling_llm.py (1)

743-745: Misleading parameter name & missing fenced formatting

add_runbook_to_user_prompt receives what you call user_prompt, but upstream you pass user_question. Consider renaming the first parameter to clarify intent and wrap the runbook in triple back-ticks (or blockquote) to avoid Markdown bleed when the runbook itself contains headings.

-def add_runbook_to_user_prompt(user_prompt: Optional[str], runbook: str) -> str:
-    return f"My instructions to check '{user_prompt}' by following the runbook:\n {runbook}"
+def add_runbook_to_user_prompt(user_question: str, runbook: str) -> str:
+    return (
+        f"My instructions to answer **{user_question}** by following the runbook below:\n"
+        f"```markdown\n{runbook}\n```"
+    )
📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between cdade5e and 8e9f9e6.

📒 Files selected for processing (2)
  • holmes/core/tool_calling_llm.py (4 hunks)
  • holmes/plugins/runbooks/catalog.json (1 hunks)
🚧 Files skipped from review as they are similar to previous changes (1)
  • holmes/plugins/runbooks/catalog.json
🧰 Additional context used
🪛 Pylint (3.3.7)
holmes/core/tool_calling_llm.py

[error] 178-178: Possibly using variable 'user_question' before assignment

(E0606)

⏰ Context from checks skipped due to timeout of 90000ms (5)
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.12)
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.12)
🔇 Additional comments (1)
holmes/core/tool_calling_llm.py (1)

118-129: Constructor change silently breaks IssueInvestigator

ToolCallingLLM.__init__ now requires a runbook_catalog_manager positional argument, but IssueInvestigator still calls super().__init__(tool_executor, max_steps, llm) (line 640). Relying on the default None value masks this behavioural change and can bite callers that forget to pass the extra argument.

Either promote the new arg to a keyword-only param or update downstream constructors explicitly:

-        super().__init__(tool_executor, max_steps, llm)
+        super().__init__(
+            tool_executor=tool_executor,
+            max_steps=max_steps,
+            llm=llm,
+            runbook_catalog_manager=None,  # or inject the real one
+        )

Comment thread holmes/core/runbooks.py Outdated
Comment thread holmes/core/runbooks.py Outdated
@github-actions

Copy link
Copy Markdown
Contributor

Results of HolmesGPT evals

Test suite Test case Status
ask_holmes 01_how_many_pods ⚠️
ask_holmes 02_what_is_wrong_with_pod ✅
ask_holmes 02_what_is_wrong_with_pod_LOKI ✅
ask_holmes 03_what_is_the_command_to_port_forward ✅
ask_holmes 04_related_k8s_events ✅
ask_holmes 05_image_version ✅
ask_holmes 06_explain_issue ✅
ask_holmes 07_high_latency ✅
ask_holmes 07_high_latency_LOKI ✅
ask_holmes 08_sock_shop_frontend ✅
ask_holmes 09_crashpod ✅
ask_holmes 10_image_pull_backoff ✅
ask_holmes 11_init_containers ✅
ask_holmes 12_job_crashing ✅
ask_holmes 12_job_crashing_CORALOGIX ⚠️
ask_holmes 12_job_crashing_LOKI ✅
ask_holmes 13_pending_node_selector ✅
ask_holmes 14_pending_resources ✅
ask_holmes 15_failed_readiness_probe ✅
ask_holmes 16_failed_no_toolset_found ✅
ask_holmes 17_oom_kill ✅
ask_holmes 18_crash_looping_v2 ✅
ask_holmes 19_detect_missing_app_details ✅
ask_holmes 20_long_log_file_search ✅
ask_holmes 20_long_log_file_search_LOKI ✅
ask_holmes 21_job_fail_curl_no_svc_account ✅
ask_holmes 22_high_latency_dbi_down ✅
ask_holmes 23_app_error_in_current_logs ✅
ask_holmes 23_app_error_in_current_logs_LOKI ✅
ask_holmes 24_misconfigured_pvc ✅
ask_holmes 25_misconfigured_ingress_class ⚠️
ask_holmes 26_multi_container_logs ✅
ask_holmes 27_permissions_error_no_helm_tools ✅
ask_holmes 28_permissions_error_helm_tools_enabled ✅
ask_holmes 29_events_from_alert_manager ✅
ask_holmes 30_basic_promql_graph_cluster_memory ✅
ask_holmes 31_basic_promql_graph_pod_memory ✅
ask_holmes 32_basic_promql_graph_pod_cpu ✅
ask_holmes 33_http_latency_graph ✅
ask_holmes 34_memory_graph ✅
ask_holmes 35_tempo ✅
ask_holmes 36_argocd_find_resource ✅
ask_holmes 37_argocd_wrong_namespace ⚠️
ask_holmes 38_rabbitmq_split_head ✅
ask_holmes 39_failed_toolset ✅
ask_holmes 40_disabled_toolset ✅
ask_holmes 41_setup_argo ✅
ask_holmes 42_dns_issues_result_all_tools ⚠️
ask_holmes 42_dns_issues_result_new_tools ⚠️
ask_holmes 42_dns_issues_result_old_tools ⚠️
ask_holmes 42_dns_issues_steps_new_all_tools ⚠️
ask_holmes 42_dns_issues_steps_new_tools ⚠️
ask_holmes 42_dns_issues_steps_old_tools ⚠️
investigate 01_oom_kill ✅
investigate 02_crashloop_backoff ✅
investigate 03_cpu_throttling ✅
investigate 04_image_pull_backoff ✅
investigate 05_crashpod ✅
investigate 05_crashpod_LOKI ✅
investigate 06_job_failure ✅
investigate 07_job_syntax_error ✅
investigate 08_memory_pressure ✅
investigate 09_high_latency ✅
investigate 10_KubeDeploymentReplicasMismatch ✅
investigate 11_KubePodCrashLooping ✅
investigate 12_KubePodNotReady ✅
investigate 13_Watchdog ✅
investigate 14_tempo ✅

Legend

  • ✅ the test was successful
  • ⚠️ the test failed but is known to be flakky or known to fail
  • ❌ the test failed and should be fixed before merging the PR

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (1)
holmes/plugins/runbooks/__init__.py (1)

63-66: Field names break Pydantic & Python naming conventions

Update_Date and KeyWords use Pascal-/Camel-Case while the rest of the codebase (and Python in general) follows snake_case. This has already been raised in a prior review and is still unresolved. Adopt snake_case field names and use alias="Update_Date" / alias="KeyWords" to preserve compatibility with the JSON.

-class RunbookCatalogEntry(BaseModel):
+class RunbookCatalogEntry(BaseModel):
@@
-    Update_Date: date
-    Description: str
-    KeyWords: list[str]
-    link: str
+    update_date: date = Field(alias="Update_Date")
+    description: str = Field(alias="Description")
+    keywords: list[str] = Field(alias="KeyWords")
+    link: str
🧹 Nitpick comments (3)
holmes/plugins/runbooks/__init__.py (1)

78-88: load_catalog() lacks error handling & memoisation

  1. A malformed catalog.json (e.g. JSON syntax error or Pydantic validation failure) will raise and crash the caller.
  2. The catalog is parsed every time load_catalog() is invoked – unnecessary I/O on hot paths.
-_catalog_cache: Optional[RunbookCatalog] = None
-
 def load_catalog() -> Optional[RunbookCatalog]:
-    dir_path = os.path.dirname(os.path.realpath(__file__))
-    catalogPath = os.path.join(dir_path, CATALOG_FILE)
-    if not os.path.isfile(catalogPath):
-        return None
-
-    with open(catalogPath) as file:
-        catalog_dict = json.load(file)
-        return RunbookCatalog(**catalog_dict)
-    return None
+    global _catalog_cache
+    if _catalog_cache is not None:
+        return _catalog_cache
+
+    catalog_path = Path(__file__).with_name(CATALOG_FILE)
+    if not catalog_path.is_file():
+        return None
+
+    try:
+        _catalog_cache = RunbookCatalog(**json.loads(catalog_path.read_text()))
+    except (json.JSONDecodeError, ValidationError) as exc:
+        logging.error("Failed to load runbook catalog: %s", exc)
+        _catalog_cache = None
+    return _catalog_cache
holmes/core/runbooks.py (2)

55-60: Inefficient string concatenation in loop

Building a large string with += in a loop is quadratic in time/allocations. Prefer "\n".join(...).

-            combined_runbooks = ""
-            for runbook_str in self.runbooks:
-                combined_runbooks += f"* {runbook_str}\n"
+            combined_runbooks = "\n".join(f"* {rb}" for rb in self.runbooks)

82-86: Unnecessary else after early return

After the if len(runbook_abs_link) == 0: branch returns, the else: block is redundant. Remove it and de-indent its body to reduce nesting.

📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 8e9f9e6 and 632203a.

📒 Files selected for processing (4)
  • holmes/config.py (2 hunks)
  • holmes/core/runbooks.py (2 hunks)
  • holmes/core/tool_calling_llm.py (4 hunks)
  • holmes/plugins/runbooks/__init__.py (3 hunks)
🚧 Files skipped from review as they are similar to previous changes (2)
  • holmes/config.py
  • holmes/core/tool_calling_llm.py
🧰 Additional context used
🪛 Pylint (3.3.7)
holmes/core/runbooks.py

[refactor] 58-58: Consider using str.join(sequence) for concatenating strings from an iterable

(R1713)


[refactor] 82-86: Unnecessary "else" after "return", remove the "else" and de-indent the code inside it

(R1705)


[refactor] 45-45: Too many return statements (7/6)

(R0911)


[refactor] 33-33: Too few public methods (1/2)

(R0903)

holmes/plugins/runbooks/__init__.py

[refactor] 57-57: Too few public methods (0/2)

(R0903)


[refactor] 69-69: Too few public methods (0/2)

(R0903)

⏰ Context from checks skipped due to timeout of 90000ms (7)
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.12)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.12)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.12)

Comment thread holmes/core/runbooks.py
@mainred mainred closed this Jun 20, 2025
arikalon1 pushed a commit that referenced this pull request Jun 23, 2025
…547)

This PR introduces a toolset Runbook to fetch the built-in runbooks, the
catalog are integrated into the system prompt. I closed the PR to let
llm pick the runbook and append it to the initial steps before executing
any tools #533, but it's
not already necessary to follow a runbook and the runbook selected might
be misleading depending only on the initial user prompt.
The result of toolset will not be appended to the user prompt, but as
the other tool result, it will be part of the assistant prompt
RoiGlinik pushed a commit that referenced this pull request May 20, 2026
…#2054)

## Summary

Documents `LITELLM_MODEL_COST_MAP_URL` and Robusta's mirror of LiteLLM's
model catalog (`model_prices_and_context_window.json`).

Customers whose egress firewalls block `raw.githubusercontent.com`
cannot let LiteLLM refresh its model catalog (which determines per-model
context windows, max output tokens, and pricing). The fix is purely
operational — LiteLLM already honors `LITELLM_MODEL_COST_MAP_URL`, and
Robusta now serves a mirror of the file at
`https://api.robusta.dev/litellm/model_prices_and_context_window.json`
with TTL caching and a stale fallback. Setting the env var via
`additionalEnvVars` in Helm is all it takes.

For fully self-hosted Robusta installs where the relay itself also
cannot reach GitHub, the relay's `LITELLM_MODEL_COST_MAP_UPSTREAM_URL`
can be pointed at Robusta's mirror to chain the lookup — documented
inline.

Relay-side endpoint:
[robusta-dev/relay#533](robusta-dev/relay#533)
(ROB-3898).

## Test plan

- [ ] Render `docs/reference/environment-variables.md` locally and
confirm the new section renders correctly
- [ ] Verify the linked relay endpoint returns valid JSON once #533 is
merged and deployed
- [ ] Confirm
`LITELLM_MODEL_COST_MAP_URL=https://api.robusta.dev/litellm/model_prices_and_context_window.json`
works end-to-end from a HolmesGPT pod that cannot reach
`raw.githubusercontent.com`


---
_Generated by [Claude
Code](https://claude.ai/code/session_01PmQBah9A7u3u4zDbyjmJ1C)_

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added guidance for a new configuration variable to override the
default LiteLLM model cost map URL.
* Documented using an alternative mirror (with caching/fallback
behavior) when direct downloads are restricted.
* Included a Helm example showing how to set the configuration for
deployments.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/HolmesGPT/holmesgpt/pull/2054?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Pick the right runbooks based on prompt (planning phase)

2 participants