Repository navigation
Conversation
WalkthroughThis update introduces a runbook catalog system for HolmesGPT, enabling dynamic selection and retrieval of operational runbooks based on user queries. It adds a catalog JSON, new runbook management classes, CLI/configuration support for custom runbooks, and documentation. The tool-calling LLM now integrates runbook guidance into its workflow. Changes
Sequence Diagram(s)sequenceDiagram
participant User
participant CLI/Config
participant ToolCallingLLM
participant RunbookCatalogManager
participant LLM
User->>CLI/Config: Provide query (optionally with custom runbooks)
CLI/Config->>ToolCallingLLM: Initialize with runbook catalog manager
ToolCallingLLM->>RunbookCatalogManager: get_runbook_by_question(user_question)
alt Custom runbooks provided
RunbookCatalogManager-->>ToolCallingLLM: Return combined custom runbooks
else No custom runbooks
RunbookCatalogManager->>LLM: Query for matching runbook link
LLM-->>RunbookCatalogManager: Return runbook link
RunbookCatalogManager->>RunbookCatalogManager: Load runbook file content
RunbookCatalogManager-->>ToolCallingLLM: Return runbook content and link
end
ToolCallingLLM->>ToolCallingLLM: Append runbook to user prompt
ToolCallingLLM->>User: Return LLM response with runbook guidance
Assessment against linked issues
Assessment against linked issues: Out-of-scope changesNo out-of-scope changes detected. Suggested reviewers
✨ Finishing Touches
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. 🪧 TipsChatThere are 3 ways to chat with CodeRabbit:
SupportNeed help? Create a ticket on our support page for assistance with any issues or questions. Note: Be mindful of the bot's finite context window. It's strongly recommended to break down tasks such as reading entire modules into smaller chunks. For a focused discussion, use review comments to chat about specific files and their changes, instead of using the PR comments. CodeRabbit Commands (Invoked using PR comments)
Other keywords and placeholders
CodeRabbit Configuration File (
|
There was a problem hiding this comment.
Actionable comments posted: 4
🧹 Nitpick comments (11)
holmes/plugins/runbooks/README.md (2)
1-3: Minor grammar – add missing article“The Runbooks folder …” reads better with the definite article:
-Runbooks folder contains operational runbooks … +The runbooks folder contains operational runbooks …
21-23: Proper noun & sentence clarity
- “markdown” → “Markdown” (proper noun).
- Slight wording tweak for clarity:
-Catalog specified in [catalog.json](catalog.json) contains a collection of runbooks written in markdown. +The catalog defined in [catalog.json](catalog.json) contains Markdown runbooks.holmes/plugins/runbooks/__init__.py (1)
78-88: Clean-up: shadowing built-ins and variable naming
catalogPath→catalog_pathfor PEP-8.- The loop variable
fileshadows the built-infile.- catalogPath = os.path.join(dir_path, CATALOG_FILE) - if not os.path.isfile(catalogPath): + catalog_path = os.path.join(dir_path, CATALOG_FILE) + if not os.path.isfile(catalog_path): return None - - with open(catalogPath) as file: - catalog_dict = json.load(file) + with open(catalog_path, "r") as fh: + catalog_dict = json.load(fh)holmes/main.py (1)
155-160: Option help text: plural vs singularThe help string says “Path to a custom runbooks”.
Grammar is off and might confuse users:-help="Path to a custom runbooks (can specify -r multiple times to add multiple runbooks)", +help="Path(s) to custom runbook files (use -r multiple times)",holmes/plugins/runbooks/networking/dns_troubleshooting_instructions.md (2)
8-10: Typo & wording“follow the troubleshoot guide” → “follow the troubleshooting guide”.
-* Instead of provide next steps to the user, you need to follow the troubleshoot guide to execute the steps. +* Instead of providing next steps to the user, follow the troubleshooting guide to execute the steps.
56-67: Markdown lint – list indentation & bare URLsIndentation levels break MD007 and bare URLs break MD034.
Example fix for one block (apply similarly to the rest):-* **CRITICAL:** ALWAYS refer to the official Kubernetes DNS debugging guide for detailed troubleshooting and solutions: - * Main guide: https://kubernetes.io/docs/tasks/administer-cluster/dns-debugging-resolution/ - * CoreDNS specific: https://kubernetes.io/docs/tasks/administer-cluster/dns-custom-nameservers/ (for CoreDNS customization which might be relevant) +* **CRITICAL:** ALWAYS refer to the official Kubernetes DNS debugging guide: + * Main guide: <https://kubernetes.io/docs/tasks/administer-cluster/dns-debugging-resolution/> + * CoreDNS specific: <https://kubernetes.io/docs/tasks/administer-cluster/dns-custom-nameservers/> (for CoreDNS customization)This resolves both MD007 (indent) and MD034 (wrap URLs in
<>).holmes/config.py (1)
269-276: Reuse the same LLM instance to avoid double initialisation
self._get_llm()is called twice – once for the catalog manager and again forToolCallingLLM.
Instantiate once and reuse; this avoids duplicate network/session setup and makes mocking easier in tests.holmes/core/runbooks.py (2)
56-60: Inefficient string build and missing separator
combined_runbooksis built via+=inside a loop.
Use"\n".join(...)– faster and clearer:- combined_runbooks = "" - for runbook_str in self.runbooks: - combined_runbooks += f"* {runbook_str}\n" + combined_runbooks = "\n".join(f"* {rb}" for rb in self.runbooks)
82-86: Unnecessaryelseafter early-returnAfter returning for the empty-string case, the
else:block is redundant – de-indent its body for cleaner flow.holmes/core/tool_calling_llm.py (2)
118-129: Constructor: propagaterunbook_catalog_managerto subclasses
IssueInvestigatorcallssuper().__init__(tool_executor, max_steps, llm)and therefore loses the catalog manager reference. Consider giving it a default ofNoneand letting derived classes pass it when relevant, or document the limitation explicitly.
745-747: Parameter naming – this builds the prompt but reads oddly
add_runbook_to_user_prompt(user_prompt, runbook)actually expects the question then the runbook. Rename first parameter toquestionor swap the arguments for clarity.
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (8)
holmes/config.py(2 hunks)holmes/core/runbooks.py(2 hunks)holmes/core/tool_calling_llm.py(6 hunks)holmes/main.py(4 hunks)holmes/plugins/runbooks/README.md(1 hunks)holmes/plugins/runbooks/__init__.py(3 hunks)holmes/plugins/runbooks/catalog.json(1 hunks)holmes/plugins/runbooks/networking/dns_troubleshooting_instructions.md(1 hunks)
🧰 Additional context used
🧬 Code Graph Analysis (2)
holmes/main.py (2)
holmes/core/prompt.py (1)
append_file_to_user_prompt(4-8)holmes/config.py (1)
create_console_toolcalling_llm(259-276)
holmes/core/runbooks.py (3)
holmes/core/issue.py (1)
Issue(13-54)holmes/core/llm.py (1)
LLM(30-58)holmes/plugins/runbooks/__init__.py (3)
Runbook(23-33)get_runbook_folder(91-92)load_catalog(78-88)
🪛 GitHub Actions: Build and test HolmesGPT
holmes/plugins/runbooks/catalog.json
[error] 11-15: End-of-file-fixer hook fixed missing newline at end of file.
holmes/core/tool_calling_llm.py
[error] 218-226: Ruff-format hook reformatted code due to line break and parentheses style changes.
[error] 579-587: Ruff-format hook reformatted code due to line break and parentheses style changes.
🪛 LanguageTool
holmes/plugins/runbooks/README.md
[uncategorized] ~2-~2: You might be missing the article “the” here.
Context: # Runbooks Runbooks folder contains operational runbooks fo...
(AI_EN_LECTOR_MISSING_DETERMINER_THE)
[uncategorized] ~8-~8: Possible missing preposition found.
Context: ... - Standardize operational processes - Enable quick onboarding for new team members -...
(AI_HYDRA_LEO_MISSING_TO)
[grammar] ~21-~21: Did you mean the formatting language “Markdown” (= proper noun)?
Context: ...ins a collection of runbooks written in markdown. During runtime, LLM will compare the r...
(MARKDOWN_NNP)
holmes/plugins/runbooks/networking/dns_troubleshooting_instructions.md
[grammar] ~8-~8: The word ‘troubleshoot’ is a verb. Did you mean the noun “troubleshooting” or “troubleshooting guide”?
Context: ...eps to the user, you need to follow the troubleshoot guide to execute the steps. * When ge...
(PREPOSITION_VERB)
[style] ~64-~64: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: ...od dnsPolicy and dnsConfig. * If NetworkPolicies are suspected, suggest ...
(ENGLISH_WORD_REPEAT_BEGINNING_RULE)
[style] ~65-~65: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: ...olicy definitions to allow DNS. * If CoreDNS configuration seems problematic...
(ENGLISH_WORD_REPEAT_BEGINNING_RULE)
[style] ~66-~66: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: ...rnetes guide on customizing it. * If upstream DNS resolution is failing, sug...
(ENGLISH_WORD_REPEAT_BEGINNING_RULE)
🪛 markdownlint-cli2 (0.17.2)
holmes/plugins/runbooks/networking/dns_troubleshooting_instructions.md
58-58: Unordered list indentation
Expected: 2; Actual: 4
(MD007, ul-indent)
58-58: Bare URL used
null
(MD034, no-bare-urls)
59-59: Unordered list indentation
Expected: 2; Actual: 4
(MD007, ul-indent)
59-59: Bare URL used
null
(MD034, no-bare-urls)
62-62: Unordered list indentation
Expected: 2; Actual: 4
(MD007, ul-indent)
63-63: Unordered list indentation
Expected: 2; Actual: 4
(MD007, ul-indent)
64-64: Unordered list indentation
Expected: 2; Actual: 4
(MD007, ul-indent)
65-65: Unordered list indentation
Expected: 2; Actual: 4
(MD007, ul-indent)
66-66: Unordered list indentation
Expected: 2; Actual: 4
(MD007, ul-indent)
🪛 Pylint (3.3.7)
holmes/plugins/runbooks/__init__.py
[refactor] 57-57: Too few public methods (0/2)
(R0903)
[refactor] 69-69: Too few public methods (0/2)
(R0903)
holmes/core/tool_calling_llm.py
[error] 178-178: Possibly using variable 'user_question' before assignment
(E0606)
[refactor] 221-229: Unnecessary "else" after "raise", remove the "else" and de-indent the code inside it
(R1720)
holmes/core/runbooks.py
[refactor] 58-58: Consider using str.join(sequence) for concatenating strings from an iterable
(R1713)
[refactor] 82-86: Unnecessary "else" after "return", remove the "else" and de-indent the code inside it
(R1705)
[refactor] 45-45: Too many return statements (7/6)
(R0911)
[refactor] 33-33: Too few public methods (1/2)
(R0903)
⏰ Context from checks skipped due to timeout of 90000ms (1)
- GitHub Check: build (3.12)
🔇 Additional comments (1)
holmes/main.py (1)
336-337: Pass-through argument name drift
custom_runbooksfrom CLI is forwarded ascustom_runbooks(snake_case) – good.
VerifyConfig.load_from_fileindeed expects the same kwarg; mismatched names silently drop data.
If it expectscustom_runbooks_from_cli, adapt here.
There was a problem hiding this comment.
Actionable comments posted: 0
♻️ Duplicate comments (1)
holmes/core/tool_calling_llm.py (1)
172-183: Runbook never reaches the LLM + possibleNameError
user_questionis not initialised before the loop – if themessageslist happens to contain no"user"role message, theNameErrorraised atget_runbook_by_question(user_question)will crash the call path.Even when a runbook is found, only the local
user_promptvariable is updated; the originalmessageslist that is sent to the LLM remains unchanged. Consequently, the LLM never receives the runbook.- for message in messages: - if message.get("role") == "user": - user_question = message.get("content", "") + # Grab the *first* user message and guard against missing ones + user_question = "" + for message in messages: + if message.get("role") == "user" and not user_question: + user_question = message.get("content", "") ... - user_prompt = add_runbook_to_user_prompt(user_question, runbook) # type: ignore + user_prompt = add_runbook_to_user_prompt(user_question, runbook) + # Inject the augmented prompt back into the message list + for msg in messages: + if msg.get("role") == "user": + msg["content"] = user_prompt + breakThis diff both prevents the potential
NameErrorand ensures the runbook content is actually forwarded to the model.
🧹 Nitpick comments (1)
holmes/core/tool_calling_llm.py (1)
743-745: Misleading parameter name & missing fenced formatting
add_runbook_to_user_promptreceives what you calluser_prompt, but upstream you passuser_question. Consider renaming the first parameter to clarify intent and wrap the runbook in triple back-ticks (or blockquote) to avoid Markdown bleed when the runbook itself contains headings.-def add_runbook_to_user_prompt(user_prompt: Optional[str], runbook: str) -> str: - return f"My instructions to check '{user_prompt}' by following the runbook:\n {runbook}" +def add_runbook_to_user_prompt(user_question: str, runbook: str) -> str: + return ( + f"My instructions to answer **{user_question}** by following the runbook below:\n" + f"```markdown\n{runbook}\n```" + )
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (2)
holmes/core/tool_calling_llm.py(4 hunks)holmes/plugins/runbooks/catalog.json(1 hunks)
🚧 Files skipped from review as they are similar to previous changes (1)
- holmes/plugins/runbooks/catalog.json
🧰 Additional context used
🪛 Pylint (3.3.7)
holmes/core/tool_calling_llm.py
[error] 178-178: Possibly using variable 'user_question' before assignment
(E0606)
⏰ Context from checks skipped due to timeout of 90000ms (5)
- GitHub Check: build (3.11)
- GitHub Check: build (3.12)
- GitHub Check: build (3.11)
- GitHub Check: build (3.10)
- GitHub Check: build (3.12)
🔇 Additional comments (1)
holmes/core/tool_calling_llm.py (1)
118-129: Constructor change silently breaksIssueInvestigator
ToolCallingLLM.__init__now requires arunbook_catalog_managerpositional argument, butIssueInvestigatorstill callssuper().__init__(tool_executor, max_steps, llm)(line 640). Relying on the defaultNonevalue masks this behavioural change and can bite callers that forget to pass the extra argument.Either promote the new arg to a keyword-only param or update downstream constructors explicitly:
- super().__init__(tool_executor, max_steps, llm) + super().__init__( + tool_executor=tool_executor, + max_steps=max_steps, + llm=llm, + runbook_catalog_manager=None, # or inject the real one + )
There was a problem hiding this comment.
Actionable comments posted: 1
♻️ Duplicate comments (1)
holmes/plugins/runbooks/__init__.py (1)
63-66: Field names break Pydantic & Python naming conventions
Update_DateandKeyWordsuse Pascal-/Camel-Case while the rest of the codebase (and Python in general) followssnake_case. This has already been raised in a prior review and is still unresolved. Adoptsnake_casefield names and usealias="Update_Date"/alias="KeyWords"to preserve compatibility with the JSON.-class RunbookCatalogEntry(BaseModel): +class RunbookCatalogEntry(BaseModel): @@ - Update_Date: date - Description: str - KeyWords: list[str] - link: str + update_date: date = Field(alias="Update_Date") + description: str = Field(alias="Description") + keywords: list[str] = Field(alias="KeyWords") + link: str
🧹 Nitpick comments (3)
holmes/plugins/runbooks/__init__.py (1)
78-88:load_catalog()lacks error handling & memoisation
- A malformed
catalog.json(e.g. JSON syntax error or Pydantic validation failure) will raise and crash the caller.- The catalog is parsed every time
load_catalog()is invoked – unnecessary I/O on hot paths.-_catalog_cache: Optional[RunbookCatalog] = None - def load_catalog() -> Optional[RunbookCatalog]: - dir_path = os.path.dirname(os.path.realpath(__file__)) - catalogPath = os.path.join(dir_path, CATALOG_FILE) - if not os.path.isfile(catalogPath): - return None - - with open(catalogPath) as file: - catalog_dict = json.load(file) - return RunbookCatalog(**catalog_dict) - return None + global _catalog_cache + if _catalog_cache is not None: + return _catalog_cache + + catalog_path = Path(__file__).with_name(CATALOG_FILE) + if not catalog_path.is_file(): + return None + + try: + _catalog_cache = RunbookCatalog(**json.loads(catalog_path.read_text())) + except (json.JSONDecodeError, ValidationError) as exc: + logging.error("Failed to load runbook catalog: %s", exc) + _catalog_cache = None + return _catalog_cacheholmes/core/runbooks.py (2)
55-60: Inefficient string concatenation in loopBuilding a large string with
+=in a loop is quadratic in time/allocations. Prefer"\n".join(...).- combined_runbooks = "" - for runbook_str in self.runbooks: - combined_runbooks += f"* {runbook_str}\n" + combined_runbooks = "\n".join(f"* {rb}" for rb in self.runbooks)
82-86: Unnecessaryelseafter early returnAfter the
if len(runbook_abs_link) == 0:branch returns, theelse:block is redundant. Remove it and de-indent its body to reduce nesting.
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (4)
holmes/config.py(2 hunks)holmes/core/runbooks.py(2 hunks)holmes/core/tool_calling_llm.py(4 hunks)holmes/plugins/runbooks/__init__.py(3 hunks)
🚧 Files skipped from review as they are similar to previous changes (2)
- holmes/config.py
- holmes/core/tool_calling_llm.py
🧰 Additional context used
🪛 Pylint (3.3.7)
holmes/core/runbooks.py
[refactor] 58-58: Consider using str.join(sequence) for concatenating strings from an iterable
(R1713)
[refactor] 82-86: Unnecessary "else" after "return", remove the "else" and de-indent the code inside it
(R1705)
[refactor] 45-45: Too many return statements (7/6)
(R0911)
[refactor] 33-33: Too few public methods (1/2)
(R0903)
holmes/plugins/runbooks/__init__.py
[refactor] 57-57: Too few public methods (0/2)
(R0903)
[refactor] 69-69: Too few public methods (0/2)
(R0903)
⏰ Context from checks skipped due to timeout of 90000ms (7)
- GitHub Check: build (3.11)
- GitHub Check: build (3.12)
- GitHub Check: build (3.10)
- GitHub Check: build (3.12)
- GitHub Check: build (3.10)
- GitHub Check: build (3.11)
- GitHub Check: build (3.12)
…547) This PR introduces a toolset Runbook to fetch the built-in runbooks, the catalog are integrated into the system prompt. I closed the PR to let llm pick the runbook and append it to the initial steps before executing any tools #533, but it's not already necessary to follow a runbook and the runbook selected might be misleading depending only on the initial user prompt. The result of toolset will not be appended to the user prompt, but as the other tool result, it will be part of the assistant prompt
…#2054) ## Summary Documents `LITELLM_MODEL_COST_MAP_URL` and Robusta's mirror of LiteLLM's model catalog (`model_prices_and_context_window.json`). Customers whose egress firewalls block `raw.githubusercontent.com` cannot let LiteLLM refresh its model catalog (which determines per-model context windows, max output tokens, and pricing). The fix is purely operational — LiteLLM already honors `LITELLM_MODEL_COST_MAP_URL`, and Robusta now serves a mirror of the file at `https://api.robusta.dev/litellm/model_prices_and_context_window.json` with TTL caching and a stale fallback. Setting the env var via `additionalEnvVars` in Helm is all it takes. For fully self-hosted Robusta installs where the relay itself also cannot reach GitHub, the relay's `LITELLM_MODEL_COST_MAP_UPSTREAM_URL` can be pointed at Robusta's mirror to chain the lookup — documented inline. Relay-side endpoint: [robusta-dev/relay#533](robusta-dev/relay#533) (ROB-3898). ## Test plan - [ ] Render `docs/reference/environment-variables.md` locally and confirm the new section renders correctly - [ ] Verify the linked relay endpoint returns valid JSON once #533 is merged and deployed - [ ] Confirm `LITELLM_MODEL_COST_MAP_URL=https://api.robusta.dev/litellm/model_prices_and_context_window.json` works end-to-end from a HolmesGPT pod that cannot reach `raw.githubusercontent.com` --- _Generated by [Claude Code](https://claude.ai/code/session_01PmQBah9A7u3u4zDbyjmJ1C)_ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added guidance for a new configuration variable to override the default LiteLLM model cost map URL. * Documented using an alternative mirror (with caching/fallback behavior) when direct downloads are restricted. * Included a Helm example showing how to set the configuration for deployments. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/HolmesGPT/holmesgpt/pull/2054?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Claude <noreply@anthropic.com> Co-authored-by: Claude <noreply@anthropic.com>
fix: #473
Changes in this PR:
holmes askimpacted function
only
holmes askCLI is expected to be impactedWhy use LLM to pick the runbook
The original thought is to introduce a runbook tool to fetch runbook from the runbook catalog folder, but considering we'll have more runbooks to be introduced and selecting the runbook by the either label or tag is limited, introducing LLM to pick the matched one should be the final solution. Also, it's best to select the runbook after customer specify their question, and the following actions should be based on runbook. With all these in mind, I decided to use LLM to pick the runbook during the early phase of the investigation. This is similar to we have achieved in
holmes investigateTests