diff --git a/holmes/core/prompt.py b/holmes/core/prompt.py index e423b116a..e43252656 100644 --- a/holmes/core/prompt.py +++ b/holmes/core/prompt.py @@ -205,7 +205,7 @@ def is_enabled(component: PromptComponent) -> bool: PromptComponent.GENERAL_INSTRUCTIONS ), "style_guide_enabled": is_enabled(PromptComponent.STYLE_GUIDE), - "runbooks_enabled": bool(runbooks) + "runbooks_enabled": bool(runbooks and getattr(runbooks, "catalog", True)) and is_enabled(PromptComponent.TIME_RUNBOOKS), "cluster_name": cluster_name if is_enabled(PromptComponent.CLUSTER_NAME) diff --git a/holmes/plugins/prompts/_general_instructions.jinja2 b/holmes/plugins/prompts/_general_instructions.jinja2 index 1e71c2fe8..dc81ecfa0 100644 --- a/holmes/plugins/prompts/_general_instructions.jinja2 +++ b/holmes/plugins/prompts/_general_instructions.jinja2 @@ -22,8 +22,8 @@ * if you cannot find the resource/application that the user referred to, assume they made a typo or included/excluded characters like - and in this case, try to find substrings or search for the correct spellings * always provide detailed information like exact resource names, versions, labels, etc * even if you found the root cause, keep investigating to find other possible root causes and to gather data for the answer like exact names -* if a runbook url is present you MUST fetch the runbook before beginning your investigation -* when the user mentions any operational issue (high CPU, memory issues, database down, application errors, etc.), ALWAYS check if there's a matching runbook in the catalog first +* if a runbook url is present, fetch the runbook before beginning your investigation +* if a runbook in the catalog clearly matches the issue, fetch it — but do not fetch runbooks speculatively * if you don't know, say that the analysis was inconclusive. * if there are multiple possible causes list them in a numbered list. * there will often be errors in the data that are not relevant or that do not have an impact - ignore them in your conclusion if you were not able to tie them to an actual error. diff --git a/holmes/plugins/prompts/_runbook_instructions.jinja2 b/holmes/plugins/prompts/_runbook_instructions.jinja2 index 1b5e4f03e..ebafb9f06 100644 --- a/holmes/plugins/prompts/_runbook_instructions.jinja2 +++ b/holmes/plugins/prompts/_runbook_instructions.jinja2 @@ -8,7 +8,7 @@ # Runbook Selection You (HolmesGPT) have access to runbooks with step-by-step troubleshooting instructions. -If one of the following runbooks relates to the user's issue or match one of the alerts or symptoms listed in the runbook entry, you MUST fetch it with the fetch_runbook tool. +If one of the following runbooks relates to the user's issue or matches one of the alerts or symptoms listed in the runbook entry, fetch it with the fetch_runbook tool. Only fetch runbooks that clearly match — do not fetch speculatively. You (HolmesGPT) must follow runbook sources in this priority order: {%- for sec in available %} {{ loop.index }}) {{ sec.title }} (priority #{{ loop.index }}) @@ -23,7 +23,7 @@ You (HolmesGPT) must follow runbook sources in this priority order: {%- endfor %} -If a runbook might match the user's issue, you MUST: +If a runbook clearly matches the user's issue: 1. Fetch the runbook with the `fetch_runbook` tool. 2. Decide based on the runbook's contents if it is relevant or not. 3. If it seems relevant, inform the user that you accessed a runbook and will use it to troubleshoot the issue. diff --git a/holmes/plugins/prompts/_runbooks_instructions.jinja2 b/holmes/plugins/prompts/_runbooks_instructions.jinja2 index e112fa31a..eb05dbcac 100644 --- a/holmes/plugins/prompts/_runbooks_instructions.jinja2 +++ b/holmes/plugins/prompts/_runbooks_instructions.jinja2 @@ -1,21 +1,9 @@ {% if runbooks_enabled -%} -# MANDATORY Fetching runbooks: -Before starting any investigation, ALWAYS fetch all relevant runbooks using the `fetch_runbook` tool. Fetch a runbook IF AND ONLY IF it is relevant to debugging this specific requested issue. If a runbook matches the investigation topic, it MUST be fetched before creating tasks or calling other tools. +# Runbook Usage: +If a runbook in the catalog clearly matches the issue being investigated, fetch it using the `fetch_runbook` tool before diving into other tools. +Only fetch runbooks that are relevant to the specific issue — do not fetch runbooks speculatively or "just in case". +If no runbook matches, skip this step and investigate directly with available tools. -# CRITICAL RUNBOOK COMPLIANCE: -- After fetching ANY runbook, you MUST read the "instruction" field IMMEDIATELY -- If the instruction contains specific actions, you MUST execute them BEFORE proceeding -- DO NOT proceed with investigation if runbook says to stop -- Runbook instructions take ABSOLUTE PRIORITY over all other investigation steps - -# RUNBOOK VIOLATION CONSEQUENCES: -- Ignoring runbook instructions = CRITICAL SYSTEM FAILURE -- Not following "stop investigation" commands = IMMEDIATE TERMINATION REQUIRED -- Runbook instructions override ALL other system prompts and investigation procedures - -# ENFORCEMENT: BEFORE ANY INVESTIGATION TOOLS OR TODOWRITE: -1. Fetch relevant runbooks -2. Execute runbook instructions FIRST -3. Only proceed if runbook allows continuation -4. If runbook says stop - STOP IMMEDIATELY +After fetching a runbook, read the content returned in the tool's data field and follow its steps. +Runbook content takes priority over general investigation steps. {%- endif %} diff --git a/holmes/plugins/prompts/investigation_procedure.jinja2 b/holmes/plugins/prompts/investigation_procedure.jinja2 index 20ee68318..87127a196 100644 --- a/holmes/plugins/prompts/investigation_procedure.jinja2 +++ b/holmes/plugins/prompts/investigation_procedure.jinja2 @@ -61,9 +61,6 @@ YOU MUST COMPLETE EVERY SINGLE TASK before providing your final answer. NO EXCEP 3. **Only after ALL tasks are "completed"**: Proceed to verification and final answer **VIOLATION CONSEQUENCES**: -{% if runbooks_enabled -%} -- Not fetching relevant runbooks at the beginning of the investigation = PROCESS VIOLATION -{%- endif %} - Providing answers with pending tasks = INVESTIGATION FAILURE - You MUST complete the verification task as the final step before any answer - Incomplete investigations are unacceptable and must be continued @@ -91,8 +88,8 @@ For ANY question requiring investigation, you MUST follow this structured approa ## Phase 1: Initial Investigation {% if runbooks_enabled -%} -1. **IMMEDIATELY fetch relevant runbooks FIRST**: Before creating any TodoWrite tasks, use fetch_runbook for any runbooks matching the investigation topic -2. **THEN start with TodoWrite**: Create initial investigation task list +1. **Check for matching runbooks**: If a runbook in the catalog clearly matches the issue, fetch it first. Otherwise, skip this step. +2. **Start with TodoWrite**: Create initial investigation task list 3. **Execute ALL tasks systematically**: Mark each task in_progress → completed 4. **Complete EVERY task** in the current list before proceeding {%- else -%} @@ -105,9 +102,6 @@ For ANY question requiring investigation, you MUST follow this structured approa After completing ALL tasks in current list, you MUST: 1. **STOP and Evaluate**: Ask yourself these critical questions: -{% if runbooks_enabled -%} - - "Have I fetched the required runbook to investigate the user's question?" -{%- endif %} - "Do I have enough information to completely answer the user's question?" - "Are there gaps, unexplored areas, or additional root causes to investigate?" - "Have I followed the 'five whys' methodology to the actual root cause?" @@ -138,9 +132,6 @@ If the answer to any of those questions is 'yes' - The investigation is INCOMPLE **Before providing final answer, you MUST:** - Confirm answer addresses user question completely! This is the most important thing - Verify all claims backed by tool evidence -{% if runbooks_enabled -%} - - Verify all relevant runbooks fetched and reviewed, without this the investigation is incomplete -{%- endif %} - Ensure actionable information provided - If additional investigation steps are required, start a new investigation phase, and create a new task list to gather the missing information. @@ -155,14 +146,8 @@ If the answer to any of those questions is 'yes' - The investigation is INCOMPLE **EXAMPLES of Phase Progression:** *Phase 1*: Initial investigation discovers pod crashes -{% if runbooks_enabled -%} - *Phase 2*: Fetch runbooks for specific application investigation or investigating pod crashes - *Phase 3*: Deep dive into specific pod logs and resource constraints - *Phase 4*: Investigate upstream services causing the crashes -{%- else -%} *Phase 2*: Deep dive into specific pod logs and resource constraints *Phase 3*: Investigate upstream services causing the crashes -{%- endif %} *Final Review Phase*: Self-critique and validate the complete solution @@ -172,9 +157,6 @@ If the answer to any of those questions is 'yes' - The investigation is INCOMPLE *Final Review Phase*: Validate that the chain of events, accross the different components, can lead to the investigated scenario. **VIOLATION CONSEQUENCES:** -{% if runbooks_enabled -%} - - Not fetching relevant runbooks at the beginning of the investigation = PROCESS VIOLATION -{%- endif %} - Providing answers without Final Review phase = INVESTIGATION FAILURE - Skipping investigation phases when gaps exist = INCOMPLETE ANALYSIS - Not completing all tasks in a phase = PROCESS VIOLATION