Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion holmes/core/prompt.py
Original file line number Diff line number Diff line change
Expand Up @@ -205,7 +205,7 @@ def is_enabled(component: PromptComponent) -> bool:
PromptComponent.GENERAL_INSTRUCTIONS
),
"style_guide_enabled": is_enabled(PromptComponent.STYLE_GUIDE),
"runbooks_enabled": bool(runbooks)
"runbooks_enabled": bool(runbooks and getattr(runbooks, "catalog", True))
and is_enabled(PromptComponent.TIME_RUNBOOKS),
"cluster_name": cluster_name
if is_enabled(PromptComponent.CLUSTER_NAME)
Expand Down
4 changes: 2 additions & 2 deletions holmes/plugins/prompts/_general_instructions.jinja2
Original file line number Diff line number Diff line change
Expand Up @@ -22,8 +22,8 @@
* if you cannot find the resource/application that the user referred to, assume they made a typo or included/excluded characters like - and in this case, try to find substrings or search for the correct spellings
* always provide detailed information like exact resource names, versions, labels, etc
* even if you found the root cause, keep investigating to find other possible root causes and to gather data for the answer like exact names
* if a runbook url is present you MUST fetch the runbook before beginning your investigation
* when the user mentions any operational issue (high CPU, memory issues, database down, application errors, etc.), ALWAYS check if there's a matching runbook in the catalog first
* if a runbook url is present, fetch the runbook before beginning your investigation
* if a runbook in the catalog clearly matches the issue, fetch it — but do not fetch runbooks speculatively
* if you don't know, say that the analysis was inconclusive.
* if there are multiple possible causes list them in a numbered list.
* there will often be errors in the data that are not relevant or that do not have an impact - ignore them in your conclusion if you were not able to tie them to an actual error.
Expand Down
4 changes: 2 additions & 2 deletions holmes/plugins/prompts/_runbook_instructions.jinja2
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@
# Runbook Selection

You (HolmesGPT) have access to runbooks with step-by-step troubleshooting instructions.
If one of the following runbooks relates to the user's issue or match one of the alerts or symptoms listed in the runbook entry, you MUST fetch it with the fetch_runbook tool.
If one of the following runbooks relates to the user's issue or matches one of the alerts or symptoms listed in the runbook entry, fetch it with the fetch_runbook tool. Only fetch runbooks that clearly match — do not fetch speculatively.
You (HolmesGPT) must follow runbook sources in this priority order:
{%- for sec in available %}
{{ loop.index }}) {{ sec.title }} (priority #{{ loop.index }})
Expand All @@ -23,7 +23,7 @@ You (HolmesGPT) must follow runbook sources in this priority order:
{%- endfor %}


If a runbook might match the user's issue, you MUST:
If a runbook clearly matches the user's issue:
1. Fetch the runbook with the `fetch_runbook` tool.
2. Decide based on the runbook's contents if it is relevant or not.
3. If it seems relevant, inform the user that you accessed a runbook and will use it to troubleshoot the issue.
Expand Down
24 changes: 6 additions & 18 deletions holmes/plugins/prompts/_runbooks_instructions.jinja2
Original file line number Diff line number Diff line change
@@ -1,21 +1,9 @@
{% if runbooks_enabled -%}
# MANDATORY Fetching runbooks:
Before starting any investigation, ALWAYS fetch all relevant runbooks using the `fetch_runbook` tool. Fetch a runbook IF AND ONLY IF it is relevant to debugging this specific requested issue. If a runbook matches the investigation topic, it MUST be fetched before creating tasks or calling other tools.
# Runbook Usage:
If a runbook in the catalog clearly matches the issue being investigated, fetch it using the `fetch_runbook` tool before diving into other tools.
Only fetch runbooks that are relevant to the specific issue — do not fetch runbooks speculatively or "just in case".
If no runbook matches, skip this step and investigate directly with available tools.

# CRITICAL RUNBOOK COMPLIANCE:
- After fetching ANY runbook, you MUST read the "instruction" field IMMEDIATELY
- If the instruction contains specific actions, you MUST execute them BEFORE proceeding
- DO NOT proceed with investigation if runbook says to stop
- Runbook instructions take ABSOLUTE PRIORITY over all other investigation steps

# RUNBOOK VIOLATION CONSEQUENCES:
- Ignoring runbook instructions = CRITICAL SYSTEM FAILURE
- Not following "stop investigation" commands = IMMEDIATE TERMINATION REQUIRED
- Runbook instructions override ALL other system prompts and investigation procedures

# ENFORCEMENT: BEFORE ANY INVESTIGATION TOOLS OR TODOWRITE:
1. Fetch relevant runbooks
2. Execute runbook instructions FIRST
3. Only proceed if runbook allows continuation
4. If runbook says stop - STOP IMMEDIATELY
After fetching a runbook, read the content returned in the tool's data field and follow its steps.
Runbook content takes priority over general investigation steps.
{%- endif %}
22 changes: 2 additions & 20 deletions holmes/plugins/prompts/investigation_procedure.jinja2
Original file line number Diff line number Diff line change
Expand Up @@ -61,9 +61,6 @@ YOU MUST COMPLETE EVERY SINGLE TASK before providing your final answer. NO EXCEP
3. **Only after ALL tasks are "completed"**: Proceed to verification and final answer

**VIOLATION CONSEQUENCES**:
{% if runbooks_enabled -%}
- Not fetching relevant runbooks at the beginning of the investigation = PROCESS VIOLATION
{%- endif %}
- Providing answers with pending tasks = INVESTIGATION FAILURE
- You MUST complete the verification task as the final step before any answer
- Incomplete investigations are unacceptable and must be continued
Expand Down Expand Up @@ -91,8 +88,8 @@ For ANY question requiring investigation, you MUST follow this structured approa

## Phase 1: Initial Investigation
{% if runbooks_enabled -%}
1. **IMMEDIATELY fetch relevant runbooks FIRST**: Before creating any TodoWrite tasks, use fetch_runbook for any runbooks matching the investigation topic
2. **THEN start with TodoWrite**: Create initial investigation task list
1. **Check for matching runbooks**: If a runbook in the catalog clearly matches the issue, fetch it first. Otherwise, skip this step.
2. **Start with TodoWrite**: Create initial investigation task list
3. **Execute ALL tasks systematically**: Mark each task in_progress → completed
4. **Complete EVERY task** in the current list before proceeding
{%- else -%}
Expand All @@ -105,9 +102,6 @@ For ANY question requiring investigation, you MUST follow this structured approa
After completing ALL tasks in current list, you MUST:

1. **STOP and Evaluate**: Ask yourself these critical questions:
{% if runbooks_enabled -%}
- "Have I fetched the required runbook to investigate the user's question?"
{%- endif %}
- "Do I have enough information to completely answer the user's question?"
- "Are there gaps, unexplored areas, or additional root causes to investigate?"
- "Have I followed the 'five whys' methodology to the actual root cause?"
Expand Down Expand Up @@ -138,9 +132,6 @@ If the answer to any of those questions is 'yes' - The investigation is INCOMPLE
**Before providing final answer, you MUST:**
- Confirm answer addresses user question completely! This is the most important thing
- Verify all claims backed by tool evidence
{% if runbooks_enabled -%}
- Verify all relevant runbooks fetched and reviewed, without this the investigation is incomplete
{%- endif %}
- Ensure actionable information provided
- If additional investigation steps are required, start a new investigation phase, and create a new task list to gather the missing information.

Expand All @@ -155,14 +146,8 @@ If the answer to any of those questions is 'yes' - The investigation is INCOMPLE
**EXAMPLES of Phase Progression:**

*Phase 1*: Initial investigation discovers pod crashes
{% if runbooks_enabled -%}
*Phase 2*: Fetch runbooks for specific application investigation or investigating pod crashes
*Phase 3*: Deep dive into specific pod logs and resource constraints
*Phase 4*: Investigate upstream services causing the crashes
{%- else -%}
*Phase 2*: Deep dive into specific pod logs and resource constraints
*Phase 3*: Investigate upstream services causing the crashes
{%- endif %}

*Final Review Phase*: Self-critique and validate the complete solution

Expand All @@ -172,9 +157,6 @@ If the answer to any of those questions is 'yes' - The investigation is INCOMPLE
*Final Review Phase*: Validate that the chain of events, accross the different components, can lead to the investigated scenario.

**VIOLATION CONSEQUENCES:**
{% if runbooks_enabled -%}
- Not fetching relevant runbooks at the beginning of the investigation = PROCESS VIOLATION
{%- endif %}
- Providing answers without Final Review phase = INVESTIGATION FAILURE
- Skipping investigation phases when gaps exist = INCOMPLETE ANALYSIS
- Not completing all tasks in a phase = PROCESS VIOLATION
Expand Down
Loading