Soften runbook fetching requirements and reduce enforcement language - #1787
Conversation
The LLM was obsessively fetching runbooks on every investigation because:
- System prompt used catastrophic language ("CRITICAL SYSTEM FAILURE",
"IMMEDIATE TERMINATION") for not fetching runbooks
- Runbook instructions appeared redundantly in both system and user prompts
- investigation_procedure.jinja2 reinforced runbook fetching at every phase
(Phase 1, evaluation, final review, violation consequences)
- runbooks_enabled was True even with empty catalogs
Changes:
- Replace catastrophic/mandatory language with conditional guidance:
"only fetch if clearly matching, skip if no match"
- Remove redundant runbook checks from phase evaluation, final review,
and violation consequences in investigation_procedure.jinja2
- Fix runbooks_enabled check to handle empty catalogs via getattr
- Soften _runbook_instructions.jinja2 (user prompt) from MUST to guidance
- Soften _general_instructions.jinja2 runbook bullets
https://claude.ai/code/session_01ShM8tiPTYRSfpwJrYsrWDN
Signed-off-by: Claude <noreply@anthropic.com>
Claude Code ReviewThis repository is configured for manual code reviews. Comment |
📂 Previous Runs📜 Run @ a7cf377 (#23120785045)✅ Results of HolmesGPT evalsAutomatically triggered by commit a7cf377 on branch Results of HolmesGPT evals
Benchmark comparison unavailable: No eval spans found in experiment 'ci-benchmark-23102181491' Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: No eval spans found in experiment 'ci-benchmark-23102181491' Benchmark experiment:
Comparison indicators:
📜 Run @ c3e50c2 (#23119065188)✅ Results of HolmesGPT evalsAutomatically triggered by commit c3e50c2 on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 73 test/model combinations loaded Benchmark experiment:
No benchmark data available for comparison. Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
📜 Run @ 28ddaf6 (#23119033936)✅ Results of HolmesGPT evalsAutomatically triggered by commit 28ddaf6 on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 73 test/model combinations loaded Benchmark experiment:
No benchmark data available for comparison. Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
✅ Results of HolmesGPT evalsAutomatically triggered by commit 32782a6 on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 73 test/model combinations loaded Benchmark experiment:
No benchmark data available for comparison. Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" Option 3: Add PR labels to include extra evals in automatic regression runs:
Examples: 🏷️ Valid tags
🤖 Valid models
Commands: CLI: |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:52a6ae4d
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:52a6ae4d me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:52a6ae4d
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:52a6ae4d
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:52a6ae4d
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:52a6ae4d me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:52a6ae4d
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:52a6ae4dPatch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:52a6ae4d \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:52a6ae4dRobusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:52a6ae4d \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:52a6ae4d |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
WalkthroughThis change tightens the runbooks_enabled computation in prompt construction and updates multiple prompt templates to make runbook fetching conditional: only fetch runbooks that clearly match the issue and avoid speculative fetching. No function signatures changed. Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. 📝 Coding Plan
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
https://claude.ai/code/session_01ShM8tiPTYRSfpwJrYsrWDN Signed-off-by: Claude <noreply@anthropic.com>
🔬 CLI Performance Benchmark🟡 Startup Time (no LLM)Measures
🟡 Full CLI with LLMMeasures
PR: |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@holmes/plugins/prompts/_runbooks_instructions.jinja2`:
- Around line 7-9: The template text in _runbooks_instructions.jinja2 wrongly
tells the agent to read an "instruction" field; update the wording to reference
the runbook content returned in the StructuredToolResult.data field (i.e., the
output of fetch_runbook), so instruct the agent to "Read the runbook content
from the tool's data field" and note that runbook_content.pretty() (Robusta) and
the markdown wrapped in <runbook> (MD runbooks) are returned via data; ensure
the template states that runbook content takes priority over general
investigation steps and remove any mention of an "instruction" field.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 05c97665-6675-4cf6-a410-ab719f5e541d
📒 Files selected for processing (5)
holmes/core/prompt.pyholmes/plugins/prompts/_general_instructions.jinja2holmes/plugins/prompts/_runbook_instructions.jinja2holmes/plugins/prompts/_runbooks_instructions.jinja2holmes/plugins/prompts/investigation_procedure.jinja2
…ion field The fetch_runbook tool returns content via StructuredToolResult.data (Robusta: pretty() YAML dump, MD: markdown in <runbook> tags). The template incorrectly told the agent to read an "instruction" field. https://claude.ai/code/session_01ShM8tiPTYRSfpwJrYsrWDN Signed-off-by: Claude <noreply@anthropic.com>
…olmesGPT#1787) ## Summary This PR relaxes the mandatory runbook fetching requirements in HolmesGPT's investigation prompts, shifting from strict enforcement language to more flexible guidance. The changes reduce unnecessary runbook fetches by encouraging selective, issue-specific retrieval rather than speculative fetching. ## Key Changes - **Softened runbook fetching language**: Changed from "MANDATORY" and "MUST fetch" to conditional "if a runbook clearly matches" language across all prompt templates - **Removed threat-based enforcement**: Eliminated dramatic consequence warnings ("CRITICAL SYSTEM FAILURE", "IMMEDIATE TERMINATION REQUIRED") that created artificial urgency around runbook compliance - **Reduced redundant checks**: Removed duplicate runbook-related violation consequences and evaluation questions that were scattered throughout the investigation procedure - **Clarified selective fetching**: Added explicit guidance to "only fetch runbooks that are relevant" and "do not fetch runbooks speculatively or 'just in case'" - **Simplified phase progression example**: Removed the extra Phase 2 runbook-fetching step from the pod crash investigation example, streamlining the investigation flow - **Updated runbook catalog check**: Modified `prompt.py` to only enable runbook instructions when both runbooks exist AND the catalog is available (`bool(runbooks and getattr(runbooks, "catalog", True))`) ## Notable Implementation Details - The changes maintain runbook functionality while making it advisory rather than mandatory - Runbook instructions still take priority when a runbook is fetched, but the decision to fetch is now more flexible - The prompt templates now emphasize matching runbooks to specific issues rather than treating runbook fetching as a prerequisite step - Removed 4 separate runbook-related violation consequence statements that were creating cognitive overload https://claude.ai/code/session_01ShM8tiPTYRSfpwJrYsrWDN <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Tightened runbook enablement so only runbooks with an affirmed catalog are considered enabled. * Fetching is now limited to runbooks that clearly match the issue. * **Refactor** * Simplified investigation flow by removing mandatory runbook-first branches. * Allow investigations to proceed with available tools when no matching runbooks exist. * **Documentation** * Reworded runbook guidance to prohibit speculative fetching and add clearer post-fetch steps. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Claude <noreply@anthropic.com> Co-authored-by: Claude <noreply@anthropic.com>
Summary
This PR relaxes the mandatory runbook fetching requirements in HolmesGPT's investigation prompts, shifting from strict enforcement language to more flexible guidance. The changes reduce unnecessary runbook fetches by encouraging selective, issue-specific retrieval rather than speculative fetching.
Key Changes
prompt.pyto only enable runbook instructions when both runbooks exist AND the catalog is available (bool(runbooks and getattr(runbooks, "catalog", True)))Notable Implementation Details
https://claude.ai/code/session_01ShM8tiPTYRSfpwJrYsrWDN
Summary by CodeRabbit
Bug Fixes
Refactor
Documentation