Skip to content

Soften runbook fetching requirements and reduce enforcement language - #1787

Merged
aantn merged 4 commits into
masterfrom
claude/fix-holmes-runbook-fetching-cvilT
Mar 16, 2026
Merged

aantn merged 4 commits into
masterfrom
claude/fix-holmes-runbook-fetching-cvilT

Conversation

@aantn

@aantn aantn commented Mar 15, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR relaxes the mandatory runbook fetching requirements in HolmesGPT's investigation prompts, shifting from strict enforcement language to more flexible guidance. The changes reduce unnecessary runbook fetches by encouraging selective, issue-specific retrieval rather than speculative fetching.

Key Changes

  • Softened runbook fetching language: Changed from "MANDATORY" and "MUST fetch" to conditional "if a runbook clearly matches" language across all prompt templates
  • Removed threat-based enforcement: Eliminated dramatic consequence warnings ("CRITICAL SYSTEM FAILURE", "IMMEDIATE TERMINATION REQUIRED") that created artificial urgency around runbook compliance
  • Reduced redundant checks: Removed duplicate runbook-related violation consequences and evaluation questions that were scattered throughout the investigation procedure
  • Clarified selective fetching: Added explicit guidance to "only fetch runbooks that are relevant" and "do not fetch runbooks speculatively or 'just in case'"
  • Simplified phase progression example: Removed the extra Phase 2 runbook-fetching step from the pod crash investigation example, streamlining the investigation flow
  • Updated runbook catalog check: Modified prompt.py to only enable runbook instructions when both runbooks exist AND the catalog is available (bool(runbooks and getattr(runbooks, "catalog", True)))

Notable Implementation Details

  • The changes maintain runbook functionality while making it advisory rather than mandatory
  • Runbook instructions still take priority when a runbook is fetched, but the decision to fetch is now more flexible
  • The prompt templates now emphasize matching runbooks to specific issues rather than treating runbook fetching as a prerequisite step
  • Removed 4 separate runbook-related violation consequence statements that were creating cognitive overload

https://claude.ai/code/session_01ShM8tiPTYRSfpwJrYsrWDN

Summary by CodeRabbit

  • Bug Fixes

    • Tightened runbook enablement so only runbooks with an affirmed catalog are considered enabled.
    • Fetching is now limited to runbooks that clearly match the issue.
  • Refactor

    • Simplified investigation flow by removing mandatory runbook-first branches.
    • Allow investigations to proceed with available tools when no matching runbooks exist.
  • Documentation

    • Reworded runbook guidance to prohibit speculative fetching and add clearer post-fetch steps.

The LLM was obsessively fetching runbooks on every investigation because:
- System prompt used catastrophic language ("CRITICAL SYSTEM FAILURE",
  "IMMEDIATE TERMINATION") for not fetching runbooks
- Runbook instructions appeared redundantly in both system and user prompts
- investigation_procedure.jinja2 reinforced runbook fetching at every phase
  (Phase 1, evaluation, final review, violation consequences)
- runbooks_enabled was True even with empty catalogs

Changes:
- Replace catastrophic/mandatory language with conditional guidance:
  "only fetch if clearly matching, skip if no match"
- Remove redundant runbook checks from phase evaluation, final review,
  and violation consequences in investigation_procedure.jinja2
- Fix runbooks_enabled check to handle empty catalogs via getattr
- Soften _runbook_instructions.jinja2 (user prompt) from MUST to guidance
- Soften _general_instructions.jinja2 runbook bullets

https://claude.ai/code/session_01ShM8tiPTYRSfpwJrYsrWDN
Signed-off-by: Claude <noreply@anthropic.com>
@claude

claude Bot commented Mar 15, 2026

Copy link
Copy Markdown
Contributor

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review.

@github-actions

github-actions Bot commented Mar 15, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ a7cf377 (#23120785045)

✅ Results of HolmesGPT evals

Automatically triggered by commit a7cf377 on branch claude/fix-holmes-runbook-fetching-cvilT

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 31.6s 7 11 $0.2517 139,584 137,617 22,916 1,967 584 114,262 23,355 — —
✅ 101_loki_historical_logs_pod_deleted 31.8s 4 9 $0.2448 84,272 81,995 23,598 2,277 955 55,916 26,079 — —
✅ 111_pod_names_contain_service 25.0s 5 8 $0.2091 93,784 92,235 20,945 1,549 632 70,447 21,788 — —
✅ 112_find_pvcs_by_uuid 18.6s 4 4 $0.1888 77,791 76,716 21,489 1,075 488 55,215 21,501 — —
✅ 12_job_crashing 28.0s 5 12 $0.2382 106,000 104,122 23,509 1,878 558 79,723 24,399 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 32.6s 6 10 $0.2506 120,547 118,451 22,843 2,096 441 93,917 24,534 — —
✅ 227_count_configmaps_per_namespace[0] 23.6s 6 10 $0.2056 109,846 108,410 20,078 1,436 616 88,318 20,092 — —
✅ 24_misconfigured_pvc 27.8s 5 14 $0.3321 102,095 99,993 22,752 2,102 700 59,875 40,118 — —
✅ 43_current_datetime_from_prompt 3.7s 1 — $0.1058 16,582 16,464 16,464 118 118 0 16,464 — —
✅ 61_exact_match_counting 13.4s 4 4 $0.1476 69,294 68,851 17,751 443 226 51,088 17,763 — —
Total 23.6s avg 4.7 avg 9.1 avg $2.1743 919,795 904,854 23,598 14,941 955 668,761 236,093 — —

Benchmark comparison unavailable: No eval spans found in experiment 'ci-benchmark-23102181491'

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No eval spans found in experiment 'ci-benchmark-23102181491'

Benchmark experiment:

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ c3e50c2 (#23119065188)

✅ Results of HolmesGPT evals

Automatically triggered by commit c3e50c2 on branch claude/fix-holmes-runbook-fetching-cvilT

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 26.2s 5 11 $0.2321 101,319 99,465 22,712 1,854 630 75,550 23,915 — —
✅ 101_loki_historical_logs_pod_deleted 43.2s 6 13 $0.2893 135,009 132,093 26,146 2,916 781 105,933 26,160 — —
✅ 111_pod_names_contain_service 22.1s 4 7 $0.1905 75,285 73,923 20,179 1,362 438 52,982 20,941 — —
✅ 112_find_pvcs_by_uuid 20.3s 4 6 $0.2034 79,929 78,614 22,314 1,315 550 55,651 22,963 — —
✅ 12_job_crashing 27.6s 5 12 $0.2374 105,861 103,937 23,461 1,924 548 79,933 24,004 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 39.7s 7 15 $0.2849 153,336 151,118 25,808 2,218 475 124,097 27,021 — —
✅ 227_count_configmaps_per_namespace[0] 16.8s 4 9 $0.1829 73,895 72,663 19,618 1,232 605 52,376 20,287 — —
✅ 24_misconfigured_pvc 38.5s 8 16 $0.2919 167,213 164,571 24,179 2,642 523 139,391 25,180 — —
✅ 43_current_datetime_from_prompt 3.5s 1 — $0.1058 16,579 16,464 16,464 115 115 0 16,464 — —
✅ 61_exact_match_counting 9.8s 3 2 $0.1313 50,752 50,452 17,129 300 169 33,312 17,140 — —
Total 24.8s avg 4.7 avg 10.1 avg $2.1494 959,178 943,300 26,146 15,878 781 719,225 224,075 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 28ddaf6 (#23119033936)

✅ Results of HolmesGPT evals

Automatically triggered by commit 28ddaf6 on branch claude/fix-holmes-runbook-fetching-cvilT

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 21.3s 4 9 $0.2057 78,825 77,273 21,950 1,552 673 54,857 22,416 — —
✅ 101_loki_historical_logs_pod_deleted 39.8s 6 11 $0.2716 130,348 127,836 25,003 2,512 796 102,590 25,246 — —
✅ 111_pod_names_contain_service 27.8s 5 9 $0.2209 98,185 96,325 21,649 1,860 448 74,294 22,031 — —
✅ 112_find_pvcs_by_uuid 24.4s 5 8 $0.2165 97,756 95,999 21,114 1,757 702 74,177 21,822 — —
✅ 12_job_crashing 34.0s 6 14 $0.2760 131,800 129,389 24,842 2,411 500 102,728 26,661 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 32.1s 6 14 $0.2628 126,310 123,982 23,538 2,328 567 98,823 25,159 — —
✅ 227_count_configmaps_per_namespace[0] 24.1s 6 10 $0.2096 110,372 108,928 20,052 1,444 434 88,050 20,878 — —
✅ 24_misconfigured_pvc 34.1s 6 19 $0.2774 127,102 124,513 24,228 2,589 683 97,819 26,694 — —
✅ 43_current_datetime_from_prompt 4.3s 1 — $0.1061 16,594 16,464 16,464 130 130 0 16,464 — —
✅ 61_exact_match_counting 10.7s 3 2 $0.1321 50,822 50,497 17,155 325 186 33,331 17,166 — —
Total 25.3s avg 4.8 avg 10.7 avg $2.1787 968,114 951,206 25,003 16,908 796 726,669 224,537 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit 32782a6 on branch claude/fix-holmes-runbook-fetching-cvilT

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 30.5s 5 10 $0.2286 102,178 100,351 22,730 1,827 718 77,158 23,193 — —
✅ 101_loki_historical_logs_pod_deleted 52.9s 7 16 $0.4311 158,038 154,855 27,254 3,183 719 106,636 48,219 — —
✅ 111_pod_names_contain_service 24.7s 4 7 $0.1906 75,504 74,150 20,445 1,354 428 53,242 20,908 — —
✅ 112_find_pvcs_by_uuid 25.3s 5 7 $0.2045 94,943 93,481 20,527 1,462 411 72,252 21,229 — —
✅ 12_job_crashing 29.9s 5 13 $0.2458 109,162 107,316 24,397 1,846 542 81,656 25,660 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 38.1s 6 15 $0.4182 127,109 124,714 25,172 2,395 490 72,701 52,013 — —
✅ 227_count_configmaps_per_namespace[0] 24.9s 6 10 $0.2066 109,872 108,462 20,103 1,410 620 88,013 20,449 — —
✅ 24_misconfigured_pvc 30.3s 5 13 $0.2352 102,031 99,985 22,681 2,046 719 76,497 23,488 — —
✅ 43_current_datetime_from_prompt 4.6s 1 — $0.1062 16,596 16,464 16,464 132 132 0 16,464 — —
✅ 61_exact_match_counting 14.1s 4 4 $0.1484 69,373 68,906 17,769 467 247 51,125 17,781 — —
Total 27.5s avg 4.8 avg 10.6 avg $2.4151 964,806 948,684 27,254 16,122 719 679,280 269,404 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-holmes-runbook-fetching-cvilT -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals in automatic regression runs:

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-holmes-runbook-fetching-cvilT -f markers=regression -f filter=

@github-actions

github-actions Bot commented Mar 15, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 52a6ae4d (built in 8m 3s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:52a6ae4d
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:52a6ae4d me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:52a6ae4d
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:52a6ae4d
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:52a6ae4d
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:52a6ae4d me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:52a6ae4d
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:52a6ae4d

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:52a6ae4d \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:52a6ae4d

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:52a6ae4d \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:52a6ae4d

@coderabbitai

coderabbitai Bot commented Mar 15, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 782c027c-6581-466a-9393-e678e8e55617

📥 Commits

Reviewing files that changed from the base of the PR and between c3e50c2 and a7cf377.

📒 Files selected for processing (1)
  • holmes/plugins/prompts/_runbooks_instructions.jinja2
🚧 Files skipped from review as they are similar to previous changes (1)
  • holmes/plugins/prompts/_runbooks_instructions.jinja2

Walkthrough

This change tightens the runbooks_enabled computation in prompt construction and updates multiple prompt templates to make runbook fetching conditional: only fetch runbooks that clearly match the issue and avoid speculative fetching. No function signatures changed.

Changes

Cohort / File(s) Summary
Core Logic
holmes/core/prompt.py
Tightened runbooks_enabled evaluation to bool(runbooks and getattr(runbooks, "catalog", True)) and AND it with is_enabled(PromptComponent.TIME_RUNBOOKS).
Prompt: general instructions
holmes/plugins/prompts/_general_instructions.jinja2
Reworded guidance to fetch runbooks only when they clearly match; removed unconditional fetch-on-detection language.
Prompt: per-runbook instructions
holmes/plugins/prompts/_runbook_instructions.jinja2
Added explicit per-runbook rendering blocks, reordered/prioritized display, tightened fetch condition to "clearly matches", and expanded post-fetch step-by-step workflow.
Prompt: runbooks flow
holmes/plugins/prompts/_runbooks_instructions.jinja2
Simplified runbook enforcement to read-and-follow fetched runbook content; removed prior mandatory multi-step enforcement and speculative/fetching directives.
Investigation procedure
holmes/plugins/prompts/investigation_procedure.jinja2
Replaced mandatory immediate runbook fetch with a conditional check; removed runbook verification steps from evaluation and final review; pruned runbook-centric conditionals.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested reviewers

  • moshemorad
  • arikalon1
  • mainred
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately reflects the main objective of the PR: relaxing mandatory runbook-fetching requirements and reducing enforcement language across the codebase.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Mar 15, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 32782a6
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69b80635e5b4a700088c6ebe
😎 Deploy Preview https://deploy-preview-1787--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Mar 15, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 11.06s 11.35s -2.6%
Warm Mean 5.10s 5.28s -3.4%
Warm Min 5.03s 5.22s
Warm Max 5.14s 5.35s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 28.84s 25.99s +10.9%
Warm Mean 9.46s 8.51s +11.1%
Warm Min 8.05s 8.07s
Warm Max 13.13s 9.52s

PR: 52a6ae4d | Master: cb78c6d5 | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@holmes/plugins/prompts/_runbooks_instructions.jinja2`:
- Around line 7-9: The template text in _runbooks_instructions.jinja2 wrongly
tells the agent to read an "instruction" field; update the wording to reference
the runbook content returned in the StructuredToolResult.data field (i.e., the
output of fetch_runbook), so instruct the agent to "Read the runbook content
from the tool's data field" and note that runbook_content.pretty() (Robusta) and
the markdown wrapped in <runbook> (MD runbooks) are returned via data; ensure
the template states that runbook content takes priority over general
investigation steps and remove any mention of an "instruction" field.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 05c97665-6675-4cf6-a410-ab719f5e541d

📥 Commits

Reviewing files that changed from the base of the PR and between 274fbf4 and c3e50c2.

📒 Files selected for processing (5)
  • holmes/core/prompt.py
  • holmes/plugins/prompts/_general_instructions.jinja2
  • holmes/plugins/prompts/_runbook_instructions.jinja2
  • holmes/plugins/prompts/_runbooks_instructions.jinja2
  • holmes/plugins/prompts/investigation_procedure.jinja2

Comment thread holmes/plugins/prompts/_runbooks_instructions.jinja2 Outdated
claude and others added 2 commits March 15, 2026 22:27
…ion field

The fetch_runbook tool returns content via StructuredToolResult.data
(Robusta: pretty() YAML dump, MD: markdown in <runbook> tags). The
template incorrectly told the agent to read an "instruction" field.

https://claude.ai/code/session_01ShM8tiPTYRSfpwJrYsrWDN
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn enabled auto-merge (squash) March 16, 2026 13:31
@aantn
aantn merged commit 405937c into master Mar 16, 2026
21 of 22 checks passed
@aantn
aantn deleted the claude/fix-holmes-runbook-fetching-cvilT branch March 16, 2026 13:55
henrikrexed pushed a commit to henrikrexed/holmesgpt that referenced this pull request Mar 17, 2026
…olmesGPT#1787)

## Summary
This PR relaxes the mandatory runbook fetching requirements in
HolmesGPT's investigation prompts, shifting from strict enforcement
language to more flexible guidance. The changes reduce unnecessary
runbook fetches by encouraging selective, issue-specific retrieval
rather than speculative fetching.

## Key Changes

- **Softened runbook fetching language**: Changed from "MANDATORY" and
"MUST fetch" to conditional "if a runbook clearly matches" language
across all prompt templates
- **Removed threat-based enforcement**: Eliminated dramatic consequence
warnings ("CRITICAL SYSTEM FAILURE", "IMMEDIATE TERMINATION REQUIRED")
that created artificial urgency around runbook compliance
- **Reduced redundant checks**: Removed duplicate runbook-related
violation consequences and evaluation questions that were scattered
throughout the investigation procedure
- **Clarified selective fetching**: Added explicit guidance to "only
fetch runbooks that are relevant" and "do not fetch runbooks
speculatively or 'just in case'"
- **Simplified phase progression example**: Removed the extra Phase 2
runbook-fetching step from the pod crash investigation example,
streamlining the investigation flow
- **Updated runbook catalog check**: Modified `prompt.py` to only enable
runbook instructions when both runbooks exist AND the catalog is
available (`bool(runbooks and getattr(runbooks, "catalog", True))`)

## Notable Implementation Details

- The changes maintain runbook functionality while making it advisory
rather than mandatory
- Runbook instructions still take priority when a runbook is fetched,
but the decision to fetch is now more flexible
- The prompt templates now emphasize matching runbooks to specific
issues rather than treating runbook fetching as a prerequisite step
- Removed 4 separate runbook-related violation consequence statements
that were creating cognitive overload

https://claude.ai/code/session_01ShM8tiPTYRSfpwJrYsrWDN

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Tightened runbook enablement so only runbooks with an affirmed catalog
are considered enabled.
  * Fetching is now limited to runbooks that clearly match the issue.

* **Refactor**
* Simplified investigation flow by removing mandatory runbook-first
branches.
* Allow investigations to proceed with available tools when no matching
runbooks exist.

* **Documentation**
* Reworded runbook guidance to prohibit speculative fetching and add
clearer post-fetch steps.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants