Skip to content

Refactor prompt templates to inline includes and improve clarity - #1905

Merged
aantn merged 7 commits into
masterfrom
claude/simplify-jinja2-prompts-1A0Jz
Apr 29, 2026
Merged

aantn merged 7 commits into
masterfrom
claude/simplify-jinja2-prompts-1A0Jz

Conversation

@aantn

@aantn aantn commented Apr 13, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR refactors the prompt template structure by inlining several included template files directly into their parent templates. This improves maintainability by reducing template fragmentation and making the prompt logic more transparent and easier to follow.

Key Changes

  • Inlined _general_instructions.jinja2 into generic_ask.jinja2: Moved all general investigation instructions, Kubernetes-specific guidance, task management rules, and tool usage guidelines directly into the main prompt template with improved conditional logic for runbooks vs TodoWrite workflows.

  • Inlined _default_log_prompt.jinja2 into _fetch_logs.jinja2: Consolidated log fetching instructions for Coralogix, K8s base, and OpenSearch toolsets by replacing the include with the actual content.

  • Inlined _runbook_instructions.jinja2 into investigation_procedure.jinja2: Moved runbook usage instructions directly into the investigation procedure template.

  • Refactored base_user_prompt.jinja2: Replaced _runbook_instructions.jinja2 include with inline runbook selection logic and replaced _current_date_time.jinja2 include with direct date/time context injection.

  • Removed template files: Deleted _general_instructions.jinja2, _default_log_prompt.jinja2, _runbook_instructions.jinja2, _permission_errors.jinja2, _current_date_time.jinja2, and _global_instructions.jinja2 as their content is now inlined.

  • Improved permission error handling: Moved permission error instructions directly into generic_ask.jinja2 with clearer formatting and context.

Implementation Details

  • The refactoring maintains all existing functionality and conditional logic (feature flags like todowrite_enabled, runbooks_enabled, etc.)
  • Runbook instructions now appear in both generic_ask.jinja2 and investigation_procedure.jinja2 with consistent messaging
  • Log fetching instructions are duplicated across multiple toolset conditions but remain identical for clarity
  • The template hierarchy is flattened, making it easier to understand the complete prompt structure without following multiple includes

https://claude.ai/code/session_01HVq6giayp3P65pUpjJMkLo

Summary by CodeRabbit

Release Notes

  • Refactor
    • Consolidated and reorganized internal prompt instruction templates for improved maintainability.
    • Transitioned to feature-flag-driven instruction delivery, enabling more flexible control over system behavior.
    • Enhanced skill-selection mechanism with priority-ordered source handling.
    • Streamlined date/time interpretation and permission-error guidance within core prompt logic.

Flattens the prompt include graph from 15 files to 8 by inlining
partials that were used in only one place, plus deletes one orphan.

Inlined into generic_ask.jinja2:
- _general_instructions.jinja2 (1 caller)
- _permission_errors.jinja2    (1 caller, 6 lines)
- _runbooks_instructions.jinja2 (mutually exclusive with investigation_procedure)

Inlined into base_user_prompt.jinja2:
- _runbook_instructions.jinja2 (1 caller)
- _current_date_time.jinja2    (1 caller, 2 lines)

Inlined into investigation_procedure.jinja2:
- _runbooks_instructions.jinja2

Inlined 3x into _fetch_logs.jinja2:
- _default_log_prompt.jinja2 (only reused inside this one parent)

Deleted as orphan (no callers anywhere):
- _global_instructions.jinja2

Kept as separate files (entry points or substantial logical units):
- generic_ask.jinja2, base_user_prompt.jinja2,
  conversation_history_compaction.jinja2, _ticket_additions.jinja2
- _ai_safety.jinja2 (partner-mandated, kept discoverable)
- _toolsets_instructions.jinja2, _fetch_logs.jinja2,
  investigation_procedure.jinja2 (sizeable data-driven units)

No semantic changes. Rendered output is byte-equivalent except for
two stripped blank lines and one trailing space (all cosmetic).
All 34 existing prompt tests pass unchanged.

Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@netlify

netlify Bot commented Apr 13, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit df33a7b
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69f23dd0d00e0900081bb9c6
😎 Deploy Preview https://deploy-preview-1905--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Apr 13, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for b3884f7e (built in 5m 16s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:b3884f7e
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:b3884f7e me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:b3884f7e
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:b3884f7e
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:b3884f7e
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:b3884f7e me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:b3884f7e
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:b3884f7e

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:b3884f7e \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:b3884f7e

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:b3884f7e \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:b3884f7e

@github-actions

github-actions Bot commented Apr 13, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 #5 · Run @ __783c682__ (#25106657786) — Apr 29, 11:45 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 783c682 on branch claude/simplify-jinja2-prompts-1A0Jz

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 33.5s 4 9 $0.2245 82,364 80,447 23,010 1,917 959 56,532 23,915 — —
✅ 101_loki_historical_logs_pod_deleted 62.6s 9 14 $0.3373 200,722 197,522 26,328 3,200 520 169,594 27,928 — —
✅ 112_find_pvcs_by_uuid 18.7s 3 3 $0.1778 60,572 59,615 21,576 957 561 38,028 21,587 — —
✅ 12_job_crashing 39.8s 6 13 $0.2684 135,990 133,825 25,075 2,165 585 108,041 25,784 — —
✅ 176_network_policy_blocking_traffic_no_skills 57.7s 8 17 $0.3401 186,976 183,919 28,231 3,057 881 153,597 30,322 — —
✅ 227_count_configmaps_per_namespace[0] 23.2s 5 9 $0.1987 95,182 93,929 20,934 1,253 580 72,982 20,947 — —
✅ 243_pod_names_contain_service 33.6s 5 8 $0.2180 99,444 97,787 21,692 1,657 420 75,449 22,338 — —
✅ 24_misconfigured_pvc 36.2s 5 13 $0.2515 104,456 102,217 23,408 2,239 660 76,622 25,595 — —
✅ 43_current_datetime_from_prompt 4.9s 1 — $0.1097 17,126 16,982 16,982 144 144 0 16,982 — —
✅ 51_logs_summarize_errors 22.8s 4 5 $0.1893 78,273 77,172 21,414 1,101 354 55,746 21,426 — —
✅ 61_exact_match_counting 12.2s 3 3 $0.1388 52,987 52,621 17,963 366 219 34,647 17,974 — —
Total 31.4s avg 4.8 avg 9.4 avg $2.4543 1,114,092 1,096,036 28,231 18,056 959 841,238 254,798 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 74 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #4 · Run @ __dfa1cc7__ (#25101501320) — Apr 29, 09:42 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit dfa1cc7 on branch claude/simplify-jinja2-prompts-1A0Jz

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 40.7s 6 11 $0.2531 126,681 124,700 23,623 1,981 498 99,868 24,832 — —
✅ 101_loki_historical_logs_pod_deleted 55.3s 7 12 $0.3123 157,896 154,998 26,612 2,898 678 126,316 28,682 — —
✅ 112_find_pvcs_by_uuid 27.8s 5 4 $0.1993 95,645 94,366 20,886 1,279 333 73,467 20,899 — —
✅ 12_job_crashing 45.7s 7 15 $0.2963 159,501 157,086 26,807 2,415 474 129,553 27,533 — —
✅ 176_network_policy_blocking_traffic_no_skills 50.3s 7 13 $0.2927 154,326 151,633 25,399 2,693 898 125,463 26,170 — —
✅ 227_count_configmaps_per_namespace[0] 25.2s 5 9 $0.2032 95,180 93,916 20,933 1,264 584 72,038 21,878 — —
✅ 243_pod_names_contain_service 31.7s 5 8 $0.2155 98,806 97,239 21,682 1,567 538 74,880 22,359 — —
✅ 24_misconfigured_pvc 43.4s 7 14 $0.2775 146,134 143,722 23,835 2,412 517 118,045 25,677 — —
✅ 43_current_datetime_from_prompt 4.6s 1 — $0.1090 17,099 16,982 16,982 117 117 0 16,982 — —
✅ 51_logs_summarize_errors 23.3s 4 5 $0.1862 77,573 76,504 21,081 1,069 329 55,411 21,093 — —
✅ 61_exact_match_counting 14.2s 3 3 $0.1419 53,272 52,808 18,055 464 317 34,742 18,066 — —
Total 32.9s avg 5.2 avg 9.4 avg $2.4871 1,182,113 1,163,954 26,807 18,159 898 909,783 254,171 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 74 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #3 · Run @ __a823241__ (#24399088767) — Apr 14, 12:42 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit a823241 on branch claude/simplify-jinja2-prompts-1A0Jz

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 36.1s 5 9 $0.2331 104,601 102,687 23,276 1,914 563 79,398 23,289 — —
✅ 101_loki_historical_logs_pod_deleted 43.8s 5 10 $0.2606 109,686 107,210 24,666 2,476 844 81,826 25,384 — —
✅ 112_find_pvcs_by_uuid 21.8s 3 3 $0.1790 60,842 59,864 21,661 978 582 38,192 21,672 — —
✅ 12_job_crashing 40.0s 5 13 $0.2562 111,854 109,714 24,859 2,140 588 83,776 25,938 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 48.7s 6 16 $0.3093 140,056 137,067 26,957 2,989 850 107,655 29,412 — —
✅ 227_count_configmaps_per_namespace[0] 29.7s 6 10 $0.2143 114,257 112,752 20,752 1,505 515 91,780 20,972 — —
✅ 243_pod_names_contain_service 38.8s 5 9 $0.2309 102,261 100,332 22,534 1,929 579 77,134 23,198 — —
✅ 24_misconfigured_pvc 45.6s 6 16 $0.2857 132,746 129,954 25,140 2,792 900 103,389 26,565 — —
✅ 43_current_datetime_from_prompt 6.1s 1 — $0.1097 17,181 17,057 17,057 124 124 0 17,057 — —
✅ 51_logs_summarize_errors 27.6s 4 5 $0.1879 77,885 76,766 21,132 1,119 356 55,622 21,144 — —
✅ 61_exact_match_counting 15.2s 3 3 $0.1423 53,495 53,035 18,132 460 313 34,892 18,143 — —
Total 32.1s avg 4.5 avg 9.4 avg $2.4090 1,024,864 1,006,438 26,957 18,426 900 753,664 252,774 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 48 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #2 · Run @ __401f3a1__ (#24338226979) — Apr 13, 10:26 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 401f3a1 on branch claude/simplify-jinja2-prompts-1A0Jz

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 32.3s 4 9 $0.2249 83,274 81,329 23,144 1,945 789 57,618 23,711 — —
✅ 101_loki_historical_logs_pod_deleted 55.1s 7 13 $0.3301 166,247 163,348 28,236 2,899 864 132,085 31,263 — —
✅ 112_find_pvcs_by_uuid 18.6s 3 3 $0.1770 60,713 59,807 21,626 906 510 38,170 21,637 — —
✅ 12_job_crashing 42.8s 6 13 $0.2739 137,017 134,780 25,290 2,237 598 108,346 26,434 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 51.5s 8 14 $0.3140 180,706 177,950 26,425 2,756 744 150,600 27,350 — —
✅ 227_count_configmaps_per_namespace[0] 27.4s 5 10 $0.2062 97,476 96,091 21,306 1,385 424 74,566 21,525 — —
✅ 243_pod_names_contain_service 34.4s 5 9 $0.2255 101,520 99,671 22,277 1,849 566 77,086 22,585 — —
✅ 24_misconfigured_pvc 34.5s 5 12 $0.2300 101,768 99,837 22,304 1,931 621 76,752 23,085 — —
✅ 43_current_datetime_from_prompt 5.0s 1 — $0.1101 17,196 17,057 17,057 139 139 0 17,057 — —
✅ 51_logs_summarize_errors 23.9s 4 5 $0.1879 77,918 76,802 21,144 1,116 353 55,646 21,156 — —
✅ 61_exact_match_counting 8.3s 2 1 $0.1244 34,799 34,533 17,467 266 197 17,056 17,477 — —
Total 30.3s avg 4.5 avg 8.9 avg $2.4039 1,058,634 1,041,205 28,236 17,429 864 787,925 253,280 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 48 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #1 · Run @ __2e725e6__ (#24337856152) — Apr 13, 10:18 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 2e725e6 on branch claude/simplify-jinja2-prompts-1A0Jz

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 31.7s 4 9 $0.2206 81,727 79,969 22,637 1,758 761 55,885 24,084 — —
✅ 101_loki_historical_logs_pod_deleted 63.6s 7 14 $0.3259 164,742 161,541 27,124 3,201 931 132,391 29,150 — —
✅ 112_find_pvcs_by_uuid 34.0s 6 5 $0.2294 119,255 117,693 22,855 1,562 346 94,824 22,869 — —
✅ 12_job_crashing 38.9s 5 11 $0.2377 107,389 105,467 23,668 1,922 494 81,612 23,855 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 42.8s 5 11 $0.2595 110,480 108,092 25,124 2,388 742 82,670 25,422 — —
✅ 227_count_configmaps_per_namespace[0] 25.9s 5 9 $0.2047 95,555 94,309 21,015 1,246 585 72,050 22,259 — —
✅ 243_pod_names_contain_service 34.5s 5 8 $0.2160 99,906 98,248 21,798 1,658 449 76,437 21,811 — —
✅ 24_misconfigured_pvc 43.3s 6 13 $0.2572 123,964 121,738 23,468 2,226 689 96,989 24,749 — —
✅ 43_current_datetime_from_prompt 5.0s 1 — $0.1099 17,188 17,057 17,057 131 131 0 17,057 — —
✅ 51_logs_summarize_errors 24.5s 4 5 $0.1881 78,153 77,066 21,269 1,087 339 55,785 21,281 — —
✅ 61_exact_match_counting 13.2s 3 3 $0.1395 53,225 52,854 18,040 371 224 34,803 18,051 — —
Total 32.5s avg 4.6 avg 8.8 avg $2.3885 1,051,584 1,034,034 27,124 17,550 931 783,446 250,588 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 48 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit df33a7b on branch claude/simplify-jinja2-prompts-1A0Jz

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 34.6s 5 10 $0.2386 104,593 102,604 23,432 1,989 574 78,532 24,072 — —
✅ 101_loki_historical_logs_pod_deleted 43.1s 5 10 $0.2501 106,316 103,911 23,434 2,405 882 79,752 24,159 — —
✅ 112_find_pvcs_by_uuid 19.6s 3 3 $0.1778 60,578 59,623 21,585 955 555 38,027 21,596 — —
✅ 12_job_crashing 35.2s 5 10 $0.2381 106,959 105,050 23,525 1,909 592 80,954 24,096 — —
✅ 176_network_policy_blocking_traffic_no_skills 48.5s 6 15 $0.3172 142,155 139,127 27,420 3,028 890 108,541 30,586 — —
✅ 227_count_configmaps_per_namespace[0] 28.4s 6 10 $0.2128 113,791 112,284 20,675 1,507 519 91,595 20,689 — —
✅ 243_pod_names_contain_service 36.4s 5 8 $0.2175 99,633 97,890 21,726 1,743 421 76,151 21,739 — —
✅ 24_misconfigured_pvc 37.9s 6 13 $0.2521 125,904 123,696 23,346 2,208 498 100,165 23,531 — —
✅ 43_current_datetime_from_prompt 5.9s 1 — $0.1100 17,136 16,982 16,982 154 154 0 16,982 — —
✅ 51_logs_summarize_errors 25.5s 4 5 $0.1880 77,859 76,751 21,200 1,108 350 55,539 21,212 — —
✅ 61_exact_match_counting 12.4s 3 3 $0.1392 53,018 52,638 17,970 380 233 34,657 17,981 — —
Total 29.8s avg 4.5 avg 8.7 avg $2.3414 1,007,942 990,556 27,420 17,386 890 743,913 246,643 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 74 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/simplify-jinja2-prompts-1A0Jz -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/simplify-jinja2-prompts-1A0Jz -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented Apr 13, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

This PR refactors the prompt template system by consolidating instruction fragments into main templates with feature-flag-driven conditional logic, removing four standalone template files, and inlining their behavior into base_user_prompt.jinja2 and generic_ask.jinja2.

Changes

Cohort / File(s) Summary
Deleted Instruction Fragments
holmes/plugins/prompts/_current_date_time.jinja2, _general_instructions.jinja2, _global_instructions.jinja2, _permission_errors.jinja2
Removed standalone template files defining date/time handling, general investigation procedures, global instruction interpretation, and permission error guidance. These behaviors are consolidated into main templates with feature-flag-driven conditionals.
Refactored Main Prompts
holmes/plugins/prompts/base_user_prompt.jinja2, generic_ask.jinja2, investigation_procedure.jinja2
Inlined instruction logic from deleted fragments as feature-flag-driven conditional blocks. base_user_prompt.jinja2 now builds priority-ordered skill lists inline and includes date/time handling; generic_ask.jinja2 expands conditional inclusion of TodoWrite, skills, safety, and permission error instructions; investigation_procedure.jinja2 conditionally includes skill-usage rules based on skills_enabled flag.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~50 minutes

Possibly related PRs

Suggested labels

codex

Suggested reviewers

  • moshemorad
  • arikalon1
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main objective of the PR: refactoring prompt templates by inlining includes and improving clarity. It directly reflects the core changes across multiple template files.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share
Review rate limit: 5/8 reviews remaining, refill in 19 minutes and 34 seconds.

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Apr 13, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟢 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 11.97s 12.55s -4.6%
Warm Mean 5.37s 5.98s -10.2%
Warm Min 5.34s 5.83s
Warm Max 5.40s 6.16s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 17.32s 16.10s +7.6%
Warm Mean 7.52s 8.35s -9.9%
Warm Min 7.32s 8.04s
Warm Max 7.81s 8.90s

PR: b3884f7e | Master: 52fddaee | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (2)
holmes/plugins/prompts/generic_ask.jinja2 (2)

49-50: Minor redundancy in runbook guidance.

These two lines about runbook fetching overlap with the dedicated runbook instructions at lines 20-28 (when runbooks_enabled && !todowrite_enabled) or in investigation_procedure.jinja2 (when todowrite_enabled). The redundancy reinforces the guidance but could be consolidated for maintainability.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/prompts/generic_ask.jinja2` around lines 49 - 50, Duplicate
runbook guidance in generic_ask.jinja2 (the two bullet lines about fetching
runbooks) overlaps with the existing runbook block used when runbooks_enabled &&
!todowrite_enabled and with investigation_procedure.jinja2 when
todowrite_enabled; remove these two lines and instead reference or include the
existing runbook guidance (or call the same macro/partial) so the logic is
centralized (look for the runbook block at lines ~20-28 and the
investigation_procedure.jinja2 template to reuse).

100-115: Potentially redundant TodoWrite instructions.

When todowrite_enabled is true, investigation_procedure.jinja2 is already included (line 18), which contains comprehensive TodoWrite task management rules (lines 4-8, 19-91 in that file). These additional instructions at lines 100-115 overlap significantly, particularly around:

  • First tool call being TodoWrite
  • Task status updates
  • Breaking down problems into tasks

Consider removing this block since investigation_procedure.jinja2 already provides more detailed guidance on the same topics.

♻️ Suggested consolidation
-{% if todowrite_enabled %}
-# MANDATORY Task Management
-
-* You MUST use the TodoWrite tool for ANY investigation requiring multiple steps
-* Your FIRST tool call MUST be TodoWrite to create your investigation plan
-* Break down ALL complex problems into smaller, manageable tasks
-* You MUST update task status (pending → in_progress → completed) as you work through your investigation
-* The TodoWrite tool will show you a formatted task list - reference this throughout your investigation
-* Mark tasks as 'in_progress' when you start them, 'completed' when finished
-* Follow ALL tasks in your plan - don't skip any tasks
-* Use task management to ensure you don't miss important investigation steps
-* If you discover additional steps during investigation, add them to your task list using TodoWrite
-* When calling TodoWrite, you may ALSO call other tools in parallel to speed things up for your users and make them happy!
-* On the first TodoWrite call, mark at least one task as in_progress, and start working on it in parallel.
-* When calling TodoWrite for the first time, mark the tasks you started working on with 'in_progress' status.
-{% endif %}

The TodoWrite guidance is already fully covered by investigation_procedure.jinja2 which is included when todowrite_enabled is true.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/prompts/generic_ask.jinja2` around lines 100 - 115, The block
in generic_ask.jinja2 guarded by todowrite_enabled duplicates rules already
provided by the included investigation_procedure.jinja2; remove or disable the
redundant TodoWrite stanza (the entire if-block content that lists the MANDATORY
Task Management rules) so that only investigation_procedure.jinja2 supplies the
TodoWrite guidance, ensuring todowrite_enabled remains the single toggle and
avoiding duplicate/conflicting instructions; update references/comments if
needed to point readers to investigation_procedure.jinja2 for the canonical
rules.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@holmes/plugins/prompts/generic_ask.jinja2`:
- Around line 49-50: Duplicate runbook guidance in generic_ask.jinja2 (the two
bullet lines about fetching runbooks) overlaps with the existing runbook block
used when runbooks_enabled && !todowrite_enabled and with
investigation_procedure.jinja2 when todowrite_enabled; remove these two lines
and instead reference or include the existing runbook guidance (or call the same
macro/partial) so the logic is centralized (look for the runbook block at lines
~20-28 and the investigation_procedure.jinja2 template to reuse).
- Around line 100-115: The block in generic_ask.jinja2 guarded by
todowrite_enabled duplicates rules already provided by the included
investigation_procedure.jinja2; remove or disable the redundant TodoWrite stanza
(the entire if-block content that lists the MANDATORY Task Management rules) so
that only investigation_procedure.jinja2 supplies the TodoWrite guidance,
ensuring todowrite_enabled remains the single toggle and avoiding
duplicate/conflicting instructions; update references/comments if needed to
point readers to investigation_procedure.jinja2 for the canonical rules.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: c1a4ac7b-9af5-424c-bccd-a4a4b2cc2be8

📥 Commits

Reviewing files that changed from the base of the PR and between a9328e8 and 2e725e6.

📒 Files selected for processing (11)
  • holmes/plugins/prompts/_current_date_time.jinja2
  • holmes/plugins/prompts/_default_log_prompt.jinja2
  • holmes/plugins/prompts/_fetch_logs.jinja2
  • holmes/plugins/prompts/_general_instructions.jinja2
  • holmes/plugins/prompts/_global_instructions.jinja2
  • holmes/plugins/prompts/_permission_errors.jinja2
  • holmes/plugins/prompts/_runbook_instructions.jinja2
  • holmes/plugins/prompts/_runbooks_instructions.jinja2
  • holmes/plugins/prompts/base_user_prompt.jinja2
  • holmes/plugins/prompts/generic_ask.jinja2
  • holmes/plugins/prompts/investigation_procedure.jinja2
💤 Files with no reviewable changes (7)
  • holmes/plugins/prompts/_global_instructions.jinja2
  • holmes/plugins/prompts/_current_date_time.jinja2
  • holmes/plugins/prompts/_runbook_instructions.jinja2
  • holmes/plugins/prompts/_permission_errors.jinja2
  • holmes/plugins/prompts/_default_log_prompt.jinja2
  • holmes/plugins/prompts/_general_instructions.jinja2
  • holmes/plugins/prompts/_runbooks_instructions.jinja2

@aantn

aantn commented Apr 13, 2026

Copy link
Copy Markdown
Collaborator Author

@claude review

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a large structural refactoring of core prompt templates — the evals pass 11/11 and both flagged issues are pre-existing, but the scope of changes to the prompt infrastructure warrants a human look to confirm no instructions were accidentally altered or dropped during inlining.

Extended reasoning...

Overview

This PR inlines 7 Jinja2 template partial files (, , , , , , ) into their parent templates (, , , ). The intent is to flatten the template hierarchy and make prompt logic more transparent. This is a purely structural refactor — no new conditional logic or instructions are introduced.

Security risks

None. These are LLM system prompt templates, not code paths that handle authentication, data access, or external input validation. The inlining adds no new attack surface.

Level of scrutiny

Medium-high. While each individual change is mechanical (copy-paste from included files), the aggregate scope is significant: the prompt templates are the primary behavioral contract for how HolmesGPT investigates issues. A reviewer needs to verify that nothing was accidentally omitted, reordered in a semantically meaningful way, or subtly modified. The 11/11 eval pass is encouraging but evals cover a fixed test set and may not catch every edge case in conditional rendering paths.

Other factors

Both bugs reported by the bug-hunting system are explicitly marked pre-existing — they existed identically in the deleted and are faithfully reproduced. The PR does not worsen the existing behavior. CodeRabbit estimated ~45 minutes of review effort (complexity score 4). The PR duplicates identical log-fetching instructions 3× in rather than using a shared include — this is intentional per the description but is a maintainability trade-off worth a human acknowledging.

Comment thread holmes/plugins/prompts/generic_ask.jinja2
Comment thread holmes/plugins/prompts/generic_ask.jinja2
claude added 2 commits April 14, 2026 12:32
…mment to runbook block

Reverts the inlining of _default_log_prompt.jinja2 into _fetch_logs.jinja2
since the 3x duplication made things worse, not better. This file is
genuinely reused across 3 elif branches (coralogix, k8s_base, opensearch).

Adds a Jinja comment block to base_user_prompt.jinja2 explaining what
the runbook selection sections/available pattern does — it dynamically
builds a priority-ordered list from whichever of 3 context variables
(runbook_catalog, custom_instructions, global_instructions) are non-empty.

Before/after rendered prompt comparison:
- System prompt: identical
- CLI user prompt: 1 cosmetic blank line diff only
- Server user prompt: 1 cosmetic blank line diff only
All 34 prompt tests pass.

Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The inline bugs are pre-existing and evals pass 11/11, but this refactor touches core prompt engineering templates — worth a human pass to verify the inlining preserved all intended semantics, especially around the conditional logic.

Extended reasoning...

Overview

This PR deletes 6 Jinja2 template fragments and inlines their content directly into 3 primary templates: generic_ask.jinja2, base_user_prompt.jinja2, and investigation_procedure.jinja2. The scope is moderate-to-large for a refactor: all changes are in prompt engineering files that directly shape LLM behavior.

Security Risks

No security-sensitive code paths are touched. The templates are LLM prompt instructions only, with no auth, crypto, or permission logic.

Level of Scrutiny

Prompt templates are production-critical because they determine how the AI investigates issues. Even a purely mechanical inlining refactor can introduce subtle differences in whitespace, rendering order, or conditional guard semantics that would not be caught by 11/11 regression evals. Three issues were flagged by the bug hunter — all pre-existing, all faithfully reproduced by inlining — which suggests the PR is accurate but also that the refactor was a missed opportunity for cleanup. A human familiar with the prompt system should confirm the conditional logic (especially the todowrite_enabled and runbooks_enabled guard boundaries) is preserved correctly.

Other Factors

All CI evals pass (11/11, 0 regressions). The PR description is clear and the intent is well-documented. CodeRabbit rated it complexity 4/5 at approximately 45 minutes review time. The inline comments from the previous review run flag pre-existing semantic inconsistencies that remain unresolved.

Comment thread holmes/plugins/prompts/base_user_prompt.jinja2 Outdated
Master renamed runbooks→skills across all prompts.
Conflicts resolved by applying the skill rename to our inlined content:
- generic_ask.jinja2: runbooks_enabled→skills_enabled, runbook→skill text
- investigation_procedure.jinja2: inlined _skills_instructions content
- base_user_prompt.jinja2: inlined _skill_instructions content
- Deleted _skill_instructions.jinja2 and _skills_instructions.jinja2
  (single-use, consistent with our inlining approach)
- Deleted _general_instructions.jinja2 (keep our deletion)

Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@holmes/plugins/prompts/base_user_prompt.jinja2`:
- Around line 9-15: The template builds "sections" from skill_catalog,
custom_instructions, and global_instructions and then shows instructions that
may tell the agent to call fetch_skill even when no actual skills exist; update
the gating logic so fetch_skill-related guidance is only emitted when a real
skills collection is present (e.g., check the actual skills/skill_catalog length
or a "skills" variable) rather than merely the presence of
custom_instructions/global_instructions; locate the sections/available logic and
the later block that references fetch_skill (also the similar block at the later
34-41 range) and add a conditional that requires non-empty skills before
including any fetch_skill instruction text.

In `@holmes/plugins/prompts/generic_ask.jinja2`:
- Around line 49-50: The template emits unconditional bullets telling the model
to fetch skills; guard those lines with the skills toggle so they only appear
when skills are enabled. Update the generic_ask.jinja2 template to check the
boolean (e.g., skills_enabled) before rendering the two bullets ("if a skill url
is present..." and "if a skill in the catalog...") so the instructions are
omitted when skills are disabled; locate the bullets in generic_ask.jinja2 and
wrap them in the existing template conditional or add one around that block.
- Line 94: In the generic_ask.jinja2 template replace the grammatical slip "you
explain say" with a clear phrasing (e.g., "you explain" or "you say") so the
sentence reads smoothly—locate the sentence near the text fragment "That is
different than - for example - fetching a pod's logs and seeing that the pod
itself has permission errors." and update it to something like "in that case,
you explain that permission errors are the cause of the problem and give
details" to fix the wording.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 3e4f5607-d72c-4539-8dcc-f70d73699bdd

📥 Commits

Reviewing files that changed from the base of the PR and between a823241 and dfa1cc7.

📒 Files selected for processing (3)
  • holmes/plugins/prompts/base_user_prompt.jinja2
  • holmes/plugins/prompts/generic_ask.jinja2
  • holmes/plugins/prompts/investigation_procedure.jinja2
🚧 Files skipped from review as they are similar to previous changes (1)
  • holmes/plugins/prompts/investigation_procedure.jinja2

Comment thread holmes/plugins/prompts/base_user_prompt.jinja2
Comment thread holmes/plugins/prompts/generic_ask.jinja2
Comment thread holmes/plugins/prompts/generic_ask.jinja2
@Sheeproid
Sheeproid self-requested a review April 29, 2026 11:27
Comment thread holmes/plugins/prompts/generic_ask.jinja2
Comment thread holmes/plugins/prompts/generic_ask.jinja2
@aantn
aantn enabled auto-merge (squash) April 29, 2026 17:20
@aantn
aantn merged commit d29f143 into master Apr 29, 2026
22 of 23 checks passed
@aantn
aantn deleted the claude/simplify-jinja2-prompts-1A0Jz branch April 29, 2026 17:24
Comment on lines +25 to +31
{%- for sec in available %}
## {{ sec.title }} (priority #{{ loop.index }})

{%- set content = (sec.content|string) -%}
{{ content.replace('\n', '\n ') }}

{%- endfor %}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟣 🟣 This is a pre-existing issue: in base_user_prompt.jinja2 (lines 25-31), the per-section render loop uses {%- set content = (sec.content|string) -%} with whitespace-strip markers on both sides. The trailing -%} eats the newline that should separate the ## Heading (priority #N) line from the first content line, producing malformed markdown like ## Skill Catalog (priority #1)CATALOG: skill1 on a single line. Sections also run together because the next iteration's {%- for %} strips the separator. The bug was inlined verbatim from the deleted _skill_instructions.jinja2, but since this PR explicitly clarifies and physically rewrites these lines, it's the natural place to fix — drop the trailing - on the set tag (e.g., {%- set content = (sec.content|string) %}).

Extended reasoning...

What the bug is and how it manifests

In holmes/plugins/prompts/base_user_prompt.jinja2 lines 25-31, the per-section rendering loop is:

{%- for sec in available %}
## {{ sec.title }} (priority #{{ loop.index }})

{%- set content = (sec.content|string) -%}
{{ content.replace('\n', '\n   ') }}

{%- endfor %}

The {%- set content = (sec.content|string) -%} has whitespace-strip markers on both sides. The leading - strips the newline + blank line above (between the title and the set tag), and the trailing -%} strips the newline that should separate the assignment from the {{ content.replace(...) }} output. The result: the heading and the first content line are emitted on the same physical line.

Step-by-step proof

I rendered the template fragment using the project's actual Jinja2 setup (no trim_blocks/lstrip_blocks per holmes/plugins/prompts/__init__.py:37) with realistic inputs: skill_catalog='CATALOG: skill1\nCATALOG: skill2', custom_instructions='* Always check CPU\n* Always check memory', global_instructions='* Always check logs\n* Always check resource limits'. The output is exactly:

## Skill Catalog (priority #1)CATALOG: skill1
   CATALOG: skill2
## Subject/Issue Skills (priority #2)* Always check CPU
   * Always check memory
## Global Instructions (priority #3)* Always check logs
   * Always check resource limits

Note (1) the heading is glued to the first content character (priority #1)CATALOG:, priority #2)* Always, priority #3)* Always) and (2) consecutive sections run together because the next iteration's {%- for %} strips the blank line.

Why existing code doesn't prevent it

There is no condition that protects the inter-line whitespace — both strip markers on {%- set ... -%} are unconditional, and the {%- for %} at the start of each iteration eats the blank line between sections.

Impact

base_user_prompt.jinja2 is rendered as the user prompt (holmes/core/prompt.py:169), so this affects production output whenever any of skill_catalog, custom_instructions, or global_instructions is non-empty (the default for many deployments). The structural intent — distinct ## headings followed by indented content — is broken: the first content character of every section ends up on the heading line. When the first content character is markdown syntax (* or a heading-like token), the corruption is even worse because it visually merges with the heading text. While LLMs are typically tolerant of malformed markdown, this is a visible defect in a prompt that explicitly instructs the LLM about source priority.

Pre-existing status

This exact pattern was in the deleted _skill_instructions.jinja2 (and earlier in _runbook_instructions.jinja2). PR #1905 inlined the template verbatim into base_user_prompt.jinja2 lines 25-31 without fixing the bug. So while the defect predates this PR, the PR physically modifies these exact lines as part of a clarity-focused refactor, making it the natural place to fix it.

How to fix

Drop the trailing strip on the set tag — change {%- set content = (sec.content|string) -%} to {%- set content = (sec.content|string) %}. That preserves the newline before the {{ content.replace(...) }} output, so the title and content end up on separate lines. Optionally also drop the leading - to keep the blank line above for readability.

aantn pushed a commit that referenced this pull request May 1, 2026
Conflicts from PR #1905 (also refactored prompts). Resolved by keeping
our simplified versions in both generic_ask.jinja2 and
investigation_procedure.jinja2.

Signed-off-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants