Simplify jinja2 prompts: inline, deduplicate, fix logic - #1970
Conversation
Flattens the prompt include graph from 15 files to 8 by inlining partials that were used in only one place, plus deletes one orphan. Inlined into generic_ask.jinja2: - _general_instructions.jinja2 (1 caller) - _permission_errors.jinja2 (1 caller, 6 lines) - _runbooks_instructions.jinja2 (mutually exclusive with investigation_procedure) Inlined into base_user_prompt.jinja2: - _runbook_instructions.jinja2 (1 caller) - _current_date_time.jinja2 (1 caller, 2 lines) Inlined into investigation_procedure.jinja2: - _runbooks_instructions.jinja2 Inlined 3x into _fetch_logs.jinja2: - _default_log_prompt.jinja2 (only reused inside this one parent) Deleted as orphan (no callers anywhere): - _global_instructions.jinja2 Kept as separate files (entry points or substantial logical units): - generic_ask.jinja2, base_user_prompt.jinja2, conversation_history_compaction.jinja2, _ticket_additions.jinja2 - _ai_safety.jinja2 (partner-mandated, kept discoverable) - _toolsets_instructions.jinja2, _fetch_logs.jinja2, investigation_procedure.jinja2 (sizeable data-driven units) No semantic changes. Rendered output is byte-equivalent except for two stripped blank lines and one trailing space (all cosmetic). All 34 existing prompt tests pass unchanged. Signed-off-by: Claude <noreply@anthropic.com>
…mment to runbook block Reverts the inlining of _default_log_prompt.jinja2 into _fetch_logs.jinja2 since the 3x duplication made things worse, not better. This file is genuinely reused across 3 elif branches (coralogix, k8s_base, opensearch). Adds a Jinja comment block to base_user_prompt.jinja2 explaining what the runbook selection sections/available pattern does — it dynamically builds a priority-ordered list from whichever of 3 context variables (runbook_catalog, custom_instructions, global_instructions) are non-empty. Before/after rendered prompt comparison: - System prompt: identical - CLI user prompt: 1 cosmetic blank line diff only - Server user prompt: 1 cosmetic blank line diff only All 34 prompt tests pass. Signed-off-by: Claude <noreply@anthropic.com>
….1:19415/git/HolmesGPT/holmesgpt into claude/simplify-jinja2-prompts-1A0Jz
Master renamed runbooks→skills across all prompts. Conflicts resolved by applying the skill rename to our inlined content: - generic_ask.jinja2: runbooks_enabled→skills_enabled, runbook→skill text - investigation_procedure.jinja2: inlined _skills_instructions content - base_user_prompt.jinja2: inlined _skill_instructions content - Deleted _skill_instructions.jinja2 and _skills_instructions.jinja2 (single-use, consistent with our inlining approach) - Deleted _general_instructions.jinja2 (keep our deletion) Signed-off-by: Claude <noreply@anthropic.com>
generic_ask.jinja2: - Remove "Use conversation history to maintain continuity" (filler) - Remove "Whenever possible you MUST first use tools" (redundant with general instructions section) - Remove "Ask for multiple tool calls at the same time" (duplicated in investigation_procedure and task management) - Merge "run as many tools...do so repeatedly" into single bullets - Remove duplicate skill-fetching bullets (already in Skill Usage block) - Collapse verbose Task Management section (14 lines → 5) - Remove "You are able to make tool calls" (the LLM already knows) investigation_procedure.jinja2 (221 → 50 lines, 77% reduction): - Fix self-contradictory phase evaluation: "yes" to "Do I have enough information?" was incorrectly triggering continuation. Now each question has explicit IF NO/IF YES direction. - Remove duplicate TASK COMPLETION ENFORCEMENT + ENFORCEMENT RULES (said the same thing twice with different headers) - Remove duplicate evaluation question lists (listed identically at lines 112 and 128) - Remove ASCII checklist art and status update examples - Remove triple VIOLATION CONSEQUENCES blocks - Collapse FINAL REVIEW PHASE from 23 bullets to one sentence - Remove INVESTIGATION PHASE TRANSITION EXAMPLES (redundant) - Remove trivial dependency examples (covered by the parallel example) _fetch_logs.jinja2: - Remove stray "Logs from newrelic" fragment with typo that was incorrectly placed inside the k8s_yaml_ts elif branch System prompt reduction: ~40% fewer chars across all configurations. Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.
Tip: disable this comment in your organization's Code Review settings.
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughThis PR removes the AI safety prompt component and its associated guardrails template from the Holmes system. The ChangesRemoval of AI Safety Component and Prompt Restructuring
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested labels
Suggested reviewers
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
There was a problem hiding this comment.
Actionable comments posted: 1
🧹 Nitpick comments (3)
holmes/plugins/prompts/base_user_prompt.jinja2 (1)
28-29: Minor: Indentation may not render as expected in Markdown.Line 29 indents continuation lines with 3 spaces (
'\n '). In standard Markdown, leading spaces don't create indentation for paragraph text—they're typically stripped. If the goal is visual indentation in the rendered prompt, this may not achieve the desired effect depending on how the LLM processes the prompt.If indentation is critical, consider using a Markdown-native approach (like blockquotes with
>). Otherwise, if this is just for human readability in logs/debugging, it's fine as-is.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@holmes/plugins/prompts/base_user_prompt.jinja2` around lines 28 - 29, The current template replaces newlines with '\n ' (see variable content and sec.content and the replace call) which may not produce visible indentation in Markdown; update the replacement to use a Markdown-native pattern such as '\n> ' (blockquote) or another explicit Markdown construct so continuation lines render as indented in LLM/Markdown contexts — change the replace invocation on content to insert a blockquote prefix (or other Markdown syntax) instead of three spaces.holmes/plugins/prompts/generic_ask.jinja2 (2)
16-24: Duplicated skill instructions betweengeneric_ask.jinja2andinvestigation_procedure.jinja2.Lines 16-24 contain skill usage instructions that are nearly identical to lines 8-16 in
investigation_procedure.jinja2. Sinceinvestigation_procedure.jinja2is included at line 14 whentodowrite_enabledis true, theelsebranch here (lines 15-25) handles the!todowrite_enabled && skills_enabledcase. However, whentodowrite_enabledis true, the includedinvestigation_procedure.jinja2also has its own skill instructions (guarded byskills_enabled).This appears intentional—one path for todowrite mode, one for non-todowrite mode—but consider extracting the skill instructions into a single partial to avoid maintaining two copies.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@holmes/plugins/prompts/generic_ask.jinja2` around lines 16 - 24, The duplicate "Skill Usage" block appears in generic_ask.jinja2 and investigation_procedure.jinja2 (both guarded by skills_enabled and influenced by todowrite_enabled); extract that repeated block into a single partial (e.g., _skill_usage.jinja2) and replace the inline blocks in both templates with an include of the partial, keeping the original conditional guards (skills_enabled and todowrite_enabled) intact so behavior for the todowrite_enabled and non-todowrite paths remains the same and references to fetch_skill and reading the tool's data field are preserved.
31-63: Consider consolidatinggeneral_instructions_enabledblocks.The
general_instructions_enabledconditional content is split across three separate blocks (lines 31-63, 79-88, and 100-106), interleaved with other conditionals. While this works, it may complicate future maintenance.If the ordering of sections is flexible, consolidating these into a contiguous block would improve readability.
Also applies to: 79-88, 100-106
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@holmes/plugins/prompts/generic_ask.jinja2` around lines 31 - 63, The template splits the general_instructions_enabled conditional into multiple non-contiguous blocks which complicates maintenance; consolidate all occurrences of the general_instructions_enabled conditional in holmes/plugins/prompts/generic_ask.jinja2 into a single contiguous block containing the "In general" section, the cluster_name sub-block, and the Kubernetes investigation section, preserving all internal content and any nested conditionals (e.g., the cluster_name check) and ensuring the relative ordering of those internal paragraphs remains correct so behavior does not change; remove the now-redundant separate general_instructions_enabled blocks at lines referenced in the PR and run a quick template-render smoke test to confirm no logic/regression changes.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@holmes/plugins/prompts/investigation_procedure.jinja2`:
- Around line 33-39: The conditional block in investigation_procedure.jinja2
uses hardcoded step numbers causing inconsistent numbering when skills_enabled
is false; update the template so numbering is generated dynamically or use
unnumbered bullets: either (A) replace manual numbers with a Jinja counter
(e.g., set a local counter and increment it when emitting each step so the final
"3. Execute ALL tasks" uses the computed value) or (B) convert steps to
unnumbered list items so the conditional branch doesn't break sequence;
reference the skills_enabled condition and the step text "Create initial
TodoWrite task list..." and "Execute ALL tasks" when applying the fix.
---
Nitpick comments:
In `@holmes/plugins/prompts/base_user_prompt.jinja2`:
- Around line 28-29: The current template replaces newlines with '\n ' (see
variable content and sec.content and the replace call) which may not produce
visible indentation in Markdown; update the replacement to use a Markdown-native
pattern such as '\n> ' (blockquote) or another explicit Markdown construct so
continuation lines render as indented in LLM/Markdown contexts — change the
replace invocation on content to insert a blockquote prefix (or other Markdown
syntax) instead of three spaces.
In `@holmes/plugins/prompts/generic_ask.jinja2`:
- Around line 16-24: The duplicate "Skill Usage" block appears in
generic_ask.jinja2 and investigation_procedure.jinja2 (both guarded by
skills_enabled and influenced by todowrite_enabled); extract that repeated block
into a single partial (e.g., _skill_usage.jinja2) and replace the inline blocks
in both templates with an include of the partial, keeping the original
conditional guards (skills_enabled and todowrite_enabled) intact so behavior for
the todowrite_enabled and non-todowrite paths remains the same and references to
fetch_skill and reading the tool's data field are preserved.
- Around line 31-63: The template splits the general_instructions_enabled
conditional into multiple non-contiguous blocks which complicates maintenance;
consolidate all occurrences of the general_instructions_enabled conditional in
holmes/plugins/prompts/generic_ask.jinja2 into a single contiguous block
containing the "In general" section, the cluster_name sub-block, and the
Kubernetes investigation section, preserving all internal content and any nested
conditionals (e.g., the cluster_name check) and ensuring the relative ordering
of those internal paragraphs remains correct so behavior does not change; remove
the now-redundant separate general_instructions_enabled blocks at lines
referenced in the PR and run a quick template-render smoke test to confirm no
logic/regression changes.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 27b48bf1-c90e-4deb-84f9-95df0b25c14f
📒 Files selected for processing (10)
holmes/plugins/prompts/_current_date_time.jinja2holmes/plugins/prompts/_fetch_logs.jinja2holmes/plugins/prompts/_general_instructions.jinja2holmes/plugins/prompts/_global_instructions.jinja2holmes/plugins/prompts/_permission_errors.jinja2holmes/plugins/prompts/_skill_instructions.jinja2holmes/plugins/prompts/_skills_instructions.jinja2holmes/plugins/prompts/base_user_prompt.jinja2holmes/plugins/prompts/generic_ask.jinja2holmes/plugins/prompts/investigation_procedure.jinja2
💤 Files with no reviewable changes (7)
- holmes/plugins/prompts/_global_instructions.jinja2
- holmes/plugins/prompts/_permission_errors.jinja2
- holmes/plugins/prompts/_skills_instructions.jinja2
- holmes/plugins/prompts/_skill_instructions.jinja2
- holmes/plugins/prompts/_general_instructions.jinja2
- holmes/plugins/prompts/_current_date_time.jinja2
- holmes/plugins/prompts/_fetch_logs.jinja2
…phases Fixes inconsistent step numbering when skills_enabled is false — steps jumped from 1 to 3. Switching to bullets avoids fragile conditional numbering entirely. Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
🧹 Nitpick comments (2)
holmes/plugins/prompts/investigation_procedure.jinja2 (2)
34-34: Drop the repeated skill-first bullet.Line 34 restates the same rule already given in Lines 10-12. Since this template is on the hot path and this PR is explicitly reducing prompt size, keeping the dedicated "Skill Usage" section and removing the duplicate would be cleaner.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@holmes/plugins/prompts/investigation_procedure.jinja2` at line 34, The template contains a duplicated "skill-first" instruction ("If a skill in the catalog clearly matches the issue, fetch it first. Otherwise, skip.") that repeats the rule already stated in the "Skill Usage" section; remove this repeated bullet so only the dedicated "Skill Usage" section contains that guidance, leaving other prompt content unchanged and ensuring no other references rely on the duplicate line.
3-3: Keep the TodoWrite trigger scoped to investigations.
generic_ask.jinja2:90-98still says TodoWrite is for investigations requiring multiple steps, but Line 3 broadens that to any multi-step question. That scope drift will push routine multi-part asks into unnecessary TodoWrite calls.Suggested wording
-For multi-step questions, start by calling TodoWrite with a `todos` parameter containing an array of task objects. Each task must have: +For multi-step investigations, start by calling TodoWrite with a `todos` parameter containing an array of task objects. Each task must have:🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@holmes/plugins/prompts/investigation_procedure.jinja2` at line 3, The TodoWrite trigger description should be restricted to investigation flows only: update the sentence in investigation_procedure.jinja2 (current "For multi-step questions...") to specify "For investigations that require multiple steps" and reference the TodoWrite trigger used by investigations; ensure consistency with generic_ask.jinja2's guidance around TodoWrite (lines referenced as generic_ask.jinja2:90-98) so ordinary multi-part questions are not routed to TodoWrite—keep the TodoWrite scope limited to investigation contexts and mirror that exact phrasing wherever TodoWrite is documented.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@holmes/plugins/prompts/investigation_procedure.jinja2`:
- Around line 3-7: The doc text in investigation_procedure.jinja2 conflicts with
examples: change the guidance for new task status to match the examples by
stating that initial tasks should use "in_progress" instead of "pending" (i.e.,
update the TodoWrite description for the todos parameter to require `status`:
"in_progress" for newly started tasks), and make the same change wherever the
initial-status rule is documented (e.g., the example payload in this template
and the related generic_ask.jinja2 examples) so all references to initial task
status are consistent.
---
Nitpick comments:
In `@holmes/plugins/prompts/investigation_procedure.jinja2`:
- Line 34: The template contains a duplicated "skill-first" instruction ("If a
skill in the catalog clearly matches the issue, fetch it first. Otherwise,
skip.") that repeats the rule already stated in the "Skill Usage" section;
remove this repeated bullet so only the dedicated "Skill Usage" section contains
that guidance, leaving other prompt content unchanged and ensuring no other
references rely on the duplicate line.
- Line 3: The TodoWrite trigger description should be restricted to
investigation flows only: update the sentence in investigation_procedure.jinja2
(current "For multi-step questions...") to specify "For investigations that
require multiple steps" and reference the TodoWrite trigger used by
investigations; ensure consistency with generic_ask.jinja2's guidance around
TodoWrite (lines referenced as generic_ask.jinja2:90-98) so ordinary multi-part
questions are not routed to TodoWrite—keep the TodoWrite scope limited to
investigation contexts and mirror that exact phrasing wherever TodoWrite is
documented.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: da3a8d40-2e20-4806-bf92-1a1b5c423031
📒 Files selected for processing (1)
holmes/plugins/prompts/investigation_procedure.jinja2
- Change "multi-step questions" → "multi-step investigations" to keep TodoWrite scoped to investigation flows, not routine multi-part asks - Remove duplicate skill-first bullet from Investigation Phases (already covered by the Skill Usage section above) - Clarify status field: "pending" for queued tasks or "in_progress" for tasks being started now, matching the example payload Signed-off-by: Claude <noreply@anthropic.com>
Conflicts from PR #1905 (also refactored prompts). Resolved by keeping our simplified versions in both generic_ask.jinja2 and investigation_procedure.jinja2. Signed-off-by: Claude <noreply@anthropic.com>
📂 Previous Runs📜 #5 · Run @ __3e7b7b1__ (#25883778899) — May 14, 20:34 UTC✅ Results of HolmesGPT evalsAutomatically triggered by commit 3e7b7b1 on branch Results of HolmesGPT evals
Benchmark comparison unavailable: No ci-benchmark experiments found Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: No ci-benchmark experiments found Comparison indicators:
📜 #4 · Run @ __5078890__ (#25883365393) — May 14, 20:26 UTC✅ Results of HolmesGPT evalsAutomatically triggered by commit 5078890 on branch Results of HolmesGPT evals
Benchmark comparison unavailable: No ci-benchmark experiments found Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: No ci-benchmark experiments found Comparison indicators:
📜 #3 · Run @ __877765e__ (#25883105057) — May 14, 20:19 UTC✅ Results of HolmesGPT evalsAutomatically triggered by commit 877765e on branch Results of HolmesGPT evals
Benchmark comparison unavailable: No ci-benchmark experiments found Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: No ci-benchmark experiments found Comparison indicators:
📜 #2 · Run @ __20eca64__ (#25313442524) — May 4, 10:21 UTC✅ Results of HolmesGPT evalsAutomatically triggered by commit 20eca64 on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 73 test/model combinations loaded Benchmark experiment:
No benchmark data available for comparison. Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
📜 #1 · Run @ __101b665__ (#25313124106) — May 4, 10:14 UTC✅ Results of HolmesGPT evalsAutomatically triggered by commit 101b665 on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 73 test/model combinations loaded Benchmark experiment:
No benchmark data available for comparison. Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
✅ Results of HolmesGPT evalsAutomatically triggered by commit 808a914 on branch Results of HolmesGPT evals
Benchmark comparison unavailable: No ci-benchmark experiments found Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: No ci-benchmark experiments found Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" Option 3: Add PR labels to include extra evals (applies to both automatic runs and
Examples: 🏷️ Valid tags
🤖 Valid models
Commands: CLI: |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:a4d4b2b7
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:a4d4b2b7 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:a4d4b2b7
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:a4d4b2b7
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:a4d4b2b7
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:a4d4b2b7 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:a4d4b2b7
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:a4d4b2b7Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:a4d4b2b7 \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:a4d4b2b7Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:a4d4b2b7 \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:a4d4b2b7 |
🔬 CLI Performance Benchmark🟡 Startup Time (no LLM)Measures
🟡 Full CLI with LLMMeasures
PR: |
…te path The "ask for multiple tool calls at the same time" instruction was only covered in the todowrite/investigation path. Simple ask mode had no explicit parallel-calling guidance. Signed-off-by: Claude <noreply@anthropic.com>
….1:56565/git/HolmesGPT/holmesgpt into claude/simplify-jinja2-prompts-1A0Jz
Partners confirmed this is no longer required. Removes: - _ai_safety.jinja2 template file - AI_SAFETY enum value from PromptComponent - ai_safety_enabled template context variable - DISABLED_BY_DEFAULT set (now empty, kept for future use) - tests/test_ai_safety_prompt.py (tested deleted content) - 3 disabled-by-default tests in test_prompt.py (tested AI_SAFETY) Signed-off-by: Claude <noreply@anthropic.com>
…_ask Single caller, no tests load it directly. Content inlined verbatim with no modifications. Reduces prompt files from 9 to 8. Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
holmes/plugins/prompts/generic_ask.jinja2 (1)
22-67: 🛠️ Refactor suggestion | 🟠 Major | ⚡ Quick winSkill Usage block is duplicated across both branches of
todowrite_enabled.Lines 22–30 and 59–67 are byte-identical. Given this PR's deduplication goal, hoist the
# Skill Usageblock out of thetodowrite_enabledif/else so it renders once for both paths. This also prevents the two copies from drifting over time.♻️ Proposed structure
{% if todowrite_enabled %} If the user's question is ambiguous or lacks critical details, ask for clarification first. Do NOT create TodoWrite tasks for unclear questions. For multi-step investigations, start by calling TodoWrite with a `todos` parameter containing an array of task objects. Each task must have: - `id`: unique identifier (string) - `content`: specific task description (string) - `status`: "pending" for queued tasks or "in_progress" for task(s) you are starting now (string) +{% endif %} -{% if skills_enabled -%} +{% if skills_enabled %} # Skill Usage: If a skill in the catalog clearly matches the issue being investigated, fetch it using the `fetch_skill` tool before diving into other tools. Only fetch skills that are relevant to the specific issue — do not fetch skills speculatively or "just in case". If no skill matches, skip this step and investigate directly with available tools. After fetching a skill, read the content returned in the tool's data field and follow its steps. Skill content takes priority over general investigation steps. -{%- endif %} +{% endif %} +{% if todowrite_enabled %} # Task Execution ... Complete ALL tasks before providing a final answer. -{% else %} -{% if skills_enabled -%} -# Skill Usage: -... -{%- endif %} {% endif %}🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@holmes/plugins/prompts/generic_ask.jinja2` around lines 22 - 67, The template duplicates the "# Skill Usage" block inside both branches of the conditional (the two identical blocks controlled by the Jinja2 if/else around skills_enabled), so hoist that block out of the conditional: remove the duplicate blocks inside the branches and place a single "# Skill Usage" block before the `{% if skills_enabled -%}` (or immediately above the todowrite_enabled conditional if that's the intended variable), preserving the original Jinja whitespace/control tags and the wording of the block so it renders once for both paths; update references to `skills_enabled` (or `todowrite_enabled` if used elsewhere) only for the remaining conditional content, not for the hoisted block.
🧹 Nitpick comments (1)
holmes/plugins/prompts/generic_ask.jinja2 (1)
129-137: ⚡ Quick win
# Task Managementoverlaps with# Task Executionand# Investigation Phases.Several rules here restate guidance from the earlier sections under the same
todowrite_enabledblock:
- "Update task status (pending → in_progress → completed) as you work." duplicates line 34.
- "Create your investigation plan as your FIRST tool call, marking initial tasks as in_progress." duplicates the intent of lines 17–20 and 47.
Only "FIRST tool call" emphasis, "call other tools in parallel with TodoWrite", and "add discovered steps to your task list" are net-new. Consider folding those one-off rules into
# Task Execution/# Investigation Phasesand dropping the# Task Managementheading entirely — it'll keep this template aligned with the deduplication goal stated in the PR description.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@holmes/plugins/prompts/generic_ask.jinja2` around lines 129 - 137, The "# Task Management" block under the todowrite_enabled branch duplicates guidance already present in "# Task Execution" and "# Investigation Phases"; remove the redundant lines ("Update task status..." and "Create your investigation plan as your FIRST tool call...") and either delete the entire "# Task Management" heading or fold only the unique rules ("FIRST tool call" emphasis, "call other tools in parallel with TodoWrite", and "add discovered steps to your task list") into the existing "# Task Execution" or "# Investigation Phases" sections so the template keeps todowrite_enabled but avoids duplication.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@holmes/plugins/prompts/generic_ask.jinja2`:
- Line 144: Replace the ambiguous phrase "test a cluster level tool" in the
generic_ask.jinja2 prompt with clearer, instructional wording such as "use a
cluster-scoped tool" or "first call a cluster-scoped lookup" so the LLM
understands to perform a cluster-level lookup rather than verify the tool;
locate the sentence containing "When searching for resources in specific
namespaces, test a cluster level tool to find the resource(s) and identify what
namespace they are part of." and update it accordingly.
---
Outside diff comments:
In `@holmes/plugins/prompts/generic_ask.jinja2`:
- Around line 22-67: The template duplicates the "# Skill Usage" block inside
both branches of the conditional (the two identical blocks controlled by the
Jinja2 if/else around skills_enabled), so hoist that block out of the
conditional: remove the duplicate blocks inside the branches and place a single
"# Skill Usage" block before the `{% if skills_enabled -%}` (or immediately
above the todowrite_enabled conditional if that's the intended variable),
preserving the original Jinja whitespace/control tags and the wording of the
block so it renders once for both paths; update references to `skills_enabled`
(or `todowrite_enabled` if used elsewhere) only for the remaining conditional
content, not for the hoisted block.
---
Nitpick comments:
In `@holmes/plugins/prompts/generic_ask.jinja2`:
- Around line 129-137: The "# Task Management" block under the todowrite_enabled
branch duplicates guidance already present in "# Task Execution" and "#
Investigation Phases"; remove the redundant lines ("Update task status..." and
"Create your investigation plan as your FIRST tool call...") and either delete
the entire "# Task Management" heading or fold only the unique rules ("FIRST
tool call" emphasis, "call other tools in parallel with TodoWrite", and "add
discovered steps to your task list") into the existing "# Task Execution" or "#
Investigation Phases" sections so the template keeps todowrite_enabled but
avoids duplication.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: b3d042a2-de91-485c-9f10-2211f6c06083
📒 Files selected for processing (2)
holmes/plugins/prompts/generic_ask.jinja2holmes/plugins/prompts/investigation_procedure.jinja2
💤 Files with no reviewable changes (1)
- holmes/plugins/prompts/investigation_procedure.jinja2
…erlap - Hoist Skill Usage block out of todowrite_enabled if/else so it renders once for both paths instead of being duplicated - Fold unique Task Management bullets (parallel tools, discovered steps) into Task Execution section; remove the redundant Task Management heading entirely - Fix "test a cluster level tool" → "first call a cluster-level tool" for clearer instructional wording Signed-off-by: Claude <noreply@anthropic.com>
….1:43595/git/HolmesGPT/holmesgpt into claude/simplify-jinja2-prompts-1A0Jz
…le block
- Revert cosmetic rewording (capitalization, bullet merging) in the
"In general" and "Kubernetes problems" sections to match master
exactly, reducing diff noise
- Move Skill Usage above the todowrite block so all todowrite content
is in a single {% if todowrite_enabled %} block
Signed-off-by: Claude <noreply@anthropic.com>
…ructions Reverts all semantic changes to investigation_procedure and general instructions content. The only changes vs master are now structural: - Files inlined (no separate partials) - Skill Usage hoisted above todowrite block (deduped) - AI safety deleted - Filler intro lines removed - Stray newrelic fragment removed from _fetch_logs - Duplicate skill-fetching bullets removed from "In general" All prompt text that existed on master is preserved word-for-word. Signed-off-by: Claude <noreply@anthropic.com>
Master updated one line in investigation_procedure.jinja2 (parallel tool call example wording). Applied that change to our inlined copy in generic_ask.jinja2 and kept the file deleted. Signed-off-by: Claude <noreply@anthropic.com>
5 iters of eval 259_loki_historical_logs_pod_deleted_docker on each side, opus-4.6 via OpenRouter: baseline cf6ddb7 (2026-04-30): 78.6s $0.407 7.8 turns 21.8 tools 194K tk current 1b61fe3 (2026-05-15): 89.4s $0.447 10.6 turns 22.0 tools 259K tk +13.7% +9.8% +35.9% +0.9% +33.6% Pass rate unchanged (5/5 vs 5/5). Tool count flat. Output tokens flat. Total/cached input tokens up ~33-41% (z > 3.8, highly significant). Turns up ~36% (z > 3) — same tool work, spread over more serialized rounds. That's the signature of a bigger system prompt driving more incremental multi-step exploration, matching the PR-#1970 hypothesis from the earlier single-iter ci-benchmark vs master-CI comparison. Includes the run_sweep.sh helper used to produce these numbers. Signed-off-by: Claude <noreply@anthropic.com>
n=5 sweep across six candidate fixes pinpointed PR #2040 as the cause. Fix-AD (full revert of PR #2040: restore the deleted '# If investigating Kubernetes problems' section in generic_ask.jinja2 AND delete holmes/plugins/skills/builtin/kubernetes-troubleshooting/) restores baseline behavior: metric baseline current fix-AD recovery time 78.6s 89.4s 83.0s 0.59 turns 7.8 10.6 8.6 0.71 total tokens 194K 259K 203K 0.86 cached 156K 220K 166K 0.85 cost $0.41 $0.45 $0.40 fully (slightly cheaper than baseline) fetch_skill (Σ5) 0 29 0 — z-score vs baseline drops from 5.11 (turns) on current to 1.79 on fix-AD — i.e. inside the noise floor. Per-iter spread on fix-AD sits entirely inside the baseline range, no outliers. Per-iter inspection showed the smoking gun: every post-#2040 iter called fetch_skill exactly once on turn 1 (pulling the kubernetes-troubleshooting skill). Baseline had zero such calls because the skill simply didn't exist. The user-prompt skill catalog block was identical between baseline and current, but only renders content when a skill is loaded — PR #2040 was what loaded one. Fix-A alone (prompt section only) recovered ~45% because the skill file still existed and still got listed. Fix-B (restoring the 'MUST use tools' one-liner from PR #1970) didn't help at all. Fix-AB and fix-AC were equivalent to fix-A. Only fix-AD closes the gap. The recommended PR diff is in analysis/2026-05-15-regression/the-fix.md. Raw 5-iter reports for all six conditions are committed alongside. Signed-off-by: Claude <noreply@anthropic.com>
Summary
_global_instructions.jinja2had zero callers)generic_ask.jinja2(e.g. "Use conversation history to maintain continuity", duplicate skill-fetching bullets, verbose task management section)investigation_procedure.jinja2from 221 → 50 lines (77% reduction), fixing a self-contradictory phase evaluation bug where "yes" to "Do I have enough information?" incorrectly triggered continuation instead of completion_fetch_logs.jinja2Files deleted (inlined or orphan)
_general_instructions.jinja2(1 caller → inlined intogeneric_ask.jinja2)_permission_errors.jinja2(1 caller, 6 lines → inlined)_runbooks_instructions.jinja2/_skills_instructions.jinja2(mutually exclusive paths → inlined)_runbook_instructions.jinja2/_skill_instructions.jinja2(1 caller → inlined intobase_user_prompt.jinja2)_current_date_time.jinja2(1 caller, 2 lines → inlined)_global_instructions.jinja2(orphan — zero callers anywhere)Files kept separate
_ai_safety.jinja2— partner-mandated, default-disabled_toolsets_instructions.jinja2,_fetch_logs.jinja2— substantial data-driven logic with own tests_default_log_prompt.jinja2— genuinely reused across 3 branches in_fetch_logs.jinja2investigation_procedure.jinja2— separate logical unitTest plan
https://claude.ai/code/session_01HVq6giayp3P65pUpjJMkLo
Generated by Claude Code
Summary by CodeRabbit
New Features
Refactor
Removed