Skip to content

Simplify jinja2 prompts: inline, deduplicate, fix logic - #1970

Merged
aantn merged 20 commits into
masterfrom
claude/simplify-jinja2-prompts-1A0Jz
May 14, 2026
Merged

aantn merged 20 commits into
masterfrom
claude/simplify-jinja2-prompts-1A0Jz

Conversation

@aantn

@aantn aantn commented Apr 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Flatten prompt include graph from 15 jinja2 files to 9 by inlining single-use partials and deleting 1 orphan (_global_instructions.jinja2 had zero callers)
  • Remove filler and duplicated instructions from generic_ask.jinja2 (e.g. "Use conversation history to maintain continuity", duplicate skill-fetching bullets, verbose task management section)
  • Rewrite investigation_procedure.jinja2 from 221 → 50 lines (77% reduction), fixing a self-contradictory phase evaluation bug where "yes" to "Do I have enough information?" incorrectly triggered continuation instead of completion
  • Remove stray newrelic fragment with typo from _fetch_logs.jinja2
  • ~40% system prompt token reduction across all configurations

Files deleted (inlined or orphan)

  • _general_instructions.jinja2 (1 caller → inlined into generic_ask.jinja2)
  • _permission_errors.jinja2 (1 caller, 6 lines → inlined)
  • _runbooks_instructions.jinja2 / _skills_instructions.jinja2 (mutually exclusive paths → inlined)
  • _runbook_instructions.jinja2 / _skill_instructions.jinja2 (1 caller → inlined into base_user_prompt.jinja2)
  • _current_date_time.jinja2 (1 caller, 2 lines → inlined)
  • _global_instructions.jinja2 (orphan — zero callers anywhere)

Files kept separate

  • _ai_safety.jinja2 — partner-mandated, default-disabled
  • _toolsets_instructions.jinja2, _fetch_logs.jinja2 — substantial data-driven logic with own tests
  • _default_log_prompt.jinja2 — genuinely reused across 3 branches in _fetch_logs.jinja2
  • investigation_procedure.jinja2 — separate logical unit

Test plan

  • All 34 prompt-specific tests pass
  • All 2078 non-LLM tests pass
  • E2E rendered prompt comparison (CLI + server paths) against master: system prompt identical except intended removals, user prompts identical except cosmetic whitespace
  • LLM eval regression tests in CI

https://claude.ai/code/session_01HVq6giayp3P65pUpjJMkLo


Generated by Claude Code

Summary by CodeRabbit

  • New Features

    • Skill-based investigation prioritized when available; expanded TodoWrite execution and task-management workflow.
  • Refactor

    • Streamlined investigation procedure and prompt flow; stronger guidance on iterative tool use, namespace/cluster discovery, and deeper kubectl exploration.
  • Removed

    • AI safety guardrails prompt and related tests.
    • New Relic-specific log-fetching guidance.

Review Change Stack

claude and others added 6 commits April 13, 2026 07:29
Flattens the prompt include graph from 15 files to 8 by inlining
partials that were used in only one place, plus deletes one orphan.

Inlined into generic_ask.jinja2:
- _general_instructions.jinja2 (1 caller)
- _permission_errors.jinja2    (1 caller, 6 lines)
- _runbooks_instructions.jinja2 (mutually exclusive with investigation_procedure)

Inlined into base_user_prompt.jinja2:
- _runbook_instructions.jinja2 (1 caller)
- _current_date_time.jinja2    (1 caller, 2 lines)

Inlined into investigation_procedure.jinja2:
- _runbooks_instructions.jinja2

Inlined 3x into _fetch_logs.jinja2:
- _default_log_prompt.jinja2 (only reused inside this one parent)

Deleted as orphan (no callers anywhere):
- _global_instructions.jinja2

Kept as separate files (entry points or substantial logical units):
- generic_ask.jinja2, base_user_prompt.jinja2,
  conversation_history_compaction.jinja2, _ticket_additions.jinja2
- _ai_safety.jinja2 (partner-mandated, kept discoverable)
- _toolsets_instructions.jinja2, _fetch_logs.jinja2,
  investigation_procedure.jinja2 (sizeable data-driven units)

No semantic changes. Rendered output is byte-equivalent except for
two stripped blank lines and one trailing space (all cosmetic).
All 34 existing prompt tests pass unchanged.

Signed-off-by: Claude <noreply@anthropic.com>
…mment to runbook block

Reverts the inlining of _default_log_prompt.jinja2 into _fetch_logs.jinja2
since the 3x duplication made things worse, not better. This file is
genuinely reused across 3 elif branches (coralogix, k8s_base, opensearch).

Adds a Jinja comment block to base_user_prompt.jinja2 explaining what
the runbook selection sections/available pattern does — it dynamically
builds a priority-ordered list from whichever of 3 context variables
(runbook_catalog, custom_instructions, global_instructions) are non-empty.

Before/after rendered prompt comparison:
- System prompt: identical
- CLI user prompt: 1 cosmetic blank line diff only
- Server user prompt: 1 cosmetic blank line diff only
All 34 prompt tests pass.

Signed-off-by: Claude <noreply@anthropic.com>
Master renamed runbooks→skills across all prompts.
Conflicts resolved by applying the skill rename to our inlined content:
- generic_ask.jinja2: runbooks_enabled→skills_enabled, runbook→skill text
- investigation_procedure.jinja2: inlined _skills_instructions content
- base_user_prompt.jinja2: inlined _skill_instructions content
- Deleted _skill_instructions.jinja2 and _skills_instructions.jinja2
  (single-use, consistent with our inlining approach)
- Deleted _general_instructions.jinja2 (keep our deletion)

Signed-off-by: Claude <noreply@anthropic.com>
generic_ask.jinja2:
- Remove "Use conversation history to maintain continuity" (filler)
- Remove "Whenever possible you MUST first use tools" (redundant with
  general instructions section)
- Remove "Ask for multiple tool calls at the same time" (duplicated in
  investigation_procedure and task management)
- Merge "run as many tools...do so repeatedly" into single bullets
- Remove duplicate skill-fetching bullets (already in Skill Usage block)
- Collapse verbose Task Management section (14 lines → 5)
- Remove "You are able to make tool calls" (the LLM already knows)

investigation_procedure.jinja2 (221 → 50 lines, 77% reduction):
- Fix self-contradictory phase evaluation: "yes" to "Do I have enough
  information?" was incorrectly triggering continuation. Now each
  question has explicit IF NO/IF YES direction.
- Remove duplicate TASK COMPLETION ENFORCEMENT + ENFORCEMENT RULES
  (said the same thing twice with different headers)
- Remove duplicate evaluation question lists (listed identically at
  lines 112 and 128)
- Remove ASCII checklist art and status update examples
- Remove triple VIOLATION CONSEQUENCES blocks
- Collapse FINAL REVIEW PHASE from 23 bullets to one sentence
- Remove INVESTIGATION PHASE TRANSITION EXAMPLES (redundant)
- Remove trivial dependency examples (covered by the parallel example)

_fetch_logs.jinja2:
- Remove stray "Logs from newrelic" fragment with typo that was
  incorrectly placed inside the k8s_yaml_ts elif branch

System prompt reduction: ~40% fewer chars across all configurations.

Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@coderabbitai

coderabbitai Bot commented Apr 29, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

This PR removes the AI safety prompt component and its associated guardrails template from the Holmes system. The PromptComponent.AI_SAFETY enum member is deleted, default-disabled components are no longer enforced, and the _ai_safety.jinja2 template is removed. Prompt templates are restructured to adjust behavior and investigation procedures without the safety guidance.

Changes

Removal of AI Safety Component and Prompt Restructuring

Layer / File(s) Summary
Enum and Config Changes
holmes/core/prompt.py
PromptComponent.AI_SAFETY enum member is removed; DISABLED_BY_DEFAULT changes from {PromptComponent.AI_SAFETY} to an empty set; ai_safety_enabled is removed from template context in build_system_prompt().
Safety Template Removal
holmes/plugins/prompts/_ai_safety.jinja2
Entire safety and guardrails template (content harms, jailbreak/UPIA/XPIA handling, IP/copyright restrictions, ungrounded response rules) is deleted.
Prompt Restructuring
holmes/plugins/prompts/generic_ask.jinja2
Mandatory "use tools first" instruction is removed; new conditional general_instructions_enabled block adds repeated investigation cycles, five-whys root-cause rules, uncertainty/hedging guidance, stricter Kubernetes/log expectations, and cluster-level namespace discovery. When todowrite_enabled, adds dedicated Task Management section.
Investigation Procedure Simplification
holmes/plugins/prompts/investigation_procedure.jinja2
Multi-section enforcement-heavy protocol is condensed to shorter control-flow; keeps clarification-first and TodoWrite kickoff; replaces strict phase transitions with brief reassessment loop; conditionally fetches and prioritizes skill steps when skills_enabled; permits in_progress task status on immediate starts.
Minor Log Fetch Update
holmes/plugins/prompts/_fetch_logs.jinja2
New Relic log tool guidance removed from the k8s_yaml_ts enabled branch.
Test Cleanup
tests/core/test_prompt.py, tests/test_ai_safety_prompt.py
Deleted TestIsComponentEnabled test cases for PromptComponent.AI_SAFETY behavior and DISABLED_BY_DEFAULT defaults; removed entire tests/test_ai_safety_prompt.py module verifying safety prompt inclusion in system prompts.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

  • HolmesGPT/holmesgpt#823: Directly related counterpart that introduced the AI safety template and its inclusions—this PR removes what that PR added.
  • HolmesGPT/holmesgpt#1452: Related historical change that introduced PromptComponent and conditional AI safety inclusion logic that is now being removed.
  • HolmesGPT/holmesgpt#1711: Modifies the same prompt templates (generic_ask.jinja2 and investigation_procedure.jinja2) to adjust investigation behavior and guidance.

Suggested labels

codex

Suggested reviewers

  • arikalon1
  • Sheeproid
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title accurately describes the primary intent: simplifying Jinja2 prompts through inlining partials, deduplicating instructions, and fixing logic bugs. It matches the changeset's main objectives.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Apr 29, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 808a914
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a0634e9cf831300080c08a0
😎 Deploy Preview https://deploy-preview-1970--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (3)
holmes/plugins/prompts/base_user_prompt.jinja2 (1)

28-29: Minor: Indentation may not render as expected in Markdown.

Line 29 indents continuation lines with 3 spaces ('\n '). In standard Markdown, leading spaces don't create indentation for paragraph text—they're typically stripped. If the goal is visual indentation in the rendered prompt, this may not achieve the desired effect depending on how the LLM processes the prompt.

If indentation is critical, consider using a Markdown-native approach (like blockquotes with >). Otherwise, if this is just for human readability in logs/debugging, it's fine as-is.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/prompts/base_user_prompt.jinja2` around lines 28 - 29, The
current template replaces newlines with '\n   ' (see variable content and
sec.content and the replace call) which may not produce visible indentation in
Markdown; update the replacement to use a Markdown-native pattern such as '\n> '
(blockquote) or another explicit Markdown construct so continuation lines render
as indented in LLM/Markdown contexts — change the replace invocation on content
to insert a blockquote prefix (or other Markdown syntax) instead of three
spaces.
holmes/plugins/prompts/generic_ask.jinja2 (2)

16-24: Duplicated skill instructions between generic_ask.jinja2 and investigation_procedure.jinja2.

Lines 16-24 contain skill usage instructions that are nearly identical to lines 8-16 in investigation_procedure.jinja2. Since investigation_procedure.jinja2 is included at line 14 when todowrite_enabled is true, the else branch here (lines 15-25) handles the !todowrite_enabled && skills_enabled case. However, when todowrite_enabled is true, the included investigation_procedure.jinja2 also has its own skill instructions (guarded by skills_enabled).

This appears intentional—one path for todowrite mode, one for non-todowrite mode—but consider extracting the skill instructions into a single partial to avoid maintaining two copies.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/prompts/generic_ask.jinja2` around lines 16 - 24, The
duplicate "Skill Usage" block appears in generic_ask.jinja2 and
investigation_procedure.jinja2 (both guarded by skills_enabled and influenced by
todowrite_enabled); extract that repeated block into a single partial (e.g.,
_skill_usage.jinja2) and replace the inline blocks in both templates with an
include of the partial, keeping the original conditional guards (skills_enabled
and todowrite_enabled) intact so behavior for the todowrite_enabled and
non-todowrite paths remains the same and references to fetch_skill and reading
the tool's data field are preserved.

31-63: Consider consolidating general_instructions_enabled blocks.

The general_instructions_enabled conditional content is split across three separate blocks (lines 31-63, 79-88, and 100-106), interleaved with other conditionals. While this works, it may complicate future maintenance.

If the ordering of sections is flexible, consolidating these into a contiguous block would improve readability.

Also applies to: 79-88, 100-106

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/prompts/generic_ask.jinja2` around lines 31 - 63, The template
splits the general_instructions_enabled conditional into multiple non-contiguous
blocks which complicates maintenance; consolidate all occurrences of the
general_instructions_enabled conditional in
holmes/plugins/prompts/generic_ask.jinja2 into a single contiguous block
containing the "In general" section, the cluster_name sub-block, and the
Kubernetes investigation section, preserving all internal content and any nested
conditionals (e.g., the cluster_name check) and ensuring the relative ordering
of those internal paragraphs remains correct so behavior does not change; remove
the now-redundant separate general_instructions_enabled blocks at lines
referenced in the PR and run a quick template-render smoke test to confirm no
logic/regression changes.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@holmes/plugins/prompts/investigation_procedure.jinja2`:
- Around line 33-39: The conditional block in investigation_procedure.jinja2
uses hardcoded step numbers causing inconsistent numbering when skills_enabled
is false; update the template so numbering is generated dynamically or use
unnumbered bullets: either (A) replace manual numbers with a Jinja counter
(e.g., set a local counter and increment it when emitting each step so the final
"3. Execute ALL tasks" uses the computed value) or (B) convert steps to
unnumbered list items so the conditional branch doesn't break sequence;
reference the skills_enabled condition and the step text "Create initial
TodoWrite task list..." and "Execute ALL tasks" when applying the fix.

---

Nitpick comments:
In `@holmes/plugins/prompts/base_user_prompt.jinja2`:
- Around line 28-29: The current template replaces newlines with '\n   ' (see
variable content and sec.content and the replace call) which may not produce
visible indentation in Markdown; update the replacement to use a Markdown-native
pattern such as '\n> ' (blockquote) or another explicit Markdown construct so
continuation lines render as indented in LLM/Markdown contexts — change the
replace invocation on content to insert a blockquote prefix (or other Markdown
syntax) instead of three spaces.

In `@holmes/plugins/prompts/generic_ask.jinja2`:
- Around line 16-24: The duplicate "Skill Usage" block appears in
generic_ask.jinja2 and investigation_procedure.jinja2 (both guarded by
skills_enabled and influenced by todowrite_enabled); extract that repeated block
into a single partial (e.g., _skill_usage.jinja2) and replace the inline blocks
in both templates with an include of the partial, keeping the original
conditional guards (skills_enabled and todowrite_enabled) intact so behavior for
the todowrite_enabled and non-todowrite paths remains the same and references to
fetch_skill and reading the tool's data field are preserved.
- Around line 31-63: The template splits the general_instructions_enabled
conditional into multiple non-contiguous blocks which complicates maintenance;
consolidate all occurrences of the general_instructions_enabled conditional in
holmes/plugins/prompts/generic_ask.jinja2 into a single contiguous block
containing the "In general" section, the cluster_name sub-block, and the
Kubernetes investigation section, preserving all internal content and any nested
conditionals (e.g., the cluster_name check) and ensuring the relative ordering
of those internal paragraphs remains correct so behavior does not change; remove
the now-redundant separate general_instructions_enabled blocks at lines
referenced in the PR and run a quick template-render smoke test to confirm no
logic/regression changes.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 27b48bf1-c90e-4deb-84f9-95df0b25c14f

📥 Commits

Reviewing files that changed from the base of the PR and between d29f143 and 8438589.

📒 Files selected for processing (10)
  • holmes/plugins/prompts/_current_date_time.jinja2
  • holmes/plugins/prompts/_fetch_logs.jinja2
  • holmes/plugins/prompts/_general_instructions.jinja2
  • holmes/plugins/prompts/_global_instructions.jinja2
  • holmes/plugins/prompts/_permission_errors.jinja2
  • holmes/plugins/prompts/_skill_instructions.jinja2
  • holmes/plugins/prompts/_skills_instructions.jinja2
  • holmes/plugins/prompts/base_user_prompt.jinja2
  • holmes/plugins/prompts/generic_ask.jinja2
  • holmes/plugins/prompts/investigation_procedure.jinja2
💤 Files with no reviewable changes (7)
  • holmes/plugins/prompts/_global_instructions.jinja2
  • holmes/plugins/prompts/_permission_errors.jinja2
  • holmes/plugins/prompts/_skills_instructions.jinja2
  • holmes/plugins/prompts/_skill_instructions.jinja2
  • holmes/plugins/prompts/_general_instructions.jinja2
  • holmes/plugins/prompts/_current_date_time.jinja2
  • holmes/plugins/prompts/_fetch_logs.jinja2

Comment thread holmes/plugins/prompts/investigation_procedure.jinja2 Outdated
…phases

Fixes inconsistent step numbering when skills_enabled is false —
steps jumped from 1 to 3. Switching to bullets avoids fragile
conditional numbering entirely.

Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (2)
holmes/plugins/prompts/investigation_procedure.jinja2 (2)

34-34: Drop the repeated skill-first bullet.

Line 34 restates the same rule already given in Lines 10-12. Since this template is on the hot path and this PR is explicitly reducing prompt size, keeping the dedicated "Skill Usage" section and removing the duplicate would be cleaner.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/prompts/investigation_procedure.jinja2` at line 34, The
template contains a duplicated "skill-first" instruction ("If a skill in the
catalog clearly matches the issue, fetch it first. Otherwise, skip.") that
repeats the rule already stated in the "Skill Usage" section; remove this
repeated bullet so only the dedicated "Skill Usage" section contains that
guidance, leaving other prompt content unchanged and ensuring no other
references rely on the duplicate line.

3-3: Keep the TodoWrite trigger scoped to investigations.

generic_ask.jinja2:90-98 still says TodoWrite is for investigations requiring multiple steps, but Line 3 broadens that to any multi-step question. That scope drift will push routine multi-part asks into unnecessary TodoWrite calls.

Suggested wording
-For multi-step questions, start by calling TodoWrite with a `todos` parameter containing an array of task objects. Each task must have:
+For multi-step investigations, start by calling TodoWrite with a `todos` parameter containing an array of task objects. Each task must have:
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/prompts/investigation_procedure.jinja2` at line 3, The
TodoWrite trigger description should be restricted to investigation flows only:
update the sentence in investigation_procedure.jinja2 (current "For multi-step
questions...") to specify "For investigations that require multiple steps" and
reference the TodoWrite trigger used by investigations; ensure consistency with
generic_ask.jinja2's guidance around TodoWrite (lines referenced as
generic_ask.jinja2:90-98) so ordinary multi-part questions are not routed to
TodoWrite—keep the TodoWrite scope limited to investigation contexts and mirror
that exact phrasing wherever TodoWrite is documented.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@holmes/plugins/prompts/investigation_procedure.jinja2`:
- Around line 3-7: The doc text in investigation_procedure.jinja2 conflicts with
examples: change the guidance for new task status to match the examples by
stating that initial tasks should use "in_progress" instead of "pending" (i.e.,
update the TodoWrite description for the todos parameter to require `status`:
"in_progress" for newly started tasks), and make the same change wherever the
initial-status rule is documented (e.g., the example payload in this template
and the related generic_ask.jinja2 examples) so all references to initial task
status are consistent.

---

Nitpick comments:
In `@holmes/plugins/prompts/investigation_procedure.jinja2`:
- Line 34: The template contains a duplicated "skill-first" instruction ("If a
skill in the catalog clearly matches the issue, fetch it first. Otherwise,
skip.") that repeats the rule already stated in the "Skill Usage" section;
remove this repeated bullet so only the dedicated "Skill Usage" section contains
that guidance, leaving other prompt content unchanged and ensuring no other
references rely on the duplicate line.
- Line 3: The TodoWrite trigger description should be restricted to
investigation flows only: update the sentence in investigation_procedure.jinja2
(current "For multi-step questions...") to specify "For investigations that
require multiple steps" and reference the TodoWrite trigger used by
investigations; ensure consistency with generic_ask.jinja2's guidance around
TodoWrite (lines referenced as generic_ask.jinja2:90-98) so ordinary multi-part
questions are not routed to TodoWrite—keep the TodoWrite scope limited to
investigation contexts and mirror that exact phrasing wherever TodoWrite is
documented.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: da3a8d40-2e20-4806-bf92-1a1b5c423031

📥 Commits

Reviewing files that changed from the base of the PR and between 8438589 and f208aac.

📒 Files selected for processing (1)
  • holmes/plugins/prompts/investigation_procedure.jinja2

Comment thread holmes/plugins/prompts/investigation_procedure.jinja2 Outdated
claude added 2 commits April 29, 2026 17:46
- Change "multi-step questions" → "multi-step investigations" to keep
  TodoWrite scoped to investigation flows, not routine multi-part asks
- Remove duplicate skill-first bullet from Investigation Phases (already
  covered by the Skill Usage section above)
- Clarify status field: "pending" for queued tasks or "in_progress" for
  tasks being started now, matching the example payload

Signed-off-by: Claude <noreply@anthropic.com>
Conflicts from PR #1905 (also refactored prompts). Resolved by keeping
our simplified versions in both generic_ask.jinja2 and
investigation_procedure.jinja2.

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented May 1, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 #5 · Run @ __3e7b7b1__ (#25883778899) — May 14, 20:34 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 3e7b7b1 on branch claude/simplify-jinja2-prompts-1A0Jz

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 29.3s 4 9 $0.2048 72,721 70,838 20,547 1,883 942 49,367 21,471 96 —
✅ 101_loki_historical_logs_pod_deleted 49.6s 6 13 $0.2739 120,942 118,102 24,181 2,840 704 92,861 25,241 132 —
✅ 112_find_pvcs_by_uuid 22.3s 4 5 $0.1754 68,860 67,552 18,679 1,308 531 48,529 19,023 95 —
✅ 12_job_crashing 30.7s 5 10 $0.2168 92,772 90,982 20,484 1,790 553 68,548 22,434 107 —
✅ 176_network_policy_blocking_traffic_no_skills 45.4s 5 14 $0.2585 101,671 99,039 23,776 2,632 1,193 73,827 25,212 130 —
✅ 227_count_configmaps_per_namespace[0] 19.5s 4 9 $0.1725 67,818 66,662 18,433 1,156 561 47,282 19,380 78 —
✅ 243_pod_names_contain_service 30.1s 5 8 $0.1990 85,875 84,147 19,345 1,728 625 64,240 19,907 95 —
✅ 24_misconfigured_pvc 29.1s 5 10 $0.2016 88,530 86,845 19,407 1,685 556 66,445 20,400 98 —
✅ 43_current_datetime_from_prompt 6.1s 1 — $0.0968 14,858 14,647 14,647 211 211 0 14,647 157 —
✅ 51_logs_summarize_errors 20.3s 4 5 $0.1645 67,683 66,699 18,514 984 319 48,173 18,526 55 —
✅ 61_exact_match_counting 8.1s 2 1 $0.1077 29,953 29,701 15,045 252 184 14,646 15,055 71 —
Total 26.4s avg 4.1 avg 8.4 avg $2.0715 811,683 795,214 24,181 16,469 1,193 573,918 221,296 1,114 —

Benchmark comparison unavailable: No ci-benchmark experiments found

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No ci-benchmark experiments found

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #4 · Run @ __5078890__ (#25883365393) — May 14, 20:26 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 5078890 on branch claude/simplify-jinja2-prompts-1A0Jz

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 29.2s 4 9 $0.2005 72,145 70,392 20,381 1,753 935 49,048 21,344 85 —
✅ 101_loki_historical_logs_pod_deleted 34.0s 4 9 $0.2103 73,350 71,089 20,124 2,261 841 50,406 20,683 115 —
✅ 112_find_pvcs_by_uuid 15.9s 3 3 $0.1571 53,030 52,160 18,997 870 507 33,152 19,008 85 —
✅ 12_job_crashing 35.8s 5 13 $0.2301 99,791 97,820 22,463 1,971 671 74,742 23,078 92 —
✅ 176_network_policy_blocking_traffic_no_skills 40.5s 6 15 $0.2739 125,519 123,039 25,233 2,480 817 96,640 26,399 97 —
✅ 227_count_configmaps_per_namespace[0] 21.1s 4 9 $0.1708 66,707 65,417 17,853 1,290 686 46,865 18,552 75 —
✅ 243_pod_names_contain_service 28.9s 4 8 $0.1920 69,973 68,205 19,319 1,768 837 48,303 19,902 104 —
✅ 24_misconfigured_pvc 33.6s 5 12 $0.2230 92,252 90,134 20,739 2,118 661 68,131 22,003 62 —
✅ 43_current_datetime_from_prompt 4.6s 1 — $0.0941 14,706 14,591 14,591 115 115 0 14,591 75 —
✅ 51_logs_summarize_errors 20.7s 4 5 $0.1681 67,709 66,574 18,507 1,135 475 48,055 18,519 55 —
✅ 61_exact_match_counting 9.4s 2 1 $0.1077 29,868 29,607 15,007 261 205 14,590 15,017 91 —
Total 24.9s avg 3.8 avg 8.4 avg $2.0275 765,050 749,028 25,233 16,022 935 529,932 219,096 936 —

Benchmark comparison unavailable: No ci-benchmark experiments found

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No ci-benchmark experiments found

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #3 · Run @ __877765e__ (#25883105057) — May 14, 20:19 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 877765e on branch claude/simplify-jinja2-prompts-1A0Jz

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 27.9s 4 9 $0.1966 71,515 69,762 19,941 1,753 855 49,098 20,664 51 —
✅ 101_loki_historical_logs_pod_deleted 40.6s 5 10 $0.2417 97,139 94,639 22,092 2,500 873 71,471 23,168 120 —
✅ 112_find_pvcs_by_uuid 16.5s 3 3 $0.1574 53,175 52,307 19,044 868 503 33,252 19,055 78 —
✅ 12_job_crashing 31.7s 5 10 $0.2108 92,855 91,089 20,496 1,766 554 69,868 21,221 103 —
✅ 176_network_policy_blocking_traffic_no_skills 36.7s 5 12 $0.2494 98,908 96,573 23,045 2,335 818 71,260 25,313 104 —
✅ 227_count_configmaps_per_namespace[0] 23.3s 5 10 $0.1856 85,236 83,915 18,833 1,321 424 64,555 19,360 76 —
✅ 243_pod_names_contain_service 30.2s 4 8 $0.1914 70,129 68,404 19,378 1,725 846 48,441 19,963 59 —
✅ 24_misconfigured_pvc 30.4s 5 11 $0.2079 89,140 87,310 19,710 1,830 594 66,442 20,868 65 —
✅ 43_current_datetime_from_prompt 5.8s 1 — $0.0973 14,878 14,646 14,646 232 232 0 14,646 168 —
✅ 51_logs_summarize_errors 20.5s 4 5 $0.1678 68,025 66,939 18,635 1,086 423 48,292 18,647 55 —
✅ 61_exact_match_counting 8.1s 2 1 $0.1074 29,931 29,688 15,033 243 175 14,645 15,043 62 —
Total 24.7s avg 3.9 avg 7.9 avg $2.0134 770,931 755,272 23,045 15,659 873 537,324 217,948 941 —

Benchmark comparison unavailable: No ci-benchmark experiments found

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No ci-benchmark experiments found

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #2 · Run @ __20eca64__ (#25313442524) — May 4, 10:21 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 20eca64 on branch claude/simplify-jinja2-prompts-1A0Jz

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 27.9s 4 8 $0.1932 70,716 69,060 19,663 1,656 837 48,459 20,601 — —
✅ 101_loki_historical_logs_pod_deleted 46.2s 5 12 $0.2490 98,026 95,280 22,170 2,746 920 71,937 23,343 — —
✅ 112_find_pvcs_by_uuid 20.2s 4 4 $0.1679 68,084 66,931 18,356 1,153 550 48,563 18,368 — —
✅ 12_job_crashing 34.1s 5 13 $0.2380 99,762 97,654 22,354 2,108 740 73,532 24,122 — —
✅ 176_network_policy_blocking_traffic_no_skills 37.4s 5 12 $0.2405 99,423 97,111 23,102 2,312 902 73,713 23,398 — —
✅ 227_count_configmaps_per_namespace[0] 21.1s 4 9 $0.1720 67,921 66,738 18,472 1,183 561 47,639 19,099 — —
✅ 243_pod_names_contain_service 29.2s 4 8 $0.1905 69,942 68,227 19,292 1,715 804 48,364 19,863 — —
✅ 24_misconfigured_pvc 38.0s 6 15 $0.2412 110,715 108,399 21,412 2,316 709 85,654 22,745 — —
✅ 43_current_datetime_from_prompt 4.2s 1 — $0.0943 14,756 14,646 14,646 110 110 0 14,646 — —
✅ 51_logs_summarize_errors 21.1s 4 5 $0.1656 67,876 66,871 18,600 1,005 351 48,259 18,612 — —
✅ 61_exact_match_counting 8.1s 2 1 $0.1080 29,971 29,708 15,053 263 195 14,645 15,063 — —
Total 26.1s avg 4.0 avg 8.7 avg $2.0602 797,192 780,625 23,102 16,567 920 560,765 219,860 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #1 · Run @ __101b665__ (#25313124106) — May 4, 10:14 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 101b665 on branch claude/simplify-jinja2-prompts-1A0Jz

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 29.8s 5 8 $0.2008 86,086 84,389 19,757 1,697 850 64,048 20,341 — —
✅ 101_loki_historical_logs_pod_deleted 46.2s 6 11 $0.2609 120,135 117,672 23,165 2,463 835 92,888 24,784 — —
✅ 112_find_pvcs_by_uuid 24.1s 4 5 $0.1848 71,409 70,072 19,942 1,337 523 49,762 20,310 — —
✅ 12_job_crashing 29.4s 5 11 $0.2181 96,271 94,629 21,354 1,642 534 71,736 22,893 — —
✅ 176_network_policy_blocking_traffic_no_skills 43.1s 6 17 $0.2733 121,953 119,142 23,522 2,811 922 93,807 25,335 — —
✅ 227_count_configmaps_per_namespace[0] 24.7s 5 10 $0.1873 85,162 83,826 18,822 1,336 424 64,163 19,663 — —
✅ 243_pod_names_contain_service 34.1s 5 9 $0.2084 89,472 87,556 19,837 1,916 644 67,135 20,421 — —
✅ 24_misconfigured_pvc 33.6s 5 12 $0.2216 91,254 89,123 20,274 2,131 787 67,242 21,881 — —
✅ 43_current_datetime_from_prompt 4.7s 1 — $0.0951 14,777 14,628 14,628 149 149 0 14,628 — —
✅ 51_logs_summarize_errors 20.8s 4 5 $0.1655 67,993 67,012 18,688 981 320 48,312 18,700 — —
✅ 61_exact_match_counting 8.1s 2 1 $0.1077 29,923 29,667 15,030 256 188 14,627 15,040 — —
Total 27.1s avg 4.4 avg 8.9 avg $2.1236 874,435 857,716 23,522 16,719 922 633,720 223,996 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit 808a914 on branch claude/simplify-jinja2-prompts-1A0Jz

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 41.1s 6 11 $0.4927 124,934 122,592 24,024 2,342 1,010 57,612 64,980 108 —
✅ 101_loki_historical_logs_pod_deleted 50.8s 7 10 $0.4081 148,416 146,067 24,421 2,349 518 97,852 48,215 114 —
✅ 112_find_pvcs_by_uuid 17.4s 3 3 $0.2979 60,474 59,528 21,527 946 556 17,004 42,524 104 —
✅ 12_job_crashing 49.7s 7 15 $0.5858 171,585 168,934 27,812 2,651 545 93,163 75,771 114 —
✅ 176_network_policy_blocking_traffic_no_skills 46.4s 7 16 $0.4264 153,627 150,901 24,821 2,726 602 101,675 49,226 118 —
✅ 227_count_configmaps_per_namespace[0] 28.3s 6 10 $0.3199 113,455 111,977 20,599 1,478 515 72,425 39,552 76 —
✅ 243_pod_names_contain_service 33.8s 4 9 $0.3160 80,760 78,826 22,275 1,934 972 39,017 39,809 63 —
✅ 24_misconfigured_pvc 43.4s 6 14 $0.5062 124,792 122,182 23,645 2,610 660 55,775 66,407 156 —
✅ 43_current_datetime_from_prompt 5.5s 1 — $0.1101 17,132 16,969 16,969 163 163 0 16,969 121 —
✅ 51_logs_summarize_errors 24.0s 4 5 $0.3011 76,823 75,713 20,660 1,110 393 34,749 40,964 87 —
✅ 61_exact_match_counting 8.9s 2 1 $0.1227 34,564 34,336 17,358 228 172 16,968 17,368 60 —
Total 31.8s avg 4.8 avg 9.4 avg $3.8870 1,106,562 1,088,025 27,812 18,537 1,010 586,240 501,785 1,121 —

Benchmark comparison unavailable: No ci-benchmark experiments found

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No ci-benchmark experiments found

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/simplify-jinja2-prompts-1A0Jz -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/simplify-jinja2-prompts-1A0Jz -f markers=regression -f filter=

@github-actions

github-actions Bot commented May 1, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for a4d4b2b7 (built in 57s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:a4d4b2b7
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:a4d4b2b7 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:a4d4b2b7
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:a4d4b2b7
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:a4d4b2b7
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:a4d4b2b7 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:a4d4b2b7
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:a4d4b2b7

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:a4d4b2b7 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:a4d4b2b7

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:a4d4b2b7 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:a4d4b2b7

@github-actions

github-actions Bot commented May 1, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 15.33s 14.57s +5.2%
Warm Mean 6.51s 6.38s +2.1%
Warm Min 6.48s 6.31s
Warm Max 6.56s 6.42s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 34.15s 34.80s -1.9%
Warm Mean 8.64s 8.54s +1.2%
Warm Min 8.19s 8.25s
Warm Max 9.61s 8.88s

PR: a4d4b2b7 | Master: bd0c5e08 | Iterations: 5

aantn and others added 5 commits May 4, 2026 13:07
…te path

The "ask for multiple tool calls at the same time" instruction was only
covered in the todowrite/investigation path. Simple ask mode had no
explicit parallel-calling guidance.

Signed-off-by: Claude <noreply@anthropic.com>
Partners confirmed this is no longer required. Removes:
- _ai_safety.jinja2 template file
- AI_SAFETY enum value from PromptComponent
- ai_safety_enabled template context variable
- DISABLED_BY_DEFAULT set (now empty, kept for future use)
- tests/test_ai_safety_prompt.py (tested deleted content)
- 3 disabled-by-default tests in test_prompt.py (tested AI_SAFETY)

Signed-off-by: Claude <noreply@anthropic.com>
…_ask

Single caller, no tests load it directly. Content inlined verbatim
with no modifications. Reduces prompt files from 9 to 8.

Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
holmes/plugins/prompts/generic_ask.jinja2 (1)

22-67: 🛠️ Refactor suggestion | 🟠 Major | ⚡ Quick win

Skill Usage block is duplicated across both branches of todowrite_enabled.

Lines 22–30 and 59–67 are byte-identical. Given this PR's deduplication goal, hoist the # Skill Usage block out of the todowrite_enabled if/else so it renders once for both paths. This also prevents the two copies from drifting over time.

♻️ Proposed structure
 {% if todowrite_enabled %}
 If the user's question is ambiguous or lacks critical details, ask for clarification first. Do NOT create TodoWrite tasks for unclear questions.

 For multi-step investigations, start by calling TodoWrite with a `todos` parameter containing an array of task objects. Each task must have:
 - `id`: unique identifier (string)
 - `content`: specific task description (string)
 - `status`: "pending" for queued tasks or "in_progress" for task(s) you are starting now (string)
+{% endif %}

-{% if skills_enabled -%}
+{% if skills_enabled %}
 # Skill Usage:
 If a skill in the catalog clearly matches the issue being investigated, fetch it using the `fetch_skill` tool before diving into other tools.
 Only fetch skills that are relevant to the specific issue — do not fetch skills speculatively or "just in case".
 If no skill matches, skip this step and investigate directly with available tools.

 After fetching a skill, read the content returned in the tool's data field and follow its steps.
 Skill content takes priority over general investigation steps.
-{%- endif %}
+{% endif %}

+{% if todowrite_enabled %}
 # Task Execution
 ...
 Complete ALL tasks before providing a final answer.
-{% else %}
-{% if skills_enabled -%}
-# Skill Usage:
-...
-{%- endif %}
 {% endif %}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@holmes/plugins/prompts/generic_ask.jinja2` around lines 22 - 67, The template
duplicates the "# Skill Usage" block inside both branches of the conditional
(the two identical blocks controlled by the Jinja2 if/else around
skills_enabled), so hoist that block out of the conditional: remove the
duplicate blocks inside the branches and place a single "# Skill Usage" block
before the `{% if skills_enabled -%}` (or immediately above the
todowrite_enabled conditional if that's the intended variable), preserving the
original Jinja whitespace/control tags and the wording of the block so it
renders once for both paths; update references to `skills_enabled` (or
`todowrite_enabled` if used elsewhere) only for the remaining conditional
content, not for the hoisted block.
🧹 Nitpick comments (1)
holmes/plugins/prompts/generic_ask.jinja2 (1)

129-137: ⚡ Quick win

# Task Management overlaps with # Task Execution and # Investigation Phases.

Several rules here restate guidance from the earlier sections under the same todowrite_enabled block:

  • "Update task status (pending → in_progress → completed) as you work." duplicates line 34.
  • "Create your investigation plan as your FIRST tool call, marking initial tasks as in_progress." duplicates the intent of lines 17–20 and 47.

Only "FIRST tool call" emphasis, "call other tools in parallel with TodoWrite", and "add discovered steps to your task list" are net-new. Consider folding those one-off rules into # Task Execution / # Investigation Phases and dropping the # Task Management heading entirely — it'll keep this template aligned with the deduplication goal stated in the PR description.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@holmes/plugins/prompts/generic_ask.jinja2` around lines 129 - 137, The "#
Task Management" block under the todowrite_enabled branch duplicates guidance
already present in "# Task Execution" and "# Investigation Phases"; remove the
redundant lines ("Update task status..." and "Create your investigation plan as
your FIRST tool call...") and either delete the entire "# Task Management"
heading or fold only the unique rules ("FIRST tool call" emphasis, "call other
tools in parallel with TodoWrite", and "add discovered steps to your task list")
into the existing "# Task Execution" or "# Investigation Phases" sections so the
template keeps todowrite_enabled but avoids duplication.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@holmes/plugins/prompts/generic_ask.jinja2`:
- Line 144: Replace the ambiguous phrase "test a cluster level tool" in the
generic_ask.jinja2 prompt with clearer, instructional wording such as "use a
cluster-scoped tool" or "first call a cluster-scoped lookup" so the LLM
understands to perform a cluster-level lookup rather than verify the tool;
locate the sentence containing "When searching for resources in specific
namespaces, test a cluster level tool to find the resource(s) and identify what
namespace they are part of." and update it accordingly.

---

Outside diff comments:
In `@holmes/plugins/prompts/generic_ask.jinja2`:
- Around line 22-67: The template duplicates the "# Skill Usage" block inside
both branches of the conditional (the two identical blocks controlled by the
Jinja2 if/else around skills_enabled), so hoist that block out of the
conditional: remove the duplicate blocks inside the branches and place a single
"# Skill Usage" block before the `{% if skills_enabled -%}` (or immediately
above the todowrite_enabled conditional if that's the intended variable),
preserving the original Jinja whitespace/control tags and the wording of the
block so it renders once for both paths; update references to `skills_enabled`
(or `todowrite_enabled` if used elsewhere) only for the remaining conditional
content, not for the hoisted block.

---

Nitpick comments:
In `@holmes/plugins/prompts/generic_ask.jinja2`:
- Around line 129-137: The "# Task Management" block under the todowrite_enabled
branch duplicates guidance already present in "# Task Execution" and "#
Investigation Phases"; remove the redundant lines ("Update task status..." and
"Create your investigation plan as your FIRST tool call...") and either delete
the entire "# Task Management" heading or fold only the unique rules ("FIRST
tool call" emphasis, "call other tools in parallel with TodoWrite", and "add
discovered steps to your task list") into the existing "# Task Execution" or "#
Investigation Phases" sections so the template keeps todowrite_enabled but
avoids duplication.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: b3d042a2-de91-485c-9f10-2211f6c06083

📥 Commits

Reviewing files that changed from the base of the PR and between 20eca64 and 877765e.

📒 Files selected for processing (2)
  • holmes/plugins/prompts/generic_ask.jinja2
  • holmes/plugins/prompts/investigation_procedure.jinja2
💤 Files with no reviewable changes (1)
  • holmes/plugins/prompts/investigation_procedure.jinja2

Comment thread holmes/plugins/prompts/generic_ask.jinja2
aantn and others added 6 commits May 14, 2026 23:16
…erlap

- Hoist Skill Usage block out of todowrite_enabled if/else so it renders
  once for both paths instead of being duplicated
- Fold unique Task Management bullets (parallel tools, discovered steps)
  into Task Execution section; remove the redundant Task Management
  heading entirely
- Fix "test a cluster level tool" → "first call a cluster-level tool"
  for clearer instructional wording

Signed-off-by: Claude <noreply@anthropic.com>
…le block

- Revert cosmetic rewording (capitalization, bullet merging) in the
  "In general" and "Kubernetes problems" sections to match master
  exactly, reducing diff noise
- Move Skill Usage above the todowrite block so all todowrite content
  is in a single {% if todowrite_enabled %} block

Signed-off-by: Claude <noreply@anthropic.com>
…ructions

Reverts all semantic changes to investigation_procedure and general
instructions content. The only changes vs master are now structural:
- Files inlined (no separate partials)
- Skill Usage hoisted above todowrite block (deduped)
- AI safety deleted
- Filler intro lines removed
- Stray newrelic fragment removed from _fetch_logs
- Duplicate skill-fetching bullets removed from "In general"

All prompt text that existed on master is preserved word-for-word.

Signed-off-by: Claude <noreply@anthropic.com>
Master updated one line in investigation_procedure.jinja2 (parallel
tool call example wording). Applied that change to our inlined copy
in generic_ask.jinja2 and kept the file deleted.

Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn enabled auto-merge (squash) May 14, 2026 20:49
@aantn
aantn merged commit 31fa24c into master May 14, 2026
19 of 22 checks passed
@aantn
aantn deleted the claude/simplify-jinja2-prompts-1A0Jz branch May 14, 2026 20:51
aantn pushed a commit that referenced this pull request May 16, 2026
5 iters of eval 259_loki_historical_logs_pod_deleted_docker on each side,
opus-4.6 via OpenRouter:

  baseline cf6ddb7 (2026-04-30): 78.6s  $0.407  7.8 turns  21.8 tools  194K tk
  current  1b61fe3 (2026-05-15): 89.4s  $0.447  10.6 turns 22.0 tools  259K tk
                                 +13.7% +9.8%   +35.9%    +0.9%      +33.6%

Pass rate unchanged (5/5 vs 5/5). Tool count flat. Output tokens flat.
Total/cached input tokens up ~33-41% (z > 3.8, highly significant).
Turns up ~36% (z > 3) — same tool work, spread over more serialized rounds.
That's the signature of a bigger system prompt driving more incremental
multi-step exploration, matching the PR-#1970 hypothesis from the earlier
single-iter ci-benchmark vs master-CI comparison.

Includes the run_sweep.sh helper used to produce these numbers.

Signed-off-by: Claude <noreply@anthropic.com>
aantn pushed a commit that referenced this pull request May 16, 2026
n=5 sweep across six candidate fixes pinpointed PR #2040 as the cause.

Fix-AD (full revert of PR #2040: restore the deleted '# If investigating
Kubernetes problems' section in generic_ask.jinja2 AND delete
holmes/plugins/skills/builtin/kubernetes-troubleshooting/) restores
baseline behavior:

  metric         baseline   current   fix-AD   recovery
  time            78.6s      89.4s    83.0s    0.59
  turns            7.8       10.6      8.6     0.71
  total tokens   194K       259K     203K     0.86
  cached         156K       220K     166K     0.85
  cost          $0.41      $0.45    $0.40    fully (slightly cheaper than baseline)
  fetch_skill (Σ5) 0          29       0      —

z-score vs baseline drops from 5.11 (turns) on current to 1.79 on
fix-AD — i.e. inside the noise floor. Per-iter spread on fix-AD sits
entirely inside the baseline range, no outliers.

Per-iter inspection showed the smoking gun: every post-#2040 iter
called fetch_skill exactly once on turn 1 (pulling the
kubernetes-troubleshooting skill). Baseline had zero such calls
because the skill simply didn't exist. The user-prompt skill catalog
block was identical between baseline and current, but only renders
content when a skill is loaded — PR #2040 was what loaded one.

Fix-A alone (prompt section only) recovered ~45% because the skill
file still existed and still got listed. Fix-B (restoring the 'MUST
use tools' one-liner from PR #1970) didn't help at all. Fix-AB and
fix-AC were equivalent to fix-A. Only fix-AD closes the gap.

The recommended PR diff is in analysis/2026-05-15-regression/the-fix.md.
Raw 5-iter reports for all six conditions are committed alongside.

Signed-off-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants