Skip to content

Add enable_todo flag to control TodoWrite feature in evals - #2127

Merged
aantn merged 2 commits into
masterfrom
claude/sleepy-hopper-VItzj
Jun 7, 2026
Merged

aantn merged 2 commits into
masterfrom
claude/sleepy-hopper-VItzj

Conversation

@aantn

@aantn aantn commented Jun 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Add a new enable_todo configuration flag to allow tests to opt-in to the TodoWrite/todos feature, which is disabled by default in evaluations. This prevents the TodoWrite tool and related prompt instructions from being offered to the LLM unless explicitly enabled by a test case.

Changes

  • Test case configuration (test_case_utils.py): Added enable_todo: bool = False field to HolmesTestCase to allow individual tests to enable the feature
  • Toolset manager (test_toolset.py):
    • Added enable_todo parameter to TestToolsetManager.__init__
    • Drop the core_investigation toolset entirely when enable_todo=False to prevent the TodoWrite tool from being offered to the LLM
  • Ask Holmes test runner (test_ask_holmes.py):
    • Import PromptComponent from prompt module
    • Pass enable_todo flag from test case to TestToolsetManager
    • Create prompt_component_overrides dict to disable TodoWrite-related prompt instructions/reminders when enable_todo=False
    • Pass prompt_component_overrides to both CLI and API investigation paths

Implementation Details

The TodoWrite feature is now controlled at three levels:

  1. Tool availability: The core_investigation toolset (which contains TodoWrite) is dropped from the toolset manager unless opted in
  2. Prompt instructions: TodoWrite-related prompt components (TODOWRITE_INSTRUCTIONS and TODOWRITE_REMINDER) are disabled via prompt_component_overrides unless opted in
  3. Test configuration: Tests can enable the feature by setting enable_todo: True in their test case definition

This ensures the LLM won't be offered or reminded about the TodoWrite tool in evaluations unless a specific test explicitly enables it.

https://claude.ai/code/session_01HeCcwBCximu5VRbgno8gqr

Summary by CodeRabbit

  • New Features
    • Added per-test enable_todo configuration option to control Todo-related functionality
    • Todo features and related tools are disabled by default and can be optionally enabled on a per-test basis

Disable the TodoWrite/todos feature in eval runs by default. Both the
TodoWrite tool (core_investigation toolset) and the related prompt
instructions/reminder are turned off so they don't influence eval
behavior or token usage.

Tests that specifically need todos can opt back in by setting
'enable_todo: true' in their test_case.yaml.

Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@github-actions

github-actions Bot commented Jun 4, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 #1 · Run @ __01c438a__ (#26970039513) — Jun 4, 18:08 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 01c438a on branch claude/sleepy-hopper-VItzj

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Src
✅ 09_crashpod 31.8s 4 8 $0.2081 68,300 66,560 19,100 1,740 709 46,896 19,664 158 — src
✅ 101_loki_historical_logs_pod_deleted 43.0s 4 8 $0.2306 67,899 65,293 19,097 2,606 1,009 45,647 19,646 425 — src
✅ 112_find_pvcs_by_uuid 15.9s 2 2 $0.1543 32,326 31,351 17,487 975 529 13,861 17,490 293 — src
✅ 12_job_crashing 29.6s 4 9 $0.2160 69,311 67,644 19,616 1,667 467 46,163 21,481 63 — src
✅ 176_network_policy_blocking_traffic_no_skills 29.7s 4 10 $0.2118 69,073 67,347 20,031 1,726 618 47,134 20,213 350 — src
✅ 227_count_configmaps_per_namespace[0] 14.4s 3 6 $0.1459 44,980 44,269 16,200 711 438 28,065 16,204 29 — src
✅ 243_pod_names_contain_service 30.3s 3 7 $0.1860 48,487 46,763 17,609 1,724 583 28,682 18,081 276 — src
✅ 24_misconfigured_pvc 28.1s 4 10 $0.2007 66,664 65,046 18,639 1,618 602 45,844 19,202 56 — src
✅ 43_current_datetime_from_prompt 3.9s 1 — $0.0968 13,883 13,819 13,819 64 64 0 13,819 24 — src
✅ 51_logs_summarize_errors 18.1s 3 2 $0.1510 45,505 44,764 16,754 741 382 28,006 16,758 29 — src
✅ 61_exact_match_counting 7.8s 2 1 $0.1113 28,238 28,015 14,187 223 154 13,825 14,190 36 — src
Total 23.0s avg 3.1 avg 6.3 avg $1.9124 554,666 540,871 20,031 13,795 1,009 344,123 196,748 1,739 —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 177 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (1d ago) Δ vs master benchmark (4d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 31.8s 39.5s ↓20% 45.0s ↓29%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 43.0s 84.3s ↓49% 85.9s ↓50%
112_find_pvcs_by_uuid (opus-4.6) 📄 15.9s 22.4s ↓29% 21.4s ↓25%
12_job_crashing (opus-4.6) 📄 29.6s 49.0s ↓40% 38.6s ↓23%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 29.7s 66.5s ↓55% 53.9s ↓45%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 14.4s 21.7s ↓33% 20.8s ↓31%
243_pod_names_contain_service (opus-4.6) 📄 30.3s 35.0s ↓13% 33.6s ±0%
24_misconfigured_pvc (opus-4.6) 📄 28.1s 45.3s ↓38% 37.0s ↓24%
43_current_datetime_from_prompt (opus-4.6) 📄 3.9s 3.5s ↑14% 3.6s ↑10%
51_logs_summarize_errors (opus-4.6) 📄 18.1s 22.9s ↓21% 21.9s ↓17%
61_exact_match_counting (opus-4.6) 📄 7.8s 11.6s ↓32% 10.0s ↓21%
Total (all, n=11) 23.0s 36.5s — 33.8s —
Comparable (m=11, b=11) 23.0s 36.5s ↓37% 33.8s ↓32%

Cost comparison:

Test case This branch master (1d ago) Δ vs master benchmark (4d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.2081 $0.2949 ↓29% $0.3318 ↓37%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.2306 $0.4098 ↓44% $0.4631 ↓50%
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.1543 $0.1861 ↓17% $0.2066 ↓25%
12_job_crashing (opus-4.6) 📄 $0.2160 $0.3288 ↓34% $0.3107 ↓30%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.2118 $0.3783 ↓44% $0.3426 ↓38%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.1459 $0.2031 ↓28% $0.2031 ↓28%
243_pod_names_contain_service (opus-4.6) 📄 $0.1860 $0.2504 ↓26% $0.2521 ↓26%
24_misconfigured_pvc (opus-4.6) 📄 $0.2007 $0.3180 ↓37% $0.2904 ↓31%
43_current_datetime_from_prompt (opus-4.6) 📄 $0.0968 $0.1190 ↓19% $0.1187 ↓18%
51_logs_summarize_errors (opus-4.6) 📄 $0.1510 $0.2049 ↓26% $0.2015 ↓25%
61_exact_match_counting (opus-4.6) 📄 $0.1113 $0.1518 ↓27% $0.1518 ↓27%
Total (all, n=11) $0.1739 $0.2587 — $0.2611 —
Comparable (m=11, b=11) $0.1739 $0.2587 ↓33% $0.2611 ↓33%

Total tokens comparison:

Test case This branch master (1d ago) Δ vs master benchmark (4d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 68,300 131,686 ↓48% 156,516 ↓56%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 67,899 171,379 ↓60% 204,760 ↓67%
112_find_pvcs_by_uuid (opus-4.6) 📄 32,326 56,309 ↓43% 61,198 ↓47%
12_job_crashing (opus-4.6) 📄 69,311 159,927 ↓57% 135,931 ↓49%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 69,073 150,279 ↓54% 140,458 ↓51%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 44,980 76,668 ↓41% 76,661 ↓41%
243_pod_names_contain_service (opus-4.6) 📄 48,487 81,854 ↓41% 82,231 ↓41%
24_misconfigured_pvc (opus-4.6) 📄 66,664 134,154 ↓50% 107,468 ↓38%
43_current_datetime_from_prompt (opus-4.6) 📄 13,883 17,001 ↓18% 16,989 ↓18%
51_logs_summarize_errors (opus-4.6) 📄 45,505 77,316 ↓41% 76,727 ↓41%
61_exact_match_counting (opus-4.6) 📄 28,238 52,721 ↓46% 52,723 ↓46%
Total (all, n=11) 50,424 100,845 — 101,060 —
Comparable (m=11, b=11) 50,424 100,845 ↓50% 101,060 ↓50%

Cached tokens comparison:

Test case This branch master (1d ago) Δ vs master benchmark (4d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 46,896 104,283 ↓55% 126,090 ↓63%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 45,647 137,233 ↓67% 166,234 ↓73%
112_find_pvcs_by_uuid (opus-4.6) 📄 13,861 35,446 ↓61% 37,847 ↓63%
12_job_crashing (opus-4.6) 📄 46,163 130,712 ↓65% 106,794 ↓57%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 47,134 115,477 ↓59% 108,926 ↓57%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 28,065 54,891 ↓49% 54,888 ↓49%
243_pod_names_contain_service (opus-4.6) 📄 28,682 56,021 ↓49% 56,547 ↓49%
24_misconfigured_pvc (opus-4.6) 📄 45,844 104,830 ↓56% 78,838 ↓42%
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 28,006 55,204 ↓49% 54,936 ↓49%
61_exact_match_counting (opus-4.6) 📄 13,825 34,482 ↓60% 34,484 ↓60%
Total (all, n=11) 31,284 75,325 — 75,053 —
Comparable (m=10, b=10) 34,412 82,858 ↓58% 82,558 ↓58%

Turns comparison:

Test case This branch master (1d ago) Δ vs master benchmark (4d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 4 6 ↓33% 7 ↓43%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 4 7 ↓43% 8 ↓50%
112_find_pvcs_by_uuid (opus-4.6) 📄 2 3 ↓33% 3 ↓33%
12_job_crashing (opus-4.6) 📄 4 7 ↓43% 6 ↓33%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 4 6 ↓33% 6 ↓33%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 3 4 ↓25% 4 ↓25%
243_pod_names_contain_service (opus-4.6) 📄 3 4 ↓25% 4 ↓25%
24_misconfigured_pvc (opus-4.6) 📄 4 6 ↓33% 5 ↓20%
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% 1 ±0%
51_logs_summarize_errors (opus-4.6) 📄 3 4 ↓25% 4 ↓25%
61_exact_match_counting (opus-4.6) 📄 2 3 ↓33% 3 ↓33%
Total (all, n=11) 3.1 4.6 — 4.6 —
Comparable (m=11, b=11) 3.1 4.6 ↓33% 4.6 ↓33%

Tool calls comparison:

Test case This branch master (1d ago) Δ vs master benchmark (4d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 8 12 ↓33% 13 ↓38%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 8 18 ↓56% 18 ↓56%
112_find_pvcs_by_uuid (opus-4.6) 📄 2 4 ↓50% 4 ↓50%
12_job_crashing (opus-4.6) 📄 9 14 ↓36% 14 ↓36%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 10 18 ↓44% 14 ↓29%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 6 9 ↓33% 9 ↓33%
243_pod_names_contain_service (opus-4.6) 📄 7 10 ↓30% 10 ↓30%
24_misconfigured_pvc (opus-4.6) 📄 10 16 ↓38% 15 ↓33%
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 2 5 ↓60% 5 ↓60%
61_exact_match_counting (opus-4.6) 📄 1 3 ↓67% 3 ↓67%
Total (all, n=11) 5.7 10.9 — 10.5 —
Comparable (m=10, b=10) 6.3 10.9 ↓42% 10.5 ↓40%

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit 4fbf2dd on branch claude/sleepy-hopper-VItzj

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Src
✅ 09_crashpod 28.2s 4 8 $0.2049 68,019 66,374 19,036 1,645 694 46,774 19,600 97 — src
✅ 101_loki_historical_logs_pod_deleted 48.0s 5 11 $0.2666 89,874 86,939 21,555 2,935 898 64,935 22,004 389 — src
✅ 112_find_pvcs_by_uuid 15.2s 2 2 $0.1526 32,265 31,352 17,488 913 465 13,861 17,491 263 — src
✅ 12_job_crashing 29.1s 5 10 $0.2208 85,899 84,143 18,743 1,756 575 63,804 20,339 89 — src
✅ 176_network_policy_blocking_traffic_no_skills 28.6s 4 10 $0.2135 68,948 67,161 19,897 1,787 614 46,913 20,248 313 — src
✅ 227_count_configmaps_per_namespace[0] 14.8s 3 6 $0.1448 45,004 44,292 16,214 712 438 28,074 16,218 36 — src
✅ 243_pod_names_contain_service 27.3s 3 7 $0.1829 47,961 46,277 17,618 1,684 742 28,456 17,821 262 — src
✅ 24_misconfigured_pvc 28.3s 4 11 $0.2071 66,810 65,068 18,770 1,742 597 45,257 19,811 49 — src
✅ 43_current_datetime_from_prompt 3.2s 1 — $0.0968 13,883 13,819 13,819 64 64 0 13,819 24 — src
✅ 51_logs_summarize_errors 17.5s 3 2 $0.1521 45,611 44,850 16,841 761 400 28,005 16,845 29 — src
✅ 61_exact_match_counting 8.2s 2 1 $0.1114 28,243 28,016 14,188 227 158 13,825 14,191 39 — src
Total 22.6s avg 3.3 avg 6.8 avg $1.9536 592,517 578,291 21,555 14,226 898 379,904 198,387 1,590 —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (1h ago) Δ vs master benchmark (5h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 28.2s 40.6s ↓31% 41.4s ↓32%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 48.0s 69.0s ↓30% 65.7s ↓27%
112_find_pvcs_by_uuid (opus-4.6) 📄 15.2s 16.8s ±0% 19.0s ↓20%
12_job_crashing (opus-4.6) 📄 29.1s 37.2s ↓22% 42.6s ↓32%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 28.6s 56.3s ↓49% 44.3s ↓35%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 14.8s 19.2s ↓23% 18.4s ↓20%
243_pod_names_contain_service (opus-4.6) 📄 27.3s 33.5s ↓19% 35.1s ↓22%
24_misconfigured_pvc (opus-4.6) 📄 28.3s 42.5s ↓33% 40.6s ↓30%
43_current_datetime_from_prompt (opus-4.6) 📄 3.2s 3.4s ±0% 3.5s ±0%
51_logs_summarize_errors (opus-4.6) 📄 17.5s 20.7s ↓15% 19.7s ↓11%
61_exact_match_counting (opus-4.6) 📄 8.2s 10.0s ↓18% 10.2s ↓20%
Total (all, n=11) 22.6s 31.7s — 31.0s —
Comparable (m=11, b=11) 22.6s 31.7s ↓29% 31.0s ↓27%

Cost comparison:

Test case This branch master (1h ago) Δ vs master benchmark (5h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.2049 $0.2971 ↓31% $0.3101 ↓34%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.2666 $0.4009 ↓33% $0.3829 ↓30%
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.1526 $0.1886 ↓19% $0.2047 ↓25%
12_job_crashing (opus-4.6) 📄 $0.2208 $0.2971 ↓26% $0.3167 ↓30%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.2135 $0.3461 ↓38% $0.3179 ↓33%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.1448 $0.2100 ↓31% $0.2048 ↓29%
243_pod_names_contain_service (opus-4.6) 📄 $0.1829 $0.2568 ↓29% $0.2662 ↓31%
24_misconfigured_pvc (opus-4.6) 📄 $0.2071 $0.3158 ↓34% $0.3110 ↓33%
43_current_datetime_from_prompt (opus-4.6) 📄 $0.0968 $0.1190 ↓19% $0.1190 ↓19%
51_logs_summarize_errors (opus-4.6) 📄 $0.1521 $0.2024 ↓25% $0.2018 ↓25%
61_exact_match_counting (opus-4.6) 📄 $0.1114 $0.1519 ↓27% $0.1518 ↓27%
Total (all, n=11) $0.1776 $0.2532 — $0.2534 —
Comparable (m=11, b=11) $0.1776 $0.2532 ↓30% $0.2534 ↓30%

Total tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (5h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 68,019 131,955 ↓48% 132,351 ↓49%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 89,874 171,732 ↓48% 144,907 ↓38%
112_find_pvcs_by_uuid (opus-4.6) 📄 32,265 58,011 ↓44% 61,219 ↓47%
12_job_crashing (opus-4.6) 📄 85,899 109,981 ↓22% 136,570 ↓37%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 68,948 138,101 ↓50% 113,514 ↓39%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 45,004 76,921 ↓41% 76,715 ↓41%
243_pod_names_contain_service (opus-4.6) 📄 47,961 82,821 ↓42% 104,698 ↓54%
24_misconfigured_pvc (opus-4.6) 📄 66,810 133,955 ↓50% 132,672 ↓50%
43_current_datetime_from_prompt (opus-4.6) 📄 13,883 17,001 ↓18% 17,001 ↓18%
51_logs_summarize_errors (opus-4.6) 📄 45,611 77,056 ↓41% 77,120 ↓41%
61_exact_match_counting (opus-4.6) 📄 28,243 52,730 ↓46% 52,716 ↓46%
Total (all, n=11) 53,865 95,479 — 95,408 —
Comparable (m=11, b=11) 53,865 95,479 ↓44% 95,408 ↓44%

Cached tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (5h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 46,774 104,416 ↓55% 102,543 ↓54%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 64,935 137,524 ↓53% 109,927 ↓41%
112_find_pvcs_by_uuid (opus-4.6) 📄 13,861 36,441 ↓62% 38,090 ↓64%
12_job_crashing (opus-4.6) 📄 63,804 80,071 ↓20% 107,265 ↓41%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 46,913 105,780 ↓56% 82,679 ↓43%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 28,074 54,088 ↓48% 54,602 ↓49%
243_pod_names_contain_service (opus-4.6) 📄 28,456 56,738 ↓50% 78,754 ↓64%
24_misconfigured_pvc (opus-4.6) 📄 45,257 104,693 ↓57% 103,729 ↓56%
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 28,005 55,110 ↓49% 55,156 ↓49%
61_exact_match_counting (opus-4.6) 📄 13,825 34,487 ↓60% 34,481 ↓60%
Total (all, n=11) 34,537 69,941 — 69,748 —
Comparable (m=10, b=10) 37,990 76,935 ↓51% 76,723 ↓50%

Turns comparison:

Test case This branch master (1h ago) Δ vs master benchmark (5h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 4 6 ↓33% 6 ↓33%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 5 7 ↓29% 6 ↓17%
112_find_pvcs_by_uuid (opus-4.6) 📄 2 3 ↓33% 3 ↓33%
12_job_crashing (opus-4.6) 📄 5 5 ±0% 6 ↓17%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 4 6 ↓33% 5 ↓20%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 3 4 ↓25% 4 ↓25%
243_pod_names_contain_service (opus-4.6) 📄 3 4 ↓25% 5 ↓40%
24_misconfigured_pvc (opus-4.6) 📄 4 6 ↓33% 6 ↓33%
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% 1 ±0%
51_logs_summarize_errors (opus-4.6) 📄 3 4 ↓25% 4 ↓25%
61_exact_match_counting (opus-4.6) 📄 2 3 ↓33% 3 ↓33%
Total (all, n=11) 3.3 4.5 — 4.5 —
Comparable (m=11, b=11) 3.3 4.5 ↓27% 4.5 ↓27%

Tool calls comparison:

Test case This branch master (1h ago) Δ vs master benchmark (5h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 8 12 ↓33% 13 ↓38%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 11 17 ↓35% 14 ↓21%
112_find_pvcs_by_uuid (opus-4.6) 📄 2 4 ↓50% 4 ↓50%
12_job_crashing (opus-4.6) 📄 10 12 ↓17% 15 ↓33%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 10 17 ↓41% 15 ↓33%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 6 9 ↓33% 9 ↓33%
243_pod_names_contain_service (opus-4.6) 📄 7 11 ↓36% 11 ↓36%
24_misconfigured_pvc (opus-4.6) 📄 11 15 ↓27% 15 ↓27%
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 2 5 ↓60% 5 ↓60%
61_exact_match_counting (opus-4.6) 📄 1 3 ↓67% 3 ↓67%
Total (all, n=11) 6.2 10.5 — 10.4 —
Comparable (m=10, b=10) 6.8 10.5 ↓35% 10.4 ↓35%

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/sleepy-hopper-VItzj -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/sleepy-hopper-VItzj -f markers=regression -f filter=

@github-actions

github-actions Bot commented Jun 4, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for a0ebf462a (built in 4m 20s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:a0ebf462a
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:a0ebf462a me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:a0ebf462a
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:a0ebf462a
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:a0ebf462a
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:a0ebf462a me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:a0ebf462a
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:a0ebf462a

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:a0ebf462a \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:a0ebf462a

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:a0ebf462a \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:a0ebf462a

@coderabbitai

coderabbitai Bot commented Jun 4, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 1a08bce5-2e99-4a91-afbb-c38ceec9dc02

📥 Commits

Reviewing files that changed from the base of the PR and between 4d0a70f and 01c438a.

📒 Files selected for processing (3)
  • tests/llm/test_ask_holmes.py
  • tests/llm/utils/test_case_utils.py
  • tests/llm/utils/test_toolset.py

Walkthrough

This PR extends the test infrastructure to conditionally gate the TodoWrite feature. A new enable_todo flag in HolmesTestCase flows through TestToolsetManager and ask_holmes, controlling whether the core_investigation toolset is available and whether TodoWrite prompt components are disabled.

Changes

Enable-todo flag propagation

Layer / File(s) Summary
Test case schema update
tests/llm/utils/test_case_utils.py
HolmesTestCase adds an enable_todo boolean field with default value False.
Toolset manager with conditional core_investigation exclusion
tests/llm/utils/test_toolset.py
TestToolsetManager.__init__ accepts enable_todo parameter; _configure_toolsets conditionally removes core_investigation toolset from the configuration when enable_todo is False.
Holmes function integration with prompt overrides
tests/llm/test_ask_holmes.py
ask_holmes forwards enable_todo to TestToolsetManager, builds prompt component overrides to disable TodoWrite instructions and reminders when enable_todo is falsy, and passes overrides to both CLI and non-CLI message builders.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • HolmesGPT/holmesgpt#1830: Refactors CoreInvestigationToolset enablement logic that aligns with this PR's conditional exclusion of core_investigation in the test harness.

Suggested reviewers

  • moshemorad
  • arikalon1
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: adding an enable_todo flag to control TodoWrite feature in evaluation tests.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Jun 4, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 4fbf2dd
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a25311847800900088dfcca
😎 Deploy Preview https://deploy-preview-2127--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@aantn
aantn enabled auto-merge (squash) June 7, 2026 08:51
@aantn
aantn merged commit 6ddd419 into master Jun 7, 2026
18 of 19 checks passed
@aantn
aantn deleted the claude/sleepy-hopper-VItzj branch June 7, 2026 08:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants