Skip to content

fix eval 43 datetime - #2130

Merged
aantn merged 6 commits into
masterfrom
claude/opus-benchmark-comparison-GMLrr
Jun 7, 2026
Merged

aantn merged 6 commits into
masterfrom
claude/opus-benchmark-comparison-GMLrr

Conversation

@aantn

@aantn aantn commented Jun 5, 2026 •

Copy link
Copy Markdown
Collaborator

No description provided.

claude added 3 commits June 5, 2026 06:48
Runs the weekly fast-benchmark eval set (regression or benchmark, 16 evals
x 3 iterations) against opus-4.6/4.7/4.8 via OpenRouter under identical
conditions, cross-referenced with Braintrust trace data from the 2026-05-31
weekly run. Documents per-failure root causes and model behavioral
differences.

Signed-off-by: Claude <noreply@anthropic.com>
…ed datetime tool

Verified from tool logs that opus-4.8 invokes the default bash toolset
(BashExecutorToolset) running 'date' to fetch the real clock, bypassing the
prompt's mocked time. opus-4.6/4.7 never shell out (0/12). Replaces the
earlier unverified 'fetches via a tool' phrasing with the measured mechanism
and per-model bash:date invocation rates.

Signed-off-by: Claude <noreply@anthropic.com>
…erify

opus-4.8 intermittently ran the whitelisted 'bash: date' command to verify the
time, which reads the real host clock and bypasses the test's Python-level date
mock, causing nondeterministic failures (~1/3 of runs). Adding an explicit
instruction to rely only on the information already in the prompt stops the
tool detour: verified opus-4.6/4.7/4.8 now make 0 bash:date calls and pass
(opus-4.8 6/6, all three 6/6).

Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@coderabbitai

coderabbitai Bot commented Jun 5, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

This PR adds evaluation documentation comparing Opus 4.6, 4.7, and 4.8 performance across weekly benchmarks, documenting metrics, test failure root causes (including a datetime mock artifact), behavioral differences, and evaluation caveats. It also refines a test fixture prompt to align with evaluation findings.

Changes

Evaluation comparison and test refinement

Layer / File(s) Summary
Opus 4.6/4.7/4.8 evaluation comparison
docs/development/evaluations/opus-4.6-4.7-4.8-comparison.md
Documents evaluation metadata, raw and adjusted pass-rate tables, root causes for failing tests including a 4.8-only datetime mock artifact with unexpected tool usage, behavioral differences across models (runbook discipline, tool-usage patterns, search persistence), and evaluation caveats around excluded tests, provider differences, and logging limitations.
Current datetime test fixture prompt refinement
tests/llm/fixtures/test_ask_holmes/43_current_datetime_from_prompt/test_case.yaml
Test fixture user prompt is updated to explicitly instruct the model to rely only on provided information without tool-based verification, reflecting documented evaluation findings.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~5 minutes

Possibly related PRs

  • HolmesGPT/holmesgpt#1776: Adds model-comparison benchmark summary documentation under the same evaluations docs area.
  • HolmesGPT/holmesgpt#555: Prior work on mocking datetime prompt injection for evals, directly related to the datetime prompt fixture refinement.

Suggested reviewers

  • Sheeproid
  • arikalon1
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title 'fix eval 43 datetime' is specific and directly related to a key change in the PR—fixing the failing eval test case 43 by modifying the datetime prompt instruction.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jun 5, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 #2 · Run @ __599ab88__ (#27087895318) — Jun 7, 09:00 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 599ab88 on branch claude/opus-benchmark-comparison-GMLrr

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Src
✅ 09_crashpod 40.9s 6 12 $0.2998 132,793 130,283 24,727 2,510 1,085 105,001 25,282 124 — src
✅ 101_loki_historical_logs_pod_deleted 76.6s 6 14 $0.4117 152,113 147,647 30,334 4,466 947 114,111 33,536 943 — src
✅ 112_find_pvcs_by_uuid 19.8s 3 4 $0.2059 61,150 59,938 21,879 1,212 619 37,806 22,132 313 — src
✅ 12_job_crashing 37.4s 5 13 $0.2949 111,886 109,389 25,000 2,497 829 82,778 26,611 182 — src
✅ 176_network_policy_blocking_traffic_no_skills 64.2s 6 17 $0.3942 153,259 149,223 30,121 4,036 964 116,954 32,269 773 — src
✅ 227_count_configmaps_per_namespace[0] 19.6s 4 9 $0.2078 76,704 75,580 20,663 1,124 593 53,978 21,602 53 — src
✅ 243_pod_names_contain_service 32.3s 4 10 $0.2484 81,784 79,677 22,583 2,107 954 56,327 23,350 208 — src
✅ 24_misconfigured_pvc 39.1s 5 14 $0.2931 108,014 105,363 24,033 2,651 966 79,269 26,094 212 — src
✅ 43_current_datetime_from_prompt 4.3s 1 — $0.1209 17,092 16,928 16,928 164 164 0 16,928 118 — src
✅ 51_logs_summarize_errors 21.2s 4 5 $0.2001 76,460 75,435 20,618 1,025 372 54,812 20,623 32 — src
✅ 61_exact_match_counting 12.2s 3 3 $0.1519 52,732 52,369 17,876 363 216 34,489 17,880 32 — src
Total 33.4s avg 4.3 avg 10.1 avg $2.8288 1,023,987 1,001,832 30,334 22,155 1,085 735,525 266,307 2,990 —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (1h ago) Δ vs master benchmark (5h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 40.9s 40.6s ±0% 41.4s ±0%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 76.6s 69.0s ↑11% 65.7s ↑17%
112_find_pvcs_by_uuid (opus-4.6) 📄 19.8s 16.8s ↑18% 19.0s ±0%
12_job_crashing (opus-4.6) 📄 37.4s 37.2s ±0% 42.6s ↓12%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 64.2s 56.3s ↑14% 44.3s ↑45%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 19.6s 19.2s ±0% 18.4s ±0%
243_pod_names_contain_service (opus-4.6) 📄 32.3s 33.5s ±0% 35.1s ±0%
24_misconfigured_pvc (opus-4.6) 📄 39.1s 42.5s ±0% 40.6s ±0%
43_current_datetime_from_prompt (opus-4.6) 📄 4.3s 3.4s ↑27% 3.5s ↑22%
51_logs_summarize_errors (opus-4.6) 📄 21.2s 20.7s ±0% 19.7s ±0%
61_exact_match_counting (opus-4.6) 📄 12.2s 10.0s ↑22% 10.2s ↑19%
Total (all, n=11) 33.4s 31.7s — 31.0s —
Comparable (m=11, b=11) 33.4s 31.7s ±0% 31.0s ±0%

Cost comparison:

Test case This branch master (1h ago) Δ vs master benchmark (5h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.2998 $0.2971 ±0% $0.3101 ±0%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.4117 $0.4009 ±0% $0.3829 ±0%
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.2059 $0.1886 ±0% $0.2047 ±0%
12_job_crashing (opus-4.6) 📄 $0.2949 $0.2971 ±0% $0.3167 ±0%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.3942 $0.3461 ↑14% $0.3179 ↑24%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.2078 $0.2100 ±0% $0.2048 ±0%
243_pod_names_contain_service (opus-4.6) 📄 $0.2484 $0.2568 ±0% $0.2662 ±0%
24_misconfigured_pvc (opus-4.6) 📄 $0.2931 $0.3158 ±0% $0.3110 ±0%
43_current_datetime_from_prompt (opus-4.6) 📄 $0.1209 $0.1190 ±0% $0.1190 ±0%
51_logs_summarize_errors (opus-4.6) 📄 $0.2001 $0.2024 ±0% $0.2018 ±0%
61_exact_match_counting (opus-4.6) 📄 $0.1519 $0.1519 ±0% $0.1518 ±0%
Total (all, n=11) $0.2572 $0.2532 — $0.2534 —
Comparable (m=11, b=11) $0.2572 $0.2532 ±0% $0.2534 ±0%

Total tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (5h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 132,793 131,955 ±0% 132,351 ±0%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 152,113 171,732 ↓11% 144,907 ±0%
112_find_pvcs_by_uuid (opus-4.6) 📄 61,150 58,011 ±0% 61,219 ±0%
12_job_crashing (opus-4.6) 📄 111,886 109,981 ±0% 136,570 ↓18%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 153,259 138,101 ↑11% 113,514 ↑35%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 76,704 76,921 ±0% 76,715 ±0%
243_pod_names_contain_service (opus-4.6) 📄 81,784 82,821 ±0% 104,698 ↓22%
24_misconfigured_pvc (opus-4.6) 📄 108,014 133,955 ↓19% 132,672 ↓19%
43_current_datetime_from_prompt (opus-4.6) 📄 17,092 17,001 ±0% 17,001 ±0%
51_logs_summarize_errors (opus-4.6) 📄 76,460 77,056 ±0% 77,120 ±0%
61_exact_match_counting (opus-4.6) 📄 52,732 52,730 ±0% 52,716 ±0%
Total (all, n=11) 93,090 95,479 — 95,408 —
Comparable (m=11, b=11) 93,090 95,479 ±0% 95,408 ±0%

Cached tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (5h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 105,001 104,416 ±0% 102,543 ±0%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 114,111 137,524 ↓17% 109,927 ±0%
112_find_pvcs_by_uuid (opus-4.6) 📄 37,806 36,441 ±0% 38,090 ±0%
12_job_crashing (opus-4.6) 📄 82,778 80,071 ±0% 107,265 ↓23%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 116,954 105,780 ↑11% 82,679 ↑41%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 53,978 54,088 ±0% 54,602 ±0%
243_pod_names_contain_service (opus-4.6) 📄 56,327 56,738 ±0% 78,754 ↓28%
24_misconfigured_pvc (opus-4.6) 📄 79,269 104,693 ↓24% 103,729 ↓24%
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 54,812 55,110 ±0% 55,156 ±0%
61_exact_match_counting (opus-4.6) 📄 34,489 34,487 ±0% 34,481 ±0%
Total (all, n=11) 66,866 69,941 — 69,748 —
Comparable (m=10, b=10) 73,552 76,935 ±0% 76,723 ±0%

Turns comparison:

Test case This branch master (1h ago) Δ vs master benchmark (5h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 6 6 ±0% 6 ±0%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 6 7 ↓14% 6 ±0%
112_find_pvcs_by_uuid (opus-4.6) 📄 3 3 ±0% 3 ±0%
12_job_crashing (opus-4.6) 📄 5 5 ±0% 6 ↓17%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 6 6 ±0% 5 ↑20%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 4 4 ±0% 4 ±0%
243_pod_names_contain_service (opus-4.6) 📄 4 4 ±0% 5 ↓20%
24_misconfigured_pvc (opus-4.6) 📄 5 6 ↓17% 6 ↓17%
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% 1 ±0%
51_logs_summarize_errors (opus-4.6) 📄 4 4 ±0% 4 ±0%
61_exact_match_counting (opus-4.6) 📄 3 3 ±0% 3 ±0%
Total (all, n=11) 4.3 4.5 — 4.5 —
Comparable (m=11, b=11) 4.3 4.5 ±0% 4.5 ±0%

Tool calls comparison:

Test case This branch master (1h ago) Δ vs master benchmark (5h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 12 12 ±0% 13 ±0%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 14 17 ↓18% 14 ±0%
112_find_pvcs_by_uuid (opus-4.6) 📄 4 4 ±0% 4 ±0%
12_job_crashing (opus-4.6) 📄 13 12 ±0% 15 ↓13%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 17 17 ±0% 15 ↑13%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 9 9 ±0% 9 ±0%
243_pod_names_contain_service (opus-4.6) 📄 10 11 ±0% 11 ±0%
24_misconfigured_pvc (opus-4.6) 📄 14 15 ±0% 15 ±0%
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 5 5 ±0% 5 ±0%
61_exact_match_counting (opus-4.6) 📄 3 3 ±0% 3 ±0%
Total (all, n=11) 9.2 10.5 — 10.4 —
Comparable (m=10, b=10) 10.1 10.5 ±0% 10.4 ±0%

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
⚠️ 1 older run truncated

Older runs were omitted to stay under GitHub's 64KB comment size limit.


✅ Results of HolmesGPT evals

Automatically triggered by commit fa0f09e on branch claude/opus-benchmark-comparison-GMLrr

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Src
✅ 09_crashpod 30.9s 4 8 $0.2073 68,263 66,566 19,149 1,697 679 46,837 19,729 164 — src
✅ 101_loki_historical_logs_pod_deleted 44.5s 4 8 $0.2283 67,839 65,265 19,153 2,574 988 45,825 19,440 517 — src
✅ 112_find_pvcs_by_uuid 16.2s 2 2 $0.1545 32,335 31,356 17,492 979 532 13,861 17,495 303 — src
✅ 12_job_crashing 28.5s 4 9 $0.2190 69,271 67,589 19,653 1,682 479 45,533 22,056 65 — src
✅ 176_network_policy_blocking_traffic_no_skills 37.3s 4 10 $0.2155 69,117 67,285 19,976 1,832 617 46,953 20,332 323 — src
✅ 227_count_configmaps_per_namespace[0] 15.3s 3 6 $0.1455 45,011 44,291 16,216 720 439 28,071 16,220 36 — src
✅ 243_pod_names_contain_service 29.2s 3 7 $0.1836 48,015 46,336 17,671 1,679 740 28,466 17,870 277 — src
✅ 24_misconfigured_pvc 30.5s 4 10 $0.2113 67,562 65,726 18,807 1,836 731 45,759 19,967 129 — src
✅ 43_current_datetime_from_prompt 5.1s 1 — $0.0984 13,966 13,848 13,848 118 118 0 13,848 78 — src
✅ 51_logs_summarize_errors 19.1s 3 2 $0.1510 45,330 44,529 16,516 801 440 28,009 16,520 29 — src
✅ 61_exact_match_counting 8.2s 2 1 $0.1111 28,229 28,010 14,182 219 150 13,825 14,185 31 — src
Total 24.1s avg 3.1 avg 6.3 avg $1.9256 554,938 540,801 19,976 14,137 988 343,139 197,662 1,952 —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (1h ago) Δ vs master benchmark (6h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 30.9s 28.1s ↑10% 41.4s ↓25%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 44.5s 41.6s ±0% 65.7s ↓32%
112_find_pvcs_by_uuid (opus-4.6) 📄 16.2s 14.5s ↑12% 19.0s ↓15%
12_job_crashing (opus-4.6) 📄 28.5s 30.2s ±0% 42.6s ↓33%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 37.3s 31.6s ↑18% 44.3s ↓16%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 15.3s 12.9s ↑19% 18.4s ↓17%
243_pod_names_contain_service (opus-4.6) 📄 29.2s 28.1s ±0% 35.1s ↓17%
24_misconfigured_pvc (opus-4.6) 📄 30.5s 26.6s ↑15% 40.6s ↓25%
43_current_datetime_from_prompt (opus-4.6) 📄 5.1s 2.8s ↑82% 3.5s ↑45%
51_logs_summarize_errors (opus-4.6) 📄 19.1s 16.6s ↑15% 19.7s ±0%
61_exact_match_counting (opus-4.6) 📄 8.2s 6.3s ↑30% 10.2s ↓20%
Total (all, n=11) 24.1s 21.8s — 31.0s —
Comparable (m=11, b=11) 24.1s 21.8s ↑11% 31.0s ↓22%

Cost comparison:

Test case This branch master (1h ago) Δ vs master benchmark (6h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.2073 $0.2101 ±0% $0.3101 ↓33%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.2283 $0.2417 ±0% $0.3829 ↓40%
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.1545 $0.1574 ±0% $0.2047 ↓25%
12_job_crashing (opus-4.6) 📄 $0.2190 $0.2172 ±0% $0.3167 ↓31%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.2155 $0.2114 ±0% $0.3179 ↓32%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.1455 $0.1449 ±0% $0.2048 ↓29%
243_pod_names_contain_service (opus-4.6) 📄 $0.1836 $0.1832 ±0% $0.2662 ↓31%
24_misconfigured_pvc (opus-4.6) 📄 $0.2113 $0.2076 ±0% $0.3110 ↓32%
43_current_datetime_from_prompt (opus-4.6) 📄 $0.0984 $0.0968 ±0% $0.1190 ↓17%
51_logs_summarize_errors (opus-4.6) 📄 $0.1510 $0.1503 ±0% $0.2018 ↓25%
61_exact_match_counting (opus-4.6) 📄 $0.1111 $0.1112 ±0% $0.1518 ↓27%
Total (all, n=11) $0.1751 $0.1756 — $0.2534 —
Comparable (m=11, b=11) $0.1751 $0.1756 ±0% $0.2534 ↓31%

Total tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (6h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 68,263 68,824 ±0% 132,351 ↓48%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 67,839 68,915 ±0% 144,907 ↓53%
112_find_pvcs_by_uuid (opus-4.6) 📄 32,335 32,628 ±0% 61,219 ↓47%
12_job_crashing (opus-4.6) 📄 69,271 69,498 ±0% 136,570 ↓49%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 69,117 68,972 ±0% 113,514 ↓39%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 45,011 45,001 ±0% 76,715 ↓41%
243_pod_names_contain_service (opus-4.6) 📄 48,015 48,098 ±0% 104,698 ↓54%
24_misconfigured_pvc (opus-4.6) 📄 67,562 67,497 ±0% 132,672 ↓49%
43_current_datetime_from_prompt (opus-4.6) 📄 13,966 13,883 ±0% 17,001 ↓18%
51_logs_summarize_errors (opus-4.6) 📄 45,330 45,446 ±0% 77,120 ↓41%
61_exact_match_counting (opus-4.6) 📄 28,229 28,229 ±0% 52,716 ↓46%
Total (all, n=11) 50,449 50,636 — 95,408 —
Comparable (m=11, b=11) 50,449 50,636 ±0% 95,408 ↓47%

Cached tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (6h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 46,837 47,354 ±0% 102,543 ↓54%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 45,825 45,648 ±0% 109,927 ↓58%
112_find_pvcs_by_uuid (opus-4.6) 📄 13,861 13,861 ±0% 38,090 ↓64%
12_job_crashing (opus-4.6) 📄 45,533 47,414 ±0% 107,265 ↓58%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 46,953 47,119 ±0% 82,679 ↓43%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 28,071 28,073 ±0% 54,602 ↓49%
243_pod_names_contain_service (opus-4.6) 📄 28,466 28,533 ±0% 78,754 ↓64%
24_misconfigured_pvc (opus-4.6) 📄 45,759 46,086 ±0% 103,729 ↓56%
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 28,009 28,007 ±0% 55,156 ↓49%
61_exact_match_counting (opus-4.6) 📄 13,825 13,825 ±0% 34,481 ↓60%
Total (all, n=11) 31,194 31,447 — 69,748 —
Comparable (m=10, b=10) 34,314 34,592 ±0% 76,723 ↓55%

Turns comparison:

Test case This branch master (1h ago) Δ vs master benchmark (6h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 4 4 ±0% 6 ↓33%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 4 4 ±0% 6 ↓33%
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% 3 ↓33%
12_job_crashing (opus-4.6) 📄 4 4 ±0% 6 ↓33%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 4 4 ±0% 5 ↓20%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 3 3 ±0% 4 ↓25%
243_pod_names_contain_service (opus-4.6) 📄 3 3 ±0% 5 ↓40%
24_misconfigured_pvc (opus-4.6) 📄 4 4 ±0% 6 ↓33%
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% 1 ±0%
51_logs_summarize_errors (opus-4.6) 📄 3 3 ±0% 4 ↓25%
61_exact_match_counting (opus-4.6) 📄 2 2 ±0% 3 ↓33%
Total (all, n=11) 3.1 3.1 — 4.5 —
Comparable (m=11, b=11) 3.1 3.1 ±0% 4.5 ↓31%

Tool calls comparison:

Test case This branch master (1h ago) Δ vs master benchmark (6h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 8 8 ±0% 13 ↓38%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 8 9 ↓11% 14 ↓43%
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% 4 ↓50%
12_job_crashing (opus-4.6) 📄 9 9 ±0% 15 ↓40%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 10 10 ±0% 15 ↓33%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 6 6 ±0% 9 ↓33%
243_pod_names_contain_service (opus-4.6) 📄 7 7 ±0% 11 ↓36%
24_misconfigured_pvc (opus-4.6) 📄 10 10 ±0% 15 ↓33%
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 2 2 ±0% 5 ↓60%
61_exact_match_counting (opus-4.6) 📄 1 1 ±0% 3 ↓67%
Total (all, n=11) 5.7 6.4 — 10.4 —
Comparable (m=10, b=10) 6.3 6.4 ±0% 10.4 ↓39%

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/opus-benchmark-comparison-GMLrr -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/opus-benchmark-comparison-GMLrr -f markers=regression -f filter=

@github-actions

github-actions Bot commented Jun 5, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for e055f1f0b (built in 4m 19s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e055f1f0b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e055f1f0b me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e055f1f0b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e055f1f0b
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:e055f1f0b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:e055f1f0b me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:e055f1f0b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:e055f1f0b

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:e055f1f0b \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:e055f1f0b

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:e055f1f0b \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:e055f1f0b

@netlify

netlify Bot commented Jun 5, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit fa0f09e
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a253d99a4e03100093b7b7a
😎 Deploy Preview https://deploy-preview-2130--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@tests/llm/fixtures/test_ask_holmes/43_current_datetime_from_prompt/test_case.yaml`:
- Line 2: Remove the added prescriptive sentence from the test prompt in the
test_ask_holmes fixture (the user_prompt in test_case.yaml) so the prompt does
not explicitly instruct the model to avoid tools; instead extend the test
harness to mock/intercept shell "date" invocations (the same mechanism used for
holmes.plugins.prompts.datetime) and return the predetermined mocked timestamp
when the model calls bash/date, ensuring the test verifies context-extraction
behavior rather than instruction-following compliance.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 6bf208c0-95c4-4852-b1c3-38678c2c0294

📥 Commits

Reviewing files that changed from the base of the PR and between 4d0a70f and eae229a.

📒 Files selected for processing (2)
  • docs/development/evaluations/opus-4.6-4.7-4.8-comparison.md
  • tests/llm/fixtures/test_ask_holmes/43_current_datetime_from_prompt/test_case.yaml

Signed-off-by: Claude <noreply@anthropic.com>
@aantn aantn changed the title Add Opus 4.6/4.7/4.8 weekly-benchmark comparison report fix eval 43 datetime Jun 7, 2026
@aantn
aantn enabled auto-merge (squash) June 7, 2026 08:53
@aantn
aantn merged commit 4c97347 into master Jun 7, 2026
18 of 19 checks passed
@aantn
aantn deleted the claude/opus-benchmark-comparison-GMLrr branch June 7, 2026 09:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants