Repository navigation
fix eval 43 datetime - #2130
fix eval 43 datetime#2130
Conversation
Runs the weekly fast-benchmark eval set (regression or benchmark, 16 evals x 3 iterations) against opus-4.6/4.7/4.8 via OpenRouter under identical conditions, cross-referenced with Braintrust trace data from the 2026-05-31 weekly run. Documents per-failure root causes and model behavioral differences. Signed-off-by: Claude <noreply@anthropic.com>
…ed datetime tool Verified from tool logs that opus-4.8 invokes the default bash toolset (BashExecutorToolset) running 'date' to fetch the real clock, bypassing the prompt's mocked time. opus-4.6/4.7 never shell out (0/12). Replaces the earlier unverified 'fetches via a tool' phrasing with the measured mechanism and per-model bash:date invocation rates. Signed-off-by: Claude <noreply@anthropic.com>
…erify opus-4.8 intermittently ran the whitelisted 'bash: date' command to verify the time, which reads the real host clock and bypasses the test's Python-level date mock, causing nondeterministic failures (~1/3 of runs). Adding an explicit instruction to rely only on the information already in the prompt stops the tool detour: verified opus-4.6/4.7/4.8 now make 0 bash:date calls and pass (opus-4.8 6/6, all three 6/6). Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.
Tip: disable this comment in your organization's Code Review settings.
WalkthroughThis PR adds evaluation documentation comparing Opus 4.6, 4.7, and 4.8 performance across weekly benchmarks, documenting metrics, test failure root causes (including a datetime mock artifact), behavioral differences, and evaluation caveats. It also refines a test fixture prompt to align with evaluation findings. ChangesEvaluation comparison and test refinement
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~5 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
📂 Previous Runs📜 #2 · Run @ __599ab88__ (#27087895318) — Jun 7, 09:00 UTC✅ Results of HolmesGPT evalsAutomatically triggered by commit 599ab88 on branch Results of HolmesGPT evals
Benchmark Comparison DetailsMaster baseline: latest master-* experiment (post-merge regression eval)
Benchmark baseline: latest ci-benchmark experiment on master
Time comparison (seconds):
Cost comparison:
Total tokens comparison:
Cached tokens comparison:
Turns comparison:
Tool calls comparison:
Comparison indicators:
|
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions | Src |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 30.9s | 4 | 8 | $0.2073 | 68,263 | 66,566 | 19,149 | 1,697 | 679 | 46,837 | 19,729 | 164 | — | src |
| ✅ | 101_loki_historical_logs_pod_deleted | 44.5s | 4 | 8 | $0.2283 | 67,839 | 65,265 | 19,153 | 2,574 | 988 | 45,825 | 19,440 | 517 | — | src |
| ✅ | 112_find_pvcs_by_uuid | 16.2s | 2 | 2 | $0.1545 | 32,335 | 31,356 | 17,492 | 979 | 532 | 13,861 | 17,495 | 303 | — | src |
| ✅ | 12_job_crashing | 28.5s | 4 | 9 | $0.2190 | 69,271 | 67,589 | 19,653 | 1,682 | 479 | 45,533 | 22,056 | 65 | — | src |
| ✅ | 176_network_policy_blocking_traffic_no_skills | 37.3s | 4 | 10 | $0.2155 | 69,117 | 67,285 | 19,976 | 1,832 | 617 | 46,953 | 20,332 | 323 | — | src |
| ✅ | 227_count_configmaps_per_namespace[0] | 15.3s | 3 | 6 | $0.1455 | 45,011 | 44,291 | 16,216 | 720 | 439 | 28,071 | 16,220 | 36 | — | src |
| ✅ | 243_pod_names_contain_service | 29.2s | 3 | 7 | $0.1836 | 48,015 | 46,336 | 17,671 | 1,679 | 740 | 28,466 | 17,870 | 277 | — | src |
| ✅ | 24_misconfigured_pvc | 30.5s | 4 | 10 | $0.2113 | 67,562 | 65,726 | 18,807 | 1,836 | 731 | 45,759 | 19,967 | 129 | — | src |
| ✅ | 43_current_datetime_from_prompt | 5.1s | 1 | — | $0.0984 | 13,966 | 13,848 | 13,848 | 118 | 118 | 0 | 13,848 | 78 | — | src |
| ✅ | 51_logs_summarize_errors | 19.1s | 3 | 2 | $0.1510 | 45,330 | 44,529 | 16,516 | 801 | 440 | 28,009 | 16,520 | 29 | — | src |
| ✅ | 61_exact_match_counting | 8.2s | 2 | 1 | $0.1111 | 28,229 | 28,010 | 14,182 | 219 | 150 | 13,825 | 14,185 | 31 | — | src |
| Total | 24.1s avg | 3.1 avg | 6.3 avg | $1.9256 | 554,938 | 540,801 | 19,976 | 14,137 | 988 | 343,139 | 197,662 | 1,952 | — |
Benchmark Comparison Details
Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded
- master-27087951262 (created: 2026-06-07)
Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded
- ci-benchmark-27081280174 (created: 2026-06-07)
Time comparison (seconds):
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 30.9s | 28.1s | ↑10% | 41.4s | ↓25% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 44.5s | 41.6s | ±0% | 65.7s | ↓32% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 16.2s | 14.5s | ↑12% | 19.0s | ↓15% |
| 12_job_crashing (opus-4.6) 📄 | 28.5s | 30.2s | ±0% | 42.6s | ↓33% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 37.3s | 31.6s | ↑18% | 44.3s | ↓16% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 15.3s | 12.9s | ↑19% | 18.4s | ↓17% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 29.2s | 28.1s | ±0% | 35.1s | ↓17% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 30.5s | 26.6s | ↑15% | 40.6s | ↓25% |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 5.1s | 2.8s | ↑82% | 3.5s | ↑45% |
| 51_logs_summarize_errors (opus-4.6) 📄 | 19.1s | 16.6s | ↑15% | 19.7s | ±0% |
| 61_exact_match_counting (opus-4.6) 📄 | 8.2s | 6.3s | ↑30% | 10.2s | ↓20% |
| Total (all, n=11) | 24.1s | 21.8s | — | 31.0s | — |
| Comparable (m=11, b=11) | 24.1s | 21.8s | ↑11% | 31.0s | ↓22% |
Cost comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | $0.2073 | $0.2101 | ±0% | $0.3101 | ↓33% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | $0.2283 | $0.2417 | ±0% | $0.3829 | ↓40% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | $0.1545 | $0.1574 | ±0% | $0.2047 | ↓25% |
| 12_job_crashing (opus-4.6) 📄 | $0.2190 | $0.2172 | ±0% | $0.3167 | ↓31% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | $0.2155 | $0.2114 | ±0% | $0.3179 | ↓32% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | $0.1455 | $0.1449 | ±0% | $0.2048 | ↓29% |
| 243_pod_names_contain_service (opus-4.6) 📄 | $0.1836 | $0.1832 | ±0% | $0.2662 | ↓31% |
| 24_misconfigured_pvc (opus-4.6) 📄 | $0.2113 | $0.2076 | ±0% | $0.3110 | ↓32% |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | $0.0984 | $0.0968 | ±0% | $0.1190 | ↓17% |
| 51_logs_summarize_errors (opus-4.6) 📄 | $0.1510 | $0.1503 | ±0% | $0.2018 | ↓25% |
| 61_exact_match_counting (opus-4.6) 📄 | $0.1111 | $0.1112 | ±0% | $0.1518 | ↓27% |
| Total (all, n=11) | $0.1751 | $0.1756 | — | $0.2534 | — |
| Comparable (m=11, b=11) | $0.1751 | $0.1756 | ±0% | $0.2534 | ↓31% |
Total tokens comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 68,263 | 68,824 | ±0% | 132,351 | ↓48% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 67,839 | 68,915 | ±0% | 144,907 | ↓53% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 32,335 | 32,628 | ±0% | 61,219 | ↓47% |
| 12_job_crashing (opus-4.6) 📄 | 69,271 | 69,498 | ±0% | 136,570 | ↓49% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 69,117 | 68,972 | ±0% | 113,514 | ↓39% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 45,011 | 45,001 | ±0% | 76,715 | ↓41% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 48,015 | 48,098 | ±0% | 104,698 | ↓54% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 67,562 | 67,497 | ±0% | 132,672 | ↓49% |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 13,966 | 13,883 | ±0% | 17,001 | ↓18% |
| 51_logs_summarize_errors (opus-4.6) 📄 | 45,330 | 45,446 | ±0% | 77,120 | ↓41% |
| 61_exact_match_counting (opus-4.6) 📄 | 28,229 | 28,229 | ±0% | 52,716 | ↓46% |
| Total (all, n=11) | 50,449 | 50,636 | — | 95,408 | — |
| Comparable (m=11, b=11) | 50,449 | 50,636 | ±0% | 95,408 | ↓47% |
Cached tokens comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 46,837 | 47,354 | ±0% | 102,543 | ↓54% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 45,825 | 45,648 | ±0% | 109,927 | ↓58% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 13,861 | 13,861 | ±0% | 38,090 | ↓64% |
| 12_job_crashing (opus-4.6) 📄 | 45,533 | 47,414 | ±0% | 107,265 | ↓58% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 46,953 | 47,119 | ±0% | 82,679 | ↓43% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 28,071 | 28,073 | ±0% | 54,602 | ↓49% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 28,466 | 28,533 | ±0% | 78,754 | ↓64% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 45,759 | 46,086 | ±0% | 103,729 | ↓56% |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | — | — | — | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 28,009 | 28,007 | ±0% | 55,156 | ↓49% |
| 61_exact_match_counting (opus-4.6) 📄 | 13,825 | 13,825 | ±0% | 34,481 | ↓60% |
| Total (all, n=11) | 31,194 | 31,447 | — | 69,748 | — |
| Comparable (m=10, b=10) | 34,314 | 34,592 | ±0% | 76,723 | ↓55% |
Turns comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 4 | 4 | ±0% | 6 | ↓33% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 4 | 4 | ±0% | 6 | ↓33% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 2 | 2 | ±0% | 3 | ↓33% |
| 12_job_crashing (opus-4.6) 📄 | 4 | 4 | ±0% | 6 | ↓33% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 4 | 4 | ±0% | 5 | ↓20% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 3 | 3 | ±0% | 4 | ↓25% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 3 | 3 | ±0% | 5 | ↓40% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 4 | 4 | ±0% | 6 | ↓33% |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 1 | 1 | ±0% | 1 | ±0% |
| 51_logs_summarize_errors (opus-4.6) 📄 | 3 | 3 | ±0% | 4 | ↓25% |
| 61_exact_match_counting (opus-4.6) 📄 | 2 | 2 | ±0% | 3 | ↓33% |
| Total (all, n=11) | 3.1 | 3.1 | — | 4.5 | — |
| Comparable (m=11, b=11) | 3.1 | 3.1 | ±0% | 4.5 | ↓31% |
Tool calls comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 8 | 8 | ±0% | 13 | ↓38% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 8 | 9 | ↓11% | 14 | ↓43% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 2 | 2 | ±0% | 4 | ↓50% |
| 12_job_crashing (opus-4.6) 📄 | 9 | 9 | ±0% | 15 | ↓40% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 10 | 10 | ±0% | 15 | ↓33% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 6 | 6 | ±0% | 9 | ↓33% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 7 | 7 | ±0% | 11 | ↓36% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 10 | 10 | ±0% | 15 | ↓33% |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | — | — | — | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 2 | 2 | ±0% | 5 | ↓60% |
| 61_exact_match_counting (opus-4.6) 📄 | 1 | 1 | ±0% | 3 | ↓67% |
| Total (all, n=11) | 5.7 | 6.4 | — | 10.4 | — |
| Comparable (m=10, b=10) | 6.3 | 6.4 | ±0% | 10.4 | ↓39% |
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
📖 Legend
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/opus-benchmark-comparison-GMLrr -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
tags: regression
Or with more options (one per line):
/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
tags: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
tags |
Pytest tags / markers (no default - runs all tests!) |
id |
Eval ID / pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):
| Label | Effect |
|---|---|
evals-tag-<name> |
Run tests with tag <name> alongside regression |
evals-id-<name> |
Run a specific eval by test ID |
evals-model-<name> |
Override the model (use model list name, e.g. sonnet-4.5) |
Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5
🏷️ Valid tags
benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs
🤖 Valid models
deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/opus-benchmark-comparison-GMLrr -f markers=regression -f filter=
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e055f1f0b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e055f1f0b me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e055f1f0b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e055f1f0b
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:e055f1f0b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:e055f1f0b me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:e055f1f0b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:e055f1f0bPatch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:e055f1f0b \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:e055f1f0bRobusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:e055f1f0b \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:e055f1f0b |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In
`@tests/llm/fixtures/test_ask_holmes/43_current_datetime_from_prompt/test_case.yaml`:
- Line 2: Remove the added prescriptive sentence from the test prompt in the
test_ask_holmes fixture (the user_prompt in test_case.yaml) so the prompt does
not explicitly instruct the model to avoid tools; instead extend the test
harness to mock/intercept shell "date" invocations (the same mechanism used for
holmes.plugins.prompts.datetime) and return the predetermined mocked timestamp
when the model calls bash/date, ensuring the test verifies context-extraction
behavior rather than instruction-following compliance.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 6bf208c0-95c4-4852-b1c3-38678c2c0294
📒 Files selected for processing (2)
docs/development/evaluations/opus-4.6-4.7-4.8-comparison.mdtests/llm/fixtures/test_ask_holmes/43_current_datetime_from_prompt/test_case.yaml
Signed-off-by: Claude <noreply@anthropic.com>
No description provided.