Conversation
Reproduces the failure mode where Holmes attributes a failing Airflow KubernetesPodOperator task to dramatic but unrelated cluster-wide infrastructure noise (AWS-CNI IP-address-exhaustion FailedCreatePodSandBox events, node draining / Karpenter churn, high pod density) instead of recognizing the task pods actually ran and failed at the application level. The 3 ephemeral task pods (log-archival-index-to-es-7q4w9z-*) were Scheduled, Pulled and Started (their Normal lifecycle events prove they ran) and were then deleted by the operator, so kubectl logs is unavailable and the real error lives in the Airflow task logs. The correct conclusion is application-level failure + unavailable pod logs, with the infra storm being unrelated context. Verified red on Claude Opus 4.6 (Score 0%, 0/1) across repeated runs: Holmes consistently pins causation on the subnet IP exhaustion (e.g. "the IP exhaustion may have caused ... inability to reach Elasticsearch") rather than treating it as unrelated noise. Tagged hard (a known-failing reasoning-quality case). Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.
Tip: disable this comment in your organization's Code Review settings.
📂 Previous Runs📜 #3 · Run @ __4725ef1__ (#27096911824) — Jun 7, 15:39 UTC✅ Results of HolmesGPT evalsAutomatically triggered by commit 4725ef1 on branch Results of HolmesGPT evals
Benchmark Comparison DetailsMaster baseline: latest master-* experiment (post-merge regression eval)
Benchmark baseline: latest ci-benchmark experiment on master
No baseline data available for comparison. Comparison indicators:
📜 #2 · Run @ __1124f7e__ (#27090599665) — Jun 7, 11:13 UTC
|
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions | Src |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 262_ambiguous_cluster_reference | 29.8s | 5 | 7 | $0.1582 | 56,711 | 55,011 | 12,626 | 1,700 | 603 | 42,124 | 12,887 | 67 | — | src |
| ❌ | 262_root_cause_buried_in_infra_noise | 128.4s | 10 | 31 | — | — | — | — | — | — | — | — | — | — | src |
| Total | 79.1s avg | 7.5 avg | 19.0 avg | $0.1582 | 56,711 | 55,011 | 12,626 | 1,700 | 603 | 42,124 | 12,887 | 67 | — |
Benchmark Comparison Details
Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded
- master-27089078550 (created: 2026-06-07)
Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded
- ci-benchmark-27081280174 (created: 2026-06-07)
No baseline data available for comparison.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 1 Failure Detected
📜 #1 · Run @ __1124f7e__ (#27090566915) — Jun 7, 11:06 UTC
✅ Results of HolmesGPT evals
Automatically triggered by commit 1124f7e on branch claude/brave-goldberg-2PpZb
Results of HolmesGPT evals
- ask_holmes: 11/14 test cases were successful, 0 regressions, 3 setup failures
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions | Src |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 29.5s | 4 | 8 | $0.2146 | 68,390 | 66,616 | 19,169 | 1,774 | 693 | 45,848 | 20,768 | 222 | — | src |
| ✅ | 101_loki_historical_logs_pod_deleted | 46.7s | 5 | 10 | $0.2467 | 87,613 | 84,979 | 19,990 | 2,634 | 680 | 64,699 | 20,280 | 377 | — | src |
| ✅ | 112_find_pvcs_by_uuid | 13.3s | 2 | 2 | $0.1502 | 32,181 | 31,356 | 17,492 | 825 | 448 | 13,861 | 17,495 | 238 | — | src |
| ✅ | 12_job_crashing | 30.9s | 5 | 10 | $0.2368 | 92,585 | 90,747 | 20,812 | 1,838 | 541 | 68,986 | 21,761 | 180 | — | src |
| ✅ | 176_network_policy_blocking_traffic_no_skills | 37.0s | 5 | 13 | $0.2434 | 89,942 | 87,636 | 20,374 | 2,306 | 676 | 66,667 | 20,969 | 375 | — | src |
| ✅ | 227_count_configmaps_per_namespace[0] | 13.8s | 3 | 6 | $0.1448 | 45,002 | 44,290 | 16,213 | 712 | 438 | 28,073 | 16,217 | 36 | — | src |
| ✅ | 243_pod_names_contain_service | 28.2s | 3 | 7 | $0.1852 | 48,429 | 46,733 | 17,629 | 1,696 | 571 | 28,634 | 18,099 | 212 | — | src |
| ✅ | 24_misconfigured_pvc | 25.8s | 4 | 10 | $0.2021 | 66,474 | 64,873 | 18,573 | 1,601 | 597 | 45,263 | 19,610 | 49 | — | src |
| 🚧 | 254_elasticsearch_dr_test_log_check | — | — | — | — | — | — | — | — | — | — | — | — | — | src |
| 🚧 | 259_wrong_cluster_logs_confusion | — | — | — | — | — | — | — | — | — | — | — | — | — | src |
| 🚧 | 260_global_es_remote_cluster_logs | — | — | — | — | — | — | — | — | — | — | — | — | — | src |
| ✅ | 43_current_datetime_from_prompt | 3.8s | 1 | — | $0.0986 | 13,970 | 13,848 | 13,848 | 122 | 122 | 0 | 13,848 | 78 | — | src |
| ✅ | 51_logs_summarize_errors | 19.7s | 3 | 2 | $0.1543 | 45,815 | 45,016 | 17,005 | 799 | 424 | 28,007 | 17,009 | 31 | — | src |
| ✅ | 61_exact_match_counting | 7.1s | 2 | 1 | $0.1113 | 28,235 | 28,012 | 14,184 | 223 | 154 | 13,825 | 14,187 | 36 | — | src |
| Total | 23.2s avg | 3.4 avg | 6.9 avg | $1.9879 | 618,636 | 604,106 | 20,812 | 14,530 | 693 | 403,863 | 200,243 | 1,834 | — |
Benchmark Comparison Details
Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded
- master-27089078550 (created: 2026-06-07)
Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded
- ci-benchmark-27081280174 (created: 2026-06-07)
Time comparison (seconds):
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (7h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 29.5s | 26.2s | ↑12% | 41.4s | ↓29% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 46.7s | 47.6s | ±0% | 65.7s | ↓29% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 13.3s | 10.6s | ↑26% | 19.0s | ↓30% |
| 12_job_crashing (opus-4.6) 📄 | 30.9s | 27.6s | ↑12% | 42.6s | ↓27% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 37.0s | 32.6s | ↑13% | 44.3s | ↓17% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 13.8s | 13.2s | ±0% | 18.4s | ↓25% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 28.2s | 28.8s | ±0% | 35.1s | ↓19% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 25.8s | 31.4s | ↓18% | 40.6s | ↓36% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | — | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | — | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | — | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 3.8s | 3.1s | ↑23% | 3.5s | ±0% |
| 51_logs_summarize_errors (opus-4.6) 📄 | 19.7s | 18.8s | ±0% | 19.7s | ±0% |
| 61_exact_match_counting (opus-4.6) 📄 | 7.1s | 7.1s | ±0% | 10.2s | ↓31% |
| Total (all, n=11) | 23.2s | 22.5s | — | 31.0s | — |
| Comparable (m=11, b=11) | 23.2s | 22.5s | ±0% | 31.0s | ↓25% |
Cost comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (7h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | $0.2146 | $0.2032 | ±0% | $0.3101 | ↓31% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | $0.2467 | $0.2642 | ±0% | $0.3829 | ↓36% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | $0.1502 | $0.1368 | ±0% | $0.2047 | ↓27% |
| 12_job_crashing (opus-4.6) 📄 | $0.2368 | $0.2173 | ±0% | $0.3167 | ↓25% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | $0.2434 | $0.2524 | ±0% | $0.3179 | ↓23% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | $0.1448 | $0.1464 | ±0% | $0.2048 | ↓29% |
| 243_pod_names_contain_service (opus-4.6) 📄 | $0.1852 | $0.1872 | ±0% | $0.2662 | ↓30% |
| 24_misconfigured_pvc (opus-4.6) 📄 | $0.2021 | $0.2240 | ±0% | $0.3110 | ↓35% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | — | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | — | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | — | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | $0.0986 | $0.0110 | ↑797% | $0.1190 | ↓17% |
| 51_logs_summarize_errors (opus-4.6) 📄 | $0.1543 | $0.1488 | ±0% | $0.2018 | ↓24% |
| 61_exact_match_counting (opus-4.6) 📄 | $0.1113 | $0.1112 | ±0% | $0.1518 | ↓27% |
| Total (all, n=14) | $0.1420 | $0.1730 | — | $0.2534 | — |
| Comparable (m=11, b=11) | $0.1807 | $0.1730 | ±0% | $0.2534 | ↓29% |
Total tokens comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (7h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 68,390 | 67,778 | ±0% | 132,351 | ↓48% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 87,613 | 88,307 | ±0% | 144,907 | ↓40% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 32,181 | 30,686 | ±0% | 61,219 | ↓47% |
| 12_job_crashing (opus-4.6) 📄 | 92,585 | 69,487 | ↑33% | 136,570 | ↓32% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 89,942 | 108,726 | ↓17% | 113,514 | ↓21% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 45,002 | 44,988 | ±0% | 76,715 | ↓41% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 48,429 | 48,612 | ±0% | 104,698 | ↓54% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 66,474 | 68,923 | ±0% | 132,672 | ↓50% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | — | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | — | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | — | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 13,970 | 13,970 | ±0% | 17,001 | ↓18% |
| 51_logs_summarize_errors (opus-4.6) 📄 | 45,815 | 45,163 | ±0% | 77,120 | ↓41% |
| 61_exact_match_counting (opus-4.6) 📄 | 28,235 | 28,231 | ±0% | 52,716 | ↓46% |
| Total (all, n=14) | 44,188 | 55,897 | — | 95,408 | — |
| Comparable (m=11, b=11) | 56,240 | 55,897 | ±0% | 95,408 | ↓41% |
Cached tokens comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (7h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 45,848 | 46,700 | ±0% | 102,543 | ↓55% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 64,699 | 64,209 | ±0% | 109,927 | ↓41% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 13,861 | 13,861 | ±0% | 38,090 | ↓64% |
| 12_job_crashing (opus-4.6) 📄 | 68,986 | 46,249 | ↑49% | 107,265 | ↓36% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 66,667 | 84,274 | ↓21% | 82,679 | ↓19% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 28,073 | 28,072 | ±0% | 54,602 | ↓49% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 28,634 | 28,690 | ±0% | 78,754 | ↓64% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 45,263 | 46,110 | ±0% | 103,729 | ↓56% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | — | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | — | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | — | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | — | 13,845 | — | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 28,007 | 28,012 | ±0% | 55,156 | ↓49% |
| 61_exact_match_counting (opus-4.6) 📄 | 13,825 | 13,825 | ±0% | 34,481 | ↓60% |
| Total (all, n=11) | 36,715 | 37,622 | — | 69,748 | — |
| Comparable (m=10, b=10) | 40,386 | 40,000 | ±0% | 76,723 | ↓47% |
Turns comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (7h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 4 | 4 | ±0% | 6 | ↓33% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 5 | 5 | ±0% | 6 | ↓17% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 2 | 2 | ±0% | 3 | ↓33% |
| 12_job_crashing (opus-4.6) 📄 | 5 | 4 | ↑25% | 6 | ↓17% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 5 | 6 | ↓17% | 5 | ±0% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 3 | 3 | ±0% | 4 | ↓25% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 3 | 3 | ±0% | 5 | ↓40% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 4 | 4 | ±0% | 6 | ↓33% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | — | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | — | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | — | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 1 | 1 | ±0% | 1 | ±0% |
| 51_logs_summarize_errors (opus-4.6) 📄 | 3 | 3 | ±0% | 4 | ↓25% |
| 61_exact_match_counting (opus-4.6) 📄 | 2 | 2 | ±0% | 3 | ↓33% |
| Total (all, n=11) | 3.4 | 3.4 | — | 4.5 | — |
| Comparable (m=11, b=11) | 3.4 | 3.4 | ±0% | 4.5 | ↓24% |
Tool calls comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (7h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 8 | 8 | ±0% | 13 | ↓38% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 10 | 11 | ±0% | 14 | ↓29% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 2 | 2 | ±0% | 4 | ↓50% |
| 12_job_crashing (opus-4.6) 📄 | 10 | 9 | ↑11% | 15 | ↓33% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 13 | 12 | ±0% | 15 | ↓13% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 6 | 6 | ±0% | 9 | ↓33% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 7 | 7 | ±0% | 11 | ↓36% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 10 | 13 | ↓23% | 15 | ↓33% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | — | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | — | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | — | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | — | — | — | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 2 | 2 | ±0% | 5 | ↓60% |
| 61_exact_match_counting (opus-4.6) 📄 | 1 | 1 | ±0% | 3 | ↓67% |
| Total (all, n=11) | 6.3 | 7.1 | — | 10.4 | — |
| Comparable (m=10, b=10) | 6.9 | 7.1 | ±0% | 10.4 | ↓34% |
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ Eval Results (with failures)
Automatically triggered by commit 4725ef1 on branch claude/brave-goldberg-2PpZb (labels: evals-id-271_root_cause_buried_in_infra_noise)
Results of HolmesGPT evals
- ask_holmes: 0/1 test cases were successful, 1 regressions
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions | Denied commands | Src |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ❌ | 271_root_cause_buried_in_infra_noise | 137.6s | 10 | 28 | — | — | — | — | — | — | — | — | — | — | — | src |
| Total | 137.6s avg | 10.0 avg | 28.0 avg | — | — | — | — | — | — | — | — | — | — | — |
Benchmark Comparison Details
Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded
- master-27089078550 (created: 2026-06-07)
Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded
- ci-benchmark-27081280174 (created: 2026-06-07)
No baseline data available for comparison.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 1 Failure Detected
📖 Legend
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/brave-goldberg-2PpZb -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
tags: regression
Or with more options (one per line):
/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
tags: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
tags |
Pytest tags / markers (no default - runs all tests!) |
id |
Eval ID / pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):
| Label | Effect |
|---|---|
evals-tag-<name> |
Run tests with tag <name> alongside regression |
evals-id-<name> |
Run a specific eval by test ID |
evals-model-<name> |
Override the model (use model list name, e.g. sonnet-4.5) |
Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5
🏷️ Valid tags
benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs
🤖 Valid models
deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/brave-goldberg-2PpZb -f markers=regression -f filter=
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:84e71ec1c
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:84e71ec1c me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:84e71ec1c
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:84e71ec1c
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:84e71ec1c
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:84e71ec1c me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:84e71ec1c
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:84e71ec1cPatch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:84e71ec1c \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:84e71ec1cRobusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:84e71ec1c \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:84e71ec1c |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (3)
WalkthroughThis PR adds a complete test fixture for scenario 271 that validates Holmes GPT's ability to identify application-layer failures in Airflow KubernetesPodOperator workflows while disregarding infrastructure noise. The fixture includes a test case specification, eight noise pods with impossible scheduling constraints, and a Bash event generation script that injects lifecycle events for task pods alongside infrastructure warning events. ChangesRoot-cause attribution with infrastructure noise test (scenario 271)
Estimated code review effort🎯 2 (Simple) | ⏱️ ~12 minutes Suggested reviewers
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In
`@tests/llm/fixtures/test_ask_holmes/260_root_cause_buried_in_infra_noise/generate_events.sh`:
- Around line 102-115: The generated Event YAML for FailedDraining (metadata
name nodedrain-${n}, kind: Event, reason: FailedDraining) lacks fixture-specific
labels, causing broad cleanup by reason; update the Event creation to add a
unique label (e.g., test-fixture: ask-holmes-260 or nodedrain: true) under
metadata.labels and update any cleanup logic to delete by that label instead of
by reason=FailedDraining so only fixture events are removed. Ensure the label is
present on every generated nodedrain-${n} Event and used in the corresponding
cleanup command.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 196b0dc8-89c2-468f-9025-403636e852f4
📒 Files selected for processing (3)
tests/llm/fixtures/test_ask_holmes/260_root_cause_buried_in_infra_noise/generate_events.shtests/llm/fixtures/test_ask_holmes/260_root_cause_buried_in_infra_noise/noise_pods.yamltests/llm/fixtures/test_ask_holmes/260_root_cause_buried_in_infra_noise/test_case.yaml
- master added 260_global_es_remote_cluster_logs and 261_time_window_gap_external_data after this branch was cut, so rename to 262 and move the namespace to app-262 to avoid a test-number / namespace collision in a shared cluster. - Per CodeRabbit review: tag the injected cluster-scoped FailedDraining node events with a fixture label (holmes-fixture=262-root-cause-noise) and clean them up by label instead of by reason=FailedDraining, so cleanup can't delete unrelated events in a shared cluster. Re-verified red on Claude Opus 4.8 (Score 0%, 0/1) in addition to Opus 4.6. Signed-off-by: Claude <noreply@anthropic.com>
PR #2125 (merged via master) added 262_ambiguous_cluster_reference and a multi-cluster eval series through 270. Move this eval to 271 (namespace app-271, fixture label 271-root-cause-noise) so the two 262_* directories no longer collide. Signed-off-by: Claude <noreply@anthropic.com>
Summary
Adds a new LLM evaluation test case (test 260) that reproduces a scenario where Holmes must correctly identify application-level failures despite significant infrastructure noise in the cluster.
Changes
generate_events.sh: Bash script that creates a realistic event landscape with:
noise_pods.yaml: 8 permanently Pending pods with impossible nodeSelector constraints to simulate high pod density and namespace congestion without consuming actual resources
test_case.yaml: RED evaluation test that validates Holmes can:
Implementation Details
harddifficulty due to the need for chain-of-causation reasoning and distinguishing signal from noisehttps://claude.ai/code/session_01TFEABQjccJGLsb9mZHoRpm
Summary by CodeRabbit