Skip to content

Add LLM eval test case for root cause buried in infra noise - #2150

Open
aantn wants to merge 4 commits into
masterfrom
claude/brave-goldberg-2PpZb
Open

aantn wants to merge 4 commits into
masterfrom
claude/brave-goldberg-2PpZb

Conversation

@aantn

@aantn aantn commented Jun 7, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds a new LLM evaluation test case (test 260) that reproduces a scenario where Holmes must correctly identify application-level failures despite significant infrastructure noise in the cluster.

Changes

  • generate_events.sh: Bash script that creates a realistic event landscape with:

    • Normal lifecycle events (Scheduled/Pulled/Started) for 3 deleted ephemeral task pods, proving they successfully ran
    • ~40 FailedCreatePodSandBox events from AWS CNI IP exhaustion on unrelated pods
    • ~10 FailedDraining events from Karpenter node churn
  • noise_pods.yaml: 8 permanently Pending pods with impossible nodeSelector constraints to simulate high pod density and namespace congestion without consuming actual resources

  • test_case.yaml: RED evaluation test that validates Holmes can:

    • Recognize that the 3 task pods (log-archival-index-to-es-7q4w9z-{1,2,3}) were successfully scheduled and started based on their Normal lifecycle events
    • Correctly attribute the failures to application-level causes (not infrastructure)
    • Explain that application logs are unavailable because the ephemeral pods were deleted
    • Distinguish between unrelated cluster noise (IP exhaustion, node churn) and the actual root cause
    • Direct users to Airflow task logs for the real error details

Implementation Details

  • The test reproduces a real-world scenario where infrastructure noise can mislead root cause analysis
  • Task pods are intentionally NOT created (only their events remain) to simulate the KubernetesPodOperator's default on-finish deletion behavior
  • Verification steps ensure both the "proof of execution" events and the noise events are properly injected
  • Marked as hard difficulty due to the need for chain-of-causation reasoning and distinguishing signal from noise

https://claude.ai/code/session_01TFEABQjccJGLsb9mZHoRpm

Summary by CodeRabbit

  • Tests
    • Added test fixtures to validate root cause identification in complex scenarios with infrastructure noise present in logs.

Reproduces the failure mode where Holmes attributes a failing Airflow
KubernetesPodOperator task to dramatic but unrelated cluster-wide
infrastructure noise (AWS-CNI IP-address-exhaustion FailedCreatePodSandBox
events, node draining / Karpenter churn, high pod density) instead of
recognizing the task pods actually ran and failed at the application level.

The 3 ephemeral task pods (log-archival-index-to-es-7q4w9z-*) were Scheduled,
Pulled and Started (their Normal lifecycle events prove they ran) and were then
deleted by the operator, so kubectl logs is unavailable and the real error
lives in the Airflow task logs. The correct conclusion is application-level
failure + unavailable pod logs, with the infra storm being unrelated context.

Verified red on Claude Opus 4.6 (Score 0%, 0/1) across repeated runs: Holmes
consistently pins causation on the subnet IP exhaustion (e.g. "the IP
exhaustion may have caused ... inability to reach Elasticsearch") rather than
treating it as unrelated noise.

Tagged hard (a known-failing reasoning-quality case).

Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@github-actions

github-actions Bot commented Jun 7, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 #3 · Run @ __4725ef1__ (#27096911824) — Jun 7, 15:39 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 4725ef1 on branch claude/brave-goldberg-2PpZb (labels: evals-id-262)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 1/1 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Denied commands Src
✅ 262_ambiguous_cluster_reference 64.5s 9 17 $0.3143 136,678 132,517 19,515 4,161 753 112,283 20,234 176 — — src
Total 64.5s avg 9.0 avg 17.0 avg $0.3143 136,678 132,517 19,515 4,161 753 112,283 20,234 176 — —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded

No baseline data available for comparison.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #2 · Run @ __1124f7e__ (#27090599665) — Jun 7, 11:13 UTC

⚠️ Eval Results (with failures)

Automatically triggered by commit 1124f7e on branch claude/brave-goldberg-2PpZb (labels: evals-id-262)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 1/2 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Src
✅ 262_ambiguous_cluster_reference 29.8s 5 7 $0.1582 56,711 55,011 12,626 1,700 603 42,124 12,887 67 — src
❌ 262_root_cause_buried_in_infra_noise 128.4s 10 31 — — — — — — — — — — src
Total 79.1s avg 7.5 avg 19.0 avg $0.1582 56,711 55,011 12,626 1,700 603 42,124 12,887 67 —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded

No baseline data available for comparison.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected

📜 #1 · Run @ __1124f7e__ (#27090566915) — Jun 7, 11:06 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 1124f7e on branch claude/brave-goldberg-2PpZb

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/14 test cases were successful, 0 regressions, 3 setup failures
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Src
✅ 09_crashpod 29.5s 4 8 $0.2146 68,390 66,616 19,169 1,774 693 45,848 20,768 222 — src
✅ 101_loki_historical_logs_pod_deleted 46.7s 5 10 $0.2467 87,613 84,979 19,990 2,634 680 64,699 20,280 377 — src
✅ 112_find_pvcs_by_uuid 13.3s 2 2 $0.1502 32,181 31,356 17,492 825 448 13,861 17,495 238 — src
✅ 12_job_crashing 30.9s 5 10 $0.2368 92,585 90,747 20,812 1,838 541 68,986 21,761 180 — src
✅ 176_network_policy_blocking_traffic_no_skills 37.0s 5 13 $0.2434 89,942 87,636 20,374 2,306 676 66,667 20,969 375 — src
✅ 227_count_configmaps_per_namespace[0] 13.8s 3 6 $0.1448 45,002 44,290 16,213 712 438 28,073 16,217 36 — src
✅ 243_pod_names_contain_service 28.2s 3 7 $0.1852 48,429 46,733 17,629 1,696 571 28,634 18,099 212 — src
✅ 24_misconfigured_pvc 25.8s 4 10 $0.2021 66,474 64,873 18,573 1,601 597 45,263 19,610 49 — src
🚧 254_elasticsearch_dr_test_log_check — — — — — — — — — — — — — src
🚧 259_wrong_cluster_logs_confusion — — — — — — — — — — — — — src
🚧 260_global_es_remote_cluster_logs — — — — — — — — — — — — — src
✅ 43_current_datetime_from_prompt 3.8s 1 — $0.0986 13,970 13,848 13,848 122 122 0 13,848 78 — src
✅ 51_logs_summarize_errors 19.7s 3 2 $0.1543 45,815 45,016 17,005 799 424 28,007 17,009 31 — src
✅ 61_exact_match_counting 7.1s 2 1 $0.1113 28,235 28,012 14,184 223 154 13,825 14,187 36 — src
Total 23.2s avg 3.4 avg 6.9 avg $1.9879 618,636 604,106 20,812 14,530 693 403,863 200,243 1,834 —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (1h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 29.5s 26.2s ↑12% 41.4s ↓29%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 46.7s 47.6s ±0% 65.7s ↓29%
112_find_pvcs_by_uuid (opus-4.6) 📄 13.3s 10.6s ↑26% 19.0s ↓30%
12_job_crashing (opus-4.6) 📄 30.9s 27.6s ↑12% 42.6s ↓27%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 37.0s 32.6s ↑13% 44.3s ↓17%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 13.8s 13.2s ±0% 18.4s ↓25%
243_pod_names_contain_service (opus-4.6) 📄 28.2s 28.8s ±0% 35.1s ↓19%
24_misconfigured_pvc (opus-4.6) 📄 25.8s 31.4s ↓18% 40.6s ↓36%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 — — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 — — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 — — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 3.8s 3.1s ↑23% 3.5s ±0%
51_logs_summarize_errors (opus-4.6) 📄 19.7s 18.8s ±0% 19.7s ±0%
61_exact_match_counting (opus-4.6) 📄 7.1s 7.1s ±0% 10.2s ↓31%
Total (all, n=11) 23.2s 22.5s — 31.0s —
Comparable (m=11, b=11) 23.2s 22.5s ±0% 31.0s ↓25%

Cost comparison:

Test case This branch master (1h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.2146 $0.2032 ±0% $0.3101 ↓31%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.2467 $0.2642 ±0% $0.3829 ↓36%
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.1502 $0.1368 ±0% $0.2047 ↓27%
12_job_crashing (opus-4.6) 📄 $0.2368 $0.2173 ±0% $0.3167 ↓25%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.2434 $0.2524 ±0% $0.3179 ↓23%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.1448 $0.1464 ±0% $0.2048 ↓29%
243_pod_names_contain_service (opus-4.6) 📄 $0.1852 $0.1872 ±0% $0.2662 ↓30%
24_misconfigured_pvc (opus-4.6) 📄 $0.2021 $0.2240 ±0% $0.3110 ↓35%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 — — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 — — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 — — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 $0.0986 $0.0110 ↑797% $0.1190 ↓17%
51_logs_summarize_errors (opus-4.6) 📄 $0.1543 $0.1488 ±0% $0.2018 ↓24%
61_exact_match_counting (opus-4.6) 📄 $0.1113 $0.1112 ±0% $0.1518 ↓27%
Total (all, n=14) $0.1420 $0.1730 — $0.2534 —
Comparable (m=11, b=11) $0.1807 $0.1730 ±0% $0.2534 ↓29%

Total tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 68,390 67,778 ±0% 132,351 ↓48%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 87,613 88,307 ±0% 144,907 ↓40%
112_find_pvcs_by_uuid (opus-4.6) 📄 32,181 30,686 ±0% 61,219 ↓47%
12_job_crashing (opus-4.6) 📄 92,585 69,487 ↑33% 136,570 ↓32%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 89,942 108,726 ↓17% 113,514 ↓21%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 45,002 44,988 ±0% 76,715 ↓41%
243_pod_names_contain_service (opus-4.6) 📄 48,429 48,612 ±0% 104,698 ↓54%
24_misconfigured_pvc (opus-4.6) 📄 66,474 68,923 ±0% 132,672 ↓50%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 — — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 — — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 — — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 13,970 13,970 ±0% 17,001 ↓18%
51_logs_summarize_errors (opus-4.6) 📄 45,815 45,163 ±0% 77,120 ↓41%
61_exact_match_counting (opus-4.6) 📄 28,235 28,231 ±0% 52,716 ↓46%
Total (all, n=14) 44,188 55,897 — 95,408 —
Comparable (m=11, b=11) 56,240 55,897 ±0% 95,408 ↓41%

Cached tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 45,848 46,700 ±0% 102,543 ↓55%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 64,699 64,209 ±0% 109,927 ↓41%
112_find_pvcs_by_uuid (opus-4.6) 📄 13,861 13,861 ±0% 38,090 ↓64%
12_job_crashing (opus-4.6) 📄 68,986 46,249 ↑49% 107,265 ↓36%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 66,667 84,274 ↓21% 82,679 ↓19%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 28,073 28,072 ±0% 54,602 ↓49%
243_pod_names_contain_service (opus-4.6) 📄 28,634 28,690 ±0% 78,754 ↓64%
24_misconfigured_pvc (opus-4.6) 📄 45,263 46,110 ±0% 103,729 ↓56%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 — — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 — — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 — — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 — 13,845 — — —
51_logs_summarize_errors (opus-4.6) 📄 28,007 28,012 ±0% 55,156 ↓49%
61_exact_match_counting (opus-4.6) 📄 13,825 13,825 ±0% 34,481 ↓60%
Total (all, n=11) 36,715 37,622 — 69,748 —
Comparable (m=10, b=10) 40,386 40,000 ±0% 76,723 ↓47%

Turns comparison:

Test case This branch master (1h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 4 4 ±0% 6 ↓33%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 5 5 ±0% 6 ↓17%
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% 3 ↓33%
12_job_crashing (opus-4.6) 📄 5 4 ↑25% 6 ↓17%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 5 6 ↓17% 5 ±0%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 3 3 ±0% 4 ↓25%
243_pod_names_contain_service (opus-4.6) 📄 3 3 ±0% 5 ↓40%
24_misconfigured_pvc (opus-4.6) 📄 4 4 ±0% 6 ↓33%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 — — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 — — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 — — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% 1 ±0%
51_logs_summarize_errors (opus-4.6) 📄 3 3 ±0% 4 ↓25%
61_exact_match_counting (opus-4.6) 📄 2 2 ±0% 3 ↓33%
Total (all, n=11) 3.4 3.4 — 4.5 —
Comparable (m=11, b=11) 3.4 3.4 ±0% 4.5 ↓24%

Tool calls comparison:

Test case This branch master (1h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 8 8 ±0% 13 ↓38%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 10 11 ±0% 14 ↓29%
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% 4 ↓50%
12_job_crashing (opus-4.6) 📄 10 9 ↑11% 15 ↓33%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 13 12 ±0% 15 ↓13%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 6 6 ±0% 9 ↓33%
243_pod_names_contain_service (opus-4.6) 📄 7 7 ±0% 11 ↓36%
24_misconfigured_pvc (opus-4.6) 📄 10 13 ↓23% 15 ↓33%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 — — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 — — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 — — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 2 2 ±0% 5 ↓60%
61_exact_match_counting (opus-4.6) 📄 1 1 ±0% 3 ↓67%
Total (all, n=11) 6.3 7.1 — 10.4 —
Comparable (m=10, b=10) 6.9 7.1 ±0% 10.4 ↓34%

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ Eval Results (with failures)

Automatically triggered by commit 4725ef1 on branch claude/brave-goldberg-2PpZb (labels: evals-id-271_root_cause_buried_in_infra_noise)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 0/1 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Denied commands Src
❌ 271_root_cause_buried_in_infra_noise 137.6s 10 28 — — — — — — — — — — — src
Total 137.6s avg 10.0 avg 28.0 avg — — — — — — — — — — —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded

No baseline data available for comparison.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/brave-goldberg-2PpZb -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/brave-goldberg-2PpZb -f markers=regression -f filter=

@github-actions

github-actions Bot commented Jun 7, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 84e71ec1c (built in 4m 52s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:84e71ec1c
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:84e71ec1c me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:84e71ec1c
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:84e71ec1c
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:84e71ec1c
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:84e71ec1c me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:84e71ec1c
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:84e71ec1c

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:84e71ec1c \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:84e71ec1c

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:84e71ec1c \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:84e71ec1c

@netlify

netlify Bot commented Jun 7, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 4725ef1
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a258f52e667110008b184b5
😎 Deploy Preview https://deploy-preview-2150--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai

coderabbitai Bot commented Jun 7, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 1fee277b-4e6b-4b33-99e8-051ea873de24

📥 Commits

Reviewing files that changed from the base of the PR and between 421bccc and 4725ef1.

📒 Files selected for processing (3)
  • tests/llm/fixtures/test_ask_holmes/271_root_cause_buried_in_infra_noise/generate_events.sh
  • tests/llm/fixtures/test_ask_holmes/271_root_cause_buried_in_infra_noise/noise_pods.yaml
  • tests/llm/fixtures/test_ask_holmes/271_root_cause_buried_in_infra_noise/test_case.yaml

Walkthrough

This PR adds a complete test fixture for scenario 271 that validates Holmes GPT's ability to identify application-layer failures in Airflow KubernetesPodOperator workflows while disregarding infrastructure noise. The fixture includes a test case specification, eight noise pods with impossible scheduling constraints, and a Bash event generation script that injects lifecycle events for task pods alongside infrastructure warning events.

Changes

Root-cause attribution with infrastructure noise test (scenario 271)

Layer / File(s) Summary
Test case specification and expected behavior
tests/llm/fixtures/test_ask_holmes/271_root_cause_buried_in_infra_noise/test_case.yaml
Test scenario for scenario 271 defining an Airflow DAG with three failed task pods. Specifies namespace setup, event generation orchestration, assertions for presence of lifecycle events (7q4w9z identifier) and infrastructure noise (FailedCreatePodSandBox), absence of ephemeral task pods, and expected output requiring Holmes to attribute failures to Airflow application logic while forbidding cluster infrastructure blame (CNI exhaustion, node draining, Karpenter).
Infrastructure noise pod fixture
tests/llm/fixtures/test_ask_holmes/271_root_cause_buried_in_infra_noise/noise_pods.yaml
Kubernetes manifest with eight Pending pods (ingest-worker-0 through ingest-worker-7) in namespace app-271, each using an impossible nodepool nodeSelector and busybox sleep container to generate scheduling-related infrastructure warning events as test noise.
Kubernetes event generation script
tests/llm/fixtures/test_ask_holmes/271_root_cause_buried_in_infra_noise/generate_events.sh
Bash script that generates and applies a Kubernetes Event list containing three lifecycle events (Scheduled/Pulled/Started) for task pods identified by 7q4w9z, multiple FailedCreatePodSandBox warnings across hosts, and FailedDraining node-scoped warnings. Script includes emit_lifecycle helper, shared UTC timestamps, and kubectl event injection.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~12 minutes

Suggested reviewers

  • moshemorad
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title clearly and specifically describes the main change: adding a new LLM eval test case for the scenario where root causes are obscured by infrastructure noise.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@tests/llm/fixtures/test_ask_holmes/260_root_cause_buried_in_infra_noise/generate_events.sh`:
- Around line 102-115: The generated Event YAML for FailedDraining (metadata
name nodedrain-${n}, kind: Event, reason: FailedDraining) lacks fixture-specific
labels, causing broad cleanup by reason; update the Event creation to add a
unique label (e.g., test-fixture: ask-holmes-260 or nodedrain: true) under
metadata.labels and update any cleanup logic to delete by that label instead of
by reason=FailedDraining so only fixture events are removed. Ensure the label is
present on every generated nodedrain-${n} Event and used in the corresponding
cleanup command.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 196b0dc8-89c2-468f-9025-403636e852f4

📥 Commits

Reviewing files that changed from the base of the PR and between 478a181 and 7b75911.

📒 Files selected for processing (3)
  • tests/llm/fixtures/test_ask_holmes/260_root_cause_buried_in_infra_noise/generate_events.sh
  • tests/llm/fixtures/test_ask_holmes/260_root_cause_buried_in_infra_noise/noise_pods.yaml
  • tests/llm/fixtures/test_ask_holmes/260_root_cause_buried_in_infra_noise/test_case.yaml

claude and others added 2 commits June 7, 2026 10:39
- master added 260_global_es_remote_cluster_logs and 261_time_window_gap_external_data
  after this branch was cut, so rename to 262 and move the namespace to app-262
  to avoid a test-number / namespace collision in a shared cluster.
- Per CodeRabbit review: tag the injected cluster-scoped FailedDraining node
  events with a fixture label (holmes-fixture=262-root-cause-noise) and clean
  them up by label instead of by reason=FailedDraining, so cleanup can't delete
  unrelated events in a shared cluster.

Re-verified red on Claude Opus 4.8 (Score 0%, 0/1) in addition to Opus 4.6.

Signed-off-by: Claude <noreply@anthropic.com>
PR #2125 (merged via master) added 262_ambiguous_cluster_reference and a
multi-cluster eval series through 270. Move this eval to 271 (namespace
app-271, fixture label 271-root-cause-noise) so the two 262_* directories
no longer collide.

Signed-off-by: Claude <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants