Skip to content

multi-cluster awareness evals (red) — tests/llm + pyproject - #2125

Merged
aantn merged 5 commits into
masterfrom
claude/multi-cluster-evals-red
Jun 7, 2026
Merged

aantn merged 5 commits into
masterfrom
claude/multi-cluster-evals-red

Conversation

@aantn

@aantn aantn commented Jun 4, 2026 •

Copy link
Copy Markdown
Collaborator

Splits PR #2042 into two: this is the red half. Tests/evals + supporting cleanup, no prompt change.

Without the companion green PR (#TBD — claude/multi-cluster-prompt-green), the new evals are expected to be RED on master. They demonstrate the desired behavior; the prompt change in the green PR turns them green.

What this PR contains

13 evals (all tagged multi-cluster)

# Test Scenario
254 (updated) elasticsearch_dr_test_log_check Similar-name service disambiguation (companyopswebjob vs Company.Ops*).
19 (updated) detect_missing_app_details personal-certs-validator vs db-certs-authenticator — transparency + usefulness both required.
259 wrong_cluster_logs_confusion Holmes on production-eu-west-2, user asks about production-us-east-2; ES has only local data with red-herring 500s.
260 global_es_remote_cluster_logs Same shape, but ES has remote-cluster data (global topology) — must investigate normally.
261 time_window_gap_external_data Data exists but only recent; user asks about older incident.
262 ambiguous_cluster_reference cluster_name="production"; user says "the US prod cluster".
263 region_suffixed_service_siblings payment-service-us/-eu/-ap; user says "payment-service".
264 wrong_env_same_region Holmes on staging-eu-west-2; user asks about production-eu-west-2.
265 multicluster_labeling_discipline Open-ended fleet question; every finding must be cluster-attributed.
266 toolset_disabled_vs_no_data Datadog APM requested, toolset disabled — must distinguish from "found nothing".
267 cluster_name_alias User says "EU prod" for acme-prod-eu-west-1 — must NOT over-correct.
268 (KIND) kubectl_wrong_cluster kubectl-only path; non-local cluster asked about.
269 (KIND+ES) mixed_sources_route_to_external Must route to external backend for non-local cluster.
270 (KIND) namespace_collision Same deployment name in 2 namespaces.

Run the whole set with:

/eval
tags: multi-cluster

Supporting changes

  • tests/llm/test_ask_holmes.py: pass test_case.cluster_name through to build_initial_ask_messages in the CLI test path. Previously it was silently dropped, so cluster_name in test fixtures had no effect on the system prompt for ask-mode evals.
  • tests/llm/utils/braintrust.py + tests/llm/utils/langfuse.py: purge legacy dataset-upload code (BraintrustEvalHelper, upload_test_cases, find_dataset_row_by_test_case, etc.). We don't use Braintrust datasets in this repo — all eval logging goes via experiment span logs. The legacy code also had a real bug (input=input where input was the Python builtin) that corrupted dataset records.
  • tests/llm/test_holmes_checks.py: drop dataset_record_id= from eval_span.log calls (we no longer link to a dataset).
  • pyproject.toml: register the new multi-cluster pytest marker.

Companion PR

Merge AFTER the green PR (claude/multi-cluster-prompt-green), which adds the cluster-awareness rule to holmes/plugins/prompts/generic_ask.jinja2. Order doesn't matter for the git merge, but red-then-green will show the red→green transition in CI; green-first will show no red phase.


Generated by Claude Code

Summary by CodeRabbit

  • Tests
    • Added and enhanced 13+ LLM test scenarios covering multi-cluster, cross-environment, ambiguous service/namespace, wrong-cluster and mixed-source routing cases; tightened expectations so responses must transparently state missing/mismatched data and clearly attribute findings.
  • Chores
    • Added a pytest marker for multi-cluster evaluations.
    • Removed external dataset utilities and simplified evaluation logging.

This PR introduces 13 new/updated evals that exercise Holmes's
multi-cluster / multi-environment / multi-source awareness, plus
supporting test framework cleanup. Without the corresponding prompt
fix (separate PR) these evals are expected to be RED on master —
they demonstrate the behavior we want Holmes to have, before the
prompt change lands.

Evals (all tagged `multi-cluster`):

  254_elasticsearch_dr_test_log_check (updated)
      Service-name disambiguation. companyopswebjob is asked about;
      only Company.Ops and Company.Ops.Radar.WebJob exist. Must say
      no exact match, may report adjacent findings if clearly labeled.

  19_detect_missing_app_details (updated)
      personal-certs-validator vs db-certs-authenticator. Both
      transparency (no exact match) AND usefulness (label the
      similar-named resource) required.

  259_wrong_cluster_logs_confusion (new)
      Holmes on production-eu-west-2, user asks about
      production-us-east-2. ES index has only eu-west-2 data with
      red-herring 500s. Must flag cluster mismatch up-front.

  260_global_es_remote_cluster_logs (new)
      Counterpart to 259 — same scenario but ES has us-east-2 data
      (global topology). Holmes must investigate normally, not
      refuse / over-correct.

  261_time_window_gap_external_data (new)
      Data exists but only for the last 30 min; user asks about an
      incident 6h ago. Must flag the time-window gap.

  262_ambiguous_cluster_reference (new)
      cluster_name="production"; user says "the US prod cluster".
      Must not silently treat local as the meant cluster.

  263_region_suffixed_service_siblings (new)
      payment-service-us / -eu / -ap. User asks about
      "payment-service" — must surface the ambiguity, not silently
      pick one.

  264_wrong_env_same_region (new)
      Holmes on staging-eu-west-2; user asks about
      production-eu-west-2. Must catch env mismatch.

  265_multicluster_labeling_discipline (new)
      Open-ended "what's broken across our fleet?" — every finding
      in the answer must be attributed to a specific cluster.

  266_toolset_disabled_vs_no_data (new)
      User asks Holmes to check Datadog APM; Datadog toolset is
      disabled. Must say "can't access" not "looked and found
      nothing."

  267_cluster_name_alias (new)
      cluster_name="acme-prod-eu-west-1"; user says "EU prod". Must
      recognise the alias as the same cluster and investigate — must
      NOT refuse / over-correct from the cluster-awareness rule.

  268_kubectl_wrong_cluster (new, KIND)
      kubectl-only path; user asks about non-local cluster. Must
      flag kubectl-local limitation.

  269_mixed_sources_route_to_external (new, KIND + ES)
      Holmes has kubectl (local only) + multi-cluster ES. User asks
      about non-local cluster. Must route to ES, not blame healthy
      local kubectl state.

  270_namespace_collision (new, KIND)
      Same-named deployment in two namespaces with different states.
      Must surface ambiguity and address both, not silently pick one.

Supporting changes:

  tests/llm/test_ask_holmes.py: pass test_case.cluster_name through
  to build_initial_ask_messages in the CLI test path. Previously it
  was silently dropped, so cluster_name in test fixtures had no
  effect on the system prompt for ask-mode evals.

  tests/llm/utils/braintrust.py + tests/llm/utils/langfuse.py: purge
  legacy Braintrust dataset upload code (BraintrustEvalHelper /
  upload_test_cases / find_dataset_row_by_test_case / etc.) and the
  dead langfuse upload helper. We don't use Braintrust datasets in
  this repo — all eval logging goes via experiment span logs. The
  legacy code also had a real bug (input=input where `input` was
  the Python builtin) that corrupted dataset records.

  tests/llm/test_holmes_checks.py: drop dataset_record_id= from
  eval_span.log calls (we no longer link to a dataset).

  pyproject.toml: register the new `multi-cluster` pytest marker.

Counterpart PR (green): the prompt-side change in
holmes/plugins/prompts/generic_ask.jinja2 that teaches Holmes about
multi-cluster scope awareness. Once that PR is merged, these evals
turn green.

Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@github-actions

github-actions Bot commented Jun 4, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

⚠️ 1 older run truncated

Older runs were omitted to stay under GitHub's 64KB comment size limit.


⚠️ Eval Results (with failures)

Automatically triggered by commit 08da31e on branch claude/multi-cluster-evals-red (labels: evals-tag-multi-cluster)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 21/24 test cases were successful, 3 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Src
✅ 09_crashpod 26.1s 4 8 $0.2032 67,826 66,217 18,927 1,609 586 46,714 19,503 158 — src
✅ 101_loki_historical_logs_pod_deleted 38.8s 4 9 $0.2251 68,396 66,078 19,545 2,318 745 45,981 20,097 212 — src
✅ 112_find_pvcs_by_uuid 14.7s 2 2 $0.1551 32,357 31,355 17,491 1,002 554 13,861 17,494 294 — src
✅ 12_job_crashing 29.6s 4 9 $0.2196 71,808 70,065 20,437 1,743 526 48,729 21,336 181 — src
✅ 176_network_policy_blocking_traffic_no_skills 27.5s 4 11 $0.2121 69,765 67,987 19,929 1,778 590 48,053 19,934 219 — src
✅ 227_count_configmaps_per_namespace[0] 15.2s 3 6 $0.1451 44,959 44,261 16,195 698 432 28,062 16,199 29 — src
✅ 243_pod_names_contain_service 27.0s 3 7 $0.1835 48,364 46,754 17,639 1,610 573 28,648 18,106 248 — src
✅ 24_misconfigured_pvc 31.3s 5 12 $0.2270 86,558 84,593 19,217 1,965 572 64,398 20,195 125 — src
✅ 254_elasticsearch_dr_test_log_check 53.0s 8 12 $0.2674 104,450 100,704 17,233 3,746 801 83,462 17,242 204 — src
❌ 259_wrong_cluster_logs_confusion 56.7s 9 12 $0.2736 121,067 117,517 17,383 3,550 650 99,846 17,671 153 — src
✅ 260_global_es_remote_cluster_logs 51.0s 8 11 $0.2544 107,504 104,375 17,050 3,129 729 86,696 17,679 258 — src
✅ 261_time_window_gap_external_data 65.5s 11 13 $0.2847 139,707 135,945 16,841 3,762 850 118,822 17,123 532 — src
✅ 262_ambiguous_cluster_reference 54.8s 8 15 $0.2817 109,723 105,985 18,098 3,738 746 86,639 19,346 172 — src
✅ 263_region_suffixed_service_siblings 48.4s 8 9 $0.2186 94,762 92,142 14,775 2,620 730 76,887 15,255 80 — src
✅ 264_wrong_env_same_region 53.6s 9 11 $0.2479 113,854 110,673 15,482 3,181 822 94,868 15,805 266 — src
✅ 265_multicluster_labeling_discipline 47.8s 8 9 $0.2314 97,723 94,769 15,497 2,954 841 79,263 15,506 308 — src
✅ 266_toolset_disabled_vs_no_data 11.3s 1 — $0.0733 9,694 9,369 9,369 325 325 0 9,369 91 — src
✅ 267_cluster_name_alias 63.5s 9 18 $0.3172 129,548 125,059 19,239 4,489 920 105,116 19,943 266 — src
❌ 268_kubectl_wrong_cluster 27.3s 3 6 $0.2013 52,335 50,599 19,619 1,736 597 30,443 20,156 466 — src
✅ 269_mixed_sources_route_to_external 60.6s 6 14 $0.2749 90,618 86,701 18,748 3,917 1,252 67,728 18,973 908 — src
❌ 270_namespace_collision 20.5s 3 6 $0.1701 47,620 46,430 17,589 1,190 505 28,643 17,787 132 — src
✅ 43_current_datetime_from_prompt 3.2s 1 — $0.0986 13,970 13,848 13,848 122 122 0 13,848 78 — src
✅ 51_logs_summarize_errors 19.1s 3 2 $0.1505 45,401 44,650 16,639 751 390 28,007 16,643 29 — src
✅ 61_exact_match_counting 8.7s 2 1 $0.1112 28,231 28,010 14,182 221 152 13,825 14,185 34 — src
Total 35.6s avg 5.2 avg 9.2 avg $5.0274 1,796,240 1,744,086 20,437 52,154 1,252 1,324,691 419,395 5,443 —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (1h ago) Δ vs master benchmark (6h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 26.1s 26.2s ±0% 41.4s ↓37%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 38.8s 47.6s ↓19% 65.7s ↓41%
112_find_pvcs_by_uuid (opus-4.6) 📄 14.7s 10.6s ↑39% 19.0s ↓23%
12_job_crashing (opus-4.6) 📄 29.6s 27.6s ±0% 42.6s ↓30%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 27.5s 32.6s ↓16% 44.3s ↓38%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 15.2s 13.2s ↑15% 18.4s ↓18%
243_pod_names_contain_service (opus-4.6) 📄 27.0s 28.8s ±0% 35.1s ↓23%
24_misconfigured_pvc (opus-4.6) 📄 31.3s 31.4s ±0% 40.6s ↓23%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 53.0s — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 56.7s — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 51.0s — — — —
261_time_window_gap_external_data (opus-4.6) 📄 65.5s — — — —
262_ambiguous_cluster_reference (opus-4.6) 📄 54.8s — — — —
263_region_suffixed_service_siblings (opus-4.6) 📄 48.4s — — — —
264_wrong_env_same_region (opus-4.6) 📄 53.6s — — — —
265_multicluster_labeling_discipline (opus-4.6) 📄 47.8s — — — —
266_toolset_disabled_vs_no_data (opus-4.6) 📄 11.3s — — — —
267_cluster_name_alias (opus-4.6) 📄 63.5s — — — —
268_kubectl_wrong_cluster (opus-4.6) 📄 27.3s — — — —
269_mixed_sources_route_to_external (opus-4.6) 📄 60.6s — — — —
270_namespace_collision (opus-4.6) 📄 20.5s — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 3.2s 3.1s ±0% 3.5s ±0%
51_logs_summarize_errors (opus-4.6) 📄 19.1s 18.8s ±0% 19.7s ±0%
61_exact_match_counting (opus-4.6) 📄 8.7s 7.1s ↑22% 10.2s ↓15%
Total (all, n=24) 35.6s 22.5s — 31.0s —
Comparable (m=11, b=11) 21.9s 22.5s ±0% 31.0s ↓29%

Cost comparison:

Test case This branch master (1h ago) Δ vs master benchmark (6h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.2032 $0.2032 ±0% $0.3101 ↓34%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.2251 $0.2642 ↓15% $0.3829 ↓41%
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.1551 $0.1368 ↑13% $0.2047 ↓24%
12_job_crashing (opus-4.6) 📄 $0.2196 $0.2173 ±0% $0.3167 ↓31%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.2121 $0.2524 ↓16% $0.3179 ↓33%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.1451 $0.1464 ±0% $0.2048 ↓29%
243_pod_names_contain_service (opus-4.6) 📄 $0.1835 $0.1872 ±0% $0.2662 ↓31%
24_misconfigured_pvc (opus-4.6) 📄 $0.2270 $0.2240 ±0% $0.3110 ↓27%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 $0.2674 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 $0.2736 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 $0.2544 — — — —
261_time_window_gap_external_data (opus-4.6) 📄 $0.2847 — — — —
262_ambiguous_cluster_reference (opus-4.6) 📄 $0.2817 — — — —
263_region_suffixed_service_siblings (opus-4.6) 📄 $0.2186 — — — —
264_wrong_env_same_region (opus-4.6) 📄 $0.2479 — — — —
265_multicluster_labeling_discipline (opus-4.6) 📄 $0.2314 — — — —
266_toolset_disabled_vs_no_data (opus-4.6) 📄 $0.0733 — — — —
267_cluster_name_alias (opus-4.6) 📄 $0.3172 — — — —
268_kubectl_wrong_cluster (opus-4.6) 📄 $0.2013 — — — —
269_mixed_sources_route_to_external (opus-4.6) 📄 $0.2749 — — — —
270_namespace_collision (opus-4.6) 📄 $0.1701 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 $0.0986 $0.0110 ↑797% $0.1190 ↓17%
51_logs_summarize_errors (opus-4.6) 📄 $0.1505 $0.1488 ±0% $0.2018 ↓25%
61_exact_match_counting (opus-4.6) 📄 $0.1112 $0.1112 ±0% $0.1518 ↓27%
Total (all, n=24) $0.2095 $0.1730 — $0.2534 —
Comparable (m=11, b=11) $0.1755 $0.1730 ±0% $0.2534 ↓31%

Total tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (6h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 67,826 67,778 ±0% 132,351 ↓49%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 68,396 88,307 ↓23% 144,907 ↓53%
112_find_pvcs_by_uuid (opus-4.6) 📄 32,357 30,686 ±0% 61,219 ↓47%
12_job_crashing (opus-4.6) 📄 71,808 69,487 ±0% 136,570 ↓47%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 69,765 108,726 ↓36% 113,514 ↓39%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 44,959 44,988 ±0% 76,715 ↓41%
243_pod_names_contain_service (opus-4.6) 📄 48,364 48,612 ±0% 104,698 ↓54%
24_misconfigured_pvc (opus-4.6) 📄 86,558 68,923 ↑26% 132,672 ↓35%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 104,450 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 121,067 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 107,504 — — — —
261_time_window_gap_external_data (opus-4.6) 📄 139,707 — — — —
262_ambiguous_cluster_reference (opus-4.6) 📄 109,723 — — — —
263_region_suffixed_service_siblings (opus-4.6) 📄 94,762 — — — —
264_wrong_env_same_region (opus-4.6) 📄 113,854 — — — —
265_multicluster_labeling_discipline (opus-4.6) 📄 97,723 — — — —
266_toolset_disabled_vs_no_data (opus-4.6) 📄 9,694 — — — —
267_cluster_name_alias (opus-4.6) 📄 129,548 — — — —
268_kubectl_wrong_cluster (opus-4.6) 📄 52,335 — — — —
269_mixed_sources_route_to_external (opus-4.6) 📄 90,618 — — — —
270_namespace_collision (opus-4.6) 📄 47,620 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 13,970 13,970 ±0% 17,001 ↓18%
51_logs_summarize_errors (opus-4.6) 📄 45,401 45,163 ±0% 77,120 ↓41%
61_exact_match_counting (opus-4.6) 📄 28,231 28,231 ±0% 52,716 ↓46%
Total (all, n=24) 74,843 55,897 — 95,408 —
Comparable (m=11, b=11) 52,512 55,897 ±0% 95,408 ↓45%

Cached tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (6h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 46,714 46,700 ±0% 102,543 ↓54%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 45,981 64,209 ↓28% 109,927 ↓58%
112_find_pvcs_by_uuid (opus-4.6) 📄 13,861 13,861 ±0% 38,090 ↓64%
12_job_crashing (opus-4.6) 📄 48,729 46,249 ±0% 107,265 ↓55%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 48,053 84,274 ↓43% 82,679 ↓42%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 28,062 28,072 ±0% 54,602 ↓49%
243_pod_names_contain_service (opus-4.6) 📄 28,648 28,690 ±0% 78,754 ↓64%
24_misconfigured_pvc (opus-4.6) 📄 64,398 46,110 ↑40% 103,729 ↓38%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 83,462 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 99,846 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 86,696 — — — —
261_time_window_gap_external_data (opus-4.6) 📄 118,822 — — — —
262_ambiguous_cluster_reference (opus-4.6) 📄 86,639 — — — —
263_region_suffixed_service_siblings (opus-4.6) 📄 76,887 — — — —
264_wrong_env_same_region (opus-4.6) 📄 94,868 — — — —
265_multicluster_labeling_discipline (opus-4.6) 📄 79,263 — — — —
266_toolset_disabled_vs_no_data (opus-4.6) 📄 — — — — —
267_cluster_name_alias (opus-4.6) 📄 105,116 — — — —
268_kubectl_wrong_cluster (opus-4.6) 📄 30,443 — — — —
269_mixed_sources_route_to_external (opus-4.6) 📄 67,728 — — — —
270_namespace_collision (opus-4.6) 📄 28,643 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 — 13,845 — — —
51_logs_summarize_errors (opus-4.6) 📄 28,007 28,012 ±0% 55,156 ↓49%
61_exact_match_counting (opus-4.6) 📄 13,825 13,825 ±0% 34,481 ↓60%
Total (all, n=24) 55,195 37,622 — 69,748 —
Comparable (m=10, b=10) 36,628 40,000 ±0% 76,723 ↓52%

Turns comparison:

Test case This branch master (1h ago) Δ vs master benchmark (6h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 4 4 ±0% 6 ↓33%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 4 5 ↓20% 6 ↓33%
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% 3 ↓33%
12_job_crashing (opus-4.6) 📄 4 4 ±0% 6 ↓33%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 4 6 ↓33% 5 ↓20%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 3 3 ±0% 4 ↓25%
243_pod_names_contain_service (opus-4.6) 📄 3 3 ±0% 5 ↓40%
24_misconfigured_pvc (opus-4.6) 📄 5 4 ↑25% 6 ↓17%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 8 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 9 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 8 — — — —
261_time_window_gap_external_data (opus-4.6) 📄 11 — — — —
262_ambiguous_cluster_reference (opus-4.6) 📄 8 — — — —
263_region_suffixed_service_siblings (opus-4.6) 📄 8 — — — —
264_wrong_env_same_region (opus-4.6) 📄 9 — — — —
265_multicluster_labeling_discipline (opus-4.6) 📄 8 — — — —
266_toolset_disabled_vs_no_data (opus-4.6) 📄 1 — — — —
267_cluster_name_alias (opus-4.6) 📄 9 — — — —
268_kubectl_wrong_cluster (opus-4.6) 📄 3 — — — —
269_mixed_sources_route_to_external (opus-4.6) 📄 6 — — — —
270_namespace_collision (opus-4.6) 📄 3 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% 1 ±0%
51_logs_summarize_errors (opus-4.6) 📄 3 3 ±0% 4 ↓25%
61_exact_match_counting (opus-4.6) 📄 2 2 ±0% 3 ↓33%
Total (all, n=24) 5.2 3.4 — 4.5 —
Comparable (m=11, b=11) 3.2 3.4 ±0% 4.5 ↓29%

Tool calls comparison:

Test case This branch master (1h ago) Δ vs master benchmark (6h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 8 8 ±0% 13 ↓38%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 9 11 ↓18% 14 ↓36%
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% 4 ↓50%
12_job_crashing (opus-4.6) 📄 9 9 ±0% 15 ↓40%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 11 12 ±0% 15 ↓27%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 6 6 ±0% 9 ↓33%
243_pod_names_contain_service (opus-4.6) 📄 7 7 ±0% 11 ↓36%
24_misconfigured_pvc (opus-4.6) 📄 12 13 ±0% 15 ↓20%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 12 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 12 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 11 — — — —
261_time_window_gap_external_data (opus-4.6) 📄 13 — — — —
262_ambiguous_cluster_reference (opus-4.6) 📄 15 — — — —
263_region_suffixed_service_siblings (opus-4.6) 📄 9 — — — —
264_wrong_env_same_region (opus-4.6) 📄 11 — — — —
265_multicluster_labeling_discipline (opus-4.6) 📄 9 — — — —
266_toolset_disabled_vs_no_data (opus-4.6) 📄 — — — — —
267_cluster_name_alias (opus-4.6) 📄 18 — — — —
268_kubectl_wrong_cluster (opus-4.6) 📄 6 — — — —
269_mixed_sources_route_to_external (opus-4.6) 📄 14 — — — —
270_namespace_collision (opus-4.6) 📄 6 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 2 2 ±0% 5 ↓60%
61_exact_match_counting (opus-4.6) 📄 1 1 ±0% 3 ↓67%
Total (all, n=24) 8.5 7.1 — 10.4 —
Comparable (m=10, b=10) 6.7 7.1 ±0% 10.4 ↓36%

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 3 Failures Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/multi-cluster-evals-red -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/multi-cluster-evals-red -f markers=regression -f filter=

@github-actions

github-actions Bot commented Jun 4, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 525f75d18 (built in 1m 4s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:525f75d18
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:525f75d18 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:525f75d18
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:525f75d18
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:525f75d18
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:525f75d18 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:525f75d18
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:525f75d18

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:525f75d18 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:525f75d18

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:525f75d18 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:525f75d18

@coderabbitai

coderabbitai Bot commented Jun 4, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: fe7e8d83-c379-4ece-abc0-54626bdb470b

📥 Commits

Reviewing files that changed from the base of the PR and between 92aa5c6 and c6c4915.

📒 Files selected for processing (1)
  • tests/llm/test_ask_holmes.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • tests/llm/test_ask_holmes.py

Walkthrough

Adds extensive multi-cluster LLM test fixtures and toolset configs, tightens existing fixture expectations for transparent labeling, forwards cluster context into test message construction, and simplifies Braintrust/langfuse logging helpers.

Changes

Multi-Cluster Holmes Evaluation Coverage

Layer / File(s) Summary
Test infra, markers, and logging simplification
pyproject.toml, tests/llm/utils/braintrust.py, tests/llm/test_holmes_checks.py, tests/llm/test_ask_holmes.py
Adds multi-cluster pytest marker; forwards test_case.cluster_name into CLI message construction; removes Braintrust/langfuse dataset helpers and omits dataset_record_id from eval span logs.
Existing fixture transparency tightening
tests/llm/fixtures/test_ask_holmes/19_detect_missing_app_details/test_case.yaml, tests/llm/fixtures/test_ask_holmes/254_elasticsearch_dr_test_log_check/test_case.yaml
Tightens expected outputs to require explicit exact-name absence statements and labeled alternative/findings disclosure.
Wrong cluster logs confusion (259)
tests/llm/fixtures/test_ask_holmes/259_wrong_cluster_logs_confusion/*
Adds ES-only fixture where Holmes is on eu-west-2 but the prompt targets us-east-2; prepares 9 eu-west-2 log records and asserts explicit cluster-mismatch reporting and labeled context usage.
Global observability topology (260)
tests/llm/fixtures/test_ask_holmes/260_global_es_remote_cluster_logs/*
Adds fixture with ES containing both us-east-2 outage logs and eu-west-2 healthy logs (8/2 split); enforces routing to and labeling of us-east-2 findings under ES-only toolsets.
Time-window gap and ambiguous cluster fixtures (261/262/264)
tests/llm/fixtures/test_ask_holmes/261_time_window_gap_external_data/*, .../262_ambiguous_cluster_reference/*, .../264_wrong_env_same_region/*
Adds fixtures for time-window gaps, ambiguous production cluster references, and wrong-environment same-region cases; each provisions ES indices and asserts transparency/no-fabrication behavior.
Service/cluster naming and attribution (263/265/267)
tests/llm/fixtures/test_ask_holmes/263_region_suffixed_service_siblings/*, .../265_multicluster_labeling_discipline/*, .../267_cluster_name_alias/*
Adds fixtures ensuring region-suffixed service variants are surfaced, fleet-wide incidents are cluster-attributed, and casual cluster-name aliases map to configured clusters; includes ES setup and toolsets.
Toolset disabled vs no-data (266)
tests/llm/fixtures/test_ask_holmes/266_toolset_disabled_vs_no_data/*
Adds fixture that validates Holmes reports disabled/unconfigured toolsets (Datadog) rather than claiming it searched and found no traces.
Kubectl-only scoping (268)
tests/llm/fixtures/test_ask_holmes/268_kubectl_wrong_cluster/*
Adds kubectl-only fixture with local Deployment manifest and harness asserting kubectl access is scoped to the local cluster and must not be used to answer about other clusters.
Mixed sources routing to external (269)
tests/llm/fixtures/test_ask_holmes/269_mixed_sources_route_to_external/*
Adds mixed-topology fixture combining local Kubernetes state and ES cross-cluster logs; enforces routing to ES for non-local incidents and accurate labeling.
Namespace collision detection (270)
tests/llm/fixtures/test_ask_holmes/270_namespace_collision/*
Adds dual-namespace manifests and a harness that validates same-name resource ambiguity and differing prod/staging pod states.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

Suggested labels

evals-id-254

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title precisely summarizes the main change: adding multi-cluster awareness evaluations to the test suite and pyproject configuration.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Jun 4, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 08da31e
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a25436e7ff2140008145685
😎 Deploy Preview https://deploy-preview-2125--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 6

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@tests/llm/fixtures/test_ask_holmes/259_wrong_cluster_logs_confusion/test_case.yaml`:
- Line 161: Update the misleading echo message so it matches the test data:
change the string echo "Test data ready: $DOC_COUNT log records (no
cluster/region field — plain app logs)" to accurately state that documents
include cluster and region (e.g. echo "Test data ready: $DOC_COUNT log records
(includes cluster and region fields)") or remove the parenthetical; ensure the
echo surrounding variable DOC_COUNT is preserved and only the descriptive text
is corrected to reflect the mapping and bulk-loaded documents that include
cluster and region values like "production-eu-west-2" and "eu-west-2".

In
`@tests/llm/fixtures/test_ask_holmes/261_time_window_gap_external_data/test_case.yaml`:
- Around line 63-67: The current curl call posts "$BULK_FILE" to
"${ELASTICSEARCH_URL}/${LOGS_INDEX}/_bulk" and discards the response, which can
hide per-item failures; change the script to capture the curl output into a
variable instead of redirecting to /dev/null, parse the returned JSON for
"errors": true or any items with error fields, and if any failures are found log
the full response (or failing items) and exit non‑zero; keep references to
BULK_FILE, ELASTICSEARCH_URL, LOGS_INDEX and the _bulk endpoint so you update
the same invocation and add the simple JSON check and failure handling after the
curl call.

In
`@tests/llm/fixtures/test_ask_holmes/265_multicluster_labeling_discipline/test_case.yaml`:
- Around line 19-23: The expected_output entry containing "payment-service
stripe-us timeouts → checkout-service circuit-breaker open / 500s" includes an
unverifiable "/ 500s" suffix; either remove or rephrase that suffix from the
expected_output list item so it matches only facts present in the fixture logs,
or alternatively add an explicit "500" / "HTTP 500" line to the test input logs
so the "/ 500s" detail becomes discoverable; update the specific expected_output
list element (the one referencing "payment-service stripe-us ...
checkout-service circuit-breaker open / 500s") accordingly.

In
`@tests/llm/fixtures/test_ask_holmes/269_mixed_sources_route_to_external/test_case.yaml`:
- Around line 39-44: The readiness loop for the checkout-service pod (for i in
$(seq 1 60) ... kubectl wait --for=condition=ready pod -l app=checkout-service
-n app-269 --timeout=5s) currently continues even if it never becomes ready;
modify the script to detect a timeout after the loop and fail-fast: after the
for-loop, check whether the pod became ready (e.g., re-run kubectl get/wait or
inspect the loop exit status) and if not, emit an error message referencing
app=checkout-service and namespace app-269 and exit with a non-zero status so
the fixture stops instead of proceeding with an invalid baseline.
- Around line 77-81: The bulk ingest currently discards the Elasticsearch _bulk
response (curl output redirected to /dev/null), which can hide partial failures;
change the invocation that posts "$BULK_FILE" to capture and parse the JSON
response (from the ${ELASTICSEARCH_URL}/${LOGS_INDEX}/_bulk call), inspect the
top-level "errors" flag and each items[*].error (or equivalent) and fail the
fixture/test (nonzero exit) if any errors are present; reference the existing
BULK_FILE, the curl POST to ${ELASTICSEARCH_URL}/${LOGS_INDEX}/_bulk, and the
_bulk response structure so the fixture explicitly validates the bulk result and
prevents silent fixture drift.

In `@tests/llm/fixtures/test_ask_holmes/270_namespace_collision/test_case.yaml`:
- Around line 33-49: The two retry loops for prod (checking STATUS via kubectl
get with label app=checkout-service -n app-270-prod and testing for
ImagePullBackOff/ErrImagePull stored in STATUS) and staging (using kubectl wait
--for=condition=ready pod -l app=checkout-service -n app-270-staging) must fail
the setup if they exhaust retries; after each loop add a check whether the loop
exited via its success condition and if not echo a clear error (including which
environment and label) and exit 1 so the test run stops instead of continuing
with unknown fixture state.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 8b6b0b60-236d-4139-abc1-94ffa074757a

📥 Commits

Reviewing files that changed from the base of the PR and between 4d0a70f and 9d2a287.

📒 Files selected for processing (33)
  • pyproject.toml
  • tests/llm/fixtures/test_ask_holmes/19_detect_missing_app_details/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/254_elasticsearch_dr_test_log_check/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/259_wrong_cluster_logs_confusion/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/259_wrong_cluster_logs_confusion/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/260_global_es_remote_cluster_logs/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/260_global_es_remote_cluster_logs/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/261_time_window_gap_external_data/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/261_time_window_gap_external_data/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/262_ambiguous_cluster_reference/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/262_ambiguous_cluster_reference/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/263_region_suffixed_service_siblings/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/263_region_suffixed_service_siblings/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/264_wrong_env_same_region/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/264_wrong_env_same_region/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/265_multicluster_labeling_discipline/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/265_multicluster_labeling_discipline/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/266_toolset_disabled_vs_no_data/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/266_toolset_disabled_vs_no_data/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/267_cluster_name_alias/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/267_cluster_name_alias/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/268_kubectl_wrong_cluster/manifest.yaml
  • tests/llm/fixtures/test_ask_holmes/268_kubectl_wrong_cluster/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/269_mixed_sources_route_to_external/manifest.yaml
  • tests/llm/fixtures/test_ask_holmes/269_mixed_sources_route_to_external/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/269_mixed_sources_route_to_external/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/270_namespace_collision/prod.yaml
  • tests/llm/fixtures/test_ask_holmes/270_namespace_collision/staging.yaml
  • tests/llm/fixtures/test_ask_holmes/270_namespace_collision/test_case.yaml
  • tests/llm/test_ask_holmes.py
  • tests/llm/test_holmes_checks.py
  • tests/llm/utils/braintrust.py
  • tests/llm/utils/langfuse.py
💤 Files with no reviewable changes (2)
  • tests/llm/utils/langfuse.py
  • tests/llm/test_holmes_checks.py

claude and others added 2 commits June 4, 2026 18:01
Address CodeRabbit nits on the new multi-cluster eval fixtures:

- 259: fix echo describing data as having no cluster/region fields when
  the documents do include them.
- 261, 269: validate the Elasticsearch _bulk response and exit non-zero
  on partial failures, matching the pattern already used in 254.
- 269, 270: fail-fast when pod-readiness retry loops exhaust without
  reaching the expected state, per the CLAUDE.md race-condition pattern.
- 265: drop the unverifiable "/ 500s" suffix from expected_output - the
  fixture logs do not contain any 500 status code, so the eval
  criterion should not reference one.

No test-semantics change; this is fixture robustness only. The combined
PR #2042 (evals + prompt fix) already runs 24/24 green with these
fixtures; these tweaks reduce the chance of misleading failures from
silent setup drift.

https://claude.ai/code/session_015scgyQbqJD23heqct31Ger
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn enabled auto-merge (squash) June 7, 2026 08:54
Sheeproid
Sheeproid previously approved these changes Jun 7, 2026
…r-evals-red

# Conflicts:
#	tests/llm/test_ask_holmes.py
@aantn
aantn merged commit 478a181 into master Jun 7, 2026
19 of 22 checks passed
@aantn
aantn deleted the claude/multi-cluster-evals-red branch June 7, 2026 10:22
aantn added a commit that referenced this pull request Jun 7, 2026
)

Splits PR #2042 into two — this is the **green** half. Just the
system-prompt change. Pair with #2125 (red) which contains the evals.

## What this PR changes

`holmes/plugins/prompts/generic_ask.jinja2`: expands the existing `* You
are running on cluster X` bullet into a multi-cluster-aware procedure.

**Old:**
```jinja2
{% if cluster_name -%}
* You are running on cluster {{ cluster_name }}.
{%- endif %}
```

**New:** distinguishes kubectl-bound data (only the local cluster) from
external observability toolsets (Elasticsearch / Datadog / Loki / etc.,
which may contain data from many clusters). When the user names a
cluster / region / env other than the local one, Holmes:

1. Investigates using external toolsets — does NOT refuse.
2. Verifies each finding's cluster against the user's named cluster (via
the data's own `cluster` / `region` / `environment` /
`kubernetes.cluster.name` field, or the index/source name).
3. If matched → investigates normally with that cluster's findings.
4. If not matched → states plainly that no data exists for the requested
cluster, labels what was found in adjacent clusters, suggests pointing
Holmes at the right data source / agent.

Two common topologies are explicitly called out:
- one Holmes per cluster (kubectl + external mostly scoped to that
cluster)
- one Holmes with a global observability backend covering many clusters

Holmes doesn't know upfront which one applies — it must look at what the
data actually contains.

## Companion PR

This turns #2125's evals green:

- 254, 19 — name disambiguation
- 259 — wrong cluster, local-only data
- 260 — global topology, remote cluster data
- 261 — time-window gap
- 262 — ambiguous cluster reference
- 263 — region-suffixed sibling services
- 264 — wrong env, same region
- 265 — labeling discipline
- 266 — toolset-disabled vs no-data
- 267 — cluster-name alias (must NOT over-correct)
- 268/269/270 — kubectl-only / mixed-source / namespace-collision

Local validation across 3 iterations × 3 models (Sonnet 4.5 / Opus 4.6 /
Opus 4.7) on the ES-based subset was clean. CI on #2125 will demonstrate
the failing state; merging this on top turns those failures green.

---
_Generated by [Claude
Code](https://claude.ai/code/session_015scgyQbqJD23heqct31Ger)_

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Improvements**
* Enhanced cluster awareness to ensure observability data is correctly
attributed to your specified cluster and prevent misinterpretation of
cross-cluster data.
* Improved transparency when exact entity data is unavailable, with
clearer reporting of similarly-named alternatives for user verification.
* Refined troubleshooting guidance to better support iterative
investigation workflows.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
aantn pushed a commit that referenced this pull request Jun 7, 2026
PR #2125 (merged via master) added 262_ambiguous_cluster_reference and a
multi-cluster eval series through 270. Move this eval to 271 (namespace
app-271, fixture label 271-root-cause-noise) so the two 262_* directories
no longer collide.

Signed-off-by: Claude <noreply@anthropic.com>
Avi-Robusta added a commit that referenced this pull request Jun 9, 2026
The multi-cluster ES evals (254/259/260, added in #2125) require
ELASTICSEARCH_URL/ELASTICSEARCH_API_KEY. Under --strict-setup-mode their
missing-env setup failures abort the whole benchmark run. Wire in the
existing repo secrets, mirroring eval-regression.yaml.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: avi@robusta.dev <avi@robusta.dev>
Avi-Robusta added a commit that referenced this pull request Jun 9, 2026
Evals 254/259/260 (added in #2125, titled 'multi-cluster awareness evals
(red)') are not expected to pass yet, but were tagged 'regression'. The
regression marker is reserved for tests that must always pass (30+ iters
reliably on Sonnet-4.5), so a red eval there pollutes the must-pass suite
and pulls these into fast-benchmark prematurely. Drop the regression tag;
they keep medium/multi-cluster/elasticsearch/transparency so they can still
be run/scored explicitly, just not as part of the regression gate.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: avi@robusta.dev <avi@robusta.dev>
aantn pushed a commit that referenced this pull request Jun 12, 2026
… eval infra bugs

Run #8 left the arms tied 49/53 with two SDK-only failures and two both-arm
failures. All four root causes identified from Braintrust traces and fixed:

- 236_image_spill_to_disk (SDK-only): the claude CLI caps MCP tool output at
  25k tokens; the eval's report intentionally pads ~100k chars before the
  image attachment, so the cap silently dropped the image. Raise
  MAX_MCP_OUTPUT_TOKENS to 100000 in the engine's CLI env. Verified locally:
  PASS.
- 259_wrong_cluster_logs_confusion (SDK-only, flaky): the agent disclosed the
  cluster mismatch as a property of the data ("index only contains eu-west-2")
  but the eval requires agent-identity framing ("this agent is connected to
  eu-west-2, not us-east-2"). One-line addition to the cluster sentence in the
  system prompt. Verified locally: PASS.
- 19_detect_missing_app_details (both arms failed): the eval (added red in
  #2125) asks Holmes to surface db-certs-authenticator as the close match for
  personal-certs-validator, but no such resource was ever deployed — nothing
  could pass it. Add before_test/after_test creating it as a crashlooping
  deployment in app-19 with a discoverable cause (missing CA bundle) in its
  logs. Verified on a live cluster (SDK engine): PASS.
- 227_count_configmaps_per_namespace (both arms failed): evals 229 and 249
  deployed into namespace app-227 (copy-paste, violating the app-<testid>
  convention), so "namespaces starting with app-227" truthfully totalled 645
  ConfigMaps against the eval's hardcoded 644 — both engines counted the live
  cluster correctly and were marked wrong. Move 229 to app-229 and 249 to
  app-249 (249's own prompt already said app-249). Verified on a live
  cluster: ground truth is 644 again.

Signed-off-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants