Repository navigation
multi-cluster awareness evals (red) — tests/llm + pyproject - #2125
Conversation
This PR introduces 13 new/updated evals that exercise Holmes's
multi-cluster / multi-environment / multi-source awareness, plus
supporting test framework cleanup. Without the corresponding prompt
fix (separate PR) these evals are expected to be RED on master —
they demonstrate the behavior we want Holmes to have, before the
prompt change lands.
Evals (all tagged `multi-cluster`):
254_elasticsearch_dr_test_log_check (updated)
Service-name disambiguation. companyopswebjob is asked about;
only Company.Ops and Company.Ops.Radar.WebJob exist. Must say
no exact match, may report adjacent findings if clearly labeled.
19_detect_missing_app_details (updated)
personal-certs-validator vs db-certs-authenticator. Both
transparency (no exact match) AND usefulness (label the
similar-named resource) required.
259_wrong_cluster_logs_confusion (new)
Holmes on production-eu-west-2, user asks about
production-us-east-2. ES index has only eu-west-2 data with
red-herring 500s. Must flag cluster mismatch up-front.
260_global_es_remote_cluster_logs (new)
Counterpart to 259 — same scenario but ES has us-east-2 data
(global topology). Holmes must investigate normally, not
refuse / over-correct.
261_time_window_gap_external_data (new)
Data exists but only for the last 30 min; user asks about an
incident 6h ago. Must flag the time-window gap.
262_ambiguous_cluster_reference (new)
cluster_name="production"; user says "the US prod cluster".
Must not silently treat local as the meant cluster.
263_region_suffixed_service_siblings (new)
payment-service-us / -eu / -ap. User asks about
"payment-service" — must surface the ambiguity, not silently
pick one.
264_wrong_env_same_region (new)
Holmes on staging-eu-west-2; user asks about
production-eu-west-2. Must catch env mismatch.
265_multicluster_labeling_discipline (new)
Open-ended "what's broken across our fleet?" — every finding
in the answer must be attributed to a specific cluster.
266_toolset_disabled_vs_no_data (new)
User asks Holmes to check Datadog APM; Datadog toolset is
disabled. Must say "can't access" not "looked and found
nothing."
267_cluster_name_alias (new)
cluster_name="acme-prod-eu-west-1"; user says "EU prod". Must
recognise the alias as the same cluster and investigate — must
NOT refuse / over-correct from the cluster-awareness rule.
268_kubectl_wrong_cluster (new, KIND)
kubectl-only path; user asks about non-local cluster. Must
flag kubectl-local limitation.
269_mixed_sources_route_to_external (new, KIND + ES)
Holmes has kubectl (local only) + multi-cluster ES. User asks
about non-local cluster. Must route to ES, not blame healthy
local kubectl state.
270_namespace_collision (new, KIND)
Same-named deployment in two namespaces with different states.
Must surface ambiguity and address both, not silently pick one.
Supporting changes:
tests/llm/test_ask_holmes.py: pass test_case.cluster_name through
to build_initial_ask_messages in the CLI test path. Previously it
was silently dropped, so cluster_name in test fixtures had no
effect on the system prompt for ask-mode evals.
tests/llm/utils/braintrust.py + tests/llm/utils/langfuse.py: purge
legacy Braintrust dataset upload code (BraintrustEvalHelper /
upload_test_cases / find_dataset_row_by_test_case / etc.) and the
dead langfuse upload helper. We don't use Braintrust datasets in
this repo — all eval logging goes via experiment span logs. The
legacy code also had a real bug (input=input where `input` was
the Python builtin) that corrupted dataset records.
tests/llm/test_holmes_checks.py: drop dataset_record_id= from
eval_span.log calls (we no longer link to a dataset).
pyproject.toml: register the new `multi-cluster` pytest marker.
Counterpart PR (green): the prompt-side change in
holmes/plugins/prompts/generic_ask.jinja2 that teaches Holmes about
multi-cluster scope awareness. Once that PR is merged, these evals
turn green.
Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.
Tip: disable this comment in your organization's Code Review settings.
📂 Previous Runs
|
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions | Src |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 26.1s | 4 | 8 | $0.2032 | 67,826 | 66,217 | 18,927 | 1,609 | 586 | 46,714 | 19,503 | 158 | — | src |
| ✅ | 101_loki_historical_logs_pod_deleted | 38.8s | 4 | 9 | $0.2251 | 68,396 | 66,078 | 19,545 | 2,318 | 745 | 45,981 | 20,097 | 212 | — | src |
| ✅ | 112_find_pvcs_by_uuid | 14.7s | 2 | 2 | $0.1551 | 32,357 | 31,355 | 17,491 | 1,002 | 554 | 13,861 | 17,494 | 294 | — | src |
| ✅ | 12_job_crashing | 29.6s | 4 | 9 | $0.2196 | 71,808 | 70,065 | 20,437 | 1,743 | 526 | 48,729 | 21,336 | 181 | — | src |
| ✅ | 176_network_policy_blocking_traffic_no_skills | 27.5s | 4 | 11 | $0.2121 | 69,765 | 67,987 | 19,929 | 1,778 | 590 | 48,053 | 19,934 | 219 | — | src |
| ✅ | 227_count_configmaps_per_namespace[0] | 15.2s | 3 | 6 | $0.1451 | 44,959 | 44,261 | 16,195 | 698 | 432 | 28,062 | 16,199 | 29 | — | src |
| ✅ | 243_pod_names_contain_service | 27.0s | 3 | 7 | $0.1835 | 48,364 | 46,754 | 17,639 | 1,610 | 573 | 28,648 | 18,106 | 248 | — | src |
| ✅ | 24_misconfigured_pvc | 31.3s | 5 | 12 | $0.2270 | 86,558 | 84,593 | 19,217 | 1,965 | 572 | 64,398 | 20,195 | 125 | — | src |
| ✅ | 254_elasticsearch_dr_test_log_check | 53.0s | 8 | 12 | $0.2674 | 104,450 | 100,704 | 17,233 | 3,746 | 801 | 83,462 | 17,242 | 204 | — | src |
| ❌ | 259_wrong_cluster_logs_confusion | 56.7s | 9 | 12 | $0.2736 | 121,067 | 117,517 | 17,383 | 3,550 | 650 | 99,846 | 17,671 | 153 | — | src |
| ✅ | 260_global_es_remote_cluster_logs | 51.0s | 8 | 11 | $0.2544 | 107,504 | 104,375 | 17,050 | 3,129 | 729 | 86,696 | 17,679 | 258 | — | src |
| ✅ | 261_time_window_gap_external_data | 65.5s | 11 | 13 | $0.2847 | 139,707 | 135,945 | 16,841 | 3,762 | 850 | 118,822 | 17,123 | 532 | — | src |
| ✅ | 262_ambiguous_cluster_reference | 54.8s | 8 | 15 | $0.2817 | 109,723 | 105,985 | 18,098 | 3,738 | 746 | 86,639 | 19,346 | 172 | — | src |
| ✅ | 263_region_suffixed_service_siblings | 48.4s | 8 | 9 | $0.2186 | 94,762 | 92,142 | 14,775 | 2,620 | 730 | 76,887 | 15,255 | 80 | — | src |
| ✅ | 264_wrong_env_same_region | 53.6s | 9 | 11 | $0.2479 | 113,854 | 110,673 | 15,482 | 3,181 | 822 | 94,868 | 15,805 | 266 | — | src |
| ✅ | 265_multicluster_labeling_discipline | 47.8s | 8 | 9 | $0.2314 | 97,723 | 94,769 | 15,497 | 2,954 | 841 | 79,263 | 15,506 | 308 | — | src |
| ✅ | 266_toolset_disabled_vs_no_data | 11.3s | 1 | — | $0.0733 | 9,694 | 9,369 | 9,369 | 325 | 325 | 0 | 9,369 | 91 | — | src |
| ✅ | 267_cluster_name_alias | 63.5s | 9 | 18 | $0.3172 | 129,548 | 125,059 | 19,239 | 4,489 | 920 | 105,116 | 19,943 | 266 | — | src |
| ❌ | 268_kubectl_wrong_cluster | 27.3s | 3 | 6 | $0.2013 | 52,335 | 50,599 | 19,619 | 1,736 | 597 | 30,443 | 20,156 | 466 | — | src |
| ✅ | 269_mixed_sources_route_to_external | 60.6s | 6 | 14 | $0.2749 | 90,618 | 86,701 | 18,748 | 3,917 | 1,252 | 67,728 | 18,973 | 908 | — | src |
| ❌ | 270_namespace_collision | 20.5s | 3 | 6 | $0.1701 | 47,620 | 46,430 | 17,589 | 1,190 | 505 | 28,643 | 17,787 | 132 | — | src |
| ✅ | 43_current_datetime_from_prompt | 3.2s | 1 | — | $0.0986 | 13,970 | 13,848 | 13,848 | 122 | 122 | 0 | 13,848 | 78 | — | src |
| ✅ | 51_logs_summarize_errors | 19.1s | 3 | 2 | $0.1505 | 45,401 | 44,650 | 16,639 | 751 | 390 | 28,007 | 16,643 | 29 | — | src |
| ✅ | 61_exact_match_counting | 8.7s | 2 | 1 | $0.1112 | 28,231 | 28,010 | 14,182 | 221 | 152 | 13,825 | 14,185 | 34 | — | src |
| Total | 35.6s avg | 5.2 avg | 9.2 avg | $5.0274 | 1,796,240 | 1,744,086 | 20,437 | 52,154 | 1,252 | 1,324,691 | 419,395 | 5,443 | — |
Benchmark Comparison Details
Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded
- master-27089078550 (created: 2026-06-07)
Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded
- ci-benchmark-27081280174 (created: 2026-06-07)
Time comparison (seconds):
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 26.1s | 26.2s | ±0% | 41.4s | ↓37% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 38.8s | 47.6s | ↓19% | 65.7s | ↓41% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 14.7s | 10.6s | ↑39% | 19.0s | ↓23% |
| 12_job_crashing (opus-4.6) 📄 | 29.6s | 27.6s | ±0% | 42.6s | ↓30% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 27.5s | 32.6s | ↓16% | 44.3s | ↓38% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 15.2s | 13.2s | ↑15% | 18.4s | ↓18% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 27.0s | 28.8s | ±0% | 35.1s | ↓23% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 31.3s | 31.4s | ±0% | 40.6s | ↓23% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 53.0s | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 56.7s | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 51.0s | — | — | — | — |
| 261_time_window_gap_external_data (opus-4.6) 📄 | 65.5s | — | — | — | — |
| 262_ambiguous_cluster_reference (opus-4.6) 📄 | 54.8s | — | — | — | — |
| 263_region_suffixed_service_siblings (opus-4.6) 📄 | 48.4s | — | — | — | — |
| 264_wrong_env_same_region (opus-4.6) 📄 | 53.6s | — | — | — | — |
| 265_multicluster_labeling_discipline (opus-4.6) 📄 | 47.8s | — | — | — | — |
| 266_toolset_disabled_vs_no_data (opus-4.6) 📄 | 11.3s | — | — | — | — |
| 267_cluster_name_alias (opus-4.6) 📄 | 63.5s | — | — | — | — |
| 268_kubectl_wrong_cluster (opus-4.6) 📄 | 27.3s | — | — | — | — |
| 269_mixed_sources_route_to_external (opus-4.6) 📄 | 60.6s | — | — | — | — |
| 270_namespace_collision (opus-4.6) 📄 | 20.5s | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 3.2s | 3.1s | ±0% | 3.5s | ±0% |
| 51_logs_summarize_errors (opus-4.6) 📄 | 19.1s | 18.8s | ±0% | 19.7s | ±0% |
| 61_exact_match_counting (opus-4.6) 📄 | 8.7s | 7.1s | ↑22% | 10.2s | ↓15% |
| Total (all, n=24) | 35.6s | 22.5s | — | 31.0s | — |
| Comparable (m=11, b=11) | 21.9s | 22.5s | ±0% | 31.0s | ↓29% |
Cost comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | $0.2032 | $0.2032 | ±0% | $0.3101 | ↓34% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | $0.2251 | $0.2642 | ↓15% | $0.3829 | ↓41% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | $0.1551 | $0.1368 | ↑13% | $0.2047 | ↓24% |
| 12_job_crashing (opus-4.6) 📄 | $0.2196 | $0.2173 | ±0% | $0.3167 | ↓31% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | $0.2121 | $0.2524 | ↓16% | $0.3179 | ↓33% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | $0.1451 | $0.1464 | ±0% | $0.2048 | ↓29% |
| 243_pod_names_contain_service (opus-4.6) 📄 | $0.1835 | $0.1872 | ±0% | $0.2662 | ↓31% |
| 24_misconfigured_pvc (opus-4.6) 📄 | $0.2270 | $0.2240 | ±0% | $0.3110 | ↓27% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | $0.2674 | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | $0.2736 | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | $0.2544 | — | — | — | — |
| 261_time_window_gap_external_data (opus-4.6) 📄 | $0.2847 | — | — | — | — |
| 262_ambiguous_cluster_reference (opus-4.6) 📄 | $0.2817 | — | — | — | — |
| 263_region_suffixed_service_siblings (opus-4.6) 📄 | $0.2186 | — | — | — | — |
| 264_wrong_env_same_region (opus-4.6) 📄 | $0.2479 | — | — | — | — |
| 265_multicluster_labeling_discipline (opus-4.6) 📄 | $0.2314 | — | — | — | — |
| 266_toolset_disabled_vs_no_data (opus-4.6) 📄 | $0.0733 | — | — | — | — |
| 267_cluster_name_alias (opus-4.6) 📄 | $0.3172 | — | — | — | — |
| 268_kubectl_wrong_cluster (opus-4.6) 📄 | $0.2013 | — | — | — | — |
| 269_mixed_sources_route_to_external (opus-4.6) 📄 | $0.2749 | — | — | — | — |
| 270_namespace_collision (opus-4.6) 📄 | $0.1701 | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | $0.0986 | $0.0110 | ↑797% | $0.1190 | ↓17% |
| 51_logs_summarize_errors (opus-4.6) 📄 | $0.1505 | $0.1488 | ±0% | $0.2018 | ↓25% |
| 61_exact_match_counting (opus-4.6) 📄 | $0.1112 | $0.1112 | ±0% | $0.1518 | ↓27% |
| Total (all, n=24) | $0.2095 | $0.1730 | — | $0.2534 | — |
| Comparable (m=11, b=11) | $0.1755 | $0.1730 | ±0% | $0.2534 | ↓31% |
Total tokens comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 67,826 | 67,778 | ±0% | 132,351 | ↓49% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 68,396 | 88,307 | ↓23% | 144,907 | ↓53% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 32,357 | 30,686 | ±0% | 61,219 | ↓47% |
| 12_job_crashing (opus-4.6) 📄 | 71,808 | 69,487 | ±0% | 136,570 | ↓47% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 69,765 | 108,726 | ↓36% | 113,514 | ↓39% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 44,959 | 44,988 | ±0% | 76,715 | ↓41% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 48,364 | 48,612 | ±0% | 104,698 | ↓54% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 86,558 | 68,923 | ↑26% | 132,672 | ↓35% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 104,450 | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 121,067 | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 107,504 | — | — | — | — |
| 261_time_window_gap_external_data (opus-4.6) 📄 | 139,707 | — | — | — | — |
| 262_ambiguous_cluster_reference (opus-4.6) 📄 | 109,723 | — | — | — | — |
| 263_region_suffixed_service_siblings (opus-4.6) 📄 | 94,762 | — | — | — | — |
| 264_wrong_env_same_region (opus-4.6) 📄 | 113,854 | — | — | — | — |
| 265_multicluster_labeling_discipline (opus-4.6) 📄 | 97,723 | — | — | — | — |
| 266_toolset_disabled_vs_no_data (opus-4.6) 📄 | 9,694 | — | — | — | — |
| 267_cluster_name_alias (opus-4.6) 📄 | 129,548 | — | — | — | — |
| 268_kubectl_wrong_cluster (opus-4.6) 📄 | 52,335 | — | — | — | — |
| 269_mixed_sources_route_to_external (opus-4.6) 📄 | 90,618 | — | — | — | — |
| 270_namespace_collision (opus-4.6) 📄 | 47,620 | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 13,970 | 13,970 | ±0% | 17,001 | ↓18% |
| 51_logs_summarize_errors (opus-4.6) 📄 | 45,401 | 45,163 | ±0% | 77,120 | ↓41% |
| 61_exact_match_counting (opus-4.6) 📄 | 28,231 | 28,231 | ±0% | 52,716 | ↓46% |
| Total (all, n=24) | 74,843 | 55,897 | — | 95,408 | — |
| Comparable (m=11, b=11) | 52,512 | 55,897 | ±0% | 95,408 | ↓45% |
Cached tokens comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 46,714 | 46,700 | ±0% | 102,543 | ↓54% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 45,981 | 64,209 | ↓28% | 109,927 | ↓58% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 13,861 | 13,861 | ±0% | 38,090 | ↓64% |
| 12_job_crashing (opus-4.6) 📄 | 48,729 | 46,249 | ±0% | 107,265 | ↓55% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 48,053 | 84,274 | ↓43% | 82,679 | ↓42% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 28,062 | 28,072 | ±0% | 54,602 | ↓49% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 28,648 | 28,690 | ±0% | 78,754 | ↓64% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 64,398 | 46,110 | ↑40% | 103,729 | ↓38% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 83,462 | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 99,846 | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 86,696 | — | — | — | — |
| 261_time_window_gap_external_data (opus-4.6) 📄 | 118,822 | — | — | — | — |
| 262_ambiguous_cluster_reference (opus-4.6) 📄 | 86,639 | — | — | — | — |
| 263_region_suffixed_service_siblings (opus-4.6) 📄 | 76,887 | — | — | — | — |
| 264_wrong_env_same_region (opus-4.6) 📄 | 94,868 | — | — | — | — |
| 265_multicluster_labeling_discipline (opus-4.6) 📄 | 79,263 | — | — | — | — |
| 266_toolset_disabled_vs_no_data (opus-4.6) 📄 | — | — | — | — | — |
| 267_cluster_name_alias (opus-4.6) 📄 | 105,116 | — | — | — | — |
| 268_kubectl_wrong_cluster (opus-4.6) 📄 | 30,443 | — | — | — | — |
| 269_mixed_sources_route_to_external (opus-4.6) 📄 | 67,728 | — | — | — | — |
| 270_namespace_collision (opus-4.6) 📄 | 28,643 | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | — | 13,845 | — | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 28,007 | 28,012 | ±0% | 55,156 | ↓49% |
| 61_exact_match_counting (opus-4.6) 📄 | 13,825 | 13,825 | ±0% | 34,481 | ↓60% |
| Total (all, n=24) | 55,195 | 37,622 | — | 69,748 | — |
| Comparable (m=10, b=10) | 36,628 | 40,000 | ±0% | 76,723 | ↓52% |
Turns comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 4 | 4 | ±0% | 6 | ↓33% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 4 | 5 | ↓20% | 6 | ↓33% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 2 | 2 | ±0% | 3 | ↓33% |
| 12_job_crashing (opus-4.6) 📄 | 4 | 4 | ±0% | 6 | ↓33% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 4 | 6 | ↓33% | 5 | ↓20% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 3 | 3 | ±0% | 4 | ↓25% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 3 | 3 | ±0% | 5 | ↓40% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 5 | 4 | ↑25% | 6 | ↓17% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 8 | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 9 | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 8 | — | — | — | — |
| 261_time_window_gap_external_data (opus-4.6) 📄 | 11 | — | — | — | — |
| 262_ambiguous_cluster_reference (opus-4.6) 📄 | 8 | — | — | — | — |
| 263_region_suffixed_service_siblings (opus-4.6) 📄 | 8 | — | — | — | — |
| 264_wrong_env_same_region (opus-4.6) 📄 | 9 | — | — | — | — |
| 265_multicluster_labeling_discipline (opus-4.6) 📄 | 8 | — | — | — | — |
| 266_toolset_disabled_vs_no_data (opus-4.6) 📄 | 1 | — | — | — | — |
| 267_cluster_name_alias (opus-4.6) 📄 | 9 | — | — | — | — |
| 268_kubectl_wrong_cluster (opus-4.6) 📄 | 3 | — | — | — | — |
| 269_mixed_sources_route_to_external (opus-4.6) 📄 | 6 | — | — | — | — |
| 270_namespace_collision (opus-4.6) 📄 | 3 | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 1 | 1 | ±0% | 1 | ±0% |
| 51_logs_summarize_errors (opus-4.6) 📄 | 3 | 3 | ±0% | 4 | ↓25% |
| 61_exact_match_counting (opus-4.6) 📄 | 2 | 2 | ±0% | 3 | ↓33% |
| Total (all, n=24) | 5.2 | 3.4 | — | 4.5 | — |
| Comparable (m=11, b=11) | 3.2 | 3.4 | ±0% | 4.5 | ↓29% |
Tool calls comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 8 | 8 | ±0% | 13 | ↓38% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 9 | 11 | ↓18% | 14 | ↓36% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 2 | 2 | ±0% | 4 | ↓50% |
| 12_job_crashing (opus-4.6) 📄 | 9 | 9 | ±0% | 15 | ↓40% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 11 | 12 | ±0% | 15 | ↓27% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 6 | 6 | ±0% | 9 | ↓33% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 7 | 7 | ±0% | 11 | ↓36% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 12 | 13 | ±0% | 15 | ↓20% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 12 | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 12 | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 11 | — | — | — | — |
| 261_time_window_gap_external_data (opus-4.6) 📄 | 13 | — | — | — | — |
| 262_ambiguous_cluster_reference (opus-4.6) 📄 | 15 | — | — | — | — |
| 263_region_suffixed_service_siblings (opus-4.6) 📄 | 9 | — | — | — | — |
| 264_wrong_env_same_region (opus-4.6) 📄 | 11 | — | — | — | — |
| 265_multicluster_labeling_discipline (opus-4.6) 📄 | 9 | — | — | — | — |
| 266_toolset_disabled_vs_no_data (opus-4.6) 📄 | — | — | — | — | — |
| 267_cluster_name_alias (opus-4.6) 📄 | 18 | — | — | — | — |
| 268_kubectl_wrong_cluster (opus-4.6) 📄 | 6 | — | — | — | — |
| 269_mixed_sources_route_to_external (opus-4.6) 📄 | 14 | — | — | — | — |
| 270_namespace_collision (opus-4.6) 📄 | 6 | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | — | — | — | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 2 | 2 | ±0% | 5 | ↓60% |
| 61_exact_match_counting (opus-4.6) 📄 | 1 | 1 | ±0% | 3 | ↓67% |
| Total (all, n=24) | 8.5 | 7.1 | — | 10.4 | — |
| Comparable (m=10, b=10) | 6.7 | 7.1 | ±0% | 10.4 | ↓36% |
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 3 Failures Detected
📖 Legend
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/multi-cluster-evals-red -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
tags: regression
Or with more options (one per line):
/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
tags: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
tags |
Pytest tags / markers (no default - runs all tests!) |
id |
Eval ID / pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):
| Label | Effect |
|---|---|
evals-tag-<name> |
Run tests with tag <name> alongside regression |
evals-id-<name> |
Run a specific eval by test ID |
evals-model-<name> |
Override the model (use model list name, e.g. sonnet-4.5) |
Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5
🏷️ Valid tags
benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs
🤖 Valid models
deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/multi-cluster-evals-red -f markers=regression -f filter=
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:525f75d18
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:525f75d18 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:525f75d18
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:525f75d18
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:525f75d18
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:525f75d18 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:525f75d18
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:525f75d18Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:525f75d18 \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:525f75d18Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:525f75d18 \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:525f75d18 |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
WalkthroughAdds extensive multi-cluster LLM test fixtures and toolset configs, tightens existing fixture expectations for transparent labeling, forwards cluster context into test message construction, and simplifies Braintrust/langfuse logging helpers. ChangesMulti-Cluster Holmes Evaluation Coverage
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related PRs
Suggested labels
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
There was a problem hiding this comment.
Actionable comments posted: 6
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In
`@tests/llm/fixtures/test_ask_holmes/259_wrong_cluster_logs_confusion/test_case.yaml`:
- Line 161: Update the misleading echo message so it matches the test data:
change the string echo "Test data ready: $DOC_COUNT log records (no
cluster/region field — plain app logs)" to accurately state that documents
include cluster and region (e.g. echo "Test data ready: $DOC_COUNT log records
(includes cluster and region fields)") or remove the parenthetical; ensure the
echo surrounding variable DOC_COUNT is preserved and only the descriptive text
is corrected to reflect the mapping and bulk-loaded documents that include
cluster and region values like "production-eu-west-2" and "eu-west-2".
In
`@tests/llm/fixtures/test_ask_holmes/261_time_window_gap_external_data/test_case.yaml`:
- Around line 63-67: The current curl call posts "$BULK_FILE" to
"${ELASTICSEARCH_URL}/${LOGS_INDEX}/_bulk" and discards the response, which can
hide per-item failures; change the script to capture the curl output into a
variable instead of redirecting to /dev/null, parse the returned JSON for
"errors": true or any items with error fields, and if any failures are found log
the full response (or failing items) and exit non‑zero; keep references to
BULK_FILE, ELASTICSEARCH_URL, LOGS_INDEX and the _bulk endpoint so you update
the same invocation and add the simple JSON check and failure handling after the
curl call.
In
`@tests/llm/fixtures/test_ask_holmes/265_multicluster_labeling_discipline/test_case.yaml`:
- Around line 19-23: The expected_output entry containing "payment-service
stripe-us timeouts → checkout-service circuit-breaker open / 500s" includes an
unverifiable "/ 500s" suffix; either remove or rephrase that suffix from the
expected_output list item so it matches only facts present in the fixture logs,
or alternatively add an explicit "500" / "HTTP 500" line to the test input logs
so the "/ 500s" detail becomes discoverable; update the specific expected_output
list element (the one referencing "payment-service stripe-us ...
checkout-service circuit-breaker open / 500s") accordingly.
In
`@tests/llm/fixtures/test_ask_holmes/269_mixed_sources_route_to_external/test_case.yaml`:
- Around line 39-44: The readiness loop for the checkout-service pod (for i in
$(seq 1 60) ... kubectl wait --for=condition=ready pod -l app=checkout-service
-n app-269 --timeout=5s) currently continues even if it never becomes ready;
modify the script to detect a timeout after the loop and fail-fast: after the
for-loop, check whether the pod became ready (e.g., re-run kubectl get/wait or
inspect the loop exit status) and if not, emit an error message referencing
app=checkout-service and namespace app-269 and exit with a non-zero status so
the fixture stops instead of proceeding with an invalid baseline.
- Around line 77-81: The bulk ingest currently discards the Elasticsearch _bulk
response (curl output redirected to /dev/null), which can hide partial failures;
change the invocation that posts "$BULK_FILE" to capture and parse the JSON
response (from the ${ELASTICSEARCH_URL}/${LOGS_INDEX}/_bulk call), inspect the
top-level "errors" flag and each items[*].error (or equivalent) and fail the
fixture/test (nonzero exit) if any errors are present; reference the existing
BULK_FILE, the curl POST to ${ELASTICSEARCH_URL}/${LOGS_INDEX}/_bulk, and the
_bulk response structure so the fixture explicitly validates the bulk result and
prevents silent fixture drift.
In `@tests/llm/fixtures/test_ask_holmes/270_namespace_collision/test_case.yaml`:
- Around line 33-49: The two retry loops for prod (checking STATUS via kubectl
get with label app=checkout-service -n app-270-prod and testing for
ImagePullBackOff/ErrImagePull stored in STATUS) and staging (using kubectl wait
--for=condition=ready pod -l app=checkout-service -n app-270-staging) must fail
the setup if they exhaust retries; after each loop add a check whether the loop
exited via its success condition and if not echo a clear error (including which
environment and label) and exit 1 so the test run stops instead of continuing
with unknown fixture state.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 8b6b0b60-236d-4139-abc1-94ffa074757a
📒 Files selected for processing (33)
pyproject.tomltests/llm/fixtures/test_ask_holmes/19_detect_missing_app_details/test_case.yamltests/llm/fixtures/test_ask_holmes/254_elasticsearch_dr_test_log_check/test_case.yamltests/llm/fixtures/test_ask_holmes/259_wrong_cluster_logs_confusion/test_case.yamltests/llm/fixtures/test_ask_holmes/259_wrong_cluster_logs_confusion/toolsets.yamltests/llm/fixtures/test_ask_holmes/260_global_es_remote_cluster_logs/test_case.yamltests/llm/fixtures/test_ask_holmes/260_global_es_remote_cluster_logs/toolsets.yamltests/llm/fixtures/test_ask_holmes/261_time_window_gap_external_data/test_case.yamltests/llm/fixtures/test_ask_holmes/261_time_window_gap_external_data/toolsets.yamltests/llm/fixtures/test_ask_holmes/262_ambiguous_cluster_reference/test_case.yamltests/llm/fixtures/test_ask_holmes/262_ambiguous_cluster_reference/toolsets.yamltests/llm/fixtures/test_ask_holmes/263_region_suffixed_service_siblings/test_case.yamltests/llm/fixtures/test_ask_holmes/263_region_suffixed_service_siblings/toolsets.yamltests/llm/fixtures/test_ask_holmes/264_wrong_env_same_region/test_case.yamltests/llm/fixtures/test_ask_holmes/264_wrong_env_same_region/toolsets.yamltests/llm/fixtures/test_ask_holmes/265_multicluster_labeling_discipline/test_case.yamltests/llm/fixtures/test_ask_holmes/265_multicluster_labeling_discipline/toolsets.yamltests/llm/fixtures/test_ask_holmes/266_toolset_disabled_vs_no_data/test_case.yamltests/llm/fixtures/test_ask_holmes/266_toolset_disabled_vs_no_data/toolsets.yamltests/llm/fixtures/test_ask_holmes/267_cluster_name_alias/test_case.yamltests/llm/fixtures/test_ask_holmes/267_cluster_name_alias/toolsets.yamltests/llm/fixtures/test_ask_holmes/268_kubectl_wrong_cluster/manifest.yamltests/llm/fixtures/test_ask_holmes/268_kubectl_wrong_cluster/test_case.yamltests/llm/fixtures/test_ask_holmes/269_mixed_sources_route_to_external/manifest.yamltests/llm/fixtures/test_ask_holmes/269_mixed_sources_route_to_external/test_case.yamltests/llm/fixtures/test_ask_holmes/269_mixed_sources_route_to_external/toolsets.yamltests/llm/fixtures/test_ask_holmes/270_namespace_collision/prod.yamltests/llm/fixtures/test_ask_holmes/270_namespace_collision/staging.yamltests/llm/fixtures/test_ask_holmes/270_namespace_collision/test_case.yamltests/llm/test_ask_holmes.pytests/llm/test_holmes_checks.pytests/llm/utils/braintrust.pytests/llm/utils/langfuse.py
💤 Files with no reviewable changes (2)
- tests/llm/utils/langfuse.py
- tests/llm/test_holmes_checks.py
Address CodeRabbit nits on the new multi-cluster eval fixtures: - 259: fix echo describing data as having no cluster/region fields when the documents do include them. - 261, 269: validate the Elasticsearch _bulk response and exit non-zero on partial failures, matching the pattern already used in 254. - 269, 270: fail-fast when pod-readiness retry loops exhaust without reaching the expected state, per the CLAUDE.md race-condition pattern. - 265: drop the unverifiable "/ 500s" suffix from expected_output - the fixture logs do not contain any 500 status code, so the eval criterion should not reference one. No test-semantics change; this is fixture robustness only. The combined PR #2042 (evals + prompt fix) already runs 24/24 green with these fixtures; these tweaks reduce the chance of misleading failures from silent setup drift. https://claude.ai/code/session_015scgyQbqJD23heqct31Ger Signed-off-by: Claude <noreply@anthropic.com>
…r-evals-red # Conflicts: # tests/llm/test_ask_holmes.py
) Splits PR #2042 into two — this is the **green** half. Just the system-prompt change. Pair with #2125 (red) which contains the evals. ## What this PR changes `holmes/plugins/prompts/generic_ask.jinja2`: expands the existing `* You are running on cluster X` bullet into a multi-cluster-aware procedure. **Old:** ```jinja2 {% if cluster_name -%} * You are running on cluster {{ cluster_name }}. {%- endif %} ``` **New:** distinguishes kubectl-bound data (only the local cluster) from external observability toolsets (Elasticsearch / Datadog / Loki / etc., which may contain data from many clusters). When the user names a cluster / region / env other than the local one, Holmes: 1. Investigates using external toolsets — does NOT refuse. 2. Verifies each finding's cluster against the user's named cluster (via the data's own `cluster` / `region` / `environment` / `kubernetes.cluster.name` field, or the index/source name). 3. If matched → investigates normally with that cluster's findings. 4. If not matched → states plainly that no data exists for the requested cluster, labels what was found in adjacent clusters, suggests pointing Holmes at the right data source / agent. Two common topologies are explicitly called out: - one Holmes per cluster (kubectl + external mostly scoped to that cluster) - one Holmes with a global observability backend covering many clusters Holmes doesn't know upfront which one applies — it must look at what the data actually contains. ## Companion PR This turns #2125's evals green: - 254, 19 — name disambiguation - 259 — wrong cluster, local-only data - 260 — global topology, remote cluster data - 261 — time-window gap - 262 — ambiguous cluster reference - 263 — region-suffixed sibling services - 264 — wrong env, same region - 265 — labeling discipline - 266 — toolset-disabled vs no-data - 267 — cluster-name alias (must NOT over-correct) - 268/269/270 — kubectl-only / mixed-source / namespace-collision Local validation across 3 iterations × 3 models (Sonnet 4.5 / Opus 4.6 / Opus 4.7) on the ES-based subset was clean. CI on #2125 will demonstrate the failing state; merging this on top turns those failures green. --- _Generated by [Claude Code](https://claude.ai/code/session_015scgyQbqJD23heqct31Ger)_ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Improvements** * Enhanced cluster awareness to ensure observability data is correctly attributed to your specified cluster and prevent misinterpretation of cross-cluster data. * Improved transparency when exact entity data is unavailable, with clearer reporting of similarly-named alternatives for user verification. * Refined troubleshooting guidance to better support iterative investigation workflows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Claude <noreply@anthropic.com> Co-authored-by: Claude <noreply@anthropic.com>
PR #2125 (merged via master) added 262_ambiguous_cluster_reference and a multi-cluster eval series through 270. Move this eval to 271 (namespace app-271, fixture label 271-root-cause-noise) so the two 262_* directories no longer collide. Signed-off-by: Claude <noreply@anthropic.com>
The multi-cluster ES evals (254/259/260, added in #2125) require ELASTICSEARCH_URL/ELASTICSEARCH_API_KEY. Under --strict-setup-mode their missing-env setup failures abort the whole benchmark run. Wire in the existing repo secrets, mirroring eval-regression.yaml. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: avi@robusta.dev <avi@robusta.dev>
Evals 254/259/260 (added in #2125, titled 'multi-cluster awareness evals (red)') are not expected to pass yet, but were tagged 'regression'. The regression marker is reserved for tests that must always pass (30+ iters reliably on Sonnet-4.5), so a red eval there pollutes the must-pass suite and pulls these into fast-benchmark prematurely. Drop the regression tag; they keep medium/multi-cluster/elasticsearch/transparency so they can still be run/scored explicitly, just not as part of the regression gate. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: avi@robusta.dev <avi@robusta.dev>
… eval infra bugs Run #8 left the arms tied 49/53 with two SDK-only failures and two both-arm failures. All four root causes identified from Braintrust traces and fixed: - 236_image_spill_to_disk (SDK-only): the claude CLI caps MCP tool output at 25k tokens; the eval's report intentionally pads ~100k chars before the image attachment, so the cap silently dropped the image. Raise MAX_MCP_OUTPUT_TOKENS to 100000 in the engine's CLI env. Verified locally: PASS. - 259_wrong_cluster_logs_confusion (SDK-only, flaky): the agent disclosed the cluster mismatch as a property of the data ("index only contains eu-west-2") but the eval requires agent-identity framing ("this agent is connected to eu-west-2, not us-east-2"). One-line addition to the cluster sentence in the system prompt. Verified locally: PASS. - 19_detect_missing_app_details (both arms failed): the eval (added red in #2125) asks Holmes to surface db-certs-authenticator as the close match for personal-certs-validator, but no such resource was ever deployed — nothing could pass it. Add before_test/after_test creating it as a crashlooping deployment in app-19 with a discoverable cause (missing CA bundle) in its logs. Verified on a live cluster (SDK engine): PASS. - 227_count_configmaps_per_namespace (both arms failed): evals 229 and 249 deployed into namespace app-227 (copy-paste, violating the app-<testid> convention), so "namespaces starting with app-227" truthfully totalled 645 ConfigMaps against the eval's hardcoded 644 — both engines counted the live cluster correctly and were marked wrong. Move 229 to app-229 and 249 to app-249 (249's own prompt already said app-249). Verified on a live cluster: ground truth is 644 again. Signed-off-by: Claude <noreply@anthropic.com>
Splits PR #2042 into two: this is the red half. Tests/evals + supporting cleanup, no prompt change.
Without the companion green PR (#TBD —
claude/multi-cluster-prompt-green), the new evals are expected to be RED on master. They demonstrate the desired behavior; the prompt change in the green PR turns them green.What this PR contains
13 evals (all tagged
multi-cluster)elasticsearch_dr_test_log_checkcompanyopswebjobvsCompany.Ops*).detect_missing_app_detailspersonal-certs-validatorvsdb-certs-authenticator— transparency + usefulness both required.wrong_cluster_logs_confusionproduction-eu-west-2, user asks aboutproduction-us-east-2; ES has only local data with red-herring 500s.global_es_remote_cluster_logstime_window_gap_external_dataambiguous_cluster_referencecluster_name="production"; user says "the US prod cluster".region_suffixed_service_siblingspayment-service-us/-eu/-ap; user says "payment-service".wrong_env_same_regionstaging-eu-west-2; user asks aboutproduction-eu-west-2.multicluster_labeling_disciplinetoolset_disabled_vs_no_datacluster_name_aliasacme-prod-eu-west-1— must NOT over-correct.kubectl_wrong_clustermixed_sources_route_to_externalnamespace_collisionRun the whole set with:
Supporting changes
tests/llm/test_ask_holmes.py: passtest_case.cluster_namethrough tobuild_initial_ask_messagesin the CLI test path. Previously it was silently dropped, socluster_namein test fixtures had no effect on the system prompt for ask-mode evals.tests/llm/utils/braintrust.py+tests/llm/utils/langfuse.py: purge legacy dataset-upload code (BraintrustEvalHelper,upload_test_cases,find_dataset_row_by_test_case, etc.). We don't use Braintrust datasets in this repo — all eval logging goes via experiment span logs. The legacy code also had a real bug (input=inputwhereinputwas the Python builtin) that corrupted dataset records.tests/llm/test_holmes_checks.py: dropdataset_record_id=fromeval_span.logcalls (we no longer link to a dataset).pyproject.toml: register the newmulti-clusterpytest marker.Companion PR
Merge AFTER the green PR (
claude/multi-cluster-prompt-green), which adds the cluster-awareness rule toholmes/plugins/prompts/generic_ask.jinja2. Order doesn't matter for the git merge, but red-then-green will show the red→green transition in CI; green-first will show no red phase.Generated by Claude Code
Summary by CodeRabbit