prompt: multi-cluster scope-awareness (green half of #2042 split) - #2126
Conversation
Expands the existing "You are running on cluster X" bullet into a
multi-cluster-aware procedure that distinguishes kubectl-bound data
(only the local cluster) from external observability toolsets
(Elasticsearch / Datadog / Loki / etc., which may contain data from
many clusters). When the user names a cluster / region / env other
than the local one, Holmes:
1. Investigates using external toolsets (does NOT refuse).
2. Verifies each finding's cluster against the user's named cluster
(via the data's own cluster/region/index/source fields).
3. If matched → investigates normally with that cluster's findings.
4. If not matched → states plainly that no data exists for the
requested cluster, labels what was found in adjacent clusters,
and suggests pointing Holmes at the right data source / agent.
Together with the new evals (separate "red" PR) this turns the
multi-cluster awareness suite green:
254: similar-name service disambiguation
259: wrong-cluster local-data-only confusion
260: global ES topology with remote data
261: time-window gap
262: ambiguous cluster reference
263: region-suffixed sibling services
264: wrong env, same region
265: multi-cluster labeling discipline
266: toolset-disabled vs no-data
267: cluster-name alias for local cluster
268/269/270: kubectl-only / mixed-source / namespace-collision
Local validation (3 iterations × 3 models, Sonnet 4.5 / Opus 4.6 /
Opus 4.7) on the elasticsearch-based subset was clean. CI on the
"red" PR will demonstrate the failing state; merging this on top
turns those failures green.
Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.
Tip: disable this comment in your organization's Code Review settings.
📂 Previous Runs
|
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions | Src |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 28.4s | 4 | 8 | $0.2129 | 69,587 | 67,859 | 19,466 | 1,728 | 783 | 47,402 | 20,457 | 183 | — | src |
| ✅ | 101_loki_historical_logs_pod_deleted | 45.2s | 5 | 9 | $0.2516 | 88,412 | 85,696 | 19,949 | 2,716 | 880 | 65,020 | 20,676 | 471 | — | src |
| ✅ | 112_find_pvcs_by_uuid | 12.7s | 2 | 2 | $0.1426 | 31,433 | 30,633 | 16,461 | 800 | 446 | 14,169 | 16,464 | 172 | — | src |
| ✅ | 12_job_crashing | 31.5s | 6 | 10 | $0.2354 | 108,305 | 106,556 | 20,004 | 1,749 | 468 | 86,113 | 20,443 | 77 | — | src |
| ✅ | 176_network_policy_blocking_traffic_no_skills | 29.1s | 5 | 11 | $0.2232 | 86,501 | 84,678 | 19,490 | 1,823 | 532 | 64,534 | 20,144 | 289 | — | src |
| ✅ | 227_count_configmaps_per_namespace[0] | 14.6s | 3 | 6 | $0.1483 | 45,895 | 45,188 | 16,503 | 707 | 432 | 28,681 | 16,507 | 29 | — | src |
| ✅ | 243_pod_names_contain_service | 25.5s | 3 | 7 | $0.1844 | 49,149 | 47,570 | 17,848 | 1,579 | 572 | 29,254 | 18,316 | 202 | — | src |
| ✅ | 24_misconfigured_pvc | 25.4s | 4 | 11 | $0.2078 | 67,954 | 66,279 | 19,069 | 1,675 | 590 | 46,174 | 20,105 | 43 | — | src |
| ✅ | 254_elasticsearch_dr_test_log_check | 54.4s | 9 | 12 | $0.2746 | 117,612 | 113,868 | 17,003 | 3,744 | 913 | 96,570 | 17,298 | 239 | — | src |
| ✅ | 259_wrong_cluster_logs_confusion | 76.9s | 9 | 12 | $0.3418 | 133,731 | 128,286 | 18,291 | 5,445 | 2,426 | 108,828 | 19,458 | 1,122 | — | src |
| ✅ | 260_global_es_remote_cluster_logs | 42.5s | 5 | 8 | $0.2123 | 67,174 | 64,385 | 15,641 | 2,789 | 1,299 | 48,481 | 15,904 | 563 | — | src |
| ✅ | 261_time_window_gap_external_data | 75.3s | 11 | 14 | $0.3217 | 154,006 | 149,536 | 18,418 | 4,470 | 945 | 131,106 | 18,430 | 804 | — | src |
| ✅ | 262_ambiguous_cluster_reference | 53.7s | 8 | 13 | $0.2724 | 110,924 | 107,358 | 17,313 | 3,566 | 897 | 88,876 | 18,482 | 531 | — | src |
| ✅ | 263_region_suffixed_service_siblings | 59.3s | 8 | 12 | $0.2630 | 105,591 | 101,877 | 16,381 | 3,714 | 1,259 | 85,242 | 16,635 | 431 | — | src |
| ✅ | 264_wrong_env_same_region | 45.1s | 6 | 7 | $0.2033 | 75,250 | 72,658 | 14,552 | 2,592 | 873 | 58,099 | 14,559 | 425 | — | src |
| ✅ | 265_multicluster_labeling_discipline | 47.2s | 8 | 8 | $0.2258 | 104,594 | 101,963 | 15,392 | 2,631 | 669 | 86,562 | 15,401 | 306 | — | src |
| ✅ | 266_toolset_disabled_vs_no_data | 8.5s | 1 | — | $0.0802 | 10,589 | 10,230 | 10,230 | 359 | 359 | 0 | 10,230 | 91 | — | src |
| ✅ | 267_cluster_name_alias | 54.4s | 10 | 14 | $0.2889 | 140,930 | 137,508 | 17,970 | 3,422 | 657 | 118,422 | 19,086 | 186 | — | src |
| ✅ | 268_kubectl_wrong_cluster | 44.5s | 4 | 7 | $0.2586 | 73,441 | 70,658 | 22,404 | 2,783 | 904 | 47,949 | 22,709 | 742 | — | src |
| ✅ | 269_mixed_sources_route_to_external | 59.9s | 10 | 13 | $0.3238 | 164,213 | 160,496 | 20,350 | 3,717 | 576 | 139,226 | 21,270 | 289 | — | src |
| ✅ | 270_namespace_collision | 23.0s | 3 | 8 | $0.1890 | 50,628 | 49,144 | 19,131 | 1,484 | 612 | 29,738 | 19,406 | 152 | — | src |
| ✅ | 43_current_datetime_from_prompt | 4.1s | 1 | — | $0.1006 | 14,274 | 14,156 | 14,156 | 118 | 118 | 0 | 14,156 | 78 | — | src |
| ✅ | 51_logs_summarize_errors | 17.6s | 3 | 2 | $0.1537 | 46,457 | 45,712 | 17,083 | 745 | 371 | 28,625 | 17,087 | 29 | — | src |
| ✅ | 61_exact_match_counting | 7.1s | 2 | 1 | $0.1136 | 28,852 | 28,629 | 14,493 | 223 | 154 | 14,133 | 14,496 | 36 | — | src |
| Total | 36.9s avg | 5.4 avg | 8.9 avg | $5.2295 | 1,945,502 | 1,890,923 | 22,404 | 54,579 | 2,426 | 1,463,204 | 427,719 | 7,490 | — |
Benchmark Comparison Details
Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded
- master-27089078550 (created: 2026-06-07)
Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded
- ci-benchmark-27081280174 (created: 2026-06-07)
Time comparison (seconds):
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (7h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 28.4s | 26.2s | ±0% | 41.4s | ↓31% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 45.2s | 47.6s | ±0% | 65.7s | ↓31% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 12.7s | 10.6s | ↑20% | 19.0s | ↓33% |
| 12_job_crashing (opus-4.6) 📄 | 31.5s | 27.6s | ↑14% | 42.6s | ↓26% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 29.1s | 32.6s | ↓11% | 44.3s | ↓34% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 14.6s | 13.2s | ±0% | 18.4s | ↓21% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 25.5s | 28.8s | ↓12% | 35.1s | ↓27% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 25.4s | 31.4s | ↓19% | 40.6s | ↓37% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 54.4s | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 76.9s | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 42.5s | — | — | — | — |
| 261_time_window_gap_external_data (opus-4.6) 📄 | 75.3s | — | — | — | — |
| 262_ambiguous_cluster_reference (opus-4.6) 📄 | 53.7s | — | — | — | — |
| 263_region_suffixed_service_siblings (opus-4.6) 📄 | 59.3s | — | — | — | — |
| 264_wrong_env_same_region (opus-4.6) 📄 | 45.1s | — | — | — | — |
| 265_multicluster_labeling_discipline (opus-4.6) 📄 | 47.2s | — | — | — | — |
| 266_toolset_disabled_vs_no_data (opus-4.6) 📄 | 8.5s | — | — | — | — |
| 267_cluster_name_alias (opus-4.6) 📄 | 54.4s | — | — | — | — |
| 268_kubectl_wrong_cluster (opus-4.6) 📄 | 44.5s | — | — | — | — |
| 269_mixed_sources_route_to_external (opus-4.6) 📄 | 59.9s | — | — | — | — |
| 270_namespace_collision (opus-4.6) 📄 | 23.0s | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 4.1s | 3.1s | ↑33% | 3.5s | ↑16% |
| 51_logs_summarize_errors (opus-4.6) 📄 | 17.6s | 18.8s | ±0% | 19.7s | ↓11% |
| 61_exact_match_counting (opus-4.6) 📄 | 7.1s | 7.1s | ±0% | 10.2s | ↓31% |
| Total (all, n=24) | 36.9s | 22.5s | — | 31.0s | — |
| Comparable (m=11, b=11) | 21.9s | 22.5s | ±0% | 31.0s | ↓29% |
Cost comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (7h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | $0.2129 | $0.2032 | ±0% | $0.3101 | ↓31% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | $0.2516 | $0.2642 | ±0% | $0.3829 | ↓34% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | $0.1426 | $0.1368 | ±0% | $0.2047 | ↓30% |
| 12_job_crashing (opus-4.6) 📄 | $0.2354 | $0.2173 | ±0% | $0.3167 | ↓26% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | $0.2232 | $0.2524 | ↓12% | $0.3179 | ↓30% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | $0.1483 | $0.1464 | ±0% | $0.2048 | ↓28% |
| 243_pod_names_contain_service (opus-4.6) 📄 | $0.1844 | $0.1872 | ±0% | $0.2662 | ↓31% |
| 24_misconfigured_pvc (opus-4.6) 📄 | $0.2078 | $0.2240 | ±0% | $0.3110 | ↓33% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | $0.2746 | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | $0.3418 | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | $0.2123 | — | — | — | — |
| 261_time_window_gap_external_data (opus-4.6) 📄 | $0.3217 | — | — | — | — |
| 262_ambiguous_cluster_reference (opus-4.6) 📄 | $0.2724 | — | — | — | — |
| 263_region_suffixed_service_siblings (opus-4.6) 📄 | $0.2630 | — | — | — | — |
| 264_wrong_env_same_region (opus-4.6) 📄 | $0.2033 | — | — | — | — |
| 265_multicluster_labeling_discipline (opus-4.6) 📄 | $0.2258 | — | — | — | — |
| 266_toolset_disabled_vs_no_data (opus-4.6) 📄 | $0.0802 | — | — | — | — |
| 267_cluster_name_alias (opus-4.6) 📄 | $0.2889 | — | — | — | — |
| 268_kubectl_wrong_cluster (opus-4.6) 📄 | $0.2586 | — | — | — | — |
| 269_mixed_sources_route_to_external (opus-4.6) 📄 | $0.3238 | — | — | — | — |
| 270_namespace_collision (opus-4.6) 📄 | $0.1890 | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | $0.1006 | $0.0110 | ↑815% | $0.1190 | ↓15% |
| 51_logs_summarize_errors (opus-4.6) 📄 | $0.1537 | $0.1488 | ±0% | $0.2018 | ↓24% |
| 61_exact_match_counting (opus-4.6) 📄 | $0.1136 | $0.1112 | ±0% | $0.1518 | ↓25% |
| Total (all, n=24) | $0.2179 | $0.1730 | — | $0.2534 | — |
| Comparable (m=11, b=11) | $0.1795 | $0.1730 | ±0% | $0.2534 | ↓29% |
Total tokens comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (7h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 69,587 | 67,778 | ±0% | 132,351 | ↓47% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 88,412 | 88,307 | ±0% | 144,907 | ↓39% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 31,433 | 30,686 | ±0% | 61,219 | ↓49% |
| 12_job_crashing (opus-4.6) 📄 | 108,305 | 69,487 | ↑56% | 136,570 | ↓21% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 86,501 | 108,726 | ↓20% | 113,514 | ↓24% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 45,895 | 44,988 | ±0% | 76,715 | ↓40% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 49,149 | 48,612 | ±0% | 104,698 | ↓53% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 67,954 | 68,923 | ±0% | 132,672 | ↓49% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 117,612 | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 133,731 | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 67,174 | — | — | — | — |
| 261_time_window_gap_external_data (opus-4.6) 📄 | 154,006 | — | — | — | — |
| 262_ambiguous_cluster_reference (opus-4.6) 📄 | 110,924 | — | — | — | — |
| 263_region_suffixed_service_siblings (opus-4.6) 📄 | 105,591 | — | — | — | — |
| 264_wrong_env_same_region (opus-4.6) 📄 | 75,250 | — | — | — | — |
| 265_multicluster_labeling_discipline (opus-4.6) 📄 | 104,594 | — | — | — | — |
| 266_toolset_disabled_vs_no_data (opus-4.6) 📄 | 10,589 | — | — | — | — |
| 267_cluster_name_alias (opus-4.6) 📄 | 140,930 | — | — | — | — |
| 268_kubectl_wrong_cluster (opus-4.6) 📄 | 73,441 | — | — | — | — |
| 269_mixed_sources_route_to_external (opus-4.6) 📄 | 164,213 | — | — | — | — |
| 270_namespace_collision (opus-4.6) 📄 | 50,628 | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 14,274 | 13,970 | ±0% | 17,001 | ↓16% |
| 51_logs_summarize_errors (opus-4.6) 📄 | 46,457 | 45,163 | ±0% | 77,120 | ↓40% |
| 61_exact_match_counting (opus-4.6) 📄 | 28,852 | 28,231 | ±0% | 52,716 | ↓45% |
| Total (all, n=24) | 81,063 | 55,897 | — | 95,408 | — |
| Comparable (m=11, b=11) | 57,893 | 55,897 | ±0% | 95,408 | ↓39% |
Cached tokens comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (7h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 47,402 | 46,700 | ±0% | 102,543 | ↓54% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 65,020 | 64,209 | ±0% | 109,927 | ↓41% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 14,169 | 13,861 | ±0% | 38,090 | ↓63% |
| 12_job_crashing (opus-4.6) 📄 | 86,113 | 46,249 | ↑86% | 107,265 | ↓20% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 64,534 | 84,274 | ↓23% | 82,679 | ↓22% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 28,681 | 28,072 | ±0% | 54,602 | ↓47% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 29,254 | 28,690 | ±0% | 78,754 | ↓63% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 46,174 | 46,110 | ±0% | 103,729 | ↓55% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 96,570 | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 108,828 | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 48,481 | — | — | — | — |
| 261_time_window_gap_external_data (opus-4.6) 📄 | 131,106 | — | — | — | — |
| 262_ambiguous_cluster_reference (opus-4.6) 📄 | 88,876 | — | — | — | — |
| 263_region_suffixed_service_siblings (opus-4.6) 📄 | 85,242 | — | — | — | — |
| 264_wrong_env_same_region (opus-4.6) 📄 | 58,099 | — | — | — | — |
| 265_multicluster_labeling_discipline (opus-4.6) 📄 | 86,562 | — | — | — | — |
| 266_toolset_disabled_vs_no_data (opus-4.6) 📄 | — | — | — | — | — |
| 267_cluster_name_alias (opus-4.6) 📄 | 118,422 | — | — | — | — |
| 268_kubectl_wrong_cluster (opus-4.6) 📄 | 47,949 | — | — | — | — |
| 269_mixed_sources_route_to_external (opus-4.6) 📄 | 139,226 | — | — | — | — |
| 270_namespace_collision (opus-4.6) 📄 | 29,738 | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | — | 13,845 | — | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 28,625 | 28,012 | ±0% | 55,156 | ↓48% |
| 61_exact_match_counting (opus-4.6) 📄 | 14,133 | 13,825 | ±0% | 34,481 | ↓59% |
| Total (all, n=24) | 60,967 | 37,622 | — | 69,748 | — |
| Comparable (m=10, b=10) | 42,410 | 40,000 | ±0% | 76,723 | ↓45% |
Turns comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (7h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 4 | 4 | ±0% | 6 | ↓33% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 5 | 5 | ±0% | 6 | ↓17% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 2 | 2 | ±0% | 3 | ↓33% |
| 12_job_crashing (opus-4.6) 📄 | 6 | 4 | ↑50% | 6 | ±0% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 5 | 6 | ↓17% | 5 | ±0% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 3 | 3 | ±0% | 4 | ↓25% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 3 | 3 | ±0% | 5 | ↓40% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 4 | 4 | ±0% | 6 | ↓33% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 9 | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 9 | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 5 | — | — | — | — |
| 261_time_window_gap_external_data (opus-4.6) 📄 | 11 | — | — | — | — |
| 262_ambiguous_cluster_reference (opus-4.6) 📄 | 8 | — | — | — | — |
| 263_region_suffixed_service_siblings (opus-4.6) 📄 | 8 | — | — | — | — |
| 264_wrong_env_same_region (opus-4.6) 📄 | 6 | — | — | — | — |
| 265_multicluster_labeling_discipline (opus-4.6) 📄 | 8 | — | — | — | — |
| 266_toolset_disabled_vs_no_data (opus-4.6) 📄 | 1 | — | — | — | — |
| 267_cluster_name_alias (opus-4.6) 📄 | 10 | — | — | — | — |
| 268_kubectl_wrong_cluster (opus-4.6) 📄 | 4 | — | — | — | — |
| 269_mixed_sources_route_to_external (opus-4.6) 📄 | 10 | — | — | — | — |
| 270_namespace_collision (opus-4.6) 📄 | 3 | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 1 | 1 | ±0% | 1 | ±0% |
| 51_logs_summarize_errors (opus-4.6) 📄 | 3 | 3 | ±0% | 4 | ↓25% |
| 61_exact_match_counting (opus-4.6) 📄 | 2 | 2 | ±0% | 3 | ↓33% |
| Total (all, n=24) | 5.4 | 3.4 | — | 4.5 | — |
| Comparable (m=11, b=11) | 3.5 | 3.4 | ±0% | 4.5 | ↓22% |
Tool calls comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (7h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 8 | 8 | ±0% | 13 | ↓38% |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 9 | 11 | ↓18% | 14 | ↓36% |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 2 | 2 | ±0% | 4 | ↓50% |
| 12_job_crashing (opus-4.6) 📄 | 10 | 9 | ↑11% | 15 | ↓33% |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 11 | 12 | ±0% | 15 | ↓27% |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 6 | 6 | ±0% | 9 | ↓33% |
| 243_pod_names_contain_service (opus-4.6) 📄 | 7 | 7 | ±0% | 11 | ↓36% |
| 24_misconfigured_pvc (opus-4.6) 📄 | 11 | 13 | ↓15% | 15 | ↓27% |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 12 | — | — | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 12 | — | — | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 8 | — | — | — | — |
| 261_time_window_gap_external_data (opus-4.6) 📄 | 14 | — | — | — | — |
| 262_ambiguous_cluster_reference (opus-4.6) 📄 | 13 | — | — | — | — |
| 263_region_suffixed_service_siblings (opus-4.6) 📄 | 12 | — | — | — | — |
| 264_wrong_env_same_region (opus-4.6) 📄 | 7 | — | — | — | — |
| 265_multicluster_labeling_discipline (opus-4.6) 📄 | 8 | — | — | — | — |
| 266_toolset_disabled_vs_no_data (opus-4.6) 📄 | — | — | — | — | — |
| 267_cluster_name_alias (opus-4.6) 📄 | 14 | — | — | — | — |
| 268_kubectl_wrong_cluster (opus-4.6) 📄 | 7 | — | — | — | — |
| 269_mixed_sources_route_to_external (opus-4.6) 📄 | 13 | — | — | — | — |
| 270_namespace_collision (opus-4.6) 📄 | 8 | — | — | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | — | — | — | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 2 | 2 | ±0% | 5 | ↓60% |
| 61_exact_match_counting (opus-4.6) 📄 | 1 | 1 | ±0% | 3 | ↓67% |
| Total (all, n=24) | 8.1 | 7.1 | — | 10.4 | — |
| Comparable (m=10, b=10) | 6.7 | 7.1 | ±0% | 10.4 | ↓36% |
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
📖 Legend
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/multi-cluster-prompt-green -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
tags: regression
Or with more options (one per line):
/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
tags: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
tags |
Pytest tags / markers (no default - runs all tests!) |
id |
Eval ID / pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):
| Label | Effect |
|---|---|
evals-tag-<name> |
Run tests with tag <name> alongside regression |
evals-id-<name> |
Run a specific eval by test ID |
evals-model-<name> |
Override the model (use model list name, e.g. sonnet-4.5) |
Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5
🏷️ Valid tags
benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs
🤖 Valid models
deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/multi-cluster-prompt-green -f markers=regression -f filter=
WalkthroughThis PR updates the Holmes generic ask prompt template with two complementary instruction enhancements: a new multi-cluster verification block ensuring observability findings are validated against the user-specified cluster, and improved general guidance enforcing transparent entity-matching behavior and iterative tool usage patterns. ChangesPrompt instruction enhancements
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:ce9069f19
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:ce9069f19 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:ce9069f19
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:ce9069f19
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:ce9069f19
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:ce9069f19 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:ce9069f19
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:ce9069f19Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:ce9069f19 \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:ce9069f19Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:ce9069f19 \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:ce9069f19 |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
There was a problem hiding this comment.
🧹 Nitpick comments (1)
holmes/plugins/prompts/generic_ask.jinja2 (1)
16-16: ⚡ Quick winAvoid premature “no data” conclusion on first mismatch.
Step 4 currently reads like a global conclusion after a mismatch. Make the “no data for requested cluster” statement conditional on checking available results/sources and finding only mismatched clusters.
Suggested wording adjustment
- 4. If verification shows the data is from a different cluster (commonly `{{ cluster_name }}` itself, because that's the cluster this Holmes is on) → you do NOT have data for the cluster the user asked about. You MUST: (a) state this plainly and prominently at the top of your answer, (b) NOT present those mismatched findings as the root cause of the user's reported incident, (c) say which cluster(s) you DID find data for and label any summary of that data accordingly, (d) suggest the user point you at the right data source / Holmes agent for the cluster they actually care about. + 4. If the data you verified is from a different cluster (commonly `{{ cluster_name }}` itself), keep checking available results/sources for the user-requested cluster. + 5. If, after those checks, you still find no verified data for the cluster the user asked about, you MUST: (a) state this plainly and prominently at the top of your answer, (b) NOT present mismatched findings as the root cause of the user's reported incident, (c) say which cluster(s) you DID find data for and label any summary of that data accordingly, (d) suggest the user point you at the right data source / Holmes agent for the cluster they actually care about.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@holmes/plugins/prompts/generic_ask.jinja2` at line 16, The step labeled "4." in holmes/plugins/prompts/generic_ask.jinja2 currently asserts a global "no data for requested cluster" conclusion when a mismatch is found; change the logic and wording so that the statement is conditional: only emit the prominent "no data for requested cluster" message if, after checking all available results/sources, you find exclusively mismatched clusters (e.g., those matching {{ cluster_name }}). Update the Step 4 text to (a) first verify whether any matching sources exist, (b) if none exist, then state the no-data conclusion prominently and follow (b)-(d) as currently written, and (c) otherwise do not treat a single mismatch as a global no-data conclusion but instead label any mismatched findings as coming from other clusters and continue searching/reporting other sources.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@holmes/plugins/prompts/generic_ask.jinja2`:
- Line 16: The step labeled "4." in holmes/plugins/prompts/generic_ask.jinja2
currently asserts a global "no data for requested cluster" conclusion when a
mismatch is found; change the logic and wording so that the statement is
conditional: only emit the prominent "no data for requested cluster" message if,
after checking all available results/sources, you find exclusively mismatched
clusters (e.g., those matching {{ cluster_name }}). Update the Step 4 text to
(a) first verify whether any matching sources exist, (b) if none exist, then
state the no-data conclusion prominently and follow (b)-(d) as currently
written, and (c) otherwise do not treat a single mismatch as a global no-data
conclusion but instead label any mismatched findings as coming from other
clusters and continue searching/reporting other sources.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: a3eb8a93-02b5-490c-9698-56b71bc278c6
📒 Files selected for processing (1)
holmes/plugins/prompts/generic_ask.jinja2
Splits PR #2042 into two — this is the green half. Just the system-prompt change. Pair with #2125 (red) which contains the evals.
What this PR changes
holmes/plugins/prompts/generic_ask.jinja2: expands the existing* You are running on cluster Xbullet into a multi-cluster-aware procedure.Old:
New: distinguishes kubectl-bound data (only the local cluster) from external observability toolsets (Elasticsearch / Datadog / Loki / etc., which may contain data from many clusters). When the user names a cluster / region / env other than the local one, Holmes:
cluster/region/environment/kubernetes.cluster.namefield, or the index/source name).Two common topologies are explicitly called out:
Holmes doesn't know upfront which one applies — it must look at what the data actually contains.
Companion PR
This turns #2125's evals green:
Local validation across 3 iterations × 3 models (Sonnet 4.5 / Opus 4.6 / Opus 4.7) on the ES-based subset was clean. CI on #2125 will demonstrate the failing state; merging this on top turns those failures green.
Generated by Claude Code
Summary by CodeRabbit
Release Notes