Skip to content

prompt: multi-cluster scope-awareness (green half of #2042 split) - #2126

Merged
aantn merged 3 commits into
masterfrom
claude/multi-cluster-prompt-green
Jun 7, 2026
Merged

aantn merged 3 commits into
masterfrom
claude/multi-cluster-prompt-green

Conversation

@aantn

@aantn aantn commented Jun 4, 2026 •

Copy link
Copy Markdown
Collaborator

Splits PR #2042 into two — this is the green half. Just the system-prompt change. Pair with #2125 (red) which contains the evals.

What this PR changes

holmes/plugins/prompts/generic_ask.jinja2: expands the existing * You are running on cluster X bullet into a multi-cluster-aware procedure.

Old:

{% if cluster_name -%}
* You are running on cluster {{ cluster_name }}.
{%- endif %}

New: distinguishes kubectl-bound data (only the local cluster) from external observability toolsets (Elasticsearch / Datadog / Loki / etc., which may contain data from many clusters). When the user names a cluster / region / env other than the local one, Holmes:

  1. Investigates using external toolsets — does NOT refuse.
  2. Verifies each finding's cluster against the user's named cluster (via the data's own cluster / region / environment / kubernetes.cluster.name field, or the index/source name).
  3. If matched → investigates normally with that cluster's findings.
  4. If not matched → states plainly that no data exists for the requested cluster, labels what was found in adjacent clusters, suggests pointing Holmes at the right data source / agent.

Two common topologies are explicitly called out:

  • one Holmes per cluster (kubectl + external mostly scoped to that cluster)
  • one Holmes with a global observability backend covering many clusters

Holmes doesn't know upfront which one applies — it must look at what the data actually contains.

Companion PR

This turns #2125's evals green:

  • 254, 19 — name disambiguation
  • 259 — wrong cluster, local-only data
  • 260 — global topology, remote cluster data
  • 261 — time-window gap
  • 262 — ambiguous cluster reference
  • 263 — region-suffixed sibling services
  • 264 — wrong env, same region
  • 265 — labeling discipline
  • 266 — toolset-disabled vs no-data
  • 267 — cluster-name alias (must NOT over-correct)
  • 268/269/270 — kubectl-only / mixed-source / namespace-collision

Local validation across 3 iterations × 3 models (Sonnet 4.5 / Opus 4.6 / Opus 4.7) on the ES-based subset was clean. CI on #2125 will demonstrate the failing state; merging this on top turns those failures green.


Generated by Claude Code

Summary by CodeRabbit

Release Notes

  • Improvements
    • Enhanced cluster awareness to ensure observability data is correctly attributed to your specified cluster and prevent misinterpretation of cross-cluster data.
    • Improved transparency when exact entity data is unavailable, with clearer reporting of similarly-named alternatives for user verification.
    • Refined troubleshooting guidance to better support iterative investigation workflows.

Expands the existing "You are running on cluster X" bullet into a
multi-cluster-aware procedure that distinguishes kubectl-bound data
(only the local cluster) from external observability toolsets
(Elasticsearch / Datadog / Loki / etc., which may contain data from
many clusters). When the user names a cluster / region / env other
than the local one, Holmes:

  1. Investigates using external toolsets (does NOT refuse).
  2. Verifies each finding's cluster against the user's named cluster
     (via the data's own cluster/region/index/source fields).
  3. If matched → investigates normally with that cluster's findings.
  4. If not matched → states plainly that no data exists for the
     requested cluster, labels what was found in adjacent clusters,
     and suggests pointing Holmes at the right data source / agent.

Together with the new evals (separate "red" PR) this turns the
multi-cluster awareness suite green:

  254: similar-name service disambiguation
  259: wrong-cluster local-data-only confusion
  260: global ES topology with remote data
  261: time-window gap
  262: ambiguous cluster reference
  263: region-suffixed sibling services
  264: wrong env, same region
  265: multi-cluster labeling discipline
  266: toolset-disabled vs no-data
  267: cluster-name alias for local cluster
  268/269/270: kubectl-only / mixed-source / namespace-collision

Local validation (3 iterations × 3 models, Sonnet 4.5 / Opus 4.6 /
Opus 4.7) on the elasticsearch-based subset was clean. CI on the
"red" PR will demonstrate the failing state; merging this on top
turns those failures green.

Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@github-actions

github-actions Bot commented Jun 4, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

⚠️ 2 older runs truncated

Older runs were omitted to stay under GitHub's 64KB comment size limit.


✅ Results of HolmesGPT evals

Automatically triggered by commit c536728 on branch claude/multi-cluster-prompt-green (labels: evals-tag-regression, evals-tag-multi-cluster)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 24/24 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Src
✅ 09_crashpod 28.4s 4 8 $0.2129 69,587 67,859 19,466 1,728 783 47,402 20,457 183 — src
✅ 101_loki_historical_logs_pod_deleted 45.2s 5 9 $0.2516 88,412 85,696 19,949 2,716 880 65,020 20,676 471 — src
✅ 112_find_pvcs_by_uuid 12.7s 2 2 $0.1426 31,433 30,633 16,461 800 446 14,169 16,464 172 — src
✅ 12_job_crashing 31.5s 6 10 $0.2354 108,305 106,556 20,004 1,749 468 86,113 20,443 77 — src
✅ 176_network_policy_blocking_traffic_no_skills 29.1s 5 11 $0.2232 86,501 84,678 19,490 1,823 532 64,534 20,144 289 — src
✅ 227_count_configmaps_per_namespace[0] 14.6s 3 6 $0.1483 45,895 45,188 16,503 707 432 28,681 16,507 29 — src
✅ 243_pod_names_contain_service 25.5s 3 7 $0.1844 49,149 47,570 17,848 1,579 572 29,254 18,316 202 — src
✅ 24_misconfigured_pvc 25.4s 4 11 $0.2078 67,954 66,279 19,069 1,675 590 46,174 20,105 43 — src
✅ 254_elasticsearch_dr_test_log_check 54.4s 9 12 $0.2746 117,612 113,868 17,003 3,744 913 96,570 17,298 239 — src
✅ 259_wrong_cluster_logs_confusion 76.9s 9 12 $0.3418 133,731 128,286 18,291 5,445 2,426 108,828 19,458 1,122 — src
✅ 260_global_es_remote_cluster_logs 42.5s 5 8 $0.2123 67,174 64,385 15,641 2,789 1,299 48,481 15,904 563 — src
✅ 261_time_window_gap_external_data 75.3s 11 14 $0.3217 154,006 149,536 18,418 4,470 945 131,106 18,430 804 — src
✅ 262_ambiguous_cluster_reference 53.7s 8 13 $0.2724 110,924 107,358 17,313 3,566 897 88,876 18,482 531 — src
✅ 263_region_suffixed_service_siblings 59.3s 8 12 $0.2630 105,591 101,877 16,381 3,714 1,259 85,242 16,635 431 — src
✅ 264_wrong_env_same_region 45.1s 6 7 $0.2033 75,250 72,658 14,552 2,592 873 58,099 14,559 425 — src
✅ 265_multicluster_labeling_discipline 47.2s 8 8 $0.2258 104,594 101,963 15,392 2,631 669 86,562 15,401 306 — src
✅ 266_toolset_disabled_vs_no_data 8.5s 1 — $0.0802 10,589 10,230 10,230 359 359 0 10,230 91 — src
✅ 267_cluster_name_alias 54.4s 10 14 $0.2889 140,930 137,508 17,970 3,422 657 118,422 19,086 186 — src
✅ 268_kubectl_wrong_cluster 44.5s 4 7 $0.2586 73,441 70,658 22,404 2,783 904 47,949 22,709 742 — src
✅ 269_mixed_sources_route_to_external 59.9s 10 13 $0.3238 164,213 160,496 20,350 3,717 576 139,226 21,270 289 — src
✅ 270_namespace_collision 23.0s 3 8 $0.1890 50,628 49,144 19,131 1,484 612 29,738 19,406 152 — src
✅ 43_current_datetime_from_prompt 4.1s 1 — $0.1006 14,274 14,156 14,156 118 118 0 14,156 78 — src
✅ 51_logs_summarize_errors 17.6s 3 2 $0.1537 46,457 45,712 17,083 745 371 28,625 17,087 29 — src
✅ 61_exact_match_counting 7.1s 2 1 $0.1136 28,852 28,629 14,493 223 154 14,133 14,496 36 — src
Total 36.9s avg 5.4 avg 8.9 avg $5.2295 1,945,502 1,890,923 22,404 54,579 2,426 1,463,204 427,719 7,490 —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 185 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (1h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 28.4s 26.2s ±0% 41.4s ↓31%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 45.2s 47.6s ±0% 65.7s ↓31%
112_find_pvcs_by_uuid (opus-4.6) 📄 12.7s 10.6s ↑20% 19.0s ↓33%
12_job_crashing (opus-4.6) 📄 31.5s 27.6s ↑14% 42.6s ↓26%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 29.1s 32.6s ↓11% 44.3s ↓34%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 14.6s 13.2s ±0% 18.4s ↓21%
243_pod_names_contain_service (opus-4.6) 📄 25.5s 28.8s ↓12% 35.1s ↓27%
24_misconfigured_pvc (opus-4.6) 📄 25.4s 31.4s ↓19% 40.6s ↓37%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 54.4s — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 76.9s — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 42.5s — — — —
261_time_window_gap_external_data (opus-4.6) 📄 75.3s — — — —
262_ambiguous_cluster_reference (opus-4.6) 📄 53.7s — — — —
263_region_suffixed_service_siblings (opus-4.6) 📄 59.3s — — — —
264_wrong_env_same_region (opus-4.6) 📄 45.1s — — — —
265_multicluster_labeling_discipline (opus-4.6) 📄 47.2s — — — —
266_toolset_disabled_vs_no_data (opus-4.6) 📄 8.5s — — — —
267_cluster_name_alias (opus-4.6) 📄 54.4s — — — —
268_kubectl_wrong_cluster (opus-4.6) 📄 44.5s — — — —
269_mixed_sources_route_to_external (opus-4.6) 📄 59.9s — — — —
270_namespace_collision (opus-4.6) 📄 23.0s — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 4.1s 3.1s ↑33% 3.5s ↑16%
51_logs_summarize_errors (opus-4.6) 📄 17.6s 18.8s ±0% 19.7s ↓11%
61_exact_match_counting (opus-4.6) 📄 7.1s 7.1s ±0% 10.2s ↓31%
Total (all, n=24) 36.9s 22.5s — 31.0s —
Comparable (m=11, b=11) 21.9s 22.5s ±0% 31.0s ↓29%

Cost comparison:

Test case This branch master (1h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.2129 $0.2032 ±0% $0.3101 ↓31%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.2516 $0.2642 ±0% $0.3829 ↓34%
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.1426 $0.1368 ±0% $0.2047 ↓30%
12_job_crashing (opus-4.6) 📄 $0.2354 $0.2173 ±0% $0.3167 ↓26%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.2232 $0.2524 ↓12% $0.3179 ↓30%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.1483 $0.1464 ±0% $0.2048 ↓28%
243_pod_names_contain_service (opus-4.6) 📄 $0.1844 $0.1872 ±0% $0.2662 ↓31%
24_misconfigured_pvc (opus-4.6) 📄 $0.2078 $0.2240 ±0% $0.3110 ↓33%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 $0.2746 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 $0.3418 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 $0.2123 — — — —
261_time_window_gap_external_data (opus-4.6) 📄 $0.3217 — — — —
262_ambiguous_cluster_reference (opus-4.6) 📄 $0.2724 — — — —
263_region_suffixed_service_siblings (opus-4.6) 📄 $0.2630 — — — —
264_wrong_env_same_region (opus-4.6) 📄 $0.2033 — — — —
265_multicluster_labeling_discipline (opus-4.6) 📄 $0.2258 — — — —
266_toolset_disabled_vs_no_data (opus-4.6) 📄 $0.0802 — — — —
267_cluster_name_alias (opus-4.6) 📄 $0.2889 — — — —
268_kubectl_wrong_cluster (opus-4.6) 📄 $0.2586 — — — —
269_mixed_sources_route_to_external (opus-4.6) 📄 $0.3238 — — — —
270_namespace_collision (opus-4.6) 📄 $0.1890 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 $0.1006 $0.0110 ↑815% $0.1190 ↓15%
51_logs_summarize_errors (opus-4.6) 📄 $0.1537 $0.1488 ±0% $0.2018 ↓24%
61_exact_match_counting (opus-4.6) 📄 $0.1136 $0.1112 ±0% $0.1518 ↓25%
Total (all, n=24) $0.2179 $0.1730 — $0.2534 —
Comparable (m=11, b=11) $0.1795 $0.1730 ±0% $0.2534 ↓29%

Total tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 69,587 67,778 ±0% 132,351 ↓47%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 88,412 88,307 ±0% 144,907 ↓39%
112_find_pvcs_by_uuid (opus-4.6) 📄 31,433 30,686 ±0% 61,219 ↓49%
12_job_crashing (opus-4.6) 📄 108,305 69,487 ↑56% 136,570 ↓21%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 86,501 108,726 ↓20% 113,514 ↓24%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 45,895 44,988 ±0% 76,715 ↓40%
243_pod_names_contain_service (opus-4.6) 📄 49,149 48,612 ±0% 104,698 ↓53%
24_misconfigured_pvc (opus-4.6) 📄 67,954 68,923 ±0% 132,672 ↓49%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 117,612 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 133,731 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 67,174 — — — —
261_time_window_gap_external_data (opus-4.6) 📄 154,006 — — — —
262_ambiguous_cluster_reference (opus-4.6) 📄 110,924 — — — —
263_region_suffixed_service_siblings (opus-4.6) 📄 105,591 — — — —
264_wrong_env_same_region (opus-4.6) 📄 75,250 — — — —
265_multicluster_labeling_discipline (opus-4.6) 📄 104,594 — — — —
266_toolset_disabled_vs_no_data (opus-4.6) 📄 10,589 — — — —
267_cluster_name_alias (opus-4.6) 📄 140,930 — — — —
268_kubectl_wrong_cluster (opus-4.6) 📄 73,441 — — — —
269_mixed_sources_route_to_external (opus-4.6) 📄 164,213 — — — —
270_namespace_collision (opus-4.6) 📄 50,628 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 14,274 13,970 ±0% 17,001 ↓16%
51_logs_summarize_errors (opus-4.6) 📄 46,457 45,163 ±0% 77,120 ↓40%
61_exact_match_counting (opus-4.6) 📄 28,852 28,231 ±0% 52,716 ↓45%
Total (all, n=24) 81,063 55,897 — 95,408 —
Comparable (m=11, b=11) 57,893 55,897 ±0% 95,408 ↓39%

Cached tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 47,402 46,700 ±0% 102,543 ↓54%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 65,020 64,209 ±0% 109,927 ↓41%
112_find_pvcs_by_uuid (opus-4.6) 📄 14,169 13,861 ±0% 38,090 ↓63%
12_job_crashing (opus-4.6) 📄 86,113 46,249 ↑86% 107,265 ↓20%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 64,534 84,274 ↓23% 82,679 ↓22%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 28,681 28,072 ±0% 54,602 ↓47%
243_pod_names_contain_service (opus-4.6) 📄 29,254 28,690 ±0% 78,754 ↓63%
24_misconfigured_pvc (opus-4.6) 📄 46,174 46,110 ±0% 103,729 ↓55%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 96,570 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 108,828 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 48,481 — — — —
261_time_window_gap_external_data (opus-4.6) 📄 131,106 — — — —
262_ambiguous_cluster_reference (opus-4.6) 📄 88,876 — — — —
263_region_suffixed_service_siblings (opus-4.6) 📄 85,242 — — — —
264_wrong_env_same_region (opus-4.6) 📄 58,099 — — — —
265_multicluster_labeling_discipline (opus-4.6) 📄 86,562 — — — —
266_toolset_disabled_vs_no_data (opus-4.6) 📄 — — — — —
267_cluster_name_alias (opus-4.6) 📄 118,422 — — — —
268_kubectl_wrong_cluster (opus-4.6) 📄 47,949 — — — —
269_mixed_sources_route_to_external (opus-4.6) 📄 139,226 — — — —
270_namespace_collision (opus-4.6) 📄 29,738 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 — 13,845 — — —
51_logs_summarize_errors (opus-4.6) 📄 28,625 28,012 ±0% 55,156 ↓48%
61_exact_match_counting (opus-4.6) 📄 14,133 13,825 ±0% 34,481 ↓59%
Total (all, n=24) 60,967 37,622 — 69,748 —
Comparable (m=10, b=10) 42,410 40,000 ±0% 76,723 ↓45%

Turns comparison:

Test case This branch master (1h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 4 4 ±0% 6 ↓33%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 5 5 ±0% 6 ↓17%
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% 3 ↓33%
12_job_crashing (opus-4.6) 📄 6 4 ↑50% 6 ±0%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 5 6 ↓17% 5 ±0%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 3 3 ±0% 4 ↓25%
243_pod_names_contain_service (opus-4.6) 📄 3 3 ±0% 5 ↓40%
24_misconfigured_pvc (opus-4.6) 📄 4 4 ±0% 6 ↓33%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 9 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 9 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 5 — — — —
261_time_window_gap_external_data (opus-4.6) 📄 11 — — — —
262_ambiguous_cluster_reference (opus-4.6) 📄 8 — — — —
263_region_suffixed_service_siblings (opus-4.6) 📄 8 — — — —
264_wrong_env_same_region (opus-4.6) 📄 6 — — — —
265_multicluster_labeling_discipline (opus-4.6) 📄 8 — — — —
266_toolset_disabled_vs_no_data (opus-4.6) 📄 1 — — — —
267_cluster_name_alias (opus-4.6) 📄 10 — — — —
268_kubectl_wrong_cluster (opus-4.6) 📄 4 — — — —
269_mixed_sources_route_to_external (opus-4.6) 📄 10 — — — —
270_namespace_collision (opus-4.6) 📄 3 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% 1 ±0%
51_logs_summarize_errors (opus-4.6) 📄 3 3 ±0% 4 ↓25%
61_exact_match_counting (opus-4.6) 📄 2 2 ±0% 3 ↓33%
Total (all, n=24) 5.4 3.4 — 4.5 —
Comparable (m=11, b=11) 3.5 3.4 ±0% 4.5 ↓22%

Tool calls comparison:

Test case This branch master (1h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 8 8 ±0% 13 ↓38%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 9 11 ↓18% 14 ↓36%
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% 4 ↓50%
12_job_crashing (opus-4.6) 📄 10 9 ↑11% 15 ↓33%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 11 12 ±0% 15 ↓27%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 6 6 ±0% 9 ↓33%
243_pod_names_contain_service (opus-4.6) 📄 7 7 ±0% 11 ↓36%
24_misconfigured_pvc (opus-4.6) 📄 11 13 ↓15% 15 ↓27%
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 12 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 12 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 8 — — — —
261_time_window_gap_external_data (opus-4.6) 📄 14 — — — —
262_ambiguous_cluster_reference (opus-4.6) 📄 13 — — — —
263_region_suffixed_service_siblings (opus-4.6) 📄 12 — — — —
264_wrong_env_same_region (opus-4.6) 📄 7 — — — —
265_multicluster_labeling_discipline (opus-4.6) 📄 8 — — — —
266_toolset_disabled_vs_no_data (opus-4.6) 📄 — — — — —
267_cluster_name_alias (opus-4.6) 📄 14 — — — —
268_kubectl_wrong_cluster (opus-4.6) 📄 7 — — — —
269_mixed_sources_route_to_external (opus-4.6) 📄 13 — — — —
270_namespace_collision (opus-4.6) 📄 8 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 2 2 ±0% 5 ↓60%
61_exact_match_counting (opus-4.6) 📄 1 1 ±0% 3 ↓67%
Total (all, n=24) 8.1 7.1 — 10.4 —
Comparable (m=10, b=10) 6.7 7.1 ±0% 10.4 ↓36%

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/multi-cluster-prompt-green -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/multi-cluster-prompt-green -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented Jun 4, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

This PR updates the Holmes generic ask prompt template with two complementary instruction enhancements: a new multi-cluster verification block ensuring observability findings are validated against the user-specified cluster, and improved general guidance enforcing transparent entity-matching behavior and iterative tool usage patterns.

Changes

Prompt instruction enhancements

Layer / File(s) Summary
Multi-cluster awareness instructions
holmes/plugins/prompts/generic_ask.jinja2
Added cluster-specific instruction block explaining local vs. external observability data interpretation, mandatory cluster identity verification for returned findings, explicit handling of mismatched cluster data, and prohibition on silent relabeling of findings to different clusters.
General instructions for iterative tools and entity matching
holmes/plugins/prompts/generic_ask.jinja2
Strengthened guidance on iterative tool-gathering with repeated distinct tool calls; inserted new "Adjacent / similarly-named entities" directive requiring transparent reporting when exact entity data is absent but similarly-named entities exist, including how to report absence and advise user verification.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • HolmesGPT/holmesgpt#1905: Refactors and inlines prompt instructions into generic_ask.jinja2; the current PR extends the same composed template with cluster verification and entity-matching enhancements.
  • HolmesGPT/holmesgpt#1801: Adds an end-to-end Elasticsearch test asserting Holmes detects cluster mismatches; this PR adds explicit cluster-mismatch verification rules to the prompt template that enforce the same behavior.

Suggested reviewers

  • Sheeproid
  • arikalon1
  • moshemorad
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'prompt: multi-cluster scope-awareness (green half of #2042 split)' directly and specifically describes the main change: expanding the prompt template to add multi-cluster awareness.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jun 4, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for ce9069f19 (built in 4m 15s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:ce9069f19
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:ce9069f19 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:ce9069f19
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:ce9069f19
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:ce9069f19
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:ce9069f19 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:ce9069f19
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:ce9069f19

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:ce9069f19 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:ce9069f19

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:ce9069f19 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:ce9069f19

@netlify

netlify Bot commented Jun 4, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit c536728
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a2546ea18dee000098eeca3
😎 Deploy Preview https://deploy-preview-2126--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
holmes/plugins/prompts/generic_ask.jinja2 (1)

16-16: ⚡ Quick win

Avoid premature “no data” conclusion on first mismatch.

Step 4 currently reads like a global conclusion after a mismatch. Make the “no data for requested cluster” statement conditional on checking available results/sources and finding only mismatched clusters.

Suggested wording adjustment
-  4. If verification shows the data is from a different cluster (commonly `{{ cluster_name }}` itself, because that's the cluster this Holmes is on) → you do NOT have data for the cluster the user asked about. You MUST: (a) state this plainly and prominently at the top of your answer, (b) NOT present those mismatched findings as the root cause of the user's reported incident, (c) say which cluster(s) you DID find data for and label any summary of that data accordingly, (d) suggest the user point you at the right data source / Holmes agent for the cluster they actually care about.
+  4. If the data you verified is from a different cluster (commonly `{{ cluster_name }}` itself), keep checking available results/sources for the user-requested cluster.
+  5. If, after those checks, you still find no verified data for the cluster the user asked about, you MUST: (a) state this plainly and prominently at the top of your answer, (b) NOT present mismatched findings as the root cause of the user's reported incident, (c) say which cluster(s) you DID find data for and label any summary of that data accordingly, (d) suggest the user point you at the right data source / Holmes agent for the cluster they actually care about.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@holmes/plugins/prompts/generic_ask.jinja2` at line 16, The step labeled "4."
in holmes/plugins/prompts/generic_ask.jinja2 currently asserts a global "no data
for requested cluster" conclusion when a mismatch is found; change the logic and
wording so that the statement is conditional: only emit the prominent "no data
for requested cluster" message if, after checking all available results/sources,
you find exclusively mismatched clusters (e.g., those matching {{ cluster_name
}}). Update the Step 4 text to (a) first verify whether any matching sources
exist, (b) if none exist, then state the no-data conclusion prominently and
follow (b)-(d) as currently written, and (c) otherwise do not treat a single
mismatch as a global no-data conclusion but instead label any mismatched
findings as coming from other clusters and continue searching/reporting other
sources.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@holmes/plugins/prompts/generic_ask.jinja2`:
- Line 16: The step labeled "4." in holmes/plugins/prompts/generic_ask.jinja2
currently asserts a global "no data for requested cluster" conclusion when a
mismatch is found; change the logic and wording so that the statement is
conditional: only emit the prominent "no data for requested cluster" message if,
after checking all available results/sources, you find exclusively mismatched
clusters (e.g., those matching {{ cluster_name }}). Update the Step 4 text to
(a) first verify whether any matching sources exist, (b) if none exist, then
state the no-data conclusion prominently and follow (b)-(d) as currently
written, and (c) otherwise do not treat a single mismatch as a global no-data
conclusion but instead label any mismatched findings as coming from other
clusters and continue searching/reporting other sources.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: a3eb8a93-02b5-490c-9698-56b71bc278c6

📥 Commits

Reviewing files that changed from the base of the PR and between 4d0a70f and 5f23336.

📒 Files selected for processing (1)
  • holmes/plugins/prompts/generic_ask.jinja2

@aantn
aantn merged commit f039113 into master Jun 7, 2026
22 of 26 checks passed
@aantn
aantn deleted the claude/multi-cluster-prompt-green branch June 7, 2026 11:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants