Repository navigation
Refactor config class handling to support multiple config classes - #1741
Conversation
…ashboard config_classes Apply the same fix from d5fd607 (Grafana Tempo) to other toolsets: - ElasticsearchBaseToolset: use self.config_classes[0] instead of hardcoding ElasticsearchConfig in prerequisites_callable, so subclasses with custom config classes work correctly - GrafanaToolset: change config_class (singular) to config_classes (plural) to match the convention used by BaseGrafanaToolset, ensuring GrafanaDashboardConfig is actually used during initialization https://claude.ai/code/session_01BDFrttpmMrSWDNcnJGwBRi Signed-off-by: Claude <noreply@anthropic.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (2)
WalkthroughChanges introduce dynamic config class selection in two toolsets. Elasticsearch now selects from a list of config classes instead of hard-coding ElasticsearchConfig. Grafana replaces a single config_class attribute with a config_classes list type to support alternate config types. Changes
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~5 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:49b88036
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:49b88036 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:49b88036
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:49b88036
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:49b88036
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:49b88036 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:49b88036
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:49b88036Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:49b88036 \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:49b88036Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:49b88036 \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:49b88036 |
📂 Previous Runs📜 Run @ d9578e4 (#23119147401)✅ Results of HolmesGPT evalsAutomatically triggered by commit d9578e4 on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 73 test/model combinations loaded Benchmark experiment:
No benchmark data available for comparison. Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
|
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 25.8s | 5 | 10 | $0.2311 | 103,988 | 102,206 | 23,186 | 1,782 | 572 | 78,546 | 23,660 | — | — |
| ✅ | 101_loki_historical_logs_pod_deleted | 31.5s | 5 | 9 | $0.2341 | 103,109 | 100,978 | 22,293 | 2,131 | 849 | 78,217 | 22,761 | — | — |
| ✅ | 111_pod_names_contain_service | 27.1s | 5 | 9 | $0.2233 | 100,713 | 98,905 | 22,007 | 1,808 | 615 | 76,434 | 22,471 | — | — |
| ✅ | 112_find_pvcs_by_uuid | 21.4s | 5 | 6 | $0.2026 | 95,950 | 94,553 | 20,631 | 1,397 | 479 | 73,523 | 21,030 | — | — |
| ✅ | 12_job_crashing | 35.5s | 8 | 16 | $0.2931 | 178,426 | 176,242 | 25,617 | 2,184 | 402 | 149,951 | 26,291 | — | — |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 41.3s | 7 | 18 | $0.3176 | 157,648 | 154,930 | 27,061 | 2,718 | 710 | 124,185 | 30,745 | — | — |
| ✅ | 183a_elasticsearch_cluster_health | 10.7s | 3 | 3 | $0.1554 | 59,411 | 58,992 | 20,069 | 419 | 175 | 38,912 | 20,080 | — | — |
| ✅ | 183b_elasticsearch_index_discovery | 10.6s | 3 | 3 | $0.1576 | 59,716 | 59,258 | 20,249 | 458 | 270 | 38,998 | 20,260 | — | — |
| ✅ | 183c_elasticsearch_log_search | 24.1s | 6 | 8 | $0.2229 | 125,932 | 124,677 | 22,458 | 1,255 | 330 | 102,205 | 22,472 | — | — |
| ✅ | 183d_elasticsearch_aggregation | 26.0s | 8 | 9 | $0.2403 | 166,524 | 165,326 | 22,192 | 1,198 | 308 | 143,118 | 22,208 | — | — |
| ✅ | 183e_elasticsearch_field_mappings | 12.2s | 3 | 3 | $0.1613 | 59,868 | 59,259 | 20,248 | 609 | 266 | 39,000 | 20,259 | — | — |
| ✅ | 183f_elasticsearch_shard_filtering | 12.4s | 3 | 3 | $0.1629 | 60,106 | 59,462 | 20,348 | 644 | 292 | 39,103 | 20,359 | — | — |
| ✅ | 183g_elasticsearch_index_stats | 12.4s | 3 | 3 | $0.1676 | 61,697 | 61,090 | 21,181 | 607 | 278 | 39,898 | 21,192 | — | — |
| ✅ | 184_elasticsearch_index_explosion | 14.4s | 3 | 3 | $0.3182 | 106,209 | 105,582 | 43,431 | 627 | 274 | 62,140 | 43,442 | — | — |
| ✅ | 185_elasticsearch_cross_region_search | 34.0s | 6 | 8 | $0.5046 | 281,540 | 279,871 | 56,158 | 1,669 | 402 | 223,699 | 56,172 | — | — |
| ✅ | 186_elasticsearch_shard_explosion | 14.6s | 3 | 3 | $0.1724 | 61,163 | 60,251 | 20,773 | 912 | 392 | 39,467 | 20,784 | — | — |
| ✅ | 187_elasticsearch_disk_space | 13.1s | 3 | 3 | $0.1660 | 61,004 | 60,357 | 20,793 | 647 | 273 | 39,553 | 20,804 | — | — |
| ❌ | 188_elasticsearch_mapping_explosion | 28.5s | 5 | 11 | $0.2554 | 114,985 | 113,103 | 26,384 | 1,882 | 517 | 86,706 | 26,397 | — | — |
| ✅ | 189_elasticsearch_timeseries_gap | 30.1s | 6 | 11 | $0.2841 | 142,689 | 140,508 | 27,136 | 2,181 | 497 | 112,650 | 27,858 | — | — |
| ✅ | 190_elasticsearch_cross_service_correlation | 17.1s | 3 | 4 | $0.1926 | 63,721 | 62,516 | 21,978 | 1,205 | 517 | 39,462 | 23,054 | — | — |
| ✅ | 191_elasticsearch_query_profile | 15.6s | 3 | 3 | $0.1742 | 62,136 | 61,295 | 21,293 | 841 | 393 | 39,991 | 21,304 | — | — |
| ✅ | 193_elasticsearch_large_mapping_search | 11.8s | 3 | 3 | $0.1592 | 59,996 | 59,512 | 20,405 | 484 | 303 | 39,096 | 20,416 | — | — |
| ✅ | 195_elasticsearch_trace_large_fields | 40.2s | 8 | 11 | $0.5661 | 307,292 | 305,369 | 63,293 | 1,923 | 413 | 241,768 | 63,601 | — | — |
| ✅ | 227_count_configmaps_per_namespace[0] | 26.6s | 7 | 11 | $0.2221 | 131,792 | 130,281 | 20,718 | 1,511 | 430 | 109,548 | 20,733 | — | — |
| ✅ | 24_misconfigured_pvc | 26.1s | 5 | 11 | $0.2333 | 101,419 | 99,475 | 22,319 | 1,944 | 551 | 75,681 | 23,794 | — | — |
| ✅ | 43_current_datetime_from_prompt | 6.4s | 1 | — | $0.1139 | 17,196 | 16,853 | 16,853 | 343 | 343 | 0 | 16,853 | — | — |
| ✅ | 61_exact_match_counting | 13.0s | 4 | 4 | $0.1509 | 70,869 | 70,416 | 18,140 | 453 | 236 | 52,264 | 18,152 | — | — |
| Total | 21.6s avg | 4.6 avg | 7.2 avg | $6.2826 | 2,915,099 | 2,881,267 | 63,293 | 33,832 | 849 | 2,184,115 | 697,152 | — | — |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 73 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 1 Failure Detected
📜 Run @ e4ebff3 (#23114099380)
✅ Results of HolmesGPT evals
Automatically triggered by commit e4ebff3 on branch claude/review-toolset-fixes-X2pPc (labels: evals-tag-elasticsearch)
Results of HolmesGPT evals
- ask_holmes: 27/27 test cases were successful, 0 regressions
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 21.3s | 4 | 8 | $0.2065 | 79,997 | 78,470 | 21,832 | 1,527 | 614 | 55,845 | 22,625 | — | — |
| ✅ | 101_loki_historical_logs_pod_deleted | 40.1s | 7 | 11 | $0.2687 | 144,987 | 142,584 | 23,378 | 2,403 | 429 | 118,562 | 24,022 | — | — |
| ✅ | 111_pod_names_contain_service | 23.7s | 5 | 8 | $0.3326 | 98,538 | 96,981 | 21,450 | 1,557 | 384 | 54,249 | 42,732 | — | — |
| ✅ | 112_find_pvcs_by_uuid | 18.3s | 5 | 4 | $0.2964 | 96,440 | 95,377 | 21,728 | 1,063 | 283 | 56,748 | 38,629 | — | — |
| ✅ | 12_job_crashing | 27.9s | 6 | 14 | $0.2634 | 128,987 | 127,125 | 24,963 | 1,862 | 498 | 99,996 | 27,129 | — | — |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 37.9s | 7 | 15 | $0.2920 | 154,520 | 152,074 | 26,207 | 2,446 | 676 | 124,943 | 27,131 | — | — |
| ✅ | 183a_elasticsearch_cluster_health | 11.0s | 3 | 3 | $0.1557 | 59,423 | 58,992 | 20,069 | 431 | 187 | 38,912 | 20,080 | — | — |
| ✅ | 183b_elasticsearch_index_discovery | 10.4s | 3 | 3 | $0.2717 | 59,698 | 59,241 | 20,237 | 457 | 270 | 19,132 | 40,109 | — | — |
| ✅ | 183c_elasticsearch_log_search | 22.6s | 6 | 6 | $0.2141 | 124,057 | 122,971 | 21,804 | 1,086 | 274 | 101,153 | 21,818 | — | — |
| ✅ | 183d_elasticsearch_aggregation | 24.6s | 7 | 9 | $0.2312 | 146,019 | 144,831 | 22,296 | 1,188 | 278 | 122,348 | 22,483 | — | — |
| ✅ | 183e_elasticsearch_field_mappings | 12.0s | 3 | 3 | $0.1608 | 59,820 | 59,227 | 20,228 | 593 | 263 | 38,988 | 20,239 | — | — |
| ✅ | 183f_elasticsearch_shard_filtering | 13.1s | 3 | 3 | $0.1650 | 60,226 | 59,502 | 20,367 | 724 | 357 | 39,124 | 20,378 | — | — |
| ✅ | 183g_elasticsearch_index_stats | 12.3s | 3 | 3 | $0.1684 | 61,780 | 61,151 | 21,225 | 629 | 288 | 39,915 | 21,236 | — | — |
| ✅ | 184_elasticsearch_index_explosion | 14.1s | 3 | 3 | $0.3178 | 106,179 | 105,568 | 43,428 | 611 | 273 | 62,129 | 43,439 | — | — |
| ✅ | 185_elasticsearch_cross_region_search | 39.9s | 7 | 10 | $0.5058 | 313,547 | 311,392 | 51,511 | 2,155 | 516 | 259,866 | 51,526 | — | — |
| ✅ | 186_elasticsearch_shard_explosion | 14.4s | 3 | 3 | $0.1710 | 60,997 | 60,124 | 20,710 | 873 | 417 | 39,403 | 20,721 | — | — |
| ✅ | 187_elasticsearch_disk_space | 13.7s | 3 | 3 | $0.1664 | 61,097 | 60,450 | 20,858 | 647 | 280 | 39,581 | 20,869 | — | — |
| ✅ | 188_elasticsearch_mapping_explosion | 25.9s | 4 | 8 | $0.3000 | 100,933 | 99,430 | 30,299 | 1,503 | 537 | 60,567 | 38,863 | — | — |
| ✅ | 189_elasticsearch_timeseries_gap | 26.9s | 5 | 9 | $0.2594 | 118,032 | 115,951 | 25,694 | 2,081 | 507 | 89,896 | 26,055 | — | — |
| ✅ | 190_elasticsearch_cross_service_correlation | 22.4s | 4 | 7 | $0.2232 | 87,846 | 86,095 | 23,551 | 1,751 | 540 | 62,326 | 23,769 | — | — |
| ✅ | 191_elasticsearch_query_profile | 15.5s | 3 | 3 | $0.1741 | 62,139 | 61,301 | 21,294 | 838 | 388 | 39,996 | 21,305 | — | — |
| ✅ | 193_elasticsearch_large_mapping_search | 11.5s | 3 | 3 | $0.1574 | 59,746 | 59,299 | 20,264 | 447 | 274 | 39,024 | 20,275 | — | — |
| ✅ | 195_elasticsearch_trace_large_fields | 34.2s | 7 | 9 | $0.4052 | 241,946 | 240,167 | 41,548 | 1,779 | 340 | 198,228 | 41,939 | — | — |
| ✅ | 227_count_configmaps_per_namespace[0] | 23.4s | 6 | 11 | $0.2132 | 113,821 | 112,388 | 20,576 | 1,433 | 434 | 91,184 | 21,204 | — | — |
| ✅ | 24_misconfigured_pvc | 29.4s | 6 | 14 | $0.2501 | 122,292 | 120,209 | 23,231 | 2,083 | 702 | 96,021 | 24,188 | — | — |
| ✅ | 43_current_datetime_from_prompt | 3.8s | 1 | — | $0.1081 | 16,964 | 16,853 | 16,853 | 111 | 111 | 0 | 16,853 | — | — |
| ✅ | 61_exact_match_counting | 12.2s | 4 | 4 | $0.1507 | 70,837 | 70,391 | 18,135 | 446 | 227 | 52,244 | 18,147 | — | — |
| Total | 20.8s avg | 4.5 avg | 6.8 avg | $6.4288 | 2,810,868 | 2,778,144 | 51,511 | 32,724 | 702 | 2,040,380 | 737,764 | — | — |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 73 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
📜 Run @ e4ebff3 (#23114091570)
✅ Results of HolmesGPT evals
Automatically triggered by commit e4ebff3 on branch claude/review-toolset-fixes-X2pPc
Results of HolmesGPT evals
- ask_holmes: 10/10 test cases were successful, 0 regressions
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 30.6s | 6 | 11 | $0.2528 | 121,197 | 119,064 | 23,544 | 2,133 | 631 | 94,516 | 24,548 | — | — |
| ✅ | 101_loki_historical_logs_pod_deleted | 34.6s | 5 | 10 | $0.2550 | 107,560 | 105,200 | 24,197 | 2,360 | 800 | 80,062 | 25,138 | — | — |
| ✅ | 111_pod_names_contain_service | 26.6s | 5 | 8 | $0.2118 | 95,871 | 94,275 | 21,337 | 1,596 | 612 | 72,472 | 21,803 | — | — |
| ✅ | 112_find_pvcs_by_uuid | 24.1s | 6 | 5 | $0.2056 | 112,760 | 111,422 | 20,248 | 1,338 | 286 | 91,160 | 20,262 | — | — |
| ✅ | 12_job_crashing | 27.9s | 5 | 11 | $0.2460 | 111,312 | 109,593 | 25,022 | 1,719 | 434 | 83,612 | 25,981 | — | — |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 34.7s | 6 | 13 | $0.2769 | 129,882 | 127,364 | 25,690 | 2,518 | 710 | 101,108 | 26,256 | — | — |
| ✅ | 227_count_configmaps_per_namespace[0] | 22.6s | 6 | 10 | $0.2082 | 112,584 | 111,172 | 20,399 | 1,412 | 435 | 90,759 | 20,413 | — | — |
| ✅ | 24_misconfigured_pvc | 32.8s | 6 | 16 | $0.3836 | 124,239 | 121,915 | 23,927 | 2,324 | 544 | 75,684 | 46,231 | — | — |
| ✅ | 43_current_datetime_from_prompt | 4.4s | 1 | — | $0.1084 | 16,977 | 16,853 | 16,853 | 124 | 124 | 0 | 16,853 | — | — |
| ✅ | 61_exact_match_counting | 9.6s | 3 | 2 | $0.1334 | 51,849 | 51,566 | 17,475 | 283 | 163 | 34,080 | 17,486 | — | — |
| Total | 24.8s avg | 4.9 avg | 9.6 avg | $2.2817 | 984,231 | 968,424 | 25,690 | 15,807 | 800 | 723,453 | 244,971 | — | — |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 73 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
📜 Run @ 98a7112 (#22950260085)
✅ Results of HolmesGPT evals
Automatically triggered by commit 98a7112 on branch claude/review-toolset-fixes-X2pPc
Results of HolmesGPT evals
- ask_holmes: 10/10 test cases were successful, 0 regressions
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Output | Cached | Non-cached | Reasoning | Max output | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 30.7s | 5 | 10 | $0.2326 | 105,284 | 103,477 | 1,807 | 79,829 | 23,648 | — | 450 | — |
| ✅ | 101_loki_historical_logs_pod_deleted | 31.4s | 4 | 8 | $0.2161 | 82,707 | 80,844 | 1,863 | 58,346 | 22,498 | — | 744 | — |
| ✅ | 111_pod_names_contain_service | 27.2s | 4 | 8 | $0.2069 | 80,564 | 79,022 | 1,542 | 56,440 | 22,582 | — | 505 | — |
| ✅ | 112_find_pvcs_by_uuid | 21.8s | 4 | 4 | $0.1905 | 79,212 | 78,032 | 1,180 | 56,819 | 21,213 | — | 577 | — |
| ✅ | 12_job_crashing | 32.7s | 5 | 12 | $0.2502 | 111,987 | 110,118 | 1,869 | 84,023 | 26,095 | — | 489 | — |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 42.4s | 7 | 14 | $0.3046 | 161,864 | 159,339 | 2,525 | 130,967 | 28,372 | — | 522 | — |
| ✅ | 227_count_configmaps_per_namespace[0] | 26.5s | 6 | 11 | $0.2202 | 117,382 | 115,896 | 1,486 | 93,990 | 21,906 | — | 429 | — |
| ✅ | 24_misconfigured_pvc | 30.8s | 5 | 13 | $0.3514 | 104,629 | 102,879 | 1,750 | 57,920 | 44,959 | — | 653 | — |
| ✅ | 43_current_datetime_from_prompt | 4.9s | 1 | — | $0.1115 | 17,481 | 17,359 | 122 | 0 | 17,359 | — | 122 | — |
| ✅ | 61_exact_match_counting | 11.3s | 3 | 2 | $0.1388 | 53,533 | 53,201 | 332 | 35,130 | 18,071 | — | 193 | — |
| Total | 26.0s avg | 4.4 avg | 9.1 avg | $2.2227 | 914,643 | 900,167 | 14,476 | 653,464 | 246,703 | — | 744 | — |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 16 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-22947901310 (created: 2026-03-11)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ Eval Results (with failures)
Automatically triggered by commit fc8d9da on branch claude/review-toolset-fixes-X2pPc (labels: evals-tag-elasticsearch)
Results of HolmesGPT evals
- ask_holmes: 25/27 test cases were successful, 2 regressions
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 28.6s | 5 | 10 | $0.2335 | 103,999 | 102,136 | 23,239 | 1,863 | 641 | 78,408 | 23,728 | — | — |
| ✅ | 101_loki_historical_logs_pod_deleted | 44.9s | 6 | 13 | $0.2821 | 133,304 | 130,722 | 26,068 | 2,582 | 726 | 104,154 | 26,568 | — | — |
| ✅ | 111_pod_names_contain_service | 25.7s | 4 | 8 | $0.2041 | 78,626 | 77,039 | 21,510 | 1,587 | 681 | 55,040 | 21,999 | — | — |
| ✅ | 112_find_pvcs_by_uuid | 22.5s | 5 | 4 | $0.1951 | 93,609 | 92,249 | 19,991 | 1,360 | 342 | 72,245 | 20,004 | — | — |
| ✅ | 12_job_crashing | 37.5s | 6 | 17 | $0.2953 | 139,949 | 137,377 | 26,535 | 2,572 | 640 | 108,683 | 28,694 | — | — |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 45.4s | 8 | 15 | $0.3072 | 175,893 | 173,233 | 26,482 | 2,660 | 575 | 146,357 | 26,876 | — | — |
| ✅ | 183a_elasticsearch_cluster_health | 11.1s | 3 | 3 | $0.1556 | 59,425 | 59,001 | 20,073 | 424 | 174 | 38,917 | 20,084 | — | — |
| ✅ | 183b_elasticsearch_index_discovery | 10.8s | 3 | 3 | $0.1573 | 59,687 | 59,235 | 20,234 | 452 | 266 | 38,990 | 20,245 | — | — |
| ✅ | 183c_elasticsearch_log_search | 23.7s | 6 | 6 | $0.2127 | 123,921 | 122,879 | 21,773 | 1,042 | 267 | 101,092 | 21,787 | — | — |
| ✅ | 183d_elasticsearch_aggregation | 29.5s | 8 | 10 | $0.3536 | 167,182 | 165,927 | 22,469 | 1,255 | 311 | 124,307 | 41,620 | — | — |
| ✅ | 183e_elasticsearch_field_mappings | 13.6s | 3 | 3 | $0.1612 | 59,861 | 59,259 | 20,249 | 602 | 260 | 38,999 | 20,260 | — | — |
| ✅ | 183f_elasticsearch_shard_filtering | 13.4s | 3 | 3 | $0.1638 | 60,159 | 59,482 | 20,357 | 677 | 320 | 39,114 | 20,368 | — | — |
| ✅ | 183g_elasticsearch_index_stats | 12.7s | 3 | 3 | $0.1674 | 61,687 | 61,085 | 21,177 | 602 | 275 | 39,897 | 21,188 | — | — |
| ✅ | 184_elasticsearch_index_explosion | 15.5s | 3 | 3 | $0.3180 | 106,169 | 105,548 | 43,410 | 621 | 271 | 62,127 | 43,421 | — | — |
| ✅ | 185_elasticsearch_cross_region_search | 39.5s | 6 | 10 | $0.5219 | 275,945 | 273,999 | 57,361 | 1,946 | 397 | 215,200 | 58,799 | — | — |
| ✅ | 186_elasticsearch_shard_explosion | 20.2s | 4 | 5 | $0.1912 | 82,982 | 81,916 | 21,478 | 1,066 | 435 | 60,426 | 21,490 | — | — |
| ✅ | 187_elasticsearch_disk_space | 14.7s | 3 | 3 | $0.1645 | 60,919 | 60,323 | 20,769 | 596 | 267 | 39,543 | 20,780 | — | — |
| ❌ | 188_elasticsearch_mapping_explosion | 26.1s | 4 | 9 | $0.2085 | 84,720 | 83,203 | 22,421 | 1,517 | 517 | 60,770 | 22,433 | — | — |
| ✅ | 189_elasticsearch_timeseries_gap | 38.2s | 7 | 11 | $0.2966 | 164,402 | 161,991 | 26,699 | 2,411 | 536 | 134,899 | 27,092 | — | — |
| ✅ | 190_elasticsearch_cross_service_correlation | 21.8s | 4 | 5 | $0.2066 | 85,392 | 84,016 | 22,641 | 1,376 | 460 | 61,363 | 22,653 | — | — |
| ✅ | 191_elasticsearch_query_profile | 16.5s | 3 | 3 | $0.1723 | 62,042 | 61,272 | 21,278 | 770 | 331 | 39,983 | 21,289 | — | — |
| ✅ | 193_elasticsearch_large_mapping_search | 13.9s | 3 | 3 | $0.1602 | 59,954 | 59,410 | 20,327 | 544 | 286 | 39,072 | 20,338 | — | — |
| ❌ | 195_elasticsearch_trace_large_fields | 38.4s | 7 | 10 | $0.4581 | 242,116 | 240,414 | 51,108 | 1,702 | 391 | 188,978 | 51,436 | — | — |
| ✅ | 227_count_configmaps_per_namespace[0] | 25.3s | 6 | 10 | $0.2089 | 112,126 | 110,723 | 20,462 | 1,403 | 617 | 90,077 | 20,646 | — | — |
| ✅ | 24_misconfigured_pvc | 33.0s | 6 | 14 | $0.2590 | 123,133 | 120,943 | 23,356 | 2,190 | 695 | 95,472 | 25,471 | — | — |
| ✅ | 43_current_datetime_from_prompt | 4.5s | 1 | — | $0.1083 | 16,971 | 16,853 | 16,853 | 118 | 118 | 0 | 16,853 | — | — |
| ✅ | 61_exact_match_counting | 13.8s | 4 | 4 | $0.1506 | 70,819 | 70,375 | 18,127 | 444 | 227 | 52,236 | 18,139 | — | — |
| Total | 23.7s avg | 4.6 avg | 7.2 avg | $6.3135 | 2,864,992 | 2,830,610 | 57,361 | 34,382 | 726 | 2,126,349 | 704,261 | — | — |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 73 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 2 Failures Detected
📖 Legend
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/review-toolset-fixes-X2pPc -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
tags: regression
Or with more options (one per line):
/eval
model: gpt-4o
tags: regression
filter: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
tags: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
tags |
Pytest tags / markers (no default - runs all tests!) |
filter |
Pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
Option 3: Add PR labels to include extra evals in automatic regression runs:
| Label | Effect |
|---|---|
evals-tag-<name> |
Run tests with tag <name> alongside regression |
evals-id-<name> |
Run a specific eval by test ID |
evals-model-<name> |
Override the model (use model list name, e.g. sonnet-4.5) |
Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5
🏷️ Valid tags
benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency
🤖 Valid models
deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/review-toolset-fixes-X2pPc -f markers=regression -f filter=
🔬 CLI Performance Benchmark🟡 Startup Time (no LLM)Measures
🟡 Full CLI with LLMMeasures
PR: |
…lmesGPT#1741) ## Summary This PR refactors the configuration class handling in toolsets to support multiple configuration classes instead of a single hardcoded class. This change enables greater flexibility in toolset configuration management. ## Key Changes - **Elasticsearch toolset**: Modified `prerequisites_callable` to dynamically select the configuration class from `self.config_classes` list, falling back to `ElasticsearchConfig` if the list is empty - **Grafana toolset**: Changed `config_class` attribute from a single `ClassVar` to `config_classes` as a list of configuration classes, updating the type annotation from `Type[GrafanaDashboardConfig]` to `list[Type[GrafanaDashboardConfig]]` ## Implementation Details - The Elasticsearch implementation now checks if `config_classes` is populated and uses the first class in the list, maintaining backward compatibility with a fallback to the default `ElasticsearchConfig` - This pattern allows toolsets to be extended with multiple configuration options while maintaining existing behavior - The change is backward compatible as long as subclasses properly initialize the `config_classes` list https://claude.ai/code/session_01BDFrttpmMrSWDNcnJGwBRi <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Enhanced Elasticsearch plugin configuration to dynamically select from available configuration classes, improving flexibility. * Updated Grafana toolset to support multiple configuration types, enabling more versatile configuration handling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Claude <noreply@anthropic.com> Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Natan Yellin <aantn@users.noreply.github.com>
Summary
This PR refactors the configuration class handling in toolsets to support multiple configuration classes instead of a single hardcoded class. This change enables greater flexibility in toolset configuration management.
Key Changes
prerequisites_callableto dynamically select the configuration class fromself.config_classeslist, falling back toElasticsearchConfigif the list is emptyconfig_classattribute from a singleClassVartoconfig_classesas a list of configuration classes, updating the type annotation fromType[GrafanaDashboardConfig]tolist[Type[GrafanaDashboardConfig]]Implementation Details
config_classesis populated and uses the first class in the list, maintaining backward compatibility with a fallback to the defaultElasticsearchConfigconfig_classeslisthttps://claude.ai/code/session_01BDFrttpmMrSWDNcnJGwBRi
Summary by CodeRabbit