Skip to content

Refactor config class handling to support multiple config classes - #1741

Merged
naomi-robusta merged 6 commits into
masterfrom
claude/review-toolset-fixes-X2pPc
Mar 16, 2026
Merged

naomi-robusta merged 6 commits into
masterfrom
claude/review-toolset-fixes-X2pPc

Conversation

@naomi-robusta

@naomi-robusta naomi-robusta commented Mar 11, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR refactors the configuration class handling in toolsets to support multiple configuration classes instead of a single hardcoded class. This change enables greater flexibility in toolset configuration management.

Key Changes

  • Elasticsearch toolset: Modified prerequisites_callable to dynamically select the configuration class from self.config_classes list, falling back to ElasticsearchConfig if the list is empty
  • Grafana toolset: Changed config_class attribute from a single ClassVar to config_classes as a list of configuration classes, updating the type annotation from Type[GrafanaDashboardConfig] to list[Type[GrafanaDashboardConfig]]

Implementation Details

  • The Elasticsearch implementation now checks if config_classes is populated and uses the first class in the list, maintaining backward compatibility with a fallback to the default ElasticsearchConfig
  • This pattern allows toolsets to be extended with multiple configuration options while maintaining existing behavior
  • The change is backward compatible as long as subclasses properly initialize the config_classes list

https://claude.ai/code/session_01BDFrttpmMrSWDNcnJGwBRi

Summary by CodeRabbit

  • Refactor
    • Enhanced Elasticsearch plugin configuration to dynamically select from available configuration classes, improving flexibility.
    • Updated Grafana toolset to support multiple configuration types, enabling more versatile configuration handling.

…ashboard config_classes

Apply the same fix from d5fd607 (Grafana Tempo) to other toolsets:

- ElasticsearchBaseToolset: use self.config_classes[0] instead of
  hardcoding ElasticsearchConfig in prerequisites_callable, so subclasses
  with custom config classes work correctly
- GrafanaToolset: change config_class (singular) to config_classes (plural)
  to match the convention used by BaseGrafanaToolset, ensuring
  GrafanaDashboardConfig is actually used during initialization

https://claude.ai/code/session_01BDFrttpmMrSWDNcnJGwBRi
Signed-off-by: Claude <noreply@anthropic.com>
@naomi-robusta
naomi-robusta requested a review from aantn March 11, 2026 11:10
@naomi-robusta
naomi-robusta enabled auto-merge (squash) March 11, 2026 11:10
@coderabbitai

coderabbitai Bot commented Mar 11, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 39fe0d27-90a1-4fab-a2bb-7814fc8cbfa1

📥 Commits

Reviewing files that changed from the base of the PR and between f0aa88c and e637666.

📒 Files selected for processing (2)
  • holmes/plugins/toolsets/elasticsearch/elasticsearch.py
  • holmes/plugins/toolsets/grafana/toolset_grafana.py

Walkthrough

Changes introduce dynamic config class selection in two toolsets. Elasticsearch now selects from a list of config classes instead of hard-coding ElasticsearchConfig. Grafana replaces a single config_class attribute with a config_classes list type to support alternate config types.

Changes

Cohort / File(s) Summary
Config class selection pattern
holmes/plugins/toolsets/elasticsearch/elasticsearch.py, holmes/plugins/toolsets/grafana/toolset_grafana.py
Elasticsearch implements dynamic config class selection via self.config_classes[0] with fallback to default ElasticsearchConfig; Grafana updates type annotation from Type[GrafanaDashboardConfig] to list[Type[GrafanaDashboardConfig]] and initializes with a list.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~5 minutes

Possibly related PRs

Suggested reviewers

  • aantn
🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main refactoring effort to support multiple config classes across Elasticsearch and Grafana toolsets.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Mar 11, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit fc8d9da
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69b7be522ee74a00086db4e0
😎 Deploy Preview https://deploy-preview-1741--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Mar 11, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 49b88036 (built in 14m 9s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:49b88036
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:49b88036 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:49b88036
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:49b88036
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:49b88036
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:49b88036 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:49b88036
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:49b88036

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:49b88036 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:49b88036

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:49b88036 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:49b88036

@github-actions

github-actions Bot commented Mar 11, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ d9578e4 (#23119147401)

✅ Results of HolmesGPT evals

Automatically triggered by commit d9578e4 on branch claude/review-toolset-fixes-X2pPc (labels: evals-tag-elasticsearch)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 26/27 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 28.0s 6 11 $0.2443 122,471 120,585 23,282 1,886 768 96,648 23,937 — —
✅ 101_loki_historical_logs_pod_deleted 43.2s 7 13 $0.2982 152,073 149,227 25,799 2,846 586 122,518 26,709 — —
✅ 111_pod_names_contain_service 23.7s 5 9 $0.2156 99,388 97,804 21,738 1,584 473 75,596 22,208 — —
✅ 112_find_pvcs_by_uuid 19.0s 5 4 $0.1961 94,819 93,611 20,703 1,208 323 72,895 20,716 — —
✅ 12_job_crashing 26.1s 5 11 $0.2345 105,529 103,789 23,411 1,740 460 79,376 24,413 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 35.8s 7 15 $0.2933 152,608 150,236 25,713 2,372 454 122,080 28,156 — —
✅ 183a_elasticsearch_cluster_health 10.6s 3 3 $0.1554 59,407 58,988 20,067 419 180 38,910 20,078 — —
✅ 183b_elasticsearch_index_discovery 10.3s 3 3 $0.1574 59,689 59,232 20,232 457 270 38,989 20,243 — —
✅ 183c_elasticsearch_log_search 23.5s 6 7 $0.3478 125,545 124,299 22,267 1,246 331 79,969 44,330 — —
✅ 183d_elasticsearch_aggregation 26.3s 8 10 $0.2463 167,696 166,405 22,541 1,291 336 143,602 22,803 — —
✅ 183e_elasticsearch_field_mappings 11.7s 3 3 $0.1609 59,836 59,243 20,241 593 260 38,991 20,252 — —
✅ 183f_elasticsearch_shard_filtering 12.5s 3 3 $0.1652 60,231 59,499 20,366 732 370 39,122 20,377 — —
✅ 183g_elasticsearch_index_stats 11.4s 3 3 $0.1674 61,615 60,999 21,126 616 261 39,862 21,137 — —
✅ 184_elasticsearch_index_explosion 13.5s 3 3 $0.3159 105,929 105,355 43,277 574 245 62,067 43,288 — —
✅ 185_elasticsearch_cross_region_search 58.4s 9 16 $1.1752 518,254 515,238 73,424 3,016 517 367,836 147,402 — —
✅ 186_elasticsearch_shard_explosion 13.5s 3 3 $0.1708 60,944 60,072 20,676 872 427 39,385 20,687 — —
✅ 187_elasticsearch_disk_space 13.2s 3 3 $0.1659 60,963 60,310 20,769 653 274 39,530 20,780 — —
✅ 188_elasticsearch_mapping_explosion 35.7s 7 14 $0.2948 164,572 162,288 26,824 2,284 539 134,948 27,340 — —
✅ 189_elasticsearch_timeseries_gap 31.6s 5 10 $0.3211 140,980 138,777 32,608 2,203 560 104,148 34,629 — —
✅ 190_elasticsearch_cross_service_correlation 20.0s 4 6 $0.2094 85,931 84,477 22,738 1,454 467 61,727 22,750 — —
✅ 191_elasticsearch_query_profile 15.7s 3 3 $0.1746 62,160 61,300 21,294 860 406 39,995 21,305 — —
✅ 193_elasticsearch_large_mapping_search 11.2s 3 3 $0.1574 59,753 59,308 20,269 445 275 39,028 20,280 — —
❌ 195_elasticsearch_trace_large_fields 35.6s 7 9 $0.4097 242,707 240,807 42,040 1,900 389 198,752 42,055 — —
✅ 227_count_configmaps_per_namespace[0] 25.4s 7 11 $0.2232 131,752 130,251 20,720 1,501 434 109,207 21,044 — —
✅ 24_misconfigured_pvc 26.9s 5 13 $0.2305 101,066 99,146 22,188 1,920 633 75,779 23,367 — —
✅ 43_current_datetime_from_prompt 4.8s 1 — $0.1096 17,024 16,853 16,853 171 171 0 16,853 — —
✅ 61_exact_match_counting 12.8s 4 4 $0.2534 70,858 70,409 18,140 449 232 34,412 35,997 — —
Total 22.2s avg 4.7 avg 7.4 avg $7.0939 3,143,800 3,108,508 73,424 35,292 768 2,295,372 813,136 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected

📜 Run @ db27844 (#23116657245)

✅ Results of HolmesGPT evals

Automatically triggered by commit db27844 on branch claude/review-toolset-fixes-X2pPc (labels: evals-tag-elasticsearch)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 26/27 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 25.8s 5 10 $0.2311 103,988 102,206 23,186 1,782 572 78,546 23,660 — —
✅ 101_loki_historical_logs_pod_deleted 31.5s 5 9 $0.2341 103,109 100,978 22,293 2,131 849 78,217 22,761 — —
✅ 111_pod_names_contain_service 27.1s 5 9 $0.2233 100,713 98,905 22,007 1,808 615 76,434 22,471 — —
✅ 112_find_pvcs_by_uuid 21.4s 5 6 $0.2026 95,950 94,553 20,631 1,397 479 73,523 21,030 — —
✅ 12_job_crashing 35.5s 8 16 $0.2931 178,426 176,242 25,617 2,184 402 149,951 26,291 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 41.3s 7 18 $0.3176 157,648 154,930 27,061 2,718 710 124,185 30,745 — —
✅ 183a_elasticsearch_cluster_health 10.7s 3 3 $0.1554 59,411 58,992 20,069 419 175 38,912 20,080 — —
✅ 183b_elasticsearch_index_discovery 10.6s 3 3 $0.1576 59,716 59,258 20,249 458 270 38,998 20,260 — —
✅ 183c_elasticsearch_log_search 24.1s 6 8 $0.2229 125,932 124,677 22,458 1,255 330 102,205 22,472 — —
✅ 183d_elasticsearch_aggregation 26.0s 8 9 $0.2403 166,524 165,326 22,192 1,198 308 143,118 22,208 — —
✅ 183e_elasticsearch_field_mappings 12.2s 3 3 $0.1613 59,868 59,259 20,248 609 266 39,000 20,259 — —
✅ 183f_elasticsearch_shard_filtering 12.4s 3 3 $0.1629 60,106 59,462 20,348 644 292 39,103 20,359 — —
✅ 183g_elasticsearch_index_stats 12.4s 3 3 $0.1676 61,697 61,090 21,181 607 278 39,898 21,192 — —
✅ 184_elasticsearch_index_explosion 14.4s 3 3 $0.3182 106,209 105,582 43,431 627 274 62,140 43,442 — —
✅ 185_elasticsearch_cross_region_search 34.0s 6 8 $0.5046 281,540 279,871 56,158 1,669 402 223,699 56,172 — —
✅ 186_elasticsearch_shard_explosion 14.6s 3 3 $0.1724 61,163 60,251 20,773 912 392 39,467 20,784 — —
✅ 187_elasticsearch_disk_space 13.1s 3 3 $0.1660 61,004 60,357 20,793 647 273 39,553 20,804 — —
❌ 188_elasticsearch_mapping_explosion 28.5s 5 11 $0.2554 114,985 113,103 26,384 1,882 517 86,706 26,397 — —
✅ 189_elasticsearch_timeseries_gap 30.1s 6 11 $0.2841 142,689 140,508 27,136 2,181 497 112,650 27,858 — —
✅ 190_elasticsearch_cross_service_correlation 17.1s 3 4 $0.1926 63,721 62,516 21,978 1,205 517 39,462 23,054 — —
✅ 191_elasticsearch_query_profile 15.6s 3 3 $0.1742 62,136 61,295 21,293 841 393 39,991 21,304 — —
✅ 193_elasticsearch_large_mapping_search 11.8s 3 3 $0.1592 59,996 59,512 20,405 484 303 39,096 20,416 — —
✅ 195_elasticsearch_trace_large_fields 40.2s 8 11 $0.5661 307,292 305,369 63,293 1,923 413 241,768 63,601 — —
✅ 227_count_configmaps_per_namespace[0] 26.6s 7 11 $0.2221 131,792 130,281 20,718 1,511 430 109,548 20,733 — —
✅ 24_misconfigured_pvc 26.1s 5 11 $0.2333 101,419 99,475 22,319 1,944 551 75,681 23,794 — —
✅ 43_current_datetime_from_prompt 6.4s 1 — $0.1139 17,196 16,853 16,853 343 343 0 16,853 — —
✅ 61_exact_match_counting 13.0s 4 4 $0.1509 70,869 70,416 18,140 453 236 52,264 18,152 — —
Total 21.6s avg 4.6 avg 7.2 avg $6.2826 2,915,099 2,881,267 63,293 33,832 849 2,184,115 697,152 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected

📜 Run @ e4ebff3 (#23114099380)

✅ Results of HolmesGPT evals

Automatically triggered by commit e4ebff3 on branch claude/review-toolset-fixes-X2pPc (labels: evals-tag-elasticsearch)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 27/27 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 21.3s 4 8 $0.2065 79,997 78,470 21,832 1,527 614 55,845 22,625 — —
✅ 101_loki_historical_logs_pod_deleted 40.1s 7 11 $0.2687 144,987 142,584 23,378 2,403 429 118,562 24,022 — —
✅ 111_pod_names_contain_service 23.7s 5 8 $0.3326 98,538 96,981 21,450 1,557 384 54,249 42,732 — —
✅ 112_find_pvcs_by_uuid 18.3s 5 4 $0.2964 96,440 95,377 21,728 1,063 283 56,748 38,629 — —
✅ 12_job_crashing 27.9s 6 14 $0.2634 128,987 127,125 24,963 1,862 498 99,996 27,129 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 37.9s 7 15 $0.2920 154,520 152,074 26,207 2,446 676 124,943 27,131 — —
✅ 183a_elasticsearch_cluster_health 11.0s 3 3 $0.1557 59,423 58,992 20,069 431 187 38,912 20,080 — —
✅ 183b_elasticsearch_index_discovery 10.4s 3 3 $0.2717 59,698 59,241 20,237 457 270 19,132 40,109 — —
✅ 183c_elasticsearch_log_search 22.6s 6 6 $0.2141 124,057 122,971 21,804 1,086 274 101,153 21,818 — —
✅ 183d_elasticsearch_aggregation 24.6s 7 9 $0.2312 146,019 144,831 22,296 1,188 278 122,348 22,483 — —
✅ 183e_elasticsearch_field_mappings 12.0s 3 3 $0.1608 59,820 59,227 20,228 593 263 38,988 20,239 — —
✅ 183f_elasticsearch_shard_filtering 13.1s 3 3 $0.1650 60,226 59,502 20,367 724 357 39,124 20,378 — —
✅ 183g_elasticsearch_index_stats 12.3s 3 3 $0.1684 61,780 61,151 21,225 629 288 39,915 21,236 — —
✅ 184_elasticsearch_index_explosion 14.1s 3 3 $0.3178 106,179 105,568 43,428 611 273 62,129 43,439 — —
✅ 185_elasticsearch_cross_region_search 39.9s 7 10 $0.5058 313,547 311,392 51,511 2,155 516 259,866 51,526 — —
✅ 186_elasticsearch_shard_explosion 14.4s 3 3 $0.1710 60,997 60,124 20,710 873 417 39,403 20,721 — —
✅ 187_elasticsearch_disk_space 13.7s 3 3 $0.1664 61,097 60,450 20,858 647 280 39,581 20,869 — —
✅ 188_elasticsearch_mapping_explosion 25.9s 4 8 $0.3000 100,933 99,430 30,299 1,503 537 60,567 38,863 — —
✅ 189_elasticsearch_timeseries_gap 26.9s 5 9 $0.2594 118,032 115,951 25,694 2,081 507 89,896 26,055 — —
✅ 190_elasticsearch_cross_service_correlation 22.4s 4 7 $0.2232 87,846 86,095 23,551 1,751 540 62,326 23,769 — —
✅ 191_elasticsearch_query_profile 15.5s 3 3 $0.1741 62,139 61,301 21,294 838 388 39,996 21,305 — —
✅ 193_elasticsearch_large_mapping_search 11.5s 3 3 $0.1574 59,746 59,299 20,264 447 274 39,024 20,275 — —
✅ 195_elasticsearch_trace_large_fields 34.2s 7 9 $0.4052 241,946 240,167 41,548 1,779 340 198,228 41,939 — —
✅ 227_count_configmaps_per_namespace[0] 23.4s 6 11 $0.2132 113,821 112,388 20,576 1,433 434 91,184 21,204 — —
✅ 24_misconfigured_pvc 29.4s 6 14 $0.2501 122,292 120,209 23,231 2,083 702 96,021 24,188 — —
✅ 43_current_datetime_from_prompt 3.8s 1 — $0.1081 16,964 16,853 16,853 111 111 0 16,853 — —
✅ 61_exact_match_counting 12.2s 4 4 $0.1507 70,837 70,391 18,135 446 227 52,244 18,147 — —
Total 20.8s avg 4.5 avg 6.8 avg $6.4288 2,810,868 2,778,144 51,511 32,724 702 2,040,380 737,764 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ e4ebff3 (#23114091570)

✅ Results of HolmesGPT evals

Automatically triggered by commit e4ebff3 on branch claude/review-toolset-fixes-X2pPc

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 30.6s 6 11 $0.2528 121,197 119,064 23,544 2,133 631 94,516 24,548 — —
✅ 101_loki_historical_logs_pod_deleted 34.6s 5 10 $0.2550 107,560 105,200 24,197 2,360 800 80,062 25,138 — —
✅ 111_pod_names_contain_service 26.6s 5 8 $0.2118 95,871 94,275 21,337 1,596 612 72,472 21,803 — —
✅ 112_find_pvcs_by_uuid 24.1s 6 5 $0.2056 112,760 111,422 20,248 1,338 286 91,160 20,262 — —
✅ 12_job_crashing 27.9s 5 11 $0.2460 111,312 109,593 25,022 1,719 434 83,612 25,981 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 34.7s 6 13 $0.2769 129,882 127,364 25,690 2,518 710 101,108 26,256 — —
✅ 227_count_configmaps_per_namespace[0] 22.6s 6 10 $0.2082 112,584 111,172 20,399 1,412 435 90,759 20,413 — —
✅ 24_misconfigured_pvc 32.8s 6 16 $0.3836 124,239 121,915 23,927 2,324 544 75,684 46,231 — —
✅ 43_current_datetime_from_prompt 4.4s 1 — $0.1084 16,977 16,853 16,853 124 124 0 16,853 — —
✅ 61_exact_match_counting 9.6s 3 2 $0.1334 51,849 51,566 17,475 283 163 34,080 17,486 — —
Total 24.8s avg 4.9 avg 9.6 avg $2.2817 984,231 968,424 25,690 15,807 800 723,453 244,971 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 98a7112 (#22950260085)

✅ Results of HolmesGPT evals

Automatically triggered by commit 98a7112 on branch claude/review-toolset-fixes-X2pPc

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 30.7s 5 10 $0.2326 105,284 103,477 1,807 79,829 23,648 — 450 —
✅ 101_loki_historical_logs_pod_deleted 31.4s 4 8 $0.2161 82,707 80,844 1,863 58,346 22,498 — 744 —
✅ 111_pod_names_contain_service 27.2s 4 8 $0.2069 80,564 79,022 1,542 56,440 22,582 — 505 —
✅ 112_find_pvcs_by_uuid 21.8s 4 4 $0.1905 79,212 78,032 1,180 56,819 21,213 — 577 —
✅ 12_job_crashing 32.7s 5 12 $0.2502 111,987 110,118 1,869 84,023 26,095 — 489 —
✅ 176_network_policy_blocking_traffic_no_runbooks 42.4s 7 14 $0.3046 161,864 159,339 2,525 130,967 28,372 — 522 —
✅ 227_count_configmaps_per_namespace[0] 26.5s 6 11 $0.2202 117,382 115,896 1,486 93,990 21,906 — 429 —
✅ 24_misconfigured_pvc 30.8s 5 13 $0.3514 104,629 102,879 1,750 57,920 44,959 — 653 —
✅ 43_current_datetime_from_prompt 4.9s 1 — $0.1115 17,481 17,359 122 0 17,359 — 122 —
✅ 61_exact_match_counting 11.3s 3 2 $0.1388 53,533 53,201 332 35,130 18,071 — 193 —
Total 26.0s avg 4.4 avg 9.1 avg $2.2227 914,643 900,167 14,476 653,464 246,703 — 744 —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 16 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ Eval Results (with failures)

Automatically triggered by commit fc8d9da on branch claude/review-toolset-fixes-X2pPc (labels: evals-tag-elasticsearch)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 25/27 test cases were successful, 2 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 28.6s 5 10 $0.2335 103,999 102,136 23,239 1,863 641 78,408 23,728 — —
✅ 101_loki_historical_logs_pod_deleted 44.9s 6 13 $0.2821 133,304 130,722 26,068 2,582 726 104,154 26,568 — —
✅ 111_pod_names_contain_service 25.7s 4 8 $0.2041 78,626 77,039 21,510 1,587 681 55,040 21,999 — —
✅ 112_find_pvcs_by_uuid 22.5s 5 4 $0.1951 93,609 92,249 19,991 1,360 342 72,245 20,004 — —
✅ 12_job_crashing 37.5s 6 17 $0.2953 139,949 137,377 26,535 2,572 640 108,683 28,694 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 45.4s 8 15 $0.3072 175,893 173,233 26,482 2,660 575 146,357 26,876 — —
✅ 183a_elasticsearch_cluster_health 11.1s 3 3 $0.1556 59,425 59,001 20,073 424 174 38,917 20,084 — —
✅ 183b_elasticsearch_index_discovery 10.8s 3 3 $0.1573 59,687 59,235 20,234 452 266 38,990 20,245 — —
✅ 183c_elasticsearch_log_search 23.7s 6 6 $0.2127 123,921 122,879 21,773 1,042 267 101,092 21,787 — —
✅ 183d_elasticsearch_aggregation 29.5s 8 10 $0.3536 167,182 165,927 22,469 1,255 311 124,307 41,620 — —
✅ 183e_elasticsearch_field_mappings 13.6s 3 3 $0.1612 59,861 59,259 20,249 602 260 38,999 20,260 — —
✅ 183f_elasticsearch_shard_filtering 13.4s 3 3 $0.1638 60,159 59,482 20,357 677 320 39,114 20,368 — —
✅ 183g_elasticsearch_index_stats 12.7s 3 3 $0.1674 61,687 61,085 21,177 602 275 39,897 21,188 — —
✅ 184_elasticsearch_index_explosion 15.5s 3 3 $0.3180 106,169 105,548 43,410 621 271 62,127 43,421 — —
✅ 185_elasticsearch_cross_region_search 39.5s 6 10 $0.5219 275,945 273,999 57,361 1,946 397 215,200 58,799 — —
✅ 186_elasticsearch_shard_explosion 20.2s 4 5 $0.1912 82,982 81,916 21,478 1,066 435 60,426 21,490 — —
✅ 187_elasticsearch_disk_space 14.7s 3 3 $0.1645 60,919 60,323 20,769 596 267 39,543 20,780 — —
❌ 188_elasticsearch_mapping_explosion 26.1s 4 9 $0.2085 84,720 83,203 22,421 1,517 517 60,770 22,433 — —
✅ 189_elasticsearch_timeseries_gap 38.2s 7 11 $0.2966 164,402 161,991 26,699 2,411 536 134,899 27,092 — —
✅ 190_elasticsearch_cross_service_correlation 21.8s 4 5 $0.2066 85,392 84,016 22,641 1,376 460 61,363 22,653 — —
✅ 191_elasticsearch_query_profile 16.5s 3 3 $0.1723 62,042 61,272 21,278 770 331 39,983 21,289 — —
✅ 193_elasticsearch_large_mapping_search 13.9s 3 3 $0.1602 59,954 59,410 20,327 544 286 39,072 20,338 — —
❌ 195_elasticsearch_trace_large_fields 38.4s 7 10 $0.4581 242,116 240,414 51,108 1,702 391 188,978 51,436 — —
✅ 227_count_configmaps_per_namespace[0] 25.3s 6 10 $0.2089 112,126 110,723 20,462 1,403 617 90,077 20,646 — —
✅ 24_misconfigured_pvc 33.0s 6 14 $0.2590 123,133 120,943 23,356 2,190 695 95,472 25,471 — —
✅ 43_current_datetime_from_prompt 4.5s 1 — $0.1083 16,971 16,853 16,853 118 118 0 16,853 — —
✅ 61_exact_match_counting 13.8s 4 4 $0.1506 70,819 70,375 18,127 444 227 52,236 18,139 — —
Total 23.7s avg 4.6 avg 7.2 avg $6.3135 2,864,992 2,830,610 57,361 34,382 726 2,126,349 704,261 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 2 Failures Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/review-toolset-fixes-X2pPc -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals in automatic regression runs:

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/review-toolset-fixes-X2pPc -f markers=regression -f filter=

@github-actions

github-actions Bot commented Mar 11, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 12.26s 10.94s +12.1%
Warm Mean 5.31s 5.05s +5.1%
Warm Min 5.22s 5.01s
Warm Max 5.37s 5.11s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 28.43s 25.06s +13.4%
Warm Mean 8.15s 8.14s +0.1%
Warm Min 8.06s 7.80s
Warm Max 8.20s 8.58s

PR: 49b88036 | Master: 9387f4cc | Iterations: 5

@aantn
aantn disabled auto-merge March 11, 2026 11:25
@naomi-robusta
naomi-robusta merged commit 539f057 into master Mar 16, 2026
21 of 22 checks passed
@naomi-robusta
naomi-robusta deleted the claude/review-toolset-fixes-X2pPc branch March 16, 2026 09:09
henrikrexed pushed a commit to henrikrexed/holmesgpt that referenced this pull request Mar 17, 2026
…lmesGPT#1741)

## Summary
This PR refactors the configuration class handling in toolsets to
support multiple configuration classes instead of a single hardcoded
class. This change enables greater flexibility in toolset configuration
management.

## Key Changes
- **Elasticsearch toolset**: Modified `prerequisites_callable` to
dynamically select the configuration class from `self.config_classes`
list, falling back to `ElasticsearchConfig` if the list is empty
- **Grafana toolset**: Changed `config_class` attribute from a single
`ClassVar` to `config_classes` as a list of configuration classes,
updating the type annotation from `Type[GrafanaDashboardConfig]` to
`list[Type[GrafanaDashboardConfig]]`

## Implementation Details
- The Elasticsearch implementation now checks if `config_classes` is
populated and uses the first class in the list, maintaining backward
compatibility with a fallback to the default `ElasticsearchConfig`
- This pattern allows toolsets to be extended with multiple
configuration options while maintaining existing behavior
- The change is backward compatible as long as subclasses properly
initialize the `config_classes` list

https://claude.ai/code/session_01BDFrttpmMrSWDNcnJGwBRi

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Enhanced Elasticsearch plugin configuration to dynamically select from
available configuration classes, improving flexibility.
* Updated Grafana toolset to support multiple configuration types,
enabling more versatile configuration handling.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Natan Yellin <aantn@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants