Skip to content

Make Grafana config class configurable in base toolset - #1732

Merged
aantn merged 3 commits into
masterfrom
claude/investigate-bug-cause-Y3mRR
Mar 10, 2026
Merged

aantn merged 3 commits into
masterfrom
claude/investigate-bug-cause-Y3mRR

Conversation

@aantn

@aantn aantn commented Mar 10, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Updated the BaseGrafanaToolset to support configurable config classes instead of hardcoding GrafanaConfig. This allows subclasses to use their own config class implementations while maintaining backward compatibility.

Key Changes

  • Modified prerequisites_callable() to dynamically select the config class from self.config_classes if available, falling back to GrafanaConfig as default
  • Enables subclasses to override the config class used during initialization without modifying the base implementation

Implementation Details

  • The change checks if self.config_classes exists and uses the first element as the config class
  • Maintains backward compatibility by defaulting to GrafanaConfig when config_classes is not defined or empty
  • This pattern allows for flexible configuration handling across different Grafana toolset implementations

https://claude.ai/code/session_01DBzG3kmGffrwmGxrWTqT84

Summary by CodeRabbit

  • New Features

    • Grafana configuration now selects and instantiates a configurable config class at runtime, allowing alternative configuration subclasses instead of a hardcoded default.
  • Tests

    • Added unit test coverage for Grafana Tempo behavior, including initialization with labels config and verification that Kubernetes filters are built correctly (producing expected regex filters).

The base class prerequisites_callable() always instantiated GrafanaConfig,
ignoring subclass config_classes (e.g. GrafanaTempoConfig which adds 'labels').
The cast() in the Tempo toolset didn't convert the object at runtime, causing
AttributeError when accessing .labels. Now uses the subclass's config_classes[0].

https://claude.ai/code/session_01DBzG3kmGffrwmGxrWTqT84
Signed-off-by: Claude <noreply@anthropic.com>
@netlify

netlify Bot commented Mar 10, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 769beb6
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69b03f675ad8530008add047
😎 Deploy Preview https://deploy-preview-1732--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

…allable

Verifies that prerequisites_callable creates GrafanaTempoConfig (with labels)
rather than plain GrafanaConfig, preventing the AttributeError regression.

https://claude.ai/code/session_01DBzG3kmGffrwmGxrWTqT84
Signed-off-by: Claude <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Mar 10, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: f1af74e2-02c2-4ddf-a0cb-66befdd9446a

📥 Commits

Reviewing files that changed from the base of the PR and between 479cfb7 and 769beb6.

📒 Files selected for processing (1)
  • tests/plugins/toolsets/grafana/test_grafana_tempo_unit.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • tests/plugins/toolsets/grafana/test_grafana_tempo_unit.py

Walkthrough

Replaces hardcoded GrafanaConfig instantiation in BaseGrafanaToolset with selection of a configurable class (uses self.config_classes[0] if present, otherwise GrafanaConfig). Adds a test exercising Tempo filters after prerequisites_callable that imports GrafanaTempoLabelsConfig and verifies generated k8s regex filters.

Changes

Cohort / File(s) Summary
Base Grafana toolset
holmes/plugins/toolsets/grafana/base_grafana_toolset.py
Selects config_class from self.config_classes[0] when available, otherwise falls back to GrafanaConfig; instantiates that class with provided config and returns health_check() or error on exception.
Grafana Tempo tests
tests/plugins/toolsets/grafana/test_grafana_tempo_unit.py
Adds GrafanaTempoLabelsConfig import and new test test_build_k8s_filters_after_prerequisites_callable that mocks GrafanaTempoAPI and verifies two regex k8s filters when use_exact_match is False.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

Suggested reviewers

  • RoiGlinik
  • arikalon1
  • moshemorad
🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 66.67% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Make Grafana config class configurable in base toolset' accurately and specifically describes the main change—making the config class selection dynamic rather than hardcoded.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
tests/plugins/toolsets/grafana/test_grafana_tempo_unit.py (1)

459-481: Well-structured regression test with clear documentation.

The docstring clearly explains the bug being prevented, and the assertions properly verify that:

  1. The config is the correct subclass type (GrafanaTempoConfig)
  2. The labels attribute exists
  3. The labels field is properly instantiated as GrafanaTempoLabelsConfig

Consider adding an assertion for the return value of prerequisites_callable() to verify the health check completes successfully (returns (True, ...)). This would make the test more comprehensive.

💡 Optional enhancement
     with patch(
         "holmes.plugins.toolsets.grafana.toolset_grafana_tempo.GrafanaTempoAPI"
     ):
-        toolset.prerequisites_callable(config)
+        result = toolset.prerequisites_callable(config)
 
+    assert result[0] is True, f"prerequisites_callable failed: {result[1]}"
     assert isinstance(toolset._grafana_config, GrafanaTempoConfig)
     assert hasattr(toolset._grafana_config, "labels")
     assert isinstance(toolset._grafana_config.labels, GrafanaTempoLabelsConfig)
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/plugins/toolsets/grafana/test_grafana_tempo_unit.py` around lines 459 -
481, Add an assertion that prerequisites_callable returns a successful
health-check tuple: call GrafanaTempoToolset.prerequisites_callable(config)
(within the same patched GrafanaTempoAPI context) and assert the returned
value's first element is True (e.g., returned_tuple[0] is True) to ensure the
health check completed successfully in addition to the existing type/attribute
assertions on toolset._grafana_config.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@tests/plugins/toolsets/grafana/test_grafana_tempo_unit.py`:
- Around line 459-481: Add an assertion that prerequisites_callable returns a
successful health-check tuple: call
GrafanaTempoToolset.prerequisites_callable(config) (within the same patched
GrafanaTempoAPI context) and assert the returned value's first element is True
(e.g., returned_tuple[0] is True) to ensure the health check completed
successfully in addition to the existing type/attribute assertions on
toolset._grafana_config.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: cc2e216e-e249-43fd-8932-35896465de2e

📥 Commits

Reviewing files that changed from the base of the PR and between 6e4b84b and 479cfb7.

📒 Files selected for processing (2)
  • holmes/plugins/toolsets/grafana/base_grafana_toolset.py
  • tests/plugins/toolsets/grafana/test_grafana_tempo_unit.py

…quisites_callable

Tests the actual user-facing behavior (building trace filters) rather than
checking internal types. Verified the test reproduces the exact original error
('GrafanaConfig' object has no attribute 'labels') when the fix is reverted.

https://claude.ai/code/session_01DBzG3kmGffrwmGxrWTqT84
Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Mar 10, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 6444c79 (#22911128333)

✅ Results of HolmesGPT evals

Automatically triggered by commit 6444c79 on branch claude/investigate-bug-cause-Y3mRR

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 28.5s 4 10 $0.2145 81,602 79,948 1,654 56,645 23,303 — 706 —
✅ 101_loki_historical_logs_pod_deleted 63.5s 9 16 $0.3586 206,839 203,623 3,216 172,275 31,348 — 611 —
✅ 111_pod_names_contain_service 35.5s 6 9 $0.2322 119,054 117,259 1,795 94,788 22,471 — 499 —
✅ 112_find_pvcs_by_uuid 21.5s 4 4 $0.1841 77,323 76,250 1,073 55,529 20,721 — 496 —
✅ 12_job_crashing 41.2s 7 14 $0.2801 156,349 154,204 2,145 128,168 26,036 — 529 —
✅ 176_network_policy_blocking_traffic_no_runbooks 51.2s 7 12 $0.4150 153,769 151,316 2,453 102,463 48,853 — 486 —
✅ 227_count_configmaps_per_namespace[0] 27.3s 6 11 $0.2136 114,650 113,227 1,423 91,991 21,236 — 432 —
✅ 24_misconfigured_pvc 44.5s 5 11 $0.2331 102,027 100,040 1,987 76,580 23,460 — 663 —
✅ 43_current_datetime_from_prompt 5.2s 1 — $0.1094 17,119 16,991 128 0 16,991 — 128 —
✅ 61_exact_match_counting 15.0s 4 4 $0.1517 71,399 70,955 444 52,667 18,288 — 227 —
Total 33.3s avg 5.3 avg 10.1 avg $2.3925 1,100,131 1,083,813 16,318 831,106 252,707 — 706 —

Benchmark comparison unavailable: No ci-benchmark experiments found

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No ci-benchmark experiments found

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⏳ HolmesGPT evals running...

Automatically triggered by commit 769beb6 on branch claude/investigate-bug-cause-Y3mRR

View workflow logs

Progress:

  • Setup HolmesGPT environment
  • Collect evals to run (11 tests)
  • Setup KIND cluster
  • Run evals
📋 Evals to run
tests/llm/test_ask_holmes.py::test_ask_holmes[61_exact_match_counting-opus-4.5-default]
tests/llm/test_ask_holmes.py::test_ask_holmes[227_count_configmaps_per_namespace[0]-opus-4.5-default]
tests/llm/test_ask_holmes.py::test_ask_holmes[43_current_datetime_from_prompt-opus-4.5-default]
tests/llm/test_ask_holmes.py::test_ask_holmes[101_loki_historical_logs_pod_deleted-opus-4.5-default]
tests/llm/test_ask_holmes.py::test_ask_holmes[12_job_crashing-opus-4.5-default]
tests/llm/test_ask_holmes.py::test_ask_holmes[09_crashpod-opus-4.5-default]
tests/llm/test_ask_holmes.py::test_ask_holmes[176_network_policy_blocking_traffic_no_runbooks-opus-4.5-default]
tests/llm/test_ask_holmes.py::test_ask_holmes[112_find_pvcs_by_uuid-opus-4.5-default]
tests/llm/test_ask_holmes.py::test_ask_holmes[111_pod_names_contain_service-opus-4.5-default]
tests/llm/test_ask_holmes.py::test_ask_holmes[24_misconfigured_pvc-opus-4.5-default]
tests/llm/utils/test_toolset.py:35
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/investigate-bug-cause-Y3mRR -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals in automatic regression runs:

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID

Examples: evals-tag-easy, evals-id-09_crashpod

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/investigate-bug-cause-Y3mRR -f markers=regression -f filter=

@github-actions

github-actions Bot commented Mar 10, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 5ccef260 (built in 7m 25s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:5ccef260
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:5ccef260 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:5ccef260
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:5ccef260
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:5ccef260
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:5ccef260 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:5ccef260
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:5ccef260

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:5ccef260 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:5ccef260

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:5ccef260 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:5ccef260

@github-actions

github-actions Bot commented Mar 10, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 12.30s 11.23s +9.6%
Warm Mean 5.43s 5.19s +4.7%
Warm Min 5.38s 5.18s
Warm Max 5.51s 5.20s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 38.66s 19.12s +102.2%
Warm Mean 8.62s 8.72s -1.1%
Warm Min 8.28s 8.26s
Warm Max 8.94s 9.29s

PR: 5ccef260 | Master: 6e4b84be | Iterations: 5

@aantn
aantn enabled auto-merge (squash) March 10, 2026 16:06
@aantn
aantn merged commit d5fd607 into master Mar 10, 2026
20 of 22 checks passed
@aantn
aantn deleted the claude/investigate-bug-cause-Y3mRR branch March 10, 2026 16:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants