Skip to content

Add max retries & timeout to grafana config - #1927

Merged
moshemorad merged 13 commits into
masterfrom
grafana_timeout
Apr 26, 2026
Merged

moshemorad merged 13 commits into
masterfrom
grafana_timeout

Conversation

@moshemorad

@moshemorad moshemorad commented Apr 19, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • Grafana and Loki requests now use configurable timeout and retry settings with exponential backoff; dashboard rendering and availability probes respect configured timeouts.
  • Documentation

    • Added guidance on class-hierarchy placement for new config and recommended HTTP mocking via the responses library.
  • Tests

    • Added tests covering timeout/retry defaults and overrides, retry/backoff behavior, and success/failure request handling.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@github-actions

github-actions Bot commented Apr 19, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 #5 · Run @ __37c40b3__ (#24832704231) — Apr 23, 11:35 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 37c40b3 on branch grafana_timeout

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 34.6s 5 10 $0.2387 104,186 102,106 23,152 2,080 865 78,393 23,713 — —
✅ 101_loki_historical_logs_pod_deleted 47.7s 6 12 $0.2892 135,741 133,059 25,578 2,682 873 105,583 27,476 — —
✅ 112_find_pvcs_by_uuid 28.3s 5 4 $0.1935 94,468 93,325 20,559 1,143 286 72,753 20,572 — —
✅ 12_job_crashing 39.6s 6 14 $0.2806 138,030 135,776 25,608 2,254 616 108,139 27,637 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 48.2s 7 15 $0.3090 156,118 153,292 26,714 2,826 794 124,785 28,507 — —
✅ 227_count_configmaps_per_namespace[0] 25.0s 5 9 $0.2040 94,913 93,656 20,899 1,257 586 71,516 22,140 — —
✅ 243_pod_names_contain_service 33.8s 5 8 $0.2180 99,364 97,581 21,662 1,783 437 75,906 21,675 — —
✅ 24_misconfigured_pvc 35.2s 5 12 $0.2393 103,107 100,943 22,709 2,164 796 77,321 23,622 — —
✅ 43_current_datetime_from_prompt 4.7s 1 — $0.1092 17,057 16,919 16,919 138 138 0 16,919 — —
✅ 51_logs_summarize_errors 23.6s 4 5 $0.1876 77,586 76,475 21,131 1,111 363 55,332 21,143 — —
✅ 61_exact_match_counting 12.3s 3 3 $0.1397 52,920 52,511 17,939 409 262 34,561 17,950 — —
Total 30.3s avg 4.7 avg 9.2 avg $2.4087 1,073,490 1,055,643 26,714 17,847 873 804,289 251,354 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 34 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #4 · Run @ __a385529__ (#24732188585) — Apr 21, 16:10 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit a385529 on branch grafana_timeout

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 35.4s 5 10 $0.2390 104,421 102,390 23,381 2,031 570 78,434 23,956 — —
✅ 101_loki_historical_logs_pod_deleted 61.1s 7 14 $0.3167 162,847 159,596 26,910 3,251 769 132,495 27,101 — —
✅ 112_find_pvcs_by_uuid 30.4s 6 6 $0.2330 121,407 119,880 23,153 1,527 390 96,327 23,553 — —
✅ 12_job_crashing 36.4s 5 12 $0.2435 109,933 108,037 24,291 1,896 576 83,209 24,828 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 40.1s 6 10 $0.2655 129,577 127,360 24,848 2,217 447 101,730 25,630 — —
✅ 227_count_configmaps_per_namespace[0] 27.6s 6 10 $0.2149 113,486 111,969 20,624 1,517 519 90,803 21,166 — —
✅ 243_pod_names_contain_service 34.5s 6 8 $0.2261 116,706 115,005 21,639 1,701 421 93,006 21,999 — —
✅ 24_misconfigured_pvc 37.4s 5 13 $0.2465 103,987 101,658 22,947 2,329 809 77,478 24,180 — —
✅ 43_current_datetime_from_prompt 5.6s 1 — $0.1093 17,063 16,919 16,919 144 144 0 16,919 — —
✅ 51_logs_summarize_errors 25.7s 4 5 $0.1904 78,322 77,197 21,494 1,125 374 55,691 21,506 — —
✅ 61_exact_match_counting 7.7s 2 1 $0.1222 34,434 34,207 17,279 227 159 16,918 17,289 — —
Total 31.1s avg 4.8 avg 8.9 avg $2.4071 1,092,183 1,074,218 26,910 17,965 809 826,091 248,127 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 34 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #3 · Run @ __9fadbb1__ (#24722963655) — Apr 21, 12:49 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 9fadbb1 on branch grafana_timeout

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 34.4s 5 11 $0.2459 105,737 103,631 23,689 2,106 871 78,780 24,851 — —
✅ 101_loki_historical_logs_pod_deleted 56.3s 7 12 $0.3006 159,231 156,365 25,742 2,866 843 130,016 26,349 — —
✅ 112_find_pvcs_by_uuid 16.2s 3 3 $0.1772 60,552 59,625 21,601 927 569 38,013 21,612 — —
✅ 12_job_crashing 37.1s 5 12 $0.2534 109,792 107,587 24,359 2,205 680 82,254 25,333 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 45.1s 6 14 $0.3100 136,225 133,564 26,892 2,661 621 101,778 31,786 — —
✅ 227_count_configmaps_per_namespace[0] 21.9s 4 9 $0.1903 76,939 75,777 20,724 1,162 563 54,112 21,665 — —
✅ 243_pod_names_contain_service 29.8s 4 8 $0.2077 79,010 77,314 21,574 1,696 824 55,167 22,147 — —
✅ 24_misconfigured_pvc 39.5s 6 14 $0.2637 127,922 125,562 23,719 2,360 593 100,618 24,944 — —
✅ 43_current_datetime_from_prompt 4.5s 1 — $0.1088 17,041 16,919 16,919 122 122 0 16,919 — —
✅ 51_logs_summarize_errors 23.4s 4 5 $0.1888 77,994 76,886 21,330 1,108 343 55,544 21,342 — —
✅ 61_exact_match_counting 11.6s 3 3 $0.1390 52,857 52,471 17,918 386 239 34,542 17,929 — —
Total 29.1s avg 4.4 avg 9.1 avg $2.3853 1,003,300 985,701 26,892 17,599 871 730,824 254,877 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 34 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #2 · Run @ __42adcd1__ (#24718095787) — Apr 21, 10:51 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 42adcd1 on branch grafana_timeout

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 33.9s 5 9 $0.2347 99,926 97,950 22,999 1,976 957 74,038 23,912 — —
✅ 101_loki_historical_logs_pod_deleted 36.5s 4 8 $0.2294 84,694 82,580 23,512 2,114 875 59,056 23,524 — —
✅ 112_find_pvcs_by_uuid 17.1s 3 3 $0.1661 57,459 56,566 20,069 893 492 36,486 20,080 — —
✅ 12_job_crashing 44.3s 6 17 $0.3070 142,415 139,571 27,295 2,844 731 110,238 29,333 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 48.8s 7 16 $0.3178 157,583 154,759 26,925 2,824 633 124,495 30,264 — —
✅ 227_count_configmaps_per_namespace[0] 23.9s 5 9 $0.2026 94,875 93,616 20,879 1,259 578 71,797 21,819 — —
✅ 243_pod_names_contain_service 31.2s 4 8 $0.2113 79,484 77,697 21,762 1,787 862 55,352 22,345 — —
✅ 24_misconfigured_pvc 37.8s 6 13 $0.2632 123,755 121,410 23,489 2,345 714 95,962 25,448 — —
✅ 43_current_datetime_from_prompt 4.8s 1 — $0.1089 17,045 16,919 16,919 126 126 0 16,919 — —
✅ 51_logs_summarize_errors 25.6s 4 5 $0.1843 76,917 75,851 20,805 1,066 365 55,034 20,817 — —
✅ 61_exact_match_counting 12.5s 3 3 $0.1388 52,839 52,459 17,912 380 233 34,536 17,923 — —
Total 28.8s avg 4.4 avg 9.1 avg $2.3639 986,992 969,378 27,295 17,614 957 716,994 252,384 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 34 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #1 · Run @ __483d9e3__ (#24708821129) — Apr 21, 07:16 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 483d9e3 on branch grafana_timeout

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 29.2s 4 9 $0.2237 82,049 80,145 22,942 1,904 970 56,288 23,857 — —
✅ 101_loki_historical_logs_pod_deleted 47.8s 6 12 $0.2838 133,034 130,171 25,145 2,863 917 104,459 25,712 — —
✅ 112_find_pvcs_by_uuid 17.8s 3 3 $0.1778 60,365 59,385 21,486 980 558 37,888 21,497 — —
✅ 12_job_crashing 34.6s 5 14 $0.2675 113,638 111,348 25,548 2,290 743 84,085 27,263 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 43.8s 7 17 $0.3080 159,778 156,844 25,978 2,934 609 129,343 27,501 — —
✅ 227_count_configmaps_per_namespace[0] 24.9s 6 10 $0.2137 113,445 111,936 20,615 1,509 519 90,986 20,950 — —
✅ 243_pod_names_contain_service 31.8s 5 8 $0.2190 99,766 97,978 21,786 1,788 461 76,179 21,799 — —
✅ 24_misconfigured_pvc 29.6s 5 10 $0.2241 99,631 97,793 21,840 1,838 484 75,146 22,647 — —
✅ 43_current_datetime_from_prompt 4.0s 1 — $0.1090 17,049 16,919 16,919 130 130 0 16,919 — —
✅ 51_logs_summarize_errors 23.4s 4 5 $0.1867 77,253 76,129 20,954 1,124 366 55,163 20,966 — —
✅ 61_exact_match_counting 6.6s 2 1 $0.1222 34,436 34,207 17,279 229 160 16,918 17,289 — —
Total 26.7s avg 4.4 avg 8.9 avg $2.3355 990,444 972,855 25,978 17,589 970 726,455 246,400 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 34 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit 18688e2 on branch grafana_timeout

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 43.8s 7 11 $0.2710 144,730 142,444 24,355 2,286 867 117,506 24,938 — —
✅ 101_loki_historical_logs_pod_deleted 47.7s 6 12 $0.2785 132,800 130,074 24,568 2,726 881 104,615 25,459 — —
✅ 112_find_pvcs_by_uuid 19.0s 3 3 $0.1777 60,401 59,432 21,512 969 575 37,909 21,523 — —
✅ 12_job_crashing 35.1s 5 11 $0.2498 109,826 107,827 24,199 1,999 600 82,131 25,696 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 57.5s 8 15 $0.3202 181,550 178,665 26,535 2,885 846 150,759 27,906 — —
✅ 227_count_configmaps_per_namespace[0] 24.4s 5 9 $0.2034 94,828 93,589 20,869 1,239 578 71,472 22,117 — —
✅ 243_pod_names_contain_service 36.6s 5 9 $0.2285 100,827 98,838 22,251 1,989 584 76,279 22,559 — —
✅ 24_misconfigured_pvc 41.7s 7 14 $0.2786 146,173 143,603 24,219 2,570 604 118,647 24,956 — —
✅ 43_current_datetime_from_prompt 5.1s 1 — $0.1088 17,042 16,919 16,919 123 123 0 16,919 — —
✅ 51_logs_summarize_errors 23.2s 4 5 $0.1868 77,385 76,276 21,027 1,109 346 55,237 21,039 — —
✅ 61_exact_match_counting 8.8s 2 1 $0.1226 34,467 34,228 17,300 239 171 16,918 17,310 — —
Total 31.2s avg 4.8 avg 9.0 avg $2.4259 1,100,029 1,081,895 26,535 18,134 881 831,473 250,422 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 34 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref grafana_timeout -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, token-limit, toolset-limitation, traces, transparency

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref grafana_timeout -f markers=regression -f filter=

Signed-off-by: Mohse Morad <moshemorad12340@gmail.com>
@github-actions

github-actions Bot commented Apr 19, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for b3edcf40 (built in 7m 8s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:b3edcf40
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:b3edcf40 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:b3edcf40
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:b3edcf40
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:b3edcf40
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:b3edcf40 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:b3edcf40
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:b3edcf40

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:b3edcf40 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:b3edcf40

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:b3edcf40 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:b3edcf40

@netlify

netlify Bot commented Apr 19, 2026

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 3455126c686db2eb98faf6dc32a214bd1ac9bc8a
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69e4bc6819addd0008c6bda3
😎 Deploy Preview https://deploy-preview-1927--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@netlify

netlify Bot commented Apr 19, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 18688e2
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69ebafc8b168d30008d7af95
😎 Deploy Preview https://deploy-preview-1927--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai

coderabbitai Bot commented Apr 19, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Adds timeout and retry configuration to Grafana configs and threads them into Tempo, Loki, and core Grafana request paths; replaces module-level retry decorators with per-request backoff, updates method signatures to accept optional timeout/retries, adds tests, and documents HTTP mocking guidance.

Changes

Cohort / File(s) Summary
Documentation
CLAUDE.md
Added "Class Hierarchy Placement" guidance and testing guideline recommending the responses library for HTTP mocking.
Config models
holmes/plugins/toolsets/grafana/common.py
Added timeout_seconds and max_retries to GrafanaConfig; reformatted GrafanaTempoLabelsConfig field declarations.
Grafana Tempo API
holmes/plugins/toolsets/grafana/grafana_tempo_api.py
_make_request now accepts optional timeout/retries (defaulting to config); per-request backoff applied; query_echo_endpoint uses config timeouts/retries and raises on non-success within retry block.
Grafana Loki toolset
holmes/plugins/toolsets/grafana/loki/toolset_grafana_loki.py
health_check() and LokiQuery._invoke() forward timeout_seconds and max_retries from config into execute_loki_query.
Grafana Loki API
holmes/plugins/toolsets/grafana/loki_api.py
Removed module-level backoff decorator; execute_loki_query gains timeout and max_retries, creates inner _make_request() wrapped with per-request backoff (giveup on <500 HTTP errors); removed exported query_loki_logs_by_label.
Grafana Toolset core
holmes/plugins/toolsets/grafana/toolset_grafana.py
Added GrafanaDashboardConfig.timeout_seconds; _make_grafana_request/_make_render_request accept Optional[int] timeouts that default to config and use per-request backoff; renderer probing uses config timeout.
Tests
tests/plugins/toolsets/grafana/test_grafana_config.py
New tests for config defaults/validation and retry/timeout propagation for Tempo and Loki using responses (call-count assertions, mocked payloads) and time.sleep patching to avoid delays.

Sequence Diagram(s)

sequenceDiagram
  participant Tool as BaseGrafanaTool / Toolset
  participant API as Grafana API wrapper (Tempo / Loki)
  participant HTTP as HTTP client (requests)
  participant Remote as Grafana / Loki service

  Tool->>API: invoke request (timeout/retries omitted)
  API->>API: resolve timeout := config.timeout_seconds\nretries := config.max_retries
  API->>HTTP: perform HTTP call with timeout
  alt transient network error / 5xx
    HTTP-->>API: raise RequestException / 5xx response
    API->>HTTP: retry per backoff (up to retries)
  else 2xx response
    HTTP-->>API: return response (JSON)
  end
  API-->>Tool: return parsed result or error
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • arikalon1
  • RoiGlinik
  • aantn
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 26.67% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: adding max retries and timeout configuration options to the Grafana config classes across the codebase.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Apr 19, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 11.92s 11.90s +0.2%
Warm Mean 5.41s 5.69s -5.0%
Warm Min 5.24s 5.59s
Warm Max 5.61s 5.82s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 11.85s 15.24s -22.2%
Warm Mean 7.51s 7.89s -4.8%
Warm Min 6.92s 7.68s
Warm Max 7.94s 8.09s

PR: b3edcf40 | Master: 2be63a92 | Iterations: 5

@moshemorad

Copy link
Copy Markdown
Collaborator Author

@claude review

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (3)
holmes/plugins/toolsets/grafana/grafana_tempo_api.py (1)

146-152: ⚠️ Potential issue | 🟠 Major

Apply max_retries to the echo endpoint too.

This request now honors timeout_seconds, but transient connection failures still get only one attempt even when config.max_retries is set.

🔁 Proposed retry wrapper
         try:
-            response = requests.get(
-                url,
-                headers=self.headers,
-                timeout=self.config.timeout_seconds,
-                verify=self.config.verify_ssl,
-            )
+            `@backoff.on_exception`(
+                backoff.expo,
+                requests.exceptions.RequestException,
+                max_tries=self.config.max_retries,
+                giveup=lambda e: isinstance(e, requests.exceptions.HTTPError)
+                and getattr(e, "response", None) is not None
+                and e.response.status_code < 500,
+            )
+            def make_request():
+                response = requests.get(
+                    url,
+                    headers=self.headers,
+                    timeout=self.config.timeout_seconds,
+                    verify=self.config.verify_ssl,
+                )
+                if response.status_code >= 500:
+                    response.raise_for_status()
+                return response
+
+            response = make_request()
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/toolsets/grafana/grafana_tempo_api.py` around lines 146 - 152,
The echo endpoint currently calls requests.get directly and only honors
timeout_seconds and verify_ssl but not self.config.max_retries; wrap the call in
a requests.Session configured with a urllib3 Retry using
total=self.config.max_retries and mount an HTTPAdapter (or reuse the existing
session/adapter if one exists) so transient connection errors are retried;
ensure you use the same headers, timeout (self.config.timeout_seconds), and
verify (self.config.verify_ssl) when calling session.get instead of requests.get
to apply the retry behavior.
holmes/plugins/toolsets/grafana/loki_api.py (1)

42-71: ⚠️ Potential issue | 🟠 Major

Preserve Loki query context and API response in retry failures.

After retries are exhausted, the raised error still omits the LogQL query, time range, limit, URL, and response body. That makes Loki failures hard for the LLM to self-correct.

🧩 Proposed error-detail fix
     params = {"query": query, "limit": limit, "start": start, "end": end}
+    url = f"{base_url}/loki/api/v1/query_range"
 
     `@backoff.on_exception`(
         backoff.expo,
         requests.exceptions.RequestException,
         max_tries=max_retries,
         giveup=lambda e: isinstance(e, requests.exceptions.HTTPError)
-        and e.response.status_code < 500,
+        and getattr(e, "response", None) is not None
+        and e.response.status_code < 500,
     )
     def _make_request():
-        url = f"{base_url}/loki/api/v1/query_range"
         response = requests.get(
             url,
             headers=build_headers(api_key=api_key, additional_headers=headers),
             params=params,  # type: ignore
             verify=verify_ssl,
@@
-    except requests.exceptions.RequestException as e:
-        raise Exception(f"Failed to query Loki logs: {str(e)}")
+    except requests.exceptions.HTTPError as e:
+        response = e.response
+        status_code = response.status_code if response is not None else "unknown"
+        response_text = response.text if response is not None else ""
+        raise Exception(
+            f"Failed to query Loki logs. URL: {url}. Query: {query}. "
+            f"Start: {start}. End: {end}. Limit: {limit}. "
+            f"HTTP status: {status_code}. Response: {response_text}. Error: {e}"
+        )
+    except requests.exceptions.RequestException as e:
+        raise Exception(
+            f"Failed to query Loki logs. URL: {url}. Query: {query}. "
+            f"Start: {start}. End: {end}. Limit: {limit}. Error: {e}"
+        )

As per coding guidelines, holmes/plugins/toolsets/**/*.{py,yaml}: “All toolsets MUST return detailed error messages from underlying APIs to enable LLM self-correction: Include the exact query/command executed, time ranges/parameters/filters used, and full API error response”.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/toolsets/grafana/loki_api.py` around lines 42 - 71, The except
block after _make_request must include the full Loki request context and API
response details when raising the final error: update the exception handling
around the _make_request call to capture and include params (query, start, end,
limit), the resolved url (built from base_url + "/loki/api/v1/query_range"), any
request headers (from build_headers with api_key and headers), and the HTTP
response details (status_code and response.text or response.content when
available) in the raised Exception message; reference the existing symbols
_make_request, params, base_url, api_key, headers, response (from
_make_request), and parse_loki_response so you add this contextual information
to the Exception(f"...") that currently raises "Failed to query Loki logs:
{str(e)}".
holmes/plugins/toolsets/grafana/toolset_grafana.py (1)

202-237: ⚠️ Potential issue | 🟠 Major

Honor max_retries for dashboard API requests.

_make_grafana_request now uses configured timeouts, but dashboard tools still make a single attempt. This leaves GrafanaDashboardConfig.max_retries ineffective for search/get/tag calls.

🔁 Proposed retry support
+import backoff
 import requests  # type: ignore
-        response = requests.get(
-            url,
-            headers=headers,
-            params=query_params,
-            timeout=timeout,
-            verify=config.verify_ssl,
-        )
-        response.raise_for_status()
+        `@backoff.on_exception`(
+            backoff.expo,
+            requests.exceptions.RequestException,
+            max_tries=config.max_retries,
+            giveup=lambda e: isinstance(e, requests.exceptions.HTTPError)
+            and getattr(e, "response", None) is not None
+            and e.response.status_code < 500,
+        )
+        def make_request():
+            response = requests.get(
+                url,
+                headers=headers,
+                params=query_params,
+                timeout=timeout,
+                verify=config.verify_ssl,
+            )
+            response.raise_for_status()
+            return response
+
+        response = make_request()
         data = response.json()

As per coding guidelines, holmes/plugins/toolsets/**/*.py: “When adding new config fields, methods, or behavior, check the class hierarchy and place changes at the most general level that applies, not scoped to specific subclasses”.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/toolsets/grafana/toolset_grafana.py` around lines 202 - 237,
The _make_grafana_request currently issues a single requests.get and ignores
GrafanaDashboardConfig.max_retries; update it to honor max_retries by creating
or reusing a requests.Session, mounting a HTTPAdapter configured with
urllib3.util.retry.Retry(total=config.max_retries, backoff_factor=...,
status_forcelist=[429,500,502,503,504],
allowed_methods=["GET","HEAD","OPTIONS"]) and then perform session.get(...) with
the same headers, params, timeout, and verify; ensure the session is properly
reused or closed and that the change is applied in the _make_grafana_request
function (referencing config = self._toolset.grafana_config and the max_retries
field) so all dashboard search/get/tag calls inherit the retry behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@holmes/plugins/toolsets/grafana/common.py`:
- Around line 54-63: The timeout_seconds and max_retries fields need strict
Pydantic validation so invalid YAML is rejected at parse time: in the Grafana
config model in holmes/plugins/toolsets/grafana/common.py, change
timeout_seconds Field to include gt=0 (timeout_seconds > 0) and change
max_retries Field to include ge=1 (max_retries >= 1); if the model currently
lacks typed constraints or uses custom parsing, add `@validator` methods for
"timeout_seconds" and "max_retries" that raise ValueError for values <=0 and <1
respectively, and keep prerequisites_callable() unchanged except rely on the
model validation to prevent invalid configs from reaching runtime.

In `@tests/plugins/toolsets/grafana/test_grafana_config.py`:
- Around line 83-124: Update the tests that instantiate
GrafanaTempoConfig/GrafanaTempoAPI to assert that the configured timeout is
actually passed to requests by adding
responses.matchers.request_kwargs_matcher({"timeout": <timeout>}) to the
rsps.add(...) calls; specifically, in
test_make_request_with_custom_config_returns_data (uses
GrafanaTempoConfig(timeout_seconds=90) and calls
GrafanaTempoAPI.search_traces_by_query) add
match=[matchers.request_kwargs_matcher({"timeout": 90})], and in
test_echo_endpoint_with_custom_config (uses timeout_seconds=120 and calls
GrafanaTempoAPI.query_echo_endpoint) add
match=[matchers.request_kwargs_matcher({"timeout": 120})]; apply the same
pattern to the other related test referenced (the one around lines 173-195) so
each rsps.add verifies the expected timeout value from the GrafanaTempoConfig.

---

Outside diff comments:
In `@holmes/plugins/toolsets/grafana/grafana_tempo_api.py`:
- Around line 146-152: The echo endpoint currently calls requests.get directly
and only honors timeout_seconds and verify_ssl but not self.config.max_retries;
wrap the call in a requests.Session configured with a urllib3 Retry using
total=self.config.max_retries and mount an HTTPAdapter (or reuse the existing
session/adapter if one exists) so transient connection errors are retried;
ensure you use the same headers, timeout (self.config.timeout_seconds), and
verify (self.config.verify_ssl) when calling session.get instead of requests.get
to apply the retry behavior.

In `@holmes/plugins/toolsets/grafana/loki_api.py`:
- Around line 42-71: The except block after _make_request must include the full
Loki request context and API response details when raising the final error:
update the exception handling around the _make_request call to capture and
include params (query, start, end, limit), the resolved url (built from base_url
+ "/loki/api/v1/query_range"), any request headers (from build_headers with
api_key and headers), and the HTTP response details (status_code and
response.text or response.content when available) in the raised Exception
message; reference the existing symbols _make_request, params, base_url,
api_key, headers, response (from _make_request), and parse_loki_response so you
add this contextual information to the Exception(f"...") that currently raises
"Failed to query Loki logs: {str(e)}".

In `@holmes/plugins/toolsets/grafana/toolset_grafana.py`:
- Around line 202-237: The _make_grafana_request currently issues a single
requests.get and ignores GrafanaDashboardConfig.max_retries; update it to honor
max_retries by creating or reusing a requests.Session, mounting a HTTPAdapter
configured with urllib3.util.retry.Retry(total=config.max_retries,
backoff_factor=..., status_forcelist=[429,500,502,503,504],
allowed_methods=["GET","HEAD","OPTIONS"]) and then perform session.get(...) with
the same headers, params, timeout, and verify; ensure the session is properly
reused or closed and that the change is applied in the _make_grafana_request
function (referencing config = self._toolset.grafana_config and the max_retries
field) so all dashboard search/get/tag calls inherit the retry behavior.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 6b0eaad3-9261-4bf8-b264-4105773490cb

📥 Commits

Reviewing files that changed from the base of the PR and between 4c6a8c6 and 960c455.

📒 Files selected for processing (7)
  • CLAUDE.md
  • holmes/plugins/toolsets/grafana/common.py
  • holmes/plugins/toolsets/grafana/grafana_tempo_api.py
  • holmes/plugins/toolsets/grafana/loki/toolset_grafana_loki.py
  • holmes/plugins/toolsets/grafana/loki_api.py
  • holmes/plugins/toolsets/grafana/toolset_grafana.py
  • tests/plugins/toolsets/grafana/test_grafana_config.py

Comment thread holmes/plugins/toolsets/grafana/common.py
Comment thread tests/plugins/toolsets/grafana/test_grafana_config.py
Comment thread holmes/plugins/toolsets/grafana/toolset_grafana.py
Comment thread holmes/plugins/toolsets/grafana/common.py
Comment thread holmes/plugins/toolsets/grafana/loki_api.py
@moshemorad

Copy link
Copy Markdown
Collaborator Author

@claude fix the comments.

Comment thread tests/plugins/toolsets/grafana/test_grafana_config.py Outdated
Comment thread holmes/plugins/toolsets/grafana/common.py
Comment thread holmes/plugins/toolsets/grafana/grafana_tempo_api.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
holmes/plugins/toolsets/grafana/toolset_grafana.py (1)

582-588: ⚠️ Potential issue | 🟡 Minor

Render requests still ignore max_retries.

grafana_render_panel and grafana_render_dashboard now honor timeout_seconds, but transient renderer failures still make only one HTTP attempt. If max_retries is meant to cover Grafana API calls broadly, route render requests through the same retry behavior or document render calls as timeout-only.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/toolsets/grafana/toolset_grafana.py` around lines 582 - 588,
The render requests in grafana_render_panel and grafana_render_dashboard call
requests.get directly (the response = requests.get(...) call) and therefore
ignore the configured max_retries; update these render code paths to use the
same retry-enabled HTTP helper used elsewhere (e.g., the module's existing
retry/session helper or a shared function like
requests_retry_session/_do_request) so render calls honor config.max_retries and
timeout_seconds, or alternatively pass a session configured with
HTTPAdapter(retries=...) into requests.get; change both grafana_render_panel and
grafana_render_dashboard to route their GET to that helper/session instead of
calling requests.get directly so transient renderer failures are retried.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@holmes/plugins/toolsets/grafana/toolset_grafana.py`:
- Around line 252-253: The call to response = _do_request() can raise
HTTPError/Timeout/ConnectionError and currently escapes without returning a
structured error; wrap the _do_request() invocation in a try/except that catches
requests.exceptions.HTTPError, requests.exceptions.Timeout,
requests.exceptions.ConnectionError (and a generic Exception fallback), and
return a structured Grafana API error object/string containing the request URL,
query params/filters, retry and timeout settings used, and the full response
body/status_code when available (use e.response or response.text/status_code for
HTTPError). Update the logic around response = _do_request() / data =
response.json() to only parse JSON when the call succeeds and to include these
detailed fields in the tool result on failure so LLMs can self-correct.
- Around line 223-237: The backoff decorator currently uses config.max_retries
as max_tries which undercounts because max_tries is total calls; update the
backoff.on_exception call to pass max_tries=config.max_retries + 1 (i.e.,
convert retry attempts to total calls) where the decorator is applied (the
backoff.on_exception block that wraps the request call and references retries,
config.max_retries, and requests.exceptions.RequestException).

---

Outside diff comments:
In `@holmes/plugins/toolsets/grafana/toolset_grafana.py`:
- Around line 582-588: The render requests in grafana_render_panel and
grafana_render_dashboard call requests.get directly (the response =
requests.get(...) call) and therefore ignore the configured max_retries; update
these render code paths to use the same retry-enabled HTTP helper used elsewhere
(e.g., the module's existing retry/session helper or a shared function like
requests_retry_session/_do_request) so render calls honor config.max_retries and
timeout_seconds, or alternatively pass a session configured with
HTTPAdapter(retries=...) into requests.get; change both grafana_render_panel and
grafana_render_dashboard to route their GET to that helper/session instead of
calling requests.get directly so transient renderer failures are retried.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 5438a269-6426-47c8-8c8a-80b6af7e73fa

📥 Commits

Reviewing files that changed from the base of the PR and between 960c455 and 5c1e354.

📒 Files selected for processing (5)
  • holmes/plugins/toolsets/grafana/common.py
  • holmes/plugins/toolsets/grafana/grafana_tempo_api.py
  • holmes/plugins/toolsets/grafana/loki_api.py
  • holmes/plugins/toolsets/grafana/toolset_grafana.py
  • tests/plugins/toolsets/grafana/test_grafana_config.py
✅ Files skipped from review due to trivial changes (1)
  • tests/plugins/toolsets/grafana/test_grafana_config.py
🚧 Files skipped from review as they are similar to previous changes (3)
  • holmes/plugins/toolsets/grafana/common.py
  • holmes/plugins/toolsets/grafana/grafana_tempo_api.py
  • holmes/plugins/toolsets/grafana/loki_api.py

Comment thread holmes/plugins/toolsets/grafana/toolset_grafana.py
Comment thread holmes/plugins/toolsets/grafana/toolset_grafana.py
Comment thread holmes/plugins/toolsets/grafana/toolset_grafana.py
arikalon1
arikalon1 previously approved these changes Apr 19, 2026

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — all previously raised issues have been addressed: backoff retry logic added to all request paths, Pydantic bounds validated, giveup lambda None guards consistent, and tests updated with sleep mocking and timeout matchers.

Extended reasoning...

Overview

This PR adds timeout_seconds and max_retries to GrafanaConfig and threads them through all Grafana toolsets (Tempo, Loki, Dashboards, Render). The implementation is now consistent across every HTTP code path.

Security risks

None. This is purely a configuration and resilience change — no auth, no crypto, no permissions changes. The new Pydantic gt=0/ge=1 constraints prevent invalid config from reaching runtime.

Level of scrutiny

Medium. This touches several HTTP request paths and adds backoff retry logic, but the changes are well-patterned and follow the existing Tempo/Loki implementations. All evals pass (11/11 across 5 runs), and the test suite now verifies timeout propagation via request_kwargs_matcher.

Prior review resolution

Every issue I raised in prior rounds was addressed: backoff added to _make_grafana_request, _make_render_request, and query_echo_endpoint; Loki giveup lambda now has a None guard matching the Tempo pattern; timeout_seconds and max_retries have Pydantic gt=0/ge=1 constraints; retry tests mock time.sleep; and request_kwargs_matcher verifies actual timeout propagation in HTTP calls. The max_retries-as-total-attempts semantic was explicitly accepted by the maintainer.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My prior inline comment about the render timeout regression at _make_render_request (line 571) was incorrect — GrafanaDashboardConfig correctly overrides timeout_seconds to 60s, which preserves the pre-PR behavior. Please disregard that concern; the actual outstanding issue is the renderer probe timeout noted in the new inline comment.

Extended reasoning...

Overview

This PR adds configurable timeout_seconds and max_retries to GrafanaConfig and propagates those settings through Grafana Tempo, Loki, and dashboard toolsets, replacing hardcoded values. It spans 7 files including config models, API wrappers, toolset classes, and a new test file.

Correction of prior review comment

My April 21 inline comment on toolset_grafana.py:571 claimed the render timeout had regressed from 60s to 30s. That was wrong. The PR adds GrafanaDashboardConfig.timeout_seconds = Field(default=60), which overrides the base class 30s default specifically for the dashboard toolset. _make_render_request now resolves to config.timeout_seconds = 60, which is identical to the old hardcoded value. No regression occurred, and no render_timeout_seconds field is needed.

Newly identified issue (probe timeout)

The new bug report correctly identifies that the renderer probe requests in _try_add_render_tools (lines 134 and 157) previously used a hardcoded timeout=10 and now use config.timeout_seconds = 60. The 6x increase is meaningful for the misconfiguration case (enable_rendering=True, no renderer installed): startup can now wait up to 120s (2×60s) instead of 20s (2×10s). The inline comment on this PR describes a concrete fix (separate probe_timeout_seconds field or capping with min()).

Security risks

No security-sensitive code is touched. This is configuration and HTTP retry/timeout logic.

Level of scrutiny

Medium-to-high: this PR touches startup behavior and all Grafana HTTP call paths. The probe timeout change is a behavioral regression that affects startup time under a common misconfiguration scenario.

Other factors

All 11 LLM eval tests pass. CodeRabbit's pydantic bounds suggestion (gt=0, ge=1) was incorporated. The time.sleep mock was added to retry tests. The outstanding issue is limited in scope (2 lines in one method).

Comment thread holmes/plugins/toolsets/grafana/toolset_grafana.py
Changed timeout value for rendering version API request.
Comment thread holmes/plugins/toolsets/grafana/toolset_grafana.py
Comment thread holmes/plugins/toolsets/grafana/grafana_tempo_api.py
Comment thread holmes/plugins/toolsets/grafana/grafana_tempo_api.py
@moshemorad
moshemorad enabled auto-merge (squash) April 23, 2026 12:19
@moshemorad
moshemorad merged commit c04448a into master Apr 26, 2026
22 of 23 checks passed
@moshemorad
moshemorad deleted the grafana_timeout branch April 26, 2026 06:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: Make Tempo toolset timeout and retries configurable via GrafanaTempoConfig

2 participants