Skip to content

Add filtering and limit parameters to list active metrics tool - #1466

Merged
aantn merged 9 commits into
masterfrom
claude/slack-fix-holmes-errors-M9eVu
Feb 7, 2026
Merged

aantn merged 9 commits into
masterfrom
claude/slack-fix-holmes-errors-M9eVu

Conversation

@aantn

@aantn aantn commented Feb 1, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Enhanced the ListActiveMetrics tool in the Datadog metrics toolset with client-side filtering and result limiting capabilities to improve usability when dealing with large metric lists.

Key Changes

  • Added ACTIVE_METRICS_DEFAULT_LIMIT constant (500) to define default result limit
  • Introduced metric_name_filter parameter to filter metrics by name prefix or substring (case-insensitive)
  • Introduced limit parameter to control maximum number of returned metrics
  • Implemented client-side metric filtering logic that applies the name filter before returning results
  • Added truncation notice in output when results exceed the limit, informing users how to retrieve more results
  • Updated tool description to document the new limit and filtering capabilities
  • Updated get_parameterized_one_liner() to include filter and limit information in the summary

Implementation Details

  • Metric name filtering is performed case-insensitively using substring matching
  • Results are sorted alphabetically before applying the limit
  • If a filter is applied but returns no matches, an error is returned with helpful guidance
  • The truncation notice shows how many metrics matched the filter vs. how many are displayed
  • Default limit of 500 metrics prevents excessively large responses while remaining configurable

https://claude.ai/code/session_01PRaYUq5hoWotLDre1UGMTU

Summary by CodeRabbit

  • New Features

    • Added regex pattern filtering for Datadog metrics (case-insensitive) to narrow results.
    • Added a limit parameter to control maximum metrics returned, defaulting to 500.
    • Shows matched vs. returned counts and a truncation notice when results are filtered or limited.
  • Tests

    • Added Datadog tag to relevant test fixtures to reflect Datadog-related behavior.

@linux-foundation-easycla

linux-foundation-easycla Bot commented Feb 1, 2026 •

Copy link
Copy Markdown

CLA Not Signed

@netlify

netlify Bot commented Feb 1, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit a5301dd
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/698752df8fe20a000891909f
😎 Deploy Preview https://deploy-preview-1466--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai

coderabbitai Bot commented Feb 1, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

This PR adds regex-based filtering and a result limit to the DataDog ListActiveMetrics tool, introduces ACTIVE_METRICS_DEFAULT_LIMIT = 500, implements client-side filtering, sorting, truncation and related validation/errors, and updates test fixtures to include a datadog tag.

Changes

Cohort / File(s) Summary
Datadog Metrics Tool
holmes/plugins/toolsets/datadog/toolset_datadog_metrics.py
Added ACTIVE_METRICS_DEFAULT_LIMIT = 500; extended ListActiveMetrics with metric_name_filter (regex) and limit parameters; implemented client-side regex validation, filtering, total_matching computation, optional truncation/sorting, truncation notice, and updated one-liner generation to include the filter and limit.
Test Fixtures
tests/llm/fixtures/test_ask_holmes/110_cpu_graph_robusta_runner/test_case.yaml, tests/llm/fixtures/test_ask_holmes/111_disabled_datadog_traces/test_case.yaml
Added datadog tag to expected_output tags in both test cases.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested reviewers

  • arikalon1
  • nherment
  • moshemorad
🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 55.56% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title directly and clearly describes the main change: adding filtering and limit parameters to the list active metrics tool, which aligns with the primary feature additions in the changeset.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

Added limit and metric_name_filter parameters to prevent the tool from
returning unbounded data that causes token overflow errors in large
Datadog environments.

Changes:
- Added `limit` parameter (default 500) to cap number of returned metrics
- Added `metric_name_filter` parameter for client-side substring filtering
- Added truncation notice when results are limited
- Updated one-liner to include filter and limit info

Slack thread: https://robustaco.slack.com/archives/C08J4QBJ2H5/p1769151924987039?thread_ts=1769149831.651569&cid=C08J4QBJ2H5

https://claude.ai/code/session_01PRaYUq5hoWotLDre1UGMTU
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn force-pushed the claude/slack-fix-holmes-errors-M9eVu branch from c107358 to a7301be Compare February 1, 2026 14:20
@github-actions

github-actions Bot commented Feb 1, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for 234331c (built in 1m 19s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:234331c
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:234331c me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:234331c
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:234331c

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:234331c

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:234331c

@github-actions

github-actions Bot commented Feb 1, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 9e12cf5 (#21781860436)

✅ Results of HolmesGPT evals

Automatically triggered by commit 9e12cf5 on branch claude/slack-fix-holmes-errors-M9eVu

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.4s 6 10 $0.2280
✅ 101_loki_historical_logs_pod_deleted 46.6s 6 11 $0.2758
✅ 111_pod_names_contain_service 36.0s 6 12 $0.2455
✅ 112_find_pvcs_by_uuid 34.5s 7 7 $0.2357
✅ 12_job_crashing 37.1s 6 12 $0.2450
✅ 176_network_policy_blocking_traffic_no_runbooks 34.6s 5 13 $0.2464
✅ 24_misconfigured_pvc 36.0s 6 13 $0.2415
✅ 43_current_datetime_from_prompt 4.9s 1 — $0.1048
✅ 61_exact_match_counting 13.9s 3 2 $0.1437
Total 30.7s avg 5.1 avg 10.0 avg $1.9664
📜 Run @ 138f951 (#21779333080)

✅ Results of HolmesGPT evals

Automatically triggered by commit 138f951 on branch claude/slack-fix-holmes-errors-M9eVu

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 41/45 test cases were successful, 2 regressions, 1 skipped, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 27.9s 4 9 $0.2052
✅ 101_loki_historical_logs_pod_deleted 35.7s 5 7 $0.2191
✅ 110_cpu_graph_robusta_runner[0] 45.5s 8 15 $0.3074
✅ 110_cpu_graph_robusta_runner[10] 34.3s 8 9 $0.2467
✅ 110_cpu_graph_robusta_runner[11] 26.0s 5 7 $0.2105
✅ 110_cpu_graph_robusta_runner[12] 30.0s 6 8 $0.2249
❌ 110_cpu_graph_robusta_runner[13] 31.3s 7 10 $0.2458
✅ 110_cpu_graph_robusta_runner[14] 34.8s 8 9 $0.2468
✅ 110_cpu_graph_robusta_runner[15] 24.9s 5 8 $0.2052
✅ 110_cpu_graph_robusta_runner[1] 25.9s 5 7 $0.2072
✅ 110_cpu_graph_robusta_runner[2] 28.8s 5 8 $0.2076
✅ 110_cpu_graph_robusta_runner[3] 28.9s 5 8 $0.2099
✅ 110_cpu_graph_robusta_runner[4] 27.9s 5 7 $0.2153
✅ 110_cpu_graph_robusta_runner[5] 26.3s 5 7 $0.2035
✅ 110_cpu_graph_robusta_runner[6] 37.1s 7 10 $0.2479
✅ 110_cpu_graph_robusta_runner[7] 32.6s 6 10 $0.2407
✅ 110_cpu_graph_robusta_runner[8] 28.0s 6 8 $0.2148
✅ 110_cpu_graph_robusta_runner[9] 32.5s 6 10 $0.2360
✅ 111_disabled_datadog_traces 8.9s 1 — $0.1121
✅ 111_pod_names_contain_service 34.9s 5 12 $0.2351
✅ 112_find_pvcs_by_uuid 33.1s 5 7 $0.2452
✅ 12_job_crashing 37.7s 6 14 $0.2674
🚧 164_datadog_traces_coupon_code[0] — — — —
✅ 176_network_policy_blocking_traffic_no_runbooks 49.8s 8 16 $0.3007
✅ 24_misconfigured_pvc 39.0s 7 15 $0.2561
✅ 43_current_datetime_from_prompt 4.4s 1 — $0.1048
✅ 61_exact_match_counting 12.5s 3 2 $0.1415
✅ 91a_datadog_metrics_no_k8s 24.5s 5 7 $0.2018
✅ 91b_datadog_metrics_pod_exists 13.0s 3 2 $0.1487
✅ 91c_datadog_metrics_deployment[0] 19.3s 4 5 $0.1791
✅ 91c_datadog_metrics_deployment[1] 23.0s 5 6 $0.1946
✅ 91c_datadog_metrics_deployment[2] 27.0s 5 8 $0.2195
✅ 91d_datadog_metrics_historical_pod[0] 17.3s 3 3 $0.1621
✅ 91d_datadog_metrics_historical_pod[1] 16.4s 3 3 $0.1641
✅ 91d_datadog_metrics_historical_pod[2] 27.1s 4 7 $0.2177
✅ 91e_datadog_custom_metrics[0] 16.1s 3 4 $0.1591
✅ 91e_datadog_custom_metrics[1] 16.0s 3 4 $0.1560
❌ 91f_datadog_logs_historical_pod 42.7s 6 12 $0.2915
✅ 91g_datadog_metrics_mismatched_pod[0] 20.0s 4 5 $0.1793
✅ 91g_datadog_metrics_mismatched_pod[1] 19.9s 4 5 $0.1805
✅ 91h_datadog_logs_empty_query_with_url 16.3s 3 3 $0.1457
➖ 91i_datadog_metrics_empty_query_with_url — — — —
✅ 92_cpu_graph_conversation[0] 10.8s 2 1 $0.1378
✅ 92_cpu_graph_conversation[1] 10.0s 2 2 $0.1359
✅ 92_cpu_graph_conversation[2] 10.8s 2 1 $0.1381
Total 25.8s avg 4.7 avg 7.3 avg $8.7691

⚠️ 2 Failures Detected

📜 Run @ a08226e (#21778850124)

✅ Results of HolmesGPT evals

Automatically triggered by commit a08226e on branch claude/slack-fix-holmes-errors-M9eVu

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.0s 5 11 $0.2383
✅ 101_loki_historical_logs_pod_deleted 36.7s 6 8 $0.2261
✅ 111_pod_names_contain_service 34.8s 6 11 $0.2322
✅ 112_find_pvcs_by_uuid 40.1s 7 9 $0.2698
✅ 12_job_crashing 43.3s 7 17 $0.2818
✅ 176_network_policy_blocking_traffic_no_runbooks 35.3s 5 15 $0.2723
✅ 24_misconfigured_pvc 38.6s 7 13 $0.2436
✅ 43_current_datetime_from_prompt 5.6s 1 — $0.0124
✅ 61_exact_match_counting 17.0s 4 4 $0.1596
Total 31.5s avg 5.3 avg 11.0 avg $1.9362
📜 Run @ e1aa79b (#21769148560)

✅ Results of HolmesGPT evals

Automatically triggered by commit e1aa79b on branch claude/slack-fix-holmes-errors-M9eVu

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 33.4s 5 11 $0.2300
✅ 101_loki_historical_logs_pod_deleted 54.0s 7 12 $0.3086
✅ 111_pod_names_contain_service 34.3s 5 11 $0.2211
✅ 112_find_pvcs_by_uuid 37.0s 7 9 $0.2554
✅ 12_job_crashing 34.0s 5 12 $0.2405
✅ 176_network_policy_blocking_traffic_no_runbooks 37.7s 5 15 $0.2600
✅ 24_misconfigured_pvc 33.4s 5 13 $0.2290
✅ 43_current_datetime_from_prompt 5.0s 1 — $0.1051
✅ 61_exact_match_counting 16.2s 4 4 $0.1589
Total 31.7s avg 4.9 avg 10.9 avg $2.0084
📜 Run @ a09dac8 (#21744194535)

✅ Results of HolmesGPT evals

Automatically triggered by commit a09dac8 on branch claude/slack-fix-holmes-errors-M9eVu

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 34.1s 5 11 $0.2286
✅ 101_loki_historical_logs_pod_deleted 50.4s 6 12 $0.2873
✅ 111_pod_names_contain_service 32.2s 5 11 $0.2207
✅ 112_find_pvcs_by_uuid 33.6s 6 7 $0.2476
✅ 12_job_crashing 34.5s 5 12 $0.2417
✅ 176_network_policy_blocking_traffic_no_runbooks 41.5s 6 16 $0.2836
✅ 24_misconfigured_pvc 30.7s 5 11 $0.2117
✅ 43_current_datetime_from_prompt 4.9s 1 — $0.1050
✅ 61_exact_match_counting 17.1s 4 4 $0.1599
Total 31.0s avg 4.8 avg 10.5 avg $1.9862

✅ Results of HolmesGPT evals

Automatically triggered by commit a5301dd on branch claude/slack-fix-holmes-errors-M9eVu

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 36.2s 6 11 $0.2353
✅ 101_loki_historical_logs_pod_deleted 37.6s 6 8 $0.2320
✅ 111_pod_names_contain_service 31.6s 5 11 $0.2170
✅ 112_find_pvcs_by_uuid 38.2s 7 8 $0.2683
✅ 12_job_crashing 26.5s 4 7 $0.1960
✅ 176_network_policy_blocking_traffic_no_runbooks 50.1s 7 18 $0.3064
✅ 24_misconfigured_pvc 34.6s 5 15 $0.2382
✅ 43_current_datetime_from_prompt 5.4s 1 — $0.1047
✅ 61_exact_match_counting 17.5s 4 4 $0.1586
Total 30.9s avg 5.0 avg 10.2 avg $1.9564
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/slack-fix-holmes-errors-M9eVu -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/slack-fix-holmes-errors-M9eVu -f markers=regression -f filter=

@github-actions

github-actions Bot commented Feb 1, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 9.93s 9.81s +1.2%
Warm Mean 4.63s 4.56s +1.5%
Warm Min 4.59s 4.49s
Warm Max 4.66s 4.75s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 23.92s 26.73s -10.5%
Warm Mean 7.36s 7.51s -2.0%
Warm Min 7.33s 7.23s
Warm Max 7.37s 7.89s

PR: 234331ca | Master: 76dbfc7e | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@holmes/plugins/toolsets/datadog/toolset_datadog_metrics.py`:
- Around line 156-161: The current limit application in the metrics retrieval
block (using params.get("limit", ACTIVE_METRICS_DEFAULT_LIMIT)) lets 0 or
negative values return all metrics; change the logic in the limit handling
inside the function that processes params (the block referencing limit,
ACTIVE_METRICS_DEFAULT_LIMIT, and metrics) to enforce limit > 0 and if limit is
missing or <= 0 fall back to ACTIVE_METRICS_DEFAULT_LIMIT, then sort and slice
metrics by that positive limit; update only the limit-check branch so invalid
limits no longer produce unbounded results.
🧹 Nitpick comments (1)
holmes/plugins/toolsets/datadog/toolset_datadog_metrics.py (1)

78-82: Nit: description says "prefix or substring" but only substring matching is implemented.

The in operator on line 145 performs substring matching, which inherently includes prefix matching. Saying "prefix or substring" might imply two distinct modes. Consider simplifying to just "substring" for accuracy.

Proposed fix
                 "metric_name_filter": ToolParameter(
-                    description="Filter metrics by name prefix or substring. Example: 'kubernetes' matches 'kubernetes.cpu.usage', 'system.kubernetes.memory'. Use this to narrow down large metric lists.",
+                    description="Filter metrics by name substring (case-insensitive). Example: 'kubernetes' matches 'kubernetes.cpu.usage', 'system.kubernetes.memory'. Use this to narrow down large metric lists.",
                     type="string",
                     required=False,
                 ),

Comment thread holmes/plugins/toolsets/datadog/toolset_datadog_metrics.py Outdated
@aantn

aantn commented Feb 6, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
marker: datadog

@github-actions

This comment was marked as outdated.

aantn and others added 2 commits February 7, 2026 10:46
#1507)

Switch all Datadog eval toolsets, helper scripts, and unit tests from
the EU endpoint (api.datadoghq.eu) to US5 (api.us5.datadoghq.com) to
match our current API keys. Remove test 93_calling_datadog which used
stale mock data files that no longer match the current toolset API.

https://claude.ai/code/session_011dbj3Cy5ApJHTWPnQ6Mpq3
Signed-off-by: Claude <noreply@anthropic.com>

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Switched default Datadog endpoints from EU to US5 across test
configurations for metrics, logs, and traces.

* **Tests**
  * Simplified and pruned multiple test fixtures and prompts.
  * Removed several Datadog conversation and log fixture files.
  * Added a new test scenario focused on Datadog metrics-only behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In
`@tests/llm/fixtures/test_ask_holmes/91a_datadog_metrics_no_k8s/test_case.yaml`:
- Line 7: The namespace "space-projects" in the bash invocation does not follow
the required app-<testid> pattern; update the first argument to "app-91a" and
make the pod names unique for this test by appending the test id (e.g., change
"spaceship-launch-counter" to "spaceship-launch-counter-91a" for both pod args
or otherwise ensure uniqueness), and update any user prompt or test metadata
that references those names to match the new "app-91a" and pod name(s).
🧹 Nitpick comments (1)
tests/llm/fixtures/test_ask_holmes/91a_datadog_metrics_no_k8s/test_case.yaml (1)

11-17: Expected output uses a generic structural pattern rather than a query-specific value.

The expected_output checks for the embed format but doesn't verify any specific metric data that can only be discovered by actually querying Datadog (e.g., a specific metric value, time range, or count). Consider adding a check for a discoverable value like the metric name container.cpu.usage in the output to strengthen the assertion.

As per coding guidelines, "For LLM eval tests, use specific values in expected_output that can only be discovered by querying (like unique IDs, injected error codes, exact counts) rather than generic output patterns."

Tagged 13 Datadog-related eval tests with 'regression' marker so they
run automatically in CI/CD regression testing.

https://claude.ai/code/session_01PRaYUq5hoWotLDre1UGMTU
Signed-off-by: Claude <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
@aantn

aantn commented Feb 7, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
marker: datadog
iterations: 5

@github-actions

This comment was marked as outdated.

…metric filter

- Remove regression tag from all Datadog eval test cases (per review feedback)
- Change metric_name_filter from substring match to regex pattern matching
- Add error handling for invalid regex patterns

https://claude.ai/code/session_01PRaYUq5hoWotLDre1UGMTU
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@holmes/plugins/toolsets/datadog/toolset_datadog_metrics.py`:
- Around line 157-162: Replace the inconsistent use of the data= field for error
messages with error= so error paths are uniform: in the block that returns a
StructuredToolResult when no metrics match (the if not metrics branch) change
the returned StructuredToolResult to set error= with the existing message
instead of data=; keep the same status=StructuredToolResultStatus.ERROR and
params=params and retain the message text referencing metric_name_filter so it
matches the regex error path that already uses error=.

Comment thread holmes/plugins/toolsets/datadog/toolset_datadog_metrics.py
When limit=0 or negative, now falls back to ACTIVE_METRICS_DEFAULT_LIMIT
instead of returning all metrics unbounded.

https://claude.ai/code/session_01PRaYUq5hoWotLDre1UGMTU
Signed-off-by: Claude <noreply@anthropic.com>
@aantn

aantn commented Feb 7, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
marker: datadog

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@holmes/plugins/toolsets/datadog/toolset_datadog_metrics.py`:
- Around line 256-260: get_parameterized_one_liner currently reads the raw limit
from params and can show invalid values (e.g., 0 or negative) even though
_invoke falls back to ACTIVE_METRICS_DEFAULT_LIMIT; modify
get_parameterized_one_liner to validate/coerce the limit the same way _invoke
does (ensure limit is a positive integer, otherwise use
ACTIVE_METRICS_DEFAULT_LIMIT) before building the one-liner string for
toolset_name_for_one_liner(self.toolset.name): List Active Metrics (...,
limit={limit}).
🧹 Nitpick comments (1)
holmes/plugins/toolsets/datadog/toolset_datadog_metrics.py (1)

142-153: Consider adding a timeout or complexity guard for user-supplied regex.

The regex pattern is compiled directly from metric_name_filter without any protection against catastrophic backtracking (ReDoS). While the input comes from the LLM rather than a direct end-user, a complex pattern like (a+)+$ applied against a large metric list could cause significant CPU usage.

A simple mitigation would be to use re.search with a timeout (Python 3.11+ has no native timeout for re, but you could use the regex library), or fall back to simple substring matching and only use regex when special characters are detected.

Given that the LLM is the source of these patterns, this is low risk but worth noting.

Comment thread holmes/plugins/toolsets/datadog/toolset_datadog_metrics.py
@github-actions

github-actions Bot commented Feb 7, 2026

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ✅ Completed successfully


🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Branch claude/slack-fix-holmes-errors-M9eVu
Model opus-4.5
Markers datadog
Iterations 1
Duration 4m 52s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 34/36 test cases were successful, 0 regressions, 1 skipped, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 110_cpu_graph_robusta_runner[0] 32.9s 6 10 $0.2355
✅ 110_cpu_graph_robusta_runner[10] 31.8s 6 9 $0.2276
✅ 110_cpu_graph_robusta_runner[11] 31.1s 6 10 $0.2293
✅ 110_cpu_graph_robusta_runner[12] 27.5s 5 8 $0.2076
✅ 110_cpu_graph_robusta_runner[13] 40.0s 9 12 $0.2711
✅ 110_cpu_graph_robusta_runner[14] 32.7s 7 8 $0.2354
✅ 110_cpu_graph_robusta_runner[15] 25.9s 5 8 $0.2081
✅ 110_cpu_graph_robusta_runner[1] 73.5s 6 9 $0.2300
✅ 110_cpu_graph_robusta_runner[2] 40.5s 8 12 $0.2677
✅ 110_cpu_graph_robusta_runner[3] 36.3s 7 11 $0.2468
✅ 110_cpu_graph_robusta_runner[4] 27.6s 5 8 $0.2105
✅ 110_cpu_graph_robusta_runner[5] 31.4s 6 9 $0.2249
✅ 110_cpu_graph_robusta_runner[6] 37.8s 7 10 $0.2468
✅ 110_cpu_graph_robusta_runner[7] 29.6s 6 8 $0.2181
✅ 110_cpu_graph_robusta_runner[8] 30.0s 6 9 $0.2202
✅ 110_cpu_graph_robusta_runner[9] 29.0s 5 8 $0.2098
✅ 111_disabled_datadog_traces 8.2s 1 — $0.1114
🚧 164_datadog_traces_coupon_code[0] — — — —
✅ 91a_datadog_metrics_no_k8s 24.9s 5 7 $0.1977
✅ 91b_datadog_metrics_pod_exists 13.6s 3 2 $0.1485
✅ 91c_datadog_metrics_deployment[0] 19.1s 4 5 $0.1792
✅ 91c_datadog_metrics_deployment[1] 24.4s 5 7 $0.1999
✅ 91c_datadog_metrics_deployment[2] 28.4s 5 7 $0.2115
✅ 91d_datadog_metrics_historical_pod[0] 16.8s 3 3 $0.1626
✅ 91d_datadog_metrics_historical_pod[1] 21.4s 4 4 $0.1759
✅ 91d_datadog_metrics_historical_pod[2] 27.5s 4 6 $0.2206
✅ 91e_datadog_custom_metrics[0] 20.8s 4 5 $0.1711
✅ 91e_datadog_custom_metrics[1] 16.7s 3 4 $0.1561
✅ 91f_datadog_logs_historical_pod 49.2s 7 12 $0.3207
✅ 91g_datadog_metrics_mismatched_pod[0] 15.0s 3 3 $0.1576
✅ 91g_datadog_metrics_mismatched_pod[1] 20.3s 4 5 $0.1792
✅ 91h_datadog_logs_empty_query_with_url 16.6s 3 3 $0.1446
➖ 91i_datadog_metrics_empty_query_with_url — — — —
✅ 92_cpu_graph_conversation[0] 10.9s 2 1 $0.1376
✅ 92_cpu_graph_conversation[1] 10.7s 2 2 $0.1390
✅ 92_cpu_graph_conversation[2] 11.6s 2 2 $0.1400
Total 26.9s avg 4.8 avg 6.9 avg $6.8426
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/slack-fix-holmes-errors-M9eVu -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/slack-fix-holmes-errors-M9eVu -f markers=regression -f filter=

@aantn
aantn enabled auto-merge (squash) February 7, 2026 15:52
@aantn
aantn merged commit d8ef34c into master Feb 7, 2026
18 of 20 checks passed
@aantn
aantn deleted the claude/slack-fix-holmes-errors-M9eVu branch February 7, 2026 16:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants