Skip to content

Improve jq filtering - #1615

Open
aantn wants to merge 19 commits into
masterfrom
claude/revert-except-jq-errors-A7N9q
Open

aantn wants to merge 19 commits into
masterfrom
claude/revert-except-jq-errors-A7N9q

Conversation

@aantn

@aantn aantn commented Feb 22, 2026 •

Copy link
Copy Markdown
Collaborator

https://claude.ai/code/session_01N5rWgtZrghkJVJtL3xdEGZ

Summary by CodeRabbit

  • Tests
    • Added a new jq-filter test marker and applied it across Elasticsearch, Grafana, Prometheus and HTTP test cases.
  • Improvements
    • JSON filtering failures now surface structured error details plus a truncated preview of the original response to aid debugging and reduce noisy output.

claude and others added 13 commits February 21, 2026 09:10
…l scenarios

Adds a new `toolsets_matrix` field to test_case.yaml that lists multiple
toolset config filenames. Each file creates a separate test variant with
the same user_prompt, expected_output, and infrastructure - only the
toolset configuration changes. This enables comparing builtin toolsets
vs HTTP toolsets vs MCP on identical scenarios.

Changes:
- HolmesTestCase: add toolsets_matrix, toolsets_config_name, toolsets_config_path fields
- MockHelper: add _expand_toolsets_matrix() post-processing step after test case loading
- MockToolsetManager: accept toolsets_config_path to override default toolsets.yaml resolution
- test_ask_holmes.py, test_investigate.py: pass toolsets_config_path through

Example test_case.yaml usage:
  toolsets_matrix:
    - toolsets_builtin.yaml
    - toolsets_http.yaml

Produces test IDs like: test_ask_holmes[01_test[builtin]-model-env]

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>
Adds toolsets_http.yaml to 9 Datadog eval tests (91a-91h, 164) with HTTP
toolset configs that hit the same Datadog APIs using the generic HTTP
toolset instead of the native Datadog toolsets. Each test now runs twice:
once with the builtin toolset and once with the HTTP toolset.

HTTP toolset auth uses DD-API-KEY header + DD-APPLICATION-KEY via
default_headers, with llm_instructions documenting the API endpoints.

Coverage:
- Metrics API (GET /api/v1/metrics, /api/v1/query, /api/v2/metrics): 91a-91e, 91g
- Logs API (POST /api/v2/logs/events/search): 91f, 91h
- Traces API (POST /api/v2/spans/events/search, /analytics/aggregate): 164

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>
…rect

When a jq filter fails (e.g. `.metrics[]` on a null field), instead of
returning an empty ERROR that gives the LLM zero information, return the
original response (truncated to depth 2) alongside the error hint. This
lets the LLM see the actual response shape and retry with a null-safe
expression like `(.metrics // [])[]`.

Also adds null-safe jq guidance to HTTP toolset instructions.

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>
The raw_response_preview is now serialized to a JSON string and
truncated at 2000 characters, preventing large API responses (e.g.
thousands of metrics) from overwhelming the LLM context window.

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>
Reduce from 17 prompt variants to one representative prompt.

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>
…trics

The 6 embed tests (91a, 91b, 91c, 91d, 91e, 91g) expect datadogql
embeds with tool_name "query_datadog_metrics", which only exists in the
builtin Python datadog/metrics toolset. The HTTP toolset has no such
tool, so the LLM loops forever trying to produce an impossible output,
causing CI to hang for 3+ hours.

Kept toolsets_matrix on 91f (logs) and 164 (traces) since those tests
don't require the embed format.

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>
Keep only jq error handling improvements:
- json_filter_mixin.py: Make jq errors non-fatal, return raw data preview + error hint
- instructions.jinja2: Add null-safe jq pattern guidance
- test_json_filter_mixin.py: Updated tests for new jq error behavior

Revert all other changes (toolsets_matrix, test framework changes) back to master.

https://claude.ai/code/session_01N5rWgtZrghkJVJtL3xdEGZ
Signed-off-by: Claude <noreply@anthropic.com>
@aantn aantn changed the title Remove toolsets matrix expansion feature Improve jq filtering Feb 22, 2026
@coderabbitai

coderabbitai Bot commented Feb 22, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Added a pytest marker jq-filter and applied it to multiple test fixtures (Grafana, Elasticsearch, Prometheus). Updated JsonFilterMixin error handling to return a structured jq-error payload including the jq expression and a depth-/size-limited preview of the original data on jq failures.

Changes

Cohort / File(s) Summary
Pytest configuration
pyproject.toml
Added pytest marker: jq-filter: Tests exercising JsonFilterMixin toolsets (jq/max_depth filtering on Elasticsearch, Grafana, Prometheus, HTTP)
Grafana test cases
tests/llm/fixtures/test_ask_holmes/177_grafana_home_dashboard/test_case.yaml, tests/llm/fixtures/test_ask_holmes/178_grafana_search_dashboard_query/test_case.yaml, tests/llm/fixtures/test_ask_holmes/179_grafana_big_dashboard_query/test_case.yaml
Added jq-filter tag to test metadata tags.
Elasticsearch test cases
tests/llm/fixtures/test_ask_holmes/183a_elasticsearch_cluster_health/test_case.yaml, tests/llm/fixtures/test_ask_holmes/183b_elasticsearch_index_discovery/test_case.yaml, tests/llm/fixtures/test_ask_holmes/183c_elasticsearch_log_search/test_case.yaml, tests/llm/fixtures/test_ask_holmes/183d_elasticsearch_aggregation/test_case.yaml, tests/llm/fixtures/test_ask_holmes/183e_elasticsearch_field_mappings/test_case.yaml, tests/llm/fixtures/test_ask_holmes/183f_elasticsearch_shard_filtering/test_case.yaml, tests/llm/fixtures/test_ask_holmes/183g_elasticsearch_index_stats/test_case.yaml, tests/llm/fixtures/test_ask_holmes/184_elasticsearch_index_explosion/test_case.yaml, tests/llm/fixtures/test_ask_holmes/185_elasticsearch_cross_region_search/test_case.yaml, tests/llm/fixtures/test_ask_holmes/186_elasticsearch_shard_explosion/test_case.yaml, tests/llm/fixtures/test_ask_holmes/187_elasticsearch_disk_space/test_case.yaml, tests/llm/fixtures/test_ask_holmes/188_elasticsearch_mapping_explosion/test_case.yaml, tests/llm/fixtures/test_ask_holmes/189_elasticsearch_timeseries_gap/test_case.yaml, tests/llm/fixtures/test_ask_holmes/190_elasticsearch_cross_service_correlation/test_case.yaml, tests/llm/fixtures/test_ask_holmes/191_elasticsearch_query_profile/test_case.yaml, tests/llm/fixtures/test_ask_holmes/193_elasticsearch_large_mapping_search/test_case.yaml, tests/llm/fixtures/test_ask_holmes/195_elasticsearch_trace_large_fields/test_case.yaml
Added jq-filter tag to expected_output/tags in each test case.
Prometheus test cases
tests/llm/fixtures/test_ask_holmes/211_prometheus_alerting_rules/test_case.yaml
Added jq-filter tag to expected_output.tags.
JsonFilterMixin
holmes/plugins/toolsets/json_filter_mixin.py
On jq evaluation error, return structured payload instead of raw error: includes jq_error, jq_expression, and a depth-/size-limited raw_response_preview of the original data; replaced previous (None, error) error return path.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

  • Json mixin for tools #1280: Modifies JsonFilterMixin jq error handling and returned payload shape (overlaps holmes/plugins/toolsets/json_filter_mixin.py changes).

Suggested labels

codex

Suggested reviewers

  • Sheeproid
  • moshemorad
🚥 Pre-merge checks | ✅ 1 | ❌ 2

❌ Failed checks (1 warning, 1 inconclusive)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
Title check ❓ Inconclusive The title 'Improve jq filtering' is vague and generic, lacking specificity about what improvement was made or which aspects were modified. Consider using a more descriptive title that specifies the improvement, such as 'Add structured error payloads to jq filtering' or 'Improve jq error handling with detailed error context'.
✅ Passed checks (1 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In
`@tests/llm/fixtures/test_ask_holmes/110_cpu_graph_robusta_runner/test_case.yaml`:
- Line 13: Update the test prompts so they explicitly request a graph and the
datadogql embed: replace the string "Show me memory metrics for robusta-runner"
with a prompt that asks for a rendered time-series graph (e.g., "Show me a
time-series graph of memory metrics for robusta-runner, render using datadogql")
and likewise change "What's the memory utilization of the robusta-runner
deployment?" to explicitly request a graph render (e.g., "Render a graph showing
memory utilization for the robusta-runner deployment using datadogql"); ensure
these prompt strings (the exact text instances above) match what the test
expects so Holmes will produce the datadogql embed.
- Line 2: The test prompt in test_case.yaml currently says "generate me a graph
of robusta-runner", which is too generic; update the prompt string to explicitly
request the CPU metric (e.g., ask for a "CPU usage graph" or "CPU metric for
robusta-runner") so the test matches the expected query_datadog_metrics embed
check; locate and replace the prompt value in the test_case.yaml fixture to
mention CPU (the prompt string to change is the one currently set to "generate
me a graph of robusta-runner").
- Around line 1-17: Rename the test folder from 110_cpu_graph_robusta_runner to
110_memory_graph_robusta_runner and update any internal references to that
folder (e.g., CI/test manifests or test discovery entries) so the test name
matches the prompts in test_case.yaml which target memory graphs; ensure the
directory name change is reflected wherever the folder is referenced
(tests/llm/fixtures/... and any sibling test index or metadata) so the suite and
naming convention remain consistent with memory-focused tests like
34_memory_graph and 70_memory_leak_detection.

Comment thread tests/llm/fixtures/test_ask_holmes/110_cpu_graph_robusta_runner/test_case.yaml Outdated
Comment thread tests/llm/fixtures/test_ask_holmes/110_cpu_graph_robusta_runner/test_case.yaml Outdated
Comment thread tests/llm/fixtures/test_ask_holmes/110_cpu_graph_robusta_runner/test_case.yaml Outdated
Base automatically changed from claude/evals-toolset-matrix-mwdXc to master February 22, 2026 08:19
New pytest marker `jq-filter` covers all 21 evals that use toolsets with
JsonFilterMixin (jq/max_depth parameters). This enables targeted testing
of jq error handling changes:

- Elasticsearch (17 evals): 183a-g, 184-191, 193, 195
  - Simple tests (negative examples): cluster health, index discovery, etc.
  - Stress tests (positive examples): index/shard/mapping explosion, large mappings
- Grafana dashboards (3 evals): 177, 178, 179
- Prometheus alerting rules (1 eval): 211

Run with: poetry run pytest -m jq-filter --no-cov

https://claude.ai/code/session_01N5rWgtZrghkJVJtL3xdEGZ
Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Feb 22, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 5b0c907 (#22807502364)

✅ Results of HolmesGPT evals

Automatically triggered by commit 5b0c907 on branch claude/revert-except-jq-errors-A7N9q (labels: evals-id-195_elasticsearch_trace_large_fields)

View workflow logs

⚠️ No eval report was generated.

📜 Run @ 1414b95 (#22280375835)

✅ Results of HolmesGPT evals

Automatically triggered by commit 1414b95 on branch claude/revert-except-jq-errors-A7N9q (labels: evals-id-195_elasticsearch_trace_large_fields)

View workflow logs

⚠️ No eval report was generated.

📜 Run @ e16755b (#22280363935)

✅ Results of HolmesGPT evals

Automatically triggered by commit e16755b on branch claude/revert-except-jq-errors-A7N9q (labels: evals-tag-jq-filter, evals-id-195_elasticsearch_trace_large_fields)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 1/1 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 195_elasticsearch_trace_large_fields 45.9s 9 13 $0.3047
Total 45.9s avg 9.0 avg 13.0 avg $0.3047
📜 Run @ e16755b (#22277829898)

✅ Results of HolmesGPT evals

Automatically triggered by commit e16755b on branch claude/revert-except-jq-errors-A7N9q (labels: evals-tag-jq-filter)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 29/30 test cases were successful, 0 regressions, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 34.4s 6 11 $0.2446
✅ 101_loki_historical_logs_pod_deleted 45.7s 6 10 $0.2682
✅ 111_pod_names_contain_service 38.6s 6 14 $0.2764
✅ 112_find_pvcs_by_uuid 38.6s 7 9 $0.2755
✅ 12_job_crashing 45.5s 8 16 $0.3257
✅ 176_network_policy_blocking_traffic_no_runbooks 44.8s 6 17 $0.3002
✅ 177_grafana_home_dashboard 14.6s 3 3 $0.1612
✅ 178_grafana_search_dashboard_query 21.0s 4 5 $0.1861
✅ 179_grafana_big_dashboard_query 20.2s 4 5 $0.1847
✅ 183a_elasticsearch_cluster_health 13.6s 3 3 $0.1612
✅ 183b_elasticsearch_index_discovery 13.0s 3 3 $0.1608
✅ 183c_elasticsearch_log_search 28.8s 6 8 $0.2319
✅ 183d_elasticsearch_aggregation 35.6s 8 10 $0.2603
✅ 183e_elasticsearch_field_mappings 14.7s 3 3 $0.1640
✅ 183f_elasticsearch_shard_filtering 16.2s 3 3 $0.1679
✅ 183g_elasticsearch_index_stats 15.7s 3 3 $0.1711
🚧 184_elasticsearch_index_explosion — — — —
✅ 185_elasticsearch_cross_region_search 116.0s 18 27 $0.6083
✅ 186_elasticsearch_shard_explosion 18.7s 3 3 $0.1753
✅ 187_elasticsearch_disk_space 15.8s 3 3 $0.1682
✅ 188_elasticsearch_mapping_explosion 26.8s 4 9 $0.2125
✅ 189_elasticsearch_timeseries_gap 39.2s 6 11 $0.2852
✅ 190_elasticsearch_cross_service_correlation 19.9s 3 4 $0.1872
✅ 191_elasticsearch_query_profile 18.9s 3 3 $0.1773
✅ 193_elasticsearch_large_mapping_search 14.8s 3 3 $0.1621
✅ 195_elasticsearch_trace_large_fields 48.6s 9 12 $0.3067
✅ 211_prometheus_alerting_rules 14.6s 3 3 $0.1825
✅ 24_misconfigured_pvc 33.3s 5 15 $0.2472
✅ 43_current_datetime_from_prompt 4.3s 1 — $0.1110
✅ 61_exact_match_counting 12.5s 3 2 $0.1500
Total 28.4s avg 4.9 avg 7.8 avg $6.5131
📜 Run @ 914ac5a (#22276894714)

✅ Results of HolmesGPT evals

Automatically triggered by commit 914ac5a on branch claude/revert-except-jq-errors-A7N9q (labels: evals-tag-jq-filter)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 30/30 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 31.5s 5 11 $0.2398
✅ 101_loki_historical_logs_pod_deleted 51.3s 6 16 $0.3224
✅ 111_pod_names_contain_service 34.4s 6 11 $0.2434
✅ 112_find_pvcs_by_uuid 32.4s 6 8 $0.2459
✅ 12_job_crashing 31.5s 5 11 $0.2376
✅ 176_network_policy_blocking_traffic_no_runbooks 48.5s 8 17 $0.2989
✅ 177_grafana_home_dashboard 14.3s 3 3 $0.1609
✅ 178_grafana_search_dashboard_query 20.6s 4 5 $0.1914
✅ 179_grafana_big_dashboard_query 19.4s 4 5 $0.1805
✅ 183a_elasticsearch_cluster_health 12.9s 3 3 $0.1590
✅ 183b_elasticsearch_index_discovery 12.2s 3 3 $0.1610
✅ 183c_elasticsearch_log_search 28.5s 6 8 $0.2280
✅ 183d_elasticsearch_aggregation 30.2s 7 9 $0.2347
✅ 183e_elasticsearch_field_mappings 14.9s 3 3 $0.1641
✅ 183f_elasticsearch_shard_filtering 14.9s 3 3 $0.1658
✅ 183g_elasticsearch_index_stats 15.2s 3 3 $0.1709
✅ 184_elasticsearch_index_explosion 16.8s 4 4 $0.1779
✅ 185_elasticsearch_cross_region_search 86.4s 15 26 $0.7803
✅ 186_elasticsearch_shard_explosion 18.1s 3 4 $0.1783
✅ 187_elasticsearch_disk_space 15.6s 3 3 $0.1686
✅ 188_elasticsearch_mapping_explosion 33.8s 6 10 $0.2608
✅ 189_elasticsearch_timeseries_gap 42.2s 7 9 $0.3189
✅ 190_elasticsearch_cross_service_correlation 19.9s 3 4 $0.1966
✅ 191_elasticsearch_query_profile 20.1s 3 3 $0.1776
✅ 193_elasticsearch_large_mapping_search 15.0s 3 3 $0.1611
✅ 195_elasticsearch_trace_large_fields 60.2s 12 15 $0.3654
✅ 211_prometheus_alerting_rules 18.4s 4 4 $0.2003
✅ 24_misconfigured_pvc 35.0s 6 14 $0.2552
✅ 43_current_datetime_from_prompt 4.0s 1 — $0.1114
✅ 61_exact_match_counting 12.5s 3 2 $0.1503
Total 27.0s avg 4.9 avg 7.6 avg $6.9069

✅ Results of HolmesGPT evals

Automatically triggered by commit 3289fb9 on branch claude/revert-except-jq-errors-A7N9q (labels: evals-id-195_elasticsearch_trace_large_fields)

View workflow logs

⚠️ No eval report was generated.

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/revert-except-jq-errors-A7N9q -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals in automatic regression runs:

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID

Examples: evals-tag-easy, evals-id-09_crashpod

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, jq-filter, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/revert-except-jq-errors-A7N9q -f markers=regression -f filter=

@github-actions

github-actions Bot commented Feb 22, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 85f07893 (built in 6m 25s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:85f07893
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:85f07893 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:85f07893
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:85f07893
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:85f07893
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:85f07893 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:85f07893
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:85f07893

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:85f07893 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:85f07893

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:85f07893 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:85f07893

@netlify

netlify Bot commented Feb 22, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 3289fb9
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69ad8684e9b9810008629c41
😎 Deploy Preview https://deploy-preview-1615--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Feb 22, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 11.24s 10.76s +4.5%
Warm Mean 5.08s 4.95s +2.6%
Warm Min 5.01s 4.91s
Warm Max 5.23s 4.99s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 29.09s 27.09s +7.4%
Warm Mean 7.51s 7.35s +2.1%
Warm Min 7.29s 7.28s
Warm Max 7.72s 7.45s

PR: 85f07893 | Master: fb98d099 | Iterations: 5

@aantn

aantn commented Feb 22, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
marker: regression or jq-filter
branch: master

@github-actions

This comment was marked as outdated.

@aantn

aantn commented Feb 22, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
filter: 177 or 178 or 179 or 183a or 183b or 183c or 183d or 183e or 183f or 183g or 184 or 185 or 186 or 187 or 188 or 189 or 190 or 191 or 193 or 195 or 211
branch: master

The null-safe jq patterns (e.g., (.key // [])[] instead of .key[])
were added as LLM hints but are being reverted as part of the
jq-error-handling split. The jq_error field now returns just the
raw error string without prescriptive advice.

https://claude.ai/code/session_01N5rWgtZrghkJVJtL3xdEGZ
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@holmes/plugins/toolsets/json_filter_mixin.py`:
- Around line 98-101: The preview generation can raise TypeError when json.dumps
is given non-serializable objects in _filter_result_data; wrap the
json.dumps(truncated, ...) call (and the length check/substring logic for
preview_str) in a try/except that catches TypeError (and optionally ValueError),
and on exception produce a safe fallback preview (e.g., use repr(truncated) or
str(truncated) truncated to max_preview_chars with the same "…(truncated)"
suffix); ensure you update references to preview_str and truncated handling so
the rest of _filter_result_data continues using the fallback string rather than
letting the exception propagate.

Comment thread holmes/plugins/toolsets/json_filter_mixin.py
@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ✅ Completed successfully


🧪 Manual Eval Results

Parameter Value
Triggered via /eval on branch master
Branch master
Model opus-4.5
Tags all LLM tests
Filter (-k) 177 or 178 or 179 or 183a or 183b or 183c or 183d or 183e or 183f or 183g or 184 or 185 or 186 or 187 or 188 or 189 or 190 or 191 or 193 or 195 or 211
Iterations 1
Duration 7m 52s
Workflow View logs | Rerun

Results of HolmesGPT evals (branch: master)

  • ask_holmes: 21/22 test cases were successful, 0 regressions, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 177_grafana_home_dashboard 15.8s 3 3 $0.1622
✅ 178_grafana_search_dashboard_query 19.3s 4 5 $0.1850
✅ 179_grafana_big_dashboard_query 24.7s 5 6 $0.1993
✅ 183a_elasticsearch_cluster_health 11.8s 3 3 $0.1597
✅ 183b_elasticsearch_index_discovery 12.7s 3 3 $0.1608
✅ 183c_elasticsearch_log_search 32.8s 7 9 $0.2512
✅ 183d_elasticsearch_aggregation 33.3s 8 10 $0.2487
✅ 183e_elasticsearch_field_mappings 14.0s 3 3 $0.1635
✅ 183f_elasticsearch_shard_filtering 14.4s 3 3 $0.1671
✅ 183g_elasticsearch_index_stats 14.4s 3 3 $0.1693
🚧 184_elasticsearch_index_explosion — — — —
✅ 185_elasticsearch_cross_region_search 97.5s 16 26 $0.8535
✅ 186_elasticsearch_shard_explosion 18.1s 3 4 $0.1783
✅ 187_elasticsearch_disk_space 15.9s 3 3 $0.1687
✅ 188_elasticsearch_mapping_explosion 31.0s 5 10 $0.2510
✅ 189_elasticsearch_timeseries_gap 40.4s 7 13 $0.3268
✅ 190_elasticsearch_cross_service_correlation 27.1s 4 6 $0.2176
✅ 191_elasticsearch_query_profile 18.1s 3 3 $0.1769
✅ 193_elasticsearch_large_mapping_search 14.8s 3 3 $0.1630
✅ 195_bash_simple_allowed 10.7s 3 3 $0.1242
✅ 195_elasticsearch_trace_large_fields 35.9s 6 8 $0.4225
✅ 211_prometheus_alerting_rules 18.9s 4 4 $0.2001
Total 24.8s avg 4.7 avg 6.2 avg $4.9494
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref master -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals in automatic regression runs:

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID

Examples: evals-tag-easy, evals-id-09_crashpod

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref master -f markers=regression -f filter=

aantn and others added 4 commits February 22, 2026 17:55
…7N9q' into claude/revert-except-jq-errors-A7N9q
Wrap the preview serialization in a try/except for TypeError/ValueError
to handle edge cases where non-JSON-serializable objects are passed
directly to _filter_result_data. Falls back to repr() on failure.

https://claude.ai/code/session_01N5rWgtZrghkJVJtL3xdEGZ
Signed-off-by: Claude <noreply@anthropic.com>
moshemorad added a commit that referenced this pull request Apr 27, 2026
…on (#1947)

## Summary

Fixes a silent-truncation bug in `JsonFilterMixin` that made six toolset
tools produce false-negative answers (e.g. `elasticsearch_list_indices`
reporting "no indices exist" against a cluster with hundreds of matching
indices). Full context, reproducer, and live-cluster symptom in #1946.

Root cause: when the LLM calls a mixin-backed tool with `max_depth=0` (a
plausible value given the old description *"0 returns only top-level
keys"*), `_truncate_to_depth` replaces the entire response with the
literal string `"...truncated at depth 0"`, but `filter_result()`
preserves `status=SUCCESS`. The LLM sees success + empty-looking data
and confidently reports the wrong answer with no retry signal.

This PR applies the minimal fix — two small edits in one file — plus
regression tests:

- **Rewrites the `max_depth` tool-schema description** so the LLM stops
choosing 0. States the valid range (`>= 1`), points at `jq` for precise
extraction, explicitly warns against `0` and negative values.
- **Fail-closed in `filter_result()` on `max_depth <= 0`** with a
self-corrective `ERROR` message the LLM can act on. The guard only fires
when the upstream call succeeded (`status == SUCCESS`), so genuine
upstream errors (e.g. HTTP 503 "cluster unreachable") are preserved
verbatim — no clobbering of real failures with a parameter error.
- Protects all six current consumer tools in a single mixin-level
change: `elasticsearch_list_indices`, `elasticsearch_mappings`,
`http_request`, `grafana_get_dashboard_by_uid`,
`grafana_get_home_dashboard`, `list_prometheus_rules`. Also preemptively
protects every toolset being added by #1695 (Datadog, ServiceNow,
MongoDB Atlas, RabbitMQ, Coralogix, New Relic, and more) once that PR
merges.

## Test plan

- [x] `poetry run pytest
tests/plugins/toolsets/test_json_filter_mixin.py -v --no-cov` → **9/9
passing** (4 pre-existing tests unchanged, 5 new regression tests)
- [x] Standalone sanity check of `_truncate_to_depth` against an
Elasticsearch `_cat/indices`–shaped payload: confirms bug at depth 0,
confirms known remaining gap at depth 1 on list-of-dicts (see below),
confirms real data returned at depth ≥ 2 and `None`
- [x] `max_depth<=0` test: returns `ERROR` with self-corrective message
- [x] `max_depth=-1` test: same (closes the undocumented
negative-means-full escape hatch)
- [x] Upstream-error-preservation test: when the mocked call returns
`ERROR` and the LLM (hypothetically) passes `max_depth=0`, the upstream
error string survives verbatim — no clobber
- [x] `max_depth` omitted: full response unchanged (no regression on the
happy path)
- [x] Description-wording regression test: the string *"0 returns only
top-level keys"* can never come back, and the description must mention
`>= 1`
- [ ] Maintainer to verify against a real Elasticsearch cluster using
the three-step reproducer in #1946

## Known remaining gap (deliberately out of scope — follow-up issue
welcome)

`max_depth=1` on a **list-of-dicts** response (the shape `_cat/indices`
returns) still produces `[sentinel, sentinel, …]` with `status=SUCCESS`
— the same silent-truncation class of bug, one level in. Fixing that
properly requires a response-envelope redesign (e.g. `{truncated: true,
max_depth_used, data, hint}`) that changes the response shape for all
six consumer tools and their tests. That is a strictly larger change and
deserves its own review and rollback surface, so it is deliberately out
of scope here. Documented in #1946 under "Known remaining gap".

## Related PRs

- #1695 extends `JsonFilterMixin` to seven additional toolsets but does
**not** touch the mixin core. If it merges before this fix, every
newly-covered toolset inherits the silent-truncation bug. Complementary
to this PR.
- #1615 improves `jq` error messaging in the same file but a different
region. No behavioral overlap; any conflict is mechanical.

## References

Closes #1946


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Bug Fixes**
* Added runtime validation for the depth parameter; values of 0 or below
now return an error with guidance instead of being processed.

* **Documentation**
* Updated depth parameter description to clarify that values must be >=
1; omit the parameter for a complete, untruncated response.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Sebastien Villanueva <sebastien.villanueva@gmail.com>
Co-authored-by: moshemorad <moshemorad12340@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants