Skip to content

Add toolsets_matrix support for comparing toolset configs on same eval scenarios - #1609

Merged
aantn merged 13 commits into
masterfrom
claude/evals-toolset-matrix-mwdXc
Feb 22, 2026
Merged

aantn merged 13 commits into
masterfrom
claude/evals-toolset-matrix-mwdXc

Conversation

@aantn

@aantn aantn commented Feb 21, 2026 •

Copy link
Copy Markdown
Collaborator

Adds a new toolsets_matrix field to test_case.yaml that lists multiple
toolset config filenames. Each file creates a separate test variant with
the same user_prompt, expected_output, and infrastructure - only the
toolset configuration changes. This enables comparing builtin toolsets
vs HTTP toolsets vs MCP on identical scenarios.

Changes:

  • HolmesTestCase: add toolsets_matrix, toolsets_config_name, toolsets_config_path fields
  • MockHelper: add _expand_toolsets_matrix() post-processing step after test case loading
  • MockToolsetManager: accept toolsets_config_path to override default toolsets.yaml resolution
  • test_ask_holmes.py, test_investigate.py: pass toolsets_config_path through

Example test_case.yaml usage:
toolsets_matrix:
- toolsets_builtin.yaml
- toolsets_http.yaml

Produces test IDs like: test_ask_holmes[01_test[builtin]-model-env]

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude noreply@anthropic.com

Summary by CodeRabbit

  • New Features

    • Added Datadog Logs REST API toolset with HTTP-based endpoints supporting log filtering, searching, and pagination with API key authentication.
  • Tests

    • Added matrix expansion support for test toolset configurations.
    • Updated test case to focus on CPU usage metrics.

…l scenarios

Adds a new `toolsets_matrix` field to test_case.yaml that lists multiple
toolset config filenames. Each file creates a separate test variant with
the same user_prompt, expected_output, and infrastructure - only the
toolset configuration changes. This enables comparing builtin toolsets
vs HTTP toolsets vs MCP on identical scenarios.

Changes:
- HolmesTestCase: add toolsets_matrix, toolsets_config_name, toolsets_config_path fields
- MockHelper: add _expand_toolsets_matrix() post-processing step after test case loading
- MockToolsetManager: accept toolsets_config_path to override default toolsets.yaml resolution
- test_ask_holmes.py, test_investigate.py: pass toolsets_config_path through

Example test_case.yaml usage:
  toolsets_matrix:
    - toolsets_builtin.yaml
    - toolsets_http.yaml

Produces test IDs like: test_ask_holmes[01_test[builtin]-model-env]

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>
Adds toolsets_http.yaml to 9 Datadog eval tests (91a-91h, 164) with HTTP
toolset configs that hit the same Datadog APIs using the generic HTTP
toolset instead of the native Datadog toolsets. Each test now runs twice:
once with the builtin toolset and once with the HTTP toolset.

HTTP toolset auth uses DD-API-KEY header + DD-APPLICATION-KEY via
default_headers, with llm_instructions documenting the API endpoints.

Coverage:
- Metrics API (GET /api/v1/metrics, /api/v1/query, /api/v2/metrics): 91a-91e, 91g
- Logs API (POST /api/v2/logs/events/search): 91f, 91h
- Traces API (POST /api/v2/spans/events/search, /analytics/aggregate): 164

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Feb 21, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ f107fa9 (#22273304553)

✅ Results of HolmesGPT evals

Automatically triggered by commit f107fa9 on branch claude/evals-toolset-matrix-mwdXc (labels: evals-tag-datadog)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 29/31 test cases were successful, 0 regressions, 1 skipped, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 30.7s 5 12 $0.2434
✅ 101_loki_historical_logs_pod_deleted 40.4s 5 10 $0.2755
✅ 110_cpu_graph_robusta_runner 25.3s 5 7 $0.2175
✅ 111_disabled_datadog_traces 8.2s 1 — $0.1196
✅ 111_pod_names_contain_service 31.0s 5 12 $0.2392
✅ 112_find_pvcs_by_uuid 32.2s 6 9 $0.2711
✅ 12_job_crashing 27.8s 5 10 $0.2368
🚧 164_datadog_traces_coupon_code[0] — — — —
✅ 176_network_policy_blocking_traffic_no_runbooks 43.9s 7 17 $0.2948
✅ 24_misconfigured_pvc 32.7s 5 16 $0.2489
✅ 43_current_datetime_from_prompt 4.0s 1 — $0.1115
✅ 61_exact_match_counting 13.9s 3 2 $0.1539
✅ 91a_datadog_metrics_no_k8s 19.5s 4 6 $0.1911
✅ 91b_datadog_metrics_pod_exists 16.3s 4 3 $0.1691
✅ 91c_datadog_metrics_deployment[0] 22.4s 5 7 $0.2080
✅ 91c_datadog_metrics_deployment[1] 18.4s 4 6 $0.1904
✅ 91c_datadog_metrics_deployment[2] 26.0s 5 7 $0.2199
✅ 91d_datadog_metrics_historical_pod[0] 19.5s 4 5 $0.1876
✅ 91d_datadog_metrics_historical_pod[1] 18.0s 4 4 $0.1820
✅ 91d_datadog_metrics_historical_pod[2] 19.5s 3 4 $0.1860
✅ 91e_datadog_custom_metrics[0] 15.8s 3 5 $0.1681
✅ 91e_datadog_custom_metrics[1] 15.2s 3 4 $0.1640
✅ 91f_datadog_logs_historical_pod[default] 38.3s 6 10 $0.2835
✅ 91f_datadog_logs_historical_pod[http] 37.4s 5 10 $0.3523
✅ 91g_datadog_metrics_mismatched_pod[0] 19.1s 4 5 $0.1917
✅ 91g_datadog_metrics_mismatched_pod[1] 19.1s 4 5 $0.1860
✅ 91h_datadog_logs_empty_query_with_url 14.7s 3 3 $0.1549
➖ 91i_datadog_metrics_empty_query_with_url — — — —
✅ 92_cpu_graph_conversation[0] 10.6s 2 2 $0.1474
✅ 92_cpu_graph_conversation[1] 11.0s 2 2 $0.1462
✅ 92_cpu_graph_conversation[2] 11.7s 3 2 $0.1559
Total 22.2s avg 4.0 avg 6.9 avg $5.8962
📜 Run @ 69825bb (#22272984724)

✅ Results of HolmesGPT evals

Automatically triggered by commit 69825bb on branch claude/evals-toolset-matrix-mwdXc (labels: evals-tag-datadog)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 29/32 test cases were successful, 1 regressions, 1 skipped, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 29.3s 5 10 $0.2320
✅ 101_loki_historical_logs_pod_deleted 29.9s 4 7 $0.2131
✅ 110_cpu_graph_robusta_runner 28.7s 5 9 $0.2326
✅ 111_disabled_datadog_traces 7.7s 1 — $0.1186
✅ 111_pod_names_contain_service 38.2s 7 12 $0.2569
✅ 112_find_pvcs_by_uuid 27.4s 5 7 $0.2255
✅ 12_job_crashing 38.7s 6 14 $0.2770
🚧 164_datadog_traces_coupon_code[0] — — — —
✅ 176_network_policy_blocking_traffic_no_runbooks 44.9s 6 15 $0.3002
✅ 24_misconfigured_pvc 31.0s 5 12 $0.2308
✅ 43_current_datetime_from_prompt 4.7s 1 — $0.1116
✅ 61_exact_match_counting 12.6s 3 2 $0.1498
✅ 91a_datadog_metrics_no_k8s 23.9s 5 7 $0.2033
✅ 91b_datadog_metrics_pod_exists 14.3s 3 3 $0.1604
✅ 91c_datadog_metrics_deployment[0] 22.3s 5 6 $0.1976
✅ 91c_datadog_metrics_deployment[1] 20.6s 4 6 $0.1917
✅ 91c_datadog_metrics_deployment[2] 25.2s 5 7 $0.2181
✅ 91d_datadog_metrics_historical_pod[0] 24.3s 5 6 $0.2032
✅ 91d_datadog_metrics_historical_pod[1] 17.7s 3 3 $0.1684
✅ 91d_datadog_metrics_historical_pod[2] 24.4s 4 6 $0.2078
✅ 91e_datadog_custom_metrics[0] 17.0s 3 4 $0.1644
✅ 91e_datadog_custom_metrics[1] 14.9s 3 4 $0.1640
✅ 91f_datadog_logs_historical_pod[default] 39.7s 6 12 $0.2864
❌ 91f_datadog_logs_historical_pod[http] 73.0s 12 31 $0.4374
✅ 91g_datadog_metrics_mismatched_pod[0] 24.0s 5 6 $0.2028
✅ 91g_datadog_metrics_mismatched_pod[1] 18.3s 4 5 $0.1830
✅ 91h_datadog_logs_empty_query_with_url[default] 17.2s 3 3 $0.1568
✅ 91h_datadog_logs_empty_query_with_url[http] 17.5s 3 3 $0.1554
➖ 91i_datadog_metrics_empty_query_with_url — — — —
✅ 92_cpu_graph_conversation[0] 13.5s 3 2 $0.1553
✅ 92_cpu_graph_conversation[1] 9.9s 2 2 $0.1456
✅ 92_cpu_graph_conversation[2] 13.3s 3 2 $0.1557
Total 24.1s avg 4.3 avg 7.4 avg $6.1054

⚠️ 1 Failure Detected

📜 Run @ 72973c8 (#22271420124)

✅ Results of HolmesGPT evals

Automatically triggered by commit 72973c8 on branch claude/evals-toolset-matrix-mwdXc (labels: evals-tag-datadog)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 30/32 test cases were successful, 0 regressions, 1 skipped, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 24.6s 4 9 $0.2104
✅ 101_loki_historical_logs_pod_deleted 39.0s 5 11 $0.2762
✅ 110_cpu_graph_robusta_runner 24.5s 5 8 $0.2139
✅ 111_disabled_datadog_traces 9.5s 1 — $0.1210
✅ 111_pod_names_contain_service 33.6s 5 12 $0.2395
✅ 112_find_pvcs_by_uuid 25.2s 5 5 $0.2099
✅ 12_job_crashing 37.6s 6 15 $0.2757
🚧 164_datadog_traces_coupon_code[0] — — — —
✅ 176_network_policy_blocking_traffic_no_runbooks 40.8s 6 16 $0.2849
✅ 24_misconfigured_pvc 37.3s 6 16 $0.2695
✅ 43_current_datetime_from_prompt 5.1s 1 — $0.1130
✅ 61_exact_match_counting 12.1s 3 2 $0.1499
✅ 91a_datadog_metrics_no_k8s 23.3s 5 6 $0.2103
✅ 91b_datadog_metrics_pod_exists 16.9s 4 3 $0.1694
✅ 91c_datadog_metrics_deployment[0] 23.0s 5 7 $0.2080
✅ 91c_datadog_metrics_deployment[1] 24.2s 5 7 $0.2073
✅ 91c_datadog_metrics_deployment[2] 23.6s 4 7 $0.2085
✅ 91d_datadog_metrics_historical_pod[0] 19.7s 4 4 $0.1823
✅ 91d_datadog_metrics_historical_pod[1] 19.5s 4 4 $0.1822
✅ 91d_datadog_metrics_historical_pod[2] 24.4s 4 6 $0.2078
✅ 91e_datadog_custom_metrics[0] 16.5s 3 4 $0.1655
✅ 91e_datadog_custom_metrics[1] 15.1s 3 4 $0.1652
✅ 91f_datadog_logs_historical_pod[default] 40.4s 6 11 $0.2715
✅ 91f_datadog_logs_historical_pod[http] 48.0s 7 14 $0.3388
✅ 91g_datadog_metrics_mismatched_pod[0] 19.3s 4 5 $0.1916
✅ 91g_datadog_metrics_mismatched_pod[1] 18.2s 4 5 $0.1889
✅ 91h_datadog_logs_empty_query_with_url[default] 16.3s 3 3 $0.1553
✅ 91h_datadog_logs_empty_query_with_url[http] 17.5s 3 3 $0.1550
➖ 91i_datadog_metrics_empty_query_with_url — — — —
✅ 92_cpu_graph_conversation[0] 10.9s 2 1 $0.1449
✅ 92_cpu_graph_conversation[1] 11.9s 3 2 $0.1550
✅ 92_cpu_graph_conversation[2] 10.1s 2 2 $0.1455
Total 22.9s avg 4.1 avg 6.9 avg $6.0170
📜 Run @ d312f85 (#22263719913)

✅ Results of HolmesGPT evals

Automatically triggered by commit d312f85 on branch claude/evals-toolset-matrix-mwdXc (labels: evals-tag-datadog)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 25/33 test cases were successful, 2 regressions, 1 skipped, 5 setup failures
Status Test case Time Turns Tools Cost
🚧 09_crashpod — — — —
❌ 101_loki_historical_logs_pod_deleted 44.0s 6 14 $0.2943
✅ 110_cpu_graph_robusta_runner 32.1s 6 10 $0.2514
✅ 111_disabled_datadog_traces 8.0s 1 — $0.1196
🚧 111_pod_names_contain_service — — — —
✅ 112_find_pvcs_by_uuid 22.6s 5 5 $0.2011
🚧 12_job_crashing — — — —
🚧 164_datadog_traces_coupon_code[0][default] — — — —
❌ 164_datadog_traces_coupon_code[0][http] 38.8s 4 13 $0.2847
🚧 176_network_policy_blocking_traffic_no_runbooks — — — —
✅ 24_misconfigured_pvc 39.7s 7 18 $0.2853
✅ 43_current_datetime_from_prompt 4.1s 1 — $0.1115
✅ 61_exact_match_counting 13.8s 3 2 $0.1533
✅ 91a_datadog_metrics_no_k8s 25.8s 6 7 $0.2166
✅ 91b_datadog_metrics_pod_exists 16.9s 4 3 $0.1697
✅ 91c_datadog_metrics_deployment[0] 18.3s 4 5 $0.1882
✅ 91c_datadog_metrics_deployment[1] 20.5s 4 6 $0.1902
✅ 91c_datadog_metrics_deployment[2] 27.2s 5 8 $0.2236
✅ 91d_datadog_metrics_historical_pod[0] 14.5s 3 3 $0.1663
✅ 91d_datadog_metrics_historical_pod[1] 22.3s 4 6 $0.1974
✅ 91d_datadog_metrics_historical_pod[2] 17.9s 3 4 $0.1838
✅ 91e_datadog_custom_metrics[0] 18.7s 3 5 $0.1676
✅ 91e_datadog_custom_metrics[1] 14.4s 3 4 $0.1656
✅ 91f_datadog_logs_historical_pod[default] 39.3s 7 11 $0.2730
✅ 91f_datadog_logs_historical_pod[http] 70.1s 12 22 $0.4608
✅ 91g_datadog_metrics_mismatched_pod[0] 23.7s 5 6 $0.2051
✅ 91g_datadog_metrics_mismatched_pod[1] 18.4s 4 5 $0.1862
✅ 91h_datadog_logs_empty_query_with_url[default] 20.0s 4 4 $0.1701
✅ 91h_datadog_logs_empty_query_with_url[http] 18.1s 3 3 $0.1557
➖ 91i_datadog_metrics_empty_query_with_url — — — —
✅ 92_cpu_graph_conversation[0] 10.7s 2 2 $0.1467
✅ 92_cpu_graph_conversation[1] 11.1s 2 2 $0.1461
✅ 92_cpu_graph_conversation[2] 10.9s 2 2 $0.1471
Total 23.0s avg 4.2 avg 6.8 avg $5.4608

⚠️ 2 Failures Detected

📜 Run @ 80cdf91 (#22260338635)

✅ Results of HolmesGPT evals

Automatically triggered by commit 80cdf91 on branch claude/evals-toolset-matrix-mwdXc (labels: evals-tag-datadog)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 29/45 test cases were successful, 14 regressions, 1 skipped, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 35.0s 6 13 $0.2640
✅ 101_loki_historical_logs_pod_deleted 36.5s 5 10 $0.2521
❌ 110_cpu_graph_robusta_runner 24.0s 5 7 $0.2106
✅ 111_disabled_datadog_traces 9.1s 1 — $0.1198
✅ 111_pod_names_contain_service 30.5s 5 12 $0.2392
✅ 112_find_pvcs_by_uuid 31.1s 6 7 $0.2475
✅ 12_job_crashing 27.4s 5 11 $0.2397
🚧 164_datadog_traces_coupon_code[0][default] — — — —
❌ 164_datadog_traces_coupon_code[0][http] 14508.2s 14 47 $1.0196
✅ 176_network_policy_blocking_traffic_no_runbooks 48.6s 8 18 $0.3276
✅ 24_misconfigured_pvc 35.3s 6 15 $0.2651
✅ 43_current_datetime_from_prompt 4.4s 1 — $0.1114
✅ 61_exact_match_counting 17.1s 4 4 $0.1701
✅ 91a_datadog_metrics_no_k8s[default] 20.1s 4 6 $0.1901
❌ 91a_datadog_metrics_no_k8s[http] 22.7s 5 7 $0.1979
✅ 91b_datadog_metrics_pod_exists[default] 12.4s 3 2 $0.1574
❌ 91b_datadog_metrics_pod_exists[http] 20.7s 3 3 $0.1666
✅ 91c_datadog_metrics_deployment[0][default] 18.5s 4 5 $0.1875
❌ 91c_datadog_metrics_deployment[0][http] 35.3s 8 9 $0.2496
✅ 91c_datadog_metrics_deployment[1][default] 22.1s 5 7 $0.2030
❌ 91c_datadog_metrics_deployment[1][http] 16.0s 4 5 $0.1680
✅ 91c_datadog_metrics_deployment[2][default] 23.3s 4 7 $0.2084
❌ 91c_datadog_metrics_deployment[2][http] 38.9s 7 11 $0.2672
✅ 91d_datadog_metrics_historical_pod[0][default] 14.7s 3 3 $0.1697
❌ 91d_datadog_metrics_historical_pod[0][http] 18.2s 3 3 $0.1676
✅ 91d_datadog_metrics_historical_pod[1][default] 21.6s 4 6 $0.1982
❌ 91d_datadog_metrics_historical_pod[1][http] 25.7s 5 8 $0.2049
✅ 91d_datadog_metrics_historical_pod[2][default] 19.0s 3 4 $0.1857
❌ 91d_datadog_metrics_historical_pod[2][http] 32.9s 6 10 $0.2413
✅ 91e_datadog_custom_metrics[0][default] 19.1s 4 5 $0.1785
❌ 91e_datadog_custom_metrics[0][http] 36.6s 7 12 $0.2540
✅ 91e_datadog_custom_metrics[1][default] 16.2s 3 4 $0.1644
❌ 91e_datadog_custom_metrics[1][http] 38.9s 8 9 $0.2481
✅ 91f_datadog_logs_historical_pod[default] 40.5s 6 12 $0.3049
✅ 91f_datadog_logs_historical_pod[http] 69.0s 11 22 $0.6674
✅ 91g_datadog_metrics_mismatched_pod[0][default] 23.3s 5 6 $0.2052
❌ 91g_datadog_metrics_mismatched_pod[0][http] 24.0s 3 4 $0.1809
✅ 91g_datadog_metrics_mismatched_pod[1][default] 19.0s 4 5 $0.1896
❌ 91g_datadog_metrics_mismatched_pod[1][http] 33.0s 5 9 $0.2377
✅ 91h_datadog_logs_empty_query_with_url[default] 15.2s 3 3 $0.1549
✅ 91h_datadog_logs_empty_query_with_url[http] 16.5s 3 3 $0.1549
➖ 91i_datadog_metrics_empty_query_with_url — — — —
✅ 92_cpu_graph_conversation[0] 16.0s 4 3 $0.1710
✅ 92_cpu_graph_conversation[1] 12.2s 3 2 $0.1561
✅ 92_cpu_graph_conversation[2] 12.2s 3 2 $0.1562
Total 361.9s avg 4.9 avg 8.3 avg $10.0538

⚠️ 14 Failures Detected


✅ Results of HolmesGPT evals

Automatically triggered by commit 1d9e3f7 on branch claude/evals-toolset-matrix-mwdXc (labels: evals-tag-datadog)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 28/31 test cases were successful, 1 regressions, 1 skipped, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 31.9s 5 12 $0.2520
✅ 101_loki_historical_logs_pod_deleted 46.8s 6 12 $0.2883
✅ 110_cpu_graph_robusta_runner 27.5s 5 8 $0.2260
✅ 111_disabled_datadog_traces 8.7s 1 — $0.1200
✅ 111_pod_names_contain_service 32.0s 5 12 $0.2429
✅ 112_find_pvcs_by_uuid 36.1s 7 8 $0.2654
❌ 12_job_crashing 23.7s 4 7 $0.2039
🚧 164_datadog_traces_coupon_code[0] — — — —
✅ 176_network_policy_blocking_traffic_no_runbooks 42.2s 6 15 $0.2818
✅ 24_misconfigured_pvc 33.2s 6 13 $0.2431
✅ 43_current_datetime_from_prompt 4.2s 1 — $0.1114
✅ 61_exact_match_counting 13.3s 3 2 $0.1499
✅ 91a_datadog_metrics_no_k8s 23.2s 5 7 $0.2044
✅ 91b_datadog_metrics_pod_exists 13.6s 3 3 $0.1586
✅ 91c_datadog_metrics_deployment[0] 19.7s 4 6 $0.1899
✅ 91c_datadog_metrics_deployment[1] 23.5s 5 7 $0.2062
✅ 91c_datadog_metrics_deployment[2] 28.6s 5 8 $0.2364
✅ 91d_datadog_metrics_historical_pod[0] 24.1s 5 6 $0.2048
✅ 91d_datadog_metrics_historical_pod[1] 15.9s 3 3 $0.1684
✅ 91d_datadog_metrics_historical_pod[2] 23.5s 4 7 $0.2080
✅ 91e_datadog_custom_metrics[0] 16.2s 3 5 $0.1680
✅ 91e_datadog_custom_metrics[1] 15.5s 3 4 $0.1647
✅ 91f_datadog_logs_historical_pod[default] 40.4s 6 11 $0.2953
✅ 91f_datadog_logs_historical_pod[http] 47.2s 7 11 $0.4248
✅ 91g_datadog_metrics_mismatched_pod[0] 24.3s 5 6 $0.2066
✅ 91g_datadog_metrics_mismatched_pod[1] 18.2s 4 5 $0.1853
✅ 91h_datadog_logs_empty_query_with_url 16.1s 3 3 $0.1567
➖ 91i_datadog_metrics_empty_query_with_url — — — —
✅ 92_cpu_graph_conversation[0] 15.5s 3 2 $0.1610
✅ 92_cpu_graph_conversation[1] 10.1s 2 1 $0.1443
✅ 92_cpu_graph_conversation[2] 11.6s 2 1 $0.1459
Total 23.7s avg 4.2 avg 6.9 avg $6.0142

⚠️ 1 Failure Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/evals-toolset-matrix-mwdXc -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals in automatic regression runs:

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID

Examples: evals-tag-easy, evals-id-09_crashpod

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/evals-toolset-matrix-mwdXc -f markers=regression -f filter=

@github-actions

github-actions Bot commented Feb 21, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for 7084e7fb (built in 1m 11s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:7084e7fb
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:7084e7fb me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:7084e7fb
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:7084e7fb

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:7084e7fb

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:7084e7fb

@coderabbitai

coderabbitai Bot commented Feb 21, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉


Walkthrough

This PR introduces a toolsets_matrix feature for test fixtures that enables running the same test case with multiple toolset configurations. It adds a new HTTP-based Datadog Logs API toolset configuration and propagates the toolsets_config_path throughout the test infrastructure to support dynamic toolset loading.

Changes

Cohort / File(s) Summary
Test Infrastructure - Matrix Expansion
tests/llm/utils/test_case_utils.py
Adds three new fields to HolmesTestCase (toolsets_matrix, toolsets_config_name, toolsets_config_path) and implements _expand_toolsets_matrix() method to generate per-config test variants with derived names and full paths. Integrates matrix expansion into test loading flow.
Mock Toolset Manager
tests/llm/utils/mock_toolset.py
Introduces optional toolsets_config_path parameter to MockToolsetManager.init and updates _initialize_toolsets() to use custom config path when provided, falling back to default location.
Test Setup Wiring
tests/llm/test_ask_holmes.py, tests/llm/test_investigate.py
Propagates toolsets_config_path parameter from test_case to MockToolsetManager instantiation in test setup code.
Datadog Logs Toolset Configuration
tests/llm/fixtures/test_ask_holmes/91f_datadog_logs_historical_pod/toolsets_http.yaml
Adds new HTTP-based datadog-logs-api toolset with REST endpoints, header-based authentication (API key and application key), 30-second timeout, and comprehensive llm_instructions covering request schema, parameter formats, and pagination.
Test Case Configuration Updates
tests/llm/fixtures/test_ask_holmes/91f_datadog_logs_historical_pod/test_case.yaml
Introduces toolsets_matrix configuration referencing two toolset files (toolsets.yaml and toolsets_http.yaml) to enable multi-variant test execution.
CPU Graph Fixture
tests/llm/fixtures/test_ask_holmes/110_cpu_graph_robusta_runner/test_case.yaml
Replaces multi-item memory-focused prompts with single CPU usage prompt and expands expected_output guidance regarding tool_call_id variations and tool output expectations.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~22 minutes

Possibly related PRs

Suggested reviewers

  • arikalon1
🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the primary change: adding toolsets_matrix support to enable comparing multiple toolset configurations on the same evaluation scenarios.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Feb 21, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 1d9e3f7
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/699abaac2131e90008509f4d
😎 Deploy Preview https://deploy-preview-1609--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
tests/llm/fixtures/test_ask_holmes/91b_datadog_metrics_pod_exists/test_case.yaml (1)

20-35: ⚠️ Potential issue | 🟠 Major

expected_output is incompatible with the toolsets_http.yaml matrix variant — HTTP variant will always fail.

The expected_output (lines 20–26) asserts:

  • An embed with type: "datadogql" and tool_name: "query_datadog_metrics"

These tokens are produced exclusively by the builtin Holmes Datadog toolset. The HTTP toolset (toolsets_http.yaml) calls the Datadog REST API directly and returns raw JSON; it has no mechanism to emit datadogql embed markers or surface a tool named query_datadog_metrics. As a result, the toolsets_http.yaml matrix variant will always fail this assertion, making it broken by design.

Options to resolve:

  1. Separate fixture: Create a dedicated test fixture for the HTTP variant with an expected_output that matches raw data presentation (e.g., asserting metric values are present in the response as text/table, without any embed requirement).
  2. Variant-specific expected outputs: If the framework supports it, define per-variant expected outputs instead of a single shared one.
  3. Relaxed assertion for shared output: If the single shared expected_output is intentional, remove the embed/tool-name assertion and instead assert on the substantive content (e.g., CPU metric values returned), so both variants can satisfy it.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In
`@tests/llm/fixtures/test_ask_holmes/91b_datadog_metrics_pod_exists/test_case.yaml`
around lines 20 - 35, The current shared expected_output requires an embed with
{"type":"datadogql","tool_name":"query_datadog_metrics"} which only the built-in
Holmes Datadog toolset emits, so update the tests by splitting the HTTP variant
into its own fixture: duplicate this test (the block containing toolsets_matrix
and expected_output), in the original keep the existing expected_output and only
include the built-in toolset (toolsets.yaml), and in the new HTTP-specific
fixture include toolsets_http.yaml and replace the expected_output to assert on
raw metric content/JSON presence (e.g., CPU metric keys or values) rather than
the datadogql embed or query_datadog_metrics tool_name so both variants can
pass.
🧹 Nitpick comments (4)
tests/llm/fixtures/test_ask_holmes/91d_datadog_metrics_historical_pod/toolsets_http.yaml (2)

1-34: Near-identical content duplicated across 9 test fixtures — optional extraction opportunity.

91d, 91e, and 91g (and the other six variants not shown here) all contain byte-for-byte the same datadog-metrics-api stanza; 91a differs only by omitting kubernetes/core. If the API host, auth scheme, or llm_instructions ever change, all nine files need updating in lockstep.

Consider a shared toolsets_http_base.yaml symlinked or referenced from each test directory, or at minimum a comment anchoring the canonical definition.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In
`@tests/llm/fixtures/test_ask_holmes/91d_datadog_metrics_historical_pod/toolsets_http.yaml`
around lines 1 - 34, The datadog-metrics-api stanza is duplicated across
multiple test fixtures; extract the repeated config (the datadog-metrics-api
block including endpoints, default_headers, timeout_seconds and
llm_instructions) into a single shared file (e.g., a toolsets_http_base.yaml)
and update each fixture to reference or symlink that shared file instead of
inlining the block; alternatively add a clear canonical comment at the top of
each duplicated fixture pointing to the shared canonical definition so future
edits happen in one place.

9-9: Hard-coded Datadog site ties all HTTP toolset tests to US5 — consider externalising via env var.

api.us5.datadoghq.com is duplicated verbatim across all nine toolsets_http.yaml fixtures in this PR. If the test account ever migrates to another region (US1, EU, AP1, …), every file needs a manual update. This is also inconsistent with the credentials themselves, which are already pulled from env vars.

♻️ Suggested change
-        - hosts: ["api.us5.datadoghq.com"]
+        - hosts: ["{{env.DATADOG_API_HOST}}"]

Then set DATADOG_API_HOST=api.us5.datadoghq.com in the CI/test environment (alongside DATADOG_API_KEY / DATADOG_APP_KEY), matching the existing DD_SITE convention used throughout Datadog tooling.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In
`@tests/llm/fixtures/test_ask_holmes/91d_datadog_metrics_historical_pod/toolsets_http.yaml`
at line 9, Replace the hard-coded Datadog host string hosts:
["api.us5.datadoghq.com"] in the toolsets_http.yaml fixtures with an
environment-driven value (e.g., read DATADOG_API_HOST with a fallback to
"api.us5.datadoghq.com"), so tests use process/CI env rather than a literal
region; update all nine toolsets_http.yaml fixtures to reference
DATADOG_API_HOST and ensure the CI/test environment defines DATADOG_API_HOST
alongside existing DATADOG_API_KEY/DATADOG_APP_KEY.
tests/llm/test_investigate.py (1)

63-65: getattr is unnecessary since toolsets_config_path is now a declared field on HolmesTestCase.

Since toolsets_config_path is a proper Pydantic field with a default of None on HolmesTestCase (and InvestigateTestCase inherits from it), you can access it directly as self._test_case.toolsets_config_path. That said, getattr is harmless and consistent with the pattern used for mock_overrides on line 62 — so this is a nit.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/llm/test_investigate.py` around lines 63 - 65, The call to
getattr(self._test_case, "toolsets_config_path", None) is unnecessary because
toolsets_config_path is a declared Pydantic field on HolmesTestCase (inherited
by InvestigateTestCase); replace the getattr usage with direct attribute access
self._test_case.toolsets_config_path in the code that constructs the test
invocation (the block referencing toolsets_config_path alongside mock_overrides)
so the code reads the field directly from the InvestigateTestCase/HolmesTestCase
instance.
tests/llm/fixtures/test_ask_holmes/91b_datadog_metrics_pod_exists/toolsets_http.yaml (1)

9-9: Hardcoded Datadog region (us5) may diverge from the test environment's actual site.

Both hosts (line 9) and the llm_instructions base URL (line 22) are hardcoded to api.us5.datadoghq.com. The builtin toolset derives its region from the configured Datadog site. If the CI environment targets a different region (e.g., US1 api.datadoghq.com), the HTTP toolset will silently fail all requests while the builtin variant succeeds, making cross-variant comparison meaningless.

Consider sourcing the region from an environment variable (e.g., {{env.DATADOG_SITE}}) and updating llm_instructions to reflect the dynamic base URL, consistent with how the builtin toolset resolves its endpoint.

Also applies to: 22-22

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In
`@tests/llm/fixtures/test_ask_holmes/91b_datadog_metrics_pod_exists/toolsets_http.yaml`
at line 9, The test hardcodes the Datadog region in the HTTP toolset causing
mismatch with the builtin variant; update the hosts entry (the hosts key
currently set to "api.us5.datadoghq.com") and the llm_instructions base URL (the
llm_instructions block) to derive the site from an environment variable (e.g.,
use DATADOG_SITE) instead of "us5" so both toolset variants use the same dynamic
base domain; ensure the hosts value and the base URL are constructed from that
env var consistently (same variable name) so CI can override the site.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In
`@tests/llm/fixtures/test_ask_holmes/91b_datadog_metrics_pod_exists/test_case.yaml`:
- Around line 20-35: The current shared expected_output requires an embed with
{"type":"datadogql","tool_name":"query_datadog_metrics"} which only the built-in
Holmes Datadog toolset emits, so update the tests by splitting the HTTP variant
into its own fixture: duplicate this test (the block containing toolsets_matrix
and expected_output), in the original keep the existing expected_output and only
include the built-in toolset (toolsets.yaml), and in the new HTTP-specific
fixture include toolsets_http.yaml and replace the expected_output to assert on
raw metric content/JSON presence (e.g., CPU metric keys or values) rather than
the datadogql embed or query_datadog_metrics tool_name so both variants can
pass.

---

Duplicate comments:
In
`@tests/llm/fixtures/test_ask_holmes/91a_datadog_metrics_no_k8s/toolsets_http.yaml`:
- Line 7: The hosts entry is hard-coded to "api.us5.datadoghq.com"; update the
template to use a configurable Datadog site variable instead (e.g., reference a
variable like datadog_site or an env var DATADOG_SITE with a sensible default)
so the hosts key in toolsets_http.yaml is not fixed to us5—locate the hosts:
["api.us5.datadoghq.com"] line and replace it with a variable reference and
ensure callers set the variable or fallback is provided.

In
`@tests/llm/fixtures/test_ask_holmes/91e_datadog_custom_metrics/toolsets_http.yaml`:
- Line 9: The hosts entry is hard-coded to "api.us5.datadoghq.com"; change it to
use a configurable variable/env placeholder (e.g., DATADOG_SITE or a templated
variable like {{ datadog_site }}) so different Datadog sites can be targeted;
update the hosts line that currently reads hosts: ["api.us5.datadoghq.com"] to
reference that variable and ensure any test harness or CI sets
DATADOG_SITE/defaults accordingly (search for the hosts array in
toolsets_http.yaml to locate the exact spot).

In
`@tests/llm/fixtures/test_ask_holmes/91g_datadog_metrics_mismatched_pod/toolsets_http.yaml`:
- Line 9: The hosts entry currently hard-codes the Datadog site as hosts:
["api.us5.datadoghq.com"]; change this to use a configurable variable (e.g.,
datadog_site or DATADOG_SITE) and construct the hosts list from that variable so
tests/configs can target different Datadog sites; update the YAML entry
referencing hosts to pull from the variable (or environment) instead of the
fixed "api.us5.datadoghq.com" and ensure any test harness that loads this
fixture supplies the variable or env var.

---

Nitpick comments:
In
`@tests/llm/fixtures/test_ask_holmes/91b_datadog_metrics_pod_exists/toolsets_http.yaml`:
- Line 9: The test hardcodes the Datadog region in the HTTP toolset causing
mismatch with the builtin variant; update the hosts entry (the hosts key
currently set to "api.us5.datadoghq.com") and the llm_instructions base URL (the
llm_instructions block) to derive the site from an environment variable (e.g.,
use DATADOG_SITE) instead of "us5" so both toolset variants use the same dynamic
base domain; ensure the hosts value and the base URL are constructed from that
env var consistently (same variable name) so CI can override the site.

In
`@tests/llm/fixtures/test_ask_holmes/91d_datadog_metrics_historical_pod/toolsets_http.yaml`:
- Around line 1-34: The datadog-metrics-api stanza is duplicated across multiple
test fixtures; extract the repeated config (the datadog-metrics-api block
including endpoints, default_headers, timeout_seconds and llm_instructions) into
a single shared file (e.g., a toolsets_http_base.yaml) and update each fixture
to reference or symlink that shared file instead of inlining the block;
alternatively add a clear canonical comment at the top of each duplicated
fixture pointing to the shared canonical definition so future edits happen in
one place.
- Line 9: Replace the hard-coded Datadog host string hosts:
["api.us5.datadoghq.com"] in the toolsets_http.yaml fixtures with an
environment-driven value (e.g., read DATADOG_API_HOST with a fallback to
"api.us5.datadoghq.com"), so tests use process/CI env rather than a literal
region; update all nine toolsets_http.yaml fixtures to reference
DATADOG_API_HOST and ensure the CI/test environment defines DATADOG_API_HOST
alongside existing DATADOG_API_KEY/DATADOG_APP_KEY.

In `@tests/llm/test_investigate.py`:
- Around line 63-65: The call to getattr(self._test_case,
"toolsets_config_path", None) is unnecessary because toolsets_config_path is a
declared Pydantic field on HolmesTestCase (inherited by InvestigateTestCase);
replace the getattr usage with direct attribute access
self._test_case.toolsets_config_path in the code that constructs the test
invocation (the block referencing toolsets_config_path alongside mock_overrides)
so the code reads the field directly from the InvestigateTestCase/HolmesTestCase
instance.

…rect

When a jq filter fails (e.g. `.metrics[]` on a null field), instead of
returning an empty ERROR that gives the LLM zero information, return the
original response (truncated to depth 2) alongside the error hint. This
lets the LLM see the actual response shape and retry with a null-safe
expression like `(.metrics // [])[]`.

Also adds null-safe jq guidance to HTTP toolset instructions.

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Feb 21, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 9.96s 10.07s -1.1%
Warm Mean 4.56s 4.61s -1.0%
Warm Min 4.55s 4.57s
Warm Max 4.58s 4.68s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 26.59s 27.26s -2.4%
Warm Mean 7.30s 7.29s +0.1%
Warm Min 6.71s 6.84s
Warm Max 7.85s 7.81s

PR: 447cc9dd | Master: 263c5cbd | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
holmes/plugins/toolsets/json_filter_mixin.py (1)

51-53: ⚠️ Potential issue | 🟡 Minor

Stale # pragma: no cover — the new test now exercises this path.

test_invalid_jq_returns_data_with_error_hint passes ".[" to _invoke, which causes jq.compile(".[") to raise a ValueError, so the except block is now reached. Remove the pragma to keep coverage reporting accurate.

🧹 Suggested cleanup
-    except Exception as exc:  # pragma: no cover - defensive
+    except Exception as exc:
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/toolsets/json_filter_mixin.py` around lines 51 - 53, The
except block in json_filter_mixin.py that catches exceptions from jq (seen in
the except Exception exc: handler which logs via logger.debug and returns
"Invalid jq expression") still has a stale "# pragma: no cover" marker; remove
that pragma so test coverage reflects the new test path (e.g.,
test_invalid_jq_returns_data_with_error_hint triggers jq.compile error and hits
this handler). Keep the existing logger.debug("Failed to apply jq filter",
exc_info=exc) and the return None, f"Invalid jq expression: {exc}" behavior, but
delete the " # pragma: no cover - defensive" comment on the except line.
🧹 Nitpick comments (2)
tests/plugins/toolsets/test_json_filter_mixin.py (1)

42-53: Optional: assert jq_expression is included in the error payload.

The structured fallback includes a jq_expression field (line 99 of json_filter_mixin.py), but the test only checks jq_error and raw_response_preview. Adding a check here pins the full contract and prevents accidental removal of the field.

💡 Optional addition
     assert "raw_response_preview" in result.data
+    assert result.data["jq_expression"] == ".["
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/plugins/toolsets/test_json_filter_mixin.py` around lines 42 - 53,
Update the test_invalid_jq_returns_data_with_error_hint to also assert the
presence (and value) of the jq_expression field in the error payload: when
invoking tool._invoke with {"uid": "abc", "jq": ".["}, assert "jq_expression" is
in result.data and that result.data["jq_expression"] == ".[" to match the
fallback constructed in json_filter_mixin.py (the jq_expression field created
around line 99).
holmes/plugins/toolsets/json_filter_mixin.py (1)

98-98: Optional: tailor the null-safe hint to runtime errors only.

The null-safe advice is always appended to the error message, even for compile-time syntax errors (e.g., ".[") where (.key // [])[] patterns are irrelevant. The LLM receives the raw {exc} text which already describes the real problem, but the appended hint can be misleading.

Consider splitting the hint:

💡 Optional refinement
-                "jq_error": f"{error}. Use null-safe patterns like (.key // [])[] instead of .key[] to handle missing/null fields.",
+                "jq_error": (
+                    f"{error}. Use null-safe patterns like (.key // [])[] instead of "
+                    ".key[] to handle missing/null fields."
+                    if "null" in str(error).lower() or "null (null)" in str(error).lower()
+                    else str(error)
+                ),

Or simply document the unconditional hint as an intentional, safe-by-default nudge and leave it as-is.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/toolsets/json_filter_mixin.py` at line 98, The message
appended to the "jq_error" field should only include the null-safe hint for
runtime missing-field errors, not for compile/parse errors; update the code that
constructs the "jq_error" string (the place building f"{error}..." in
json_filter_mixin.py, e.g., inside JsonFilterMixin or the method that catches jq
exceptions) to detect parse/syntax faults by inspecting the exception type or
text (check for keywords like "parse", "syntax", "Unrecognized", or an exception
attribute indicating a parse error) and only append "Use null-safe patterns like
(.key // [])[] ..." when it is not a parse/compile error; leave the original
exception text ({exc}/{error}) intact in all cases.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@holmes/plugins/toolsets/json_filter_mixin.py`:
- Around line 51-53: The except block in json_filter_mixin.py that catches
exceptions from jq (seen in the except Exception exc: handler which logs via
logger.debug and returns "Invalid jq expression") still has a stale "# pragma:
no cover" marker; remove that pragma so test coverage reflects the new test path
(e.g., test_invalid_jq_returns_data_with_error_hint triggers jq.compile error
and hits this handler). Keep the existing logger.debug("Failed to apply jq
filter", exc_info=exc) and the return None, f"Invalid jq expression: {exc}"
behavior, but delete the " # pragma: no cover - defensive" comment on the except
line.

---

Nitpick comments:
In `@holmes/plugins/toolsets/json_filter_mixin.py`:
- Line 98: The message appended to the "jq_error" field should only include the
null-safe hint for runtime missing-field errors, not for compile/parse errors;
update the code that constructs the "jq_error" string (the place building
f"{error}..." in json_filter_mixin.py, e.g., inside JsonFilterMixin or the
method that catches jq exceptions) to detect parse/syntax faults by inspecting
the exception type or text (check for keywords like "parse", "syntax",
"Unrecognized", or an exception attribute indicating a parse error) and only
append "Use null-safe patterns like (.key // [])[] ..." when it is not a
parse/compile error; leave the original exception text ({exc}/{error}) intact in
all cases.

In `@tests/plugins/toolsets/test_json_filter_mixin.py`:
- Around line 42-53: Update the test_invalid_jq_returns_data_with_error_hint to
also assert the presence (and value) of the jq_expression field in the error
payload: when invoking tool._invoke with {"uid": "abc", "jq": ".["}, assert
"jq_expression" is in result.data and that result.data["jq_expression"] == ".["
to match the fallback constructed in json_filter_mixin.py (the jq_expression
field created around line 99).

The raw_response_preview is now serialized to a JSON string and
truncated at 2000 characters, preventing large API responses (e.g.
thousands of metrics) from overwhelming the LLM context window.

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>
Reduce from 17 prompt variants to one representative prompt.

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/llm/fixtures/test_ask_holmes/110_cpu_graph_robusta_runner/test_case.yaml (1)

2-8: Expected output is too generic — it only validates embed format, not metric content.

Per coding guidelines, eval tests should be specific: e.g., verifying the query targets memory metrics (system.mem.* or equivalent) and covers the requested 24-hour window. The current assertion would pass even if Holmes queried the wrong metric or wrong time range, as long as any datadogql embed is present.

Consider tightening the assertion to include, for example, that the metric name in the query relates to memory (e.g., container.memory), consistent with the user prompt requesting memory usage. Based on learnings, eval test prompts should test exact values, not just structural format.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In
`@tests/llm/fixtures/test_ask_holmes/110_cpu_graph_robusta_runner/test_case.yaml`
around lines 2 - 8, Update the test's expected_output assertion (the YAML key
expected_output) to validate not just the presence of a datadogql embed but that
the embedded query targets memory metrics and a 24-hour window; specifically
require the embed JSON to contain "type":"datadogql",
"tool_name":"query_datadog_metrics", and that the query string includes
memory-related metric names (e.g., "system.mem", "container.memory", or
"container.memory.usage") and a time range or aggregate indicating 24h (e.g.,
"from: now-24h" or "last 24 hours"); adjust the text so the test will fail if
the embed queries unrelated metrics or omits the 24-hour window while keeping
the same embed format checks.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In
`@tests/llm/fixtures/test_ask_holmes/110_cpu_graph_robusta_runner/test_case.yaml`:
- Line 1: The fixture's user_prompt asks for "memory usage" but the fixture
folder is named 110_cpu_graph_robusta_runner indicating a CPU graph test; fix
the semantic mismatch by either updating the user_prompt value to request "CPU
usage" (change the user_prompt string to "Show me robusta-runner CPU usage over
the last 24 hours") or renaming the fixture folder from
110_cpu_graph_robusta_runner to reflect a memory-metric scenario; ensure the
change is applied to the user_prompt field (symbol: user_prompt) or the fixture
folder name (symbol: 110_cpu_graph_robusta_runner) so intent and naming are
consistent.

---

Nitpick comments:
In
`@tests/llm/fixtures/test_ask_holmes/110_cpu_graph_robusta_runner/test_case.yaml`:
- Around line 2-8: Update the test's expected_output assertion (the YAML key
expected_output) to validate not just the presence of a datadogql embed but that
the embedded query targets memory metrics and a 24-hour window; specifically
require the embed JSON to contain "type":"datadogql",
"tool_name":"query_datadog_metrics", and that the query string includes
memory-related metric names (e.g., "system.mem", "container.memory", or
"container.memory.usage") and a time range or aggregate indicating 24h (e.g.,
"from: now-24h" or "last 24 hours"); adjust the text so the test will fail if
the embed queries unrelated metrics or omits the 24-hour window while keeping
the same embed format checks.

Comment thread tests/llm/fixtures/test_ask_holmes/110_cpu_graph_robusta_runner/test_case.yaml Outdated
aantn and others added 5 commits February 21, 2026 20:33
…trics

The 6 embed tests (91a, 91b, 91c, 91d, 91e, 91g) expect datadogql
embeds with tool_name "query_datadog_metrics", which only exists in the
builtin Python datadog/metrics toolset. The HTTP toolset has no such
tool, so the LLM loops forever trying to produce an impossible output,
causing CI to hang for 3+ hours.

Kept toolsets_matrix on 91f (logs) and 164 (traces) since those tests
don't require the embed format.

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn enabled auto-merge (squash) February 22, 2026 08:03
…folder name

- Reverted json_filter_mixin.py, instructions.jinja2, and
  test_json_filter_mixin.py to master (jq error handling goes in PR2)
- Changed 110 user_prompt from "memory usage" to "CPU usage" to match
  the folder name 110_cpu_graph_robusta_runner

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn merged commit e6ee525 into master Feb 22, 2026
17 of 18 checks passed
@aantn
aantn deleted the claude/evals-toolset-matrix-mwdXc branch February 22, 2026 08:19
moshemorad pushed a commit that referenced this pull request Feb 22, 2026
…l scenarios (#1609)

Adds a new `toolsets_matrix` field to test_case.yaml that lists multiple
toolset config filenames. Each file creates a separate test variant with
the same user_prompt, expected_output, and infrastructure - only the
toolset configuration changes. This enables comparing builtin toolsets
vs HTTP toolsets vs MCP on identical scenarios.

Changes:
- HolmesTestCase: add toolsets_matrix, toolsets_config_name,
toolsets_config_path fields
- MockHelper: add _expand_toolsets_matrix() post-processing step after
test case loading
- MockToolsetManager: accept toolsets_config_path to override default
toolsets.yaml resolution
- test_ask_holmes.py, test_investigate.py: pass toolsets_config_path
through

Example test_case.yaml usage:
  toolsets_matrix:
    - toolsets_builtin.yaml
    - toolsets_http.yaml

Produces test IDs like: test_ask_holmes[01_test[builtin]-model-env]

https://claude.ai/code/session_014iLLnCruoc4xRE3r8aiXJh
Signed-off-by: Claude <noreply@anthropic.com>

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* HTTP integrations for Datadog Traces, Metrics, and Logs with auth and
usage guidance.
* Test matrix support to run tests against multiple toolset
configurations.
* Improved JSON filtering: null-safe guidance and structured error hints
when filters fail.

* **Tests**
* Added and expanded fixtures and test cases to cover the new Datadog
integrations and matrixed runs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Mohse Morad <moshemorad12340@gmail.com>
@coderabbitai coderabbitai Bot mentioned this pull request Feb 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants