Skip to content

experimental: add generic curl tool for whitelisted domains - #1405

Closed
aantn wants to merge 19 commits into
masterfrom
claude/generic-curl-tool-Q555c
Closed

aantn wants to merge 19 commits into
masterfrom
claude/generic-curl-tool-Q555c

Conversation

@aantn

@aantn aantn commented Jan 22, 2026 •

Copy link
Copy Markdown
Collaborator

and try to replace confluence tool with it

Summary by CodeRabbit

  • New Features

    • Introduced a generic HTTP toolset for whitelisted HTTP requests (supports basic, bearer, and header auth); Confluence flows now use it.
  • Documentation

    • Added OpenRouter/OpenAI fallback guidance and expanded SSL verification troubleshooting for sandbox TLS errors.
    • Added HTTP tool usage guidance and Confluence REST API instructions for classic and service-account flows.
  • Tests

    • Added comprehensive tests and new fixtures covering HTTP auth types, endpoint matching, health checks, and Confluence service-account scenarios.

✏️ Tip: You can customize this high-level summary in your review settings.

Introduces a new HTTP toolset that allows making requests to
user-configured API endpoints. This provides a generic alternative
to creating dedicated toolsets for simple API integrations.

Features:
- Host-based whitelisting with wildcard support (*.example.com)
- Optional path restrictions per endpoint
- Configurable HTTP methods (GET by default, POST opt-in)
- Multiple auth types: basic, bearer, custom header
- JSON filtering via JsonFilterMixin (jq, max_depth)
- Environment variable support for secrets ({{ env.FOO }})

Configuration example:
  http:
    endpoints:
      - host: "*.atlassian.net"
        auth:
          type: basic
          username: "{{ env.CONFLUENCE_USER }}"
          password: "{{ env.CONFLUENCE_API_KEY }}"

This toolset can replace simpler curl-based toolsets like Confluence
while offering more flexibility for ad-hoc API access.

Signed-off-by: Claude <noreply@anthropic.com>
- Remove dedicated confluence.yaml toolset (single curl command)
- Update Confluence evals (208, 209, 210) to use the HTTP toolset
- Add Confluence API patterns to HTTP instructions for LLM guidance
- Fix env var name: CONFLUENCE_USER_NAME (not CONFLUENCE_USER)

The HTTP toolset provides more flexibility:
- LLM can construct any valid Confluence API URL
- Supports search by title (not just page ID fetching)
- Can expand to other Atlassian APIs with same auth

This validates the HTTP toolset as a replacement for simple
curl-based toolsets while offering more capability.

Signed-off-by: Claude <noreply@anthropic.com>
Add documentation for using OpenRouter when OPENAI_API_KEY is not
available. This helps Claude Code know to use OPENROUTER_API_KEY
as a fallback for running LLM evaluation tests.

Includes:
- OpenRouter model format examples
- Command to check available API keys
- Note about using same model for CLASSIFIER_MODEL

Signed-off-by: Claude <noreply@anthropic.com>
- Use Opus 4.5 as the primary recommended model
- Add stronger emphasis to always try OpenRouter before giving up
- Add "check available keys" step at the beginning
- Make it clear this is the fallback when OpenAI key is missing

Signed-off-by: Claude <noreply@anthropic.com>
@netlify

netlify Bot commented Jan 22, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit f967155
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/697db2b12a183700086965fe
😎 Deploy Preview https://deploy-preview-1405--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@linux-foundation-easycla

linux-foundation-easycla Bot commented Jan 22, 2026 •

Copy link
Copy Markdown

CLA Signed

The committers listed above are authorized under a signed CLA.

  • ✅ login: aantn / name: Natan Yellin (f967155)

@github-actions

github-actions Bot commented Jan 22, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 5f49f60 (#21248353706)

✅ Results of HolmesGPT evals

Automatically triggered by commit 5f49f60 on branch claude/generic-curl-tool-Q555c

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.3s ±0% 5 12 $0.2277
✅ 101_loki_historical_logs_pod_deleted 63.0s ±0% 9 16 $0.3149
✅ 111_pod_names_contain_service 29.1s ±0% 5 10 $0.2032
✅ 12_job_crashing 35.1s ↓14% 5 13 $0.2422
✅ 162_get_runbooks 46.8s ↑11% 7 12 $0.2841
✅ 176_network_policy_blocking_traffic_no_runbooks 43.6s ±0% 6 15 $0.2665
✅ 24_misconfigured_pvc 37.0s ±0% 6 14 $0.2325
✅ 43_current_datetime_from_prompt 4.3s ±0% 1 — $0.0942
✅ 61_exact_match_counting 15.5s ±0% 4 4 $0.1501
Total 34.1s avg 5.3 avg 12.0 avg $2.0155

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/generic-curl-tool-Q555c'

Status: Success - 36 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ c17970d (#21247763693)

✅ Results of HolmesGPT evals

Automatically triggered by commit c17970d on branch claude/generic-curl-tool-Q555c

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 28.1s ↓10% 5 9 $0.2008
✅ 101_loki_historical_logs_pod_deleted 51.8s ↓19% 7 13 $0.2714
✅ 111_pod_names_contain_service 37.6s ↑30% 6 12 $0.2340
✅ 12_job_crashing 37.3s ±0% 5 16 $0.2453
✅ 162_get_runbooks 41.8s ±0% 6 11 $0.2689
✅ 176_network_policy_blocking_traffic_no_runbooks 42.5s ±0% 6 15 $0.2693
✅ 24_misconfigured_pvc 38.5s ±0% 6 15 $0.2472
✅ 43_current_datetime_from_prompt 4.5s ±0% 1 — $0.0942
✅ 61_exact_match_counting 16.0s ±0% 4 4 $0.1501
Total 33.1s avg 5.1 avg 11.9 avg $1.9813

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/generic-curl-tool-Q555c'

Status: Success - 36 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ bcb2960 (#21247433547)

✅ Results of HolmesGPT evals

Automatically triggered by commit bcb2960 on branch claude/generic-curl-tool-Q555c

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 26.4s ↓16% 5 9 $0.2038
✅ 101_loki_historical_logs_pod_deleted 68.7s ±0% 10 18 $0.3472
✅ 111_pod_names_contain_service 27.0s ±0% 5 8 $0.1897
✅ 12_job_crashing 36.8s ±0% 5 15 $0.2544
✅ 162_get_runbooks 43.1s ±0% 7 12 $0.2971
✅ 176_network_policy_blocking_traffic_no_runbooks 40.4s ↓14% 6 15 $0.2528
✅ 24_misconfigured_pvc 35.5s ±0% 6 14 $0.2381
✅ 43_current_datetime_from_prompt 4.4s ±0% 1 — $0.0942
✅ 61_exact_match_counting 15.6s ±0% 4 4 $0.1502
Total 33.1s avg 5.4 avg 11.9 avg $2.0274

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/generic-curl-tool-Q555c'

Status: Success - 36 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ fe74b51 (#21245829669)

✅ Results of HolmesGPT evals

Automatically triggered by commit fe74b51 on branch claude/generic-curl-tool-Q555c

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 7/9 test cases were successful, 0 regressions, 2 skipped
Status Test case Time Turns Tools Cost
✅ 09_crashpod 39.2s ↑25% 5 9 $0.2017
✅ 101_loki_historical_logs_pod_deleted 67.9s ±0% 10 18 $0.3530
✅ 111_pod_names_contain_service 40.4s ↑39% 5 8 $0.1955
✅ 12_job_crashing 52.3s ↑28% 5 15 $0.3952
✅ 162_get_runbooks 77.4s ↑84% 7 14 $0.4248
:minus: 176_network_policy_blocking_traffic_no_runbooks-bedrock/eu.anthropic.claude-opus-4-5-20251101-v1: — — — —
:minus: 24_misconfigured_pvc-bedrock/eu.anthropic.claude-opus-4-5-20251101-v1: — — — —
✅ 43_current_datetime_from_prompt 13.6s ↑218% 1 — $0.0942
✅ 61_exact_match_counting 29.0s ↑76% 4 4 $0.1708
Total 45.7s avg 5.3 avg 11.3 avg $1.8353

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/generic-curl-tool-Q555c'

Status: Success - 36 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 0597077 (#21245668606)

✅ Results of HolmesGPT evals

Automatically triggered by commit 0597077 on branch claude/generic-curl-tool-Q555c

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 28.0s ↓11% 5 8 $0.1922
✅ 101_loki_historical_logs_pod_deleted 61.7s ±0% 9 16 $0.3271
✅ 111_pod_names_contain_service 28.8s ±0% 5 9 $0.1964
✅ 12_job_crashing 36.4s ↓11% 5 13 $0.2430
✅ 162_get_runbooks 40.0s ±0% 6 13 $0.2623
✅ 176_network_policy_blocking_traffic_no_runbooks 33.2s ↓29% 5 15 $0.2517
✅ 24_misconfigured_pvc 36.7s ±0% 6 14 $0.2439
✅ 43_current_datetime_from_prompt 4.7s ±0% 1 — $0.0942
✅ 61_exact_match_counting 16.6s ±0% 4 4 $0.1502
Total 31.8s avg 5.1 avg 11.5 avg $1.9611

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/generic-curl-tool-Q555c'

Status: Success - 36 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit f967155 on branch claude/generic-curl-tool-Q555c

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.4s ±0% 5 11 $0.2318
✅ 101_loki_historical_logs_pod_deleted 29.9s ↓28% 4 8 $0.2119
✅ 111_pod_names_contain_service 34.4s ±0% 5 12 $0.2280
✅ 12_job_crashing 35.8s ↑12% 5 12 $0.2446
✅ 162_get_runbooks 43.4s ±0% 7 12 $0.2939
✅ 176_network_policy_blocking_traffic_no_runbooks 50.5s ↑14% 7 16 $0.2981
✅ 24_misconfigured_pvc 32.3s ±0% 5 14 $0.2292
✅ 43_current_datetime_from_prompt 4.7s ↓10% 1 — $0.1052
✅ 61_exact_match_counting 13.0s ↓15% 3 2 $0.1419
Total 30.7s avg 4.7 avg 10.9 avg $1.9845

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/generic-curl-tool-Q555c'

Status: Success - 9 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence-service-account, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=

@github-actions

github-actions Bot commented Jan 22, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for e8b2470 (built in 4m 46s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e8b2470
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e8b2470 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e8b2470
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e8b2470

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:e8b2470

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:e8b2470

@aantn

aantn commented Jan 22, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
markers: confluence

@coderabbitai

coderabbitai Bot commented Jan 22, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Rate limit exceeded

@aantn has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 19 minutes and 41 seconds before requesting another review.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

Walkthrough

Adds a new generic HTTP toolset (implementation, template, registration, and tests), migrates Confluence fixtures to use the HTTP toolset (classic and service-account flows), removes the Confluence YAML toolset, and extends CLAUDE.md with OpenRouter/OpenAI fallback and SSL_VERIFY sandbox guidance.

Changes

Cohort / File(s) Summary
HTTP Toolset Core
holmes/plugins/toolsets/http/http_toolset.py
New implementation: typed config (AuthConfig, EndpointConfig, HttpToolsetConfig), endpoint whitelist (host/path matching, wildcards), auth types (none/basic/bearer/header), health checks, header/basic-auth helpers, and HttpRequest tool executing requests with timeout/SSL handling and structured results.
Toolset Registration & Packaging
holmes/plugins/toolsets/__init__.py,
holmes/plugins/toolsets/http/__init__.py
Registers HttpToolset in the Python toolset loader and adds package initializer for the HTTP toolset.
Docs & Template
holmes/plugins/toolsets/http/instructions.jinja2,
CLAUDE.md
New Jinja2 instructions template for the HTTP tool; CLAUDE.md extended with OpenRouter/OpenAI fallback guidance and expanded SSL_VERIFY sandbox TLS troubleshooting notes.
Confluence Toolset Removal
holmes/plugins/toolsets/confluence.yaml
Deleted Confluence-specific toolset configuration and example curl command.
Tests — HTTP Toolset Unit Tests
tests/plugins/toolsets/http/__init__.py,
tests/plugins/toolsets/http/test_http_toolset.py
New test package and extensive unit tests covering config validation, host/path matching, method enforcement, auth/header behavior, health checks, and HttpRequest invocation/error handling.
Tests — Fixture Migration (Confluence → HTTP)
tests/llm/fixtures/test_ask_holmes/208_*/toolsets.yaml, .../208_*/test_case.yaml,
.../209_*/**, .../210_*/**, .../211_*/**, .../212_*/**, .../213_*/**
Migrated Confluence fixtures to http toolset (classic and service-account flows). Updated env vars (CONFLUENCE_USER → CONFLUENCE_USER_NAME), added endpoint definitions/auth (basic/bearer), added instructions, and updated expectations to require http_request usage.
Pytest Markers
pyproject.toml
Updated confluence marker description and added confluence-service-account marker.
Minor docs tweak
docs/development/evaluations/history/results_20260120_143709.md
Small formatting change only.

Sequence Diagram(s)

sequenceDiagram
    participant User as User/Holmes
    participant HttpTool as HttpRequest
    participant Toolset as HttpToolset
    participant Executor as HTTP Executor
    participant Service as External HTTP Service

    User->>HttpTool: Invoke(url, method, body, headers)
    HttpTool->>Toolset: match_endpoint(url)
    Toolset->>Toolset: parse host & path, match host pattern, match path patterns
    Toolset-->>HttpTool: return EndpointConfig or error
    HttpTool->>HttpTool: validate method allowed
    HttpTool->>Toolset: build_headers(endpoint, extra_headers)
    Toolset-->>HttpTool: return headers and basic auth tuple (if any)
    HttpTool->>Executor: execute request(method, url, headers, auth, timeout, verify)
    Executor->>Service: send HTTP request
    Service-->>Executor: response
    Executor-->>HttpTool: parse JSON or text, handle errors/timeouts
    HttpTool-->>User: return StructuredToolResult(status, data/error)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

  • Datadog general tool #911 — Similar change adding a new Python toolset class to holmes/plugins/toolsets/__init__.py (toolset registration patterns overlap).
  • Updating URL's for API usage #900 — Prior edits to Confluence toolset configuration; directly related to converting/removing Confluence YAML in this PR.
  • Create CLAUDE.md #657 — Related documentation edits to CLAUDE.md (fallback guidance and SSL notes).

Suggested reviewers

  • moshemorad
  • arikalon1
🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 32.43% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: adding a generic HTTP/curl tool for whitelisted domains, which aligns with the substantial additions of the HttpToolset implementation and related infrastructure changes.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ⚠️ Completed with 3 failures


🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Branch claude/generic-curl-tool-Q555c
Model bedrock/eu.anthropic.claude-sonnet-4-5-20250929-v1:0
Markers confluence
Iterations 1
Duration 1m 31s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 0/3 test cases were successful, 3 regressions
Status Test case Time Turns Tools Cost
❌ 208_confluence_page_fetch 14.6s 3 3 $0.0767
❌ 209_confluence_url_lookup 14.8s 3 3 $0.0766
❌ 210_confluence_runbook_fetch 15.8s 3 3 $0.0783
Total 15.1s avg 3.0 avg 3.0 avg $0.2316

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'master'

Status: Success - 39 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 3 Failures Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Fix all issues with AI agents
In `@holmes/plugins/toolsets/http/http_toolset.py`:
- Around line 304-307: The _invoke method declares a context parameter that is
unused and triggers Ruff ARG002; to fix, explicitly silence it inside _invoke
(e.g., assign context to _ or del context) so the name remains present for the
API but Ruff no longer flags it; update the method body in the _invoke function
to add a single statement like "_ = context" or "del context" immediately after
the signature.
- Around line 338-346: Move the json import out of the function to module scope
and validate the parsed headers are a JSON object before using them: in the
block handling extra_headers_str (variables extra_headers and
extra_headers_str), replace the local import with the top-level import of json
and after json.loads(...) ensure the result is a dict (e.g., isinstance(result,
dict)); if it is not, return a StructuredToolResult error indicating invalid
header format so that headers.update(...) won't receive a non-dict.
- Around line 379-395: When building the StructuredToolResult in the HTTP call
path (the block that parses response into data and returns self.filter_result),
populate the result.error field on non-OK responses with a concise string
containing the HTTP status code, the response body (or parsed JSON), and the
exact invocation (include url and params) so the LLM can self-correct; keep
using StructuredToolResultStatus.ERROR when not response.ok and leave error
empty for SUCCESS, then return self.filter_result(result, params) as before
(refer to StructuredToolResult, StructuredToolResultStatus, response, params,
url, and filter_result to locate the code).
🧹 Nitpick comments (3)
holmes/plugins/toolsets/http/http_toolset.py (1)

105-127: Add endpoint health checks during prerequisites.

prerequisites_callable validates config but never checks endpoint reachability. Consider adding an optional per-endpoint health path (or lightweight HEAD/GET check) so misconfigured or unreachable endpoints fail fast.

As per coding guidelines, include a health check in prerequisites_callable().

tests/plugins/toolsets/http/test_http_toolset.py (2)

167-212: LGTM! Thorough header construction validation.

The tests correctly validate:

  • Bearer and custom header authentication in headers
  • Basic auth properly separated into tuple format (correct for requests library usage, per learnings)
  • Extra headers override behavior

Note: Static analysis warnings about hardcoded passwords on lines 194, 203 are false positives—these are test fixtures.

Optional: Consider adding explicit test for default_headers

While default_headers are tested indirectly via build_headers, an explicit test would improve clarity:

def test_default_headers_included(self):
    ts = HttpToolset()
    ts._http_config = HttpToolsetConfig(
        endpoints=[EndpointConfig(host="example.com")],
        default_headers={"X-Custom": "value"}
    )
    endpoint = EndpointConfig(host="example.com", auth=AuthConfig(type="none"))
    headers = ts.build_headers(endpoint)
    assert headers["X-Custom"] == "value"

214-242: LGTM! Prerequisites validation covers key scenarios.

The tests validate configuration validation, empty endpoint detection, and authentication error handling. The flexible message assertions appropriately verify key information without being overly brittle.

Optional: Consider adding test for multiple endpoints

To verify the endpoint/host counting logic more thoroughly:

def test_multiple_endpoints(self):
    toolset = HttpToolset()
    success, message = toolset.prerequisites_callable(
        {
            "endpoints": [
                {"host": "example.com", "auth": {"type": "none"}},
                {"host": ["api1.com", "api2.com"], "auth": {"type": "none"}},
            ]
        }
    )
    assert success is True
    assert "2 endpoint" in message
    assert "3 host pattern" in message

Comment thread holmes/plugins/toolsets/http/http_toolset.py
Comment thread holmes/plugins/toolsets/http/http_toolset.py
Comment thread holmes/plugins/toolsets/http/http_toolset.py
claude and others added 6 commits January 22, 2026 07:55
- Fix unused context parameter with _ = context
- Remove duplicate json import (move to module top)
- Add dict type validation for headers JSON parameter
- Add error field when HTTP response is not OK
- Add optional health_check_url field for auth validation at init
- Add comprehensive tests for new functionality (50 tests total)

Signed-off-by: Claude <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Use /wiki/rest/api/space?limit=1 instead of /user/current for health
check - this endpoint only requires read access to spaces, avoiding
potential scope issues with service account tokens.

Signed-off-by: Claude <noreply@anthropic.com>
@aantn

aantn commented Jan 22, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
markers: confluence

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ✅ Completed successfully


🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Branch claude/generic-curl-tool-Q555c
Model bedrock/eu.anthropic.claude-opus-4-5-20251101-v1:0
Markers confluence
Iterations 1
Duration 1m 15s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 0/3 test cases were successful, 0 regressions, 3 setup failures
Status Test case Time Turns Tools Cost
🚧 208_confluence_page_fetch — — — —
🚧 209_confluence_url_lookup — — — —
🚧 210_confluence_runbook_fetch — — — —
Total — avg — avg — avg —

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'master'

Status: Success - 36 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=

The API token has content-read permissions but not space-list permissions,
causing /wiki/rest/api/space to return 403. Since health check is optional
and actual API calls will fail with descriptive errors anyway, removing
the health check is the simplest fix.

Signed-off-by: Claude <noreply@anthropic.com>
When health checks fail, include a curl command (with secrets redacted)
to help users troubleshoot authentication issues manually.

Signed-off-by: Claude <noreply@anthropic.com>
This endpoint uses the same scope (read:confluence-content.summary) that
the evals need for reading pages. The previous /space endpoint required
a different scope (read:confluence-space.summary) that wasn't granted.

Signed-off-by: Claude <noreply@anthropic.com>
@aantn

aantn commented Jan 22, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
markers: confluence

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ✅ Completed successfully


🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Branch claude/generic-curl-tool-Q555c
Model bedrock/eu.anthropic.claude-opus-4-5-20251101-v1:0
Markers confluence
Iterations 1
Duration 1m 38s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 0/3 test cases were successful, 0 regressions, 3 setup failures
Status Test case Time Turns Tools Cost
🚧 208_confluence_page_fetch — — — —
🚧 209_confluence_url_lookup — — — —
🚧 210_confluence_runbook_fetch — — — —
Total — avg — avg — avg —

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'master'

Status: Success - 36 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=

Two auth methods are now supported for Confluence:

1. Classic Auth (208-210): Personal API tokens
   - URL: https://{site}.atlassian.net/wiki/rest/api/...
   - Auth: Basic auth (email:token)
   - Env: CONFLUENCE_USER_NAME, CONFLUENCE_API_KEY

2. Service Account Auth (211-213): Scoped tokens
   - URL: https://api.atlassian.com/ex/confluence/{cloudId}/wiki/rest/api/...
   - Auth: Bearer token
   - Env: CONFLUENCE_CLOUD_ID, CONFLUENCE_SERVICE_ACCOUNT_TOKEN

Also updated HTTP toolset instructions to document both patterns.

Signed-off-by: Claude <noreply@anthropic.com>
Instead of requiring CONFLUENCE_CLOUD_ID env var, the LLM now discovers
the cloudId automatically by:

1. Calling /oauth/token/accessible-resources to get list of sites
2. Matching the Confluence URL from the user's question
3. Using the cloudId to construct the gateway URL

This makes service account setup easier - users only need:
- CONFLUENCE_SERVICE_ACCOUNT_TOKEN

Updated:
- Service account eval configs (211-213) to whitelist accessible-resources
- HTTP toolset instructions with cloudId discovery workflow
- Expected outputs to verify cloudId discovery

Signed-off-by: Claude <noreply@anthropic.com>
- Add instructions field to HttpToolsetConfig to allow per-config custom LLM instructions
- Make instructions.jinja2 generic (remove Confluence-specific content)
- Custom instructions are appended under "## API-Specific Instructions" header
- Update all Confluence eval toolsets with appropriate instructions:
  - Classic auth (208-210): Site-specific URL patterns
  - Service account auth (211-213): CloudId discovery and gateway URL patterns

Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In `@holmes/plugins/toolsets/http/http_toolset.py`:
- Around line 266-275: The hostname used for whitelist matching is being taken
from parsed.netloc which can include a port and cause mismatches; update the
code that assigns host (in the URL parsing block where parsed = urlparse(url)
and host = parsed.netloc) to use parsed.hostname (falling back to parsed.netloc
if hostname is None) so you normalize hostnames and avoid port leakage when
comparing against host patterns.
- Around line 401-505: The _invoke method returns error StructuredToolResult
objects without the required invocation details; update every error path in
_invoke (e.g., URL validation branch after self._toolset.match_endpoint,
unsupported method branch, method-not-allowed branch using
self._toolset.is_method_allowed, headers validation/JSONDecode branches, and the
non-OK HTTP response branch) to include an invocation field containing the exact
invocation string (f"{method} {url}" or equivalent). Ensure the invocation is
added to the StructuredToolResult constructions for all error returns and
preserved when calling self.filter_result(result, params).

Comment on lines +266 to +275
try:
parsed = urlparse(url)
except Exception as e:
return None, f"Invalid URL: {e}"

if not parsed.scheme or not parsed.netloc:
return None, f"Invalid URL format: {url}"

host = parsed.netloc
path = parsed.path or "/"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Normalize hostnames before whitelist matching.
urlparse(...).netloc includes ports (e.g., example.com:8443), which can cause false negatives against host patterns. Use parsed.hostname to avoid port leakage in comparisons.

🐛 Proposed fix
-        host = parsed.netloc
+        host = parsed.hostname or ""
+        if not host:
+            return None, f"Invalid URL format: {url}"
🧰 Tools
🪛 Ruff (0.14.13)

268-268: Do not catch blind exception: Exception

(BLE001)

🤖 Prompt for AI Agents
In `@holmes/plugins/toolsets/http/http_toolset.py` around lines 266 - 275, The
hostname used for whitelist matching is being taken from parsed.netloc which can
include a port and cause mismatches; update the code that assigns host (in the
URL parsing block where parsed = urlparse(url) and host = parsed.netloc) to use
parsed.hostname (falling back to parsed.netloc if hostname is None) so you
normalize hostnames and avoid port leakage when comparing against host patterns.

Comment on lines +401 to +505
def _invoke(self, params: dict, context: ToolInvokeContext) -> StructuredToolResult:
_ = context # Required by interface but not used
url = params.get("url", "")
method = params.get("method", "GET").upper()
body = params.get("body")
extra_headers_str = params.get("headers")

# Validate URL against whitelist
endpoint, error = self._toolset.match_endpoint(url)
if error or endpoint is None:
return StructuredToolResult(
status=StructuredToolResultStatus.ERROR,
error=error or "URL not matched",
params=params,
url=url,
)

# Validate method
if method not in ("GET", "POST"):
return StructuredToolResult(
status=StructuredToolResultStatus.ERROR,
error=f"Unsupported HTTP method: {method}. Only GET and POST are supported.",
params=params,
url=url,
)

# Check if method is allowed for this endpoint
if not self._toolset.is_method_allowed(method, endpoint):
return StructuredToolResult(
status=StructuredToolResultStatus.ERROR,
error=f"Method {method} not allowed for this endpoint. Allowed methods: {endpoint.get_methods()}",
params=params,
url=url,
)

# Parse extra headers if provided
extra_headers = None
if extra_headers_str:
try:
extra_headers = json.loads(extra_headers_str)
if not isinstance(extra_headers, dict):
return StructuredToolResult(
status=StructuredToolResultStatus.ERROR,
error="Headers must be a JSON object, not a list or primitive",
params=params,
url=url,
)
except json.JSONDecodeError as e:
return StructuredToolResult(
status=StructuredToolResultStatus.ERROR,
error=f"Invalid headers JSON: {e}",
params=params,
url=url,
)

# Build headers
headers = self._toolset.build_headers(endpoint, extra_headers)

# Get auth
basic_auth = self._toolset.get_basic_auth(endpoint)

# Make request
try:
if method == "GET":
response = requests.get(
url,
headers=headers,
auth=basic_auth,
timeout=self._toolset.http_config.timeout_seconds,
verify=self._toolset.http_config.verify_ssl,
)
else: # POST
response = requests.post(
url,
headers=headers,
auth=basic_auth,
data=body,
timeout=self._toolset.http_config.timeout_seconds,
verify=self._toolset.http_config.verify_ssl,
)

# Return raw response (status + body)
try:
data = response.json()
except Exception:
data = response.text

if response.ok:
result = StructuredToolResult(
status=StructuredToolResultStatus.SUCCESS,
data={"status_code": response.status_code, "body": data},
params=params,
url=url,
)
else:
result = StructuredToolResult(
status=StructuredToolResultStatus.ERROR,
error=f"HTTP {response.status_code}: {data}",
data={"status_code": response.status_code, "body": data},
params=params,
url=url,
)

# Apply JSON filtering from mixin
return self.filter_result(result, params)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Include invocation details in error results for LLM self‑correction.
Error responses should include the exact invocation (method + URL) so the LLM can self-correct. Currently, error paths (including non‑OK HTTP responses and validation failures) omit invocation.

🔧 Suggested fix
     def _invoke(self, params: dict, context: ToolInvokeContext) -> StructuredToolResult:
         _ = context  # Required by interface but not used
         url = params.get("url", "")
         method = params.get("method", "GET").upper()
+        invocation = self.get_parameterized_one_liner(params)
         body = params.get("body")
         extra_headers_str = params.get("headers")
@@
             return StructuredToolResult(
                 status=StructuredToolResultStatus.ERROR,
                 error=error or "URL not matched",
                 params=params,
                 url=url,
+                invocation=invocation,
             )
@@
                 result = StructuredToolResult(
                     status=StructuredToolResultStatus.SUCCESS,
                     data={"status_code": response.status_code, "body": data},
                     params=params,
                     url=url,
+                    invocation=invocation,
                 )
             else:
                 result = StructuredToolResult(
                     status=StructuredToolResultStatus.ERROR,
                     error=f"HTTP {response.status_code}: {data}",
                     data={"status_code": response.status_code, "body": data},
                     params=params,
                     url=url,
+                    invocation=invocation,
                 )
As per coding guidelines, error results must include the exact invocation details.
🧰 Tools
🪛 Ruff (0.14.13)

485-485: Do not catch blind exception: Exception

(BLE001)

🤖 Prompt for AI Agents
In `@holmes/plugins/toolsets/http/http_toolset.py` around lines 401 - 505, The
_invoke method returns error StructuredToolResult objects without the required
invocation details; update every error path in _invoke (e.g., URL validation
branch after self._toolset.match_endpoint, unsupported method branch,
method-not-allowed branch using self._toolset.is_method_allowed, headers
validation/JSONDecode branches, and the non-OK HTTP response branch) to include
an invocation field containing the exact invocation string (f"{method} {url}" or
equivalent). Ensure the invocation is added to the StructuredToolResult
constructions for all error returns and preserved when calling
self.filter_result(result, params).

@aantn

aantn commented Jan 31, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
markers: confluence

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ⚠️ Completed with 6 failures


🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Branch claude/generic-curl-tool-Q555c
Model opus-4.5
Markers confluence
Iterations 1
Duration 1m 40s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 0/6 test cases were successful, 6 regressions
Status Test case Time Turns Tools Cost
❌ 208_confluence_page_fetch 21.8s 4 5 $0.1457
❌ 209_confluence_url_lookup 21.4s 4 4 $0.1402
❌ 210_confluence_runbook_fetch 16.1s 3 3 $0.1280
❌ 211_confluence_page_fetch_service_account 18.4s 3 3 $0.1257
❌ 212_confluence_url_lookup_service_account 18.6s 3 3 $0.1252
❌ 213_confluence_runbook_fetch_service_account 18.8s 3 3 $0.1253
Total 19.2s avg 3.3 avg 3.5 avg $0.7902

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'master'

Status: Success - 18 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 6 Failures Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence-service-account, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=

@arikalon1

Copy link
Copy Markdown
Collaborator

/eval
markers: confluence

@github-actions

github-actions Bot commented Feb 8, 2026

Copy link
Copy Markdown
Contributor

@arikalon1 Your eval run has finished. ⚠️ Completed with 6 failures


🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Branch claude/generic-curl-tool-Q555c
Model opus-4.5
Markers confluence
Iterations 1
Duration 2m 55s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 0/6 test cases were successful, 6 regressions
Status Test case Time Turns Tools Cost
❌ 208_confluence_page_fetch 15.8s 3 3 $0.1282
❌ 209_confluence_url_lookup 86.4s 3 3 $0.1280
❌ 210_confluence_runbook_fetch 21.4s 4 5 $0.1489
❌ 211_confluence_page_fetch_service_account 16.8s 3 3 $0.1230
❌ 212_confluence_url_lookup_service_account 17.4s 3 3 $0.1242
❌ 213_confluence_runbook_fetch_service_account 16.1s 3 3 $0.1217
Total 29.0s avg 3.2 avg 3.3 avg $0.7740

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'master'

Status: Success - 9 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 6 Failures Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence-service-account, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=

@aantn aantn closed this Feb 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants