Conversation
Introduces a new HTTP toolset that allows making requests to
user-configured API endpoints. This provides a generic alternative
to creating dedicated toolsets for simple API integrations.
Features:
- Host-based whitelisting with wildcard support (*.example.com)
- Optional path restrictions per endpoint
- Configurable HTTP methods (GET by default, POST opt-in)
- Multiple auth types: basic, bearer, custom header
- JSON filtering via JsonFilterMixin (jq, max_depth)
- Environment variable support for secrets ({{ env.FOO }})
Configuration example:
http:
endpoints:
- host: "*.atlassian.net"
auth:
type: basic
username: "{{ env.CONFLUENCE_USER }}"
password: "{{ env.CONFLUENCE_API_KEY }}"
This toolset can replace simpler curl-based toolsets like Confluence
while offering more flexibility for ad-hoc API access.
Signed-off-by: Claude <noreply@anthropic.com>
- Remove dedicated confluence.yaml toolset (single curl command) - Update Confluence evals (208, 209, 210) to use the HTTP toolset - Add Confluence API patterns to HTTP instructions for LLM guidance - Fix env var name: CONFLUENCE_USER_NAME (not CONFLUENCE_USER) The HTTP toolset provides more flexibility: - LLM can construct any valid Confluence API URL - Supports search by title (not just page ID fetching) - Can expand to other Atlassian APIs with same auth This validates the HTTP toolset as a replacement for simple curl-based toolsets while offering more capability. Signed-off-by: Claude <noreply@anthropic.com>
Add documentation for using OpenRouter when OPENAI_API_KEY is not available. This helps Claude Code know to use OPENROUTER_API_KEY as a fallback for running LLM evaluation tests. Includes: - OpenRouter model format examples - Command to check available API keys - Note about using same model for CLASSIFIER_MODEL Signed-off-by: Claude <noreply@anthropic.com>
- Use Opus 4.5 as the primary recommended model - Add stronger emphasis to always try OpenRouter before giving up - Add "check available keys" step at the beginning - Make it clear this is the fallback when OpenAI key is missing Signed-off-by: Claude <noreply@anthropic.com>
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
|
|
📂 Previous Runs📜 Run @ 5f49f60 (#21248353706)✅ Results of HolmesGPT evalsAutomatically triggered by commit 5f49f60 on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/generic-curl-tool-Q555c' Status: Success - 36 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📜 Run @ c17970d (#21247763693)✅ Results of HolmesGPT evalsAutomatically triggered by commit c17970d on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/generic-curl-tool-Q555c' Status: Success - 36 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📜 Run @ bcb2960 (#21247433547)✅ Results of HolmesGPT evalsAutomatically triggered by commit bcb2960 on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/generic-curl-tool-Q555c' Status: Success - 36 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📜 Run @ fe74b51 (#21245829669)✅ Results of HolmesGPT evalsAutomatically triggered by commit fe74b51 on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/generic-curl-tool-Q555c' Status: Success - 36 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📜 Run @ 0597077 (#21245668606)✅ Results of HolmesGPT evalsAutomatically triggered by commit 0597077 on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/generic-curl-tool-Q555c' Status: Success - 36 test/model combinations loaded Experiments compared (30):
Comparison indicators:
✅ Results of HolmesGPT evalsAutomatically triggered by commit f967155 on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/generic-curl-tool-Q555c' Status: Success - 9 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
Commands: CLI: |
|
✅ Docker image ready for
Use this tag to pull the image for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e8b2470
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e8b2470 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e8b2470
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e8b2470Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:e8b2470Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:e8b2470 |
|
/eval |
|
Warning Rate limit exceeded
⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. WalkthroughAdds a new generic HTTP toolset (implementation, template, registration, and tests), migrates Confluence fixtures to use the HTTP toolset (classic and service-account flows), removes the Confluence YAML toolset, and extends CLAUDE.md with OpenRouter/OpenAI fallback and SSL_VERIFY sandbox guidance. Changes
Sequence Diagram(s)sequenceDiagram
participant User as User/Holmes
participant HttpTool as HttpRequest
participant Toolset as HttpToolset
participant Executor as HTTP Executor
participant Service as External HTTP Service
User->>HttpTool: Invoke(url, method, body, headers)
HttpTool->>Toolset: match_endpoint(url)
Toolset->>Toolset: parse host & path, match host pattern, match path patterns
Toolset-->>HttpTool: return EndpointConfig or error
HttpTool->>HttpTool: validate method allowed
HttpTool->>Toolset: build_headers(endpoint, extra_headers)
Toolset-->>HttpTool: return headers and basic auth tuple (if any)
HttpTool->>Executor: execute request(method, url, headers, auth, timeout, verify)
Executor->>Service: send HTTP request
Service-->>Executor: response
Executor-->>HttpTool: parse JSON or text, handle errors/timeouts
HttpTool-->>User: return StructuredToolResult(status, data/error)
Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
@aantn Your eval run has finished. 🧪 Manual Eval Results
Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'master' Status: Success - 39 test/model combinations loaded Experiments compared (30):
Comparison indicators:
|
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
markers: regression
Or with more options (one per line):
/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
markers: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
markers |
Pytest markers (no default - runs all tests!) |
filter |
Pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
🏷️ Valid markers
benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=
There was a problem hiding this comment.
Actionable comments posted: 3
🤖 Fix all issues with AI agents
In `@holmes/plugins/toolsets/http/http_toolset.py`:
- Around line 304-307: The _invoke method declares a context parameter that is
unused and triggers Ruff ARG002; to fix, explicitly silence it inside _invoke
(e.g., assign context to _ or del context) so the name remains present for the
API but Ruff no longer flags it; update the method body in the _invoke function
to add a single statement like "_ = context" or "del context" immediately after
the signature.
- Around line 338-346: Move the json import out of the function to module scope
and validate the parsed headers are a JSON object before using them: in the
block handling extra_headers_str (variables extra_headers and
extra_headers_str), replace the local import with the top-level import of json
and after json.loads(...) ensure the result is a dict (e.g., isinstance(result,
dict)); if it is not, return a StructuredToolResult error indicating invalid
header format so that headers.update(...) won't receive a non-dict.
- Around line 379-395: When building the StructuredToolResult in the HTTP call
path (the block that parses response into data and returns self.filter_result),
populate the result.error field on non-OK responses with a concise string
containing the HTTP status code, the response body (or parsed JSON), and the
exact invocation (include url and params) so the LLM can self-correct; keep
using StructuredToolResultStatus.ERROR when not response.ok and leave error
empty for SUCCESS, then return self.filter_result(result, params) as before
(refer to StructuredToolResult, StructuredToolResultStatus, response, params,
url, and filter_result to locate the code).
🧹 Nitpick comments (3)
holmes/plugins/toolsets/http/http_toolset.py (1)
105-127: Add endpoint health checks during prerequisites.
prerequisites_callablevalidates config but never checks endpoint reachability. Consider adding an optional per-endpoint health path (or lightweight HEAD/GET check) so misconfigured or unreachable endpoints fail fast.As per coding guidelines, include a health check in
prerequisites_callable().tests/plugins/toolsets/http/test_http_toolset.py (2)
167-212: LGTM! Thorough header construction validation.The tests correctly validate:
- Bearer and custom header authentication in headers
- Basic auth properly separated into tuple format (correct for requests library usage, per learnings)
- Extra headers override behavior
Note: Static analysis warnings about hardcoded passwords on lines 194, 203 are false positives—these are test fixtures.
Optional: Consider adding explicit test for default_headers
While default_headers are tested indirectly via
build_headers, an explicit test would improve clarity:def test_default_headers_included(self): ts = HttpToolset() ts._http_config = HttpToolsetConfig( endpoints=[EndpointConfig(host="example.com")], default_headers={"X-Custom": "value"} ) endpoint = EndpointConfig(host="example.com", auth=AuthConfig(type="none")) headers = ts.build_headers(endpoint) assert headers["X-Custom"] == "value"
214-242: LGTM! Prerequisites validation covers key scenarios.The tests validate configuration validation, empty endpoint detection, and authentication error handling. The flexible message assertions appropriately verify key information without being overly brittle.
Optional: Consider adding test for multiple endpoints
To verify the endpoint/host counting logic more thoroughly:
def test_multiple_endpoints(self): toolset = HttpToolset() success, message = toolset.prerequisites_callable( { "endpoints": [ {"host": "example.com", "auth": {"type": "none"}}, {"host": ["api1.com", "api2.com"], "auth": {"type": "none"}}, ] } ) assert success is True assert "2 endpoint" in message assert "3 host pattern" in message
- Fix unused context parameter with _ = context - Remove duplicate json import (move to module top) - Add dict type validation for headers JSON parameter - Add error field when HTTP response is not OK - Add optional health_check_url field for auth validation at init - Add comprehensive tests for new functionality (50 tests total) Signed-off-by: Claude <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Use /wiki/rest/api/space?limit=1 instead of /user/current for health check - this endpoint only requires read access to spaces, avoiding potential scope issues with service account tokens. Signed-off-by: Claude <noreply@anthropic.com>
|
/eval |
|
@aantn Your eval run has finished. ✅ Completed successfully 🧪 Manual Eval Results
Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'master' Status: Success - 36 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
Commands: CLI: |
The API token has content-read permissions but not space-list permissions, causing /wiki/rest/api/space to return 403. Since health check is optional and actual API calls will fail with descriptive errors anyway, removing the health check is the simplest fix. Signed-off-by: Claude <noreply@anthropic.com>
When health checks fail, include a curl command (with secrets redacted) to help users troubleshoot authentication issues manually. Signed-off-by: Claude <noreply@anthropic.com>
This endpoint uses the same scope (read:confluence-content.summary) that the evals need for reading pages. The previous /space endpoint required a different scope (read:confluence-space.summary) that wasn't granted. Signed-off-by: Claude <noreply@anthropic.com>
|
/eval |
|
@aantn Your eval run has finished. ✅ Completed successfully 🧪 Manual Eval Results
Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'master' Status: Success - 36 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
Commands: CLI: |
Two auth methods are now supported for Confluence:
1. Classic Auth (208-210): Personal API tokens
- URL: https://{site}.atlassian.net/wiki/rest/api/...
- Auth: Basic auth (email:token)
- Env: CONFLUENCE_USER_NAME, CONFLUENCE_API_KEY
2. Service Account Auth (211-213): Scoped tokens
- URL: https://api.atlassian.com/ex/confluence/{cloudId}/wiki/rest/api/...
- Auth: Bearer token
- Env: CONFLUENCE_CLOUD_ID, CONFLUENCE_SERVICE_ACCOUNT_TOKEN
Also updated HTTP toolset instructions to document both patterns.
Signed-off-by: Claude <noreply@anthropic.com>
Instead of requiring CONFLUENCE_CLOUD_ID env var, the LLM now discovers the cloudId automatically by: 1. Calling /oauth/token/accessible-resources to get list of sites 2. Matching the Confluence URL from the user's question 3. Using the cloudId to construct the gateway URL This makes service account setup easier - users only need: - CONFLUENCE_SERVICE_ACCOUNT_TOKEN Updated: - Service account eval configs (211-213) to whitelist accessible-resources - HTTP toolset instructions with cloudId discovery workflow - Expected outputs to verify cloudId discovery Signed-off-by: Claude <noreply@anthropic.com>
- Add instructions field to HttpToolsetConfig to allow per-config custom LLM instructions - Make instructions.jinja2 generic (remove Confluence-specific content) - Custom instructions are appended under "## API-Specific Instructions" header - Update all Confluence eval toolsets with appropriate instructions: - Classic auth (208-210): Site-specific URL patterns - Service account auth (211-213): CloudId discovery and gateway URL patterns Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Fix all issues with AI agents
In `@holmes/plugins/toolsets/http/http_toolset.py`:
- Around line 266-275: The hostname used for whitelist matching is being taken
from parsed.netloc which can include a port and cause mismatches; update the
code that assigns host (in the URL parsing block where parsed = urlparse(url)
and host = parsed.netloc) to use parsed.hostname (falling back to parsed.netloc
if hostname is None) so you normalize hostnames and avoid port leakage when
comparing against host patterns.
- Around line 401-505: The _invoke method returns error StructuredToolResult
objects without the required invocation details; update every error path in
_invoke (e.g., URL validation branch after self._toolset.match_endpoint,
unsupported method branch, method-not-allowed branch using
self._toolset.is_method_allowed, headers validation/JSONDecode branches, and the
non-OK HTTP response branch) to include an invocation field containing the exact
invocation string (f"{method} {url}" or equivalent). Ensure the invocation is
added to the StructuredToolResult constructions for all error returns and
preserved when calling self.filter_result(result, params).
| try: | ||
| parsed = urlparse(url) | ||
| except Exception as e: | ||
| return None, f"Invalid URL: {e}" | ||
|
|
||
| if not parsed.scheme or not parsed.netloc: | ||
| return None, f"Invalid URL format: {url}" | ||
|
|
||
| host = parsed.netloc | ||
| path = parsed.path or "/" |
There was a problem hiding this comment.
Normalize hostnames before whitelist matching.
urlparse(...).netloc includes ports (e.g., example.com:8443), which can cause false negatives against host patterns. Use parsed.hostname to avoid port leakage in comparisons.
🐛 Proposed fix
- host = parsed.netloc
+ host = parsed.hostname or ""
+ if not host:
+ return None, f"Invalid URL format: {url}"🧰 Tools
🪛 Ruff (0.14.13)
268-268: Do not catch blind exception: Exception
(BLE001)
🤖 Prompt for AI Agents
In `@holmes/plugins/toolsets/http/http_toolset.py` around lines 266 - 275, The
hostname used for whitelist matching is being taken from parsed.netloc which can
include a port and cause mismatches; update the code that assigns host (in the
URL parsing block where parsed = urlparse(url) and host = parsed.netloc) to use
parsed.hostname (falling back to parsed.netloc if hostname is None) so you
normalize hostnames and avoid port leakage when comparing against host patterns.
| def _invoke(self, params: dict, context: ToolInvokeContext) -> StructuredToolResult: | ||
| _ = context # Required by interface but not used | ||
| url = params.get("url", "") | ||
| method = params.get("method", "GET").upper() | ||
| body = params.get("body") | ||
| extra_headers_str = params.get("headers") | ||
|
|
||
| # Validate URL against whitelist | ||
| endpoint, error = self._toolset.match_endpoint(url) | ||
| if error or endpoint is None: | ||
| return StructuredToolResult( | ||
| status=StructuredToolResultStatus.ERROR, | ||
| error=error or "URL not matched", | ||
| params=params, | ||
| url=url, | ||
| ) | ||
|
|
||
| # Validate method | ||
| if method not in ("GET", "POST"): | ||
| return StructuredToolResult( | ||
| status=StructuredToolResultStatus.ERROR, | ||
| error=f"Unsupported HTTP method: {method}. Only GET and POST are supported.", | ||
| params=params, | ||
| url=url, | ||
| ) | ||
|
|
||
| # Check if method is allowed for this endpoint | ||
| if not self._toolset.is_method_allowed(method, endpoint): | ||
| return StructuredToolResult( | ||
| status=StructuredToolResultStatus.ERROR, | ||
| error=f"Method {method} not allowed for this endpoint. Allowed methods: {endpoint.get_methods()}", | ||
| params=params, | ||
| url=url, | ||
| ) | ||
|
|
||
| # Parse extra headers if provided | ||
| extra_headers = None | ||
| if extra_headers_str: | ||
| try: | ||
| extra_headers = json.loads(extra_headers_str) | ||
| if not isinstance(extra_headers, dict): | ||
| return StructuredToolResult( | ||
| status=StructuredToolResultStatus.ERROR, | ||
| error="Headers must be a JSON object, not a list or primitive", | ||
| params=params, | ||
| url=url, | ||
| ) | ||
| except json.JSONDecodeError as e: | ||
| return StructuredToolResult( | ||
| status=StructuredToolResultStatus.ERROR, | ||
| error=f"Invalid headers JSON: {e}", | ||
| params=params, | ||
| url=url, | ||
| ) | ||
|
|
||
| # Build headers | ||
| headers = self._toolset.build_headers(endpoint, extra_headers) | ||
|
|
||
| # Get auth | ||
| basic_auth = self._toolset.get_basic_auth(endpoint) | ||
|
|
||
| # Make request | ||
| try: | ||
| if method == "GET": | ||
| response = requests.get( | ||
| url, | ||
| headers=headers, | ||
| auth=basic_auth, | ||
| timeout=self._toolset.http_config.timeout_seconds, | ||
| verify=self._toolset.http_config.verify_ssl, | ||
| ) | ||
| else: # POST | ||
| response = requests.post( | ||
| url, | ||
| headers=headers, | ||
| auth=basic_auth, | ||
| data=body, | ||
| timeout=self._toolset.http_config.timeout_seconds, | ||
| verify=self._toolset.http_config.verify_ssl, | ||
| ) | ||
|
|
||
| # Return raw response (status + body) | ||
| try: | ||
| data = response.json() | ||
| except Exception: | ||
| data = response.text | ||
|
|
||
| if response.ok: | ||
| result = StructuredToolResult( | ||
| status=StructuredToolResultStatus.SUCCESS, | ||
| data={"status_code": response.status_code, "body": data}, | ||
| params=params, | ||
| url=url, | ||
| ) | ||
| else: | ||
| result = StructuredToolResult( | ||
| status=StructuredToolResultStatus.ERROR, | ||
| error=f"HTTP {response.status_code}: {data}", | ||
| data={"status_code": response.status_code, "body": data}, | ||
| params=params, | ||
| url=url, | ||
| ) | ||
|
|
||
| # Apply JSON filtering from mixin | ||
| return self.filter_result(result, params) |
There was a problem hiding this comment.
Include invocation details in error results for LLM self‑correction.
Error responses should include the exact invocation (method + URL) so the LLM can self-correct. Currently, error paths (including non‑OK HTTP responses and validation failures) omit invocation.
🔧 Suggested fix
def _invoke(self, params: dict, context: ToolInvokeContext) -> StructuredToolResult:
_ = context # Required by interface but not used
url = params.get("url", "")
method = params.get("method", "GET").upper()
+ invocation = self.get_parameterized_one_liner(params)
body = params.get("body")
extra_headers_str = params.get("headers")
@@
return StructuredToolResult(
status=StructuredToolResultStatus.ERROR,
error=error or "URL not matched",
params=params,
url=url,
+ invocation=invocation,
)
@@
result = StructuredToolResult(
status=StructuredToolResultStatus.SUCCESS,
data={"status_code": response.status_code, "body": data},
params=params,
url=url,
+ invocation=invocation,
)
else:
result = StructuredToolResult(
status=StructuredToolResultStatus.ERROR,
error=f"HTTP {response.status_code}: {data}",
data={"status_code": response.status_code, "body": data},
params=params,
url=url,
+ invocation=invocation,
)🧰 Tools
🪛 Ruff (0.14.13)
485-485: Do not catch blind exception: Exception
(BLE001)
🤖 Prompt for AI Agents
In `@holmes/plugins/toolsets/http/http_toolset.py` around lines 401 - 505, The
_invoke method returns error StructuredToolResult objects without the required
invocation details; update every error path in _invoke (e.g., URL validation
branch after self._toolset.match_endpoint, unsupported method branch,
method-not-allowed branch using self._toolset.is_method_allowed, headers
validation/JSONDecode branches, and the non-OK HTTP response branch) to include
an invocation field containing the exact invocation string (f"{method} {url}" or
equivalent). Ensure the invocation is added to the StructuredToolResult
constructions for all error returns and preserved when calling
self.filter_result(result, params).
|
/eval |
|
@aantn Your eval run has finished. 🧪 Manual Eval Results
Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'master' Status: Success - 18 test/model combinations loaded Experiments compared (30):
Comparison indicators:
|
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
markers: regression
Or with more options (one per line):
/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
markers: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
markers |
Pytest markers (no default - runs all tests!) |
filter |
Pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
🏷️ Valid markers
benchmark, chain-of-causation, compaction, confluence-service-account, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=
|
/eval |
|
@arikalon1 Your eval run has finished. 🧪 Manual Eval Results
Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'master' Status: Success - 9 test/model combinations loaded Experiments compared (30):
Comparison indicators:
|
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
markers: regression
Or with more options (one per line):
/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
markers: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
markers |
Pytest markers (no default - runs all tests!) |
filter |
Pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
🏷️ Valid markers
benchmark, chain-of-causation, compaction, confluence-service-account, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/generic-curl-tool-Q555c -f markers=regression -f filter=
and try to replace confluence tool with it
Summary by CodeRabbit
New Features
Documentation
Tests
✏️ Tip: You can customize this high-level summary in your review settings.