Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
34 commits
Select commit Hold shift + click to select a range
48a850f
Pass MCP ImageContent through to vision-capable LLMs
claude Mar 17, 2026
5e7c478
Fix MCP toolset silently dropped when description missing, fix trunca…
claude Mar 17, 2026
13092d0
Fix eval 233 auth (BYOT OAuth), detect silently dropped toolsets in t…
claude Mar 18, 2026
4fa75d3
fix(eval-233): switch mcp-atlassian from OAuth to basic auth
claude Mar 19, 2026
edeac5f
fix(eval-233): use static image, basic auth, and SA setup creds
claude Mar 19, 2026
fcc43e8
fix(eval-233): use SA creds via API gateway, add /wiki to URL
claude Mar 19, 2026
b294905
test: add real MCP server image passthrough integration test
claude Mar 19, 2026
ae416aa
feat(eval-234): add local MCP image attachment eval (no external deps)
claude Mar 20, 2026
f9567c9
Merge remote-tracking branch 'origin/master' into claude/fix-eval-233…
claude Mar 20, 2026
165bb4c
chore(eval-234): add mcp tag after merging master
claude Mar 20, 2026
607d3f3
feat: spill images to disk when tool results exceed context window
claude Mar 20, 2026
98181da
feat: rename spill function, add context management docs, add eval 236
claude Mar 20, 2026
abe4452
feat: add Grafana vision rendering, image embed hints, and compaction…
claude Mar 20, 2026
b6df5d5
chore: add 'images' tag to all image-related evals for easy filtering
claude Mar 20, 2026
4cedc73
Merge branch 'master' into claude/fix-eval-233-7keOj
aantn Mar 20, 2026
cfdbf8b
fix: move truncate_tool_messages import to module scope
claude Mar 21, 2026
67a17de
Merge branch 'claude/fix-eval-233-7keOj' of http://127.0.0.1:38447/gi…
claude Mar 21, 2026
03b59b8
feat: smart image handling in compaction — keep if they fit, strip if…
claude Mar 21, 2026
80b3b41
fix: resolve duplicate eval test numbers by renaming to 239-253
claude Mar 21, 2026
9bf8e5f
remove truncate_tool_messages, improve compaction, update docs and evals
claude Mar 21, 2026
579dc35
rename limit_input_context_window to compact_if_necessary, improve gr…
claude Mar 21, 2026
67f0e22
fix: GrafanaToolset._tools AttributeError and duplicate health_check
claude Mar 21, 2026
cdde68c
feat: add tool call images to Braintrust logs and update eval 238 prompt
claude Mar 21, 2026
5f14eee
fix: address PR review comments for image handling edge cases
claude Mar 21, 2026
8f9c6ed
fix: move save_images inside file_path guard in spill_oversized_tool_…
claude Mar 21, 2026
62c9bc8
Merge branch 'master' into claude/fix-eval-233-7keOj
aantn Mar 21, 2026
b3d254f
fix: address review comments - render height, spill loop, path restri…
claude Mar 21, 2026
8d93be9
Merge branch 'claude/fix-eval-233-7keOj' of http://127.0.0.1:42101/gi…
claude Mar 21, 2026
df0b666
fix: update ReadImageFile tests for storage path restriction
claude Mar 21, 2026
0348cc9
feat: add date/time and run numbers to previous eval runs
claude Mar 21, 2026
9e1526c
fix: log image metadata to Braintrust tool spans
claude Mar 21, 2026
818afdc
fix: address CodeRabbit review — namespace, MIME, and vision issues
claude Mar 21, 2026
63fd0aa
fix: default to vision=True for unrecognized models
claude Mar 21, 2026
dd53a94
simplify: always enable vision, add HOLMES_DISABLE_VISION env var
claude Mar 21, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 36 additions & 6 deletions .github/scripts/eval-comment-helpers.js
Original file line number Diff line number Diff line change
Expand Up @@ -31,8 +31,11 @@ function parseRunHistory(body) {
const historyRegex = /<details>\s*<summary>📜\s*(.+?)<\/summary>\s*([\s\S]*?)<!-- END_HISTORY_RUN -->\s*<\/details>/g;
let match;
while ((match = historyRegex.exec(body)) !== null) {
// Strip existing run number prefix (e.g., "#3 · ") to avoid double-numbering on re-render
let summary = match[1].trim();
summary = summary.replace(/^#\d+\s*·\s*/, '');
runs.push({
summary: match[1].trim(),
summary: summary,
content: match[2].trim()
});
}
Expand Down Expand Up @@ -62,7 +65,8 @@ function extractCurrentRun(body) {
}

// Find the header line (## ✅ Results... or ## ⏳ HolmesGPT evals running...)
const headerMatch = cleanBody.match(/^(## [^\n]+)/);
// Use multiline flag since content may start with hidden HTML comments (e.g., eval-timestamp)
const headerMatch = cleanBody.match(/^(## [^\n]+)/m);
if (!headerMatch) return null;

const header = headerMatch[1];
Expand Down Expand Up @@ -100,13 +104,29 @@ function extractCurrentRun(body) {
const runUrlMatch = currentSection.match(/\[View workflow logs\]\(([^)]+)\)/);
const runUrl = runUrlMatch ? runUrlMatch[1] : '';

// Extract embedded timestamp if available
const timestampMatch = currentSection.match(/<!-- eval-timestamp: (\S+) -->/);
const evalTimestamp = timestampMatch ? timestampMatch[1] : '';

// Format date for display (e.g., "Mar 21, 14:32 UTC")
let dateStr = '';
if (evalTimestamp) {
try {
const d = new Date(evalTimestamp);
const months = ['Jan','Feb','Mar','Apr','May','Jun','Jul','Aug','Sep','Oct','Nov','Dec'];
dateStr = `${months[d.getUTCMonth()]} ${d.getUTCDate()}, ${String(d.getUTCHours()).padStart(2,'0')}:${String(d.getUTCMinutes()).padStart(2,'0')} UTC`;
} catch (e) {
// ignore parse errors
}
}

// Build a descriptive summary with trigger info
let summary = 'Previous Run';
if (trigger) {
// Extract commit and branch info from trigger like "commit abc1234 on branch `feature`"
const commitMatch = trigger.match(/commit ([a-f0-9]+)/);
const commit = commitMatch ? commitMatch[1] : '';
summary = commit ? `Run @ ${commit}` : `Run: ${trigger.substring(0, 50)}`;
summary = commit ? `Run @ __${commit}__` : `Run: ${trigger.substring(0, 50)}`;
}
if (runUrl) {
// Extract run ID from URL for reference
Expand All @@ -115,6 +135,9 @@ function extractCurrentRun(body) {
summary += ` (#${runIdMatch[1]})`;
}
}
if (dateStr) {
summary += ` — ${dateStr}`;
}

// Build content for when this run becomes collapsed
const content = cleanBody.substring(0, endPos).trim();
Expand Down Expand Up @@ -149,9 +172,13 @@ function buildAutoCommentWithHistory(currentContent, previousRuns, footer, maxHi
if (runsToShow.length > 0) {
historySection = '## 📂 Previous Runs\n\n';

for (const run of runsToShow) {
for (let i = 0; i < runsToShow.length; i++) {
const run = runsToShow[i];
// Number runs: most recent = N, oldest = 1 (descending, newest first)
const runNumber = runsToShow.length - i;
const numberedSummary = `#${runNumber} · ${run.summary}`;
// Use END_HISTORY_RUN marker to properly delimit content (handles nested <details> tags in reports)
const historyEntry = `<details>\n<summary>📜 ${run.summary}</summary>\n\n${run.content}\n\n${HISTORY_RUN_END_MARKER}\n</details>\n\n`;
const historyEntry = `<details>\n<summary>📜 ${numberedSummary}</summary>\n\n${run.content}\n\n${HISTORY_RUN_END_MARKER}\n</details>\n\n`;

// Check if adding this entry would exceed the limit
const projectedSize = body.length + historySection.length + historyEntry.length + currentContent.length + footer.length;
Expand Down Expand Up @@ -248,7 +275,10 @@ function renderParamsTable(p, context = null) {
* @returns {string} Markdown body
*/
function buildBody(p, progressSteps, extras = {}) {
let body = p.isManual
// Embed timestamp as hidden comment for history display
const timestamp = new Date().toISOString();
let body = `<!-- eval-timestamp: ${timestamp} -->\n`;
body += p.isManual
? `## ${extras.icon || '🚀'} ${extras.title || 'Manual Eval Running...'}\n\n` +
renderParamsTable(p, extras.context)
: `## ${extras.icon || '⏳'} ${extras.title || 'HolmesGPT evals running...'}\n\n` +
Expand Down
3 changes: 2 additions & 1 deletion conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -128,7 +128,8 @@ def _patched_openai_init(self, *args, **kwargs):
with urllib.request.urlopen(tenant_url, timeout=10) as resp:
cloud_id = json.loads(resp.read())["cloudId"]
os.environ["CONFLUENCE_SA_BASE_URL"] = f"https://api.atlassian.com/ex/confluence/{cloud_id}"
logging.info(f"Auto-derived CONFLUENCE_SA_BASE_URL from cloud ID {cloud_id}")
os.environ["CONFLUENCE_CLOUD_ID"] = cloud_id
logging.info(f"Auto-derived CONFLUENCE_SA_BASE_URL and CONFLUENCE_CLOUD_ID from cloud ID {cloud_id}")
except Exception as e:
logging.warning(f"Could not auto-derive CONFLUENCE_SA_BASE_URL: {e}")

Expand Down
165 changes: 129 additions & 36 deletions docs/data-sources/builtin-toolsets/grafanadashboards.md
Original file line number Diff line number Diff line change
@@ -1,13 +1,15 @@
# Grafana Dashboards

Connect HolmesGPT to Grafana for dashboard analysis, query extraction, and understanding your monitoring setup. This integration enables investigation of dashboard configurations and extraction of Prometheus queries for deeper analysis.
Connect HolmesGPT to Grafana for dashboard analysis, visual rendering, query extraction, and understanding your monitoring setup. When the [Grafana Image Renderer](https://grafana.com/grafana/plugins/grafana-image-renderer/) is installed, HolmesGPT can visually render dashboards and panels to detect anomalies like spikes, trends, and outliers.

## Prerequisites

A [Grafana service account token](https://grafana.com/docs/grafana/latest/administration/service-accounts/) with the following permissions:

- Basic role → Viewer

For visual rendering, the [Grafana Image Renderer](https://grafana.com/grafana/plugins/grafana-image-renderer/) plugin must be installed on your Grafana instance and `enable_rendering: true` must be set in the config. HolmesGPT auto-detects the renderer — if it's not installed, visual rendering tools are simply not registered and everything else works normally.

## Configuration

=== "Holmes CLI"
Expand Down Expand Up @@ -49,68 +51,159 @@ A [Grafana service account token](https://grafana.com/docs/grafana/latest/admini
# X-Custom-Header: "custom-value"
```

## Capabilities
## Visual Rendering

When the Grafana Image Renderer is available, HolmesGPT can take screenshots of dashboards and panels and analyze them using the LLM's vision capabilities. This is useful for:

- Spotting anomalous spikes or patterns across many panels at once
- Analyzing visual dashboard layouts without parsing raw query data
- Investigating dashboards that use complex visualizations (heatmaps, gauges, etc.)

The LLM controls all rendering parameters — time range, dimensions, theme, timezone, and template variables — so it can zoom in on specific time windows or adjust the view as needed during investigation.

Rendering is **disabled by default**. To enable it, add `enable_rendering: true` to your config:

=== "Holmes CLI"

```yaml
toolsets:
grafana/dashboards:
enabled: true
config:
api_url: <your grafana url>
api_key: <your api key>
enable_rendering: true
```

=== "Holmes Helm Chart"

```yaml
toolsets:
grafana/dashboards:
enabled: true
config:
api_url: <your grafana url>
api_key: <your api key>
enable_rendering: true
```

=== "Robusta Helm Chart"

| Tool Name | Description |
|-----------|-------------|
| grafana_search_dashboards | Search for dashboards and folders by query, tags, UIDs, or folder locations |
| grafana_get_dashboard_by_uid | Retrieve complete dashboard JSON including all panels and queries |
| grafana_get_home_dashboard | Get the home dashboard configuration |
| grafana_get_dashboard_tags | List all tags used across dashboards for categorization |
```yaml
holmes:
toolsets:
grafana/dashboards:
enabled: true
config:
api_url: <your grafana url>
api_key: <your api key>
enable_rendering: true
```

When rendering a full dashboard, HolmesGPT captures the entire page (all rows) so that panels at the bottom are not cropped.

## Advanced Configuration

### SSL Verification

For self-signed certificates, you can disable SSL verification:

```yaml
toolsets:
grafana/dashboards:
enabled: true
config:
api_url: https://grafana.internal
api_key: <your api key>
verify_ssl: false # Disable SSL verification (default: true)
```
=== "Holmes CLI"

```yaml
toolsets:
grafana/dashboards:
enabled: true
config:
api_url: https://grafana.internal
api_key: <your api key>
verify_ssl: false # Disable SSL verification (default: true)
```

=== "Holmes Helm Chart"

```yaml
toolsets:
grafana/dashboards:
enabled: true
config:
api_url: https://grafana.internal
api_key: <your api key>
verify_ssl: false
```

=== "Robusta Helm Chart"

```yaml
holmes:
toolsets:
grafana/dashboards:
enabled: true
config:
api_url: https://grafana.internal
api_key: <your api key>
verify_ssl: false
```

### External URL

If HolmesGPT accesses Grafana through an internal URL but you want clickable links in results to use a different URL:

```yaml
toolsets:
grafana/dashboards:
enabled: true
config:
api_url: http://grafana.internal:3000 # Internal URL for API calls
external_url: https://grafana.example.com # URL for links in results
api_key: <your api key>
```
=== "Holmes CLI"

```yaml
toolsets:
grafana/dashboards:
enabled: true
config:
api_url: http://grafana.internal:3000 # Internal URL for API calls
external_url: https://grafana.example.com # URL for links in results
api_key: <your api key>
```

## How it Works
=== "Holmes Helm Chart"

### Dashboard Query Extraction
```yaml
toolsets:
grafana/dashboards:
enabled: true
config:
api_url: http://grafana.internal:3000
external_url: https://grafana.example.com
api_key: <your api key>
```

When HolmesGPT retrieves a dashboard, it can extract and analyze Prometheus queries from dashboard panels. This is particularly useful for:
=== "Robusta Helm Chart"

- Understanding what metrics a dashboard monitors
- Extracting queries for further investigation with the Prometheus toolset
- Analyzing dashboard time ranges and variable usage
```yaml
holmes:
toolsets:
grafana/dashboards:
enabled: true
config:
api_url: http://grafana.internal:3000
external_url: https://grafana.example.com
api_key: <your api key>
```

### Example Usage
## Common Use Cases

**Finding dashboards by tag:**
```bash
holmes ask "Find all dashboards tagged with 'production' or 'kubernetes'"
```

**Analyzing a specific dashboard:**
```bash
holmes ask "Show me what metrics the 'Node Exporter' dashboard monitors"
```

**Extracting queries for investigation:**
```bash
holmes ask "Get the CPU usage queries from the Kubernetes cluster dashboard and check if any nodes are throttling"
```

```bash
holmes ask "Look at the Platform Services dashboard and tell me if any panels show anomalous spikes"
```

```bash
holmes ask "Render the checkout latency panel from the last 24 hours and analyze the trend"
```
1 change: 1 addition & 0 deletions docs/reference/.nav.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,4 +7,5 @@ nav:
- HTTP API: http-api.md
- Runbooks: runbooks.md
- Slash Commands: slash-commands.md
- Context Management: context-management.md
- Troubleshooting: troubleshooting.md
69 changes: 69 additions & 0 deletions docs/reference/context-management.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
# Context Management

HolmesGPT uses two mechanisms to keep conversations within the LLM's context window.
They run at different points in the pipeline and serve different purposes.

## 1. Single Tool Result Spill-to-Disk

**Function:** `spill_oversized_tool_result()` in `holmes/core/tools_utils/tool_context_window_limiter.py`

**When:** Immediately after each tool call returns, before the result is added to conversation history.

**What it does:**

- Counts the tokens in the single tool result.
- If it exceeds `max_token_count_for_single_tool` (configured via `TOOL_MAX_ALLOCATED_CONTEXT_WINDOW_PCT`):
- Saves the full text result to a file on disk.
- If the result contains images, saves them as separate files on disk.
- Replaces the in-conversation result with a pointer message containing the file path, a preview, and instructions for the LLM to `cat` the file or use `read_image_file` to load images back.
- If disk storage is unavailable, drops the data entirely and returns an error asking the LLM to narrow its query.

**Called from:** `tool_calling_llm.py` → `_invoke_llm_tool_call()`, after tool execution.

**Scope:** One tool result at a time. Does not look at the overall conversation size.

## 2. Conversation History Compaction

**Function:** `compact_conversation_history()` in `holmes/core/truncation/compaction.py`, orchestrated by `compact_if_necessary()` in `holmes/core/truncation/input_context_window_limiter.py`

**When:** Before each LLM call in the agentic loop, if the total conversation tokens exceed a compaction threshold.

**What it does:**

- Checks if `(total_tokens + max_output_tokens) > (context_window_size * threshold_pct / 100)`.
- If so, sends the conversation history to the LLM with a compaction prompt, asking it to produce a concise summary.
- Replaces the old messages with: system prompt + compacted summary + last user message.
- Tracks compaction cost in `RequestStats`.

**Guard:** Controlled by `ENABLE_CONVERSATION_HISTORY_COMPACTION` env var (defaults to true).

**Called from:** `tool_calling_llm.py` → `call_stream()`, at the top of each agentic loop iteration.

**Scope:** The entire conversation history. Uses an LLM call (costs tokens/money).

## How They Interact

```
Tool executes
│
▼
┌─────────────────────────────┐
│ 1. spill_oversized_tool_result │ ← caps single tool result
└─────────────────────────────┘
│
▼
Tool result added to conversation
│
▼
┌─────────────────────────────┐
│ 2. compaction (if needed) │ ← summarizes full conversation via LLM
└─────────────────────────────┘
│
▼
LLM called with messages
```

In practice:

- Mechanism 1 prevents any single tool from blowing up the context.
- Mechanism 2 prevents the cumulative conversation from growing unbounded.
Loading
Loading