Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
1252b29
Add comprehensive documentation improvement analysis
claude Jan 7, 2026
253d6e6
Improve documentation clarity and conciseness
claude Jan 9, 2026
bdc1448
Fix documentation organization and MkDocs formatting
claude Jan 20, 2026
06a7b5f
Merge branch 'master' into claude/improve-docs-clarity-fu94h
aantn Jan 21, 2026
9be4850
Merge branch 'master' into claude/improve-docs-clarity-fu94h
aantn Jan 21, 2026
84fd585
Fix Anthropic docs: move CLI parameter info inside Holmes CLI tab
claude Jan 21, 2026
dbbf09c
Remove redundant Getting Started section and add Prometheus URL link
claude Jan 21, 2026
5af9f57
Reorganize Prometheus docs: validation, remove fluff, group providers
claude Jan 21, 2026
3324e9c
Remove fluff from Prometheus intro sentence
claude Jan 21, 2026
7f004f3
Move validation into Configuration section on Prometheus page
claude Jan 21, 2026
3cd121a
Simplify Prometheus docs: move test command inline, remove Helm valid…
claude Jan 21, 2026
5d21956
Add CLI/Helm tabs to Prometheus configuration section
claude Jan 21, 2026
53fb18f
Change Prometheus config to 3 tabs: Holmes CLI, Holmes Helm, Robusta …
claude Jan 21, 2026
adac44f
Use yaml-toolset-config fence for Prometheus CLI tab
claude Jan 21, 2026
6069e57
Revert to regular yaml fence in Prometheus CLI tab
claude Jan 21, 2026
baec590
Simplify Prometheus docs: use yaml-toolset-config, remove validation
claude Jan 21, 2026
fae01e7
Merge branch 'master' into claude/improve-docs-clarity-fu94h
aantn Jan 21, 2026
a61eb7e
Move Advanced Configuration and Capabilities after Specific Providers
claude Jan 21, 2026
da03888
Fix Azure Managed Prometheus section hierarchy
claude Jan 21, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 6 additions & 10 deletions docs/ai-providers/anthropic.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,12 @@ Get an [Anthropic API key](https://support.anthropic.com/en/articles/8114521-how

**Note**: You can use any Anthropic model by changing the model name. See [Claude Models Overview](https://docs.claude.com/en/docs/about-claude/models/overview#latest-models-comparison){:target="_blank"} for available model names.

You can also pass the API key directly as a command-line parameter:

```bash
holmes ask "what pods are failing?" --model="anthropic/claude-sonnet-4-5" --api-key="your-api-key"
```

=== "Holmes Helm Chart"

**Create Kubernetes Secret:**
Expand Down Expand Up @@ -96,16 +102,6 @@ Get an [Anthropic API key](https://support.anthropic.com/en/articles/8114521-how
model: "claude-sonnet-4" # This refers to the key name in modelList above
```

## Using CLI Parameters

You can also pass the API key directly as a command-line parameter:

```bash
holmes ask "what pods are failing?" --model="anthropic/claude-sonnet-4-5" --api-key="your-api-key"
```

**Note**: You can use any Anthropic model by changing the model name. See [Claude Models Overview](https://docs.claude.com/en/docs/about-claude/models/overview#latest-models-comparison){:target="_blank"} for available model names.

## Prompt Caching

HolmesGPT adds Anthropic's prompt caching feature, which can significantly reduce costs and latency for repeated API calls with similar prompts.
Expand Down
2 changes: 2 additions & 0 deletions docs/data-sources/builtin-toolsets/gcp.md
Original file line number Diff line number Diff line change
Expand Up @@ -368,6 +368,7 @@ The GCP MCP servers require a service account with appropriate read-only permiss
The setup script grants ~50 optimized read-only roles designed for incident response and troubleshooting:

**What's Included:**

- ✅ Complete audit log visibility (who changed what)
- ✅ Full networking troubleshooting (firewalls, load balancers, SSL)
- ✅ Database and BigQuery metadata (schemas, configurations)
Expand All @@ -376,6 +377,7 @@ The setup script grants ~50 optimized read-only roles designed for incident resp
- ✅ Monitoring, logging, and tracing

**Security Boundaries:**

- ❌ NO actual data access (cannot read storage objects or BigQuery data)
- ❌ NO secret values (only metadata)
- ❌ NO write permissions
Expand Down
1 change: 1 addition & 0 deletions docs/data-sources/builtin-toolsets/github.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,6 +94,7 @@ This tool modifies files line-by-line and can create or update pull requests. It
2. Then run with `dry_run=false` to commit changes

**Key Parameters:**

- `line`: Line number where the change occurs
- `command`: Operation type (`insert`, `update`, or `remove`)
- `commit_pr`: Either a PR title (for new PRs) or PR number like `#123` (for updating existing PRs)
Expand Down
1 change: 1 addition & 0 deletions docs/data-sources/builtin-toolsets/grafanaloki.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,7 @@ Choose one of the following methods:
### Option 1: Through Grafana (Recommended)

**Required:**

- [Grafana service account token](https://grafana.com/docs/grafana/latest/administration/service-accounts/) with Viewer role
- Loki datasource UID from Grafana

Expand Down
8 changes: 0 additions & 8 deletions docs/data-sources/builtin-toolsets/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,11 +43,3 @@ HolmesGPT includes pre-built integrations for popular monitoring and observabili
- [:simple-grafana:{ .lg .middle } **Tempo**](grafanatempo.md)

</div>

## Getting Started

1. **Choose toolsets** that match your infrastructure (Prometheus, Grafana, etc.)
2. **Configure authentication** - some need API keys, others work automatically
3. **Run a test investigation** to verify data access

💡 **Tip**: Start with [Kubernetes](kubernetes.md) and [Prometheus](prometheus.md) for basic cluster monitoring.
160 changes: 70 additions & 90 deletions docs/data-sources/builtin-toolsets/prometheus.md
Original file line number Diff line number Diff line change
@@ -1,11 +1,11 @@
# Prometheus

Connect HolmesGPT to Prometheus for metrics analysis and query generation. This integration enables detection of memory leaks, CPU throttling, queue backlogs, and performance issues.
Connect HolmesGPT to Prometheus for metrics analysis and query generation.

## Prerequisites

- A running and accessible Prometheus server
- Ensure HolmesGPT can connect to the Prometheus endpoint
- Ensure HolmesGPT can connect to the Prometheus endpoint (see [Finding your Prometheus URL](#finding-your-prometheus-url))

## Configuration

Expand All @@ -21,23 +21,6 @@ toolsets:
# Authorization: "Basic <base_64_encoded_string>"
```


💡 **Alternative**: Set environment variables instead of using the config file:
- `PROMETHEUS_URL`: The Prometheus server URL
- `PROMETHEUS_AUTH_HEADER`: Optional authorization header value (e.g., "Bearer token123")

## Validation

To test your connection, run:

```bash
holmes ask "Show me the CPU usage for the last hour"
```

## Troubleshooting



### Finding your Prometheus URL

There are several ways to find your Prometheus URL:
Expand All @@ -63,73 +46,9 @@ kubectl get svc --all-namespaces -o jsonpath='{range .items[*]}{.metadata.name}{

This will print all possible Prometheus service URLs in your cluster. Pick the one that matches your deployment.

### Common Issues

- **Connection refused**: Check if the Prometheus URL is accessible from HolmesGPT.
- **Authentication errors**: Verify the headers configuration for secured Prometheus endpoints.
- **No metrics returned**: Ensure that Prometheus is scraping your targets.


## Advanced Configuration
## Specific Providers

You can further customize the Prometheus toolset with the following options:

```yaml
toolsets:
prometheus/metrics:
enabled: true
config:
prometheus_url: http://prometheus-server.monitoring.svc.cluster.local:9090
headers:
Authorization: "Basic <base64_encoded_credentials>"

# Discovery settings
discover_metrics_from_last_hours: 1 # Only return metrics with data in last N hours (default: 1)

# Timeout configuration
query_timeout_seconds_default: 20 # Default timeout for PromQL queries (default: 20)
query_timeout_seconds_hard_max: 180 # Maximum allowed timeout for PromQL queries (default: 180)
metadata_timeout_seconds_default: 20 # Default timeout for metadata/discovery APIs (default: 20)
metadata_timeout_seconds_hard_max: 60 # Maximum allowed timeout for metadata APIs (default: 60)

# Other options
rules_cache_duration_seconds: 1800 # Cache duration for Prometheus rules (default: 30 minutes)
verify_ssl: true # Enable SSL verification (default: true)
tool_calls_return_data: true # If false, disables returning Prometheus data (default: true)
additional_labels: # Additional labels to add to all queries
cluster: "production"
```

**Config option explanations:**

- `prometheus_url`: The base URL for Prometheus. Should include protocol and port.
- `headers`: Extra headers for all Prometheus HTTP requests (e.g., for authentication).
- `discover_metrics_from_last_hours`: Only return metrics that have data in the last N hours when using discovery APIs (get_metric_names, get_label_values, etc.). Default: 1 hour. Increase if you have metrics that report infrequently.
- `query_timeout_seconds_default`: Default timeout for PromQL queries. Can be overridden per query. Default: 20.
- `query_timeout_seconds_hard_max`: Maximum allowed timeout for PromQL queries. Default: 180.
- `metadata_timeout_seconds_default`: Default timeout for metadata/discovery API calls. Default: 20.
- `metadata_timeout_seconds_hard_max`: Maximum allowed timeout for metadata API calls. Default: 60.
- `rules_cache_duration_seconds`: How long to cache Prometheus rules. Set to `null` to disable caching. Default: 1800 (30 minutes).
- `verify_ssl`: Enable SSL certificate verification. Default: true.
- `tool_calls_return_data`: If `false`, disables returning Prometheus data to HolmesGPT (useful if you hit token limits). Default: true.
- `additional_labels`: Dictionary of labels to add to all queries (currently only implemented for AWS/AMP).

## Capabilities

| Tool Name | Description |
|-----------|-------------|
| list_prometheus_rules | List all defined Prometheus rules with descriptions and annotations |
| get_metric_names | Get list of metric names (fastest discovery method) - requires match filter |
| get_label_values | Get all values for a specific label (e.g., pod names, namespaces) |
| get_all_labels | Get list of all label names available in Prometheus |
| get_series | Get time series matching a selector (returns full label sets) |
| get_metric_metadata | Get metadata (type, description, unit) for metrics |
| execute_prometheus_instant_query | Execute an instant PromQL query (single point in time) |
| execute_prometheus_range_query | Execute a range PromQL query for time series data with graph generation |

---

## Coralogix Prometheus Configuration
### Coralogix Prometheus

To use a Coralogix PromQL endpoint with HolmesGPT:

Expand Down Expand Up @@ -164,7 +83,7 @@ To use a Coralogix PromQL endpoint with HolmesGPT:

---

## AWS Managed Prometheus (AMP) Configuration
### AWS Managed Prometheus (AMP)

To connect HolmesGPT to AWS Managed Prometheus:

Expand Down Expand Up @@ -193,7 +112,7 @@ holmes:

---

## Google Managed Prometheus Configuration
### Google Managed Prometheus

Before configuring Holmes, make sure you have:

Expand All @@ -220,14 +139,14 @@ holmes:
* No additional headers or credentials are required
* The Prometheus Frontend endpoint must be accessible from the cluster

## Azure Managed Prometheus Configuration
### Azure Managed Prometheus

Before configuring Holmes, make sure you have:

* An Azure Monitor workspace with Managed Prometheus enabled
* A service principal (or managed identity) that has access to the workspace

### Using a service principal (client secret)
#### Using a service principal (client secret)

```yaml
holmes:
Expand All @@ -250,7 +169,7 @@ holmes:
- No extra headers are required; authentication is handled through Azure AD (service principal or managed identity).
- SSL is enabled by default (`verify_ssl: true`). Disable only if you know you need to trust a custom cert.

## Grafana Cloud (Mimir) Configuration
### Grafana Cloud (Mimir)

To connect HolmesGPT to Grafana Cloud's Prometheus/Mimir endpoint:

Expand Down Expand Up @@ -282,3 +201,64 @@ To connect HolmesGPT to Grafana Cloud's Prometheus/Mimir endpoint:

- Use the proxy endpoint URL format `/api/datasources/proxy/uid/` - this handles authentication and routing to Mimir automatically
- The toolset automatically detects and uses the most appropriate APIs for discovery

---

## Advanced Configuration

You can further customize the Prometheus toolset with the following options:

```yaml
toolsets:
prometheus/metrics:
enabled: true
config:
prometheus_url: http://prometheus-server.monitoring.svc.cluster.local:9090
headers:
Authorization: "Basic <base64_encoded_credentials>"

# Discovery settings
discover_metrics_from_last_hours: 1 # Only return metrics with data in last N hours (default: 1)

# Timeout configuration
query_timeout_seconds_default: 20 # Default timeout for PromQL queries (default: 20)
query_timeout_seconds_hard_max: 180 # Maximum allowed timeout for PromQL queries (default: 180)
metadata_timeout_seconds_default: 20 # Default timeout for metadata/discovery APIs (default: 20)
metadata_timeout_seconds_hard_max: 60 # Maximum allowed timeout for metadata APIs (default: 60)

# Other options
rules_cache_duration_seconds: 1800 # Cache duration for Prometheus rules (default: 30 minutes)
verify_ssl: true # Enable SSL verification (default: true)
tool_calls_return_data: true # If false, disables returning Prometheus data (default: true)
additional_labels: # Additional labels to add to all queries
cluster: "production"
```

**Configuration options:**

| Option | Default | Description |
|--------|---------|-------------|
| `prometheus_url` | - | Prometheus server URL (include protocol and port) |
| `headers` | `{}` | Authentication headers (e.g., `Authorization: Bearer token`) |
| `discover_metrics_from_last_hours` | `1` | Only discover metrics with data in last N hours |
| `query_timeout_seconds_default` | `20` | Default PromQL query timeout |
| `query_timeout_seconds_hard_max` | `180` | Maximum query timeout |
| `metadata_timeout_seconds_default` | `20` | Default metadata/discovery API timeout |
| `metadata_timeout_seconds_hard_max` | `60` | Maximum metadata API timeout |
| `rules_cache_duration_seconds` | `1800` | Cache duration for rules (set to `null` to disable) |
| `verify_ssl` | `true` | Enable SSL certificate verification |
| `tool_calls_return_data` | `true` | Return Prometheus data (disable if hitting token limits) |
| `additional_labels` | `{}` | Labels to add to all queries (AWS/AMP only) |

## Capabilities

| Tool Name | Description |
|-----------|-------------|
| list_prometheus_rules | List all defined Prometheus rules with descriptions and annotations |
| get_metric_names | Get list of metric names (fastest discovery method) - requires match filter |
| get_label_values | Get all values for a specific label (e.g., pod names, namespaces) |
| get_all_labels | Get list of all label names available in Prometheus |
| get_series | Get time series matching a selector (returns full label sets) |
| get_metric_metadata | Get metadata (type, description, unit) for metrics |
| execute_prometheus_instant_query | Execute an instant PromQL query (single point in time) |
| execute_prometheus_range_query | Execute a range PromQL query for time series data with graph generation |
31 changes: 0 additions & 31 deletions docs/reference/environment-variables.md
Original file line number Diff line number Diff line change
Expand Up @@ -159,37 +159,6 @@ Controls message building flow in ask_holmes tests:
- `cli` (default) - Uses CLI-style message building
- `server` - Uses server-style message building with ChatRequest

## Usage Examples

### Basic Setup
```bash
# Set up OpenAI
export OPENAI_API_KEY="sk-..."
export HOLMES_LOG_LEVEL="INFO"

# Run HolmesGPT
holmes ask "what pods are failing?"
```

### Gemini Configuration
```bash
# Configure for Gemini models
export GEMINI_API_KEY="your-key"
export TOOL_SCHEMA_NO_PARAM_OBJECT_IF_NO_PARAMS=true

holmes ask "analyze cluster health" --model="gemini/gemini-1.5-pro"
```

### Testing with Strict Tool Calling
```bash
# Enable strict tool calling for additional models
export LLMS_WITH_STRICT_TOOL_CALLS="azure/gpt-4.1,openai/*,custom/model-*"
export RUN_LIVE=true
export MODEL="custom/model-v2"

poetry run pytest tests/llm/ -n 6
```

## See Also

- [AI Providers](../ai-providers/index.md) - Detailed configuration for each AI provider
Expand Down
Loading