Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions docs/ai-providers/aws-bedrock.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,9 @@

Configure HolmesGPT to use AWS Bedrock foundation models.

!!! tip "Which Model to Use"
We highly recommend using Sonnet 4.0 or Sonnet 4.5 as they give the best results by far. See examples below for configuration.

## Setup

### Prerequisites
Expand Down
54 changes: 19 additions & 35 deletions docs/ai-providers/ollama.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
Configure HolmesGPT to use local models with Ollama.

!!! warning
Ollama support is experimental. Tool-calling capabilities are limited and may produce inconsistent results. Only [LiteLLM supported Ollama models](https://docs.litellm.ai/docs/providers/ollama#ollama-models){:target="_blank"} work with HolmesGPT.
Ollama support is experimental and can be tricky to configure correctly. We recommend trying HolmesGPT with a hosted model first (like Claude or OpenAI) to ensure everything works before switching to Ollama. Tool-calling capabilities are limited and may produce inconsistent results. Only [LiteLLM supported Ollama models](https://docs.litellm.ai/docs/providers/ollama#ollama-models){:target="_blank"} work with HolmesGPT.
Comment thread
aantn marked this conversation as resolved.

## Setup

Expand All @@ -18,6 +18,24 @@ Configure HolmesGPT to use local models with Ollama.
```bash
export OLLAMA_API_BASE="http://localhost:11434"
holmes ask "what pods are failing?" --model="ollama_chat/<your-ollama-model>"

# Or use MODEL environment variable instead of --model flag
export MODEL="ollama_chat/<your-ollama-model>"
holmes ask "what pods are failing?"
```

**Alternative (OpenAI-compatible gateway)**

If you hit compatibility issues with certain Ollama models via LiteLLM, you can use Ollama's OpenAI-compatible API endpoint:

```bash
export OPENAI_API_BASE="http://localhost:11434/v1"
export OPENAI_API_KEY="dummy-key" # Required but can be any value
holmes ask "what pods are failing?" --model="openai/<your-ollama-model>"

# Or use MODEL environment variable instead of --model flag
export MODEL="openai/<your-ollama-model>"
holmes ask "what pods are failing?"
```

=== "Holmes Helm Chart"
Expand Down Expand Up @@ -126,40 +144,6 @@ Configure HolmesGPT to use local models with Ollama.
model: "ollama-alt"
```

### Using Environment Variables

```bash
export OLLAMA_API_BASE="http://localhost:11434"
export MODEL="ollama_chat/<your-ollama-model>"
holmes ask "what pods are failing?"
```

Alternative via an OpenAI-compatible gateway:

```bash
export OPENAI_API_BASE="http://localhost:11434/v1"
export OPENAI_API_KEY="YOUR_BEARER_TOKEN_HERE"
export MODEL="openai/OLLAMA_MODEL_NAME"
holmes ask "what pods are failing?"
```

### Using CLI Parameters

You can also specify the model directly as a command-line parameter:

```bash
export OLLAMA_API_BASE="http://localhost:11434"
holmes ask "what pods are failing?" --model="ollama_chat/<your-ollama-model>"
```

Using the OpenAI-compatible alternative:

```bash
export OPENAI_API_BASE="http://ollama-service:11434/v1"
export OPENAI_API_KEY="YOUR_BEARER_TOKEN_HERE"
holmes ask "what pods are failing?" --model="openai/OLLAMA_MODEL_NAME"
```

## Additional Resources

HolmesGPT uses the LiteLLM API to support Ollama provider. Refer to [LiteLLM Ollama docs](https://docs.litellm.ai/docs/providers/ollama){:target="_blank"} for more details.
85 changes: 46 additions & 39 deletions docs/development/evaluations/adding-evals.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,39 @@

Create test cases that measure HolmesGPT's diagnostic accuracy and help track improvements over time.

## Test Types
## Prerequisites

- **Ask Holmes**: Chat-like Q&A interactions
- **Investigation**: AlertManager event analysis
Install HolmesGPT python dependencies:

```bash
poetry install --with=dev
```

## Quick Start: Running Your First Eval

Try running an existing eval to understand how the system works. We'll use [eval 80_pvc_storage_class_mismatch](https://github.com/robusta-dev/holmesgpt/tree/master/tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch) as an example:

```bash
# Run eval #80 with Claude Sonnet 4.5 (this specific eval passes reliably with Sonnet 4.5)
RUN_LIVE=true MODEL=anthropic/claude-sonnet-4-20250514 \
CLASSIFIER_MODEL=gpt-4.1 \
poetry run pytest tests/llm/test_ask_holmes.py -k "80_pvc_storage_class_mismatch"

# Compare with GPT-4o (may not pass as reliably)
RUN_LIVE=true MODEL=gpt-4o \
poetry run pytest tests/llm/test_ask_holmes.py -k "80_pvc_storage_class_mismatch"

# Compare with GPT-4.1 (may not pass as reliably)
RUN_LIVE=true MODEL=gpt-4.1 \
poetry run pytest tests/llm/test_ask_holmes.py -k "80_pvc_storage_class_mismatch"

# Test multiple models at once to compare performance
RUN_LIVE=true MODEL=gpt-4o,gpt-4.1,anthropic/claude-sonnet-4-20250514 \
CLASSIFIER_MODEL=gpt-4.1 \
poetry run pytest tests/llm/test_ask_holmes.py -k "80_pvc_storage_class_mismatch"
```

**Note:** Eval #80 demonstrates how different models perform differently - Sonnet 4.5 passes this specific eval reliably while weaker models like GPT-4o and GPT-4.1 may struggle with this scenario.

## Quick Start

Expand Down Expand Up @@ -37,10 +66,21 @@ spec:

4. Run test:
```bash
poetry run pytest tests/llm/test_ask_holmes.py -k "99_your_test" -v
# With GPT-4.1
RUN_LIVE=true MODEL=gpt-4.1 \
poetry run pytest tests/llm/test_ask_holmes.py -k "99_your_test" -v

# With Claude Sonnet 4.5 (must set CLASSIFIER_MODEL since Anthropic models can't be used as classifiers)
RUN_LIVE=true MODEL=anthropic/claude-sonnet-4-20250514 \
CLASSIFIER_MODEL=gpt-4.1 \
poetry run pytest tests/llm/test_ask_holmes.py -k "99_your_test" -v
```

## Test Configuration
**Note on CLASSIFIER_MODEL:** An LLM judges whether tests pass. Only OpenAI models (like `gpt-4.1`) work as classifiers. Set `CLASSIFIER_MODEL=gpt-4.1` explicitly when using Anthropic models. For OpenAI models, it defaults to `MODEL`.

## test_case.yaml Configuration

Configure your test by defining these fields in `test_case.yaml`:

### Required Fields
- `user_prompt`: Question for Holmes
Expand Down Expand Up @@ -80,26 +120,6 @@ poetry run pytest tests/llm/test_ask_holmes.py -k "99_your_test" -v

**Live evaluations (`RUN_LIVE=true`) are strongly preferred** because they're more reliable and accurate.

### Why Live Evaluations Are Preferred

**LLMs can take multiple paths to reach the same conclusion.** When using mock data:

- The LLM might call tools in a different order than when mocks were generated
- It might use different tool combinations to diagnose the same issue
- It might ask for additional information not captured in the mocks
- Mock data represents only one possible investigation path

With live evaluations, the LLM can explore any path it chooses, making tests more robust and realistic.

### When Mock Data Is Necessary

Mock data is sometimes unavoidable:
- CI/CD environments without Kubernetes cluster access
- Testing specific edge cases that require controlled responses
- Reproducing exact historical scenarios

**Important**: Even when using mocks, always validate with `RUN_LIVE=true` in a real environment.

### Generating Mock Data

```bash
Expand All @@ -115,6 +135,7 @@ Mock files are named: `{tool_name}_{context}.txt`
### Mock Data Guidelines

When creating mock data:

- Never generate mock data manually - always use `--generate-mocks` with live execution
- Mock data should match real-world responses exactly
- Include all fields that would be present in actual responses
Expand Down Expand Up @@ -247,17 +268,3 @@ Some examples
- `synthetic` - Tests that use manually generated mock data (cannot be run live)
- `datetime` - Tests date/time handling and interpretation
- etc.

### Using Tags in Test Cases

Add tags to your `test_case.yaml`:

```yaml
user_prompt: "Show me the logs for the pod `robusta-holmes` since last Thursday"
tags:
- logs
- datetime
expected_output:
- Database unavailable
- Memory pressure
```
143 changes: 55 additions & 88 deletions docs/development/evaluations/running-evals.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,9 +25,35 @@ Install HolmesGPT python dependencies:
poetry install --with=dev
```

### Quick Start: Using run_benchmarks_local.sh
### Quick Start: Running Your First Eval

The easiest way to run benchmarks locally is using the `run_benchmarks_local.sh` script, which mirrors the exact behavior of our CI/CD workflow:
Try running a single eval to understand how the system works. We'll use [eval 80_pvc_storage_class_mismatch](https://github.com/robusta-dev/holmesgpt/tree/master/tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch) as an example:

```bash
# Run eval #80 with Claude Sonnet 4.5 (this specific eval passes reliably with Sonnet 4.5)
RUN_LIVE=true MODEL=anthropic/claude-sonnet-4-20250514 \
CLASSIFIER_MODEL=gpt-4.1 \
poetry run pytest tests/llm/test_ask_holmes.py -k "80_pvc_storage_class_mismatch"

# Compare with GPT-4o (may not pass as reliably)
RUN_LIVE=true MODEL=gpt-4o \
poetry run pytest tests/llm/test_ask_holmes.py -k "80_pvc_storage_class_mismatch"

# Compare with GPT-4.1 (may not pass as reliably)
RUN_LIVE=true MODEL=gpt-4.1 \
poetry run pytest tests/llm/test_ask_holmes.py -k "80_pvc_storage_class_mismatch"

# Test multiple models at once to compare performance
RUN_LIVE=true MODEL=gpt-4o,gpt-4.1,anthropic/claude-sonnet-4-20250514 \
CLASSIFIER_MODEL=gpt-4.1 \
poetry run pytest tests/llm/test_ask_holmes.py -k "80_pvc_storage_class_mismatch"
```

**Note:** This eval demonstrates how different models perform differently - Sonnet 4.5 passes this specific eval reliably while weaker models like GPT-4o and GPT-4.1 may struggle with this scenario.

### Running Full Benchmark Suite

Once you're comfortable running individual evals, you can run the full benchmark suite to test all important evals at once. The easiest way to do this locally is using the `run_benchmarks_local.sh` script, which mirrors the exact behavior of our CI/CD workflow:

```bash
# Run with defaults (easy tests, default models, 1 iteration)
Expand All @@ -46,28 +72,43 @@ The easiest way to run benchmarks locally is using the `run_benchmarks_local.sh`
./run_benchmarks_local.sh 'gpt-4o' 'easy' 1 '' 6
```

The script automatically:
- Sets up all required environment variables
- Runs tests with JSON report generation
- Generates a markdown benchmark report with dashboard, costs, and timings
- Saves a historical copy in `docs/development/evaluations/history/`
- Shows how to commit results
## Environment Variables

### Manual Commands
Essential variables for controlling test behavior:

For more control, you can run pytest directly:
| Variable | Purpose | Example |
|----------|---------|---------|
| `RUN_LIVE` | Use real tools instead of mocks | `RUN_LIVE=true` |
| `ITERATIONS` | Run each test N times | `ITERATIONS=10` |
| `MODEL` | LLM to test | `MODEL=gpt-4.1` |
| `CLASSIFIER_MODEL` | LLM for scoring (needed for Anthropic) | `CLASSIFIER_MODEL=gpt-4.1` |

## Advanced Usage

### Selecting Which Evals to Run

For more control over which evals to run, you can use pytest directly with markers (tags) or test name patterns:

```bash
# Run all easy evals - these should always pass assuming you have a kubernetes cluster with sufficient resources and a 'good enough model' (e.g. gpt-4.1, claude-opus-4-1)
# Run all easy evals (regression tests - should always pass)
RUN_LIVE=true poetry run pytest -m 'llm and easy' --no-cov

# Run a specific eval
RUN_LIVE=true poetry run pytest tests/llm/test_ask_holmes.py -k "01_how_many_pods"
# Run challenging tests
RUN_LIVE=true poetry run pytest -m 'llm and medium' --no-cov

# Run evals with a specific tag
# Run evals with a specific tag (e.g., tests involving logs)
RUN_LIVE=true poetry run pytest -m "llm and logs" --no-cov

# Run a specific eval by name
RUN_LIVE=true poetry run pytest tests/llm/test_ask_holmes.py -k "01_how_many_pods"
```

**Available markers:** See `pyproject.toml` for all available markers. Common ones include:
- `easy` - Regression tests that should always pass
- `medium` - More challenging scenarios
- `logs` - Tests involving log analysis
- `kubernetes` - Kubernetes-specific tests

### Testing Different Models

The `MODEL` environment variable is equivalent to the `--model` flag on the `holmes ask` CLI command. You can test HolmesGPT with different LLM providers:
Expand Down Expand Up @@ -139,42 +180,6 @@ Some evals support mock-data and don't need a live Kubernetes cluster to run. Ho

This is important because LLMs can take multiple paths to reach conclusions, and mock data only captures one path.

## Environment Variables

Essential variables for controlling test behavior:

| Variable | Purpose | Example |
|----------|---------|---------|
| `RUN_LIVE` | Use real tools instead of mocks | `RUN_LIVE=true` |
| `ITERATIONS` | Run each test N times | `ITERATIONS=10` |
| `MODEL` | LLM to test | `MODEL=gpt-4.1` |
| `CLASSIFIER_MODEL` | LLM for scoring (needed for Anthropic) | `CLASSIFIER_MODEL=gpt-4.1` |
| `ASK_HOLMES_TEST_TYPE` | Message building flow (`cli` or `server`) | `ASK_HOLMES_TEST_TYPE=server` |

### ASK_HOLMES_TEST_TYPE Details

The `ASK_HOLMES_TEST_TYPE` environment variable controls how messages are built in ask_holmes tests:

- **`cli` (default)**: Uses `build_initial_ask_messages` like the CLI ask() command. This mode:
- Simulates the CLI interface behavior
- Does not support conversation history tests (will skip them)
- Includes runbook loading and system prompts as done in the CLI

- **`server`**: Uses `build_chat_messages` with ChatRequest for server-style flow. This mode:
- Simulates the API/server interface behavior
- Supports conversation history tests
- Uses the ChatRequest model for message building

```bash
# Test with CLI-style message building (default)
RUN_LIVE=true poetry run pytest -k "test_name"

# Test with server-style message building
RUN_LIVE=true ASK_HOLMES_TEST_TYPE=server poetry run pytest -k "test_name"
```

## Advanced Usage

### Parallel Execution

Speed up test runs with parallel workers:
Expand Down Expand Up @@ -262,41 +267,3 @@ export BRAINTRUST_ORG=your-org
# Then run any evaluation command - results will be tracked automatically
RUN_LIVE=true MODEL=gpt-4o,anthropic/claude-sonnet-4-20250514 pytest -m 'llm and easy'
```

### CI/CD Benchmarking

HolmesGPT includes a [GitHub Actions workflow](https://github.com/robusta-dev/holmesgpt/blob/main/.github/workflows/eval-benchmarks.yml) for automated benchmarking that runs:

**Automatically:**
Coming soon: automation to run weekly.

**Manually:**
- Navigate to Actions tab → "LLM Evaluation Benchmarks" → Run workflow
- Customize models, markers, and iterations

**Workflow Parameters:**
- `models`: Comma-separated list (e.g., `gpt-4o,anthropic/claude-sonnet-4-20250514`)
- `test_markers`: Pytest markers (e.g., `llm and easy`, `llm and medium`, `llm`)
- `iterations`: Number of test iterations

Benchmark results are automatically:
- Published to `docs/development/evaluations/latest-results.md`
- Archived with timestamps in `docs/development/evaluations/history/`
- Posted as PR comments when tests run on pull requests

## Test Markers

Filter tests by functionality:

```bash
# Regression tests (should always pass)
RUN_LIVE=true ITERATIONS=10 poetry run pytest -m "llm and easy"

# Challenging tests
RUN_LIVE=true ITERATIONS=10 poetry run pytest -m "llm and medium"

# Tests involving logs
RUN_LIVE=true poetry run pytest -m "llm and logs"
```

See `pyproject.toml` for all available markers.
2 changes: 1 addition & 1 deletion docs/installation/.nav.yml
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
nav:
- Install CLI: cli-installation.md
- Install UI/TUI: ui-installation.md
- Install UI/Slack/K9s: ui-installation.md
- Install Helm Chart: kubernetes-installation.md
- Install Python SDK: python-installation.md
Loading
Loading