Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions .github/workflows/eval-benchmarks.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -115,11 +115,19 @@ jobs:
- name: Generate benchmark report
if: always()
run: |
# Generate latest results
poetry run python scripts/generate_eval_report.py \
--json-file eval_results.json \
--output-file docs/development/evaluations/latest-results.md \
--models "${{ steps.test-command.outputs.models }}"

# Also generate timestamped version for history
TIMESTAMP=$(date +%Y%m%d_%H%M%S)
poetry run python scripts/generate_eval_report.py \
--json-file eval_results.json \
--output-file "docs/development/evaluations/history/results_${TIMESTAMP}.md" \
--models "${{ steps.test-command.outputs.models }}"

- name: Upload eval results
if: always()
uses: actions/upload-artifact@v4
Expand Down
9 changes: 9 additions & 0 deletions docs/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
nav:
- index.md
- Installation: installation
- Walkthrough: walkthrough
- AI Providers: ai-providers
- Data Sources: data-sources
- Development: development
- Reference: reference
- Community: community.md
12 changes: 12 additions & 0 deletions docs/ai-providers/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
nav:
- index.md
- Anthropic: anthropic.md
- AWS Bedrock: aws-bedrock.md
- Azure OpenAI: azure-openai.md
- Gemini: gemini.md
- Google Vertex AI: google-vertex-ai.md
- Ollama: ollama.md
- OpenAI: openai.md
- OpenAI-Compatible: openai-compatible.md
- Robusta AI: robusta-ai.md
- Using Multiple Providers: using-multiple-providers.md
6 changes: 6 additions & 0 deletions docs/data-sources/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
nav:
- index.md
- Built-in Toolsets: builtin-toolsets
- Custom Toolsets: custom-toolsets.md
- Remote MCP Servers: remote-mcp-servers.md
- Adding Permissions for Additional Resources: permissions.md
29 changes: 29 additions & 0 deletions docs/data-sources/builtin-toolsets/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
nav:
- index.md
- AKS Node Health: aks-node-health.md
- ArgoCD: argocd.md
- AWS: aws.md
- Azure Kubernetes Service: aks.md
- Azure SQL Database: azure-sql.md
- Confluence: confluence.md
- Coralogix logs: coralogix-logs.md
- DataDog: datadog.md
- Datetime: datetime.md
- Docker: docker.md
- GitHub: github.md
- Loki: grafanaloki.md
- Tempo: grafanatempo.md
- Helm: helm.md
- Internet: internet.md
- Kafka: kafka.md
- Kubernetes: kubernetes.md
- MongoDB Atlas: mongodb-atlas.md
- New Relic: newrelic.md
- Notion: notion.md
- OpenSearch logs: opensearch-logs.md
- OpenSearch status: opensearch-status.md
- Prometheus: prometheus.md
- RabbitMQ: rabbitmq.md
- Robusta: robusta.md
- ServiceNow: servicenow.md
- Slab: slab.md
3 changes: 3 additions & 0 deletions docs/development/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
nav:
- index.md
- Evaluations: evaluations
7 changes: 7 additions & 0 deletions docs/development/evaluations/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
nav:
- index.md
- Latest Results: latest-results.md
- Historical Results: history
- Running Evaluations: running-evals.md
- Adding New Evaluations: adding-evals.md
- Reporting with Braintrust: reporting.md
3 changes: 3 additions & 0 deletions docs/development/evaluations/history/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
nav:
- index.md
- "*"
24 changes: 1 addition & 23 deletions docs/development/evaluations/history/index.md
Original file line number Diff line number Diff line change
@@ -1,27 +1,5 @@
# Historical Evaluation Results

This directory contains archived evaluation results from previous runs. These results help track HolmesGPT's performance over time and identify regressions or improvements.

## Available Results

- [Results from 2025-09-08 21:47:56](./results_20250908_214756.md)
- [Results from 2025-09-08 18:19:29](./results_20250908_181929.md)
- [Results from 2025-09-08 15:35:35](./results_20250908_153535.md)
- [Results from 2025-09-03 22:11:55](./results_20250903_221155.md)

## Understanding the Results

Each result file contains:
- Model performance metrics (pass rate, success rate)
- Individual test outcomes
- Execution times and resource usage
- Comparison across different models when applicable

## Automated Archiving

New results are automatically added here when:
- Weekly benchmarks run (every Sunday at 2 AM UTC)
- Manual benchmark runs are completed
- Significant model updates are tested

For the most recent results, see the [latest results](../latest-results.md) page.
All benchmark results are listed in the navigation sidebar.
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
# HolmesGPT LLM Evaluation Benchmark Results
# September 28, 2025 - 00:14:34

**Generated**: 2025-09-28 00:14 UTC
**Generated**: 2025-09-29 10:49 UTC
**Total Duration**: 1h 4m 41s
**Iterations**: 1
**Judge (classifier) model**: gpt-4o
**Judge (classifier) model**: gpt-4.1

## About this Benchmark

Expand Down
4 changes: 2 additions & 2 deletions docs/development/evaluations/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,7 +66,7 @@ Our CI/CD pipeline runs evaluations automatically:
- **Pull Requests** - When eval-related files are modified (quick validation)
- **On-demand** - Via GitHub Actions UI

Results are published here and archived in [history/](./history/).
Results are published here and archived in [history](./history/index.md).

## Model Comparison

Expand All @@ -86,4 +86,4 @@ See the [latest results](./latest-results.md) for current model performance comp
- **[Running Evaluations](./running-evals.md)** - Complete guide to running tests
- **[Adding New Evaluations](./adding-evals.md)** - Contribute test scenarios
- **[Reporting with Braintrust](./reporting.md)** - Analyze results in detail
- **[Historical Results](./history/)** - Past benchmark data
- **[Historical Results](./history/index.md)** - Past benchmark data
6 changes: 3 additions & 3 deletions docs/development/evaluations/latest-results.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
# HolmesGPT LLM Evaluation Benchmark Results
# Latest Benchmark Results

**Generated**: 2025-09-28 00:14 UTC
**Generated**: 2025-09-29 10:51 UTC
**Total Duration**: 1h 4m 41s
**Iterations**: 1
**Judge (classifier) model**: gpt-4o
**Judge (classifier) model**: gpt-4.1

## About this Benchmark

Expand Down
5 changes: 5 additions & 0 deletions docs/installation/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
nav:
- Install CLI: cli-installation.md
- Install UI/TUI: ui-installation.md
- Install Helm Chart: kubernetes-installation.md
- Install Python SDK: python-installation.md
6 changes: 6 additions & 0 deletions docs/reference/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
nav:
- Environment Variables: environment-variables.md
- Helm Configuration: helm-configuration.md
- HTTP API: http-api.md
- Slash Commands: slash-commands.md
- Troubleshooting: troubleshooting.md
6 changes: 6 additions & 0 deletions docs/walkthrough/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
nav:
- index.md
- Interactive Mode: interactive-mode.md
- CI/CD Troubleshooting: cicd-troubleshooting.md
- Investigating Prometheus Alerts: investigating-prometheus-alerts.md
- Investigating using AKS MCP Server: investigating-using-aks-mcp-server.md
84 changes: 1 addition & 83 deletions mkdocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,7 @@ theme:
logo: assets/logo.png

plugins:
- awesome-nav
- search:
separator: '[\s\-,:!=\[\]()"`/]+|\.(?!\d)|&[lg]t;|(?!\b)(?=[A-Z][a-z])'
- glightbox
Expand Down Expand Up @@ -134,86 +135,3 @@ extra:
link: https://bit.ly/robusta-slack
- icon: fontawesome/brands/twitter
link: https://twitter.com/RobustaDev

nav:
- Installation:
- index.md
- Install CLI: installation/cli-installation.md
- Install UI/TUI: installation/ui-installation.md
- Install Helm Chart: installation/kubernetes-installation.md
- Install Python SDK: installation/python-installation.md

- Walkthrough:
- walkthrough/index.md
- Interactive Mode: walkthrough/interactive-mode.md
- CI/CD Troubleshooting: walkthrough/cicd-troubleshooting.md
- Investigating Prometheus Alerts: walkthrough/investigating-prometheus-alerts.md
- Investigating using AKS MCP Server: walkthrough/investigating-using-aks-mcp-server.md

- AI Providers:
- ai-providers/index.md
- Anthropic: ai-providers/anthropic.md
- AWS Bedrock: ai-providers/aws-bedrock.md
- Azure OpenAI: ai-providers/azure-openai.md
- Gemini: ai-providers/gemini.md
- Google Vertex AI: ai-providers/google-vertex-ai.md
- Ollama: ai-providers/ollama.md
- OpenAI: ai-providers/openai.md
- OpenAI-Compatible: ai-providers/openai-compatible.md
- Robusta AI: ai-providers/robusta-ai.md
- Using Multiple Providers: ai-providers/using-multiple-providers.md

- Data Sources:
- data-sources/index.md
- Built-in Toolsets:
- data-sources/builtin-toolsets/index.md
- AKS Node Health: data-sources/builtin-toolsets/aks-node-health.md
- ArgoCD: data-sources/builtin-toolsets/argocd.md
- AWS: data-sources/builtin-toolsets/aws.md
- Azure Kubernetes Service: data-sources/builtin-toolsets/aks.md
- Azure SQL Database: data-sources/builtin-toolsets/azure-sql.md
- Confluence: data-sources/builtin-toolsets/confluence.md
- Coralogix logs: data-sources/builtin-toolsets/coralogix-logs.md
- DataDog: data-sources/builtin-toolsets/datadog.md
- Datetime: data-sources/builtin-toolsets/datetime.md
- Docker: data-sources/builtin-toolsets/docker.md
- GitHub: data-sources/builtin-toolsets/github.md
- Loki: data-sources/builtin-toolsets/grafanaloki.md
- Tempo: data-sources/builtin-toolsets/grafanatempo.md
- Helm: data-sources/builtin-toolsets/helm.md
- Internet: data-sources/builtin-toolsets/internet.md
- Kafka: data-sources/builtin-toolsets/kafka.md
- Kubernetes: data-sources/builtin-toolsets/kubernetes.md
- MongoDB Atlas: data-sources/builtin-toolsets/mongodb-atlas.md
- New Relic: data-sources/builtin-toolsets/newrelic.md
- Notion: data-sources/builtin-toolsets/notion.md
- OpenSearch logs: data-sources/builtin-toolsets/opensearch-logs.md
- OpenSearch status: data-sources/builtin-toolsets/opensearch-status.md
- Prometheus: data-sources/builtin-toolsets/prometheus.md
- RabbitMQ: data-sources/builtin-toolsets/rabbitmq.md
- Robusta: data-sources/builtin-toolsets/robusta.md
- ServiceNow: data-sources/builtin-toolsets/servicenow.md
- Slab: data-sources/builtin-toolsets/slab.md

- Custom Toolsets: data-sources/custom-toolsets.md
- Remote MCP Servers: data-sources/remote-mcp-servers.md
- Adding Permissions for Additional Resources: data-sources/permissions.md

- Development:
- development/index.md
- Evaluations:
- development/evaluations/index.md
- Latest Results: development/evaluations/latest-results.md
- Running Evaluations: development/evaluations/running-evals.md
- Adding New Evaluations: development/evaluations/adding-evals.md
- Reporting with Braintrust: development/evaluations/reporting.md
- Historical Results: development/evaluations/history/index.md

- Reference:
- Environment Variables: reference/environment-variables.md
- reference/helm-configuration.md
- HTTP API: reference/http-api.md
- Slash Commands: reference/slash-commands.md
- Troubleshooting: reference/troubleshooting.md

- Community: community.md
Loading
Loading