Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions .github/workflows/eval-benchmarks.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -115,11 +115,19 @@ jobs:
- name: Generate benchmark report
if: always()
run: |
# Generate latest results
poetry run python scripts/generate_eval_report.py \
--json-file eval_results.json \
--output-file docs/development/evaluations/latest-results.md \
--models "${{ steps.test-command.outputs.models }}"

# Also generate timestamped version for history
TIMESTAMP=$(date +%Y%m%d_%H%M%S)
poetry run python scripts/generate_eval_report.py \
--json-file eval_results.json \
--output-file "docs/development/evaluations/history/results_${TIMESTAMP}.md" \
--models "${{ steps.test-command.outputs.models }}"

- name: Upload eval results
if: always()
uses: actions/upload-artifact@v4
Expand Down
9 changes: 9 additions & 0 deletions docs/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
nav:
- index.md
- Installation: installation
- Walkthrough: walkthrough
- AI Providers: ai-providers
- Data Sources: data-sources
- Development: development
- Reference: reference
- Community: community.md
12 changes: 12 additions & 0 deletions docs/ai-providers/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
nav:
- index.md
- Anthropic: anthropic.md
- AWS Bedrock: aws-bedrock.md
- Azure OpenAI: azure-openai.md
- Gemini: gemini.md
- Google Vertex AI: google-vertex-ai.md
- Ollama: ollama.md
- OpenAI: openai.md
- OpenAI-Compatible: openai-compatible.md
- Robusta AI: robusta-ai.md
- Using Multiple Providers: using-multiple-providers.md
6 changes: 6 additions & 0 deletions docs/data-sources/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
nav:
- index.md
- Built-in Toolsets: builtin-toolsets
- Custom Toolsets: custom-toolsets.md
- Remote MCP Servers: remote-mcp-servers.md
- Adding Permissions for Additional Resources: permissions.md
29 changes: 29 additions & 0 deletions docs/data-sources/builtin-toolsets/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
nav:
- index.md
- AKS Node Health: aks-node-health.md
- ArgoCD: argocd.md
- AWS: aws.md
- Azure Kubernetes Service: aks.md
- Azure SQL Database: azure-sql.md
- Confluence: confluence.md
- Coralogix logs: coralogix-logs.md
- DataDog: datadog.md
- Datetime: datetime.md
- Docker: docker.md
- GitHub: github.md
- Loki: grafanaloki.md
- Tempo: grafanatempo.md
- Helm: helm.md
- Internet: internet.md
- Kafka: kafka.md
- Kubernetes: kubernetes.md
- MongoDB Atlas: mongodb-atlas.md
- New Relic: newrelic.md
- Notion: notion.md
- OpenSearch logs: opensearch-logs.md
- OpenSearch status: opensearch-status.md
- Prometheus: prometheus.md
- RabbitMQ: rabbitmq.md
- Robusta: robusta.md
- ServiceNow: servicenow.md
- Slab: slab.md
3 changes: 3 additions & 0 deletions docs/development/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
nav:
- index.md
- Evaluations: evaluations
7 changes: 7 additions & 0 deletions docs/development/evaluations/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
nav:
- index.md
- Latest Results: latest-results.md
- Historical Results: history
- Running Evaluations: running-evals.md
- Adding New Evaluations: adding-evals.md
- Reporting with Braintrust: reporting.md
3 changes: 3 additions & 0 deletions docs/development/evaluations/history/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
nav:
- index.md
- "*"
24 changes: 1 addition & 23 deletions docs/development/evaluations/history/index.md
Original file line number Diff line number Diff line change
@@ -1,27 +1,5 @@
# Historical Evaluation Results

This directory contains archived evaluation results from previous runs. These results help track HolmesGPT's performance over time and identify regressions or improvements.

## Available Results

- [Results from 2025-09-08 21:47:56](./results_20250908_214756.md)
- [Results from 2025-09-08 18:19:29](./results_20250908_181929.md)
- [Results from 2025-09-08 15:35:35](./results_20250908_153535.md)
- [Results from 2025-09-03 22:11:55](./results_20250903_221155.md)

## Understanding the Results

Each result file contains:
- Model performance metrics (pass rate, success rate)
- Individual test outcomes
- Execution times and resource usage
- Comparison across different models when applicable

## Automated Archiving

New results are automatically added here when:
- Weekly benchmarks run (every Sunday at 2 AM UTC)
- Manual benchmark runs are completed
- Significant model updates are tested

For the most recent results, see the [latest results](../latest-results.md) page.
All benchmark results are listed in the navigation sidebar.
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
# HolmesGPT LLM Evaluation Benchmark Results
# September 28, 2025 - 00:14:34

**Generated**: 2025-09-28 00:14 UTC
**Generated**: 2025-09-29 10:49 UTC
**Total Duration**: 1h 4m 41s
**Iterations**: 1
**Judge (classifier) model**: gpt-4o
**Judge (classifier) model**: gpt-4.1

## About this Benchmark

Expand Down
305 changes: 305 additions & 0 deletions docs/development/evaluations/history/results_20250930_085923.md

Large diffs are not rendered by default.

4 changes: 2 additions & 2 deletions docs/development/evaluations/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,7 +66,7 @@ Our CI/CD pipeline runs evaluations automatically:
- **Pull Requests** - When eval-related files are modified (quick validation)
- **On-demand** - Via GitHub Actions UI

Results are published here and archived in [history/](./history/).
Results are published here and archived in [history](./history/index.md).

## Model Comparison

Expand All @@ -86,4 +86,4 @@ See the [latest results](./latest-results.md) for current model performance comp
- **[Running Evaluations](./running-evals.md)** - Complete guide to running tests
- **[Adding New Evaluations](./adding-evals.md)** - Contribute test scenarios
- **[Reporting with Braintrust](./reporting.md)** - Analyze results in detail
- **[Historical Results](./history/)** - Past benchmark data
- **[Historical Results](./history/index.md)** - Past benchmark data
Loading
Loading