Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/eval-benchmarks.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -125,7 +125,7 @@ jobs:
TIMESTAMP=$(date +%Y%m%d_%H%M%S)
poetry run python scripts/generate_eval_report.py \
--json-file eval_results.json \
--output-file "docs/development/evaluations/history/results_${TIMESTAMP}.md" \
--output-file "docs/development/evaluations/history/weekly/results_${TIMESTAMP}.md" \
--models "${{ steps.test-command.outputs.models }}"

- name: Upload eval results
Expand Down
3 changes: 2 additions & 1 deletion docs/development/evaluations/history/.nav.yml
Original file line number Diff line number Diff line change
@@ -1,3 +1,4 @@
nav:
- index.md
- "*"
- Weekly: weekly/
- Special: special/
17 changes: 15 additions & 2 deletions docs/development/evaluations/history/index.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,18 @@
# Historical Evaluation Results

## Available Results
## Weekly Runs

All benchmark results are listed in the navigation sidebar.
Weekly benchmark runs with a standard set of models.

See the **Weekly** section in the navigation sidebar for all weekly benchmark results.

## Special Benchmark Runs

One-off benchmark runs for specific purposes such as:

- Comparing self-hosted models
- Testing new model versions
- Performance analysis for specific scenarios
- Custom model comparisons

See the **Special** section in the navigation sidebar for all special benchmark runs.
5 changes: 5 additions & 0 deletions docs/development/evaluations/history/special/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
sort:
direction: desc
nav:
- index.md
- "*"
7 changes: 7 additions & 0 deletions docs/development/evaluations/history/special/index.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# Special Benchmark Runs

One-off benchmark runs for specific purposes such as comparing self-hosted models, testing new model versions, or custom performance analysis.

## Available Results

All special benchmark results are listed in the navigation sidebar.
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
# September 30, 2025
# Claude 4.0 vs 4.5 (n=5)

**Generated**: 2025-09-30 15:37 UTC

Expand Down
Original file line number Diff line number Diff line change
@@ -1,8 +1,11 @@
# HolmesGPT LLM Evaluation Benchmark Results
# Self-Hosted Models v1

**Generated**: 2025-10-08 05:37 UTC

**Total Duration**: 5h 38m 14s

**Iterations**: 1

**Judge (classifier) model**: gpt-4o

## About this Benchmark
Expand Down
5 changes: 5 additions & 0 deletions docs/development/evaluations/history/weekly/.nav.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
sort:
direction: desc
nav:
- index.md
- "*"
7 changes: 7 additions & 0 deletions docs/development/evaluations/history/weekly/index.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# Weekly Benchmark Runs

Weekly benchmark runs with a standard set of models.

## Available Results

All weekly benchmark results are listed in the navigation sidebar.
Original file line number Diff line number Diff line change
@@ -1,8 +1,11 @@
# HolmesGPT LLM Evaluation Benchmark Results
# October 12, 2025

**Generated**: 2025-10-12 17:03 UTC

**Total Duration**: 1h 48m 23s

**Iterations**: 1

**Judge (classifier) model**: azure/gpt-4.1

## About this Benchmark
Expand Down
2 changes: 1 addition & 1 deletion docs/development/evaluations/running-evals.md
Original file line number Diff line number Diff line change
Expand Up @@ -137,7 +137,7 @@ RUN_LIVE=true poetry run pytest -m 'llm and easy' --no-cov

Some evals support mock-data and don't need a live Kubernetes cluster to run. However, for the most accurate evaluation you should set `RUN_LIVE=true` which tests HolmesGPT with a live Kubernetes cluster not mock data.

This is important because LLMs can take multiple paths to reach conclusions, and mock data only captures one path. See [Using Mock Data](../../using-mock-data.md) for rare cases when mocks are necessary.
This is important because LLMs can take multiple paths to reach conclusions, and mock data only captures one path.

## Environment Variables

Expand Down
22 changes: 12 additions & 10 deletions run_benchmarks_local.sh
Original file line number Diff line number Diff line change
Expand Up @@ -147,11 +147,22 @@ echo "================================"
echo ""
echo "Generating benchmark report..."
if [ -f "scripts/generate_eval_report.py" ]; then
# Generate latest results
poetry run python scripts/generate_eval_report.py \
--json-file eval_results.json \
--output-file docs/development/evaluations/latest-results.md \
--models "$MODELS"
echo "✅ Report generated: docs/development/evaluations/latest-results.md"

# Also generate timestamped version for history (always in weekly/)
mkdir -p docs/development/evaluations/history/weekly
TIMESTAMP=$(date +%Y%m%d_%H%M%S)
HISTORY_FILE="docs/development/evaluations/history/weekly/results_${TIMESTAMP}.md"
poetry run python scripts/generate_eval_report.py \
--json-file eval_results.json \
--output-file "$HISTORY_FILE" \
--models "$MODELS"
echo "📁 Saved historical copy: $HISTORY_FILE"
else
echo "⚠️ Report generation script not found: scripts/generate_eval_report.py"
fi
Expand All @@ -162,16 +173,7 @@ echo "Generated files:"
[ -f "eval_results.json" ] && echo " ✓ eval_results.json ($(wc -l < eval_results.json) lines)"
[ -f "evals_report.md" ] && echo " ✓ evals_report.md ($(wc -l < evals_report.md) lines)"
[ -f "docs/development/evaluations/latest-results.md" ] && echo " ✓ docs/development/evaluations/latest-results.md ($(wc -l < docs/development/evaluations/latest-results.md) lines)"

# Save historical copy
if [ -f "docs/development/evaluations/latest-results.md" ]; then
mkdir -p docs/development/evaluations/history
TIMESTAMP=$(date +%Y%m%d_%H%M%S)
HISTORY_FILE="docs/development/evaluations/history/results_${TIMESTAMP}.md"
cp docs/development/evaluations/latest-results.md "$HISTORY_FILE"
echo ""
echo "📁 Saved historical copy: $HISTORY_FILE"
fi
[ -f "$HISTORY_FILE" ] && echo " ✓ $HISTORY_FILE ($(wc -l < "$HISTORY_FILE") lines)"

echo ""
echo "=============================================="
Expand Down
Loading