Skip to content

Test cli performance - #1415

Merged
aantn merged 29 commits into
masterfrom
claude/holmes-performance-benchmarks-yKW8b
Feb 1, 2026
Merged

aantn merged 29 commits into
masterfrom
claude/holmes-performance-benchmarks-yKW8b

Conversation

@aantn

@aantn aantn commented Jan 24, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • Added an automated CLI performance benchmark tool to measure startup and optional end-to-end behavior.
    • Introduced a workflow that runs benchmarks on PRs and master, compares results, and posts a comparison report.
  • Chores

    • CI now fails when startup time regresses beyond the configured threshold (20%), helping prevent performance regressions.

✏️ Tip: You can customize this high-level summary in your review settings.

claude added 22 commits January 23, 2026 14:35
Add black-box performance testing that measures wall time of
`holmes ask` with a simple prompt. This helps detect performance
regressions in PRs by comparing against the master branch.

Components:
- scripts/cli_performance_benchmark.py: Standalone benchmark script
  - Measures wall time of CLI execution
  - Multiple iterations for statistical reliability
  - Outputs JSON for easy comparison
  - Supports baseline comparison with markdown report

- .github/workflows/cli-performance.yaml: GitHub Actions workflow
  - Runs on every PR to master
  - Benchmarks both PR and master branches
  - Posts comparison comment to PR
  - Fails CI if >10% regression detected

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
Focus on deterministic startup time measurement:
- `--startup-only`: Measures `holmes version` (imports + init only)
- `--e2e-only`: Measures full `holmes ask` (startup + LLM call)
- Combined mode: Reports both metrics separately

Startup benchmark is the primary metric for tracking import overhead
regressions. E2E serves as a sanity check that CLI commands work.

Workflow changes:
- Primary job: startup-benchmark (no API key needed)
- Secondary job: e2e-sanity-check (requires API key)
- Tighter regression threshold for startup (20% vs 10%)
- PR comment highlights startup-specific issues to check

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
Changes:
- Report both cold start (first run) and warm start (subsequent runs)
- Cold start = no bytecode cache, no fs cache
- Warm start = caches populated from previous runs
- Include min/max/stdev in warm start stats

Persistence:
- Store baseline in GitHub Actions cache
- Update baseline on push to master
- PRs compare against cached baseline
- Cache key includes runner OS for consistency

This enables tracking both first-run experience (cold) and typical
usage (warm) separately, as they can regress independently.

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
Remove push-to-master trigger and cache-based persistence.
Instead, benchmark both branches in a single PR workflow run.

Simpler and always fresh comparison (no cache staleness issues).

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
The benchmark script doesn't exist on master yet, so we need to
copy it before checking out master branch. This allows the first
PR that adds the benchmark to still compare against master.

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
The multi-line f-string in heredoc confused YAML parser because
unindented lines looked like YAML keys. Using list of strings
and joining them avoids this issue.

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
Git won't overwrite untracked files during checkout, so we need
to remove the benchmark script we copied to master before switching
back to the PR branch.

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
The report content may have conflicted with the 'EOF' delimiter.
Using a unique delimiter to avoid parsing issues.

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
Fixes:
- Output master_benchmark.json to /tmp to preserve across checkout
- Copy both benchmark files back from /tmp after checkout
- Use random delimiter for GITHUB_OUTPUT (recommended approach)

Naming clarifications:
- "Benchmark: Startup Time (no LLM)" - measures holmes version
- "Benchmark: Full CLI (with LLM)" - measures holmes ask

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
Avoid heredoc delimiter issues by reading comparison_report.md
directly in the JavaScript step using fs.readFileSync.

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
Changes:
- Use OPENROUTER_API_KEY instead of OpenAI
- Use model openrouter/anthropic/claude-3-5-haiku-latest
- If API key not available, still update PR comment saying so
- LLM job runs after no-LLM job and appends to same comment

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
- Add stdout/stderr to BenchmarkResult and print on failure
- Fix deprecation warning: use datetime.now(timezone.utc)
- Use correct OpenRouter model alias: claude-haiku-4.5

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
- Combine startup and LLM benchmarks into single job (DRY)
- Both benchmarks now compare PR vs master with same table format
- Consistent information hierarchy with single top-level heading
- LLM benchmark shows full comparison table when API key available

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
- Remove ~200 lines of unused code (compare_results, format_comparison_report, etc.)
- Unify run_command and run_benchmark to eliminate duplication
- Both startup and e2e benchmarks now use same code path
- Simplify CLI - remove unused --compare and --fail-on-regression options
- Fix PR comment search to find both old and new format names

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
The previous benchmark showed 117% cold start difference between PR and
master with no actual code changes - this was because master's "cold"
start was actually warm (benefiting from PR's cached bytecode and venv).

Now we clear __pycache__ and .venv before benchmarking master to get
a true cold start measurement.

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
Shows a banner at the top of the existing comment while benchmark runs,
with a link to the workflow logs. Old results remain visible until the
new benchmark completes and replaces them entirely.

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
Previous attempts (clearing bytecode + venv) still showed ~43% cold start
differences due to OS page cache keeping Python interpreter and shared
libraries in memory.

Now we also run `sync && sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'`
to clear page cache, dentries, and inodes before master benchmark.

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
- Clear Python bytecode (__pycache__, *.pyc)
- Remove .venv
- Clear Holmes cache (~/.holmes)
- Clear pip/poetry caches (~/.cache/pip, ~/.cache/pypoetry)
- Drop OS page cache
- Memory pressure: read 6GB random data to evict cached files from RAM
- Drop caches again after memory pressure

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
- Add console.log for debugging
- Use hardcoded github.com URL instead of process.env
- Add continue-on-error so workflow doesn't fail if banner update fails
- Fix regex with 's' flag for multiline

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
…atch

- Fix bot detection: use `github-actions[bot]` login instead of `type === 'Bot'`
  (prevents accidentally editing CodeRabbit's comment)
- Fix script injection: save github.head_ref to file via env var, read it back
- Add workflow_dispatch guard: only run on pull_request events
- Add --sync flag to poetry install for clean dependency management

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
@linux-foundation-easycla

linux-foundation-easycla Bot commented Jan 24, 2026 •

Copy link
Copy Markdown

CLA Signed

The committers listed above are authorized under a signed CLA.

  • ✅ login: moshemorad / name: moshemorad (fde5b9a)

@netlify

netlify Bot commented Jan 24, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit fde5b9a
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/697f593f892dd600086cc33f
😎 Deploy Preview https://deploy-preview-1415--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai

coderabbitai Bot commented Jan 24, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Adds a GitHub Actions workflow and a Python benchmarking script to measure CLI startup and optional LLM end-to-end performance on PRs and master, compare results, post a comparison report to PRs, and fail when startup regression exceeds 20%.

Changes

Cohort / File(s) Summary
CI/CD Workflow Configuration
.github/workflows/cli-performance.yaml
New GitHub Actions workflow "CLI Performance Benchmark" that runs PR and master benchmarks, conditionally runs LLM benchmarks when an API key is present, compares results, posts/updates a PR comment with a comparison report, and fails on >20% startup regression.
Benchmarking Utility
scripts/cli_performance_benchmark.py
New Python CLI benchmarking tool that measures startup and optional e2e timings, defines BenchmarkResult and BenchmarkSummary, captures git metadata, runs subprocess iterations, computes cold/warm stats, validates iterations, outputs JSON and human-readable summaries.

Sequence Diagram

sequenceDiagram
    actor GitHub as GitHub Events
    participant Workflow as "GitHub Actions\nWorkflow"
    participant BenchScript as "Benchmark\nScript"
    participant CLI as "Holmes CLI"
    participant Comparison as "Comparison\nLogic"
    participant Bot as "PR Comment\nUpdater"

    GitHub->>Workflow: Trigger (PR or manual)
    Workflow->>Workflow: Checkout PR branch
    Workflow->>BenchScript: Run startup benchmark (PR)
    BenchScript->>CLI: Execute `holmes version` (multiple iterations)
    CLI-->>BenchScript: Timing & exit results
    BenchScript-->>Workflow: pr_startup.json

    alt OPENROUTER_API_KEY set
        Workflow->>BenchScript: Run LLM benchmark (PR)
        BenchScript->>CLI: Execute `holmes ask` (multiple iterations)
        CLI-->>BenchScript: Timing & exit results
        BenchScript-->>Workflow: pr_llm.json
    end

    Workflow->>Workflow: Checkout master, sparse-copy PR script
    Workflow->>BenchScript: Run startup benchmark (master)
    BenchScript->>CLI: Execute `holmes version` (multiple iterations)
    CLI-->>BenchScript: Timing & exit results
    BenchScript-->>Workflow: master_startup.json

    alt OPENROUTER_API_KEY set
        Workflow->>BenchScript: Run LLM benchmark (master)
        BenchScript->>CLI: Execute `holmes ask` (multiple iterations)
        CLI-->>BenchScript: Timing & exit results
        BenchScript-->>Workflow: master_llm.json
    end

    Workflow->>Comparison: Generate comparison report
    Comparison->>Comparison: Load PR & master results, compute diffs, classify
    Comparison-->>Workflow: comparison_report.md + status
    Workflow->>Bot: Post/update PR comment with report
    Workflow->>Workflow: Enforce regression threshold (>20% fails)
    Workflow-->>GitHub: Job result (pass/fail)
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~30 minutes

🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 inconclusive)
Check name Status Explanation Resolution
Title check ❓ Inconclusive The title 'Test cli performance' is vague and generic, lacking specificity about what is being added or tested. Revise the title to be more descriptive, such as 'Add CLI performance benchmarking workflow and utilities' to clearly convey the main changes being introduced.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Docstring Coverage ✅ Passed Docstring coverage is 83.33% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jan 24, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 25ad5c7 (#21563205706)

✅ Results of HolmesGPT evals

Automatically triggered by commit 25ad5c7 on branch claude/holmes-performance-benchmarks-yKW8b

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 29.5s ±0% 5 10 $0.2220
✅ 101_loki_historical_logs_pod_deleted 44.8s ±0% 6 11 $0.2690
✅ 111_pod_names_contain_service 32.2s ±0% 5 11 $0.2290
✅ 12_job_crashing 36.9s ↑12% 6 14 $0.2559
✅ 162_get_runbooks 37.9s ±0% 6 11 $0.2828
✅ 176_network_policy_blocking_traffic_no_runbooks 42.9s ±0% 7 16 $0.2765
✅ 24_misconfigured_pvc 36.0s ±0% 6 15 $0.2430
✅ 43_current_datetime_from_prompt 5.3s ±0% 1 — $0.1060
✅ 61_exact_match_counting 15.8s ±0% 4 4 $0.1590
Total 31.3s avg 5.1 avg 11.5 avg $2.0432

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/holmes-performance-benchmarks-yKW8b'

Status: Success - 38 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ f2108c1 (#21562585752)

✅ Results of HolmesGPT evals

Automatically triggered by commit f2108c1 on branch claude/holmes-performance-benchmarks-yKW8b

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 30.3s ±0% 5 11 $0.2281
✅ 101_loki_historical_logs_pod_deleted 37.8s ↓11% 6 8 $0.2308
✅ 111_pod_names_contain_service 34.9s ±0% 5 12 $0.2413
✅ 12_job_crashing 39.3s ↑14% 6 14 $0.2692
✅ 162_get_runbooks 37.9s ±0% 6 11 $0.2665
✅ 176_network_policy_blocking_traffic_no_runbooks 38.0s ↓11% 6 12 $0.2546
✅ 24_misconfigured_pvc 33.9s ±0% 5 14 $0.2336
✅ 43_current_datetime_from_prompt 5.5s ±0% 1 — $0.1063
✅ 61_exact_match_counting 16.6s ±0% 4 4 $0.1589
Total 30.5s avg 4.9 avg 10.8 avg $1.9893

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/holmes-performance-benchmarks-yKW8b'

Status: Success - 38 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ b4ed214 (#21514930829)

✅ Results of HolmesGPT evals

Automatically triggered by commit b4ed214 on branch claude/holmes-performance-benchmarks-yKW8b

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 35.2s ↑13% 6 11 $0.2492
✅ 101_loki_historical_logs_pod_deleted 39.5s ↓11% 5 10 $0.2334
✅ 111_pod_names_contain_service 33.8s ±0% 5 11 $0.2230
✅ 12_job_crashing 33.5s ±0% 5 11 $0.2355
✅ 162_get_runbooks 38.7s ±0% 5 10 $0.2472
✅ 176_network_policy_blocking_traffic_no_runbooks 47.3s ±0% 6 18 $0.2994
✅ 24_misconfigured_pvc 37.9s ↑13% 6 14 $0.2473
✅ 43_current_datetime_from_prompt 5.5s ±0% 1 — $0.1053
✅ 61_exact_match_counting 12.7s ↓13% 3 2 $0.1418
Total 31.6s avg 4.7 avg 10.9 avg $1.9820

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/holmes-performance-benchmarks-yKW8b'

Status: Success - 13 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 87a210a (#21329273939)

✅ Results of HolmesGPT evals

Automatically triggered by commit 87a210a on branch claude/holmes-performance-benchmarks-yKW8b

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 26.7s ↑10% 5 9 $0.1164
✅ 101_loki_historical_logs_pod_deleted 38.9s ↓27% 6 11 $0.1616
✅ 111_pod_names_contain_service 21.8s ↓24% 4 7 $0.0835
✅ 12_job_crashing 32.3s ↓13% 5 12 $0.1466
✅ 162_get_runbooks 36.6s ±0% 6 11 $0.1739
✅ 176_network_policy_blocking_traffic_no_runbooks 42.8s ±0% 6 17 $0.1894
✅ 24_misconfigured_pvc 30.9s ↓19% 6 13 $0.1358
✅ 43_current_datetime_from_prompt 3.5s ±0% 1 — $0.0106
✅ 61_exact_match_counting 15.4s ±0% 4 4 $0.0655
Total 27.6s avg 4.8 avg 10.5 avg $1.0833

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/holmes-performance-benchmarks-yKW8b'

Status: Success - 20 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ fb408b6 (#21321240745)

✅ Results of HolmesGPT evals

Automatically triggered by commit fb408b6 on branch claude/holmes-performance-benchmarks-yKW8b

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 21.4s ↓14% 4 7 $0.1670
✅ 101_loki_historical_logs_pod_deleted 42.8s ↓25% 7 12 $0.2731
✅ 111_pod_names_contain_service 23.9s ↓21% 4 7 $0.1667
✅ 12_job_crashing 34.9s ±0% 5 14 $0.2391
✅ 162_get_runbooks 42.8s ±0% 7 12 $0.2782
✅ 176_network_policy_blocking_traffic_no_runbooks 40.9s ±0% 7 17 $0.2752
✅ 24_misconfigured_pvc 35.2s ±0% 6 15 $0.2375
✅ 43_current_datetime_from_prompt 3.9s ±0% 1 — $0.0909
✅ 61_exact_match_counting 15.7s ±0% 4 4 $0.1459
Total 29.1s avg 5.0 avg 11.0 avg $1.8735

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/holmes-performance-benchmarks-yKW8b'

Status: Success - 19 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit fde5b9a on branch claude/holmes-performance-benchmarks-yKW8b

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 30.6s ±0% 5 10 $0.2181
✅ 101_loki_historical_logs_pod_deleted 35.2s ↓17% 5 9 $0.2336
✅ 111_pod_names_contain_service 33.9s ±0% 5 12 $0.2382
✅ 12_job_crashing 41.1s ↑19% 7 15 $0.2799
✅ 162_get_runbooks 40.4s ±0% 6 11 $0.2693
✅ 176_network_policy_blocking_traffic_no_runbooks 48.7s ±0% 7 16 $0.2903
✅ 24_misconfigured_pvc 33.4s ±0% 5 15 $0.2387
✅ 43_current_datetime_from_prompt 4.8s ±0% 1 — $0.1054
✅ 61_exact_match_counting 16.8s ↑12% 4 4 $0.1597
Total 31.6s avg 5.0 avg 11.5 avg $2.0332

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/holmes-performance-benchmarks-yKW8b'

Status: Success - 38 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/holmes-performance-benchmarks-yKW8b -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/holmes-performance-benchmarks-yKW8b -f markers=regression -f filter=

@github-actions

github-actions Bot commented Jan 24, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for d502f03 (built in 1m 8s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:d502f03
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:d502f03 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:d502f03
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:d502f03

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:d502f03

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:d502f03

@github-actions

github-actions Bot commented Jan 24, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟢 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 8.73s 9.98s -12.5%
Warm Mean 3.97s 4.50s -11.8%
Warm Min 3.93s 4.46s
Warm Max 4.01s 4.55s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 24.47s 30.27s -19.2%
Warm Mean 7.55s 6.91s +9.1%
Warm Min 6.36s 6.87s
Warm Max 8.69s 7.03s

PR: d502f03b | Master: 19bb6165 | Iterations: 5

GitHub already shows when comments are updated.

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
8GB wasn't consistently clearing caches, try 16GB (2x runner RAM).

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In @.github/workflows/cli-performance.yaml:
- Around line 21-22: The workflow uses a workflow_dispatch trigger but the job
has a hard condition "if: github.event_name == 'pull_request'", which prevents
manual runs and makes the "iterations" input unusable; fix by either removing
the workflow_dispatch trigger or changing the job condition to allow manual runs
(e.g., allow github.event_name == 'workflow_dispatch' as well) and update the PR
comment steps to guard against missing context.issue.number when event is
workflow_dispatch (skip or conditionally run comment steps for non-PR runs).

Comment thread .github/workflows/cli-performance.yaml Outdated
claude and others added 4 commits January 25, 2026 07:56
Instead of running both benchmarks on the same runner and trying to
clear caches (which was unreliable), use separate jobs:

1. benchmark-pr: Fresh VM, benchmarks PR branch
2. benchmark-master: Fresh VM, benchmarks master branch
3. compare: Combines results and posts PR comment

Benefits:
- True cold start isolation (separate VMs)
- Jobs run in parallel (faster)
- No cache clearing workarounds needed
- Simpler, more reliable

Removed:
- Running banner (complexity not worth it with separate jobs)
- All cache clearing logic (bytecode, venv, memory pressure, etc.)

https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn enabled auto-merge (squash) February 1, 2026 12:52
@aantn
aantn merged commit a4b1825 into master Feb 1, 2026
19 of 20 checks passed
@aantn
aantn deleted the claude/holmes-performance-benchmarks-yKW8b branch February 1, 2026 13:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants