Repository navigation
Test cli performance - #1415
Test cli performance#1415
Conversation
Add black-box performance testing that measures wall time of `holmes ask` with a simple prompt. This helps detect performance regressions in PRs by comparing against the master branch. Components: - scripts/cli_performance_benchmark.py: Standalone benchmark script - Measures wall time of CLI execution - Multiple iterations for statistical reliability - Outputs JSON for easy comparison - Supports baseline comparison with markdown report - .github/workflows/cli-performance.yaml: GitHub Actions workflow - Runs on every PR to master - Benchmarks both PR and master branches - Posts comparison comment to PR - Fails CI if >10% regression detected https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
Focus on deterministic startup time measurement: - `--startup-only`: Measures `holmes version` (imports + init only) - `--e2e-only`: Measures full `holmes ask` (startup + LLM call) - Combined mode: Reports both metrics separately Startup benchmark is the primary metric for tracking import overhead regressions. E2E serves as a sanity check that CLI commands work. Workflow changes: - Primary job: startup-benchmark (no API key needed) - Secondary job: e2e-sanity-check (requires API key) - Tighter regression threshold for startup (20% vs 10%) - PR comment highlights startup-specific issues to check https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
Changes: - Report both cold start (first run) and warm start (subsequent runs) - Cold start = no bytecode cache, no fs cache - Warm start = caches populated from previous runs - Include min/max/stdev in warm start stats Persistence: - Store baseline in GitHub Actions cache - Update baseline on push to master - PRs compare against cached baseline - Cache key includes runner OS for consistency This enables tracking both first-run experience (cold) and typical usage (warm) separately, as they can regress independently. https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
Remove push-to-master trigger and cache-based persistence. Instead, benchmark both branches in a single PR workflow run. Simpler and always fresh comparison (no cache staleness issues). https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
The benchmark script doesn't exist on master yet, so we need to copy it before checking out master branch. This allows the first PR that adds the benchmark to still compare against master. https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
The multi-line f-string in heredoc confused YAML parser because unindented lines looked like YAML keys. Using list of strings and joining them avoids this issue. https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
Git won't overwrite untracked files during checkout, so we need to remove the benchmark script we copied to master before switching back to the PR branch. https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
The report content may have conflicted with the 'EOF' delimiter. Using a unique delimiter to avoid parsing issues. https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
Fixes: - Output master_benchmark.json to /tmp to preserve across checkout - Copy both benchmark files back from /tmp after checkout - Use random delimiter for GITHUB_OUTPUT (recommended approach) Naming clarifications: - "Benchmark: Startup Time (no LLM)" - measures holmes version - "Benchmark: Full CLI (with LLM)" - measures holmes ask https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
Avoid heredoc delimiter issues by reading comparison_report.md directly in the JavaScript step using fs.readFileSync. https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
Changes: - Use OPENROUTER_API_KEY instead of OpenAI - Use model openrouter/anthropic/claude-3-5-haiku-latest - If API key not available, still update PR comment saying so - LLM job runs after no-LLM job and appends to same comment https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
- Add stdout/stderr to BenchmarkResult and print on failure - Fix deprecation warning: use datetime.now(timezone.utc) - Use correct OpenRouter model alias: claude-haiku-4.5 https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
- Combine startup and LLM benchmarks into single job (DRY) - Both benchmarks now compare PR vs master with same table format - Consistent information hierarchy with single top-level heading - LLM benchmark shows full comparison table when API key available https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
- Remove ~200 lines of unused code (compare_results, format_comparison_report, etc.) - Unify run_command and run_benchmark to eliminate duplication - Both startup and e2e benchmarks now use same code path - Simplify CLI - remove unused --compare and --fail-on-regression options - Fix PR comment search to find both old and new format names https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
The previous benchmark showed 117% cold start difference between PR and master with no actual code changes - this was because master's "cold" start was actually warm (benefiting from PR's cached bytecode and venv). Now we clear __pycache__ and .venv before benchmarking master to get a true cold start measurement. https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
Shows a banner at the top of the existing comment while benchmark runs, with a link to the workflow logs. Old results remain visible until the new benchmark completes and replaces them entirely. https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
Previous attempts (clearing bytecode + venv) still showed ~43% cold start differences due to OS page cache keeping Python interpreter and shared libraries in memory. Now we also run `sync && sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'` to clear page cache, dentries, and inodes before master benchmark. https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
- Clear Python bytecode (__pycache__, *.pyc) - Remove .venv - Clear Holmes cache (~/.holmes) - Clear pip/poetry caches (~/.cache/pip, ~/.cache/pypoetry) - Drop OS page cache - Memory pressure: read 6GB random data to evict cached files from RAM - Drop caches again after memory pressure https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
- Add console.log for debugging - Use hardcoded github.com URL instead of process.env - Add continue-on-error so workflow doesn't fail if banner update fails - Fix regex with 's' flag for multiline https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
…atch - Fix bot detection: use `github-actions[bot]` login instead of `type === 'Bot'` (prevents accidentally editing CodeRabbit's comment) - Fix script injection: save github.head_ref to file via env var, read it back - Add workflow_dispatch guard: only run on pull_request events - Add --sync flag to poetry install for clean dependency management https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
|
|
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
WalkthroughAdds a GitHub Actions workflow and a Python benchmarking script to measure CLI startup and optional LLM end-to-end performance on PRs and master, compare results, post a comparison report to PRs, and fail when startup regression exceeds 20%. Changes
Sequence DiagramsequenceDiagram
actor GitHub as GitHub Events
participant Workflow as "GitHub Actions\nWorkflow"
participant BenchScript as "Benchmark\nScript"
participant CLI as "Holmes CLI"
participant Comparison as "Comparison\nLogic"
participant Bot as "PR Comment\nUpdater"
GitHub->>Workflow: Trigger (PR or manual)
Workflow->>Workflow: Checkout PR branch
Workflow->>BenchScript: Run startup benchmark (PR)
BenchScript->>CLI: Execute `holmes version` (multiple iterations)
CLI-->>BenchScript: Timing & exit results
BenchScript-->>Workflow: pr_startup.json
alt OPENROUTER_API_KEY set
Workflow->>BenchScript: Run LLM benchmark (PR)
BenchScript->>CLI: Execute `holmes ask` (multiple iterations)
CLI-->>BenchScript: Timing & exit results
BenchScript-->>Workflow: pr_llm.json
end
Workflow->>Workflow: Checkout master, sparse-copy PR script
Workflow->>BenchScript: Run startup benchmark (master)
BenchScript->>CLI: Execute `holmes version` (multiple iterations)
CLI-->>BenchScript: Timing & exit results
BenchScript-->>Workflow: master_startup.json
alt OPENROUTER_API_KEY set
Workflow->>BenchScript: Run LLM benchmark (master)
BenchScript->>CLI: Execute `holmes ask` (multiple iterations)
CLI-->>BenchScript: Timing & exit results
BenchScript-->>Workflow: master_llm.json
end
Workflow->>Comparison: Generate comparison report
Comparison->>Comparison: Load PR & master results, compute diffs, classify
Comparison-->>Workflow: comparison_report.md + status
Workflow->>Bot: Post/update PR comment with report
Workflow->>Workflow: Enforce regression threshold (>20% fails)
Workflow-->>GitHub: Job result (pass/fail)
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~30 minutes 🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 inconclusive)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
📂 Previous Runs📜 Run @ 25ad5c7 (#21563205706)✅ Results of HolmesGPT evalsAutomatically triggered by commit 25ad5c7 on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/holmes-performance-benchmarks-yKW8b' Status: Success - 38 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📜 Run @ f2108c1 (#21562585752)✅ Results of HolmesGPT evalsAutomatically triggered by commit f2108c1 on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/holmes-performance-benchmarks-yKW8b' Status: Success - 38 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📜 Run @ b4ed214 (#21514930829)✅ Results of HolmesGPT evalsAutomatically triggered by commit b4ed214 on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/holmes-performance-benchmarks-yKW8b' Status: Success - 13 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📜 Run @ 87a210a (#21329273939)✅ Results of HolmesGPT evalsAutomatically triggered by commit 87a210a on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/holmes-performance-benchmarks-yKW8b' Status: Success - 20 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📜 Run @ fb408b6 (#21321240745)✅ Results of HolmesGPT evalsAutomatically triggered by commit fb408b6 on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/holmes-performance-benchmarks-yKW8b' Status: Success - 19 test/model combinations loaded Experiments compared (30):
Comparison indicators:
✅ Results of HolmesGPT evalsAutomatically triggered by commit fde5b9a on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/holmes-performance-benchmarks-yKW8b' Status: Success - 38 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
Commands: CLI: |
|
✅ Docker image ready for
Use this tag to pull the image for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:d502f03
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:d502f03 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:d502f03
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:d502f03Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:d502f03Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:d502f03 |
🔬 CLI Performance Benchmark🟢 Startup Time (no LLM)Measures
🟡 Full CLI with LLMMeasures
PR: |
GitHub already shows when comments are updated. https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
8GB wasn't consistently clearing caches, try 16GB (2x runner RAM). https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Fix all issues with AI agents
In @.github/workflows/cli-performance.yaml:
- Around line 21-22: The workflow uses a workflow_dispatch trigger but the job
has a hard condition "if: github.event_name == 'pull_request'", which prevents
manual runs and makes the "iterations" input unusable; fix by either removing
the workflow_dispatch trigger or changing the job condition to allow manual runs
(e.g., allow github.event_name == 'workflow_dispatch' as well) and update the PR
comment steps to guard against missing context.issue.number when event is
workflow_dispatch (skip or conditionally run comment steps for non-PR runs).
Instead of running both benchmarks on the same runner and trying to clear caches (which was unreliable), use separate jobs: 1. benchmark-pr: Fresh VM, benchmarks PR branch 2. benchmark-master: Fresh VM, benchmarks master branch 3. compare: Combines results and posts PR comment Benefits: - True cold start isolation (separate VMs) - Jobs run in parallel (faster) - No cache clearing workarounds needed - Simpler, more reliable Removed: - Running banner (complexity not worth it with separate jobs) - All cache clearing logic (bytecode, venv, memory pressure, etc.) https://claude.ai/code/session_01BbRNhdcHhMSAte8mSdJqcz Signed-off-by: Claude <noreply@anthropic.com>
Summary by CodeRabbit
New Features
Chores
✏️ Tip: You can customize this high-level summary in your review settings.