Skip to content

Only run CLI Performance Benchmark on source changes - #1601

Merged
aantn merged 2 commits into
masterfrom
claude/cli-benchmark-on-changes-FRj3B
Feb 20, 2026
Merged

aantn merged 2 commits into
masterfrom
claude/cli-benchmark-on-changes-FRj3B

Conversation

@aantn

@aantn aantn commented Feb 20, 2026 •

Copy link
Copy Markdown
Collaborator

Switch from paths-ignore to paths so the benchmark workflow only
triggers when files that could affect CLI performance are modified:
holmes/**, pyproject.toml, poetry.lock, the benchmark script, or the
workflow itself.

https://claude.ai/code/session_01XrNrjxLPQ4viTbe9bMVKFx
Signed-off-by: Claude noreply@anthropic.com

Summary by CodeRabbit

  • Chores
    • Optimized performance benchmark CI/CD workflow to trigger only on relevant code changes, reducing unnecessary workflow executions.

Switch from paths-ignore to paths so the benchmark workflow only
triggers when files that could affect CLI performance are modified:
holmes/**, pyproject.toml, poetry.lock, the benchmark script, or the
workflow itself.

https://claude.ai/code/session_01XrNrjxLPQ4viTbe9bMVKFx
Signed-off-by: Claude <noreply@anthropic.com>
@netlify

netlify Bot commented Feb 20, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit ab9b802
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69986dc8e1982600088396f7
😎 Deploy Preview https://deploy-preview-1601--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Feb 20, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 0aae1a3 (#22218924963)

✅ Results of HolmesGPT evals

Automatically triggered by commit 0aae1a3 on branch claude/cli-benchmark-on-changes-FRj3B

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 34.4s 5 12 $0.2458
✅ 101_loki_historical_logs_pod_deleted 45.3s 5 10 $0.2621
✅ 111_pod_names_contain_service 35.9s 5 11 $0.2363
✅ 112_find_pvcs_by_uuid 36.6s 6 8 $0.2744
✅ 12_job_crashing 34.4s 5 11 $0.2434
✅ 176_network_policy_blocking_traffic_no_runbooks 49.7s 7 17 $0.3023
✅ 24_misconfigured_pvc 42.3s 8 14 $0.2721
✅ 43_current_datetime_from_prompt 5.0s 1 — $0.1115
✅ 61_exact_match_counting 12.9s 3 2 $0.1492
Total 33.0s avg 5.0 avg 10.6 avg $2.0973

✅ Results of HolmesGPT evals

Automatically triggered by commit ab9b802 on branch claude/cli-benchmark-on-changes-FRj3B

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.0s 5 11 $0.2332
✅ 101_loki_historical_logs_pod_deleted 33.5s 4 8 $0.2220
✅ 111_pod_names_contain_service 35.7s 5 12 $0.2408
✅ 112_find_pvcs_by_uuid 30.9s 5 7 $0.2268
✅ 12_job_crashing 34.1s 5 11 $0.2453
✅ 176_network_policy_blocking_traffic_no_runbooks 45.9s 6 17 $0.2874
✅ 24_misconfigured_pvc 38.8s 6 14 $0.2539
✅ 43_current_datetime_from_prompt 5.2s 1 — $0.0118
✅ 61_exact_match_counting 15.9s 3 2 $0.1553
Total 30.2s avg 4.4 avg 10.2 avg $1.8764
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/cli-benchmark-on-changes-FRj3B -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals in automatic regression runs:

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID

Examples: evals-tag-easy, evals-id-09_crashpod

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/cli-benchmark-on-changes-FRj3B -f markers=regression -f filter=

@github-actions

github-actions Bot commented Feb 20, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for 916b4773 (built in 1m 4s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:916b4773
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:916b4773 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:916b4773
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:916b4773

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:916b4773

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:916b4773

@coderabbitai

coderabbitai Bot commented Feb 20, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Modified GitHub Actions workflow path filtering in .github/workflows/cli-performance.yaml from an ignore-list approach (excluding docs, markdown, and mkdocs config) to an explicit allow-list approach that triggers only on changes to holmes/**, pyproject.toml, poetry.lock, scripts/cli_performance_benchmark.py, or the workflow itself.

Changes

Cohort / File(s) Summary
Workflow Configuration
.github/workflows/cli-performance.yaml
Changed path-based trigger from deny-list to allow-list, narrowing CLI performance workflow activation to specific files relevant to Holmes CLI and its dependencies.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~5 minutes

Possibly related PRs

Suggested reviewers

  • pavangudiwada
  • moshemorad
  • arikalon1
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: converting the workflow from an ignore-list to an explicit include-list that only runs on source code changes.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Feb 20, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 10.65s 11.03s -3.5%
Warm Mean 4.59s 4.80s -4.4%
Warm Min 4.50s 4.71s
Warm Max 4.79s 4.85s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 17.69s 30.02s -41.1%
Warm Mean 7.25s 7.70s -5.9%
Warm Min 6.85s 7.21s
Warm Max 7.64s 7.97s

PR: 916b4773 | Master: d3996c5c | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/cli-performance.yaml:
- Around line 9-14: The workflow paths allowlist is missing two top-level
CLI-related files; update the paths array in
.github/workflows/cli-performance.yaml to include 'holmes_cli.py' and
'tempo_cli.py' alongside the existing entries (e.g., keep 'holmes/**',
'pyproject.toml', 'poetry.lock', 'scripts/cli_performance_benchmark.py', and
'.github/workflows/cli-performance.yaml') so changes to holmes_cli.py and
tempo_cli.py will trigger the workflow; if either file is truly unrelated to the
CLI startup, you may omit it, otherwise add both to be safe.

Comment thread .github/workflows/cli-performance.yaml
@aantn
aantn enabled auto-merge (squash) February 20, 2026 10:32
@aantn
aantn merged commit 2b640e3 into master Feb 20, 2026
16 of 19 checks passed
@aantn
aantn deleted the claude/cli-benchmark-on-changes-FRj3B branch February 20, 2026 14:26
arikalon1 pushed a commit that referenced this pull request Feb 21, 2026
Switch from paths-ignore to paths so the benchmark workflow only
triggers when files that could affect CLI performance are modified:
holmes/**, pyproject.toml, poetry.lock, the benchmark script, or the
workflow itself.

https://claude.ai/code/session_01XrNrjxLPQ4viTbe9bMVKFx
Signed-off-by: Claude <noreply@anthropic.com>

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Optimized performance benchmark CI/CD workflow to trigger only on
relevant code changes, reducing unnecessary workflow executions.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Arik Alon <alon.arik@gmail.com>
moshemorad pushed a commit that referenced this pull request Feb 22, 2026
Switch from paths-ignore to paths so the benchmark workflow only
triggers when files that could affect CLI performance are modified:
holmes/**, pyproject.toml, poetry.lock, the benchmark script, or the
workflow itself.

https://claude.ai/code/session_01XrNrjxLPQ4viTbe9bMVKFx
Signed-off-by: Claude <noreply@anthropic.com>

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Optimized performance benchmark CI/CD workflow to trigger only on
relevant code changes, reducing unnecessary workflow executions.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Mohse Morad <moshemorad12340@gmail.com>
moshemorad pushed a commit that referenced this pull request Feb 22, 2026
Switch from paths-ignore to paths so the benchmark workflow only
triggers when files that could affect CLI performance are modified:
holmes/**, pyproject.toml, poetry.lock, the benchmark script, or the
workflow itself.

https://claude.ai/code/session_01XrNrjxLPQ4viTbe9bMVKFx
Signed-off-by: Claude <noreply@anthropic.com>

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Optimized performance benchmark CI/CD workflow to trigger only on
relevant code changes, reducing unnecessary workflow executions.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Mohse Morad <moshemorad12340@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants