Skip to content

Skip CI workflows for documentation-only changes - #1588

Closed
aantn wants to merge 2 commits into
masterfrom
claude/skip-evals-docs-only-FuMiy
Closed

aantn wants to merge 2 commits into
masterfrom
claude/skip-evals-docs-only-FuMiy

Conversation

@aantn

@aantn aantn commented Feb 19, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR adds paths-ignore filters to three GitHub Actions workflows to prevent unnecessary CI runs when only documentation files are modified.

Key Changes

  • cli-performance.yaml: Added paths-ignore filter to skip workflow on docs changes
  • docker-dev-images.yaml: Added paths-ignore filter to skip workflow on docs changes
  • eval-regression.yaml: Added paths-ignore filter to skip workflow on docs changes

All three workflows now ignore changes to:

  • docs/** directory
  • **/*.md markdown files
  • mkdocs.yml configuration file

Benefits

  • Reduces unnecessary CI resource consumption for documentation-only PRs
  • Faster feedback loop for contributors making doc updates
  • Maintains CI runs for actual code changes that require testing

https://claude.ai/code/session_01JGwBxWqzQyD3Wwd5K4HtV8

Summary by CodeRabbit

  • Chores
    • CI/CD workflows now skip benchmarks and builds when only documentation files are changed.
    • Added code-change detection to conditionally gate performance benchmarking, Docker image builds, and regression evaluations.
    • Performance and evaluation jobs now run only when code changes are detected, improving pipeline efficiency.

Add paths-ignore filters to eval-regression, cli-performance, and
docker-dev-images workflows so they don't run when a PR only changes
documentation files (docs/**, *.md, mkdocs.yml). Other triggers like
workflow_dispatch, issue_comment, push to master, and schedule are
unaffected since GitHub Actions paths filters only apply to push and
pull_request events.

https://claude.ai/code/session_01JGwBxWqzQyD3Wwd5K4HtV8
Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Feb 19, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 5f5d156 (#22171587015)

✅ Results of HolmesGPT evals

Automatically triggered by commit 5f5d156 on branch claude/skip-evals-docs-only-FuMiy

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.7s 5 12 $0.2442
✅ 101_loki_historical_logs_pod_deleted 35.1s 5 10 $0.2340
✅ 111_pod_names_contain_service 32.5s 5 11 $0.2349
✅ 112_find_pvcs_by_uuid 35.8s 7 9 $0.2809
✅ 12_job_crashing 40.1s 6 14 $0.2762
✅ 176_network_policy_blocking_traffic_no_runbooks 40.1s 6 15 $0.2793
✅ 24_misconfigured_pvc 35.5s 6 14 $0.2515
✅ 43_current_datetime_from_prompt 5.7s 1 — $0.0118
✅ 61_exact_match_counting 15.8s 4 4 $0.1668
Total 30.4s avg 5.0 avg 11.1 avg $1.9796

✅ Results of HolmesGPT evals

Automatically triggered by commit 895f57f on branch claude/skip-evals-docs-only-FuMiy

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.3s 5 12 $0.2453
✅ 101_loki_historical_logs_pod_deleted 53.9s 8 12 $0.3199
✅ 111_pod_names_contain_service 34.3s 5 12 $0.2405
✅ 112_find_pvcs_by_uuid 30.8s 6 7 $0.2365
✅ 12_job_crashing 34.6s 6 12 $0.2546
✅ 176_network_policy_blocking_traffic_no_runbooks 47.9s 8 17 $0.3198
✅ 24_misconfigured_pvc 31.1s 5 12 $0.2316
✅ 43_current_datetime_from_prompt 5.0s 1 — $0.0119
✅ 61_exact_match_counting 12.6s 3 2 $0.1490
Total 31.4s avg 5.2 avg 10.8 avg $2.0092
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/skip-evals-docs-only-FuMiy -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/skip-evals-docs-only-FuMiy -f markers=regression -f filter=

@github-actions

github-actions Bot commented Feb 19, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for e61425b (built in 55s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e61425b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:e61425b me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e61425b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:e61425b

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:e61425b

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:e61425b

@netlify

netlify Bot commented Feb 19, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 895f57f
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6996b368b99dd700087b68cb
😎 Deploy Preview https://deploy-preview-1588--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Feb 19, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 10.33s 10.44s -1.0%
Warm Mean 4.88s 5.17s -5.6%
Warm Min 4.78s 5.04s
Warm Max 5.02s 5.29s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 32.32s 26.55s +21.8%
Warm Mean 7.37s 8.02s -8.2%
Warm Min 6.80s 7.82s
Warm Max 7.65s 8.30s

PR: e61425b5 | Master: 217bba1d | Iterations: 5

paths-ignore prevents the workflow from triggering entirely, which causes
required status checks to stay in "expected" state and block PR merges.

Instead, always trigger the workflow but add a lightweight check-changes
job using dorny/paths-filter that detects docs-only PRs. Expensive jobs
(evals, benchmarks, docker build) depend on it and get cleanly "skipped"
when only docs/md/mkdocs.yml files changed - which satisfies required
status checks.

For eval-regression.yaml, check-changes only runs on pull_request events.
Non-PR triggers (push, workflow_dispatch, issue_comment) are unaffected
because the output defaults to empty (not 'false'), so llm_evals still
runs.

https://claude.ai/code/session_01JGwBxWqzQyD3Wwd5K4HtV8
Signed-off-by: Claude <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Feb 19, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

The PR adds a new check-changes job using dorny/paths-filter to detect code changes across three GitHub Actions workflows, gating subsequent benchmark, build, and evaluation jobs to skip when only documentation changes are present.

Changes

Cohort / File(s) Summary
Workflow Code Change Detection
.github/workflows/cli-performance.yaml, .github/workflows/docker-dev-images.yaml, .github/workflows/eval-regression.yaml
Added check-changes job using dorny/paths-filter to detect code changes (excluding docs, Markdown, and config files). Gated dependent jobs—benchmark-pr, benchmark-master, compare in cli-performance; build in docker-dev-images; llm_evals in eval-regression—to skip execution when only documentation changes are detected.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

Suggested reviewers

  • arikalon1
  • Sheeproid
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The pull request title accurately describes the main objective: skipping CI workflows when only documentation changes occur, which is the core purpose of the changeset.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
.github/workflows/eval-regression.yaml (1)

139-150: Well-crafted gating condition — the != 'false' check is key.

The use of != 'false' (rather than == 'true') is the correct approach here: when check-changes is skipped for non-PR events, its output is empty, and '' != 'false' evaluates to true, allowing push/workflow_dispatch/issue_comment triggers to proceed unimpeded. Combined with !cancelled() to override the default success() evaluation on needs, this correctly skips evals only for docs-only PRs.

Consider adding a brief inline comment (e.g., # != 'false' (not == 'true') so non-PR triggers pass through when check-changes is skipped) for future maintainers, since this pattern is easy to inadvertently "fix" into == 'true'.

💡 Suggested inline comment
   if: |
     !cancelled() &&
+    # Use != 'false' (not == 'true') so that non-PR triggers pass through
+    # when check-changes is skipped and output is empty
     needs.check-changes.outputs.has_code_changes != 'false' &&
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/eval-regression.yaml around lines 139 - 150, Add a short
inline comment explaining why the gating uses
"needs.check-changes.outputs.has_code_changes != 'false'" (instead of "==
'true'") next to the condition in the llm_evals job and mention the role of
"!cancelled()" in overriding default needs evaluation; locate the condition that
contains "!cancelled() && needs.check-changes.outputs.has_code_changes !=
'false' && (...)" and insert a one-line comment such as "# use != 'false' (not
== 'true') so non-PR triggers pass when check-changes is skipped" to prevent
future accidental changes.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In @.github/workflows/eval-regression.yaml:
- Around line 139-150: Add a short inline comment explaining why the gating uses
"needs.check-changes.outputs.has_code_changes != 'false'" (instead of "==
'true'") next to the condition in the llm_evals job and mention the role of
"!cancelled()" in overriding default needs evaluation; locate the condition that
contains "!cancelled() && needs.check-changes.outputs.has_code_changes !=
'false' && (...)" and insert a one-line comment such as "# use != 'false' (not
== 'true') so non-PR triggers pass when check-changes is skipped" to prevent
future accidental changes.

@aantn aantn closed this Feb 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants