Skip to content

Add evals-skip-default PR label to skip regression marker in eval runs - #1821

Merged
aantn merged 5 commits into
masterfrom
claude/add-evals-skip-default-SMTR5
Mar 21, 2026
Merged

aantn merged 5 commits into
masterfrom
claude/add-evals-skip-default-SMTR5

Conversation

@aantn

@aantn aantn commented Mar 21, 2026 •

Copy link
Copy Markdown
Collaborator

When the evals-skip-default label is added to a PR, automatic eval runs
will no longer include the regression marker by default. This allows
running all LLM tests (or only the tags specified by other evals-tag-*
labels) without being restricted to regression-tagged tests.

The label works in combination with other eval labels:

  • evals-skip-default alone: runs all LLM tests
  • evals-skip-default + evals-tag-X: runs only tag X (no regression)
  • evals-skip-default + evals-tag-X + evals-id-Y: tag X filtered by id Y

https://claude.ai/code/session_013bRMaSy1AYRR2Zg1mYBZqB
Signed-off-by: Claude noreply@anthropic.com

Summary by CodeRabbit

  • New Features

    • Add a PR label to control automated regression evaluations: it can suppress the default regression marker, let tag/id labels determine markers, or—if no eval labels are present—skip the eval run entirely.
    • Surface a run-level skip flag so subsequent steps exit early when evals are skipped, preventing result generation and comments.
  • Chores

    • Ensure the skip label is preserved in surfaced trigger metadata.

When the `evals-skip-default` label is added to a PR, automatic eval runs
will no longer include the `regression` marker by default. This allows
running all LLM tests (or only the tags specified by other evals-tag-*
labels) without being restricted to regression-tagged tests.

The label works in combination with other eval labels:
- evals-skip-default alone: runs all LLM tests
- evals-skip-default + evals-tag-X: runs only tag X (no regression)
- evals-skip-default + evals-tag-X + evals-id-Y: tag X filtered by id Y

https://claude.ai/code/session_013bRMaSy1AYRR2Zg1mYBZqB
Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review.

@coderabbitai

coderabbitai Bot commented Mar 21, 2026 •

Copy link
Copy Markdown
Contributor

Caution

Review failed

Pull request was closed or merged during review

Walkthrough

Adds evals-skip-default handling to .github/workflows/eval-regression.yaml: computes skipEval from PR labels, exports steps.eval-params.outputs.skip_eval, alters marker/filter generation to omit the default regression in certain label combinations, and short-circuits eval execution and downstream reporting when skipping is requested.

Changes

Cohort / File(s) Summary
Eval regression workflow
​.github/workflows/eval-regression.yaml
Add detection of evals-skip-default; compute skipEval and export as steps.eval-params.outputs.skip_eval; adjust markers/filter generation to produce tag-only markers when appropriate and retain evals-skip-default in surfaced labels; write should-run=false and exit early when skip_eval is true; gate post-run steps (posting results, regression checks) on steps.check-tests.outputs.should-run == 'true'.

Sequence Diagram(s)

sequenceDiagram
    participant PR as Pull Request (labels)
    participant GH as GitHub Actions Runner
    participant Params as Determine Eval Params
    participant Guard as Check if tests should run
    participant Eval as Eval jobs
    participant Reporter as Post results / Regression checks

    PR->>GH: open/update PR (labels)
    GH->>Params: read PR labels, compute markers/filter, set skip_eval
    Params-->>GH: outputs (markers, filter, skip_eval)
    GH->>Guard: evaluate skip_eval
    alt skip_eval == true
        Guard-->>GH: should-run=false (exit early)
        GH->>Reporter: skip posting/regression checks (gated)
    else skip_eval == false
        Guard-->>GH: should-run=true
        GH->>Eval: run evaluations
        Eval-->>Reporter: produce results and regression checks
    end
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested labels

evals-tag-elasticsearch

Suggested reviewers

  • arikalon1
  • Sheeproid
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: adding an evals-skip-default PR label to skip the regression marker in eval runs, which aligns with the primary purpose described in the PR objectives.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Mar 21, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ f2ad9de (#23378486697)

✅ Results of HolmesGPT evals

Automatically triggered by commit f2ad9de on branch claude/add-evals-skip-default-SMTR5 (labels: evals-tag-grafana, evals-skip-default)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 3/3 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 177_grafana_home_dashboard 12.7s 3 3 $0.1591 58,858 58,271 20,049 587 243 38,211 20,060 — —
✅ 178_grafana_search_dashboard_query 18.8s 4 5 $0.3035 80,966 79,777 21,089 1,189 445 39,090 40,687 — —
✅ 179_grafana_big_dashboard_query 20.1s 4 5 $0.1861 80,524 79,481 20,917 1,043 441 58,552 20,929 — —
Total 17.2s avg 3.7 avg 4.3 avg $0.6488 220,348 217,529 21,089 2,819 445 135,853 81,676 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit b820e8c on branch claude/add-evals-skip-default-SMTR5 (labels: evals-tag-grafana, evals-skip-default)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 3/3 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 177_grafana_home_dashboard 12.8s 3 3 $0.1604 58,913 58,276 20,055 637 299 38,210 20,066 — —
✅ 178_grafana_search_dashboard_query 19.0s 4 5 $0.1928 81,039 79,772 21,081 1,267 456 58,679 21,093 — —
✅ 179_grafana_big_dashboard_query 21.8s 5 5 $0.2201 106,218 105,070 24,145 1,148 308 80,912 24,158 — —
Total 17.9s avg 4.0 avg 4.3 avg $0.5734 246,170 243,118 24,145 3,052 456 177,801 65,317 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/add-evals-skip-default-SMTR5 -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, integration, kafka, kubernetes, leaked-information, logs, loki, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/add-evals-skip-default-SMTR5 -f markers=regression -f filter=

@github-actions

github-actions Bot commented Mar 21, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for f5af68ab (built in 7m 26s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:f5af68ab
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:f5af68ab me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:f5af68ab
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:f5af68ab
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:f5af68ab
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:f5af68ab me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:f5af68ab
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:f5af68ab

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:f5af68ab \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:f5af68ab

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:f5af68ab \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:f5af68ab

@netlify

netlify Bot commented Mar 21, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 8686a55
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69be870ba8e59c0008b534bc
😎 Deploy Preview https://deploy-preview-1821--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

Previously, the evals-skip-default label with no other evals-tag-* labels
would set markers to '' which expanded to all 244 LLM tests (marker_expr
= 'llm'). This caused eval runs to never finish.

Now when evals-skip-default is the only eval label, the eval run is
skipped entirely via a skip_eval flag. When combined with evals-tag-*
or evals-id-* labels, it still runs only those specific tags without
prepending 'regression'.

https://claude.ai/code/session_013bRMaSy1AYRR2Zg1mYBZqB
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/eval-regression.yaml:
- Around line 580-585: The final reporting steps still run even when the
workflow was intentionally skipped; update the "Post evaluation results" and
"Check test results" job steps to gate them on the should-run flag emitted by
the check-tests step by changing their if condition from always() to include
steps.check-tests.outputs.should-run == 'true' (e.g. if: ${{ always() &&
steps.check-tests.outputs.should-run == 'true' }}), referencing the step that
echoes "should-run" (check-tests) so the post-reporting only runs when
should-run is true.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 4bcf01d8-85a8-4f25-84f2-8b32e878a784

📥 Commits

Reviewing files that changed from the base of the PR and between e394b64 and 95937e0.

📒 Files selected for processing (1)
  • .github/workflows/eval-regression.yaml

Comment thread .github/workflows/eval-regression.yaml
@github-actions

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 95937e0 on branch claude/add-evals-skip-default-SMTR5 (labels: evals-skip-default)

View workflow logs

⚠️ No eval report was generated.

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/add-evals-skip-default-SMTR5 -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

(loading...)

🤖 Valid models

(loading...)


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/add-evals-skip-default-SMTR5 -f markers=regression -f filter=

…ipped

When evals-skip-default skips the eval run, the "Post evaluation results"
and "Check test results" steps still ran due to `if: always()`, posting
misleading comments. Add should-run check to these steps.

https://claude.ai/code/session_013bRMaSy1AYRR2Zg1mYBZqB
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn enabled auto-merge (squash) March 21, 2026 11:15

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
.github/workflows/eval-regression.yaml (1)

513-525: ⚠️ Potential issue | 🟡 Minor

Update the shared rerun footer for the new label behavior.

Line 513 starts surfacing evals-skip-default in automatic-run metadata, but the footer builder in .github/scripts/eval-comment-helpers.js still hardcodes markers=regression and documents evals-tag-* as running alongside regression, with no mention of evals-skip-default. PR comments will therefore tell users to rerun the wrong thing.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/eval-regression.yaml around lines 513 - 525, The rerun
footer still hardcodes "markers=regression" and references evals-tag-* behavior
while the code now surfaces evals-skip-default and builds evalLabels (see
evalLabels, displayModelSafe, displayLabelSafe, triggerSource); update the
footer builder logic (in the footer creation that uses evalLabels and
triggerSource) to dynamically reflect the actual labels and special flags
including evals-skip-default instead of always using markers=regression, e.g.,
generate the rerun query from evalLabels and include evals-skip-default when
present so the PR comment instructs users to rerun the correct set of tests.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/eval-regression.yaml:
- Around line 474-493: The branch that handles evals-skip-default currently sets
skipEval = true which incorrectly aborts collection; change that branch to clear
the default marker instead by setting markers = '' (do not set skipEval) so when
skipDefault is true and no idLabels/tagLabels are present the run broadens to
all LLM tests; ensure the logic around skipDefault, idLabels, tagLabels, markers
and filter retains existing behavior for combinations (i.e., when tagLabels or
idLabels exist still apply the tag/id logic) and remove the short-circuit that
prevents the job from reaching collection.

---

Outside diff comments:
In @.github/workflows/eval-regression.yaml:
- Around line 513-525: The rerun footer still hardcodes "markers=regression" and
references evals-tag-* behavior while the code now surfaces evals-skip-default
and builds evalLabels (see evalLabels, displayModelSafe, displayLabelSafe,
triggerSource); update the footer builder logic (in the footer creation that
uses evalLabels and triggerSource) to dynamically reflect the actual labels and
special flags including evals-skip-default instead of always using
markers=regression, e.g., generate the rerun query from evalLabels and include
evals-skip-default when present so the PR comment instructs users to rerun the
correct set of tests.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 8a15eb12-543d-445b-a116-f705309bb2f9

📥 Commits

Reviewing files that changed from the base of the PR and between 95937e0 and f2ad9de.

📒 Files selected for processing (1)
  • .github/workflows/eval-regression.yaml

Comment thread .github/workflows/eval-regression.yaml Outdated
claude added 2 commits March 21, 2026 11:53
…ping

When evals-skip-default is the only eval label (no evals-tag-* or
evals-id-*), set markers to '' (all LLM tests) instead of aborting the
run. Remove the skipEval variable, its output, and the check-tests
short-circuit. Revert reporting step guards to their original conditions.

https://claude.ai/code/session_013bRMaSy1AYRR2Zg1mYBZqB
Signed-off-by: Claude <noreply@anthropic.com>
When evals-skip-default is the only eval label, skip the run (run
nothing) instead of broadening to all 244 LLM tests. Re-adds skipEval
flag, check-tests gate, and reporting step guards.

https://claude.ai/code/session_013bRMaSy1AYRR2Zg1mYBZqB
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn merged commit 2b470f9 into master Mar 21, 2026
16 of 19 checks passed
@aantn
aantn deleted the claude/add-evals-skip-default-SMTR5 branch March 21, 2026 11:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants