Repository navigation
Add evals-skip-default PR label to skip regression marker in eval runs - #1821
Conversation
When the `evals-skip-default` label is added to a PR, automatic eval runs will no longer include the `regression` marker by default. This allows running all LLM tests (or only the tags specified by other evals-tag-* labels) without being restricted to regression-tagged tests. The label works in combination with other eval labels: - evals-skip-default alone: runs all LLM tests - evals-skip-default + evals-tag-X: runs only tag X (no regression) - evals-skip-default + evals-tag-X + evals-id-Y: tag X filtered by id Y https://claude.ai/code/session_013bRMaSy1AYRR2Zg1mYBZqB Signed-off-by: Claude <noreply@anthropic.com>
|
Caution Review failedPull request was closed or merged during review WalkthroughAdds Changes
Sequence Diagram(s)sequenceDiagram
participant PR as Pull Request (labels)
participant GH as GitHub Actions Runner
participant Params as Determine Eval Params
participant Guard as Check if tests should run
participant Eval as Eval jobs
participant Reporter as Post results / Regression checks
PR->>GH: open/update PR (labels)
GH->>Params: read PR labels, compute markers/filter, set skip_eval
Params-->>GH: outputs (markers, filter, skip_eval)
GH->>Guard: evaluate skip_eval
alt skip_eval == true
Guard-->>GH: should-run=false (exit early)
GH->>Reporter: skip posting/regression checks (gated)
else skip_eval == false
Guard-->>GH: should-run=true
GH->>Eval: run evaluations
Eval-->>Reporter: produce results and regression checks
end
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested labels
Suggested reviewers
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
📂 Previous Runs📜 Run @ f2ad9de (#23378486697)✅ Results of HolmesGPT evalsAutomatically triggered by commit f2ad9de on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 73 test/model combinations loaded Benchmark experiment:
No benchmark data available for comparison. Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
✅ Results of HolmesGPT evalsAutomatically triggered by commit b820e8c on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 73 test/model combinations loaded Benchmark experiment:
No benchmark data available for comparison. Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" Option 3: Add PR labels to include extra evals (applies to both automatic runs and
Examples: 🏷️ Valid tags
🤖 Valid models
Commands: CLI: |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:f5af68ab
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:f5af68ab me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:f5af68ab
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:f5af68ab
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:f5af68ab
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:f5af68ab me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:f5af68ab
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:f5af68abPatch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:f5af68ab \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:f5af68abRobusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:f5af68ab \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:f5af68ab |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
Previously, the evals-skip-default label with no other evals-tag-* labels would set markers to '' which expanded to all 244 LLM tests (marker_expr = 'llm'). This caused eval runs to never finish. Now when evals-skip-default is the only eval label, the eval run is skipped entirely via a skip_eval flag. When combined with evals-tag-* or evals-id-* labels, it still runs only those specific tags without prepending 'regression'. https://claude.ai/code/session_013bRMaSy1AYRR2Zg1mYBZqB Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In @.github/workflows/eval-regression.yaml:
- Around line 580-585: The final reporting steps still run even when the
workflow was intentionally skipped; update the "Post evaluation results" and
"Check test results" job steps to gate them on the should-run flag emitted by
the check-tests step by changing their if condition from always() to include
steps.check-tests.outputs.should-run == 'true' (e.g. if: ${{ always() &&
steps.check-tests.outputs.should-run == 'true' }}), referencing the step that
echoes "should-run" (check-tests) so the post-reporting only runs when
should-run is true.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 4bcf01d8-85a8-4f25-84f2-8b32e878a784
📒 Files selected for processing (1)
.github/workflows/eval-regression.yaml
✅ Results of HolmesGPT evalsAutomatically triggered by commit 95937e0 on branch 📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" Option 3: Add PR labels to include extra evals (applies to both automatic runs and
Examples: 🏷️ Valid tags(loading...) 🤖 Valid models(loading...) Commands: CLI: |
…ipped When evals-skip-default skips the eval run, the "Post evaluation results" and "Check test results" steps still ran due to `if: always()`, posting misleading comments. Add should-run check to these steps. https://claude.ai/code/session_013bRMaSy1AYRR2Zg1mYBZqB Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
.github/workflows/eval-regression.yaml (1)
513-525:⚠️ Potential issue | 🟡 MinorUpdate the shared rerun footer for the new label behavior.
Line 513 starts surfacing
evals-skip-defaultin automatic-run metadata, but the footer builder in.github/scripts/eval-comment-helpers.jsstill hardcodesmarkers=regressionand documentsevals-tag-*as running alongside regression, with no mention ofevals-skip-default. PR comments will therefore tell users to rerun the wrong thing.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In @.github/workflows/eval-regression.yaml around lines 513 - 525, The rerun footer still hardcodes "markers=regression" and references evals-tag-* behavior while the code now surfaces evals-skip-default and builds evalLabels (see evalLabels, displayModelSafe, displayLabelSafe, triggerSource); update the footer builder logic (in the footer creation that uses evalLabels and triggerSource) to dynamically reflect the actual labels and special flags including evals-skip-default instead of always using markers=regression, e.g., generate the rerun query from evalLabels and include evals-skip-default when present so the PR comment instructs users to rerun the correct set of tests.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In @.github/workflows/eval-regression.yaml:
- Around line 474-493: The branch that handles evals-skip-default currently sets
skipEval = true which incorrectly aborts collection; change that branch to clear
the default marker instead by setting markers = '' (do not set skipEval) so when
skipDefault is true and no idLabels/tagLabels are present the run broadens to
all LLM tests; ensure the logic around skipDefault, idLabels, tagLabels, markers
and filter retains existing behavior for combinations (i.e., when tagLabels or
idLabels exist still apply the tag/id logic) and remove the short-circuit that
prevents the job from reaching collection.
---
Outside diff comments:
In @.github/workflows/eval-regression.yaml:
- Around line 513-525: The rerun footer still hardcodes "markers=regression" and
references evals-tag-* behavior while the code now surfaces evals-skip-default
and builds evalLabels (see evalLabels, displayModelSafe, displayLabelSafe,
triggerSource); update the footer builder logic (in the footer creation that
uses evalLabels and triggerSource) to dynamically reflect the actual labels and
special flags including evals-skip-default instead of always using
markers=regression, e.g., generate the rerun query from evalLabels and include
evals-skip-default when present so the PR comment instructs users to rerun the
correct set of tests.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 8a15eb12-543d-445b-a116-f705309bb2f9
📒 Files selected for processing (1)
.github/workflows/eval-regression.yaml
…ping When evals-skip-default is the only eval label (no evals-tag-* or evals-id-*), set markers to '' (all LLM tests) instead of aborting the run. Remove the skipEval variable, its output, and the check-tests short-circuit. Revert reporting step guards to their original conditions. https://claude.ai/code/session_013bRMaSy1AYRR2Zg1mYBZqB Signed-off-by: Claude <noreply@anthropic.com>
When evals-skip-default is the only eval label, skip the run (run nothing) instead of broadening to all 244 LLM tests. Re-adds skipEval flag, check-tests gate, and reporting step guards. https://claude.ai/code/session_013bRMaSy1AYRR2Zg1mYBZqB Signed-off-by: Claude <noreply@anthropic.com>
When the
evals-skip-defaultlabel is added to a PR, automatic eval runswill no longer include the
regressionmarker by default. This allowsrunning all LLM tests (or only the tags specified by other evals-tag-*
labels) without being restricted to regression-tagged tests.
The label works in combination with other eval labels:
https://claude.ai/code/session_013bRMaSy1AYRR2Zg1mYBZqB
Signed-off-by: Claude noreply@anthropic.com
Summary by CodeRabbit
New Features
Chores