Repository navigation
Conversation
- Remove 'regression' as default marker for /eval comments and workflow_dispatch triggers - they now run all LLM tests by default - Keep 'regression' as default only for automatic triggers (PR/push) - Add details section with list of valid markers and example test names - Update marker_expr to handle empty markers (just 'llm' instead of 'llm and ()') Signed-off-by: Claude <noreply@anthropic.com>
- Add test preview step that runs pytest --collect-only to show which tests will run before actually running them - Update initial comment with test count and expandable test list - Add warning that manual re-runs have no default markers and will run all LLM tests (~100+) which can take 1+ hours - Update example /eval command to include markers: regression - Update markers description to emphasize no default Signed-off-by: Claude <noreply@anthropic.com>
|
✅ Docker image ready for
Use this tag to pull the image for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:5bba825
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:5bba825 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:5bba825
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:5bba825Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:5bba825Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:5bba825 |
WalkthroughThis change updates the eval-regression GitHub Actions workflow to track who triggered runs, change marker defaulting behavior based on trigger type, add pre-collection of evals (test_count/test_preview), and surface richer PR comments and outputs (including triggered_by and marker_expr). Changes
Sequence DiagramsequenceDiagram
participant Trigger as Trigger (comment / dispatch)
participant WF as Workflow
participant Parser as eval-params step
participant Collector as Collect evals step
participant Commenter as Update PR Comment
participant Executor as Run tests step
participant Notifier as Notify user step
Trigger->>WF: start (manual or automatic)
WF->>Parser: parse inputs
Parser->>Parser: set triggered_by
Parser->>Parser: apply conditional marker default (auto only)
Parser->>Parser: derive marker_expr
rect rgb(220,240,255)
Parser->>Collector: request tests (marker_expr, filter)
Collector->>Collector: collect tests
Collector->>Commenter: return test_count & test_preview
end
Commenter->>Trigger: post detailed PR comment (trigger, model, markers, filter, iterations, test_count, preview, run URL)
Parser->>Executor: pass eval params
Executor->>Executor: run tests
rect rgb(240,220,255)
Executor->>Notifier: send results
Notifier->>Trigger: post completion comment (manual runs include triggered_by)
end
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested reviewers
Pre-merge checks✅ Passed checks (3 passed)
📜 Recent review detailsConfiguration used: Organization UI Review profile: CHILL Plan: Pro 📒 Files selected for processing (1)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
🔇 Additional comments (11)
Comment |
Results of HolmesGPT evalsDuration: 4m 6s | View workflow logs Results of HolmesGPT evals
Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 📋 Valid eval names and markersValid markers: Example test names (use with Full list: Run |
When a user triggers evals manually via /eval comment or workflow_dispatch, they now receive a notification when the run completes. The notification: - @mentions the user who triggered the eval - Shows success or regression count status - Points to the updated results comment above This ensures users get a GitHub notification instead of having to poll the PR for the updated comment. Signed-off-by: Claude <noreply@anthropic.com>
Results of HolmesGPT evalsDuration: 4m 27s | View workflow logs Results of HolmesGPT evals
Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 📋 Valid eval names and markersValid markers: Example test names (use with Full list: Run |
- Rename step names and summary text from "Test preview" to "Evals to run" - Remove the 20 test limit to show all evals that will run Signed-off-by: Claude <noreply@anthropic.com>
Results of HolmesGPT evalsDuration: N/A | View workflow logs 🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 📋 Valid eval names and markersValid markers: Example test names (use with Full list: Run |
There was a problem hiding this comment.
Actionable comments posted: 0
🧹 Nitpick comments (1)
.github/workflows/eval-regression.yaml (1)
239-259: Consider handling pytest collection failures gracefully.If
pytest --collect-onlyfails (e.g., due to a syntax error in a test file), the step will fail and block the workflow. The actual test run step (line 333) uses|| trueto continue on failure.Consider adding similar error handling:
🔎 Suggested improvement
# Collect test names - TEST_LIST=$(poetry run pytest "${PYTEST_ARGS[@]}" 2>/dev/null | grep -E "^tests/llm/" || echo "") + TEST_LIST=$(poetry run pytest "${PYTEST_ARGS[@]}" 2>/dev/null | grep -E "^tests/llm/" || true) TEST_COUNT=$(echo "$TEST_LIST" | grep -c "^tests/llm/" || echo "0")
📜 Review details
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (1)
.github/workflows/eval-regression.yaml
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
- GitHub Check: build
- GitHub Check: build (3.11)
- GitHub Check: build (3.10)
- GitHub Check: build (3.12)
🔇 Additional comments (6)
.github/workflows/eval-regression.yaml (6)
18-21: LGTM!The input description and default changes correctly reflect the new behavior where manual triggers (workflow_dispatch, /eval) will run all LLM tests unless markers are explicitly specified. The warning in the PR comment (lines 384-385) appropriately alerts users about this behavior.
123-157: LGTM!The
triggeredBytracking is correctly implemented for each trigger type. Usingcontext.payload.comment.user.loginfor issue comments andcontext.actorfor workflow_dispatch ensures the right user is notified upon completion.
159-173: LGTM!The marker defaulting logic correctly differentiates between automatic and manual triggers. The
marker_exproutput properly handles both cases: usingllm and (markers)when markers are specified, or justllmto run all LLM tests when empty.
261-303: LGTM!The comment update step is well-structured with proper fallbacks for missing outputs. The collapsible details section for test preview is a nice UX improvement. The step conditions correctly ensure it only runs when all prerequisites are met.
384-405: LGTM!The updated help text provides clear guidance about the behavior change. The warning about manual re-runs having no default markers is prominent, and the expanded marker/test name examples help users construct appropriate
/evalcommands.
420-440: LGTM!The notification step appropriately pings the user who triggered the manual eval, with a clear status message and pointer to the results. The conditions correctly ensure this only runs for manual triggers where the user is known.
- Remove "Results will appear here when complete." text from both initial and running status comments since it's confusing - For manual triggers, don't create a new comment if comment_id is missing - the initial comment should always exist for manual runs Signed-off-by: Claude <noreply@anthropic.com>
Results of HolmesGPT evalsDuration: 3m 55s | View workflow logs Results of HolmesGPT evals
Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 📋 Valid eval names and markersValid markers: Example test names (use with Full list: Run |
Summary by CodeRabbit
New Features
Changes
✏️ Tip: You can customize this high-level summary in your review settings.