Repository navigation
Minor tweaks to evals - #1266
Minor tweaks to evals #1266
Conversation
- Remove 'regression' as default marker for /eval comments and workflow_dispatch triggers - they now run all LLM tests by default - Keep 'regression' as default only for automatic triggers (PR/push) - Add details section with list of valid markers and example test names - Update marker_expr to handle empty markers (just 'llm' instead of 'llm and ()') Signed-off-by: Claude <noreply@anthropic.com>
- Add test preview step that runs pytest --collect-only to show which tests will run before actually running them - Update initial comment with test count and expandable test list - Add warning that manual re-runs have no default markers and will run all LLM tests (~100+) which can take 1+ hours - Update example /eval command to include markers: regression - Update markers description to emphasize no default Signed-off-by: Claude <noreply@anthropic.com>
When a user triggers evals manually via /eval comment or workflow_dispatch, they now receive a notification when the run completes. The notification: - @mentions the user who triggered the eval - Shows success or regression count status - Points to the updated results comment above This ensures users get a GitHub notification instead of having to poll the PR for the updated comment. Signed-off-by: Claude <noreply@anthropic.com>
- Rename step names and summary text from "Test preview" to "Evals to run" - Remove the 20 test limit to show all evals that will run Signed-off-by: Claude <noreply@anthropic.com>
- Remove "Results will appear here when complete." text from both initial and running status comments since it's confusing - For manual triggers, don't create a new comment if comment_id is missing - the initial comment should always exist for manual runs Signed-off-by: Claude <noreply@anthropic.com>
- Split into two separate <details> sections: Valid markers and Valid eval names - List each marker and eval name on its own line - Include the complete list of all 140+ ask_holmes evals and 17 investigate evals Signed-off-by: Claude <noreply@anthropic.com>
grep -c returns exit code 1 when no matches found, even though it outputs "0". Using || echo "0" would result in "0\n0" being captured. Changed to || true which prevents failure without adding extra output. Signed-off-by: Claude <noreply@anthropic.com>
- For PR triggers: show "PR #123 (abc1234)" - For push triggers: show "push to branch-name (abc1234)" - Include trigger info in both initial and updated comments Signed-off-by: Claude <noreply@anthropic.com>
- Add progress checklist showing completion status of each step - Update comment after each major step (HolmesGPT setup, collect evals, KIND ready) - Reorder steps: HolmesGPT env -> Collect evals -> KIND cluster for optimization - Show completed checklist in final results comment Signed-off-by: Claude <noreply@anthropic.com>
When no model is specified in /eval comment, MODEL env var was set to empty string instead of being unset. This caused get_models() to return an empty list (since it only uses DEFAULT_MODEL when MODEL is unset, not when it's empty string), resulting in 0 tests being collected. Fix: Unset MODEL in bash if it's empty so Python uses DEFAULT_MODEL. Signed-off-by: Claude <noreply@anthropic.com>
Replace hardcoded lists of valid markers and eval names with dynamic collection at runtime: - Markers: extracted from pyproject.toml [tool.pytest.ini_options] section - Eval names: listed from tests/llm/fixtures/test_ask_holmes/ and test_investigate/ This ensures the comment always shows the current list of available markers and tests without manual updates when new tests are added. Signed-off-by: Claude <noreply@anthropic.com>
Refactor the 5 comment-generating steps to use consistent helper functions: - renderProgress(): Progress checklist renderer - renderParamsTable(): Parameters table for manual runs - buildBody(): Main comment body builder with icon/title customization - buildRerunFooter(): Re-run instructions for automatic runs (results only) This ensures consistent styling between manual and automatic eval runs, and between initial/progress/results comments. The helpers are defined inline in each step (GitHub Actions limitation) but follow the same pattern. Signed-off-by: Claude <noreply@anthropic.com>
- Create .github/scripts/eval-comment-helpers.js with reusable functions - Update all 5 workflow steps to require() the helpers instead of copy-paste - Remove redundant emoji from progress checklist items (⏳ and ✅ were redundant since checkbox state already indicates completion) Signed-off-by: Claude <noreply@anthropic.com>
- Change PR trigger source from "PR #N (sha)" to just "commit sha" (PR number is redundant since comment is on the PR) - Remove eyes reaction when workflow completes, leaving only hooray Signed-off-by: Claude <noreply@anthropic.com>
- Add buildParams() helper to centralize type conversions and defaults - Update all 5 workflow steps to use buildParams() for cleaner code - All transformations (=== 'true', || 'default', parseInt) in one place Signed-off-by: Claude <noreply@anthropic.com>
- Output base_params JSON from eval-params step containing common params - Each step now spreads base_params and only adds step-specific extras - Reduces ~7 lines of repeated parameter passing per step Signed-off-by: Claude <noreply@anthropic.com>
When triggering via GitHub Actions UI without specifying a PR number, automatically find the open PR associated with the selected branch. This makes it easier to run evals - just select the branch and go. Signed-off-by: Claude <noreply@anthropic.com>
- Hide Progress section once run is complete (pass null to buildBody) - Consolidate legend into single collapsible <details> section - Change "Regression" to "Failure" throughout for clarity Signed-off-by: Claude <noreply@anthropic.com>
- Add new Legend <details> section with table explaining each status icon - Restore original 3 separate <details> sections (Re-run, Markers, Eval names) Signed-off-by: Claude <noreply@anthropic.com>
…atic runs Signed-off-by: Claude <noreply@anthropic.com>
Keep branch improvements: - Legend section in its own details tag - Separate details sections for Re-run, Markers, and Eval names - Footer shown for both manual and automatic runs
✅ Results of HolmesGPT evalsAutomatically triggered by commit 8d85979 Results of HolmesGPT evals
Legend
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
📋 Valid eval names (use with filter)test_ask_holmes:
test_investigate:
|
|
✅ Docker image ready for
Use this tag to pull the image for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:4c2aef2
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:4c2aef2 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:4c2aef2
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:4c2aef2Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:4c2aef2Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:4c2aef2 |
WalkthroughbuildRerunFooter’s Markdown was restructured (legend table, nested details, updated placeholders). The eval workflow imports and appends that rerun footer in more places (previous isManual gating removed) and threads expanded eval parameter fields between steps. The test report legend was removed. Changes
Sequence Diagram(s)sequenceDiagram
autonumber
participant WF as GitHub Actions workflow
participant Helper as eval-comment-helpers.js
participant Builder as buildBody (comment builder)
participant GH as GitHub API (Comments)
Note over WF,Helper: workflow now calls buildRerunFooter in more places
WF->>Helper: buildRerunFooter(p, context)
Helper-->>WF: Markdown footer (legend table + nested details)
WF->>Builder: buildBody(..., footer, expanded params)
Builder-->>WF: assembled comment body
WF->>GH: create/update comment with assembled body
GH-->>WF: comment created/updated
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested reviewers
Pre-merge checks❌ Failed checks (1 inconclusive)
✅ Passed checks (2 passed)
📜 Recent review detailsConfiguration used: Organization UI Review profile: CHILL Plan: Pro 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
Comment |
- Add buildRerunFooter to all comment update steps - Include valid_markers/evals data when available (after test-preview step) - Earlier steps show placeholder text for markers/evals Signed-off-by: Claude <noreply@anthropic.com>
✅ Results of HolmesGPT evalsAutomatically triggered by commit ce87e8e Results of HolmesGPT evals
Legend
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
📋 Valid eval names (use with filter)test_ask_holmes:
test_investigate:
|
Show "Loading... or see [source]" instead of "No markers found" Signed-off-by: Claude <noreply@anthropic.com>
✅ Results of HolmesGPT evalsAutomatically triggered by commit 865f0ce Results of HolmesGPT evals
Legend
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
📋 Valid eval names (use with filter)test_ask_holmes:
test_investigate:
|
Legend is now only in the footer details section Signed-off-by: Claude <noreply@anthropic.com>
✅ Results of HolmesGPT evalsAutomatically triggered by commit c49f969 Results of HolmesGPT evals
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
📋 Valid eval names (use with filter)test_ask_holmes:
test_investigate:
|
Signed-off-by: Claude <noreply@anthropic.com>
✅ Results of HolmesGPT evalsAutomatically triggered by commit dc8628e Results of HolmesGPT evals
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
📋 Valid eval names (use with filter)test_ask_holmes:
test_investigate:
|
|
/eval |
🧪 Manual Eval Results
Results of HolmesGPT evals
|
|
@aantn Your eval run has finished. ✅ Completed successfully 👆 See updated results above. |
Summary by CodeRabbit
✏️ Tip: You can customize this high-level summary in your review settings.