Improvements to manual evals run - #1265
Conversation
- Remove 'regression' as default marker for /eval comments and workflow_dispatch triggers - they now run all LLM tests by default - Keep 'regression' as default only for automatic triggers (PR/push) - Add details section with list of valid markers and example test names - Update marker_expr to handle empty markers (just 'llm' instead of 'llm and ()') Signed-off-by: Claude <noreply@anthropic.com>
- Add test preview step that runs pytest --collect-only to show which tests will run before actually running them - Update initial comment with test count and expandable test list - Add warning that manual re-runs have no default markers and will run all LLM tests (~100+) which can take 1+ hours - Update example /eval command to include markers: regression - Update markers description to emphasize no default Signed-off-by: Claude <noreply@anthropic.com>
When a user triggers evals manually via /eval comment or workflow_dispatch, they now receive a notification when the run completes. The notification: - @mentions the user who triggered the eval - Shows success or regression count status - Points to the updated results comment above This ensures users get a GitHub notification instead of having to poll the PR for the updated comment. Signed-off-by: Claude <noreply@anthropic.com>
- Rename step names and summary text from "Test preview" to "Evals to run" - Remove the 20 test limit to show all evals that will run Signed-off-by: Claude <noreply@anthropic.com>
- Remove "Results will appear here when complete." text from both initial and running status comments since it's confusing - For manual triggers, don't create a new comment if comment_id is missing - the initial comment should always exist for manual runs Signed-off-by: Claude <noreply@anthropic.com>
- Split into two separate <details> sections: Valid markers and Valid eval names - List each marker and eval name on its own line - Include the complete list of all 140+ ask_holmes evals and 17 investigate evals Signed-off-by: Claude <noreply@anthropic.com>
grep -c returns exit code 1 when no matches found, even though it outputs "0". Using || echo "0" would result in "0\n0" being captured. Changed to || true which prevents failure without adding extra output. Signed-off-by: Claude <noreply@anthropic.com>
- For PR triggers: show "PR #123 (abc1234)" - For push triggers: show "push to branch-name (abc1234)" - Include trigger info in both initial and updated comments Signed-off-by: Claude <noreply@anthropic.com>
- Add progress checklist showing completion status of each step - Update comment after each major step (HolmesGPT setup, collect evals, KIND ready) - Reorder steps: HolmesGPT env -> Collect evals -> KIND cluster for optimization - Show completed checklist in final results comment Signed-off-by: Claude <noreply@anthropic.com>
When no model is specified in /eval comment, MODEL env var was set to empty string instead of being unset. This caused get_models() to return an empty list (since it only uses DEFAULT_MODEL when MODEL is unset, not when it's empty string), resulting in 0 tests being collected. Fix: Unset MODEL in bash if it's empty so Python uses DEFAULT_MODEL. Signed-off-by: Claude <noreply@anthropic.com>
Replace hardcoded lists of valid markers and eval names with dynamic collection at runtime: - Markers: extracted from pyproject.toml [tool.pytest.ini_options] section - Eval names: listed from tests/llm/fixtures/test_ask_holmes/ and test_investigate/ This ensures the comment always shows the current list of available markers and tests without manual updates when new tests are added. Signed-off-by: Claude <noreply@anthropic.com>
Refactor the 5 comment-generating steps to use consistent helper functions: - renderProgress(): Progress checklist renderer - renderParamsTable(): Parameters table for manual runs - buildBody(): Main comment body builder with icon/title customization - buildRerunFooter(): Re-run instructions for automatic runs (results only) This ensures consistent styling between manual and automatic eval runs, and between initial/progress/results comments. The helpers are defined inline in each step (GitHub Actions limitation) but follow the same pattern. Signed-off-by: Claude <noreply@anthropic.com>
WalkthroughWorkflow and comment tooling updated: eval-regression GitHub Actions now compute trigger context, conditional marker expressions, and incremental progress comments; they auto-detect PRs from branches for manual runs and expose new outputs. A new helper script builds normalized params and dynamic comment bodies/footers. Changes
Sequence Diagram(s)sequenceDiagram
autonumber
actor User
participant GH as "GitHub (events/API)"
participant WF as "eval-regression workflow"
participant Helper as "eval-comment-helpers.js"
participant Runner as "Actions runner"
participant KIND as "KIND cluster"
participant Evals as "Eval jobs"
User->>GH: trigger (issue_comment / workflow_dispatch / push / PR)
GH->>WF: start workflow (context)
WF->>WF: compute trigger_source, triggered_by, pr_number, marker_expr, base_params
WF->>Helper: buildBody(progress=init, params)
Helper-->>WF: initial comment body
WF->>GH: create/update PR comment
WF->>Runner: setup HolmesGPT env
Runner-->>WF: env ready
WF->>Helper: buildBody(progress=collected, params)
Helper-->>WF: updated body
WF->>KIND: setup KIND cluster
KIND-->>WF: KIND ready
WF->>Evals: run evals (apply marker_expr, unset MODEL if empty)
Evals-->>WF: results
WF->>Helper: buildBody(progress=done, extras=results)
Helper-->>WF: final comment body (and rerun footer if automatic)
WF->>GH: update comment, remove/add reactions, post completion comment if manual
Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes Possibly related PRs
Suggested reviewers
Pre-merge checks❌ Failed checks (1 inconclusive)
✅ Passed checks (2 passed)
📜 Recent review detailsConfiguration used: Organization UI Review profile: CHILL Plan: Pro 📒 Files selected for processing (2)
🧰 Additional context used🧠 Learnings (1)📚 Learning: 2025-12-21T13:17:57.170ZApplied to files:
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
🔇 Additional comments (19)
Comment |
✅ Results of HolmesGPT evalsAutomatically triggered by PR #1265 (efee49a) Progress:
Results of HolmesGPT evals
Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
📋 Valid eval names (use with filter)test_ask_holmes:
test_investigate:
|
|
✅ Docker image ready for
Use this tag to pull the image for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:fe7a182
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:fe7a182 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:fe7a182
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:fe7a182Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:fe7a182Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:fe7a182 |
There was a problem hiding this comment.
Actionable comments posted: 1
🧹 Nitpick comments (2)
.github/workflows/eval-regression.yaml (2)
204-241: Consider extracting helper functions to avoid duplication.These helper functions (
renderProgress,renderParamsTable,buildBody) are duplicated verbatim in 5 steps throughout the workflow (lines 284-312, 394-422, 463-491, 569-597). While GitHub Actions JavaScript steps don't support sharing code directly, consider:
- Composite action: Create a reusable composite action that handles progress updates
- Inline template string: Define the helpers once and pass them via outputs (limited but possible)
- External script: Move to a
.github/scripts/file and useactions/github-scriptwithscriptfrom fileThis would reduce ~200 lines of duplication and make future updates less error-prone.
358-364: Fragile marker extraction from pyproject.toml.The sed/grep pipeline on line 360 assumes a specific format for
[tool.pytest.ini_options]inpyproject.toml. If the file format changes (e.g., markers in a different section, different quoting), this could silently fail.Consider adding a fallback or validation:
- VALID_MARKERS=$(sed -n '/^\[tool\.pytest\.ini_options\]/,/^\[/p' pyproject.toml | grep -E '"[a-zA-Z]' | sed 's/.*"\([^:]*\):.*/\1/' | grep -v '^llm$' | sort | awk '{print "- \x60" $0 "\x60"}') + VALID_MARKERS=$(sed -n '/^\[tool\.pytest\.ini_options\]/,/^\[/p' pyproject.toml | grep -E '"[a-zA-Z]' | sed 's/.*"\([^:]*\):.*/\1/' | grep -v '^llm$' | sort | awk '{print "- \x60" $0 "\x60"}') + if [[ -z "$VALID_MARKERS" ]]; then + echo "::warning::Could not extract markers from pyproject.toml" + fi
📜 Review details
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (1)
.github/workflows/eval-regression.yaml
🧰 Additional context used
🧠 Learnings (3)
📚 Learning: 2025-12-21T13:17:57.170Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Workflow section with numbered sequential steps containing Action, Function Description, Parameters, Expected Output, and Success/Failure Criteria
Applied to files:
.github/workflows/eval-regression.yaml
📚 Learning: 2025-12-21T13:17:57.170Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include verification steps in workflow to confirm each diagnostic action was successful before proceeding
Applied to files:
.github/workflows/eval-regression.yaml
📚 Learning: 2025-12-21T13:17:57.170Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Use generic function descriptions in workflow steps (e.g., 'execute a command to test network connectivity') rather than tool-specific names to enable mapping to available tools
Applied to files:
.github/workflows/eval-regression.yaml
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
- GitHub Check: llm_evals
- GitHub Check: build (3.12)
- GitHub Check: build (3.10)
- GitHub Check: build (3.11)
- GitHub Check: build
🔇 Additional comments (8)
.github/workflows/eval-regression.yaml (8)
18-21: Clear input description update.The updated description accurately reflects the new behavior where empty markers runs all LLM tests. This aligns with the changed defaulting semantics for manual vs automatic triggers.
110-121: Consistent default handling in parseComment.The change to initialize
markers: ''aligns with the new behavior where manual triggers don't automatically get the regression marker.
123-167: Well-structured trigger attribution logic.The
triggeredByvariable correctly captures the initiating user for manual runs (comment author or workflow actor) while leaving it empty for automatic triggers. The descriptive trigger sources with SHA snippets provide useful context.
169-184: Intentional behavior shift for manual runs.The conditional defaulting ensures manual
/evalandworkflow_dispatchdon't automatically filter to regression tests. The warning in the re-run footer (line 603-604) appropriately alerts users about the broader test scope.
544-546: Consistent MODEL handling.Good addition to unset
MODELwhen empty, matching the behavior in the "Collect evals" step. This ensuresget_models()properly falls back toDEFAULT_MODEL.
599-624: Excellent user guidance in re-run footer.The detailed instructions with warnings about running all tests, options table, and collapsible sections for valid markers/evals significantly improve the user experience for re-running evals.
677-697: Helpful completion notification for manual runs.The @mention notification ensures the triggering user is alerted when their eval run completes, with clear status indication. The condition chain correctly limits this to manual runs with a known triggerer.
660-663: Appropriate conditional for re-run footer.The re-run instructions are correctly shown only for automatic runs. Users who manually triggered the eval already know the process.
|
/eval |
🧪 Manual Eval Results
Results of HolmesGPT evals
Legend
|
- Create .github/scripts/eval-comment-helpers.js with reusable functions - Update all 5 workflow steps to require() the helpers instead of copy-paste - Remove redundant emoji from progress checklist items (⏳ and ✅ were redundant since checkbox state already indicates completion) Signed-off-by: Claude <noreply@anthropic.com>
✅ Results of HolmesGPT evalsAutomatically triggered by PR #1265 (1536571) Progress:
Results of HolmesGPT evals
Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
📋 Valid eval names (use with filter)test_ask_holmes:
test_investigate:
|
- Change PR trigger source from "PR #N (sha)" to just "commit sha" (PR number is redundant since comment is on the PR) - Remove eyes reaction when workflow completes, leaving only hooray Signed-off-by: Claude <noreply@anthropic.com>
✅ Results of HolmesGPT evalsAutomatically triggered by commit 7bc70b3 Progress:
Results of HolmesGPT evals
Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
📋 Valid eval names (use with filter)test_ask_holmes:
test_investigate:
|
There was a problem hiding this comment.
Actionable comments posted: 1
📜 Review details
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (2)
.github/scripts/eval-comment-helpers.js.github/workflows/eval-regression.yaml
🧰 Additional context used
🧠 Learnings (1)
📓 Common learnings
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include conditional logic (IF/ELSE) in workflow steps when branching is required based on diagnostic findings
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Workflow section with numbered sequential steps containing Action, Function Description, Parameters, Expected Output, and Success/Failure Criteria
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Use generic function descriptions in workflow steps (e.g., 'execute a command to test network connectivity') rather than tool-specific names to enable mapping to available tools
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
- GitHub Check: llm_evals
- GitHub Check: build (3.11)
- GitHub Check: build (3.12)
- GitHub Check: build (3.10)
🔇 Additional comments (12)
.github/scripts/eval-comment-helpers.js (4)
11-15: LGTM!The progress rendering logic is clean and straightforward.
22-31: LGTM!The conditional parameter table rendering handles optional fields appropriately.
40-56: LGTM!The body builder correctly differentiates between manual and automatic runs and handles optional test previews.
64-88: Function logic is sound.The rerun footer construction is well-structured and provides comprehensive instructions.
Note: The past review comment regarding misleading footer messages in fork PRs still applies to lines 80, 84, and 86 where the fallback text may be displayed even when collection was skipped.
.github/workflows/eval-regression.yaml (8)
123-183: LGTM!The parameter extraction and validation logic correctly handles different trigger sources (comment, workflow_dispatch, automatic) and properly constructs the
triggered_byandmarker_exproutputs.
204-273: LGTM!The integration with
buildBodyis correct, and the reordering of setup steps (HolmesGPT before collecting evals) is logical since pytest is required for collection.
274-323: Test collection logic is sound.The collection of test names, markers, and eval names from the filesystem is well-implemented.
Note: Line 275's condition
&& steps.eval-params.outputs.pr_number != ''means this step skips for push events and fork PRs. The outputs are then used in the footer (lines 460-462) with|| ''fallbacks, which relates to the existing review comment about misleading footer messages.
324-397: LGTM!The incremental progress updates provide good user feedback and correctly incorporate test counts and conditional test previews.
441-498: Results posting logic is well-structured.The conditional logic for manual vs automatic runs is correct, and the rerun footer is appropriately limited to automatic runs only.
Note: Lines 460-462 are specifically mentioned in the existing review comment regarding misleading footer messages in fork PRs.
500-521: LGTM!The user notification step provides helpful feedback by mentioning the user who triggered the manual eval and summarizing the results.
527-543: LGTM!The reaction cleanup logic provides better UX by removing the temporary "eyes" reaction before adding the completion "hooray" reaction.
423-424: No action needed—Python code correctly handles unset MODEL.The
os.environ.get("MODEL", DEFAULT_MODEL)intests/llm/utils/test_env_vars.pyalready provides the fallback toDEFAULT_MODELwhen theMODELenvironment variable is unset. Theget_models()function uses this variable and works as intended.
- Add buildParams() helper to centralize type conversions and defaults - Update all 5 workflow steps to use buildParams() for cleaner code - All transformations (=== 'true', || 'default', parseInt) in one place Signed-off-by: Claude <noreply@anthropic.com>
✅ Results of HolmesGPT evalsAutomatically triggered by commit ecfe41e Progress:
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
📋 Valid eval names (use with filter)test_ask_holmes:
test_investigate:
|
- Output base_params JSON from eval-params step containing common params - Each step now spreads base_params and only adds step-specific extras - Reduces ~7 lines of repeated parameter passing per step Signed-off-by: Claude <noreply@anthropic.com>
✅ Results of HolmesGPT evalsAutomatically triggered by commit 452b283 Progress:
Results of HolmesGPT evals
Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
📋 Valid eval names (use with filter)test_ask_holmes:
test_investigate:
|
When triggering via GitHub Actions UI without specifying a PR number, automatically find the open PR associated with the selected branch. This makes it easier to run evals - just select the branch and go. Signed-off-by: Claude <noreply@anthropic.com>
✅ Results of HolmesGPT evalsAutomatically triggered by commit 9b8a501 Progress:
Results of HolmesGPT evals
Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
📋 Valid eval names (use with filter)test_ask_holmes:
test_investigate:
|
🧪 Manual Eval Results
Progress:
Results of HolmesGPT evals
Legend
|
|
@aantn Your eval run has finished. 👆 See updated results above. |
There was a problem hiding this comment.
Actionable comments posted: 0
♻️ Duplicate comments (2)
.github/workflows/eval-regression.yaml (2)
19-21: This breaking change has already been flagged.As noted in the previous review, the empty default for
markerscombined with the conditional logic on line 194 means manual runs with no markers specified will execute all 100+ LLM tests. This is a significant change in behavior.
285-334: Undefined outputs when collection step is skipped.The "Collect evals to run" step (line 286) only runs when
steps.eval-params.outputs.pr_number != '', but the "Post evaluation results" step (line 435) runs withalways()and references collection outputs (lines 445-447) that may not exist if collection was skipped.This is related to the fork PR issue flagged in previous reviews. When
pr_numberis empty (e.g., workflow_dispatch without auto-detected PR), the collection step skips, but posting still attempts to usesteps.test-preview.outputs.valid_markers,ask_holmes_evals, andinvestigate_evals, which will be undefined.🔎 Proposed fix to handle missing outputs
Add fallback handling in the "Post evaluation results" step:
const p = buildParams({ ...JSON.parse(${{ toJSON(steps.eval-params.outputs.base_params) }}), comment_id: ${{ toJSON(steps.initial-comment.outputs.comment_id) }}, duration: ${{ toJSON(steps.evals.outputs.duration) }}, - valid_markers: ${{ toJSON(steps.test-preview.outputs.valid_markers) }}, - ask_holmes_evals: ${{ toJSON(steps.test-preview.outputs.ask_holmes_evals) }}, - investigate_evals: ${{ toJSON(steps.test-preview.outputs.investigate_evals) }} + valid_markers: ${{ toJSON(steps.test-preview.outputs.valid_markers) }} || '', + ask_holmes_evals: ${{ toJSON(steps.test-preview.outputs.ask_holmes_evals) }} || '', + investigate_evals: ${{ toJSON(steps.test-preview.outputs.investigate_evals) }} || '' });Or, better yet, only pass these when collection ran:
const p = buildParams({ ...JSON.parse(${{ toJSON(steps.eval-params.outputs.base_params) }}), comment_id: ${{ toJSON(steps.initial-comment.outputs.comment_id) }}, duration: ${{ toJSON(steps.evals.outputs.duration) }}, - valid_markers: ${{ toJSON(steps.test-preview.outputs.valid_markers) }}, - ask_holmes_evals: ${{ toJSON(steps.test-preview.outputs.ask_holmes_evals) }}, - investigate_evals: ${{ toJSON(steps.test-preview.outputs.investigate_evals) }} + ...${{ steps.test-preview.outcome == 'success' }} && { + valid_markers: ${{ toJSON(steps.test-preview.outputs.valid_markers) }}, + ask_holmes_evals: ${{ toJSON(steps.test-preview.outputs.ask_holmes_evals) }}, + investigate_evals: ${{ toJSON(steps.test-preview.outputs.investigate_evals) }} + } });
🧹 Nitpick comments (1)
.github/workflows/eval-regression.yaml (1)
148-169: Consider failing early when multiple PRs are detected.When multiple open PRs are found for a branch (lines 162-164), the workflow logs a warning but continues with
prNumber = null, which means no comments will be posted. Consider failing the workflow with a clear error message asking the user to specifypr_numberexplicitly, rather than silently continuing without PR detection.🔎 Proposed enhancement for multiple PR detection
} else if (prs.data.length > 1) { - core.warning(`Multiple open PRs found for branch ${branch}, please specify pr_number`); - prNumber = null; + core.setFailed(`Multiple open PRs found for branch ${branch}. Please specify pr_number explicitly in workflow inputs. Found PRs: ${prs.data.map(p => `#${p.number}`).join(', ')}`); + return; } else {
📜 Review details
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (1)
.github/workflows/eval-regression.yaml
🧰 Additional context used
🧠 Learnings (1)
📚 Learning: 2025-12-21T13:17:57.170Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include verification steps in workflow to confirm each diagnostic action was successful before proceeding
Applied to files:
.github/workflows/eval-regression.yaml
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
- GitHub Check: llm_evals
- GitHub Check: build (3.10)
- GitHub Check: build (3.11)
- GitHub Check: build (3.12)
- GitHub Check: llm_evals
🔇 Additional comments (5)
.github/workflows/eval-regression.yaml (5)
208-218: LGTM! Good DRY improvement.The
base_paramsJSON output effectively reduces parameter repetition across subsequent steps while maintaining proper type safety and validation.
262-284: LGTM! Progress tracking improves user experience.The new progress update steps provide clear feedback at each stage of the evaluation workflow. The conditional logic properly ensures updates only occur when a comment exists.
Also applies to: 335-359, 367-390
416-418: LGTM! Proper handling of empty MODEL.Unsetting the
MODELenvironment variable when empty allows the Python code to use its default model selection logic, which is more robust than passing an empty string.
479-500: LGTM! Good user experience for manual runs.The completion notification with @ mention provides clear feedback to users who triggered manual eval runs, including regression status.
506-522: LGTM! Reaction handling provides good feedback.The logic to remove the "eyes" reaction and add "hooray" properly indicates workflow completion. The code defensively handles the case where the eyes reaction might not exist.
Note: The
listForIssueCommentAPI might paginate with many reactions, but this is unlikely to be an issue in practice.
- Hide Progress section once run is complete (pass null to buildBody) - Consolidate legend into single collapsible <details> section - Change "Regression" to "Failure" throughout for clarity Signed-off-by: Claude <noreply@anthropic.com>
✅ Results of HolmesGPT evalsAutomatically triggered by commit 959f740 Results of HolmesGPT evals
Legend
📖 Legend🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
📋 Valid eval names (use with filter)test_ask_holmes:
test_investigate:
|
🧪 Manual Eval Results
Results of HolmesGPT evals
Legend
|
|
@aantn Your eval run has finished. 👆 See updated results above. |
Summary by CodeRabbit
New Features
Improvements
✏️ Tip: You can customize this high-level summary in your review settings.