Repository navigation
Allow manually triggering evals - #1258
Conversation
Features: - Add /eval command support in PR comments with options: --model/-m: Specify model(s) to test --markers/-M: Pytest markers (default: regression) --filter/-k: Pytest -k filter --iterations/-i: Number of iterations (max 10) - Add workflow_dispatch trigger for GitHub UI triggering - Post initial "running" comment when evals start - Update same comment with results when complete - Each run preserves its own comment (history maintained) - Input validation to prevent shell injection - Add manual trigger instructions in automatic run comments
- Consolidate PR number logic into eval-params step - Remove duplicate validation (parse step extracts, eval-params validates) - Add run_url output to avoid recalculating - Use shared params object in comment steps - Condense duration formatting in bash - 621 → 387 lines
Before: /eval --model gpt-4o --filter "my test" After: /eval model: gpt-4o filter: my test Much simpler parsing (split on newlines, split on colon) and easier for humans to read/write in PR comments.
- Simplify parse-eval-command job to just: get PR SHA + add reaction - Move parseComment() function into eval-params step - Remove 4 outputs (model, markers, filter, iterations) from job 1 - Eliminates duplicate passing of values between jobs - 388 → 357 lines
- Consolidate parse-eval-command job into llm_evals as a step - Add permission check: only OWNER, MEMBER, COLLABORATOR can trigger /eval - Prevents random external users from triggering expensive evals - Simplifies workflow structure (1 job instead of 2)
- Fix workflow dispatch link to point to specific workflow file - Change rocket reaction to eyes for clearer "processing" signal - Simplify secrets check (repo has all or none) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> Signed-off-by: Natan Yellin <aantny@gmail.com>
…/HolmesGPT/holmesgpt into claude/manual-eval-trigger-CZGZL Signed-off-by: Natan Yellin <aantny@gmail.com>
WalkthroughAdds manual and comment-triggered eval runs to the existing eval-regression workflow: supports Changes
Sequence Diagram(s)sequenceDiagram
participant User as GitHub User
participant GH as GitHub (Events)
participant Actions as GitHub Actions Runner
participant Repo as Repository Checkout
participant Test as Test Runner / pytest
participant PR as Pull Request (comments)
rect rgba(200,230,255,0.3)
User->>GH: /eval comment on PR (or triggers workflow_dispatch)
GH->>Actions: dispatch workflow (issue_comment or workflow_dispatch)
end
rect rgba(230,255,200,0.25)
Actions->>GH: validate commenter permissions (if issue_comment)
GH-->>Actions: permission result
Actions->>GH: fetch PR head SHA (if comment)
GH-->>Actions: PR head SHA
Actions->>Repo: checkout appropriate ref (PR SHA or ref)
Repo-->>Actions: code checked out
end
rect rgba(255,240,200,0.25)
Actions->>Actions: Determine eval parameters (parse inputs, validate regex, set defaults)
Actions->>Actions: check AZURE_API_KEY -> decide should-run
Actions->>Actions: setup KIND & environment (conditional)
end
rect rgba(255,220,220,0.15)
Actions->>Test: run pytest with dynamic args (MODEL, ITERATIONS, marker expr, filter)
Test-->>Actions: test results, optional report/regressions
Actions->>GH: post or update PR comment with results (manual vs automatic formatting)
Actions->>GH: add reactions (eyes, completion) for comment-triggered runs
end
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Suggested reviewers
Pre-merge checks✅ Passed checks (3 passed)
📜 Recent review detailsConfiguration used: Organization UI Review profile: CHILL Plan: Pro 📒 Files selected for processing (1)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
🔇 Additional comments (12)
Comment |
|
✅ Docker image ready for
Use this tag to pull the image for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:c1ea6b1
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:c1ea6b1 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:c1ea6b1
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:c1ea6b1Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:c1ea6b1Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:c1ea6b1 |
There was a problem hiding this comment.
Actionable comments posted: 0
🧹 Nitpick comments (1)
.github/workflows/eval-regression.yaml (1)
171-181: Secrets check is functional but could use env var indirection.The pattern works and secrets are masked, but passing via
env:would be slightly safer against edge-case log leaks.🔎 Optional improvement using env var
- name: Check if tests should run id: check-tests if: github.event_name != 'issue_comment' || steps.eval-comment.outcome == 'success' shell: bash + env: + HAS_SECRETS: ${{ secrets.AZURE_API_KEY != '' }} run: | - # Check one secret as proxy for secrets access (repo has all or none) - if [[ -n "${{ secrets.AZURE_API_KEY }}" ]]; then + if [[ "$HAS_SECRETS" == "true" ]]; then echo "should-run=true" >> $GITHUB_OUTPUT else echo "should-run=false" >> $GITHUB_OUTPUT fi
📜 Review details
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (1)
.github/workflows/eval-regression.yaml
🧰 Additional context used
🪛 actionlint (1.7.9)
.github/workflows/eval-regression.yaml
87-87: unexpected end of input while parsing variable access, function call, null, bool, int, float or string. expecting "IDENT", "(", "INTEGER", "FLOAT", "STRING"
(expression)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
- GitHub Check: build (3.12)
- GitHub Check: build (3.11)
- GitHub Check: build (3.10)
- GitHub Check: build
🔇 Additional comments (12)
.github/workflows/eval-regression.yaml (12)
8-31: Well-structured workflow_dispatch and issue_comment triggers.The input definitions with clear descriptions and sensible defaults are good for usability. The
issue_commenttrigger scoped tocreatedtype is appropriate.
33-36: Permissions are appropriately scoped.The added
issues: writepermission is required for the reactions and comment functionality on issue_comment triggers.
41-45: Job condition correctly gates the workflow.The multi-condition check properly ensures that issue_comment events only trigger for PR comments that start with
/eval.
48-75: Good permission validation and user feedback.The authorization check via
author_associationis a solid security practice. The 'eyes' reaction provides immediate feedback to the user that their comment was recognized.
77-80: Checkout logic handles all trigger types correctly.The conditional checkout with fallback from PR SHA to
github.refappropriately handles both comment-triggered and other event types.
82-169: Well-designed input validation with whitelist patterns.The security approach is solid: whitelist regex patterns combined with
toJSON()for proper escaping. The parameter parsing for/evalcomments is cleanly implemented.The static analysis warning on line 87 is a false positive —
toJSON()expressions are valid withinactions/github-scriptblocks and produce properly quoted JavaScript string literals.
183-218: Initial comment provides good visibility into running evals.The conditional formatting for manual vs automatic runs and the use of a markdown table for parameters provides clear status reporting.
220-232: Setup steps properly gated by secrets availability check.
234-271: Secure test execution with properly quoted variables.The test step correctly follows the security guidance: user inputs are passed via environment variables and properly quoted in bash. The duration formatting provides useful human-readable output.
273-335: Results posting handles all cases gracefully.The conditional update vs create logic is correct, and the detailed re-run instructions in automatic runs are helpful. The
fs.existsSyncchecks prevent errors when report files are missing.
337-345: Completion reaction provides clear feedback.The 'hooray' reaction gives users visible confirmation that the eval completed.
347-355: Regressions are reported but don't fail the workflow.This step logs regressions without causing a workflow failure. If regressions should block PRs, you may want to add
exit 1when regressions are detected.Is the current behavior (report-only, no failure) intentional? If regressions should block merging, consider:
if [[ -f "regressions.txt" ]]; then echo "⚠️ There are regressions in the evals. Please check the evals file for details." cat regressions.txt + exit 1 else echo "✅ All tests passed without regressions." fi
Remove `${{ }}` from JS comment - GitHub parses expressions even in comments.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Signed-off-by: Natan Yellin <aantny@gmail.com>
Results of HolmesGPT evalsDuration: 4m 5s | View workflow logs Results of HolmesGPT evals
Legend
🔄 Re-run evals manuallyOption 1: Comment on this PR with Or with options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" |
Results of HolmesGPT evalsDuration: 4m 3s | View workflow logs Results of HolmesGPT evals
Legend
🔄 Re-run evals manuallyOption 1: Comment on this PR with Or with options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" |
Results of HolmesGPT evalsDuration: 3m 58s | View workflow logs Results of HolmesGPT evals
Legend
🔄 Re-run evals manuallyOption 1: Comment on this PR with Or with options (one per line):
Option 2: Trigger via GitHub Actions UI → "Run workflow" |
Summary by CodeRabbit
/evalPR comment, including a re-run UI in results.✏️ Tip: You can customize this high-level summary in your review settings.