Repository navigation
/evals dx improvements - #1303
Conversation
When clicking the "Trigger via GitHub Actions UI" link, the branch dropdown is now pre-filled with the current PR/push branch using the ref query parameter. Signed-off-by: Claude <noreply@anthropic.com>
The params table now shows "View logs | Rerun" in the Workflow row. The Rerun link goes to the workflow dispatch page with the branch pre-selected, making it easy to trigger a new run. Signed-off-by: Claude <noreply@anthropic.com>
- Change valid markers and eval names from bullet lists to comma-separated clickable links that go to the workflow dispatch page - Add prominent warning that /eval comments always use the workflow from master, with link to trigger from PR branch instead - This helps users who modify the GitHub Action (e.g., adding secrets) understand why their changes don't take effect with /eval Signed-off-by: Claude <noreply@anthropic.com>
|
|
Warning Rate limit exceeded@aantn has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 1 minutes and 16 seconds before requesting another review. ⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. 📒 Files selected for processing (1)
WalkthroughComment-rendering helpers were updated to accept an optional context so workflow URLs and rerun links can include branch refs; markers/eval-name outputs were collapsed to a comma-separated Changes
Sequence Diagram(s)sequenceDiagram
participant User as PR commenter
participant GitHub as GitHub API
participant Workflow as GH Actions (eval-regression)
participant Helper as eval-comment-helpers.js
participant Runner as Action Runner
User->>GitHub: Post trigger comment (/eval or /list)
GitHub->>Workflow: Start workflow (with event/context)
Workflow->>Runner: Run jobs (list_evals / eval jobs)
Runner->>Helper: buildBody(..., { context }) / buildRerunFooter(..., { context })
Helper-->>Runner: Rendered comment body (includes workflow URL + rerun link if ref present)
Runner->>GitHub: Post/Update comment on PR with rendered body
Note over Helper,GitHub: Valid markers produced as comma-separated output for helpers
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested reviewers
Pre-merge checks✅ Passed checks (3 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
✅ Docker image ready for
Use this tag to pull the image for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:907ba60
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:907ba60 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:907ba60
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:907ba60Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:907ba60Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:907ba60 |
There was a problem hiding this comment.
Actionable comments posted: 0
🧹 Nitpick comments (2)
.github/scripts/eval-comment-helpers.js (2)
97-112: Consider handling empty items after filtering.The function handles
!itemsbut could return an empty string if all items are empty after splitting/trimming (e.g., input is just commas). This would render as nothing in the markdown.🔎 Suggested defensive check
function formatAsLinks(items, workflowUrl, paramType) { if (!items) return '_(loading...)_'; - return items.split(',').map(item => { + const links = items.split(',').map(item => { const trimmed = item.trim(); if (!trimmed) return ''; // Link to workflow dispatch - user will need to enter the value manually return `[\`${trimmed}\`](${workflowUrl})`; }).filter(Boolean).join(', '); + return links || '_(none found)_'; }
104-111: TheparamTypeparameter is unused.The function accepts
paramType(documented as'markers'or'filter') but never uses it. Either remove it or implement the intended differentiation.🔎 Option 1: Remove unused parameter
-function formatAsLinks(items, workflowUrl, paramType) { +function formatAsLinks(items, workflowUrl) {Then update the call sites (lines 125-127) to remove the third argument.
📜 Review details
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (2)
.github/scripts/eval-comment-helpers.js.github/workflows/eval-regression.yaml
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
- GitHub Check: build (3.12)
- GitHub Check: build (3.11)
- GitHub Check: build (3.10)
- GitHub Check: build
- GitHub Check: llm_evals
🔇 Additional comments (8)
.github/scripts/eval-comment-helpers.js (4)
52-68: LGTM - Context-aware rerun link addition.The optional context parameter with default
nullmaintains backward compatibility. URL construction properly usesencodeURIComponentfor the branch name to prevent URL injection issues.
77-80: LGTM - Context threading through buildBody.Properly passes
extras.contexttorenderParamsTablefor manual runs, enabling the rerun link feature.
120-127: LGTM - Consistent URL construction with renderParamsTable.The base workflow URL and branch handling logic matches
renderParamsTable, ensuring consistent behavior across the comment.
140-143: LGTM - Clear warning about workflow source.The warning about
/evalusing the workflow from master is helpful for developers who modify the GitHub Action and expect their changes to take effect..github/workflows/eval-regression.yaml (4)
338-338: LGTM - Context properly passed to buildBody.The
{ context }object is correctly passed as the extras parameter, enabling the rerun link generation in the comment.
393-406: LGTM - Comma-separated format for JS consumption.The bash commands correctly produce comma-separated lists that the
formatAsLinkshelper expects. Thetr '\n' ','followed bysed 's/,$//'properly handles the trailing comma.
456-456: LGTM - Consistent context threading.Both
testPreviewandcontextare properly passed in the extras object for the progress update.
564-568: LGTM - Final results include context for rerun links.The context is correctly included in the buildBody call for the final evaluation results, ensuring the rerun link appears in the completed comment.
- Add /list command to show available eval names on demand - Replace workflow link with gh CLI command in the master warning - Change markers from linked list to plain comma-separated format - Remove eval names from footer (use /list instead) - Clean up unused ask_holmes_evals and investigate_evals params Signed-off-by: Claude <noreply@anthropic.com>
✅ Results of HolmesGPT evalsAutomatically triggered by commit fd5992c on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/prepopulate-workflow-branch-Ig8Te' Status: Success - 16 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
|
- Add includeLegend option to buildRerunFooter (default false) - Only show Legend section when final results are displayed - Bold "Historical Comparison Details" for consistency with other collapsible sections Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
🧹 Nitpick comments (1)
.github/workflows/eval-regression.yaml (1)
447-470: Consider adding format validation for marker parsing.The marker parsing assumes
pyproject.tomlmarkers follow the"name:description"format and splits on:. While this is fine for internal configuration, consider whether the script should validate the marker format or handle edge cases (e.g., markers without colons).Optional: Add defensive parsing
VALID_MARKERS=$(python3 << 'PYEOF' import tomllib with open('pyproject.toml', 'rb') as f: config = tomllib.load(f) markers = config.get('tool', {}).get('pytest', {}).get('ini_options', {}).get('markers', []) -names = [m.split(':')[0] for m in sorted(markers) if m.split(':')[0] != 'llm'] +names = [] +for m in sorted(markers): + parts = m.split(':', 1) + if parts[0] and parts[0] != 'llm': + names.append(parts[0]) print(','.join(names)) PYEOF )This validates that the marker name exists before adding it to the list.
📜 Review details
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (3)
.github/scripts/eval-comment-helpers.js.github/workflows/eval-regression.yamltests/llm/utils/reporting/github_reporter.py
🧰 Additional context used
📓 Path-based instructions (2)
**/*.py
📄 CodeRabbit inference engine (CLAUDE.md)
**/*.py: Use Ruff for formatting and linting (configured in pyproject.toml)
Type hints required (mypy configuration in pyproject.toml)
ALWAYS place Python imports at the top of the file, not inside functions or methods
Files:
tests/llm/utils/reporting/github_reporter.py
tests/**/*.py
📄 CodeRabbit inference engine (CLAUDE.md)
Tests: match source structure under
tests/
Files:
tests/llm/utils/reporting/github_reporter.py
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
- GitHub Check: llm_evals
- GitHub Check: build (3.11)
- GitHub Check: build (3.10)
- GitHub Check: build (3.12)
🔇 Additional comments (6)
tests/llm/utils/reporting/github_reporter.py (1)
67-67: LGTM!The bold formatting improves visual hierarchy in the collapsible details section.
.github/scripts/eval-comment-helpers.js (2)
47-66: LGTM!The context-aware workflow link generation is implemented correctly, with proper URL encoding for the branch parameter.
116-171: LGTM!The context-aware rerun footer implementation is well-structured:
- Proper optional parameters with defaults
- Conditional legend rendering improves UX
- The warning about
/evalusing master workflow is valuable for users testing workflow changes.github/workflows/eval-regression.yaml (3)
39-93: LGTM!The new
/listcommand implementation is well-designed:
- Proper permission checks restrict access to authorized users
- Graceful error handling for missing directories
- Clear output format with code-styled eval names
393-393: LGTM!The context parameter is consistently propagated through all
buildBodyandbuildRerunFootercalls. The conditionalincludeLegendoption is correctly applied only to the final results comment where the legend is relevant.Also applies to: 425-425, 495-495, 528-528, 599-606
483-483: LGTM!The
valid_markersoutput is consistently passed tobuildParamsacross all progress updates and final results, ensuring markers are available for display in the rerun footer.Also applies to: 516-516, 585-585
✅ Results of HolmesGPT evalsAutomatically triggered by commit f3b506f on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/prepopulate-workflow-branch-Ig8Te' Status: Success - 18 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
|
Signed-off-by: Claude <noreply@anthropic.com>
✅ Results of HolmesGPT evalsAutomatically triggered by commit a68a22f on branch Results of HolmesGPT evals
Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%. Historical Comparison DetailsFilter: excluding branch 'claude/prepopulate-workflow-branch-Ig8Te' Status: Success - 18 test/model combinations loaded Experiments compared (30):
Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" 🏷️ Valid markers
Commands: |
Summary by CodeRabbit
New Features
Improvements
Style
✏️ Tip: You can customize this high-level summary in your review settings.