Add branch-specific eval source links to dashboard - #2049
Conversation
Adds a 📄 emoji next to each eval name in the markdown report tables that links to the eval's test_case.yaml on the branch the run was executed from (GITHUB_REF_NAME / BUILDKITE_BRANCH, falling back to master). The existing bold-name link to master is preserved as the canonical reference. Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.
Tip: disable this comment in your organization's Code Review settings.
📂 Previous Runs📜 #4 · Run @ __3f3275c__ (#25930968115) — May 15, 17:17 UTC✅ Results of HolmesGPT evalsAutomatically triggered by commit 3f3275c on branch Results of HolmesGPT evals
Benchmark comparison unavailable: No ci-benchmark experiments found Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: No ci-benchmark experiments found Comparison indicators:
📜 #3 · Run @ __db109ab__ (#25930374521) — May 15, 17:06 UTC✅ Results of HolmesGPT evalsAutomatically triggered by commit db109ab on branch Results of HolmesGPT evals
Benchmark comparison unavailable: No ci-benchmark experiments found Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: No ci-benchmark experiments found Comparison indicators:
📜 #2 · Run @ __8e45250__ (#25929901763) — May 15, 16:58 UTC✅ Results of HolmesGPT evalsAutomatically triggered by commit 8e45250 on branch Results of HolmesGPT evals
Benchmark comparison unavailable: No ci-benchmark experiments found Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: No ci-benchmark experiments found Comparison indicators:
📜 #1 · Run @ __a17981e__ (#25929779500) — May 15, 16:51 UTC✅ Results of HolmesGPT evalsAutomatically triggered by commit a17981e on branch Results of HolmesGPT evals
Benchmark comparison unavailable: No ci-benchmark experiments found Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: No ci-benchmark experiments found Comparison indicators:
✅ Results of HolmesGPT evalsAutomatically triggered by commit 3e0bc94 on branch Results of HolmesGPT evals
Benchmark comparison unavailable: No ci-benchmark experiments found Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: No ci-benchmark experiments found Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" Option 3: Add PR labels to include extra evals (applies to both automatic runs and
Examples: 🏷️ Valid tags
🤖 Valid models
Commands: CLI: |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
WalkthroughAdds a helper mapping test types to fixture directories and builds branch-aware GitHub blob URLs to ChangesEval Source URL Integration
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:2d83c822
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:2d83c822 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:2d83c822
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:2d83c822
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:2d83c822
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:2d83c822 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:2d83c822
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:2d83c822Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:2d83c822 \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:2d83c822Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:2d83c822 \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:2d83c822 |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
The github_reporter is what generates the actual GitHub PR comment (evals_report.md). Adds a 📄 link next to each eval name in both the main results table and the comparison sub-tables, pointing to the eval's test_case.yaml on the branch this run was executed from (EVAL_BRANCH / GITHUB_REF_NAME / BUILDKITE_BRANCH, falling back to master). Maps test_type → fixture directory (ask → test_ask_holmes, investigate → test_investigate). Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
♻️ Duplicate comments (1)
tests/generate_eval_report.py (1)
884-896:⚠️ Potential issue | 🔴 Critical | 🏗️ Heavy liftMissing test_type information at call site.
Same issue as the heatmap table: the call to
get_eval_source_url(eval_case)on line 887 doesn't pass test_type, so URLs will be incorrect for non-"ask" test types. The data structure needs to be updated to track test_type for each eval_case.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/generate_eval_report.py` around lines 884 - 896, The call to get_eval_source_url(eval_case) doesn't pass test_type so non-"ask" URLs are wrong; update the data structure that iterates eval_case to also carry its test_type and change the call sites here to get_eval_source_url(eval_case, test_type) (and likewise pass test_type into get_braintrust_eval_filter_url if needed) so name_cell is built with the correct branch_source_url and eval_filter_url; search for other uses of get_eval_source_url/get_braintrust_eval_filter_url in this file and adjust their signatures/usages to accept the new test_type parameter and update the data construction where eval_case entries are created to include test_type.
🧹 Nitpick comments (1)
tests/generate_eval_report.py (1)
241-252: ⚡ Quick winConsider extracting URL generation logic to a shared utility.
The URL generation logic for eval source files is now duplicated between
generate_eval_report.pyandgithub_reporter.py(lines 28-47 in that file), with subtle differences in implementation. Consider extracting this to a shared utility module (e.g.,tests/llm/utils/eval_urls.py) to maintain consistency and reduce duplication.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/generate_eval_report.py` around lines 241 - 252, The URL-building logic in get_eval_source_url duplicates similar code in github_reporter (see its URL generation block) causing subtle inconsistencies; extract this logic into a single shared utility (e.g., tests/llm/utils/eval_urls.py) that exposes a function (e.g., build_eval_source_url or get_eval_source_url) and update both get_eval_source_url in generate_eval_report.py and the corresponding code in github_reporter to call the new utility so they use identical encoding/ref-fallback behavior and path construction.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/generate_eval_report.py`:
- Around line 773-786: The URL is wrong because get_eval_source_url(eval_case)
lacks the test_type; update the data collection so each eval_case entry in
eval_model_stats includes its test_type (add a test_type field where
eval_model_stats is populated), change calls to get_eval_source_url to pass that
test_type (e.g., get_eval_source_url(eval_case, test_type)), and update
get_eval_source_url's signature/logic to accept and use test_type; then retrieve
test_type from eval_model_stats when building name_cell so branch_source_url is
constructed correctly for non-"ask" tests.
---
Duplicate comments:
In `@tests/generate_eval_report.py`:
- Around line 884-896: The call to get_eval_source_url(eval_case) doesn't pass
test_type so non-"ask" URLs are wrong; update the data structure that iterates
eval_case to also carry its test_type and change the call sites here to
get_eval_source_url(eval_case, test_type) (and likewise pass test_type into
get_braintrust_eval_filter_url if needed) so name_cell is built with the correct
branch_source_url and eval_filter_url; search for other uses of
get_eval_source_url/get_braintrust_eval_filter_url in this file and adjust their
signatures/usages to accept the new test_type parameter and update the data
construction where eval_case entries are created to include test_type.
---
Nitpick comments:
In `@tests/generate_eval_report.py`:
- Around line 241-252: The URL-building logic in get_eval_source_url duplicates
similar code in github_reporter (see its URL generation block) causing subtle
inconsistencies; extract this logic into a single shared utility (e.g.,
tests/llm/utils/eval_urls.py) that exposes a function (e.g.,
build_eval_source_url or get_eval_source_url) and update both
get_eval_source_url in generate_eval_report.py and the corresponding code in
github_reporter to call the new utility so they use identical
encoding/ref-fallback behavior and path construction.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 38f9942d-8eac-4c9e-b256-6a2641127e33
📒 Files selected for processing (2)
tests/generate_eval_report.pytests/llm/utils/reporting/github_reporter.py
This reverts commit a17981e. Signed-off-by: Claude <noreply@anthropic.com>
- On pull_request events GITHUB_REF_NAME is the virtual "<num>/merge" ref (e.g. 2049/merge) which doesn't exist as a browsable branch. Prefer GITHUB_HEAD_REF (the PR head branch) when set, falling back to GITHUB_REF_NAME for push events. - Move the 📄 link to the start of the eval name cell so the column is easier to scan. Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/llm/utils/reporting/github_reporter.py`:
- Around line 53-55: The GitHub URL construction inserts test_case_name raw
which can break links; update the code that builds the URL (the string
containing encoded_ref, fixture_dir and test_case_name) to percent-encode
test_case_name just like encoded_ref (e.g., create encoded_test_case via
urllib.parse.quote(..., safe='/') or similar) and use encoded_test_case in the
f-string instead of test_case_name so reserved characters are escaped.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 352ae3a7-d0d1-4e65-91ca-a05181d5b819
📒 Files selected for processing (1)
tests/llm/utils/reporting/github_reporter.py
The PR comment report only runs in GitHub Actions; we don't use Buildkite for evals. Removes the dead fallback. Signed-off-by: Claude <noreply@anthropic.com>
Summary
Add links to eval test case source files on the branch where the evaluation was executed, in addition to the existing master branch links. This helps users quickly access the exact eval configuration that was run.
Key Changes
New function
get_eval_source_url(): Generates GitHub URLs totest_case.yamlfiles on a specific git ref (branch/tag/SHA), with fallback logic:GITHUB_REF_NAMEorBUILDKITE_BRANCHenvironment variables (set by CI systems)Updated dashboard generation:
Updated legend: Added documentation for the new 📄 source link icon in the dashboard legend
Implementation Details
get_eval_source_url()function usesurllib.parse.quote()to safely encode git refs (handles branch names with slashes, special characters, etc.)https://claude.ai/code/session_01GNAbUvZiXhAi6nBBVUSAGA
Summary by CodeRabbit