Skip to content

Add branch-specific eval source links to dashboard - #2049

Merged
aantn merged 5 commits into
masterfrom
claude/add-eval-source-emoji-DsZft
May 15, 2026
Merged

aantn merged 5 commits into
masterfrom
claude/add-eval-source-emoji-DsZft

Conversation

@aantn

@aantn aantn commented May 15, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Add links to eval test case source files on the branch where the evaluation was executed, in addition to the existing master branch links. This helps users quickly access the exact eval configuration that was run.

Key Changes

  • New function get_eval_source_url(): Generates GitHub URLs to test_case.yaml files on a specific git ref (branch/tag/SHA), with fallback logic:

    • Uses GITHUB_REF_NAME or BUILDKITE_BRANCH environment variables (set by CI systems)
    • Falls back to "master" when no ref is available
    • Properly URL-encodes the ref for special characters
  • Updated dashboard generation:

    • Added 📄 icon links to branch-specific eval sources in both summary and detailed breakdown tables
    • Maintains existing master branch links (🔗 for Braintrust, canonical source)
    • Refactored cell construction to reduce duplication when building name cells with multiple links
  • Updated legend: Added documentation for the new 📄 source link icon in the dashboard legend

Implementation Details

  • The get_eval_source_url() function uses urllib.parse.quote() to safely encode git refs (handles branch names with slashes, special characters, etc.)
  • Applied consistently to both heatmap summary table and detailed breakdown table
  • Maintains backward compatibility - works in environments without CI environment variables by defaulting to master

https://claude.ai/code/session_01GNAbUvZiXhAi6nBBVUSAGA

Summary by CodeRabbit

  • New Features
    • Test case names in reports now include a clickable document link (📄) to the source test definitions on GitHub, added to both comparison rows and main results tables for faster navigation and improved traceability.

Review Change Stack

Adds a 📄 emoji next to each eval name in the markdown report tables
that links to the eval's test_case.yaml on the branch the run was
executed from (GITHUB_REF_NAME / BUILDKITE_BRANCH, falling back to
master). The existing bold-name link to master is preserved as the
canonical reference.

Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@github-actions

github-actions Bot commented May 15, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 #4 · Run @ __3f3275c__ (#25930968115) — May 15, 17:17 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 3f3275c on branch claude/add-eval-source-emoji-DsZft

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 📄 09_crashpod 43.6s 6 11 $0.2990 130,428 127,876 24,124 2,552 989 102,552 25,324 239 —
✅ 📄 101_loki_historical_logs_pod_deleted 87.8s 9 19 $0.4667 233,513 228,194 30,809 5,319 931 197,375 30,819 1,257 —
✅ 📄 112_find_pvcs_by_uuid 22.8s 3 4 $0.2066 61,414 60,151 21,938 1,263 687 38,209 21,942 321 —
✅ 📄 12_job_crashing 40.8s 6 14 $0.3065 135,405 132,850 25,128 2,555 1,044 106,874 25,976 97 —
✅ 📄 176_network_policy_blocking_traffic_no_skills 45.6s 5 13 $0.3118 111,127 108,293 25,982 2,834 952 80,306 27,987 420 —
✅ 📄 227_count_configmaps_per_namespace[0] 20.7s 4 9 $0.2065 76,825 75,698 20,680 1,127 593 54,403 21,295 53 —
✅ 📄 243_pod_names_contain_service 36.6s 5 10 $0.2648 105,450 103,327 23,197 2,123 892 79,553 23,774 172 —
✅ 📄 24_misconfigured_pvc 50.9s 6 15 $0.3173 132,835 129,802 24,683 3,033 977 103,839 25,963 365 —
✅ 📄 43_current_datetime_from_prompt 4.9s 1 — $0.1195 17,049 16,940 16,940 109 109 0 16,940 68 —
✅ 📄 51_logs_summarize_errors 25.1s 4 5 $0.2063 77,465 76,325 21,019 1,140 416 55,301 21,024 42 —
✅ 📄 61_exact_match_counting 10.8s 3 3 $0.1522 52,853 52,490 17,918 363 216 34,568 17,922 32 —
Total 35.4s avg 4.7 avg 10.3 avg $2.8570 1,134,364 1,111,946 30,809 22,418 1,044 852,980 258,966 3,066 —

Benchmark comparison unavailable: No ci-benchmark experiments found

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No ci-benchmark experiments found

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #3 · Run @ __db109ab__ (#25930374521) — May 15, 17:06 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit db109ab on branch claude/add-eval-source-emoji-DsZft

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 📄 38.6s 5 11 $0.2767 107,704 105,383 23,690 2,321 947 80,671 24,712 202 —
✅ 101_loki_historical_logs_pod_deleted 📄 73.3s 7 17 $0.4181 172,530 167,859 29,576 4,671 1,023 136,219 31,640 843 —
✅ 112_find_pvcs_by_uuid 📄 19.4s 3 4 $0.2047 61,101 59,921 21,820 1,180 586 37,846 22,075 246 —
✅ 12_job_crashing 📄 40.5s 6 14 $0.3305 150,213 147,716 27,896 2,497 935 118,992 28,724 95 —
✅ 176_network_policy_blocking_traffic_no_skills 📄 60.3s 6 18 $0.3614 141,119 137,540 27,189 3,579 800 107,253 30,287 677 —
✅ 227_count_configmaps_per_namespace[0] 📄 19.8s 4 9 $0.2104 76,963 75,824 20,747 1,139 597 53,825 21,999 53 —
✅ 243_pod_names_contain_service 📄 33.2s 4 10 $0.2534 83,461 81,326 23,091 2,135 955 57,452 23,874 259 —
✅ 24_misconfigured_pvc 📄 43.5s 5 15 $0.2954 108,224 105,418 24,082 2,806 1,049 79,726 25,692 199 —
✅ 43_current_datetime_from_prompt 📄 5.5s 1 — $0.1204 17,085 16,940 16,940 145 145 0 16,940 99 —
✅ 51_logs_summarize_errors 📄 23.6s 4 5 $0.2103 77,922 76,687 21,201 1,235 451 55,481 21,206 81 —
✅ 61_exact_match_counting 📄 11.6s 3 3 $0.1522 52,847 52,484 17,914 363 216 34,566 17,918 32 —
Total 33.6s avg 4.4 avg 10.6 avg $2.8334 1,049,169 1,027,098 29,576 22,071 1,049 762,031 265,067 2,786 —

Benchmark comparison unavailable: No ci-benchmark experiments found

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No ci-benchmark experiments found

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #2 · Run @ __8e45250__ (#25929901763) — May 15, 16:58 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 8e45250 on branch claude/add-eval-source-emoji-DsZft

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 📄 36.3s 4 11 $0.2851 90,298 87,884 25,696 2,414 1,051 60,600 27,284 291 —
✅ 101_loki_historical_logs_pod_deleted 📄 72.4s 7 17 $0.4044 170,100 165,531 28,222 4,569 923 135,448 30,083 994 —
✅ 112_find_pvcs_by_uuid 📄 21.6s 3 4 $0.2088 61,504 60,222 22,005 1,282 687 37,963 22,259 333 —
✅ 12_job_crashing 📄 45.1s 6 15 $0.3395 145,140 142,311 27,158 2,829 1,058 112,796 29,515 171 —
✅ 176_network_policy_blocking_traffic_no_skills 📄 62.5s 7 15 $0.3714 164,916 161,074 27,244 3,842 885 132,848 28,226 845 —
✅ 227_count_configmaps_per_namespace[0] 📄 21.5s 4 9 $0.2096 76,851 75,726 20,696 1,125 592 53,792 21,934 54 —
✅ 243_pod_names_contain_service 📄 42.5s 5 11 $0.2866 110,572 108,084 24,473 2,488 944 82,816 25,268 253 —
✅ 24_misconfigured_pvc 📄 40.1s 5 12 $0.2746 104,852 102,421 22,974 2,431 816 78,208 24,213 256 —
✅ 43_current_datetime_from_prompt 📄 5.2s 1 — $0.1200 17,069 16,940 16,940 129 129 0 16,940 86 —
✅ 51_logs_summarize_errors 📄 25.1s 4 5 $0.2099 77,817 76,582 21,146 1,235 512 55,431 21,151 39 —
✅ 61_exact_match_counting 📄 10.8s 3 3 $0.1522 52,853 52,490 17,918 363 216 34,568 17,922 32 —
Total 34.8s avg 4.5 avg 10.2 avg $2.8622 1,071,972 1,049,265 28,222 22,707 1,058 784,470 264,795 3,354 —

Benchmark comparison unavailable: No ci-benchmark experiments found

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No ci-benchmark experiments found

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #1 · Run @ __a17981e__ (#25929779500) — May 15, 16:51 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit a17981e on branch claude/add-eval-source-emoji-DsZft

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 49.3s 6 13 $0.3202 141,889 139,248 26,580 2,641 963 112,091 27,157 228 —
✅ 101_loki_historical_logs_pod_deleted 70.2s 6 15 $0.3553 139,799 135,969 26,581 3,830 1,030 107,957 28,012 648 —
✅ 112_find_pvcs_by_uuid 22.8s 3 4 $0.2063 61,247 60,030 21,905 1,217 592 37,868 22,162 289 —
✅ 12_job_crashing 41.8s 5 11 $0.2934 120,344 118,273 26,759 2,071 759 91,076 27,197 121 —
✅ 176_network_policy_blocking_traffic_no_skills 57.3s 6 14 $0.3255 136,981 133,919 26,071 3,062 765 107,294 26,625 485 —
✅ 227_count_configmaps_per_namespace[0] 23.1s 4 9 $0.2069 76,905 75,775 20,720 1,130 599 54,427 21,348 61 —
✅ 243_pod_names_contain_service 44.8s 5 11 $0.2839 110,092 107,723 24,352 2,369 883 82,257 25,466 253 —
✅ 24_misconfigured_pvc 56.0s 6 15 $0.3187 132,623 129,547 24,643 3,076 1,013 103,498 26,049 378 —
✅ 43_current_datetime_from_prompt 5.4s 1 — $0.1200 17,069 16,940 16,940 129 129 0 16,940 86 —
✅ 51_logs_summarize_errors 28.0s 4 5 $0.2102 78,220 77,033 21,372 1,187 405 55,656 21,377 84 —
✅ 61_exact_match_counting 13.3s 3 3 $0.1522 52,855 52,492 17,918 363 216 34,570 17,922 32 —
Total 37.5s avg 4.5 avg 10.0 avg $2.7927 1,068,024 1,046,949 26,759 21,075 1,030 786,694 260,255 2,665 —

Benchmark comparison unavailable: No ci-benchmark experiments found

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No ci-benchmark experiments found

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit 3e0bc94 on branch claude/add-eval-source-emoji-DsZft

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 📄 09_crashpod 39.1s 5 11 $0.2897 110,254 107,794 25,088 2,460 865 81,887 25,907 312 —
✅ 📄 101_loki_historical_logs_pod_deleted 86.7s 7 18 $0.4582 187,044 181,959 31,981 5,085 1,023 146,746 35,213 1,335 —
✅ 📄 112_find_pvcs_by_uuid 19.2s 3 4 $0.2044 61,155 59,999 21,873 1,156 590 37,868 22,131 251 —
✅ 📄 12_job_crashing 40.1s 6 13 $0.3407 148,539 146,208 27,517 2,331 767 114,216 31,992 108 —
✅ 📄 176_network_policy_blocking_traffic_no_skills 53.5s 6 17 $0.3532 141,773 138,181 27,337 3,592 1,130 109,735 28,446 576 —
✅ 📄 227_count_configmaps_per_namespace[0] 20.5s 4 9 $0.2038 76,863 75,727 20,697 1,136 604 55,025 20,702 63 —
✅ 📄 243_pod_names_contain_service 36.5s 4 10 $0.2667 85,678 83,509 23,800 2,169 938 57,560 25,949 189 —
✅ 📄 24_misconfigured_pvc 36.4s 5 12 $0.2693 104,622 102,323 23,028 2,299 981 78,448 23,875 218 —
✅ 📄 43_current_datetime_from_prompt 4.5s 1 — $0.1200 17,069 16,940 16,940 129 129 0 16,940 86 —
✅ 📄 51_logs_summarize_errors 21.6s 4 5 $0.2068 77,804 76,695 21,203 1,109 385 55,487 21,208 42 —
✅ 📄 61_exact_match_counting 11.4s 3 3 $0.1522 52,851 52,488 17,916 363 216 34,568 17,920 32 —
Total 33.6s avg 4.4 avg 10.2 avg $2.8651 1,063,652 1,041,823 31,981 21,829 1,130 771,540 270,283 3,212 —

Benchmark comparison unavailable: No ci-benchmark experiments found

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No ci-benchmark experiments found

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/add-eval-source-emoji-DsZft -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/add-eval-source-emoji-DsZft -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented May 15, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: c8bcd779-4a36-492e-a8d0-fd33c0f82d2b

📥 Commits

Reviewing files that changed from the base of the PR and between 3f3275c and 3e0bc94.

📒 Files selected for processing (1)
  • tests/llm/utils/reporting/github_reporter.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • tests/llm/utils/reporting/github_reporter.py

Walkthrough

Adds a helper mapping test types to fixture directories and builds branch-aware GitHub blob URLs to tests/llm/fixtures/.../test_case.yaml, then inserts a 📄 markdown link into comparison and main results tables when a fixture URL exists.

Changes

Eval Source URL Integration

Layer / File(s) Summary
Source URL helper and import
tests/llm/utils/reporting/github_reporter.py
Adds quote import; introduces _TEST_TYPE_TO_FIXTURE_DIR mapping for fixture-backed test types (ask, investigate) and _get_eval_source_url(test_type, test_case_name) which resolves Git ref from EVAL_BRANCH → GITHUB_HEAD_REF → GITHUB_REF_NAME → "master", URL-encodes the ref, and returns a GitHub blob URL to the test_case.yaml file or None.
Report table rendering updates
tests/llm/utils/reporting/github_reporter.py
Comparison-table and main results-table rendering compute source_url via _get_eval_source_url(...) and append (comparison) or prefix (main results) a 📄 markdown link to the displayed test case name when a URL is available.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

  • HolmesGPT/holmesgpt#1267: Modifies tests/llm/utils/reporting/github_reporter.py rendering logic and overlaps with changes to the test-case column and report formatting.

Suggested reviewers

  • Sheeproid
  • arikalon1
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Add branch-specific eval source links to dashboard' directly and clearly describes the main change: adding branch-specific links to evaluation source files in the dashboard interface.
Docstring Coverage ✅ Passed Docstring coverage is 87.50% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented May 15, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 2d83c822 (built in 1m 18s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:2d83c822
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:2d83c822 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:2d83c822
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:2d83c822
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:2d83c822
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:2d83c822 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:2d83c822
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:2d83c822

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:2d83c822 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:2d83c822

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:2d83c822 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:2d83c822

@netlify

netlify Bot commented May 15, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 3e0bc94
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a0754ab9cae28000808c122
😎 Deploy Preview https://deploy-preview-2049--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

The github_reporter is what generates the actual GitHub PR comment
(evals_report.md). Adds a 📄 link next to each eval name in both the
main results table and the comparison sub-tables, pointing to the
eval's test_case.yaml on the branch this run was executed from
(EVAL_BRANCH / GITHUB_REF_NAME / BUILDKITE_BRANCH, falling back to
master). Maps test_type → fixture directory (ask → test_ask_holmes,
investigate → test_investigate).

Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (1)
tests/generate_eval_report.py (1)

884-896: ⚠️ Potential issue | 🔴 Critical | 🏗️ Heavy lift

Missing test_type information at call site.

Same issue as the heatmap table: the call to get_eval_source_url(eval_case) on line 887 doesn't pass test_type, so URLs will be incorrect for non-"ask" test types. The data structure needs to be updated to track test_type for each eval_case.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/generate_eval_report.py` around lines 884 - 896, The call to
get_eval_source_url(eval_case) doesn't pass test_type so non-"ask" URLs are
wrong; update the data structure that iterates eval_case to also carry its
test_type and change the call sites here to get_eval_source_url(eval_case,
test_type) (and likewise pass test_type into get_braintrust_eval_filter_url if
needed) so name_cell is built with the correct branch_source_url and
eval_filter_url; search for other uses of
get_eval_source_url/get_braintrust_eval_filter_url in this file and adjust their
signatures/usages to accept the new test_type parameter and update the data
construction where eval_case entries are created to include test_type.
🧹 Nitpick comments (1)
tests/generate_eval_report.py (1)

241-252: ⚡ Quick win

Consider extracting URL generation logic to a shared utility.

The URL generation logic for eval source files is now duplicated between generate_eval_report.py and github_reporter.py (lines 28-47 in that file), with subtle differences in implementation. Consider extracting this to a shared utility module (e.g., tests/llm/utils/eval_urls.py) to maintain consistency and reduce duplication.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/generate_eval_report.py` around lines 241 - 252, The URL-building logic
in get_eval_source_url duplicates similar code in github_reporter (see its URL
generation block) causing subtle inconsistencies; extract this logic into a
single shared utility (e.g., tests/llm/utils/eval_urls.py) that exposes a
function (e.g., build_eval_source_url or get_eval_source_url) and update both
get_eval_source_url in generate_eval_report.py and the corresponding code in
github_reporter to call the new utility so they use identical
encoding/ref-fallback behavior and path construction.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/generate_eval_report.py`:
- Around line 773-786: The URL is wrong because get_eval_source_url(eval_case)
lacks the test_type; update the data collection so each eval_case entry in
eval_model_stats includes its test_type (add a test_type field where
eval_model_stats is populated), change calls to get_eval_source_url to pass that
test_type (e.g., get_eval_source_url(eval_case, test_type)), and update
get_eval_source_url's signature/logic to accept and use test_type; then retrieve
test_type from eval_model_stats when building name_cell so branch_source_url is
constructed correctly for non-"ask" tests.

---

Duplicate comments:
In `@tests/generate_eval_report.py`:
- Around line 884-896: The call to get_eval_source_url(eval_case) doesn't pass
test_type so non-"ask" URLs are wrong; update the data structure that iterates
eval_case to also carry its test_type and change the call sites here to
get_eval_source_url(eval_case, test_type) (and likewise pass test_type into
get_braintrust_eval_filter_url if needed) so name_cell is built with the correct
branch_source_url and eval_filter_url; search for other uses of
get_eval_source_url/get_braintrust_eval_filter_url in this file and adjust their
signatures/usages to accept the new test_type parameter and update the data
construction where eval_case entries are created to include test_type.

---

Nitpick comments:
In `@tests/generate_eval_report.py`:
- Around line 241-252: The URL-building logic in get_eval_source_url duplicates
similar code in github_reporter (see its URL generation block) causing subtle
inconsistencies; extract this logic into a single shared utility (e.g.,
tests/llm/utils/eval_urls.py) that exposes a function (e.g.,
build_eval_source_url or get_eval_source_url) and update both
get_eval_source_url in generate_eval_report.py and the corresponding code in
github_reporter to call the new utility so they use identical
encoding/ref-fallback behavior and path construction.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 38f9942d-8eac-4c9e-b256-6a2641127e33

📥 Commits

Reviewing files that changed from the base of the PR and between 31fa24c and 8e45250.

📒 Files selected for processing (2)
  • tests/generate_eval_report.py
  • tests/llm/utils/reporting/github_reporter.py

Comment thread tests/generate_eval_report.py Outdated
claude added 2 commits May 15, 2026 16:56
This reverts commit a17981e.

Signed-off-by: Claude <noreply@anthropic.com>
- On pull_request events GITHUB_REF_NAME is the virtual "<num>/merge"
  ref (e.g. 2049/merge) which doesn't exist as a browsable branch.
  Prefer GITHUB_HEAD_REF (the PR head branch) when set, falling back
  to GITHUB_REF_NAME for push events.
- Move the 📄 link to the start of the eval name cell so the column
  is easier to scan.

Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/llm/utils/reporting/github_reporter.py`:
- Around line 53-55: The GitHub URL construction inserts test_case_name raw
which can break links; update the code that builds the URL (the string
containing encoded_ref, fixture_dir and test_case_name) to percent-encode
test_case_name just like encoded_ref (e.g., create encoded_test_case via
urllib.parse.quote(..., safe='/') or similar) and use encoded_test_case in the
f-string instead of test_case_name so reserved characters are escaped.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 352ae3a7-d0d1-4e65-91ca-a05181d5b819

📥 Commits

Reviewing files that changed from the base of the PR and between 8e45250 and 3f3275c.

📒 Files selected for processing (1)
  • tests/llm/utils/reporting/github_reporter.py

Comment thread tests/llm/utils/reporting/github_reporter.py
The PR comment report only runs in GitHub Actions; we don't use
Buildkite for evals. Removes the dead fallback.

Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn enabled auto-merge (squash) May 15, 2026 17:17
@aantn
aantn merged commit f21be71 into master May 15, 2026
18 of 19 checks passed
@aantn
aantn deleted the claude/add-eval-source-emoji-DsZft branch May 15, 2026 17:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants