Skip to content

/evals dx improvements - #1303

Merged
aantn merged 6 commits into
masterfrom
claude/prepopulate-workflow-branch-Ig8Te
Jan 1, 2026
Merged

aantn merged 6 commits into
masterfrom
claude/prepopulate-workflow-branch-Ig8Te

Conversation

@aantn

@aantn aantn commented Jan 1, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • Added rerun links in evaluation comment tables for quick workflow re-execution
    • New command to list available evaluations and post a formatted comment
  • Improvements

    • Markers and eval names formatted for clearer, code-styled display and navigation
    • Evaluation comments and progress updates now include contextual info and an optional legend
  • Style

    • Historical comparison section header now renders in bold for clearer readability

✏️ Tip: You can customize this high-level summary in your review settings.

claude added 3 commits January 1, 2026 21:22
When clicking the "Trigger via GitHub Actions UI" link, the branch
dropdown is now pre-filled with the current PR/push branch using
the ref query parameter.

Signed-off-by: Claude <noreply@anthropic.com>
The params table now shows "View logs | Rerun" in the Workflow row.
The Rerun link goes to the workflow dispatch page with the branch
pre-selected, making it easy to trigger a new run.

Signed-off-by: Claude <noreply@anthropic.com>
- Change valid markers and eval names from bullet lists to comma-separated
  clickable links that go to the workflow dispatch page
- Add prominent warning that /eval comments always use the workflow from
  master, with link to trigger from PR branch instead
- This helps users who modify the GitHub Action (e.g., adding secrets)
  understand why their changes don't take effect with /eval

Signed-off-by: Claude <noreply@anthropic.com>
@linux-foundation-easycla

linux-foundation-easycla Bot commented Jan 1, 2026 •

Copy link
Copy Markdown

CLA Not Signed

@coderabbitai

coderabbitai Bot commented Jan 1, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Rate limit exceeded

@aantn has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 1 minutes and 16 seconds before requesting another review.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

📥 Commits

Reviewing files that changed from the base of the PR and between f3b506f and a68a22f.

📒 Files selected for processing (1)
  • .github/scripts/eval-comment-helpers.js

Walkthrough

Comment-rendering helpers were updated to accept an optional context so workflow URLs and rerun links can include branch refs; markers/eval-name outputs were collapsed to a comma-separated valid_markers output and workflows were adjusted to pass context into comment construction.

Changes

Cohort / File(s) Change Summary
Comment rendering helpers
.github/scripts/eval-comment-helpers.js
renderParamsTable(p) → renderParamsTable(p, context = null); buildRerunFooter(p, context) → buildRerunFooter(p, context, options = {}). Removed askHolmesEvals / investigateEvals from params. Added formatAsCodes helper, dynamic workflow URL/ref handling, conditional Legend via includeLegend, and render changes to include Rerun links and formatted marker lists.
Workflow: eval/regression
.github/workflows/eval-regression.yaml
Added list_evals job to produce marker list on /list. Consolidated per-eval outputs into a single comma-separated valid_markers output (removed separate ask/investigate outputs). Threaded context into multiple buildBody and buildRerunFooter calls; updated steps order and progress comments to use context-aware messages and include legend where requested.
Test reporting minor rendering
tests/llm/utils/reporting/github_reporter.py
Adjusted Markdown/HTML rendering for the collapsible "Historical Comparison Details" header to render the summary text in bold (UI-only change).

Sequence Diagram(s)

sequenceDiagram
  participant User as PR commenter
  participant GitHub as GitHub API
  participant Workflow as GH Actions (eval-regression)
  participant Helper as eval-comment-helpers.js
  participant Runner as Action Runner

  User->>GitHub: Post trigger comment (/eval or /list)
  GitHub->>Workflow: Start workflow (with event/context)
  Workflow->>Runner: Run jobs (list_evals / eval jobs)
  Runner->>Helper: buildBody(..., { context }) / buildRerunFooter(..., { context })
  Helper-->>Runner: Rendered comment body (includes workflow URL + rerun link if ref present)
  Runner->>GitHub: Post/Update comment on PR with rendered body
  Note over Helper,GitHub: Valid markers produced as comma-separated output for helpers
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested reviewers

  • Sheeproid
  • arikalon1

Pre-merge checks

✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately reflects the main changes: improvements to the developer experience of the /evals workflow, including enhanced linking, formatting, and context-aware features across the evaluation infrastructure.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jan 1, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for 907ba60 (built in 45s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:907ba60
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:907ba60 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:907ba60
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:907ba60

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:907ba60

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:907ba60

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

🧹 Nitpick comments (2)
.github/scripts/eval-comment-helpers.js (2)

97-112: Consider handling empty items after filtering.

The function handles !items but could return an empty string if all items are empty after splitting/trimming (e.g., input is just commas). This would render as nothing in the markdown.

🔎 Suggested defensive check
 function formatAsLinks(items, workflowUrl, paramType) {
   if (!items) return '_(loading...)_';
-  return items.split(',').map(item => {
+  const links = items.split(',').map(item => {
     const trimmed = item.trim();
     if (!trimmed) return '';
     // Link to workflow dispatch - user will need to enter the value manually
     return `[\`${trimmed}\`](${workflowUrl})`;
   }).filter(Boolean).join(', ');
+  return links || '_(none found)_';
 }

104-111: The paramType parameter is unused.

The function accepts paramType (documented as 'markers' or 'filter') but never uses it. Either remove it or implement the intended differentiation.

🔎 Option 1: Remove unused parameter
-function formatAsLinks(items, workflowUrl, paramType) {
+function formatAsLinks(items, workflowUrl) {

Then update the call sites (lines 125-127) to remove the third argument.

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 0584639 and aeb29ba.

📒 Files selected for processing (2)
  • .github/scripts/eval-comment-helpers.js
  • .github/workflows/eval-regression.yaml
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
  • GitHub Check: build (3.12)
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.10)
  • GitHub Check: build
  • GitHub Check: llm_evals
🔇 Additional comments (8)
.github/scripts/eval-comment-helpers.js (4)

52-68: LGTM - Context-aware rerun link addition.

The optional context parameter with default null maintains backward compatibility. URL construction properly uses encodeURIComponent for the branch name to prevent URL injection issues.


77-80: LGTM - Context threading through buildBody.

Properly passes extras.context to renderParamsTable for manual runs, enabling the rerun link feature.


120-127: LGTM - Consistent URL construction with renderParamsTable.

The base workflow URL and branch handling logic matches renderParamsTable, ensuring consistent behavior across the comment.


140-143: LGTM - Clear warning about workflow source.

The warning about /eval using the workflow from master is helpful for developers who modify the GitHub Action and expect their changes to take effect.

.github/workflows/eval-regression.yaml (4)

338-338: LGTM - Context properly passed to buildBody.

The { context } object is correctly passed as the extras parameter, enabling the rerun link generation in the comment.


393-406: LGTM - Comma-separated format for JS consumption.

The bash commands correctly produce comma-separated lists that the formatAsLinks helper expects. The tr '\n' ',' followed by sed 's/,$//' properly handles the trailing comma.


456-456: LGTM - Consistent context threading.

Both testPreview and context are properly passed in the extras object for the progress update.


564-568: LGTM - Final results include context for rerun links.

The context is correctly included in the buildBody call for the final evaluation results, ensuring the rerun link appears in the completed comment.

@github-actions

github-actions Bot commented Jan 1, 2026

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit aeb29ba on branch claude/prepopulate-workflow-branch-Ig8Te

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 29.7s ↓18% 5 13 $0.1567
✅ 101_loki_historical_logs_pod_deleted 46.8s ↓12% 7 15 $0.1932
✅ 111_pod_names_contain_service 37.3s ±0% 7 14 $0.1719
✅ 12_job_crashing 39.6s ↓19% 7 16 $0.1854
✅ 162_get_runbooks 49.3s ±0% 8 16 $0.2320
✅ 176_network_policy_blocking_traffic_no_runbooks 34.3s ↓14% 5 13 $0.1646
✅ 24_misconfigured_pvc 35.7s ±0% 6 16 $0.1711
✅ 43_current_datetime_from_prompt 3.3s ±0% 1 — $0.0621
✅ 61_exact_match_counting 10.6s ±0% 3 3 $0.0860
Total 31.9s avg 5.4 avg 13.2 avg $1.4230

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/prepopulate-workflow-branch-Ig8Te'

Status: Success - 16 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use Trigger via GitHub Actions UI instead, which runs the workflow from your PR branch.


Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /last to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, context_window, coralogix, counting, database, datadog, datetime, easy, embeds, grafana-dashboard, hard, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency

📋 Valid eval names (use with filter)

test_ask_holmes: 01_how_many_pods, 02_what_is_wrong_with_pod, 03_what_is_the_command_to_port_forward, 04_related_k8s_events, 05_image_version, 06_explain_issue, 07_high_latency, 08_sock_shop_frontend, 09_crashpod, 100a_loki_historical_logs, 101_loki_historical_logs_pod_deleted, 102_loki_label_discovery, 102a_loki_logs_transparency, 102b_loki_multiple_pods, 103_logs_transparency_default_limit, 104a_postgres_root_issue, 104b_postgres_missing_index_pgstat, 104c_postgres_minimal_missing_index, 105_redis_wrong_data_structure, 107_log_filter_http_status_code, 108_logs_nearby_lines, 109_logs_transparency_not_found, 10_image_pull_backoff, 110_cpu_graph_robusta_runner, 110_k8s_events_image_pull, 111_disabled_datadog_traces, 111_pod_names_contain_service, 111_tool_hallucination, 112_find_pvcs_by_uuid, 114_checkout_latency_tracing_rebuild, 115_checkout_errors_tracing, 117_new_relic_tracing, 117b_new_relic_block_embed, 118_new_relic_logs, 119_new_relic_metrics, 11_init_containers, 120_new_relic_traces2, 121_new_relic_checkout_errors_tracing, 122_new_relic_checkout_latency_tracing_rebuild, 123_new_relic_checkout_errors_tracing, 124_checkout_latency_prometheus, 12_job_crashing, 13a_pending_node_selector_basic, 13b_pending_node_selector_detailed, 14_pending_resources, 151_disabled_toolsets_fallback_only, 156_kafka_opensearch_latency, 157_disk_full_statefulset, 158_slack_chat_correct_date, 159_prometheus_high_cardinality_cpu, 15_failed_readiness_probe, 160_electricity_market_bidding_bug, 160a_cpu_per_namespace_graph, 160b_cpu_per_namespace_graph_with_prom_truncation, 160c_cpu_per_namespace_graph_with_global_truncation, 161_bidding_version_performance, 161_conversation_compaction, 162_get_runbooks, 163_compaction_follow_up, 164_datadog_traces_coupon_code, 165_alert_with_multiple_runbooks, 16_failed_no_toolset_found, 173_coralogix_logs, 174_coralogix_traces_ad, 175_coralogix_metrics_frontend, 176_network_policy_blocking_traffic_no_runbooks, 177_grafana_home_dashboard, 178_grafana_search_dashboard_query, 179_grafana_big_dashboard_query, 17_oom_kill, 180_connectivity_check_tcp, 181_connectivity_check_http, 182_connectivity_check_http_url, 18_oom_kill_from_issues_history, 19_detect_missing_app_details, 20_long_log_file_search, 21_job_fail_curl_no_svc_account, 22_high_latency_dbi_down, 23_app_error_in_current_logs, 24_misconfigured_pvc, 25_misconfigured_ingress_class, 26_page_render_times, 27a_multi_container_logs, 27b_multi_container_logs, 28_permissions_error, 30_basic_promql_graph_cluster_memory, 32_basic_promql_graph_pod_cpu, 33_cpu_metrics_discovery, 34_memory_graph, 35_tempo, 36_argocd_find_resource, 37_argocd_wrong_namespace, 38_rabbitmq_split_head, 39_failed_toolset, 41_setup_argo, 42_dns_issues_result_all_tools, 42_dns_issues_result_new_tools, 42_dns_issues_result_new_tools_no_runbook, 42_dns_issues_result_old_tools, 42_dns_issues_steps_new_all_tools, 42_dns_issues_steps_new_tools, 42_dns_issues_steps_old_tools, 43_current_datetime_from_prompt, 43_slack_deployment_logs, 44_slack_statefulset_logs, 45_fetch_deployment_logs_simple, 46_job_crashing_no_longer_exists, 47_truncated_logs_context_window, 48_logs_since_thursday, 49_logs_since_last_week, 50_logs_since_specific_date, 50a_logs_since_last_specific_month, 51_logs_summarize_errors, 52_logs_login_issues, 53_logs_find_term, 54_azure_sql, 54_not_truncated_when_getting_pods, 55_kafka_runbook, 57_cluster_name_confusion, 57_wrong_namespace, 58_counting_pods_by_status, 59_label_based_counting, 60_count_less_than, 61_exact_match_counting, 62_fetch_error_logs_with_errors, 63_fetch_error_logs_no_errors, 64_keda_vs_hpa_confusion, 65_health_check_followup, 66_http_error_needle, 67_performance_degradation, 68_cascading_failures, 69_rate_limit_exhaustion, 70_memory_leak_detection, 71_connection_pool_starvation, 73a_time_window_anomaly, 73b_time_window_anomaly, 74_config_change_impact, 75_network_flapping, 76_service_discovery_issue, 77_liveness_probe_misconfiguration, 78a_missing_cpu_limits, 78b_cpu_quota_exceeded, 79_configmap_mount_issue, 80_pvc_storage_class_mismatch, 81_service_account_permission_denied, 82_pod_anti_affinity_conflict, 83_secret_not_found, 84_network_policy_blocking_traffic, 85_hpa_not_scaling, 86_configmap_like_but_secret, 89_runbook_missing_cloudwatch, 90_runbook_basic_selection, 91a_datadog_metrics_missing_namespace, 91b_datadog_metrics_pod_exists, 91c_datadog_metrics_deployment, 91d_datadog_metrics_historical_pod, 91e_datadog_custom_metrics, 91f_datadog_logs_historical_pod, 91g_datadog_metrics_mismatched_pod, 91h_datadog_logs_empty_query_with_url, 91i_datadog_metrics_empty_query_with_url, 92_cpu_graph_conversation, 93_calling_datadog, 93_events_since_specific_date, 94_runbook_transparency, 95_runbook_memory_leak_detection, 96_no_matching_runbook, 97_logs_clarification_needed, 99_logs_transparency_custom_time

test_investigate: 01_oom_kill, 02_crashloop_backoff, 03_cpu_throttling, 04_image_pull_backoff, 05_crashpod, 06_job_failure, 07_job_syntax_error, 08_memory_pressure, 09_high_latency, 10_KubeDeploymentReplicasMismatch, 11_KubePodCrashLooping, 12_KubePodNotReady, 13_Watchdog, 14_tempo, 15_dns_resolution, 16_dns_resolution_no_tool, 17_investigate_correct_date

- Add /list command to show available eval names on demand
- Replace workflow link with gh CLI command in the master warning
- Change markers from linked list to plain comma-separated format
- Remove eval names from footer (use /list instead)
- Clean up unused ask_holmes_evals and investigate_evals params

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Jan 1, 2026

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit fd5992c on branch claude/prepopulate-workflow-branch-Ig8Te

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 36.1s ±0% 6 13 $0.1639
✅ 101_loki_historical_logs_pod_deleted 50.9s ±0% 7 15 $0.2086
✅ 111_pod_names_contain_service 44.8s ±0% 7 15 $0.1774
✅ 12_job_crashing 41.9s ↓15% 6 16 $0.1805
✅ 162_get_runbooks 54.9s ↑10% 8 17 $0.2405
✅ 176_network_policy_blocking_traffic_no_runbooks 33.7s ↓16% 5 14 $0.1558
✅ 24_misconfigured_pvc 46.5s ↑22% 8 18 $0.1940
✅ 43_current_datetime_from_prompt 3.8s ±0% 1 — $0.0621
✅ 61_exact_match_counting 12.5s ±0% 3 3 $0.0860
Total 36.1s avg 5.7 avg 13.9 avg $1.4689

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/prepopulate-workflow-branch-Ig8Te'

Status: Success - 16 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/prepopulate-workflow-branch-Ig8Te -f markers=regression

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /last to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, context_window, coralogix, counting, database, datadog, datetime, easy, embeds, grafana-dashboard, hard, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency

- Add includeLegend option to buildRerunFooter (default false)
- Only show Legend section when final results are displayed
- Bold "Historical Comparison Details" for consistency with other
  collapsible sections

Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
.github/workflows/eval-regression.yaml (1)

447-470: Consider adding format validation for marker parsing.

The marker parsing assumes pyproject.toml markers follow the "name:description" format and splits on :. While this is fine for internal configuration, consider whether the script should validate the marker format or handle edge cases (e.g., markers without colons).

Optional: Add defensive parsing
 VALID_MARKERS=$(python3 << 'PYEOF'
 import tomllib
 with open('pyproject.toml', 'rb') as f:
     config = tomllib.load(f)
 markers = config.get('tool', {}).get('pytest', {}).get('ini_options', {}).get('markers', [])
-names = [m.split(':')[0] for m in sorted(markers) if m.split(':')[0] != 'llm']
+names = []
+for m in sorted(markers):
+    parts = m.split(':', 1)
+    if parts[0] and parts[0] != 'llm':
+        names.append(parts[0])
 print(','.join(names))
 PYEOF
 )

This validates that the marker name exists before adding it to the list.

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between aeb29ba and f3b506f.

📒 Files selected for processing (3)
  • .github/scripts/eval-comment-helpers.js
  • .github/workflows/eval-regression.yaml
  • tests/llm/utils/reporting/github_reporter.py
🧰 Additional context used
📓 Path-based instructions (2)
**/*.py

📄 CodeRabbit inference engine (CLAUDE.md)

**/*.py: Use Ruff for formatting and linting (configured in pyproject.toml)
Type hints required (mypy configuration in pyproject.toml)
ALWAYS place Python imports at the top of the file, not inside functions or methods

Files:

  • tests/llm/utils/reporting/github_reporter.py
tests/**/*.py

📄 CodeRabbit inference engine (CLAUDE.md)

Tests: match source structure under tests/

Files:

  • tests/llm/utils/reporting/github_reporter.py
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
  • GitHub Check: llm_evals
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.12)
🔇 Additional comments (6)
tests/llm/utils/reporting/github_reporter.py (1)

67-67: LGTM!

The bold formatting improves visual hierarchy in the collapsible details section.

.github/scripts/eval-comment-helpers.js (2)

47-66: LGTM!

The context-aware workflow link generation is implemented correctly, with proper URL encoding for the branch parameter.


116-171: LGTM!

The context-aware rerun footer implementation is well-structured:

  • Proper optional parameters with defaults
  • Conditional legend rendering improves UX
  • The warning about /eval using master workflow is valuable for users testing workflow changes
.github/workflows/eval-regression.yaml (3)

39-93: LGTM!

The new /list command implementation is well-designed:

  • Proper permission checks restrict access to authorized users
  • Graceful error handling for missing directories
  • Clear output format with code-styled eval names

393-393: LGTM!

The context parameter is consistently propagated through all buildBody and buildRerunFooter calls. The conditional includeLegend option is correctly applied only to the final results comment where the legend is relevant.

Also applies to: 425-425, 495-495, 528-528, 599-606


483-483: LGTM!

The valid_markers output is consistently passed to buildParams across all progress updates and final results, ensuring markers are available for display in the rerun footer.

Also applies to: 516-516, 585-585

Comment thread .github/scripts/eval-comment-helpers.js
@github-actions

github-actions Bot commented Jan 1, 2026

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit f3b506f on branch claude/prepopulate-workflow-branch-Ig8Te

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 27.7s ↓22% 5 12 $0.1509
✅ 101_loki_historical_logs_pod_deleted 63.6s ↑23% 10 24 $0.2852
✅ 111_pod_names_contain_service 35.1s ↓13% 7 15 $0.1725
✅ 12_job_crashing 33.2s ↓32% 6 14 $0.1642
✅ 162_get_runbooks 45.1s ±0% 8 16 $0.2229
✅ 176_network_policy_blocking_traffic_no_runbooks 31.9s ↓18% 5 13 $0.1105
✅ 24_misconfigured_pvc 32.0s ↓12% 6 17 $0.1671
✅ 43_current_datetime_from_prompt 3.1s ↓13% 1 — $0.0621
✅ 61_exact_match_counting 9.9s ↓11% 3 3 $0.0872
Total 31.3s avg 5.7 avg 14.2 avg $1.4225

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/prepopulate-workflow-branch-Ig8Te'

Status: Success - 18 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/prepopulate-workflow-branch-Ig8Te -f markers=regression

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /last to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, context_window, coralogix, counting, database, datadog, datetime, easy, embeds, grafana-dashboard, hard, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency

Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn enabled auto-merge (squash) January 1, 2026 22:25
@github-actions

github-actions Bot commented Jan 1, 2026

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit a68a22f on branch claude/prepopulate-workflow-branch-Ig8Te

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 40.1s ↑13% 6 14 $0.1664
✅ 101_loki_historical_logs_pod_deleted 56.1s ±0% 8 18 $0.2204
✅ 111_pod_names_contain_service 43.4s ±0% 7 15 $0.1759
✅ 12_job_crashing 44.4s ±0% 7 17 $0.1842
✅ 162_get_runbooks 61.5s ↑26% 7 22 $0.2389
✅ 176_network_policy_blocking_traffic_no_runbooks 47.4s ↑22% 7 15 $0.1938
✅ 24_misconfigured_pvc 43.1s ↑18% 7 17 $0.1823
✅ 43_current_datetime_from_prompt 4.0s ↑12% 1 — $0.0621
✅ 61_exact_match_counting 12.3s ±0% 3 3 $0.0860
Total 39.1s avg 5.9 avg 15.1 avg $1.5101

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/prepopulate-workflow-branch-Ig8Te'

Status: Success - 18 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/prepopulate-workflow-branch-Ig8Te -f markers=regression

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /last to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, context_window, coralogix, counting, database, datadog, datetime, easy, embeds, grafana-dashboard, hard, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /last · /list

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants