Skip to content

Improvements to manual evals run - #1265

Merged
aantn merged 18 commits into
masterfrom
claude/fix-eval-regression-marker-XeZ7i
Dec 29, 2025
Merged

aantn merged 18 commits into
masterfrom
claude/fix-eval-regression-marker-XeZ7i

Conversation

@aantn

@aantn aantn commented Dec 29, 2025 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • Richer evaluation comments: who triggered the run, incremental progress checklist, test preview, parameter table, dynamic rerun guidance, and dedicated completion comments with reaction updates.
    • Explicit KIND setup and clearer run sequencing with progress updates.
  • Improvements

    • Auto-detect PR from branch for manual runs; clearer manual vs automatic behavior.
    • Smarter marker/default handling, unified outputs (base params, triggered_by, marker expression) and consistent result posting.

✏️ Tip: You can customize this high-level summary in your review settings.

- Remove 'regression' as default marker for /eval comments and
  workflow_dispatch triggers - they now run all LLM tests by default
- Keep 'regression' as default only for automatic triggers (PR/push)
- Add details section with list of valid markers and example test names
- Update marker_expr to handle empty markers (just 'llm' instead of
  'llm and ()')

Signed-off-by: Claude <noreply@anthropic.com>
- Add test preview step that runs pytest --collect-only to show which
  tests will run before actually running them
- Update initial comment with test count and expandable test list
- Add warning that manual re-runs have no default markers and will run
  all LLM tests (~100+) which can take 1+ hours
- Update example /eval command to include markers: regression
- Update markers description to emphasize no default

Signed-off-by: Claude <noreply@anthropic.com>
When a user triggers evals manually via /eval comment or workflow_dispatch,
they now receive a notification when the run completes. The notification:
- @mentions the user who triggered the eval
- Shows success or regression count status
- Points to the updated results comment above

This ensures users get a GitHub notification instead of having to poll
the PR for the updated comment.

Signed-off-by: Claude <noreply@anthropic.com>
- Rename step names and summary text from "Test preview" to "Evals to run"
- Remove the 20 test limit to show all evals that will run

Signed-off-by: Claude <noreply@anthropic.com>
- Remove "Results will appear here when complete." text from both
  initial and running status comments since it's confusing
- For manual triggers, don't create a new comment if comment_id is
  missing - the initial comment should always exist for manual runs

Signed-off-by: Claude <noreply@anthropic.com>
- Split into two separate <details> sections: Valid markers and Valid eval names
- List each marker and eval name on its own line
- Include the complete list of all 140+ ask_holmes evals and 17 investigate evals

Signed-off-by: Claude <noreply@anthropic.com>
grep -c returns exit code 1 when no matches found, even though it
outputs "0". Using || echo "0" would result in "0\n0" being captured.
Changed to || true which prevents failure without adding extra output.

Signed-off-by: Claude <noreply@anthropic.com>
- For PR triggers: show "PR #123 (abc1234)"
- For push triggers: show "push to branch-name (abc1234)"
- Include trigger info in both initial and updated comments

Signed-off-by: Claude <noreply@anthropic.com>
- Add progress checklist showing completion status of each step
- Update comment after each major step (HolmesGPT setup, collect evals, KIND ready)
- Reorder steps: HolmesGPT env -> Collect evals -> KIND cluster for optimization
- Show completed checklist in final results comment

Signed-off-by: Claude <noreply@anthropic.com>
When no model is specified in /eval comment, MODEL env var was set to
empty string instead of being unset. This caused get_models() to return
an empty list (since it only uses DEFAULT_MODEL when MODEL is unset,
not when it's empty string), resulting in 0 tests being collected.

Fix: Unset MODEL in bash if it's empty so Python uses DEFAULT_MODEL.
Signed-off-by: Claude <noreply@anthropic.com>
Replace hardcoded lists of valid markers and eval names with dynamic
collection at runtime:
- Markers: extracted from pyproject.toml [tool.pytest.ini_options] section
- Eval names: listed from tests/llm/fixtures/test_ask_holmes/ and test_investigate/

This ensures the comment always shows the current list of available
markers and tests without manual updates when new tests are added.

Signed-off-by: Claude <noreply@anthropic.com>
Refactor the 5 comment-generating steps to use consistent helper functions:
- renderProgress(): Progress checklist renderer
- renderParamsTable(): Parameters table for manual runs
- buildBody(): Main comment body builder with icon/title customization
- buildRerunFooter(): Re-run instructions for automatic runs (results only)

This ensures consistent styling between manual and automatic eval runs,
and between initial/progress/results comments. The helpers are defined
inline in each step (GitHub Actions limitation) but follow the same pattern.

Signed-off-by: Claude <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

Walkthrough

Workflow and comment tooling updated: eval-regression GitHub Actions now compute trigger context, conditional marker expressions, and incremental progress comments; they auto-detect PRs from branches for manual runs and expose new outputs. A new helper script builds normalized params and dynamic comment bodies/footers.

Changes

Cohort / File(s) Summary
Eval workflow YAML
​.github/workflows/eval-regression.yaml
Reworked inputs/defaults (markers default empty), added base_params and triggered_by outputs, conditional marker_expr generation, improved trigger/pr detection (auto-detect PR from branch), reordered steps (HolmesGPT setup, collect evals, KIND), dynamic progress/comment updates, unset MODEL when empty, and completion reactions/comments.
Comment helpers script
​.github/scripts/eval-comment-helpers.js
New module exporting buildParams, renderProgress, renderParamsTable, buildBody, and buildRerunFooter to normalize action inputs/outputs and generate dynamic comment bodies, progress checklists, parameter tables, test previews, and rerun instructions.

Sequence Diagram(s)

sequenceDiagram
  autonumber
  actor User
  participant GH as "GitHub (events/API)"
  participant WF as "eval-regression workflow"
  participant Helper as "eval-comment-helpers.js"
  participant Runner as "Actions runner"
  participant KIND as "KIND cluster"
  participant Evals as "Eval jobs"

  User->>GH: trigger (issue_comment / workflow_dispatch / push / PR)
  GH->>WF: start workflow (context)
  WF->>WF: compute trigger_source, triggered_by, pr_number, marker_expr, base_params
  WF->>Helper: buildBody(progress=init, params)
  Helper-->>WF: initial comment body
  WF->>GH: create/update PR comment
  WF->>Runner: setup HolmesGPT env
  Runner-->>WF: env ready
  WF->>Helper: buildBody(progress=collected, params)
  Helper-->>WF: updated body
  WF->>KIND: setup KIND cluster
  KIND-->>WF: KIND ready
  WF->>Evals: run evals (apply marker_expr, unset MODEL if empty)
  Evals-->>WF: results
  WF->>Helper: buildBody(progress=done, extras=results)
  Helper-->>WF: final comment body (and rerun footer if automatic)
  WF->>GH: update comment, remove/add reactions, post completion comment if manual
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • Sheeproid
  • moshemorad

Pre-merge checks

❌ Failed checks (1 inconclusive)
Check name Status Explanation Resolution
Title check ❓ Inconclusive The title 'Improvements to manual evals run' is vague and generic, using non-descriptive terms that don't convey specific technical changes. Consider a more specific title that highlights the main change, such as: 'Refactor eval workflow to dynamically generate PR comments' or 'Update eval regression workflow with helper functions for comment generation'.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

📜 Recent review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 9b8a501 and 959f740.

📒 Files selected for processing (2)
  • .github/scripts/eval-comment-helpers.js
  • .github/workflows/eval-regression.yaml
🧰 Additional context used
🧠 Learnings (1)
📚 Learning: 2025-12-21T13:17:57.170Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include verification steps in workflow to confirm each diagnostic action was successful before proceeding

Applied to files:

  • .github/workflows/eval-regression.yaml
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
  • GitHub Check: llm_evals
  • GitHub Check: build (3.12)
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.10)
🔇 Additional comments (19)
.github/scripts/eval-comment-helpers.js (4)

12-30: LGTM!

The type conversions and defaults are handled correctly. The parseInt calls properly specify radix 10, and the boolean conversion from GitHub Actions string output is appropriate.


37-41: LGTM!

The progress rendering logic is correct and produces proper markdown checkbox syntax.


48-57: LGTM!

The conditional rows and fallback text for empty markers align correctly with the workflow behavior.


66-84: LGTM!

The conditional rendering logic correctly handles manual vs automatic runs, optional progress steps (null = hide), and optional test preview sections.

.github/workflows/eval-regression.yaml (15)

110-121: LGTM!

The comment parsing logic correctly extracts key-value pairs and applies validation via the validate() function.


123-136: LGTM!

The triggeredBy field correctly captures the comment author's login for later use in the completion notification.


148-169: Auto-detection works for same-repo branches.

The PR auto-detection logic correctly queries for open PRs matching the current branch. Note that line 156 constructs the head as owner:branch, which means it only finds PRs from the same repository (not forks). This is acceptable since fork PRs typically cannot trigger workflow_dispatch due to secret access restrictions.


177-189: LGTM!

The descriptive trigger_source values improve observability by showing commit SHAs and branch names for automatic runs.


204-218: LGTM!

The marker_expr conditional logic correctly constructs pytest marker expressions, and the base_params JSON output improves DRY by centralizing parameter passing to subsequent steps.


238-252: LGTM!

The dynamic comment generation using the helper module is correctly integrated. The progress steps clearly outline the evaluation workflow stages.


254-261: LGTM!

The reordering makes sense—HolmesGPT setup must occur before pytest can collect test names in the next step.


262-284: LGTM!

The progress update correctly marks the HolmesGPT setup step as complete and updates the existing comment.


285-334: LGTM!

The collection logic correctly gathers test names, valid markers, and eval names. The heredoc delimiters use random values to prevent injection, which is a good security practice.


335-391: LGTM!

The progress updates correctly show incremental completion and conditionally include the test preview when tests are found.


416-432: LGTM!

Unsetting empty MODEL ensures the default model is used as intended, and the duration formatting logic correctly handles various time ranges.


434-471: LGTM!

The results posting logic correctly hides the progress section (line 451 passes null) for completed runs and only adds the rerun footer for automatic runs. The "Failure" terminology aligns with the PR objectives.


473-494: LGTM!

The completion notification provides good UX by directly mentioning the user who triggered the manual run and summarizing the result status.


495-516: LGTM!

The reaction updates provide clear visual feedback: removing "eyes" (processing) and adding "hooray" (completion) helps users track the eval lifecycle.


518-526: LGTM!

The terminology update from "regressions" to "failures" aligns with the PR objectives for improved clarity.


Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by PR #1265 (efee49a)

View workflow logs

Progress:

  • Setup HolmesGPT environment
  • Collect evals to run
  • Setup KIND cluster
  • Run evals ✅

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅
ask 101_loki_historical_logs_pod_deleted ✅
ask 111_pod_names_contain_service ✅
ask 12_job_crashing ✅
ask 162_get_runbooks ✅
ask 176_network_policy_blocking_traffic_no_runbooks ✅
ask 24_misconfigured_pvc ✅
ask 43_current_datetime_from_prompt ✅
ask 61_exact_match_counting ✅

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • log_cli_level = "INFO"
  • log_file = "tests.log"
  • log_file_level = "INFO" # Changed from DEBUG to reduce noise in HTML report
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 100b_loki_historical_logs_nonstandard_label
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for fe7a182 (built in 47s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:fe7a182
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:fe7a182 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:fe7a182
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:fe7a182

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:fe7a182

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:fe7a182

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (2)
.github/workflows/eval-regression.yaml (2)

204-241: Consider extracting helper functions to avoid duplication.

These helper functions (renderProgress, renderParamsTable, buildBody) are duplicated verbatim in 5 steps throughout the workflow (lines 284-312, 394-422, 463-491, 569-597). While GitHub Actions JavaScript steps don't support sharing code directly, consider:

  1. Composite action: Create a reusable composite action that handles progress updates
  2. Inline template string: Define the helpers once and pass them via outputs (limited but possible)
  3. External script: Move to a .github/scripts/ file and use actions/github-script with script from file

This would reduce ~200 lines of duplication and make future updates less error-prone.


358-364: Fragile marker extraction from pyproject.toml.

The sed/grep pipeline on line 360 assumes a specific format for [tool.pytest.ini_options] in pyproject.toml. If the file format changes (e.g., markers in a different section, different quoting), this could silently fail.

Consider adding a fallback or validation:

-          VALID_MARKERS=$(sed -n '/^\[tool\.pytest\.ini_options\]/,/^\[/p' pyproject.toml | grep -E '"[a-zA-Z]' | sed 's/.*"\([^:]*\):.*/\1/' | grep -v '^llm$' | sort | awk '{print "- \x60" $0 "\x60"}')
+          VALID_MARKERS=$(sed -n '/^\[tool\.pytest\.ini_options\]/,/^\[/p' pyproject.toml | grep -E '"[a-zA-Z]' | sed 's/.*"\([^:]*\):.*/\1/' | grep -v '^llm$' | sort | awk '{print "- \x60" $0 "\x60"}')
+          if [[ -z "$VALID_MARKERS" ]]; then
+            echo "::warning::Could not extract markers from pyproject.toml"
+          fi
📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e6ae36b and efee49a.

📒 Files selected for processing (1)
  • .github/workflows/eval-regression.yaml
🧰 Additional context used
🧠 Learnings (3)
📚 Learning: 2025-12-21T13:17:57.170Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Workflow section with numbered sequential steps containing Action, Function Description, Parameters, Expected Output, and Success/Failure Criteria

Applied to files:

  • .github/workflows/eval-regression.yaml
📚 Learning: 2025-12-21T13:17:57.170Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include verification steps in workflow to confirm each diagnostic action was successful before proceeding

Applied to files:

  • .github/workflows/eval-regression.yaml
📚 Learning: 2025-12-21T13:17:57.170Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Use generic function descriptions in workflow steps (e.g., 'execute a command to test network connectivity') rather than tool-specific names to enable mapping to available tools

Applied to files:

  • .github/workflows/eval-regression.yaml
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
  • GitHub Check: llm_evals
  • GitHub Check: build (3.12)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.11)
  • GitHub Check: build
🔇 Additional comments (8)
.github/workflows/eval-regression.yaml (8)

18-21: Clear input description update.

The updated description accurately reflects the new behavior where empty markers runs all LLM tests. This aligns with the changed defaulting semantics for manual vs automatic triggers.


110-121: Consistent default handling in parseComment.

The change to initialize markers: '' aligns with the new behavior where manual triggers don't automatically get the regression marker.


123-167: Well-structured trigger attribution logic.

The triggeredBy variable correctly captures the initiating user for manual runs (comment author or workflow actor) while leaving it empty for automatic triggers. The descriptive trigger sources with SHA snippets provide useful context.


169-184: Intentional behavior shift for manual runs.

The conditional defaulting ensures manual /eval and workflow_dispatch don't automatically filter to regression tests. The warning in the re-run footer (line 603-604) appropriately alerts users about the broader test scope.


544-546: Consistent MODEL handling.

Good addition to unset MODEL when empty, matching the behavior in the "Collect evals" step. This ensures get_models() properly falls back to DEFAULT_MODEL.


599-624: Excellent user guidance in re-run footer.

The detailed instructions with warnings about running all tests, options table, and collapsible sections for valid markers/evals significantly improve the user experience for re-running evals.


677-697: Helpful completion notification for manual runs.

The @mention notification ensures the triggering user is alerted when their eval run completes, with clear status indication. The condition chain correctly limits this to manual runs with a known triggerer.


660-663: Appropriate conditional for re-run footer.

The re-run instructions are correctly shown only for automatic runs. Users who manually triggered the eval already know the process.

Comment thread .github/workflows/eval-regression.yaml Outdated
@aantn

aantn commented Dec 29, 2025

Copy link
Copy Markdown
Collaborator Author

/eval
markers: regression
filter: 09_crashpod
iterations: 1

@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Model bedrock/eu.anthropic.claude-sonnet-4-5-20250929-v1:0
Markers regression
Filter (-k) 09_crashpod
Iterations 1
Duration 1m 42s
Workflow View logs

Results of HolmesGPT evals

  • ask_holmes: 1/1 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

- Create .github/scripts/eval-comment-helpers.js with reusable functions
- Update all 5 workflow steps to require() the helpers instead of copy-paste
- Remove redundant emoji from progress checklist items (⏳ and ✅ were redundant
  since checkbox state already indicates completion)

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by PR #1265 (1536571)

View workflow logs

Progress:

  • Setup HolmesGPT environment
  • Collect evals to run
  • Setup KIND cluster
  • Run evals

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅
ask 101_loki_historical_logs_pod_deleted ✅
ask 111_pod_names_contain_service ✅
ask 12_job_crashing ✅
ask 162_get_runbooks ✅
ask 176_network_policy_blocking_traffic_no_runbooks ✅
ask 24_misconfigured_pvc ✅
ask 43_current_datetime_from_prompt ✅
ask 61_exact_match_counting ✅

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • log_cli_level = "INFO"
  • log_file = "tests.log"
  • log_file_level = "INFO" # Changed from DEBUG to reduce noise in HTML report
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 100b_loki_historical_logs_nonstandard_label
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

- Change PR trigger source from "PR #N (sha)" to just "commit sha"
  (PR number is redundant since comment is on the PR)
- Remove eyes reaction when workflow completes, leaving only hooray

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 7bc70b3

View workflow logs

Progress:

  • Setup HolmesGPT environment
  • Collect evals to run
  • Setup KIND cluster
  • Run evals

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅
ask 101_loki_historical_logs_pod_deleted ✅
ask 111_pod_names_contain_service ✅
ask 12_job_crashing ✅
ask 162_get_runbooks ✅
ask 176_network_policy_blocking_traffic_no_runbooks ✅
ask 24_misconfigured_pvc ✅
ask 43_current_datetime_from_prompt ✅
ask 61_exact_match_counting ✅

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • log_cli_level = "INFO"
  • log_file = "tests.log"
  • log_file_level = "INFO" # Changed from DEBUG to reduce noise in HTML report
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 100b_loki_historical_logs_nonstandard_label
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between efee49a and 7bc70b3.

📒 Files selected for processing (2)
  • .github/scripts/eval-comment-helpers.js
  • .github/workflows/eval-regression.yaml
🧰 Additional context used
🧠 Learnings (1)
📓 Common learnings
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include conditional logic (IF/ELSE) in workflow steps when branching is required based on diagnostic findings
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Workflow section with numbered sequential steps containing Action, Function Description, Parameters, Expected Output, and Success/Failure Criteria
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Use generic function descriptions in workflow steps (e.g., 'execute a command to test network connectivity') rather than tool-specific names to enable mapping to available tools
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
  • GitHub Check: llm_evals
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.12)
  • GitHub Check: build (3.10)
🔇 Additional comments (12)
.github/scripts/eval-comment-helpers.js (4)

11-15: LGTM!

The progress rendering logic is clean and straightforward.


22-31: LGTM!

The conditional parameter table rendering handles optional fields appropriately.


40-56: LGTM!

The body builder correctly differentiates between manual and automatic runs and handles optional test previews.


64-88: Function logic is sound.

The rerun footer construction is well-structured and provides comprehensive instructions.

Note: The past review comment regarding misleading footer messages in fork PRs still applies to lines 80, 84, and 86 where the fallback text may be displayed even when collection was skipped.

.github/workflows/eval-regression.yaml (8)

123-183: LGTM!

The parameter extraction and validation logic correctly handles different trigger sources (comment, workflow_dispatch, automatic) and properly constructs the triggered_by and marker_expr outputs.


204-273: LGTM!

The integration with buildBody is correct, and the reordering of setup steps (HolmesGPT before collecting evals) is logical since pytest is required for collection.


274-323: Test collection logic is sound.

The collection of test names, markers, and eval names from the filesystem is well-implemented.

Note: Line 275's condition && steps.eval-params.outputs.pr_number != '' means this step skips for push events and fork PRs. The outputs are then used in the footer (lines 460-462) with || '' fallbacks, which relates to the existing review comment about misleading footer messages.


324-397: LGTM!

The incremental progress updates provide good user feedback and correctly incorporate test counts and conditional test previews.


441-498: Results posting logic is well-structured.

The conditional logic for manual vs automatic runs is correct, and the rerun footer is appropriately limited to automatic runs only.

Note: Lines 460-462 are specifically mentioned in the existing review comment regarding misleading footer messages in fork PRs.


500-521: LGTM!

The user notification step provides helpful feedback by mentioning the user who triggered the manual eval and summarizing the results.


527-543: LGTM!

The reaction cleanup logic provides better UX by removing the temporary "eyes" reaction before adding the completion "hooray" reaction.


423-424: No action needed—Python code correctly handles unset MODEL.

The os.environ.get("MODEL", DEFAULT_MODEL) in tests/llm/utils/test_env_vars.py already provides the fallback to DEFAULT_MODEL when the MODEL environment variable is unset. The get_models() function uses this variable and works as intended.

Comment thread .github/workflows/eval-regression.yaml
- Add buildParams() helper to centralize type conversions and defaults
- Update all 5 workflow steps to use buildParams() for cleaner code
- All transformations (=== 'true', || 'default', parseInt) in one place

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit ecfe41e

View workflow logs

Progress:

  • Setup HolmesGPT environment
  • Collect evals to run
  • Setup KIND cluster
  • Run evals

⚠️ No eval report was generated.


🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • log_cli_level = "INFO"
  • log_file = "tests.log"
  • log_file_level = "INFO" # Changed from DEBUG to reduce noise in HTML report
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 100b_loki_historical_logs_nonstandard_label
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

- Output base_params JSON from eval-params step containing common params
- Each step now spreads base_params and only adds step-specific extras
- Reduces ~7 lines of repeated parameter passing per step

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 452b283

View workflow logs

Progress:

  • Setup HolmesGPT environment
  • Collect evals to run
  • Setup KIND cluster
  • Run evals

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅
ask 101_loki_historical_logs_pod_deleted ✅
ask 111_pod_names_contain_service ✅
ask 12_job_crashing ✅
ask 162_get_runbooks ✅
ask 176_network_policy_blocking_traffic_no_runbooks ✅
ask 24_misconfigured_pvc ✅
ask 43_current_datetime_from_prompt ✅
ask 61_exact_match_counting ✅

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • log_cli_level = "INFO"
  • log_file = "tests.log"
  • log_file_level = "INFO" # Changed from DEBUG to reduce noise in HTML report
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 100b_loki_historical_logs_nonstandard_label
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

When triggering via GitHub Actions UI without specifying a PR number,
automatically find the open PR associated with the selected branch.
This makes it easier to run evals - just select the branch and go.

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 9b8a501

View workflow logs

Progress:

  • Setup HolmesGPT environment
  • Collect evals to run
  • Setup KIND cluster
  • Run evals

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅
ask 101_loki_historical_logs_pod_deleted ✅
ask 111_pod_names_contain_service ✅
ask 12_job_crashing ✅
ask 162_get_runbooks ✅
ask 176_network_policy_blocking_traffic_no_runbooks ✅
ask 24_misconfigured_pvc ✅
ask 43_current_datetime_from_prompt ✅
ask 61_exact_match_counting ✅

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • log_cli_level = "INFO"
  • log_file = "tests.log"
  • log_file_level = "INFO" # Changed from DEBUG to reduce noise in HTML report
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 100b_loki_historical_logs_nonstandard_label
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

🧪 Manual Eval Results

Parameter Value
Triggered via workflow_dispatch
Model bedrock/eu.anthropic.claude-sonnet-4-5-20250929-v1:0
Markers all LLM tests
Filter (-k) grafana
Iterations 2
Duration 2m 40s
Workflow View logs

Progress:

  • Setup HolmesGPT environment
  • Collect evals to run
  • Setup KIND cluster
  • Run evals

Results of HolmesGPT evals

  • ask_holmes: 4/6 test cases were successful, 2 regressions
Test suite Test case Status
ask 177_grafana_home_dashboard ✅
ask 177_grafana_home_dashboard ✅
ask 178_grafana_search_dashboard_query ✅
ask 178_grafana_search_dashboard_query ✅
ask 179_grafana_big_dashboard_query ❌
ask 179_grafana_big_dashboard_query ❌

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

⚠️ 2 Regressions Detected

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ⚠️ Completed with 2 regressions

👆 See updated results above.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

♻️ Duplicate comments (2)
.github/workflows/eval-regression.yaml (2)

19-21: This breaking change has already been flagged.

As noted in the previous review, the empty default for markers combined with the conditional logic on line 194 means manual runs with no markers specified will execute all 100+ LLM tests. This is a significant change in behavior.


285-334: Undefined outputs when collection step is skipped.

The "Collect evals to run" step (line 286) only runs when steps.eval-params.outputs.pr_number != '', but the "Post evaluation results" step (line 435) runs with always() and references collection outputs (lines 445-447) that may not exist if collection was skipped.

This is related to the fork PR issue flagged in previous reviews. When pr_number is empty (e.g., workflow_dispatch without auto-detected PR), the collection step skips, but posting still attempts to use steps.test-preview.outputs.valid_markers, ask_holmes_evals, and investigate_evals, which will be undefined.

🔎 Proposed fix to handle missing outputs

Add fallback handling in the "Post evaluation results" step:

            const p = buildParams({
              ...JSON.parse(${{ toJSON(steps.eval-params.outputs.base_params) }}),
              comment_id: ${{ toJSON(steps.initial-comment.outputs.comment_id) }},
              duration: ${{ toJSON(steps.evals.outputs.duration) }},
-              valid_markers: ${{ toJSON(steps.test-preview.outputs.valid_markers) }},
-              ask_holmes_evals: ${{ toJSON(steps.test-preview.outputs.ask_holmes_evals) }},
-              investigate_evals: ${{ toJSON(steps.test-preview.outputs.investigate_evals) }}
+              valid_markers: ${{ toJSON(steps.test-preview.outputs.valid_markers) }} || '',
+              ask_holmes_evals: ${{ toJSON(steps.test-preview.outputs.ask_holmes_evals) }} || '',
+              investigate_evals: ${{ toJSON(steps.test-preview.outputs.investigate_evals) }} || ''
            });

Or, better yet, only pass these when collection ran:

            const p = buildParams({
              ...JSON.parse(${{ toJSON(steps.eval-params.outputs.base_params) }}),
              comment_id: ${{ toJSON(steps.initial-comment.outputs.comment_id) }},
              duration: ${{ toJSON(steps.evals.outputs.duration) }},
-              valid_markers: ${{ toJSON(steps.test-preview.outputs.valid_markers) }},
-              ask_holmes_evals: ${{ toJSON(steps.test-preview.outputs.ask_holmes_evals) }},
-              investigate_evals: ${{ toJSON(steps.test-preview.outputs.investigate_evals) }}
+              ...${{ steps.test-preview.outcome == 'success' }} && {
+                valid_markers: ${{ toJSON(steps.test-preview.outputs.valid_markers) }},
+                ask_holmes_evals: ${{ toJSON(steps.test-preview.outputs.ask_holmes_evals) }},
+                investigate_evals: ${{ toJSON(steps.test-preview.outputs.investigate_evals) }}
+              }
            });
🧹 Nitpick comments (1)
.github/workflows/eval-regression.yaml (1)

148-169: Consider failing early when multiple PRs are detected.

When multiple open PRs are found for a branch (lines 162-164), the workflow logs a warning but continues with prNumber = null, which means no comments will be posted. Consider failing the workflow with a clear error message asking the user to specify pr_number explicitly, rather than silently continuing without PR detection.

🔎 Proposed enhancement for multiple PR detection
                } else if (prs.data.length > 1) {
-                  core.warning(`Multiple open PRs found for branch ${branch}, please specify pr_number`);
-                  prNumber = null;
+                  core.setFailed(`Multiple open PRs found for branch ${branch}. Please specify pr_number explicitly in workflow inputs. Found PRs: ${prs.data.map(p => `#${p.number}`).join(', ')}`);
+                  return;
                } else {
📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between ecfe41e and 9b8a501.

📒 Files selected for processing (1)
  • .github/workflows/eval-regression.yaml
🧰 Additional context used
🧠 Learnings (1)
📚 Learning: 2025-12-21T13:17:57.170Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.170Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include verification steps in workflow to confirm each diagnostic action was successful before proceeding

Applied to files:

  • .github/workflows/eval-regression.yaml
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
  • GitHub Check: llm_evals
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.12)
  • GitHub Check: llm_evals
🔇 Additional comments (5)
.github/workflows/eval-regression.yaml (5)

208-218: LGTM! Good DRY improvement.

The base_params JSON output effectively reduces parameter repetition across subsequent steps while maintaining proper type safety and validation.


262-284: LGTM! Progress tracking improves user experience.

The new progress update steps provide clear feedback at each stage of the evaluation workflow. The conditional logic properly ensures updates only occur when a comment exists.

Also applies to: 335-359, 367-390


416-418: LGTM! Proper handling of empty MODEL.

Unsetting the MODEL environment variable when empty allows the Python code to use its default model selection logic, which is more robust than passing an empty string.


479-500: LGTM! Good user experience for manual runs.

The completion notification with @ mention provides clear feedback to users who triggered manual eval runs, including regression status.


506-522: LGTM! Reaction handling provides good feedback.

The logic to remove the "eyes" reaction and add "hooray" properly indicates workflow completion. The code defensively handles the case where the eyes reaction might not exist.

Note: The listForIssueComment API might paginate with many reactions, but this is unlikely to be an issue in practice.

- Hide Progress section once run is complete (pass null to buildBody)
- Consolidate legend into single collapsible <details> section
- Change "Regression" to "Failure" throughout for clarity

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 959f740

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅
ask 101_loki_historical_logs_pod_deleted ✅
ask 111_pod_names_contain_service ✅
ask 12_job_crashing ✅
ask 162_get_runbooks ✅
ask 176_network_policy_blocking_traffic_no_runbooks ✅
ask 24_misconfigured_pvc ✅
ask 43_current_datetime_from_prompt ✅
ask 61_exact_match_counting ✅

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

📖 Legend

🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • log_cli_level = "INFO"
  • log_file = "tests.log"
  • log_file_level = "INFO" # Changed from DEBUG to reduce noise in HTML report
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency

📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 100b_loki_historical_logs_nonstandard_label
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

@aantn
aantn enabled auto-merge (squash) December 29, 2025 19:49
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

🧪 Manual Eval Results

Parameter Value
Triggered via workflow_dispatch
Model bedrock/eu.anthropic.claude-sonnet-4-5-20250929-v1:0
Markers all LLM tests
Filter (-k) grafana
Iterations 1
Duration 2m 19s
Workflow View logs

Results of HolmesGPT evals

  • ask_holmes: 2/3 test cases were successful, 1 regressions
Test suite Test case Status
ask 177_grafana_home_dashboard ✅
ask 178_grafana_search_dashboard_query ✅
ask 179_grafana_big_dashboard_query ❌

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

⚠️ 1 Failure Detected

@aantn
aantn merged commit 89c093f into master Dec 29, 2025
11 of 12 checks passed
@aantn
aantn deleted the claude/fix-eval-regression-marker-XeZ7i branch December 29, 2025 19:53
@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ⚠️ Completed with 1 failure

👆 See updated results above.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants