Skip to content

Minor tweaks to evals - #1266

Merged
aantn merged 25 commits into
masterfrom
claude/fix-eval-regression-marker-XeZ7i
Dec 29, 2025
Merged

aantn merged 25 commits into
masterfrom
claude/fix-eval-regression-marker-XeZ7i

Conversation

@aantn

@aantn aantn commented Dec 29, 2025 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • Improvements
    • Comments now include a fuller Legend table clarifying icon meanings; the trailing legend was removed from generated reports.
    • Re-run footer and manual re-run instructions appear more broadly and are reorganized into nested, clearer sections.
    • Marker and eval-name lists show clearer collection placeholders (e.g., "Collecting from pyproject.toml..." and "Collecting from tests/llm/fixtures/...") while data loads.

✏️ Tip: You can customize this high-level summary in your review settings.

- Remove 'regression' as default marker for /eval comments and
  workflow_dispatch triggers - they now run all LLM tests by default
- Keep 'regression' as default only for automatic triggers (PR/push)
- Add details section with list of valid markers and example test names
- Update marker_expr to handle empty markers (just 'llm' instead of
  'llm and ()')

Signed-off-by: Claude <noreply@anthropic.com>
- Add test preview step that runs pytest --collect-only to show which
  tests will run before actually running them
- Update initial comment with test count and expandable test list
- Add warning that manual re-runs have no default markers and will run
  all LLM tests (~100+) which can take 1+ hours
- Update example /eval command to include markers: regression
- Update markers description to emphasize no default

Signed-off-by: Claude <noreply@anthropic.com>
When a user triggers evals manually via /eval comment or workflow_dispatch,
they now receive a notification when the run completes. The notification:
- @mentions the user who triggered the eval
- Shows success or regression count status
- Points to the updated results comment above

This ensures users get a GitHub notification instead of having to poll
the PR for the updated comment.

Signed-off-by: Claude <noreply@anthropic.com>
- Rename step names and summary text from "Test preview" to "Evals to run"
- Remove the 20 test limit to show all evals that will run

Signed-off-by: Claude <noreply@anthropic.com>
- Remove "Results will appear here when complete." text from both
  initial and running status comments since it's confusing
- For manual triggers, don't create a new comment if comment_id is
  missing - the initial comment should always exist for manual runs

Signed-off-by: Claude <noreply@anthropic.com>
- Split into two separate <details> sections: Valid markers and Valid eval names
- List each marker and eval name on its own line
- Include the complete list of all 140+ ask_holmes evals and 17 investigate evals

Signed-off-by: Claude <noreply@anthropic.com>
grep -c returns exit code 1 when no matches found, even though it
outputs "0". Using || echo "0" would result in "0\n0" being captured.
Changed to || true which prevents failure without adding extra output.

Signed-off-by: Claude <noreply@anthropic.com>
- For PR triggers: show "PR #123 (abc1234)"
- For push triggers: show "push to branch-name (abc1234)"
- Include trigger info in both initial and updated comments

Signed-off-by: Claude <noreply@anthropic.com>
- Add progress checklist showing completion status of each step
- Update comment after each major step (HolmesGPT setup, collect evals, KIND ready)
- Reorder steps: HolmesGPT env -> Collect evals -> KIND cluster for optimization
- Show completed checklist in final results comment

Signed-off-by: Claude <noreply@anthropic.com>
When no model is specified in /eval comment, MODEL env var was set to
empty string instead of being unset. This caused get_models() to return
an empty list (since it only uses DEFAULT_MODEL when MODEL is unset,
not when it's empty string), resulting in 0 tests being collected.

Fix: Unset MODEL in bash if it's empty so Python uses DEFAULT_MODEL.
Signed-off-by: Claude <noreply@anthropic.com>
Replace hardcoded lists of valid markers and eval names with dynamic
collection at runtime:
- Markers: extracted from pyproject.toml [tool.pytest.ini_options] section
- Eval names: listed from tests/llm/fixtures/test_ask_holmes/ and test_investigate/

This ensures the comment always shows the current list of available
markers and tests without manual updates when new tests are added.

Signed-off-by: Claude <noreply@anthropic.com>
Refactor the 5 comment-generating steps to use consistent helper functions:
- renderProgress(): Progress checklist renderer
- renderParamsTable(): Parameters table for manual runs
- buildBody(): Main comment body builder with icon/title customization
- buildRerunFooter(): Re-run instructions for automatic runs (results only)

This ensures consistent styling between manual and automatic eval runs,
and between initial/progress/results comments. The helpers are defined
inline in each step (GitHub Actions limitation) but follow the same pattern.

Signed-off-by: Claude <noreply@anthropic.com>
- Create .github/scripts/eval-comment-helpers.js with reusable functions
- Update all 5 workflow steps to require() the helpers instead of copy-paste
- Remove redundant emoji from progress checklist items (⏳ and ✅ were redundant
  since checkbox state already indicates completion)

Signed-off-by: Claude <noreply@anthropic.com>
- Change PR trigger source from "PR #N (sha)" to just "commit sha"
  (PR number is redundant since comment is on the PR)
- Remove eyes reaction when workflow completes, leaving only hooray

Signed-off-by: Claude <noreply@anthropic.com>
- Add buildParams() helper to centralize type conversions and defaults
- Update all 5 workflow steps to use buildParams() for cleaner code
- All transformations (=== 'true', || 'default', parseInt) in one place

Signed-off-by: Claude <noreply@anthropic.com>
- Output base_params JSON from eval-params step containing common params
- Each step now spreads base_params and only adds step-specific extras
- Reduces ~7 lines of repeated parameter passing per step

Signed-off-by: Claude <noreply@anthropic.com>
When triggering via GitHub Actions UI without specifying a PR number,
automatically find the open PR associated with the selected branch.
This makes it easier to run evals - just select the branch and go.

Signed-off-by: Claude <noreply@anthropic.com>
- Hide Progress section once run is complete (pass null to buildBody)
- Consolidate legend into single collapsible <details> section
- Change "Regression" to "Failure" throughout for clarity

Signed-off-by: Claude <noreply@anthropic.com>
- Add new Legend <details> section with table explaining each status icon
- Restore original 3 separate <details> sections (Re-run, Markers, Eval names)

Signed-off-by: Claude <noreply@anthropic.com>
…atic runs

Signed-off-by: Claude <noreply@anthropic.com>
Keep branch improvements:
- Legend section in its own details tag
- Separate details sections for Re-run, Markers, and Eval names
- Footer shown for both manual and automatic runs
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 8d85979

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅
ask 101_loki_historical_logs_pod_deleted ✅
ask 111_pod_names_contain_service ✅
ask 12_job_crashing ✅
ask 162_get_runbooks ✅
ask 176_network_policy_blocking_traffic_no_runbooks ✅
ask 24_misconfigured_pvc ✅
ask 43_current_datetime_from_prompt ✅
ask 61_exact_match_counting ✅

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • log_cli_level = "INFO"
  • log_file = "tests.log"
  • log_file_level = "INFO" # Changed from DEBUG to reduce noise in HTML report
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 100b_loki_historical_logs_nonstandard_label
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for 4c2aef2 (built in 40s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:4c2aef2
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:4c2aef2 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:4c2aef2
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:4c2aef2

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:4c2aef2

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:4c2aef2

@coderabbitai

coderabbitai Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

Walkthrough

buildRerunFooter’s Markdown was restructured (legend table, nested details, updated placeholders). The eval workflow imports and appends that rerun footer in more places (previous isManual gating removed) and threads expanded eval parameter fields between steps. The test report legend was removed.

Changes

Cohort / File(s) Summary
Eval comment footer restructuring
.github/scripts/eval-comment-helpers.js
buildRerunFooter output rewritten: adds a Legend table, replaces previous placeholder lines with “Collecting…” placeholders, nests “Re-run evals manually”, “Valid markers”, and “Valid eval names” in collapsible <details> blocks, and reorganizes detail tag placement. No public API/signature changes.
Workflow footer application & param propagation
.github/workflows/eval-regression.yaml
buildRerunFooter is destructured/imported and appended to more comment/update calls (removing prior p.isManual gating). Eval parameter payloads/outputs expanded to include valid_markers, ask_holmes_evals, investigate_evals, and test_preview, with these fields propagated across steps; a “Collect evals to run” step was removed/reordered.
Test reporting legend removal
tests/llm/utils/reporting/github_reporter.py
The final legend emoji/descriptions block appended after the results table was removed; core table/report generation logic unchanged.

Sequence Diagram(s)

sequenceDiagram
    autonumber
    participant WF as GitHub Actions workflow
    participant Helper as eval-comment-helpers.js
    participant Builder as buildBody (comment builder)
    participant GH as GitHub API (Comments)

    Note over WF,Helper: workflow now calls buildRerunFooter in more places
    WF->>Helper: buildRerunFooter(p, context)
    Helper-->>WF: Markdown footer (legend table + nested details)
    WF->>Builder: buildBody(..., footer, expanded params)
    Builder-->>WF: assembled comment body
    WF->>GH: create/update comment with assembled body
    GH-->>WF: comment created/updated
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested reviewers

  • Sheeproid
  • arikalon1

Pre-merge checks

❌ Failed checks (1 inconclusive)
Check name Status Explanation Resolution
Title check ❓ Inconclusive The title 'Minor tweaks to evals' is vague and generic, using non-descriptive language that does not convey meaningful information about the actual changes in the changeset. Provide a more specific title that describes the primary change, such as 'Refactor eval footer formatting and legend display' or 'Reorganize eval regression comment structure'.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

📜 Recent review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between c49f969 and dc8628e.

📒 Files selected for processing (1)
  • .github/scripts/eval-comment-helpers.js
🚧 Files skipped from review as they are similar to previous changes (1)
  • .github/scripts/eval-comment-helpers.js
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.12)
  • GitHub Check: llm_evals

Comment @coderabbitai help to get the list of available commands and usage tips.

- Add buildRerunFooter to all comment update steps
- Include valid_markers/evals data when available (after test-preview step)
- Earlier steps show placeholder text for markers/evals

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit ce87e8e

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅
ask 101_loki_historical_logs_pod_deleted ✅
ask 111_pod_names_contain_service ✅
ask 12_job_crashing ✅
ask 162_get_runbooks ✅
ask 176_network_policy_blocking_traffic_no_runbooks ✅
ask 24_misconfigured_pvc ✅
ask 43_current_datetime_from_prompt ✅
ask 61_exact_match_counting ✅

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • log_cli_level = "INFO"
  • log_file = "tests.log"
  • log_file_level = "INFO" # Changed from DEBUG to reduce noise in HTML report
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 100b_loki_historical_logs_nonstandard_label
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

Show "Loading... or see [source]" instead of "No markers found"

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 865f0ce

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅
ask 101_loki_historical_logs_pod_deleted ✅
ask 111_pod_names_contain_service ✅
ask 12_job_crashing ✅
ask 162_get_runbooks ✅
ask 176_network_policy_blocking_traffic_no_runbooks ✅
ask 24_misconfigured_pvc ✅
ask 43_current_datetime_from_prompt ✅
ask 61_exact_match_counting ✅

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • log_cli_level = "INFO"
  • log_file = "tests.log"
  • log_file_level = "INFO" # Changed from DEBUG to reduce noise in HTML report
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 100b_loki_historical_logs_nonstandard_label
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

Legend is now only in the footer details section

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit c49f969

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅
ask 101_loki_historical_logs_pod_deleted ✅
ask 111_pod_names_contain_service ✅
ask 12_job_crashing ✅
ask 162_get_runbooks ✅
ask 176_network_policy_blocking_traffic_no_runbooks ✅
ask 24_misconfigured_pvc ✅
ask 43_current_datetime_from_prompt ✅
ask 61_exact_match_counting ✅

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • log_cli_level = "INFO"
  • log_file = "tests.log"
  • log_file_level = "INFO" # Changed from DEBUG to reduce noise in HTML report
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 100b_loki_historical_logs_nonstandard_label
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit dc8628e

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅
ask 101_loki_historical_logs_pod_deleted ✅
ask 111_pod_names_contain_service ✅
ask 12_job_crashing ✅
ask 162_get_runbooks ✅
ask 176_network_policy_blocking_traffic_no_runbooks ✅
ask 24_misconfigured_pvc ✅
ask 43_current_datetime_from_prompt ✅
ask 61_exact_match_counting ✅

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • log_cli_level = "INFO"
  • log_file = "tests.log"
  • log_file_level = "INFO" # Changed from DEBUG to reduce noise in HTML report
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 100b_loki_historical_logs_nonstandard_label
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

@aantn

aantn commented Dec 29, 2025

Copy link
Copy Markdown
Collaborator Author

/eval
markers: regression
filter: 09_crashpod
iterations: 2

@aantn
aantn enabled auto-merge (squash) December 29, 2025 20:33
@github-actions

github-actions Bot commented Dec 29, 2025 •

Copy link
Copy Markdown
Contributor

🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Model bedrock/eu.anthropic.claude-sonnet-4-5-20250929-v1:0
Markers regression
Filter (-k) 09_crashpod
Iterations 2
Duration 1m 35s
Workflow View logs

Results of HolmesGPT evals

  • ask_holmes: 2/2 test cases were successful, 0 regressions
Test suite Test case Status
ask 09_crashpod ✅
ask 09_crashpod ✅

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ✅ Completed successfully

👆 See updated results above.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants