Skip to content

Remove post-processing support - #1279

Merged
aantn merged 5 commits into
masterfrom
codex/linear-mention-rob-98-remove-post-processing-capability
Jan 1, 2026
Merged

aantn merged 5 commits into
masterfrom
codex/linear-mention-rob-98-remove-post-processing-capability

Conversation

@aantn

@aantn aantn commented Dec 31, 2025 •

Copy link
Copy Markdown
Collaborator

Summary

  • remove CLI, server, and LLM handling for the deprecated post-processing prompt path
  • clean up Helm values/templates and documentation that referenced post-processing configuration
  • update tests and add evals_per_branch.json with the recommended LLM regression eval for this branch

Testing

  • python -m compileall holmes tests server.py

Codex Task

Summary by CodeRabbit

  • Documentation

    • Removed post-processing prompt examples and guidance from reference and toolset docs.
  • Configuration

    • Removed post-processing prompt settings from CLI options and Helm values; environment variable no longer documented.
  • Chores

    • Cleaned up code paths by removing post-processing prompt plumbing.
  • Tests

    • Updated tests and test helpers to reflect removal of the post-processing prompt parameter.

✏️ Tip: You can customize this high-level summary in your review settings.

Signed-off-by: Codex <codex@openai.com>
@coderabbitai

coderabbitai Bot commented Dec 31, 2025 •

Copy link
Copy Markdown
Contributor

Walkthrough

This PR removes the post-processing prompt feature by deleting its environment variable, Helm values, documentation, function parameters, call sites, and tests that supplied it, and by adding a small test config file. No new runtime behavior beyond omission of post-processing remains.

Changes

Cohort / File(s) Summary
Documentation
docs/data-sources/custom-toolsets.md, docs/reference/environment-variables.md, docs/reference/helm-configuration.md
Removed references and sample text for the post-processing prompt and related Helm/config docs.
Helm chart & values
helm/holmes/templates/holmes.yaml, helm/holmes/values.yaml
Removed enablePostProcessing / postProcessingPrompt keys and the conditional env var injection for HOLMES_POST_PROCESSING_PROMPT.
Environment variables
holmes/common/env_vars.py
Removed the HOLMES_POST_PROCESSING_PROMPT environment variable declaration and default.
Core LLM / investigation logic
holmes/core/tool_calling_llm.py, holmes/core/investigation.py
Removed post-processing plumbing: parameters, helpers, and call sites; updated signatures for ToolCallingLLM methods and IssueInvestigator.investigate; added cost-extraction helpers in LLM call path.
CLI & interactive
holmes/main.py, holmes/interactive.py
Removed CLI option/parameter post_processing_prompt and all call-site usages (commands and interactive loop).
Server
server.py
Removed import/use of HOLMES_POST_PROCESSING_PROMPT and corresponding argument from ai.prompt_call in workload_health_check.
Tests
tests/test_interactive.py, tests/test_issue_investigator.py
Removed post_processing_prompt arguments and related imports from tests; adjusted call sites accordingly.
Test config
evals_per_branch.json
Added new JSON with a pytest command entry for LLM-marked easy tests.

Sequence Diagram(s)

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • Add runbook alerts to prompt #1205: Also changes holmes/core/investigation.py and holmes/core/tool_calling_llm.py, updating IssueInvestigator.investigate signatures and call sites.
  • Enhance prompt caching #916: Overlaps on LLM call flow and CLI/integration changes; touches the same modules and introduces cost/call-path adjustments.
  • improve cli output #572: Modifies run_interactive_loop and interactive-mode behavior, intersecting with the interactive signature changes here.

Suggested reviewers

  • moshemorad
  • arikalon1
  • Sheeproid

Pre-merge checks

✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Remove post-processing support' accurately and concisely describes the main change of systematically removing post-processing functionality across CLI, server, LLM handling, configuration, and tests.
Docstring Coverage ✅ Passed Docstring coverage is 81.82% which is sufficient. The required threshold is 80.00%.

📜 Recent review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between a8ca0aa and 4fda527.

📒 Files selected for processing (5)
  • helm/holmes/templates/holmes.yaml
  • helm/holmes/values.yaml
  • holmes/core/investigation.py
  • holmes/core/tool_calling_llm.py
  • server.py
💤 Files with no reviewable changes (4)
  • helm/holmes/templates/holmes.yaml
  • helm/holmes/values.yaml
  • holmes/core/investigation.py
  • server.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • holmes/core/tool_calling_llm.py
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
  • GitHub Check: llm_evals
  • GitHub Check: build
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.12)
  • GitHub Check: build (3.11)

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Dec 31, 2025 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 4fda527

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 29.5s 5 12 $0.0978
✅ 101_loki_historical_logs_pod_deleted 41.8s 6 11 $0.1222
✅ 111_pod_names_contain_service 35.2s 6 14 $0.1083
✅ 12_job_crashing 40.2s 7 16 $0.1307
✅ 162_get_runbooks 53.2s 8 18 $0.1754
✅ 176_network_policy_blocking_traffic_no_runbooks 46.3s 8 16 $0.1570
✅ 24_misconfigured_pvc 36.9s 7 15 $0.1175
✅ 43_current_datetime_from_prompt 3.4s 1 — $0.0085
✅ 61_exact_match_counting 10.5s 3 3 $0.0324
Total 33.0s avg 5.7 avg 13.1 avg $0.9498

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

@github-actions

github-actions Bot commented Dec 31, 2025 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for 5267acd (built in 36s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:5267acd
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:5267acd me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:5267acd
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:5267acd

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:5267acd

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:5267acd

@aantn

aantn commented Dec 31, 2025

Copy link
Copy Markdown
Collaborator Author

@copilot remove evals_per_branch.json

Copilot AI mentioned this pull request Dec 31, 2025
7 tasks

Copilot AI commented Dec 31, 2025

Copy link
Copy Markdown
Contributor

@aantn I've opened a new pull request, #1291, to work on those changes. Once the pull request is ready, I'll request review from you.

@github-actions

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit af79aa8 on branch codex/linear-mention-rob-98-remove-post-processing-capability

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 34.0s ±0% 5 11 $0.0985
✅ 101_loki_historical_logs_pod_deleted 48.4s ±0% 6 13 $0.1811
✅ 111_pod_names_contain_service 42.5s ±0% 7 14 $0.1173
✅ 12_job_crashing 44.1s ±0% 7 14 $0.1295
✅ 162_get_runbooks 55.5s ±0% 8 17 $0.2334
✅ 176_network_policy_blocking_traffic_no_runbooks 42.6s ±0% 6 14 $0.1787
✅ 24_misconfigured_pvc 32.5s ↓16% 6 15 $0.1569
✅ 43_current_datetime_from_prompt 7.5s ↑126% 1 — $0.0086
✅ 61_exact_match_counting 14.6s ↑38% 3 3 $0.0326
Total 35.8s avg 5.4 avg 12.6 avg $1.1366

Time/Cost columns compare each test+model pair against its own historical average from other branches (↑slower/costlier, ↓faster/cheaper). Historical data available for 14 unique test+model pairs.

Historical Comparison Details

Filter: excluding branch 'codex/linear-mention-rob-98-remove-post-processing-capability'

Status: Success - 14 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 180_connectivity_check_tcp
  • 181_connectivity_check_http
  • 182_connectivity_check_http_url
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

Removed `evals_per_branch.json` file that was added as part of the
post-processing removal PR but is no longer needed.

Checklist for the toolset:

- [ ] Toolset has unit tests where relevant
- [ ] Toolset has both ask_holmes and investigate evals
- [ ] Toolset has a documentation
- [ ] Toolset has the correct is_default flag
- [ ] Toolset returns a correct `get_example_config`
- [ ] Toolset does a live health check in addition to checking for
correct configuration
- [ ] Create a demo video (if relevant)

<!-- START COPILOT CODING AGENT TIPS -->
---

💡 You can make Copilot smarter by setting up custom instructions,
customizing its development environment and configuring Model Context
Protocol (MCP) servers. Learn more [Copilot coding agent
tips](https://gh.io/copilot-coding-agent-tips) in the docs.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: aantn <494087+aantn@users.noreply.github.com>
@CLAassistant

CLAassistant commented Dec 31, 2025 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
2 out of 3 committers have signed the CLA.

✅ aantn
✅ moshemorad
❌ Copilot
You have signed the CLA already but the status is still pending? Let us recheck it.

@github-actions

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit d7ea020 on branch codex/linear-mention-rob-98-remove-post-processing-capability

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 34.7s ±0% 6 13 $0.1643
✅ 101_loki_historical_logs_pod_deleted 69.7s ↑46% 11 24 $0.2809
✅ 111_pod_names_contain_service 43.3s ↑12% 8 16 $0.1883
✅ 12_job_crashing 35.9s ↓16% 5 14 $0.1703
✅ 162_get_runbooks 36.3s ↓31% 5 15 $0.1821
✅ 176_network_policy_blocking_traffic_no_runbooks 42.6s ±0% 7 15 $0.1908
✅ 24_misconfigured_pvc 39.0s ±0% 7 17 $0.1781
✅ 43_current_datetime_from_prompt 3.3s ±0% 1 — $0.0621
✅ 61_exact_match_counting 10.9s ±0% 3 3 $0.0872
Total 35.1s avg 5.9 avg 14.6 avg $1.5040

Time/Cost columns compare each test+model pair against its own historical average from other branches (↑slower/costlier, ↓faster/cheaper). Historical data available for 12 unique test+model pairs.

Historical Comparison Details

Filter: excluding branch 'codex/linear-mention-rob-98-remove-post-processing-capability'

Status: Success - 12 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 180_connectivity_check_tcp
  • 181_connectivity_check_http
  • 182_connectivity_check_http_url
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

@aantn

aantn commented Dec 31, 2025

Copy link
Copy Markdown
Collaborator Author

@copilot is additionalInstructions in the custom toolsets actually used anywhere? or should it be removed entirely

Copilot AI commented Dec 31, 2025

Copy link
Copy Markdown
Contributor

@aantn I've opened a new pull request, #1295, to work on those changes. Once the pull request is ready, I'll request review from you.

@linux-foundation-easycla

Copy link
Copy Markdown

CLA Not Signed

@aantn
aantn enabled auto-merge (squash) January 1, 2026 10:27
@aantn
aantn merged commit bec4d12 into master Jan 1, 2026
8 of 12 checks passed
@aantn
aantn deleted the codex/linear-mention-rob-98-remove-post-processing-capability branch January 1, 2026 10:28
@github-actions

github-actions Bot commented Jan 1, 2026

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit bfae6d8 on branch codex/linear-mention-rob-98-remove-post-processing-capability

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 41.0s ↑28% 6 13 $0.1684
✅ 101_loki_historical_logs_pod_deleted 61.1s ↑54% 9 20 $0.2228
✅ 111_pod_names_contain_service 49.0s ↑17% 7 17 $0.1773
✅ 12_job_crashing 68.2s ↑48% 9 21 $0.2408
✅ 162_get_runbooks 45.0s ±0% 6 16 $0.1976
✅ 176_network_policy_blocking_traffic_no_runbooks 44.7s ↑16% 7 15 $0.1949
✅ 24_misconfigured_pvc 45.2s ↑30% 7 16 $0.1717
✅ 43_current_datetime_from_prompt 4.3s ↑28% 1 — $0.0621
✅ 61_exact_match_counting 13.4s ↑29% 3 3 $0.0860
Total 41.3s avg 6.1 avg 15.1 avg $1.5215

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'codex/linear-mention-rob-98-remove-post-processing-capability'

Status: Success - 18 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: Manual re-runs have NO default markers and will run ALL LLM tests (~100+), which can take 1+ hours. Use markers: regression or filter: test_name to limit scope.

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /last to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers
  • chain-of-causation
  • compaction
  • context_window
  • coralogix
  • counting
  • database
  • datadog
  • datetime
  • easy
  • embeds
  • grafana-dashboard
  • hard
  • kafka
  • kubernetes
  • leaked-information
  • logs
  • loki
  • medium
  • metrics
  • network
  • newrelic
  • no-cicd
  • numerical
  • one-test
  • port-forward
  • prometheus
  • question-answer
  • regression
  • runbooks
  • slackbot
  • storage
  • toolset-limitation
  • traces
  • transparency
📋 Valid eval names (use with filter)

test_ask_holmes:

  • 01_how_many_pods
  • 02_what_is_wrong_with_pod
  • 03_what_is_the_command_to_port_forward
  • 04_related_k8s_events
  • 05_image_version
  • 06_explain_issue
  • 07_high_latency
  • 08_sock_shop_frontend
  • 09_crashpod
  • 100a_loki_historical_logs
  • 101_loki_historical_logs_pod_deleted
  • 102_loki_label_discovery
  • 102a_loki_logs_transparency
  • 102b_loki_multiple_pods
  • 103_logs_transparency_default_limit
  • 104a_postgres_root_issue
  • 104b_postgres_missing_index_pgstat
  • 104c_postgres_minimal_missing_index
  • 105_redis_wrong_data_structure
  • 107_log_filter_http_status_code
  • 108_logs_nearby_lines
  • 109_logs_transparency_not_found
  • 10_image_pull_backoff
  • 110_cpu_graph_robusta_runner
  • 110_k8s_events_image_pull
  • 111_disabled_datadog_traces
  • 111_pod_names_contain_service
  • 111_tool_hallucination
  • 112_find_pvcs_by_uuid
  • 114_checkout_latency_tracing_rebuild
  • 115_checkout_errors_tracing
  • 117_new_relic_tracing
  • 117b_new_relic_block_embed
  • 118_new_relic_logs
  • 119_new_relic_metrics
  • 11_init_containers
  • 120_new_relic_traces2
  • 121_new_relic_checkout_errors_tracing
  • 122_new_relic_checkout_latency_tracing_rebuild
  • 123_new_relic_checkout_errors_tracing
  • 124_checkout_latency_prometheus
  • 12_job_crashing
  • 13a_pending_node_selector_basic
  • 13b_pending_node_selector_detailed
  • 14_pending_resources
  • 151_disabled_toolsets_fallback_only
  • 156_kafka_opensearch_latency
  • 157_disk_full_statefulset
  • 158_slack_chat_correct_date
  • 159_prometheus_high_cardinality_cpu
  • 15_failed_readiness_probe
  • 160_electricity_market_bidding_bug
  • 160a_cpu_per_namespace_graph
  • 160b_cpu_per_namespace_graph_with_prom_truncation
  • 160c_cpu_per_namespace_graph_with_global_truncation
  • 161_bidding_version_performance
  • 161_conversation_compaction
  • 162_get_runbooks
  • 163_compaction_follow_up
  • 164_datadog_traces_coupon_code
  • 165_alert_with_multiple_runbooks
  • 16_failed_no_toolset_found
  • 173_coralogix_logs
  • 174_coralogix_traces_ad
  • 175_coralogix_metrics_frontend
  • 176_network_policy_blocking_traffic_no_runbooks
  • 177_grafana_home_dashboard
  • 178_grafana_search_dashboard_query
  • 179_grafana_big_dashboard_query
  • 17_oom_kill
  • 180_connectivity_check_tcp
  • 181_connectivity_check_http
  • 182_connectivity_check_http_url
  • 18_oom_kill_from_issues_history
  • 19_detect_missing_app_details
  • 20_long_log_file_search
  • 21_job_fail_curl_no_svc_account
  • 22_high_latency_dbi_down
  • 23_app_error_in_current_logs
  • 24_misconfigured_pvc
  • 25_misconfigured_ingress_class
  • 26_page_render_times
  • 27a_multi_container_logs
  • 27b_multi_container_logs
  • 28_permissions_error
  • 30_basic_promql_graph_cluster_memory
  • 32_basic_promql_graph_pod_cpu
  • 33_cpu_metrics_discovery
  • 34_memory_graph
  • 35_tempo
  • 36_argocd_find_resource
  • 37_argocd_wrong_namespace
  • 38_rabbitmq_split_head
  • 39_failed_toolset
  • 41_setup_argo
  • 42_dns_issues_result_all_tools
  • 42_dns_issues_result_new_tools
  • 42_dns_issues_result_new_tools_no_runbook
  • 42_dns_issues_result_old_tools
  • 42_dns_issues_steps_new_all_tools
  • 42_dns_issues_steps_new_tools
  • 42_dns_issues_steps_old_tools
  • 43_current_datetime_from_prompt
  • 43_slack_deployment_logs
  • 44_slack_statefulset_logs
  • 45_fetch_deployment_logs_simple
  • 46_job_crashing_no_longer_exists
  • 47_truncated_logs_context_window
  • 48_logs_since_thursday
  • 49_logs_since_last_week
  • 50_logs_since_specific_date
  • 50a_logs_since_last_specific_month
  • 51_logs_summarize_errors
  • 52_logs_login_issues
  • 53_logs_find_term
  • 54_azure_sql
  • 54_not_truncated_when_getting_pods
  • 55_kafka_runbook
  • 57_cluster_name_confusion
  • 57_wrong_namespace
  • 58_counting_pods_by_status
  • 59_label_based_counting
  • 60_count_less_than
  • 61_exact_match_counting
  • 62_fetch_error_logs_with_errors
  • 63_fetch_error_logs_no_errors
  • 64_keda_vs_hpa_confusion
  • 65_health_check_followup
  • 66_http_error_needle
  • 67_performance_degradation
  • 68_cascading_failures
  • 69_rate_limit_exhaustion
  • 70_memory_leak_detection
  • 71_connection_pool_starvation
  • 73a_time_window_anomaly
  • 73b_time_window_anomaly
  • 74_config_change_impact
  • 75_network_flapping
  • 76_service_discovery_issue
  • 77_liveness_probe_misconfiguration
  • 78a_missing_cpu_limits
  • 78b_cpu_quota_exceeded
  • 79_configmap_mount_issue
  • 80_pvc_storage_class_mismatch
  • 81_service_account_permission_denied
  • 82_pod_anti_affinity_conflict
  • 83_secret_not_found
  • 84_network_policy_blocking_traffic
  • 85_hpa_not_scaling
  • 86_configmap_like_but_secret
  • 89_runbook_missing_cloudwatch
  • 90_runbook_basic_selection
  • 91a_datadog_metrics_missing_namespace
  • 91b_datadog_metrics_pod_exists
  • 91c_datadog_metrics_deployment
  • 91d_datadog_metrics_historical_pod
  • 91e_datadog_custom_metrics
  • 91f_datadog_logs_historical_pod
  • 91g_datadog_metrics_mismatched_pod
  • 91h_datadog_logs_empty_query_with_url
  • 91i_datadog_metrics_empty_query_with_url
  • 92_cpu_graph_conversation
  • 93_calling_datadog
  • 93_events_since_specific_date
  • 94_runbook_transparency
  • 95_runbook_memory_leak_detection
  • 96_no_matching_runbook
  • 97_logs_clarification_needed
  • 99_logs_transparency_custom_time

test_investigate:

  • 01_oom_kill
  • 02_crashloop_backoff
  • 03_cpu_throttling
  • 04_image_pull_backoff
  • 05_crashpod
  • 06_job_failure
  • 07_job_syntax_error
  • 08_memory_pressure
  • 09_high_latency
  • 10_KubeDeploymentReplicasMismatch
  • 11_KubePodCrashLooping
  • 12_KubePodNotReady
  • 13_Watchdog
  • 14_tempo
  • 15_dns_resolution
  • 16_dns_resolution_no_tool
  • 17_investigate_correct_date

Sheeproid added a commit that referenced this pull request Feb 2, 2026
 Post-processing support was removed in PR #1279 (bec4d12)

Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants