Skip to content

Optimize LLM calls by batching final answer with todo completion - #1574

Open
aantn wants to merge 9 commits into
masterfrom
claude/fix-holmes-llm-call-RZFf6
Open

aantn wants to merge 9 commits into
masterfrom
claude/fix-holmes-llm-call-RZFf6

Conversation

@aantn

@aantn aantn commented Feb 16, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR optimizes the LLM interaction flow by allowing the model to provide its final answer in the same response as marking the last tasks as completed, eliminating unnecessary extra LLM calls.

Key Changes

  • Early return optimization: Added _all_todos_completed() helper function to detect when all tasks in a TodoWrite call are marked as completed
  • Skip redundant LLM calls: Modified both call() and call_stream() methods in ToolCallingLLM to return early when:
    • All todos are marked as completed
    • The LLM has already provided a text response in the same turn
    • This prevents an unnecessary extra LLM call that would only be used to generate the final answer
  • Updated prompts: Enhanced instructions across multiple prompt templates to explicitly encourage batching the final answer with the last TodoWrite call:
    • _general_instructions.jinja2
    • _noflag_general_instructions.jinja2
    • investigator_instructions.jinja2
    • investigation_procedure.jinja2

Implementation Details

The optimization works by:

  1. Checking if any TodoWrite tool call in the current batch has all tasks marked as "completed"
  2. Verifying that the LLM response includes text content alongside the tool calls
  3. Returning the result immediately instead of looping for another LLM call
  4. Properly accounting for token usage and metadata before returning

The prompt updates make it clear to the LLM that it should include the final answer text in the same response as the final TodoWrite call, rather than waiting for a separate turn.

https://claude.ai/code/session_01D3bToEw4CfTUkPWhVGsHCN

Summary by CodeRabbit

  • Refactor
    • Streamlined finalization flow so the last task’s response includes the complete final answer, removing an extra completion step and reducing response latency.
    • Harmonized efficiency guidance across task systems to consistently handle the final task and improve end-to-end response timing.

When the LLM marks all TodoWrite tasks as completed and includes its
final answer text in the same response, we now return immediately
instead of making an unnecessary extra LLM call. This saves 3-5 seconds
per Holmes run.

Changes:
- Prompt updates across 4 template files instructing the LLM to batch
  its final answer with the last TodoWrite call
- Code-level early return in both call() and call_stream() that detects
  all-tasks-completed and uses the existing text response

https://claude.ai/code/session_01D3bToEw4CfTUkPWhVGsHCN
Signed-off-by: Claude <noreply@anthropic.com>
@netlify

netlify Bot commented Feb 16, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit c4dafba
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69981ab777ff1e00083d462b
😎 Deploy Preview https://deploy-preview-1574--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Feb 16, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 60fd7b0 (#22129859294)

✅ Results of HolmesGPT evals

Automatically triggered by commit 60fd7b0 on branch claude/fix-holmes-llm-call-RZFf6

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 31.0s 5 11 $0.2279
✅ 101_loki_historical_logs_pod_deleted 34.5s 5 7 $0.2138
✅ 111_pod_names_contain_service 33.0s 5 11 $0.2237
✅ 112_find_pvcs_by_uuid 26.4s 5 6 $0.2030
✅ 12_job_crashing 29.4s 4 10 $0.2171
✅ 176_network_policy_blocking_traffic_no_runbooks 38.2s 6 14 $0.2717
✅ 24_misconfigured_pvc 30.3s 4 13 $0.2130
✅ 43_current_datetime_from_prompt 6.0s 1 — $0.1104
✅ 61_exact_match_counting 14.4s 3 2 $0.1519
Total 27.0s avg 4.2 avg 9.2 avg $1.8325
📜 Run @ 2582494 (#22097414575)

✅ Results of HolmesGPT evals

Automatically triggered by commit 2582494 on branch claude/fix-holmes-llm-call-RZFf6

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 33.5s 5 11 $0.2325
✅ 101_loki_historical_logs_pod_deleted 29.4s 3 7 $0.1921
✅ 111_pod_names_contain_service 29.2s 4 10 $0.2046
✅ 112_find_pvcs_by_uuid 31.1s 5 7 $0.2362
✅ 12_job_crashing 33.9s 5 12 $0.2336
✅ 176_network_policy_blocking_traffic_no_runbooks 45.6s 6 17 $0.2841
✅ 24_misconfigured_pvc 33.8s 5 13 $0.2255
✅ 43_current_datetime_from_prompt 5.2s 1 — $0.1067
✅ 61_exact_match_counting 13.7s 3 3 $0.1483
Total 28.4s avg 4.1 avg 10.0 avg $1.8633
📜 Run @ f0c4240 (#22079178276)

✅ Results of HolmesGPT evals

Automatically triggered by commit f0c4240 on branch claude/fix-holmes-llm-call-RZFf6

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 28.8s 4 10 $0.2073
✅ 101_loki_historical_logs_pod_deleted 37.9s 5 9 $0.2329
✅ 111_pod_names_contain_service 34.8s 5 11 $0.2292
✅ 112_find_pvcs_by_uuid 31.0s 5 6 $0.2193
✅ 12_job_crashing 27.6s 4 9 $0.2069
✅ 176_network_policy_blocking_traffic_no_runbooks 44.7s 6 17 $0.2863
✅ 24_misconfigured_pvc 37.1s 6 14 $0.2443
✅ 43_current_datetime_from_prompt 5.3s 1 — $0.1068
✅ 61_exact_match_counting 13.2s 3 2 $0.1437
Total 28.9s avg 4.3 avg 9.8 avg $1.8767
📜 Run @ b5d9cac (#22077454812)

✅ Results of HolmesGPT evals

Automatically triggered by commit b5d9cac on branch claude/fix-holmes-llm-call-RZFf6

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.2s 5 11 $0.2282
✅ 101_loki_historical_logs_pod_deleted 53.2s 6 13 $0.2825
✅ 111_pod_names_contain_service 33.3s 5 11 $0.2277
✅ 112_find_pvcs_by_uuid 40.5s 7 9 $0.2688
✅ 12_job_crashing 30.8s 5 10 $0.2240
✅ 176_network_policy_blocking_traffic_no_runbooks 46.8s 6 17 $0.2799
✅ 24_misconfigured_pvc 33.7s 5 14 $0.2319
✅ 43_current_datetime_from_prompt 5.7s 1 — $0.1069
✅ 61_exact_match_counting 13.2s 3 2 $0.1437
Total 32.1s avg 4.8 avg 10.9 avg $1.9937
📜 Run @ 9c7e80e (#22068355617)

✅ Results of HolmesGPT evals

Automatically triggered by commit 9c7e80e on branch claude/fix-holmes-llm-call-RZFf6

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 33.7s 5 11 $0.2364
✅ 101_loki_historical_logs_pod_deleted 48.8s 6 11 $0.2659
✅ 111_pod_names_contain_service 36.3s 5 11 $0.2255
✅ 112_find_pvcs_by_uuid 38.8s 7 8 $0.2530
✅ 12_job_crashing 40.7s 6 14 $0.2625
✅ 176_network_policy_blocking_traffic_no_runbooks 55.2s 8 17 $0.3168
✅ 24_misconfigured_pvc 35.0s 5 14 $0.2382
✅ 43_current_datetime_from_prompt 5.1s 1 — $0.1081
✅ 61_exact_match_counting 18.0s 4 4 $0.1637
Total 34.6s avg 5.2 avg 11.2 avg $2.0701

✅ Results of HolmesGPT evals

Automatically triggered by commit c4dafba on branch claude/fix-holmes-llm-call-RZFf6

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 28.9s 5 9 $0.2158
✅ 101_loki_historical_logs_pod_deleted 37.8s 5 10 $0.2470
✅ 111_pod_names_contain_service 28.8s 4 10 $0.2087
✅ 112_find_pvcs_by_uuid 23.1s 4 6 $0.2005
✅ 12_job_crashing 21.4s 3 6 $0.1783
✅ 176_network_policy_blocking_traffic_no_runbooks 40.5s 5 15 $0.2605
✅ 24_misconfigured_pvc 30.6s 5 14 $0.2239
✅ 43_current_datetime_from_prompt 5.0s 1 — $0.1092
✅ 61_exact_match_counting 14.7s 3 2 $0.1523
Total 25.6s avg 3.9 avg 9.0 avg $1.7963
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-holmes-llm-call-RZFf6 -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-holmes-llm-call-RZFf6 -f markers=regression -f filter=

@github-actions

github-actions Bot commented Feb 16, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for 344c262 (built in 3m 36s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:344c262
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:344c262 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:344c262
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:344c262

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:344c262

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:344c262

@coderabbitai

coderabbitai Bot commented Feb 16, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Updates four prompt templates to add an EFFICIENCY rule: when the last task is completed, the assistant should provide the final answer directly (embedding it in or replacing the final TodoWrite call), eliminating an extra TodoWrite/final-turn step.

Changes

Cohort / File(s) Summary
General instruction templates
holmes/plugins/prompts/_general_instructions.jinja2, holmes/plugins/prompts/_noflag_general_instructions.jinja2
Added EFFICIENCY guidance to instruct embedding the final answer in the last TodoWrite response or providing the final answer directly instead of issuing a final TodoWrite; also fixed quote style for the in_progress token.
Investigator procedure templates
holmes/plugins/prompts/investigation_procedure.jinja2, holmes/plugins/toolsets/investigator/investigator_instructions.jinja2
Inserted an EFFICIENCY rule to omit the final TodoWrite call when a single remaining task is finished; adjusted final-review/verifications wording and phase references to reflect the streamlined finalization flow.

Sequence Diagram(s)

sequenceDiagram
    participant User
    participant Assistant
    participant TodoWriteTool as TodoWrite
    Note right of Assistant: New EFFICIENCY flow
    User->>Assistant: Request / task list
    Assistant->>TodoWrite: Create/Update tasks (multiple rounds)
    TodoWrite-->>Assistant: Task statuses
    alt Multiple tasks remain
        Assistant->>TodoWrite: Mark task complete / create next task
        TodoWrite-->>Assistant: Ack
    else Last task completed
        Assistant-->>User: Provide final answer (no final TodoWrite call)
    end
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~8 minutes

Suggested reviewers

  • moshemorad
🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main optimization: batching the final answer with todo completion to reduce LLM calls, which is the core focus of all file changes.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Feb 16, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 10.77s 9.98s +8.0%
Warm Mean 4.65s 4.88s -4.7%
Warm Min 4.59s 4.77s
Warm Max 4.70s 5.03s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 34.43s 26.61s +29.4%
Warm Mean 7.84s 7.74s +1.2%
Warm Min 6.77s 7.43s
Warm Max 8.72s 8.15s

PR: 344c2620 | Master: 23219be6 | Iterations: 5

Remove the _all_todos_completed helper and early-return logic from both
call() and call_stream() in tool_calling_llm.py. The prompt changes
alone should be sufficient to get the LLM to batch its final answer
with the last TodoWrite call, avoiding the extra LLM round trip.

https://claude.ai/code/session_01D3bToEw4CfTUkPWhVGsHCN
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@holmes/core/tool_calling_llm.py`:
- Around line 239-251: Update the type hint for the function
_all_todos_completed so tools_to_call is annotated as
list[ChatCompletionMessageToolCall] instead of the unparameterized list; locate
the _all_todos_completed definition and replace its parameter type accordingly
(using the already-imported ChatCompletionMessageToolCall) to satisfy complete
typing requirements.
🧹 Nitpick comments (1)
holmes/core/tool_calling_llm.py (1)

51-51: New cross-module import couples tool_calling_llm to the investigator toolset.

Importing TODO_WRITE_TOOL_NAME from holmes.plugins.toolsets.investigator.core_investigation creates a dependency from the core module to a specific plugin/toolset. This is a minor architectural concern — if the investigator toolset is ever removed or restructured, this core module would break.

Consider defining the constant in a shared location (e.g., holmes.core.constants) or passing it as configuration. Not blocking, but worth noting for future maintainability.

Comment thread holmes/core/tool_calling_llm.py Outdated
@aantn

aantn commented Feb 16, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
branch: master
iterations: 3

@aantn

This comment was marked as outdated.

@github-actions

This comment was marked as outdated.

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ⚠️ Completed with 90 failures


🧪 Manual Eval Results

Parameter Value
Triggered via /eval on branch master
Branch master
Model opus-4.5
Markers all LLM tests
Iterations 3
Duration 53m 4s
Workflow View logs | Rerun

Results of HolmesGPT evals (branch: master)

  • ask_holmes: 332/702 test cases were successful, 83 regressions, 63 skipped, 219 setup failures, 5 mock failures
  • investigate: 42/51 test cases were successful, 7 regressions
Status Test case Time Turns Tools Cost
🚧 01_how_many_pods — — — —
🚧 01_how_many_pods — — — —
🚧 01_how_many_pods — — — —
✅ 02_what_is_wrong_with_pod 31.2s 4 6 $0.2909
✅ 02_what_is_wrong_with_pod 29.2s 4 6 $0.2933
✅ 02_what_is_wrong_with_pod 30.4s 4 6 $0.2954
✅ 03_what_is_the_command_to_port_forward 26.6s 5 6 $0.4447
✅ 03_what_is_the_command_to_port_forward 26.2s 5 6 $0.4063
✅ 03_what_is_the_command_to_port_forward 27.2s 5 6 $0.4234
✅ 04_related_k8s_events 34.3s 5 6 $0.4605
✅ 04_related_k8s_events 31.5s 5 6 $0.4345
✅ 04_related_k8s_events 33.2s 5 6 $0.4365
✅ 05_image_version 34.8s 6 6 $0.5767
✅ 05_image_version 33.2s 6 6 $0.5806
✅ 05_image_version 36.2s 6 6 $0.6387
🔧 06_explain_issue 26.7s 4 5 —
🔧 06_explain_issue 29.9s 5 5 —
✅ 06_explain_issue 15.2s 2 1 $0.2134
✅ 07_high_latency 44.1s 5 9 $0.3769
✅ 07_high_latency 40.5s 5 10 $0.3917
✅ 07_high_latency 55.9s 8 13 $0.4640
➖ 08_sock_shop_frontend — — — —
➖ 08_sock_shop_frontend — — — —
➖ 08_sock_shop_frontend — — — —
✅ 09_crashpod 37.3s 5 10 $0.4743
✅ 09_crashpod 42.2s 6 11 $0.5033
✅ 09_crashpod 43.3s 6 11 $0.5413
🚧 100a_loki_historical_logs — — — —
🚧 100a_loki_historical_logs — — — —
🚧 100a_loki_historical_logs — — — —
🚧 101_loki_historical_logs_pod_deleted — — — —
🚧 101_loki_historical_logs_pod_deleted — — — —
🚧 101_loki_historical_logs_pod_deleted — — — —
🚧 102_loki_label_discovery — — — —
🚧 102_loki_label_discovery — — — —
🚧 102_loki_label_discovery — — — —
🚧 102a_loki_logs_transparency — — — —
🚧 102a_loki_logs_transparency — — — —
🚧 102a_loki_logs_transparency — — — —
🚧 102b_loki_multiple_pods — — — —
🚧 102b_loki_multiple_pods — — — —
🚧 102b_loki_multiple_pods — — — —
✅ 103_logs_transparency_default_limit 34.3s 6 7 $0.3253
✅ 103_logs_transparency_default_limit 36.6s 6 8 $0.5149
✅ 103_logs_transparency_default_limit 33.7s 6 5 $0.4565
❌ 104a_postgres_root_issue 39.9s 5 12 $0.3514
❌ 104a_postgres_root_issue 50.2s 6 11 $0.4581
❌ 104a_postgres_root_issue 51.1s 6 12 $0.4566
➖ 104b_postgres_missing_index_pgstat — — — —
➖ 104b_postgres_missing_index_pgstat — — — —
➖ 104b_postgres_missing_index_pgstat — — — —
➖ 104c_postgres_minimal_missing_index — — — —
➖ 104c_postgres_minimal_missing_index — — — —
➖ 104c_postgres_minimal_missing_index — — — —
➖ 105_redis_wrong_data_structure — — — —
➖ 105_redis_wrong_data_structure — — — —
➖ 105_redis_wrong_data_structure — — — —
✅ 107_log_filter_http_status_code 58.7s 6 17 $0.6415
✅ 107_log_filter_http_status_code 58.0s 6 17 $0.7592
✅ 107_log_filter_http_status_code 63.9s 7 20 $0.5922
🚧 108_logs_nearby_lines — — — —
🚧 108_logs_nearby_lines — — — —
🚧 108_logs_nearby_lines — — — —
✅ 109_logs_transparency_not_found 36.5s 6 8 $0.4600
✅ 109_logs_transparency_not_found 30.6s 5 7 $0.2998
✅ 109_logs_transparency_not_found 23.5s 4 5 $0.2607
✅ 10_image_pull_backoff 36.8s 5 10 $0.3327
✅ 10_image_pull_backoff 46.4s 5 13 $0.4823
✅ 10_image_pull_backoff 36.1s 5 9 $0.3262
✅ 110_cpu_graph_robusta_runner[0] 31.6s 5 8 $0.4338
✅ 110_cpu_graph_robusta_runner[0] 51.2s 8 14 $0.5534
✅ 110_cpu_graph_robusta_runner[0] 43.2s 7 12 $0.5394
✅ 110_cpu_graph_robusta_runner[10] 30.5s 5 6 $0.4303
✅ 110_cpu_graph_robusta_runner[10] 51.4s 9 11 $0.7840
✅ 110_cpu_graph_robusta_runner[10] 39.5s 7 9 $0.5423
✅ 110_cpu_graph_robusta_runner[11] 34.3s 6 8 $0.4943
✅ 110_cpu_graph_robusta_runner[11] 37.2s 6 11 $0.5140
✅ 110_cpu_graph_robusta_runner[11] 66.4s 10 18 $0.6894
✅ 110_cpu_graph_robusta_runner[12] 45.6s 7 10 $0.5761
✅ 110_cpu_graph_robusta_runner[12] 37.8s 6 8 $0.6434
✅ 110_cpu_graph_robusta_runner[12] 35.9s 6 9 $0.4911
❌ 110_cpu_graph_robusta_runner[13] 44.9s 8 10 $0.5331
✅ 110_cpu_graph_robusta_runner[13] 51.7s 9 12 $0.6480
✅ 110_cpu_graph_robusta_runner[13] 66.5s 12 14 $0.6745
✅ 110_cpu_graph_robusta_runner[14] 34.4s 6 7 $0.4750
✅ 110_cpu_graph_robusta_runner[14] 32.0s 5 7 $0.4640
✅ 110_cpu_graph_robusta_runner[14] 33.2s 5 8 $0.4273
✅ 110_cpu_graph_robusta_runner[15] 41.3s 7 10 $0.5328
✅ 110_cpu_graph_robusta_runner[15] 36.8s 6 10 $0.4933
✅ 110_cpu_graph_robusta_runner[15] 38.6s 6 10 $0.4889
✅ 110_cpu_graph_robusta_runner[1] 37.5s 6 8 $0.5280
✅ 110_cpu_graph_robusta_runner[1] 39.2s 6 9 $0.4927
✅ 110_cpu_graph_robusta_runner[1] 65.8s 10 15 $0.8145
✅ 110_cpu_graph_robusta_runner[2] 51.8s 8 11 $0.7093
✅ 110_cpu_graph_robusta_runner[2] 85.5s 13 22 $0.7785
✅ 110_cpu_graph_robusta_runner[2] 57.2s 9 12 $0.7540
✅ 110_cpu_graph_robusta_runner[3] 61.8s 9 13 $0.6816
✅ 110_cpu_graph_robusta_runner[3] 57.1s 9 13 $0.6351
✅ 110_cpu_graph_robusta_runner[3] 46.5s 7 10 $0.6194
✅ 110_cpu_graph_robusta_runner[4] 38.7s 6 9 $0.5100
✅ 110_cpu_graph_robusta_runner[4] 31.2s 5 8 $0.4647
✅ 110_cpu_graph_robusta_runner[4] 41.2s 6 10 $0.5238
✅ 110_cpu_graph_robusta_runner[5] 35.2s 6 8 $0.5182
✅ 110_cpu_graph_robusta_runner[5] 49.4s 8 12 $0.6834
✅ 110_cpu_graph_robusta_runner[5] 67.8s 10 15 $0.8380
✅ 110_cpu_graph_robusta_runner[6] 49.1s 7 11 $0.6404
✅ 110_cpu_graph_robusta_runner[6] 40.4s 6 9 $0.4676
✅ 110_cpu_graph_robusta_runner[6] 36.2s 5 8 $0.4661
✅ 110_cpu_graph_robusta_runner[7] 50.2s 9 11 $0.5614
✅ 110_cpu_graph_robusta_runner[7] 34.1s 5 7 $0.4932
✅ 110_cpu_graph_robusta_runner[7] 42.6s 6 9 $0.6668
✅ 110_cpu_graph_robusta_runner[8] 38.1s 6 9 $0.4688
✅ 110_cpu_graph_robusta_runner[8] 39.7s 7 10 $0.5280
✅ 110_cpu_graph_robusta_runner[8] 32.8s 5 7 $0.4487
✅ 110_cpu_graph_robusta_runner[9] 103.5s 17 29 $0.9455
✅ 110_cpu_graph_robusta_runner[9] 53.2s 8 12 $0.6122
✅ 110_cpu_graph_robusta_runner[9] 59.7s 9 14 $0.6135
✅ 110_k8s_events_image_pull 33.4s 5 7 $0.4248
✅ 110_k8s_events_image_pull 23.5s 4 6 $0.2759
✅ 110_k8s_events_image_pull 26.4s 4 6 $0.2800
✅ 111_disabled_datadog_traces 12.4s 1 — $0.1956
✅ 111_disabled_datadog_traces 11.0s 1 — $0.1938
✅ 111_disabled_datadog_traces 10.9s 1 — $0.1931
🚧 111_pod_names_contain_service — — — —
🚧 111_pod_names_contain_service — — — —
🚧 111_pod_names_contain_service — — — —
🔧 111_tool_hallucination 21.0s 3 2 —
✅ 111_tool_hallucination 11.8s 2 1 $0.1715
✅ 111_tool_hallucination 12.8s 2 1 $0.1723
✅ 112_find_pvcs_by_uuid 42.2s 7 6 $0.3806
✅ 112_find_pvcs_by_uuid 38.4s 6 7 $0.3650
✅ 112_find_pvcs_by_uuid 37.2s 6 5 $0.3492
🚧 114_checkout_latency_tracing_rebuild[0] — — — —
🚧 114_checkout_latency_tracing_rebuild[0] — — — —
🚧 114_checkout_latency_tracing_rebuild[0] — — — —
🚧 115_checkout_errors_tracing[0] — — — —
🚧 115_checkout_errors_tracing[0] — — — —
🚧 115_checkout_errors_tracing[0] — — — —
🚧 117_new_relic_tracing[0] — — — —
🚧 117_new_relic_tracing[0] — — — —
🚧 117_new_relic_tracing[0] — — — —
❌ 117b_new_relic_block_embed[0] 19.5s 3 2 $0.2375
❌ 117b_new_relic_block_embed[0] 15.1s 1 — $0.2012
❌ 117b_new_relic_block_embed[0] 11.8s 1 — $0.1973
🚧 118_new_relic_logs[0] — — — —
🚧 118_new_relic_logs[0] — — — —
🚧 118_new_relic_logs[0] — — — —
🚧 119_new_relic_metrics[0] — — — —
🚧 119_new_relic_metrics[0] — — — —
🚧 119_new_relic_metrics[0] — — — —
❌ 11_init_containers 93.8s 14 25 $1.0378
❌ 11_init_containers 139.4s 20 44 $1.4787
✅ 11_init_containers 62.4s 9 15 $0.7520
➖ 120_new_relic_traces2[0] — — — —
➖ 120_new_relic_traces2[0] — — — —
➖ 120_new_relic_traces2[0] — — — —
➖ 121_new_relic_checkout_errors_tracing[0] — — — —
➖ 121_new_relic_checkout_errors_tracing[0] — — — —
➖ 121_new_relic_checkout_errors_tracing[0] — — — —
➖ 122_new_relic_checkout_latency_tracing_rebuild[0] — — — —
➖ 122_new_relic_checkout_latency_tracing_rebuild[0] — — — —
➖ 122_new_relic_checkout_latency_tracing_rebuild[0] — — — —
➖ 123_new_relic_checkout_errors_tracing[0] — — — —
➖ 123_new_relic_checkout_errors_tracing[0] — — — —
➖ 123_new_relic_checkout_errors_tracing[0] — — — —
🚧 124_checkout_latency_prometheus[0] — — — —
🚧 124_checkout_latency_prometheus[0] — — — —
🚧 124_checkout_latency_prometheus[0] — — — —
❌ 124a_new_relic_multi_account_account_name[0] 9.4s 1 — $0.1764
❌ 124a_new_relic_multi_account_account_name[0] 10.8s 1 — $0.1786
❌ 124a_new_relic_multi_account_account_name[0] 10.7s 1 — $0.1781
❌ 124b_new_relic_multi_account_alert_prompt 53.0s 6 11 $0.4453
❌ 124b_new_relic_multi_account_alert_prompt 56.6s 7 14 $0.4678
❌ 124b_new_relic_multi_account_alert_prompt 54.3s 6 13 $0.4601
❌ 124c_new_relic_multi_account_default[0] 19.1s 3 2 $0.2207
❌ 124c_new_relic_multi_account_default[0] 18.9s 3 2 $0.2212
❌ 124c_new_relic_multi_account_default[0] 20.3s 3 2 $0.2213
🚧 12_job_crashing — — — —
🚧 12_job_crashing — — — —
🚧 12_job_crashing — — — —
❌ 13a_pending_node_selector_basic 88.3s 11 27 $1.1291
✅ 13a_pending_node_selector_basic 72.2s 9 21 $0.8621
✅ 13a_pending_node_selector_basic 57.0s 8 19 $0.7478
❌ 13b_pending_node_selector_detailed 130.1s 19 42 $1.3447
❌ 13b_pending_node_selector_detailed 93.2s 12 30 $1.0443
❌ 13b_pending_node_selector_detailed 92.2s 12 30 $1.0100
✅ 14_pending_resources 50.5s 6 14 $0.5289
✅ 14_pending_resources 46.2s 6 11 $0.5234
✅ 14_pending_resources 47.3s 6 12 $0.5464
🚧 151_disabled_toolsets_fallback_only — — — —
🚧 151_disabled_toolsets_fallback_only — — — —
🚧 151_disabled_toolsets_fallback_only — — — —
➖ 156_kafka_opensearch_latency — — — —
➖ 156_kafka_opensearch_latency — — — —
➖ 156_kafka_opensearch_latency — — — —
🚧 157_disk_full_statefulset — — — —
🚧 157_disk_full_statefulset — — — —
🚧 157_disk_full_statefulset — — — —
🚧 158_slack_chat_correct_date[0] — — — —
🚧 158_slack_chat_correct_date[0] — — — —
🚧 158_slack_chat_correct_date[0] — — — —
🚧 159_prometheus_high_cardinality_cpu[0] — — — —
🚧 159_prometheus_high_cardinality_cpu[0] — — — —
🚧 159_prometheus_high_cardinality_cpu[0] — — — —
🚧 159_prometheus_high_cardinality_cpu[1] — — — —
🚧 159_prometheus_high_cardinality_cpu[1] — — — —
🚧 159_prometheus_high_cardinality_cpu[1] — — — —
🚧 159_prometheus_high_cardinality_cpu[2] — — — —
🚧 159_prometheus_high_cardinality_cpu[2] — — — —
🚧 159_prometheus_high_cardinality_cpu[2] — — — —
✅ 15_failed_readiness_probe 78.1s 12 17 $0.9878
✅ 15_failed_readiness_probe 66.6s 10 15 $0.8115
✅ 15_failed_readiness_probe 55.3s 8 12 $0.7501
🚧 160_electricity_market_bidding_bug[0] — — — —
🚧 160_electricity_market_bidding_bug[0] — — — —
🚧 160_electricity_market_bidding_bug[0] — — — —
🚧 160a_cpu_per_namespace_graph — — — —
🚧 160a_cpu_per_namespace_graph — — — —
🚧 160a_cpu_per_namespace_graph — — — —
➖ 160b_cpu_per_namespace_graph_with_prom_truncation — — — —
➖ 160b_cpu_per_namespace_graph_with_prom_truncation — — — —
➖ 160b_cpu_per_namespace_graph_with_prom_truncation — — — —
➖ 160c_cpu_per_namespace_graph_with_global_truncation — — — —
➖ 160c_cpu_per_namespace_graph_with_global_truncation — — — —
➖ 160c_cpu_per_namespace_graph_with_global_truncation — — — —
🚧 161_bidding_version_performance[0] — — — —
🚧 161_bidding_version_performance[0] — — — —
🚧 161_bidding_version_performance[0] — — — —
🚧 161_conversation_compaction — — — —
🚧 161_conversation_compaction — — — —
🚧 161_conversation_compaction — — — —
🚧 162_get_runbooks — — — —
🚧 162_get_runbooks — — — —
🚧 162_get_runbooks — — — —
🚧 163_compaction_follow_up — — — —
🚧 163_compaction_follow_up — — — —
🚧 163_compaction_follow_up — — — —
🚧 164_datadog_traces_coupon_code[0] — — — —
🚧 164_datadog_traces_coupon_code[0] — — — —
🚧 164_datadog_traces_coupon_code[0] — — — —
✅ 165_alert_with_multiple_runbooks 35.1s 5 8 $0.3278
✅ 165_alert_with_multiple_runbooks 41.6s 6 10 $0.3610
✅ 165_alert_with_multiple_runbooks 34.6s 5 8 $0.3262
✅ 16_failed_no_toolset_found 12.5s 1 — $0.1944
✅ 16_failed_no_toolset_found 14.1s 1 — $0.1965
✅ 16_failed_no_toolset_found 17.2s 1 — $0.2035
🚧 173_coralogix_logs[0] — — — —
🚧 173_coralogix_logs[0] — — — —
🚧 173_coralogix_logs[0] — — — —
🚧 174_coralogix_traces_ad[0] — — — —
🚧 174_coralogix_traces_ad[0] — — — —
🚧 174_coralogix_traces_ad[0] — — — —
🚧 175_coralogix_metrics_frontend[0] — — — —
🚧 175_coralogix_metrics_frontend[0] — — — —
🚧 175_coralogix_metrics_frontend[0] — — — —
🚧 176_network_policy_blocking_traffic_no_runbooks — — — —
🚧 176_network_policy_blocking_traffic_no_runbooks — — — —
🚧 176_network_policy_blocking_traffic_no_runbooks — — — —
🚧 177_grafana_home_dashboard — — — —
🚧 177_grafana_home_dashboard — — — —
🚧 177_grafana_home_dashboard — — — —
🚧 178_grafana_search_dashboard_query — — — —
🚧 178_grafana_search_dashboard_query — — — —
🚧 178_grafana_search_dashboard_query — — — —
🚧 179_grafana_big_dashboard_query — — — —
🚧 179_grafana_big_dashboard_query — — — —
🚧 179_grafana_big_dashboard_query — — — —
✅ 17_oom_kill 47.7s 6 12 $0.5509
✅ 17_oom_kill 41.3s 5 11 $0.3540
✅ 17_oom_kill 50.6s 6 13 $0.5644
✅ 180_connectivity_check_tcp 15.5s 3 4 $0.2364
✅ 180_connectivity_check_tcp 16.3s 3 4 $0.2365
✅ 180_connectivity_check_tcp 15.5s 3 4 $0.2333
❌ 181_connectivity_check_http 19.4s 3 4 $0.2355
❌ 181_connectivity_check_http 18.4s 3 5 $0.2404
❌ 181_connectivity_check_http 18.0s 3 4 $0.2366
❌ 182_connectivity_check_http_url 20.2s 3 3 $0.2403
❌ 182_connectivity_check_http_url 19.5s 3 3 $0.2379
❌ 182_connectivity_check_http_url 23.2s 3 3 $0.2431
✅ 183a_elasticsearch_cluster_health 14.8s 3 3 $0.2441
✅ 183a_elasticsearch_cluster_health 16.6s 3 3 $0.2504
✅ 183a_elasticsearch_cluster_health 16.0s 3 3 $0.2481
🚧 183b_elasticsearch_index_discovery — — — —
🚧 183b_elasticsearch_index_discovery — — — —
🚧 183b_elasticsearch_index_discovery — — — —
✅ 183c_elasticsearch_log_search 31.8s 6 6 $0.3250
✅ 183c_elasticsearch_log_search 30.3s 6 6 $0.3233
✅ 183c_elasticsearch_log_search 32.6s 6 6 $0.3262
✅ 183d_elasticsearch_aggregation 35.0s 6 8 $0.3352
✅ 183d_elasticsearch_aggregation 37.5s 6 8 $0.3402
✅ 183d_elasticsearch_aggregation 37.2s 6 8 $0.3390
✅ 183e_elasticsearch_field_mappings 16.3s 3 3 $0.2494
✅ 183e_elasticsearch_field_mappings 16.7s 3 3 $0.2503
✅ 183e_elasticsearch_field_mappings 15.8s 3 3 $0.2489
🚧 183f_elasticsearch_shard_filtering — — — —
🚧 183f_elasticsearch_shard_filtering — — — —
🚧 183f_elasticsearch_shard_filtering — — — —
❌ 183g_elasticsearch_index_stats 20.4s 3 3 $0.2683
❌ 183g_elasticsearch_index_stats 16.3s 3 3 $0.2549
❌ 183g_elasticsearch_index_stats 19.4s 3 3 $0.2646
🚧 184_elasticsearch_index_explosion — — — —
🚧 184_elasticsearch_index_explosion — — — —
🚧 184_elasticsearch_index_explosion — — — —
✅ 185_elasticsearch_cross_region_search 65.1s 9 12 $0.4969
✅ 185_elasticsearch_cross_region_search 62.4s 9 13 $0.4684
✅ 185_elasticsearch_cross_region_search 67.1s 10 13 $0.6739
🚧 186_elasticsearch_shard_explosion — — — —
🚧 186_elasticsearch_shard_explosion — — — —
🚧 186_elasticsearch_shard_explosion — — — —
✅ 187_elasticsearch_disk_space 19.1s 3 3 $0.2571
✅ 187_elasticsearch_disk_space 18.8s 3 3 $0.2554
✅ 187_elasticsearch_disk_space 18.2s 3 3 $0.2533
✅ 188_elasticsearch_mapping_explosion 43.4s 5 10 $0.3628
✅ 188_elasticsearch_mapping_explosion 33.6s 4 9 $0.3107
✅ 188_elasticsearch_mapping_explosion 37.9s 6 9 $0.3452
✅ 189_elasticsearch_timeseries_gap 62.4s 8 11 $0.4476
✅ 189_elasticsearch_timeseries_gap 38.8s 5 7 $0.3616
✅ 189_elasticsearch_timeseries_gap 54.0s 8 9 $0.4544
✅ 18_oom_kill_from_issues_history 44.6s 6 11 $0.4905
✅ 18_oom_kill_from_issues_history 59.0s 6 14 $0.5604
✅ 18_oom_kill_from_issues_history 61.7s 8 12 $0.5928
✅ 190_elasticsearch_cross_service_correlation 30.0s 4 4 $0.3031
✅ 190_elasticsearch_cross_service_correlation 38.5s 5 7 $0.3429
✅ 190_elasticsearch_cross_service_correlation 30.6s 4 5 $0.3106
🚧 191_elasticsearch_query_profile — — — —
🚧 191_elasticsearch_query_profile — — — —
🚧 191_elasticsearch_query_profile — — — —
✅ 193_elasticsearch_large_mapping_search 16.1s 3 3 $0.2443
✅ 193_elasticsearch_large_mapping_search 21.0s 4 4 $0.2692
✅ 193_elasticsearch_large_mapping_search 15.6s 3 3 $0.2443
❌ 194_frontend_chart_embed[0] 25.0s 4 4 $0.2772
❌ 194_frontend_chart_embed[0] 26.0s 4 4 $0.2790
❌ 194_frontend_chart_embed[0] 29.4s 4 5 $0.2890
🚧 195_bash_simple_allowed — — — —
🚧 195_bash_simple_allowed — — — —
🚧 195_bash_simple_allowed — — — —
✅ 195_elasticsearch_trace_large_fields 64.4s 12 13 $0.4829
✅ 195_elasticsearch_trace_large_fields 82.1s 15 15 $0.5584
✅ 195_elasticsearch_trace_large_fields 62.9s 11 12 $0.4688
🚧 196_bash_composed_allowed — — — —
🚧 196_bash_composed_allowed — — — —
🚧 196_bash_composed_allowed — — — —
✅ 197_bash_secrets_denied 8.1s 1 — $0.1712
✅ 197_bash_secrets_denied 7.3s 1 — $0.1712
✅ 197_bash_secrets_denied 7.6s 1 — $0.1714
✅ 198_bash_hardcoded_block 19.8s 3 2 $0.2214
✅ 198_bash_hardcoded_block 7.6s 1 — $0.1726
✅ 198_bash_hardcoded_block 9.4s 1 — $0.1738
🚧 199_bash_prefix_extraction — — — —
🚧 199_bash_prefix_extraction — — — —
🚧 199_bash_prefix_extraction — — — —
❌ 19_detect_missing_app_details — — — —
✅ 19_detect_missing_app_details 67.7s 9 14 $0.8513
✅ 19_detect_missing_app_details 166.3s 27 31 $1.7839
✅ 200_bash_prefix_resource_type 12.0s 2 1 $0.2020
✅ 200_bash_prefix_resource_type 11.3s 2 1 $0.2019
✅ 200_bash_prefix_resource_type 12.2s 2 1 $0.2031
✅ 201_bash_read_only_prefix_suggested 11.2s 2 1 $0.2036
✅ 201_bash_read_only_prefix_suggested 11.2s 2 1 $0.2064
✅ 201_bash_read_only_prefix_suggested 13.2s 2 1 $0.2072
✅ 202_bash_committed_prefix_shorter 11.7s 2 1 $0.1903
✅ 202_bash_committed_prefix_shorter 10.0s 2 1 $0.1901
✅ 202_bash_committed_prefix_shorter 11.2s 2 1 $0.1907
✅ 203_bash_delete_prefix_shorter 9.2s 2 1 $0.2041
✅ 203_bash_delete_prefix_shorter 9.9s 2 1 $0.2045
✅ 203_bash_delete_prefix_shorter 11.1s 2 1 $0.2057
✅ 204_bash_allow_list_visibility 10.2s 1 — $0.1883
✅ 204_bash_allow_list_visibility 9.6s 1 — $0.1886
✅ 204_bash_allow_list_visibility 9.0s 1 — $0.1874
🚧 205_bash_deployment_logs_all_pods — — — —
🚧 205_bash_deployment_logs_all_pods — — — —
🚧 205_bash_deployment_logs_all_pods — — — —
🚧 206_bash_deployment_logs_all_pods_all_toolsets — — — —
🚧 206_bash_deployment_logs_all_pods_all_toolsets — — — —
🚧 206_bash_deployment_logs_all_pods_all_toolsets — — — —
🚧 208_confluence_page_fetch — — — —
🚧 208_confluence_page_fetch — — — —
🚧 208_confluence_page_fetch — — — —
🚧 209_confluence_url_lookup — — — —
🚧 209_confluence_url_lookup — — — —
🚧 209_confluence_url_lookup — — — —
❌ 20_long_log_file_search 65.0s 10 15 $0.5993
❌ 20_long_log_file_search 54.6s 9 13 $0.5907
❌ 20_long_log_file_search 47.5s 8 11 $0.5488
🚧 210_confluence_runbook_fetch — — — —
🚧 210_confluence_runbook_fetch — — — —
🚧 210_confluence_runbook_fetch — — — —
🚧 211_prometheus_alerting_rules — — — —
🚧 211_prometheus_alerting_rules — — — —
🚧 211_prometheus_alerting_rules — — — —
🚧 212_kubevela_app_diagnosis — — — —
🚧 212_kubevela_app_diagnosis — — — —
🚧 212_kubevela_app_diagnosis — — — —
✅ 212_large_configmap_needle 27.4s 5 5 $0.2830
✅ 212_large_configmap_needle 28.3s 5 5 $0.2846
✅ 212_large_configmap_needle 35.4s 7 7 $0.3265
🚧 213_confluence_space_info — — — —
🚧 213_confluence_space_info — — — —
🚧 213_confluence_space_info — — — —
🚧 214_confluence_page_content — — — —
🚧 214_confluence_page_content — — — —
🚧 214_confluence_page_content — — — —
🚧 215_confluence_cql_search — — — —
🚧 215_confluence_cql_search — — — —
🚧 215_confluence_cql_search — — — —
🚧 216_confluence_page_hierarchy — — — —
🚧 216_confluence_page_hierarchy — — — —
🚧 216_confluence_page_hierarchy — — — —
✅ 21_job_fail_curl_no_svc_account 45.6s 7 10 $0.5370
✅ 21_job_fail_curl_no_svc_account 44.1s 6 10 $0.5244
✅ 21_job_fail_curl_no_svc_account 42.3s 6 10 $0.5158
➖ 22_high_latency_dbi_down — — — —
➖ 22_high_latency_dbi_down — — — —
➖ 22_high_latency_dbi_down — — — —
🚧 23_app_error_in_current_logs — — — —
🚧 23_app_error_in_current_logs — — — —
🚧 23_app_error_in_current_logs — — — —
✅ 24_misconfigured_pvc 50.3s 6 17 $0.5578
✅ 24_misconfigured_pvc 43.6s 6 14 $0.5308
✅ 24_misconfigured_pvc 56.3s 7 17 $0.6213
✅ 25_misconfigured_ingress_class 60.8s 8 15 $0.6941
✅ 25_misconfigured_ingress_class 63.0s 8 18 $0.5374
✅ 25_misconfigured_ingress_class 61.8s 8 20 $0.4373
✅ 26_page_render_times 44.1s 5 6 $0.5324
✅ 26_page_render_times 41.9s 6 7 $0.4910
✅ 26_page_render_times 44.3s 6 7 $0.5844
✅ 27a_multi_container_logs 28.0s 4 6 $0.3196
✅ 27a_multi_container_logs 26.5s 4 6 $0.3155
✅ 27a_multi_container_logs 30.3s 5 7 $0.3479
❌ 27b_multi_container_logs 22.2s 4 5 $0.2603
❌ 27b_multi_container_logs 21.7s 4 5 $0.2613
❌ 27b_multi_container_logs 22.4s 4 5 $0.2610
🚧 28_permissions_error — — — —
🚧 28_permissions_error — — — —
🚧 28_permissions_error — — — —
🚧 30_basic_promql_graph_cluster_memory — — — —
🚧 30_basic_promql_graph_cluster_memory — — — —
🚧 30_basic_promql_graph_cluster_memory — — — —
🚧 32_basic_promql_graph_pod_cpu — — — —
🚧 32_basic_promql_graph_pod_cpu — — — —
🚧 32_basic_promql_graph_pod_cpu — — — —
🚧 33_cpu_metrics_discovery — — — —
🚧 33_cpu_metrics_discovery — — — —
🚧 33_cpu_metrics_discovery — — — —
🚧 34_memory_graph — — — —
🚧 34_memory_graph — — — —
🚧 34_memory_graph — — — —
🚧 35_tempo — — — —
🚧 35_tempo — — — —
🚧 35_tempo — — — —
🚧 36_argocd_find_resource — — — —
🚧 36_argocd_find_resource — — — —
🚧 36_argocd_find_resource — — — —
🚧 37_argocd_wrong_namespace — — — —
🚧 37_argocd_wrong_namespace — — — —
🚧 37_argocd_wrong_namespace — — — —
❌ 38_rabbitmq_split_head 60.1s 9 18 $0.7260
❌ 38_rabbitmq_split_head 91.9s 13 28 $1.0744
❌ 38_rabbitmq_split_head 56.3s 8 22 $0.6072
✅ 39_failed_toolset 41.0s 5 10 $0.6154
✅ 39_failed_toolset 52.5s 6 16 $0.7043
✅ 39_failed_toolset 70.3s 10 20 $0.8330
✅ 41_setup_argo 12.1s 1 — $0.1941
✅ 41_setup_argo 11.7s 1 — $0.1936
✅ 41_setup_argo 11.5s 1 — $0.1938
✅ 42_dns_issues_result_new_tools 60.5s 6 15 $0.4722
✅ 42_dns_issues_result_new_tools 54.3s 6 16 $0.4136
✅ 42_dns_issues_result_new_tools 69.6s 7 19 $0.5410
✅ 42_dns_issues_result_new_tools_no_runbook 49.6s 6 16 $0.4040
✅ 42_dns_issues_result_new_tools_no_runbook 51.1s 6 13 $0.4554
✅ 42_dns_issues_result_new_tools_no_runbook 59.2s 7 15 $0.4417
✅ 42_dns_issues_result_old_tools 54.2s 5 13 $0.8150
✅ 42_dns_issues_result_old_tools 62.8s 7 15 $1.0962
✅ 42_dns_issues_result_old_tools 58.2s 6 13 $0.8082
✅ 42_dns_issues_steps_new_all_tools 68.2s 7 19 $0.8984
✅ 42_dns_issues_steps_new_all_tools 68.1s 6 18 $0.8292
✅ 42_dns_issues_steps_new_all_tools 78.1s 9 15 $0.8935
✅ 42_dns_issues_steps_new_tools 51.0s 6 14 $0.3825
✅ 42_dns_issues_steps_new_tools 67.7s 6 17 $0.5120
✅ 42_dns_issues_steps_new_tools 58.7s 6 16 $0.4193
✅ 42_dns_issues_steps_old_tools 67.7s 7 15 $0.8313
✅ 42_dns_issues_steps_old_tools 61.7s 7 13 $0.6256
✅ 42_dns_issues_steps_old_tools 59.5s 7 14 $0.9087
✅ 43_current_datetime_from_prompt 5.5s 1 — $0.1850
✅ 43_current_datetime_from_prompt 5.9s 1 — $0.0180
✅ 43_current_datetime_from_prompt 4.9s 1 — $0.1844
➖ 43_slack_deployment_logs — — — —
➖ 43_slack_deployment_logs — — — —
➖ 43_slack_deployment_logs — — — —
➖ 44_slack_statefulset_logs — — — —
➖ 44_slack_statefulset_logs — — — —
➖ 44_slack_statefulset_logs — — — —
✅ 45_fetch_deployment_logs_simple 37.1s 6 8 $0.4675
✅ 45_fetch_deployment_logs_simple 35.6s 6 8 $0.4812
✅ 45_fetch_deployment_logs_simple 35.8s 6 8 $0.4740
➖ 46_job_crashing_no_longer_exists — — — —
➖ 46_job_crashing_no_longer_exists — — — —
➖ 46_job_crashing_no_longer_exists — — — —
➖ 47_truncated_logs_context_window — — — —
➖ 47_truncated_logs_context_window — — — —
➖ 47_truncated_logs_context_window — — — —
➖ 48_logs_since_thursday — — — —
➖ 48_logs_since_thursday — — — —
➖ 48_logs_since_thursday — — — —
🔧 49_logs_since_last_week 27.0s 5 6 —
🔧 49_logs_since_last_week 27.2s 5 6 —
✅ 49_logs_since_last_week 25.0s 4 5 $0.2689
✅ 50_logs_since_specific_date 24.2s 3 3 $0.2432
✅ 50_logs_since_specific_date 24.7s 3 3 $0.2426
✅ 50_logs_since_specific_date 23.8s 3 3 $0.2444
➖ 50a_logs_since_last_specific_month — — — —
➖ 50a_logs_since_last_specific_month — — — —
➖ 50a_logs_since_last_specific_month — — — —
🚧 51_logs_summarize_errors — — — —
🚧 51_logs_summarize_errors — — — —
🚧 51_logs_summarize_errors — — — —
✅ 52_logs_login_issues 69.0s 10 11 $0.8311
✅ 52_logs_login_issues 58.2s 7 8 $0.7231
✅ 52_logs_login_issues 81.3s 11 15 $0.9250
🚧 53_logs_find_term — — — —
🚧 53_logs_find_term — — — —
🚧 53_logs_find_term — — — —
✅ 54_not_truncated_when_getting_pods 37.4s 5 9 $0.3263
✅ 54_not_truncated_when_getting_pods 29.2s 4 8 $0.2990
✅ 54_not_truncated_when_getting_pods 33.9s 4 9 $0.3104
➖ 55_kafka_runbook — — — —
➖ 55_kafka_runbook — — — —
➖ 55_kafka_runbook — — — —
✅ 57_cluster_name_confusion 55.6s 7 11 $0.7019
✅ 57_cluster_name_confusion 91.6s 15 18 $0.9903
✅ 57_cluster_name_confusion 82.6s 13 19 $0.9572
❌ 57_wrong_namespace 45.0s 7 9 $0.6826
❌ 57_wrong_namespace 40.4s 7 8 $0.5639
✅ 57_wrong_namespace 37.8s 7 8 $0.5228
🚧 58_counting_pods_by_status — — — —
🚧 58_counting_pods_by_status — — — —
🚧 58_counting_pods_by_status — — — —
🚧 59_label_based_counting — — — —
🚧 59_label_based_counting — — — —
🚧 59_label_based_counting — — — —
🚧 60_count_less_than — — — —
🚧 60_count_less_than — — — —
🚧 60_count_less_than — — — —
🚧 61_exact_match_counting — — — —
🚧 61_exact_match_counting — — — —
🚧 61_exact_match_counting — — — —
❌ 62_fetch_error_logs_with_errors 28.1s 5 6 $0.4115
❌ 62_fetch_error_logs_with_errors 29.9s 5 6 $0.4255
❌ 62_fetch_error_logs_with_errors 28.6s 5 6 $0.4685
🚧 63_fetch_error_logs_no_errors — — — —
🚧 63_fetch_error_logs_no_errors — — — —
🚧 63_fetch_error_logs_no_errors — — — —
🚧 64_keda_vs_hpa_confusion — — — —
🚧 64_keda_vs_hpa_confusion — — — —
🚧 64_keda_vs_hpa_confusion — — — —
❌ 65_health_check_followup 57.9s 7 18 $0.4511
❌ 65_health_check_followup 66.9s 7 20 $0.5963
❌ 65_health_check_followup 65.1s 7 20 $0.4801
❌ 66_http_error_needle 56.1s 6 11 $0.5453
❌ 66_http_error_needle 55.1s 6 11 $0.5536
❌ 66_http_error_needle 64.0s 8 10 $0.5985
❌ 67_performance_degradation 47.1s 6 11 $0.6067
❌ 67_performance_degradation 41.4s 6 9 $0.5111
❌ 67_performance_degradation 55.7s 7 11 $0.6478
❌ 68_cascading_failures 56.1s 7 11 $0.6432
❌ 68_cascading_failures 55.6s 6 12 $0.6083
❌ 68_cascading_failures 56.1s 7 11 $0.6572
❌ 69_rate_limit_exhaustion 57.6s 7 12 $0.6600
❌ 69_rate_limit_exhaustion 53.8s 6 12 $0.6682
❌ 69_rate_limit_exhaustion 57.0s 7 12 $0.6117
❌ 70_memory_leak_detection 56.3s 7 11 $0.6425
❌ 70_memory_leak_detection 54.4s 7 11 $0.6175
❌ 70_memory_leak_detection 51.3s 7 11 $0.6459
❌ 71_connection_pool_starvation 53.1s 7 11 $0.6270
❌ 71_connection_pool_starvation 44.5s 6 10 $0.4377
❌ 71_connection_pool_starvation 54.6s 7 11 $0.6004
🚧 73a_time_window_anomaly — — — —
🚧 73a_time_window_anomaly — — — —
🚧 73a_time_window_anomaly — — — —
🚧 73b_time_window_anomaly — — — —
🚧 73b_time_window_anomaly — — — —
🚧 73b_time_window_anomaly — — — —
❌ 74_config_change_impact 56.2s 7 13 $0.6422
❌ 74_config_change_impact 58.2s 7 12 $0.5994
❌ 74_config_change_impact 59.3s 6 12 $0.6279
❌ 75_network_flapping 56.4s 7 11 $0.6459
❌ 75_network_flapping 53.2s 7 11 $0.6536
❌ 75_network_flapping 40.1s 6 9 $0.4634
✅ 76_service_discovery_issue 43.2s 6 11 $0.3609
✅ 76_service_discovery_issue 46.1s 6 11 $0.5243
✅ 76_service_discovery_issue 51.1s 5 15 $0.5497
✅ 77_liveness_probe_misconfiguration 44.3s 6 10 $0.4888
✅ 77_liveness_probe_misconfiguration 55.9s 7 10 $0.5615
✅ 77_liveness_probe_misconfiguration 35.7s 5 8 $0.3268
✅ 78a_missing_cpu_limits 43.8s 6 13 $0.3697
✅ 78a_missing_cpu_limits 34.2s 5 9 $0.3286
✅ 78a_missing_cpu_limits 38.0s 5 10 $0.3471
✅ 78b_cpu_quota_exceeded 37.5s 5 10 $0.3345
✅ 78b_cpu_quota_exceeded 36.7s 5 11 $0.3443
✅ 78b_cpu_quota_exceeded 34.0s 5 9 $0.3237
❌ 79_configmap_mount_issue 54.2s 7 11 $0.6085
❌ 79_configmap_mount_issue 52.2s 7 11 $0.6419
❌ 79_configmap_mount_issue 49.5s 6 10 $0.5993
✅ 80_pvc_storage_class_mismatch 48.8s 7 15 $0.5698
✅ 80_pvc_storage_class_mismatch 46.1s 7 15 $0.3899
✅ 80_pvc_storage_class_mismatch 43.0s 6 15 $0.3630
✅ 81_service_account_permission_denied 58.3s 9 15 $0.6104
✅ 81_service_account_permission_denied 50.5s 7 14 $0.5432
✅ 81_service_account_permission_denied 72.5s 11 16 $0.7164
❌ 82_pod_anti_affinity_conflict 44.4s 5 9 $0.4292
✅ 82_pod_anti_affinity_conflict 59.2s 7 13 $0.6782
✅ 82_pod_anti_affinity_conflict 42.9s 5 9 $0.4440
✅ 83_secret_not_found 39.1s 6 10 $0.4810
✅ 83_secret_not_found 39.5s 6 10 $0.5217
✅ 83_secret_not_found 40.8s 5 10 $0.5346
❌ 84_network_policy_blocking_traffic 55.4s 6 12 $0.4677
❌ 84_network_policy_blocking_traffic 59.3s 7 13 $0.5049
✅ 84_network_policy_blocking_traffic 57.4s 8 15 $0.4310
✅ 85_hpa_not_scaling 37.4s 5 10 $0.3451
✅ 85_hpa_not_scaling 32.3s 4 8 $0.3099
✅ 85_hpa_not_scaling 35.6s 5 10 $0.3381
✅ 86_configmap_like_but_secret 34.5s 5 10 $0.3330
✅ 86_configmap_like_but_secret 36.5s 5 10 $0.3316
✅ 86_configmap_like_but_secret 43.8s 6 13 $0.3630
✅ 89_runbook_missing_cloudwatch 39.2s 5 7 $0.3310
✅ 89_runbook_missing_cloudwatch 33.9s 4 6 $0.3047
✅ 89_runbook_missing_cloudwatch 31.7s 4 5 $0.3064
✅ 90_runbook_basic_selection 63.3s 8 16 $0.5603
✅ 90_runbook_basic_selection 127.3s 14 33 $1.6329
✅ 90_runbook_basic_selection 97.8s 11 28 $1.2775
✅ 91a_datadog_metrics_no_k8s 36.3s 6 8 $0.5015
✅ 91a_datadog_metrics_no_k8s 42.2s 7 9 $0.6871
✅ 91a_datadog_metrics_no_k8s 37.6s 6 8 $0.4819
✅ 91b_datadog_metrics_pod_exists 22.4s 4 3 $0.4043
✅ 91b_datadog_metrics_pod_exists 23.6s 4 4 $0.4175
✅ 91b_datadog_metrics_pod_exists 23.9s 4 5 $0.3859
✅ 91c_datadog_metrics_deployment[0] 28.1s 5 7 $0.3032
✅ 91c_datadog_metrics_deployment[0] 33.7s 6 7 $0.4623
✅ 91c_datadog_metrics_deployment[0] 31.0s 5 8 $0.3116
✅ 91c_datadog_metrics_deployment[1] 32.8s 5 8 $0.4588
✅ 91c_datadog_metrics_deployment[1] 31.3s 5 7 $0.4634
✅ 91c_datadog_metrics_deployment[1] 29.0s 4 6 $0.4321
✅ 91c_datadog_metrics_deployment[2] 29.9s 5 7 $0.3145
✅ 91c_datadog_metrics_deployment[2] 35.9s 6 9 $0.3438
✅ 91c_datadog_metrics_deployment[2] 33.7s 5 8 $0.4836
✅ 91d_datadog_metrics_historical_pod[0] 23.0s 4 4 $0.2711
✅ 91d_datadog_metrics_historical_pod[0] 18.1s 3 3 $0.2536
✅ 91d_datadog_metrics_historical_pod[0] 18.0s 3 3 $0.2532
✅ 91d_datadog_metrics_historical_pod[1] 24.2s 4 5 $0.2882
✅ 91d_datadog_metrics_historical_pod[1] 20.2s 4 4 $0.2694
✅ 91d_datadog_metrics_historical_pod[1] 23.7s 4 5 $0.2824
✅ 91d_datadog_metrics_historical_pod[2] 29.1s 4 7 $0.3003
✅ 91d_datadog_metrics_historical_pod[2] 22.5s 3 4 $0.2684
✅ 91d_datadog_metrics_historical_pod[2] 29.3s 4 6 $0.3003
✅ 91e_datadog_custom_metrics[0] 34.7s 5 9 $0.4400
✅ 91e_datadog_custom_metrics[0] 35.7s 5 9 $0.4809
✅ 91e_datadog_custom_metrics[0] 29.8s 4 7 $0.4224
✅ 91e_datadog_custom_metrics[1] 18.2s 3 3 $0.2476
✅ 91e_datadog_custom_metrics[1] 23.6s 4 5 $0.3982
✅ 91e_datadog_custom_metrics[1] 25.3s 4 5 $0.4008
✅ 91f_datadog_logs_historical_pod 59.0s 7 11 $0.5784
✅ 91f_datadog_logs_historical_pod 53.4s 7 10 $0.5313
✅ 91f_datadog_logs_historical_pod 58.5s 6 11 $0.5909
✅ 91g_datadog_metrics_mismatched_pod[0] 26.7s 5 6 $0.3036
✅ 91g_datadog_metrics_mismatched_pod[0] 24.0s 4 5 $0.2837
✅ 91g_datadog_metrics_mismatched_pod[0] 24.0s 4 5 $0.2816
✅ 91g_datadog_metrics_mismatched_pod[1] 22.9s 4 5 $0.2812
✅ 91g_datadog_metrics_mismatched_pod[1] 21.4s 4 5 $0.2794
✅ 91g_datadog_metrics_mismatched_pod[1] 21.1s 4 5 $0.2760
✅ 91h_datadog_logs_empty_query_with_url 18.9s 3 3 $0.2421
✅ 91h_datadog_logs_empty_query_with_url 17.4s 3 3 $0.2387
✅ 91h_datadog_logs_empty_query_with_url 17.2s 3 3 $0.2394
➖ 91i_datadog_metrics_empty_query_with_url — — — —
➖ 91i_datadog_metrics_empty_query_with_url — — — —
➖ 91i_datadog_metrics_empty_query_with_url — — — —
✅ 92_cpu_graph_conversation[0] 12.0s 2 1 $0.2245
✅ 92_cpu_graph_conversation[0] 12.3s 2 1 $0.2244
✅ 92_cpu_graph_conversation[0] 11.3s 2 1 $0.2228
✅ 92_cpu_graph_conversation[1] 27.5s 5 6 $0.2949
✅ 92_cpu_graph_conversation[1] 14.0s 3 2 $0.2412
✅ 92_cpu_graph_conversation[1] 12.0s 2 2 $0.2256
✅ 92_cpu_graph_conversation[2] 17.1s 3 3 $0.2462
✅ 92_cpu_graph_conversation[2] 12.5s 2 2 $0.2263
✅ 92_cpu_graph_conversation[2] 12.6s 2 2 $0.2259
➖ 93_events_since_specific_date — — — —
➖ 93_events_since_specific_date — — — —
➖ 93_events_since_specific_date — — — —
✅ 94_runbook_transparency 78.8s 9 22 $0.8628
✅ 94_runbook_transparency 76.5s 9 19 $0.8853
✅ 94_runbook_transparency 63.3s 6 13 $0.5830
❌ 95_runbook_memory_leak_detection 66.0s 7 13 $0.6430
✅ 95_runbook_memory_leak_detection 70.0s 8 15 $0.6321
✅ 95_runbook_memory_leak_detection 62.2s 6 12 $0.5952
✅ 96_no_matching_runbook 59.2s 6 14 $0.6402
✅ 96_no_matching_runbook 52.6s 6 12 $0.5292
✅ 96_no_matching_runbook 54.6s 6 14 $0.4874
✅ 97_logs_clarification_needed 6.5s 1 — $0.1867
✅ 97_logs_clarification_needed 6.5s 1 — $0.1871
✅ 97_logs_clarification_needed 5.9s 1 — $0.1861
✅ 99_logs_transparency_custom_time 33.1s 5 7 $0.4186
✅ 99_logs_transparency_custom_time 32.0s 5 9 $0.4508
✅ 99_logs_transparency_custom_time 34.3s 5 7 $0.4552
✅ 01_oom_kill 59.5s 7 16 —
✅ 01_oom_kill 58.4s 7 18 —
✅ 01_oom_kill 58.8s 7 18 —
✅ 02_crashloop_backoff 37.9s 4 8 —
✅ 02_crashloop_backoff 39.0s 4 9 —
✅ 02_crashloop_backoff 31.2s 3 7 —
✅ 03_cpu_throttling 45.9s 5 9 —
✅ 03_cpu_throttling 59.3s 7 12 —
✅ 03_cpu_throttling 52.7s 5 11 —
✅ 04_image_pull_backoff 42.0s 5 8 —
✅ 04_image_pull_backoff 41.1s 4 9 —
✅ 04_image_pull_backoff 45.5s 6 9 —
✅ 05_crashpod 64.1s 6 18 —
✅ 05_crashpod 67.4s 7 17 —
✅ 05_crashpod 60.2s 7 16 —
✅ 06_job_failure 49.7s 5 13 —
✅ 06_job_failure 56.3s 6 15 —
✅ 06_job_failure 43.9s 5 12 —
✅ 07_job_syntax_error 44.9s 5 14 —
✅ 07_job_syntax_error 50.5s 5 14 —
✅ 07_job_syntax_error 51.4s 6 14 —
✅ 08_memory_pressure 65.7s 9 17 —
✅ 08_memory_pressure 56.8s 8 16 —
✅ 08_memory_pressure 68.9s 11 18 —
✅ 09_high_latency 50.7s 5 10 —
✅ 09_high_latency 59.4s 6 13 —
✅ 09_high_latency 57.8s 6 13 —
✅ 10_KubeDeploymentReplicasMismatch 51.2s 5 10 —
✅ 10_KubeDeploymentReplicasMismatch 49.2s 5 10 —
✅ 10_KubeDeploymentReplicasMismatch 48.8s 5 10 —
✅ 11_KubePodCrashLooping 43.3s 5 11 —
✅ 11_KubePodCrashLooping 44.2s 5 12 —
✅ 11_KubePodCrashLooping 39.5s 4 10 —
✅ 12_KubePodNotReady 50.7s 5 12 —
✅ 12_KubePodNotReady 52.2s 6 13 —
✅ 12_KubePodNotReady 53.9s 6 13 —
✅ 13_Watchdog 42.9s 6 10 —
✅ 13_Watchdog 45.4s 6 11 —
✅ 13_Watchdog 37.7s 5 10 —
❌ 14_tempo 92.7s 12 33 —
❌ 14_tempo 88.2s 11 36 —
❌ 14_tempo 99.0s 13 34 —
✅ 15_dns_resolution 61.2s 8 14 —
⚠️ 15_dns_resolution 61.8s 6 17 —
⚠️ 15_dns_resolution 54.5s 5 15 —
❌ 16_dns_resolution_no_tool 49.2s 6 12 —
❌ 16_dns_resolution_no_tool 49.1s 6 12 —
❌ 16_dns_resolution_no_tool 50.4s 6 13 —
✅ 17_investigate_correct_date 62.8s 8 19 —
✅ 17_investigate_correct_date 66.2s 8 19 —
❌ 17_investigate_correct_date 61.9s 7 19 —
Total 41.7s avg 5.7 avg 10.2 avg $191.1961

⚠️ 90 Failures Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref master -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref master -f markers=regression -f filter=

claude and others added 3 commits February 16, 2026 22:20
Instead of asking the LLM to batch text with tool calls (unreliable),
tell it to simply not call TodoWrite when the last task is done.
The final answer itself signals completion, saving one full LLM
round trip (3-5 seconds).

https://claude.ai/code/session_01D3bToEw4CfTUkPWhVGsHCN
Signed-off-by: Claude <noreply@anthropic.com>
@aantn

aantn commented Feb 17, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
iterations: 3

@github-actions

This comment was marked as outdated.

@aantn

aantn commented Feb 17, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
marker: regression
branch: master

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ✅ Completed successfully


🧪 Manual Eval Results

Parameter Value
Triggered via /eval on branch master
Branch master
Model opus-4.5
Markers regression
Iterations 1
Duration 4m 4s
Workflow View logs | Rerun

Results of HolmesGPT evals (branch: master)

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 27.4s 4 9 $0.2002
✅ 101_loki_historical_logs_pod_deleted 47.8s 5 10 $0.2513
✅ 111_pod_names_contain_service 32.3s 5 11 $0.2225
✅ 112_find_pvcs_by_uuid 37.5s 7 8 $0.2490
✅ 12_job_crashing 57.4s 5 12 $0.2420
✅ 176_network_policy_blocking_traffic_no_runbooks 53.9s 9 16 $0.3268
✅ 24_misconfigured_pvc 33.2s 5 13 $0.2315
✅ 43_current_datetime_from_prompt 5.9s 1 — $0.1064
✅ 61_exact_match_counting 13.6s 3 2 $0.1422
Total 34.3s avg 4.9 avg 10.1 avg $1.9719
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref master -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref master -f markers=regression -f filter=

The "Final Review Phase" / "Investigation Verification" was a separate
tracked task that forced an extra TodoWrite round trip just to mark it
in_progress then completed. Convert it to a lightweight mental check
("Mentally verify, do NOT create a task for this") that the LLM does
before answering. This eliminates the extra LLM call the verification
task was causing.

https://claude.ai/code/session_01D3bToEw4CfTUkPWhVGsHCN
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
holmes/plugins/prompts/investigation_procedure.jinja2 (1)

52-77: ⚠️ Potential issue | 🟠 Major

Enforcement rules at lines 54, 61, and 77 directly contradict the new Efficiency Rule — update the carve-out.

The new EFFICIENCY RULE (line 62) tells the LLM to skip the final TodoWrite completion call and provide the answer directly. However, the three unchanged enforcement statements directly block this path:

Line Statement Conflict
54 "Verify ALL tasks show 'completed' status" Last task never reaches completed
61 "Only after ALL tasks are 'completed': Proceed to…final answer" Same — hard gate on all-completed
77 "If you see ANY [ ] pending or [~] in_progress tasks, DO NOT provide final answer" Last task is still in_progress; this check explicitly blocks the answer

An LLM trying to reconcile these will either silently fall back to the old behaviour (making the optimisation ineffective) or violate the enforcement rules unpredictably.

The enforcement block and the Task Status Check example (lines 71–77) should be updated to carve out the single-remaining-task exception, e.g.:

✏️ Suggested patch
-1. **Check TodoWrite status**: Verify ALL tasks show "completed" status
-2. **If ANY task is "pending" or "in_progress"**:
+1. **Check TodoWrite status**: Verify ALL tasks show "completed" status — **EXCEPT** when the single last task is in_progress (see EFFICIENCY RULE `#4` below).
+2. **If ANY task is "pending" or "in_progress"** (and there is MORE THAN ONE remaining task):
-If you see ANY `[ ] pending` or `[~] in_progress` tasks, DO NOT provide final answer.
+If you see ANY `[ ] pending` or `[~] in_progress` tasks (and more than one task remains), DO NOT provide final answer.
+If only ONE task is in_progress and you have completed the work, apply the EFFICIENCY RULE: provide your final answer directly.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/plugins/prompts/investigation_procedure.jinja2` around lines 52 - 77,
Update the enforcement block in investigation_procedure.jinja2 to explicitly
carve out the EFFICIENCY RULE exception: modify the three enforcement statements
that require "ALL tasks show 'completed' status", "Only after ALL tasks are
'completed': Proceed...", and "If you see ANY [ ] pending or [~] in_progress
tasks, DO NOT provide final answer" so they allow a single last-task exception
when the EFFICIENCY RULE applies (i.e., if exactly one task remains and it is
in_progress AND the agent has finished the work, the agent may provide the final
answer instead of calling TodoWrite to mark it completed). Also update the "Task
Status Check Example" block to reflect this exception (show an example where the
last task is [~] in_progress but final answer is allowed under the EFFICIENCY
RULE). Reference the enforcement block and the "Task Status Check Example" text
in this template when making the edits.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@holmes/plugins/prompts/investigation_procedure.jinja2`:
- Around line 52-77: Update the enforcement block in
investigation_procedure.jinja2 to explicitly carve out the EFFICIENCY RULE
exception: modify the three enforcement statements that require "ALL tasks show
'completed' status", "Only after ALL tasks are 'completed': Proceed...", and "If
you see ANY [ ] pending or [~] in_progress tasks, DO NOT provide final answer"
so they allow a single last-task exception when the EFFICIENCY RULE applies
(i.e., if exactly one task remains and it is in_progress AND the agent has
finished the work, the agent may provide the final answer instead of calling
TodoWrite to mark it completed). Also update the "Task Status Check Example"
block to reflect this exception (show an example where the last task is [~]
in_progress but final answer is allowed under the EFFICIENCY RULE). Reference
the enforcement block and the "Task Status Check Example" text in this template
when making the edits.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants