Skip to content

Add three LLM overconfidence test cases for Elasticsearch - #1831

Open
aantn wants to merge 1 commit into
masterfrom
claude/overconfidence-evals-0mQcy
Open

aantn wants to merge 1 commit into
masterfrom
claude/overconfidence-evals-0mQcy

Conversation

@aantn

@aantn aantn commented Mar 22, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR adds three new test cases to the Holmes LLM test suite that validate the system's ability to express appropriate uncertainty and avoid overconfident conclusions when analyzing incomplete or ambiguous data.

Key Changes

New Test Cases

  1. Test 212 - Incomplete Evidence: Validates that Holmes acknowledges HTTP 500 errors but expresses uncertainty about root cause when only HTTP access logs are available (no application logs, traces, or diagnostics)

  2. Test 213 - Ambiguous Root Cause: Tests that Holmes presents multiple hypotheses rather than claiming a single definitive root cause when three services fail simultaneously with no dependency topology information

  3. Test 214 - Correlation Not Causation: Ensures Holmes doesn't assume a deployment caused errors in another service just because they occurred in temporal proximity, without evidence of a causal link

Implementation Details

  • Each test includes:

    • A test_case.yaml file with setup/teardown logic that creates Elasticsearch indices with realistic log data
    • A toolsets.yaml configuration enabling Elasticsearch data and cluster tools while disabling Kubernetes-specific tools
    • Clear comments explaining the test scenario and what constitutes overconfidence vs. well-calibrated responses
  • Test data is carefully crafted to:

    • Provide sufficient information to identify the problem
    • Deliberately omit diagnostic information needed to determine root cause
    • Create scenarios where multiple explanations are equally plausible
  • Expected outputs use specific language requirements to validate appropriate uncertainty:

    • Hedging language ("possible", "could be", "might", "unclear")
    • Acknowledgment of insufficient data
    • Recommendations for additional investigation

These tests help ensure Holmes maintains calibrated confidence levels and avoids the common LLM pitfall of fabricating explanations when data is incomplete or ambiguous.

https://claude.ai/code/session_01JqXGb7TpzRuZgnKDrzgMXf

Summary by CodeRabbit

  • Tests
    • Added a pytest marker "overconfidence" to test configuration.
    • Added three Elasticsearch-backed test scenarios: "Incomplete Evidence", "Ambiguous Cause", and "Correlation vs. Causation" that require the model to acknowledge uncertainty, avoid asserting unsupported root causes, and recommend further investigation.
    • Added test toolset configurations to enable Elasticsearch access and disable Kubernetes/Helm dependencies for these fixtures.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review.

@coderabbitai

coderabbitai Bot commented Mar 22, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 54958075-a01e-4f86-b76c-666ca84e8245

📥 Commits

Reviewing files that changed from the base of the PR and between 920936a8f1147479035ec3fade26c2eb64520060 and 9a4af76.

📒 Files selected for processing (7)
  • pyproject.toml
  • tests/llm/fixtures/test_ask_holmes/212_elasticsearch_overconfidence_incomplete_evidence/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/212_elasticsearch_overconfidence_incomplete_evidence/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/213_elasticsearch_overconfidence_ambiguous_cause/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/213_elasticsearch_overconfidence_ambiguous_cause/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/214_elasticsearch_overconfidence_correlation_not_causation/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/214_elasticsearch_overconfidence_correlation_not_causation/toolsets.yaml
✅ Files skipped from review due to trivial changes (6)
  • pyproject.toml
  • tests/llm/fixtures/test_ask_holmes/212_elasticsearch_overconfidence_incomplete_evidence/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/214_elasticsearch_overconfidence_correlation_not_causation/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/214_elasticsearch_overconfidence_correlation_not_causation/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/213_elasticsearch_overconfidence_ambiguous_cause/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/213_elasticsearch_overconfidence_ambiguous_cause/test_case.yaml
🚧 Files skipped from review as they are similar to previous changes (1)
  • tests/llm/fixtures/test_ask_holmes/212_elasticsearch_overconfidence_incomplete_evidence/test_case.yaml

Walkthrough

Added an overconfidence pytest marker and three new Elasticsearch-backed LLM test fixtures (212, 213, 214) with corresponding toolset configs to validate model hedging/uncertainty behavior; each fixture includes ES setup, data indexing, verification, and teardown.

Changes

Cohort / File(s) Summary
Pytest Configuration
pyproject.toml
Added overconfidence entry to [tool.pytest.ini_options].markers (minor list formatting tweak).
Test Fixture 212 — Incomplete Evidence
tests/llm/fixtures/test_ask_holmes/212_elasticsearch_overconfidence_incomplete_evidence/test_case.yaml, .../toolsets.yaml
New fixture that creates app-212-access-logs with access-log mappings and six status_code: 500 entries; prompts model to hedge/express uncertainty; toolsets enable Elasticsearch (env-configured) and disable Kubernetes/Helm.
Test Fixture 213 — Ambiguous Root Cause
tests/llm/fixtures/test_ask_holmes/213_elasticsearch_overconfidence_ambiguous_cause/test_case.yaml, .../toolsets.yaml
New fixture inserting multi-service logs (18 docs across three services) expecting the model to present multiple hypotheses and request further diagnostics; toolsets enable Elasticsearch and disable Kubernetes/Helm.
Test Fixture 214 — Correlation Not Causation
tests/llm/fixtures/test_ask_holmes/214_elasticsearch_overconfidence_correlation_not_causation/test_case.yaml, .../toolsets.yaml
New fixture with deployment and downstream error logs (18 docs) requiring the model to avoid asserting causation from temporal proximity; includes ES setup/verification and teardown; toolsets enable Elasticsearch and disable Kubernetes/Helm.

Sequence Diagram(s)

sequenceDiagram
    participant Tester as Tester (test harness)
    participant ES as Elasticsearch
    participant Holmes as Holmes LLM
    Tester->>ES: delete index (if exists) / create index with mappings
    Tester->>ES: bulk index test documents
    Tester->>ES: refresh index and run count/validation queries
    Tester->>Holmes: invoke test prompt (query logs via ES toolset)
    Holmes->>ES: query logs (via elasticsearch/data or cluster toolset)
    Holmes-->>Tester: return analysis (must hedge / provide hypotheses)
    Tester->>ES: delete test index (teardown)
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested reviewers

  • moshemorad
  • Sheeproid
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The pull request title 'Add three LLM overconfidence test cases for Elasticsearch' accurately and directly summarizes the main change—adding three new test cases (212–214) to validate Holmes LLM behavior on overconfidence scenarios.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Mar 22, 2026 •

Copy link
Copy Markdown

❌ Deploy Preview for holmes-docs failed. Why did it fail? →

Name Link
🔨 Latest commit 9a4af76
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69c06ed65ecbc2000804c9ba

@github-actions

github-actions Bot commented Mar 22, 2026 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 9a4af76 on branch claude/overconfidence-evals-0mQcy

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/10 test cases were successful, 0 regressions, 1 setup failures
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 24.2s 5 9 $0.2267 102,946 101,184 22,739 1,762 669 78,149 23,035 — —
✅ 101_loki_historical_logs_pod_deleted 34.5s 5 10 $0.2658 110,348 107,877 24,900 2,471 880 81,433 26,444 — —
✅ 112_find_pvcs_by_uuid 14.8s 3 4 $0.1817 61,228 60,232 21,834 996 601 38,135 22,097 — —
✅ 12_job_crashing 27.4s 5 12 $0.2506 111,471 109,415 24,677 2,056 577 84,158 25,257 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 30.9s 6 9 $0.2622 129,706 127,575 24,639 2,131 480 102,156 25,419 — —
✅ 227_count_configmaps_per_namespace[0] 20.0s 5 9 $0.2032 95,541 94,297 20,990 1,244 580 72,365 21,932 — —
🚧 243_pod_names_contain_service — — — — — — — — — — — — —
✅ 24_misconfigured_pvc 29.0s 6 12 $0.2509 124,029 121,870 22,885 2,159 446 98,001 23,869 — —
✅ 43_current_datetime_from_prompt 3.6s 1 — $0.0117 17,203 17,078 17,078 125 125 17,068 10 — —
✅ 61_exact_match_counting 10.0s 3 3 $0.1411 53,410 52,992 18,098 418 271 34,883 18,109 — —
Total 21.6s avg 4.3 avg 8.5 avg $1.7940 805,882 792,520 24,900 13,362 880 606,348 186,172 — —

Benchmark comparison unavailable: No ci-benchmark experiments found

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No ci-benchmark experiments found

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/overconfidence-evals-0mQcy -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, overconfidence, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/overconfidence-evals-0mQcy -f markers=regression -f filter=

@github-actions

github-actions Bot commented Mar 22, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 79e569ff (built in 1m 4s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:79e569ff
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:79e569ff me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:79e569ff
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:79e569ff
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:79e569ff
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:79e569ff me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:79e569ff
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:79e569ff

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:79e569ff \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:79e569ff

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:79e569ff \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:79e569ff

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
tests/llm/fixtures/test_ask_holmes/212_elasticsearch_overconfidence_incomplete_evidence/test_case.yaml (1)

127-141: Consider adding total document count verification for consistency.

Tests 213 and 214 verify both total document count and specific filtered counts, but this test only verifies the error count. While verifying ERROR_COUNT=6 is sufficient for test data integrity, adding a total count check (expecting 15 documents) would improve consistency across the test suite and catch potential bulk insert issues.

Optional: Add total document count verification
   echo "Test index created with $DOC_COUNT total logs ($ERROR_COUNT with 500 status)"
 
+  if [ "$DOC_COUNT" != "15" ]; then
+    echo "Expected 15 total logs but found: $DOC_COUNT"
+    exit 1
+  fi
+
   if [ "$ERROR_COUNT" != "6" ]; then
     echo "Expected 6 error logs but found: $ERROR_COUNT"
     exit 1
   fi
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In
`@tests/llm/fixtures/test_ask_holmes/212_elasticsearch_overconfidence_incomplete_evidence/test_case.yaml`
around lines 127 - 141, Add a verification that the total document count
(DOC_COUNT) equals 15 before or alongside the existing ERROR_COUNT check: after
computing DOC_COUNT (from the curl to
"${ELASTICSEARCH_URL}/${HOLMES_ES_TEST_INDEX}/_count"), compare it to "15" and
if it does not match, echo a descriptive message and exit 1; keep existing
ERROR_COUNT logic intact so both DOC_COUNT and ERROR_COUNT (6) are validated for
the test.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In
`@tests/llm/fixtures/test_ask_holmes/212_elasticsearch_overconfidence_incomplete_evidence/test_case.yaml`:
- Around line 127-141: Add a verification that the total document count
(DOC_COUNT) equals 15 before or alongside the existing ERROR_COUNT check: after
computing DOC_COUNT (from the curl to
"${ELASTICSEARCH_URL}/${HOLMES_ES_TEST_INDEX}/_count"), compare it to "15" and
if it does not match, echo a descriptive message and exit 1; keep existing
ERROR_COUNT logic intact so both DOC_COUNT and ERROR_COUNT (6) are validated for
the test.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 8f9dcca8-8cee-4441-ad06-5a772ee69ae6

📥 Commits

Reviewing files that changed from the base of the PR and between 4110a0a and 552a58e.

📒 Files selected for processing (7)
  • pyproject.toml
  • tests/llm/fixtures/test_ask_holmes/212_elasticsearch_overconfidence_incomplete_evidence/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/212_elasticsearch_overconfidence_incomplete_evidence/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/213_elasticsearch_overconfidence_ambiguous_cause/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/213_elasticsearch_overconfidence_ambiguous_cause/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/214_elasticsearch_overconfidence_correlation_not_causation/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/214_elasticsearch_overconfidence_correlation_not_causation/toolsets.yaml

@github-actions

github-actions Bot commented Mar 22, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟢 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 11.34s 12.12s -6.4%
Warm Mean 5.14s 5.77s -10.9%
Warm Min 5.09s 5.62s
Warm Max 5.18s 5.85s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 18.46s 19.12s -3.5%
Warm Mean 8.30s 8.54s -2.8%
Warm Min 7.83s 8.48s
Warm Max 9.58s 8.65s

PR: 79e569ff | Master: 4110a0a4 | Iterations: 5

Three new Elasticsearch-based evals testing whether the LLM expresses
appropriate uncertainty when evidence is insufficient or ambiguous:

- 212: Incomplete evidence - HTTP access logs show 500 errors but contain
  no application-level diagnostics. LLM should not fabricate root causes.
  Fails 3/3 on Opus 4.5 (model confidently claims "upstream timeout" as
  root cause despite having only status codes and latencies).

- 213: Ambiguous cause - Three services fail simultaneously with no
  dependency topology info. LLM should present multiple hypotheses.
  Fails 3/3 on Opus 4.5 (model picks one service as THE definitive
  root cause based on 1-second timestamp differences).

- 214: Correlation not causation - Deployment of one service coincides
  temporally with errors in an unrelated service. LLM should note
  correlation without claiming causation.
  Passes 3/3 on Opus 4.5 (included as calibration baseline).

Also adds 'overconfidence' pytest marker to pyproject.toml.

https://claude.ai/code/session_01JqXGb7TpzRuZgnKDrzgMXf
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn force-pushed the claude/overconfidence-evals-0mQcy branch from 920936a to 9a4af76 Compare March 22, 2026 22:36

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants