Add cluster mismatch detection test for Holmes - #1801
Conversation
Tests that Holmes correctly identifies when it's connected to the wrong cluster (ci-cd) while investigating an alert from a different cluster (us-prod-eks). Uses Elasticsearch to simulate a deployment registry and change history. The test expects Holmes to refuse providing a root cause when the target app doesn't exist in the current cluster. This is a TDD "red" test - it should initially fail because Holmes currently provides analysis even when it detects a cluster mismatch. https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg Signed-off-by: Claude <noreply@anthropic.com>
Claude Code ReviewThis repository is configured for manual code reviews. Comment |
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughAdds a new Elasticsearch-backed test fixture and a toolsets reference to validate detection of alerts that reference a different cluster than the connected Elasticsearch data (cluster mismatch), including ES index setup, CI/CD-only test data, and expected response assertions. Changes
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. 📝 Coding Plan
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
📂 Previous Runs📜 Run @ 73ebf76 (#23292352513)✅ Results of HolmesGPT evalsAutomatically triggered by commit 73ebf76 on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 73 test/model combinations loaded Benchmark experiment:
No benchmark data available for comparison. Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
📜 Run @ 86bf58d (#23290503429)
|
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ❌ | 235_elasticsearch_cluster_mismatch | 124.4s | 19 | 30 | — | — | — | — | — | — | — | — | — | — |
| Total | 124.4s avg | 19.0 avg | 30.0 avg | — | — | — | — | — | — | — | — | — | — |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 73 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 1 Failure Detected
📜 Run @ ce2808e (#23290302902)
⚠️ Eval Results (with failures)
Automatically triggered by commit ce2808e on branch claude/tdd-cluster-mismatch-test-Bwa0K (labels: evals-id-235)
Results of HolmesGPT evals
- ask_holmes: 0/1 test cases were successful, 1 regressions
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ❌ | 235_elasticsearch_cluster_mismatch | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Total | — avg | — avg | — avg | — | — | — | — | — | — | — | — | — | — |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 73 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 1 Failure Detected
📜 Run @ fc2a315 (#23289012717)
⚠️ Eval Results (with failures)
Automatically triggered by commit fc2a315 on branch claude/tdd-cluster-mismatch-test-Bwa0K (labels: evals-id-235)
Results of HolmesGPT evals
- ask_holmes: 0/1 test cases were successful, 1 regressions
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ❌ | 235_elasticsearch_cluster_mismatch | 98.3s | 14 | 26 | — | — | — | — | — | — | — | — | — | — |
| Total | 98.3s avg | 14.0 avg | 26.0 avg | — | — | — | — | — | — | — | — | — | — |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 7 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 1 Failure Detected
📜 Run @ 02d1779 (#23288784512)
⚠️ Eval Results (with failures)
Automatically triggered by commit 02d1779 on branch claude/tdd-cluster-mismatch-test-Bwa0K (labels: evals-id-235)
Results of HolmesGPT evals
- ask_holmes: 0/1 test cases were successful, 1 regressions
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ❌ | 235_elasticsearch_cluster_mismatch | 44.9s | 7 | 14 | $0.3217 | 173,939 | 170,900 | 27,509 | 3,039 | 843 | 142,923 | 27,977 | — | — |
| Total | 44.9s avg | 7.0 avg | 14.0 avg | $0.3217 | 173,939 | 170,900 | 27,509 | 3,039 | 843 | 142,923 | 27,977 | — | — |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 73 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 1 Failure Detected
⚠️ Eval Results (with failures)
Automatically triggered by commit 062e55c on branch claude/tdd-cluster-mismatch-test-Bwa0K (labels: evals-id-235)
Results of HolmesGPT evals
- ask_holmes: 0/1 test cases were successful, 1 regressions
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ❌ | 235_elasticsearch_cluster_mismatch | 117.5s | 17 | 32 | $0.5476 | 379,554 | 372,434 | 30,494 | 7,120 | 745 | 340,155 | 32,279 | — | — |
| Total | 117.5s avg | 17.0 avg | 32.0 avg | $0.5476 | 379,554 | 372,434 | 30,494 | 7,120 | 745 | 340,155 | 32,279 | — | — |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 73 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 1 Failure Detected
📖 Legend
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tdd-cluster-mismatch-test-Bwa0K -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
tags: regression
Or with more options (one per line):
/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
tags: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
tags |
Pytest tags / markers (no default - runs all tests!) |
id |
Eval ID / pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):
| Label | Effect |
|---|---|
evals-tag-<name> |
Run tests with tag <name> alongside regression |
evals-id-<name> |
Run a specific eval by test ID |
evals-model-<name> |
Override the model (use model list name, e.g. sonnet-4.5) |
Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5
🏷️ Valid tags
benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, integration, kafka, kubernetes, leaked-information, logs, loki, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency
🤖 Valid models
deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tdd-cluster-mismatch-test-Bwa0K -f markers=regression -f filter=
There was a problem hiding this comment.
🧹 Nitpick comments (1)
tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml (1)
71-74: Consider using-sSinstead of-sffor initial cleanup.The
-fflag causes curl to fail silently on HTTP errors, but combined with|| true, the error is already suppressed. Using-sSwould still suppress progress output while showing errors (though they're intentionally ignored here via|| true). This is a minor consistency point - the current approach works correctly.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml` around lines 71 - 74, Replace the curl invocations that use the `-sf` flags with `-sS` so errors are still reported (while keeping output silent) — update the two lines invoking curl with `"${ELASTICSEARCH_URL}/${INVENTORY_INDEX}"` and `"${ELASTICSEARCH_URL}/${CHANGES_INDEX}"` to use `-sS` instead of `-sf`, and keep the existing `-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" || true` behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Nitpick comments:
In
`@tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml`:
- Around line 71-74: Replace the curl invocations that use the `-sf` flags with
`-sS` so errors are still reported (while keeping output silent) — update the
two lines invoking curl with `"${ELASTICSEARCH_URL}/${INVENTORY_INDEX}"` and
`"${ELASTICSEARCH_URL}/${CHANGES_INDEX}"` to use `-sS` instead of `-sf`, and
keep the existing `-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" || true`
behavior.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 08969301-b9f8-4624-a9fd-c4f5dc1f61a3
📒 Files selected for processing (2)
tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yamltests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/toolsets.yaml
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:6e8d851b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:6e8d851b me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:6e8d851b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:6e8d851b
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:6e8d851b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:6e8d851b me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:6e8d851b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:6e8d851bPatch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:6e8d851b \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:6e8d851bRobusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:6e8d851b \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:6e8d851b |
The eval must simulate the real scenario where Holmes can only see data from the current cluster (ci-cd). Previously the ES data included us-prod-eks deployment records, which let Holmes find the app and provide analysis — the wrong failure mode. Now the ES data contains only: - ci-cd deployments (jenkins, argocd, gitlab-runner, flux, tekton) - An empty us-prod-eks namespace (red herring) - Change history with only CI/CD tool changes (no alert-related changes) Holmes should search for lambda-dashboard-monitoring/clientorg-alpha, find nothing, and conclude it's in the wrong cluster — not provide a speculative root cause. https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 2
🧹 Nitpick comments (1)
tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml (1)
68-70: Enable fail-fast shell mode before setup calls
set -eis applied aftersourceandes_setup. If either fails, the script may continue in a partially initialized state.Proposed diff
before_test: | + set -euo pipefail source ../../shared/es_test_utils.sh es_setup - set -e🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml` around lines 68 - 70, The script enables fail-fast mode too late—move the guard so that `set -e` (or `set -euo pipefail` if desired) is invoked before sourcing `../../shared/es_test_utils.sh` and before the `es_setup` call so any failure in `source` or `es_setup` aborts immediately; ensure `set -e` is placed at the top (right after the shebang) so `source ../../shared/es_test_utils.sh` and `es_setup` cannot leave the script in a partially initialized state.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In
`@tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml`:
- Around line 53-59: Update the expected_output to require at least one concrete
seeded value and explicit zero-hit evidence: include the seeded identifier
BUILD-7X4Q9M (and its app/version context) as a required discovery and also
require an explicit statement that searching for the alerting app returned zero
hits (e.g., "no matching results for alerting app"). Modify the expected_output
array entries (the list under expected_output) so one item asserts the
discovered BUILD-7X4Q9M value and another asserts explicit zero-hit evidence for
the alerting app, while keeping the existing cluster-mismatch and
investigation-recommendation checks intact.
- Around line 72-74: The test uses fixed Elasticsearch index names
INVENTORY_INDEX and CHANGES_INDEX in the fixture and its before_test/after_test
hooks, which can collide across parallel runs; update the before_test to create
and the test to reference run-scoped index names (e.g., append a unique run ID
or timestamp to INVENTORY_INDEX and CHANGES_INDEX) and change after_test to only
delete those run-scoped indices (or skip deletion of shared names) so setup is
idempotent and teardown doesn’t remove shared resources; apply the same change
pattern to the other affected fixtures/tests (cases 227–232).
---
Nitpick comments:
In
`@tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml`:
- Around line 68-70: The script enables fail-fast mode too late—move the guard
so that `set -e` (or `set -euo pipefail` if desired) is invoked before sourcing
`../../shared/es_test_utils.sh` and before the `es_setup` call so any failure in
`source` or `es_setup` aborts immediately; ensure `set -e` is placed at the top
(right after the shebang) so `source ../../shared/es_test_utils.sh` and
`es_setup` cannot leave the script in a partially initialized state.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 4c491c3a-b02e-4888-b8db-5375bf1f240b
📒 Files selected for processing (1)
tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml
- Require Holmes to discover BUILD-7X4Q9M (Jenkins 2.426.3-lts) from the change history index, proving it actually queried ES. This CI/CD change is unrelated to the alert and should be noted as irrelevant. - Require explicit zero-hit evidence: Holmes must state that searching for lambda-dashboard-monitoring/clientorg-alpha returned no results. - Keep existing cluster-mismatch and no-diagnosis checks intact. https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg Signed-off-by: Claude <noreply@anthropic.com>
…mismatch-test-Bwa0K
234 is now taken by mcp_additional_properties from master. Renamed to 235, updated index names (app-234→app-235), and made after_test reentrant by skipping index deletion (parallel CI runs share the same ES cluster; before_test handles cleanup idempotently). https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg Signed-off-by: Claude <noreply@anthropic.com>
…-Bwa0K' into claude/tdd-cluster-mismatch-test-Bwa0K
There was a problem hiding this comment.
🧹 Nitpick comments (1)
tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/test_case.yaml (1)
125-125: Minor portability note:sed -isyntax differs on macOS.The
sed -i 's/^ //' "$BULK_FILE"command works on Linux but fails on macOS, which requiressed -i ''. Since CI likely runs Linux, this is a low-priority concern, but for local development consistency you could use:sed 's/^ //' "$BULK_FILE" > "${BULK_FILE}.tmp" && mv "${BULK_FILE}.tmp" "$BULK_FILE"Or ensure the heredoc content has no leading whitespace to begin with.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/test_case.yaml` at line 125, Replace the non-portable in-place sed invocation `sed -i 's/^ //' "$BULK_FILE"` in the test case with a portable approach: either write the sanitized output to a temp file and atomically move it back (i.e., run sed without -i and mv the tmp file over "$BULK_FILE") or remove the leading indentation upstream in the heredoc that generates BULK_FILE; update the line containing `sed -i 's/^ //' "$BULK_FILE"` accordingly so macOS and Linux both work.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Nitpick comments:
In
`@tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/test_case.yaml`:
- Line 125: Replace the non-portable in-place sed invocation `sed -i 's/^ //'
"$BULK_FILE"` in the test case with a portable approach: either write the
sanitized output to a temp file and atomically move it back (i.e., run sed
without -i and mv the tmp file over "$BULK_FILE") or remove the leading
indentation upstream in the heredoc that generates BULK_FILE; update the line
containing `sed -i 's/^ //' "$BULK_FILE"` accordingly so macOS and Linux both
work.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 23bd6d48-3a27-45c5-9ad1-60fdc2b26ae5
📒 Files selected for processing (2)
tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/test_case.yamltests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/toolsets.yaml
✅ Files skipped from review due to trivial changes (1)
- tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/toolsets.yaml
Replace hardcoded dates with shell-computed relative timestamps (e.g. 6 hours ago, 2 hours ago) so ES test data always looks fresh regardless of when the eval runs. https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg Signed-off-by: Claude <noreply@anthropic.com>
The cluster mismatch eval is ES-only and doesn't need kubectl/helm. Replace the symlink to shared/elasticsearch_toolsets.yaml with a standalone file that also disables kubernetes/core, kubernetes/logs, and helm/core to prevent ToolsetPrerequisiteError in sandbox envs. https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg Signed-off-by: Claude <noreply@anthropic.com>
…ernatives Replace `date -u -d` (GNU-only) with a Python helper function and `sed -i` with portable sed+mv pattern in the test 235 before_test script. https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg Signed-off-by: Claude <noreply@anthropic.com>
The verbose expected_output entries caused the evaluator LLM to exceed output token limits, producing truncated JSON that failed to parse. https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg Signed-off-by: Claude <noreply@anthropic.com>
- Disable robusta and bash toolsets so Holmes must use elasticsearch - Update prompt to reference ES index names (app-235-cluster-inventory, app-235-change-history) so Holmes knows where to search - Holmes now queries ES, finds CI/CD data but nothing for the alert's app, yet still provides speculative diagnosis instead of flagging the cluster mismatch — correct TDD red failure https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg Signed-off-by: Claude <noreply@anthropic.com>
Summary
Add a comprehensive test case that validates Holmes' ability to detect when it's investigating alerts from a different cluster than the one it's connected to, and ensures it refuses to provide speculative root cause analysis in such scenarios.
Key Changes
New test case:
234_elasticsearch_cluster_mismatch/test_case.yamlci-cdcluster but alert originates fromus-prod-ekslambda-dashboard-monitoring) doesn't exist in the current clusterTest infrastructure:
234_elasticsearch_cluster_mismatch/toolsets.yamlImplementation Details
us-prod-eksandci-cdclusterselasticsearch,mediumcomplexity, andtransparencycategoryhttps://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg
Summary by CodeRabbit