Skip to content

Add cluster mismatch detection test for Holmes - #1801

Merged
naomi-robusta merged 13 commits into
masterfrom
claude/tdd-cluster-mismatch-test-Bwa0K
Mar 19, 2026
Merged

naomi-robusta merged 13 commits into
masterfrom
claude/tdd-cluster-mismatch-test-Bwa0K

Conversation

@naomi-robusta

@naomi-robusta naomi-robusta commented Mar 17, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Add a comprehensive test case that validates Holmes' ability to detect when it's investigating alerts from a different cluster than the one it's connected to, and ensures it refuses to provide speculative root cause analysis in such scenarios.

Key Changes

  • New test case: 234_elasticsearch_cluster_mismatch/test_case.yaml

    • Tests scenario where user is connected to ci-cd cluster but alert originates from us-prod-eks
    • Validates that Holmes identifies the app (lambda-dashboard-monitoring) doesn't exist in the current cluster
    • Ensures Holmes recognizes the cluster mismatch and declines to provide diagnosis
    • Includes anti-hallucination validation: Holmes must discover specific deployment ID (DEP-K8M9X2) and version (v2.4.7-rc3) by querying Elasticsearch, proving actual data search
    • Marked as TDD test (expected to fail initially, pass when feature is implemented)
  • Test infrastructure: 234_elasticsearch_cluster_mismatch/toolsets.yaml

    • References shared Elasticsearch toolsets for test execution

Implementation Details

  • Sets up two Elasticsearch indices:
    • Cluster Inventory: Contains 10 deployment records across us-prod-eks and ci-cd clusters
    • Change History: Contains 6 change records with recent deployment (DEP-K8M9X2 v2.4.7-rc3) in us-prod-eks
  • Includes comprehensive validation checks to ensure test data integrity before execution
  • Properly cleans up test indices after execution
  • Tagged as elasticsearch, medium complexity, and transparency category

https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg

Summary by CodeRabbit

  • Tests
    • Added a new end-to-end test fixture covering cluster-mismatch detection.
    • Validates that alerts from an inaccessible cluster yield no false results, that irrelevant change history is identified, and that the system refuses to provide root-cause when the correct cluster is unreachable.
    • Confirms user-facing guidance recommends connecting to the missing cluster.

Tests that Holmes correctly identifies when it's connected to the wrong
cluster (ci-cd) while investigating an alert from a different cluster
(us-prod-eks). Uses Elasticsearch to simulate a deployment registry and
change history. The test expects Holmes to refuse providing a root cause
when the target app doesn't exist in the current cluster.

This is a TDD "red" test - it should initially fail because Holmes
currently provides analysis even when it detects a cluster mismatch.

https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg
Signed-off-by: Claude <noreply@anthropic.com>
@claude

claude Bot commented Mar 17, 2026

Copy link
Copy Markdown
Contributor

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review.

@coderabbitai

coderabbitai Bot commented Mar 17, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Adds a new Elasticsearch-backed test fixture and a toolsets reference to validate detection of alerts that reference a different cluster than the connected Elasticsearch data (cluster mismatch), including ES index setup, CI/CD-only test data, and expected response assertions.

Changes

Cohort / File(s) Summary
Elasticsearch cluster-mismatch fixture
tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/test_case.yaml
New test YAML that: initializes Elasticsearch, deletes/creates indices, inserts inventory and change-history documents scoped to ci-cd cluster (with a distractor us-prod-eks namespace doc), refreshes indices, asserts preconditions (inventory counts, zero records for the alert app), enables tool calls, and declares expected assistant responses about cluster mismatch and inability to root-cause.
Toolsets ref
tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/toolsets.yaml
New file pointing to shared Elasticsearch toolset (../../shared/elasticsearch_toolsets.yaml).

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

Suggested reviewers

  • aantn
  • arikalon1
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the primary change: adding a new cluster mismatch detection test for Holmes. It is clear, concise, and directly reflects the main objective of the PR.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Mar 17, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 062e55c
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69bbf163e44aa500088d5ed6
😎 Deploy Preview https://deploy-preview-1801--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Mar 17, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 73ebf76 (#23292352513)

✅ Results of HolmesGPT evals

Automatically triggered by commit 73ebf76 on branch claude/tdd-cluster-mismatch-test-Bwa0K (labels: evals-id-235)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 1/1 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 235_elasticsearch_cluster_mismatch 80.9s 12 26 $0.5158 246,716 240,983 26,045 5,733 745 196,971 44,012 — —
Total 80.9s avg 12.0 avg 26.0 avg $0.5158 246,716 240,983 26,045 5,733 745 196,971 44,012 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 86bf58d (#23290503429)

⚠️ Eval Results (with failures)

Automatically triggered by commit 86bf58d on branch claude/tdd-cluster-mismatch-test-Bwa0K (labels: evals-id-235)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 0/1 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
❌ 235_elasticsearch_cluster_mismatch 124.4s 19 30 — — — — — — — — — —
Total 124.4s avg 19.0 avg 30.0 avg — — — — — — — — — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected

📜 Run @ ce2808e (#23290302902)

⚠️ Eval Results (with failures)

Automatically triggered by commit ce2808e on branch claude/tdd-cluster-mismatch-test-Bwa0K (labels: evals-id-235)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 0/1 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
❌ 235_elasticsearch_cluster_mismatch — — — — — — — — — — — — —
Total — avg — avg — avg — — — — — — — — — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected

📜 Run @ fc2a315 (#23289012717)

⚠️ Eval Results (with failures)

Automatically triggered by commit fc2a315 on branch claude/tdd-cluster-mismatch-test-Bwa0K (labels: evals-id-235)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 0/1 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
❌ 235_elasticsearch_cluster_mismatch 98.3s 14 26 — — — — — — — — — —
Total 98.3s avg 14.0 avg 26.0 avg — — — — — — — — — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 7 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected

📜 Run @ 02d1779 (#23288784512)

⚠️ Eval Results (with failures)

Automatically triggered by commit 02d1779 on branch claude/tdd-cluster-mismatch-test-Bwa0K (labels: evals-id-235)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 0/1 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
❌ 235_elasticsearch_cluster_mismatch 44.9s 7 14 $0.3217 173,939 170,900 27,509 3,039 843 142,923 27,977 — —
Total 44.9s avg 7.0 avg 14.0 avg $0.3217 173,939 170,900 27,509 3,039 843 142,923 27,977 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected


⚠️ Eval Results (with failures)

Automatically triggered by commit 062e55c on branch claude/tdd-cluster-mismatch-test-Bwa0K (labels: evals-id-235)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 0/1 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
❌ 235_elasticsearch_cluster_mismatch 117.5s 17 32 $0.5476 379,554 372,434 30,494 7,120 745 340,155 32,279 — —
Total 117.5s avg 17.0 avg 32.0 avg $0.5476 379,554 372,434 30,494 7,120 745 340,155 32,279 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tdd-cluster-mismatch-test-Bwa0K -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, integration, kafka, kubernetes, leaked-information, logs, loki, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tdd-cluster-mismatch-test-Bwa0K -f markers=regression -f filter=

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml (1)

71-74: Consider using -sS instead of -sf for initial cleanup.

The -f flag causes curl to fail silently on HTTP errors, but combined with || true, the error is already suppressed. Using -sS would still suppress progress output while showing errors (though they're intentionally ignored here via || true). This is a minor consistency point - the current approach works correctly.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In
`@tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml`
around lines 71 - 74, Replace the curl invocations that use the `-sf` flags with
`-sS` so errors are still reported (while keeping output silent) — update the
two lines invoking curl with `"${ELASTICSEARCH_URL}/${INVENTORY_INDEX}"` and
`"${ELASTICSEARCH_URL}/${CHANGES_INDEX}"` to use `-sS` instead of `-sf`, and
keep the existing `-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" || true`
behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In
`@tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml`:
- Around line 71-74: Replace the curl invocations that use the `-sf` flags with
`-sS` so errors are still reported (while keeping output silent) — update the
two lines invoking curl with `"${ELASTICSEARCH_URL}/${INVENTORY_INDEX}"` and
`"${ELASTICSEARCH_URL}/${CHANGES_INDEX}"` to use `-sS` instead of `-sf`, and
keep the existing `-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" || true`
behavior.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 08969301-b9f8-4624-a9fd-c4f5dc1f61a3

📥 Commits

Reviewing files that changed from the base of the PR and between 2319767 and a2aa374.

📒 Files selected for processing (2)
  • tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/toolsets.yaml

@github-actions

github-actions Bot commented Mar 17, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 6e8d851b (built in 1m 8s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:6e8d851b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:6e8d851b me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:6e8d851b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:6e8d851b
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:6e8d851b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:6e8d851b me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:6e8d851b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:6e8d851b

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:6e8d851b \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:6e8d851b

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:6e8d851b \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:6e8d851b

The eval must simulate the real scenario where Holmes can only see data
from the current cluster (ci-cd). Previously the ES data included
us-prod-eks deployment records, which let Holmes find the app and
provide analysis — the wrong failure mode.

Now the ES data contains only:
- ci-cd deployments (jenkins, argocd, gitlab-runner, flux, tekton)
- An empty us-prod-eks namespace (red herring)
- Change history with only CI/CD tool changes (no alert-related changes)

Holmes should search for lambda-dashboard-monitoring/clientorg-alpha,
find nothing, and conclude it's in the wrong cluster — not provide a
speculative root cause.

https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml (1)

68-70: Enable fail-fast shell mode before setup calls

set -e is applied after source and es_setup. If either fails, the script may continue in a partially initialized state.

Proposed diff
 before_test: |
+  set -euo pipefail
   source ../../shared/es_test_utils.sh
   es_setup
-  set -e
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In
`@tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml`
around lines 68 - 70, The script enables fail-fast mode too late—move the guard
so that `set -e` (or `set -euo pipefail` if desired) is invoked before sourcing
`../../shared/es_test_utils.sh` and before the `es_setup` call so any failure in
`source` or `es_setup` aborts immediately; ensure `set -e` is placed at the top
(right after the shebang) so `source ../../shared/es_test_utils.sh` and
`es_setup` cannot leave the script in a partially initialized state.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In
`@tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml`:
- Around line 53-59: Update the expected_output to require at least one concrete
seeded value and explicit zero-hit evidence: include the seeded identifier
BUILD-7X4Q9M (and its app/version context) as a required discovery and also
require an explicit statement that searching for the alerting app returned zero
hits (e.g., "no matching results for alerting app"). Modify the expected_output
array entries (the list under expected_output) so one item asserts the
discovered BUILD-7X4Q9M value and another asserts explicit zero-hit evidence for
the alerting app, while keeping the existing cluster-mismatch and
investigation-recommendation checks intact.
- Around line 72-74: The test uses fixed Elasticsearch index names
INVENTORY_INDEX and CHANGES_INDEX in the fixture and its before_test/after_test
hooks, which can collide across parallel runs; update the before_test to create
and the test to reference run-scoped index names (e.g., append a unique run ID
or timestamp to INVENTORY_INDEX and CHANGES_INDEX) and change after_test to only
delete those run-scoped indices (or skip deletion of shared names) so setup is
idempotent and teardown doesn’t remove shared resources; apply the same change
pattern to the other affected fixtures/tests (cases 227–232).

---

Nitpick comments:
In
`@tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml`:
- Around line 68-70: The script enables fail-fast mode too late—move the guard
so that `set -e` (or `set -euo pipefail` if desired) is invoked before sourcing
`../../shared/es_test_utils.sh` and before the `es_setup` call so any failure in
`source` or `es_setup` aborts immediately; ensure `set -e` is placed at the top
(right after the shebang) so `source ../../shared/es_test_utils.sh` and
`es_setup` cannot leave the script in a partially initialized state.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 4c491c3a-b02e-4888-b8db-5375bf1f240b

📥 Commits

Reviewing files that changed from the base of the PR and between a2aa374 and 398ffbe.

📒 Files selected for processing (1)
  • tests/llm/fixtures/test_ask_holmes/234_elasticsearch_cluster_mismatch/test_case.yaml

claude and others added 2 commits March 19, 2026 08:12
- Require Holmes to discover BUILD-7X4Q9M (Jenkins 2.426.3-lts) from the
  change history index, proving it actually queried ES. This CI/CD change
  is unrelated to the alert and should be noted as irrelevant.
- Require explicit zero-hit evidence: Holmes must state that searching for
  lambda-dashboard-monitoring/clientorg-alpha returned no results.
- Keep existing cluster-mismatch and no-diagnosis checks intact.

https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg
Signed-off-by: Claude <noreply@anthropic.com>
claude added 3 commits March 19, 2026 09:11
234 is now taken by mcp_additional_properties from master.
Renamed to 235, updated index names (app-234→app-235), and made
after_test reentrant by skipping index deletion (parallel CI runs
share the same ES cluster; before_test handles cleanup idempotently).

https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg
Signed-off-by: Claude <noreply@anthropic.com>
…-Bwa0K' into claude/tdd-cluster-mismatch-test-Bwa0K

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/test_case.yaml (1)

125-125: Minor portability note: sed -i syntax differs on macOS.

The sed -i 's/^ //' "$BULK_FILE" command works on Linux but fails on macOS, which requires sed -i ''. Since CI likely runs Linux, this is a low-priority concern, but for local development consistency you could use:

sed 's/^  //' "$BULK_FILE" > "${BULK_FILE}.tmp" && mv "${BULK_FILE}.tmp" "$BULK_FILE"

Or ensure the heredoc content has no leading whitespace to begin with.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In
`@tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/test_case.yaml`
at line 125, Replace the non-portable in-place sed invocation `sed -i 's/^  //'
"$BULK_FILE"` in the test case with a portable approach: either write the
sanitized output to a temp file and atomically move it back (i.e., run sed
without -i and mv the tmp file over "$BULK_FILE") or remove the leading
indentation upstream in the heredoc that generates BULK_FILE; update the line
containing `sed -i 's/^  //' "$BULK_FILE"` accordingly so macOS and Linux both
work.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In
`@tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/test_case.yaml`:
- Line 125: Replace the non-portable in-place sed invocation `sed -i 's/^  //'
"$BULK_FILE"` in the test case with a portable approach: either write the
sanitized output to a temp file and atomically move it back (i.e., run sed
without -i and mv the tmp file over "$BULK_FILE") or remove the leading
indentation upstream in the heredoc that generates BULK_FILE; update the line
containing `sed -i 's/^  //' "$BULK_FILE"` accordingly so macOS and Linux both
work.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 23bd6d48-3a27-45c5-9ad1-60fdc2b26ae5

📥 Commits

Reviewing files that changed from the base of the PR and between 7083ce1 and 8e670ed.

📒 Files selected for processing (2)
  • tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/toolsets.yaml
✅ Files skipped from review due to trivial changes (1)
  • tests/llm/fixtures/test_ask_holmes/235_elasticsearch_cluster_mismatch/toolsets.yaml

claude and others added 6 commits March 19, 2026 09:39
Replace hardcoded dates with shell-computed relative timestamps
(e.g. 6 hours ago, 2 hours ago) so ES test data always looks fresh
regardless of when the eval runs.

https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg
Signed-off-by: Claude <noreply@anthropic.com>
The cluster mismatch eval is ES-only and doesn't need kubectl/helm.
Replace the symlink to shared/elasticsearch_toolsets.yaml with a
standalone file that also disables kubernetes/core, kubernetes/logs,
and helm/core to prevent ToolsetPrerequisiteError in sandbox envs.

https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg
Signed-off-by: Claude <noreply@anthropic.com>
…ernatives

Replace `date -u -d` (GNU-only) with a Python helper function and
`sed -i` with portable sed+mv pattern in the test 235 before_test script.

https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg
Signed-off-by: Claude <noreply@anthropic.com>
The verbose expected_output entries caused the evaluator LLM to exceed
output token limits, producing truncated JSON that failed to parse.

https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg
Signed-off-by: Claude <noreply@anthropic.com>
- Disable robusta and bash toolsets so Holmes must use elasticsearch
- Update prompt to reference ES index names (app-235-cluster-inventory,
  app-235-change-history) so Holmes knows where to search
- Holmes now queries ES, finds CI/CD data but nothing for the alert's
  app, yet still provides speculative diagnosis instead of flagging the
  cluster mismatch — correct TDD red failure

https://claude.ai/code/session_01TM2EJsWGCSHCpfgNAg92Kg
Signed-off-by: Claude <noreply@anthropic.com>
@naomi-robusta
naomi-robusta merged commit 7ca1c20 into master Mar 19, 2026
18 of 19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants