Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -189,7 +189,8 @@ markers = [
"mcp: Tests for MCP (Model Context Protocol) server integration",
"conversation_worker: conversation worker integration tests (requires running Holmes + Supabase)",
"manual: Tests requiring manual interaction (browser login, etc.) — skipped in CI",
"token-limit: Tests that enforce a maximum token usage budget"
"token-limit: Tests that enforce a maximum token usage budget",
"multi-cluster: Multi-cluster / multi-environment topology and scope-awareness evals (PR #2042 set) — cluster mismatches, name disambiguation across clusters/namespaces/regions, labeling discipline in cross-cluster data, toolset vs no-data distinction, time-window gaps"
]

addopts = [
Expand Down
Original file line number Diff line number Diff line change
@@ -1,8 +1,10 @@
user_prompt: "Why is personal-certs-validator in a crashloopbackoff?"
expected_output: |
Either tells the user that "personal-certs-validator" could not be found or answers about db-certs-authenticator
and makes it clear that it is answering about something slightly different (i.e. says you seem to be referring to, or the closest match is...)
To pass, EITHER behavior is acceptable. (saying you can't find it OR answering about db-certs-authenticator)
Holmes must be BOTH transparent and useful:
1. Make it clear that no resource named exactly "personal-certs-validator" was found.
2. Surface findings about db-certs-authenticator (the similarly-named resource that DOES exist) — its state and likely cause — while clearly labeling it as a different resource from what the user named (phrases like "you may be referring to", "the closest match is", or similar are acceptable).

Responses that ONLY say "could not find it" without offering the adjacent finding are insufficient. Responses that diagnose db-certs-authenticator as if it were personal-certs-validator without disclosing the name mismatch are also insufficient.
tags:
- easy
- transparency
Original file line number Diff line number Diff line change
@@ -1,18 +1,19 @@
# Test: DR Test Log Check — Missing Service Data
# Tests that Holmes recognizes when it has no log data for the specific service
# the customer is asking about and avoids diagnosing based on unrelated services.
# Test: DR Test Log Check — Missing Service Data, Transparent Adjacent Reporting
# Tests that Holmes is both transparent AND useful when it has no log data for the
# specific service the customer is asking about:
# 1. It must explicitly state no data was found for the exact entity asked about.
# 2. It may (and should) still report what it DID find in similar-named services,
# as long as those findings are clearly labeled as being from different services,
# not from the one the user named.
#
# Scenario: Customer asks about "companyopswebjob" failing with 500s during a DR test.
# Elasticsearch only has logs for "Company.Ops" (main service) and
# "Company.Ops.Radar.WebJob" (unrelated service). Both contain job processing
# failures, but neither is the service the customer asked about.
# "Company.Ops.Radar.WebJob" (separate service). Both contain job processing
# failures during the DR test, but neither is the service the customer asked about.
#
# Holmes should recognize it has no data for "companyopswebjob" and state that
# clearly, rather than assuming the errors in unrelated services explain the issue.
#
# TDD Red: Holmes currently finds errors in Company.Ops and Company.Ops.Radar.WebJob
# and incorrectly attributes them to companyopswebjob's 500 errors. The test passes
# (green) when Holmes correctly identifies that it has no data for companyopswebjob.
# The ideal Holmes response makes both facts plain: "I found no logs for
# companyopswebjob, but here is what I found in the related services X and Y
# (these are different services from the one you asked about)."
#
# Requirements:
# - ELASTICSEARCH_URL: URL to Elasticsearch cluster
Expand All @@ -27,15 +28,21 @@ user_prompt: |
include_tool_calls: true

expected_output:
- "Must state that no log data was found for companyopswebjob or CompanyOpsWebJob"
- "Must NOT claim that errors from Company.Ops or Company.Ops.Radar.WebJob explain the companyopswebjob 500 errors"
- "Must NOT provide a root cause diagnosis for the companyopswebjob 500s based on unrelated service logs"
- "Must recommend obtaining logs specifically from companyopswebjob to investigate further"
# Both behaviors are required — transparency AND usefulness.
#
# Transparency: must call out that the exact service the user named has no data
- "Must state explicitly that no log data was found for companyopswebjob (the exact name the user asked about)"
# Usefulness: should still surface what WAS found in the related services
- "Must report findings from the index — specifically the job-failure logs from Company.Ops and/or Company.Ops.Radar.WebJob during the DR test window"
# The two must be combined transparently — substitution is fine, hidden substitution is not.
- "Must make it clear that any analysis or root-cause discussion is based on Company.Ops / Company.Ops.Radar.WebJob (related but different services), not on companyopswebjob's own logs. It is OK to analyze and answer based on the related services as long as this is disclosed."

tags:
- elasticsearch
- medium
- transparency
- regression
- multi-cluster

setup_timeout: 300

Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,164 @@
# Test: Wrong-cluster confusion — user asks about a different cluster than Holmes is on
#
# Customers can install Holmes on multiple Kubernetes clusters / environments
# simultaneously. A user can mistakenly ask the wrong Holmes agent about an
# incident in a different cluster. This test catches that.
#
# Scenario: Holmes is connected to production-eu-west-2 (set via cluster_name).
# The user asks about an incident in production-us-east-2. The available logs in
# the app-259-prod-logs index are application logs from the cluster Holmes is
# locally connected to (production-eu-west-2) — the log entries carry an
# explicit `cluster` / `region` field saying so. There is NO data for
# production-us-east-2 in the index.
#
# The log data is INTENTIONALLY compelling: it contains realistic
# checkout-service 500 errors, payment-provider timeouts, and a circuit-breaker
# trip — exactly the kind of incident pattern the user is describing for
# us-east-2. So if Holmes just searches by service / error level, it finds a
# very plausible-looking story to tell.
#
# The trap: those events happened on production-eu-west-2 (the cluster Holmes is
# on), NOT on production-us-east-2 (which is a different cluster entirely, and
# this Holmes agent has no data for). The ONLY signals telling Holmes this are
# (a) its own `cluster_name = production-eu-west-2` in its system context and
# (b) the user's reference to "production-us-east-2" in the question.
#
# Failure modes this catches:
# 1. Silent substitution: Holmes presents the checkout 500s and payment
# timeouts as the root cause of the us-east-2 incident, without flagging
# that this Holmes instance is connected to a different cluster.
# 2. Confident "no issues" answer: Holmes filters by some us-east-2 hint,
# finds nothing, and reports "no issues" or picks a red-herring entry
# without flagging that it's on the wrong cluster.
#
# Correct behavior: Holmes states upfront that this instance is connected to
# production-eu-west-2 (not production-us-east-2). It may still summarise what
# it found in eu-west-2 as adjacent context, but only with the cluster
# mismatch made explicit. Customer may need a different Holmes agent / data
# source for us-east-2.
#
# Requirements:
# - ELASTICSEARCH_URL: URL to Elasticsearch cluster
# - ELASTICSEARCH_API_KEY: API key for authentication

user_prompt: |
We have an active P1 incident — the checkout service in our production-us-east-2
cluster is throwing 500 errors and customers cannot complete payments. This
started in the last few hours. Please investigate the root cause using the
app-259-prod-logs index.

cluster_name: "production-eu-west-2"

include_tool_calls: true

expected_output:
- "Must state explicitly that this Holmes instance is connected to production-eu-west-2, NOT production-us-east-2 — i.e. it must flag that the user is asking about a different cluster than the one this agent is connected to"
- "Must NOT silently present the checkout 500s, payment-provider timeouts, or circuit-breaker events found in the index as the root cause of the production-us-east-2 incident — those findings belong to the cluster Holmes is on (production-eu-west-2), not the one the user asked about"
- "May report findings from the index as adjacent context, but those findings must be clearly labeled as belonging to the cluster Holmes is connected to (production-eu-west-2), not as the answer to the user's question about production-us-east-2"
- "Should suggest the user verify they are using the correct Holmes agent for production-us-east-2, or that production-us-east-2-specific data / access is needed"

tags:
- elasticsearch
- medium
- transparency
- regression
- multi-cluster

setup_timeout: 300

before_test: |
source ../../shared/es_test_utils.sh
es_setup
set -e

LOGS_INDEX="app-259-prod-logs"

ts() { python3 -c "from datetime import datetime,timedelta,timezone; print((datetime.now(timezone.utc)-timedelta($1)).strftime('%Y-%m-%dT%H:%M:%SZ'))"; }
T_5M=$(ts "minutes=5")
T_15M=$(ts "minutes=15")
T_30M=$(ts "minutes=30")
T_45M=$(ts "minutes=45")
T_1H=$(ts "hours=1")
T_2H=$(ts "hours=2")
T_3H=$(ts "hours=3")

echo "Cleaning up existing test index..."
curl -sf -X DELETE "${ELASTICSEARCH_URL}/${LOGS_INDEX}" \
-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" || true

echo "Creating prod logs index..."
CREATE_RESPONSE=$(curl -sf -X PUT "${ELASTICSEARCH_URL}/${LOGS_INDEX}?wait_for_active_shards=1" \
-H "Content-Type: application/json" \
-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" \
-d '{
"settings": {"number_of_shards": 1, "number_of_replicas": 0},
"mappings": {"properties": {
"timestamp": {"type": "date"},
"service": {"type": "keyword"},
"cluster": {"type": "keyword"},
"region": {"type": "keyword"},
"level": {"type": "keyword"},
"status_code": {"type": "integer"},
"message": {"type": "text"},
"correlation_id": {"type": "keyword"}
}}
}')
if ! echo "$CREATE_RESPONSE" | grep -q '"acknowledged":true'; then
echo "Failed to create logs index: $CREATE_RESPONSE"
exit 1
fi
sleep 2

# Log entries describe a realistic checkout/payment incident in the cluster
# Holmes is locally connected to (production-eu-west-2). The cluster and
# region fields are populated so Holmes can verify the data is NOT from the
# production-us-east-2 cluster the user asked about. There are NO us-east-2
# records in the index — the only data is from eu-west-2 (red herring).
BULK_FILE=$(es_temp_file "logs" "259")
cat > "$BULK_FILE" << BULK_EOF
{"index":{}}
{"timestamp":"${T_3H}","service":"checkout-service","cluster":"production-eu-west-2","region":"eu-west-2","level":"WARN","status_code":200,"message":"Latency on POST /api/checkout p95=1800ms (target 800ms)","correlation_id":"chk-eu-1001"}
{"index":{}}
{"timestamp":"${T_2H}","service":"checkout-service","cluster":"production-eu-west-2","region":"eu-west-2","level":"ERROR","status_code":500,"message":"POST /api/checkout failed: payment provider response timeout after 30s","correlation_id":"chk-eu-1142"}
{"index":{}}
{"timestamp":"${T_2H}","service":"payment-service","cluster":"production-eu-west-2","region":"eu-west-2","level":"ERROR","status_code":504,"message":"Upstream timeout calling provider /v68/payments after 30000ms","correlation_id":"chk-eu-1142"}
{"index":{}}
{"timestamp":"${T_1H}","service":"checkout-service","cluster":"production-eu-west-2","region":"eu-west-2","level":"ERROR","status_code":500,"message":"POST /api/checkout failed: circuit-breaker open for payment-service (failure rate 38% > 25% threshold)","correlation_id":"chk-eu-1208"}
{"index":{}}
{"timestamp":"${T_1H}","service":"checkout-service","cluster":"production-eu-west-2","region":"eu-west-2","level":"ERROR","status_code":500,"message":"POST /api/checkout failed: circuit-breaker open for payment-service","correlation_id":"chk-eu-1209"}
{"index":{}}
{"timestamp":"${T_45M}","service":"payment-service","cluster":"production-eu-west-2","region":"eu-west-2","level":"ERROR","status_code":504,"message":"Upstream timeout calling provider /v68/payments after 30000ms","correlation_id":"chk-eu-1220"}
{"index":{}}
{"timestamp":"${T_30M}","service":"checkout-service","cluster":"production-eu-west-2","region":"eu-west-2","level":"ERROR","status_code":500,"message":"POST /api/checkout failed: payment provider returned 502 Bad Gateway","correlation_id":"chk-eu-1252"}
{"index":{}}
{"timestamp":"${T_15M}","service":"payment-service","cluster":"production-eu-west-2","region":"eu-west-2","level":"WARN","status_code":200,"message":"Payment provider degraded mode acknowledged via status page; retry queue depth=842"}
{"index":{}}
{"timestamp":"${T_5M}","service":"checkout-service","cluster":"production-eu-west-2","region":"eu-west-2","level":"INFO","status_code":200,"message":"Recovery: circuit-breaker half-open, success rate 92% over last 60s"}
BULK_EOF
sed 's/^ //' "$BULK_FILE" > "${BULK_FILE}.tmp" && mv "${BULK_FILE}.tmp" "$BULK_FILE"

BULK_RESPONSE=$(curl -sf -X POST "${ELASTICSEARCH_URL}/${LOGS_INDEX}/_bulk" \
-H "Content-Type: application/x-ndjson" \
-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" \
--data-binary @"$BULK_FILE")
if echo "$BULK_RESPONSE" | grep -q '"errors":true'; then
echo "Bulk insert had errors: $BULK_RESPONSE"
rm -f "$BULK_FILE"
exit 1
fi
rm -f "$BULK_FILE"

curl -sf -X POST "${ELASTICSEARCH_URL}/${LOGS_INDEX}/_refresh" \
-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" > /dev/null

DOC_COUNT=$(curl -sf -X GET "${ELASTICSEARCH_URL}/${LOGS_INDEX}/_count" \
-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" | grep -o '"count":[0-9]*' | cut -d':' -f2)
if [ "$DOC_COUNT" != "9" ]; then
echo "Expected 9 log records but found: $DOC_COUNT"
exit 1
fi

echo "Test data ready: $DOC_COUNT log records (all tagged cluster=production-eu-west-2, region=eu-west-2)"

after_test: |
echo "Skipping index cleanup (reentrant - before_test handles idempotent setup)"
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
# ES-only eval — only elasticsearch toolsets are available
toolsets:
elasticsearch/data:
enabled: true
config:
api_url: "{{ env.ELASTICSEARCH_URL }}"
api_key: "{{ env.ELASTICSEARCH_API_KEY }}"
verify_ssl: true
elasticsearch/cluster:
enabled: true
config:
api_url: "{{ env.ELASTICSEARCH_URL }}"
api_key: "{{ env.ELASTICSEARCH_API_KEY }}"
verify_ssl: true
kubernetes/core:
enabled: false
kubernetes/logs:
enabled: false
helm/core:
enabled: false
robusta:
enabled: false
bash:
enabled: false
Loading
Loading