Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .github/workflows/eval-regression.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -1212,6 +1212,12 @@ jobs:
DATADOG_API_KEY: ${{ secrets.DATADOG_API_KEY }}
DATADOG_APP_KEY: ${{ secrets.DATADOG_APP_KEY }}
BRAINTRUST_API_KEY: ${{ secrets.BRAINTRUST_API_KEY }}
# Needed by tests/llm/utils/braintrust_history.py to authenticate the
# GitHub Actions API lookup that finds recent master-* / ci-benchmark-*
# workflow runs. Without it the unauthenticated 60 req/h rate limit is
# exhausted within minutes on shared runners and the "vs master" /
# "vs benchmark" columns silently come back empty.
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
MODEL_LIST_FILE_LOCATION: /tmp/model_list.yaml
UPLOAD_DATASET: "true"
EXPERIMENT_ID: github-${{ github.run_id }}.${{ github.run_number }}.${{ github.run_attempt }}
Expand Down
6 changes: 6 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -305,6 +305,12 @@ raw = buf.getvalue() # Contains full ANSI escape sequences
- Known Rich 13.9.4 bug: `Live.refresh()` calls `console.print(Control())` with default `end="\\n"`, adding a trailing newline not counted in `LiveRender._shape`. When the terminal has room below the display, each frame leaks 1 ghost line. When the display is at the bottom (common case), the `\\n` causes scrolling and `height-1` cursor-ups is correct.
- Workaround: subclass `Live` and override `refresh()` to pass `end=""`. Do NOT patch `position_cursor` — that over-erases when the display is at the terminal bottom (the common case).

## Investigating Eval Regressions / Holmes Behavior Changes

**Understand the behavior from trace data before designing a fix.** Braintrust (or the local evals_report.md) holds the rendered prompts, per-LLM-call metrics, and tool-call sequences for each iter. Pull traces for one baseline run and one current run for the same `(test, model)` before reading any source diffs — file-level diffs (prompts, code, config) routinely mislead about what actually changed at runtime (e.g. a jinja2 template can grow in source lines and shrink in rendered output). Look at individual runs first; aggregate statistics over n iters can hide deterministic per-iter differences under variance.

When reverting or fixing a suspect PR, read its full diff (`git show --stat <commit>`) before deciding what to change — a PR often has multiple effects, and reverting only the file you noticed leaves the others still acting on the model.

## Security Notes

- All tools have read-only access by design
Expand Down
15 changes: 15 additions & 0 deletions holmes/plugins/prompts/generic_ask.jinja2
Original file line number Diff line number Diff line change
Expand Up @@ -270,6 +270,21 @@ If the answer to any of those questions is 'yes' - The investigation is INCOMPLE
* Treat error messages as exact diagnostic evidence. `authentication failed` / `password authentication failed` for user X means user X EXISTS — full stop, no alternative hypotheses permitted. `role does not exist` / `user not found` means the user is absent. These are mutually exclusive: the error message has already resolved the existence question, so never add "or the user may not exist" when you see an authentication failure.
* Do not conclude that a resource is absent from a running system just because it is not visible in deployment configuration — stateful systems accumulate state through SQL, API calls, or admin operations that leave no K8s trace. If you cannot read a value (e.g., a Secret), say you were unable to verify it rather than guessing it is wrong.
* ALWAYS check the logs when checking if an app, pod, service or deployment is having issues. Something "running" and reporting healthy does not mean it is without issues.

# If investigating Kubernetes problems

* run as many kubectl commands as you need to gather more information, then respond.
* if possible, do so repeatedly on different Kubernetes objects.
* for example, for deployments first run kubectl on the deployment then a replicaset inside it, then a pod inside that.
* when investigating a pod that crashed or application errors, always run kubectl_describe and fetch the logs
* Do check both the status of the kubernetes resources and the application runtime as well, by investigating logs
* do not give an answer like "The pod is pending" as that doesn't state why the pod is pending and how to fix it.
* do not give an answer like "Pod's node affinity/selector doesn't match any available nodes" because that doesn't include data on WHICH label doesn't match
* if investigating an issue on many pods, there is no need to check more than 3 individual pods in the same deployment. pick up to a representative 3 from each deployment if relevant
* if the user says something isn't working, ALWAYS:
** use kubectl_describe on the owner workload + individual pods and look for any transient issues they might have been referring to
** look for misconfigured ingresses/services etc
** check the application logs because there may be runtime issues
{% endif %}

{% if toolset_instructions_enabled %}
Expand Down
38 changes: 0 additions & 38 deletions holmes/plugins/skills/builtin/kubernetes-troubleshooting/SKILL.md

This file was deleted.

Original file line number Diff line number Diff line change
@@ -0,0 +1,125 @@
#!/usr/bin/env python3
"""Generate the same historical log timeline as 101_loki_historical_logs_pod_deleted
and push it directly to a local Loki instance (no Promtail, no Kubernetes).

Usage:
python generate_logs.py <loki_url>

Pushes logs to <loki_url>/loki/api/v1/push with the same labels Holmes expects.
"""

import json
import random
import sys
import urllib.error
import urllib.request
from datetime import datetime, timedelta

random.seed(100)

NAMESPACE = "app-259"
POD_NAME = "payment-api-259-d8f7b9c4-abc12"
SERVICE = "payment-api"
BATCH_SIZE = 200


def push(loki_url, streams_by_level):
"""Push grouped streams to Loki."""
Comment thread
aantn marked this conversation as resolved.
streams = []
for level, values in streams_by_level.items():
if not values:
continue
streams.append(
{
"stream": {
"job": "payment-api",
"namespace": NAMESPACE,
"pod_name": POD_NAME,
"service": SERVICE,
"level": level,
},
"values": values,
}
)
if not streams:
return
payload = {"streams": streams}
req = urllib.request.Request(
loki_url + "/loki/api/v1/push",
data=json.dumps(payload).encode("utf-8"),
headers={"Content-Type": "application/json"},
)
try:
urllib.request.urlopen(req, timeout=30)
except urllib.error.HTTPError as e:
body = e.read().decode("utf-8", errors="replace")[:500]
raise SystemExit(f"Loki push failed {e.code}: {body}")


def log_entry(level, message, **extra):
return {
"level": level,
"message": message,
"service": SERVICE,
"pod": POD_NAME,
**extra,
}


def generate(loki_url):
"""Generate the same logs as 101 (and 100a's incident period), push in batches."""
problem_start = datetime(2025, 8, 2, 13, 45, 0)
problem_end = datetime(2025, 8, 2, 14, 45, 0)

current = datetime(2025, 8, 1, 12, 0, 0)
scenario_end = datetime(2025, 8, 4, 14, 0, 0)
Comment thread
aantn marked this conversation as resolved.

streams_by_level: dict[str, list[list[str]]] = {}
total = 0

def add(ts: datetime, level: str, data: dict):
ts_nano = str(int(ts.timestamp() * 1e9))
streams_by_level.setdefault(level, []).append([ts_nano, json.dumps(data)])

while current < scenario_end:
if random.random() < 0.05:
add(
current,
"INFO",
log_entry(
"INFO",
"Payment processed successfully",
payment_id=f"PAY-{random.randint(1000, 9999)}",
),
)
total += 1

if problem_start <= current <= problem_end:
if random.random() < 0.4:
add(
current,
"ERROR",
log_entry(
"ERROR",
"Failed to acquire database connection - pool exhausted",
wait_time_ms=random.randint(1000, 5000),
queue_length=random.randint(5, 15),
),
)
total += 1

current += timedelta(minutes=random.randint(1, 5))

# Flush in batches
pending = sum(len(v) for v in streams_by_level.values())
if pending >= BATCH_SIZE:
push(loki_url, streams_by_level)
streams_by_level = {}

push(loki_url, streams_by_level)
print(f"Pushed {total} historical log entries to {loki_url}")


if __name__ == "__main__":
url = sys.argv[1] if len(sys.argv) > 1 else "http://localhost:3100"
generate(url.rstrip("/"))
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
auth_enabled: false
server:
http_listen_port: 3100
ingester:
query_store_max_look_back_period: -1
max_chunk_age: 17520h
wal:
enabled: false
lifecycler:
address: 127.0.0.1
ring:
kvstore:
store: inmemory
replication_factor: 1
final_sleep: 0s
chunk_store_config:
max_look_back_period: 17520h
schema_config:
configs:
- from: 2020-01-01
store: boltdb
object_store: filesystem
schema: v11
index:
prefix: index_
period: 168h
storage_config:
boltdb:
directory: /tmp/loki/index
filesystem:
directory: /tmp/loki/chunks
limits_config:
enforce_metric_name: false
reject_old_samples: false
reject_old_samples_max_age: 104w
max_entries_limit_per_query: 5000
max_query_lookback: 17520h
Original file line number Diff line number Diff line change
@@ -0,0 +1,77 @@
user_prompt: "The payment-api pod in namespace app-259 had issues on August 2, 2025 around 13:45 UTC. What happened?"

description: |
Docker-only variant of 101_loki_historical_logs_pod_deleted.

Stands up a Loki container locally (no Kubernetes), pushes a historical
log timeline directly to Loki's HTTP API, and lets Holmes investigate
using only the grafana/loki toolset. The "pod was deleted" framing is preserved
by simply never running the pod — Holmes only has Loki to work with.

Use this variant to reproduce the regression on a developer laptop, or in any
environment where Kubernetes is unavailable (e.g. cgroup-v1 sandboxes).

expected_output:
- Between 2025-08-02 13:45 UTC and 14:45 UTC, the payment-api pod logged the ERROR "Failed to acquire database connection - pool exhausted"
- The root cause is database connection pool exhaustion

include_tool_calls: true

tags:
- logs
- easy
- loki
- fast

setup_timeout: 180

before_test: |
set -e
CONTAINER=holmes-eval-loki-259
PORT=3259

# Clean up any leftover container from a previous run
docker rm -f "$CONTAINER" >/dev/null 2>&1 || true

# Start Loki with config that allows ingestion of old samples
docker run -d --name "$CONTAINER" -p "$PORT:3100" \
-v "$(pwd)/loki-config.yaml:/etc/loki/local-config.yaml:ro" \
grafana/loki:2.9.0 -config.file=/etc/loki/local-config.yaml >/dev/null

# Wait for Loki to be ready
READY=false
for i in $(seq 1 60); do
if curl -sf "http://localhost:$PORT/ready" 2>/dev/null | grep -q ready; then
READY=true
break
fi
sleep 1
done
if [ "$READY" != "true" ]; then
echo "ERROR: Loki not ready after 60s"
docker logs "$CONTAINER" | tail -40
exit 1
fi

# Push the historical log timeline directly via Loki's push API
python3 generate_logs.py "http://localhost:$PORT"

# Verify the incident logs landed in Loki
for i in $(seq 1 30); do
ERROR_COUNT=$(curl -sG "http://localhost:$PORT/loki/api/v1/query_range" \
--data-urlencode 'query={namespace="app-259",level="ERROR"}' \
--data-urlencode 'start=2025-08-02T13:00:00Z' \
--data-urlencode 'end=2025-08-02T15:00:00Z' \
--data-urlencode 'limit=1' 2>/dev/null | grep -o '"values"' | wc -l)
if [ "$ERROR_COUNT" -gt "0" ]; then
echo "Historical ERROR logs ready in Loki"
exit 0
fi
sleep 1
done

echo "ERROR: no historical ERROR logs found in Loki after push"
exit 1

after_test: |
docker rm -f holmes-eval-loki-259 >/dev/null 2>&1 || true
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
toolsets:
# Disable everything Kubernetes/Helm-related — this test runs against a plain
# Docker Loki container with no cluster to query.
kubernetes/logs:
enabled: false
kubernetes/core:
enabled: false
helm/core:
enabled: false
connectivity_check:
enabled: false
grafana/loki:
enabled: true
config:
api_url: http://localhost:3259
api_key: ""
labels:
pod: pod_name
namespace: namespace
8 changes: 7 additions & 1 deletion tests/llm/utils/braintrust_history.py
Original file line number Diff line number Diff line change
Expand Up @@ -331,7 +331,13 @@ def _extract_metrics(span: Dict[str, Any]) -> Optional[BenchmarkMetrics]:
scores = span.get("scores") or {}
metrics = span.get("metrics") or {}

test_id = metadata.get("eval_id") or metadata.get("test_id", "")
# Prefer test_id over eval_id: for parameterized tests (e.g.
# "227_count_configmaps_per_namespace[0]") eval_id strips the [N] suffix
# while test_id keeps it. The report joins baseline rows on test_case_name
# which always carries the parameterization, so eval_id-keyed rows would
# silently fail to match. Fall back to eval_id for older spans that only
# set the latter.
test_id = metadata.get("test_id") or metadata.get("eval_id", "")
model = metadata.get("model", "")

if not test_id or not model:
Expand Down
11 changes: 11 additions & 0 deletions tests/llm/utils/classifiers.py
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,17 @@ def get_classifier_model_params() -> ClassifierModelParams:
client_api_key = llm.api_key
client_base_url = llm.api_base
client_api_version = llm.api_version

# The classifier talks to OpenAI/OpenRouter directly via the openai SDK
# (autoevals doesn't go through litellm), so litellm-style provider
# prefixes in the model name must be stripped before we send the call.
if model_for_api and model_for_api.startswith("openrouter/"):
model_for_api = model_for_api.split("/", 1)[1]
# Fall back to OpenRouter's public endpoint if model_list didn't pin one.
if not client_base_url:
client_base_url = (
OPENROUTER_API_BASE or "https://openrouter.ai/api/v1"
)
else:
if not OPENAI_API_KEY and not AZURE_API_KEY and not OPENROUTER_API_KEY:
raise ValueError(
Expand Down
Loading
Loading