Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
48 commits
Select commit Hold shift + click to select a range
cec2801
Add tool_suggestions matrix to eval runs
claude May 1, 2026
823c0b2
Drop spurious f-string prefixes in tool_suggestions_config
claude May 3, 2026
bb2bbf2
Reframe suggest_runbooks toward access patterns (Hermes-style)
claude May 14, 2026
7b39a04
Narrow suggest_runbooks to tool-call corrections only
claude May 14, 2026
c431218
Add memories_generated check + two tool-call-correction evals
claude May 14, 2026
fab38bc
Merge branch 'master' into claude/add-tool-suggestions-matrix-av1Uv
aantn May 14, 2026
ef3d97d
Fix report status when post-judge assertions fail
claude May 14, 2026
5334d6d
Hoist json import to module level in property_manager
claude May 14, 2026
d41ab0c
Replace weak 259/260 evals with three forced-correction evals
claude May 14, 2026
4bc00ca
Enable elasticsearch toolset for 261/262 fixtures
claude May 15, 2026
c8b975b
Enrich suggest_runbooks prompt + add Kafka metric-name eval
claude May 15, 2026
b8e3913
Bias 261 prompt toward the default level= call shape
claude May 15, 2026
d20627a
Add explicit workflow checklist to suggest_runbooks prompt
claude May 15, 2026
4070836
Require suggest_runbooks emission to be parallel with final answer
claude May 15, 2026
0f0d8ec
Make suggest_runbooks emission a pre-answer step, not parallel
claude May 15, 2026
72b470b
Closed-loop memory replay: rerun_with_memory flag + report row
claude May 15, 2026
c259fc3
Fix 262 setup: switch exporter to python http.server + safe verify
claude May 15, 2026
91aea68
Show full stats on the replay row in the GitHub eval report
claude May 19, 2026
388d300
Replay row: red status when skill was not loaded
claude May 19, 2026
e15dad8
Re-trigger CI evals (no functional change)
claude May 19, 2026
8291b26
Merge branch 'origin/master' into claude/add-tool-suggestions-matrix-…
claude Jun 3, 2026
0e1999b
Separate primary-row status from replay-row status
claude Jun 3, 2026
a45ccc1
Add 264 ES numeric severity eval; lead skill description with when_to…
claude Jun 3, 2026
250136a
Bias 263 prompt + add 265 ES timestamp-keyword eval
claude Jun 3, 2026
214ba38
Soften 265 prompt: drop the explicit @timestamp instruction
claude Jun 3, 2026
360b1f2
Strengthen 261/264/265 prompt bias to force the wrong-call first
claude Jun 3, 2026
bf5797a
Fix SUGGEST_RUNBOOKS prompt — don't treat field-name discovery as gen…
claude Jun 3, 2026
888dc16
SKILL.md template: lead with "What to do" directive, then working call
claude Jun 3, 2026
0848dbd
264: clarify expected_output to match agent phrasing
claude Jun 3, 2026
89ea7c8
Revert SKILL.md "What to do" directive — it backfired in CI
claude Jun 3, 2026
5a7d67b
Tag the closed-loop skills baseline evals with 'skills'
claude Jun 3, 2026
4e9bbd9
collapse tool_suggestions matrix to always-on
claude Jun 4, 2026
6015a1e
add skills net-win summary, un-bias prompts, bad-memory regression eval
claude Jun 4, 2026
245f25b
fix 267 ES toolset; revert 261/263 prompt softening
claude Jun 4, 2026
356689a
add replay_user_prompt to decouple emission from skill-fetch on replay
claude Jun 4, 2026
4d27a91
revert replay_user_prompt on 265 — it regressed correctness
claude Jun 4, 2026
c5c87c9
consolidate captured skills by data-source domain (one skill per sour…
claude Jun 5, 2026
e61477d
268: make date references current-day relative so 'from today' works
claude Jun 5, 2026
52b74a7
268: disambiguate severity_num convention via msg-tag hints
claude Jun 5, 2026
ca1d9c9
update-existing skill path: cross-investigation accumulation
claude Jun 5, 2026
fe5e3a1
269: split expected_output by phase (primary vs replay)
claude Jun 5, 2026
700dd43
Add require_skill_load_on_replay opt-out for merge-focused evals
claude Jun 5, 2026
9f5c665
Make the consolidated skill discoverable on replay
claude Jun 5, 2026
d94a161
Split Memories column into Skill Generated / Skills Read
claude Jun 5, 2026
46f651a
Skills summary: drop interpretation, wrap stats in <details>
claude Jun 5, 2026
6eba53a
Merge origin/master into claude/consolidated-skills-per-domain
claude Jun 10, 2026
b5102f4
Add kind=discovery: capture schema/exploration learnings, not just fa…
claude Jun 10, 2026
f03e45a
Rename suggest_runbooks tool to suggest_skills
claude Jun 10, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
38 changes: 38 additions & 0 deletions .goal_criteria.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
# Sub-agent improvement criteria (set 2026-06-03, before any iteration)

## Metric definitions
- **Token cost** = LLMResult.total_tokens (prompt + completion + reasoning)
- **Accuracy** = LLMResult correctness score from the judge (0 or 1), AND
the user-facing answer matches expected_output.
- **Replay efficacy** = on rerun_with_memory evals, replay must:
(a) call fetch_skill (skill was deemed relevant), AND
(b) keep accuracy = 1.

## Improvement threshold (the goal's "30% on ≥5 evals")
For ≥5 distinct evals, in the same fixture/data state, the **replay run**
(memory captured then injected as a skill) must show:
- **total_tokens ≥30% lower** than the same eval's suggest=on primary
run (i.e. replay tokens ≤ 70% of primary tokens).
- **Accuracy = 1** on the replay.
- **fetch_skill called** on the replay.
Baseline for the primary run is the same eval's suggest=on numbers
captured under the same model (opus-4.6).

## Regression guard
The 11 existing regression evals (memories_generated: false) — i.e.
09_crashpod, 12_job_crashing, 24_misconfigured_pvc, 43_current_datetime_from_prompt,
51_logs_summarize_errors, 61_exact_match_counting,
101_loki_historical_logs_pod_deleted, 112_find_pvcs_by_uuid,
176_network_policy_blocking_traffic_no_skills,
227_count_configmaps_per_namespace, 243_pod_names_contain_service.
- Average accuracy across the 11 must not drop.
- Average token cost across the 11 must stay within ±10% of the
pre-iteration baseline.

## Verification model
- All measurements use **opus-4.6 via OpenRouter** (the project's main model).

## Source of truth
- Local runs in this sandbox produce the iteration numbers.
- CI on the PR is the final gate — every iteration must also produce
green CI (or be the candidate for green CI on the next push).
36 changes: 36 additions & 0 deletions tests/llm/conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -789,6 +789,8 @@ def _collect_test_results_from_stats(terminalreporter):
"braintrust_root_span_id": None,
"clean_test_case_id": None, # Not available for skipped tests
"env_config": "default", # Not available for skipped tests
"memories_count": 0,
"suggested_memories": [],
}
continue
elif when != "call":
Expand Down Expand Up @@ -863,6 +865,40 @@ def _collect_test_results_from_stats(terminalreporter):
), # Any throttling during execution
"model": user_props.get("model", "Unknown"),
"env_config": user_props.get("env_config", "default"),
"memories_count": user_props.get("memories_count", 0),
"suggested_memories": user_props.get("suggested_memories", []),
# Whether the primary pass (everything before the replay
# block) passed. Used by the GitHub reporter to keep the
# primary row ✅ even if the replay assertion fails the
# whole test.
"primary_passed": user_props.get("primary_passed", False),
# Replay-with-memory results (only populated when the test
# set rerun_with_memory: true AND a memory was actually
# captured on the first pass).
"replay_attempted": user_props.get("replay_attempted", False),
"replay_skill_loaded": user_props.get("replay_skill_loaded", False),
"replay_correctness": user_props.get("replay_correctness"),
"replay_turns": user_props.get("replay_turns"),
"replay_tool_calls_count": user_props.get("replay_tool_calls_count"),
"replay_skill_count": user_props.get("replay_skill_count"),
"replay_duration": user_props.get("replay_duration"),
"replay_total_cost": user_props.get("replay_total_cost"),
"replay_total_tokens": user_props.get("replay_total_tokens"),
"replay_prompt_tokens": user_props.get("replay_prompt_tokens"),
"replay_completion_tokens": user_props.get(
"replay_completion_tokens"
),
"replay_cached_tokens": user_props.get("replay_cached_tokens"),
"replay_reasoning_tokens": user_props.get(
"replay_reasoning_tokens"
),
"replay_max_completion_tokens_per_call": user_props.get(
"replay_max_completion_tokens_per_call"
),
"replay_max_prompt_tokens_per_call": user_props.get(
"replay_max_prompt_tokens_per_call"
),
"replay_num_compactions": user_props.get("replay_num_compactions"),
"clean_test_case_id": user_props.get("clean_test_case_id"),
"braintrust_span_id": user_props.get("braintrust_span_id"),
"braintrust_root_span_id": user_props.get("braintrust_root_span_id"),
Expand Down
2 changes: 2 additions & 0 deletions tests/llm/fixtures/test_ask_holmes/09_crashpod/test_case.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -53,3 +53,5 @@ tags:
- kubernetes
- one-test
- regression

memories_generated: false
Original file line number Diff line number Diff line change
Expand Up @@ -87,3 +87,5 @@ before_test: |
after_test: |
# Delete namespace in background to avoid hanging
kubectl delete namespace app-101 --wait=false || true

memories_generated: false
Original file line number Diff line number Diff line change
Expand Up @@ -30,3 +30,5 @@ before_test: |
fi
after_test: |
kubectl delete -f ./manifest.yaml

memories_generated: false
Original file line number Diff line number Diff line change
Expand Up @@ -38,3 +38,5 @@ after_test: |
tags:
- easy
- regression

memories_generated: false
Original file line number Diff line number Diff line change
Expand Up @@ -36,3 +36,5 @@ before_test: |
fi
after_test: |
kubectl delete namespace app-176 || true

memories_generated: false
Original file line number Diff line number Diff line change
Expand Up @@ -63,3 +63,5 @@ after_test: |
for NS in payments inventory shipping analytics notifications; do
kubectl delete namespace "app-227-$NS" || true
done

memories_generated: false
Original file line number Diff line number Diff line change
Expand Up @@ -28,3 +28,5 @@ tags:
- kubernetes
- medium
- regression

memories_generated: false
Original file line number Diff line number Diff line change
Expand Up @@ -30,3 +30,5 @@ before_test: |
after_test: |
kubectl delete -f manifest.yaml -n app-24
kubectl delete namespace app-24 || true

memories_generated: false
Original file line number Diff line number Diff line change
@@ -0,0 +1,138 @@
# Test: ES log severity field is named "severity" not "level"
#
# This eval forces a real failed→succeeded tool-call correction. The test
# index contains 10 ERROR-severity documents, but they use a custom
# "severity" field instead of the standard "level" field used by most
# log schemas. A fresh LLM defaulting to a level:"ERROR" query returns
# zero hits; the only way to find the data is to inspect the mapping or
# a sample doc, then re-query with severity:"ERROR".
#
# The LLM should capture this lesson as an env-specific memory
# ("this index uses 'severity' for log levels, not 'level'").

# Phrasing matters: by saying "where level is ERROR" we describe the
# question using the conventional log-field terminology, which biases
# the agent into running `level:"ERROR"` as its first query. That call
# returns zero hits (this index has only `severity`, no `level`) and
# the agent has to read the mapping and re-query with `severity:"ERROR"`
# to recover. THAT correction is the durable env-specific lesson the
# suggest_skills tool should capture.
#
# An earlier iteration tried a softer phrasing ("how many ERROR-level
# entries are there") and discovered the agent solves it on the first
# try (presumably from mapping enrichment) without ever issuing the
# wrong call — and therefore correctly emits zero memories. That's a
# different eval (whether emission happens on natural queries) than
# what this fixture is meant to test (whether the correction shape is
# captured when one is forced). Keep the biased phrasing here so the
# wrong→right correction is reliably triggered.
user_prompt: 'In the elasticsearch index ''app-261-logs-trk8m2p9'', run a `term: { level: "ERROR" }` query and tell me how many hits come back. If that returns zero, find the right field for log level and re-run the query.'

# Replay prompt: a natural, un-biased version of the same question. The
# primary pass uses the biased phrasing above to force the wrong→right
# correction (and reliably emit a memory). The replay simulates a future
# investigation asking the question naturally — which is when the
# captured skill should actually short-circuit the failed call. Without
# this softer phrasing the agent on replay just follows the explicit
# recovery instructions in the primary prompt and never bothers
# fetching the skill.
replay_user_prompt: "In the elasticsearch index 'app-261-logs-trk8m2p9', how many ERROR-level log entries are there for the checkout-261 service?"

expected_output:
- "There are 10 entries with ERROR severity / level in app-261-logs-trk8m2p9"

memories_generated: true
rerun_with_memory: true

tags:
- elasticsearch
- question-answer
- medium
- regression
- skills

setup_timeout: 300

before_test: |
source ../../shared/es_test_utils.sh
es_setup
set -e

export HOLMES_ES_TEST_INDEX="app-261-logs-trk8m2p9"
echo "Using test index: $HOLMES_ES_TEST_INDEX"

# Clean up any existing index first
curl -sf -X DELETE "${ELASTICSEARCH_URL}/${HOLMES_ES_TEST_INDEX}" \
-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" || true

echo "⏳ Creating test index with explicit mapping (severity, not level)..."

# Create the index with explicit mapping — note: NO "level" field exists,
# only "severity". This is the env-specific quirk this eval teaches.
CREATE_RESPONSE=$(curl -sf -X PUT "${ELASTICSEARCH_URL}/${HOLMES_ES_TEST_INDEX}?wait_for_active_shards=1" \
-H "Content-Type: application/json" \
-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" \
-d '{
"settings": { "number_of_shards": 1, "number_of_replicas": 0 },
"mappings": {
"properties": {
"severity": { "type": "keyword" },
"message": { "type": "text" },
"service": { "type": "keyword" },
"@timestamp": { "type": "date" }
}
}
}')

if ! echo "$CREATE_RESPONSE" | grep -q '"acknowledged":true'; then
echo "❌ Failed to create index: $CREATE_RESPONSE"
exit 1
fi
sleep 2

# 10 ERROR, 5 WARN, 15 INFO — all using the "severity" field. Pad the
# seconds field so i >= 10 still produces a valid ISO timestamp; an
# earlier version of this fixture concatenated "0${i}" which produced
# "12:00:010" for i=10 and ES rejected the doc.
BULK_DATA=""
for i in $(seq 1 10); do
TS=$(printf '2026-05-14T12:00:%02dZ' "$i")
BULK_DATA="${BULK_DATA}{\"index\":{}}\n{\"severity\":\"ERROR\",\"service\":\"checkout-261\",\"message\":\"DB connection refused $i\",\"@timestamp\":\"${TS}\"}\n"
done
for i in $(seq 1 5); do
TS=$(printf '2026-05-14T12:01:%02dZ' "$i")
BULK_DATA="${BULK_DATA}{\"index\":{}}\n{\"severity\":\"WARN\",\"service\":\"checkout-261\",\"message\":\"Slow query $i\",\"@timestamp\":\"${TS}\"}\n"
done
for i in $(seq 1 15); do
TS=$(printf '2026-05-14T12:02:%02dZ' "$i")
BULK_DATA="${BULK_DATA}{\"index\":{}}\n{\"severity\":\"INFO\",\"service\":\"checkout-261\",\"message\":\"Health check $i\",\"@timestamp\":\"${TS}\"}\n"
done

BULK_RESPONSE=$(echo -e "$BULK_DATA" | curl -sf -X POST "${ELASTICSEARCH_URL}/${HOLMES_ES_TEST_INDEX}/_bulk" \
-H "Content-Type: application/x-ndjson" \
-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" \
--data-binary @-)

if echo "$BULK_RESPONSE" | grep -q '"errors":true'; then
echo "❌ Bulk insert had errors: $BULK_RESPONSE"
exit 1
fi

curl -sf -X POST "${ELASTICSEARCH_URL}/${HOLMES_ES_TEST_INDEX}/_refresh" \
-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" > /dev/null

DOC_COUNT=$(curl -sf -X GET "${ELASTICSEARCH_URL}/${HOLMES_ES_TEST_INDEX}/_count" \
-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" | grep -o '"count":[0-9]*' | cut -d':' -f2)

if [ "$DOC_COUNT" = "30" ]; then
echo "✅ Index created with $DOC_COUNT docs (10 ERROR / 5 WARN / 15 INFO) using 'severity' field"
else
echo "❌ Expected 30 documents, found: $DOC_COUNT"
exit 1
fi

after_test: |
echo "⏳ Cleaning up test index: app-261-logs-trk8m2p9"
curl -sf -X DELETE "${ELASTICSEARCH_URL}/app-261-logs-trk8m2p9" \
-H "Authorization: ApiKey ${ELASTICSEARCH_API_KEY}" || true
echo "✅ Cleanup complete"
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# Shared Elasticsearch toolset configuration for eval tests
# Copy this content to test toolsets.yaml files
toolsets:
elasticsearch/data:
enabled: true
config:
api_url: "{{ env.ELASTICSEARCH_URL }}"
api_key: "{{ env.ELASTICSEARCH_API_KEY }}"
verify_ssl: true
elasticsearch/cluster:
enabled: true
config:
api_url: "{{ env.ELASTICSEARCH_URL }}"
api_key: "{{ env.ELASTICSEARCH_API_KEY }}"
verify_ssl: true
Loading
Loading