Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/llm-evaluation.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ jobs:
EXPERIMENT_ID: github-${{ github.run_id }}.${{ github.run_number }}.${{ github.run_attempt }}
GENERATE_REGRESSIONS_FILE: "true"
run: |
poetry run pytest --no-cov tests/llm/test_ask_holmes.py tests/llm/test_investigate.py -m 'not skip' -n 6
poetry run pytest --no-cov tests/llm/test_ask_holmes.py tests/llm/test_investigate.py -n 6
- uses: actions/github-script@v7
if: always()
with:
Expand Down
4 changes: 3 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -169,7 +169,6 @@ poetry run pytest -m "llm and not network" --collect-only -q
- `network`: Network-dependent tests
- `runbooks`: Runbook functionality
- `misleading-history`: Misleading data scenarios
- `k8s-misconfig`: Kubernetes misconfigurations
- `chain-of-causation`: Causation analysis
- `slackbot`: Slack integration
- `counting`: Resource counting tests
Expand Down Expand Up @@ -232,3 +231,6 @@ poetry run pytest -m "llm and not network" --collect-only -q
- No secrets should be committed to repository
- Use environment variables or config files for API keys
- RBAC permissions are respected for Kubernetes access

## Eval Notes
- You can run evals with --skip-cleanup or --skip-setup if you are debugging the eval itself
7 changes: 4 additions & 3 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -101,12 +101,13 @@ markers = [
"slackbot: Tests involving Slack bot functionality",
"understanding-real-intent: Tests that require understanding the user's real intent behind a question, not just giving a technically correct answer that is not useful",
"counting: Ask holmes to count kubernetes/cloud resources",
"reproducible: Tests with before_test setup that can be reliably reproduced",
"kubernetes: Tests for Kubernetes-specific troubleshooting scenarios",
"hard: Tests that are hard and holmes does not pass today",
"easy: Tests that are supposed to pass - if these fail, it indicates a regression",
"medium: Tests that we want to focus on in the near future and get holmes to pass them",
"hard: Tests that are hard and holmes does not pass today and might not pass in the near future",
"missing-tool: Verify holmes knows how to communicate this to the user",
"kafka: Tests involving Kafka functionality"
"kafka: Tests involving Kafka functionality",
"leaked-information: cases where the eval is accidentally leaking information that should not be available to the LLM - e.g. if the names of environment variables or the docker image gives away part of the test"
]
Comment thread
aantn marked this conversation as resolved.

addopts = [
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,6 @@ before_test: |
sleep 60
after_test: |
kubectl delete -f manifests.yaml
evaluation:
correctness: 1
tags:
- counting
- reproducible
- easy

This file was deleted.

This file was deleted.

This file was deleted.

Original file line number Diff line number Diff line change
@@ -1,5 +1,40 @@
user_prompt: 'What is wrong with << { "type": "pod", "name": "robusta-runner-6958c5bdd8-69gtn" } >> ?'
user_prompt: 'What is wrong with << { "type": "pod", "name": "giant-narwhal-6958c5bdd8-69gtn" } >> ?'
expected_output:
- The pod was killed due to it being out of memory
evaluation:
correctness: 1
tags:
- kubernetes
# need to check, but according to coderabbit this wont work
# CodeRabbit: BusyBox shell can’t handle array syntax – pod will CrashLoop, not OOMKill
# busybox’s /bin/sh (ash) does not support Bash-style arrays (arr=(); arr+=($i)), so
# the container exits immediately with a syntax error instead of allocating memory.
# Consequently the pod goes into CrashLoopBackOff, never hitting the memory limit,
# and the test’s correctness assertion (“OOMKilled”) will fail.
# Alternative: switch to stress/stress-ng for a cleaner, deterministic OOM trigger.
before_test: |
kubectl create namespace app-02
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: giant-narwhal-6958c5bdd8-69gtn
namespace: app-02
labels:
app: ocean-explorer
spec:
containers:
- name: main
image: busybox:1.35
command: ["sh", "-c", "while true; do echo 'Processing data...'; arr=(); for i in $(seq 1 100000000); do arr+=($i); done; done"]
resources:
limits:
memory: "50Mi"
requests:
memory: "50Mi"
Comment thread
aantn marked this conversation as resolved.
EOF
# Wait for pod to be OOMKilled
sleep 30

after_test: |
kubectl delete namespace app-02 --force --grace-period=0
Original file line number Diff line number Diff line change
@@ -1,7 +1,32 @@
user_prompt: 'what is the command to port-forward to << { "type": "pod", "name": "my_grafana_4j981" } >>'
#user_prompt: "what is the command to port-forward to my grafana service?"
user_prompt: 'port forward to grafana pod in app-03'
expected_output:
- A kubectl port-forward command
- The pod name in the port forward command must be my_grafana_4j981
evaluation:
correctness: 1
- Must find the actual grafana pod name (grafana-7f8b9c6d5-x2m4n)
- Must include the correct port (3000)
- "Full kubectl port-forward command like: kubectl port-forward -n app-03 grafana-7f8b9c6d5-x2m4n 3000:3000"
- Should NOT just say "find grafana in your cluster" or give generic instructions
tags:
- kubernetes
- easy
before_test: |
kubectl create namespace app-03
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: grafana-7f8b9c6d5-x2m4n
namespace: app-03
labels:
app: grafana
spec:
containers:
- name: grafana
image: busybox:1.35
command: ['sh', '-c', 'while true; do nc -l -p 3000; done']
ports:
- containerPort: 3000
name: http
EOF
kubectl wait --for=condition=Ready pod/grafana-7f8b9c6d5-x2m4n -n app-03 --timeout=60s

after_test: |
kubectl delete namespace app-03 --force --grace-period=0
Loading