Repository navigation
Conversation
New eval tests targeting real-world latency scenarios where scaling can't keep up with demand: - 212: ResourceQuota blocks HPA scale-up, pods can't be created - 213: Pending pods due to insufficient node CPU resources - 214: OOMKill during scaling creates restart loops - 215: HPA at maxReplicas with Prometheus metrics showing saturation All tests use neutral service names (order-gateway, payment-processor, inventory-service, notification-service) and vague latency prompts to test causal chain reasoning from surface symptom to root cause. https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok Signed-off-by: Claude <noreply@anthropic.com>
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
|
📂 Previous Runs📜 Run @ 252eaa2 (#21768849052)✅ Results of HolmesGPT evalsAutomatically triggered by commit 252eaa2 on branch Results of HolmesGPT evals
|
| Status | Test case | Time | Turns | Tools | Cost |
|---|---|---|---|---|---|
| ✅ | 09_crashpod | 33.2s | 5 | 11 | $0.2286 |
| ✅ | 101_loki_historical_logs_pod_deleted | 54.2s | 7 | 14 | $0.3005 |
| ✅ | 111_pod_names_contain_service | 35.6s | 5 | 11 | $0.2308 |
| ✅ | 112_find_pvcs_by_uuid | 31.1s | 5 | 7 | $0.2536 |
| ✅ | 12_job_crashing | 34.0s | 5 | 11 | $0.2271 |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 47.4s | 6 | 18 | $0.2903 |
| ✅ | 212_resource_quota_scaling_blocked[0] | 63.0s | 8 | 21 | $0.4196 |
| ✅ | 213_pending_pods_node_pressure[0] | 43.1s | 5 | 13 | $0.2845 |
| ✅ | 214_oomkill_during_scaling[0] | 46.9s | 6 | 16 | $0.2944 |
| ❌ | 215_hpa_maxed_out[0] | 138.0s | 18 | 55 | $1.0281 |
| ❌ | 216_hpa_ceiling_latency[0] | 88.5s | 13 | 37 | $0.6256 |
| ✅ | 24_misconfigured_pvc | 33.5s | 5 | 12 | $0.2211 |
| ✅ | 43_current_datetime_from_prompt | 5.1s | 1 | — | $0.1050 |
| ✅ | 61_exact_match_counting | 17.5s | 4 | 4 | $0.1595 |
| Total | 47.9s avg | 6.6 avg | 17.7 avg | $4.6687 |
⚠️ 2 Failures Detected
📜 Run @ e87aa10 (#21767736258)
✅ Results of HolmesGPT evals
Automatically triggered by commit e87aa10 on branch claude/latency-scenario-evals-MHcZ3
Results of HolmesGPT evals
- ask_holmes: 11/14 test cases were successful, 2 regressions, 1 setup failures
| Status | Test case | Time | Turns | Tools | Cost |
|---|---|---|---|---|---|
| ✅ | 09_crashpod | 29.6s | 5 | 10 | $0.2210 |
| ✅ | 101_loki_historical_logs_pod_deleted | 46.1s | 7 | 10 | $0.2646 |
| ✅ | 111_pod_names_contain_service | 34.4s | 6 | 11 | $0.2356 |
| ✅ | 112_find_pvcs_by_uuid | 31.8s | 6 | 7 | $0.2333 |
| ✅ | 12_job_crashing | 32.6s | 5 | 12 | $0.2418 |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 41.4s | 6 | 17 | $0.2870 |
| 🚧 | 212_resource_quota_scaling_blocked[0] | — | — | — | — |
| ✅ | 213_pending_pods_node_pressure[0] | 45.4s | 5 | 15 | $0.3168 |
| ✅ | 214_oomkill_during_scaling[0] | 38.6s | 5 | 13 | $0.2520 |
| ❌ | 215_hpa_maxed_out[0] | 113.2s | 14 | 52 | $0.8710 |
| ❌ | 216_hpa_ceiling_latency[0] | 57.7s | 7 | 18 | $0.3739 |
| ✅ | 24_misconfigured_pvc | 38.0s | 7 | 14 | $0.2517 |
| ✅ | 43_current_datetime_from_prompt | 4.9s | 1 | — | $0.1058 |
| ✅ | 61_exact_match_counting | 16.3s | 4 | 4 | $0.1600 |
| Total | 40.8s avg | 6.0 avg | 15.2 avg | $3.8146 |
⚠️ 2 Failures Detected
📜 Run @ 4eaad45 (#21758709544)
✅ Results of HolmesGPT evals
Automatically triggered by commit 4eaad45 on branch claude/latency-scenario-evals-MHcZ3
Results of HolmesGPT evals
- ask_holmes: 9/14 test cases were successful, 2 regressions, 3 setup failures
| Status | Test case | Time | Turns | Tools | Cost |
|---|---|---|---|---|---|
| ✅ | 09_crashpod | 27.6s | 4 | 10 | $0.2093 |
| ✅ | 101_loki_historical_logs_pod_deleted | 54.2s | 8 | 11 | $0.3024 |
| ✅ | 111_pod_names_contain_service | 33.7s | 5 | 11 | $0.2204 |
| ✅ | 112_find_pvcs_by_uuid | 38.6s | 7 | 8 | $0.2538 |
| ✅ | 12_job_crashing | 40.9s | 6 | 15 | $0.2742 |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 50.8s | 6 | 18 | $0.2925 |
| 🚧 | 212_resource_quota_scaling_blocked[0] | — | — | — | — |
| ✅ | 213_pending_pods_node_pressure[0] | 47.1s | 5 | 16 | $0.3052 |
| ❌ | 214_oomkill_during_scaling[0] | 80.2s | 10 | 26 | $0.4753 |
| ❌ | 215_hpa_maxed_out[0] | 154.0s | 20 | 53 | $1.2317 |
| 🚧 | 216_hpa_ceiling_latency[0] | — | — | — | — |
| ✅ | 24_misconfigured_pvc | 35.5s | 6 | 12 | $0.2336 |
| ✅ | 43_current_datetime_from_prompt | 4.9s | 1 | — | $0.1048 |
| 🚧 | 61_exact_match_counting | — | — | — | — |
| Total | 51.6s avg | 7.1 avg | 18.0 avg | $3.9032 |
⚠️ 2 Failures Detected
📜 Run @ f7fd3aa (#21755906121)
✅ Results of HolmesGPT evals
Automatically triggered by commit f7fd3aa on branch claude/latency-scenario-evals-MHcZ3
Results of HolmesGPT evals
- ask_holmes: 12/13 test cases were successful, 1 regressions
| Status | Test case | Time | Turns | Tools | Cost |
|---|---|---|---|---|---|
| ✅ | 09_crashpod | 35.2s | 5 | 11 | $0.2378 |
| ✅ | 101_loki_historical_logs_pod_deleted | 36.7s | 5 | 9 | $0.2390 |
| ✅ | 111_pod_names_contain_service | 31.8s | 5 | 10 | $0.2203 |
| ✅ | 112_find_pvcs_by_uuid | 33.7s | 6 | 8 | $0.2474 |
| ✅ | 12_job_crashing | 39.1s | 6 | 14 | $0.2617 |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 45.3s | 6 | 16 | $0.2949 |
| ✅ | 212_resource_quota_scaling_blocked[0] | 43.1s | 6 | 16 | $0.2793 |
| ✅ | 213_pending_pods_node_pressure[0] | 40.3s | 5 | 12 | $0.2654 |
| ✅ | 214_oomkill_during_scaling[0] | 43.1s | 5 | 14 | $0.2570 |
| ❌ | 215_hpa_maxed_out[0] | 136.3s | 22 | 45 | $0.9846 |
| ✅ | 24_misconfigured_pvc | 31.8s | 5 | 13 | $0.2290 |
| ✅ | 43_current_datetime_from_prompt | 4.4s | 1 | — | $0.1051 |
| ✅ | 61_exact_match_counting | 17.2s | 4 | 4 | $0.1599 |
| Total | 41.4s avg | 6.2 avg | 14.3 avg | $3.7815 |
⚠️ 1 Failure Detected
✅ Results of HolmesGPT evals
Automatically triggered by commit 6c65772 on branch claude/latency-scenario-evals-MHcZ3
Results of HolmesGPT evals
- ask_holmes: 9/11 test cases were successful, 2 regressions
| Status | Test case | Time | Turns | Tools | Cost |
|---|---|---|---|---|---|
| ✅ | 09_crashpod | 33.0s | 5 | 11 | $0.2293 |
| ✅ | 101_loki_historical_logs_pod_deleted | 31.9s | 4 | 7 | $0.2048 |
| ✅ | 111_pod_names_contain_service | 27.6s | 4 | 9 | $0.1969 |
| ✅ | 112_find_pvcs_by_uuid | 34.6s | 6 | 8 | $0.2576 |
| ✅ | 12_job_crashing | 32.0s | 5 | 11 | $0.2327 |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 43.2s | 7 | 17 | $0.2993 |
| ❌ | 215_hpa_maxed_out[0] | 102.3s | 17 | 35 | $0.7137 |
| ❌ | 216_hpa_ceiling_latency[0] | 106.5s | 9 | 29 | $0.5008 |
| ✅ | 24_misconfigured_pvc | 39.8s | 7 | 14 | $0.2559 |
| ✅ | 43_current_datetime_from_prompt | 4.7s | 1 | — | $0.0112 |
| ✅ | 61_exact_match_counting | 13.3s | 3 | 2 | $0.1427 |
| Total | 42.6s avg | 6.2 avg | 14.3 avg | $3.0450 |
⚠️ 2 Failures Detected
📖 Legend
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/latency-scenario-evals-MHcZ3 -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
markers: regression
Or with more options (one per line):
/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
markers: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
markers |
Pytest markers (no default - runs all tests!) |
filter |
Pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
🏷️ Valid markers
benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/latency-scenario-evals-MHcZ3 -f markers=regression -f filter=
|
✅ Docker image ready for
Use this tag to pull the image for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:89d2589
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:89d2589 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:89d2589
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:89d2589Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:89d2589Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:89d2589 |
WalkthroughAdds multiple Kubernetes end-to-end test fixtures (namespaces app-212–app-216) that simulate HPA and resource constraint scenarios: manifests (Deployments, Services, Secrets, ConfigMaps, HPAs), traffic generators, Prometheus configs, and orchestrated test_case.yaml scripts for automated setup, load, observation, and teardown. Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok Signed-off-by: Claude <noreply@anthropic.com>
🔬 CLI Performance Benchmark🟡 Startup Time (no LLM)Measures
🟡 Full CLI with LLMMeasures
PR: |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Fix all issues with AI agents
In
`@tests/llm/fixtures/test_ask_holmes/212_resource_quota_scaling_blocked/test_case.yaml`:
- Around line 53-66: The quota-event wait loop (the for loop that checks kubectl
get events -n app-212 for reason=FailedCreate and sets EVENTS_FOUND) currently
only logs a warning when no "exceeded quota" events are found; make this failure
fatal like the pod readiness check by calling exit 1 when EVENTS_FOUND is still
false after the loop (or, alternatively, add an explicit comment if this
non-fatal behavior is deliberate). Locate the loop that greps for "exceeded
quota" and the subsequent if [ "$EVENTS_FOUND" = false ]; then block and replace
the warning-only branch with a clear failing path that echoes a descriptive
error and runs exit 1 so the test fails fast on missing quota events.
In `@tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/cpu-stress.yaml`:
- Around line 3-7: The resource name cpu-stress in the manifest should be
changed to a neutral, realistic service name (e.g., analytics-worker or
request-client) to avoid revealing test intent; update the metadata.name value
and also update any matching labels/selectors (labels.app, pod template labels,
and any Deployment/Service selector references) so they remain consistent with
the new name (look for occurrences of "cpu-stress", metadata.name, labels.app
and selector/template labels in this file and associated manifests and rename
them to the chosen neutral identifier).
🧹 Nitpick comments (3)
tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/prometheus-config.yaml (1)
19-32: Redundant relabel rule on lines 23-27 — immediately overwritten by lines 28-32.The second relabel rule sets
__address__to the bare port annotation value (e.g.,"8080"), but the third rule immediately overwrites it withpod_ip:port. The second rule is dead config and could leave an invalid__address__(just a port number) if__meta_kubernetes_pod_ipwere ever unset.Remove the intermediate rule:
Suggested fix
relabel_configs: - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape] action: keep regex: true - - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port] - action: replace - target_label: __address__ - regex: (.+) - replacement: ${1} - source_labels: [__meta_kubernetes_pod_ip, __meta_kubernetes_pod_annotation_prometheus_io_port] action: replace target_label: __address__ regex: (.+);(.+) replacement: ${1}:${2}tests/llm/fixtures/test_ask_holmes/214_oomkill_during_scaling/inventory-service.yaml (1)
37-57: Consider adding asecurityContextto avoid running as root.Checkov flags
CKV_K8S_20andCKV_K8S_23for missingallowPrivilegeEscalation: falseandrunAsNonRoot. For a test fixture using busybox, this is low priority, but adding it would silence static analysis and follow K8s hardening best practices.🔒 Optional securityContext addition
containers: - name: inventory-service image: busybox:1.36 command: ["/bin/sh", "/app/app.sh"] + securityContext: + allowPrivilegeEscalation: false + runAsNonRoot: true + runAsUser: 1000 ports: - containerPort: 8080tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/notification-service.yaml (1)
78-105: Optional: addsecurityContextto silence Checkov warnings.Same as the inventory-service manifest — Checkov flags
CKV_K8S_20andCKV_K8S_23. Low priority for test fixtures but good hygiene.🔒 Optional securityContext addition
containers: - name: notification-service image: me-west1-docker.pkg.dev/robusta-development/development/python-flask-otel:2.2 imagePullPolicy: IfNotPresent command: ["python", "/app/app.py"] + securityContext: + allowPrivilegeEscalation: false ports: - containerPort: 8080
| metadata: | ||
| name: cpu-stress | ||
| namespace: app-215 | ||
| labels: | ||
| app: cpu-stress |
There was a problem hiding this comment.
🛠️ Refactor suggestion | 🟠 Major
Resource name cpu-stress hints at the problem — use a neutral name.
The name directly reveals this deployment's role in the test scenario. Consider a real-world service name (e.g., analytics-worker, request-client, or similar) to keep the fixture realistic and not leak test intent to the LLM under evaluation.
Example rename
metadata:
- name: cpu-stress
+ name: analytics-worker
namespace: app-215
labels:
- app: cpu-stress
+ app: analytics-worker(Also update selector/template labels accordingly.)
As per coding guidelines, tests/llm/fixtures/test_ask_holmes/**/*.yaml: "Use real-world, neutral resource names that don't hint at problems (avoid names like 'broken-pod', 'crashloop-app')". Based on learnings: "Use real-world, neutral resource names that don't hint at problems."
🤖 Prompt for AI Agents
In `@tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/cpu-stress.yaml` around
lines 3 - 7, The resource name cpu-stress in the manifest should be changed to a
neutral, realistic service name (e.g., analytics-worker or request-client) to
avoid revealing test intent; update the metadata.name value and also update any
matching labels/selectors (labels.app, pod template labels, and any
Deployment/Service selector references) so they remain consistent with the new
name (look for occurrences of "cpu-stress", metadata.name, labels.app and
selector/template labels in this file and associated manifests and rename them
to the chosen neutral identifier).
|
/eval |
|
@aantn Your eval run has finished. 🧪 Manual Eval Results
Results of HolmesGPT evals
|
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/latency-scenario-evals-MHcZ3 -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
markers: regression
Or with more options (one per line):
/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
markers: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
markers |
Pytest markers (no default - runs all tests!) |
filter |
Pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
🏷️ Valid markers
benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/latency-scenario-evals-MHcZ3 -f markers=regression -f filter=
212 (ResourceQuota): Replace manual scale with real Flask app + traffic generator that drives HPA organically. Add decoy services (session-cache, order-db) consuming quota to make the failure less obvious. Holmes must dig into ReplicaSet events to find the quota error. 213 (Pending pods): Add 3 supporting services (payment-cache, payment-queue, payment-db) as noise in kubectl output. Add fraud-detection service that logs connection pool warnings as a red herring. Holmes must distinguish noisy-but-running pods from the actual Pending pods. 214 (OOMKill): Replace single deployment with v1→v2 rolling update. V1 is stable and works fine. V2 has a memory leak that OOMKills. Holmes sees a mid-rollout state and must figure out the new version has a memory bug, not just a slow rollout. Add supporting services. 216 (NEW - HPA ceiling): Clean variant of 215 using gunicorn instead of Flask dev server, eliminating the competing red herring. Add decoy services. The ONLY path to the answer is noticing HPA at maxReplicas. All tests use include_tool_calls: true to verify Holmes actually queries the right resources rather than guessing from domain knowledge. https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Fix all issues with AI agents
In
`@tests/llm/fixtures/test_ask_holmes/213_pending_pods_node_pressure/test_case.yaml`:
- Around line 47-54: The supporting-services wait loop for pods labeled
tier=infra in namespace app-213 currently falls through silently if pods never
become ready; update the loop (the for ...; do ... kubectl wait
--for=condition=ready pod -l tier=infra -n app-213 --timeout=5s ...) to mirror
the payment-processor pattern by adding a post-loop check that detects failure
(e.g., if the kubectl wait never succeeded) and prints an error like "❌
Supporting services failed to become ready" and exits non-zero (exit 1) so the
test fails fast and clearly.
In
`@tests/llm/fixtures/test_ask_holmes/214_oomkill_during_scaling/inventory-service-v2.yaml`:
- Around line 15-23: The loop's DATA variable is overwritten each iteration so
memory doesn't truly accumulate; update the script to either append to DATA (use
DATA="$DATA$(dd ... | base64)") to force deterministic accumulation and trigger
OOMKill, or change the comment near BATCH/DATA/while true to remove the
misleading claim about additive growth and note that OOM is caused by transient
peak during the dd|base64 capture; modify the block referencing DATA, BATCH and
the while loop accordingly.
🧹 Nitpick comments (3)
tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/test_case.yaml (1)
98-114: HPA wait loop does not fail the setup on timeout.The HPA wait loop logs a warning but does not
exit 1when the HPA fails to reachmaxReplicaswithin 90 iterations. This is fine if the test is designed to be resilient to partial setup, but note that the dispatch-service starts withreplicas: 3and the HPA hasminReplicas: 3/maxReplicas: 3, so this condition should be trivially satisfied almost immediately without needing traffic. If this loop consistently times out, it may indicate a deeper issue with the HPA controller not reportingcurrentReplicas.tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/dispatch-service.yaml (1)
38-52: Nit:import hashlibinside the loop.The
hashlibimport on line 44 (within theapp.pystring) is inside theforloop. While CPython caches modules insys.modulesso this is functionally harmless, moving it to the top of the file alongside the other imports would be more idiomatic.Proposed fix
from flask import Flask, Response, jsonify from prometheus_client import Counter, Histogram, generate_latest, CONTENT_TYPE_LATEST import time import random import os + import hashlib ... for _ in range(300): - import hashlib data = hashlib.sha256(data.encode()).hexdigest()tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/prometheus-config.yaml (1)
23-27: Redundant relabel rule — immediately overwritten by the next rule.Rule at lines 23–27 sets
__address__to just the port value (e.g.,"8080"), but the next rule (lines 28–32) unconditionally overwrites__address__withpod_ip:port. This first replacement has no effect and adds confusion.Remove the redundant rule
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape] action: keep regex: true - - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port] - action: replace - target_label: __address__ - regex: (.+) - replacement: ${1} - source_labels: [__meta_kubernetes_pod_ip, __meta_kubernetes_pod_annotation_prometheus_io_port] action: replace target_label: __address__ regex: (.+);(.+) replacement: ${1}:${2}
- 212: Add exit 1 on quota event fallback failure (was silently continuing) - 213: Add exit 1 on supporting services wait failure (was falling through) - 214: Fix memory accumulation bug (CACHE grows across iterations instead of overwriting DATA each iteration), add exit 1 on OOMKill/infra waits - 215: Rename cpu-stress to analytics-worker (name revealed scenario), fix redundant Prometheus relabel rule - 216: Remove gunicorn (not in image), use Flask threaded mode instead, move hashlib import to top level, fix redundant Prometheus relabel rule https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 3
🤖 Fix all issues with AI agents
In `@tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/test_case.yaml`:
- Around line 100-103: The current HPA_MAXED conditional only emits a warning;
change it to a hard failure so the test fails when HPA doesn't reach
maxReplicas: replace the soft echo/inspection branch that checks HPA_MAXED with
a failing branch that prints a descriptive error (including kubectl get hpa -n
app-215 output for diagnostics) and exits non-zero (e.g., exit 1) similar to the
pod readiness checks; update any test documentation only if you intentionally
keep it soft, but prefer converting the HPA_MAXED check into a hard failure to
ensure the core assertion is enforced.
- Around line 75-85: The wait loop that checks for analytics-worker stress pods
(loop using RUNNING, STRESS_READY, label app=analytics-worker in namespace
app-215) lacks an exit-on-failure like the other loops; after the for loop
finishes, check if STRESS_READY is still false and if so print an error and call
exit 1 so the test fails early; update the block that sets STRESS_READY and the
surrounding loop to mirror the behavior used for notification-service and
Prometheus waits.
In
`@tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/dispatch-service.yaml`:
- Around line 18-19: The code retrieves the Werkzeug logger into the variable
cli but never configures it, so the intended suppression is a no-op; update the
code that obtains the logger (cli = logging.getLogger('werkzeug')) to explicitly
set an appropriate level (e.g., cli.setLevel(logging.ERROR) or logging.WARNING)
and/or attach a NullHandler or desired handler to suppress the Flask dev server
warning (use cli.addHandler(logging.NullHandler()) or similar) so the logger is
actually silenced during tests.
🧹 Nitpick comments (3)
tests/llm/fixtures/test_ask_holmes/214_oomkill_during_scaling/inventory-service-v2.yaml (1)
51-67: Consider adding a minimalsecurityContextto the container spec.Static analysis flags that the container runs without
allowPrivilegeEscalation: falseandrunAsNonRoot: true. While this is a test fixture, adding a security context is low-effort and aligns with best practices. The busybox script doesn't need root.🔒 Proposed securityContext addition
- name: inventory-service image: busybox:1.36 command: ["/bin/sh", "/app/run.sh"] + securityContext: + allowPrivilegeEscalation: false + runAsNonRoot: true + runAsUser: 1000 ports: - containerPort: 8080tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/dispatch-service.yaml (1)
66-137: Security context not set on the container (Checkov CKV_K8S_20, CKV_K8S_23).Static analysis flags that
allowPrivilegeEscalationisn't explicitly disabled and the container may run as root. For test fixtures this is low-risk, but adding a minimalsecurityContextis a good habit and silences the warnings.Suggested securityContext addition
resources: requests: memory: "64Mi" cpu: "50m" limits: memory: "128Mi" cpu: "100m" + securityContext: + allowPrivilegeEscalation: false + runAsNonRoot: true + runAsUser: 1000tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/analytics-worker.yaml (1)
19-19: Container name "stress" may hint at the problem to the LLM.The resource name
analytics-workeris appropriately neutral, but the container namestresscould leak intokubectl describeoutput that Holmes inspects, potentially hinting at the root cause. Consider renaming it to something neutral likeworkeroranalytics-worker.Based on learnings: "Use real-world, neutral resource names that don't hint at problems."
Suggested rename
- - name: stress + - name: worker
| echo "⏳ Waiting for stress job pods to start..." | ||
| STRESS_READY=false | ||
| for i in {1..60}; do | ||
| RUNNING=$(kubectl get pods -n app-215 -l app=analytics-worker --no-headers 2>/dev/null | grep -c "Running" || echo "0") | ||
| if [ "$RUNNING" -gt "0" ]; then | ||
| echo "✅ Stress generators running!" | ||
| STRESS_READY=true | ||
| break | ||
| fi | ||
| sleep 1 | ||
| done |
There was a problem hiding this comment.
Missing exit-on-failure when stress pods fail to start.
The other wait loops (notification-service pods at Line 47, Prometheus at Line 63) exit with exit 1 on failure. This block silently continues if no analytics-worker pods reach Running state. If the stress generator never starts, HPA won't scale and the test precondition won't hold.
This is inconsistent with the commit message that mentions "Added explicit exit-on-failure behavior."
Proposed fix
sleep 1
done
+ if [ "$STRESS_READY" = false ]; then
+ echo "❌ Stress generators failed to start"
+ kubectl get pods -n app-215
+ exit 1
+ fi
echo "⏳ Waiting for HPA to hit maxReplicas and stabilize..."📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| echo "⏳ Waiting for stress job pods to start..." | |
| STRESS_READY=false | |
| for i in {1..60}; do | |
| RUNNING=$(kubectl get pods -n app-215 -l app=analytics-worker --no-headers 2>/dev/null | grep -c "Running" || echo "0") | |
| if [ "$RUNNING" -gt "0" ]; then | |
| echo "✅ Stress generators running!" | |
| STRESS_READY=true | |
| break | |
| fi | |
| sleep 1 | |
| done | |
| echo "⏳ Waiting for stress job pods to start..." | |
| STRESS_READY=false | |
| for i in {1..60}; do | |
| RUNNING=$(kubectl get pods -n app-215 -l app=analytics-worker --no-headers 2>/dev/null | grep -c "Running" || echo "0") | |
| if [ "$RUNNING" -gt "0" ]; then | |
| echo "✅ Stress generators running!" | |
| STRESS_READY=true | |
| break | |
| fi | |
| sleep 1 | |
| done | |
| if [ "$STRESS_READY" = false ]; then | |
| echo "❌ Stress generators failed to start" | |
| kubectl get pods -n app-215 | |
| exit 1 | |
| fi | |
| echo "⏳ Waiting for HPA to hit maxReplicas and stabilize..." |
🤖 Prompt for AI Agents
In `@tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/test_case.yaml` around
lines 75 - 85, The wait loop that checks for analytics-worker stress pods (loop
using RUNNING, STRESS_READY, label app=analytics-worker in namespace app-215)
lacks an exit-on-failure like the other loops; after the for loop finishes,
check if STRESS_READY is still false and if so print an error and call exit 1 so
the test fails early; update the block that sets STRESS_READY and the
surrounding loop to mirror the behavior used for notification-service and
Prometheus waits.
| if [ "$HPA_MAXED" = false ]; then | ||
| echo "⚠️ HPA did not reach maxReplicas within timeout" | ||
| kubectl get hpa -n app-215 | ||
| fi |
There was a problem hiding this comment.
HPA not reaching maxReplicas is only a warning, not a failure.
If HPA doesn't reach maxReplicas, the test's core assertion (that Holmes detects HPA is maxed out) cannot pass. Consider making this a hard failure like the pod readiness checks, or at minimum document why this is intentionally a soft warning.
🤖 Prompt for AI Agents
In `@tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/test_case.yaml` around
lines 100 - 103, The current HPA_MAXED conditional only emits a warning; change
it to a hard failure so the test fails when HPA doesn't reach maxReplicas:
replace the soft echo/inspection branch that checks HPA_MAXED with a failing
branch that prints a descriptive error (including kubectl get hpa -n app-215
output for diagnostics) and exits non-zero (e.g., exit 1) similar to the pod
readiness checks; update any test documentation only if you intentionally keep
it soft, but prefer converting the HPA_MAXED check into a hard failure to ensure
the core assertion is enforced.
| # Suppress Flask dev server warning — this is intentional for test purposes | ||
| cli = logging.getLogger('werkzeug') |
There was a problem hiding this comment.
Werkzeug logger retrieved but never configured.
The comment says "Suppress Flask dev server warning" but the logger level is never set. This is a no-op.
Proposed fix
# Suppress Flask dev server warning — this is intentional for test purposes
- cli = logging.getLogger('werkzeug')
+ cli = logging.getLogger('werkzeug')
+ cli.setLevel(logging.ERROR)📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| # Suppress Flask dev server warning — this is intentional for test purposes | |
| cli = logging.getLogger('werkzeug') | |
| # Suppress Flask dev server warning — this is intentional for test purposes | |
| cli = logging.getLogger('werkzeug') | |
| cli.setLevel(logging.ERROR) |
🤖 Prompt for AI Agents
In
`@tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/dispatch-service.yaml`
around lines 18 - 19, The code retrieves the Werkzeug logger into the variable
cli but never configures it, so the intended suppression is a no-op; update the
code that obtains the logger (cli = logging.getLogger('werkzeug')) to explicitly
set an appropriate level (e.g., cli.setLevel(logging.ERROR) or logging.WARNING)
and/or attach a NullHandler or desired handler to suppress the Flask dev server
warning (use cli.addHandler(logging.NullHandler()) or similar) so the logger is
actually silenced during tests.
…erations The HPA organic scaling loop (120 × 2s = 240s) consumed nearly the entire 300s setup timeout before the forced fallback could run. In Kind CI/CD, metrics-server is slow to provide CPU metrics, so the HPA rarely scales organically. Reducing to 30 iterations (60s) leaves ample time for the reliable forced kubectl scale fallback. https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok Signed-off-by: Claude <noreply@anthropic.com>
All three tests (ResourceQuota scaling blocked, Pending pods node pressure, OOMKill during scaling) pass consistently on latest models. Only keeping tests 215 and 216 which are hard enough to fail. https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok Signed-off-by: Claude <noreply@anthropic.com>
Summary
This PR adds five new test case fixtures for the
test_ask_holmessuite, covering common Kubernetes scaling and resource constraint failure scenarios. These fixtures enable testing of the Holmes LLM's ability to diagnose complex chain-of-causation issues in Kubernetes environments.Key Changes
New Test Cases Added
Test 212: Resource Quota Scaling Blocked
Test 213: Pending Pods Due to Node Pressure
Test 214: OOMKill During Scaling
Test 215: HPA Maxed Out
Implementation Details
Each test case includes:
test_case.yaml: User prompt, expected outputs, and setup/teardown scriptstoolsets.yaml: Enabled diagnostic tools (kubernetes/core, kubernetes/logs, prometheus/metrics for test 215)Test 215 includes a shared Prometheus deployment reference for metrics collection
All tests use realistic application scenarios (order-gateway, payment-processor, inventory-service, notification-service)
Setup scripts include detailed logging and error handling for debugging
Tags
All tests are tagged as
kubernetes,mediumcomplexity, andchain-of-causationto reflect their focus on multi-step failure diagnosis.https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok
Summary by CodeRabbit