Repository navigation
Improve evals - #709
Improve evals#709
Conversation
|
Warning Rate limit exceeded@aantn has exceeded the limit for the number of commits or files that can be reviewed per hour. Please wait 3 minutes and 5 seconds before requesting another review. ⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. 📒 Files selected for processing (8)
## Walkthrough
This update introduces a large set of new and revised Kubernetes troubleshooting test cases, manifests, and log generators for an LLM evaluation framework. It adds detailed runbook documentation for pod scheduling failures and replica mismatches, expands the pytest marker set, and updates or creates numerous YAML fixtures, manifests, and Python scripts to simulate realistic Kubernetes scenarios, resource constraints, and log patterns.
## Changes
| File(s) | Change Summary |
|---------------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------|
| CLAUDE.md, pyproject.toml | Updated test marker documentation: removed `k8s-misconfig` from docs; added `leaked-information` marker to pytest config. |
| holmes/plugins/runbooks/catalog.json | Added two new runbook entries for Kubernetes pod scheduling failures and replica mismatch troubleshooting. |
| holmes/plugins/runbooks/kubernetes/pod_scheduling_failures_instructions.md | Added comprehensive runbook for diagnosing and remediating Kubernetes pod scheduling failures. |
| holmes/plugins/runbooks/kubernetes/replica_mismatch_troubleshooting.md | Added detailed runbook for troubleshooting Kubernetes replica mismatch issues. |
| tests/llm/fixtures/test_ask_holmes/14_pending_resources/test_case.yaml | Added "easy" and "kubernetes" tags to test metadata. |
| tests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_new_tools/test_case.yaml | Simplified expected output to directly state DNS failure cause; removed step-by-step instructions. |
| tests/llm/fixtures/test_ask_holmes/47_truncated_logs_context_window/manifest.yaml | Reduced memory and CPU requests/limits for `long-logs-app` container. |
| tests/llm/fixtures/test_ask_holmes/50_logs_since_specific_date/test_case.yaml | Removed comments, added "easy" tag, updated correctness to 1. |
| tests/llm/fixtures/test_ask_holmes/51_logs_summarize_errors/{manifest.yaml,test_case.yaml} | Reduced resource requests/limits in manifest; added "easy" and "leaked-information" tags in test case; set correctness=1. |
| tests/llm/fixtures/test_ask_holmes/52_logs_login_issues/{manifest.yaml,test_case.yaml} | Reduced resource requests/limits; added "leaked-information" tag and comment in test case. |
| tests/llm/fixtures/test_ask_holmes/53_logs_find_term/{manifest.yaml,test_case.yaml} | Reduced resource requests/limits; added "easy" and "leaked-information" tags, updated correctness to 1. |
| tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/test_case.yaml | Swapped script names in setup/teardown; added "easy", "kubernetes", "context_window" tags. |
| tests/llm/fixtures/test_ask_holmes/57_wrong_namespace/manifest.yaml | Increased memory requests/limits, decreased CPU request, removed CPU limit. |
| tests/llm/fixtures/test_ask_holmes/60_count_less_than/{manifests.yaml,test_case.yaml} | Renamed pod metadata names; added "easy" tag to test case. |
| tests/llm/fixtures/test_ask_holmes/60_time_based_filtering/* | Deleted two test fixture files containing shell scripts for pod filtering/counting. |
| tests/llm/fixtures/test_ask_holmes/61_exact_match_counting/test_case.yaml | Added "easy" tag to evaluation section. |
| tests/llm/fixtures/test_ask_holmes/62_fetch_error_logs_with_errors/{manifest.yaml,test_case.yaml} | Increased memory, decreased CPU, removed CPU limit in manifest; added "leaked-information" tag and comment in test case. |
| tests/llm/fixtures/test_ask_holmes/63_fetch_error_logs_no_errors/{app.py,manifest.yaml,test_case.yaml} | Added Python log generator; switched manifest to Python app via secret volume; expanded setup/teardown; added "easy" tag. |
| tests/llm/fixtures/test_ask_holmes/64_keda_vs_hpa_confusion/manifest.yaml | Decreased CPU, increased memory requests/limits, removed CPU limit. |
| tests/llm/fixtures/test_ask_holmes/66_http_error_needle/{generate_logs.py,manifest.yaml,test_case.yaml} | Added log generator script, manifest for web-server pod, and new test case for error diagnosis. |
| tests/llm/fixtures/test_ask_holmes/67_performance_degradation/{generate_logs.py,manifest.yaml,test_case.yaml} | Added log generator, manifest, and test case simulating performance degradation. |
| tests/llm/fixtures/test_ask_holmes/68_cascading_failures/{generate_logs.py,manifest.yaml,test_case.yaml}| Added log generator, manifest, and test case for cascading failure scenario. |
| tests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/{generate_logs.py,manifest.yaml,test_case.yaml} | Added log generator, manifest, and test case for rate limit exhaustion scenario. |
| tests/llm/fixtures/test_ask_holmes/70_memory_leak_detection/{generate_logs.py,manifest.yaml,test_case.yaml} | Added log generator, manifest, and test case for memory leak detection scenario. |
| tests/llm/fixtures/test_ask_holmes/71_connection_pool_starvation/{generate_logs.py,manifest.yaml,test_case.yaml} | Added log generator, manifest, and test case for connection pool exhaustion. |
| tests/llm/fixtures/test_ask_holmes/73_time_window_anomaly/{generate_logs.py,manifest.yaml,test_case.yaml} | Added log generator, manifest, and test case for time-based anomaly detection. |
| tests/llm/fixtures/test_ask_holmes/74_config_change_impact/{generate_logs.py,manifest.yaml,test_case.yaml} | Added log generator, manifest, and test case for config change impact on cache hit rate. |
| tests/llm/fixtures/test_ask_holmes/75_network_flapping/{generate_logs.py,manifest.yaml,test_case.yaml} | Added log generator, manifest, and test case for network flapping scenario. |
| tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/{manifest.yaml,test_case.yaml} | Added manifest and test case for service selector mismatch scenario. |
| tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/{manifest.yaml,test_case.yaml} | Added manifest and test case for liveness probe misconfiguration scenario. |
| tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/{manifest.yaml,test_case.yaml} | Added manifest and test case for resource quota exceeded scenario. |
| tests/llm/fixtures/test_ask_holmes/79_configmap_mount_issue/{manifest.yaml,test_case.yaml} | Added manifest and test case for ConfigMap mount issue scenario. |
| tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/{manifest.yaml,test_case.yaml} | Added manifest and test case for PVC requesting non-existent storage class. |
| tests/llm/fixtures/test_ask_holmes/81_service_account_permission_denied/{manifest.yaml,test_case.yaml} | Added manifest and test case for service account permission denied scenario. |
| tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/{manifest.yaml,test_case.yaml} | Added manifest and test case for pod anti-affinity scheduling conflict. |
| tests/llm/fixtures/test_ask_holmes/83_secret_not_found/{manifest.yaml,test_case.yaml} | Added manifest and test case for missing secret in pod environment. |
| tests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/{manifest.yaml,test_case.yaml} | Added manifest and test case for network policy blocking frontend-backend traffic. |
| tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/{manifest.yaml,test_case.yaml} | Added manifest and test case for HPA not scaling due to missing resource requests. |
| tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/{manifest.yaml,test_case.yaml} | Added manifest and test case for missing secret mistaken as ConfigMap issue. |
| tests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/{manifest.yaml,test_case.yaml} | Added manifest and test case for taint/toleration scheduling issue with zone constraints. |
| tests/llm/fixtures/test_ask_holmes/02_what_is_wrong_with_pod/test_case.yaml | Updated pod name in prompt; added setup/teardown creating pod designed to trigger OOMKilled; no change to expected output. |
| tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml | Changed prompt to natural language; added setup/teardown to deploy grafana pod; refined expected output requirements. |
| tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/test_case.yaml | Changed prompt to natural language; generalized expected output summaries; added setup/teardown applying nginx deployment. |
| tests/llm/fixtures/test_ask_holmes/05_image_version/test_case.yaml | Changed pod name and image version in prompt and expected output; added setup/teardown deploying pod with specified image. |
| tests/llm/fixtures/test_ask_holmes/06_explain_issue/test_case.yaml | Added `mock_policy: always_mock` key to test case metadata. |
| tests/llm/fixtures/test_ask_holmes/27_permissions_error_no_helm_tools/test_case.yaml | Changed expected output format to numbered list; removed setup/teardown commands; replaced `tags` with `mock_policy`. |
| tests/llm/fixtures/test_ask_holmes/33_http_latency_graph/test_case.yaml | Changed manifest paths in setup/teardown; removed `reproducible` tag; added `skip: true` with explanation. |
| tests/llm/fixtures/test_ask_holmes/48_logs_since_thursday/test_case.yaml | Added "easy" tag. |
| tests/llm/fixtures/test_ask_holmes/55_kafka_runbook/test_case.yaml | Replaced `reproducible` tag with `hard` tag and added `skip: true` with reason about port-forward dependency. |
| tests/llm/fixtures/test_ask_holmes/56_kafka_runbook_no_tool/test_case.yaml | Replaced `reproducible` tag with `easy`. |
| tests/llm/fixtures/test_ask_holmes/01_how_many_pods/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/07_high_latency/test_case.yaml | Removed `reproducible` tag and emptied tags list. |
| tests/llm/fixtures/test_ask_holmes/08_sock_shop_frontend/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/09_crashpod/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/10_image_pull_backoff/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/11_init_containers/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/12_job_crashing/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/13_pending_node_selector/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/15_failed_readiness_probe/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/17_oom_kill/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/18_crash_looping_v2/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/20_long_log_file_search/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/21_job_fail_curl_no_svc_account/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/22_high_latency_dbi_down/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/23_app_error_in_current_logs/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/24_misconfigured_pvc/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/25_misconfigured_ingress_class/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/28_permissions_error_helm_tools_enabled/test_case.yaml | Removed trailing empty line and `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/38_rabbitmq_split_head/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_all_tools/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_new_tools/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_new_tools_no_runbook/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_old_tools/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_new_all_tools/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_old_tools/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/43_slack_deployment_logs/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/45_fetch_deployment_logs_simple/test_case.yaml | Removed `tags` section containing `reproducible`. |
| tests/llm/fixtures/test_ask_holmes/46_job_crashing_no_longer_exists/test_case.yaml | Removed `reproducible` tag; test remains skipped. |
| tests/llm/fixtures/test_ask_holmes/57_wrong_namespace/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/58_counting_pods_by_status/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/59_label_based_counting/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/64_keda_vs_hpa_confusion/test_case.yaml | Removed `reproducible` tag. |
| tests/llm/fixtures/test_ask_holmes/65_health_check_followup/test_case.yaml | Removed extra empty line and `reproducible` tag, leaving only `understanding-real-intent`. |
| .github/workflows/llm-evaluation.yaml | Removed pytest marker filter `-m 'not skip'`, running all tests without marker exclusion. |
## Sequence Diagram(s)
```mermaid
sequenceDiagram
participant User
participant TestRunner
participant KubernetesCluster
participant HolmesLLM
participant RunbookDocs
User->>TestRunner: Trigger test case (e.g., pod scheduling failure)
TestRunner->>KubernetesCluster: Apply manifest/setup resources
KubernetesCluster-->>TestRunner: Pods/services/events/logs
TestRunner->>HolmesLLM: Provide scenario context and cluster state
HolmesLLM->>RunbookDocs: Reference runbook for diagnosis/remediation
RunbookDocs-->>HolmesLLM: Steps, root cause analysis, remediation
HolmesLLM-->>TestRunner: Diagnosis and remediation suggestions
TestRunner-->>User: Present results and evaluationEstimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes
Possibly related PRs
Suggested labels
Suggested reviewers
|
There was a problem hiding this comment.
Actionable comments posted: 8
🔭 Outside diff range comments (3)
tests/llm/fixtures/test_ask_holmes/57_wrong_namespace/manifest.yaml (2)
40-45: Consider restoring a sensible CPU limit to avoid noisy-neighbor problemsDropping the
limits.cpufield means the container can burst without bound, which is fine for some synthetic tests but risky for shared clusters and can give misleading scheduling behaviour. A modest limit (e.g. equal to the request or 2-3 × it) keeps the scenario realistic while still exercising low-CPU scheduling.limits: memory: "64Mi" + cpu: "20m" # keep within a single core while allowing short bursts
21-38: Harden the container with a minimal securityContextStatic-analysis flags (
CKV_K8S_20,CKV_K8S_23) indicate the container can run as root and escalate privileges. Even in test fixtures, adding an explicitsecurityContextimproves fidelity and avoids teaching insecure defaults:spec: containers: - name: app-container image: busybox:1.35 + securityContext: + runAsNonRoot: true + allowPrivilegeEscalation: false + readOnlyRootFilesystem: truetests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_new_tools/test_case.yaml (1)
15-17:expected_scoreshould be1now that a concreteexpected_outputis providedWith a non-empty
expected_output, leavingexpected_score: 0means the test will mark every answer incorrect irrespective of matching content.- expected_score: 0 + expected_score: 1
♻️ Duplicate comments (1)
tests/llm/fixtures/test_ask_holmes/53_logs_find_term/manifest.yaml (1)
38-42: Repeat note on absent CPU limitSame recommendation as earlier: confirm the cluster doesn’t mandate
cpulimits, or add one.
🧹 Nitpick comments (26)
tests/llm/fixtures/test_ask_holmes/47_truncated_logs_context_window/manifest.yaml (1)
34-40: Extremely low resources may cause OOM / throttling in log generator
long-logs-appis configured to spit out 5 M tokens every 100 ms yet is allotted only 64 Mi memory and 10 m CPU. In practice the container will thrash or be OOM-killed, which may mask the log-truncation behaviour you want to test.Consider bumping requests/limits to a more realistic floor (e.g. 256 Mi / 100 m) or add a comment explaining the intentional constraint.
tests/llm/fixtures/test_ask_holmes/51_logs_summarize_errors/manifest.yaml (1)
36-42: Same resource pattern as file 52 – ensure global consistency with any cluster LimitRangesSee earlier comment about missing CPU limits. If a cluster-wide
LimitRangeenforces both requests and limits, this deployment will be rejected.If intentional, add a brief comment in the YAML so future readers know the absence of
cpulimit is deliberate.tests/llm/fixtures/test_ask_holmes/64_keda_vs_hpa_confusion/manifest.yaml (1)
29-35: HPA + no CPU limit can distort utilisation calculationsHPA bases
% Utilizationon requested CPU. Withrequest: 10mand no limit, the pod can burst well above 10 m, quickly reading as 1000 % and scaling up aggressively.
If the purpose of the scenario isn’t to test that edge case, consider settinglimit: 50mor raising the request to something closer to expected steady-state to avoid runaway scaling.tests/llm/fixtures/test_ask_holmes/53_logs_find_term/manifest.yaml (1)
24-26: Spurious blank line introduces visual noiseLine 25 is an empty line between
name:andimage:. YAML permits it, but removing keeps the manifest tidy and consistent with the other fixtures.- name: main-container - image: us-central1-docker.pkg.dev/genuine-flight-317411/devel/multiple-errors-in-logs:v1tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/manifest.yaml (1)
30-37: Add an explanatory comment for the intentionally-missing SecretThe volume references
frontend-api-keys, which is not defined in this manifest (onlybackend-api-keysis). A short YAML comment will prevent future readers from “fixing” the intentional gap and breaking the test logic.- volumes: - - name: api-credentials - secret: - secretName: frontend-api-keys - optional: false + volumes: + - name: api-credentials + # NOTE: This secret is *deliberately* absent – the test expects pods to fail. + secret: + secretName: frontend-api-keys + optional: falseholmes/plugins/runbooks/kubernetes/pod_scheduling_failures_instructions.md (2)
69-70: Escape literal dots in grep patternUnescaped
.acts as “any character”, sogrep -E "(cloud.google.com|eks.amazonaws.com|kubernetes.azure.com)"may match unintended labels. Escaping improves precision.-grep -E "(cloud.google.com|eks.amazonaws.com|kubernetes.azure.com)" +grep -E "(cloud\.google\.com|eks\.amazonaws\.com|kubernetes\.azure\.com)"
102-117: Specify a language for the fenced blockMarkdown linters (MD040) complain about language-less fences. Since this is plain text, set it to
textfor consistency.-``` +```texttests/llm/fixtures/test_ask_holmes/53_logs_find_term/test_case.yaml (1)
5-15: Typo in expected output and changedcorrectnessscore may invalidate the test
- Line 11: “volation” ➜ “violation”.
evaluation.correctnesswas bumped from 0 to 1. Make sure the grading harness now treats the provided expected output as definitive; otherwise previously failing solutions will erroneously pass.- 2. DB query issues due unicity constraint volation + 2. DB query issues due to uniqueness-constraint violationtests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/test_case.yaml (1)
6-6: YAML bullet formatting: indent two spaces for consistency.Minor nit: use two-space indent for list items under
expected_outputto match surrounding fixtures.tests/llm/fixtures/test_ask_holmes/62_fetch_error_logs_with_errors/manifest.yaml (2)
41-45: CPU limit removed – evaluate whether unrestricted CPU is intentional.Requests dropped from 50 m → 10 m and the limit was removed. On burst-heavy clusters this container could monopolise CPU and skew tests. If unbounded CPU is not required, keep a modest limit (e.g.,
100m) to stay predictable.
40-44: Static analysis flags root/priv-escalation; consider tightening securityContext.Although not changed in this diff, Checkov flagged CKV_K8S_20/23 for this container. Adding:
+ securityContext: + runAsNonRoot: true + allowPrivilegeEscalation: falsewould silence the warnings without impacting the scenario.
holmes/plugins/runbooks/catalog.json (1)
14-22: Consider grouping new entries by domain for quicker lookupPlacing the two Kubernetes runbooks next to each other (or under a dedicated
"kubernetes"subsection) keeps the catalog logically grouped and improves discoverability as the list grows. No functional change, purely an ordering tweak.tests/llm/fixtures/test_ask_holmes/63_fetch_error_logs_no_errors/test_case.yaml (1)
3-9: Minor shell-quoting & idempotency improvements in the setup/teardown
- Quote
./manifest.yamlto avoid glob expansion accidents.- Prefix the teardown with
set -eso failures surface early (e.g., if the secret creation fails).-before_test: | - kubectl create secret generic app-code-63 -n staging-63 --from-file=app.py=./app.py --dry-run=client -o yaml | kubectl apply -f - - kubectl apply -f ./manifest.yaml +before_test: | + set -e + kubectl create secret generic app-code-63 -n staging-63 \ + --from-file=app.py=./app.py \ + --dry-run=client -o yaml | kubectl apply -f - + kubectl apply -f "./manifest.yaml" @@ -after_test: | - kubectl delete -f ./manifest.yaml - kubectl delete secret app-code-63 -n staging-63 --ignore-not-found +after_test: | + set -e + kubectl delete -f "./manifest.yaml" + kubectl delete secret app-code-63 -n staging-63 --ignore-not-foundtests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/test_case.yaml (1)
12-14: Alignevaluationschema with other test cases for consistencyElsewhere the file nests
correctnessoptions (e.g.,expected_score,type). Consider adopting the same structure for easier tooling/parsing.-evaluation: - correctness: 1 +evaluation: + correctness: + expected_score: 1 + type: "strict"tests/llm/fixtures/test_ask_holmes/83_secret_not_found/manifest.yaml (1)
33-53: Consider adding security best practices for completeness.While this is a test fixture, consider adding security best practices to make it more realistic:
spec: + securityContext: + runAsNonRoot: true + runAsUser: 999 containers: - name: database image: postgres:alpine + securityContext: + allowPrivilegeEscalation: false + readOnlyRootFilesystem: false + runAsNonRoot: truetests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/manifest.yaml (1)
36-45: Consider adding security best practices for completeness.While this is a test fixture, adding security configurations would make it more realistic:
spec: + securityContext: + runAsNonRoot: true + runAsUser: 101 containers: - name: api-server image: nginx:alpine + securityContext: + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true + runAsNonRoot: truetests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/manifest.yaml (2)
22-29: MissingnamespaceSelectoron anti-affinity may catch unrelated podsIf another team deploys
app: cache-serverpods outsidenamespace-82, they will also block scheduling here. AddnamespaceSelector:to confine the rule to the current namespace unless cross-namespace blocking is intended.- - labelSelector: + - namespaceSelector: + matchNames: + - namespace-82 + labelSelector:
30-38: No container security context – fails common CIS & Checkov controls
allowPrivilegeEscalationdefaults to true. Even in test fixtures it is a good habit to lock this down to avoid noisy scanner findings.- name: cache-server image: redis:alpine + securityContext: + allowPrivilegeEscalation: false resources:tests/llm/fixtures/test_ask_holmes/87_resource_like_but_ephemeral_storage/manifest.yaml (2)
26-31:ephemeral-storagerequest without a matching limitKubernetes only enforces quota & eviction on the higher of
requestsandlimits. Omittinglimitshere means a runaway container can consume more than 15 Gi and mask the scheduling failure the test intends to surface. Add an identical limit to keep the signal clean.ephemeral-storage: "15Gi" + limits: + ephemeral-storage: "15Gi"
50-54: Repeat the storage request pattern for consistency
api-serversets a 10 Gi request but again no limit. Mirror the earlier advice for consistency and to avoid noisy evictions during test runs.tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/manifest.yaml (1)
20-28: Intentional storage-class typo – add an explanatory commentFuture maintainers may “fix” the obvious typo and invalidate the test. Add a clarifying comment to lock the behaviour in place.
- storageClassName: fast-ssd # This storage class doesn't exist! + # Deliberately incorrect to trigger Pending PVC in the test. + storageClassName: fast-ssdtests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/test_case.yaml (2)
68-76: Minor mismatch in taint string – may trip brittle parsers
kubectl_describe_pod_database_pendingshows theFailedSchedulingevent with
{node-role.kubernetes.io/control-plane: }(missingNoSchedule) whereas the
actual taint on the nodes (line 95) isnode-role.kubernetes.io/control-plane:NoSchedule.If downstream assertion logic extracts the full taint key + effect from the
event text, this difference can cause false-negatives.
Recommend making the event snippet consistent with the real taint string.
70-72: Topology spread constraint wording deviates from kubectl output
topology.kubernetes.io/zone:ScheduleAnyway when max skew 1 is exceeded …
is not howkubectl describe podprints the constraint (it emits
topologyKey=topology.kubernetes.io/zone, whenUnsatisfiable=ScheduleAnyway).Keeping the mock closer to the real CLI output helps future-proof the test.
tests/llm/fixtures/test_ask_holmes/87_resource_like_but_ephemeral_storage/test_case.yaml (2)
60-68: Event reason good, but consider echoing exact scheduler phrasingThe mock uses
0/4 nodes are available: 4 Insufficient ephemeral-storage.Recent Kubernetes versions output
0/4 nodes had insufficient ephemeral-storage.Aligning the string avoids fragile substring checks in evaluation code.
78-87:kubectl top nodesoutput lacks storage columnsBecause
topomits ephemeral-storage, you already rely ondescribe nodes
for capacity/allocatable. Consider omittingkubectl top nodesfrom this
fixture (or add a comment that it’s only for CPU/memory) to avoid implying
thattopwould reveal the storage shortage.holmes/plugins/runbooks/kubernetes/replica_mismatch_troubleshooting.md (1)
80-81: Minor: command comment repeats “events”The comment
# Get events for specific pods (more targeted)appears twice in
Step 1 (lines 8 and 11). Removing the duplicate tidies the doc.
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (42)
CLAUDE.md(0 hunks)holmes/plugins/runbooks/catalog.json(1 hunks)holmes/plugins/runbooks/kubernetes/pod_scheduling_failures_instructions.md(1 hunks)holmes/plugins/runbooks/kubernetes/replica_mismatch_troubleshooting.md(1 hunks)pyproject.toml(1 hunks)tests/llm/fixtures/test_ask_holmes/14_pending_resources/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_new_tools/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/47_truncated_logs_context_window/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/50_logs_since_specific_date/test_case.yaml(2 hunks)tests/llm/fixtures/test_ask_holmes/51_logs_summarize_errors/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/51_logs_summarize_errors/test_case.yaml(2 hunks)tests/llm/fixtures/test_ask_holmes/52_logs_login_issues/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/52_logs_login_issues/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/53_logs_find_term/manifest.yaml(2 hunks)tests/llm/fixtures/test_ask_holmes/53_logs_find_term/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/57_wrong_namespace/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/60_count_less_than/manifests.yaml(4 hunks)tests/llm/fixtures/test_ask_holmes/60_count_less_than/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/60_time_based_filtering/kubernetes_countitems_select_.metadata.namespace_test-60_and_.status.containerStatuses_.restartCount_3_.metadata.name_Pod.txt(0 hunks)tests/llm/fixtures/test_ask_holmes/60_time_based_filtering/kubernetes_countitems_select_.metadata.namespace_test-60_and_.status.startTime_fromdateiso8601_now_-_45_.metadata.name_pod.txt(0 hunks)tests/llm/fixtures/test_ask_holmes/61_exact_match_counting/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/62_fetch_error_logs_with_errors/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/62_fetch_error_logs_with_errors/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/63_fetch_error_logs_no_errors/app.py(1 hunks)tests/llm/fixtures/test_ask_holmes/63_fetch_error_logs_no_errors/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/63_fetch_error_logs_no_errors/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/64_keda_vs_hpa_confusion/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/83_secret_not_found/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/83_secret_not_found/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/87_resource_like_but_ephemeral_storage/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/87_resource_like_but_ephemeral_storage/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/test_case.yaml(1 hunks)
💤 Files with no reviewable changes (3)
- CLAUDE.md
- tests/llm/fixtures/test_ask_holmes/60_time_based_filtering/kubernetes_countitems_select_.metadata.namespace_test-60_and_.status.containerStatuses_.restartCount_3_.metadata.name_Pod.txt
- tests/llm/fixtures/test_ask_holmes/60_time_based_filtering/kubernetes_countitems_select_.metadata.namespace_test-60_and_.status.startTime_fromdateiso8601_now_-45.metadata.name_pod.txt
🧰 Additional context used
🪛 Checkov (3.2.334)
tests/llm/fixtures/test_ask_holmes/57_wrong_namespace/manifest.yaml
[MEDIUM] 6-44: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-44: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/63_fetch_error_logs_no_errors/manifest.yaml
[MEDIUM] 6-38: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-38: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/62_fetch_error_logs_with_errors/manifest.yaml
[MEDIUM] 6-44: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-44: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/manifest.yaml
[MEDIUM] 7-38: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 7-38: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/87_resource_like_but_ephemeral_storage/manifest.yaml
[MEDIUM] 6-31: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-31: Minimize the admission of root containers
(CKV_K8S_23)
[MEDIUM] 32-55: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 32-55: Minimize the admission of root containers
(CKV_K8S_23)
[MEDIUM] 56-78: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 56-78: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/manifest.yaml
[MEDIUM] 19-65: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 19-65: Minimize the admission of root containers
(CKV_K8S_23)
[MEDIUM] 66-88: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 66-88: Minimize the admission of root containers
(CKV_K8S_23)
[MEDIUM] 89-110: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 89-110: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/manifest.yaml
[MEDIUM] 6-38: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-38: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/manifest.yaml
[MEDIUM] 21-45: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 21-45: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/83_secret_not_found/manifest.yaml
[MEDIUM] 18-53: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 18-53: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/manifest.yaml
[MEDIUM] 21-55: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 21-55: Minimize the admission of root containers
(CKV_K8S_23)
🪛 Gitleaks (8.27.2)
tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/manifest.yaml
46-46: Detected a Generic API Key, potentially exposing access to various services and sensitive operations.
(generic-api-key)
48-48: Detected a Generic API Key, potentially exposing access to various services and sensitive operations.
(generic-api-key)
58-58: Detected a Generic API Key, potentially exposing access to various services and sensitive operations.
(generic-api-key)
40-46: Possible Kubernetes Secret detected, posing a risk of leaking credentials/tokens from your deployments
(kubernetes-secret-yaml)
51-58: Possible Kubernetes Secret detected, posing a risk of leaking credentials/tokens from your deployments
(kubernetes-secret-yaml)
61-67: Possible Kubernetes Secret detected, posing a risk of leaking credentials/tokens from your deployments
(kubernetes-secret-yaml)
tests/llm/fixtures/test_ask_holmes/83_secret_not_found/manifest.yaml
15-15: Detected a Generic API Key, potentially exposing access to various services and sensitive operations.
(generic-api-key)
8-15: Possible Kubernetes Secret detected, posing a risk of leaking credentials/tokens from your deployments
(kubernetes-secret-yaml)
🪛 markdownlint-cli2 (0.17.2)
holmes/plugins/runbooks/kubernetes/pod_scheduling_failures_instructions.md
111-111: Trailing punctuation in heading
Punctuation: ':'
(MD026, no-trailing-punctuation)
112-112: Fenced code blocks should have a language specified
(MD040, fenced-code-language)
128-128: Trailing punctuation in heading
Punctuation: ':'
(MD026, no-trailing-punctuation)
135-135: Trailing punctuation in heading
Punctuation: ':'
(MD026, no-trailing-punctuation)
164-164: Emphasis used instead of a heading
(MD036, no-emphasis-as-heading)
182-182: Emphasis used instead of a heading
(MD036, no-emphasis-as-heading)
200-200: Emphasis used instead of a heading
(MD036, no-emphasis-as-heading)
219-219: Emphasis used instead of a heading
(MD036, no-emphasis-as-heading)
238-238: Emphasis used instead of a heading
(MD036, no-emphasis-as-heading)
256-256: Emphasis used instead of a heading
(MD036, no-emphasis-as-heading)
275-275: Emphasis used instead of a heading
(MD036, no-emphasis-as-heading)
295-295: Emphasis used instead of a heading
(MD036, no-emphasis-as-heading)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (1)
- GitHub Check: Pre-commit checks
🔇 Additional comments (21)
tests/llm/fixtures/test_ask_holmes/52_logs_login_issues/manifest.yaml (1)
34-40: Verify or Add a CPU Limit in the ManifestA missing CPU
limitsentry may trigger pod rejections on clusters with aLimitRangepolicy. I didn’t find anyLimitRangedefinitions in the repo, so please confirm your test-cluster defaults or explicitly set a small CPU limit.• File:
tests/llm/fixtures/test_ask_holmes/52_logs_login_issues/manifest.yaml
Lines: 34–40Suggested diff:
resources: requests: memory: "64Mi" cpu: "10m" limits: memory: "64Mi" + cpu: "50m"tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/manifest.yaml (1)
40-58: Double-check that sample secrets are non-production placeholdersStatic scanners flagged these base64 blobs as “generic API keys”. If they are synthetic (they decode to
api-key-value,dbpassword, etc.) leave them; otherwise replace them with obvious dummy strings (e.g.dW5zZWN1cmUtcGxhY2Vob2xkZXI=).tests/llm/fixtures/test_ask_holmes/60_count_less_than/test_case.yaml (1)
14-14: Tag addition LGTMAdding the
easytag improves test filtering without changing semantics.tests/llm/fixtures/test_ask_holmes/61_exact_match_counting/test_case.yaml (1)
15-15: Tag addition LGTMSame rationale as above – no issues spotted.
tests/llm/fixtures/test_ask_holmes/62_fetch_error_logs_with_errors/test_case.yaml (1)
25-26:leaked-informationtag & inline YAML comment – confirm parser & marker support
- The inline comment (
# Hardcoded …) sits inside the sequence. YAML parsers are permissive, but some linters (e.g. yamllintseq-spaces) flag this placement.- Ensure the new
leaked-informationmarker is defined inpytestconfiguration; otherwise marker-unknown warnings appear.tests/llm/fixtures/test_ask_holmes/52_logs_login_issues/test_case.yaml (1)
5-6: Confirm consistency betweenleaked-informationtag and remaining evaluation metadataThe new marker signals that the manifest discloses the root cause, yet
evaluation.correctnessis still 0. Double-check that this is intentional; other tests transitioned to 1 after similar edits.tests/llm/fixtures/test_ask_holmes/50_logs_since_specific_date/test_case.yaml (1)
5-5: Updated toeasyandcorrectness: 1— ensure alignment with grading logicChanging both the difficulty tag and correctness flag alters scoring. Verify upstream reporting dashboards and any difficulty-based weighting use the new values.
Also applies to: 14-14
tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/test_case.yaml (2)
12-14: Tag additions look good – confirm harness awareness of the new markers.The extra tags improve filtering, but ensure the test-runner’s tag registry (pytest markers / metadata parser) recognises
easy,kubernetes, andcontext_window; otherwise they will be silently ignored.
5-6: Incorrect namespace mismatch – no changes needed
Thewait_for_replicas.shcalls target the Deployment inask-holmes-namespace-54a(where it’s defined), while theuser_promptand subsequentkubectlsteps run inask-holmes-namespace-54bfor the fast-fail pod. Both namespaces are declared in the manifest, so the fixture is correct as-is.Likely an incorrect or invalid review comment.
tests/llm/fixtures/test_ask_holmes/51_logs_summarize_errors/test_case.yaml (2)
5-7: New “leaked-information” marker may need code-base registration.Adding the marker is fine, but pytest will raise
PytestUnknownMarkWarningunless it is declared inpytest.ini/pyproject.toml. Please confirm it was added there alongside the other markers.
17-17: Correctness bumped to 1 – double-check expected_output still matches implementation.Switching the grading implies the LLM output is now considered right. Make sure CI has a failing test that now passes; otherwise the change masks a real issue.
tests/llm/fixtures/test_ask_holmes/60_count_less_than/manifests.yaml (1)
83-96: No lingeringflaky-references found in fixtures
I ran a global search acrosstests/llm/fixturesand didn’t find any remaining “flaky-” pod names. Test expectations and helper scripts appear to have been updated correctly—no further changes needed.tests/llm/fixtures/test_ask_holmes/83_secret_not_found/test_case.yaml (1)
1-14: Well-structured test case for Kubernetes secret troubleshooting.The test case is correctly configured to simulate a database pod failure due to a missing secret reference. The setup, expected output, and teardown are appropriately defined.
tests/llm/fixtures/test_ask_holmes/63_fetch_error_logs_no_errors/app.py (1)
1-42: Well-implemented application simulator for log testing.The Python script effectively simulates a realistic application that generates various log levels without errors. The code structure is clean with proper timestamping, appropriate use of
flush=Truefor container environments, and realistic timing patterns.tests/llm/fixtures/test_ask_holmes/83_secret_not_found/manifest.yaml (1)
1-54: Effective test manifest for secret resolution failure scenario.The manifest correctly simulates a realistic Kubernetes secret resolution failure by creating a secret with one name (
db-credentials) while the deployment references a different secret name (database-secret).tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/test_case.yaml (1)
1-15: Well-designed test case for PVC storage class troubleshooting.The test case effectively covers both problem identification and solution recommendation for PVC storage class mismatches. The expected output appropriately includes suggesting alternative storage classes, which demonstrates comprehensive troubleshooting capabilities.
tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/manifest.yaml (1)
1-46: Effective resource quota violation test scenario.The manifest correctly demonstrates resource quota enforcement by creating a deployment that exceeds the namespace's CPU quota (90m requested vs 20m allowed). The resource configuration is well-structured with matching requests and limits.
tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/manifest.yaml (1)
12-18: Anti-affinity scenario may pass on large clustersThe test relies on having fewer than 30 schedulable nodes, otherwise all pods will schedule and the scenario won’t reproduce. To make the failure deterministic, consider constraining the Deployment with an additional
nodeSelector/topologySpreadConstraintsor loweringreplicastonode_count + 1during setup.tests/llm/fixtures/test_ask_holmes/87_resource_like_but_ephemeral_storage/manifest.yaml (1)
70-78: Front-end pod may schedule when nodes free <8 GiThe test’s intent is “cache-server pods fail, others succeed”. If a node has only 7 Gi free the front-end pod will also stay Pending, muddying the scenario. Consider lowering its request or adding a comment explaining the expected node capacity.
tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/test_case.yaml (1)
148-157: Test expectations are clear and accurateReplica count, error reason, and remediation guidance line up with the mocked kubectl output. No changes needed.
holmes/plugins/runbooks/kubernetes/replica_mismatch_troubleshooting.md (1)
59-60:greppattern won’t OR-match due to quoted pipeInside double quotes the
|is taken literally.
grep -A10 "Liveness\|Readiness"therefore matches the actual string
Liveness|Readiness, not either word.-grep -A10 "Liveness\|Readiness" +grep -A10 -E 'Liveness|Readiness'Apply the same fix for the other
grep … "Volumes\|Mounts"pattern.Likely an incorrect or invalid review comment.
There was a problem hiding this comment.
Actionable comments posted: 7
♻️ Duplicate comments (3)
tests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/manifest.yaml (3)
11-17: RWO PVC + 4 replicas ⇒ inevitableMulti-Attachfailure
ReadWriteOncevolumes can only be mounted by one node at a time. Withreplicas: 4, the first pod will bind the claim and the remaining three will stay Pending withMulti-Attacherrors, masking the taint scenario this fixture is supposed to test.- replicas: 4 + # Option A – simplest: keep a single writer + replicas: 1Other options: switch the claim to an RWX-capable storageClass, or use StatefulSet so each replica gets its own PVC.
Also applies to: 25-25
34-40: Zone-level topology spread conflicts with single-AZ PVCBecause the PVC will bind to a single-AZ EBS volume,
topology.kubernetes.io/zonespreading inevitably breaks (pods in other zones cannot mount the volume). Drop the constraint or scope it to the same zone label you know the volume lands in.- topologySpreadConstraints: - - maxSkew: 1 - topologyKey: topology.kubernetes.io/zone - whenUnsatisfiable: ScheduleAnyway - labelSelector: - matchLabels: - app: database-primary +# (remove or adapt this block)
60-64: Missing toleration fordedicated=database:NoScheduletaintThe comment already notes it; without the toleration the pods will be rejected before we hit any affinity logic.
volumes: - name: postgres-storage persistentVolumeClaim: claimName: postgres-pvc-us-east-1a + tolerations: + - key: "dedicated" + operator: "Equal" + value: "database" + effect: "NoSchedule"
🧹 Nitpick comments (33)
tests/llm/fixtures/test_ask_holmes/71_connection_pool_starvation/manifest.yaml (1)
6-40: Address security concerns in container configuration.The deployment configuration is functional but has security implications flagged by static analysis:
- Container runs as root by default - Consider adding a non-root user specification
- No explicit privilege escalation prevention - Should explicitly disable privilege escalation
For a test fixture, these may be acceptable, but consider adding security context for better practices:
spec: + securityContext: + runAsNonRoot: true + runAsUser: 1000 containers: - name: backend-service image: python:3.9-slim + securityContext: + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true command: ["python"]tests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/manifest.yaml (1)
23-40: Consider adding security context for better container security practices.While this is a test fixture, adding a security context would demonstrate better practices and address the static analysis security concerns about privilege escalation and root containers.
Consider adding a security context to the container spec:
containers: - name: api-limiter image: python:3.9-slim command: ["python"] args: ["/scripts/generate_logs.py"] + securityContext: + allowPrivilegeEscalation: false + runAsNonRoot: true + runAsUser: 1000 + readOnlyRootFilesystem: true volumeMounts: - name: script-volume mountPath: /scriptstests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/generate_logs.py (2)
31-31: Address unused loop variables as flagged by static analysis.The loop control variables are not used within the loop bodies. Following the linter suggestion will improve code clarity.
Apply these changes to rename unused loop variables:
- for i in range(80000): + for _ in range(80000): print( generate_log_entry( status=200, req_per_sec=random.randint(8, 12), message="Request processed successfully", ) ) - for i in range(60000): + for _ in range(60000): print( generate_log_entry( status=200, req_per_sec=1000, client_ip=spike_client, message="Request processed (high rate)", ) ) - for i in range(5000): + for _ in range(5000): print( generate_log_entry( status=429, req_per_sec=1000, client_ip=spike_client, message="Too Many Requests - Rate limit exceeded (100/second)", ) ) - for i in range(5000): + for _ in range(5000): print( generate_log_entry( status=200, req_per_sec=random.randint(8, 12), message="Request processed successfully", ) )Also applies to: 59-59, 70-70, 81-81
91-92: Consider adding graceful shutdown handling.While the infinite sleep keeps the container running as intended, adding signal handling would allow for cleaner container shutdown during testing.
Consider adding signal handling for graceful shutdown:
+import signal +import sys + +def signal_handler(sig, frame): + print('Shutting down log generator...') + sys.exit(0) def main(): # ... existing log generation code ... # Keep pod running + signal.signal(signal.SIGINT, signal_handler) + signal.signal(signal.SIGTERM, signal_handler) while True: time.sleep(3600)tests/llm/fixtures/test_ask_holmes/75_network_flapping/manifest.yaml (2)
22-40: Consider adding security context for better security posture.While this is a test fixture, consider adding a security context to follow security best practices:
spec: + securityContext: + runAsNonRoot: true + runAsUser: 1000 + fsGroup: 1000 containers: - name: frontend image: python:3.9-slim + securityContext: + allowPrivilegeEscalation: false + readOnlyRootFilesystem: false + capabilities: + drop: + - ALLThis addresses the static analysis concerns about privilege escalation and root containers.
31-34: Consider adding resource limits for better resource management.While the resource requests are appropriate for a test scenario, adding limits helps prevent resource exhaustion:
resources: requests: memory: "64Mi" cpu: "10m" + limits: + memory: "128Mi" + cpu: "50m"tests/llm/fixtures/test_ask_holmes/75_network_flapping/generate_logs.py (1)
154-174: Fix unused loop variable in Phase 4.The loop variable
iis declared but not used within the loop body. Since Phase 4 uses random probability instead of index-based conditions, the variable is unnecessary.- for i in range(25000): + for _ in range(25000): total_requests += 1 if random.random() < 0.6: # 60% failure ratetests/llm/fixtures/test_ask_holmes/67_performance_degradation/manifest.yaml (1)
23-40: Consider adding security context for production-like testing.The container specification lacks a security context, which means it runs as root by default. While this is acceptable for a test environment, consider adding security hardening to make the test more realistic:
containers: - name: api-gateway image: python:3.9-slim command: ["python"] args: ["/scripts/generate_logs.py"] + securityContext: + runAsNonRoot: true + runAsUser: 1001 + allowPrivilegeEscalation: false volumeMounts:tests/llm/fixtures/test_ask_holmes/66_http_error_needle/generate_logs.py (1)
26-50: Fix unused loop variables as indicated by static analysis.The main function logic is correct for the testing scenario, but the loop variables should be renamed to indicate they're intentionally unused.
Apply this diff to address the static analysis hints:
- for i in range(50000): + for _ in range(50000): print(generate_log_entry()) # The critical error in the middle print( generate_log_entry( status=500, message="Database connection timeout: Could not acquire connection from pool after 30s", path="/api/orders", ) ) # Another 50,000 successful requests - for i in range(50000): + for _ in range(50000): print(generate_log_entry())tests/llm/fixtures/test_ask_holmes/66_http_error_needle/manifest.yaml (1)
1-41: Consider adding security hardening for best practices.The manifest structure is correct and appropriate for a test environment. However, consider adding security best practices even in testing scenarios.
Apply this diff to improve security posture:
containers: - name: web-server image: python:3.9-slim command: ["python"] args: ["/scripts/generate_logs.py"] + securityContext: + allowPrivilegeEscalation: false + runAsNonRoot: true + runAsUser: 1000 + readOnlyRootFilesystem: true volumeMounts: - name: script-volume mountPath: /scripts resources: requests: memory: "64Mi" cpu: "10m"Note: If
readOnlyRootFilesystem: truecauses issues with the Python runtime, you may need to add a temporary volume mount for/tmp.tests/llm/fixtures/test_ask_holmes/81_service_account_permission_denied/manifest.yaml (2)
60-78: Harden the test container with a minimalsecurityContextStatic analysis flags (CKV_K8S_20, CKV_K8S_23) point out that the pod may run as root and with
allowPrivilegeEscalationenabled. Even for fixtures it’s trivial to mitigate:containers: - name: monitoring-agent image: bitnami/kubectl:latest + securityContext: + allowPrivilegeEscalation: false + runAsNonRoot: true + readOnlyRootFilesystem: trueThis keeps the fixture functionally identical while demonstrating best practices.
64-64: Pin thebitnami/kubectlimage to a specific tag/digestUsing
:latestmakes the test nondeterministic—future upstream image changes could break or slow the CI run. Recommend referencing an immutable tag (e.g.,bitnami/kubectl:1.30.1) or digest.tests/llm/fixtures/test_ask_holmes/73_time_window_anomaly/generate_logs.py (2)
39-39: Simplify the time window condition.The condition can be simplified for better readability.
- if hour == 3 and minute >= 0 and minute <= 5: + if hour == 3 and minute <= 5:Since
minute >= 0is always true, it can be omitted.
93-94: Consider making the randomization configurable.While the randomization adds realism, consider making it configurable through environment variables for more predictable testing scenarios.
# Add at the top of main(): randomize_timing = os.getenv('RANDOMIZE_TIMING', 'true').lower() == 'true' # Then modify the randomization block: if randomize_timing and random.random() < 0.1: current_time += timedelta(seconds=random.randint(1, 5))tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/manifest.yaml (2)
24-35: Harden the backend container: drop privileges & set CPU limitThe static-analysis hints (CKV_K8S_20 / CKV_K8S_23) flag the lack of basic securityContext settings, and the spec defines a CPU request but no limit. Adding both reduces risk and avoids CPU throttling surprises.
ports: - containerPort: 80 + securityContext: + allowPrivilegeEscalation: false + runAsNonRoot: true + capabilities: + drop: ["ALL"] resources: requests: memory: "64Mi" cpu: "10m" limits: memory: "64Mi" + cpu: "20m"
66-88: Apply the same hardening to the frontend containerMirror the securityContext and CPU limit changes here for consistency.
image: busybox @@ sleep 10 done resources: requests: memory: "64Mi" cpu: "10m" limits: memory: "64Mi" + cpu: "20m" + securityContext: + allowPrivilegeEscalation: false + runAsNonRoot: true + capabilities: + drop: ["ALL"]tests/llm/fixtures/test_ask_holmes/79_configmap_mount_issue/manifest.yaml (3)
1-5: Namespace lacks identifying labelsTagging the test namespace with something like
app.kubernetes.io/part-of: llm-fixtures(and/or anenvironment: testlabel) makes bulk cleanup and dashboard filtering simpler while having zero impact on the fixture’s behaviour.
22-36: Harden the container spec to silence security scannersCheckov flags CKV_K8S_20 & CKV_K8S_23 because the container runs as root with privilege escalation allowed.
A minimal securityContext keeps the fixture functionally identical while staying compliant:image: busybox command: ["/bin/sh"] args: ["-c", "cat /config/app.properties && sleep 3600"] + securityContext: + runAsNonRoot: true + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true
38-40: Intentional ConfigMap mismatch — suppress linter noiseThe wrong name is the whole point of the scenario. To avoid endless “ConfigMap not found” warnings from static linters, consider an inline suppression, e.g.
# kube-linter.io/disabled-checks: "configmap-volume-mount" name: app-config # This ConfigMap doesn't exist!tests/llm/fixtures/test_ask_holmes/68_cascading_failures/generate_logs.py (1)
83-115: Consider optimizing log volume for test performanceThe additional error logs and log burial strategy effectively simulate realistic scenarios where critical errors are hidden in high-volume log streams. However, generating 100,000 total log entries might impact test execution time and resource usage.
Consider reducing the log volumes while maintaining the diagnostic challenge:
- # More error logs showing the cascade - for i in range(100): + # More error logs showing the cascade + for i in range(50): service = random.choice(services) # ... existing logic ... - # Generate more normal logs to bury the errors - for i in range(89900): + # Generate more normal logs to bury the errors + for i in range(5000): service = random.choice(services) # ... existing logic ...The infinite sleep loop is correctly implemented to keep the container running for log collection.
tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/manifest.yaml (1)
37-51: Harden container security context – fail static analysis warningsCheckov flags CKV_K8S_20 and CKV_K8S_23 because the container runs as root and allows privilege escalation by default.
Tightening the security context keeps the fixture realistic while not interfering with the PVC-mismatch focus.- containers: - - name: database - image: busybox - command: ["/bin/sh"] - args: ["-c", "echo 'Database started' && sleep 3600"] + containers: + - name: database + image: busybox + command: ["/bin/sh"] + args: ["-c", "echo 'Database started' && sleep 3600"] + securityContext: + runAsNonRoot: true + runAsUser: 1000 + allowPrivilegeEscalation: falseThis small addition silences the warnings without impacting the test’s purpose.
tests/llm/fixtures/test_ask_holmes/74_config_change_impact/manifest.yaml (1)
40-40: Remove redundant restartPolicy.The
restartPolicy: Alwaysat the pod template level is redundant since Deployments automatically manage pod restarts. This is the default behavior for Deployment-managed pods.- restartPolicy: Alwaystests/llm/fixtures/test_ask_holmes/74_config_change_impact/generate_logs.py (1)
42-58: Replace unused loop variable with underscore.The static analysis correctly identifies that the loop control variable
iis not used within the loop bodies. This is a common pattern where you only need to iterate a specific number of times.- for i in range(50000): + for _ in range(50000):Apply the same change to the second loop:
- for i in range(50000): + for _ in range(50000):Also applies to: 78-94
tests/llm/fixtures/test_ask_holmes/70_memory_leak_detection/generate_logs.py (2)
40-46: Consider making the memory growth pattern more realistic.The current linear growth based on iteration count may not accurately simulate real memory leaks. Consider using a more realistic exponential or step-wise growth pattern.
- # Calculate current memory usage (exponential growth) - memory_mb = int(base_memory + (i / records_per_gb) * 1000) + # Calculate current memory usage (more realistic exponential growth) + growth_factor = 1 + (i / 50000) # Exponential growth over time + memory_mb = int(base_memory * growth_factor)
89-91: Document the indefinite sleep purpose.The indefinite sleep is necessary to keep the Kubernetes pod running for log observation, but this should be documented for clarity.
- # Keep pod running + # Keep pod running indefinitely for Kubernetes pod to remain active + # This allows log observation and prevents pod restart while True: time.sleep(3600)tests/llm/fixtures/test_ask_holmes/70_memory_leak_detection/manifest.yaml (1)
31-34: Consider adding resource limits.While minimal requests are appropriate for this test scenario, adding limits would demonstrate good resource management practices.
resources: requests: memory: "64Mi" cpu: "10m" + limits: + memory: "128Mi" + cpu: "50m"tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/manifest.yaml (1)
37-42: Harden the container with a minimal securityContextStatic analysis (CKV_K8S_20 / 23) flags the default root user and unchecked privilege escalation. Even for fixtures it’s good to embed sane defaults so examples don’t propagate insecure patterns.
- image: nginx:alpine + image: nginx:alpine + securityContext: + runAsNonRoot: true + allowPrivilegeEscalation: false resources:This keeps the manifest pedagogically sound while remaining harmless to the quota-violation scenario.
tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/manifest.yaml (3)
23-23: Pin the image by digest for deterministic test runsAn unpinned
nginx:alpinetag may change over time, breaking historical reproducibility of the fixture.- image: nginx:alpine +# 💡 pin to a specific digest to avoid future tag drift + image: nginx@sha256:<replace-with-actual-digest>
39-44: Add acpulimit to match the memory limitOnly
cpurequest is set. Omitting a limit means the pod can burst to any CPU, which makes results nondeterministic under load and triggers Checkov advice.limits: memory: 64Mi + cpu: 20m # keep modest but finite
21-24: Harden the container securityContext (CKV_K8S_20 / 23)Static analysis points out missing
allowPrivilegeEscalationand root safeguards. Even in fixtures it is cheap to model good practice.containers: - name: web-app image: nginx:alpine + securityContext: + allowPrivilegeEscalation: false + runAsNonRoot: true + runAsUser: 101 # nginx non-root uidtests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/manifest.yaml (2)
24-34: Harden backend container: addsecurityContextand set a CPU limitCheckov correctly warns that the pod can run as root and with privilege escalation. Even for fixtures it’s cheap to model good practice—your generated YAML gets copied into real clusters sooner or later.
containers: - name: backend image: nginx:alpine + securityContext: + runAsNonRoot: true + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true + capabilities: + drop: ["ALL"] ports: - containerPort: 80 resources: requests: memory: "64Mi" cpu: "10m" limits: memory: "64Mi" + cpu: "50m"
65-87: Frontend container: pin the image tag and tighten securityBusyBox without an explicit tag is mutable and can break reproducibility; same securityContext gaps as the backend.
- name: frontend - image: busybox + image: busybox:1.36 # pin exact version for reproducibility command: ["/bin/sh"] @@ resources: requests: memory: "64Mi" cpu: "10m" limits: memory: "64Mi" + securityContext: + runAsNonRoot: true + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true + capabilities: + drop: ["ALL"]tests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/manifest.yaml (1)
41-59: Harden containers with an explicitsecurityContextNone of the deployments set
runAsNonRootor disable privilege escalation, triggering CKV_K8S_20/23 findings. Add a minimalsecurityContext(replicate to the other two deployments as well):env: - name: POSTGRES_PASSWORD value: supersecret + securityContext: + runAsNonRoot: true + allowPrivilegeEscalation: false + capabilities: + drop: ["ALL"]
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (45)
pyproject.toml(1 hunks)tests/llm/fixtures/test_ask_holmes/66_http_error_needle/generate_logs.py(1 hunks)tests/llm/fixtures/test_ask_holmes/66_http_error_needle/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/66_http_error_needle/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/67_performance_degradation/generate_logs.py(1 hunks)tests/llm/fixtures/test_ask_holmes/67_performance_degradation/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/67_performance_degradation/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/68_cascading_failures/generate_logs.py(1 hunks)tests/llm/fixtures/test_ask_holmes/68_cascading_failures/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/68_cascading_failures/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/generate_logs.py(1 hunks)tests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/70_memory_leak_detection/generate_logs.py(1 hunks)tests/llm/fixtures/test_ask_holmes/70_memory_leak_detection/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/70_memory_leak_detection/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/71_connection_pool_starvation/generate_logs.py(1 hunks)tests/llm/fixtures/test_ask_holmes/71_connection_pool_starvation/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/71_connection_pool_starvation/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/73_time_window_anomaly/generate_logs.py(1 hunks)tests/llm/fixtures/test_ask_holmes/73_time_window_anomaly/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/73_time_window_anomaly/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/74_config_change_impact/generate_logs.py(1 hunks)tests/llm/fixtures/test_ask_holmes/74_config_change_impact/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/74_config_change_impact/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/75_network_flapping/generate_logs.py(1 hunks)tests/llm/fixtures/test_ask_holmes/75_network_flapping/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/75_network_flapping/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/79_configmap_mount_issue/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/79_configmap_mount_issue/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/81_service_account_permission_denied/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/81_service_account_permission_denied/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/test_case.yaml(1 hunks)
✅ Files skipped from review due to trivial changes (15)
- tests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/68_cascading_failures/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/79_configmap_mount_issue/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml
- tests/llm/fixtures/test_ask_holmes/67_performance_degradation/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/71_connection_pool_starvation/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/81_service_account_permission_denied/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/73_time_window_anomaly/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/66_http_error_needle/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/70_memory_leak_detection/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/75_network_flapping/test_case.yaml
🚧 Files skipped from review as they are similar to previous changes (3)
- pyproject.toml
- tests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/test_case.yaml
🧰 Additional context used
🪛 Checkov (3.2.334)
tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/manifest.yaml
[MEDIUM] 6-44: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-44: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/81_service_account_permission_denied/manifest.yaml
[MEDIUM] 46-78: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 46-78: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/manifest.yaml
[MEDIUM] 7-35: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 7-35: Minimize the admission of root containers
(CKV_K8S_23)
[MEDIUM] 51-87: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 51-87: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/manifest.yaml
[MEDIUM] 7-34: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 7-34: Minimize the admission of root containers
(CKV_K8S_23)
[MEDIUM] 49-87: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 49-87: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/66_http_error_needle/manifest.yaml
[MEDIUM] 6-40: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-40: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/67_performance_degradation/manifest.yaml
[MEDIUM] 6-40: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-40: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/68_cascading_failures/manifest.yaml
[MEDIUM] 6-40: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-40: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/manifest.yaml
[MEDIUM] 6-40: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-40: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/70_memory_leak_detection/manifest.yaml
[MEDIUM] 6-40: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-40: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/71_connection_pool_starvation/manifest.yaml
[MEDIUM] 6-40: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-40: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/73_time_window_anomaly/manifest.yaml
[MEDIUM] 6-40: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-40: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/74_config_change_impact/manifest.yaml
[MEDIUM] 6-40: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-40: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/75_network_flapping/manifest.yaml
[MEDIUM] 6-40: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-40: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/manifest.yaml
[MEDIUM] 19-42: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 19-42: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/79_configmap_mount_issue/manifest.yaml
[MEDIUM] 7-41: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 7-41: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/manifest.yaml
[MEDIUM] 21-54: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 21-54: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/manifest.yaml
[MEDIUM] 19-65: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 19-65: Minimize the admission of root containers
(CKV_K8S_23)
[MEDIUM] 66-90: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 66-90: Minimize the admission of root containers
(CKV_K8S_23)
[MEDIUM] 91-114: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 91-114: Minimize the admission of root containers
(CKV_K8S_23)
🪛 Ruff (0.12.2)
tests/llm/fixtures/test_ask_holmes/75_network_flapping/generate_logs.py
154-154: Loop control variable i not used within loop body
(B007)
tests/llm/fixtures/test_ask_holmes/66_http_error_needle/generate_logs.py
28-28: Loop control variable i not used within loop body
Rename unused i to _i
(B007)
41-41: Loop control variable i not used within loop body
Rename unused i to _i
(B007)
tests/llm/fixtures/test_ask_holmes/74_config_change_impact/generate_logs.py
42-42: Loop control variable i not used within loop body
(B007)
tests/llm/fixtures/test_ask_holmes/71_connection_pool_starvation/generate_logs.py
47-47: Loop control variable i not used within loop body
(B007)
109-109: Loop control variable i not used within loop body
(B007)
tests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/generate_logs.py
31-31: Loop control variable i not used within loop body
Rename unused i to _i
(B007)
59-59: Loop control variable i not used within loop body
Rename unused i to _i
(B007)
70-70: Loop control variable i not used within loop body
Rename unused i to _i
(B007)
81-81: Loop control variable i not used within loop body
Rename unused i to _i
(B007)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (3)
- GitHub Check: llm_evals
- GitHub Check: Pre-commit checks
- GitHub Check: Pre-commit checks
🔇 Additional comments (41)
tests/llm/fixtures/test_ask_holmes/71_connection_pool_starvation/manifest.yaml (1)
1-4: LGTM: Namespace definition is clean and follows conventions.The namespace definition is straightforward and appropriate for a test fixture environment.
tests/llm/fixtures/test_ask_holmes/71_connection_pool_starvation/generate_logs.py (5)
8-42: LGTM: Well-structured log entry generation function.The function effectively generates realistic connection pool monitoring logs with appropriate thresholds and detailed metadata. The conditional logic for different warning levels (80+ and 95+ connections) provides good granularity for testing log analysis scenarios.
45-77: Excellent simulation phases with realistic progression.The phased approach effectively simulates a realistic connection pool exhaustion scenario:
- Normal operation baseline (5-20 connections)
- Gradual increase during peak hours
- Introduction of wait times as pressure builds
The progression from 20 to 80 connections with corresponding wait time increases creates authentic monitoring data for testing.
79-107: Effective critical phase simulation with realistic error patterns.The critical exhaustion phase accurately represents real-world database connection pool failures:
- Pool exhaustion at 100/100 connections
- Realistic timeout errors (30000ms)
- Mix of failed requests with occasional successful connections
- Authentic error messages that would appear in production logs
This provides excellent test data for validating LLM diagnostic capabilities.
47-47: Static analysis false positive: Loop variables are intentionally unused.The unused loop variables
iflagged by Ruff are intentional and appropriate for this log generation context. The loops are used purely for iteration count to generate the desired number of log entries, not for indexing or calculations.In log generation scenarios, this pattern is common and correct. No changes needed.
Also applies to: 109-109
114-115: Appropriate container lifecycle management.The infinite sleep loop correctly keeps the container running for the test environment, allowing log collection and analysis. This is the expected pattern for containerized log generation scripts.
tests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/manifest.yaml (2)
1-4: LGTM: Namespace definition is clean and follows conventions.The namespace definition is straightforward and appropriate for test isolation.
31-34: LGTM: Resource requests are appropriate for a log generation test.The minimal resource allocation (64Mi memory, 10m CPU) is suitable for this synthetic log generation scenario.
tests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/generate_logs.py (2)
8-26: LGTM: Log entry generation function is well-structured.The function correctly generates JSON log entries with appropriate fields for rate limiting simulation. The conditional logic for error levels based on status codes is clean.
29-92: LGTM: Log generation sequence effectively simulates rate limiting scenario.The script creates a realistic progression from normal traffic → spike detection → rate limiting → recovery. The volume and timing of logs (150,000 total entries) should provide sufficient data for LLM testing scenarios.
tests/llm/fixtures/test_ask_holmes/75_network_flapping/manifest.yaml (1)
1-41: Well-structured manifest for network flapping test scenario.The manifest correctly defines the namespace and deployment structure. The use of a secret-mounted script with executable permissions is appropriate for this test fixture, and the minimal resource allocation is suitable for log generation.
tests/llm/fixtures/test_ask_holmes/75_network_flapping/generate_logs.py (3)
34-43: Excellent design for progressive network degradation simulation.The four-phase approach effectively simulates realistic network flapping scenarios with clear progression from 0.1% to 60% timeout rates. This provides comprehensive test data for LLM-based diagnostics.
8-31: Well-structured log entry generation function.The function correctly generates structured JSON logs with appropriate log levels, timestamps, and contextual information. The conditional logic for error vs. success states is sound.
175-177: Appropriate keep-alive mechanism for test pod.The infinite sleep loop correctly keeps the pod running for log collection during testing scenarios.
tests/llm/fixtures/test_ask_holmes/67_performance_degradation/manifest.yaml (2)
31-34: Resource allocation is appropriate for the test scenario.The resource requests (64Mi memory, 10m CPU) are well-suited for a log generation script that simulates performance degradation without consuming excessive cluster resources.
36-39: Secret volume mounting is correctly configured.The secret volume mount with executable permissions (0755) is properly configured to run the Python script from the mounted location.
tests/llm/fixtures/test_ask_holmes/67_performance_degradation/generate_logs.py (4)
43-49: Effective exponential degradation model.The mathematical approach using
math.exp(i / 20000)creates a realistic performance degradation pattern, starting at 100ms and gradually increasing to timeout levels. The capping at 6000ms ensures the simulation doesn't produce unrealistic values.
11-25: Well-designed log level progression.The conditional logic effectively maps response times to appropriate log levels:
- Normal operation (≤3000ms): INFO
- Performance degradation (3000-5000ms): WARN
- Timeout conditions (>5000ms): ERROR
This creates a realistic troubleshooting scenario for the LLM evaluation.
31-34: Realistic resource usage simulation.The CPU and memory calculations provide correlated metrics that would help in diagnosing performance issues:
- CPU usage increases proportionally with response time
- Memory usage also scales, simulating resource pressure
- Values are bounded to prevent unrealistic metrics
54-59: Proper container lifecycle management.The throttling mechanism (sleep every 1000 entries) simulates real-time log generation, and the infinite sleep loop at the end keeps the container running for the test duration, which is essential for the Kubernetes test scenario.
tests/llm/fixtures/test_ask_holmes/66_http_error_needle/generate_logs.py (1)
8-23: LGTM! Well-designed log generation function.The function effectively generates realistic HTTP log entries with appropriate randomization and automatic error level detection based on status codes. The use of default parameters and private IP ranges is suitable for testing scenarios.
tests/llm/fixtures/test_ask_holmes/81_service_account_permission_denied/manifest.yaml (1)
18-24: Intentional omission oflistverb looks correct for the negative-test scenarioThe Role intentionally grants only
getonpods, which will trigger the expected “permission denied – cannot list pods” error exercised by the test. Looks good and aligns with the fixture’s goal.tests/llm/fixtures/test_ask_holmes/73_time_window_anomaly/generate_logs.py (1)
26-102: LGTM! Well-structured log generation script.The script effectively simulates a realistic time-window anomaly scenario with:
- Proper 24-hour log generation
- Targeted error injection in the 03:00-03:05 window
- Realistic timing variations and job types
- Appropriate JSON log formatting
- Correct container lifecycle management
The implementation supports the test case objective of detecting time-window-specific scheduling issues.
tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/manifest.yaml (1)
36-46: Confirm the selector mismatch is truly intentionalThe selector deliberately targets
version: v2, so no backend Pod will ever be selected and the Service will remain without endpoints.
If that’s the objective of the test fixture, great. Otherwise, switch the selector toversion: v1(or remove the version key) to restore connectivity.tests/llm/fixtures/test_ask_holmes/68_cascading_failures/manifest.yaml (2)
1-4: LGTM: Clean namespace definitionThe namespace definition follows standard Kubernetes conventions and aligns with the test case numbering scheme.
6-40: LGTM: Well-structured test deployment with appropriate security contextThe deployment configuration is well-suited for a test fixture generating synthetic logs. The static analysis warnings about privilege escalation and root containers are acceptable in this test context since:
- Test fixtures run in isolated environments where security constraints can be relaxed
- The container executes a specific Python script for log generation
- Minimal resource allocation is appropriate for the workload
The secret volume mounting with executable permissions (0755) is correctly configured for script execution.
tests/llm/fixtures/test_ask_holmes/68_cascading_failures/generate_logs.py (3)
8-19: LGTM: Well-designed log entry generation functionThe function creates realistic structured log entries with proper timestamp handling, random instance IDs for authenticity, and conditional trace ID correlation. The ISO timestamp format with UTC is appropriate for log analysis scenarios.
22-49: LGTM: Excellent cascading failure simulation setupThe initial normal log generation and root cause implementation are well-designed:
- Realistic baseline: 10,000 normal logs establish expected behavior
- Clear root cause: Redis connection failure is a common real-world scenario
- Proper timing: Timestamp offsets create chronological log sequence
The Redis connection error message is realistic and provides clear diagnostic information.
50-81: LGTM: Realistic cascading failure propagationThe failure sequence excellently demonstrates microservice dependency chains:
- Logical progression: auth-service → user-service → order-service → payment-processor
- Realistic timing: 5-second intervals between failures
- Trace correlation: Shared trace_id enables failure tracking
- Diagnostic messages: Each error clearly indicates the upstream dependency failure
This provides excellent training data for LLM analysis of distributed system failures.
tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/manifest.yaml (1)
15-18: Intentional storage-class mismatch acknowledgedThe
storageClassName: fast-ssdline is the crux of this negative test; leaving it as an invalid value is correct for the scenario you’re validating.tests/llm/fixtures/test_ask_holmes/74_config_change_impact/manifest.yaml (1)
31-34: Resource allocation is appropriate for test scenario.The minimal resource requests (64Mi memory, 10m CPU) are well-suited for a log generation script in a test environment.
tests/llm/fixtures/test_ask_holmes/74_config_change_impact/test_case.yaml (2)
1-21: Well-structured test case with appropriate expectations.The test case effectively defines a scenario for evaluating cache service performance diagnosis. The expected outputs align well with the log patterns that will be generated by the Python script, and the setup/teardown commands properly manage the test environment.
12-14: Good use of kubectl dry-run for secret creation.Using
--dry-run=client -o yaml | kubectl apply -f -is a best practice that ensures the secret is created or updated safely without conflicts.tests/llm/fixtures/test_ask_holmes/74_config_change_impact/generate_logs.py (2)
8-34: Well-designed log entry generation function.The
generate_log_entryfunction effectively creates realistic JSON log entries with appropriate conditional logic for warning levels and metadata inclusion. The hit rate calculation and random TTL metadata inclusion add realism to the simulation.
37-76: Effective simulation of performance degradation scenario.The two-phase approach (high performance followed by poor performance after configuration reload) creates a clear before/after pattern that aligns perfectly with the test case expectations. The counter reset between phases effectively demonstrates the impact of the configuration change.
tests/llm/fixtures/test_ask_holmes/70_memory_leak_detection/generate_logs.py (2)
7-31: LGTM! Well-structured log entry generator.The
generate_log_entryfunction correctly creates structured JSON logs with appropriate log levels based on memory usage thresholds (WARN >3000MB, ERROR for exceptions). The timestamp formatting and memory percentage calculations are accurate.
70-87: OOM condition logic is sound.Simulation shows
memory_mbreaches 4096 at iteration 99900 (> 99000), so the error block is reliably triggered. No changes required.tests/llm/fixtures/test_ask_holmes/70_memory_leak_detection/manifest.yaml (1)
1-5: LGTM! Proper namespace isolation.The namespace creation follows Kubernetes best practices for test isolation.
tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/manifest.yaml (1)
1-17: Scenario manifest looks correct and purpose-built to exceed the quotaThe namespace,
ResourceQuota, and deployment skeleton all line up with the goal of provoking quota violations in the test case. No functional issues spotted in these lines.tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/manifest.yaml (1)
26-30: Intentional probe mismatch is clear ‑ no action requiredThe liveness probe deliberately targets port 8080 instead of 80 to exercise the “mis-configured probe” scenario the fixture is meant to test. LGTM as-is.
tests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/manifest.yaml (1)
95-107: NetworkPolicy logic looks correct—ingress restricted exactly as the test expectsThe policy selects backend pods and only allows ingress from peers with
tier=backend, so traffic from the frontend (tier=frontend) is indeed denied. Nice concise reproduction of a common mis-configuration.
There was a problem hiding this comment.
Actionable comments posted: 2
🧹 Nitpick comments (7)
tests/llm/fixtures/test_ask_holmes/27_permissions_error_no_helm_tools/test_case.yaml (1)
2-6: Expected-output text is still quite prescriptive and may create brittle gradingNumbered sentences make the judge expect those exact lines, order and wording. From prior feedback (
03_what_is_the_command_to_port_forward, see retrieved learnings) we try to keep expectations flexible to avoid false negatives when the LLM answers with equivalent wording or a different ordering. Consider switching to a looser matcher – e.g. a YAML list of key substrings or a regex list – and omit ordering cues.-expected_output: | - user is told: - 1. holmes lacks permission - 2. is given an example of configuration to add permissions - 3. is told to run `helm upgrade ...` to apply that configuration +expected_output: + - lacks permission + - example configuration to grant permissions + - suggest running `helm upgrade`This keeps the essence while letting wording/order vary.
tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/test_case.yaml (1)
7-9: Replace fixedsleep 30with a deterministic readiness checkBlind sleeps make the test either slower than needed or flaky on congested clusters.
kubectl waitgives a deterministic, faster setup.-before_test: | - kubectl apply -f manifest.yaml - sleep 30 +before_test: | + kubectl apply -f manifest.yaml + kubectl -n namespace-04 wait --for=condition=available deployment/nginx --timeout=60stests/llm/fixtures/test_ask_holmes/05_image_version/test_case.yaml (3)
1-1: Clarify pod selector by including namespace in the entity referenceIf another
sea-turtlepod ever exists in a different namespace, the LLM could return an unexpected image. Explicitly passing the namespace keeps the prompt unambiguous:-user_prompt: 'what is the image version of << { "type": "pod", "name": "sea-turtle" } >>' +user_prompt: 'what is the image version of << { "type": "pod", "name": "sea-turtle", "namespace": "app-05" } >>'
8-26: Remove leading indentation in the heredoc to avoid malformed YAMLThe extra two-space prefix on every manifest line is preserved by the heredoc and becomes part of the file, so
apiVersionis indented instead of starting at column 0. Most kubectl parsers tolerate it, but a few strict linters and CI jobs flag this as a formatting error. Dropping the indentation keeps the manifest canonical.- cat <<EOF | kubectl apply -f - - apiVersion: v1 - kind: Pod - metadata: - name: sea-turtle - namespace: app-05 - labels: - app: analytics-engine - spec: - containers: - - name: web - image: nginx:1.23.4 - env: - - name: VERSION - value: "1.23.4" - EOF + cat <<EOF | kubectl apply -f - +apiVersion: v1 +kind: Pod +metadata: + name: sea-turtle + namespace: app-05 + labels: + app: analytics-engine +spec: + containers: + - name: web + image: nginx:1.23.4 + env: + - name: VERSION + value: "1.23.4" +EOF
29-29: Add--ignore-not-foundto make teardown idempotentIf the namespace is already gone (e.g., a flaky previous cleanup),
kubectl deleteexits non-zero and can fail the test run. Adding the flag makes the teardown resilient:- kubectl delete namespace app-05 --force --grace-period=0 + kubectl delete namespace app-05 --force --grace-period=0 --ignore-not-foundtests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml (2)
2-6: Expected-output is too prescriptive; relax matching to reduce false negativesThe bullets hard-code the exact pod name (
grafana-7f8b9c6d5-x2m4n) and port number. While deterministic for this fixture, it contradicts the previous learning that HolmesGPT tests should reward semantically correct answers rather than string-perfect matches. Any future change to the pod hash or port will break the test even though the LLM’s answer is still correct.Consider switching to a regex-style matcher or placeholder wording (e.g. “
grafana-<pod-hash>”, “<port>”) and validate with a regex in the evaluator instead of literal string comparison.
23-24:busybox:1.35+nc -l -p 3000can be flaky across BusyBox buildsSome BusyBox variants omit
nc, or require the BSD-style-l -pflags to be combined as-lp. A missing or crashingncwill make the pod stayCrashLoopBackOff, causing thekubectl waitstep to time out and the whole test to fail.A more robust one-liner is to use Alpine + a simple HTTP server, or even keep BusyBox but replace the loop with an inert process:
- image: busybox:1.35 - command: ['sh', '-c', 'while true; do nc -l -p 3000; done'] + image: alpine:3.19 + command: ['sh', '-c', 'while true; do sleep 3600; done']The pod comes up instantly and holds port 3000 open regardless of netcat availability, making the fixture less brittle.
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (19)
tests/llm/fixtures/test_ask_holmes/02_what_is_wrong_with_pod/kubectl_describe.txt(0 hunks)tests/llm/fixtures/test_ask_holmes/02_what_is_wrong_with_pod/kubectl_find_resource.txt(0 hunks)tests/llm/fixtures/test_ask_holmes/02_what_is_wrong_with_pod/kubectl_logs.txt(0 hunks)tests/llm/fixtures/test_ask_holmes/02_what_is_wrong_with_pod/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/kubectl_describe.txt(0 hunks)tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/kubectl_events.txt(0 hunks)tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/kubectl_find_resource.txt(0 hunks)tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/kubectl_lineage_parents.txt(0 hunks)tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/05_image_version/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/06_explain_issue/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/27_permissions_error_no_helm_tools/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/33_http_latency_graph/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/48_logs_since_thursday/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/55_kafka_runbook/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/56_kafka_runbook_no_tool/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/73_time_window_anomaly/test_case.yaml(1 hunks)
💤 Files with no reviewable changes (7)
- tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/kubectl_find_resource.txt
- tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/kubectl_events.txt
- tests/llm/fixtures/test_ask_holmes/02_what_is_wrong_with_pod/kubectl_find_resource.txt
- tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/kubectl_lineage_parents.txt
- tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/kubectl_describe.txt
- tests/llm/fixtures/test_ask_holmes/02_what_is_wrong_with_pod/kubectl_logs.txt
- tests/llm/fixtures/test_ask_holmes/02_what_is_wrong_with_pod/kubectl_describe.txt
✅ Files skipped from review due to trivial changes (5)
- tests/llm/fixtures/test_ask_holmes/56_kafka_runbook_no_tool/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/48_logs_since_thursday/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/55_kafka_runbook/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/06_explain_issue/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/73_time_window_anomaly/test_case.yaml
🧰 Additional context used
🧠 Learnings (6)
tests/llm/fixtures/test_ask_holmes/33_http_latency_graph/test_case.yaml (1)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
tests/llm/fixtures/test_ask_holmes/27_permissions_error_no_helm_tools/test_case.yaml (1)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
tests/llm/fixtures/test_ask_holmes/05_image_version/test_case.yaml (1)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/test_case.yaml (1)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
tests/llm/fixtures/test_ask_holmes/02_what_is_wrong_with_pod/test_case.yaml (1)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml (1)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
🪛 Checkov (3.2.334)
tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/manifest.yaml
[MEDIUM] 6-31: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-31: Minimize the admission of root containers
(CKV_K8S_23)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (3)
- GitHub Check: Pre-commit checks
- GitHub Check: llm_evals
- GitHub Check: Pre-commit checks
🔇 Additional comments (4)
tests/llm/fixtures/test_ask_holmes/27_permissions_error_no_helm_tools/test_case.yaml (1)
13-14: Verify harness support formock_policy
mock_policy: always_mockis new to this fixture. Ensure the evaluation framework actually consumes this key; if not, the flag will be silently ignored and the test may behave unexpectedly.tests/llm/fixtures/test_ask_holmes/33_http_latency_graph/test_case.yaml (2)
10-10: Confirm that the “skip” mechanism is implemented as a tagThe test harnesses I’ve seen either expect a top-level key (
skip: true) or a marker in the filename; using a tag value (- skip) is uncommon. Double-check that the runner actually inspectstagsforskip, otherwise the case will still execute.
4-5: Validation complete: referenced manifest existsThe
./app/manifest.yamlfile referenced in
tests/llm/fixtures/test_ask_holmes/33_http_latency_graph/test_case.yaml
resolves to
tests/llm/fixtures/test_ask_holmes/33_http_latency_graph/app/manifest.yaml.
No further action required.tests/llm/fixtures/test_ask_holmes/02_what_is_wrong_with_pod/test_case.yaml (1)
1-1: Include namespace in the object reference to avoid ambiguityDown-stream evaluators often rely on the triple <type/namespace/name>. Omitting the namespace may cause the LLM to mis-resolve the reference when multiple pods with the same name exist in different namespaces.
-user_prompt: 'What is wrong with << { "type": "pod", "name": "giant-narwhal-6958c5bdd8-69gtn" } >> ?' +user_prompt: 'What is wrong with << { "type": "pod", "namespace": "app-02", "name": "giant-narwhal-6958c5bdd8-69gtn" } >> ?'
There was a problem hiding this comment.
Actionable comments posted: 1
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (73)
.github/workflows/llm-evaluation.yaml(1 hunks)CLAUDE.md(1 hunks)pyproject.toml(1 hunks)tests/llm/fixtures/test_ask_holmes/01_how_many_pods/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/07_high_latency/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/08_sock_shop_frontend/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/09_crashpod/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/10_image_pull_backoff/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/11_init_containers/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/12_job_crashing/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/13_pending_node_selector/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/14_pending_resources/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/15_failed_readiness_probe/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/17_oom_kill/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/18_crash_looping_v2/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/20_long_log_file_search/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/21_job_fail_curl_no_svc_account/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/22_high_latency_dbi_down/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/23_app_error_in_current_logs/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/24_misconfigured_pvc/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/25_misconfigured_ingress_class/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/27_permissions_error_no_helm_tools/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/28_permissions_error_helm_tools_enabled/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/33_http_latency_graph/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/38_rabbitmq_split_head/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_all_tools/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_new_tools/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_new_tools_no_runbook/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_old_tools/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_new_all_tools/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_new_tools/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_old_tools/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/43_slack_deployment_logs/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/45_fetch_deployment_logs_simple/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/46_job_crashing_no_longer_exists/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/51_logs_summarize_errors/test_case.yaml(2 hunks)tests/llm/fixtures/test_ask_holmes/52_logs_login_issues/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/53_logs_find_term/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/55_kafka_runbook/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/56_kafka_runbook_no_tool/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/57_wrong_namespace/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/58_counting_pods_by_status/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/59_label_based_counting/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/60_count_less_than/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/61_exact_match_counting/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/62_fetch_error_logs_with_errors/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/63_fetch_error_logs_no_errors/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/64_keda_vs_hpa_confusion/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/65_health_check_followup/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/66_http_error_needle/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/67_performance_degradation/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/68_cascading_failures/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/70_memory_leak_detection/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/71_connection_pool_starvation/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/73_time_window_anomaly/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/74_config_change_impact/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/75_network_flapping/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/79_configmap_mount_issue/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/81_service_account_permission_denied/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/83_secret_not_found/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/test_case.yaml(1 hunks)
💤 Files with no reviewable changes (32)
- tests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_old_tools/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/18_crash_looping_v2/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/38_rabbitmq_split_head/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/01_how_many_pods/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_old_tools/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/45_fetch_deployment_logs_simple/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/08_sock_shop_frontend/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/64_keda_vs_hpa_confusion/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/22_high_latency_dbi_down/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/28_permissions_error_helm_tools_enabled/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/23_app_error_in_current_logs/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_new_tools/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/10_image_pull_backoff/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/24_misconfigured_pvc/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/25_misconfigured_ingress_class/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/12_job_crashing/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_all_tools/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/15_failed_readiness_probe/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/46_job_crashing_no_longer_exists/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/17_oom_kill/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/20_long_log_file_search/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_new_tools_no_runbook/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/43_slack_deployment_logs/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/21_job_fail_curl_no_svc_account/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/13_pending_node_selector/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/59_label_based_counting/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/11_init_containers/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_new_all_tools/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/58_counting_pods_by_status/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/57_wrong_namespace/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/65_health_check_followup/test_case.yaml
✅ Files skipped from review due to trivial changes (9)
- tests/llm/fixtures/test_ask_holmes/07_high_latency/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/09_crashpod/test_case.yaml
- CLAUDE.md
- tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/81_service_account_permission_denied/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_new_tools/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/71_connection_pool_starvation/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/test_case.yaml
🚧 Files skipped from review as they are similar to previous changes (31)
- tests/llm/fixtures/test_ask_holmes/14_pending_resources/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/62_fetch_error_logs_with_errors/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/56_kafka_runbook_no_tool/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/60_count_less_than/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/61_exact_match_counting/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/53_logs_find_term/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/52_logs_login_issues/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/27_permissions_error_no_helm_tools/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/33_http_latency_graph/test_case.yaml
- pyproject.toml
- tests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/51_logs_summarize_errors/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/63_fetch_error_logs_no_errors/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/55_kafka_runbook/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/73_time_window_anomaly/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/83_secret_not_found/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/68_cascading_failures/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/67_performance_degradation/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/66_http_error_needle/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/70_memory_leak_detection/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/75_network_flapping/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/79_configmap_mount_issue/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/88_affinity_like_but_taints/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/69_rate_limit_exhaustion/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/74_config_change_impact/test_case.yaml
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
- GitHub Check: Pre-commit checks
- GitHub Check: Pre-commit checks
There was a problem hiding this comment.
Actionable comments posted: 1
🧹 Nitpick comments (3)
tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml (2)
3-6: Consider relaxing the exact-match requirement on pod name & port to keep the test resilientHard-coding the full pod name (
grafana-7f8b9c6d5-x2m4n) and port3000in the grading rubric makes the evaluation very brittle:
- If the manifest or generator ever changes the random suffix, the fixture will fail even when the LLM gives a perfectly valid answer (
kubectl port-forward -n app-03 $(kubectl get pod -l app=grafana -n app-03 -o name | cut -d/ -f2) 3000:3000, etc.).- Past feedback (see Sheeproid learning) recommends descriptive checks for flexible answers to avoid unnecessary failures.
A more robust approach is to assert that the answer contains:
kubectl port-forward- the correct namespace (
app-03)- a variable pod name matching
grafana-*(regex)- the port mapping
3000:3000Example diff (YAML anchors/regex supported by most internal graders):
- - Must find the actual grafana pod name (grafana-7f8b9c6d5-x2m4n) - - Must include the correct port (3000) - - "Full kubectl port-forward command like: kubectl port-forward -n app-03 grafana-7f8b9c6d5-x2m4n 3000:3000" + - Contains "kubectl port-forward" + - Includes namespace "app-03" + - References a pod matching /grafana-[a-z0-9-]+/ + - Maps port 3000:3000
11-29: Pod setup works but can be simplified & made lighterThe busybox command spins an endless
ncloop which is unnecessary for a port-forward test and may consume CPU on constrained runners. A lighter (and less error-prone) option is:command: ["sleep", "infinity"] ports: - containerPort: 3000 name: httpThis keeps the container Ready without requiring
nc(which is not compiled into all BusyBox builds).
If traffic testing is required later, you can add a readinessProbe instead.Minor bonus: consider adding
---separators after the namespace creation to avoid accidental YAML concatenation when manifests grow.tests/llm/fixtures/test_ask_holmes/79_configmap_mount_issue/test_case.yaml (1)
4-6: Replace fixedsleep 20with a deterministickubectl waitStatic sleeps slow the test suite and can still be flaky if cluster latency is higher or lower than expected.
Leveragekubectl waiton the Deployment/Pod readiness with a timeout—this is faster when the resource is ready early and safer when it takes longer.- sleep 20 +# wait up to 20 s for pod containers to start (non-blocking once ready) + kubectl wait --for=condition=ContainersReady pod -l app=app-server -n namespace-79 --timeout=20s || true
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (56)
pyproject.toml(1 hunks)tests/llm/fixtures/test_ask_holmes/01_how_many_pods/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/05_image_version/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/08_sock_shop_frontend/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/09_crashpod/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/10_image_pull_backoff/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/11_init_containers/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/12_job_crashing/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/13_pending_node_selector/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/14_pending_resources/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/15_failed_readiness_probe/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/16_failed_no_toolset_found/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/17_oom_kill/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/18_crash_looping_v2/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/20_long_log_file_search/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/23_app_error_in_current_logs/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/24_misconfigured_pvc/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/28_permissions_error_helm_tools_enabled/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/29_events_from_alert_manager/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/39_failed_toolset/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/40_disabled_toolset/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/41_setup_argo/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_all_tools/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_new_tools/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/43_current_datetime_from_prompt/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/43_slack_deployment_logs/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/conversation_history/01_assistant.md(1 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/fetch_pod_logs_alertmanager-0_ask-holmes-slack-statefulset-logs_-10800.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/fetch_pod_logs_alertmanager-0_ask-holmes-slack-statefulset-logs_2025-06-12.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/fetch_pod_logs_alertmanager-0_ask-holmes-slack-statefulset-logs_2025-06-12_2025-06-12.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/fetch_pod_logs_alertmanager-alertmanager-0_ask-holmes-slack-statefulset-logs_-10800.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/kubectl_find_resource_statefulset_alertmanager.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/kubectl_get_by_kind_in_namespace_pod_ask-holmes-slack-statefulset-logs.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/kubectl_get_by_kind_in_namespace_statefulset_ask-holmes-slack-statefulset-logs.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/kubectl_get_by_name_statefulset_alertmanager_ask-holmes-slack-statefulset-logs.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/kubectl_get_yaml_statefulset_alertmanager_ask-holmes-slack-statefulset-logs.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/kubectl_lineage_children_statefulset_alertmanager_ask-holmes-slack-statefulset-logs.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/manifest.yaml(11 hunks)tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/45_fetch_deployment_logs_simple/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/48_logs_since_thursday/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/51_logs_summarize_errors/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/53_logs_find_term/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/kubectl_describepod_fast-fail-pod_ask-holmes-namespace-54b.txt(2 hunks)tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/kubectl_get_by_kind_in_namespacepod_ask-holmes-namespace-54b.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/manifest.yaml(2 hunks)tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/59_label_based_counting/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/61_exact_match_counting/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/62_fetch_error_logs_with_errors/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/79_configmap_mount_issue/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/81_service_account_permission_denied/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/83_secret_not_found/test_case.yaml(1 hunks)
✅ Files skipped from review due to trivial changes (26)
- tests/llm/fixtures/test_ask_holmes/39_failed_toolset/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/41_setup_argo/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/kubectl_get_by_kind_in_namespacepod_ask-holmes-namespace-54b.txt
- tests/llm/fixtures/test_ask_holmes/16_failed_no_toolset_found/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/kubectl_find_resource_statefulset_alertmanager.txt
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/fetch_pod_logs_alertmanager-0_ask-holmes-slack-statefulset-logs_2025-06-12_2025-06-12.txt
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/kubectl_get_by_kind_in_namespace_pod_ask-holmes-slack-statefulset-logs.txt
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/conversation_history/01_assistant.md
- tests/llm/fixtures/test_ask_holmes/01_how_many_pods/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/fetch_pod_logs_alertmanager-0_ask-holmes-slack-statefulset-logs_2025-06-12.txt
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/kubectl_get_by_name_statefulset_alertmanager_ask-holmes-slack-statefulset-logs.txt
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/kubectl_lineage_children_statefulset_alertmanager_ask-holmes-slack-statefulset-logs.txt
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/kubectl_get_by_kind_in_namespace_statefulset_ask-holmes-slack-statefulset-logs.txt
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/kubectl_get_yaml_statefulset_alertmanager_ask-holmes-slack-statefulset-logs.txt
- tests/llm/fixtures/test_ask_holmes/29_events_from_alert_manager/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/48_logs_since_thursday/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/11_init_containers/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/40_disabled_toolset/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/43_current_datetime_from_prompt/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/fetch_pod_logs_alertmanager-alertmanager-0_ask-holmes-slack-statefulset-logs_-10800.txt
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/fetch_pod_logs_alertmanager-0_ask-holmes-slack-statefulset-logs_-10800.txt
- tests/llm/fixtures/test_ask_holmes/42_dns_issues_result_all_tools/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/manifest.yaml
- tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/kubectl_describepod_fast-fail-pod_ask-holmes-namespace-54b.txt
- tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/manifest.yaml
🚧 Files skipped from review as they are similar to previous changes (28)
- tests/llm/fixtures/test_ask_holmes/28_permissions_error_helm_tools_enabled/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/08_sock_shop_frontend/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/24_misconfigured_pvc/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/61_exact_match_counting/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/23_app_error_in_current_logs/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/09_crashpod/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/43_slack_deployment_logs/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/20_long_log_file_search/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/53_logs_find_term/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/45_fetch_deployment_logs_simple/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/12_job_crashing/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/13_pending_node_selector/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/18_crash_looping_v2/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/42_dns_issues_steps_new_tools/test_case.yaml
- pyproject.toml
- tests/llm/fixtures/test_ask_holmes/54_not_truncated_when_getting_pods/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/17_oom_kill/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/51_logs_summarize_errors/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/10_image_pull_backoff/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/15_failed_readiness_probe/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/62_fetch_error_logs_with_errors/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/05_image_version/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/81_service_account_permission_denied/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/59_label_based_counting/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/83_secret_not_found/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/14_pending_resources/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/04_related_k8s_events/test_case.yaml
🧰 Additional context used
🧠 Learnings (2)
tests/llm/fixtures/test_ask_holmes/79_configmap_mount_issue/test_case.yaml (1)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml (1)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (6)
- GitHub Check: build (3.12)
- GitHub Check: build (3.11)
- GitHub Check: build (3.10)
- GitHub Check: build (3.12)
- GitHub Check: build (3.10)
- GitHub Check: build (3.11)
🔇 Additional comments (1)
tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml (1)
32-32: Namespace deletion is forceful; verify other tests don’t still rely on it
kubectl delete namespace app-03 --force --grace-period=0is fine for speed, but will abruptly terminate resources. Ensure no parallel tests expect the namespace to persist or rely on graceful shutdown hooks.
No description provided.