Repository navigation
small eval fixes - #755
small eval fixes#755
Conversation
WalkthroughThis update introduces new and modified test fixtures for Kubernetes troubleshooting scenarios within the LLM test suite. Changes include new test cases for pod scheduling failures due to nodeSelector mismatches, PVC misconfigurations, and HPA scaling issues, as well as metadata adjustments such as tag updates and evaluation section removals in several existing test cases. Additionally, namespace parameters in several fixtures were adjusted for accuracy, and a deployment manifest was updated for the HPA scaling test. Changes
Sequence Diagram(s)sequenceDiagram
participant User
participant LLM Test Suite
participant Kubernetes Cluster
User->>LLM Test Suite: Submit troubleshooting prompt (e.g., "What is the issue with redis?")
LLM Test Suite->>Kubernetes Cluster: Apply test setup manifests
LLM Test Suite->>Kubernetes Cluster: Execute kubectl commands (describe, get, etc.)
Kubernetes Cluster-->>LLM Test Suite: Return resource status and logs
LLM Test Suite->>LLM Test Suite: Analyze responses and match expected output
LLM Test Suite-->>User: Return diagnostic explanation and suggested resolution
LLM Test Suite->>Kubernetes Cluster: Teardown test resources
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~15 minutes Suggested labels
Suggested reviewers
Note ⚡️ Unit Test Generation is now available in beta!Learn more here, or try it out under "Finishing Touches" below. 📜 Recent review detailsConfiguration used: CodeRabbit UI 📒 Files selected for processing (5)
🚧 Files skipped from review as they are similar to previous changes (4)
🧰 Additional context used📓 Path-based instructions (1)tests/llm/fixtures/**/*📄 CodeRabbit Inference Engine (CLAUDE.md)
Files:
🧠 Learnings (2)📓 Common learningstests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_describe_persistentvolumeclaim_redis-standard-pvc_app-24a.txt (4)Learnt from: CR Learnt from: CR Learnt from: Sheeproid Learnt from: CR ⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (3)
🔇 Additional comments (1)
✨ Finishing Touches🧪 Generate unit tests
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. 🪧 TipsChatThere are 3 ways to chat with CodeRabbit:
SupportNeed help? Create a ticket on our support page for assistance with any issues or questions. Note: Be mindful of the bot's finite context window. It's strongly recommended to break down tasks such as reading entire modules into smaller chunks. For a focused discussion, use review comments to chat about specific files and their changes, instead of using the PR comments. CodeRabbit Commands (Invoked using PR comments)
Other keywords and placeholders
CodeRabbit Configuration File (
|
There was a problem hiding this comment.
Actionable comments posted: 6
🧹 Nitpick comments (4)
tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml (1)
21-25: Harden the container: drop root & disable privilege escalationStatic analysis flags CKV_K8S_20 / 23. Add a minimal securityContext to the test fixture without affecting the HPA scenario:
- name: worker image: busybox command: ["sh", "-c", "while true; do echo 'Working...'; sleep 1; done"] + securityContext: + runAsNonRoot: true + allowPrivilegeEscalation: falsetests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/test_case.yaml (1)
5-6: Long inline comment clutters thetagslistThe explanatory sentence after
#makes the tag line hard to scan. Consider moving the remark to a separate YAML comment line above theeasytag.tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/test_case.yaml (1)
7-8: Inline explanation inside tag is noisyLike in 24a, move the explanatory note (
# requires deeper analysis…) to a standalone YAML comment line to keep the tag list clean.tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_describe.txt (1)
33-39: Minor: extremely generic label key/value.Using
label=someLabelworks for the test but is less realistic than the conventional<key>=<value>style (e.g.environment=production).
Not a blocker, but consider switching to a more production-like pair to keep fixtures closer to real clusters.
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (16)
tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_describe.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_find_resource.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_get_by_kind_in_cluster.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/26_multi_container_logs/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/test_case.yaml(0 hunks)tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/test_case.yaml(1 hunks)
💤 Files with no reviewable changes (1)
- tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/test_case.yaml
🧰 Additional context used
📓 Path-based instructions (1)
tests/llm/fixtures/**/*
📄 CodeRabbit Inference Engine (CLAUDE.md)
tests/llm/fixtures/**/*: Mock data: tests/llm/fixtures/{test_name}/
All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Files:
tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/test_case.yamltests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/test_case.yamltests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/test_case.yamltests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/test_case.yamltests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_find_resource.txttests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/test_case.yamltests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_describe.txttests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/test_case.yamltests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/test_case.yamltests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_get_by_kind_in_cluster.txttests/llm/fixtures/test_ask_holmes/26_multi_container_logs/test_case.yamltests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/test_case.yamltests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/test_case.yamltests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/test_case.yamltests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml
🧠 Learnings (16)
📓 Common learnings
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/**/*.py : Complex investigations should have LLM evaluation tests
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Mock data: tests/llm/fixtures/{test_name}/
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: LLM evaluation tests run automatically in CI
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/**/*.py : Each LLM test must use a dedicated namespace app-<testid> (e.g., app-01, app-02) to prevent conflicts when tests run simultaneously
tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/test_case.yaml (1)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/test_case.yaml (2)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
tests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/test_case.yaml (1)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/test_case.yaml (3)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_find_resource.txt (3)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Mock data: tests/llm/fixtures/{test_name}/
tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/test_case.yaml (2)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_describe.txt (4)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Mock data: tests/llm/fixtures/{test_name}/
tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/test_case.yaml (4)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Mock data: tests/llm/fixtures/{test_name}/
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/test_case.yaml (1)
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_get_by_kind_in_cluster.txt (4)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Mock data: tests/llm/fixtures/{test_name}/
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
tests/llm/fixtures/test_ask_holmes/26_multi_container_logs/test_case.yaml (5)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/**/*.py : Complex investigations should have LLM evaluation tests
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: LLM evaluation tests run automatically in CI
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: nherment
PR: #408
File: holmes/plugins/toolsets/kubernetes_logs.py:90-97
Timestamp: 2025-05-15T05:13:43.169Z
Learning: In the Kubernetes logs toolset for Holmes, both current and previous logs are intentionally fetched and combined for each pod, even though this requires more API calls. This design ensures all logs are captured even when pods restart but retain their name, providing complete diagnostic information.
tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/test_case.yaml (4)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Mock data: tests/llm/fixtures/{test_name}/
tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/test_case.yaml (4)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Mock data: tests/llm/fixtures/{test_name}/
tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/test_case.yaml (4)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Mock data: tests/llm/fixtures/{test_name}/
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml (1)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
🪛 Checkov (3.2.334)
tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml
[MEDIUM] 6-26: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-26: Minimize the admission of root containers
(CKV_K8S_23)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (3)
- GitHub Check: llm_evals
- GitHub Check: Pre-commit checks
- GitHub Check: Pre-commit checks
🔇 Additional comments (16)
tests/llm/fixtures/test_ask_holmes/86_configmap_like_but_secret/test_case.yaml (3)
7-9: Trim the block scalar or switch to single-line style
expected_outputis declared with the literal block indicator (|), but it only contains a single short sentence.
The extra newline emitted by|is harmless, yet it creates noisy diffs and can trip fragile string-equality checks in the grading harness.-expected_output: | - The issue should be identified as a missing Secret mount (frontend-api-keys), +expected_output: > + The issue should be identified as a missing Secret mount (frontend-api-keys)Using
>(folded) or an inline string keeps the file concise and avoids the trailing newline.
[ suggest_nitpick ]
14-16: Confirm tag downgrade is reflected in the rubricThe difficulty tag has been changed from presumably
mediumtoeasy. Make sure the scoring rubric and any CI filters that rely on difficulty levels are updated accordingly; otherwise this fixture may be skipped or mis-weighted during evaluation.
1-6: Double-check resource-name neutralityThe description still references a
frontenddeployment and a Secret calledfrontend-api-keys. While these are fairly generic, ensure they don’t leak hints (per coding guidelines) across other fixtures. Verify that no other test uses the same deployment or Secret names to retain cross-fixture uniqueness.tests/llm/fixtures/test_ask_holmes/77_liveness_probe_misconfiguration/test_case.yaml (1)
9-11: Tag change looks good
medium→easybetter reflects the relative difficulty of this scenario and doesn’t impact test behaviour.
No other issues spotted in this hunk.tests/llm/fixtures/test_ask_holmes/26_multi_container_logs/test_case.yaml (2)
22-22: Difficulty tag addition aligns with taxonomy – LGTM
The newmediumtag is consistent with the existing tagging convention and fills the gap between simple and advanced scenarios.
1-23: Double-check CI tolerance to droppedevaluationblocks
Theevaluation:section has been removed (not shown in diff). Please confirm the test loader no longer expects this key to avoid YAML-parsing or scoring regressions in CI.tests/llm/fixtures/test_ask_holmes/76_service_discovery_issue/test_case.yaml (2)
4-4: Tag enrichment looks good
Addingmediumcorrectly classifies the scenario’s complexity without affecting test behaviour.
1-12: Verify framework compatibility after removingevaluation
Ensure the harness that previously read theevaluation.correctnessfield is updated; otherwise tests may silently skip scoring or fail on missing keys.tests/llm/fixtures/test_ask_holmes/84_network_policy_blocking_traffic/test_case.yaml (2)
5-5: Difficulty tag addition approved
mediumaccurately reflects the investigative effort required.
1-13: Confirm removal of scoring metadata is reflected in CI
Please verify that automated evaluation pipelines no longer rely on the deletedevaluationstanza.tests/llm/fixtures/test_ask_holmes/78_resource_quota_exceeded/test_case.yaml (2)
4-4:mediumtag addition looks correct
No other behavioural impact detected.
1-12: Sanity-check test runner after evaluation removal
Validate that score aggregation scripts remain functional with the missingevaluationsection.tests/llm/fixtures/test_ask_holmes/80_pvc_storage_class_mismatch/test_case.yaml (2)
4-4: Tagging update accepted
The newmediumtag properly categorises this PVC troubleshooting case.
1-13: Please ensure scoring logic adapts to metadata change
Confirm that droppingevaluationdoes not break downstream analytics or result reporting.tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/test_case.yaml (1)
4-4: "medium" tag addition looks correct
Consistent with existing taxonomy (easy/medium/hard) used across other fixtures.tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_get_by_kind_in_cluster.txt (1)
3-5: LGTM – fixture content looks correct and neutral.Node names and label sets are realistic and don’t leak hints about the scheduling issue.
No review action required.
There was a problem hiding this comment.
Actionable comments posted: 7
🔭 Outside diff range comments (1)
tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml (1)
1-4: Rename namespace to app-85 for consistencyVerified that no other fixture uses
app-85, so it’s safe to update. Please rename the namespace and any references accordingly.• tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml: change
namespace-85→app-85
• tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/test_case.yaml: update anynamespace-85references toapp-85--- a/tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml +++ b/tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml @@ metadata: - name: namespace-85 + name: app-85
🧹 Nitpick comments (2)
tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml (1)
24-25: Inline comment leaks the root causeThe fixture comment explicitly explains the missing resource requests/limits, potentially giving away the investigation’s answer when someone reads the manifest. Remove the explanatory comment to keep fixtures neutral.
- # Note: Missing resource requests/limits - this is intentional to match the expected output + # intentionally omit resource requests/limits for test purposestests/llm/fixtures/test_ask_holmes/13a_pending_node_selector_basic/test_case.yaml (1)
8-9: Cleanup could be simplified
Because everything lives inapp-13a, deleting the namespace already garbage-collects namespaced resources.
You can drop the explicitkubectl delete -f …step and save test time:- kubectl delete -f https://raw.githubusercontent.com/robusta-dev/kubernetes-demos/main/pending_pods/pending_pod_node_selector.yaml -n app-13a || true - kubectl delete namespace app-13a || true + kubectl delete namespace app-13a || true
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (21)
tests/llm/fixtures/test_ask_holmes/13a_pending_node_selector_basic/kubectl_describe.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/13a_pending_node_selector_basic/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_describe.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_describe_PersistentVolumeClaim_redis-standard-pvc_app-24a_2.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_describe_persistentvolumeclaim_redis-standard-pvc_app-24a.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_describe_pod_redis-747ffc844f-f9ghd_app-24a.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_get_by_name_PersistentVolumeClaim_redis-standard-pvc_app-24a_2.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_get_by_name_persistentvolumeclaim_redis-standard-pvc_app-24a.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_describe_PersistentVolumeClaim_redis-standard-pvc_app-24b_2.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_describe_persistentvolumeclaim_redis-standard-pvc_app-24b.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_describe_pod_redis-747ffc844f-f9ghd_app-24b.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_find_resource_pod_redis.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_get_by_name_PersistentVolumeClaim_redis-standard-pvc_app-24b_2.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_get_by_name_persistentvolumeclaim_redis-standard-pvc_app-24b.txt(1 hunks)tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/test_case.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml(1 hunks)tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/test_case.yaml(1 hunks)
✅ Files skipped from review due to trivial changes (11)
- tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_describe_pod_redis-747ffc844f-f9ghd_app-24a.txt
- tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/82_pod_anti_affinity_conflict/manifest.yaml
- tests/llm/fixtures/test_ask_holmes/13a_pending_node_selector_basic/kubectl_describe.txt
- tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_find_resource_pod_redis.txt
- tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_get_by_name_persistentvolumeclaim_redis-standard-pvc_app-24b.txt
- tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_describe_persistentvolumeclaim_redis-standard-pvc_app-24b.txt
- tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_describe_pod_redis-747ffc844f-f9ghd_app-24b.txt
- tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_describe_PersistentVolumeClaim_redis-standard-pvc_app-24b_2.txt
🚧 Files skipped from review as they are similar to previous changes (3)
- tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/test_case.yaml
- tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/kubectl_describe.txt
- tests/llm/fixtures/test_ask_holmes/13b_pending_node_selector_detailed/test_case.yaml
🧰 Additional context used
📓 Path-based instructions (1)
tests/llm/fixtures/**/*
📄 CodeRabbit Inference Engine (CLAUDE.md)
tests/llm/fixtures/**/*: Mock data: tests/llm/fixtures/{test_name}/
All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Files:
tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_describe_PersistentVolumeClaim_redis-standard-pvc_app-24a_2.txttests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_get_by_name_persistentvolumeclaim_redis-standard-pvc_app-24a.txttests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_describe_persistentvolumeclaim_redis-standard-pvc_app-24a.txttests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_get_by_name_PersistentVolumeClaim_redis-standard-pvc_app-24a_2.txttests/llm/fixtures/test_ask_holmes/13a_pending_node_selector_basic/test_case.yamltests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_get_by_name_PersistentVolumeClaim_redis-standard-pvc_app-24b_2.txttests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml
🧠 Learnings (8)
📓 Common learnings
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Mock data: tests/llm/fixtures/{test_name}/
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/**/*.py : Each LLM test must use a dedicated namespace app-<testid> (e.g., app-01, app-02) to prevent conflicts when tests run simultaneously
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/**/*.py : Complex investigations should have LLM evaluation tests
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: LLM evaluation tests run automatically in CI
tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_describe_PersistentVolumeClaim_redis-standard-pvc_app-24a_2.txt (1)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_get_by_name_persistentvolumeclaim_redis-standard-pvc_app-24a.txt (2)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_describe_persistentvolumeclaim_redis-standard-pvc_app-24a.txt (1)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
tests/llm/fixtures/test_ask_holmes/24a_misconfigured_pvc_basic/kubectl_get_by_name_PersistentVolumeClaim_redis-standard-pvc_app-24a_2.txt (3)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
tests/llm/fixtures/test_ask_holmes/13a_pending_node_selector_basic/test_case.yaml (5)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/**/*.py : Each LLM test must use a dedicated namespace app- (e.g., app-01, app-02) to prevent conflicts when tests run simultaneously
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Mock data: tests/llm/fixtures/{test_name}/
tests/llm/fixtures/test_ask_holmes/24b_misconfigured_pvc_detailed/kubectl_get_by_name_PersistentVolumeClaim_redis-standard-pvc_app-24b_2.txt (3)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Mock data: tests/llm/fixtures/{test_name}/
tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml (5)
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : All pod names must be unique across tests (e.g., giant-narwhal, blue-whale, sea-turtle) - never reuse pod names between tests
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Never use resource names that hint at the problem or expected behavior (e.g., avoid broken-pod, test-project-that-does-not-exist, crashloop-app). Use neutral names that don't give away what the LLM should discover
Learnt from: Sheeproid
PR: #586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/fixtures/**/* : Mock data: tests/llm/fixtures/{test_name}/
Learnt from: CR
PR: robusta-dev/holmesgpt#0
File: CLAUDE.md:0-0
Timestamp: 2025-07-27T14:10:31.590Z
Learning: Applies to tests/llm/**/*.py : Each LLM test must use a dedicated namespace app- (e.g., app-01, app-02) to prevent conflicts when tests run simultaneously
🪛 Checkov (3.2.334)
tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml
[MEDIUM] 6-26: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[MEDIUM] 6-26: Minimize the admission of root containers
(CKV_K8S_23)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (3)
- GitHub Check: llm_evals
- GitHub Check: Pre-commit checks
- GitHub Check: Pre-commit checks
🔇 Additional comments (2)
tests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling/manifest.yaml (2)
8-19: Good choice of unique, neutral deployment name
golden-eaglesatisfies the “unique-across-tests” and “no-hints” rules.
30-36: No duplicate “golden-eagle” identifiers found across tests
Search for “golden-eagle” returned only entries withintests/llm/fixtures/test_ask_holmes/85_hpa_not_scaling, confirming that both the deployment (golden-eagle) and the HPA (golden-eagle-hpa) names are unique across the suite.
…/kubectl_describe_PersistentVolumeClaim_redis-standard-pvc_app-24a_2.txt Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
…/kubectl_describe_persistentvolumeclaim_redis-standard-pvc_app-24a.txt Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
…/kubectl_get_by_name_persistentvolumeclaim_redis-standard-pvc_app-24a.txt Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
…/kubectl_get_by_name_PersistentVolumeClaim_redis-standard-pvc_app-24a_2.txt Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
…led/kubectl_get_by_name_PersistentVolumeClaim_redis-standard-pvc_app-24b_2.txt Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
No description provided.