Repository navigation
ad 176_network_policy_blocking_traffic_no_runbooks based on test 84 - #1236
Conversation
WalkthroughAdds a new Kubernetes test fixture that deploys a backend and probing frontend in namespace app-176, plus a NetworkPolicy that restricts ingress to the backend; includes test orchestration to validate that the frontend experiences connection timeouts and configuration to disable runbook/internet tools. Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Suggested reviewers
Pre-merge checks❌ Failed checks (1 inconclusive)
✅ Passed checks (2 passed)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 7
📜 Review details
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
📥 Commits
Reviewing files that changed from the base of the PR and between 9202490 and 0da6436e717184f6704069b2e33ac36fdb74e6df.
📒 Files selected for processing (5)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
🧰 Additional context used
📓 Path-based instructions (4)
tests/**
📄 CodeRabbit inference engine (CLAUDE.md)
Tests: Match source structure under tests/
Files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
tests/llm/**/*.yaml
📄 CodeRabbit inference engine (CLAUDE.md)
tests/llm/**/*.yaml: ALWAYS use Secrets for scripts in eval tests, not inline manifests or ConfigMaps, to prevent code visibility with kubectl describe
All pod names must be unique across LLM tests - never reuse pod names between tests
Never use resource names that hint at the problem or expected behavior in evals - use neutral names that don't give away what the LLM should discover
Files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
tests/llm/**
📄 CodeRabbit inference engine (CLAUDE.md)
tests/llm/**: Each LLM test must use a dedicated namespace app- to prevent conflicts when tests run simultaneously
No fake/obvious logs in eval scenarios - avoid logs like 'Memory usage stabilized at 800MB'
No hints in eval filenames - use realistic names like 'training_pipeline.py' instead of 'disk_consumer.py'
Implement full architecture even if complex in evals (e.g., use Loki for log aggregation properly, not simplified alternatives)
Files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
**/*toolsets.yaml
📄 CodeRabbit inference engine (CLAUDE.md)
All toolset-specific configuration must go under a 'config' field in toolsets.yaml, not at the top level
Files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
🧠 Learnings (11)
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to holmes/plugins/toolsets/** : Toolsets: holmes/plugins/toolsets/{name}.yaml or {name}/ directory structure
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Use generic function descriptions in workflow steps (e.g., 'execute a command to test network connectivity') rather than tool-specific names to enable mapping to available tools
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to **/*toolsets.yaml : All toolset-specific configuration must go under a 'config' field in toolsets.yaml, not at the top level
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Place runbook files in category folders under holmes/plugins/runbooks/ and use consistent lowercase filenames with hyphens (e.g., dns-resolution-troubleshooting.md)
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Workflow section with numbered sequential steps containing Action, Function Description, Parameters, Expected Output, and Success/Failure Criteria
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Recommended Remediation Steps section with Immediate Actions, Permanent Solutions, Verification Steps, Documentation References, Escalation Criteria, and Post-Remediation Monitoring
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include conditional logic (IF/ELSE) in workflow steps when branching is required based on diagnostic findings
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Goal section with Primary Objective, Scope, Agent Mandate, and Expected Outcome
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include verification steps in workflow to confirm each diagnostic action was successful before proceeding
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Synthesize Findings section with Data Correlation, Pattern Recognition, Prioritization Logic, Evidence Requirements, and Example Scenarios
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/before_test.sh : Never use bare kubectl wait immediately after creating resources - implement retry loop to handle race conditions
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
🪛 Checkov (3.2.334)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml
[medium] 7-34: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[medium] 7-34: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
[medium] 7-34: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[medium] 7-34: Minimize the admission of root containers
(CKV_K8S_23)
[medium] 49-87: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[medium] 49-87: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
[medium] 2-39: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[medium] 2-39: Minimize the admission of root containers
(CKV_K8S_23)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
- GitHub Check: llm_evals
- GitHub Check: build (3.12)
- GitHub Check: build (3.10)
- GitHub Check: build (3.11)
- GitHub Check: build
🔇 Additional comments (3)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml (1)
1-39: LGTM!The frontend deployment correctly uses namespace app-176, has neutral resource naming, and implements the connectivity test scenario appropriately. The specific error message format is necessary for test log validation in test_case.yaml.
Note: Static analysis warnings about privilege escalation and root containers are false positives in the context of minimal test fixtures.
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml (1)
1-37: LGTM!The test orchestration properly implements retry loops for resource readiness, includes robust error detection with explicit success/failure criteria, and handles cleanup appropriately. The before_test script follows best practices from learnings by avoiding bare kubectl wait and implementing proper polling.
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml (1)
1-5: Structure is correct and follows the established pattern.The
enabledflag belongs at the toolset level alongsideconfig(when present), not nested within it. All toolsets.yaml files in the codebase follow this structure whereenabledandconfigare siblings under each toolset name. No changes needed.
|
Dev Docker images are ready for this commit: Use this tag to pull the image for testing. gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:8f79788
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:8f79788 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:8f79788
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:8f79788Patch Helm values in one line (choose the chart you use):
helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:8f79788
helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:8f79788 |
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
0da6436 to
497a8b7
Compare
There was a problem hiding this comment.
Actionable comments posted: 1
♻️ Duplicate comments (5)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml (2)
48-48: Remove hint comment that reveals the problem.The comment "Network Policy that blocks frontend->backend traffic" explicitly reveals what Holmes should discover independently. Per coding guidelines, eval scenarios must not include hints about the problem.
🔎 Proposed fix
-# Network Policy that blocks frontend->backend traffic +# Network Policy
64-64: Remove hint comment that reveals the policy restriction.The inline comment "Only allows traffic from pods with tier=backend" explains exactly why traffic is blocked, undermining the eval objective. Holmes should diagnose this from the NetworkPolicy configuration itself.
🔎 Proposed fix
- tier: backend # Only allows traffic from pods with tier=backend + tier: backendtests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml (3)
88-88: Remove hint comment that reveals the problem.The comment "Network Policy that blocks frontend->backend traffic" explicitly states what Holmes should discover independently. Per coding guidelines, eval scenarios must not include hints about the problem.
🔎 Proposed fix
-# Network Policy that blocks frontend->backend traffic +# Network Policy
104-104: Remove hint comment that reveals the policy restriction.The inline comment "Only allows traffic from pods with tier=backend" explains exactly why traffic is blocked, undermining the eval objective. Holmes should diagnose this from the NetworkPolicy configuration itself.
🔎 Proposed fix
- tier: backend # Only allows traffic from pods with tier=backend + tier: backend
1-107: Clarify the purpose of manifest.yaml or remove it.The test_case.yaml applies backend.yaml and frontend.yaml separately and does not reference manifest.yaml. This combined manifest appears unused and creates confusion about which files are actually deployed during the test. Either remove this file or add a clear comment explaining its purpose (e.g., for manual testing or alternative test execution paths).
📜 Review details
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
📥 Commits
Reviewing files that changed from the base of the PR and between 0da6436e717184f6704069b2e33ac36fdb74e6df and 497a8b7.
📒 Files selected for processing (5)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
🧰 Additional context used
📓 Path-based instructions (4)
tests/**
📄 CodeRabbit inference engine (CLAUDE.md)
Tests: Match source structure under tests/
Files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
tests/llm/**/*.yaml
📄 CodeRabbit inference engine (CLAUDE.md)
tests/llm/**/*.yaml: ALWAYS use Secrets for scripts in eval tests, not inline manifests or ConfigMaps, to prevent code visibility with kubectl describe
All pod names must be unique across LLM tests - never reuse pod names between tests
Never use resource names that hint at the problem or expected behavior in evals - use neutral names that don't give away what the LLM should discover
Files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
tests/llm/**
📄 CodeRabbit inference engine (CLAUDE.md)
tests/llm/**: Each LLM test must use a dedicated namespace app- to prevent conflicts when tests run simultaneously
No fake/obvious logs in eval scenarios - avoid logs like 'Memory usage stabilized at 800MB'
No hints in eval filenames - use realistic names like 'training_pipeline.py' instead of 'disk_consumer.py'
Implement full architecture even if complex in evals (e.g., use Loki for log aggregation properly, not simplified alternatives)
Files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yamltests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
**/*toolsets.yaml
📄 CodeRabbit inference engine (CLAUDE.md)
All toolset-specific configuration must go under a 'config' field in toolsets.yaml, not at the top level
Files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
🧠 Learnings (15)
📓 Common learnings
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/before_test.sh : Never use bare kubectl wait immediately after creating resources - implement retry loop to handle race conditions
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to holmes/plugins/toolsets/** : Toolsets: holmes/plugins/toolsets/{name}.yaml or {name}/ directory structure
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Use generic function descriptions in workflow steps (e.g., 'execute a command to test network connectivity') rather than tool-specific names to enable mapping to available tools
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Place runbook files in category folders under holmes/plugins/runbooks/ and use consistent lowercase filenames with hyphens (e.g., dns-resolution-troubleshooting.md)
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Workflow section with numbered sequential steps containing Action, Function Description, Parameters, Expected Output, and Success/Failure Criteria
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Recommended Remediation Steps section with Immediate Actions, Permanent Solutions, Verification Steps, Documentation References, Escalation Criteria, and Post-Remediation Monitoring
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include conditional logic (IF/ELSE) in workflow steps when branching is required based on diagnostic findings
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Goal section with Primary Objective, Scope, Agent Mandate, and Expected Outcome
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include verification steps in workflow to confirm each diagnostic action was successful before proceeding
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Synthesize Findings section with Data Correlation, Pattern Recognition, Prioritization Logic, Evidence Requirements, and Example Scenarios
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/before_test.sh : Never use bare kubectl wait immediately after creating resources - implement retry loop to handle race conditions
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/** : Each LLM test must use a dedicated namespace app-<testid> to prevent conflicts when tests run simultaneously
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/*.yaml : All pod names must be unique across LLM tests - never reuse pod names between tests
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/*.yaml : ALWAYS use Secrets for scripts in eval tests, not inline manifests or ConfigMaps, to prevent code visibility with kubectl describe
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/*.yaml : Never use resource names that hint at the problem or expected behavior in evals - use neutral names that don't give away what the LLM should discover
Applied to files:
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
🪛 Checkov (3.2.334)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml
[medium] 7-34: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[medium] 7-34: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
[medium] 7-34: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[medium] 7-34: Minimize the admission of root containers
(CKV_K8S_23)
[medium] 49-87: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[medium] 49-87: Minimize the admission of root containers
(CKV_K8S_23)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
[medium] 2-39: Containers should not run with allowPrivilegeEscalation
(CKV_K8S_20)
[medium] 2-39: Minimize the admission of root containers
(CKV_K8S_23)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
- GitHub Check: llm_evals
- GitHub Check: build (3.11)
- GitHub Check: build (3.10)
- GitHub Check: build (3.12)
- GitHub Check: build
🔇 Additional comments (4)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml (1)
1-5: LGTM! Toolset configuration is correct.The toolsets configuration properly disables both runbook and internet tools for this eval scenario. The comment on line 4 provides useful context about Holmes' fallback behavior without revealing the actual network policy issue.
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml (3)
1-7: LGTM! User prompt and test configuration are well-structured.The user prompt is neutral and doesn't reveal the issue. The expected output appropriately describes the NetworkPolicy behavior for test validation. Tags correctly categorize this as a hard Kubernetes networking scenario.
8-35: LGTM! Test orchestration follows best practices.The before_test script correctly:
- Applies backend resources first and waits for readiness using a retry loop (avoiding bare kubectl wait per learnings)
- Applies frontend resources after backend is ready
- Validates that the expected timeout error appears in logs before proceeding
- Provides clear success/failure feedback with appropriate diagnostics
36-37: LGTM! Cleanup is appropriate.The after_test cleanup properly deletes the test namespace with || true to prevent cleanup failures from affecting test results.
|
Dev Docker images are ready for this commit:
Use either tag to pull the image for testing. |
…1236) Signed-off-by: Mohse Morad <moshemorad12340@gmail.com>
…olmesGPT#1236) Signed-off-by: Filip Grebowski <grebowskifilip@gmail.com>
✏️ Tip: You can customize this high-level summary in your review settings.