Skip to content

ad 176_network_policy_blocking_traffic_no_runbooks based on test 84 - #1236

Merged
Sheeproid merged 3 commits into
masterfrom
eval-networkpolicy-harder
Dec 24, 2025
Merged

Sheeproid merged 3 commits into
masterfrom
eval-networkpolicy-harder

Conversation

@Sheeproid

@Sheeproid Sheeproid commented Dec 24, 2025 •

Copy link
Copy Markdown
Collaborator
  • Tests
    • Added an end-to-end test that deploys the resources, verifies the frontend’s repeated connection timeouts to the backend, and cleans up the namespace.

✏️ Tip: You can customize this high-level summary in your review settings.

@Sheeproid
Sheeproid requested a review from moshemorad December 24, 2025 08:33
@coderabbitai

coderabbitai Bot commented Dec 24, 2025 •

Copy link
Copy Markdown
Contributor

Walkthrough

Adds a new Kubernetes test fixture that deploys a backend and probing frontend in namespace app-176, plus a NetworkPolicy that restricts ingress to the backend; includes test orchestration to validate that the frontend experiences connection timeouts and configuration to disable runbook/internet tools.

Changes

Cohort / File(s) Summary
Kubernetes Manifests
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml, tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml, tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
New manifests: Namespace app-176; backend Deployment (nginx:alpine, 1 replica, resources), backend-service Service (port 80); frontend Deployment (busybox loop that repeatedly curls backend-service:80, resource limits). NetworkPolicy backend-network-policy restricts ingress to pods labeled tier: backend on TCP/80.
Test Orchestration
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
New test case that applies backend, waits for pod readiness, applies frontend, polls frontend logs for the exact timeout error string (ERROR: Connection timeout to backend-service!) within retries, and deletes namespace on teardown.
Tool Configuration
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
New toolsets config disabling runbook and internet toolsets (both enabled: false).

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Suggested reviewers

  • moshemorad
  • aantn

Pre-merge checks

❌ Failed checks (1 inconclusive)
Check name Status Explanation Resolution
Title check ❓ Inconclusive The title is partially related to the changeset. It mentions adding test case 176 about network policy blocking traffic without runbooks and references test 84, but uses a typo ('ad' instead of 'Add') and lacks clarity about the main purpose. Clarify the title by fixing the typo and being more explicit about what the test fixture demonstrates, such as 'Add test fixture 176: network policy blocking traffic scenario without runbooks' or similar.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

Comment @coderabbitai help to get the list of available commands and usage tips.

@Sheeproid
Sheeproid enabled auto-merge (squash) December 24, 2025 08:33

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 7

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 9202490 and 0da6436e717184f6704069b2e33ac36fdb74e6df.

📒 Files selected for processing (5)
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
🧰 Additional context used
📓 Path-based instructions (4)
tests/**

📄 CodeRabbit inference engine (CLAUDE.md)

Tests: Match source structure under tests/

Files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
tests/llm/**/*.yaml

📄 CodeRabbit inference engine (CLAUDE.md)

tests/llm/**/*.yaml: ALWAYS use Secrets for scripts in eval tests, not inline manifests or ConfigMaps, to prevent code visibility with kubectl describe
All pod names must be unique across LLM tests - never reuse pod names between tests
Never use resource names that hint at the problem or expected behavior in evals - use neutral names that don't give away what the LLM should discover

Files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
tests/llm/**

📄 CodeRabbit inference engine (CLAUDE.md)

tests/llm/**: Each LLM test must use a dedicated namespace app- to prevent conflicts when tests run simultaneously
No fake/obvious logs in eval scenarios - avoid logs like 'Memory usage stabilized at 800MB'
No hints in eval filenames - use realistic names like 'training_pipeline.py' instead of 'disk_consumer.py'
Implement full architecture even if complex in evals (e.g., use Loki for log aggregation properly, not simplified alternatives)

Files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
**/*toolsets.yaml

📄 CodeRabbit inference engine (CLAUDE.md)

All toolset-specific configuration must go under a 'config' field in toolsets.yaml, not at the top level

Files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
🧠 Learnings (11)
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to holmes/plugins/toolsets/** : Toolsets: holmes/plugins/toolsets/{name}.yaml or {name}/ directory structure

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Use generic function descriptions in workflow steps (e.g., 'execute a command to test network connectivity') rather than tool-specific names to enable mapping to available tools

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to **/*toolsets.yaml : All toolset-specific configuration must go under a 'config' field in toolsets.yaml, not at the top level

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Place runbook files in category folders under holmes/plugins/runbooks/ and use consistent lowercase filenames with hyphens (e.g., dns-resolution-troubleshooting.md)

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Workflow section with numbered sequential steps containing Action, Function Description, Parameters, Expected Output, and Success/Failure Criteria

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Recommended Remediation Steps section with Immediate Actions, Permanent Solutions, Verification Steps, Documentation References, Escalation Criteria, and Post-Remediation Monitoring

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include conditional logic (IF/ELSE) in workflow steps when branching is required based on diagnostic findings

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Goal section with Primary Objective, Scope, Agent Mandate, and Expected Outcome

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include verification steps in workflow to confirm each diagnostic action was successful before proceeding

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Synthesize Findings section with Data Correlation, Pattern Recognition, Prioritization Logic, Evidence Requirements, and Example Scenarios

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/before_test.sh : Never use bare kubectl wait immediately after creating resources - implement retry loop to handle race conditions

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
🪛 Checkov (3.2.334)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml

[medium] 7-34: Containers should not run with allowPrivilegeEscalation

(CKV_K8S_20)


[medium] 7-34: Minimize the admission of root containers

(CKV_K8S_23)

tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml

[medium] 7-34: Containers should not run with allowPrivilegeEscalation

(CKV_K8S_20)


[medium] 7-34: Minimize the admission of root containers

(CKV_K8S_23)


[medium] 49-87: Containers should not run with allowPrivilegeEscalation

(CKV_K8S_20)


[medium] 49-87: Minimize the admission of root containers

(CKV_K8S_23)

tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml

[medium] 2-39: Containers should not run with allowPrivilegeEscalation

(CKV_K8S_20)


[medium] 2-39: Minimize the admission of root containers

(CKV_K8S_23)

⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
  • GitHub Check: llm_evals
  • GitHub Check: build (3.12)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.11)
  • GitHub Check: build
🔇 Additional comments (3)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml (1)

1-39: LGTM!

The frontend deployment correctly uses namespace app-176, has neutral resource naming, and implements the connectivity test scenario appropriately. The specific error message format is necessary for test log validation in test_case.yaml.

Note: Static analysis warnings about privilege escalation and root containers are false positives in the context of minimal test fixtures.

tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml (1)

1-37: LGTM!

The test orchestration properly implements retry loops for resource readiness, includes robust error detection with explicit success/failure criteria, and handles cleanup appropriately. The before_test script follows best practices from learnings by avoiding bare kubectl wait and implementing proper polling.

tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml (1)

1-5: Structure is correct and follows the established pattern.

The enabled flag belongs at the toolset level alongside config (when present), not nested within it. All toolsets.yaml files in the codebase follow this structure where enabled and config are siblings under each toolset name. No changes needed.

@github-actions

github-actions Bot commented Dec 24, 2025 •

Copy link
Copy Markdown
Contributor

Dev Docker images are ready for this commit:

Use this tag to pull the image for testing.

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:8f79788
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:8f79788 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:8f79788
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:8f79788

Patch Helm values in one line (choose the chart you use):

  • HolmesGPT chart:
helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:8f79788
  • Robusta wrapper chart:
helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:8f79788

@Sheeproid
Sheeproid disabled auto-merge December 24, 2025 09:21
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
@Sheeproid
Sheeproid force-pushed the eval-networkpolicy-harder branch from 0da6436 to 497a8b7 Compare December 24, 2025 09:24
@Sheeproid
Sheeproid enabled auto-merge (squash) December 24, 2025 09:30

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (5)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml (2)

48-48: Remove hint comment that reveals the problem.

The comment "Network Policy that blocks frontend->backend traffic" explicitly reveals what Holmes should discover independently. Per coding guidelines, eval scenarios must not include hints about the problem.

🔎 Proposed fix
-# Network Policy that blocks frontend->backend traffic
+# Network Policy

64-64: Remove hint comment that reveals the policy restriction.

The inline comment "Only allows traffic from pods with tier=backend" explains exactly why traffic is blocked, undermining the eval objective. Holmes should diagnose this from the NetworkPolicy configuration itself.

🔎 Proposed fix
-          tier: backend  # Only allows traffic from pods with tier=backend
+          tier: backend
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml (3)

88-88: Remove hint comment that reveals the problem.

The comment "Network Policy that blocks frontend->backend traffic" explicitly states what Holmes should discover independently. Per coding guidelines, eval scenarios must not include hints about the problem.

🔎 Proposed fix
-# Network Policy that blocks frontend->backend traffic
+# Network Policy

104-104: Remove hint comment that reveals the policy restriction.

The inline comment "Only allows traffic from pods with tier=backend" explains exactly why traffic is blocked, undermining the eval objective. Holmes should diagnose this from the NetworkPolicy configuration itself.

🔎 Proposed fix
-          tier: backend  # Only allows traffic from pods with tier=backend
+          tier: backend

1-107: Clarify the purpose of manifest.yaml or remove it.

The test_case.yaml applies backend.yaml and frontend.yaml separately and does not reference manifest.yaml. This combined manifest appears unused and creates confusion about which files are actually deployed during the test. Either remove this file or add a clear comment explaining its purpose (e.g., for manual testing or alternative test execution paths).

📜 Review details

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 0da6436e717184f6704069b2e33ac36fdb74e6df and 497a8b7.

📒 Files selected for processing (5)
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
🧰 Additional context used
📓 Path-based instructions (4)
tests/**

📄 CodeRabbit inference engine (CLAUDE.md)

Tests: Match source structure under tests/

Files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
tests/llm/**/*.yaml

📄 CodeRabbit inference engine (CLAUDE.md)

tests/llm/**/*.yaml: ALWAYS use Secrets for scripts in eval tests, not inline manifests or ConfigMaps, to prevent code visibility with kubectl describe
All pod names must be unique across LLM tests - never reuse pod names between tests
Never use resource names that hint at the problem or expected behavior in evals - use neutral names that don't give away what the LLM should discover

Files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
tests/llm/**

📄 CodeRabbit inference engine (CLAUDE.md)

tests/llm/**: Each LLM test must use a dedicated namespace app- to prevent conflicts when tests run simultaneously
No fake/obvious logs in eval scenarios - avoid logs like 'Memory usage stabilized at 800MB'
No hints in eval filenames - use realistic names like 'training_pipeline.py' instead of 'disk_consumer.py'
Implement full architecture even if complex in evals (e.g., use Loki for log aggregation properly, not simplified alternatives)

Files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml
**/*toolsets.yaml

📄 CodeRabbit inference engine (CLAUDE.md)

All toolset-specific configuration must go under a 'config' field in toolsets.yaml, not at the top level

Files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
🧠 Learnings (15)
📓 Common learnings
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/before_test.sh : Never use bare kubectl wait immediately after creating resources - implement retry loop to handle race conditions
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to holmes/plugins/toolsets/** : Toolsets: holmes/plugins/toolsets/{name}.yaml or {name}/ directory structure

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Use generic function descriptions in workflow steps (e.g., 'execute a command to test network connectivity') rather than tool-specific names to enable mapping to available tools

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Place runbook files in category folders under holmes/plugins/runbooks/ and use consistent lowercase filenames with hyphens (e.g., dns-resolution-troubleshooting.md)

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Workflow section with numbered sequential steps containing Action, Function Description, Parameters, Expected Output, and Success/Failure Criteria

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Recommended Remediation Steps section with Immediate Actions, Permanent Solutions, Verification Steps, Documentation References, Escalation Criteria, and Post-Remediation Monitoring

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include conditional logic (IF/ELSE) in workflow steps when branching is required based on diagnostic findings

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Goal section with Primary Objective, Scope, Agent Mandate, and Expected Outcome

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Include verification steps in workflow to confirm each diagnostic action was successful before proceeding

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:57.161Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:57.161Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : Runbook must include Synthesize Findings section with Data Correlation, Pattern Recognition, Prioritization Logic, Evidence Requirements, and Example Scenarios

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/before_test.sh : Never use bare kubectl wait immediately after creating resources - implement retry loop to handle race conditions

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/** : Each LLM test must use a dedicated namespace app-<testid> to prevent conflicts when tests run simultaneously

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/*.yaml : All pod names must be unique across LLM tests - never reuse pod names between tests

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/*.yaml : ALWAYS use Secrets for scripts in eval tests, not inline manifests or ConfigMaps, to prevent code visibility with kubectl describe

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
📚 Learning: 2025-12-21T13:17:48.366Z
Learnt from: CR
Repo: HolmesGPT/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-21T13:17:48.366Z
Learning: Applies to tests/llm/**/*.yaml : Never use resource names that hint at the problem or expected behavior in evals - use neutral names that don't give away what the LLM should discover

Applied to files:

  • tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml
🪛 Checkov (3.2.334)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/backend.yaml

[medium] 7-34: Containers should not run with allowPrivilegeEscalation

(CKV_K8S_20)


[medium] 7-34: Minimize the admission of root containers

(CKV_K8S_23)

tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/manifest.yaml

[medium] 7-34: Containers should not run with allowPrivilegeEscalation

(CKV_K8S_20)


[medium] 7-34: Minimize the admission of root containers

(CKV_K8S_23)


[medium] 49-87: Containers should not run with allowPrivilegeEscalation

(CKV_K8S_20)


[medium] 49-87: Minimize the admission of root containers

(CKV_K8S_23)

tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/frontend.yaml

[medium] 2-39: Containers should not run with allowPrivilegeEscalation

(CKV_K8S_20)


[medium] 2-39: Minimize the admission of root containers

(CKV_K8S_23)

⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (5)
  • GitHub Check: llm_evals
  • GitHub Check: build (3.11)
  • GitHub Check: build (3.10)
  • GitHub Check: build (3.12)
  • GitHub Check: build
🔇 Additional comments (4)
tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/toolsets.yaml (1)

1-5: LGTM! Toolset configuration is correct.

The toolsets configuration properly disables both runbook and internet tools for this eval scenario. The comment on line 4 provides useful context about Holmes' fallback behavior without revealing the actual network policy issue.

tests/llm/fixtures/test_ask_holmes/176_network_policy_blocking_traffic_no_runbooks/test_case.yaml (3)

1-7: LGTM! User prompt and test configuration are well-structured.

The user prompt is neutral and doesn't reveal the issue. The expected output appropriately describes the NetworkPolicy behavior for test validation. Tags correctly categorize this as a hard Kubernetes networking scenario.


8-35: LGTM! Test orchestration follows best practices.

The before_test script correctly:

  • Applies backend resources first and waits for readiness using a retry loop (avoiding bare kubectl wait per learnings)
  • Applies frontend resources after backend is ready
  • Validates that the expected timeout error appears in logs before proceeding
  • Provides clear success/failure feedback with appropriate diagnostics

36-37: LGTM! Cleanup is appropriate.

The after_test cleanup properly deletes the test namespace with || true to prevent cleanup failures from affecting test results.

@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

Results of HolmesGPT evals

  • ask_holmes: 30/37 test cases were successful, 1 regressions, 2 setup failures, 3 mock failures
Test suite Test case Status
ask 01_how_many_pods ✅
ask 02_what_is_wrong_with_pod ✅
ask 04_related_k8s_events ✅
ask 05_image_version ✅
ask 09_crashpod ✅
ask 10_image_pull_backoff ✅
ask 110_k8s_events_image_pull ✅
ask 11_init_containers ✅
ask 13a_pending_node_selector_basic ✅
ask 14_pending_resources ✅
ask 15_failed_readiness_probe ✅
ask 163_compaction_follow_up ✅
ask 17_oom_kill 🚧
ask 18_oom_kill_from_issues_history ✅
ask 19_detect_missing_app_details ✅
ask 20_long_log_file_search ✅
ask 24_misconfigured_pvc ✅
ask 24a_misconfigured_pvc_basic ✅
ask 28_permissions_error 🚧
ask 39_failed_toolset ✅
ask 41_setup_argo ✅
ask 42_dns_issues_steps_new_tools ⚠️
ask 43_current_datetime_from_prompt ✅
ask 45_fetch_deployment_logs_simple ✅
ask 51_logs_summarize_errors ✅
ask 53_logs_find_term ✅
ask 54_not_truncated_when_getting_pods ✅
ask 59_label_based_counting ✅
ask 60_count_less_than ✅
ask 61_exact_match_counting ✅
ask 63_fetch_error_logs_no_errors ✅
ask 79_configmap_mount_issue ❌
ask 83_secret_not_found ✅
ask 86_configmap_like_but_secret ✅
ask 93_calling_datadog[0] 🔧
ask 93_calling_datadog[1] 🔧
ask 93_calling_datadog[2] 🔧

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

@Sheeproid
Sheeproid merged commit 4697741 into master Dec 24, 2025
10 of 11 checks passed
@Sheeproid
Sheeproid deleted the eval-networkpolicy-harder branch December 24, 2025 12:09
moshemorad pushed a commit that referenced this pull request Dec 25, 2025
…1236)

Signed-off-by: Mohse Morad <moshemorad12340@gmail.com>
FilipGrebowski pushed a commit to FilipGrebowski/holmesgpt that referenced this pull request Dec 27, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants