Skip to content

Add 5 bash resource checker evals for kubectl exec scenarios - #1456

Open
aantn wants to merge 8 commits into
masterfrom
claude/bash-resource-checker-evals-fSHoE
Open

aantn wants to merge 8 commits into
masterfrom
claude/bash-resource-checker-evals-fSHoE

Conversation

@aantn

@aantn aantn commented Jan 30, 2026 •

Copy link
Copy Markdown
Collaborator

Implements 5 new eval tests that require bash scripting with kubectl exec:

  • 211: Process memory audit (single pod, 100 processes)
  • 212: Container disk usage audit (single pod, 100 files)
  • 213: Network connection audit (single pod, 100 connections)
  • 214: Multi-pod disk check (10 pods)
  • 215: Multi-pod connectivity test (10 pods + NetworkPolicy)

These tests verify Holmes can use the bash toolset to inspect in-container
state that no existing toolset can access (processes, filesystem, network).

Also adds bash-resource-checker pytest marker to pyproject.toml.

https://claude.ai/code/session_01H2cRJG4a2oz7qB2gKPhrQj
Signed-off-by: Claude noreply@anthropic.com

Summary by CodeRabbit

  • Tests
    • Added a pytest marker for bash resource-checker tests.
    • Added end-to-end fixtures for bash-based Kubernetes validations: process memory audits, container disk audits, multi-pod disk checks, network connection audits, and multi-pod connectivity checks.
    • Fixtures include automated setup/cleanup and enforce restricted bash toolset usage (kubectl exec/get).

✏️ Tip: You can customize this high-level summary in your review settings.

Implements 5 new eval tests that require bash scripting with kubectl exec:
- 211: Process memory audit (single pod, 100 processes)
- 212: Container disk usage audit (single pod, 100 files)
- 213: Network connection audit (single pod, 100 connections)
- 214: Multi-pod disk check (10 pods)
- 215: Multi-pod connectivity test (10 pods + NetworkPolicy)

These tests verify Holmes can use the bash toolset to inspect in-container
state that no existing toolset can access (processes, filesystem, network).

Also adds bash-resource-checker pytest marker to pyproject.toml.

https://claude.ai/code/session_01H2cRJG4a2oz7qB2gKPhrQj
Signed-off-by: Claude <noreply@anthropic.com>
The bash toolset explicitly doesn't support loop constructs (for/while/until).
Removing 'for' from allow lists since it's not a valid command prefix anyway.

https://claude.ai/code/session_01H2cRJG4a2oz7qB2gKPhrQj
Signed-off-by: Claude <noreply@anthropic.com>
211: Use Python processes that actually allocate memory in their address
space so ps aux shows real memory usage. Changed from Alpine to
python:3.11-alpine image.

213: Fix netcat command syntax for openbsd-netcat. Use while loop for
persistent listeners and sleep|nc pattern for established connections.

https://claude.ai/code/session_01H2cRJG4a2oz7qB2gKPhrQj
Signed-off-by: Claude <noreply@anthropic.com>
@linux-foundation-easycla

linux-foundation-easycla Bot commented Jan 30, 2026 •

Copy link
Copy Markdown

CLA Signed

The committers listed above are authorized under a signed CLA.

  • ✅ login: aantn / name: Natan Yellin (ce7cc19)

@netlify

netlify Bot commented Jan 30, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit ce7cc19
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/697f3b5147d83300081546d1
😎 Deploy Preview https://deploy-preview-1456--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Jan 30, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 26a5f9a (#21516345255)

✅ Results of HolmesGPT evals

Automatically triggered by commit 26a5f9a on branch claude/bash-resource-checker-evals-fSHoE

View workflow logs

⚠️ No eval report was generated.

📜 Run @ c73a813 (#21515965999)

✅ Results of HolmesGPT evals

Automatically triggered by commit c73a813 on branch claude/bash-resource-checker-evals-fSHoE

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/14 test cases were successful, 0 regressions, 3 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.9s ±0% 5 11 $0.2279
✅ 101_loki_historical_logs_pod_deleted 53.5s ↑20% 6 14 $0.2922
✅ 111_pod_names_contain_service 38.1s ↑22% 6 11 $0.2309
✅ 12_job_crashing 33.7s ±0% 5 10 $0.2297
✅ 162_get_runbooks 35.6s ↓14% 5 10 $0.2413
🚧 176_network_policy_blocking_traffic_no_runbooks — — — —
✅ 211_bash_process_memory_audit 39.5s 6 10 $0.2339
🚧 212_bash_container_disk_audit — — — —
✅ 213_bash_network_connection_audit 38.2s 6 10 $0.2335
✅ 214_bash_multipod_disk_check 100.2s 8 105 $1.4588
✅ 215_bash_multipod_connectivity 68.8s 9 25 $0.5847
✅ 24_misconfigured_pvc 41.9s ↑24% 7 14 $0.2521
✅ 43_current_datetime_from_prompt 4.9s ±0% 1 — $0.1050
🚧 61_exact_match_counting — — — —
Total 44.3s avg 5.8 avg 22.0 avg $4.0900

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/bash-resource-checker-evals-fSHoE'

Status: Success - 9 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 3c4d80c (#21515467494)

✅ Results of HolmesGPT evals

Automatically triggered by commit 3c4d80c on branch claude/bash-resource-checker-evals-fSHoE

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/14 test cases were successful, 3 regressions, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 33.7s ±0% 5 11 $0.2326
✅ 101_loki_historical_logs_pod_deleted 40.3s ±0% 5 10 $0.2476
✅ 111_pod_names_contain_service 34.3s ↑10% 5 11 $0.2202
✅ 12_job_crashing 30.1s ±0% 4 9 $0.2093
✅ 162_get_runbooks 42.0s ±0% 6 11 $0.2689
✅ 176_network_policy_blocking_traffic_no_runbooks 41.6s ±0% 6 14 $0.2653
✅ 211_bash_process_memory_audit 38.3s 5 9 $0.2198
🚧 212_bash_container_disk_audit — — — —
❌ 213_bash_network_connection_audit 29.3s 5 6 $0.1724
❌ 214_bash_multipod_disk_check 16.3s 2 1 $0.1267
❌ 215_bash_multipod_connectivity 21.8s 3 4 $0.1500
✅ 24_misconfigured_pvc 35.3s ±0% 5 14 $0.2350
✅ 43_current_datetime_from_prompt 5.0s ±0% 1 — $0.1052
✅ 61_exact_match_counting 14.6s ±0% 3 2 $0.1453
Total 29.4s avg 4.2 avg 8.5 avg $2.5982

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/bash-resource-checker-evals-fSHoE'

Status: Success - 9 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 3 Failures Detected

📜 Run @ b9c2240 (#21515345224)

✅ Results of HolmesGPT evals

Automatically triggered by commit b9c2240 on branch claude/bash-resource-checker-evals-fSHoE

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 13/14 test cases were successful, 0 regressions, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 41.6s ↑34% 7 12 $0.2589
✅ 101_loki_historical_logs_pod_deleted 58.1s ↑31% 7 13 $0.3092
✅ 111_pod_names_contain_service 33.6s ±0% 5 11 $0.2210
✅ 12_job_crashing 36.3s ↑11% 5 12 $0.2485
✅ 162_get_runbooks 42.0s ±0% 6 11 $0.2684
✅ 176_network_policy_blocking_traffic_no_runbooks 45.7s ±0% 5 16 $0.2689
✅ 211_bash_process_memory_audit 35.9s 5 8 $0.2272
🚧 212_bash_container_disk_audit — — — —
✅ 213_bash_network_connection_audit 44.6s 5 6 $0.2029
✅ 214_bash_multipod_disk_check 35.7s 5 24 $0.3350
✅ 215_bash_multipod_connectivity 43.3s 6 18 $0.3518
✅ 24_misconfigured_pvc 32.9s ±0% 5 13 $0.2223
✅ 43_current_datetime_from_prompt 5.4s ±0% 1 — $0.1056
✅ 61_exact_match_counting 17.4s ↑21% 4 4 $0.1590
Total 36.3s avg 5.1 avg 12.3 avg $3.1785

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/bash-resource-checker-evals-fSHoE'

Status: Success - 9 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 23a455f (#21514820310)

✅ Results of HolmesGPT evals

Automatically triggered by commit 23a455f on branch claude/bash-resource-checker-evals-fSHoE

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 13/14 test cases were successful, 0 regressions, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 35.5s ↑21% 6 12 $0.2514
✅ 101_loki_historical_logs_pod_deleted 46.3s ↑18% 6 11 $0.2806
✅ 111_pod_names_contain_service 33.0s ↑15% 5 11 $0.2237
✅ 12_job_crashing 33.8s ±0% 5 12 $0.2380
✅ 162_get_runbooks 46.2s ↑12% 6 12 $0.3013
✅ 176_network_policy_blocking_traffic_no_runbooks 48.6s ↑18% 6 17 $0.2881
✅ 211_bash_process_memory_audit 36.9s 6 9 $0.2184
🚧 212_bash_container_disk_audit — — — —
✅ 213_bash_network_connection_audit 30.6s 5 6 $0.2045
✅ 214_bash_multipod_disk_check 35.6s 5 24 $0.3522
✅ 215_bash_multipod_connectivity 63.4s 8 29 $0.3408
✅ 24_misconfigured_pvc 32.4s ±0% 5 13 $0.2232
✅ 43_current_datetime_from_prompt 5.1s ±0% 1 — $0.1053
✅ 61_exact_match_counting 17.8s ↑41% 4 4 $0.1594
Total 35.8s avg 5.2 avg 13.3 avg $3.1870

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/bash-resource-checker-evals-fSHoE'

Status: Success - 11 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit ce7cc19 on branch claude/bash-resource-checker-evals-fSHoE

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/14 test cases were successful, 0 regressions, 3 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 26.9s ↓15% 4 9 $0.2019
✅ 101_loki_historical_logs_pod_deleted 33.3s ↓23% 4 8 $0.2149
✅ 111_pod_names_contain_service 32.8s ±0% 5 11 $0.2197
✅ 12_job_crashing 45.2s ↑37% 7 16 $0.2945
✅ 162_get_runbooks 40.3s ±0% 6 11 $0.2820
✅ 176_network_policy_blocking_traffic_no_runbooks 51.4s ↑20% 7 20 $0.3316
✅ 211_bash_process_memory_audit 37.2s 6 11 $0.2397
🚧 212_bash_container_disk_audit — — — —
✅ 213_bash_network_connection_audit 36.6s 7 8 $0.2331
🚧 214_bash_multipod_disk_check — — — —
✅ 215_bash_multipod_connectivity 82.4s 12 36 $0.5380
✅ 24_misconfigured_pvc 34.3s ±0% 5 14 $0.2363
✅ 43_current_datetime_from_prompt 5.0s ±0% 1 — $0.1062
🚧 61_exact_match_counting — — — —
Total 38.7s avg 5.8 avg 14.4 avg $2.8978

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'claude/bash-resource-checker-evals-fSHoE'

Status: Success - 60 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/bash-resource-checker-evals-fSHoE -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

bash-resource-checker, benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/bash-resource-checker-evals-fSHoE -f markers=regression -f filter=

@github-actions

github-actions Bot commented Jan 30, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for 91f7fc7 (built in 4m 23s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:91f7fc7
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:91f7fc7 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:91f7fc7
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:91f7fc7

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:91f7fc7

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:91f7fc7

@coderabbitai

coderabbitai Bot commented Jan 30, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Rate limit exceeded

@aantn has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 3 minutes and 22 seconds before requesting another review.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

Walkthrough

Adds a pytest marker bash-resource-checker and introduces five new Holmes LLM test fixtures (211–215) under tests/llm/fixtures/test_ask_holmes/, each including Kubernetes manifests, test_case.yaml, and toolsets.yaml to validate bash-based container inspections via kubectl exec/kubectl get.

Changes

Cohort / File(s) Summary
Pytest Configuration
pyproject.toml
Added pytest marker bash-resource-checker to MARKERS.
Fixture 211 — Process Memory Audit
tests/llm/fixtures/test_ask_holmes/211_bash_process_memory_audit/manifest.yaml, .../test_case.yaml, .../toolsets.yaml
New Pod manifest spawning many sleepers and Python processes to simulate memory usage; test enforces kubectl exec bash inspection to identify top memory consumers; toolset allows kubectl exec/get.
Fixture 212 — Container Disk Audit
tests/llm/fixtures/test_ask_holmes/212_bash_container_disk_audit/manifest.yaml, .../test_case.yaml, .../toolsets.yaml
Pod generates 90 small and 10 large files in /var/data; test requires identifying files >10MB via kubectl exec; toolset permits kubectl exec/get.
Fixture 213 — Network Connection Audit
tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit/manifest.yaml, .../test_case.yaml, .../toolsets.yaml
Pod installs netcat and creates 50 listeners and 50 clients (ports 8001–8050); test expects kubectl exec evidence of LISTEN/ESTABLISHED states; toolset configured accordingly.
Fixture 214 — Multi-Pod Disk Check
tests/llm/fixtures/test_ask_holmes/214_bash_multipod_disk_check/manifest.yaml, .../test_case.yaml, .../toolsets.yaml
StatefulSet with 100 replicas where pods 90–99 write large files (>50MB); test asks to identify high-disk pods across all replicas; toolset enables kubectl exec/get.
Fixture 215 — Multi-Pod Connectivity
tests/llm/fixtures/test_ask_holmes/215_bash_multipod_connectivity/manifest.yaml, .../test_case.yaml, .../toolsets.yaml
Adds DB pod, Service, NetworkPolicy and many worker pods (some labeled to allow DB access, others unlabeled and blocked); test validates in-cluster connectivity and policy effects via kubectl exec; toolset allows kubectl exec/get.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested labels

enhancement

Suggested reviewers

  • Sheeproid
  • moshemorad
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: adding 5 new bash resource checker evaluation tests for kubectl exec scenarios.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🤖 Fix all issues with AI agents
In
`@tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit/manifest.yaml`:
- Around line 3-7: The manifest uses a technical, non-unique resource name/label
"connection-pool" which can collide across tests; rename the metadata.name and
labels.app to a neutral, unique identifier that includes the test id (e.g.,
"app-213-<neutral-name>") and update any references to it in the associated
test_case.yaml so they match the new metadata.name and labels.app; ensure the
new name follows resource naming rules and is used consistently across the
manifest and test_case.yaml.
- Around line 12-44: The inline multi-line shell in the pod/container "command"
should be moved into a Kubernetes Secret and mounted as a file; create a Secret
(e.g., name orders-worker-213-script) with stringData.entrypoint.sh containing
the current script body, update the manifest to remove the inline "command"
block and instead mount that Secret as a volume (readOnly) and set the
container's command to run the mounted /path/to/entrypoint.sh (e.g., /bin/sh -c
/mnt/entrypoint.sh), ensure the Secret volume is referenced in the pod spec and
the script file has executable/newline-safe content and appropriate permissions
at runtime.

In
`@tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit/test_case.yaml`:
- Line 1: Update the user_prompt and any kubectl references that mention the old
pod name "connection-pool" to use the neutral, unique pod name used in the
manifest (the renamed pod referenced elsewhere in this test); locate the
user_prompt entry in
test_ask_holmes/213_bash_network_connection_audit/test_case.yaml and replace
occurrences of "connection-pool" (and any other hard-coded references in the
same file) so they exactly match the manifest's pod name used by functions that
query pods (e.g., any kubectl commands or variables that reference the pod), and
apply the same rename to the other occurrences noted in this test (the other
similar prompt/command entries).
- Around line 3-10: Update the test fixture by replacing the generic
expected_output items with concrete, discoverable assertions (e.g., exact TCP
state counts and explicit port listings such as "LISTEN on ports 8001-8005: 3
entries" or "ESTABLISHED connections: 7" and specific kubectl exec invocation
text), and then remove or set include_tool_calls to false; reference the keys
expected_output and include_tool_calls in
test_ask_holmes/213_bash_network_connection_audit/test_case.yaml to locate and
change the entries.
- Around line 45-48: The LISTEN_COUNT assignment in the loop uses `grep -c
LISTEN || echo "0"` which can produce duplicate zeros when grep exits non-zero;
update the command substitution used to set LISTEN_COUNT (the line assigning
LISTEN_COUNT in the for loop that runs `kubectl exec -n app-213 connection-pool
-- ss -tln`) to suppress the non-zero exit without emitting an extra "0" by
replacing `grep -c LISTEN || echo "0"` with `grep -c LISTEN || true`, so
LISTEN_COUNT contains only the grep output and the subsequent numeric comparison
(`[ "$LISTEN_COUNT" -ge 40 ]`) won't fail.
🧹 Nitpick comments (7)
tests/llm/fixtures/test_ask_holmes/212_bash_container_disk_audit/manifest.yaml (2)

12-32: Consider using a Kubernetes Secret for the initialization script.

As per coding guidelines, scripts in test infrastructure should be stored in Kubernetes Secrets rather than inline in manifests. While the current inline approach works, extracting the script to a Secret would improve consistency with the project's established patterns.


8-39: Optional: Add security context to harden the test pod.

The static analysis flags allowPrivilegeEscalation and running as root. For a test fixture in an isolated namespace, this is low risk, but adding a security context would align with security best practices:

🛡️ Optional security hardening
 spec:
+  securityContext:
+    runAsNonRoot: true
+    runAsUser: 1000
   containers:
     - name: worker
       image: alpine:3.19
+      securityContext:
+        allowPrivilegeEscalation: false
+        readOnlyRootFilesystem: false
       command:
tests/llm/fixtures/test_ask_holmes/211_bash_process_memory_audit/manifest.yaml (2)

12-35: Consider using a Kubernetes Secret for the initialization script.

Same as the 212 fixture - the guideline recommends using Secrets for scripts in test infrastructure rather than inline manifests.


8-42: Optional: Add security context to harden the test pod.

Same security hardening opportunity as the 212 fixture. For test pods, this is optional but aligns with best practices flagged by static analysis (CKV_K8S_20, CKV_K8S_23).

🛡️ Optional security hardening
 spec:
+  securityContext:
+    runAsNonRoot: true
+    runAsUser: 1000
   containers:
     - name: worker
       image: python:3.11-alpine
+      securityContext:
+        allowPrivilegeEscalation: false
       command:
tests/llm/fixtures/test_ask_holmes/214_bash_multipod_disk_check/manifest.yaml (1)

10-18: Optional: Add security context to harden test pods.

Static analysis flags that pods can run with privilege escalation and as root. While acceptable for ephemeral test fixtures, adding a security context would align with best practices.

🛡️ Example security context for one pod
 spec:
   containers:
     - name: worker
       image: alpine:3.19
       command: ["/bin/sh", "-c", "dd if=/dev/urandom of=/tmp/cache.dat bs=1K count=100 2>/dev/null; sleep 3600"]
       resources:
         limits:
           memory: "64Mi"
           cpu: "100m"
+      securityContext:
+        allowPrivilegeEscalation: false
+        runAsNonRoot: true
+        runAsUser: 1000
tests/llm/fixtures/test_ask_holmes/215_bash_multipod_connectivity/manifest.yaml (1)

11-26: Optional: Add security context to harden test pods.

Same as other manifests - static analysis flags privilege escalation and root containers. Consider adding securityContext for best practices, though acceptable for ephemeral test fixtures.

tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit/manifest.yaml (1)

1-8: Consider adding small shared infra components for a more realistic topology.
This fixture currently provisions a single pod; if the suite expects realistic architecture, add a small log aggregation component (e.g., Loki) with separation of concerns.

As per coding guidelines: Implement realistic full architecture in test infrastructure (e.g., use Loki for log aggregation) with proper separation of concerns.

Comment on lines +3 to +7
metadata:
name: connection-pool
namespace: app-213
labels:
app: connection-pool

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Rename the pod/label to a neutral, unique identifier.
connection-pool is a hint‑giving technical name and may collide across tests; use a neutral business‑context name with the test id and update references in test_case.yaml accordingly.

🔁 Suggested rename
 metadata:
-  name: connection-pool
+  name: orders-worker-213
   namespace: app-213
   labels:
-    app: connection-pool
+    app: orders-worker-213

As per coding guidelines: All resource names must be unique across tests to prevent conflicts when tests run simultaneously; Never use obvious or hint-giving resource names in tests - use neutral, business-context names instead of technical indicators.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
metadata:
name: connection-pool
namespace: app-213
labels:
app: connection-pool
metadata:
name: orders-worker-213
namespace: app-213
labels:
app: orders-worker-213
🤖 Prompt for AI Agents
In
`@tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit/manifest.yaml`
around lines 3 - 7, The manifest uses a technical, non-unique resource
name/label "connection-pool" which can collide across tests; rename the
metadata.name and labels.app to a neutral, unique identifier that includes the
test id (e.g., "app-213-<neutral-name>") and update any references to it in the
associated test_case.yaml so they match the new metadata.name and labels.app;
ensure the new name follows resource naming rules and is used consistently
across the manifest and test_case.yaml.

Comment on lines +12 to +44
command:
- /bin/sh
- -c
- |
# Install netcat and iproute2 for ss
apk add --no-cache netcat-openbsd iproute2 >/dev/null 2>&1

# Create 50 TCP listeners on ports 8001-8050
for port in $(seq 8001 8050); do
# Use while loop to keep listener alive after connection
while true; do nc -l -p $port; done &
done

echo "Created 50 TCP listeners"

# Wait a moment for listeners to start
sleep 2

# Create connections to each listener (these will be ESTABLISHED)
for port in $(seq 8001 8050); do
# Keep connection open by sleeping after connect
(sleep 3600 | nc 127.0.0.1 $port) &
done

echo "Created 50 client connections"

# Show connection summary
sleep 2
echo "Connection states:"
ss -tan | head -20

# Keep container running
sleep 3600

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

cd tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit && \
cat -n manifest.yaml

Repository: HolmesGPT/holmesgpt

Length of output: 1731


Move the startup script to a Kubernetes Secret.
The inline script in the container command violates test infrastructure policy. Create a Secret and mount it as a volume instead.

Suggested structure
 spec:
+  volumes:
+    - name: entrypoint
+      secret:
+        secretName: orders-worker-213-script
   containers:
     - name: worker
       image: alpine:3.19
-      command:
-        - /bin/sh
-        - -c
-        - |
-          # Install netcat and iproute2 for ss
-          apk add --no-cache netcat-openbsd iproute2 >/dev/null 2>&1
-          ...
+      command: ["/bin/sh", "/scripts/entrypoint.sh"]
+      volumeMounts:
+        - name: entrypoint
+          mountPath: /scripts
+          readOnly: true
apiVersion: v1
kind: Secret
metadata:
  name: orders-worker-213-script
  namespace: app-213
stringData:
  entrypoint.sh: |
    # (move the current script body here)
🤖 Prompt for AI Agents
In
`@tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit/manifest.yaml`
around lines 12 - 44, The inline multi-line shell in the pod/container "command"
should be moved into a Kubernetes Secret and mounted as a file; create a Secret
(e.g., name orders-worker-213-script) with stringData.entrypoint.sh containing
the current script body, update the manifest to remove the inline "command"
block and instead mount that Secret as a volume (readOnly) and set the
container's command to run the mounted /path/to/entrypoint.sh (e.g., /bin/sh -c
/mnt/entrypoint.sh), ensure the Secret volume is referenced in the pod spec and
the script file has executable/newline-safe content and appropriate permissions
at runtime.

@@ -0,0 +1,64 @@
user_prompt: "The connection-pool pod in namespace app-213 may have connection leaks. Check how many TCP connections it has and group them by state (ESTABLISHED, TIME_WAIT, LISTEN, etc.)."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Align pod references with the neutral, unique name.
Update the prompt and kubectl references to match the renamed pod from the manifest.

🔁 Suggested rename alignment
-user_prompt: "The connection-pool pod in namespace app-213 may have connection leaks. Check how many TCP connections it has and group them by state (ESTABLISHED, TIME_WAIT, LISTEN, etc.)."
+user_prompt: "The orders-worker-213 pod in namespace app-213 may have connection leaks. Check how many TCP connections it has and group them by state (ESTABLISHED, TIME_WAIT, LISTEN, etc.)."

-    if kubectl wait --for=condition=ready pod/connection-pool -n app-213 --timeout=2s 2>/dev/null; then
+    if kubectl wait --for=condition=ready pod/orders-worker-213 -n app-213 --timeout=2s 2>/dev/null; then

-    kubectl describe pod/connection-pool -n app-213
+    kubectl describe pod/orders-worker-213 -n app-213

-    kubectl exec -n app-213 connection-pool -- ss -tan
+    kubectl exec -n app-213 orders-worker-213 -- ss -tan

As per coding guidelines: All resource names must be unique across tests to prevent conflicts when tests run simultaneously; Never use obvious or hint-giving resource names in tests - use neutral, business-context names instead of technical indicators.

Also applies to: 28-28, 39-39, 58-58

🤖 Prompt for AI Agents
In
`@tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit/test_case.yaml`
at line 1, Update the user_prompt and any kubectl references that mention the
old pod name "connection-pool" to use the neutral, unique pod name used in the
manifest (the renamed pod referenced elsewhere in this test); locate the
user_prompt entry in
test_ask_holmes/213_bash_network_connection_audit/test_case.yaml and replace
occurrences of "connection-pool" (and any other hard-coded references in the
same file) so they exactly match the manifest's pod name used by functions that
query pods (e.g., any kubectl commands or variables that reference the pod), and
apply the same rename to the other occurrences noted in this test (the other
similar prompt/command entries).

Comment on lines +3 to +10
include_tool_calls: true

expected_output:
- "Must call bash tool with kubectl exec"
- "Must report TCP connection states"
- "Must identify LISTEN connections on ports 8001-8050"
- "Must identify ESTABLISHED connections"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Make expected_output specific; then drop include_tool_calls.
The current expectations are generic. Encode concrete counts/port ranges to be hallucination‑proof; once specific, include_tool_calls can be false.

✅ Example expectations
-include_tool_calls: true
+include_tool_calls: false

 expected_output:
-  - "Must call bash tool with kubectl exec"
-  - "Must report TCP connection states"
-  - "Must identify LISTEN connections on ports 8001-8050"
-  - "Must identify ESTABLISHED connections"
+  - "Uses kubectl exec with bash against pod orders-worker-213"
+  - "Reports LISTEN sockets on ports 8001-8050 (≈50)"
+  - "Reports ESTABLISHED connections to 127.0.0.1:8001-8050 (≈50)"
+  - "Groups results by state (LISTEN/ESTABLISHED/TIME_WAIT)"

As per coding guidelines: Use specific, discoverable values in expected_output instead of generic patterns to rule out hallucinations; Use 'include_tool_calls: true' in test configuration only when expected output values are too generic to be hallucination-proof.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
include_tool_calls: true
expected_output:
- "Must call bash tool with kubectl exec"
- "Must report TCP connection states"
- "Must identify LISTEN connections on ports 8001-8050"
- "Must identify ESTABLISHED connections"
include_tool_calls: false
expected_output:
- "Uses kubectl exec with bash against pod orders-worker-213"
- "Reports LISTEN sockets on ports 8001-8050 (≈50)"
- "Reports ESTABLISHED connections to 127.0.0.1:8001-8050 (≈50)"
- "Groups results by state (LISTEN/ESTABLISHED/TIME_WAIT)"
🤖 Prompt for AI Agents
In
`@tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit/test_case.yaml`
around lines 3 - 10, Update the test fixture by replacing the generic
expected_output items with concrete, discoverable assertions (e.g., exact TCP
state counts and explicit port listings such as "LISTEN on ports 8001-8005: 3
entries" or "ESTABLISHED connections: 7" and specific kubectl exec invocation
text), and then remove or set include_tool_calls to false; reference the keys
expected_output and include_tool_calls in
test_ask_holmes/213_bash_network_connection_audit/test_case.yaml to locate and
change the entries.

Comment on lines +45 to +48
for i in {1..30}; do
LISTEN_COUNT=$(kubectl exec -n app-213 connection-pool -- ss -tln 2>/dev/null | grep -c LISTEN || echo "0")
echo "Attempt $i: $LISTEN_COUNT LISTEN connections"
if [ "$LISTEN_COUNT" -ge 40 ]; then

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

🧩 Analysis chain

🏁 Script executed:

cat -n tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit/test_case.yaml | head -60

Repository: HolmesGPT/holmesgpt

Length of output: 2380


🏁 Script executed:

find tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit -type f -name "*.yaml" -o -name "*.yml"

Repository: HolmesGPT/holmesgpt

Length of output: 312


🏁 Script executed:

# Check if there's a fixture definition or infrastructure file for this test
find tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit -type f

Repository: HolmesGPT/holmesgpt

Length of output: 312


🏁 Script executed:

cat tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit/manifest.yaml

Repository: HolmesGPT/holmesgpt

Length of output: 1374


🏁 Script executed:

# Let me verify the grep -c behavior to understand if the claim about double-zero output is accurate
cat > /tmp/test_grep.sh << 'EOF'
#!/bin/bash

# Test 1: grep -c on non-empty output
echo "Test 1: grep -c with matches"
echo -e "LISTEN\nLISTEN\nESTABLISHED" | grep -c "LISTEN"
echo "---"

# Test 2: grep -c on empty output
echo "Test 2: grep -c with no matches (should print 0)"
echo -e "ESTABLISHED\nTIME_WAIT" | grep -c "LISTEN"
echo "---"

# Test 3: Original command pattern with || echo "0"
echo "Test 3: With || echo '0' when grep returns 0"
LISTEN_COUNT=$(echo -e "ESTABLISHED\nTIME_WAIT" | grep -c "LISTEN" || echo "0")
echo "LISTEN_COUNT='$LISTEN_COUNT'"
echo "---"

# Test 4: Check if there are multiple lines
echo "Test 4: Checking raw output with set -x"
set -x
LISTEN_COUNT=$(echo -e "ESTABLISHED\nTIME_WAIT" | grep -c "LISTEN" || echo "0")
set +x
echo "Final value: $LISTEN_COUNT"
echo "---"

# Test 5: Check with failed command
echo "Test 5: With command that fails"
LISTEN_COUNT=$(/bin/false | grep -c "LISTEN" || echo "0")
echo "LISTEN_COUNT='$LISTEN_COUNT'"
EOF

bash /tmp/test_grep.sh

Repository: HolmesGPT/holmesgpt

Length of output: 443


🏁 Script executed:

cat tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit/manifest.yaml | grep -A 5 "metadata:" | grep "name:"

Repository: HolmesGPT/holmesgpt

Length of output: 86


Remove redundant echo fallback in LISTEN_COUNT assignment.

The pipe grep -c LISTEN || echo "0" outputs double zero (0\n0) because grep -c exits with code 1 when no matches are found, triggering the fallback echo. This causes [ "$LISTEN_COUNT" -ge 40 ] to fail with "integer expression expected". Use || true instead to suppress the non-zero exit without duplicating output:

🔧 Fixed assignment
-    LISTEN_COUNT=$(kubectl exec -n app-213 connection-pool -- ss -tln 2>/dev/null | grep -c LISTEN || echo "0")
+    LISTEN_COUNT=$(kubectl exec -n app-213 connection-pool -- ss -tln 2>/dev/null | grep -c LISTEN || true)
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
for i in {1..30}; do
LISTEN_COUNT=$(kubectl exec -n app-213 connection-pool -- ss -tln 2>/dev/null | grep -c LISTEN || echo "0")
echo "Attempt $i: $LISTEN_COUNT LISTEN connections"
if [ "$LISTEN_COUNT" -ge 40 ]; then
for i in {1..30}; do
LISTEN_COUNT=$(kubectl exec -n app-213 connection-pool -- ss -tln 2>/dev/null | grep -c LISTEN || true)
echo "Attempt $i: $LISTEN_COUNT LISTEN connections"
if [ "$LISTEN_COUNT" -ge 40 ]; then
🤖 Prompt for AI Agents
In
`@tests/llm/fixtures/test_ask_holmes/213_bash_network_connection_audit/test_case.yaml`
around lines 45 - 48, The LISTEN_COUNT assignment in the loop uses `grep -c
LISTEN || echo "0"` which can produce duplicate zeros when grep exits non-zero;
update the command substitution used to set LISTEN_COUNT (the line assigning
LISTEN_COUNT in the for loop that runs `kubectl exec -n app-213 connection-pool
-- ss -tln`) to suppress the non-zero exit without emitting an extra "0" by
replacing `grep -c LISTEN || echo "0"` with `grep -c LISTEN || true`, so
LISTEN_COUNT contains only the grep output and the subsequent numeric comparison
(`[ "$LISTEN_COUNT" -ge 40 ]`) won't fail.

Using /dev/urandom to create 150MB of files is slow and likely caused
setup timeout. /dev/zero is much faster and still works for testing
file size detection.

https://claude.ai/code/session_01H2cRJG4a2oz7qB2gKPhrQj
Signed-off-by: Claude <noreply@anthropic.com>
Update tests 214 and 215 to explicitly require bash for-loops in the command.
These tests should FAIL until the bash toolset supports loop constructs.

- User prompt now explicitly asks for a bash script with for-loop
- Expected output checks that command contains loop syntax
- Expected output fails if Holmes makes separate tool calls per pod

This is TDD: tests define the desired behavior and fail until implemented.

https://claude.ai/code/session_01H2cRJG4a2oz7qB2gKPhrQj
Signed-off-by: Claude <noreply@anthropic.com>
Tests 214 & 215 now deploy 100 pods instead of 10. Making 100 individual
kubectl exec calls is impractical, so Holmes must write a bash loop.
Since bash toolset blocks loops, tests will fail naturally.

- Natural prompts (no artificial "write a for-loop" requirement)
- 214: 100 cache-worker pods (90 small, 10 large /tmp)
- 215: 100 api-worker pods (90 allowed, 10 blocked by NetworkPolicy)
- Uses StatefulSet for efficient deployment

https://claude.ai/code/session_01H2cRJG4a2oz7qB2gKPhrQj
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Fix all issues with AI agents
In
`@tests/llm/fixtures/test_ask_holmes/214_bash_multipod_disk_check/manifest.yaml`:
- Around line 6-29: The manifest uses the technical, non-unique resource name
"cache-worker" in multiple places (metadata.name, StatefulSet name, serviceName,
selector app label and template labels); rename all occurrences to a neutral,
test-unique name that includes the test id (for example a business-context token
plus "214") so names are globally unique, and update the selector keys
(matchLabels/app), pod template labels, and any hostname parsing or references
in the test_case fixture to use the new name consistently (ensure metadata.name,
spec.serviceName, spec.selector.matchLabels, and template.metadata.labels all
match the new identifier).
- Around line 34-44: The startup shell script is inlined in the container args
(the block that sets POD_ID, writes /tmp/cache.dat, and sleeps) which violates
the guideline; extract that script into a Kubernetes Secret (e.g., key
startup.sh containing the POD_ID logic and dd commands writing /tmp/cache.dat),
update the Pod spec to mount the Secret as a volume and run the mounted script
from the container (replace the inline args with a command that executes the
mounted script), and ensure the test runner has RBAC to create and mount Secrets
in the app-214 namespace so the Secret can be created and the volume mounted at
runtime.

In
`@tests/llm/fixtures/test_ask_holmes/215_bash_multipod_connectivity/test_case.yaml`:
- Around line 9-12: Update the test description string that contains the
contradictory sentence "Holmes needs to write a bash script with a loop, but
bash toolset blocks loops." to clearly state the intended scenario: indicate
whether this is an intentionally failing TDD test, whether the bash toolset
actually allows loops (remove “blocks loops” if outdated), or describe the
alternative approach Holmes should use (e.g., parallel kubectl exec via xargs or
a provided multipod helper). Edit the YAML description field so it unambiguously
reflects one of the three options and mention the expected outcome for pods 0-89
vs 90-99.
🧹 Nitpick comments (4)
tests/llm/fixtures/test_ask_holmes/215_bash_multipod_connectivity/manifest.yaml (1)

103-273: Consider using a second StatefulSet for blocked pods.

These 10 identical pod definitions could be consolidated into a StatefulSet without the db-access label, reducing ~170 lines to ~30 lines while maintaining the same test behavior.

♻️ Suggested refactor using StatefulSet
-# 10 individual pods WITHOUT db-access label (blocked by NetworkPolicy)
-# Named api-worker-90 through api-worker-99 to continue the sequence
-apiVersion: v1
-kind: Pod
-metadata:
-  name: api-worker-90
-  namespace: app-215
-  labels:
-    app: api-worker
-spec:
-  containers:
-    - name: worker
-      image: alpine:3.19
-      command: ["/bin/sh", "-c", "apk add --no-cache netcat-openbsd >/dev/null 2>&1; sleep 3600"]
-      resources:
-        limits:
-          memory: "64Mi"
-          cpu: "50m"
-# ... (remaining 9 pods removed)
+# StatefulSet: 10 worker pods WITHOUT db-access label (blocked by NetworkPolicy)
+apiVersion: v1
+kind: Service
+metadata:
+  name: blocked-worker
+  namespace: app-215
+spec:
+  clusterIP: None
+  selector:
+    app: blocked-worker
+---
+apiVersion: apps/v1
+kind: StatefulSet
+metadata:
+  name: blocked-worker
+  namespace: app-215
+spec:
+  serviceName: blocked-worker
+  replicas: 10
+  podManagementPolicy: Parallel
+  selector:
+    matchLabels:
+      app: blocked-worker
+  template:
+    metadata:
+      labels:
+        app: blocked-worker
+    spec:
+      containers:
+        - name: worker
+          image: alpine:3.19
+          command: ["/bin/sh", "-c", "apk add --no-cache netcat-openbsd >/dev/null 2>&1; sleep 3600"]
+          resources:
+            limits:
+              memory: "64Mi"
+              cpu: "50m"
+            requests:
+              memory: "32Mi"
+              cpu: "10m"

Note: This would require updating test_case.yaml to reference blocked-worker-0 through blocked-worker-9 instead of api-worker-90 through api-worker-99.

tests/llm/fixtures/test_ask_holmes/214_bash_multipod_disk_check/manifest.yaml (2)

31-33: Add a restrictive securityContext to avoid PSA/OPA rejections.

If your test clusters enforce restricted policies, lack of runAsNonRoot / allowPrivilegeEscalation: false may block pod admission.

🔒 Suggested hardening
       containers:
         - name: worker
           image: alpine:3.19
+          securityContext:
+            allowPrivilegeEscalation: false
+            runAsNonRoot: true
+            runAsUser: 1000
           command: ["/bin/sh", "-c"]

45-51: Consider lowering CPU/memory limits for 100 pods.

This workload writes files then sleeps; with 100 replicas, trimming limits/requests reduces test footprint.

♻️ Example reduction
           resources:
             limits:
-              memory: "128Mi"
-              cpu: "50m"
+              memory: "64Mi"
+              cpu: "25m"
             requests:
-              memory: "32Mi"
-              cpu: "10m"
+              memory: "16Mi"
+              cpu: "5m"

As per coding guidelines, Use minimal resource footprints in test services by reducing memory and CPU allocations.

tests/llm/fixtures/test_ask_holmes/214_bash_multipod_disk_check/test_case.yaml (1)

1-8: Make expected_output more concrete so include_tool_calls can be dropped.

Right now the assertions are high-level; consider listing specific pod/size outputs (or a representative subset) so the check is hallucination-resistant and you can set include_tool_calls: false unless you truly need tool-call inspection.

✍️ Example tightening
-include_tool_calls: true
+include_tool_calls: false

 expected_output:
-  - "Must identify cache-worker-90 through cache-worker-99 as having high disk usage"
-  - "Must check all 100 pods"
+  - "cache-worker-90 through cache-worker-99 should each report ~60MB in /tmp (e.g., cache-worker-90: 60M)"
+  - "cache-worker-0 through cache-worker-89 should each report ~100KB in /tmp"
+  - "Output should confirm total pods checked: 100"

Based on learnings, Use specific, discoverable values in expected_output instead of generic patterns to rule out hallucinations; Use 'include_tool_calls: true' in test configuration only when expected output values are too generic to be hallucination-proof.

Comment on lines +6 to +29
metadata:
name: cache-worker
namespace: app-214
spec:
clusterIP: None
selector:
app: cache-worker
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: cache-worker
namespace: app-214
spec:
serviceName: cache-worker
replicas: 100
podManagementPolicy: Parallel
selector:
matchLabels:
app: cache-worker
template:
metadata:
labels:
app: cache-worker

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Rename resources to neutral, test-unique names (include the test id).

cache-worker is both technical/hinting and not unique across tests. Please switch to a neutral, business-context name with the test id and update selectors plus hostname parsing accordingly (and then update references in the test_case fixture).

🛠️ Example rename (adjust consistently across this fixture)
 metadata:
-  name: cache-worker
+  name: ledger-worker-214
   namespace: app-214
 spec:
   clusterIP: None
   selector:
-    app: cache-worker
+    app: ledger-worker-214
 ---
 apiVersion: apps/v1
 kind: StatefulSet
 metadata:
-  name: cache-worker
+  name: ledger-worker-214
   namespace: app-214
 spec:
-  serviceName: cache-worker
+  serviceName: ledger-worker-214
   replicas: 100
   podManagementPolicy: Parallel
   selector:
     matchLabels:
-      app: cache-worker
+      app: ledger-worker-214
   template:
     metadata:
       labels:
-        app: cache-worker
+        app: ledger-worker-214
     spec:
       containers:
         - name: worker
           image: alpine:3.19
           command: ["/bin/sh", "-c"]
           args:
             - |
-              POD_ID=$(hostname | sed 's/cache-worker-//')
+              POD_ID=$(hostname | sed 's/ledger-worker-214-//')

As per coding guidelines, All resource names must be unique across tests to prevent conflicts when tests run simultaneously; Never use obvious or hint-giving resource names in tests - use neutral, business-context names instead of technical indicators.

Also applies to: 37-43

🤖 Prompt for AI Agents
In
`@tests/llm/fixtures/test_ask_holmes/214_bash_multipod_disk_check/manifest.yaml`
around lines 6 - 29, The manifest uses the technical, non-unique resource name
"cache-worker" in multiple places (metadata.name, StatefulSet name, serviceName,
selector app label and template labels); rename all occurrences to a neutral,
test-unique name that includes the test id (for example a business-context token
plus "214") so names are globally unique, and update the selector keys
(matchLabels/app), pod template labels, and any hostname parsing or references
in the test_case fixture to use the new name consistently (ensure metadata.name,
spec.serviceName, spec.selector.matchLabels, and template.metadata.labels all
match the new identifier).

Comment on lines +34 to +44
command: ["/bin/sh", "-c"]
args:
- |
POD_ID=$(hostname | sed 's/cache-worker-//')
# Pods 90-99 get large files (60MB), others get small (100KB)
if [ "$POD_ID" -ge 90 ]; then
dd if=/dev/zero of=/tmp/cache.dat bs=1M count=60 2>/dev/null
else
dd if=/dev/zero of=/tmp/cache.dat bs=1K count=100 2>/dev/null
fi
sleep 3600

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

# First, let's check if the file exists and read its content
cat tests/llm/fixtures/test_ask_holmes/214_bash_multipod_disk_check/manifest.yaml

Repository: HolmesGPT/holmesgpt

Length of output: 1343


Store the startup script in a Kubernetes Secret instead of inline args.

This manifest violates the test-infrastructure guideline: scripts must be stored in Secrets, not inline. The current implementation also makes the script harder to reuse and test independently.

Create a Secret with the script content and mount it as a volume, then verify the test runner's RBAC allows Secret creation/mounting in the app-214 namespace.

🛠️ Proposed fix (Secret-mounted script)
+apiVersion: v1
+kind: Secret
+metadata:
+  name: cache-worker-script
+  namespace: app-214
+type: Opaque
+stringData:
+  start.sh: |
+    POD_ID=$(hostname | sed 's/cache-worker-//')
+    if [ "$POD_ID" -ge 90 ]; then
+      dd if=/dev/zero of=/tmp/cache.dat bs=1M count=60 2>/dev/null
+    else
+      dd if=/dev/zero of=/tmp/cache.dat bs=1K count=100 2>/dev/null
+    fi
+    sleep 3600
+---
 apiVersion: v1
 kind: Service
 metadata:
   name: cache-worker
   namespace: app-214
 spec:
   clusterIP: None
   selector:
     app: cache-worker
 ---
 apiVersion: apps/v1
 kind: StatefulSet
 metadata:
   name: cache-worker
   namespace: app-214
 spec:
   serviceName: cache-worker
   replicas: 100
   podManagementPolicy: Parallel
   selector:
     matchLabels:
       app: cache-worker
   template:
     metadata:
       labels:
         app: cache-worker
     spec:
       containers:
         - name: worker
           image: alpine:3.19
-          command: ["/bin/sh", "-c"]
-          args:
-            - |
-              POD_ID=$(hostname | sed 's/cache-worker-//')
-              # Pods 90-99 get large files (60MB), others get small (100KB)
-              if [ "$POD_ID" -ge 90 ]; then
-                dd if=/dev/zero of=/tmp/cache.dat bs=1M count=60 2>/dev/null
-              else
-                dd if=/dev/zero of=/tmp/cache.dat bs=1K count=100 2>/dev/null
-              fi
-              sleep 3600
+          command: ["/bin/sh", "/scripts/start.sh"]
+          volumeMounts:
+            - name: script
+              mountPath: /scripts
+              readOnly: true
           resources:
             limits:
               memory: "128Mi"
               cpu: "50m"
             requests:
               memory: "32Mi"
               cpu: "10m"
+      volumes:
+        - name: script
+          secret:
+            secretName: cache-worker-script
+            defaultMode: 0555
🤖 Prompt for AI Agents
In
`@tests/llm/fixtures/test_ask_holmes/214_bash_multipod_disk_check/manifest.yaml`
around lines 34 - 44, The startup shell script is inlined in the container args
(the block that sets POD_ID, writes /tmp/cache.dat, and sleeps) which violates
the guideline; extract that script into a Kubernetes Secret (e.g., key
startup.sh containing the POD_ID logic and dd commands writing /tmp/cache.dat),
update the Pod spec to mount the Secret as a volume and run the mounted script
from the container (replace the inline args with a command that executes the
mounted script), and ensure the test runner has RBAC to create and mount Secrets
in the app-214 namespace so the Secret can be created and the volume mounted at
runtime.

Comment on lines +9 to +12
description: |
TDD test: 100 worker pods deployed - too many for individual kubectl exec calls.
Holmes needs to write a bash script with a loop, but bash toolset blocks loops.
Pods 0-89 can reach database, pods 90-99 are blocked by NetworkPolicy.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Clarify the contradictory description.

Line 11 states: "Holmes needs to write a bash script with a loop, but bash toolset blocks loops."

This reads as contradictory—if loops are blocked, how can Holmes accomplish the task? Please clarify whether:

  1. This is intentionally a failing TDD test scenario
  2. The toolset configuration allows loops (and the description is outdated)
  3. There's an alternative approach Holmes should use
🤖 Prompt for AI Agents
In
`@tests/llm/fixtures/test_ask_holmes/215_bash_multipod_connectivity/test_case.yaml`
around lines 9 - 12, Update the test description string that contains the
contradictory sentence "Holmes needs to write a bash script with a loop, but
bash toolset blocks loops." to clearly state the intended scenario: indicate
whether this is an intentionally failing TDD test, whether the bash toolset
actually allows loops (remove “blocks loops” if outdated), or describe the
alternative approach Holmes should use (e.g., parallel kubectl exec via xargs or
a provided multipod helper). Edit the YAML description field so it unambiguously
reflects one of the three options and mention the expected outcome for pods 0-89
vs 90-99.

claude and others added 2 commits January 30, 2026 12:47
Bad pods now use formula (ordinal*7+3)%11==0, creating scattered pattern:
pods 1, 12, 23, 34, 45, 56, 67, 78, 89 (9 pods spread across 100)

Holmes CANNOT:
- Guess which pods are bad
- Sample strategically
- Check just high-numbered pods

Holmes MUST check every pod to find all 9 scattered bad pods.
Without bash loops, this will fail/timeout.

https://claude.ai/code/session_01H2cRJG4a2oz7qB2gKPhrQj
Signed-off-by: Claude <noreply@anthropic.com>
@aantn

aantn commented Feb 1, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
marker: bash-resource-checker

@github-actions

github-actions Bot commented Feb 1, 2026

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ✅ Completed successfully


🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Branch claude/bash-resource-checker-evals-fSHoE
Model opus-4.5
Markers bash-resource-checker
Iterations 1
Duration 6m 28s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 3/5 test cases were successful, 0 regressions, 2 setup failures
Status Test case Time Turns Tools Cost
✅ 211_bash_process_memory_audit 37.6s 5 9 $0.2286
🚧 212_bash_container_disk_audit — — — —
✅ 213_bash_network_connection_audit 28.8s 5 6 $0.2022
🚧 214_bash_multipod_disk_check — — — —
✅ 215_bash_multipod_connectivity 66.3s 9 25 $0.5295
Total 44.3s avg 6.3 avg 13.3 avg $0.9604

Time/Cost columns show % change vs historical average (↑slower/costlier, ↓faster/cheaper). Changes under 10% shown as ±0%.

Historical Comparison Details

Filter: excluding branch 'master'

Status: Success - 21 test/model combinations loaded

Experiments compared (30):

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/bash-resource-checker-evals-fSHoE -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

bash-resource-checker, benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/bash-resource-checker-evals-fSHoE -f markers=regression -f filter=

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants