Skip to content

Add test fixtures for Kubernetes scaling and resource constraint scenarios - #1510

Open
aantn wants to merge 7 commits into
masterfrom
claude/latency-scenario-evals-MHcZ3
Open

aantn wants to merge 7 commits into
masterfrom
claude/latency-scenario-evals-MHcZ3

Conversation

@aantn

@aantn aantn commented Feb 6, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR adds five new test case fixtures for the test_ask_holmes suite, covering common Kubernetes scaling and resource constraint failure scenarios. These fixtures enable testing of the Holmes LLM's ability to diagnose complex chain-of-causation issues in Kubernetes environments.

Key Changes

New Test Cases Added

  1. Test 212: Resource Quota Scaling Blocked

    • Scenario: HPA wants to scale up but ResourceQuota limits prevent new pods from being created
    • Tests diagnosis of quota exhaustion blocking horizontal scaling
    • Includes tight CPU/memory quota constraints and scaling attempts
  2. Test 213: Pending Pods Due to Node Pressure

    • Scenario: Pods remain in Pending state due to insufficient node resources
    • Tests identification of resource pressure preventing pod scheduling
    • Simulates high CPU request requirements exceeding node capacity
  3. Test 214: OOMKill During Scaling

    • Scenario: Pods are repeatedly killed due to memory limits being exceeded
    • Tests detection of OOMKill events and memory pressure diagnosis
    • Includes a memory-hungry application with tight memory limits
  4. Test 215: HPA Maxed Out

    • Scenario: HPA reaches maximum replica count while demand continues to increase
    • Tests identification of HPA saturation and scaling ceiling
    • Includes Prometheus monitoring setup for metrics-based diagnosis
    • Generates CPU load to demonstrate inability to scale further

Implementation Details

  • Each test case includes:

    • test_case.yaml: User prompt, expected outputs, and setup/teardown scripts
    • toolsets.yaml: Enabled diagnostic tools (kubernetes/core, kubernetes/logs, prometheus/metrics for test 215)
    • Kubernetes manifests: Deployments, Services, HPAs, ResourceQuotas, and Secrets as needed
    • Comprehensive setup scripts with validation checkpoints and timeout handling
    • Proper namespace cleanup in after_test hooks
  • Test 215 includes a shared Prometheus deployment reference for metrics collection

  • All tests use realistic application scenarios (order-gateway, payment-processor, inventory-service, notification-service)

  • Setup scripts include detailed logging and error handling for debugging

Tags

All tests are tagged as kubernetes, medium complexity, and chain-of-causation to reflect their focus on multi-step failure diagnosis.

https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok

Summary by CodeRabbit

  • Tests
    • Added five end-to-end Kubernetes troubleshooting test scenarios (namespaces app-212–app-216) covering resource-quota constrained scaling, pending pods/node pressure, OOMKilled during scaling, HPA maxed-out, and HPA ceiling latency. Each fixture includes deployments, HPAs, services, secrets (embedded app code), traffic generators, decoy/supporting workloads, Prometheus configs, and orchestration scripts for setup, monitoring, and teardown.

New eval tests targeting real-world latency scenarios where scaling
can't keep up with demand:

- 212: ResourceQuota blocks HPA scale-up, pods can't be created
- 213: Pending pods due to insufficient node CPU resources
- 214: OOMKill during scaling creates restart loops
- 215: HPA at maxReplicas with Prometheus metrics showing saturation

All tests use neutral service names (order-gateway, payment-processor,
inventory-service, notification-service) and vague latency prompts
to test causal chain reasoning from surface symptom to root cause.

https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok
Signed-off-by: Claude <noreply@anthropic.com>
@netlify

netlify Bot commented Feb 6, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 6c65772
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69874bb8ef72e00008ba81c9
😎 Deploy Preview https://deploy-preview-1510--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@linux-foundation-easycla

linux-foundation-easycla Bot commented Feb 6, 2026 •

Copy link
Copy Markdown

CLA Not Signed

@github-actions

github-actions Bot commented Feb 6, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 252eaa2 (#21768849052)

✅ Results of HolmesGPT evals

Automatically triggered by commit 252eaa2 on branch claude/latency-scenario-evals-MHcZ3

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 12/14 test cases were successful, 2 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 34.7s 6 11 $0.2413
✅ 101_loki_historical_logs_pod_deleted 38.3s 5 9 $0.2411
✅ 111_pod_names_contain_service 33.2s 5 11 $0.2246
✅ 112_find_pvcs_by_uuid 33.5s 6 6 $0.2179
✅ 12_job_crashing 40.8s 5 13 $0.2547
✅ 176_network_policy_blocking_traffic_no_runbooks 44.9s 6 15 $0.2591
✅ 212_resource_quota_scaling_blocked[0] 45.3s 5 13 $0.2850
✅ 213_pending_pods_node_pressure[0] 47.7s 6 14 $0.3013
✅ 214_oomkill_during_scaling[0] 39.3s 5 13 $0.2668
❌ 215_hpa_maxed_out[0] 159.9s 24 58 $1.1789
❌ 216_hpa_ceiling_latency[0] 66.0s 6 21 $0.3878
✅ 24_misconfigured_pvc 38.8s 7 13 $0.2518
✅ 43_current_datetime_from_prompt 4.3s 1 — $0.0112
✅ 61_exact_match_counting 18.0s 4 4 $0.1592
Total 46.1s avg 6.5 avg 15.5 avg $4.2805

⚠️ 2 Failures Detected

📜 Run @ 9138a8a (#21768687662)

✅ Results of HolmesGPT evals

Automatically triggered by commit 9138a8a on branch claude/latency-scenario-evals-MHcZ3

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 12/14 test cases were successful, 2 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 33.2s 5 11 $0.2286
✅ 101_loki_historical_logs_pod_deleted 54.2s 7 14 $0.3005
✅ 111_pod_names_contain_service 35.6s 5 11 $0.2308
✅ 112_find_pvcs_by_uuid 31.1s 5 7 $0.2536
✅ 12_job_crashing 34.0s 5 11 $0.2271
✅ 176_network_policy_blocking_traffic_no_runbooks 47.4s 6 18 $0.2903
✅ 212_resource_quota_scaling_blocked[0] 63.0s 8 21 $0.4196
✅ 213_pending_pods_node_pressure[0] 43.1s 5 13 $0.2845
✅ 214_oomkill_during_scaling[0] 46.9s 6 16 $0.2944
❌ 215_hpa_maxed_out[0] 138.0s 18 55 $1.0281
❌ 216_hpa_ceiling_latency[0] 88.5s 13 37 $0.6256
✅ 24_misconfigured_pvc 33.5s 5 12 $0.2211
✅ 43_current_datetime_from_prompt 5.1s 1 — $0.1050
✅ 61_exact_match_counting 17.5s 4 4 $0.1595
Total 47.9s avg 6.6 avg 17.7 avg $4.6687

⚠️ 2 Failures Detected

📜 Run @ e87aa10 (#21767736258)

✅ Results of HolmesGPT evals

Automatically triggered by commit e87aa10 on branch claude/latency-scenario-evals-MHcZ3

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/14 test cases were successful, 2 regressions, 1 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 29.6s 5 10 $0.2210
✅ 101_loki_historical_logs_pod_deleted 46.1s 7 10 $0.2646
✅ 111_pod_names_contain_service 34.4s 6 11 $0.2356
✅ 112_find_pvcs_by_uuid 31.8s 6 7 $0.2333
✅ 12_job_crashing 32.6s 5 12 $0.2418
✅ 176_network_policy_blocking_traffic_no_runbooks 41.4s 6 17 $0.2870
🚧 212_resource_quota_scaling_blocked[0] — — — —
✅ 213_pending_pods_node_pressure[0] 45.4s 5 15 $0.3168
✅ 214_oomkill_during_scaling[0] 38.6s 5 13 $0.2520
❌ 215_hpa_maxed_out[0] 113.2s 14 52 $0.8710
❌ 216_hpa_ceiling_latency[0] 57.7s 7 18 $0.3739
✅ 24_misconfigured_pvc 38.0s 7 14 $0.2517
✅ 43_current_datetime_from_prompt 4.9s 1 — $0.1058
✅ 61_exact_match_counting 16.3s 4 4 $0.1600
Total 40.8s avg 6.0 avg 15.2 avg $3.8146

⚠️ 2 Failures Detected

📜 Run @ 4eaad45 (#21758709544)

✅ Results of HolmesGPT evals

Automatically triggered by commit 4eaad45 on branch claude/latency-scenario-evals-MHcZ3

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/14 test cases were successful, 2 regressions, 3 setup failures
Status Test case Time Turns Tools Cost
✅ 09_crashpod 27.6s 4 10 $0.2093
✅ 101_loki_historical_logs_pod_deleted 54.2s 8 11 $0.3024
✅ 111_pod_names_contain_service 33.7s 5 11 $0.2204
✅ 112_find_pvcs_by_uuid 38.6s 7 8 $0.2538
✅ 12_job_crashing 40.9s 6 15 $0.2742
✅ 176_network_policy_blocking_traffic_no_runbooks 50.8s 6 18 $0.2925
🚧 212_resource_quota_scaling_blocked[0] — — — —
✅ 213_pending_pods_node_pressure[0] 47.1s 5 16 $0.3052
❌ 214_oomkill_during_scaling[0] 80.2s 10 26 $0.4753
❌ 215_hpa_maxed_out[0] 154.0s 20 53 $1.2317
🚧 216_hpa_ceiling_latency[0] — — — —
✅ 24_misconfigured_pvc 35.5s 6 12 $0.2336
✅ 43_current_datetime_from_prompt 4.9s 1 — $0.1048
🚧 61_exact_match_counting — — — —
Total 51.6s avg 7.1 avg 18.0 avg $3.9032

⚠️ 2 Failures Detected

📜 Run @ f7fd3aa (#21755906121)

✅ Results of HolmesGPT evals

Automatically triggered by commit f7fd3aa on branch claude/latency-scenario-evals-MHcZ3

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 12/13 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 35.2s 5 11 $0.2378
✅ 101_loki_historical_logs_pod_deleted 36.7s 5 9 $0.2390
✅ 111_pod_names_contain_service 31.8s 5 10 $0.2203
✅ 112_find_pvcs_by_uuid 33.7s 6 8 $0.2474
✅ 12_job_crashing 39.1s 6 14 $0.2617
✅ 176_network_policy_blocking_traffic_no_runbooks 45.3s 6 16 $0.2949
✅ 212_resource_quota_scaling_blocked[0] 43.1s 6 16 $0.2793
✅ 213_pending_pods_node_pressure[0] 40.3s 5 12 $0.2654
✅ 214_oomkill_during_scaling[0] 43.1s 5 14 $0.2570
❌ 215_hpa_maxed_out[0] 136.3s 22 45 $0.9846
✅ 24_misconfigured_pvc 31.8s 5 13 $0.2290
✅ 43_current_datetime_from_prompt 4.4s 1 — $0.1051
✅ 61_exact_match_counting 17.2s 4 4 $0.1599
Total 41.4s avg 6.2 avg 14.3 avg $3.7815

⚠️ 1 Failure Detected


✅ Results of HolmesGPT evals

Automatically triggered by commit 6c65772 on branch claude/latency-scenario-evals-MHcZ3

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/11 test cases were successful, 2 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 33.0s 5 11 $0.2293
✅ 101_loki_historical_logs_pod_deleted 31.9s 4 7 $0.2048
✅ 111_pod_names_contain_service 27.6s 4 9 $0.1969
✅ 112_find_pvcs_by_uuid 34.6s 6 8 $0.2576
✅ 12_job_crashing 32.0s 5 11 $0.2327
✅ 176_network_policy_blocking_traffic_no_runbooks 43.2s 7 17 $0.2993
❌ 215_hpa_maxed_out[0] 102.3s 17 35 $0.7137
❌ 216_hpa_ceiling_latency[0] 106.5s 9 29 $0.5008
✅ 24_misconfigured_pvc 39.8s 7 14 $0.2559
✅ 43_current_datetime_from_prompt 4.7s 1 — $0.0112
✅ 61_exact_match_counting 13.3s 3 2 $0.1427
Total 42.6s avg 6.2 avg 14.3 avg $3.0450

⚠️ 2 Failures Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/latency-scenario-evals-MHcZ3 -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/latency-scenario-evals-MHcZ3 -f markers=regression -f filter=

@github-actions

github-actions Bot commented Feb 6, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for 89d2589 (built in 5m 17s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:89d2589
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:89d2589 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:89d2589
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:89d2589

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:89d2589

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:89d2589

@coderabbitai

coderabbitai Bot commented Feb 6, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Adds multiple Kubernetes end-to-end test fixtures (namespaces app-212–app-216) that simulate HPA and resource constraint scenarios: manifests (Deployments, Services, Secrets, ConfigMaps, HPAs), traffic generators, Prometheus configs, and orchestrated test_case.yaml scripts for automated setup, load, observation, and teardown.

Changes

Cohort / File(s) Summary
Test Fixture 212: Resource Quota Scaling Blocked
tests/llm/fixtures/test_ask_holmes/212_resource_quota_scaling_blocked/*
Adds ResourceQuota, HPA, order-gateway Deployment/Service/Secret (Flask app), decoy services, traffic-generator, toolsets, and test_case.yaml orchestrating namespace creation, load, HPA apply, quota observation, and teardown.
Test Fixture 213: Pending Pods Node Pressure
tests/llm/fixtures/test_ask_holmes/213_pending_pods_node_pressure/*
Adds payment-processor Deployment/Service, HPA, supporting-services, noisy-sidecar (secret + deployment), toolsets, and test_case.yaml that scales to induce Pending pods and collects pod/event state.
Test Fixture 214: OOMKill During Scaling
tests/llm/fixtures/test_ask_holmes/214_oomkill_during_scaling/*
Adds inventory-service v1/v2 (Secrets + Deployments simulating memory leak), supporting-services, HPA, toolsets, and test_case.yaml to rollout v2 under HPA and detect OOMKilled events.
Test Fixture 215: HPA Maxed Out
tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/*
Adds notification-service (Secret, Deployment with Prometheus metrics, Service), Prometheus ConfigMap, HPA (min=max=3), analytics-worker stress Deployment, toolsets, and test_case.yaml to drive HPA to maxReplicas and collect metrics/state.
Test Fixture 216: HPA Ceiling Latency
tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/*
Adds dispatch-service (Secret with Flask app and Prometheus metrics, Deployment, Service), Prometheus ConfigMap, HPA (max 3), supporting-services, traffic-generator, toolsets, and test_case.yaml to saturate HPA and observe latency/metrics.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

Suggested reviewers

  • moshemorad
  • RoiGlinik
  • Sheeproid
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: adding test fixtures for Kubernetes scaling and resource constraint scenarios across five new test cases (212-216).
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Feb 6, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 10.11s 9.67s +4.5%
Warm Mean 4.72s 4.59s +2.7%
Warm Min 4.63s 4.57s
Warm Max 4.89s 4.65s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 24.88s 15.79s +57.6%
Warm Mean 7.05s 7.04s +0.1%
Warm Min 6.74s 6.53s
Warm Max 7.40s 7.63s

PR: 89d2589e | Master: 76dbfc7e | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In
`@tests/llm/fixtures/test_ask_holmes/212_resource_quota_scaling_blocked/test_case.yaml`:
- Around line 53-66: The quota-event wait loop (the for loop that checks kubectl
get events -n app-212 for reason=FailedCreate and sets EVENTS_FOUND) currently
only logs a warning when no "exceeded quota" events are found; make this failure
fatal like the pod readiness check by calling exit 1 when EVENTS_FOUND is still
false after the loop (or, alternatively, add an explicit comment if this
non-fatal behavior is deliberate). Locate the loop that greps for "exceeded
quota" and the subsequent if [ "$EVENTS_FOUND" = false ]; then block and replace
the warning-only branch with a clear failing path that echoes a descriptive
error and runs exit 1 so the test fails fast on missing quota events.

In `@tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/cpu-stress.yaml`:
- Around line 3-7: The resource name cpu-stress in the manifest should be
changed to a neutral, realistic service name (e.g., analytics-worker or
request-client) to avoid revealing test intent; update the metadata.name value
and also update any matching labels/selectors (labels.app, pod template labels,
and any Deployment/Service selector references) so they remain consistent with
the new name (look for occurrences of "cpu-stress", metadata.name, labels.app
and selector/template labels in this file and associated manifests and rename
them to the chosen neutral identifier).
🧹 Nitpick comments (3)
tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/prometheus-config.yaml (1)

19-32: Redundant relabel rule on lines 23-27 — immediately overwritten by lines 28-32.

The second relabel rule sets __address__ to the bare port annotation value (e.g., "8080"), but the third rule immediately overwrites it with pod_ip:port. The second rule is dead config and could leave an invalid __address__ (just a port number) if __meta_kubernetes_pod_ip were ever unset.

Remove the intermediate rule:

Suggested fix
       relabel_configs:
       - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
         action: keep
         regex: true
-      - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port]
-        action: replace
-        target_label: __address__
-        regex: (.+)
-        replacement: ${1}
       - source_labels: [__meta_kubernetes_pod_ip, __meta_kubernetes_pod_annotation_prometheus_io_port]
         action: replace
         target_label: __address__
         regex: (.+);(.+)
         replacement: ${1}:${2}
tests/llm/fixtures/test_ask_holmes/214_oomkill_during_scaling/inventory-service.yaml (1)

37-57: Consider adding a securityContext to avoid running as root.

Checkov flags CKV_K8S_20 and CKV_K8S_23 for missing allowPrivilegeEscalation: false and runAsNonRoot. For a test fixture using busybox, this is low priority, but adding it would silence static analysis and follow K8s hardening best practices.

🔒 Optional securityContext addition
       containers:
       - name: inventory-service
         image: busybox:1.36
         command: ["/bin/sh", "/app/app.sh"]
+        securityContext:
+          allowPrivilegeEscalation: false
+          runAsNonRoot: true
+          runAsUser: 1000
         ports:
         - containerPort: 8080
tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/notification-service.yaml (1)

78-105: Optional: add securityContext to silence Checkov warnings.

Same as the inventory-service manifest — Checkov flags CKV_K8S_20 and CKV_K8S_23. Low priority for test fixtures but good hygiene.

🔒 Optional securityContext addition
       containers:
       - name: notification-service
         image: me-west1-docker.pkg.dev/robusta-development/development/python-flask-otel:2.2
         imagePullPolicy: IfNotPresent
         command: ["python", "/app/app.py"]
+        securityContext:
+          allowPrivilegeEscalation: false
         ports:
         - containerPort: 8080

Comment on lines +3 to +7
metadata:
name: cpu-stress
namespace: app-215
labels:
app: cpu-stress

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion | 🟠 Major

Resource name cpu-stress hints at the problem — use a neutral name.

The name directly reveals this deployment's role in the test scenario. Consider a real-world service name (e.g., analytics-worker, request-client, or similar) to keep the fixture realistic and not leak test intent to the LLM under evaluation.

Example rename
 metadata:
-  name: cpu-stress
+  name: analytics-worker
   namespace: app-215
   labels:
-    app: cpu-stress
+    app: analytics-worker

(Also update selector/template labels accordingly.)

As per coding guidelines, tests/llm/fixtures/test_ask_holmes/**/*.yaml: "Use real-world, neutral resource names that don't hint at problems (avoid names like 'broken-pod', 'crashloop-app')". Based on learnings: "Use real-world, neutral resource names that don't hint at problems."

🤖 Prompt for AI Agents
In `@tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/cpu-stress.yaml` around
lines 3 - 7, The resource name cpu-stress in the manifest should be changed to a
neutral, realistic service name (e.g., analytics-worker or request-client) to
avoid revealing test intent; update the metadata.name value and also update any
matching labels/selectors (labels.app, pod template labels, and any
Deployment/Service selector references) so they remain consistent with the new
name (look for occurrences of "cpu-stress", metadata.name, labels.app and
selector/template labels in this file and associated manifests and rename them
to the chosen neutral identifier).

@aantn

aantn commented Feb 6, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
marker: regression
model: sonnet-4.5

@github-actions

github-actions Bot commented Feb 6, 2026

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ⚠️ Completed with 1 failure


🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Branch claude/latency-scenario-evals-MHcZ3
Model sonnet-4.5
Markers regression
Iterations 1
Duration 5m 36s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 12/13 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 30.8s 5 11 $0.1578
✅ 101_loki_historical_logs_pod_deleted 32.1s 5 9 $0.1575
✅ 111_pod_names_contain_service 40.7s 8 16 $0.2068
✅ 112_find_pvcs_by_uuid 24.7s 5 7 $0.1567
✅ 12_job_crashing 35.0s 7 14 $0.1815
✅ 176_network_policy_blocking_traffic_no_runbooks 48.9s 9 22 $0.2477
✅ 212_resource_quota_scaling_blocked[0] 46.2s 7 17 $0.2202
✅ 213_pending_pods_node_pressure[0] 56.4s 8 19 $0.2400
✅ 214_oomkill_during_scaling[0] 40.0s 6 14 $0.1945
❌ 215_hpa_maxed_out[0] 110.2s 21 44 $0.5214
✅ 24_misconfigured_pvc 31.5s 6 14 $0.1686
✅ 43_current_datetime_from_prompt 4.9s 1 — $0.0704
✅ 61_exact_match_counting 18.8s 3 2 $0.1071
Total 40.0s avg 7.0 avg 15.8 avg $2.6301

⚠️ 1 Failure Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/latency-scenario-evals-MHcZ3 -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/latency-scenario-evals-MHcZ3 -f markers=regression -f filter=

212 (ResourceQuota): Replace manual scale with real Flask app + traffic
generator that drives HPA organically. Add decoy services (session-cache,
order-db) consuming quota to make the failure less obvious. Holmes must
dig into ReplicaSet events to find the quota error.

213 (Pending pods): Add 3 supporting services (payment-cache, payment-queue,
payment-db) as noise in kubectl output. Add fraud-detection service that
logs connection pool warnings as a red herring. Holmes must distinguish
noisy-but-running pods from the actual Pending pods.

214 (OOMKill): Replace single deployment with v1→v2 rolling update.
V1 is stable and works fine. V2 has a memory leak that OOMKills.
Holmes sees a mid-rollout state and must figure out the new version
has a memory bug, not just a slow rollout. Add supporting services.

216 (NEW - HPA ceiling): Clean variant of 215 using gunicorn instead of
Flask dev server, eliminating the competing red herring. Add decoy
services. The ONLY path to the answer is noticing HPA at maxReplicas.

All tests use include_tool_calls: true to verify Holmes actually queries
the right resources rather than guessing from domain knowledge.

https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In
`@tests/llm/fixtures/test_ask_holmes/213_pending_pods_node_pressure/test_case.yaml`:
- Around line 47-54: The supporting-services wait loop for pods labeled
tier=infra in namespace app-213 currently falls through silently if pods never
become ready; update the loop (the for ...; do ... kubectl wait
--for=condition=ready pod -l tier=infra -n app-213 --timeout=5s ...) to mirror
the payment-processor pattern by adding a post-loop check that detects failure
(e.g., if the kubectl wait never succeeded) and prints an error like "❌
Supporting services failed to become ready" and exits non-zero (exit 1) so the
test fails fast and clearly.

In
`@tests/llm/fixtures/test_ask_holmes/214_oomkill_during_scaling/inventory-service-v2.yaml`:
- Around line 15-23: The loop's DATA variable is overwritten each iteration so
memory doesn't truly accumulate; update the script to either append to DATA (use
DATA="$DATA$(dd ... | base64)") to force deterministic accumulation and trigger
OOMKill, or change the comment near BATCH/DATA/while true to remove the
misleading claim about additive growth and note that OOM is caused by transient
peak during the dd|base64 capture; modify the block referencing DATA, BATCH and
the while loop accordingly.
🧹 Nitpick comments (3)
tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/test_case.yaml (1)

98-114: HPA wait loop does not fail the setup on timeout.

The HPA wait loop logs a warning but does not exit 1 when the HPA fails to reach maxReplicas within 90 iterations. This is fine if the test is designed to be resilient to partial setup, but note that the dispatch-service starts with replicas: 3 and the HPA has minReplicas: 3 / maxReplicas: 3, so this condition should be trivially satisfied almost immediately without needing traffic. If this loop consistently times out, it may indicate a deeper issue with the HPA controller not reporting currentReplicas.

tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/dispatch-service.yaml (1)

38-52: Nit: import hashlib inside the loop.

The hashlib import on line 44 (within the app.py string) is inside the for loop. While CPython caches modules in sys.modules so this is functionally harmless, moving it to the top of the file alongside the other imports would be more idiomatic.

Proposed fix
     from flask import Flask, Response, jsonify
     from prometheus_client import Counter, Histogram, generate_latest, CONTENT_TYPE_LATEST
     import time
     import random
     import os
+    import hashlib
     ...
         for _ in range(300):
-            import hashlib
             data = hashlib.sha256(data.encode()).hexdigest()
tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/prometheus-config.yaml (1)

23-27: Redundant relabel rule — immediately overwritten by the next rule.

Rule at lines 23–27 sets __address__ to just the port value (e.g., "8080"), but the next rule (lines 28–32) unconditionally overwrites __address__ with pod_ip:port. This first replacement has no effect and adds confusion.

Remove the redundant rule
       - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
         action: keep
         regex: true
-      - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port]
-        action: replace
-        target_label: __address__
-        regex: (.+)
-        replacement: ${1}
       - source_labels: [__meta_kubernetes_pod_ip, __meta_kubernetes_pod_annotation_prometheus_io_port]
         action: replace
         target_label: __address__
         regex: (.+);(.+)
         replacement: ${1}:${2}

Comment thread tests/llm/fixtures/test_ask_holmes/213_pending_pods_node_pressure/test_case.yaml Outdated
- 212: Add exit 1 on quota event fallback failure (was silently continuing)
- 213: Add exit 1 on supporting services wait failure (was falling through)
- 214: Fix memory accumulation bug (CACHE grows across iterations instead
  of overwriting DATA each iteration), add exit 1 on OOMKill/infra waits
- 215: Rename cpu-stress to analytics-worker (name revealed scenario),
  fix redundant Prometheus relabel rule
- 216: Remove gunicorn (not in image), use Flask threaded mode instead,
  move hashlib import to top level, fix redundant Prometheus relabel rule

https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Fix all issues with AI agents
In `@tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/test_case.yaml`:
- Around line 100-103: The current HPA_MAXED conditional only emits a warning;
change it to a hard failure so the test fails when HPA doesn't reach
maxReplicas: replace the soft echo/inspection branch that checks HPA_MAXED with
a failing branch that prints a descriptive error (including kubectl get hpa -n
app-215 output for diagnostics) and exits non-zero (e.g., exit 1) similar to the
pod readiness checks; update any test documentation only if you intentionally
keep it soft, but prefer converting the HPA_MAXED check into a hard failure to
ensure the core assertion is enforced.
- Around line 75-85: The wait loop that checks for analytics-worker stress pods
(loop using RUNNING, STRESS_READY, label app=analytics-worker in namespace
app-215) lacks an exit-on-failure like the other loops; after the for loop
finishes, check if STRESS_READY is still false and if so print an error and call
exit 1 so the test fails early; update the block that sets STRESS_READY and the
surrounding loop to mirror the behavior used for notification-service and
Prometheus waits.

In
`@tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/dispatch-service.yaml`:
- Around line 18-19: The code retrieves the Werkzeug logger into the variable
cli but never configures it, so the intended suppression is a no-op; update the
code that obtains the logger (cli = logging.getLogger('werkzeug')) to explicitly
set an appropriate level (e.g., cli.setLevel(logging.ERROR) or logging.WARNING)
and/or attach a NullHandler or desired handler to suppress the Flask dev server
warning (use cli.addHandler(logging.NullHandler()) or similar) so the logger is
actually silenced during tests.
🧹 Nitpick comments (3)
tests/llm/fixtures/test_ask_holmes/214_oomkill_during_scaling/inventory-service-v2.yaml (1)

51-67: Consider adding a minimal securityContext to the container spec.

Static analysis flags that the container runs without allowPrivilegeEscalation: false and runAsNonRoot: true. While this is a test fixture, adding a security context is low-effort and aligns with best practices. The busybox script doesn't need root.

🔒 Proposed securityContext addition
       - name: inventory-service
         image: busybox:1.36
         command: ["/bin/sh", "/app/run.sh"]
+        securityContext:
+          allowPrivilegeEscalation: false
+          runAsNonRoot: true
+          runAsUser: 1000
         ports:
         - containerPort: 8080
tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/dispatch-service.yaml (1)

66-137: Security context not set on the container (Checkov CKV_K8S_20, CKV_K8S_23).

Static analysis flags that allowPrivilegeEscalation isn't explicitly disabled and the container may run as root. For test fixtures this is low-risk, but adding a minimal securityContext is a good habit and silences the warnings.

Suggested securityContext addition
         resources:
           requests:
             memory: "64Mi"
             cpu: "50m"
           limits:
             memory: "128Mi"
             cpu: "100m"
+        securityContext:
+          allowPrivilegeEscalation: false
+          runAsNonRoot: true
+          runAsUser: 1000
tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/analytics-worker.yaml (1)

19-19: Container name "stress" may hint at the problem to the LLM.

The resource name analytics-worker is appropriately neutral, but the container name stress could leak into kubectl describe output that Holmes inspects, potentially hinting at the root cause. Consider renaming it to something neutral like worker or analytics-worker.

Based on learnings: "Use real-world, neutral resource names that don't hint at problems."

Suggested rename
-      - name: stress
+      - name: worker

Comment on lines +75 to +85
echo "⏳ Waiting for stress job pods to start..."
STRESS_READY=false
for i in {1..60}; do
RUNNING=$(kubectl get pods -n app-215 -l app=analytics-worker --no-headers 2>/dev/null | grep -c "Running" || echo "0")
if [ "$RUNNING" -gt "0" ]; then
echo "✅ Stress generators running!"
STRESS_READY=true
break
fi
sleep 1
done

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Missing exit-on-failure when stress pods fail to start.

The other wait loops (notification-service pods at Line 47, Prometheus at Line 63) exit with exit 1 on failure. This block silently continues if no analytics-worker pods reach Running state. If the stress generator never starts, HPA won't scale and the test precondition won't hold.

This is inconsistent with the commit message that mentions "Added explicit exit-on-failure behavior."

Proposed fix
     sleep 1
   done
+  if [ "$STRESS_READY" = false ]; then
+    echo "❌ Stress generators failed to start"
+    kubectl get pods -n app-215
+    exit 1
+  fi
 
   echo "⏳ Waiting for HPA to hit maxReplicas and stabilize..."
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
echo "⏳ Waiting for stress job pods to start..."
STRESS_READY=false
for i in {1..60}; do
RUNNING=$(kubectl get pods -n app-215 -l app=analytics-worker --no-headers 2>/dev/null | grep -c "Running" || echo "0")
if [ "$RUNNING" -gt "0" ]; then
echo "✅ Stress generators running!"
STRESS_READY=true
break
fi
sleep 1
done
echo "⏳ Waiting for stress job pods to start..."
STRESS_READY=false
for i in {1..60}; do
RUNNING=$(kubectl get pods -n app-215 -l app=analytics-worker --no-headers 2>/dev/null | grep -c "Running" || echo "0")
if [ "$RUNNING" -gt "0" ]; then
echo "✅ Stress generators running!"
STRESS_READY=true
break
fi
sleep 1
done
if [ "$STRESS_READY" = false ]; then
echo "❌ Stress generators failed to start"
kubectl get pods -n app-215
exit 1
fi
echo "⏳ Waiting for HPA to hit maxReplicas and stabilize..."
🤖 Prompt for AI Agents
In `@tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/test_case.yaml` around
lines 75 - 85, The wait loop that checks for analytics-worker stress pods (loop
using RUNNING, STRESS_READY, label app=analytics-worker in namespace app-215)
lacks an exit-on-failure like the other loops; after the for loop finishes,
check if STRESS_READY is still false and if so print an error and call exit 1 so
the test fails early; update the block that sets STRESS_READY and the
surrounding loop to mirror the behavior used for notification-service and
Prometheus waits.

Comment on lines +100 to +103
if [ "$HPA_MAXED" = false ]; then
echo "⚠️ HPA did not reach maxReplicas within timeout"
kubectl get hpa -n app-215
fi

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

HPA not reaching maxReplicas is only a warning, not a failure.

If HPA doesn't reach maxReplicas, the test's core assertion (that Holmes detects HPA is maxed out) cannot pass. Consider making this a hard failure like the pod readiness checks, or at minimum document why this is intentionally a soft warning.

🤖 Prompt for AI Agents
In `@tests/llm/fixtures/test_ask_holmes/215_hpa_maxed_out/test_case.yaml` around
lines 100 - 103, The current HPA_MAXED conditional only emits a warning; change
it to a hard failure so the test fails when HPA doesn't reach maxReplicas:
replace the soft echo/inspection branch that checks HPA_MAXED with a failing
branch that prints a descriptive error (including kubectl get hpa -n app-215
output for diagnostics) and exits non-zero (e.g., exit 1) similar to the pod
readiness checks; update any test documentation only if you intentionally keep
it soft, but prefer converting the HPA_MAXED check into a hard failure to ensure
the core assertion is enforced.

Comment on lines +18 to +19
# Suppress Flask dev server warning — this is intentional for test purposes
cli = logging.getLogger('werkzeug')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Werkzeug logger retrieved but never configured.

The comment says "Suppress Flask dev server warning" but the logger level is never set. This is a no-op.

Proposed fix
     # Suppress Flask dev server warning — this is intentional for test purposes
-    cli = logging.getLogger('werkzeug')
+    cli = logging.getLogger('werkzeug')
+    cli.setLevel(logging.ERROR)
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
# Suppress Flask dev server warning — this is intentional for test purposes
cli = logging.getLogger('werkzeug')
# Suppress Flask dev server warning — this is intentional for test purposes
cli = logging.getLogger('werkzeug')
cli.setLevel(logging.ERROR)
🤖 Prompt for AI Agents
In
`@tests/llm/fixtures/test_ask_holmes/216_hpa_ceiling_latency/dispatch-service.yaml`
around lines 18 - 19, The code retrieves the Werkzeug logger into the variable
cli but never configures it, so the intended suppression is a no-op; update the
code that obtains the logger (cli = logging.getLogger('werkzeug')) to explicitly
set an appropriate level (e.g., cli.setLevel(logging.ERROR) or logging.WARNING)
and/or attach a NullHandler or desired handler to suppress the Flask dev server
warning (use cli.addHandler(logging.NullHandler()) or similar) so the logger is
actually silenced during tests.

claude and others added 3 commits February 6, 2026 22:49
…erations

The HPA organic scaling loop (120 × 2s = 240s) consumed nearly the entire
300s setup timeout before the forced fallback could run. In Kind CI/CD,
metrics-server is slow to provide CPU metrics, so the HPA rarely scales
organically. Reducing to 30 iterations (60s) leaves ample time for the
reliable forced kubectl scale fallback.

https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok
Signed-off-by: Claude <noreply@anthropic.com>
All three tests (ResourceQuota scaling blocked, Pending pods node pressure,
OOMKill during scaling) pass consistently on latest models. Only keeping
tests 215 and 216 which are hard enough to fail.

https://claude.ai/code/session_01DUAGjg9ovrc7BX2TWLXwok
Signed-off-by: Claude <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants