Skip to content

Parallelize KIND cluster setup with HolmesGPT environment - #1687

Merged
aantn merged 15 commits into
masterfrom
claude/speed-up-regression-evals-TfJav
Mar 8, 2026
Merged

aantn merged 15 commits into
masterfrom
claude/speed-up-regression-evals-TfJav

Conversation

@aantn

@aantn aantn commented Mar 6, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Optimize CI/CD pipeline performance by running KIND cluster setup in parallel with HolmesGPT environment setup, reducing overall workflow execution time by approximately 70 seconds.

Key Changes

  • Parallel KIND Setup: KIND cluster initialization now runs in the background as a separate shell script (/tmp/kind-setup.sh) instead of blocking on a dedicated action step

    • Starts immediately after the test check step
    • Runs concurrently with HolmesGPT environment setup
    • Includes complete cluster configuration: KIND installation, Calico CNI, Prometheus stack, and metrics-server
  • Background Process Management:

    • KIND setup script runs in background with output logged to /tmp/kind-setup.log
    • Process ID stored in /tmp/kind-setup-pid for tracking
    • New "Wait for KIND cluster setup" step polls for completion and validates success via log sentinel
  • Improved Helm Installations: Prometheus and metrics-server installations now run in parallel within the KIND setup script using background processes (&) to reduce setup time

  • Dynamic Test Parallelism:

    • Pytest worker count now scales based on actual test count (min 6, max 20, capped at CPU count)
    • Replaces hardcoded -n10 with dynamic -n"$N_WORKERS" calculation
    • Prevents resource contention on systems with fewer cores
  • Enhanced Setup Action:

    • Added Python dependency caching layer in setup-holmes-env action using GitHub Actions cache
    • Caches pip and pypoetry directories keyed by Python version and poetry.lock hash
    • Reduces dependency installation time on cache hits

Implementation Details

  • KIND setup script uses heredoc syntax to avoid indentation issues and is written to disk before execution
  • Calico CNI polling intervals reduced from 5s to 2s for faster convergence
  • Process synchronization uses wait command for parallel Helm installations
  • Cluster readiness validation includes node status, kube-system pods, and calico-system pods checks

https://claude.ai/code/session_01ReC3SkRaWTNe3BZyuhLRyx

Summary by CodeRabbit

  • Chores
    • CI Python setup: install Poetry, switch to an in-repo virtual environment, add the virtualenv to PATH, and cache the virtualenv to speed and stabilize dependency installation.
    • CI environment provisioning and test runs: background cluster provisioning with cacheable artifacts, sentinel-based readiness checks, parallelized component installs, improved logging, and increased test parallelism (pytest -n20) for faster, more reliable CI.

- Add pip/poetry dependency caching to reduce Python install time
- Run KIND cluster setup in background, parallel with Holmes env setup
  (saves ~70s by overlapping the two independent setup phases)
- Install kube-prometheus-stack and metrics-server in parallel within
  KIND setup (saves ~26s)
- Reduce Calico CNI wait loop intervals from 5s to 2s
- Remove unnecessary test pod verification step from cluster readiness
- Scale pytest parallelism dynamically based on test count (6-20 workers)

https://claude.ai/code/session_01ReC3SkRaWTNe3BZyuhLRyx
Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Mar 6, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 586b55d (#22806196276)

✅ Results of HolmesGPT evals

Automatically triggered by commit 586b55d on branch claude/speed-up-regression-evals-TfJav

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 30.0s 5 11 $0.1464 106,910 104,997 1,913 96,721 8,276 — 494 —
✅ 101_loki_historical_logs_pod_deleted 37.1s 5 11 $0.1719 111,573 109,133 2,440 99,129 10,004 — 797 —
✅ 111_pod_names_contain_service 30.1s 5 12 $0.1446 105,894 103,909 1,985 96,192 7,717 — 489 —
✅ 112_find_pvcs_by_uuid 32.4s 6 10 $0.1906 140,174 138,156 2,018 125,351 12,805 — 460 —
✅ 12_job_crashing 28.5s 5 12 $0.1431 109,082 107,260 1,822 99,625 7,635 — 453 —
✅ 176_network_policy_blocking_traffic_no_runbooks 35.8s 6 16 $0.1929 139,226 136,812 2,414 125,276 11,536 — 634 —
✅ 24_misconfigured_pvc 35.5s 7 13 $0.1570 144,736 142,873 1,863 135,932 6,941 — 453 —
✅ 43_current_datetime_from_prompt 5.0s 1 — $0.0124 17,535 17,388 147 17,379 9 — 147 —
✅ 61_exact_match_counting 15.7s 4 4 $0.0715 77,302 76,748 554 73,394 3,354 — 256 —
Total 27.8s avg 4.9 avg 11.1 avg $1.2304 952,432 937,276 15,156 868,999 68,277 — 797 —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 61 test/model combinations loaded

Benchmark experiment:

Time comparison (seconds):

Test case This branch master Diff
09_crashpod (opus-4.5) 30.0s 29.5s ±0%
101_loki_historical_logs_pod_deleted (opus-4.5) 37.1s — —
111_pod_names_contain_service (opus-4.5) 30.1s 34.2s ↓12%
112_find_pvcs_by_uuid (opus-4.5) 32.4s 31.0s ±0%
12_job_crashing (opus-4.5) 28.5s 29.7s ±0%
176_network_policy_blocking_traffic_no_runbooks (opus-4.5) 35.8s 42.7s ↓16%
24_misconfigured_pvc (opus-4.5) 35.5s 38.7s ±0%
43_current_datetime_from_prompt (opus-4.5) 5.0s — —
61_exact_match_counting (opus-4.5) 15.7s — —

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ b28cc18 (#22806170809)

✅ Results of HolmesGPT evals

Automatically triggered by commit b28cc18 on branch claude/speed-up-regression-evals-TfJav

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 28.7s 5 11 $0.1406 107,267 105,487 1,780 97,822 7,665 — 545 —
✅ 101_loki_historical_logs_pod_deleted 37.9s 5 10 $0.1613 109,916 107,608 2,308 98,745 8,863 — 795 —
✅ 111_pod_names_contain_service 31.1s 5 12 $0.1416 106,282 104,351 1,931 97,044 7,307 — 611 —
✅ 112_find_pvcs_by_uuid 31.6s 6 8 $0.1555 132,596 130,876 1,720 122,695 8,181 — 449 —
✅ 12_job_crashing 35.8s 6 14 $0.1726 134,833 132,672 2,161 123,349 9,323 — 501 —
✅ 176_network_policy_blocking_traffic_no_runbooks 43.6s 7 17 $0.2098 159,793 157,208 2,585 145,340 11,868 — 704 —
✅ 24_misconfigured_pvc 30.5s 5 14 $0.1484 106,243 104,282 1,961 95,635 8,647 — 675 —
✅ 43_current_datetime_from_prompt 7.8s 1 — $0.0221 17,787 17,388 399 16,789 599 — 399 —
✅ 61_exact_match_counting 11.8s 3 2 $0.0530 56,334 55,960 374 53,226 2,734 — 187 —
Total 28.8s avg 4.8 avg 11.0 avg $1.2049 931,051 915,832 15,219 850,645 65,187 — 795 —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 61 test/model combinations loaded

Benchmark experiment:

Time comparison (seconds):

Test case This branch master Diff
09_crashpod (opus-4.5) 28.7s 29.5s ±0%
101_loki_historical_logs_pod_deleted (opus-4.5) 37.9s — —
111_pod_names_contain_service (opus-4.5) 31.1s 34.2s ±0%
112_find_pvcs_by_uuid (opus-4.5) 31.6s 31.0s ±0%
12_job_crashing (opus-4.5) 35.8s 29.7s ↑21%
176_network_policy_blocking_traffic_no_runbooks (opus-4.5) 43.6s 42.7s ±0%
24_misconfigured_pvc (opus-4.5) 30.5s 38.7s ↓21%
43_current_datetime_from_prompt (opus-4.5) 7.8s — —
61_exact_match_counting (opus-4.5) 11.8s — —

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 37d23da (#22805896168)

✅ Results of HolmesGPT evals

Automatically triggered by commit 37d23da on branch claude/speed-up-regression-evals-TfJav

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 29.0s 5 12 $0.1460 108,048 106,170 1,878 98,013 8,157 — 543 —
✅ 101_loki_historical_logs_pod_deleted 51.3s 8 14 $0.2222 182,782 179,923 2,859 169,208 10,715 — 446 —
✅ 111_pod_names_contain_service 30.0s 5 11 $0.1354 104,519 102,681 1,838 95,888 6,793 — 485 —
✅ 112_find_pvcs_by_uuid 26.6s 5 7 $0.1217 102,155 100,677 1,478 94,611 6,066 — 363 —
✅ 12_job_crashing 30.2s 5 12 $0.1499 109,604 107,783 1,821 98,712 9,071 — 440 —
✅ 176_network_policy_blocking_traffic_no_runbooks 45.3s 8 18 $0.2237 183,716 181,048 2,668 169,157 11,891 — 674 —
✅ 24_misconfigured_pvc 28.7s 5 13 $0.1338 104,631 102,894 1,737 95,941 6,953 — 550 —
✅ 43_current_datetime_from_prompt 4.8s 1 — $0.0115 17,499 17,388 111 17,379 9 — 111 —
✅ 61_exact_match_counting 14.7s 4 4 $0.0709 77,202 76,665 537 73,339 3,326 — 239 —
Total 29.0s avg 5.1 avg 11.4 avg $1.2151 990,156 975,229 14,927 912,248 62,981 — 674 —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 61 test/model combinations loaded

Benchmark experiment:

Time comparison (seconds):

Test case This branch master Diff
09_crashpod (opus-4.5) 29.0s 29.5s ±0%
101_loki_historical_logs_pod_deleted (opus-4.5) 51.3s — —
111_pod_names_contain_service (opus-4.5) 30.0s 34.2s ↓12%
112_find_pvcs_by_uuid (opus-4.5) 26.6s 31.0s ↓14%
12_job_crashing (opus-4.5) 30.2s 29.7s ±0%
176_network_policy_blocking_traffic_no_runbooks (opus-4.5) 45.3s 42.7s ±0%
24_misconfigured_pvc (opus-4.5) 28.7s 38.7s ↓26%
43_current_datetime_from_prompt (opus-4.5) 4.8s — —
61_exact_match_counting (opus-4.5) 14.7s — —

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 8f7ea3f (#22795410737)

✅ Results of HolmesGPT evals

Automatically triggered by commit 8f7ea3f on branch claude/speed-up-regression-evals-TfJav

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 29.0s 5 11 $0.2355 106,963 105,251 1,712 80,728 24,523 — 548 —
✅ 101_loki_historical_logs_pod_deleted 36.8s 6 8 $0.2395 122,212 120,223 1,989 97,625 22,598 — 461 —
✅ 111_pod_names_contain_service 27.8s 4 11 $0.2190 83,363 81,603 1,760 58,130 23,473 — 614 —
✅ 112_find_pvcs_by_uuid 26.7s 6 5 $0.2122 117,948 116,734 1,214 95,264 21,470 — 343 —
✅ 12_job_crashing 32.3s 5 12 $0.2467 110,124 108,234 1,890 82,741 25,493 — 560 —
✅ 176_network_policy_blocking_traffic_no_runbooks 42.8s 7 16 $0.2952 155,825 153,417 2,408 125,576 27,841 — 467 —
✅ 24_misconfigured_pvc 32.3s 5 15 $0.2461 107,059 104,958 2,101 80,196 24,762 — 725 —
✅ 43_current_datetime_from_prompt 4.2s 1 — $0.1114 17,499 17,388 111 0 17,388 — 111 —
✅ 61_exact_match_counting 11.7s 3 2 $0.1484 56,173 55,828 345 36,369 19,459 — 157 —
Total 27.1s avg 4.7 avg 10.0 avg $1.9538 877,166 863,636 13,530 656,629 207,007 — 725 —
📜 Run @ 59bd1bf (#22764384077)

✅ Results of HolmesGPT evals

Automatically triggered by commit 59bd1bf on branch claude/speed-up-regression-evals-TfJav

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Tokens Compactions
✅ 09_crashpod 34.4s 5 12 $0.1547 109,248 —
✅ 101_loki_historical_logs_pod_deleted 77.7s 8 17 $0.3520 187,543 —
✅ 111_pod_names_contain_service 28.7s 4 10 $0.1196 82,235 —
✅ 112_find_pvcs_by_uuid 37.1s 8 7 $0.1678 167,536 —
✅ 12_job_crashing 31.5s 5 11 $0.1446 108,867 —
✅ 176_network_policy_blocking_traffic_no_runbooks 38.5s 6 13 $0.2601 128,594 —
✅ 24_misconfigured_pvc 34.0s 5 15 $0.1466 106,973 —
✅ 43_current_datetime_from_prompt 5.0s 1 — $0.0151 17,508 —
✅ 61_exact_match_counting 15.7s 4 4 $0.0704 77,164 —
Total 33.6s avg 5.1 avg 11.1 avg $1.4310 985,668 —

✅ Results of HolmesGPT evals

Automatically triggered by commit 613a3fe on branch claude/speed-up-regression-evals-TfJav

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 30.7s 5 11 $0.2378 107,253 105,471 1,782 80,870 24,601 — 584 —
✅ 101_loki_historical_logs_pod_deleted 43.8s 5 12 $0.2992 118,423 115,244 3,179 87,052 28,192 — 834 —
✅ 111_pod_names_contain_service 33.0s 5 12 $0.2434 106,289 104,231 2,058 79,691 24,540 — 534 —
✅ 112_find_pvcs_by_uuid 20.8s 4 5 $0.1972 80,178 78,950 1,228 56,774 22,176 — 455 —
✅ 12_job_crashing 36.1s 6 13 $0.2655 134,349 132,187 2,162 106,819 25,368 — 489 —
✅ 176_network_policy_blocking_traffic_no_runbooks 37.2s 5 16 $0.2761 115,695 113,170 2,525 85,850 27,320 — 699 —
✅ 24_misconfigured_pvc 35.4s 6 15 $0.2639 129,789 127,534 2,255 102,282 25,252 — 601 —
✅ 43_current_datetime_from_prompt 4.7s 1 — $0.1114 17,496 17,388 108 0 17,388 — 108 —
✅ 61_exact_match_counting 15.2s 4 4 $0.1677 77,244 76,699 545 56,571 20,128 — 247 —
Total 28.5s avg 4.6 avg 11.0 avg $2.0622 886,716 870,874 15,842 655,909 214,965 — 834 —

Benchmark comparison unavailable: No ci-benchmark experiments found

Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: No ci-benchmark experiments found

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/speed-up-regression-evals-TfJav -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals in automatic regression runs:

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID

Examples: evals-tag-easy, evals-id-09_crashpod

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/speed-up-regression-evals-TfJav -f markers=regression -f filter=

@github-actions

github-actions Bot commented Mar 6, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 38d334dd (built in 5m 41s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:38d334dd
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:38d334dd me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:38d334dd
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:38d334dd
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:38d334dd
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:38d334dd me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:38d334dd
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:38d334dd

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:38d334dd \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:38d334dd

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:38d334dd \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:38d334dd

@netlify

netlify Bot commented Mar 6, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 613a3fe
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69ad3256520ec100097f6f99
😎 Deploy Preview https://deploy-preview-1687--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai

coderabbitai Bot commented Mar 6, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Installs Poetry via a curl installer and caches an in-project .venv; moves KIND cluster setup into a background script which caches KIND artifacts and emits a sentinel; adds a polling/wait step for cluster readiness; increases pytest xdist parallelism from -n10 to -n20.

Changes

Cohort / File(s) Summary
Poetry setup & virtualenv caching
​.github/actions/setup-holmes-env/action.yml
Replace combined "Install Python dependencies and Poetry" with a curl-based Poetry installer (v1.4.0); switch Poetry to create an in-project .venv; add step to add .venv/bin to PATH; add cache step for .venv and Poetry caches keyed by python-version + poetry.lock.
Background KIND setup, caching & test parallelism
​.github/workflows/eval-regression.yaml
Run KIND cluster setup as a background script that saves/loads node & workload images and caches KIND binaries/Helm/manifests; write KIND_SETUP_COMPLETE sentinel on success and add a polling/wait step that validates sentinel and cluster state; increase pytest xdist workers from -n10 to -n20; add related logging and synchronization.

Sequence Diagram(s)

sequenceDiagram
    participant GH as "GitHub Actions Runner"
    participant SetupAction as "setup-holmes-env Action"
    participant Cache as "actions/cache"
    participant BGScript as "Background KIND setup script"
    participant WaitStep as "Wait for KIND setup"
    participant TestRunner as "pytest (-n20)"
    participant KIND as "KIND Cluster"

    GH->>SetupAction: run environment setup
    SetupAction->>SetupAction: install Poetry (curl), create .venv, add .venv/bin to PATH
    SetupAction->>Cache: restore/save caches (python/.venv, poetry caches)
    GH->>BGScript: start KIND setup (background)
    BGScript->>KIND: create cluster, load/save node & workload images, apply manifests
    BGScript-->>GH: emit KIND_SETUP_COMPLETE sentinel on success
    GH->>WaitStep: poll logs for KIND_SETUP_COMPLETE
    WaitStep->>KIND: validate cluster status (kubectl)
    GH->>TestRunner: run pytest with -n20
    TestRunner->>KIND: execute tests
    TestRunner-->>GH: upload results/logs
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • moshemorad
  • RoiGlinik
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: parallelizing KIND cluster setup by running it in the background while HolmesGPT environment setup occurs concurrently, which is the core optimization objective.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
.github/actions/setup-holmes-env/action.yml (1)

31-32: ⚠️ Potential issue | 🟠 Major

Update Poetry to a current stable version; 1.4.0 (early 2023) is severely outdated.

Poetry 1.4.0 is nearly 3 years old. Current stable is 2.3.2 (Feb 2026), with 1.8.5 also available (Dec 2024). The outdated version lacks critical bug fixes and security patches. No compatibility constraint is documented in pyproject.toml or elsewhere in the codebase. Update to the latest stable version unless there's a specific compatibility requirement.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/actions/setup-holmes-env/action.yml around lines 31 - 32, The
install step pins Poetry to an outdated version (curl ... | python3 - --version
1.4.0); update that install invocation to a current stable release (e.g.,
replace "--version 1.4.0" with "--version 2.3.2" or use a workflow
input/variable like POETRY_VERSION and set it to the latest stable) so the
action installs a maintained Poetry release; ensure the change is applied to the
install line that runs "curl -sSL https://install.python-poetry.org | python3 -
--version ..." in action.yml.
🧹 Nitpick comments (1)
.github/workflows/eval-regression.yaml (1)

687-688: Pin Helm CLI version for reproducible builds.

The Helm installation pulls from the main branch without version pinning, which could introduce unexpected breaking changes. Consider using a versioned URL similar to how KIND (v0.31.0) and Calico (v3.31.3) are pinned.

♻️ Suggested fix
           # Install Helm
-          curl https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash
+          curl https://raw.githubusercontent.com/helm/helm/v3.17.2/scripts/get-helm-3 | DESIRED_VERSION=v3.17.2 bash
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/eval-regression.yaml around lines 687 - 688, Replace the
unpinned Helm install curl line that fetches the script from the main branch by
pinning to a specific Helm release: set the desired version (e.g., v3.x.y) and
use the release-specific URL or the script's VERSION environment variable so the
pipeline always installs that tagged Helm release instead of latest; update the
line that currently runs "curl
https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash" to
reference the chosen tag (or export VERSION=<tag> before running the script) to
ensure reproducible builds.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/eval-regression.yaml:
- Line 1007: Replace the hardcoded "-n20" in PYTEST_ARGS with a dynamic worker
count derived from the previously-collected TEST_COUNT: compute a worker count
(for example min(TEST_COUNT, 20) or another cap you prefer) using bash
arithmetic (e.g. compute WORKER_COUNT=$(( TEST_COUNT < 20 ? TEST_COUNT : 20 ))
or similar) and then use "-n${WORKER_COUNT}" when building PYTEST_ARGS; update
the PYTEST_ARGS assignment to reference that computed WORKER_COUNT (keep
EVAL_MARKER_EXPR and existing test list unchanged).

---

Outside diff comments:
In @.github/actions/setup-holmes-env/action.yml:
- Around line 31-32: The install step pins Poetry to an outdated version (curl
... | python3 - --version 1.4.0); update that install invocation to a current
stable release (e.g., replace "--version 1.4.0" with "--version 2.3.2" or use a
workflow input/variable like POETRY_VERSION and set it to the latest stable) so
the action installs a maintained Poetry release; ensure the change is applied to
the install line that runs "curl -sSL https://install.python-poetry.org |
python3 - --version ..." in action.yml.

---

Nitpick comments:
In @.github/workflows/eval-regression.yaml:
- Around line 687-688: Replace the unpinned Helm install curl line that fetches
the script from the main branch by pinning to a specific Helm release: set the
desired version (e.g., v3.x.y) and use the release-specific URL or the script's
VERSION environment variable so the pipeline always installs that tagged Helm
release instead of latest; update the line that currently runs "curl
https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash" to
reference the chosen tag (or export VERSION=<tag> before running the script) to
ensure reproducible builds.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: a68177b9-a020-4ac8-925d-a97592414a72

📥 Commits

Reviewing files that changed from the base of the PR and between 777eb56 and 79f18cd.

📒 Files selected for processing (2)
  • .github/actions/setup-holmes-env/action.yml
  • .github/workflows/eval-regression.yaml

Comment thread .github/workflows/eval-regression.yaml

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (1)
.github/workflows/eval-regression.yaml (1)

1005-1007: ⚠️ Potential issue | 🟡 Minor

Still hardcoded at -n20.

steps.test-preview.outputs.test_count is collected earlier but never used here, so small/manual selections still fan out 20 workers against the same shared test infrastructure. That misses the dynamic scaling described in the PR and keeps the contention risk around the shared port-forward setup.

Suggested change
+          TEST_COUNT=${{ steps.test-preview.outputs.test_count }}
+          CPU_COUNT=$(nproc)
+          if (( TEST_COUNT == 0 )); then
+            WORKER_COUNT=1
+          else
+            WORKER_COUNT=$TEST_COUNT
+            (( WORKER_COUNT < 6 )) && WORKER_COUNT=6
+            (( WORKER_COUNT > 20 )) && WORKER_COUNT=20
+            (( WORKER_COUNT > CPU_COUNT )) && WORKER_COUNT=$CPU_COUNT
+          fi
-          PYTEST_ARGS=(--no-cov tests/llm/test_ask_holmes.py tests/llm/test_investigate.py -s -n20 -m "$EVAL_MARKER_EXPR")
+          PYTEST_ARGS=(--no-cov tests/llm/test_ask_holmes.py tests/llm/test_investigate.py -s -n"$WORKER_COUNT" -m "$EVAL_MARKER_EXPR")
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/eval-regression.yaml around lines 1005 - 1007, The
PYTEST_ARGS variable currently hardcodes "-n20", which ignores the previously
collected steps.test-preview.outputs.test_count; update the PYTEST_ARGS
construction (the PYTEST_ARGS assignment) to use the dynamic test count output
instead of "-n20" by injecting the test_count value from
steps.test-preview.outputs.test_count (or an env var fed from that output) so
parallel workers scale to the discovered test_count rather than always using 20.
🧹 Nitpick comments (1)
.github/workflows/eval-regression.yaml (1)

605-606: Enable pipefail in the generated KIND setup script.

set -e will not fail on the left side of curl … | bash, so a transient download error can surface later as a much less obvious Helm install failure instead of stopping at the real root cause.

Suggested change
-          set -e
+          set -euo pipefail

Also applies to: 688-688

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/eval-regression.yaml around lines 605 - 606, The generated
KIND setup script currently uses "#!/bin/bash" followed by "set -e" which
doesn't propagate failures through pipelines; update the script to enable
pipefail (e.g., replace or augment the "set -e" invocation with a form that
enables pipefail, such as adding "set -o pipefail" or using "set -euo pipefail")
so that a failed download in a pipeline like "curl … | bash" causes the script
to fail immediately; apply the same change to the other occurrence of the
script.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Duplicate comments:
In @.github/workflows/eval-regression.yaml:
- Around line 1005-1007: The PYTEST_ARGS variable currently hardcodes "-n20",
which ignores the previously collected steps.test-preview.outputs.test_count;
update the PYTEST_ARGS construction (the PYTEST_ARGS assignment) to use the
dynamic test count output instead of "-n20" by injecting the test_count value
from steps.test-preview.outputs.test_count (or an env var fed from that output)
so parallel workers scale to the discovered test_count rather than always using
20.

---

Nitpick comments:
In @.github/workflows/eval-regression.yaml:
- Around line 605-606: The generated KIND setup script currently uses
"#!/bin/bash" followed by "set -e" which doesn't propagate failures through
pipelines; update the script to enable pipefail (e.g., replace or augment the
"set -e" invocation with a form that enables pipefail, such as adding "set -o
pipefail" or using "set -euo pipefail") so that a failed download in a pipeline
like "curl … | bash" causes the script to fail immediately; apply the same
change to the other occurrence of the script.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: b1c1ebac-f828-4174-bcc0-086a983586f6

📥 Commits

Reviewing files that changed from the base of the PR and between 79f18cd and 569a437.

📒 Files selected for processing (1)
  • .github/workflows/eval-regression.yaml

@aantn
aantn enabled auto-merge (squash) March 6, 2026 11:59
Cache KIND binary, Helm binary, KIND node Docker image, Calico/metrics-server
manifests, Helm chart repository data, and all workload container images
(Calico, Prometheus, kube-state-metrics, node-exporter, metrics-server).

On cache hit:
- Binaries loaded from cache instead of downloading (~10s saved)
- KIND node image loaded via docker load (~15-20s saved)
- Workload images pre-loaded into KIND via kind load image-archive,
  so pods start without pulling from registries (~40-60s saved)
- Helm charts served from local cache (~10s saved)
- Manifests read from local files (~5s saved)

On cache miss (first run): artifacts are saved for subsequent runs.
Cache key includes all component versions to auto-bust on upgrades.

https://claude.ai/code/session_01ReC3SkRaWTNe3BZyuhLRyx
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
.github/workflows/eval-regression.yaml (1)

755-788: Parallel background jobs may orphan on partial failure.

If the Prometheus install (PROM_PID) fails, wait $PROM_PID triggers set -e and exits the script. However, the metrics-server subshell (METRICS_PID) continues running as an orphaned process, potentially leaving the cluster in an inconsistent state.

🛠️ Proposed fix: Add trap to clean up background jobs
+          # Track background PIDs for cleanup
+          BACKGROUND_PIDS=()
+          cleanup_background() {
+            for pid in "${BACKGROUND_PIDS[@]}"; do
+              kill "$pid" 2>/dev/null || true
+            done
+          }
+          trap cleanup_background EXIT
+
           (
             helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
             # ... rest of prometheus install
           ) &
           PROM_PID=$!
+          BACKGROUND_PIDS+=($PROM_PID)

           (
             kubectl apply -f "$METRICS_MANIFEST"
             # ... rest of metrics-server install
           ) &
           METRICS_PID=$!
+          BACKGROUND_PIDS+=($METRICS_PID)

           wait $PROM_PID
           wait $METRICS_PID
+          trap - EXIT  # Clear trap after successful completion
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/eval-regression.yaml around lines 755 - 788, The two
background subshells assigned to PROM_PID and METRICS_PID can be orphaned if one
fails; add a cleanup trap before starting them that kills/terminates PROM_PID
and METRICS_PID on EXIT or ERR and waits for them to finish, then after
launching the subshells use robust waiting and error checking (e.g., capture
each wait exit code rather than relying on set -e) so the script explicitly
kills the other job if one fails; reference PROM_PID, METRICS_PID, the
background subshells (the (...) & blocks), and the wait calls to locate where to
add the trap and the explicit kill/wait logic.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In @.github/workflows/eval-regression.yaml:
- Around line 755-788: The two background subshells assigned to PROM_PID and
METRICS_PID can be orphaned if one fails; add a cleanup trap before starting
them that kills/terminates PROM_PID and METRICS_PID on EXIT or ERR and waits for
them to finish, then after launching the subshells use robust waiting and error
checking (e.g., capture each wait exit code rather than relying on set -e) so
the script explicitly kills the other job if one fails; reference PROM_PID,
METRICS_PID, the background subshells (the (...) & blocks), and the wait calls
to locate where to add the trap and the explicit kill/wait logic.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 1e2a0a4b-d91a-423e-9540-8eabe7fe0287

📥 Commits

Reviewing files that changed from the base of the PR and between 569a437 and e09f672.

📒 Files selected for processing (1)
  • .github/workflows/eval-regression.yaml

The previous cache only stored ~/.cache/pip and ~/.cache/pypoetry (download
caches). With virtualenvs.create=false, poetry install still had to unpack
and install every package to system site-packages on every run.

Now uses in-project virtualenv (.venv/) and caches it directly. On cache
hit, poetry install is near-instant since all packages are already installed.
Also adds .venv/bin to PATH so scripts with #!/usr/bin/env python3 work.

https://claude.ai/code/session_01ReC3SkRaWTNe3BZyuhLRyx
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (2)
.github/actions/setup-holmes-env/action.yml (2)

40-44: Consider adding restore-keys for partial cache restoration.

When poetry.lock changes, the exact cache key won't match and there's no fallback. Adding restore-keys allows restoring a previous virtualenv, and poetry install will then only update the diff—often faster than a full install.

♻️ Proposed fix to add restore-keys
     - name: Cache Python virtualenv
       uses: actions/cache@v4
       with:
         path: ${{ inputs.working-directory }}/.venv
         key: venv-${{ inputs.python-version }}-${{ hashFiles(format('{0}/poetry.lock', inputs.working-directory)) }}
+        restore-keys: |
+          venv-${{ inputs.python-version }}-
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/actions/setup-holmes-env/action.yml around lines 40 - 44, The cache
step "Cache Python virtualenv" (uses: actions/cache@v4) lacks restore-keys, so
when the key built from venv-${{ inputs.python-version }}-${{
hashFiles(format('{0}/poetry.lock', inputs.working-directory)) }} misses there
is no fallback; add a restore-keys entry (e.g. a prefix fallback like venv-${{
inputs.python-version }}- or venv-) so the action can restore a partial cache
and speed up subsequent poetry installs, keeping path: ${{
inputs.working-directory }}/.venv and the existing key intact.

26-38: Consider upgrading Poetry to a more recent stable version for long-term maintenance.

Poetry 1.4.0 is available on PyPI but is significantly outdated—the latest stable version is 2.3.2 (as of March 2026). While 1.4.0 remains functional, evaluate whether upgrading to a newer stable release would provide better long-term compatibility and access to recent improvements and security patches.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/actions/setup-holmes-env/action.yml around lines 26 - 38, Update the
"Install Poetry" step to install a newer stable Poetry release instead of
pinning to 1.4.0: change the install command that pipes to python3 (the curl |
python3 - --version 1.4.0 invocation) to use a current stable version (e.g.,
2.3.2) or make the version configurable via an environment variable/input so
future upgrades are easier; ensure the PATH echo and poetry config
virtualenvs.in-project true lines remain unchanged and verify the workflow uses
the same binary path ($HOME/.local/bin/poetry) after the change.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In @.github/actions/setup-holmes-env/action.yml:
- Around line 40-44: The cache step "Cache Python virtualenv" (uses:
actions/cache@v4) lacks restore-keys, so when the key built from venv-${{
inputs.python-version }}-${{ hashFiles(format('{0}/poetry.lock',
inputs.working-directory)) }} misses there is no fallback; add a restore-keys
entry (e.g. a prefix fallback like venv-${{ inputs.python-version }}- or venv-)
so the action can restore a partial cache and speed up subsequent poetry
installs, keeping path: ${{ inputs.working-directory }}/.venv and the existing
key intact.
- Around line 26-38: Update the "Install Poetry" step to install a newer stable
Poetry release instead of pinning to 1.4.0: change the install command that
pipes to python3 (the curl | python3 - --version 1.4.0 invocation) to use a
current stable version (e.g., 2.3.2) or make the version configurable via an
environment variable/input so future upgrades are easier; ensure the PATH echo
and poetry config virtualenvs.in-project true lines remain unchanged and verify
the workflow uses the same binary path ($HOME/.local/bin/poetry) after the
change.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 775e0410-059d-4559-8f89-42ea84fa8386

📥 Commits

Reviewing files that changed from the base of the PR and between e09f672 and c26ec28.

📒 Files selected for processing (1)
  • .github/actions/setup-holmes-env/action.yml

claude and others added 3 commits March 6, 2026 12:39
ctr images export fails with "content digest not found" when containerd
has multi-arch manifest lists but only the linux/amd64 layers are present.
Fix by using --platform linux/amd64 flag, with fallback to per-image
export if bulk export still fails.

https://claude.ai/code/session_01ReC3SkRaWTNe3BZyuhLRyx
Signed-off-by: Claude <noreply@anthropic.com>
…export

ctr images export fails on every image in KIND because containerd stores
multi-arch manifest lists without all platform layers. Switch to pulling
images on the host Docker daemon and using docker save, which produces a
tar that kind load image-archive can consume on cache hit.

Also make the entire caching block non-fatal (|| true) so image cache
failures never break the build.

https://claude.ai/code/session_01ReC3SkRaWTNe3BZyuhLRyx
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/eval-regression.yaml:
- Line 1236: Remove or complete the orphaned trailing comment "# CI speed
optimizations" in the workflow file: either delete that standalone comment line
or replace it with the intended documentation/details about the CI optimization
steps so the meaning is clear; locate the comment by searching for the exact
text "# CI speed optimizations" and update or remove it accordingly.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 10547331-1e65-4086-ae6e-9935451c35d9

📥 Commits

Reviewing files that changed from the base of the PR and between c26ec28 and 8f7ea3f.

📒 Files selected for processing (1)
  • .github/workflows/eval-regression.yaml

Comment thread .github/workflows/eval-regression.yaml Outdated
@aantn
aantn merged commit ab74f40 into master Mar 8, 2026
18 of 19 checks passed
@aantn
aantn deleted the claude/speed-up-regression-evals-TfJav branch March 8, 2026 08:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants