Skip to content

fix(multi_instance): treat instances: [] as zero instances, not flat default - #2174

Open
aantn wants to merge 1 commit into
masterfrom
claude/multi-instance-empty-list
Open

aantn wants to merge 1 commit into
masterfrom
claude/multi-instance-empty-list

Conversation

@aantn

@aantn aantn commented Jun 10, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

_parse_instances in holmes/plugins/toolsets/multi_instance.py used if not raw: to detect the flat (single-instance) config shape. Because an empty list is falsy, this treats an explicit instances: [] identically to a missing instances key — synthesizing a default child from the top-level globals.

That bypasses the "No instances configured" path in _aggregate() and can leave a routable endpoint silently enabled when the config explicitly declared zero instances.

Fix

Distinguish absence from emptiness:

  • raw is None (key missing) → flat default instance (unchanged, backwards compatible)
  • raw == [] (explicit empty list) → zero instances, flows to _aggregate()'s "No instances configured" result
 raw = config.get("instances")
-if not raw:
+if raw is None:
     flat = {k: v for k, v in config.items() if k != "instances"}
     return [("default", flat)]
 if not isinstance(raw, list):
     raise ValueError("`instances` must be a list")
+if not raw:
+    return []

Tests

Added two cases to TestHealthAggregation:

  • test_explicit_empty_instances_yields_zero_instances — instances: [] (with top-level globals present) → prereqs fail with "No instances configured", no children, no tools.
  • test_missing_instances_key_still_flat_default — absent key keeps the flat default shape.

All 24 tests in tests/plugins/toolsets/test_multi_instance.py pass.

Context

Found by CodeRabbit while reviewing #2148 (it surfaced in that PR's diff only because of a master merge). The line is unrelated to #2148's approval refactor and originates in #2115, so it's split out here as a focused fix.

https://claude.ai/code/session_01FoqM3sjnrRgPjdqYdqutzK


Generated by Claude Code

Summary by CodeRabbit

  • Bug Fixes

    • Fixed configuration parsing to properly distinguish between explicitly empty instances (instances: []) and missing instances key. Empty instances now correctly indicate "no instances configured," while maintaining backward compatibility.
  • Tests

    • Added test cases for multi-instance configuration edge cases.

…t default

`_parse_instances` used `if not raw:`, which treats an explicit empty
`instances: []` the same as a missing `instances` key and synthesizes a
`default` child from the top-level globals. That bypasses the "No instances
configured" path in `_aggregate()` and can leave a routable endpoint enabled
when the config explicitly declared zero instances.

Distinguish absence (`raw is None` -> flat `default`) from an explicit empty
list (`raw == []` -> zero instances). Add tests for both shapes.

Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Jun 10, 2026 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit 6571983 on branch claude/multi-instance-empty-list

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Denied commands Src
✅ 09_crashpod 21.7s 3 6 $0.1802 49,572 48,322 18,417 1,250 530 29,353 18,969 111 — — src
✅ 101_loki_historical_logs_pod_deleted 47.2s 4 9 $0.2493 71,427 68,561 20,679 2,866 1,163 47,159 21,402 448 — — src
✅ 112_find_pvcs_by_uuid 14.6s 2 2 $0.1472 32,266 31,455 17,133 811 542 14,319 17,136 91 — — src
✅ 12_job_crashing 34.6s 5 10 $0.2380 92,865 90,925 20,724 1,940 512 69,490 21,435 200 — — src
✅ 176_network_policy_blocking_traffic_no_skills 34.2s 4 11 $0.2260 72,106 70,026 20,688 2,080 598 49,333 20,693 477 — — src
✅ 227_count_configmaps_per_namespace[0] 14.8s 3 6 $0.1490 46,341 45,634 16,651 707 432 28,979 16,655 29 — — src
✅ 243_pod_names_contain_service 30.7s 3 7 $0.1935 50,244 48,478 18,265 1,766 659 29,441 19,037 263 — — src
✅ 24_misconfigured_pvc 31.1s 5 9 $0.2239 91,708 90,216 20,143 1,492 351 68,898 21,318 108 — — src
✅ 43_current_datetime_from_prompt 4.5s 1 — $0.1017 14,428 14,306 14,306 122 122 0 14,306 78 — — src
✅ 51_logs_summarize_errors 19.9s 3 2 $0.1541 46,619 45,815 16,880 804 437 28,931 16,884 33 — — src
✅ 61_exact_match_counting 8.5s 2 1 $0.1146 29,149 28,928 14,642 221 152 14,283 14,645 34 — — src
Total 23.8s avg 3.2 avg 6.3 avg $1.9775 596,725 582,666 20,724 14,059 1,163 380,186 202,480 1,872 — —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 17 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (11h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 21.7s 29.7s ↓27% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 47.2s 53.1s ↓11% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 14.6s 15.5s ±0% — —
12_job_crashing (opus-4.6) 📄 34.6s 36.5s ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 34.2s 33.4s ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 14.8s 15.6s ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 30.7s 27.3s ↑12% — —
24_misconfigured_pvc (opus-4.6) 📄 31.1s 31.1s ±0% — —
43_current_datetime_from_prompt (opus-4.6) 📄 4.5s 3.9s ↑15% — —
51_logs_summarize_errors (opus-4.6) 📄 19.9s 22.2s ↓10% — —
61_exact_match_counting (opus-4.6) 📄 8.5s 8.3s ±0% — —
Total (all, n=11) 23.8s 25.1s — — —
Comparable (m=11, b=0) 23.8s 25.1s ±0% — —

Cost comparison:

Test case This branch master (11h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.1802 $0.2105 ↓14% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.2493 $0.2724 ±0% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.1472 $0.1555 ±0% — —
12_job_crashing (opus-4.6) 📄 $0.2380 $0.2355 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.2260 $0.2340 ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.1490 $0.1494 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 $0.1935 $0.1794 ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 $0.2239 $0.2290 ±0% — —
43_current_datetime_from_prompt (opus-4.6) 📄 $0.1017 $0.0112 ↑805% — —
51_logs_summarize_errors (opus-4.6) 📄 $0.1541 $0.1586 ±0% — —
61_exact_match_counting (opus-4.6) 📄 $0.1146 $0.1147 ±0% — —
Total (all, n=11) $0.1798 $0.1773 — — —
Comparable (m=11, b=0) $0.1798 $0.1773 ±0% — —

Total tokens comparison:

Test case This branch master (11h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 49,572 69,874 ↓29% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 71,427 91,813 ↓22% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 32,266 33,166 ±0% — —
12_job_crashing (opus-4.6) 📄 92,865 91,375 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 72,106 107,045 ↓33% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 46,341 46,360 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 50,244 49,378 ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 91,708 71,330 ↑29% — —
43_current_datetime_from_prompt (opus-4.6) 📄 14,428 14,428 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 46,619 47,068 ±0% — —
61_exact_match_counting (opus-4.6) 📄 29,149 29,157 ±0% — —
Total (all, n=11) 54,248 59,181 — — —
Comparable (m=11, b=0) 54,248 59,181 ±0% — —

Cached tokens comparison:

Test case This branch master (11h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 29,353 47,758 ↓39% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 47,159 66,223 ↓29% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 14,319 14,319 ±0% — —
12_job_crashing (opus-4.6) 📄 69,490 68,085 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 49,333 84,955 ↓42% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 28,979 28,986 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 29,441 29,650 ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 68,898 47,757 ↑44% — —
43_current_datetime_from_prompt (opus-4.6) 📄 — 14,303 — — —
51_logs_summarize_errors (opus-4.6) 📄 28,931 28,923 ±0% — —
61_exact_match_counting (opus-4.6) 📄 14,283 14,283 ±0% — —
Total (all, n=11) 34,562 40,477 — — —
Comparable (m=10, b=0) 38,019 43,094 ↓12% — —

Turns comparison:

Test case This branch master (11h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 3 4 ↓25% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 4 5 ↓20% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% — —
12_job_crashing (opus-4.6) 📄 5 5 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 4 6 ↓33% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 3 3 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 3 3 ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 5 4 ↑25% — —
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 3 3 ±0% — —
61_exact_match_counting (opus-4.6) 📄 2 2 ±0% — —
Total (all, n=11) 3.2 3.5 — — —
Comparable (m=11, b=0) 3.2 3.5 ±0% — —

Tool calls comparison:

Test case This branch master (11h ago) Δ vs master benchmark (7h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 6 8 ↓25% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 9 10 ↓10% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% — —
12_job_crashing (opus-4.6) 📄 10 11 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 11 12 ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 6 6 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 7 6 ↑17% — —
24_misconfigured_pvc (opus-4.6) 📄 9 14 ↓36% — —
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 2 2 ±0% — —
61_exact_match_counting (opus-4.6) 📄 1 1 ±0% — —
Total (all, n=11) 5.7 7.2 — — —
Comparable (m=10, b=0) 6.3 7.2 ↓13% — —

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/multi-instance-empty-list -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, fable-5, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, gpt-5.5, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, opus-4.8, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/multi-instance-empty-list -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented Jun 10, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: de6a0ff3-0ddd-443d-8d28-e5e7ad959b90

📥 Commits

Reviewing files that changed from the base of the PR and between cbaef1a and 6571983.

📒 Files selected for processing (2)
  • holmes/plugins/toolsets/multi_instance.py
  • tests/plugins/toolsets/test_multi_instance.py

Walkthrough

The PR refines multi-instance configuration parsing in the toolset to handle an edge case: distinguishing an explicitly empty instances: [] configuration (treated as zero instances) from a missing instances key (which retains backwards-compatible single-instance behavior). Implementation changes and corresponding test cases validate both scenarios.

Changes

Multi-Instance Configuration Parsing

Layer / File(s) Summary
Instance configuration parsing refinement
holmes/plugins/toolsets/multi_instance.py
_parse_instances now returns [] when instances: [] is explicitly provided, instead of falling back to single-instance defaults, while preserving backwards compatibility for missing instances keys.
Empty and missing instances configuration tests
tests/plugins/toolsets/test_multi_instance.py
Two new test methods validate that explicit instances: [] fails prerequisites and produces no children, while omitting instances creates a single default child with passing prerequisites.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Suggested reviewers

  • naomi-robusta
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: fixing the handling of an explicit empty instances list to distinguish it from a missing key, which is exactly what the changeset implements.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jun 10, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 78b3e9d9d (built in 4m 37s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:78b3e9d9d
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:78b3e9d9d me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:78b3e9d9d
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:78b3e9d9d
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:78b3e9d9d
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:78b3e9d9d me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:78b3e9d9d
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:78b3e9d9d

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:78b3e9d9d \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:78b3e9d9d

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:78b3e9d9d \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:78b3e9d9d

@netlify

netlify Bot commented Jun 10, 2026

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 6571983
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a29c1888fac66000811e2b3
😎 Deploy Preview https://deploy-preview-2174--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants