Skip to content

[ROB-321] Restore regression tag on ES multi-cluster evals + wire ES secrets in… - #2179

Merged
Avi-Robusta merged 1 commit into
masterfrom
fix-es-multicluster-benchmark-regression
Jun 11, 2026
Merged

Avi-Robusta merged 1 commit into
masterfrom
fix-es-multicluster-benchmark-regression

Conversation

@Avi-Robusta

@Avi-Robusta Avi-Robusta commented Jun 11, 2026 •

Copy link
Copy Markdown
Collaborator

Fixing regression tests

  • Re-add regression tag to 254/259/260.
  • Add ELASTICSEARCH_URL/ELASTICSEARCH_API_KEY to eval-master.yaml, mirroring eval-regression.yaml and eval-benchmarks.yaml.

Summary by CodeRabbit

  • Chores

    • Updated regression evaluation workflow to include additional infrastructure credentials.
  • Tests

    • Added regression test tags to multiple test cases to strengthen test coverage.

…to master workflow

The multi-cluster Elasticsearch evals (254/259/260) had their `regression`
tag removed in cbaef1a because they aborted eval runs under
--strict-setup-mode when ELASTICSEARCH_URL/ELASTICSEARCH_API_KEY were unset.

The PR-triggered regression workflow (eval-regression.yaml) had the ES
secrets, so the tests passed there. The benchmark workflow got them in
cbaef1a. But eval-master.yaml ("Run eval regression on master") still
lacked them, so these tests failed setup and aborted the whole master run.

- Re-add `regression` tag to 254/259/260.
- Add ELASTICSEARCH_URL/ELASTICSEARCH_API_KEY to eval-master.yaml,
  mirroring eval-regression.yaml and eval-benchmarks.yaml.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: avi@robusta.dev <avi@robusta.dev>
@github-actions

github-actions Bot commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit b87b849 on branch fix-es-multicluster-benchmark-regression

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 14/14 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Denied commands Src
✅ 09_crashpod 21.2s 3 6 $0.1801 49,639 48,387 18,465 1,252 530 29,360 19,027 103 — — src
✅ 101_loki_historical_logs_pod_deleted 53.2s 5 11 $0.2721 91,950 88,994 21,789 2,956 923 66,394 22,600 421 — — src
✅ 112_find_pvcs_by_uuid 14.0s 2 2 $0.1432 31,704 30,919 16,597 785 435 14,319 16,600 156 — — src
✅ 12_job_crashing 36.1s 5 10 $0.2369 92,992 91,057 20,777 1,935 511 69,841 21,216 186 — — src
✅ 176_network_policy_blocking_traffic_no_skills 35.3s 6 13 $0.2561 112,370 110,392 21,634 1,978 499 87,918 22,474 133 — — src
✅ 227_count_configmaps_per_namespace[0] 16.2s 3 6 $0.1491 46,350 45,643 16,660 707 432 28,979 16,664 29 — — src
✅ 243_pod_names_contain_service 28.1s 3 6 $0.1816 48,976 47,459 17,757 1,517 578 29,229 18,230 266 — — src
✅ 24_misconfigured_pvc 33.4s 4 12 $0.2249 69,496 67,493 19,552 2,003 747 46,027 21,466 156 — — src
✅ 254_elasticsearch_dr_test_log_check 60.7s 8 12 $0.2780 105,557 101,550 17,692 4,007 1,533 83,849 17,701 574 — — src
✅ 259_wrong_cluster_logs_confusion 73.7s 9 11 $0.2925 126,840 122,693 17,540 4,147 1,431 105,143 17,550 819 — — src
✅ 260_global_es_remote_cluster_logs 51.8s 7 11 $0.2346 99,367 96,674 16,737 2,693 796 79,639 17,035 137 — — src
✅ 43_current_datetime_from_prompt 4.7s 1 — $0.1017 14,428 14,306 14,306 122 122 0 14,306 78 — — src
✅ 51_logs_summarize_errors 23.3s 3 2 $0.1616 47,461 46,572 17,639 889 509 28,929 17,643 33 — — src
✅ 61_exact_match_counting 9.2s 2 1 $0.1146 29,148 28,927 14,641 221 152 14,283 14,644 34 — — src
Total 32.9s avg 4.4 avg 7.9 avg $2.8270 966,278 941,066 21,789 25,212 1,533 683,910 257,156 3,125 — —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 17 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (1d ago) Δ vs master benchmark (23h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 21.2s 29.7s ↓29% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 53.2s 53.1s ±0% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 14.0s 15.5s ±0% — —
12_job_crashing (opus-4.6) 📄 36.1s 36.5s ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 35.3s 33.4s ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 16.2s 15.6s ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 28.1s 27.3s ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 33.4s 31.1s ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 60.7s — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 73.7s — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 51.8s — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 4.7s 3.9s ↑20% — —
51_logs_summarize_errors (opus-4.6) 📄 23.3s 22.2s ±0% — —
61_exact_match_counting (opus-4.6) 📄 9.2s 8.3s ↑11% — —
Total (all, n=14) 32.9s 25.1s — — —
Comparable (m=11, b=0) 25.0s 25.1s ±0% — —

Cost comparison:

Test case This branch master (1d ago) Δ vs master benchmark (23h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.1801 $0.2105 ↓14% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.2721 $0.2724 ±0% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.1432 $0.1555 ±0% — —
12_job_crashing (opus-4.6) 📄 $0.2369 $0.2355 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.2561 $0.2340 ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.1491 $0.1494 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 $0.1816 $0.1794 ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 $0.2249 $0.2290 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 $0.2780 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 $0.2925 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 $0.2346 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 $0.1017 $0.0112 ↑805% — —
51_logs_summarize_errors (opus-4.6) 📄 $0.1616 $0.1586 ±0% — —
61_exact_match_counting (opus-4.6) 📄 $0.1146 $0.1147 ±0% — —
Total (all, n=14) $0.2019 $0.1773 — — —
Comparable (m=11, b=0) $0.1838 $0.1773 ±0% — —

Total tokens comparison:

Test case This branch master (1d ago) Δ vs master benchmark (23h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 49,639 69,874 ↓29% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 91,950 91,813 ±0% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 31,704 33,166 ±0% — —
12_job_crashing (opus-4.6) 📄 92,992 91,375 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 112,370 107,045 ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 46,350 46,360 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 48,976 49,378 ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 69,496 71,330 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 105,557 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 126,840 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 99,367 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 14,428 14,428 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 47,461 47,068 ±0% — —
61_exact_match_counting (opus-4.6) 📄 29,148 29,157 ±0% — —
Total (all, n=14) 69,020 59,181 — — —
Comparable (m=11, b=0) 57,683 59,181 ±0% — —

Cached tokens comparison:

Test case This branch master (1d ago) Δ vs master benchmark (23h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 29,360 47,758 ↓39% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 66,394 66,223 ±0% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 14,319 14,319 ±0% — —
12_job_crashing (opus-4.6) 📄 69,841 68,085 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 87,918 84,955 ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 28,979 28,986 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 29,229 29,650 ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 46,027 47,757 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 83,849 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 105,143 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 79,639 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 — 14,303 — — —
51_logs_summarize_errors (opus-4.6) 📄 28,929 28,923 ±0% — —
61_exact_match_counting (opus-4.6) 📄 14,283 14,283 ±0% — —
Total (all, n=14) 48,851 40,477 — — —
Comparable (m=10, b=0) 41,528 43,094 ±0% — —

Turns comparison:

Test case This branch master (1d ago) Δ vs master benchmark (23h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 3 4 ↓25% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 5 5 ±0% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% — —
12_job_crashing (opus-4.6) 📄 5 5 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 6 6 ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 3 3 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 3 3 ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 4 4 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 8 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 9 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 7 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 3 3 ±0% — —
61_exact_match_counting (opus-4.6) 📄 2 2 ±0% — —
Total (all, n=14) 4.4 3.5 — — —
Comparable (m=11, b=0) 3.4 3.5 ±0% — —

Tool calls comparison:

Test case This branch master (1d ago) Δ vs master benchmark (23h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 6 8 ↓25% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 11 10 ↑10% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% — —
12_job_crashing (opus-4.6) 📄 10 11 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 13 12 ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 6 6 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 6 6 ±0% — —
24_misconfigured_pvc (opus-4.6) 📄 12 14 ↓14% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 12 — — — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 11 — — — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 11 — — — —
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 2 2 ±0% — —
61_exact_match_counting (opus-4.6) 📄 1 1 ±0% — —
Total (all, n=14) 7.4 7.2 — — —
Comparable (m=10, b=0) 6.9 7.2 ±0% — —

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref fix-es-multicluster-benchmark-regression -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, fable-5, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, gpt-5.5, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, opus-4.8, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref fix-es-multicluster-benchmark-regression -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 03d453a4-23e3-4f1d-8be6-e43f9a637c9f

📥 Commits

Reviewing files that changed from the base of the PR and between cbaef1a and b87b849.

📒 Files selected for processing (4)
  • .github/workflows/eval-master.yaml
  • tests/llm/fixtures/test_ask_holmes/254_elasticsearch_dr_test_log_check/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/259_wrong_cluster_logs_confusion/test_case.yaml
  • tests/llm/fixtures/test_ask_holmes/260_global_es_remote_cluster_logs/test_case.yaml

Walkthrough

The PR configures Elasticsearch credential environment variables in the master branch regression evaluation workflow and tags three specific Elasticsearch-dependent test fixtures as regression tests to enable automated regression testing with Elasticsearch.

Changes

Add Elasticsearch regression tests

Layer / File(s) Summary
CI regression evaluation with Elasticsearch credentials
.github/workflows/eval-master.yaml
The regression evaluation step now receives ELASTICSEARCH_URL and ELASTICSEARCH_API_KEY from secrets, enabling Elasticsearch connectivity during regression testing.
Mark test fixtures as regression tests
tests/llm/fixtures/test_ask_holmes/254_elasticsearch_dr_test_log_check/test_case.yaml, tests/llm/fixtures/test_ask_holmes/259_wrong_cluster_logs_confusion/test_case.yaml, tests/llm/fixtures/test_ask_holmes/260_global_es_remote_cluster_logs/test_case.yaml
Three Elasticsearch-dependent test cases are tagged with regression to be included in automated regression test runs.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~5 minutes

Possibly related PRs

Suggested reviewers

  • aantn
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main changes: restoring the regression tag on ES multi-cluster evals and wiring ES secrets in the workflow.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 79989c0de (built in 4m 24s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:79989c0de
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:79989c0de me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:79989c0de
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:79989c0de
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:79989c0de
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:79989c0de me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:79989c0de
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:79989c0de

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:79989c0de \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:79989c0de

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:79989c0de \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:79989c0de

@Avi-Robusta
Avi-Robusta enabled auto-merge (squash) June 11, 2026 11:13
@Avi-Robusta
Avi-Robusta merged commit 30a5e2b into master Jun 11, 2026
14 of 15 checks passed
@Avi-Robusta
Avi-Robusta deleted the fix-es-multicluster-benchmark-regression branch June 11, 2026 11:15
@netlify

netlify Bot commented Jun 11, 2026

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit b87b849
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a2a97e1c08be30008083578
😎 Deploy Preview https://deploy-preview-2179--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants