Skip to content

Add ENV_CONFIGS support for evaluation tests & fix kubectl tabular query filtering - #1471

Merged
Sheeproid merged 32 commits into
masterfrom
env-configs-feature
Feb 3, 2026
Merged

Sheeproid merged 32 commits into
masterfrom
env-configs-feature

Conversation

@Sheeproid

@Sheeproid Sheeproid commented Feb 3, 2026 •

Copy link
Copy Markdown
Collaborator

Title:
Add ENV_CONFIGS support for evaluation tests & fix kubectl tabular query filtering

Description:

Summary

  • Add support for running evaluation tests with different environment variable configurations
  • Fix kubectl tabular query tool to preserve header when filtering

ENV_CONFIGS Feature

Enables comparing test runs across different environment configurations (similar to multi-model support).

Format:

ENV_CONFIGS='config1:VAR1=val1;VAR2=val2|config2:VAR1=val3'                                                                                                                                                                                                                                                      
                                                                                                                                                                                                                                                                                                                 
Example:                                                                                                                                                                                                                                                                                                         
ENV_CONFIGS='baseline:|minimal:ENABLED_PROMPTS=none' MODEL=opus-4.5 poetry run pytest -m llm -k "test_name"                                                                                                                                                                                                      
                                                                                                                                                                                                                                                                                                                 
- Creates Cartesian product: 2 models × 2 configs = 4 test runs                                                                                                                                                                                                                                                  
- Adds ENV CONFIG COMPARISON table to test output                                                                                                                                                                                                                                                                
- Tracks env_config in test properties for reporting                                                                                                                                                                                                                                                             
                                                                                                                                                                                                                                                                                                                 
kubectl Tabular Query Fix                                                                                                                                                                                                                                                                                        
                                                                                                                                                                                                                                                                                                                 
Fixed kubernetes_tabular_query tool to preserve header line when using filter_pattern:                                                                                                                                                                                                                           
- Previous: Used Jinja2 split filter which doesn't exist                                                                                                                                                                                                                                                         
- New: Uses (read -r header; echo "$header"; grep -E 'pattern') to preserve header                                                                                                                                                                                                                               
                                                                                                                                                                                                                                                                                                                 
Test plan                                                                                                                                                                                                                                                                                                        
                                                                                                                                                                                                                                                                                                                 
- Verified ENV_CONFIGS with multiple configurations                                                                                                                                                                                                                                                              
- Verified kubectl tabular query preserves headers when filtering                                                                                                                                                                                                                                                
                                                                                                                                                                                                                                                                                 
                                              

<!-- This is an auto-generated comment: release notes by coderabbit.ai -->

## Summary by CodeRabbit

* **New Features**
* Added support for parameterized test execution with different environment configurations.
* Enhanced test reporting with environment configuration comparisons and detailed statistics.

* **Bug Fixes**
* Improved Kubernetes query filtering to better preserve data accuracy.
* Marked several tests as skipped due to inconsistencies with mocked data or date-dependent failures.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
 Post-processing support was removed in PR #1279 (bec4d12)

Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
Signed-off-by: Tomer Keshet <tomer@robusta.dev>
@linux-foundation-easycla

Copy link
Copy Markdown

CLA Not Signed

@github-actions

github-actions Bot commented Feb 3, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for 09b55fc (built in 3m 52s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:09b55fc
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:09b55fc me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:09b55fc
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:09b55fc

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:09b55fc

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:09b55fc

@github-actions

github-actions Bot commented Feb 3, 2026 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit e6d63fa on branch env-configs-feature

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 33.9s 5 11 $0.2293
✅ 101_loki_historical_logs_pod_deleted 30.9s 4 8 $0.2062
✅ 111_pod_names_contain_service 43.2s 5 12 $0.2383
✅ 112_find_pvcs_by_uuid 37.3s 6 8 $0.2557
✅ 12_job_crashing 34.0s 5 10 $0.2241
✅ 176_network_policy_blocking_traffic_no_runbooks 52.0s 7 18 $0.2928
✅ 24_misconfigured_pvc 36.7s 6 13 $0.2348
✅ 43_current_datetime_from_prompt 5.1s 1 — $0.1050
✅ 61_exact_match_counting 13.5s 3 2 $0.1416
Total 31.8s avg 4.7 avg 10.2 avg $1.9279
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref env-configs-feature -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref env-configs-feature -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented Feb 3, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

The pull request introduces environment configuration support to the test suite, allowing tests to run against multiple environment configurations. It refactors test utilities to support parameterized env_config, updates reporting to track and compare results across configurations, and marks several test fixtures as skipped or updates their metadata. Additionally, Kubernetes toolset filtering is improved to preserve table headers.

Changes

Cohort / File(s) Summary
Kubernetes Toolset Filtering
holmes/plugins/toolsets/kubernetes.yaml
Improved kubectl tabular output filtering to preserve header lines while applying extended regex filtering instead of simple grep-based filtering.
Test Infrastructure - Environment Config
tests/llm/utils/env_config.py
New module introducing EnvConfig dataclass, parse_env_configs() parser for environment configuration strings, and get_env_configs() helper to retrieve configs from ENV_CONFIGS environment variable.
Test Infrastructure - Utilities
tests/llm/utils/commands.py
Refactored environment variable management with new _temporary_env_vars helper for consistent save/apply/restore logic, and introduced apply_env_config() context manager to apply environment variables from EnvConfig objects.
Test Infrastructure - Properties & Reporting
tests/llm/utils/property_manager.py
Extended set_initial_properties() to accept optional env_config parameter and record it in test properties.
Test Execution
tests/llm/conftest.py
Added env_config field to test result records, populated from user_props with "default" fallback for skipped and executed tests.
Test Files - Parametrization
tests/llm/test_ask_holmes.py, tests/llm/test_investigate.py
Parameterized tests with env_config using pytest.mark.parametrize, added helper _get_env_config_ids() for test ID generation, integrated env_config into trace names and metadata, and updated context management to apply environment configurations via ExitStack.
Test Reporting & Statistics
tests/llm/utils/reporting/terminal_reporter.py
Significantly expanded TestStatistics to support three-level nesting by test_case → model → env_config; added env_config-aware query methods (get_stats_with_env_config, get_env_config_summary, get_env_config_costs, get_env_config_times); added env_config comparison table rendering; exposed env_configs property and has_multiple_env_configs flag.
Test Fixtures - Skipped Tests
tests/llm/fixtures/test_ask_holmes/{44_slack_statefulset_logs,48_logs_since_thursday,50a_logs_since_last_specific_month,93_events_since_specific_date}/test_case.yaml
Marked test cases as skipped with skip_reason explaining test failures due to toolset inconsistency, broken date logic, or missing mocked data.
Test Fixtures - Tag Updates
tests/llm/fixtures/test_ask_holmes/{100a_loki_historical_logs,108_logs_nearby_lines,112_find_pvcs_by_uuid,162_get_runbooks,42_dns_issues_*.../91f_datadog_logs_historical_pod}/test_case.yaml
Added, removed, or commented tags (benchmark, regression, easy, medium, hard); updated readiness check in 112_find_pvcs_by_uuid to actively poll for 4 PVCs instead of static sleep; added explanatory comments.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~35 minutes

Possibly related PRs

Suggested reviewers

  • aantn
  • arikalon1
  • moshemorad
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately captures the two main changes in the changeset: adding ENV_CONFIGS support for evaluation tests and fixing kubectl tabular query filtering.
Docstring Coverage ✅ Passed Docstring coverage is 84.85% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Feb 3, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit e6d63fa
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69821394d2b52700077b8631
😎 Deploy Preview https://deploy-preview-1471--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@Sheeproid
Sheeproid enabled auto-merge (squash) February 3, 2026 15:27
@github-actions

github-actions Bot commented Feb 3, 2026

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 10.20s 10.10s +1.0%
Warm Mean 4.42s 4.78s -7.6%
Warm Min 4.40s 4.71s
Warm Max 4.44s 4.91s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 30.34s 32.95s -7.9%
Warm Mean 7.42s 7.05s +5.2%
Warm Min 7.18s 6.68s
Warm Max 7.80s 7.74s

PR: 09b55fc9 | Master: 73febf3d | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
tests/llm/fixtures/test_ask_holmes/50a_logs_since_last_specific_month/test_case.yaml (1)

16-22: ⚠️ Potential issue | 🟡 Minor

Fix grammar in skip_reason.

Minor wording tweak for clarity.

✏️ Proposed fix
-skip_reason: "this test fails now as the toolset are not consistent with the mocked data"
+skip_reason: "this test fails now as the toolset is not consistent with the mocked data"
tests/llm/conftest.py (1)

845-862: ⚠️ Potential issue | 🟡 Minor

Capture env_config via pytest hooks instead of hardcoding "default"
Skipped tests always report "default", which hides real parameters. In pytest_runtest_makereport append item.callspec.params to report.user_properties, then in the skipped-test branch extract env_config (falling back to "default").

Example
# in pytest_runtest_makereport
report.user_properties.append(("callspec_params", getattr(item.callspec, "params", {})))

# in pytest_runtest_logreport for skipped tests
params = dict(report.user_properties).get("callspec_params", {})
env_val = params.get("env_config", "default")
env_config = getattr(env_val, "name", str(env_val))
"env_config": env_config
🤖 Fix all issues with AI agents
In `@holmes/plugins/toolsets/kubernetes.yaml`:
- Line 239: The kubectl command line in the command field that uses {{
filter_pattern }} currently lets grep's exit code 1 (no matches) fail the whole
tool; update the subshell that runs read/echo/grep so that grep returning no
matches is treated as success by making grep non-fatal (for example, append an
OR that forces a zero exit status on grep failures) — change the command that
includes {{ kind }}, {{ columns }} and {{ filter_pattern }} to wrap the grep
accordingly so only real errors (not "no matches") cause the tool to fail.

In
`@tests/llm/fixtures/test_ask_holmes/91f_datadog_logs_historical_pod/test_case.yaml`:
- Around line 20-21: The YAML comment references a non-existent pytest marker
"benchmark"; either remove the commented lines in test_case.yaml that reference
"benchmark" or add a proper marker declaration for "benchmark" in pyproject.toml
under the pytest markers (e.g., add "benchmark: description" to the markers list
in [tool.pytest.ini_options]) so the marker is valid for future use.

In
`@tests/llm/fixtures/test_ask_holmes/93_events_since_specific_date/test_case.yaml`:
- Around line 10-13: Fix the typo in the YAML test metadata: update the
skip_reason value for the test (the YAML key skip_reason) to correct the
duplicated word "to to" -> "to" so the string reads "this sometimes fails due to
missing mock errors"; ensure only the value for skip_reason is changed and
formatting remains valid YAML.

In `@tests/llm/utils/commands.py`:
- Around line 8-15: Add the missing typing import and return annotations for the
context managers: import Iterator from typing (e.g. change the top import to
"from typing import TYPE_CHECKING, Dict, Optional, Iterator") and update each
context manager function in this file to declare a return type of "->
Iterator[None]" (ensure functions that use yield are the ones annotated). Keep
the rest of the imports (is_run_live_enabled, HolmesTestCase) unchanged and only
add the Iterator import and the "-> Iterator[None]" annotations to the context
manager definitions.
🧹 Nitpick comments (4)
tests/llm/fixtures/test_ask_holmes/44_slack_statefulset_logs/test_case.yaml (1)

14-15: Consider linking the skip to a tracking issue or expiry.
Skipping is fine short-term, but a reference helps ensure the eval is revisited.

tests/llm/utils/reporting/terminal_reporter.py (1)

191-242: Hoist env_config uniqueness check out of the row loop.

The env_config set is recomputed for every row; precomputing once reduces repeated work and keeps the logic centralized.

🔧 Suggested refactor
-    # Add rows to table
-    for result in sorted_results:
+    # Add rows to table
+    unique_env_configs = {r.get("env_config", "default") for r in sorted_results}
+    show_env_config = len(unique_env_configs) > 1
+    for result in sorted_results:
         status = TestStatus(result)
@@
-        unique_env_configs = {r.get("env_config", "default") for r in sorted_results}
-        if len(unique_env_configs) > 1:
+        if show_env_config:
             parts.append(env_config)
tests/llm/test_investigate.py (1)

85-87: Add a return type and clarify the docstring intent.

This keeps type-hint coverage consistent and documents why the IDs exist (stable pytest names).

✍️ Suggested update
-def _get_env_config_ids():
-    """Generate ids for env_config parameterization."""
+def _get_env_config_ids() -> list[str]:
+    """Keep pytest parametrized IDs stable by using env_config names."""
     return [ec.name for ec in get_env_configs()]

As per coding guidelines, "Use type hints throughout Python code and run 'mypy' for type checking" and "Write clear, concise comments that explain 'why' rather than 'what'."

tests/llm/test_ask_holmes.py (1)

53-55: Add a return type and clarify the docstring intent.

This keeps typing consistent and explains why the IDs are derived from env_config names.

✍️ Suggested update
-def _get_env_config_ids():
-    """Generate ids for env_config parameterization."""
+def _get_env_config_ids() -> list[str]:
+    """Keep pytest parametrized IDs stable by using env_config names."""
     return [ec.name for ec in get_env_configs()]

As per coding guidelines, "Use type hints throughout Python code and run 'mypy' for type checking" and "Write clear, concise comments that explain 'why' rather than 'what'."

Comment thread holmes/plugins/toolsets/kubernetes.yaml
Comment thread tests/llm/utils/commands.py

@arikalon1 arikalon1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice work

@Sheeproid
Sheeproid merged commit 8941b73 into master Feb 3, 2026
18 of 20 checks passed
@Sheeproid
Sheeproid deleted the env-configs-feature branch February 3, 2026 19:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants