Skip to content

Remove duplicate test fixtures and consolidate context_window tags - #1585

Merged
aantn merged 3 commits into
masterfrom
claude/review-context-window-evals-hSITx
Feb 19, 2026
Merged

aantn merged 3 commits into
masterfrom
claude/review-context-window-evals-hSITx

Conversation

@aantn

@aantn aantn commented Feb 19, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR removes duplicate test fixtures for ask_holmes tests and consolidates context_window-related test tags across the test suite.

Key Changes

  • Removed duplicate test fixtures:

    • Deleted 47_truncated_logs_context_window/ directory containing test case, manifest, toolsets configuration, and fixture files for a logs context window truncation test
    • Deleted 160c_cpu_per_namespace_graph_with_global_truncation/ directory containing test case and toolsets configuration for a Prometheus metrics test with global truncation
  • Added context_window tag to existing tests:

    • Added context_window tag to 103_logs_transparency_default_limit test case
    • Added context_window tag to 160a_cpu_per_namespace_graph test case
    • Added context_window tag to 160b_cpu_per_namespace_graph_with_prom_truncation test case

Rationale

The changes consolidate test coverage by removing redundant test fixtures while ensuring existing tests that validate context window behavior are properly tagged for organization and filtering purposes.

https://claude.ai/code/session_01TceJLidFkDBiUpG1qmroPZ

Summary by CodeRabbit

  • Tests
    • Added context_window tag to test cases for improved categorization
    • Removed test fixtures for truncated logs and global truncation scenarios
    • Updated test configurations across multiple test cases

Delete two non-functional context window evaluation tests:
- 47_truncated_logs_context_window: Eval was broken and marked as skip=true
- 160c_cpu_per_namespace_graph_with_global_truncation: Env var override doesn't work

Both tests had skip: true in their test_case.yaml files and were not
providing value to the test suite.
Add missing context_window tag to three tests that validate context
window/truncation behavior:
- 103_logs_transparency_default_limit: Tests log truncation transparency
- 160a_cpu_per_namespace_graph: Tests handling of large Prometheus result sets
- 160b_cpu_per_namespace_graph_with_prom_truncation: Tests query_response_size_limit

All context_window related evals now have consistent tagging.
@netlify

netlify Bot commented Feb 19, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit b1f1596
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6996966a37791b00088fd6e7
😎 Deploy Preview https://deploy-preview-1585--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Feb 19, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ ef4501b (#22168967572)

✅ Results of HolmesGPT evals

Automatically triggered by commit ef4501b on branch claude/review-context-window-evals-hSITx

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 31.1s 5 11 $0.2386
✅ 101_loki_historical_logs_pod_deleted 39.6s 6 9 $0.2476
✅ 111_pod_names_contain_service 33.2s 5 12 $0.2394
✅ 112_find_pvcs_by_uuid 37.4s 6 8 $0.2646
✅ 12_job_crashing 36.8s 6 13 $0.2650
✅ 176_network_policy_blocking_traffic_no_runbooks 45.6s 6 18 $0.3002
✅ 24_misconfigured_pvc 36.3s 6 13 $0.2467
✅ 43_current_datetime_from_prompt 5.2s 1 — $0.1112
✅ 61_exact_match_counting 16.3s 4 4 $0.1676
Total 31.3s avg 5.0 avg 11.0 avg $2.0809

✅ Results of HolmesGPT evals

Automatically triggered by commit b1f1596 on branch claude/review-context-window-evals-hSITx

View workflow logs

⚠️ No eval report was generated.

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/review-context-window-evals-hSITx -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/review-context-window-evals-hSITx -f markers=regression -f filter=

@github-actions

github-actions Bot commented Feb 19, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for 21b087f (built in 5m 6s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:21b087f
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:21b087f me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:21b087f
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:21b087f

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:21b087f

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:21b087f

@coderabbitai

coderabbitai Bot commented Feb 19, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

The PR adds a "context_window" tag to three existing LLM test fixtures and removes two complete test case directories along with their associated configurations, Kubernetes manifests, and fixture files.

Changes

Cohort / File(s) Summary
Add context_window tag
tests/llm/fixtures/test_ask_holmes/103_logs_transparency_default_limit/test_case.yaml, tests/llm/fixtures/test_ask_holmes/160a_cpu_per_namespace_graph/test_case.yaml, tests/llm/fixtures/test_ask_holmes/160b_cpu_per_namespace_graph_with_prom_truncation/test_case.yaml
Added context_window tag to tags list in test_case.yaml files.
Remove test case 160c
tests/llm/fixtures/test_ask_holmes/160c_cpu_per_namespace_graph_with_global_truncation/test_case.yaml, tests/llm/fixtures/test_ask_holmes/160c_cpu_per_namespace_graph_with_global_truncation/toolsets.yaml
Deleted complete test case 160c configuration and prometheus metrics toolset.
Remove test case 47
tests/llm/fixtures/test_ask_holmes/47_truncated_logs_context_window/test_case.yaml, tests/llm/fixtures/test_ask_holmes/47_truncated_logs_context_window/toolsets.yaml, tests/llm/fixtures/test_ask_holmes/47_truncated_logs_context_window/manifest.yaml, tests/llm/fixtures/test_ask_holmes/47_truncated_logs_context_window/kubectl_...
Deleted complete test case 47 including test configuration, kubernetes toolsets, deployment manifest, and kubectl fixture files.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~15 minutes

Possibly related PRs

  • Big yaml eval #1527: Both PRs modify LLM test fixtures' metadata by adding/using the "context_window" tag alongside related test_case and toolsets changes.
  • add tags to failing tests #597: Both PRs modify test_case.yaml tags metadata, affecting how test cases are categorized and filtered.

Suggested reviewers

  • Sheeproid
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main changes: removing duplicate test fixtures (47_truncated_logs_context_window and 160c_cpu_per_namespace_graph_with_global_truncation) and consolidating context_window tags across related tests.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@aantn
aantn enabled auto-merge (squash) February 19, 2026 04:49
@github-actions

github-actions Bot commented Feb 19, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 10.65s 10.19s +4.5%
Warm Mean 4.76s 4.76s +0.1%
Warm Min 4.74s 4.69s
Warm Max 4.80s 4.80s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 27.76s 32.23s -13.9%
Warm Mean 7.65s 7.67s -0.2%
Warm Min 7.57s 7.38s
Warm Max 7.71s 8.04s

PR: 21b087f8 | Master: 218194c8 | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
tests/llm/fixtures/test_ask_holmes/160b_cpu_per_namespace_graph_with_prom_truncation/test_case.yaml (1)

5-6: ⚠️ Potential issue | 🟡 Minor

Adding a tag to a permanently-skipped test may create misleading filter results.

This test is marked skip: true with a skip_reason stating it is "no longer relevant." Adding context_window to it means that any tag-filtered run using context_window will surface this fixture but still skip it silently, potentially obscuring coverage gaps. If the test is truly obsolete, consider removing it entirely (as was done for 47_truncated_logs_context_window and 160c_cpu_per_namespace_graph_with_global_truncation in this same PR); if it still has value, the skip should be lifted.

Also applies to: 18-18

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In
`@tests/llm/fixtures/test_ask_holmes/160b_cpu_per_namespace_graph_with_prom_truncation/test_case.yaml`
around lines 5 - 6, The test fixture
"160b_cpu_per_namespace_graph_with_prom_truncation" is marked skip: true but was
updated to include the context_window tag, which can mislead tag-filtered test
runs; either remove this obsolete fixture entirely (as done for similar tests)
or unskip it by removing skip: true and skip_reason so it runs with the
context_window tag—locate the YAML for
160b_cpu_per_namespace_graph_with_prom_truncation and either delete the file or
delete the skip: true/skip_reason entries to enable the test, ensuring
tag-filtered executions reflect reality.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In
`@tests/llm/fixtures/test_ask_holmes/160b_cpu_per_namespace_graph_with_prom_truncation/test_case.yaml`:
- Around line 5-6: The test fixture
"160b_cpu_per_namespace_graph_with_prom_truncation" is marked skip: true but was
updated to include the context_window tag, which can mislead tag-filtered test
runs; either remove this obsolete fixture entirely (as done for similar tests)
or unskip it by removing skip: true and skip_reason so it runs with the
context_window tag—locate the YAML for
160b_cpu_per_namespace_graph_with_prom_truncation and either delete the file or
delete the skip: true/skip_reason entries to enable the test, ensuring
tag-filtered executions reflect reality.

---

Duplicate comments:
In
`@tests/llm/fixtures/test_ask_holmes/160a_cpu_per_namespace_graph/test_case.yaml`:
- Line 12: The YAML contains an unverified test tag "context_window" which may
not be registered in pyproject.toml and will cause pytest collection failures;
open pyproject.toml, locate the pytest markers/allowed-tags section (the same
registration used by the verification script referenced in
103_logs_transparency_default_limit/test_case.yaml) and add or confirm the
"context_window" tag is listed there, then re-run the verification script to
ensure the tag is valid for the test fixture.

@aantn
aantn merged commit aa2618a into master Feb 19, 2026
16 of 19 checks passed
@aantn
aantn deleted the claude/review-context-window-evals-hSITx branch February 19, 2026 06:21
arikalon1 pushed a commit that referenced this pull request Feb 21, 2026
…1585)

## Summary
This PR removes duplicate test fixtures for ask_holmes tests and
consolidates context_window-related test tags across the test suite.

## Key Changes
- **Removed duplicate test fixtures:**
- Deleted `47_truncated_logs_context_window/` directory containing test
case, manifest, toolsets configuration, and fixture files for a logs
context window truncation test
- Deleted `160c_cpu_per_namespace_graph_with_global_truncation/`
directory containing test case and toolsets configuration for a
Prometheus metrics test with global truncation

- **Added context_window tag to existing tests:**
- Added `context_window` tag to `103_logs_transparency_default_limit`
test case
- Added `context_window` tag to `160a_cpu_per_namespace_graph` test case
- Added `context_window` tag to
`160b_cpu_per_namespace_graph_with_prom_truncation` test case

## Rationale
The changes consolidate test coverage by removing redundant test
fixtures while ensuring existing tests that validate context window
behavior are properly tagged for organization and filtering purposes.

https://claude.ai/code/session_01TceJLidFkDBiUpG1qmroPZ

---------

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Arik Alon <alon.arik@gmail.com>
moshemorad pushed a commit that referenced this pull request Feb 22, 2026
…1585)

## Summary
This PR removes duplicate test fixtures for ask_holmes tests and
consolidates context_window-related test tags across the test suite.

## Key Changes
- **Removed duplicate test fixtures:**
- Deleted `47_truncated_logs_context_window/` directory containing test
case, manifest, toolsets configuration, and fixture files for a logs
context window truncation test
- Deleted `160c_cpu_per_namespace_graph_with_global_truncation/`
directory containing test case and toolsets configuration for a
Prometheus metrics test with global truncation

- **Added context_window tag to existing tests:**
- Added `context_window` tag to `103_logs_transparency_default_limit`
test case
- Added `context_window` tag to `160a_cpu_per_namespace_graph` test case
- Added `context_window` tag to
`160b_cpu_per_namespace_graph_with_prom_truncation` test case

## Rationale
The changes consolidate test coverage by removing redundant test
fixtures while ensuring existing tests that validate context window
behavior are properly tagged for organization and filtering purposes.

https://claude.ai/code/session_01TceJLidFkDBiUpG1qmroPZ

---------

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Mohse Morad <moshemorad12340@gmail.com>
moshemorad pushed a commit that referenced this pull request Feb 22, 2026
…1585)

## Summary
This PR removes duplicate test fixtures for ask_holmes tests and
consolidates context_window-related test tags across the test suite.

## Key Changes
- **Removed duplicate test fixtures:**
- Deleted `47_truncated_logs_context_window/` directory containing test
case, manifest, toolsets configuration, and fixture files for a logs
context window truncation test
- Deleted `160c_cpu_per_namespace_graph_with_global_truncation/`
directory containing test case and toolsets configuration for a
Prometheus metrics test with global truncation

- **Added context_window tag to existing tests:**
- Added `context_window` tag to `103_logs_transparency_default_limit`
test case
- Added `context_window` tag to `160a_cpu_per_namespace_graph` test case
- Added `context_window` tag to
`160b_cpu_per_namespace_graph_with_prom_truncation` test case

## Rationale
The changes consolidate test coverage by removing redundant test
fixtures while ensuring existing tests that validate context window
behavior are properly tagged for organization and filtering purposes.

https://claude.ai/code/session_01TceJLidFkDBiUpG1qmroPZ

---------

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Mohse Morad <moshemorad12340@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants