Skip to content

Prometheus: spill full query results to disk instead of summary-only; remove old-model steering - #2168

Open
aantn wants to merge 5 commits into
claude/tool-limitations-logs-spillfrom
claude/tool-limitations-prometheus-spill
Open

aantn wants to merge 5 commits into
claude/tool-limitations-logs-spillfrom
claude/tool-limitations-prometheus-spill

Conversation

@aantn

@aantn aantn commented Jun 10, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #2167 (needs ToolInvokeContext.tool_results_dir from that PR). Once #2167 merges, this PR's base will retarget to master.

Problem

When a PromQL query result exceeded the inline token budget, execute_prometheus_instant_query / execute_prometheus_range_query returned only a cardinality summary plus a topk(5, ...) suggestion — the actual data was dropped. The instructions then told the model "NEVER EVER EVER answer a question based on Prometheus data that was truncated… prefer telling the user you can't answer", leaving refusal as the designed outcome.

That steering existed because older models answered from truncated, unordered data and gave wrong answers (e.g. CPU rankings from an arbitrary subset). The fix here keeps the safety property (don't answer from partial data) while giving the model an actual path to the full data.

Changes

  • summarize_large_query_result() (shared by instant + range queries): keeps the cardinality summary for orientation, and when spill-to-disk is available also saves the complete result JSON to disk. The summary gains full_data_file + how_to_read_full_data (cat/jq examples) so the model can compute exact answers without re-querying. Summary-only behavior is preserved when storage is unavailable.
  • MAX_GRAPH_POINTS_HARD_LIMIT default raised 600 → 3000 (10× MAX_GRAPH_POINTS instead of 2×). High-resolution results that don't fit inline now land on disk instead of being dropped, so the cap mainly bounds Prometheus load. Still env-overridable.
  • prometheus_instructions.jinja2 rewrite:
    • "Retrying queries that return too much data" → describes the summary + spilled-file mechanism; tells the model to read the file for exact answers, or refine the query when no file exists.
    • Drops "NEVER EVER EVER…" / "prefer telling the user you can't answer" / "CRITICAL: ALWAYS use topk()". Keeps topk as the recommended pattern for user-facing graphs, warns it hides everything outside the top-k, and still requires disclosing when an answer is based on a partial view.
    • "The toolcall will return no data to you" is now conditional on tool_calls_return_data: false — it was unconditional and wrong for the default configuration.

Testing

  • New unit tests for summarize_large_query_result (spill + fallback paths, instant + range).
  • tests/plugins/toolsets/ + tests/core/: 1262 passed.
  • Instruction template render-tested for both tool_calls_return_data modes.
  • Evals: will trigger /eval with prometheus + regression tags on this PR (opus-4.6) — results to be posted below.

https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq


Generated by Claude Code

Summary by CodeRabbit

  • New Features

    • Improved handling of large Prometheus query results with automatic summarization and optional on-disk spill of full data
    • Increased default upper bound for data points on range queries
  • Documentation

    • Updated Prometheus guidance for interpreting, retrying, and refining large or truncated query results; clarified visualization expectations
  • Tests

    • Added unit tests covering large-query summarization, on-disk spill behavior, and related expectations

When a PromQL query result exceeded the inline token budget, the tools
returned only a cardinality summary plus a topk() suggestion and the
data was dropped — the model had to re-query and often hit the same
wall. The instructions then forbade answering from partial data,
leaving 'tell the user you can't answer' as the only path. This was
designed for older models that answered from truncated, unordered data.

- New summarize_large_query_result() used by both instant and range
  queries: keeps the summary for orientation, and when spill-to-disk is
  available also saves the complete result JSON to a file the model can
  read back with cat/jq (full_data_file + how_to_read_full_data in the
  summary). Falls back to summary-only when storage is unavailable.
- Raise MAX_GRAPH_POINTS_HARD_LIMIT default from 2x to 10x of
  MAX_GRAPH_POINTS (600 -> 3000): oversized results now land on disk
  instead of being dropped, so high-resolution requests are safe.
- Rewrite prometheus_instructions.jinja2: drop the 'NEVER EVER answer /
  prefer refusing / ALWAYS topk(5)' steering in favor of guidance that
  points the model at the spilled file for exact answers, keeps topk
  for user-facing graphs, and still requires disclosing partial views.
- Make 'the toolcall will return no data to you' conditional on
  tool_calls_return_data=false (it was unconditional and wrong for the
  default configuration).

https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq
Signed-off-by: Claude <noreply@anthropic.com>

aantn commented Jun 10, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
tags: prometheus


Generated by Claude Code

@coderabbitai

coderabbitai Bot commented Jun 10, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 5930c1e3-b08c-4a3f-9a44-00b9430149c5

📥 Commits

Reviewing files that changed from the base of the PR and between c37ec7e and ba296ba.

📒 Files selected for processing (1)
  • tests/llm/fixtures/test_ask_holmes/159_prometheus_high_cardinality_cpu/test_case.yaml

Walkthrough

Refactors Prometheus oversized-result handling by centralizing summarization into a helper, increases the graph-points hard cap, updates execution paths to use the helper, refines PromQL instructions for large results, and adds unit tests validating spill-to-disk and summary behavior.

Changes

Prometheus Large-Result Handling Refactor

Layer / File(s) Summary
Hard limit configuration adjustment
holmes/common/env_vars.py
MAX_GRAPH_POINTS_HARD_LIMIT default multiplier increased from 2x to 10x with explanatory comments, raising the upper bound for model-requested max-points overrides on range queries.
Large-result summarization helper refactor
holmes/plugins/toolsets/prometheus/prometheus.py
New summarize_large_query_result helper function standardizes oversized-response handling by clearing inline response data, building cardinality summaries, optionally spilling full results to disk via save_large_result, and augmenting responses with file pointers and debug info.
Call sites updated to use helper
holmes/plugins/toolsets/prometheus/prometheus.py
ExecuteInstantQuery and ExecuteRangeQuery refactored to call summarize_large_query_result(...) in the token_count > token_limit path instead of constructing data_summary/_debug_info inline.
Updated guidance for handling large results
holmes/plugins/toolsets/prometheus/prometheus_instructions.jinja2
Instruction text refined to describe data_summary responses with on-disk full_data_file fallback, retry strategies with tighter queries and max_points adjustment, prohibition against drawing conclusions from partial/truncated data, and conditional notes when tool_calls_return_data is disabled.
Test suite for summarization logic
tests/plugins/toolsets/test_prometheus_unit.py
New TestSummarizeLargeQueryResult suite verifies full raw data is written to disk when tool_results_dir is available and matches original result_data, and confirms summary-only output without disk spill when storage unavailable, covering both range and instant query scenarios.
Test fixture expectation tweak
tests/llm/fixtures/test_ask_holmes/159_prometheus_high_cardinality_cpu/test_case.yaml
Clarifies that an embedded PromQL placeholder produced by the Prometheus range query tool is the visualization the test expects.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • HolmesGPT/holmesgpt#1027: Modifies how the Prometheus token limit (query_response_size_limit_pct) is determined; related to when the summarization path is triggered.
  • HolmesGPT/holmesgpt#1566: Also changes MAX_GRAPH_POINTS_HARD_LIMIT default multiplier affecting max-points clamping used by Prometheus query overrides.
  • HolmesGPT/holmesgpt#975: Adjusts Prometheus oversized-result handling and standardizes data_summary behavior; overlaps with this refactor.

Suggested reviewers

  • moshemorad
  • nherment
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 22.22% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title directly describes the main changes: spilling full query results to disk and removing old-model steering from Prometheus toolset.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jun 10, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for f9f6bd72c (built in 1m 16s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:f9f6bd72c
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:f9f6bd72c me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:f9f6bd72c
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:f9f6bd72c
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:f9f6bd72c
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:f9f6bd72c me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:f9f6bd72c
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:f9f6bd72c

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:f9f6bd72c \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:f9f6bd72c

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:f9f6bd72c \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:f9f6bd72c

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/plugins/toolsets/test_prometheus_unit.py (1)

293-312: ⚡ Quick win

Assert instant-path inline data clearing explicitly.

test_instant_query_summary_with_spill checks summary fields, but it doesn’t verify the core contract that oversized responses clear response.data. Add the same assertion used in the range-path tests to prevent instant-path regressions.

Suggested patch
     def test_instant_query_summary_with_spill(self, tmp_path):
         response, result_data = self._make_response_and_data()
         context = create_mock_tool_invoke_context(
             tool_name="execute_prometheus_instant_query",
             tool_results_dir=tmp_path,
         )

         summarize_large_query_result(
             response_data=response,
             result_data=result_data,
             query=response.query,
             token_count=100_000,
             token_limit=25_000,
             is_range_query=False,
             context=context,
         )

+        assert response.data is None
         assert response.data_summary["result_count"] == 20
         assert response.data_summary["full_data_file"]
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/plugins/toolsets/test_prometheus_unit.py` around lines 293 - 312, The
test test_instant_query_summary_with_spill must assert that oversized
instant-query responses clear the inline payload: after calling
summarize_large_query_result with token_count=100_000 and token_limit=25_000,
add an assertion that response.data is empty/cleared (the same check used in the
range-path tests) to ensure the instant-path contract
(summarize_large_query_result and response.data) is enforced.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@holmes/plugins/toolsets/prometheus/prometheus.py`:
- Around line 885-891: The example written to
response_data.data_summary["how_to_read_full_data"] in
summarize_large_query_result always uses a range-oriented jq snippet
(.values[-1]) even for instant/vector results; update
summarize_large_query_result to detect is_range_query and emit an instant-aware
jq example when is_range_query is False (use .value or the correct instant path)
and keep the .values[-1] example when is_range_query is True, ensuring the
constructed string uses the appropriate snippet in both branches.

---

Nitpick comments:
In `@tests/plugins/toolsets/test_prometheus_unit.py`:
- Around line 293-312: The test test_instant_query_summary_with_spill must
assert that oversized instant-query responses clear the inline payload: after
calling summarize_large_query_result with token_count=100_000 and
token_limit=25_000, add an assertion that response.data is empty/cleared (the
same check used in the range-path tests) to ensure the instant-path contract
(summarize_large_query_result and response.data) is enforced.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: d11a6731-95d8-4692-a8a2-c96d18eac9ac

📥 Commits

Reviewing files that changed from the base of the PR and between 734b514 and a585b33.

📒 Files selected for processing (4)
  • holmes/common/env_vars.py
  • holmes/plugins/toolsets/prometheus/prometheus.py
  • holmes/plugins/toolsets/prometheus/prometheus_instructions.jinja2
  • tests/plugins/toolsets/test_prometheus_unit.py

Comment thread holmes/plugins/toolsets/prometheus/prometheus.py
CodeRabbit review: instant (vector) results have a single 'value' per
series, not 'values', so the suggested jq snippet differed per query
type. Also assert inline data is cleared in the instant-path test.

https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq
Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ⚠️ Completed with 1 failure


⚠️ Eval Results (with failures)

Parameter Value
Triggered via /eval comment
Branch claude/tool-limitations-prometheus-spill
Model opus-4.6
Tags prometheus
Iterations 1
Duration 15m 22s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 14/16 test cases were successful, 1 regressions, 1 skipped
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Denied commands Src
✅ 124_checkout_latency_prometheus[0] 92.0s 7 17 $0.6338 235,513 231,252 59,312 4,261 951 169,058 62,194 858 — — src
✅ 151_disabled_toolsets_fallback_only 46.2s 4 12 $0.2615 87,169 84,865 23,484 2,304 712 60,676 24,189 282 — — src
❌ 159_prometheus_high_cardinality_cpu[0] 24.5s 2 1 $0.3241 62,527 61,826 42,822 701 387 19,001 42,825 105 — — src
✅ 159_prometheus_high_cardinality_cpu[1] 25.8s 3 2 $0.2188 65,263 64,372 25,121 891 384 39,247 25,125 50 — — src
✅ 159_prometheus_high_cardinality_cpu[2] 23.7s 3 2 $0.2097 63,580 62,704 23,889 876 440 38,811 23,893 49 — — src
✅ 160_electricity_market_bidding_bug[0] 151.7s 11 22 $0.9394 453,052 445,970 68,089 7,082 1,151 364,002 81,968 1,874 — — src
✅ 160a_cpu_per_namespace_graph 16.0s 2 1 $0.2160 47,276 46,746 27,781 530 284 18,962 27,784 28 — — src
➖ 160b_cpu_per_namespace_graph_with_prom_truncation — — — — — — — — — — — — — — src
✅ 161_bidding_version_performance[0] 84.2s 5 12 $0.5450 169,927 166,279 55,300 3,648 989 109,226 57,053 933 — — src
✅ 211_prometheus_alerting_rules 21.7s 2 1 $0.1593 39,235 38,767 19,781 468 256 18,983 19,784 30 — — src
✅ 233_compaction_prometheus_data 141.3s 9 34 $1.5664 602,773 594,787 146,498 7,986 1,153 429,490 165,297 799 — — src
✅ 257_victoriametrics_mysql_handlers_spike 103.5s 7 11 $0.5999 217,492 212,469 47,314 5,023 1,472 155,924 56,545 1,229 — — src
✅ 30_basic_promql_graph_cluster_memory 25.1s 2 2 $0.1959 42,948 41,946 22,966 1,002 752 18,977 22,969 179 — — src
✅ 32_basic_promql_graph_pod_cpu 26.0s 3 2 $0.2477 69,009 68,212 29,766 797 339 38,442 29,770 36 — — src
✅ 33_cpu_metrics_discovery 21.7s 2 1 $0.1812 40,139 38,914 19,935 1,225 1,039 18,976 19,938 43 — — src
✅ 34_memory_graph 21.8s 2 1 $0.2591 52,972 52,250 33,277 722 401 18,970 33,280 54 — — src
Total 55.0s avg 4.3 avg 8.1 avg $6.5579 2,248,875 2,211,359 146,498 37,516 1,472 1,518,745 692,614 6,549 — —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 34 test/model combinations loaded

No baseline data available for comparison.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tool-limitations-prometheus-spill -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, gpt-5.5, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, opus-4.8, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tool-limitations-prometheus-spill -f markers=regression -f filter=

aantn commented Jun 10, 2026

Copy link
Copy Markdown
Collaborator Author

Analysis of the one failure above (159_prometheus_high_cardinality_cpu[0], 2/3 iterations passed): pulled the Braintrust trace — the judge confirmed expected element 1 passed (Holmes correctly identified the z-cpu-heavy-* pods with data-grounded values: ~500m each vs ~0.2m for the 37 low-cpu pods), and failed it only on element 2, reasoning that the promql embed << {"type": "promql", "tool_name": "execute_prometheus_range_query", ...} >> "does not equate to a visualization being included".

That embed IS the documented graph-visualization mechanism (post-processing renders it into a chart), and the identical format passed in iterations [1], [2] and in 30/32/34/160a in the same run. So this is judge flakiness on the embed convention, not a behavior regression from this PR. Re-running below to confirm. The skipped 160b is skip: true on master (pre-existing, unrelated).


Generated by Claude Code

aantn commented Jun 10, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
id: 159_prometheus_high_cardinality_cpu
iterations: 2


Generated by Claude Code

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ⚠️ Completed with 1 failure


⚠️ Eval Results (with failures)

Parameter Value
Triggered via /eval comment
Branch claude/tool-limitations-prometheus-spill
Model opus-4.6
Tags all LLM tests
ID (-k) 159_prometheus_high_cardinality_cpu
Iterations 2
Duration 7m 24s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 5/6 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Denied commands Src
✅ 159_prometheus_high_cardinality_cpu[0] 34.5s 2 1 $0.2830 55,865 54,939 35,935 926 500 19,001 35,938 280 — — src
❌ 159_prometheus_high_cardinality_cpu[0] 34.7s 2 1 $0.2797 55,864 55,095 36,091 769 437 19,001 36,094 152 — — src
✅ 159_prometheus_high_cardinality_cpu[1] 26.2s 2 1 $0.1833 42,598 42,094 23,121 504 311 18,970 23,124 37 — — src
✅ 159_prometheus_high_cardinality_cpu[1] 24.6s 2 1 $0.1879 43,358 42,884 23,911 474 311 18,970 23,914 37 — — src
✅ 159_prometheus_high_cardinality_cpu[2] 35.0s 3 2 $0.1952 61,661 60,845 22,025 816 370 38,816 22,029 37 — — src
✅ 159_prometheus_high_cardinality_cpu[2] 34.4s 3 2 $0.1960 61,687 60,842 22,023 845 399 38,815 22,027 37 — — src
Total 31.6s avg 2.3 avg 1.3 avg $1.3251 321,033 316,699 36,091 4,334 500 153,573 163,126 580 — —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 34 test/model combinations loaded

No baseline data available for comparison.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tool-limitations-prometheus-spill -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, gpt-5.5, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, opus-4.8, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tool-limitations-prometheus-spill -f markers=regression -f filter=

The LLM judge intermittently failed iterations whose output embedded
the graph via the standard promql placeholder, ruling the placeholder
'does not equate to a visualization' while accepting the identical
format in other iterations of the same run. State explicitly in
expected_output that the embed IS the graph (post-processing renders
it into a chart), which is how Holmes is instructed to show graphs.

https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq
Signed-off-by: Claude <noreply@anthropic.com>

aantn commented Jun 10, 2026

Copy link
Copy Markdown
Collaborator Author

Rerun analysis confirms judge flakiness, not a regression: the two 159[0] iterations produced essentially identical outputs (same promql embed, same z-cpu-heavy ranking with correct values), and the judge passed one ("placeholder satisfies the requirement") while failing the other ("no explicit rendering of a graph"). Prompts [1] and [2] passed all 4 iterations.

Pushed ba296ba clarifying in the fixture's expected_output that the promql embed IS the graph visualization (post-processing renders it into a chart — exactly what the prometheus instructions tell Holmes to do). Verifying with one more targeted run:

/eval
id: 159_prometheus_high_cardinality_cpu
iterations: 2


Generated by Claude Code

aantn commented Jun 10, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
id: 159_prometheus_high_cardinality_cpu
iterations: 2


Generated by Claude Code

aantn added a commit that referenced this pull request Jun 10, 2026
Documentation-only change, no behavior change.

The `llm_summarize` transformer predates the spill-to-disk mechanism, is
disabled by default, and never worked well in practice: summarization is
lossy (the original tool output is unrecoverable afterwards) and it adds
latency and cost to every large tool call. Modern models do better
working from the full data spilled to disk.

This marks it as legacy in the module/class docstrings and adds a
warning admonition to `docs/development/transformers.md`, so future
contributors don't build on it and users don't enable it expecting good
results. Kept for backwards compatibility with existing configs that
reference it.

Part of a series of targeted fixes removing old-model data limitations
from built-in tools (see #2167, #2168, #2169).

https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq

---
_Generated by [Claude
Code](https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq)_

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added notices marking the `llm_summarize` transformer as a legacy
feature that is disabled by default. This feature is not recommended for
new configurations due to lossy summarization, increased latency, and
associated costs. Documentation has been updated to recommend the
spill-to-disk mechanism as the preferred alternative for handling
oversized tool results.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ✅ Completed successfully


🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Branch claude/tool-limitations-prometheus-spill
Model opus-4.6
Tags all LLM tests
ID (-k) 159_prometheus_high_cardinality_cpu
Iterations 2
Duration 27m 59s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 3/6 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Denied commands Src
🚫 159_prometheus_high_cardinality_cpu[0] — — — — — — — — — — — — — — src
🚫 159_prometheus_high_cardinality_cpu[0] — — — — — — — — — — — — — — src
✅ 159_prometheus_high_cardinality_cpu[1] 25.1s 2 1 $0.2994 59,555 59,072 40,099 483 303 18,970 40,102 39 — — src
✅ 159_prometheus_high_cardinality_cpu[1] 26.9s 2 1 $0.2999 59,587 59,091 40,118 496 310 18,970 40,121 36 — — src
🚫 159_prometheus_high_cardinality_cpu[2] — — — — — — — — — — — — — — src
✅ 159_prometheus_high_cardinality_cpu[2] 41.1s 3 2 $0.2539 69,815 68,874 30,056 941 495 38,814 30,060 57 — — src
Total 31.1s avg 2.3 avg 1.3 avg $0.8532 188,957 187,037 40,118 1,920 495 76,754 110,283 132 — —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 34 test/model combinations loaded

No baseline data available for comparison.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tool-limitations-prometheus-spill -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, gpt-5.5, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, opus-4.8, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tool-limitations-prometheus-spill -f markers=regression -f filter=

aantn commented Jun 10, 2026

Copy link
Copy Markdown
Collaborator Author

/rerun


Generated by Claude Code

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ✅ Completed successfully


🧪 Manual Eval Results

Parameter Value
Triggered via /rerun → /eval
Branch claude/tool-limitations-prometheus-spill
Model opus-4.6
Tags all LLM tests
ID (-k) 159_prometheus_high_cardinality_cpu
Iterations 2
Duration 28m 4s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 6/6 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Denied commands Src
✅ 159_prometheus_high_cardinality_cpu[0] 30.0s 2 1 $0.2796 55,885 55,128 36,124 757 411 19,001 36,127 126 — — src
✅ 159_prometheus_high_cardinality_cpu[0] 37.7s 2 1 $0.3080 58,978 57,878 38,874 1,100 638 19,001 38,877 402 — — src
✅ 159_prometheus_high_cardinality_cpu[1] 25.4s 2 1 $0.1836 42,650 42,147 23,174 503 310 18,970 23,177 36 — — src
✅ 159_prometheus_high_cardinality_cpu[1] 27.2s 2 1 $0.1919 43,822 43,305 24,332 517 303 18,970 24,335 39 — — src
✅ 159_prometheus_high_cardinality_cpu[2] 34.8s 3 2 $0.1954 61,653 60,830 22,017 823 377 38,809 22,021 37 — — src
✅ 159_prometheus_high_cardinality_cpu[2] 23.8s 3 2 $0.2543 69,978 69,069 30,251 909 473 38,814 30,255 49 — — src
Total 29.8s avg 2.3 avg 1.3 avg $1.4129 332,966 328,357 38,874 4,609 638 153,565 174,792 689 — —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 34 test/model combinations loaded

No baseline data available for comparison.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tool-limitations-prometheus-spill -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, gpt-5.5, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, opus-4.8, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tool-limitations-prometheus-spill -f markers=regression -f filter=

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants