Repository navigation
Prometheus: spill full query results to disk instead of summary-only; remove old-model steering - #2168
Prometheus: spill full query results to disk instead of summary-only; remove old-model steering#2168aantn wants to merge 5 commits into
Conversation
When a PromQL query result exceeded the inline token budget, the tools returned only a cardinality summary plus a topk() suggestion and the data was dropped — the model had to re-query and often hit the same wall. The instructions then forbade answering from partial data, leaving 'tell the user you can't answer' as the only path. This was designed for older models that answered from truncated, unordered data. - New summarize_large_query_result() used by both instant and range queries: keeps the summary for orientation, and when spill-to-disk is available also saves the complete result JSON to a file the model can read back with cat/jq (full_data_file + how_to_read_full_data in the summary). Falls back to summary-only when storage is unavailable. - Raise MAX_GRAPH_POINTS_HARD_LIMIT default from 2x to 10x of MAX_GRAPH_POINTS (600 -> 3000): oversized results now land on disk instead of being dropped, so high-resolution requests are safe. - Rewrite prometheus_instructions.jinja2: drop the 'NEVER EVER answer / prefer refusing / ALWAYS topk(5)' steering in favor of guidance that points the model at the spilled file for exact answers, keeps topk for user-facing graphs, and still requires disclosing partial views. - Make 'the toolcall will return no data to you' conditional on tool_calls_return_data=false (it was unconditional and wrong for the default configuration). https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq Signed-off-by: Claude <noreply@anthropic.com>
|
/eval Generated by Claude Code |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (1)
WalkthroughRefactors Prometheus oversized-result handling by centralizing summarization into a helper, increases the graph-points hard cap, updates execution paths to use the helper, refines PromQL instructions for large results, and adds unit tests validating spill-to-disk and summary behavior. ChangesPrometheus Large-Result Handling Refactor
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:f9f6bd72c
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:f9f6bd72c me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:f9f6bd72c
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:f9f6bd72c
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:f9f6bd72c
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:f9f6bd72c me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:f9f6bd72c
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:f9f6bd72cPatch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:f9f6bd72c \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:f9f6bd72cRobusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:f9f6bd72c \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:f9f6bd72c |
There was a problem hiding this comment.
Actionable comments posted: 1
🧹 Nitpick comments (1)
tests/plugins/toolsets/test_prometheus_unit.py (1)
293-312: ⚡ Quick winAssert instant-path inline data clearing explicitly.
test_instant_query_summary_with_spillchecks summary fields, but it doesn’t verify the core contract that oversized responses clearresponse.data. Add the same assertion used in the range-path tests to prevent instant-path regressions.Suggested patch
def test_instant_query_summary_with_spill(self, tmp_path): response, result_data = self._make_response_and_data() context = create_mock_tool_invoke_context( tool_name="execute_prometheus_instant_query", tool_results_dir=tmp_path, ) summarize_large_query_result( response_data=response, result_data=result_data, query=response.query, token_count=100_000, token_limit=25_000, is_range_query=False, context=context, ) + assert response.data is None assert response.data_summary["result_count"] == 20 assert response.data_summary["full_data_file"]🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/plugins/toolsets/test_prometheus_unit.py` around lines 293 - 312, The test test_instant_query_summary_with_spill must assert that oversized instant-query responses clear the inline payload: after calling summarize_large_query_result with token_count=100_000 and token_limit=25_000, add an assertion that response.data is empty/cleared (the same check used in the range-path tests) to ensure the instant-path contract (summarize_large_query_result and response.data) is enforced.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@holmes/plugins/toolsets/prometheus/prometheus.py`:
- Around line 885-891: The example written to
response_data.data_summary["how_to_read_full_data"] in
summarize_large_query_result always uses a range-oriented jq snippet
(.values[-1]) even for instant/vector results; update
summarize_large_query_result to detect is_range_query and emit an instant-aware
jq example when is_range_query is False (use .value or the correct instant path)
and keep the .values[-1] example when is_range_query is True, ensuring the
constructed string uses the appropriate snippet in both branches.
---
Nitpick comments:
In `@tests/plugins/toolsets/test_prometheus_unit.py`:
- Around line 293-312: The test test_instant_query_summary_with_spill must
assert that oversized instant-query responses clear the inline payload: after
calling summarize_large_query_result with token_count=100_000 and
token_limit=25_000, add an assertion that response.data is empty/cleared (the
same check used in the range-path tests) to ensure the instant-path contract
(summarize_large_query_result and response.data) is enforced.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: d11a6731-95d8-4692-a8a2-c96d18eac9ac
📒 Files selected for processing (4)
holmes/common/env_vars.pyholmes/plugins/toolsets/prometheus/prometheus.pyholmes/plugins/toolsets/prometheus/prometheus_instructions.jinja2tests/plugins/toolsets/test_prometheus_unit.py
CodeRabbit review: instant (vector) results have a single 'value' per series, not 'values', so the suggested jq snippet differed per query type. Also assert inline data is cleared in the instant-path test. https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq Signed-off-by: Claude <noreply@anthropic.com>
|
@aantn Your eval run has finished.
|
| Parameter | Value |
|---|---|
| Triggered via | /eval comment |
| Branch | claude/tool-limitations-prometheus-spill |
| Model | opus-4.6 |
| Tags | prometheus |
| Iterations | 1 |
| Duration | 15m 22s |
| Workflow | View logs | Rerun |
Results of HolmesGPT evals
- ask_holmes: 14/16 test cases were successful, 1 regressions, 1 skipped
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions | Denied commands | Src |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 124_checkout_latency_prometheus[0] | 92.0s | 7 | 17 | $0.6338 | 235,513 | 231,252 | 59,312 | 4,261 | 951 | 169,058 | 62,194 | 858 | — | — | src |
| ✅ | 151_disabled_toolsets_fallback_only | 46.2s | 4 | 12 | $0.2615 | 87,169 | 84,865 | 23,484 | 2,304 | 712 | 60,676 | 24,189 | 282 | — | — | src |
| ❌ | 159_prometheus_high_cardinality_cpu[0] | 24.5s | 2 | 1 | $0.3241 | 62,527 | 61,826 | 42,822 | 701 | 387 | 19,001 | 42,825 | 105 | — | — | src |
| ✅ | 159_prometheus_high_cardinality_cpu[1] | 25.8s | 3 | 2 | $0.2188 | 65,263 | 64,372 | 25,121 | 891 | 384 | 39,247 | 25,125 | 50 | — | — | src |
| ✅ | 159_prometheus_high_cardinality_cpu[2] | 23.7s | 3 | 2 | $0.2097 | 63,580 | 62,704 | 23,889 | 876 | 440 | 38,811 | 23,893 | 49 | — | — | src |
| ✅ | 160_electricity_market_bidding_bug[0] | 151.7s | 11 | 22 | $0.9394 | 453,052 | 445,970 | 68,089 | 7,082 | 1,151 | 364,002 | 81,968 | 1,874 | — | — | src |
| ✅ | 160a_cpu_per_namespace_graph | 16.0s | 2 | 1 | $0.2160 | 47,276 | 46,746 | 27,781 | 530 | 284 | 18,962 | 27,784 | 28 | — | — | src |
| ➖ | 160b_cpu_per_namespace_graph_with_prom_truncation | — | — | — | — | — | — | — | — | — | — | — | — | — | — | src |
| ✅ | 161_bidding_version_performance[0] | 84.2s | 5 | 12 | $0.5450 | 169,927 | 166,279 | 55,300 | 3,648 | 989 | 109,226 | 57,053 | 933 | — | — | src |
| ✅ | 211_prometheus_alerting_rules | 21.7s | 2 | 1 | $0.1593 | 39,235 | 38,767 | 19,781 | 468 | 256 | 18,983 | 19,784 | 30 | — | — | src |
| ✅ | 233_compaction_prometheus_data | 141.3s | 9 | 34 | $1.5664 | 602,773 | 594,787 | 146,498 | 7,986 | 1,153 | 429,490 | 165,297 | 799 | — | — | src |
| ✅ | 257_victoriametrics_mysql_handlers_spike | 103.5s | 7 | 11 | $0.5999 | 217,492 | 212,469 | 47,314 | 5,023 | 1,472 | 155,924 | 56,545 | 1,229 | — | — | src |
| ✅ | 30_basic_promql_graph_cluster_memory | 25.1s | 2 | 2 | $0.1959 | 42,948 | 41,946 | 22,966 | 1,002 | 752 | 18,977 | 22,969 | 179 | — | — | src |
| ✅ | 32_basic_promql_graph_pod_cpu | 26.0s | 3 | 2 | $0.2477 | 69,009 | 68,212 | 29,766 | 797 | 339 | 38,442 | 29,770 | 36 | — | — | src |
| ✅ | 33_cpu_metrics_discovery | 21.7s | 2 | 1 | $0.1812 | 40,139 | 38,914 | 19,935 | 1,225 | 1,039 | 18,976 | 19,938 | 43 | — | — | src |
| ✅ | 34_memory_graph | 21.8s | 2 | 1 | $0.2591 | 52,972 | 52,250 | 33,277 | 722 | 401 | 18,970 | 33,280 | 54 | — | — | src |
| Total | 55.0s avg | 4.3 avg | 8.1 avg | $6.5579 | 2,248,875 | 2,211,359 | 146,498 | 37,516 | 1,472 | 1,518,745 | 692,614 | 6,549 | — | — |
Benchmark Comparison Details
Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded
- master-27089078550 (created: 2026-06-07)
Benchmark baseline: latest ci-benchmark experiment on master
Status: 34 test/model combinations loaded
- ci-benchmark-27257681111 (created: 2026-06-10)
No baseline data available for comparison.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 1 Failure Detected
📖 Legend
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tool-limitations-prometheus-spill -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
tags: regression
Or with more options (one per line):
/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
tags: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
tags |
Pytest tags / markers (no default - runs all tests!) |
id |
Eval ID / pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):
| Label | Effect |
|---|---|
evals-tag-<name> |
Run tests with tag <name> alongside regression |
evals-id-<name> |
Run a specific eval by test ID |
evals-model-<name> |
Override the model (use model list name, e.g. sonnet-4.5) |
Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5
🏷️ Valid tags
benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs
🤖 Valid models
deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, gpt-5.5, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, opus-4.8, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tool-limitations-prometheus-spill -f markers=regression -f filter=
|
Analysis of the one failure above ( That embed IS the documented graph-visualization mechanism (post-processing renders it into a chart), and the identical format passed in iterations [1], [2] and in 30/32/34/160a in the same run. So this is judge flakiness on the embed convention, not a behavior regression from this PR. Re-running below to confirm. The skipped Generated by Claude Code |
|
/eval Generated by Claude Code |
https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq Signed-off-by: Claude <noreply@anthropic.com>
…mitations-prometheus-spill Signed-off-by: Claude <noreply@anthropic.com>
|
@aantn Your eval run has finished.
|
| Parameter | Value |
|---|---|
| Triggered via | /eval comment |
| Branch | claude/tool-limitations-prometheus-spill |
| Model | opus-4.6 |
| Tags | all LLM tests |
| ID (-k) | 159_prometheus_high_cardinality_cpu |
| Iterations | 2 |
| Duration | 7m 24s |
| Workflow | View logs | Rerun |
Results of HolmesGPT evals
- ask_holmes: 5/6 test cases were successful, 1 regressions
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions | Denied commands | Src |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 159_prometheus_high_cardinality_cpu[0] | 34.5s | 2 | 1 | $0.2830 | 55,865 | 54,939 | 35,935 | 926 | 500 | 19,001 | 35,938 | 280 | — | — | src |
| ❌ | 159_prometheus_high_cardinality_cpu[0] | 34.7s | 2 | 1 | $0.2797 | 55,864 | 55,095 | 36,091 | 769 | 437 | 19,001 | 36,094 | 152 | — | — | src |
| ✅ | 159_prometheus_high_cardinality_cpu[1] | 26.2s | 2 | 1 | $0.1833 | 42,598 | 42,094 | 23,121 | 504 | 311 | 18,970 | 23,124 | 37 | — | — | src |
| ✅ | 159_prometheus_high_cardinality_cpu[1] | 24.6s | 2 | 1 | $0.1879 | 43,358 | 42,884 | 23,911 | 474 | 311 | 18,970 | 23,914 | 37 | — | — | src |
| ✅ | 159_prometheus_high_cardinality_cpu[2] | 35.0s | 3 | 2 | $0.1952 | 61,661 | 60,845 | 22,025 | 816 | 370 | 38,816 | 22,029 | 37 | — | — | src |
| ✅ | 159_prometheus_high_cardinality_cpu[2] | 34.4s | 3 | 2 | $0.1960 | 61,687 | 60,842 | 22,023 | 845 | 399 | 38,815 | 22,027 | 37 | — | — | src |
| Total | 31.6s avg | 2.3 avg | 1.3 avg | $1.3251 | 321,033 | 316,699 | 36,091 | 4,334 | 500 | 153,573 | 163,126 | 580 | — | — |
Benchmark Comparison Details
Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded
- master-27089078550 (created: 2026-06-07)
Benchmark baseline: latest ci-benchmark experiment on master
Status: 34 test/model combinations loaded
- ci-benchmark-27257681111 (created: 2026-06-10)
No baseline data available for comparison.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 1 Failure Detected
📖 Legend
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tool-limitations-prometheus-spill -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
tags: regression
Or with more options (one per line):
/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
tags: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
tags |
Pytest tags / markers (no default - runs all tests!) |
id |
Eval ID / pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):
| Label | Effect |
|---|---|
evals-tag-<name> |
Run tests with tag <name> alongside regression |
evals-id-<name> |
Run a specific eval by test ID |
evals-model-<name> |
Override the model (use model list name, e.g. sonnet-4.5) |
Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5
🏷️ Valid tags
benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs
🤖 Valid models
deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, gpt-5.5, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, opus-4.8, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tool-limitations-prometheus-spill -f markers=regression -f filter=
The LLM judge intermittently failed iterations whose output embedded the graph via the standard promql placeholder, ruling the placeholder 'does not equate to a visualization' while accepting the identical format in other iterations of the same run. State explicitly in expected_output that the embed IS the graph (post-processing renders it into a chart), which is how Holmes is instructed to show graphs. https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq Signed-off-by: Claude <noreply@anthropic.com>
|
Rerun analysis confirms judge flakiness, not a regression: the two Pushed /eval Generated by Claude Code |
|
/eval Generated by Claude Code |
Documentation-only change, no behavior change. The `llm_summarize` transformer predates the spill-to-disk mechanism, is disabled by default, and never worked well in practice: summarization is lossy (the original tool output is unrecoverable afterwards) and it adds latency and cost to every large tool call. Modern models do better working from the full data spilled to disk. This marks it as legacy in the module/class docstrings and adds a warning admonition to `docs/development/transformers.md`, so future contributors don't build on it and users don't enable it expecting good results. Kept for backwards compatibility with existing configs that reference it. Part of a series of targeted fixes removing old-model data limitations from built-in tools (see #2167, #2168, #2169). https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq --- _Generated by [Claude Code](https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq)_ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added notices marking the `llm_summarize` transformer as a legacy feature that is disabled by default. This feature is not recommended for new configurations due to lossy summarization, increased latency, and associated costs. Documentation has been updated to recommend the spill-to-disk mechanism as the preferred alternative for handling oversized tool results. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Claude <noreply@anthropic.com> Co-authored-by: Claude <noreply@anthropic.com>
|
@aantn Your eval run has finished. ✅ Completed successfully 🧪 Manual Eval Results
Results of HolmesGPT evals
Benchmark Comparison DetailsMaster baseline: latest master-* experiment (post-merge regression eval)
Benchmark baseline: latest ci-benchmark experiment on master
No baseline data available for comparison. Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" Option 3: Add PR labels to include extra evals (applies to both automatic runs and
Examples: 🏷️ Valid tags
🤖 Valid models
Commands: CLI: |
|
/rerun Generated by Claude Code |
|
@aantn Your eval run has finished. ✅ Completed successfully 🧪 Manual Eval Results
Results of HolmesGPT evals
Benchmark Comparison DetailsMaster baseline: latest master-* experiment (post-merge regression eval)
Benchmark baseline: latest ci-benchmark experiment on master
No baseline data available for comparison. Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" Option 3: Add PR labels to include extra evals (applies to both automatic runs and
Examples: 🏷️ Valid tags
🤖 Valid models
Commands: CLI: |
Problem
When a PromQL query result exceeded the inline token budget,
execute_prometheus_instant_query/execute_prometheus_range_queryreturned only a cardinality summary plus atopk(5, ...)suggestion — the actual data was dropped. The instructions then told the model "NEVER EVER EVER answer a question based on Prometheus data that was truncated… prefer telling the user you can't answer", leaving refusal as the designed outcome.That steering existed because older models answered from truncated, unordered data and gave wrong answers (e.g. CPU rankings from an arbitrary subset). The fix here keeps the safety property (don't answer from partial data) while giving the model an actual path to the full data.
Changes
summarize_large_query_result()(shared by instant + range queries): keeps the cardinality summary for orientation, and when spill-to-disk is available also saves the complete result JSON to disk. The summary gainsfull_data_file+how_to_read_full_data(cat/jq examples) so the model can compute exact answers without re-querying. Summary-only behavior is preserved when storage is unavailable.MAX_GRAPH_POINTS_HARD_LIMITdefault raised 600 → 3000 (10×MAX_GRAPH_POINTSinstead of 2×). High-resolution results that don't fit inline now land on disk instead of being dropped, so the cap mainly bounds Prometheus load. Still env-overridable.prometheus_instructions.jinja2rewrite:topkas the recommended pattern for user-facing graphs, warns it hides everything outside the top-k, and still requires disclosing when an answer is based on a partial view.tool_calls_return_data: false— it was unconditional and wrong for the default configuration.Testing
summarize_large_query_result(spill + fallback paths, instant + range).tests/plugins/toolsets/+tests/core/: 1262 passed.tool_calls_return_datamodes./evalwith prometheus + regression tags on this PR (opus-4.6) — results to be posted below.https://claude.ai/code/session_01BwJeGAGBLoby5rShhQADwq
Generated by Claude Code
Summary by CodeRabbit
New Features
Documentation
Tests