Skip to content

Increase default max_points to 500 and allow LLM override up to 1000 - #1566

Merged
aantn merged 11 commits into
masterfrom
claude/fix-prometheus-downsampling-jZZnp
Feb 16, 2026
Merged

aantn merged 11 commits into
masterfrom
claude/fix-prometheus-downsampling-jZZnp

Conversation

@aantn

@aantn aantn commented Feb 15, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR improves Prometheus query resolution by increasing the default maximum data points from 100 to 500, and allowing the LLM to request up to 1000 points (2x the default) for high-resolution analysis of low-cardinality queries. The change includes updated logic, comprehensive tests, and improved documentation.

Key Changes

  • Increased default MAX_GRAPH_POINTS: Changed from 100 to 500 to provide better default resolution for time series queries
  • Implemented hard limit for max_points override: LLM can now request up to 2x MAX_GRAPH_POINTS (1000 points) for higher resolution, with a hard cap to prevent excessive data retrieval
  • Updated step calculation logic: When no step is provided, the default now targets the configured max_points instead of a fixed 60-point target
  • Enhanced parameter descriptions: Updated tool parameter documentation to clarify:
    • step parameter now explains the relationship between step size and data points
    • max_points parameter now includes guidance on when to increase (low-cardinality queries) vs decrease (overview graphs, high-cardinality queries)
  • Added comprehensive test coverage: New test class TestMaxPointsOverride with 5 test cases covering:
    • Override above default (allowed)
    • Override capped at hard limit
    • Override below default (allowed)
    • Invalid override fallback behavior
    • Interaction between explicit step and max_points override
  • Updated Prometheus instructions: Added new section on query resolution and data points with guidance on:
    • Default 500-point limit per series
    • When to increase max_points for spike/anomaly detection
    • Cardinality considerations
    • Alternative approach of narrowing time range for higher resolution

Implementation Details

  • The hard limit is calculated as MAX_GRAPH_POINTS * 2 to allow flexibility while maintaining safety
  • Invalid overrides (< 1) fall back to the default MAX_GRAPH_POINTS
  • When both step and max_points_override are provided, the step is adjusted if it would exceed the max_points target
  • Logging messages updated to reflect the new behavior and hard limits
  • All existing tests updated to use monkeypatch for MAX_GRAPH_POINTS to ensure test isolation

https://claude.ai/code/session_0174s8zNS3peCtHc2tg9iJ8y

Summary by CodeRabbit

  • New Features

    • Increased default maximum graph data points from 100 to 300.
    • Added a configurable hard limit (default 2× the max) that clamps and warns on excessive max-points.
  • Improvements

    • Step-size calculation now targets the configured data-point count; clearer messaging about limits and overrides.
    • Expanded parameter guidance for resolution vs. data points.
  • Documentation

    • Added query-resolution and data-point guidance with examples for Prometheus tools.
  • Tests

    • Added and updated tests covering step adjustments, overrides, and hard-limit behavior.

The default MAX_GRAPH_POINTS was 100 with a hardcoded default of 60
data points, making graphs look completely different from Grafana.
For a 6-hour range, Holmes showed 60 points (1 per 6 min) while
Grafana shows ~1440 points at 15s resolution.

Changes:
- Raise MAX_GRAPH_POINTS default from 100 to 500
- Default step now targets max_points (not hardcoded 60)
- Allow LLM to request higher resolution (up to 5x default)
  for simple single-series queries
- Token-based truncation remains as safety net for large responses

https://claude.ai/code/session_0174s8zNS3peCtHc2tg9iJ8y
Signed-off-by: Claude <noreply@anthropic.com>
…flows

For high-cardinality queries (many time series), a 5x override could
generate responses that exceed the tool call token budget. Reducing to
2x still allows meaningful resolution increase for low-cardinality
queries while keeping token usage reasonable. The tool description now
explicitly guides the LLM to only increase above default for queries
returning 1-3 series.

https://claude.ai/code/session_0174s8zNS3peCtHc2tg9iJ8y
Signed-off-by: Claude <noreply@anthropic.com>
The LLM had no guidance on how resolution affects data quality.
Added instructions section explaining:
- Default 500 points per series, controllable via max_points
- Increase max_points (up to 1000) when investigating spikes/anomalies
  since they can disappear at lower resolution
- Only increase for low-cardinality queries (1-3 series)
- Narrow the time range for even higher resolution
Also improved the step parameter description.

https://claude.ai/code/session_0174s8zNS3peCtHc2tg9iJ8y
Signed-off-by: Claude <noreply@anthropic.com>
@netlify

netlify Bot commented Feb 15, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 13dfd9f
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/699303df1fb29f0008d63694
😎 Deploy Preview https://deploy-preview-1566--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Feb 15, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 22cd0c0 (#22060892440)

✅ Results of HolmesGPT evals

Automatically triggered by commit 22cd0c0 on branch claude/fix-prometheus-downsampling-jZZnp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.7s 5 11 $0.2346
✅ 101_loki_historical_logs_pod_deleted 51.3s 7 12 $0.3019
✅ 111_pod_names_contain_service 34.6s 5 11 $0.2279
✅ 112_find_pvcs_by_uuid 37.1s 7 9 $0.2685
✅ 12_job_crashing 25.0s 4 7 $0.1989
✅ 176_network_policy_blocking_traffic_no_runbooks 48.1s 6 17 $0.2911
✅ 24_misconfigured_pvc 31.6s 5 13 $0.2283
✅ 43_current_datetime_from_prompt 4.8s 1 — $0.1068
✅ 61_exact_match_counting 13.6s 3 2 $0.1461
Total 31.0s avg 4.8 avg 10.2 avg $2.0041
📜 Run @ 4b5c604 (#22039013933)

✅ Results of HolmesGPT evals

Automatically triggered by commit 4b5c604 on branch claude/fix-prometheus-downsampling-jZZnp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 31.6s 5 10 $0.2245
✅ 101_loki_historical_logs_pod_deleted 38.4s 5 10 $0.2372
✅ 111_pod_names_contain_service 33.3s 5 11 $0.2275
✅ 112_find_pvcs_by_uuid 30.2s 5 8 $0.2372
✅ 12_job_crashing 36.9s 6 12 $0.2534
✅ 176_network_policy_blocking_traffic_no_runbooks 42.2s 6 14 $0.2690
✅ 24_misconfigured_pvc 35.5s 6 14 $0.2476
✅ 43_current_datetime_from_prompt 5.0s 1 — $0.0114
✅ 61_exact_match_counting 17.3s 4 4 $0.1639
Total 30.1s avg 4.8 avg 10.4 avg $1.8716
📜 Run @ 9f6de64 (#22038847415)

✅ Results of HolmesGPT evals

Automatically triggered by commit 9f6de64 on branch claude/fix-prometheus-downsampling-jZZnp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 34.2s 5 11 $0.2299
✅ 101_loki_historical_logs_pod_deleted 42.9s 5 10 $0.2504
✅ 111_pod_names_contain_service 33.7s 5 11 $0.2256
✅ 112_find_pvcs_by_uuid 31.9s 6 7 $0.2430
✅ 12_job_crashing 34.8s 5 12 $0.2449
✅ 176_network_policy_blocking_traffic_no_runbooks 44.4s 6 17 $0.2968
✅ 24_misconfigured_pvc 41.9s 7 15 $0.2656
✅ 43_current_datetime_from_prompt 6.3s 1 — $0.1080
✅ 61_exact_match_counting 17.6s 4 4 $0.1622
Total 32.0s avg 4.9 avg 10.9 avg $2.0264
📜 Run @ 49ccaa4 (#22038698023)

✅ Results of HolmesGPT evals

Automatically triggered by commit 49ccaa4 on branch claude/fix-prometheus-downsampling-jZZnp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 27.8s 4 10 $0.2216
✅ 101_loki_historical_logs_pod_deleted 54.7s 8 16 $0.3396
✅ 111_pod_names_contain_service 37.4s 6 14 $0.2631
✅ 112_find_pvcs_by_uuid 31.7s 6 6 $0.2336
✅ 12_job_crashing 44.6s 7 16 $0.3046
✅ 176_network_policy_blocking_traffic_no_runbooks 39.1s 6 13 $0.2656
✅ 24_misconfigured_pvc 32.8s 5 13 $0.2326
✅ 43_current_datetime_from_prompt 4.7s 1 — $0.0114
✅ 61_exact_match_counting 12.5s 3 2 $0.1459
Total 31.7s avg 5.1 avg 11.2 avg $2.0180
📜 Run @ 0080e46 (#22034151619)

✅ Results of HolmesGPT evals

Automatically triggered by commit 0080e46 on branch claude/fix-prometheus-downsampling-jZZnp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 31.8s 5 10 $0.2266
✅ 101_loki_historical_logs_pod_deleted 39.5s 5 10 $0.2477
✅ 111_pod_names_contain_service 33.7s 5 11 $0.2307
✅ 112_find_pvcs_by_uuid 33.2s 6 7 $0.2374
✅ 12_job_crashing 38.8s 6 14 $0.2744
✅ 176_network_policy_blocking_traffic_no_runbooks 47.6s 7 16 $0.2982
✅ 24_misconfigured_pvc 32.5s 5 13 $0.2286
✅ 43_current_datetime_from_prompt 5.3s 1 — $0.1072
✅ 61_exact_match_counting 12.8s 3 2 $0.1441
Total 30.6s avg 4.8 avg 10.4 avg $1.9950

✅ Results of HolmesGPT evals

Automatically triggered by commit 13dfd9f on branch claude/fix-prometheus-downsampling-jZZnp

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 32.0s 5 10 $0.2261
✅ 101_loki_historical_logs_pod_deleted 60.7s 8 13 $0.3290
✅ 111_pod_names_contain_service 35.6s 5 11 $0.2286
✅ 112_find_pvcs_by_uuid 35.5s 6 7 $0.2460
✅ 12_job_crashing 41.7s 6 15 $0.2768
✅ 176_network_policy_blocking_traffic_no_runbooks 49.8s 7 17 $0.2910
✅ 24_misconfigured_pvc 36.8s 6 14 $0.2466
✅ 43_current_datetime_from_prompt 5.9s 1 — $0.1070
✅ 61_exact_match_counting 20.2s 4 4 $0.1656
Total 35.4s avg 5.3 avg 11.4 avg $2.1168
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-prometheus-downsampling-jZZnp -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-prometheus-downsampling-jZZnp -f markers=regression -f filter=

@github-actions

github-actions Bot commented Feb 15, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for dd1293b (built in 42s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:dd1293b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:dd1293b me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:dd1293b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:dd1293b

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:dd1293b

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:dd1293b

@coderabbitai

coderabbitai Bot commented Feb 15, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉


Walkthrough

Default MAX_GRAPH_POINTS increased from 100 to 300 and a new MAX_GRAPH_POINTS_HARD_LIMIT was added; Prometheus step-calculation now supports a max_points override clamped by a hard limit and prompt rendering switched to load_and_render_prompt; prompts and tests updated accordingly. (49 words)

Changes

Cohort / File(s) Summary
Configuration & Constants
holmes/common/env_vars.py
Default MAX_GRAPH_POINTS changed from 100 to 300 and new MAX_GRAPH_POINTS_HARD_LIMIT added (defaults to env var or 2× MAX_GRAPH_POINTS).
Prometheus Query Logic & Prompts
holmes/plugins/toolsets/prometheus/prometheus.py, holmes/plugins/toolsets/prometheus/prometheus_instructions.jinja2
Imported MAX_GRAPH_POINTS_HARD_LIMIT and load_and_render_prompt; adjust_step_for_max_points now accepts/clamps max_points overrides to a hard limit, defaults step calculation to target max_points, updated logging/docstrings; prompt template expanded with guidance on resolution, max_points, and hard_max_points.
Tests — Grafana Tempo
tests/plugins/toolsets/grafana/test_grafana_tempo_tools.py
Several test methods updated to accept monkeypatch and monkeypatch MAX_GRAPH_POINTS to a lower value to make step auto-calculation deterministic.
Tests — Prometheus Unit
tests/plugins/toolsets/test_prometheus_unit.py
Added extensive unit tests for adjust_step_for_max_points covering default behavior, overrides (above, below, invalid), hard-limit capping, and interaction with explicit step (tests use monkeypatch).

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

Suggested reviewers

  • moshemorad
  • nherment
🚥 Pre-merge checks | ✅ 1 | ❌ 3

❌ Failed checks (3 warnings)

Check name Status Explanation Resolution
Title check ⚠️ Warning The PR title claims 'Increase default max_points to 500' but the actual implementation sets it to 300, creating a direct mismatch between title and code. Update the PR title to reflect the actual default value of 300, e.g., 'Increase default max_points to 300 and allow LLM override up to 600'.
Docstring Coverage ⚠️ Warning Docstring coverage is 73.33% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
Merge Conflict Detection ⚠️ Warning ❌ Merge conflicts detected (14 files):

⚔️ docs/data-sources/builtin-toolsets/.nav.yml (content)
⚔️ docs/data-sources/builtin-toolsets/index.md (content)
⚔️ docs/reference/.nav.yml (content)
⚔️ holmes/common/env_vars.py (content)
⚔️ holmes/core/openai_formatting.py (content)
⚔️ holmes/core/tools.py (content)
⚔️ holmes/plugins/toolsets/__init__.py (content)
⚔️ holmes/plugins/toolsets/prometheus/prometheus.py (content)
⚔️ holmes/plugins/toolsets/prometheus/prometheus_instructions.jinja2 (content)
⚔️ mkdocs.yml (content)
⚔️ tests/llm/utils/mock_toolset.py (content)
⚔️ tests/plugins/toolsets/grafana/test_grafana_tempo_tools.py (content)
⚔️ tests/plugins/toolsets/test_prometheus_unit.py (content)
⚔️ tests/test_mcp_toolset.py (content)

These conflicts must be resolved before merging into master.
Resolve conflicts locally and push changes to this branch.
✅ Passed checks (1 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@aantn

aantn commented Feb 15, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
marker: prometheus

@github-actions

github-actions Bot commented Feb 15, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 9.96s 9.95s +0.1%
Warm Mean 4.66s 4.72s -1.4%
Warm Min 4.60s 4.70s
Warm Max 4.76s 4.75s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 24.59s 28.04s -12.3%
Warm Mean 7.34s 7.10s +3.4%
Warm Min 6.81s 6.84s
Warm Max 7.88s 7.25s

PR: dd1293bf | Master: c1d0a210 | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
holmes/plugins/toolsets/prometheus/prometheus.py (1)

481-498: ⚠️ Potential issue | 🟡 Minor

Potential ZeroDivisionError if MAX_GRAPH_POINTS is set to 0.

If someone sets the environment variable MAX_GRAPH_POINTS=0, both hard_limit and max_points will be 0, causing a ZeroDivisionError on line 507 (time_range_seconds / max_points) and line 516. Consider adding a guard at the top of the function.

🛡️ Proposed fix
     hard_limit = MAX_GRAPH_POINTS * 2
+    if hard_limit < 2:
+        hard_limit = 1000  # sensible fallback
 
     # Use override if provided and valid, otherwise use default
-    max_points = MAX_GRAPH_POINTS
+    max_points = MAX_GRAPH_POINTS if MAX_GRAPH_POINTS >= 1 else 500
🧹 Nitpick comments (3)
holmes/plugins/toolsets/prometheus/prometheus_instructions.jinja2 (1)

31-35: Hardcoded data point values (500, 1000) may drift from MAX_GRAPH_POINTS.

The instructions reference specific numbers (500 default, 1000 max) that are actually derived from the MAX_GRAPH_POINTS environment variable. If an operator customizes MAX_GRAPH_POINTS, these instructions will be inaccurate. Consider templating these values if the variable is accessible in the template context.

tests/plugins/toolsets/test_prometheus_unit.py (1)

56-59: Repeated local imports of prom_module could be hoisted to module level.

The import holmes.plugins.toolsets.prometheus.prometheus as prom_module is duplicated in every test function/method. Moving it to the top of the file would comply with the coding guideline and reduce repetition. The monkeypatch will still work correctly since it patches the module's attribute.

As per coding guidelines, "ALWAYS place Python imports at the top of the file, not inside functions or methods".

Also applies to: 103-105, 116-118, 131-133, 148-150, 163-165, 179-181

holmes/plugins/toolsets/prometheus/prometheus.py (1)

457-522: Two adjust_step_for_max_points implementations with different signatures exist and are intentionally used by different toolsets.

The utils.py version (time_range_seconds: int, max_points: int, step: Optional[int]) is used by Tempo (grafana), while the Prometheus version (start_timestamp: str, end_timestamp: str, step: Optional[float], max_points_override: Optional[float]) is used by Prometheus itself. The Prometheus implementation is more sophisticated, handling timestamp parsing and max_points override validation with hard limits and logging. This differentiation is appropriate for their respective toolsets, but worth noting for future maintenance if changes are made to either implementation.

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ⚠️ Completed with 2 failures


🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Branch claude/fix-prometheus-downsampling-jZZnp
Model opus-4.5
Markers prometheus
Iterations 1
Duration 12m 53s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 12/16 test cases were successful, 1 regressions, 2 skipped, 1 mock failures
  • investigate: 4/5 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost
🔧 111_tool_hallucination 18.4s 3 2 —
✅ 124_checkout_latency_prometheus[0] 64.2s 8 22 $0.5656
✅ 151_disabled_toolsets_fallback_only 58.6s 8 23 $0.4328
✅ 159_prometheus_high_cardinality_cpu[0] 26.0s 3 4 $0.3791
✅ 159_prometheus_high_cardinality_cpu[1] 26.6s 4 5 $0.3116
❌ 159_prometheus_high_cardinality_cpu[2] 17.5s 3 3 $0.1863
✅ 160_electricity_market_bidding_bug[0] 50.9s 7 12 $0.5494
✅ 160a_cpu_per_namespace_graph 19.6s 3 3 $0.2533
➖ 160b_cpu_per_namespace_graph_with_prom_truncation — — — —
➖ 160c_cpu_per_namespace_graph_with_global_truncation — — — —
✅ 161_bidding_version_performance[0] 87.0s 8 27 $0.7356
✅ 211_prometheus_alerting_rules 43.6s 8 10 $0.4023
✅ 30_basic_promql_graph_cluster_memory 23.4s 4 5 $0.2196
✅ 32_basic_promql_graph_pod_cpu 24.0s 4 5 $0.2853
✅ 33_cpu_metrics_discovery 20.1s 3 3 $0.1913
✅ 34_memory_graph 21.5s 3 3 $0.2993
✅ 03_cpu_throttling 46.8s 5 10 —
✅ 08_memory_pressure 61.2s 9 16 —
✅ 10_KubeDeploymentReplicasMismatch 49.0s 5 10 —
❌ 14_tempo 73.0s 7 34 —
✅ 17_investigate_correct_date 70.9s 10 20 —
Total 42.2s avg 5.5 avg 11.4 avg $4.8118

⚠️ 2 Failures Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-prometheus-downsampling-jZZnp -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-prometheus-downsampling-jZZnp -f markers=regression -f filter=

arikalon1
arikalon1 previously approved these changes Feb 15, 2026
@aantn

aantn commented Feb 15, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
marker: prometheus
branch: master

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ⚠️ Completed with 2 failures


🧪 Manual Eval Results

Parameter Value
Triggered via /eval on branch master
Branch master
Model opus-4.5
Markers prometheus
Iterations 1
Duration 13m 51s
Workflow View logs | Rerun

Results of HolmesGPT evals (branch: master)

  • ask_holmes: 12/16 test cases were successful, 1 regressions, 2 skipped, 1 mock failures
  • investigate: 4/5 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost
🔧 111_tool_hallucination 19.2s 3 2 —
✅ 124_checkout_latency_prometheus[0] 71.1s 7 27 $0.4428
✅ 151_disabled_toolsets_fallback_only 45.7s 6 15 $0.3251
✅ 159_prometheus_high_cardinality_cpu[0] 29.4s 4 4 $0.2503
✅ 159_prometheus_high_cardinality_cpu[1] 23.3s 4 5 $0.2378
❌ 159_prometheus_high_cardinality_cpu[2] 20.8s 3 4 $0.2199
✅ 160_electricity_market_bidding_bug[0] 80.2s 12 21 $0.4472
✅ 160a_cpu_per_namespace_graph 18.1s 3 3 $0.1865
➖ 160b_cpu_per_namespace_graph_with_prom_truncation — — — —
➖ 160c_cpu_per_namespace_graph_with_global_truncation — — — —
✅ 161_bidding_version_performance[0] 81.9s 7 26 $0.5360
✅ 211_prometheus_alerting_rules 36.2s 6 8 $0.3441
✅ 30_basic_promql_graph_cluster_memory 23.8s 4 5 $0.2052
✅ 32_basic_promql_graph_pod_cpu 22.3s 4 5 $0.2134
✅ 33_cpu_metrics_discovery 21.5s 3 3 $0.1907
✅ 34_memory_graph 18.3s 3 3 $0.2254
✅ 03_cpu_throttling 43.8s 4 11 —
✅ 08_memory_pressure 61.6s 8 16 —
✅ 10_KubeDeploymentReplicasMismatch 46.6s 5 11 —
❌ 14_tempo 93.2s 11 36 —
✅ 17_investigate_correct_date 60.2s 7 17 —
Total 43.0s avg 5.5 avg 11.7 avg $3.8244

⚠️ 2 Failures Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref master -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref master -f markers=regression -f filter=

aantn and others added 3 commits February 15, 2026 17:59
… numbers

- Template now uses {{ default_max_points }} and {{ hard_max_points }} from
  the MAX_GRAPH_POINTS env var so instructions stay accurate if the default changes
- Updated max_points tool description: removed "overview graphs" wording,
  clarified decrease is to avoid hitting the data point limit

https://claude.ai/code/session_0174s8zNS3peCtHc2tg9iJ8y
Signed-off-by: Claude <noreply@anthropic.com>
Co-locate with MAX_GRAPH_POINTS for better locality. Defaults to
MAX_GRAPH_POINTS * 2 but can now be configured independently.

https://claude.ai/code/session_0174s8zNS3peCtHc2tg9iJ8y
Signed-off-by: Claude <noreply@anthropic.com>
@aantn

aantn commented Feb 15, 2026

Copy link
Copy Markdown
Collaborator Author

/eval
marker: prometheus

@github-actions

Copy link
Copy Markdown
Contributor

@aantn Your eval run has finished. ⚠️ Completed with 2 failures


🧪 Manual Eval Results

Parameter Value
Triggered via /eval comment
Branch claude/fix-prometheus-downsampling-jZZnp
Model opus-4.5
Markers prometheus
Iterations 1
Duration 13m 45s
Workflow View logs | Rerun

Results of HolmesGPT evals

  • ask_holmes: 12/16 test cases were successful, 1 regressions, 2 skipped, 1 mock failures
  • investigate: 4/5 test cases were successful, 1 regressions
Status Test case Time Turns Tools Cost
🔧 111_tool_hallucination 20.1s 3 2 —
✅ 124_checkout_latency_prometheus[0] 81.1s 9 27 $0.6680
✅ 151_disabled_toolsets_fallback_only 73.2s 10 26 $0.4519
✅ 159_prometheus_high_cardinality_cpu[0] 30.6s 4 5 $0.2838
✅ 159_prometheus_high_cardinality_cpu[1] 26.6s 4 5 $0.3201
❌ 159_prometheus_high_cardinality_cpu[2] 16.0s 3 3 $0.1825
✅ 160_electricity_market_bidding_bug[0] 54.7s 7 12 $0.5473
✅ 160a_cpu_per_namespace_graph 20.5s 3 3 $0.2604
➖ 160b_cpu_per_namespace_graph_with_prom_truncation — — — —
➖ 160c_cpu_per_namespace_graph_with_global_truncation — — — —
✅ 161_bidding_version_performance[0] 100.9s 13 28 $0.7278
✅ 211_prometheus_alerting_rules 19.0s 4 4 $0.1931
✅ 30_basic_promql_graph_cluster_memory 22.4s 4 5 $0.2253
✅ 32_basic_promql_graph_pod_cpu 24.4s 4 5 $0.2548
✅ 33_cpu_metrics_discovery 20.2s 3 3 $0.1909
✅ 34_memory_graph 19.9s 3 3 $0.2987
✅ 03_cpu_throttling 50.2s 6 10 —
✅ 08_memory_pressure 52.7s 8 14 —
✅ 10_KubeDeploymentReplicasMismatch 48.7s 5 9 —
❌ 14_tempo 88.7s 10 35 —
✅ 17_investigate_correct_date 62.8s 7 17 —
Total 43.8s avg 5.8 avg 11.4 avg $4.6046

⚠️ 2 Failures Detected

📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-prometheus-downsampling-jZZnp -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-prometheus-downsampling-jZZnp -f markers=regression -f filter=

claude and others added 2 commits February 16, 2026 11:26
Also fix tests to monkeypatch MAX_GRAPH_POINTS_HARD_LIMIT alongside
MAX_GRAPH_POINTS so they remain independent of the default values.

https://claude.ai/code/session_0174s8zNS3peCtHc2tg9iJ8y
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn enabled auto-merge (squash) February 16, 2026 11:47
@aantn
aantn merged commit 9d82775 into master Feb 16, 2026
19 of 20 checks passed
@aantn
aantn deleted the claude/fix-prometheus-downsampling-jZZnp branch February 16, 2026 11:52
moshemorad pushed a commit that referenced this pull request Feb 22, 2026
…1566)

## Summary
This PR improves Prometheus query resolution by increasing the default
maximum data points from 100 to 500, and allowing the LLM to request up
to 1000 points (2x the default) for high-resolution analysis of
low-cardinality queries. The change includes updated logic,
comprehensive tests, and improved documentation.

## Key Changes

- **Increased default MAX_GRAPH_POINTS**: Changed from 100 to 500 to
provide better default resolution for time series queries
- **Implemented hard limit for max_points override**: LLM can now
request up to 2x MAX_GRAPH_POINTS (1000 points) for higher resolution,
with a hard cap to prevent excessive data retrieval
- **Updated step calculation logic**: When no step is provided, the
default now targets the configured max_points instead of a fixed
60-point target
- **Enhanced parameter descriptions**: Updated tool parameter
documentation to clarify:
- `step` parameter now explains the relationship between step size and
data points
- `max_points` parameter now includes guidance on when to increase
(low-cardinality queries) vs decrease (overview graphs, high-cardinality
queries)
- **Added comprehensive test coverage**: New test class
`TestMaxPointsOverride` with 5 test cases covering:
  - Override above default (allowed)
  - Override capped at hard limit
  - Override below default (allowed)
  - Invalid override fallback behavior
  - Interaction between explicit step and max_points override
- **Updated Prometheus instructions**: Added new section on query
resolution and data points with guidance on:
  - Default 500-point limit per series
  - When to increase max_points for spike/anomaly detection
  - Cardinality considerations
  - Alternative approach of narrowing time range for higher resolution

## Implementation Details

- The hard limit is calculated as `MAX_GRAPH_POINTS * 2` to allow
flexibility while maintaining safety
- Invalid overrides (< 1) fall back to the default MAX_GRAPH_POINTS
- When both step and max_points_override are provided, the step is
adjusted if it would exceed the max_points target
- Logging messages updated to reflect the new behavior and hard limits
- All existing tests updated to use monkeypatch for MAX_GRAPH_POINTS to
ensure test isolation

https://claude.ai/code/session_0174s8zNS3peCtHc2tg9iJ8y

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Increased default maximum graph data points from 100 to 300.
* Added a configurable hard limit (default 2× the max) that clamps and
warns on excessive max-points.

* **Improvements**
* Step-size calculation now targets the configured data-point count;
clearer messaging about limits and overrides.
  * Expanded parameter guidance for resolution vs. data points.

* **Documentation**
* Added query-resolution and data-point guidance with examples for
Prometheus tools.

* **Tests**
* Added and updated tests covering step adjustments, overrides, and
hard-limit behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Mohse Morad <moshemorad12340@gmail.com>
moshemorad pushed a commit that referenced this pull request Feb 22, 2026
…1566)

## Summary
This PR improves Prometheus query resolution by increasing the default
maximum data points from 100 to 500, and allowing the LLM to request up
to 1000 points (2x the default) for high-resolution analysis of
low-cardinality queries. The change includes updated logic,
comprehensive tests, and improved documentation.

## Key Changes

- **Increased default MAX_GRAPH_POINTS**: Changed from 100 to 500 to
provide better default resolution for time series queries
- **Implemented hard limit for max_points override**: LLM can now
request up to 2x MAX_GRAPH_POINTS (1000 points) for higher resolution,
with a hard cap to prevent excessive data retrieval
- **Updated step calculation logic**: When no step is provided, the
default now targets the configured max_points instead of a fixed
60-point target
- **Enhanced parameter descriptions**: Updated tool parameter
documentation to clarify:
- `step` parameter now explains the relationship between step size and
data points
- `max_points` parameter now includes guidance on when to increase
(low-cardinality queries) vs decrease (overview graphs, high-cardinality
queries)
- **Added comprehensive test coverage**: New test class
`TestMaxPointsOverride` with 5 test cases covering:
  - Override above default (allowed)
  - Override capped at hard limit
  - Override below default (allowed)
  - Invalid override fallback behavior
  - Interaction between explicit step and max_points override
- **Updated Prometheus instructions**: Added new section on query
resolution and data points with guidance on:
  - Default 500-point limit per series
  - When to increase max_points for spike/anomaly detection
  - Cardinality considerations
  - Alternative approach of narrowing time range for higher resolution

## Implementation Details

- The hard limit is calculated as `MAX_GRAPH_POINTS * 2` to allow
flexibility while maintaining safety
- Invalid overrides (< 1) fall back to the default MAX_GRAPH_POINTS
- When both step and max_points_override are provided, the step is
adjusted if it would exceed the max_points target
- Logging messages updated to reflect the new behavior and hard limits
- All existing tests updated to use monkeypatch for MAX_GRAPH_POINTS to
ensure test isolation

https://claude.ai/code/session_0174s8zNS3peCtHc2tg9iJ8y

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Increased default maximum graph data points from 100 to 300.
* Added a configurable hard limit (default 2× the max) that clamps and
warns on excessive max-points.

* **Improvements**
* Step-size calculation now targets the configured data-point count;
clearer messaging about limits and overrides.
  * Expanded parameter guidance for resolution vs. data points.

* **Documentation**
* Added query-resolution and data-point guidance with examples for
Prometheus tools.

* **Tests**
* Added and updated tests covering step adjustments, overrides, and
hard-limit behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Mohse Morad <moshemorad12340@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants