Skip to content

[ROB-3883] Loki returns raw response on parse failure - #2035

Merged
Avi-Robusta merged 8 commits into
masterfrom
0.26.0-loki-debug
May 17, 2026
Merged

Avi-Robusta merged 8 commits into
masterfrom
0.26.0-loki-debug

Conversation

@Avi-Robusta

@Avi-Robusta Avi-Robusta commented May 14, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • Bug Fixes

    • Improved Loki query error handling: connection failures and response-processing failures are now reported separately with clearer, more diagnostic messages (including raw response snippets and content details).
  • Tests

    • Added tests covering malformed Loki responses and refined test skipping for Grafana/Loki health checks to ensure robust behavior.

Review Change Stack

Signed-off-by: avi@robusta.dev <avi@robusta.dev>
When Loki returns a malformed JSON body, the toolset previously surfaced
only the parse error message (e.g. "Expecting ',' delimiter: line 1
column 2811"), giving no insight into what was actually returned.

This is particularly painful when a proxy (e.g. nginx with
body_filter_by_lua scrubbing) is corrupting the response in transit —
the error position is on the corrupted body, not the original.

Catch the ValueError from response.json() and re-raise with the raw
response text, length, and content-type included so the LLM (and
operator) can see the actual bytes that came back.

Signed-off-by: avi@robusta.dev <avi@robusta.dev>
The previous commit only included the raw response when JSON parsing
failed. Widen the net to any exception raised after the HTTP response
is in hand (e.g. unexpected response shape in parse_loki_response,
KeyError on missing fields), so the raw bytes are always available for
debugging — not just for JSONDecodeError.

The split into two try/except blocks keeps the RequestException path
(no response object yet) separate from the response-handling path
(response.text is available).

Signed-off-by: avi@robusta.dev <avi@robusta.dev>
This reverts commit ef07978.

Signed-off-by: avi@robusta.dev <avi@robusta.dev>
Signed-off-by: avi@robusta.dev <avi@robusta.dev>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@github-actions

github-actions Bot commented May 14, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 #3 · Run @ __fbd82a1__ (#25988798786) — May 17, 10:59 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit fbd82a1 on branch 0.26.0-loki-debug

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Src
✅ 09_crashpod 39.9s 6 12 $0.3013 134,314 131,803 24,814 2,511 1,081 106,413 25,390 310 — src
✅ 101_loki_historical_logs_pod_deleted 81.5s 9 20 $0.4593 225,346 220,282 29,458 5,064 916 188,288 31,994 1,023 — src
✅ 112_find_pvcs_by_uuid 19.5s 3 4 $0.2028 60,750 59,565 21,733 1,185 595 37,828 21,737 287 — src
✅ 12_job_crashing 45.0s 6 15 $0.3555 154,549 151,704 29,019 2,845 898 120,617 31,087 171 — src
✅ 176_network_policy_blocking_traffic_no_skills 42.2s 5 13 $0.3078 113,035 110,180 25,940 2,855 780 83,315 26,865 451 — src
✅ 227_count_configmaps_per_namespace[0] 20.0s 4 9 $0.2068 76,261 75,134 20,540 1,127 593 53,667 21,467 53 — src
✅ 243_pod_names_contain_service 31.0s 4 8 $0.2341 78,935 77,113 21,694 1,822 877 54,527 22,586 246 — src
✅ 24_misconfigured_pvc 42.0s 5 14 $0.2942 107,410 104,621 23,983 2,789 1,051 78,971 25,650 279 — src
✅ 43_current_datetime_from_prompt 4.0s 1 — $0.1183 16,900 16,798 16,798 102 102 0 16,798 61 — src
✅ 51_logs_summarize_errors 24.4s 4 5 $0.2049 77,098 75,996 20,999 1,102 392 54,992 21,004 73 — src
✅ 61_exact_match_counting 10.8s 3 3 $0.1510 52,420 52,058 17,771 362 215 34,283 17,775 31 — src
Total 32.7s avg 4.5 avg 10.3 avg $2.8361 1,097,018 1,075,254 29,458 21,764 1,081 812,901 262,353 2,985 —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 147 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (1h ago) Δ vs master benchmark (14d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 39.9s 37.8s ±0% 32.5s ↑23%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 81.5s 79.5s ±0% 51.7s ↑58%
112_find_pvcs_by_uuid (opus-4.6) 📄 19.5s 20.7s ±0% 18.1s ±0%
12_job_crashing (opus-4.6) 📄 45.0s 41.9s ±0% 42.4s ±0%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 42.2s 46.8s ±0% 35.7s ↑18%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 20.0s 19.9s ±0% 19.7s ±0%
243_pod_names_contain_service (opus-4.6) 📄 31.0s 34.2s ±0% 27.4s ↑13%
24_misconfigured_pvc (opus-4.6) 📄 42.0s 37.6s ↑12% 35.8s ↑17%
43_current_datetime_from_prompt (opus-4.6) 📄 4.0s 3.4s ↑17% — —
51_logs_summarize_errors (opus-4.6) 📄 24.4s 22.4s ±0% 23.1s ±0%
61_exact_match_counting (opus-4.6) 📄 10.8s 10.3s ±0% 10.8s ±0%
Total (all, n=11) 32.7s 32.2s — 29.7s —
Comparable (m=11, b=10) 32.7s 32.2s ±0% 29.7s ↑20%

Cost comparison:

Test case This branch master (1h ago) Δ vs master benchmark (14d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.3013 $0.2829 ±0% $0.2616 ↑15%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.4593 $0.4365 ±0% $0.3371 ↑36%
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.2028 $0.2049 ±0% $0.2014 ±0%
12_job_crashing (opus-4.6) 📄 $0.3555 $0.2999 ↑19% $0.3076 ↑16%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.3078 $0.3208 ±0% $0.2914 ±0%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.2068 $0.2038 ±0% $0.2059 ±0%
243_pod_names_contain_service (opus-4.6) 📄 $0.2341 $0.2566 ±0% $0.2280 ±0%
24_misconfigured_pvc (opus-4.6) 📄 $0.2942 $0.2858 ±0% $0.2831 ±0%
43_current_datetime_from_prompt (opus-4.6) 📄 $0.1183 $0.1182 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 $0.2049 $0.2025 ±0% $0.2072 ±0%
61_exact_match_counting (opus-4.6) 📄 $0.1510 $0.1511 ±0% $0.1522 ±0%
Total (all, n=11) $0.2578 $0.2512 — $0.2475 —
Comparable (m=11, b=10) $0.2578 $0.2512 ±0% $0.2475 ±0%

Total tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (14d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 134,314 108,034 ↑24% 103,497 ↑30%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 225,346 195,566 ↑15% 138,670 ↑63%
112_find_pvcs_by_uuid (opus-4.6) 📄 60,750 60,781 ±0% 61,169 ±0%
12_job_crashing (opus-4.6) 📄 154,549 133,114 ↑16% 133,893 ↑15%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 113,035 114,683 ±0% 111,145 ±0%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 76,261 76,264 ±0% 76,945 ±0%
243_pod_names_contain_service (opus-4.6) 📄 78,935 101,645 ↓22% 79,525 ±0%
24_misconfigured_pvc (opus-4.6) 📄 107,410 106,626 ±0% 108,047 ±0%
43_current_datetime_from_prompt (opus-4.6) 📄 16,900 16,898 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 77,098 76,361 ±0% 77,707 ±0%
61_exact_match_counting (opus-4.6) 📄 52,420 52,424 ±0% 52,942 ±0%
Total (all, n=11) 99,729 94,763 — 94,354 —
Comparable (m=11, b=10) 99,729 94,763 ±0% 94,354 ↑14%

Cached tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (14d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 106,413 80,131 ↑33% 77,391 ↑38%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 188,288 159,018 ↑18% 106,565 ↑77%
112_find_pvcs_by_uuid (opus-4.6) 📄 37,828 37,590 ±0% 38,002 ±0%
12_job_crashing (opus-4.6) 📄 120,617 105,176 ↑15% 104,761 ↑15%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 83,315 84,084 ±0% 81,519 ±0%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 53,667 54,279 ±0% 54,503 ±0%
243_pod_names_contain_service (opus-4.6) 📄 54,527 76,327 ↓29% 55,513 ±0%
24_misconfigured_pvc (opus-4.6) 📄 78,971 78,821 ±0% 80,270 ±0%
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 54,992 54,618 ±0% 55,443 ±0%
61_exact_match_counting (opus-4.6) 📄 34,283 34,283 ±0% 34,632 ±0%
Total (all, n=11) 73,900 69,484 — 68,860 —
Comparable (m=10, b=10) 81,290 76,433 ±0% 68,860 ↑18%

Turns comparison:

Test case This branch master (1h ago) Δ vs master benchmark (14d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 6 5 ↑20% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 9 8 ↑12% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 3 3 ±0% — —
12_job_crashing (opus-4.6) 📄 6 6 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 5 5 ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 4 4 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 4 5 ↓20% — —
24_misconfigured_pvc (opus-4.6) 📄 5 5 ±0% — —
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 4 4 ±0% — —
61_exact_match_counting (opus-4.6) 📄 3 3 ±0% — —
Total (all, n=11) 4.5 4.5 — — —
Comparable (m=11, b=0) 4.5 4.5 ±0% — —

Tool calls comparison:

Test case This branch master (1h ago) Δ vs master benchmark (14d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 12 11 ±0% 10 ↑20%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 20 18 ↑11% 14 ↑43%
112_find_pvcs_by_uuid (opus-4.6) 📄 4 4 ±0% 4 ±0%
12_job_crashing (opus-4.6) 📄 15 13 ↑15% 14 ±0%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 13 13 ±0% 13 ±0%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 9 9 ±0% 9 ±0%
243_pod_names_contain_service (opus-4.6) 📄 8 9 ↓11% 8 ±0%
24_misconfigured_pvc (opus-4.6) 📄 14 15 ±0% 14 ±0%
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 5 5 ±0% 5 ±0%
61_exact_match_counting (opus-4.6) 📄 3 3 ±0% 3 ±0%
Total (all, n=11) 9.4 10.0 — 9.4 —
Comparable (m=10, b=10) 10.3 10.0 ±0% 9.4 ±0%

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
⚠️ 2 older runs truncated

Older runs were omitted to stay under GitHub's 64KB comment size limit.


✅ Results of HolmesGPT evals

Automatically triggered by commit cb1bd9b on branch 0.26.0-loki-debug

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Src
✅ 09_crashpod 35.1s 5 10 $0.2675 104,901 102,822 23,604 2,079 652 78,313 24,509 196 — src
✅ 101_loki_historical_logs_pod_deleted 76.4s 7 18 $0.4440 178,284 173,245 31,057 5,039 922 139,415 33,830 925 — src
✅ 112_find_pvcs_by_uuid 21.3s 3 5 $0.2141 61,904 60,533 22,233 1,371 777 37,800 22,733 264 — src
✅ 12_job_crashing 38.3s 6 12 $0.2989 131,831 129,597 24,133 2,234 776 102,720 26,877 170 — src
✅ 176_network_policy_blocking_traffic_no_skills 54.5s 7 17 $0.3753 166,953 163,312 27,902 3,641 842 133,614 29,698 667 — src
✅ 227_count_configmaps_per_namespace[0] 19.0s 4 9 $0.2085 76,280 75,153 20,551 1,127 593 53,356 21,797 53 — src
✅ 243_pod_names_contain_service 30.0s 4 8 $0.2330 78,925 77,083 21,683 1,842 872 54,833 22,250 265 — src
✅ 24_misconfigured_pvc 38.3s 5 14 $0.2862 106,562 103,984 23,608 2,578 1,057 78,591 25,393 222 — src
✅ 43_current_datetime_from_prompt 3.9s 1 — $0.1182 16,898 16,798 16,798 100 100 0 16,798 60 — src
✅ 51_logs_summarize_errors 21.2s 4 5 $0.2044 76,671 75,523 20,757 1,148 368 54,761 20,762 84 — src
✅ 61_exact_match_counting 10.3s 3 3 $0.1511 52,433 52,071 17,780 362 215 34,287 17,784 31 — src
Total 31.7s avg 4.5 avg 10.1 avg $2.8013 1,051,642 1,030,121 31,057 21,521 1,057 767,690 262,431 2,937 —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 11 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 147 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (1h ago) Δ vs master benchmark (14d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 35.1s 37.8s ±0% 32.5s ±0%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 76.4s 79.5s ±0% 51.7s ↑48%
112_find_pvcs_by_uuid (opus-4.6) 📄 21.3s 20.7s ±0% 18.1s ↑18%
12_job_crashing (opus-4.6) 📄 38.3s 41.9s ±0% 42.4s ±0%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 54.5s 46.8s ↑16% 35.7s ↑53%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 19.0s 19.9s ±0% 19.7s ±0%
243_pod_names_contain_service (opus-4.6) 📄 30.0s 34.2s ↓12% 27.4s ±0%
24_misconfigured_pvc (opus-4.6) 📄 38.3s 37.6s ±0% 35.8s ±0%
43_current_datetime_from_prompt (opus-4.6) 📄 3.9s 3.4s ↑16% — —
51_logs_summarize_errors (opus-4.6) 📄 21.2s 22.4s ±0% 23.1s ±0%
61_exact_match_counting (opus-4.6) 📄 10.3s 10.3s ±0% 10.8s ±0%
Total (all, n=11) 31.7s 32.2s — 29.7s —
Comparable (m=11, b=10) 31.7s 32.2s ±0% 29.7s ↑16%

Cost comparison:

Test case This branch master (1h ago) Δ vs master benchmark (14d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.2675 $0.2829 ±0% $0.2616 ±0%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.4440 $0.4365 ±0% $0.3371 ↑32%
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.2141 $0.2049 ±0% $0.2014 ±0%
12_job_crashing (opus-4.6) 📄 $0.2989 $0.2999 ±0% $0.3076 ±0%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.3753 $0.3208 ↑17% $0.2914 ↑29%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.2085 $0.2038 ±0% $0.2059 ±0%
243_pod_names_contain_service (opus-4.6) 📄 $0.2330 $0.2566 ±0% $0.2280 ±0%
24_misconfigured_pvc (opus-4.6) 📄 $0.2862 $0.2858 ±0% $0.2831 ±0%
43_current_datetime_from_prompt (opus-4.6) 📄 $0.1182 $0.1182 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 $0.2044 $0.2025 ±0% $0.2072 ±0%
61_exact_match_counting (opus-4.6) 📄 $0.1511 $0.1511 ±0% $0.1522 ±0%
Total (all, n=11) $0.2547 $0.2512 — $0.2475 —
Comparable (m=11, b=10) $0.2547 $0.2512 ±0% $0.2475 ±0%

Total tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (14d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 104,901 108,034 ±0% 103,497 ±0%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 178,284 195,566 ±0% 138,670 ↑29%
112_find_pvcs_by_uuid (opus-4.6) 📄 61,904 60,781 ±0% 61,169 ±0%
12_job_crashing (opus-4.6) 📄 131,831 133,114 ±0% 133,893 ±0%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 166,953 114,683 ↑46% 111,145 ↑50%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 76,280 76,264 ±0% 76,945 ±0%
243_pod_names_contain_service (opus-4.6) 📄 78,925 101,645 ↓22% 79,525 ±0%
24_misconfigured_pvc (opus-4.6) 📄 106,562 106,626 ±0% 108,047 ±0%
43_current_datetime_from_prompt (opus-4.6) 📄 16,898 16,898 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 76,671 76,361 ±0% 77,707 ±0%
61_exact_match_counting (opus-4.6) 📄 52,433 52,424 ±0% 52,942 ±0%
Total (all, n=11) 95,604 94,763 — 94,354 —
Comparable (m=11, b=10) 95,604 94,763 ±0% 94,354 ±0%

Cached tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (14d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 78,313 80,131 ±0% 77,391 ±0%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 139,415 159,018 ↓12% 106,565 ↑31%
112_find_pvcs_by_uuid (opus-4.6) 📄 37,800 37,590 ±0% 38,002 ±0%
12_job_crashing (opus-4.6) 📄 102,720 105,176 ±0% 104,761 ±0%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 133,614 84,084 ↑59% 81,519 ↑64%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 53,356 54,279 ±0% 54,503 ±0%
243_pod_names_contain_service (opus-4.6) 📄 54,833 76,327 ↓28% 55,513 ±0%
24_misconfigured_pvc (opus-4.6) 📄 78,591 78,821 ±0% 80,270 ±0%
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 54,761 54,618 ±0% 55,443 ±0%
61_exact_match_counting (opus-4.6) 📄 34,287 34,283 ±0% 34,632 ±0%
Total (all, n=11) 69,790 69,484 — 68,860 —
Comparable (m=10, b=10) 76,769 76,433 ±0% 68,860 ↑11%

Turns comparison:

Test case This branch master (1h ago) Δ vs master benchmark (14d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 5 5 ±0% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 7 8 ↓12% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 3 3 ±0% — —
12_job_crashing (opus-4.6) 📄 6 6 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 7 5 ↑40% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 4 4 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 4 5 ↓20% — —
24_misconfigured_pvc (opus-4.6) 📄 5 5 ±0% — —
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 4 4 ±0% — —
61_exact_match_counting (opus-4.6) 📄 3 3 ±0% — —
Total (all, n=11) 4.5 4.5 — — —
Comparable (m=11, b=0) 4.5 4.5 ±0% — —

Tool calls comparison:

Test case This branch master (1h ago) Δ vs master benchmark (14d ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 10 11 ±0% 10 ±0%
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 18 18 ±0% 14 ↑29%
112_find_pvcs_by_uuid (opus-4.6) 📄 5 4 ↑25% 4 ↑25%
12_job_crashing (opus-4.6) 📄 12 13 ±0% 14 ↓14%
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 17 13 ↑31% 13 ↑31%
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 9 9 ±0% 9 ±0%
243_pod_names_contain_service (opus-4.6) 📄 8 9 ↓11% 8 ±0%
24_misconfigured_pvc (opus-4.6) 📄 14 15 ±0% 14 ±0%
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 5 5 ±0% 5 ±0%
61_exact_match_counting (opus-4.6) 📄 3 3 ±0% 3 ±0%
Total (all, n=11) 9.2 10.0 — 9.4 —
Comparable (m=10, b=10) 10.1 10.0 ±0% 9.4 ±0%

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref 0.26.0-loki-debug -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref 0.26.0-loki-debug -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented May 14, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

This pull request refactors Loki query error handling to separate request failures (caught and re-raised immediately) from response JSON/processing failures (caught separately and raised with raw body, length, and Content-Type). Tests are updated to use per-test skipping and include a new malformed-JSON assertion.

Changes

Loki Query Error Handling and Tests

Layer / File(s) Summary
Two-phase exception handling for Loki queries
holmes/plugins/toolsets/grafana/loki_api.py
Request exceptions are caught immediately after the HTTP request and re-raised as "Failed to query Loki logs". JSON/response processing failures are caught separately and raised as "Failed to process Loki response" with diagnostics: raw body, character length, and Content-Type.
Per-test Grafana skip marker and malformed-JSON test
tests/plugins/toolsets/grafana/test_grafana.py
Replaces module-level pytestmark skip with needs_grafana = pytest.mark.skipif(...) applied to individual tests, and adds test_execute_loki_query_includes_raw_body_on_malformed_json which mocks malformed Loki JSON and asserts the raised exception includes the raw body, content-type, and body length.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Suggested reviewers

  • Sheeproid
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 14.29% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically describes the main change: improving Loki error handling to include raw response data when JSON parsing fails.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented May 14, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for a48eeca3 (built in 1m 1s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:a48eeca3
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:a48eeca3 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:a48eeca3
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:a48eeca3
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:a48eeca3
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:a48eeca3 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:a48eeca3
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:a48eeca3

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:a48eeca3 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:a48eeca3

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:a48eeca3 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:a48eeca3

@netlify

netlify Bot commented May 14, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit cb1bd9b
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a099f85c982010007434b4d
😎 Deploy Preview https://deploy-preview-2035--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 13.57s 13.28s +2.1%
Warm Mean 6.61s 6.22s +6.3%
Warm Min 6.51s 6.20s
Warm Max 6.68s 6.25s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 16.14s 12.79s +26.1%
Warm Mean 8.65s 8.11s +6.6%
Warm Min 8.31s 7.83s
Warm Max 9.35s 8.43s

PR: 41bcb8a5 | Master: e18f575f | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@holmes/plugins/toolsets/grafana/loki_api.py`:
- Around line 77-83: The exception currently embeds the full response.text (raw)
which can leak sensitive or unbounded data; modify the error construction in the
Loki response handling (where raw, response, and e are used) to include a
sanitized, bounded preview: define a MAX_PREVIEW (e.g. 1000) and build a preview
by truncating raw to MAX_PREVIEW and appending an explicit "(truncated)" marker,
strip or replace non-printable/control characters (or collapse long whitespace)
to avoid log injection, and report the preview length alongside
response.headers.get('Content-Type', 'unknown') and the original exception e
instead of inserting the entire raw payload. Ensure the final message still
gives context but never includes the full unbounded raw string.
- Around line 68-69: Replace the bare raise Exception in the requests exception
handler with raise RuntimeError(... ) from e to preserve the original traceback
(keep the existing message string but use RuntimeError and include "from e"),
and in the response processing block that calls response.json() and
parse_loki_response, narrow the except clauses to catch json.JSONDecodeError and
ValueError (for tuple unpacking) instead of a broad except, wrapping them in
RuntimeError(... ) from e; reference the existing
requests.exceptions.RequestException handler and the parse_loki_response call to
locate where to change the exception types and add exception chaining.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 6ff22c00-0c4a-43ae-b2d8-08ed2bfbbdea

📥 Commits

Reviewing files that changed from the base of the PR and between e18f575 and e653f8e.

📒 Files selected for processing (1)
  • holmes/plugins/toolsets/grafana/loki_api.py

Comment thread holmes/plugins/toolsets/grafana/loki_api.py
Comment thread holmes/plugins/toolsets/grafana/loki_api.py
moshemorad
moshemorad previously approved these changes May 17, 2026
Avi-Robusta and others added 2 commits May 17, 2026 13:51
Add a unit test that mocks the Loki query_range endpoint with a malformed
JSON body and verifies the exception raised by execute_loki_query includes
both the JSON parser error and the raw response text (with content-type
and byte count), matching the debug behavior added earlier on this branch.

Also scope the existing module-level "Grafana not running" skip to the
integration tests via a @needs_grafana marker, so the new unit test is
not skipped on machines without a local Grafana.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: avi@robusta.dev <avi@robusta.dev>
@Avi-Robusta
Avi-Robusta enabled auto-merge (squash) May 17, 2026 11:02

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/plugins/toolsets/grafana/test_grafana.py`:
- Around line 14-18: The skip guard currently uses _grafana_skip_reason and
needs_grafana (from check_service_running("Grafana", 3000)) but
test_loki_toolset_direct_health_check calls Loki on port 3100; add a separate
_loki_skip_reason = check_service_running("Loki", 3100) and a needs_loki
pytest.mark.skipif using that reason (or replace the existing needs_grafana
marker on test_loki_toolset_direct_health_check with needs_loki) so the
Loki-specific test is skipped when Loki is unavailable; reference the symbols
_grafana_skip_reason, needs_grafana, _loki_skip_reason, needs_loki,
check_service_running, and test_loki_toolset_direct_health_check to locate and
update the code.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: c13b141b-67ef-4201-8bc9-40da88c37b9c

📥 Commits

Reviewing files that changed from the base of the PR and between e653f8e and cb1bd9b.

📒 Files selected for processing (1)
  • tests/plugins/toolsets/grafana/test_grafana.py

Comment thread tests/plugins/toolsets/grafana/test_grafana.py
@Avi-Robusta
Avi-Robusta merged commit 6e0c0f2 into master May 17, 2026
17 of 22 checks passed
@Avi-Robusta
Avi-Robusta deleted the 0.26.0-loki-debug branch May 17, 2026 11:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants