Skip to content

Add support for image content in MCP tool results - #1803

Merged
aantn merged 34 commits into
masterfrom
claude/fix-eval-233-7keOj
Mar 21, 2026
Merged

aantn merged 34 commits into
masterfrom
claude/fix-eval-233-7keOj

Conversation

@aantn

@aantn aantn commented Mar 17, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR adds support for extracting and handling image content from MCP (Model Context Protocol) tool results, enabling multimodal responses in tool calls.

Key Changes

  • StructuredToolResult model: Added optional images field to store image data with MIME types
  • MCP tool invocation: Updated RemoteMCPTool._invoke_async() to extract ImageContent blocks from MCP responses and populate the images field
  • Result formatting: Modified format_tool_result_data() to return multimodal content (text + image_url blocks in OpenAI vision format) when images are present, while maintaining backward compatibility with plain text responses
  • LLM message generation: Tool call results with images now produce properly formatted multimodal messages with data URIs for image content
  • Test coverage: Added comprehensive tests for image extraction, text-only results, and multimodal message formatting

Implementation Details

  • Images are extracted from MCP CallToolResult.content blocks where type == "image"
  • Image data is stored as a list of dictionaries with data (base64) and mimeType fields
  • When formatting for LLM consumption, images are converted to OpenAI vision format data URIs: data:{mimeType};base64,{data}
  • The images field is None when no images are present, preserving backward compatibility
  • Text content from MCP responses is merged and included alongside images in multimodal results

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt

Summary by CodeRabbit

  • New Features

    • Tool results now support multimodal content: text plus image blocks; string-only results remain supported.
    • Tools can optionally include image data with results.
  • Behavior Changes

    • Message truncation now preserves image blocks and truncates only text within multimodal content.
  • Tests

    • Added tests for image extraction, multimodal formatting, message conversion, and truncation; MCP toolset image handling validated.
  • Chores

    • Test fixtures and environment handling updated; toolset credential sourcing adjusted.

MCP tool results can contain ImageContent blocks (base64 images), but
Holmes was silently dropping them by only extracting TextContent. This
adds an images field to StructuredToolResult and converts MCP
ImageContent into OpenAI vision-format image_url blocks with data URIs,
so any MCP server returning images (Confluence, etc.) can be interpreted
by the LLM.

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
@claude

claude Bot commented Mar 17, 2026

Copy link
Copy Markdown
Contributor

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review.

@github-actions

github-actions Bot commented Mar 17, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 818afdc (#23380457439)

✅ Results of HolmesGPT evals

Automatically triggered by commit 818afdc on branch claude/fix-eval-233-7keOj (labels: evals-tag-grafana)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 15/16 test cases were successful, 0 regressions, 1 setup failures
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 30.4s 5 10 $0.2428 105,471 103,418 23,521 2,053 888 78,891 24,527 — —
✅ 101_loki_historical_logs_pod_deleted 42.4s 6 11 $0.2687 126,357 123,681 23,975 2,676 761 99,234 24,447 — —
✅ 112_find_pvcs_by_uuid 23.6s 5 4 $0.2028 97,718 96,604 22,018 1,114 239 74,573 22,031 — —
✅ 12_job_crashing 33.4s 6 14 $0.2778 136,691 134,589 25,178 2,102 587 106,489 28,100 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 39.1s 6 14 $0.4421 134,342 131,639 26,718 2,703 985 77,656 53,983 — —
✅ 177_grafana_home_dashboard 12.9s 3 3 $0.2848 61,688 61,092 20,991 596 250 19,465 41,627 — —
✅ 178_grafana_search_dashboard_query 19.9s 4 5 $0.2012 84,892 83,594 22,064 1,298 444 61,518 22,076 — —
✅ 179_grafana_big_dashboard_query 22.6s 4 5 $0.2180 90,964 89,826 25,146 1,138 490 64,668 25,158 — —
✅ 212_grafana_render_vision 52.1s 5 8 $0.2688 121,848 119,644 26,749 2,204 532 92,882 26,762 — —
✅ 213_grafana_render_spike_detection 82.7s 7 10 $0.3229 183,973 181,522 29,701 2,451 486 151,806 29,716 — —
✅ 227_count_configmaps_per_namespace[0] 22.2s 5 9 $0.3009 95,675 94,414 21,037 1,261 585 55,637 38,777 — —
✅ 237_grafana_large_dashboard_spike 60.6s 5 8 $0.2773 126,044 123,883 28,041 2,161 532 95,829 28,054 — —
🚧 243_pod_names_contain_service — — — — — — — — — — — — —
✅ 24_misconfigured_pvc 34.0s 6 15 $0.2689 126,019 123,486 24,015 2,533 724 98,202 25,284 — —
✅ 43_current_datetime_from_prompt 4.7s 1 — $0.1099 17,204 17,078 17,078 126 126 0 17,078 — —
✅ 61_exact_match_counting 11.4s 3 3 $0.1397 53,291 52,919 18,062 372 225 34,846 18,073 — —
Total 32.8s avg 4.7 avg 8.5 avg $3.8264 1,562,177 1,537,389 29,701 24,788 985 1,111,696 425,693 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 23 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 9e1526c (#23380280116)

✅ Results of HolmesGPT evals

Automatically triggered by commit 9e1526c on branch claude/fix-eval-233-7keOj (labels: evals-tag-grafana)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 15/16 test cases were successful, 0 regressions, 1 setup failures
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 25.3s 4 9 $0.2202 82,352 80,544 22,985 1,808 948 56,984 23,560 — —
✅ 101_loki_historical_logs_pod_deleted 58.6s 8 16 $0.3801 198,468 194,821 29,734 3,647 874 160,516 34,305 — —
✅ 112_find_pvcs_by_uuid 17.3s 3 3 $0.1803 61,158 60,166 21,798 992 598 38,357 21,809 — —
✅ 12_job_crashing 29.0s 5 12 $0.2524 110,969 109,047 24,430 1,922 688 82,549 26,498 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 39.9s 6 14 $0.2882 139,158 136,524 26,135 2,634 815 109,538 26,986 — —
✅ 177_grafana_home_dashboard 12.1s 3 3 $0.1663 61,699 61,102 21,000 597 247 40,091 21,011 — —
✅ 178_grafana_search_dashboard_query 19.5s 4 5 $0.2008 84,827 83,533 22,029 1,294 457 61,492 22,041 — —
✅ 179_grafana_big_dashboard_query 21.0s 4 5 $0.2159 90,459 89,343 24,912 1,116 467 64,419 24,924 — —
✅ 212_grafana_render_vision 35.8s 4 6 $0.2265 92,839 91,270 24,627 1,569 505 66,631 24,639 — —
✅ 213_grafana_render_spike_detection 58.8s 5 8 $0.2778 126,547 124,433 28,291 2,114 525 96,129 28,304 — —
✅ 227_count_configmaps_per_namespace[0] 24.7s 6 10 $0.2149 114,393 112,872 20,771 1,521 519 91,882 20,990 — —
✅ 237_grafana_large_dashboard_spike 74.8s 6 8 $0.4403 152,999 150,799 28,240 2,200 510 96,902 53,897 — —
🚧 243_pod_names_contain_service — — — — — — — — — — — — —
✅ 24_misconfigured_pvc 30.0s 6 11 $0.2411 120,545 118,553 22,402 1,992 415 95,433 23,120 — —
✅ 43_current_datetime_from_prompt 4.7s 1 — $0.1099 17,204 17,078 17,078 126 126 0 17,078 — —
✅ 61_exact_match_counting 8.6s 2 1 $0.1241 34,811 34,560 17,473 251 183 17,077 17,483 — —
Total 30.7s avg 4.5 avg 7.9 avg $3.5387 1,488,428 1,464,645 29,734 23,783 948 1,078,000 386,645 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 36 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 0348cc9 (#23380208120)

⚠️ Eval Results (with failures)

Automatically triggered by commit 0348cc9 on branch claude/fix-eval-233-7keOj (labels: evals-tag-grafana)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 14/16 test cases were successful, 1 regressions, 1 setup failures
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 31.8s 5 11 $0.2469 106,421 104,209 23,816 2,212 809 79,825 24,384 — —
✅ 101_loki_historical_logs_pod_deleted 44.7s 6 13 $0.3096 137,242 134,295 26,710 2,947 873 104,199 30,096 — —
✅ 112_find_pvcs_by_uuid 20.1s 4 4 $0.1902 78,353 77,119 20,996 1,234 613 56,111 21,008 — —
✅ 12_job_crashing 29.8s 5 12 $0.3705 111,487 109,535 24,726 1,952 566 62,717 46,818 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 43.2s 8 18 $0.4666 186,320 183,188 27,548 3,132 639 131,416 51,772 — —
✅ 177_grafana_home_dashboard 12.5s 3 3 $0.1659 61,660 61,072 20,979 588 245 40,082 20,990 — —
✅ 178_grafana_search_dashboard_query 19.5s 4 5 $0.1990 84,669 83,441 22,008 1,228 405 61,421 22,020 — —
✅ 179_grafana_big_dashboard_query 20.7s 4 5 $0.2155 90,446 89,345 24,913 1,101 470 64,420 24,925 — —
✅ 212_grafana_render_vision 36.1s 4 6 $0.2251 92,693 91,169 24,577 1,524 494 66,580 24,589 — —
✅ 213_grafana_render_spike_detection 62.9s 5 8 $0.2757 125,797 123,691 28,029 2,106 522 95,649 28,042 — —
✅ 227_count_configmaps_per_namespace[0] 22.3s 5 9 $0.2048 95,608 94,357 21,010 1,251 577 72,107 22,250 — —
❌ 237_grafana_large_dashboard_spike 29.5s 5 6 $0.2309 110,167 108,460 23,290 1,707 503 85,157 23,303 — —
🚧 243_pod_names_contain_service — — — — — — — — — — — — —
✅ 24_misconfigured_pvc 29.7s 5 12 $0.3361 101,866 99,895 22,304 1,971 603 58,478 41,417 — —
✅ 43_current_datetime_from_prompt 4.4s 1 — $0.1099 17,206 17,078 17,078 128 128 0 17,078 — —
✅ 61_exact_match_counting 11.0s 3 3 $0.1395 53,276 52,908 18,057 368 221 34,840 18,068 — —
Total 27.9s avg 4.5 avg 8.2 avg $3.6863 1,453,211 1,429,762 28,029 23,449 873 1,013,002 416,760 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected

📜 Run @ 0348cc9 (#23379939789)

✅ Results of HolmesGPT evals

Automatically triggered by commit 0348cc9 on branch claude/fix-eval-233-7keOj (labels: evals-tag-images)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 16/17 test cases were successful, 0 regressions, 1 setup failures
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 27.2s 4 10 $0.2295 83,876 81,827 23,431 2,049 975 57,815 24,012 — —
✅ 101_loki_historical_logs_pod_deleted 43.3s 6 12 $0.2881 133,189 130,353 24,849 2,836 874 103,476 26,877 — —
✅ 112_find_pvcs_by_uuid 15.6s 3 3 $0.1790 60,856 59,875 21,655 981 554 38,209 21,666 — —
✅ 12_job_crashing 31.6s 5 13 $0.2557 113,107 110,952 25,194 2,155 601 85,421 25,531 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 44.9s 7 14 $0.3028 160,135 157,504 26,453 2,631 536 129,685 27,819 — —
✅ 212_grafana_render_vision 36.2s 4 6 $0.2241 92,556 91,056 24,518 1,500 484 66,526 24,530 — —
✅ 213_grafana_render_spike_detection 96.4s 6 10 $0.3153 157,264 154,736 30,063 2,528 481 124,259 30,477 — —
✅ 227_count_configmaps_per_namespace[0] 22.8s 5 9 $0.2024 95,675 94,420 21,037 1,255 585 72,748 21,672 — —
✅ 234_local_mcp_image_attachment 18.1s 4 5 $0.1562 66,522 65,623 17,533 899 403 48,078 17,545 — —
✅ 236_image_spill_to_disk 21.0s 4 5 $0.2194 93,457 92,670 26,665 787 293 65,993 26,677 — —
✅ 237_grafana_large_dashboard_spike 63.0s 5 8 $0.3946 125,504 123,367 27,858 2,137 507 74,755 48,612 — —
✅ 238_image_compaction_overflow 126.8s 7 6 $0.6264 148,883 139,813 18,000 9,070 1,652 79,803 60,010 — 7
🚧 243_pod_names_contain_service — — — — — — — — — — — — —
✅ 24_misconfigured_pvc 33.6s 6 15 $0.2698 125,919 123,372 24,107 2,547 906 97,961 25,411 — —
✅ 251_mcp_confluence_image_attachment 25.3s 4 5 $0.1913 84,631 83,736 22,082 895 412 61,642 22,094 — —
✅ 43_current_datetime_from_prompt 4.3s 1 — $0.1099 17,204 17,078 17,078 126 126 0 17,078 — —
✅ 61_exact_match_counting 8.7s 2 1 $0.1244 34,832 34,571 17,484 261 192 17,077 17,494 — —
Total 38.7s avg 4.6 avg 8.1 avg $4.0888 1,593,610 1,560,953 30,063 32,657 1,652 1,123,448 437,505 — 7
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ df0b666 (#23379801486)

⚠️ Eval Results (with failures)

Automatically triggered by commit df0b666 on branch claude/fix-eval-233-7keOj (labels: evals-tag-images)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 15/17 test cases were successful, 1 regressions, 1 setup failures
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 25.3s 4 9 $0.2262 83,082 81,103 23,239 1,979 982 57,293 23,810 — —
✅ 101_loki_historical_logs_pod_deleted 49.9s 6 14 $0.3092 137,785 134,362 26,390 3,423 832 106,917 27,445 — —
✅ 112_find_pvcs_by_uuid 21.1s 5 4 $0.2042 97,931 96,777 22,066 1,154 282 74,698 22,079 — —
✅ 12_job_crashing 32.0s 6 15 $0.2769 137,743 135,463 25,557 2,280 583 108,747 26,716 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 41.3s 7 16 $0.3105 160,792 157,942 27,437 2,850 533 129,951 27,991 — —
❌ 212_grafana_render_vision 23.1s 4 5 $0.2147 87,207 85,639 23,074 1,568 523 62,553 23,086 — —
✅ 213_grafana_render_spike_detection 110.5s 9 12 $0.3919 245,094 241,777 32,119 3,317 526 208,911 32,866 — —
✅ 227_count_configmaps_per_namespace[0] 19.8s 4 9 $0.1900 77,562 76,402 20,872 1,160 563 54,901 21,501 — —
✅ 234_local_mcp_image_attachment 17.5s 4 5 $0.1550 66,328 65,453 17,449 875 393 47,992 17,461 — —
✅ 236_image_spill_to_disk 21.4s 4 5 $0.2198 93,494 92,692 26,674 802 296 66,006 26,686 — —
✅ 237_grafana_large_dashboard_spike 61.2s 5 8 $0.2774 126,129 123,984 28,133 2,145 515 95,838 28,146 — —
✅ 238_image_compaction_overflow 134.1s 7 8 $0.6615 151,111 141,085 18,069 10,026 2,337 80,066 61,019 — 7
🚧 243_pod_names_contain_service — — — — — — — — — — — — —
✅ 24_misconfigured_pvc 30.1s 6 13 $0.2562 126,752 124,590 23,521 2,162 485 100,050 24,540 — —
✅ 251_mcp_confluence_image_attachment 21.3s 4 5 $0.1886 84,251 83,419 21,918 832 396 61,489 21,930 — —
✅ 43_current_datetime_from_prompt 4.5s 1 — $0.1096 17,195 17,078 17,078 117 117 0 17,078 — —
✅ 61_exact_match_counting 11.8s 3 3 $0.1397 53,287 52,915 18,061 372 225 34,843 18,072 — —
Total 39.1s avg 4.9 avg 8.7 avg $4.1315 1,745,743 1,710,681 32,119 35,062 2,337 1,290,255 420,426 — 7
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

⚠️ 1 Failure Detected


✅ Results of HolmesGPT evals

Automatically triggered by commit dd53a94 on branch claude/fix-eval-233-7keOj (labels: evals-tag-grafana)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 15/16 test cases were successful, 0 regressions, 1 setup failures
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 30.0s 5 10 $0.2446 105,846 103,784 23,804 2,062 608 79,035 24,749 — —
✅ 101_loki_historical_logs_pod_deleted 51.3s 6 14 $0.4449 143,495 140,297 28,863 3,198 909 88,282 52,015 — —
✅ 112_find_pvcs_by_uuid 22.7s 5 4 $0.1981 95,844 94,643 20,988 1,201 287 73,642 21,001 — —
✅ 12_job_crashing 36.7s 6 14 $0.2712 132,253 129,902 24,594 2,351 594 103,971 25,931 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 35.8s 6 12 $0.2713 128,572 126,134 24,699 2,438 736 100,276 25,858 — —
✅ 177_grafana_home_dashboard 12.8s 3 3 $0.1662 61,723 61,134 21,017 589 255 40,106 21,028 — —
✅ 178_grafana_search_dashboard_query 19.1s 4 5 $0.1944 84,055 82,945 21,759 1,110 374 61,174 21,771 — —
✅ 179_grafana_big_dashboard_query 21.3s 4 5 $0.3365 91,061 89,930 25,212 1,131 467 44,136 45,794 — —
✅ 212_grafana_render_vision 45.1s 6 7 $0.3847 140,842 138,891 25,595 1,951 539 92,548 46,343 — —
✅ 213_grafana_render_spike_detection 318.9s 9 27 $0.5369 271,059 264,924 40,803 6,135 1,263 220,463 44,461 — —
✅ 227_count_configmaps_per_namespace[0] 23.3s 5 9 $0.1998 95,698 94,437 21,046 1,261 585 73,378 21,059 — —
✅ 237_grafana_large_dashboard_spike 64.1s 5 8 $0.2762 126,075 123,985 28,159 2,090 509 95,813 28,172 — —
🚧 243_pod_names_contain_service — — — — — — — — — — — — —
✅ 24_misconfigured_pvc 32.0s 6 12 $0.2480 123,618 121,504 22,723 2,114 466 97,946 23,558 — —
✅ 43_current_datetime_from_prompt 4.3s 1 — $0.1097 17,198 17,078 17,078 120 120 0 17,078 — —
✅ 61_exact_match_counting 11.9s 3 3 $0.1395 53,275 52,909 18,057 366 219 34,841 18,068 — —
Total 48.6s avg 4.9 avg 9.5 avg $4.0219 1,670,614 1,642,497 40,803 28,117 1,263 1,205,611 436,886 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-eval-233-7keOj -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-eval-233-7keOj -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented Mar 17, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Adds optional image support to tool results: StructuredToolResult gains an images field; MCP toolset extracts image blocks into that field; format_tool_result_data can return a plain string or a multimodal content list (text + image blocks); truncation and tests updated to handle multimodal content.

Changes

Cohort / File(s) Summary
Data Structure & Formatting
holmes/core/tools.py, holmes/core/models.py
Added images: Optional[List[Dict[str, str]]] to StructuredToolResult; changed format_tool_result_data return type to Union[str, List[Dict[str, Any]]] and produce multimodal content (text block + image_url/data-URI blocks) when images exist.
MCP Toolset Integration
holmes/plugins/toolsets/mcp/toolset_mcp.py
Extracts type == "image" blocks from MCP CallToolResult content into descriptors (data, mimeType) and passes them as images when constructing StructuredToolResult; added description field to RemoteMCPToolset.
Conversation Truncation
holmes/core/conversations.py
truncate_tool_messages now recognizes multimodal content lists and truncates only text blocks while preserving non-text blocks (e.g., images).
Tests & Fixtures
tests/test_mcp_toolset.py, tests/test_structured_toolcall_result.py, tests/llm/fixtures/.../toolsets.yaml, tests/llm/utils/test_toolset.py
Added tests for image extraction, multimodal formatting (format_tool_result_data and to_llm_message), and truncation behavior; updated toolset fixture credentials and strengthened toolset-loading validation.
Test Setup
conftest.py, tests/llm/fixtures/.../test_case.yaml
When auto-deriving Confluence cloud ID, set CONFLUENCE_CLOUD_ID env var and adjust log; test fixture separated credentials and switched to a pre-existing credentials.png image.

Sequence Diagram

sequenceDiagram
    participant MCP as "MCP Tool Server"
    participant Toolset as "MCP Toolset"
    participant Result as "StructuredToolResult"
    participant Formatter as "format_tool_result_data"
    participant LLM as "LLM"

    MCP->>Toolset: Return CallToolResult (content: text + image blocks)
    activate Toolset
    Toolset->>Toolset: Extract image blocks -> [{data, mimeType}, ...]
    Toolset->>Result: Create StructuredToolResult(..., images=[...])
    deactivate Toolset

    Result->>Formatter: format_tool_result_data(StructuredToolResult)
    activate Formatter
    alt result.images present
        Formatter->>Formatter: Build multimodal list (text block + image_url/data-URI blocks)
        Formatter->>LLM: Return List[Dict] (multimodal)
    else no images
        Formatter->>LLM: Return plain string
    end
    deactivate Formatter
Loading

Estimated Code Review Effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly Related PRs

Suggested Reviewers

  • arikalon1
  • nherment
🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 60.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main purpose of the changeset: adding support for image content in MCP tool results.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Mar 17, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit dd53a94
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69be9d5485fe450008b62163
😎 Deploy Preview https://deploy-preview-1803--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Mar 17, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 88f0a87b (built in 1m 5s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:88f0a87b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:88f0a87b me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:88f0a87b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:88f0a87b
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:88f0a87b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:88f0a87b me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:88f0a87b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:88f0a87b

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:88f0a87b \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:88f0a87b

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:88f0a87b \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:88f0a87b

@github-actions

github-actions Bot commented Mar 17, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 11.34s 12.78s -11.3%
Warm Mean 5.27s 5.68s -7.2%
Warm Min 5.25s 5.64s
Warm Max 5.30s 5.73s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 26.02s 36.81s -29.3%
Warm Mean 8.19s 8.47s -3.4%
Warm Min 8.05s 8.37s
Warm Max 8.34s 8.59s

PR: 88f0a87b | Master: 2b470f9c | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
holmes/core/models.py (1)

48-76: ⚠️ Potential issue | 🟠 Major

Breaking change in downstream truncate_tool_messages function.

The return type change to Union[str, List[Dict[str, Any]]] introduces a compatibility issue with truncate_tool_messages in holmes/core/conversations.py:

def truncate_tool_messages(conversation_history: list, tool_size: int) -> None:
    for message in conversation_history:
        if message.get("role") == "tool":
            message["content"] = message["content"][:tool_size]

When content is a list (images present), the slice operation [:tool_size] will truncate the list (keeping the first N content blocks) rather than truncating the text. This breaks the truncation behavior for multimodal results:

  • If tool_size is small, early image blocks may be discarded
  • Text truncation no longer works as intended; instead, content blocks are removed
Proposed fix for truncate_tool_messages
 def truncate_tool_messages(conversation_history: list, tool_size: int) -> None:
     for message in conversation_history:
         if message.get("role") == "tool":
-            message["content"] = message["content"][:tool_size]
+            content = message["content"]
+            if isinstance(content, str):
+                message["content"] = content[:tool_size]
+            elif isinstance(content, list):
+                # Truncate text within the first text block, preserve image blocks
+                truncated = []
+                for item in content:
+                    if item.get("type") == "text":
+                        truncated.append({**item, "text": item.get("text", "")[:tool_size]})
+                    else:
+                        truncated.append(item)
+                message["content"] = truncated
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/core/models.py` around lines 48 - 76, truncate_tool_messages must be
updated to handle multimodal content returned by the tool code (which now
returns either a str tool_response or a List[Dict] containing a
{"type":"text","text":...} block alongside images); change
truncate_tool_messages (in holmes/core/conversations.py) so that for messages
where message["content"] is a list it does NOT slice the list, but instead finds
the text block (dict with "type" == "text") and truncates its "text" value to
tool_size, leaving image blocks intact, while preserving the existing behavior
of slicing when message["content"] is a plain string (tool_response) or when no
text block is present.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@holmes/core/models.py`:
- Around line 48-76: truncate_tool_messages must be updated to handle multimodal
content returned by the tool code (which now returns either a str tool_response
or a List[Dict] containing a {"type":"text","text":...} block alongside images);
change truncate_tool_messages (in holmes/core/conversations.py) so that for
messages where message["content"] is a list it does NOT slice the list, but
instead finds the text block (dict with "type" == "text") and truncates its
"text" value to tool_size, leaving image blocks intact, while preserving the
existing behavior of slicing when message["content"] is a plain string
(tool_response) or when no text block is present.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 04dd7b36-b56c-46a7-97bb-e6b617d973a8

📥 Commits

Reviewing files that changed from the base of the PR and between 6ada1c2 and 48a850f.

📒 Files selected for processing (5)
  • holmes/core/models.py
  • holmes/core/tools.py
  • holmes/plugins/toolsets/mcp/toolset_mcp.py
  • tests/test_mcp_toolset.py
  • tests/test_structured_toolcall_result.py

…te_tool_messages for multimodal content

Root cause of CI failure: RemoteMCPToolset inherits Toolset which requires
`description: str` with no default. When toolsets.yaml omits description,
RemoteMCPToolset(**config) raises ValidationError, caught silently by the
generic except in load_toolsets_from_config. The toolset is dropped and the
test proceeds with no MCP tools available.

Fix: Add default description to RemoteMCPToolset so users don't need to
provide it in YAML config.

Also fixes truncate_tool_messages (CodeRabbit finding): when content is a
list (multimodal with images), the old code sliced the list by element count
instead of truncating the text. Now it finds text blocks and truncates their
text value while preserving image blocks.

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@tests/test_structured_toolcall_result.py`:
- Around line 335-347: The tests import truncate_tool_messages inside the test
functions test_truncate_tool_messages_string_content and
test_truncate_tool_messages_list_content_preserves_images; move the import to
module scope by adding a single top-level import for truncate_tool_messages near
the other imports at the top of the test file so both tests (and any others) use
the module-level symbol instead of function-scoped imports.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 8935c76c-abec-4e96-adb2-72830a84aec2

📥 Commits

Reviewing files that changed from the base of the PR and between 48a850f and 5e7c478.

📒 Files selected for processing (3)
  • holmes/core/conversations.py
  • holmes/plugins/toolsets/mcp/toolset_mcp.py
  • tests/test_structured_toolcall_result.py

Comment thread tests/test_structured_toolcall_result.py Outdated
…est framework

Three fixes:

1. Auth: The SA API token is a scoped token that only works with Bearer auth
   on the Atlassian API gateway, not basic auth on the direct URL. Switched
   mcp-atlassian to BYOT OAuth mode (ATLASSIAN_OAUTH_CLOUD_ID +
   ATLASSIAN_OAUTH_ACCESS_TOKEN + ATLASSIAN_OAUTH_ENABLE=true).

2. Error visibility: TestToolsetManager._load_custom_toolsets now detects
   when explicitly enabled toolsets are silently dropped during loading
   (e.g., validation errors) and raises ToolsetPrerequisiteError with a
   clear message pointing to common causes. Previously, the test would
   silently proceed with no tools.

3. Cloud ID: conftest.py now also exports CONFLUENCE_CLOUD_ID alongside
   CONFLUENCE_SA_BASE_URL for use by MCP toolset configs.

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
tests/llm/utils/test_toolset.py (1)

101-106: Add defensive YAML shape validation (optional improvement).

All current toolsets.yaml files use proper mapping structures, but the code should guard against malformed YAML to prevent unhandled AttributeError exceptions. Consider the proposed fix to validate that the root and section levels are dictionaries before calling .get() and .items().

Suggested improvement
-        with open(config_path) as f:
-            raw = yaml.safe_load(f) or {}
+        with open(config_path, encoding="utf-8") as f:
+            raw_loaded = yaml.safe_load(f) or {}
+        if not isinstance(raw_loaded, dict):
+            if not self.allow_toolset_failures:
+                raise ToolsetPrerequisiteError(
+                    toolset_name="<config>",
+                    error_detail=f"Invalid YAML structure in {config_path}: expected a mapping at top-level.",
+                )
+            return loaded
+        raw = raw_loaded
         expected_names = set()
         for section_key in ("toolsets", "mcp_servers"):
-            for name, cfg in (raw.get(section_key) or {}).items():
+            section = raw.get(section_key) or {}
+            if not isinstance(section, dict):
+                continue
+            for name, cfg in section.items():
                 if isinstance(cfg, dict) and cfg.get("enabled", True):
                     expected_names.add(name)
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/llm/utils/test_toolset.py` around lines 101 - 106, The test currently
assumes YAML loads to mappings and uses raw.get(...).items(), which can raise
AttributeError for malformed YAML; update the logic around the variables raw and
the per-section value to validate shapes: after loading config_path into raw,
ensure raw = raw if isinstance(raw, dict) else {}; when iterating section_key,
retrieve section = raw.get(section_key) and skip or treat as {} if not
isinstance(section, dict) before calling .items(); reference the variables raw,
expected_names, section_key, and the loop that does for name, cfg in
(raw.get(section_key) or {}).items() to implement these guards.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@tests/llm/utils/test_toolset.py`:
- Around line 101-106: The test currently assumes YAML loads to mappings and
uses raw.get(...).items(), which can raise AttributeError for malformed YAML;
update the logic around the variables raw and the per-section value to validate
shapes: after loading config_path into raw, ensure raw = raw if isinstance(raw,
dict) else {}; when iterating section_key, retrieve section =
raw.get(section_key) and skip or treat as {} if not isinstance(section, dict)
before calling .items(); reference the variables raw, expected_names,
section_key, and the loop that does for name, cfg in (raw.get(section_key) or
{}).items() to implement these guards.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 6aa38fc0-3c30-463b-94de-a7add461453b

📥 Commits

Reviewing files that changed from the base of the PR and between 5e7c478 and 13092d0.

📒 Files selected for processing (3)
  • conftest.py
  • tests/llm/fixtures/test_ask_holmes/233_mcp_confluence_image_attachment/toolsets.yaml
  • tests/llm/utils/test_toolset.py

claude added 11 commits March 19, 2026 01:39
OAuth mode has scope limitations that prevent downloading Confluence
attachments. Switch to basic auth with the personal API key, which
uses the v1 API path and doesn't have these restrictions.

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
Three fixes for eval 233:
1. Replace Pillow image generation with pre-generated credentials.png
   (Pillow is not a project dependency, fails in CI)
2. Switch MCP server to basic auth with personal API key (OAuth has
   upstream bug with v2 attachment API)
3. Keep SA Bearer auth for before_test setup (creates space/page/attachment)

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
Three fixes:
1. Add /wiki to CONFLUENCE_URL — mcp-atlassian was building attachment
   download URLs as /rest/api/content/... instead of /wiki/rest/api/content/...,
   causing the 404 error
2. Switch MCP server from personal API key to SA credentials via the
   Atlassian Cloud API gateway (CONFLUENCE_SA_BASE_URL). SA creds work
   with basic auth on the gateway and are available in all environments
3. Remove CONFLUENCE_USER/CONFLUENCE_API_KEY dependency — only SA creds
   needed now (for both setup and MCP server runtime)

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
Adds a get_test_image tool to the local stdio MCP test server and a
test_everything_stdio_image_passthrough test that exercises the full
pipeline: real MCP server → ImageContent extraction → StructuredToolResult.images
→ format_tool_result_data multimodal content → truncation preserves images.

This validates the core eval-233 fix locally without needing external
network access (Confluence).

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
Adds eval 234 as the locally-runnable counterpart of eval 233. Uses a
mock Confluence MCP server (stdio) that returns ImageContent blocks,
validating the full MCP image pipeline end-to-end:
  MCP server → ImageContent → StructuredToolResult.images →
  multimodal LLM content → vision reads code from image

Tested with OpenRouter Opus 4.6: 100% pass, 20.9s, $0.15.

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
When tool results are too large for the context window, images are now
saved to disk alongside the text data. A new read_image_file tool
(part of the bash toolset) lets the LLM load them back for analysis.

- Add save_images() to filesystem_result_storage.py
- Update prevent_overly_big_tool_response() to save images and include
  file paths in the pointer message
- Add ReadImageFile tool to bash toolset (gated by bash toolset enabled)
- Add unit tests for all new functionality

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
- Rename prevent_overly_big_tool_response → spill_oversized_tool_result
- Add docs/reference/context-management.md explaining the three mechanisms:
  spill-to-disk, compaction, and chat message truncation
- Add doc-linking comments to key source files
- Add eval 236: image spill-to-disk recovery via read_image_file with
  local MCP server that returns oversized report + image

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
… image safety

- Merge Grafana render tools from grafana-vision branch: RenderPanel and
  RenderDashboard with renderer auto-detection (disabled by default)
- Add tool-image:// embed syntax so LLMs can embed rendered images inline
  in responses (frontend resolves references against tool_calls array)
- Move multimodal assembly to to_llm_message() — format_tool_result_data()
  stays a pure string function, cleaner separation of concerns
- Fix compaction to strip images before sending to LLM — prevents context
  overflow when many images accumulate in conversation history
- Add Jinja2 prompt template teaching LLM visual investigation workflow
- Add evals: 212 (render vision), 213 (spike detection), 237 (large
  dashboard spike), 238 (image compaction overflow)
- Add shared grafana-with-renderer.yaml K8s manifest
- Update Grafana dashboards docs with visual rendering section

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
@aantn

aantn commented Mar 20, 2026

Copy link
Copy Markdown
Collaborator Author

@claude review

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Code review skipped — your organization's overage spend limit has been reached.

Code review is billed via overage credits. To resume reviews, an organization admin can raise the monthly limit in Settings → Usage.

Once credits are available, comment @claude review on this pull request to trigger a review.

Addresses CodeRabbit review comment on PR #1803.

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
Comment thread holmes/plugins/toolsets/grafana/toolset_grafana.py Outdated
Comment thread holmes/core/truncation/compaction.py
claude added 3 commits March 21, 2026 10:58
- Collect base64 images from tool call results and attach them as
  Braintrust Attachment objects in metadata for visual inspection
- Update eval 238 prompt to tell LLM to keep fetching pages until
  compaction triggers (it will be notified), rather than hardcoding
  "fetch all 3 pages"
- Add eval criteria to verify compaction actually happened

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
- Clear images on error path in spill_oversized_tool_result() to prevent
  large base64 images on ERROR-status results being sent to the LLM
- Don't extract images from ERROR MCP tool results in _invoke_async()
  since they bypass the spill/size-limiting logic
- Add deduplication guard in _try_add_render_tools() to prevent duplicate
  RenderPanel/RenderDashboard tools on repeated prerequisites_callable calls
- Include compaction instruction tokens in the image-keep overflow guard
  to prevent the compaction LLM call from exceeding its context window

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
…result

save_images() was called before confirming save_large_result() succeeded,
causing orphaned image files on disk when the text save failed. Move it
inside the if file_path: block so images are only saved when the text
file was successfully written.

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn enabled auto-merge (squash) March 21, 2026 11:14
Comment thread holmes/core/tools_utils/tool_context_window_limiter.py
Comment thread holmes/plugins/toolsets/grafana/toolset_grafana.py Outdated
Comment thread holmes/plugins/toolsets/grafana/toolset_grafana.py Outdated
Comment thread holmes/plugins/toolsets/bash/bash_toolset.py
Comment thread holmes/plugins/toolsets/grafana/toolset_grafana.py
aantn and others added 5 commits March 21, 2026 14:21
…ction, var validation

- Remove height=-1 from RenderDashboard (Puppeteer requires positive integers)
- Guard against infinite read_image_file spill loop in tool_context_window_limiter
- Restrict ReadImageFile to HOLMES_TOOL_RESULT_STORAGE_PATH for defense-in-depth
- Validate var- prefix in _build_render_query_params to prevent param overwrites
- Add early-return guard in _try_add_render_tools to skip redundant HTTP probes
- Add debug logging for renderer detection to help diagnose eval 212/237 failures

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
Tests now create files inside HOLMES_TOOL_RESULT_STORAGE_PATH instead of
tmp_path to work with the new path restriction. Added test for path
outside storage being rejected.

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
- Embed timestamp as hidden HTML comment in buildBody()
- Extract and display date in "Mar 21, 14:32 UTC" format in run summaries
- Add sequential run numbers (#1, #2, ...) to previous runs (newest = highest)
- Strip existing run numbers on re-parse to avoid double-numbering
- Bold commit SHA with __underscores__ for visual clarity
- Handle multiline header matching for timestamp comment prefix

Example: 📜 #3 · Run @ __8d93be9__ (#23379683267) — Mar 21, 14:32 UTC

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
Comment thread holmes/core/tools_utils/filesystem_result_storage.py
Comment thread holmes/core/models.py
claude added 4 commits March 21, 2026 13:03
Images were invisible in Braintrust because:
1. spill_oversized_tool_result() clears images before logging
2. _log_tool_call_result() only logged result.data, ignoring images

Now captures image count before spilling and includes it in span
metadata. When images are still present (non-oversized results),
logs image mime types and data lengths in the output dict.

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
1. Fix test 244 namespace mismatch: setup.sh and conversation_history.json
   still referenced app-161 while test_case.yaml used app-244, causing
   the test to always fail and leak the app-161 namespace.

2. Fix save_images() MIME fallback: unsupported MIME types (e.g. image/bmp)
   were silently saved with .png extension, causing LLM API errors when
   read back. Now skips unsupported types with a warning.

3. Add vision capability check: to_llm_message() now accepts supports_vision
   parameter. ToolCallingLLM checks litellm.supports_vision() before sending
   multimodal content, falling back to text-only for non-vision models.

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
litellm.supports_vision() returns False for custom gateway models
(robusta/..., openrouter/...) since they're not in its registry.
This would silently drop images for all Robusta gateway users.

Changed to use get_model_info() and only disable vision when litellm
explicitly sets supports_vision=False. Unknown models and None values
default to True — it's better to get a clear API error than silently
lose images.

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
Remove litellm.get_model_info() check which doesn't work reliably
with custom gateways. Vision is now always enabled by default.
Set HOLMES_DISABLE_VISION=true to opt out.

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn merged commit 82ba9aa into master Mar 21, 2026
21 of 22 checks passed
@aantn
aantn deleted the claude/fix-eval-233-7keOj branch March 21, 2026 13:42
n0ct1s-k8sh pushed a commit to n0ct1s-k8sh/holmesgpt that referenced this pull request Mar 23, 2026
This PR adds support for extracting and handling image content from MCP
(Model Context Protocol) tool results, enabling multimodal responses in
tool calls.

- **StructuredToolResult model**: Added optional `images` field to store
image data with MIME types
- **MCP tool invocation**: Updated `RemoteMCPTool._invoke_async()` to
extract `ImageContent` blocks from MCP responses and populate the
`images` field
- **Result formatting**: Modified `format_tool_result_data()` to return
multimodal content (text + image_url blocks in OpenAI vision format)
when images are present, while maintaining backward compatibility with
plain text responses
- **LLM message generation**: Tool call results with images now produce
properly formatted multimodal messages with data URIs for image content
- **Test coverage**: Added comprehensive tests for image extraction,
text-only results, and multimodal message formatting

- Images are extracted from MCP `CallToolResult.content` blocks where
`type == "image"`
- Image data is stored as a list of dictionaries with `data` (base64)
and `mimeType` fields
- When formatting for LLM consumption, images are converted to OpenAI
vision format data URIs: `data:{mimeType};base64,{data}`
- The `images` field is `None` when no images are present, preserving
backward compatibility
- Text content from MCP responses is merged and included alongside
images in multimodal results

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

* **New Features**
* Tool results now support multimodal content: text plus image blocks;
string-only results remain supported.
  * Tools can optionally include image data with results.

* **Behavior Changes**
* Message truncation now preserves image blocks and truncates only text
within multimodal content.

* **Tests**
* Added tests for image extraction, multimodal formatting, message
conversion, and truncation; MCP toolset image handling validated.

* **Chores**
* Test fixtures and environment handling updated; toolset credential
sourcing adjusted.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
@n0ct1s-k8sh

Copy link
Copy Markdown

Hi, which release will contain these changes?

@aantn

aantn commented Mar 23, 2026

Copy link
Copy Markdown
Collaborator Author

Hi @n0ct1s-k8sh next release! (version 0.23.0)
We don't have a specific date for release yet, but it will almost certainly be in the next 2 weeks

@n0ct1s-k8sh

Copy link
Copy Markdown

Ok then. I've backported this PR merge commit onto 0.22.0 in my fork and it works flawlessly.

Thank you for the amazing work you are doing!

@aantn

aantn commented Mar 30, 2026

Copy link
Copy Markdown
Collaborator Author

Ok then. I've backported this PR merge commit onto 0.22.0 in my fork and it works flawlessly.

Thank you for the amazing work you are doing!

With pleasure!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants