Add support for image content in MCP tool results - #1803
Conversation
MCP tool results can contain ImageContent blocks (base64 images), but Holmes was silently dropping them by only extracting TextContent. This adds an images field to StructuredToolResult and converts MCP ImageContent into OpenAI vision-format image_url blocks with data URIs, so any MCP server returning images (Confluence, etc.) can be interpreted by the LLM. https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
Claude Code ReviewThis repository is configured for manual code reviews. Comment |
📂 Previous Runs📜 Run @ 818afdc (#23380457439)✅ Results of HolmesGPT evalsAutomatically triggered by commit 818afdc on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 23 test/model combinations loaded Benchmark experiment:
No benchmark data available for comparison. Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
📜 Run @ 9e1526c (#23380280116)✅ Results of HolmesGPT evalsAutomatically triggered by commit 9e1526c on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 36 test/model combinations loaded Benchmark experiment:
No benchmark data available for comparison. Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
📜 Run @ 0348cc9 (#23380208120)
|
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 31.8s | 5 | 11 | $0.2469 | 106,421 | 104,209 | 23,816 | 2,212 | 809 | 79,825 | 24,384 | — | — |
| ✅ | 101_loki_historical_logs_pod_deleted | 44.7s | 6 | 13 | $0.3096 | 137,242 | 134,295 | 26,710 | 2,947 | 873 | 104,199 | 30,096 | — | — |
| ✅ | 112_find_pvcs_by_uuid | 20.1s | 4 | 4 | $0.1902 | 78,353 | 77,119 | 20,996 | 1,234 | 613 | 56,111 | 21,008 | — | — |
| ✅ | 12_job_crashing | 29.8s | 5 | 12 | $0.3705 | 111,487 | 109,535 | 24,726 | 1,952 | 566 | 62,717 | 46,818 | — | — |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 43.2s | 8 | 18 | $0.4666 | 186,320 | 183,188 | 27,548 | 3,132 | 639 | 131,416 | 51,772 | — | — |
| ✅ | 177_grafana_home_dashboard | 12.5s | 3 | 3 | $0.1659 | 61,660 | 61,072 | 20,979 | 588 | 245 | 40,082 | 20,990 | — | — |
| ✅ | 178_grafana_search_dashboard_query | 19.5s | 4 | 5 | $0.1990 | 84,669 | 83,441 | 22,008 | 1,228 | 405 | 61,421 | 22,020 | — | — |
| ✅ | 179_grafana_big_dashboard_query | 20.7s | 4 | 5 | $0.2155 | 90,446 | 89,345 | 24,913 | 1,101 | 470 | 64,420 | 24,925 | — | — |
| ✅ | 212_grafana_render_vision | 36.1s | 4 | 6 | $0.2251 | 92,693 | 91,169 | 24,577 | 1,524 | 494 | 66,580 | 24,589 | — | — |
| ✅ | 213_grafana_render_spike_detection | 62.9s | 5 | 8 | $0.2757 | 125,797 | 123,691 | 28,029 | 2,106 | 522 | 95,649 | 28,042 | — | — |
| ✅ | 227_count_configmaps_per_namespace[0] | 22.3s | 5 | 9 | $0.2048 | 95,608 | 94,357 | 21,010 | 1,251 | 577 | 72,107 | 22,250 | — | — |
| ❌ | 237_grafana_large_dashboard_spike | 29.5s | 5 | 6 | $0.2309 | 110,167 | 108,460 | 23,290 | 1,707 | 503 | 85,157 | 23,303 | — | — |
| 🚧 | 243_pod_names_contain_service | — | — | — | — | — | — | — | — | — | — | — | — | — |
| ✅ | 24_misconfigured_pvc | 29.7s | 5 | 12 | $0.3361 | 101,866 | 99,895 | 22,304 | 1,971 | 603 | 58,478 | 41,417 | — | — |
| ✅ | 43_current_datetime_from_prompt | 4.4s | 1 | — | $0.1099 | 17,206 | 17,078 | 17,078 | 128 | 128 | 0 | 17,078 | — | — |
| ✅ | 61_exact_match_counting | 11.0s | 3 | 3 | $0.1395 | 53,276 | 52,908 | 18,057 | 368 | 221 | 34,840 | 18,068 | — | — |
| Total | 27.9s avg | 4.5 avg | 8.2 avg | $3.6863 | 1,453,211 | 1,429,762 | 28,029 | 23,449 | 873 | 1,013,002 | 416,760 | — | — |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 73 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 1 Failure Detected
📜 Run @ 0348cc9 (#23379939789)
✅ Results of HolmesGPT evals
Automatically triggered by commit 0348cc9 on branch claude/fix-eval-233-7keOj (labels: evals-tag-images)
Results of HolmesGPT evals
- ask_holmes: 16/17 test cases were successful, 0 regressions, 1 setup failures
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 27.2s | 4 | 10 | $0.2295 | 83,876 | 81,827 | 23,431 | 2,049 | 975 | 57,815 | 24,012 | — | — |
| ✅ | 101_loki_historical_logs_pod_deleted | 43.3s | 6 | 12 | $0.2881 | 133,189 | 130,353 | 24,849 | 2,836 | 874 | 103,476 | 26,877 | — | — |
| ✅ | 112_find_pvcs_by_uuid | 15.6s | 3 | 3 | $0.1790 | 60,856 | 59,875 | 21,655 | 981 | 554 | 38,209 | 21,666 | — | — |
| ✅ | 12_job_crashing | 31.6s | 5 | 13 | $0.2557 | 113,107 | 110,952 | 25,194 | 2,155 | 601 | 85,421 | 25,531 | — | — |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 44.9s | 7 | 14 | $0.3028 | 160,135 | 157,504 | 26,453 | 2,631 | 536 | 129,685 | 27,819 | — | — |
| ✅ | 212_grafana_render_vision | 36.2s | 4 | 6 | $0.2241 | 92,556 | 91,056 | 24,518 | 1,500 | 484 | 66,526 | 24,530 | — | — |
| ✅ | 213_grafana_render_spike_detection | 96.4s | 6 | 10 | $0.3153 | 157,264 | 154,736 | 30,063 | 2,528 | 481 | 124,259 | 30,477 | — | — |
| ✅ | 227_count_configmaps_per_namespace[0] | 22.8s | 5 | 9 | $0.2024 | 95,675 | 94,420 | 21,037 | 1,255 | 585 | 72,748 | 21,672 | — | — |
| ✅ | 234_local_mcp_image_attachment | 18.1s | 4 | 5 | $0.1562 | 66,522 | 65,623 | 17,533 | 899 | 403 | 48,078 | 17,545 | — | — |
| ✅ | 236_image_spill_to_disk | 21.0s | 4 | 5 | $0.2194 | 93,457 | 92,670 | 26,665 | 787 | 293 | 65,993 | 26,677 | — | — |
| ✅ | 237_grafana_large_dashboard_spike | 63.0s | 5 | 8 | $0.3946 | 125,504 | 123,367 | 27,858 | 2,137 | 507 | 74,755 | 48,612 | — | — |
| ✅ | 238_image_compaction_overflow | 126.8s | 7 | 6 | $0.6264 | 148,883 | 139,813 | 18,000 | 9,070 | 1,652 | 79,803 | 60,010 | — | 7 |
| 🚧 | 243_pod_names_contain_service | — | — | — | — | — | — | — | — | — | — | — | — | — |
| ✅ | 24_misconfigured_pvc | 33.6s | 6 | 15 | $0.2698 | 125,919 | 123,372 | 24,107 | 2,547 | 906 | 97,961 | 25,411 | — | — |
| ✅ | 251_mcp_confluence_image_attachment | 25.3s | 4 | 5 | $0.1913 | 84,631 | 83,736 | 22,082 | 895 | 412 | 61,642 | 22,094 | — | — |
| ✅ | 43_current_datetime_from_prompt | 4.3s | 1 | — | $0.1099 | 17,204 | 17,078 | 17,078 | 126 | 126 | 0 | 17,078 | — | — |
| ✅ | 61_exact_match_counting | 8.7s | 2 | 1 | $0.1244 | 34,832 | 34,571 | 17,484 | 261 | 192 | 17,077 | 17,494 | — | — |
| Total | 38.7s avg | 4.6 avg | 8.1 avg | $4.0888 | 1,593,610 | 1,560,953 | 30,063 | 32,657 | 1,652 | 1,123,448 | 437,505 | — | 7 |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 73 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
📜 Run @ df0b666 (#23379801486)
⚠️ Eval Results (with failures)
Automatically triggered by commit df0b666 on branch claude/fix-eval-233-7keOj (labels: evals-tag-images)
Results of HolmesGPT evals
- ask_holmes: 15/17 test cases were successful, 1 regressions, 1 setup failures
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 25.3s | 4 | 9 | $0.2262 | 83,082 | 81,103 | 23,239 | 1,979 | 982 | 57,293 | 23,810 | — | — |
| ✅ | 101_loki_historical_logs_pod_deleted | 49.9s | 6 | 14 | $0.3092 | 137,785 | 134,362 | 26,390 | 3,423 | 832 | 106,917 | 27,445 | — | — |
| ✅ | 112_find_pvcs_by_uuid | 21.1s | 5 | 4 | $0.2042 | 97,931 | 96,777 | 22,066 | 1,154 | 282 | 74,698 | 22,079 | — | — |
| ✅ | 12_job_crashing | 32.0s | 6 | 15 | $0.2769 | 137,743 | 135,463 | 25,557 | 2,280 | 583 | 108,747 | 26,716 | — | — |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 41.3s | 7 | 16 | $0.3105 | 160,792 | 157,942 | 27,437 | 2,850 | 533 | 129,951 | 27,991 | — | — |
| ❌ | 212_grafana_render_vision | 23.1s | 4 | 5 | $0.2147 | 87,207 | 85,639 | 23,074 | 1,568 | 523 | 62,553 | 23,086 | — | — |
| ✅ | 213_grafana_render_spike_detection | 110.5s | 9 | 12 | $0.3919 | 245,094 | 241,777 | 32,119 | 3,317 | 526 | 208,911 | 32,866 | — | — |
| ✅ | 227_count_configmaps_per_namespace[0] | 19.8s | 4 | 9 | $0.1900 | 77,562 | 76,402 | 20,872 | 1,160 | 563 | 54,901 | 21,501 | — | — |
| ✅ | 234_local_mcp_image_attachment | 17.5s | 4 | 5 | $0.1550 | 66,328 | 65,453 | 17,449 | 875 | 393 | 47,992 | 17,461 | — | — |
| ✅ | 236_image_spill_to_disk | 21.4s | 4 | 5 | $0.2198 | 93,494 | 92,692 | 26,674 | 802 | 296 | 66,006 | 26,686 | — | — |
| ✅ | 237_grafana_large_dashboard_spike | 61.2s | 5 | 8 | $0.2774 | 126,129 | 123,984 | 28,133 | 2,145 | 515 | 95,838 | 28,146 | — | — |
| ✅ | 238_image_compaction_overflow | 134.1s | 7 | 8 | $0.6615 | 151,111 | 141,085 | 18,069 | 10,026 | 2,337 | 80,066 | 61,019 | — | 7 |
| 🚧 | 243_pod_names_contain_service | — | — | — | — | — | — | — | — | — | — | — | — | — |
| ✅ | 24_misconfigured_pvc | 30.1s | 6 | 13 | $0.2562 | 126,752 | 124,590 | 23,521 | 2,162 | 485 | 100,050 | 24,540 | — | — |
| ✅ | 251_mcp_confluence_image_attachment | 21.3s | 4 | 5 | $0.1886 | 84,251 | 83,419 | 21,918 | 832 | 396 | 61,489 | 21,930 | — | — |
| ✅ | 43_current_datetime_from_prompt | 4.5s | 1 | — | $0.1096 | 17,195 | 17,078 | 17,078 | 117 | 117 | 0 | 17,078 | — | — |
| ✅ | 61_exact_match_counting | 11.8s | 3 | 3 | $0.1397 | 53,287 | 52,915 | 18,061 | 372 | 225 | 34,843 | 18,072 | — | — |
| Total | 39.1s avg | 4.9 avg | 8.7 avg | $4.1315 | 1,745,743 | 1,710,681 | 32,119 | 35,062 | 2,337 | 1,290,255 | 420,426 | — | 7 |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 73 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
⚠️ 1 Failure Detected
✅ Results of HolmesGPT evals
Automatically triggered by commit dd53a94 on branch claude/fix-eval-233-7keOj (labels: evals-tag-grafana)
Results of HolmesGPT evals
- ask_holmes: 15/16 test cases were successful, 0 regressions, 1 setup failures
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 30.0s | 5 | 10 | $0.2446 | 105,846 | 103,784 | 23,804 | 2,062 | 608 | 79,035 | 24,749 | — | — |
| ✅ | 101_loki_historical_logs_pod_deleted | 51.3s | 6 | 14 | $0.4449 | 143,495 | 140,297 | 28,863 | 3,198 | 909 | 88,282 | 52,015 | — | — |
| ✅ | 112_find_pvcs_by_uuid | 22.7s | 5 | 4 | $0.1981 | 95,844 | 94,643 | 20,988 | 1,201 | 287 | 73,642 | 21,001 | — | — |
| ✅ | 12_job_crashing | 36.7s | 6 | 14 | $0.2712 | 132,253 | 129,902 | 24,594 | 2,351 | 594 | 103,971 | 25,931 | — | — |
| ✅ | 176_network_policy_blocking_traffic_no_runbooks | 35.8s | 6 | 12 | $0.2713 | 128,572 | 126,134 | 24,699 | 2,438 | 736 | 100,276 | 25,858 | — | — |
| ✅ | 177_grafana_home_dashboard | 12.8s | 3 | 3 | $0.1662 | 61,723 | 61,134 | 21,017 | 589 | 255 | 40,106 | 21,028 | — | — |
| ✅ | 178_grafana_search_dashboard_query | 19.1s | 4 | 5 | $0.1944 | 84,055 | 82,945 | 21,759 | 1,110 | 374 | 61,174 | 21,771 | — | — |
| ✅ | 179_grafana_big_dashboard_query | 21.3s | 4 | 5 | $0.3365 | 91,061 | 89,930 | 25,212 | 1,131 | 467 | 44,136 | 45,794 | — | — |
| ✅ | 212_grafana_render_vision | 45.1s | 6 | 7 | $0.3847 | 140,842 | 138,891 | 25,595 | 1,951 | 539 | 92,548 | 46,343 | — | — |
| ✅ | 213_grafana_render_spike_detection | 318.9s | 9 | 27 | $0.5369 | 271,059 | 264,924 | 40,803 | 6,135 | 1,263 | 220,463 | 44,461 | — | — |
| ✅ | 227_count_configmaps_per_namespace[0] | 23.3s | 5 | 9 | $0.1998 | 95,698 | 94,437 | 21,046 | 1,261 | 585 | 73,378 | 21,059 | — | — |
| ✅ | 237_grafana_large_dashboard_spike | 64.1s | 5 | 8 | $0.2762 | 126,075 | 123,985 | 28,159 | 2,090 | 509 | 95,813 | 28,172 | — | — |
| 🚧 | 243_pod_names_contain_service | — | — | — | — | — | — | — | — | — | — | — | — | — |
| ✅ | 24_misconfigured_pvc | 32.0s | 6 | 12 | $0.2480 | 123,618 | 121,504 | 22,723 | 2,114 | 466 | 97,946 | 23,558 | — | — |
| ✅ | 43_current_datetime_from_prompt | 4.3s | 1 | — | $0.1097 | 17,198 | 17,078 | 17,078 | 120 | 120 | 0 | 17,078 | — | — |
| ✅ | 61_exact_match_counting | 11.9s | 3 | 3 | $0.1395 | 53,275 | 52,909 | 18,057 | 366 | 219 | 34,841 | 18,068 | — | — |
| Total | 48.6s avg | 4.9 avg | 9.5 avg | $4.0219 | 1,670,614 | 1,642,497 | 40,803 | 28,117 | 1,263 | 1,205,611 | 436,886 | — | — |
Benchmark Comparison Details
Baseline: latest ci-benchmark experiment on master
Status: Success - 73 test/model combinations loaded
Benchmark experiment:
- ci-benchmark-23102181491 (created: 2026-03-15)
No benchmark data available for comparison.
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
📖 Legend
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-eval-233-7keOj -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
tags: regression
Or with more options (one per line):
/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
tags: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
tags |
Pytest tags / markers (no default - runs all tests!) |
id |
Eval ID / pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):
| Label | Effect |
|---|---|
evals-tag-<name> |
Run tests with tag <name> alongside regression |
evals-id-<name> |
Run a specific eval by test ID |
evals-model-<name> |
Override the model (use model list name, e.g. sonnet-4.5) |
Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5
🏷️ Valid tags
benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency
🤖 Valid models
deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-eval-233-7keOj -f markers=regression -f filter=
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughAdds optional image support to tool results: Changes
Sequence DiagramsequenceDiagram
participant MCP as "MCP Tool Server"
participant Toolset as "MCP Toolset"
participant Result as "StructuredToolResult"
participant Formatter as "format_tool_result_data"
participant LLM as "LLM"
MCP->>Toolset: Return CallToolResult (content: text + image blocks)
activate Toolset
Toolset->>Toolset: Extract image blocks -> [{data, mimeType}, ...]
Toolset->>Result: Create StructuredToolResult(..., images=[...])
deactivate Toolset
Result->>Formatter: format_tool_result_data(StructuredToolResult)
activate Formatter
alt result.images present
Formatter->>Formatter: Build multimodal list (text block + image_url/data-URI blocks)
Formatter->>LLM: Return List[Dict] (multimodal)
else no images
Formatter->>LLM: Return plain string
end
deactivate Formatter
Estimated Code Review Effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly Related PRs
Suggested Reviewers
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:88f0a87b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:88f0a87b me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:88f0a87b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:88f0a87b
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:88f0a87b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:88f0a87b me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:88f0a87b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:88f0a87bPatch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:88f0a87b \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:88f0a87bRobusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:88f0a87b \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:88f0a87b |
🔬 CLI Performance Benchmark🟡 Startup Time (no LLM)Measures
🟡 Full CLI with LLMMeasures
PR: |
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
holmes/core/models.py (1)
48-76:⚠️ Potential issue | 🟠 MajorBreaking change in downstream
truncate_tool_messagesfunction.The return type change to
Union[str, List[Dict[str, Any]]]introduces a compatibility issue withtruncate_tool_messagesinholmes/core/conversations.py:def truncate_tool_messages(conversation_history: list, tool_size: int) -> None: for message in conversation_history: if message.get("role") == "tool": message["content"] = message["content"][:tool_size]When
contentis a list (images present), the slice operation[:tool_size]will truncate the list (keeping the first N content blocks) rather than truncating the text. This breaks the truncation behavior for multimodal results:
- If
tool_sizeis small, early image blocks may be discarded- Text truncation no longer works as intended; instead, content blocks are removed
Proposed fix for truncate_tool_messages
def truncate_tool_messages(conversation_history: list, tool_size: int) -> None: for message in conversation_history: if message.get("role") == "tool": - message["content"] = message["content"][:tool_size] + content = message["content"] + if isinstance(content, str): + message["content"] = content[:tool_size] + elif isinstance(content, list): + # Truncate text within the first text block, preserve image blocks + truncated = [] + for item in content: + if item.get("type") == "text": + truncated.append({**item, "text": item.get("text", "")[:tool_size]}) + else: + truncated.append(item) + message["content"] = truncated🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@holmes/core/models.py` around lines 48 - 76, truncate_tool_messages must be updated to handle multimodal content returned by the tool code (which now returns either a str tool_response or a List[Dict] containing a {"type":"text","text":...} block alongside images); change truncate_tool_messages (in holmes/core/conversations.py) so that for messages where message["content"] is a list it does NOT slice the list, but instead finds the text block (dict with "type" == "text") and truncates its "text" value to tool_size, leaving image blocks intact, while preserving the existing behavior of slicing when message["content"] is a plain string (tool_response) or when no text block is present.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@holmes/core/models.py`:
- Around line 48-76: truncate_tool_messages must be updated to handle multimodal
content returned by the tool code (which now returns either a str tool_response
or a List[Dict] containing a {"type":"text","text":...} block alongside images);
change truncate_tool_messages (in holmes/core/conversations.py) so that for
messages where message["content"] is a list it does NOT slice the list, but
instead finds the text block (dict with "type" == "text") and truncates its
"text" value to tool_size, leaving image blocks intact, while preserving the
existing behavior of slicing when message["content"] is a plain string
(tool_response) or when no text block is present.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 04dd7b36-b56c-46a7-97bb-e6b617d973a8
📒 Files selected for processing (5)
holmes/core/models.pyholmes/core/tools.pyholmes/plugins/toolsets/mcp/toolset_mcp.pytests/test_mcp_toolset.pytests/test_structured_toolcall_result.py
…te_tool_messages for multimodal content Root cause of CI failure: RemoteMCPToolset inherits Toolset which requires `description: str` with no default. When toolsets.yaml omits description, RemoteMCPToolset(**config) raises ValidationError, caught silently by the generic except in load_toolsets_from_config. The toolset is dropped and the test proceeds with no MCP tools available. Fix: Add default description to RemoteMCPToolset so users don't need to provide it in YAML config. Also fixes truncate_tool_messages (CodeRabbit finding): when content is a list (multimodal with images), the old code sliced the list by element count instead of truncating the text. Now it finds text blocks and truncates their text value while preserving image blocks. https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@tests/test_structured_toolcall_result.py`:
- Around line 335-347: The tests import truncate_tool_messages inside the test
functions test_truncate_tool_messages_string_content and
test_truncate_tool_messages_list_content_preserves_images; move the import to
module scope by adding a single top-level import for truncate_tool_messages near
the other imports at the top of the test file so both tests (and any others) use
the module-level symbol instead of function-scoped imports.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 8935c76c-abec-4e96-adb2-72830a84aec2
📒 Files selected for processing (3)
holmes/core/conversations.pyholmes/plugins/toolsets/mcp/toolset_mcp.pytests/test_structured_toolcall_result.py
…est framework Three fixes: 1. Auth: The SA API token is a scoped token that only works with Bearer auth on the Atlassian API gateway, not basic auth on the direct URL. Switched mcp-atlassian to BYOT OAuth mode (ATLASSIAN_OAUTH_CLOUD_ID + ATLASSIAN_OAUTH_ACCESS_TOKEN + ATLASSIAN_OAUTH_ENABLE=true). 2. Error visibility: TestToolsetManager._load_custom_toolsets now detects when explicitly enabled toolsets are silently dropped during loading (e.g., validation errors) and raises ToolsetPrerequisiteError with a clear message pointing to common causes. Previously, the test would silently proceed with no tools. 3. Cloud ID: conftest.py now also exports CONFLUENCE_CLOUD_ID alongside CONFLUENCE_SA_BASE_URL for use by MCP toolset configs. https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
There was a problem hiding this comment.
🧹 Nitpick comments (1)
tests/llm/utils/test_toolset.py (1)
101-106: Add defensive YAML shape validation (optional improvement).All current
toolsets.yamlfiles use proper mapping structures, but the code should guard against malformed YAML to prevent unhandledAttributeErrorexceptions. Consider the proposed fix to validate that the root and section levels are dictionaries before calling.get()and.items().Suggested improvement
- with open(config_path) as f: - raw = yaml.safe_load(f) or {} + with open(config_path, encoding="utf-8") as f: + raw_loaded = yaml.safe_load(f) or {} + if not isinstance(raw_loaded, dict): + if not self.allow_toolset_failures: + raise ToolsetPrerequisiteError( + toolset_name="<config>", + error_detail=f"Invalid YAML structure in {config_path}: expected a mapping at top-level.", + ) + return loaded + raw = raw_loaded expected_names = set() for section_key in ("toolsets", "mcp_servers"): - for name, cfg in (raw.get(section_key) or {}).items(): + section = raw.get(section_key) or {} + if not isinstance(section, dict): + continue + for name, cfg in section.items(): if isinstance(cfg, dict) and cfg.get("enabled", True): expected_names.add(name)🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@tests/llm/utils/test_toolset.py` around lines 101 - 106, The test currently assumes YAML loads to mappings and uses raw.get(...).items(), which can raise AttributeError for malformed YAML; update the logic around the variables raw and the per-section value to validate shapes: after loading config_path into raw, ensure raw = raw if isinstance(raw, dict) else {}; when iterating section_key, retrieve section = raw.get(section_key) and skip or treat as {} if not isinstance(section, dict) before calling .items(); reference the variables raw, expected_names, section_key, and the loop that does for name, cfg in (raw.get(section_key) or {}).items() to implement these guards.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Nitpick comments:
In `@tests/llm/utils/test_toolset.py`:
- Around line 101-106: The test currently assumes YAML loads to mappings and
uses raw.get(...).items(), which can raise AttributeError for malformed YAML;
update the logic around the variables raw and the per-section value to validate
shapes: after loading config_path into raw, ensure raw = raw if isinstance(raw,
dict) else {}; when iterating section_key, retrieve section =
raw.get(section_key) and skip or treat as {} if not isinstance(section, dict)
before calling .items(); reference the variables raw, expected_names,
section_key, and the loop that does for name, cfg in (raw.get(section_key) or
{}).items() to implement these guards.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 6aa38fc0-3c30-463b-94de-a7add461453b
📒 Files selected for processing (3)
conftest.pytests/llm/fixtures/test_ask_holmes/233_mcp_confluence_image_attachment/toolsets.yamltests/llm/utils/test_toolset.py
OAuth mode has scope limitations that prevent downloading Confluence attachments. Switch to basic auth with the personal API key, which uses the v1 API path and doesn't have these restrictions. https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
Three fixes for eval 233: 1. Replace Pillow image generation with pre-generated credentials.png (Pillow is not a project dependency, fails in CI) 2. Switch MCP server to basic auth with personal API key (OAuth has upstream bug with v2 attachment API) 3. Keep SA Bearer auth for before_test setup (creates space/page/attachment) https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
Three fixes: 1. Add /wiki to CONFLUENCE_URL — mcp-atlassian was building attachment download URLs as /rest/api/content/... instead of /wiki/rest/api/content/..., causing the 404 error 2. Switch MCP server from personal API key to SA credentials via the Atlassian Cloud API gateway (CONFLUENCE_SA_BASE_URL). SA creds work with basic auth on the gateway and are available in all environments 3. Remove CONFLUENCE_USER/CONFLUENCE_API_KEY dependency — only SA creds needed now (for both setup and MCP server runtime) https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
Adds a get_test_image tool to the local stdio MCP test server and a test_everything_stdio_image_passthrough test that exercises the full pipeline: real MCP server → ImageContent extraction → StructuredToolResult.images → format_tool_result_data multimodal content → truncation preserves images. This validates the core eval-233 fix locally without needing external network access (Confluence). https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
Adds eval 234 as the locally-runnable counterpart of eval 233. Uses a mock Confluence MCP server (stdio) that returns ImageContent blocks, validating the full MCP image pipeline end-to-end: MCP server → ImageContent → StructuredToolResult.images → multimodal LLM content → vision reads code from image Tested with OpenRouter Opus 4.6: 100% pass, 20.9s, $0.15. https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
When tool results are too large for the context window, images are now saved to disk alongside the text data. A new read_image_file tool (part of the bash toolset) lets the LLM load them back for analysis. - Add save_images() to filesystem_result_storage.py - Update prevent_overly_big_tool_response() to save images and include file paths in the pointer message - Add ReadImageFile tool to bash toolset (gated by bash toolset enabled) - Add unit tests for all new functionality https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
- Rename prevent_overly_big_tool_response → spill_oversized_tool_result - Add docs/reference/context-management.md explaining the three mechanisms: spill-to-disk, compaction, and chat message truncation - Add doc-linking comments to key source files - Add eval 236: image spill-to-disk recovery via read_image_file with local MCP server that returns oversized report + image https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
… image safety - Merge Grafana render tools from grafana-vision branch: RenderPanel and RenderDashboard with renderer auto-detection (disabled by default) - Add tool-image:// embed syntax so LLMs can embed rendered images inline in responses (frontend resolves references against tool_calls array) - Move multimodal assembly to to_llm_message() — format_tool_result_data() stays a pure string function, cleaner separation of concerns - Fix compaction to strip images before sending to LLM — prevents context overflow when many images accumulate in conversation history - Add Jinja2 prompt template teaching LLM visual investigation workflow - Add evals: 212 (render vision), 213 (spike detection), 237 (large dashboard spike), 238 (image compaction overflow) - Add shared grafana-with-renderer.yaml K8s manifest - Update Grafana dashboards docs with visual rendering section https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
|
@claude review |
There was a problem hiding this comment.
Code review is billed via overage credits. To resume reviews, an organization admin can raise the monthly limit in Settings → Usage.
Once credits are available, comment @claude review on this pull request to trigger a review.
Addresses CodeRabbit review comment on PR #1803. https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
- Collect base64 images from tool call results and attach them as Braintrust Attachment objects in metadata for visual inspection - Update eval 238 prompt to tell LLM to keep fetching pages until compaction triggers (it will be notified), rather than hardcoding "fetch all 3 pages" - Add eval criteria to verify compaction actually happened https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
- Clear images on error path in spill_oversized_tool_result() to prevent large base64 images on ERROR-status results being sent to the LLM - Don't extract images from ERROR MCP tool results in _invoke_async() since they bypass the spill/size-limiting logic - Add deduplication guard in _try_add_render_tools() to prevent duplicate RenderPanel/RenderDashboard tools on repeated prerequisites_callable calls - Include compaction instruction tokens in the image-keep overflow guard to prevent the compaction LLM call from exceeding its context window https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
…result save_images() was called before confirming save_large_result() succeeded, causing orphaned image files on disk when the text save failed. Move it inside the if file_path: block so images are only saved when the text file was successfully written. https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
…ction, var validation - Remove height=-1 from RenderDashboard (Puppeteer requires positive integers) - Guard against infinite read_image_file spill loop in tool_context_window_limiter - Restrict ReadImageFile to HOLMES_TOOL_RESULT_STORAGE_PATH for defense-in-depth - Validate var- prefix in _build_render_query_params to prevent param overwrites - Add early-return guard in _try_add_render_tools to skip redundant HTTP probes - Add debug logging for renderer detection to help diagnose eval 212/237 failures https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
…t/HolmesGPT/holmesgpt into claude/fix-eval-233-7keOj
Tests now create files inside HOLMES_TOOL_RESULT_STORAGE_PATH instead of tmp_path to work with the new path restriction. Added test for path outside storage being rejected. https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
- Embed timestamp as hidden HTML comment in buildBody() - Extract and display date in "Mar 21, 14:32 UTC" format in run summaries - Add sequential run numbers (#1, #2, ...) to previous runs (newest = highest) - Strip existing run numbers on re-parse to avoid double-numbering - Bold commit SHA with __underscores__ for visual clarity - Handle multiline header matching for timestamp comment prefix Example: 📜 #3 · Run @ __8d93be9__ (#23379683267) — Mar 21, 14:32 UTC https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
Images were invisible in Braintrust because: 1. spill_oversized_tool_result() clears images before logging 2. _log_tool_call_result() only logged result.data, ignoring images Now captures image count before spilling and includes it in span metadata. When images are still present (non-oversized results), logs image mime types and data lengths in the output dict. https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
1. Fix test 244 namespace mismatch: setup.sh and conversation_history.json still referenced app-161 while test_case.yaml used app-244, causing the test to always fail and leak the app-161 namespace. 2. Fix save_images() MIME fallback: unsupported MIME types (e.g. image/bmp) were silently saved with .png extension, causing LLM API errors when read back. Now skips unsupported types with a warning. 3. Add vision capability check: to_llm_message() now accepts supports_vision parameter. ToolCallingLLM checks litellm.supports_vision() before sending multimodal content, falling back to text-only for non-vision models. https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
litellm.supports_vision() returns False for custom gateway models (robusta/..., openrouter/...) since they're not in its registry. This would silently drop images for all Robusta gateway users. Changed to use get_model_info() and only disable vision when litellm explicitly sets supports_vision=False. Unknown models and None values default to True — it's better to get a clear API error than silently lose images. https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
Remove litellm.get_model_info() check which doesn't work reliably with custom gateways. Vision is now always enabled by default. Set HOLMES_DISABLE_VISION=true to opt out. https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
This PR adds support for extracting and handling image content from MCP
(Model Context Protocol) tool results, enabling multimodal responses in
tool calls.
- **StructuredToolResult model**: Added optional `images` field to store
image data with MIME types
- **MCP tool invocation**: Updated `RemoteMCPTool._invoke_async()` to
extract `ImageContent` blocks from MCP responses and populate the
`images` field
- **Result formatting**: Modified `format_tool_result_data()` to return
multimodal content (text + image_url blocks in OpenAI vision format)
when images are present, while maintaining backward compatibility with
plain text responses
- **LLM message generation**: Tool call results with images now produce
properly formatted multimodal messages with data URIs for image content
- **Test coverage**: Added comprehensive tests for image extraction,
text-only results, and multimodal message formatting
- Images are extracted from MCP `CallToolResult.content` blocks where
`type == "image"`
- Image data is stored as a list of dictionaries with `data` (base64)
and `mimeType` fields
- When formatting for LLM consumption, images are converted to OpenAI
vision format data URIs: `data:{mimeType};base64,{data}`
- The `images` field is `None` when no images are present, preserving
backward compatibility
- Text content from MCP responses is merged and included alongside
images in multimodal results
https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
* **New Features**
* Tool results now support multimodal content: text plus image blocks;
string-only results remain supported.
* Tools can optionally include image data with results.
* **Behavior Changes**
* Message truncation now preserves image blocks and truncates only text
within multimodal content.
* **Tests**
* Added tests for image extraction, multimodal formatting, message
conversion, and truncation; MCP toolset image handling validated.
* **Chores**
* Test fixtures and environment handling updated; toolset credential
sourcing adjusted.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
Hi, which release will contain these changes? |
|
Hi @n0ct1s-k8sh next release! (version 0.23.0) |
|
Ok then. I've backported this PR merge commit onto 0.22.0 in my fork and it works flawlessly. Thank you for the amazing work you are doing! |
With pleasure! |
Summary
This PR adds support for extracting and handling image content from MCP (Model Context Protocol) tool results, enabling multimodal responses in tool calls.
Key Changes
imagesfield to store image data with MIME typesRemoteMCPTool._invoke_async()to extractImageContentblocks from MCP responses and populate theimagesfieldformat_tool_result_data()to return multimodal content (text + image_url blocks in OpenAI vision format) when images are present, while maintaining backward compatibility with plain text responsesImplementation Details
CallToolResult.contentblocks wheretype == "image"data(base64) andmimeTypefieldsdata:{mimeType};base64,{data}imagesfield isNonewhen no images are present, preserving backward compatibilityhttps://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Summary by CodeRabbit
New Features
Behavior Changes
Tests
Chores