Skip to content

ROB-3756 - Extract text from MCP resource content blocks in tool results - #1961

Merged
naomi-robusta merged 4 commits into
masterfrom
claude/fix-github-get-file-yw3Lx
Apr 29, 2026
Merged

naomi-robusta merged 4 commits into
masterfrom
claude/fix-github-get-file-yw3Lx

Conversation

@naomi-robusta

@naomi-robusta naomi-robusta commented Apr 28, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR fixes a bug where MCP tool results containing file contents in EmbeddedResource blocks were silently dropped, preventing the LLM from accessing the actual data. The fix extracts text from all MCP content block types and properly handles different resource formats.

Key Changes

  • Added _extract_text_from_content_block() method to RemoteMCPTool that handles extraction from:

    • TextContent: Direct text passthrough
    • EmbeddedResource with TextResourceContents: Extracts the text field
    • EmbeddedResource with BlobResourceContents: Base64-decodes when mimeType indicates text (text/*, application/json, application/xml, etc.), otherwise returns a placeholder with resource metadata
    • ResourceLink: Surfaces the URI as a hint for large files (>1MB)
  • Updated _invoke_async() method to use the new extraction logic instead of only filtering for TextContent blocks

  • Added comprehensive test coverage with four new test cases:

    • test_invoke_async_extracts_text_resource_contents: Verifies text extraction from TextResourceContents
    • test_invoke_async_decodes_text_blob_resource: Verifies base64 decoding of text-like blobs
    • test_invoke_async_keeps_binary_blob_as_placeholder: Verifies binary blobs emit a placeholder instead of failing
    • test_invoke_async_surfaces_resource_link: Verifies ResourceLink URIs are surfaced

Implementation Details

  • The extraction method uses attribute introspection (getattr) to safely handle different content block types without strict type checking
  • Text-like MIME types are identified by checking for text/ prefix, specific types like application/json, and +json/+xml suffixes
  • Base64 decoding uses errors='replace' to gracefully handle malformed UTF-8
  • Binary resources are represented as [binary resource uri=... mimeType=... base64_size=...] to provide context without attempting text conversion
  • Resource links are formatted as [resource_link label: uri] or [resource_link: uri] depending on availability of name/title

This resolves the issue where tools like GitHub's get_file_contents would appear to succeed but deliver no usable content to the LLM.

https://claude.ai/code/session_01M3RAEYpyaVQuRQifYUhwEf

Summary by CodeRabbit

  • New Features

    • MCP tools now extract text from additional content block types: embedded text, embedded resources, and resource links (including URI hints).
  • Bug Fixes

    • More reliable assembly of text from MCP tool results, reducing missed or incomplete text in outputs.
  • Tests

    • Added unit tests covering embedded text, base64-decoded text resources, binary placeholders, and resource link URIs.

_invoke_async only collected TextContent blocks, dropping any
EmbeddedResource the server returned. The github MCP server's
get_file_contents returns the file body inside a ResourceContents
(EmbeddedResource), so Holmes saw only the "successfully downloaded
text file (SHA: ...)" preamble and the LLM never received the file
content. Same problem for files >= 1MB, which come back as a
ResourceLink, and for any other MCP server using the resource pattern.

Now extract:
- TextResourceContents.text directly
- BlobResourceContents.blob (base64) decoded when mimeType is text-like
- ResourceLink uri/name as a hint

Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@coderabbitai

coderabbitai Bot commented Apr 28, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

Update MCP tool result extraction to support content block types text, resource, and resource_link via a new static extractor that merges embedded text, base64-decodes text-like blobs, emits placeholders for binary blobs, and includes resource URIs for downstream checks.

Changes

Cohort / File(s) Summary
MCP Content Block Extraction
holmes/plugins/toolsets/mcp/toolset_mcp.py
Added RemoteMCPTool._extract_text_from_content_block() (static) and updated result text extraction to include text blocks, extract embedded text from resource blocks, base64-decode blob text-like MIME types, emit diagnostic placeholders for binary blobs, and include resource_link URIs.
MCP Tool Result Tests
tests/test_mcp_toolset.py
Added tests and helpers for _invoke_async content extraction covering: EmbeddedResource with TextResourceContents, BlobResourceContents decoded for text-like MIME types, binary blob handling (placeholder), and ResourceLink URI inclusion.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested labels

evals-tag-grafana

Suggested reviewers

  • moshemorad
  • arikalon1
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title accurately describes the main change: extracting text from MCP resource content blocks in tool results, which is the core bug fix and feature addition.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@naomi-robusta naomi-robusta changed the title Extract text from MCP resource content blocks in tool results ROB-3756 - Extract text from MCP resource content blocks in tool results Apr 28, 2026
@netlify

netlify Bot commented Apr 28, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 5b87478
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69f1e8850f8cc20008ebc2d2
😎 Deploy Preview https://deploy-preview-1961--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Apr 28, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 #3 · Run @ __96b0fe8__ (#25056178169) — Apr 28, 13:47 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 96b0fe8 on branch claude/fix-github-get-file-yw3Lx

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 51.7s 7 11 $0.3694 145,748 143,482 24,038 2,266 788 101,309 42,173 — —
✅ 101_loki_historical_logs_pod_deleted 77.5s 6 13 $0.4309 140,486 137,529 27,293 2,957 941 86,973 50,556 — —
✅ 112_find_pvcs_by_uuid 25.0s 3 3 $0.2643 57,717 56,835 20,171 882 510 19,636 37,199 — —
✅ 12_job_crashing 51.9s 6 16 $0.3936 141,851 139,145 27,233 2,706 623 94,460 44,685 — —
✅ 176_network_policy_blocking_traffic_no_skills 75.9s 6 15 $0.3006 140,212 137,332 26,855 2,880 842 109,245 28,087 — —
✅ 227_count_configmaps_per_namespace[0] 32.3s 5 9 $0.2004 95,199 93,937 20,943 1,262 585 72,667 21,270 — —
✅ 243_pod_names_contain_service 45.1s 5 10 $0.3345 102,953 100,829 22,921 2,124 604 60,573 40,256 — —
✅ 24_misconfigured_pvc 42.2s 6 15 $0.2649 125,290 122,872 23,900 2,418 858 97,721 25,151 — —
✅ 43_current_datetime_from_prompt 11.4s 1 — $0.1095 17,118 16,982 16,982 136 136 0 16,982 — —
✅ 51_logs_summarize_errors 30.5s 4 5 $0.2831 77,399 76,340 20,996 1,059 324 38,351 37,989 — —
✅ 61_exact_match_counting 13.6s 2 1 $0.2212 34,626 34,371 17,380 255 187 0 34,371 — —
Total 41.5s avg 4.6 avg 9.8 avg $3.1726 1,078,599 1,059,654 27,293 18,945 941 680,935 378,719 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 74 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #2 · Run @ __dd6e4c6__ (#25053290326) — Apr 28, 12:57 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit dd6e4c6 on branch claude/fix-github-get-file-yw3Lx

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 48.5s 6 11 $0.2598 127,251 125,017 23,758 2,234 777 100,176 24,841 — —
✅ 101_loki_historical_logs_pod_deleted 56.6s 5 10 $0.3742 108,098 105,676 24,239 2,422 858 59,876 45,800 — —
✅ 112_find_pvcs_by_uuid 30.0s 5 4 $0.1975 95,376 94,145 20,796 1,231 323 73,336 20,809 — —
✅ 12_job_crashing 47.1s 6 13 $0.3641 136,181 133,983 25,128 2,198 580 91,857 42,126 — —
✅ 176_network_policy_blocking_traffic_no_skills 70.3s 7 18 $0.4337 162,235 158,979 27,871 3,256 728 111,389 47,590 — —
✅ 227_count_configmaps_per_namespace[0] 29.5s 4 9 $0.1952 77,578 76,338 20,944 1,240 584 54,150 22,188 — —
✅ 243_pod_names_contain_service 36.1s 4 7 $0.2035 78,526 76,917 21,270 1,609 632 55,078 21,839 — —
✅ 24_misconfigured_pvc 40.9s 6 13 $0.2633 123,920 121,593 23,517 2,327 563 96,045 25,548 — —
✅ 43_current_datetime_from_prompt 9.9s 1 — $0.1091 17,100 16,982 16,982 118 118 0 16,982 — —
✅ 51_logs_summarize_errors 32.7s 4 5 $0.1907 77,916 76,693 21,175 1,223 468 55,506 21,187 — —
✅ 61_exact_match_counting 18.0s 3 3 $0.1391 53,018 52,641 17,971 377 230 34,659 17,982 — —
Total 38.1s avg 4.6 avg 9.3 avg $2.7303 1,057,199 1,038,964 27,871 18,235 858 732,072 306,892 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 74 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #1 · Run @ __6448abe__ (#25051867509) — Apr 28, 12:17 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 6448abe on branch claude/fix-github-get-file-yw3Lx

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 37.4s 5 11 $0.2501 105,244 103,041 23,810 2,203 941 77,765 25,276 — —
✅ 101_loki_historical_logs_pod_deleted 57.7s 7 15 $0.3124 159,428 156,143 25,588 3,285 881 129,428 26,715 — —
✅ 112_find_pvcs_by_uuid 17.8s 3 3 $0.1665 57,640 56,746 20,127 894 491 36,608 20,138 — —
✅ 12_job_crashing 44.4s 6 15 $0.2938 136,833 134,154 25,842 2,679 833 105,834 28,320 — —
✅ 176_network_policy_blocking_traffic_no_skills 46.0s 6 14 $0.2997 137,849 135,226 26,218 2,623 667 105,502 29,724 — —
✅ 227_count_configmaps_per_namespace[0] 25.5s 5 9 $0.2003 95,206 93,951 20,949 1,255 585 72,671 21,280 — —
✅ 243_pod_names_contain_service 33.7s 5 8 $0.2174 98,864 97,219 21,671 1,645 537 74,872 22,347 — —
✅ 24_misconfigured_pvc 36.2s 5 12 $0.2383 103,400 101,306 22,598 2,094 642 77,522 23,784 — —
✅ 43_current_datetime_from_prompt 4.7s 1 — $0.1091 17,101 16,982 16,982 119 119 0 16,982 — —
✅ 51_logs_summarize_errors 23.4s 4 5 $0.1885 78,020 76,914 21,283 1,106 350 55,619 21,295 — —
✅ 61_exact_match_counting 11.6s 3 3 $0.1389 52,996 52,627 17,966 369 222 34,650 17,977 — —
Total 30.8s avg 4.5 avg 9.5 avg $2.4150 1,042,581 1,024,309 26,218 18,272 941 770,471 253,838 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 74 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit 5b87478 on branch claude/fix-github-get-file-yw3Lx

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 38.9s 6 11 $0.2568 127,741 125,552 23,718 2,189 660 101,194 24,358 — —
✅ 101_loki_historical_logs_pod_deleted 42.1s 5 9 $0.2420 104,661 102,414 22,866 2,247 913 78,851 23,563 — —
✅ 112_find_pvcs_by_uuid 19.3s 3 3 $0.1775 60,510 59,561 21,560 949 508 37,990 21,571 — —
✅ 12_job_crashing 35.8s 5 12 $0.2468 110,817 108,890 24,563 1,927 577 83,668 25,222 — —
✅ 176_network_policy_blocking_traffic_no_skills 58.4s 7 16 $0.3202 163,360 160,364 27,936 2,996 958 131,434 28,930 — —
✅ 227_count_configmaps_per_namespace[0] 24.1s 5 9 $0.1982 95,114 93,879 20,919 1,235 576 72,947 20,932 — —
✅ 243_pod_names_contain_service 33.0s 5 8 $0.2191 99,504 97,802 21,695 1,702 419 75,454 22,348 — —
✅ 24_misconfigured_pvc 41.0s 6 14 $0.2567 125,857 123,516 23,352 2,341 592 99,695 23,821 — —
✅ 43_current_datetime_from_prompt 5.0s 1 — $0.1094 17,111 16,982 16,982 129 129 0 16,982 — —
✅ 51_logs_summarize_errors 23.1s 4 5 $0.1867 77,513 76,415 21,036 1,098 347 55,367 21,048 — —
✅ 61_exact_match_counting 8.1s 2 1 $0.1231 34,601 34,359 17,368 242 174 16,981 17,378 — —
Total 29.9s avg 4.5 avg 8.8 avg $2.3364 1,016,789 999,734 27,936 17,055 958 753,581 246,153 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 74 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-github-get-file-yw3Lx -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-github-get-file-yw3Lx -f markers=regression -f filter=

@github-actions

github-actions Bot commented Apr 28, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for ce1f3796 (built in 5m 10s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:ce1f3796
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:ce1f3796 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:ce1f3796
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:ce1f3796
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:ce1f3796
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:ce1f3796 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:ce1f3796
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:ce1f3796

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:ce1f3796 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:ce1f3796

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:ce1f3796 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:ce1f3796

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
tests/test_mcp_toolset.py (1)

1232-1232: Move base64 imports to module scope

These function-local imports should be hoisted to the top-level import block.

♻️ Suggested cleanup
 import asyncio
+import base64
 import copy
 import logging
@@
-        import base64 as _b64
-
         file_body = '{"hello": "world"}'
-        encoded = _b64.b64encode(file_body.encode("utf-8")).decode("ascii")
+        encoded = base64.b64encode(file_body.encode("utf-8")).decode("ascii")
@@
-        import base64 as _b64
-
-        encoded = _b64.b64encode(b"\x89PNG\r\n\x1a\n").decode("ascii")
+        encoded = base64.b64encode(b"\x89PNG\r\n\x1a\n").decode("ascii")

As per coding guidelines "ALWAYS place Python imports at the top of the file, not inside functions or methods".

Also applies to: 1256-1256

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/test_mcp_toolset.py` at line 1232, Hoist the function-local "import
base64 as _b64" statements into the module-level import block (keeping the alias
_b64) and remove the local imports inside the test functions; ensure the
top-level imports appear with the other imports and that any references to _b64
in the functions still work without local imports.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@holmes/plugins/toolsets/mcp/toolset_mcp.py`:
- Around line 471-475: The code currently fully decodes text-like base64 blobs
(base64.b64decode(blob).decode("utf-8", errors="replace")) and can return
arbitrarily large strings; add a decode-size guard (e.g., MAX_DECODE_BYTES) and
only decode and return up to that limit, include a clear truncation indicator
and the original size/uri/mime in the returned string, and for blobs larger than
the limit return a short summary like the existing "[binary resource ...]"
message with base64_size and a note that the decoded text was truncated; update
the handling around the base64.b64decode(...).decode(...) call and the fallback
return that uses uri, mime, and len(blob) to reflect truncation.
- Around line 476-482: The code returns raw resource_link URIs which can expose
sensitive query params; update the resource_link branch (the block_type ==
"resource_link" code that uses variables uri, name, title, label) to redact
query strings and fragments before returning the URI — e.g., parse the uri and
drop any query and fragment components (or split at '?'/'#') so only the
scheme/host/path remain, then construct the return string using that sanitized
URI and the existing label logic.

---

Nitpick comments:
In `@tests/test_mcp_toolset.py`:
- Line 1232: Hoist the function-local "import base64 as _b64" statements into
the module-level import block (keeping the alias _b64) and remove the local
imports inside the test functions; ensure the top-level imports appear with the
other imports and that any references to _b64 in the functions still work
without local imports.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 319e96eb-5072-4e5a-989d-78d9ca75ea7f

📥 Commits

Reviewing files that changed from the base of the PR and between d52ad42 and 6448abe.

📒 Files selected for processing (2)
  • holmes/plugins/toolsets/mcp/toolset_mcp.py
  • tests/test_mcp_toolset.py

Comment thread holmes/plugins/toolsets/mcp/toolset_mcp.py
Comment thread holmes/plugins/toolsets/mcp/toolset_mcp.py
@github-actions

github-actions Bot commented Apr 28, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🔴 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 12.75s 10.85s +17.5%
Warm Mean 5.70s 4.65s +22.4%
Warm Min 5.64s 4.57s
Warm Max 5.74s 4.74s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 19.28s 15.16s +27.2%
Warm Mean 7.11s 6.66s +6.8%
Warm Min 7.05s 6.32s
Warm Max 7.27s 7.47s

PR: ce1f3796 | Master: cade9596 | Iterations: 5

CLAUDE.md requires imports at the top of the file.

Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@tests/test_mcp_toolset.py`:
- Around line 1251-1294: Update the two tests to assert the exact placeholder
formats rather than loose substring checks: in
test_invoke_async_keeps_binary_blob_as_placeholder, replace the two loose
asserts against result.data with a single assertion that result.data equals (or
contains) the full binary placeholder emitted by the extractor (include the
exact placeholder token + the MIME and the resource URI as produced by the code
that handles BlobResourceContents/EmbeddedResource); in
test_invoke_async_surfaces_resource_link, replace the loose URI and message
checks with an assertion that result.data contains the exact ResourceLink
placeholder format (include the URI and the filename/name field from
ResourceLink and the surrounding structured wrapper the extractor emits). Locate
these tests by name (_run_invoke_with_content, BlobResourceContents,
ResourceLink, EmbeddedResource, StructuredToolResultStatus) and assert the full
expected placeholder strings instead of partial substrings.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 2fd20d43-2452-423e-a5d3-458f17c52321

📥 Commits

Reviewing files that changed from the base of the PR and between 6448abe and dd6e4c6.

📒 Files selected for processing (1)
  • tests/test_mcp_toolset.py

Comment thread tests/test_mcp_toolset.py
Substring checks in the binary-blob and resource_link tests would pass
even if the wrapper format degraded: the resource_link test was
asserting "too large to display" — text emitted by the upstream
TextContent block, not the resource_link extractor — so a broken
[resource_link <name>: <uri>] wrapper would not have been caught.
Tighten both to full-string equality.

Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
tests/test_mcp_toolset.py (1)

2-2: Use a more descriptive base64 import alias.

_b64 is terse for a module alias in test code; base64 (or b64) would be clearer and still concise.

As per coding guidelines "Use semantic, descriptive names for variables, functions, and components".

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/test_mcp_toolset.py` at line 2, The import alias `_b64` in
tests/test_mcp_toolset.py is non-descriptive; replace the import "import base64
as _b64" with a clearer alias such as "import base64" or "import base64 as b64"
and then update every usage of `_b64` in the file to the new name (e.g.,
base64.b64encode/base64.b64decode or b64.b64encode/b64.b64decode) so all
references match the updated import.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@tests/test_mcp_toolset.py`:
- Line 2: The import alias `_b64` in tests/test_mcp_toolset.py is
non-descriptive; replace the import "import base64 as _b64" with a clearer alias
such as "import base64" or "import base64 as b64" and then update every usage of
`_b64` in the file to the new name (e.g., base64.b64encode/base64.b64decode or
b64.b64encode/b64.b64decode) so all references match the updated import.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 6c82d630-ccca-4e97-b52b-e9f20657d251

📥 Commits

Reviewing files that changed from the base of the PR and between dd6e4c6 and 96b0fe8.

📒 Files selected for processing (1)
  • tests/test_mcp_toolset.py

@naomi-robusta

Copy link
Copy Markdown
Collaborator Author

us-central1-docker.pkg.dev/genuine-flight-317411/devel/holmes:github-mcp-get-file-content-fix

@naomi-robusta
naomi-robusta merged commit 22090df into master Apr 29, 2026
18 of 22 checks passed
@naomi-robusta
naomi-robusta deleted the claude/fix-github-get-file-yw3Lx branch April 29, 2026 11:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants