fix(security): scan the text a tool_result carries - #13101
Merged
diegosouzapw merged 2 commits intoSep 10, 2026
Merged
diegosouzapw merged 2 commits into
diegosouzapw merged 2 commits into
Conversation
extractMessageContents() reads a content part's `text`. A tool_result block does not have one: it carries its payload on `content`, as a string or as a nested block list. The repo's own Claude translator reads exactly that (`providers/xai/translators/claude.ts`), and redactBody() in this same file already rewrites the string form -- only the extractor never looked. So nothing the model was handed back by a tool was ever scanned. Measured on release/v3.8.51 with one payload in three carriers: plain text block -> 2 detections tool_result (string) -> 0, nothing extracted tool_result (blocks) -> 0, nothing extracted That is the carrier that matters most for this guard: user text is written by the caller, while tool output is fetched from outside the conversation, which is where indirect prompt injection arrives from. Collect a part's `content` alongside its `text`, in messages and in system blocks, and let redactBody() reach a nested block list too -- redaction only runs once detection has fired, so the two have to see the same bytes or PII would be detected, logged, and forwarded anyway. Issue diegosouzapw#8094 closed the coverage holes it listed (prompt, instructions, query, documents); the tool_result carrier was not among them.
Contributor
Author
|
One note on the That file arrived with #12637 ( The fragment this PR adds ( |
This was referenced Sep 9, 2026
diegosouzapw
merged commit Sep 10, 2026
567abb5
into
diegosouzapw:release/v3.8.51
9 of 16 checks passed
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
Boarded with 13 sibling PRs into one worktree off release/v3.8.51 and validated as a set: 132 focused tests pass across all 15 test files in the batch, typecheck:core is clean, check-changelog-integrity reports no lost base bullets, and check-file-size is green. Your PR merged without conflict against its siblings. Thank you — the write-up made this reviewable: measuring the behaviour on the release tip and showing the before/after table meant the defect could be confirmed rather than taken on faith.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The bug
extractMessageContents()reads a content part'stext. Atool_resultblock does not have one — it carries its payload oncontent, as a string or as a nested block list:So nothing a tool handed back to the model was ever scanned.
detectInjection()runs on the joined output of that extractor, and the extractor returned an empty list.Measured on
release/v3.8.51One payload, three carriers, same request otherwise:
{ type: "text", text }system_override,system_prompt_leak, both high){ type: "tool_result", content: "…" }[]{ type: "tool_result", content: [{ type: "text", text }] }[]This is the carrier that matters most for this particular guard. User text is written by the caller; tool output is fetched from outside the conversation, which is where indirect prompt injection arrives from. With
INPUT_SANITIZER_MODE=block, the same sentence is a 400 in a user turn and a pass-through in tool output.The file already believed a part can carry text there
redactBody(), sixty lines down in the same file, rewrites it:Only the extractor never looked. And because redaction runs only after detection fires, that branch could not do anything on its own: PII inside a tool result was neither detected nor redacted.
The change
collectPartText()— one helper, so the extractor andredactBody()agree on what a part carries. It takestextandcontent(string, or a nested list of strings/{text}blocks).systemblocks —redactBody()already rewrotecontenton both.redactBody()now also rewrites a nested block list, so detection and redaction reach the same bytes. Without that half, the new detection would report PII, log it, and still forward it.No pattern changes, no new carriers beyond the one shape, and
MAX_INJECTION_SCAN_BYTESstill caps the scan.Tests
tests/unit/guardrails/injection-extraction-tool-result.test.ts, 10 tests, extending the coverage style ofinjection-extraction.test.ts. They run the real pipeline throughsanitizeRequest, not just the extractor, and include a no-duplicate test and a "part with neither field" test so the helper cannot pass by over-collecting.Three mutations, each killing a different set:
contentredactBodystops rewriting a nested block listExisting suites —
guardrails/injection-extraction,guardrails/injection-route-coverage,env-docs-input-sanitizer-8093,chatcore-sanitization,guardrails-registry,injection-guard-nonchat-route-logging: 45 passed.eslintclean; the pre-commit gates (docs-sync, any-budget, tracked-artifacts) all pass.Relation to earlier work
Issue #8094 closed the
redactBody()coverage holes it listed —prompt,instructions,query,documents. Thetool_resultcarrier was not among them, and it is the only one whose bytes originate outside the conversation.