Skip to content

fix(security): safety layer bypass via output truncation [HIGH] - #1851

Merged
ilblackdragon merged 3 commits into
nearai:stagingfrom
lycheepuppy:fix/safety-layer-truncation-bypass
Apr 5, 2026
Merged

ilblackdragon merged 3 commits into
nearai:stagingfrom
lycheepuppy:fix/safety-layer-truncation-bypass

Conversation

@lycheepuppy

Copy link
Copy Markdown
Contributor

Summary

  • Fix safety layer bypass where sanitize_tool_output() returned immediately after truncating oversized tool output, skipping all three safety checks (leak detection, policy enforcement, injection scanning)
  • Restructured truncation path so truncated content flows through the full safety pipeline before being returned
  • Added regression tests proving truncated output is now scanned for injection patterns

Security Finding: Safety Layer Bypass via Output Truncation

Severity: HIGH
Reported by: FailSafe Security Researcher
Component: crates/ironclaw_safety/src/lib.rs — sanitize_tool_output()

Description

The sanitize_tool_output() method truncated oversized tool output and returned it immediately — before leak detection, policy enforcement, or injection scanning ever ran. This meant any tool output exceeding max_output_length was delivered to the LLM with zero safety checks applied.

Order-of-operations problem (before fix):

  1. Length check → if oversized, truncate to UTF-8 boundary
  2. Early return → truncated content returned with only a low-severity output_too_large warning
  3. Leak detection → SKIPPED for truncated output
  4. Policy enforcement → SKIPPED for truncated output
  5. Injection scanning → SKIPPED for truncated output

Attack Vector

A malicious tool (compromised MCP server, poisoned web page via HTTP tool, sandbox job processing attacker-controlled input) could craft output where:

  • The first N bytes (within the truncation boundary) contain prompt injection payloads, credential exfiltration instructions, or policy-violating content
  • The total output exceeds max_output_length, triggering the early-return path
  • The adversarial content reaches the LLM without any safety scanning

The attacker does not need to know the exact truncation threshold — they only need to ensure the total output is large enough while placing the payload near the beginning.

Impact

  • Any tool output exceeding max_output_length bypassed all three safety layers
  • Undermines the entire safety architecture — the safety layer is the single choke point for tool output
  • The bypass is trivially triggered by any tool that returns large output (web fetch, file read, shell command, MCP tool)

Fix

Removed the early return after truncation. Truncation now produces (content, was_modified, extra_warnings) and falls through to the same leak detection → policy enforcement → injection scanning pipeline as non-truncated content. Truncation warnings are preserved and merged into the final output.

Test plan

  • cargo test -p ironclaw_safety — existing adversarial truncation tests still pass
  • New test truncated_output_still_scanned_for_injection — verifies injection patterns are detected in oversized output
  • New test truncated_output_preserves_truncation_warning — verifies truncation warning survives the full pipeline
  • cargo clippy -p ironclaw_safety -- -D warnings — zero warnings

🤖 Generated with Claude Code

Previously, `sanitize_tool_output()` returned immediately after
truncating oversized output, skipping leak detection, policy
enforcement, and injection scanning entirely. This allowed an attacker
to embed malicious payloads in the first N bytes of oversized tool
output and have them delivered unsanitized to the LLM.

Restructure the truncation path so it feeds into the same safety
pipeline as non-truncated content: leak detection, policy checks,
and Aho-Corasick injection scanning all run on the (possibly
truncated) content before it is returned.

Adds regression tests to verify truncated output is still scanned.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@github-actions github-actions Bot added size: M 50-199 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: new First-time contributor labels Apr 1, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the sanitize_tool_output function to ensure that truncated output is still processed by subsequent safety layers, such as leak detection and injection scanning, instead of returning early. It also adds regression tests to confirm that security checks are applied to truncated content. Feedback was provided to correct an invalid range in the truncation warning that could cause out-of-bounds errors and to append low-severity warnings to the end of the list to maintain proper sorting.

Comment thread crates/ironclaw_safety/src/lib.rs Outdated
vec![InjectionWarning {
pattern: "output_too_large".to_string(),
severity: Severity::Low,
location: 0..output.len(),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The location range 0..output.len() is incorrect because it refers to the original, untruncated string. Since sanitize_tool_output returns a modified string (truncated prefix plus notice), any consumer using this range will encounter an out-of-bounds panic. Additionally, when truncating UTF-8 strings, ensure the split occurs at a valid character boundary using is_char_boundary to prevent panics on multi-byte characters.

Suggested change
location: 0..output.len(),
location: cut..(cut + notice.len()),
References
  1. When truncating a UTF-8 string at a byte boundary, walk backwards from the desired length until a valid character boundary is found using is_char_boundary to prevent panics.
  2. Always truncate tool output for previews or status updates to a reasonable maximum length to prevent excessive memory/bandwidth usage.

Comment on lines +128 to +129
extra_warnings.append(&mut sanitized.warnings);
sanitized.warnings = extra_warnings;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Merging extra_warnings at the beginning of the list breaks the severity-based sort order (Critical/High first) established by the Sanitizer. Since the truncation warning is Severity::Low, it should be appended to the end of the sanitized.warnings list. This ensures that more urgent security warnings remain at the top of the list for the caller. Additionally, using extend is more idiomatic than the current append and reassignment pattern.

Suggested change
extra_warnings.append(&mut sanitized.warnings);
sanitized.warnings = extra_warnings;
sanitized.warnings.extend(extra_warnings);

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review -- APPROVE

Legitimate HIGH severity security fix. The vulnerability is real: sanitize_tool_output() has an early return after truncation that skips leak detection, policy enforcement, and injection scanning. An attacker only needs to pad output past max_output_length while placing malicious content in the first N bytes.

The fix is correct -- truncation now produces a tuple that falls through to the full safety pipeline. Truncated content is scanned (which is right, since it's what the LLM sees). Warning merging preserves both truncation and downstream warnings.

No bypass vectors in the fix. Low regression risk -- non-truncated output path is unchanged.

Suggestions (non-blocking)

  • The policy Block path discards all warnings including truncation warning -- consider preserving for logging (pre-existing issue)
  • Consider adding a test where truncated output contains a leaked secret pattern (currently only injection scanning is tested)

zmanian
zmanian previously approved these changes Apr 3, 2026

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review -- APPROVE (post-approval commit check)

The new commit (4e2ff0b) is a merge of staging into the feature branch. It brings in unrelated staging changes (channels, config, e2e tests, github tool, etc.) but does not touch crates/ironclaw_safety/ at all -- zero files in the safety crate were modified by the merge.

The security fix from the original commit (a3fa6df) is intact and unchanged:

  • Truncation no longer early-returns; content flows through the full safety pipeline
  • Warning merging preserves truncation warnings alongside downstream findings
  • Regression tests cover injection scanning on truncated output

Gemini's comments (not addressed, both non-blocking)

  1. location: 0..output.len() range -- refers to original string length, not truncated content. Pre-existing pattern, not introduced by this PR. Low risk since location is metadata on the warning struct, not used for content slicing.

  2. Warning ordering -- truncation warnings (Low severity) end up before injection warnings (higher severity) due to the append + reassign pattern. Cosmetic; callers should sort/filter by severity if ordering matters.

Neither weakens the security fix. Safe to merge.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@ilblackdragon
ilblackdragon merged commit f303638 into nearai:staging Apr 5, 2026
14 checks passed
@claude

claude Bot commented Apr 5, 2026

Copy link
Copy Markdown

Code review

Found 2 issues:

  1. [HIGH:95] Truncation warnings may be silently lost due to append order at lines 127-128

https://github.com/anthropics/ironclaw/blob/e8b1989ae24bb6917f1e14478dc60ad6cee7df55/crates/ironclaw_safety/src/lib.rs#L124-L129

         if self.config.injection_check_enabled || force_sanitize {
             let mut sanitized = self.sanitizer.sanitize(&content);
             sanitized.was_modified = sanitized.was_modified || was_modified;
+            extra_warnings.append(&mut sanitized.warnings);
+            sanitized.warnings = extra_warnings;
             sanitized

The code appends sanitized.warnings into extra_warnings, then replaces sanitized.warnings with the merged vector. This loses truncation warnings when no injection patterns are detected in the branch below. The append should be reversed: sanitized.warnings.append(&mut extra_warnings) to prepend truncation warnings to injection warnings in the correct order.

  1. [MEDIUM:70] Missing documentation of was_modified contract in bypass path

https://github.com/anthropics/ironclaw/blob/e8b1989ae24bb6917f1e14478dc60ad6cee7df55/crates/ironclaw_safety/src/lib.rs#L130-L140

         } else {
             SanitizedOutput {
                 content,
-                warnings: vec![],
+                warnings: extra_warnings,
                 was_modified,
             }
         }

When injection_check_enabled=false, the was_modified flag reflects only truncation and policy actions, not injection pattern modifications. This coupling is fragile for future changes. Consider documenting: "was_modified reflects truncation and policy actions only, not injection pattern modification when injection_check is disabled."


Security & Safety: ✅ No issues — fix correctly prevents truncation from bypassing all downstream checks (leak detection, policy, injection scanning).

Bug Scan: ✅ No obvious logic errors — refactoring preserves safety check execution order.

Performance: Append order at lines 127-128 could be optimized (see issue #1 above).

serrrfirat pushed a commit that referenced this pull request Apr 5, 2026
* fix(security): run safety checks on truncated tool output

Previously, `sanitize_tool_output()` returned immediately after
truncating oversized output, skipping leak detection, policy
enforcement, and injection scanning entirely. This allowed an attacker
to embed malicious payloads in the first N bytes of oversized tool
output and have them delivered unsanitized to the LLM.

Restructure the truncation path so it feeds into the same safety
pipeline as non-truncated content: leak detection, policy checks,
and Aho-Corasick injection scanning all run on the (possibly
truncated) content before it is returned.

Adds regression tests to verify truncated output is still scanned.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: apply rustfmt to fix CI formatting check

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Wui <wui@Wui-Work-2.local>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
…ai#1851)

* fix(security): run safety checks on truncated tool output

Previously, `sanitize_tool_output()` returned immediately after
truncating oversized output, skipping leak detection, policy
enforcement, and injection scanning entirely. This allowed an attacker
to embed malicious payloads in the first N bytes of oversized tool
output and have them delivered unsanitized to the LLM.

Restructure the truncation path so it feeds into the same safety
pipeline as non-truncated content: leak detection, policy checks,
and Aho-Corasick injection scanning all run on the (possibly
truncated) content before it is returned.

Adds regression tests to verify truncated output is still scanned.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: apply rustfmt to fix CI formatting check

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Wui <wui@Wui-Work-2.local>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
@ironclaw-ci ironclaw-ci Bot mentioned this pull request Apr 10, 2026
JZKK720 pushed a commit to JZKK720/ironclaw that referenced this pull request Apr 13, 2026
…ai#1851)

* fix(security): run safety checks on truncated tool output

Previously, `sanitize_tool_output()` returned immediately after
truncating oversized output, skipping leak detection, policy
enforcement, and injection scanning entirely. This allowed an attacker
to embed malicious payloads in the first N bytes of oversized tool
output and have them delivered unsanitized to the LLM.

Restructure the truncation path so it feeds into the same safety
pipeline as non-truncated content: leak detection, policy checks,
and Aho-Corasick injection scanning all run on the (possibly
truncated) content before it is returned.

Adds regression tests to verify truncated output is still scanned.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: apply rustfmt to fix CI formatting check

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Wui <wui@Wui-Work-2.local>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
(cherry picked from commit f303638)
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
…ai#1851)

* fix(security): run safety checks on truncated tool output

Previously, `sanitize_tool_output()` returned immediately after
truncating oversized output, skipping leak detection, policy
enforcement, and injection scanning entirely. This allowed an attacker
to embed malicious payloads in the first N bytes of oversized tool
output and have them delivered unsanitized to the LLM.

Restructure the truncation path so it feeds into the same safety
pipeline as non-truncated content: leak detection, policy checks,
and Aho-Corasick injection scanning all run on the (possibly
truncated) content before it is returned.

Adds regression tests to verify truncated output is still scanned.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: apply rustfmt to fix CI formatting check

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Wui <wui@Wui-Work-2.local>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: new First-time contributor risk: low Changes to docs, tests, or low-risk modules size: M 50-199 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants