Skip to content

fix: sanitize tool error results before llm injection - #1639

Merged
serrrfirat merged 7 commits into
stagingfrom
fix/tool-error-sanitization-staging
Mar 27, 2026
Merged

serrrfirat merged 7 commits into
stagingfrom
fix/tool-error-sanitization-staging

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

  • route failed tool results through the same sanitize-and-wrap safety path as successful tool outputs
  • preserve wrapped tool errors when rebuilding LLM context from persisted tool-call history
  • unwrap wrapped tool errors for web history display and avoid an extra allocation on successful tool results

Verification

  • cargo test test_rebuild_chat_messages_preserves_wrapped_tool_error --lib
  • cargo test test_build_turns_unwrap_wrapped_tool_error_for_display --lib
  • cargo test test_process_tool_result_error --lib
  • cargo test test_tool_error_format_includes_tool_name --lib
  • cargo check --lib

@github-actions github-actions Bot added scope: agent Agent core (agent loop, router, scheduler) scope: channel/web Web gateway channel size: L 200-499 changed lines risk: medium Business logic, config, or moderate-risk modules contributor: experienced 6-19 merged PRs labels Mar 25, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request enhances the robustness and security of tool output handling by ensuring that all tool results, including errors, undergo a consistent sanitization and wrapping process before being presented to the LLM. It also refines the display of these errors in the web UI and ensures their proper persistence and retrieval from history, mitigating potential LLM injection vulnerabilities and improving user experience.

Highlights

  • Consistent Error Sanitization: Failed tool results are now routed through the same sanitize-and-wrap safety path as successful tool outputs, preventing LLM injection vulnerabilities.
  • Preserved Wrapped Errors: Wrapped tool errors are now correctly preserved when rebuilding the LLM context from persisted tool-call history.
  • Improved Web Display: Wrapped tool errors are unwrapped for web history display, providing a cleaner user experience and avoiding extra allocations for successful tool results.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors tool result and error processing to ensure consistent formatting and safety for LLM consumption and UI display. It introduces a centralized process_tool_result function that sanitizes and wraps both successful tool outputs and errors. The changes also update the ChatDelegate to utilize this new processing, and modify the database message rebuilding and UI display logic to correctly handle and unwrap these standardized tool error messages. A review comment suggests simplifying the error handling in rebuild_chat_messages_from_db to avoid redundant "Error: " prefixes for legacy errors and improve consistency.

Comment thread src/agent/thread_ops.rs Outdated
Comment on lines +1856 to +1860
if ironclaw_safety::SafetyLayer::unwrap_tool_output(err).is_some() {
err.to_string()
} else {
format!("Error: {}", err)
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

For consistency with how new, wrapped errors are handled, it would be better to use the legacy error string from the database as-is, rather than prepending Error: .

The legacy error strings stored in the database (from result_content before this PR) were already descriptive messages like Tool 'http' failed: .... Prepending Error: makes it redundant (e.g., Error: Tool 'http' failed: ...).

Using err.to_string() for the else branch would make the handling of legacy errors consistent with the unwrapped content of new errors. With this change, both branches of the if become identical, so the entire conditional can be simplified.

                                err.to_string()
References
  1. Avoid coupling log messages to implementation details. The underlying error message should provide sufficient context on its own, making redundant prefixes like 'Error: ' unnecessary.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch — addressed in bd0d605. Simplified the conditional to err.to_string() for both branches since legacy errors already contain descriptive text. Updated the test assertion accordingly.

- Simplify legacy error handling in rebuild_chat_messages_from_db:
  remove redundant "Error: " prefix since legacy errors already contain
  descriptive text (e.g. "Tool 'http' failed: timeout"). Both wrapped
  (new) and plain (legacy) errors now pass through as-is.
- Update existing test assertion to match simplified format.
- Restore error-path doc line on process_tool_result.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
zmanian
zmanian previously approved these changes Mar 25, 2026

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: fix: sanitize tool error results before LLM injection

Good security fix. The prompt injection vector via unsanitized tool error messages is real -- an attacker-controlled error could inject </tool_output><system>...</system> to break the safety boundary. Unifying error and success paths through process_tool_result() is the right approach.

Low

  1. Missed error path in builder/core.rs:746-749: The builder's tool execution loop still formats errors raw without sanitization. Lower risk (sandboxed context), but inconsistent with the PR's goal.

  2. Duplicated logic in routine_engine.rs: The routine engine does its own inline sanitize+wrap. Not a vulnerability (it does sanitize), but should use process_tool_result() for consistency.

Positive

  • Good backward compat handling for old unwrapped error formats in DB round-trip
  • tool_error_for_display() correctly strips XML wrapping for web UI
  • Strong test coverage including direct injection attack test
  • PreflightOutcome enum moved to module scope is cleaner

Approve. Consider addressing the builder path in this PR or fast follow-up.

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: tool error sanitization

Primary finding (builder/core.rs unsanitized path) is fixed -- process_builder_tool_result now delegates to process_tool_result with proper sanitize+wrap. Tests confirm injection neutralization.

One LOW item remains: routine_engine.rs still has inline sanitize+wrap logic rather than calling process_tool_result. Functionally correct (does sanitize) but a maintenance divergence risk. Fine as follow-up.

Approve.

@serrrfirat
serrrfirat merged commit 2f4eb08 into staging Mar 27, 2026
14 checks passed
@serrrfirat
serrrfirat deleted the fix/tool-error-sanitization-staging branch March 27, 2026 07:49
DougAnderson444 pushed a commit to DougAnderson444/ironclaw that referenced this pull request Mar 29, 2026
* fix: sanitize tool error results before llm injection

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* fix: wrap preflight tool rejection errors for llm safety

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* style: apply rustfmt to error-path regressions

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* fix: preserve wrapped tool errors in history replay

* fix: address review findings on PR nearai#1639

- Simplify legacy error handling in rebuild_chat_messages_from_db:
  remove redundant "Error: " prefix since legacy errors already contain
  descriptive text (e.g. "Tool 'http' failed: timeout"). Both wrapped
  (new) and plain (legacy) errors now pass through as-is.
- Update existing test assertion to match simplified format.
- Restore error-path doc line on process_tool_result.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: satisfy clippy on builder tool safety helper

---------

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
* fix: sanitize tool error results before llm injection

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* fix: wrap preflight tool rejection errors for llm safety

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* style: apply rustfmt to error-path regressions

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* fix: preserve wrapped tool errors in history replay

* fix: address review findings on PR nearai#1639

- Simplify legacy error handling in rebuild_chat_messages_from_db:
  remove redundant "Error: " prefix since legacy errors already contain
  descriptive text (e.g. "Tool 'http' failed: timeout"). Both wrapped
  (new) and plain (legacy) errors now pass through as-is.
- Update existing test assertion to match simplified format.
- Restore error-path doc line on process_tool_result.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: satisfy clippy on builder tool safety helper

---------

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: experienced 6-19 merged PRs risk: medium Business logic, config, or moderate-risk modules scope: agent Agent core (agent loop, router, scheduler) scope: channel/web Web gateway channel scope: tool/builder Dynamic tool builder size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants