Skip to content

fix(llm): filter XML tool-call recovery by context - #1641

Merged
serrrfirat merged 2 commits into
stagingfrom
fix/xml-tool-call-recovery-context-filtering-staging
Mar 26, 2026
Merged

serrrfirat merged 2 commits into
stagingfrom
fix/xml-tool-call-recovery-context-filtering-staging

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Mar 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • prevent recover_tool_calls_from_content from converting XML tool-call examples inside fenced code blocks, inline code, or markdown blockquotes into real tool calls
  • keep XML recovery working for legitimate standalone tool-call payloads emitted by models that use text-based tool calling
  • add regression tests covering ignored quoted/code contexts and a preserved multiline positive case

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

  • None

Validation

  • cargo fmt
  • cargo clippy --all --benches --tests --examples --all-features
  • Relevant tests pass: cargo test test_recover_ --lib, cargo test llm::reasoning --lib
  • Manual testing: None

Security Impact

  • Hardens text-based tool-call recovery so quoted or code-formatted attacker-controlled XML is not reinterpreted as executable tool calls. This only affects the LLM tool recovery path in src/llm/reasoning.rs.

Database Impact

  • None

Blast Radius

  • src/llm/reasoning.rs XML tool-call recovery and its unit tests.
  • Could affect local-model compatibility only if a model emits XML tool calls inline with other prose on the same line rather than as isolated tool-call blocks. Existing standalone recovery behavior remains covered by tests.

Rollback Plan

  • Revert commit 8724178 to restore the previous recovery behavior.

Review track: C (security/runtime/DB/CI)

@github-actions github-actions Bot added scope: llm LLM integration size: M 50-199 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: experienced 6-19 merged PRs labels Mar 25, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the robustness and security of the LLM tool-call recovery mechanism. It introduces logic to prevent the system from inadvertently parsing XML-formatted tool calls that appear within markdown code blocks, inline code, or blockquotes. This change ensures that only legitimate, standalone tool-call payloads are processed, mitigating potential security risks associated with misinterpreting user-provided content as executable commands.

Highlights

  • Contextual Filtering for Tool Calls: Introduced new helper functions (overlaps_code_region, line_bounds, is_recoverable_tool_call_segment) to identify and filter out XML tool calls that appear within markdown code blocks, inline code, or blockquotes.
  • Enhanced Tool Call Recovery Logic: Modified the recover_tool_calls_from_content function to utilize this new filtering logic, preventing the misinterpretation of code examples or quoted snippets as executable tool calls.
  • Security Hardening: Enhanced security by hardening text-based tool-call recovery, ensuring that attacker-controlled XML in quoted or code-formatted contexts is not processed as executable commands.
  • Comprehensive Regression Tests: Added comprehensive regression tests to cover scenarios where tool calls should be ignored (fenced code, inline code, blockquotes) and to confirm that legitimate multiline tool calls are still correctly recovered.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request enhances the recover_tool_calls_from_content function by introducing new helper functions (overlaps_code_region, line_bounds, is_recoverable_tool_call_segment) to prevent the parsing of XML-style tool calls that appear within markdown code blocks, inline code, or blockquotes. This ensures that code examples or quoted snippets are not mistakenly interpreted as executable tool calls. New test cases have been added to validate these scenarios. Feedback indicates that the logic within is_recoverable_tool_call_segment for checking blockquotes and surrounding text is confusing and conflates concerns, suggesting a refactoring for improved clarity and robustness.

Comment thread src/llm/reasoning.rs Outdated
@serrrfirat
serrrfirat requested a review from zmanian March 25, 2026 13:50

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: fix(llm): filter XML tool-call recovery by context

Well-scoped security hardening. The core idea -- preventing XML tool-call recovery inside code blocks, inline code, and blockquotes -- is correct and the implementation is clean.

Low

  1. Blockquote check only examines first line: Multi-line tool calls could have the opening tag outside a blockquote. Acceptable in practice since the "isolated on own line" check already constrains the format tightly.

  2. Bracket-format recovery not filtered: The [Called tool ...] recovery at lines 1359-1402 doesn't go through is_recoverable_tool_call_segment. This seems intentional since bracket format comes from flatten_tool_messages (internal), not model output. Confirm this is deliberate.

Suggestions (non-blocking)

  • Add a test for inline-with-prose rejection (e.g., "Here is <tool_call>tool_list</tool_call>")
  • Consider unit tests for line_bounds directly -- it does byte-level string indexing where off-by-one errors hide

Good refactoring of the search loop from remaining slice to search_from offset. Test coverage hits the key cases. Approve.

@serrrfirat
serrrfirat merged commit b3fbef5 into staging Mar 26, 2026
14 checks passed
@serrrfirat
serrrfirat deleted the fix/xml-tool-call-recovery-context-filtering-staging branch March 26, 2026 06:38
bkutasi pushed a commit to bkutasi/ironclaw that referenced this pull request Mar 28, 2026
* fix(llm): filter XML tool-call recovery by context

* fix: address review comments on PR nearai#1641
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
* fix(llm): filter XML tool-call recovery by context

* fix: address review comments on PR nearai#1641
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: experienced 6-19 merged PRs risk: low Changes to docs, tests, or low-risk modules scope: llm LLM integration size: M 50-199 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants