Skip to content

fix(tools) raw tool-call output in routine Telegram notifications - #2033

Merged
ilblackdragon merged 1 commit into
nearai:stagingfrom
serrrfirat:firat/feat-routine-raw-tool-call-sxj
Apr 8, 2026
Merged

ilblackdragon merged 1 commit into
nearai:stagingfrom
serrrfirat:firat/feat-routine-raw-tool-call-sxj

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

  • strip leaked internal tool-call markers from lightweight routine summaries before persistence and notification delivery
  • steer lightweight routine prompts to return the primary notification as assistant text instead of using the message tool by default
  • add regression tests for both the prompt guidance and marker-only summary fallback

Closes #1995

Testing

  • cargo fmt --all -- --check
  • cargo test build_lightweight_prompt_explains_delivery_and_disabled_tools
  • cargo test handle_text_response_strips_internal_tool_markers
  • cargo test handle_text_response_replaces_marker_only_text

@github-actions github-actions Bot added scope: agent Agent core (agent loop, router, scheduler) size: M 50-199 changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Apr 5, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request enhances user-facing notifications by adding prompt instructions that guide the LLM to provide direct text responses and avoid using the message tool for primary delivery. It also introduces a strip_internal_tool_call_text function to remove internal execution markers from the output, along with comprehensive unit tests. A review comment suggests refactoring this stripping logic into a shared utility to avoid code duplication with the dispatcher module.

Comment on lines +1764 to +1767
!((trimmed.starts_with("[Called tool ") && trimmed.ends_with(']'))
|| (trimmed.starts_with("[Tool ")
&& trimmed.contains(" returned:")
&& trimmed.ends_with(']')))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The logic for stripping tool markers is duplicated in strip_internal_tool_call_text within src/agent/routine_engine.rs and src/agent/dispatcher.rs. This duplication should be refactored into a shared utility function to ensure consistency and maintainability. When implementing the shared utility, ensure that distinct parsing conditions are separated into individual checks to improve code clarity and robustness.

References
  1. Separate checks for distinct conditions to improve code clarity and robustness, particularly when parsing text.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 96a8c5d — extracted to src/agent/text_util.rs with separate is_called_tool_marker() / is_tool_result_marker() helpers. Duplicate tests in dispatcher.rs removed in 2af4fa1.

@serrrfirat serrrfirat changed the title Fix raw tool-call output in routine Telegram notifications fix(tools) raw tool-call output in routine Telegram notifications Apr 5, 2026

@ilblackdragon ilblackdragon left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated review — Needs changes (small)

TL;DR: Logic is correct and well-tested, but `strip_internal_tool_call_text()` is duplicated from `dispatcher.rs:1394`. Extract to a shared util before merge — marker format drift would silently break both copies.

Findings

# Severity File:Line Issue Suggestion
1 Medium `routine_engine.rs:1755` vs `dispatcher.rs:1394` `strip_internal_tool_call_text()` exists identically in both files; only the empty-text fallback string differs Extract to `src/agent/text_util.rs` (or `src/agent/util.rs`); caller passes the fallback string. Update `dispatcher.rs` to call the shared version.
2 Low `routine_engine.rs:1763-1768` Combined `starts_with("[Called tool ")` + `starts_with("[Tool ")` is dense (already flagged by gemini-code-assist) Split into `is_internal_call_marker(trimmed) -> bool` helper
3 Low `routine_engine.rs:1722` Prompt guidance correctly reserves the `message` tool for non-primary use cases No action
4 Very Low Tests No test for multi-line input with mixed markers and text Add `"Line1\n[Called tool]\nLine2"` case to verify `fold()` with `push('\n')` preserves line breaks

Behavioral changes

No unintended behavioral changes detected. Telegram output will no longer include raw `[Called tool \`http\`]` / `[Tool ... returned: ...]` markers, and the LLM is now guided to return text directly instead of wrapping everything in the `message` tool.

Open questions

  1. Dual fallback messages: Why do `dispatcher.rs:1416` ("I wasn't able to complete that request...") and `routine_engine.rs:1781` ("I wasn't able to produce a user-facing routine summary...") use different fallbacks? Intentional (job vs. lightweight routine context) or should be unified?
  2. Marker format stability: Is the `[Called tool ...]` / `[Tool ... returned: ...]` format stable? If it changes, both copies will break. A shared util with a clear contract reduces this risk.
  3. Prompt injection risk: Any concern that a malicious routine context could trick the LLM into outputting marker-shaped text that then gets stripped incorrectly? Low risk but worth a thought.

Copy link
Copy Markdown
Collaborator Author

Addressed the latest review feedback in 96a8c5da.

Changes:

  • extracted the duplicated tool-marker stripping logic into shared helper src/agent/text_util.rs
  • updated both dispatcher and routine-engine call sites to use the shared helper while preserving their distinct fallback messages
  • split the marker detection into explicit helper checks for readability
  • added multiline coverage for text mixed with internal tool markers

Validation:

  • cargo fmt --check
  • cargo test strip_internal_tool_call_text --lib --target-dir /tmp/pr2033-target
  • cargo test handle_text_response --lib --target-dir /tmp/pr2033-target

@serrrfirat
serrrfirat requested a review from ilblackdragon April 6, 2026 09:25
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Thanks for the review — all findings addressed:

Findings 1 & 2: Extracted to src/agent/text_util.rs with split is_called_tool_marker() / is_tool_result_marker() / is_internal_tool_marker() helpers (96a8c5d). Removed duplicate tests from dispatcher.rs (2af4fa1).

Finding 4: Added strip_internal_tool_call_text_preserves_multiline_text_around_markers test covering mixed markers and text with line breaks (96a8c5d).

Open questions:

  1. Dual fallback messages — intentional. Chat context ("Could you try rephrasing...") assumes interactive conversation; routine context ("I wasn't able to produce a user-facing routine summary") is for async/unattended delivery. Added a doc comment on the fallback param explaining this (2af4fa1).
  2. Marker format stability — the shared util is the single source of truth now, so format drift can't silently break one copy. If the marker format changes, only text_util.rs needs updating.
  3. Prompt injection risk — agreed it's low. The markers are structural patterns ([Called tool ... with balanced ]) that are unlikely to appear in natural user content. The worst case is legitimate text being stripped, not a security issue.

@ilblackdragon ilblackdragon left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Overview

Fixes #1995 where lightweight routines were sending raw internal [Called tool …] / [Tool … returned: …] markers to Telegram. Two-pronged fix:

  1. Prompt steering — build_lightweight_prompt now instructs the LLM to return the user-facing notification as plain assistant text rather than via the message tool.
  2. Defensive sanitization — handle_text_response runs a new strip_internal_tool_call_text filter that drops marker lines, falling back to a placeholder if everything was stripped.

Three regression tests cover the prompt change, the partial-strip path, and the marker-only fallback.

Issues

Significant — Code duplication of an existing helper

A virtually identical function already lives at src/agent/dispatcher.rs:1390 (fn strip_internal_tool_call_text). The new copy in src/agent/routine_engine.rs:1758 has the same line-filter logic and the same [Called tool …] / [Tool … returned: …] patterns — only the empty-fallback string differs:

  • dispatcher.rs:1416 → \"I wasn't able to complete that request. Could you try rephrasing…\"
  • routine_engine.rs:1779 → \"I wasn't able to produce a user-facing routine summary.\"

Per the project's review-discipline rule ("Fix the pattern, not just the instance"), this should be extracted into a shared helper (e.g. an pub(crate) function in src/agent/mod.rs or a small src/agent/text_sanitize.rs) that takes the fallback message as a parameter, with both call sites using it. Future patches to the marker grammar otherwise need to be made in two places — and this exact bracket grammar already shows up in src/llm/reasoning.rs:1597-1707, so the team has hit "two implementations is too many" before.

Minor — Filter only catches whole-line markers

Both copies of the function only strip a line if the trimmed line starts with [Called tool (or [Tool … returned:) and ends with ]. If the LLM emits the marker mid-line (Done. [Called tool http with arguments: {...}]), or if the JSON args span multiple lines so the closing ] lands on its own line, the marker survives. src/llm/reasoning.rs:1698-1707 already has a more thorough scanner — worth reusing in the shared helper rather than re-inventing the weaker line-based one.

Not a regression introduced by this PR (dispatcher.rs has the same gap), but worth noting since you're touching the area.

Minor — Fallback path still produces a notification

When the LLM emits only markers (the actual #1995 failure mode), every line is filtered, the fallback \"I wasn't able to produce a user-facing routine summary.\" is substituted, and the routine still completes with RunStatus::Attention — meaning the user is notified with a content-free apology instead of nothing. Consider whether returning RunStatus::Ok with None (treat marker-only output as "nothing useful to deliver") would be a better UX, or alternatively a failure status so the routine retry path kicks in. As written, this replaces one form of noise with a slightly nicer form of noise.

Nit — Re-allocations in handle_text_response

let content = strip_internal_tool_call_text(content); // alloc 1
let content = content.trim();                          // borrow
…
Some(content.to_string())                              // alloc 2

The trimmed &str borrow is fine, but strip_internal_tool_call_text could .trim() internally and return one String. Trivial.

What's good

  • Prompt change is the right primary fix — addressing root cause (model misuse of message tool) instead of only patching the symptom.
  • Test cases cover both the partial-strip and total-strip paths, and assert the specific fallback string.
  • RunStatus::Attention and token totals are preserved correctly through the new strip pass.
  • Regression tests are included alongside the fix, per the project's review-discipline rule.

Recommendation

The code-duplication issue is the only real blocker. Once strip_internal_tool_call_text is consolidated into a single helper (parameterized by fallback message) and dispatcher.rs:1394 switches to call it, this is good to merge. The "marker-only → fallback message vs. silent skip" question is worth a follow-up discussion but doesn't have to block this PR.

@ilblackdragon ilblackdragon left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verdict: Approve with nits

  • routine_engine.rs:1759 strip_internal_tool_call_text — only filters lines where the marker is the entire trimmed line. Inline markers (Result: [Tool http returned: …]) and multi-line JSON args still leak. Recommend a regex-based scrub spanning the whole string.
  • The ends_with(']') check is fragile: an args payload like {"a":"]"} ends with ] and gets dropped; a trailing space or ]. passes through. Tighten or anchor to a known marker grammar.
  • The fallback string is hardcoded English; fine, but worth a constant for reuse/i18n.
  • Tests cover three cases but miss [Tool … returned: end-to-end, multi-line JSON args, and multiple markers in one response.
  • All .expect() calls are inside #[cfg(test)]. Style matches surrounding code; uses fold instead of join to avoid intermediate Vec allocation — reasonable.

Ship-able as a mitigation; recommend a follow-up to harden the sanitizer against inline/multiline marker variants.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: agent Agent core (agent loop, router, scheduler) size: M 50-199 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Routine sends raw tool-call output to Telegram instead of human-readable message

2 participants