feat(llm): thread per-tool reasoning through provider/tool-call path - #456
panosAthDBX wants to merge 3 commits into
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request integrates per-tool reasoning into the core LLM and provider layers. By extending the Highlights
Changelog
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request successfully threads per-tool reasoning through the LLM provider and tool-calling pathways, a significant enhancement for tool-use transparency. The changes are well-implemented, with consistent use of a normalization function for reasoning strings and good test coverage for the new logic. A notable addition is the robust panic handling mechanism in the rig_adapter, which improves application stability by catching and managing panics from the underlying rig-core dependency. While this is a valuable improvement, one part of its implementation could be more resilient, as detailed in the specific comment.
| fn should_suppress_rig_openai_usage_panic_details( | ||
| message: Option<&str>, | ||
| location_file: Option<&str>, | ||
| ) -> bool { | ||
| let Some(message) = message else { | ||
| return false; | ||
| }; | ||
| if !message.contains("attempt to subtract with overflow") { | ||
| return false; | ||
| } | ||
|
|
||
| let Some(file) = location_file else { | ||
| return false; | ||
| }; | ||
| file.contains("/src/providers/openai/completion/mod.rs") && file.contains("/rig-core-") | ||
| } |
There was a problem hiding this comment.
The current implementation for detecting the panic location is brittle as it relies on a hardcoded Unix-style path separator (/). This check will fail on Windows environments where the path separator is \. To make this logic more robust and cross-platform, I suggest normalizing the path separators to a consistent format before performing the string containment check.
| fn should_suppress_rig_openai_usage_panic_details( | |
| message: Option<&str>, | |
| location_file: Option<&str>, | |
| ) -> bool { | |
| let Some(message) = message else { | |
| return false; | |
| }; | |
| if !message.contains("attempt to subtract with overflow") { | |
| return false; | |
| } | |
| let Some(file) = location_file else { | |
| return false; | |
| }; | |
| file.contains("/src/providers/openai/completion/mod.rs") && file.contains("/rig-core-") | |
| } | |
| fn should_suppress_rig_openai_usage_panic_details( | |
| message: Option<&str>, | |
| location_file: Option<&str>, | |
| ) -> bool { | |
| let Some(message) = message else { | |
| return false; | |
| }; | |
| if !message.contains("attempt to subtract with overflow") { | |
| return false; | |
| } | |
| let Some(file) = location_file else { | |
| return false; | |
| }; | |
| file.replace('\\', "/").contains("src/providers/openai/completion/mod.rs") && file.contains("rig-core-") | |
| } |
ilblackdragon
left a comment
There was a problem hiding this comment.
Code Review
Overview
This PR extends ToolCall with a reasoning: String field so each tool call carries its own rationale (instead of sharing the response-level content). It adds a normalize_tool_reasoning() helper that falls back to a default string when providers don't supply reasoning. Additionally, it introduces a panic-catching wrapper around rig-core completions to handle a known upstream panic in rig-core's OpenAI usage tracking.
Two distinct concerns are bundled here: (1) per-tool reasoning plumbing, and (2) rig-core panic suppression. These should ideally be separate PRs, but it's not a blocker.
Positives
- Clean, well-scoped type change to
ToolCallwith consistent propagation - Good test coverage for the new normalization logic and reasoning threading
- The
StaticToolCompletionProvidermock is well-designed for testingReasoning::select_tools - Proper use of
normalize_tool_reasoning()at everyToolCallconstruction site
Issues & Suggestions
1. Global panic hook is risky (High concern)
The install_rig_panic_suppression_hook_once() function replaces the global panic hook. This is a process-wide side effect with several problems:
- Race condition with other hooks: The warning comment acknowledges this but doesn't mitigate it. If any other component (tracing-subscriber panic layer, test harness, etc.) also sets a hook, behavior is unpredictable.
Ordering::Relaxedon the depth counter: Inshould_suppress_rig_openai_usage_panic, the depth check usesRelaxedordering. Since panic hooks can run on any thread, this should useOrdering::Acquire/Release(or at minimumSeqCst) to ensure the hook sees the updated depth from the thread that entered the guard.- Scope creep: This panic suppression is a workaround for a bug in
rig-core. Consider filing an upstream issue and pinning to a fixed version instead. If the workaround must stay, it deserves its own module/file and PR.
2. AssertUnwindSafe on async futures (Medium concern)
let response = AssertUnwindSafe(fut)
.catch_unwind()
.awaitThe // SAFETY comment claims no mutable aliases are retained, but AssertUnwindSafe on an arbitrary CompletionModel::completion future is a strong assertion. If the model implementation holds &mut self or interior mutability across .await points, a caught panic could leave the model in an inconsistent state. The RigAdapter would then be reused on the next call with potentially corrupted internal state.
Suggestion: At minimum, document that RigAdapter instances should be considered poisoned after a caught panic, or wrap the model in an Option that gets taken on panic.
3. Unnecessary allocation on every tool call (Low concern)
normalize_tool_reasoning("") allocates a new String from DEFAULT_TOOL_RATIONALE every time. Since most providers don't supply reasoning yet, this happens on virtually every tool call.
Suggestion: Consider using Cow<'static, str> as the return type, or store reasoning as Option<String> on ToolCall (where None means "use default"). This avoids cloning the static string repeatedly and is more idiomatic for "maybe has a value" semantics.
4. Behavioral change in select_tools (Medium concern)
Previously, ToolSelection::reasoning was set to the response-level content (the model's overall thinking). Now it's set to per-tool reasoning from tool_call.reasoning. But since all current providers populate reasoning as "", every selection will now get DEFAULT_TOOL_RATIONALE instead of the model's actual thinking text.
This is a regression in reasoning quality for the select_tools path until providers actually populate per-tool reasoning. The old behavior (using response.content) provided more useful context.
Suggestion: Fall back to response.content when per-tool reasoning is empty, rather than the generic default string:
reasoning: if tool_call.reasoning.trim().is_empty() {
reasoning_from_content.clone() // preserve old behavior
} else {
tool_call.reasoning
}5. Consider extracting panic suppression (Nit)
Once, AtomicUsize, and the suppression machinery add non-trivial complexity to what was described as a "reasoning plumbing" PR. Consider extracting to a separate module like src/llm/rig_panic_guard.rs.
Summary
| Area | Assessment |
|---|---|
| Correctness | The reasoning field plumbing is correct. The panic suppression works but has ordering and safety concerns. |
| Conventions | Follows project patterns. Test coverage is good. |
| Performance | Minor: unnecessary string allocations on every tool call. |
| Security | No concerns. |
| Risk | Medium — the global panic hook and AssertUnwindSafe on async futures are the main risks. The behavioral change in select_tools reasoning is a subtle regression. |
Recommendation: Address items #1 (memory ordering) and #4 (reasoning fallback regression) before merging. Consider splitting the panic suppression into a separate PR.
|
Implemented all ilblackdragon review recommendations in follow-up commit 2bab285 (ported from local split/worker-orchestrator-streaming-v2):
|
ilblackdragon
left a comment
There was a problem hiding this comment.
Code Review
Overview
This PR adds a reasoning: String field to ToolCall so each tool call carries its own rationale, with normalize_tool_reasoning() ensuring a non-empty fallback. The field is threaded through Reasoning::select_tools, RespondResult::ToolCalls, and all adapter/provider construction sites. A second, unrelated change catches panics from a known rig-core bug in the OpenAI usage tracker.
Key Issues
1. Behavioral regression in select_tools (High)
src/llm/reasoning.rs:372-379 — Previously ToolSelection::reasoning used response.content (the model's actual thinking). Now it uses normalize_tool_reasoning(&tool_call.reasoning), which resolves to "Tool selected to satisfy the current subtask." for every provider since none populate per-tool reasoning yet.
This is a downgrade — callers that consumed ToolSelection::reasoning now get a useless static string instead of the model's chain-of-thought. Suggestion: fall back to response.content when per-tool reasoning is empty:
let fallback_reasoning = response.content.unwrap_or_default();
// ...
reasoning: if tool_call.reasoning.trim().is_empty() {
fallback_reasoning.clone()
} else {
tool_call.reasoning.trim().to_string()
},2. Global panic hook is a footgun (High)
src/llm/rig_adapter.rs:399-462 — install_rig_panic_suppression_hook_once() replaces the global panic hook. Concerns:
Ordering::RelaxedonRIG_PANIC_SUPPRESSION_DEPTHis incorrect. The hook runs on whatever thread panicked;Relaxeddoesn't guarantee visibility of thefetch_addfrom the calling thread. UseAcquire/Releaseat minimum.- Hook ordering fragility — any subsequent
set_hookcall silently removes this suppression. - Scope: This is a workaround for a specific rig-core bug. Is there an upstream issue filed? If rig-core ≥ 0.31 fixes it, a version bump is cleaner. If the workaround is necessary, it should live in its own module (
src/llm/rig_panic_guard.rs), not inline in the adapter.
3. AssertUnwindSafe on async futures (Medium)
src/llm/rig_adapter.rs:489-499 — Wrapping an arbitrary CompletionModel future in AssertUnwindSafe is technically unsound if the model holds &mut state across .await points. After a caught panic, the RigAdapter is reused with potentially corrupted model state.
Consider marking the adapter as poisoned after a caught panic, or documenting this as a known risk scoped to specific model implementations.
4. Unnecessary allocations (Low)
normalize_tool_reasoning("") allocates a new String on every tool call, and currently all providers pass "". Consider Cow<'static, str> or Option<String> on ToolCall to avoid this.
Positives
- Clean, consistent propagation of the new field to all
ToolCallconstruction sites - Good test coverage:
StaticToolCompletionProvidermock, normalization unit tests, threading integration tests - Thorough test updates — no missed construction sites
Nits
tests/openai_compat_integration.rs:95uses a raw string instead ofnormalize_tool_reasoning(...), inconsistent with the rest of the PR- The panic suppression machinery (
futures::FutureExt,Once,AtomicUsize, guard struct) adds significant complexity to a "reasoning plumbing" PR — consider splitting into a separate PR
Recommendation
Fix the select_tools reasoning regression (#1) and memory ordering (#2) before merge. Consider splitting panic suppression into a separate PR.
zmanian
left a comment
There was a problem hiding this comment.
Review: feat(llm): thread per-tool reasoning through provider/tool-call path
The type-level change and normalization approach are sound. normalize_tool_reasoning with DEFAULT_TOOL_RATIONALE fallback is clean, and the test coverage for the new helpers and the select_tools / respond threading is solid.
However, this PR will not compile as-is. Adding reasoning: String to ToolCall is a breaking struct change, and several production call sites that construct ToolCall were not updated:
Missing reasoning field (compilation errors)
Production providers (same layer as the files in this PR):
src/llm/bedrock.rs~line 518 --ToolCallinextract_responseequivalentsrc/llm/anthropic_oauth.rs~line 550 --ToolCallconstruction in response parsing
Agent/worker layer (production code):
src/agent/session.rs~line 349 --rebuild_messages()constructsToolCallfrom turn historysrc/agent/thread_ops.rs~line 1660 -- tool call reconstruction from stored JSONsrc/worker/job.rs~line 871 and ~line 1325 (selections_to_tool_calls) -- job worker tool dispatch
Test code (also won't compile):
src/agent/agentic_loop.rs~line 388src/agent/session.rstest blocks (~lines 1191, 1221, 1286)src/worker/job.rs~line 871 testsrc/llm/bedrock.rstest blocks (~lines 755, 760, 798, 821, 985, 990)src/llm/nearai_chat.rshas some in-PR fixes but check ~lines 1456, 1505, 2125 as well
Rig panic suppression: out of scope
The run_completion_guarded / install_rig_panic_suppression_hook_once addition (~120 lines) is unrelated to the per-tool reasoning feature. It catches a known rig-core overflow panic and converts it to LlmError. While potentially useful, it:
- Replaces the global panic hook, which affects the entire process -- this deserves its own PR with dedicated review and testing.
- Uses
AssertUnwindSafeacross an async boundary, which needs careful scrutiny beyond a "reasoning threading" PR. - The
// SAFETYcomment is appreciated butAssertUnwindSafeon a future is not a soundness concern -- it's a logical correctness concern (are invariants maintained after partial execution?). The comment should address that.
Recommend splitting this into its own PR.
Minor
- In
tests/openai_compat_integration.rs, the mock uses a raw string"tool selected for integration test coverage"instead ofnormalize_tool_reasoning(...). This is fine functionally but inconsistent with the pattern established elsewhere in the PR. Consider using the normalizer for consistency, or at least noting that non-empty strings pass through unchanged.
Verdict
The core design (field on ToolCall, normalization helper, threading through reasoning/provider) is good. But the PR needs to update ALL ToolCall construction sites to compile. Run cargo check across the full workspace to confirm. The panic suppression should be split out.
…, SSE, and DB Add end-to-end agent reasoning summaries so users can see *why* the agent chose specific tools, not just what it did. - Add `reasoning: Option<String>` to `ToolCall` (all providers) - Populate from LLM response content in `Reasoning::respond_with_tools` and `select_tools`, with per-tool override when providers supply it - Extend `Turn` with `narrative` and `TurnToolCall` with `rationale` + `tool_call_id` for identity-based result matching - Persist reasoning in DB via existing tool_calls JSON (no migration) - Add `StatusUpdate::ReasoningUpdate` and `SseEvent::ReasoningUpdate` + `SseEvent::JobReasoning` for real-time streaming - Emit reasoning events in both chat dispatcher and worker job path - Add `/reasoning [N|all]` command for inspecting turn reasoning - Surface `narrative` and `rationale` in HTTP `/api/chat/history` Based on the design from #361 and #456, reconstructed cleanly with Option<String> to minimize blast radius (vs mandatory String that broke compilation in #456). Closes #456 Co-Authored-By: panosAthDBX <47406510+panosAthDBX@users.noreply.github.com> Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
This has been implemented in #1513 which covers the full feature end-to-end:
All 3500+ tests pass with zero clippy warnings. Thanks for starting this work, @panosAthDBX! |
… all surfaces (#1513) * feat(agent): thread per-tool reasoning from LLM through to REPL, HTTP, SSE, and DB Add end-to-end agent reasoning summaries so users can see *why* the agent chose specific tools, not just what it did. - Add `reasoning: Option<String>` to `ToolCall` (all providers) - Populate from LLM response content in `Reasoning::respond_with_tools` and `select_tools`, with per-tool override when providers supply it - Extend `Turn` with `narrative` and `TurnToolCall` with `rationale` + `tool_call_id` for identity-based result matching - Persist reasoning in DB via existing tool_calls JSON (no migration) - Add `StatusUpdate::ReasoningUpdate` and `SseEvent::ReasoningUpdate` + `SseEvent::JobReasoning` for real-time streaming - Emit reasoning events in both chat dispatcher and worker job path - Add `/reasoning [N|all]` command for inspecting turn reasoning - Surface `narrative` and `rationale` in HTTP `/api/chat/history` Based on the design from #361 and #456, reconstructed cleanly with Option<String> to minimize blast radius (vs mandatory String that broke compilation in #456). Closes #456 Co-Authored-By: panosAthDBX <47406510+panosAthDBX@users.noreply.github.com> Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review feedback from Gemini and Copilot - Fix `_ => Ok(None)` in agent_loop.rs to avoid accidental shutdown - Fix fallback in record_tool_result_for/record_tool_error_for to use first pending call instead of last_mut (parallel execution safety) - Include per-tool decisions in WASM channel reasoning messages - Apply truncate_at_tool_tags + clean_response to shared_reasoning in select_tools (parity with respond_with_tools) - Persist turn-level narrative to DB in tool_calls JSON wrapper - Parse both old (array) and new (object) tool_calls formats in build_turns_from_db_messages for backward compatibility - Populate reasoning from action.reasoning in execute_plan ToolCalls [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address second round of review comments + merge fixes - Add reasoning: None to new github_copilot.rs ToolCall sites (from staging merge) - Run cargo fmt on 4 files with formatting diffs - Truncate narrative to 1000 chars before DB persistence - Clone turn data and drop session lock in /reasoning command - Extract ToolDecisionDto::from_json_array shared helper (deduplicate worker/job.rs and orchestrator/api.rs) - Add unit tests for wrapped tool_calls JSON format with narrative [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address third round of review comments (Copilot + serrrfirat) - Reword ToolCall.reasoning docstring to reflect provider-supplied or fallback contract - Sanitize narrative through SafetyLayer before storage/emission - Clean per-tool reasoning via truncate_at_tool_tags + clean_response in select_tools (parity with shared reasoning) - Convert 4 approval-path recording sites in thread_ops.rs to identity-based record_tool_result_for/record_tool_error_for - Preserve tool_call_id and reasoning through restore_from_messages - Fix has_result/has_error to reject JSON null values - Truncate tool_call_id to 128 chars before DB persistence - Add 4 unit tests for record_tool_result_for/error_for edge cases Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address zmanian review — sanitize JobDelegate reasoning + warn on dropped results - Sanitize narrative and per-tool rationale through SafetyLayer in JobDelegate reasoning events (parity with ChatDelegate) - Add tracing::warn when record_tool_result_for/error_for drops a result because no matching or pending tool call exists - Add 3 unit tests for reasoning normalization (thinking tags, tool tags, empty-after-cleaning) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address 4 remaining unreplied review comments - Clean per-tool reasoning in respond_with_tools via truncate_at_tool_tags + clean_response (parity with select_tools) - Handle wrapped JSON format in rebuild_chat_messages_from_db so cold hydration works after persist_tool_calls format change - Update persist_tool_calls doc comment to describe new JSON shape - Sanitize per-tool rationale through SafetyLayer in ChatDelegate before emission and storage (parity with JobDelegate) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address zmanian review round 2 - Add tracing::debug on fallback-to-pending path in record_tool_result_for and record_tool_error_for (item 1) - Add comment explaining why /reasoning is special-cased in agent_loop.rs (item 4) - Items 2 (narrative persistence), 3 (rationale sanitization), and 5 (catch-all fix) were already addressed in prior commits Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: panosAthDBX <47406510+panosAthDBX@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
… all surfaces (nearai#1513) * feat(agent): thread per-tool reasoning from LLM through to REPL, HTTP, SSE, and DB Add end-to-end agent reasoning summaries so users can see *why* the agent chose specific tools, not just what it did. - Add `reasoning: Option<String>` to `ToolCall` (all providers) - Populate from LLM response content in `Reasoning::respond_with_tools` and `select_tools`, with per-tool override when providers supply it - Extend `Turn` with `narrative` and `TurnToolCall` with `rationale` + `tool_call_id` for identity-based result matching - Persist reasoning in DB via existing tool_calls JSON (no migration) - Add `StatusUpdate::ReasoningUpdate` and `SseEvent::ReasoningUpdate` + `SseEvent::JobReasoning` for real-time streaming - Emit reasoning events in both chat dispatcher and worker job path - Add `/reasoning [N|all]` command for inspecting turn reasoning - Surface `narrative` and `rationale` in HTTP `/api/chat/history` Based on the design from nearai#361 and nearai#456, reconstructed cleanly with Option<String> to minimize blast radius (vs mandatory String that broke compilation in nearai#456). Closes nearai#456 Co-Authored-By: panosAthDBX <47406510+panosAthDBX@users.noreply.github.com> Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review feedback from Gemini and Copilot - Fix `_ => Ok(None)` in agent_loop.rs to avoid accidental shutdown - Fix fallback in record_tool_result_for/record_tool_error_for to use first pending call instead of last_mut (parallel execution safety) - Include per-tool decisions in WASM channel reasoning messages - Apply truncate_at_tool_tags + clean_response to shared_reasoning in select_tools (parity with respond_with_tools) - Persist turn-level narrative to DB in tool_calls JSON wrapper - Parse both old (array) and new (object) tool_calls formats in build_turns_from_db_messages for backward compatibility - Populate reasoning from action.reasoning in execute_plan ToolCalls [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address second round of review comments + merge fixes - Add reasoning: None to new github_copilot.rs ToolCall sites (from staging merge) - Run cargo fmt on 4 files with formatting diffs - Truncate narrative to 1000 chars before DB persistence - Clone turn data and drop session lock in /reasoning command - Extract ToolDecisionDto::from_json_array shared helper (deduplicate worker/job.rs and orchestrator/api.rs) - Add unit tests for wrapped tool_calls JSON format with narrative [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address third round of review comments (Copilot + serrrfirat) - Reword ToolCall.reasoning docstring to reflect provider-supplied or fallback contract - Sanitize narrative through SafetyLayer before storage/emission - Clean per-tool reasoning via truncate_at_tool_tags + clean_response in select_tools (parity with shared reasoning) - Convert 4 approval-path recording sites in thread_ops.rs to identity-based record_tool_result_for/record_tool_error_for - Preserve tool_call_id and reasoning through restore_from_messages - Fix has_result/has_error to reject JSON null values - Truncate tool_call_id to 128 chars before DB persistence - Add 4 unit tests for record_tool_result_for/error_for edge cases Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address zmanian review — sanitize JobDelegate reasoning + warn on dropped results - Sanitize narrative and per-tool rationale through SafetyLayer in JobDelegate reasoning events (parity with ChatDelegate) - Add tracing::warn when record_tool_result_for/error_for drops a result because no matching or pending tool call exists - Add 3 unit tests for reasoning normalization (thinking tags, tool tags, empty-after-cleaning) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address 4 remaining unreplied review comments - Clean per-tool reasoning in respond_with_tools via truncate_at_tool_tags + clean_response (parity with select_tools) - Handle wrapped JSON format in rebuild_chat_messages_from_db so cold hydration works after persist_tool_calls format change - Update persist_tool_calls doc comment to describe new JSON shape - Sanitize per-tool rationale through SafetyLayer in ChatDelegate before emission and storage (parity with JobDelegate) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address zmanian review round 2 - Add tracing::debug on fallback-to-pending path in record_tool_result_for and record_tool_error_for (item 1) - Add comment explaining why /reasoning is special-cased in agent_loop.rs (item 4) - Items 2 (narrative persistence), 3 (rationale sanitization), and 5 (catch-all fix) were already addressed in prior commits Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: panosAthDBX <47406510+panosAthDBX@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
… all surfaces (nearai#1513) * feat(agent): thread per-tool reasoning from LLM through to REPL, HTTP, SSE, and DB Add end-to-end agent reasoning summaries so users can see *why* the agent chose specific tools, not just what it did. - Add `reasoning: Option<String>` to `ToolCall` (all providers) - Populate from LLM response content in `Reasoning::respond_with_tools` and `select_tools`, with per-tool override when providers supply it - Extend `Turn` with `narrative` and `TurnToolCall` with `rationale` + `tool_call_id` for identity-based result matching - Persist reasoning in DB via existing tool_calls JSON (no migration) - Add `StatusUpdate::ReasoningUpdate` and `SseEvent::ReasoningUpdate` + `SseEvent::JobReasoning` for real-time streaming - Emit reasoning events in both chat dispatcher and worker job path - Add `/reasoning [N|all]` command for inspecting turn reasoning - Surface `narrative` and `rationale` in HTTP `/api/chat/history` Based on the design from nearai#361 and nearai#456, reconstructed cleanly with Option<String> to minimize blast radius (vs mandatory String that broke compilation in nearai#456). Closes nearai#456 Co-Authored-By: panosAthDBX <47406510+panosAthDBX@users.noreply.github.com> Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review feedback from Gemini and Copilot - Fix `_ => Ok(None)` in agent_loop.rs to avoid accidental shutdown - Fix fallback in record_tool_result_for/record_tool_error_for to use first pending call instead of last_mut (parallel execution safety) - Include per-tool decisions in WASM channel reasoning messages - Apply truncate_at_tool_tags + clean_response to shared_reasoning in select_tools (parity with respond_with_tools) - Persist turn-level narrative to DB in tool_calls JSON wrapper - Parse both old (array) and new (object) tool_calls formats in build_turns_from_db_messages for backward compatibility - Populate reasoning from action.reasoning in execute_plan ToolCalls [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address second round of review comments + merge fixes - Add reasoning: None to new github_copilot.rs ToolCall sites (from staging merge) - Run cargo fmt on 4 files with formatting diffs - Truncate narrative to 1000 chars before DB persistence - Clone turn data and drop session lock in /reasoning command - Extract ToolDecisionDto::from_json_array shared helper (deduplicate worker/job.rs and orchestrator/api.rs) - Add unit tests for wrapped tool_calls JSON format with narrative [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address third round of review comments (Copilot + serrrfirat) - Reword ToolCall.reasoning docstring to reflect provider-supplied or fallback contract - Sanitize narrative through SafetyLayer before storage/emission - Clean per-tool reasoning via truncate_at_tool_tags + clean_response in select_tools (parity with shared reasoning) - Convert 4 approval-path recording sites in thread_ops.rs to identity-based record_tool_result_for/record_tool_error_for - Preserve tool_call_id and reasoning through restore_from_messages - Fix has_result/has_error to reject JSON null values - Truncate tool_call_id to 128 chars before DB persistence - Add 4 unit tests for record_tool_result_for/error_for edge cases Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address zmanian review — sanitize JobDelegate reasoning + warn on dropped results - Sanitize narrative and per-tool rationale through SafetyLayer in JobDelegate reasoning events (parity with ChatDelegate) - Add tracing::warn when record_tool_result_for/error_for drops a result because no matching or pending tool call exists - Add 3 unit tests for reasoning normalization (thinking tags, tool tags, empty-after-cleaning) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address 4 remaining unreplied review comments - Clean per-tool reasoning in respond_with_tools via truncate_at_tool_tags + clean_response (parity with select_tools) - Handle wrapped JSON format in rebuild_chat_messages_from_db so cold hydration works after persist_tool_calls format change - Update persist_tool_calls doc comment to describe new JSON shape - Sanitize per-tool rationale through SafetyLayer in ChatDelegate before emission and storage (parity with JobDelegate) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address zmanian review round 2 - Add tracing::debug on fallback-to-pending path in record_tool_result_for and record_tool_error_for (item 1) - Add comment explaining why /reasoning is special-cased in agent_loop.rs (item 4) - Items 2 (narrative persistence), 3 (rationale sanitization), and 5 (catch-all fix) were already addressed in prior commits Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: panosAthDBX <47406510+panosAthDBX@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Summary
First split PR from #361: core reasoning-summary plumbing in the LLM/provider layer.
This PR is intentionally scoped to provider/types + LLM reasoning integration, matching requested split direction:
src/llm/provider.rssrc/llm/reasoning.rssrc/llm/rig_adapter.rsToolCallChanges
ToolCallwithreasoning: String.DEFAULT_TOOL_RATIONALEnormalize_tool_reasoning(...)Reasoning::select_toolsandRespondResult::ToolCallsconstruction.ToolCallis constructed directly.Why this split
Everything downstream (dispatcher/session, worker streaming, web UI reasoning surfaces) depends on this type-level and provider-level contract. Landing this first reduces risk and makes follow-up PRs mechanical and reviewable.
Testing
Ran locally before opening:
cargo clippy --all --all-featurescargo test