Skip to content

refactor: unify three agentic loops into single AgenticLoop engine - #879

Closed
zmanian wants to merge 9 commits into
stagingfrom
fix/pr-800-rebase
Closed

zmanian wants to merge 9 commits into
stagingfrom
fix/pr-800-rebase

Conversation

@zmanian

@zmanian zmanian commented Mar 10, 2026

Copy link
Copy Markdown
Collaborator

Summary

Rebased version of #800 with merge conflicts resolved against current staging.

  • Resolves conflicts in src/agent/dispatcher.rs (delegate-based loop already includes cost guardrails and nudge from staging)
  • Accepts deletion of src/worker/runtime.rs (worker runtime moved into unified loop)

Original PR: #800 by @qbit-glitch
Supersedes: #800

Test plan

  • All existing tests pass (verified locally in previous review)
  • CI should run full suite on this branch

Generated with Claude Code

qbit-glitch and others added 9 commits March 10, 2026 09:54
)

Replace three independent copy-pasted agentic loops (dispatcher, worker,
container runtime) with a single shared engine in `agentic_loop.rs` that
all consumers customize via the `LoopDelegate` trait.

Phase 1 — Shared engine (`src/agent/agentic_loop.rs`, 205 lines):
  - `run_agentic_loop()` owns the core LLM → tool exec → repeat cycle
  - `LoopDelegate` trait (Send + Sync, &dyn dispatch) with 6 hook points
  - Tool intent nudge logic consolidated (was duplicated in 3 files)
  - Iteration limit + force-text behavior preserved

Phase 2 — Three delegate implementations:
  - `ChatDelegate` (dispatcher.rs): 3-phase approval flow, hooks, cost
    guard, context compaction, skill attenuation, interruption
  - `JobDelegate` (worker/job.rs): planning pre-loop phase, parallel
    JoinSet exec, mark_completed/stuck/failed, SSE streaming, self-repair
  - `ContainerDelegate` (worker/container.rs): sequential tool exec,
    HTTP-proxied LLM, container-safe tools, credential injection

Phase 3 — File moves and cleanup:
  - Delete `src/agent/worker.rs` — job logic moved to `src/worker/job.rs`
  - Rename `src/worker/runtime.rs` → `src/worker/container.rs`
  - Re-export `Worker`/`WorkerDeps` from `crate::worker` in `agent/mod.rs`
  - Update `scheduler.rs` imports to new worker location

Shared helpers (`src/tools/execute.rs`):
  - `execute_tool_with_safety()` replaces 4 copies of validate → timeout
    → execute → serialize
  - `process_tool_result()` replaces 3 copies of sanitize → wrap →
    ChatMessage (also used by thread_ops.rs approval resume paths)

Net result: -2,408 lines, zero duplicated loop logic, single code path
for tool intent nudge and completion detection.

Closes #654

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
1. scheduler.rs: Replace `unwrap_or` fallback with proper error
   propagation when parsing tool output JSON — surfaces bugs instead
   of silently changing the output type.

2. worker/job.rs: Drop MutexGuard before the cancellation `.await` in
   `check_signals()` to avoid holding a lock across an async I/O call
   (prevents `await_holding_lock` lint).

3. worker/job.rs: Restore consecutive rate-limit counter
   (MAX_CONSECUTIVE_RATE_LIMITS = 10) so sustained rate limiting marks
   the job stuck with "Persistent rate limiting" instead of silently
   burning through max_iterations.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Merge staging's changes into the refactored JobDelegate:
- Add token budget tracking in call_llm (update_context/add_tokens)
- mark_stuck → mark_failed for iteration cap and rate-limit exhaustion
  (aligns with staging's #788 fix)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Address all 6 review points from zmanian on PR #800:

1. Replace LoopOutcome::Custom(Box<dyn Any>) with typed
   LoopOutcome::NeedApproval(Box<PendingApproval>) — eliminates
   type erasure and downcast, resolves clippy large_enum_variant.

2. Remove dead max_tool_iterations field from ChatDelegate struct.

3. Add on_tool_intent_nudge() hook to LoopDelegate trait with
   implementations in Job and Container delegates for observability.

4. Fix SSE events in job worker to emit raw sanitized content
   instead of XML-wrapped <tool_output> tags.

5. Remove 4 duplicate completion tests from job.rs that were
   already covered by the shared util module.

6. Avoid logging full tool results — use result_size_bytes in
   debug logs (execute.rs, job.rs).

Also updates path references in CLAUDE.md, COVERAGE_PLAN.md,
and add-sse-event.md command.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add 16 tests covering the two new critical shared modules:

agentic_loop.rs (10 tests):
- Text response exits loop immediately
- Tool call → text response continuation
- LoopSignal::Stop exits before LLM call
- LoopSignal::InjectMessage adds user message to context
- Max iterations terminates with LoopOutcome::MaxIterations
- Tool intent nudge fires twice then caps
- before_llm_call early exit bypasses LLM
- truncate_for_preview: short string, long string, multibyte safety

execute.rs (6 tests):
- execute_tool_with_safety success path
- Missing tool returns ToolError::NotFound
- Tool execution failure propagates
- Per-tool timeout enforcement (50ms)
- process_tool_result XML wrapping on success
- process_tool_result error formatting

All 2,777 unit tests pass, 0 clippy warnings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…ntic-loops

# Conflicts:
#	src/main.rs
#	src/worker/mod.rs
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…container

CRITICAL fixes:
- Rate-limit exhaustion now returns Err(LlmError::RateLimited) instead of
  Ok(Text("")), stopping the loop immediately with no ghost iteration.
  Below-threshold retries still use Text("") with an explicit empty-string
  guard in handle_text_response to skip injection.
- check_signals drains the entire message channel before returning,
  prioritizing Stop over UserMessage. Previously returned early on first
  UserMessage, silently dropping any queued Stop or additional messages.
- check_signals now detects all non-progressing job states (Cancelled,
  Failed, Stuck, Completed, Submitted, Accepted) instead of only
  Cancelled and Failed.

HIGH fixes:
- Error path in process_tool_result_job applies truncate_for_preview to
  bound error strings in SSE/DB events (was unbounded).
- Document Send+Sync lifetime constraint on LoopDelegate trait.
- Test mock before_llm_call refactored from double-lock to single lock
  acquisition, eliminating deadlock risk on refactor.

MEDIUM fixes:
- CompletionReport includes actual iteration count via shared
  Arc<Mutex<u32>> tracker (was hardcoded 0).
- process_tool_result_job return type changed from Result<bool> to
  Result<()> — the bool was always false (dead API).
- Deduplicate truncate in container.rs; now uses truncate_for_preview
  from agentic_loop.

Verified: 0 clippy warnings, 2781 tests pass, cargo fmt clean.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Resolve conflicts:
- dispatcher.rs: take PR version (delegate-based loop already includes cost guardrails and nudge)
- runtime.rs: accept deletion (worker runtime moved into unified loop)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@github-actions github-actions Bot added scope: agent Agent core (agent loop, router, scheduler) scope: tool Tool infrastructure scope: worker Container worker scope: docs Documentation size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Mar 10, 2026
@zmanian

zmanian commented Mar 10, 2026

Copy link
Copy Markdown
Collaborator Author

Superseded by combined PR

@zmanian zmanian closed this Mar 10, 2026
@zmanian
zmanian deleted the fix/pr-800-rebase branch March 10, 2026 18:13
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly refactors the core agentic reasoning and execution logic by unifying three previously distinct loops (for chat, background jobs, and container workers) into a single, reusable AgenticLoop engine. This change introduces a LoopDelegate trait, allowing each context to customize its specific behaviors while benefiting from a shared, robust foundation for LLM interaction and tool execution. The refactoring improves code consistency, reduces duplication, and enhances maintainability across the system's various agentic components.

Highlights

  • Unified Agentic Loop Engine: Introduced a new AgenticLoop engine (src/agent/agentic_loop.rs) that consolidates the core LLM call, tool execution, and context management logic previously duplicated across chat, job, and container worker loops. This engine uses a LoopDelegate trait to allow different contexts to customize behavior.
  • Refactored Chat Dispatcher: The src/agent/dispatcher.rs module, responsible for conversational turns, was refactored to utilize the new AgenticLoop via a ChatDelegate implementation, streamlining its internal logic.
  • Refactored Job Worker: The background job worker logic, previously in src/agent/worker.rs, was moved and refactored into src/worker/job.rs. It now implements a JobDelegate to leverage the shared AgenticLoop engine, including handling rate-limiting and signal processing.
  • New Container Worker Runtime: A new src/worker/container.rs module was added, implementing a ContainerDelegate to run the AgenticLoop within Docker containers, providing real-time event streaming and completion detection.
  • Shared Tool Execution Pipeline: A new module src/tools/execute.rs was created to centralize the tool execution pipeline, including validation, timeout handling, and result serialization, ensuring consistent and safe tool invocation across all agentic loops.
  • Documentation Updates: Key documentation files (CLAUDE.md, src/agent/CLAUDE.md, COVERAGE_PLAN.md, .claude/commands/add-sse-event.md) were updated to reflect the new modular structure and the unified agentic loop architecture.
Changelog
  • .claude/commands/add-sse-event.md
    • Updated the recommended location for triggering events from src/agent/worker.rs to src/worker/job.rs.
  • CLAUDE.md
    • Updated the directory structure description for the worker/ module, replacing runtime.rs with container.rs and job.rs.
    • Added descriptions for container.rs and job.rs in the worker/ module.
  • COVERAGE_PLAN.md
    • Updated file paths in coverage tables, replacing src/agent/worker.rs with src/worker/job.rs and src/worker/runtime.rs with src/worker/container.rs.
  • src/agent/CLAUDE.md
    • Updated the description for worker.rs to indicate its relocation to src/worker/job.rs and its new role as JobDelegate.
    • Added a new entry for agentic_loop.rs, describing it as the shared engine for all three execution paths.
  • src/agent/agentic_loop.rs
    • Added a new module containing the unified run_agentic_loop function and the LoopDelegate trait, along with associated enums and structs for loop control and outcomes.
    • Implemented comprehensive unit tests for the AgenticLoop functionality, covering text responses, tool calls, signals, iteration limits, and tool intent nudges.
  • src/agent/dispatcher.rs
    • Removed direct implementation of the agentic loop, replacing it with a ChatDelegate that uses the new run_agentic_loop.
    • Updated imports to include AgenticLoopConfig, LoopDelegate, LoopOutcome, LoopSignal, and TextAction.
    • Refactored tool execution logic to delegate to the shared execute_tool_with_safety function in src/tools/execute.rs.
  • src/agent/mod.rs
    • Added pub mod agentic_loop; to expose the new module.
    • Updated scheduler module visibility to pub(crate).
    • Removed pub use worker::{Worker, WorkerDeps}; as these are now in src/worker/job.rs.
  • src/agent/scheduler.rs
    • Updated imports to use crate::worker::job::{Worker, WorkerDeps} instead of crate::agent::worker::{Worker, WorkerDeps}.
    • Refactored execute_tool_task to use the shared crate::tools::execute::execute_tool_with_safety function, removing duplicated tool validation and execution logic.
  • src/agent/thread_ops.rs
    • Updated tool result processing in resume_tool_approval and execute_deferred_tool_calls to use the shared crate::tools::execute::process_tool_result function.
  • src/tools/execute.rs
    • Added a new module providing execute_tool_with_safety for robust tool execution (lookup, validation, timeout, serialization) and process_tool_result for consistent result handling and sanitization.
    • Included unit tests for tool execution success, failure, missing tools, and timeouts.
  • src/tools/mod.rs
    • Added pub mod execute; to expose the new shared tool execution module.
  • src/util.rs
    • Updated a comment in llm_signals_completion to reflect the new file paths for worker modules.
  • src/worker/container.rs
    • Added a new module implementing WorkerRuntime and ContainerDelegate for running the AgenticLoop within Docker containers.
    • Implemented logic for reporting status, polling for prompts, and streaming events to the orchestrator.
  • src/worker/job.rs
    • Renamed from src/agent/worker.rs to src/worker/job.rs.
    • Refactored the execution_loop to use the new run_agentic_loop via a JobDelegate implementation.
    • Removed redundant tool execution and result processing logic, delegating to shared functions in src/tools/execute.rs.
    • Updated signal handling to integrate with LoopDelegate::check_signals and manage rate-limiting within the delegate.
  • src/worker/mod.rs
    • Added pub mod container; and pub mod job; to include the new worker modules.
    • Updated pub use statements to export WorkerRuntime from container and Worker, WorkerDeps from job.
    • Modified run_worker to use container::WorkerConfig.
  • src/worker/runtime.rs
    • Removed the file, as its functionality has been replaced by src/worker/container.rs and the shared AgenticLoop.
Activity
  • The pull request is a rebased version of a previous PR (refactor: unify three agentic loops into single AgenticLoop engine #800), indicating prior development and review cycles.
  • Merge conflicts in src/agent/dispatcher.rs were resolved, specifically noting that the delegate-based loop already includes cost guardrails and nudges from staging.
  • The deletion of src/worker/runtime.rs was accepted, as its functionality was moved into the unified loop.
  • The author verified locally that all existing tests pass, and expects the CI to run the full test suite.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a significant and valuable refactoring by unifying the three agentic loops (chat, job, container) into a single engine driven by a LoopDelegate trait, greatly reducing code duplication and improving overall architecture. The introduction of a shared execute_tool_with_safety function is a good step towards consolidation. My review focuses on ensuring this new abstraction is applied consistently and correctly, particularly in areas concerning duplicated tool execution logic and precise job state management, aligning with established repository rules for code quality and resource handling.

Note: Security Review did not run due to the size of the PR.

Comment thread src/worker/job.rs
Comment on lines +601 to 609
let result_size = serde_json::to_string(&output.result)
.map(|s| s.len())
.unwrap_or(0);
tracing::debug!(
tool = %tool_name,
elapsed_ms = elapsed.as_millis() as u64,
result = %result_str,
result_size_bytes = result_size,
"Tool call succeeded"
);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This function (execute_tool_inner) duplicates a significant amount of logic from the new execute_tool_with_safety function in src/tools/execute.rs. To align with the goal of this PR to unify logic, this function should be refactored to use the shared tool execution logic after performing its job-specific pre-flight checks (approval, rate limiting, hooks). This might require adjusting execute_tool_with_safety to return more information (like the ToolOutput and execution duration) to allow for action recording.

References
  1. When an issue is found in duplicated code, prefer refactoring into a shared function over applying localized fixes.

Comment thread src/worker/job.rs
Comment on lines +1125 to +1126
| JobState::Submitted
| JobState::Accepted

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Including JobState::Submitted and JobState::Accepted in this check seems incorrect. A worker should only be active for a job in an InProgress state. Stopping the loop for jobs in these initial states might mask a potential logic error in the scheduler. This check should likely only include terminal or stuck states (Cancelled, Failed, Stuck, Completed). This aligns with the principle of distinguishing between truly active states and terminal or intermediate states to ensure proper resource management and prevent orphaned resources.

References
  1. When reaping resources based on job state, distinguish between truly active states (e.g., InProgress, Stuck) and terminal or intermediate states (e.g., Completed, Submitted). Do not skip cleanup for jobs in terminal or intermediate states, as their resources may be orphaned.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: agent Agent core (agent loop, router, scheduler) scope: docs Documentation scope: tool Tool infrastructure scope: worker Container worker size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants