Skip to content

feat(agent-loop): output-aware no-progress detection (PR3) - #5022

Merged
serrrfirat merged 7 commits into
mainfrom
firat/no-progress-pr3-output-aware
Jun 17, 2026
Merged

serrrfirat merged 7 commits into
mainfrom
firat/no-progress-pr3-output-aware

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Stacked on #5000 (PR2), which is stacked on #4993 (PR1). Review against base firat/no-progress-pr2-content-digest so the diff shows only this change; retarget down the stack as the lower PRs merge.

What

The keystone of the no-progress redesign. PR2 added an inert ContentDigest of each completed capability output; this PR consumes it so "no progress" finally means the output actually repeated, and it takes failures off the no-progress axis.

  • Derive real progress in append_completed_capability_result: if the same call (signature) produced an output we've already seen this run → NoChange; a first-seen output → MadeProgress; no digest (synthetic/older host) → the host-reported progress. The membership check runs before recording the observation (or a first occurrence would look "seen"). This flips PR2 from inert population to active output-aware detection.
  • Failures off the axis in stop.rs: only NoChange counts toward the no-progress escape (Blocked dropped), and the repeated-call warning terminalizes only on a genuine NoChange repeat (new no_change_signatures track). Failing/blocked tools route through recovery and the budget/iteration limit — never the no-progress give-up.

Why this is the keystone

Forensics on the PinchBench taxonomy run showed the no-progress give-up killed tasks that were making real progress (e.g. task_selector_fix: repeated apply_patch with changing output, 10/19 rubric checks already green) and tasks where tools were failing (apply_patch/http errors mislabeled as "looping"). After PR3:

  • same call + changing output (polling/pagination) → progress, never trips
  • same call + identical output, repeated → genuine stagnation → honest typed failure
  • failing/blocked tools → recovery + budget, not a fake "I'm repeating myself"

Net effect (PR1 + PR2 + PR3)

No-progress now fires only on genuine output-stagnation, as an honest Failed(NoProgressDetected) (or a real nudged answer where #4837 is enabled) — never a faked completion, never on a progressing model, never on failing tools.

Tests

  • Flipped PR2's inertness test → repeated_identical_output_digest_trips_no_progress (identical output now trips).
  • New changing_output_digests_do_not_trip_no_progress — the load-bearing polling counterpart (same call, advancing output, completes normally).
  • Blocked-failure executor/integration tests now assert the run recovers and completes instead of a no-progress escape; trailing_no_progress_results == 0.
  • Unknown progress no longer terminalizes the repeated-call warning.

Validation

fmt clean · cargo clippy -p ironclaw_agent_loop --all-targets --all-features 0 warnings · cargo test -p ironclaw_agent_loop all green (318 lib + integration binaries). Changes are confined to ironclaw_agent_loop internals (no public API/struct change), so the full-workspace gate + the PinchBench benchmark on the PR1+PR2+PR3 stack run in CI.

Implementation note: the Codex handoff hung repeatedly in a disk-exhausted local environment; this was implemented directly against the agreed spec once the disk was reclaimed.

🤖 Generated with Claude Code

serrrfirat and others added 3 commits June 16, 2026 22:57
…mpletion

When the runaway-loop safety guard fires (StopKind::NoProgressDetected), the
executor finalized a canned "I stopped because I was repeating the same step"
assistant reply and returned the run as Completed — a runtime control decision
leaking into the conversation as a fake successful turn. This hid real blockers
(auth, broken connector, empty search) and marked incomplete tasks as done.

The ExitStage NoProgressDetected arm now:
- keeps the #4837 final-answer-nudge path bit-for-bit: when the gate is enabled
  and the model synthesizes a real closing answer, complete with that answer
  (PinchBench path unchanged);
- otherwise writes the Final checkpoint and returns a typed
  failed_exit(LoopFailureKind::NoProgressDetected) instead of the canned reply.
  The product layer already maps "no_progress_detected" to deterministic copy,
  so the user sees an honest failure on every channel.

Deletes NO_PROGRESS_FALLBACK_REPLY and finalize_no_progress_fallback.

PR1 of the output-aware no-progress redesign (exit honesty). PR2 (content-digest
progress signal) and PR3 (failures off the no-progress axis) follow.

Wire behavior: gate-off no-progress runs now arrive as Failed{no_progress_detected}
instead of Completed with a canned reply.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…PR2)

Inert plumbing for output-aware no-progress detection (PR3 will consume it).
Adds a ContentDigest of each completed capability output so a later change can
treat "the output moved" as real progress. This PR does NOT change any
detection/stop behavior — it is observably inert.

- ContentDigest newtype (keyed-blake3, mirrors ArgsHash) in ironclaw_turns;
  normalize_for_hash moved there and re-exported from agent_loop::strategies so
  CapabilityCallSignature hashing is unchanged (no checkpoint-signature break).
- output_digest: Option<ContentDigest> on CapabilityResultMessage (serde default
  None = back-compat); computed host-side in the write_capability_result
  chokepoint alongside byte_len. Best-effort: a digest-compute failure degrades
  to None and never fails an otherwise-successful capability write.
- seen_capability_output_digests: BoundedRing<_, 64> on LoopExecutionState,
  populated at append_completed_capability_result but NOT consumed. Additive
  #[serde(default)] checkpoint field (legacy checkpoints decode to an empty
  ring; round-trip + legacy-decode tests added).
- record_result still receives the host's progress unchanged — detection is
  byte-identical (inertness test drives the executor and asserts no behavior
  change for repeated identical-output completed calls).

Review fixes folded in (multi-agent review): from_output is fail-open; a
caller-level test asserts the digest is recorded into the ring through a real
run; boundary note documents why the digest impl lives in ironclaw_turns.

PR2 of the no-progress redesign. PR1 (#4993) = honest typed failure;
PR3 = consume the digest (NoChange when output repeats) + take failures off the
no-progress axis. Benchmark-gated at PR3.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Consumes the PR2 content-digest plumbing to make "no progress" mean the tool
OUTPUT actually repeated, and takes failures off the no-progress axis.

- append_completed_capability_result now DERIVES progress from the digest: if the
  same call (signature) produced an output already seen this run -> NoChange; a
  first-seen output -> MadeProgress; no digest -> host-reported progress. The
  membership check runs BEFORE recording the observation. This flips PR2 from
  inert population to active output-aware detection.
- stop.rs counts ONLY NoChange toward the no-progress escape (Blocked dropped),
  and the repeated-call warning terminalizes only on a genuine NoChange repeat
  (new no_change_signatures track). Failing/blocked tools therefore route through
  recovery and the budget/iteration limit, never the no-progress give-up.

Net (PR1+PR2+PR3): no-progress fires on genuine output-stagnation (same call,
same output, repeated) as an honest typed failure; it no longer fires on a model
progressing with changing output (polling/pagination) or on failing tools, and
it never fakes a completion.

Tests: flipped PR2's inertness test (identical output now trips); added the
polling counterpart (changing output does not trip); blocked-failure tests now
assert the run recovers instead of a no-progress escape; Unknown progress no
longer terminalizes the repeated-call warning.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Jun 17, 2026 •

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

Pull request was closed or merged during review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 4ee05270-f8f4-4782-8b5d-0eb7d11c7516

📥 Commits

Reviewing files that changed from the base of the PR and between a6f7395 and 967641f.

📒 Files selected for processing (4)
  • crates/ironclaw_agent_loop/src/executor/capabilities.rs
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_agent_loop/src/strategies/stop.rs
  • crates/ironclaw_agent_loop/tests/safety_nets.rs

📝 Walkthrough

Summary by CodeRabbit

  • Bug Fixes

    • Enhanced no-progress detection to accurately identify repeated capability calls with identical outputs.
    • Blocked or failed capability outcomes no longer incorrectly trigger no-progress escape logic.
  • Refactor

    • Improved capability progress tracking to distinguish between unchanged outputs and blocked/failed states, providing more accurate agent loop termination behavior.

Walkthrough

Refines no-progress detection so only capability results with identical output digests (mapped to CapabilityProgress::NoChange) contribute to the repeated-call termination path. Adds no_change_signatures to CapabilityBatchTurnSummary, updates append_completed_capability_result to derive progress via (signature, output_digest) deduplication, gates warning terminalization on that new field, and updates unit and integration tests throughout.

Changes

No-progress detection via output-digest deduplication

Layer / File(s) Summary
CapabilityBatchTurnSummary field and recording
crates/ironclaw_agent_loop/src/strategies/stop.rs
Adds no_change_signatures: Vec<CapabilityCallSignature> field; initializes empty; record_result now increments no_progress_count and populates no_change_signatures only for NoChange — Blocked no longer counts.
Repeated-call warning termination logic
crates/ironclaw_agent_loop/src/strategies/stop.rs
Replaces signature_observed helper with signature_no_change; TerminalReady transition now requires the repeated signature to be present in no_change_signatures; Unknown progress keeps warning in Rendered and returns Continue.
Executor digest deduplication
crates/ironclaw_agent_loop/src/executor/capabilities.rs
append_completed_capability_result checks seen_capability_output_digests for (signature, output_digest); first-seen → MadeProgress + stored observation; repeat → NoChange; absent digest falls back to result.progress.
Executor unit tests
crates/ironclaw_agent_loop/src/executor/tests.rs
Adds no_change_signatures: Vec::new() to TurnSummary literals; replaces repeated-failure → NoProgressDetected test with two new tests asserting LoopExit::Completed and trailing_no_progress_results == 0 for blocked/failed batches.
Safety-net integration tests
crates/ironclaw_agent_loop/tests/safety_nets.rs
Three test rewrites: identical digests still trip LoopFailureKind::NoProgressDetected; changing digests complete normally; blocked results complete normally with recovered reply.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • nearai/ironclaw#4993: Changes the executor's NoProgressDetected exit to a typed Failed path — the no-progress termination path this PR refines at the stop-strategy level.
  • nearai/ironclaw#5000: Adds the output_digest / seen_capability_output_digests plumbing that append_completed_capability_result now reads to derive NoChange vs MadeProgress.

Poem

🔁 Repeated calls with the same old hash,
Get stamped NoChange — no progress, no dash.
But Blocked tools get a gentler pass,
The loop recovers, avoids the impasse.
Digests now rule the no-progress gate. 🦀

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed Title follows Conventional Commits style (feat(scope): summary) and clearly describes the PR's core change: output-aware progress detection.
Description check ✅ Passed Description includes What/Why summary, validates all cargo checks and test passes, documents changes across three stacked commits, and declares confined scope with zero public API changes.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions github-actions Bot added the size: M 50-199 changed lines label Jun 17, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5022 June 17, 2026 07:33 Destroyed
@github-actions github-actions Bot added the risk: low Changes to docs, tests, or low-risk modules label Jun 17, 2026
@railway-app

railway-app Bot commented Jun 17, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5022 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw 🕒 Building (View Logs) Web Jun 17, 2026 at 10:47 pm

@github-actions github-actions Bot added the contributor: core 20+ merged PRs label Jun 17, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the agent loop's progress tracking to be output-aware, ensuring that only genuine output stagnation (repeated identical output digests) triggers the no-progress guard, while tool failures and advancing outputs are allowed to continue or recover. The feedback suggests utilizing exhaustive enum matching instead of conditional checks when handling the capability progress in stop.rs to leverage compiler-enforced safety and avoid redundant cloning.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +138 to 143
if progress == CapabilityProgress::NoChange {
self.no_progress_count = self.no_progress_count.saturating_add(1);
push_unique_signature(&mut self.no_change_signatures, signature.clone());
}
if progress == CapabilityProgress::MadeProgress {
push_unique_signature(&mut self.made_progress_signatures, signature);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Since CapabilityProgress is an enum, we should prefer exhaustive enum matching over runtime property-based checks (like if / else if chains) to leverage compiler-enforced safety. By using a match statement, we can safely move signature in both branches without needing to clone it, avoiding redundant allocations while keeping the code robust.

        match progress {
            CapabilityProgress::NoChange => {
                self.no_progress_count = self.no_progress_count.saturating_add(1);
                push_unique_signature(&mut self.no_change_signatures, signature);
            }
            CapabilityProgress::MadeProgress => {
                push_unique_signature(&mut self.made_progress_signatures, signature);
            }
        }
References
  1. Prefer exhaustive enum matching over runtime property-based checks for classifying variants when the set is small and known, as it leverages compiler-enforced safety and prevents loosening type constraints.

serrrfirat and others added 3 commits June 17, 2026 15:22
…nto firat/no-progress-pr2-content-digest

# Conflicts:
#	crates/ironclaw_product_workflow/tests/inbound_turn_contract.rs
#	crates/ironclaw_reborn/tests/loop_driver_host.rs
#	crates/ironclaw_reborn_composition/tests/product_live_adapters.rs
#	tests/support/reborn/harness.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5022 June 17, 2026 12:41 Destroyed
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

/benchmark pinchbench26 --framework ironclaw-reborn

@github-actions

Copy link
Copy Markdown
Contributor

🧪 Started pinchbench26 on ironclaw-reborn against ironclaw d562f9d9ff — watch run.

@github-actions

Copy link
Copy Markdown
Contributor

🧪 nearai-bench pinchbench26 — run complete (no baseline)

Ironclaw d562f9d9ff on ironclaw-reborn: 38.5% pass, avg score 0.780 across 26 tasks. No baseline exists under baselines/pinchbench26/ to compare against — add one (e.g. via the nightly refresh job) to enable regression detection.

🔍 browse run + per-task trajectories · download results

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

/benchmark pinchbench --framework ironclaw-reborn

@github-actions

Copy link
Copy Markdown
Contributor

🧪 Started pinchbench on ironclaw-reborn against ironclaw d562f9d9ff — watch run.

@github-actions

Copy link
Copy Markdown
Contributor

🧪 nearai-bench pinchbench — run complete (no baseline)

Ironclaw d562f9d9ff on ironclaw-reborn: 41.5% pass, avg score 0.738 across 147 tasks. No baseline exists under baselines/pinchbench/ to compare against — add one (e.g. via the nightly refresh job) to enable regression detection.

🔍 browse run + per-task trajectories · download results

Base automatically changed from firat/no-progress-pr2-content-digest to main June 17, 2026 22:24
…i-fix

# Conflicts:
#	crates/ironclaw_agent_loop/src/executor/capabilities.rs
#	crates/ironclaw_agent_loop/src/executor/loop_exit.rs
#	crates/ironclaw_agent_loop/src/executor/tests.rs
#	crates/ironclaw_agent_loop/tests/safety_nets.rs
#	crates/ironclaw_reborn_composition/src/product_live_adapters.rs
#	crates/ironclaw_reborn_composition/src/runtime/local_dev.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5022 June 17, 2026 22:47 Destroyed
@serrrfirat
serrrfirat merged commit d01df23 into main Jun 17, 2026
41 of 64 checks passed
@serrrfirat
serrrfirat deleted the firat/no-progress-pr3-output-aware branch June 17, 2026 22:50
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
* fix(agent-loop): no-progress stop fails honestly instead of faking completion

When the runaway-loop safety guard fires (StopKind::NoProgressDetected), the
executor finalized a canned "I stopped because I was repeating the same step"
assistant reply and returned the run as Completed — a runtime control decision
leaking into the conversation as a fake successful turn. This hid real blockers
(auth, broken connector, empty search) and marked incomplete tasks as done.

The ExitStage NoProgressDetected arm now:
- keeps the nearai#4837 final-answer-nudge path bit-for-bit: when the gate is enabled
  and the model synthesizes a real closing answer, complete with that answer
  (PinchBench path unchanged);
- otherwise writes the Final checkpoint and returns a typed
  failed_exit(LoopFailureKind::NoProgressDetected) instead of the canned reply.
  The product layer already maps "no_progress_detected" to deterministic copy,
  so the user sees an honest failure on every channel.

Deletes NO_PROGRESS_FALLBACK_REPLY and finalize_no_progress_fallback.

PR1 of the output-aware no-progress redesign (exit honesty). PR2 (content-digest
progress signal) and PR3 (failures off the no-progress axis) follow.

Wire behavior: gate-off no-progress runs now arrive as Failed{no_progress_detected}
instead of Completed with a canned reply.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(agent-loop): content-digest plumbing for output-aware progress (PR2)

Inert plumbing for output-aware no-progress detection (PR3 will consume it).
Adds a ContentDigest of each completed capability output so a later change can
treat "the output moved" as real progress. This PR does NOT change any
detection/stop behavior — it is observably inert.

- ContentDigest newtype (keyed-blake3, mirrors ArgsHash) in ironclaw_turns;
  normalize_for_hash moved there and re-exported from agent_loop::strategies so
  CapabilityCallSignature hashing is unchanged (no checkpoint-signature break).
- output_digest: Option<ContentDigest> on CapabilityResultMessage (serde default
  None = back-compat); computed host-side in the write_capability_result
  chokepoint alongside byte_len. Best-effort: a digest-compute failure degrades
  to None and never fails an otherwise-successful capability write.
- seen_capability_output_digests: BoundedRing<_, 64> on LoopExecutionState,
  populated at append_completed_capability_result but NOT consumed. Additive
  #[serde(default)] checkpoint field (legacy checkpoints decode to an empty
  ring; round-trip + legacy-decode tests added).
- record_result still receives the host's progress unchanged — detection is
  byte-identical (inertness test drives the executor and asserts no behavior
  change for repeated identical-output completed calls).

Review fixes folded in (multi-agent review): from_output is fail-open; a
caller-level test asserts the digest is recorded into the ring through a real
run; boundary note documents why the digest impl lives in ironclaw_turns.

PR2 of the no-progress redesign. PR1 (nearai#4993) = honest typed failure;
PR3 = consume the digest (NoChange when output repeats) + take failures off the
no-progress axis. Benchmark-gated at PR3.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(agent-loop): output-aware no-progress detection (PR3)

Consumes the PR2 content-digest plumbing to make "no progress" mean the tool
OUTPUT actually repeated, and takes failures off the no-progress axis.

- append_completed_capability_result now DERIVES progress from the digest: if the
  same call (signature) produced an output already seen this run -> NoChange; a
  first-seen output -> MadeProgress; no digest -> host-reported progress. The
  membership check runs BEFORE recording the observation. This flips PR2 from
  inert population to active output-aware detection.
- stop.rs counts ONLY NoChange toward the no-progress escape (Blocked dropped),
  and the repeated-call warning terminalizes only on a genuine NoChange repeat
  (new no_change_signatures track). Failing/blocked tools therefore route through
  recovery and the budget/iteration limit, never the no-progress give-up.

Net (PR1+PR2+PR3): no-progress fires on genuine output-stagnation (same call,
same output, repeated) as an honest typed failure; it no longer fires on a model
progressing with changing output (polling/pagination) or on failing tools, and
it never fakes a completion.

Tests: flipped PR2's inertness test (identical output now trips); added the
polling counterpart (changing output does not trip); blocked-failure tests now
assert the run recovers instead of a no-progress escape; Unknown progress no
longer terminalizes the repeated-call warning.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5022 — 967641fa Deployed Jun 17, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: M 50-199 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant