Skip to content

Fix checkpointless pre-model recovery - #6841

Merged
serrrfirat merged 12 commits into
mainfrom
codex/ws6-pre-model-recovery
Jul 29, 2026
Merged

serrrfirat merged 12 commits into
mainfrom
codex/ws6-pre-model-recovery

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Feed capability recovery strategies the exact bounded, provenance-tagged structured observation exposed to the model.
  • Automatically re-drive the same run for verified transient failures before the first BeforeModel checkpoint.
  • Bound checkpointless re-drive with the existing durable claim counter while preserving cancellation, checkpoint, identity, and failure-cause invariants.
  • Preserve the newly claimed lease when an older same-run executor task completes, so shutdown can safely relinquish the active attempt.
  • Centralize claimed-run requeue mutations, document the contract, and add lifecycle, durable-reopen, libSQL, and optional PostgreSQL parity regressions.

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Related #6284

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings — Not run repo-wide; cargo clippy -p ironclaw_runner -p ironclaw_turns --all-targets -- -D warnings passed after the review fixes, and the original touched-crate clippy pass included ironclaw_agent_loop.
  • cargo build — Not run; targeted crate tests and all-target clippy compiled every changed Rust path.
  • Relevant tests pass: the full scheduler contract and ironclaw_turns crate suites, including lifecycle, crash/reopen, and libSQL parity coverage.
  • cargo test --features integration if database-backed or integration behavior changed — Not applicable: ironclaw_turns has no integration feature and the integration harness seam did not change. Real libSQL coverage runs in the crate suite; the PostgreSQL parity test is opt-in when a test database URL is configured.
  • Manual testing — Not applicable: deterministic scheduler/store behavior is covered at production seams.
  • If a coding agent was used and supports it, review-pr or pr-shepherd --fix was run before requesting review — Eight-lens code-review-multi completed with full packet coverage; all six deduplicated findings were fixed.

Test Strategy

User behavior: A first-iteration transient input, prompt/context, or capability-surface construction failure is retried automatically as the same run. Persistent failure stops at the bounded attempt limit with the original safe cause. Cancellation and any recorded loop checkpoint prevent scratch re-drive. Shutdown relinquishes the currently active same-run attempt rather than losing its lease identity.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: structured capability observation through CanonicalAgentLoopExecutor; pre-model category allowlist; same-run scheduler recovery; bounded exhaustion; cancellation precedence; checkpoint side-effect guard; exact lease identity during same-run redrive shutdown.
  • Reborn integration: Not applicable: the integration harness has no injectable pre-prompt transient seam; caller behavior is covered through the real canonical executor and scheduler/store production ports.
  • Recorded fixture: Not applicable: no provider payload, transcript fixture, or replay format changed.
  • Browser E2E: Not applicable: no UI or browser behavior changed.
  • Backend or runtime: row-store graceful reopen preserves queued identity and claim count; checkpoint and claim-bound terminal guards survive reopen; queued lifecycle publication is asserted; real libSQL parity covers requeue/reopen/exhaustion; matching PostgreSQL parity runs when a test database is configured.
  • Live canary: Not applicable: no external service or credential-dependent behavior changed.

What the tests prove:

  • Recovery strategy decisions receive the same structured, bounded observation appended for the model.
  • Eligible first-iteration failures re-drive the identical run and accepted_message_ref without creating a duplicate run.
  • The existing durable claim counter bounds automatic retry and exhaustion preserves category plus redacted detail.
  • Cancellation wins over retry, and any durable loop checkpoint blocks scratch re-execution.
  • Queued recovery identity, lifecycle classification, and retry bounds survive graceful reopen.
  • Completion of an older same-run task cannot erase the newer lease used for shutdown relinquishment.
  • libSQL preserves the same requeue/reopen/identity/exhaustion contract; PostgreSQL has the same opt-in parity test.

Commands run:

  • cargo test -p ironclaw_agent_loop invalid_provider_tool_failure_appends_structured_model_observation
  • cargo test -p ironclaw_agent_loop --lib strategies::recovery::tests
  • cargo test -p ironclaw_runner --test turn_scheduler_contract — 34 passed
  • cargo test -p ironclaw_turns — all unit, contract, crash-consistency, and libSQL parity tests passed
  • cargo test -p ironclaw_runner --test turn_scheduler_contract shutdown_relinquishes_the_active_lease_after_same_run_redrive -- --exact — also repeated 50 times after the fix
  • cargo clippy -p ironclaw_agent_loop -p ironclaw_turns -p ironclaw_runner --all-targets -- -D warnings — original change
  • cargo clippy -p ironclaw_runner -p ironclaw_turns --all-targets -- -D warnings — review fixes
  • cargo fmt --all -- --check
  • git diff --check

Security Impact

None. The change does not expand authority, permissions, network access, secret handling, or sandbox policy. Recovery receives the existing sanitized model-visible observation rather than raw host/provider detail.

Reborn Trust-Boundary Checklist

  • Public policy/evidence/trust-bearing types: RunnerFailureRecovery is selected by the trusted scheduler; the store remains authoritative for cancellation, checkpoint, lease, and bound validation.
  • Untrusted content enters prompts only through an envelope/escaping primitive. No new prompt ingress is added; the strategy receives the already-bounded model-visible observation.
  • Hashes declare purpose; trust/binding/authenticity uses SHA-256/BLAKE3 or separate authenticity check. N/A: no hashes changed.
  • New/changed status, exit, policy, runtime, or error variants: downstream match sites audited. Command/output: rg -n "RecordRunnerFailureRequest|RunnerFailureRecovery|RedriveIfCheckpointless" crates and all constructors/production forwarders were reviewed.
  • Security/durability serde(default) fields fail closed or have migration tests. RunnerFailureRecovery defaults to Terminal; the request is not persisted as durable state.
  • Queues/maps/buffers/counters have bounds and overflow-safe arithmetic. Re-drive reuses the existing durable claim_count < max_crash_recovery_reclaims bound.
  • Driver/operator-visible errors have stable class semantics (Transient, Permanent, Misconfigured, PolicyDenied or equivalent). Only the three stable pre-model host_stage_unavailable_* categories opt in; exhaustion preserves the original sanitized category/detail.
  • Sandbox/native/host names accurately describe trust boundary. N/A: no sandbox or native-host boundary changed.

Database Impact

No migration or schema change. Durable row-store transition semantics change for eligible checkpointless runner failures; libSQL parity passed locally and matching PostgreSQL parity is included for configured CI/developer environments.

Blast Radius

Touches capability recovery strategy inputs, runner executor-failure settlement, scheduler active-lease tracking, and turn-state row-store transitions. The principal risks are accidental scratch execution after work has begun, duplicate queueing, lost identity, stale lifecycle projection, or unbounded retries; cancellation/checkpoint/identity/bound/event/reopen/backend tests cover those seams.

Rollback Plan

Revert commits 1074646aa and 0f5ad2f32. Existing behavior will return: graceful pre-model executor failures terminalize immediately, while manual retry and expired-lease checkpointless recovery remain available.

Review Follow-Through

Eight isolated reviewers covered both complete diff packets. The review found one high-confidence same-run lease race plus test, contract, backend-parity, and duplication gaps; all were fixed. Reviewer judgment remains useful on the deliberately narrow pre-model category allowlist and reuse of max_crash_recovery_reclaims as the shared automatic-redrive bound. No epic boxes were edited; the epic's literal “no retry path” premise is stale after #6295 and #6376, while this narrower graceful-failure gap remained live.


Review track: C (runtime/persistence recovery behavior)

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@coderabbitai

coderabbitai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c6debf2c-9e6c-433f-81c7-ac7ae000c53c

📥 Commits

Reviewing files that changed from the base of the PR and between 89e9a00 and b1e3f52.

📒 Files selected for processing (7)
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_reborn_composition/src/runtime/tests/core.rs
  • crates/ironclaw_turns/src/turn_state_row_store/row_store/commit.rs
  • crates/ironclaw_turns/src/turn_state_row_store/row_store/traits.rs
  • crates/ironclaw_turns/tests/row_store_crash_consistency.rs
  • crates/ironclaw_turns/tests/runner_failure_backend_parity.rs
  • scripts/no_panics_reborn_baseline.txt
💤 Files with no reviewable changes (1)
  • scripts/no_panics_reborn_baseline.txt

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Added safe, bounded re-drive for eligible transient failures occurring before model processing.
    • Preserves run identity and accepted messages during recovery, including across shutdowns and restarts.
    • Prevents re-drive after checkpoints and prioritizes cancellation when applicable.
    • Recovery strategies can now use structured tool observations for improved error handling.
  • Bug Fixes

    • Prompt-stage failures now retain specific categories and sanitized diagnostics.
    • Permanent prompt errors are reported accurately instead of being treated as temporary unavailability.
    • Runner failure states and lease handling are now more consistent during retries and shutdowns.
  • Documentation

    • Documented checkpointless pre-model failure re-drive rules and safeguards.

Walkthrough

The PR adds structured capability observations to recovery callbacks, preserves typed and sanitized prompt-stage diagnostics, and introduces bounded checkpointless runner-failure redrive with durable identity, lease, cancellation, checkpoint, and backend coverage.

Changes

Recovery flow

Layer / File(s) Summary
Capability observation and prompt diagnostics
crates/ironclaw_agent_loop/src/strategies/recovery.rs, crates/ironclaw_agent_loop/src/executor/*, crates/ironclaw_agent_loop/src/executor/tests/*
Recovery callbacks receive optional model-visible observations; prompt failures preserve typed diagnostics, cancellation, and redaction behavior.
Prompt error classification
crates/ironclaw_runner/src/planned_driver.rs, crates/ironclaw_runner/tests/planned_driver_e2e.rs
Prompt-stage policy and invalid-invocation failures map to permanent driver categories while other stages retain unavailable mapping.
Runner recovery state transitions
crates/ironclaw_turns/src/runner.rs, crates/ironclaw_turns/src/turn_state_row_store/**/*
Runner failure requests carry terminal or checkpointless-redrive policy; eligible failures requeue while checkpointed and cancelled runs remain terminal.
Scheduler recovery policy
crates/ironclaw_runner/src/turn_scheduler.rs, crates/ironclaw_runner/src/turn_run_executor.rs, crates/ironclaw_runner/src/loop_exit_applier/tests/support.rs
The scheduler selects redrive for verified pre-model categories, records the policy, and preserves runner and lease identity through shutdown handling.
Redrive and compatibility validation
crates/ironclaw_runner/tests/turn_scheduler_contract.rs, crates/ironclaw_turns/tests/*, docs/reborn/contracts/turn-runner.md, crates/ironclaw_reborn_composition/..., scripts/no_panics_reborn_baseline.txt
Tests and contract documentation cover bounded retries, identity preservation, checkpoint and cancellation precedence, reopen durability, backend parity, and lifecycle events.

Estimated code review effort: 4 (Complex) | ~60 minutes

Possibly related issues

Possibly related PRs

  • nearai/ironclaw#6792 — Constructs denial-specific model-visible observations consumed by this recovery plumbing.
  • nearai/ironclaw#6840 — Modifies adjacent recovery behavior using model-visible observations.
  • nearai/ironclaw#6467 — Changes adjacent capability-error recovery handling at the same call site.

Suggested reviewers: ilblackdragon

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly matches the PR’s main change: checkpointless pre-model recovery.
Description check ✅ Passed The description is complete and matches the repository template with summary, validation, test strategy, impact, and rollout details.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6841 July 29, 2026 08:29 Destroyed
@github-actions github-actions Bot added size: M 50-199 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 29, 2026
@ironloopai

ironloopai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #6841

🔴 Failed

Execution result is invalid

The structured result could not be verified.

Automatic · PR opened · attempt 1 of 3 · failed after 3m 33s

Failure details
  • Repository: nearai/ironclaw
  • Base: main at a0e91d1
  • Head: codex/ws6-pre-model-recovery at 0f5ad2f
  • Created: Jul 29, 2026, 8:34 AM UTC
  • Updated: Jul 29, 2026, 8:37 AM UTC
  • Run: 82a2085d-be89-47d9-9452-6590a5373dfb
  • Latest attempt: 1 · Completed · 02920582-bbda-42be-82c0-a607e9bbca3b
  • Failed during: Verification
  • Retryable: No
  • Failure: 6530067a-7d4c-4fc1-bd16-2758f17c706d

@railway-app

railway-app Bot commented Jul 29, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6841 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 29, 2026 at 7:34 pm

@serrrfirat
serrrfirat marked this pull request as ready for review July 29, 2026 11:01
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6841 July 29, 2026 11:01 Destroyed
@ironloopai

ironloopai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #6841

🟢 Completed · Review submitted

Submitted review →

Reviewed the complete trusted base-to-head comparison. No concrete correctness, security, concurrency, persistence, architecture, or test-coverage defects were found. The checkpointless re-drive is limited to explicit pre-model failure categories and remains lease-validated, cancellation-first, checkpoint-gated, same-run, and durably bounded by claim count.

Automatic · PR opened · attempt 1 of 3 · completed in 1m 29s

Run details
  • Repository: nearai/ironclaw
  • Base: main at de34247
  • Head: codex/ws6-pre-model-recovery at 9bdb802
  • Created: Jul 29, 2026, 11:06 AM UTC
  • Updated: Jul 29, 2026, 11:08 AM UTC
  • Run: a4e867f5-afd1-4234-adec-a264828f59ec
  • Latest attempt: 1 · Completed · 3dfa3018-c3a1-4371-89bf-68898f1bbf81

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Review complete · PR #6841

✅ No actionable findings

Reviewed the complete trusted base-to-head comparison. No concrete correctness, security, concurrency, persistence, architecture, or test-coverage defects were found. The checkpointless re-drive is limited to explicit pre-model failure categories and remains lease-validated, cancellation-first, checkpoint-gated, same-run, and durably bounded by claim count.

Validation and technical details
  • Verified trusted comparison refs: de34247..9bdb802.
  • Inspected all 15 changed files and surrounding scheduler, executor, lease-retirement, row-store commit, checkpoint, queue, lifecycle-event, snapshot-reopen, and recovery-limit code.
  • Enumerated all RecordRunnerFailureRequest constructors and verified explicit terminal/re-drive disposition at production and test call sites.
  • Verified added tests cover structured model observation, category allowlisting, same-run identity, bounded exhaustion with preserved failure cause, cancellation precedence, checkpoint prevention, and durable reopen.
  • git diff --check refs/ironloop/base..refs/ironloop/head passed.
  • Scoped cargo tests could not be executed in this review environment because cargo is unavailable (/bin/bash: cargo: command not found).
  • Base: main
  • Head: codex/ws6-pre-model-recovery at 9bdb802
  • Run: a4e867f5-afd1-4234-adec-a264828f59ec

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_turns/src/turn_state_row_store/row_store/traits.rs`:
- Around line 755-764: The runner-leaving bookkeeping in record_runner_failure
currently derives retired_status from request.recovery, which incorrectly forces
RedriveIfCheckpointless to Queued. Update the record_runner_failure transition
flow to derive the status passed to apply_run_state_transition from the
TurnRunState returned by runner_failure_transition, preserving resolved Failed
or Cancelled outcomes. Cover checkpointed, exhausted, and cancel-requested
RedriveIfCheckpointless cases across row-store and partition paths.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 230fa7aa-95ab-469c-a506-b6f4050cb46f

📥 Commits

Reviewing files that changed from the base of the PR and between de34247 and 9bdb802.

📒 Files selected for processing (15)
  • crates/ironclaw_agent_loop/src/executor/capabilities.rs
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_agent_loop/src/executor/tests/support.rs
  • crates/ironclaw_agent_loop/src/strategies/recovery.rs
  • crates/ironclaw_runner/src/loop_exit_applier/tests/support.rs
  • crates/ironclaw_runner/src/turn_run_executor.rs
  • crates/ironclaw_runner/src/turn_scheduler.rs
  • crates/ironclaw_runner/src/turn_scheduler/tests.rs
  • crates/ironclaw_runner/tests/turn_scheduler_contract.rs
  • crates/ironclaw_turns/src/runner.rs
  • crates/ironclaw_turns/src/turn_state_row_store/row_store/traits.rs
  • crates/ironclaw_turns/src/turn_state_row_store/turn_state_engine/mod.rs
  • crates/ironclaw_turns/src/turn_state_row_store/turn_state_engine/transitions.rs
  • crates/ironclaw_turns/tests/row_store_crash_consistency.rs
  • crates/ironclaw_turns/tests/turn_coordinator_contract.rs

Comment thread crates/ironclaw_turns/src/turn_state_row_store/row_store/traits.rs Outdated
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6841 July 29, 2026 11:17 Destroyed
@github-actions github-actions Bot added size: M 50-199 changed lines scope: docs Documentation size: L 200-499 changed lines and removed size: M 50-199 changed lines labels Jul 29, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@crates/ironclaw_turns/src/turn_state_row_store/turn_state_engine/transitions.rs`:
- Around line 965-979: Update the TurnStatus::CancelRequested transition to call
the existing release_terminal_lease helper instead of manually performing
clear_runner_lease, cursor advancement, active-lock release, and queued-run
removal. Preserve the existing cancellation status, failure reset, event, and
terminal-marking behavior while reusing the helper’s shared terminal-transition
tail.

In `@crates/ironclaw_turns/tests/runner_failure_backend_parity.rs`:
- Around line 48-71: Update build_postgres_scoped so only the initial
environment-variable lookup can return None when PostgreSQL is not configured.
After a URL is present, replace the .ok()? error suppression for URL parsing,
pool construction, and root.run_migrations() with explicit fail-loud handling
such as expect messages, while preserving the existing optional skip behavior
for missing configuration.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: e3950054-0f11-4f93-bcc8-5f0ceac789ab

📥 Commits

Reviewing files that changed from the base of the PR and between 9bdb802 and 1074646.

📒 Files selected for processing (7)
  • crates/ironclaw_runner/src/turn_scheduler.rs
  • crates/ironclaw_runner/tests/turn_scheduler_contract.rs
  • crates/ironclaw_turns/src/turn_state_row_store/turn_state_engine/transitions.rs
  • crates/ironclaw_turns/tests/row_store_crash_consistency.rs
  • crates/ironclaw_turns/tests/runner_failure_backend_parity.rs
  • crates/ironclaw_turns/tests/turn_coordinator_contract.rs
  • docs/reborn/contracts/turn-runner.md

Comment thread crates/ironclaw_turns/tests/runner_failure_backend_parity.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6841 July 29, 2026 12:11 Destroyed
@github-actions github-actions Bot removed the size: M 50-199 changed lines label Jul 29, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_agent_loop/src/executor/tests.rs`:
- Around line 2129-2162: Extend the PromptStage tests around
prompt_stage_preserves_policy_denied_kind_from_prompt_bundle to cover
prompt_host_error’s remaining branches: verify a prompt-bundle
AgentLoopHostErrorKind::Cancelled becomes AgentLoopExecutorError::Cancelled, and
inject a full host error whose safe_summary fallback is rejected to verify the
sanitized error behavior. Update MockHost only as needed to supply that full
AgentLoopHostError, while preserving the existing policy-denied case.

In `@crates/ironclaw_runner/src/planned_driver.rs`:
- Around line 348-356: The permanent failure mapping around
permanent_host_stage_failure_category must be scoped to the intended
prompt-stage cases instead of treating every non-Model stage as terminal. Update
the stage/kind handling in the surrounding error path to preserve capability,
checkpoint, transcript, and input failure kinds and durable state, or define and
test an explicit stage/kind matrix covering them.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 79ec6ce8-2f3f-4e08-8512-853671405fd7

📥 Commits

Reviewing files that changed from the base of the PR and between 1074646 and ba2f1a2.

📒 Files selected for processing (5)
  • crates/ironclaw_agent_loop/src/executor/prompt.rs
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_agent_loop/src/executor/tests/support.rs
  • crates/ironclaw_reborn_composition/src/runtime/tests/core.rs
  • crates/ironclaw_runner/src/planned_driver.rs

Comment thread crates/ironclaw_agent_loop/src/executor/tests.rs
Comment thread crates/ironclaw_runner/src/planned_driver.rs Outdated
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6841 July 29, 2026 12:27 Destroyed
@github-actions

github-actions Bot commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.82% (316906 / 369259 lines)
  floor:    80.81% (tolerance 0.5pp -> effective floor 80.31%)
  denominator: 369259 lines now vs 377084 at floor capture (-7825 lines, -2.08%) — not a material change

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.82% — 316906 / 369259 lines

Per-crate breakdown (61 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_process_sandbox 33.91% 118 / 348
ironclaw_host_ingress 42.5% 17 / 40
ironclaw_event_projections 43.51% 684 / 1572
ironclaw_libsql_runtime 59.91% 127 / 212
ironclaw_observability 61.54% 16 / 26
ironclaw_authorization 63.02% 610 / 968
ironclaw_memory 64.41% 959 / 1489
ironclaw_telegram_v2_adapter 69.69% 731 / 1049
ironclaw_trust 73.21% 664 / 907
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_extractors 74.72% 538 / 720
ironclaw_filesystem 75.03% 4975 / 6631
ironclaw_capabilities 75.3% 2878 / 3822
ironclaw_projects 76.48% 400 / 523
ironclaw_mcp 76.6% 779 / 1017
ironclaw_reborn_cli 78.36% 10782 / 13760
ironclaw_wasm 79.72% 735 / 922
ironclaw_llm 80.14% 22266 / 27783
ironclaw_memory_native 81.02% 3299 / 4072
ironclaw_auth 81.88% 6679 / 8157
ironclaw_first_party_extensions 82.38% 6682 / 8111
ironclaw_processes 83.3% 933 / 1120
ironclaw_host_api 83.7% 9722 / 11615
ironclaw_reborn_identity 83.8% 450 / 537
ironclaw_events 84.06% 1687 / 2007
ironclaw_operator 84.41% 5561 / 6588
ironclaw_secrets 84.56% 2798 / 3309
ironclaw_network 85.07% 957 / 1125
ironclaw_reborn_config 85.23% 2101 / 2465
ironclaw_skills 85.27% 4493 / 5269
ironclaw_telegram_extension 85.4% 1299 / 1521
ironclaw_extension_host 85.6% 19504 / 22786
ironclaw_run_state 85.77% 458 / 534
ironclaw_reborn_composition 86% 24847 / 28893
ironclaw_triggers 86.11% 2803 / 3255
ironclaw_webui 86.37% 11171 / 12934
ironclaw_hooks 86.72% 9949 / 11472
ironclaw_common 86.99% 1772 / 2037
ironclaw_approvals 87.07% 1542 / 1771
ironclaw_reborn_event_store 87.19% 1327 / 1522
ironclaw_threads 87.36% 4912 / 5623
ironclaw_extensions 87.45% 4927 / 5634
ironclaw_product 87.57% 20357 / 23247
ironclaw_reborn_traces 88.13% 11987 / 13601
ironclaw_turns 88.45% 14734 / 16658
ironclaw_host_runtime 89% 20499 / 23032
ironclaw_slack_extension 89.26% 2027 / 2271
ironclaw_reborn_openai_compat 89.32% 3780 / 4232
ironclaw_conversations 90.03% 3171 / 3522
ironclaw_resources 90.84% 4474 / 4925
ironclaw_event_streams 91.24% 1063 / 1165
ironclaw_runner 91.66% 17726 / 19338
ironclaw_loop_host 92.01% 16650 / 18095
ironclaw_attachments 93.06% 630 / 677
ironclaw_outbound 93.91% 4101 / 4367
ironclaw_agent_loop 94.52% 10081 / 10665
ironclaw_safety 95.22% 3941 / 4139
ironclaw_first_party_extension_ports 95.62% 3672 / 3840
ironclaw_runtime_policy 96.56% 814 / 843

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6841 July 29, 2026 13:16 Destroyed
@github-actions github-actions Bot added size: XL 500+ changed lines and removed size: L 200-499 changed lines labels Jul 29, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6841 July 29, 2026 13:23 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_turns/src/turn_state_row_store/turn_state_engine/transitions.rs (1)

891-947: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Checkpointless redrive discards the triggering failure — no audit trail until retries exhaust.

runner_failure_transition's RedriveIfCheckpointless branch (937-943) calls requeue_checkpointless_runner_failure(record), but that function (905-916) never takes the failure: SanitizedFailure the caller received — it's simply dropped. requeue_claimed_record (891-903) always pushes TurnEventKind::RunnerHeartbeat with None, None for category/detail (line 902). So a runner-reported transient crash that triggers a redrive is durably indistinguishable from a normal heartbeat — operators get zero signal about why a run is being retried, until (if ever) retries exhaust into Failed. This directly undercuts "preserve observable failure kinds ... audit error_kind" — the failure detail exists (SanitizedFailure is right there in the caller) and is being thrown away rather than recorded.

fail_claimed_record/terminal_transition already thread failure.into_category() / failure.detail() into push_event's 3rd/4th params — reuse that pattern here instead of inventing a new event kind.

🐛 Proposed fix — carry failure category/detail through the requeue event
-    fn requeue_claimed_record(&mut self, record: &mut RunRecord, now: DateTime<Utc>) {
+    fn requeue_claimed_record(
+        &mut self,
+        record: &mut RunRecord,
+        now: DateTime<Utc>,
+        redrive_cause: Option<&SanitizedFailure>,
+    ) {
         let transition = record.status.set(TurnStatus::Queued);
         self.apply_status_transition(transition, record);
         record.failure = None;
         clear_runner_lease(record);
         record.event_cursor = self.next_cursor();
         self.update_active_lock(record, now);
         self.queued_runs.push_back(record.run_id);
-        // Running → Queued uses the same lifecycle classification for graceful
-        // relinquish, checkpointless failure re-drive, and expired-lease
-        // recovery so the durable log and publishing wrapper stay aligned.
-        self.push_event(record, TurnEventKind::RunnerHeartbeat, None, None);
+        // Running → Queued keeps the same event KIND for graceful relinquish,
+        // checkpointless failure re-drive, and expired-lease recovery so the
+        // publishing wrapper stays aligned, but a failure-driven redrive still
+        // records its cause via category/detail.
+        let (category, detail) = redrive_cause
+            .map(|failure| (Some(failure.category().to_string()), failure.detail().map(str::to_string)))
+            .unwrap_or((None, None));
+        self.push_event(record, TurnEventKind::RunnerHeartbeat, category, detail);
     }
 
     fn requeue_checkpointless_runner_failure(
         &mut self,
         mut record: RunRecord,
+        failure: SanitizedFailure,
     ) -> AppliedLoopTransition {
-        self.requeue_claimed_record(&mut record, Utc::now());
+        self.requeue_claimed_record(&mut record, Utc::now(), Some(&failure));
         ...
     }

Other two call sites (recover_expired_leases, relinquish_transition) pass None and keep today's behavior unchanged.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@crates/ironclaw_turns/src/turn_state_row_store/turn_state_engine/transitions.rs`
around lines 891 - 947, Preserve the triggering SanitizedFailure when
checkpointless redrive occurs: update requeue_checkpointless_runner_failure and
requeue_claimed_record to accept and forward the failure category and detail,
and have the requeue event use failure.into_category() and failure.detail()
instead of None values. Update only the runner_failure_transition call site for
this failure-aware path; keep recover_expired_leases and relinquish_transition
behavior unchanged.

Source: Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In
`@crates/ironclaw_turns/src/turn_state_row_store/turn_state_engine/transitions.rs`:
- Around line 891-947: Preserve the triggering SanitizedFailure when
checkpointless redrive occurs: update requeue_checkpointless_runner_failure and
requeue_claimed_record to accept and forward the failure category and detail,
and have the requeue event use failure.into_category() and failure.detail()
instead of None values. Update only the runner_failure_transition call site for
this failure-aware path; keep recover_expired_leases and relinquish_transition
behavior unchanged.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 5a44e920-b42b-4efd-9fb8-de2a18243ada

📥 Commits

Reviewing files that changed from the base of the PR and between 3dfa00d and b4d345e.

📒 Files selected for processing (7)
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_agent_loop/src/executor/tests/support.rs
  • crates/ironclaw_runner/src/planned_driver.rs
  • crates/ironclaw_turns/src/turn_state_row_store/row_store/commit.rs
  • crates/ironclaw_turns/src/turn_state_row_store/row_store/traits.rs
  • crates/ironclaw_turns/src/turn_state_row_store/turn_state_engine/transitions.rs
  • crates/ironclaw_turns/tests/runner_failure_backend_parity.rs

…ecovery

# Conflicts:
#	crates/ironclaw_agent_loop/src/strategies/recovery.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6841 July 29, 2026 14:07 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6841 July 29, 2026 15:12 Destroyed
…ecovery

# Conflicts:
#	crates/ironclaw_agent_loop/src/executor/capabilities.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6841 July 29, 2026 15:38 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
crates/ironclaw_runner/src/planned_driver.rs (2)

346-354: 🔒 Security & Privacy | 🔵 Trivial | ⚡ Quick win

Centralize the diagnostic-scrubbing path.

This branch duplicates the fallback, scrub_model_visible_detail, and Failed construction already used above. Extract one helper so model- and prompt-stage failures cannot diverge on redaction behavior.

As per coding guidelines, extract helpers when logic is reused.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_runner/src/planned_driver.rs` around lines 346 - 354, Extract
a shared helper for constructing scrubbed AgentLoopDriverError::Failed values,
including the detail fallback to safe_summary and
ironclaw_loop_host::scrub_model_visible_detail. Update this prompt-stage failure
branch and the existing model-failure path above to use the helper, preserving
their respective category and diagnostic inputs.

Source: Coding guidelines


373-379: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win

Complete regression coverage for the new terminal mappings.

The added tests cover Prompt/PolicyDenied and non-Prompt rejection, but do not cover RecoverySequenceExhausted or the positive Prompt mappings for InvalidInvocation, Invalid, and ScopeMismatch. Add caller-level assertions for these contracts.

As per coding guidelines, new production-wired behavior requires caller-level regression tests.

Also applies to: 759-815

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_runner/src/planned_driver.rs` around lines 373 - 379, Add
caller-level regression tests for the terminal mappings in the planned driver
flow: assert RecoverySequenceExhausted produces Failed with reason_kind
"driver_bug", and assert Prompt requests map InvalidInvocation, Invalid, and
ScopeMismatch to their expected positive outcomes. Extend the existing tests
around the caller handling these AgentLoopExecutorError variants, preserving the
current Prompt/PolicyDenied and non-Prompt rejection coverage.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@crates/ironclaw_runner/src/planned_driver.rs`:
- Around line 346-354: Extract a shared helper for constructing scrubbed
AgentLoopDriverError::Failed values, including the detail fallback to
safe_summary and ironclaw_loop_host::scrub_model_visible_detail. Update this
prompt-stage failure branch and the existing model-failure path above to use the
helper, preserving their respective category and diagnostic inputs.
- Around line 373-379: Add caller-level regression tests for the terminal
mappings in the planned driver flow: assert RecoverySequenceExhausted produces
Failed with reason_kind "driver_bug", and assert Prompt requests map
InvalidInvocation, Invalid, and ScopeMismatch to their expected positive
outcomes. Extend the existing tests around the caller handling these
AgentLoopExecutorError variants, preserving the current Prompt/PolicyDenied and
non-Prompt rejection coverage.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 65fc8d9a-3c3b-4a4e-b606-b18521f0e241

📥 Commits

Reviewing files that changed from the base of the PR and between 3350371 and 89e9a00.

📒 Files selected for processing (4)
  • crates/ironclaw_agent_loop/src/executor/capabilities.rs
  • crates/ironclaw_agent_loop/src/executor/checkpoint.rs
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_runner/src/planned_driver.rs

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6841 July 29, 2026 19:22 Destroyed
@serrrfirat
serrrfirat merged commit 2cc9408 into main Jul 29, 2026
65 checks passed
@serrrfirat
serrrfirat deleted the codex/ws6-pre-model-recovery branch July 29, 2026 20:22
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
* fix checkpointless pre-model recovery

* fix(runner): preserve active redrive lease identity

* fix(runner): keep permanent prompt failures terminal

* test(runner): expect preserved prompt summary

* fix(runner): address coderabbit recovery review (nearai#6841)

* chore(ci): ratchet removed runner panic baseline

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6841 — b1e3f52b Deployed Jul 29, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant