From 293ade150f9e6512db75d1ecf8b98908362dad94 Mon Sep 17 00:00:00 2001 From: italic-jinxin <106428113+italic-jinxin@users.noreply.github.com> Date: Wed, 22 Jul 2026 20:15:40 +0800 Subject: [PATCH 1/3] docs(reborn): reconcile self-repair lease recovery --- .../tests/row_store_crash_consistency.rs | 17 ++--- docs/reborn/contracts/turn-runner.md | 43 +++++++++-- .../tier-b-self-repair-reconciliation.md | 75 +++++++++++++++++++ 3 files changed, 115 insertions(+), 20 deletions(-) create mode 100644 docs/reborn/tier-b-self-repair-reconciliation.md diff --git a/crates/ironclaw_turns/tests/row_store_crash_consistency.rs b/crates/ironclaw_turns/tests/row_store_crash_consistency.rs index 179686b2c00..4252f9d819d 100644 --- a/crates/ironclaw_turns/tests/row_store_crash_consistency.rs +++ b/crates/ironclaw_turns/tests/row_store_crash_consistency.rs @@ -2015,18 +2015,11 @@ async fn byte_state_fork_recovers_to_model_snapshot() { /// it deterministically (identically to the never-crashed reference model), /// with any cause preserved. /// -/// NOTE on the literal coordinator ask ("recover_expired_leases returns the run -/// to a *claimable* state; never Failed"): the shared turn engine does NOT -/// re-queue an expired lease — the store`s `recover_expired_leases` -/// terminates the abandoned run as `Failed(lease_expired)` (a -/// resumable-checkpointed run keeps its checkpoint and is retryable; a -/// checkpoint-less one does not — see the ignored reproducer below). That is -/// pre-existing engine semantics, identical in the direct authority and the row -/// store, and is unaffected by whether a crash occurred. This suite's charge is -/// row-store *crash-consistency*, so the assertion here is the defensible one: -/// the crash neither loses the run nor diverges its recovered outcome from the -/// model. The lifecycle question (should lease expiry be terminal at all under -/// #6284?) is captured, reproducibly, in the ignored test below. +/// The crash itself preserves the durable `Running` state. A later lease-recovery +/// pass resolves both the row store and its reference model identically. The +/// caller-level tests below separately prove the two current #6284 outcomes: +/// checkpointless work is requeued while its reclaim budget remains, and it +/// terminal-fails with `crash_retry_exhausted` once that bound is reached. #[tokio::test] async fn crash_mid_run_recovers_identically_to_model_and_preserves_cause() { let backend = Arc::new(FaultBackend::new(InMemoryBackend::new())); diff --git a/docs/reborn/contracts/turn-runner.md b/docs/reborn/contracts/turn-runner.md index 998b94d6d13..8fd2de4b16d 100644 --- a/docs/reborn/contracts/turn-runner.md +++ b/docs/reborn/contracts/turn-runner.md @@ -2,13 +2,14 @@ **Status:** Contract-freeze draft **Date:** 2026-05-06 +**Last reconciled:** 2026-07-22 (#6455) **Depends on:** [`turn-persistence.md`](turn-persistence.md), [`turns-agent-loop.md`](turns-agent-loop.md), [`loop-exit.md`](loop-exit.md), [`runtime-profiles.md`](runtime-profiles.md) --- ## 1. Purpose -`TurnRunner` is the trusted worker-side control plane for executable turn runs. It claims queued runs, maintains leases while model/tool work is active, records safe checkpoint/block/terminal transitions, and moves abandoned work to explicit recovery instead of blindly retrying side effects. +`TurnRunner` is the trusted worker-side control plane for executable turn runs. It claims queued runs, maintains leases while model/tool work is active, records safe checkpoint/block/terminal transitions, and moves abandoned work through checkpoint-aware, bounded recovery instead of blindly retrying uncertain side effects. Product adapters must continue to use `TurnCoordinator`. Runner transition APIs are trusted-worker APIs and remain under `ironclaw_turns::runner`. Driver-facing loop exits remain distinct from trusted runner outcomes; see [`loop-exit.md`](loop-exit.md). @@ -19,7 +20,7 @@ Product adapters must continue to use `TurnCoordinator`. Runner transition APIs - `submit_turn` creates a queued `TurnRunId` and active-thread lock, but no model/tool side effects may run before a runner claim succeeds. - `claim_next_run` atomically moves one matching `Queued` run to `Running`. - A successful claim stores `runner_id`, `lease_token`, `last_heartbeat_at`, `lease_expires_at`, increments `claim_count`, updates the active lock, and emits `RunnerClaimed`. -- `heartbeat` requires the matching `runner_id` and `lease_token`, only refreshes actively `Running` work, and rejects leases whose `lease_expires_at` has already passed. Once cancellation is requested, heartbeats no longer extend the lease; the runner must complete cancellation before the existing lease expires or the reconciler moves the run to recovery. On success, heartbeat refreshes durable `last_heartbeat_at` and extends durable `lease_expires_at`; adapters may touch active-lock freshness and emit/coalesce `RunnerHeartbeat` lifecycle events, but consumers must use lease metadata as the liveness source of truth. +- `heartbeat` requires the matching `runner_id` and `lease_token`, only refreshes actively `Running` work, and rejects leases whose `lease_expires_at` has already passed. Once cancellation is requested, heartbeats no longer extend the lease; the runner must complete cancellation before the existing lease expires or the reconciler terminalizes the run as `Cancelled`. On success, heartbeat refreshes durable `last_heartbeat_at` and extends durable `lease_expires_at`; adapters may touch active-lock freshness and emit/coalesce `RunnerHeartbeat` lifecycle events, but consumers must use lease metadata as the liveness source of truth. - Pull-based claims are authoritative. Wake notifications are optimization hints only. - After `TurnCoordinator` durably accepts a submitted run or requeues a resumed/retried run, it may emit a redacted queued-run wake hint containing only the canonical scope, `TurnRunId`, queued status, and event cursor. Wake delivery is best-effort, is not a source of truth, must not fail the durable adapter call, and duplicate hints must be harmless. @@ -28,10 +29,13 @@ Product adapters must continue to use `TurnCoordinator`. Runner transition APIs ## 3. Expired lease recovery - A reconciler scans runner-owned `Running` and `CancelRequested` leases using durable `lease_expires_at` metadata. -- Expired `Running` or `CancelRequested` leases transition to `RecoveryRequired`, clear current runner ownership, emit a redacted `RecoveryRequired` event with reason `lease_expired`, and keep the same canonical-thread active lock. -- `RecoveryRequired` runs are not returned by the normal `claim_next_run` path. The system must not auto-retry uncertain side-effecting work. -- A duplicate/new submit for the same canonical thread remains `ThreadBusy` while recovery is required. -- Explicit cancellation of `RecoveryRequired` is terminal `Cancelled` and releases the active lock so a new turn can be submitted. +- An expired `CancelRequested` lease becomes terminal `Cancelled`, clears runner ownership, and releases the canonical-thread active lock. +- An expired `Running` lease with any loop checkpoint becomes terminal `Failed(lease_expired)`. The latest resumable checkpoint is attached when one exists, making the normal explicit retry path available; a run with only a non-resumable checkpoint is not retried from scratch. +- An expired checkpointless `Running` lease is safe to re-drive because it has not crossed the first loop checkpoint. Recovery clears runner ownership and requeues it as `Queued`, while retaining the canonical-thread active lock and preserving `claim_count`. +- Checkpointless re-drive is bounded by `max_crash_recovery_reclaims`. Once `claim_count` reaches the bound, recovery produces terminal `Failed(crash_retry_exhausted)` and releases the active lock. +- A requeued run is claimable through the normal `claim_next_run` path. A duplicate/new submit for the same canonical thread remains `ThreadBusy` while that active run is queued or running. +- Recovery emits the lifecycle event for the resulting state. Because recovered runs leave `Running`/`CancelRequested`, later reconciliation scans do not repeat the same transition. +- `RecoveryRequired` remains readable as legacy status vocabulary, but current lease recovery does not produce it. --- @@ -49,11 +53,34 @@ Agent-loop drivers return `LoopExit` claims. `TurnRunner` validates those claims - valid completed exits require host-verified durable reply/result refs and map to `TurnRunnerOutcome::Completed`; - valid blocked exits require host-verified checkpoint + gate refs and map to `TurnRunnerOutcome::Blocked`; -- valid cancelled exits require observed host cancellation/interrupt and map to `TurnRunnerOutcome::Cancelled`; a missing final checkpoint is allowed for host-initiated cancellation because the host can preempt the driver before checkpointing; runner-side application then consults durable run state in one transition-port operation, terminalizing only recorded `CancelRequested` runs and mapping observed interrupts that race ahead of recorded cancellation to recovery instead of terminal cancellation; +- valid cancelled exits require observed host cancellation/interrupt and map to `TurnRunnerOutcome::Cancelled`; a missing final checkpoint is allowed for host-initiated cancellation because the host can preempt the driver before checkpointing; runner-side application then consults durable run state in one transition-port operation, terminalizing only recorded `CancelRequested` runs and mapping observed interrupts that race ahead of recorded cancellation to a sanitized terminal failure instead of terminal cancellation; - valid failed exits require host-verified evidence that the failure is safe to terminalize, then map stable sanitized failure kinds or sanitized safe summaries to `TurnRunnerOutcome::Failed`; failed outcomes may include host-verified explanation refs and a retry checkpoint id admitted by the checkpoint policy; -- invalid exits map either to sanitized terminal failure or runner/system-derived `RecoveryRequired` depending on side-effect safety evidence; +- invalid exits map to a sanitized terminal failure. The loop-exit policy still names its unsafe internal mapping `RecoveryRequired`, but the current persistence transition terminalizes that mapping as `Failed` rather than persisting a new `RecoveryRequired` state; - runner-side loop-exit application must call trusted transition-port methods, not mutate durable run state directly. ## 6. Deferred work The current slices define the core lease/recovery state machine, initial PostgreSQL/libSQL persistence adapters, pure `LoopExit` validation/mapping types, trusted `LoopExitApplier` policy derivation from host-owned evidence, host-runtime production scheduler wiring, and failed-run retry persistence plus runner resume execution through `RetryTurnRequest`/`RetryTurnResponse`. Durable exit-id replay storage, transcript draft validation, side-effect boundary checkpoint cadence inside the loop, and safe explicit fork UX remain follow-up slices. + +Broken-tool detection and automatic rebuild are not turn-runner responsibilities. Their disposition is recorded separately in [`../tier-b-self-repair-reconciliation.md`](../tier-b-self-repair-reconciliation.md). + +## 7. Verification + +The caller-level and persistence-contract evidence for this behavior is: + +- scheduler heartbeats, heartbeat failure handling, cancellation, and periodic lease reconciliation: `crates/ironclaw_runner/tests/turn_scheduler_contract.rs`; +- checkpointless requeue and bounded exhaustion across a durable-store crash: `crates/ironclaw_turns/tests/row_store_crash_consistency.rs`; +- checkpoint-preserving failed-run retry: `crates/ironclaw_turns/tests/retry_failed_turn_store_contract.rs`; +- terminal state, lifecycle event, and active-lock release: `crates/ironclaw_turns/tests/turn_coordinator_contract.rs`; +- a production-shaped scheduler/tool-dispatch wedge: `tests/integration/lease_wedge.rs`. + +Run the focused evidence with: + +```bash +cargo test -p ironclaw_runner --test turn_scheduler_contract +cargo test -p ironclaw_turns --test row_store_crash_consistency lease_expiry_ +cargo test -p ironclaw_turns --test retry_failed_turn_store_contract lease_recovery_preserves_retryability +cargo test -p ironclaw_turns --test turn_coordinator_contract expired_running_lease_fails_and_releases_thread_lock +cargo test -p ironclaw_turns --test turn_coordinator_contract expired_cancel_requested_lease_cancels_and_releases_thread_lock +cargo test --test reborn_integration_lease_wedge +``` diff --git a/docs/reborn/tier-b-self-repair-reconciliation.md b/docs/reborn/tier-b-self-repair-reconciliation.md new file mode 100644 index 00000000000..c456798c503 --- /dev/null +++ b/docs/reborn/tier-b-self-repair-reconciliation.md @@ -0,0 +1,75 @@ +# Tier B Self-Repair Reconciliation + +Issues: #6455, parent #6369 +Date: 2026-07-22 + +## Decision + +The v1 self-repair module combined two independent features: + +1. recovery of jobs that stopped making progress; and +2. detection and automatic rebuilding of repeatedly failing dynamic tools. + +Reborn already satisfies the stuck-run part through its durable runner lease, +heartbeat, checkpoint, retry, and lifecycle contracts. This issue therefore +does not add a second scheduler, background worker, recovery state machine, or +configuration default. The broken-tool rebuild part is not satisfied by lease +recovery and remains separate future work. + +The historical comparison point is +`5b307e2c920e2b6e1e6f219a8776103d1906257f:src/agent/self_repair.rs`, the final +v1 tree before the legacy monolith was removed. + +## Stuck-run mapping + +| v1 responsibility | Reborn owner and behavior | Evidence | Disposition | +| --- | --- | --- | --- | +| Detect an in-progress job that stopped making progress | `TurnRunScheduler` heartbeats active executors; the durable turn store identifies expired `Running` and `CancelRequested` leases from `lease_expires_at`; the scheduler periodically invokes `recover_expired_leases`. | `crates/ironclaw_runner/src/turn_scheduler.rs`; `scheduler_heartbeats_long_running_executor_until_completion`; `wedged_tool_call_is_reaped_by_lease_expiry_not_left_running_forever` | Satisfied | +| Recover work that is safe to run again | A checkpointless expired run is cleared of lease ownership and requeued as `Queued`; normal claim processing re-drives it. Work that crossed a loop checkpoint is not blindly restarted. | `crates/ironclaw_turns/src/filesystem_store/turn_state_engine/transitions.rs`; `lease_expiry_requeues_checkpointless_run_as_redrivable` | Satisfied | +| Bound repeated recovery attempts | Every claim increments durable `claim_count`; checkpointless recovery stops at `max_crash_recovery_reclaims` (default `5`) and terminal-fails with `crash_retry_exhausted`. | `crates/ironclaw_turns/src/filesystem_store/turn_state_engine/limits.rs`; `lease_expiry_crash_retry_bound_fails_with_crash_retry_exhausted`; `expired_lease_reconciler_fails_running_run_at_crash_retry_bound` | Satisfied | +| Preserve safe progress instead of restarting uncertain work | A run with any loop checkpoint becomes `Failed(lease_expired)`; the latest resumable checkpoint is attached when available and can be used by the explicit failed-run retry path. | `assert_lease_recovery_preserves_retryability` in `crates/ironclaw_turns/tests/retry_failed_turn_store_contract.rs` | Satisfied | +| Stop and report terminal recovery outcomes | Cancellation becomes `Cancelled`; exhausted pre-checkpoint recovery becomes `Failed(crash_retry_exhausted)`; checkpointed expiry becomes `Failed(lease_expired)`. Terminal transitions release the thread lock and publish sanitized lifecycle state. | `expired_running_lease_fails_and_releases_thread_lock`; `expired_cancel_requested_lease_cancels_and_releases_thread_lock`; `tests/integration/lease_wedge.rs` | Satisfied | +| Avoid repeated repair notifications for the same stuck item | Reconciliation only selects `Running`/`CancelRequested`; each result leaves those states, so later scans cannot repeat the same recovery transition. Requeued work can expire again only after a new claim, and that cycle is bounded. | `recover_expired_leases` transition code and the bounded reclaim tests above | Satisfied | + +No additional caller-level regression test is needed for #6455: the scheduler +contract covers heartbeat and reconciliation, the row-store contract covers +requeue and exhaustion, the retry contract covers checkpoints, and the root +integration test wedges the real tool-dispatch caller path. + +## Broken-tool auto-rebuild split + +Lease recovery answers whether a turn run is still owned by a live runner. It +does not establish that a tool artifact is defective, decide whether source is +trusted, compile replacement code, install an artifact, or activate it. + +The v1 behavior that counted repeated tool failures and invoked +`SoftwareBuilder`/`ToolRegistry` to rebuild non-built-in tools is therefore +**not** covered by Reborn lease recovery and is intentionally out of scope for +#6455. It should be tracked in a separate future issue, for example: + +> Reconcile v1 broken-tool detection and auto-rebuild with the Reborn extension lifecycle + +That follow-up must first define failure attribution, trusted source and build +isolation, approval policy, bounded attempts, artifact verification, activation +rollback, and operator-visible evidence. It must not be implemented as an +implicit side effect of turn lease recovery. + +## Compatibility and rollback + +- Runtime behavior and defaults are unchanged. +- Persistence schemas and serialized status vocabulary are unchanged; + `RecoveryRequired` remains readable for legacy compatibility. +- The corrected contract describes behavior already shipped and tested under + #6284. +- Rollback is a documentation revert only. + +## Validation + +```bash +cargo test -p ironclaw_runner --test turn_scheduler_contract +cargo test -p ironclaw_turns --test row_store_crash_consistency lease_expiry_ +cargo test -p ironclaw_turns --test retry_failed_turn_store_contract lease_recovery_preserves_retryability +cargo test -p ironclaw_turns --test turn_coordinator_contract expired_running_lease_fails_and_releases_thread_lock +cargo test -p ironclaw_turns --test turn_coordinator_contract expired_cancel_requested_lease_cancels_and_releases_thread_lock +cargo test --test reborn_integration_lease_wedge +``` From 405b5675fac02c8ee93f39786a634505b3b1586c Mon Sep 17 00:00:00 2001 From: italic-jinxin <106428113+italic-jinxin@users.noreply.github.com> Date: Wed, 22 Jul 2026 20:40:03 +0800 Subject: [PATCH 2/3] docs(reborn): address lease recovery review feedback --- .../tests/row_store_crash_consistency.rs | 11 ++++++----- docs/reborn/contracts/turn-persistence.md | 9 ++++++--- docs/reborn/contracts/turn-runner.md | 2 +- docs/reborn/contracts/turns-agent-loop.md | 9 ++++++--- docs/reborn/tier-b-self-repair-reconciliation.md | 2 +- 5 files changed, 20 insertions(+), 13 deletions(-) diff --git a/crates/ironclaw_turns/tests/row_store_crash_consistency.rs b/crates/ironclaw_turns/tests/row_store_crash_consistency.rs index 4252f9d819d..61a5d3df9c2 100644 --- a/crates/ironclaw_turns/tests/row_store_crash_consistency.rs +++ b/crates/ironclaw_turns/tests/row_store_crash_consistency.rs @@ -2015,11 +2015,12 @@ async fn byte_state_fork_recovers_to_model_snapshot() { /// it deterministically (identically to the never-crashed reference model), /// with any cause preserved. /// -/// The crash itself preserves the durable `Running` state. A later lease-recovery -/// pass resolves both the row store and its reference model identically. The -/// caller-level tests below separately prove the two current #6284 outcomes: -/// checkpointless work is requeued while its reclaim budget remains, and it -/// terminal-fails with `crash_retry_exhausted` once that bound is reached. +/// The crash itself preserves the durable `Running` state. This test reopens the +/// row store, reads that state, and compares its durable prefix with the +/// never-crashed reference model. The caller-level tests below separately prove +/// the two current #6284 lease-recovery outcomes: checkpointless work is requeued +/// while its reclaim budget remains, and it terminal-fails with +/// `crash_retry_exhausted` once that bound is reached. #[tokio::test] async fn crash_mid_run_recovers_identically_to_model_and_preserves_cause() { let backend = Arc::new(FaultBackend::new(InMemoryBackend::new())); diff --git a/docs/reborn/contracts/turn-persistence.md b/docs/reborn/contracts/turn-persistence.md index 716e2753138..f0c0efba2c1 100644 --- a/docs/reborn/contracts/turn-persistence.md +++ b/docs/reborn/contracts/turn-persistence.md @@ -99,7 +99,7 @@ A duplicate idempotency key must replay prior accepted submit and admission-reje - Submit admission policy checks that can reject unauthorized/profile-invalid requests run before returning same-thread busy metadata; same-thread busy is still checked before capacity reservation and never consumes admission slots. - Capacity denial returns one deterministic safe `AdmissionRejected` payload with axis kind, total/class bucket, admission class when applicable, limit, active count, and optional retry hint. It must not expose foreign bucket IDs or raw provider internals. - Missing limits mean unlimited. A non-AllowAll provider that is unavailable fails closed with `AdmissionRejectionReason::Unavailable` and creates no run/reservation. -- Queued, running, blocked, cancel-requested, and recovery-required runs keep reservations. Resume reuses the existing reservation. +- Queued, running, blocked, and cancel-requested runs keep reservations. Resume reuses the existing reservation. - Terminal transitions (`Completed`, `Failed`, `Cancelled`, and future terminal states) release reservations exactly once. Released reservation evidence is retained only while the corresponding terminal run remains within the bounded terminal-record retention window; active capacity accounting must not scan unbounded released history. - Limit changes do not evict existing runs; new admissions are denied until active reservations drop below the configured limit. - Snapshot/DB loaders must synthesize unreleased reservation evidence for legacy non-terminal runs that predate persisted reservation rows so active capacity is not bypassed after migration/restart. @@ -109,9 +109,12 @@ A duplicate idempotency key must replay prior accepted submit and admission-reje ## 6. Runner lease and checkpoint rules - Claiming a queued run atomically moves it to `Running`, stores runner ID/lease token, increments `claim_count`, records `last_heartbeat_at`, records `lease_expires_at`, and updates active-lock metadata. -- Heartbeats only renew metadata for matching, unexpired runner ID/lease token; successful heartbeats refresh `last_heartbeat_at` and extend `lease_expires_at`. +- Heartbeats only renew metadata for matching, unexpired runner ID/lease token on actively `Running` work; heartbeat requests are rejected once the run is `CancelRequested`. Successful heartbeats refresh `last_heartbeat_at` and extend `lease_expires_at`. - Physical adapters may split high-churn runner lease metadata from lower-churn turn snapshots/tables, as long as all read, recovery, and terminal transition APIs expose one logical run state. Liveness decisions must use durable lease metadata, not require one lifecycle event per heartbeat. -- Expired `Running` and `CancelRequested` leases transition to `RecoveryRequired`, clear current runner ownership, emit a redacted recovery event, and keep the active lock so uncertain side-effecting work is not auto-retried. +- An expired `CancelRequested` lease becomes terminal `Cancelled`, clears runner ownership, and releases its active lock and admission reservation. +- An expired `Running` lease with any loop checkpoint becomes terminal `Failed(lease_expired)`, with its latest resumable checkpoint attached when one exists; it clears runner ownership and releases its active lock and admission reservation. +- An expired checkpointless `Running` lease below `max_crash_recovery_reclaims` clears runner ownership and returns to `Queued`, retaining its active lock, admission reservation, and `claim_count`. At the reclaim bound it instead becomes terminal `Failed(crash_retry_exhausted)` and releases the lock and reservation. +- `RecoveryRequired` remains readable as legacy terminal status vocabulary, but current lease recovery does not produce it. - Blocking a running run requires a matching, unexpired lease, writes a checkpoint record, stores the latest checkpoint/gate refs on the run, clears current lease ownership, and keeps the active lock. - Loop-driver resume payloads are staged in a host-owned `CheckpointStateStore` before a public checkpoint record is written. The store returns an opaque `LoopCheckpointStateRef`; callers cannot choose arbitrary refs for durable records. - Checkpoint-state records are scoped by `TurnScope`, `TurnId`, and `TurnRunId`. Reads with a matching ref but foreign scope or run return no state, preserving tenant/thread/run isolation. diff --git a/docs/reborn/contracts/turn-runner.md b/docs/reborn/contracts/turn-runner.md index 8fd2de4b16d..7081f6a9187 100644 --- a/docs/reborn/contracts/turn-runner.md +++ b/docs/reborn/contracts/turn-runner.md @@ -20,7 +20,7 @@ Product adapters must continue to use `TurnCoordinator`. Runner transition APIs - `submit_turn` creates a queued `TurnRunId` and active-thread lock, but no model/tool side effects may run before a runner claim succeeds. - `claim_next_run` atomically moves one matching `Queued` run to `Running`. - A successful claim stores `runner_id`, `lease_token`, `last_heartbeat_at`, `lease_expires_at`, increments `claim_count`, updates the active lock, and emits `RunnerClaimed`. -- `heartbeat` requires the matching `runner_id` and `lease_token`, only refreshes actively `Running` work, and rejects leases whose `lease_expires_at` has already passed. Once cancellation is requested, heartbeats no longer extend the lease; the runner must complete cancellation before the existing lease expires or the reconciler terminalizes the run as `Cancelled`. On success, heartbeat refreshes durable `last_heartbeat_at` and extends durable `lease_expires_at`; adapters may touch active-lock freshness and emit/coalesce `RunnerHeartbeat` lifecycle events, but consumers must use lease metadata as the liveness source of truth. +- `heartbeat` requires the matching `runner_id` and `lease_token`, only refreshes actively `Running` work, and rejects leases whose `lease_expires_at` has already passed. Once cancellation is requested and the run is durably `CancelRequested`, heartbeat calls are rejected with an error; the runner cannot extend the lease and must complete cancellation before the existing lease expires or the reconciler terminalizes the run as `Cancelled`. On success, heartbeat refreshes durable `last_heartbeat_at` and extends durable `lease_expires_at`; adapters may touch active-lock freshness and emit/coalesce `RunnerHeartbeat` lifecycle events, but consumers must use lease metadata as the liveness source of truth. - Pull-based claims are authoritative. Wake notifications are optimization hints only. - After `TurnCoordinator` durably accepts a submitted run or requeues a resumed/retried run, it may emit a redacted queued-run wake hint containing only the canonical scope, `TurnRunId`, queued status, and event cursor. Wake delivery is best-effort, is not a source of truth, must not fail the durable adapter call, and duplicate hints must be harmless. diff --git a/docs/reborn/contracts/turns-agent-loop.md b/docs/reborn/contracts/turns-agent-loop.md index 69bd7b61ffd..8b48b03260f 100644 --- a/docs/reborn/contracts/turns-agent-loop.md +++ b/docs/reborn/contracts/turns-agent-loop.md @@ -107,10 +107,11 @@ blocked_approval blocked_auth waiting_tool waiting_process -recovery_required +cancel_requested completed failed cancelled +recovery_required (legacy compatibility only; no new transition produces it) ``` Transitions: @@ -120,8 +121,10 @@ accepted -> queued -> running running -> blocked_approval -> running running -> waiting_tool -> running running -> waiting_process -> running -running -> recovery_required -recovery_required -> cancelled +running -> cancel_requested -> cancelled +running(no checkpoint, expired lease, reclaim budget remains) -> queued +running(any checkpoint, expired lease) -> failed +running(no checkpoint, expired lease, reclaim budget exhausted) -> failed running -> completed|failed|cancelled ``` diff --git a/docs/reborn/tier-b-self-repair-reconciliation.md b/docs/reborn/tier-b-self-repair-reconciliation.md index c456798c503..b68bbdb6191 100644 --- a/docs/reborn/tier-b-self-repair-reconciliation.md +++ b/docs/reborn/tier-b-self-repair-reconciliation.md @@ -28,7 +28,7 @@ v1 tree before the legacy monolith was removed. | Recover work that is safe to run again | A checkpointless expired run is cleared of lease ownership and requeued as `Queued`; normal claim processing re-drives it. Work that crossed a loop checkpoint is not blindly restarted. | `crates/ironclaw_turns/src/filesystem_store/turn_state_engine/transitions.rs`; `lease_expiry_requeues_checkpointless_run_as_redrivable` | Satisfied | | Bound repeated recovery attempts | Every claim increments durable `claim_count`; checkpointless recovery stops at `max_crash_recovery_reclaims` (default `5`) and terminal-fails with `crash_retry_exhausted`. | `crates/ironclaw_turns/src/filesystem_store/turn_state_engine/limits.rs`; `lease_expiry_crash_retry_bound_fails_with_crash_retry_exhausted`; `expired_lease_reconciler_fails_running_run_at_crash_retry_bound` | Satisfied | | Preserve safe progress instead of restarting uncertain work | A run with any loop checkpoint becomes `Failed(lease_expired)`; the latest resumable checkpoint is attached when available and can be used by the explicit failed-run retry path. | `assert_lease_recovery_preserves_retryability` in `crates/ironclaw_turns/tests/retry_failed_turn_store_contract.rs` | Satisfied | -| Stop and report terminal recovery outcomes | Cancellation becomes `Cancelled`; exhausted pre-checkpoint recovery becomes `Failed(crash_retry_exhausted)`; checkpointed expiry becomes `Failed(lease_expired)`. Terminal transitions release the thread lock and publish sanitized lifecycle state. | `expired_running_lease_fails_and_releases_thread_lock`; `expired_cancel_requested_lease_cancels_and_releases_thread_lock`; `tests/integration/lease_wedge.rs` | Satisfied | +| Stop and report terminal recovery outcomes | An expired recorded `CancelRequested` lease becomes `Cancelled`; exhausted pre-checkpoint recovery becomes `Failed(crash_retry_exhausted)`; checkpointed expiry becomes `Failed(lease_expired)`. An observed interrupt that races ahead of persisted cancellation becomes sanitized terminal `Failed`, not `Cancelled`. Terminal transitions release the thread lock and publish sanitized lifecycle state. | `expired_running_lease_fails_and_releases_thread_lock`; `expired_cancel_requested_lease_cancels_and_releases_thread_lock`; `tests/integration/lease_wedge.rs` | Satisfied | | Avoid repeated repair notifications for the same stuck item | Reconciliation only selects `Running`/`CancelRequested`; each result leaves those states, so later scans cannot repeat the same recovery transition. Requeued work can expire again only after a new claim, and that cycle is bounded. | `recover_expired_leases` transition code and the bounded reclaim tests above | Satisfied | No additional caller-level regression test is needed for #6455: the scheduler From acfb2ec9ac401ac9e4f12de2f73d2917752a97fc Mon Sep 17 00:00:00 2001 From: italic-jinxin <106428113+italic-jinxin@users.noreply.github.com> Date: Wed, 22 Jul 2026 20:52:17 +0800 Subject: [PATCH 3/3] docs(reborn): clarify legacy recovery lock behavior --- crates/ironclaw_turns/tests/row_store_crash_consistency.rs | 6 +++--- docs/reborn/contracts/turn-persistence.md | 7 ++++--- 2 files changed, 7 insertions(+), 6 deletions(-) diff --git a/crates/ironclaw_turns/tests/row_store_crash_consistency.rs b/crates/ironclaw_turns/tests/row_store_crash_consistency.rs index 61a5d3df9c2..411e2a9647b 100644 --- a/crates/ironclaw_turns/tests/row_store_crash_consistency.rs +++ b/crates/ironclaw_turns/tests/row_store_crash_consistency.rs @@ -2017,9 +2017,9 @@ async fn byte_state_fork_recovers_to_model_snapshot() { /// /// The crash itself preserves the durable `Running` state. This test reopens the /// row store, reads that state, and compares its durable prefix with the -/// never-crashed reference model. The caller-level tests below separately prove -/// the two current #6284 lease-recovery outcomes: checkpointless work is requeued -/// while its reclaim budget remains, and it terminal-fails with +/// never-crashed reference model. The persistence-level tests below separately +/// prove the two current #6284 lease-recovery outcomes: checkpointless work is +/// requeued while its reclaim budget remains, and it terminal-fails with /// `crash_retry_exhausted` once that bound is reached. #[tokio::test] async fn crash_mid_run_recovers_identically_to_model_and_preserves_cause() { diff --git a/docs/reborn/contracts/turn-persistence.md b/docs/reborn/contracts/turn-persistence.md index f0c0efba2c1..7a14b819446 100644 --- a/docs/reborn/contracts/turn-persistence.md +++ b/docs/reborn/contracts/turn-persistence.md @@ -72,8 +72,9 @@ verify the row projection before enabling row-store-only production traffic. - Active-lock key is the canonical `TurnScope`: tenant, agent, optional project, and thread. - The key excludes `TurnActor.user_id`, channel IDs, source binding refs, and reply binding refs. - A lock stores the current owning `TurnRunId`, explicit `TurnStatus`, monotonically increasing `TurnLockVersion`, `acquired_at`, and `updated_at`. -- Queued, running, cancel-requested, blocked, and recovery-required runs keep the lock. -- Terminal runs release the lock exactly once. +- Queued, running, cancel-requested, and blocked runs keep the lock. +- Current terminal transitions release their owned lock exactly once through `Inner::release_active_lock` in `crates/ironclaw_turns/src/filesystem_store/turn_state_engine/transitions.rs`. +- A persisted legacy `RecoveryRequired` run is terminal and does not keep effective active-lock ownership. `Inner::from_persistence_snapshot` in `crates/ironclaw_turns/src/filesystem_store/turn_state_engine/snapshot.rs` rehydrates the status; `TurnStatus::keeps_active_lock` and `Inner::thread_busy` then make any stale matching lock row non-blocking. New submissions may replace that stale row. - Runner claim/resume/block/cancel-request transitions update the lock status/version while keeping ownership with the same run. --- @@ -114,7 +115,7 @@ A duplicate idempotency key must replay prior accepted submit and admission-reje - An expired `CancelRequested` lease becomes terminal `Cancelled`, clears runner ownership, and releases its active lock and admission reservation. - An expired `Running` lease with any loop checkpoint becomes terminal `Failed(lease_expired)`, with its latest resumable checkpoint attached when one exists; it clears runner ownership and releases its active lock and admission reservation. - An expired checkpointless `Running` lease below `max_crash_recovery_reclaims` clears runner ownership and returns to `Queued`, retaining its active lock, admission reservation, and `claim_count`. At the reclaim bound it instead becomes terminal `Failed(crash_retry_exhausted)` and releases the lock and reservation. -- `RecoveryRequired` remains readable as legacy terminal status vocabulary, but current lease recovery does not produce it. +- `RecoveryRequired` remains readable as a legacy terminal status, follows the non-blocking legacy-row rule in §3, and is not produced by current lease recovery. - Blocking a running run requires a matching, unexpired lease, writes a checkpoint record, stores the latest checkpoint/gate refs on the run, clears current lease ownership, and keeps the active lock. - Loop-driver resume payloads are staged in a host-owned `CheckpointStateStore` before a public checkpoint record is written. The store returns an opaque `LoopCheckpointStateRef`; callers cannot choose arbitrary refs for durable records. - Checkpoint-state records are scoped by `TurnScope`, `TurnId`, and `TurnRunId`. Reads with a matching ref but foreign scope or run return no state, preserving tenant/thread/run isolation.