Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 6 additions & 12 deletions crates/ironclaw_turns/tests/row_store_crash_consistency.rs
Original file line number Diff line number Diff line change
Expand Up @@ -2015,18 +2015,12 @@ async fn byte_state_fork_recovers_to_model_snapshot() {
/// it deterministically (identically to the never-crashed reference model),
/// with any cause preserved.
///
/// NOTE on the literal coordinator ask ("recover_expired_leases returns the run
/// to a *claimable* state; never Failed"): the shared turn engine does NOT
/// re-queue an expired lease — the store`s `recover_expired_leases`
/// terminates the abandoned run as `Failed(lease_expired)` (a
/// resumable-checkpointed run keeps its checkpoint and is retryable; a
/// checkpoint-less one does not — see the ignored reproducer below). That is
/// pre-existing engine semantics, identical in the direct authority and the row
/// store, and is unaffected by whether a crash occurred. This suite's charge is
/// row-store *crash-consistency*, so the assertion here is the defensible one:
/// the crash neither loses the run nor diverges its recovered outcome from the
/// model. The lifecycle question (should lease expiry be terminal at all under
/// #6284?) is captured, reproducibly, in the ignored test below.
/// The crash itself preserves the durable `Running` state. This test reopens the
/// row store, reads that state, and compares its durable prefix with the
/// never-crashed reference model. The persistence-level tests below separately
/// prove the two current #6284 lease-recovery outcomes: checkpointless work is
/// requeued while its reclaim budget remains, and it terminal-fails with
/// `crash_retry_exhausted` once that bound is reached.
Comment thread
coderabbitai[bot] marked this conversation as resolved.
#[tokio::test]
async fn crash_mid_run_recovers_identically_to_model_and_preserves_cause() {
let backend = Arc::new(FaultBackend::new(InMemoryBackend::new()));
Expand Down
14 changes: 9 additions & 5 deletions docs/reborn/contracts/turn-persistence.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,8 +72,9 @@ verify the row projection before enabling row-store-only production traffic.
- Active-lock key is the canonical `TurnScope`: tenant, agent, optional project, and thread.
- The key excludes `TurnActor.user_id`, channel IDs, source binding refs, and reply binding refs.
- A lock stores the current owning `TurnRunId`, explicit `TurnStatus`, monotonically increasing `TurnLockVersion`, `acquired_at`, and `updated_at`.
- Queued, running, cancel-requested, blocked, and recovery-required runs keep the lock.
- Terminal runs release the lock exactly once.
- Queued, running, cancel-requested, and blocked runs keep the lock.
- Current terminal transitions release their owned lock exactly once through `Inner::release_active_lock` in `crates/ironclaw_turns/src/filesystem_store/turn_state_engine/transitions.rs`.
- A persisted legacy `RecoveryRequired` run is terminal and does not keep effective active-lock ownership. `Inner::from_persistence_snapshot` in `crates/ironclaw_turns/src/filesystem_store/turn_state_engine/snapshot.rs` rehydrates the status; `TurnStatus::keeps_active_lock` and `Inner::thread_busy` then make any stale matching lock row non-blocking. New submissions may replace that stale row.
- Runner claim/resume/block/cancel-request transitions update the lock status/version while keeping ownership with the same run.

---
Expand All @@ -99,7 +100,7 @@ A duplicate idempotency key must replay prior accepted submit and admission-reje
- Submit admission policy checks that can reject unauthorized/profile-invalid requests run before returning same-thread busy metadata; same-thread busy is still checked before capacity reservation and never consumes admission slots.
- Capacity denial returns one deterministic safe `AdmissionRejected` payload with axis kind, total/class bucket, admission class when applicable, limit, active count, and optional retry hint. It must not expose foreign bucket IDs or raw provider internals.
- Missing limits mean unlimited. A non-AllowAll provider that is unavailable fails closed with `AdmissionRejectionReason::Unavailable` and creates no run/reservation.
- Queued, running, blocked, cancel-requested, and recovery-required runs keep reservations. Resume reuses the existing reservation.
- Queued, running, blocked, and cancel-requested runs keep reservations. Resume reuses the existing reservation.
- Terminal transitions (`Completed`, `Failed`, `Cancelled`, and future terminal states) release reservations exactly once. Released reservation evidence is retained only while the corresponding terminal run remains within the bounded terminal-record retention window; active capacity accounting must not scan unbounded released history.
- Limit changes do not evict existing runs; new admissions are denied until active reservations drop below the configured limit.
- Snapshot/DB loaders must synthesize unreleased reservation evidence for legacy non-terminal runs that predate persisted reservation rows so active capacity is not bypassed after migration/restart.
Expand All @@ -109,9 +110,12 @@ A duplicate idempotency key must replay prior accepted submit and admission-reje
## 6. Runner lease and checkpoint rules

- Claiming a queued run atomically moves it to `Running`, stores runner ID/lease token, increments `claim_count`, records `last_heartbeat_at`, records `lease_expires_at`, and updates active-lock metadata.
- Heartbeats only renew metadata for matching, unexpired runner ID/lease token; successful heartbeats refresh `last_heartbeat_at` and extend `lease_expires_at`.
- Heartbeats only renew metadata for matching, unexpired runner ID/lease token on actively `Running` work; heartbeat requests are rejected once the run is `CancelRequested`. Successful heartbeats refresh `last_heartbeat_at` and extend `lease_expires_at`.
- Physical adapters may split high-churn runner lease metadata from lower-churn turn snapshots/tables, as long as all read, recovery, and terminal transition APIs expose one logical run state. Liveness decisions must use durable lease metadata, not require one lifecycle event per heartbeat.
- Expired `Running` and `CancelRequested` leases transition to `RecoveryRequired`, clear current runner ownership, emit a redacted recovery event, and keep the active lock so uncertain side-effecting work is not auto-retried.
- An expired `CancelRequested` lease becomes terminal `Cancelled`, clears runner ownership, and releases its active lock and admission reservation.
- An expired `Running` lease with any loop checkpoint becomes terminal `Failed(lease_expired)`, with its latest resumable checkpoint attached when one exists; it clears runner ownership and releases its active lock and admission reservation.
- An expired checkpointless `Running` lease below `max_crash_recovery_reclaims` clears runner ownership and returns to `Queued`, retaining its active lock, admission reservation, and `claim_count`. At the reclaim bound it instead becomes terminal `Failed(crash_retry_exhausted)` and releases the lock and reservation.
- `RecoveryRequired` remains readable as a legacy terminal status, follows the non-blocking legacy-row rule in §3, and is not produced by current lease recovery.
- Blocking a running run requires a matching, unexpired lease, writes a checkpoint record, stores the latest checkpoint/gate refs on the run, clears current lease ownership, and keeps the active lock.
- Loop-driver resume payloads are staged in a host-owned `CheckpointStateStore` before a public checkpoint record is written. The store returns an opaque `LoopCheckpointStateRef`; callers cannot choose arbitrary refs for durable records.
- Checkpoint-state records are scoped by `TurnScope`, `TurnId`, and `TurnRunId`. Reads with a matching ref but foreign scope or run return no state, preserving tenant/thread/run isolation.
Expand Down
43 changes: 35 additions & 8 deletions docs/reborn/contracts/turn-runner.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,13 +2,14 @@

**Status:** Contract-freeze draft
**Date:** 2026-05-06
**Last reconciled:** 2026-07-22 (#6455)
**Depends on:** [`turn-persistence.md`](turn-persistence.md), [`turns-agent-loop.md`](turns-agent-loop.md), [`loop-exit.md`](loop-exit.md), [`runtime-profiles.md`](runtime-profiles.md)

---

## 1. Purpose

`TurnRunner` is the trusted worker-side control plane for executable turn runs. It claims queued runs, maintains leases while model/tool work is active, records safe checkpoint/block/terminal transitions, and moves abandoned work to explicit recovery instead of blindly retrying side effects.
`TurnRunner` is the trusted worker-side control plane for executable turn runs. It claims queued runs, maintains leases while model/tool work is active, records safe checkpoint/block/terminal transitions, and moves abandoned work through checkpoint-aware, bounded recovery instead of blindly retrying uncertain side effects.

Product adapters must continue to use `TurnCoordinator`. Runner transition APIs are trusted-worker APIs and remain under `ironclaw_turns::runner`. Driver-facing loop exits remain distinct from trusted runner outcomes; see [`loop-exit.md`](loop-exit.md).

Expand All @@ -19,7 +20,7 @@ Product adapters must continue to use `TurnCoordinator`. Runner transition APIs
- `submit_turn` creates a queued `TurnRunId` and active-thread lock, but no model/tool side effects may run before a runner claim succeeds.
- `claim_next_run` atomically moves one matching `Queued` run to `Running`.
- A successful claim stores `runner_id`, `lease_token`, `last_heartbeat_at`, `lease_expires_at`, increments `claim_count`, updates the active lock, and emits `RunnerClaimed`.
- `heartbeat` requires the matching `runner_id` and `lease_token`, only refreshes actively `Running` work, and rejects leases whose `lease_expires_at` has already passed. Once cancellation is requested, heartbeats no longer extend the lease; the runner must complete cancellation before the existing lease expires or the reconciler moves the run to recovery. On success, heartbeat refreshes durable `last_heartbeat_at` and extends durable `lease_expires_at`; adapters may touch active-lock freshness and emit/coalesce `RunnerHeartbeat` lifecycle events, but consumers must use lease metadata as the liveness source of truth.
- `heartbeat` requires the matching `runner_id` and `lease_token`, only refreshes actively `Running` work, and rejects leases whose `lease_expires_at` has already passed. Once cancellation is requested and the run is durably `CancelRequested`, heartbeat calls are rejected with an error; the runner cannot extend the lease and must complete cancellation before the existing lease expires or the reconciler terminalizes the run as `Cancelled`. On success, heartbeat refreshes durable `last_heartbeat_at` and extends durable `lease_expires_at`; adapters may touch active-lock freshness and emit/coalesce `RunnerHeartbeat` lifecycle events, but consumers must use lease metadata as the liveness source of truth.
- Pull-based claims are authoritative. Wake notifications are optimization hints only.
- After `TurnCoordinator` durably accepts a submitted run or requeues a resumed/retried run, it may emit a redacted queued-run wake hint containing only the canonical scope, `TurnRunId`, queued status, and event cursor. Wake delivery is best-effort, is not a source of truth, must not fail the durable adapter call, and duplicate hints must be harmless.

Expand All @@ -28,10 +29,13 @@ Product adapters must continue to use `TurnCoordinator`. Runner transition APIs
## 3. Expired lease recovery

- A reconciler scans runner-owned `Running` and `CancelRequested` leases using durable `lease_expires_at` metadata.
Comment thread
italic-jinxin marked this conversation as resolved.
- Expired `Running` or `CancelRequested` leases transition to `RecoveryRequired`, clear current runner ownership, emit a redacted `RecoveryRequired` event with reason `lease_expired`, and keep the same canonical-thread active lock.
- `RecoveryRequired` runs are not returned by the normal `claim_next_run` path. The system must not auto-retry uncertain side-effecting work.
- A duplicate/new submit for the same canonical thread remains `ThreadBusy` while recovery is required.
- Explicit cancellation of `RecoveryRequired` is terminal `Cancelled` and releases the active lock so a new turn can be submitted.
- An expired `CancelRequested` lease becomes terminal `Cancelled`, clears runner ownership, and releases the canonical-thread active lock.
- An expired `Running` lease with any loop checkpoint becomes terminal `Failed(lease_expired)`. The latest resumable checkpoint is attached when one exists, making the normal explicit retry path available; a run with only a non-resumable checkpoint is not retried from scratch.
- An expired checkpointless `Running` lease is safe to re-drive because it has not crossed the first loop checkpoint. Recovery clears runner ownership and requeues it as `Queued`, while retaining the canonical-thread active lock and preserving `claim_count`.
- Checkpointless re-drive is bounded by `max_crash_recovery_reclaims`. Once `claim_count` reaches the bound, recovery produces terminal `Failed(crash_retry_exhausted)` and releases the active lock.
- A requeued run is claimable through the normal `claim_next_run` path. A duplicate/new submit for the same canonical thread remains `ThreadBusy` while that active run is queued or running.
- Recovery emits the lifecycle event for the resulting state. Because recovered runs leave `Running`/`CancelRequested`, later reconciliation scans do not repeat the same transition.
- `RecoveryRequired` remains readable as legacy status vocabulary, but current lease recovery does not produce it.

---

Expand All @@ -49,11 +53,34 @@ Agent-loop drivers return `LoopExit` claims. `TurnRunner` validates those claims

- valid completed exits require host-verified durable reply/result refs and map to `TurnRunnerOutcome::Completed`;
- valid blocked exits require host-verified checkpoint + gate refs and map to `TurnRunnerOutcome::Blocked`;
- valid cancelled exits require observed host cancellation/interrupt and map to `TurnRunnerOutcome::Cancelled`; a missing final checkpoint is allowed for host-initiated cancellation because the host can preempt the driver before checkpointing; runner-side application then consults durable run state in one transition-port operation, terminalizing only recorded `CancelRequested` runs and mapping observed interrupts that race ahead of recorded cancellation to recovery instead of terminal cancellation;
- valid cancelled exits require observed host cancellation/interrupt and map to `TurnRunnerOutcome::Cancelled`; a missing final checkpoint is allowed for host-initiated cancellation because the host can preempt the driver before checkpointing; runner-side application then consults durable run state in one transition-port operation, terminalizing only recorded `CancelRequested` runs and mapping observed interrupts that race ahead of recorded cancellation to a sanitized terminal failure instead of terminal cancellation;
- valid failed exits require host-verified evidence that the failure is safe to terminalize, then map stable sanitized failure kinds or sanitized safe summaries to `TurnRunnerOutcome::Failed`; failed outcomes may include host-verified explanation refs and a retry checkpoint id admitted by the checkpoint policy;
- invalid exits map either to sanitized terminal failure or runner/system-derived `RecoveryRequired` depending on side-effect safety evidence;
- invalid exits map to a sanitized terminal failure. The loop-exit policy still names its unsafe internal mapping `RecoveryRequired`, but the current persistence transition terminalizes that mapping as `Failed` rather than persisting a new `RecoveryRequired` state;
- runner-side loop-exit application must call trusted transition-port methods, not mutate durable run state directly.

## 6. Deferred work

The current slices define the core lease/recovery state machine, initial PostgreSQL/libSQL persistence adapters, pure `LoopExit` validation/mapping types, trusted `LoopExitApplier` policy derivation from host-owned evidence, host-runtime production scheduler wiring, and failed-run retry persistence plus runner resume execution through `RetryTurnRequest`/`RetryTurnResponse`. Durable exit-id replay storage, transcript draft validation, side-effect boundary checkpoint cadence inside the loop, and safe explicit fork UX remain follow-up slices.

Broken-tool detection and automatic rebuild are not turn-runner responsibilities. Their disposition is recorded separately in [`../tier-b-self-repair-reconciliation.md`](../tier-b-self-repair-reconciliation.md).

## 7. Verification

The caller-level and persistence-contract evidence for this behavior is:

- scheduler heartbeats, heartbeat failure handling, cancellation, and periodic lease reconciliation: `crates/ironclaw_runner/tests/turn_scheduler_contract.rs`;
- checkpointless requeue and bounded exhaustion across a durable-store crash: `crates/ironclaw_turns/tests/row_store_crash_consistency.rs`;
- checkpoint-preserving failed-run retry: `crates/ironclaw_turns/tests/retry_failed_turn_store_contract.rs`;
- terminal state, lifecycle event, and active-lock release: `crates/ironclaw_turns/tests/turn_coordinator_contract.rs`;
- a production-shaped scheduler/tool-dispatch wedge: `tests/integration/lease_wedge.rs`.

Run the focused evidence with:

```bash
cargo test -p ironclaw_runner --test turn_scheduler_contract
cargo test -p ironclaw_turns --test row_store_crash_consistency lease_expiry_
cargo test -p ironclaw_turns --test retry_failed_turn_store_contract lease_recovery_preserves_retryability
cargo test -p ironclaw_turns --test turn_coordinator_contract expired_running_lease_fails_and_releases_thread_lock
cargo test -p ironclaw_turns --test turn_coordinator_contract expired_cancel_requested_lease_cancels_and_releases_thread_lock
cargo test --test reborn_integration_lease_wedge
```
9 changes: 6 additions & 3 deletions docs/reborn/contracts/turns-agent-loop.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,10 +107,11 @@ blocked_approval
blocked_auth
waiting_tool
waiting_process
recovery_required
cancel_requested
completed
failed
cancelled
recovery_required (legacy compatibility only; no new transition produces it)
```

Transitions:
Expand All @@ -120,8 +121,10 @@ accepted -> queued -> running
running -> blocked_approval -> running
running -> waiting_tool -> running
running -> waiting_process -> running
running -> recovery_required
recovery_required -> cancelled
running -> cancel_requested -> cancelled
running(no checkpoint, expired lease, reclaim budget remains) -> queued
running(any checkpoint, expired lease) -> failed
running(no checkpoint, expired lease, reclaim budget exhausted) -> failed
running -> completed|failed|cancelled
```

Expand Down
Loading
Loading