Skip to content

docs(reborn): design — write-behind lease durability + side-effect gate - #5249

Closed
serrrfirat wants to merge 2 commits into
mainfrom
firat/write-behind-lease-design
Closed

serrrfirat wants to merge 2 commits into
mainfrom
firat/write-behind-lease-design

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

What

A design doc (no code) for eliminating the deployed-Postgres lease-expiry / scheduler_heartbeat_failed churn that #5232 reduced but did not remove. Adds docs/plans/2026-06-25-write-behind-lease-durability.md.

Why

#5232 moved runner heartbeats to a per-run durable sidecar, removing per-user contention — but the heartbeat is still a durable per-run CAS write through Postgres, so it's bound by per-write latency / pool pressure / cross-region RTT, and the scheduler still insta-fails a run on a single failed heartbeat. On the deployed Supabase instance this shows up as "constantly getting lease expired"; local libSQL (in-memory) never reproduces it.

The design

Split durable writes into two planes:

  • Liveness (heartbeats) → per-replica write-behind cache + coalescing 5s drain; safe to lose on crash (loss only makes recovery more eager — the safe direction). The hot path becomes a non-blocking in-memory renew(), so the single-heartbeat insta-fail disappears.
  • Correctness (claim / complete / fail / block / idempotency keys / tool-side-effect commits) → always synchronously durable.

The load-bearing safety mechanism is a side-effect gate: write-behind alone is not safe under a Postgres partition (a stale owner keeps executing tools while flushes fail, then another replica reclaims and re-runs → double side-effects). So, immediately before each side-effecting call, synchronously check the durable lease runway and stop the run rather than dispatch without it — plus a run-level self-stop at 45s of durable-lag. Both are pure-local checks needing no Postgres round-trip, so they hold during a partition.

Includes quantified recovery bounds (no false reclaim while clock-skew < 15s; dead-run strand ≤ TTL+poll+skew), a fault/latency-injection test plan driven through the scheduler, and a dark-mode, feature-flagged, reversible rollout safe on both Postgres and libSQL.

Provenance

Synthesized via a two-model design council (Opus 4.8 + GPT 5.5 XHigh): independent drafts → adversarial cross-review (which caught the partition/double-side-effect gap) → fusion → signoff. Final state: both slots accepted, no blockers. Agreement ledger is in §10 of the doc.

Status / next steps

Design only — accepted, not yet implemented. Two gates before any enable (both reviewers insisted): (1) decide explicitly whether the sidecar persists a fence or commits to runner_id-only CAS; (2) measure real Railway↔Supabase clock skew and confirm < 15s. Implementation would be a follow-up PR (starter: lease write-behind + side-effect gate; then the system-wide DurabilityClass layer).

🤖 Generated with Claude Code

Council-synthesized design (Opus 4.8 + GPT 5.5 XHigh, fusion-design-council)
for removing the deployed-Postgres lease-expiry / "scheduler_heartbeat_failed"
churn that #5232 reduced but did not eliminate.

Core: split durable writes into a liveness plane (heartbeats → per-replica
write-behind cache + coalescing drain, safe to lose on crash) and a correctness
plane (claim/complete/fail/block/idempotency/tool-side-effect commits → always
synchronously durable). The load-bearing safety mechanism is a side-effect gate:
before each side-effecting call, synchronously check the DURABLE lease runway and
stop the run rather than dispatch without it, plus a run-level self-stop at 45s of
durable-lag — both pure-local checks that hold during a Postgres partition, so a
stale owner cannot keep producing side effects a reclaiming replica then repeats.

Includes quantified recovery bounds (no false reclaim while skew<15s; dead-run
strand <= TTL+poll+skew), a fault/latency-injection test plan driven through the
scheduler, and a dark-mode, feature-flagged, reversible rollout safe on both
Postgres and libSQL. Status: accepted by both council slots, no blockers.

Design doc only — no code changes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5249 June 25, 2026 15:12 Destroyed
@github-actions github-actions Bot added scope: docs Documentation size: XS < 10 changed lines (excluding docs) risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jun 25, 2026
@coderabbitai

coderabbitai Bot commented Jun 25, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 88fa29b5-3e00-4ca6-b0c6-c5a73075f154

📥 Commits

Reviewing files that changed from the base of the PR and between 2a41f18 and 1dfc4f9.

📒 Files selected for processing (1)
  • docs/plans/2026-06-25-write-behind-lease-durability.md

📝 Walkthrough

Summary by CodeRabbit

  • Documentation
    • Added an architecture proposal for improving scheduler lease durability during heartbeat handling.
    • Described partition-safe safeguards to reduce duplicate external side effects.
    • Documented durability/timing requirements, validation scenarios, rollout/migration guidance, and risk/mitigation notes for the planned change.

Walkthrough

Adds an ADR-style plan for write-behind lease durability, covering the problem statement, timing constraints, buffered heartbeat writes, dispatch-time side-effect gating, dual-backend expectations, validation and rollout steps, and an agreement ledger.

Changes

Write-behind lease durability plan

Layer / File(s) Summary
Scope and mechanism
docs/plans/2026-06-25-write-behind-lease-durability.md
Defines the problem statement, goals, non-goals, timing bounds, durability classification, buffered lease cache, and dispatch-time gating model.
Backend behavior and rollout
docs/plans/2026-06-25-write-behind-lease-durability.md
Records backend expectations, rejected alternatives, durability metadata implications, runtime behavior notes, and rollout gating decisions.
Validation and risks
docs/plans/2026-06-25-write-behind-lease-durability.md
Lays out scheduler and dispatch validation cases, parity checks, chaos coverage, and the associated risk and mitigation notes.
Agreement and appendix
docs/plans/2026-06-25-write-behind-lease-durability.md
Adds the agreement ledger, states there are no unresolved blockers, and appends the broader durable-write contention landscape and sequencing notes.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~2 minutes

Poem

A lease was penned with careful care,
With heartbeats buffered through the air.
Gates and fences held their line,
So side effects stay well-defined.
A plan now hums, precise and bright ✨

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description is useful but misses most required template sections, including Summary bullets, Change Type, Linked Issue, Validation, and review metadata. Restructure it to the repo template and add the missing sections: summary bullets, change type, linked issue, validation, security/DB impacts, rollback, and review track.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title matches the design doc and broadly follows conventional-commits style, with a clear scope and summary.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a design document for 'Write-Behind Lease Durability + Side-Effect Gate' to address runner heartbeat latency issues on Postgres/Supabase. The design splits writes into SyncCritical and AsyncLossTolerantCoalesced planes, introduces a LeaseWriteBehind cache with a background drain task, and implements a side-effect gate to prevent double execution during network partitions. The review feedback highlights several critical areas for refinement: ensuring the drain task's CAS skip condition also checks for fence changes, using safe/checked arithmetic for the freshness gate subtraction to prevent panics, running the voluntary self-stop check on a frequent cadence to avoid sampling delays, and correcting a mathematical inconsistency in the safety buffer calculation.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

`LeaseWriteBehind` — per-replica `Arc` singleton:
- `dirty: std::sync::Mutex<HashMap<RunId, LeaseIntent>>`, `LeaseIntent{ expires_at, fence, last_durable_flush_succeeded_at, last_flushed_expiry, enqueued_at, consecutive_flush_failures }`.
- **Hot path `renew(run, fence, now)`** (30s heartbeat tick): `want = now + T_lease`; lock std mutex (NO `.await`, NO I/O); `if fence >= e.fence { e.fence = fence; e.expires_at = e.expires_at.max(want) }`. Non-blocking; cannot fail from Postgres latency. **The heartbeat tick no longer does a durable write → the single-heartbeat insta-fail is removed** (G1, G2).
- **Drain task** (one supervised `tokio::spawn` per replica): every `T_flush`, snapshot+coalesce (one row per run; **skip CAS if `expires_at == last_flushed_expiry`**), write via the existing `put_with_cas` sidecar with bounded concurrency; success → record `last_durable_flush_succeeded_at` + `last_flushed_expiry`; transient failure → keep dirty + bump `consecutive_flush_failures`; **CAS/fence conflict → mark reclaimed → cancel the local executor**. Durable write rate falls to ≤ 1 / `T_hb` per run (write *less*, not just *async*).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

When skipping the CAS write in the drain task, ensure the skip condition checks both expires_at == last_flushed_expiry and that the fence (or any other lease metadata) has not changed. If a fence update or other critical metadata changes but expires_at remains identical (e.g., due to rapid successive renewals or clock resolution limits), skipping the CAS write would prevent the new fence from being persisted, potentially causing subsequent correctness writes to fail CAS validation against the stale database fence.

### 4.3 The side-effect gate (load-bearing correctness mechanism)
Reactive fencing alone is insufficient: between flush attempts a partitioned owner can keep dispatching real side effects until another replica reclaims → double execution. Bound exposure with two purely-local (no Postgres round-trip) checks:

1. **Per-dispatch freshness gate (primary).** Synchronously, **immediately before each side-effecting tool/model call**: `remaining = last_flushed_expiry − now`; require `remaining > expected_op_duration + cancel_grace + S`. Else attempt **one synchronous flush**; if still no runway → **abort/pause the run** (`lease_degraded`). Evaluated **on the dispatch path itself**, off the 30s/5s cadence. **Uses the DURABLE `last_flushed_expiry`, never the in-memory `expires_at`** — a `renew` that never flushed grants NO runway.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

When implementing the per-dispatch freshness gate, ensure that the subtraction last_flushed_expiry - now is performed safely to prevent panics or underflows. If last_flushed_expiry is less than now (which can easily happen during database lag or network partitions), a naive subtraction using unsigned duration types or std::time::Instant will panic. Use saturating subtraction or checked arithmetic (e.g., checked_sub or saturating_duration_since in Rust) to safely default remaining to zero or a negative duration, triggering the synchronous flush or abort path gracefully.

Reactive fencing alone is insufficient: between flush attempts a partitioned owner can keep dispatching real side effects until another replica reclaims → double execution. Bound exposure with two purely-local (no Postgres round-trip) checks:

1. **Per-dispatch freshness gate (primary).** Synchronously, **immediately before each side-effecting tool/model call**: `remaining = last_flushed_expiry − now`; require `remaining > expected_op_duration + cancel_grace + S`. Else attempt **one synchronous flush**; if still no runway → **abort/pause the run** (`lease_degraded`). Evaluated **on the dispatch path itself**, off the 30s/5s cadence. **Uses the DURABLE `last_flushed_expiry`, never the in-memory `expires_at`** — a `renew` that never flushed grants NO runway.
2. **Run-level voluntary self-stop (backstop).** If `now − last_durable_flush_succeeded_at > D_stop (45s)`, cancel the run. `D_stop=45 < 60` (heartbeat→durable-expiry margin) leaves a `15s − S` buffer to stop **before** any other replica could legitimately reclaim.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

To ensure the run-level voluntary self-stop is enforced promptly at D_stop (45s), the check should be evaluated on a frequent cadence (such as the 5s drain task or the executor's main loop) rather than only on the 30s heartbeat tick. If evaluated only during the 30s heartbeat, a flush failure occurring shortly after a heartbeat could experience up to a 30s sampling delay before the self-stop is triggered, potentially pushing the actual stop time past the safe D_stop window.

A monotonic `fence` (generation) per run persisted on the lease sidecar AND validated on **every** correctness write — `complete`, `fail`, `block`, **tool-result/side-effect commit**, and recovery `reclaim` — not only the heartbeat CAS. A reclaim bumps/tombstones the fence; any stale-fence write fails CAS. **Idempotency records (SyncCritical) are the third line.**

### 4.5 Bounds (explicit, named variables)
- **No false reclaim of a live run:** owner self-stops at `D_stop=45s`; durable lease valid `≥ 60s` past the last flushed heartbeat; `45 + S < 60` ⇒ owner stops before reclaim is possible (requires `S < 15s`).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

There appears to be a minor mathematical inconsistency in the safety buffer calculation: If T_lease = 90s and the owner self-stops at D_stop = 45s of durable lag (measured from the last successful flush), the owner stops at t = 45s relative to that flush. Since the lease expires at t = 90s on the database clock, the other replica can reclaim at t = 90s - S. Therefore, the actual safety buffer before a legitimate reclaim is (90 - S) - 45 = 45s - S, rather than 15s - S. The 15s - S buffer would only apply if the lease duration was 60s or if D_stop was 75s.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/plans/2026-06-25-write-behind-lease-durability.md`:
- Around line 3-4: The plan status is too strong because the fence-vs-runner_id
durability decision is still unresolved in the document. Update the status block
in the accepted section to either reflect that the contract is not yet closed or
explicitly name the committed fallback/choice, and make sure the sections
referenced by the write-behind lease design are consistent with that decision so
the “Accepted”/“No unresolved blockers” wording is accurate.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 453723cf-bff4-4970-9976-122501a610bd

📥 Commits

Reviewing files that changed from the base of the PR and between a38119f and 2a41f18.

📒 Files selected for processing (1)
  • docs/plans/2026-06-25-write-behind-lease-durability.md

Comment on lines +3 to +4
**Status:** Accepted (council signoff — anthropic-slot Opus 4.8: ACCEPT; openai-slot GPT 5.5 XHigh: ACCEPT_WITH_NONBLOCKING_NOTES). No unresolved blockers.
**Origin:** fusion-design-council (Opus 4.8 + GPT 5.5 XHigh), 1 draft round + 1 cross-review round + signoff.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Status overstates closure.

§6 still leaves the fence-vs-runner_id durability choice open (“MUST be explicit”, “Commit to one before enable”), so Status: Accepted / No unresolved blockers is misleading. Please either name the committed fallback here or downgrade the status until that enablement-critical contract is fixed.

Also applies to: 67-68, 114-115

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/plans/2026-06-25-write-behind-lease-durability.md` around lines 3 - 4,
The plan status is too strong because the fence-vs-runner_id durability decision
is still unresolved in the document. Update the status block in the accepted
section to either reflect that the contract is not yet closed or explicitly name
the committed fallback/choice, and make sure the sections referenced by the
write-behind lease design are consistent with that decision so the
“Accepted”/“No unresolved blockers” wording is accurate.

@henrypark133

Copy link
Copy Markdown
Collaborator

Strong design — the durability-plane split (write-behind heartbeats vs synchronously-durable ownership/side-effect writes) is exactly right. A few things that may be worth folding in, from tracing the Postgres heartbeat path while debugging the "heartbeat could not be recorded" cascade:

1. The connection pool itself is likely the dominant bottleneck, and the doc does not address it.
DEFAULT_POSTGRES_POOL_MAX_SIZE = 2 (ironclaw_reborn_event_store/src/lib.rs:55) — a 2-connection pool shared across all Postgres FS I/O (LLM output, tool results, turn/run state, event log, lease). deadpool has no checkout timeout, so a starved heartbeat checkout blocks until the 15s FILESYSTEM_APPLY_TIMEOUT fires, then returns Err → scheduler_heartbeat_failed → run killed. Write-behind removes the heartbeat from this contention, but the drain task + every ownership-critical write still share those 2 slots. Suggest the design also cover: (a) raising the pool size for the hosted/cross-region profile, and/or a dedicated small connection reserved for lease/critical writes; (b) a checkout timeout shorter than the apply timeout so a starved write fails fast and retries instead of burning 15s. This is a cheap mitigation that helps before the write-behind work lands.

2. Consider generalizing the coalescing drain beyond heartbeats to the append-only event log.
Each put/append_file is its own INSERT round-trip (postgres.rs:167,505); there are many event appends per turn. The event log is append-only and ordered by SeqNo with no CAS precondition, so it micro-batches cleanly (one multi-row INSERT or tokio-postgres pipelining per drain window) at much higher QPS savings than heartbeats alone — and the same write-behind buffer/drain machinery you are building covers it. (Strictly liveness/projection/event writes only — never the side-effect-gating writes, per your split.)

3. Relationship to #5234 (remove per-record lock convoys via shared cas_update).
The doc references #5232 but not #5234. #5234 deletes the per-record tokio::Mutex held across the PG .await on the 5 turn-state stores — it removes the in-process convoy amplification, complementary to this PR’s pool-latency fix. Worth noting the write-behind cache should sit on top of the post-#5234 CAS path, not the old mutex one, so they compose cleanly.

4. Drain/critical-write round-trips. Where the drain flushes a batch or commits a set of correctness writes, tokio-postgres pipelining (or a single transaction) would cut the cross-region round-trip count further.

None of these block the design — just additive coverage so the implementation PR closes the cascade on all three layers (in-process convoy → #5234, pool starvation → pool size/timeout, sync heartbeat hot path → this PR).

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

@henrypark133 DEFAULT_POSTGRES_POOL_MAX_SIZE = 2 this is overwritten in our deployments as 16.
I am working generalising this design to cover every write that is not hard required to be atomic.

Fold in PR review (henrypark133) + a durable-write hot-path audit. Adds §12
framing the deployed lease-expiry cascade as three independent layers:
- Layer A: in-process lock convoy → #5234 (open); write-behind composes on the
  post-#5234 CAS path.
- Layer B: pool starvation — DEFAULT_POSTGRES_POOL_MAX_SIZE=2 shared across all
  Postgres FS I/O. Notes the existing 30s checkout guard (closes the infinite
  hang) but flags pool-too-small + checkout(30s)>apply(15s); cheap mitigations
  (raise pool, reserve a critical connection, align checkout<apply) prior to and
  complementary with write-behind.
- Layer C: the synchronous hot-path write map (events / governor / thread-append
  / lease / memory) with the write-behind-vs-batch-coalesce distinction. Events
  are the top batch-coalesce target (highest churn, O(1) INSERT no CAS); must
  stay DURABLE (source of truth), per-step flush to preserve live SSE. Memory
  stays synchronous (FTS-only, no embedding write; agent-initiated, low churn).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5249 June 25, 2026 16:12 Destroyed
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Thanks — this is exactly the systems-level framing the doc was missing. I verified each point against current main and folded all four into a new §12 (broader durable-write contention landscape), framed as the three-layer cascade you described (in-process convoy → pool starvation → sync hot path).

1. Pool — confirmed and amplified, with one correction. DEFAULT_POSTGRES_POOL_MAX_SIZE = 2 is real (ironclaw_reborn_event_store/src/lib.rs:8), shared across all Postgres FS I/O — agreed this is likely the dominant bottleneck, and write-behind only removes the heartbeat from it while the drain + every SyncCritical write still share the 2 slots. One correction: there is a checkout timeout now — POOL_CHECKOUT_TIMEOUT = 30s via .wait_timeout/.create_timeout/.recycle_timeout (lib.rs:561, 628–639), added explicitly as a deadlock guard "well under the 90s runner lease." So the infinite Pool::get() hang is already closed. But your deeper point stands: (a) 2 is far too small for a hosted cross-region profile, and (b) the 30s checkout is longer than the 15s FILESYSTEM_APPLY_TIMEOUT, so the apply timeout still wins on a starved write. Doc now recommends raise-pool + a reserved connection/small pool for lease+critical writes + aligning checkout below the apply timeout, as the cheap mitigation to do first.

2. Generalize coalescing to the event log — strongly agreed; an independent audit reached the same conclusion. The event-append plane is the highest-churn durable path per turn (~2M+2C+ appends), and each is a clean O(1) INSERT … RETURNING id (BIGSERIAL, no CAS, no head scan) — the safest thing in the system to micro-batch. One nuance I'd flag on framing: events are source of truth ("LLM data is never deleted"), so this should be a durable batch-coalesce (multi-row INSERT / pipeline / txn, committed before the step completes) — not the lossy write-behind buffer used for heartbeats, where a crash legitimately drops the value. Same drain shape, different durability contract. Kept the flush window per-step so live SSE isn't stalled. (If a specific event class turns out genuinely loss-tolerant, only those could move to lossy write-behind.)

3. #5234 — noted. Added as Layer A; doc now states the write-behind cache must compose on the post-#5234 cas_update path, not the old per-record-mutex path.

4. Pipelining/transactions — added to Layer C for both the drain flush and any batched SyncCritical commit set, to cut cross-region round-trips.

Net sequencing in the doc: (B) pool size + reserved connection [cheap, now] → land #5234 (A) → lease write-behind (this PR, C) → batch-coalesce events (C, top) → governor shard+coalesce → thread-append coalesce. Memory stays synchronous (verified FTS-only, no embedding write, agent-initiated). Pushed as 1dfc4f91c.

@railway-app

railway-app Bot commented Jun 25, 2026

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5249 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jun 25, 2026 at 4:13 pm

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5249 — 1dfc4f91 Deployed Jun 25, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: XS < 10 changed lines (excluding docs)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants