Skip to content

feat(reborn): flip inmemory-turn-state profile onto the durable row store + block-persistence migration (#6263 Step 4) - #6305

Merged
ilblackdragon merged 1 commit into
refactor/turn-state-write-behindfrom
refactor/turn-state-profile-write-behind
Jul 20, 2026
Merged

ilblackdragon merged 1 commit into
refactor/turn-state-write-behindfrom
refactor/turn-state-profile-write-behind

Conversation

@ilblackdragon

Copy link
Copy Markdown
Member

Step 4 of the turn-state consolidation (#6263). Stacked on #6298 (base is refactor/turn-state-write-behind; retarget to main after it merges). Moves the live inmemory-turn-state production profile off the raw InMemoryTurnStateStore authority onto the durable FilesystemTurnStateRowStore.

Ships WriteThrough, not WriteBehind — deliberately

Validating the WriteBehind flip surfaced a runtime-breaking read-your-writes defect (the durable-read query paths bypassed the hot cache; a read right after an async submit returned ScopeNotFound, failing every turn). That's fixed in #6298 (the read-side barrier commit + IronLoop's hardening). This PR lands the profile on WriteThrough, which is already strictly more durable than the old raw in-memory authority + persist-on-block: every transition is synchronously durable, crash-recoverable, and rehydrates from its own rows — with no per-user state.json CAS livelock (journal/row model). Flipping this same profile to WriteBehind is the follow-on (Step 5) now that the read path is correct.

What

  • Type unification: ComposedTurnStateStore is now FilesystemTurnStateStoreKind unconditionally — the inmemory-turn-state feature no longer selects the store type, only (eventually) the durability policy. Ripple was minimal since the non-feature arm already used Kind; no downstream consumer changed.
  • Factory flip: the inmemory-turn-state arm builds the durable row store (WriteThrough) instead of InMemoryTurnStateStore + FilesystemTurnStateBlockPersistence (block-persistence load / from_persistence_snapshot / with_block_persistence all removed).
  • Migration — verified automatic: FilesystemTurnStateBlockPersistence writes /turns/state.json via io::snapshot_path; the row store's migrate_legacy_blob_if_needed reads the same path + format, so an existing deployment's gate-parked/approval snapshot imports as the first delta on boot — no turn lost. Pinned by a regression test that writes a real BlockedApproval snapshot exactly as the old sink did and asserts the row store recovers it.
  • drain() seam: added FilesystemTurnStateRowStore::drain() (flush the WriteBehind ack window; no-op under WriteThrough/empty) forwarded through Kind; runtime.rs shutdown() repointed from the removed InMemoryTurnStateStore::flush() to Kind::drain(). No-op today (WriteThrough), load-bearing once WriteBehind lands.
  • Composition integration test drives a real turn to Completed over the flipped store via build_reborn_runtime and asserts graceful shutdown() drains.
  • libSQL stress: added --turn-state-durability; on a real libSQL file WriteBehind beats WriteThrough at every concurrency with no livelock (results under tools/ironclaw_stress/results/2026-07-20-turn-state-write-behind-libsql/).

Operator note

No drain-before-deploy requirement (WriteThrough persists every transition synchronously; the shutdown drain is a no-op). The one operational item is the automatic migration on first boot after upgrade — verified, but a backup of /turns/state.json before upgrading is a cheap safety net.

Tests

ironclaw_turns green; composition full suite under the production feature set (webui-v2-beta,slack-v2-host-beta,libsql,postgres,inmemory-turn-state,test-support) — 0 failures; cargo check across production / libsql,postgres / no-default / CLI-default; clippy -D warnings; fmt; pre-commit-safety exit 0. Ratchet unchanged (InMemoryTurnStateStore stays in FROZEN_DEBT; Step 5 eliminates the type).

🤖 Generated with Claude Code

…tore (WriteThrough); type unification, block-persistence migration, drain seam (#6263 Step 4)

Ships WriteThrough (NOT WriteBehind) because Step 4 validation found a
runtime-breaking read-path defect in the WriteBehind capability (#6298):
get_run_state and the other durable-read query methods bypass the hot
cache, so under WriteBehind a read immediately after an async submit
returns ScopeNotFound and every runtime turn fails. Fix is a separate
follow-up on the capability; this flip lands on WriteThrough, which is
still strictly more durable than the old raw in-memory authority +
persist-on-block.

- ComposedTurnStateStore unified to FilesystemTurnStateStoreKind
  unconditionally (feature no longer selects the store type; downstream
  already supported Kind so ripple is minimal).
- inmemory-turn-state arm now builds the durable row store (WriteThrough)
  instead of InMemoryTurnStateStore + FilesystemTurnStateBlockPersistence.
- Migration verified automatic: block-persistence writes /turns/state.json
  via io::snapshot_path; the row store's migrate_legacy_blob_if_needed
  reads the SAME path+format, so an existing deployment's gate-parked
  snapshot imports as the first delta on boot (regression test added).
- Row store drain() + Kind forwarding; runtime shutdown() repointed from
  the removed InMemoryTurnStateStore::flush() to Kind::drain() (no-op
  under WriteThrough; the graceful-restart seam WriteBehind will use).
- Composition integration test drives a real turn to Completed over the
  flipped store and asserts graceful shutdown drains.
- libSQL stress: added --turn-state-durability; WriteBehind beats
  WriteThrough at every concurrency with no livelock (results under
  tools/ironclaw_stress/results/2026-07-20-turn-state-write-behind-libsql).
- Ratchet unchanged (InMemoryTurnStateStore stays in FROZEN_DEBT; the
  follow-up eliminates the type).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Jul 20, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • staging
  • reborn-integration

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9536bdd8-a73f-4e09-b48d-1e98db34aad2

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6305 July 20, 2026 04:15 Destroyed
@ironloopai

ironloopai Bot commented Jul 20, 2026 •

Copy link
Copy Markdown
Contributor

🔎 IronLoop Review Status

Head: 0a198d2408ba6c0fd731ce62904e6b88d58afebe
Result: 1 blocking finding across 1 reviewer.
Next: Address the blocking findings, push fixes, then re-run the relevant reviewer.
Updated: 2026-07-20T04:20:45.291Z

Current reviewers:

Reviewer State Verdict Findings Last update
ironloop/common-reviewer (reviewer) Completed Changes requested 1 blocking finding / 1 note 2026-07-20T04:20:45.281Z
Reviewer summaries
Reviewer Detail
ironloop/common-reviewer (reviewer) Changes requested; 1 blocking finding; The profile flip has a data-loss migration race during overlapping old/new deployments. The PR also includes a massive ignored vendor payload.
Recent activity
Time Reviewer State Detail
2026-07-20T04:15:47.474Z ironloop/common-reviewer (reviewer) Queued Accepted review request for head 0a198d2.
2026-07-20T04:15:47.474Z ironloop/common-reviewer (reviewer) Queued Waiting for this reviewer lane to become available.
2026-07-20T04:15:48.423Z ironloop/common-reviewer (reviewer) Started Reviewer worker started.
2026-07-20T04:15:53.643Z ironloop/common-reviewer (reviewer) Workspace ready Prepared isolated checkout (merge_ref) at a7fe33e.
2026-07-20T04:20:45.281Z ironloop/common-reviewer (reviewer) Result captured Changes requested; 1 blocking finding.
2026-07-20T04:20:45.281Z ironloop/common-reviewer (reviewer) Completed Review completed and terminal status was persisted.
Available commands
  • @ironloopai help
  • @ironloopai agents
  • @ironloopai review
  • @ironloopai review --agent <agent>
Run metadata

Admission: webhook accepted the request and IronLoop persisted reviewer state before this projection.

@github-actions github-actions Bot added size: XL 500+ changed lines scope: docs Documentation risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs and removed size: XL 500+ changed lines labels Jul 20, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request transitions the turn-state storage architecture to use the production row-based filesystem store (FilesystemTurnStateStoreKind) unconditionally across all deployments, replacing the legacy standalone in-memory store. It introduces a drain mechanism to safely flush the asynchronous WriteBehind queue during graceful shutdowns and provides automatic migration of legacy block-persistence snapshots on first boot. Additionally, a critical issue was raised regarding the accidental commitment of the entire frontend node_modules directory, which contains absolute local paths and must be removed from the repository.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +1 to +28
#!/bin/sh
basedir=$(dirname "$(echo "$0" | sed -e 's,\\,/,g')")
basedir_win="$basedir"
exe=""
msys=""

case `uname -a` in
*CYGWIN*|*MINGW*|*MSYS*)
if command -v cygpath > /dev/null 2>&1; then
basedir_win=`cygpath -w "$basedir"`
fi
exe=".exe"
msys="true"
;;
*WSL2*)
if command -v wslpath > /dev/null 2>&1; then
basedir_win="$(wslpath -w "$basedir" 2> /dev/null)"
if [ $? -ne 0 ] || [ -z "$basedir_win" ]; then
basedir_win="$basedir"
else
exe=".exe"
fi
fi
;;
esac

if [ -z "$NODE_PATH" ]; then
export NODE_PATH="/data/illia/ironclaw5/crates/ironclaw_webui_v2/frontend/node_modules/.pnpm/marked@17.0.2/node_modules/marked/node_modules:/data/illia/ironclaw5/crates/ironclaw_webui_v2/frontend/node_modules/.pnpm/marked@17.0.2/node_modules:/data/illia/ironclaw5/crates/ironclaw_webui_v2/frontend/node_modules/.pnpm/node_modules"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

This file, along with many others in this pull request, appears to be part of a node_modules directory. Committing node_modules to version control is generally discouraged for several reasons:

  1. Repository Size: It significantly inflates the repository size with dependencies that can be fetched on-demand using a package manager like pnpm, npm, or yarn.
  2. Platform Incompatibility: It can lead to platform-specific binaries being checked in, causing build failures for other developers on different operating systems.
  3. Local Paths: These files contain absolute local paths specific to your machine (e.g., /data/illia/ironclaw5/...), which will break the build for anyone else who clones the repository.

It is strongly recommended to:

  • Remove the entire crates/ironclaw_webui_v2/frontend/node_modules/ directory from this pull request.
  • Add node_modules/ to the .gitignore file in the crates/ironclaw_webui_v2/frontend/ directory to prevent these files from being committed in the future.

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

❌ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
❌ Changes requested 1 1 1 0a198d2408ba

Head: 0a198d2408ba6c0fd731ce62904e6b88d58afebe
Next: Fix the blocking findings, push the PR branch, then re-run this reviewer.

Run details

Status: Current
Needs human: no
Needs validation: no

Summary

The profile flip has a data-loss migration race during overlapping old/new deployments. The PR also includes a massive ignored vendor payload.

Findings

Blocking: 1 / Notes: 1

Blocking findings

1. ❌ [HIGH] Require an exclusive legacy-to-row-store cutover

Location: crates/ironclaw_reborn_composition/src/factory.rs:2519-2525
The new production path starts the row store with automatic blob migration, but it does not enforce the migration contract's “no live turn writers” cutover. If an old inmemory-turn-state process is still serving while this instance imports /turns/state.json, a later old-process gate transition or graceful-shutdown flush overwrites that blob after rows exist. The row store intentionally ignores blobs once it has row data, so that newly persisted blocked/in-flight turn is never imported and is lost when the old process exits. Gate startup behind an exclusive migration/quiescence handoff (or fail closed until the old writer is stopped) and add an overlap regression test.

Non-blocking notes (1)
1. 💬 [LOW] Remove the ignored frontend node_modules payload

Location: crates/ironclaw_webui_v2/frontend/node_modules
This PR adds the ignored node_modules tree under an otherwise removed frontend path (4,058 vendor files and roughly 1.38 million inserted lines) without source-manifest or lockfile changes. It is unrelated to the turn-state migration, makes the dependency payload unauditable in this review, and unnecessarily ships platform binaries. Remove these generated files from the change.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.

Arc::new(store.with_block_persistence(block_persistence))
};
let turn_state = Arc::new(
production_turn_state_store(Arc::clone(&turn_state_filesystem), turn_state_store_limits)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This automatic migration is only safe after the legacy writer is quiesced. An old process can persist a newer /turns/state.json after this instance imports it; once rows exist, the importer intentionally ignores the blob, so that old-process gate/in-flight turn is never copied and is lost when the old process exits. Enforce an exclusive cutover (or fail closed until the old writer stops) and cover old/new overlap in a regression test.

@railway-app

railway-app Bot commented Jul 20, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6305 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 20, 2026 at 4:31 am

@ilblackdragon

Copy link
Copy Markdown
Member Author

Reviewed — two items, and a note that this is gated on #6298 landing first (it's stacked on that branch).

HIGH — exclusive legacy→row-store cutover (factory.rs:2519-2525): Real migration-safety gap and, like #6271's gate-migration, a deployment-strategy decision rather than a mechanical fix. The row store imports /turns/state.json on startup then ignores blobs once it has rows, so an old inmemory-turn-state process still serving during the import can flush a blob after rows exist and silently lose a just-persisted blocked/in-flight turn. Two sound shapes, your call:

  • Drain/quiesce handoff — stop the old writer (or take an exclusive migration lock) before the new instance imports; the new instance fails closed until it can prove no live legacy writer.
  • One-shot import barrier — import under an exclusive lease that the old process must release, and refuse to serve turns until the import lease is held.
    Either way this needs an overlap regression test (old-writer flush during/after row import must not lose the turn). Happy to implement whichever you pick.

LOW — accidental crates/ironclaw_webui_v2/frontend/node_modules payload (~4,058 files / 1.38M lines): unrelated vendored node_modules committed with no lockfile/manifest change — makes the diff unreviewable and ships platform binaries. Remove with git rm -r --cached crates/ironclaw_webui_v2/frontend/node_modules and add it to .gitignore. I can do this cleanup if you'd like, but flagging first since it's your WIP branch and the removal is a large history-touching commit.

Which cutover shape do you want? I'll wire it + the overlap test once #6298 is settled.

@ilblackdragon
ilblackdragon merged commit 06ee5ad into refactor/turn-state-write-behind Jul 20, 2026
24 checks passed
@ilblackdragon
ilblackdragon deleted the refactor/turn-state-profile-write-behind branch July 20, 2026 04:58
ilblackdragon added a commit that referenced this pull request Jul 20, 2026
…6298/#6305 IronLoop)

`crates/ironclaw_webui_v2/frontend/node_modules` was committed with no source,
manifest, lockfile, or Cargo.toml — the directory holds *only* vendored deps and
the crate is not even a workspace member. The 4,058-file / ~1.38M-line payload
dominated the diff and caused IronLoop to decline the review as unauditable.

Remove all 4,058 tracked node_modules files (working tree untouched) and add a
`.gitignore` rule so it can't be re-added. No source or build change: nothing
references or builds this vendored tree.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ilblackdragon added a commit that referenced this pull request Jul 20, 2026
…tate row store (#6263 Step 3) (#6298)

* feat(turns): opt-in async write-behind durability mode for the turn-state row store (#6263 Step 3)

Adds TurnStateDurabilityPolicy { WriteThrough (default), WriteBehind } to
FilesystemTurnStateRowStore. WriteThrough is byte-for-byte today's behavior
(every mutation awaits the durable ack) — default deployments get zero
change and no crash-loss window. This PR adds the capability + proves it +
measures it; wiring the volatile profile onto WriteBehind (retiring the raw
direct authority) is a separate follow-on.

WriteBehind: a mutation whose durable delta is NOT recoverability-critical
returns Ok immediately after enqueue (background flush); a critical
transition (is_recoverability_critical = is_blocked() || is_terminal(),
now a production predicate in status.rs the crash suite references) awaits
the ack. Because the journal is a strictly sequential single-writer,
awaiting a critical op's ack implies the whole prior async tail is durable
— critical ops are natural barriers, so a run can only ever lose trailing
non-critical churn, never a gate-park or terminal a human/model waits on.

Write-behind robustness (write-through gets these for free):
- Backpressure: bounded pending-un-acked delta window
  (max_pending_write_behind_deltas, default 128); at the cap a non-critical
  op awaits the oldest ack. Bounds memory + the crash-loss window.
- Append-failure halt: the flusher latches degraded + closes on append
  failure in write-behind (write-through still continues), so no durable
  gap; mutations then fail fast and reads reload from the rolled-back
  durable point. Pre-append cross-store-CAS reservations are disabled in
  write-behind (incompatible: a critical op's reservation can't find rows
  the async ops never wrote synchronously).
- Cache safety: a rejected mutation rebuilds the engine from the cached
  (unflushed-inclusive) snapshot instead of reload-from-durable, so acked
  non-critical ops the caller was told succeeded aren't silently dropped.
- Journal abort-on-drop so a 'crashed' store can't flush its queued tail
  post-crash and race a store reopened over the same backend.

Crash suite parameterized over both modes (19 tests, 0 ignored):
WriteThrough keeps the strict acked-durable projection diff; WriteBehind
keeps assert_recoverability_critical_survives strict (gate-park + terminal
never lost) and replaces the diff with a legal-prefix check PLUS an
anti-cheat convergence test — fork the durable bytes before K acked
non-critical ops, recover, verify prefix, then re-apply the K lost ops and
assert convergence to the model (loss = redoable work, not corruption).
Barrier, append-failure-halt/degrade/recover, and backpressure tests added.

Measured (store-isolated per-transition, InMemoryBackend): WriteBehind
non-critical writes ~45-190us (vs WriteThrough ~2ms c=1 / ~20ms c=32
durable-ack floor — 10-400x), approaching the direct authority; critical
transitions stay synchronously durable in both.

Follow-on: graceful drain-before-shutdown when the profile is wired onto
WriteBehind (abort-on-drop currently discards the un-flushed tail at drop).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(turns): close 3 write-behind hazards IronLoop found (#6298 f1/f2/f3)

All three are write-behind-only; WriteThrough is unaffected. The crash suite
(now 22 tests) passes in both modes; each fix ships a regression test verified
to fail before the fix.

f1 [HIGH] — CancelRequested must be a durability barrier. `request_cancel`
reports success once the transition commits, so a write-behind crash that
reverts a Running run's acked CancelRequested to Running re-executes work the
caller was told was cancelled (and drops its idempotency record). Add
`CancelRequested` to `is_recoverability_critical` (the single boundary the row
store and the crash-suite oracle both key on), so the cancel awaits its durable
ack. Test: `write_behind_cancel_of_running_run_survives_crash` (without the fix
the whole un-flushed tail is lost → ScopeNotFound).

f2 [MED] — read-your-writes. `get_run_state` read only durable rows, but a
non-critical write-behind submit returns Ok after updating the hot snapshot and
before its durable append, so an immediate same-store read returned
ScopeNotFound/stale. Serve `get_run_state` from the hot snapshot under healthy
write-behind (the single-writer authority); write-through and the degraded path
still read durable. Renames `read_run_state_for_cancellation` →
`read_run_state_from_hot_cache` (it was never cancellation-specific). Test:
`write_behind_get_run_state_reflects_unflushed_submit`.

f3 [MED] — backpressure was applied AFTER an unbounded enqueue: concurrent
callers each enqueued into the journal channel, then serialized on the pending
window, so a stalled flusher let the channel grow without bound. Reserve the
pending-window slot and track the ack BEFORE/around the enqueue, both inside
`apply` under the `snapshot_state` lock that serializes enqueue, so the channel
can never exceed the cap. `register_write_behind_commit` is removed;
`commit_pending` short-circuits on the now-`None` ack. Test:
`write_behind_concurrent_writers_under_cap_stay_consistent` exercises the
concurrent reserve→enqueue→track path (the strict peak-depth bound is structural
— the journal channel length is not externally observable).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(turns): route put_loop_checkpoint through the write-behind async path (#6298 IronLoop f4)

The f3 backpressure fix added a `commit_pending` debug assertion that a
non-critical write-behind commit is never awaited there — it must be
reserved+tracked in `apply`. But `put_loop_checkpoint` enqueues under
`snapshot_state` and hands a live ack straight to `commit_pending` with
`critical: false`, bypassing that flow. Under WriteBehind it therefore tripped
the assertion in debug/test builds, and in release waited synchronously —
contradicting the intended lazy flush for a non-critical checkpoint (and leaving
the same unbounded-channel gap f3 closed elsewhere).

Give the checkpoint enqueue the same reserve→enqueue→track flow every other
async commit uses: reserve the pending-window slot before enqueue (under
`snapshot_state`), then `track_write_behind_ack_if_async` returns `None` so
`commit_pending` lazy-flushes. WriteThrough awaits the ack unchanged.

Regression: `write_behind_put_loop_checkpoint_takes_async_path` drives the real
`put_loop_checkpoint` under WriteBehind — verified it panics on the assertion
before this fix. Crash suite now 23 tests / 0 ignored, green in both modes;
clippy clean (all-features + default).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(turns): cargo fmt the checkpoint-test import block (#6298)

Import reordering only — the f4 regression-test imports were added out of sorted
order and tripped the Formatting / Code Style lane. No behavior change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(turns): read-your-writes barrier for WriteBehind durable-read query paths (#6263 Step 3)

The runtime does submit_turn -> get_run_state, but get_run_state (and
read_turn_events_after / get_loop_checkpoint) read materialized durable
rows. Under WriteBehind a submit's delta materializes asynchronously, so
the immediate read saw stale/missing rows and returned ScopeNotFound —
every runtime turn would fail. (WriteThrough was unaffected: durable ==
cache there.) The crash suite never caught it because it only ever read
AFTER recovery, never live after an async write.

The naive fix (serve reads from the hot cache) is WRONG and loses data:
the cache is a bounded window, not a superset of durable state —
terminal runs/events are evicted (max_terminal_records / max_events)
while durable rows retain them, and cross-writer freshness needs the
durable read. Four WriteThrough contract tests pin exactly this.

Fix: a read-side barrier. flush_pending_write_behind_for_read() drains
the enqueued-but-un-acked non-critical pending window (awaiting those
durable appends) at the top of the three durable-read helpers, so durable
rows catch up to the just-acked writes BEFORE the read — read-your-writes
consistent, while the durable read keeps its exact eviction /
cross-writer-freshness / scope / cursor / retention semantics. No-op
under WriteThrough (window always empty), so that path is byte-for-byte
unchanged. get_run_state_for_cancellation and get_run_record already read
the cache and were WB-safe; traits.rs untouched.

Coverage gap closed: live read-your-writes tests in BOTH modes (submit ->
immediate get_run_state/get_run_record/read_turn_events_after,
put -> get_loop_checkpoint) plus a property test asserting live
get_run_state tracks the model after every op. Reverting only the source
reproduces model=Queued vs store=ScopeNotFound under WriteBehind.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(reborn): flip inmemory-turn-state profile onto the durable row store (WriteThrough); type unification, block-persistence migration, drain seam (#6263 Step 4) (#6305)

Ships WriteThrough (NOT WriteBehind) because Step 4 validation found a
runtime-breaking read-path defect in the WriteBehind capability (#6298):
get_run_state and the other durable-read query methods bypass the hot
cache, so under WriteBehind a read immediately after an async submit
returns ScopeNotFound and every runtime turn fails. Fix is a separate
follow-up on the capability; this flip lands on WriteThrough, which is
still strictly more durable than the old raw in-memory authority +
persist-on-block.

- ComposedTurnStateStore unified to FilesystemTurnStateStoreKind
  unconditionally (feature no longer selects the store type; downstream
  already supported Kind so ripple is minimal).
- inmemory-turn-state arm now builds the durable row store (WriteThrough)
  instead of InMemoryTurnStateStore + FilesystemTurnStateBlockPersistence.
- Migration verified automatic: block-persistence writes /turns/state.json
  via io::snapshot_path; the row store's migrate_legacy_blob_if_needed
  reads the SAME path+format, so an existing deployment's gate-parked
  snapshot imports as the first delta on boot (regression test added).
- Row store drain() + Kind forwarding; runtime shutdown() repointed from
  the removed InMemoryTurnStateStore::flush() to Kind::drain() (no-op
  under WriteThrough; the graceful-restart seam WriteBehind will use).
- Composition integration test drives a real turn to Completed over the
  flipped store and asserts graceful shutdown drains.
- libSQL stress: added --turn-state-durability; WriteBehind beats
  WriteThrough at every concurrency with no livelock (results under
  tools/ironclaw_stress/results/2026-07-20-turn-state-write-behind-libsql).
- Ratchet unchanged (InMemoryTurnStateStore stays in FROZEN_DEBT; the
  follow-up eliminates the type).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* fix(turns): 2 more write-behind read/checkpoint hazards IronLoop found (#6298 f5/f6)

f5 [HIGH] — `BeforeSideEffect` checkpoints must be synchronous. f4 made ALL
`put_loop_checkpoint` writes async under write-behind, but a `BeforeSideEffect`
checkpoint is written immediately before a capability's external side effect,
and expired-lease recovery treats the ABSENCE of a durable checkpoint as proof
no side effect ran (and requeues). If it lazy-flushed, a crash after the
capability ran but before the flush would replay a non-idempotent side effect.
Key the checkpoint's criticality on its kind: `BeforeSideEffect` → critical
(synchronous, durable before the capability executes); other kinds stay async
(a gate-park's checkpoint is flushed by the block transition's barrier; a lost
BeforeModel is redoable). Test:
`write_behind_before_side_effect_checkpoint_survives_crash`.

f6 [MED] — evicted-but-durable terminals hidden. f2 routed healthy-write-behind
`get_run_state` through the bounded hot snapshot, which evicts OLD TERMINAL runs
while their durable rows persist and must stay queryable (the eviction
contract). Once `max_terminal_records` is exceeded the hot-cache read returned
`ScopeNotFound` for those durable terminals. On a hot-cache miss, flush the
pending write-behind tail (read-your-writes) then fall back to the durable-row
lookup. Test:
`write_behind_get_run_state_finds_evicted_terminal_via_durable_fallback`.

Both verified failing before the fix. Crash suite now 27 tests / 0 ignored,
green in both modes; clippy clean (all-features + default); fmt clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: remove accidentally-committed webui_v2 frontend node_modules (#6298/#6305 IronLoop)

`crates/ironclaw_webui_v2/frontend/node_modules` was committed with no source,
manifest, lockfile, or Cargo.toml — the directory holds *only* vendored deps and
the crate is not even a workspace member. The 4,058-file / ~1.38M-line payload
dominated the diff and caused IronLoop to decline the review as unauditable.

Remove all 4,058 tracked node_modules files (working tree untouched) and add a
`.gitignore` rule so it can't be re-added. No source or build change: nothing
references or builds this vendored tree.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: broaden .gitignore to ignore node_modules at any depth (#6298)

Vendored node_modules trees kept getting accidentally committed (the
webui_v2 frontend payload; likely more as frontends land). Replace the
single-path rule with a bare `node_modules/` (no leading slash), which
git matches at ANY depth — so every current and future frontend/tooling
node_modules is ignored, not just the one path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(turns): preserve pending write-behind acks across cancellation (#6298 IronLoop f7)

`reserve_write_behind_slot` (backpressure) and `flush_pending_write_behind_for_read`
(read barrier) removed the oldest ack from the pending window BEFORE awaiting it —
`pop_front()` and `drain(..).collect()` respectively. Both run under a caller that
can cancel them (apply's outer timeout; a dropped read future). A cancellation
then dropped the removed ack permanently, so: (1) later writes saw an empty
window and enqueued behind a stalled append — the channel/loss window unbounded
again; and (2) a subsequent flush found nothing and falsely reported success while
the acknowledged write was still un-appended (lost on the next store drop).

Await each ack IN PLACE instead — peek the front, await it by `&mut` (new
`DeltaJournal::await_ack_ref`; a `oneshot::Receiver` is `Future + Unpin`, so
awaiting `&mut` it doesn't consume it), and remove it only once it resolves. A
cancelled await now leaves the ack tracked. The flush holds the window lock across
its awaits (bounded by the pending cap; the journal flusher doesn't take that
lock, so no deadlock), which also keeps its set to the writes present when the
flush began.

Regression: `write_behind_cancelled_flush_preserves_pending_ack` — under a stalled
flusher (new FaultBackend append gate), a timed-out first flush must leave the ack
so a second flush still blocks rather than falsely succeeding. Verified failing
before the fix. Crash suite 29/0-ignored green in both modes; clippy clean
(all-features + default); fmt clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6305 — 0a198d24 Deployed Jul 20, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant