Skip to content

[DRAFT] Optimize hosted Postgres turn-state latency - #5667

Closed
serrrfirat wants to merge 36 commits into
mainfrom
codex/root-filesystem-resource-governor
Closed

serrrfirat wants to merge 36 commits into
mainfrom
codex/root-filesystem-resource-governor

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 5, 2026 •

Copy link
Copy Markdown
Collaborator

Human comment: I will chop this up.

Summary

This PR is the hosted-single-tenant Postgres latency cycle. It moves the hot paths away from blob-style persistence and per-reservation Postgres transactions toward RootFilesystem-backed append/row stores with in-process authority where the deployment contract allows it.

Main changes:

  • Adds the RootFilesystem-backed resource governor with in-memory per-account authority and grouped durability.
  • Adds the filesystem row turn-state path, targeted row deltas, indexed replay, and grouped delta flushing via append_batch.
  • Switches libSQL hosted latency runs onto the same row turn-state path, removing the libSQL blob cliff.
  • Splits row-store turn state into durable Tier 2 run-record rows (turns, runs, events) and bounded hot-cache state; terminal runs/events are evicted from memory without deleting durable rows.
  • Adds ironclaw_stress --scenario turn-lifecycle-churn for c100 submit/claim/complete churn and RSS/cache validation.

Journey / starting point

Representative starting signals from the loop:

  • c100 full mixed flow: op p95 ~388.1ms, resource governor p95 ~213.3ms.
  • c100 turn lifecycle: p95 ~10.56s, effectively unusable.
  • libSQL c4 turn lifecycle on the blob path: p95 ~7,816ms versus Postgres row-store ~35ms, a 200x pathological gap.
  • Resource-only c100: p95 ~158.9ms, throughput ~859.5 ops/sec.

Main bottlenecks found:

  • Turn state and thread state were behaving like blob stores, so reads/writes grew with historical state size.
  • The Postgres resource governor serialized through per-reservation row-lock transactions.
  • The libSQL latency runner was still exercising the blob turn-state path while Postgres used row state.
  • Existing row-store cache limits were still able to encode terminal run/event pruning as durable deletion.

Ending point

Latest local gates on this branch:

  • libSQL c4 focused turn lifecycle: p95 30.9ms.
  • Postgres c4 focused turn lifecycle: p95 27.6ms.
  • Postgres c100 mixed flow, pool 32, filesystem-row, 100 users, 1,000 operations:
    • op p95 118.3ms, p99 120.3ms, throughput 1,587.9 ops/sec.
    • turn-store p95 35.9ms.
    • resource-governor p95 10.7ms.
    • thread-store writes p95 64.2ms.
  • Postgres c100 turn-lifecycle-churn, pool 32, filesystem-row, 10s measured after 3s warmup:
    • 18,057/18,057 succeeded.
    • op/turn-store p95 93.1ms.
    • throughput 1,798.6 ops/sec.

Notes:

  • The c100 duration churn run still reports RSS growth in the stress process, but that measurement is confounded by the harness retaining every sample plus Postgres client allocations. The row-store contract test now verifies the actual hot snapshot is bounded while old failed terminal runs and lifecycle events remain queryable after eviction and restart.
  • Row-store persistence currently has no materialized row compactor; this PR bounds active memory and prevents Tier 2 durable data loss, but it does not destructively rewrite old delta journal history.

Verification

Ran locally:

  • cargo check -p ironclaw_turns
  • cargo test -p ironclaw_turns --test filesystem_turn_state_contract
  • cargo test -p ironclaw_stress
  • cargo test -p ironclaw_filesystem --features libsql,postgres --test db_root_filesystem_contract
  • cargo test -p ironclaw_reborn_composition --features libsql,postgres --test libsql_substrate --test postgres_substrate
  • harness/latency/score.sh --dev plus focused c1/c4 reruns during the cycle
  • ironclaw_stress c100 mixed-flow and turn-lifecycle-churn runs listed above

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5667 July 5, 2026 22:13 Destroyed
@github-actions github-actions Bot added scope: docs Documentation scope: dependencies Dependency updates size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules labels Jul 5, 2026
@coderabbitai

coderabbitai Bot commented Jul 5, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Adds a StorageTxn::reserve_sequence primitive across libSQL/Postgres/scoped filesystems; replaces PersistentResourceGovernor with a new journaling FilesystemResourceGovernor everywhere it is wired; adds a PostgresSecretStore and a filesystem row-store turn-state backend; hardens trigger repositories; and introduces a latency benchmarking harness with goal/spec docs.

Changes

Filesystem reserve_sequence primitive

Layer / File(s) Summary
Trait contract and default
crates/ironclaw_filesystem/src/backend.rs
StorageTxn::reserve_sequence added with a default Unsupported implementation.
Scoped delegation
crates/ironclaw_filesystem/src/scoped.rs
ScopedStorageTxn enforces permission/mount checks then forwards to inner transaction.
Postgres implementation, migrations, stat/query caching
crates/ironclaw_filesystem/src/postgres.rs
reserve_sequence delegates to a shared SQL helper; run_migrations gains advisory-lock keying and legacy index cleanup; stat/query use prepared/cached statements; shared projection index naming added with tests.
libSQL connection retry/pragma handling
crates/ironclaw_filesystem/src/libsql.rs
connect_with_retry applies PRAGMAs via execute_batch with retry and clarifies failure messages.
Thread message sequence assignment
crates/ironclaw_threads/src/filesystem_service.rs
Legacy per-thread counter path split from native txn.reserve_sequence, handling Unsupported.

FilesystemResourceGovernor rollout

Layer / File(s) Summary
Core governor implementation
crates/ironclaw_resources/src/filesystem_governor.rs, .../lib.rs, .../filesystem_store.rs, tests/resource_governor_contract.rs
New journaling FilesystemResourceGovernor<F> with delta journal, compaction, sharded authority; snapshot schema bumped to v3; internal visibility widened; tests added.
Host runtime wiring
crates/ironclaw_host_runtime/src/services*.rs
Builder constructs FilesystemResourceGovernor directly, replacing persistent governor+store.
Reborn composition production/local-dev wiring
crates/ironclaw_reborn_composition/src/factory.rs, lib.rs, observability/budget.rs
Local-dev and production governor construction switched; Postgres migration keying/locking added; production builders generalized over governor type; FilesystemTurnStateStoreKind used for production turn state.
Stress tool governor/Postgres wiring
tools/ironclaw_stress/src/main.rs
governor_from_root builds FilesystemResourceGovernor; Postgres backend builder returns the pool.
CAS snapshot worker pool
crates/ironclaw_resources/src/cas_snapshot.rs
Single dedicated worker replaced with round-robin AsyncStorageWorkerPool.
Supporting cleanup
crates/ironclaw_resources/Cargo.toml, crates/ironclaw_reborn_composition/src/outbound/mod.rs, docs/reborn/contracts/resources.md
Unused optional deps removed; outbound re-exports test-gated; contract docs updated.

Postgres secret store & secret index scoping

Layer / File(s) Summary
PostgresSecretStore
crates/ironclaw_secrets/src/postgres_store.rs, lib.rs, Cargo.toml
New Postgres-backed SecretStore with migrations, encryption, and lease lifecycle.
Fixed-root tenant index scoping
crates/ironclaw_secrets/src/filesystem_store.rs
ensure_tenant_id_index_secret/broker now target a fixed /secrets root instead of per-scope prefixes.

Trigger repository hardening

Layer / File(s) Summary
libSQL operation gate
crates/ironclaw_triggers/src/libsql.rs
Per-repository mutex serializes all trigger operations; busy-timeout uses execute_batch.
Postgres cached statements
crates/ironclaw_triggers/src/postgres.rs
Upsert/list/scoped-list queries use new cached_query/cached_execute helpers.

Filesystem turn-state row store & lifecycle churn

Layer / File(s) Summary
Store kind selector
crates/ironclaw_turns/src/filesystem_store.rs, lib.rs
FilesystemTurnStateStoreKind selects between blob and row implementations.
Row store engine
crates/ironclaw_turns/src/filesystem_store/row_store.rs
New delta-journaled, append-log turn-state backend implementing all turn-state traits.
Projection & runner-lease support
crates/ironclaw_turns/src/filesystem_store/projection.rs, runner_lease.rs
Adds run-state/loop-checkpoint projections and lease overlay/seed helpers.
In-memory accessor helpers
crates/ironclaw_turns/src/memory/mod.rs
New accessor methods and a terminal-pruning fix for orphaned turn records.
Contract tests
crates/ironclaw_turns/tests/*.rs
New/updated tests validating row-store durability, eviction, and heartbeat behavior.
Stress tool churn scenario
tools/ironclaw_stress/src/*.rs
TurnLifecycleChurn scenario and FilesystemRow backend added, with stage timing extensions.

RebornBuildInput Postgres constructor

Layer / File(s) Summary
hosted_single_tenant_postgres constructor
crates/ironclaw_reborn_composition/src/input.rs
New Postgres-only builder validating profile and constructing hosted single-tenant storage input.

Hosted single-tenant Postgres latency harness

Layer / File(s) Summary
Goal and spec docs
goal.md, spec.md
Define acceptance criteria, hard-fail conditions, and cycle protocol.
Harness scripts
harness/latency/{score,lint,probe,status}.sh, README.md
Score/lint/probe/status entrypoints and usage docs.
Rust latency runner
harness/latency/runner/**
New crate running storage/control-plane/turn-lifecycle/WebUI/hosted-build workloads and comparing libSQL vs Postgres.

Estimated code review effort: 5 (Critical) | ~150 minutes

Possibly related issues

Possibly related PRs

  • nearai/ironclaw#5313: Adds a stress harness with "reserve-sequence" storage workloads depending on the reserve_sequence API introduced here.
  • nearai/ironclaw#5451: Also modifies connect_with_retry PRAGMA application in crates/ironclaw_filesystem/src/libsql.rs.
  • nearai/ironclaw#5455: Also implements reserve_sequence across Postgres/scoped filesystem plumbing.

Suggested reviewers: henrypark133, ilblackdragon, italic-jinxin

Advisory-xact-lock migration path (postgres.rs run_migrations): verify lock released on early-return error paths — CLAUDE.md transaction-safety invariant, not caught by clippy. Delta journal ack propagation on flusher-thread panic (filesystem_governor.rs): confirm pending oneshot receivers don't hang forever — sandbox/trust concern (silent stall vs. hard fail). PostgresSecretStore AAD derivation from scope/handle: confirm no cross-tenant AAD collision — secrets invariant, needs explicit test, not typechecked. Row store durable-delta suppression for Tier-2 cache evictions: confirm this can't cause silent data loss on process crash before compaction — flag as durability risk, not style.

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning Summary, verification, and journey notes are present, but most required template sections and the trust-boundary checklist are missing. Add the missing template sections: Change Type, Linked Issue, Security Impact, Reborn Trust-Boundary Checklist, Database Impact, Blast Radius, Rollback Plan, Review Follow-Through, and Review track.
✅ Passed checks (3 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed Title is directly related to the PR’s hosted Postgres turn-state latency work, though it is not Conventional Commits style.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added the contributor: core 20+ merged PRs label Jul 5, 2026
@serrrfirat
serrrfirat marked this pull request as draft July 5, 2026 22:14
@serrrfirat serrrfirat changed the title Optimize hosted Postgres turn-state latency [DRAFT] Optimize hosted Postgres turn-state latency Jul 5, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a journaled filesystem-backed resource governor (FilesystemResourceGovernor) and a filesystem turn-state row store (FilesystemTurnStateRowStore) to optimize performance and reduce contention on single-blob state documents. It also adds a comprehensive latency harness to compare libSQL and PostgreSQL backends across several production-shaped workloads. The code review identified three key issues: a correctness bug where historical usage from previous periods is incorrectly restored upon restart or compaction, a race condition in turn_store_for_context that can instantiate duplicate row stores for the same user, and an O(N) read complexity bottleneck in read_delta_log that should be documented for future optimization.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +711 to +714
fn rebuild_tallies_from_reservations(state: &mut ResourceState) {
state.reserved_by_account.clear();
state.usage_by_account.clear();
for record in state.reservations.values() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Correctness Issue: Historical usage from previous periods is incorrectly restored upon restart/compaction

In rebuild_tallies_from_reservations, the tallies for usage_by_account are cleared and rebuilt by iterating over all reservations in state.reservations and adding their usage:

state.usage_by_account.entry(account.clone()).or_default().add_assign(&spent);

However, this does not take into account the period boundaries or whether a reservation was reconciled in a previous, already-rolled-over budget period.

When the governor restarts or performs compaction, rebuild_tallies_from_reservations will blindly add the usage of all historical reconciled reservations to the current period's usage_by_account. Since the current time now is already within the active period, advance_period_if_rolled_over will not trigger a rollover, meaning the old usage from previous periods will incorrectly inflate the current period's tally. This can lead to premature quota exhaustion and false-positive LimitExceeded errors.

Recommendation:
Only include reconciled reservations in the tally if they fall within the current active budget period for that account (i.e., after the start of the current period as defined by the account's period anchor/bounds), or avoid clearing and rebuilding the historical usage_by_account from the raw reservations map if the snapshot already contains the correct rolled-over tallies.

Comment on lines +1385 to +1420
if let Some(store) = self
.row_turn_stores
.lock()
.map_err(|_| {
OperationFailure::new(
"turn_store_lock_poisoned",
"turn_store",
"row turn-store cache lock poisoned",
)
})?
.get(&key)
.cloned()
{
return Ok(store);
}

let view = user_turn_mount_view(&self.run_id, &resource_scope)
.map_err(|error| OperationFailure::invalid_request("turn_store", error))?;
let scoped = Arc::new(ScopedFilesystem::with_fixed_view(
Arc::clone(&self.root),
view,
));
let store = Arc::new(
FilesystemTurnStateRowStore::new(scoped).with_limits(self.turn_state_limits),
) as Arc<dyn StressTurnStore>;
self.row_turn_stores
.lock()
.map_err(|_| {
OperationFailure::new(
"turn_store_lock_poisoned",
"turn_store",
"row turn-store cache lock poisoned",
)
})?
.insert(key, Arc::clone(&store));
Ok(store)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Concurrency Bug: Race condition in turn_store_for_context can create duplicate FilesystemTurnStateRowStore instances for the same user

The TurnStateBackend::FilesystemRow branch of turn_store_for_context releases the lock on self.row_turn_stores after checking for an existing store, and then re-acquires it to insert the newly created store. This creates a classic check-then-act race condition. If two concurrent threads call turn_store_for_context for the same user/tenant, both can see a cache miss, initialize separate FilesystemTurnStateRowStore instances, and insert them into the map (with one overwriting the other).

As a result, the two threads will operate on different in-memory hot caches (snapshot_state), completely breaking the in-process cache consistency and leading to stale reads, lost updates, or transaction conflicts.

Recommendation:
Since creating a FilesystemTurnStateRowStore is extremely fast and does not perform any blocking I/O, you should hold the lock for the entire duration of the check-and-insert operation.

                let mut stores = self
                    .row_turn_stores
                    .lock()
                    .map_err(|_|
                        OperationFailure::new(
                            "turn_store_lock_poisoned",
                            "turn_store",
                            "row turn-store cache lock poisoned",
                        )
                    )?;
                if let Some(store) = stores.get(&key).cloned() {
                    return Ok(store);
                }

                let view = user_turn_mount_view(&self.run_id, &resource_scope)
                    .map_err(|error| OperationFailure::invalid_request("turn_store", error))?;
                let scoped = Arc::new(ScopedFilesystem::with_fixed_view(
                    Arc::clone(&self.root),
                    view,
                ));
                let store = Arc::new(
                    FilesystemTurnStateRowStore::new(scoped).with_limits(self.turn_state_limits),
                ) as Arc<dyn StressTurnStore>;
                stores.insert(key, Arc::clone(&store));
                Ok(store)
References
  1. Ensure atomicity of read-modify-write operations on shared state by holding a lock across the entire sequence (read, validate, and write). All related operations must share the same lock to prevent race conditions and inconsistent interleaving in concurrent environments.

Comment on lines +527 to +532
async fn read_delta_log(&self) -> Result<Vec<SnapshotDelta>, TurnError> {
let path = delta_log_path()?;
let records = match self
.filesystem
.tail(&ResourceScope::system(), &path, SeqNo::ZERO)
.await

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Performance/Scalability Issue: O(N) complexity in read_delta_log due to full scan from SeqNo::ZERO

The read_delta_log method reads the entire delta log from the beginning of time (SeqNo::ZERO):

let records = match self.filesystem.tail(&ResourceScope::system(), &path, SeqNo::ZERO).await

Since there is currently no materialized row compactor or truncation mechanism for the delta log, this log will grow indefinitely. As a result, load_snapshot_from_rows (called at startup or when the cache is cleared) will take progressively longer to complete, directly impacting startup latency.

Recommendation:
Because this snapshot-adapter pattern necessitates loading the full state, we should defer immediate optimization of this read path. Instead, please document the requirement for targeted read paths or a background compaction/checkpointing mechanism as a follow-up task.

References
  1. If a specific design pattern (e.g., snapshot-adapter) necessitates an inefficient implementation (e.g., loading full state), defer optimization and document the requirement for targeted read paths as a follow-up task.

@github-actions

github-actions Bot commented Jul 5, 2026

Copy link
Copy Markdown
Contributor

⚠️ 9 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_embeddings, ironclaw_gateway, ironclaw_hooks, ironclaw_oauth, ironclaw_process_sandbox, ironclaw_prompt_envelope, ironclaw_scripts, ironclaw_skill_learning, ironclaw_tui

Reborn integration-tier coverage

Line coverage (Reborn crates): 28.15% — 49476 / 175731 lines

Per-crate breakdown (62 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_embeddings 0% 0 / 337
ironclaw_gateway 0% 0 / 283
ironclaw_hooks 0% 0 / 4468
ironclaw_oauth 0% 0 / 155
ironclaw_process_sandbox 0% 0 / 795
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 347
ironclaw_skill_learning 0% 0 / 61
ironclaw_tui 0% 0 / 4776
ironclaw_outbound 0.22% 3 / 1339
ironclaw_event_projections 0.4% 6 / 1489
ironclaw_reborn_event_store 0.66% 6 / 913
ironclaw_reborn_config 1.36% 15 / 1101
ironclaw_llm 3.62% 437 / 12075
ironclaw_event_streams 3.87% 40 / 1034
ironclaw_product_adapter_registry 5.38% 25 / 465
ironclaw_extractors 6.18% 26 / 421
ironclaw_wasm_sandbox_core 7.37% 7 / 95
ironclaw_product_workflow 7.97% 768 / 9635
ironclaw_webui_v2 8.5% 228 / 2683
ironclaw_processes 8.61% 98 / 1138
ironclaw_common 10.22% 74 / 724
ironclaw_events 12.45% 143 / 1149
ironclaw_product_adapters 12.53% 280 / 2234
ironclaw_network 13.25% 66 / 498
ironclaw_skills 14.58% 377 / 2585
ironclaw_first_party_extensions 22.46% 1125 / 5010
ironclaw_triggers 23.05% 666 / 2889
ironclaw_reborn_traces 23.24% 1492 / 6420
ironclaw_secrets 25.92% 442 / 1705
ironclaw_reborn 29.03% 2538 / 8742
ironclaw_reborn_composition 30.21% 9231 / 30561
ironclaw_capabilities 32.97% 580 / 1759
ironclaw_auth 33.09% 667 / 2016
ironclaw_runtime_policy 33.2% 80 / 241
ironclaw_memory_native 37.02% 857 / 2315
ironclaw_filesystem 38.71% 1419 / 3666
ironclaw_host_api 39.9% 942 / 2361
ironclaw_host_runtime 41.18% 6170 / 14983
ironclaw_threads 41.65% 1326 / 3184
ironclaw_trust 42.56% 326 / 766
ironclaw_loop_support 42.62% 3142 / 7372
ironclaw_resources 46.01% 1351 / 2936
ironclaw_turns 46.87% 5206 / 11108
ironclaw_memory 47.47% 357 / 752
ironclaw_first_party_extension_ports 48.74% 637 / 1307
ironclaw_wasm 48.79% 363 / 744
ironclaw_projects 50% 147 / 294
ironclaw_extensions 51.4% 1211 / 2356
ironclaw_agent_loop 51.5% 2400 / 4660
ironclaw_run_state 52.73% 222 / 421
ironclaw_authorization 53.54% 461 / 861
ironclaw_safety 59.38% 1035 / 1743
ironclaw_observability 61.54% 16 / 26
ironclaw_conversations 66.13% 937 / 1417
ironclaw_approvals 66.63% 549 / 824
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_mcp 67.42% 569 / 844
ironclaw_reborn_identity 70.91% 156 / 220
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_product_context 78.57% 11 / 14
ironclaw_attachments 84.92% 107 / 126

This signal is informational: coverage never gates the PR — not the percentage, not the per-crate holes, not the 0-coverage callout.

Exemptions (0 file(s) excluded from the accounting above)

No exemptions configured.

@serrrfirat serrrfirat closed this Jul 5, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 13

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_triggers/src/postgres.rs (1)

162-178: 🚀 Performance & Scalability | 🔵 Trivial | ⚡ Quick win

Route get_trigger through the cached query helper
This lookup still calls query_opt directly, so it bypasses prepare_cached while the rest of the read paths use the shared wrapper. Switching it over keeps trigger fetches consistent and avoids an extra prepare on a hot path.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_triggers/src/postgres.rs` around lines 162 - 178, The
get_trigger read path still bypasses the shared cached query wrapper by calling
query_opt directly. Update the get_trigger logic in postgres.rs to route this
lookup through the existing cached query helper used by the other read paths, so
it benefits from prepare_cached and stays consistent with the rest of the
trigger queries. Keep the row_to_record mapping and backend_error handling
intact while swapping the direct query call to the shared helper.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_resources/src/filesystem_governor.rs`:
- Around line 692-706: The restoration path in filesystem_governor.rs is
rebuilding v3 tallies from historical reservations via
rebuild_tallies_from_reservations(), which can resurrect expired spend after
rollover. Update the snapshot recovery logic to keep snapshot.state as-is for v3
and only replay deltas after journal_seq in the delta log processing flow around
rebuild_tallies_from_reservations, delta.apply_to, and the records loop. If
legacy repair still needs rebuilding, isolate it behind a separate explicit path
instead of using it during normal snapshot replay.
- Around line 173-193: The resource-governor mutation paths are publishing
in-memory state before the durable journal append is confirmed, then releasing
serialization too early, which can let replay observe operations out of order.
Update the relevant mutation flows around persist_delta(),
write_accounts_from_state(), and the reserve/reconcile/account snapshot handlers
to stage changes first, append the delta while the affected
accounts/reservations are still locked, and only publish the state after the
append succeeds. If needed, funnel all mutations through a single journaled
command pipeline so ordering is preserved across callers. Add a concurrent
reserve/reconcile/restart test that proves replay cannot see B-before-A.
- Line 25: The background compactor/journal diagnostics in filesystem_governor
should use debug! instead of warn! to avoid breaking the REPL/TUI logging
invariant. Update the tracing::warn import and replace the warn! calls in the
internal background paths with debug! in the filesystem governor routines that
emit these messages, including the affected compactor/journal code paths.
- Around line 694-699: Treat FilesystemError::Unsupported in the filesystem.tail
handling as a hard failure instead of returning an empty Vec, because delta-log
tailing cannot be proven safe in that case. Update the match in the tail path to
propagate Unsupported through fs_error(error) like other unexpected errors,
while keeping the NotFound case as an empty result. Use the existing
filesystem.tail and fs_error flow in filesystem_governor to locate the branch
and preserve the current error context propagation.

In `@crates/ironclaw_secrets/src/postgres_store.rs`:
- Around line 691-695: The error mapping in `postgres_error` (and the matching
`serde_to_store_error` path) is leaking raw backend `Display` text into
`SecretStoreError::StoreUnavailable.reason`. Update the fallible paths in
`postgres_store` so they log the original `tokio_postgres::Error` /
`serde_json::Error` internally at debug/error level, but return a fixed
sanitized reason string without backend details. Apply the same pattern to the
store operations that call these helpers (`put`, `read`, `delete`, `lease_once`,
`consume`, `revoke`, `migrate`) and check `FilesystemSecretStore`’s
`fs_to_secret_store_error` for the same treatment.
- Around line 20-24: Extend the secret-store contract coverage to
PostgresSecretStore by adding caller-level tests that exercise the persistence
side effects in lease_once, consume, and revoke. Use the PostgresSecretStore
type and its related methods to verify FOR UPDATE contention behavior and expiry
write-back against a real Postgres-backed setup, following the contract-test
style already expected for secret stores.

In `@crates/ironclaw_threads/src/filesystem_service.rs`:
- Around line 516-531: The thread record update in `FilesystemService` should
not fail immediately on `VersionMismatch` from the `txn.put` for `thread_entry`;
this is a CAS conflict that can happen when another writer advances
`thread.json`. Update the write path to treat this as a retryable conflict in
the same message-write flow, so the outer retry loop re-reads and re-applies the
legacy-counter update instead of returning an error. Use the
`stored.next_sequence` / `thread_version` update block in `FilesystemService` as
the place to handle this conflict before calling `absent_put_error`.

In `@crates/ironclaw_turns/src/filesystem_store/row_store.rs`:
- Around line 744-822: The `apply` flow in `row_store.rs` publishes the updated
snapshot cache before `await_delta_ack` confirms the durable write, allowing a
later caller to build on unconfirmed state. Update `apply` (and the matching
`apply_with_targeted_delta` path referenced in the review) so
`snapshot_state`/`guard` remains protected until the ack succeeds, or add a
fencing/version validation that prevents a write from committing on a baseline
that was not durably confirmed. Ensure the cache is only updated after the
durable delta is acknowledged and clear the cache on any ack failure.
- Around line 391-404: get_run_state and get_loop_checkpoint are still doing
linear scans over the Vec-backed TurnPersistenceSnapshot instead of using the
existing O(1) indexes. Update the lookup path used by
projection::run_state_parts and projection::loop_checkpoint so it reads through
RowSnapshotState/RowSnapshotIndexes or the InMemoryTurnStateStore accessors
(run_record/turn_record) rather than relying on with_cached_snapshot alone. Keep
with_cached_snapshot as the cache entry point, but change the callers in the
read path to use the indexed lookup helpers already exercised by
submit_turn_targeted_delta and apply_with_targeted_delta’s
RunnerLeaseOverlay::Run branch.
- Around line 410-461: The delta-log replay in
load_snapshot_from_rows/replay_deltas is unbounded because it always tails from
SeqNo::ZERO. Update the replay path to start from the current retention floor or
the latest durable checkpoint metadata instead of the beginning, using the
existing RowSnapshotState/TurnPersistenceSnapshot flow to seed the starting
SeqNo. Also make sure read_delta_log, read_run_state_from_durable_rows, and
read_turn_events_from_durable_rows use the same bounded starting point so
durable reads do not scan the full history.

In `@crates/ironclaw_turns/src/memory/mod.rs`:
- Around line 479-506: overlay_runner_lease_record duplicates the same
lease-fencing and stale-heartbeat checks already used by
runner_lease::apply_runner_lease_overlay and run_can_use_external_lease. Extract
a shared predicate/helper over the common record fields (status, runner_id,
lease_token, last_heartbeat_at) and have overlay_runner_lease_record call it
instead of re-implementing the logic, so both paths stay in sync.

In `@harness/latency/lint.sh`:
- Around line 24-33: The temporary file handling in the latency lint script is
using a predictable, shell-PID-based path, which is racy and unsafe. Update the
logic in the lint script to create the scratch file with mktemp, store that path
in a variable, and use the same variable for both rg reads and cleanup. Keep the
rest of the flow intact around the rg checks and the final rm, but ensure the
temp file name is unpredictable and atomically created.

In `@tools/ironclaw_stress/src/user_turn.rs`:
- Around line 757-820: The TurnLifecycleChurn branch in user_turn.rs is
discarding the ClaimedTurnRun returned by claim_next_run() and then calling
complete_run() with the submit-time run_id instead of the claimed state. Update
this path to capture the claimed value from turn_store.claim_next_run in the
existing time_stage call, and use claimed.state.run_id, claimed.runner_id, and
claimed.lease_token when completing the run, matching the other call sites and
preserving the actual lease ownership.

---

Outside diff comments:
In `@crates/ironclaw_triggers/src/postgres.rs`:
- Around line 162-178: The get_trigger read path still bypasses the shared
cached query wrapper by calling query_opt directly. Update the get_trigger logic
in postgres.rs to route this lookup through the existing cached query helper
used by the other read paths, so it benefits from prepare_cached and stays
consistent with the rest of the trigger queries. Keep the row_to_record mapping
and backend_error handling intact while swapping the direct query call to the
shared helper.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 30de446b-ce32-48f9-8ec9-fed1a0fe80af

📥 Commits

Reviewing files that changed from the base of the PR and between 771c1fe and fab13d8.

⛔ Files ignored due to path filters (2)
  • Cargo.lock is excluded by !**/*.lock, !**/Cargo.lock
  • harness/latency/runner/Cargo.lock is excluded by !**/*.lock, !**/Cargo.lock
📒 Files selected for processing (56)
  • LOG.md
  • crates/ironclaw_filesystem/src/backend.rs
  • crates/ironclaw_filesystem/src/libsql.rs
  • crates/ironclaw_filesystem/src/postgres.rs
  • crates/ironclaw_filesystem/src/scoped.rs
  • crates/ironclaw_host_runtime/src/services.rs
  • crates/ironclaw_host_runtime/src/services/builder.rs
  • crates/ironclaw_reborn_composition/src/factory.rs
  • crates/ironclaw_reborn_composition/src/input.rs
  • crates/ironclaw_reborn_composition/src/lib.rs
  • crates/ironclaw_reborn_composition/src/observability/budget.rs
  • crates/ironclaw_reborn_composition/src/outbound/mod.rs
  • crates/ironclaw_reborn_event_store/src/lib.rs
  • crates/ironclaw_resources/Cargo.toml
  • crates/ironclaw_resources/src/cas_snapshot.rs
  • crates/ironclaw_resources/src/filesystem_governor.rs
  • crates/ironclaw_resources/src/filesystem_store.rs
  • crates/ironclaw_resources/src/lib.rs
  • crates/ironclaw_resources/tests/resource_governor_contract.rs
  • crates/ironclaw_secrets/Cargo.toml
  • crates/ironclaw_secrets/src/filesystem_store.rs
  • crates/ironclaw_secrets/src/lib.rs
  • crates/ironclaw_secrets/src/postgres_store.rs
  • crates/ironclaw_threads/src/filesystem_service.rs
  • crates/ironclaw_triggers/src/libsql.rs
  • crates/ironclaw_triggers/src/postgres.rs
  • crates/ironclaw_turns/src/filesystem_store.rs
  • crates/ironclaw_turns/src/filesystem_store/projection.rs
  • crates/ironclaw_turns/src/filesystem_store/row_store.rs
  • crates/ironclaw_turns/src/filesystem_store/runner_lease.rs
  • crates/ironclaw_turns/src/lib.rs
  • crates/ironclaw_turns/src/memory/mod.rs
  • crates/ironclaw_turns/tests/filesystem_turn_state_contract.rs
  • crates/ironclaw_turns/tests/loop_checkpoint_store_contract.rs
  • docs/reborn/contracts/resources.md
  • goal.md
  • harness/latency/README.md
  • harness/latency/lint.sh
  • harness/latency/probe.sh
  • harness/latency/runner/.gitignore
  • harness/latency/runner/Cargo.toml
  • harness/latency/runner/src/main.rs
  • harness/latency/score.sh
  • harness/latency/status.sh
  • spec.md
  • tools/ironclaw_stress/Cargo.toml
  • tools/ironclaw_stress/README.md
  • tools/ironclaw_stress/src/analysis.rs
  • tools/ironclaw_stress/src/human.rs
  • tools/ironclaw_stress/src/main.rs
  • tools/ironclaw_stress/src/process_pressure.rs
  • tools/ironclaw_stress/src/report.rs
  • tools/ironclaw_stress/src/resource_ops.rs
  • tools/ironclaw_stress/src/summary.rs
  • tools/ironclaw_stress/src/tests.rs
  • tools/ironclaw_stress/src/user_turn.rs

use ironclaw_filesystem::{FilesystemError, RootFilesystem, ScopedFilesystem, SeqNo};
use ironclaw_host_api::{ReservationStatus, ResourceReservationId, ResourceScope, ScopedPath};
use serde::{Deserialize, Serialize};
use tracing::warn;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Use debug! for background diagnostics.

These compactor/journal messages are internal background diagnostics; warn! can corrupt REPL/TUI paths under the repo logging invariant.

Proposed fix
-use tracing::warn;
+use tracing::debug;
...
-                    warn!(reason = %error, "resource governor compaction write failed");
+                    debug!(reason = %error, "resource governor compaction write failed");
...
-            warn!(reason = %error, "resource governor delta journal thread failed to start");
+            debug!(reason = %error, "resource governor delta journal thread failed to start");

As per path instructions, “REPL/TUI logging: info!/warn! corrupt the terminal UI — internal diagnostics use debug!.”

Also applies to: 145-155, 559-563

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_resources/src/filesystem_governor.rs` at line 25, The
background compactor/journal diagnostics in filesystem_governor should use
debug! instead of warn! to avoid breaking the REPL/TUI logging invariant. Update
the tracing::warn import and replace the warn! calls in the internal background
paths with debug! in the filesystem governor routines that emit these messages,
including the affected compactor/journal code paths.

Source: Path instructions

Comment on lines +173 to +193
let (tally, changed) = {
let mut locked = authority.lock_accounts(std::slice::from_ref(account))?;
let before = locked.account_parts(account);
let mut state =
locked.state_for_accounts(std::slice::from_ref(account), HashMap::new());
advance_period_if_rolled_over(&mut state, account, now);
let tally = state
.reserved_by_account
.get(account)
.cloned()
.unwrap_or_default();
locked.write_accounts_from_state(std::slice::from_ref(account), &state);
let after = locked.account_parts(account);
(tally, before != after)
};
if changed {
let delta = ResourceGovernorDelta::AccountSnapshot {
account: account.clone(),
at: now,
};
if let Err(error) = self.persist_delta(&authority, delta) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🔴 Critical | 🏗️ Heavy lift

Serialize authority mutation with the durable journal append.

These paths publish in-memory changes before the delta append is acknowledged, then release the account/reservation locks before persist_delta(). A second operation can observe A’s state, append B first, and make restart replay see B-before-A—for example Reconcile before the corresponding Reserve.

Stage the state, append while the affected accounts/reservation are still serialized, then publish the state only after the append succeeds; or route all mutations through one journaled command pipeline. Add a concurrent reserve/reconcile/restart caller test. This violates the resource-governor fail-closed durability contract documented for production persistence.

Also applies to: 204-224, 244-257, 282-323, 345-383, 400-437, 454-469

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_resources/src/filesystem_governor.rs` around lines 173 - 193,
The resource-governor mutation paths are publishing in-memory state before the
durable journal append is confirmed, then releasing serialization too early,
which can let replay observe operations out of order. Update the relevant
mutation flows around persist_delta(), write_accounts_from_state(), and the
reserve/reconcile/account snapshot handlers to stage changes first, append the
delta while the affected accounts/reservations are still locked, and only
publish the state after the append succeeds. If needed, funnel all mutations
through a single journaled command pipeline so ordering is preserved across
callers. Add a concurrent reserve/reconcile/restart test that proves replay
cannot see B-before-A.

Comment on lines +692 to +706
rebuild_tallies_from_reservations(&mut state);
let path = delta_log_path()?;
let records = match filesystem.tail(&ResourceScope::system(), &path, from).await {
Ok(records) => records,
Err(FilesystemError::NotFound { .. }) | Err(FilesystemError::Unsupported { .. }) => {
Vec::new()
}
Err(error) => return Err(fs_error(error)),
};
let mut latest = from;
for record in records {
latest = record.seq;
let delta: ResourceGovernorDelta = serde_json::from_slice(&record.payload)
.map_err(|error| storage_error(format!("decode resource governor delta: {error}")))?;
delta.apply_to(&mut state)?;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Do not rebuild v3 period ledgers from historical reservations.

rebuild_tallies_from_reservations() clears the snapshot’s authoritative ledgers and re-adds every reconciled reservation. After a periodic budget rolls over and compaction stores the cleared ledger with an advanced cursor, restart will resurrect old spend because ReservationRecord has no period/reconcile timestamp to filter by.

For v3 snapshots, replay from snapshot.state as-is and only apply deltas after journal_seq; reserve rebuilds for an explicit legacy repair path if still needed.

Suggested direction
 async fn replay_journal<F>(
     filesystem: Arc<ScopedFilesystem<F>>,
     mut state: ResourceState,
     from: SeqNo,
 ) -> Result<(ResourceState, SeqNo), ResourceError>
 where
     F: RootFilesystem,
 {
-    rebuild_tallies_from_reservations(&mut state);
     let path = delta_log_path()?;

Also applies to: 711-736

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_resources/src/filesystem_governor.rs` around lines 692 - 706,
The restoration path in filesystem_governor.rs is rebuilding v3 tallies from
historical reservations via rebuild_tallies_from_reservations(), which can
resurrect expired spend after rollover. Update the snapshot recovery logic to
keep snapshot.state as-is for v3 and only replay deltas after journal_seq in the
delta log processing flow around rebuild_tallies_from_reservations,
delta.apply_to, and the records loop. If legacy repair still needs rebuilding,
isolate it behind a separate explicit path instead of using it during normal
snapshot replay.

Comment on lines +694 to +699
let records = match filesystem.tail(&ResourceScope::system(), &path, from).await {
Ok(records) => records,
Err(FilesystemError::NotFound { .. }) | Err(FilesystemError::Unsupported { .. }) => {
Vec::new()
}
Err(error) => return Err(fs_error(error)),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Fail closed when delta-log tail is unsupported.

Unsupported means recovery cannot prove whether deltas after the snapshot cursor exist. Treating it as an empty log can admit costed work from stale quota state.

Proposed fix
     let records = match filesystem.tail(&ResourceScope::system(), &path, from).await {
         Ok(records) => records,
-        Err(FilesystemError::NotFound { .. }) | Err(FilesystemError::Unsupported { .. }) => {
-            Vec::new()
-        }
+        Err(FilesystemError::NotFound { .. }) => Vec::new(),
         Err(error) => return Err(fs_error(error)),
     };

As per path instructions, “Fail loud: flag silent-failure patterns … Errors propagate with ? into thiserror types with context.”

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
let records = match filesystem.tail(&ResourceScope::system(), &path, from).await {
Ok(records) => records,
Err(FilesystemError::NotFound { .. }) | Err(FilesystemError::Unsupported { .. }) => {
Vec::new()
}
Err(error) => return Err(fs_error(error)),
let records = match filesystem.tail(&ResourceScope::system(), &path, from).await {
Ok(records) => records,
Err(FilesystemError::NotFound { .. }) => Vec::new(),
Err(error) => return Err(fs_error(error)),
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_resources/src/filesystem_governor.rs` around lines 694 - 699,
Treat FilesystemError::Unsupported in the filesystem.tail handling as a hard
failure instead of returning an empty Vec, because delta-log tailing cannot be
proven safe in that case. Update the match in the tail path to propagate
Unsupported through fs_error(error) like other unexpected errors, while keeping
the NotFound case as an empty result. Use the existing filesystem.tail and
fs_error flow in filesystem_governor to locate the branch and preserve the
current error context propagation.

Source: Path instructions

Comment on lines +20 to +24
pub struct PostgresSecretStore {
pool: Pool,
crypto: Arc<SecretsCrypto>,
lease_ttl: Duration,
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚀 Performance & Scalability | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
fd . crates/ironclaw_secrets -e rs | xargs rg -ln "PostgresSecretStore"

Repository: nearai/ironclaw

Length of output: 234


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '== files ==\n'
git ls-files crates/ironclaw_secrets | sed 's#^`#-` #'

printf '\n== outline: postgres_store.rs ==\n'
ast-grep outline crates/ironclaw_secrets/src/postgres_store.rs --view expanded || true

printf '\n== outline: lib.rs ==\n'
ast-grep outline crates/ironclaw_secrets/src/lib.rs --view expanded || true

printf '\n== tests mentioning PostgresSecretStore ==\n'
rg -n "PostgresSecretStore|SecretStore|lease|FOR UPDATE|expiry" crates/ironclaw_secrets -g '*test*' -g '*.rs' || true

printf '\n== file sizes ==\n'
wc -l crates/ironclaw_secrets/src/postgres_store.rs crates/ironclaw_secrets/src/lib.rs

Repository: nearai/ironclaw

Length of output: 50371


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '== postgres store references in tests ==\n'
rg -n "PostgresSecretStore|run_migrations|with_lease_ttl|lease_once\\(|consume\\(|revoke\\(" crates/ironclaw_secrets/tests crates/ironclaw_secrets/src -g '*.rs' || true

printf '\n== secret_store_contract.rs (relevant sections) ==\n'
sed -n '1,260p' crates/ironclaw_secrets/tests/secret_store_contract.rs

printf '\n== boundary_contract.rs (relevant sections) ==\n'
sed -n '1,220p' crates/ironclaw_secrets/tests/boundary_contract.rs

printf '\n== AGENTS and CLAUDE for crates/ironclaw_secrets ==\n'
sed -n '1,220p' crates/ironclaw_secrets/AGENTS.md
printf '\n---\n'
sed -n '1,220p' crates/ironclaw_secrets/CLAUDE.md

Repository: nearai/ironclaw

Length of output: 23406


Add Postgres-backed contract tests
crates/ironclaw_secrets/AGENTS.md asks for caller-level tests when a helper gates persistence side effects. Extend the secret-store contract coverage to PostgresSecretStore, especially lease_once/consume/revoke, FOR UPDATE contention, and expiry write-back.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_secrets/src/postgres_store.rs` around lines 20 - 24, Extend
the secret-store contract coverage to PostgresSecretStore by adding caller-level
tests that exercise the persistence side effects in lease_once, consume, and
revoke. Use the PostgresSecretStore type and its related methods to verify FOR
UPDATE contention behavior and expiry write-back against a real Postgres-backed
setup, following the contract-test style already expected for secret stores.

Source: Path instructions

Comment on lines +410 to +461
async fn load_snapshot_from_rows(&self) -> Result<RowSnapshotState, TurnError> {
let meta = self.read_meta().await?;
let turns = self.read_row_collection(TURN_ROWS).await?;
let runs = self.read_row_collection(RUN_ROWS).await?;
let active_locks = self.read_row_collection(ACTIVE_LOCK_ROWS).await?;
let checkpoints = self.read_row_collection(CHECKPOINT_ROWS).await?;
let loop_checkpoints = self.read_row_collection(LOOP_CHECKPOINT_ROWS).await?;
let idempotency_records = self.read_row_collection(IDEMPOTENCY_ROWS).await?;
let events = self.read_row_collection(EVENT_ROWS).await?;
let admission_reservations = self.read_row_collection(ADMISSION_RESERVATION_ROWS).await?;
let spawn_tree_reservations = self
.read_row_collection(SPAWN_TREE_RESERVATION_ROWS)
.await?;

let mut snapshot = TurnPersistenceSnapshot {
turns,
runs,
active_locks,
checkpoints,
loop_checkpoints,
idempotency_records,
events,
event_retention_floor: meta.event_retention_floor,
admission_reservations,
spawn_tree_reservations,
};
self.replay_deltas(&mut snapshot).await?;
let snapshot = row_store_hot_cache_snapshot(snapshot, self.limits);
let store = self.build_in_memory_store(snapshot)?;
let snapshot = store.persistence_snapshot();
RowSnapshotState::new(snapshot, Arc::new(store))
}

async fn replay_deltas(&self, snapshot: &mut TurnPersistenceSnapshot) -> Result<(), TurnError> {
let path = delta_log_path()?;
let records = match self
.filesystem
.tail(&ResourceScope::system(), &path, SeqNo::ZERO)
.await
{
Ok(records) => records,
Err(FilesystemError::NotFound { .. }) | Err(FilesystemError::Unsupported { .. }) => {
Vec::new()
}
Err(error) => return Err(fs_error(error)),
};
for record in records {
let delta: SnapshotDelta = deserialize_row(&record.payload, "turn-state delta")?;
apply_delta(snapshot, delta)?;
}
Ok(())
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚀 Performance & Scalability | 🟠 Major | 🏗️ Heavy lift

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Map the file and inspect the relevant regions first.
ast-grep outline crates/ironclaw_turns/src/filesystem_store/row_store.rs --view expanded || true

echo "---- lines around load/replay ----"
sed -n '360,560p' crates/ironclaw_turns/src/filesystem_store/row_store.rs | cat -n

echo "---- search for meta_path writes and delta-log truncation/compaction ----"
rg -n "meta_path\\(|event_retention_floor|write_meta|read_meta|delta_log|tail\\(&ResourceScope::system\\(\\), &path, SeqNo::ZERO|SeqNo::ZERO|truncate|compact|retention_floor" crates/ironclaw_turns/src/filesystem_store/row_store.rs crates/ironclaw_turns/src/filesystem_store -S

Repository: nearai/ironclaw

Length of output: 28437


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "---- write path around the delta journal ----"
sed -n '240,360p' crates/ironclaw_turns/src/filesystem_store/row_store.rs | cat -n
sed -n '820,900p' crates/ironclaw_turns/src/filesystem_store/row_store.rs | cat -n

echo "---- delta-log / meta-path writes across the crate ----"
rg -n "meta_path\\(|delta_log_path\\(|META_FILE|DELTA_LOG|write\\(|append\\(|truncate\\(|compact\\(|remove\\(|delete\\(" crates/ironclaw_turns/src/filesystem_store -S

echo "---- broader repo search for meta_path writes ----"
rg -n "meta_path\\(" crates src tests -S

Repository: nearai/ironclaw

Length of output: 11414


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "---- durable-delta shaping and retention-floor handling ----"
sed -n '2128,2410p' crates/ironclaw_turns/src/filesystem_store/row_store.rs | cat -n

echo "---- durable query paths ----"
sed -n '540,640p' crates/ironclaw_turns/src/filesystem_store/row_store.rs | cat -n

echo "---- all row_store references to delta-log / meta write APIs ----"
rg -n "\.append_batch\(|\.append\(|\.truncate\(|\.delete\(|\.remove\(|meta_path\(|delta_log_path\(|read_meta\(|read_delta_log\(" crates/ironclaw_turns/src/filesystem_store/row_store.rs -S

Repository: nearai/ironclaw

Length of output: 16914


Delta-log replay is unbounded on cold load and durable reads
load_snapshot_from_rows, read_delta_log, read_run_state_from_durable_rows, and read_turn_events_from_durable_rows all tail deltas/log from SeqNo::ZERO. row_store_durable_delta() strips deletes and retention-floor updates, and meta_path() is never written in this file, so replay cost grows with the full write history of the store.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_turns/src/filesystem_store/row_store.rs` around lines 410 -
461, The delta-log replay in load_snapshot_from_rows/replay_deltas is unbounded
because it always tails from SeqNo::ZERO. Update the replay path to start from
the current retention floor or the latest durable checkpoint metadata instead of
the beginning, using the existing RowSnapshotState/TurnPersistenceSnapshot flow
to seed the starting SeqNo. Also make sure read_delta_log,
read_run_state_from_durable_rows, and read_turn_events_from_durable_rows use the
same bounded starting point so durable reads do not scan the full history.

Comment on lines +744 to +822
async fn apply<T, A, Fut>(
&self,
overlay: RunnerLeaseOverlay,
mut apply: A,
) -> Result<T, TurnError>
where
A: FnMut(Arc<InMemoryTurnStateStore>) -> Fut + Send,
Fut: std::future::Future<Output = Result<T, TurnError>> + Send,
T: Send,
{
let critical = async {
let mut guard = self.snapshot_state.lock().await;
if guard.is_none() {
*guard = Some(self.load_snapshot_from_rows().await?);
}
let store = match (overlay, guard.as_ref()) {
(RunnerLeaseOverlay::None, Some(state)) => Arc::clone(&state.store),
(_, Some(state)) => {
let snapshot = state.snapshot.clone();
let (overlaid_snapshot, _) = self
.runner_lease_store()
.overlay((snapshot, None), overlay)
.await?;
Arc::new(self.build_in_memory_store(overlaid_snapshot)?)
}
(_, None) => unreachable!("row snapshot cache is initialized above"),
};
let baseline = guard
.as_ref()
.map(|state| state.snapshot.clone())
.unwrap_or_default();
let outcome = apply(Arc::clone(&store)).await;
let mut new_snapshot = store.persistence_snapshot();
preserve_loop_checkpoints(&baseline, &mut new_snapshot);
let value = match outcome {
Ok(value) => value,
Err(error) => {
*guard = None;
return Err(error);
}
};
if new_snapshot == baseline {
return Ok((None, value));
}

let delta = match snapshot_delta(&baseline, &new_snapshot) {
Ok(delta) => delta,
Err(RowPersistError::Turn(error)) => {
*guard = None;
return Err(error);
}
};
let persist_delta = row_store_durable_delta(delta);
let ack = match self.enqueue_delta(persist_delta) {
Ok(ack) => ack,
Err(RowPersistError::Turn(error)) => {
*guard = None;
return Err(error);
}
};
*guard = Some(RowSnapshotState::new(new_snapshot, store)?);
Ok((ack, value))
};

let (ack, value) = match tokio::time::timeout(self.apply_timeout, critical).await {
Ok(result) => result?,
Err(_) => {
self.clear_snapshot_cache().await;
return Err(TurnError::Unavailable {
reason: "turn state row-store apply timed out".to_string(),
});
}
};
if let Err(error) = self.await_delta_ack(ack).await {
self.clear_snapshot_cache().await;
return Err(error.into_turn());
}
Ok(value)
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🔴 Critical | 🏗️ Heavy lift

Cache is published before the durable delta is acknowledged — a concurrent writer can build on unconfirmed state.

In both apply and apply_with_targeted_delta, guard (the snapshot_state lock) is local to the critical async block and is dropped when critical returns — i.e. *guard = Some(new_snapshot) (or state.store = store) is visible to the next caller before self.await_delta_ack(ack).await runs. If a second write (B) acquires the lock in that window, it computes its delta against A's not-yet-durable snapshot. If A's delta later fails to persist, A calls clear_snapshot_cache(), but B's delta — computed on top of A's phantom state — may already be enqueued/acked in a separate journal batch and persisted durably. The append log then contains B without the A it logically depends on: permanent corruption on the next replay, and a direct violation of goal.md's own hard-fail bar ("skipped durable writes, lost state transitions"). This is squarely in the blast radius of the new turn-lifecycle-churn stress scenario (concurrent submit/claim/complete against one shared store instance).

Hold the lock across the ack (accepting serialized durable writes per store), or add a fencing/version check that rejects a write whose baseline was never confirmed durable.

🔒 Sketch: keep the guard held until the ack resolves
-        let (ack, value) = match tokio::time::timeout(self.apply_timeout, critical).await {
-            Ok(result) => result?,
-            Err(_) => {
-                self.clear_snapshot_cache().await;
-                return Err(TurnError::Unavailable {
-                    reason: "turn state row-store apply timed out".to_string(),
-                });
-            }
-        };
-        if let Err(error) = self.await_delta_ack(ack).await {
-            self.clear_snapshot_cache().await;
-            return Err(error.into_turn());
-        }
-        Ok(value)
+        // Move the ack-await inside `critical`, still holding `guard`, so no
+        // other writer can observe `new_snapshot` before it is durably
+        // confirmed. On ack failure, do not commit `*guard` at all.

Also applies to: 852-940

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_turns/src/filesystem_store/row_store.rs` around lines 744 -
822, The `apply` flow in `row_store.rs` publishes the updated snapshot cache
before `await_delta_ack` confirms the durable write, allowing a later caller to
build on unconfirmed state. Update `apply` (and the matching
`apply_with_targeted_delta` path referenced in the review) so
`snapshot_state`/`guard` remains protected until the ack succeeds, or add a
fencing/version validation that prevents a write from committing on a baseline
that was not durably confirmed. Ensure the cache is only updated after the
durable delta is acknowledged and clear the cache on any ack failure.

Comment on lines +479 to +506
pub(crate) fn overlay_runner_lease_record(
&self,
overlaid: TurnRunRecord,
) -> Result<(), TurnError> {
let mut inner = self.lock_inner()?;
let Some(record) = inner.records.get_mut(&overlaid.run_id) else {
return Err(TurnError::ScopeNotFound);
};
if !matches!(
record.status.get(),
TurnStatus::Running | TurnStatus::CancelRequested
) || record.runner_id != overlaid.runner_id
|| record.lease_token != overlaid.lease_token
|| record.runner_id.is_none()
|| record.lease_token.is_none()
{
return Ok(());
}
if let (Some(current), Some(incoming)) =
(record.last_heartbeat_at, overlaid.last_heartbeat_at)
&& incoming < current
{
return Ok(());
}
record.last_heartbeat_at = overlaid.last_heartbeat_at;
record.lease_expires_at = overlaid.lease_expires_at;
Ok(())
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicates runner_lease.rs::apply_runner_lease_overlay staleness logic.

overlay_runner_lease_record re-implements the same "status must be Running/CancelRequested + runner/lease-token match + reject stale heartbeat" check that already lives in runner_lease::apply_runner_lease_overlay / run_can_use_external_lease. Two independent copies of a concurrency-safety fencing check will drift silently. Extract a shared predicate (operating on the common fields) both sites call.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_turns/src/memory/mod.rs` around lines 479 - 506,
overlay_runner_lease_record duplicates the same lease-fencing and
stale-heartbeat checks already used by runner_lease::apply_runner_lease_overlay
and run_can_use_external_lease. Extract a shared predicate/helper over the
common record fields (status, runner_id, lease_token, last_heartbeat_at) and
have overlay_runner_lease_record call it instead of re-implementing the logic,
so both paths stay in sync.

Comment thread harness/latency/lint.sh
Comment on lines +24 to +33
if rg -n "LATENCY_|latency|benchmark|bench" crates src \
-g '*.rs' >/tmp/ironclaw-latency-lint.$$ 2>/dev/null; then
if rg -n "sleep|tokio::time::sleep|std::thread::sleep|mock readiness|fast path|fast-path" \
/tmp/ironclaw-latency-lint.$$ >/dev/null 2>&1; then
rm -f /tmp/ironclaw-latency-lint.$$
echo "VOID: constraint violation"
exit 1
fi
fi
rm -f /tmp/ironclaw-latency-lint.$$

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

Predictable temp file path — use mktemp.

/tmp/ironclaw-latency-lint.$$ is guessable/racy; a local attacker can pre-create/symlink it. Use mktemp for an atomically-created, unpredictable path.

🔒️ Proposed fix
+tmpfile="$(mktemp)"
+trap 'rm -f "$tmpfile"' EXIT
 if rg -n "LATENCY_|latency|benchmark|bench" crates src \
-  -g '*.rs' >/tmp/ironclaw-latency-lint.$$ 2>/dev/null; then
+  -g '*.rs' >"$tmpfile" 2>/dev/null; then
   if rg -n "sleep|tokio::time::sleep|std::thread::sleep|mock readiness|fast path|fast-path" \
-    /tmp/ironclaw-latency-lint.$$ >/dev/null 2>&1; then
-    rm -f /tmp/ironclaw-latency-lint.$$
+    "$tmpfile" >/dev/null 2>&1; then
     echo "VOID: constraint violation"
     exit 1
   fi
 fi
-rm -f /tmp/ironclaw-latency-lint.$$
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if rg -n "LATENCY_|latency|benchmark|bench" crates src \
-g '*.rs' >/tmp/ironclaw-latency-lint.$$ 2>/dev/null; then
if rg -n "sleep|tokio::time::sleep|std::thread::sleep|mock readiness|fast path|fast-path" \
/tmp/ironclaw-latency-lint.$$ >/dev/null 2>&1; then
rm -f /tmp/ironclaw-latency-lint.$$
echo "VOID: constraint violation"
exit 1
fi
fi
rm -f /tmp/ironclaw-latency-lint.$$
tmpfile="$(mktemp)"
trap 'rm -f "$tmpfile"' EXIT
if rg -n "LATENCY_|latency|benchmark|bench" crates src \
-g '*.rs' >"$tmpfile" 2>/dev/null; then
if rg -n "sleep|tokio::time::sleep|std::thread::sleep|mock readiness|fast path|fast-path" \
"$tmpfile" >/dev/null 2>&1; then
echo "VOID: constraint violation"
exit 1
fi
fi
🧰 Tools
🪛 ast-grep (0.44.1)

[warning] 24-24: Building a temp file path in a world-writable directory from the PID ($$) or `` is predictable and racy: an attacker can pre-create or guess the name and win a symlink/race attack. Use mktemp (e.g. `f=$(mktemp)` or `f=$(mktemp /tmp/myapp.XXXXXX)`) so the kernel atomically creates a unique, unpredictable file.
Context: /tmp/ironclaw-latency-lint.$$
Note: [CWE-377] Insecure Temporary File.

(tmp-file-pid-name-bash)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@harness/latency/lint.sh` around lines 24 - 33, The temporary file handling in
the latency lint script is using a predictable, shell-PID-based path, which is
racy and unsafe. Update the logic in the lint script to create the scratch file
with mktemp, store that path in a variable, and use the same variable for both
rg reads and cleanup. Keep the rest of the flow intact around the rg checks and
the final rm, but ensure the temp file name is unpredictable and atomically
created.

Source: Linters/SAST tools

Comment on lines +757 to +820
if matches!(args.scenario, Scenario::TurnLifecycleChurn) {
let operation_ref = turn_operation_ref(args, worker_index, operation_index, 0, 1);
let SubmitTurnResponse::Accepted { run_id, .. } = time_stage(
&mut stages.submit_turn,
turn_coordinator.submit_turn(SubmitTurnRequest {
scope: context.turn_scope.clone(),
actor: TurnActor::new(context.user_id.clone()),
accepted_message_ref: AcceptedMessageRef::new(format!(
"message:{operation_ref}"
))
.map_err(|error| OperationFailure::invalid_request("submit_turn", error))?,
source_binding_ref: SourceBindingRef::new(source_binding)
.map_err(|error| OperationFailure::invalid_request("submit_turn", error))?,
reply_target_binding_ref: ReplyTargetBindingRef::new(reply_target)
.map_err(|error| OperationFailure::invalid_request("submit_turn", error))?,
requested_run_profile: None,
idempotency_key: IdempotencyKey::new(format!(
"ironclaw-stress:{operation_ref}"
))
.map_err(|error| OperationFailure::invalid_request("submit_turn", error))?,
received_at: Utc::now(),
requested_run_id: None,
parent_run_id: None,
subagent_depth: 0,
spawn_tree_root_run_id: None,
product_context: None,
}),
)
.await
.map_err(|error| turn_failure("submit_turn", error))?;

let runner_id = TurnRunnerId::new();
let lease_token = TurnLeaseToken::new();
time_stage(
&mut stages.claim_run,
turn_store.claim_next_run(ClaimRunRequest {
runner_id,
lease_token,
scope_filter: Some(context.turn_scope.clone()),
}),
)
.await
.map_err(|error| turn_failure("claim_run", error))?
.ok_or_else(|| {
OperationFailure::new(
"turn_claim_miss",
"claim_run",
"submitted run was not claimable",
)
})?;

time_stage(
&mut stages.complete_run,
turn_store.complete_run(CompleteRunRequest {
run_id,
runner_id,
lease_token,
}),
)
.await
.map_err(|error| turn_failure("complete_run", error))?;

return Ok(());
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '## user_turn.rs around TurnLifecycleChurn\n'
sed -n '720,840p' tools/ironclaw_stress/src/user_turn.rs

printf '\n## Find claim/complete request and response types\n'
rg -n "struct ClaimRunRequest|struct CompleteRunRequest|enum SubmitTurnResponse|struct ClaimedTurnRun|claim_next_run|complete_run" -S crates tools -g '!target'

printf '\n## Comparable call sites in user_turn.rs\n'
rg -n "claimed\.state\.run_id|claimed\.runner_id|claimed\.lease_token|complete_run\(CompleteRunRequest" tools/ironclaw_stress/src/user_turn.rs -n -A3 -B3

Repository: nearai/ironclaw

Length of output: 50372


🏁 Script executed:

#!/bin/bash
set -euo pipefail

sed -n '1,260p' tools/ironclaw_stress/src/user_turn.rs

Repository: nearai/ironclaw

Length of output: 10461


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '## runner request/response types\n'
sed -n '1,220p' crates/ironclaw_turns/src/runner.rs

printf '\n## claim_next_run implementation contract in memory store\n'
sed -n '1741,1845p' crates/ironclaw_turns/src/memory/mod.rs

printf '\n## claim_next_run / complete_run in filesystem row store\n'
sed -n '1361,1495p' crates/ironclaw_turns/src/filesystem_store/row_store.rs

printf '\n## TurnLifecycleChurn scenario wiring / operation counts\n'
rg -n "TurnLifecycleChurn|turn-lifecycle-churn|threads_per_owner|operations_per_thread|turns_per_operation" tools/ironclaw_stress/src/user_turn.rs tools/ironclaw_stress/src/*.rs -A4 -B4

Repository: nearai/ironclaw

Length of output: 50371


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '## synthetic user turn context / scope derivation\n'
sed -n '1,220p' tools/ironclaw_stress/src/synthetic.rs

printf '\n## turn scope and thread ownership references\n'
rg -n "turn_scope|thread_scope|thread_id|thread_owner|threads_per_owner|scope_filter" tools/ironclaw_stress/src/synthetic.rs tools/ironclaw_stress/src/user_turn.rs tools/ironclaw_stress/src/main.rs -A4 -B4

Repository: nearai/ironclaw

Length of output: 49021


🏁 Script executed:

#!/bin/bash
set -euo pipefail

rg -n "fn user_turn_context|turn_scope|thread_scope|thread_owner_user_id|scope_filter" tools/ironclaw_stress/src/synthetic.rs -A30 -B10

Repository: nearai/ironclaw

Length of output: 6944


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '## TurnScope definition and resource conversion\n'
rg -n "struct TurnScope|impl TurnScope|to_resource_scope|new_with_owner" crates/ironclaw_turns crates/ironclaw_threads -A40 -B10

printf '\n## user_turn_mount_view / row-store key inputs\n'
rg -n "user_turn_mount_view|row_turn_store_key|to_resource_scope\\(" tools/ironclaw_stress/src/user_turn.rs crates/ironclaw_turns/src -A20 -B10

Repository: nearai/ironclaw

Length of output: 50372


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '## TurnScope struct and state-path helpers\n'
rg -n "pub struct TurnScope|impl TurnScope|thread_root_string|turn_root_string|state.json|turns/state" crates/ironclaw_turns crates/ironclaw_threads -A40 -B10

printf '\n## claim/submit scope filtering helpers\n'
rg -n "pop_matching_queued_run|scope_filter|new_with_owner|TurnScope::new_with_owner" crates/ironclaw_turns/src -A30 -B10

Repository: nearai/ironclaw

Length of output: 50374


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '## TurnScope definition\n'
sed -n '1,120p' crates/ironclaw_turns/src/contract.rs

printf '\n## partitioned_worker_index helper\n'
sed -n '220,320p' tools/ironclaw_stress/src/synthetic.rs

Repository: nearai/ironclaw

Length of output: 258


🏁 Script executed:

#!/bin/bash
set -euo pipefail

sed -n '220,320p' tools/ironclaw_stress/src/synthetic.rs

Repository: nearai/ironclaw

Length of output: 336


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '## validate_args checks for active_thread_count / threads_per_owner / TurnLifecycleChurn\n'
rg -n "active_thread_count|threads_per_owner|TurnLifecycleChurn|turn-lifecycle-churn|requires --processes 1|one thread per owner" tools/ironclaw_stress/src/main.rs tools/ironclaw_stress/src/tests.rs -A6 -B6

printf '\n## user_turn_task loop / concurrency shape\n'
rg -n "run_user_turn_tasks|operations_per_thread|for operation_index|worker_index" tools/ironclaw_stress/src/user_turn.rs tools/ironclaw_stress/src/main.rs -A10 -B10

Repository: nearai/ironclaw

Length of output: 50371


Use the claimed run here. claim_next_run() returns a ClaimedTurnRun, but this branch drops it and completes the submit-time run_id. In the shared-hot-thread mode this workload supports (--active-thread-count 1), multiple workers can queue runs under the same turn_scope, so complete_run() can hit the wrong lease and leave the actual claim running. Use claimed.state.run_id / claimed.runner_id / claimed.lease_token like the other call sites.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tools/ironclaw_stress/src/user_turn.rs` around lines 757 - 820, The
TurnLifecycleChurn branch in user_turn.rs is discarding the ClaimedTurnRun
returned by claim_next_run() and then calling complete_run() with the
submit-time run_id instead of the claimed state. Update this path to capture the
claimed value from turn_store.claim_next_run in the existing time_stage call,
and use claimed.state.run_id, claimed.runner_id, and claimed.lease_token when
completing the run, matching the other call sites and preserving the actual
lease ownership.

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5667 — fab13d88 Deployed Jul 5, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: dependencies Dependency updates scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant