Skip to content

[HST Postgres 1/4] RootFilesystem latency substrate - #5688

Closed
serrrfirat wants to merge 2 commits into
mainfrom
codex/hst-postgres-01-rootfs
Closed

serrrfirat wants to merge 2 commits into
mainfrom
codex/hst-postgres-01-rootfs

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 6, 2026 •

Copy link
Copy Markdown
Collaborator

Stack

1/4 for hosted-single-tenant Postgres latency parity.

Base: main
Next: codex/hst-postgres-02-turn-state (#5689)

This first PR is intentionally the reviewer entry point for the whole stack. The code here is only the RootFilesystem substrate slice, but the journey, experiments, and final gains are summarized here so reviewers do not have to start from the large final harness PR.

Stack Map

What This PR Changes

  • Optimizes the Postgres RootFilesystem substrate used by later stores:
    • serializes root filesystem migrations with an advisory lock and process-local migrated-schema cache
    • moves projection indexes to shared names and includes path in exact/prefix indexes
    • uses prepared statements for RootFilesystem query paths
    • exposes transaction-level sequence reservation for append-heavy stores
  • Adds a small libSQL connection-init retry fix for transient setup races.
  • Adds an event-store helper for wrapping an already-migrated RootFilesystem.

Hosted-Volume Safety

No hosted profile wiring changes here. This PR should not switch hosted-single-tenant-volume to any new store layout by itself.

Live-data safety gates across the stack:

Starting Point

The production-shaped hosted-volume reference the stack had to beat:

  • c32 mixed flow: op p95 154.4ms, resource-governor p95 52.2ms
  • c100 mixed flow: op p95 388.1ms, resource-governor p95 213.3ms
  • resource-only c100: p95 158.9ms, throughput 859.5 ops/sec

Other important baselines discovered during the cycles:

  • Hosted substrate build before migration memoization: libSQL p95 31.87ms, Postgres pool-1 p95 68.77ms, pool-2 p95 52.97ms.
  • Blob turn-state size growth: chat-turn turn-store p95 climbed by quartile from 32.7ms -> 60.9ms -> 78.8ms -> 114.4ms as /turns/state.json grew.
  • Locked turn lifecycle blob signal: libSQL c4 p95 about 5.6-8.4s, Postgres blob/early row path in seconds under c4/c8/c100 pressure.
  • High-concurrency row turn-state before group/index work: Postgres-only c100 turn lifecycle p95 29.96s, then 15.04s, then 10.56s across early row cycles.
  • Thread list before index rows: libSQL cold 11.31s/warm 10.75s; Postgres cold 1.35s/warm 1.23s for 1000 threads.
  • Full-flow c100 after row turn-state but before governor fix: op p95 467.7ms, resource-governor p95 271.0ms, turn-store p95 85.5ms.

Ending Point

Final recorded push-gate numbers after the stack:

  • Postgres c100 mixed-user-session, pool32, filesystem-row: op p95 105.7ms, p99 173.7ms, throughput 2,246.6 ops/sec.
  • Attribution for that final c100 run: thread-store writes p95 56.0ms, turn-store p95 43.0ms, resource-governor p95 6.2ms, context reads p95 5.1ms.
  • Postgres c100 reserve-reconcile: p95 8.8ms, p99 15.7ms, throughput 18,134.4 ops/sec.
  • Postgres c100 turn-lifecycle-churn: p95 76.0ms, p99 132.7ms, throughput 5,966.9 ops/sec.
  • Clean libSQL vs Postgres c4 turn lifecycle after row-store alignment: libSQL p95 30.9ms, Postgres pool-2 p95 27.6ms.
  • Indexed thread list latest: libSQL op p95 5.9ms, Postgres op p95 6.0ms; warm reads are about 185us and 211us respectively.
  • Postgres c100 secret consume latest: p95 19.3ms, p99 21.7ms, throughput 10,656.2 ops/sec.

Gains

Area Before After Gain
c100 mixed flow op p95 388.1ms hosted-volume reference 105.7ms 3.7x lower p95
c100 mixed flow governor p95 213.3ms hosted-volume reference 6.2ms attribution 34.4x lower p95
resource-only c100 p95 158.9ms 8.8ms 18.1x lower p95
resource-only throughput 859.5 ops/sec 18,134.4 ops/sec 21.1x higher throughput
c100 turn lifecycle p95 about 10.56s after initial row indexing cycle, 29.96s before lifecycle delta work 76.0ms 139x to 394x lower p95 depending baseline
libSQL c4 turn lifecycle about 7.8s when still on blob path 30.9ms about 253x lower p95
Postgres c4 turn lifecycle about 35ms after row alignment 27.6ms holds parity
libSQL 1000-thread cold list 11.31s 5.7ms about 1984x lower cold-list time
Postgres 1000-thread cold list 1.35s 5.8ms about 233x lower cold-list time
Postgres c100 secret consume temporary/direct-store path removed; unified FS store latest 19.3ms p95 unified path still clears c100 gate

Experiment Ledger

The detailed journey is summarized here in PR text; the raw scratch log is intentionally not committed to the repo.

  • Cycle 0, Harness Bootstrap: built the first RootFilesystem latency harness; found Postgres append/tail slower than libSQL and acceptance gaps in turns/triggers/resources/secrets.
  • Cycle 1, Invalid Pool-Sizing Detour: detected that raising Postgres pool size made scores look better but violated the hosted pool constraint; reverted the detour.
  • Cycle 2, Encode Pool-Cap Scoring: made the scorer compare one libSQL baseline against fixed Postgres pool sizes so pool inflation could not hide regressions.
  • Cycle 3, LFD Goal Scaffold: wrote the loss/constraint scaffold and made acceptance gaps explicit.
  • Cycle 4, Hosted Substrate Build Workload: measured real hosted runtime construction; Postgres pool-2 p95 was 52.97ms vs libSQL 31.87ms.
  • Cycle 5, Postgres Migration Memoization: memoized repeated root migrations; hosted substrate dropped to Postgres pool-1 p95 18.05ms and pool-2 p95 20.89ms.
  • Cycle 6, Trigger Coverage: stabilized trigger setup and added durable trigger latency coverage; trigger rows passed and Postgres was faster in focused runs.
  • Cycle 7, Control-Plane Snapshot Coverage: added approval/secret/resource snapshot coverage; exposed secret lease consume CAS retry storms.
  • Cycle 8, Temporary Postgres Secret Rows: proved per-secret/per-lease row shape removed secret retry storms, but this direct store was later deleted in favor of RootFilesystem rows.
  • Cycle 9, Temporary Postgres Resource Rows: proved row-shaped governor state could beat libSQL on control-plane paths, but still bypassed RootFilesystem.
  • Cycle 10, Production Resource Wiring Detour: wired native Postgres resource governor as a diagnostic production path; later replaced by the filesystem governor.
  • Cycle 11, Shared Postgres Query Indexes: replaced per-prefix projection indexes with shared indexes and prepared query paths; removed stable query hard failures.
  • Cycle 12, LibSQL Trigger PRAGMA Drain: fixed transient libSQL trigger baseline failures by retrying connection PRAGMA setup.
  • Cycle 13, Single-Query Postgres Stat: collapsed Postgres stat to one cached query; control-plane probe stayed faster than libSQL.
  • Cycle 14, Resource Migration Memoization: memoized resource migrations; improved substrate path but did not yet hit full latency target.
  • Cycle 15, Root Migration Front-Guard: added composition-level migration front guard; hosted substrate moved to about 10.9-11.8ms p50/p95 vs libSQL about 13.6-15.2ms.
  • Cycle 16, Stress E2E Pool Deadlock: switched to ironclaw_stress; found and fixed nested Postgres pool checkout deadlock via transaction-local sequence reservation.
  • Cycle 17, Governor Worker Serialization: widened the temporary Postgres resource governor worker path; c16 pool-2 op p95 improved 104.0ms -> 86.0ms.
  • Cycle 18, Holdout LibSQL Connection Flake: retried libSQL PRAGMA setup and got holdout scoring clean; E2E c16 Postgres passed with p95 76.7ms on pool-2.
  • Cycle 19, Turn-State Attribution: instrumented full turn flow and proved blob CAS cost grows with state size; retention caps helped but could not solve the design.
  • Cycle 20, Filesystem Turn-State Row Layout: introduced typed row/append turn state behind RootFilesystem; validated shape but first row versions were still too slow.
  • Cycle 21, Turn Lifecycle Harness Signal: added locked turn_lifecycle_blob workload so blob contention became visible in the scorer.
  • Cycle 22, Postgres Row Turn-State Wiring: wired Postgres to row turn state and made the dev scorer pass, while c32/c100 still showed same-user serialization.
  • Cycle 23, High-Concurrency Resource Pressure: measured c32/c100 full flow and identified resource governor as the top full-flow bottleneck.
  • Cycle 24, Diagnostic Backend Filter: added diagnostic-only backend filtering so Postgres c100 rows could be measured even when libSQL crashed first.
  • Cycle 25, WebUI Session Path: added real Axum /api/webchat/v2/session workload; c1/c4 passed, c100 showed middleware/read-path pressure.
  • Cycle 26, Loop Checkpoint Targeted Deltas: made loop checkpoint writes targeted row deltas; c100 turn lifecycle improved but still sat in seconds.
  • Cycle 27, Lifecycle Targeted Deltas: made block/resume/cancel/request-cancel targeted; Postgres c100 turn lifecycle improved 29.96s -> 15.04s.
  • Cycle 28, Resource Shared-Row Contention: optimized temporary native governor unlimited path; captured hosted-volume reference c32/c100 numbers used as baseline.
  • Cycle 29, Checkpoint Readback Projection: removed full snapshot rebuild from checkpoint reads; c100 turn lifecycle improved 14.72s -> 13.23s.
  • Cycle 30, Run-State Readback Projection: projected a single run directly; c100 turn lifecycle improved 13.23s -> 12.50s.
  • Cycle 31, Event Tail Tracking: cached latest lifecycle event cursor; mostly flat, so we stopped tuning that lever.
  • Cycle 32, In-Place Delta Apply: replaced vector rebuilds with in-place row updates; c100 improved 12.43s -> 10.56s.
  • Cycle 33, Pool Sweep and Single-Run Lease Prep: confirmed pool size was not the main lever and prepared targeted lease overlay work.
  • Cycle 34, Direct Loop Checkpoint Row Deltas: continued checkpoint delta cleanup; held semantics while reducing row-store churn.
  • Cycle 35, Claim Lease Seeding: removed more full-snapshot clone/rebuild from lease paths; c100 improved to about 10.20s.
  • Cycle 36, Single-Run Overlay: applied runner lease overlay directly to one run; c100 improved 10.20s -> 8.67s.
  • Cycle 37, Sparse Delta Experiment: tried sparse JSON delta encoding; rejected it because p95 did not improve.
  • Cycle 38, Group-Commit Delta Journal: added the single flusher and append_batch; c100 turn lifecycle improved 8.67s -> 3.52s.
  • Cycle 39, Indexed Row Snapshot Apply: added row-keyed hot indexes; c100 turn lifecycle improved 3.52s -> 790ms and dev c4 Postgres reached about 35ms.
  • Cycle 40, RootFilesystem Journaled Governor: replaced direct Postgres governor with in-process authority plus RootFilesystem delta journal; c100 governor p95 dropped 271.0ms -> 31.1ms.
  • Cycle 41, Post-Governor Attribution: after governor fix, top groups moved to thread-store writes and turn-store rather than governor.
  • Cycle 42, LibSQL Turn-State Cliff Diagnostic: diagnosed 7.8s libSQL c4 as the harness still using blob turn state; fixed libSQL to row store and got c4 p95 30.9ms.
  • Cycle 43, Two-Tier Turn-State Lifecycle: made terminal run eviction a hot-cache policy, not durable row deletion; c100 full mixed flow p95 118.3ms.
  • Cycle 44, Thread Index Rows: added derivable per-thread index rows and warm cache; 1000-thread list fell from seconds to about 23ms, then later to about 6ms.
  • Cycle 45, Remove Direct Postgres Secret Store: deleted the temporary direct DB secret store and kept secrets on unified per-record RootFilesystem layout; c100 Postgres secret consume stayed around 20ms p95.
  • Cycle 46, Thermo Cleanup: split oversized modules and reran gates; mixed-flow c100 improved to 102.3ms p95 in that run, thread list to about 5ms.
  • Cycle 47, Review Fixes: fixed real review findings in checkpoint concurrency, stale thread-index delete, partial bootstrap, and no-op index touch; reran benchmarks and found a turn-lifecycle regression.
  • Cycle 48, Governor/Turn Lock Boundary: released governor and turn commit gates before durable ack waits while preserving enqueue ordering; final c100 mixed p95 105.7ms, governor p95 6.2ms, turn-lifecycle p95 76.0ms.
  • Cycle 49, Turn Blob-to-Row Migration Gate: added the legacy /turns/state.json import gate and stale-blob no-remigrate tests so the stack has a live hosted-volume migration path before go-live.

Verification For This PR

  • cargo check -p ironclaw_filesystem --features libsql,postgres
  • cargo check -p ironclaw_reborn_event_store --features libsql,postgres

Stack-wide verification is listed in the dependent PRs and summarized above.

@ironloopai

ironloopai Bot commented Jul 6, 2026 •

Copy link
Copy Markdown
Contributor

✅ IronLoop Review Status

Head: 01db07815ab950adfe9df26d68514f43eaa8927f
Result: 1/1 reviewers completed without blocking findings.
Next: Ready for normal human review and CI checks.
Updated: 2026-07-06T13:29:51.298Z

Current reviewers:

Reviewer State Verdict Findings Last update
ironloop/common-reviewer (reviewer) Completed Approved 0 blocking findings / 0 notes 2026-07-06T13:29:51.270Z
Reviewer summaries
Reviewer Detail
ironloop/common-reviewer (reviewer) Approved; 0 blocking findings; No concrete blocking issues found in the reviewed diff. The changes keep the filesystem/event-store behavior scoped and add reasonable fail-closed/default handling for transaction…
Recent activity
Time Reviewer State Detail
2026-07-06T13:27:13.120Z ironloop/common-reviewer (reviewer) Queued Accepted review request for head 01db078.
2026-07-06T13:27:13.120Z ironloop/common-reviewer (reviewer) Queued Waiting for this reviewer lane to become available.
2026-07-06T13:27:13.145Z ironloop/common-reviewer (reviewer) Queued Added to the local review work handoff.
2026-07-06T13:27:15.156Z ironloop/common-reviewer (reviewer) Started Reviewer worker started attempt 1.
2026-07-06T13:27:19.103Z ironloop/common-reviewer (reviewer) Workspace ready Prepared isolated checkout (merge_ref) at e0cd2bd.
2026-07-06T13:29:32.476Z ironloop/common-reviewer (reviewer) Running Codex is reviewing; process live; elapsed 2m 15s; timeout in 17m 45s; last heartbeat 2026-07-06T13:29:32.476Z. Activity (stderr): ...ration::ReserveSeq)?; let result = self.root.reserve_sequence(&virtual_path).await; trace_fs_latency("reserve_sequen….
2026-07-06T13:29:51.270Z ironloop/common-reviewer (reviewer) Result captured Approved; 0 blocking findings.
2026-07-06T13:29:51.270Z ironloop/common-reviewer (reviewer) Completed Review completed and terminal status was persisted.
Available commands
  • @ironloop agents
  • @ironloop review
  • @ironloop review --agent <agent-id-or-alias>
  • @ironloop status
Run metadata

Admission: webhook accepted the request and IronLoop persisted review state before this projection.

@coderabbitai

coderabbitai Bot commented Jul 6, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • New Features
    • Added sequence reservation support to storage transactions, including scoped access/path enforcement.
    • Added a feature-gated helper to build reborn event stores from an existing root filesystem for supported backends.
  • Bug Fixes
    • Improved libSQL connection retry behavior so PRAGMA initialization failures are retried as transient errors.
    • Enhanced PostgreSQL migration reliability, shared projection index handling, and optimized query/stat/sequence operations.
  • Tests
    • Added regression coverage for libSQL retry recovery and a unit check for default “unsupported” sequence reservation behavior.

Walkthrough

Adds reserve_sequence to filesystem transactions with scoped and Postgres support, hardens Postgres migrations and query paths, fixes libSQL PRAGMA retry handling, and exposes a RootFilesystem-backed event-store builder.

Changes

Filesystem transaction and backend updates

Layer / File(s) Summary
StorageTxn reserve_sequence contract and test
crates/ironclaw_filesystem/src/backend.rs
Adds reserve_sequence to StorageTxn with a default Unsupported implementation and validates the default behavior with a DummyTxn test.
Scoped reserve_sequence forwarding
crates/ironclaw_filesystem/src/scoped.rs
ScopedStorageTxn checks ReserveSeq permission and mount-prefix containment before delegating reserve_sequence to the inner transaction.
RootFilesystem event store builder
crates/ironclaw_reborn_event_store/src/lib.rs
Adds a cfg-gated public builder that wraps a RootFilesystem into RebornEventStores.

Postgres and libSQL backend behavior

Layer / File(s) Summary
libSQL retry and PRAGMA initialization
crates/ironclaw_filesystem/src/libsql.rs
Splits retry setup, makes PRAGMA failures retryable, updates the terminal error text, and adds a regression test for transient PRAGMA failure.
Postgres migration locking and legacy cleanup
crates/ironclaw_filesystem/src/postgres.rs
Adds migration-key memoization, advisory-lock protection, and legacy projection-index cleanup during migrations.
Postgres index, query, stat, and sequence helpers
crates/ironclaw_filesystem/src/postgres.rs
Updates shared projection index naming, caches generated queries, and centralizes stat and reserve_sequence SQL in shared helpers with tests for naming and quoting.

Estimated code review effort: 4 (Complex) | ~60 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Caller
  participant ScopedStorageTxn
  participant PostgresStorageTxn
  participant Postgres

  Caller->>ScopedStorageTxn: reserve_sequence(path)
  ScopedStorageTxn->>ScopedStorageTxn: check ReserveSeq permission
  ScopedStorageTxn->>ScopedStorageTxn: check mount_prefix containment
  ScopedStorageTxn->>PostgresStorageTxn: reserve_sequence(path)
  PostgresStorageTxn->>Postgres: postgres_reserve_sequence_with_client (INSERT...ON CONFLICT...RETURNING)
  Postgres-->>PostgresStorageTxn: reserved value
  PostgresStorageTxn-->>ScopedStorageTxn: SeqNo
  ScopedStorageTxn-->>Caller: SeqNo
Loading

Possibly related PRs

  • nearai/ironclaw#5451: Both PRs change crates/ironclaw_filesystem/src/libsql.rs around connect_with_retry PRAGMA application and retry behavior.
  • nearai/ironclaw#5455: Both PRs extend the filesystem transaction surface with ReserveSeq plumbing through scoped/backed implementations.

Suggested reviewers: think-in-universe, henrypark133

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description is detailed, but it does not follow the required template and omits several mandatory sections like Change Type, Linked Issue, and Rollback Plan. Reformat it to the repo template: add Summary bullets, Change Type, Linked Issue, Validation checkboxes, Security Impact, DB Impact, Blast Radius, Rollback Plan, and Review track.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title is clearly about the RootFilesystem substrate work and matches the main changeset, though it is not Conventional Commits style.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5688 July 6, 2026 12:26 Destroyed
@github-actions github-actions Bot added the size: L 200-499 changed lines label Jul 6, 2026
@github-actions github-actions Bot added risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 6, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a reserve_sequence method to the StorageTxn trait (with implementations for Postgres and Scoped transactions), refactors Postgres migrations to use advisory locks and track migrated schemas, optimizes Postgres stat queries, and adds a public helper to build event stores from a root filesystem. Feedback on the changes suggests addressing a potential collision risk in postgres_shared_projection_index_name where concatenating keys with a simple _ delimiter could cause different key sets to produce identical index names; using an injective length-prefixed encoding is recommended to prevent this.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +1686 to +1692
let keys = spec
.keys
.iter()
.map(|key| key.as_str())
.collect::<Vec<_>>()
.join("_");
sql_index_name(&format!("/shared/{kind}/{keys}"), spec.name.as_str())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The generation of the shared projection index name concatenates the keys using a simple _ delimiter. This can lead to collisions where different sets of keys (e.g., ["a", "b"] vs ["a_b"]) produce the same index name, causing CREATE INDEX IF NOT EXISTS to silently skip creating the second index.\n\nTo prevent separator-collision attacks or accidental collisions, use an injective length-prefixed encoding or another injective scheme when combining the keys and kind before generating the index name.

References
  1. When generating deterministic identifiers or hashes from multiple string components, use an injective length-prefixed encoding rather than simple concatenation with a delimiter to eliminate the risk of separator-collision attacks or accidental collisions.

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ IronLoop Review: reviewer

Verdict: ✅ Approved
Findings: 0 blocking / 0 notes
Next: No reviewer action needed.
Head: 5b15dce48899c1b443ab21d1f80f986306f03ef5

Run details

Status: Current
Needs human: no
Needs validation: no

**Inline candidates:** 0

Summary

No concrete blocking issues found in the changed filesystem backend and event-store wrapper code. The PR keeps the changes scoped to backend migration/index/sequence behavior and preserves libSQL/Postgres parity expectations in the reviewed paths.

Findings

None.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloop review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloop review when the fix may affect multiple areas.
  4. Use @ironloop status to check queued/running/completed/stale/stalled state while reviewers run.

@github-actions

github-actions Bot commented Jul 6, 2026 •

Copy link
Copy Markdown
Contributor

⚠️ 9 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_embeddings, ironclaw_gateway, ironclaw_hooks, ironclaw_oauth, ironclaw_process_sandbox, ironclaw_prompt_envelope, ironclaw_scripts, ironclaw_skill_learning, ironclaw_tui

Reborn integration-tier coverage

Line coverage (Reborn crates): 28.52% — 49276 / 172796 lines

Per-crate breakdown (62 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_embeddings 0% 0 / 337
ironclaw_gateway 0% 0 / 283
ironclaw_hooks 0% 0 / 4468
ironclaw_oauth 0% 0 / 155
ironclaw_process_sandbox 0% 0 / 795
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 347
ironclaw_skill_learning 0% 0 / 61
ironclaw_tui 0% 0 / 4776
ironclaw_outbound 0.22% 3 / 1339
ironclaw_event_projections 0.4% 6 / 1489
ironclaw_reborn_event_store 0.66% 6 / 913
ironclaw_reborn_config 1.36% 15 / 1101
ironclaw_llm 3.62% 437 / 12075
ironclaw_event_streams 3.87% 40 / 1034
ironclaw_product_adapter_registry 5.38% 25 / 465
ironclaw_extractors 6.18% 26 / 421
ironclaw_wasm_sandbox_core 7.37% 7 / 95
ironclaw_product_workflow 7.97% 768 / 9635
ironclaw_webui_v2 8.5% 228 / 2683
ironclaw_processes 8.61% 98 / 1138
ironclaw_common 10.22% 74 / 724
ironclaw_events 12.45% 143 / 1149
ironclaw_product_adapters 12.53% 280 / 2234
ironclaw_network 13.25% 66 / 498
ironclaw_skills 14.58% 377 / 2585
ironclaw_first_party_extensions 22.46% 1125 / 5010
ironclaw_triggers 23.1% 663 / 2870
ironclaw_reborn_traces 23.24% 1492 / 6420
ironclaw_secrets 26.22% 450 / 1716
ironclaw_reborn 28.98% 2542 / 8771
ironclaw_reborn_composition 30.36% 9256 / 30483
ironclaw_capabilities 32.97% 580 / 1759
ironclaw_auth 33.09% 667 / 2016
ironclaw_runtime_policy 33.2% 80 / 241
ironclaw_memory_native 37.02% 857 / 2315
ironclaw_filesystem 38.37% 1411 / 3677
ironclaw_host_api 39.9% 942 / 2361
ironclaw_host_runtime 41.16% 6170 / 14989
ironclaw_threads 41.98% 1326 / 3159
ironclaw_trust 42.56% 326 / 766
ironclaw_loop_support 42.68% 3142 / 7362
ironclaw_memory 47.47% 357 / 752
ironclaw_first_party_extension_ports 48.74% 637 / 1307
ironclaw_wasm 48.79% 363 / 744
ironclaw_projects 50% 147 / 294
ironclaw_extensions 51.4% 1211 / 2356
ironclaw_agent_loop 51.49% 2400 / 4661
ironclaw_resources 51.63% 1109 / 2148
ironclaw_run_state 52.73% 222 / 421
ironclaw_authorization 53.54% 461 / 861
ironclaw_turns 57.8% 5222 / 9035
ironclaw_safety 59.38% 1035 / 1743
ironclaw_observability 61.54% 16 / 26
ironclaw_conversations 66.13% 937 / 1417
ironclaw_approvals 66.63% 549 / 824
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_mcp 67.42% 569 / 844
ironclaw_reborn_identity 70.91% 156 / 220
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_product_context 78.57% 11 / 14
ironclaw_attachments 84.92% 107 / 126

This signal is informational: coverage never gates the PR — not the percentage, not the per-crate holes, not the 0-coverage callout.

Exemptions (0 file(s) excluded from the accounting above)

No exemptions configured.

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Cycle 38 is the goat.

@serrrfirat
serrrfirat marked this pull request as ready for review July 6, 2026 12:39

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_filesystem/src/libsql.rs (1)

160-176: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate retry/backoff logic across both Err arms.

The execute_batch failure arm (162-167) and the open() failure arm (170-175) both do last_error = Some(error); if attempt + 1 < LIBSQL_CONNECT_ATTEMPTS { sleep(...) }. Worth extracting into a small helper so the two retry sources can't drift.

♻️ Proposed refactor
+fn record_retry(
+    last_error: &mut Option<libsql::Error>,
+    error: libsql::Error,
+) -> Option<std::time::Duration> {
+    *last_error = Some(error);
+    None // caller checks attempt bound and calls connect_backoff
+}

Simplest form: keep the if attempt + 1 < LIBSQL_CONNECT_ATTEMPTS { sleep(...) }.await inline at the call site but move last_error = Some(error) + bound-check into one closure invoked from both arms, e.g.:

-                match conn.execute_batch(LIBSQL_CONNECTION_PRAGMAS).await {
-                    Ok(_) => return Ok(conn),
-                    Err(error) => {
-                        last_error = Some(error);
-                        if attempt + 1 < LIBSQL_CONNECT_ATTEMPTS {
-                            tokio::time::sleep(connect_backoff(attempt)).await;
-                        }
-                    }
-                }
+                match conn.execute_batch(LIBSQL_CONNECTION_PRAGMAS).await {
+                    Ok(_) => return Ok(conn),
+                    Err(error) => last_error = Some(error),
+                }

and hoist a single if attempt + 1 < LIBSQL_CONNECT_ATTEMPTS && last_error.is_some() { sleep(...).await } after the outer match once both arms just record the error.

As per coding guidelines, "Keep functions focused and extract helpers when logic is reused."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_filesystem/src/libsql.rs` around lines 160 - 176, The retry
handling in the libsql connection loop duplicates the same “record error and
maybe sleep” logic in both the `execute_batch` and `open()` failure branches.
Refactor the `connect` flow in `libsql.rs` so both `Err` arms only capture the
error source, then delegate the shared `last_error = Some(error)` plus `attempt
+ 1 < LIBSQL_CONNECT_ATTEMPTS` backoff check to a single helper or closure used
by both paths, keeping the retry behavior identical.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_filesystem/src/libsql.rs`:
- Around line 160-168: The retry branch in the libsql connection flow is not
covered when `execute_batch(LIBSQL_CONNECTION_PRAGMAS)` fails, so add a
regression test that forces PRAGMA execution to fail on an initial attempt and
then succeed on a later retry. Extend the existing
`connect_retries_transient_open_failures_before_succeeding` coverage or add a
nearby test around the same `libsql::connect` path to verify the retry logic and
eventual success after a PRAGMA failure, using the same
`LIBSQL_CONNECTION_PRAGMAS` and `LIBSQL_CONNECT_ATTEMPTS` behavior.

---

Outside diff comments:
In `@crates/ironclaw_filesystem/src/libsql.rs`:
- Around line 160-176: The retry handling in the libsql connection loop
duplicates the same “record error and maybe sleep” logic in both the
`execute_batch` and `open()` failure branches. Refactor the `connect` flow in
`libsql.rs` so both `Err` arms only capture the error source, then delegate the
shared `last_error = Some(error)` plus `attempt + 1 < LIBSQL_CONNECT_ATTEMPTS`
backoff check to a single helper or closure used by both paths, keeping the
retry behavior identical.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 5255f9e1-f57a-45f9-94a8-5c579b1ccfa0

📥 Commits

Reviewing files that changed from the base of the PR and between 0606cd5 and 5b15dce.

📒 Files selected for processing (5)
  • crates/ironclaw_filesystem/src/backend.rs
  • crates/ironclaw_filesystem/src/libsql.rs
  • crates/ironclaw_filesystem/src/postgres.rs
  • crates/ironclaw_filesystem/src/scoped.rs
  • crates/ironclaw_reborn_event_store/src/lib.rs

Comment thread crates/ironclaw_filesystem/src/libsql.rs Outdated
@railway-app

railway-app Bot commented Jul 6, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5688 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 6, 2026 at 1:39 pm

@ironloopai

ironloopai Bot commented Jul 6, 2026

Copy link
Copy Markdown
Contributor

🗂️ Archived IronLoop Review: reviewer

This result is from an older PR head and is no longer the active review.

Field Value
Status Superseded
Verdict ✅ Approved
Findings 0 blocking / 0 notes
Reviewed head 5b15dce48899
Archived summary

No concrete blocking issues found in the changed filesystem backend and event-store wrapper code. The PR keeps the changes scoped to backend migration/index/sequence behavior and preserves libSQL/Postgres parity expectations in the reviewed paths.

@github-actions github-actions Bot added size: XL 500+ changed lines and removed size: L 200-499 changed lines labels Jul 6, 2026

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ IronLoop Review: reviewer

Verdict: ✅ Approved
Findings: 0 blocking / 0 notes
Next: No reviewer action needed.
Head: 01db07815ab950adfe9df26d68514f43eaa8927f

Run details

Status: Current
Needs human: no
Needs validation: no

**Inline candidates:** 0

Summary

No concrete blocking issues found in the reviewed diff. The changes keep the filesystem/event-store behavior scoped and add reasonable fail-closed/default handling for transaction sequence reservation and Postgres/libSQL backend paths.

Findings

None.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloop review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloop review when the fix may affect multiple areas.
  4. Use @ironloop status to check queued/running/completed/stale/stalled state while reviewers run.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5688 July 6, 2026 13:33 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_filesystem/src/postgres.rs (1)

466-471: 🚀 Performance & Scalability | 🟡 Minor | ⚡ Quick win

Avoid prepare_cached() for this query. crates/ironclaw_filesystem/src/postgres.rs:466-471 The statement cache is unbounded per connection, and this SQL text varies with the filter tree and LIMIT/OFFSET placeholder positions, so hot connections will accumulate prepared statements. Use prepare() here or bound the cache around a normalized query shape.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_filesystem/src/postgres.rs` around lines 466 - 471, The query
path in the Postgres filesystem code is using prepare_cached for SQL generated
from varying filter trees and pagination placeholders, which can grow the
per-connection statement cache without bound. Update the logic in the
query-building flow around the prepare_cached call to use prepare instead, or
otherwise constrain caching to a normalized query shape, while keeping the
existing db_error handling and subsequent client.query invocation intact.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@crates/ironclaw_filesystem/src/postgres.rs`:
- Around line 466-471: The query path in the Postgres filesystem code is using
prepare_cached for SQL generated from varying filter trees and pagination
placeholders, which can grow the per-connection statement cache without bound.
Update the logic in the query-building flow around the prepare_cached call to
use prepare instead, or otherwise constrain caching to a normalized query shape,
while keeping the existing db_error handling and subsequent client.query
invocation intact.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: b9b48a6b-47d9-4266-b301-6780c78d0cee

📥 Commits

Reviewing files that changed from the base of the PR and between 5b15dce and 01db078.

📒 Files selected for processing (2)
  • crates/ironclaw_filesystem/src/libsql.rs
  • crates/ironclaw_filesystem/src/postgres.rs

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Closing in favor of the re-sliced v2 stack rebuilt from the latest fixed head and rebased onto current main: #5724, #5725, #5726, #5727. The runtime fixes that restored the Cycle 49 gates have been folded into the owning subsystem PRs.

@serrrfirat serrrfirat closed this Jul 6, 2026

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5688 — 01db0781 Deployed Jul 6, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant