Skip to content

[HST Postgres v2 1/4] RootFilesystem latency substrate - #5724

Merged
serrrfirat merged 2 commits into
mainfrom
codex/hst-postgres-v2-01-rootfs
Jul 7, 2026
Merged

serrrfirat merged 2 commits into
mainfrom
codex/hst-postgres-v2-01-rootfs

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 6, 2026 •

Copy link
Copy Markdown
Collaborator

Replacement Stack

Replaces the old HST Postgres latency stack #5688, #5689, #5690, and #5691.

This v2 stack is rebuilt from the latest fixed head and rebased onto current main. The fixes that restored the Cycle 49 benchmark gates have been folded down into the subsystem PRs that own them:

What This PR Changes

  • Optimizes the Postgres RootFilesystem substrate used by later stores.
  • Serializes root filesystem migrations and memoizes completed schema setup.
  • Improves exact/prefix query indexes and prepared query paths.
  • Adds transaction-level sequence reservation support for append-heavy stores.
  • Includes the later libSQL RootFilesystem recovery fix and contract coverage.
  • Adds an event-store helper for an already-migrated RootFilesystem.

Hosted-Volume Safety

No hosted profile wiring changes here. This PR should not switch hosted-single-tenant-volume to any new store layout by itself.

Starting Point

The production-shaped hosted-volume reference the stack had to beat:

  • c32 mixed flow: op p95 154.4ms, resource-governor p95 52.2ms
  • c100 mixed flow: op p95 388.1ms, resource-governor p95 213.3ms
  • resource-only c100: p95 158.9ms, throughput 859.5 ops/sec

Other important baselines discovered during the cycles:

  • Hosted substrate build before migration memoization: libSQL p95 31.87ms, Postgres pool-1 p95 68.77ms, pool-2 p95 52.97ms.
  • Blob turn-state size growth: chat-turn turn-store p95 climbed by quartile from 32.7ms -> 60.9ms -> 78.8ms -> 114.4ms as /turns/state.json grew.
  • Locked turn lifecycle blob signal: libSQL c4 p95 about 5.6-8.4s, Postgres blob/early row path in seconds under c4/c8/c100 pressure.
  • High-concurrency row turn-state before group/index work: Postgres-only c100 turn lifecycle p95 29.96s, then 15.04s, then 10.56s across early row cycles.
  • Thread list before index rows: libSQL cold 11.31s/warm 10.75s; Postgres cold 1.35s/warm 1.23s for 1000 threads.
  • Full-flow c100 after row turn-state but before governor fix: op p95 467.7ms, resource-governor p95 271.0ms, turn-store p95 85.5ms.

Ending Point

Final recorded push-gate numbers after the stack:

  • Postgres c100 mixed-user-session, pool32, filesystem-row: op p95 105.7ms, p99 173.7ms, throughput 2,246.6 ops/sec.
  • Attribution for that final c100 run: thread-store writes p95 56.0ms, turn-store p95 43.0ms, resource-governor p95 6.2ms, context reads p95 5.1ms.
  • Postgres c100 reserve-reconcile: p95 8.8ms, p99 15.7ms, throughput 18,134.4 ops/sec.
  • Postgres c100 turn-lifecycle-churn: p95 76.0ms, p99 132.7ms, throughput 5,966.9 ops/sec.
  • Clean libSQL vs Postgres c4 turn lifecycle after row-store alignment: libSQL p95 30.9ms, Postgres pool-2 p95 27.6ms.
  • Indexed thread list latest: libSQL op p95 5.9ms, Postgres op p95 6.0ms; warm reads are about 185us and 211us respectively.
  • Postgres c100 secret consume latest: p95 19.3ms, p99 21.7ms, throughput 10,656.2 ops/sec.

Gains

Area Before After Gain
c100 mixed flow op p95 388.1ms hosted-volume reference 105.7ms 3.7x lower p95
c100 mixed flow governor p95 213.3ms hosted-volume reference 6.2ms attribution 34.4x lower p95
resource-only c100 p95 158.9ms 8.8ms 18.1x lower p95
resource-only throughput 859.5 ops/sec 18,134.4 ops/sec 21.1x higher throughput
c100 turn lifecycle p95 about 10.56s after initial row indexing cycle, 29.96s before lifecycle delta work 76.0ms 139x to 394x lower p95 depending baseline
libSQL c4 turn lifecycle about 7.8s when still on blob path 30.9ms about 253x lower p95
Postgres c4 turn lifecycle about 35ms after row alignment 27.6ms holds parity
libSQL 1000-thread cold list 11.31s 5.7ms about 1984x lower cold-list time
Postgres 1000-thread cold list 1.35s 5.8ms about 233x lower cold-list time
Postgres c100 secret consume temporary/direct-store path removed; unified FS store latest 19.3ms p95 unified path still clears c100 gate

Experiment Ledger

The detailed journey is summarized here in PR text; the raw scratch log is intentionally not committed to the repo.

  • Cycle 0, Harness Bootstrap: built the first RootFilesystem latency harness; found Postgres append/tail slower than libSQL and acceptance gaps in turns/triggers/resources/secrets.
  • Cycle 1, Invalid Pool-Sizing Detour: detected that raising Postgres pool size made scores look better but violated the hosted pool constraint; reverted the detour.
  • Cycle 2, Encode Pool-Cap Scoring: made the scorer compare one libSQL baseline against fixed Postgres pool sizes so pool inflation could not hide regressions.
  • Cycle 3, LFD Goal Scaffold: wrote the loss/constraint scaffold and made acceptance gaps explicit.
  • Cycle 4, Hosted Substrate Build Workload: measured real hosted runtime construction; Postgres pool-2 p95 was 52.97ms vs libSQL 31.87ms.
  • Cycle 5, Postgres Migration Memoization: memoized repeated root migrations; hosted substrate dropped to Postgres pool-1 p95 18.05ms and pool-2 p95 20.89ms.
  • Cycle 6, Trigger Coverage: stabilized trigger setup and added durable trigger latency coverage; trigger rows passed and Postgres was faster in focused runs.
  • Cycle 7, Control-Plane Snapshot Coverage: added approval/secret/resource snapshot coverage; exposed secret lease consume CAS retry storms.
  • Cycle 8, Temporary Postgres Secret Rows: proved per-secret/per-lease row shape removed secret retry storms, but this direct store was later deleted in favor of RootFilesystem rows.
  • Cycle 9, Temporary Postgres Resource Rows: proved row-shaped governor state could beat libSQL on control-plane paths, but still bypassed RootFilesystem.
  • Cycle 10, Production Resource Wiring Detour: wired native Postgres resource governor as a diagnostic production path; later replaced by the filesystem governor.
  • Cycle 11, Shared Postgres Query Indexes: replaced per-prefix projection indexes with shared indexes and prepared query paths; removed stable query hard failures.
  • Cycle 12, LibSQL Trigger PRAGMA Drain: fixed transient libSQL trigger baseline failures by retrying connection PRAGMA setup.
  • Cycle 13, Single-Query Postgres Stat: collapsed Postgres stat to one cached query; control-plane probe stayed faster than libSQL.
  • Cycle 14, Resource Migration Memoization: memoized resource migrations; improved substrate path but did not yet hit full latency target.
  • Cycle 15, Root Migration Front-Guard: added composition-level migration front guard; hosted substrate moved to about 10.9-11.8ms p50/p95 vs libSQL about 13.6-15.2ms.
  • Cycle 16, Stress E2E Pool Deadlock: switched to ironclaw_stress; found and fixed nested Postgres pool checkout deadlock via transaction-local sequence reservation.
  • Cycle 17, Governor Worker Serialization: widened the temporary Postgres resource governor worker path; c16 pool-2 op p95 improved 104.0ms -> 86.0ms.
  • Cycle 18, Holdout LibSQL Connection Flake: retried libSQL PRAGMA setup and got holdout scoring clean; E2E c16 Postgres passed with p95 76.7ms on pool-2.
  • Cycle 19, Turn-State Attribution: instrumented full turn flow and proved blob CAS cost grows with state size; retention caps helped but could not solve the design.
  • Cycle 20, Filesystem Turn-State Row Layout: introduced typed row/append turn state behind RootFilesystem; validated shape but first row versions were still too slow.
  • Cycle 21, Turn Lifecycle Harness Signal: added locked turn_lifecycle_blob workload so blob contention became visible in the scorer.
  • Cycle 22, Postgres Row Turn-State Wiring: wired Postgres to row turn state and made the dev scorer pass, while c32/c100 still showed same-user serialization.
  • Cycle 23, High-Concurrency Resource Pressure: measured c32/c100 full flow and identified resource governor as the top full-flow bottleneck.
  • Cycle 24, Diagnostic Backend Filter: added diagnostic-only backend filtering so Postgres c100 rows could be measured even when libSQL crashed first.
  • Cycle 25, WebUI Session Path: added real Axum /api/webchat/v2/session workload; c1/c4 passed, c100 showed middleware/read-path pressure.
  • Cycle 26, Loop Checkpoint Targeted Deltas: made loop checkpoint writes targeted row deltas; c100 turn lifecycle improved but still sat in seconds.
  • Cycle 27, Lifecycle Targeted Deltas: made block/resume/cancel/request-cancel targeted; Postgres c100 turn lifecycle improved 29.96s -> 15.04s.
  • Cycle 28, Resource Shared-Row Contention: optimized temporary native governor unlimited path; captured hosted-volume reference c32/c100 numbers used as baseline.
  • Cycle 29, Checkpoint Readback Projection: removed full snapshot rebuild from checkpoint reads; c100 turn lifecycle improved 14.72s -> 13.23s.
  • Cycle 30, Run-State Readback Projection: projected a single run directly; c100 turn lifecycle improved 13.23s -> 12.50s.
  • Cycle 31, Event Tail Tracking: cached latest lifecycle event cursor; mostly flat, so we stopped tuning that lever.
  • Cycle 32, In-Place Delta Apply: replaced vector rebuilds with in-place row updates; c100 improved 12.43s -> 10.56s.
  • Cycle 33, Pool Sweep and Single-Run Lease Prep: confirmed pool size was not the main lever and prepared targeted lease overlay work.
  • Cycle 34, Direct Loop Checkpoint Row Deltas: continued checkpoint delta cleanup; held semantics while reducing row-store churn.
  • Cycle 35, Claim Lease Seeding: removed more full-snapshot clone/rebuild from lease paths; c100 improved to about 10.20s.
  • Cycle 36, Single-Run Overlay: applied runner lease overlay directly to one run; c100 improved 10.20s -> 8.67s.
  • Cycle 37, Sparse Delta Experiment: tried sparse JSON delta encoding; rejected it because p95 did not improve.
  • Cycle 38, Group-Commit Delta Journal: added the single flusher and append_batch; c100 turn lifecycle improved 8.67s -> 3.52s.
  • Cycle 39, Indexed Row Snapshot Apply: added row-keyed hot indexes; c100 turn lifecycle improved 3.52s -> 790ms and dev c4 Postgres reached about 35ms.
  • Cycle 40, RootFilesystem Journaled Governor: replaced direct Postgres governor with in-process authority plus RootFilesystem delta journal; c100 governor p95 dropped 271.0ms -> 31.1ms.
  • Cycle 41, Post-Governor Attribution: after governor fix, top groups moved to thread-store writes and turn-store rather than governor.
  • Cycle 42, LibSQL Turn-State Cliff Diagnostic: diagnosed 7.8s libSQL c4 as the harness still using blob turn state; fixed libSQL to row store and got c4 p95 30.9ms.
  • Cycle 43, Two-Tier Turn-State Lifecycle: made terminal run eviction a hot-cache policy, not durable row deletion; c100 full mixed flow p95 118.3ms.
  • Cycle 44, Thread Index Rows: added derivable per-thread index rows and warm cache; 1000-thread list fell from seconds to about 23ms, then later to about 6ms.
  • Cycle 45, Remove Direct Postgres Secret Store: deleted the temporary direct DB secret store and kept secrets on unified per-record RootFilesystem layout; c100 Postgres secret consume stayed around 20ms p95.
  • Cycle 46, Thermo Cleanup: split oversized modules and reran gates; mixed-flow c100 improved to 102.3ms p95 in that run, thread list to about 5ms.
  • Cycle 47, Review Fixes: fixed real review findings in checkpoint concurrency, stale thread-index delete, partial bootstrap, and no-op index touch; reran benchmarks and found a turn-lifecycle regression.
  • Cycle 48, Governor/Turn Lock Boundary: released governor and turn commit gates before durable ack waits while preserving enqueue ordering; final c100 mixed p95 105.7ms, governor p95 6.2ms, turn-lifecycle p95 76.0ms.
  • Cycle 49, Turn Blob-to-Row Migration Gate: added the legacy /turns/state.json import gate and stale-blob no-remigrate tests so the stack has a live hosted-volume migration path before go-live.

Verification For This PR

  • cargo check -p ironclaw_filesystem --features libsql,postgres
  • cargo check -p ironclaw_reborn_event_store --features libsql,postgres

Stack-wide verification is listed in the dependent PRs and summarized above.

@ironloopai

ironloopai Bot commented Jul 6, 2026 •

Copy link
Copy Markdown
Contributor

✅ IronLoop Review Status

Head: 94671744d09a806f0b511f16bccd48371e3c9bb1
Result: 1/1 reviewers completed without blocking findings.
Next: Ready for normal human review and CI checks.
Updated: 2026-07-06T21:26:09.475Z

Current reviewers:

Reviewer State Verdict Findings Last update
ironloop/common-reviewer (reviewer) Completed Approved 0 blocking findings / 0 notes 2026-07-06T21:26:09.357Z
Reviewer summaries
Reviewer Detail
ironloop/common-reviewer (reviewer) Approved; 0 blocking findings; Reviewed the filesystem/event-store persistence changes. I did not find a concrete correctness, security, or test-coverage issue that should block the PR.
Recent activity
Time Reviewer State Detail
2026-07-06T21:24:05.844Z ironloop/common-reviewer (reviewer) Queued Accepted review request for head 9467174.
2026-07-06T21:24:05.844Z ironloop/common-reviewer (reviewer) Queued Waiting for this reviewer lane to become available.
2026-07-06T21:24:05.950Z ironloop/common-reviewer (reviewer) Queued Added to the local review work handoff.
2026-07-06T21:24:08.912Z ironloop/common-reviewer (reviewer) Started Reviewer worker started attempt 1.
2026-07-06T21:24:15.259Z ironloop/common-reviewer (reviewer) Workspace ready Prepared isolated checkout (merge_ref) at 82b82b0.
2026-07-06T21:25:57.386Z ironloop/common-reviewer (reviewer) Running Codex is reviewing; process live; elapsed 1m 44s; timeout in 18m 16s; last heartbeat 2026-07-06T21:25:57.386Z. Activity (stderr): ...tReason::SpecMismatch, }); } let fts_key = spec.keys[0].as_str(); // GIN expression index on a to_tsvector(...) ov….
2026-07-06T21:26:09.357Z ironloop/common-reviewer (reviewer) Result captured Approved; 0 blocking findings.
2026-07-06T21:26:09.357Z ironloop/common-reviewer (reviewer) Completed Review completed and terminal status was persisted.
Available commands
  • @ironloop agents
  • @ironloop review
  • @ironloop review --agent <agent-id-or-alias>
  • @ironloop status
Run metadata

Admission: webhook accepted the request and IronLoop persisted review state before this projection.

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ IronLoop Review: reviewer

Verdict: ✅ Approved
Findings: 0 blocking / 0 notes
Next: No reviewer action needed.
Head: 19224b0fe8a7dc212956ce6123990a691cc96246

Run details

Status: Current
Needs human: no
Needs validation: no

**Inline candidates:** 0

Summary

No concrete blocking issues found in the reviewed diff. The changes are scoped to RootFilesystem latency/concurrency improvements, transaction sequence reservation plumbing, and event-store filesystem wrapping. I could not run Rust tests because cargo is not installed in this environment.

Findings

None.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloop review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloop review when the fix may affect multiple areas.
  4. Use @ironloop status to check queued/running/completed/stale/stalled state while reviewers run.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces several robustness and concurrency improvements to the ironclaw filesystem backends. Key changes include refactoring the libSQL backend to use explicit "BEGIN IMMEDIATE" transactions to prevent concurrent write races, implementing advisory locks and schema migration tracking for PostgreSQL, adding support for reserving sequences within storage transactions, and optimizing query performance via cached statements. Feedback on the changes highlights a potential conversion error in the PostgreSQL migration key generation if "current_schema()" returns NULL, which can be resolved by applying a COALESCE fallback.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

.query_one(
"SELECT \
current_database(), \
current_schema(), \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

In PostgreSQL, current_schema() can return NULL if the search path is empty or contains no valid schemas. If it returns NULL, row.get(1) will fail with a conversion error when attempting to retrieve it as a String. Consider using COALESCE(current_schema(), 'public') to ensure a non-null fallback schema name is always returned.

Suggested change
current_schema(), \
COALESCE(current_schema(), 'public'), \

@coderabbitai

coderabbitai Bot commented Jul 6, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Added filesystem transaction support for reserving sequence numbers, including scoped permission and mount-prefix validation.
    • Added a feature-gated helper to build durable event stores from an existing initialized root filesystem.
  • Bug Fixes

    • Improved LibSQL connection initialization by retrying transient connection/PRAGMA failures with exponential backoff.
    • Strengthened concurrent LibSQL directory and record writes to serialize safely and produce deterministic results.
  • Refactor

    • Enhanced PostgreSQL migration handling with advisory locking and deterministic shared projection index cleanup/naming.
    • Optimized database queries and reused SQL helpers for more consistent execution.

Walkthrough

Adds a default-failing reserve_sequence on StorageTxn, wires it through scoped and Postgres backends, hardens LibSQL connection and write transaction handling, updates Postgres migration/index/query behavior, and adds a feature-gated root-filesystem event-store constructor.

Changes

reserve_sequence and filesystem hardening

Layer / File(s) Summary
StorageTxn default reserve_sequence
crates/ironclaw_filesystem/src/backend.rs
Adds default reserve_sequence returning Unsupported/ReserveSeq with path, plus a DummyTxn unit test.
Scoped txn reserve_sequence
crates/ironclaw_filesystem/src/scoped.rs
ScopedStorageTxn implements reserve_sequence with permission and mount-prefix checks before delegating.
libsql connection retry hardening
crates/ironclaw_filesystem/src/libsql.rs
Generalizes retry to cover PRAGMA failures via connect_with_retry_and_pragmas, updates error messages, adds a retry test.
libsql explicit transactions for put/create_dir_all
crates/ironclaw_filesystem/src/libsql.rs
Wraps put/create_dir_all in BEGIN IMMEDIATE/COMMIT/ROLLBACK, extracts put_libsql_inner, create_dir_all_libsql_inner, and connection-taking helper functions.
libsql concurrency contract tests
crates/ironclaw_filesystem/tests/db_root_filesystem_contract.rs
Adds tests spawning 32 concurrent create_dir_all/put tasks to validate transactional serialization.
Postgres migration idempotency and cleanup
crates/ironclaw_filesystem/src/postgres.rs
run_migrations uses advisory locks and a process-wide registry to skip redundant migrations, plus drops legacy indexes.
Postgres shared index naming and query helpers
crates/ironclaw_filesystem/src/postgres.rs
Deterministic shared projection index names, prepare_cached query optimization, shared stat/reserve_sequence helpers, and naming/quoting tests.
Reborn event store constructor
crates/ironclaw_reborn_event_store/src/lib.rs
Adds build_reborn_event_stores_from_root_filesystem to build event stores from an existing RootFilesystem.

Estimated code review effort: 4 (Complex) | ~60 minutes

Possibly related PRs

  • nearai/ironclaw#5451: Both modify crates/ironclaw_filesystem/src/libsql.rs, specifically connection initialization and PRAGMA retry handling.

Suggested reviewers: henrypark133, ilblackdragon

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description is substantial but skips several required template sections, including Change Type, Linked Issue, Security/DB Impact, Rollback Plan, and Review track. Add the missing template sections and fill in the required linkage, validation, impact, rollback, and review-track details.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed It clearly summarizes the main change: the HST Postgres v2 RootFilesystem latency substrate.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Jul 6, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.28% (273440 / 320643 lines)
  floor:    85.3% (tolerance 0.5pp -> effective floor 84.8%)
  denominator: 320643 lines now vs 320188 at floor capture (+455 lines, +0.14%) — not a material change

⚠️ 3 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts, ironclaw_skill_learning

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.28% — 273440 / 320643 lines

Per-crate breakdown (65 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 347
ironclaw_skill_learning 0% 0 / 61
ironclaw_wasm_sandbox_core 7.37% 7 / 95
ironclaw_runtime_policy 33.2% 80 / 241
ironclaw_event_projections 43.34% 673 / 1553
ironclaw_run_state 52.73% 222 / 421
ironclaw_authorization 53.54% 461 / 861
ironclaw_triggers 60.32% 1736 / 2878
ironclaw_observability 61.54% 16 / 26
ironclaw_reborn_cli 64.58% 3988 / 6175
ironclaw_filesystem 65.11% 3449 / 5297
ironclaw_webui_v2 65.47% 2391 / 3652
ironclaw_reborn_migration 67.01% 1172 / 1749
ironclaw_memory 67.12% 747 / 1113
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_mcp 67.42% 569 / 844
ironclaw_trust 72.88% 661 / 907
ironclaw_reborn_event_store 73.01% 944 / 1293
ironclaw_capabilities 74.08% 1658 / 2238
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_extractors 74.72% 538 / 720
ironclaw_first_party_extensions 77.62% 5410 / 6970
ironclaw_llm 77.67% 19044 / 24519
ironclaw_product_context 78.57% 11 / 14
ironclaw_network 79.82% 621 / 778
ironclaw_wasm_product_adapters 80.58% 1510 / 1874
ironclaw_process_sandbox 80.65% 671 / 832
ironclaw_reborn_openai_compat 80.95% 956 / 1181
ironclaw_memory_native 81.86% 3226 / 3941
ironclaw_secrets 82.33% 2716 / 3299
ironclaw_wasm 82.54% 950 / 1151
ironclaw_events 83.44% 1759 / 2108
ironclaw_processes 84.06% 965 / 1148
ironclaw_host_api 84.5% 3119 / 3691
ironclaw_threads 84.9% 3251 / 3829
ironclaw_turns 85.63% 9819 / 11467
ironclaw_slack_v2_adapter 85.82% 1786 / 2081
ironclaw_projects 85.92% 659 / 767
ironclaw_product_workflow 86.26% 10656 / 12354
ironclaw_auth 86.32% 2727 / 3159
ironclaw_common 86.59% 1472 / 1700
ironclaw_product_adapters 86.66% 3152 / 3637
ironclaw_reborn_config 86.98% 1730 / 1989
ironclaw_hooks 87.25% 9782 / 11211
ironclaw_reborn_traces 87.35% 10325 / 11820
ironclaw_skills 87.36% 4335 / 4962
ironclaw_product_adapter_registry 87.96% 526 / 598
ironclaw_extensions 88.26% 2631 / 2981
ironclaw_reborn_identity 88.43% 344 / 389
ironclaw_reborn_composition 89% 68411 / 76866
ironclaw_host_runtime 89.12% 17250 / 19355
ironclaw_conversations 90.11% 2924 / 3245
ironclaw_approvals 90.51% 1507 / 1665
ironclaw_reborn 91.17% 17238 / 18908
ironclaw_event_streams 91.48% 1009 / 1103
ironclaw_reborn_webui_ingress 91.68% 2094 / 2284
ironclaw_loop_support 92.34% 14093 / 15262
ironclaw_attachments 93.06% 630 / 677
ironclaw_telegram_v2_adapter 94.01% 2447 / 2603
ironclaw_resources 94.25% 3625 / 3846
ironclaw_agent_loop 94.49% 8290 / 8773
ironclaw_safety 94.78% 3668 / 3870
ironclaw_first_party_extension_ports 95% 3094 / 3257
ironclaw_outbound 95.59% 3556 / 3720

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (4 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_oauth v1-only: consumed only by root ironclaw (src/auth/oauth.rs); no crates/* dependents. Crate's own doc comment confirms v1-only. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@railway-app

railway-app Bot commented Jul 6, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5724 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 6, 2026 at 9:31 pm

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5724 July 6, 2026 21:24 Destroyed
@ironloopai

ironloopai Bot commented Jul 6, 2026

Copy link
Copy Markdown
Contributor

🗂️ Archived IronLoop Review: reviewer

This result is from an older PR head and is no longer the active review.

Field Value
Status Superseded
Verdict ✅ Approved
Findings 0 blocking / 0 notes
Reviewed head 19224b0fe8a7
Archived summary

No concrete blocking issues found in the reviewed diff. The changes are scoped to RootFilesystem latency/concurrency improvements, transaction sequence reservation plumbing, and event-store filesystem wrapping. I could not run Rust tests because cargo is not installed in this environment.

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ IronLoop Review: reviewer

Verdict: ✅ Approved
Findings: 0 blocking / 0 notes
Next: No reviewer action needed.
Head: 94671744d09a806f0b511f16bccd48371e3c9bb1

Run details

Status: Current
Needs human: no
Needs validation: no

**Inline candidates:** 0

Summary

Reviewed the filesystem/event-store persistence changes. I did not find a concrete correctness, security, or test-coverage issue that should block the PR.

Findings

None.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloop review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloop review when the fix may affect multiple areas.
  4. Use @ironloop status to check queued/running/completed/stale/stalled state while reviewers run.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_filesystem/src/libsql.rs`:
- Around line 1319-1328: The INSERT in the create_dir_all flow is currently
reporting errors against the full target path instead of the per-segment prefix
being written, which makes diagnostics misleading for intermediate failures.
Update the error mapping around the conn.execute call in libsql.rs so it uses
the same prefix-based context as the earlier SELECT/CREATE steps, keeping
libsql_db_error aligned with the current segment being processed rather than
path.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 9b7977f0-fe64-4f66-8462-09897a42cfa3

📥 Commits

Reviewing files that changed from the base of the PR and between ca88418 and 9467174.

📒 Files selected for processing (6)
  • crates/ironclaw_filesystem/src/backend.rs
  • crates/ironclaw_filesystem/src/libsql.rs
  • crates/ironclaw_filesystem/src/postgres.rs
  • crates/ironclaw_filesystem/src/scoped.rs
  • crates/ironclaw_filesystem/tests/db_root_filesystem_contract.rs
  • crates/ironclaw_reborn_event_store/src/lib.rs

Comment on lines +1319 to +1328
conn.execute(
r#"
INSERT INTO root_filesystem_entries (path, contents, is_dir, updated_at)
VALUES (?1, X'', 1, strftime('%Y-%m-%dT%H:%M:%fZ', 'now'))
ON CONFLICT (path) DO NOTHING
"#,
libsql::params![prefix.as_str()],
)
.await
.map_err(|error| libsql_db_error(path.clone(), FilesystemOperation::CreateDirAll, error))?;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟡 Minor | ⚡ Quick win

Per-prefix INSERT error attributes to path, not prefix. The SELECT above (1303/1306/1309) reports errors against prefix, but this INSERT reports against the full target path. In a multi-level create_dir_all, a failure on an intermediate segment will surface the leaf path, misdirecting diagnostics.

🩹 Align error context with the prefix under write
         .await
-        .map_err(|error| libsql_db_error(path.clone(), FilesystemOperation::CreateDirAll, error))?;
+        .map_err(|error| {
+            libsql_db_error(prefix.clone(), FilesystemOperation::CreateDirAll, error)
+        })?;
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
conn.execute(
r#"
INSERT INTO root_filesystem_entries (path, contents, is_dir, updated_at)
VALUES (?1, X'', 1, strftime('%Y-%m-%dT%H:%M:%fZ', 'now'))
ON CONFLICT (path) DO NOTHING
"#,
libsql::params![prefix.as_str()],
)
.await
.map_err(|error| libsql_db_error(path.clone(), FilesystemOperation::CreateDirAll, error))?;
conn.execute(
r#"
INSERT INTO root_filesystem_entries (path, contents, is_dir, updated_at)
VALUES (?1, X'', 1, strftime('%Y-%m-%dT%H:%M:%fZ', 'now'))
ON CONFLICT (path) DO NOTHING
"#,
libsql::params![prefix.as_str()],
)
.await
.map_err(|error| {
libsql_db_error(prefix.clone(), FilesystemOperation::CreateDirAll, error)
})?;
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_filesystem/src/libsql.rs` around lines 1319 - 1328, The
INSERT in the create_dir_all flow is currently reporting errors against the full
target path instead of the per-segment prefix being written, which makes
diagnostics misleading for intermediate failures. Update the error mapping
around the conn.execute call in libsql.rs so it uses the same prefix-based
context as the earlier SELECT/CREATE steps, keeping libsql_db_error aligned with
the current segment being processed rather than path.

@serrrfirat
serrrfirat merged commit ff8079b into main Jul 7, 2026
61 checks passed
@serrrfirat
serrrfirat deleted the codex/hst-postgres-v2-01-rootfs branch July 7, 2026 09:35
henrypark133 added a commit that referenced this pull request Jul 7, 2026
- Delete the pub current_version wrapper: main's #5724 independently
  removed it in favor of the free-function current_version_libsql used
  inside the transactional put_libsql_inner path, confirming it had no
  external caller. The pool-typed current_version_with_conn duplicate
  is dead after adopting that transactional structure; deleted too.
- Add a libsql contract test mirroring
  postgres_put_cas_version_on_missing_path_reports_no_found_version:
  CasExpectation::Version against a missing path must report
  VersionMismatch { found: None }.
- Add a pool checkout-timeout test via a new build_libsql_pool_with_config
  seam (tiny size-1/short-timeout pool) asserting the timeout maps to a
  FilesystemOperation::Connect infrastructure error through connect()'s
  debug!-logged fallback arm.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
henrypark133 added a commit that referenced this pull request Jul 7, 2026
…E_MISUSE (#5466) (#5751)

* fix(filesystem): pool libSQL connections to stop concurrent-CAS SQLITE_MISUSE (#5466)

LibSqlRootFilesystem opened a fresh connection (sqlite3_open_v2 + PRAGMA
batch) for every RootFilesystem operation. Under genuinely parallel CAS
storms against one WAL database that unbounded open/PRAGMA/close churn
intermittently fails inside the C library with SQLITE_MISUSE ("bad
parameter or other API misuse") or spurious disk I/O errors — the ~10%
failure / SIGABRT reported in #5466. A single shared connection is also
wrong: the CAS rows-affected readback is per-connection state, and two
tasks interleaving statements on one connection corrupt compare-and-swap
into silent lost updates (reproduced during diagnosis).

Fix: a bounded deadpool-managed pool (same pooling core the Postgres
backend already uses) in the new libsql_pool module — each operation
checks out one PRAGMA-initialized connection for exclusive use and
returns it on drop; recycle() rejects connections left mid-transaction.
put()'s three CasExpectation arms now drop their checkout before the
nested current_version readback, upholding the documented
one-checkout-per-call-stack invariant.

Regression test: tests/concurrent_cas_storm.rs drives 16 spawned writers
x 100 cas_update increments on a multi-thread runtime against in-memory,
libSQL, and (env-gated) Postgres backends, asserting zero backend errors
and an exact final count. Mutation-verified: re-injecting per-op churn
(recycle always discarding) goes RED with the exact SQLITE_MISUSE
signature; the shared-connection probe goes RED on the lost-update
assertion.

Closes #5466

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(filesystem): address PR #5751 review findings on libSQL pool

- put()'s CAS-mismatch/success version readbacks reuse the already
  checked-out connection via a new current_version_with_conn helper
  instead of drop-then-recheckout, making the one-checkout-per-call-
  stack invariant structural for that call site.
- Add pool-internal tests: recycle rejects a connection returned
  mid-transaction, and connect_with_retry surfaces a Connect error
  with the final cause after exhausting its retry budget.
- Log libSQL pool checkout failures at debug level, mirroring the
  Postgres backend's shape, for trace correlation on checkout
  timeouts.
- Annotate the four silent-skip fallbacks in the Postgres CAS-storm
  test with per-site silent-ok rationale.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(filesystem): address libSQL pool review findings on PR #5751

- Delete the pub current_version wrapper: main's #5724 independently
  removed it in favor of the free-function current_version_libsql used
  inside the transactional put_libsql_inner path, confirming it had no
  external caller. The pool-typed current_version_with_conn duplicate
  is dead after adopting that transactional structure; deleted too.
- Add a libsql contract test mirroring
  postgres_put_cas_version_on_missing_path_reports_no_found_version:
  CasExpectation::Version against a missing path must report
  VersionMismatch { found: None }.
- Add a pool checkout-timeout test via a new build_libsql_pool_with_config
  seam (tiny size-1/short-timeout pool) asserting the timeout maps to a
  FilesystemOperation::Connect infrastructure error through connect()'s
  debug!-logged fallback arm.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(filesystem): cover permanent PRAGMA-failure path in connect_with_retry

Every open succeeds but every PRAGMA batch fails across all retry
attempts; final error must be FilesystemOperation::Connect carrying the
PRAGMA cause. Addresses PR #5751 round-3 review finding.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai coderabbitai Bot mentioned this pull request Jul 10, 2026
9 of 14 tasks

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5724 — 94671744 Deployed Jul 6, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant