Skip to content

fix: serialize concurrent inbound-message writes per thread - #6096

Closed
henrypark133 wants to merge 2 commits into
mainfrom
fix/issue-6047-msg-order
Closed

henrypark133 wants to merge 2 commits into
mainfrom
fix/issue-6047-msg-order

Conversation

@henrypark133

Copy link
Copy Markdown
Collaborator

Summary

Fixes #6047 — two chat messages sent in quick succession to the same thread could be persisted, displayed, and acted on out of chronological order.

Verified bug

Reproduced with a failing test on origin/main @ d68073336 (before this fix), and confirmed the fix turns it green:

cargo test -p ironclaw_threads --test filesystem_session_thread_contract \
  filesystem_accept_inbound_message_serializes_concurrent_calls_for_same_thread -- --nocapture

Root cause

FilesystemSessionThreadService::accept_inbound_message (crates/ironclaw_threads/src/filesystem_service.rs) falls back to a non-transactional path on backends without multi-key transaction support — this is the shape of the production libSQL backend (crates/ironclaw_filesystem/src/libsql.rs has no begin() override). On that path, the function:

  1. reserves a sequence number, then
  2. writes the message record as a separate, unsynchronized put.

Nothing serialized two concurrent accept_inbound_message calls against the same thread, so a message that reserved a later sequence number could finish its write before a message that reserved an earlier one — making the later message durably visible (and readable by the UI/agent) ahead of the earlier one. The transactional path (backends with begin() support) is unaffected: its CAS-retry loop commits reserve+write atomically, and a losing writer retries into a fresh, strictly-later sequence number rather than reusing a stale one.

Fix

  • Added a per-(scope, thread_id) in-process async lock, held only around the ordering-sensitive step (reserve_sequence + a new put_new_message_record — just the message-record CAS put), not the best-effort sequence-index/lookup-index/cache work that follows it.
  • Split write_new_message into put_new_message_record (the CAS put) and finish_new_message_indexes (the rest), so both the lock-guarded fallback path and write_new_message's three other callers share the same code with no duplication.
  • Extracted a shared weak_keyed_lock helper (get-or-create a weak-referenced per-key async lock) used by both the new lock and the pre-existing thread_index_load_lock singleflight cache-fill lock in the same file, instead of duplicating that logic.
  • Renamed one_shot_context_window_cache_key → thread_scope_key since it's now shared by two call sites, with a doc comment distinguishing it from the differently-shaped thread_index_record_cache_key in the neighboring module.
  • The lock is process-local by design — this crate's filesystem-backed services run under a single mounted, single-owner view per instance (matching the with_fixed_view boot-owner model used elsewhere in Reborn), and the pre-existing thread_index_load_lock already relies on the same mechanism. A durable cross-process ordering guarantee would need the write itself to become one atomic domain-level operation (gap-aware reads with a bounded recovery contract, or transactional support on backends that lack it) — a materially larger change tracked as a follow-up, not folded into this fix.

Regression tests

All in crates/ironclaw_threads/tests/filesystem_session_thread_contract.rs:

  • filesystem_accept_inbound_message_serializes_concurrent_calls_for_same_thread — the core proof: message B (called second) must not reserve a sequence or become visible while message A (called first) is still mid-write.
  • filesystem_accept_inbound_message_write_lock_releases_after_failed_write — the lock releases via RAII even when the guarded write errors, so a subsequent call to the same thread doesn't hang forever.
  • filesystem_accept_inbound_message_does_not_serialize_across_different_threads — the lock is scoped per-thread, not globally/per-tenant; unrelated threads never block each other.
  • weak_keyed_lock_reclaims_entry_after_guard_dropped (inline unit test) — the shared lock registry doesn't leak an entry per distinct key forever.

Verification

cargo test -p ironclaw_threads       # 78+8+49+56 = 191 passed, 0 failed
cargo clippy -p ironclaw_threads --tests --all-features   # clean
cargo fmt --check -p ironclaw_threads                      # clean

Review

Went through an 8-agent code review + two thermo-nuclear maintainability passes (plan and final diff). Findings addressed: narrowed the lock's critical section (was originally wider), fixed a lock-accessor signature inconsistency, disambiguated the renamed cache-key helper from a same-purpose-but-different-shape sibling, and added the two missing regression tests above. Security, bugs, maintainability, and approach reviewers returned zero findings.

Deferred (out of scope for this fix)

  • A structurally similar TOCTOU exists between accept_inbound_message (durable, sequenced) and submit_turn's busy-check in crates/ironclaw_product_workflow/src/inbound_turn.rs — a second message can win the "claim the active run" race ahead of an earlier-accepted one. Bigger lift (touches ironclaw_turns' TurnActiveLockKey machinery); flagging for a follow-up issue rather than folding into this PR.
  • append_assistant_draft and append_finalized_assistant_message use the same reserve-then-write two-step pattern for assistant messages on the same thread — out of scope since Task messages are processed and displayed out of chronological order #6047 specifically reports user-submitted messages, but worth a maintainer's awareness.

On the non-transactional fallback path used by the production libSQL
backend, accept_inbound_message reserved a sequence number and wrote the
message record as two separate, unsynchronized steps. Two concurrent
messages to the same thread could commit out of call order, so a later
message could become durably visible (and get acted on) before an
earlier one.

Add a per-(scope, thread_id) in-process lock around just the reserve +
write step (not the best-effort index/cache work after it), sharing the
weak-referenced lock-registry primitive already used by
thread_index_load_lock. Split write_new_message into put_new_message_record
(the ordering-sensitive CAS put) and finish_new_message_indexes (the rest)
so the lock covers only what needs it.

Closes #6047

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 14, 2026 17:09

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6096 July 14, 2026 17:09 Destroyed
@github-actions github-actions Bot added size: M 50-199 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 14, 2026
@coderabbitai

coderabbitai Bot commented Jul 14, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: b7741a06-9e2b-468b-a4c5-240c851d6af2

📥 Commits

Reviewing files that changed from the base of the PR and between d236be5 and a86c579.

📒 Files selected for processing (2)
  • crates/ironclaw_threads/src/filesystem_service.rs
  • crates/ironclaw_threads/tests/filesystem_session_thread_contract.rs

📝 Walkthrough

Summary by CodeRabbit

  • Bug Fixes
    • Improved inbound message visibility and sequencing during concurrent accepts on filesystem backends that don’t support multi-key transactions.
    • Prevented potential hangs when inbound message writes fail partway through (including sequence/index steps).
    • Reduced contention by scoping ordering coordination to the specific thread rather than blocking unrelated threads.
  • Tests
    • Added regression coverage (Issue #6047) with deterministic concurrency to validate ordering, failure-handling, and per-thread lock scoping behavior.

Walkthrough

Changes

Unsupported filesystem backends now serialize sequence reservation and message-record visibility per thread. Shared weak-keyed locks and injective thread keys support the path, while index/cache work remains outside the lock. Tests cover ordering, failures, lock scope, and reclamation.

Inbound message ordering

Layer / File(s) Summary
Shared thread lock infrastructure
crates/ironclaw_threads/src/filesystem_service.rs, crates/ironclaw_threads/src/filesystem_service/thread_index.rs
Adds reusable weak-keyed locks and shared thread-scoped keys for inbound writes, context caching, and thread-index loading.
Message record write separation
crates/ironclaw_threads/src/filesystem_service.rs
Separates CAS-put message visibility from sequence, index, lookup, and cache updates.
Unsupported transaction fallback
crates/ironclaw_threads/src/filesystem_service.rs
Protects sequence reservation and message-record persistence with a per-thread lock, then performs finishing work after release.
Concurrency regression coverage
crates/ironclaw_threads/tests/filesystem_session_thread_contract.rs, crates/ironclaw_threads/src/filesystem_service.rs
Adds deterministic tests for ordering, failure recovery, cross-thread progress, key injectivity, and weak-lock reclamation.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant FilesystemSessionThreadService
  participant ThreadWriteLock
  participant FilesystemBackend
  participant IndexAndCache
  Client->>FilesystemSessionThreadService: accept_inbound_message
  FilesystemSessionThreadService->>ThreadWriteLock: acquire thread-scoped lock
  ThreadWriteLock->>FilesystemBackend: reserve sequence
  ThreadWriteLock->>FilesystemBackend: CAS-put message record
  FilesystemBackend-->>FilesystemSessionThreadService: write result
  FilesystemSessionThreadService->>IndexAndCache: finish indexes and invalidate cache
  FilesystemSessionThreadService-->>Client: acceptance result
Loading

Possibly related issues

Possibly related PRs

Suggested reviewers: copilot, ilblackdragon, serrrfirat

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed Conventional Commits-style title accurately summarizes the fix for per-thread serialization on inbound message writes.
Description check ✅ Passed The PR description is mostly complete and covers summary, root cause, fix, tests, verification, and follow-up items.
Linked Issues check ✅ Passed The patch satisfies #6047 by serializing same-thread inbound writes on the non-transactional path so messages stay chronological.
Out of Scope Changes check ✅ Passed No material out-of-scope code changes are evident; the lock helper, cache-key rename, and tests directly support the ordering fix.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request addresses Issue #6047 by introducing a process-local, weak-keyed async lock (inbound_message_write_locks) to serialize the sequence reservation and message write steps for non-transactional backends, preventing out-of-order message visibility. It refactors message writing to isolate the ordering-sensitive step and adds robust regression tests. The review feedback suggests improving the robustness of thread_scope_key by using a length-prefixed encoding to prevent separator collisions, and extending the per-thread write lock to other message types (such as assistant drafts and tool results) to ensure complete write ordering across all roles.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +3084 to 3090
fn thread_scope_key(scope: &ThreadScope, thread_id: &ThreadId) -> String {
format!(
"{}:{}",
scope.tenant_id.as_str(),
thread_root_string(scope, thread_id)
)
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security-high high

The thread_scope_key function constructs a cache and lock key by simply concatenating scope.tenant_id and thread_root_string with a : delimiter. This introduces a separator collision risk (e.g., if a tenant ID contains a : or if there is any overlap in the dynamic string components), which could lead to cross-tenant cache pollution or lock contention.

To eliminate this risk and ensure complete injectivity, use a length-prefixed encoding for each component as recommended by the general rules.

fn thread_scope_key(scope: &ThreadScope, thread_id: &ThreadId) -> String {
    let tenant = scope.tenant_id.as_str();
    let root = thread_root_string(scope, thread_id);
    format!("{}:{}:{}:{}", tenant.len(), tenant, root.len(), root)
}
References
  1. When generating deterministic identifiers or hashes from multiple string components, use an injective length-prefixed encoding rather than simple concatenation with a delimiter to eliminate the risk of separator-collision attacks or accidental collisions.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed real — validate_scope_id forbids / and \\ but not :. Fixed in a86c579: thread_scope_key now length-prefixes each component. One nuance worth recording: I verified the specific collision isn't reachable through the current concrete ThreadId/TenantId shapes (thread_root_string always starts with a fixed /threads prefix, and / is forbidden in every id component, so the delimiter can't actually be repositioned today) — so this was defense-in-depth against a fragile implicit invariant rather than a live exploit, but the fix is correct either way and a regression test now proves the new encoding's injectivity directly rather than relying on that invariant.

Comment on lines +473 to +484
async fn write_new_message(
&self,
scope: &ThreadScope,
thread_id: &ThreadId,
message: &ThreadMessageRecord,
description: &'static str,
) -> Result<(), SessionThreadError> {
self.put_new_message_record(scope, thread_id, message, description)
.await?;
self.finish_new_message_indexes(scope, thread_id, message)
.await
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

While this PR successfully serializes user messages in accept_inbound_message, other message types (such as assistant drafts, finalized assistant messages, and tool result references) still use the un-serialized reserve_sequence followed by write_new_message pattern.

If a user message is sent concurrently with an assistant response or tool result write on the same thread, they can still race and be written/displayed out of chronological order. To ensure complete thread-level write ordering and prevent cross-role message reordering, consider acquiring the per-thread write lock around the reserve-then-write sequence for assistant and tool messages as well.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed real and in-scope for a follow-up, not this PR — #6047 is specifically about user messages (that's the reported bug and what the regression tests prove), and append_assistant_draft/append_finalized_assistant_message/append_tool_result_reference do share the same unserialized two-step pattern. Filed #6101 to extend the same lock mechanism to those call sites.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d236be5a6c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1757 to 1758
self.finish_new_message_indexes(&scope, &thread_id, &message)
.await?;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep one-shot cache updates inside the thread write lock

On non-transactional backends this runs after _write_guard is dropped, but finish_new_message_indexes invalidates the one-shot context cache and the sequence == 1 branch below may re-seed it. If message A (seq 1) stalls in the sequence-index/idempotency work after dropping the lock, message B (seq 2) can complete and invalidate the cache, then A resumes and seeds a one-message context; the next load_context_window takes that stale cache and omits B even though B is already durable. Keep the cache seed/invalidate (or the full finish) serialized with the per-thread write order.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified this is real, but it predates this PR — the if sequence == 1 { seed } else { invalidate } block already existed on origin/main unchanged, outside any lock/transaction, for both the transactional and non-transactional write paths. The #6047 fix's lock-narrowing (from the first review round) didn't open this window since that decision point was never covered by any lock before or after. Filed #6100 as its own follow-up rather than folding it into this PR's scope.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_threads/src/filesystem_service.rs`:
- Around line 3096-3112: Amortize stale-lock cleanup in weak_keyed_lock instead
of calling retain on every lock acquisition. Preserve reuse and insertion of
keyed locks, but trigger periodic cleanup based on an established counter or
time interval so accept_inbound_message avoids an O(n) map scan on each call.
- Around line 3096-3112: Fix weak_keyed_lock in
crates/ironclaw_threads/src/filesystem_service.rs:3096-3112 to recover the
poisoned Mutex guard via into_inner instead of returning a fresh unshared lock,
preserving shared mutual exclusion for all keys. No direct change is needed at
thread_index_load_lock in
crates/ironclaw_threads/src/filesystem_service/thread_index.rs:497-499; it is
corrected automatically by the shared helper fix.
- Around line 193-219: Add the repository-required arch-exempt annotation for
the process-local lock used by inbound_message_write_locks and _write_guard
across reserve_sequence and put_new_message_record I/O. Include a valid plan or
follow-up tracking link documenting the intended durable/CAS-based replacement,
while preserving the existing lock scope and behavior.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 40187cfc-68e9-4f41-82ab-a7a33268afc6

📥 Commits

Reviewing files that changed from the base of the PR and between 15f06e1 and d236be5.

📒 Files selected for processing (3)
  • crates/ironclaw_threads/src/filesystem_service.rs
  • crates/ironclaw_threads/src/filesystem_service/thread_index.rs
  • crates/ironclaw_threads/tests/filesystem_session_thread_contract.rs

Comment thread crates/ironclaw_threads/src/filesystem_service.rs
Comment thread crates/ironclaw_threads/src/filesystem_service.rs
@github-actions

github-actions Bot commented Jul 14, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.59% (303487 / 354600 lines)
  floor:    85.3% (tolerance 0.5pp -> effective floor 84.8%)
  denominator: 354600 lines now vs 320188 at floor capture (+34412 lines, +10.75%) — material change (>5%)

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.59% — 303487 / 354600 lines

Per-crate breakdown (63 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_runtime_policy 31.75% 80 / 252
ironclaw_event_projections 43.31% 673 / 1554
ironclaw_run_state 53.07% 225 / 424
ironclaw_authorization 53.89% 464 / 861
ironclaw_triggers 59.89% 1792 / 2992
ironclaw_observability 61.54% 16 / 26
ironclaw_webui_v2 62.96% 2682 / 4260
ironclaw_mcp 63.03% 578 / 917
ironclaw_reborn_cli 66.18% 4488 / 6781
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_filesystem 67.69% 3932 / 5809
ironclaw_memory 69.2% 773 / 1117
ironclaw_reborn_migration 71.57% 1551 / 2167
ironclaw_trust 72.88% 661 / 907
ironclaw_capabilities 74.39% 1685 / 2265
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_reborn_event_store 74.67% 958 / 1283
ironclaw_extractors 74.72% 538 / 720
ironclaw_llm 78.36% 20328 / 25941
ironclaw_product_context 78.57% 11 / 14
ironclaw_first_party_extensions 78.81% 5576 / 7075
ironclaw_process_sandbox 80.65% 671 / 832
ironclaw_wasm_product_adapters 80.71% 1448 / 1794
ironclaw_memory_native 81.22% 3205 / 3946
ironclaw_secrets 82.7% 2791 / 3375
ironclaw_events 82.86% 1765 / 2130
ironclaw_reborn_identity 83.59% 433 / 518
ironclaw_wasm 83.97% 1011 / 1204
ironclaw_auth 83.99% 3147 / 3747
ironclaw_reborn_config 84.06% 1814 / 2158
ironclaw_processes 84.44% 993 / 1176
ironclaw_common 84.85% 1490 / 1756
ironclaw_turns 85% 13676 / 16090
ironclaw_host_api 85.13% 2663 / 3128
ironclaw_product_workflow 85.78% 10886 / 12691
ironclaw_projects 85.92% 659 / 767
ironclaw_network 86.12% 670 / 778
ironclaw_slack_v2_adapter 86.79% 1806 / 2081
ironclaw_threads 86.89% 4651 / 5353
ironclaw_product_adapters 87.18% 3265 / 3745
ironclaw_skills 87.58% 4470 / 5104
ironclaw_hooks 87.78% 9921 / 11302
ironclaw_product_adapter_registry 88.06% 531 / 603
ironclaw_reborn_traces 88.19% 11946 / 13546
ironclaw_host_runtime 88.59% 17395 / 19635
ironclaw_reborn_composition 89.29% 80244 / 89866
ironclaw_extensions 89.38% 2971 / 3324
ironclaw_approvals 89.41% 1587 / 1775
ironclaw_runner 89.5% 16989 / 18983
ironclaw_reborn_openai_compat 89.55% 3798 / 4241
ironclaw_conversations 90.33% 3121 / 3455
ironclaw_event_streams 90.82% 1009 / 1111
ironclaw_loop_host 92.25% 15043 / 16307
ironclaw_resources 92.69% 5134 / 5539
ironclaw_attachments 93.06% 630 / 677
ironclaw_reborn_webui_ingress 93.19% 2217 / 2379
ironclaw_telegram_v2_adapter 93.62% 2511 / 2682
ironclaw_agent_loop 94.8% 9200 / 9705
ironclaw_safety 95.04% 3677 / 3869
ironclaw_first_party_extension_ports 95.24% 3343 / 3510
ironclaw_outbound 95.59% 3556 / 3720

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@railway-app

railway-app Bot commented Jul 14, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6096 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 14, 2026 at 9:25 pm

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review (multi-agent)

Intent: Serialize concurrent inbound-message writes per thread to preserve chronological persistence, display, and processing order on non-transactional backends.

Stats: 1 finding (from 1 retained finding) across 1 file; 8 review lenses completed using the exact-head snapshot, with parent-context fallback after the environment rejected parallel reviewer-agent dispatch at the thread limit. Body-only: 0. Existing current-head threads were deduplicated.

Tests

  1. Medium Missing sequence-index failure coverage (crates/ironclaw_threads/src/filesystem_service.rs:490-498, confidence 80) — anchor: crates/ironclaw_threads/src/filesystem_service.rs:496

    The non-transactional fallback can return a sequence-index write error after the message record is already durable and the write lock has been released. The current regression tests cover message-record failure and best-effort lookup-index failure, but do not exercise this sequence-index failure path or prove that a subsequent same-thread inbound message still proceeds.

    Fix: Add filesystem_accept_inbound_message_sequence_index_failure_releases_lock, injecting a sequence-index failure after the record succeeds, then asserting the next same-thread inbound message completes.

Existing unresolved current-head coverage not duplicated

The live PR already has unresolved threads for the process-local mutex across backend I/O, stale one-shot cache reseeding, delimiter-based thread-key collisions, the O(n) lock-registry sweep, and poisoned-lock fallback behavior. Those were excluded from this review to avoid duplicate feedback.

/// CAS put — not ordering-sensitive, so callers that need the put itself
/// under a narrower critical section (the issue #6047 write lock in
/// `accept_inbound_message`) run this afterward, unlocked.
async fn finish_new_message_indexes(

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — Missing sequence-index failure coverage.

The non-transactional fallback propagates finish_new_message_indexes errors after the message record is durable and after the write lock is released. Existing tests cover message-record failure and best-effort lookup-index failure, but none inject a sequence-index write failure to verify the caller-level error and subsequent same-thread call behavior.

Fix: tests::filesystem_accept_inbound_message_sequence_index_failure_releases_lock covering a sequence-index write error after the message record succeeds, followed by a successful same-thread inbound message

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added filesystem_accept_inbound_message_sequence_index_failure_releases_lock in a86c579, following the same pattern as the existing message-record-failure test — injects a sequence-index write failure after the record succeeds, asserts the caller gets Err, then asserts a same-thread follow-up message completes (proving the lock, already released before this failure point, doesn't wedge).

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up finding

A delayed concurrency review completed after the initial batch and identified one additional issue not covered by the existing current-head threads. The review remains scoped to exact head d236be5a6c10132d6c5bca8c0de75dc2889e8b44.

// which doesn't affect visibility and doesn't need to wait.
{
let write_lock =
self.inbound_message_write_lock(&thread_scope_key(&scope, &thread_id));

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

High — Lock is not shared across service instances.

The registry is stored inside each FilesystemSessionThreadService, so two service instances sharing the same ScopedFilesystem or backend have independent locks. Concurrent inbound calls through those instances can still reserve and write sequences out of order, making the per-thread ordering guarantee depend on an undocumented singleton-service assumption.

Fix: Move the per-thread lock registry into shared filesystem/backend state, or otherwise ensure every wrapper over a store shares the same registry.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This one I can't fully resolve with confidence from static reading — traced the wiring in ironclaw_reborn_composition/src/factory.rs and confirmed it's genuinely not a simple singleton (two separate build functions each construct their own instance, and there's 'reopen' language suggesting a store graph can be rebuilt mid-process). I don't have enough visibility into that reopen path's draining/lifecycle guarantees to say whether overlap is actually possible. Strengthened the doc comment on inbound_message_write_locks in a86c579 to state the assumption explicitly rather than leave it implicit, and filed #6102 to get a real answer on the reopen path before deciding whether the registry needs to move into shared backend state.

- weak_keyed_lock: recover a poisoned std::sync::Mutex instead of
  fabricating a fresh, unshared lock (was silently disabling mutual
  exclusion for every key once poisoned — coderabbit).
- weak_keyed_lock: skip the retain() sweep on the live-key hit path so
  repeated lookups for an already-live key stay O(1) instead of
  rescanning the whole map on every accept_inbound_message call
  (coderabbit).
- thread_scope_key: length-prefix each component instead of naive `:`
  concatenation, so it's injective regardless of what characters end up
  inside a tenant/thread-id component (gemini-code-assist; verified
  validate_scope_id permits `:`, though not reachable through the
  current concrete ThreadId/TenantId shapes today).
- Add filesystem_accept_inbound_message_sequence_index_failure_releases_lock
  covering the sequence-index-write-failure path, alongside the existing
  message-record-failure coverage (review feedback from henrypark133).
- Document the inbound_message_write_locks singleton-instance assumption
  more precisely and flag it as an open verification item against the
  composition factory's "reopen" store-graph path (henrypark133, High).

Two items triaged as out-of-scope follow-ups rather than folded in here
(reasoning posted on the PR): the one-shot context-window cache's
seed/invalidate race, which predates this PR and affects both storage
backends; and extending per-thread write serialization to assistant/
tool-result message writes, which is a real but separately-scoped gap
(#6047 is specifically about user messages).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 14, 2026 21:13
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6096 July 14, 2026 21:13 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@henrypark133

Copy link
Copy Markdown
Collaborator Author

Closing as stale — no activity in over three weeks. The branch is untouched; reopen if this is still needed.

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6096 — a86c579b Deployed Jul 14, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Task messages are processed and displayed out of chronological order

2 participants