Skip to content

fix(loop): preserve pageable result_read continuation references - #7135

Merged
think-in-universe merged 9 commits into
mainfrom
codex/fix-result-read-continuation-ref
Aug 6, 2026
Merged

think-in-universe merged 9 commits into
mainfrom
codex/fix-result-read-continuation-ref

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Aug 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Return the original pageable result reference from builtin.result_read instead of exposing the fresh inline-only staging write reference.
  • Preserve inline chunk persistence, output digests/evidence, byte offsets, preview sanitization, and replay metadata.
  • Key durable tool-result dedup by provider call identity so distinct pages can share one source reference while exact replays remain idempotent.
  • Pin subagent settlement updates to the original spawn provider call so later paged-read rows sharing the reference cannot be overwritten.
  • Add a two-page caller-path regression plus in-memory and filesystem dedup contracts.

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Related #5838

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings — passed in CI via Clippy (all-features); focused all-target/all-feature clippy also passed for both changed crates.
  • cargo build — covered by the focused test and clippy builds; no separate workspace build.
  • Relevant tests pass: exact two-page result_read integration (1), in-memory result-reference contracts (9), filesystem result-reference contracts (3), spawn identity (1), await-edge suite (10), focused composition cancellation test (1), and all affected Reborn buckets/integration suites on final head 048c4c348.
  • cargo test --features integration if database-backed or integration behavior changed — not applicable: no database schema/backend integration changed; filesystem persistence is covered directly.
  • Manual testing: Railway preview passed on exact final head 048c4c3482968ca1c215748e55853cbb7f473cd1; two expanded continuation cards reused the original reference at offsets 24576 and 49152, replay survived refresh, and cleanup was verified.
  • If a coding agent was used and supports it, review-pr or pr-shepherd --fix was run before requesting review — not run; normal review remains pending.

Test Strategy

User behavior:

When a large durable tool result requires multiple result_read calls, the model receives one usable continuation identity: the original pageable result reference. Feeding that reference and next_offset into the next call returns the next byte-exact page without replay or dedup conflicts.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: in-memory and filesystem thread contracts prove exact provider-call replay deduplicates, distinct page calls sharing one result reference persist separately, and settlement updates only the original provider-call row.
  • Reborn integration: the existing result_read scenario now performs two pages and feeds page one's surfaced reference and offset verbatim into page two.
  • Recorded fixture: Not applicable: deterministic scripted calls prove the continuation contract without depending on model tool selection.
  • Browser E2E: Railway preview PASS through authenticated WebUI chat; two continuation calls reused one original result reference, the terminal marker rendered, and the deep-linked conversation survived refresh.
  • Backend or runtime: filesystem transcript lookup and subagent await-edge propagation are covered by the filesystem contract, spawn identity test, and runner await-edge suite.
  • Live canary: Railway preview PASS on the exact PR head; deterministic tests remain the authoritative byte-exact regression evidence.

What the tests prove:

The completed tool outcome and persisted replay envelope retain the original durable reference; each page has the correct offsets, total length, digest/evidence path, and sanitized observation; exact call replay remains idempotent while two provider calls may share one continuation reference. Railway additionally proves the production WebUI caller can reuse the same reference across offsets 24576 and 49152, render the terminal marker, and replay the completed history after refresh.

Commands run:

cargo fmt --all -- --check
git diff origin/main...HEAD --check
cargo test -p ironclaw_reborn_composition cancel_run_propagates_to_subagent_children --lib
RUST_MIN_STACK=16777216 CARGO_INCREMENTAL=0 CARGO_PROFILE_DEV_DEBUG=0 CARGO_PROFILE_TEST_DEBUG=0 cargo test -j 2 -p ironclaw_threads --test session_thread_contract append_tool_result_reference -- --nocapture
RUST_MIN_STACK=16777216 CARGO_INCREMENTAL=0 CARGO_PROFILE_DEV_DEBUG=0 CARGO_PROFILE_TEST_DEBUG=0 cargo test -j 2 -p ironclaw_threads --test filesystem_session_thread_contract filesystem_tool_result -- --nocapture
cargo test -p ironclaw_loop_host invoke_spawn_submits_child_run_through_spawn_tree_port --lib -- --nocapture
cargo test -p ironclaw_runner await_edge --lib -- --nocapture
RUST_MIN_STACK=16777216 CARGO_INCREMENTAL=0 CARGO_PROFILE_DEV_DEBUG=0 CARGO_PROFILE_TEST_DEBUG=0 cargo test -j 2 -p ironclaw_reborn_integration_tests --test reborn_integration_tool_call result_read -- --nocapture
CARGO_INCREMENTAL=0 CARGO_PROFILE_DEV_DEBUG=0 CARGO_PROFILE_TEST_DEBUG=0 cargo clippy -j 2 -p ironclaw_loop_host -p ironclaw_threads --all-targets --all-features -- -D warnings
CARGO_INCREMENTAL=0 CARGO_PROFILE_DEV_DEBUG=0 CARGO_PROFILE_TEST_DEBUG=0 cargo test -j 2 -p ironclaw_architecture
/Users/firatsertgoz/.codex/skills/railway-test/scripts/preview_state.sh 7135 nearai/ironclaw

Railway browser evidence for exact final head: #7135 (comment)

Security Impact

None. Permission, authorization, redaction, secret mediation, sandboxing, and external network behavior are unchanged. Credential-like inline previews remain suppressed.

Reborn Trust-Boundary Checklist

  • Public policy/evidence/trust-bearing types: no new public types; the existing validated LoopResultRef now carries the original durable continuation authority through completion.
  • Untrusted content enters prompts only through an envelope/escaping primitive: existing observation validation, text sanitization, and credential-preview suppression remain in place.
  • Hashes declare purpose; trust/binding/authenticity uses SHA-256/BLAKE3 or separate authenticity check: existing output digest/evidence generation is preserved unchanged.
  • New/changed status, exit, policy, runtime, or error variants: none added or changed.
  • Security/durability serde(default) fields fail closed or have migration tests: the additive optional spawn provider-call identity defaults to legacy lookup only for pre-existing durable edges; fresh edges persist exact identity, and reconstruction tests cover both shapes.
  • Queues/maps/buffers/counters have bounds and overflow-safe arithmetic: no loop budgets, counters, or buffer limits changed; existing bounded page limits remain intact.
  • Driver/operator-visible errors have stable class semantics: no error class changed; validated-reference reconstruction maps only impossible internal drift to Internal.
  • Sandbox/native/host names accurately describe trust boundary: no runtime lane or trust-boundary naming changed.

Database Impact

None. No database migration or schema change. Filesystem transcript lookup adds provider-call-specific index keys alongside the existing generic index and retains compatibility fallback for pre-existing rows.

Blast Radius

builtin.result_read completion identity and thread transcript result-reference indexing. A regression could affect multi-page model continuations or replay lookup; inline result persistence, bytes, offsets, and output evidence stay on the existing writer path.

Rollback Plan

Revert the review-fix commit, then the original continuation-reference commit if a full rollback is required. Existing data requires no migration rollback: generic result-reference indexes remain written and readable, and provider-call-specific indexes are additive derived lookup entries.

Review Follow-Through

Please focus on the provider-call-specific dedup key and legacy generic-index fallback, especially the distinction between exact replay and a distinct page invocation sharing the same durable source reference. Railway exact-head caller-path and refresh-replay evidence is recorded in the PR conversation.


Review track: C (runtime/persistence behavior)

Return the original pageable result reference from result_read while retaining inline-only chunk persistence and evidence. Key transcript dedup by provider call so multiple pages can safely share one durable source reference, with regression coverage for replay and a two-page continuation.
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@railway-app

railway-app Bot commented Aug 4, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-7135 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Aug 6, 2026 at 10:03 am

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7135 August 4, 2026 10:53 Destroyed
@github-actions github-actions Bot added size: L 200-499 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Aug 4, 2026
@ironloopai

ironloopai Bot commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #7135

🟢 Completed · Review submitted

Submitted review →

Reviewed the complete trusted base-to-head comparison. The change consistently preserves the original pageable result reference, distinguishes transcript entries by provider call identity, retains legacy filesystem lookup compatibility, and adds appropriate backend and caller-path regressions. No actionable findings identified.

Automatic · PR opened · attempt 1 of 3 · completed in 1m 46s

Run details
  • Repository: nearai/ironclaw
  • Base: main at 1e2a294
  • Head: codex/fix-result-read-continuation-ref at 11eda63
  • Created: Aug 4, 2026, 10:58 AM UTC
  • Updated: Aug 4, 2026, 10:59 AM UTC
  • Run: 12b00d73-29af-4fb0-b1e9-54fc158a48bd
  • Latest attempt: 1 · Completed · aed616b8-ab64-4ad0-a241-d8aa3864cdf6

@coderabbitai

coderabbitai Bot commented Aug 4, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: de214d8a-8315-44ce-8605-6e7972f31959

📥 Commits

Reviewing files that changed from the base of the PR and between 362b462 and 70d962c.

📒 Files selected for processing (1)
  • tests/integration/subagent_await_edge.rs

📝 Walkthrough

Summary by CodeRabbit

  • Bug Fixes

    • Improved paginated result continuation, preserving references across pages and ensuring contiguous, byte-exact content with accurate total sizes.
    • Prevented credential-bearing previews from being persisted or exposed.
    • Improved updates and replay when multiple provider calls share a result reference.
    • Preserved provider-call identity during subagent execution and recovery.
  • Tests

    • Expanded coverage for pagination, persistence, deduplication, replay, and provider-call-specific updates.

Walkthrough

Changes

Durable result flow

Layer / File(s) Summary
Preserve continuation references
crates/loop/ironclaw_loop_host/src/result_read.rs, crates/app/ironclaw_composition/src/runtime/..., tests/integration/tool_call.rs
result_read returns the original durable reference. Tests verify staged output, offsets, byte-exact pages, and persisted metadata.
Provider-call lookup indexes
crates/domains/ironclaw_threads/src/contract.rs, crates/domains/ironclaw_threads/src/filesystem_service/...
Tool-result requests accept an optional provider-call ID. Lookup uses provider-specific indexes with legacy fallback support.
Provider-aware persistence and validation
crates/domains/ironclaw_threads/src/filesystem_service.rs, crates/domains/ironclaw_threads/src/in_memory.rs
Append and update operations validate provider metadata and target matching records during deduplication and CAS updates.
Provider-call contract coverage
crates/domains/ironclaw_threads/tests/*
Contract tests cover exact-row updates, same-call replay, distinct calls with shared result references, and metadata conflicts.

Subagent provider-call identity

Layer / File(s) Summary
Authorize and record spawn identity
crates/loop/ironclaw_loop_host/src/subagent_spawn_port.rs, crates/loop/ironclaw_loop_host/src/subagent_spawn_port/tests.rs
Spawn authorization stores and validates provider-call IDs, then forwards them to child metadata and awaited-child records.
Propagate identity through await edges
crates/loop/ironclaw_turn_runner/src/subagent/await_edge/*, crates/app/ironclaw_composition/src/runtime/tests/core.rs, tests/integration/subagent_await_edge.rs
AwaitEdge and legacy metadata preserve the optional provider-call ID. Reconstruction and settlement updates pass it into parent tool-result reference updates.

Estimated code review effort: 4 (Complex) | ~60 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant result_read
  participant ThreadHistory
  participant AwaitEdge
  Client->>result_read: request result page with durable reference and offset
  result_read->>ThreadHistory: resolve and persist result-read metadata
  ThreadHistory-->>result_read: return durable reference and next offset
  result_read-->>Client: return page and continuation metadata
  AwaitEdge->>ThreadHistory: update parent tool-result reference with provider call ID
Loading

Suggested reviewers: ilblackdragon

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title follows Conventional Commits style and accurately identifies the pageable result_read continuation-reference fix.
Description check ✅ Passed The description follows the repository template and documents scope, validation, risks, security, database impact, rollback, and review follow-through.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Review complete · PR #7135

✅ No actionable findings

Reviewed the complete trusted base-to-head comparison. The change consistently preserves the original pageable result reference, distinguishes transcript entries by provider call identity, retains legacy filesystem lookup compatibility, and adds appropriate backend and caller-path regressions. No actionable findings identified.

Validation and technical details
  • Confirmed comparison refs: base 1e2a294 to head 11eda63.
  • Inspected all seven changed files and surrounding result-writing, transcript lookup, update, replay, indexing, and model-observation code.
  • git diff --check completed successfully.
  • Checked changed production additions for prohibited unwrap/expect calls, unsafe code, and hardcoded temporary paths; none found.
  • Focused Cargo tests could not be executed because cargo is unavailable in the review environment (/bin/bash: cargo: command not found).
  • Base: main
  • Head: codex/fix-result-read-continuation-ref at 11eda63
  • Run: 12b00d73-29af-4fb0-b1e9-54fc158a48bd

@serrrfirat
serrrfirat marked this pull request as ready for review August 4, 2026 11:34
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_loop_host/src/result_read.rs`:
- Around line 242-248: Update the LoopResultRef::new call in result_read to
preserve and propagate the original validation error while adding contextual
information; replace the map_err closure’s discarded error binding with an
error-aware conversion. Keep the existing AgentLoopHostError classification and
message context, and do not introduce a separate validation path.

In `@crates/ironclaw_threads/src/filesystem_service.rs`:
- Around line 3294-3307: The provider-keyed lookup and legacy fallback currently
share matches_tool_result_reference_invocation, allowing fallback matches on
rows with provider metadata. Split the predicates in filesystem_service.rs at
matches_tool_result_reference_invocation and the corresponding in_memory.rs
site: keep provider_call_id matching for the keyed lookup, but make the fallback
match only records whose tool_result_provider_call is absent.
- Line 2202: Update update_tool_result_reference to preserve and use the
specific message_id returned by append_tool_result_reference for each tool
result/reference row, rather than performing a generic lookup with no
provider-call ID. Ensure the read-modify-write follows the shared bounded CAS
update path and cannot target another row when multiple writers share the same
result reference. Add a regression test covering multiple writers and verifying
each summary update applies to its original entry.

In `@crates/ironclaw_threads/tests/filesystem_session_thread_contract.rs`:
- Around line 122-138: Extend the filesystem session thread contract test to
assert each retained ToolResultReference row’s provider_call_id matches its
originating provider call, not just message identity and count. Add a mixed case
appending page one with no provider call and page two with Some(call_2), then
verify both rows remain distinct and retain their correct provider metadata
through list_thread_history. Use the existing legacy-fallback scenario and
history assertions as the test location.

In `@tests/integration/tool_call.rs`:
- Around line 736-746: Update the envelope selection in the integration test
around persisted_tool_result_envelopes() to filter entries by
envelope.result_ref == result_ref, rather than taking the last two rows. Assert
that exactly two matching result_read envelopes exist, then select page one and
page two using their stable order while preserving the existing assertions.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: e977d814-daa1-44b2-a745-2175475f957a

📥 Commits

Reviewing files that changed from the base of the PR and between 1e2a294 and 11eda63.

📒 Files selected for processing (7)
  • crates/ironclaw_loop_host/src/result_read.rs
  • crates/ironclaw_threads/src/filesystem_service.rs
  • crates/ironclaw_threads/src/filesystem_service/message_lookup_index.rs
  • crates/ironclaw_threads/src/in_memory.rs
  • crates/ironclaw_threads/tests/filesystem_session_thread_contract.rs
  • crates/ironclaw_threads/tests/session_thread_contract.rs
  • tests/integration/tool_call.rs

Comment thread crates/loop/ironclaw_loop_host/src/result_read.rs
Comment thread crates/ironclaw_threads/src/filesystem_service.rs Outdated
Comment thread crates/domains/ironclaw_threads/src/filesystem_service.rs
Comment thread tests/integration/tool_call.rs Outdated
@ironloopai

ironloopai Bot commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #7135

⚫ Cancelled · Target changed

The target changed before this Run could finish.

Automatic · PR opened · attempt 0 of 3 · cancelled after 1s

Run details
  • Repository: nearai/ironclaw
  • Base: main at 1e2a294
  • Head: codex/fix-result-read-continuation-ref at 11eda63
  • Created: Aug 4, 2026, 11:39 AM UTC
  • Updated: Aug 4, 2026, 11:39 AM UTC
  • Run: 6cae1ad3-6824-4249-a061-eba27a175774

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Railway preview QA — PASS

Tested the exact deployed PR head:

Given / When / Then evidence

Case Evidence Result
Multi-page durable result Created a non-sensitive 1,500-line fixture through the preview, then invoked read_file. The durable result reported 67,053 bytes, truncated at 24,576, and surfaced a continuation reference. PASS
First continuation The first expanded builtin__result_read card succeeded with max_bytes=24576, offset=24576, and the original read_file result reference. PASS
Second continuation using surfaced identity The second expanded builtin__result_read card succeeded with max_bytes=24576, offset=49152, and the same result reference byte-for-byte as the first call. No fresh inline-only reference was exposed. PASS
Terminal caller-visible outcome The rendered response reported 2 result_read continuation calls and contained both line-0000 and terminal marker line-1499. PASS
Durable replay/read-back Reloaded the deep-linked conversation. History restored the final markers and the five-tool activity chain (shell, read_file, two builtin__result_read calls, final ranged read_file). PASS
Cleanup Deleted only railway-test-result-read-7135.txt; the preview confirmed it no longer existed and no other files were touched. PASS

Exact regression result

The production WebUI caller successfully reused one original pageable identity across two distinct result_read provider calls with increasing offsets. Both calls persisted and replayed after refresh, directly exercising the continuation-reference and provider-call-aware dedup behavior in this PR.

Exploratory attempts not counted as evidence

  • An initial fixture setup attempted Python, which is unavailable in the preview; that run was cancelled before the target path was exercised.
  • A direct large shell-output case completed from its safe preview and correctly made no result_read call, so it was excluded from the regression evidence.

Skipped as unrelated

Permissions/role denial, responsive layout, uploads, auth lifecycle beyond preview login, and streaming cadence were not tested because this PR changes backend result-reference continuation and durable replay behavior, not those surfaces.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7135 August 4, 2026 12:15 Destroyed
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

CI follow-up for commit 2076227b8:

  • Root cause: four composition tests still treated result_read's surfaced continuation ref as the fresh inline-only write ref. The corrected contract intentionally surfaces the original pageable ref, while retaining a distinct internal inline evidence write.
  • Updated the tests to assert that the completed/model-visible ref equals the original durable ref, retrieve inline evidence independently, and continue verifying that the internal chunk write is not durably persisted.

Local evidence:

  • cargo fmt --check — pass
  • cargo test -j 2 -p ironclaw_reborn_composition standalone_result_read -- --nocapture — pass (5/5)
  • cargo test -j 2 -p ironclaw_reborn_composition standalone_runtime_safe_preview_observer_receives_bounded_payload -- --nocapture — pass (1/1)
  • GitHub affected-2-equivalent hermetic bucket (composition, product, loop_host, Slack, Telegram; declared features; all targets): the corrected loop-host/product/composition unit suites passed, including composition 550/550 and all four former failures. The local run later stopped in unrelated adapter fixtures because the workstation exhausted disk space after isolated rebuilds (no storage space), not because of a test assertion. Generated Cargo artifacts were cleaned afterward. The fresh GitHub runner is the authoritative complete-bucket check.

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

CI is green on 2076227b8. The formerly failing Test Reborn crate bucket (affected-2) passed in 9m49s, all affected Reborn buckets passed, all-features clippy passed, Railway preview deployment passed, and the required aggregate checks Code Style (fmt + clippy) and Tests (Reborn) are both green.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7135 August 4, 2026 13:04 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (3)
crates/ironclaw_runner/src/subagent/await_edge/resolver.rs (1)

1572-1640: 📐 Maintainability & Code Quality | 🟠 Major | ⚡ Quick win

Add a collision regression for spawn_provider_call_id.

The drain test gives each child a unique result_ref. It passes if provider-call propagation is removed because legacy generic lookup can still find each row. Use one shared result_ref with distinct provider call IDs. Assert that each placeholder receives its own terminal summary.

  • crates/ironclaw_runner/src/subagent/await_edge/resolver.rs#L1572-L1640: use a shared result reference and assert provider-call-specific transcript updates.
  • crates/ironclaw_runner/src/subagent/await_edge/resolver.rs#L1062-L1100: assert that reconstruction retains spawn_provider_call_id.
  • crates/ironclaw_runner/src/subagent/await_edge/store.rs#L489-L503: assert that legacy conversion retains spawn_provider_call_id.

As per coding guidelines, “New or changed production-wired behavior must have a caller-level test” and “Every bug fix must include a regression test that fails before the fix.”

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_runner/src/subagent/await_edge/resolver.rs` around lines 1572
- 1640, Strengthen the collision regression: in
crates/ironclaw_runner/src/subagent/await_edge/resolver.rs#L1572-L1640, give all
drained children one shared result_ref while keeping distinct
spawn_provider_call_id values, then assert each placeholder receives its own
terminal summary. In
crates/ironclaw_runner/src/subagent/await_edge/resolver.rs#L1062-L1100, assert
reconstruction preserves AwaitedChildSetRecord.spawn_provider_call_id. In
crates/ironclaw_runner/src/subagent/await_edge/store.rs#L489-L503, assert legacy
conversion also preserves spawn_provider_call_id.

Source: Coding guidelines

tests/integration/tool_call.rs (1)

799-806: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Assert that the second continuation is terminal.

page_two_detail["next_offset"] is only compared with page_two["next_offset"]. A regression that emits a third-page offset preserves that equality and passes this test. Assert that page_two["next_offset"] is null to verify the terminal marker required by this continuation scenario.

Proposed regression assertion
     assert_eq!(
         page_two_detail["next_offset"].as_u64(),
         page_two["next_offset"].as_u64(),
         "second-page replay metadata must match the second page output"
     );
+    assert!(
+        page_two["next_offset"].is_null(),
+        "the second continuation must be terminal"
+    );
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/integration/tool_call.rs` around lines 799 - 806, Update the
second-page assertions in the continuation scenario to verify that
page_two["next_offset"] is null, confirming the second continuation is terminal.
Retain the existing equality check with page_two_detail["next_offset"] so replay
metadata remains validated.
crates/ironclaw_threads/src/filesystem_service.rs (1)

1989-2054: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Make provider-call replay lookups the write authority.

append_tool_result_reference uses find_tool_result_reference_message and then persists separately: write_new_message skips a populated lookup key instead of returning the winner, so concurrent identical replays can create duplicate records. apply_message_update also writes provider indexes best-effort after the message CAS, so a failed tool_result_provider_call index write can cause a later generic-row replay to add another row. Store the provider-call key in the same transaction/CAS as the message or lookup failure, and update the generic row as a conflict or validation failure rather than a replay duplicate, with concurrent-replay and index-write-failure regression tests.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_threads/src/filesystem_service.rs` around lines 1989 - 2054,
Make provider-call replay lookup authoritative in append_tool_result_reference
and the related write/update flow: atomically claim or persist the provider-call
key with the message CAS, and have write_new_message return the existing winner
when the key is already populated instead of creating another row. Treat
provider-index write failures as lookup/validation failures and update the
existing generic row as a conflict rather than appending a duplicate. Add
regression coverage for concurrent identical replays and provider-index write
failures.

Sources: Coding guidelines, Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@crates/ironclaw_runner/src/subagent/await_edge/resolver.rs`:
- Around line 1572-1640: Strengthen the collision regression: in
crates/ironclaw_runner/src/subagent/await_edge/resolver.rs#L1572-L1640, give all
drained children one shared result_ref while keeping distinct
spawn_provider_call_id values, then assert each placeholder receives its own
terminal summary. In
crates/ironclaw_runner/src/subagent/await_edge/resolver.rs#L1062-L1100, assert
reconstruction preserves AwaitedChildSetRecord.spawn_provider_call_id. In
crates/ironclaw_runner/src/subagent/await_edge/store.rs#L489-L503, assert legacy
conversion also preserves spawn_provider_call_id.

In `@crates/ironclaw_threads/src/filesystem_service.rs`:
- Around line 1989-2054: Make provider-call replay lookup authoritative in
append_tool_result_reference and the related write/update flow: atomically claim
or persist the provider-call key with the message CAS, and have
write_new_message return the existing winner when the key is already populated
instead of creating another row. Treat provider-index write failures as
lookup/validation failures and update the existing generic row as a conflict
rather than appending a duplicate. Add regression coverage for concurrent
identical replays and provider-index write failures.

In `@tests/integration/tool_call.rs`:
- Around line 799-806: Update the second-page assertions in the continuation
scenario to verify that page_two["next_offset"] is null, confirming the second
continuation is terminal. Retain the existing equality check with
page_two_detail["next_offset"] so replay metadata remains validated.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 8b92839d-6411-4e30-8e93-1072cb9887bf

📥 Commits

Reviewing files that changed from the base of the PR and between 2076227 and a2d2205.

📒 Files selected for processing (13)
  • crates/ironclaw_loop_host/src/result_read.rs
  • crates/ironclaw_loop_host/src/subagent_spawn_port.rs
  • crates/ironclaw_loop_host/src/subagent_spawn_port/tests.rs
  • crates/ironclaw_runner/src/subagent/await_edge/mod.rs
  • crates/ironclaw_runner/src/subagent/await_edge/resolver.rs
  • crates/ironclaw_runner/src/subagent/await_edge/store.rs
  • crates/ironclaw_runner/src/subagent/prompt_material.rs
  • crates/ironclaw_threads/src/contract.rs
  • crates/ironclaw_threads/src/filesystem_service.rs
  • crates/ironclaw_threads/src/in_memory.rs
  • crates/ironclaw_threads/tests/filesystem_session_thread_contract.rs
  • crates/ironclaw_threads/tests/session_thread_contract.rs
  • tests/integration/tool_call.rs

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7135 August 4, 2026 13:21 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_reborn_composition/src/runtime/tests/core.rs (1)

1038-1040: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Assert against the original result_ref.

result_ref is captured from the original observation at Lines 974-977 and passed to result_read at Lines 996-997. This assertion compares two fields from the returned payload. If both fields contain the staging-write reference, the test still passes. Compare both returned fields with the captured result_ref so the caller-path regression test enforces the corrected continuation contract.

Suggested assertion
+            let expected_result_ref = result_ref.clone();
             assert_eq!(
-                detail["result_ref"], observation["result_ref"],
-                "result_read replay must expose only the original pageable result reference"
+                detail["result_ref"].as_str(),
+                Some(expected_result_ref.as_str()),
+                "result_read detail must preserve the original pageable result reference"
             );
+            assert_eq!(
+                observation["result_ref"].as_str(),
+                Some(expected_result_ref.as_str()),
+                "result_read observation must preserve the original pageable result reference"
+            );
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_reborn_composition/src/runtime/tests/core.rs` around lines
1038 - 1040, Update the assertion in the result_read replay test to compare both
returned payload fields against the captured original result_ref from the
initial observation, rather than comparing the returned fields to each other.
Preserve the existing regression coverage for the corrected continuation
contract.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@crates/ironclaw_reborn_composition/src/runtime/tests/core.rs`:
- Around line 1038-1040: Update the assertion in the result_read replay test to
compare both returned payload fields against the captured original result_ref
from the initial observation, rather than comparing the returned fields to each
other. Preserve the existing regression coverage for the corrected continuation
contract.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: d9bf8aa4-c7c8-4aeb-899f-5788144062fa

📥 Commits

Reviewing files that changed from the base of the PR and between a2d2205 and 048c4c3.

📒 Files selected for processing (1)
  • crates/ironclaw_reborn_composition/src/runtime/tests/core.rs

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Railway preview QA — PASS (review-fix head)

Tested the exact deployed PR head:

Given / When / Then evidence

Case Evidence Result
Large durable result Created a non-sensitive 1,500-line fixture and invoked read_file without a range. The durable result reported 67,013 bytes and surfaced next_offset=24576. PASS
First continuation Expanded builtin__result_read showed offset=24576, max_bytes=24576, and the original read_file result reference; the card status was succeeded and surfaced the next byte offset 49152. PASS
Second continuation from surfaced values Expanded the next builtin__result_read card and verified offset=49152, max_bytes=24576, and the same original result reference byte-for-byte. The card status was succeeded; no fresh inline-only identity was used. PASS
Output markers line-0000 was present in the initial page and line-1499 was read back from the fixture. PASS
Durable replay Reloaded the deep-linked conversation. The 67,013-byte report, original result reference, offsets 24576/49152, final markers, and six-tool activity history were restored. PASS
Cleanup Deleted only /workspace/railway-test-result-read-7135-048c.txt; read-back reported that exact path no longer exists and no other file was touched. PASS

Exact regression assertion

The production WebUI caller fed the first surfaced next_offset and original result reference into page one, then fed page one’s surfaced offset 49152 and that same reference into page two. Both persisted provider-call cards succeeded and replayed after refresh.

Exploratory call excluded from the assertion

After the two target byte-page calls succeeded, the preview assistant made one extra exploratory result_read call by interpreting nested line-pagination metadata as a byte offset. It was not needed for, and is not counted as, the continuation-identity assertion above. Terminal content was independently read back before cleanup.

Skipped as unrelated

Permissions/role denial, responsive layout, uploads, and auth lifecycle beyond preview login were not retested because this PR changes backend result-reference continuation and durable replay behavior.

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Final verification — green

  • Exact head: 048c4c3482968ca1c215748e55853cbb7f473cd1
  • GitHub Actions: all required checks pass; no failures or pending checks.
  • Former failure: Test Reborn crate bucket (affected-2) passes in 9m35s.
  • Other key gates: affected buckets 1/3, Reborn integration tests, recorded fixtures, all-features clippy, code style, fast deterministic checks, WebUI build/E2E partitions, hooks parity, platform compatibility, CodeRabbit, and Railway all pass.
  • Review audit: 0 unresolved review threads; all five addressed comments received reviewer/bot confirmation.
  • Local CI-fix regression: cargo test -p ironclaw_reborn_composition cancel_run_propagates_to_subagent_children --lib → 1 passed.
  • Railway exact-head browser evidence: fix(loop): preserve pageable result_read continuation references #7135 (comment)
  • Repository state: worktree clean and local/remote heads aligned.

BenKurrek
BenKurrek previously approved these changes Aug 5, 2026
@serrrfirat
serrrfirat added this pull request to the merge queue Aug 5, 2026
@serrrfirat
serrrfirat removed this pull request from the merge queue due to a manual request Aug 5, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7135 August 5, 2026 14:20 Destroyed
@coderabbitai

coderabbitai Bot commented Aug 5, 2026

Copy link
Copy Markdown

Note

GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer.

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Railway preview QA — PASS

Given / When / Then evidence

Case Evidence Result
Exact deployed head Railway reported the PR preview healthy for the tested SHA before browser testing began. The frontend asset remained assets/app-x_bie4az.js, acceptable for this backend-only change. PASS
Authenticated caller path Bearer login succeeded and the protected chat surface rendered. PASS
Pageable source result A read-only deterministic shell command produced one 16,678-byte result and surfaced one pageable result reference. PASS
Continuation identity and offsets The caller reused that same surfaced reference for result_read at offsets 0, 4000, 8000, 12000, and 16000; every call returned success, with next offsets 4000, 8000, 12000, 16000, then terminal/no next offset. PASS
Byte completeness Returned chunks totaled 16,678 bytes, exactly matching the source result; the final chunk was 853 bytes. PASS
Regression condition No invalid-reference or boundary error occurred while feeding the surfaced reference and each returned continuation offset into the next call. PASS
Refresh/read-back Reloading /chat restored the completed conversation with the same five pagination steps, terminal state, and exact byte total. PASS

Exact regression result: the model-visible continuation identity remained usable through all pages and terminated correctly; no fresh inline-only reference displaced continuation authority.

Cleanup: browser tabs were finalized. No files or external systems were mutated. One clearly identifiable acceptance-test conversation remains only in the ephemeral PR preview; it was retained to preserve reviewable read-back evidence.

Skipped: permissions/denied-role and unrelated UI routes, because this PR changes backend result-reference continuation rather than authorization or frontend behavior. External model selection was used only to drive the real caller; the deterministic pagination assertions above are based on rendered tool outcomes and byte counts.

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Merge-queue remediation — PASS

Exact head: 80d38bc748a3ab06591d460be9cd9cc7c9a91281

  • Merged current main without history rewriting; GitHub now reports mergeable: MERGEABLE and the history-sharing check passes.
  • Fixed the post-merge architecture failure by relocating the test-only latest_result_output helper from the production module into test support; runtime behavior is unchanged.
  • All non-skipped exact-head checks are green, including required Code Style (fmt + clippy) and Tests (Reborn), affected crate buckets, integration tests, WebUI E2E shards, stress, CodeRabbit, and Railway.
  • Local focused evidence:
    • two-page/end-to-end durable continuation regression: 1 passed
    • thread result-reference contracts: 9 passed
    • filesystem restart/scope/dedup contracts: 3 passed
    • capability-host suite: 41 passed
    • spawn and cancellation compatibility: 1 passed each
    • full ironclaw_architecture_tests: PASS
    • affected crates Clippy with -D warnings: PASS

Railway browser evidence: #7135 (comment)

The PR is CI-green and conflict-free. GitHub still shows REVIEW_REQUIRED because the prior human approval was dismissed when the new head was pushed; that is the only remaining merge gate.

@serrrfirat
serrrfirat enabled auto-merge August 5, 2026 15:49
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7135 August 5, 2026 16:55 Destroyed
@coderabbitai

coderabbitai Bot commented Aug 5, 2026

Copy link
Copy Markdown

Note

GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/loop/ironclaw_turn_runner/src/subagent/await_edge/resolver.rs (1)

1572-1641: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Make provider-call identity observable in these regression tests.

The mixed-status test uses a distinct result_ref for each child. update_tool_result_reference can therefore select the correct placeholder without provider_call_id. Use one shared result_ref and assert that each provider-call envelope receives its own terminal summary.

  • crates/loop/ironclaw_turn_runner/src/subagent/await_edge/resolver.rs#L1572-L1641: use the same result reference for both parent placeholders, then assert updates select the matching provider call ID.
  • crates/loop/ironclaw_turn_runner/src/subagent/await_edge/resolver.rs#L1062-L1101: assert reconstruct_edge preserves spawn_provider_call_id.
  • crates/loop/ironclaw_turn_runner/src/subagent/await_edge/store.rs#L478-L504: assert legacy metadata projection preserves spawn_provider_call_id.

This violates the Test through the caller invariant for a helper that gates transcript updates. As per coding guidelines, “New or changed production-wired behavior must have a caller-level test at the nearest meaningful seam.” As per path instructions, “Test through the caller” requires a test that drives the real call site for a helper that gates a side effect.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/loop/ironclaw_turn_runner/src/subagent/await_edge/resolver.rs` around
lines 1572 - 1641, Update the mixed-status regression test around the await-edge
setup in
crates/loop/ironclaw_turn_runner/src/subagent/await_edge/resolver.rs:1572-1641
to reuse one result_ref for both parent placeholders and assert each terminal
update selects the matching provider_call_id and summary. In
crates/loop/ironclaw_turn_runner/src/subagent/await_edge/resolver.rs:1062-1101,
extend the reconstruct_edge assertions to preserve spawn_provider_call_id. In
crates/loop/ironclaw_turn_runner/src/subagent/await_edge/store.rs:478-504,
assert the legacy metadata projection also preserves spawn_provider_call_id.

Sources: Coding guidelines, Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/integration/tool_call.rs`:
- Around line 598-607: Update the result_read-related documentation entry in
tests/CLAUDE.md to describe the re-scoped tool_call.rs integration scenario: two
continuation turns on the same thread, both persisted result_read envelopes, and
continuation via result_ref and next_offset while preserving the durable
read_file serialization. Keep the existing Tools reference and make the
documentation reflect the current test behavior.

---

Outside diff comments:
In `@crates/loop/ironclaw_turn_runner/src/subagent/await_edge/resolver.rs`:
- Around line 1572-1641: Update the mixed-status regression test around the
await-edge setup in
crates/loop/ironclaw_turn_runner/src/subagent/await_edge/resolver.rs:1572-1641
to reuse one result_ref for both parent placeholders and assert each terminal
update selects the matching provider_call_id and summary. In
crates/loop/ironclaw_turn_runner/src/subagent/await_edge/resolver.rs:1062-1101,
extend the reconstruct_edge assertions to preserve spawn_provider_call_id. In
crates/loop/ironclaw_turn_runner/src/subagent/await_edge/store.rs:478-504,
assert the legacy metadata projection also preserves spawn_provider_call_id.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 2853faf2-8418-43ab-9d50-39b1fde774e2

📥 Commits

Reviewing files that changed from the base of the PR and between 80af973 and f1ba429.

📒 Files selected for processing (16)
  • crates/app/ironclaw_composition/src/runtime/capability_host/tests.rs
  • crates/app/ironclaw_composition/src/runtime/tests/core.rs
  • crates/domains/ironclaw_threads/src/contract.rs
  • crates/domains/ironclaw_threads/src/filesystem_service.rs
  • crates/domains/ironclaw_threads/src/filesystem_service/message_lookup_index.rs
  • crates/domains/ironclaw_threads/src/in_memory.rs
  • crates/domains/ironclaw_threads/tests/filesystem_session_thread_contract.rs
  • crates/domains/ironclaw_threads/tests/session_thread_contract.rs
  • crates/loop/ironclaw_loop_host/src/result_read.rs
  • crates/loop/ironclaw_loop_host/src/subagent_spawn_port.rs
  • crates/loop/ironclaw_loop_host/src/subagent_spawn_port/tests.rs
  • crates/loop/ironclaw_turn_runner/src/subagent/await_edge/mod.rs
  • crates/loop/ironclaw_turn_runner/src/subagent/await_edge/resolver.rs
  • crates/loop/ironclaw_turn_runner/src/subagent/await_edge/store.rs
  • crates/loop/ironclaw_turn_runner/src/subagent/prompt_material.rs
  • tests/integration/tool_call.rs

Comment on lines +598 to +607
/// `result_read` continuation (issue #5838): two subsequent scripted turns on
/// the SAME thread page the durable `read_file` result. Page two is invoked
/// exclusively with the `result_ref` and `next_offset` surfaced by page one,
/// proving that model-visible continuation metadata retains the original
/// pageable identity instead of exposing the fresh `InlineOnly` write ref.
/// Both chunks continue byte-exactly through the SAME canonical serialization
/// `tool_result_output` returns for `read_file` — no gap, no overlap — and
/// report the durable record's true `total_bytes`. Page one's chunk contains a
/// credential marker, so its inline preview is suppressed; the continuation
/// identity and offset must survive independently of preview content.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Description: Check that tests/CLAUDE.md describes the re-scoped result_read continuation scenario.
set -euo pipefail

fd -t f 'CLAUDE.md' tests

# Test: the scenario name and its documented shape. Expect a description matching two continuation pages.
rg -n -C 6 'result_read_continues_a_durable_result_byte_exactly|tool_call' tests/CLAUDE.md || true

# Confirm the change touched the doc in this branch.
git log --name-only -1 --pretty=format: | rg -n 'tests/CLAUDE\.md' || echo "tests/CLAUDE.md not modified in HEAD commit"

Repository: nearai/ironclaw

Length of output: 2811


🏁 Script executed:

#!/bin/bash
set -euo pipefail

python3 - <<'PY'
from pathlib import Path
import re

md_path = Path("tests/CL AUDE.md")
if md_path.exists():
    text = md_path.read_text()
    lines = []
    for i, line in enumerate(text.splitlines(), 1):
        if "tool_call.rs" in line or "result_read_continues_a_durable_result_byte_exactly" in line:
            start=max(1,i-8); end=min(len(text.splitlines()),i+9)
            rows=text.splitlines()
            print(f"\n--- context around line {i} ---")
            for j in range(start,end+1):
                print(f"{j}:{rows[j-1]}")
else:
    print("tests/CL AUDE.md missing")

test_path = Path("tests/integration/tool_call.rs")
if test_path.exists():
    txt=test_path.read_text().splitlines()
    for idx,line in enumerate(txt,1):
        if "result_read_continues_a_durable_result_byte_exactly" in line:
            print(f"\n--- test symbol at line {idx} ---")
            for j in range(max(1,idx-5), min(len(txt), idx+14)+1):
                print(f"{j}:{txt[j-1]}")
PY

# List changed tracked files in the PR diff, if available.
git diff --name-only -- tests/CLAUDE.md tests/integration/tool_call.rs || true

Repository: nearai/ironclaw

Length of output: 1236


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "--- tests CLAUDE.md tool sections ---"
rg -n -C 3 'tool_call\.rs|result_read_continues_a_durable_result_byte_exactly|result_read' tests/CLAUDE.md || echo "no matches"

echo "--- files tracked matching CLAUDE.md ---"
git ls-files | rg '(^|/)CLAUDE\.md$|CLAUDE|claude|AGENTS' || true

echo "--- test function surrounding lines ---"
sed -n '588,650p' tests/integration/tool_call.rs

Repository: nearai/ironclaw

Length of output: 9657


Add the re-scoped result_read scenario to tests/CLAUDE.md.

This integration test now expects two continuation turns and checks both persisted result_read envelopes. Path invariants require the same-commit tests/CLAUDE.md update for materially re-scoped Reborn integration tests. The current entry only references tool_call.rs under Tools.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/integration/tool_call.rs` around lines 598 - 607, Update the
result_read-related documentation entry in tests/CLAUDE.md to describe the
re-scoped tool_call.rs integration scenario: two continuation turns on the same
thread, both persisted result_read envelopes, and continuation via result_ref
and next_offset while preserving the durable read_file serialization. Keep the
existing Tools reference and make the documentation reflect the current test
behavior.

Source: Path instructions

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7135 August 5, 2026 20:57 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (3)
tests/integration/tool_call.rs (3)

406-408: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Key the HTTP response by URL.

with_real_egress_response_bodies([source_bytes]) binds the body by response order. Register source_bytes with ScriptedHttpResponse for HTTP_TOOL_URL. This prevents a later request-order change from serving the fixture to the wrong request.

As per path instructions, “Use ScriptedHttpResponse as the canonical keyed HTTP scripting API; add new per-URL responses there rather than creating a parallel matcher.”

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/integration/tool_call.rs` around lines 406 - 408, Update the
RebornIntegrationHarness setup to register source_bytes through
ScriptedHttpResponse keyed by HTTP_TOOL_URL instead of using
with_real_egress_response_bodies. Preserve the existing real egress pipeline
while ensuring the fixture is selected by URL rather than response order.

Source: Path instructions


357-400: 📐 Maintainability & Code Quality | 🟠 Major | 🏗️ Heavy lift

Move the large fixture and assertions out of the test body.

json_queries_scoped_file_and_adjacent_array_indices spans about 180 lines and embeds large nested serde_json::json! fixtures. Extract the fixture, scripted replies, and repeated assertions into focused helpers. Keep the visible test body as build → submit_turn → assert.

As per path instructions, “Keep individual tests approximately 3–12 lines with no nested structs in the body.”

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/integration/tool_call.rs` around lines 357 - 400, Refactor
json_queries_scoped_file_and_adjacent_array_indices so its body only builds the
scenario, calls submit_turn, and performs the high-level assertion. Move the
large serde_json fixture, scripted replies, repeated assertions, and related
setup into focused helper functions, ensuring the test body contains no nested
structs and remains approximately 3–12 lines.

Source: Path instructions


486-503: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Assert these query results through the scoped result helper.

assert_tool_result_contains scans every recorded capability result with any(), so "$", "value-15" or "1.740..." can satisfy the assertion from another tool result in the same turn. Use tool_result_output("builtin.json") plus a scoped assertion that checks the intended query/path/output, or add a scoped *_since variant for the JSON query output.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/integration/tool_call.rs` around lines 486 - 503, Replace the broad
assert_tool_result_contains calls for the JSON query expectations with scoped
assertions against tool_result_output("builtin.json"), ensuring each expected
path/query value is validated only within the builtin.json result. If the
existing helper cannot express this, add and use a JSON-specific *_since scoped
variant while preserving the invalid-input error assertion’s intended behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@tests/integration/tool_call.rs`:
- Around line 406-408: Update the RebornIntegrationHarness setup to register
source_bytes through ScriptedHttpResponse keyed by HTTP_TOOL_URL instead of
using with_real_egress_response_bodies. Preserve the existing real egress
pipeline while ensuring the fixture is selected by URL rather than response
order.
- Around line 357-400: Refactor
json_queries_scoped_file_and_adjacent_array_indices so its body only builds the
scenario, calls submit_turn, and performs the high-level assertion. Move the
large serde_json fixture, scripted replies, repeated assertions, and related
setup into focused helper functions, ensuring the test body contains no nested
structs and remains approximately 3–12 lines.
- Around line 486-503: Replace the broad assert_tool_result_contains calls for
the JSON query expectations with scoped assertions against
tool_result_output("builtin.json"), ensuring each expected path/query value is
validated only within the builtin.json result. If the existing helper cannot
express this, add and use a JSON-specific *_since scoped variant while
preserving the invalid-input error assertion’s intended behavior.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: ea8bc08f-fb88-467b-a124-de8bd5a1e87b

📥 Commits

Reviewing files that changed from the base of the PR and between f1ba429 and 362b462.

📒 Files selected for processing (1)
  • tests/integration/tool_call.rs

@think-in-universe

Copy link
Copy Markdown
Collaborator

@ironloopai review

@ironloopai

ironloopai Bot commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #7135

🟢 Completed · Review submitted

Submitted review →

Reviewed the complete trusted base-to-head comparison, including the large mainline merge and the feature-specific result paging, provider-call indexing, replay, and subagent-settlement paths. No concrete, actionable defect was identified. The implementation preserves the original pageable reference while retaining inline evidence writes, separates exact provider-call replay from distinct page calls sharing a reference, and carries spawn identity through durable await-edge recovery.

Manual command by @think-in-universe · attempt 1 of 3 · completed in 2m 19s

Run details
  • Repository: nearai/ironclaw
  • Base: main at 1f6b56d
  • Head: codex/fix-result-read-continuation-ref at 362b462
  • Created: Aug 6, 2026, 8:53 AM UTC
  • Updated: Aug 6, 2026, 8:55 AM UTC
  • Run: e9c0fe98-4c05-49df-a3c4-5ab6fd74f654
  • Latest attempt: 1 · Completed · 4635651e-5461-4040-a5ab-88d7d357aa2b

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Review complete · PR #7135

✅ No actionable findings

Reviewed the complete trusted base-to-head comparison, including the large mainline merge and the feature-specific result paging, provider-call indexing, replay, and subagent-settlement paths. No concrete, actionable defect was identified. The implementation preserves the original pageable reference while retaining inline evidence writes, separates exact provider-call replay from distinct page calls sharing a reference, and carries spawn identity through durable await-edge recovery.

Validation and technical details
  • Verified refs/ironloop/base (1f6b56d) against refs/ironloop/head (362b462): 390 files changed, +5,591/-17,693, including merged mainline changes.
  • Inspected the complete changed-area inventory and repository/crate guidance; the codebase graph was unavailable, so live-code tracing and targeted rg searches were used as the prescribed fallback.
  • Traced result_read completion through inline result staging, sanitized model observation, output evidence, transcript persistence, provider-call-specific lookup indexes, filesystem CAS updates, and in-memory parity.
  • Traced spawn provider-call identity from registration and authorization through SubagentThreadMetadata, AwaitedChildSetRecord, AwaitEdge persistence/reconstruction, and exact parent transcript settlement.
  • Reviewed the added in-memory, filesystem, composition, runner, loop-host, and two-page integration assertions covering exact replay, shared-reference pagination, legacy fallback, restart behavior, and settlement targeting.
  • git diff --check refs/ironloop/base refs/ironloop/head completed successfully, and the worktree remained unchanged.
  • Focused Rust tests could not be executed in this review environment because cargo is not installed (/bin/bash: cargo: command not found); test source and supplied CI context were inspected instead.
  • Base: main
  • Head: codex/fix-result-read-continuation-ref at 362b462
  • Run: e9c0fe98-4c05-49df-a3c4-5ab6fd74f654

@serrrfirat
serrrfirat added this pull request to the merge queue Aug 6, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Aug 6, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7135 August 6, 2026 09:56 Destroyed
@think-in-universe
think-in-universe added this pull request to the merge queue Aug 6, 2026
Merged via the queue into main with commit 77a1287 Aug 6, 2026
48 checks passed
@think-in-universe
think-in-universe deleted the codex/fix-result-read-continuation-ref branch August 6, 2026 10:24
@serrrfirat serrrfirat mentioned this pull request Aug 10, 2026
29 tasks
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…rai#7135)

* fix(loop): preserve result_read continuation reference

Return the original pageable result reference from result_read while retaining inline-only chunk persistence and evidence. Key transcript dedup by provider call so multiple pages can safely share one durable source reference, with regression coverage for replay and a two-page continuation.

* test(composition): align result read continuation assertions

* fix(loop): pin result updates to provider calls

* fix(ci): update subagent metadata fixture

* test(composition): keep result lookup helper test-only

* test(reborn): restore await-edge fixture compatibility

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-7135 — 70d962ca Deployed Aug 6, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants