Skip to content

fix(reborn): make model-visible failures recoverable - #6437

Merged
serrrfirat merged 13 commits into
mainfrom
codex/reborn-error-recoverability-item1
Jul 22, 2026
Merged

serrrfirat merged 13 commits into
mainfrom
codex/reborn-error-recoverability-item1

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 21, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Route model-fixable request, sandbox-plan, WASM guest, and capability failures through typed recovery or model-visible outcomes instead of opaque executor failures.
  • Preserve precise failure categories and loop-exit violation details across runner, projection, retry, and durable checkpoint boundaries.
  • Address review feedback by distinguishing typed stale requests from deterministic invalid requests, retaining host-vs-guest WASM provenance, and preserving corrective sandbox diagnostics after canonical secret scrubbing and an untrusted-data fence.

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Related #6284 (item 1 only; this PR does not close the epic).

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings — ran the stronger workspace-wide command, plus the host-runtime command after the final module extraction.
  • cargo build — not run separately; all changed crates and the workspace were compiled by tests and clippy.
  • Relevant tests pass: agent-loop, loop-host, runner, host-runtime contracts, architecture, failure/retry/resume, and the Reborn architecture E2E gate.
  • cargo test --features integration — not applicable: no database-backed behavior or schema changed.
  • Manual testing — not applicable: behavior is covered through deterministic runner, host-runtime, and Reborn harness seams.
  • Multi-agent review and comment-fix audit were run before requesting re-review.

Test Strategy

User behavior: Model-correctable failures remain model-visible and can lead to a changed action in the same run. The new deterministic scenarios prove recovery for typed stale model requests, generic invalid capability input, Exa content-fetch failure with fallback search, and a real GitHub WASM guest operation_failed response. Corrective sandbox validation details still reach the model only after canonical secret scrubbing and an explicit untrusted-data fence.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: ExternalTool gate contract, no-progress cancellation, stale-request checkpoint round-trip, invalid/stale/scope-mismatch model mapping, sandbox spawn/resume diagnostics, canonical redaction/fencing, WASM host-vs-guest provenance, public runtime cause propagation, and GitHub non-validation 422 classification.
  • Reborn integration: a stale provider request re-drives and completes in one run; invalid capability input reaches the next model request and succeeds after changed input; Exa get_content failure reaches the model as typed corrective context and recovers through web-access.search; a real GitHub WASM guest 422 reaches the model as operation_failed and recovers through github.get_repo.
  • Recorded model fixture: Not applicable: these tests use deterministic scripted model turns to verify recovery routing, not stochastic provider choice or request-shape drift.
  • Browser E2E: Not applicable: no WebUI behavior changed.
  • Backend or runtime: full failure/retry/resume, Web Access, and GitHub WASM parity binaries; exact loop-host sandbox producer/scrubber contracts; focused runner mapping; host-runtime GitHub classification; architecture suite; and pre-commit safety.
  • Live canary: Not applicable: hermetic external-service doubles trigger the exact failure classes without live credentials or provider drift.

What the tests prove: Eligible failures cross the real loop/runner/checkpoint seams with stable categories, appear in the next model request as safe typed context, permit the model to choose a different action, and can still persist a final reply. Invalid gate outcomes fail closed; sandbox diagnostics remain useful, secret-scrubbed, and fenced; and host infrastructure failures cannot be mislabeled as guest tool failures.

Commands run:

  • cargo fmt --all -- --check
  • cargo test -p ironclaw_agent_loop --no-fail-fast -q
  • cargo test -p ironclaw_loop_host --no-fail-fast
  • cargo test -p ironclaw_runner --no-fail-fast -q
  • cargo test -p ironclaw_host_runtime --no-fail-fast
  • cargo test -p ironclaw_host_runtime services::wasm_execution::tests --lib -q
  • cargo test -p ironclaw_host_runtime --test host_runtime_services_contract -q
  • cargo test --test reborn_failure_retry_resume_e2e
  • cargo test --test reborn_integration_web_access
  • cargo test --test reborn_trace_wasm_github_fixture_parity
  • cargo test --test reborn_trace_error_path_parity reborn_trace_invalid_input_recovers_with_changed_action -- --nocapture
  • cargo test -p ironclaw_loop_host process_sandbox_capability_maps_runtime_invalid_plan_failure_to_model -- --nocapture
  • cargo test -p ironclaw_loop_host process_sandbox_rejection_keeps_scrubbed_fenced_diagnostic_model_visible -- --nocapture
  • cargo test -p ironclaw_host_runtime --test github_wasm_runtime_contract host_runtime_services_keeps_github_non_validation_422_as_operation_failed -- --nocapture
  • cargo test -p ironclaw_runner capability_model_request_errors_preserve_stale_distinction -- --nocapture
  • cargo test -p ironclaw_architecture
  • CARGO_TEST_ARGS='-q' scripts/reborn-e2e-rust.sh architecture
  • cargo clippy --workspace --all-targets --all-features -- -D warnings
  • cargo clippy -p ironclaw_host_runtime --all-targets --all-features -- -D warnings
  • scripts/pre-commit-safety.sh
  • git diff --check

Local-only limitations:

  • The full reborn_trace_error_path_parity binary was stopped after multiple existing host-runtime-backed cases hung while building Reborn services; an untouched existing process-profile test reproduced the same harness issue. The new recording-seam invalid-input recovery test passes when run directly, and the other three affected binaries pass in full.
  • A literal invalid process-sandbox plan followed by a corrected process execution cannot yet complete in the same turn: a valid system.process_sandbox.run yields Suspension::Process, and the canonical executor currently fails closed on unsupported process-wait resume. This PR therefore pairs the exact sandbox producer/scrubber contracts with a whole-turn capability-neutral InvalidInput recovery test instead of adding a test-only adapter that would misrepresent production behavior.
  • The host-runtime library target previously passed 338/340 tests. Two untouched Trace Commons tests expected an unenrolled system scope, but this workstation has an existing enrollment, so they reached the deliberately absent test egress and returned NetworkDenied. Every host-runtime integration target passed, including 114/114 host-runtime service contracts; the focused final WASM module run passed 16/16.
  • No browser test was run because this patch has no WebUI behavior. The prior PR-head CI had an unrelated WebUI E2E failure; the pushed commit will receive a fresh CI run.

Security Impact

This changes error handling at model, capability, sandbox, and gate boundaries. Process-sandbox request validation now has one production owner in host-runtime; its corrective cause is passed through the existing leak detector, model-visible text sanitizer, injection scanner, compact untrusted-data fence, legacy delimiter normalization, and a UTF-8-safe 400-byte bound before reaching the model. Authorization, approvals, network mediation, credential handling, and capability visibility are unchanged. Invalid gate outcomes continue to fail closed as DriverBug.

Reborn Trust-Boundary Checklist

  • Public policy/evidence/trust-bearing types: no new public constructors; host-owned evidence adapters remain the only production trust source.
  • Untrusted content enters prompts only through an envelope/escaping primitive: sandbox diagnostics use the canonical scrubber plus an explicit compact untrusted-data fence.
  • Hashes declare purpose; trust/binding/authenticity uses SHA-256/BLAKE3 or separate authenticity check: replay-relevant family identities use updated BLAKE3 fingerprints.
  • New/changed status, exit, policy, runtime, or error variants: all HostManagedModelErrorKind mappings were audited; StaleRequest is additive and maps only to StaleSurface.
  • Security/durability serde(default) fields fail closed or have migration tests: stale-request recovery state survives checkpoint serialization and reload.
  • Queues/maps/buffers/counters have bounds and overflow-safe arithmetic: existing bounded recovery state is retained; sandbox diagnostic truncation is UTF-8 safe.
  • Driver/operator-visible errors have stable class semantics: stale request, deterministic invalid request, guest operation failure, host executor failure, and driver bug remain distinct.
  • Sandbox/native/host names accurately describe trust boundary: host-owned WASM blocking failures carry executor provenance; confirmed guest failures retain guest provenance.

Database Impact

None. No migration, schema, PostgreSQL, or libSQL behavior changed.

Blast Radius

Agent-loop recovery, host-managed model error serialization/mapping, runner retry disposition, process-sandbox failure diagnostics, WASM blocking error provenance, gate contract enforcement, and recovery checkpoint tests. Compatibility risk is limited to the additive stale_request serialized variant and the intentional behavior change that generic invalid model requests no longer retry.

Rollback Plan

Revert this PR; there is no database or persistence-schema migration. Reverting restores generic invalid-request retries, the prior WASM error classification, and the previous sandbox diagnostic path. If rollback occurs during mixed-version operation, older readers will not understand newly persisted stale_request values, so drain or complete affected in-flight runs first.

Review Follow-Through


Review track: C (runtime and trust-boundary error semantics)

claude and others added 5 commits July 21, 2026 20:59
…ons + mapping re-bucket (checkpoint)

Checkpoint of in-flight subagent work; will be reorganized into reviewed
commits before any PR.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017n7vtfDLvAD9KUZLMTxWVm
…odel-visible operation failures

A guest trap and a script timeout are driven by the composed request, not
host infrastructure — an identical retry fails identically. Both previously
classified into the retryable-infra Backend bucket, burning the availability
retry budget before the model ever saw the error. Re-bucket both to
OperationFailed so they surface as immediate model-visible tool errors
(#6284 item 1; follows the audit's §6.1 re-bucket pattern).

Red-first: flipped the pinned mapping expectations
(dispatch_kind_to_failure_pins_every_runtime_dispatch_error_kind, new
script_timeout_classifies_as_model_visible_operation_failure), watched both
fail on the old mapping, then changed the two production arms.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017n7vtfDLvAD9KUZLMTxWVm
…checkpoint 3)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017n7vtfDLvAD9KUZLMTxWVm
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@ironloopai

ironloopai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown
Contributor

🔎 IronLoop Review Status

Head: 0d898b00ebc9de2965eb94d2e5e94a04ba4b6a62
Result: One or more review results were superseded by a newer PR head.
Next: Run @ironloopai review on the latest PR head.
Updated: 2026-07-22T13:16:49.645Z

Current reviewers:

Reviewer State Verdict Findings Last update
ironloop/common-reviewer (reviewer) Superseded N/A N/A 2026-07-21T22:08:10.992Z
Reviewer summaries
Reviewer Detail
ironloop/common-reviewer (reviewer) Superseded by a newer PR head. New head: 16f9955. Previous verdict: Changes requested.
Recent activity
Time Reviewer State Detail
2026-07-21T20:32:18.740Z ironloop/common-reviewer (reviewer) Superseded A newer PR head replaced this review (99cb2dc).
2026-07-21T21:53:39.924Z ironloop/common-reviewer (reviewer) Queued Accepted review request for head 99cb2dc.
2026-07-21T21:53:39.924Z ironloop/common-reviewer (reviewer) Queued Waiting for this reviewer lane to become available.
2026-07-21T21:53:40.887Z ironloop/common-reviewer (reviewer) Started Reviewer worker started.
2026-07-21T21:53:43.577Z ironloop/common-reviewer (reviewer) Workspace ready Prepared isolated checkout (merge_ref) at f132131.
2026-07-21T22:01:13.197Z ironloop/common-reviewer (reviewer) Result captured Changes requested; 1 blocking finding.
2026-07-21T22:01:13.197Z ironloop/common-reviewer (reviewer) Completed Review completed and terminal status was persisted.
2026-07-21T22:08:10.992Z ironloop/common-reviewer (reviewer) Superseded A newer PR head replaced this review (16f9955).
Available commands
  • @ironloopai help
  • @ironloopai agents
  • @ironloopai review
  • @ironloopai review --agent <agent>
Run metadata

Admission: webhook accepted the request and IronLoop persisted reviewer state before this projection.

@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 2ed1ad24-5d3d-4ec8-9784-77158f0dc2af

📥 Commits

Reviewing files that changed from the base of the PR and between 60afcf8 and e2073a3.

📒 Files selected for processing (2)
  • crates/ironclaw_loop_host/src/capability_port.rs
  • tests/reborn_failure_retry_resume_e2e.rs

📝 Walkthrough

Summary by CodeRabbit

  • Bug Fixes
    • Improved recovery from stale model requests by retrying at iteration scope and adding the model_stale_request failure category.
    • Authentication, checkpoint rejection, and transcript write failures now abort immediately with clearer, persisted failure details.
    • No-progress termination now attaches a best-effort explanation when available.
    • Invalid planner gate outcomes now fail safely, preserving durable loop-exit violation details.
    • Sandbox plan validation failures now surface actionable, redacted, model-visible diagnostics.
  • Performance
    • Added host-side concurrency throttling for WASM execution and preparation.

Walkthrough

The PR updates executor recovery and gate validation, adds stale-request classification and retry handling, preserves failure details, improves sandbox diagnostics, bounds WASM blocking work, and expands integration and replay coverage.

Changes

Executor recovery and gate enforcement

Layer / File(s) Summary
Gate validation and model recovery
crates/ironclaw_agent_loop/src/executor/*, crates/ironclaw_agent_loop/src/strategies/*, crates/ironclaw_agent_loop/src/state/*, crates/ironclaw_agent_loop/src/families/*
Invalid gate outcomes become DriverBug failures; stale requests retry at iteration scope; terminal model errors receive precise classifications; no-progress failures attach best-effort explanations; family fingerprints are updated.
Executor regression coverage
crates/ironclaw_agent_loop/src/executor/tests.rs, crates/ironclaw_agent_loop/tests/*
Tests cover explanation references, nudge accounting, stale-surface rebuilding, retry exhaustion, cancellation, and invalid gate outcomes.

Sandbox and WASM runtime handling

Layer / File(s) Summary
Sandbox diagnostics
crates/ironclaw_host_runtime/src/production.rs, crates/ironclaw_loop_host/src/capability_port.rs, crates/ironclaw_loop_host/src/model_visible_scrub.rs
Sandbox validation moves to runtime spawn/resume paths and emits scrubbed, model-visible corrective diagnostics.
WASM blocking and error mapping
crates/ironclaw_host_runtime/src/services/*
WASM execution and preparation use bounded blocking semaphores while preserving executor and runtime error provenance.

Runner failure records

Layer / File(s) Summary
Failure categories and durable details
crates/ironclaw_runner/src/*, crates/ironclaw_turns/src/loop_exit*
model_stale_request becomes auto-retriable, interruption categories are preserved, and loop-exit violations persist specific detail strings.

Integration and replay contracts

Layer / File(s) Summary
Typed provider failures and replay behavior
tests/integration/*, tests/reborn_failure_retry_resume_e2e.rs, tests/reborn_trace_error_path_parity.rs
Harnesses script authentication and stale-request failures; recovery, cancellation, malformed-input, and unadvertised-capability behavior is asserted.
Parity, checkpoint, and network wiring
tests/support/reborn_parity_qa/*, tests/integration/support/*
Test harnesses wire checkpoint evidence, exact network responses, invalid-input fixtures, and capability recovery paths.

Estimated code review effort: 4 (Complex) | ~60 minutes

Possibly related issues

Possibly related PRs

Suggested reviewers: ilblackdragon

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title uses conventional-commits style and accurately summarizes the PR’s main recovery-and-failure-visible changes.
Description check ✅ Passed The description matches the required template and fills the key sections: summary, change type, linked issue, validation, testing, security, and rollback.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6437 July 21, 2026 20:26 Destroyed
@github-actions github-actions Bot added scope: docs Documentation size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 21, 2026

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⏭️ IronLoop Review Declined: reviewer

Review at a glance

Disposition Head
⏭️ Review declined 277223b701ac

Head: 277223b701ac3e7cfdedeb27d8731f13de1dd2f8
Reason: The stated base commit is not an ancestor of head (merge base: b9f86de). The supplied comparison spans 159 files and 3,583 additions across recovery semantics plus unrelated composition, migration, CLI, CI, Docker, harness, and documentation areas, so the intended PR layer cannot be isolated reliably.
Next: Rebase or provide the intended stack-layer base so it is an ancestor of head, then request review with unrelated composition/migration/CLI/CI changes split into separate PRs or explicitly identify the cumulative-stack review target.

Run details

Status: Current
Trustworthy review produced: no

Summary

Skipped: the supplied base-to-head comparison is a mega, cross-cutting diff that cannot be reviewed reliably within the configured focused-review scope.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6437 July 21, 2026 20:32 Destroyed

@serrrfirat serrrfirat left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review (multi-agent)

Intent: Make model-correctable Reborn failures recoverable and preserve precise failure evidence across execution, retry, projection, replay, and checkpoint boundaries.
Snapshot: 99cb2dc8f8094c160636448cad90c5b5005ba3e0 (refreshed after the PR merged current main)
Coverage: complete — 28 files, 6 packets, 8 specialist reviewers, 0 reviewer failures.
Stats: 9 deduplicated findings (1 High, 7 Medium, 1 Low) from 16 raw findings.

Findings

  1. High — Sandbox diagnostics bypass the required secret and injection scrub (crates/ironclaw_loop_host/src/capability_port.rs:2845-2864, confidence 100)
    safe_sandbox_plan_cause converts raw serde/ProcessSandboxPlanError text into a model-visible diagnostic after stripping only delimiters and control characters. It bypasses this module's existing model_visible_diagnostic_text path, which applies the full secret registry and prompt-injection fencing. Several sandbox errors interpolate model-supplied host, path, and environment fields, so credential-shaped text or injected instructions can survive into the model-visible detail. The repository error-boundary rule forbids paths and credential material crossing the boundary without the established redaction contract.

  2. Medium — External-tool invalid skip outcome lacks stage coverage (crates/ironclaw_agent_loop/src/executor/gates.rs:33-45, confidence 100)
    The new enforcement contract explicitly rejects SkipAndContinue for Approval, AwaitDependentRun, and ExternalTool gates. The changed tests exercise Approval and AwaitDependentRun, but no GateStage test drives ExternalTool through this enforcement. A regression that exempts or misroutes ExternalTool would therefore pass while silently skipping a gated external call.

  3. Medium — No-progress explanation cancellation path is untested (crates/ironclaw_agent_loop/src/executor/loop_exit.rs:284-288, confidence 100)
    The new NoProgressDetected branch propagates the only error from attach_failure_explanation via await?: cancellation during prompt/model/finalization. Existing cancellation tests exercise the older Aborted explanation path, while the new no-progress tests cover successful and fail-soft explanations only. They would not catch this branch writing a Final checkpoint or returning Failed instead of Cancelled after cancellation.

  4. Medium — Spawn callers do not verify the new model-visible cause (crates/ironclaw_host_runtime/src/production.rs:1587-1619, confidence 100)
    Helper tests assert that malformed and invalid plans initially receive model_visible_cause, but the existing public spawn_capability and resume_spawn_capability contract tests assert only failure kind, disposition, and summary. They do not prove the newly added cause survives the production caller's failure finalization and scrubbing path, so wiring that drops the corrective diagnostic would still pass.

  5. Medium — WASM host failures are mislabeled as guest operation failures (crates/ironclaw_host_runtime/src/production.rs:1878-1884, confidence 100)
    RuntimeDispatchErrorKind::Guest is not limited to guest traps. wasm_error_kind maps every WasmError::ExecutionFailed to it, while run_wasm_execution_blocking and run_wasm_prepare_blocking construct that variant for a closed execution/preparation gate and for blocking-task panics. This new blanket mapping turns those host-runtime failures into model-visible OperationFailed results, removing the backend retry path and incorrectly asking the model to change its tool call.

  6. Medium — Deterministic invalid model requests are treated as stale (crates/ironclaw_agent_loop/src/executor/mapping.rs:116-118, confidence 98)
    InvalidInvocation is the loop-host mapping for every HostManagedModelErrorKind::InvalidRequest, whose contract includes malformed requests and unknown tools. Production producers also use it for invalid routes, provider/model identity mismatches, invalid replay metadata, and missing replay content. Rebuilding the same iteration cannot repair those deterministic faults, but this mapping performs repeated model calls, labels exhaustion model_stale_request, and assigns the Auto retry disposition. Only the genuinely stale-surface case is safe to retry this way.

  7. Medium — Delete the duplicate sandbox-plan validation path (crates/ironclaw_loop_host/src/capability_port.rs:2791-2807, confidence 96)
    The loop-host adapter now parses and validates SandboxProcessPlan, wraps failures in a provider-argument error, and sanitizes its own diagnostic even though DefaultHostRuntime::spawn_capability performs the same parse and validation and already returns a model-visible RuntimeCapabilityFailure. Production spawns therefore validate and serialize the plan twice, while validation rules, summaries, and diagnostic handling must remain synchronized across two crates. This also places runtime-specific request-shape logic in the upper adapter instead of the runtime owner.

  8. Medium — New stale-request attempt class lacks checkpoint round-trip coverage (crates/ironclaw_agent_loop/src/state/slots.rs:427-436, confidence 75)
    ModelStaleRequest is added to the serialized RecoveryAttemptClass used as a persistent BTreeMap key. Existing checkpoint lifecycle coverage round-trips ModelInvalidOutput and the generic state builder uses ModelTransient; no test serializes and reloads the new wire value model_stale_request. Replay/checkpoint compatibility for the new retry counter is therefore unverified.

  9. Low — Gate outcome documentation still recommends a now-invalid skip (crates/ironclaw_agent_loop/src/executor/gates.rs:26-32, confidence 100)
    The new enforcement makes SkipAndContinue invalid for Approval, AwaitDependentRun, and ExternalTool gates, but the owning GateOutcome documentation still says this variant is intended for tools where a missing approval is non-fatal. A strategy author following that local documentation will now produce a DriverBug abort. The new comment also calls the conversion to the more severe Abort outcome a downgrade, which obscures the actual behavior.

Posted from the validated local multi-agent review; detailed fixes are attached inline.

/// is bounded. Returns `None` when nothing legible remains.
fn safe_sandbox_plan_cause(raw: &str) -> Option<String> {
const MAX_BYTES: usize = 400;
let sanitized: String = raw

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

High — Sandbox diagnostics bypass the required secret and injection scrub

safe_sandbox_plan_cause converts raw serde/ProcessSandboxPlanError text into a model-visible diagnostic after stripping only delimiters and control characters. It bypasses this module's existing model_visible_diagnostic_text path, which applies the full secret registry and prompt-injection fencing. Several sandbox errors interpolate model-supplied host, path, and environment fields, so credential-shaped text or injected instructions can survive into the model-visible detail. The repository error-boundary rule forbids paths and credential material crossing the boundary without the established redaction contract.

Fix: Pass the cause through model_visible_diagnostic_text before applying any additional SafeSummary-compatible delimiter normalization, and add a regression case containing credential-shaped/injection text.

Also flagged by: tests/Medium, approach/Medium, bugs/Medium, security/Medium

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 16f9955. Sandbox validation causes now pass through the canonical leak detector, model-visible sanitizer, and injection scanner before a compact untrusted-data fence, SafeSummary delimiter normalization, and the existing UTF-8-safe 400-byte bound. The caller test proves credential-shaped text is redacted while corrective detail remains model-visible. Verified by the full loop-host suite and workspace Clippy.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed; no code change: ALREADY ADDRESSED — sandbox validation causes use the canonical compact scrubber, secret redaction, and injection fence on the current head. Verification: full ironclaw_loop_host suite and workspace clippy passed.

/// mask a driver bug behind a normal completion. The invalid outcome is
/// downgraded to `Abort` with the validator's failure kind (`DriverBug`) so
/// the run fails through the standard abort path.
fn enforce_gate_outcome_contract(outcome: GateOutcome, kind: GateKind) -> GateOutcome {

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — External-tool invalid skip outcome lacks stage coverage

The new enforcement contract explicitly rejects SkipAndContinue for Approval, AwaitDependentRun, and ExternalTool gates. The changed tests exercise Approval and AwaitDependentRun, but no GateStage test drives ExternalTool through this enforcement. A regression that exempts or misroutes ExternalTool would therefore pass while silently skipping a gated external call.

Fix: tests::executor::external_tool_gate_skip_and_continue_fails_as_driver_bug covering an ExternalTool GateStage outcome of SkipAndContinue and asserting Failed(DriverBug) with no subsequent model turn

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 16f9955. Added a full executor caller-path test for ExternalTool plus SkipAndContinue; it asserts Failed(DriverBug), a checkpoint, and no subsequent model turn. The agent-loop suite passes.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed; no code change: ALREADY ADDRESSED — the full executor test covers ExternalTool plus SkipAndContinue and asserts DriverBug with no subsequent model turn. Verification: current-head test audit and workspace clippy passed.

let explanation_message_ref = if completed {
None
} else {
attach_failure_explanation(ctx, &mut state, LoopFailureKind::NoProgressDetected)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — No-progress explanation cancellation path is untested

The new NoProgressDetected branch propagates the only error from attach_failure_explanation via await?: cancellation during prompt/model/finalization. Existing cancellation tests exercise the older Aborted explanation path, while the new no-progress tests cover successful and fail-soft explanations only. They would not catch this branch writing a Final checkpoint or returning Failed instead of Cancelled after cancellation.

Fix: tests::executor::no_progress_explanation_cancellation_returns_cancelled_before_final_checkpoint covering cancellation during the NoProgressDetected explanation call and asserting AgentLoopExecutorError::Cancelled without a final failed checkpoint

Also flagged by: conventions/Medium

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 16f9955. Added the no-progress explanation cancellation test and asserted AgentLoopExecutorError::Cancelled, no final checkpoint, and no finalized assistant message. The agent-loop suite passes.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed; no code change: ALREADY ADDRESSED — the no-progress cancellation test asserts Cancelled, no final checkpoint, and no finalized assistant message. Verification: current-head test audit and workspace clippy passed.

// The parse cause ("missing field `run`", …) rides the
// model-visible Diagnostic channel — scrubbed at the loop
// seam — so the model can correct the plan shape on retry.
.with_model_visible_cause(error.to_string()),

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — Spawn callers do not verify the new model-visible cause

Helper tests assert that malformed and invalid plans initially receive model_visible_cause, but the existing public spawn_capability and resume_spawn_capability contract tests assert only failure kind, disposition, and summary. They do not prove the newly added cause survives the production caller's failure finalization and scrubbing path, so wiring that drops the corrective diagnostic would still pass.

Fix: tests::host_runtime_services_contract::host_runtime_spawn_and_resume_invalid_plan_preserve_model_visible_cause covering both public spawn paths and asserting the returned failure cause names the missing field or failed validation rule

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 16f9955. Both public spawn_capability and resume_spawn_capability contract tests now assert that the returned model-visible cause retains run command must not be empty. The host-runtime services contract passes 114/114.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed; no code change: ALREADY ADDRESSED — both public spawn and resume contracts assert the corrective sandbox cause survives. Verification: current-head contract audit and workspace clippy passed.

// same guest invocation as host infrastructure cannot repair it.
// Surface it as an operation failure so the model can change
// approach or report the broken extension.
DispatchFailureKind::Runtime(RuntimeDispatchErrorKind::Guest) => {

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — WASM host failures are mislabeled as guest operation failures

RuntimeDispatchErrorKind::Guest is not limited to guest traps. wasm_error_kind maps every WasmError::ExecutionFailed to it, while run_wasm_execution_blocking and run_wasm_prepare_blocking construct that variant for a closed execution/preparation gate and for blocking-task panics. This new blanket mapping turns those host-runtime failures into model-visible OperationFailed results, removing the backend retry path and incorrectly asking the model to change its tool call.

Fix: Give gate/join failures an executor/backend-specific WASM error kind, or classify them before this mapping, and map only confirmed guest execution traps to OperationFailed.

Also flagged by: tests/Medium

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 16f9955. Added a typed WasmBlockingError provenance boundary: semaphore acquire and blocking-task join failures map to Executor, while runtime-returned guest execution failures retain Guest. The boundary was extracted into wasm_blocking.rs; focused WASM tests pass 16/16 and host-runtime Clippy is clean.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed; no code change: ALREADY ADDRESSED — WasmBlockingError preserves Executor provenance for host gate/join failures and runtime provenance for guest failures. Verification: current-head code/test audit and workspace clippy passed.

| AgentLoopHostErrorKind::PolicyDenied
| AgentLoopHostErrorKind::CheckpointRejected
| AgentLoopHostErrorKind::TranscriptWriteFailed => None,
| AgentLoopHostErrorKind::Invalid => Some(ModelErrorClass::StaleRequest),

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — Deterministic invalid model requests are treated as stale

InvalidInvocation is the loop-host mapping for every HostManagedModelErrorKind::InvalidRequest, whose contract includes malformed requests and unknown tools. Production producers also use it for invalid routes, provider/model identity mismatches, invalid replay metadata, and missing replay content. Rebuilding the same iteration cannot repair those deterministic faults, but this mapping performs repeated model calls, labels exhaustion model_stale_request, and assigns the Auto retry disposition. Only the genuinely stale-surface case is safe to retry this way.

Fix: Introduce a distinct stale-request error kind across the model gateway boundary and retry only that kind; leave generic InvalidInvocation/Invalid terminal with their accurate category.

Also flagged by: performance/Medium

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 16f9955, with replay-double alignment in 7e64d5d. Added the distinct serialized StaleRequest kind and retry only StaleSurface; generic InvalidRequest/InvalidInvocation is terminal and does not consume another response. Provider output outside the capability surface remains correctly typed InvalidOutput, so the model receives the bounded repair instruction without broadening invalid-request retries. The retry/resume E2E and trace parity targets pass.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed; no code change: ALREADY ADDRESSED — only typed StaleRequest maps to StaleSurface retry; deterministic invalid requests remain terminal. Verification: retry/resume E2E passed 21 tests with 1 credential-gated test ignored.

let plan = serde_json::from_value::<SandboxProcessPlan>(input).map_err(|_| {
AgentLoopHostError::new(
AgentLoopHostErrorKind::InvalidInvocation,
let plan = serde_json::from_value::<SandboxProcessPlan>(input).map_err(|error| {

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — Delete the duplicate sandbox-plan validation path

The loop-host adapter now parses and validates SandboxProcessPlan, wraps failures in a provider-argument error, and sanitizes its own diagnostic even though DefaultHostRuntime::spawn_capability performs the same parse and validation and already returns a model-visible RuntimeCapabilityFailure. Production spawns therefore validate and serialize the plan twice, while validation rules, summaries, and diagnostic handling must remain synchronized across two crates. This also places runtime-specific request-shape logic in the upper adapter instead of the runtime owner.

Fix: Remove host_runtime_input_for_capability, sandbox_plan_input_error, and safe_sandbox_plan_cause; pass the provider-normalized JSON directly to HostRuntime::spawn_capability and let the existing runtime-failure mapper turn the host runtime's InvalidInput outcome and model-visible cause into the loop result. Test doubles should implement that HostRuntime contract rather than requiring a second production preflight.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 16f9955. Removed loop-host's duplicate sandbox parse/validation helpers and now pass provider-normalized JSON directly to host-runtime. The loop-host test double implements the runtime contract and asserts one spawn attempt, preserving caller coverage without a second production validator.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed; no code change: ALREADY ADDRESSED — duplicate loop-host sandbox validation helpers are absent and runtime remains the validation owner. Verification: full ironclaw_loop_host suite and workspace clippy passed.

ModelInvalidOutput,
ModelUnavailable,
ModelInternal,
ModelStaleRequest,

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — New stale-request attempt class lacks checkpoint round-trip coverage

ModelStaleRequest is added to the serialized RecoveryAttemptClass used as a persistent BTreeMap key. Existing checkpoint lifecycle coverage round-trips ModelInvalidOutput and the generic state builder uses ModelTransient; no test serializes and reloads the new wire value model_stale_request. Replay/checkpoint compatibility for the new retry counter is therefore unverified.

Fix: tests::state_lifecycle::model_stale_request_recovery_attempts_survive_checkpoint_reload covering a BeforeModel checkpoint containing RecoveryAttemptClass::ModelStaleRequest and asserting its count after deserialization

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 16f9955. Added a BeforeModel checkpoint serialization/reload test with two ModelStaleRequest attempts and asserted the exact count after restoration. The agent-loop lifecycle and full crate suites pass.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed; no code change: ALREADY ADDRESSED — checkpoint lifecycle coverage round-trips ModelStaleRequest attempts and asserts the restored count. Verification: current-head test audit and workspace clippy passed.

/// strategy outcome (§5a.1, docs/plans/2026-07-03-loop-failure-matrix.md): a
/// `SkipAndContinue` on an Approval / AwaitDependentRun / ExternalTool gate is
/// a strategy-contract violation, and silently skipping the gated call would
/// mask a driver bug behind a normal completion. The invalid outcome is

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Low — Gate outcome documentation still recommends a now-invalid skip

The new enforcement makes SkipAndContinue invalid for Approval, AwaitDependentRun, and ExternalTool gates, but the owning GateOutcome documentation still says this variant is intended for tools where a missing approval is non-fatal. A strategy author following that local documentation will now produce a DriverBug abort. The new comment also calls the conversion to the more severe Abort outcome a downgrade, which obscures the actual behavior.

Fix: Update the GateOutcome::SkipAndContinue documentation in strategies/gate.rs to name Auth and Resource as the valid skip kinds and explicitly state that Approval, AwaitDependentRun, and ExternalTool skips are converted to Abort DriverBug; describe the conversion here as enforcement rather than a downgrade.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 16f9955. The owning docs now state that SkipAndContinue is valid only for Auth and Resource gates, name the three enforced Abort(DriverBug) cases, and describe the conversion as enforcement rather than a downgrade.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed; no code change: ALREADY ADDRESSED — GateOutcome documentation names Auth and Resource as valid skip kinds and the three enforced DriverBug cases. Verification: current-head source audit.

@railway-app

railway-app Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6437 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 22, 2026 at 1:26 pm

@serrrfirat
serrrfirat marked this pull request as ready for review July 21, 2026 21:53
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

❌ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
❌ Changes requested 1 0 1 99cb2dc8f809

Head: 99cb2dc8f8094c160636448cad90c5b5005ba3e0
Next: Fix the blocking findings, push the PR branch, then re-run this reviewer.

Run details

Status: Current
Needs human: no
Needs validation: no

Summary

Changes requested: sandbox validation diagnostics bypass the canonical secret and prompt-injection scrubber before becoming model-visible.

Findings

Blocking: 1 / Notes: 0

Blocking findings

1. ❌ [MEDIUM] Scrub sandbox validation causes before exposing them to the model

Location: crates/ironclaw_loop_host/src/capability_port.rs:2847-2854
ProcessSandboxPlanError embeds model-supplied host/path/env strings, but this sanitizer only replaces delimiters before passing the text to CapabilityFailureDetail::Diagnostic. That bypasses scrub_model_visible_detail's leak-detector and injection fencing. A hostile or credential-shaped invalid value can reach model context unredacted, or cause the later transcript validator to drop the entire observation and lose the corrective detail. Route this through the canonical scrubber before applying SafeSummary formatting, and add hostile host/path coverage.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.

/// is bounded. Returns `None` when nothing legible remains.
fn safe_sandbox_plan_cause(raw: &str) -> Option<String> {
const MAX_BYTES: usize = 400;
let sanitized: String = raw

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This turns a ProcessSandboxPlanError containing model-supplied host/path/env values into a model-visible diagnostic with only character replacement. It bypasses the canonical secret/prompt-injection scrubber, so hostile or credential-shaped input can be exposed or make the later transcript validator drop the recovery detail. Please route the cause through scrub_model_visible_detail before publishing it.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 16f9955. The character-only sandbox helper was removed with the duplicate loop-host validator. Runtime-owned validation causes now flow through the canonical secret scrubber and injection scanner, then a compact explicit untrusted-data fence and the legacy SafeSummary bound. A regression test covers both an api_key token and injection-shaped text.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed; no code change: DUPLICATE/ALREADY ADDRESSED — runtime-owned sandbox causes cross the canonical secret scrubber and injection fence, with hostile-input caller coverage. Verification: full ironclaw_loop_host suite passed.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/reborn_trace_error_path_parity.rs`:
- Around line 126-136: Update the assertions in the completed-run test to
inspect the second entry from harness.model_requests() and verify it contains a
ToolResult carrying the unadvertised-capability rejection diagnostic. Keep the
existing invocation, response-count, and request-count assertions, matching the
malformed-input test’s assertion style.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: f43b57ac-1d01-4eb7-9d43-7cb1264f4a0e

📥 Commits

Reviewing files that changed from the base of the PR and between 9c9c77b and 99cb2dc.

📒 Files selected for processing (28)
  • crates/ironclaw_agent_loop/src/executor/gates.rs
  • crates/ironclaw_agent_loop/src/executor/loop_exit.rs
  • crates/ironclaw_agent_loop/src/executor/mapping.rs
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_agent_loop/src/executor/tests/failure_matrix.rs
  • crates/ironclaw_agent_loop/src/families/mod.rs
  • crates/ironclaw_agent_loop/src/families/subagent.rs
  • crates/ironclaw_agent_loop/src/state/slots.rs
  • crates/ironclaw_agent_loop/src/strategies/recovery.rs
  • crates/ironclaw_agent_loop/tests/safety_nets.rs
  • crates/ironclaw_host_runtime/src/production.rs
  • crates/ironclaw_host_runtime/src/services/runtime_adapters.rs
  • crates/ironclaw_loop_host/src/capability_port.rs
  • crates/ironclaw_runner/src/failure_lane.rs
  • crates/ironclaw_runner/src/failure_summary.rs
  • crates/ironclaw_runner/src/loop_exit_applier/tests/mod.rs
  • crates/ironclaw_runner/src/retry_disposition.rs
  • crates/ironclaw_runner/src/turn_runner.rs
  • crates/ironclaw_turns/src/loop_exit.rs
  • crates/ironclaw_turns/src/loop_exit/tests/mod.rs
  • docs/plans/2026-07-03-loop-failure-matrix.md
  • tests/integration/cancel.rs
  • tests/integration/support/builder.rs
  • tests/integration/support/group.rs
  • tests/integration/support/scripted_provider.rs
  • tests/reborn_failure_retry_resume_e2e.rs
  • tests/reborn_trace_error_path_parity.rs
  • tests/support/reborn_parity_qa/binary_e2e.rs

Comment thread tests/reborn_trace_error_path_parity.rs Outdated
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6437 July 21, 2026 22:08 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6437 July 21, 2026 22:15 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_host_runtime/src/services/wasm_blocking.rs`:
- Around line 24-30: Update the WASM execution and preparation admission flow
around WASM_EXEC_SEMAPHORE, WASM_PREPARE_SEMAPHORE, and their
acquire_owned().await calls to prevent indefinite process-wide queueing: add a
bounded acquire timeout and map timeout failures to the existing retryable
dispatch error, or implement equivalent per-scope fairness using the existing
ResourceGovernor scope context. Preserve the current concurrency limits while
ensuring one tenant cannot indefinitely starve other scopes.

In `@crates/ironclaw_runner/src/model_gateway.rs`:
- Around line 2761-2785: The existing test only validates
map_capability_host_error, so add a regression test through the caller that
invokes issue_host_prompt_bundle with a surface-version mismatch and asserts the
resulting gateway error is StaleRequest. Extend the
capability_model_request_errors_preserve_stale_distinction table with
ScopeMismatch if it should continue mapping to InvalidRequest.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 3d995348-a804-4f6c-a55e-485444dff7b8

📥 Commits

Reviewing files that changed from the base of the PR and between 99cb2dc and 16f9955.

📒 Files selected for processing (15)
  • crates/ironclaw_agent_loop/src/executor/gates.rs
  • crates/ironclaw_agent_loop/src/executor/mapping.rs
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_agent_loop/src/strategies/gate.rs
  • crates/ironclaw_agent_loop/tests/state_lifecycle.rs
  • crates/ironclaw_host_runtime/src/services.rs
  • crates/ironclaw_host_runtime/src/services/runtime_adapters.rs
  • crates/ironclaw_host_runtime/src/services/wasm_blocking.rs
  • crates/ironclaw_host_runtime/src/services/wasm_execution.rs
  • crates/ironclaw_host_runtime/tests/host_runtime_services_contract.rs
  • crates/ironclaw_loop_host/src/capability_port.rs
  • crates/ironclaw_loop_host/src/lib.rs
  • crates/ironclaw_loop_host/src/model_visible_scrub.rs
  • crates/ironclaw_runner/src/model_gateway.rs
  • tests/reborn_failure_retry_resume_e2e.rs

Comment on lines +24 to +30
/// Process-wide gate over concurrent native WASM execution.
pub(super) static WASM_EXEC_SEMAPHORE: std::sync::LazyLock<Arc<Semaphore>> =
std::sync::LazyLock::new(|| Arc::new(Semaphore::new(MAX_CONCURRENT_WASM_EXEC)));

/// Process-wide gate over concurrent WASM component compilation.
pub(super) static WASM_PREPARE_SEMAPHORE: std::sync::LazyLock<Arc<Semaphore>> =
std::sync::LazyLock::new(|| Arc::new(Semaphore::new(MAX_CONCURRENT_WASM_PREPARE)));

@coderabbitai coderabbitai Bot Jul 21, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚀 Performance & Scalability | 🔵 Trivial

Global WASM concurrency gates are process-wide, not tenant-scoped, and have no acquire timeout.

WASM_EXEC_SEMAPHORE/WASM_PREPARE_SEMAPHORE bound total concurrency well, but a single tenant's burst of long-running WASM calls can exhaust all 64/16 slots process-wide and starve every other tenant queued on acquire_owned().await indefinitely — there's no fairness partitioning and no deadline, so the failure mode under sustained overload is unbounded queuing rather than load-shedding. Worth considering per-scope fairness (or at least a bounded acquire timeout that maps to a retryable dispatch error) given the existing scope-aware ResourceGovernor reservation already tracks tenant/scope for the same call.

Also applies to: 66-89

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_host_runtime/src/services/wasm_blocking.rs` around lines 24 -
30, Update the WASM execution and preparation admission flow around
WASM_EXEC_SEMAPHORE, WASM_PREPARE_SEMAPHORE, and their acquire_owned().await
calls to prevent indefinite process-wide queueing: add a bounded acquire timeout
and map timeout failures to the existing retryable dispatch error, or implement
equivalent per-scope fairness using the existing ResourceGovernor scope context.
Preserve the current concurrency limits while ensuring one tenant cannot
indefinitely starve other scopes.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Skipped for this PR after verification. The process-wide semaphores and unbounded acquire_owned() waits are real, but there is no existing retryable dispatch error that can represent an admission timeout without changing runtime contracts and recovery semantics across layers. This PR intentionally keeps timeout provenance in the executor lane until that typed contract exists; adding per-scope fairness or load-shedding belongs in a dedicated resource-governance change. The existing 64/16 concurrency bounds remain unchanged.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@serrrfirat, agreed—this remains a valid capacity-isolation concern, but a timeout without a typed retryable admission-failure contract would be an incomplete cross-layer change. Deferring it keeps this PR focused and preserves the existing 64/16 bounds.

Would you like me to open a follow-up GitHub issue for tenant-aware WASM admission control / bounded load-shedding?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed; no code change: QUESTION/NON-ACTIONABLE for this PR — bounded admission or tenant fairness requires a typed cross-layer load-shedding contract and would materially expand the change; the thread already agreed to defer it. Existing 64/16 concurrency bounds are unchanged.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@serrrfirat, acknowledged. The concern is deferred as non-actionable for #6437: the current process-wide 64/16 bounds are preserved, and introducing timeout/load-shedding or tenant-aware admission correctly requires a dedicated typed cross-layer contract.

Comment thread crates/ironclaw_runner/src/model_gateway.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6437 July 21, 2026 23:14 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/reborn_failure_retry_resume_e2e.rs`:
- Around line 32-34: Revise the documentation comment above the affected test to
describe recovery from a stale model request injected at the model gateway,
rather than claiming recovery of a stale capability-surface mismatch. Remove
assertions about cross-layer capability recovery unless the test is expanded to
exercise an actual surface change.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: d237626c-5a11-4092-aa53-ddaba8a08f5e

📥 Commits

Reviewing files that changed from the base of the PR and between 7e64d5d and 73665bc.

📒 Files selected for processing (10)
  • crates/ironclaw_runner/src/model_gateway.rs
  • tests/integration/support/doubles/recording_network_http_egress.rs
  • tests/integration/support/doubles/recording_test_capability_port.rs
  • tests/integration/support/harness/mod.rs
  • tests/integration/support/harness/recorder.rs
  • tests/integration/web_access.rs
  • tests/reborn_failure_retry_resume_e2e.rs
  • tests/reborn_trace_error_path_parity.rs
  • tests/reborn_trace_wasm_github_fixture_parity.rs
  • tests/support/reborn_parity_qa/binary_e2e.rs

Comment thread tests/reborn_failure_retry_resume_e2e.rs Outdated
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6437 July 22, 2026 09:01 Destroyed
…coverability-item1

# Conflicts:
#	crates/ironclaw_host_runtime/src/services.rs
#	crates/ironclaw_host_runtime/src/services/runtime_adapters.rs
#	crates/ironclaw_host_runtime/src/services/wasm_execution.rs
#	crates/ironclaw_loop_host/src/capability_port.rs
#	tests/integration/support/doubles/recording_network_http_egress.rs
#	tests/integration/support/doubles/recording_test_capability_port.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6437 July 22, 2026 09:23 Destroyed
@github-actions

github-actions Bot commented Jul 22, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 86.26% (305059 / 353671 lines)
  floor:    85.3% (tolerance 0.5pp -> effective floor 84.8%)
  denominator: 353671 lines now vs 320188 at floor capture (+33483 lines, +10.46%) — material change (>5%)

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 86.26% — 305059 / 353671 lines

Per-crate breakdown (62 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_telegram_extension 34.88% 60 / 172
ironclaw_event_projections 43.31% 673 / 1554
ironclaw_dispatcher 60% 72 / 120
ironclaw_observability 61.54% 16 / 26
ironclaw_authorization 62.98% 609 / 967
ironclaw_memory 69.2% 773 / 1117
ironclaw_trust 72.88% 661 / 907
ironclaw_filesystem 73.5% 4543 / 6181
ironclaw_capabilities 74.07% 2717 / 3668
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_extractors 74.72% 538 / 720
ironclaw_projects 76.48% 400 / 523
ironclaw_triggers 77.33% 2531 / 3273
ironclaw_mcp 77.56% 736 / 949
ironclaw_reborn_cli 78.25% 10379 / 13264
ironclaw_product_context 78.57% 11 / 14
ironclaw_llm 78.63% 20824 / 26485
ironclaw_wasm 79.72% 735 / 922
ironclaw_process_sandbox 80.46% 671 / 834
ironclaw_memory_native 81.17% 3195 / 3936
ironclaw_events 81.95% 1594 / 1945
ironclaw_first_party_extensions 82.58% 6608 / 8002
ironclaw_reborn_event_store 83.03% 1169 / 1408
ironclaw_telegram_v2_adapter 83.07% 2017 / 2428
ironclaw_product_adapter_registry 83.43% 574 / 688
ironclaw_reborn_identity 83.59% 433 / 518
ironclaw_processes 83.78% 940 / 1122
ironclaw_secrets 83.8% 2550 / 3043
ironclaw_reborn_config 84.17% 1962 / 2331
ironclaw_product_workflow 84.79% 13158 / 15518
ironclaw_auth 85.01% 4011 / 4718
ironclaw_common 85.18% 1741 / 2044
ironclaw_product_adapters 85.29% 3351 / 3929
ironclaw_run_state 85.61% 458 / 535
ironclaw_network 85.97% 913 / 1062
ironclaw_hooks 86.6% 9931 / 11468
ironclaw_extensions 87.02% 3795 / 4361
ironclaw_threads 87.17% 4849 / 5563
ironclaw_host_api 87.52% 5394 / 6163
ironclaw_skills 87.58% 4470 / 5104
ironclaw_reborn_traces 88.11% 11972 / 13587
ironclaw_reborn_composition 88.38% 57807 / 65404
ironclaw_turns 88.55% 14473 / 16345
ironclaw_host_runtime 88.66% 18271 / 20608
ironclaw_webui 88.95% 8064 / 9066
ironclaw_reborn_openai_compat 89.03% 3627 / 4074
ironclaw_extension_host 89.59% 2856 / 3188
ironclaw_slack_extension 89.7% 2439 / 2719
ironclaw_approvals 90.18% 1598 / 1772
ironclaw_conversations 90.39% 3123 / 3455
ironclaw_resources 90.85% 4477 / 4928
ironclaw_runner 91.2% 16944 / 18579
ironclaw_event_streams 91.24% 1063 / 1165
ironclaw_loop_host 92.26% 16258 / 17622
ironclaw_attachments 93.06% 630 / 677
ironclaw_agent_loop 94.93% 9603 / 10116
ironclaw_safety 95.15% 3749 / 3940
ironclaw_first_party_extension_ports 95.62% 3672 / 3840
ironclaw_outbound 95.77% 3513 / 3668
ironclaw_runtime_policy 96.55% 811 / 840

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_loop_host/src/capability_port.rs`:
- Around line 3437-3464: Add a focused unit test for
sandbox_model_visible_diagnostic_text that supplies a diagnostic longer than 400
bytes containing multi-byte UTF-8 characters, then verifies the result is
truncated to at most 400 bytes and remains valid UTF-8 without splitting a
character. Keep the test targeted to the function’s byte-limit and
is_char_boundary walk-back behavior.

In `@crates/ironclaw_loop_host/src/lib.rs`:
- Around line 2343-2354: Replace or supplement
model_gateway_error_preserves_stale_request_kind with a caller-driven regression
test that invokes the production executor path using a stale model request.
Verify that the resulting StaleSurface triggers iteration-scoped retries and is
classified as terminal after retry exhaustion, rather than testing
model_gateway_error directly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 76d8065d-e7b8-4868-a48c-a4fe66590335

📥 Commits

Reviewing files that changed from the base of the PR and between ed0237c and 60afcf8.

📒 Files selected for processing (8)
  • crates/ironclaw_agent_loop/tests/safety_nets.rs
  • crates/ironclaw_host_runtime/src/production.rs
  • crates/ironclaw_host_runtime/src/services.rs
  • crates/ironclaw_host_runtime/src/services/runtime_adapters.rs
  • crates/ironclaw_host_runtime/src/services/wasm_execution.rs
  • crates/ironclaw_host_runtime/tests/host_runtime_services_contract.rs
  • crates/ironclaw_loop_host/src/capability_port.rs
  • crates/ironclaw_loop_host/src/lib.rs

Comment thread crates/ironclaw_loop_host/src/capability_port.rs
Comment thread crates/ironclaw_loop_host/src/lib.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6437 July 22, 2026 09:49 Destroyed
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

/canary

@github-actions

Copy link
Copy Markdown
Contributor

Started Reborn WebUI v2 live canary for codex/reborn-error-recoverability-item1 at e2073a31e3 with cases all: https://github.com/nearai/ironclaw/actions/runs/29910684485

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6437 July 22, 2026 13:16 Destroyed
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

/canary

@github-actions

Copy link
Copy Markdown
Contributor

Started Reborn WebUI v2 live canary for codex/reborn-error-recoverability-item1 at 0d898b00eb with cases all: https://github.com/nearai/ironclaw/actions/runs/29923209039

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6437 — 0d898b00 Deployed Jul 22, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants