test(reborn): LoopFailureKind fault-coverage matrix (100% error types + main comparison) - #5613
serrrfirat wants to merge 4 commits into
Conversation
…ive category table Phase 1 of the per-error coverage harness (docs/plans/2026-07-03-loop-failure-matrix.md): - turns: all_failure_kinds category table now exhaustive (13/13 — adds the previously-missing CheckpointUnavailable + CompactionUnavailable) with a same-crate exhaustive-match guard so a new variant breaks compilation. - agent_loop: new table-driven executor failure matrix (executor/tests/failure_matrix.rs) driving every executor-reachable LoopFailureKind at its real origin via MockHost seams, asserting per row: P1 (reason_kind + sanitized category/safe_summary), P3 (no fabricated final assistant reply), and explanation_message_refs presence per the explainable set. New fail_transcript_with test knob on MockHost + DriverMockHost (test code only). - Four divergences found and documented (doc §5a), asserted as actual behavior, none silently fixed: Approval+SkipAndContinue completes (gate enforcement gap), NoProgressDetected missing its explanation attach, single Denied recovers-and-completes (no-borking working as designed), TranscriptWriteFailed/CheckpointRejected legacy-only enum origins. Validated (bounded): cargo test -p ironclaw_turns all_failure_kinds; cargo test -p ironclaw_agent_loop failure_matrix; check/clippy -D warnings/ fmt on ironclaw_agent_loop — all green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Phase 2 of the per-error coverage harness (docs/plans/2026-07-03-loop-failure-matrix.md):
- planned_driver: resume with missing checkpoint payload asserts
LoopFailureKind::CheckpointUnavailable + "checkpoint_unavailable";
in-flight model Cancelled (no cooperative cancel signal) asserts
map_executor_error yields "interrupted_unexpectedly".
- e2e: binary-level divergence lock — the same in-flight Cancelled run
projects "driver_failed" at the runner boundary (category overwritten;
doc §5a.5, candidate follow-up to preserve the driver-mapped category).
- e2e: non-model P4 row — capability-stage invocation failure is
retryable ("host_stage_unavailable_capability", checkpoint preserved,
no fabricated reply) and retry_run resumes to completion, so P4 is no
longer proven only through the model stage. New scripted
capability-invocation-error mode in the test harness (test support only).
Validated (bounded): targeted planned_driver tests + full
reborn_failure_retry_resume_e2e (15 passed) + clippy -D warnings + fmt.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
#5390) This matrix PR sits on top of the recoverability stack (main←#4841←#5389←#5390←#5403), so it validates the FIXED + classified system rather than #4841's base behavior: - Add executor rows for #5389's model-fixable capability failures that are now RECOVERABLE (InvalidInput / InvalidOutput / PolicyDenied): where the base branch terminated the run, the stack now surfaces a model-visible tool error and the loop completes. Rows assert the recovered outcome. - Add FailureLane / RetryDisposition alignment (binary/e2e layer, since ironclaw_agent_loop can't depend on ironclaw_reborn_composition): each real failure path's (category, retryable) is asserted to map to the expected #5390 FailureLane bucket + RetryDisposition — proving the real paths feed the classifier correctly. Complements #5390's classifier unit tests (which test the functions directly) rather than duplicating them. - Doc: §6 relationship to #5390; §5a marks divergences RESOLVED by the stack (capability recoverability) vs still-open (Approval SkipAndContinue completes; NoProgressDetected lacks explanation; Transcript/Checkpoint are planned-executor host errors; interrupted_unexpectedly projected as driver_failed at the runner boundary). The cherry-picked base matrix assertions still passed unchanged on the stack — they assert actual behavior, so the additive fixes did not break them; this commit adds the stack-specific coverage on top. Validated (bounded, on the stack): cargo test ironclaw_turns all_failure_kinds; cargo test ironclaw_agent_loop failure_matrix; cargo test --test reborn_failure_retry_resume_e2e (19 passed); clippy -D warnings; fmt --check. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. 🗂️ Base branches to auto review (2)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Code Review
This pull request introduces a comprehensive failure matrix test suite for the AgentLoopExecutor to verify how various failure scenarios are handled, sanitized, and mapped. It adds a new failure_matrix.rs test file, updates mock hosts to support simulating transcript write failures, introduces exhaustiveness guards for LoopFailureKind variants, and adds end-to-end tests for failure lane alignment. The review feedback identifies a compilation error in planned_driver.rs due to a missing detail field in AgentLoopDriverError::Failed, points out an unused import in the new failure matrix test file, and suggests refactoring the sequential matrix test execution into individual test cases for better diagnostics.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| assert_eq!( | ||
| result, | ||
| Err(AgentLoopDriverError::Failed { | ||
| reason_kind: "interrupted_unexpectedly".to_string() | ||
| }) | ||
| ); |
There was a problem hiding this comment.
The initialization of AgentLoopDriverError::Failed is missing the detail field, which will cause a compilation error. Struct variants of enums in Rust require all fields to be specified during construction. Please add detail: None to the struct initializer.
assert_eq!(
result,
Err(AgentLoopDriverError::Failed {
reason_kind: "interrupted_unexpectedly".to_string(),
detail: None,
})
);References
- Avoid using generic
is_err()assertions in tests when verifying that an operation fails. Instead, assert against the specific expected error message or variant to prevent infrastructure or harness-level failures from causing false positives.
| CanonicalAgentLoopExecutor, CapabilityFailureKind, CapabilityOutcome, CapabilityResultMessage, | ||
| CheckpointKind, DefaultCompactionStrategy, FixedReplyAdmissionPolicy, GateOutcome, HostStage, | ||
| LoopCheckpointKind, LoopCompactionError, LoopExecutionState, LoopExit, LoopFailureKind, | ||
| LoopGateRef, LoopResultRef, LoopRunInfoPort, LoopSafeSummary, MockHost, |
There was a problem hiding this comment.
The import LoopRunInfoPort is not used anywhere in this file. Under strict compiler flags or #![deny(unused_imports)], this will cause a compilation failure. Please remove the unused import.
| LoopGateRef, LoopResultRef, LoopRunInfoPort, LoopSafeSummary, MockHost, | |
| LoopGateRef, LoopResultRef, LoopSafeSummary, MockHost, |
| #[tokio::test] | ||
| async fn executor_layer_failure_matrix() { | ||
| for row in ROWS { | ||
| let observed = run_setup(row.setup).await; | ||
| assert_expected_terminal(row, &observed); | ||
| } | ||
| } |
There was a problem hiding this comment.
Running all matrix rows sequentially in a single test case means that if any row fails, the test execution halts immediately, preventing subsequent rows from being evaluated. It also makes it harder to identify which specific row failed from the test runner's top-level output. Consider using a macro to generate individual test cases for each row, or at least executing them in parallel/reporting all failures.
❌ IronLoop Review StatusHead:
Configuration errorMessage: Unable to load trusted agent config from .ironloop/agents.yaml. Available commands
Run metadataOrigin: |
|
/canary all |
|
Started Reborn WebUI v2 live canary for |
|
⛔️ Superseded by #5692 — the recoverability stack was collapsed onto the refreshed #4841 head as a single PR (#5692). Please do not review this branch; it is stale (~139 commits behind main and un-restacked). Kept open only as a fallback until #5692 lands, after which this will be closed as merged-via-#5692. |
… + #5389/#5390/#5403/#5613) (#5692) * reborn: add failure explanations and retryable failed runs * test(loop_support): set inline_messages on the contract-test LoopModelRequest #4841 added the `inline_messages` field (serde default) to LoopModelRequest but missed this one construction in the thread_loop_support_contract integration test, breaking that test target's compile. Production builds default it to Vec::new(); match that. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(reborn): resolve main-merge CI breaks on #4841 (retry_turn stub + finer failure categories) The main→#4841 merge surfaced two semantic conflicts the auto-merge missed: - Clippy: main added `TurnCoordinator::retry_turn`; the StaticTurnCoordinator test stub in openai_compat_serve/tests.rs didn't implement it. Add the stub (returns Unavailable, matching its other methods). - Test ironclaw_reborn: main's chaos tests (#5296) assert the coarse failure categories "driver_unavailable"/"model_error", but #4841 refined production to finer, accurate categories — a full checkpoint-state disk now yields "host_stage_unavailable_checkpoint" and an offline model provider yields "model_unavailable" (both deliberate named categories with dedicated failure_summary messages). Update the two stale assertions to match #4841's intended categorization. Verified: the two turn_runner_worker_full_reborn_fails_* tests pass; clippy ironclaw_reborn_composition --all-features clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Persist retry busy idempotency records * test(reborn): expect specific model unavailable failure * fix(reborn): preserve same-run checkpoint refs * fix(ci): update reborn compile drift * fix(ci): update composition test drift * feat(reborn): make model-fixable capability failures recoverable (batch 1) Turns recoverable→bork mis-mappings into model-visible tool errors so the agent self-corrects instead of the run dying. On top of #4841. - agent_loop keystone: capability_error_class re-buckets Dispatcher / InvalidOutput / Unknown(_) / non-exhaustive default from Permanent (Abort) to OperationFailed (ToolErrorResult), aligning with the host_runtime disposition layer (which never intends a capability failure to abort). Cancelled / Permanent stay terminal. This makes "model called a nonexistent tool" (UnknownCapability/UnknownProvider -> InvalidOutput) recoverable. - outbound_delivery: outbound_delivery_outcome routes recoverable RebornServicesErrorCode to Ok(Failed/Denied) (only Internal -> Err); fixed the safe_summary that interpolated the model-supplied target_id (Invariant 2); expired/not-yet-approved approval-lease arms -> Ok(Denied) instead of terminal. - host_runtime: malformed model-supplied SandboxProcessPlan -> recoverable Failed{InvalidInput} outcome (defense-in-depth; the live gate in loop_support's host_runtime_input_for_capability is fixed in a follow-up). Lib tests green: agent_loop 354, host_runtime 297, loop_support 337, reborn 259; outbound_delivery 26. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(loop_support): malformed sandbox plan is recoverable, not run-ending Completes the sandbox-plan fix. The live terminal gate is loop_support's host_runtime_input_for_capability: a malformed/invalid model-supplied SandboxProcessPlan returned AgentLoopHostError::InvalidInvocation, which capability_host_error maps to terminal HostUnavailable{Capability} (run dies). Now the invoke path downgrades that InvalidInvocation to a model-visible Ok(CapabilityOutcome::Failed{InvalidInput}) so the agent can correct the plan and the run continues. The helper only emits InvalidInvocation for the sandbox-plan parse/validation case; its host-internal serialization failure keeps its Internal Err. Updated both locked tests to assert the recoverable outcome. loop_support lib: 337 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(llm): provider error fidelity for accurate recover/explain (batch 2) Maps provider failures to the right LlmError variant so model-call errors are explained accurately and context overflow recovers via context-shrink instead of borking. - rig_adapter (OpenAI/Anthropic/Ollama/Tinfoil/openai_compatible): map_rig_error now detects auth failures (401/403/invalid key) -> AuthFailed (non-retryable, non-breaker-tripping) instead of generic RequestFailed, so a bad key surfaces as a credentials problem rather than wasted retries + opaque run-bork. - Codex (openai_codex_provider + codex_chatgpt): a stream ending without response.completed is now a retryable InvalidResponse/EmptyResponse instead of a silent successful Stop; codex_chatgpt now maps SSE error/response.failed events; both detect 413/context-overflow -> ContextLengthExceeded. - github_copilot + anthropic_oauth: detect 413 (and 400+context body) -> ContextLengthExceeded so context-shrink recovery fires (401/429/5xx untouched). ironclaw_llm lib: 915 passed, 0 failed. clippy + fmt clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(host_runtime): unknown method/capability is immediate model-visible error (batch 3) Sweep finding: a method/capability the model named that does not exist (RuntimeDispatchErrorKind::MethodMissing / UndeclaredCapability) mapped to RuntimeFailureKind::Backend -> RetrySameCall, so it burned the retry budget before becoming model-visible. Retrying never resolves a nonexistent target. Now maps to InvalidInput -> ModelVisibleToolError: the model gets an immediate "no such method/capability" tool error and self-corrects. Updated the pinning table entries. host_runtime lib: 297 passed. Batch-3 sweep conclusion: after the keystone + batches 1-2, no remaining recoverable->bork CORRECTNESS defects exist (tool backends fully clean; no hard-Err bypass on model-fixable conditions; nothing mapped to terminal Cancelled/Permanent). This was the last shape-#3 quality nit worth fixing now. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(reborn): FailureLane classifier + two-bucket enforcement test Item #2 foundation, built ON TOP of #4841's failure-surfacing machinery (reuses category + FailureExplanationProvider + retryable rather than a parallel RunFailureReason taxonomy). - FailureLane enum (Retriable | Explainable | Security), wire-stable snake_case. - failure_lane(category, retryable): retryable -> Retriable, else Explainable. Security is reserved for the ingress safety/leak refusal path (minimal security-stop policy) and is never produced at the run boundary; the match on category is the seam for a future mid-run safety-abort category. - ALL_RUN_FAILURE_CATEGORIES: canonical list of every category the run boundary can produce. - ENFORCEMENT TEST (every_failure_category_is_explainable_and_classified): locks the two-bucket invariant — every failure category resolves to a SPECIFIC user explanation (never the generic fallback) AND a definite lane. A new category that forgets its sentence, or regresses to the generic fallback, fails here. Plus canonical_list_covers_loop_failure_kinds guards against list drift. reborn_composition failure_lane: 5 passed. clippy clean (the one pre-existing needless_return in local_runtime_profile.rs is unrelated). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(reborn): retry-disposition policy (hybrid retry core) Operationalizes the hybrid retry decision on top of the FailureLane classifier. Pure decision function; the auto-redrive scheduler is its consumer. - RetryDisposition { Auto | UserInitiated | NoRetry }, wire-stable snake_case. - retry_disposition(category, retryable): no checkpoint -> NoRetry; transient host/lease/store/provider/tool faults -> Auto (silent re-drive from checkpoint, bounded by the scheduler); model/provider/config/model-fixable faults -> UserInitiated (retry affordance; a silent re-drive would just re-fail). Conservative Auto allowlist (anything not clearly transient -> UserInitiated). - RetryDisposition::failure_lane() ties it back to FailureLane; a test asserts the two layers agree for every category in ALL_RUN_FAILURE_CATEGORIES. reborn_composition retry_disposition: 5 passed. clippy + fmt clean. Follow-up: the scheduler wiring that calls retry_disposition() to auto-requeue (the behavior-flipping "Auto" half) — this is its tested decision core. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(reborn): carry secret-scrubbed raw cause to the model Reborn over-sanitized capability/host failures: a real cause like `missing input_schema_ref at /system/extensions/.../list_calendars.input.v1.json` was collapsed to the generic "host runtime rejected capability request" because there was no model-visible field to carry the raw cause and the summary validator rejected any string containing `/`. Policy shift: redact secret VALUES only; let paths, codes, schema refs, and raw error text reach the model so it can retry or explain. Foundation + Tier-1 vertical: - AgentLoopHostError gains an optional model-visible `detail: Option<String>` channel (+ `with_detail`). - CapabilityFailureDetail gains a free-text `Diagnostic { text }` variant. - ToolObservationDetail::GenericFailure gains a bounded, leniently-validated `detail` (allows `/ { } [ ] < >`, rejects NUL/control + caps length) — the channel that already reaches the model and bypasses the strict summary validator. - Relax ONLY the false-positive word bans in validate_loop_safe_summary and validate_tool_result_safe_summary (drop "provider error", "stack trace", "tool input", "traceback", "host path", "raw runtime", "invalid api key"); keep the delimiter ban, control-char ban, length cap, and credential markers. - Tier-1 producers stop dropping the cause: raw_agent_loop_host_error threads the value-scrubbed raw_detail into AgentLoopHostError.detail; the runtime model-visible failure path carries a value-scrubbed Diagnostic when the strict summary validator drops the reason; capability_helpers forwards the diagnostic into the model-visible observation. - Boxed ProviderArgumentError.error to keep result_large_err quiet after the AgentLoopHostError/CapabilityFailureDetail size growth. Tests cover the anchor (path string reaches the model-visible detail), secret value redaction, the relaxed/retained summary markers, and legacy GenericFailure JSON round-trip. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(reborn): MCP per-cause error tokens + explainer detail plumbing (tiers 2a, 3.2) - ironclaw_mcp: replace the flat "response_error"/"request_denied" literals with per-cause diagnostic tokens (mcp_http_status_<code>, mcp_jsonrpc_error code=..., mcp_parse_failed, ...), bounded + control-char-stripped, no public signature change. The model now learns the real HTTP status / JSON-RPC code. - reborn_composition: FailureExplanationInput gains a `detail` field rendered into the failure-explanation prompt (secret-scrubbed via sanitize_model_visible_text). Wired end-to-end in the projection; sourced once TurnLifecycleEvent carries detail (upstream chain in a follow-up commit). mcp --lib 18 passed; reborn_composition --lib failure_explanation tests pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(reborn): thread secret-scrubbed failure detail to the explainer (tiers 2b, 3.1) Complete the model-error-detail chain so the failure explainer (and the model) receive the real cause of a model/provider/driver fault instead of only a sanitized category. Only secret VALUES are withheld (scrubbed via the existing value-level redactors); the descriptive cause now flows end-to-end. Carrier `detail: Option<String>` (serde default + skip_serializing_if, so pre-detail persisted rows rehydrate as None) added and threaded through: - ironclaw_loop_support: HostManagedModelError.detail + with_detail; threaded in model_gateway_error into AgentLoopHostError.detail. - ironclaw_agent_loop: AgentLoopExecutorError::HostUnavailableWithDiagnostics gains detail; model-stage construction carries error.detail. - ironclaw_turns: AgentLoopDriverError::Failed.detail; TurnLifecycleEvent.detail (Failed events only, via failure_detail_for_event in the runner/memory path). - ironclaw_reborn: map_provider_error puts the scrubbed provider reason into HostManagedModelError.detail; planned_driver carries HostUnavailable detail into AgentLoopDriverError::Failed; turn_runner/turn_run_executor carry it into the failure record. - ironclaw_reborn_composition: detail_for_turn_event sources from event.detail, feeding the FailureExplanationInput.detail already rendered in the explainer prompt. Construction-site churn: detail added to TurnLifecycleEvent / AgentLoopDriverError test fixtures and the event_projections pending-gate test support. Verified per crate (--lib): turns 355, loop_support 342, agent_loop 261, reborn 175, event_projections 26 — all green; composition --lib 1042 passed (1 pre-existing live_progress_stream failure, unrelated). turns integration contracts compile; clippy clean across the chain. Pre-existing base-branch breakage in loop_support thread_loop_support_contract (inline_messages) is unrelated to this change. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(turns): allow secret-scrubbed model-visible detail on Failed events The detail channel (TurnLifecycleEvent.detail) intentionally carries a secret-scrubbed description of the real failure cause to the model/explainer; update the guardrail so the spec matches the behavior. Only secret values are withheld; raw unscrubbed backend strings still stay behind host adapters. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): carry detail on AgentLoopDriverError::Failed in integration tests The tier-2b/3.1 detail field on AgentLoopDriverError::Failed broke construction and pattern sites in reborn's integration test targets (concurrent_workers, loop_driver_host) that the original --lib gate never compiled. Constructions get detail: None; the driver_host_error helper carries error.detail; exhaustive match patterns bind detail: _. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): update integration contracts for per-cause MCP tokens + detail channel Two integration test targets asserted pre-refinement model-visible strings that the --lib gate never compiled (test-through-the-caller gap): - mcp_adapter_contract: 5 assertions expected the flat "response_error"/ "request_denied" tokens; update them to the per-cause tokens the Tier 2a change now emits (mcp_invalid_protocol_version, mcp_jsonrpc_id_mismatch, mcp_invalid_session_id, mcp_http_status_500, mcp_denied_credential_source). - llm_gateway: the offline-provider test asserted the error Debug leaked NO provider detail at all. Tier 2b deliberately surfaces the secret-scrubbed non-secret cause on the detail channel. Rewrite (and rename) the test to the current policy: assert the non-secret reason ("connection refused", endpoint URL) reaches the model via `detail`, while the credential token (sk-provider-secret) is scrubbed from both `detail` and the full Debug. This strengthens the secret-scrubbing guard. mcp full suite green; llm_gateway full suite green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(host_runtime): missing first-party handler is InvalidInput, not Backend first_party_missing_handler_fails_closed_without_side_effect_handler asserted the dispatch failure kind was Backend, but #5389 deliberately reclassified an UndeclaredCapability/MethodMissing dispatch failure (a capability the model named that has no registered handler) to InvalidInput — a model-fixable, model-visible tool error that must not burn the retry budget on a call that can never resolve by retrying (see the From<DispatchFailureKind> mapping in production.rs). The test still fails closed (Failed outcome) and still carries "dispatch failed: UndeclaredCapability"; only the kind assertion is updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): LoopFailureKind fault matrix — executor layer + exhaustive category table Phase 1 of the per-error coverage harness (docs/plans/2026-07-03-loop-failure-matrix.md): - turns: all_failure_kinds category table now exhaustive (13/13 — adds the previously-missing CheckpointUnavailable + CompactionUnavailable) with a same-crate exhaustive-match guard so a new variant breaks compilation. - agent_loop: new table-driven executor failure matrix (executor/tests/failure_matrix.rs) driving every executor-reachable LoopFailureKind at its real origin via MockHost seams, asserting per row: P1 (reason_kind + sanitized category/safe_summary), P3 (no fabricated final assistant reply), and explanation_message_refs presence per the explainable set. New fail_transcript_with test knob on MockHost + DriverMockHost (test code only). - Four divergences found and documented (doc §5a), asserted as actual behavior, none silently fixed: Approval+SkipAndContinue completes (gate enforcement gap), NoProgressDetected missing its explanation attach, single Denied recovers-and-completes (no-borking working as designed), TranscriptWriteFailed/CheckpointRejected legacy-only enum origins. Validated (bounded): cargo test -p ironclaw_turns all_failure_kinds; cargo test -p ironclaw_agent_loop failure_matrix; check/clippy -D warnings/ fmt on ironclaw_agent_loop — all green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): LoopFailureKind fault matrix — binary/driver layer rows Phase 2 of the per-error coverage harness (docs/plans/2026-07-03-loop-failure-matrix.md): - planned_driver: resume with missing checkpoint payload asserts LoopFailureKind::CheckpointUnavailable + "checkpoint_unavailable"; in-flight model Cancelled (no cooperative cancel signal) asserts map_executor_error yields "interrupted_unexpectedly". - e2e: binary-level divergence lock — the same in-flight Cancelled run projects "driver_failed" at the runner boundary (category overwritten; doc §5a.5, candidate follow-up to preserve the driver-mapped category). - e2e: non-model P4 row — capability-stage invocation failure is retryable ("host_stage_unavailable_capability", checkpoint preserved, no fabricated reply) and retry_run resumes to completion, so P4 is no longer proven only through the model stage. New scripted capability-invocation-error mode in the test harness (test support only). Validated (bounded): targeted planned_driver tests + full reborn_failure_retry_resume_e2e (15 passed) + clippy -D warnings + fmt. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): reconcile fault matrix to the recoverability stack (#5389/#5390) This matrix PR sits on top of the recoverability stack (main←#4841←#5389←#5390←#5403), so it validates the FIXED + classified system rather than #4841's base behavior: - Add executor rows for #5389's model-fixable capability failures that are now RECOVERABLE (InvalidInput / InvalidOutput / PolicyDenied): where the base branch terminated the run, the stack now surfaces a model-visible tool error and the loop completes. Rows assert the recovered outcome. - Add FailureLane / RetryDisposition alignment (binary/e2e layer, since ironclaw_agent_loop can't depend on ironclaw_reborn_composition): each real failure path's (category, retryable) is asserted to map to the expected #5390 FailureLane bucket + RetryDisposition — proving the real paths feed the classifier correctly. Complements #5390's classifier unit tests (which test the functions directly) rather than duplicating them. - Doc: §6 relationship to #5390; §5a marks divergences RESOLVED by the stack (capability recoverability) vs still-open (Approval SkipAndContinue completes; NoProgressDetected lacks explanation; Transcript/Checkpoint are planned-executor host errors; interrupted_unexpectedly projected as driver_failed at the runner boundary). The cherry-picked base matrix assertions still passed unchanged on the stack — they assert actual behavior, so the additive fixes did not break them; this commit adds the stack-specific coverage on top. Validated (bounded, on the stack): cargo test ironclaw_turns all_failure_kinds; cargo test ironclaw_agent_loop failure_matrix; cargo test --test reborn_failure_retry_resume_e2e (19 passed); clippy -D warnings; fmt --check. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): address gemini failure matrix comments (#5613) * style: cargo fmt on batch-1 recoverable-error changes Reproduces the stack's skipped fmt commit (c19db45) against the restacked base. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tests): restack reconciliation — retry_run stub + superseded IssueCode import - webui_v2_router_smoke's MinimalWebuiServices gained the retry_run rejecting stub the trait now requires (base #4841 break: #5633's smoke fake predates #4841's retry_run addition; every other RebornServicesApi fake already has it). - drop CapabilityInputIssueCode from the ironclaw_turns re-export and ironclaw_loop_support import: the restacked base's CapabilityInputIssue carries DispatchInputIssueCode instead, and the stack's name was import-only on this branch. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(restack): outbound set-target routes through outbound_delivery_outcome; MCP tests use Option error info - set-target handler: replace the superseded #5445 NotFound special-case + outbound_delivery_host_error (deleted by the recoverability batch) with the outbound_delivery_outcome disposition the stack pins in unit tests; matches the list handler. - parse_mcp_response framing tests (base-side, written against the old `error: bool`): assert against the stack's richer `Option<JsonRpcErrorInfo>` — same intent, error presence still pinned. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(restack): turn_stream_auth fake event carries the new detail field Base-side projection test predates the stack's TurnLifecycleEvent.detail addition; None matches the auth-gate fixture's intent. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): align failure_category_demasked pin to the fidelity taxonomy TraceLlm exhaustion (gateway cannot serve the call) maps to ModelErrorClass::Unavailable -> "model_unavailable" under the batch-2 provider-error fidelity mapping; the scenario's "model_error" pin predated it. The scenario's intent — the de-masked TRUE category survives, never the "driver_protocol_violation" sentinel — is unchanged and still asserted exactly. Capability-surface direction (not a security boundary): both categories are classified, retriable lanes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(agent_loop): run fault-matrix rows on 16MiB threads The restacked executor's future (failure explanations + digests + the recovery re-entry) outgrows the 2MiB default test-thread stack in debug on the in-run recovery rows. Production loop threads run 8MiB stacks (ironclaw_reborn_cli serve runtime); mirror the repo's big-stack test-thread pattern (traces/tests.rs, process_port.rs) per row. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(host_runtime): egress contract pins the per-cause MCP denial token The SecretStoreLease-over-production-egress denial now surfaces mcp_denied_credential_source (McpRequestDeniedCause::DeniedCredentialSource) instead of the flat request_denied. Deny-before-transport is unchanged and still asserted (zero recorded requests); ironclaw_mcp's own adapter contract pins the same token. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(review): strip model-visible failure detail from public run-state + align feature-gated tests Security (IronLoop HIGH): RebornGetRunStateResponse forwarded SanitizedFailure.detail — free-form, model-visible backend cause text, scrubbed only for secret VALUES — straight to the browser. Add SanitizedFailure::public_projection() (keeps category, drops detail) and project the public WebUI shape through it. Regression tests: the strip helper (status.rs) and the caller (get_run_state contract test now programs a detail-bearing failure and asserts the DTO omits it). Inherent to the stack, not the reconciliation (original top-of-stack forwarded it raw too). Feature-gated test alignments (only run under --all-features/libsql, so the default-feature local sweep missed them; CI crate buckets caught them): - factory web-access: missing first-party handler is InvalidInput, not Backend (#5389 reclassify; capability still fails closed, only disposition changed). - outbound local_dev (x2): set-target routes through outbound_delivery_outcome (matches original top-of-stack), so the missing-target summary is the fixed "invalid outbound delivery request"; error_kind stays recoverable InvalidInput. - ironclaw_mcp parse_mcp_response_rejects_empty: per-cause tokens replace the flat "response_error" (mcp_parse_failed / mcp_no_payload). Also: cargo fmt (capability_port import reflow from the restack edit). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(review): drop untrusted MCP server error message from model-visible reason + align tool_call category Security (IronLoop HIGH): parse_json_rpc_error_info copied the untrusted MCP server JSON-RPC error.message into the McpClientError::Client reason (model-visible, 'stable sanitized reason') with only length-bounding — not redaction. MCP servers can echo request args, paths, provider diagnostics, or credential-shaped values. Remove message end-to-end (JsonRpcErrorInfo field, parse, cause variant, render); keep only the standardized protocol code=<n>, which is the safe diagnostic. No redaction util is reachable from this leaf crate, and the stable-reason surface should carry stable tokens, not free text. Regression: the former 'reason carries message' test now asserts the message does NOT leak. Feature-gated integration test (ran only under --features libsql): tests/integration/tool_call.rs disabled-spawn-subagent-called-anyway now asserts 'model_unavailable' (InvalidOutput -> Unavailable fidelity category), not the stale 'model_error'. The security property (disabled capability never dispatched) is unchanged and still asserted. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): cancel-path provider error is model_context_overflow, not model_error fail_model() -> ErrLlm -> LlmError::ContextLengthExceeded, which the batch-2 provider fidelity mapping now categorizes as the accurate model_context_overflow (was the generic model_error). Both cancel tests still pin the load-bearing behavior: reaches Failed after bounded context-shrink recovery (no retry-forever), and the per-thread busy lock releases on Failed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(merge): implement retry_turn on TurnCoordinator doubles added by main Merging origin/main brought two new TurnCoordinator test doubles (UnusedTurnCoordinator in src/runtime.rs, SpyTurnCoordinator in tests/runtime.rs) that predate this stack's retry_turn addition to the TurnCoordinator trait. Add the impls (unimplemented!/delegate, matching each double's existing method style) so the merged tree compiles. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(deny): ignore RUSTSEC-2026-0204 (crossbeam Debug-fmt invalid deref) Newly published advisory (after this branch and main), transitive via crossbeam-epoch. The affected path is the `fmt::Pointer`/`Debug` impl for `Atomic`/`Shared` when the pointer is already invalid — a formatting path we do not exercise. Ignore with justification per the existing advisories convention; remove when the fixed crossbeam-utils release propagates. Verified `cargo deny check advisories` = ok locally. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(reborn): preserve scrubbed failure detail into TurnRunExecutorError (IronLoop) The driver-failed Err path in execute_claimed_run converted the computed SanitizedFailure back to TurnRunExecutorError::new(category), dropping the scrubbed model-visible detail. The scheduler records error.failure(), so production driver failures persisted only the category and TurnLifecycleEvent.detail stayed empty — the failure explainer got the fallback summary instead of the real provider/model cause. Add TurnRunExecutorError::from_failure(SanitizedFailure) (the struct already holds a full SanitizedFailure) and use it at the call site so detail survives across the host-runtime boundary. Caller regression test: driver Failed{detail: Some(..)} -> execute_claimed_run -> err.failure().detail() is preserved. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(reborn): JSON-frame untrusted failure detail in the explainer prompt (IronLoop) detail is untrusted provider/tool/runtime error text (e.g. MCP server or provider bodies). sanitize_model_visible_text redacts credential tokens but keeps newlines/instructions, so appending it raw let a crafted error inject extra prompt fields or directives (a fake fallback_summary:, an 'ignore previous instructions') into the failure explainer — whose output becomes the public failure_summary. That is a prompt-injection path into user-visible messaging, widened by the detail-preservation fix. Frame detail as data: JSON-string-escape it so newlines/quotes are escaped and it stays a single quoted 'detail: "..."' value. failure_category and fallback_summary are host-authored (category-derived) and unchanged. Regression test: a detail embedding newline+fallback_summary+directive is neutralized (exactly one real fallback_summary line, no directive line, detail present as an escaped quoted value). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Stacked on #5403 (top of the recoverability stack: main ← #4841 ← #5389 ← #5390 ← #5403). A pure test + docs PR — no production changes — giving the "no run-borking errors" work a per-error-type coverage net and a documented behavior comparison vs
main.What & why
We had no test proving recovery/retry/explanation for each
LoopFailureKind. This adds one, built from a 4-way parallel audit (taxonomy, coverage, main-delta, injection seams) synthesized indocs/plans/2026-07-03-loop-failure-matrix.md.1. Exhaustive category table (
ironclaw_turns)all_failure_kinds_produce_stable_sanitized_category_stringsnow covers 13/13 variants (adds the previously-missingCheckpointUnavailable+CompactionUnavailable) with a same-crate exhaustive-match guard, so a new variant fails compilation until it's added.2. Table-driven executor matrix (
ironclaw_agent_loop, newexecutor/tests/failure_matrix.rs)Drives every executor-reachable variant at its real origin via
MockHostseams (+ a newfail_transcript_withknob), asserting per row: reason_kind + sanitized category (P1), no fabricated final reply (P3), andexplanation_message_refspresence per the explainable set. Includes recovery rows for #5389's now-model-fixable capability failures (InvalidInput / InvalidOutput / PolicyDenied) that complete via a model-visible tool error instead of terminating.3. Binary/driver rows + FailureLane alignment (
tests/reborn_failure_retry_resume_e2e.rs)CheckpointUnavailableat the resume-decode path; in-flightCancelled→interrupted_unexpectedlymapping.retry_runresumes to completion — so retry is no longer proven only through the model stage.FailureLane/RetryDispositionalignment: each real failure path's(category, retryable)is asserted to map to the expected bucket — proving the real paths feed feat(reborn): FailureLane classifier + two-bucket enforcement test #5390's classifier correctly. Complements feat(reborn): FailureLane classifier + two-bucket enforcement test #5390's classifier unit tests (which test the functions directly); does not duplicate them.vs
main(doc §2)mainshows only a category + a lazy generic blurb and every failure is terminal for the user. The stack adds run-context explanations persisted to the transcript, theretry_runpath, and #5390's recoverable/run-ending classification. The matrix asserts these as CI facts (e.g. "a single policy denial recovers and completes").Findings surfaced (doc §5a) — documented, not silently fixed
Resolved by the stack: model-fixable capability failures now recover. Still open (candidate follow-ups, flagged for owners): Approval-gate
SkipAndContinuecompletes instead ofDriverBug;NoProgressDetectednever attaches its explanation (must not disturb the PinchBench-load-bearing nudge);TranscriptWriteFailed/CheckpointRejectedare legacy-driver-only enum origins;interrupted_unexpectedlyis overwritten todriver_failedat the runner boundary.Validation (bounded, on the stack)
cargo test ironclaw_turns all_failure_kinds·cargo test ironclaw_agent_loop failure_matrix·cargo test --test reborn_failure_retry_resume_e2e(19 passed) ·clippy -D warnings·fmt --check— all green.🤖 Generated with Claude Code