Skip to content

fix(reborn): complete model error recovery contract - #6845

Merged
serrrfirat merged 10 commits into
mainfrom
codex/error-recoverability-ws2
Jul 30, 2026
Merged

serrrfirat merged 10 commits into
mainfrom
codex/error-recoverability-ws2

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

  • Split model-host capacity failures into spend-budget exhaustion, input/context overflow, and output truncation so each follows its own recovery contract.
  • Route FinishReason::Length through a model-visible continue-or-condense turn without consuming context-shrink attempts, and prevent truncated textual tool calls from being dispatched.
  • Add durable, bounded observations for recoverable availability, stale-request, and budget-accounting failures while preserving sanctioned terminal boundaries.
  • Carry the new typed identities through runner categories, summaries, retry dispositions, host adapters, checkpoint state, and caller-path integration tests.

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Fixes #6700

Related #6284

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings
    • Targeted equivalent passed for all touched production crates: cargo clippy -p ironclaw_turns -p ironclaw_agent_loop -p ironclaw_loop_host -p ironclaw_runner --tests -- -D warnings.
  • cargo build
    • Not run separately: the targeted tests and clippy compiled the touched dependency graph.
  • Relevant tests pass: agent loop (441), turns host contract (83), loop host (399), runner unit tests (358 after rebase), real gateway caller tests (80), and model-recovery integration tests (17).
  • cargo test --features integration if database-backed or integration behavior changed
    • The monolithic integration suite was not selected by the playbook; the changed Reborn target passed with cargo test -p ironclaw_reborn_integration_tests --test reborn_integration_model_recovery.
  • Manual testing: Not applicable; behavior is deterministic and covered through the real gateway caller plus durable integration seams.
  • If a coding agent was used and supports it, review-pr or pr-shepherd --fix was run before requesting review
    • Not run; focused diff/forbidden-pattern review, rustfmt, clippy, and selected test tiers were completed locally.

Test Strategy

User behavior: A provider-truncated response now gets one actionable request to continue or condense, without losing context-shrink budget or dispatching an incomplete textual tool call. Recoverable provider availability/stale errors and a budget-accounting failure receive the allowed model-visible observation/final turn.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: exhaustive error-kind projection matrices; durable observation/warning round trips; recovery-budget accounting; failure categories, lanes, summaries, and retry dispositions; truncated textual-tool-call suppression.
  • Reborn integration: scripted provider returns FinishReason::Length through the real gateway caller and the run survives to a durable final response without shrinking input context.
  • Recorded fixture: Not applicable: no provider wire protocol or HTTP fixture changed.
  • Browser E2E: Not applicable: no browser or frontend behavior changed.
  • Backend or runtime: Not applicable: no external backend/runtime lane changed; in-memory host and runner seams exercise the contract.
  • Live canary: Not applicable: deterministic provider behavior is covered without external credentials or spend.

What the tests prove: Typed failure identity survives every producer/projection seam; output truncation cannot consume ShrinkContext; partial tool-like output cannot escape as a side effect; observations and the budget-accounting warning are bounded and survive checkpoint serialization; and the provider caller can recover to a durable completed run.

Commands run:

CARGO_PROFILE_TEST_DEBUG=0 CARGO_PROFILE_DEV_DEBUG=0 cargo test -p ironclaw_agent_loop --lib --quiet
CARGO_PROFILE_TEST_DEBUG=0 CARGO_PROFILE_DEV_DEBUG=0 cargo test -p ironclaw_turns --test agent_loop_host_contract --quiet
CARGO_PROFILE_TEST_DEBUG=0 CARGO_PROFILE_DEV_DEBUG=0 cargo test -p ironclaw_loop_host --lib --quiet
CARGO_PROFILE_TEST_DEBUG=0 CARGO_PROFILE_DEV_DEBUG=0 cargo test -p ironclaw_runner --lib --quiet
CARGO_PROFILE_TEST_DEBUG=0 CARGO_PROFILE_DEV_DEBUG=0 cargo test -p ironclaw_runner --test llm_gateway --quiet
CARGO_PROFILE_TEST_DEBUG=0 CARGO_PROFILE_DEV_DEBUG=0 cargo test -p ironclaw_reborn_integration_tests --test reborn_integration_model_recovery
CARGO_PROFILE_TEST_DEBUG=0 CARGO_PROFILE_DEV_DEBUG=0 cargo clippy -p ironclaw_turns -p ironclaw_agent_loop -p ironclaw_loop_host -p ironclaw_runner --tests -- -D warnings
cargo fmt --all -- --check
git diff --check origin/main...HEAD

Security Impact

None. No permissions, network boundaries, secrets, file access, runtime policy, or sandbox behavior changed. Recovery instructions are host-authored typed observations; provider partial text is discarded on truncation and is not persisted or dispatched.

Reborn Trust-Boundary Checklist

  • Public policy/evidence/trust-bearing types: the changed public error variants are constructed by the owning host/gateway producers and projected without re-deriving identity from strings.
  • Untrusted content enters prompts only through an envelope/escaping primitive. No new untrusted prompt content was introduced; the new observations are fixed host-authored control text.
  • Hashes declare purpose; trust/binding/authenticity uses SHA-256/BLAKE3 or separate authenticity check. N/A: no new hash purpose or authenticity mechanism; existing recovery-family digests were extended with typed state.
  • New/changed status, exit, policy, runtime, or error variants: downstream match sites audited. Command/output: rg -n "AgentLoopHostErrorKind::|HostManagedModelErrorKind::" crates tests; exhaustive matrices, touched-crate clippy, and caller tests passed.
  • Security/durability serde(default) fields fail closed or have migration tests. No permissive default was added; new serialized enum variants and checkpoint-carried warning/observation state have round-trip tests.
  • Queues/maps/buffers/counters have bounds and overflow-safe arithmetic. Observation and warning budgets are one-shot/bounded and attempt arithmetic remains saturating.
  • Driver/operator-visible errors have stable class semantics (Transient, Permanent, Misconfigured, PolicyDenied or equivalent). Spend, context, truncation, and accounting categories remain distinct through summaries and retry disposition.
  • Sandbox/native/host names accurately describe trust boundary. N/A: no sandbox, native, or host trust-boundary names changed.

Database Impact

None. No migration or database schema changed.

Blast Radius

The change touches Reborn agent-loop recovery, loop-host error producers, turns host/model contracts, runner failure projection, and the scripted integration provider. Regressions could affect model retry/finalization behavior, failure telemetry categories, or deserialization of an in-flight checkpoint containing one of the newly serialized variants.

Rollback Plan

Revert this commit. Existing pre-change checkpoints remain readable, but do not roll an affected deployment back to an older binary while runs with the new serialized warning/observation variants are in flight; drain or restart those runs first.

Review Follow-Through

Reviewer judgment is requested on the retracted CheckpointRejected warning box in #6284: the live rejection occurs while writing the pre-model checkpoint, so requesting another model turn would cross the durable-event boundary that failed. This PR therefore preserves checkpoint rejection as terminal and adds the bounded final warning only for BudgetAccountingFailed.

Unauthorized and TranscriptWriteFailed also remain sanctioned terminal/unreachable model-observation classes. Additionally, rig-core does not expose streaming finish reasons for every provider; this fixes every caller path where FinishReason::Length is currently reachable.


Review track: C (runtime and durable checkpoint contract)

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6845 July 29, 2026 09:54 Destroyed
@coderabbitai

coderabbitai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 1744a795-4366-42b7-a297-2c30b2e8559e

📥 Commits

Reviewing files that changed from the base of the PR and between 08b487a and ad80dc9.

📒 Files selected for processing (1)
  • scripts/reborn-e2e-rust.sh

📝 Walkthrough

Summary by CodeRabbit

  • New Features
    • Added dedicated handling for spend-budget exhaustion, context overflow, and output truncation, including observation-assisted “observe then abort” recovery behavior.
    • Preserved provider-reported token usage across model failures for more accurate budget accounting.
    • Introduced new typed terminal/recovery observations and failure categories for clearer run outcomes.
  • Bug Fixes
    • Improved fallback-route and checkpoint-evidence validation; checkpoint rejections now produce deterministic host-authored explanations without unintended model/capability execution.
  • Documentation
    • Updated contracts for checkpoint rejection before trustworthy exit and deterministic turn-runner behavior.

Walkthrough

This PR separates spend-budget, context-overflow, and output-truncation errors; adds typed observation-based recovery and provider-usage reconciliation; validates fallback-route evidence; and introduces non-retryable typed checkpoint rejection with bounded host-authored failure details.

Changes

Typed model error recovery

Layer / File(s) Summary
Error and observation contracts
crates/ironclaw_agent_loop/src/state/*, crates/ironclaw_turns/src/run_profile/{host,error.rs,host/progress.rs,model.rs}, crates/ironclaw_agent_loop/src/families/*
Adds typed model observations, output-truncation recovery identity, budget-accounting warnings, precise host error kinds, provider usage fields, and updated replay fingerprints.
Host and gateway mappings
crates/ironclaw_loop_host/src/*, crates/ironclaw_runner/src/{model_gateway.rs,model_gateway_error_mapping.rs,model_failure_mapping.rs,text_loop_driver.rs,failure_*.rs}
Maps spend, context, and truncation outcomes distinctly, preserves provider usage on failures, and validates optional fallback-route evidence.
Observation-assisted recovery and execution
crates/ironclaw_agent_loop/src/{executor/*,strategies/recovery.rs}, tests/integration/*, tests/reborn_failure_retry_resume_e2e.rs
Adds retry-then-observe behavior for availability, stale requests, and truncation; adds accounting-failure retry handling; and validates fallback advancement and attempt limits.
Checkpoint rejection and terminal projection
crates/ironclaw_agent_loop/src/{executor.rs,executor/checkpoint.rs,executor/tests/*}, crates/ironclaw_runner/src/{planned_driver.rs,failure_summary.rs}, crates/ironclaw_product/src/projection/*, docs/reborn/contracts/*, scripts/*
Adds typed checkpoint rejection, safe-summary validation and fallback, host-authored terminal explanations, no-explainer projection, and contract-test coverage.

Estimated code review effort: 4 (Complex) | ~75 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Provider
  participant ModelStage
  participant RecoveryStrategy
  Provider->>ModelStage: FinishReason::Length
  ModelStage->>RecoveryStrategy: OutputTruncated
  RecoveryStrategy->>ModelStage: typed continue-or-condense observation
  ModelStage->>Provider: recovery request
  Provider->>ModelStage: finalized reply
Loading
sequenceDiagram
  participant Executor
  participant CheckpointHost
  participant PlannedDriver
  participant ProductProjection
  Executor->>CheckpointHost: write checkpoint
  CheckpointHost-->>Executor: CheckpointRejected
  Executor->>PlannedDriver: typed rejection
  PlannedDriver->>ProductProjection: host-authored failure detail
  ProductProjection->>ProductProjection: validate detail without explainer
Loading

Possibly related issues

Possibly related PRs

Suggested reviewers: think-in-universe

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed Conventional Commits-style title matches the PR's main model-recovery and error-splitting work.
Description check ✅ Passed Required sections are present and mostly filled, including summary, linked issue, validation, test strategy, and rollout details.
Linked Issues check ✅ Passed Length now maps to OutputTruncated, truncation uses model-visible recovery, the real-gateway regression is added, and spend-budget/context-overflow are split.
Out of Scope Changes check ✅ Passed The diff stays focused on recovery-contract, error-projection, and regression coverage; no clearly unrelated code paths stand out.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 29, 2026
@ironloopai

ironloopai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #6845

🟢 Completed · Review submitted

Submitted review →

Reviewed the complete trusted comparison a0e91d1..2cc0ae7 across all 31 changed files. The typed error projections, bounded recovery state, truncation handling, budget-accounting path, runner categories, checkpoint serialization, and integration support are internally consistent. No concrete actionable findings were identified.

Automatic · PR opened · attempt 1 of 3 · completed in 2m 17s

Run details
  • Repository: nearai/ironclaw
  • Base: main at a0e91d1
  • Head: codex/error-recoverability-ws2 at 2cc0ae7
  • Created: Jul 29, 2026, 9:59 AM UTC
  • Updated: Jul 29, 2026, 10:01 AM UTC
  • Run: 7d2f8c37-4be1-4626-a8f5-86ba97b155ca
  • Latest attempt: 1 · Completed · 4735dae1-7b05-4a00-80a5-55b00b6135e0

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Review complete · PR #6845

✅ No actionable findings

Reviewed the complete trusted comparison a0e91d1..2cc0ae7 across all 31 changed files. The typed error projections, bounded recovery state, truncation handling, budget-accounting path, runner categories, checkpoint serialization, and integration support are internally consistent. No concrete actionable findings were identified.

Validation and technical details
  • Inspected every changed production and test area plus surrounding model gateway, executor, checkpoint, prompt, budget-accountant, failure-mapping, and runner call paths.
  • Verified the trusted base and head refs and used their complete comparison rather than the checkout HEAD.
  • git diff --check refs/ironloop/base..refs/ironloop/head passed.
  • Reviewed exhaustive error-kind/category matrices and tests covering output truncation, stale/availability recovery, spend-budget exhaustion, accounting warnings, checkpoint round trips, and suppression of truncated textual tool calls.
  • Could not execute Cargo tests because the review environment does not provide the cargo executable; this is an environment limitation, not a human-decision blocker.
  • Base: main
  • Head: codex/error-recoverability-ws2 at 2cc0ae7
  • Run: 7d2f8c37-4be1-4626-a8f5-86ba97b155ca

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_runner/src/turn_runner.rs (1)

51-58: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Preserve terminal context/truncation categories.

model_context_overflow and model_output_truncated are registered in failure_lane.rs and have dedicated summaries/retry dispositions, but this allowlist omits both. When recovery finally returns either category, this branch rewrites it to driver_failed, losing the typed durable/public failure contract. Preserve both categories here and add regression cases alongside the spend-budget test.

As per coding guidelines, preserve observable failure kinds and durable state exactly.

Also applies to: 175-188

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_runner/src/turn_runner.rs` around lines 51 - 58, Extend the
allowlist in the base-category selection logic of turn_runner to include
model_context_overflow and model_output_truncated. Preserve these categories
unchanged when recovery returns them, and add regression coverage alongside the
existing spend-budget test verifying their durable/public failure kinds remain
intact.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_loop_host/src/budget_accountant.rs`:
- Around line 1040-1046: Update the approval-threshold test setup to set
max_output_tokens to 27, producing a $9.10 total below the $10 hard cap, and
change the assertion around err.kind to require
AgentLoopHostErrorKind::BudgetApprovalRequired exclusively. Add or update a
caller-level test at the meaningful contract seam to verify this approval-path
behavior.

---

Outside diff comments:
In `@crates/ironclaw_runner/src/turn_runner.rs`:
- Around line 51-58: Extend the allowlist in the base-category selection logic
of turn_runner to include model_context_overflow and model_output_truncated.
Preserve these categories unchanged when recovery returns them, and add
regression coverage alongside the existing spend-budget test verifying their
durable/public failure kinds remain intact.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: bce3aabe-d03d-45e8-9cf7-57b29679fd60

📥 Commits

Reviewing files that changed from the base of the PR and between a0e91d1 and 2cc0ae7.

📒 Files selected for processing (31)
  • crates/ironclaw_agent_loop/src/executor/mapping.rs
  • crates/ironclaw_agent_loop/src/executor/model.rs
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_agent_loop/src/executor/tests/failure_matrix.rs
  • crates/ironclaw_agent_loop/src/families/mod.rs
  • crates/ironclaw_agent_loop/src/families/subagent.rs
  • crates/ironclaw_agent_loop/src/state/model_recovery.rs
  • crates/ironclaw_agent_loop/src/state/slots.rs
  • crates/ironclaw_agent_loop/src/state/terminal_warning.rs
  • crates/ironclaw_agent_loop/src/strategies/recovery.rs
  • crates/ironclaw_loop_host/src/budget_accountant.rs
  • crates/ironclaw_loop_host/src/identity_context.rs
  • crates/ironclaw_loop_host/src/lib.rs
  • crates/ironclaw_loop_host/src/skill_context.rs
  • crates/ironclaw_loop_host/src/system_inference.rs
  • crates/ironclaw_runner/src/failure_categories.rs
  • crates/ironclaw_runner/src/failure_lane.rs
  • crates/ironclaw_runner/src/failure_summary.rs
  • crates/ironclaw_runner/src/model_failure_mapping.rs
  • crates/ironclaw_runner/src/model_gateway.rs
  • crates/ironclaw_runner/src/retry_disposition.rs
  • crates/ironclaw_runner/src/text_loop_driver.rs
  • crates/ironclaw_runner/src/turn_runner.rs
  • crates/ironclaw_runner/tests/llm_gateway.rs
  • crates/ironclaw_turns/src/run_profile/host/error.rs
  • crates/ironclaw_turns/src/run_profile/model.rs
  • crates/ironclaw_turns/tests/agent_loop_host_contract.rs
  • tests/integration/model_recovery.rs
  • tests/integration/support/assertions.rs
  • tests/integration/support/builder.rs
  • tests/integration/support/scripted_provider.rs

Comment thread crates/ironclaw_loop_host/src/budget_accountant.rs Outdated
@railway-app

railway-app Bot commented Jul 29, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6845 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 30, 2026 at 8:53 am

…ility-ws2

# Conflicts:
#	crates/ironclaw_agent_loop/src/executor/tests.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6845 July 29, 2026 16:51 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6845 July 29, 2026 16:55 Destroyed
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Addressed the remaining CodeRabbit findings in 60e742bd6:

  • Made the budget-approval threshold test deterministic at $9.10 and require BudgetApprovalRequired; the executor caller-path test verifies the resource gate and BeforeBlock checkpoint.
  • Preserved model_context_overflow and model_output_truncated through sanitized_driver_failure, with a regression test covering both durable/public categories.

Focused validation: runner category regression passed; formatting and diff checks passed. The accountant test compilation was attempted locally but the shared macOS build volume exhausted disk; the pushed CI run is the clean-runner verification.

@github-actions

github-actions Bot commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.55% (315441 / 368740 lines)
  floor:    85.11% (tolerance 0.5pp -> effective floor 84.61%)
  denominator: 368740 lines now vs 360961 at floor capture (+7779 lines, +2.16%) — not a material change

RATCHET PASS: ironclaw_runner
  observed: 86.87% (14669 / 16887 lines)
  floor:    83.58% (tolerance 0.5pp -> effective floor 83.08%)
  floor_covered_lines: 13083 (tolerance 20 lines -> effective floor 13063)
  denominator: 16887 lines now vs 15653 at floor capture (+1234 lines, +7.88%) — material change (>5%)

RATCHET PASS: ironclaw_processes
  observed: 88.05% (5838 / 6630 lines)
  floor:    82.29% (tolerance 0.5pp -> effective floor 81.79%)
  floor_covered_lines: 5140 (tolerance 20 lines -> effective floor 5120)
  denominator: 6630 lines now vs 6246 at floor capture (+384 lines, +6.15%) — material change (>5%)

RATCHET PASS: ironclaw_turns
  observed: 86.21% (9453 / 10965 lines)
  floor:    85.36% (tolerance 0.5pp -> effective floor 84.86%)
  floor_covered_lines: 9181 (tolerance 20 lines -> effective floor 9161)
  denominator: 10965 lines now vs 10755 at floor capture (+210 lines, +1.95%) — not a material change

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.55% — 315441 / 368740 lines

Per-crate breakdown (60 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_process_sandbox 34.38% 120 / 349
ironclaw_host_ingress 42.5% 17 / 40
ironclaw_event_projections 43.51% 684 / 1572
ironclaw_libsql_runtime 59.91% 127 / 212
ironclaw_observability 61.54% 16 / 26
ironclaw_authorization 63.02% 610 / 968
ironclaw_memory 64.41% 959 / 1489
ironclaw_telegram_v2_adapter 69.69% 731 / 1049
ironclaw_trust 73.21% 664 / 907
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_extractors 74.72% 538 / 720
ironclaw_capabilities 74.95% 2882 / 3845
ironclaw_projects 76.48% 400 / 523
ironclaw_mcp 76.6% 779 / 1017
ironclaw_filesystem 77.62% 5995 / 7724
ironclaw_reborn_cli 78.37% 10795 / 13774
ironclaw_wasm 79.72% 735 / 922
ironclaw_llm 80.82% 23319 / 28854
ironclaw_memory_native 81.02% 3299 / 4072
ironclaw_auth 81.88% 6679 / 8157
ironclaw_first_party_extensions 82.38% 6682 / 8111
ironclaw_host_api 83.68% 9728 / 11625
ironclaw_events 83.8% 1536 / 1833
ironclaw_reborn_identity 83.8% 450 / 537
ironclaw_operator 84.41% 5561 / 6588
ironclaw_secrets 84.53% 2797 / 3309
ironclaw_network 85.09% 959 / 1127
ironclaw_reborn_config 85.23% 2101 / 2465
ironclaw_skills 85.27% 4493 / 5269
ironclaw_telegram_extension 85.4% 1299 / 1521
ironclaw_extension_host 85.52% 19926 / 23299
ironclaw_approvals 86.08% 1818 / 2112
ironclaw_triggers 86.11% 2803 / 3255
ironclaw_turns 86.21% 9453 / 10965
ironclaw_reborn_composition 86.23% 22037 / 25556
ironclaw_webui 86.37% 11171 / 12934
ironclaw_hooks 86.73% 9950 / 11473
ironclaw_runner 86.87% 14669 / 16887
ironclaw_common 86.99% 1772 / 2037
ironclaw_reborn_event_store 87.19% 1327 / 1522
ironclaw_product 87.44% 21340 / 24406
ironclaw_extensions 87.45% 4927 / 5634
ironclaw_reborn_openai_compat 88% 3724 / 4232
ironclaw_processes 88.05% 5838 / 6630
ironclaw_reborn_traces 88.13% 11987 / 13601
ironclaw_threads 88.57% 4943 / 5581
ironclaw_host_runtime 89.09% 20539 / 23054
ironclaw_slack_extension 89.26% 2027 / 2271
ironclaw_conversations 90.03% 3171 / 3522
ironclaw_resources 90.84% 4474 / 4925
ironclaw_loop_host 90.93% 17735 / 19505
ironclaw_event_streams 91.24% 1063 / 1165
ironclaw_attachments 93.06% 630 / 677
ironclaw_outbound 93.91% 4101 / 4367
ironclaw_agent_loop 94.42% 10422 / 11038
ironclaw_safety 95.22% 3941 / 4139
ironclaw_first_party_extension_ports 95.71% 3837 / 4009
ironclaw_runtime_policy 96.56% 814 / 843

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

* fix(reborn): explain rejected checkpoints durably

* Format reconciled checkpoint projection imports
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6845 July 29, 2026 19:16 Destroyed
@github-actions github-actions Bot added the scope: docs Documentation label Jul 29, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_product/src/projection/tests/failure_explanation.rs`:
- Around line 218-221: Remove the duplicate
MODEL_SPEND_BUDGET_EXHAUSTED_CATEGORY entry from the expected
failure-explanation table, retaining the earlier row and its summary so the
table contains each category exactly once.

In `@crates/ironclaw_runner/src/failure_summary.rs`:
- Line 44: Update the LoopSafeSummary::new call in the failure-summary
construction to avoid .ok()?; explicitly handle its validation error while
preserving the existing None fallback and retaining diagnostic context for the
rejected persisted detail. Ensure production Rust does not silently discard this
error.

In `@crates/ironclaw_turns/src/run_profile/host/refs.rs`:
- Around line 199-203: Update CheckpointRejected::checkpoint_rejected to use
cause-neutral fallback wording rather than implying validation failure, while
preserving its role for host-rejected checkpoints without a valid bounded cause.
Add a caller-level regression test through checkpoint_host_error covering an
invalid producer summary and asserting the neutral fallback message.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: abf3619b-81d1-4c58-ae4a-f8bb4a8b55de

📥 Commits

Reviewing files that changed from the base of the PR and between 60e742b and 1ac7545.

📒 Files selected for processing (17)
  • crates/ironclaw_agent_loop/src/executor.rs
  • crates/ironclaw_agent_loop/src/executor/checkpoint.rs
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_agent_loop/src/executor/tests/cancellation.rs
  • crates/ironclaw_agent_loop/src/executor/tests/failure_matrix.rs
  • crates/ironclaw_product/src/projection/tests/failure_explanation.rs
  • crates/ironclaw_product/src/projection/turn_events.rs
  • crates/ironclaw_runner/src/failure_categories.rs
  • crates/ironclaw_runner/src/failure_summary.rs
  • crates/ironclaw_runner/src/planned_driver.rs
  • crates/ironclaw_runner/src/turn_runner.rs
  • crates/ironclaw_runner/tests/loop_driver_host.rs
  • crates/ironclaw_turns/src/run_profile/host/refs.rs
  • crates/ironclaw_turns/src/turn_state_row_store/turn_state_engine/transitions.rs
  • docs/reborn/contracts/loop-exit.md
  • docs/reborn/contracts/turn-runner.md
  • scripts/reborn-e2e-rust.sh

Comment thread crates/ironclaw_product/src/projection/tests/failure_explanation.rs Outdated
Comment thread crates/ironclaw_runner/src/failure_summary.rs Outdated
Comment thread crates/ironclaw_turns/src/run_profile/host/refs.rs
…ility-ws2

# Conflicts:
#	crates/ironclaw_agent_loop/src/strategies/recovery.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6845 July 29, 2026 20:44 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (3)
crates/ironclaw_runner/src/model_gateway.rs (1)

2143-2146: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Account for truncated completion usage before recovery.

FinishReason::Length drops usage when it becomes OutputTruncated. The post-call accountant then sees a failure and releases the reservation, so tokens consumed by the truncated attempt are not charged before the continuation call. Preserve/reconcile usage on this failure path in both text and tool response conversion, with a caller-path regression test.

As per coding guidelines, “Test through the caller” applies when a transform gates model dispatch and accounting side effects.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_runner/src/model_gateway.rs` around lines 2143 - 2146, Update
both text and tool response conversion paths handling FinishReason::Length so
they preserve and reconcile the attempt’s usage before returning
OutputTruncated, allowing post-call accounting to charge consumed tokens before
recovery. Add a regression test through the caller path that exercises the
truncation, continuation, and accounting side effects rather than testing the
conversion helper alone.

Source: Coding guidelines

crates/ironclaw_loop_host/src/lib.rs (1)

1820-1822: 🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

Require and validate fallback-route evidence.

effective_fallback_index is declared authoritative, but a missing serialized field becomes 0; this passes the existing default-route equality check. System inference also consumes gateway responses without checking the reported route at all. Fail closed instead of accepting absent or mismatched runtime-selection evidence.

  • crates/ironclaw_loop_host/src/lib.rs#L1820-L1822: remove the zero default for authoritative response evidence, or represent absence explicitly and reject it before accepting a response.
  • crates/ironclaw_loop_host/src/system_inference.rs#L160-L166: retain the requested index and reject a response whose effective index differs before reading its output.

As per coding guidelines, “Fail closed for ... runtime selection” is required.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_loop_host/src/lib.rs` around lines 1820 - 1822, Require
explicit fallback-route evidence: in crates/ironclaw_loop_host/src/lib.rs lines
1820-1822, remove the default zero from effective_fallback_index or make absence
explicit and reject it before accepting the response; in
crates/ironclaw_loop_host/src/system_inference.rs lines 160-166, retain the
requested index and reject any response whose effective index differs before
reading its output.

Sources: Coding guidelines, Path instructions

crates/ironclaw_runner/src/planned_driver.rs (1)

350-358: 🔒 Security & Privacy | 🟠 Major | 🏗️ Heavy lift

Do not forward raw prompt-stage detail to the driver boundary.

Line 351 can expose AgentLoopHostError.detail; the current scrubber preserves path-like values, while this new prompt-stage path returns that detail in AgentLoopDriverError::Failed. Keep the actionable LoopSafeSummary, but drop raw detail or map it to a stable redacted explanation.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_runner/src/planned_driver.rs` around lines 350 - 358, Update
the prompt-stage failure handling in permanent_prompt_stage_failure_category so
AgentLoopDriverError::Failed never receives raw AgentLoopHostError.detail.
Preserve the actionable LoopSafeSummary via the existing safe_summary path, and
either omit detail or replace it with a stable redacted explanation before
constructing the Failed result.

Sources: Coding guidelines, Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_agent_loop/src/strategies/recovery.rs`:
- Around line 388-405: Preserve the host-selected fallback index through
recovery: in crates/ironclaw_agent_loop/src/strategies/recovery.rs lines
388-405, pass next_fallback_index as payload in RetryAlteration instead of using
payload-free AdvanceFallback; in
crates/ironclaw_agent_loop/src/executor/model.rs lines 462-473, validate the
requested index advances monotonically and assign that exact index to
state.model_state.fallback_index rather than incrementing the local index.

In `@tests/integration/model_recovery.rs`:
- Around line 187-228: Extend
output_truncation_recovers_without_shrinking_input_context, or add a
capability-backed regression alongside it, using recognizable partial textual
tool-call content instead of only the default echo response and empty
tool_calls. Configure a capability that would visibly record invocation or
egress, then assert it is never invoked before the recovery turn while
preserving the existing successful recovery assertions.

---

Outside diff comments:
In `@crates/ironclaw_loop_host/src/lib.rs`:
- Around line 1820-1822: Require explicit fallback-route evidence: in
crates/ironclaw_loop_host/src/lib.rs lines 1820-1822, remove the default zero
from effective_fallback_index or make absence explicit and reject it before
accepting the response; in crates/ironclaw_loop_host/src/system_inference.rs
lines 160-166, retain the requested index and reject any response whose
effective index differs before reading its output.

In `@crates/ironclaw_runner/src/model_gateway.rs`:
- Around line 2143-2146: Update both text and tool response conversion paths
handling FinishReason::Length so they preserve and reconcile the attempt’s usage
before returning OutputTruncated, allowing post-call accounting to charge
consumed tokens before recovery. Add a regression test through the caller path
that exercises the truncation, continuation, and accounting side effects rather
than testing the conversion helper alone.

In `@crates/ironclaw_runner/src/planned_driver.rs`:
- Around line 350-358: Update the prompt-stage failure handling in
permanent_prompt_stage_failure_category so AgentLoopDriverError::Failed never
receives raw AgentLoopHostError.detail. Preserve the actionable LoopSafeSummary
via the existing safe_summary path, and either omit detail or replace it with a
stable redacted explanation before constructing the Failed result.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 2db1b57d-d88d-4e41-83a1-fecf3b45cddb

📥 Commits

Reviewing files that changed from the base of the PR and between 1ac7545 and 4931832.

📒 Files selected for processing (25)
  • crates/ironclaw_agent_loop/src/executor.rs
  • crates/ironclaw_agent_loop/src/executor/mapping.rs
  • crates/ironclaw_agent_loop/src/executor/model.rs
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_agent_loop/src/executor/tests/cancellation.rs
  • crates/ironclaw_agent_loop/src/families/mod.rs
  • crates/ironclaw_agent_loop/src/strategies/recovery.rs
  • crates/ironclaw_loop_host/src/budget_accountant.rs
  • crates/ironclaw_loop_host/src/lib.rs
  • crates/ironclaw_loop_host/src/system_inference.rs
  • crates/ironclaw_runner/src/model_failure_mapping.rs
  • crates/ironclaw_runner/src/model_gateway.rs
  • crates/ironclaw_runner/src/planned_driver.rs
  • crates/ironclaw_runner/src/text_loop_driver.rs
  • crates/ironclaw_runner/tests/llm_gateway.rs
  • crates/ironclaw_runner/tests/loop_driver_host.rs
  • crates/ironclaw_turns/src/run_profile/host/error.rs
  • crates/ironclaw_turns/src/run_profile/model.rs
  • crates/ironclaw_turns/src/turn_state_row_store/turn_state_engine/transitions.rs
  • crates/ironclaw_turns/tests/agent_loop_host_contract.rs
  • docs/reborn/contracts/turn-runner.md
  • tests/integration/model_recovery.rs
  • tests/integration/support/assertions.rs
  • tests/integration/support/builder.rs
  • tests/integration/support/scripted_provider.rs

Comment thread crates/ironclaw_agent_loop/src/strategies/recovery.rs
Comment thread tests/integration/model_recovery.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6845 July 29, 2026 21:02 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
crates/ironclaw_runner/src/failure_summary.rs (2)

167-167: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Remove the unreachable duplicate checkpoint branch.

reborn_failure_summary_for_category returns at Lines 68–70 when pinned_failure_summary_for_category recognizes CHECKPOINT_REJECTED_CATEGORY at Lines 303–309. Therefore the match arm at Line 167 cannot execute; keeping both definitions creates two sources of truth that can drift.

Also applies to: 303-309

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_runner/src/failure_summary.rs` at line 167, Remove the
unreachable CHECKPOINT_REJECTED_CATEGORY match arm from
reborn_failure_summary_for_category and remove the corresponding pinned fallback
definition in pinned_failure_summary_for_category, leaving a single
authoritative checkpoint rejection mapping.

329-346: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win

Test the newly added invalid-cause rejection path.

The production change at Lines 44–50 validates the checkpoint cause, but this test covers only a valid round trip and an unknown stage. Add a valid-stage envelope with a validator-rejected cause and assert None.

As per path instructions, every changed failure-handling path needs regression coverage at the nearest meaningful seam.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_runner/src/failure_summary.rs` around lines 329 - 346, Add
regression coverage in
checkpoint_rejection_explanation_is_bounded_and_provenance_validated for a valid
CheckpointKind envelope containing a cause rejected by the validator, and assert
checkpoint_rejection_host_explanation_from_detail returns None. Keep the
existing valid round-trip and unknown-stage assertions unchanged.

Source: Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@crates/ironclaw_runner/src/failure_summary.rs`:
- Line 167: Remove the unreachable CHECKPOINT_REJECTED_CATEGORY match arm from
reborn_failure_summary_for_category and remove the corresponding pinned fallback
definition in pinned_failure_summary_for_category, leaving a single
authoritative checkpoint rejection mapping.
- Around line 329-346: Add regression coverage in
checkpoint_rejection_explanation_is_bounded_and_provenance_validated for a valid
CheckpointKind envelope containing a cause rejected by the validator, and assert
checkpoint_rejection_host_explanation_from_detail returns None. Keep the
existing valid round-trip and unknown-stage assertions unchanged.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 14b48dd5-2b4d-4812-93dc-e2bea5c89c94

📥 Commits

Reviewing files that changed from the base of the PR and between 4931832 and 8b1e952.

📒 Files selected for processing (4)
  • crates/ironclaw_agent_loop/src/executor/checkpoint.rs
  • crates/ironclaw_product/src/projection/tests/failure_explanation.rs
  • crates/ironclaw_runner/src/failure_summary.rs
  • crates/ironclaw_turns/src/run_profile/host/refs.rs
💤 Files with no reviewable changes (1)
  • crates/ironclaw_product/src/projection/tests/failure_explanation.rs

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6845 July 29, 2026 21:11 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6845 July 29, 2026 21:30 Destroyed
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Addressed the three outside-diff findings from CodeRabbit review 4812851195 in commit 7c4bfe4:

  • Preserve provider-reported usage when a Length response is rejected, carry it across every model-error boundary, and reconcile the failed call instead of releasing its reservation. The full truncation-recovery integration now proves 11 input / 7 output tokens are charged while partial output and textual tool syntax remain undispatched.
  • Treat missing fallback-route evidence as explicit absence and reject missing or mismatched evidence in both the main loop model port and system inference before accepting output.
  • Stop forwarding raw host detail for permanent prompt-stage failures; only the validated safe summary reaches the model-visible driver failure.

Verification: full ironclaw_turns, ironclaw_loop_host, and ironclaw_runner test suites; all 18 reborn_integration_model_recovery tests; warning-free clippy for affected crates and the integration target; dependent product/parity targets compiled.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_loop_host/src/lib.rs (1)

1275-1283: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Preserve provider usage on every rejected model call.

The accountant treats failures without usage as releasable reservations. These paths can have provider-reported tokens, so dropping usage undercharges spend and can bypass durable budget limits.

  • crates/ironclaw_loop_host/src/lib.rs#L1275-L1283: attach the destructured usage when rejecting missing or mismatched fallback evidence.
  • crates/ironclaw_runner/src/model_gateway.rs#L1815-L1824: attach usage to every tool-response failure arm, not only FinishReason::Length.
  • crates/ironclaw_runner/src/model_gateway.rs#L2149-L2153: apply the same handling to text-only failures, including empty stop responses, content filters, unsupported tool use, and unknown finishes.

As per path instructions, fail-loud handling must preserve the real accounting outcome; the PR objective also requires provider usage to survive rejected calls.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_loop_host/src/lib.rs` around lines 1275 - 1283, Preserve
provider-reported usage on every rejected model call. In
crates/ironclaw_loop_host/src/lib.rs lines 1275-1283, update the
fallback-evidence rejection around the response destructuring to attach usage to
the error. In crates/ironclaw_runner/src/model_gateway.rs lines 1815-1824 and
2149-2153, attach usage in every tool-response and text-only failure arm,
including empty stop responses, content filters, unsupported tool use, unknown
finishes, and non-length failures; retain fail-loud accounting behavior.

Source: Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_loop_host/src/budget_accountant.rs`:
- Around line 612-624: The ModelCallOutcome::Failure reconciliation path
currently derives pricing from the request preference or run default instead of
the model that actually handled the call. Persist the gateway-selected effective
model alongside failure usage, then have the Some(usage) branch use that
persisted model when calling usage_for_reported_usage; add a regression covering
a fallback model with different costs.

---

Outside diff comments:
In `@crates/ironclaw_loop_host/src/lib.rs`:
- Around line 1275-1283: Preserve provider-reported usage on every rejected
model call. In crates/ironclaw_loop_host/src/lib.rs lines 1275-1283, update the
fallback-evidence rejection around the response destructuring to attach usage to
the error. In crates/ironclaw_runner/src/model_gateway.rs lines 1815-1824 and
2149-2153, attach usage in every tool-response and text-only failure arm,
including empty stop responses, content filters, unsupported tool use, unknown
finishes, and non-length failures; retain fail-loud accounting behavior.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 4b8f5aa3-2685-4df9-83ba-2f0fc4aaf86c

📥 Commits

Reviewing files that changed from the base of the PR and between 4feb175 and 7c4bfe4.

📒 Files selected for processing (16)
  • crates/ironclaw_loop_host/src/budget_accountant.rs
  • crates/ironclaw_loop_host/src/lib.rs
  • crates/ironclaw_loop_host/src/system_inference.rs
  • crates/ironclaw_loop_host/tests/thread_loop_host_contract.rs
  • crates/ironclaw_product/tests/support/planned_agent_loop.rs
  • crates/ironclaw_runner/src/model_gateway.rs
  • crates/ironclaw_runner/src/model_gateway_error_mapping.rs
  • crates/ironclaw_runner/src/planned_driver.rs
  • crates/ironclaw_runner/tests/llm_gateway.rs
  • crates/ironclaw_runner/tests/loop_driver_host.rs
  • crates/ironclaw_turns/src/run_profile/host/error.rs
  • crates/ironclaw_turns/src/run_profile/model.rs
  • tests/integration/model_recovery.rs
  • tests/integration/support/assertions.rs
  • tests/integration/support/scripted_provider.rs
  • tests/support/reborn_parity_qa/binary_e2e.rs

Comment on lines +612 to +624
ModelCallOutcome::Failure(error) => match error.usage {
Some(usage) => {
let effective_model = request
.model_preference
.as_ref()
.unwrap_or(&context.resolved_run_profile.model_profile_id);
PendingAccounting::Reconcile(usage_for_reported_usage(
usage,
0,
self.cost_table.as_ref(),
effective_model,
&self.default_cost,
))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Preserve the failed call’s effective model for reconciliation.

Line 614 prices failure usage from request.model_preference or the run default. That is not necessarily the route that ran: a fallback-indexed call may consume tokens on a differently priced model. Persist the gateway-selected effective model with failure usage and use it here; add a differing-cost fallback regression.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_loop_host/src/budget_accountant.rs` around lines 612 - 624,
The ModelCallOutcome::Failure reconciliation path currently derives pricing from
the request preference or run default instead of the model that actually handled
the call. Persist the gateway-selected effective model alongside failure usage,
then have the Some(usage) branch use that persisted model when calling
usage_for_reported_usage; add a regression covering a fallback model with
different costs.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6845 July 30, 2026 08:40 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6845 July 30, 2026 08:53 Destroyed
@serrrfirat
serrrfirat merged commit a378b39 into main Jul 30, 2026
65 checks passed
@serrrfirat
serrrfirat deleted the codex/error-recoverability-ws2 branch July 30, 2026 09:18
serrrfirat added a commit that referenced this pull request Jul 30, 2026
* fix(reborn): complete model error recovery contract

* fix(reborn): preserve terminal model error explanations

* fix(reborn): address terminal error review findings

* fix(reborn): address recovery review findings (#6845)

* fix(reborn): close rejected checkpoints durably (#6861)

* fix(reborn): explain rejected checkpoints durably

* Format reconciled checkpoint projection imports

* fix(reborn): retry idempotent transcript writes

* fix(reborn): address checkpoint recovery review comments

* fix(reborn): preserve selected model fallback

* fix(reborn): harden model failure accounting

* fix(reborn): address terminal error review

* fix(agent-loop): use honest transcript fallback

* test(loop-host): exercise transcript error classification

* test(threads): cover backend error classification

* fix(ci): exclude Rust unit tests from coverage gate
serrrfirat added a commit that referenced this pull request Aug 7, 2026
… SSE transport

event-source-plus (fetch/ReadableStream). The legacy suites still faked
window.EventSource, so the app never opened a stream and every
duration-1/duration-4 legacy test that emitted frames failed with "no
EventSource stream is open".

- Extract the smoke suite's proven fetch fake into
  install_fake_v2_event_stream() in reborn_webui_harness, extended to
  record request URLs and headers for reconnect assertions.
- Port all seven legacy scenario files onto it, updating cursor/token
  assertions to the header contract (Authorization bearer,
  Last-Event-ID) instead of the retired token/after_cursor query
  params.
- legacy_skills delete: use the shared in-app confirmation dialog
  instead of a native browser dialog.
- legacy_dom_resource_limits reconnect-timer: assert the pending
  reconnect is cancelled when the tab hides (the fetch transport
  schedules retries internally).
- legacy_rendering: assert no live onerror/iframe/img nodes instead of
  substring-scanning escaped text.
- extensions_api: restore the #6520 wire contract (retired
  authenticated/active/needs_setup/has_auth/onboarding_state booleans
  must be absent).
- tool_execution truncated-tool test: expect model_output_truncated
  failure per #6845's no-recovery contract instead of an assistant
  recovery message.
- streaming_run_control_api: drop the stream=true assertion for the
  OpenAI-compatible mock, which rides the buffered fallback since
  #7120 (rig-core cannot distinguish a complete stream from a truncated
  one).
serrrfirat added a commit that referenced this pull request Aug 8, 2026
… SSE transport

event-source-plus (fetch/ReadableStream). The legacy suites still faked
window.EventSource, so the app never opened a stream and every
duration-1/duration-4 legacy test that emitted frames failed with "no
EventSource stream is open".

- Extract the smoke suite's proven fetch fake into
  install_fake_v2_event_stream() in reborn_webui_harness, extended to
  record request URLs and headers for reconnect assertions.
- Port all seven legacy scenario files onto it, updating cursor/token
  assertions to the header contract (Authorization bearer,
  Last-Event-ID) instead of the retired token/after_cursor query
  params.
- legacy_skills delete: use the shared in-app confirmation dialog
  instead of a native browser dialog.
- legacy_dom_resource_limits reconnect-timer: assert the pending
  reconnect is cancelled when the tab hides (the fetch transport
  schedules retries internally).
- legacy_rendering: assert no live onerror/iframe/img nodes instead of
  substring-scanning escaped text.
- extensions_api: restore the #6520 wire contract (retired
  authenticated/active/needs_setup/has_auth/onboarding_state booleans
  must be absent).
- tool_execution truncated-tool test: expect model_output_truncated
  failure per #6845's no-recovery contract instead of an assistant
  recovery message.
- streaming_run_control_api: drop the stream=true assertion for the
  OpenAI-compatible mock, which rides the buffered fallback since
  #7120 (rig-core cannot distinguish a complete stream from a truncated
  one).
serrrfirat added a commit that referenced this pull request Aug 8, 2026
… SSE transport

event-source-plus (fetch/ReadableStream). The legacy suites still faked
window.EventSource, so the app never opened a stream and every
duration-1/duration-4 legacy test that emitted frames failed with "no
EventSource stream is open".

- Extract the smoke suite's proven fetch fake into
  install_fake_v2_event_stream() in reborn_webui_harness, extended to
  record request URLs and headers for reconnect assertions.
- Port all seven legacy scenario files onto it, updating cursor/token
  assertions to the header contract (Authorization bearer,
  Last-Event-ID) instead of the retired token/after_cursor query
  params.
- legacy_skills delete: use the shared in-app confirmation dialog
  instead of a native browser dialog.
- legacy_dom_resource_limits reconnect-timer: assert the pending
  reconnect is cancelled when the tab hides (the fetch transport
  schedules retries internally).
- legacy_rendering: assert no live onerror/iframe/img nodes instead of
  substring-scanning escaped text.
- extensions_api: restore the #6520 wire contract (retired
  authenticated/active/needs_setup/has_auth/onboarding_state booleans
  must be absent).
- tool_execution truncated-tool test: expect model_output_truncated
  failure per #6845's no-recovery contract instead of an assistant
  recovery message.
- streaming_run_control_api: drop the stream=true assertion for the
  OpenAI-compatible mock, which rides the buffered fallback since
  #7120 (rig-core cannot distinguish a complete stream from a truncated
  one).
pull Bot pushed a commit to Stars1233/ironclaw that referenced this pull request Aug 10, 2026
* fix(composition): read landed attachments through the per-caller workspace mount

/projects/workspace/tenants/{tenant}/users/{user}, but the loop-host
attachment_read_port still read through the shared read-only fixed view
(services.workspace_filesystem), which resolves the workspace root. A
landed image therefore came back NotFound at model-gateway time and was
silently dropped, so vision-capable model payloads lost every inline
image (the duration-4 Playwright attachment failure).

Wire the read port over the same per-caller scoped handle the WebUI
lander uses (runtime_mounts::read_write_workspace_filesystem), mirroring
what nearai#7062 already did for the channel-host assembly. Under the Shared
policy the handle is byte-identical to the old fixed view; under
PerCaller it now resolves the caller's subtree.

* test(playwright): reconcile legacy WebUI v2 suites to the fetch-based SSE transport

event-source-plus (fetch/ReadableStream). The legacy suites still faked
window.EventSource, so the app never opened a stream and every
duration-1/duration-4 legacy test that emitted frames failed with "no
EventSource stream is open".

- Extract the smoke suite's proven fetch fake into
  install_fake_v2_event_stream() in reborn_webui_harness, extended to
  record request URLs and headers for reconnect assertions.
- Port all seven legacy scenario files onto it, updating cursor/token
  assertions to the header contract (Authorization bearer,
  Last-Event-ID) instead of the retired token/after_cursor query
  params.
- legacy_skills delete: use the shared in-app confirmation dialog
  instead of a native browser dialog.
- legacy_dom_resource_limits reconnect-timer: assert the pending
  reconnect is cancelled when the tab hides (the fetch transport
  schedules retries internally).
- legacy_rendering: assert no live onerror/iframe/img nodes instead of
  substring-scanning escaped text.
- extensions_api: restore the nearai#6520 wire contract (retired
  authenticated/active/needs_setup/has_auth/onboarding_state booleans
  must be absent).
- tool_execution truncated-tool test: expect model_output_truncated
  failure per nearai#6845's no-recovery contract instead of an assistant
  recovery message.
- streaming_run_control_api: drop the stream=true assertion for the
  OpenAI-compatible mock, which rides the buffered fallback since
  nearai#7120 (rig-core cannot distinguish a complete stream from a truncated
  one).
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
* fix(composition): read landed attachments through the per-caller workspace mount

/projects/workspace/tenants/{tenant}/users/{user}, but the loop-host
attachment_read_port still read through the shared read-only fixed view
(services.workspace_filesystem), which resolves the workspace root. A
landed image therefore came back NotFound at model-gateway time and was
silently dropped, so vision-capable model payloads lost every inline
image (the duration-4 Playwright attachment failure).

Wire the read port over the same per-caller scoped handle the WebUI
lander uses (runtime_mounts::read_write_workspace_filesystem), mirroring
what nearai#7062 already did for the channel-host assembly. Under the Shared
policy the handle is byte-identical to the old fixed view; under
PerCaller it now resolves the caller's subtree.

* test(playwright): reconcile legacy WebUI v2 suites to the fetch-based SSE transport

event-source-plus (fetch/ReadableStream). The legacy suites still faked
window.EventSource, so the app never opened a stream and every
duration-1/duration-4 legacy test that emitted frames failed with "no
EventSource stream is open".

- Extract the smoke suite's proven fetch fake into
  install_fake_v2_event_stream() in reborn_webui_harness, extended to
  record request URLs and headers for reconnect assertions.
- Port all seven legacy scenario files onto it, updating cursor/token
  assertions to the header contract (Authorization bearer,
  Last-Event-ID) instead of the retired token/after_cursor query
  params.
- legacy_skills delete: use the shared in-app confirmation dialog
  instead of a native browser dialog.
- legacy_dom_resource_limits reconnect-timer: assert the pending
  reconnect is cancelled when the tab hides (the fetch transport
  schedules retries internally).
- legacy_rendering: assert no live onerror/iframe/img nodes instead of
  substring-scanning escaped text.
- extensions_api: restore the nearai#6520 wire contract (retired
  authenticated/active/needs_setup/has_auth/onboarding_state booleans
  must be absent).
- tool_execution truncated-tool test: expect model_output_truncated
  failure per nearai#6845's no-recovery contract instead of an assistant
  recovery message.
- streaming_run_control_api: drop the stream=true assertion for the
  OpenAI-compatible mock, which rides the buffered fallback since
  nearai#7120 (rig-core cannot distinguish a complete stream from a truncated
  one).
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
* fix(composition): read landed attachments through the per-caller workspace mount

/projects/workspace/tenants/{tenant}/users/{user}, but the loop-host
attachment_read_port still read through the shared read-only fixed view
(services.workspace_filesystem), which resolves the workspace root. A
landed image therefore came back NotFound at model-gateway time and was
silently dropped, so vision-capable model payloads lost every inline
image (the duration-4 Playwright attachment failure).

Wire the read port over the same per-caller scoped handle the WebUI
lander uses (runtime_mounts::read_write_workspace_filesystem), mirroring
what nearai#7062 already did for the channel-host assembly. Under the Shared
policy the handle is byte-identical to the old fixed view; under
PerCaller it now resolves the caller's subtree.

* test(playwright): reconcile legacy WebUI v2 suites to the fetch-based SSE transport

event-source-plus (fetch/ReadableStream). The legacy suites still faked
window.EventSource, so the app never opened a stream and every
duration-1/duration-4 legacy test that emitted frames failed with "no
EventSource stream is open".

- Extract the smoke suite's proven fetch fake into
  install_fake_v2_event_stream() in reborn_webui_harness, extended to
  record request URLs and headers for reconnect assertions.
- Port all seven legacy scenario files onto it, updating cursor/token
  assertions to the header contract (Authorization bearer,
  Last-Event-ID) instead of the retired token/after_cursor query
  params.
- legacy_skills delete: use the shared in-app confirmation dialog
  instead of a native browser dialog.
- legacy_dom_resource_limits reconnect-timer: assert the pending
  reconnect is cancelled when the tab hides (the fetch transport
  schedules retries internally).
- legacy_rendering: assert no live onerror/iframe/img nodes instead of
  substring-scanning escaped text.
- extensions_api: restore the nearai#6520 wire contract (retired
  authenticated/active/needs_setup/has_auth/onboarding_state booleans
  must be absent).
- tool_execution truncated-tool test: expect model_output_truncated
  failure per nearai#6845's no-recovery contract instead of an assistant
  recovery message.
- streaming_run_control_api: drop the stream=true assertion for the
  OpenAI-compatible mock, which rides the buffered fallback since
  nearai#7120 (rig-core cannot distinguish a complete stream from a truncated
  one).
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
* fix(reborn): complete model error recovery contract

* fix(reborn): address recovery review findings (nearai#6845)

* fix(reborn): close rejected checkpoints durably (nearai#6861)

* fix(reborn): explain rejected checkpoints durably

* Format reconciled checkpoint projection imports

* fix(reborn): address checkpoint recovery review comments

* fix(reborn): preserve selected model fallback

* fix(reborn): harden model failure accounting

* Fix Reborn E2E checkpoint retry gate
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
* fix(reborn): complete model error recovery contract

* fix(reborn): preserve terminal model error explanations

* fix(reborn): address terminal error review findings

* fix(reborn): address recovery review findings (nearai#6845)

* fix(reborn): close rejected checkpoints durably (nearai#6861)

* fix(reborn): explain rejected checkpoints durably

* Format reconciled checkpoint projection imports

* fix(reborn): retry idempotent transcript writes

* fix(reborn): address checkpoint recovery review comments

* fix(reborn): preserve selected model fallback

* fix(reborn): harden model failure accounting

* fix(reborn): address terminal error review

* fix(agent-loop): use honest transcript fallback

* test(loop-host): exercise transcript error classification

* test(threads): cover backend error classification

* fix(ci): exclude Rust unit tests from coverage gate
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
* fix(composition): read landed attachments through the per-caller workspace mount

/projects/workspace/tenants/{tenant}/users/{user}, but the loop-host
attachment_read_port still read through the shared read-only fixed view
(services.workspace_filesystem), which resolves the workspace root. A
landed image therefore came back NotFound at model-gateway time and was
silently dropped, so vision-capable model payloads lost every inline
image (the duration-4 Playwright attachment failure).

Wire the read port over the same per-caller scoped handle the WebUI
lander uses (runtime_mounts::read_write_workspace_filesystem), mirroring
what nearai#7062 already did for the channel-host assembly. Under the Shared
policy the handle is byte-identical to the old fixed view; under
PerCaller it now resolves the caller's subtree.

* test(playwright): reconcile legacy WebUI v2 suites to the fetch-based SSE transport

event-source-plus (fetch/ReadableStream). The legacy suites still faked
window.EventSource, so the app never opened a stream and every
duration-1/duration-4 legacy test that emitted frames failed with "no
EventSource stream is open".

- Extract the smoke suite's proven fetch fake into
  install_fake_v2_event_stream() in reborn_webui_harness, extended to
  record request URLs and headers for reconnect assertions.
- Port all seven legacy scenario files onto it, updating cursor/token
  assertions to the header contract (Authorization bearer,
  Last-Event-ID) instead of the retired token/after_cursor query
  params.
- legacy_skills delete: use the shared in-app confirmation dialog
  instead of a native browser dialog.
- legacy_dom_resource_limits reconnect-timer: assert the pending
  reconnect is cancelled when the tab hides (the fetch transport
  schedules retries internally).
- legacy_rendering: assert no live onerror/iframe/img nodes instead of
  substring-scanning escaped text.
- extensions_api: restore the nearai#6520 wire contract (retired
  authenticated/active/needs_setup/has_auth/onboarding_state booleans
  must be absent).
- tool_execution truncated-tool test: expect model_output_truncated
  failure per nearai#6845's no-recovery contract instead of an assistant
  recovery message.
- streaming_run_control_api: drop the stream=true assertion for the
  OpenAI-compatible mock, which rides the buffered fallback since
  nearai#7120 (rig-core cannot distinguish a complete stream from a truncated
  one).

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6845 — ad80dc9e Deployed Jul 30, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Output truncation is recovered by shrinking the input, and overloads BudgetExceeded a third time

1 participant