Skip to content

feat(turns): RunFailureReason funnel foundation (phase 1 of 4) - #5954

Closed
ilblackdragon wants to merge 1 commit into
mainfrom
run-failure-reason-funnel
Closed

ilblackdragon wants to merge 1 commit into
mainfrom
run-failure-reason-funnel

Conversation

@ilblackdragon

Copy link
Copy Markdown
Member

What

Phase 1 of the "every terminal run failure passes one classifier" funnel — the target architecture from the error-recoverability audit, and the run-boundary companion to the exhaustive-classification keystone (#5651). Additive and consumed by nothing yet — zero behavior change.

Adds ironclaw_turns::run_failure:

  • RunFailureReason — all-private fields, only two constructors: classify() and security_stop(). Once the run-exit contracts carry it (phase 2), a terminal failure cannot be recorded without passing a funnel — an unsurfaced terminal error becomes structurally unrepresentable.
  • RunFailureCategory — a wildcard-free enum over every terminal category (LoopFailureKind, protocol violations, driver/scheduler paths, model-provider categories). classify matches it exhaustively, so a new category can't be added without deciding its lane/retry/message (compile error otherwise — ties into refactor(errors): static enforcement that failures surface, not swallow #5651).
  • FailureLane { SecurityStop, Retriable, Explainable } + RetryPolicy. classify never yields SecurityStop; the only path to it is security_stop, gated behind SafetyStopEvidence (mintable only by the safety layer — wired in phase 4).
  • UserMessage — a non-empty newtype, so "a terminal failure with no user-facing explanation" is unrepresentable. The per-category message table is moved in from reborn_composition::failure_summary.
  • category() returns the wire-stable SanitizedFailure so persistence stays byte-identical (lane/retry/message are recomputable from the category on read).

Tests

cargo test -p ironclaw_turns — the module's 5 tests pass: exhaustive per-category classification (non-empty message, never SecurityStop), category-string round-trip, lane/retry agreement, and the security_stop path.

Follow-up PRs (this is 1 of 4)

  1. Wire the funnel through the recording paths — change the four in-process contract failure fields (FailRunRequest, RecordRunnerFailureRequest, TurnRunnerOutcome::Failed, LoopExitMapping::RecoveryRequired) to RunFailureReason; extract .category() where TurnRunState.failure is written (durable persistence stays byte-identical); route the three producer paths through classify. Preserves the exactly-once fail_run invariant.
  2. Consolidate surfacing — source the projection's baseline user_message from classify() (keep FailureExplanationProvider as the model-enrichment override); add lane/retry_policy to ProductProjectionItem::RunStatus (the UI retry signal).
  3. SecurityStop producer — the ironclaw_safety → ironclaw_turns edge, mint SafetyStopEvidence at the leak-block site, and a grep-based origin test locking SecurityStop to the safety layer.

🤖 Generated with Claude Code

Copilot AI review requested due to automatic review settings July 10, 2026 17:14
@ironloopai

ironloopai Bot commented Jul 10, 2026 •

Copy link
Copy Markdown
Contributor

🔎 IronLoop Review Status

Head: 805cb13b2a39f3adc18ef483db0eb5c15d8a159d
Result: One or more review results were superseded by a newer PR head.
Next: Run @ironloopai review on the latest PR head.
Updated: 2026-07-10T21:05:45.506Z

Current reviewers:

Reviewer State Verdict Findings Last update
ironloop/common-reviewer (reviewer) Superseded N/A N/A 2026-07-10T21:05:45.489Z
Reviewer summaries
Reviewer Detail
ironloop/common-reviewer (reviewer) Superseded by a newer PR head. New head: 805cb13. Previous verdict: Changes requested.
Recent activity
Time Reviewer State Detail
2026-07-10T17:14:49.630Z ironloop/common-reviewer (reviewer) Queued Accepted review request for head a14ab58.
2026-07-10T17:14:49.630Z ironloop/common-reviewer (reviewer) Queued Waiting for this reviewer lane to become available.
2026-07-10T17:14:50.342Z ironloop/common-reviewer (reviewer) Started Reviewer worker started attempt 1.
2026-07-10T17:14:52.747Z ironloop/common-reviewer (reviewer) Workspace ready Prepared isolated checkout (head_ref) at a14ab58.
2026-07-10T17:18:32.987Z ironloop/common-reviewer (reviewer) Superseded Old-head reviewer is still running after newer head 805cb13 replaced it. Codex is reviewing; process live; elapsed 3m 41s; timeout in 16m 19s; last heartbeat 2026-07-10T17:18:32.987Z. Codex emitted stderr output at 2026-07-10T17:17:22.301Z.
2026-07-10T17:18:52.221Z ironloop/common-reviewer (reviewer) Result captured Changes requested; 2 blocking findings.
2026-07-10T17:18:52.221Z ironloop/common-reviewer (reviewer) Completed Review completed and terminal status was persisted.
2026-07-10T21:05:45.489Z ironloop/common-reviewer (reviewer) Superseded A newer PR head replaced this review (805cb13).
Available commands
  • @ironloopai help
  • @ironloopai agents
  • @ironloopai review
  • @ironloopai review --agent <agent>
  • @ironloopai status
Run metadata

Admission: webhook accepted the request and IronLoop persisted reviewer state before this projection.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5954 July 10, 2026 17:14 Destroyed
@github-actions github-actions Bot added size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 10, 2026
@coderabbitai

coderabbitai Bot commented Jul 10, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • New Features
    • Added standardized run-failure classification with stable categories, retry guidance, failure sources, and user-facing messages.
    • Added dedicated handling for safety-related stops with sanitized evidence.
    • Added validation to prevent empty or whitespace-only failure messages.
    • Exposed failure details such as category, retry policy, processing lane, and correlation ID for consistent reporting.

Walkthrough

Adds a public run_failure module containing typed terminal-failure classification, stable categories, retry metadata, user messages, safety-stop evidence, construction funnels, accessors, and unit tests.

Changes

Run failure classification

Layer / File(s) Summary
Failure metadata contracts
crates/ironclaw_turns/src/run_failure.rs
Defines failure lanes, retry policies, failure sources, and non-empty validated user messages.
Classification and terminal carrier
crates/ironclaw_turns/src/run_failure.rs
Adds wire-stable categories, persisted-string conversion, host-authored messages, safety-stop evidence, and restricted RunFailureReason constructors.
Public exports and invariants
crates/ironclaw_turns/src/lib.rs, crates/ironclaw_turns/src/run_failure.rs
Exposes the module and types, with tests covering category round-trips, retry mappings, message validation, and exclusive security-stop construction.

Estimated code review effort: 4 (Complex) | ~45 minutes

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description covers the change and tests, but it misses required template sections like Summary, Change Type, Linked Issue, and Security Impact. Rewrite it to follow the template: add the required headings, change type, linked issue, validation, security impact, rollback, and review follow-through.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title uses conventional commits and accurately summarizes the new run-failure funnel foundation.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
⚔️ Resolve merge conflicts
  • Resolve merge conflict in branch run-failure-reason-funnel

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new run_failure module to provide structured, single-funnel classification for terminal run failures, mapping them to specific recovery lanes, retry policies, and user-facing messages. The review feedback highlights a critical issue in RunFailureReason::from_category_str where unrecognized or newer category strings are mapped to UnknownFailure, causing the original category string to be lost and overwritten with 'unknown_failure'. A code suggestion is provided to preserve the original category string while maintaining the fallback classification behavior.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +439 to +450
/// Classify from a raw persisted/wire category string (recompute on read).
pub fn from_category_str(
category: &str,
source: FailureSource,
correlation_id: TurnRunId,
) -> Self {
Self::classify(
RunFailureCategory::from_category_str(category),
source,
correlation_id,
)
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

When reconstructing a RunFailureReason from a persisted category string using from_category_str, any unrecognized/newer category string (e.g., "some_future_category") is mapped to RunFailureCategory::UnknownFailure. This causes Self::classify to set the internal category field to "unknown_failure", completely losing the original category string.

Since the PR design aims for byte-identical persistence (where the category is recomputed on read and written back as-is), this loss of fidelity violates that guarantee and mutates newer categories to "unknown_failure" upon write-back.

We should preserve the original category string (if it is a valid SanitizedFailure) while still classifying it under the UnknownFailure rules.

    /// Classify from a raw persisted/wire category string (recompute on read).
    pub fn from_category_str(
        category: &str,
        source: FailureSource,
        correlation_id: TurnRunId,
    ) -> Self {
        let parsed = RunFailureCategory::from_category_str(category);
        let (lane, retry_policy) = parsed.lane_and_policy();
        let sanitized_category = SanitizedFailure::new(category)
            .unwrap_or_else(|_| SanitizedFailure::from_trusted_static(parsed.as_str()));
        Self {
            lane,
            retry_policy,
            user_message: parsed.user_message(),
            correlation_id,
            category: sanitized_category,
        }
    }

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

❌ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
❌ Changes requested 2 0 2 a14ab587380d

Head: a14ab587380d147f66313a03d34f1525c20d5f05
Next: Fix the blocking findings, push the PR branch, then re-run this reviewer.

Run details

Status: Current
Needs human: no
Needs validation: no

Summary

Found two blocking issues in the new run-failure funnel API: the safety-stop evidence token is publicly mintable, and persisted security stops cannot round-trip back to the SecurityStop lane.

Findings

Blocking: 2 / Notes: 0

Blocking findings

1. ❌ [HIGH] Safety-stop evidence can be minted by any caller

Location: crates/ironclaw_turns/src/run_failure.rs:393
SafetyStopEvidence::from_safety_layer is a public constructor on a public type that is also re-exported from ironclaw_turns, so any crate that can depend on ironclaw_turns can mint this token and call RunFailureReason::security_stop. That makes the promised safety-only origin false and lets non-safety code classify arbitrary terminal failures as SecurityStop. Move the proof/token construction to a safety-owned boundary, restrict the constructor, or otherwise enforce this with a boundary/compile-fail test before exposing the lane.

2. ❌ [MEDIUM] Persisted security stops rehydrate as ordinary unknown failures

Location: crates/ironclaw_turns/src/run_failure.rs:440-449
security_stop persists evidence.category() as the only stored category, but from_category_str maps any category outside RunFailureCategory to UnknownFailure and then classify rewrites it to unknown_failure with the Explainable lane. A real safety category such as prompt_injection_blocked would therefore lose both its original category and its SecurityStop lane when recomputed from persisted state. Store the lane, reserve/parse security-stop categories or prefixes, or add a security-aware persisted form and cover the round trip in tests.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.
  4. Use @ironloopai status to check queued/running/completed/failed/superseded state while reviewers run.

/// Mint security-stop evidence from the safety layer. The caller is
/// `ironclaw_safety` at a real block decision (prompt-injection detection or
/// a confirmed secret leak); `category` is the sanitized block category.
pub fn from_safety_layer(category: SanitizedFailure) -> Self {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This constructor is public on a public, re-exported type, so every ironclaw_turns consumer can mint SafetyStopEvidence and call RunFailureReason::security_stop. That does not enforce the documented safety-layer-only origin; please move/restrict the minting boundary and add a boundary/compile-fail guard.

}

/// Classify from a raw persisted/wire category string (recompute on read).
pub fn from_category_str(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This read-side recomputation cannot reconstruct SecurityStop: security_stop stores an arbitrary safety category, while unknown strings here become UnknownFailure and are re-persisted as unknown_failure with the Explainable lane. Please add a security-aware persisted representation or parser and a round-trip test.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the initial “single-funnel” data model for terminal run failures in ironclaw_turns, introducing a typed classification layer intended to make unclassified terminal failures structurally unrepresentable once wired into run-exit contracts in later phases.

Changes:

  • Introduces run_failure module with RunFailureReason, RunFailureCategory, lane/policy enums, and a non-empty UserMessage newtype.
  • Adds parsing/round-trip helpers for wire-stable failure category strings and a baseline per-category user message table.
  • Exposes the new types from ironclaw_turns’ public API via lib.rs re-exports.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 3 comments.

File Description
crates/ironclaw_turns/src/run_failure.rs Adds the failure funnel foundation: categories, classification into lane/retry/message, security-stop pathway, and tests.
crates/ironclaw_turns/src/lib.rs Exposes the new run_failure module and re-exports its public types.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +389 to +395
impl SafetyStopEvidence {
/// Mint security-stop evidence from the safety layer. The caller is
/// `ironclaw_safety` at a real block decision (prompt-injection detection or
/// a confirmed secret leak); `category` is the sanitized block category.
pub fn from_safety_layer(category: SanitizedFailure) -> Self {
Self { category }
}
Comment on lines +28 to +32
/// Which of the three terminal outcomes a run failure resolves to.
///
/// This is the two-bucket end state: a security-related failure stops the run;
/// everything else is user-explainable and, where safe, retriable.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
Comment on lines +126 to +130
/// Provenance of a terminal failure — which recording path produced it.
///
/// Carried alongside the category so a future classifier can nuance lane/retry
/// by source without re-deriving it. Today the category alone determines the
/// classification; `source` is retained for diagnostics and forward flexibility.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_turns/src/run_failure.rs`:
- Around line 389-400: Make the minting guarantee accurate: either introduce an
unforgeable/private capability required by SafetyStopEvidence::from_safety_layer
and ensure only the safety layer can obtain it, or soften the method and related
documentation to explicitly describe this as a convention rather than an
enforced boundary. Review SanitizedFailure::new, the SafetyStopEvidence
re-export, and RunFailureReason::security_stop so the public API does not claim
downstream crates are unable to forge tokens.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 6a8ff907-2c41-4252-a5cc-09a4053a2187

📥 Commits

Reviewing files that changed from the base of the PR and between d2bb9d5 and a14ab58.

📒 Files selected for processing (2)
  • crates/ironclaw_turns/src/lib.rs
  • crates/ironclaw_turns/src/run_failure.rs

Comment on lines +389 to +400
impl SafetyStopEvidence {
/// Mint security-stop evidence from the safety layer. The caller is
/// `ironclaw_safety` at a real block decision (prompt-injection detection or
/// a confirmed secret leak); `category` is the sanitized block category.
pub fn from_safety_layer(category: SanitizedFailure) -> Self {
Self { category }
}

fn category(&self) -> SanitizedFailure {
self.category.clone()
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# How is SafetyStopEvidence minted across the workspace, and is there a real seal?
rg -nP '\bfrom_safety_layer\s*\(' --type=rust
rg -nP 'SafetyStopEvidence' --type=rust -C2
# Any sealed-trait / crate-private mechanism claimed by the "boundary test"?
rg -nP 'security_stop|single.?origin|boundary test' --type=rust -C2

Repository: nearai/ironclaw

Length of output: 153


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo '--- run_failure.rs outline ---'
ast-grep outline crates/ironclaw_turns/src/run_failure.rs --view expanded || true

echo '--- run_failure.rs lines 340-430 ---'
sed -n '340,430p' crates/ironclaw_turns/src/run_failure.rs | cat -n

echo '--- search SafetyStopEvidence / SanitizedFailure / security_stop ---'
rg -n 'SafetyStopEvidence|SanitizedFailure|security_stop|from_safety_layer|single.?origin|boundary test|seal|sealed' crates/ironclaw_turns/src -C 3 || true

echo '--- find SanitizedFailure definition ---'
rg -n 'struct SanitizedFailure|enum SanitizedFailure|impl SanitizedFailure' crates/ironclaw_turns/src -C 3 || true

Repository: nearai/ironclaw

Length of output: 38529


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo '--- call sites of RunFailureReason::security_stop / SafetyStopEvidence::from_safety_layer ---'
rg -n 'RunFailureReason::security_stop|SafetyStopEvidence::from_safety_layer|security_stop\(' crates src tests -C 2 || true

echo '--- workspace members / published surface clues ---'
sed -n '1,220p' Cargo.toml | cat -n

echo '--- ironclaw_turns lib visibility ---'
sed -n '1,160p' crates/ironclaw_turns/src/lib.rs | cat -n

Repository: nearai/ironclaw

Length of output: 22409


SafetyStopEvidence is forgeable from downstream crates
from_safety_layer is pub, SanitizedFailure::new is pub, and SafetyStopEvidence is re-exported from crates/ironclaw_turns/src/lib.rs, so any crate that can import ironclaw_turns can mint a security-stop token and call RunFailureReason::security_stop(...). The doc’s “only ironclaw_safety can mint” claim is too strong unless there’s a real seal elsewhere. Either add an unforgeable guard or soften the API/docs to make this a convention, not a compile-time boundary.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_turns/src/run_failure.rs` around lines 389 - 400, Make the
minting guarantee accurate: either introduce an unforgeable/private capability
required by SafetyStopEvidence::from_safety_layer and ensure only the safety
layer can obtain it, or soften the method and related documentation to
explicitly describe this as a convention rather than an enforced boundary.
Review SanitizedFailure::new, the SafetyStopEvidence re-export, and
RunFailureReason::security_stop so the public API does not claim downstream
crates are unable to forge tokens.

Sources: Coding guidelines, Path instructions

@railway-app

railway-app Bot commented Jul 10, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5954 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 10, 2026 at 9:15 pm

Add `ironclaw_turns::run_failure` — the single classification funnel every
terminal run failure will route through (target architecture from the
error-recoverability audit; follows the exhaustive-classification keystone

Phase 1 is additive and consumed by nothing yet (no behavior change):

- `RunFailureReason` with ALL-private fields and only two constructors,
  `classify()` and `security_stop()`. Once the run-exit contracts carry it
  (phase 2), a terminal failure cannot be recorded without passing a funnel.
- `RunFailureCategory`: a wildcard-free enum over every terminal category
  (LoopFailureKind, protocol violations, driver/scheduler paths, model
  provider categories). `classify` matches it exhaustively, so a new
  category can't be added without deciding its lane/retry/message.
- `FailureLane { SecurityStop, Retriable, Explainable }` + `RetryPolicy`.
  `classify` never yields SecurityStop; the only path to it is
  `security_stop`, gated behind `SafetyStopEvidence` (mintable only by the
  safety layer — wired in phase 4).
- `UserMessage`: non-empty newtype, so "a terminal failure with no
  user-facing explanation" is unrepresentable. The per-category message
  table is moved in from `reborn_composition::failure_summary`.
- `category()` returns the wire-stable `SanitizedFailure` so persistence
  stays byte-identical (lane/retry/message are recomputable on read).

Tests: exhaustive per-category classification (non-empty message, never
SecurityStop), category-string round-trip, lane/retry agreement, and the
security_stop path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 10, 2026 21:05
@ilblackdragon
ilblackdragon force-pushed the run-failure-reason-funnel branch from a14ab58 to 805cb13 Compare July 10, 2026 21:05
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5954 July 10, 2026 21:05 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated 7 comments.

Comment on lines +20 to +23
//! 3. **`SecurityStop` originates only in the safety layer.** `classify` never
//! yields [`FailureLane::SecurityStop`]; the only path to it is
//! [`RunFailureReason::security_stop`], gated behind a [`SafetyStopEvidence`]
//! token that only `ironclaw_safety` can mint.
Comment on lines +28 to +31
/// Which of the three terminal outcomes a run failure resolves to.
///
/// This is the two-bucket end state: a security-related failure stops the run;
/// everything else is user-explainable and, where safe, retriable.
Comment on lines +126 to +130
/// Provenance of a terminal failure — which recording path produced it.
///
/// Carried alongside the category so a future classifier can nuance lane/retry
/// by source without re-deriving it. Today the category alone determines the
/// classification; `source` is retained for diagnostics and forward flexibility.
/// reachable only from the safety layer. The single-origin property is
/// additionally locked by a boundary test.
#[derive(Debug, Clone)]
pub struct SafetyStopEvidence {
/// Mint security-stop evidence from the safety layer. The caller is
/// `ironclaw_safety` at a real block decision (prompt-injection detection or
/// a confirmed secret leak); `category` is the sanitized block category.
pub fn from_safety_layer(category: SanitizedFailure) -> Self {

/// The security-stop funnel: the only path to [`FailureLane::SecurityStop`].
/// Requires [`SafetyStopEvidence`], which only the safety layer can mint.
pub fn security_stop(evidence: SafetyStopEvidence, correlation_id: TurnRunId) -> Self {
Comment on lines +94 to +97
pub use run_failure::{
FailureLane, FailureSource, RetryPolicy, RunFailureCategory, RunFailureReason,
SafetyStopEvidence, UserMessage, UserMessageError,
};
@github-actions

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.18% (290621 / 341200 lines)
  floor:    85.3% (tolerance 0.5pp -> effective floor 84.8%)
  denominator: 341200 lines now vs 320188 at floor capture (+21012 lines, +6.56%) — material change (>5%)

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.18% — 290621 / 341200 lines

Per-crate breakdown (63 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_runtime_policy 31.75% 80 / 252
ironclaw_event_projections 43.31% 673 / 1554
ironclaw_run_state 52.36% 222 / 424
ironclaw_authorization 53.66% 462 / 861
ironclaw_triggers 59.89% 1792 / 2992
ironclaw_observability 61.54% 16 / 26
ironclaw_reborn_cli 62.5% 3766 / 6026
ironclaw_webui_v2 62.62% 2632 / 4203
ironclaw_mcp 63.03% 578 / 917
ironclaw_reborn_migration 66.93% 1168 / 1745
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_filesystem 67.2% 3833 / 5704
ironclaw_memory 69.2% 773 / 1117
ironclaw_trust 72.88% 661 / 907
ironclaw_capabilities 74.39% 1685 / 2265
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_reborn_event_store 74.67% 958 / 1283
ironclaw_extractors 74.72% 538 / 720
ironclaw_first_party_extensions 77.66% 5400 / 6953
ironclaw_llm 78.31% 20258 / 25870
ironclaw_product_context 78.57% 11 / 14
ironclaw_process_sandbox 80.65% 671 / 832
ironclaw_wasm_product_adapters 80.71% 1448 / 1794
ironclaw_reborn_openai_compat 81.16% 978 / 1205
ironclaw_memory_native 81.22% 3205 / 3946
ironclaw_secrets 82.7% 2791 / 3375
ironclaw_wasm 82.72% 996 / 1204
ironclaw_events 82.86% 1765 / 2130
ironclaw_auth 83.87% 3078 / 3670
ironclaw_reborn_config 84.33% 1814 / 2151
ironclaw_turns 84.38% 13601 / 16118
ironclaw_processes 84.44% 993 / 1176
ironclaw_host_api 84.9% 2608 / 3072
ironclaw_product_workflow 85.19% 10762 / 12633
ironclaw_threads 85.88% 4226 / 4921
ironclaw_projects 85.92% 659 / 767
ironclaw_network 86.12% 670 / 778
ironclaw_common 86.46% 1514 / 1751
ironclaw_slack_v2_adapter 86.79% 1806 / 2081
ironclaw_product_adapters 86.98% 3207 / 3687
ironclaw_reborn_identity 87.03% 557 / 640
ironclaw_skills 87.58% 4470 / 5104
ironclaw_hooks 87.75% 9917 / 11302
ironclaw_product_adapter_registry 88.06% 531 / 603
ironclaw_reborn_traces 88.19% 11946 / 13546
ironclaw_extensions 88.35% 2638 / 2986
ironclaw_host_runtime 88.37% 17066 / 19311
ironclaw_reborn_composition 88.91% 74447 / 83732
ironclaw_approvals 89.24% 1584 / 1775
ironclaw_runner 89.28% 16692 / 18697
ironclaw_conversations 90.33% 3120 / 3454
ironclaw_event_streams 90.82% 1009 / 1111
ironclaw_loop_support 92.49% 14749 / 15946
ironclaw_resources 92.81% 4722 / 5088
ironclaw_attachments 93.06% 630 / 677
ironclaw_reborn_webui_ingress 93.19% 2217 / 2379
ironclaw_telegram_v2_adapter 93.62% 2511 / 2682
ironclaw_agent_loop 94.63% 8811 / 9311
ironclaw_safety 94.8% 3668 / 3869
ironclaw_first_party_extension_ports 95.24% 3343 / 3510
ironclaw_outbound 95.59% 3556 / 3720

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@ilblackdragon

Copy link
Copy Markdown
Member Author

Closing — superseded by already-merged #5692 ("reborn: no run-borking failures — collapsed recoverability stack").

This PR was Phase 1 of a RunFailureReason funnel that would classify terminal run failures into lane/retry/user-message. While it was in flight, #5692 merged that entire architecture onto main: ironclaw_reborn_composition now has FailureLane { Retriable, Explainable, Security }, RetryDisposition { Auto, UserInitiated, NoRetry }, the per-category reborn_failure_summary_for_category table, and an enforcement test (every_failure_category_is_explainable_and_classified) locking every category to a specific explanation + lane.

Continuing would introduce a duplicate FailureLane type (type-placement violation), a duplicate message table, and a duplicate enforcement test. The only delta this funnel offered — classification at the recording boundary with compile-time exhaustiveness instead of #5692's projection-layer, runtime-tested version — is incremental hardening, not new capability, and isn't worth a multi-crate reshape that duplicates just-merged code.

The static no-swallow enforcement that motivated this whole effort is delivered by the merged #5651 (exhaustive capability-error classification) and #5652 (unused_must_use deny).

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5954 — 805cb13b Deployed Jul 10, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants