refactor(errors): static enforcement that failures surface, not swallow - #5651
Conversation
Make the capability-failure classification path fail-closed against the "a new error kind silently aborts the run" class of bug (the keystone finding of docs/plans/2026-06-28-reborn-error-recoverability-audit.md §6.1). Static enforcement: - Drop `#[non_exhaustive]` from `CapabilityFailureKind` and `RuntimeFailureKind`. Both already carry an open-set escape hatch (`Unknown`), so the attribute was redundant belt-and-suspenders whose only effect was to force classifiers to keep a wildcard `_ =>` arm that silently buckets any newly-added *named* variant. - Remove the wildcard arms from every fate-deciding classifier (`capability_error_class`, `capability_failure_kind`, `generic_failure_recovery`, `runtime_failure_kind_to_loop`). They are now exhaustive, so a new named variant fails to compile until it is deliberately classified rather than defaulting into a run-aborting or wrong bucket. Behavior is unchanged for all existing variants. Stop propagating "unknown": - Delete `RuntimeFailureKind::Unknown` (internal, not serialized, and its only production source was a dead chain). Collapse the fail-safe redaction bucket `RuntimeDispatchErrorKind::Unknown` to `RuntimeFailureKind::Internal` in the dispatch->runtime `From` — an uncategorized dispatch error now surfaces as a retryable Internal failure instead of riding an opaque `Unknown` to a dedicated abort. - Keep `RuntimeDispatchErrorKind::Unknown` as the outermost fail-safe redaction category (so redaction never fails closed) and `CapabilityFailureKind::Unknown(String)` as the string-carrying deserialization hatch. Test: `every_capability_failure_kind_has_a_deliberate_recovery_class` locks each variant's recovery class and asserts only genuinely-terminal kinds may reach the run-aborting `Permanent` class (guards against silent re-bucketing, complementing the compile-time exhaustiveness). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
📝 WalkthroughSummary by CodeRabbit
WalkthroughRemoves ChangesExhaustive failure-kind classification
Estimated code review effort: 3 (Moderate) | ~25 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 3 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (3 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Code Review
This pull request removes the #[non_exhaustive] attribute from RuntimeFailureKind and CapabilityFailureKind enums to enforce exhaustive compile-time matching. This ensures that any newly added variants must be explicitly classified rather than silently falling into wildcard fallback arms. Additionally, the RuntimeFailureKind::Unknown variant has been removed, with uncategorized dispatch errors now mapping to RuntimeFailureKind::Internal. Tests and mappings across the codebase have been updated accordingly. I have no further feedback to provide as there are no review comments.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
Coverage ratchetReborn integration-tier coverageLine coverage (Reborn crates): 85.32% — 283704 / 332530 lines Per-crate breakdown (65 crates, lowest-covered first)
This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors. Exemptions (4 entry/entries excluded from the accounting above)
|
|
🚅 Deployed to the ironclaw-pr-5651 environment in ironclaw-ci-preview
|
There was a problem hiding this comment.
Pull request overview
This PR strengthens “no silent swallow” guarantees in the Reborn error pipeline by removing #[non_exhaustive] + wildcard match arms in fate-deciding classifiers, making new error variants a compile-time forcing function for deliberate classification/recovery behavior.
Changes:
- Removed
#[non_exhaustive]and wildcard_ =>arms across failure-kind classifiers to enforce exhaustive matching at compile time. - Collapsed dispatch’s redaction bucket (
RuntimeDispatchErrorKind::Unknown) intoRuntimeFailureKind::Internal, and removed the deadRuntimeFailureKind::Unknownvariant. - Added a classification-locking test to pin
CapabilityFailureKind -> CapabilityErrorClassmappings and detect re-bucketing regressions.
Reviewed changes
Copilot reviewed 6 out of 6 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
| crates/ironclaw_turns/src/run_profile/host.rs | Drops #[non_exhaustive] on CapabilityFailureKind and documents the “Unknown-value” escape hatch to keep classifiers exhaustive. |
| crates/ironclaw_loop_support/src/capability_port.rs | Makes runtime_failure_kind_to_loop exhaustive by removing Unknown/wildcard mapping paths and updates tests accordingly. |
| crates/ironclaw_host_runtime/src/production.rs | Maps RuntimeDispatchErrorKind::Unknown to RuntimeFailureKind::Internal and updates pinning tests for the new behavior. |
| crates/ironclaw_host_runtime/src/lib.rs | Removes #[non_exhaustive] and the Unknown variant from RuntimeFailureKind; keeps the enum match surfaces exhaustive. |
| crates/ironclaw_agent_loop/src/executor/mapping.rs | Removes wildcard classification fallback and adds a regression test that locks recovery class per CapabilityFailureKind variant. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| /// Stable, sanitized failure categories. | ||
| /// | ||
| // Deliberately NOT `#[non_exhaustive]`: the `Unknown` variant is the open-set | ||
| // escape hatch for unrecognized runtime failures, so the attribute would only | ||
| // force classifiers to keep a wildcard arm that silently buckets a new named | ||
| // variant. Without it, disposition/classification matches are exhaustive and a | ||
| // new named variant fails to compile until classified. See | ||
| // `docs/plans/2026-06-28-reborn-error-recoverability-audit.md` §6.1. |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@crates/ironclaw_host_runtime/src/lib.rs`:
- Around line 697-703: Update the comment attached to RuntimeFailureKind so it
no longer references an Unknown variant that was removed; the current rationale
is incorrect and belongs to CapabilityFailureKind. Keep only the valid
explanation for omitting #[non_exhaustive] in this enum, using the
RuntimeFailureKind symbol to locate the block, and ensure the comment reflects
that matches remain exhaustive and new variants fail to compile until
classified. Also consider removing the comment entirely if the remaining
behavior is obvious enough to satisfy the Rust commenting guideline.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 8487d6f4-b705-4267-bbd8-c67dce48cb17
📒 Files selected for processing (6)
crates/ironclaw_agent_loop/src/executor/capability_helpers.rscrates/ironclaw_agent_loop/src/executor/mapping.rscrates/ironclaw_host_runtime/src/lib.rscrates/ironclaw_host_runtime/src/production.rscrates/ironclaw_loop_support/src/capability_port.rscrates/ironclaw_turns/src/run_profile/host.rs
💤 Files with no reviewable changes (2)
- crates/ironclaw_agent_loop/src/executor/capability_helpers.rs
- crates/ironclaw_loop_support/src/capability_port.rs
| /// | ||
| // Deliberately NOT `#[non_exhaustive]`: the `Unknown` variant is the open-set | ||
| // escape hatch for unrecognized runtime failures, so the attribute would only | ||
| // force classifiers to keep a wildcard arm that silently buckets a new named | ||
| // variant. Without it, disposition/classification matches are exhaustive and a | ||
| // new named variant fails to compile until classified. See | ||
| // `docs/plans/2026-06-28-reborn-error-recoverability-audit.md` §6.1. |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
Comment contradicts the enum it documents. RuntimeFailureKind no longer has an Unknown variant (this PR removed it), yet the block asserts "the Unknown variant is the open-set escape hatch for unrecognized runtime failures." That rationale belongs to CapabilityFailureKind, which keeps Unknown(_); here it's just wrong and will mislead the next reader. The valid reason for dropping #[non_exhaustive] is only the "matches stay exhaustive, new variants fail to compile until classified" clause.
📝 Suggested rewrite
-///
-// Deliberately NOT `#[non_exhaustive]`: the `Unknown` variant is the open-set
-// escape hatch for unrecognized runtime failures, so the attribute would only
-// force classifiers to keep a wildcard arm that silently buckets a new named
-// variant. Without it, disposition/classification matches are exhaustive and a
-// new named variant fails to compile until classified. See
-// `docs/plans/2026-06-28-reborn-error-recoverability-audit.md` §6.1.
+///
+// Deliberately NOT `#[non_exhaustive]` and intentionally has no open-set
+// `Unknown` variant: uncategorized dispatch errors collapse to `Internal`.
+// Without the attribute, disposition/classification matches stay exhaustive and
+// a new named variant fails to compile until it is explicitly classified. See
+// `docs/plans/2026-06-28-reborn-error-recoverability-audit.md` §6.1.As per coding guidelines: "Add comments only for non-obvious logic in Rust code."
📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| /// | |
| // Deliberately NOT `#[non_exhaustive]`: the `Unknown` variant is the open-set | |
| // escape hatch for unrecognized runtime failures, so the attribute would only | |
| // force classifiers to keep a wildcard arm that silently buckets a new named | |
| // variant. Without it, disposition/classification matches are exhaustive and a | |
| // new named variant fails to compile until classified. See | |
| // `docs/plans/2026-06-28-reborn-error-recoverability-audit.md` §6.1. | |
| /// | |
| // Deliberately NOT `#[non_exhaustive]` and intentionally has no open-set | |
| // `Unknown` variant: uncategorized dispatch errors collapse to `Internal`. | |
| // Without the attribute, disposition/classification matches stay exhaustive and | |
| // a new named variant fails to compile until it is explicitly classified. See | |
| // `docs/plans/2026-06-28-reborn-error-recoverability-audit.md` §6.1. |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/ironclaw_host_runtime/src/lib.rs` around lines 697 - 703, Update the
comment attached to RuntimeFailureKind so it no longer references an Unknown
variant that was removed; the current rationale is incorrect and belongs to
CapabilityFailureKind. Keep only the valid explanation for omitting
#[non_exhaustive] in this enum, using the RuntimeFailureKind symbol to locate
the block, and ensure the comment reflects that matches remain exhaustive and
new variants fail to compile until classified. Also consider removing the
comment entirely if the remaining behavior is obvious enough to satisfy the Rust
commenting guideline.
Source: Coding guidelines
…error-surfacing # Conflicts: # crates/ironclaw_agent_loop/src/executor/mapping.rs
⏳ IronLoop Review StatusHead: Current reviewers:
Reviewer summaries
Recent activity
Available commands
Run metadataAdmission: webhook accepted the request and IronLoop persisted review state before this projection. |
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (2)
crates/ironclaw_host_runtime/src/lib.rs (1)
727-746: 🩺 Stability & Availability | 🔵 TrivialRemoving the
"unknown"tracing/metric token, not renaming it, is an observability contract change.Per the pinning test comment elsewhere ("Pin the public metric/tracing tokens; renaming any of these is a breaking observability contract change"), any dashboard/alert filtering on
runtime_failure_kind="unknown"will silently stop matching once unclassified dispatch errors report as"internal"instead. Worth a release note / dashboard sweep before this ships.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/ironclaw_host_runtime/src/lib.rs` around lines 727 - 746, The observability token set is changing in `as_str` on the runtime failure kind enum, and dropping the `"unknown"` token is a breaking contract. Keep `"unknown"` available as a stable tracing/metric value (or add an explicit compatibility mapping alongside the new `Internal` variant) so existing dashboards and alerts keep matching, and if the rename is intentional, update any consumers and release notes accordingly.crates/ironclaw_host_runtime/src/production.rs (1)
937-959: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick winSettle the blocked resume record here
resume_spawn_capabilityalready settles every other preflight failure throughfail_matching_blocked_resume_on_preflight_error(...); thisModelInputRejectedearly return skips that path, so theBlockedApprovalrecord for thisapproval_request_idcan remain stale.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/ironclaw_host_runtime/src/production.rs` around lines 937 - 959, The ModelInputRejected branch in resume_spawn_capability is returning early without settling the blocked resume record, leaving the BlockedApproval for approval_request_id stale. Route this failure through the same preflight cleanup path used elsewhere in resume_spawn_capability by calling fail_matching_blocked_resume_on_preflight_error(...) before returning the Failed outcome, so the blocked resume state is always resolved consistently.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@crates/ironclaw_host_runtime/src/lib.rs`:
- Around line 727-746: The observability token set is changing in `as_str` on
the runtime failure kind enum, and dropping the `"unknown"` token is a breaking
contract. Keep `"unknown"` available as a stable tracing/metric value (or add an
explicit compatibility mapping alongside the new `Internal` variant) so existing
dashboards and alerts keep matching, and if the rename is intentional, update
any consumers and release notes accordingly.
In `@crates/ironclaw_host_runtime/src/production.rs`:
- Around line 937-959: The ModelInputRejected branch in resume_spawn_capability
is returning early without settling the blocked resume record, leaving the
BlockedApproval for approval_request_id stale. Route this failure through the
same preflight cleanup path used elsewhere in resume_spawn_capability by calling
fail_matching_blocked_resume_on_preflight_error(...) before returning the Failed
outcome, so the blocked resume state is always resolved consistently.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: a96086ea-8cbd-44cd-b5e8-acaa457a6ee4
📒 Files selected for processing (5)
crates/ironclaw_agent_loop/src/executor/capability_helpers.rscrates/ironclaw_agent_loop/src/executor/mapping.rscrates/ironclaw_host_runtime/src/lib.rscrates/ironclaw_host_runtime/src/production.rscrates/ironclaw_loop_support/src/capability_port.rs
…eshold Reborn Playwright and IronClaw Stress were failing nightly with no alerting at all — same silent-failure class the deep-CI revival fixed, in two more places. Diagnosis of the standing failures: - IronClaw Stress (red every retained scheduled run): the nightly bottleneck suite caps p95 at 1500ms, but its model-tail case injects a synthetic 2.0s model wait by design — the ceiling was structurally unsatisfiable (measured p95 2.05s = the 2.0s wait + ~50ms real work; every case at 0.00% failures). Raised to 2500ms with a comment; the PR-mode variant already ran at 3000ms. Per-case ceilings in ironclaw_stress are the proper follow-up. - Reborn Playwright (flaky-red): last night's failure was test_reborn_legacy_looping_tool_calls_stop_at_low_iteration_boundary asserting failure_category == "driver_protocol_violation" while the runtime now emits "iteration_limit" — already realigned on main by the error-classification refactor (#5651); no change needed here. Alerting changes: - ironclaw-stress.yml + reborn-playwright.yml: schedule-only alert jobs driving .github/scripts/nightly-alert-issue.sh, same contract as the Nightly E2E / Nightly Deep CI alerts. - nightly-watchdog.yml: generalized to a matrix over all four nightlies (Deep CI, E2E, Playwright, Stress) — startup failures and cron-never-fired now alarm for every nightly, not just Deep CI. - nightly-alert-issue.sh: optional Slack mirror. When the SLACK_CI_ALERTS_WEBHOOK_URL repo secret is set (Slack incoming webhook), every failure posts a one-liner with run + issue links and every recovery posts a close-out; absent secret = silent no-op. Delivery is best-effort and never fails the alert job. All five call sites pass the secret through. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…E2E requirable (#5841) * ci: revive the nightly deep tier + make Platform & Compat and Reborn E2E requirable Nightly Deep CI has startup-failed every night since 2026-06-20 with zero jobs: reborn-tests.yml references secrets.SCCACHE_* (sccache-dist) and nightly's call site did not pass secrets, which fails workflow validation at trigger time. The in-run nightly-alert job dies with the run, so nothing reported it. Separately, Platform & Compat's deep jobs (Windows build, bench compile, docker build) were gated on github.event_name == 'workflow_call', which never matches — in a reusable workflow event_name reflects the caller's event — so nightly's "deep reuse" of those jobs silently skipped (all three are 'skipped' in the last nightly run that executed jobs at all). - nightly-deep-ci.yml: pass `secrets: inherit` to the reborn-tests call - platform-and-compat.yml: add a `deep` workflow_call marker input (default true, materializes only under workflow_call) and gate windows-build / wasm-wit-compat / bench-compile / docker-build on it instead of the never-true event_name comparison - nightly-watchdog.yml (new): dead-man's switch that inspects the latest scheduled Nightly Deep CI run from outside it — startup failures and never-fired crons now raise/update the same "Nightly Deep CI failed" issue via .github/scripts/nightly-alert-issue.sh - platform-and-compat.yml: add a stable "Platform & Compat" roll-up job (skip-tolerant, requirable as a status check), delete the vestigial matrix-config job (test_matrix had no consumer; windows_matrix's SLIM branch was unreachable because windows-build never runs on PR or merge_group), and build Docker images in the merge queue when the merge group touches Dockerfile inputs - reborn-e2e.yml: run in the merge queue — merge_group trigger plus a changes job mirroring the pull_request/push paths filters (merge_group does not support paths), with the "Reborn E2E" roll-up reporting on every queue entry so it can become a required check - .github/workflows/README.md (new): the CI tier contract, required-check inventory, deep-tier gotchas, and deliberately accepted gaps Verified with actionlint (no findings beyond pre-existing SC2129 style nits). Reborn E2E queue cost is ~5-9 min based on recent main runs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(nightly): grant caller-side permissions the called workflows require A dispatched validation run of the previous commit still startup-failed: secrets: inherit fixed one violation, but called-workflow jobs that declare job-level permissions beyond the caller's grant also fail validation at trigger time (even when those jobs are event-gated off schedule runs). platform-and-compat's version-check declares pull-requests/issues read; reborn-tests' coverage-report declares pull-requests: write. Grant those supersets at the call sites. Also corrects the incident window in the comments and README: retained history shows zero successful Nightly Deep CI runs since its creation on 2026-05-06 (65 of 74 runs are startup_failures), not merely since 2026-06-20 — the permissions violations date to day one, the secrets one to the 2026-07-03 sccache rollout. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(nightly): alert every nightly, mirror to Slack, fix the stress threshold Reborn Playwright and IronClaw Stress were failing nightly with no alerting at all — same silent-failure class the deep-CI revival fixed, in two more places. Diagnosis of the standing failures: - IronClaw Stress (red every retained scheduled run): the nightly bottleneck suite caps p95 at 1500ms, but its model-tail case injects a synthetic 2.0s model wait by design — the ceiling was structurally unsatisfiable (measured p95 2.05s = the 2.0s wait + ~50ms real work; every case at 0.00% failures). Raised to 2500ms with a comment; the PR-mode variant already ran at 3000ms. Per-case ceilings in ironclaw_stress are the proper follow-up. - Reborn Playwright (flaky-red): last night's failure was test_reborn_legacy_looping_tool_calls_stop_at_low_iteration_boundary asserting failure_category == "driver_protocol_violation" while the runtime now emits "iteration_limit" — already realigned on main by the error-classification refactor (#5651); no change needed here. Alerting changes: - ironclaw-stress.yml + reborn-playwright.yml: schedule-only alert jobs driving .github/scripts/nightly-alert-issue.sh, same contract as the Nightly E2E / Nightly Deep CI alerts. - nightly-watchdog.yml: generalized to a matrix over all four nightlies (Deep CI, E2E, Playwright, Stress) — startup failures and cron-never-fired now alarm for every nightly, not just Deep CI. - nightly-alert-issue.sh: optional Slack mirror. When the SLACK_CI_ALERTS_WEBHOOK_URL repo secret is set (Slack incoming webhook), every failure posts a one-liner with run + issue links and every recovery posts a close-out; absent secret = silent no-op. Delivery is best-effort and never fails the alert job. All five call sites pass the secret through. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(nightly): Slack-only failure alerting through a single watchdog path Per review direction: failures post to Slack, nothing else, one code path, no GitHub issues. - nightly-watchdog.yml is now the only alerting mechanism: at 08:00 UTC it checks each nightly's latest scheduled run (Nightly Deep CI, Nightly E2E, Reborn Playwright, IronClaw Stress) and posts failures — workflow, conclusion, failed job names, run link — to the Slack channel behind the existing secrets.SLACK_WEBHOOK_URL (the same webhook live-canary reports through; no new secret needed). Missing runs, stale runs (>26h, cron never fired), and startup_failures alarm too — the cases an in-run alert job structurally cannot see. A detected failure turns the watchdog matrix job red so its run history doubles as the failure record. Successes post nothing. - Removed the GitHub-issue alerting entirely: nightly-alert-issue.sh and its test harness are deleted, and the in-run nightly-alert jobs are removed from nightly-deep-ci.yml and nightly-e2e.yml along with the round-2 stress/playwright alert jobs. Housekeeping after merge: close the open "Nightly E2E failed" issue (#4108) manually — nothing auto-closes it now. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci: address review — merge-queue docker gate + drop dead scope output - docker-build: the merge_group arm no longer requires has_legacy_tests. The scope classifier matches only the literal `Dockerfile`, so a Dockerfile.reborn/.dockerignore-only merge group reported has_legacy_tests=false and skipped the Docker build exactly when it should run (IronLoop blocking finding). push/deep arms keep the gate. - changes: drop has_engine_replay_risk — no consumer in this workflow, and its workflow_call arm could never fire (github.event_name is the caller's event in reusable workflows). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(nightly): freeze the legacy v1 suite — do not invoke test.yml from nightly Deliberate freeze pending v1 (src/) removal, per team decision: nightly no longer calls the Legacy Tests workflow, leaving test.yml invoked nowhere. Documented loudly in the workflow header and the CI README — including the consequence that test.yml is the only place the root `ironclaw` package's tests run, and that a v1 fix landing before src/ is deleted should temporarily restore the call job. This is an explicit freeze with a paper trail, not the silent-death mode this workflow's history is infamous for; delete test.yml together with src/. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(nightly): remove the v1 browser suite from the nightly fleet Completes the legacy freeze: nightly-e2e.yml (the scheduler for e2e.yml's full v1 browser suite) is deleted and Nightly E2E is dropped from the watchdog matrix. Like Nightly Deep CI before its revival, it had zero successful runs in retained history — its own alert issue notes there is no prior green run on main to attribute against — and its standing failures (v1 /api/chat auth-gate SSE dedup and approval tests) are v1 work the team has frozen pending src/ removal. e2e.yml itself stays (workflow_call/workflow_dispatch), frozen alongside test.yml; both are documented in the CI README to be deleted together with src/. The nightly fleet is now Nightly Deep CI, Reborn Playwright, and IronClaw Stress — all three validated green today — with the watchdog covering exactly those three. After merge: manually close the open "Nightly E2E failed" issue #4108 (the freeze resolves it; nothing auto-closes it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(ci): record the full-path Emulate coverage gap left by the v1 freeze The Emulate-backed full-path tests are the only scheduled coverage for install -> OAuth -> model-routed tool call -> provider mutation, and per tests/e2e/CLAUDE.md they boot the legacy gateway binary, so they froze with v1. Note in the CI README that a Reborn-native port through `ironclaw-reborn serve` is the follow-up that restores this tier — deliberately NOT re-homed into the Reborn nightlies as-is, which would have smuggled the legacy binary back in. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
#5652) A discarded `Result` / `#[must_use]` value is a swallowed error. This promotes the warn-by-default `unused_must_use` to a workspace-wide deny, so a dropped `Result` fails the build instead of silently hiding a failure. Verified zero current fires across `cargo check --workspace --tests`, so this changes no existing code — it only guards against *future* silent drops. Companion swallow-idiom lints are deferred to their own PRs because each needs a real cleanup first (measured on this tree): - clippy::let_underscore_must_use — 67 `let _ = <must_use>` sites (mix of safe discards and genuine swallows, e.g. an ignored async delete); repo treats clippy warnings as errors, so enabling it requires fixing all 67. - clippy::map_err_ignore — ~861 sites; a blanket deny is wrong since error-handling.md permits map_err to a specific typed error and only forbids cause-dropping ones. Part of the error-surfacing enforcement effort (keystone: #5651). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: firat.sertgoz <firat.sertgoz@near.ai>
…_ = drops Convert 90 `let _ = <fallible>` sites — where a `Result`/`#[must_use]` was silently discarded — into one of two explicit forms, across the runtime error paths that matter most (host_runtime, reborn_composition, reborn, hooks, plus filesystem/common/openai_compat stragglers): - Real fallible best-effort ops (temp-file/secret/manifest cleanup, child.kill/.wait, event emit, service teardown, scheduler nudges) now SURFACE their error to a `tracing::debug!` line instead of swallowing it, so a failed cleanup is diagnosable. `debug!` only — never info!/warn!, which corrupt the REPL/TUI. - Genuinely-intentional discards (infallible `write!` to a String, fire-and-forget oneshot sends where a dropped receiver is expected, detached-thread join results, `OnceLock::set` idempotency) keep `let _ =` but carry an explicit `#[allow(clippy::let_underscore_must_use)]` with a one-line reason stating why discarding is correct. This does NOT enable a workspace-wide `let_underscore_must_use = deny`: a fresh full-workspace clippy shows ~998 fires across ~29 crates, the large majority benign intentional discards, so blanket enforcement is disproportionate churn. The static no-swallow *enforcement* is already delivered by the exhaustive-classification keystone (#5651) and the `unused_must_use` deny (#5652); this PR is the targeted surfacing cleanup for the highest-value runtime paths. Verified: touched-crate suites green (2020 tests in the default-feature group; reborn_composition green under its full feature set). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…_ drops (90 sites) (#5662) * refactor(errors): surface best-effort failures instead of silent let _ = drops Convert 90 `let _ = <fallible>` sites — where a `Result`/`#[must_use]` was silently discarded — into one of two explicit forms, across the runtime error paths that matter most (host_runtime, reborn_composition, reborn, hooks, plus filesystem/common/openai_compat stragglers): - Real fallible best-effort ops (temp-file/secret/manifest cleanup, child.kill/.wait, event emit, service teardown, scheduler nudges) now SURFACE their error to a `tracing::debug!` line instead of swallowing it, so a failed cleanup is diagnosable. `debug!` only — never info!/warn!, which corrupt the REPL/TUI. - Genuinely-intentional discards (infallible `write!` to a String, fire-and-forget oneshot sends where a dropped receiver is expected, detached-thread join results, `OnceLock::set` idempotency) keep `let _ =` but carry an explicit `#[allow(clippy::let_underscore_must_use)]` with a one-line reason stating why discarding is correct. This does NOT enable a workspace-wide `let_underscore_must_use = deny`: a fresh full-workspace clippy shows ~998 fires across ~29 crates, the large majority benign intentional discards, so blanket enforcement is disproportionate churn. The static no-swallow *enforcement* is already delivered by the exhaustive-classification keystone (#5651) and the `unused_must_use` deny (#5652); this PR is the targeted surfacing cleanup for the highest-value runtime paths. Verified: touched-crate suites green (2020 tests in the default-feature group; reborn_composition green under its full feature set). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: address cleanup review feedback * fix: address review feedback on best-effort failure surfacing Follow-up to the let-underscore refactor, addressing actionable review threads: - llm_config_service: delete_provider now fails closed. The stored key is deleted *before* the provider definition (delete is idempotent), so a key-cleanup failure returns an error with the definition intact rather than reporting deletion success while orphaning a secret. Regression test added. - extension_lifecycle: lifecycle-disable and manifest rollback failures during activation/install compensation now propagate through the existing compensation_failure path instead of being debug-logged, so a compound failure surfaces the orphaned state (fail loud) rather than silently poisoning future retries. Single-fault and happy paths are unchanged (covered by existing tests). - shell.rs: extracted the duplicated saved-output cleanup-and-log block into a single closure, matching the sibling trace_commons.rs fix. - wasm/runtime.rs: log the epoch-ticker join panic at debug on drop instead of a pure discard, dropping the let_underscore_must_use allow. - turn_scheduler tests: assert the first enqueue succeeds to establish the saturated-queue precondition before matching DeliveryUnavailable. Already addressed on-branch and verified: SecretStore cleanup logs use stable_reason(), purge_secret_handle consolidates the six durable product-auth cleanup sites, and trace_commons.rs has the cleanup_temp helper. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
What & why
Follows the design discussion on #5383 (error-recoverability audit): can we make it a compile-time requirement that every error surfaces to the model/user instead of getting silently swallowed?
Full end-to-end delivery can't be proven by types (Rust has no effect/taint system), but misclassification and swallowing can be made build failures. This PR lands that enforcement, keystone-first.
Committed in this PR so far
Keystone — exhaustive error classification (no silent swallow):
#[non_exhaustive]fromCapabilityFailureKind+RuntimeFailureKind(both already carry anUnknownopen-set hatch, so the attribute was redundant belt-and-suspenders that only forced classifiers to keep a swallowing_ =>arm).capability_error_class,capability_failure_kind,generic_failure_recovery,runtime_failure_kind_to_loop). They're now exhaustive → a new named error variant fails to compile until deliberately classified. Behavior unchanged for existing variants.RuntimeFailureKind::Unknownand collapsed the fail-safe redaction bucketRuntimeDispatchErrorKind::Unknown→RuntimeFailureKind::Internal(surfaces + retryable) instead of riding an opaqueUnknownto a dedicated abort. Kept the outermost fail-safe redaction category and the string-carrying deserialization hatch.every_capability_failure_kind_has_a_deliberate_recovery_classlocking each variant's class and asserting only genuinely-terminal kinds may reach the run-abortingPermanentclass.Validated:
cargo check --workspace --testsclean; touched-crate suites green (host_runtime 337, turns 295, loop_support 191, agent_loop + classification-lock test).Still TODO in this PR (tracked; larger + higher-risk — see PR discussion)
RunFailureReasonsingle funnel with a private constructor at the run boundary (planned_driver/turn_run_executor), so an unsurfaced terminal error becomes structurally unrepresentable. Major run-boundary change.SecurityStoponly if from the safety/leak layer; else Retriable/Explainable with a non-empty user message.unused_must_use/let_underscore_must_use/map_err_ignore. Measured blast radius is large (map_err(|_|~861 sites), so the cleanup is a sweep, not a lint flip; scoping under discussion.🤖 Generated with Claude Code