Skip to content

fix(ci): stop safety ReDoS timing guards flaking under coverage - #6181

Merged
ilblackdragon merged 1 commit into
mainfrom
fix/ci-safety-redos-timing-flake
Jul 17, 2026
Merged

ilblackdragon merged 1 commit into
mainfrom
fix/ci-safety-redos-timing-flake

Conversation

@ilblackdragon

Copy link
Copy Markdown
Member

Problem

The Coverage (default) job on main failed (run 29555194141):

thread 'leak_detector::tests::adversarial::anthropic_key_pattern_100kb_near_miss'
panicked at crates/ironclaw_safety/src/leak_detector.rs:1144:13:
anthropic_api_key pattern took 101ms on 100KB near-miss

This is not a real regression — it's a flaky timing assertion. The *_100kb_near_miss adversarial tests in ironclaw_safety are ReDoS guards: they feed a ~100 KB near-miss payload to a regex scan and assert it finishes fast enough to rule out catastrophic backtracking (which would take seconds or hang). They used a hard 100 ms bound.

Under cargo llvm-cov instrumentation on a shared CI runner the linear scan is ~100× slower, so it landed at 101 ms — 1 ms over. Only the instrumented coverage job is affected; all-features, libsql-only, and e2e all passed.

Fix

Replace the per-test hard thresholds (100 ms in leak_detector/validator/sanitizer, 500 ms already in policy) with one documented shared constant in the crate root:

#[cfg(test)]
pub(crate) const REDOS_SCAN_BUDGET_MS: u128 = 2000;

2 s keeps a large margin below any real ReDoS (seconds→minutes/hang) while tolerating instrumentation and scheduling overhead. The guards still fail loudly on genuine catastrophic backtracking. 32 assertions across the four files now reference the constant; also fixed a stale (threshold: 100ms) message to print the actual budget.

Verification

  • cargo test -p ironclaw_safety --lib adversarial → 100 passed
  • cargo clippy -p ironclaw_safety --all-targets --all-features -- -D warnings → clean

Note

Test-only change; no production behavior is touched, so the commit carries [skip-regression-check].

🤖 Generated with Claude Code

…-regression-check]

The `*_100kb_near_miss` adversarial tests in ironclaw_safety are ReDoS
guards: they feed a ~100 KB near-miss payload to a regex scan and assert
it finishes fast enough to rule out catastrophic backtracking (which
would take seconds or hang). They used a hard 100 ms bound.

Under `cargo llvm-cov` instrumentation on shared CI runners the linear
scan is ~100x slower, so the Coverage (default) job flaked with
"anthropic_api_key pattern took 101ms on 100KB near-miss" — 1 ms over
the threshold. Only the instrumented coverage job is affected.

Replace the per-test hard thresholds (100 ms in leak_detector/validator/
sanitizer, 500 ms already in policy) with one documented shared constant
REDOS_SCAN_BUDGET_MS = 2000 in the crate root. 2 s keeps a large margin
below any real ReDoS while tolerating instrumentation overhead, and the
guards still fail loudly on genuine catastrophic backtracking.

Test-only change; no production behavior touched — hence the
regression-check skip.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@ironloopai

ironloopai Bot commented Jul 17, 2026 •

Copy link
Copy Markdown
Contributor

🔎 IronLoop Review Status

Head: d691ad91781b35965d45164a9f9dbe49cf942344
Result: 1/1 reviewers completed without blocking findings.
Next: Ready for normal human review and CI checks.
Updated: 2026-07-17T06:00:06.299Z

Current reviewers:

Reviewer State Verdict Findings Last update
ironloop/common-reviewer (reviewer) Completed Approved 0 blocking findings / 0 notes 2026-07-17T06:00:06.288Z
Reviewer summaries
Reviewer Detail
ironloop/common-reviewer (reviewer) Approved; 0 blocking findings; Focused test-only change. The shared 2-second budget is correctly gated by cfg(test), referenced by all 32 affected timing assertions, and does not alter production behavior. No c…
Recent activity
Time Reviewer State Detail
2026-07-17T05:58:00.109Z ironloop/common-reviewer (reviewer) Queued Accepted review request for head d691ad9.
2026-07-17T05:58:00.109Z ironloop/common-reviewer (reviewer) Queued Waiting for this reviewer lane to become available.
2026-07-17T05:58:01.034Z ironloop/common-reviewer (reviewer) Started Reviewer worker started.
2026-07-17T05:58:03.671Z ironloop/common-reviewer (reviewer) Workspace ready Prepared isolated checkout (merge_ref) at 221419d.
2026-07-17T06:00:06.288Z ironloop/common-reviewer (reviewer) Result captured Approved; 0 blocking findings.
2026-07-17T06:00:06.288Z ironloop/common-reviewer (reviewer) Completed Review completed and terminal status was persisted.
Available commands
  • @ironloopai help
  • @ironloopai agents
  • @ironloopai review
  • @ironloopai review --agent <agent>
Run metadata

Admission: webhook accepted the request and IronLoop persisted reviewer state before this projection.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6181 July 17, 2026 05:58 Destroyed
@github-actions github-actions Bot added size: M 50-199 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 17, 2026
@coderabbitai

coderabbitai Bot commented Jul 17, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 884b7766-7f5e-4a45-826a-9d3348827cbb

📥 Commits

Reviewing files that changed from the base of the PR and between af0b9de and d691ad9.

📒 Files selected for processing (5)
  • crates/ironclaw_safety/src/leak_detector.rs
  • crates/ironclaw_safety/src/lib.rs
  • crates/ironclaw_safety/src/policy.rs
  • crates/ironclaw_safety/src/sanitizer.rs
  • crates/ironclaw_safety/src/validator.rs

📝 Walkthrough

Summary by CodeRabbit

  • Tests
    • Standardized adversarial performance tests on a shared, configurable timing budget.
    • Improved consistency of large-input and near-miss safety checks across validation, sanitization, policy, and leak detection.
    • Updated timeout failure messages to include the configured budget.

Walkthrough

The safety crate adds a documented test-only scan budget and replaces hardcoded timing thresholds in adversarial performance tests across leak detection, policy, sanitizer, and validator modules.

Changes

ReDoS timing budget

Layer / File(s) Summary
Define shared scan budget
crates/ironclaw_safety/src/lib.rs
Adds the documented REDOS_SCAN_BUDGET_MS test constant.
Apply budget to adversarial tests
crates/ironclaw_safety/src/leak_detector.rs, crates/ironclaw_safety/src/policy.rs, crates/ironclaw_safety/src/sanitizer.rs, crates/ironclaw_safety/src/validator.rs
Replaces fixed 100 ms and 500 ms timing assertions with the shared budget and updates one failure message to report it.

Estimated code review effort: 2 (Simple) | ~10 minutes

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description explains the bug, fix, and verification, but it does not follow the repository template or include several required sections. Rewrite the PR description to match the template, including Summary, Change Type, Linked Issue, Validation, Security Impact, and the remaining required sections.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title uses Conventional Commits style and accurately summarizes the test-only fix for flaky safety ReDoS timing guards.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a centralized wall-clock budget constant, REDOS_SCAN_BUDGET_MS (set to 2000 ms), for ReDoS and catastrophic-backtracking regression tests within the ironclaw_safety crate. It updates various test assertions across leak_detector.rs, policy.rs, sanitizer.rs, and validator.rs to use this new constant instead of hardcoded thresholds (like 100 ms or 500 ms), preventing test flakiness under instrumentation or heavy CI runner load. There are no review comments to address, and I have no additional feedback to provide.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
✅ Approved 0 0 0 d691ad91781b

Head: d691ad91781b35965d45164a9f9dbe49cf942344
Next: No reviewer action needed.

Run details

Status: Current
Needs human: no
Needs validation: no

Summary

Focused test-only change. The shared 2-second budget is correctly gated by cfg(test), referenced by all 32 affected timing assertions, and does not alter production behavior. No concrete correctness, security, or maintainability issues found.

Findings

None.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.

@railway-app

railway-app Bot commented Jul 17, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6181 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 17, 2026 at 6:09 am

@github-actions

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.69% (305198 / 356184 lines)
  floor:    85.3% (tolerance 0.5pp -> effective floor 84.8%)
  denominator: 356184 lines now vs 320188 at floor capture (+35996 lines, +11.24%) — material change (>5%)

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.69% — 305198 / 356184 lines

Per-crate breakdown (63 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_runtime_policy 31.75% 80 / 252
ironclaw_event_projections 43.31% 673 / 1554
ironclaw_run_state 53.07% 225 / 424
ironclaw_authorization 53.89% 464 / 861
ironclaw_observability 61.54% 16 / 26
ironclaw_webui_v2 62.98% 2684 / 4262
ironclaw_mcp 64.89% 595 / 917
ironclaw_triggers 65.44% 2142 / 3273
ironclaw_reborn_cli 66.18% 4488 / 6781
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_filesystem 67.69% 3932 / 5809
ironclaw_memory 69.2% 773 / 1117
ironclaw_reborn_migration 71.57% 1551 / 2167
ironclaw_trust 72.88% 661 / 907
ironclaw_capabilities 74.39% 1685 / 2265
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_reborn_event_store 74.67% 958 / 1283
ironclaw_extractors 74.72% 538 / 720
ironclaw_llm 78.25% 20193 / 25805
ironclaw_product_context 78.57% 11 / 14
ironclaw_first_party_extensions 78.86% 5579 / 7075
ironclaw_process_sandbox 80.65% 671 / 832
ironclaw_wasm_product_adapters 80.71% 1448 / 1794
ironclaw_memory_native 81.22% 3205 / 3946
ironclaw_secrets 82.79% 2794 / 3375
ironclaw_events 82.86% 1765 / 2130
ironclaw_reborn_identity 83.59% 433 / 518
ironclaw_wasm 83.97% 1011 / 1204
ironclaw_reborn_config 84.06% 1814 / 2158
ironclaw_processes 84.44% 993 / 1176
ironclaw_auth 84.79% 3312 / 3906
ironclaw_turns 85.04% 13722 / 16136
ironclaw_host_api 85.2% 2665 / 3128
ironclaw_product_workflow 85.79% 10917 / 12725
ironclaw_projects 85.92% 659 / 767
ironclaw_network 86.12% 670 / 778
ironclaw_common 86.66% 1741 / 2009
ironclaw_slack_v2_adapter 86.79% 1806 / 2081
ironclaw_threads 86.93% 4708 / 5416
ironclaw_product_adapters 87.18% 3265 / 3745
ironclaw_skills 87.58% 4470 / 5104
ironclaw_hooks 87.78% 9921 / 11302
ironclaw_product_adapter_registry 88.06% 531 / 603
ironclaw_reborn_traces 88.19% 11946 / 13546
ironclaw_host_runtime 88.53% 17437 / 19697
ironclaw_extensions 89.38% 2971 / 3324
ironclaw_reborn_composition 89.4% 80727 / 90297
ironclaw_approvals 89.41% 1587 / 1775
ironclaw_runner 89.5% 16990 / 18983
ironclaw_reborn_openai_compat 89.5% 3778 / 4221
ironclaw_conversations 90.39% 3123 / 3455
ironclaw_event_streams 90.82% 1009 / 1111
ironclaw_loop_host 92.24% 15043 / 16308
ironclaw_resources 92.69% 5134 / 5539
ironclaw_attachments 93.06% 630 / 677
ironclaw_reborn_webui_ingress 93.19% 2217 / 2379
ironclaw_telegram_v2_adapter 93.91% 2592 / 2760
ironclaw_agent_loop 94.81% 9201 / 9705
ironclaw_safety 95.04% 3677 / 3869
ironclaw_outbound 95.59% 3556 / 3720
ironclaw_first_party_extension_ports 95.62% 3672 / 3840

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@ilblackdragon
ilblackdragon merged commit 7ed0313 into main Jul 17, 2026
64 checks passed
@ilblackdragon
ilblackdragon deleted the fix/ci-safety-redos-timing-flake branch July 17, 2026 06:48
@ironclaw-ci ironclaw-ci Bot mentioned this pull request Jul 17, 2026
ilblackdragon added a commit that referenced this pull request Jul 17, 2026
Reconciles the 6 new main commits (serve-webui-at-root #6152, workspace
download failures #6150, theme controls #6148, toast lifecycle #6151, serve
durability e2e #5523, safety ReDoS CI #6181) with the WebUI host-stack merge +
`ironclaw_webui` rename.

Conflict resolutions:
- Frontend (app/auth/sidebar/toast/theme/i18n + 3 new tests): rename detection
  auto-applied main's edits to the moved `crates/ironclaw_webui/frontend/` path.
- `webui_serve.rs`: took main's #6152 root-serving surface —
  `static_router_with_config(static_router_config)` replaces the old
  `mount_at_prefix("/v2")`, plus the new `WebuiServeError` root-namespace
  variants and `static_router_config_from_descriptors` validator — adapted to
  this crate's module paths (`crate::webui_v2::`, `crate::webui_rate_limit::`).
- `webui_v2/mod.rs`: took main's static-router exports (`StaticRouterConfig`,
  `StaticRouterConfigError`, `static_router_with_config`; dropped
  `mount_at_prefix`) without the `webui-v2-beta` cfg gate, since the folded
  module is unconditional here.
- Deleted `crates/ironclaw_webui_v2/{CLAUDE.md,Cargo.toml}` (main modified the
  now-removed crate); dropped composition's optional `ironclaw_webui_v2` dep.
- `composition/tests/webui_v2_serve.rs`: imports of the moved
  `Webui*`/`webui_v2_app` (incl. main's newly-used `WebuiServeError`) now come
  from `ironclaw_webui`.
- Cargo.lock regenerated.

Verified: `ironclaw_webui` builds + full test suite green (175 lib + integration,
0 failures) under default and `--all-features`; `reborn_cli` builds with
`slack-v2-host-beta,openai-compat-beta` (wiring survived the merge).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
BenKurrek added a commit that referenced this pull request Jul 18, 2026
13-commit window af0b9de..c1a9807 re-expressed onto the unified generic
extension architecture. Dominant event: #6194's webui consolidation
(ingress + webui_v2 + composition webui_serve -> single ironclaw_webui)
adopted with the retired Slack host-beta lane excised; #6173 runtime.rs
test extraction adopted with this branch's test module landing in
runtime/tests/core.rs; #6195/#6197 §4.3 filesystem-backed store
refactors adopted; #5978 run_id threaded through this branch's
dispatcher/engine construction sites; #6172/#5523/#6152/#6148/#6150/
#6151/#6181 present-verbatim. Full dispositions in the PR-body fifth-fold
ledger.

Verified locally: workspace check + clippy -D warnings (all targets, all
features) green; architecture 43/0; composition CI-bucket 1533/0;
product_workflow 633/0; runner 673/0 (CI features); cli 432/0 (CI
features); webui 378/0; loop_host 492/0; host_runtime 344/0 + 42
sandbox tests with Docker; auth 128/0; dispatcher/capabilities/
extension_host/authorization/approvals all green; frontend tsc + vitest
773/0; oauth_connect integration 21/0 with colima up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
theredspoon added a commit to theredspoon/ironclaw that referenced this pull request Jul 18, 2026
* chore(ci): dev metrics + composition mass ratchet gate (#6167)

* chore(ci): dev metrics + composition mass ratchet gate

Adds a three-tier development-metrics tool and a guardrail that stops the
ironclaw_reborn_composition crate from accreting more of the codebase.

scripts/dev_metrics.py — three tiers from git + GitHub + working tree:
  - Tier 1 flow/speed: PR lead time, size distribution, merge cadence
  - Tier 2 quality/stability: change-failure proxy, rework, test share
  - Tier 3 codebase health: composition mass, v1 src burndown, file sprawl,
    abstraction density, boundary-test coverage

Composition mass ratchet — the dependency-boundary tests police edges
*between* crates but are blind to mass piling up *inside* one crate.
ironclaw_reborn_composition is charter-bound to assembly-only wiring yet is
now ~26.7% of all production crate code. This gate is that missing guard:
  - scripts/ci/composition-budget.toml — committed ceiling (enforce +
    tolerance), modeled on the existing coverage-floor ratchet
  - scripts/ci/check-composition-budget.sh — pure-bash gate; one-directional
    (fails only on growth past the ceiling), emits a down-ratchet nudge as
    carve-outs free up slack
  - scripts/ci/test-check-composition-budget.sh — 22 assertions / 10 fixture
    cases incl. a guard that the real tree passes the committed budget

Wiring:
  - CI: new composition-budget job in code_style.yml (runs the gate + self-
    tests it, registered in the aggregating code-style gate)
  - Local: pre-commit-safety.sh runs the gate when composition or the gate
    itself is staged; dev-setup.sh install message updated

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ci): address review — production-only metric, script hardening, dev-metrics tests

Review feedback on #6167 (gemini, ironloopai, coderabbit):

Blocking — gate counted test-only code despite its documented "tests
excluded" contract. Exclude test-only FILES (tests.rs/test_*.rs/*_tests.rs
and /tests/ dirs) from both numerator and denominator; rebaseline the
ceiling 2670 -> 2398 bp (26.70% -> 23.98%). Inline #[cfg(test)] modules
remain a documented, symmetric residual (a line-counter can't parse them).
Added a regression case proving test files are excluded.

check-composition-budget.sh: toml_get no longer aborts under set -e +
pipefail when a key is missing (|| true) so schema validation is reached;
added a missing-key regression case.

test-check-composition-budget.sh: set -euo pipefail (repo invariant);
SIGPIPE-safe capture + fixture generation; pure-bash asserts (no pipes).

dev_metrics.py: bound `gh` with a 30s timeout and treat non-JSON output as
unavailable; fix the trait-impl density regex to count `impl<T> ... for`
generics; harden find/grep/wc probes with pipefail + rc checks (no more
false-zero metrics); UTF-8 file writes; surface the gate-aligned production
share as the ratchet metric and relabel the byte-based trend as a distinct,
coarser measurement; extract a pure classify_commit helper.

New scripts/test_dev_metrics.py — caller-level unit tests for
classification, percentiles, change-failure bucketing, rendering, and the
test-file/impl regexes; wired into the composition-budget CI job.

pre-commit hook: trigger on any staged crates/**.rs change (the metric is a
ratio) and document the working-tree/CI-authoritative limitation.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ci): harden PR classifier against transient GitHub API flakes

The classify job (#6167 CI) failed with `invalid character '<' looking
for beginning of value`: a transient API error returned an HTML page,
`gh --jq` aborted, and under `set -e` the whole labels-only job failed
and blocked the PR.

pr-labeler.sh now:
- routes every gh call through a `gh_retry` wrapper (retry + linear
  backoff), and
- treats each classifier as best-effort — a step that still can't fetch
  after retries only emits a `::warning::` and the script exits 0, so
  labeling never gates a merge.

Two bash traps fixed along the way, both caught by the new test:
- a bare `if cmd; then …; fi` resets `$?` to 0 after `fi`, so gh_retry's
  give-up looked like success — capture rc in the `else`;
- `set -e` is suppressed inside a function on the left of `||`, so the
  classifiers check their own fetches explicitly instead of relying on
  errexit.

Regression test: .github/scripts/test-pr-labeler.sh (retry/backoff,
give-up, and end-to-end non-fatal + happy-path via a fake `gh`), wired
into the code_style "Static-check self-tests" step and the has_code
path filter so it runs when the labeler or its test changes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(ci): add dispatch (Arc<dyn>) ratchet + dev-metrics dispatch signals

Companion to the mass ratchet for the "reduce traits & dispatch" goal
(#6168 / runtime-decomposition plan #4471).

check-composition-budget.sh now enforces TWO metrics: composition's share of
production crate code (existing) AND its Arc<dyn> dispatch count. The dispatch
count is scoped to composition production files EXCLUDING src/slack and
src/extension_host — those are owned by the separate channel/extension
refactor, so this gate must not govern or trip on their work. One-directional
like the mass ratchet: only trips on growth; nudges when slack accrues.

composition-budget.toml: arc_dyn_ceiling = 1093 (current governed count),
tolerance 15.

test harness: +6 dispatch cases (within / breach / dry-run / slack+extension
exclusion / missing-key schema error); budget() helper carries the dispatch
keys; count_arc_dyn tolerates no-match under set -e + pipefail. 36 cases pass.

dev_metrics.py: Tier-3 reports governed Arc<dyn> count and distinct dyn-trait
count (the dispatch-breadth trend), matching the ratchet scope.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn-cli): background service install (launchd/systemd) + service restart (#6172)

* feat(reborn-cli): background service install (launchd/systemd) for ironclaw-reborn

Extracted from #6157 (service half only; TUI stays parked there). Adds
`service install/uninstall/start/stop/status`, the serve-invocation
plist/unit contract (IRONCLAW_REBORN_HOME only, no secrets), and the
`full` feature bundle with libsql as default storage.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(reborn-cli): add `service restart` verb

Composes the existing stop+start through the shared ServiceCommandRunner
dispatch. Stopped service starts cleanly; uninstalled service errors with
guidance; a failed start after a successful stop reports the service as
stopped rather than half-restarted.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(reborn-cli): pin service-PR surface invariants

Extend help_mentions_reborn_commands to assert `service` is listed
under webui-v2-beta and that no `tui` subcommand exists; add
service_help_lists_all_verbs pinning the six service verbs
(install/start/stop/status/restart/uninstall).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(reborn-cli): dedupe service restart, normalize status output

Extracts the shared restart decision tree into restart_generic (fn-pointer
seam; platforms keep only their own detection), normalizes `service status`
to one running/stopped/not-installed vocabulary on both platforms, hoists
write_atomic so launchd plist writes are crash-safe, and fixes a stale
verb-count doc.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn-cli): preserve raw systemd ActiveState as a status detail line

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn-cli): harden service module per PR review (PID parse, systemctl parsing, perms, reload, orphan status, rollback, preflight)

Addresses 9 verified findings from PR #6172 review:
1. launchd service_running misread `-` (loaded-but-stopped) PID as
   running; mirrors operator_service_lifecycle's launchd_status_from_line
   shape, split into service_running (has PID) vs service_loaded (any
   status) since uninstall/install genuinely need the latter.
2. systemd query_unit_state now parses Key=Value lines (order-independent,
   no --value) and errors on a missing required key instead of
   unwrap_or_default(), which silently read as enabled=false.
3. write_atomic sets 0600 on unix before create_new, matching
   operator_service_lifecycle's write_service_file.
4. launchd install now unloads/reloads a currently-loaded job after
   rewriting the plist, so a reinstall actually picks up the new
   ProgramArguments/EnvironmentVariables.
5. status now queries the manager unconditionally on both platforms so
   an orphaned unit (file removed out-of-band, still loaded/enabled)
   reports installed; two tests that pinned the old skip-when-absent
   behavior were pinning the orphan-hiding bug and are updated/renamed.
6. systemd uninstall's remove_file failure now routes through the same
   rollback path (restore file + reload + re-enable) as a daemon-reload
   failure, via an extracted rollback_uninstall helper.
7. preflight_warnings gained a webui_token_file_is_valid check; doc
   comment corrected to describe what's actually checked.
8. Documented (not built) that launchd's StandardOutPath/StandardErrorPath
   logs are unrotated, in both a code comment and `service install --help`.
9. Cargo.toml `full` feature now includes root-llm-provider so
   `--no-default-features --features full` stays self-contained.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(reborn-cli): adopt the canonical service identity (com.ironclaw.reborn)

The CLI service surface and the WebUI operator facade
(RebornLocalServiceLifecycle) now deliberately share one unit name and
launchd label. CLI installs atomically replace a facade-installed unit —
a security improvement, since the facade bakes the WebUI token into the
unit file while the CLI unit is secret-free. Consolidating the two
implementations is a documented follow-up.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn-cli): CI mock fidelity, token-file hygiene, review fixes

1. mod.rs: add the "systemctl show unit state" arm to the shared
   SuccessfulServiceCommandRunner mock (install_with_runner now queries
   unit state pre-write). Production code and the strict parser are
   correct; only the mock modeled reality incompletely.
2. webui_token.rs: propagate real I/O errors instead of treating them
   as "absent" (was silently overwriting unreadable tokens); reject
   symlinked/oversized token files; repair (not reject) a wrongly
   permissioned but valid token on accept; serve.rs no longer collapses
   VarError::NotUnicode into "unset".
3. launchd.rs/mod.rs: suppress the "keeps the OLD definition" advisory
   when install already reloaded a loaded job in place (the definition
   is live immediately in that case); systemd's advisory is unaffected.
4. mod.rs: preflight-warning coverage now drives service install
   (runner-injectable, warnings returned) instead of only unit-testing
   the preflight_warnings helper directly.
5. systemd.rs: uninstall's remove_file step is now injectable so its
   rollback test forces a deterministic failure, replacing the
   chmod-0o555 approach that silently no-ops under a root test runner.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn-cli): preserve error source chain in service rollback failures

combined_failure flattened the primary error and rollback outcomes into
one anyhow!() string, losing the source chain. Use .context() so the
primary stays inspectable via source()/{:#} beneath the rollback text.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn-cli): single-handle token read closes TOCTOU window

read_token_file_checked previously stat'd then read the token file as
two separate syscalls, letting a symlink/FIFO/oversized file be swapped
in between them; a FIFO also passed the length check and could block
serve startup indefinitely. Now opens once with O_NOFOLLOW|O_NONBLOCK,
checks type/size via fstat on that handle, and bounds the read to
MAX_BYTES+1 from the same fd. Non-unix keeps the prior best-effort path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn-cli): uninstall disable rollback, hermetic verb dispatch, dry-run coverage, docs

Four verified findings from PR #6172 review round: (1) systemd uninstall's
disable failure now rolls back like every sibling failure branch instead of
propagating with a bare `?`; (2) a smoke test pins that a directory at the
webui-token path fails `onboard --dry-run` non-zero without mutating home;
(3) start/stop/restart/status/uninstall get the same runner-injectable split
`install` already had (`ServicePlatform::*_with_runner`), with one
consolidated clap-dispatch test instead of duplicating per-verb coverage;
(4) FEATURE_PARITY.md and CHANGELOG.md reflect the shipped service-install
feature.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn-cli): single status line per service restart

restart_generic's stop/start fn pointers called the loud
start_with_runner/stop_with_runner, which each print their own
"Service started"/"Service stopped" line in addition to
restart_generic's own summary line, so `service restart` printed two
lines. Add quiet variants (start_with_runner_quiet/
stop_with_runner_quiet) that skip the println, used only by
restart_with_runner; the public start/stop commands keep printing as
before.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn-cli): stop/restart honor manager-loaded state like status/uninstall

launchd `stop` gated on service_running alone, leaving a loaded-but-not-
running KeepAlive job (`-` PID) registered for respawn; `restart` derived
`was_running` the same way, so it bare-loaded an already-loaded label
(which launchd errors on) instead of reloading. systemd `stop` gated only
on unit-file existence, silently no-opping on a unit removed out-of-band
while still loaded/enabled. All three now check manager state the way
status/uninstall already do.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn-cli): honor XDG_CONFIG_HOME for systemd units, guard launchd start on loaded labels

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(reborn-cli): make service tests hermetic over XDG_CONFIG_HOME

commit 5b3f39eb9 made config_home() honor $XDG_CONFIG_HOME, which
unit_path() now reads. Service tests that fake $HOME into a tempdir
never cleared XDG_CONFIG_HOME, so on hosts where it's set (CI runners
observed setting it to $HOME/.config), unit_path() resolved to the
real path instead of the tempdir — causing
systemd::tests::restart_not_installed_errors_with_install_guidance
and
tests::install_then_uninstall_linux_writes_and_removes_unit_file to
fail. Production config_home() behavior is unchanged and correct.

Extends the TempHomeGuard helpers in mod.rs and systemd.rs (new,
mirroring mod.rs's) to also clear/restore XDG_CONFIG_HOME, and
switches all HOME-faking tests in systemd.rs onto the guard instead of
manual set/restore blocks.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* fix(ci): stop safety ReDoS timing guards flaking under coverage [skip-regression-check] (#6181)

The `*_100kb_near_miss` adversarial tests in ironclaw_safety are ReDoS
guards: they feed a ~100 KB near-miss payload to a regex scan and assert
it finishes fast enough to rule out catastrophic backtracking (which
would take seconds or hang). They used a hard 100 ms bound.

Under `cargo llvm-cov` instrumentation on shared CI runners the linear
scan is ~100x slower, so the Coverage (default) job flaked with
"anthropic_api_key pattern took 101ms on 100KB near-miss" — 1 ms over
the threshold. Only the instrumented coverage job is affected.

Replace the per-test hard thresholds (100 ms in leak_detector/validator/
sanitizer, 500 ms already in policy) with one documented shared constant
REDOS_SCAN_BUDGET_MS = 2000 in the crate root. 2 s keeps a large margin
below any real ReDoS while tolerating instrumentation overhead, and the
guards still fail loudly on genuine catastrophic backtracking.

Test-only change; no production behavior touched — hence the
regression-check skip.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(e2e): black-box smoke for ironclaw-reborn serve — restart + kill-9 durability (#5523)

The in-process Reborn integration harness cannot prove real process
startup, real HTTP end-to-end, or process-death durability — its
new_at_path() reopen approximates a restart but never actually kills a
process. Add a thin, permanent black-box smoke suite that boots the
real ironclaw-reborn binary and drives it purely over HTTP:

- boot -> /api/health -> scripted chat round-trip
- tool-call turn executes and finalizes a reply
- graceful restart (SIGINT) preserves thread history
- kill -9 durability: on-disk libsql state survives an unclean death,
  server comes back healthy, no leaked child processes
- bearer-auth boundary (401 without token, 200 with)

Fixture design: reuses the existing reborn_v2_restartable_server
fixture (already restart-capable against a persistent home dir) rather
than porting the legacy ManagedIronclawServer class. Extends its
stop() closure with a `hard: bool = False` flag for SIGKILL, so the
fixture's tuple shape and existing consumer are untouched. Promotes
the capability-preview polling helpers out of
test_reborn_webui_v2_legacy_tool_execution.py into the shared harness
(now used by both files) instead of adding a second copy for the new
suite.

Mutation-verified the durability scenario: temporarily pointed
restart() at a fresh home dir per call, confirmed the kill-9 test goes
red for the right reason (persisted thread missing after "restart"),
then reverted to a clean diff.

Wires a `blackbox-smoke` CI job into the existing reborn-e2e.yml
job-per-file pattern (mirrors webui-v2-smoke's build step, no
Playwright/Node needed since this suite is HTTP-only).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* fix(webui-v2): improve toast lifecycle and accessibility (#6151)

* fix(webui): improve toast lifecycle and accessibility

* test(e2e): cover toast lifecycle and stacking

* fix(webui): type toast presentation mappings

* test(e2e): fast-forward toast hover timing

* test(e2e): align toast clock requirements

* fix(webui-v2): add theme selection controls to Appearance settings (#6148)

* fix(webui): add theme controls to appearance settings

* test(e2e): cover appearance theme persistence

* fix(webui): address appearance accessibility review

* fix(webui): type appearance theme controls

* fix(webui): use native theme radios

* feat(reborn): serve webui at root path instead of `v2` (#6152)

* feat(reborn): serve the WebUI from root paths

* test(e2e): cover root-mounted Reborn WebUI

* fix(webui): reject noncanonical SPA paths

* fix(webui): address root-mount review feedback

* refactor(webui): derive static router config errors

* fix(webui-v2): surface workspace download failures (#6150)

* fix(webui-v2): surface workspace download failures

* test(e2e): cover workspace download failure feedback

* test(e2e): centralize workspace download selectors

* test(e2e): navigate workspace downloads through UI

* docs(reborn): propose architecture simplification — fewer DTOs, less dyn, no local-specific structs (#6175)

* docs(reborn): propose architecture simplification — fewer DTOs, less dyn, no local-specific structs

Design note proposing a fundamental simplification of the Reborn host/runtime
internals, grounded in a code audit and a cross-reference against the last ~30
days of PRs/issues.

Thesis: DTO proliferation (~14 mirror structs per capability call), dyn
proliferation (~6 hot-path trait objects, most single-impl), and local-specific
store structs are three symptoms of one decision — treating every crate boundary
as a trust boundary when Reborn has exactly one (loop <-> host).

Proposes: one canonical payload type in ironclaw_host_api (Invocation/Authority/
Outcome), authority as a single fold, a closed RuntimeLane enum instead of dyn
RuntimeAdapter, backend-generic stores (RowBackend) to delete the InMemory*/
Filesystem* tree, and DeploymentConfig-as-data instead of composition-mode
struct families. Preserves all security invariants; incremental migration with
a first-party-lane proof-of-concept slice. Complements #6168.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): show the field-level "why" behind the ~14 re-wraps

Fold the mechanistic root cause into §1.1: a hop-by-hop field diff of the five
request types, showing only three are genuinely distinct states and the other
two are duplication forced by the crate DAG plus dead transitional fields
(trust_decision is ignored by DefaultHostRuntime; idempotency_key is
unimplemented). Names the four mechanisms and quantifies the ~40% that is pure
duplication.

Add §3.1 mapping the five request types onto the three real states
(Invocation -> +Authority -> +resolved handles), showing how each mechanism is
eliminated or made explicit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): show impl/store-level "why" for the dyn and stores sections

§1.3: replace the flat dyn table with prod-vs-test-double counts and storage
(Arc<dyn>), and name the three mechanisms — trait-as-test-seam (HostRuntime: 1
prod + 6 doubles; CapabilityDispatcher 1 + 2), speculative replaceability, and
generic-and-dyn double indirection. Correct a mislabel: CapabilityHost is a
concrete generic struct, not a trait; the dyn on that path is the dispatcher it
holds. RuntimeAdapter = 5 impls (4 lanes + resolver), a closed set.

§1.4: quantify the per-backend, per-domain store duplication with LOC (turns
~4,260 in-memory vs ~1,710 filesystem) and a domain table (turns/processes/
approvals/authorization/run_state), and name the two mechanisms — logic welded
to storage, multiplied by the TurnRun/processes lifecycle split.

§4.2: drop CapabilityHost from the "delete trait" list (already concrete) and
route the test-seam need through generics/one boundary fake.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): OS mechanism/policy framing — kill in-memory stores and Local*

Fold in five directives:

- §2.1: the operating-system lens — kernel = mechanism (small, frozen, feature-
  agnostic vocabulary + a few real seams); everything that varies by feature or
  deployment is policy resolved to data at the edge. Adding a feature must not
  change the kernel.
- §4.3 (rewritten): the storage seam already exists — RootFilesystem, with a
  first-class InMemoryBackend. Delete every hand-written InMemory*Store; tests use
  FilesystemXStore<InMemoryBackend>. No RowBackend to invent.
- §4.4 (new): local-dev is a policy config (a DeploymentConfig value), NOT an
  implementation. The ~66-identifier LocalDev* shadow runtime across 42 composition
  files collapses to one config literal selecting shared substrates. Rename the two
  genuine resource types (LocalFilesystem->DiskFilesystem, LocalHostProcessPort->
  HostProcessPort). Enforce with a no-"Local*"-type-names boundary test.
- §4.5 (new): enumerate and freeze the kernel boundary — host_api's ~124 types +
  the ~13 AgentLoopDriverHost ports + the mediators. Move runtime_policy (mode
  enums) out of the vocabulary; freeze the neutral authority survivors by test.
- §5/§7/refs updated: before/after rows for in-memory stores and Local*; migration
  resequenced by risk (delete-in-memory and Local*->config are the low-risk first
  slices); new evidence pointers.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): audit result — does DeploymentConfig express everything?

Four-cluster audit of the ~40 LocalDev* types (policies, stores, capability
wiring, trust/evidence) against "config not code". Adds §4.4.1 with the verified
verdict: DeploymentConfig expresses every local-dev *selection*, but the LocalDev*
family is three things, and zero-LocalDev is three moves not one —

1. Already config: the capability policy is literally a TOML file; stores are the
   same shared types prod uses (production_turn_state_store<F> called by both),
   backend-selected; LocalDevOverride trust seam is inert.
2. Mis-prefixed shared substrate (gate-evidence readers, lease-terms provider,
   auth read-model, capability IO): genuine code but not local — de-prefix and
   share, not configify.
3. Genuine local-only mechanism (capability-port decorator stack: synthetic tools,
   surface disclosure, mid-run refresh): behavior stays code, but config-GATED
   shared middleware, not a LocalDev* factory. Synthetic product-ops
   (project_create/skill_activate/result_read/outbound_delivery) should become
   first-party capabilities on the normal lane.

Security note: no trust/approval bypass found — override inert, provider trust
only user_trusted, gate-evidence readers fail closed. §8 Q3 answered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): add target-structure/interfaces section + shell-escape case study

§5 (new) — Target structure: the minimal kernel and clean interfaces. Component
table (kernel = authority/recovery; substrates = mechanism behind ports; loops
and products = replaceable userland). Interface sketches: the one generic
ProductSurface (open/submit_turn/events/reply/resolve_gate/cancel — feature-
agnostic), the kernel authorize/dispatch, the AgentLoopHost trust membrane, the
substrate ports (RootFilesystem, ProcessSandbox with scope-only SandboxMount,
SecretBroker, NetworkPolicy), and DeploymentConfig-as-data. Structure diagram.

§6 (new) — Case study: the shell cross-tenant escape (#6170). Verified root cause
(shell is a real OS subprocess the virtual FS doesn't bound; unsafe host port is
the default; HostedSingleTenant -> LocalSingleUser -> LocalHost) and how the §5
structure makes it structurally impossible (ProcessSandbox as the only path,
unconstructible host port, mode-from-fact, fail-closed, two-user containment test).

Renumbered subsequent sections (7 before/after, 8 invariants, 9 migration,
10 open questions); added #6170 to references.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(rules): add process/shell tenant-isolation invariant; point architecture rule at the plan

safety-and-sandbox.md: new "Process and shell execution: real OS isolation, per
tenant" section — the standing invariant issue #6170 violated. Codifies that the
virtual ScopedFilesystem does not contain a subprocess; multi-user/served
deployments must route process spawns through TenantSandboxProcessPort with a
scope-derived mount (never LocalHostProcessPort); deployment mode must reflect
multi-user serving; fail closed (no sandbox => no shell, never host shell); and
requires a two-user cross-tenant escape test for changes to process ports,
planner backend rules, or the profile->mode mapping.

architecture.md: add a "Direction" pointer to the simplification design doc as
the owning plan for the DTO/dyn/InMemory*/LocalDev* debt the smells describe, and
cross-reference type-placement.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): §5.8 — products are adapters over ProductSurface, collapse composition-split surfaces

Every product owns its whole host side (protocol + transport + identity) as one
adapter consuming the kernel ProductSurface; composition holds no product/transport
code. Quantifies the split: WebUI across 4 places (webui_v2 + webui_ingress +
static + composition/webui), and ~108K LOC of product code in the god-crate
(slack 40.6K, product_auth 32.7K, runtime 14.5K, llm_admin 8.5K, automation 6.1K,
webui 4.6K, outbound 1.8K). Telegram is the closest-to-clean reference shape.
Invariant enforceable by an ironclaw_architecture test banning slack/webui/
telegram/openai/transport identifiers from composition. Ties to §4.4 and #6168.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): §10 — enforcement / anti-slippage ratchets pulling together all checks

Consolidates the per-axis static checks into one table: process isolation (#6170
two-user escape test), mirror DTOs (check-type-duplicates.py), dyn mediators,
InMemory*Store, Local* types, host_api freeze, and products-in-composition — each
with its home, the addable-now ratchet (freeze current count/allowlist, fail on
new), and the hard ban that is also its definition-of-done when the axis lands.
Notes the guardrail self-test + two-hook-path requirement and that Local*/
InMemory*/composition bans must start as frozen allowlists (can't hard-ban today).
Renumbered Open questions to §11.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): add §11 testing (state-machine invariants, idempotency from any state); distinguish crate seams from trust boundaries

§11 Testing (new): specifies the behavioral suite — the state machines to pin
(turn/run, capability invoke, lease, gate/resume), the invariants that must hold
from ANY reachable state, the idempotency contract, and how to reach arbitrary
states (model-based stateful property tests, exhaustive state×op enumeration,
fault injection), cross-backend parity, fail-closed/adversarial, interface
conformance harnesses, and integration-first tiering. Design only.

Trust-boundary correction (per review): the doc overstated "exactly one trust
boundary." Reborn has several — the untrusted loop, untrusted runtime-lane
execution (WASM/script/MCP/containers/external services), and untrusted
runtime-supplied data (egress/worker output) — all mediated and all preserved.
What collapses is the trusted mediation-chain CRATE SEAMS, not a trust boundary.
Fixed intro, §2 (retitled + enumerates the boundaries), §5.4, §8 invariants
(added lane + data boundaries; fixed stale RowBackend/TurnStore refs to
RootFilesystem), and §11.6 (lane/worker/egress adversarial). Renumbered Open
questions to §12.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): fix last 'the one trust membrane' → the loop's (one of several)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): consistency pass + address PR review feedback

End-to-end review fixes (also addressing gemini/coderabbit comments; note some
reviewed a superseded head that still proposed RowBackend/TurnStore<B>):

- authorize() no longer takes a separate `scope` that can diverge from
  Invocation.scope — derive from inv.scope (§3, §5.3). [coderabbit]
- Clarify the result contract: Outcome carries success OR recoverable failure;
  the seam is Result<Outcome, Blocked>; no separate Err(terminate) (§3). [coderabbit]
- Mark `Authority` SEALED — private fields, host-only construction via authorize()
  (§3, §5.3); resolve open-question 2 accordingly. [gemini + coderabbit]
- Align §4.4 LOCAL_DEV process field with §5.6/§6: HostUnsandboxed(LocalOnly),
  gated by a local-only token a served boot can't mint.
- Retire the RowBackend framing (superseded by RootFilesystem) and note deleting
  InMemory*Store has zero persistence-compat impact; durable backends untouched
  (§4.3). [gemini/coderabbit]
- Add file:line evidence for the five request types + audit date/window (§1.1).
  [coderabbit]
- Intro: "most with exactly one prod impl" -> "several" (only 2 of 5).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): §12 — performance-critical paths (locking, remote-store latency, event fan-out)

Adds a performance section grounded in the actual hotspots: consolidating stores
onto RootFilesystem (§4.3) makes the backend latency profile the kernel's, and it
can be remote (libSQL/Postgres). Ranked critical-path table (per-turn ~11-store
fan-out; heartbeat vs store-lock with lease-TTL coupling; libSQL BEGIN IMMEDIATE
single-writer + #5751/#6089 contention; authorize() per-tool-call reads; active-
thread lock; event append/projection fan-out; recovery poll). Plus the
no-lock-across-remote-I/O rule (turn_scheduler.rs:787/881), what the refactor
helps vs risks (centralizing latency/writer contention onto one seam), and design
guidance (batched per-transition write, isolated heartbeat, cached read-mostly
authority, async coalesced events, writer sharding). Renumbered Open questions
to §13.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): §13 — answer the open questions directly (now Decisions)

1. TurnRun vs ironclaw_processes: converge the mechanism (shared LeasedWorkUnit —
   §4.3 already collapses the store layer), keep the policy distinct (turn resume-
   from-checkpoint vs process terminal+re-spawn); bonus, the shared lease-recovery
   gives processes the reconciler they lack.
2. Authority: one sealed value (host-only construction via authorize()); narrowed
   read projection if an adapter needs a field.
3. DeploymentConfig: surface-disclosure derived from process:HostUnsandboxed;
   mid-run refresh gated by session:LongLived|PerRun; synthetic tools promoted to
   first-party capabilities (default hidden in hosted) — a security improvement
   (they gain authorize/scope-binding). Remaining items are tuning knobs, not
   architecture.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): cross-reference the in-progress Unified Extension Runtime (BenKurrek gist)

Adds §5.9 mapping this doc against the URT extension/adapter/auth design: strong,
independent convergence (no-product-code-in-composition; config-not-code as
recipe+engine auth; runtime-kind=closed lane set; built-ins on the identical
pipeline; the trust boundaries; ProductSurface above the host pipelines).
Complementary scope: URT is the deep extension/adapter/auth axis, this doc the
broader kernel refactor; they compose (URT's dispatcher pipeline = authorize+
dispatch; adapter invoke/deliver = RuntimeLane execution).

Adopts two URT refinements: (a) product_auth collapses to recipe data + one host
AuthEngine, not per-adapter code (§5.8); (b) the Deletion/Addition/retired-
taxonomy tests are the products-in-composition ratchet (§10). Clarifies WebUI
consumes ProductSurface directly and is not a ChannelAdapter. Adds a reference.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): §5.9 — RuntimeLane reconciliation + ToolPorts↔Authority integration points

Name the two seams where this doc and the Unified Extension Runtime must agree:

1. One closed execution enum RuntimeLane = {FirstParty|Wasm|Mcp|Process}; the URT's
   extension-declarable runtime *kinds* (first_party/wasm/mcp) are a strict subset.
   Process (OS-subprocess/script sandbox) is host-only — no manifest can select it;
   only host built-ins (shell/script) dispatch to it via ProcessSandbox. Load kind
   (URT) vs execution lane (this doc) are different axes; don't merge them.

2. ToolPorts is derived from Authority, never independent: dispatch() materializes
   egress (NetworkPolicy + host-side SecretBroker lease), state (ScopedFilesystem =
   Authority.mounts), logging from (&Invocation, &Authority, descriptor). ToolPorts
   can't be wider than Authority grants; the adapter never sees Authority itself. So
   URT's ToolAdapter::invoke(call, ports) IS the body of dispatch(inv, auth, lane).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): address CodeRabbit review findings on the latest head

Nine substantive design-contract fixes:

- §3 core model: LoopRequest (loop pre-trust, input-by-ref) resolved to Invocation
  at the membrane; Authorized = sealed AND invocation-bound (actor/scope/activity_id
  provenance) so dispatch can't be handed a mismatched (inv,auth); activity_id IS
  the invocation idempotency identity (idempotency_key unified with it, not deleted,
  satisfying §11.3); three distinct outcome channels Blocked | HostFailure | Outcome
  (no Ok(Failed)/Err ambiguity). Type count 3→4. §3.1/§5.3/§5.4/§11.1 aligned;
  Authority→Authorized throughout.
- §11.2/§6: scope cross-tenant isolation to multi-user/served deployments (matrix
  test), not "any deployment state" (single-user local legitimately allows host proc).
- §9/§10: quarantine the known-red two-user test (#[ignore]/expected-fail until the
  fix merges); ratchets freeze checked-in symbol allowlists (set membership), not
  aggregate counts (a swap evades a count).
- §12: durable event append is atomic with the state transition (same tx/outbox);
  only subscriber fan-out is decoupled.
- §5.8: adapters resolved via a product-neutral ExtensionId-keyed factory registry
  passed to composition as input; config lists ids, not types.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): approval stores over RootFilesystem, delete InMemory*Store (§4.3) (#6195)

* refactor(reborn): approval stores over RootFilesystem, delete InMemory*Store (§4.3)

First slice of the architecture-simplification note
(docs/reborn/2026-07-17-architecture-simplification-dto-dyn-local.md §4.3):
"in-memory" stops being a bespoke store and becomes a filesystem backend, so
each approval domain has one production Filesystem*Store<F> exercised over the
in-memory backend in tests and libSQL/Postgres in production — no parallel
hand-written implementation to keep in lock-step.

Deletes the three hand-written approval stores in ironclaw_approvals
(InMemoryAutoApproveSettingStore, InMemoryPersistentApprovalPolicyStore,
InMemoryCapabilityPermissionOverrideStore, plus the InMemoryToolPermissionOverrideStore
alias). Everything now runs the existing Filesystem*Store<F>:

- ironclaw_approvals: adds a `test-support`-gated helper module with
  in_memory_backed_* constructors (the production store over a fresh
  InMemoryBackend mounted at /approvals). The stores' own unit tests move onto
  Filesystem*Store<InMemoryBackend>, proving it covers the deleted stores' cases.
- composition factory.rs: the LocalDev* approval-store aliases collapse to one
  unconditional Filesystem*Store<LocalDevRootFilesystem>; the no-durable-features
  local-dev builder wires them over the composite root filesystem (in-memory
  backed) via the existing scoped-filesystem path instead of the deleted
  InMemory* stores. wrap_scoped / invocation_mount_view and the /approvals mount
  machinery are un-gated so both builders share one path.
- host_runtime production-wiring guard: the fail-closed LocalOnly classification
  now keys on FilesystemPersistentApprovalPolicyStore<InMemoryBackend> instead of
  the deleted InMemory type. Production (<LibSql>/<Postgres>) and durable-local-dev
  (<Composite>) classifications are unchanged; the guard contract test is
  repointed and still asserts LocalOnly.
- downstream test suites (host_runtime, composition, product_workflow) repoint to
  the test-support helpers; the affected crates enable ironclaw_approvals/test-support
  in [dev-dependencies].

Net subtractive (−136 LOC). No trust boundary or persistence-compatibility change:
the in-memory approval stores were volatile/local; the durable libSQL/Postgres
backends are untouched. Boundary tests (ironclaw_architecture) stay green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): address review — migrate root harness, honest volatile approval-store type

Two review findings on the approvals-store consolidation:

1. Root integration harness left uncompilable. `tests/integration/support/
   harness/mod.rs` still constructed the deleted `InMemory*Store`s (as
   `Arc<dyn …>` defaults). Repoint to the `in_memory_backed_*` helpers and enable
   `ironclaw_approvals/test-support` in the root `[dev-dependencies]`.

2. Guard weakened for the no-durable composition. The composite-unified alias made
   the no-durable-features build wire
   `FilesystemPersistentApprovalPolicyStore<CompositeRootFilesystem>`, whose
   TypeId misses the guard's `<InMemoryBackend>` branch, so the volatile store was
   classified `ProductionCandidate`. Fix by making the store type honestly reflect
   its volatility: the no-durable build now backs the three approval stores with a
   dedicated `InMemoryBackend` directly (via `wrap_scoped`), so the concrete type
   is `Filesystem*Store<InMemoryBackend>` — which the production-wiring guard
   classifies `LocalOnly`, exactly as the sibling `InMemoryRunStateStore` /
   `InMemoryCapabilityLeaseStore` are. Durable builds keep the composite-backed
   type (distinct, correctly a production candidate). The `LocalDev*` approval
   aliases go back to cfg-split (InMemoryBackend vs composite); the guard contract
   test now documents that it exercises the exact type the no-durable composition
   wires. `local_dev_scoped_filesystem` is re-gated to durable-only (the no-durable
   builder no longer uses it); `wrap_scoped`/`invocation_mount_view` stay ungated
   since the no-durable builder now calls `wrap_scoped` directly.

Verified: composition compiles + clippy clean on default (no-durable) and libSQL;
guard contract test green; local_dev_authorization tests green; root
reborn_integration_* targets compile.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): consolidate WebUI host stack into a single ironclaw_webui crate (+ Slack/OpenAI-compat wiring) (#6194)

* refactor(reborn): merge WebUI host stack into ironclaw_reborn_webui_ingress

Fold `ironclaw_webui_v2` (route surface + SPA bundle) and composition's WebUI
middleware/assembly into `ironclaw_reborn_webui_ingress` so the whole WebUI host
stack is one crate above composition, and composition shrinks.

Move-only for behavior; the composed `webui_v2_app` Router, middleware order,
descriptors, and security invariants are unchanged (locked by the moved contract
tests + the composition/ingress router tests, all green under default features).

Structure:
- `ironclaw_webui_v2/src/*` -> ingress `src/webui_v2/` (public module,
  unconditional); `build.rs` + `frontend/` moved to ingress; crate deleted and
  removed from workspace members (68 -> 67).
- Composition WebUI middleware (`webui_body_limit`, `webui_operator_auth`,
  `webui_rate_limit`, `webui_route_match`, `webui_ws_origin`) + `webui_serve.rs`
  -> ingress `src/`.
- `webui_serve.rs` split: `WebuiServeConfig`/`webui_v2_app`/`WebuiV2App`/
  `Webui{Serve,Config}Error`/`WebuiAuthenticator`/`WebuiAuthentication` move to
  ingress; the mount vocabulary (`PublicRouteMount`/`ProtectedRouteMount`/
  `PublicRouteDrain(s)`) stays in composition (`webui/route_mounts.rs`) because
  nearai/openai/runtime construct it — moving it up would cycle.
- Product-auth decoupled: `ProductAuthRouteState`, `product_auth_route_mount`,
  `ProductAuthRouteMount` exposed `pub` + re-exported from composition root;
  ingress imports them (+ `RebornWebuiBundle`, `GoogleOAuthRouteConfig`) via the
  composition facade. Composition no longer depends on `ironclaw_webui_v2`.
- Callers repointed: composition tests, ingress tests, reborn_cli
  (serve/webui_auth), root v1 int-tier tests + dev-dep, Dockerfile.reborn +
  smoke test frontend path, and the ironclaw_architecture boundary spec.

Deferred (out of scope, feature-gated off by default): the
`slack-v2-host-beta` / `openai-compat-beta` blocks in `webui_serve.rs` still
reference composition-internal surfaces and compile out under default features
(declared as known cfgs). Wiring the Slack/OpenAI-compat host surface through
ingress is a follow-up; composition's slack feature will not build until then.

[skip-regression-check] move-only refactor; behavior covered by relocated
contract tests and existing composition/ingress router suites.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): make Slack + OpenAI-compat host-beta build and wire after WebUI merge

The WebUI host-stack merge (parent commit) hoisted `webui_v2_app` + its config,
authenticator, and middleware surface from `ironclaw_reborn_composition` up into
`ironclaw_reborn_webui_ingress`, but left the `slack-v2-host-beta` /
`openai-compat-beta` blocks in the moved `webui_serve.rs` pointing at
`crate::slack::*` / composition internals that don't exist in ingress. Those
features were declared only as known-cfgs and compiled out, so:

- composition failed to build under `slack-v2-host-beta` (two mount-vocabulary
  imports still on the old `webui::webui_serve` path);
- the ingress serve blocks were permanently dead, so the CLI's slack/openai
  features forwarded to composition but never mounted the Slack routes —
  a functional parity break, not just a compile break;
- composition's slack-gated tests still imported the moved `webui_v2_app`.

Wiring (behavior-preserving; restores pre-merge parity):
- ingress now defines real `slack-v2-host-beta` / `openai-compat-beta` features
  that forward to composition (+ optional `ironclaw_reborn_openai_compat`); the
  moved `webui_serve.rs` reaches Slack setup/route types and the
  OpenAI-compat bearer-evidence helper through composition's public facade
  (`ironclaw_reborn_composition::{SlackPersonalSetupServiceSlot,
  SlackChannelRouteAdminRouteConfig, slack_channel_route_admin_route_mount,
  SlackPersonalOAuthBindingConfig, mark_bearer_token_verified_for_tenant}`).
  Ingress does NOT depend on `ironclaw_product_adapters` directly — the
  architecture boundary (`reborn_dependency_boundaries.rs`) forbids it, so the
  evidence helper is re-exported from composition instead.
- composition promotes `slack_channel_route_admin_route_mount` + its
  `SlackChannelRouteAdminRouteMount` return type to `pub` (its sole caller,
  `webui_v2_app`, moved up), mirroring the already-public `ProtectedRouteMount`.
- CLI forwards `slack-v2-host-beta` / `openai-compat-beta` to the ingress crate
  as well as composition, so the serve blocks compile in and the routes mount.

Tests:
- The 7 composition slack unit tests that drove the now-relocated `webui_v2_app`
  move to `ironclaw_reborn_webui_ingress/tests/slack_host_beta_webui_v2.rs`.
  They use only composition's public builders, so ingress (which normal-deps
  composition — single crate copy, no dev-dep cycle) is their correct home; a
  composition lib-test cannot call the ingress `webui_v2_app` without cargo
  building two incompatible copies of composition. composition's ingress
  dev-dep gains `slack-v2-host-beta` so its own `webui_v2_product_auth*` tests'
  `with_slack_*` blocks compile.

Verified (clean env): composition/ingress/cli build + clippy `-D warnings` under
both beta features; ingress `--all-features` suite green incl. the 7 relocated
tests; composition lib (1589) + router (webui_v2_serve 44 / product_auth 51) +
cli (440 incl. Dockerfile smoke) green; `ironclaw_architecture` boundaries hold.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): rename crate ironclaw_reborn_webui_ingress -> ironclaw_webui + doc pass

Now that the crate owns the whole WebUI host stack (route surface + SPA +
gateway assembly/middleware + serve loop + host auth), "reborn_webui_ingress"
undersells it. Rename the crate to `ironclaw_webui` and refresh its docs to
describe the composed subsystems.

Rename (pure identifier swap, no behavior change):
- `git mv crates/ironclaw_reborn_webui_ingress crates/ironclaw_webui`; package
  `name` + workspace members + root dep alias updated.
- Every `ironclaw_reborn_webui_ingress` reference repointed across Rust, Cargo
  manifests, Cargo.lock, Dockerfile.reborn, CI scripts (.sh/.py), the
  `ironclaw_architecture` boundary spec (crate_name / forbidden lists / layer
  exception / source-path prefixes), root + crate CLAUDE/AGENTS docs, .claude
  rules & skills, and the security-parity docs. `openwiki/` (auto-generated) and
  `docs/plans/` (historical) intentionally left for their own regen/record.

Docs (README.md new; AGENTS.md + CLAUDE.md restructured):
- README.md: human-facing overview with the three-piece fold-in map
  (route surface + SPA from the former `ironclaw_webui_v2`; gateway assembly +
  middleware from `ironclaw_reborn_composition::webui`; serve loop + host auth
  from this crate's original scope), layering/boundaries, feature flags, build/test.
- AGENTS.md: replaced the stale "deliberately small" framing with an accurate
  agent map — composed subsystems, do-not-move-in, allowed deps, how to add a
  route / authenticator / OAuth provider.
- CLAUDE.md: reframed opening (it no longer is a "counterpart to webui_v2_app" —
  that fn lives here now); Surface table gains the route/gateway symbols
  (`webui_v2_router`, `webui_v2_routes`, `WebUiV2State`, `WebUiV2HttpError`,
  `webui_v2_app`, `WebuiServeConfig`); folded in the WebChat v2 route table +
  streaming/SSE model + SPA build detail; test layout now lists the
  route-surface/gateway suites. OAuth login security contract retained verbatim.

Verified: `cargo metadata` resolves; `cargo build -p ironclaw_webui` and
`-p ironclaw_reborn_cli --features slack-v2-host-beta,openai-compat-beta` green;
`cargo test -p ironclaw_architecture reborn` (boundaries, new name) green;
`cargo test -p ironclaw_webui --features slack-v2-host-beta --test
slack_host_beta_webui_v2` green; 0 stale `ironclaw_reborn_webui_ingress` refs
outside openwiki/docs-plans.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): address PR review findings + stale ironclaw_webui_v2 refs post-merge

Review findings on PR #6194 (gemini-code-assist + ironloopai) and the leftover
references the `ironclaw_webui_v2` → `ironclaw_webui` fold-in left behind.

CI / build path migration (ironloopai "path migration incomplete"):
- Repointed the deleted `crates/ironclaw_webui_v2/frontend` build path to
  `crates/ironclaw_webui/frontend` across all workflows (code_style, coverage,
  ironclaw-stress, platform-and-compat, reborn-e2e, reborn-playwright),
  `.dockerignore`, `scripts/run-reborn-webui.sh`, `scripts/ci/quality_gate_strict.sh`,
  and the `regression-test-check.yml` frontend-test detector.
- Test bucketing: dropped the dead `ironclaw_webui_v2` entries from
  `reborn-crate-test-buckets.sh` + `package-feature-flags.sh` (the renamed
  `ironclaw_webui` entries already exist), repointed `classify-test-scope.sh`,
  and widened the `reborn-tests.yml` jq filter to `startswith("ironclaw_webui")`
  so the folded crate's tests still land in the webui bucket.
- QA inventory (`scripts/reborn_qa_matrix/audit_surface_inventory.py`) now reads
  `crates/ironclaw_webui/src/webui_v2/descriptors.rs`.
- Regenerated `harness/latency/runner/Cargo.lock` (transitively referenced the
  deleted crate via composition's `webui-v2-beta`).

Broken doc links / stale comments (gemini):
- `nearai_login_serve.rs` + `runtime.rs`: the broken intra-doc link
  `crate::webui::route_mounts::WebuiServeConfig` (type moved out of composition)
  is now a plain code span `ironclaw_webui::WebuiServeConfig::with_public_route_mount`
  — composition cannot link into `ironclaw_webui` (not a dependency).
- `webui/facade.rs`: comment now says routing/auth/static/SSE live in
  `ironclaw_webui`; only the route-mount vocabulary stays in `route_mounts`.

build.rs frontend opt-out (gemini):
- `SKIP_FRONTEND_BUILD=1` skips the Node/pnpm frontend build for backend-only
  dev / docs.rs / minimal CI images (`webui_enabled = env::var_os(...).is_none()`).

Guidance docs (ironloopai "update AGENTS/CLAUDE + crates/AGENTS.md"):
- `crates/AGENTS.md`: rewrote the `ironclaw_webui` row to the whole WebUI host
  stack, removed the deleted `ironclaw_webui_v2` row, repointed cross-refs.
- Refreshed `ironclaw_webui_v2` → `ironclaw_webui` across living guidance
  (`.claude/` rules/skills/commands, `crates/README.md`, `crates/Architecture.md`,
  `crates/ironclaw_projects/CLAUDE.md`, `ironclaw_reborn_composition/CLAUDE.md`,
  product_workflow comments, security-parity docs) and the `-p ironclaw_webui_v2
  --features webui-v2-beta` commands. `ironclaw_webui_v2_static` (a distinct,
  still-live v1 crate) left untouched.

Already addressed earlier in this PR, confirmed still green post-merge:
- Slack / OpenAI-compat host-beta compile + wiring (the `slack-v2-host-beta` /
  `openai-compat-beta` findings) — commit `a2ed602`.
- The obsolete `/v2` SPA mount — replaced by main's root-serving
  `static_router_with_config` in the merge (`c000a16`).

Verified: clippy `-D warnings` on composition (`webui-v2-beta`) and
`ironclaw_webui` (`--all-features`); root int-tier webui tests compile;
QA-inventory path resolves; harness lock clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ci): rustfmt import ordering after crate rename + refresh composition pub-use snapshot

Two CI failures on PR #6194:

- **Formatting / Code Style (fmt+clippy)**: the `ironclaw_reborn_webui_ingress`
  → `ironclaw_webui` rename shifted where the crate sorts in `use` blocks, so
  rustfmt wanted to reorder imports across ~26 files. I had wrongly reverted
  those fmt-only files during the rename commit (assuming rustfmt-version
  drift); the reordering is deterministic and CI's gate caught it. Ran
  `cargo fmt --all`.

- **Test Reborn crate bucket (adapters-misc)** →
  `composition_public_pub_use_surface_matches_snapshot`: this PR intentionally
  changed composition's public facade — `webui_serve`/`Webui*`/`webui_v2_app`
  moved out to `ironclaw_webui` (so composition's `webui` re-export is now just
  `route_mounts::*`), the product-auth mount builders were exposed, and the
  Slack channel-route mount + `mark_bearer_token_verified_for_tenant` were
  promoted. Regenerated `docs/plans/composition-pubuse.snapshot` by replaying
  the test's own `extract_pub_use_surface` extraction; the diff is exactly those
  intended facade changes.

Verified: `cargo fmt --all -- --check` clean; `cargo test -p
ironclaw_architecture --test reborn_composition_boundaries` green (8 passed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(composition): extract runtime.rs inline test module (Phase 0) (#6173)

* refactor(composition): extract runtime.rs inline test module to sibling file

Moves runtime.rs's trailing `#[cfg(test)] mod tests { … }` (~6.9k lines) into
`runtime/tests/core.rs` via the crate's existing `#[path = "runtime/tests/…"]`
convention. Pure move — the module keeps its identity (`crate::runtime::tests`),
so all `super::`/`crate::` refs resolve unchanged; cargo fmt de-indented the
relocated items. runtime.rs: 11,673 -> 4,709 lines.

Phase 0 of the composition decomposition (parent #4471, plan #6168): single-
crate, zero cross-crate coupling, does not touch slack/ or extension_host/.
Fixes the crate's worst file-size violation and — because the inline test block
no longer counts as production LOC (it's now a test-only file, excluded) —
ticks the composition mass ratchet down (~23.98% -> ~23.2%).

Verified: `cargo test -p ironclaw_reborn_composition --all-features --no-run`
compiles all relocated tests unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(composition): rustfmt runtime test module + merge main

Remove stray leading blank line in runtime/tests/core.rs flagged by the
Formatting CI check, and merge origin/main to bring the branch current.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Require read-before-edit and reject stale edits in reborn coding tools (#5978)

* Ride out provider outages and drop the 32-call turn cap in the reborn loop

Two failure modes discovered via claw-swe-bench-lite run 9ca133e5 (30% vs
hermes 65% on the same model) discarded hours of agent work:

- A transient provider 5xx storm aborted the whole run after 2 quick
  retries (max_attempts_per_class=2, backoff capped at 5s). Availability-
  class model errors (transient/unavailable/internal) now retry on their
  own deeper budget: max_model_availability_attempts=12 with a 1s..60s
  exponential backoff, riding out ~7 minutes of sustained provider
  failure. MAX_MODEL_RETRIES raised 8 -> 16 to let the strategy govern.

- DefaultBudgetStrategy's iteration_limit=32 failed closed mid-task with
  no synthesis (llm_calls in failed bench tasks clustered at exactly
  63/64/127/128). The default is now DEFAULT_ITERATION_BACKSTOP=1024
  (subagent 16 -> 256), documented as a runaway backstop: operational
  bounds are the resource budget system and stop-condition strategy.

New seam mirroring IRONCLAW_REBORN_PLANNED_DEFAULT_ITERATION_LIMIT:
IRONCLAW_REBORN_MODEL_AVAILABILITY_RETRY_ATTEMPTS ->
DefaultPlannedRuntimeConfig.planned_model_availability_retry_attempts ->
families::default_with_overrides. The integration group harness pins
attempts=1 so scenarios that deliberately script provider failures
(failure_category_demasked) reach Failed in seconds, not minutes; the
availability-retry tests run under paused tokio time.

Family fingerprint digests regenerated for the new strategy parameters.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Surface tool-failure reasons to the model for shell and coding tools

Benchmark traces showed the model retrying identical failing calls blind:
builtin.shell parameter errors and coding-tool path rejections reached it
as a bare category ("the tool input could not be encoded") because the
handlers built FirstPartyCapabilityError/CodingCapabilityError with no
safe_summary — the model-visible Diagnostic detail channel downstream was
already wired but starved (one agent burned 13 apply_patch calls against
an out-of-scope /testbed path with empty errors).

- shell.rs: shell_error/process_error now carry the concrete reason
  ("missing 'command' parameter", timeout duration, spawn failure),
  bounded to 512 chars. The strict safe-summary validator still falls
  back to the fixed category string; the reason always survives on the
  secret-scrubbed diagnostic channel.
- coding/paths.rs: scoped-path rejections name the offending path and
  the available scoped roots; permission rejections say the operation is
  not permitted on that mount.

Covered at the dispatch tier (coding state dispatch, host-runtime
invoke_capability) per test-through-the-caller.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Run agent_loop scenario test targets under paused tokio time

The deep availability-retry backoff added for provider-outage ride-out
made outage-scripting scenario tests sleep for real: safety_nets alone
took ~423s (the exact cumulative backoff schedule) because scripted or
script-exhausted model errors now retry for minutes. Pause the clock on
all executor scenario targets — they drive the in-process mock host
exclusively, so timers auto-advance and the suites return to seconds.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Fail fast when no LLM provider is configured instead of riding availability retries

The placeholder unconfigured provider's RequestFailed was mapped through the
catch-all Unavailable arm, so the new deep availability retry budget rode a
permanent configuration fault through ~7 minutes of exponential backoff.
Users with no LLM configured waited minutes for an error that retrying can
never fix, and the Reborn CLI smoke tests that pin fast nonzero exits timed
out (the 4 failures on CI run 29136954176).

Map errors carrying the shared UNCONFIGURED_PROVIDER_ID to
CredentialUnavailable, which is unclassified in loop recovery and therefore
terminal on first sight; the Settings → Inference hint travels on the
scrubbed detail channel. The provider id moves to a shared constant in
ironclaw_llm so the composition placeholder and the runner mapping cannot
drift.

Regression tests: unconfigured_provider_error_maps_to_credential_unavailable_
not_availability and unconfigured_provider_detection_requires_the_placeholder_
provider_id in model_gateway.rs; the existing smoke tests
(*_exits_nonzero_when_runtime_does_not_produce_reply) pin the fast-fail at
the caller tier.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Derive override-built default-family replay identity from resolved config

families::default_with_overrides swapped budget/recovery strategies but
reused the planner's static version digest, so an overridden composition
carried the pure-default replay identity — violating the component-identity
contract (family.rs: the digest identifies replay-relevant configuration).

- Turn the cfg(test) fingerprint const into a runtime
  default_family_fingerprint(iteration_limit, model_availability_attempts)
  builder; override-built families hash it with their resolved values at
  composition time (BLAKE3, same path as the pinned const). The pure-default
  composition keeps the static DEFAULT_FAMILY_DIGEST, and overrides spelling
  out the production defaults hash to that same digest.
- Collapse the two Option args into a FamilyOverrides struct and drop the
  now-dead (None, None) branch in the runner's registry factory.
- Tests: digest differs per override knob and is deterministic; explicit
  production defaults reproduce the static digest; an attempts=1 override
  reaches the composed recovery strategy (one retry then abort).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Let the recovery strategy own the model retry guard; wake backoff on cancel

Two model-stage fixes from the PR 5959 review:

MAX_MODEL_RETRIES=16 silently capped any configured availability budget of
16+: the retry loop fell through to a generic ModelError exit with
FailedExitDetails::default() — no failure category, no diagnostic ref —
before the strategy could reach its own Abort. The executor now derives the
loop bound from the composed strategy via
RecoveryStrategy::max_total_model_attempts() (DefaultRecoveryStrategy
computes it from its per-class + availability budgets with margin), so
every accepted override reaches the strategy's abort boundary. The
contract-bug fall-through now carries the last observed model error's
category and diagnostic ref instead of empty details.

The availability backoff sleep (up to 60s per attempt) was not
cancellation-aware: a cancel request could wait out the full delay. The
sleep now selects over the host's cancellation_requested() future (same
pattern as the prompt-compaction and failure-explanation waits), and a
boundary cancel check right after the alteration turns the wake into a
checkpointed Cancelled exit without issuing another model call.

Tests (paused tokio time): an availability budget of 20 — past the old
executor cap — fails with the strategy's model_unavailable category and
diagnostic ref after exactly 21 model calls; cancellation requested during
the first 1s backoff exits Cancelled without riding out the sleep.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Clarify DEFAULT_ITERATION_BACKSTOP doc: resource budgets are not yet enforced

The doc claimed operational bounds come from the resource budget system,
but ResourceBudgetPolicy.max_model_calls and the wall-clock cap are defined
and not applied; until they are, this backstop and the stop-condition
strategy are the only live ceilings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Carry tool-failure reasons to the model past the strict summary validator

PR 5959's headline feature (model-visible tool-failure reasons) never
reached the model for path-bearing reasons: LoopSafeSummary rejects
path/payload delimiters and newlines, and dispatch_failure_message
silently degraded every such reason to the generic category sentence
before it could reach the diagnostic channel.

- production.rs (failure_from): a host-authored safe_summary that fails
  LoopSafeSummary validation is preserved as the new
  DispatchFailureDetail::Diagnostic instead of being dropped; the
  message keeps the fixed category sentence (host-authored, Invariant 2).
- capability_port.rs: maps the Diagnostic detail into the model-visible
  CapabilityFailureDetail::Diagnostic, scrubbing secret values and
  normalizing control characters the observation validator rejects (so
  one stray escape byte cannot drop the whole observation); newlines
  are preserved. The RetrySameCall arm now forwards structured detail
  too.
- coding/paths.rs: scoped-path rejection summaries render the path and
  available roots delimiter-free ("path testbed replacer.go",
  "available roots: workspace") so they pass the strict validator —
  FilesystemDenied surfaces as a Denied loop outcome whose only
  model-visible channel is the summary itself.
- shell.rs: bounded_failure_reason documents the (now real) diagnostic
  flow; truncation remains char-based (no byte-boundary panics).

Regre…

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6181 — d691ad91 Deployed Jul 17, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: M 50-199 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant