Conversation
…contract (WS3)
Deletes the two `-> ironclaw_extensions` layer-matrix exceptions
(`ironclaw_mcp`, `ironclaw_scripts`) by giving the runtimes-layer lanes a
contracts home for the descriptors they read, instead of the registry crate
they may not depend on. Exceptions 13 -> 11; baseline lowered in the same
change.
Moved to `ironclaw_extension_contracts`:
- `runtime::{ExtensionRuntime, ExtensionAssetPath, ExtensionAssetPathError}`
- `hosted_mcp::{HostedMcpDiscoveredTool, HostedMcpDiscoveredToolAnnotations}`
`ExtensionPackage`/`ExtensionManifest` deliberately stay in
`ironclaw_extensions`: they carry the whole parsed manifest tree and a
`PackageRootBinding` typed on `ironclaw_filesystem::VirtualPath`, which the
§11.2.3 contracts-purity allowlist (`{ironclaw_host_api}` only) forbids the
contracts crate from naming. Measured instead: both lanes read exactly three
things off the package — `id`, `capabilities`, `manifest.runtime` — so the
lane request structs now take those three and the caller (which owns the
package) projects them.
Also repointed `ResourceReceipt` to its real owner: `ironclaw_resources`
only re-exports `ironclaw_host_api::resource::ResourceReceipt`, so the lanes'
import was a §11.2.4 two-import-paths hop, not a dependency.
No `pub use` shims (§11.3): every consumer is repointed in this change, and
`resolve_under` becomes the free function `ironclaw_extensions::resolve_asset_under`
because the orphan rule forbids an inherent impl on the moved type.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Creates `ironclaw_sandbox` (runtimes) from the three halves of "run an already-authorized command away from the host", and deletes the two crates PROPOSAL §6.6.4 marks for merge: - `ironclaw_process_sandbox` (plan contract) -> `src/plan.rs`, `src/validation.rs` - `ironclaw_host_runtime::sandbox_process` -> `src/sandbox_process/**` - `ironclaw_scripts` (script lane + Docker path) -> `src/script.rs` The kernel sheds the Docker/CA cone: `bollard`, `rcgen`, `x509-parser` and `time` are gone from `ironclaw_host_runtime`'s manifest, and `bollard`/`rcgen` are now declared by exactly one crate in the workspace. Two migration details PROPOSAL §6.6.4 and CHECKLIST WS10 call load-bearing: - `PROCESS_SANDBOX_CAPABILITY_ID` -> `ironclaw_host_api::capability`, so `ironclaw_loop_host` drops its lane dependency (production dep gone; a dev-dep remains for the tests that build plans). - `SandboxCommandTransport` -> `ironclaw_host_api::process`, with the shapes it names (`CommandExecutionRequest`/`Output`, `RuntimeProcessError`, `SavedCommandOutput`, `SavedCommandOutputSanitization`). Without this the runtimes-layer lane could not implement what the kernel consumes. Enumerating gates were repointed, never relaxed: the specificity carve-outs and the struct/test-support ratchet entries moved with their files (both baselines unchanged at 129 and their prior values), the panic-gate baseline row moved, `reborn-crate-test-buckets.sh` registers the new crate, and the three `reborn-e2e-rust.sh` script selectors follow the tests (plus `docker_security`, which had no selector before). One gate would have gone silently vacuous and was fixed rather than moved: the script-lane surface scan in `reborn_dependency_boundaries.rs` read a hardcoded `src/lib.rs`, which after the merge no longer holds the lane. It now scans the whole crate source tree with a fatal-read walk and a non-vacuity assertion. One deletion, recorded: `RebornScopedSandboxCommandTransport::into_process_port` returned a kernel type a runtimes crate may not name. It had zero callers workspace-wide; the kernel wraps the transport, which is the direction the port inversion requires. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ence Three dated amendments, each quoting the text it replaces: 1. CHECKLIST WS3 sandbox row + PROPOSAL §6.6.4 — "all pieces currently unwired/test-only" is REFUTED. Three production paths cross the merged crate (spawn-path plan validation, the process_executor routing check, and the saved-command-output scope digest). The accurate claim is narrower: no production *execution backend*. Behavior preservation is therefore argued at the diff (11 of 26 moved files byte-identical, 9 more differing by one import line, +63/-36 overall), not inferred from deadness. 2. CHECKLIST WS3 mcp row + PROPOSAL §6.6.3 — the prior wave's "structurally blocked" finding is half right, and the wrong half is load-bearing: only `ExtensionPackage` is un-absorbable, and no lane ever needed it (both read `id`, `capabilities`, `manifest.runtime` and nothing else). The registry half of the flip is done; the `resources` half is refuted as phrased — the estimate/usage vocabulary the row asks about is already in `host_api::resource` and already imported from there, while the real blocker is the `ResourceGovernor` authority port and `ResourceError`'s denial cone. 3. Recorded as a structural finding, not a note: the sandbox row and the mcp row are ONE problem. `ironclaw_scripts` imports the identical DTO set, so the merge alone deletes zero exceptions and only the mcp carve-out lets either lane shed the registry edge. Also reconciled: PROPOSAL §6.1.2's as-built inventory gains the two modules WS3 landed (and states why `ExtensionPackage` stayed); §2's package count 66 -> 65; the §9 disposition rows for `ironclaw_scripts`/`ironclaw_process_sandbox`/ `ironclaw_mcp`; the §11.2.2 ratchet rows (13 -> 11); the WS3 verify row; the stale WS1.3 sentence asserting the blocker as settled fact; and `reborn_restructure_baselines.rs`'s doc table, which still read 15. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`process_port.rs` no longer names `MountView` or `thiserror::Error` (both went to `host_api::process` with the types that used them), and `sandbox_process.rs` no longer needs `sync::Arc` after `into_process_port` was deleted. Found by per-crate `clippy --all-targets --all-features -D warnings`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tions Three fail-closed gaps in `reborn_pr_test_plan.py`, all hit by this PR and all live on `main` today — any PR with the same change shape is unplannable. 1. `.claude/**` was unclassified, so the planner refused outright. It is agent guidance in exactly the sense `docs/**` is human guidance: no Rust test reads either as data (the only in-tree references are prose citations in test doc comments). Added to `IGNORED_PREFIXES`. Without this, "guidance travels with the change" — the restructure's own discipline — cannot be satisfied in a single PR. 2. `crates/AGENTS.md`, `crates/README.md`, `crates/Architecture.md` raised "unmapped crate path": they sit under `crates/` but belong to no package. Now classified as crate-tree prose, matched by "Markdown no package directory owns" so a genuinely unmapped crate path is unaffected. 3. An unmapped crate path used to raise. `git diff` reports a deleted crate's old paths and CI feeds the planner that diff, so **every crate deletion or rename was unplannable** — including the six deletions PROPOSAL §2 plans. It now widens to the exhaustive plan. This is a semantic change and it is the safe direction: the full plan is a superset of any narrowing, so an unattributable path can never cause under-selection, whereas refusing to plan blocks the PR instead of protecting it. Malformed input is still rejected by the unclassified-path branch. Each lands with fixtures per WS10's rule, positive and negative: guidance paths select nothing while non-guidance paths still fail closed; crate-tree prose selects nothing while crate *code* under the same unmapped directory widens to `full` (so the Markdown carve-out cannot swallow code). The pre-existing `test_unmapped_crate_path_fails_fast` is renamed and rewritten to pin the new contract rather than deleted. Verified against this PR's real 130-path diff: the planner returns `mode: full`, and the workflow's own exhaustiveness guard passes on that output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… a wave Review (#7065) caught that both surviving `-> ironclaw_resources` exceptions declared `removes_in = "WS3"` — the wave this PR *is*, which does not remove them. That is precisely the defect §11.2.2 already records against `conversations -> turns` ("`removes_in = "WS5"` and WS5 has partly shipped without it falling"), and it would have been repeated here. Both now point at issue #7067, which owns the design work that actually clears them: replacing the `ResourceGovernor` dependency with a narrow reserve/reconcile/release port. The issue carries the measurements — 3 of 10 methods used, zero implementors, and the `ResourceError` denial cone — plus the two open questions (error shape, port home) that make it a design slice rather than a move. An owning issue is also what §11.2.2 asks for and what the ratchet still cannot enforce (there is no `owning_issue` field yet), so this is the strongest form currently expressible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…on_contracts `validate_asset_path` moved here with `ExtensionAssetPath`, the type it constructs. In `ironclaw_extensions` it was only ever reached indirectly through manifest parsing, so its six rejection branches had no direct test — and a contracts crate that carries validation owes that validation one. Two tests: every reject branch with its exact reason and `Display` output (empty, NUL/control, URL, absolute, Windows drive and backslash, and the empty/`.`/`..` segment cases) plus the manifest-relative shapes that must keep being accepted; and `ExtensionRuntime::kind()` over all five variants, since that projection is what every lane uses to reject a runtime it does not serve. Also removes a changed-line coverage risk this PR would otherwise carry into the merge queue: the gate does not run on ordinary PRs (#7036), so ~100 newly-added lines of validator would first be measured where a failure is expensive to diagnose. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…andbox lane
`RATCHET FAIL: ironclaw_host_runtime` — observed 18854 covered vs a
`floor_covered_lines` of 20538. This is the shrinkage case the ratchet's own
"To fix" text describes, not a coverage regression: `sandbox_process/**` moved
to `ironclaw_sandbox`, so the crate's denominator fell 23277 -> 21267 (-2010
instrumented lines) and its covered lines fell with it.
The percentage floor is **raised, not lowered**: observed 88.65% against an old
floor of 88.23%, so the entry now reads 88.65. Only the absolute line count
moves down, and it must — those lines are no longer in this crate.
To keep that from being a net loss of protection, `ironclaw_sandbox` is floored
on arrival at its observed 87.09% (3185 / 3657). This is a net *increase* in
ratchet coverage: neither `ironclaw_scripts` nor `ironclaw_process_sandbox` was
ever floored, and the `sandbox_process` half was protected only as part of
host_runtime's line count, which this PR necessarily reduces. Floored crates
16 -> 17.
Verified by replaying the ratchet arithmetic against CI's observed numbers:
both crates pass on percentage and on covered lines. Numbers taken from the
failing run's own report (job 91740733521), which is the authority for this
gate.
The `Tests (Reborn)` roll-up failed solely on this sub-job
("coverage-report result 'failure' did not match planned=true"); no other lane
failed — 50 pass, 2 fail, both this root cause and its roll-up.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…itive gate WS3 hit a gate no move row had named. `tests/integration/coverage-floor.toml` is keyed on crate identity plus absolute covered-line counts, so it is invisible to WS10's path-keyed gate audit and yet it fails on every crate move, merge, rename, or family `git mv` that shifts instrumented lines between crates — as it did here, while the percentage floor was *improving*. Recorded on WS10 with the three rules WS7 will need: re-capture in the same PR, raise the percentage floor rather than leaving it, and floor the destination crate or the move silently drops that code out of the ratchet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Brings in #6780 (ironhub deep-link register/install + the REBORN_COV_COLLECT hermetic-env allowlist fix), #7050, and #7033. Two conflicts, both resolved as a union with each side's contribution verified present afterwards: - `crates/ironclaw_extension_host/src/available_extension_import.rs` — main added `use ironclaw_extension_contracts::recipe::VendorAuthRecipe;` at the same import position this branch added `use ironclaw_extension_contracts::runtime::ExtensionAssetPath;`. Both kept. - `scripts/ci/test_reborn_pr_test_plan.py` — main added `test_selected_integration_lane_keeps_msrv_override` and kept `test_unmapped_crate_path_fails_fast`; this branch had replaced the latter with `test_unmapped_crate_path_widens_instead_of_refusing`. Resolution keeps main's new test verbatim and this branch's rewrite, and drops the superseded original — it asserts the exact behavior this branch deliberately changed (an unmapped crate path now widens instead of refusing, so crate deletions are plannable). Keeping both would have been contradictory, not a union. 37 tests pass (36 here + main's 1). Post-merge verification of both sides: exceptions 11 with baseline 11 and `ironclaw_scripts` absent from the matrix (this branch); 9 #7033 decision markers and the `REBORN_COV_COLLECT` allowlist entry present (main). Note on the coverage floors this branch re-captured: #6780's fix stops the hermetic env filter stripping `REBORN_COV_COLLECT`, which gates *whether* a lane collects coverage. This branch's numbers were captured on a `full` plan, where every lane collects either way, so they are expected to hold — CI re-measures and will say so. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…Path A semantic conflict the merge could not see: #6780 landed `ironhub/{package,catalog}.rs` importing `ExtensionAssetPath` from `ironclaw_extensions`, while this branch moved that type to `ironclaw_extension_contracts::runtime`. Different files, so git auto-merged cleanly and the breakage surfaced only at `cargo check`. Repointed both sites to the contracts crate (no shim, per §11.3). The manifest already named `ironclaw_extension_contracts`, so this is imports only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…gate The changed-lines coverage gate went red on four files while changed-line coverage was 95.35% against a 90% floor: the failure was its two fail-closed STRUCTURAL assertions, not any percentage. Every line below was derived by replaying scripts/ci/reborn_changed_coverage.py against this PR's own merged lcov (run 30831658659) with the base lcov the gate itself resolved (run 30828540055 @ b89fcd3), until the replay reproduced the CI verdict byte-identically. Line numbers come from the gate's own `candidate_lines - mechanically_uninstrumentable_lines()`, not from the log. - host_api/src/process.rs (31 lines): new placement-neutral process vocabulary with no function body anywhere in the file; rustc emits no LCOV record for it at all. Same shape already exempted for product_contracts/loop_contracts. - extension_contracts/src/hosted_mcp.rs (12): field declarations of the two new tools/list descriptor structs. The file is plainly instrumented (191 DA, 164 hit), so this is a no-region artifact, not an instrumentation gap. - host_runtime/src/services/runtime_adapters.rs (13): continuation lines of three rewritten calls, all PROVEN EXECUTING by their region-start heads (lines 380/434/977 score 24/16/63 hits). The four genuinely-uncovered lines in the same rewrite are deliberately NOT exempted -- the gate already subtracts them as pre-existing debt inherited from base. - composition capability_host_tests/approval_gates.rs (6): type positions in a test double whose body region scores 1 hit. The last one is a finding, not just a waiver: that file is 100% test code behind `#[cfg(test)] mod capability_host_tests;`, but the gate's test_only_path() recognises /tests/, /test_support/, */tests.rs and *_tests.rs and NOT a cfg(test) module DIRECTORY, so it measures it as production. It is the only such directory in crates/ today. Docs (target-architecture, same PR per the docs-truth rule): - CHECKLIST WS10 gains the changed-lines gate beside the ratchet row, cross- referencing the WS2.1 note rather than restating it: percentages are not what fail a move; derive lines by byte-identical replay (--fetch-base-coverage silently degrades without --github-repo); and a stranded exemption path is an ABORT with no verdict, not a loud failure. - CHECKLIST WS10 exception-ratchet row: the constant was cited at :4063 and sits at :4164 -- corrected by removing the line pin, since the file is edited every wave. Records that the baseline is a UNION across parallel WS3 lanes. - families/contracts.md: records extension_contracts' new ownership of the runtime descriptor vocabulary -- the carve-out that let BOTH lanes drop the registry edge -- and the orphan-rule seam that keeps resolve_asset_under in the registry crate. - families/lanes.md: two "Never" claims were reading as satisfied when they are not. ironclaw_mcp's "never depends on the resource-governor crate directly" is refuted (the compiled edge survives; #7067 tracks the narrow port), and ironclaw_sandbox's "no direct process spawning outside the transport seam" is aspirational -- script.rs:454 still builds Command::new("docker"). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tion cost Two review findings verified against the tree; three refuted with evidence in the PR threads. Valid — the sandbox wiring inventory was self-contradictory. `CLAUDE.md` said "Two production call paths ... and both are plan validation" directly above a list of THREE bullets, and `lib.rs` omitted the third entirely. The third is real and is not validation: `host_runtime/src/process_output.rs:482` derives the scoped saved-output directory through `RebornSandboxScopeKey::from_scope`. That inventory is what tells a future agent which paths are live, so an undercount invites deleting a production path as dead code. Both surfaces now say three and no longer claim they are all plan validation (the `loop_host` capability-id comparison never was either). Valid, and recorded rather than redesigned — the registry carve-out cost a type-level invariant. Replacing `package: &ExtensionPackage` with independent `extension` / `capabilities` / `runtime` borrows is what deleted the `mcp -> extensions` and `scripts -> extensions` exceptions, but it also means the type no longer guarantees the three came from one package. `execute_extension_json` re-checks the descriptor half (`descriptor.provider == extension`); the runtime half cannot be re-derived, because nothing in an `&ExtensionRuntime` names its owning extension. No caller can trip it today -- there is exactly one production caller (`runtime_adapters`) and it projects all three from one package in one expression -- so this is a latent structural weakening, not a live defect. Restoring the compile-time binding needs a sealed projection minted by the package owner; a check inside the lane cannot express it, and re-taking the registry edge would undo the carve-out. Both request types now carry the caller obligation in their field docs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…pport (WS3)
WS3's first-party-tools row, family 1 of 6: skill management / URL install.
`skill_url_install.rs` and its `bundle`/`github`/`zip_bundle` submodules,
plus the install-input normalizer, move out of
`ironclaw_host_runtime::first_party_tools` into
`ironclaw_extension_support::skills::{url_install, resolve_install_input}`,
where the skill executor half already lived. Move-only: no behavior change,
no test edited for content.
`ironclaw_host_runtime -> ironclaw_skills` is deleted from
LAYER_MATRIX_EXCEPTIONS — the edge is gone, not waived (exceptions 13 -> 12,
WS0_LAYER_MATRIX_EXCEPTION_BASELINE drops with it). `ironclaw_skills` and
`zip` survive as dev-dependencies for host_runtime's own tests; dev edges are
outside the matrix by construction.
Two doc ambiguities are resolved in the same diff, as dated PROPOSAL
amendments quoting the text they replace:
- §6.8.4's "the builtin first-party tool handlers absorbed from
host_runtime/first_party_tools" contradicted §8.2's "kernel: ✗ (ports only)"
row and the enforced BoundaryRule. Resolution: the seam splits executor from
adapter — the executor moves behind a neutral request/error pair, the
FirstPartyCapabilityHandler / CapabilityManifest / registry wiring stay
host-side. Same shape the groupware and web-access tools already ship.
- §8.2's "ports only" cell now says what it means: contracts-layer ports the
kernel also consumes, not permission to name a kernel trait.
Two cost corrections recorded for the remaining families:
`host_runtime -> extension_support` is not divisible family-by-family (mod.rs
holds it via `extension_support::coding`), and
`host_runtime -> ironclaw_extensions` is not reachable by this row at all.
PATH_TERM_COLLISIONS shrinks by two: the installer's github carve-outs now sit
inside a scan-exempt crate.
Test accounting (un-masking discipline), unfiltered `--list` over both crates:
1398 -> 1398, with exactly two tests renamed by module path and none lost.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…nothing Review asked why the migrated docker_security test can pass with no daemon. The skip is pre-existing (the file differs from its pre-merge original by one import line); WS3 only enrolled it in the required Rust e2e lane, where it was not run at all before. The real defect the question surfaced is worse and also pre-existing: this crate's tests/support/docker_gate.rs states that IRONCLAW_REQUIRE_DOCKER_TESTS=1 makes a missing daemon a hard failure and that "CI sets this" -- and nothing sets it. Repo-wide the name occurs only in docker_gate.rs and attribution_tests.rs, here and on main. So every real-Docker test in the crate skips-and-passes everywhere, which is exactly the gap the gate's own comment says let sandbox security bugs ship unnoticed. docker_security.rs additionally open-codes its own check rather than using the gate, so it would stay fail-open even once something did set the variable. Recorded rather than fixed: setting the variable is a CI-behavior change that would hard-fail any lane without a daemon or the ironclaw-worker image, which is not verifiable from inside a move PR whose evidence claim is behavior preservation. Filed as the #6945 guardrail-claim-vs-reality class with the two-part fix stated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The crate's CLAUDE.md said "first-party runtime tools belong under `first_party_tools/`" without saying that only the host half does. WS3 moves each tool's executor into `ironclaw_extension_support`, which may not name this crate, so the rule now names both halves and points at the skill-install family as the worked example. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The moved executor returns `SkillManagementCapabilityError`, and routing it through `skill_management_error` would have added a `debug!` line to a path that had none before the move. A move-only change must not add one, so the install-input arm maps the kind directly and the `dispatch` arm keeps the record it already had. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…move The ratchet does not run on `pull_request` (`reborn_pr_test_plan.py:21`; issue #7036), so this PR's green checks were not evidence on this axis. A full-plan `workflow_dispatch` run on this exact head reported: RATCHET FAIL: ironclaw_host_runtime observed: 88.59% (20485 / 23124 lines) floor: 88.23% ... floor_covered_lines: 20538 (effective floor 20518) The percentage went UP while `floor_covered_lines` went DOWN — shedding well-covered code lowers the absolute numerator, which is a separate assertion from the percentage one. Re-captured to the observed numbers (floor raised 88.23 -> 88.59, not merely held). Verified locally against that run's own merged lcov artifact: ENFORCING mode, 17 PASS / 0 FAIL, exit 0. run: https://github.com/nearai/ironclaw/actions/runs/30858257594 head: e07b3b0 The destination crate is deliberately not floored, because it cannot be: every crate under `crates/extensions/` is invisible to the coverage tooling — `reborn_coverage_lcov.py:19`'s CRATE_RE still requires a crate directory directly under `crates/`, which #7037's colocation broke. Filed as #7083 with the measurement; the global floor is left alone rather than re-captured onto that hole. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CHECKLIST WS4 + WS10 `wit/` rows. `wit/{tool,channel}.wit` moves from the
repo root to `crates/ironclaw_wasm/wit/` — the crate that owns the ABI —
per PROPOSAL §6.6.1. Behavior-free: same bytes, same generated bindings.
Wave-3 coordinates: the docs write the destination as
`crates/lanes/ironclaw_wasm/wit/`, but `crates/lanes/` does not exist until
WS7. Because the files now sit *inside* the crate, the WS7 family move
carries them with no further path edit anywhere — which is the whole point
of putting them there.
Ten wit-bindgen `path:` args repointed (the host plus nine guests: six under
`crates/extensions/packages/*/wasm-src/`, three under `test-tools/*/wasm-src/`
— the CHECKLIST row said six). All nine guests verified building against the
moved WIT on wasm32-wasip2.
The four `include_str!` readers of the ABI text do NOT get repointed
literals. Doing that would turn the two `ironclaw_host_runtime` sites from
repo-root reach-ins into *cross-crate* ones — §11.2.7's strict class, the
one WS2 turns into hard failures — taking the scan from 19 to 21 while
ticking a box that says "§11.2.7 scan passes". Instead the ABI text gets one
owner, `ironclaw_wasm::TOOL_WIT` (`src/config.rs`, beside `WIT_TOOL_VERSION`),
and all four sites read the const over cargo edges that already exist.
Measured with the scan: 133 -> 129 escaping sites, cross-crate 19 -> 19,
zero `wit/` entries remaining.
Path-keyed gates repointed: `scripts/check-version-bumps.sh` (both ABI
paths), `.githooks/pre-commit`, and `platform-and-compat.yml`'s
`has_direct_wasm_abi_risk` filter — where the bare `wit/` alternative is
*deleted* rather than rewritten, because the filter's existing
`crates/([^/]+/)*ironclaw_wasm/` alternative already matches both the
Wave-3 and the WS7 location. `scripts/ci/ws12_workflow_contracts.py`
anchored on that deleted string, so its anchor moves to
`build-wasm-extensions` and its in-scope probe now pins both locations.
`Dockerfile` loses two `COPY wit/ wit/` lines in the planner and builder
stages: both already run `COPY crates/ crates/`, so the files arrive with
the crate and the old line would COPY a path that no longer exists.
Docs: the WS4 row's `crates/lanes/wit/` destination was the only doc site
placing the directory beside the crate rather than inside it; corrected
there and in README's tree, with dated amendments in CHECKLIST, PROPOSAL
§6.6.1 and PLAN Wave 3 recording what the move found.
Test accounting (unfiltered `--list`, name-by-name, quiescent tree):
ironclaw_wasm 51 -> 51, ironclaw_host_runtime 1246 -> 1246,
ironclaw_architecture 198 -> 198. Zero diff, no test edited for content.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Forced by the previous commit, not incidental to it. `scripts/ci/check-wasm-artifact-freshness.py` keys each package's committed `wasm/<name>.wasm` to a digest of the `wasm-src/` tree that produced it, so editing a guest's `wit_bindgen::generate!` `path:` — which the `wit/` move requires in all six shipped guests — invalidates the recorded digest and fails the gate. The gate's own contract forbids the shortcut: "Re-record only after `./scripts/build-wasm-extensions.sh --first-party` and committing the rebuilt artifact — the digest asserts a claim about the artifact, and updating it without rebuilding launders a stale one." So the artifacts are genuinely rebuilt (`--first-party`, exit 0, 6 OK / 2 host-native SKIP), not re-recorded in place. Byte sizes move by more than the source change accounts for because these builds are not reproducible by design — the guests pin no toolchain and resolve their own `Cargo.lock` at build time, which is the documented reason the gate hashes sources rather than artifact bytes. Verified: `check-wasm-artifact-freshness.py` OK (6 packages), and `cargo test -p ironclaw_extension_support` green (102/46/4) — that crate `include_bytes!`s these artifacts, so it exercises the rebuilt components. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… edits The `wit/` move had to rebuild six shipped WASM binaries because `check-wasm-artifact-freshness.py` digests each guest's whole `wasm-src/` tree. WS7 hits the same wall from the other direction: the six package guests reach the ABI across two trees, so moving either `ironclaw_wasm` or `extensions/packages` rewrites all six `path:` literals and forces the same rebuild. Recorded on CHECKLIST WS10's `wit/` row (point 6), on the loud-path-pattern row that owns the WS7 repoint (also corrected six -> nine guests there), and on PLAN's Wave 5 block with the cheap mitigation: move the two crates in one PR and pay it once. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
# Conflicts: # docs/reborn/target-architecture/CHECKLIST.md
Reconciles this lane with #7064, which landed the parallel WS3/WS4 runner sheds and edited the same coordination files. Five conflicts, all in shared coordination surfaces — no code move in this PR was altered (all 117 PR-only files are byte-identical to the pre-merge tip; the 5 apparent diffs are deletions absent on both sides). - `reborn_dependency_boundaries.rs`: the array auto-merged to the union of both sides' removals; only `WS0_LAYER_MATRIX_EXCEPTION_BASELINE` conflicted. Recomputed as `len()` of the merged list — 13 base, minus #7064's three (`hooks -> wasm_limiter`, `runner -> agent_loop`, `runner -> loop_host`) and this PR's three (`mcp -> extensions`, `scripts -> extensions`, `scripts -> resources`), plus this PR's justified `ironclaw_sandbox -> ironclaw_resources` = **8**. Counted by parsing only the entries between the const and its closing `];`, so the struct definition and the four test fixtures are excluded. - `loop_host/Cargo.toml`: both sides added a `[dev-dependencies]` line; kept both (`http` from main, `ironclaw_sandbox` from this PR). - `CHECKLIST.md`: kept both dated amendments in date order — #7064's `13 -> 10` and this PR's, with its count corrected from the authored-in-isolation `11` to the merged `8` exactly as the §11.2.2 row instructs. Also kept this PR's two new coverage-gate rows beside main's amended loud-inventory row. - `reborn_pr_test_plan.py`: both sides made the same `.claude/` fix; took main's landed wording (`startswith` makes tuple order irrelevant). - `test_reborn_pr_test_plan.py`: union of both sides' new tests, no name collisions — 46 tests pass. Also repointed the one PR-authored `changed-coverage-exemptions.toml` entry the merge shifted: `hosted_mcp.rs` moved +1 because main added a doc-comment line, so its line-keyed exemption now resolves to byte-identical source lines. The other stale entries in that manifest are inherited and already stale on main; left untouched. Verified: architecture suite 206 passed / 0 failed, `cargo check --all-targets` clean, `cargo fmt` a no-op, zero conflict markers, and both sides' `coverage-floor.toml` recaptures intact (runner 82.53 from #7064; host_runtime 88.65 and the new ironclaw_sandbox 87.09 from this PR). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reconciles this lane with #7064, which landed the parallel WS3/WS4 runner sheds and edited the same coordination files. One conflict: `WS0_LAYER_MATRIX_EXCEPTION_BASELINE` in `reborn_dependency_boundaries.rs`. The array itself auto-merged to the union of both sides' removals — #7064's three (`hooks -> wasm_limiter`, `runner -> agent_loop`, `runner -> loop_host`) and this PR's one (`host_runtime -> skills`). Recomputed the constant as `len()` of the merged list: main is at 10, so this slice takes it to **9**. Counted by parsing only the entries between the const and its closing `];`, which excludes the struct definition and the four test fixtures. This slice was authored off 13 and computed `13 -> 12` in isolation; the ratchet narrative and the two doc rows that quoted that figure (PLAN's "First-party tools" bullet and CHECKLIST's W7-progress row) now read `10 -> 9` and record why, per the union rule on the CHECKLIST §11.2.2 row. The edge deleted is unchanged; only the total moved. Everything else auto-merged and was verified rather than assumed: both sides' `coverage-floor.toml` recaptures are intact (runner 82.53 and the new `ironclaw_loop_host` 90.89 from #7064; `host_runtime` 88.59 from this PR), and both sides' dated amendments survive in CHECKLIST, PROPOSAL and PLAN. All 14 PR-only files are byte-identical to the pre-merge tip, so no code move was altered. Verified: architecture suite 206 passed / 0 failed, `cargo check --all-targets` clean, `cargo fmt` a no-op, zero conflict markers, and the changed-coverage manifest validates with no exemption stranded or shifted by this merge. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`Detect Reborn test scope` exits 1 on any pull request whose diff holds a path `reborn_pr_test_plan.py` has no rule for, which made this PR unmergeable: it must edit `Dockerfile` (the moved directory's `COPY wit/ wit/` no longer resolves) and `scripts/check-version-bumps.sh` (the ABI gate would otherwise grep dead paths and silently stop enforcing). 18 of its 46 paths were unclassified. Same class as the `.claude/` gap #7064 fixed, and classified the same way — one rule per class, recorded beside the constant: * `Dockerfile` / `.dockerignore` — `platform-and-compat.yml` keys `has_docker_risk` off exactly this pair and owns the image build. * `.githooks/**` — Code Style triggers on the tree and lints its contents (`test-ci-comm-locale-pin.sh`); no Reborn lane runs a hook. * `scripts/{build-wasm-extensions,check-version-bumps}.sh` — `platform-and-compat.yml`'s `has_direct_wasm_abi_risk` classifier both scopes and runs them. * markdown owned by no crate (`crates/AGENTS.md`, `test-tools/README.md`) — prose, like `docs/` and `.claude/`. A crate-resident doc still selects its own crate's lane. The first-party extension package assets are deliberately NOT ignored. `crates/extensions/packages/*/wasm/*.wasm` is a shipped artifact that `ironclaw_extension_support` embeds with `include_bytes!`, and `test-tools/*/manifest.toml` is `include_str!`d by `ironclaw_extension_host`. Calling either prose would convert today's loud failure into a silent under-schedule of a change to production output — the WS10 failure mode. `EMBEDDED_ASSET_OWNERS` routes each tree to the crate that compiles it instead, so this PR now additionally schedules `ironclaw_extension_{support,host,manager}`: the crates that consume the six rebuilt WASM artifacts. Also fixes #7085 in a file this PR already touches. The WIT version extractors used the GNU-only BRE `\+`, so on BSD sed (macOS) they matched nothing, and because the `WIT_TOOL_VERSION` cross-check is guarded on a non-empty version the hook printed "All version checks passed" having compared nothing. `[[:space:]][[:space:]]*` is identical under GNU sed, so the enforced Linux CI lane is unchanged; verified on BSD sed that both `wit/tool.wit` (0.3.0) and `wit/channel.wit` (0.3.1) now extract. Regression tests: every classified class gets a case in `test_reborn_pr_test_plan.py`, including the paired assertion that the embedded assets *select a lane* rather than merely being accepted (the inverse of the `.claude/` prose test), and a staleness pin that fails if an asset tree or its owning crate moves. All ten new cases fail against the planner on `main`. `test_unclassified_build_input_fails_fast` moves off `Dockerfile` onto a still-undecided input so the fail-closed arm stays exercised. Refs #7087, #7085 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ners (WS3)
`crates/ironclaw_host_runtime/src/obligations.rs` was 3,122 lines fusing the
three owners PROPOSAL §6.5.9 charters separately, held apart only by an
`// arch-exempt: large_file` waiver. It is now one module per owner:
- `obligations::handler` — which obligations apply and what each does
before/after dispatch, plus the audit/redaction/ceiling/mount validation.
- `obligations::staged_handoffs` — material staged for a later consumer:
the runtime-secret and network-policy stores and the credential-account
resolver port.
- `obligations::process_store` — post-start handoff discard and reservation
reconciliation.
- `obligations::mod` — only `BuiltinObligationServices`, the assembly seam,
and deliberately the one place naming all three at once.
Every module is under the 1,500-line gate, so the waiver is deleted rather
than carried forward: re-fusing the owners now trips `pre-commit-safety.sh`.
`mod obligations;` stays private and the crate's `pub use obligations::{…}`
names are unchanged, so no consumer outside the crate sees this.
Behavior-free. Cross-owner access is `pub(super)` (three methods), not
`pub(crate)`. The split revealed one narrowing in the other direction:
`secret_present` was `pub(crate)` with no caller outside its own file and is
now private.
Also from the same CHECKLIST row, the bounded half of "shrink
`services/builder.rs` toward composition-facing factories": three builder
methods whose only callers are inside the crate's `src` narrow to
`pub(crate)`. The rest of that clause is measured and deferred in the
CHECKLIST amendment — 17 methods need a `test-support` cargo feature, three
are callerless and belong to WS8, and the remaining 33 are a redesign of the
fluent surface rather than a shrink of it. `+production_wiring` is refuted
there: it is readiness diagnostics, not assembly.
Two loud path-keyed gates fired and were repointed, not relaxed:
`reborn_host_runtime_services_do_not_expose_lower_substrate_handles` now
scans the whole `obligations/` directory and asserts it read ≥ 4 files
(`collect_runtime_rs` returns a count; both its callers now assert non-zero),
and `reborn_struct_test_support_ratchet`'s frozen per-file count moves to
`staged_handoffs.rs` with its count unchanged at 1.
Test accounting (un-masking discipline): `cargo test -p ironclaw_host_runtime
--all-targets -- --list` is 1,246 before and 1,246 after, name-by-name
identical — zero added, removed or renamed. `LAYER_MATRIX_EXCEPTIONS` is 10
before and after; an intra-crate split cannot move the register.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…t_contracts port (WS3) `ironclaw_operator` is a products-tier crate and held `ironclaw_secrets`, the substrate that owns CAS one-shot leases, AAD/crypto and the OS keychain master key. PROPOSAL §8.2's product row says the products tier loses that edge, and §12.1b requires the port replacement to land before the edge is removed. Both happen here, in that order. - Port: `ironclaw_product_contracts::operator_secrets::OperatorSecretValueStore`. - Implementor: `ironclaw_reborn_composition::RuntimeOperatorSecretValueStore`, the same placement as `OperatorStatusService` — assembly is the only layer that may name both a products-tier port and a substrate. Registered in `INVERTED_PORTS` beside it. - `ironclaw_secrets` is gone from the operator manifest under every dependency kind, and `"ironclaw_secrets"` is now in the crate's `boundary_rules()` forbidden list. That gate's comment previously said the entry was deliberately absent because "the row owns it"; the row now owns it. The port is deliberately narrower than the substrate, so this is a tightening rather than a relocation: it takes no `ResourceScope` (the implementor fixes the operator scope, where the caller used to pass one), exposes no lease/consume protocol, and carries only a `&'static str` classification instead of the substrate's error `Display` — asserted, including that the backend message and the handle name are both absent from what crosses. Two tests travelled with the behavior rather than being pointed at a fake: `read_is_repeatable_across_reloads` (repeatability is a property of the lease protocol) and the #4673 production-store reproduction (its value is wiring the store exactly as production does, which now means the real store *behind the adapter*). Two `FaultInjecting`-over-real-store fixtures became per-operation port fakes, with the substrate error mapping re-pinned at the adapter; a third assertion got stronger — batched-vs-N+1 stored-key lookup is now observed at the port rather than by counting filesystem ops. Test accounting: operator 154 -> 153, product_contracts 142 -> 143, composition 937 -> 942 with zero removed; name-by-name diffs on a quiescent tree. Two findings the row could not have anticipated, both recorded in the CHECKLIST amendment: - The `webui` half of the row was already closed and was never a production edge. `ironclaw_secrets` has been a dev-dependency of `ironclaw_webui` since the commit that added it (#6619), both src mentions are `#[cfg(test)]`, and webui's boundary rule already forbade it. - `ironclaw_extension_manager` (layer `products`) still holds a normal `ironclaw_secrets` edge in `admin_configuration.rs`. §8.2 covers it; the row does not, because the crate landed with WS2.4 after the row was written, and the substrate sits in the service's type parameters so it is not a like-for-like swap. Filed as #7095. `LAYER_MATRIX_EXCEPTIONS` is 10 before and after: `products -> substrates` is matrix-legal, so this edge was always an §8.2 rule and never a layer exception. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Review asked why the required Rust e2e lane can report `docker_security` as passing with no daemon. Half of that is #7081 (nothing sets IRONCLAW_REQUIRE_DOCKER_TESTS=1, so the switch is inert) and is not fixable from here -- arming it hard-fails any lane lacking a daemon or the worker image, which needs a runner guaranteed to have both. The other half is fixable here and is fixed: docker_security.rs open-coded its own `docker version` / `image inspect` checks with three bare `return`s, so it sat entirely outside docker_gate and would have stayed fail-open even once something did set the variable. It now takes both preconditions from docker_gate::{docker_available, docker_image_available} and skips with the visible `SKIP:` line that gate's module doc requires. Measured, same machine, image absent: before, IRONCLAW_REQUIRE_DOCKER_TESTS=1 -> "skipping ..." / 1 passed after, IRONCLAW_REQUIRE_DOCKER_TESTS=1 -> panic at docker_gate.rs:74 / FAILED after, variable unset -> "SKIP: ..." / 1 passed The third line is the no-op proof: the variable is set nowhere in this tree or on main, so no lane's behavior changes today. The daemon-down path already reached the image check and skipped there, so the outcome is identical; only the branch it takes differs. Two stale comments in docker_gate.rs corrected with it (they claimed docker_security used its own gate, and that docker_image_available had no consumer), and the crate's Known debt entry now splits the done half from the #7081 half instead of describing both as open. cargo test -p ironclaw_sandbox: 193 passed, 0 failed cargo clippy -p ironclaw_sandbox --tests --all-features -- -D warnings: exit 0 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two review findings, both correct, both artifacts of this PR's own renames.
1. engine-v2-to-reborn-parity.md note 4 read "a native script/software
execution lane (`ironclaw_sandbox`, `RuntimeKind::Script`) sandboxed via
`ironclaw_sandbox`" -- self-referential after the merge collapsed
ironclaw_scripts and ironclaw_process_sandbox into one crate, and it
contradicts note 5 four paragraphs down ("no production execution backend
is wired for it"). Re-stated as the typed runtime contract it is, citing
the measurement: `with_script_runtime` has zero production callers
(`rg` finds only the builder itself, docs, and 30 test call sites).
2. CHECKLIST WS10 ratchet note 2 said "raise the percentage floor ...; only
the line count should fall". That generalises WS3's sandbox merge, where
observed coverage happened to rise. It is wrong as guidance for WS7, and
the counterexample is in this same file: the 2026-08-03 entry from #7064
records ironclaw_runner falling 85.55% -> 82.53% because the shed removed
the crate's better-covered half, holding the floor, and RATCHET FAILing in
the merge queue. Note 2 now says re-capture from the merged artifact, and
lower only with that entry's move-not-regression counterfactual (add the
moved files back, confirm the union clears the old floor, plus a zero-tests-
lost name set-diff).
cargo test -p ironclaw_architecture: 32 targets, 206 passed, 0 failed
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three review findings on the `wit/` move, each verified before it was acted on.
1. `ws12_workflow_contracts.py` probed `crates/ironclaw_wasm/wit/host.wit` and
its nested twin. No `host.wit` exists in this repository — `git ls-files
'*.wit'` returns only `tool.wit` and `channel.wit` — so both probes sat
under the `crates/([^/]+/)*ironclaw_wasm/` alternative and re-asserted the
crate-name term while saying nothing about the canonical ABI contracts. In
a validator whose stated design is "probe derived from reality rather than
from a guessed layout", a fabricated filename is a defect on its own terms.
Replaced with a `crate_globs` entry, `("ironclaw_wasm", "wit/*.wit")`, which
discovers the contracts on disk, requires each in scope, and synthesises the
nested WS7 form — so a third contract, or the directory leaving the crate,
fails the pin instead of passing on a stale name. Verified non-vacuous:
narrowing the workflow alternative to `.../ironclaw_wasm/src/` now reports
`tool.wit`, `channel.wit` and the nested probe as out of scope.
2. The embedded-asset routing test substituted `alpha`/`beta` owners so it
could reuse the synthetic workspace. That exercised the real prefix strings
through the real routing, but left the prefix->owner *pairing* — the table's
entire semantic content — asserted nowhere: swapping
`ironclaw_extension_support` and `ironclaw_extension_host` passed. Fixed in
two halves. The routing test now drives the real `EMBEDDED_ASSET_OWNERS`
against a workspace carrying the real owners' names and real manifest paths
(the synthetic one could not: `build_plan` rejects a changed package outside
the canonical set), asserting the real owner is selected. And the not-stale
test now derives the same pairing from the tree instead of restating the
constant: it resolves every literal `include_str!`/`include_bytes!` in every
workspace crate through `crate_tree`, keeps the targets no crate owns — the
ones that actually reach the table — and asserts that every crate compiling
one of them is the routed owner or a dependent of it.
That surfaced a property worth pinning: `crates/extensions/packages/` is
embedded by four crates, not one. `ironclaw_extension_host`,
`ironclaw_extension_manager` and `ironclaw_reborn_composition` reach into it
alongside `ironclaw_extension_support`, and routing to the support crate
covers them only because each depends on it. If that edge goes, a shipped
artifact change stops scheduling a crate that embeds it — the silent
under-schedule the table exists to prevent.
Regression coverage verified red by sabotage, all three wrong tables:
owners swapped (7 failures), `packages/` -> `ironclaw_llm` ("embeds nothing
from it"), and the hardest case, `packages/` -> `ironclaw_reborn_composition`
— a real embedder that the other embedders do not depend on
("...does not depend on..., so routing there never schedules it").
3. CHECKLIST WS10 claimed each of the nine `wit_bindgen` guest edits forces a
committed WASM artifact rebuild. Only six do:
`scripts/ci/check-wasm-artifact-freshness.py` scans
`crates/extensions/packages/*/wasm-src` alone, `wasm-src-digests.toml` holds
exactly six entries, and `git ls-files '*.wasm'` returns exactly those six.
The three `test-tools/*/wasm-src/` guests commit no artifact; the tenth site
is the host's `bindings.rs`, not a guest. Corrected, and the `wit/` row now
states the boundary rather than implying it.
Guest paths, `wit/` contents and the six rebuilt artifacts are untouched.
Verified: `test_reborn_pr_test_plan.py` 46/46, `test_ws12_workflow_contracts.py`
25/25, `ws12_workflow_contracts.py` green on the real tree,
`cargo test -p ironclaw_architecture` 206/206 across 32 binaries.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This was referenced Aug 5, 2026
Closed
pull Bot
pushed a commit
to bhardwajRahul/ironclaw
that referenced
this pull request
Aug 5, 2026
…factory port (nearai#7202) * refactor(contracts): move extension runtime descriptors to a neutral contract (WS3) Deletes the two `-> ironclaw_extensions` layer-matrix exceptions (`ironclaw_mcp`, `ironclaw_scripts`) by giving the runtimes-layer lanes a contracts home for the descriptors they read, instead of the registry crate they may not depend on. Exceptions 13 -> 11; baseline lowered in the same change. Moved to `ironclaw_extension_contracts`: - `runtime::{ExtensionRuntime, ExtensionAssetPath, ExtensionAssetPathError}` - `hosted_mcp::{HostedMcpDiscoveredTool, HostedMcpDiscoveredToolAnnotations}` `ExtensionPackage`/`ExtensionManifest` deliberately stay in `ironclaw_extensions`: they carry the whole parsed manifest tree and a `PackageRootBinding` typed on `ironclaw_filesystem::VirtualPath`, which the §11.2.3 contracts-purity allowlist (`{ironclaw_host_api}` only) forbids the contracts crate from naming. Measured instead: both lanes read exactly three things off the package — `id`, `capabilities`, `manifest.runtime` — so the lane request structs now take those three and the caller (which owns the package) projects them. Also repointed `ResourceReceipt` to its real owner: `ironclaw_resources` only re-exports `ironclaw_host_api::resource::ResourceReceipt`, so the lanes' import was a §11.2.4 two-import-paths hop, not a dependency. No `pub use` shims (§11.3): every consumer is repointed in this change, and `resolve_under` becomes the free function `ironclaw_extensions::resolve_asset_under` because the orphan rule forbids an inherent impl on the moved type. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(sandbox): merge the sandbox lane into one crate (WS3) Creates `ironclaw_sandbox` (runtimes) from the three halves of "run an already-authorized command away from the host", and deletes the two crates PROPOSAL §6.6.4 marks for merge: - `ironclaw_process_sandbox` (plan contract) -> `src/plan.rs`, `src/validation.rs` - `ironclaw_host_runtime::sandbox_process` -> `src/sandbox_process/**` - `ironclaw_scripts` (script lane + Docker path) -> `src/script.rs` The kernel sheds the Docker/CA cone: `bollard`, `rcgen`, `x509-parser` and `time` are gone from `ironclaw_host_runtime`'s manifest, and `bollard`/`rcgen` are now declared by exactly one crate in the workspace. Two migration details PROPOSAL §6.6.4 and CHECKLIST WS10 call load-bearing: - `PROCESS_SANDBOX_CAPABILITY_ID` -> `ironclaw_host_api::capability`, so `ironclaw_loop_host` drops its lane dependency (production dep gone; a dev-dep remains for the tests that build plans). - `SandboxCommandTransport` -> `ironclaw_host_api::process`, with the shapes it names (`CommandExecutionRequest`/`Output`, `RuntimeProcessError`, `SavedCommandOutput`, `SavedCommandOutputSanitization`). Without this the runtimes-layer lane could not implement what the kernel consumes. Enumerating gates were repointed, never relaxed: the specificity carve-outs and the struct/test-support ratchet entries moved with their files (both baselines unchanged at 129 and their prior values), the panic-gate baseline row moved, `reborn-crate-test-buckets.sh` registers the new crate, and the three `reborn-e2e-rust.sh` script selectors follow the tests (plus `docker_security`, which had no selector before). One gate would have gone silently vacuous and was fixed rather than moved: the script-lane surface scan in `reborn_dependency_boundaries.rs` read a hardcoded `src/lib.rs`, which after the merge no longer holds the lane. It now scans the whole crate source tree with a fatal-read walk and a non-vacuity assertion. One deletion, recorded: `RebornScopedSandboxCommandTransport::into_process_port` returned a kernel type a runtimes crate may not name. It had zero callers workspace-wide; the kernel wraps the transport, which is the direction the port inversion requires. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): record the WS3 corrections with their evidence Three dated amendments, each quoting the text it replaces: 1. CHECKLIST WS3 sandbox row + PROPOSAL §6.6.4 — "all pieces currently unwired/test-only" is REFUTED. Three production paths cross the merged crate (spawn-path plan validation, the process_executor routing check, and the saved-command-output scope digest). The accurate claim is narrower: no production *execution backend*. Behavior preservation is therefore argued at the diff (11 of 26 moved files byte-identical, 9 more differing by one import line, +63/-36 overall), not inferred from deadness. 2. CHECKLIST WS3 mcp row + PROPOSAL §6.6.3 — the prior wave's "structurally blocked" finding is half right, and the wrong half is load-bearing: only `ExtensionPackage` is un-absorbable, and no lane ever needed it (both read `id`, `capabilities`, `manifest.runtime` and nothing else). The registry half of the flip is done; the `resources` half is refuted as phrased — the estimate/usage vocabulary the row asks about is already in `host_api::resource` and already imported from there, while the real blocker is the `ResourceGovernor` authority port and `ResourceError`'s denial cone. 3. Recorded as a structural finding, not a note: the sandbox row and the mcp row are ONE problem. `ironclaw_scripts` imports the identical DTO set, so the merge alone deletes zero exceptions and only the mcp carve-out lets either lane shed the registry edge. Also reconciled: PROPOSAL §6.1.2's as-built inventory gains the two modules WS3 landed (and states why `ExtensionPackage` stayed); §2's package count 66 -> 65; the §9 disposition rows for `ironclaw_scripts`/`ironclaw_process_sandbox`/ `ironclaw_mcp`; the §11.2.2 ratchet rows (13 -> 11); the WS3 verify row; the stale WS1.3 sentence asserting the blocker as settled fact; and `reborn_restructure_baselines.rs`'s doc table, which still read 15. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(sandbox): drop imports the merge left unused `process_port.rs` no longer names `MountView` or `thiserror::Error` (both went to `host_api::process` with the types that used them), and `sandbox_process.rs` no longer needs `sync::Arc` after `into_process_port` was deleted. Found by per-crate `clippy --all-targets --all-features -D warnings`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): let the Reborn PR planner plan guidance edits and crate deletions Three fail-closed gaps in `reborn_pr_test_plan.py`, all hit by this PR and all live on `main` today — any PR with the same change shape is unplannable. 1. `.claude/**` was unclassified, so the planner refused outright. It is agent guidance in exactly the sense `docs/**` is human guidance: no Rust test reads either as data (the only in-tree references are prose citations in test doc comments). Added to `IGNORED_PREFIXES`. Without this, "guidance travels with the change" — the restructure's own discipline — cannot be satisfied in a single PR. 2. `crates/AGENTS.md`, `crates/README.md`, `crates/Architecture.md` raised "unmapped crate path": they sit under `crates/` but belong to no package. Now classified as crate-tree prose, matched by "Markdown no package directory owns" so a genuinely unmapped crate path is unaffected. 3. An unmapped crate path used to raise. `git diff` reports a deleted crate's old paths and CI feeds the planner that diff, so **every crate deletion or rename was unplannable** — including the six deletions PROPOSAL §2 plans. It now widens to the exhaustive plan. This is a semantic change and it is the safe direction: the full plan is a superset of any narrowing, so an unattributable path can never cause under-selection, whereas refusing to plan blocks the PR instead of protecting it. Malformed input is still rejected by the unclassified-path branch. Each lands with fixtures per WS10's rule, positive and negative: guidance paths select nothing while non-guidance paths still fail closed; crate-tree prose selects nothing while crate *code* under the same unmapped directory widens to `full` (so the Markdown carve-out cannot swallow code). The pre-existing `test_unmapped_crate_path_fails_fast` is renamed and rewritten to pin the new contract rather than deleted. Verified against this PR's real 130-path diff: the planner returns `mode: full`, and the workflow's own exhaustiveness guard passes on that output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(arch): give the retained resource exceptions an owning issue, not a wave Review (#7065) caught that both surviving `-> ironclaw_resources` exceptions declared `removes_in = "WS3"` — the wave this PR *is*, which does not remove them. That is precisely the defect §11.2.2 already records against `conversations -> turns` ("`removes_in = "WS5"` and WS5 has partly shipped without it falling"), and it would have been repeated here. Both now point at issue #7067, which owns the design work that actually clears them: replacing the `ResourceGovernor` dependency with a narrow reserve/reconcile/release port. The issue carries the measurements — 3 of 10 methods used, zero implementors, and the `ResourceError` denial cone — plus the two open questions (error shape, port home) that make it a design slice rather than a move. An owning issue is also what §11.2.2 asks for and what the ratchet still cannot enforce (there is no `owning_issue` field yet), so this is the strongest form currently expressible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(contracts): pin the asset-path validator that moved into extension_contracts `validate_asset_path` moved here with `ExtensionAssetPath`, the type it constructs. In `ironclaw_extensions` it was only ever reached indirectly through manifest parsing, so its six rejection branches had no direct test — and a contracts crate that carries validation owes that validation one. Two tests: every reject branch with its exact reason and `Display` output (empty, NUL/control, URL, absolute, Windows drive and backslash, and the empty/`.`/`..` segment cases) plus the manifest-relative shapes that must keep being accepted; and `ExtensionRuntime::kind()` over all five variants, since that projection is what every lane uses to reject a runtime it does not serve. Also removes a changed-line coverage risk this PR would otherwise carry into the merge queue: the gate does not run on ordinary PRs (#7036), so ~100 newly-added lines of validator would first be measured where a failure is expensive to diagnose. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(coverage): re-capture the host_runtime floor and floor the new sandbox lane `RATCHET FAIL: ironclaw_host_runtime` — observed 18854 covered vs a `floor_covered_lines` of 20538. This is the shrinkage case the ratchet's own "To fix" text describes, not a coverage regression: `sandbox_process/**` moved to `ironclaw_sandbox`, so the crate's denominator fell 23277 -> 21267 (-2010 instrumented lines) and its covered lines fell with it. The percentage floor is **raised, not lowered**: observed 88.65% against an old floor of 88.23%, so the entry now reads 88.65. Only the absolute line count moves down, and it must — those lines are no longer in this crate. To keep that from being a net loss of protection, `ironclaw_sandbox` is floored on arrival at its observed 87.09% (3185 / 3657). This is a net *increase* in ratchet coverage: neither `ironclaw_scripts` nor `ironclaw_process_sandbox` was ever floored, and the `sandbox_process` half was protected only as part of host_runtime's line count, which this PR necessarily reduces. Floored crates 16 -> 17. Verified by replaying the ratchet arithmetic against CI's observed numbers: both crates pass on percentage and on covered lines. Numbers taken from the failing run's own report (job 91740733521), which is the authority for this gate. The `Tests (Reborn)` roll-up failed solely on this sub-job ("coverage-report result 'failure' did not match planned=true"); no other lane failed — 50 pass, 2 fail, both this root cause and its roll-up. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): record the coverage ratchet as a move-sensitive gate WS3 hit a gate no move row had named. `tests/integration/coverage-floor.toml` is keyed on crate identity plus absolute covered-line counts, so it is invisible to WS10's path-keyed gate audit and yet it fails on every crate move, merge, rename, or family `git mv` that shifts instrumented lines between crates — as it did here, while the percentage floor was *improving*. Recorded on WS10 with the three rules WS7 will need: re-capture in the same PR, raise the percentage floor rather than leaving it, and floor the destination crate or the move silently drops that code out of the ratchet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(extension-manager): repoint ironhub onto the moved ExtensionAssetPath A semantic conflict the merge could not see: #6780 landed `ironhub/{package,catalog}.rs` importing `ExtensionAssetPath` from `ironclaw_extensions`, while this branch moved that type to `ironclaw_extension_contracts::runtime`. Different files, so git auto-merged cleanly and the breakage surfaced only at `cargo check`. Repointed both sites to the contracts crate (no shim, per §11.3). The manifest already named `ironclaw_extension_contracts`, so this is imports only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(coverage): exempt the WS3 move's no-region lines and record the gate The changed-lines coverage gate went red on four files while changed-line coverage was 95.35% against a 90% floor: the failure was its two fail-closed STRUCTURAL assertions, not any percentage. Every line below was derived by replaying scripts/ci/reborn_changed_coverage.py against this PR's own merged lcov (run 30831658659) with the base lcov the gate itself resolved (run 30828540055 @ b89fcd3575), until the replay reproduced the CI verdict byte-identically. Line numbers come from the gate's own `candidate_lines - mechanically_uninstrumentable_lines()`, not from the log. - host_api/src/process.rs (31 lines): new placement-neutral process vocabulary with no function body anywhere in the file; rustc emits no LCOV record for it at all. Same shape already exempted for product_contracts/loop_contracts. - extension_contracts/src/hosted_mcp.rs (12): field declarations of the two new tools/list descriptor structs. The file is plainly instrumented (191 DA, 164 hit), so this is a no-region artifact, not an instrumentation gap. - host_runtime/src/services/runtime_adapters.rs (13): continuation lines of three rewritten calls, all PROVEN EXECUTING by their region-start heads (lines 380/434/977 score 24/16/63 hits). The four genuinely-uncovered lines in the same rewrite are deliberately NOT exempted -- the gate already subtracts them as pre-existing debt inherited from base. - composition capability_host_tests/approval_gates.rs (6): type positions in a test double whose body region scores 1 hit. The last one is a finding, not just a waiver: that file is 100% test code behind `#[cfg(test)] mod capability_host_tests;`, but the gate's test_only_path() recognises /tests/, /test_support/, */tests.rs and *_tests.rs and NOT a cfg(test) module DIRECTORY, so it measures it as production. It is the only such directory in crates/ today. Docs (target-architecture, same PR per the docs-truth rule): - CHECKLIST WS10 gains the changed-lines gate beside the ratchet row, cross- referencing the WS2.1 note rather than restating it: percentages are not what fail a move; derive lines by byte-identical replay (--fetch-base-coverage silently degrades without --github-repo); and a stranded exemption path is an ABORT with no verdict, not a loud failure. - CHECKLIST WS10 exception-ratchet row: the constant was cited at :4063 and sits at :4164 -- corrected by removing the line pin, since the file is edited every wave. Records that the baseline is a UNION across parallel WS3 lanes. - families/contracts.md: records extension_contracts' new ownership of the runtime descriptor vocabulary -- the carve-out that let BOTH lanes drop the registry edge -- and the orphan-rule seam that keeps resolve_asset_under in the registry crate. - families/lanes.md: two "Never" claims were reading as satisfied when they are not. ironclaw_mcp's "never depends on the resource-governor crate directly" is refuted (the compiled edge survives; #7067 tracks the narrow port), and ironclaw_sandbox's "no direct process spawning outside the transport seam" is aspirational -- script.rs:454 still builds Command::new("docker"). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sandbox,mcp): correct the wiring inventory and record the projection cost Two review findings verified against the tree; three refuted with evidence in the PR threads. Valid — the sandbox wiring inventory was self-contradictory. `CLAUDE.md` said "Two production call paths ... and both are plan validation" directly above a list of THREE bullets, and `lib.rs` omitted the third entirely. The third is real and is not validation: `host_runtime/src/process_output.rs:482` derives the scoped saved-output directory through `RebornSandboxScopeKey::from_scope`. That inventory is what tells a future agent which paths are live, so an undercount invites deleting a production path as dead code. Both surfaces now say three and no longer claim they are all plan validation (the `loop_host` capability-id comparison never was either). Valid, and recorded rather than redesigned — the registry carve-out cost a type-level invariant. Replacing `package: &ExtensionPackage` with independent `extension` / `capabilities` / `runtime` borrows is what deleted the `mcp -> extensions` and `scripts -> extensions` exceptions, but it also means the type no longer guarantees the three came from one package. `execute_extension_json` re-checks the descriptor half (`descriptor.provider == extension`); the runtime half cannot be re-derived, because nothing in an `&ExtensionRuntime` names its owning extension. No caller can trip it today -- there is exactly one production caller (`runtime_adapters`) and it projects all three from one package in one expression -- so this is a latent structural weakening, not a live defect. Restoring the compile-time binding needs a sealed projection minted by the package owner; a check inside the lane cannot express it, and re-taking the registry edge would undo the carve-out. Both request types now carry the caller obligation in their field docs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(extensions): move the skill-install executor to extension_support (WS3) WS3's first-party-tools row, family 1 of 6: skill management / URL install. `skill_url_install.rs` and its `bundle`/`github`/`zip_bundle` submodules, plus the install-input normalizer, move out of `ironclaw_host_runtime::first_party_tools` into `ironclaw_extension_support::skills::{url_install, resolve_install_input}`, where the skill executor half already lived. Move-only: no behavior change, no test edited for content. `ironclaw_host_runtime -> ironclaw_skills` is deleted from LAYER_MATRIX_EXCEPTIONS — the edge is gone, not waived (exceptions 13 -> 12, WS0_LAYER_MATRIX_EXCEPTION_BASELINE drops with it). `ironclaw_skills` and `zip` survive as dev-dependencies for host_runtime's own tests; dev edges are outside the matrix by construction. Two doc ambiguities are resolved in the same diff, as dated PROPOSAL amendments quoting the text they replace: - §6.8.4's "the builtin first-party tool handlers absorbed from host_runtime/first_party_tools" contradicted §8.2's "kernel: ✗ (ports only)" row and the enforced BoundaryRule. Resolution: the seam splits executor from adapter — the executor moves behind a neutral request/error pair, the FirstPartyCapabilityHandler / CapabilityManifest / registry wiring stay host-side. Same shape the groupware and web-access tools already ship. - §8.2's "ports only" cell now says what it means: contracts-layer ports the kernel also consumes, not permission to name a kernel trait. Two cost corrections recorded for the remaining families: `host_runtime -> extension_support` is not divisible family-by-family (mod.rs holds it via `extension_support::coding`), and `host_runtime -> ironclaw_extensions` is not reachable by this row at all. PATH_TERM_COLLISIONS shrinks by two: the installer's github carve-outs now sit inside a scan-exempt crate. Test accounting (un-masking discipline), unfiltered `--list` over both crates: 1398 -> 1398, with exactly two tests renamed by module path and none lost. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sandbox): record that the Docker fail-closed switch is wired to nothing Review asked why the migrated docker_security test can pass with no daemon. The skip is pre-existing (the file differs from its pre-merge original by one import line); WS3 only enrolled it in the required Rust e2e lane, where it was not run at all before. The real defect the question surfaced is worse and also pre-existing: this crate's tests/support/docker_gate.rs states that IRONCLAW_REQUIRE_DOCKER_TESTS=1 makes a missing daemon a hard failure and that "CI sets this" -- and nothing sets it. Repo-wide the name occurs only in docker_gate.rs and attribution_tests.rs, here and on main. So every real-Docker test in the crate skips-and-passes everywhere, which is exactly the gap the gate's own comment says let sandbox security bugs ship unnoticed. docker_security.rs additionally open-codes its own check rather than using the gate, so it would stay fail-open even once something did set the variable. Recorded rather than fixed: setting the variable is a CI-behavior change that would hard-fail any lane without a daemon or the ironclaw-worker image, which is not verifiable from inside a move PR whose evidence claim is behavior preservation. Filed as the #6945 guardrail-claim-vs-reality class with the two-part fix stated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(host_runtime): record the executor/adapter seam in crate guidance The crate's CLAUDE.md said "first-party runtime tools belong under `first_party_tools/`" without saying that only the host half does. WS3 moves each tool's executor into `ironclaw_extension_support`, which may not name this crate, so the rule now names both halves and points at the skill-install family as the worked example. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(host_runtime): keep the install-input error path log-free The moved executor returns `SkillManagementCapabilityError`, and routing it through `skill_management_error` would have added a `debug!` line to a path that had none before the move. A move-only change must not add one, so the install-input arm maps the kind directly and the `dispatch` arm keeps the record it already had. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci(coverage): re-capture the host_runtime floor for the WS3 executor move The ratchet does not run on `pull_request` (`reborn_pr_test_plan.py:21`; issue #7036), so this PR's green checks were not evidence on this axis. A full-plan `workflow_dispatch` run on this exact head reported: RATCHET FAIL: ironclaw_host_runtime observed: 88.59% (20485 / 23124 lines) floor: 88.23% ... floor_covered_lines: 20538 (effective floor 20518) The percentage went UP while `floor_covered_lines` went DOWN — shedding well-covered code lowers the absolute numerator, which is a separate assertion from the percentage one. Re-captured to the observed numbers (floor raised 88.23 -> 88.59, not merely held). Verified locally against that run's own merged lcov artifact: ENFORCING mode, 17 PASS / 0 FAIL, exit 0. run: https://github.com/nearai/ironclaw/actions/runs/30858257594 head: e07b3b0299b0add11117e9591da71d46d7a7c832 The destination crate is deliberately not floored, because it cannot be: every crate under `crates/extensions/` is invisible to the coverage tooling — `reborn_coverage_lcov.py:19`'s CRATE_RE still requires a crate directory directly under `crates/`, which #7037's colocation broke. Filed as #7083 with the measurement; the global floor is left alone rather than re-captured onto that hole. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(wasm): move wit/ inside its owning crate (Wave 3) CHECKLIST WS4 + WS10 `wit/` rows. `wit/{tool,channel}.wit` moves from the repo root to `crates/ironclaw_wasm/wit/` — the crate that owns the ABI — per PROPOSAL §6.6.1. Behavior-free: same bytes, same generated bindings. Wave-3 coordinates: the docs write the destination as `crates/lanes/ironclaw_wasm/wit/`, but `crates/lanes/` does not exist until WS7. Because the files now sit *inside* the crate, the WS7 family move carries them with no further path edit anywhere — which is the whole point of putting them there. Ten wit-bindgen `path:` args repointed (the host plus nine guests: six under `crates/extensions/packages/*/wasm-src/`, three under `test-tools/*/wasm-src/` — the CHECKLIST row said six). All nine guests verified building against the moved WIT on wasm32-wasip2. The four `include_str!` readers of the ABI text do NOT get repointed literals. Doing that would turn the two `ironclaw_host_runtime` sites from repo-root reach-ins into *cross-crate* ones — §11.2.7's strict class, the one WS2 turns into hard failures — taking the scan from 19 to 21 while ticking a box that says "§11.2.7 scan passes". Instead the ABI text gets one owner, `ironclaw_wasm::TOOL_WIT` (`src/config.rs`, beside `WIT_TOOL_VERSION`), and all four sites read the const over cargo edges that already exist. Measured with the scan: 133 -> 129 escaping sites, cross-crate 19 -> 19, zero `wit/` entries remaining. Path-keyed gates repointed: `scripts/check-version-bumps.sh` (both ABI paths), `.githooks/pre-commit`, and `platform-and-compat.yml`'s `has_direct_wasm_abi_risk` filter — where the bare `wit/` alternative is *deleted* rather than rewritten, because the filter's existing `crates/([^/]+/)*ironclaw_wasm/` alternative already matches both the Wave-3 and the WS7 location. `scripts/ci/ws12_workflow_contracts.py` anchored on that deleted string, so its anchor moves to `build-wasm-extensions` and its in-scope probe now pins both locations. `Dockerfile` loses two `COPY wit/ wit/` lines in the planner and builder stages: both already run `COPY crates/ crates/`, so the files arrive with the crate and the old line would COPY a path that no longer exists. Docs: the WS4 row's `crates/lanes/wit/` destination was the only doc site placing the directory beside the crate rather than inside it; corrected there and in README's tree, with dated amendments in CHECKLIST, PROPOSAL §6.6.1 and PLAN Wave 3 recording what the move found. Test accounting (unfiltered `--list`, name-by-name, quiescent tree): ironclaw_wasm 51 -> 51, ironclaw_host_runtime 1246 -> 1246, ironclaw_architecture 198 -> 198. Zero diff, no test edited for content. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * build(wasm): rebuild first-party artifacts for the moved wit/ path Forced by the previous commit, not incidental to it. `scripts/ci/check-wasm-artifact-freshness.py` keys each package's committed `wasm/<name>.wasm` to a digest of the `wasm-src/` tree that produced it, so editing a guest's `wit_bindgen::generate!` `path:` — which the `wit/` move requires in all six shipped guests — invalidates the recorded digest and fails the gate. The gate's own contract forbids the shortcut: "Re-record only after `./scripts/build-wasm-extensions.sh --first-party` and committing the rebuilt artifact — the digest asserts a claim about the artifact, and updating it without rebuilding launders a stale one." So the artifacts are genuinely rebuilt (`--first-party`, exit 0, 6 OK / 2 host-native SKIP), not re-recorded in place. Byte sizes move by more than the source change accounts for because these builds are not reproducible by design — the guests pin no toolchain and resolve their own `Cargo.lock` at build time, which is the documented reason the gate hashes sources rather than artifact bytes. Verified: `check-wasm-artifact-freshness.py` OK (6 packages), and `cargo test -p ironclaw_extension_support` green (102/46/4) — that crate `include_bytes!`s these artifacts, so it exercises the rebuilt components. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-arch): record the WS7 artifact-rebuild cost of guest path edits The `wit/` move had to rebuild six shipped WASM binaries because `check-wasm-artifact-freshness.py` digests each guest's whole `wasm-src/` tree. WS7 hits the same wall from the other direction: the six package guests reach the ABI across two trees, so moving either `ironclaw_wasm` or `extensions/packages` rewrites all six `path:` literals and forces the same rebuild. Recorded on CHECKLIST WS10's `wit/` row (point 6), on the loud-path-pattern row that owns the WS7 repoint (also corrected six -> nine guests there), and on PLAN's Wave 5 block with the cheap mitigation: move the two crates in one PR and pay it once. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci(planner): classify the path classes that blocked the wit/ move `Detect Reborn test scope` exits 1 on any pull request whose diff holds a path `reborn_pr_test_plan.py` has no rule for, which made this PR unmergeable: it must edit `Dockerfile` (the moved directory's `COPY wit/ wit/` no longer resolves) and `scripts/check-version-bumps.sh` (the ABI gate would otherwise grep dead paths and silently stop enforcing). 18 of its 46 paths were unclassified. Same class as the `.claude/` gap #7064 fixed, and classified the same way — one rule per class, recorded beside the constant: * `Dockerfile` / `.dockerignore` — `platform-and-compat.yml` keys `has_docker_risk` off exactly this pair and owns the image build. * `.githooks/**` — Code Style triggers on the tree and lints its contents (`test-ci-comm-locale-pin.sh`); no Reborn lane runs a hook. * `scripts/{build-wasm-extensions,check-version-bumps}.sh` — `platform-and-compat.yml`'s `has_direct_wasm_abi_risk` classifier both scopes and runs them. * markdown owned by no crate (`crates/AGENTS.md`, `test-tools/README.md`) — prose, like `docs/` and `.claude/`. A crate-resident doc still selects its own crate's lane. The first-party extension package assets are deliberately NOT ignored. `crates/extensions/packages/*/wasm/*.wasm` is a shipped artifact that `ironclaw_extension_support` embeds with `include_bytes!`, and `test-tools/*/manifest.toml` is `include_str!`d by `ironclaw_extension_host`. Calling either prose would convert today's loud failure into a silent under-schedule of a change to production output — the WS10 failure mode. `EMBEDDED_ASSET_OWNERS` routes each tree to the crate that compiles it instead, so this PR now additionally schedules `ironclaw_extension_{support,host,manager}`: the crates that consume the six rebuilt WASM artifacts. Also fixes #7085 in a file this PR already touches. The WIT version extractors used the GNU-only BRE `\+`, so on BSD sed (macOS) they matched nothing, and because the `WIT_TOOL_VERSION` cross-check is guarded on a non-empty version the hook printed "All version checks passed" having compared nothing. `[[:space:]][[:space:]]*` is identical under GNU sed, so the enforced Linux CI lane is unchanged; verified on BSD sed that both `wit/tool.wit` (0.3.0) and `wit/channel.wit` (0.3.1) now extract. Regression tests: every classified class gets a case in `test_reborn_pr_test_plan.py`, including the paired assertion that the embedded assets *select a lane* rather than merely being accepted (the inverse of the `.claude/` prose test), and a staleness pin that fails if an asset tree or its owning crate moves. All ten new cases fail against the planner on `main`. `test_unclassified_build_input_fails_fast` moves off `Dockerfile` onto a still-undecided input so the fail-closed arm stays exercised. Refs #7087, #7085 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(host-runtime): split obligations into its three chartered owners (WS3) `crates/ironclaw_host_runtime/src/obligations.rs` was 3,122 lines fusing the three owners PROPOSAL §6.5.9 charters separately, held apart only by an `// arch-exempt: large_file` waiver. It is now one module per owner: - `obligations::handler` — which obligations apply and what each does before/after dispatch, plus the audit/redaction/ceiling/mount validation. - `obligations::staged_handoffs` — material staged for a later consumer: the runtime-secret and network-policy stores and the credential-account resolver port. - `obligations::process_store` — post-start handoff discard and reservation reconciliation. - `obligations::mod` — only `BuiltinObligationServices`, the assembly seam, and deliberately the one place naming all three at once. Every module is under the 1,500-line gate, so the waiver is deleted rather than carried forward: re-fusing the owners now trips `pre-commit-safety.sh`. `mod obligations;` stays private and the crate's `pub use obligations::{…}` names are unchanged, so no consumer outside the crate sees this. Behavior-free. Cross-owner access is `pub(super)` (three methods), not `pub(crate)`. The split revealed one narrowing in the other direction: `secret_present` was `pub(crate)` with no caller outside its own file and is now private. Also from the same CHECKLIST row, the bounded half of "shrink `services/builder.rs` toward composition-facing factories": three builder methods whose only callers are inside the crate's `src` narrow to `pub(crate)`. The rest of that clause is measured and deferred in the CHECKLIST amendment — 17 methods need a `test-support` cargo feature, three are callerless and belong to WS8, and the remaining 33 are a redesign of the fluent surface rather than a shrink of it. `+production_wiring` is refuted there: it is readiness diagnostics, not assembly. Two loud path-keyed gates fired and were repointed, not relaxed: `reborn_host_runtime_services_do_not_expose_lower_substrate_handles` now scans the whole `obligations/` directory and asserts it read ≥ 4 files (`collect_runtime_rs` returns a count; both its callers now assert non-zero), and `reborn_struct_test_support_ratchet`'s frozen per-file count moves to `staged_handoffs.rs` with its count unchanged at 1. Test accounting (un-masking discipline): `cargo test -p ironclaw_host_runtime --all-targets -- --list` is 1,246 before and 1,246 after, name-by-name identical — zero added, removed or renamed. `LAYER_MATRIX_EXCEPTIONS` is 10 before and after; an intra-crate split cannot move the register. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(operator,contracts): route operator secrets through a product_contracts port (WS3) `ironclaw_operator` is a products-tier crate and held `ironclaw_secrets`, the substrate that owns CAS one-shot leases, AAD/crypto and the OS keychain master key. PROPOSAL §8.2's product row says the products tier loses that edge, and §12.1b requires the port replacement to land before the edge is removed. Both happen here, in that order. - Port: `ironclaw_product_contracts::operator_secrets::OperatorSecretValueStore`. - Implementor: `ironclaw_reborn_composition::RuntimeOperatorSecretValueStore`, the same placement as `OperatorStatusService` — assembly is the only layer that may name both a products-tier port and a substrate. Registered in `INVERTED_PORTS` beside it. - `ironclaw_secrets` is gone from the operator manifest under every dependency kind, and `"ironclaw_secrets"` is now in the crate's `boundary_rules()` forbidden list. That gate's comment previously said the entry was deliberately absent because "the row owns it"; the row now owns it. The port is deliberately narrower than the substrate, so this is a tightening rather than a relocation: it takes no `ResourceScope` (the implementor fixes the operator scope, where the caller used to pass one), exposes no lease/consume protocol, and carries only a `&'static str` classification instead of the substrate's error `Display` — asserted, including that the backend message and the handle name are both absent from what crosses. Two tests travelled with the behavior rather than being pointed at a fake: `read_is_repeatable_across_reloads` (repeatability is a property of the lease protocol) and the #4673 production-store reproduction (its value is wiring the store exactly as production does, which now means the real store *behind the adapter*). Two `FaultInjecting`-over-real-store fixtures became per-operation port fakes, with the substrate error mapping re-pinned at the adapter; a third assertion got stronger — batched-vs-N+1 stored-key lookup is now observed at the port rather than by counting filesystem ops. Test accounting: operator 154 -> 153, product_contracts 142 -> 143, composition 937 -> 942 with zero removed; name-by-name diffs on a quiescent tree. Two findings the row could not have anticipated, both recorded in the CHECKLIST amendment: - The `webui` half of the row was already closed and was never a production edge. `ironclaw_secrets` has been a dev-dependency of `ironclaw_webui` since the commit that added it (#6619), both src mentions are `#[cfg(test)]`, and webui's boundary rule already forbade it. - `ironclaw_extension_manager` (layer `products`) still holds a normal `ironclaw_secrets` edge in `admin_configuration.rs`. §8.2 covers it; the row does not, because the crate landed with WS2.4 after the row was written, and the substrate sits in the service's type parameters so it is not a like-for-like swap. Filed as #7095. `LAYER_MATRIX_EXCEPTIONS` is 10 before and after: `products -> substrates` is matrix-legal, so this edge was always an §8.2 rule and never a layer exception. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(sandbox): put the Docker security check behind the fail-closed gate Review asked why the required Rust e2e lane can report `docker_security` as passing with no daemon. Half of that is #7081 (nothing sets IRONCLAW_REQUIRE_DOCKER_TESTS=1, so the switch is inert) and is not fixable from here -- arming it hard-fails any lane lacking a daemon or the worker image, which needs a runner guaranteed to have both. The other half is fixable here and is fixed: docker_security.rs open-coded its own `docker version` / `image inspect` checks with three bare `return`s, so it sat entirely outside docker_gate and would have stayed fail-open even once something did set the variable. It now takes both preconditions from docker_gate::{docker_available, docker_image_available} and skips with the visible `SKIP:` line that gate's module doc requires. Measured, same machine, image absent: before, IRONCLAW_REQUIRE_DOCKER_TESTS=1 -> "skipping ..." / 1 passed after, IRONCLAW_REQUIRE_DOCKER_TESTS=1 -> panic at docker_gate.rs:74 / FAILED after, variable unset -> "SKIP: ..." / 1 passed The third line is the no-op proof: the variable is set nowhere in this tree or on main, so no lane's behavior changes today. The daemon-down path already reached the image check and skipped there, so the outcome is identical; only the branch it takes differs. Two stale comments in docker_gate.rs corrected with it (they claimed docker_security used its own gate, and that docker_image_available had no consumer), and the crate's Known debt entry now splits the done half from the #7081 half instead of describing both as open. cargo test -p ironclaw_sandbox: 193 passed, 0 failed cargo clippy -p ironclaw_sandbox --tests --all-features -- -D warnings: exit 0 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(reborn): stop calling the unwired script lane an execution lane Two review findings, both correct, both artifacts of this PR's own renames. 1. engine-v2-to-reborn-parity.md note 4 read "a native script/software execution lane (`ironclaw_sandbox`, `RuntimeKind::Script`) sandboxed via `ironclaw_sandbox`" -- self-referential after the merge collapsed ironclaw_scripts and ironclaw_process_sandbox into one crate, and it contradicts note 5 four paragraphs down ("no production execution backend is wired for it"). Re-stated as the typed runtime contract it is, citing the measurement: `with_script_runtime` has zero production callers (`rg` finds only the builder itself, docs, and 30 test call sites). 2. CHECKLIST WS10 ratchet note 2 said "raise the percentage floor ...; only the line count should fall". That generalises WS3's sandbox merge, where observed coverage happened to rise. It is wrong as guidance for WS7, and the counterexample is in this same file: the 2026-08-03 entry from #7064 records ironclaw_runner falling 85.55% -> 82.53% because the shed removed the crate's better-covered half, holding the floor, and RATCHET FAILing in the merge queue. Note 2 now says re-capture from the merged artifact, and lower only with that entry's move-not-regression counterfactual (add the moved files back, confirm the union clears the old floor, plus a zero-tests- lost name set-diff). cargo test -p ironclaw_architecture: 32 targets, 206 passed, 0 failed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): pin the WIT scope probes and the embedded-asset owner pairing Three review findings on the `wit/` move, each verified before it was acted on. 1. `ws12_workflow_contracts.py` probed `crates/ironclaw_wasm/wit/host.wit` and its nested twin. No `host.wit` exists in this repository — `git ls-files '*.wit'` returns only `tool.wit` and `channel.wit` — so both probes sat under the `crates/([^/]+/)*ironclaw_wasm/` alternative and re-asserted the crate-name term while saying nothing about the canonical ABI contracts. In a validator whose stated design is "probe derived from reality rather than from a guessed layout", a fabricated filename is a defect on its own terms. Replaced with a `crate_globs` entry, `("ironclaw_wasm", "wit/*.wit")`, which discovers the contracts on disk, requires each in scope, and synthesises the nested WS7 form — so a third contract, or the directory leaving the crate, fails the pin instead of passing on a stale name. Verified non-vacuous: narrowing the workflow alternative to `.../ironclaw_wasm/src/` now reports `tool.wit`, `channel.wit` and the nested probe as out of scope. 2. The embedded-asset routing test substituted `alpha`/`beta` owners so it could reuse the synthetic workspace. That exercised the real prefix strings through the real routing, but left the prefix->owner *pairing* — the table's entire semantic content — asserted nowhere: swapping `ironclaw_extension_support` and `ironclaw_extension_host` passed. Fixed in two halves. The routing test now drives the real `EMBEDDED_ASSET_OWNERS` against a workspace carrying the real owners' names and real manifest paths (the synthetic one could not: `build_plan` rejects a changed package outside the canonical set), asserting the real owner is selected. And the not-stale test now derives the same pairing from the tree instead of restating the constant: it resolves every literal `include_str!`/`include_bytes!` in every workspace crate through `crate_tree`, keeps the targets no crate owns — the ones that actually reach the table — and asserts that every crate compiling one of them is the routed owner or a dependent of it. That surfaced a property worth pinning: `crates/extensions/packages/` is embedded by four crates, not one. `ironclaw_extension_host`, `ironclaw_extension_manager` and `ironclaw_reborn_composition` reach into it alongside `ironclaw_extension_support`, and routing to the support crate covers them only because each depends on it. If that edge goes, a shipped artifact change stops scheduling a crate that embeds it — the silent under-schedule the table exists to prevent. Regression coverage verified red by sabotage, all three wrong tables: owners swapped (7 failures), `packages/` -> `ironclaw_llm` ("embeds nothing from it"), and the hardest case, `packages/` -> `ironclaw_reborn_composition` — a real embedder that the other embedders do not depend on ("...does not depend on..., so routing there never schedules it"). 3. CHECKLIST WS10 claimed each of the nine `wit_bindgen` guest edits forces a committed WASM artifact rebuild. Only six do: `scripts/ci/check-wasm-artifact-freshness.py` scans `crates/extensions/packages/*/wasm-src` alone, `wasm-src-digests.toml` holds exactly six entries, and `git ls-files '*.wasm'` returns exactly those six. The three `test-tools/*/wasm-src/` guests commit no artifact; the tenth site is the host's `bindings.rs`, not a guest. Corrected, and the `wit/` row now states the boundary rather than implying it. Guest paths, `wit/` contents and the six rebuilt artifacts are untouched. Verified: `test_reborn_pr_test_plan.py` 46/46, `test_ws12_workflow_contracts.py` 25/25, `ws12_workflow_contracts.py` green on the real tree, `cargo test -p ironclaw_architecture` 206/206 across 32 binaries. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(host-runtime): state the obligation visibility rule as it holds Review catch (#7090): the guardrail sentence promised "cross-owner access is `pub(super)`, never `pub(crate)`", which is stronger than the code. Verified: `RuntimeSecretInjectionStore::{insert, take, clone_material, discard_for_capability}`, `NetworkObligationPolicyStore::{insert, get, take, discard_for_capability}` and both constructors are `pub(crate)` and must stay so — `src/egress/{mod,host_port,credential}.rs` call them, and that is host-runtime composition outside `obligations/`. The rule is restated as the property that actually holds: a method whose only callers are inside `obligations/` is `pub(super)` (the three that are), and `pub(crate)` is what the stores expose to the egress pipeline they exist to serve. A future agent reading the old sentence would have read the existing `pub(crate)` methods as violations. Guidance-only; no code change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(architecture): put the operator secrets boundary entry on the right rule Review catch (#7096), and it is the serious kind: the `"ironclaw_secrets"` entry landed in `ironclaw_extension_contracts`'s forbidden vector, not `ironclaw_operator`'s. The suite still passed, because `extension_contracts` has no such dependency and `ironclaw_operator` then had no entry at all — so the guard this row exists to add was inert, and a green architecture suite was evidence of nothing. Reintroducing the edge would have passed every check. Moved to `ironclaw_operator`'s vector; `extension_contracts` restored to its `origin/main` content byte-for-byte. Negative-probed rather than assumed. With `ironclaw_secrets` temporarily re-added to `crates/ironclaw_operator/Cargo.toml`: reborn_crate_dependency_boundaries_hold ... FAILED ironclaw_operator must not have a normal dependency on ironclaw_secrets and with the manifest restored, 35/35 pass. Two further review findings, both verified before being accepted: - `ironclaw_extension_manager` **does** have a `boundary_rules()` entry (`:3543-3556`, added with WS2.4). The CHECKLIST residue note and PROPOSAL §8.2's 2026-08-02 amendment both said it had none; §8.2's sentence is stale and is marked superseded. The real gap is narrower and now stated: the rule exists and simply does not forbid `ironclaw_secrets` (#7095). - `ironclaw_product_contracts`'s guide claimed "twenty-four shipped modules". Measured: `src/lib.rs` has 26 shipped (27 `pub mod` less the gated `test_support`), and the table was missing `ironhub` **before** this branch touched it. Count corrected to twenty-six and the missing `ironhub` row added, so the inventory matches `lib.rs`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sandbox): state the Docker-gate claim as the search that checks it Review caught a false inventory in the Known debt entry, and the previous commit is what made it false: "the name appears only in docker_gate.rs and attribution_tests.rs" stopped holding the moment docker_security.rs gained a module doc naming the variable, and CLAUDE.md itself was already a third counterexample. The narrower claim is the one that was always meant and is the one that matters, so it now carries its own reproduction: no workflow, script, env file or manifest mentions the name at all -- `git grep` over *.yml/*.yaml/*.sh/ *.toml/*.py/*.json/.env* is empty here and on main -- and the sole code reference is a read, std::env::var(...) at docker_gate.rs:23. Every other occurrence is a doc comment or a panic message. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(coverage): re-anchor the exemptions the merge shifted tests/integration/changed-coverage-exemptions.toml is exact-line-keyed and auto-merges silently. #7096's additions to ironclaw_reborn_composition moved four entries' subject lines by +2 without anything flagging it; a stranded entry makes the changed-coverage validator abort with no verdict at all. Re-anchored by content (difflib line map from the #7065 tree, which the file was validated against, to the union) rather than by arithmetic: runtime.rs [4068..4073, 4082, 4083] -> [4070..4075, 4084, 4085] runtime.rs [3701] -> [3703] ; runtime.rs [3433] -> [3435] lib.rs [616] -> [618] All 142 entries / 1124 line references re-verified against the merged tree: 0 drift, 0 out-of-bounds, 0 missing paths. * refactor(layers): re-layer processes -> kernel and skills -> substrates (WS3/WS4) Two CHECKLIST rows, both of which were a one-line manifest correction rather than a code move: the family docs already placed both crates where the rows want them and only `Cargo.toml`'s `layer =` disagreed. processes -> kernel (WS3). families/kernel.md already lists ironclaw_processes among the kernel crates. The re-layer makes processes -> resources a kernel -> kernel edge, so its LAYER_MATRIX_EXCEPTION went STALE and the gate said so itself: Stale IronClaw crate layer matrix exceptions: ironclaw_processes -> ironclaw_resources from 2026-07-09 should be removed in W7: runtime process management still depends on resource contracts currently classed with kernel behavior That is the gate's verdict, not a judgement call - deleting the entry is the only way to make it pass. Baseline 5 -> 4, recomputed as len(merged list). Checked the direction both ways: all nine crates that take a normal dependency on processes (capabilities, turns, host_runtime, extension_host, loop_host, extension_manager, runner, reborn_composition, stress) are kernel or above, so the move legalizes an edge without forbidding an existing one. skills -> substrates (WS4 SS3.D). families/domains.md already lists ironclaw_skills under 'Layer(s): substrates'. Its only two normal dependencies are ironclaw_filesystem (substrates) and ironclaw_host_api (contracts), both at or below substrates, and its six consumers are all loops or above. No exception moves in either direction. cargo test -p ironclaw_architecture: 206 passed, 0 failed. cargo check --workspace --all-targets: clean. * docs(target-arch): close the WS3/WS4 rows this work satisfies, with evidence Every tick was verified against the merged tree, never against a PR title. TICKED: - sandbox lane merge: ironclaw_sandbox exists, ironclaw_scripts and ironclaw_process_sandbox absent, bollard/rcgen declared by exactly one manifest in the workspace. - mcp drops the registry dep: ironclaw_extensions is [dev-dependencies] only, 0 production ironclaw_extensions:: refs in src/. - skills -> substrates: landed here. - hooks libSQL/Postgres [decision]: ADR recorded - keep both, with the four rejected alternatives and the evidence they are already converged on one trait plus a shared conformance suite. #6945 read first as the row demands, and explicitly NOT discharged: this PR changes nothing in the dispatch path. - WS3 verify row: the row conflated Wave 3 with Wave 5 work (9 of its 10 exceptions carried removes_in = W7). Corrected with the replaced text quoted, the Wave-3 half satisfied edge by edge, and the Wave-5 remainder named with its owning field value. Ticked on the corrected condition. LEFT OPEN OR PARTIAL, each with measurements rather than a hand-wave: - first_party_tools: 1 of 6 families moved; 15 modules still in host_runtime. Ticking would be false. - processes/capabilities row: re-layer DONE; the capabilities/host.rs split is deferred with every module boundary already computed (4,560 lines, the six workflow ranges, and the arch-exempt waiver that must be deleted with it). - host_runtime binding/catalog-defaults: binding half REFUTED (moving it needs RuntimeLaneExecutor/RuntimeLaneRequest made pub, contradicting the same section's Keeps clause; zero external references to either). Catalog half cannot go to extension_host at all - host_runtime is itself a production consumer at memory_native_extension.rs:96,101, so the move is a kernel -> products edge and a Cargo cycle. Correct destination is downward. - network test_rewrite: NOT executed. Recorded the security shape (production binaries compile the seam and honour the rewrite env var at runtime) and the full 6-step plan, because the env var is how the entire E2E suite redirects vendor traffic through the production binary and the change needs feature forwarding into CI lanes I cannot verify here. cargo test -p ironclaw_architecture: 206 passed, 0 failed. * ci(coverage): recapture the two composed floors from a real measurement The provisional values were arithmetic - the sum of the two slices' recorded deltas - and the dispatch caught them, which is the whole reason the brief demanded a measurement rather than a reconciliation. Dispatch run 30907774036 at 4512e03e28f1df15b419d2e36f9f38f8f55d62fd: 26 success / 1 skipped / 2 failure, judged by per-job tally per #6978. The one skip is the pull_request-gated mutation gate; the two failures are the coverage report and the roll-up it drags down, i.e. this file doing its job. ironclaw_host_runtime: predicted 89.05% (18801 / 21114), MEASURED 88.63% (17562 / 19814). The composition was wrong by 1300 denominator lines because both slices measured their delta under the pre-#7083 aggregator, which could not see crates/extensions/** at all - lines leaving host_runtime for extension_support vanished from the tree it could measure, so neither branch's recorded delta describes the post-#7094 world. ironclaw_extension_support: MEASURED 75.31% (7142 / 9484) against #7094's 82.64% (6826 / 8260), captured before #7080's executor lines arrived. floor_percent FALLS 7.33pp and that is flagged in the file for an owner's eye rather than written quietly. Evidence it is composition and not lost tests: floor_covered_lines RISES 6826 -> 7142, so the crate is protected by more absolute lines than before, and #7080's un-masking accounting was 1398 -> 1398 with zero test names lost. Same shape as #7094's own ironclaw_runner recapture. ironclaw_sandbox passed unchanged at its arrival capture (87.09%, 3185 / 3657). The [global] entry is untouched: both moves are crate-to-crate inside the set the fixed aggregator sees. * fix(network): compile the test rewrite seam out of production builds (WS3) Closes the WS3 network row. Also RETRACTS an overstatement I made in this row's earlier annotation. CORRECTION FIRST. The earlier note claimed production binaries compile the seam and honour IRONCLAW_REBORN_TEST_HTTP_REWRITE_MAP at runtime, so anyone able to set it could redirect all credentialed vendor egress. That was WRONG. RewriteNetworkTransport::from_env_value already returned UnavailableInRelease when !cfg!(debug_assertions) (test_rewrite.rs:150), and neither [profile.release] nor [profile.dist] sets debug-assertions, so a shipped binary with the variable set REFUSES TO BOOT. It was fail-closed before this PR. I had read the ungated `mod test_rewrite;` declaration as an ungated runtime path. What was genuinely wrong, and is fixed: 1. The guard was a RUNTIME check keyed on cfg!(debug_assertions) - a profile proxy, not a build-kind guarantee. A release profile with debug-assertions turned on (normal when chasing a production bug) silently re-arms it. 2. The refusal arm had NO TEST. The one guard between a shipped binary and redirectable vendor egress was unpinned. Fix: compile-time exclusion instead of a runtime check. mod test_rewrite and its four re-exports are now cfg(any(debug_assertions, feature=test-support)), and default_host_http_egress is a compile-time pair - production builds PolicyNetworkHttpEgress<ReqwestNetworkTransport> directly, with the rewrite wrapper absent from the binary. The runtime check stays as defence in depth. E2E needs no change: those harnesses build DEBUG binaries, so they satisfy debug_assertions and keep redirecting with no feature flag and no workflow edit. The feature-forwarding-into-CI risk I flagged earlier does not arise. test-support is still forwarded composition -> network for a release-PROFILE build that needs the seam. Both halves proven rather than assumed: (a) release refuses - new regression test a_set_rewrite_map_activates_only_in_debug_and_is_refused_in_release feeds a well-formed map and asserts on profile. Under 'cargo test --release -p ironclaw_network --features test-support' it passes on the UnavailableInRelease branch; under debug 'cargo test -p ironclaw_network' it passes on the active branch. 56 passed, 0 failed. (b) production compiles without the seam - 'cargo check --release -p ironclaw_reborn_composition' (no test-support) is clean, which only compiles if the cfg(not(..)) arm is right. Also: WS0_EXTENSION_SPECIFICITY_ALLOWLIST_BASELINE 129 -> 127. The constant had drifted ABOVE the real list length; the ratchet is shrink-only so it passed silently while buying back two unearned slots. Measured off the compiler (set baseline to 0, read the reported length), identical on main and on every slice, so pre-existing drift rather than something this PR caused. cargo test -p ironclaw_architecture: 206 passed, 0 failed. cargo check --workspace --all-targets: clean. * docs(coverage): verify the extension_support floor drop is composition, independently The 82.64 -> 75.31 recapture carried a rationale that was recorded but explicitly NOT verified. Re-derived it from scratch between the two capture refs (f946a93fae -> 939af4847d) rather than inheriting the claim: - 0 test names lost in the crate (158 -> 160 test fns; both new names belong to the arriving executor). - 0 test names lost WORKSPACE-WIDE (13836 -> 13843 test fns, 13752 -> 13759 unique). This is the check that separates a relocation from a deletion: host_runtime's roster drops 156 names over the same range and every one reappears in another crate. - Exactly four files arrived, 1367 source lines, all of them the family-1 skill-install executor (src/skills/url_install.rs + url_install/{github, zip_bundle,bundle}.rs). No pre-existing file left the crate. - The arithmetic closes with the pre-existing numerator held CONSTANT: (6826+316)/(8260+1224) = 75.31% exactly, so the pre-existing code lost zero covered lines. The arriving block's own coverage is 316/1224 = 25.82%. Composition, confirmed rather than assumed. No test regression to fix; the 25.82% arrival is what earns the follow-up already recorded above the entry. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(host_runtime): collapse a duplicated obligation predicate and quiet a background warn! Three verified review findings from the #7141 round. Each was confirmed against the code before being acted on; nothing was changed on assertion alone. 1. obligations/handler.rs — `obligation_supported_before_dispatch` and `obligation_supported_after_dispatch` had BYTE-IDENTICAL 19-line bodies (verified by exact line-by-line comparison). Both were private, each called exactly once, both taking the same `phase` argument. The two names asserted a pre/post-dispatch distinction the code never implemented, while the pair gates admission of RedactOutput, EnforceOutputLimit and EnforceResourceCeiling — so editing one copy alone would have left the other stage accepting an obligation the host cannot honour (a fail-open). Collapsed to one `obligation_supported`, with the reasoning recorded so the pair is not reintroduced. 2. obligations/process_store.rs — `cleanup_terminal` is reached from `observe_process_commit` (an async background journal callback, call sites at :363/:379/:394), so its `tracing::warn!` violates the repo rule that background tasks never use info!/warn! — they corrupt the REPL/TUI display. Lowered to `debug!`; the error is still returned to the caller on the next line, so nothing is swallowed. 3. reborn_restructure_baselines.rs — the doc table said the LAYER_MATRIX_EXCEPTIONS count was "now 11". Recomputed on this ref by anchoring on the `= &[` of the value (the `&[LayerMatrixException]` type annotation opens a bracket on the same line and silently yields 0): the real count is 4, matching WS0_LAYER_MATRIX_EXCEPTION_BASELINE = 4. Corrected. Verification: cargo check --all-targets -p ironclaw_host_runtime exit 0; obligation tests 13+26 passed, 0 failed; reborn_restructure_baselines 1 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): a shipped package prompt is an asset, not prose — it was selecting no lane Review finding on #7141, confirmed empirically before acting. The Markdown prose carve-out in the planner ran BEFORE the `EMBEDDED_ASSET_OWNERS` lookup. A prompt is a `.md` file that no package *directory* owns, so a change to `crates/extensions/packages/*/prompts/**.md` took the prose arm and planned: mode=none crate_buckets=[] "crate-tree guidance changed: ..." while its sibling `manifest.toml` in the same package planned `mode=selected` onto ironclaw_extension_support + ironclaw_extension_host. Prompts are shipped production output that `ironclaw_extension_support` compiles in, and the comment above `EMBEDDED_ASSET_OWNERS` names "manifests, prompts, schemas and built wasm/*.wasm" as exactly what that table owns — so this was the "silent under-schedule of a change to production output" that comment forbids. 145 of the 149 `.md` files under `packages/` are prompts. The rule is keyed on the `prompts/` path segment, not on the asset prefixes. That distinction is load-bearing: the first attempt yielded to the asset prefixes wholesale and broke `test-tools/README.md`, which is documentation of the fixture bundles and is deliberately pinned as prose. Of the four asset kinds the table owns, only a prompt is Markdown (manifests are .toml, schemas .json, wasm .wasm), so `.md` asset <=> prompt is exact. Sabotage-tested in both directions: * `_is_package_prompt` -> False (reinstates the bug): RED, "AssertionError: 'none' != 'selected'". * `_is_package_prompt` -> any .md under an asset prefix (over-broad): RED on both the new test and the pre-existing `test_markdown_owned_by_no_crate_is_prose`, at `test-tools/README.md`. * restored: 52 passed, 51 subtests, green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(harness): refresh the latency-runner lockfile after the sandbox consolidation Review finding on #7141, reproduced before fixing. The latency harness keeps its own committed `Cargo.lock`, separate from the workspace lockfile, and the crate consolidation that replaced `ironclaw_scripts` + `ironclaw_process_sandbox` with `ironclaw_sandbox` never regenerated it. It still carried entries for both removed packages (lines 3244 and 3602) and the old host-runtime/loop-host dependency graphs. Reproduced exactly as reported: $ cargo metadata --locked --manifest-path harness/latency/runner/Cargo.toml error: cannot update the lock file ... because --locked was passed exit 101 so any reproducible invocation of the harness was broken, while the documented unlocked command silently rewrote the lockfile as a side effect of running. Regenerated with `cargo update --workspace`, which re-resolves the path dependencies. Verified after: `--locked` exits 0, the two removed packages are gone (0 entries), and `ironclaw_sandbox` is present (1 entry). Note: the re-resolve also carried three registry deps forward (wasmtime-wasi 46.0.1 -> 47.0.3, wasmtime-wasi-io likewise, wit-parser 0.251.0 -> 0.252.0). That is contained — this lockfile governs only the standalone benchmark harness and is not the workspace lockfile, and it was already unusable under `--locked` before this change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(skills): stop …
This was referenced Aug 5, 2026
pranavraja99
added a commit
that referenced
this pull request
Aug 5, 2026
One conflict, in `first_party_tools/skill_management.rs`, and one decision behind it. WS3 moved install-input normalization out of the host runtime into `ironclaw_extension_support::skills::resolve_install_input`, while this branch was editing the host-runtime copy (`skill_install_input`). Resolution takes main's side of the move wholesale -- the function is gone from this crate, and the branch's `passthrough` unit module goes with it, since it asserted against a hand-copy of the logic rather than the logic. The behaviour change this branch exists for is re-applied at the new location instead. That change is: an inline install may carry `files`. main's resolver grouped `files` with `source`/`source_url` as fields "a caller may never supply", citing review on #7141. Splitting them apart: * `source`/`source_url` are provenance. A caller supplying them is claiming its own output was fetched from somewhere trusted. Still refused, hard, with nothing written -- and now also refused when a bundle is attached, which nothing pinned before. * `files` is ordinary caller content. Refusing it did not drop the file, it failed the WHOLE install, so an agent attaching the script its skill exists to preserve got an `InputEncode` and nothing else. Measured on the 31-task SkillsBench subset (nearai/benchmarks#287): 18 correctly-shaped `{path, text}` entries across 9 calls, all refused; 0 of 27 agent-authored skills shipped a resource file against 18 of 31 human-curated ones. What makes accepting `files` safe is not the shape check that used to sit there: each entry's path is normalized and confined to the skill's own directory downstream (`install_bundle::normalize_safe_relative_path`). The url arm is untouched -- it still rebuilds from the fetched payload and drops caller files. The four capability-level fixtures added in the commit before this merge were written to pass on both sides of it. Against a mechanically-resolved tree that kept main's tightened arm verbatim, the two acceptance cases failed with `InputEncode` and the two refusal cases passed, which is the reason to trust them: the relaxation is load-bearing and the security invariant is independent of it. All 33 skill-install tests pass after, including main's own url-path suite. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
pranavraja99
added a commit
that referenced
this pull request
Aug 5, 2026
The comment claimed a review decision had been "reversed on evidence", which overstates what happened and would read as overriding the team. What actually happened: the refusal predates #7141 entirely -- it lived in the host-runtime copy of this resolver, and #7141 carried it across the move to this crate verbatim, declining a reviewer's suggestion to relax it there. That was the right call for a move-only refactor. This PR is where the behavior change belongs, and it is made deliberately with the measurement attached. Comment only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This was referenced Aug 5, 2026
Merged
pull Bot
pushed a commit
to Stars1233/ironclaw
that referenced
this pull request
Aug 7, 2026
…tallable, and complete (nearai#6745) * fix(reborn): inject skill bodies by default, not a one-line listing Reborn defaulted `SkillInjectionMode` to `Listing`, where a non-activated skill contributes only `- name: description` to context and its body loads only on an explicit `$name` mention or a `builtin.skill_activate` call. The intent was to save context budget. Benchmarking shows the model reads the menu and then never opens the skill. Over 30 runs with human-curated skills installed (SkillsBench/SkillLearnBench subset, `deepseek-v4-flash`, nearai/benchmarks#287): builtin.skill_list called in 30/30 runs builtin.skill_activate called in 3/30 runs a skill body actually read 0/30 runs So installed skills were effectively inert. Same 31 tasks, same skills, same model, varying only this default: no skills 78.5% curated skills, Listing 79.8% (+1.3pp -- skills bought almost nothing) curated skills, Full 85.6% (+7.1pp) For reference, harnesses that inject skill bodies unconditionally (Hermes, Claude Code) score 91.5% on these tasks with the same skills, so `Full` closes most but not all of that gap; the remainder is loop/verification behavior on a handful of multi-output tasks and is tracked separately. `Full` is already the library default in `SkillActivationSelectorConfig`; only the Reborn composition seam opted out. This restores it and adds a guard test so a revert is deliberate. `IRONCLAW_REBORN_SKILL_INJECTION=listing` still selects the previous behavior where context budget matters more than skills being used. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(skills): hot-swappable activation strategies so agent-authored skills are reusable Adds `skill.activation.v1`, a swappable-provider module in the shape of the memory-provider binding (`ironclaw_host_runtime::memory_binding`): named strategies, fail-closed resolution, behavior-preserving default, and a composition seam so nothing downstream names a concrete implementation. ## The bug it addresses `selector::score_skill` accumulates score ONLY from `activation.keywords` (+10/+5), `activation.tags` (+3) and `activation.patterns` (+20). A skill's `name` and `description` contribute nothing, and `select_skills` keeps a skill only `if score > 0`. That is fine for curated skills, which ship an `activation` block. It is fatal for skills an agent writes for itself: measured across the 31-task SkillsBench/SkillLearnBench subset in nearai/benchmarks#287, **0 of 30** agent-authored skills contained an `activation` block. Every one scored 0 and was permanently unselectable — the agent could create a skill via `builtin.skill_install` and then never reuse it, which makes self-improvement structurally impossible rather than merely weak. Claude Code has no such requirement: a skill is selectable from name and description alone. `ActivationStrategy::NameAndDescription` ports that contract. ## Design * `CriteriaOnly` (default) — today's rule, byte-identical. * `NameAndDescription` — whole-word name/description fallback, applied ONLY when the criteria pass scored 0, so a curated skill's explicit keywords always decide ordering and this can never reorder two skills that both declare metadata. `NAME_WORD_SCORE` (8) is deliberately below the selector's exact-keyword award (10). * `Disabled` — explicit mention / `skill_activate` only. * `ThirdParty { extension_id }` — production requires an admin override. Whole-word matching and a `MAX_FALLBACK_SCORE` cap keep it from over-selecting; over-selection is the failure mode that makes injecting an unrelated skill bank harmful (a whole-catalog injection took `xlsx_recover_data` 1.000 -> 0.271). ## Default stays behavior-preserving Reborn's default remains `CriteriaOnly`, opt in with `IRONCLAW_REBORN_SKILL_ACTIVATION=name_and_description`. Flipping the default changes three existing local-dev expectations (setup-marker suppression, the webui listing candidate, `skill_activate` context loading), so the strategy ships opt-in — the same discipline as the memory work, where the bundled native provider stays the default. ## Tests `cargo test -p ironclaw_skills --lib` — 239 passed, including: * `agent_authored_skill_unreachable_by_default_but_selected_under_name_strategy` — end-to-end via `prefilter_skills_with_options`: the same no-activation skill is dropped under `CriteriaOnly` and selected under `NameAndDescription`. * `name_strategy_does_not_select_an_irrelevant_skill` — no over-selection. * `name_hit_outranked_by_an_explicit_curated_keyword`, `whole_word_only_...`, `fallback_is_capped_...`, `stop_words_do_not_accumulate_score`. `cargo test -p ironclaw_first_party_extension_ports --lib` — 58 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(reborn): ship the Full skill-injection default as opt-in, not a flip The measurement in the previous commit stands: `Listing` leaves installed skills unread (`skill_list` 30/30 runs, a body actually opened 0/30) and `Full` is worth 79.8% -> 85.6% on the 31-task SkillsBench subset. But flipping the product default HANGS three existing local-dev tests, which drive a mock that expects the one-line listing candidate: * `local_dev_skill_activate_tool_loads_selected_skill_context` * `local_dev_webui_bundle_records_selectable_filesystem_skill_context` * `local_dev_runtime_wires_filesystem_skills_by_default_to_model_calls` Verified by bisect: all three hang on the previous commit alone, and pass with the default restored — the activation-strategy work is not implicated. Changing a documented product default in a way that turns CI red is a maintainer call, not something to force through, so `DEFAULT_SKILL_INJECTION_MODE` returns to `Listing` and `Full` ships as `IRONCLAW_REBORN_SKILL_INJECTION=full`. Both switches in this PR are now opt-in with the evidence attached, matching the memory-provider discipline where the bundled default is preserved. The guard test is retargeted to assert the current default, verify the opt-in path still resolves, and name the three tests that must be updated alongside a future flip. cargo test -p ironclaw_reborn_composition --lib -- skill_injection_mode \ local_dev_selector_config skill_activation # 14 passed, 0 failed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(threads): raise the result_read cap to 64 KiB, env-tunable A small per-request `result_read` cap turns one large file into a paging loop. On `manufacturing_equipment_maintenance` (nearai/benchmarks#287) reborn made 8 `read_file` calls and ZERO shell calls, hit the 24 KiB cap, then spent the whole turn paging — `result_read` at offset 24576, `handbook.pdf` at offsets 400/800/1200 — and never computed anything (`outputs_exist=0.00`). hermes, using shell to sample the same data, scored 0.522. * `TOOL_RESULT_RECORD_READ_MAX_BYTES` 24 KiB -> 64 KiB. This is the compile-time ceiling the model-observation envelope in `tool_result_reference.rs` is derived from (`* 2`, asserted at compile time), so 64 KiB here means a 128 KiB envelope — the reason not to go higher. * `TOOL_RESULT_RECORD_READ_DEFAULT_MAX_BYTES` = 64 KiB — the effective default. Enough that a typical data file or document page arrives in one read instead of a paging loop. * `IRONCLAW_TOOL_RESULT_READ_MAX_BYTES` overrides it, clamped to `[4, ceiling]`, so an override can never outgrow the envelope. Unparseable values fall back to the default rather than failing the run — a malformed tuning knob must not take down an agent. Unlike the skill-injection and skill-activation switches in this branch, this one does move the default: the paging loop is a silent capability loss rather than a behavior preference, and the knob exists for deployments that want the old size. cargo test -p ironclaw_threads --lib # 88 passed (85 existing + 3 new) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(skills): add always_available activation, Claude Code's actual contract `skill.activation.v1` gains a third binding, `always_available`: every installed skill is a candidate regardless of what it matches. This is what Claude Code and Hermes actually do. In both, a skill is a file in a directory the agent can read, so there is no gate for a correctly-installed skill to fail. Reborn's selector instead scores only `activation.keywords`/`tags`/ `patterns` and drops anything scoring 0 -- and `name_and_description` (this branch's earlier binding) only WIDENS that gate: it still needs a lexical hit, so an applicable skill phrased differently from the prompt is still discarded. The new test pins exactly that case -- a skill described as "cyclical component / growth path" against a prompt saying "hp filter" is dropped by both `criteria_only` AND `name_and_description`, and kept by `always_available`. Why it matters, measured on the 31-task SkillsBench/SkillLearnBench subset in nearai/benchmarks#287: 0 of 30 agent-authored skills contained an `activation` block, so under `criteria_only` a self-authored skill could never be selected again -- self-improvement was structurally impossible. Implementation is deliberately tiny: a `floor_score()` of 1 for this binding, applied via `.max()` in the selector's existing scoring loop. Ordering is untouched (a real keyword match still outranks a floor skill, so the context budget spends on the relevant skill first), and the existing budget -- not the score filter -- decides what is injected, which is also how Claude Code behaves. `floor_score()` is 0 for every other binding, so non-adopters are byte-identical. Default remains `criteria_only`; opt in with IRONCLAW_REBORN_SKILL_ACTIVATION=always_available. cargo test -p ironclaw_skills --lib # 241 passed (239 existing + 2 new) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(threads): drop the now-unused ceiling import Validation bounds against `contract::effective_tool_result_read_max_bytes()` (which applies the env override), so the compile-time ceiling is no longer referenced here. Removes an unused-import warning introduced by the 64 KiB cap commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * revert(threads): default result_read back to 24 KiB, keep the knob The raise to 64 KiB was never isolated: it shipped in a measurement arm alongside two other switches (skill activation, tool disclosure), so there is no evidence it changed anything. Defaulting it back keeps this crate byte-identical to pre-PR behavior. The paging trace that motivated it is real (`manufacturing_equipment_maintenance`, nearai/benchmarks#287: 8 `read_file` calls, zero shell calls, `result_read` at offset 24576, nothing computed) — but a real trace is not a measured fix, so the larger cap stays opt-in via IRONCLAW_TOOL_RESULT_READ_MAX_BYTES for whoever wants to measure it properly. The compile-time ceiling stays 64 KiB: it now bounds only how far the env override may reach, and still pins the derived model-observation envelope at 128 KiB. Net effect of this commit plus its parent: a new env knob, no default change. cargo test -p ironclaw_threads --lib # 88 passed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(skills): design for agent-authored multi-file skill bundles @henrypark133 pushed back on "move skills to the filesystem" as an overhaul that a single aggregate result did not justify. He was right, and stratifying the data shows why: the entire filesystem gain sits in skills that ship files besides SKILL.md. ships resource files (n=16): inject 81.0% -> files 94.2% (+13.2pp, CI [+0.3, +26.2]) SKILL.md-only (n=11): inject 91.5% -> files 84.7% (-6.9pp, CI [-20.4, +6.7]) So filesystem-for-everything is a REGRESSION on 13 of 31 tasks, paid to fix the other 18. The mechanism is not "models prefer filesystems": 81 of the resources are executable (you cannot run pasted Python -- citation_check scored 0.000 with the script absent, 0.833 with it present), and the text resources are too large to inline (exceltable_in_ppt would be ~262k tokens folded into SKILL.md). The design therefore keeps storage, discovery and selection exactly as they are and adds ONE extension holding the already-existing `/skills` read_write mount: skill_write_file / skill_read_file / skill_list_files. Discovery already lists from the same root that mount writes to, so nothing needs plumbing. Executing a bundled script copies that one file into `/workspace`, which the agent already mounts. Documents two things the implementation must not miss: SkillBundleDescriptor exposes only `skill_md_path`, so bundle resources are un-advertisable without skill_list_files; and `FilesystemSkillBundleRoot::user` marks bundles Trusted, so an agent that can write executable scripts there needs a distinct trust level -- the real open question. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(skills): state explicitly that creation, discovery and indexing are unchanged The crux of @henrypark133's objection. Spells out, per concern, that skill creation stays on the `skill_install` tool, discovery stays on the storage-agnostic `SkillBundleSource` trait with no new impl / trait method / descriptor change, and that there is no session-start index to migrate at all (selection is per-request; the only cache is a 5-minute TTL on catalog search). The single behavioral change remains the opt-in `always_available` selection predicate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(skills): the write tool needs an authoring prompt that asks for code skill_write_file makes multi-file skills possible; it does not elicit them. Measured: 6 of 31 tasks finished with ZERO skill_install calls despite 'Saving the skill is required', and the authoring request only ever asks for prose (method, conventions, output contract). An agent following it writes prose whether or not a write tool exists. Adds the elicitation requirement and a falsifiable success criterion: agent-authored bundles are currently 100% prose (0 of 27 ship a resource file) against 18 of 31 curated skills. If that ratio does not move once the tool ships, the bottleneck was elicitation rather than capability and the tool alone will not move scores. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(skills): let an agent install skill bundles, not just prose Agents could only ever author the PROSE half of a skill. Measured on the 31-task SkillsBench/SkillLearnBench subset (nearai/benchmarks#287): **0 of 27** agent-authored skills shipped a single file besides SKILL.md, against **18 of 31** human-curated ones (79 .py scripts, 78 .xsd schemas, 84 .md references). So every later run re-derived the method from prose and could re-make the same mistake -- lake_warming's self-authored skill described its regression procedure in prose, the next run recomputed it slightly differently and missed the grader's p<0.05 threshold. This was NOT a missing capability. `install_skill` has always taken `files: &[SkillInstallFile]`, and `parse_install_files` has always read an `input["files"]` array. Two things made it unreachable: 1. `schemas/builtin/skill_install.input.v1.json` advertised only `name`/`content`/`url` AND set `additionalProperties: false` -- so a model sending `files` was not merely uninformed, it was REJECTED. Across 112 observed skill_install calls, 111 used exactly `['content','name']`, which is what the schema permits. 2. The only encodings were `bytes_base64` and a JSON array of byte integers. A bundle file an agent writes is a script, a reference doc or a schema fragment -- all UTF-8. Making those go through base64 costs ~33% more tokens and turns one encoding slip into an InputEncode failure of the whole install. Changes: - `parse_install_files` accepts `text` (UTF-8) alongside `bytes_base64`/`bytes`. `text` takes precedence when both are given, matching the documented preference. Binary payloads are unaffected. - the schema advertises `files` with `path` + `text`/`bytes_base64`, and the description tells the model WHY to use it: put a reusable computation in a script rather than describing it in prose, and have SKILL.md name the files it relies on. That last part matters because `SkillBundleDescriptor` exposes only `skill_md_path`, so a bundle cannot advertise its own resources. - prose-only installs are untouched: no `files` key still parses to an empty vec. cargo test -p ironclaw_first_party_extensions --lib install_files_encoding # 4 passed cargo test -p ironclaw_host_runtime --test tool_surface_contract # 43 passed cargo test -p ironclaw_reborn_composition --test product_live_adapters skill_install # 1 passed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(skills): stop rejecting an install that carries both content and files `skill_install_input` gated the direct-install arm on `!object.contains_key("files")`, so `content` + `files` matched NO arm and fell through to `_ => Err(InputEncode)`. An agent attaching a script had its ENTIRE install refused. `files` was reachable only on the URL-fetch arm, which builds the array itself. This was the third of three stacked gates hiding the same capability, and the one that actually bit. With the schema fixed to advertise `files` and a `text` encoding available, the model on the 31-task SkillsBench subset (nearai/benchmarks#287) immediately sent 18 correctly-shaped `{path, text}` entries across 9 calls -- `scripts/verify_bib.py`, `references/fake_patterns.json` -- and every one was rejected here. That is the real reason 0 of 27 agent-authored skills shipped a resource file while 18 of 31 human-curated ones do: not a missing capability, and not the model failing to try. `source`/`source_url` stay excluded from the direct arm: those record provenance and are set by the URL path, so an agent must not be able to forge them. cargo test -p ironclaw_host_runtime --lib skill_install_input # 4 passed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * style: rustfmt the skill-bundle and activation changes Test modules were appended programmatically without rustfmt, which is why Formatting, Code Style and Clippy all went red on this PR. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(skills): rewrite to match what was measured, not the abandoned design The doc recommended a three-tool extension plus a resource-gate. Both are superseded: the tools turned out to be redundant (install_skill already accepted files -- three stacked gates were hiding it), and the gate MEASURED WORSE than always advertising a readable path (-25.7pp on self-creation, -40.6pp vs claude-code), because an agent-authored skill is usually SKILK.md-only so the gate suppresses the one route the selector had not already closed. Rewritten around the durable findings: the three gates and how each masked the next, the 0-of-27 vs 18-of-31 measurement, the SkillBundleDescriptor enumeration gap, and the trust question. The gate is kept in the doc as a recorded negative result, since its stratified justification is persuasive and will be proposed again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(skills): correct why always_available is not the default The previous note claimed the floor score overrides setup-marker suppression. It does not: `prefilter_skills_with_options` returns None for a satisfied marker BEFORE scoring, and the host-side filter in activation.rs already removed the candidate. What actually fails: all 32 bundled skills reach floor 1, so 3-4 unrelated ones land in plan.activations() in ActivationCriteria mode -- a mode that injects nothing under Listing. The defect exposed is that a criteria activation which injects no body is still recorded as an activation, so the count assertions stop being meaningful. Also records the sequencing against epic nearai#6565 (Slice 0 first; Slice 5's bounded-shortlist rule constrains what an unbounded floor may do) and the measured detail that under Listing a zero-scoring skill is still listed -- the model just called skill_activate in only 3 of 30 runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(skills): a floor-only skill is ranked, not activated Three defects, all surfaced by trying to make `always_available` the default. It failed 8 tests in ironclaw_reborn_composition; all 8 now pass with the flag on AND off. 1. A criteria selection that injects nothing was still recorded as an activation. Under `SkillInjectionMode::Listing` an `ActivationCriteria` entry contributes no body -- `body_eligible_bundle_ids` already ignores that mode -- so with a floor score every installed skill "activated" on every turn. Concretely: all 32 bundled skills reach floor 1, and 3 of them (6000 token budget / 2000 default per-skill cost) landed in each plan, chosen by descriptor order because the score sort is stable. `SelectionOutcome` now returns those separately as `ranked_only`, and the activation path does not iterate them. They still reach the model through the listing, which is where they belonged. 2. `AlwaysAvailable` also enabled the name/description fallback, which manufactured fake merit: a bundled skill whose description shares one word with the message scored above zero and was reported as a genuine activation. Under `AlwaysAvailable` the fallback adds no reach at all (the floor already admits everything), so it is now scoped to `NameAndDescription`, where widening the match is the entire point. This is what kept `local_dev_runtime_suppresses_explicit_setup_skill_when_workspace_marker_exists` failing after (1). 3. Raising TOOL_RESULT_RECORD_READ_MAX_BYTES to 64 KiB was NOT the no-op this PR claimed. `tool_result_reference.rs` derives MAX_MODEL_OBSERVATION_BYTES from it (* 2), so the observation envelope silently doubled 48 KiB -> 128 KiB and preview truncation changed for every caller. It broke three tests whose fixtures are sized against the envelope ("fixture must exceed the preview cap"), independently of any activation setting. The contract ceiling is back to 24 KiB and the env override is bounded by a new TOOL_RESULT_READ_ENV_CEILING_BYTES that nothing is derived from -- so the knob can raise a single read without moving anyone else's behavior. Correcting the record on an earlier comment in this PR: the failures were never the setup-marker interaction. Marker suppression returns None before scoring, so a floor score cannot revive a suppressed skill. cargo test -p ironclaw_reborn_composition --lib # 634 passed IRONCLAW_REBORN_SKILL_ACTIVATION=always_available cargo test -p ironclaw_reborn_composition --lib # 634 passed cargo test -p ironclaw_skills --lib # 241 cargo test -p ironclaw_threads --lib # 88 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * revert(skills): remove the always_available strategy, it bought nothing Verified against this branch: `AlwaysAvailable` was a no-op for everything the model can observe, and it carried a regression. Removing it rather than wiring a compensating half. Why it bought nothing. Listing membership is decided by VISIBILITY, not selection (extension_ports/activation.rs partitions candidates on body-eligibility, and everything not body-eligible still goes into the listing), so the model was ALREADY shown every visible skill before this strategy existed. The floor score never added reach -- pre-C1 its only effect was listing ORDER, and after C1 excluded floor-only skills from activations even the ordering effect was gone, because the ranking input is derived from the activation list. `SelectionOutcome::ranked_only` had no production reader at all: allocated, populated, returned, dropped. Under `Full` a floor-only skill could never be injected either, since `context_candidates_for_plan` renders only activated bundles. The regression. The floor-only bookkeeping ran for every non-merit entry BEFORE `try_select`, so under this strategy a chain-loaded companion got its own loop iteration, was recorded as floor-only, and was then partitioned OUT of `selected` -- i.e. `A requires B` activated only `A`, where `CriteriaOnly` activates both. Strictly worse than the default for any bundle with companions, and order-dependent. The comment claiming this could not happen was wrong. Also removed: ~29 "budget exhausted" notes per turn that reached `feedback` and fired a SkillActivation live-projection event with empty skill_names, because floor-only skills still ran the budget loop and `BudgetFull` continues rather than breaks. Kept: `NameAndDescription`, which has a real effect (matching on name/description, not only `activation.keywords`/`tags`/`patterns`), and the `skill.activation.v1` seam. Corrects the record in two places that argued the opposite: the runtime.rs doc comment and docs/skills/agent_authored_bundles.md. The measured reachability gap is elicitation, not filtering -- `builtin.skill_activate` was called in 3 of 30 runs and a body read in 0 of 30 -- so the next step is the listing header, not a scoring change. Note the parity numbers in nearai/benchmarks#327 never depended on this strategy: those arms ran with it off. cargo test -p ironclaw_reborn_composition --lib # 634 passed cargo test -p ironclaw_skills --lib # 240 passed cargo test -p ironclaw_threads --lib # 88 passed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(skills): show that an agent-authored skill carries scripts This PR is what lets an agent author a skill containing a script. The Skills page could not show that it did: `skill_info` hardcoded `has_requirements: false` and `has_scripts: false`, so a scripted skill was indistinguishable from a prose-only one. The WebUI has rendered the chips since nearai#6194 and the wire fields have existed since nearai#7002 -- only the server never populated them. It was not just agent-authored skills. `portfolio`, a BUNDLED skill, ships four Python scripts (`weekly_report.py`, `backtest_strategy.py`, `concentration_warning.py`, `alert_if_health_below.py`) and has always displayed as prose-only. - `SkillSummary::has_scripts`, from one stat on the bundle's sibling `scripts` path. Absent is the common case and is not an error, so only a genuine backend failure is logged -- a skill listing must never fail because a skill has no scripts. - The bundled-summary path reads it from the embedded bundle files, so `portfolio` reports correctly there too. - `has_requirements` comes from `requires_skills`, which was already on the summary. Verified on a live production server: 33 skills listed, `portfolio` the one reporting `has_scripts`, six reporting `has_requirements`. Note: skill scripts still cannot EXECUTE under hosted multi-tenant -- `HostedMultiTenant` + `SecureDefault` resolves to `ProcessBackendKind::None`, which strips `builtin.shell`. That is deliberate pending the tenant sandbox. So the chip tells a multi-tenant user their skill has scripts the agent can read but not run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(skills): pin the direct-install input contract at the capability boundary Four cases covering what a caller may and may not put in a `builtin.skill_install` input, asserted through runtime dispatch rather than against whichever helper currently normalizes the input. The normalizer has already moved once (host runtime -> ironclaw_extension_support, WS3) and is about to be merged across that move again. Written so the same four pass on both sides: a resolution that quietly re-tightens the inline arm fails the first one instead of silently dropping the capability this PR adds. The dividing line these pin is provenance, not shape: - `content` + `files` installs, and the script lands on disk verbatim - `bytes_base64` works on the direct arm too, not only the rewritten URL payload - `content` + `files` + `source`/`source_url` is still refused whole - a `../..` bundle path is refused and writes nothing outside the skill directory Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(skills): repair the post-merge build and close the review findings Build breakage from the merge: main added two `SkillSummary` construction sites (`skill_learning`, `lifecycle_product_service`) that this branch's new `has_scripts` field left incomplete. Both are fixtures whose assertions do not depend on it, so both set `false` with a note saying why. CI, all under `-D warnings`: - Four `result_read` cap items in `ironclaw_threads::contract` were `pub` in a private module (`unreachable_pub`). Nothing outside the crate reads them, so they are `pub(crate)`. - Four constant-value asserts (two in the same contract module, two in `activation_strategy`) move into `const {}` blocks, which is what they always meant: they are compile-time invariants, not runtime checks. Review findings: - The stale-default doc note (coderabbit, ironloop) was against `ecabcb5fe`, before `4951d76bb` reverted the flip. Docs and `DEFAULT_SKILL_INJECTION_MODE` both say `Listing` with `full` as the opt-in, so there is nothing left to correct. - The unset-env branch is now reachable from a test (coderabbit). `skill_injection_mode_from_env_value` takes the lookup's `Result`, so the product default can be asserted without `remove_var` racing every other test in this binary. Covered along with `full`, trimming/case, empty, unrecognized, and non-unicode. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(skills): state accurately why the inline arm now accepts a bundle The comment claimed a review decision had been "reversed on evidence", which overstates what happened and would read as overriding the team. What actually happened: the refusal predates nearai#7141 entirely -- it lived in the host-runtime copy of this resolver, and nearai#7141 carried it across the move to this crate verbatim, declining a reviewer's suggestion to relax it there. That was the right call for a move-only refactor. This PR is where the behavior change belongs, and it is made deliberately with the measurement attached. Comment only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: raise the composition mass ceiling for this layer's assembly wiring CI's "Check composition mass budget" step reds this branch by 62 LOC. Not because this branch is large: main sits only 54 LOC under the effective ceiling (40,499 + 150 tolerance against 40,595 observed), so the gate currently trips on any PR adding more than that to composition, and this one adds skill-summary and product-surface assembly. Raised to the measured 40,711 in both places the gate pairs — `[gate].loc_ceiling` in the manifest and `COMPOSITION_ABSOLUTE_SRC_LOC` in `reborn_restructure_baselines.rs`, since a second ratchet fails when they disagree, which is how it enforces recording the change in the PR that causes it. Measured with `check-composition-budget.sh --print`, set to current rather than padded, per the manifest's own protocol. A raise is a reviewed decision by that file's rules, not routine wiring, so it is flagged here and in the PR body rather than left in a diff. The next wave close should re-ratchet. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(threads): make the result_read env knob actually raise the cap serrrfirat, High: `IRONCLAW_TOOL_RESULT_READ_MAX_BYTES` was inert. It widened `validate_tool_result_record_read`, which sits DOWNSTREAM, while the caller-facing gate in `result_read.rs` stayed pinned to the compile-time `TOOL_RESULT_RECORD_READ_MAX_BYTES` (24 KiB) and the advertised schema still said `maximum: 24576`. A larger read was rejected before it could reach the widened validator, so setting the variable changed nothing. The gate and the schema now resolve `effective_tool_result_read_max_bytes()` per request. This also corrects a fix I made earlier in this PR for the wrong reason. Clippy flagged `effective_tool_result_read_max_bytes` as `unreachable_pub` and I narrowed it to `pub(crate)` — but it had no cross-crate caller precisely BECAUSE the wiring was missing. The lint was reporting the bug, not dead code. It is `pub` again, with the caller it was always supposed to have. Tested as a wiring identity (gate == effective cap, schema == gate) rather than by setting the env var: these tests run in-process and in parallel, so mutating process environment races every other test reading it, and the identity is exactly what regressed. Verified: workspace `cargo check --all-targets` clean, 12 test binaries green across loop_host and threads, fmt clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(threads): satisfy the production-target lint lane CI's PR lane lints the DEFAULT target set (lib + bins, no tests or examples) with \`--all-features\`; I had been running \`--all --tests --examples\`, and the extra targets masked both of these. - \`TOOL_RESULT_READ_ENV_CEILING_BYTES\` was widened to \`pub\` alongside \`effective_tool_result_read_max_bytes\` in the previous commit, but only the function is re-exported from \`lib.rs\`, so the constant was unreachable-pub. Only that function reads it, so it is \`pub(crate)\`. - \`result_read.rs\` no longer reads \`TOOL_RESULT_RECORD_READ_MAX_BYTES\` now that the gate resolves the effective cap, so the import goes. Verified with the lane CI actually runs (\`cargo clippy --workspace --all-features -- -D warnings\`, no test targets), plus fmt and the loop_host/threads suites. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(skills): cut the comment bloat on this layer Comment share of this PR's diff: 33% -> 26%, 471 -> 330 lines. The measurements that justify a constant stay; the retellings of how we got there go, since they are already in the commit history. One of these was a duplicate rather than verbosity: `DEFAULT_SKILL_ACTIVATION` carried an earlier draft stacked directly above its own replacement, so the file gave two competing accounts of the same default. I had removed that copy at the top of the stack only; removing it here means all three layers carry one version instead of conflicting on every merge. Also corrects a doc that contradicted the code, which Copilot flagged: `effective_tool_result_read_max_bytes` was documented as clamping to `TOOL_RESULT_RECORD_READ_MAX_BYTES` when it clamps to `TOOL_RESULT_READ_ENV_CEILING_BYTES` — the whole point of the separate ceiling. Comments only; no code touched. `--all-features` clippy on the production target set (the lane CI runs) clean, fmt clean, 19 test binaries green across skills / threads / loop_host / extension_support. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(skills): move the design doc under the fenced internal tree `docs: enforce the docs/ publication boundary` landed on main at 00:53 and dequeued this PR: every file under `docs/` must be either published (referenced from `docs.json` navigation) or fenced (by `docs/.mintignore`), and `docs/skills/agent_authored_bundles.md` was neither — so it would have been deployed to the public docs site and indexed. That is exactly what the gate exists to catch, and the doc predates the rule rather than breaking it. `.mintignore` is frozen by that same change ("all new internal material goes under internal/"), so the fix is the move, not a new fence entry. `docs/internal/` is already fenced. `scripts/ci/test_docs_publication_boundary.py`: 21/21 pass. Workspace `--all-features` clippy on the production target set clean, fmt clean, panic gate clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This was referenced Aug 7, 2026
l3ocifer
pushed a commit
to l3ocifer/frick-ironclaw
that referenced
this pull request
Sep 3, 2026
…te the eight doc-truth corrections (nearai#7155) * test(architecture): itemize extension_host's product references as a frozen ledger The products -> loops re-layer has been sized five times from proxies and was wrong five times (D-A's one-file seam, nearai#7092's twelve files, three more since). The trait residue is trait-shaped and cannot see constants, free functions, or inline concrete construction; the manifest biconditional sees the sum but only as a boolean. This adds the itemization: EXTENSION_HOST_PRODUCTION_FILES_STILL_NAMING_PRODUCT — exact-match in both directions, shrink-only under a baseline ceiling, one reason per file, on a whole-token crate matcher (the raw-substring helper would count ironclaw_product_contracts importers). A ledger<->manifest consistency assert keeps the itemization and the biconditional agreeing about whether the edge exists, so the ledger cannot read empty while the manifest still carries the dependency. Sabotage-verified before trusting it: a planted production file naming ironclaw_product reds the gate naming that file; a planted stale row reds the stale direction; comment and string-literal mentions do not register (channel_delivery.rs and skill_learning.rs are the standing comment-only exclusions, and channel_subject_routes.rs's usage is cfg(test)-only). Part of the WS2 re-layer re-scope (nearai#7145). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): amend D-A — re-cite its precedent, record unstarted execution The ruling's two measured legs stand. Its option-(b) refutation cited ChannelWorkflowStateFactory as the in-file precedent; measured, that trait is a sole-impl same-file convenience no architecture test names — the load-bearing precedent is the landed nearai#7004 operator inversion and the ten INVERTED_PORT_IMPLEMENTORS ports, so the amendment re-cites it. Also records that the factory port exists on no ref (decision, not partial execution; shape still open) and that the residue is now mechanically itemized by the reference ledger. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): correct the Wave 2 closeout's twelve-file sizing Six of the twelve were re-export repoints, executed in nearai#7143. The enforced remainder is four reference classes (trait residue, adapter-registry, product free functions, D-A assembly), now itemized mechanically by the reference ledger, with the inventory carried by nearai#7145. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): correct the Wave 3 milestone and the W7 label reading The milestone was wrong three ways: the start was 13 (live register is 6); the ratchet is ceiling-only and pins nothing; and zero is WS12's gate, not this wave's reachable exit — the lane edges are nearai#7067's (whose measurement refutes the WS3 mcp row's vocabulary premise) and conversations->turns is WS5's. Also records that removes_in=W7 is a retired July-train label (nearai#5852 era, introduced 2026-07-09), not Wave 5 or WS7. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): give conversations->turns an owning WS5 row; key the register to it The exception's removal condition lived only inside WS1's verify-row explanation — no row owned it, which is how a milestone silently expires (its removes_in=WS5 date already passed once without it falling, as PROPOSAL 8.3's 2026-08-02 amendment records while asking for exactly this re-milestone). Adds the owning WS5 slice row, re-keys the register entry to it, re-keys host_runtime->extension_support to WS3, documents that W7 is the retired July-train label (not Wave 5 or WS7), and marks the WS5 product-narrows adapter_registry clause as a prerequisite of the extension_host re-layer (nearai#7145). The two entries nearai#7141 deletes and the two it re-keys to nearai#7067 are deliberately left untouched here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): refute the stale 'delete reasoning.rs (dead)' claims Sections 9 (row 29) and 12.4 still said delete-it-outright while WS8's own execution (nearai#6964) deleted only the dead half and the surviving module is live on main (mod reasoning; + re-exports in llm/src/lib.rs). Acting on the rows as written would have deleted production surface. 6.4.13's own line is amended by the in-flight nearai#7128 and deliberately not touched here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): amend 8.3 — row 7's blocker refuted by nearai#7067; live register is 6 Row 7 promised the lane->resources edges dissolve as vocabulary; nearai#7067's measurement (raised on nearai#7065) shows the vocabulary is already in host_api::resource and imported from there — the real holders are ResourceGovernor (3 of 10 methods used) and the ResourceError cone, whose relocation is an authority carve-out. Also refreshes the live register to 6 post-nearai#7094 and notes the conversations->turns re-milestone landed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
l3ocifer
pushed a commit
to l3ocifer/frick-ironclaw
that referenced
this pull request
Sep 3, 2026
…ations sever + WS10 inventory keying + enforcement gates (nearai#7170) * refactor(contracts): move extension runtime descriptors to a neutral contract (WS3) Deletes the two `-> ironclaw_extensions` layer-matrix exceptions (`ironclaw_mcp`, `ironclaw_scripts`) by giving the runtimes-layer lanes a contracts home for the descriptors they read, instead of the registry crate they may not depend on. Exceptions 13 -> 11; baseline lowered in the same change. Moved to `ironclaw_extension_contracts`: - `runtime::{ExtensionRuntime, ExtensionAssetPath, ExtensionAssetPathError}` - `hosted_mcp::{HostedMcpDiscoveredTool, HostedMcpDiscoveredToolAnnotations}` `ExtensionPackage`/`ExtensionManifest` deliberately stay in `ironclaw_extensions`: they carry the whole parsed manifest tree and a `PackageRootBinding` typed on `ironclaw_filesystem::VirtualPath`, which the §11.2.3 contracts-purity allowlist (`{ironclaw_host_api}` only) forbids the contracts crate from naming. Measured instead: both lanes read exactly three things off the package — `id`, `capabilities`, `manifest.runtime` — so the lane request structs now take those three and the caller (which owns the package) projects them. Also repointed `ResourceReceipt` to its real owner: `ironclaw_resources` only re-exports `ironclaw_host_api::resource::ResourceReceipt`, so the lanes' import was a §11.2.4 two-import-paths hop, not a dependency. No `pub use` shims (§11.3): every consumer is repointed in this change, and `resolve_under` becomes the free function `ironclaw_extensions::resolve_asset_under` because the orphan rule forbids an inherent impl on the moved type. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(sandbox): merge the sandbox lane into one crate (WS3) Creates `ironclaw_sandbox` (runtimes) from the three halves of "run an already-authorized command away from the host", and deletes the two crates PROPOSAL §6.6.4 marks for merge: - `ironclaw_process_sandbox` (plan contract) -> `src/plan.rs`, `src/validation.rs` - `ironclaw_host_runtime::sandbox_process` -> `src/sandbox_process/**` - `ironclaw_scripts` (script lane + Docker path) -> `src/script.rs` The kernel sheds the Docker/CA cone: `bollard`, `rcgen`, `x509-parser` and `time` are gone from `ironclaw_host_runtime`'s manifest, and `bollard`/`rcgen` are now declared by exactly one crate in the workspace. Two migration details PROPOSAL §6.6.4 and CHECKLIST WS10 call load-bearing: - `PROCESS_SANDBOX_CAPABILITY_ID` -> `ironclaw_host_api::capability`, so `ironclaw_loop_host` drops its lane dependency (production dep gone; a dev-dep remains for the tests that build plans). - `SandboxCommandTransport` -> `ironclaw_host_api::process`, with the shapes it names (`CommandExecutionRequest`/`Output`, `RuntimeProcessError`, `SavedCommandOutput`, `SavedCommandOutputSanitization`). Without this the runtimes-layer lane could not implement what the kernel consumes. Enumerating gates were repointed, never relaxed: the specificity carve-outs and the struct/test-support ratchet entries moved with their files (both baselines unchanged at 129 and their prior values), the panic-gate baseline row moved, `reborn-crate-test-buckets.sh` registers the new crate, and the three `reborn-e2e-rust.sh` script selectors follow the tests (plus `docker_security`, which had no selector before). One gate would have gone silently vacuous and was fixed rather than moved: the script-lane surface scan in `reborn_dependency_boundaries.rs` read a hardcoded `src/lib.rs`, which after the merge no longer holds the lane. It now scans the whole crate source tree with a fatal-read walk and a non-vacuity assertion. One deletion, recorded: `RebornScopedSandboxCommandTransport::into_process_port` returned a kernel type a runtimes crate may not name. It had zero callers workspace-wide; the kernel wraps the transport, which is the direction the port inversion requires. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): record the WS3 corrections with their evidence Three dated amendments, each quoting the text it replaces: 1. CHECKLIST WS3 sandbox row + PROPOSAL §6.6.4 — "all pieces currently unwired/test-only" is REFUTED. Three production paths cross the merged crate (spawn-path plan validation, the process_executor routing check, and the saved-command-output scope digest). The accurate claim is narrower: no production *execution backend*. Behavior preservation is therefore argued at the diff (11 of 26 moved files byte-identical, 9 more differing by one import line, +63/-36 overall), not inferred from deadness. 2. CHECKLIST WS3 mcp row + PROPOSAL §6.6.3 — the prior wave's "structurally blocked" finding is half right, and the wrong half is load-bearing: only `ExtensionPackage` is un-absorbable, and no lane ever needed it (both read `id`, `capabilities`, `manifest.runtime` and nothing else). The registry half of the flip is done; the `resources` half is refuted as phrased — the estimate/usage vocabulary the row asks about is already in `host_api::resource` and already imported from there, while the real blocker is the `ResourceGovernor` authority port and `ResourceError`'s denial cone. 3. Recorded as a structural finding, not a note: the sandbox row and the mcp row are ONE problem. `ironclaw_scripts` imports the identical DTO set, so the merge alone deletes zero exceptions and only the mcp carve-out lets either lane shed the registry edge. Also reconciled: PROPOSAL §6.1.2's as-built inventory gains the two modules WS3 landed (and states why `ExtensionPackage` stayed); §2's package count 66 -> 65; the §9 disposition rows for `ironclaw_scripts`/`ironclaw_process_sandbox`/ `ironclaw_mcp`; the §11.2.2 ratchet rows (13 -> 11); the WS3 verify row; the stale WS1.3 sentence asserting the blocker as settled fact; and `reborn_restructure_baselines.rs`'s doc table, which still read 15. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(sandbox): drop imports the merge left unused `process_port.rs` no longer names `MountView` or `thiserror::Error` (both went to `host_api::process` with the types that used them), and `sandbox_process.rs` no longer needs `sync::Arc` after `into_process_port` was deleted. Found by per-crate `clippy --all-targets --all-features -D warnings`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): let the Reborn PR planner plan guidance edits and crate deletions Three fail-closed gaps in `reborn_pr_test_plan.py`, all hit by this PR and all live on `main` today — any PR with the same change shape is unplannable. 1. `.claude/**` was unclassified, so the planner refused outright. It is agent guidance in exactly the sense `docs/**` is human guidance: no Rust test reads either as data (the only in-tree references are prose citations in test doc comments). Added to `IGNORED_PREFIXES`. Without this, "guidance travels with the change" — the restructure's own discipline — cannot be satisfied in a single PR. 2. `crates/AGENTS.md`, `crates/README.md`, `crates/Architecture.md` raised "unmapped crate path": they sit under `crates/` but belong to no package. Now classified as crate-tree prose, matched by "Markdown no package directory owns" so a genuinely unmapped crate path is unaffected. 3. An unmapped crate path used to raise. `git diff` reports a deleted crate's old paths and CI feeds the planner that diff, so **every crate deletion or rename was unplannable** — including the six deletions PROPOSAL §2 plans. It now widens to the exhaustive plan. This is a semantic change and it is the safe direction: the full plan is a superset of any narrowing, so an unattributable path can never cause under-selection, whereas refusing to plan blocks the PR instead of protecting it. Malformed input is still rejected by the unclassified-path branch. Each lands with fixtures per WS10's rule, positive and negative: guidance paths select nothing while non-guidance paths still fail closed; crate-tree prose selects nothing while crate *code* under the same unmapped directory widens to `full` (so the Markdown carve-out cannot swallow code). The pre-existing `test_unmapped_crate_path_fails_fast` is renamed and rewritten to pin the new contract rather than deleted. Verified against this PR's real 130-path diff: the planner returns `mode: full`, and the workflow's own exhaustiveness guard passes on that output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(arch): give the retained resource exceptions an owning issue, not a wave Review (#7065) caught that both surviving `-> ironclaw_resources` exceptions declared `removes_in = "WS3"` — the wave this PR *is*, which does not remove them. That is precisely the defect §11.2.2 already records against `conversations -> turns` ("`removes_in = "WS5"` and WS5 has partly shipped without it falling"), and it would have been repeated here. Both now point at issue #7067, which owns the design work that actually clears them: replacing the `ResourceGovernor` dependency with a narrow reserve/reconcile/release port. The issue carries the measurements — 3 of 10 methods used, zero implementors, and the `ResourceError` denial cone — plus the two open questions (error shape, port home) that make it a design slice rather than a move. An owning issue is also what §11.2.2 asks for and what the ratchet still cannot enforce (there is no `owning_issue` field yet), so this is the strongest form currently expressible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(contracts): pin the asset-path validator that moved into extension_contracts `validate_asset_path` moved here with `ExtensionAssetPath`, the type it constructs. In `ironclaw_extensions` it was only ever reached indirectly through manifest parsing, so its six rejection branches had no direct test — and a contracts crate that carries validation owes that validation one. Two tests: every reject branch with its exact reason and `Display` output (empty, NUL/control, URL, absolute, Windows drive and backslash, and the empty/`.`/`..` segment cases) plus the manifest-relative shapes that must keep being accepted; and `ExtensionRuntime::kind()` over all five variants, since that projection is what every lane uses to reject a runtime it does not serve. Also removes a changed-line coverage risk this PR would otherwise carry into the merge queue: the gate does not run on ordinary PRs (#7036), so ~100 newly-added lines of validator would first be measured where a failure is expensive to diagnose. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(coverage): re-capture the host_runtime floor and floor the new sandbox lane `RATCHET FAIL: ironclaw_host_runtime` — observed 18854 covered vs a `floor_covered_lines` of 20538. This is the shrinkage case the ratchet's own "To fix" text describes, not a coverage regression: `sandbox_process/**` moved to `ironclaw_sandbox`, so the crate's denominator fell 23277 -> 21267 (-2010 instrumented lines) and its covered lines fell with it. The percentage floor is **raised, not lowered**: observed 88.65% against an old floor of 88.23%, so the entry now reads 88.65. Only the absolute line count moves down, and it must — those lines are no longer in this crate. To keep that from being a net loss of protection, `ironclaw_sandbox` is floored on arrival at its observed 87.09% (3185 / 3657). This is a net *increase* in ratchet coverage: neither `ironclaw_scripts` nor `ironclaw_process_sandbox` was ever floored, and the `sandbox_process` half was protected only as part of host_runtime's line count, which this PR necessarily reduces. Floored crates 16 -> 17. Verified by replaying the ratchet arithmetic against CI's observed numbers: both crates pass on percentage and on covered lines. Numbers taken from the failing run's own report (job 91740733521), which is the authority for this gate. The `Tests (Reborn)` roll-up failed solely on this sub-job ("coverage-report result 'failure' did not match planned=true"); no other lane failed — 50 pass, 2 fail, both this root cause and its roll-up. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): record the coverage ratchet as a move-sensitive gate WS3 hit a gate no move row had named. `tests/integration/coverage-floor.toml` is keyed on crate identity plus absolute covered-line counts, so it is invisible to WS10's path-keyed gate audit and yet it fails on every crate move, merge, rename, or family `git mv` that shifts instrumented lines between crates — as it did here, while the percentage floor was *improving*. Recorded on WS10 with the three rules WS7 will need: re-capture in the same PR, raise the percentage floor rather than leaving it, and floor the destination crate or the move silently drops that code out of the ratchet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(extension-manager): repoint ironhub onto the moved ExtensionAssetPath A semantic conflict the merge could not see: #6780 landed `ironhub/{package,catalog}.rs` importing `ExtensionAssetPath` from `ironclaw_extensions`, while this branch moved that type to `ironclaw_extension_contracts::runtime`. Different files, so git auto-merged cleanly and the breakage surfaced only at `cargo check`. Repointed both sites to the contracts crate (no shim, per §11.3). The manifest already named `ironclaw_extension_contracts`, so this is imports only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(coverage): exempt the WS3 move's no-region lines and record the gate The changed-lines coverage gate went red on four files while changed-line coverage was 95.35% against a 90% floor: the failure was its two fail-closed STRUCTURAL assertions, not any percentage. Every line below was derived by replaying scripts/ci/reborn_changed_coverage.py against this PR's own merged lcov (run 30831658659) with the base lcov the gate itself resolved (run 30828540055 @ b89fcd3575), until the replay reproduced the CI verdict byte-identically. Line numbers come from the gate's own `candidate_lines - mechanically_uninstrumentable_lines()`, not from the log. - host_api/src/process.rs (31 lines): new placement-neutral process vocabulary with no function body anywhere in the file; rustc emits no LCOV record for it at all. Same shape already exempted for product_contracts/loop_contracts. - extension_contracts/src/hosted_mcp.rs (12): field declarations of the two new tools/list descriptor structs. The file is plainly instrumented (191 DA, 164 hit), so this is a no-region artifact, not an instrumentation gap. - host_runtime/src/services/runtime_adapters.rs (13): continuation lines of three rewritten calls, all PROVEN EXECUTING by their region-start heads (lines 380/434/977 score 24/16/63 hits). The four genuinely-uncovered lines in the same rewrite are deliberately NOT exempted -- the gate already subtracts them as pre-existing debt inherited from base. - composition capability_host_tests/approval_gates.rs (6): type positions in a test double whose body region scores 1 hit. The last one is a finding, not just a waiver: that file is 100% test code behind `#[cfg(test)] mod capability_host_tests;`, but the gate's test_only_path() recognises /tests/, /test_support/, */tests.rs and *_tests.rs and NOT a cfg(test) module DIRECTORY, so it measures it as production. It is the only such directory in crates/ today. Docs (target-architecture, same PR per the docs-truth rule): - CHECKLIST WS10 gains the changed-lines gate beside the ratchet row, cross- referencing the WS2.1 note rather than restating it: percentages are not what fail a move; derive lines by byte-identical replay (--fetch-base-coverage silently degrades without --github-repo); and a stranded exemption path is an ABORT with no verdict, not a loud failure. - CHECKLIST WS10 exception-ratchet row: the constant was cited at :4063 and sits at :4164 -- corrected by removing the line pin, since the file is edited every wave. Records that the baseline is a UNION across parallel WS3 lanes. - families/contracts.md: records extension_contracts' new ownership of the runtime descriptor vocabulary -- the carve-out that let BOTH lanes drop the registry edge -- and the orphan-rule seam that keeps resolve_asset_under in the registry crate. - families/lanes.md: two "Never" claims were reading as satisfied when they are not. ironclaw_mcp's "never depends on the resource-governor crate directly" is refuted (the compiled edge survives; #7067 tracks the narrow port), and ironclaw_sandbox's "no direct process spawning outside the transport seam" is aspirational -- script.rs:454 still builds Command::new("docker"). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sandbox,mcp): correct the wiring inventory and record the projection cost Two review findings verified against the tree; three refuted with evidence in the PR threads. Valid — the sandbox wiring inventory was self-contradictory. `CLAUDE.md` said "Two production call paths ... and both are plan validation" directly above a list of THREE bullets, and `lib.rs` omitted the third entirely. The third is real and is not validation: `host_runtime/src/process_output.rs:482` derives the scoped saved-output directory through `RebornSandboxScopeKey::from_scope`. That inventory is what tells a future agent which paths are live, so an undercount invites deleting a production path as dead code. Both surfaces now say three and no longer claim they are all plan validation (the `loop_host` capability-id comparison never was either). Valid, and recorded rather than redesigned — the registry carve-out cost a type-level invariant. Replacing `package: &ExtensionPackage` with independent `extension` / `capabilities` / `runtime` borrows is what deleted the `mcp -> extensions` and `scripts -> extensions` exceptions, but it also means the type no longer guarantees the three came from one package. `execute_extension_json` re-checks the descriptor half (`descriptor.provider == extension`); the runtime half cannot be re-derived, because nothing in an `&ExtensionRuntime` names its owning extension. No caller can trip it today -- there is exactly one production caller (`runtime_adapters`) and it projects all three from one package in one expression -- so this is a latent structural weakening, not a live defect. Restoring the compile-time binding needs a sealed projection minted by the package owner; a check inside the lane cannot express it, and re-taking the registry edge would undo the carve-out. Both request types now carry the caller obligation in their field docs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(extensions): move the skill-install executor to extension_support (WS3) WS3's first-party-tools row, family 1 of 6: skill management / URL install. `skill_url_install.rs` and its `bundle`/`github`/`zip_bundle` submodules, plus the install-input normalizer, move out of `ironclaw_host_runtime::first_party_tools` into `ironclaw_extension_support::skills::{url_install, resolve_install_input}`, where the skill executor half already lived. Move-only: no behavior change, no test edited for content. `ironclaw_host_runtime -> ironclaw_skills` is deleted from LAYER_MATRIX_EXCEPTIONS — the edge is gone, not waived (exceptions 13 -> 12, WS0_LAYER_MATRIX_EXCEPTION_BASELINE drops with it). `ironclaw_skills` and `zip` survive as dev-dependencies for host_runtime's own tests; dev edges are outside the matrix by construction. Two doc ambiguities are resolved in the same diff, as dated PROPOSAL amendments quoting the text they replace: - §6.8.4's "the builtin first-party tool handlers absorbed from host_runtime/first_party_tools" contradicted §8.2's "kernel: ✗ (ports only)" row and the enforced BoundaryRule. Resolution: the seam splits executor from adapter — the executor moves behind a neutral request/error pair, the FirstPartyCapabilityHandler / CapabilityManifest / registry wiring stay host-side. Same shape the groupware and web-access tools already ship. - §8.2's "ports only" cell now says what it means: contracts-layer ports the kernel also consumes, not permission to name a kernel trait. Two cost corrections recorded for the remaining families: `host_runtime -> extension_support` is not divisible family-by-family (mod.rs holds it via `extension_support::coding`), and `host_runtime -> ironclaw_extensions` is not reachable by this row at all. PATH_TERM_COLLISIONS shrinks by two: the installer's github carve-outs now sit inside a scan-exempt crate. Test accounting (un-masking discipline), unfiltered `--list` over both crates: 1398 -> 1398, with exactly two tests renamed by module path and none lost. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sandbox): record that the Docker fail-closed switch is wired to nothing Review asked why the migrated docker_security test can pass with no daemon. The skip is pre-existing (the file differs from its pre-merge original by one import line); WS3 only enrolled it in the required Rust e2e lane, where it was not run at all before. The real defect the question surfaced is worse and also pre-existing: this crate's tests/support/docker_gate.rs states that IRONCLAW_REQUIRE_DOCKER_TESTS=1 makes a missing daemon a hard failure and that "CI sets this" -- and nothing sets it. Repo-wide the name occurs only in docker_gate.rs and attribution_tests.rs, here and on main. So every real-Docker test in the crate skips-and-passes everywhere, which is exactly the gap the gate's own comment says let sandbox security bugs ship unnoticed. docker_security.rs additionally open-codes its own check rather than using the gate, so it would stay fail-open even once something did set the variable. Recorded rather than fixed: setting the variable is a CI-behavior change that would hard-fail any lane without a daemon or the ironclaw-worker image, which is not verifiable from inside a move PR whose evidence claim is behavior preservation. Filed as the #6945 guardrail-claim-vs-reality class with the two-part fix stated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(host_runtime): record the executor/adapter seam in crate guidance The crate's CLAUDE.md said "first-party runtime tools belong under `first_party_tools/`" without saying that only the host half does. WS3 moves each tool's executor into `ironclaw_extension_support`, which may not name this crate, so the rule now names both halves and points at the skill-install family as the worked example. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(host_runtime): keep the install-input error path log-free The moved executor returns `SkillManagementCapabilityError`, and routing it through `skill_management_error` would have added a `debug!` line to a path that had none before the move. A move-only change must not add one, so the install-input arm maps the kind directly and the `dispatch` arm keeps the record it already had. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci(coverage): re-capture the host_runtime floor for the WS3 executor move The ratchet does not run on `pull_request` (`reborn_pr_test_plan.py:21`; issue #7036), so this PR's green checks were not evidence on this axis. A full-plan `workflow_dispatch` run on this exact head reported: RATCHET FAIL: ironclaw_host_runtime observed: 88.59% (20485 / 23124 lines) floor: 88.23% ... floor_covered_lines: 20538 (effective floor 20518) The percentage went UP while `floor_covered_lines` went DOWN — shedding well-covered code lowers the absolute numerator, which is a separate assertion from the percentage one. Re-captured to the observed numbers (floor raised 88.23 -> 88.59, not merely held). Verified locally against that run's own merged lcov artifact: ENFORCING mode, 17 PASS / 0 FAIL, exit 0. run: https://github.com/nearai/ironclaw/actions/runs/30858257594 head: e07b3b0299b0add11117e9591da71d46d7a7c832 The destination crate is deliberately not floored, because it cannot be: every crate under `crates/extensions/` is invisible to the coverage tooling — `reborn_coverage_lcov.py:19`'s CRATE_RE still requires a crate directory directly under `crates/`, which #7037's colocation broke. Filed as #7083 with the measurement; the global floor is left alone rather than re-captured onto that hole. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(wasm): move wit/ inside its owning crate (Wave 3) CHECKLIST WS4 + WS10 `wit/` rows. `wit/{tool,channel}.wit` moves from the repo root to `crates/ironclaw_wasm/wit/` — the crate that owns the ABI — per PROPOSAL §6.6.1. Behavior-free: same bytes, same generated bindings. Wave-3 coordinates: the docs write the destination as `crates/lanes/ironclaw_wasm/wit/`, but `crates/lanes/` does not exist until WS7. Because the files now sit *inside* the crate, the WS7 family move carries them with no further path edit anywhere — which is the whole point of putting them there. Ten wit-bindgen `path:` args repointed (the host plus nine guests: six under `crates/extensions/packages/*/wasm-src/`, three under `test-tools/*/wasm-src/` — the CHECKLIST row said six). All nine guests verified building against the moved WIT on wasm32-wasip2. The four `include_str!` readers of the ABI text do NOT get repointed literals. Doing that would turn the two `ironclaw_host_runtime` sites from repo-root reach-ins into *cross-crate* ones — §11.2.7's strict class, the one WS2 turns into hard failures — taking the scan from 19 to 21 while ticking a box that says "§11.2.7 scan passes". Instead the ABI text gets one owner, `ironclaw_wasm::TOOL_WIT` (`src/config.rs`, beside `WIT_TOOL_VERSION`), and all four sites read the const over cargo edges that already exist. Measured with the scan: 133 -> 129 escaping sites, cross-crate 19 -> 19, zero `wit/` entries remaining. Path-keyed gates repointed: `scripts/check-version-bumps.sh` (both ABI paths), `.githooks/pre-commit`, and `platform-and-compat.yml`'s `has_direct_wasm_abi_risk` filter — where the bare `wit/` alternative is *deleted* rather than rewritten, because the filter's existing `crates/([^/]+/)*ironclaw_wasm/` alternative already matches both the Wave-3 and the WS7 location. `scripts/ci/ws12_workflow_contracts.py` anchored on that deleted string, so its anchor moves to `build-wasm-extensions` and its in-scope probe now pins both locations. `Dockerfile` loses two `COPY wit/ wit/` lines in the planner and builder stages: both already run `COPY crates/ crates/`, so the files arrive with the crate and the old line would COPY a path that no longer exists. Docs: the WS4 row's `crates/lanes/wit/` destination was the only doc site placing the directory beside the crate rather than inside it; corrected there and in README's tree, with dated amendments in CHECKLIST, PROPOSAL §6.6.1 and PLAN Wave 3 recording what the move found. Test accounting (unfiltered `--list`, name-by-name, quiescent tree): ironclaw_wasm 51 -> 51, ironclaw_host_runtime 1246 -> 1246, ironclaw_architecture 198 -> 198. Zero diff, no test edited for content. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * build(wasm): rebuild first-party artifacts for the moved wit/ path Forced by the previous commit, not incidental to it. `scripts/ci/check-wasm-artifact-freshness.py` keys each package's committed `wasm/<name>.wasm` to a digest of the `wasm-src/` tree that produced it, so editing a guest's `wit_bindgen::generate!` `path:` — which the `wit/` move requires in all six shipped guests — invalidates the recorded digest and fails the gate. The gate's own contract forbids the shortcut: "Re-record only after `./scripts/build-wasm-extensions.sh --first-party` and committing the rebuilt artifact — the digest asserts a claim about the artifact, and updating it without rebuilding launders a stale one." So the artifacts are genuinely rebuilt (`--first-party`, exit 0, 6 OK / 2 host-native SKIP), not re-recorded in place. Byte sizes move by more than the source change accounts for because these builds are not reproducible by design — the guests pin no toolchain and resolve their own `Cargo.lock` at build time, which is the documented reason the gate hashes sources rather than artifact bytes. Verified: `check-wasm-artifact-freshness.py` OK (6 packages), and `cargo test -p ironclaw_extension_support` green (102/46/4) — that crate `include_bytes!`s these artifacts, so it exercises the rebuilt components. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-arch): record the WS7 artifact-rebuild cost of guest path edits The `wit/` move had to rebuild six shipped WASM binaries because `check-wasm-artifact-freshness.py` digests each guest's whole `wasm-src/` tree. WS7 hits the same wall from the other direction: the six package guests reach the ABI across two trees, so moving either `ironclaw_wasm` or `extensions/packages` rewrites all six `path:` literals and forces the same rebuild. Recorded on CHECKLIST WS10's `wit/` row (point 6), on the loud-path-pattern row that owns the WS7 repoint (also corrected six -> nine guests there), and on PLAN's Wave 5 block with the cheap mitigation: move the two crates in one PR and pay it once. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci(planner): classify the path classes that blocked the wit/ move `Detect Reborn test scope` exits 1 on any pull request whose diff holds a path `reborn_pr_test_plan.py` has no rule for, which made this PR unmergeable: it must edit `Dockerfile` (the moved directory's `COPY wit/ wit/` no longer resolves) and `scripts/check-version-bumps.sh` (the ABI gate would otherwise grep dead paths and silently stop enforcing). 18 of its 46 paths were unclassified. Same class as the `.claude/` gap #7064 fixed, and classified the same way — one rule per class, recorded beside the constant: * `Dockerfile` / `.dockerignore` — `platform-and-compat.yml` keys `has_docker_risk` off exactly this pair and owns the image build. * `.githooks/**` — Code Style triggers on the tree and lints its contents (`test-ci-comm-locale-pin.sh`); no Reborn lane runs a hook. * `scripts/{build-wasm-extensions,check-version-bumps}.sh` — `platform-and-compat.yml`'s `has_direct_wasm_abi_risk` classifier both scopes and runs them. * markdown owned by no crate (`crates/AGENTS.md`, `test-tools/README.md`) — prose, like `docs/` and `.claude/`. A crate-resident doc still selects its own crate's lane. The first-party extension package assets are deliberately NOT ignored. `crates/extensions/packages/*/wasm/*.wasm` is a shipped artifact that `ironclaw_extension_support` embeds with `include_bytes!`, and `test-tools/*/manifest.toml` is `include_str!`d by `ironclaw_extension_host`. Calling either prose would convert today's loud failure into a silent under-schedule of a change to production output — the WS10 failure mode. `EMBEDDED_ASSET_OWNERS` routes each tree to the crate that compiles it instead, so this PR now additionally schedules `ironclaw_extension_{support,host,manager}`: the crates that consume the six rebuilt WASM artifacts. Also fixes #7085 in a file this PR already touches. The WIT version extractors used the GNU-only BRE `\+`, so on BSD sed (macOS) they matched nothing, and because the `WIT_TOOL_VERSION` cross-check is guarded on a non-empty version the hook printed "All version checks passed" having compared nothing. `[[:space:]][[:space:]]*` is identical under GNU sed, so the enforced Linux CI lane is unchanged; verified on BSD sed that both `wit/tool.wit` (0.3.0) and `wit/channel.wit` (0.3.1) now extract. Regression tests: every classified class gets a case in `test_reborn_pr_test_plan.py`, including the paired assertion that the embedded assets *select a lane* rather than merely being accepted (the inverse of the `.claude/` prose test), and a staleness pin that fails if an asset tree or its owning crate moves. All ten new cases fail against the planner on `main`. `test_unclassified_build_input_fails_fast` moves off `Dockerfile` onto a still-undecided input so the fail-closed arm stays exercised. Refs #7087, #7085 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(host-runtime): split obligations into its three chartered owners (WS3) `crates/ironclaw_host_runtime/src/obligations.rs` was 3,122 lines fusing the three owners PROPOSAL §6.5.9 charters separately, held apart only by an `// arch-exempt: large_file` waiver. It is now one module per owner: - `obligations::handler` — which obligations apply and what each does before/after dispatch, plus the audit/redaction/ceiling/mount validation. - `obligations::staged_handoffs` — material staged for a later consumer: the runtime-secret and network-policy stores and the credential-account resolver port. - `obligations::process_store` — post-start handoff discard and reservation reconciliation. - `obligations::mod` — only `BuiltinObligationServices`, the assembly seam, and deliberately the one place naming all three at once. Every module is under the 1,500-line gate, so the waiver is deleted rather than carried forward: re-fusing the owners now trips `pre-commit-safety.sh`. `mod obligations;` stays private and the crate's `pub use obligations::{…}` names are unchanged, so no consumer outside the crate sees this. Behavior-free. Cross-owner access is `pub(super)` (three methods), not `pub(crate)`. The split revealed one narrowing in the other direction: `secret_present` was `pub(crate)` with no caller outside its own file and is now private. Also from the same CHECKLIST row, the bounded half of "shrink `services/builder.rs` toward composition-facing factories": three builder methods whose only callers are inside the crate's `src` narrow to `pub(crate)`. The rest of that clause is measured and deferred in the CHECKLIST amendment — 17 methods need a `test-support` cargo feature, three are callerless and belong to WS8, and the remaining 33 are a redesign of the fluent surface rather than a shrink of it. `+production_wiring` is refuted there: it is readiness diagnostics, not assembly. Two loud path-keyed gates fired and were repointed, not relaxed: `reborn_host_runtime_services_do_not_expose_lower_substrate_handles` now scans the whole `obligations/` directory and asserts it read ≥ 4 files (`collect_runtime_rs` returns a count; both its callers now assert non-zero), and `reborn_struct_test_support_ratchet`'s frozen per-file count moves to `staged_handoffs.rs` with its count unchanged at 1. Test accounting (un-masking discipline): `cargo test -p ironclaw_host_runtime --all-targets -- --list` is 1,246 before and 1,246 after, name-by-name identical — zero added, removed or renamed. `LAYER_MATRIX_EXCEPTIONS` is 10 before and after; an intra-crate split cannot move the register. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(operator,contracts): route operator secrets through a product_contracts port (WS3) `ironclaw_operator` is a products-tier crate and held `ironclaw_secrets`, the substrate that owns CAS one-shot leases, AAD/crypto and the OS keychain master key. PROPOSAL §8.2's product row says the products tier loses that edge, and §12.1b requires the port replacement to land before the edge is removed. Both happen here, in that order. - Port: `ironclaw_product_contracts::operator_secrets::OperatorSecretValueStore`. - Implementor: `ironclaw_reborn_composition::RuntimeOperatorSecretValueStore`, the same placement as `OperatorStatusService` — assembly is the only layer that may name both a products-tier port and a substrate. Registered in `INVERTED_PORTS` beside it. - `ironclaw_secrets` is gone from the operator manifest under every dependency kind, and `"ironclaw_secrets"` is now in the crate's `boundary_rules()` forbidden list. That gate's comment previously said the entry was deliberately absent because "the row owns it"; the row now owns it. The port is deliberately narrower than the substrate, so this is a tightening rather than a relocation: it takes no `ResourceScope` (the implementor fixes the operator scope, where the caller used to pass one), exposes no lease/consume protocol, and carries only a `&'static str` classification instead of the substrate's error `Display` — asserted, including that the backend message and the handle name are both absent from what crosses. Two tests travelled with the behavior rather than being pointed at a fake: `read_is_repeatable_across_reloads` (repeatability is a property of the lease protocol) and the #4673 production-store reproduction (its value is wiring the store exactly as production does, which now means the real store *behind the adapter*). Two `FaultInjecting`-over-real-store fixtures became per-operation port fakes, with the substrate error mapping re-pinned at the adapter; a third assertion got stronger — batched-vs-N+1 stored-key lookup is now observed at the port rather than by counting filesystem ops. Test accounting: operator 154 -> 153, product_contracts 142 -> 143, composition 937 -> 942 with zero removed; name-by-name diffs on a quiescent tree. Two findings the row could not have anticipated, both recorded in the CHECKLIST amendment: - The `webui` half of the row was already closed and was never a production edge. `ironclaw_secrets` has been a dev-dependency of `ironclaw_webui` since the commit that added it (#6619), both src mentions are `#[cfg(test)]`, and webui's boundary rule already forbade it. - `ironclaw_extension_manager` (layer `products`) still holds a normal `ironclaw_secrets` edge in `admin_configuration.rs`. §8.2 covers it; the row does not, because the crate landed with WS2.4 after the row was written, and the substrate sits in the service's type parameters so it is not a like-for-like swap. Filed as #7095. `LAYER_MATRIX_EXCEPTIONS` is 10 before and after: `products -> substrates` is matrix-legal, so this edge was always an §8.2 rule and never a layer exception. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(sandbox): put the Docker security check behind the fail-closed gate Review asked why the required Rust e2e lane can report `docker_security` as passing with no daemon. Half of that is #7081 (nothing sets IRONCLAW_REQUIRE_DOCKER_TESTS=1, so the switch is inert) and is not fixable from here -- arming it hard-fails any lane lacking a daemon or the worker image, which needs a runner guaranteed to have both. The other half is fixable here and is fixed: docker_security.rs open-coded its own `docker version` / `image inspect` checks with three bare `return`s, so it sat entirely outside docker_gate and would have stayed fail-open even once something did set the variable. It now takes both preconditions from docker_gate::{docker_available, docker_image_available} and skips with the visible `SKIP:` line that gate's module doc requires. Measured, same machine, image absent: before, IRONCLAW_REQUIRE_DOCKER_TESTS=1 -> "skipping ..." / 1 passed after, IRONCLAW_REQUIRE_DOCKER_TESTS=1 -> panic at docker_gate.rs:74 / FAILED after, variable unset -> "SKIP: ..." / 1 passed The third line is the no-op proof: the variable is set nowhere in this tree or on main, so no lane's behavior changes today. The daemon-down path already reached the image check and skipped there, so the outcome is identical; only the branch it takes differs. Two stale comments in docker_gate.rs corrected with it (they claimed docker_security used its own gate, and that docker_image_available had no consumer), and the crate's Known debt entry now splits the done half from the #7081 half instead of describing both as open. cargo test -p ironclaw_sandbox: 193 passed, 0 failed cargo clippy -p ironclaw_sandbox --tests --all-features -- -D warnings: exit 0 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(reborn): stop calling the unwired script lane an execution lane Two review findings, both correct, both artifacts of this PR's own renames. 1. engine-v2-to-reborn-parity.md note 4 read "a native script/software execution lane (`ironclaw_sandbox`, `RuntimeKind::Script`) sandboxed via `ironclaw_sandbox`" -- self-referential after the merge collapsed ironclaw_scripts and ironclaw_process_sandbox into one crate, and it contradicts note 5 four paragraphs down ("no production execution backend is wired for it"). Re-stated as the typed runtime contract it is, citing the measurement: `with_script_runtime` has zero production callers (`rg` finds only the builder itself, docs, and 30 test call sites). 2. CHECKLIST WS10 ratchet note 2 said "raise the percentage floor ...; only the line count should fall". That generalises WS3's sandbox merge, where observed coverage happened to rise. It is wrong as guidance for WS7, and the counterexample is in this same file: the 2026-08-03 entry from #7064 records ironclaw_runner falling 85.55% -> 82.53% because the shed removed the crate's better-covered half, holding the floor, and RATCHET FAILing in the merge queue. Note 2 now says re-capture from the merged artifact, and lower only with that entry's move-not-regression counterfactual (add the moved files back, confirm the union clears the old floor, plus a zero-tests- lost name set-diff). cargo test -p ironclaw_architecture: 32 targets, 206 passed, 0 failed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): pin the WIT scope probes and the embedded-asset owner pairing Three review findings on the `wit/` move, each verified before it was acted on. 1. `ws12_workflow_contracts.py` probed `crates/ironclaw_wasm/wit/host.wit` and its nested twin. No `host.wit` exists in this repository — `git ls-files '*.wit'` returns only `tool.wit` and `channel.wit` — so both probes sat under the `crates/([^/]+/)*ironclaw_wasm/` alternative and re-asserted the crate-name term while saying nothing about the canonical ABI contracts. In a validator whose stated design is "probe derived from reality rather than from a guessed layout", a fabricated filename is a defect on its own terms. Replaced with a `crate_globs` entry, `("ironclaw_wasm", "wit/*.wit")`, which discovers the contracts on disk, requires each in scope, and synthesises the nested WS7 form — so a third contract, or the directory leaving the crate, fails the pin instead of passing on a stale name. Verified non-vacuous: narrowing the workflow alternative to `.../ironclaw_wasm/src/` now reports `tool.wit`, `channel.wit` and the nested probe as out of scope. 2. The embedded-asset routing test substituted `alpha`/`beta` owners so it could reuse the synthetic workspace. That exercised the real prefix strings through the real routing, but left the prefix->owner *pairing* — the table's entire semantic content — asserted nowhere: swapping `ironclaw_extension_support` and `ironclaw_extension_host` passed. Fixed in two halves. The routing test now drives the real `EMBEDDED_ASSET_OWNERS` against a workspace carrying the real owners' names and real manifest paths (the synthetic one could not: `build_plan` rejects a changed package outside the canonical set), asserting the real owner is selected. And the not-stale test now derives the same pairing from the tree instead of restating the constant: it resolves every literal `include_str!`/`include_bytes!` in every workspace crate through `crate_tree`, keeps the targets no crate owns — the ones that actually reach the table — and asserts that every crate compiling one of them is the routed owner or a dependent of it. That surfaced a property worth pinning: `crates/extensions/packages/` is embedded by four crates, not one. `ironclaw_extension_host`, `ironclaw_extension_manager` and `ironclaw_reborn_composition` reach into it alongside `ironclaw_extension_support`, and routing to the support crate covers them only because each depends on it. If that edge goes, a shipped artifact change stops scheduling a crate that embeds it — the silent under-schedule the table exists to prevent. Regression coverage verified red by sabotage, all three wrong tables: owners swapped (7 failures), `packages/` -> `ironclaw_llm` ("embeds nothing from it"), and the hardest case, `packages/` -> `ironclaw_reborn_composition` — a real embedder that the other embedders do not depend on ("...does not depend on..., so routing there never schedules it"). 3. CHECKLIST WS10 claimed each of the nine `wit_bindgen` guest edits forces a committed WASM artifact rebuild. Only six do: `scripts/ci/check-wasm-artifact-freshness.py` scans `crates/extensions/packages/*/wasm-src` alone, `wasm-src-digests.toml` holds exactly six entries, and `git ls-files '*.wasm'` returns exactly those six. The three `test-tools/*/wasm-src/` guests commit no artifact; the tenth site is the host's `bindings.rs`, not a guest. Corrected, and the `wit/` row now states the boundary rather than implying it. Guest paths, `wit/` contents and the six rebuilt artifacts are untouched. Verified: `test_reborn_pr_test_plan.py` 46/46, `test_ws12_workflow_contracts.py` 25/25, `ws12_workflow_contracts.py` green on the real tree, `cargo test -p ironclaw_architecture` 206/206 across 32 binaries. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(host-runtime): state the obligation visibility rule as it holds Review catch (#7090): the guardrail sentence promised "cross-owner access is `pub(super)`, never `pub(crate)`", which is stronger than the code. Verified: `RuntimeSecretInjectionStore::{insert, take, clone_material, discard_for_capability}`, `NetworkObligationPolicyStore::{insert, get, take, discard_for_capability}` and both constructors are `pub(crate)` and must stay so — `src/egress/{mod,host_port,credential}.rs` call them, and that is host-runtime composition outside `obligations/`. The rule is restated as the property that actually holds: a method whose only callers are inside `obligations/` is `pub(super)` (the three that are), and `pub(crate)` is what the stores expose to the egress pipeline they exist to serve. A future agent reading the old sentence would have read the existing `pub(crate)` methods as violations. Guidance-only; no code change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(architecture): put the operator secrets boundary entry on the right rule Review catch (#7096), and it is the serious kind: the `"ironclaw_secrets"` entry landed in `ironclaw_extension_contracts`'s forbidden vector, not `ironclaw_operator`'s. The suite still passed, because `extension_contracts` has no such dependency and `ironclaw_operator` then had no entry at all — so the guard this row exists to add was inert, and a green architecture suite was evidence of nothing. Reintroducing the edge would have passed every check. Moved to `ironclaw_operator`'s vector; `extension_contracts` restored to its `origin/main` content byte-for-byte. Negative-probed rather than assumed. With `ironclaw_secrets` temporarily re-added to `crates/ironclaw_operator/Cargo.toml`: reborn_crate_dependency_boundaries_hold ... FAILED ironclaw_operator must not have a normal dependency on ironclaw_secrets and with the manifest restored, 35/35 pass. Two further review findings, both verified before being accepted: - `ironclaw_extension_manager` **does** have a `boundary_rules()` entry (`:3543-3556`, added with WS2.4). The CHECKLIST residue note and PROPOSAL §8.2's 2026-08-02 amendment both said it had none; §8.2's sentence is stale and is marked superseded. The real gap is narrower and now stated: the rule exists and simply does not forbid `ironclaw_secrets` (#7095). - `ironclaw_product_contracts`'s guide claimed "twenty-four shipped modules". Measured: `src/lib.rs` has 26 shipped (27 `pub mod` less the gated `test_support`), and the table was missing `ironhub` **before** this branch touched it. Count corrected to twenty-six and the missing `ironhub` row added, so the inventory matches `lib.rs`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sandbox): state the Docker-gate claim as the search that checks it Review caught a false inventory in the Known debt entry, and the previous commit is what made it false: "the name appears only in docker_gate.rs and attribution_tests.rs" stopped holding the moment docker_security.rs gained a module doc naming the variable, and CLAUDE.md itself was already a third counterexample. The narrower claim is the one that was always meant and is the one that matters, so it now carries its own reproduction: no workflow, script, env file or manifest mentions the name at all -- `git grep` over *.yml/*.yaml/*.sh/ *.toml/*.py/*.json/.env* is empty here and on main -- and the sole code reference is a read, std::env::var(...) at docker_gate.rs:23. Every other occurrence is a doc comment or a panic message. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(coverage): re-anchor the exemptions the merge shifted tests/integration/changed-coverage-exemptions.toml is exact-line-keyed and auto-merges silently. #7096's additions to ironclaw_reborn_composition moved four entries' subject lines by +2 without anything flagging it; a stranded entry makes the changed-coverage validator abort with no verdict at all. Re-anchored by content (difflib line map from the #7065 tree, which the file was validated against, to the union) rather than by arithmetic: runtime.rs [4068..4073, 4082, 4083] -> [4070..4075, 4084, 4085] runtime.rs [3701] -> [3703] ; runtime.rs [3433] -> [3435] lib.rs [616] -> [618] All 142 entries / 1124 line references re-verified against the merged tree: 0 drift, 0 out-of-bounds, 0 missing paths. * refactor(layers): re-layer processes -> kernel and skills -> substrates (WS3/WS4) Two CHECKLIST rows, both of which were a one-line manifest correction rather than a code move: the family docs already placed both crates where the rows want them and only `Cargo.toml`'s `layer =` disagreed. processes -> kernel (WS3). families/kernel.md already lists ironclaw_processes among the kernel crates. The re-layer makes processes -> resources a kernel -> kernel edge, so its LAYER_MATRIX_EXCEPTION went STALE and the gate said so itself: Stale IronClaw crate layer matrix exceptions: ironclaw_processes -> ironclaw_resources from 2026-07-09 should be removed in W7: runtime process management still depends on resource contracts currently classed with kernel behavior That is the gate's verdict, not a judgement call - deleting the entry is the only way to make it pass. Baseline 5 -> 4, recomputed as len(merged list). Checked the direction both ways: all nine crates that take a normal dependency on processes (capabilities, turns, host_runtime, extension_host, loop_host, extension_manager, runner, reborn_composition, stress) are kernel or above, so the move legalizes an edge without forbidding an existing one. skills -> substrates (WS4 SS3.D). families/domains.md already lists ironclaw_skills under 'Layer(s): substrates'. Its only two normal dependencies are ironclaw_filesystem (substrates) and ironclaw_host_api (contracts), both at or below substrates, and its six consumers are all loops or above. No exception moves in either direction. cargo test -p ironclaw_architecture: 206 passed, 0 failed. cargo check --workspace --all-targets: clean. * docs(target-arch): close the WS3/WS4 rows this work satisfies, with evidence Every tick was verified against the merged tree, never against a PR title. TICKED: - sandbox lane merge: ironclaw_sandbox exists, ironclaw_scripts and ironclaw_process_sandbox absent, bollard/rcgen declared by exactly one manifest in the workspace. - mcp drops the registry dep: ironclaw_extensions is [dev-dependencies] only, 0 production ironclaw_extensions:: refs in src/. - skills -> substrates: landed here. - hooks libSQL/Postgres [decision]: ADR recorded - keep both, with the four rejected alternatives and the evidence they are already converged on one trait plus a shared conformance suite. #6945 read first as the row demands, and explicitly NOT discharged: this PR changes nothing in the dispatch path. - WS3 verify row: the row conflated Wave 3 with Wave 5 work (9 of its 10 exceptions carried removes_in = W7). Corrected with the replaced text quoted, the Wave-3 half satisfied edge by edge, and the Wave-5 remainder named with its owning field value. Ticked on the corrected condition. LEFT OPEN OR PARTIAL, each with measurements rather than a hand-wave: - first_party_tools: 1 of 6 families moved; 15 modules still in host_runtime. Ticking would be false. - processes/capabilities row: re-layer DONE; the capabilities/host.rs split is deferred with every module boundary already computed (4,560 lines, the six workflow ranges, and the arch-exempt waiver that must be deleted with it). - host_runtime binding/catalog-defaults: binding half REFUTED (moving it needs RuntimeLaneExecutor/RuntimeLaneRequest made pub, contradicting the same section's Keeps clause; zero external references to either). Catalog half cannot go to extension_host at all - host_runtime is itself a production consumer at memory_native_extension.rs:96,101, so the move is a kernel -> products edge and a Cargo cycle. Correct destination is downward. - network test_rewrite: NOT executed. Recorded the security shape (production binaries compile the seam and honour the rewrite env var at runtime) and the full 6-step plan, because the env var is how the entire E2E suite redirects vendor traffic through the production binary and the change needs feature forwarding into CI lanes I cannot verify here. cargo test -p ironclaw_architecture: 206 passed, 0 failed. * ci(coverage): recapture the two composed floors from a real measurement The provisional values were arithmetic - the sum of the two slices' recorded deltas - and the dispatch caught them, which is the whole reason the brief demanded a measurement rather than a reconciliation. Dispatch run 30907774036 at 4512e03e28f1df15b419d2e36f9f38f8f55d62fd: 26 success / 1 skipped / 2 failure, judged by per-job tally per #6978. The one skip is the pull_request-gated mutation gate; the two failures are the coverage report and the roll-up it drags down, i.e. this file doing its job. ironclaw_host_runtime: predicted 89.05% (18801 / 21114), MEASURED 88.63% (17562 / 19814). The composition was wrong by 1300 denominator lines because both slices measured their delta under the pre-#7083 aggregator, which could not see crates/extensions/** at all - lines leaving host_runtime for extension_support vanished from the tree it could measure, so neither branch's recorded delta describes the post-#7094 world. ironclaw_extension_support: MEASURED 75.31% (7142 / 9484) against #7094's 82.64% (6826 / 8260), captured before #7080's executor lines arrived. floor_percent FALLS 7.33pp and that is flagged in the file for an owner's eye rather than written quietly. Evidence it is composition and not lost tests: floor_covered_lines RISES 6826 -> 7142, so the crate is protected by more absolute lines than before, and #7080's un-masking accounting was 1398 -> 1398 with zero test names lost. Same shape as #7094's own ironclaw_runner recapture. ironclaw_sandbox passed unchanged at its arrival capture (87.09%, 3185 / 3657). The [global] entry is untouched: both moves are crate-to-crate inside the set the fixed aggregator sees. * fix(network): compile the test rewrite seam out of production builds (WS3) Closes the WS3 network row. Also RETRACTS an overstatement I made in this row's earlier annotation. CORRECTION FIRST. The earlier note claimed production binaries compile the seam and honour IRONCLAW_REBORN_TEST_HTTP_REWRITE_MAP at runtime, so anyone able to set it could redirect all credentialed vendor egress. That was WRONG. RewriteNetworkTransport::from_env_value already returned UnavailableInRelease when !cfg!(debug_assertions) (test_rewrite.rs:150), and neither [profile.release] nor [profile.dist] sets debug-assertions, so a shipped binary with the variable set REFUSES TO BOOT. It was fail-closed before this PR. I had read the ungated `mod test_rewrite;` declaration as an ungated runtime path. What was genuinely wrong, and is fixed: 1. The guard was a RUNTIME check keyed on cfg!(debug_assertions) - a profile proxy, not a build-kind guarantee. A release profile with debug-assertions turned on (normal when chasing a production bug) silently re-arms it. 2. The refusal arm had NO TEST. The one guard between a shipped binary and redirectable vendor egress was unpinned. Fix: compile-time exclusion instead of a runtime check. mod test_rewrite and its four re-exports are now cfg(any(debug_assertions, feature=test-support)), and default_host_http_egress is a compile-time pair - production builds PolicyNetworkHttpEgress<ReqwestNetworkTransport> directly, with the rewrite wrapper absent from the binary. The runtime check stays as defence in depth. E2E needs no change: those harnesses build DEBUG binaries, so they satisfy debug_assertions and keep redirecting with no feature flag and no workflow edit. The feature-forwarding-into-CI risk I flagged earlier does not arise. test-support is still forwarded composition -> network for a release-PROFILE build that needs the seam. Both halves proven rather than assumed: (a) release refuses - new regression test a_set_rewrite_map_activates_only_in_debug_and_is_refused_in_release feeds a well-formed map and asserts on profile. Under 'cargo test --release -p ironclaw_network --features test-support' it passes on the UnavailableInRelease branch; under debug 'cargo test -p ironclaw_network' it passes on the active branch. 56 passed, 0 failed. (b) production compiles without the seam - 'cargo check --release -p ironclaw_reborn_composition' (no test-support) is clean, which only compiles if the cfg(not(..)) arm is right. Also: WS0_EXTENSION_SPECIFICITY_ALLOWLIST_BASELINE 129 -> 127. The constant had drifted ABOVE the real list length; the ratchet is shrink-only so it passed silently while buying back two unearned slots. Measured off the compiler (set baseline to 0, read the reported length), identical on main and on every slice, so pre-existing drift rather than something this PR caused. cargo test -p ironclaw_architecture: 206 passed, 0 failed. cargo check --workspace --all-targets: clean. * docs(coverage): verify the extension_support floor drop is composition, independently The 82.64 -> 75.31 recapture carried a rationale that was recorded but explicitly NOT verified. Re-derived it from scratch between the two capture refs (f946a93fae -> 939af4847d) rather than inheriting the claim: - 0 test names lost in the crate (158 -> 160 test fns; both new names belong to the arriving executor). - 0 test names lost WORKSPACE-WIDE (13836 -> 13843 test fns, 13752 -> 13759 unique). This is the check that separates a relocation from a deletion: host_runtime's roster drops 156 names over the same range and every one reappears in another crate. - Exactly four files arrived, 1367 source lines, all of them the family-1 skill-install executor (src/skills/url_install.rs + url_install/{github, zip_bundle,bundle}.rs). No pre-existing file left the crate. - The arithmetic closes with the pre-existing numerator held CONSTANT: (6826+316)/(8260+1224) = 75.31% exactly, so the pre-existing code lost zero covered lines. The arriving block's own coverage is 316/1224 = 25.82%. Composition, confirmed rather than assumed. No test regression to fix; the 25.82% arrival is what earns the follow-up already recorded above the entry. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(host_runtime): collapse a duplicated obligation predicate and quiet a background warn! Three verified review findings from the #7141 round. Each was confirmed against the code before being acted on; nothing was changed on assertion alone. 1. obligations/handler.rs — `obligation_supported_before_dispatch` and `obligation_supported_after_dispatch` had BYTE-IDENTICAL 19-line bodies (verified by exact line-by-line comparison). Both were private, each called exactly once, both taking the same `phase` argument. The two names asserted a pre/post-dispatch distinction the code never implemented, while the pair gates admission of RedactOutput, EnforceOutputLimit and EnforceResourceCeiling — so editing one copy alone would have left the other stage accepting an obligation the host cannot honour (a fail-open). Collapsed to one `obligation_supported`, with the reasoning recorded so the pair is not reintroduced. 2. obligations/process_store.rs — `cleanup_terminal` is reached from `observe_process_commit` (an async background journal callback, call sites at :363/:379/:394), so its `tracing::warn!` violates the repo rule that background tasks never use info!/warn! — they corrupt the REPL/TUI display. Lowered to `debug!`; the error is still returned to the caller on the next line, so nothing is swallowed. 3. reborn_restructure_baselines.rs — the doc table said the LAYER_MATRIX_EXCEPTIONS count was "now 11". Recomputed on this ref by anchoring on the `= &[` of the value (the `&[LayerMatrixException]` type annotation opens a bracket on the same line and silently yields 0): the real count is 4, matching WS0_LAYER_MATRIX_EXCEPTION_BASELINE = 4. Corrected. Verification: cargo check --all-targets -p ironclaw_host_runtime exit 0; obligation tests 13+26 passed, 0 failed; reborn_restructure_baselines 1 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): a shipped package prompt is an asset, not prose — it was selecting no lane Review finding on #7141, confirmed empirically before acting. The Markdown prose carve-out in the planner ran BEFORE the `EMBEDDED_ASSET_OWNERS` lookup. A prompt is a `.md` file that no package *directory* owns, so a change to `crates/extensions/packages/*/prompts/**.md` took the prose arm and planned: mode=none crate_buckets=[] "crate-tree guidance changed: ..." while its sibling `manifest.toml` in the same package planned `mode=selected` onto ironclaw_extension_support + ironclaw_extension_host. Prompts are shipped production output that `ironclaw_extension_support` compiles in, and the comment above `EMBEDDED_ASSET_OWNERS` names "manifests, prompts, schemas and built wasm/*.wasm" as exactly what that table owns — so this was the "silent under-schedule of a change to production output" that comment forbids. 145 of the 149 `.md` files under `packages/` are prompts. The rule is keyed on the `prompts/` path segment, not on the asset prefixes. That distinction is load-bearing: the first attempt yielded to the asset prefixes wholesale and broke `test-tools/README.md`, which is documentation of the fixture bundles and is deliberately pinned as prose. Of the four asset kinds the table owns, only a prompt is Markdown (manifests are .toml, schemas .json, wasm .wasm), so `.md` asset <=> prompt is exact. Sabotage-tested in both directions: * `_is_package_prompt` -> False (reinstates the bug): RED, "AssertionError: 'none' != 'selected'". * `_is_package_prompt` -> any .md under an asset prefix (over-broad): RED on both the new test and the pre-existing `test_markdown_owned_by_no_crate_is_prose`, at `test-tools/README.md`. * restored: 52 passed, 51 subtests, green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(harness): refresh the latency-runner lockfile after the sandbox consolidation Review finding on #7141, reproduced before fixing. The latency harness keeps its own committed `Cargo.lock`, separate from the workspace lockfile, and the crate consolidation that replaced `ironclaw_scripts` + `ironclaw_process_sandbox` with `ironclaw_sandbox` never regenerated it. It still carried entries for both removed packages (lines 3244 and 3602) and the old host-runtime/loop-host dependency graphs. Reproduced exactly as reported: $ cargo metadata --locked --manifest-path harness/latency/runner/Cargo.toml error: cannot update the lock file ... because --locked was passed exit 101 so any reproducible invocation of the harness was broken, while the documented unlocked command silently rewrote the lockfile as a side effect of running. Regenerated with `cargo update --workspace`, which re-resolves the path dependencies. Verified after: `--locked` exits 0, the two removed packages are gone (0 entries), and `ironclaw_sandbox` is present (1 entry). Note: the re-resolve also carried three registry deps forward (wasmtime-wasi 46.0.1 -> 47.0.3, wasmtime-wasi-io likewise, wit-parser 0.251.0 -> 0.252.0). That is contained — this lockfile governs only the standalone benchmark harness and is not the workspace lockfile, and it was already unusable under `--locked` before this change. Co-Authored-By: Claude Opus 5…
l3ocifer
pushed a commit
to l3ocifer/frick-ironclaw
that referenced
this pull request
Sep 3, 2026
…isions (accumulating the fleet) (nearai#7181) * refactor(contracts): move extension runtime descriptors to a neutral contract (WS3) Deletes the two `-> ironclaw_extensions` layer-matrix exceptions (`ironclaw_mcp`, `ironclaw_scripts`) by giving the runtimes-layer lanes a contracts home for the descriptors they read, instead of the registry crate they may not depend on. Exceptions 13 -> 11; baseline lowered in the same change. Moved to `ironclaw_extension_contracts`: - `runtime::{ExtensionRuntime, ExtensionAssetPath, ExtensionAssetPathError}` - `hosted_mcp::{HostedMcpDiscoveredTool, HostedMcpDiscoveredToolAnnotations}` `ExtensionPackage`/`ExtensionManifest` deliberately stay in `ironclaw_extensions`: they carry the whole parsed manifest tree and a `PackageRootBinding` typed on `ironclaw_filesystem::VirtualPath`, which the §11.2.3 contracts-purity allowlist (`{ironclaw_host_api}` only) forbids the contracts crate from naming. Measured instead: both lanes read exactly three things off the package — `id`, `capabilities`, `manifest.runtime` — so the lane request structs now take those three and the caller (which owns the package) projects them. Also repointed `ResourceReceipt` to its real owner: `ironclaw_resources` only re-exports `ironclaw_host_api::resource::ResourceReceipt`, so the lanes' import was a §11.2.4 two-import-paths hop, not a dependency. No `pub use` shims (§11.3): every consumer is repointed in this change, and `resolve_under` becomes the free function `ironclaw_extensions::resolve_asset_under` because the orphan rule forbids an inherent impl on the moved type. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(sandbox): merge the sandbox lane into one crate (WS3) Creates `ironclaw_sandbox` (runtimes) from the three halves of "run an already-authorized command away from the host", and deletes the two crates PROPOSAL §6.6.4 marks for merge: - `ironclaw_process_sandbox` (plan contract) -> `src/plan.rs`, `src/validation.rs` - `ironclaw_host_runtime::sandbox_process` -> `src/sandbox_process/**` - `ironclaw_scripts` (script lane + Docker path) -> `src/script.rs` The kernel sheds the Docker/CA cone: `bollard`, `rcgen`, `x509-parser` and `time` are gone from `ironclaw_host_runtime`'s manifest, and `bollard`/`rcgen` are now declared by exactly one crate in the workspace. Two migration details PROPOSAL §6.6.4 and CHECKLIST WS10 call load-bearing: - `PROCESS_SANDBOX_CAPABILITY_ID` -> `ironclaw_host_api::capability`, so `ironclaw_loop_host` drops its lane dependency (production dep gone; a dev-dep remains for the tests that build plans). - `SandboxCommandTransport` -> `ironclaw_host_api::process`, with the shapes it names (`CommandExecutionRequest`/`Output`, `RuntimeProcessError`, `SavedCommandOutput`, `SavedCommandOutputSanitization`). Without this the runtimes-layer lane could not implement what the kernel consumes. Enumerating gates were repointed, never relaxed: the specificity carve-outs and the struct/test-support ratchet entries moved with their files (both baselines unchanged at 129 and their prior values), the panic-gate baseline row moved, `reborn-crate-test-buckets.sh` registers the new crate, and the three `reborn-e2e-rust.sh` script selectors follow the tests (plus `docker_security`, which had no selector before). One gate would have gone silently vacuous and was fixed rather than moved: the script-lane surface scan in `reborn_dependency_boundaries.rs` read a hardcoded `src/lib.rs`, which after the merge no longer holds the lane. It now scans the whole crate source tree with a fatal-read walk and a non-vacuity assertion. One deletion, recorded: `RebornScopedSandboxCommandTransport::into_process_port` returned a kernel type a runtimes crate may not name. It had zero callers workspace-wide; the kernel wraps the transport, which is the direction the port inversion requires. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): record the WS3 corrections with their evidence Three dated amendments, each quoting the text it replaces: 1. CHECKLIST WS3 sandbox row + PROPOSAL §6.6.4 — "all pieces currently unwired/test-only" is REFUTED. Three production paths cross the merged crate (spawn-path plan validation, the process_executor routing check, and the saved-command-output scope digest). The accurate claim is narrower: no production *execution backend*. Behavior preservation is therefore argued at the diff (11 of 26 moved files byte-identical, 9 more differing by one import line, +63/-36 overall), not inferred from deadness. 2. CHECKLIST WS3 mcp row + PROPOSAL §6.6.3 — the prior wave's "structurally blocked" finding is half right, and the wrong half is load-bearing: only `ExtensionPackage` is un-absorbable, and no lane ever needed it (both read `id`, `capabilities`, `manifest.runtime` and nothing else). The registry half of the flip is done; the `resources` half is refuted as phrased — the estimate/usage vocabulary the row asks about is already in `host_api::resource` and already imported from there, while the real blocker is the `ResourceGovernor` authority port and `ResourceError`'s denial cone. 3. Recorded as a structural finding, not a note: the sandbox row and the mcp row are ONE problem. `ironclaw_scripts` imports the identical DTO set, so the merge alone deletes zero exceptions and only the mcp carve-out lets either lane shed the registry edge. Also reconciled: PROPOSAL §6.1.2's as-built inventory gains the two modules WS3 landed (and states why `ExtensionPackage` stayed); §2's package count 66 -> 65; the §9 disposition rows for `ironclaw_scripts`/`ironclaw_process_sandbox`/ `ironclaw_mcp`; the §11.2.2 ratchet rows (13 -> 11); the WS3 verify row; the stale WS1.3 sentence asserting the blocker as settled fact; and `reborn_restructure_baselines.rs`'s doc table, which still read 15. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(sandbox): drop imports the merge left unused `process_port.rs` no longer names `MountView` or `thiserror::Error` (both went to `host_api::process` with the types that used them), and `sandbox_process.rs` no longer needs `sync::Arc` after `into_process_port` was deleted. Found by per-crate `clippy --all-targets --all-features -D warnings`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): let the Reborn PR planner plan guidance edits and crate deletions Three fail-closed gaps in `reborn_pr_test_plan.py`, all hit by this PR and all live on `main` today — any PR with the same change shape is unplannable. 1. `.claude/**` was unclassified, so the planner refused outright. It is agent guidance in exactly the sense `docs/**` is human guidance: no Rust test reads either as data (the only in-tree references are prose citations in test doc comments). Added to `IGNORED_PREFIXES`. Without this, "guidance travels with the change" — the restructure's own discipline — cannot be satisfied in a single PR. 2. `crates/AGENTS.md`, `crates/README.md`, `crates/Architecture.md` raised "unmapped crate path": they sit under `crates/` but belong to no package. Now classified as crate-tree prose, matched by "Markdown no package directory owns" so a genuinely unmapped crate path is unaffected. 3. An unmapped crate path used to raise. `git diff` reports a deleted crate's old paths and CI feeds the planner that diff, so **every crate deletion or rename was unplannable** — including the six deletions PROPOSAL §2 plans. It now widens to the exhaustive plan. This is a semantic change and it is the safe direction: the full plan is a superset of any narrowing, so an unattributable path can never cause under-selection, whereas refusing to plan blocks the PR instead of protecting it. Malformed input is still rejected by the unclassified-path branch. Each lands with fixtures per WS10's rule, positive and negative: guidance paths select nothing while non-guidance paths still fail closed; crate-tree prose selects nothing while crate *code* under the same unmapped directory widens to `full` (so the Markdown carve-out cannot swallow code). The pre-existing `test_unmapped_crate_path_fails_fast` is renamed and rewritten to pin the new contract rather than deleted. Verified against this PR's real 130-path diff: the planner returns `mode: full`, and the workflow's own exhaustiveness guard passes on that output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(arch): give the retained resource exceptions an owning issue, not a wave Review (#7065) caught that both surviving `-> ironclaw_resources` exceptions declared `removes_in = "WS3"` — the wave this PR *is*, which does not remove them. That is precisely the defect §11.2.2 already records against `conversations -> turns` ("`removes_in = "WS5"` and WS5 has partly shipped without it falling"), and it would have been repeated here. Both now point at issue #7067, which owns the design work that actually clears them: replacing the `ResourceGovernor` dependency with a narrow reserve/reconcile/release port. The issue carries the measurements — 3 of 10 methods used, zero implementors, and the `ResourceError` denial cone — plus the two open questions (error shape, port home) that make it a design slice rather than a move. An owning issue is also what §11.2.2 asks for and what the ratchet still cannot enforce (there is no `owning_issue` field yet), so this is the strongest form currently expressible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(contracts): pin the asset-path validator that moved into extension_contracts `validate_asset_path` moved here with `ExtensionAssetPath`, the type it constructs. In `ironclaw_extensions` it was only ever reached indirectly through manifest parsing, so its six rejection branches had no direct test — and a contracts crate that carries validation owes that validation one. Two tests: every reject branch with its exact reason and `Display` output (empty, NUL/control, URL, absolute, Windows drive and backslash, and the empty/`.`/`..` segment cases) plus the manifest-relative shapes that must keep being accepted; and `ExtensionRuntime::kind()` over all five variants, since that projection is what every lane uses to reject a runtime it does not serve. Also removes a changed-line coverage risk this PR would otherwise carry into the merge queue: the gate does not run on ordinary PRs (#7036), so ~100 newly-added lines of validator would first be measured where a failure is expensive to diagnose. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(coverage): re-capture the host_runtime floor and floor the new sandbox lane `RATCHET FAIL: ironclaw_host_runtime` — observed 18854 covered vs a `floor_covered_lines` of 20538. This is the shrinkage case the ratchet's own "To fix" text describes, not a coverage regression: `sandbox_process/**` moved to `ironclaw_sandbox`, so the crate's denominator fell 23277 -> 21267 (-2010 instrumented lines) and its covered lines fell with it. The percentage floor is **raised, not lowered**: observed 88.65% against an old floor of 88.23%, so the entry now reads 88.65. Only the absolute line count moves down, and it must — those lines are no longer in this crate. To keep that from being a net loss of protection, `ironclaw_sandbox` is floored on arrival at its observed 87.09% (3185 / 3657). This is a net *increase* in ratchet coverage: neither `ironclaw_scripts` nor `ironclaw_process_sandbox` was ever floored, and the `sandbox_process` half was protected only as part of host_runtime's line count, which this PR necessarily reduces. Floored crates 16 -> 17. Verified by replaying the ratchet arithmetic against CI's observed numbers: both crates pass on percentage and on covered lines. Numbers taken from the failing run's own report (job 91740733521), which is the authority for this gate. The `Tests (Reborn)` roll-up failed solely on this sub-job ("coverage-report result 'failure' did not match planned=true"); no other lane failed — 50 pass, 2 fail, both this root cause and its roll-up. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): record the coverage ratchet as a move-sensitive gate WS3 hit a gate no move row had named. `tests/integration/coverage-floor.toml` is keyed on crate identity plus absolute covered-line counts, so it is invisible to WS10's path-keyed gate audit and yet it fails on every crate move, merge, rename, or family `git mv` that shifts instrumented lines between crates — as it did here, while the percentage floor was *improving*. Recorded on WS10 with the three rules WS7 will need: re-capture in the same PR, raise the percentage floor rather than leaving it, and floor the destination crate or the move silently drops that code out of the ratchet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(extension-manager): repoint ironhub onto the moved ExtensionAssetPath A semantic conflict the merge could not see: #6780 landed `ironhub/{package,catalog}.rs` importing `ExtensionAssetPath` from `ironclaw_extensions`, while this branch moved that type to `ironclaw_extension_contracts::runtime`. Different files, so git auto-merged cleanly and the breakage surfaced only at `cargo check`. Repointed both sites to the contracts crate (no shim, per §11.3). The manifest already named `ironclaw_extension_contracts`, so this is imports only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(coverage): exempt the WS3 move's no-region lines and record the gate The changed-lines coverage gate went red on four files while changed-line coverage was 95.35% against a 90% floor: the failure was its two fail-closed STRUCTURAL assertions, not any percentage. Every line below was derived by replaying scripts/ci/reborn_changed_coverage.py against this PR's own merged lcov (run 30831658659) with the base lcov the gate itself resolved (run 30828540055 @ b89fcd3575), until the replay reproduced the CI verdict byte-identically. Line numbers come from the gate's own `candidate_lines - mechanically_uninstrumentable_lines()`, not from the log. - host_api/src/process.rs (31 lines): new placement-neutral process vocabulary with no function body anywhere in the file; rustc emits no LCOV record for it at all. Same shape already exempted for product_contracts/loop_contracts. - extension_contracts/src/hosted_mcp.rs (12): field declarations of the two new tools/list descriptor structs. The file is plainly instrumented (191 DA, 164 hit), so this is a no-region artifact, not an instrumentation gap. - host_runtime/src/services/runtime_adapters.rs (13): continuation lines of three rewritten calls, all PROVEN EXECUTING by their region-start heads (lines 380/434/977 score 24/16/63 hits). The four genuinely-uncovered lines in the same rewrite are deliberately NOT exempted -- the gate already subtracts them as pre-existing debt inherited from base. - composition capability_host_tests/approval_gates.rs (6): type positions in a test double whose body region scores 1 hit. The last one is a finding, not just a waiver: that file is 100% test code behind `#[cfg(test)] mod capability_host_tests;`, but the gate's test_only_path() recognises /tests/, /test_support/, */tests.rs and *_tests.rs and NOT a cfg(test) module DIRECTORY, so it measures it as production. It is the only such directory in crates/ today. Docs (target-architecture, same PR per the docs-truth rule): - CHECKLIST WS10 gains the changed-lines gate beside the ratchet row, cross- referencing the WS2.1 note rather than restating it: percentages are not what fail a move; derive lines by byte-identical replay (--fetch-base-coverage silently degrades without --github-repo); and a stranded exemption path is an ABORT with no verdict, not a loud failure. - CHECKLIST WS10 exception-ratchet row: the constant was cited at :4063 and sits at :4164 -- corrected by removing the line pin, since the file is edited every wave. Records that the baseline is a UNION across parallel WS3 lanes. - families/contracts.md: records extension_contracts' new ownership of the runtime descriptor vocabulary -- the carve-out that let BOTH lanes drop the registry edge -- and the orphan-rule seam that keeps resolve_asset_under in the registry crate. - families/lanes.md: two "Never" claims were reading as satisfied when they are not. ironclaw_mcp's "never depends on the resource-governor crate directly" is refuted (the compiled edge survives; #7067 tracks the narrow port), and ironclaw_sandbox's "no direct process spawning outside the transport seam" is aspirational -- script.rs:454 still builds Command::new("docker"). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sandbox,mcp): correct the wiring inventory and record the projection cost Two review findings verified against the tree; three refuted with evidence in the PR threads. Valid — the sandbox wiring inventory was self-contradictory. `CLAUDE.md` said "Two production call paths ... and both are plan validation" directly above a list of THREE bullets, and `lib.rs` omitted the third entirely. The third is real and is not validation: `host_runtime/src/process_output.rs:482` derives the scoped saved-output directory through `RebornSandboxScopeKey::from_scope`. That inventory is what tells a future agent which paths are live, so an undercount invites deleting a production path as dead code. Both surfaces now say three and no longer claim they are all plan validation (the `loop_host` capability-id comparison never was either). Valid, and recorded rather than redesigned — the registry carve-out cost a type-level invariant. Replacing `package: &ExtensionPackage` with independent `extension` / `capabilities` / `runtime` borrows is what deleted the `mcp -> extensions` and `scripts -> extensions` exceptions, but it also means the type no longer guarantees the three came from one package. `execute_extension_json` re-checks the descriptor half (`descriptor.provider == extension`); the runtime half cannot be re-derived, because nothing in an `&ExtensionRuntime` names its owning extension. No caller can trip it today -- there is exactly one production caller (`runtime_adapters`) and it projects all three from one package in one expression -- so this is a latent structural weakening, not a live defect. Restoring the compile-time binding needs a sealed projection minted by the package owner; a check inside the lane cannot express it, and re-taking the registry edge would undo the carve-out. Both request types now carry the caller obligation in their field docs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(extensions): move the skill-install executor to extension_support (WS3) WS3's first-party-tools row, family 1 of 6: skill management / URL install. `skill_url_install.rs` and its `bundle`/`github`/`zip_bundle` submodules, plus the install-input normalizer, move out of `ironclaw_host_runtime::first_party_tools` into `ironclaw_extension_support::skills::{url_install, resolve_install_input}`, where the skill executor half already lived. Move-only: no behavior change, no test edited for content. `ironclaw_host_runtime -> ironclaw_skills` is deleted from LAYER_MATRIX_EXCEPTIONS — the edge is gone, not waived (exceptions 13 -> 12, WS0_LAYER_MATRIX_EXCEPTION_BASELINE drops with it). `ironclaw_skills` and `zip` survive as dev-dependencies for host_runtime's own tests; dev edges are outside the matrix by construction. Two doc ambiguities are resolved in the same diff, as dated PROPOSAL amendments quoting the text they replace: - §6.8.4's "the builtin first-party tool handlers absorbed from host_runtime/first_party_tools" contradicted §8.2's "kernel: ✗ (ports only)" row and the enforced BoundaryRule. Resolution: the seam splits executor from adapter — the executor moves behind a neutral request/error pair, the FirstPartyCapabilityHandler / CapabilityManifest / registry wiring stay host-side. Same shape the groupware and web-access tools already ship. - §8.2's "ports only" cell now says what it means: contracts-layer ports the kernel also consumes, not permission to name a kernel trait. Two cost corrections recorded for the remaining families: `host_runtime -> extension_support` is not divisible family-by-family (mod.rs holds it via `extension_support::coding`), and `host_runtime -> ironclaw_extensions` is not reachable by this row at all. PATH_TERM_COLLISIONS shrinks by two: the installer's github carve-outs now sit inside a scan-exempt crate. Test accounting (un-masking discipline), unfiltered `--list` over both crates: 1398 -> 1398, with exactly two tests renamed by module path and none lost. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sandbox): record that the Docker fail-closed switch is wired to nothing Review asked why the migrated docker_security test can pass with no daemon. The skip is pre-existing (the file differs from its pre-merge original by one import line); WS3 only enrolled it in the required Rust e2e lane, where it was not run at all before. The real defect the question surfaced is worse and also pre-existing: this crate's tests/support/docker_gate.rs states that IRONCLAW_REQUIRE_DOCKER_TESTS=1 makes a missing daemon a hard failure and that "CI sets this" -- and nothing sets it. Repo-wide the name occurs only in docker_gate.rs and attribution_tests.rs, here and on main. So every real-Docker test in the crate skips-and-passes everywhere, which is exactly the gap the gate's own comment says let sandbox security bugs ship unnoticed. docker_security.rs additionally open-codes its own check rather than using the gate, so it would stay fail-open even once something did set the variable. Recorded rather than fixed: setting the variable is a CI-behavior change that would hard-fail any lane without a daemon or the ironclaw-worker image, which is not verifiable from inside a move PR whose evidence claim is behavior preservation. Filed as the #6945 guardrail-claim-vs-reality class with the two-part fix stated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(host_runtime): record the executor/adapter seam in crate guidance The crate's CLAUDE.md said "first-party runtime tools belong under `first_party_tools/`" without saying that only the host half does. WS3 moves each tool's executor into `ironclaw_extension_support`, which may not name this crate, so the rule now names both halves and points at the skill-install family as the worked example. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(host_runtime): keep the install-input error path log-free The moved executor returns `SkillManagementCapabilityError`, and routing it through `skill_management_error` would have added a `debug!` line to a path that had none before the move. A move-only change must not add one, so the install-input arm maps the kind directly and the `dispatch` arm keeps the record it already had. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci(coverage): re-capture the host_runtime floor for the WS3 executor move The ratchet does not run on `pull_request` (`reborn_pr_test_plan.py:21`; issue #7036), so this PR's green checks were not evidence on this axis. A full-plan `workflow_dispatch` run on this exact head reported: RATCHET FAIL: ironclaw_host_runtime observed: 88.59% (20485 / 23124 lines) floor: 88.23% ... floor_covered_lines: 20538 (effective floor 20518) The percentage went UP while `floor_covered_lines` went DOWN — shedding well-covered code lowers the absolute numerator, which is a separate assertion from the percentage one. Re-captured to the observed numbers (floor raised 88.23 -> 88.59, not merely held). Verified locally against that run's own merged lcov artifact: ENFORCING mode, 17 PASS / 0 FAIL, exit 0. run: https://github.com/nearai/ironclaw/actions/runs/30858257594 head: e07b3b0299b0add11117e9591da71d46d7a7c832 The destination crate is deliberately not floored, because it cannot be: every crate under `crates/extensions/` is invisible to the coverage tooling — `reborn_coverage_lcov.py:19`'s CRATE_RE still requires a crate directory directly under `crates/`, which #7037's colocation broke. Filed as #7083 with the measurement; the global floor is left alone rather than re-captured onto that hole. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(wasm): move wit/ inside its owning crate (Wave 3) CHECKLIST WS4 + WS10 `wit/` rows. `wit/{tool,channel}.wit` moves from the repo root to `crates/ironclaw_wasm/wit/` — the crate that owns the ABI — per PROPOSAL §6.6.1. Behavior-free: same bytes, same generated bindings. Wave-3 coordinates: the docs write the destination as `crates/lanes/ironclaw_wasm/wit/`, but `crates/lanes/` does not exist until WS7. Because the files now sit *inside* the crate, the WS7 family move carries them with no further path edit anywhere — which is the whole point of putting them there. Ten wit-bindgen `path:` args repointed (the host plus nine guests: six under `crates/extensions/packages/*/wasm-src/`, three under `test-tools/*/wasm-src/` — the CHECKLIST row said six). All nine guests verified building against the moved WIT on wasm32-wasip2. The four `include_str!` readers of the ABI text do NOT get repointed literals. Doing that would turn the two `ironclaw_host_runtime` sites from repo-root reach-ins into *cross-crate* ones — §11.2.7's strict class, the one WS2 turns into hard failures — taking the scan from 19 to 21 while ticking a box that says "§11.2.7 scan passes". Instead the ABI text gets one owner, `ironclaw_wasm::TOOL_WIT` (`src/config.rs`, beside `WIT_TOOL_VERSION`), and all four sites read the const over cargo edges that already exist. Measured with the scan: 133 -> 129 escaping sites, cross-crate 19 -> 19, zero `wit/` entries remaining. Path-keyed gates repointed: `scripts/check-version-bumps.sh` (both ABI paths), `.githooks/pre-commit`, and `platform-and-compat.yml`'s `has_direct_wasm_abi_risk` filter — where the bare `wit/` alternative is *deleted* rather than rewritten, because the filter's existing `crates/([^/]+/)*ironclaw_wasm/` alternative already matches both the Wave-3 and the WS7 location. `scripts/ci/ws12_workflow_contracts.py` anchored on that deleted string, so its anchor moves to `build-wasm-extensions` and its in-scope probe now pins both locations. `Dockerfile` loses two `COPY wit/ wit/` lines in the planner and builder stages: both already run `COPY crates/ crates/`, so the files arrive with the crate and the old line would COPY a path that no longer exists. Docs: the WS4 row's `crates/lanes/wit/` destination was the only doc site placing the directory beside the crate rather than inside it; corrected there and in README's tree, with dated amendments in CHECKLIST, PROPOSAL §6.6.1 and PLAN Wave 3 recording what the move found. Test accounting (unfiltered `--list`, name-by-name, quiescent tree): ironclaw_wasm 51 -> 51, ironclaw_host_runtime 1246 -> 1246, ironclaw_architecture 198 -> 198. Zero diff, no test edited for content. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * build(wasm): rebuild first-party artifacts for the moved wit/ path Forced by the previous commit, not incidental to it. `scripts/ci/check-wasm-artifact-freshness.py` keys each package's committed `wasm/<name>.wasm` to a digest of the `wasm-src/` tree that produced it, so editing a guest's `wit_bindgen::generate!` `path:` — which the `wit/` move requires in all six shipped guests — invalidates the recorded digest and fails the gate. The gate's own contract forbids the shortcut: "Re-record only after `./scripts/build-wasm-extensions.sh --first-party` and committing the rebuilt artifact — the digest asserts a claim about the artifact, and updating it without rebuilding launders a stale one." So the artifacts are genuinely rebuilt (`--first-party`, exit 0, 6 OK / 2 host-native SKIP), not re-recorded in place. Byte sizes move by more than the source change accounts for because these builds are not reproducible by design — the guests pin no toolchain and resolve their own `Cargo.lock` at build time, which is the documented reason the gate hashes sources rather than artifact bytes. Verified: `check-wasm-artifact-freshness.py` OK (6 packages), and `cargo test -p ironclaw_extension_support` green (102/46/4) — that crate `include_bytes!`s these artifacts, so it exercises the rebuilt components. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-arch): record the WS7 artifact-rebuild cost of guest path edits The `wit/` move had to rebuild six shipped WASM binaries because `check-wasm-artifact-freshness.py` digests each guest's whole `wasm-src/` tree. WS7 hits the same wall from the other direction: the six package guests reach the ABI across two trees, so moving either `ironclaw_wasm` or `extensions/packages` rewrites all six `path:` literals and forces the same rebuild. Recorded on CHECKLIST WS10's `wit/` row (point 6), on the loud-path-pattern row that owns the WS7 repoint (also corrected six -> nine guests there), and on PLAN's Wave 5 block with the cheap mitigation: move the two crates in one PR and pay it once. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci(planner): classify the path classes that blocked the wit/ move `Detect Reborn test scope` exits 1 on any pull request whose diff holds a path `reborn_pr_test_plan.py` has no rule for, which made this PR unmergeable: it must edit `Dockerfile` (the moved directory's `COPY wit/ wit/` no longer resolves) and `scripts/check-version-bumps.sh` (the ABI gate would otherwise grep dead paths and silently stop enforcing). 18 of its 46 paths were unclassified. Same class as the `.claude/` gap #7064 fixed, and classified the same way — one rule per class, recorded beside the constant: * `Dockerfile` / `.dockerignore` — `platform-and-compat.yml` keys `has_docker_risk` off exactly this pair and owns the image build. * `.githooks/**` — Code Style triggers on the tree and lints its contents (`test-ci-comm-locale-pin.sh`); no Reborn lane runs a hook. * `scripts/{build-wasm-extensions,check-version-bumps}.sh` — `platform-and-compat.yml`'s `has_direct_wasm_abi_risk` classifier both scopes and runs them. * markdown owned by no crate (`crates/AGENTS.md`, `test-tools/README.md`) — prose, like `docs/` and `.claude/`. A crate-resident doc still selects its own crate's lane. The first-party extension package assets are deliberately NOT ignored. `crates/extensions/packages/*/wasm/*.wasm` is a shipped artifact that `ironclaw_extension_support` embeds with `include_bytes!`, and `test-tools/*/manifest.toml` is `include_str!`d by `ironclaw_extension_host`. Calling either prose would convert today's loud failure into a silent under-schedule of a change to production output — the WS10 failure mode. `EMBEDDED_ASSET_OWNERS` routes each tree to the crate that compiles it instead, so this PR now additionally schedules `ironclaw_extension_{support,host,manager}`: the crates that consume the six rebuilt WASM artifacts. Also fixes #7085 in a file this PR already touches. The WIT version extractors used the GNU-only BRE `\+`, so on BSD sed (macOS) they matched nothing, and because the `WIT_TOOL_VERSION` cross-check is guarded on a non-empty version the hook printed "All version checks passed" having compared nothing. `[[:space:]][[:space:]]*` is identical under GNU sed, so the enforced Linux CI lane is unchanged; verified on BSD sed that both `wit/tool.wit` (0.3.0) and `wit/channel.wit` (0.3.1) now extract. Regression tests: every classified class gets a case in `test_reborn_pr_test_plan.py`, including the paired assertion that the embedded assets *select a lane* rather than merely being accepted (the inverse of the `.claude/` prose test), and a staleness pin that fails if an asset tree or its owning crate moves. All ten new cases fail against the planner on `main`. `test_unclassified_build_input_fails_fast` moves off `Dockerfile` onto a still-undecided input so the fail-closed arm stays exercised. Refs #7087, #7085 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(host-runtime): split obligations into its three chartered owners (WS3) `crates/ironclaw_host_runtime/src/obligations.rs` was 3,122 lines fusing the three owners PROPOSAL §6.5.9 charters separately, held apart only by an `// arch-exempt: large_file` waiver. It is now one module per owner: - `obligations::handler` — which obligations apply and what each does before/after dispatch, plus the audit/redaction/ceiling/mount validation. - `obligations::staged_handoffs` — material staged for a later consumer: the runtime-secret and network-policy stores and the credential-account resolver port. - `obligations::process_store` — post-start handoff discard and reservation reconciliation. - `obligations::mod` — only `BuiltinObligationServices`, the assembly seam, and deliberately the one place naming all three at once. Every module is under the 1,500-line gate, so the waiver is deleted rather than carried forward: re-fusing the owners now trips `pre-commit-safety.sh`. `mod obligations;` stays private and the crate's `pub use obligations::{…}` names are unchanged, so no consumer outside the crate sees this. Behavior-free. Cross-owner access is `pub(super)` (three methods), not `pub(crate)`. The split revealed one narrowing in the other direction: `secret_present` was `pub(crate)` with no caller outside its own file and is now private. Also from the same CHECKLIST row, the bounded half of "shrink `services/builder.rs` toward composition-facing factories": three builder methods whose only callers are inside the crate's `src` narrow to `pub(crate)`. The rest of that clause is measured and deferred in the CHECKLIST amendment — 17 methods need a `test-support` cargo feature, three are callerless and belong to WS8, and the remaining 33 are a redesign of the fluent surface rather than a shrink of it. `+production_wiring` is refuted there: it is readiness diagnostics, not assembly. Two loud path-keyed gates fired and were repointed, not relaxed: `reborn_host_runtime_services_do_not_expose_lower_substrate_handles` now scans the whole `obligations/` directory and asserts it read ≥ 4 files (`collect_runtime_rs` returns a count; both its callers now assert non-zero), and `reborn_struct_test_support_ratchet`'s frozen per-file count moves to `staged_handoffs.rs` with its count unchanged at 1. Test accounting (un-masking discipline): `cargo test -p ironclaw_host_runtime --all-targets -- --list` is 1,246 before and 1,246 after, name-by-name identical — zero added, removed or renamed. `LAYER_MATRIX_EXCEPTIONS` is 10 before and after; an intra-crate split cannot move the register. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(operator,contracts): route operator secrets through a product_contracts port (WS3) `ironclaw_operator` is a products-tier crate and held `ironclaw_secrets`, the substrate that owns CAS one-shot leases, AAD/crypto and the OS keychain master key. PROPOSAL §8.2's product row says the products tier loses that edge, and §12.1b requires the port replacement to land before the edge is removed. Both happen here, in that order. - Port: `ironclaw_product_contracts::operator_secrets::OperatorSecretValueStore`. - Implementor: `ironclaw_reborn_composition::RuntimeOperatorSecretValueStore`, the same placement as `OperatorStatusService` — assembly is the only layer that may name both a products-tier port and a substrate. Registered in `INVERTED_PORTS` beside it. - `ironclaw_secrets` is gone from the operator manifest under every dependency kind, and `"ironclaw_secrets"` is now in the crate's `boundary_rules()` forbidden list. That gate's comment previously said the entry was deliberately absent because "the row owns it"; the row now owns it. The port is deliberately narrower than the substrate, so this is a tightening rather than a relocation: it takes no `ResourceScope` (the implementor fixes the operator scope, where the caller used to pass one), exposes no lease/consume protocol, and carries only a `&'static str` classification instead of the substrate's error `Display` — asserted, including that the backend message and the handle name are both absent from what crosses. Two tests travelled with the behavior rather than being pointed at a fake: `read_is_repeatable_across_reloads` (repeatability is a property of the lease protocol) and the #4673 production-store reproduction (its value is wiring the store exactly as production does, which now means the real store *behind the adapter*). Two `FaultInjecting`-over-real-store fixtures became per-operation port fakes, with the substrate error mapping re-pinned at the adapter; a third assertion got stronger — batched-vs-N+1 stored-key lookup is now observed at the port rather than by counting filesystem ops. Test accounting: operator 154 -> 153, product_contracts 142 -> 143, composition 937 -> 942 with zero removed; name-by-name diffs on a quiescent tree. Two findings the row could not have anticipated, both recorded in the CHECKLIST amendment: - The `webui` half of the row was already closed and was never a production edge. `ironclaw_secrets` has been a dev-dependency of `ironclaw_webui` since the commit that added it (#6619), both src mentions are `#[cfg(test)]`, and webui's boundary rule already forbade it. - `ironclaw_extension_manager` (layer `products`) still holds a normal `ironclaw_secrets` edge in `admin_configuration.rs`. §8.2 covers it; the row does not, because the crate landed with WS2.4 after the row was written, and the substrate sits in the service's type parameters so it is not a like-for-like swap. Filed as #7095. `LAYER_MATRIX_EXCEPTIONS` is 10 before and after: `products -> substrates` is matrix-legal, so this edge was always an §8.2 rule and never a layer exception. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(sandbox): put the Docker security check behind the fail-closed gate Review asked why the required Rust e2e lane can report `docker_security` as passing with no daemon. Half of that is #7081 (nothing sets IRONCLAW_REQUIRE_DOCKER_TESTS=1, so the switch is inert) and is not fixable from here -- arming it hard-fails any lane lacking a daemon or the worker image, which needs a runner guaranteed to have both. The other half is fixable here and is fixed: docker_security.rs open-coded its own `docker version` / `image inspect` checks with three bare `return`s, so it sat entirely outside docker_gate and would have stayed fail-open even once something did set the variable. It now takes both preconditions from docker_gate::{docker_available, docker_image_available} and skips with the visible `SKIP:` line that gate's module doc requires. Measured, same machine, image absent: before, IRONCLAW_REQUIRE_DOCKER_TESTS=1 -> "skipping ..." / 1 passed after, IRONCLAW_REQUIRE_DOCKER_TESTS=1 -> panic at docker_gate.rs:74 / FAILED after, variable unset -> "SKIP: ..." / 1 passed The third line is the no-op proof: the variable is set nowhere in this tree or on main, so no lane's behavior changes today. The daemon-down path already reached the image check and skipped there, so the outcome is identical; only the branch it takes differs. Two stale comments in docker_gate.rs corrected with it (they claimed docker_security used its own gate, and that docker_image_available had no consumer), and the crate's Known debt entry now splits the done half from the #7081 half instead of describing both as open. cargo test -p ironclaw_sandbox: 193 passed, 0 failed cargo clippy -p ironclaw_sandbox --tests --all-features -- -D warnings: exit 0 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(reborn): stop calling the unwired script lane an execution lane Two review findings, both correct, both artifacts of this PR's own renames. 1. engine-v2-to-reborn-parity.md note 4 read "a native script/software execution lane (`ironclaw_sandbox`, `RuntimeKind::Script`) sandboxed via `ironclaw_sandbox`" -- self-referential after the merge collapsed ironclaw_scripts and ironclaw_process_sandbox into one crate, and it contradicts note 5 four paragraphs down ("no production execution backend is wired for it"). Re-stated as the typed runtime contract it is, citing the measurement: `with_script_runtime` has zero production callers (`rg` finds only the builder itself, docs, and 30 test call sites). 2. CHECKLIST WS10 ratchet note 2 said "raise the percentage floor ...; only the line count should fall". That generalises WS3's sandbox merge, where observed coverage happened to rise. It is wrong as guidance for WS7, and the counterexample is in this same file: the 2026-08-03 entry from #7064 records ironclaw_runner falling 85.55% -> 82.53% because the shed removed the crate's better-covered half, holding the floor, and RATCHET FAILing in the merge queue. Note 2 now says re-capture from the merged artifact, and lower only with that entry's move-not-regression counterfactual (add the moved files back, confirm the union clears the old floor, plus a zero-tests- lost name set-diff). cargo test -p ironclaw_architecture: 32 targets, 206 passed, 0 failed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): pin the WIT scope probes and the embedded-asset owner pairing Three review findings on the `wit/` move, each verified before it was acted on. 1. `ws12_workflow_contracts.py` probed `crates/ironclaw_wasm/wit/host.wit` and its nested twin. No `host.wit` exists in this repository — `git ls-files '*.wit'` returns only `tool.wit` and `channel.wit` — so both probes sat under the `crates/([^/]+/)*ironclaw_wasm/` alternative and re-asserted the crate-name term while saying nothing about the canonical ABI contracts. In a validator whose stated design is "probe derived from reality rather than from a guessed layout", a fabricated filename is a defect on its own terms. Replaced with a `crate_globs` entry, `("ironclaw_wasm", "wit/*.wit")`, which discovers the contracts on disk, requires each in scope, and synthesises the nested WS7 form — so a third contract, or the directory leaving the crate, fails the pin instead of passing on a stale name. Verified non-vacuous: narrowing the workflow alternative to `.../ironclaw_wasm/src/` now reports `tool.wit`, `channel.wit` and the nested probe as out of scope. 2. The embedded-asset routing test substituted `alpha`/`beta` owners so it could reuse the synthetic workspace. That exercised the real prefix strings through the real routing, but left the prefix->owner *pairing* — the table's entire semantic content — asserted nowhere: swapping `ironclaw_extension_support` and `ironclaw_extension_host` passed. Fixed in two halves. The routing test now drives the real `EMBEDDED_ASSET_OWNERS` against a workspace carrying the real owners' names and real manifest paths (the synthetic one could not: `build_plan` rejects a changed package outside the canonical set), asserting the real owner is selected. And the not-stale test now derives the same pairing from the tree instead of restating the constant: it resolves every literal `include_str!`/`include_bytes!` in every workspace crate through `crate_tree`, keeps the targets no crate owns — the ones that actually reach the table — and asserts that every crate compiling one of them is the routed owner or a dependent of it. That surfaced a property worth pinning: `crates/extensions/packages/` is embedded by four crates, not one. `ironclaw_extension_host`, `ironclaw_extension_manager` and `ironclaw_reborn_composition` reach into it alongside `ironclaw_extension_support`, and routing to the support crate covers them only because each depends on it. If that edge goes, a shipped artifact change stops scheduling a crate that embeds it — the silent under-schedule the table exists to prevent. Regression coverage verified red by sabotage, all three wrong tables: owners swapped (7 failures), `packages/` -> `ironclaw_llm` ("embeds nothing from it"), and the hardest case, `packages/` -> `ironclaw_reborn_composition` — a real embedder that the other embedders do not depend on ("...does not depend on..., so routing there never schedules it"). 3. CHECKLIST WS10 claimed each of the nine `wit_bindgen` guest edits forces a committed WASM artifact rebuild. Only six do: `scripts/ci/check-wasm-artifact-freshness.py` scans `crates/extensions/packages/*/wasm-src` alone, `wasm-src-digests.toml` holds exactly six entries, and `git ls-files '*.wasm'` returns exactly those six. The three `test-tools/*/wasm-src/` guests commit no artifact; the tenth site is the host's `bindings.rs`, not a guest. Corrected, and the `wit/` row now states the boundary rather than implying it. Guest paths, `wit/` contents and the six rebuilt artifacts are untouched. Verified: `test_reborn_pr_test_plan.py` 46/46, `test_ws12_workflow_contracts.py` 25/25, `ws12_workflow_contracts.py` green on the real tree, `cargo test -p ironclaw_architecture` 206/206 across 32 binaries. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(host-runtime): state the obligation visibility rule as it holds Review catch (#7090): the guardrail sentence promised "cross-owner access is `pub(super)`, never `pub(crate)`", which is stronger than the code. Verified: `RuntimeSecretInjectionStore::{insert, take, clone_material, discard_for_capability}`, `NetworkObligationPolicyStore::{insert, get, take, discard_for_capability}` and both constructors are `pub(crate)` and must stay so — `src/egress/{mod,host_port,credential}.rs` call them, and that is host-runtime composition outside `obligations/`. The rule is restated as the property that actually holds: a method whose only callers are inside `obligations/` is `pub(super)` (the three that are), and `pub(crate)` is what the stores expose to the egress pipeline they exist to serve. A future agent reading the old sentence would have read the existing `pub(crate)` methods as violations. Guidance-only; no code change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(architecture): put the operator secrets boundary entry on the right rule Review catch (#7096), and it is the serious kind: the `"ironclaw_secrets"` entry landed in `ironclaw_extension_contracts`'s forbidden vector, not `ironclaw_operator`'s. The suite still passed, because `extension_contracts` has no such dependency and `ironclaw_operator` then had no entry at all — so the guard this row exists to add was inert, and a green architecture suite was evidence of nothing. Reintroducing the edge would have passed every check. Moved to `ironclaw_operator`'s vector; `extension_contracts` restored to its `origin/main` content byte-for-byte. Negative-probed rather than assumed. With `ironclaw_secrets` temporarily re-added to `crates/ironclaw_operator/Cargo.toml`: reborn_crate_dependency_boundaries_hold ... FAILED ironclaw_operator must not have a normal dependency on ironclaw_secrets and with the manifest restored, 35/35 pass. Two further review findings, both verified before being accepted: - `ironclaw_extension_manager` **does** have a `boundary_rules()` entry (`:3543-3556`, added with WS2.4). The CHECKLIST residue note and PROPOSAL §8.2's 2026-08-02 amendment both said it had none; §8.2's sentence is stale and is marked superseded. The real gap is narrower and now stated: the rule exists and simply does not forbid `ironclaw_secrets` (#7095). - `ironclaw_product_contracts`'s guide claimed "twenty-four shipped modules". Measured: `src/lib.rs` has 26 shipped (27 `pub mod` less the gated `test_support`), and the table was missing `ironhub` **before** this branch touched it. Count corrected to twenty-six and the missing `ironhub` row added, so the inventory matches `lib.rs`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sandbox): state the Docker-gate claim as the search that checks it Review caught a false inventory in the Known debt entry, and the previous commit is what made it false: "the name appears only in docker_gate.rs and attribution_tests.rs" stopped holding the moment docker_security.rs gained a module doc naming the variable, and CLAUDE.md itself was already a third counterexample. The narrower claim is the one that was always meant and is the one that matters, so it now carries its own reproduction: no workflow, script, env file or manifest mentions the name at all -- `git grep` over *.yml/*.yaml/*.sh/ *.toml/*.py/*.json/.env* is empty here and on main -- and the sole code reference is a read, std::env::var(...) at docker_gate.rs:23. Every other occurrence is a doc comment or a panic message. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(coverage): re-anchor the exemptions the merge shifted tests/integration/changed-coverage-exemptions.toml is exact-line-keyed and auto-merges silently. #7096's additions to ironclaw_reborn_composition moved four entries' subject lines by +2 without anything flagging it; a stranded entry makes the changed-coverage validator abort with no verdict at all. Re-anchored by content (difflib line map from the #7065 tree, which the file was validated against, to the union) rather than by arithmetic: runtime.rs [4068..4073, 4082, 4083] -> [4070..4075, 4084, 4085] runtime.rs [3701] -> [3703] ; runtime.rs [3433] -> [3435] lib.rs [616] -> [618] All 142 entries / 1124 line references re-verified against the merged tree: 0 drift, 0 out-of-bounds, 0 missing paths. * refactor(layers): re-layer processes -> kernel and skills -> substrates (WS3/WS4) Two CHECKLIST rows, both of which were a one-line manifest correction rather than a code move: the family docs already placed both crates where the rows want them and only `Cargo.toml`'s `layer =` disagreed. processes -> kernel (WS3). families/kernel.md already lists ironclaw_processes among the kernel crates. The re-layer makes processes -> resources a kernel -> kernel edge, so its LAYER_MATRIX_EXCEPTION went STALE and the gate said so itself: Stale IronClaw crate layer matrix exceptions: ironclaw_processes -> ironclaw_resources from 2026-07-09 should be removed in W7: runtime process management still depends on resource contracts currently classed with kernel behavior That is the gate's verdict, not a judgement call - deleting the entry is the only way to make it pass. Baseline 5 -> 4, recomputed as len(merged list). Checked the direction both ways: all nine crates that take a normal dependency on processes (capabilities, turns, host_runtime, extension_host, loop_host, extension_manager, runner, reborn_composition, stress) are kernel or above, so the move legalizes an edge without forbidding an existing one. skills -> substrates (WS4 SS3.D). families/domains.md already lists ironclaw_skills under 'Layer(s): substrates'. Its only two normal dependencies are ironclaw_filesystem (substrates) and ironclaw_host_api (contracts), both at or below substrates, and its six consumers are all loops or above. No exception moves in either direction. cargo test -p ironclaw_architecture: 206 passed, 0 failed. cargo check --workspace --all-targets: clean. * docs(target-arch): close the WS3/WS4 rows this work satisfies, with evidence Every tick was verified against the merged tree, never against a PR title. TICKED: - sandbox lane merge: ironclaw_sandbox exists, ironclaw_scripts and ironclaw_process_sandbox absent, bollard/rcgen declared by exactly one manifest in the workspace. - mcp drops the registry dep: ironclaw_extensions is [dev-dependencies] only, 0 production ironclaw_extensions:: refs in src/. - skills -> substrates: landed here. - hooks libSQL/Postgres [decision]: ADR recorded - keep both, with the four rejected alternatives and the evidence they are already converged on one trait plus a shared conformance suite. #6945 read first as the row demands, and explicitly NOT discharged: this PR changes nothing in the dispatch path. - WS3 verify row: the row conflated Wave 3 with Wave 5 work (9 of its 10 exceptions carried removes_in = W7). Corrected with the replaced text quoted, the Wave-3 half satisfied edge by edge, and the Wave-5 remainder named with its owning field value. Ticked on the corrected condition. LEFT OPEN OR PARTIAL, each with measurements rather than a hand-wave: - first_party_tools: 1 of 6 families moved; 15 modules still in host_runtime. Ticking would be false. - processes/capabilities row: re-layer DONE; the capabilities/host.rs split is deferred with every module boundary already computed (4,560 lines, the six workflow ranges, and the arch-exempt waiver that must be deleted with it). - host_runtime binding/catalog-defaults: binding half REFUTED (moving it needs RuntimeLaneExecutor/RuntimeLaneRequest made pub, contradicting the same section's Keeps clause; zero external references to either). Catalog half cannot go to extension_host at all - host_runtime is itself a production consumer at memory_native_extension.rs:96,101, so the move is a kernel -> products edge and a Cargo cycle. Correct destination is downward. - network test_rewrite: NOT executed. Recorded the security shape (production binaries compile the seam and honour the rewrite env var at runtime) and the full 6-step plan, because the env var is how the entire E2E suite redirects vendor traffic through the production binary and the change needs feature forwarding into CI lanes I cannot verify here. cargo test -p ironclaw_architecture: 206 passed, 0 failed. * ci(coverage): recapture the two composed floors from a real measurement The provisional values were arithmetic - the sum of the two slices' recorded deltas - and the dispatch caught them, which is the whole reason the brief demanded a measurement rather than a reconciliation. Dispatch run 30907774036 at 4512e03e28f1df15b419d2e36f9f38f8f55d62fd: 26 success / 1 skipped / 2 failure, judged by per-job tally per #6978. The one skip is the pull_request-gated mutation gate; the two failures are the coverage report and the roll-up it drags down, i.e. this file doing its job. ironclaw_host_runtime: predicted 89.05% (18801 / 21114), MEASURED 88.63% (17562 / 19814). The composition was wrong by 1300 denominator lines because both slices measured their delta under the pre-#7083 aggregator, which could not see crates/extensions/** at all - lines leaving host_runtime for extension_support vanished from the tree it could measure, so neither branch's recorded delta describes the post-#7094 world. ironclaw_extension_support: MEASURED 75.31% (7142 / 9484) against #7094's 82.64% (6826 / 8260), captured before #7080's executor lines arrived. floor_percent FALLS 7.33pp and that is flagged in the file for an owner's eye rather than written quietly. Evidence it is composition and not lost tests: floor_covered_lines RISES 6826 -> 7142, so the crate is protected by more absolute lines than before, and #7080's un-masking accounting was 1398 -> 1398 with zero test names lost. Same shape as #7094's own ironclaw_runner recapture. ironclaw_sandbox passed unchanged at its arrival capture (87.09%, 3185 / 3657). The [global] entry is untouched: both moves are crate-to-crate inside the set the fixed aggregator sees. * fix(network): compile the test rewrite seam out of production builds (WS3) Closes the WS3 network row. Also RETRACTS an overstatement I made in this row's earlier annotation. CORRECTION FIRST. The earlier note claimed production binaries compile the seam and honour IRONCLAW_REBORN_TEST_HTTP_REWRITE_MAP at runtime, so anyone able to set it could redirect all credentialed vendor egress. That was WRONG. RewriteNetworkTransport::from_env_value already returned UnavailableInRelease when !cfg!(debug_assertions) (test_rewrite.rs:150), and neither [profile.release] nor [profile.dist] sets debug-assertions, so a shipped binary with the variable set REFUSES TO BOOT. It was fail-closed before this PR. I had read the ungated `mod test_rewrite;` declaration as an ungated runtime path. What was genuinely wrong, and is fixed: 1. The guard was a RUNTIME check keyed on cfg!(debug_assertions) - a profile proxy, not a build-kind guarantee. A release profile with debug-assertions turned on (normal when chasing a production bug) silently re-arms it. 2. The refusal arm had NO TEST. The one guard between a shipped binary and redirectable vendor egress was unpinned. Fix: compile-time exclusion instead of a runtime check. mod test_rewrite and its four re-exports are now cfg(any(debug_assertions, feature=test-support)), and default_host_http_egress is a compile-time pair - production builds PolicyNetworkHttpEgress<ReqwestNetworkTransport> directly, with the rewrite wrapper absent from the binary. The runtime check stays as defence in depth. E2E needs no change: those harnesses build DEBUG binaries, so they satisfy debug_assertions and keep redirecting with no feature flag and no workflow edit. The feature-forwarding-into-CI risk I flagged earlier does not arise. test-support is still forwarded composition -> network for a release-PROFILE build that needs the seam. Both halves proven rather than assumed: (a) release refuses - new regression test a_set_rewrite_map_activates_only_in_debug_and_is_refused_in_release feeds a well-formed map and asserts on profile. Under 'cargo test --release -p ironclaw_network --features test-support' it passes on the UnavailableInRelease branch; under debug 'cargo test -p ironclaw_network' it passes on the active branch. 56 passed, 0 failed. (b) production compiles without the seam - 'cargo check --release -p ironclaw_reborn_composition' (no test-support) is clean, which only compiles if the cfg(not(..)) arm is right. Also: WS0_EXTENSION_SPECIFICITY_ALLOWLIST_BASELINE 129 -> 127. The constant had drifted ABOVE the real list length; the ratchet is shrink-only so it passed silently while buying back two unearned slots. Measured off the compiler (set baseline to 0, read the reported length), identical on main and on every slice, so pre-existing drift rather than something this PR caused. cargo test -p ironclaw_architecture: 206 passed, 0 failed. cargo check --workspace --all-targets: clean. * docs(coverage): verify the extension_support floor drop is composition, independently The 82.64 -> 75.31 recapture carried a rationale that was recorded but explicitly NOT verified. Re-derived it from scratch between the two capture refs (f946a93fae -> 939af4847d) rather than inheriting the claim: - 0 test names lost in the crate (158 -> 160 test fns; both new names belong to the arriving executor). - 0 test names lost WORKSPACE-WIDE (13836 -> 13843 test fns, 13752 -> 13759 unique). This is the check that separates a relocation from a deletion: host_runtime's roster drops 156 names over the same range and every one reappears in another crate. - Exactly four files arrived, 1367 source lines, all of them the family-1 skill-install executor (src/skills/url_install.rs + url_install/{github, zip_bundle,bundle}.rs). No pre-existing file left the crate. - The arithmetic closes with the pre-existing numerator held CONSTANT: (6826+316)/(8260+1224) = 75.31% exactly, so the pre-existing code lost zero covered lines. The arriving block's own coverage is 316/1224 = 25.82%. Composition, confirmed rather than assumed. No test regression to fix; the 25.82% arrival is what earns the follow-up already recorded above the entry. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(host_runtime): collapse a duplicated obligation predicate and quiet a background warn! Three verified review findings from the #7141 round. Each was confirmed against the code before being acted on; nothing was changed on assertion alone. 1. obligations/handler.rs — `obligation_supported_before_dispatch` and `obligation_supported_after_dispatch` had BYTE-IDENTICAL 19-line bodies (verified by exact line-by-line comparison). Both were private, each called exactly once, both taking the same `phase` argument. The two names asserted a pre/post-dispatch distinction the code never implemented, while the pair gates admission of RedactOutput, EnforceOutputLimit and EnforceResourceCeiling — so editing one copy alone would have left the other stage accepting an obligation the host cannot honour (a fail-open). Collapsed to one `obligation_supported`, with the reasoning recorded so the pair is not reintroduced. 2. obligations/process_store.rs — `cleanup_terminal` is reached from `observe_process_commit` (an async background journal callback, call sites at :363/:379/:394), so its `tracing::warn!` violates the repo rule that background tasks never use info!/warn! — they corrupt the REPL/TUI display. Lowered to `debug!`; the error is still returned to the caller on the next line, so nothing is swallowed. 3. reborn_restructure_baselines.rs — the doc table said the LAYER_MATRIX_EXCEPTIONS count was "now 11". Recomputed on this ref by anchoring on the `= &[` of the value (the `&[LayerMatrixException]` type annotation opens a bracket on the same line and silently yields 0): the real count is 4, matching WS0_LAYER_MATRIX_EXCEPTION_BASELINE = 4. Corrected. Verification: cargo check --all-targets -p ironclaw_host_runtime exit 0; obligation tests 13+26 passed, 0 failed; reborn_restructure_baselines 1 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): a shipped package prompt is an asset, not prose — it was selecting no lane Review finding on #7141, confirmed empirically before acting. The Markdown prose carve-out in the planner ran BEFORE the `EMBEDDED_ASSET_OWNERS` lookup. A prompt is a `.md` file that no package *directory* owns, so a change to `crates/extensions/packages/*/prompts/**.md` took the prose arm and planned: mode=none crate_buckets=[] "crate-tree guidance changed: ..." while its sibling `manifest.toml` in the same package planned `mode=selected` onto ironclaw_extension_support + ironclaw_extension_host. Prompts are shipped production output that `ironclaw_extension_support` compiles in, and the comment above `EMBEDDED_ASSET_OWNERS` names "manifests, prompts, schemas and built wasm/*.wasm" as exactly what that table owns — so this was the "silent under-schedule of a change to production output" that comment forbids. 145 of the 149 `.md` files under `packages/` are prompts. The rule is keyed on the `prompts/` path segment, not on the asset prefixes. That distinction is load-bearing: the first attempt yielded to the asset prefixes wholesale and broke `test-tools/README.md`, which is documentation of the fixture bundles and is deliberately pinned as prose. Of the four asset kinds the table owns, only a prompt is Markdown (manifests are .toml, schemas .json, wasm .wasm), so `.md` asset <=> prompt is exact. Sabotage-tested in both directions: * `_is_package_prompt` -> False (reinstates the bug): RED, "AssertionError: 'none' != 'selected'". * `_is_package_prompt` -> any .md under an asset prefix (over-broad): RED on both the new test and the pre-existing `test_markdown_owned_by_no_crate_is_prose`, at `test-tools/README.md`. * restored: 52 passed, 51 subtests, green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(harness): refresh the latency-runner lockfile after the sandbox consolidation Review finding on #7141, reproduced before fixing. The latency harness keeps its own committed `Cargo.lock`, separate from the workspace lockfile, and the crate consolidation that replaced `ironclaw_scripts` + `ironclaw_process_sandbox` with `ironclaw_sandbox` never regenerated it. It still carried entries for both removed packages (lines 3244 and 3602) and the old host-runtime/loop-host dependency graphs. Reproduced exactly as reported: $ cargo metadata --locked --manifest-path harness/latency/runner/Cargo.toml error: cannot update the lock file ... because --locked was passed exit 101 so any reproducible invocation of the harness was broken, while the documented unlocked command silently rewrote the lockfile as a side effect of running. Regenerated with `cargo update --workspace`, which re-resolves the path dependencies. Verified after: `--locked` exits 0, the two removed packages are gone (0 entries), and `ironclaw_sandbox` is present (1 entry). Note: the re-resolve also carried three registry deps forward (wasmtime-wasi 46.0.1 -> 47.0.3, wasmtime-wasi-io likewise, wit-parser 0.251.0 -> 0.252.0). That is contained — this lockfile governs only the standalone benchmark harness and is not the workspace lockfile, and it was already unusable under `--locked` before this change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> …
l3ocifer
pushed a commit
to l3ocifer/frick-ironclaw
that referenced
this pull request
Sep 3, 2026
…factory port (nearai#7202) * refactor(contracts): move extension runtime descriptors to a neutral contract (WS3) Deletes the two `-> ironclaw_extensions` layer-matrix exceptions (`ironclaw_mcp`, `ironclaw_scripts`) by giving the runtimes-layer lanes a contracts home for the descriptors they read, instead of the registry crate they may not depend on. Exceptions 13 -> 11; baseline lowered in the same change. Moved to `ironclaw_extension_contracts`: - `runtime::{ExtensionRuntime, ExtensionAssetPath, ExtensionAssetPathError}` - `hosted_mcp::{HostedMcpDiscoveredTool, HostedMcpDiscoveredToolAnnotations}` `ExtensionPackage`/`ExtensionManifest` deliberately stay in `ironclaw_extensions`: they carry the whole parsed manifest tree and a `PackageRootBinding` typed on `ironclaw_filesystem::VirtualPath`, which the §11.2.3 contracts-purity allowlist (`{ironclaw_host_api}` only) forbids the contracts crate from naming. Measured instead: both lanes read exactly three things off the package — `id`, `capabilities`, `manifest.runtime` — so the lane request structs now take those three and the caller (which owns the package) projects them. Also repointed `ResourceReceipt` to its real owner: `ironclaw_resources` only re-exports `ironclaw_host_api::resource::ResourceReceipt`, so the lanes' import was a §11.2.4 two-import-paths hop, not a dependency. No `pub use` shims (§11.3): every consumer is repointed in this change, and `resolve_under` becomes the free function `ironclaw_extensions::resolve_asset_under` because the orphan rule forbids an inherent impl on the moved type. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(sandbox): merge the sandbox lane into one crate (WS3) Creates `ironclaw_sandbox` (runtimes) from the three halves of "run an already-authorized command away from the host", and deletes the two crates PROPOSAL §6.6.4 marks for merge: - `ironclaw_process_sandbox` (plan contract) -> `src/plan.rs`, `src/validation.rs` - `ironclaw_host_runtime::sandbox_process` -> `src/sandbox_process/**` - `ironclaw_scripts` (script lane + Docker path) -> `src/script.rs` The kernel sheds the Docker/CA cone: `bollard`, `rcgen`, `x509-parser` and `time` are gone from `ironclaw_host_runtime`'s manifest, and `bollard`/`rcgen` are now declared by exactly one crate in the workspace. Two migration details PROPOSAL §6.6.4 and CHECKLIST WS10 call load-bearing: - `PROCESS_SANDBOX_CAPABILITY_ID` -> `ironclaw_host_api::capability`, so `ironclaw_loop_host` drops its lane dependency (production dep gone; a dev-dep remains for the tests that build plans). - `SandboxCommandTransport` -> `ironclaw_host_api::process`, with the shapes it names (`CommandExecutionRequest`/`Output`, `RuntimeProcessError`, `SavedCommandOutput`, `SavedCommandOutputSanitization`). Without this the runtimes-layer lane could not implement what the kernel consumes. Enumerating gates were repointed, never relaxed: the specificity carve-outs and the struct/test-support ratchet entries moved with their files (both baselines unchanged at 129 and their prior values), the panic-gate baseline row moved, `reborn-crate-test-buckets.sh` registers the new crate, and the three `reborn-e2e-rust.sh` script selectors follow the tests (plus `docker_security`, which had no selector before). One gate would have gone silently vacuous and was fixed rather than moved: the script-lane surface scan in `reborn_dependency_boundaries.rs` read a hardcoded `src/lib.rs`, which after the merge no longer holds the lane. It now scans the whole crate source tree with a fatal-read walk and a non-vacuity assertion. One deletion, recorded: `RebornScopedSandboxCommandTransport::into_process_port` returned a kernel type a runtimes crate may not name. It had zero callers workspace-wide; the kernel wraps the transport, which is the direction the port inversion requires. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): record the WS3 corrections with their evidence Three dated amendments, each quoting the text it replaces: 1. CHECKLIST WS3 sandbox row + PROPOSAL §6.6.4 — "all pieces currently unwired/test-only" is REFUTED. Three production paths cross the merged crate (spawn-path plan validation, the process_executor routing check, and the saved-command-output scope digest). The accurate claim is narrower: no production *execution backend*. Behavior preservation is therefore argued at the diff (11 of 26 moved files byte-identical, 9 more differing by one import line, +63/-36 overall), not inferred from deadness. 2. CHECKLIST WS3 mcp row + PROPOSAL §6.6.3 — the prior wave's "structurally blocked" finding is half right, and the wrong half is load-bearing: only `ExtensionPackage` is un-absorbable, and no lane ever needed it (both read `id`, `capabilities`, `manifest.runtime` and nothing else). The registry half of the flip is done; the `resources` half is refuted as phrased — the estimate/usage vocabulary the row asks about is already in `host_api::resource` and already imported from there, while the real blocker is the `ResourceGovernor` authority port and `ResourceError`'s denial cone. 3. Recorded as a structural finding, not a note: the sandbox row and the mcp row are ONE problem. `ironclaw_scripts` imports the identical DTO set, so the merge alone deletes zero exceptions and only the mcp carve-out lets either lane shed the registry edge. Also reconciled: PROPOSAL §6.1.2's as-built inventory gains the two modules WS3 landed (and states why `ExtensionPackage` stayed); §2's package count 66 -> 65; the §9 disposition rows for `ironclaw_scripts`/`ironclaw_process_sandbox`/ `ironclaw_mcp`; the §11.2.2 ratchet rows (13 -> 11); the WS3 verify row; the stale WS1.3 sentence asserting the blocker as settled fact; and `reborn_restructure_baselines.rs`'s doc table, which still read 15. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(sandbox): drop imports the merge left unused `process_port.rs` no longer names `MountView` or `thiserror::Error` (both went to `host_api::process` with the types that used them), and `sandbox_process.rs` no longer needs `sync::Arc` after `into_process_port` was deleted. Found by per-crate `clippy --all-targets --all-features -D warnings`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): let the Reborn PR planner plan guidance edits and crate deletions Three fail-closed gaps in `reborn_pr_test_plan.py`, all hit by this PR and all live on `main` today — any PR with the same change shape is unplannable. 1. `.claude/**` was unclassified, so the planner refused outright. It is agent guidance in exactly the sense `docs/**` is human guidance: no Rust test reads either as data (the only in-tree references are prose citations in test doc comments). Added to `IGNORED_PREFIXES`. Without this, "guidance travels with the change" — the restructure's own discipline — cannot be satisfied in a single PR. 2. `crates/AGENTS.md`, `crates/README.md`, `crates/Architecture.md` raised "unmapped crate path": they sit under `crates/` but belong to no package. Now classified as crate-tree prose, matched by "Markdown no package directory owns" so a genuinely unmapped crate path is unaffected. 3. An unmapped crate path used to raise. `git diff` reports a deleted crate's old paths and CI feeds the planner that diff, so **every crate deletion or rename was unplannable** — including the six deletions PROPOSAL §2 plans. It now widens to the exhaustive plan. This is a semantic change and it is the safe direction: the full plan is a superset of any narrowing, so an unattributable path can never cause under-selection, whereas refusing to plan blocks the PR instead of protecting it. Malformed input is still rejected by the unclassified-path branch. Each lands with fixtures per WS10's rule, positive and negative: guidance paths select nothing while non-guidance paths still fail closed; crate-tree prose selects nothing while crate *code* under the same unmapped directory widens to `full` (so the Markdown carve-out cannot swallow code). The pre-existing `test_unmapped_crate_path_fails_fast` is renamed and rewritten to pin the new contract rather than deleted. Verified against this PR's real 130-path diff: the planner returns `mode: full`, and the workflow's own exhaustiveness guard passes on that output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(arch): give the retained resource exceptions an owning issue, not a wave Review (#7065) caught that both surviving `-> ironclaw_resources` exceptions declared `removes_in = "WS3"` — the wave this PR *is*, which does not remove them. That is precisely the defect §11.2.2 already records against `conversations -> turns` ("`removes_in = "WS5"` and WS5 has partly shipped without it falling"), and it would have been repeated here. Both now point at issue #7067, which owns the design work that actually clears them: replacing the `ResourceGovernor` dependency with a narrow reserve/reconcile/release port. The issue carries the measurements — 3 of 10 methods used, zero implementors, and the `ResourceError` denial cone — plus the two open questions (error shape, port home) that make it a design slice rather than a move. An owning issue is also what §11.2.2 asks for and what the ratchet still cannot enforce (there is no `owning_issue` field yet), so this is the strongest form currently expressible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(contracts): pin the asset-path validator that moved into extension_contracts `validate_asset_path` moved here with `ExtensionAssetPath`, the type it constructs. In `ironclaw_extensions` it was only ever reached indirectly through manifest parsing, so its six rejection branches had no direct test — and a contracts crate that carries validation owes that validation one. Two tests: every reject branch with its exact reason and `Display` output (empty, NUL/control, URL, absolute, Windows drive and backslash, and the empty/`.`/`..` segment cases) plus the manifest-relative shapes that must keep being accepted; and `ExtensionRuntime::kind()` over all five variants, since that projection is what every lane uses to reject a runtime it does not serve. Also removes a changed-line coverage risk this PR would otherwise carry into the merge queue: the gate does not run on ordinary PRs (#7036), so ~100 newly-added lines of validator would first be measured where a failure is expensive to diagnose. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(coverage): re-capture the host_runtime floor and floor the new sandbox lane `RATCHET FAIL: ironclaw_host_runtime` — observed 18854 covered vs a `floor_covered_lines` of 20538. This is the shrinkage case the ratchet's own "To fix" text describes, not a coverage regression: `sandbox_process/**` moved to `ironclaw_sandbox`, so the crate's denominator fell 23277 -> 21267 (-2010 instrumented lines) and its covered lines fell with it. The percentage floor is **raised, not lowered**: observed 88.65% against an old floor of 88.23%, so the entry now reads 88.65. Only the absolute line count moves down, and it must — those lines are no longer in this crate. To keep that from being a net loss of protection, `ironclaw_sandbox` is floored on arrival at its observed 87.09% (3185 / 3657). This is a net *increase* in ratchet coverage: neither `ironclaw_scripts` nor `ironclaw_process_sandbox` was ever floored, and the `sandbox_process` half was protected only as part of host_runtime's line count, which this PR necessarily reduces. Floored crates 16 -> 17. Verified by replaying the ratchet arithmetic against CI's observed numbers: both crates pass on percentage and on covered lines. Numbers taken from the failing run's own report (job 91740733521), which is the authority for this gate. The `Tests (Reborn)` roll-up failed solely on this sub-job ("coverage-report result 'failure' did not match planned=true"); no other lane failed — 50 pass, 2 fail, both this root cause and its roll-up. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-architecture): record the coverage ratchet as a move-sensitive gate WS3 hit a gate no move row had named. `tests/integration/coverage-floor.toml` is keyed on crate identity plus absolute covered-line counts, so it is invisible to WS10's path-keyed gate audit and yet it fails on every crate move, merge, rename, or family `git mv` that shifts instrumented lines between crates — as it did here, while the percentage floor was *improving*. Recorded on WS10 with the three rules WS7 will need: re-capture in the same PR, raise the percentage floor rather than leaving it, and floor the destination crate or the move silently drops that code out of the ratchet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(extension-manager): repoint ironhub onto the moved ExtensionAssetPath A semantic conflict the merge could not see: #6780 landed `ironhub/{package,catalog}.rs` importing `ExtensionAssetPath` from `ironclaw_extensions`, while this branch moved that type to `ironclaw_extension_contracts::runtime`. Different files, so git auto-merged cleanly and the breakage surfaced only at `cargo check`. Repointed both sites to the contracts crate (no shim, per §11.3). The manifest already named `ironclaw_extension_contracts`, so this is imports only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(coverage): exempt the WS3 move's no-region lines and record the gate The changed-lines coverage gate went red on four files while changed-line coverage was 95.35% against a 90% floor: the failure was its two fail-closed STRUCTURAL assertions, not any percentage. Every line below was derived by replaying scripts/ci/reborn_changed_coverage.py against this PR's own merged lcov (run 30831658659) with the base lcov the gate itself resolved (run 30828540055 @ b89fcd3575), until the replay reproduced the CI verdict byte-identically. Line numbers come from the gate's own `candidate_lines - mechanically_uninstrumentable_lines()`, not from the log. - host_api/src/process.rs (31 lines): new placement-neutral process vocabulary with no function body anywhere in the file; rustc emits no LCOV record for it at all. Same shape already exempted for product_contracts/loop_contracts. - extension_contracts/src/hosted_mcp.rs (12): field declarations of the two new tools/list descriptor structs. The file is plainly instrumented (191 DA, 164 hit), so this is a no-region artifact, not an instrumentation gap. - host_runtime/src/services/runtime_adapters.rs (13): continuation lines of three rewritten calls, all PROVEN EXECUTING by their region-start heads (lines 380/434/977 score 24/16/63 hits). The four genuinely-uncovered lines in the same rewrite are deliberately NOT exempted -- the gate already subtracts them as pre-existing debt inherited from base. - composition capability_host_tests/approval_gates.rs (6): type positions in a test double whose body region scores 1 hit. The last one is a finding, not just a waiver: that file is 100% test code behind `#[cfg(test)] mod capability_host_tests;`, but the gate's test_only_path() recognises /tests/, /test_support/, */tests.rs and *_tests.rs and NOT a cfg(test) module DIRECTORY, so it measures it as production. It is the only such directory in crates/ today. Docs (target-architecture, same PR per the docs-truth rule): - CHECKLIST WS10 gains the changed-lines gate beside the ratchet row, cross- referencing the WS2.1 note rather than restating it: percentages are not what fail a move; derive lines by byte-identical replay (--fetch-base-coverage silently degrades without --github-repo); and a stranded exemption path is an ABORT with no verdict, not a loud failure. - CHECKLIST WS10 exception-ratchet row: the constant was cited at :4063 and sits at :4164 -- corrected by removing the line pin, since the file is edited every wave. Records that the baseline is a UNION across parallel WS3 lanes. - families/contracts.md: records extension_contracts' new ownership of the runtime descriptor vocabulary -- the carve-out that let BOTH lanes drop the registry edge -- and the orphan-rule seam that keeps resolve_asset_under in the registry crate. - families/lanes.md: two "Never" claims were reading as satisfied when they are not. ironclaw_mcp's "never depends on the resource-governor crate directly" is refuted (the compiled edge survives; #7067 tracks the narrow port), and ironclaw_sandbox's "no direct process spawning outside the transport seam" is aspirational -- script.rs:454 still builds Command::new("docker"). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sandbox,mcp): correct the wiring inventory and record the projection cost Two review findings verified against the tree; three refuted with evidence in the PR threads. Valid — the sandbox wiring inventory was self-contradictory. `CLAUDE.md` said "Two production call paths ... and both are plan validation" directly above a list of THREE bullets, and `lib.rs` omitted the third entirely. The third is real and is not validation: `host_runtime/src/process_output.rs:482` derives the scoped saved-output directory through `RebornSandboxScopeKey::from_scope`. That inventory is what tells a future agent which paths are live, so an undercount invites deleting a production path as dead code. Both surfaces now say three and no longer claim they are all plan validation (the `loop_host` capability-id comparison never was either). Valid, and recorded rather than redesigned — the registry carve-out cost a type-level invariant. Replacing `package: &ExtensionPackage` with independent `extension` / `capabilities` / `runtime` borrows is what deleted the `mcp -> extensions` and `scripts -> extensions` exceptions, but it also means the type no longer guarantees the three came from one package. `execute_extension_json` re-checks the descriptor half (`descriptor.provider == extension`); the runtime half cannot be re-derived, because nothing in an `&ExtensionRuntime` names its owning extension. No caller can trip it today -- there is exactly one production caller (`runtime_adapters`) and it projects all three from one package in one expression -- so this is a latent structural weakening, not a live defect. Restoring the compile-time binding needs a sealed projection minted by the package owner; a check inside the lane cannot express it, and re-taking the registry edge would undo the carve-out. Both request types now carry the caller obligation in their field docs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(extensions): move the skill-install executor to extension_support (WS3) WS3's first-party-tools row, family 1 of 6: skill management / URL install. `skill_url_install.rs` and its `bundle`/`github`/`zip_bundle` submodules, plus the install-input normalizer, move out of `ironclaw_host_runtime::first_party_tools` into `ironclaw_extension_support::skills::{url_install, resolve_install_input}`, where the skill executor half already lived. Move-only: no behavior change, no test edited for content. `ironclaw_host_runtime -> ironclaw_skills` is deleted from LAYER_MATRIX_EXCEPTIONS — the edge is gone, not waived (exceptions 13 -> 12, WS0_LAYER_MATRIX_EXCEPTION_BASELINE drops with it). `ironclaw_skills` and `zip` survive as dev-dependencies for host_runtime's own tests; dev edges are outside the matrix by construction. Two doc ambiguities are resolved in the same diff, as dated PROPOSAL amendments quoting the text they replace: - §6.8.4's "the builtin first-party tool handlers absorbed from host_runtime/first_party_tools" contradicted §8.2's "kernel: ✗ (ports only)" row and the enforced BoundaryRule. Resolution: the seam splits executor from adapter — the executor moves behind a neutral request/error pair, the FirstPartyCapabilityHandler / CapabilityManifest / registry wiring stay host-side. Same shape the groupware and web-access tools already ship. - §8.2's "ports only" cell now says what it means: contracts-layer ports the kernel also consumes, not permission to name a kernel trait. Two cost corrections recorded for the remaining families: `host_runtime -> extension_support` is not divisible family-by-family (mod.rs holds it via `extension_support::coding`), and `host_runtime -> ironclaw_extensions` is not reachable by this row at all. PATH_TERM_COLLISIONS shrinks by two: the installer's github carve-outs now sit inside a scan-exempt crate. Test accounting (un-masking discipline), unfiltered `--list` over both crates: 1398 -> 1398, with exactly two tests renamed by module path and none lost. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sandbox): record that the Docker fail-closed switch is wired to nothing Review asked why the migrated docker_security test can pass with no daemon. The skip is pre-existing (the file differs from its pre-merge original by one import line); WS3 only enrolled it in the required Rust e2e lane, where it was not run at all before. The real defect the question surfaced is worse and also pre-existing: this crate's tests/support/docker_gate.rs states that IRONCLAW_REQUIRE_DOCKER_TESTS=1 makes a missing daemon a hard failure and that "CI sets this" -- and nothing sets it. Repo-wide the name occurs only in docker_gate.rs and attribution_tests.rs, here and on main. So every real-Docker test in the crate skips-and-passes everywhere, which is exactly the gap the gate's own comment says let sandbox security bugs ship unnoticed. docker_security.rs additionally open-codes its own check rather than using the gate, so it would stay fail-open even once something did set the variable. Recorded rather than fixed: setting the variable is a CI-behavior change that would hard-fail any lane without a daemon or the ironclaw-worker image, which is not verifiable from inside a move PR whose evidence claim is behavior preservation. Filed as the #6945 guardrail-claim-vs-reality class with the two-part fix stated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(host_runtime): record the executor/adapter seam in crate guidance The crate's CLAUDE.md said "first-party runtime tools belong under `first_party_tools/`" without saying that only the host half does. WS3 moves each tool's executor into `ironclaw_extension_support`, which may not name this crate, so the rule now names both halves and points at the skill-install family as the worked example. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(host_runtime): keep the install-input error path log-free The moved executor returns `SkillManagementCapabilityError`, and routing it through `skill_management_error` would have added a `debug!` line to a path that had none before the move. A move-only change must not add one, so the install-input arm maps the kind directly and the `dispatch` arm keeps the record it already had. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci(coverage): re-capture the host_runtime floor for the WS3 executor move The ratchet does not run on `pull_request` (`reborn_pr_test_plan.py:21`; issue #7036), so this PR's green checks were not evidence on this axis. A full-plan `workflow_dispatch` run on this exact head reported: RATCHET FAIL: ironclaw_host_runtime observed: 88.59% (20485 / 23124 lines) floor: 88.23% ... floor_covered_lines: 20538 (effective floor 20518) The percentage went UP while `floor_covered_lines` went DOWN — shedding well-covered code lowers the absolute numerator, which is a separate assertion from the percentage one. Re-captured to the observed numbers (floor raised 88.23 -> 88.59, not merely held). Verified locally against that run's own merged lcov artifact: ENFORCING mode, 17 PASS / 0 FAIL, exit 0. run: https://github.com/nearai/ironclaw/actions/runs/30858257594 head: e07b3b0299b0add11117e9591da71d46d7a7c832 The destination crate is deliberately not floored, because it cannot be: every crate under `crates/extensions/` is invisible to the coverage tooling — `reborn_coverage_lcov.py:19`'s CRATE_RE still requires a crate directory directly under `crates/`, which #7037's colocation broke. Filed as #7083 with the measurement; the global floor is left alone rather than re-captured onto that hole. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(wasm): move wit/ inside its owning crate (Wave 3) CHECKLIST WS4 + WS10 `wit/` rows. `wit/{tool,channel}.wit` moves from the repo root to `crates/ironclaw_wasm/wit/` — the crate that owns the ABI — per PROPOSAL §6.6.1. Behavior-free: same bytes, same generated bindings. Wave-3 coordinates: the docs write the destination as `crates/lanes/ironclaw_wasm/wit/`, but `crates/lanes/` does not exist until WS7. Because the files now sit *inside* the crate, the WS7 family move carries them with no further path edit anywhere — which is the whole point of putting them there. Ten wit-bindgen `path:` args repointed (the host plus nine guests: six under `crates/extensions/packages/*/wasm-src/`, three under `test-tools/*/wasm-src/` — the CHECKLIST row said six). All nine guests verified building against the moved WIT on wasm32-wasip2. The four `include_str!` readers of the ABI text do NOT get repointed literals. Doing that would turn the two `ironclaw_host_runtime` sites from repo-root reach-ins into *cross-crate* ones — §11.2.7's strict class, the one WS2 turns into hard failures — taking the scan from 19 to 21 while ticking a box that says "§11.2.7 scan passes". Instead the ABI text gets one owner, `ironclaw_wasm::TOOL_WIT` (`src/config.rs`, beside `WIT_TOOL_VERSION`), and all four sites read the const over cargo edges that already exist. Measured with the scan: 133 -> 129 escaping sites, cross-crate 19 -> 19, zero `wit/` entries remaining. Path-keyed gates repointed: `scripts/check-version-bumps.sh` (both ABI paths), `.githooks/pre-commit`, and `platform-and-compat.yml`'s `has_direct_wasm_abi_risk` filter — where the bare `wit/` alternative is *deleted* rather than rewritten, because the filter's existing `crates/([^/]+/)*ironclaw_wasm/` alternative already matches both the Wave-3 and the WS7 location. `scripts/ci/ws12_workflow_contracts.py` anchored on that deleted string, so its anchor moves to `build-wasm-extensions` and its in-scope probe now pins both locations. `Dockerfile` loses two `COPY wit/ wit/` lines in the planner and builder stages: both already run `COPY crates/ crates/`, so the files arrive with the crate and the old line would COPY a path that no longer exists. Docs: the WS4 row's `crates/lanes/wit/` destination was the only doc site placing the directory beside the crate rather than inside it; corrected there and in README's tree, with dated amendments in CHECKLIST, PROPOSAL §6.6.1 and PLAN Wave 3 recording what the move found. Test accounting (unfiltered `--list`, name-by-name, quiescent tree): ironclaw_wasm 51 -> 51, ironclaw_host_runtime 1246 -> 1246, ironclaw_architecture 198 -> 198. Zero diff, no test edited for content. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * build(wasm): rebuild first-party artifacts for the moved wit/ path Forced by the previous commit, not incidental to it. `scripts/ci/check-wasm-artifact-freshness.py` keys each package's committed `wasm/<name>.wasm` to a digest of the `wasm-src/` tree that produced it, so editing a guest's `wit_bindgen::generate!` `path:` — which the `wit/` move requires in all six shipped guests — invalidates the recorded digest and fails the gate. The gate's own contract forbids the shortcut: "Re-record only after `./scripts/build-wasm-extensions.sh --first-party` and committing the rebuilt artifact — the digest asserts a claim about the artifact, and updating it without rebuilding launders a stale one." So the artifacts are genuinely rebuilt (`--first-party`, exit 0, 6 OK / 2 host-native SKIP), not re-recorded in place. Byte sizes move by more than the source change accounts for because these builds are not reproducible by design — the guests pin no toolchain and resolve their own `Cargo.lock` at build time, which is the documented reason the gate hashes sources rather than artifact bytes. Verified: `check-wasm-artifact-freshness.py` OK (6 packages), and `cargo test -p ironclaw_extension_support` green (102/46/4) — that crate `include_bytes!`s these artifacts, so it exercises the rebuilt components. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(target-arch): record the WS7 artifact-rebuild cost of guest path edits The `wit/` move had to rebuild six shipped WASM binaries because `check-wasm-artifact-freshness.py` digests each guest's whole `wasm-src/` tree. WS7 hits the same wall from the other direction: the six package guests reach the ABI across two trees, so moving either `ironclaw_wasm` or `extensions/packages` rewrites all six `path:` literals and forces the same rebuild. Recorded on CHECKLIST WS10's `wit/` row (point 6), on the loud-path-pattern row that owns the WS7 repoint (also corrected six -> nine guests there), and on PLAN's Wave 5 block with the cheap mitigation: move the two crates in one PR and pay it once. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci(planner): classify the path classes that blocked the wit/ move `Detect Reborn test scope` exits 1 on any pull request whose diff holds a path `reborn_pr_test_plan.py` has no rule for, which made this PR unmergeable: it must edit `Dockerfile` (the moved directory's `COPY wit/ wit/` no longer resolves) and `scripts/check-version-bumps.sh` (the ABI gate would otherwise grep dead paths and silently stop enforcing). 18 of its 46 paths were unclassified. Same class as the `.claude/` gap #7064 fixed, and classified the same way — one rule per class, recorded beside the constant: * `Dockerfile` / `.dockerignore` — `platform-and-compat.yml` keys `has_docker_risk` off exactly this pair and owns the image build. * `.githooks/**` — Code Style triggers on the tree and lints its contents (`test-ci-comm-locale-pin.sh`); no Reborn lane runs a hook. * `scripts/{build-wasm-extensions,check-version-bumps}.sh` — `platform-and-compat.yml`'s `has_direct_wasm_abi_risk` classifier both scopes and runs them. * markdown owned by no crate (`crates/AGENTS.md`, `test-tools/README.md`) — prose, like `docs/` and `.claude/`. A crate-resident doc still selects its own crate's lane. The first-party extension package assets are deliberately NOT ignored. `crates/extensions/packages/*/wasm/*.wasm` is a shipped artifact that `ironclaw_extension_support` embeds with `include_bytes!`, and `test-tools/*/manifest.toml` is `include_str!`d by `ironclaw_extension_host`. Calling either prose would convert today's loud failure into a silent under-schedule of a change to production output — the WS10 failure mode. `EMBEDDED_ASSET_OWNERS` routes each tree to the crate that compiles it instead, so this PR now additionally schedules `ironclaw_extension_{support,host,manager}`: the crates that consume the six rebuilt WASM artifacts. Also fixes #7085 in a file this PR already touches. The WIT version extractors used the GNU-only BRE `\+`, so on BSD sed (macOS) they matched nothing, and because the `WIT_TOOL_VERSION` cross-check is guarded on a non-empty version the hook printed "All version checks passed" having compared nothing. `[[:space:]][[:space:]]*` is identical under GNU sed, so the enforced Linux CI lane is unchanged; verified on BSD sed that both `wit/tool.wit` (0.3.0) and `wit/channel.wit` (0.3.1) now extract. Regression tests: every classified class gets a case in `test_reborn_pr_test_plan.py`, including the paired assertion that the embedded assets *select a lane* rather than merely being accepted (the inverse of the `.claude/` prose test), and a staleness pin that fails if an asset tree or its owning crate moves. All ten new cases fail against the planner on `main`. `test_unclassified_build_input_fails_fast` moves off `Dockerfile` onto a still-undecided input so the fail-closed arm stays exercised. Refs #7087, #7085 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(host-runtime): split obligations into its three chartered owners (WS3) `crates/ironclaw_host_runtime/src/obligations.rs` was 3,122 lines fusing the three owners PROPOSAL §6.5.9 charters separately, held apart only by an `// arch-exempt: large_file` waiver. It is now one module per owner: - `obligations::handler` — which obligations apply and what each does before/after dispatch, plus the audit/redaction/ceiling/mount validation. - `obligations::staged_handoffs` — material staged for a later consumer: the runtime-secret and network-policy stores and the credential-account resolver port. - `obligations::process_store` — post-start handoff discard and reservation reconciliation. - `obligations::mod` — only `BuiltinObligationServices`, the assembly seam, and deliberately the one place naming all three at once. Every module is under the 1,500-line gate, so the waiver is deleted rather than carried forward: re-fusing the owners now trips `pre-commit-safety.sh`. `mod obligations;` stays private and the crate's `pub use obligations::{…}` names are unchanged, so no consumer outside the crate sees this. Behavior-free. Cross-owner access is `pub(super)` (three methods), not `pub(crate)`. The split revealed one narrowing in the other direction: `secret_present` was `pub(crate)` with no caller outside its own file and is now private. Also from the same CHECKLIST row, the bounded half of "shrink `services/builder.rs` toward composition-facing factories": three builder methods whose only callers are inside the crate's `src` narrow to `pub(crate)`. The rest of that clause is measured and deferred in the CHECKLIST amendment — 17 methods need a `test-support` cargo feature, three are callerless and belong to WS8, and the remaining 33 are a redesign of the fluent surface rather than a shrink of it. `+production_wiring` is refuted there: it is readiness diagnostics, not assembly. Two loud path-keyed gates fired and were repointed, not relaxed: `reborn_host_runtime_services_do_not_expose_lower_substrate_handles` now scans the whole `obligations/` directory and asserts it read ≥ 4 files (`collect_runtime_rs` returns a count; both its callers now assert non-zero), and `reborn_struct_test_support_ratchet`'s frozen per-file count moves to `staged_handoffs.rs` with its count unchanged at 1. Test accounting (un-masking discipline): `cargo test -p ironclaw_host_runtime --all-targets -- --list` is 1,246 before and 1,246 after, name-by-name identical — zero added, removed or renamed. `LAYER_MATRIX_EXCEPTIONS` is 10 before and after; an intra-crate split cannot move the register. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(operator,contracts): route operator secrets through a product_contracts port (WS3) `ironclaw_operator` is a products-tier crate and held `ironclaw_secrets`, the substrate that owns CAS one-shot leases, AAD/crypto and the OS keychain master key. PROPOSAL §8.2's product row says the products tier loses that edge, and §12.1b requires the port replacement to land before the edge is removed. Both happen here, in that order. - Port: `ironclaw_product_contracts::operator_secrets::OperatorSecretValueStore`. - Implementor: `ironclaw_reborn_composition::RuntimeOperatorSecretValueStore`, the same placement as `OperatorStatusService` — assembly is the only layer that may name both a products-tier port and a substrate. Registered in `INVERTED_PORTS` beside it. - `ironclaw_secrets` is gone from the operator manifest under every dependency kind, and `"ironclaw_secrets"` is now in the crate's `boundary_rules()` forbidden list. That gate's comment previously said the entry was deliberately absent because "the row owns it"; the row now owns it. The port is deliberately narrower than the substrate, so this is a tightening rather than a relocation: it takes no `ResourceScope` (the implementor fixes the operator scope, where the caller used to pass one), exposes no lease/consume protocol, and carries only a `&'static str` classification instead of the substrate's error `Display` — asserted, including that the backend message and the handle name are both absent from what crosses. Two tests travelled with the behavior rather than being pointed at a fake: `read_is_repeatable_across_reloads` (repeatability is a property of the lease protocol) and the #4673 production-store reproduction (its value is wiring the store exactly as production does, which now means the real store *behind the adapter*). Two `FaultInjecting`-over-real-store fixtures became per-operation port fakes, with the substrate error mapping re-pinned at the adapter; a third assertion got stronger — batched-vs-N+1 stored-key lookup is now observed at the port rather than by counting filesystem ops. Test accounting: operator 154 -> 153, product_contracts 142 -> 143, composition 937 -> 942 with zero removed; name-by-name diffs on a quiescent tree. Two findings the row could not have anticipated, both recorded in the CHECKLIST amendment: - The `webui` half of the row was already closed and was never a production edge. `ironclaw_secrets` has been a dev-dependency of `ironclaw_webui` since the commit that added it (#6619), both src mentions are `#[cfg(test)]`, and webui's boundary rule already forbade it. - `ironclaw_extension_manager` (layer `products`) still holds a normal `ironclaw_secrets` edge in `admin_configuration.rs`. §8.2 covers it; the row does not, because the crate landed with WS2.4 after the row was written, and the substrate sits in the service's type parameters so it is not a like-for-like swap. Filed as #7095. `LAYER_MATRIX_EXCEPTIONS` is 10 before and after: `products -> substrates` is matrix-legal, so this edge was always an §8.2 rule and never a layer exception. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(sandbox): put the Docker security check behind the fail-closed gate Review asked why the required Rust e2e lane can report `docker_security` as passing with no daemon. Half of that is #7081 (nothing sets IRONCLAW_REQUIRE_DOCKER_TESTS=1, so the switch is inert) and is not fixable from here -- arming it hard-fails any lane lacking a daemon or the worker image, which needs a runner guaranteed to have both. The other half is fixable here and is fixed: docker_security.rs open-coded its own `docker version` / `image inspect` checks with three bare `return`s, so it sat entirely outside docker_gate and would have stayed fail-open even once something did set the variable. It now takes both preconditions from docker_gate::{docker_available, docker_image_available} and skips with the visible `SKIP:` line that gate's module doc requires. Measured, same machine, image absent: before, IRONCLAW_REQUIRE_DOCKER_TESTS=1 -> "skipping ..." / 1 passed after, IRONCLAW_REQUIRE_DOCKER_TESTS=1 -> panic at docker_gate.rs:74 / FAILED after, variable unset -> "SKIP: ..." / 1 passed The third line is the no-op proof: the variable is set nowhere in this tree or on main, so no lane's behavior changes today. The daemon-down path already reached the image check and skipped there, so the outcome is identical; only the branch it takes differs. Two stale comments in docker_gate.rs corrected with it (they claimed docker_security used its own gate, and that docker_image_available had no consumer), and the crate's Known debt entry now splits the done half from the #7081 half instead of describing both as open. cargo test -p ironclaw_sandbox: 193 passed, 0 failed cargo clippy -p ironclaw_sandbox --tests --all-features -- -D warnings: exit 0 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(reborn): stop calling the unwired script lane an execution lane Two review findings, both correct, both artifacts of this PR's own renames. 1. engine-v2-to-reborn-parity.md note 4 read "a native script/software execution lane (`ironclaw_sandbox`, `RuntimeKind::Script`) sandboxed via `ironclaw_sandbox`" -- self-referential after the merge collapsed ironclaw_scripts and ironclaw_process_sandbox into one crate, and it contradicts note 5 four paragraphs down ("no production execution backend is wired for it"). Re-stated as the typed runtime contract it is, citing the measurement: `with_script_runtime` has zero production callers (`rg` finds only the builder itself, docs, and 30 test call sites). 2. CHECKLIST WS10 ratchet note 2 said "raise the percentage floor ...; only the line count should fall". That generalises WS3's sandbox merge, where observed coverage happened to rise. It is wrong as guidance for WS7, and the counterexample is in this same file: the 2026-08-03 entry from #7064 records ironclaw_runner falling 85.55% -> 82.53% because the shed removed the crate's better-covered half, holding the floor, and RATCHET FAILing in the merge queue. Note 2 now says re-capture from the merged artifact, and lower only with that entry's move-not-regression counterfactual (add the moved files back, confirm the union clears the old floor, plus a zero-tests- lost name set-diff). cargo test -p ironclaw_architecture: 32 targets, 206 passed, 0 failed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): pin the WIT scope probes and the embedded-asset owner pairing Three review findings on the `wit/` move, each verified before it was acted on. 1. `ws12_workflow_contracts.py` probed `crates/ironclaw_wasm/wit/host.wit` and its nested twin. No `host.wit` exists in this repository — `git ls-files '*.wit'` returns only `tool.wit` and `channel.wit` — so both probes sat under the `crates/([^/]+/)*ironclaw_wasm/` alternative and re-asserted the crate-name term while saying nothing about the canonical ABI contracts. In a validator whose stated design is "probe derived from reality rather than from a guessed layout", a fabricated filename is a defect on its own terms. Replaced with a `crate_globs` entry, `("ironclaw_wasm", "wit/*.wit")`, which discovers the contracts on disk, requires each in scope, and synthesises the nested WS7 form — so a third contract, or the directory leaving the crate, fails the pin instead of passing on a stale name. Verified non-vacuous: narrowing the workflow alternative to `.../ironclaw_wasm/src/` now reports `tool.wit`, `channel.wit` and the nested probe as out of scope. 2. The embedded-asset routing test substituted `alpha`/`beta` owners so it could reuse the synthetic workspace. That exercised the real prefix strings through the real routing, but left the prefix->owner *pairing* — the table's entire semantic content — asserted nowhere: swapping `ironclaw_extension_support` and `ironclaw_extension_host` passed. Fixed in two halves. The routing test now drives the real `EMBEDDED_ASSET_OWNERS` against a workspace carrying the real owners' names and real manifest paths (the synthetic one could not: `build_plan` rejects a changed package outside the canonical set), asserting the real owner is selected. And the not-stale test now derives the same pairing from the tree instead of restating the constant: it resolves every literal `include_str!`/`include_bytes!` in every workspace crate through `crate_tree`, keeps the targets no crate owns — the ones that actually reach the table — and asserts that every crate compiling one of them is the routed owner or a dependent of it. That surfaced a property worth pinning: `crates/extensions/packages/` is embedded by four crates, not one. `ironclaw_extension_host`, `ironclaw_extension_manager` and `ironclaw_reborn_composition` reach into it alongside `ironclaw_extension_support`, and routing to the support crate covers them only because each depends on it. If that edge goes, a shipped artifact change stops scheduling a crate that embeds it — the silent under-schedule the table exists to prevent. Regression coverage verified red by sabotage, all three wrong tables: owners swapped (7 failures), `packages/` -> `ironclaw_llm` ("embeds nothing from it"), and the hardest case, `packages/` -> `ironclaw_reborn_composition` — a real embedder that the other embedders do not depend on ("...does not depend on..., so routing there never schedules it"). 3. CHECKLIST WS10 claimed each of the nine `wit_bindgen` guest edits forces a committed WASM artifact rebuild. Only six do: `scripts/ci/check-wasm-artifact-freshness.py` scans `crates/extensions/packages/*/wasm-src` alone, `wasm-src-digests.toml` holds exactly six entries, and `git ls-files '*.wasm'` returns exactly those six. The three `test-tools/*/wasm-src/` guests commit no artifact; the tenth site is the host's `bindings.rs`, not a guest. Corrected, and the `wit/` row now states the boundary rather than implying it. Guest paths, `wit/` contents and the six rebuilt artifacts are untouched. Verified: `test_reborn_pr_test_plan.py` 46/46, `test_ws12_workflow_contracts.py` 25/25, `ws12_workflow_contracts.py` green on the real tree, `cargo test -p ironclaw_architecture` 206/206 across 32 binaries. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(host-runtime): state the obligation visibility rule as it holds Review catch (#7090): the guardrail sentence promised "cross-owner access is `pub(super)`, never `pub(crate)`", which is stronger than the code. Verified: `RuntimeSecretInjectionStore::{insert, take, clone_material, discard_for_capability}`, `NetworkObligationPolicyStore::{insert, get, take, discard_for_capability}` and both constructors are `pub(crate)` and must stay so — `src/egress/{mod,host_port,credential}.rs` call them, and that is host-runtime composition outside `obligations/`. The rule is restated as the property that actually holds: a method whose only callers are inside `obligations/` is `pub(super)` (the three that are), and `pub(crate)` is what the stores expose to the egress pipeline they exist to serve. A future agent reading the old sentence would have read the existing `pub(crate)` methods as violations. Guidance-only; no code change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(architecture): put the operator secrets boundary entry on the right rule Review catch (#7096), and it is the serious kind: the `"ironclaw_secrets"` entry landed in `ironclaw_extension_contracts`'s forbidden vector, not `ironclaw_operator`'s. The suite still passed, because `extension_contracts` has no such dependency and `ironclaw_operator` then had no entry at all — so the guard this row exists to add was inert, and a green architecture suite was evidence of nothing. Reintroducing the edge would have passed every check. Moved to `ironclaw_operator`'s vector; `extension_contracts` restored to its `origin/main` content byte-for-byte. Negative-probed rather than assumed. With `ironclaw_secrets` temporarily re-added to `crates/ironclaw_operator/Cargo.toml`: reborn_crate_dependency_boundaries_hold ... FAILED ironclaw_operator must not have a normal dependency on ironclaw_secrets and with the manifest restored, 35/35 pass. Two further review findings, both verified before being accepted: - `ironclaw_extension_manager` **does** have a `boundary_rules()` entry (`:3543-3556`, added with WS2.4). The CHECKLIST residue note and PROPOSAL §8.2's 2026-08-02 amendment both said it had none; §8.2's sentence is stale and is marked superseded. The real gap is narrower and now stated: the rule exists and simply does not forbid `ironclaw_secrets` (#7095). - `ironclaw_product_contracts`'s guide claimed "twenty-four shipped modules". Measured: `src/lib.rs` has 26 shipped (27 `pub mod` less the gated `test_support`), and the table was missing `ironhub` **before** this branch touched it. Count corrected to twenty-six and the missing `ironhub` row added, so the inventory matches `lib.rs`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sandbox): state the Docker-gate claim as the search that checks it Review caught a false inventory in the Known debt entry, and the previous commit is what made it false: "the name appears only in docker_gate.rs and attribution_tests.rs" stopped holding the moment docker_security.rs gained a module doc naming the variable, and CLAUDE.md itself was already a third counterexample. The narrower claim is the one that was always meant and is the one that matters, so it now carries its own reproduction: no workflow, script, env file or manifest mentions the name at all -- `git grep` over *.yml/*.yaml/*.sh/ *.toml/*.py/*.json/.env* is empty here and on main -- and the sole code reference is a read, std::env::var(...) at docker_gate.rs:23. Every other occurrence is a doc comment or a panic message. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(coverage): re-anchor the exemptions the merge shifted tests/integration/changed-coverage-exemptions.toml is exact-line-keyed and auto-merges silently. #7096's additions to ironclaw_reborn_composition moved four entries' subject lines by +2 without anything flagging it; a stranded entry makes the changed-coverage validator abort with no verdict at all. Re-anchored by content (difflib line map from the #7065 tree, which the file was validated against, to the union) rather than by arithmetic: runtime.rs [4068..4073, 4082, 4083] -> [4070..4075, 4084, 4085] runtime.rs [3701] -> [3703] ; runtime.rs [3433] -> [3435] lib.rs [616] -> [618] All 142 entries / 1124 line references re-verified against the merged tree: 0 drift, 0 out-of-bounds, 0 missing paths. * refactor(layers): re-layer processes -> kernel and skills -> substrates (WS3/WS4) Two CHECKLIST rows, both of which were a one-line manifest correction rather than a code move: the family docs already placed both crates where the rows want them and only `Cargo.toml`'s `layer =` disagreed. processes -> kernel (WS3). families/kernel.md already lists ironclaw_processes among the kernel crates. The re-layer makes processes -> resources a kernel -> kernel edge, so its LAYER_MATRIX_EXCEPTION went STALE and the gate said so itself: Stale IronClaw crate layer matrix exceptions: ironclaw_processes -> ironclaw_resources from 2026-07-09 should be removed in W7: runtime process management still depends on resource contracts currently classed with kernel behavior That is the gate's verdict, not a judgement call - deleting the entry is the only way to make it pass. Baseline 5 -> 4, recomputed as len(merged list). Checked the direction both ways: all nine crates that take a normal dependency on processes (capabilities, turns, host_runtime, extension_host, loop_host, extension_manager, runner, reborn_composition, stress) are kernel or above, so the move legalizes an edge without forbidding an existing one. skills -> substrates (WS4 SS3.D). families/domains.md already lists ironclaw_skills under 'Layer(s): substrates'. Its only two normal dependencies are ironclaw_filesystem (substrates) and ironclaw_host_api (contracts), both at or below substrates, and its six consumers are all loops or above. No exception moves in either direction. cargo test -p ironclaw_architecture: 206 passed, 0 failed. cargo check --workspace --all-targets: clean. * docs(target-arch): close the WS3/WS4 rows this work satisfies, with evidence Every tick was verified against the merged tree, never against a PR title. TICKED: - sandbox lane merge: ironclaw_sandbox exists, ironclaw_scripts and ironclaw_process_sandbox absent, bollard/rcgen declared by exactly one manifest in the workspace. - mcp drops the registry dep: ironclaw_extensions is [dev-dependencies] only, 0 production ironclaw_extensions:: refs in src/. - skills -> substrates: landed here. - hooks libSQL/Postgres [decision]: ADR recorded - keep both, with the four rejected alternatives and the evidence they are already converged on one trait plus a shared conformance suite. #6945 read first as the row demands, and explicitly NOT discharged: this PR changes nothing in the dispatch path. - WS3 verify row: the row conflated Wave 3 with Wave 5 work (9 of its 10 exceptions carried removes_in = W7). Corrected with the replaced text quoted, the Wave-3 half satisfied edge by edge, and the Wave-5 remainder named with its owning field value. Ticked on the corrected condition. LEFT OPEN OR PARTIAL, each with measurements rather than a hand-wave: - first_party_tools: 1 of 6 families moved; 15 modules still in host_runtime. Ticking would be false. - processes/capabilities row: re-layer DONE; the capabilities/host.rs split is deferred with every module boundary already computed (4,560 lines, the six workflow ranges, and the arch-exempt waiver that must be deleted with it). - host_runtime binding/catalog-defaults: binding half REFUTED (moving it needs RuntimeLaneExecutor/RuntimeLaneRequest made pub, contradicting the same section's Keeps clause; zero external references to either). Catalog half cannot go to extension_host at all - host_runtime is itself a production consumer at memory_native_extension.rs:96,101, so the move is a kernel -> products edge and a Cargo cycle. Correct destination is downward. - network test_rewrite: NOT executed. Recorded the security shape (production binaries compile the seam and honour the rewrite env var at runtime) and the full 6-step plan, because the env var is how the entire E2E suite redirects vendor traffic through the production binary and the change needs feature forwarding into CI lanes I cannot verify here. cargo test -p ironclaw_architecture: 206 passed, 0 failed. * ci(coverage): recapture the two composed floors from a real measurement The provisional values were arithmetic - the sum of the two slices' recorded deltas - and the dispatch caught them, which is the whole reason the brief demanded a measurement rather than a reconciliation. Dispatch run 30907774036 at 4512e03e28f1df15b419d2e36f9f38f8f55d62fd: 26 success / 1 skipped / 2 failure, judged by per-job tally per #6978. The one skip is the pull_request-gated mutation gate; the two failures are the coverage report and the roll-up it drags down, i.e. this file doing its job. ironclaw_host_runtime: predicted 89.05% (18801 / 21114), MEASURED 88.63% (17562 / 19814). The composition was wrong by 1300 denominator lines because both slices measured their delta under the pre-#7083 aggregator, which could not see crates/extensions/** at all - lines leaving host_runtime for extension_support vanished from the tree it could measure, so neither branch's recorded delta describes the post-#7094 world. ironclaw_extension_support: MEASURED 75.31% (7142 / 9484) against #7094's 82.64% (6826 / 8260), captured before #7080's executor lines arrived. floor_percent FALLS 7.33pp and that is flagged in the file for an owner's eye rather than written quietly. Evidence it is composition and not lost tests: floor_covered_lines RISES 6826 -> 7142, so the crate is protected by more absolute lines than before, and #7080's un-masking accounting was 1398 -> 1398 with zero test names lost. Same shape as #7094's own ironclaw_runner recapture. ironclaw_sandbox passed unchanged at its arrival capture (87.09%, 3185 / 3657). The [global] entry is untouched: both moves are crate-to-crate inside the set the fixed aggregator sees. * fix(network): compile the test rewrite seam out of production builds (WS3) Closes the WS3 network row. Also RETRACTS an overstatement I made in this row's earlier annotation. CORRECTION FIRST. The earlier note claimed production binaries compile the seam and honour IRONCLAW_REBORN_TEST_HTTP_REWRITE_MAP at runtime, so anyone able to set it could redirect all credentialed vendor egress. That was WRONG. RewriteNetworkTransport::from_env_value already returned UnavailableInRelease when !cfg!(debug_assertions) (test_rewrite.rs:150), and neither [profile.release] nor [profile.dist] sets debug-assertions, so a shipped binary with the variable set REFUSES TO BOOT. It was fail-closed before this PR. I had read the ungated `mod test_rewrite;` declaration as an ungated runtime path. What was genuinely wrong, and is fixed: 1. The guard was a RUNTIME check keyed on cfg!(debug_assertions) - a profile proxy, not a build-kind guarantee. A release profile with debug-assertions turned on (normal when chasing a production bug) silently re-arms it. 2. The refusal arm had NO TEST. The one guard between a shipped binary and redirectable vendor egress was unpinned. Fix: compile-time exclusion instead of a runtime check. mod test_rewrite and its four re-exports are now cfg(any(debug_assertions, feature=test-support)), and default_host_http_egress is a compile-time pair - production builds PolicyNetworkHttpEgress<ReqwestNetworkTransport> directly, with the rewrite wrapper absent from the binary. The runtime check stays as defence in depth. E2E needs no change: those harnesses build DEBUG binaries, so they satisfy debug_assertions and keep redirecting with no feature flag and no workflow edit. The feature-forwarding-into-CI risk I flagged earlier does not arise. test-support is still forwarded composition -> network for a release-PROFILE build that needs the seam. Both halves proven rather than assumed: (a) release refuses - new regression test a_set_rewrite_map_activates_only_in_debug_and_is_refused_in_release feeds a well-formed map and asserts on profile. Under 'cargo test --release -p ironclaw_network --features test-support' it passes on the UnavailableInRelease branch; under debug 'cargo test -p ironclaw_network' it passes on the active branch. 56 passed, 0 failed. (b) production compiles without the seam - 'cargo check --release -p ironclaw_reborn_composition' (no test-support) is clean, which only compiles if the cfg(not(..)) arm is right. Also: WS0_EXTENSION_SPECIFICITY_ALLOWLIST_BASELINE 129 -> 127. The constant had drifted ABOVE the real list length; the ratchet is shrink-only so it passed silently while buying back two unearned slots. Measured off the compiler (set baseline to 0, read the reported length), identical on main and on every slice, so pre-existing drift rather than something this PR caused. cargo test -p ironclaw_architecture: 206 passed, 0 failed. cargo check --workspace --all-targets: clean. * docs(coverage): verify the extension_support floor drop is composition, independently The 82.64 -> 75.31 recapture carried a rationale that was recorded but explicitly NOT verified. Re-derived it from scratch between the two capture refs (f946a93fae -> 939af4847d) rather than inheriting the claim: - 0 test names lost in the crate (158 -> 160 test fns; both new names belong to the arriving executor). - 0 test names lost WORKSPACE-WIDE (13836 -> 13843 test fns, 13752 -> 13759 unique). This is the check that separates a relocation from a deletion: host_runtime's roster drops 156 names over the same range and every one reappears in another crate. - Exactly four files arrived, 1367 source lines, all of them the family-1 skill-install executor (src/skills/url_install.rs + url_install/{github, zip_bundle,bundle}.rs). No pre-existing file left the crate. - The arithmetic closes with the pre-existing numerator held CONSTANT: (6826+316)/(8260+1224) = 75.31% exactly, so the pre-existing code lost zero covered lines. The arriving block's own coverage is 316/1224 = 25.82%. Composition, confirmed rather than assumed. No test regression to fix; the 25.82% arrival is what earns the follow-up already recorded above the entry. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(host_runtime): collapse a duplicated obligation predicate and quiet a background warn! Three verified review findings from the #7141 round. Each was confirmed against the code before being acted on; nothing was changed on assertion alone. 1. obligations/handler.rs — `obligation_supported_before_dispatch` and `obligation_supported_after_dispatch` had BYTE-IDENTICAL 19-line bodies (verified by exact line-by-line comparison). Both were private, each called exactly once, both taking the same `phase` argument. The two names asserted a pre/post-dispatch distinction the code never implemented, while the pair gates admission of RedactOutput, EnforceOutputLimit and EnforceResourceCeiling — so editing one copy alone would have left the other stage accepting an obligation the host cannot honour (a fail-open). Collapsed to one `obligation_supported`, with the reasoning recorded so the pair is not reintroduced. 2. obligations/process_store.rs — `cleanup_terminal` is reached from `observe_process_commit` (an async background journal callback, call sites at :363/:379/:394), so its `tracing::warn!` violates the repo rule that background tasks never use info!/warn! — they corrupt the REPL/TUI display. Lowered to `debug!`; the error is still returned to the caller on the next line, so nothing is swallowed. 3. reborn_restructure_baselines.rs — the doc table said the LAYER_MATRIX_EXCEPTIONS count was "now 11". Recomputed on this ref by anchoring on the `= &[` of the value (the `&[LayerMatrixException]` type annotation opens a bracket on the same line and silently yields 0): the real count is 4, matching WS0_LAYER_MATRIX_EXCEPTION_BASELINE = 4. Corrected. Verification: cargo check --all-targets -p ironclaw_host_runtime exit 0; obligation tests 13+26 passed, 0 failed; reborn_restructure_baselines 1 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): a shipped package prompt is an asset, not prose — it was selecting no lane Review finding on #7141, confirmed empirically before acting. The Markdown prose carve-out in the planner ran BEFORE the `EMBEDDED_ASSET_OWNERS` lookup. A prompt is a `.md` file that no package *directory* owns, so a change to `crates/extensions/packages/*/prompts/**.md` took the prose arm and planned: mode=none crate_buckets=[] "crate-tree guidance changed: ..." while its sibling `manifest.toml` in the same package planned `mode=selected` onto ironclaw_extension_support + ironclaw_extension_host. Prompts are shipped production output that `ironclaw_extension_support` compiles in, and the comment above `EMBEDDED_ASSET_OWNERS` names "manifests, prompts, schemas and built wasm/*.wasm" as exactly what that table owns — so this was the "silent under-schedule of a change to production output" that comment forbids. 145 of the 149 `.md` files under `packages/` are prompts. The rule is keyed on the `prompts/` path segment, not on the asset prefixes. That distinction is load-bearing: the first attempt yielded to the asset prefixes wholesale and broke `test-tools/README.md`, which is documentation of the fixture bundles and is deliberately pinned as prose. Of the four asset kinds the table owns, only a prompt is Markdown (manifests are .toml, schemas .json, wasm .wasm), so `.md` asset <=> prompt is exact. Sabotage-tested in both directions: * `_is_package_prompt` -> False (reinstates the bug): RED, "AssertionError: 'none' != 'selected'". * `_is_package_prompt` -> any .md under an asset prefix (over-broad): RED on both the new test and the pre-existing `test_markdown_owned_by_no_crate_is_prose`, at `test-tools/README.md`. * restored: 52 passed, 51 subtests, green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(harness): refresh the latency-runner lockfile after the sandbox consolidation Review finding on #7141, reproduced before fixing. The latency harness keeps its own committed `Cargo.lock`, separate from the workspace lockfile, and the crate consolidation that replaced `ironclaw_scripts` + `ironclaw_process_sandbox` with `ironclaw_sandbox` never regenerated it. It still carried entries for both removed packages (lines 3244 and 3602) and the old host-runtime/loop-host dependency graphs. Reproduced exactly as reported: $ cargo metadata --locked --manifest-path harness/latency/runner/Cargo.toml error: cannot update the lock file ... because --locked was passed exit 101 so any reproducible invocation of the harness was broken, while the documented unlocked command silently rewrote the lockfile as a side effect of running. Regenerated with `cargo update --workspace`, which re-resolves the path dependencies. Verified after: `--locked` exits 0, the two removed packages are gone (0 entries), and `ironclaw_sandbox` is present (1 entry). Note: the re-resolve also carried three registry deps forward (wasmtime-wasi 46.0.1 -> 47.0.3, wasmtime-wasi-io likewise, wit-parser 0.251.0 -> 0.252.0). That is contained — this lockfile governs only the standalone benchmark harness and is not the workspace lockfile, and it was already unusable under `--locked` before this change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(skills): stop …
l3ocifer
pushed a commit
to l3ocifer/frick-ironclaw
that referenced
this pull request
Sep 3, 2026
…tallable, and complete (nearai#6745) * fix(reborn): inject skill bodies by default, not a one-line listing Reborn defaulted `SkillInjectionMode` to `Listing`, where a non-activated skill contributes only `- name: description` to context and its body loads only on an explicit `$name` mention or a `builtin.skill_activate` call. The intent was to save context budget. Benchmarking shows the model reads the menu and then never opens the skill. Over 30 runs with human-curated skills installed (SkillsBench/SkillLearnBench subset, `deepseek-v4-flash`, nearai/benchmarks#287): builtin.skill_list called in 30/30 runs builtin.skill_activate called in 3/30 runs a skill body actually read 0/30 runs So installed skills were effectively inert. Same 31 tasks, same skills, same model, varying only this default: no skills 78.5% curated skills, Listing 79.8% (+1.3pp -- skills bought almost nothing) curated skills, Full 85.6% (+7.1pp) For reference, harnesses that inject skill bodies unconditionally (Hermes, Claude Code) score 91.5% on these tasks with the same skills, so `Full` closes most but not all of that gap; the remainder is loop/verification behavior on a handful of multi-output tasks and is tracked separately. `Full` is already the library default in `SkillActivationSelectorConfig`; only the Reborn composition seam opted out. This restores it and adds a guard test so a revert is deliberate. `IRONCLAW_REBORN_SKILL_INJECTION=listing` still selects the previous behavior where context budget matters more than skills being used. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(skills): hot-swappable activation strategies so agent-authored skills are reusable Adds `skill.activation.v1`, a swappable-provider module in the shape of the memory-provider binding (`ironclaw_host_runtime::memory_binding`): named strategies, fail-closed resolution, behavior-preserving default, and a composition seam so nothing downstream names a concrete implementation. ## The bug it addresses `selector::score_skill` accumulates score ONLY from `activation.keywords` (+10/+5), `activation.tags` (+3) and `activation.patterns` (+20). A skill's `name` and `description` contribute nothing, and `select_skills` keeps a skill only `if score > 0`. That is fine for curated skills, which ship an `activation` block. It is fatal for skills an agent writes for itself: measured across the 31-task SkillsBench/SkillLearnBench subset in nearai/benchmarks#287, **0 of 30** agent-authored skills contained an `activation` block. Every one scored 0 and was permanently unselectable — the agent could create a skill via `builtin.skill_install` and then never reuse it, which makes self-improvement structurally impossible rather than merely weak. Claude Code has no such requirement: a skill is selectable from name and description alone. `ActivationStrategy::NameAndDescription` ports that contract. ## Design * `CriteriaOnly` (default) — today's rule, byte-identical. * `NameAndDescription` — whole-word name/description fallback, applied ONLY when the criteria pass scored 0, so a curated skill's explicit keywords always decide ordering and this can never reorder two skills that both declare metadata. `NAME_WORD_SCORE` (8) is deliberately below the selector's exact-keyword award (10). * `Disabled` — explicit mention / `skill_activate` only. * `ThirdParty { extension_id }` — production requires an admin override. Whole-word matching and a `MAX_FALLBACK_SCORE` cap keep it from over-selecting; over-selection is the failure mode that makes injecting an unrelated skill bank harmful (a whole-catalog injection took `xlsx_recover_data` 1.000 -> 0.271). ## Default stays behavior-preserving Reborn's default remains `CriteriaOnly`, opt in with `IRONCLAW_REBORN_SKILL_ACTIVATION=name_and_description`. Flipping the default changes three existing local-dev expectations (setup-marker suppression, the webui listing candidate, `skill_activate` context loading), so the strategy ships opt-in — the same discipline as the memory work, where the bundled native provider stays the default. ## Tests `cargo test -p ironclaw_skills --lib` — 239 passed, including: * `agent_authored_skill_unreachable_by_default_but_selected_under_name_strategy` — end-to-end via `prefilter_skills_with_options`: the same no-activation skill is dropped under `CriteriaOnly` and selected under `NameAndDescription`. * `name_strategy_does_not_select_an_irrelevant_skill` — no over-selection. * `name_hit_outranked_by_an_explicit_curated_keyword`, `whole_word_only_...`, `fallback_is_capped_...`, `stop_words_do_not_accumulate_score`. `cargo test -p ironclaw_first_party_extension_ports --lib` — 58 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(reborn): ship the Full skill-injection default as opt-in, not a flip The measurement in the previous commit stands: `Listing` leaves installed skills unread (`skill_list` 30/30 runs, a body actually opened 0/30) and `Full` is worth 79.8% -> 85.6% on the 31-task SkillsBench subset. But flipping the product default HANGS three existing local-dev tests, which drive a mock that expects the one-line listing candidate: * `local_dev_skill_activate_tool_loads_selected_skill_context` * `local_dev_webui_bundle_records_selectable_filesystem_skill_context` * `local_dev_runtime_wires_filesystem_skills_by_default_to_model_calls` Verified by bisect: all three hang on the previous commit alone, and pass with the default restored — the activation-strategy work is not implicated. Changing a documented product default in a way that turns CI red is a maintainer call, not something to force through, so `DEFAULT_SKILL_INJECTION_MODE` returns to `Listing` and `Full` ships as `IRONCLAW_REBORN_SKILL_INJECTION=full`. Both switches in this PR are now opt-in with the evidence attached, matching the memory-provider discipline where the bundled default is preserved. The guard test is retargeted to assert the current default, verify the opt-in path still resolves, and name the three tests that must be updated alongside a future flip. cargo test -p ironclaw_reborn_composition --lib -- skill_injection_mode \ local_dev_selector_config skill_activation # 14 passed, 0 failed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(threads): raise the result_read cap to 64 KiB, env-tunable A small per-request `result_read` cap turns one large file into a paging loop. On `manufacturing_equipment_maintenance` (nearai/benchmarks#287) reborn made 8 `read_file` calls and ZERO shell calls, hit the 24 KiB cap, then spent the whole turn paging — `result_read` at offset 24576, `handbook.pdf` at offsets 400/800/1200 — and never computed anything (`outputs_exist=0.00`). hermes, using shell to sample the same data, scored 0.522. * `TOOL_RESULT_RECORD_READ_MAX_BYTES` 24 KiB -> 64 KiB. This is the compile-time ceiling the model-observation envelope in `tool_result_reference.rs` is derived from (`* 2`, asserted at compile time), so 64 KiB here means a 128 KiB envelope — the reason not to go higher. * `TOOL_RESULT_RECORD_READ_DEFAULT_MAX_BYTES` = 64 KiB — the effective default. Enough that a typical data file or document page arrives in one read instead of a paging loop. * `IRONCLAW_TOOL_RESULT_READ_MAX_BYTES` overrides it, clamped to `[4, ceiling]`, so an override can never outgrow the envelope. Unparseable values fall back to the default rather than failing the run — a malformed tuning knob must not take down an agent. Unlike the skill-injection and skill-activation switches in this branch, this one does move the default: the paging loop is a silent capability loss rather than a behavior preference, and the knob exists for deployments that want the old size. cargo test -p ironclaw_threads --lib # 88 passed (85 existing + 3 new) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(skills): add always_available activation, Claude Code's actual contract `skill.activation.v1` gains a third binding, `always_available`: every installed skill is a candidate regardless of what it matches. This is what Claude Code and Hermes actually do. In both, a skill is a file in a directory the agent can read, so there is no gate for a correctly-installed skill to fail. Reborn's selector instead scores only `activation.keywords`/`tags`/ `patterns` and drops anything scoring 0 -- and `name_and_description` (this branch's earlier binding) only WIDENS that gate: it still needs a lexical hit, so an applicable skill phrased differently from the prompt is still discarded. The new test pins exactly that case -- a skill described as "cyclical component / growth path" against a prompt saying "hp filter" is dropped by both `criteria_only` AND `name_and_description`, and kept by `always_available`. Why it matters, measured on the 31-task SkillsBench/SkillLearnBench subset in nearai/benchmarks#287: 0 of 30 agent-authored skills contained an `activation` block, so under `criteria_only` a self-authored skill could never be selected again -- self-improvement was structurally impossible. Implementation is deliberately tiny: a `floor_score()` of 1 for this binding, applied via `.max()` in the selector's existing scoring loop. Ordering is untouched (a real keyword match still outranks a floor skill, so the context budget spends on the relevant skill first), and the existing budget -- not the score filter -- decides what is injected, which is also how Claude Code behaves. `floor_score()` is 0 for every other binding, so non-adopters are byte-identical. Default remains `criteria_only`; opt in with IRONCLAW_REBORN_SKILL_ACTIVATION=always_available. cargo test -p ironclaw_skills --lib # 241 passed (239 existing + 2 new) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(threads): drop the now-unused ceiling import Validation bounds against `contract::effective_tool_result_read_max_bytes()` (which applies the env override), so the compile-time ceiling is no longer referenced here. Removes an unused-import warning introduced by the 64 KiB cap commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * revert(threads): default result_read back to 24 KiB, keep the knob The raise to 64 KiB was never isolated: it shipped in a measurement arm alongside two other switches (skill activation, tool disclosure), so there is no evidence it changed anything. Defaulting it back keeps this crate byte-identical to pre-PR behavior. The paging trace that motivated it is real (`manufacturing_equipment_maintenance`, nearai/benchmarks#287: 8 `read_file` calls, zero shell calls, `result_read` at offset 24576, nothing computed) — but a real trace is not a measured fix, so the larger cap stays opt-in via IRONCLAW_TOOL_RESULT_READ_MAX_BYTES for whoever wants to measure it properly. The compile-time ceiling stays 64 KiB: it now bounds only how far the env override may reach, and still pins the derived model-observation envelope at 128 KiB. Net effect of this commit plus its parent: a new env knob, no default change. cargo test -p ironclaw_threads --lib # 88 passed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(skills): design for agent-authored multi-file skill bundles @henrypark133 pushed back on "move skills to the filesystem" as an overhaul that a single aggregate result did not justify. He was right, and stratifying the data shows why: the entire filesystem gain sits in skills that ship files besides SKILL.md. ships resource files (n=16): inject 81.0% -> files 94.2% (+13.2pp, CI [+0.3, +26.2]) SKILL.md-only (n=11): inject 91.5% -> files 84.7% (-6.9pp, CI [-20.4, +6.7]) So filesystem-for-everything is a REGRESSION on 13 of 31 tasks, paid to fix the other 18. The mechanism is not "models prefer filesystems": 81 of the resources are executable (you cannot run pasted Python -- citation_check scored 0.000 with the script absent, 0.833 with it present), and the text resources are too large to inline (exceltable_in_ppt would be ~262k tokens folded into SKILL.md). The design therefore keeps storage, discovery and selection exactly as they are and adds ONE extension holding the already-existing `/skills` read_write mount: skill_write_file / skill_read_file / skill_list_files. Discovery already lists from the same root that mount writes to, so nothing needs plumbing. Executing a bundled script copies that one file into `/workspace`, which the agent already mounts. Documents two things the implementation must not miss: SkillBundleDescriptor exposes only `skill_md_path`, so bundle resources are un-advertisable without skill_list_files; and `FilesystemSkillBundleRoot::user` marks bundles Trusted, so an agent that can write executable scripts there needs a distinct trust level -- the real open question. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(skills): state explicitly that creation, discovery and indexing are unchanged The crux of @henrypark133's objection. Spells out, per concern, that skill creation stays on the `skill_install` tool, discovery stays on the storage-agnostic `SkillBundleSource` trait with no new impl / trait method / descriptor change, and that there is no session-start index to migrate at all (selection is per-request; the only cache is a 5-minute TTL on catalog search). The single behavioral change remains the opt-in `always_available` selection predicate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(skills): the write tool needs an authoring prompt that asks for code skill_write_file makes multi-file skills possible; it does not elicit them. Measured: 6 of 31 tasks finished with ZERO skill_install calls despite 'Saving the skill is required', and the authoring request only ever asks for prose (method, conventions, output contract). An agent following it writes prose whether or not a write tool exists. Adds the elicitation requirement and a falsifiable success criterion: agent-authored bundles are currently 100% prose (0 of 27 ship a resource file) against 18 of 31 curated skills. If that ratio does not move once the tool ships, the bottleneck was elicitation rather than capability and the tool alone will not move scores. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(skills): let an agent install skill bundles, not just prose Agents could only ever author the PROSE half of a skill. Measured on the 31-task SkillsBench/SkillLearnBench subset (nearai/benchmarks#287): **0 of 27** agent-authored skills shipped a single file besides SKILL.md, against **18 of 31** human-curated ones (79 .py scripts, 78 .xsd schemas, 84 .md references). So every later run re-derived the method from prose and could re-make the same mistake -- lake_warming's self-authored skill described its regression procedure in prose, the next run recomputed it slightly differently and missed the grader's p<0.05 threshold. This was NOT a missing capability. `install_skill` has always taken `files: &[SkillInstallFile]`, and `parse_install_files` has always read an `input["files"]` array. Two things made it unreachable: 1. `schemas/builtin/skill_install.input.v1.json` advertised only `name`/`content`/`url` AND set `additionalProperties: false` -- so a model sending `files` was not merely uninformed, it was REJECTED. Across 112 observed skill_install calls, 111 used exactly `['content','name']`, which is what the schema permits. 2. The only encodings were `bytes_base64` and a JSON array of byte integers. A bundle file an agent writes is a script, a reference doc or a schema fragment -- all UTF-8. Making those go through base64 costs ~33% more tokens and turns one encoding slip into an InputEncode failure of the whole install. Changes: - `parse_install_files` accepts `text` (UTF-8) alongside `bytes_base64`/`bytes`. `text` takes precedence when both are given, matching the documented preference. Binary payloads are unaffected. - the schema advertises `files` with `path` + `text`/`bytes_base64`, and the description tells the model WHY to use it: put a reusable computation in a script rather than describing it in prose, and have SKILL.md name the files it relies on. That last part matters because `SkillBundleDescriptor` exposes only `skill_md_path`, so a bundle cannot advertise its own resources. - prose-only installs are untouched: no `files` key still parses to an empty vec. cargo test -p ironclaw_first_party_extensions --lib install_files_encoding # 4 passed cargo test -p ironclaw_host_runtime --test tool_surface_contract # 43 passed cargo test -p ironclaw_reborn_composition --test product_live_adapters skill_install # 1 passed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(skills): stop rejecting an install that carries both content and files `skill_install_input` gated the direct-install arm on `!object.contains_key("files")`, so `content` + `files` matched NO arm and fell through to `_ => Err(InputEncode)`. An agent attaching a script had its ENTIRE install refused. `files` was reachable only on the URL-fetch arm, which builds the array itself. This was the third of three stacked gates hiding the same capability, and the one that actually bit. With the schema fixed to advertise `files` and a `text` encoding available, the model on the 31-task SkillsBench subset (nearai/benchmarks#287) immediately sent 18 correctly-shaped `{path, text}` entries across 9 calls -- `scripts/verify_bib.py`, `references/fake_patterns.json` -- and every one was rejected here. That is the real reason 0 of 27 agent-authored skills shipped a resource file while 18 of 31 human-curated ones do: not a missing capability, and not the model failing to try. `source`/`source_url` stay excluded from the direct arm: those record provenance and are set by the URL path, so an agent must not be able to forge them. cargo test -p ironclaw_host_runtime --lib skill_install_input # 4 passed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * style: rustfmt the skill-bundle and activation changes Test modules were appended programmatically without rustfmt, which is why Formatting, Code Style and Clippy all went red on this PR. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(skills): rewrite to match what was measured, not the abandoned design The doc recommended a three-tool extension plus a resource-gate. Both are superseded: the tools turned out to be redundant (install_skill already accepted files -- three stacked gates were hiding it), and the gate MEASURED WORSE than always advertising a readable path (-25.7pp on self-creation, -40.6pp vs claude-code), because an agent-authored skill is usually SKILK.md-only so the gate suppresses the one route the selector had not already closed. Rewritten around the durable findings: the three gates and how each masked the next, the 0-of-27 vs 18-of-31 measurement, the SkillBundleDescriptor enumeration gap, and the trust question. The gate is kept in the doc as a recorded negative result, since its stratified justification is persuasive and will be proposed again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(skills): correct why always_available is not the default The previous note claimed the floor score overrides setup-marker suppression. It does not: `prefilter_skills_with_options` returns None for a satisfied marker BEFORE scoring, and the host-side filter in activation.rs already removed the candidate. What actually fails: all 32 bundled skills reach floor 1, so 3-4 unrelated ones land in plan.activations() in ActivationCriteria mode -- a mode that injects nothing under Listing. The defect exposed is that a criteria activation which injects no body is still recorded as an activation, so the count assertions stop being meaningful. Also records the sequencing against epic nearai#6565 (Slice 0 first; Slice 5's bounded-shortlist rule constrains what an unbounded floor may do) and the measured detail that under Listing a zero-scoring skill is still listed -- the model just called skill_activate in only 3 of 30 runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(skills): a floor-only skill is ranked, not activated Three defects, all surfaced by trying to make `always_available` the default. It failed 8 tests in ironclaw_reborn_composition; all 8 now pass with the flag on AND off. 1. A criteria selection that injects nothing was still recorded as an activation. Under `SkillInjectionMode::Listing` an `ActivationCriteria` entry contributes no body -- `body_eligible_bundle_ids` already ignores that mode -- so with a floor score every installed skill "activated" on every turn. Concretely: all 32 bundled skills reach floor 1, and 3 of them (6000 token budget / 2000 default per-skill cost) landed in each plan, chosen by descriptor order because the score sort is stable. `SelectionOutcome` now returns those separately as `ranked_only`, and the activation path does not iterate them. They still reach the model through the listing, which is where they belonged. 2. `AlwaysAvailable` also enabled the name/description fallback, which manufactured fake merit: a bundled skill whose description shares one word with the message scored above zero and was reported as a genuine activation. Under `AlwaysAvailable` the fallback adds no reach at all (the floor already admits everything), so it is now scoped to `NameAndDescription`, where widening the match is the entire point. This is what kept `local_dev_runtime_suppresses_explicit_setup_skill_when_workspace_marker_exists` failing after (1). 3. Raising TOOL_RESULT_RECORD_READ_MAX_BYTES to 64 KiB was NOT the no-op this PR claimed. `tool_result_reference.rs` derives MAX_MODEL_OBSERVATION_BYTES from it (* 2), so the observation envelope silently doubled 48 KiB -> 128 KiB and preview truncation changed for every caller. It broke three tests whose fixtures are sized against the envelope ("fixture must exceed the preview cap"), independently of any activation setting. The contract ceiling is back to 24 KiB and the env override is bounded by a new TOOL_RESULT_READ_ENV_CEILING_BYTES that nothing is derived from -- so the knob can raise a single read without moving anyone else's behavior. Correcting the record on an earlier comment in this PR: the failures were never the setup-marker interaction. Marker suppression returns None before scoring, so a floor score cannot revive a suppressed skill. cargo test -p ironclaw_reborn_composition --lib # 634 passed IRONCLAW_REBORN_SKILL_ACTIVATION=always_available cargo test -p ironclaw_reborn_composition --lib # 634 passed cargo test -p ironclaw_skills --lib # 241 cargo test -p ironclaw_threads --lib # 88 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * revert(skills): remove the always_available strategy, it bought nothing Verified against this branch: `AlwaysAvailable` was a no-op for everything the model can observe, and it carried a regression. Removing it rather than wiring a compensating half. Why it bought nothing. Listing membership is decided by VISIBILITY, not selection (extension_ports/activation.rs partitions candidates on body-eligibility, and everything not body-eligible still goes into the listing), so the model was ALREADY shown every visible skill before this strategy existed. The floor score never added reach -- pre-C1 its only effect was listing ORDER, and after C1 excluded floor-only skills from activations even the ordering effect was gone, because the ranking input is derived from the activation list. `SelectionOutcome::ranked_only` had no production reader at all: allocated, populated, returned, dropped. Under `Full` a floor-only skill could never be injected either, since `context_candidates_for_plan` renders only activated bundles. The regression. The floor-only bookkeeping ran for every non-merit entry BEFORE `try_select`, so under this strategy a chain-loaded companion got its own loop iteration, was recorded as floor-only, and was then partitioned OUT of `selected` -- i.e. `A requires B` activated only `A`, where `CriteriaOnly` activates both. Strictly worse than the default for any bundle with companions, and order-dependent. The comment claiming this could not happen was wrong. Also removed: ~29 "budget exhausted" notes per turn that reached `feedback` and fired a SkillActivation live-projection event with empty skill_names, because floor-only skills still ran the budget loop and `BudgetFull` continues rather than breaks. Kept: `NameAndDescription`, which has a real effect (matching on name/description, not only `activation.keywords`/`tags`/`patterns`), and the `skill.activation.v1` seam. Corrects the record in two places that argued the opposite: the runtime.rs doc comment and docs/skills/agent_authored_bundles.md. The measured reachability gap is elicitation, not filtering -- `builtin.skill_activate` was called in 3 of 30 runs and a body read in 0 of 30 -- so the next step is the listing header, not a scoring change. Note the parity numbers in nearai/benchmarks#327 never depended on this strategy: those arms ran with it off. cargo test -p ironclaw_reborn_composition --lib # 634 passed cargo test -p ironclaw_skills --lib # 240 passed cargo test -p ironclaw_threads --lib # 88 passed Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(skills): show that an agent-authored skill carries scripts This PR is what lets an agent author a skill containing a script. The Skills page could not show that it did: `skill_info` hardcoded `has_requirements: false` and `has_scripts: false`, so a scripted skill was indistinguishable from a prose-only one. The WebUI has rendered the chips since nearai#6194 and the wire fields have existed since nearai#7002 -- only the server never populated them. It was not just agent-authored skills. `portfolio`, a BUNDLED skill, ships four Python scripts (`weekly_report.py`, `backtest_strategy.py`, `concentration_warning.py`, `alert_if_health_below.py`) and has always displayed as prose-only. - `SkillSummary::has_scripts`, from one stat on the bundle's sibling `scripts` path. Absent is the common case and is not an error, so only a genuine backend failure is logged -- a skill listing must never fail because a skill has no scripts. - The bundled-summary path reads it from the embedded bundle files, so `portfolio` reports correctly there too. - `has_requirements` comes from `requires_skills`, which was already on the summary. Verified on a live production server: 33 skills listed, `portfolio` the one reporting `has_scripts`, six reporting `has_requirements`. Note: skill scripts still cannot EXECUTE under hosted multi-tenant -- `HostedMultiTenant` + `SecureDefault` resolves to `ProcessBackendKind::None`, which strips `builtin.shell`. That is deliberate pending the tenant sandbox. So the chip tells a multi-tenant user their skill has scripts the agent can read but not run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(skills): pin the direct-install input contract at the capability boundary Four cases covering what a caller may and may not put in a `builtin.skill_install` input, asserted through runtime dispatch rather than against whichever helper currently normalizes the input. The normalizer has already moved once (host runtime -> ironclaw_extension_support, WS3) and is about to be merged across that move again. Written so the same four pass on both sides: a resolution that quietly re-tightens the inline arm fails the first one instead of silently dropping the capability this PR adds. The dividing line these pin is provenance, not shape: - `content` + `files` installs, and the script lands on disk verbatim - `bytes_base64` works on the direct arm too, not only the rewritten URL payload - `content` + `files` + `source`/`source_url` is still refused whole - a `../..` bundle path is refused and writes nothing outside the skill directory Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(skills): repair the post-merge build and close the review findings Build breakage from the merge: main added two `SkillSummary` construction sites (`skill_learning`, `lifecycle_product_service`) that this branch's new `has_scripts` field left incomplete. Both are fixtures whose assertions do not depend on it, so both set `false` with a note saying why. CI, all under `-D warnings`: - Four `result_read` cap items in `ironclaw_threads::contract` were `pub` in a private module (`unreachable_pub`). Nothing outside the crate reads them, so they are `pub(crate)`. - Four constant-value asserts (two in the same contract module, two in `activation_strategy`) move into `const {}` blocks, which is what they always meant: they are compile-time invariants, not runtime checks. Review findings: - The stale-default doc note (coderabbit, ironloop) was against `ecabcb5fe`, before `4951d76bb` reverted the flip. Docs and `DEFAULT_SKILL_INJECTION_MODE` both say `Listing` with `full` as the opt-in, so there is nothing left to correct. - The unset-env branch is now reachable from a test (coderabbit). `skill_injection_mode_from_env_value` takes the lookup's `Result`, so the product default can be asserted without `remove_var` racing every other test in this binary. Covered along with `full`, trimming/case, empty, unrecognized, and non-unicode. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(skills): state accurately why the inline arm now accepts a bundle The comment claimed a review decision had been "reversed on evidence", which overstates what happened and would read as overriding the team. What actually happened: the refusal predates nearai#7141 entirely -- it lived in the host-runtime copy of this resolver, and nearai#7141 carried it across the move to this crate verbatim, declining a reviewer's suggestion to relax it there. That was the right call for a move-only refactor. This PR is where the behavior change belongs, and it is made deliberately with the measurement attached. Comment only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: raise the composition mass ceiling for this layer's assembly wiring CI's "Check composition mass budget" step reds this branch by 62 LOC. Not because this branch is large: main sits only 54 LOC under the effective ceiling (40,499 + 150 tolerance against 40,595 observed), so the gate currently trips on any PR adding more than that to composition, and this one adds skill-summary and product-surface assembly. Raised to the measured 40,711 in both places the gate pairs — `[gate].loc_ceiling` in the manifest and `COMPOSITION_ABSOLUTE_SRC_LOC` in `reborn_restructure_baselines.rs`, since a second ratchet fails when they disagree, which is how it enforces recording the change in the PR that causes it. Measured with `check-composition-budget.sh --print`, set to current rather than padded, per the manifest's own protocol. A raise is a reviewed decision by that file's rules, not routine wiring, so it is flagged here and in the PR body rather than left in a diff. The next wave close should re-ratchet. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(threads): make the result_read env knob actually raise the cap serrrfirat, High: `IRONCLAW_TOOL_RESULT_READ_MAX_BYTES` was inert. It widened `validate_tool_result_record_read`, which sits DOWNSTREAM, while the caller-facing gate in `result_read.rs` stayed pinned to the compile-time `TOOL_RESULT_RECORD_READ_MAX_BYTES` (24 KiB) and the advertised schema still said `maximum: 24576`. A larger read was rejected before it could reach the widened validator, so setting the variable changed nothing. The gate and the schema now resolve `effective_tool_result_read_max_bytes()` per request. This also corrects a fix I made earlier in this PR for the wrong reason. Clippy flagged `effective_tool_result_read_max_bytes` as `unreachable_pub` and I narrowed it to `pub(crate)` — but it had no cross-crate caller precisely BECAUSE the wiring was missing. The lint was reporting the bug, not dead code. It is `pub` again, with the caller it was always supposed to have. Tested as a wiring identity (gate == effective cap, schema == gate) rather than by setting the env var: these tests run in-process and in parallel, so mutating process environment races every other test reading it, and the identity is exactly what regressed. Verified: workspace `cargo check --all-targets` clean, 12 test binaries green across loop_host and threads, fmt clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(threads): satisfy the production-target lint lane CI's PR lane lints the DEFAULT target set (lib + bins, no tests or examples) with \`--all-features\`; I had been running \`--all --tests --examples\`, and the extra targets masked both of these. - \`TOOL_RESULT_READ_ENV_CEILING_BYTES\` was widened to \`pub\` alongside \`effective_tool_result_read_max_bytes\` in the previous commit, but only the function is re-exported from \`lib.rs\`, so the constant was unreachable-pub. Only that function reads it, so it is \`pub(crate)\`. - \`result_read.rs\` no longer reads \`TOOL_RESULT_RECORD_READ_MAX_BYTES\` now that the gate resolves the effective cap, so the import goes. Verified with the lane CI actually runs (\`cargo clippy --workspace --all-features -- -D warnings\`, no test targets), plus fmt and the loop_host/threads suites. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(skills): cut the comment bloat on this layer Comment share of this PR's diff: 33% -> 26%, 471 -> 330 lines. The measurements that justify a constant stay; the retellings of how we got there go, since they are already in the commit history. One of these was a duplicate rather than verbosity: `DEFAULT_SKILL_ACTIVATION` carried an earlier draft stacked directly above its own replacement, so the file gave two competing accounts of the same default. I had removed that copy at the top of the stack only; removing it here means all three layers carry one version instead of conflicting on every merge. Also corrects a doc that contradicted the code, which Copilot flagged: `effective_tool_result_read_max_bytes` was documented as clamping to `TOOL_RESULT_RECORD_READ_MAX_BYTES` when it clamps to `TOOL_RESULT_READ_ENV_CEILING_BYTES` — the whole point of the separate ceiling. Comments only; no code touched. `--all-features` clippy on the production target set (the lane CI runs) clean, fmt clean, 19 test binaries green across skills / threads / loop_host / extension_support. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(skills): move the design doc under the fenced internal tree `docs: enforce the docs/ publication boundary` landed on main at 00:53 and dequeued this PR: every file under `docs/` must be either published (referenced from `docs.json` navigation) or fenced (by `docs/.mintignore`), and `docs/skills/agent_authored_bundles.md` was neither — so it would have been deployed to the public docs site and indexed. That is exactly what the gate exists to catch, and the doc predates the rule rather than breaking it. `.mintignore` is frozen by that same change ("all new internal material goes under internal/"), so the fix is the move, not a new fence entry. `docs/internal/` is already fenced. `scripts/ci/test_docs_publication_boundary.py`: 21/21 pass. Workspace `--all-features` clippy on the production target set clean, fmt clean, panic gate clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Supersedes #7065, #7080, #7084, #7090, #7096 — one PR at the owner's request, plus the Wave 3/4 rows none of them covered. The five source PRs are left open; closing them is the owner's call.
Merged, never rebased and never squashed: rebase has already silently reverted a fail-closed scanner in this program, and squash would destroy the rename detection that #7084's 100%-rename
wit/moves and #7065's crate merge depend on.Merge order, and why
origin/main→ #7065 → #7084 → #7080 → #7096 → #7090, thenorigin/mainmerged down as it moved.#7065 is the base: 138 files, and it touches 12 of the 21 files that more than one branch touches. #7084 goes second because it is the only branch carrying 100%-rename moves and six rebuilt
.wasmbinaries, so it lands against the least accumulated history. #7080 follows because itscoverage-floor.tomlandfirst_party_tools/edits compose directly with #7065's. #7096 and #7090 are the least entangled and go last.Union verification — 0 unexplained
247 (branch, file) pairs across the five branches:
Strengthened beyond a category check: for all 54 differing files I diffed every line each branch added against the union. Only 6 files lose any added line, all of them files I deliberately reconciled, and every lost line is accounted for below. No test function, no assertion and no production line was lost anywhere else.
Numbers, not just files — this is where the cross-slice failures were
LAYER_MATRIX_EXCEPTIONS: 6 to 4, recomputed three timesNever inherited from a slice. Counted in Python between
const LAYER_MATRIX_EXCEPTIONSand its closing];, anchored on the= &[of the value — the&[LayerMatrixException]type annotation also opens a bracket, and a naive scan silently returns 0 entries.LayerMatrixException {also appears 5 more times file-wide (the struct definition plus 4 test fixtures), which is the trap the count must avoid.The arithmetic changed twice under me and no slice's stored figure was correct:
ironclaw_extensionsto substrates, taking main to 6 and deleting two of refactor(sandbox,contracts): merge the sandbox lane and flip mcp onto contracts (WS3) #7065's three removals before this PR touches them. refactor(sandbox,contracts): merge the sandbox lane and flip mcp onto contracts (WS3) #7065's remaining effect is now only thescripts → resources⇒sandbox → resourcesrename (one out, one in, net zero); refactor(extensions): move the skill-install executor to extension_support (WS3) #7080'shost_runtime → skillsis the single net drop. Union = 5.processes→ kernel re-layer then madeprocesses → resourceslegal, so the gate reported it stale and it is deleted. Final = 4.Verified on the pushed ref: in-array count ==
WS0_LAYER_MATRIX_EXCEPTION_BASELINE== 4. Survivors:host_runtime → extension_support(W7),conversations → turns(WS5),mcp → resourcesandsandbox → resources(both #7067).tests/integration/coverage-floor.toml— the one number I could not measure#7065 and #7080 both rewrite
ironclaw_host_runtimefrom the same base (88.23 / 20538 / 23277), and their moves compose:sandbox_process/**→ironclaw_sandboxextension_supportNeither stored triple is correct for the union. The value in this PR is that composition — arithmetic, not measurement — and it is flagged as such in the file. Note the failure mode this file exists to catch, which #7080's own dispatch run demonstrated: its percentage went UP while the absolute count fell through the floor. Watch
floor_covered_linesseparately from the percentage.#7094 added a third recapture obligation. It fixed #7083, so
crates/extensions/**is no longer invisible to the aggregator, and it flooredironclaw_extension_support(82.64 / 6826 / 8260) — from a measurement taken before #7080's executor lines arrive there. #7080's recorded "the destination cannot be floored" is superseded. The[global]entry is left exactly as #7094 captured it: both moves are crate-to-crate inside the set the fixed aggregator can see, so lines leave one visible denominator and enter another.changed-coverage-exemptions.toml— a silent auto-merge that had really driftedExact-line-keyed, and it auto-merged with no conflict. #7096's additions to
ironclaw_reborn_compositionhad moved four entries' subject lines by +2 with nothing flagging it; a stranded entry makes the validator abort with no verdict at all. Re-anchored by content (a difflib line map from the tree the file was validated against), not by arithmetic:All 142 entries / 1124 line references then re-verified against the merged tree — and again after each
origin/mainmerge — resolving to byte-identical source text: 0 drift, 0 out-of-bounds, 0 missing paths.reborn_extension_specificity.rsRead off the compiler rather than counted by eye (temporarily set the baseline to 0 and let the ratchet report it): the merged
ALLOWLISTis 127. Identical on main, #7065, #7080 and the union — this PR changes it by zero. Finding, pre-existing and not introduced here: the baseline constant says 129 while the list is 127; the ratchet is shrink-only, so it passes silently. Left alone as out of scope, flagged for a follow-up.PATH_TERM_COLLISIONSis the array that actually moved, and its union is exact: 82 − 2 = 80, with #7080's twoskill_url_installremovals gone, #7065's twosandbox_processpath renames present, and both survivingfirst_party_toolssiblings kept.scripts/ci/reborn_pr_test_plan.py— the union, with #7084's routing intact#7065 and #7084 both rewrote the same arm. The union takes #7084's
CRATE_OR_ASSET_PREFIXES(sotest-tools/**is classified) and #7065's crate-tree-prose carve-out, and orders the unmapped branch soEMBEDDED_ASSET_OWNERSis consulted BEFORE #7065's widening — otherwise a shipped-artifact change would widen tofullinstead of scheduling the crate that compiles it, degrading #7084's precise routing.crates/extensions/packages/**→ironclaw_extension_supportandtest-tools/**→ironclaw_extension_hostboth verified present and exercised.Test roster reconciles exactly: 43 (main) − 1 (the
fails_fasttest #7065 deliberately rewrote) + 5 + 4 = 51, all passing.Two cross-slice fixture collisions, both genuine, both reconciled rather than silenced:
Dockerfileis unclassified; refactor(wasm): move wit/ inside its owning crate (Wave 3) #7084 classified it. Moved toMakefile— the same fix refactor(wasm): move wit/ inside its owning crate (Wave 3) #7084 already applied to its owntest_unclassified_build_input_fails_fast— so the fail-closed arm stays exercised.raisesunmapped crate path; refactor(sandbox,contracts): merge the sandbox lane and flip mcp onto contracts (WS3) #7065 replaced that fallback with the exhaustive plan. Re-pinned asmode == "full"— a superset of any narrowing, so the property (an unattributable asset path can never under-schedule) is preserved, and it can never resolve tonone.Append-only ledgers
CHECKLIST.md/PROPOSAL.md/PLAN.mdkeep every side's dated amendments; all five slices' text verified present after every merge. Where a slice's amendment was invalidated by a sibling I corrected it rather than leaving it: #7065's verify row still listedhost_runtime → skillsas open (#7080 closes it), #7080's rows said10 → 9, #7090's said "the register is 10 before and after".What each slice contributes (evidence preserved)
mcpcontracts flip. The docker-gate fail-closed fix; the "no production behavior" claim refuted with three live call paths; behaviour argued at the diff (26 moved files: 11 byte-identical, 9 differing by one import line, total +63/−36); and the script-lane surface scan that would have gone silently vacuous, fixed to walk the whole crate tree with fatal reads and a non-vacuity assertion.wit/inside its owning crate.EMBEDDED_ASSET_OWNERSplus its four planner tests and the six genuinely rebuilt.wasmartifacts; theinclude_str!measurement that would have taken §11.2.7 from 19 to 21 cross-crate reach-ins while ticking a box that says the scan passes.extension_support. Its floor rationale (percentage up, absolute count down) is what made the composed recapture necessary; plus the executor/adapter seam rule.product_contracts::operator_secretsport and theironclaw_operatorboundary rule, kept in the merged forbidden list beside refactor(sandbox,contracts): merge the sandbox lane and flip mcp onto contracts (WS3) #7065'sscripts→sandboxrename.arch-exempt: large_filewaiver, so re-fusing the owners re-trips the gate.Beyond the consolidation — Wave 3/4 rows
Landed here:
processes→ kernel (WS3). A one-line manifest correction:families/kernel.mdalready listed the crate as kernel. It madeprocesses → resourcesa kernel→kernel edge, so the gate itself reported the exception stale — deleting it is the gate's verdict, not a judgement call. All nine crates depending onprocessesare kernel or above, so no existing edge becomes illegal.skills→ substrates (WS4 §3.D). Same shape;families/domains.mdalready said substrates. Verified both directions: deps arefilesystem(substrates) andhost_api(contracts); all six consumers are loops or above.hookslibSQL/Postgres[decision]— ADR: keep both. ironclaw_hooks: cross-run hook-isolation semantic has no regression test (doc previously cited nonexistent tests) #6945 read first, as the row demands. They are already converged on the seam that matters: onePredicateStateBackendtrait with three real impls, one sharedpredicate_state::contractsuite behindtest-support, and per-backend conformance plus parity suites. Both are live deployment shapes chosen by profile, so convergence would delete a shipped shape. Four rejected alternatives recorded. ironclaw_hooks: cross-run hook-isolation semantic has no regression test (doc previously cited nonexistent tests) #6945 is explicitly not discharged — this PR changes nothing in the dispatch path.removes_in = "W7"). Corrected with the replaced text quoted, satisfied edge by edge, and ticked on the corrected condition. Thebollard/rcgenclause clears stronger than asked: exactly one declaring manifest workspace-wide.Honestly not closed — each annotated in CHECKLIST with measurements rather than a hand-wave:
first_party_tools(1 of 6 families), thecapabilities/host.rscharter split,host_runtimebinding/catalog-defaults (binding half refuted; catalog half is a Cargo cycle as worded), andnetwork'stest_rewritegating.Verification
cargo check --workspace --all-targets— clean, 0 errors, 0 warnings.cargo test -p ironclaw_architecture— 206 passed, 0 failed, re-run after every step.python3 scripts/ci/test_reborn_pr_test_plan.py— 51/51.cargo fmt --allbefore every commit.workflow_dispatchofreborn-tests.ymlis required for a coverage verdict —_full_plan()is structurally unreachable whenevent == "pull_request"(Changed-coverage gate does not run on ordinary PRs — first verdict now lands in the merge queue #7036), so this PR's green checks carry no coverage information. Theironclaw_host_runtimeandironclaw_extension_supportfloors in this PR are arithmetic, not measurement, and must be replaced from that run's per-crate table before merge. Judge that run by the per-job tally, never the roll-up (reborn-tests.yml: workflow_dispatch runs structurally fail the Tests (Reborn) roll-up (critical-mutation skipped but disallowed) #6978).🤖 Generated with Claude Code