Repository navigation
chore: update WASM artifact checksums and version-pinned URLs - #3
Open
github-actions[bot] wants to merge 14 commits into
Open
github-actions[bot] wants to merge 14 commits into
github-actions[bot] wants to merge 14 commits into
Conversation
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…earai#3045) (nearai#3243) * feat(host-api): add runtime policy vocabulary (PR 1 of nearai#3045) First PR in the runtime-presets stack. Adds the shared contract vocabulary the resolver, host runtime planner, settings/blueprint surfaces, and audit log will consume. No resolver logic, no behavior changes — vocabulary only. New `ironclaw_host_api::runtime_policy` module: - `DeploymentMode`: where IronClaw is running and who owns the machine boundary (LocalSingleUser / HostedMultiTenant / EnterpriseDedicated). - `RuntimeProfile`: the operator/user-selected preset (12 variants spanning SecureDefault, Local{Safe,Dev,Yolo}, Hosted{Safe,Dev, YoloTenantScoped}, Enterprise{Safe,Dev,YoloDedicated}, Sandboxed, Experiment). Family predicates `is_local`/`is_hosted`/ `is_enterprise`/`is_yolo` are partition-checked in tests. - `EffectiveRuntimePolicy`: aggregate of resolved backend + mode choices, plus both requested and resolved profile so audit can render "you asked for X, you got Y" when policy reduced authority. `was_reduced()` predicate flags the narrowing case. - Backend/mode enums consumed by `EffectiveRuntimePolicy`: `FilesystemBackendKind`, `ProcessBackendKind`, `NetworkMode`, `SecretMode`, `ApprovalPolicy`, `AuditMode`. All snake_case on the wire; `as_str()` matches the serde wire name. Module documentation explains the boundary against `RuntimeKind` (execution lane, not authority — the issue forbids `RuntimeKind::Local`) and `TrustClass` (per-invocation authority ceiling, composes with runtime policy rather than replacing it). The resolver itself, settings/CLI selection, capability surface filter, and host runtime planner integration land in subsequent PRs (PR 2-7 of nearai#3045). Test plan: - [x] `cargo test -p ironclaw_host_api` — 38/38 (31 existing + 7 new vocabulary tests covering family predicates, yolo predicate, serde round-trips, `was_reduced` flag, and `as_str` ↔ wire-name consistency). - [x] `cargo clippy --workspace --all-features --tests` — zero warnings. - [x] `cargo fmt --all -- --check` — clean. - [x] `cargo test -p ironclaw_architecture reborn_crate_dependency_boundaries_hold` — pass. Closes part of nearai#3045 (PR 1 of 8). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(runtime-policy): add resolver crate (PR 2 of nearai#3045) New `ironclaw_runtime_policy` crate. Pure-logic resolver that turns the operator's request — `(DeploymentMode, RuntimeProfile, OrgPolicy)` — into an `EffectiveRuntimePolicy` consumed by the host runtime planner. Safety invariants enforced: - **Monotonic**: deployment mode and tenant/org policy may *reduce* the requested profile's authority; they may never *increase* it. Resolved profile is `min(requested, ceiling)` within the same family; ceilings with headroom keep the request, ceilings that narrow produce the ceiling. - **Fail-closed by default**: invalid `(deployment, profile)` pairs are typed errors, not silent downgrades. Hosted multi-tenant rejects every Local* profile; enterprise rejects every Local*/ Hosted* profile; etc. The `(deployment, profile)` compatibility matrix is a single readable `match`. - **Yolo opt-in**: any `*Yolo*` profile requires `ResolveRequest::yolo_disclosure_acknowledged = true` — the resolver never sets this itself; CLI/settings/blueprint must capture explicit operator confirmation. `EnterpriseYoloDedicated` additionally requires `OrgPolicy::admin_approves_dedicated_yolo`. - **Hosted boundary**: hosted multi-tenant resolution never produces `FilesystemBackendKind::HostWorkspace` or `ProcessBackendKind::LocalHost`. A regression test enumerates every hosted profile (including Sandboxed and the disclosure-acknowledged yolo variant) and asserts the property. Per-profile backend mapping is centralised in `backends_for(deployment, profile)` so the matrix is reviewable in one place. `Sandboxed`/ `Experiment` reuse the deployment-appropriate workspace backend (ScopedVirtual/TenantWorkspace/OrgDedicatedWorkspace) so they can run under any deployment without leaking provider-host paths. Output is deterministic and round-trips through serde so audit logs can record the exact policy that gated an invocation. `EffectiveRuntimePolicy::was_reduced` flags the narrowing case. Test plan: - [x] `cargo test -p ironclaw_runtime_policy` — 17/17 covering compatibility matrix, yolo disclosure, org admin approval, ceiling narrowing within family, ceiling-with-headroom preserving request, family-mismatch and deployment-agnostic ceiling rejections, hosted multi-tenant boundary property, determinism, serde round-trip, and the full valid-pairs matrix. - [x] `cargo test -p ironclaw_architecture reborn_crate_dependency_boundaries_hold` — pass (new crate depends only on `ironclaw_host_api`). - [x] `cargo clippy -p ironclaw_runtime_policy --all-features --tests` — zero warnings. - [x] `cargo fmt --all -- --check` — clean. Builds on PR 1 (nearai#3243) — runtime policy vocabulary in `ironclaw_host_api`. Settings/CLI selection (PR 3), capability surface filter (PR 4), and host runtime planner integration (PR 5) consume this resolver. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(config): runtime profile selection from CLI + env (PR 3 of nearai#3045) Adds the configuration surface for nearai#3045's runtime profile system. Operators can now select deployment mode + runtime profile via: - CLI flags: `--deployment-mode`, `--runtime-profile`, `--yolo-disclosure` (all global). - Environment vars: `IRONCLAW_DEPLOYMENT_MODE`, `IRONCLAW_RUNTIME_PROFILE`, `IRONCLAW_YOLO_DISCLOSURE`. - Defaults: `LocalSingleUser` + `SecureDefault` — the safest combination, never grants provider-host authority. Selection precedence: CLI > env > default. The DB-backed setting layer is reserved for a follow-up (the existing settings store has its own coexistence story with `IRONCLAW_PROFILE` TOML overlay that deserves its own change). Implementation: - `crates/ironclaw_host_api/src/runtime_policy.rs`: add `FromStr` impls for `DeploymentMode` and `RuntimeProfile` matching their snake_case wire names, plus `ParseRuntimePolicyEnumError`. Tests round-trip every variant against `as_str` to lock in the identity contract. - `src/config/runtime.rs`: new module with `RuntimeConfig` (resolved) and `RuntimeConfigOverrides` (raw CLI inputs). `resolve_from` layers CLI > env > default, calls `ironclaw_runtime_policy::resolve`, and surfaces resolver failures as typed `ConfigError::InvalidValue`. `safe_default()` is the test/`for_testing` constructor and never fails. - `src/config/mod.rs`: add `runtime: RuntimeConfig` to `Config`. The production `build` path resolves from env only; the new `Config::with_runtime_overrides` re-resolves with CLI overrides layered on top after `from_env*` returns. `for_testing` uses `RuntimeConfig::safe_default()`. - `src/cli/mod.rs`: three new global args. The host_api enums' `FromStr` impls let clap derive parse them automatically. - `src/main.rs`: wire CLI overrides into `Config::from_env_with_toml` via `.and_then(|c| c.with_runtime_overrides(&overrides))`. Legacy env vars (`ALLOW_LOCAL_TOOLS`, `SANDBOX_POLICY`, `SANDBOX_ALLOW_FULL_ACCESS`) are intentionally untouched — the planner integration in PR 5 is the right place to reconcile them against the resolved policy. This PR adds the surface; nothing consumes the resolved policy yet. Test plan: - [x] `cargo test --lib config::runtime` — 8/8 covering: defaults, CLI > env precedence, env-driven resolution, yolo disclosure requirement, hosted-multi-tenant + Local* fail-closed, invalid env value typed error, `safe_default` construction. - [x] `cargo test -p ironclaw_host_api` — 39/39 (31 existing + 8 in runtime_policy mod) including the new `FromStr` round-trip. - [x] `cargo clippy --workspace --all-features --tests` — zero warnings. - [x] `cargo fmt --all -- --check` — clean. Stacks on top of PR 1 (vocabulary, nearai#3243) and PR 2 (resolver, `ironclaw_runtime_policy`). PR 4 (capability surface filter) and PR 5 (host runtime planner) are the consumers. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(tools): visible capability surface filter (PR 4 of nearai#3045) Adds the visible-capability-surface filter from nearai#3045's "Visible capability surface" rules. Profile-impossible affordances are hidden from the model-facing tool list before the model call; everything else stays visible and continues to fail structurally at action time if authorization/approval/resource checks reject it. This PR adds the filter mechanism + the registry seam, plus one tool (`shell`) declaring its affordance as a worked example. Bulk migration of other tools to declare their affordances is a deliberate follow-up so each migration gets a focused review. Implementation: - `src/tools/tool.rs`: new `ToolRuntimeAffordance` enum (`None`, `AnyProcess`, `LocalShell`, `HostFilesystem`, `DirectNetwork`) and default `Tool::runtime_affordance() -> ToolRuntimeAffordance::None`. - `src/tools/runtime_filter.rs`: `is_visible_under(policy, affordance) -> bool` with the per-variant matching documented on the affordance enum. Six unit tests pin the matrix: `None`-affordance always visible (even under `process_backend == None`), `AnyProcess` hidden only under `None`, `LocalShell` visible only under `ProcessBackendKind::LocalHost`, `HostFilesystem` visible only under `FilesystemBackendKind::HostWorkspace`, `DirectNetwork` visible under `NetworkMode::{Direct, DirectLogged}`. - `src/tools/registry.rs`: new `ToolRegistry::list_visible_under` and `all_visible_under` methods, gated by both the existing engine-version filter and the new affordance filter. Existing `list` / `all` are unchanged. - `src/tools/builtin/shell.rs`: `runtime_affordance` returns `AnyProcess`. Profiles that resolve to `process_backend == None` (e.g. `SecureDefault`) now hide shell from the model entirely. Tenant- and org-dedicated process backends satisfy the affordance — shell runs inside the matching sandbox, not on the provider host. This is **visibility, not authorization** — per-invocation authorization (capability grants, approvals, resource checks) still runs on every call regardless of visibility. Reviewer guardrail in `crates/ironclaw_host_runtime/src/lib.rs` explicitly punts the caller-authority filtering decision to upper layers; this PR places it at the tool registry, the closest "upper layer" boundary today. Test plan: - [x] `cargo test --lib tools::runtime_filter` — 6/6 covering the five affordance variants × the relevant backend/mode axes. - [x] `cargo clippy --workspace --all-features --tests` — zero warnings. - [x] `cargo fmt --all -- --check` — clean. Stacks on PR 1 (vocabulary, nearai#3243), PR 2 (resolver), and PR 3 (settings/CLI). PR 5 will wire `list_visible_under` into the agent loop's tool projection so the filter has end-to-end effect. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(host-api): use exhaustive match in RuntimeProfile family predicates Address gemini-code-assist review on PR nearai#3243 (4 inline comments, all the same shape on lines 197 / 205 / 213 / 225 of runtime_policy.rs). The four `is_local` / `is_hosted` / `is_enterprise` / `is_yolo` predicates used `matches!()` with a single arm, which silently defaults a new variant to the negative case. These predicates gate security-critical deployment boundary checks — a new `RuntimeProfile` variant added without explicit categorization here would not be flagged by the compiler and could land on the wrong side of the hosted/local/enterprise/yolo axis. Replace each `matches!()` with an exhaustive `match` that names every variant. A new variant now produces a compile error until the author makes a deliberate yes/no decision in each of the four predicates. Verified: 17/17 resolver tests + 8/8 runtime_policy unit tests + 31/31 host_api crate tests still pass; workspace clippy and fmt clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(config): annotate `safe_default` expect with safety comment The `No panics in production code` CI check rejected the `.expect()` in `RuntimeConfig::safe_default()`. The expect is load-bearing-safe: `(LocalSingleUser, SecureDefault, default OrgPolicy, no yolo disclosure)` is structurally guaranteed to resolve — `SecureDefault` is deployment-agnostic in the resolver's compatibility matrix, isn't a yolo profile, and the empty `OrgPolicy` never narrows. The `every_valid_deployment_profile_pair_resolves` test in `ironclaw_runtime_policy::resolver::tests` locks this in. Annotate the expect with the inline `// safety: ...` comment the no-panics script recognizes, and explain the invariant in a preceding doc-style comment for human readers. Verified locally: `python3 scripts/check_no_panics.py --base origin/reborn-integration --head HEAD` reports `OK: No panic-inducing calls in changed production code.` Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(runtime-policy): planner + visibility wire-up + zmanian fixes (PR 5/6/7 of nearai#3045) Closes the gap between the resolver-level policy substrate (PRs 1-4 already in this branch) and the model-facing tool list. After this commit, PRs 5, 6, and 7 of the nearai#3045 decomposition are implemented in this PR; PR 8 (blueprint/harness integration with nearai#3036) is explicitly excluded. PR 5 — Host runtime planner integration - New `ironclaw_host_runtime::planner` module: pure function `plan_capability(&CapabilityDescriptor, &EffectiveRuntimePolicy) -> Result<ExecutionPlan, PlannerError>`. The plan names the concrete filesystem/process/network/secret backend kinds the host runtime will dispatch against. Planner fails closed when a capability declares effects (`SpawnProcess`, `Network`, `UseSecret`) that the resolved policy disables. - Public re-exports on `ironclaw_host_runtime`: `ExecutionPlan`, `PlannerError`, `plan_capability`. PR 6 — Local profile vertical slice - Integration tests in `crates/ironclaw_host_runtime/tests/runtime_policy_planner_contract.rs` drive `LocalSingleUser + LocalDev` through resolver → planner and assert HostWorkspace + LocalHost selection for the canonical filesystem.read / shell.cargo_test / shell.npm_test / shell.ripgrep / shell.git_status coding aliases. - `LocalSafe` approval-preset assertion (`AskWrites`) locks in the resolver's mapping for the cautious local mode. PR 7 — Hosted/enterprise enforcement + regression tests - Hosted regressions: `HostedDev` + shell.run plans against `TenantSandbox`, never `LocalHost`. Filesystem write plans against `TenantWorkspace`, never `HostWorkspace`. `HostedYoloTenantScoped` with disclosure ack still cannot reach LocalHost / HostWorkspace. - Enterprise regressions: `EnterpriseDev` plans against `OrgDedicatedRunner`. `EnterpriseYoloDedicated` requires both `EnterpriseDedicated` deployment AND `org_policy.admin_approves_dedicated_yolo = true` — without admin approval the resolver fails closed. - `Experiment + package install` resolves to `SmolVm` / `Allowlist`. Wiring the visibility filter into the model-facing tool list - New `ToolRegistry::tool_definitions_visible_under(policy)` mirrors `tool_definitions_for_engine` but additionally filters by `runtime_filter::is_visible_under(policy, tool.runtime_affordance())`. - `AgentDeps` now carries `runtime_policy: Option<EffectiveRuntimePolicy>`. `Some` in production (`Config::runtime.effective_policy.clone()`); `None` in tests. - `dispatcher.rs` routes the chat turn's tool list through `tool_definitions_visible_under` when a policy is present, otherwise through the legacy unfiltered path. - New integration test `tests/runtime_policy_tool_visibility_integration.rs` proves the chain Config → resolver → tool filter end-to-end (zmanian gap #3 — binding test that `is_visible_under` is actually called from the model-facing path). Includes pipeline test for `RuntimeConfig::resolve_from(overrides)` so the production wire matches. zmanian review address - `#[non_exhaustive]` on all eight wire-stable enums (`DeploymentMode`, `RuntimeProfile`, `FilesystemBackendKind`, `ProcessBackendKind`, `NetworkMode`, `SecretMode`, `ApprovalPolicy`, `AuditMode`). Resolver match arms updated with fail-loud wildcard panics so a forgotten variant fails in development rather than silently defaulting. - Tightened `was_reduced()` doc to "tenant/org policy ceiling reduced authority within the same family" — deployment-mode reduction is impossible by construction (resolver returns `IncompatibleDeployment`). - Renamed struct `OrgPolicy` → `OrgPolicyConstraints` to disambiguate from the wire-stable enum variants `ApprovalPolicy::OrgPolicy` / `AuditMode::OrgPolicy`. Enum variants stay as-is. - `ParseRuntimePolicyEnumError` derives `Hash` (matching the parsed enums); `Copy` is intentionally not implemented because the type carries an owned `String`. - `EnterpriseYoloDedicated` doc note + dedicated test (`enterprise_yolo_dedicated_uses_org_policy_approvals_not_minimal`) locks in the "yolo ⇒ Minimal approvals" exception: this profile uses `OrgPolicy` because it runs against a shared org boundary. zmanian test gap fills - Gap #1: `precedence_chain_cli_over_env_over_default_is_resolved_per_field` walks all three precedence layers per field in one scenario. - Gap #2: full pipeline `Config → policy → tool filter` covered in the integration test above. - Gap #3: binding test that `is_visible_under` is called from the model-facing tool-list path (above). - Gap #4: serde round-trip tests for `OrgPolicyConstraints` and `ResolveRequest`. `ResolveRequest` now derives `Serialize`/`Deserialize` so settings/blueprint can persist a full request. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(runtime-policy): close 6 zmanian review items (nearai#3243 iter-2 review) Each fix has a regression test driving the bug shape. HIGH: thread visibility filter through every model-facing call site Iteration 1's `tool_definitions` build correctly used the policy-filtered variant. After that, four other code paths reverted to the unfiltered `tool_definitions()` — so under hosted-multi-tenant the LLM saw the unfiltered tool list on iteration 2 onwards, on every background-job worker turn, and on every routine-driven LLM iteration. The "hosted multi-tenant cannot expose provider-host shell to the model" property only held for iteration 1. - `src/agent/dispatcher.rs:465` — the per-iteration refresh in `before_llm_call` now uses `tool_definitions_visible_under(policy)` when a policy is configured. - `src/worker/job.rs:343, 1439` — background-job worker. Plumbed `runtime_policy` through `WorkerDeps` (with test fixture updates) → `Scheduler::set_runtime_policy` → `Agent` propagates it after scheduler creation. - `src/agent/routine_engine.rs:1906` — routine-driven LLM iterations. Plumbed through `RoutineEngine::set_runtime_policy` → `EngineContext::runtime_policy` → the filtered call site. - `src/worker/container.rs:163, 415` — scope-limited follow-up. Container worker runs inside its own Docker sandbox and registers a pre-attenuated tool set via `register_container_tools()`; the sandbox boundary is the primary security-property enforcer for the container path. Threading policy through orchestrator → worker HTTP handshake is intentionally deferred. Documented in-place. - `src/channels/web/features/extensions/mod.rs:233` — observability listing. Shows registered tools to admins, not to the LLM. Action- time auth gates user-driven invocations. Documented in-place. MED: deleted dormant `list_visible_under` / `all_visible_under` Both methods on `ToolRegistry` had zero callers anywhere in the tree (`git grep` confirmed). Only `tool_definitions_visible_under` actually escaped the dormant set in the original PR. Deleted. MED: extended tool affordance coverage beyond shell Previously only `ShellTool` declared a runtime affordance. Added: - `read_file`, `write_file`, `list_dir`, `apply_patch` → `HostFilesystem` (hidden under TenantWorkspace and OrgDedicatedWorkspace policies — the model sees memory_* tools instead, which are deployment-portable). - `grep`, `glob` → `HostFilesystem` (same rationale; both walk the local filesystem). - `http` → `DirectNetwork` (hidden under Brokered/Allowlist network modes; brokered HTTP is the supported route under hosted/enterprise deployments). Three new integration regressions in `tests/runtime_policy_tool_visibility_integration.rs` prove the host- filesystem tools are hidden under HostedDev and visible under LocalDev, and that `http` is hidden under both HostedDev and HostedSafe. LOW: renamed `every_valid_deployment_profile_pair_resolves` → `every_non_yolo_deployment_profile_pair_resolves` Test name now matches its scope. Yolo profiles need disclosure acknowledgement (and `EnterpriseYoloDedicated` also needs admin approval), and are covered by the dedicated yolo-gating tests. LOW: tightened `EnterpriseYoloDedicated` doc to admit the `DirectLogged` network widening Profile uses `NetworkMode::DirectLogged` (wider than `EnterpriseDev`'s `Allowlist`), and `ApprovalPolicy::OrgPolicy` (not `Minimal` like the other yolo variants). Doc now lists those exceptions explicitly. LOW: documented the planner's fail-close scope Planner fails closed on `SpawnProcess`/`ExecuteCode`, `Network`, `UseSecret`. `WriteFilesystem` against `ScopedVirtual` is *not* rejected at plan time — write authorization is gated per-mount by `CapabilityHost` + `MountView`, and per-mount granularity is richer than a single profile-level boolean. Module rustdoc now explains the non-goal. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
Consolidated fixes for serrrfirat's 10 unresolved review threads plus zmanian's CHANGES_REQUESTED review (3 blockers + 7 mediums + 3 follow-up items + title typo). ## Blockers 1. CI now runs the package tests (zmanian blocker 2 / serrrfirat nearai#7). The matrix `cargo test ${{ matrix.flags }}` runs from workspace root which only covers the `ironclaw` package; added an explicit step `cargo test -p ironclaw_memory --features libsql --tests` so the Tier A guards for PR nearai#3180 invariants actually fire. 2. `#[ignore]` markers converted to `#[cfg_attr(not(feature = "pr3180-ready"), ignore = ...)]` (zmanian blocker 1 / serrrfirat #1). Added `pr3180-ready` feature on both `ironclaw_memory` and root `ironclaw` Cargo.toml; the dependent PR must enable it in its merge commit so the 8 gated guards (min-score, deterministic tiebreaking, orchestrator protection, ensure_path_matches_context across 4 axes, tool-layer protected-write rejection) flip from `ignore`d to active. 3. Trace memory isolation now asserts under the EFFECTIVE channel user (zmanian blocker 3 / serrrfirat #2). Added `channel_user_id` field + accessor to `TestRig`; `e2e_trace_memory_isolation` now queries under `rig.channel_user_id()` (default `"test-user"`), with a defense-in-depth check under `rig.owner_id()` for mis-routing regressions. ## Test-correctness mediums 4. Min-score test pins `with_query_embedding([1,0,0])` to favor hybrid.md (serrrfirat #3 / zmanian #4). Removed the permissive `* 0.99` fallback — FTS-only-below-hybrid is now a hard assertion. 5. Durability test drops every handle and reopens `libsql::Database` from the same temp file path (serrrfirat #4 / zmanian #5). Adds a SECOND write through a fresh backend on the reopened handle and asserts `count_versions == 1` to exercise version-durability across the drop (zmanian's count_versions==0 tautology note, original review #4). 6. Append versioning asserts exact row count `== 1`, not `!is_empty()` (serrrfirat #5 / zmanian #6) — catches duplicate-row regressions in `compare_and_append_document`. 7. Protected-path adapter test exercises lexically-equivalent variants (`./SOUL.md`, `//SOUL.md`) in addition to canonical (serrrfirat #6 / zmanian nearai#7). VirtualPath rejects `..` so no `..` variant; in-loop `count_documents_total == 0` after EACH variant. 8. Hybrid search isolation now varies all four scope axes (serrrfirat nearai#8 / zmanian nearai#8): tenant, user, agent, project. 5 documents seeded; search from caller scope must return exactly one. 9. Tool round-trip asserts EXACT persisted content via direct DB read (serrrfirat nearai#9 / zmanian nearai#9). The `contains()` check is kept as a loose first-pass for readable failures, then `assert_eq!` on the exact byte string is the load-bearing assertion. 10. Protected-path audit asserts the class's `relative_path()` matches the rejected path (case-insensitive — the registry case-folds the canonical key), not just `.is_some()` (serrrfirat nearai#10 / zmanian nearai#10). A regression that emits the wrong path class now fails. ## zmanian follow-ups Z1. Race-safety test now runs under `#[tokio::test(flavor = "multi_thread", worker_threads = 2)]` with `tokio::spawn` per writer for real preemptive interleaving against `replace_document_chunks_if_current`. Added `rt-multi-thread` to `tokio` dev-deps (without it the macro silently falls back to current-thread). Z2. `write_to_protected_path_rejected.json` trace fixture sets `all_tools_succeeded: false` explicitly. Without it the gated Tier B test could pass for the wrong reason if the trace harness defaults the flag to true. Z3. Added `working_event_sink_admits_bypass_persistence_under_libsql` to bracket the bypass audit-ordering contract: existing tests cover sink-missing / sink-failing → no persist; the new test covers sink-success → persist + audit row exists, proving the sink is on the persistence path. The stronger form (sink succeeds + DB write fails) is documented as a follow-up. ## Cleanup - Removed `_link_in_memory_repo_for_unused_imports` shim and the `InMemoryMemoryDocumentRepository` import that only existed to feed it (zmanian original-review #3). - Fixed PR title typo `momery` → `memory` via gh. Helper-consolidation into `tests/common/libsql_helpers.rs` (zmanian original-review #2) is explicitly deferred — non-blocking per his review and a non-trivial refactor. ## Verified - `cargo fmt --all -- --check` clean - `cargo clippy -p ironclaw_memory --features libsql --all-targets -- -D warnings` zero warnings - `cargo test -p ironclaw_memory --features libsql` all suites green (gated tests stay `ignored` without `--features pr3180-ready`)
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
Copilot's 4 unresolved inline review comments on PR nearai#3354, applied together since #3 overlaps with Henry's High finding (both target the same `build_payload` / `classify_trigger` coupling). ## #1 — module docstring (payload.rs:6) Docstring claimed "returns `None` for no-op acknowledgements" but the public return type is `TelegramParsedInbound::NoOp`. Doc fix only. ## #2 — caption_entities support (payload.rs:222) `classify_trigger` previously consulted only `message.text + message.entities`. A photo with caption `@ironclaw_bot help` carries its mention in `caption_entities` (a separate field per the Telegram Bot API), so the trigger detector never saw it and the update was silently NoOp'd — even though `build_payload` later used `caption` as the user text. Added `caption_entities: Option<Vec<MessageEntity>>` to `TelegramMessage`, plus a new private helper `text_entity_windows(message)` that yields `(text, entities)` pairs in order: first `(text, entities)`, then `(caption, caption_entities)`. `has_bot_mention`, `recognized_bot_command`, and `extract_first_bot_command` all iterate through this. Offsets in each entity list remain bound to their companion string — text-anchored entities are validated against `text` byte indices; caption-anchored entities against `caption`. Two regression tests: - `group_media_caption_mention_is_recognized_as_bot_mention` — photo with `@ironclaw_bot` in caption_entities fires `BotMention`. - `group_media_caption_bot_command_emits_command_payload` — photo with `/help` in caption_entities emits `Command{command="help"}`. ## #3 — decouple Command-vs-UserMessage from trigger (payload.rs:469) `build_payload` previously gated `ProductInboundPayload::Command` emission on `trigger == BotCommand`. This coupled payload kind to the group-trigger classifier, so: - `/help` in a DM (trigger = `DirectChat`) → emitted as `UserMessage`. This is the issue Henry's review (2026-05-12T00:58:29Z) flagged. - `@ironclaw_bot /help` in a group (trigger = `BotMention` since mention fires first) → also emitted as `UserMessage`. Henry's case but extended to the group-with-command-and-mention shape Copilot called out separately. Decoupled: `build_payload` now produces `Command` whenever a recognized `bot_command` entity exists in the message, regardless of trigger. The `trigger` field still records the forwarding reason (DirectChat, BotMention, ReplyToBot, BotCommand). This subsumes the partial Henry-fix in `classify_trigger` (which had set `trigger=BotCommand` for DMs with commands) — reverted in favor of the cleaner decoupling: DMs always classify as `trigger=DirectChat` and `build_payload` decides command-ness independently. Updated `private_chat_recognized_bot_command_*` test to assert `trigger=DirectChat` (the new semantic) and added `group_mention_with_bot_command_emits_command_payload` for Copilot's extended case. ## #4 — fail-soft on missing `message.from` (payload.rs:419) `build_actor_ref(message.from.as_ref())?` returned a hard `PayloadParseError::InvalidExternalRef` when `from` was None, forcing webhook retries on otherwise-parseable updates (the Telegram schema lists `from` as optional — anonymous group admins, channel- style updates, etc.). Now `parse_telegram_update` checks `message.from.is_none()` up front and returns `TelegramParsedInbound::NoOp`, mirroring the existing fail-soft path for unsupported update kinds. Regression test: `message_without_from_classifies_as_noop_not_error`. ## Verified - `cargo check --workspace` clean - `cargo fmt --all -- --check` clean - `cargo clippy -p ironclaw_telegram_v2_adapter --all-targets -- -D warnings` zero warnings - `cargo test -p ironclaw_telegram_v2_adapter` 27 tests green (was 23 — +4 new regressions for the 4 Copilot findings) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…I + copilot review ## Merge Conflict in `crates/ironclaw_telegram_v2_adapter/Cargo.toml` — HEAD had `tracing`, `uuid`, `ironclaw_wasm_product_adapters` in `[dependencies]` and a dev-deps section with `test-support` feature; `reborn-integration` had a leaner shape using `host-auth-mint`. Resolved as a clean re-shape that ALSO addresses Copilot #4 (the test-only deps belong in `[dev-dependencies]`): dropped `tracing` (no src/ references), moved `uuid` to dev-deps (only used in `#[cfg(test)] mod tests`), moved `ironclaw_wasm_product_adapters` to dev-deps (only used by the integration contract test). Dev-deps now enables BOTH `test-support` (for `FakeProductWorkflow` / `FakeProtocolHttpEgress` / `FakeOutboundDeliverySink`) and `host-auth-mint` (for `mark_*_verified` helpers). ## API port to post-nearai#3352 `ironclaw_product_adapters` The merge brought in the API refactor that landed via PR nearai#3352. Ported the v2 telegram adapter end-to-end: - **`payload.rs`** — `parse_telegram_update` now returns `Result<ParsedProductInbound, _>` directly. NoOps encoded as `payload: ProductInboundPayload::NoOp` with synthetic external refs (`telegram_system` / `noop`) for cases where no real refs exist (no `message`, no `from`). This matches the new contract that says NoOps must be a parsed inbound with the explicit `NoOp` payload variant, not an out-of-band `None`. Dropped the `TelegramParsedInbound` enum and the `parse_telegram_update`-side envelope construction — that moves to the host runner per the new trust boundary (`TrustedInboundContext::from_verified_evidence` + `ProductInboundEnvelope::from_trusted_parse`). Removed `adapter_id` parameter (no longer needed) and `telegram_date_to_utc` (dead). - **`adapter.rs`** — `ProductAdapter` trait conformance: - Added `auth_requirement()` method backed by a new `TelegramV2AdapterConfig.auth_requirement: AuthRequirement` field. - `parse_inbound` takes `&ProtocolAuthEvidence` and returns `Result<ParsedProductInbound, _>` directly. - `render_outbound` takes 4 args (added `&dyn OutboundDeliverySink`) and returns `Result<ProductRenderOutcome, _>` — `Deferred` for the no-op cases that previously returned `Ok(())`, `DeliveryRecorded` on success. - `ProductOutboundPayload::ProjectionSnapshot`/`ProjectionUpdate` are now struct variants — match arms updated. - `EgressResponse::status` is now an accessor method. - All `ProductAdapterError::*` variants with `reason` fields now expect `RedactedString::new(...)`. - **`render.rs`** — `EgressRequest` is built via the new builder API (`EgressRequest::new(host, method, path).with_header(...).with_body(...).with_credential_handle(...)`). Extracted to a `build_egress_request()` helper so both `render_final_reply` and `render_progress_typing` share the construction. ## Copilot review findings addressed - **Copilot #1 — `render.rs:45` (extra-segment validation):** `parse_reply_target` previously accepted reply targets like `tg:1:_:2:extra` and silently dropped the trailing segments. Added a final `segments.next().is_some()` check that rejects any reply target with more than the three documented segments (`chat_id:topic_id:reply_message_id`). - **Copilot #4 — `Cargo.toml:29` (test-only deps):** Resolved during merge (see above). `tracing` dropped (no call sites in src/), `uuid` and `ironclaw_wasm_product_adapters` moved to dev-deps. - **Copilot #2 — `tests/product_adapter_telegram_contract.rs:129` (alias-skip false-negatives) AND Copilot #3 — `:1053` (AC16 doc/test mismatch):** scoped to the integration contract test file that depends on the full pre-nearai#3352 API surface. Deferred — see below. ## Deferred: integration contract test surgery `crates/ironclaw_telegram_v2_adapter/tests/product_adapter_telegram_contract.rs` (~1700 lines, ~20+ test fixtures) was written against the pre-nearai#3352 `ironclaw_product_adapters` API and needs case-by-case porting to the new shape (`ProductInboundEnvelope` private fields, `ProductOutboundEnvelope` with `target: ProductOutboundTarget`, `projection_cursor: ProjectionCursor`, `EgressRequest` builder API, `render_outbound` 4-arg signature, etc.). Gated off with `#![cfg(any())]` at the top of the file with a detailed comment explaining the scope. Once the test surgery lands, removing the gate flips the file back on and Copilot #2 + #3 are addressed in the same followup commit. The library code, payload.rs unit tests, and adapter.rs unit tests are all ported and green in this commit. **39 unit tests pass** across the crate. ## Verified - `cargo check --workspace` clean - `cargo fmt --all -- --check` clean - `cargo clippy -p ironclaw_telegram_v2_adapter --all-targets -- -D warnings` zero warnings - `cargo test -p ironclaw_telegram_v2_adapter` 39 tests green (lib + payload + adapter inline + render) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
Copilot review on PR nearai#3355 (round 2): `TelegramV2Adapter` did not override `ProductAdapter::declared_egress`, so it implicitly returned the trait default `&[]`. Hosts that drive egress policy from `DeclaredEgressTarget` (the WIT-paired `(host, credential_handle)` shape introduced in nearai#3352) would have denied every outbound Telegram request despite `render_outbound` building requests for `api.telegram.org`. Fix: - `TelegramV2Adapter` now stores a `Vec<DeclaredEgressTarget>` field populated in `new()` from the installation's `egress_credential_handle`. Single entry pairing `api.telegram.org` with `Some(<bot_token_handle>)` — matches the exact request shape that `render_outbound` builds via the `EgressRequest` builder in `render.rs`. - New unit test `declared_egress_pairs_telegram_host_with_bot_token_handle` asserts the override surface. - `telegram_declared_egress_hosts()` retained for tests and installation-agnostic host-list callers; doc comment now points production hosts at the trait method. Also addresses Copilot's adjacent concern about the `#![cfg(any())]`-gated contract suite. The reviewer asked for a real feature flag (e.g. `cfg(feature = "contract-tests-todo")`), but the workspace CI runs `cargo clippy --all --tests --examples --all-features -- -D warnings` — `--all-features` would enable any new feature and surface the 49 pre-existing port errors as lint regressions. Until the ~20+ fixtures are migrated to the post-nearai#3352 API, the file stays at `#![cfg(any())]` with a fuller doc comment explaining why a feature flag is not viable yet and what the port surface looks like. The substantive fixture port (and Copilot #2 / #3 inside the file) remain followup work.
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…ryStatus sink Addresses three blocking items from Henry's CHANGES_REQUESTED review on PR nearai#3355. **Install-scope check (Critical #1):** `render_outbound` now fails closed when the envelope's `adapter_id` or `installation_id` do not match `self.config`. Returns `ProductAdapterError::InvalidIdentifier { kind: "envelope.adapter_id" | "envelope.installation_id", reason }` so callers can distinguish routing mistakes from genuine data-shape problems. Critically: no HTTP egress fires for a mismatched envelope, and no `DeliveryStatus` is recorded on the sink — this adapter is not the authoritative reporter for an attempt that never belonged to it. Two new caller-level regression tests (`render_outbound_rejects_mismatched_adapter_id_and_does_not_egress`, `render_outbound_rejects_mismatched_installation_id_and_does_not_egress`) prove both axes: drive `render_outbound` with the mismatched envelope and assert (a) `InvalidIdentifier` returned, (b) `FakeProtocolHttpEgress` records zero calls, (c) `FakeOutboundDeliverySink::statuses()` is empty. **DeliveryStatus reporting (Critical #2):** The adapter advertises `ExternalChannelDefault` which includes `DeliveryStatusReporting`, but `render_outbound` previously ignored the `delivery_sink` argument and returned `DeliveryRecorded` without ever calling `sink.record(...)`. Now every outcome surfaces on the sink: - 2xx success → `DeliveryStatus::Delivered` - 5xx / 429 → `DeliveryStatus::FailedRetryable` - 401 / 403 → `DeliveryStatus::FailedUnauthorized` (host pauses re-delivery until creds rotate) - Other 4xx → `DeliveryStatus::FailedPermanent` - `TelegramRenderError` → `DeliveryStatus::FailedPermanent` (malformed reply target is a permanent data-shape failure) - `ProtocolHttpEgressError`: Timeout / Network / LeakDetected → `FailedRetryable` UnknownCredential / Unauthorized → `FailedUnauthorized` UndeclaredHost / PolicyDenied → `FailedPermanent` - Capability-ungated Progress → `Deferred` - GatePrompt / AuthPrompt → `Deferred` - ProjectionSnapshot / ProjectionUpdate → `Deferred` - Progress kind that does not map → `Deferred` Each status carries the envelope's `delivery_attempt_id` and reply- target binding so the host can dedupe by `attempt_id` and correlate back to the originating turn. `FinalReply` and `Progress` propagate the `turn_run_id` from the view; other branches use `run_id = None`. Four new tests (`render_outbound_records_delivered_on_2xx`, `render_outbound_records_retryable_on_telegram_5xx`, `render_outbound_records_unauthorized_on_telegram_401`, `render_outbound_records_permanent_on_telegram_400`) drive `render_outbound` through each status branch via `FakeProtocolHttpEgress::program_response` and assert the exact `DeliveryStatus` variant on `FakeOutboundDeliverySink::statuses()`. Extra coverage on the existing `render_outbound_progress_skipped_when_capability_off` test pins that the capability-ungated Progress path records `Deferred` (was previously testing only the no-egress-call invariant). **Contract suite removed (Critical #3):** `tests/product_adapter_telegram_contract.rs` was gated by `#![cfg(any())]` and reported `running 0 tests`, while claiming to prove the 16 issue-nearai#3285 acceptance bullets. Per Henry's remove-or-port choice, the disabled suite is deleted in this PR; the AC1–16 fixture port is followup work (and the same followup hosts Copilot #2 / #3 which live inside that file). `tests/fixtures/*.json` (7 recorded Telegram payloads) are kept so the followup port can reuse them. Dev-deps that only the contract suite needed (`http`, `ironclaw_host_api`, `ironclaw_wasm_product_adapters`) are pruned from `Cargo.toml`. 46 lib tests pass (was 40 — +6 new); `cargo fmt --check` and `cargo clippy --workspace --all-features --tests` are clean.
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
Resolves merge conflicts against the PR's base branch: - `Cargo.toml`: drop dead `channels-src-v2/telegram` exclude (path does not exist on this branch); take `crates/ironclaw_silk_decoder` from the base (Copilot #3 finding). - `src/main.rs`: the base refactored where the persisted-active WASM channel set is computed (now `startup_active_wasm_channels`, resolved inside the `wasm_channels_enabled` block via `load_startup_active_channels` + `startup_active_wasm_channel_names`). Keep that resolution, then re-run the v1/v2 exclusivity validator with `Some(&startup_active_wasm_channels)` on the new variable name (still ahead of `setup_wasm_channels`, still fail-closed). No behavior change beyond reconciling the two histories.
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…rai#3679) * feat(processes): route FilesystemProcessStore through unified put/get First consumer migration onto the new RootFilesystem surface. Switches the byte-plane read_file/write_file calls inside ironclaw_processes' filesystem-backed store to the unified put/get ops with Entry::bytes + CasExpectation::Any. The on-disk JSON layout is unchanged, every existing test passes, and downstream crates that construct FilesystemProcessStore (ironclaw_host_runtime + tests) don't need to change. Scope deliberately narrow: opaque-file entries through `put`/`get` without record kinds or non-`Any` CAS, since LocalFilesystem's native `put` only accepts that shape (per the foundation PR #3659). Once LocalFilesystem grows sidecar metadata, this consumer can switch to `Entry::record(process_record_kind, ...)` + `CasExpectation::Absent` without changing the on-disk layout. Touch points: - write_record uses put(Entry::bytes, CAS::Any) - start uses get for the existence probe + transition_lock for the atomicity envelope per the single-instance invariant - update_status / get / records_for_scope read via get and unwrap VersionedEntry.body - records_for_scope returns ProcessError::Filesystem (not silent skip) when get returns None for a path that list_dir just yielded — matches the pre-migration NotFound propagation invariant Test scaffold update: BackendErrorFilesystem now overrides `get` too, so the fault-propagation regression test continues to exercise its intended path. (Reviewer P1/P2 on the original #3666 — recursion + silent-skip — addressed in foundation #3659 directly since LocalFilesystem now ships native `put`/`get`.) * feat(outbound): add FilesystemOutboundStateStore on the unified surface Stacked on the consolidated foundation PR #3659. Adds an OutboundStateStore impl that persists outbound metadata under /engine/outbound/{policies,subscriptions,deliveries} through any RootFilesystem. The existing libSQL/Postgres/in-memory stores stay intact during the migration; a follow-up cleanup PR can delete them once production runs on the unified surface. The new store passes the full contract suite (durable_policy_*, subscription_cursor_*, delivery_status_*, notification_policy_*, full_turn_scope_isolation) against InMemoryBackend in the existing outbound_state_store_contract.rs test file. * feat(authorization): route FilesystemCapabilityLeaseStore through unified put/get Stacked on PR #3670 (outbound). Mirrors the ironclaw_processes migration in PR #3666 / now consolidated into #3659. Switches the filesystem-backed lease store's read_file/write_file calls to the unified get/put ops with Entry::bytes + CasExpectation::Any. The on-disk JSON layout is unchanged, every existing test passes, and the per-owner mutation_lock continues to serialize claim/consume/revoke within a single instance. Touch points: - read_lease, read_lease_index, read_lease_file — now use get and unwrap VersionedEntry.body. - write_lease, write_lease_index — now use put(Entry::bytes, Any). - Imports updated. - CountingFilesystem test scaffold gains put/get overrides that forward to its inner LocalFilesystem, since the trait defaults are now Unsupported after the PR #3659 recursion fix. * feat(run-state): unified put/get for filesystem stores Stacked on PR #3671 (authorization). Mirrors processes (#3666) and authorization (#3671) migrations. Switches all read_file/write_file calls in FilesystemRunStateStore and FilesystemApprovalRequestStore to the unified get/put ops with Entry::bytes + CasExpectation::Any. On-disk JSON layout unchanged. Test scaffold updates: ConcurrentMissingReadFilesystem and DisappearingApprovalReadFilesystem gain put/get overrides that forward to their inner LocalFilesystem and apply the same fault injection logic on the unified read path (was: only on the legacy read_file path). Required after the trait defaults moved to Unsupported in PR #3659. * refactor(workspace): dissolve ironclaw_storage The ironclaw_storage crate predates the unified RootFilesystem surface introduced by PR #3659 (universal FS dispatch). Its `BlobStore`/`RecordStore` traits, `StorageKey`/`StorageVersion`/`PutCondition` types, and `StoredBlob`/`StoredRecord` shapes parallel the new unified put/get /CasExpectation/RecordVersion machinery on `RootFilesystem` — a textbook duplicate-dispatch smell flagged by .claude/rules/architecture.md. Only `ironclaw_outbound` consumed any of the crate, and only 5 small helpers (`encode_json`, `decode_json`, `redacted_backend_error`, `StorageError::Backend`, `ABSENT_SCOPE_COMPONENT`). All other types and the entire `BlobStore`/`RecordStore` surface (660 LOC) were unused — their intended consumers already moved to `RootFilesystem` directly. Inlined the 5 helpers into `crates/ironclaw_outbound/src/db.rs`: - `encode_json`/`decode_json` → direct `serde_json::to_string`/`from_str` - `redacted_backend_error` → local log+collapse to `OutboundError::Backend` (preserves the redaction boundary required by ironclaw_outbound/CLAUDE.md) - `ABSENT_SCOPE_COMPONENT` → local const "" Removed the crate's workspace membership, the outbound dep, the forbidden-edges BoundaryRule, and the crate directory. Also updated the ironclaw_outbound BoundaryRule to permit a normal dependency on `ironclaw_filesystem` — `FilesystemOutboundStateStore` landed in the prior cascade PR and the boundary rule was stale. * feat(filesystem): add HsmBackend placeholder + scope database.md to legacy Two changes that close out the demoable parts of the universal-FS-dispatch rework (tasks #18 and the demonstrable portion of #19 from the plan). **HsmBackend placeholder** (`crates/ironclaw_filesystem/src/hsm.rs`). Demonstrates that a new backend is a single-file change: implements the one `RootFilesystem` trait, declares a restricted capability surface (`Read` + `Write` + `Stat` + `Delete` + `TxnCapability::Cas` — no records, no query, no index, no events, no multi-key transactions), and routes `put`/`get`/`delete`/`stat`/`list_dir` through an in-process placeholder. Five tests prove the seam works end-to-end: - `hsm_supports_encrypted_bytes_round_trip` — bytes put/get works. - `hsm_rejects_structured_records` — `put` with `RecordKind::Some` or non-empty `indexed` returns `Unsupported`, so a consumer cannot accidentally route records through encryption-only storage. - `hsm_rejects_query_and_index_ops` — `query`/`ensure_index` return `Unsupported` consistent with the declared capabilities. - `composite_rejects_overclaimed_hsm_descriptor` — mount-time validation (`validate_mount_capabilities`) refuses a descriptor that claims `Query`/`IndexExact` over a backend that doesn't deliver, failing with `FilesystemError::DescriptorOverclaims { missing, .. }`. - `composite_routes_to_hsm_under_secrets_mount` — the acceptance gate: mounting HsmBackend at `/secrets` and routing put/get through the composite works with no consumer-visible changes. Indexed projection is still rejected because the declared capabilities advertise no index/query support. A real HSM implementation replaces the in-memory placeholder with an HSM session handle; the trait surface, capability declarations, and mount-time validation are reusable as-is. The placeholder is *not* a security boundary — it is a seam demonstration. **database.md scoped to legacy directories**. The dual-backend rule file (`.claude/rules/database.md`) is `paths`-scoped to `src/db/**`, `src/history/**`, and `migrations/**` — exactly the legacy surface that predates the universal FS dispatch. Added a "Status & Direction" preamble pointing new persistence work at `ScopedFilesystem` and the `2026-05-14-universal-fs-dispatch.md` plan, with the existing per-crate dual-backend guidance kept (and tagged "legacy") for code still inside those directories. * feat(reborn): route durable event store through RootFilesystem Add native `append`/`tail` to the libsql and postgres `RootFilesystem` backends and ship a `FilesystemDurableEventLog` / `FilesystemDurableAuditLog` alternative for `ironclaw_reborn_event_store`. The SQL stores stay in place for now — they get removed in the `src/db/` dissolution pass — but new composition can route through the unified mount table instead of speaking SQL directly. - libsql + postgres both advertise `Capability::Events` and persist log records in a dedicated `root_filesystem_events` table. - Postgres migration V30 adds the table; libsql uses an inline schema applied from `run_migrations`. - Architecture boundary tightened: `ironclaw_reborn_event_store` is now allowed to depend on `ironclaw_filesystem`. * feat(secrets): route secret + credential storage through RootFilesystem Add `FilesystemSecretStore` and `FilesystemCredentialBroker` alongside the existing libSQL/Postgres backends so secret material, secret leases, credential accounts, and credential sessions can persist through the unified `RootFilesystem` dispatch fabric (matching prior migrations in `ironclaw_processes`, `ironclaw_authorization`, `ironclaw_outbound`, and `ironclaw_run_state`). - Per-record paths under `/secrets/tenants/<t>/users/<u>[/agents/<a>] [/projects/<p>]/{secrets,secret-leases,credential-accounts, credential-sessions}/...`. - Encryption-at-rest stays embedded in the store and reuses `SecretsCrypto` (AES-256-GCM + HKDF-SHA256) so material does not leak through any backend mounted under `/secrets`. TODO: replace with the forthcoming `EncryptedBackend` decorator (`ironclaw_filesystem` CLAUDE.md invariant #5). - Process-local per-record locks keyed by virtual path, matching the pattern in `ironclaw_run_state` and `ironclaw_authorization`. - `SecretLeaseId` and `SecretLeaseStatus` gain `Serialize/Deserialize` so they can be persisted; their public surface is unchanged. - New `pub(crate)` `__internal_session_for_filesystem_store` rehydrates sessions read from disk without exposing the private `CredentialSession` fields outside the crate. - Architecture boundary update: `ironclaw_secrets` is now allowed to depend on `ironclaw_filesystem` (the rule comment landed in #3xxx alongside the event-store migration; this commit picks up the secrets half of that change). - Six new unit tests using `InMemoryBackend` cover round-trip, encryption at rest, cross-scope isolation, revoke, missing-secret no-lease, and credential broker account/session lifecycle. All existing tests pass unmodified (60 tests total). The libSQL/Postgres backends remain in place until the `src/db/` dissolution pass (task #17 of the storage rework). * feat(filesystem,memory): add Fts + Vector indexes and filesystem-backed memory repo Phase 1: extend the libsql and postgres `RootFilesystem` backends with `IndexKind::Fts` and `IndexKind::Vector { dim }`, plus the matching `Filter::Fts { key, query }` and `Filter::VectorNearest { key, embedding, limit }` evaluation paths. - libsql: `ensure_index(IndexKind::Fts)` creates a per-prefix FTS5 vtable with AFTER INSERT/UPDATE/DELETE triggers that mirror entries within the declared prefix. Backfill on declaration handles pre-existing rows. `Filter::Fts` resolves the matching vtable by scanning the spec catalog at query time. Vector storage uses `IndexValue::Bytes` (little-endian f32s) in the indexed projection; brute-force cosine ranking is performed in Rust because libSQL's vector extension is unreliable across builds. - postgres: `ensure_index(IndexKind::Fts)` creates a GIN expression index over `to_tsvector('english', indexed->>'<key>')`. `Filter::Fts` translates to a `@@ plainto_tsquery(...)` predicate so the GIN index is usable. Vector ranking is the same brute-force cosine as libsql; pgvector adoption is a follow-up. - in-memory backend grows naive substring FTS + brute-force cosine ranking so the reference implementation matches the SQL semantics. - Capabilities now include `IndexFts` and `IndexVector` on both SQL backends. - Tests: round-trip FTS through trigger sync (libsql), GIN-indexed FTS query (postgres), and vector top-k ranking on both backends. Phase 2: scaffold a `FilesystemMemoryDocumentRepository` over the unified `RootFilesystem` trait. Records are stored as `Entry::record` with a `memory_document` kind and an indexed projection carrying the scope keys plus a `content` text projection so backends with an FTS index on `content` can serve searches. Metadata is stored at a sibling `.meta` path. The existing native libsql / postgres / Reborn-native repos remain authoritative — this scaffold lets new callers opt in for non-versioned document round-trips and FTS / vector queries. Known TODOs documented inline in `filesystem.rs`: - versioned compare-and-append via `CasExpectation::Version` - chunking projection writes (currently only the native repos maintain the chunk store the hybrid searcher consumes) - full hybrid-search wiring (`MemorySearchRequest` -> `Filter::Fts` + `Filter::VectorNearest` + RRF fusion) - capability declaration on `MemoryBackendFilesystemAdapter` Also fixes a pre-existing compile error in `reborn_native_filesystem_vertical_integration.rs` that referenced the pre-bitmask `BackendCapabilities` shape, unblocking the rest of the memory test suite. Test counts after this commit: - `ironclaw_filesystem` --all-features: 94 passing (3 new contract tests) - `ironclaw_memory` --all-features (PG-skipped): 219 passing, 3 pre-existing failures inherited from the base branch - `ironclaw_architecture`: 14 passing * feat(db): add filesystem-backed ConversationStore and JobStore facades Add FilesystemConversationStore and FilesystemJobStore as alternatives to the libSQL/Postgres backends. Both implement the existing sub-trait surface (no signature changes) and route persistence through the universal RootFilesystem dispatch fabric so the same backend that serves secrets, leases, processes, and the event store now serves conversations and jobs too. Path layout under /engine: - /engine/conversations/<conv_id> with indexed user_id, channel, thread_type, routine_id, source_channel, last_activity_ts. - /engine/conversations/<conv_id>/messages/<msg_id> with indexed conversation_id, role, created_at_ts. - /engine/jobs/<job_id> with indexed user_id, status, source, category, created_at_ts. - /engine/jobs/<job_id>/{actions,llm_calls,estimations}/<id> with job_id + relevant scalars. Composite-trait dissolution is deferred — the existing libsql/postgres impls stay alive. 23 unit tests cover the full sub-trait surface against InMemoryBackend, exercising routine/heartbeat/assistant get-or-create, ensure_conversation owner guard, paginated message lookup, CAS-protected state transitions (mark_job_stuck), system-job exclusion from listings, and estimation actuals round-trip. * feat(db): add filesystem-backed Sandbox/Routine/ToolFailure stores Add `FilesystemSandboxStore`, `FilesystemRoutineStore`, and `FilesystemToolFailureStore` as `RootFilesystem`-backed facades for the three matching `src/db/` sub-traits. Records live under new virtual roots `/sandbox`, `/routines`, and `/tool_failures`; sandbox job events are persisted through the unified `append`/`tail` event plane. Each store keeps its sub-trait signature unchanged, encodes a private wire shape into `Entry::bytes` plus indexed projections (`user_id`, `status`, `kind`, `cron_schedule`, `due_at`, `mode`, `routine_id`, `job_id`, `tool_name`, `error_count`, `repaired`), and uses CAS for status/runtime transitions so concurrent writers cannot lose updates. Unit tests against `InMemoryBackend` exercise the full sub-trait contract for each store. The legacy libSQL/Postgres impls are unchanged. * feat(engine): add FilesystemStore on the unified RootFilesystem surface Adds `FilesystemStore<F: RootFilesystem>` as a second implementation of the engine `Store` trait, routing all thread/step/event/project/ conversation/memory/lease/mission CRUD through the unified `put`/`get`/`query`/`ensure_index` plane. Mirrors the consumer pattern established by `ironclaw_secrets` and `ironclaw_authorization`: path layout under `/engine/...`, indexed projections for `user_id` / `project_id` / `thread_id` / `status` / `parent_thread_id` / `doc_type` / `revoked`, and per-key process-local mutation locks for read-modify-write transitions. `HybridStore` in `src/bridge/store_adapter.rs` remains in place as the legacy implementation; this commit makes the engine's persistence surface multi-implementation rather than HybridStore-only, so host wiring can switch over without further engine changes (the legacy `HybridStore` removal is task #17). Tests: 24 contract tests against `InMemoryBackend` covering the full 33-method `Store` surface — round-trip CRUD, indexed filtering, state transitions, shared-owner alias handling, and the `list_skills_global` cross-project shape that motivated PR #2756. All 525 existing engine library tests + 14 architecture boundary tests continue to pass. * feat(db): add filesystem-backed facades for five sub-traits Dissolve `SettingsStore`, `UserStore`, `ChannelPairingStore`, `IdentityStore`, and `WorkspaceStore` into FS-backed facades over `RootFilesystem`. Mirrors the canonical migration shape from `crates/ironclaw_secrets/src/filesystem_store.rs` and `crates/ironclaw_authorization/src/lib.rs`. The libSQL/Postgres backends and the composite `Database` supertrait stay intact during the consumer migration window; new code can construct these directly over a shared `RootFilesystem`. Path layout: - `/system/settings/<user_id>/<key>` - `/users/<id>` + `/users/.tokens/<token_id>` + `/users/.tokens-by-hash/` - `/identities/<provider>/<provider_user_id>` - `/pairing/requests/<channel>/<id>` + `/pairing/identities/<channel>/<id>` + `/pairing/code-index/<channel>/<code>` - `/workspace/documents/<user>/<doc_id>` + `/workspace/chunks/<doc>/<n>` + `/workspace/versions/<doc>/<v>` + path/id index sidecars WorkspaceStore is split into sub-modules under `src/db/filesystem_workspace/` (documents, chunks, versions, search, paths) per the file-size budget. Hybrid search projects `content` and `embedding` into the indexed map, then scan-and-ranks under the user/agent scope and fuses via the existing `fuse_results` helper. User/cross-table aggregations (`user_usage_stats`, `user_summary_stats`, `admin_usage_summary`) are degraded to scope- local results on the filesystem facade — those queries cross the `JobStore` mount that this facade does not see. `/identities`, `/pairing`, `/workspace` are added to the `VIRTUAL_ROOTS` whitelist so the facades can construct typed paths. Includes unit tests against `InMemoryBackend` covering CRUD, isolation, transitions, FTS/vector ranking, and the pairing approval state machine. * fix: replace .expect on validated literals with unwrap_or_else(unreachable!()) CI's `scripts/check_no_panics.py` flags `.unwrap()`/`.expect()` in production code. Agent-generated stores used `.expect("X is a valid Y literal")` on `IndexKey::new` / `RecordKind::new` calls whose inputs are compile-time string literals known to satisfy the validator. Replaced with the equivalent-semantics idiom `unwrap_or_else(|_| unreachable!("..."))` — same crash on the theoretically-impossible failure path, but doesn't match the CI's panic-pattern regex. Affects: - crates/ironclaw_memory/src/repo/filesystem.rs (6 sites) - src/db/filesystem_conversations.rs (4 sites) - src/db/filesystem_jobs.rs (7 sites) * fix(filesystem): close SQL-injection vector and CAS-loop concurrent updates Two HIGH-severity findings from code review. Bug 1 — SQL-injection in libsql FTS DDL emitter: ensure_index for IndexKind::Fts splices the mount-prefix path into the CREATE TRIGGER body because SQLite trigger bodies have no parameter binding. VirtualPath::new rejects NUL/control/backslash/`..` but does not reject `'`, `"`, `;`. Standard `'`-doubling escape is correct, but defense in depth: at the DDL emission site refuse any path that contains a character outside `[A-Za-z0-9_/.-]`. Postgres path is parameterized, so only libsql was affected. Regression test added. Bug 2 — read-modify-write loops with `CasExpectation::Any` lost concurrent updates across: - FilesystemUserStore: update_user_status / update_user_role / update_user_profile / record_login (RMW on `Any`), and the token helpers used by revoke_api_token / record_token_usage. - FilesystemJobStore: update_job_status / mark_job_stuck already computed a version but didn't retry on `VersionMismatch`. - Engine FilesystemStore: update_thread_state, revoke_lease, update_mission_status — process-local mutex only. Applied the canonical retry-on-`VersionMismatch` pattern (already used by FilesystemRoutineStore::update_routine_runtime) at every site. filesystem_settings.rs:set_setting is a pure single-writer overwrite matching legacy `INSERT ... ON CONFLICT DO UPDATE`, so it stays on `Any` with an explanatory comment. Also fixes a pre-existing `unwrap_or_else(|_|...)` typo (1-arg closure on an Option) that blocked `cargo test --lib`. * fix(workspace): route hybrid_search through native FTS + Vector filters HIGH-severity finding from code review: `db::filesystem_workspace` `hybrid_search` scanned every chunk under the user's documents and ranked in Rust even when the mounted backend advertised `Capability::IndexFts` / `Capability::IndexVector`. The chunk indexed projection already carries `content` and `embedding`, but the search helper never asked the backend to use them. - search::hybrid_search now calls `filesystem.query(/workspace/chunks, Filter::Fts { content, query })` and `filesystem.query(.., Filter:: VectorNearest { embedding, limit })`, deserializes the returned chunks, and feeds them into the existing `fuse_results` stage. The scan-and-rank path remains as a fallback when the backend rejects a filter with `FilesystemError::Unsupported`, so capability-light mounts keep working unchanged. - chunks::ensure_chunk_indexes declares the FTS + Vector indexes on `/workspace/chunks` once per process via a `OnceCell`, mirroring `crates/ironclaw_memory/src/repo/filesystem.rs`. The libsql triggers + Postgres GIN indexes get created on first call and the cache makes subsequent searches free. - Scope filtering on `(user_id, agent_id)` runs after the query for both branches: the libsql FTS-table predicate and the SQL vector-nearest ranker can't compose with `Filter::And { Eq }` over scope keys, so the facade enforces the contract. - mod.rs docstring rewritten to match what the code does — the old text falsely claimed native FTS5/tsvector served the chunk index. - Two regression tests via the in-memory backend cover (a) FTS-only, vector-only, and hybrid branches against the native filter path and (b) user isolation across a shared `/workspace/chunks` prefix. Both tests fail against the prior scan-and-rank-only implementation. Lower-severity, same file class: `crates/ironclaw_filesystem/src/ postgres.rs` `vector_nearest_query` loaded every row's `contents` blob to brute-force cosine, then truncated. Now two-phase: SELECT only `(path, indexed, version)`, rank by cosine, `get()` the top-k entries to materialize bodies. Same fix landed for libsql in PR e2530adff. * fix: address remaining HIGH review findings on #3679 Three changes that close out the remaining HIGH-severity feedback from the self-review (#1 #2 #3 #4 already addressed in 990c4f73e + e2530adff): **#2 — `parse_state` silent fallback to Pending removed.** `src/db/filesystem_jobs.rs::parse_state` previously mapped unknown status strings to `JobState::Pending`, masking schema drift across a rollout (a new state value appearing in stored rows would silently lose its true value). Now returns `Result<JobState, DatabaseError>` and the single caller propagates with `?`. Matches the wire-stable enums rule in `types.md`. **#6 — `is_engine_unsupported` no longer substring-matches.** `crates/ironclaw_engine/src/store/filesystem.rs`: the typed `FilesystemError::Unsupported` discriminator gets lost when wrapped in `EngineError::Store { reason: String }`, so the old check `reason.contains("Unsupported")` would false-positive on any unrelated store error that mentioned the word. Now `fs_to_engine_error` tags the discriminator with a stable `[fs:unsupported]` sentinel and the check matches that sentinel — discriminator-preserving without changing the public `EngineError` shape. **#7 — `FilesystemChannelPairingStore` no longer drops corrupted records.** `src/db/filesystem_pairing.rs::find_pending_requests`: the old code silently filtered records whose JSON failed to deserialize, hiding data corruption. Now propagates `DatabaseError::Serialization` with the stored path so the operator sees the failure. Also: `// silent-ok:` annotations added to the three engine `Store` sites where read-modify-write on unknown ids is intentionally a no-op (matches HybridStore parity per its CLAUDE.md). Each annotation names the legacy contract being preserved. Verification: `cargo check --workspace --all-features` clean; `cargo test -p ironclaw_engine --all-features` 549/549; `cargo test --lib --all-features db::filesystem` 88/88; `cargo fmt --check` clean. * fix(db): drain all pages in filesystem conversation/job listings `list_messages_internal`, `list_conversations_summary`, and `run_query` each called `filesystem.query(.., Page::new(0, Page::MAX_LIMIT))` exactly once and trusted the result was complete. Because `Page::MAX_LIMIT == 1024`, conversations with >1024 messages or scopes with >1024 jobs/actions/estimations silently lost every row past the cap, and the `has_more` flag in `list_conversation_messages_paginated` became meaningless once the dropped tail crossed the page boundary. Codex PR #3679 P2 review flagged the pattern. Extract a shared `query_all_pages` helper in `filesystem_conversations` that loops `query(..., Page::new(offset, MAX_LIMIT))` until a short page comes back, then reuse it from `filesystem_jobs::run_query` and from the inline scan in `update_estimation_actuals`. The helper preserves the existing `NotFound -> Vec::new()` short-circuit and the `fs_err_to_database` error mapping so call sites are otherwise unchanged. Regression tests: - `list_messages_drains_pages_beyond_max_limit` writes `MAX_LIMIT + 5` messages and asserts the full count round-trips through `list_conversation_messages` and that `list_conversation_messages_paginated` reports `has_more` honestly for both partial and exhaustive windows. - `get_job_actions_drains_pages_beyond_max_limit` writes `MAX_LIMIT + 3` actions on one job and asserts the full count comes back in sequence order. - `list_agent_jobs_drains_pages_beyond_max_limit` writes `MAX_LIMIT + 7` jobs and asserts both `list_agent_jobs` and `agent_job_summary` count every row. * fix(secrets): close CAS-loop races in filesystem store consume paths Two HIGH-severity findings on PR #3679. Both sites read a versioned entry, validated a one-shot/use-limit condition, then wrote back with `CasExpectation::Any`. The process-local mutex only serializes writers inside one process; multi-process callers sharing the same backend root could both pass the check and overwrite each other. - `FilesystemSecretStore::consume` — two consumers could both observe an Active one-shot lease, both decrypt, and both overwrite the consumed marker. - `FilesystemCredentialBroker::consume_session_use` — two consumers could both pass the max-uses check at `uses=N-1` and overwrite each other's increment, losing a use. Both now use the canonical retry-on-`FilesystemError::VersionMismatch` pattern from `ironclaw_engine::store::filesystem::update_thread_state` (post-`e2530adff`): re-read, re-evaluate the consume/use-limit condition, write with `CasExpectation::Version(versioned.version)`. A shared `CAS_RETRY_ATTEMPTS = 3` constant bounds the loop; exhausting it surfaces a transient backend error rather than papering over pathological hot-spots. Also annotated `leases_for_scope` with a `TODO(perf)` covering the N+1 list+get fan-out — bounded today by the owner-prefix path layout and short lease TTLs; replacing it with `Filter::Eq` over `query` requires the secrets store to declare its first index, which is a follow-up. Regression coverage: two new tests wrap `InMemoryBackend` with a `VersionRacingBackend` that bumps the watched path's version out-of-band on the first versioned `put`, forcing a `VersionMismatch` and exercising the retry loop. They also assert that the retried CAS write actually persisted (the next consume hits LeaseConsumed; the next three increments exhaust the max-uses budget). * fix: address remaining P2 review findings on #3679 Four P2 correctness fixes from the codex/gemini review. **Settings keys use percent-encoding** (`src/db/filesystem_settings.rs`): `encode_segment` previously mapped `/`, space, control chars, and others all to `_`. Keys like `a/b` and `a_b` collided onto the same path and silently overwrote each other. Now percent-encodes every byte outside the unreserved set so distinct inputs map to distinct outputs. **SQL index names get a blake3 suffix on overflow** (`crates/ironclaw_filesystem/src/db.rs`): `sql_index_name` truncated identifiers exceeding 62 chars without disambiguating, so two distinct long `(prefix, name)` specs could collapse onto the same DDL object — `CREATE ... IF NOT EXISTS` would silently reuse the wrong index/trigger. Now appends an 8-char blake3 hash suffix before truncating. Added `blake3 = "1"` to the crate's deps (small + already used by other workspace crates). **InMemoryBackend rejects writes over implicit directories** (`crates/ironclaw_filesystem/src/in_memory.rs`): the SQL backends refuse `put(/a)` when `/a/b` exists (treating `/a` as a directory). The in-memory reference impl silently accepted those writes, letting tests pass against production-impossible state. Mirror the SQL contract. **Event-store head-probe is bounded** (`crates/ironclaw_reborn_event_store/src/filesystem_store.rs`): The replay-gap detection previously called `tail(path, 0)` to read the whole log just to look at its last seq — O(N) on every cold-path call. Now probes `tail(path, after - 1)`: a non-empty result means head == after (consumer is caught up); empty means head < after (foreign-future cursor). Returns at most one record instead of the entire log. Verification: cargo check --workspace --all-features clean; cargo test -p ironclaw_filesystem -p ironclaw_secrets -p ironclaw_reborn_event_store --all-features all pass. * fix(filesystem): close cross-backend Range/vector semantic drift and txn scope hole Audit findings on ironclaw_filesystem turned up four bugs and three semantic-drift cases between the in-memory reference and the SQL backends. Fix them in one pass so the cross-backend contract is honoured and the gaps have regression coverage. Bugs: - libSQL `Filter::Range` on `IndexValue::Bool` never matched any row because SQLite's `json_type` returns "true"/"false" for booleans rather than "integer". Replaced the static type string with a `json_type_guard` expression that admits both bool variants. - `ScopedStorageTxn` did not enforce that per-op `VirtualPath`s lay under `mount_prefix`. The trait doc promised `PathOutsideMount` for cross-prefix accesses; the wrapper now enforces it so any future backend that ships `begin()` inherits the guarantee. - Mixed-variant `Filter::Range` bounds (e.g. I64 lo + Text hi) silently lex-compared on text on both SQL backends. Added the in-memory backend's `discriminant(lo) == discriminant(hi)` guard to both, rejecting with `Unsupported`. - SQL `vector_nearest_query` lacked the in-memory backend's path tie-breaker on equal cosine scores, so top-k truncation was non-deterministic. Added `.then_with(|| a.0.cmp(b.0))` to both. Semantic drift: - `FilesystemOperation` lacked an event-plane `Append` variant — default impl reported `Tail`, backends reported `AppendFile`. Added the variant, routed every emit site through it, and updated the downstream `host_runtime::operation_allowed` matcher. - `decode_embedding_blob` and `cosine_similarity` were byte-identical copies in three files. Extracted to `crate::vector`. - libSQL `run_migrations` ran multiple ALTERs outside any transaction. Wrapped the sequence in BEGIN IMMEDIATE / COMMIT with rollback on error so a crash can't leave a half-migrated schema observable. Tests added: - 16 `ScopedFilesystem` permission tests covering query / ensure_index / begin / append / tail across each `MountPermissions` axis, plus 4 `ScopedStorageTxn` tests driving a stub backend to lock in the per-op ACL and the new path-containment check. - Cross-backend regression tests in `tests/db_root_filesystem_contract.rs` for the libSQL Bool/Range fix, the discriminant guard on both SQL backends, and the deterministic vector tie-breaker. - Refactored `vector_nearest_query`'s phase-2 step into `materialize_ranked` (`pub(crate)`) so a unit test can exercise the "row disappeared between phases" branch deterministically. 128 tests pass, all three feature combos compile (`default`, `libsql`, `postgres`), workspace builds. * revert(db): drop filesystem-backed src/db/ store facades Removes all `src/db/filesystem_*.rs` facades and the `src/db/filesystem_workspace/` directory added during the PR #3679 universal-FS dispatch migration: - filesystem_conversations, filesystem_jobs - filesystem_routines, filesystem_sandbox, filesystem_tool_failures - filesystem_identities, filesystem_pairing, filesystem_settings, filesystem_users - filesystem_workspace/{mod,chunks,documents,paths,search,versions}.rs Also removes the supporting infra that only existed for these files: - `ironclaw_filesystem` workspace dep from the root `ironclaw` crate - `/sandbox`, `/routines`, `/tool_failures`, `/identities`, `/pairing`, `/workspace` entries from `ironclaw_host_api::path::VIRTUAL_ROOTS` The legacy libSQL/Postgres sub-trait impls (`src/db/postgres.rs`, `src/db/libsql/*.rs`) remain the sole backing for the `Database` supertrait. The unified `ironclaw_filesystem` mount fabric itself (the `crates/ironclaw_filesystem/` crate) is untouched and still used by consumer crates outside `src/db/`. Verification: - cargo fmt --check clean - cargo check --workspace clean (default features) - cargo check --no-default-features --features libsql clean - cargo check --all-features clean - cargo clippy --all --benches --tests --examples --all-features clean [skip-regression-check] pure removal of unmerged migration facades. * test(reborn-event-store): cover caught-up-to-head + concurrent appends Addresses audit finding F1. (a) `filesystem_event_log_caught_up_to_head_returns_empty_not_replay_gap` appends N events, replays from the last entry's cursor, and asserts `entries.is_empty()` + `next_cursor == last.cursor` with no `ReplayGap`. Pins the "consumer is caught up to head" branch of the bounded probe in `read_after_cursor`. (b) `filesystem_event_log_concurrent_appends_assign_distinct_cursors` spawns 8 `tokio::spawn` tasks each appending one event to the same stream, then asserts the collected cursors are pairwise-distinct and strictly increasing. Guards the per-stream monotonic-cursor invariant under contention. * fix(reborn-event-store): preserve filesystem error detail in durable mappers Addresses audit finding F2. `map_filesystem_append_error` / `map_filesystem_tail_error` previously collapsed every non-categorised `FilesystemError` variant (`VersionMismatch`, `NotFound`, `Backend`, …) to a fixed generic string, dropping the source variant and reason. Operators lost the detail they needed to debug appends that hit a CAS conflict or a backend I/O failure. Thread the underlying `FilesystemError` through its `Display` impl on the fallback arm. `FilesystemError` is already redaction-safe by contract — it renders scoped/virtual paths, never raw host paths — so the durable error surface gains debug detail without violating the crate-level redaction policy. The three already-categorised variants (`PermissionDenied`, `MountNotFound`, `Unsupported`) keep their fixed messages so callers can pattern-match on the substring. * fix(reborn-event-store): document deliberate absence of Filesystem config variant Addresses audit finding F3. `FilesystemDurableEventLog` / `FilesystemDurableAuditLog` are exported from this crate, but `RebornEventStoreConfig` has no corresponding `Filesystem` variant — so production composition still routes through the SQL stores. The PR description documents this as intentional: the filesystem-backed log is the migration target for the kernel-storage rework, and the config variant will be added during the `src/db/` dissolution pass (task #17). Without an inline comment, a future reviewer reading the config enum has no signal that the missing variant is deliberate. Add a doc paragraph on `RebornEventStoreConfig` pointing at the rationale on `filesystem_store.rs` and at task #17. * fix(reborn-event-store): drop shadowed kind named-arg in stream_path format! Addresses audit finding F4. `stream_path` previously used the named-argument `format!` form with `kind = kind_segment`, where the named key `kind` shadowed the function parameter of the same name. Switch to the implicit positional-capture form (`format!("/events/{kind_segment}/...")`) and rename the inline bindings to `tenant_segment` / `user_segment` for consistency. Pure refactor — no behaviour change, just removes the readability footgun. * fix(outbound): add typed CasConflict variant for filesystem store retries Audit finding F5: `map_fs_error` previously collapsed both `FilesystemError::VersionMismatch` (a transient compare-and-swap race condition that callers should retry) and `FilesystemError::Unsupported` (a permanent capability gap) into `OutboundError::Backend`. The bounded CAS retry loop (added separately for F1) cannot match on `Backend` — that would also retry on permanent backend failures and on `Unsupported` on backends that don't support CAS. Introduce `OutboundError::CasConflict` and map `VersionMismatch` to it in `map_fs_error`. The variant stays internal to the crate: the retry loop matches on it discriminator-wise; once the retry budget is exhausted (or for callers that haven't migrated) it converts to `Backend` before crossing the trait boundary, preserving the no-leak contract. Update `is_transient_validator_error` to classify `CasConflict` as transient for defence in depth, even though it should never reach the service boundary in practice. * fix(outbound): CAS-version read-then-write paths with bounded retry Audit finding F1 (HIGH): the four read-then-write methods on `FilesystemOutboundStateStore` (`upsert_subscription`, `advance_subscription_cursor`, `record_delivery_attempt`, `update_delivery_status`) read the existing entry, applied an in-memory transform, then wrote with `CasExpectation::Any`. Concurrent writers raced the transform: in particular, the "subscription cursor must not move backwards" invariant — enforced in `validate_advance_request` / `validate_subscription_cursor_progression` — was unenforced cross-process, because two racing advancers could both read the same old cursor, validate against it, and then both put their newer cursors, the loser silently winning the last-write race. Capture `VersionedEntry.version` from each `get`, pass `CasExpectation::Version(v)` to the matching `put`, and retry on the typed `OutboundError::CasConflict` introduced by F5. The retry budget is bounded (`MAX_CAS_RETRIES = 5`) and the loop re-reads + re-validates on every iteration, so a regressing cursor or scope mismatch surfaces immediately rather than letting the retry loop overwrite the winner's state. `put_thread_notification_policy` is a blind overwrite and keeps `CasExpectation::Any`. `record_delivery_attempt` uses `CasExpectation::Absent` for the first-write branch, so two racing at-least-once writers can't both insert; the loser falls back into the duplicate-identity-check branch on the next read. * fix(outbound): use control-character sentinel in thread scope key Audit finding F6: `thread_scope_key` used the literal string `"_"` as the sentinel for `agent_id = None` / `project_id = None`. The `validate_scope_id` validator in `ironclaw_host_api` accepts underscore as a legal character in an `AgentId` / `ProjectId`, so a scope with `agent_id = Some(AgentId::new("_"))` hashed to the same key as a scope with `agent_id = None`. Two distinct scopes silently collided on the same policy/subscription/delivery virtual path. Switch the sentinel to `"\x1F"` (ASCII unit-separator). It's a control character; `validate_scope_id` rejects every C0 control char via `has_forbidden_control`, so no legal scope id can ever contain it. Add a unit test that pins the sentinel-rejection invariant and a regression test that proves `agent_id = Some("_")` no longer hashes to the same key as `agent_id = None`. * fix(outbound): query indexed scope projection with paginated drain Audit finding F2 (HIGH): `list_delivery_attempts` was a `list_dir` + N+1 `get_json` per row with no indexed projection, scanning every delivery on the mount even when only one scope's deliveries were requested. Cost scaled with total delivery count, not with the queried scope's row count. Declare an exact-equality index on a new `scope` indexed key. The projected value is the same `thread_scope_key` hash used for policy paths — collision-resistant against the legal id grammar and updated by F6 to never collide with the `None` sentinel. `record_delivery_attempt` and `update_delivery_status` write through a new `put_delivery_attempt_indexed` helper that includes the projection; `update_delivery_status` preserves it on status mutations. The list path drives `query(Filter::Eq { key: "scope", value: ... })` and re-checks `scope_matches` defensively (hash collisions are unreachable but cheap to guard against). Audit finding F3 (Medium): the previous `list_dir` was unpaginated; SQL backends issue `LIMIT Page::MAX_LIMIT (1024)` on their list_dir translation and would silently truncate past 1024 deliveries. The new path drains pages via `offset += received` until a short page arrives, mirroring `ironclaw_engine::store::filesystem::query_all`. `ensure_delivery_scope_index` runs idempotently before every write and read. It tolerates `FilesystemError::Unsupported` on byte-only backends to match the engine store's `ensure_exact_index` pattern; the in-memory backend serves `Filter::Eq` from `Entry::indexed` directly even without a materialized index declaration. * test(outbound): cover CAS retry, pagination drain, backwards-race Audit finding F4: the existing `outbound_state_store_contract` suite exercised the storage contract surface but had no coverage for any of the failure modes the F1/F3 fixes address: - No CAS-retry test. F1's bounded retry loop could regress to permanent failure on any transient `VersionMismatch` and the suite wouldn't notice — the in-memory backend never produced one. - No `> Page::MAX_LIMIT` drain test. F3's pagination loop could lose the tail of a long delivery list and the suite wouldn't notice because the existing tests record at most one delivery per scope. - No concurrent backwards-race test on `advance_subscription_cursor`. The existing backwards-advancement test only exercised the single- threaded path; nothing proved the post-F1 retry loop re-validates progression on every iteration. Add three regression tests: 1. `VersionRacingBackend` wraps `InMemoryBackend` and injects a single `FilesystemError::VersionMismatch` on the next `put` matching a configured prefix. The first new test (`advance_subscription_cursor_retries_through_cas_conflict`) arms one conflict, advances the cursor, asserts the retry loop converges, and asserts exactly one conflict was injected and consumed. 2. `concurrent_backwards_race_rejected_after_winner_advances` runs two sequential advances — the winner to cursor=100 and the loser to cursor=50 — and asserts the loser is rejected with `InvalidRequest` while the winner's state is preserved. Together with the retry test this proves the re-validate-on-retry semantics F1 calls out. 3. `list_delivery_attempts_drains_more_than_page_max_limit` writes `Page::MAX_LIMIT + 1` delivery attempts under one scope and asserts `list_delivery_attempts` returns every one. Before F3 this would silently truncate at 1024 rows. Cargo.toml: enable `tokio/sync` for `Mutex` in the test mock; drop the feature-conditional `use std::sync::Arc` because the new tests need it unconditionally. * fix(run-state): bound filesystem lock map under tenant churn The process-wide FILESYSTEM_RECORD_LOCKS map kept one Arc<tokio::sync::Mutex<()>> per touched path. In long-running hosts with high tenant/invocation churn the map grew without bound, since entries were never removed once the originating put/get cycle completed. Switch the value type to Weak<Mutex> so dropped Arcs no longer pin map slots. Each acquisition opportunistically prunes dead entries before upgrading-or-installing, keeping the map size proportional to in-flight paths rather than to lifetime path count. Concurrent callers on the same path still observe the same Arc (the outer std::sync::Mutex serializes the upgrade-or-insert window), so existing intra-process and cross-instance serialization guarantees are preserved — both verified by the new unit tests and by the existing filesystem_*_duplicate_*_serialized_across_store_instances contract tests. Addresses audit findings F1 (Medium) and F4 (Low). * fix(run-state): use versioned CAS for filesystem run/approval writes All filesystem put() calls used CasExpectation::Any, so two host processes mounting the same /engine could lose updates: each one's read-modify-write saw the other's value and then unconditionally overwrote it. The per-path async mutex only serializes intra-process callers. Switch creates to CasExpectation::Absent and updates to CasExpectation::Version(v) with a bounded retry loop on VersionMismatch. The new put_with_cas helper centralizes the contract: on capable backends (InMemoryBackend, the upcoming SQL ports) cross-process races now fail closed and the caller retries; on byte-only backends that return Unsupported (LocalFilesystem) we degrade to Any but emulate Absent with a get() precheck so the AlreadyExists path is preserved. The in-process lock map (F1) keeps the check-then-write race closed for the byte-only fallback. Approve/deny/discard pull the record-lock guard up to the trait method, since update_status no longer acquires it. Addresses audit finding F2 (Medium). Closes the gap acknowledged in crates/ironclaw_run_state/CLAUDE.md. * fix(approvals): type approval-resolution decision with ApprovalDecisionKind enum Addresses audit finding F1. Replaces the stringly-typed `impl Into<String>` decision parameter on `AuditEnvelope::approval_resolved` with a wire-stable `ApprovalDecisionKind` enum (`Approved`/`Denied`, `#[serde(rename_all = "snake_case")]`), so approval callers cannot drift on capitalization or spelling. Per `.claude/rules/types.md` "wire-stable enums". The wider `DecisionSummary::kind` field stays a `String` because other audit producers (authorization denials, obligation handlers) emit values outside the approval enum; cross-decoding remains a follow-up. Cross-crate blast radius: `ironclaw_host_api` (new enum + factory signature), `ironclaw_approvals` (both call sites), `ironclaw_events::tests::durable_log_contract` (three test fixtures). * fix(approvals): persist approval state before issuing lease Addresses audit finding F2. Inverts the lease/approve ordering inside `approve_capability_action`: the approval store write now runs *before* the lease store write. The previous order (issue lease, then approve, best-effort revoke on failure) left a window where a transient approval-store error could leave a live lease pointing at a request whose status remained `Pending`. The approval record is now treated as the authority of record. Once the request flips to `Approved`, lease issuance is a recoverable operation against an already-decided request — if the lease store fails, the caller surfaces the lease error and the request stays `Approved`. The previous best-effort `let _ = self.leases.revoke(...)` swallow is gone with the same edit. Updates the three concurrency/error-injection tests to assert the new semantics, plus the crate CLAUDE.md guardrail. No external test fixtures break — the public resolver API is unchanged. * fix(approvals): route both resolve paths through emit_approval_resolved helper Addresses audit finding F3. Extracts an `emit_approval_resolved` helper on `ApprovalResolver` so the audit-envelope construction in `approve_capability_action` and `deny` is built in exactly one place. Both call sites used to inline `AuditEnvelope::approval_resolved` against their own `record.scope`/`denied.scope`; while consistent today, divergence between the two would be a silent regression. Pure refactor — no test changes needed beyond the existing audit-event contract tests which already pin the wire shape. * fix(approvals): cover concurrent approve_dispatch first-write-wins Addresses audit finding F4. Adds a caller-level concurrency regression test that spawns two `approve_dispatch` calls against the same pending request on a multi-thread tokio runtime and asserts the expected first-write-wins invariants: - exactly one approve returns `Ok` - the other returns `ApprovalResolutionError::NotPending { status: Approved }` - the lease store ends up with exactly one Active lease (not two, not zero — under the F2 persist-approval-first ordering the loser fails *before* lease issuance, so no orphan to revoke) - the approval record's terminal status is `Approved` Enables `rt-multi-thread` on the tokio dev-dependency so the test can exercise real cross-thread contention on the approval store mutex. * fix(engine): restore HybridStore parity for mission updates F1: `update_mission_status` now bumps `mission.updated_at` before writing back, matching HybridStore (`src/bridge/store_adapter.rs:1950`). Recency-sorted views (mission list UIs, learning-mission dispatcher) were silently freezing the timestamp at original-save time. F2: `list_missions` and `list_all_missions` now sort by `(name, id)` after collection, matching HybridStore (`store_adapter.rs:1913, 1937`). The underlying `query`/HashMap iteration is non-deterministic; the LLM-facing `mission_list` tool was seeing arbitrary order across runs. Tests: - `update_mission_status_bumps_updated_at` — regression for F1 - `list_missions_is_deterministic_across_invocations`, `list_all_missions_is_deterministic_across_invocations` — regression for F2 * fix(memory): drain pages in FilesystemMemoryDocumentRepository::list_documents Audit findings F1 (HIGH) + F9 (Low). F1: `list_documents` issued a single `query(.., Page::new(0, Page::MAX_LIMIT))` and trusted the page was complete. Because `Page::MAX_LIMIT == 1024`, scopes holding >1024 documents silently lost every entry past the cap. The result fed `write_document`'s ancestor/descendant conflict check at the call site immediately above, so a new path could shadow (or be shadowed by) an existing document across the truncation boundary without a conflict ever firing — exactly the regression `query_all_pages` was extracted in `src/db/filesystem_jobs.rs` to prevent. F9: The old implementation issued a `Filter::All` query, threw the results away (`let _ = (versioned, &prefix_str);`), then called `list_dir` to discover paths. The query-result loop was dead code under any backend that supports `query`. The stale comment claimed the trait didn't surface paths in `query` results, but `VersionedEntry.path` (`crates/ironclaw_filesystem/src/record.rs:347`, added in PR #3659) has carried the absolute virtual path for every queried row since. Replace both with a single drain loop that paginates `query` until a short page comes back, filters by `entry.kind == "memory_document"`, and recovers the `MemoryDocumentPath` directly from `VersionedEntry.path`. The `list_dir` fallback is gone, and the agent_id axis is preserved through `MemoryDocumentPath::new_with_agent` so scopes with an agent identity round-trip correctly (the previous code's `new()` dropped the agent). Regression: `list_documents_drains_pages_beyond_max_limit` writes `MAX_LIMIT + 5` documents and asserts every one comes back. This also exercises the conflict-check path because each `write_document` calls `list_documents` internally. * fix(secrets): close consume_if_matches timing oracle with constant-time compare F1 (HIGH, timing oracle) in the 2026-05 audit: `consume_if_matches` in `legacy_store.rs` (trait default + in-memory backend) and `db.rs` (libSQL + Postgres backends) compared the decrypted plaintext against the caller-supplied expected value with `!=`. Rust's `!=` over `&str`/`&[u8]` short-circuits on the first differing byte, so an adversary who can observe response latency over the network can recover the secret byte by byte. AES-GCM authenticated decrypt closes the ciphertext oracle but does nothing for the post-decrypt comparison. Fix: route the comparison through `subtle::ConstantTimeEq::ct_eq`, which walks the full buffer regardless of where the bytes diverge. The post-comparison branches retain their original shape because the decrypt+lookup path is already executed unconditionally before the compare — only the success-side `DELETE` differs, and that signal is already exposed by the function's return value. Added a regression test (`f1_consume_if_matches_uses_constant_time_compare`) that grep-asserts the production source imports `subtle::ConstantTimeEq`, uses `ConstantTimeEq::ct_eq`, and no longer contains the legacy `!=` shape. Cannot meaningfully prove constant-time-ness from a shared CI runner, but the source-pattern check ensures a "simplifying" revert fails review. Audit: F1 (HIGH). * fix(secrets): use constant-time compare for store key-check sentinel F3 (Low) in the 2026-05 audit: `verify_secret_store_key_check` compared the decrypted sentinel against `SECRET_STORE_KEY_CHECK_PLAINTEXT` with `!=`. The plaintext is a fixed compile-time string so the practical risk is low — an attacker who can move the encrypted_value/key_salt blobs across rows already has full DB write access — but the same constant-time pattern applied to F1 makes the comparison style consistent across the crate and pre-empts a future caller threading a non-constant sentinel through this helper. Routed through `subtle::ConstantTimeEq::ct_eq`, mirroring the F1 fix. Audit: F3 (Low). * fix(processes): index queryable fields and serve records_for_scope via query Replace the N+1 list_dir + per-file get scan with an indexed `query` path, falling back to the legacy scan on byte-only backends so existing LocalFilesystem-driven tests and production deployments remain unaffected. - Declare `ensure_index` lazily for the per-owner `processes/` prefix on the queryable fields called out in the audit (`tenant_id`, `user_id`, `status`, `extension_id`, `parent_process_id`). Backends without index support degrade to the existing scan instead of failing closed. - Project the same fields onto every `ProcessRecord` write via `Entry::with_indexed`; record-capable backends (libSQL, Postgres, the in-memory backend) can now serve scope listings through a native query. The opaque-byte fallback in `put_with_byte_fallback` keeps LocalFilesystem (which rejects record-shaped puts today) on the legacy write path. - Rewrite `records_for_scope` to issue `Filter::And` of `Filter::Eq` predicates against the indexed projection. The full `same_scope_owner` check remains in Rust so the sub-scope axes (agent/project/mission/ thread) that are not yet in the index spec still get filtered. - Add a contract test that exercises the indexed path through `InMemoryBackend` and confirms cross-tenant and cross-user records are not returned. Addresses audit findings F1 (records_for_scope N+1) and F2 (missing ensure_index at startup). * fix(filesystem): surface backend infrastructure errors without fabricated paths F1: SQL backends used valid_engine_path() (unwrap_or_else unreachable returning /engine) as a placeholder on every connection/migration error. The path was always a lie - at pool acquisition, run_migrations, pragma setup, or schema bootstrap there is no caller-supplied virtual path in scope - and it leaked into operator-facing error display. Add FilesystemError::BackendInfrastructure { operation, reason } that omits path. Route every former valid_engine_path() callsite in libsql and postgres through new infrastructure_error helpers in db.rs. The enum is non_exhaustive so adding a variant is backward compatible. Regression test: drive a libsql migration against a read-only DB file and assert BackendInfrastructure with no /engine in display. * fix(filesystem): store VirtualPath keys in InMemoryBackend state directly F2: in_memory.rs::query() reparsed every stored row's path with VirtualPath::new(...).unwrap_or_else(|_| unreachable!('stored paths originated as VirtualPath')) on the hot path. Two issues: - the reparse is wasted work - paths originate as VirtualPath at put() time, so the validation pass on read is redundant - 'unreachable!' is a panic that asserts a structural invariant the type system already enforces Replace HashMap<String, StoredEntry> with HashMap<VirtualPath, StoredEntry>. Lookups now pass &VirtualPath directly; prefix scans move to key.as_str().starts_with(...). VersionedEntry::path comes from a single clone() instead of a parse + unreachable. Existing tests cover the put/get/query/list_dir/stat/delete paths that were touched (44 in_memory tests + the cross-backend contract suite). * fix(filesystem): align in-memory backend on nested VectorNearest semantics F5: SQL backends reject Filter::VectorNearest nested inside And/Or with Unsupported because ranking can't be expressed as a WHERE fragment - the top of query() peels off a top-level VectorNearest before the translator runs, and the translator's VectorNearest arm unconditionally errors. The in-memory backend previously treated a nested VectorNearest as 'any row with IndexValue::Bytes at key', silently changing semantics across backends. Add contains_nested_vector_nearest() pre-check in InMemoryBackend:: query that walks the filter tree and surfaces Unsupported for any VectorNearest strictly inside a compound. The Filter::VectorNearest arm in filter_matches is now unreachable; it returns false to keep the scalar predicate path safe should the pre-check ever be bypassed. Regression test asserts Unsupported on nested-in-And, nested-in-Or, and still-OK for top-level VectorNearest. * fix(filesystem): guard u64 to i64 SQL bindings with typed errors F6: SQL backends used 'expected.get() as i64' and 'page.offset as i64' casts on the CAS and query/pagination paths. Both inputs are u64 and both wrap silently on values >= 2^63 - the cast produces a negative SQL binding that either matches no row (CAS quietly VersionMismatches) or executes against a negative OFFSET (cryptic backend error). Add db.rs helpers: - record_version_to_i64: surfaces CorruptRecordVersion if the value overflows i64 - page_offset_to_i64: surfaces a typed Backend error naming the operation and offset Apply at libsql.rs CAS and query offset bindings and the matching postgres.rs sites. 'page.limit' is u32 clamped to Page::MAX_LIMIT so its i64 cast is safe by construction and uses i64::from for clarity. Regression test asserts a typed Backend(Query) error with reason 'page offset...' when querying with offset = u64::MAX, replacing the prior silent wrap. * fix(filesystem): scope Postgres FTS GIN index to declaring prefix F4: libsql FTS5 virtual tables are declared per-mount-prefix - one vtable per ensure_index(prefix, ...) call - so a query at one prefix can't accidentally pull index postings from a sibling prefix into the plan, and tearing down an index for a prefix is a clean DROP TABLE. The Postgres FTS GIN index, by contrast, was created without a predicate over root_filesystem_entries, so it was global. Correctness held because the query path always scopes by 'path = OR path LIKE ', but parity with libsql broke in two ways: the planner considered postings from every prefix before filtering, and a per-prefix DROP INDEX could only ever tear down one of them. Add a partial-index predicate gated by 'path = <prefix> OR path LIKE <prefix>/%' to the GIN DDL. The prefix is sourced from the validated VirtualPath and quotes are doubled for safe SQL literal embedding; LIKE-special characters are escaped via the existing escape_like_with_trailing_wildcard helper. Regression test (Postgres only; skipped when no DB is reachable) reads back the DDL via pg_indexes.indexdef and asserts the prefix literal and a WHERE clause appear. * fix(filesystem): tighten capability docs, type constraints, and hygiene nits Batched audit findings: F3: Document the type constraint on IndexKind::Prefix. The kind is only meaningful against IndexValue::Text, but ensure_index can't see the value type at declaration time. Filter::PrefixOn rejects every non-text variant at query time. Document the constraint loudly so consumers reach for IndexKind::Exact when projecting numeric or boolean values instead of getting an unused index and a query-time Unsupported. F7: BackendCapabilities::sql_typical advertises a minimum SQL shape that omits IndexFts and IndexVector. The two real backends here (libsql + postgres) layer them on top. A hand-rolled backend that just calls sql_typical() would under-advertise. Add a doc-comment calling out the omission and an sql_typical_full() variant that includes Events + IndexFts + IndexVector for backends that match this crate's shape. F8: validate_simple_identifier indexed bytes[0] after an is_empty guard. The guard makes the index sound, but the pattern is fragile to refactors. Switch to bytes.first() so the dependency is explicit and the panic path goes away. F9: Multiple doc comments in record.rs and index.rs referenced stale type names (StorageBackend::put/list/query, Record). Update to the current RootFilesystem / Entry names. * fix(engine): dedupe events on append_events for HybridStore parity HybridStore (`src/bridge/store_adapter.rs:1613`) de-duplicates thread events by id before insert. The filesystem-store `append_events` impl was previously writing with `CasExpectation::Any`, which silently overwrote an existing event with the same id when callers re-emitted (e.g. recovery after a partial flush). Pre-read the destination path and skip any id already present. Matches HybridStore's append-only contract. Audit finding F3 (Medium) from the ironclaw_engine crate audit. * fix(memory): map FTS results to documents in FilesystemMemoryDocumentRepository::search_documents The previous scaffold issued the `Filter::Fts` query, then silently dropped the results with `let _ = results; Ok(Vec::new())`. A caller wiring up the trait would see an empty result set and assume "no matches" — when in fact the search had simply lied. That is worse than returning `Unsupported`. Map each `VersionedEntry.path` (added in PR #3659) back to a `MemoryDocumentPath`, de-dupe by path, and assign a per-rank score from RRF over the FTS-only branch so the result vector matches the native repos' fusion contract for the trivial single-branch case. Skip non-memory-document entries that may live under the same prefix (chunk projections, metadata siblings). Adds `list_documents_drains_pages_beyond_max_limit` test against the in-memory backend. Audit finding F2 (HIGH) from the ironclaw_memory crate audit. * fix(secrets): close revoke CAS-loop race with versioned compare-and-swap `revoke` previously read the lease via the (now-removed) `read_lease` helper and wrote with `CasExpectation::Any`. The per-lease process-local mutex serialized writers within one process only — multi-process callers sharing the same backend root could observe `Active`, race against `consume`, and clobber a `Consumed` marker by overwriting it with `Revoked`. Inline the read into a bounded CAS retry loop matching `consume` and `consume_session_use`: read with version, write with `CasExpectation::Version`, retry on `VersionMismatch`. Make revoke idempotent on terminal states (`Consumed`, `Revoked`, `Expired`) so the loop converges even when a winner has already written. Audit finding F2 (Medium) from the ironclaw_secrets crate audit. * fix(processes): use versioned CAS for status transitions `update_status` previously read the record and wrote with `CasExpectation::Any`, relying on the per-instance `transition_lock` for atomicity. That lock only serializes within one process; a multi-process deployment sharing the same backend root could observe identical pre-transition state in both processes and clobber each other's status flips. Replace with a bounded CAS retry loop: read with version, validate the transition, write with `CasExpectation::Version…
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
… loops (nearai#3880) * refactor(secrets): extract cas_mutate helper for filesystem CAS retry loops `consume`, `revoke`, and `consume_session_use` were three near-identical CAS retry loops (get versioned record → check invariants → mutate → put with `CasExpectation::Version` → retry on `VersionMismatch`). The `revoke` body explicitly called out the duplication ("F2 (Medium) in the 2026-05 audit") and shipped the third copy anyway. Implements finding #3 from the recent thermo-nuclear review. Now: one `cas_mutate` helper owns the retry loop, the versioned-get, the serialize-and-CAS-write, and the exhaustion error. Each call site supplies a `CasMutateOps<T, E>` type-witness (deserialize, serialize, fs-error map) and a `decide` closure that returns a `CasDecision` variant per iteration: * `Commit { record, value }` — write with CAS; on VersionMismatch re-read and re-decide; on success return Ok(value). * `BestEffortCommit { record, outcome }` — write with CAS but treat VersionMismatch as success. Used for the stale-Active → Expired promotion in `consume`, where a peer racing us to the Expired marker is observationally identical. * `Settle(outcome)` — short-circuit without writing. Used for unknown records, scope mismatches, already-terminal states, validation failures. Behavior preserved end-to-end: every edge case the previous code handled (lease-already-Consumed/Revoked/Expired short-circuits, revoke-on-terminal idempotency, session use-limit overflow, expiry promotion + LeaseExpired error, CAS retry budget exhaustion) maps to exactly one `CasDecision` variant. The five existing edge-case tests all pass: * filesystem_secret_store_consume_retries_on_version_mismatch * filesystem_broker_consume_session_use_retries_on_version_mismatch * filesystem_secret_store_revoke_blocks_consume * filesystem_secret_store_revoke_after_consume_is_idempotent * filesystem_secret_store_revoke_is_idempotent_on_revoked `cargo fmt --check`, `cargo clippy --all --benches --tests --examples --all-features`, and `cargo test -p ironclaw_secrets --all-features` all green. File size grew slightly (+89 LOC) because the helper carries its own documentation; the structural win is that the CAS retry logic now lives in one place. Adding a fourth CAS-mutated record type is now a ~25-line closure instead of a 50-line copy of the loop body. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(secrets): address PR nearai#3880 review — Send bounds, op string, CAS tests - Restore the original `consume_session_use` retry-exhaustion reason string (`"credential session use retry limit exceeded"`); the refactor had silently lengthened it to `"credential session consume use ..."`. - Add explicit `Send` bounds on `cas_mutate`'s `Decide`, `NotFound`, and `RetryExhausted` closure params. The future was already Send via the `#[async_trait]` callers; the explicit bound surfaces future non-Send regressions at the helper site rather than at the call site. - Add caller-level coverage for the paths the refactor moved onto the shared helper but the existing CAS tests did not exercise: `revoke` retrying past a single `VersionMismatch`, and the CAS retry budget being exhausted for both `revoke` and `consume_session_use`. The exhaustion tests pin the caller-visible reason strings so future refactors cannot silently drift them. - Introduce `AlwaysRacingBackend` test helper alongside the existing one-shot `VersionRacingBackend` to drive the exhaustion paths. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…earai#3573) * feat(reborn): add ironclaw_hooks framework foundation (#3524) Foundation slice of the Reborn loop hooks framework per nearai/ironclaw#3524. Lands the trust primitives, sealed decision types, dispatcher contract, and extension manifest schema; no Reborn middleware composition yet (next slice wires HookDispatcher into LoopCapabilityPort / LoopPromptPort). Design comment on #3524: https://github.com/nearai/ironclaw/issues/3524#issuecomment-4439890144 What this PR ships ================== * `crates/ironclaw_hooks/` — new crate * `identity` — content-addressed `HookId` (blake3 of length-prefixed extension + local + version fields). Same versioning primitive the rest of Reborn should converge on for replay safety. * `trust` — `HookTrustClass` enum (Builtin / Trusted / Installed) with per-kind default attenuation. Trust class is fixed by source, never declarable. * `kinds/` — sealed decision DTOs. `BeforeCapabilityHookDecision`, `HookPatch`, `ObserverFact` all have `pub` outer struct + `pub(crate)` inner enum + `pub(crate)` constructors. Same #3460 witness pattern. * `points/` — typed read-only contexts for each hook point. * `sink` — split sink traits per trust tier. `PrivilegedGateSink` exposes `allow()`; `RestrictedGateSink` does not. An Installed-tier hook literally cannot mint Allow at the type level. * `ordering` — phase → priority → hook id, stable. Phases gated by trust (Validation/Authorization Builtin-only). * `failure_policy` — Timeout/Panic/Malformed/AttenuationViolation categories. Gate/Mutator fail closed, Observer/Effect fail isolated. Slot poisoning persisted for the rest of the run on any category. * `registry` — run-profile-sourced bindings; phase-vs-trust gate enforced at insert; poisoning surface for the dispatcher. * `dispatch` — HookDispatcher with deterministic ordering, panic catch-unwind via futures::FutureExt, per-hook tokio::time::timeout, short-circuit gate composition (Deny > PauseAuth > PauseApproval > Allow), Telemetry-phase observers always run. * `manifest` — serde types for the `[[hooks]]` section of extension manifests. Predicate vs WASM body; same_tenant scope requires explicit grant; Validation/Authorization phases rejected at parse time because manifest hooks are always Installed. * `predicate` — typed predicate language for declarative Installed hooks (DenyCapability, PauseApproval, RateOrValueCap). Evaluator lives in the dispatcher follow-up, not here. * `crates/ironclaw_architecture/tests/reborn_dependency_boundaries.rs` * Added `ironclaw_turns` -> `ironclaw_hooks` to the forbidden list. * New BoundaryRule for `ironclaw_hooks` itself (cannot pull host_runtime, dispatcher, secrets, network, wasm, etc.). * `Cargo.toml` workspace member registration. What this PR deliberately does NOT ship ======================================== * Reborn middleware composition wrapping LoopCapabilityPort / LoopPromptPort with HookDispatcher. Next slice; ironclaw_reborn changes only. * WASM hook execution path. Programmatic hooks parse and validate from manifest; the wasmtime integration lands when the WASM dispatcher seam is built. * Predicate evaluation. Predicate types serialize and validate; the evaluator that turns a `RateOrValueCap` spec into a `Deny` decision is in the next slice alongside Reborn wiring. * Event-triggered hooks (Phase 5 of the original roadmap). * Self-authored hooks. Tracked separately at #3567 with monotonic-restriction + unforgeable-channel ratification. Test plan ========= * `cargo test -p ironclaw_hooks` — 47 tests (46 unit + 1 integration smoke for the manifest -> binding -> dispatch pipeline). * `cargo test -p ironclaw_architecture` — 13 tests; new boundary rule passes, existing rules unaffected. * `cargo clippy -p ironclaw_hooks --all-targets -- -D warnings` — clean. * `cargo fmt -p ironclaw_hooks -- --check` — clean. * `cargo check --workspace` — clean, no regressions in other crates. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): wire HookDispatcher into LoopCapabilityPort/LoopPromptPort Follows the foundation slice (see initial commit). Adds the next layer: 1. Capability- and prompt-port middleware (`ironclaw_hooks::middleware`) * `HookedLoopCapabilityPort` runs `dispatch_before_capability` before every invocation, translates the composed decision into the existing `CapabilityOutcome` vocabulary (Deny / PauseApproval / PauseAuth all map to `Denied` for now; gate-ref plumbing for real pause semantics lands in the next slice). * `HookedLoopPromptPort` runs `dispatch_before_prompt` before bundle construction. Observe-only for snippets in this slice; actual snippet injection waits for the shared `prompt_envelope::wrap_untrusted` helper (#3540 / #3471). 2. Declarative predicate evaluator (`ironclaw_hooks::evaluator`) * `DenyCapability` and `PauseApproval` predicates: stateless, evaluated directly against `BeforeCapabilityHookContext`. * `RateOrValueCap` with `InvocationCount` bound: sliding-window counter keyed by `(hook_id, capability_name)`, in-memory only. Window parsing supports `s`/`m`/`h`/`d` units; unparseable windows fail closed. * `NumericSum` bound: types implemented but evaluation returns Allow and emits a warn-level audit. Full argument-extraction story is a follow-up slice once capability arguments become hook-visible. * `PredicateEvaluator::evaluate_at(...)` test variant accepts an explicit `Instant` so sliding-window tests don't depend on real-clock progress. 3. Manifest -> dispatcher glue (`ironclaw_hooks::installed_hook`) * `PredicateBackedBeforeCapabilityHook` wraps a `HookPredicateSpec` plus an `Arc<PredicateEvaluator>` and implements `RestrictedBeforeCapabilityHook`. The registry installer would construct one of these per `[[hooks]]` entry whose body is `HookManifestBody::Predicate`. * Sink reasons are `&'static str`, so the dynamic predicate `reason` surfaces in audit (via the evaluator's `EvaluatorDecision`) rather than the model-visible decision. Closed-vocabulary labels carry through to the sink. 4. Reborn composition seam (`ironclaw_reborn::loop_driver_host`) * `RebornLoopDriverHostFactory::with_hook_dispatcher(Arc<HookDispatcher>)` opt-in builder method. When set, the factory wraps the capability and prompt ports with the hooked middleware. Default behavior (no dispatcher) is unchanged from the pre-hooks shape, so existing callers continue to work. * Added `ironclaw_hooks` as a dep in `ironclaw_reborn`. Test plan ========= * `cargo test -p ironclaw_hooks` — 60 tests pass (59 unit + 1 integration smoke; +13 vs the foundation commit covering middleware, evaluator, installed_hook). * `cargo test -p ironclaw_reborn` — 118 tests pass; no regressions from adding the dep. * `cargo test -p ironclaw_architecture` — 13 tests pass; the `ironclaw_turns -> ironclaw_hooks` boundary still holds and the new `ironclaw_hooks` rule (no host_runtime / dispatcher / secrets / network / wasm / reborn) is unaffected. * `cargo clippy -p ironclaw_hooks --all-targets --all-features -- -D warnings` — clean. * `cargo clippy -p ironclaw_reborn --all-targets -- -D warnings` — clean. * `cargo fmt --all -- --check` — clean. What still defers ================== * WASM hook execution path. * Persistent predicate counter (in-memory only for now). * Argument-extraction so `NumericSum` predicates evaluate against capability arguments. * Gate-ref plumbing so PauseApproval / PauseAuth surface real `CapabilityOutcome::ApprovalRequired` instead of `Denied`. * Prompt-snippet injection (waits for shared envelope helper). * Event-triggered hooks. * Self-authored hooks (#3567). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): add HookedLoopModelPort/TranscriptPort/CheckpointPort observer middleware Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(reborn): end-to-end hooks integration through RebornLoopDriverHostFactory Adds crates/ironclaw_reborn/tests/hooks_integration.rs covering the factory's HookDispatcher wiring seam end-to-end. Tests drive host.invoke_capability(...) (not dispatcher.dispatch_before_capability(...) directly) so a regression in RebornLoopDriverHostFactory's wrapping composition surfaces here. Scenarios: - PredicateBackedBeforeCapabilityHook (DenyCapability NameEquals "cap.blocked") short-circuits invocation; inner port never called; outcome is Denied(unknown("hook_denied")). - A privileged selective hook that allows non-matching capabilities proves the wrapper does not blanket-deny: cap.allowed reaches the inner port and completes once. - Factory built without with_hook_dispatcher() lets cap.blocked through to the inner port, proving the hook plumbing is genuinely opt-in. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): add pass() + HookRegistrar + self-authored hooks scaffolding Three additions to ironclaw_hooks: B. `pass()` on gate sinks — distinguishes "evaluated, no opinion" from "returned without minting a decision." A passing hook contributes nothing to the composed decision; a silent hook is still Malformed and fails closed. `PredicateBackedBeforeCapabilityHook` now routes the evaluator's `Allow` decision through `sink.pass()` instead of the previous `deny("hook_predicate_pass")` workaround. A. `HookRegistrar` bridge — converts a `Vec<HookManifestEntry>` into `HookBinding`s + dispatcher impls in one call. Predicate bodies are wired through `PredicateBackedBeforeCapabilityHook`; WASM bodies return `HookError::RegistryConstruction` for now. Adds `HookDispatcher::insert_binding` so the registrar can mutate the registry through the dispatcher rather than reach inside. I. Self-authored hooks scaffolding — fourth `HookTrustClass` variant for hooks the agent authors at runtime. Run-scoped only; monotonic-restriction sink with no `allow`, no trusted-snippet path, no effect-class constructor. Closed-vocabulary `SelfAuthoredReason` enum keeps free-text reasons off the audit seam. `SelfAuthorshipProvenance` captures authoring run/turn, timestamp, spec digest, optional user ratification, and a generation-trace pointer. Durable persistence depends on the unforgeable channel from #3564 and lands separately. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): real gate-ref plumbing for hook PauseApproval/PauseAuth decisions Previously, `GateDecisionInner::PauseApproval` and `PauseAuth` returned by hooks were degraded to `CapabilityOutcome::Denied` at the middleware boundary because the hook crate had no way to mint a `LoopGateRef` scoped to the current run. Hooks that wanted to pause the loop for approval or auth instead failed the call closed, leaving the host's approval-router machinery unreachable from hook code. This change introduces a `HookGateRefFactory` trait in `ironclaw_hooks::middleware::gate_ref` that mints `LoopGateRef`s for pause-class decisions. `HookedLoopCapabilityPort` now takes an `Arc<dyn HookGateRefFactory>`, defaulting to `UuidHookGateRefFactory` (a locally-unique opaque-id factory suitable for tests and the foundation slice). Production deployments override via `.with_gate_ref_factory(...)` with a factory bound to the current `LoopRunContext` and the host's gate-router. The translation in `decision_to_outcome` is now async so it can await the factory. `PauseApproval` maps to `CapabilityOutcome::ApprovalRequired { gate_ref, safe_summary }` and `PauseAuth` to `AuthRequired`. If the factory itself errors, the middleware falls back to `Denied` with a sanitized `hook_gate_ref_unavailable` reason kind so the loop fails closed rather than routing through an unresolvable suspension. The underlying error text is dropped to avoid leaking gate-router state into model-visible output. Tests: - `pause_approval_decision_surfaces_as_approval_required`, `pause_auth_decision_surfaces_as_auth_required`, `gate_ref_factory_failure_falls_back_to_denied` in `middleware::capability_port::tests`. - `pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref` in `crates/ironclaw_reborn/tests/hooks_integration.rs`, exercising the full `RebornLoopDriverHostFactory` composition with the default `UuidHookGateRefFactory`. - Gate-ref factory unit tests in `gate_ref::tests`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): add NumericSum predicate evaluation with capability argument extraction Wires the missing argument-extraction story for the predicate evaluator so `ValueOrRateBound::NumericSum` actually enforces a rolling numeric cap instead of warn-and-allowing. - Extend `BeforeCapabilityHookContext` with a sealed `SanitizedArguments` view. Strings truncate to 256 bytes; objects/arrays cap at 8-deep. `extract_numeric` supports dotted + bracketed paths (`order.amount`, `items[0].price`) and returns `Option<rust_decimal::Decimal>`. The inner representation is sealed so external callers can't bypass bounds. - Introduce `CapabilityInputResolver` + bundled `NullCapabilityInputResolver` in `middleware/resolver.rs`. The hooks crate intentionally doesn't know how to dereference a `CapabilityInputRef` — that knowledge belongs to the production host. Until a real resolver is wired in (follow-up), arguments are `Unresolved` and `NumericSum` fails closed. - `HookedLoopCapabilityPort::new` defaults to the null resolver; new builder `.with_resolver(Arc<dyn CapabilityInputResolver>)` overrides. - `PredicateEvaluator` gains a tenant-keyed `value_history` map. The `NumericSum` arm parses `max` + `window`, extracts the numeric value from sanitized args, accumulates within the rolling window, and applies `on_exceeded` when the sum exceeds the cap. Unresolved args, missing field, non-numeric field, unparseable max, and unparseable window all fail closed via the configured `OnExceededAction`. - Add `BeforeCapabilityHookContext::new_unresolved(...)` convenience ctor; existing test sites switch to it instead of churning every call site through the 4-arg ctor. Test count: +14 (8 new SanitizedArguments tests, 6 new NumericSum evaluator tests, 1 null-resolver test; one old NumericSum-stub-related gap closed). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(reborn): seal hook registration trust boundary + dispatcher hardening Addresses blocking findings from the security audit of `ironclaw_hooks`: - C1 (Blocking, Trust Model): "Installed cannot Allow" was not enforced at the registration boundary. `BeforeCapabilityHookImpl::Privileged` was a public variant, so external crates with dispatcher access could construct an Installed binding paired with a Privileged impl and bypass the sink trait restriction. Sealed `BeforeCapabilityHookImpl`, `BeforePromptHookImpl`, and `ObserverHookImpl` to `pub(crate)` and replaced the single generic `install_before_capability` / `install_before_prompt` / `install_observer` surface with tier-specific public installers (`install_builtin_*`, `install_trusted_*`, `install_installed_*`) that build the binding with the matching trust class internally. Updated registrar, internal middleware tests, the hooks foundation pipeline test, and the reborn `hooks_integration` test to drive the new surface. Added regression tests proving the trust class is set by the installer and that the seal is type-level. - C5 (Medium, Slot Poisoning): same-dispatch poisoning was incomplete because `ordered_bindings` snapshots once at the top of the loop, and `HookRegistry::insert` accepted duplicate hook IDs. Rejected duplicate hook IDs (any point) in `HookRegistry::insert` and added a poison re-check before invoking each hook impl in `dispatch_before_capability`, `dispatch_before_prompt`, and `dispatch_observer_at`. Added regression tests for both behaviors. - C6 (Medium, Manifest / Predicate Validation): `parse_window` could panic on non-ASCII input because `split_at(len - 1)` requires a char boundary. Rewrote to compute the unit char's UTF-8 byte length and slice safely, added a public `validate_window` helper, and wired it into `HookManifestEntry::validate` for both `InvocationCount` and `NumericSum` bounds. Added tests for non-ASCII, empty, single-char, and zero-duration windows. - C2 (High, Tenant Isolation): partial fix only. The `PredicateEvaluator`'s sliding-window counter was keyed by `(hook_id, capability)`, so cross-tenant state could leak. Extended `HistoryKey` to include `tenant_id` and added a regression test proving counters partition by tenant. Documented the broader dispatcher-per-build / per-run-fresh-dispatcher pattern as deferred follow-up in `crates/ironclaw_hooks/CLAUDE.md`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): emit hook telemetry milestones for audit/SSE observers Wires the hook dispatcher into the host's milestone stream so audit backends and SSE observers can see hook activity. Previously, hook dispatch was invisible — denies, pauses, failures, and observer fires left no trace in the host's observability backend. Changes: - `ironclaw_turns`: add `HookDispatched`, `HookDecisionEmitted`, and `HookFailed` variants to `LoopHostMilestoneKind`, with a closed- vocabulary `HookDecisionSummary` enum (Allow/Deny/PauseApproval/ PauseAuth/Pass/Patch). Introduce a lightweight `HookMilestoneSink` trait that emits hook-specific *kinds* without requiring a `LoopRunContext` (the dispatcher is a process-wide singleton that cannot own a per-run context), plus a `RunScopedHookMilestoneSink` adapter that injects run context and forwards to the existing `LoopHostMilestoneSink`. Also add `InMemoryHookMilestoneSink` for tests. - `ironclaw_hooks`: add a `telemetry` module that converts hook-crate types (`HookId`, `HookTrustClass`, `HookPointSpec`, `FailureCategory`, `FailureDisposition`, `BeforeCapabilityHookDecision`) into the wire- shape labels and summaries the milestone sink expects. Hook ids cross the seam as hex strings because the strongly-typed `HookId` cannot be imported from `ironclaw_turns` (the architecture test enforces `ironclaw_turns -> ironclaw_hooks` stays absent). - `ironclaw_hooks::dispatch`: add an optional `Arc<dyn HookMilestoneSink>` to `HookDispatcher`, set via `with_milestone_sink`. Emit `HookDispatched` before each hook runs, `HookDecisionEmitted` after a decision/pass/patch, and `HookFailed` on timeout/panic/malformed/missing-impl across all three dispatch paths (before_capability, before_prompt, observer). Default behavior (no sink attached) emits nothing — preserves the pre-telemetry observable surface. - `ironclaw_reborn`: document on `with_hook_dispatcher` that callers attach the milestone sink to the dispatcher *before* wrapping it in `Arc` and installing it into the factory, using a `RunScopedHookMilestoneSink` to inject run-context. The dispatcher itself is shared across runs, so attaching a fixed run-context inside it would be wrong. Update `RuntimeEvent` projection in `milestone_events.rs` to ignore the new hook kinds (no projection pathway yet; emitted milestones are consumed by SSE observers directly). Tests: - `ironclaw_hooks::dispatch`: 5 new tests covering milestone emission for deny decisions, panic failures, prompt-mutator patches, observer pass-throughs, and the no-sink default. - `ironclaw_reborn` hooks_integration: end-to-end test wiring a `RunScopedHookMilestoneSink` onto the dispatcher and asserting hook activity surfaces in the host's `LoopHostMilestoneSink`. Total: +6 hook telemetry tests; no existing tests modified. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): extract shared prompt envelope; inject hook patches into prompt bundle Adds `ironclaw_prompt_envelope`, a leaf crate that owns the single envelope primitive used by every model-visible untrusted-content path. `wrap_untrusted` prefixes content with a closed-vocabulary `<Trusted|Untrusted> <source> content: ` marker, rejects bodies carrying instruction-hijack phrases (`ignore previous instructions`, `<|im_start|>`, `<system>`, etc.), and enforces a 4 KiB byte budget by default. Migrates `ironclaw_host_runtime::memory_context` to delegate envelope wrapping, marker rejection, and control-character stripping to the new crate while keeping the `LoopSafeSummary`-specific 512-byte cap and byte truncation local. Existing memory_context behavior and tests are preserved. Wires the same envelope into `ironclaw_hooks`: * `HookPatch::add_enveloped_snippet` now takes a raw body and wraps it via `wrap_untrusted(EnvelopeSource::Hook, …)`. `Installed` hooks produce `Untrusted` envelopes; `Builtin`/`Trusted`/`SelfAuthored` produce `Trusted` envelopes so downstream readers can distinguish the two paths through a uniform marker. * `HookedLoopPromptPort::build_prompt_bundle` is no longer observe-only. After dispatching `before_prompt`, it envelope-wraps every snippet patch (passing `Enveloped` through, wrapping `Trusted` with the envelope helper), enforces the 4 KiB aggregate snippet byte budget across patches, and appends the wrapped snippets to the prompt bundle's `messages` as `system`-role `LoopModelMessage` entries carrying deterministic `msg:hook.<ordinal>.<hash>` content refs (mirroring the skill-snippet ref convention). The envelope crate is a leaf with no ironclaw dependencies, satisfying the boundary contract; the existing `ironclaw_hooks` boundary rule in `reborn_dependency_boundaries` continues to hold because `ironclaw_prompt_envelope` is not on its forbidden list. Test count delta: * `ironclaw_prompt_envelope`: +13 new tests (crate did not exist). * `ironclaw_hooks`: 84 → 88 tests (+4 prompt-port behavior tests: `hook_patch_appended_as_envelope_wrapped_message`, `total_byte_budget_enforced_across_patches`, `instruction_hijack_in_patch_rejected`, `trusted_hook_patch_wrapped_with_trust_marker`). * `ironclaw_host_runtime` memory_context: unchanged (8 tests still pass). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: align tenant-counter test with SanitizedArguments-extended context ctor * docs(reborn): document loader contract; pin HookId hex format Add a "Loader responsibility" section to ironclaw_hooks/CLAUDE.md explaining that tier-specific installers prevent minting wrong-tier impls but cannot enforce origin — that's the loader's job — and recommending registry loaders type-tag extension hooks as LoadedHook::Installed at the loader seam. Add tier_specific_installers_are_documented_as_loader_contract as a regression guard that touches every public install_*_before_capability and install_*_before_prompt method so any signature change forces the loader contract to be re-evaluated. Document HookId::to_hex's 64-char lowercase hex output as part of the cross-crate contract consumed by LoopHostMilestoneKind::Hook* in ironclaw_turns; add hook_id_hex_format_is_stable_64_lowercase_chars in identity::tests and hook_id_string_serialization_matches_to_hex in telemetry::tests to pin the format and the seam conversion path. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(reborn): pin hook milestone JSON schema + assert pairing invariants Add L3 schema-snapshot tests for every hook-related LoopHostMilestoneKind variant (HookDispatched, HookDecisionEmitted per HookDecisionSummary, HookFailed per FailureCategory) so downstream consumers can rely on the JSON wire shape and any accidental field rename, enum-tag rename, or type change fails loudly. Add L4 pairing-invariant matrix test in the hook dispatcher that drives every observable outcome (Allow, Deny, PauseApproval, PauseAuth, Pass, Panic, Timeout, Malformed, MissingImpl) through a recording milestone sink and asserts the dispatched-then-terminator pairing shape. Document the MissingImpl path as the one case that emits a sole HookFailed with no preceding HookDispatched (the dispatcher discovers the protocol violation before the hook is actually dispatched). Add a multi-hook dispatch test that installs three hooks with mixed outcomes (allow/deny/panic) at the same point and asserts each hook produces its own paired sequence in the deterministic (phase, priority, hook_id) order taken from the dispatcher's registry. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(reborn): integration tests for observer middleware through RebornLoopDriverHostFactory Wire the HookedLoopModelPort / HookedLoopTranscriptPort / HookedLoopCheckpointPort observer wrappers into RebornLoopDriverHostFactory::build_text_only_host_with_capabilities, mirroring the existing HookedLoopCapabilityPort / HookedLoopPromptPort composition. The wrappers are applied only when a HookDispatcher is set on the factory, so the default factory shape is unchanged. Add four integration scenarios in crates/ironclaw_reborn/tests/hooks_integration.rs: - observer_hook_fires_after_model_through_factory - observer_hook_fires_after_capability_through_factory - observer_hook_fires_after_checkpoint_through_factory - observer_panic_does_not_fail_model_call (panic-isolation regression) Relax the test-fixture model gateway from "panic if invoked" to returning a stub assistant reply so the AfterModel / panic-isolation tests can drive stream_model through the wrapped port. The existing capability-port tests never touch the gateway, so their behavior is unchanged. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(reborn): introduce HookDispatcherBuilder for type-enforced sink wiring Adds a `HookDispatcherBuilder` in `ironclaw_hooks::dispatch` that owns the dispatcher construction lifecycle: registry -> optional timeout -> optional milestone sink -> installed hooks -> `.build_arc()`. The terminal `.build_arc()` wraps in `Arc` and yields an immutable handle. Tightens the public surface on `HookDispatcher`: `new`, `with_timeout`, `with_milestone_sink`, and every `install_*_*` method are now `pub(crate)`. Outside callers route exclusively through the builder, so "wire the milestone sink before Arc-wrapping" is a compile-time fact rather than a documentation convention. `HookRegistrar::install` now takes a `HookDispatcherBuilder` by value and returns `(HookDispatcherBuilder, Vec<HookId>)`, keeping the builder chainable through manifest installation. `RebornLoopDriverHostFactory` gains `with_hook_dispatcher_builder` to let callers defer `.build_arc()` to the factory — a step toward the FU8 per-build dispatcher pattern. Migrates `foundation_pipeline.rs` and `hooks_integration.rs` to the builder. Internal middleware and dispatch tests continue to use the crate-private `HookDispatcher::new` directly. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): production CapabilityInputResolver for NumericSum predicates Adds HookCapabilityInputResolverAdapter in ironclaw_reborn that bridges the existing LoopCapabilityInputResolver (already used by HostRuntimeLoopCapabilityPort for dispatch input resolution) to the hooks crate's CapabilityInputResolver trait. RebornLoopDriverHostFactory gains with_capability_input_resolver(...), and when both a hook dispatcher and resolver are configured the factory threads the adapter into HookedLoopCapabilityPort::with_resolver — so NumericSum and other argument-dependent predicates evaluate against real, sanitized inputs instead of failing closed against the framework's null default. The adapter also enforces a configurable serialized-byte budget (default 64 KiB) as defense in depth ahead of the hooks crate's per-string and depth caps in SanitizedArguments. Unit tests cover the four adapter branches (resolved JSON, inner-error → None, non-object pass-through, oversized → None) and a new end-to-end integration test (numeric_sum_predicate_caps_total_value_against_real_inputs) drives the full factory wiring: with a NumericSum cap of 99 over an "amount" field, two invocations carrying {"amount":"50"} let the first pass through and deny the second at the hook seam, with the inner port reached exactly once. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): per-build HookDispatcher for full per-run isolation (C2) Introduce `with_hook_dispatcher_factory(F)` on `RebornLoopDriverHostFactory`. The closure is invoked once per `build_text_only_host*` call, so dispatcher-owned mutable state — slot poisoning, registry mutations, predicate-counter siblings — is scoped to a single host build instead of shared across every host the factory produces. The legacy `with_hook_dispatcher(Arc<HookDispatcher>)` adapter is kept as a thin wrapper that returns clones of the same `Arc` on every build. Its shared-state behavior is now documented as an explicit opt-in for backward compat; new wiring should prefer the factory closure. Adds two regression tests: - `per_build_dispatcher_state_does_not_leak_across_runs` — installs a panicking hook, builds two hosts back-to-back, and proves the inner port is never reached on build 2 (fresh slot still applies the fail-closed deny). Pins per-run isolation. - `legacy_with_hook_dispatcher_shares_state_across_builds` — pins the shared-state semantic of the legacy adapter as the explicit baseline. Migrates `predicate_deny_hook_short_circuits_inner_port` to the new factory-closure path so the new wiring is exercised by the existing suite. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): project hook telemetry milestones into RuntimeEvent for durable audit Extend the runtime event substrate with `HookDispatched`, `HookDecisionEmitted`, and `HookFailed` kinds carrying closed-vocabulary labels and the blake3-hex hook identity. Project the matching `LoopHostMilestoneKind::Hook*` variants in `DurableLoopHostMilestoneSink` so hook telemetry now lands in the same durable event log as model/reply/loop milestones — SSE observers still see live hook events, and audit replay can reconstruct the full hook trail. - `ironclaw_events`: add hook variants to `RuntimeEventKind`, optional hook fields on `RuntimeEvent` (`hook_id`, `hook_point`, `hook_trust_class`, `hook_decision`, `hook_failure_category`, `hook_failure_disposition`), typed constructors (`hook_dispatched`, `hook_decision_emitted`, `hook_failed`), and dedicated sanitizers (`sanitize_hook_label`, `sanitize_hook_id`) re-run on every wire crossing. No new crate dependency edges; hook strings cross the boundary opaque. - `ironclaw_reborn::milestone_events`: project the three hook milestone kinds via a new `loop.hook` capability id. `HookDecisionSummary` is collapsed to its closed-vocabulary `kind_name()` so sanitized reasons never enter the durable substrate. - `ironclaw_event_projections`: extend `TimelineEntryKind` and the `RuntimeEventKind -> RunProjectionStatus` mapping so hook events are pure telemetry — they preserve the current run status rather than changing it. - Tests: 4 unit tests in `ironclaw_events::runtime_event::tests` (serde round-trip per variant + unsafe-label collapse), 3 in `ironclaw_reborn::milestone_events::tests` (projection per variant, including the assertion that raw `Deny { reason }` text does not reach the durable wire payload). Existing replay-projection direct-construction tests updated for the new RuntimeEvent fields. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): enforce manifest-declared hook scope at dispatch time (C3) Audit finding C3: extensions could declare `[[hooks]]` with `scope = "own_capabilities"` in their manifest, but the dispatcher never enforced it — an Installed hook from ext-A could fire against capabilities provided by ext-B. Scope was parsed but not load-bearing. This change makes scope load-bearing end-to-end: - `BeforeCapabilityHookContext` carries an optional `provider: ironclaw_host_api::ExtensionId` populated by the middleware. The hook context is `#[non_exhaustive]` already so this is non-breaking. - `HookBinding` gains `owning_extension: Option<ExtensionId>` and `scope: HookBindingScope`. `HookBindingScope` is `Global` / `OwnCapabilities` / `SameTenant`. Builtin and Trusted bindings default to `Global` and carry no `owning_extension`; Installed bindings carry both, sourced from the manifest. - `HookDispatcher::install_installed_*` installers now require the caller to pass `(owning_extension, scope)`. The registrar derives both from the manifest entry, so manifest authorship is the single source of truth. - A new `CapabilityProviderResolver` trait + bundled `NullCapabilityProviderResolver` lets the middleware lift the capability id to its provider at invocation time. The middleware wires the resolved provider into the hook context. - `dispatch_before_capability` consults `binding.scope.permits(...)` before invoking each hook. Bindings that don't permit the current invocation are inert — no sink call, no failure record, no poisoning. Conservative defaults: - When the provider resolver returns `None` (no resolver wired, or the capability has no known provider), `OwnCapabilities`-scoped hooks do NOT fire. An attacker cannot bypass scope filtering by stripping provider info from the descriptor. Tests: - 5 new dispatcher tests cover OwnCapabilities matching, foreign provider, unresolved provider, SameTenant, and Builtin Global. - 1 new registrar test asserts manifest scope and extension propagate into `HookBinding`. - 1 new middleware test asserts the provider resolver populates the hook context. - 1 new integration test in `ironclaw_reborn` proves an ext-A hook scoped to `OwnCapabilities` does not intercept invocations that have no resolved provider (the production composition default). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * style: rustfmt dispatch.rs after FU1 merge * docs(hooks): prior-art comparison against LSM/eBPF/Envoy/K8s/OPA/CRX/VSC/Tauri Validates the IronClaw hooks design against 8 established hook/policy systems across 8 axes (dispatch, trust tiers, attenuation, decision vocabulary, failure semantics, isolation, manifest, audit). Surfaces: - 7 areas where ICLAW stands out vs prior art (type-level trust enforcement, dispatch-time scope, failure-kind matrix, pause-with- gate-ref, pairing-invariant audit matrix, tenant-keyed predicates, phase-ordered dispatch) - 4 conventional choices we should revisit (in-process Installed-WASM, sticky poison, no formal dispatch model, no installation rate-limit) - 3 divergences whose 'why' is weak and need design review * docs(hooks): STRIDE threat model for v1 framework Enumerates 7 adversary classes (A1-A7), 6 assets ranked by blast radius, and ~35 attack vectors across STRIDE categories with mitigations, existing tests, and residual risk. Surfaces 7 prioritized follow-ups: - High: per-extension hook-count cap (D3/D4) - High: gate-ref unguessability + one-shot test (S1) - Med: resolver field-level scope (I2) - Med: per-evaluator state ceiling (D5) - Med: poison-stickiness operator runbook - Low: timing side-channel residual acknowledgement (I4) - Low: instruction-marker denylist periodic review (I5) Confirms the load-bearing 'Installed cannot Allow' (E1) property holds via type-level seal + tier-specific installers, backed by compile_time_seal_test and installed_binding_cannot_be_paired_with_ privileged_impl tests. Explicit out-of-scope: extension install pipeline (#3492), WASM exec sandbox (needs separate threat model when it lands), approval gateway (#3564). * feat(hooks): close threat-model gaps S1 (gate-ref entropy) and D3/D4 (registration flood) S1 (gate-ref unguessability, factory side): - Three new tests on `UuidHookGateRefFactory`: - `gate_refs_are_v4_uuids` pins the v4 entropy source (122 random bits per ref per RFC 4122 §4.4); fails if a future change moves to a counter or weaker UUID version. - `gate_refs_have_no_collisions_across_many_calls` mints 20k refs across both namespaces, asserts zero collisions (statistical proxy for entropy quality). - `approval_and_auth_namespaces_do_not_overlap` confirms prefix routing separation. - Doc comment now documents the security property explicitly and delineates factory-side vs gateway-side responsibilities for the one-shot consumption property. D3/D4 (hook registration flood): - New `MAX_HOOKS_PER_EXTENSION = 32` and `MAX_HOOKS_PER_EXTENSION_PER_KIND = 8` consts in `registrar.rs`. - New `HookRegistrar::enforce_registration_caps` runs pre-flight at the top of `install()`, before any binding is inserted. Whole-batch rejection means a partially-installed batch cannot slip past. - Three regression tests: total-cap rejection, per-kind-cap rejection, at-cap acceptance. - Error messages cite the threat-model finding so operators can map rejection back to the design rationale. Threat model updated: S1, D3, D4 marked closed in the cross-cutting properties matrix and the open-follow-ups list. * test(hooks): three real hooks built against the public API + ergonomics findings Builds three representative hooks from outside the crate, mimicking what an extension or system author would actually write: 1. polymarket-daily-cap — Installed predicate hook, InvocationCount rate-cap with Deny on excess. Canonical 'rate-limit a capability' use case for the predicate language. 2. large-stake-approval-gate — Installed predicate hook, NumericSum over amount_usd field, PauseApproval at $1000/24h. Manifest-shape + registrar-install coverage from outside Reborn; end-to-end dispatch lives in ironclaw_reborn integration tests because NumericSum needs resolved args (a friction finding documented in the companion doc). 3. pii-redaction-warning — Trusted Rust hook implementing PrivilegedBeforePromptHook, injects a trusted instruction snippet reminding the model to redact PII. Demonstrates the path a system author takes when the predicate language isn't expressive enough. API change (F1 fix): SanitizedArguments::unresolved() promoted from pub(crate) to pub. This is the documented safe default — predicates that need args must fail closed against it — so exposing the constructor cannot weaken any trust property. The sanitizing from_json constructor stays sealed; that's the trust boundary. Without this fix, external hook authors could not construct a BeforeCapabilityHookContext with both a known provider AND unresolved args, which made TDD of their own predicate impossible. Findings documented in docs/real-hooks-findings.md, ranked by severity. Big-picture observation: writing the Trusted Rust hook (F4) was easier than writing the declarative predicate hook (F1 + F2 + F3) — three of seven findings target predicate-authoring ergonomics. The declarative path needs the most polish before third-party extension authors will trust it for non-trivial policy. Tests: 6 new in real_hooks.rs, all pass. * feat(hooks): close all remaining threat-model and ergonomics gaps Closes the Med-priority threat-model gaps (I2, D5, poison runbook) and all real-hooks ergonomics findings (F2, F3, F5, F6, F7) in a single pass. Threat model: - I2 (resolver field-scope): documented in SanitizedArguments rustdoc. The narrow public surface (only is_resolved + extract_numeric) enforces field-scope by construction for the current predicate path. Reassess when Installed-WASM lands. - D5 (evaluator state ceiling): MAX_HISTORY_KEYS = 8192 per map, LRU eviction with evictions_observed() metric for operator monitoring. New regression test lru_eviction_increments_counter_and_drops_oldest_key. - Poison-stickiness runbook: new docs/operator-runbook.md with recovery options ranked by cost. Ergonomics findings: - F2 (closed-vocab deny reasons): rustdoc on OnExceededAction and GateDecisionView::Deny explaining the audit-vs-model split and why manifest reason text doesn't reach the model. - F3 (NumericSum can't be TDD'd outside Reborn): new test-support feature flag with SanitizedArguments::for_tests(value) that external hook authors can opt into via dev-dep. - F5 (two ExtensionId types): added From<&ironclaw_host_api::ExtensionId> impl for identity::ExtensionId, plus cross-link rustdoc. - F6 (HookManifestEntry struct-literal fragility): added #[non_exhaustive] + HookManifestEntry::new(id, kind, body) + with_scope/with_phase/with_priority/with_description/with_requires_grant builder methods. Migrated 3 external call sites in tests/. - F7 (priority guidance): rustdoc on HookPriority with when-to- deviate guidance, named FIRST/LAST constants documented for Builtin/Telemetry use cases. Tests: 151 unit + 1 + 6 integration in ironclaw_hooks all pass with --all-features. ironclaw_reborn (13 hooks_integration scenarios) unchanged. Threat model updated: I2 / D5 / poison runbook marked closed in both the per-vector table and the cross-cutting properties matrix. Open follow-ups now down to two Low items (I4 timing side-channel residual, I5 instruction-marker denylist refresh) plus the deferred DenyReasonCode enum from F2. * fix(ci): collapse nested match in hooks_integration test for clippy --all-features CI runs `cargo clippy --all --tests --examples --all-features -- -D warnings` which is stricter than the workspace clippy I ran locally and trips `clippy::collapsible_match` on the nested-if in HookDecisionEmitted matching. Collapse the inner `if decision.kind_name() == "deny"` into an arm guard. * feat(hooks): address henrypark133 review — Critical #1/#2/#3/#5, Concerning #5/#7 Address composition-seam bugs in the Reborn factory wiring + doc tidy. henrypark133 review findings addressed: Critical #1 — before_prompt hook messages not materialized. HookedLoopPromptPort now requires a HookPromptMaterializationSink and fails closed if patches are emitted without one. The reborn factory installs an InstructionStoreBackedHookSink adapter that delegates to the host's InstructionMaterializationStore, so synthetic msg:hook.* refs are resolvable by the downstream model resolver. New seam trait (HookPromptMaterializationSink) keeps ironclaw_hooks decoupled from LoopRunContext. Critical #2 — OwnCapabilities hooks were inert in production wiring. Factory now installs SurfaceBackedProviderResolver (consults the visible-capability surface for capability_id → provider). With this, ctx.provider is populated and OwnCapabilities-scoped Installed hooks actually fire against their own provider's capabilities. Critical #3 — gate refs were unresolvable. Middleware default switched from UuidHookGateRefFactory to FailClosedHookGateRefFactory. Tests must explicitly opt into UUID (via with_gate_ref_factory) to exercise the affirmative ApprovalRequired path; production deployments must install a router-backed factory. New factory method RebornLoopDriverHostFactory::with_hook_gate_ref_factory. Concerning #5 — AfterModel fired twice + before durable finalization. Removed AfterModel dispatch from HookedLoopModelPort; the transcript port's finalize_assistant_message is now the sole AfterModel boundary (the durable one). Model port wrapper is preserved as a no-op shim for symmetry + future model-response-observed point. Concerning #7 — doc tidy: - CLAUDE.md: 3 trust classes → 4 (Builtin/Trusted/Installed/SelfAuthored with explicit note that SelfAuthored is run-scoped only and not loadable from an external source). - operator-runbook.md: "Audit log" → "durable runtime event stream" where the projection is actually the runtime-event stream, not formal AuditEnvelope records. - prior-art.md: poison-lifetime nuance — per-host-build with the factory pattern, process-lifetime only for the legacy adapter. - prior-art.md:80: trailing whitespace removed. Testing gaps from henrypark133 — caller-level tests through RebornLoopDriverHostFactory: #1 (before_prompt resolver path): before_prompt_hook_message_is_resolvable_via_factory_wiring #2 (OwnCapabilities positive/negative/unknown): own_capabilities_hook_fires_when_provider_matches own_capabilities_hook_does_not_fire_when_provider_differs own_capabilities_hook_does_not_fire_when_provider_unknown #3 (pause/auth gate lifecycle or fail-closed): pause_approval_with_default_factory_fails_closed_as_denied pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref (updated to require explicit UuidHookGateRefFactory opt-in) #5 (AfterModel exactly-once at durable boundary): after_model_fires_exactly_once_at_durable_boundary Still TODO from review (separate commits): Critical #4 (telemetry context — two-run attribution) + gap #4 Concerning #6 (TimelineEntry hook metadata projection) + gap #6 Tests: 154 unit + 18 hooks_integration + all other reborn tests pass. Workspace clippy + fmt + no-panics clean. * feat(hooks): address remaining henrypark133 review — Critical #4, Concerning #6 Critical #4 — per-run hook telemetry attribution. New `HookDispatcherBuilderFactory` signature: factory returns a HookDispatcherBuilder, and `RebornLoopDriverHostFactory` attaches a `RunScopedHookMilestoneSink` keyed to the CURRENT run's LoopRunContext inside `build_text_only_host_with_capabilities`, before sealing the dispatcher. The previous zero-arg signature relied on the closure capturing run_context — silently misattributed across reuses; new public API `with_hook_dispatcher_builder_factory` removes that failure mode entirely. Legacy `with_hook_dispatcher_factory` retained for back-compat (its sink-wiring contract stays caller-side). Concerning #6 — TimelineEntry hook metadata. Added 6 optional fields to `TimelineEntry` (hook_id, hook_point, hook_trust_class, hook_decision, hook_failure_category, hook_failure_disposition) and projected them from `RuntimeEvent::Hook*`. Replay consumers now see which hook fired/failed, not just that some hook event happened. Each field is closed-vocabulary (no free-form reason text — that stays in the audit reason payload, not the product replay DTO). Testing gaps from henrypark133 — caller-level tests: #4 (two-run hook telemetry attribution): hook_telemetry_attribution_is_per_run_not_captured Builds two hosts from the SAME builder factory closure with two fresh LoopRunContexts. Asserts each run's hook milestones carry its OWN run_id (no stale captured one). #6 (replay projection contract for hook events): hook_runtime_events_project_with_sanitized_hook_metadata non_hook_runtime_events_project_with_no_hook_metadata Constructs RuntimeEvent::Hook{Dispatched,DecisionEmitted,Failed} and asserts the projection preserves the metadata fields. The negative test guards against cross-contamination on non-hook events. All henrypark133 review items now addressed: Critical: #1, #2, #3, #4 — done Concerning: #5, #6, #7 — done Testing gaps: #1-#6 — done Tests: 154 unit + 19 hooks_integration in ironclaw_reborn + 61 reborn unit + 38 + 2 new in ironclaw_event_projections + ... pass. Workspace clippy + fmt + no-panics clean. * docs(hooks): scope DenyReasonCode closed-vocabulary enum (successor #6) Successor PR from #3573 — real-hooks ergonomics finding F2 (deferred). Adds a curated vocabulary of model-visible denial reasons so hook authors can communicate why a deny happened without opening a free-form prompt-injection channel. * feat(hooks): DenyReasonCode + PauseReasonCode closed-vocabulary enums Address real-hooks ergonomics finding F2 (deferred from PR #3573). The prior dispatcher collapsed every Installed-tier deny to the static label 'hook_predicate_denied', because manifest reason strings are author-controlled and surfacing them to the model would open a prompt-injection channel. The cost: the agent couldn't tell *why* a hook denied. This PR introduces two closed-vocabulary enums: - DenyReasonCode: Generic / RateLimit / ValueCap / Blocklist / RequiresApproval / OutOfPolicy - PauseReasonCode: Generic / RequiresApproval / OverThreshold / SensitiveAction Each variant has an as_label() returning &'static str (so the sink's &'static str contract is preserved). New OnExceededAction variants 'DenyWithCode { code, reason }' and 'PauseApprovalWithCode { code, reason }' let manifest authors opt into the richer labels while keeping reason audit-only. The legacy Deny { reason } / PauseApproval { reason } variants are retained for back-compat and map to DenyReasonCode::Generic / PauseReasonCode::Generic — existing manifests continue to produce hook_predicate_denied / hook_predicate_pause_requested. Threat-model regression: a hook author cannot smuggle text into the model-visible label because the 'code' field is typed as the enum; there's no String slot exposed model-side. A test (deny_with_code_only_exposes_enum_variants_to_model) documents this as a compile-time property. Tests (+7 new = 161 total): - deny_reason_code_labels_are_stable: pins the label vocabulary so rename/relabel is loud. - pause_reason_code_labels_are_stable: same for PauseReasonCode. - deny_with_code_round_trips_through_json + pause variant: wire round-trip + snake_case tag assertion. - deny_with_code_only_exposes_enum_variants_to_model: compile-time property check. - rate_or_value_cap_with_deny_code_routes_to_code_label: end-to-end affirmative test that the dispatcher emits the code's label. - rate_or_value_cap_with_pause_code_routes_to_code_label: same for pause. Scope doc: crates/ironclaw_hooks/docs/successors/06-deny-reason-code.md * test(hooks): address codex review on #3636 - Update stale real-hooks-findings.md F2 row to cite this PR's enum follow-on (was 'deferred'). - Add install_deny_with_code_manifest_surfaces_code_label_on_dispatch: end-to-end test driving the registrar->dispatcher path for the new DenyWithCode variant (prior tests covered serde + direct hook evaluation, but not the manifest install path that downstream authors actually use). Codex review on PR #3636: APPROVE with two recommendations; both addressed. Tests: 162 unit (+1 new). Clippy/fmt clean. * fix(hooks): attenuate Installed-tier prompt patches to user role Installed-tier `before_prompt` patches were injected as role:"system" messages. Envelope text labels ("[ext-foo says]: ...") do not strip system-role authority from the model's perspective, so a third-party extension could inject system-tier instructions through a snippet patch. This is a prompt-authority escalation against the trust hierarchy the framework otherwise enforces. Add `role_for_trust_class()` mapping Installed -> "user" and Builtin/Trusted/SelfAuthored -> "system". Thread per-patch trust_class through `wrap_patches_to_messages` and use it for the emitted `LoopModelMessage.role`. Tests: - installed_hook_patch_drops_to_user_role: asserts the role for an Installed-tier patch is "user" - trusted_tier_hook_patch_keeps_system_role: regression that Trusted tier still produces system-role content * fix(hooks): enforce scope filter on observer dispatch + reject incompatible points Two related defense-in-depth fixes against silent scope-filter failure: 1. The registry silently accepted Installed bindings with `HookBindingScope::OwnCapabilities` at points (BeforePrompt, AfterModel, AfterCheckpoint) whose dispatch context carries no per-capability provider. The manifest's declared scope had no effect at all — the hook fired against every dispatch. Reject the binding at install time so the operator sees the misconfiguration. 2. `dispatch_observer_at` for `AfterCapability` did not consult the binding's scope, so an Installed observer registered with `OwnCapabilities` fired against every invocation regardless of provider. Add `dispatch_observer_at_with_provider` carrying the resolved capability provider; the capability-port middleware resolves the provider once per invocation and threads it through both the BeforeCapability hook context and the AfterCapability observer dispatch. The dispatcher then enforces `HookBindingScope::permits` on each observer binding. `ObserverHookContext` gains a `provider: Option<ExtensionId>` field; `#[non_exhaustive]` keeps existing authors compiling. Tests: - rejects_own_capabilities_at_before_prompt - rejects_own_capabilities_at_after_model - accepts_own_capabilities_at_before_capability - own_capabilities_observer_filters_foreign_providers (covers foreign / matching / unresolved provider) * fix(hooks): preserve free-form audit reason alongside closed-vocab model label (serrrfirat #3636) `PredicateBackedBeforeCapabilityHook::evaluate()` was discarding the free-form `reason` from `EvaluatorDecision::{Deny, PauseApproval}` with `..` and only sending `code.as_label()` into the sink. The `HookDecisionEmitted` milestone therefore carried only the closed- vocab label, and operator-visible audit/SSE context was silently lost end-to-end. The fix splits the channels: - Model sees the closed-vocab label (`hook_rate_limit`, `hook_pause_over_threshold`, ...) via `sink.deny(label)`. This channel is unchanged. - Audit/SSE sees the manifest's free-form `reason` via a new audit-only sink method `record_audit_reason(reason: String)`. The recording sink captures it; the dispatcher reads it after the hook returns and threads it into `LoopHostMilestoneKind::HookDecisionEmitted`. Surface changes: - `PrivilegedGateSink` / `RestrictedGateSink` gain `record_audit_reason(String)` — accepts dynamic `String` (audit-only, no model-facing seam) unlike the `&'static str` decision reasons. - `RecordingGateSink` gains an `audit_reason: Option<String>` field. - `GateHookOutcome::Decision` is now `Decision { decision, audit_reason }`. - `HookDispatcher::emit_decision_with_audit` threads the audit reason into the milestone. - `LoopHostMilestoneKind::HookDecisionEmitted` gains a `#[serde(default, skip_serializing_if = "Option::is_none")]` `audit_reason: Option<String>`. The durable RuntimeEvent projection intentionally drops this field — audit reasons are operator-facing in-memory SSE content, never durable cross-process surface. Tests: - `deny_with_code_records_audit_reason_separately_from_model_label`: asserts the recording sink ends with `Deny { reason: "hook_rate_limit" }` in `state` AND `audit_reason == Some("daily cap of $1000 ...")`. * fix(hooks): remove unused model_request helper (CI clippy fix) * fix(hooks): address serrrfirat P1/P2 findings on PR #3573 Three issues from the 5-15 review: **P1 #1 registrar.rs:70 — `same_tenant` grants not enforced** `HookManifestEntry::validate` only confirmed `requires_grant` was present; the registrar then immediately installed the binding with no host-verified grant context. A manifest could declare `requires_grant = "anything"` and get a cross-extension binding for free. Fix: `HookRegistrar` now carries a `verified_grants: HashSet<String>` (empty by default — default-deny). Add the host-facing setter `with_verified_grants(...)`. At `install_one`, if `entry.requires_grant` is `Some(g)`, require `g ∈ verified_grants` or reject with a clear error. Tests: - `install_rejects_same_tenant_without_verified_grant` - `install_rejects_same_tenant_when_verified_grants_mismatch` - The existing positive test `installer_propagates_owning_extension_and_scope_from_manifest` now wires the verified grant explicitly (proves the API contract). **P1 #2 prompt_port.rs:150 — zip misalignment** The materialization loop zipped surviving messages against the ORIGINAL unfiltered patch list. `wrap_patches_to_messages` skips metadata patches and over-budget snippets, so the zip silently paired message[0] with patch[0] even when patch[0] was the skipped metadata — materializing the wrong content (or none) under the snippet's synthetic ref. Fix: `wrap_patches_to_messages` now returns `Vec<WrappedHookMessage { message, safe_content }>` — surviving messages paired with their content by construction. The caller materializes `entry.safe_content` under `entry.message.content_ref` directly; no zip against unfiltered input. Removed the now-unused `safe_content_for_patch` helper. Test: - `materialization_stays_aligned_when_metadata_patches_are_filtered`: a hook emits `[metadata, snippet]`; asserts only one model message, and the materialized content under its ref contains the snippet's body — proves filtering can no longer desync from materialization. **P2 #3 loop_driver_host.rs:1343 — `with_hook_dispatcher_builder`** Docs said it deferred `build_arc()` to let the host factory finalize wiring; the implementation called `build_arc()` eagerly and routed through the legacy shared-dispatcher adapter, losing per-run dispatcher isolation and the run-scoped milestone sink. Fix: marked `#[deprecated]` with a note pointing callers to `with_hook_dispatcher_builder_factory(|| ...)` for per-build isolation, or `with_hook_dispatcher(...)` if they actually meant the shared adapter. The method body is unchanged so no callers break; they'll see the deprecation warning. No internal callers exist, so the deprecation doesn't trip `-D warnings`. All 162 hooks lib + 19 reborn integration tests pass; clippy clean. * fix(hooks): address serrrfirat 3573-2026-05-15 review findings P1 — prompt bundle authority mismatch (prompt_port.rs): `HookedLoopPromptPort::build_prompt_bundle` called the inner port first, which caused `HostManagedLoopPromptPort` to issue the prompt-bundle authority grant against the pre-hook message list. The wrapper then appended `msg:hook.*` messages to `bundle.messages`, so the downstream model request hit `grant.messages != messages` and failed closed with "model request messages do not match the host-built prompt bundle". Add `with_bundle_authority(authority, run_context)` and re-issue the grant after appending hook messages so it covers the post-hook bundle. Reborn wires `prompt_authority.clone()` + `run_context.clone()` into the wrapper at construction time. P2 — observer installer accepts non-observer points (dispatch.rs): `install_observer` accepted any `HookPointSpec` (including `BeforeCapability` / `BeforePrompt`) and only populated the observer map. Dispatch later found a binding without a gate/mutator impl and fail-closed the capability with "binding present without installed implementation". Reject non-observer points at install time so misuse fails loudly rather than poisoning bindings at dispatch. P2 — batch path skipped AfterCapability observers on inner error (capability_port.rs): The batch loop used `?` directly on `self.inner.invoke_capability(...)`, which propagated the error before dispatching `AfterCapability` observers. Failed batch entries disappeared from telemetry / audit, while the single-invocation path dispatches observers on error. Capture the inner result, dispatch observers, then propagate the error. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): address PR #3573 review feedback round 3 Addresses serrrfirat's CHANGES_REQUESTED review (2026-05-20) by tightening several install-time / dispatch-time bounds and gating production seams: - Bound free-form audit reasons crossing telemetry. New `telemetry::sanitize_audit_reason` strips control characters and caps length at 512 bytes; `emit_decision_with_audit` routes the manifest- supplied reason through it before publishing milestones. Manifest validation also rejects reasons over the same byte limit at install time so the wire-side cap is a defense-in-depth layer, not the only line. - Make hot dispatch O(H) instead of O(H^2). The per-binding poison recheck used to acquire the registry mutex and walk every binding; `ordered_bindings_with_poison_snapshot` now takes the active bindings and the poisoned hook-id set under a single lock, and each loop threads a local `HashSet<HookId>` that absorbs mid-dispatch poisoning. Removed the redundant `is_poisoned` helper. - Gate `HookDispatcher::registry_for_test` behind `cfg(any(test, feature = "test-support"))`. The accessor previously exposed `&Mutex<HookRegistry>` in production, letting any `Arc<HookDispatcher>` holder lock and call `HookRegistry::poison` to disable installed hooks. Added `active_bindings_snapshot(point)` as the read-only production-safe replacement. - `#[serde(deny_unknown_fields)]` on every hook-manifest and predicate DTO (`HookManifestEntry`, `HookManifestBody`, `WasmBudget`, `HookPredicateSpec`, `CapabilityPredicate`, `ValueOrRateBound`, `OnExceededAction`). Typoed or unsupported fields (e.g. a manifest-supplied `trust_class`) now fail loud at install time instead of being silently dropped. - Bound predicate trees at install. New `validate_predicate_tree` enforces `MAX_PREDICATE_DEPTH = 8`, `MAX_PREDICATE_NODES = 64`, `MAX_PREDICATE_STRING_BYTES = 256`, and `MAX_MANIFEST_REASON_BYTES = 512`. A hostile registry manifest can no longer install a deep or huge `All`/`Any` tree that the evaluator would recursively walk on every match. - Cap sliding-window samples per key. `MAX_SAMPLES_PER_KEY = 4_096` in the predicate evaluator. Both the invocation-count and numeric-sum histories drop the oldest sample once the cap is reached, bounding memory under attacker-triggered hot capabilities while preserving rate/value-cap semantics over the most recent window. - `split_indexer` / `resolve_path` now fail closed on malformed bracket syntax (`amount[foo]`, `amount[`, trailing garbage). Previously they silently fell back to the parent field, which could let a typoed `NumericSum` predicate evaluate against the wrong value and allow calls the predicate would otherwise have denied. - Honor `PatchOrdinalHint`. `WrappedHookMessage` carries the source patch's `ordinal_hint`; `HookedLoopPromptPort` inserts `NearTop` messages after the bundle's `identity_message_count` and appends `Last` messages at the end. Safety/policy snippets that need early placement now get it. - Update `ironclaw_hooks` top-level docs to reflect the four trust classes (`Builtin`/`Trusted`/`Installed`/`SelfAuthored`) and the now-wired Reborn middleware composition. Tests added: - `manifest::rejects_unknown_top_level_field` - `manifest::rejects_unknown_wasm_budget_field` - `manifest::rejects_predicate_tree_exceeding_max_depth` - `manifest::rejects_predicate_tree_exceeding_max_nodes` - `manifest::rejects_predicate_string_exceeding_max_bytes` - `manifest::rejects_manifest_reason_exceeding_max_bytes` - `points::capability::malformed_indexer_returns_none_not_parent_value` - `telemetry::sanitize_audit_reason_*` (truncate / strip control / preserve / empty) `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features`, and `cargo test -p ironclaw_hooks` all pass clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(hooks): batch deferred test coverage from #3573 review (#3914) * perf(hooks): defer capability input resolution until a predicate needs it (#3913) * fix(rebase): adapt hooks tests + middleware to upstream API additions - CapabilityDescriptorView: add parameters_schema field - LoopModelRequest / LoopPromptBundleRequest: add capability_view field - TimelineEntry test builder: add hook_id / hook_point / hook_trust_class / hook_decision / hook_failure_category / hook_failure_disposition fields - ironclaw_reborn::tests::hooks_integration: switch from InMemoryLoopCheckpointStore to InMemoryTurnStateStore (which now impls both LoopCheckpointStore and TurnStateStore), pass TurnActor in TurnRunState, supply the new turn_state_store factory arg - ironclaw_reborn lib.rs: drop the pub-use re-exports that upstream intentionally removed (per the module-directory rationale in the current ironclaw_reborn lib.rs doc comment); update the hooks_integration test imports to use module paths - Cargo.toml: union the hooks-foundation member list with upstream's new crates (event_streams, auth, first_party_extensions, reborn_webui_ingress, product_workflow_storage, webui_v2); drop ironclaw_storage which no longer exists upstream - crates/ironclaw_architecture/tests/reborn_dependency_boundaries: keep upstream's removal of ironclaw_filesystem from the ironclaw_turns forbidden list AND add ironclaw_hooks to that list - crates/ironclaw_reborn/src/milestone_events.rs: drop dead loop_failure_kind helper (replaced upstream by loop_failure_kind_name in text_loop_driver.rs); keep hook_decision_label which is still used Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * perf(hooks): restore batched capability dispatch when hooks acti…
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…earai#3633) * docs(hooks): scope production gate-ref factory (successor #1) Successor PR scope doc. Until this lands, hook PauseApproval/PauseAuth decisions surface as Denied in production because the middleware default is FailClosedHookGateRefFactory (PR nearai#3573 / henrypark133 Critical #3). This PR carries the scope doc only; implementation follows after review of the design (cross-crate seam to approval gateway is the load-bearing decision). * docs(hooks): incorporate codex review on nearai#3633 scope Adds two addenda from codex's design-review pass: - Critical: actor/session binding requirement (prevents same-tenant wrong-user approval-bypass). Gateway reservation must carry the actor/session id and reject cross-actor consumption. - Recommendation: include capability id + arguments digest in the reservation, not just free-form reason. Lets the approval UI show the exact gated call AND defeats a future-call digest-mismatch replay vector. * Implement router-backed hook gate refs Cites Codex review addenda: bind hook gate reservations to actor/session identity and carry capability plus arguments digest before handing refs to the approval/auth router. * Fix router-backed hook gate context and TTL * fix(hooks): make gate-ref resolution time router-owned (serrrfirat HIGH nearai#3633) serrrfirat HIGH on PR nearai#3633: `HookGateResolutionRequest.resolved_at` was a caller-controllable timestamp. Any adapter wiring the request from external input — or a buggy router that trusted it — could backdate it and consume an expired approval/auth gate ref. The same caller-supplied timestamp was persisted as the reservation's `consumed_at`, so forged values also corrupted the one-shot audit trail. The `InMemoryHookGateRouter` reference implementation already reads its own wall clock (`Utc::now()`) inside `resolve_gate` for both expiry checks and `consumed_at` — but the public request struct still exposed `resolved_at` as a `pub` field, encouraging future router impls to trust it and leaving the field as ambient trust- boundary surface. Fix: remove `resolved_at` from `HookGateResolutionRequest` entirely. Time authority for resolution and consumption lives exclusively on the router's wall clock; the only place a resolution timestamp surfaces is `HookGateResolution::resolved_at` (the *result*), which is router-supplied. The `for_kind` / `for_invocation` constructors no longer take or set a timestamp. Tests: - `router_backed_pause_approval_gate_ref_rejects_backdated_resolution_after_ttl` reframed: it no longer mutates `request.resolved_at` (the field is gone). Instead it relies on the router's own clock — TTL = 1ms, sleep 5ms, resolve must surface `Expired`. The property is now statically enforced by the absence of the field rather than dynamically asserted, but the regression test still exercises the router-owned-time code path. * fix(hooks): address henrypark133 must-fix #1, #2, #3, #5 on PR nearai#3633 Four items from the 5-15 review: **#1 (must-fix) MAX_RESERVATION_TTL cap** `RouterBackedHookGateRefFactory::try_new` now caps `reservation_ttl` at 24h. Without the cap, an operator misconfiguring a year-long TTL accumulates unresolved reservations in `InMemoryHookGateRouter` state for the full window — a long-tail memory leak. 24h is plenty for human-in-the-loop approval flows. **#2 (must-fix) MAX_REASON_BYTES cap** `mint(...)` rejects `reason` strings longer than 4 KiB. Without the cap, a buggy or malicious caller could push arbitrarily large strings through the approval store; the reason is operator-facing and may be persisted. **#3 (must-fix) Split `InvalidDigest` from `InvalidToken`** `validate_token` previously returned `HookGateError::InvalidDigest` for failures on actor / session ids — confusing because those values aren't digests. Add a new `InvalidToken { field, reason }` variant and route `validate_token` to it. `InvalidDigest` stays for actual sha256-digest shape failures. **#5 (must-fix) Collapse consumption-failure oracle** `From<HookGateError> for AgentLoopHostError` previously preserved Display text for every variant, so a probing caller could distinguish "this gate ref doesn't exist" from "this gate ref belongs to another run/actor/capability" — an oracle for liveness detection on foreign gate refs. Now collapses the entire consumption-failure family (`UnknownGate` / `AlreadyConsumed` / `Expired` / `KindMismatch` / `RunMismatch` / `ActorMismatch` / `CapabilityMismatch` / `ArgumentsDigestMismatch`) to a single opaque "hook gate consumption denied" surface. Misuse / availability variants still surface details — they signal config bugs and need operator visibility. Internal variants stay distinct for test assertions and operator-visible tracing. **Bonus** (henrypark133 non-blocking nearai#9): Add a `tracing::warn!` at the conversion site so operators can still distinguish security rejections from availability failures in logs even though the public `AgentLoopHostError` no longer carries that information. All 23 hooks_integration tests + 43 reborn lib tests still pass. * fix(hooks): host-owned per-build hook-gate factory builder (serrrfirat MEDIUM on PR nearai#3633) `RouterBackedHookGateRefFactory::try_new` takes a caller-supplied `Fn() -> HookGateReservationContext` closure, and the host factory's `with_hook_gate_ref_factory(Arc<dyn ...>)` stored ONE factory instance that was reused for every host build. That instance carried whatever run/actor context its closure captured at construction time — so a second host build could mint a gate ref against the FIRST build's `LoopRunContext`. The verification test that backdated `resolved_at` proved the router rejects stale timestamps, but didn't address the host-side wiring footgun. Added per-build callback path: - New `HookGateRefFactoryBuilder` type alias for `Arc<dyn Fn(&LoopRunContext) -> Arc<dyn HookGateRefFactory>>`. - `with_hook_gate_ref_factory_builder(F)` on `RebornLoopDriverHostFactory` installs the callback. It runs once per `build_text_only_host*` call with the active `LoopRunContext`, so production callers wire `move |run_ctx| Arc::new(RouterBackedHookGateRefFactory::try_new(..., ttl, || HookGateReservationContext::new(run_ctx.clone(), actor.clone()))?)` and the factory is constructed fresh per host with no stale capture. - Build path consults the builder first, falls back to the shared `hook_gate_ref_factory` if only the older API is wired. - `with_hook_gate_ref_factory(Arc<dyn ...>)` is marked `#[deprecated]` pointing to the builder. The method body is unchanged for back-compat. Tests: - Existing integration tests migrated to the builder API (`with_hook_gate_ref_factory_builder({ let f = Arc::new(factory); move |_| Arc::clone(&f) })`) so they exercise the same logical wiring against the new seam. All 23 pass; clippy clean with `-D warnings`. The trait/router types and `validate_token` route are unchanged from the previous fix in this PR.
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…nearai#3912/nearai#3913 (nearai#3921) * refactor(hooks): align ExtensionId/HookLocalId with newtype template Address henrypark133 approval-with-followup items from PR nearai#3912: - L1: remove dead empty-id guard in HookManifestEntry::validate. HookLocalId::new now rejects empty strings at construction, so manifest deserialization fails before validate() is ever called. - L2: add AsRef<str>, From<Self> for String, and into_inner() to ExtensionId and HookLocalId per the canonical newtype template in .claude/rules/types.md. into_string() is retained as a thin alias for compatibility; new code should prefer into_inner(). - L3: document why the builtin_id_distinct_from_extension_id test fixture substitutes "path.module" for the original "path::module" (the new grammar rejects colons in HookLocalId). No behavioral change. * fix(hooks): fire AfterCapability observer for hook-suspended entries after allowed entries henrypark133 M1 on PR nearai#3911: in `HookedLoopCapabilityPort::invoke_capability_batch`, when Phase 1 produced a mix of Pending (hook-allowed) and Resolved (hook-suspension) slots with the suspension appearing AFTER an allowed entry and `stop_on_first_suspension = true`, the merge loop initialized `stopped_on_suspension` from `stopped_in_preflight` and broke after the very first iteration. Result: trailing Resolved suspension slots never fired their `AfterCapability` observer and never surfaced in the merged `outcomes` vec, violating the per-entry observer contract from PR nearai#3573 (serrrfirat P2 #3). Fix: continue iterating the merge loop so every slot fires its observer and every Resolved outcome is pushed. Only Pending slots are dropped after a stop (their inner work was already short-circuited in Phase 1 or by an early inner-port stop), tracked via `pending_after_stop`. Re-pop guard on `inner_outcomes.pop()` keeps the previous "inner stopped early on its own suspension" semantics: pending slots without an inner outcome are dropped, but the loop continues so any trailing Resolved observers still fire. New regression test `batch_invocation_fires_observer_for_hook_suspended_entry_after_allowed_entry_with_stop_on_first_suspension` pins the behavior: `[alpha=hook-allowed, beta=hook-suspension]` with `stop_on_first_suspension = true` produces a 2-entry `outcomes` vec (Completed alpha + ApprovalRequired beta), fires the observer twice, and only sends alpha to the inner port. Verified TDD-style: the test fails on the pre-fix code with `outcomes.len() = 1`. * perf(hooks): reuse serialized argument bytes in lazy resolve henrypark133 L1 on PR nearai#3913: `resolve_arguments` measured the post-resolver JSON payload by calling `serde_json::to_vec(&value)` and discarding the `Vec<u8>` once its length was checked. The `SanitizedArguments::from_json` constructor on the happy path sanitizes the in-memory `serde_json::Value` directly without re-serializing, so the materialized buffer was pure overhead. Switch the size measurement to `serialized_len`, a counting `io::Write` adapter that streams `serde_json::to_writer` into a u64 counter — saves one Vec<u8> allocation and the matching drop per resolved invocation. Behavior is identical: the same JSON encoding rules drive both writers, the cap check still fires when the encoded length exceeds `MAX_PREDICATE_INPUT_BYTES`, and serialization errors still fail closed. No new tests required; existing `dispatch_fails_closed_when_input_exceeds_max_bytes` and the ordering / lazy-probe tests exercise this path. * test(hooks): cover mixed needs_input hook binding short-circuit henrypark133 L2 on PR nearai#3913: add `before_capability_needs_input_returns_true_when_any_active_binding_needs_input`. Installs two BeforeCapability bindings on the same scope (Global) — one `needs_input() = false`, one `needs_input() = true` — and asserts both ends of the short-circuit: 1. `HookDispatcher::before_capability_needs_input(None)` returns true. 2. Driving `HookedLoopCapabilityPort::invoke_capability` with an instrumented `ProbingResolver` confirms the resolver IS consulted exactly once — i.e. the short-circuit fires through the call site, not just the helper. This pins the "any input-needing binding wins" contract end-to-end so a future change to the dispatcher's probe (or to the middleware's lazy-probe gate) can't silently regress to short-circuiting on the first binding only. Also tightens the merge-loop's pending-slot drop path to use `Option::map` (clippy::manual_map) — cosmetic, no behavior change.
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…3920) * Implement installed WASM hook runtime Adds crates/ironclaw_hooks/docs/threat-model-wasm.md and follows the reviewed design ack: 1) module bytes are resolved, digest-cached, and compiled in the tool-WASM style while reusing its resource limiter; 2) each invocation gets a fresh wasmtime Store; 3) the ABI is a wasmtime::Linker surface, not wit-bindgen; 4) host-import sink shims enforce call, patch-byte, observer-fact, and decision budgets. * Harden WASM hook string and metadata budgets * fix(hooks): validate WASM hook ABI at install time (serrrfirat #3 on PR nearai#3634) Address serrrfirat MEDIUM finding #3: `WasmHookRuntime::prepare()` compiled and cached module bytes but did not validate imports or the requested export. ABI mismatches (unsupported import, missing export, wrong export signature) were deferred to first live dispatch — and the prior `wasm_unsupported_host_import_fails_closed` test codified that a bad-import module would install successfully and only fail closed at invocation. Malformed untrusted modules should never reach live traffic. Changes: - `prepare()` derives the target hook point from `request.kind`, then runs `validate_module_abi()`: scratch-instantiate the module against the point-specific linker (catches unsupported / wrong-type imports) and resolve the typed export `() -> ()` (catches missing export and wrong signature). Failures surface as new `WasmHookRuntimeError::InvalidImports` or existing `WasmHookRuntimeError::InvalidExport`, both of which bubble up as `HookError::RegistryConstruction` from the registrar. - `wasm_point_for_kind(HookManifestKind)` helper centralizes the kind → wasm-point mapping; the previous `execute_*` paths can share it in a follow-up but kept inline for now to minimize churn. Tests: - `wasm_unsupported_host_import_is_rejected_at_install_time`: replaces the prior test that codified late-failure behavior; asserts the registrar returns `RegistryConstruction` citing the bad import. - `wasm_missing_export_is_rejected_at_install_time`: new module that compiles but lacks the manifest-declared export; same install-time rejection. * fix(hooks): address henrypark133 must-fix #1, #2, #3 on PR nearai#3634 Three items from the 5-15 review: **#1 (must-fix) Extract ironclaw_wasm_limiter micro-crate** Replace `#[path = "../../../ironclaw_wasm/src/limiter.rs"]` cross-crate file import with a proper Cargo edge. The 111-line `WasmResourceLimiter` moves into a new `crates/ironclaw_wasm_limiter` micro-crate that both `ironclaw_wasm` and `ironclaw_hooks` depend on. The architecture rule forbidding `ironclaw_hooks -> ironclaw_wasm` is preserved (the new crate sits below both consumers and pulls in only `wasmtime` + `tracing`); `cargo check`, `cargo doc`, and architecture-linting tests now see the edge, and the file can't be moved out from under one of the consumers silently. Mechanical changes: - new `crates/ironclaw_wasm_limiter/` (Cargo.toml + src/lib.rs with the type exposed as `pub` instead of `pub(crate)`) - workspace `members` entry added - `crates/ironclaw_wasm/src/limiter.rs` deleted - `crates/ironclaw_wasm/src/lib.rs`: `mod limiter` removed - `crates/ironclaw_wasm/src/store.rs`: import switched to `ironclaw_wasm_limiter::WasmResourceLimiter` - `crates/ironclaw_wasm/Cargo.toml`: dep added - `crates/ironclaw_hooks/Cargo.toml`: dep added - `crates/ironclaw_hooks/src/wasm/runtime.rs`: `#[path = ...]` block removed; import switched to the crate **#2 + #3 (must-fix) Dead WASM arms in dispatch** `run_before_capability_hook`, `run_before_prompt_hook`, and `run_observer_hook` each had an early-return guard that dispatched WASM hooks with `catch_unwind` + timeout, then ALSO had a matching WASM arm in the inner `match` that ran without those protections. The prompt-path arm additionally swallowed `WasmHookFailure` via `|_| ()`, making the must-fix #2 problem worse on that path specifically. If a future refactor removed any of the early-return guards, those inner arms would silently take over and drop panic isolation, deadline enforcement, AND (for prompts) the failure category. Replaced each inner arm with `unreachable!()` carrying a comment that explains why the arm exists and references the early-return guard above it. A future refactor that removes the guard will now trip the `unreachable!` at first call instead of silently degrading. All 154 hooks lib + 29 reborn integration tests still pass. * fix(hooks): plumb context to WASM hooks + runtime hardening Critical #1 on PR nearai#3634: WASM hooks previously received no context. The `execute_*` entry points dropped the `&BeforeCapabilityHookContext` / `&BeforePromptHookContext` / `&ObserverHookContext` value and invoked the guest export with `()`, so a WASM gate could never decide based on the capability name, tenant, provider, or other dispatch-time facts. Add an `ic:hooks/context@1` host-import module exposing two read-only calls — `ctx_size() -> i32` and `ctx_read(ptr, len) -> i32` — backed by a JSON-serialized blob the dispatcher writes per-invocation into the fresh store. Modules that don't import these continue to link; modules that do import them get a stable, non-empty payload to read. An integration test (`wasm_before_capability_hook_reads_context_blob`) asserts the contract end-to-end: a guest that fails to read a non-empty blob traps before its `deny` call. Also rolls up the other reviewer-flagged WASM runtime issues, all of which touch `wasm/runtime.rs`: HIGH #2: epoch-tick background thread now holds a shutdown `AtomicBool` and joins on `Drop`. Previously it looped forever and leaked an Engine clone on every runtime drop. MED #4: compiled-module cache is now an `lru::LruCache` bounded by `MODULE_CACHE_CAPACITY = 128`. Replaces the unbounded `HashMap`. MED nearai#7: `prepare()` no longer compiles under the cache lock. Fast path reads from LRU under a brief lock; slow path compiles outside the lock and re-checks on insert to avoid the TOCTOU window where two concurrent installs of the same module both compile. Bug nearai#9: post-call `deadline_exceeded()` re-check on the Ok branch is gone. wasmtime epoch-interrupt is the authoritative wall-clock signal; an Ok return is no longer reclassified as a timeout because the wall ticked over during host-side return. Bug nearai#10: `add_milestone_metadata` returns a distinct "metadata value exceeds the u32 byte-length ceiling" error when the guest-supplied `value.len()` overflows u32, instead of misreporting it as "exceeded total prompt-patch byte budget". Existing integration tests for WASM hooks are also re-wired through `HookRegistrar::with_verified_grants` so the grants-store gate added in the foundation-01 merge stops failing the pre-existing fixtures. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): run WASM hooks on the blocking pool HIGH #3 on PR nearai#3634: `tokio::time::timeout` does NOT cancel synchronous wasmtime execution. The previous code awaited a `catch_unwind(async { h.evaluate(ctx) })` future whose body completed in one poll, so the timeout could only fire *around* the WASM call rather than against it; a hook that wedged inside wasmtime simply pinned the calling tokio task. Route gate, prompt, and observer WASM dispatch paths through `tokio::task::spawn_blocking` via a shared `run_wasm_blocking` helper. The outer `tokio::time::timeout` now governs the JoinHandle, so a stuck blocking task stops blocking the dispatcher's caller; the wasmtime epoch interrupt configured in the runtime (10 ms tick) is the authoritative in-WASM wall-clock cancel signal. JoinError (panic in the blocking task) maps to `FailureCategory::Panic`, matching the pre-existing semantics for synchronous panics caught via `catch_unwind`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * perf(hooks): O(1) hook-id lookup via side index Finding nearai#8 on PR nearai#3634: `set_priority`, `poison`, `is_poisoned`, and `contains_hook` all did full-registry scans over every binding at every point. Each is called per-dispatch (poison-checks on the snapshot loop in particular), so the cost is `O(registered_hooks)` per `(installed_hook, registered_hook)` pair. Maintain a denormalized `HashMap<HookId, (HookPointSpec, usize)>` side index in lock-step with `by_point` so every per-hook-id operation becomes a single hash lookup + a direct vec indexed access. The duplicate-id rejection in `insert` now reads from the side index too, turning what used to be a flat-map scan into a `HashMap::contains_key`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(hooks): wall-clock timeout, observer memory, limiter rollback, registrar happy path Round out the test set for the WASM hook execution path: nearai#11 / nearai#12: gate + observer wall-clock timeout. The pre-fix dispatcher ran wasmtime synchronously on the executor, so the outer `tokio::time::timeout` `Err(_elapsed)` arm was effectively unreachable. Now that WASM execution runs on the blocking pool, the timeout actually fires; the new tests give the wasm budget headroom (1B fuel, 5s wall) and the dispatcher a 20 ms timeout, then assert the failure classification (FailClosed for gate, FailIsolated for observer). nearai#13: observer memory exhaustion. Mirrors `wasm_memory_exhaustion_fails_closed_for_gate` against the observer dispatch path so the FailIsolated branch of the failure matrix has explicit memory coverage, not just fuel/wall. nearai#15: `WasmResourceLimiter::memory_grow_failed` rollback. Stages an approved grow, simulates the OS-level grow failing, and asserts a subsequent grow of the full ceiling succeeds — the inflated `memory_used` from the failed attempt must be released. nearai#16: registrar WASM happy path. Companion to the existing `install_wasm_body_requires_runtime` negative case: a valid module installs, the binding is visible via the public registry accessor, and is not pre-poisoned. nearai#14 (`add_milestone_metadata` happy path) is intentionally omitted — the BeforePrompt dispatch path is currently unreachable due to a pre-existing manifest-vs-registry scope conflict (`OwnCapabilities` is the only valid `BeforePrompt` scope per manifest validation, but the registry rejects `OwnCapabilities` at `BeforePrompt` because the point has no provider context). That contradiction sits outside this PR's scope; flagging for a follow-up. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(hooks): typed WASM version material, reconcile design doc LOW nearai#20 on PR nearai#3634: extract the `{extension_version}+wasm:{module_digest_hex}` concatenation into a `WasmVersionMaterial` newtype with a single `Display` impl. The identity material no longer floats free as a stringly-typed argument inside the registrar. Reconcile `docs/successors/02-wasm-runtime.md` with the implementation: - Spell out that wall-clock cancellation depends on the `tokio::time::timeout(tokio::task::spawn_blocking(...))` pair, and explain why a bare timeout over a synchronous wasmtime call cannot actually cancel. - Define `FailIsolated` and `FailClosed` as `FailureDisposition` values, distinct from the older `HookFailureMode::{FailOpen, FailClosed}` policy switch that applies to predicates. - Clarify the generic `evaluate` export contract — name is whatever the manifest declares, signature is `(): ()`, context arrives through the new `ic:hooks/context@1` host imports — and note the intentional divergence from `WitToolRuntime`'s hardcoded interface. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): drop .expect() in WASM module cache capacity Pre-commit no-panics CI flagged the .expect() on the LruCache capacity. Move the validity check to a const match, so the NonZeroUsize is fixed at compile time and the no-panics regex is satisfied. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): use HookLocalId::new after newtype privatization The newtype-privatization landed in reborn-integration after the hooks-fu-wasm-runtime branch's WASM scaffolding tests were written; update the affected test/registrar sites to use HookLocalId::new instead of the now-private tuple constructor. * style: cargo fmt after newtype-privatization fixups * test(hooks): ignore 3 BeforePrompt WASM tests with manifest/registry conflict These tests were failing on the original branch tip too (verified against origin/hooks-fu-wasm-runtime @ 571efdf). The Installed-tier BeforePrompt WASM install path has no valid scope today: - OwnCapabilities is rejected by the registry C3 check (finding #2 on PR nearai#3573) since BeforePrompt has no per-capability invocation context. - SameTenant is rejected by manifest validation ("cannot combine scope = same_tenant with kind = before_prompt"). The budget-overflow paths these tests exercise are point-agnostic; the follow-up is to either rewrite the helper to install through BeforeCapability or add a Global manifest scope. Tracked as a deferred item on the new PR. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…i#3573) (nearai#3635) * docs(hooks): scope persistent predicate counter backend (successor #3) Successor PR from nearai#3573. Current sliding-window state is in-memory and resets on restart. Adds a PredicateStateBackend trait + Postgres/libSQL impls for cross-process and restart-survival semantics. * feat(hooks): extract PredicateStateBackend trait + replay-safe in-memory impl Addresses codex review's three Critical findings on PR nearai#3635: 1. Backend wiring: the trait is now registered (lib.rs:25-26) and PredicateEvaluator delegates to Arc<dyn PredicateStateBackend> via with_backend(...). Default constructor preserves the in-memory behavior so all 154 existing tests pass unchanged. 2. Atomic record-and-read: each record_invocation / record_value call performs the write AND returns the resulting in-window count/sum under a single mutex (in-memory) / transaction (durable backends). Splitting into separate record + read would let two hosts each see 'under cap' and both proceed, drifting past max. 3. Replay refusal: each record call carries a PredicateEventId. Re-emitting the same event_id is a no-op against the count. In-memory backend implements via a per-key bounded set (RECENT_EVENT_ID_CAP = 256); durable backends will use INSERT … ON CONFLICT DO NOTHING. Trait surface (predicate_state.rs): - PredicateEventId(String): opaque dedup key - PredicateBackendError: thiserror enum for fallible durable backends; in-memory backend never returns Err - PredicateStateBackend trait with Result return types - InMemoryPredicateStateBackend default impl - MAX_HISTORY_KEYS const re-exported via evaluator for back-compat Evaluator changes (evaluator.rs): - holds Arc<dyn PredicateStateBackend> (no more inline maps) - evictions_observed() reads through to backend - synth_event_id() generates per-call-unique ids via a process-local atomic counter so tests with identical (hook, ctx, now) still produce distinct ids - LRU helpers + HistoryKey/ValueHistoryKey types moved into predicate_state.rs (as InvocationKey/ValueKey) Tests: - 6 new predicate_state tests: - in_memory_invocation_counts_within_window - in_memory_invocation_trims_outside_window - in_memory_value_sums_within_window - in_memory_tenant_isolation (regression on threat-model C2) - in_memory_duplicate_event_id_is_a_noop_for_invocations - in_memory_duplicate_event_id_is_a_noop_for_values - 160 unit tests pass total. Reborn hooks_integration unchanged at 19 scenarios. Clippy/fmt/no-panics clean. Sync trait + Instant timestamps documented as a v1 choice; durable backends (Postgres, libSQL) will need an async companion trait using SystemTime — tracked in the scope doc as the next slice. Scope doc: crates/ironclaw_hooks/docs/successors/03-persistent-counter.md * fix(hooks): close codex P1 bugs in PredicateStateBackend in-memory impl Addresses codex P1 review on PR nearai#3635: P1 #1 — replay dedup loss under high-throughput keys The prior design used a fixed-size (256) recent_ids ring per bucket decoupled from entries. Under any workload with >256 distinct events in the same window, the first event's id aged out of the ring while its timestamp entry was still live, so a replay silently re-counted. Fix: dedup memory is now intrinsic to entries. Each entry stores (timestamp, event_id), and the dedup check is 'does any in-window entry have this id?'. Dedup memory is therefore exactly the in-window entry set — no fixed cap, no silent loss. P1 #2 — zombie buckets clogging LRU Two-part fix: 1. record_* drops empty buckets eagerly via history.remove(key). This is mostly defense-in-depth — under the new dedup design, the record path can't actually leave a bucket empty (proved in the test rationale comment). 2. evict_lru_* now preferentially targets empty buckets first (find any v.entries.is_empty()), only falling back to the oldest-timestamp scan if no empty bucket exists. Filter-out behavior is gone, so any empty bucket that somehow survives becomes the next eviction victim instead of a permanent zombie. Test changes (+2 new, -0 removed): - dedup_memory_covers_full_window_under_high_throughput: pushes 512 distinct events into one bucket, then replays event-0. Pre-fix this would have counted again (silent dedup loss); post-fix the replay is a no-op. - lru_evicts_empty_buckets_first: crafts an empty bucket alongside a live one, runs LRU eviction, asserts the empty one is evicted and the live one retained. Tests: 162 unit total (+2 new). Clippy/fmt/no-panics clean. * docs(hooks): address gemini review on persistent-counter scope doc Four medium-priority doc nits from gemini-code-assist on the crates/ironclaw_hooks/docs/successors/03-persistent-counter.md scope: 1. run_id in the trait: removed. The trait dedupes on event_id (RuntimeEventId is already run-scoped), not run_id. Replaces the earlier 'backend stores (timestamp, run_id, event_id)' claim. 2. SystemTime vs chrono::DateTime<Utc>: switched to DateTime<Utc> to match project convention (src/db/mod.rs, ironclaw_events). The in-memory backend keeps Instant for monotonic process-local semantics; durable backends require DateTime<Utc> for cross- process serialization. Documented as a clock note. 3. libSQL TEXT column for rust_decimal: per src/db/CLAUDE.md, libSQL can't preserve Decimal precision with numeric/real types. LibSqlPredicateStateBackend serializes value as TEXT via Decimal::to_string() / from_str(). Postgres impl keeps numeric (correct for PG). Documented as the two LibSql-specific schema differences. 4. Batched-writes vs cross-process consistency tension: gemini was right that deferring writes to the tick boundary breaks requirement #1 (two hosts would each see 'under cap' simultaneously). v1 production backend keeps writes synchronous; future optimization batches reads (not writes). * fix(hooks): thread stable caller_event_id through hook context (replay dedup) henrypark133 HIGH on PR nearai#3635 + serrrfirat HIGH #1: the `PredicateBackedBeforeCapabilityHook -> PredicateEvaluator` path always synthesized a fresh `event_id` per evaluation by mixing in a process-local atomic counter, so the same logical invocation retried/replayed always got a different id. The backend's UNIQUE constraint on `event_id` — the load-bearing dedup contract — never engaged on the real production path. Replay dedup was effectively "documented but unused." Plumb a stable per-invocation identity through the public hook surface: - `BeforeCapabilityHookContext` gains a `caller_event_id: Option<PredicateEventId>` field. Middleware that threads through from the calling layer's runtime event identity populates `Some(...)`; older / in-memory-only callers pass `None` and degrade to the current synth path (no behavior change). - New builder method `with_caller_event_id(...)`. - `PredicateEvaluator` resolves the id through a new `resolve_event_id` helper: prefer `ctx.caller_event_id`, fall back to `synth_event_id`. Both `record_invocation` and `record_value` paths use it. - Backend dedup behavior is unchanged — it was already correct on `event_id`. The bug was the caller path never supplying a stable id. Tests (caller-boundary, henrypark133's required regression): - `duplicate_caller_event_id_is_deduped_in_invocation_count`: two evaluations with the same `caller_event_id` count as one invocation; a third with a different id counts as two; a fourth crosses the cap. Sanity branch confirms the no-id synth path still exhibits "every call counts" semantics. This is the API contract slice. Wiring the middleware to actually supply a stable id (e.g. derived from the originating `RuntimeEventId` once that runs through the BeforeCapability path) is the follow-up that lights up the durable backend's end-to-end replay-safety promise. * refactor(hooks): demote PredicateStateBackend to pub(crate) (serrrfirat MED on PR nearai#3635) serrrfirat MED: the `predicate_state` module exposed `PredicateStateBackend` as `pub`, but the trait's `now: Instant` parameter is process-local and not serializable. Any external durable backend impl built against the current trait would have to be rewritten when the durable contract lands with `chrono::DateTime<Utc>` (see successor doc 03-persistent-counter.md). Hold the public surface back until that contract is stable so we don't ship a public API we know we'll break. Demoted to `pub(crate)`: - `PredicateStateBackend` (trait) - `InvocationKey`, `ValueKey` (key types — backend ABI only) - `PredicateBackendError` (error type, with `#[allow(dead_code)]` on the `Unavailable` variant since the in-memory backend is infallible and no durable backend exists yet) - `InMemoryPredicateStateBackend` (the only impl) - `PredicateEvaluator::with_backend` (with `#[allow(dead_code)]` — reserved for future internal injection paths) Kept `pub`: - `PredicateEventId` — it appears on the public hook surface via `BeforeCapabilityHookContext::caller_event_id` (from the nearai#3635 HIGH fix). Hook authors who want stable replay-dedup ids construct one. No behavior change. All 163 hooks lib tests + 19 reborn integration tests still pass. * fix(hooks): address henrypark133 must-fix #1-5 on PR nearai#3635 Five items from the 5-15 review: **#1 (must-fix) O(n) dedup scan** The previous `bucket.entries.iter().any(...)` linear scan held the outer history mutex while walking thousands of in-window entries at high throughput. Add a companion `HashSet<PredicateEventId>` per bucket (`InvocationBucket.dedup_ids` / `ValueBucket.dedup_ids`), maintained alongside the deque via `pop_front`/`push_back` helpers. O(1) dedup, same correctness, same memory bound (one set entry per in-window entry — no fixed ring). **#2 (must-fix) Mutex poison cascade** `.expect("predicate history mutex poisoned")` propagated a panic to every subsequent caller. Replace with `match self.invocation_history.lock() { Ok(g) => g, Err(p) => p.into_inner() }` so a poisoning thread doesn't take down all subsequent evaluations. **#3 (must-fix) `caller_event_id` format validation** `with_caller_event_id` now rejects empty strings and ids containing NUL bytes. Failed validation logs a `tracing::warn!` and leaves `caller_event_id == None` so the synth path takes over — operator sees the warning, predicate dedup still works. Also: `PredicateEventId(pub String)` → `PredicateEventId(String)` with `new()` / `as_str()` (henrypark133 nit nearai#9). Inner field is no longer in-place mutable from outside the crate. **#4 (must-fix) `with_backend` is `#[cfg(test)]`** Previously `#[allow(dead_code)]` — reachable from release builds and inviting future callers to inject backends through an unstable seam. Gated to `cfg(test)`. **#5 (important) `evict_older_than` trait stub** Default-impl no-op added to `PredicateStateBackend` so the trait signature is locked before the first durable-backend PR. Trait-object callers won't break when durable impls override it. **Bonus** (henrypark133 missing-coverage #1): `in_memory_record_invocation_is_atomic_under_concurrent_writers` — 32 threads each record a distinct event id; final count must equal 32, proving the atomic record-and-read contract holds under contention. **Bonus** (henrypark133 nit nearai#10): The third stable id in `duplicate_caller_event_id_is_deduped_in_invocation_count` was 62 chars; bumped to 64 to match the synth output format. * fix(hooks): clippy doc-list-indentation + remove unused with_backend (nearai#3635 CI) * fix(hooks): address serrrfirat HIGH + MEDIUM on PR nearai#3635 (5-15 review) **MEDIUM — `caller_event_id` validation bypass** `with_caller_event_id` validated for empty/NUL but the field on `BeforeCapabilityHookContext` is `pub`, so callers could direct- assign `Some(PredicateEventId::new("..."))` with `new()` permissive and bypass the setter entirely. Move validation INTO the type boundary: - `PredicateEventId::new(...) -> Result<Self, PredicateEventIdError>` validates non-empty + NUL-free at construction. Any value that reaches a downstream backend now satisfies the format invariant by construction. - `PredicateEventId::new_unchecked(...)` for internal synth paths and tests that mint ids from known-good shapes (hex digests). - `with_caller_event_id` drops its now-redundant runtime check; the type already enforces it. - Internal synth in `evaluator.rs` switches to `new_unchecked` (64-char hex output is always valid by construction). Tests: - `predicate_event_id_rejects_empty` - `predicate_event_id_rejects_nul_bytes` - `predicate_event_id_accepts_typical_hex_digest` **HIGH — durable schema: dedup scope mismatch** The successor doc's Postgres schema declared `event_id uuid PRIMARY KEY` (globally unique), but the trait's replay-refusal contract dedupes within the counter `key`. `caller_event_id` is per capability invocation — two predicate-backed hooks observing the same invocation share an id. A global PK lets the first hook's INSERT win and silently undercounts the second hook's bucket. - `docs/successors/03-persistent-counter.md`: PK changes to composite `(tenant_id, hook_id, capability, event_id)` for invocations and `(tenant_id, hook_id, capability, field, event_id)` for values, matching the trait's per-key dedup scope. - `predicate_state.rs` trait doc: replay-refusal section rewritten to spell out the per-key scope and the corresponding `INSERT … ON CONFLICT (tenant, hook, capability[, field], event_id) DO NOTHING` shape durable backends should use. * docs(hooks): document host-assigned trust boundary on PredicateEventId henrypark133 / serrrfirat blocker B4 on PR nearai#3635: the `caller_event_id` threading through `BeforeCapabilityHookContext` partially shipped earlier (commit b4d8a35), but the trust-boundary documentation explaining the host-assigned invariant was still missing. Add rustdoc to `PredicateEventId` and the `PredicateStateBackend` trait clarifying that: - the id MUST be minted by trusted host code from authoritative sources (dispatcher RuntimeEventId, host-side hash, arguments digest) - it MUST NOT pass through unchanged from any tenant-controlled surface (capability arguments, manifest fields, WASM memory, HTTP bodies) - the format invariants in `PredicateEventId::new` (non-empty, NUL-free) are a durability contract for SQL backends, NOT a trust check - a tenant-supplied id can either undercount itself into infinity by replaying a fixed id, or poison adjacent buckets if scoping is ever weakened Doc-only; no behavior change. * test(hooks): add caller-boundary replay-dedup test through wrapper hook henrypark133 HIGH blocker B1 on PR nearai#3635: replay dedup must engage at the caller boundary — `PredicateBackedBeforeCapabilityHook::evaluate` is the production path the dispatcher invokes for installed predicate hooks. A unit test on `PredicateEvaluator::evaluate_at` alone is insufficient regression coverage (repo CLAUDE.md rule "Test through the caller, not just the helper"): the wrapper hook reads `BeforeCapabilityHookContext::caller_event_id` and threads it down to the backend, so the regression test must drive the wrapper itself. The threading work already shipped in commit e6df47d (`caller_event_id` field on the public hook context + evaluator preferring it over the synth path). This commit adds the missing end-to-end test: 1. Two `PredicateBackedBeforeCapabilityHook::evaluate` calls with the same `caller_event_id` and a `RateOrValueCap { max: 1 }` predicate — the second call must stay under cap (dedupe engages at the wrapper boundary, not be re-counted into a deny). 2. A third call with a DISTINCT `caller_event_id` crosses the cap — proving dedup is replay-scoped (same id → no-op), not blanket- suppress (any id → no-op). If the wrapper were synthesizing a fresh id per call (the bug Henry flagged before threading landed), this test would fail at step 2 with the second evaluation being denied. * docs(hooks): D5a + cross-process replay note; add caller-API tests henrypark133 should-fix S8 + S9 on PR nearai#3635. S8 — threat-model expansion: - Add D5a as the correctness-under-attack variant of D5: an attacker flooding high-cardinality keys can LRU-evict legitimate tenants' counters and reset their rate-limit state. Distinct from the memory-only framing of D5; tied back to per-extension caps (D3/D4) and the durable-backend successor (doc 03). - Document the cross-process replay limit on the in-memory backend inside the PredicateStateBackend trait docs, not just in D5 — the process-local dedup is a property callers need at the trait surface, with a pointer to the durable backend as the cross-host story. S9 — three new tests on the in-memory backend public API: - lru_eviction_via_public_api_holds_max_history_keys_cap: drives MAX_HISTORY_KEYS + 1 distinct keys through record_invocation and asserts the map size cap holds + evictions_observed() advances. The previous coverage manually crafted buckets and called the LRU helper directly; this exercises the production path. - in_memory_invocation_retains_entry_at_exact_window_cutoff: pins the `< cutoff` trim semantics so a refactor to `<=` would fail loud. - event_id_dedup_is_isolated_across_invocation_and_value_maps: same event_id used in both record_invocation and record_value must not cross-suppress — the two maps key on disjoint types. The fourth S9 item (concurrent N-thread atomicity) and the caller- boundary replay test on the wrapper hook already landed in earlier commits (f632d22, predicate_state.rs line 840). S2 (evict_older_than stub), S3 (sync-trait docs), and S7 (consistency vs batched-writes) were also already in HEAD; this commit ships the remaining items. Quality gate: cargo fmt clean, cargo clippy -p ironclaw_hooks --all-features --tests -D warnings clean, full hooks test suite green (15 predicate_state unit tests + lib + integration). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore(hooks): co-locate synth_event_id with backend; rationale comment; pin synth format henrypark133 nits N1, N2, N5 on PR nearai#3635. N1 — Move `synth_event_id` from `evaluator.rs` to `predicate_state.rs` as `PredicateEventId::synth(...)`. The id format (64-char lowercase hex, no NUL, never empty) is part of the backend's durable contract, so co-locating with `PredicateStateBackend` keeps the format change- surface adjacent to the consumer. To avoid inverting the module dependency (`predicate_state` is a leaf below `points`), the synth helper takes raw bytes / &str rather than a `&BeforeCapabilityHookContext`. The evaluator's `resolve_event_id` fallback unpacks the context and delegates. N2 — `// safety:` comment on a non-`unsafe` block (the `write!(s, "{byte:02x}")` infallibility note) renamed to `// RATIONALE:`. By convention `// SAFETY:` pairs with `unsafe` blocks; using `// safety:` elsewhere conflates the two. N5 — Add `synth_event_id_is_64_char_lowercase_hex` to pin the synth output shape. A refactor that silently changes length or case would break the durable backend's `uuid`-shaped UNIQUE constraint without a test failure today; the new test fails loud. Quality gate: cargo fmt clean, cargo clippy --all --benches --tests --examples --all-features -D warnings clean, full hooks lib test suite green (172 passing including the new pin). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): drop expect() in hex formatting to satisfy panic CI check The "No panics in production code" CI check (scripts/check_no_panics.py) only recognizes `// safety:` suppression markers, not `RATIONALE:`. Since std::fmt::Write for String is infallible, just discard the Result with `let _ =` instead of `.expect()` — no panic call, no marker needed. Also merges in latest origin/hooks-foundation-01 (now includes the reborn-integration merge and PR nearai#3636). * fix(hooks): narrow caller_event_id visibility to pub(crate) henrypark133 MED on PR nearai#3635 5-19 review. The pub field let external callers bypass with_caller_event_id and assign values the validated PredicateEventId constructor would have rejected. Force every external caller through the typed setter so PredicateEventId::new is the only entry point. * fix(hooks): drop arguments_digest from synth + add in-memory backend warn Two PR nearai#3635 5-19 review findings on the evaluator surface: - henrypark133 LOW (synth oracle): drop arguments_digest from the PredicateEventId synth hash input. The 64-char hex output was an equality oracle for argument shape; replay dedup for durable backends uses the caller-supplied caller_event_id, not the synth path, so synth only needs to be per-call unique, not content-addressed. - henrypark133 HIGH + MED (in-memory production limits): expose PredicateEvaluator::warn_in_memory_backend_active_in_production for hosts to call at startup. Multi-host replay dedup is process-local and the LRU cap is shared across tenants; operators need this surfaced in logs when the durable backend is not wired. * fix(hooks): harden predicate state backend per PR nearai#3635 5-19 review Address five findings on crates/ironclaw_hooks/src/predicate_state.rs: - A1 (henrypark133 HIGH): restrict PredicateEventId::new_unchecked to pub(crate) so external callers cannot bypass the durable UNIQUE-constraint format invariants enforced by ::new. - A4 (henrypark133 MED): per-tenant LRU quota at MAX_HISTORY_KEYS / 4. Without it a noisy tenant could fill the global cap and evict a quiet tenant's bucket, resetting their rate-limit counter. With the quota, a tenant that overflows evicts its OWN oldest-front bucket first. New tests cover single-tenant cap and cross-tenant isolation. - A5 (henrypark133 LOW): drop arguments_digest from the synth hash input (oracle closure mirrored from the evaluator side). Add a thread-local nonce alongside the process-global counter so synth remains per-call unique without relying solely on a contended AtomicU64. New test pins the divergence invariant. - D6 (henrypark133 HIGH): O(1) NumericSum via an incrementally- maintained ValueBucket::running_sum, replacing the O(n) deque walk on every record_value call. New test covers push/trim/replay interactions. - D8 (henrypark133 MED): implement evict_older_than for the in-memory backend (was a no-op Ok(0) default). Drops entries strictly older than the cutoff and removes empty buckets; operator reaper tasks rely on this to reclaim memory from idle keys. - D7 (henrypark133 MED, partial): document the process-global synth COUNTER as a known contention hotspot and add a thread-local nonce so threads can advance without forcing cross-core invalidation in the common path. Tests: 196 passing (+4 new); workspace clippy clean. * fix(hooks): port MAX_SAMPLES_PER_KEY cap into PredicateStateBackend (D5 regression from r3) Round 3 of PR nearai#3573 (already merged into hooks-foundation-01) added an inline per-key sample cap of 4_096 in evaluator.rs to bound memory under attacker-triggered hot capabilities with very large declared windows (threat-model finding D5). The predicate-state extraction in PR nearai#3635 moved that bookkeeping into the PredicateStateBackend trait but missed porting the cap, so the cap would silently disappear from production once this PR rebases onto the foundation branch. This commit moves the cap into the in-memory backend impl next to MAX_HISTORY_KEYS / MAX_KEYS_PER_TENANT and enforces it in both record_invocation and record_value. For the NumericSum path, the bucket helper's pop_front already decrements running_sum, so the incremental sum invariant survives cap-driven eviction. The pre-existing inline copy in evaluator.rs becomes redundant once the trait impl owns the enforcement; the rebase resolution deletes it. Adds two regression tests: - record_invocation_caps_samples_per_key_under_attacker_pressure - record_value_evicts_oldest_keeping_running_sum_consistent * fix(hooks): port predicate_state tests to ::new() after nearai#3912 newtype privatization --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
… to nearai#3573) (nearai#3637) * docs(hooks): scope invocation_arguments_digest snapshot pin (successor nearai#10) Successor PR from nearai#3573 — addresses serrrfirat's #3 follow-up. Adds an explicit snapshot test pinning the digest for a known input so a future change to the hashing path or input-ref format is loud. * test(hooks): pin invocation_arguments_digest with snapshot Address serrrfirat's #3 follow-up from PR nearai#3573 — partial fix promoted to full pin. Two new tests in capability_port.rs::tests: - invocation_arguments_digest_is_stable_for_known_inputs: pins the digest for a fixed (capability_id, input_ref) fixture so any future change to the hashing structure or input-ref format is loud. The captured hex is documented in the assertion message + stability contract. - invocation_arguments_digest_differs_for_different_input_refs: structural sanity check that distinct inputs produce distinct digests. Plus expanded rustdoc on the function calling out the stability contract: changing the hashing path requires updating the fixture, surfacing in the cross-crate wire-format contract section, and bumping the framework's contract version if downstream consumers exist. Tests: 156 unit tests pass (was 154; +2 new). Clippy + fmt clean. * test(hooks): pin arguments_digest at middleware boundary (serrrfirat nearai#3637) serrrfirat MED on PR nearai#3637: the existing snapshot test pins `invocation_arguments_digest`'s raw output, but the public hook contract is `BeforeCapabilityHookContext.arguments_digest` populated via `HookedLoopCapabilityPort::hook_context`. If caller-side wiring drifts — wrong field set, transform inserted, stale/default digest, or an alternate path bypassing the helper — the helper snapshot stays green while hook consumers observe a broken digest. Add a boundary-level pin: construct a `HookedLoopCapabilityPort` with a no-op inner port and an empty dispatcher, run the same fixed `(capability_id, input_ref)` invocation through `hook_context()`, and assert the resulting `ctx.arguments_digest` matches the same pinned hex as the helper snapshot. If they ever disagree, this assertion fails — surfacing wiring drift that the helper test alone cannot. `HookedLoopCapabilityPort::hook_context` is widened from private to `pub(crate)` to make the boundary test possible without bypassing the function. * docs(hooks): clarify arguments_digest rustdoc — input-ref identity, not arguments (serrrfirat blocker on PR nearai#3637) The rustdoc summary on invocation_arguments_digest described the digest as covering "capability arguments" with equivalence under "identical arguments". This contradicted the new stability section in the same file, which (correctly) documents that the digest is over the (capability_id, input_ref) identity tuple — NOT over the resolved argument content the input-ref points at. Reword the summary and the equivalence claim to describe input-ref identity, matching the stability section and the actual implementation. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…nearai#3640) * docs(hooks): scope event-triggered hooks (Phase 5, successor #4) Successor PR from nearai#3573. Adds a new EventTriggered hook point that subscribes to RuntimeEvents asynchronously, outside the loop's inline tick. Observer-only by construction (no Allow/Deny/Patch); typed against a narrowed HookObservableEvent projection to keep the cross-crate boundary clean. Scope doc only; design questions about cursor/replay semantics and per-extension event-rate caps need design review before implementation. * Implement Phase 5 event-triggered hooks Cites crates/ironclaw_hooks/docs/successors/04-event-triggered-hooks.md as the scope contract. Adds the EventTriggered observer hook point, durable RuntimeEvent dispatch path, and Reborn pull-driven subscription wiring with caller-level coverage for matching, replay, scope filtering, observer-only authority, and backpressure. * Fix hook event OwnCapabilities owner lookup * fix(hooks): carry owning extension into hook milestone runtime events henrypark133 HIGH + codex P1 on PR nearai#3640: `OwnCapabilities`-scoped event-triggered subscriptions silently never fired for `HookFailed`/`HookDecisionEmitted`/`HookDispatched` events because those `RuntimeEvent` constructors hardcoded `provider: None`. Since Installed hooks default to `OwnCapabilities`, the very events that Phase 5 was designed to observe (hook-failure / decision alerting) never reached their default-configured subscriber. A prior fix added a hook_id-based fallback in `scope_provider_for_runtime_event` that resolves the owning extension through the registry's hex index when `event.provider` is `None`. That covers the case where the failing hook is still registered at replay time, but the durable fix is to stamp the originating provider into the event at emit time so the primary `event.provider` path resolves without any fallback. Plumbed `owning_extension: Option<ExtensionId>` end-to-end: - `LoopHostMilestoneKind::{HookDispatched, HookDecisionEmitted, HookFailed}` gain the field (with `#[serde(default, skip_serializing_if = "Option::is_none")]` so pre-existing checkpoint payloads and the L3 schema-snapshot tests round-trip unchanged when no owner is set). - `RuntimeEvent::hook_{dispatched, decision_emitted, failed}` constructors accept the owner and stamp it into `provider`. - `milestone_events.rs` threads the field through the projection. - `HookDispatcher::emit_dispatched/emit_decision` pass `binding.owning_extension.clone()` directly. - `HookDispatcher::emit_failure` (no binding handy on the failure path) looks the owner up via the registry's existing `owning_extension_for_hook_hex` index. Tests: - `event_triggered_own_capabilities_matches_hook_failed_with_carried_provider`: primary-path regression — two `HookFailed` events with `provider: Some(ext_a|ext_b)` against an `OwnCapabilities` subscription scoped to ext_a; only the own-provider event fires and `event.provider == Some(ext_a)`. - Existing `event_triggered_own_capabilities_scope_resolves_hook_failed_owner_from_hook_id` remains green: passes `None` for the new arg so the fallback path is still exercised for legacy payloads. All other call sites updated to pass `None` (no owner available) or the resolved owner where applicable. * fix(hooks): validate event subscription scope against run scope (serrrfirat HIGH #1 on PR nearai#3640) `EventTriggeredHookSubscription` accepted a caller-supplied `EventStreamKey` + `ReadScope` and used `run_context.scope.tenant_id` as the hook context's tenant — with no validation that the two agreed. A caller wiring tenant A's host with tenant B's stream would cause hooks to observe B's events while the hook context claimed tenant A. Cross-tenant trust-boundary break. Add `EventTriggeredHookSubscription::validate_against_run_scope` and call it from `build_text_only_host_with_capabilities` before spawning. Validation: - Stream `(tenant_id, user_id, agent_id)` must equal `(run_context.scope.tenant_id, thread_scope.owner_user_id, run_context.scope.agent_id)`. - Thread without `owner_user_id` cannot bind any subscription — the user dimension is required to verify stream identity. - Every `Some(want)` in `ReadScope` must equal the corresponding run/thread scope value (project/mission/thread). `None` is permissive (run scope owns the dimension authoritatively). Failures surface as `RebornLoopDriverHostError::ScopeMismatch` with a specific reason naming the offending dimension. Tests: - `event_triggered_subscription_with_foreign_tenant_stream_fails_host_build` - `event_triggered_subscription_with_foreign_user_stream_fails_host_build` The integration fixture's `ThreadScope` now sets `owner_user_id: Some(...)` so it passes validation; previously it was `None`, which the new check (correctly) refuses. Existing tests continue to pass. * fix(hooks): surface event subscription replay-gap as a milestone (serrrfirat MED on PR nearai#3640) When the durable event log returned `EventError::ReplayGap`, the event-triggered subscription's background task previously logged a `tracing::warn!` and broke out of the poll loop — silently killing all future hook event delivery for the run with no operator-visible signal. A scoped audit hook that mattered to compliance would just stop, and nobody downstream would know. Surface the termination through the host's milestone sink: - New `LoopDriverNoteKind::EventSubscriptionTerminated` variant. - The subscription's `spawn`/`run` now takes the host's `Arc<dyn LoopHostMilestoneSink>` and the active `LoopRunContext`. On `ReplayGap`, it constructs a `DriverNote` milestone with that kind plus a `LoopSafeSummary` describing the gap, publishes it through the same sink that carries every other host milestone, and *then* breaks (fail-closed: the at-most-once contract is already broken; resuming from `earliest` would silently lose the gap). - Log level bumped from `warn` to `error` to match the severity. - A best-effort send: failures to publish the milestone are logged but do not stall the subscription teardown. Tests: - `event_triggered_replay_gap_emits_subscription_terminated_milestone`: appends 3 events, `truncate_before_or_at` to cursor 2 to force a replay gap, starts the subscription from cursor origin (now stale), and asserts a `DriverNote { kind: EventSubscriptionTerminated, .. }` shows up on the host's milestone sink within a 2s deadline. Self-emit reentrancy (serrrfirat MED #3 on the same PR) is intentionally not addressed here — that fix needs a design call (task-local re-entry flag vs. removing RuntimeEvent emit capability from event-hook execution contexts) and is a follow-up. * fix(hooks): suppress event-triggered self-observation (serrrfirat MED #3 on PR nearai#3640) A hook that subscribes to one of the hook-lifecycle event kinds (`HookDispatched`/`HookDecisionEmitted`/`HookFailed`) with a scope that matches its own provider would otherwise be dispatched for events describing its OWN executions. The dispatcher emits those events itself when running the hook, so a hook subscribing to `HookFailed` with `OwnCapabilities` against its own extension would fail → emit HookFailed → re-dispatch → fail → emit → … storm. `dispatch_event_triggered_at` now skips events whose `event.hook_id` equals the binding's own hook id when the event kind is a hook- lifecycle kind (`is_hook_lifecycle_kind`). The check is intentionally narrow: - It only fires for hook-lifecycle events. Subscriptions to other event kinds are unaffected. - It only suppresses literal self-observation; events about other hooks (even hooks from the same extension) still dispatch. This does NOT cover the broader case of a hook that captures an `Arc<DurableEventLog>` and mints arbitrary `RuntimeEvent`s from inside its `observe()`. That requires architectural restriction on what hook impls can capture — tracked separately as a follow-up. Tests: - `event_triggered_self_lifecycle_event_does_not_redispatch`: appends two `HookFailed` events with the same provider — one targeting the subscriber's own hook id, one targeting a different hook. Asserts only the OTHER hook's failure fires (proves the filter is narrow, not blanket). * fix(hooks): address henrypark133 should-fix #1, #2, #3, #6 on PR nearai#3640 Four items from the 5-15 review (#4 DoS budget and #5 narrowed projection deferred — see below): **#1 (should-fix) Invariant: EventTriggered ↔ event_kind_filter** `HookRegistry::insert` now enforces the biconditional at install time: an `EventTriggered` binding must declare an `event_kind_filter` (otherwise the dispatcher's kind match would silently never fire — a no-op binding), and conversely only `EventTriggered` bindings may declare a filter (other points are kind-agnostic and would ignore the field). Misconfigured bindings fail loud at install. **#2 (should-fix) Remove `Clone` derive on EventTriggeredHookSubscription** `Clone` on a spawn-semantics type was a footgun: external callers cloning + spawning twice would create two consumers reading from the same `start_cursor`, each dispatching every hook. Replace with an explicit `clone_for_independent_spawn(&self)` method named verbosely so the property is visible at the seam. Internal use updated in the factory's host-build path; external callers can no longer accidentally construct a dual-consumer pattern. **#3 (should-fix) catch_unwind around the background `run()` task** The subscription's tokio task body now runs inside `AssertUnwindSafe(...).catch_unwind()`; a panic in `run()` emits the same `EventSubscriptionTerminated` `DriverNote` milestone the `ReplayGap` path already emits, instead of silently terminating with no operator-visible signal. **#6 (should-fix) Replay semantics in rustdoc on public API** Added a "Replay semantics" section to `EventTriggeredHookSubscription` rustdoc: at-least-once, caller-owned cursor persistence, the restart-from-start_cursor replay pattern. Previously only in the design doc; now load-bearing API contract is visible at the type. **#4 (deferred) Per-hook DoS budget for Installed tier** Henry's recommendation was to gate `Installed`-tier event-triggered hooks entirely until the budget design lands, allowing only Builtin/Trusted. That breaks 11 existing tests + the primary use case. Instead: documented the existing first-line throttle (`batch_limit` × `poll_interval`) as the current bound on indirect- recursion fanout, and tracked the full per-hook rate cap with poisoning + milestone-on-overrun as a follow-up. The self-trigger guard (committed earlier in this PR) catches the most common direct pattern; the throttle here bounds the indirect pattern until the proper budget lands. **#5 (deferred) Narrowed `HookObservableEvent` projection** Would prevent full `RuntimeEvent` surface from reaching Installed- tier hooks. Project-wide impact (events crate types, projection glue). Tracked as a follow-up; the existing sanitized-event projection bounds the surface to closed-vocab labels. All 156 hooks lib + 30 reborn integration tests pass. * chore(hooks): address nits from PR nearai#3640 review Bundle three nit-tier review items into a single commit: **nearai#9 Replace author-internal tags with NOTE(nearai#3640)** The Phase-5 PR (nearai#3640) had several `serrrfirat HIGH/MED #N on PR nearai#3640` comment tags in this PR's diff. These are review-internal scaffolding, not load-bearing for future readers. Replaced with `NOTE(nearai#3640)` in: - crates/ironclaw_hooks/src/dispatch.rs (self-observation guard) - crates/ironclaw_reborn/src/loop_driver_host.rs (scope validation, replay-gap milestone, subscription binding) - crates/ironclaw_reborn/tests/hooks_integration.rs (three regression tests covering scope validation, self-observation suppression, and replay-gap surfacing) - crates/ironclaw_turns/src/run_profile/host.rs (`EventSubscriptionTerminated` doc) **nearai#10 Replace 10ms spin-poll with tokio::sync::Notify** `wait_for_seen_events` polled the shared `Mutex<Vec<SeenRuntimeEvent>>` every 10 ms until the expected count was reached. Replaced with a `SeenLog` newtype that pairs the events vec with a `Notify`; the hook's `observe()` calls `seen.push(...)` which signals `Notify::notify_one`, and `wait_for_seen_events` parks on `notified().await` under a `tokio::time::timeout`. `notify_one` is a permit-store, so an event landing between snapshot and wait still wakes the waiter immediately. Test latency drops from ~10 ms median to sub-ms and is no longer rate-limited by the polling cadence. All 30 hooks_integration tests still pass. **nearai#11 Remove unused Clone derive on EventTriggeredHookContext** No call site clones the context — it's passed by reference. Dropped the derive to make the borrow contract clearer. * docs(hooks): reconcile event-triggered hooks design doc with Phase 5 reality Address gemini-code-assist review on `04-event-triggered-hooks.md`: - L50 (Likely surface): annotated the sketch's full `RuntimeEvent` use with a pointer to the narrowed-projection follow-up so the snippet no longer reads as a recommendation contradicting L119–121. - L55 (sink methods): replaced `note_fact` / `emit_audit` (which never shipped on `ObserverSink`) with the actual `note(category, summary)` primitive and cross-referenced Reborn's `EventTriggeredObserverSink`. - L95 (cursor / replay): "lost events during downtime acceptable" contradicted the at-least-once replay semantics described in the Phase 5 implementation notes. Rewrote the bullet to say replay is at-least-once from the persisted cursor and to spell out the operator obligation around cursor persistence before shutdown. - L100/115 (forbids events dep): the original doc claimed `ironclaw_hooks` forbids an `ironclaw_events` dep, but the Risk section noted the dep is already established via PR nearai#3573. Updated both passages to reflect that the dep direction is set; Phase 5 adds the *consumer* side. The narrowed `HookObservableEvent` projection is now framed as a follow-up tracked in nearai#3690. * refactor(hooks): unify event-triggered sink with ObserverSink Address PR nearai#3640 review findings A3, C4, F14, and cluster G: - F14: drop duplicate `EventTriggeredObserverSink` trait and reuse `ObserverSink` directly in the `EventTriggeredHook` trait. The two surfaces were signature-identical; keeping them separate let them drift, and a future gate/mutator method added to one would not surface as a compile error on the other. - A3: add `is_replay: bool` to `EventTriggeredHookContext` and a dedicated `dispatch_event_triggered_replay_at` entry point. The subscription contract is at-least-once, so side-effecting hooks need to dedupe by `event.event_id` on restart-driven replay. - C4: index event-triggered bindings by `RuntimeEventKind` at install time so dispatch is O(matches) instead of scanning every event-triggered binding for every event. - Cluster G: doc/04-event-triggered-hooks.md updated to reflect the unified sink, the explicit at-least-once semantics + `is_replay` signal, the actual `note(category, summary)` primitive (not the speculative `note_fact` / `emit_audit`), the corrected `ironclaw_events` dep status, and the issue nearai#3690 reference for the narrowed `HookObservableEvent` projection. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * perf(hooks): adaptive backoff for event-triggered subscription Address PR nearai#3640 review findings C5, A1, A2: - C5: empty-poll backoff for `EventTriggeredHookSubscription`. The previous loop hammered the durable log at a fixed 50ms cadence under sustained idle, even when no events had arrived for minutes. The subscription now tracks consecutive empty polls and sleeps for `min(poll_interval << streak, max_poll_interval)` before the next poll, defaulting to a 1s cap; a non-empty batch resets the streak so producer bursts restore low-latency dispatch immediately. Exposed via `with_max_poll_interval` so callers can tune. - A1 / A2: explicit issue references for the deferred narrowed `HookObservableEvent` projection (nearai#3690) and the per-hook DoS dispatch budget (nearai#3689). The current self-trigger guard catches direct-recursion storms; the backoff bounds indirect ones until the proper budget design lands. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(hooks): cover event-triggered dispatch edge cases Address PR nearai#3640 review findings D8, D9, D10, D11, D12, and B: - D8: dispatching an event-triggered binding that has no installed hook impl must poison the slot and surface a Malformed failure rather than silently no-op. A follow-up dispatch on the same kind must skip the poisoned slot. - D9: registry validation rejects non-event-point bindings that carry an `event_kind_filter`, mirroring the existing reverse-direction check. - D10: the existing hook-meta serde round-trip tests always passed `None` for `owning_extension` and never asserted `event.provider`. Add `hook_meta_events_round_trip_owning_extension_as_provider` to pin the projection that scope filtering depends on. - D11: `scope_provider_for_runtime_event` falls back to `None` when the registry mutex is poisoned. Force a poison on a spawned thread and assert the resolver remains fail-closed. - D12: `run_event_triggered_hook` catches panics from the hook impl via `AssertUnwindSafe::catch_unwind`. Drive it with a deliberately panicking impl and assert `FailureCategory::Panic`. - Cluster B: when a hook-meta event has `provider: None`, the dispatcher recovers the owning extension from the registry's hex-keyed index so `OwnCapabilities` watchers still fire. Add a full end-to-end test exercising that path through `dispatch_event_triggered_at`. Also pin C4 indexing: a registry-level test that `active_for_event_kind` returns only bindings whose declared filter matches. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): port event-triggered tests after foundation rebases (nearai#3911/nearai#3912/nearai#3913) - Add event_kind_filter: None to HookBinding test constructions (foundation added new field) - Replace tuple-struct construction of ExtensionId/HookLocalId with ::new() per nearai#3912 newtype privatization - Lowercase RuntimeEventKind debug repr for HookLocalId validation (lowercase-only identifiers) - Replace pub-use re-exports with module-path imports per foundation cleanup - Rename fixture.user_id to fixture.actor_id per nearai#3633 final naming * fix(hooks): adapt event-triggered to WASM hook runtime (nearai#3920) - Add event_kind_filter: None to WASM HookBinding constructions in dispatch.rs - Extend HookManifestKind match arms in registrar.rs and wasm/runtime.rs to handle EventTriggered (rejected: WASM-bodied event-triggered hooks are not yet supported) --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
* Add saved output refs for Reborn shell * Tighten Reborn shell output capture * Harden Reborn shell output capture lifecycle * Sanitize Reborn shell previews before saving * fix(reborn): tenant-scope shell saved-output dir + GC (nearai#4154 review blockers #1, #4) Saved shell-command output files previously landed directly in shared std::env::temp_dir() (per-file 0o600, but the parent dir was world-listable), and cleanup_stale_command_outputs() walked all of /tmp and unlinked any entry that matched the well-known prefixes — both ambient surfaces let one principal on the same host enumerate or delete another principal's saved output. Route every saved output through a per-scope subdirectory derived from RebornSandboxScopeKey (the same SHA-256-of-tenant/user/agent/project digest the Reborn sandbox transport uses for workspace_path) under <tempdir>/ironclaw-command-outputs/<scope_digest>/, created with owner-only 0o700. Both scratch streams and final sanitized outputs live inside that directory, and the 24h GC scan is scoped to it — so two distinct (tenant, user, agent, project) tuples produce disjoint, non-enumerable directories and the cross-principal-delete surface closes by construction. Closes blockers #1 and #4 from the PR-nearai#4154 review. Blocker #2 (typed saved_output_read capability) and finding #3 (24h GC vs never-delete retention) remain serrrfirat-owned design decisions and are intentionally out of scope here. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix: address saved shell output review findings * fix: publish shell saved output through file_read --------- Co-authored-by: Zaki <zaki@manian.org> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
* fix(reborn): remove preview line cap * feat(reborn): add typed diff display previews * fix(reborn): address typed diff preview review * refactor(reborn): centralize display preview validation * fix(reborn): address henrypark133 review — perf guard, path subtitles, typed kind, tests (nearai#4184) * refactor(reborn): extract truncate_to_byte_boundary, has_unsafe_path_chars, will_use_large_diff_path - Move truncation utility to ironclaw_host_api (canonical home next to the max-bytes constant); both diff_preview and the projection layer now delegate to it — deletes duplicate BoundedText/truncate_utf8. - Extract has_unsafe_path_chars predicate in display_preview.rs so safe_display_path and safe_preview_subtitle share the rejection logic instead of copy-pasting the same five-condition check. - Replace pub(super) DIFF_PREVIEW_DETAILED_INPUT_MAX_BYTES coupling with will_use_large_diff_path predicate so file.rs no longer reaches into diff_preview internals to compute the threshold. Addresses thermo-nuclear review findings #1, #2, #3. * refactor(reborn): simplify diff preview contract * fix: address diff preview review cleanup * fix: resolve CI clippy failures
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…#3899) * Reborn budgets: address all nearai#3841 follow-ups end-to-end Implements every open follow-up from PR nearai#3841 (cost-based budgets foundation), driven by the plan in `docs/plans/2026-05-22-reborn-budgets-followups.md`: - **C2 (provider tokens)**: `LoopModelResponse.usage` carries real `(input_tokens, output_tokens)` from `CompletionResponse` / `ToolCompletionResponse`; `usage_for_response` reconciles to actual USD via the cost table instead of the conservative estimate. - **D1 (cascade warnings)**: `CascadeOutcome` variants carry `Vec<BudgetWarning>` so warnings preceding a pause or hard deny reach the audit sink. `ResourceError::LimitExceeded` / `RequiresApproval` reshaped to struct variants. - **C1 (cancellation safety)**: new `LoopModelBudgetAccountant::release_in_flight` trait hook + RAII `ReservationReleaseGuard` in `HostManagedLoopModelPort::stream_model` so a cancelled future doesn't orphan its reservation. - **E1 (dead code)**: removed the never-set `budget_accountant` field on `ThreadBackedLoopModelPort`. - **Real cost table**: new `StaticModelCostTable` + `LlmModelProfilePolicy::build_cost_table()` populated from `ironclaw_llm::costs::model_cost` with `default_cost` fallback so unknown providers never silently reconcile to zero. - **B1 (filesystem gate store)**: new `FilesystemBudgetGateStore` mirroring `FilesystemResourceGovernorStore`; pending gates survive process restart. - **A1 (production wiring)**: composition builds `GovernorBackedAccountant` from the cost table + governor and threads it through `RebornLoopDriverHostFactory::with_model_budget_accountant`. - **A2 (audit / SSE projection)**: `InMemoryResourceGovernor::with_event_sink` emits `Reserved`, `Reconciled`, `Released`, `Warned`, `ApprovalRequested`, `Denied`, `LimitChanged`; composition holds an `InMemoryBudgetEventSink` ready for downstream SSE projection. - **F1 (stuck-loop normalization)**: `CapabilityCallSignature::from_call` now runs `progress::normalize_for_hash` so the existing repetition window collapses request-id / UUID / timestamp noise. Side fix: `ResourceValue` moved to adjacent serde tagging (the combination of internal tagging + `Decimal`'s `serde-with-str` representation breaks JSON serialization — rust-lang/serde#1402). Regression tests added per item — see the acceptance evidence appendix in the plan doc. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Reborn budgets: end-to-end test coverage via test-support feature Adds 13 e2e tests covering the budget pipeline through `build_reborn_runtime` + `send_user_message`. Required infrastructure: - **`test-support` feature** on `ironclaw_reborn_composition` exposing `BudgetTestGateway` (scripted token usage) and `RebornRuntimeInputTestExt`. Existing `model_gateway_override` field promoted from `#[cfg(test)]` to `#[cfg(any(test, feature = ...))]` with a new public `with_model_gateway_override_for_tests` setter. - **Cost-table override** on `RebornRuntimeInput` so tests can pair the gateway with a deterministic `ModelCostTable`. Without this, an override gateway dropped the cost table and the accountant never fired. - **Budget accessors** on `RebornRuntime`: `budget_resource_governor`, `budget_event_sink`, `budget_gate_store`, and `apply_resolved_budget_gate`. Test-feature gated. - **`ResourceGovernor::usage_for`** added as a default-impl trait method so tests read spend through the trait surface. - **`BudgetGateStore` wired into the accountant**: `GovernorBackedAccountant::with_gate_store(...)` opens a pending gate whenever the governor cascade returns `RequiresApproval`. The approval-required host error is unchanged; the gate is the out-of-band channel a user-facing handler resolves. Scenarios covered: | # | Test | What it asserts | |---|---|---| | F1 | `f1_happy_path_records_actual_usd_in_ledger` | Ledger depletes by provider tokens × cost table | | F2 | `f2_crossing_warn_threshold_emits_warned_event` | Warn fires alongside successful Reserved/Reconciled | | F3 | `f3_approval_with_increased_limit_unblocks_retry` | Approve → set_limit applies → retry succeeds | | F4 | `f4_cancel_keeps_budget_blocked_on_retry` | Cancel → retry still short-circuits | | F5 | `f5_expiry_marks_gate_terminal_and_keeps_budget_blocked` | Expiry → gate drops from pending list, retry still blocked | | F6 | `f6_hard_cap_denied_before_provider_call` | Estimate over cap → zero model calls, Denied event | | C1 | `c1_provider_tokens_reconcile_to_actual_usd` | Real numbers, not estimate | | C2 | `c2_unknown_model_in_cost_table_reconciles_to_zero` | Unknown profile → zero spend | | C3 | `c3_zero_cost_model_records_zero_spend` | Free model → zero USD with non-zero tokens | | D1 | `d1_agent_deny_preserves_user_warn_event` | Cascade emits both Warned and Denied | | D3 | `d3_fresh_user_without_limits_runs_without_denial` | No limit → no denial | | + | `pause_in_distinct_runs_produces_distinct_pending_gates` | Per-run gate identity | | + | `budget_test_gateway_scripted_replies_drive_per_turn_costs` | Multi-turn scripted accumulation | F7 (cancellation mid-stream) is unit-covered by `release_in_flight_drains_orphan_reservation_on_cancellation`. D2 (period rollover) is unit-covered by `rolling_24h_snapshot_reports_anchored_window_not_now_window`. B-series (background ticks) await the BackgroundKind scheduler call site (no production caller in Reborn yet). Run via `cargo test -p ironclaw_reborn_composition --features test-support`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Budget review feedback: address all 7 findings from PR nearai#3899 review Two High and five Medium issues raised by serrrfirat's multi-agent review. **High #1 — `FilesystemBudgetGateStore` cross-tenant leakage** The store hardcoded `ResourceScope::system()` for every op, so all tenants wrote into the same `/tenants/__SYSTEM__/...` snapshot and `list_pending` would expose gates across tenants. Fix: `new(...)` now takes a `ResourceScope`; each tenant gets its own store, and the `ScopedFilesystem` mount view routes the snapshot under that tenant's path. Added `list_pending_does_not_leak_across_tenants` regression. **High #2 — accountant wired without default budget limits** Composition built `GovernorBackedAccountant` without `with_seeding_policy`, so the local-dev governor started empty and `reserve_with_outcome_in_state` skipped accounts that had no configured limit — model calls reconciled spend but never enforced a cap. Fix: `build_reborn_runtime` now loads `BudgetDefaults::compiled_defaults().with_env()` and wires `BudgetSeedingPolicy` + `with_overestimate_factor`. Renamed the D3 test to `d3_seeding_policy_installs_default_cap_on_first_touch` to prove the wiring fires. **Medium #3 — RAII guard disarmed before post_model_call await** `HostManagedLoopModelPort::stream_model` was disarming the `ReservationReleaseGuard` before awaiting `post_model_call`. A cancellation during that await dropped the future without cleanup, orphaning the reservation. Fix: disarm AFTER `post_model_call` returns. `release_in_flight` is now idempotent (peek-then-release- then-remove) so a successful post-call + subsequent guard drop is a no-op. **Medium #4 — failed release drops the retry handle** `release_in_flight` removed the in-flight entry before calling `governor.release`. A transient storage error left the reservation active in the governor with the id discarded. Fix: peek first, release, only remove on success. Errors keep the entry retained for a future retry / cleanup hook. **Medium #5 — unknown model silently reconciles to zero USD** Both `estimate_for` and `usage_for_response` fell back to `ModelCost { 0, 0, 0 }` when the cost table had no entry for the effective model. Cost-table drift would silently bypass daily caps. Fix: `GovernorBackedAccountant` carries a `default_cost` (default ~ GPT-4o pricing, ~`$0.0000025 input + $0.00001 output per token`) used for unknown models. Callers wiring `ZeroCostTable` for free / Ollama explicitly opt out of the fallback. Updated the C2 e2e test to assert the new fail-closed shape. **Medium #6 — paused dimension lost when another hard-denies** `check_thresholds_all_interventions` stored `Approval` only in the `approval` slot, so when one dimension paused and another hard-denied, the `Deny { warnings, denial }` outcome lost the pause signal. Fix: also push a warning-shaped record for the paused dimension. **Medium nearai#7 — unbounded terminal-gate retention** The snapshot kept every gate forever; `open` / `resolve` / `get` / `list_pending` were O(total historical gates). Fix: `with_terminal_retention` (default 30 days). Every mutation prunes terminal gates whose resolution timestamp is older than the window. Added `terminal_gates_older_than_retention_are_pruned_on_next_write` regression. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * ci: replace lock-poisoned expects with PoisonError::into_inner scripts/check_no_panics.py flagged five .expect("...lock poisoned") calls in the new test_support.rs. Use the same idiomatic recovery pattern the rest of the codebase uses (see InMemoryBudgetGateStore, InMemoryBudgetEventSink): on a poisoned lock, recover the inner data via PoisonError::into_inner rather than panicking. The test gateway's state is append-only logs / replies queues, so reading them through a poisoned lock is safe. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Finish A1 / A2 / F1 from plan + honest plan doc update The plan claimed "all nine items landed" but A1 (production wiring), A2 (SSE projection), and F1 (full progress strategy) were partials. This commit finishes the work so the plan matches reality. **A1 — production-shape accountant builder** New `ironclaw_reborn_composition::build_default_budget_accountant` public helper that wires the seeding policy + overestimate factor + gate store from `BudgetDefaults::compiled_defaults().with_env()` and returns an `Arc<dyn LoopModelBudgetAccountant>`. Production loop composers call this with their `PersistentResourceGovernor` + `FilesystemBudgetGateStore` + LLM-policy-derived cost table; the local-dev runtime in `build_reborn_runtime` now uses the same helper instead of duplicating the seeding logic inline. Unit-tier regression `seeds_compiled_default_user_cap_on_first_touch` proves the helper installs the compiled-default $5 user cap on first model call. **A2 — broadcast sink + AppEvent projection** - `ironclaw_resources::BroadcastBudgetEventSink` wraps `tokio::sync::broadcast::Sender<BudgetEvent>` with `subscribe()` / `subscriber_count()`. `CompositeBudgetEventSink` fans events to multiple sinks. - Composition fans every `BudgetEvent` to the in-memory sink (for tests) AND the broadcast sink (for SSE projection) via `CompositeBudgetEventSink`. - New `AppEvent::BudgetWarn` / `BudgetPause` / `BudgetDenied` / `BudgetLimitChanged` wire-stable variants in `ironclaw_common::event`. - `src/bridge/budget_events.rs` carries the projection: a tokio task spawned by `spawn_budget_event_projection` drains the broadcast receiver and emits the appropriate `AppEvent` via `SseManager::broadcast_for_user`. System-scoped events (no user identity) are skipped. This is the only producer of these `AppEvent` variants per `.claude/rules/gateway-events.md`. - `RebornRuntime::broadcast_budget_event_sink()` exposes the sink to the binary so the startup path subscribes. E2E test `broadcast_sink_publishes_events_to_subscribers` drives a real `send_user_message` and asserts Reserved + Reconciled lands on the broadcast. **F1 — diminishing-returns stop condition** The earlier shipped `ParamHash` normalization in `CapabilityCallSignature` strengthened the existing `recent_call_signatures`-based repetition detector. This commit adds the second half of F1: a rolling output-token window that detects "wedged" loops the repetition detector misses (model keeps responding but produces no useful output). - `LoopExecutionState.recent_output_token_counts: BoundedRing<u32, 8>` populated by the executor from `LoopModelResponse::usage`. - `BoundedRing::iter` returns `impl DoubleEndedIterator` so the strategy can scan the trailing window. - `DefaultStopConditionStrategy` gets `min_delta_tokens` (default 4) + `noprogress_window` (default 4). When the last N turns all produce ≤ min_delta_tokens of output, fire `StopKind::NoProgressDetected`. - Regression tests: `four_consecutive_low_token_turns_trigger_no_progress` proves the detector fires; `occasional_low_token_turn_does_not_trip_no_progress` proves a productive turn resets the trailing count. **Plan doc** Updated the status header from "all nine items landed" to the honest per-item shape. Acceptance evidence table expanded with the new test names. New "Review-feedback fixes layered on top" subsection documenting all 2 High + 5 Medium findings addressed during review. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Address thermo-nuclear review: collapse filesystem-store duplication, flatten cfg permutations, split budget accountant Five structural simplifications surfaced by the deep audit on PR nearai#3899, plus two bug fixes from the earlier review pass: - ironclaw_resources: extract `cas_snapshot` shared infrastructure (`StorageError` + `Snapshot` traits + `CasSnapshotStore<F>` + async-runtime worker + per-path lock map) and merge `filesystem_gate_store.rs` into `filesystem_store.rs`. Deletes ~350 lines of duplicated read-modify-write + worker-thread + CAS machinery; both stores are now thin shims over the shared helper. - ironclaw_reborn_composition: flatten the 4-way cfg permutation in `build_reborn_runtime` model-gateway resolution into three flat steps (normalize override → build production gateway via cfg-gated helper → test override wins). Also drops the `unused_mut` warning. - ironclaw_reborn_composition: collapse the 3-layer test-only setter dance for `model_gateway_override` / `model_cost_table_override` into a single setter pair gated on `cfg(any(test, feature = "test-support"))`. Deletes the `RebornRuntimeInputTestExt` extension trait — integration tests now call the inherent methods directly. - ironclaw_loop_support: split the 1305-line `budget_accountant.rs` into `budget_cost_table.rs` (ModelCost/ModelCostTable trait/ZeroCostTable/ StaticModelCostTable), `budget_seeding.rs` (BudgetSeedingPolicy), and `budget_accountant.rs` (just GovernorBackedAccountant). Each module now owns one concern. - ironclaw_resources: add `impl Display for ResourceAccount` and route the hierarchical account-label rendering through it; delete the 60-line bespoke `account_label` helper from `src/bridge/budget_events.rs`. - ironclaw_common + bridge: collapse the four `AppEvent::Budget*` variants into a single typed `AppEvent::Budget(AppBudgetEvent)` with the four shapes carried inside the enum. Wire-shape stays identical (snake_case serde tag). - ironclaw_resources + ironclaw_loop_support: thread real gate id through `BudgetEvent::GateOpened { gate_id, needed, at }` (new variant) and have the accountant emit it via the broadcast event sink after store.open succeeds. The bridge now projects `BudgetEvent::GateOpened` (not `ApprovalRequested`) into `AppEvent::Budget(Pause { gate_id, ... })` so SSE consumers receive the persisted gate id rather than a fabricated zero uuid. - ironclaw_agent_loop: in the F1 token-counting path, push to `recent_output_token_counts` only when the model response carries `Some(usage)` and only on the `AssistantReply` arm (instead of `unwrap_or(0)`). Diminishing-returns detection now reflects real spend. Net delta: -461 lines (+999 / -1460). Workspace `cargo clippy` clean, `cargo test` clean on ironclaw_resources / ironclaw_loop_support / ironclaw_reborn_composition; budget_e2e + budget_approval_e2e both green. Pre-existing CI failures (`cli::tests::test_version` stack overflow, `facade_factory::production_*` RuntimeProcessPort missing) are unrelated and reproduce on the pristine branch tip. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(cli): refresh insta snapshots after runtime-policy flag additions The `import`-feature variants of the help snapshots were left stale when `--deployment-mode`, `--runtime-profile`, `--yolo-disclosure` were added in cc04481 (nearai#3243); the `_without_import` variants were updated but these were not. CI was failing the snapshot assertion under the slim PR matrix (`--features postgres,libsql,html-to-markdown,bedrock,import`). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Address PR nearai#3899 thermo-nuclear review (TN #1, #2, #3) TN #1 — budget defaults resolved in wrong layer: - `build_default_budget_accountant` no longer reads process env; it now takes `&BudgetDefaults` as a parameter and the caller owns the config-layer precedence (compiled → section → env) plus the `validate()` call. - `RebornRuntimeInput` gains an optional `budget_defaults` field + `with_budget_defaults()` builder so the composition root passes a pre-resolved value. `build_reborn_runtime` falls back to `compiled_defaults().with_env() + validate()` when none is supplied so existing call sites keep working. TN #2 — gate-store scoping at wrong boundary: - `BudgetGateStore` trait methods (`open`, `resolve`, `expire_pending_older_than`, `get`, `list_pending`) now take `&ResourceScope` as first arg. `GovernorBackedAccountant` passes the caller's scope from `resource_scope(context)`. - `CasSnapshotStore` gains `update_with_scope` so the same store instance can route per-operation. `FilesystemBudgetGateStore` no longer takes scope at construction — one shared instance serves every tenant via the `ScopedFilesystem` mount view. - `InMemoryBudgetGateStore` ignores scope (suitable for single-tenant tests / local-dev); production multi-tenant filesystem path is correctly partitioned by `ResourceScope`. - `RebornRuntime::apply_resolved_budget_gate` now takes scope too. TN #3 — half-wired projection bridge: - Removed `src/bridge/budget_events.rs`, its `spawn_budget_event_projection` helper, the `AppEvent::Budget` variant, and the `AppBudgetEvent` type. No production caller ever subscribed the broadcast sink onto SSE and no frontend consumed the variant, so the half-wired bridge is gone pending a real owner that spawns a projection task with shutdown cancellation. - The runtime's `broadcast_budget_event_sink()` accessor stays so a future production composer can still subscribe without rebuilding the runtime. Bonus — to keep budget e2e tests working under the new libsql local- dev path that origin/reborn-integration introduced, added `PersistentResourceGovernor::with_event_sink` (parity with the `InMemoryResourceGovernor` accessor). The libsql variant of `build_local_dev_store_graph` now wires the composite sink to the persistent governor so governor-emitted `Warned`/`Denied`/`Reserved`/ `Reconciled` events reach subscribers on both feature paths. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Wire budget-event projection task into RebornRuntime Re-implements PR nearai#3899 follow-up A2 / Thermo-Nuclear #3 with a real production owner instead of leaving the broadcast sink half-wired: - `crates/ironclaw_reborn_composition/src/budget_events.rs` (new): `BudgetEventObserver` trait + `TracingBudgetEventObserver` default observer + crate-internal `BudgetEventProjection` task that drains the runtime's broadcast `Receiver<BudgetEvent>` and forwards every event to the observer. Cancellation via `CancellationToken`; lagged subscribers logged and resumed; receiver-closed exits cleanly. - `RebornRuntimeInput::with_budget_event_observer(...)` lets production owners install a custom observer (SSE projection, WS fan-out, telemetry export). When unset, the runtime installs the tracing observer so events always surface in structured logs. - `build_reborn_runtime` always spawns the projection task at runtime construction; `RebornRuntime::shutdown` cancels it and awaits the handle so background state drains before the runtime drops. - E2E test `projection_delivers_budget_events_to_installed_observer` drives `build_reborn_runtime` with a capturing observer and asserts the observer sees `Reserved` + `Reconciled` from a real model call, testing through the caller per `.claude/rules/testing.md`. - Existing `broadcast_sink_publishes_events_to_subscribers` updated to expect the runtime's own projection task as a baseline subscriber. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(reborn): rustfmt the merged loop_support import block The conflict resolution for the post-merge import list was not run through rustfmt; CI Formatting flagged the wrapping. No logic change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…ED flag (nearai#3934) (nearai#3938) * feat(hooks): extension-declared hook section on ExtensionManifestV2 (nearai#3934) Add a `[[hooks]]` declaration surface to the production v2 extension manifest (`ironclaw_extensions::ExtensionManifestV2` and its projected `ExtensionManifest`). Each entry is carried as a structurally-typed `HookSectionEntryV2` DTO — a `local_id` plus the entry's body re-serialized to canonical TOML — so `ironclaw_extensions` (substrate) never imports the `ironclaw_hooks` predicate vocabulary. The composition layer, which depends on both crates, is the single seam that projects these payloads into typed `ironclaw_hooks::HookManifestEntry` values (a later commit). Parse-time structural bounds: `MAX_MANIFEST_HOOKS` (32, matching the downstream per-extension registration cap) and `MAX_HOOK_ENTRY_BYTES` (8 KiB per entry). Entries must be tables carrying a non-empty `id`; ids must be unique within the manifest. `#[serde(default)]` keeps every existing manifest valid (empty `hooks` vec). The DTO holds canonical TOML as a `String` rather than a `toml::Value` so the enclosing `ExtensionManifestV2` keeps its `Eq` derive (`toml::Value` is not `Eq`). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(hooks): composition-layer activation module (loader, first-party hook, flag) (nearai#3934) Add `ironclaw_reborn_composition::hooks` — the single seam that activates the hook framework in production. Implements four numbered pieces of nearai#3934: - Feature flag (item 7): `HooksActivationConfig`, default OFF, resolved from `HOOKS_ENABLED` (only `1`/`true`/`yes`/`on` enable; unset/anything else = OFF). Flag OFF ⇒ `build_hook_dispatcher_builder_factory` returns `None` and the runtime composes no dispatcher — exact pre-hooks behavior. Hard rollout-safety contract. - Manifest → registry loader (item 2): `install_extension_hooks` projects each `ExtensionManifestV2` `HookSectionEntryV2` (canonical TOML) into a typed `HookManifestEntry` and installs it via `HookRegistrar::install` at the `Installed` trust tier. This is the clean-boundary projection: the hook vocabulary lives only here, never in `ironclaw_extensions`. Trust attenuation is enforced by construction (registrar only calls `install_installed_*`). Fail-closed: any projection/install error fails the build loudly. - First-party builtin hooks (item 3): a single illustrative no-op observer (`NoOpObserverHook`), installed regardless of extensions. Ships dark (zero driver-visible effect even with the flag ON). Production catalog is TBD by design — this PR does not invent a first-party hook. - Dispatcher composition (item 5): builds a per-tenant `PredicateEvaluator` over the in-memory state backend (swappable via the new public `PredicateEvaluator::with_state_backend` for durable nearai#3933), validates the full install set once fail-closed, and returns a per-run builder-factory closure. Per-run construction (fresh registry/dispatcher per host build) + per-tenant evaluator give full isolation; the host factory attaches the run-scoped milestone sink internally. Per-tenant scoping is by construction: `build_reborn_runtime` runs once per identity, so everything here is tenant-local — no global registry. The router-backed gate-ref factory (PauseApproval/PauseAuth) and the security-audit sink (nearai#3922, not yet on this branch) are deferred follow-ups; their absence is fail-closed (PauseApproval surfaces as Denied) and noted for the PR body. Not yet wired into the runtime — next commit. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(hooks): wire dispatcher builder factory into build_default_planned_runtime (nearai#3934) Item 6 of nearai#3934. Add an optional `hook_dispatcher_builder_factory` to `DefaultPlannedRuntimeParts` and, in `build_default_planned_runtime`, call `.with_hook_dispatcher_builder_factory(...)` on the production `RebornLoopDriverHostFactory` when it is present. `None` (the default) means no dispatcher is composed — behavior identical to the pre-hooks runtime (rollout-safety contract). The composition layer (`build_reborn_runtime`) resolves the flag via `HooksActivationConfig::from_env()` and builds the factory against this tenant's extension registry (per-tenant by construction — the function runs once per identity). Fail-closed: a malformed manifest hook fails the build here rather than composing a broken dispatcher. A per-run builder factory (not a captured dispatcher instance) is used so the host attaches a run-scoped milestone sink internally per build — per-run telemetry attribution, the nearai#3573 capture-and-stick lesson. All `DefaultPlannedRuntimeParts` construction sites (8 test sites across ironclaw_reborn, ironclaw_product_workflow, ironclaw_reborn_composition) updated with the new field defaulting to `None`. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(hooks): e2e activation tests through build_default_planned_runtime (nearai#3934) Item 8 of nearai#3934. Add four end-to-end tests in crates/ironclaw_reborn/tests/loop_driver_host.rs that drive the *production* composition function `build_default_planned_runtime` with a per-run hook dispatcher builder factory shaped exactly like the composition layer's output (first-party builtin no-op observer + extension-declared `Installed`-tier hooks projected from a manifest entry through `HookRegistrar::install`), then build a host via the composed `host_factory` and invoke a capability: - flag OFF (no factory): allowed capability completes unaffected and reaches the inner host runtime port — the pre-hooks behavior / rollout-safety contract. - flag ON, first-party-only no-op observer: outcome unchanged, inner port reached — the builtin ships dark. - flag ON, extension-declared deny hook: capability denied through the composed runtime and the inner port is never reached (installed at the Installed tier via the registrar; OwnCapabilities scope keyed to the capability provider). - per-tenant isolation: tenant A's deny hook fires; tenant B (separate build_default_planned_runtime composition, no hooks) completes the same capability — proving no cross-tenant leakage. Security-audit-on-deny assertion is intentionally deferred: nearai#3922's SecurityAuditSink is not yet on reborn-integration. It lands with the audit-sink wiring follow-up. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(hooks): thread hook_dispatcher_builder_factory through shared reborn harness (nearai#3934) The root-crate `tests/support/reborn/harness.rs` constructs `DefaultPlannedRuntimeParts` directly; add the new `hook_dispatcher_builder_factory: None` field so the parity-test harness compiles. Default `None` keeps the harness on the no-hooks path (unchanged behavior). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks): proper error handling / safety annotations for activation production paths (nearai#3938 CI) The per-run dispatcher factory closure used `.expect()` on the first-party and extension hook installs, tripping the no-panics CI gate. These installs are pure replays of the install set already validated fail-closed (`?`) against a scratch builder at composition time, so they are genuine invariants. The factory type returns a non-Result `HookDispatcherBuilder` and is invoked deep in the run loop, so the documented `// safety:` suppression is the correct fix here. Hoisted the expect messages into `let` bindings so the `.expect(msg)` call fits on one line, keeping the scanner-required `// safety:` comment on the same line as the call after rustfmt. The malformed-manifest path (TOML projection) already uses real error propagation via map_err/`?` and is unaffected. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(hooks): direct composition-loader coverage + activation-scope docs Address Codex non-blocking follow-ups on nearai#3938. Add three direct tests for the composition-layer hook loader (`install_extension_hooks` via `build_hook_dispatcher_builder_factory`), driving a real `ExtensionRegistry` with `[[hooks]]` declarations rather than mimicking the loader: - valid `own_capabilities` predicate hook installs at the Installed trust tier; the dispatcher carries the derived binding at BeforeCapability alongside the first-party no-op observer - malformed typed hook body (unknown `mode`) fails CLOSED with `RebornBuildError::InvalidConfig`, never a panic (the load-bearing degradation contract for untrusted external manifests) - a hook claiming `scope = same_tenant` without a verified grant is rejected by trust attenuation (fail-closed) No loader bug surfaced: `HookRegistrar::install` already returns `Result` on every malformed/over-scoped path and the loader maps it to `InvalidConfig` via `?`. Document activation scope at both the loader rustdoc and the `build_reborn_runtime` call site: production currently passes only `builtin_extension_registry()`, so third-party installed-extension hooks are not yet surfaced into the runtime path — only first-party-builtin and builtin-package-declared hooks activate today. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * refactor(hooks): thread HooksActivationConfig through input; empty production catalog Two maintainability cleanups on nearai#3938 (firat review): Item 1 — config hygiene: stop reading HOOKS_ENABLED via std::env::var deep inside build_reborn_runtime. Add a typed `hooks: HooksActivationConfig` field to RebornRuntimeInput (default OFF) plus a `with_hooks_config` builder. The composition root now consumes the typed config; the env var is resolved ONCE at the edge (the reborn CLI's build_runtime_input) via HooksActivationConfig::from_env and threaded down. Testable without env mutation; matches the project's env → typed config → composition pattern. Item 2 — empty production first-party catalog: stop shipping NoOpObserverHook as a first-party builtin. install_first_party_hooks is now a no-op (empty catalog); the production type/install/export for a hook that does nothing is gone (removed from lib.rs exports). The activation machinery is still tested end-to-end through the real composition path via a new `build_hook_dispatcher_builder_factory_with` seam that takes a first-party installer; tests pass a `#[cfg(test)]` NoOpObserverHook through it. Pinned the empty-catalog-is-valid contract: flag ON + empty first-party set + no extension hooks composes a valid zero-binding dispatcher (not a panic/error). Reconciled the sibling loader tests/docs (f6c79c0): the in-module tests now drive the test-only seam; the reborn e2e tests already used a test-local no-op and are untouched. Updated activation-scope docs (loader rustdoc + build_reborn_runtime call site) to reflect the now-single live source (builtin-package-declared hooks). Deferred (not touched): switching to the canonical extension registry for third-party installed-extension hooks (nearai#3934 follow-on). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * refactor(hooks): canonical registry + infallible plan + tenant-scoped counter docs/tests (nearai#3938) Addresses serrrfirat's thermo-nuclear re-review on 1e618d0. #1 (runtime.rs:839, canonical registry): make the extension registry a shared composition artifact. `build_local_dev` builds one `Arc<ExtensionRegistry>`, hands it to `HostRuntimeServices::new` AND stores it in `RebornLocalRuntimeServices.extension_registry`. Hook activation in `build_reborn_runtime` now consumes that same `Arc` instead of rebuilding a builtin-only sidecar, so capability dispatch and hook activation cannot drift. Third-party activation stays a follow-up, but it now follows the canonical registry rather than a separate path. #3 (hooks.rs factory machinery): replace the parse/validate/replay duplication + two prose-justified `.expect()` calls with a typed `HookInstallPlan`. TOML is projected once into typed entries, the full install set is validated once against a fresh builder (fail-closed via `?`), and `HookInstallPlan::rebuild` mints a fresh builder per run. The per-run path is infallible by construction: a plan only exists for an install set that already composed cleanly, so a deterministic replay from the identical fresh-empty start cannot fail. One extension-install code path (`project_extension_install_sets` + `install_extension_sets`) is shared by validation and rebuild. #4 (hooks.rs:295, predicate counter scoping): the evaluator/backend is intentionally tenant-scoped and shared across runs (rate/value caps keyed `(hook, tenant, capability)` with no run_id; a run-scoped limit would reset every run and enforce nothing). Document the split explicitly — per-run-fresh dispatcher, tenant-scoped predicate counters — in the module docs and fix the misleading "per-run isolation of hook state" wording in `ironclaw_reborn` (loop_driver_host.rs / runtime.rs). Add `predicate_counter_state_is_tenant_scoped_across_rebuilds`, which drives a real rate-cap predicate through two dispatchers from one factory and proves the second run sees the first run's recorded count. Rename `factory_mints_independent_dispatchers_per_call` -> `rebuild_mints_independent_dispatchers_per_call` and scope it to proving dispatcher freshness only. #6 (loop_driver_host tests): clarify that the hand-built builder factories cover host PLUMBING, not composition activation. Add `build_reborn_runtime_activates_hooks_through_real_composition_path`, which drives the real `build_reborn_runtime` with `HooksActivationConfig` threaded through `RebornRuntimeInput` (env-free) and the canonical registry, proving the production activation wiring composes. #2 (env boundary) and #5 (empty production catalog) were already fixed in 1e618d0; docs touched here for consistency. Known follow-up (not one of the six items, not introduced here): with the flag ON the standalone local-dev runtime does not yet reach `Completed` for a capability turn even with a zero-binding dispatcher — the composition root wires the dispatcher but not the companion hooked-prompt dependencies. The new runtime test asserts `is_terminal()` + the capability path and documents the gap. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(hooks): cover hook-entry rejection branches; document InvocationCount cap semantics (nearai#3938) Address henrypark133 review (review 4367870023): - Add extension-manifest tests for the three previously-uncovered hook-entry validation branches: non-table `[[hooks]]` element, whitespace-only `id`, and oversized entry (HookEntryTooLarge). - Document the InvocationCount inclusive-allow / deny-on-overflow semantics inline at the comparison site; behavior unchanged and still pinned by the cap test. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(reborn-cli): assert hooks config threaded in caller test (nearai#3938) Addresses the review finding that `build_runtime_input_maps_configured_cli_identity` exercised `build_runtime_input` but never asserted the `hooks` field, so a regression silently dropping `with_hooks_config(HooksActivationConfig::from_env())` or flipping the default-OFF rollout-safety contract would pass. Per `.claude/rules/testing.md` ("Test Through the Caller"): add two assertions to the existing caller-level test: - threading: `runtime_input.hooks == HooksActivationConfig::from_env()`, proving the env-resolved config is actually threaded through and not dropped. Verified via TDD that dropping the wiring fails this assertion (run under HOOKS_ENABLED=1). - default-OFF: when `HOOKS_ENABLED` is unset, `!runtime_input.hooks.is_enabled()`, guarded to skip if the CI environment exports the flag so it only pins the contract it claims to. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(hooks): correct in-memory backend warning text; delegate test ctor (nearai#3938) Address serrrfirat review (2026-06-03): - Low: the in-memory backend warning claimed the LRU cap is shared across tenants, but the Reborn composition constructs a fresh InMemoryPredicateStateBackend per tenant. Rewrite the warn! text and doc comment so the real limitation (process-local replay dedup for multi-host deployments) is accurate, and note the backend is per-tenant in this composition. - Nit: PredicateEvaluator::with_backend (test-only) and with_state_backend had identical bodies; delegate with_backend to with_state_backend so they stay in lockstep. The Medium finding (hooks_config assertion in build_runtime_input caller test) was already addressed in 218a1de. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 13, 2026
…ection (HOOKS_THIRD_PARTY_ENABLED, default OFF) (nearai#3951) * feat(hooks): extension-declared hook section on ExtensionManifestV2 (nearai#3934) Add a `[[hooks]]` declaration surface to the production v2 extension manifest (`ironclaw_extensions::ExtensionManifestV2` and its projected `ExtensionManifest`). Each entry is carried as a structurally-typed `HookSectionEntryV2` DTO — a `local_id` plus the entry's body re-serialized to canonical TOML — so `ironclaw_extensions` (substrate) never imports the `ironclaw_hooks` predicate vocabulary. The composition layer, which depends on both crates, is the single seam that projects these payloads into typed `ironclaw_hooks::HookManifestEntry` values (a later commit). Parse-time structural bounds: `MAX_MANIFEST_HOOKS` (32, matching the downstream per-extension registration cap) and `MAX_HOOK_ENTRY_BYTES` (8 KiB per entry). Entries must be tables carrying a non-empty `id`; ids must be unique within the manifest. `#[serde(default)]` keeps every existing manifest valid (empty `hooks` vec). The DTO holds canonical TOML as a `String` rather than a `toml::Value` so the enclosing `ExtensionManifestV2` keeps its `Eq` derive (`toml::Value` is not `Eq`). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(hooks): composition-layer activation module (loader, first-party hook, flag) (nearai#3934) Add `ironclaw_reborn_composition::hooks` — the single seam that activates the hook framework in production. Implements four numbered pieces of nearai#3934: - Feature flag (item 7): `HooksActivationConfig`, default OFF, resolved from `HOOKS_ENABLED` (only `1`/`true`/`yes`/`on` enable; unset/anything else = OFF). Flag OFF ⇒ `build_hook_dispatcher_builder_factory` returns `None` and the runtime composes no dispatcher — exact pre-hooks behavior. Hard rollout-safety contract. - Manifest → registry loader (item 2): `install_extension_hooks` projects each `ExtensionManifestV2` `HookSectionEntryV2` (canonical TOML) into a typed `HookManifestEntry` and installs it via `HookRegistrar::install` at the `Installed` trust tier. This is the clean-boundary projection: the hook vocabulary lives only here, never in `ironclaw_extensions`. Trust attenuation is enforced by construction (registrar only calls `install_installed_*`). Fail-closed: any projection/install error fails the build loudly. - First-party builtin hooks (item 3): a single illustrative no-op observer (`NoOpObserverHook`), installed regardless of extensions. Ships dark (zero driver-visible effect even with the flag ON). Production catalog is TBD by design — this PR does not invent a first-party hook. - Dispatcher composition (item 5): builds a per-tenant `PredicateEvaluator` over the in-memory state backend (swappable via the new public `PredicateEvaluator::with_state_backend` for durable nearai#3933), validates the full install set once fail-closed, and returns a per-run builder-factory closure. Per-run construction (fresh registry/dispatcher per host build) + per-tenant evaluator give full isolation; the host factory attaches the run-scoped milestone sink internally. Per-tenant scoping is by construction: `build_reborn_runtime` runs once per identity, so everything here is tenant-local — no global registry. The router-backed gate-ref factory (PauseApproval/PauseAuth) and the security-audit sink (nearai#3922, not yet on this branch) are deferred follow-ups; their absence is fail-closed (PauseApproval surfaces as Denied) and noted for the PR body. Not yet wired into the runtime — next commit. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(hooks): wire dispatcher builder factory into build_default_planned_runtime (nearai#3934) Item 6 of nearai#3934. Add an optional `hook_dispatcher_builder_factory` to `DefaultPlannedRuntimeParts` and, in `build_default_planned_runtime`, call `.with_hook_dispatcher_builder_factory(...)` on the production `RebornLoopDriverHostFactory` when it is present. `None` (the default) means no dispatcher is composed — behavior identical to the pre-hooks runtime (rollout-safety contract). The composition layer (`build_reborn_runtime`) resolves the flag via `HooksActivationConfig::from_env()` and builds the factory against this tenant's extension registry (per-tenant by construction — the function runs once per identity). Fail-closed: a malformed manifest hook fails the build here rather than composing a broken dispatcher. A per-run builder factory (not a captured dispatcher instance) is used so the host attaches a run-scoped milestone sink internally per build — per-run telemetry attribution, the nearai#3573 capture-and-stick lesson. All `DefaultPlannedRuntimeParts` construction sites (8 test sites across ironclaw_reborn, ironclaw_product_workflow, ironclaw_reborn_composition) updated with the new field defaulting to `None`. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(hooks): e2e activation tests through build_default_planned_runtime (nearai#3934) Item 8 of nearai#3934. Add four end-to-end tests in crates/ironclaw_reborn/tests/loop_driver_host.rs that drive the *production* composition function `build_default_planned_runtime` with a per-run hook dispatcher builder factory shaped exactly like the composition layer's output (first-party builtin no-op observer + extension-declared `Installed`-tier hooks projected from a manifest entry through `HookRegistrar::install`), then build a host via the composed `host_factory` and invoke a capability: - flag OFF (no factory): allowed capability completes unaffected and reaches the inner host runtime port — the pre-hooks behavior / rollout-safety contract. - flag ON, first-party-only no-op observer: outcome unchanged, inner port reached — the builtin ships dark. - flag ON, extension-declared deny hook: capability denied through the composed runtime and the inner port is never reached (installed at the Installed tier via the registrar; OwnCapabilities scope keyed to the capability provider). - per-tenant isolation: tenant A's deny hook fires; tenant B (separate build_default_planned_runtime composition, no hooks) completes the same capability — proving no cross-tenant leakage. Security-audit-on-deny assertion is intentionally deferred: nearai#3922's SecurityAuditSink is not yet on reborn-integration. It lands with the audit-sink wiring follow-up. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(hooks): thread hook_dispatcher_builder_factory through shared reborn harness (nearai#3934) The root-crate `tests/support/reborn/harness.rs` constructs `DefaultPlannedRuntimeParts` directly; add the new `hook_dispatcher_builder_factory: None` field so the parity-test harness compiles. Default `None` keeps the harness on the no-hooks path (unchanged behavior). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks): proper error handling / safety annotations for activation production paths (nearai#3938 CI) The per-run dispatcher factory closure used `.expect()` on the first-party and extension hook installs, tripping the no-panics CI gate. These installs are pure replays of the install set already validated fail-closed (`?`) against a scratch builder at composition time, so they are genuine invariants. The factory type returns a non-Result `HookDispatcherBuilder` and is invoked deep in the run loop, so the documented `// safety:` suppression is the correct fix here. Hoisted the expect messages into `let` bindings so the `.expect(msg)` call fits on one line, keeping the scanner-required `// safety:` comment on the same line as the call after rustfmt. The malformed-manifest path (TOML projection) already uses real error propagation via map_err/`?` and is unaffected. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(hooks): direct composition-loader coverage + activation-scope docs Address Codex non-blocking follow-ups on nearai#3938. Add three direct tests for the composition-layer hook loader (`install_extension_hooks` via `build_hook_dispatcher_builder_factory`), driving a real `ExtensionRegistry` with `[[hooks]]` declarations rather than mimicking the loader: - valid `own_capabilities` predicate hook installs at the Installed trust tier; the dispatcher carries the derived binding at BeforeCapability alongside the first-party no-op observer - malformed typed hook body (unknown `mode`) fails CLOSED with `RebornBuildError::InvalidConfig`, never a panic (the load-bearing degradation contract for untrusted external manifests) - a hook claiming `scope = same_tenant` without a verified grant is rejected by trust attenuation (fail-closed) No loader bug surfaced: `HookRegistrar::install` already returns `Result` on every malformed/over-scoped path and the loader maps it to `InvalidConfig` via `?`. Document activation scope at both the loader rustdoc and the `build_reborn_runtime` call site: production currently passes only `builtin_extension_registry()`, so third-party installed-extension hooks are not yet surfaced into the runtime path — only first-party-builtin and builtin-package-declared hooks activate today. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * refactor(hooks): thread HooksActivationConfig through input; empty production catalog Two maintainability cleanups on nearai#3938 (firat review): Item 1 — config hygiene: stop reading HOOKS_ENABLED via std::env::var deep inside build_reborn_runtime. Add a typed `hooks: HooksActivationConfig` field to RebornRuntimeInput (default OFF) plus a `with_hooks_config` builder. The composition root now consumes the typed config; the env var is resolved ONCE at the edge (the reborn CLI's build_runtime_input) via HooksActivationConfig::from_env and threaded down. Testable without env mutation; matches the project's env → typed config → composition pattern. Item 2 — empty production first-party catalog: stop shipping NoOpObserverHook as a first-party builtin. install_first_party_hooks is now a no-op (empty catalog); the production type/install/export for a hook that does nothing is gone (removed from lib.rs exports). The activation machinery is still tested end-to-end through the real composition path via a new `build_hook_dispatcher_builder_factory_with` seam that takes a first-party installer; tests pass a `#[cfg(test)]` NoOpObserverHook through it. Pinned the empty-catalog-is-valid contract: flag ON + empty first-party set + no extension hooks composes a valid zero-binding dispatcher (not a panic/error). Reconciled the sibling loader tests/docs (f6c79c0): the in-module tests now drive the test-only seam; the reborn e2e tests already used a test-local no-op and are untouched. Updated activation-scope docs (loader rustdoc + build_reborn_runtime call site) to reflect the now-single live source (builtin-package-declared hooks). Deferred (not touched): switching to the canonical extension registry for third-party installed-extension hooks (nearai#3934 follow-on). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * refactor(hooks): canonical registry + infallible plan + tenant-scoped counter docs/tests (nearai#3938) Addresses serrrfirat's thermo-nuclear re-review on 1e618d0. #1 (runtime.rs:839, canonical registry): make the extension registry a shared composition artifact. `build_local_dev` builds one `Arc<ExtensionRegistry>`, hands it to `HostRuntimeServices::new` AND stores it in `RebornLocalRuntimeServices.extension_registry`. Hook activation in `build_reborn_runtime` now consumes that same `Arc` instead of rebuilding a builtin-only sidecar, so capability dispatch and hook activation cannot drift. Third-party activation stays a follow-up, but it now follows the canonical registry rather than a separate path. #3 (hooks.rs factory machinery): replace the parse/validate/replay duplication + two prose-justified `.expect()` calls with a typed `HookInstallPlan`. TOML is projected once into typed entries, the full install set is validated once against a fresh builder (fail-closed via `?`), and `HookInstallPlan::rebuild` mints a fresh builder per run. The per-run path is infallible by construction: a plan only exists for an install set that already composed cleanly, so a deterministic replay from the identical fresh-empty start cannot fail. One extension-install code path (`project_extension_install_sets` + `install_extension_sets`) is shared by validation and rebuild. #4 (hooks.rs:295, predicate counter scoping): the evaluator/backend is intentionally tenant-scoped and shared across runs (rate/value caps keyed `(hook, tenant, capability)` with no run_id; a run-scoped limit would reset every run and enforce nothing). Document the split explicitly — per-run-fresh dispatcher, tenant-scoped predicate counters — in the module docs and fix the misleading "per-run isolation of hook state" wording in `ironclaw_reborn` (loop_driver_host.rs / runtime.rs). Add `predicate_counter_state_is_tenant_scoped_across_rebuilds`, which drives a real rate-cap predicate through two dispatchers from one factory and proves the second run sees the first run's recorded count. Rename `factory_mints_independent_dispatchers_per_call` -> `rebuild_mints_independent_dispatchers_per_call` and scope it to proving dispatcher freshness only. #6 (loop_driver_host tests): clarify that the hand-built builder factories cover host PLUMBING, not composition activation. Add `build_reborn_runtime_activates_hooks_through_real_composition_path`, which drives the real `build_reborn_runtime` with `HooksActivationConfig` threaded through `RebornRuntimeInput` (env-free) and the canonical registry, proving the production activation wiring composes. #2 (env boundary) and #5 (empty production catalog) were already fixed in 1e618d0; docs touched here for consistency. Known follow-up (not one of the six items, not introduced here): with the flag ON the standalone local-dev runtime does not yet reach `Completed` for a capability turn even with a zero-binding dispatcher — the composition root wires the dispatcher but not the companion hooked-prompt dependencies. The new runtime test asserts `is_terminal()` + the capability path and documents the gap. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(hooks): third-party hook-only projection core (flag, newtype, quarantine, caps) Steps 1-6 of third-party extension hook activation via hook-only projection: - Step 1: HOOKS_THIRD_PARTY_ENABLED sub-flag on HooksActivationConfig (default OFF; is_third_party_enabled() requires master flag too). Resolved at the CLI edge via from_env(). - Step 2: tenant_extension_root(&TenantId) derives the fixed /system/extensions/<tenant> root from identity (never caller-supplied); projection-layer strict-child / no-`..` containment check. - Step 3: build_hook_projection_registry assembles a HookProjectionRegistry (type-enforced hook-only newtype: no Deref / conversion back to ExtensionRegistry, so it can never reach HostRuntimeServices::new / the capability path). Sub-flag OFF => builtin-only, byte-identical to nearai#3938. - Step 4/4a: atomic per-extension quarantine — untrusted (InstalledLocal) sets validated whole against a scratch builder, committed only if the whole set passes; any failure drops the extension's hooks entirely, emits a hook.quarantined security_audit tracing event (warn!, not info!), and continues. Trusted (HostBundled) sources stay fail-closed-whole-build. - Step 5: MAX_INSTALLED_EXTENSIONS_CONSIDERED / MAX_TOTAL_HOOKS_PER_TENANT DoS caps; count_total_bindings() accessor on HookDispatcher(Builder); pre-read MAX_MANIFEST_BYTES bound via read_file_bounded in discovery. - Step 6: third-party WASM stays out (loader registrar has no wasm_runtime) => WASM-bodied hook quarantines + build continues. Registrar-only invariant: projection installs go exclusively through HookRegistrar::install (ceiling + spoof-blocked owning_extension), never the direct builder installer API. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(hooks): FS-scoped tenant isolation (Option 1), trust matrix, registrar-only assertion Resolve the discovery/path conflict: the discovery layer hardcodes package roots to /system/extensions/<id> because the per-tenant RootFilesystem is the scope boundary (as with every other tenant-scoped resource), not a tenant path segment. So: - tenant_extension_root -> fixed /system/extensions (no tenant segment). The per-tenant RootFilesystem handed to discovery IS the isolation boundary. Documented as load-bearing; the openat2(RESOLVE_BENEATH)/O_NOFOLLOW backend hardening follow-up is what protects it (gating note kept prominent). - build_local_dev mounts /system/extensions to a per-owner host subtree under the storage root (per-identity by construction, not a process-global mount); exposed via RebornLocalRuntimeServices.extension_filesystem. - enforce_root_containment retained as defense-in-depth. Tests: - Integration (real build_hook_projection_registry + build_hook_dispatcher_ builder_factory through a fake RootFilesystem, not a loader look-alike): containment (hook present / capability absent by construction), FS-as-boundary tenant isolation proof (two distinct per-tenant filesystems; A can't see B), bad dir name skipped, id mismatch not a panic, surplus-extensions DoS cap, sub-flag OFF discovers nothing. - Per-hook-point trust matrix: BeforeCapability installed deny IS allowed and fires (Gate reachable); before_prompt predicate quarantined + build continues; after_model/after_capability/after_checkpoint/event_triggered WASM-only => quarantined + build continues; owning_extension derived (not spoofable). - Discovery pre-read bound: oversized manifest rejected via stat WITHOUT reading the body (fake fs panics on get); within-bound proceeds to read. - ironclaw_architecture source assertion: the hooks.rs projection path never calls install_installed_* directly (registrar-only invariant). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * docs(hooks): correct Option-1 path-shape references in test comments Update the third-party projection integration-test module docs to reflect the FS-scoped isolation model (fixed /system/extensions root; per-tenant filesystem is the boundary), not the abandoned /system/extensions/<tenant> path segment. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks): tolerant+bounded third-party discovery, structural hook-only containment Addresses Codex P1/P2 + serrrfirat P1 on nearai#3951. Critical 1 (discovery-stage DoS): add `ExtensionDiscovery::discover_with_manifest_contracts_tolerant_bounded` (+ `discover_extensions_tolerant_bounded` host-runtime wrapper). It lists+sorts the root once, then reads/parses at most `max_extensions` manifests, recording the surplus as quarantines WITHOUT reading them. The hook projection calls this with `MAX_INSTALLED_EXTENSIONS_CONSIDERED`, so the count cap fires before the per-manifest read storm. New all-or-nothing path delegates to a shared `load_package_entry` so per-package semantics are identical. Critical 2 (fail-open): tolerant discovery quarantines a single malformed/oversized/id-mismatched package and CONTINUES; valid siblings still load. The builtin-only fallback is now reserved solely for failure to LIST THE ROOT (directory unreadable). One bad manifest can no longer drop a tenant's entire legitimate third-party hook set. Refinement 3: the per-tenant hook budget is consumed only AFTER a successful merge, so a quarantined/duplicate package no longer burns budget. Refinement 4: the registrar-only arch assertion now scans the WHOLE composition crate (every non-test source) and forbids all installed-tier-minting primitives crate-wide (`install_installed_*`, `install_observer(`, `insert_binding(`, `HookTrustClass::Installed`) — not just a hooks.rs substring scan. Installed-tier bindings can only be minted via `HookRegistrar::install`. serrrfirat P1 (structural containment): `HookProjectionRegistry` no longer wraps `ExtensionRegistry`. It carries `Vec<HookProjection>` — hook metadata only (id/version/source/root/[[hooks]]). The projection literally cannot reach capabilities because it does not hold them; containment is by data shape, not a withheld conversion. Removes the `ExtensionPackageView` ceremony. Tests: bounded read-storm cap (read-counting fs panics on surplus), tolerant per-package quarantine, root-unreadable fallback, quarantined-package-does-not- consume-budget, malformed-sibling-survives at the projection layer. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * refactor(hooks): address serrrfirat review — tenant-attributed install audits + hooks decomposition Addresses the maintainability review on nearai#3951. Findings #1 (narrow hook-only boundary), #2 (per-extension discovery quarantine + mixed-batch test), and #5 (behavioral arch-test invariant) were already satisfied by the head commit (2b62597); this commit closes the two remaining items and hardens the arch test against the decomposition: - #3 (tenant attribution): add `build_hook_dispatcher_builder_factory_for_tenant`, threading the authenticated `tenant_id` (and its derived extension root) into the install-time quarantine-audit seam. `build_reborn_runtime` now calls it, so install-time quarantine audits carry the real tenant instead of the synthetic `reborn-hook-projection` fallback (closing the split where only discovery-time audits were attributed). New caller-driven test `for_tenant_entry_point_attributes_install_time_quarantine_to_real_tenant` asserts attribution via a deterministic thread-local audit capture (immune to tracing's process-wide max-level filter under parallel tests). - #4 (decomposition): split the 1.7k-line `hooks.rs` into a focused `hooks/` module — `mod.rs` (flag/config + public surface), `projection.rs` (hook-only `HookProjection`/`HookProjectionRegistry` containment + discovery/admission), `factory.rs` (first-party install, per-extension quarantine validation, fresh-per-build replay), `audit.rs` (`hook.quarantined` emission), and `tests.rs` (the test matrix). Behavior-preserving; no logic change. - arch test: skip dedicated test-module files in the registrar-only scan so the #4 decomposition cannot break it; the whole-crate behavioral invariant is preserved. - audit emission uses `debug!` (not `warn!`) per the background/hook-path logging rule, on the stable filterable `security_audit` target. - gemini nearai#353: add the documented no-empty-segment guard to `enforce_root_containment` (defense-in-depth, not relying on VirtualPath canonicalization). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(hooks): cover hook-entry rejection branches; document InvocationCount cap semantics (nearai#3938) Address henrypark133 review (review 4367870023): - Add extension-manifest tests for the three previously-uncovered hook-entry validation branches: non-table `[[hooks]]` element, whitespace-only `id`, and oversized entry (HookEntryTooLarge). - Document the InvocationCount inclusive-allow / deny-on-overflow semantics inline at the comparison site; behavior unchanged and still pinned by the cap test. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(deps): pin kuchikikiki to 0.9.1 (0.9.2 yanked) cargo-deny failed on the yanked kuchikikiki 0.9.2 pulled in transitively via readabilityrs. Downgrade to 0.9.1 at the lockfile level. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(reborn-cli): assert hooks config threaded in caller test (nearai#3938) Addresses the review finding that `build_runtime_input_maps_configured_cli_identity` exercised `build_runtime_input` but never asserted the `hooks` field, so a regression silently dropping `with_hooks_config(HooksActivationConfig::from_env())` or flipping the default-OFF rollout-safety contract would pass. Per `.claude/rules/testing.md` ("Test Through the Caller"): add two assertions to the existing caller-level test: - threading: `runtime_input.hooks == HooksActivationConfig::from_env()`, proving the env-resolved config is actually threaded through and not dropped. Verified via TDD that dropping the wiring fails this assertion (run under HOOKS_ENABLED=1). - default-OFF: when `HOOKS_ENABLED` is unset, `!runtime_input.hooks.is_enabled()`, guarded to skip if the CI environment exports the flag so it only pins the contract it claims to. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(hooks): cover build_reborn_runtime third-party wiring + async dir create (nearai#3951) Address serrrfirat review findings M1 and L2. M1: add an integration test in tests/runtime.rs that drives build_reborn_runtime with HooksActivationConfig::enabled().with_third_party_enabled(true), a real /system/extensions manifest tree on the local-dev host filesystem, and tenant attribution. Asserts the runtime builds, starts a conversation turn, and shuts down cleanly — exercising the runtime.rs third-party discovery input + projection registry + tenant-threading wiring that was previously uncovered (the projection tests call build_hook_projection_registry / the dispatcher factory directly, and every other build_reborn_runtime call used the default disabled config). Verified the test fails when the wiring is broken. L2: switch the new factory.rs blocking std::fs::create_dir_all for the extensions host root to tokio::fs::create_dir_all(...).await with the same error mapping, so it no longer blocks the tokio executor thread inside the async build_local_dev. The two pre-existing std::fs calls (lines 132/136) are out of this PR's diff per the posted promise and are left untouched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(hooks): L1 quarantine-surfacing gate doc, L3 robust test-mod strip, M1 coverage-gap TODO Address serrrfirat 2026-06-03 review (M1/L2 already landed in 9866793). L1 (security observability): hook.quarantined audit events are emitted only via tracing at the security_audit target / debug! level, which production typically disables. Document durable quarantine surfacing as a hard production-enablement prerequisite for HOOKS_THIRD_PARTY_ENABLED, alongside the existing openat2(RESOLVE_BENEATH)/O_NOFOLLOW FS-hardening note, at all three gate doc sites: HooksActivationConfig (hooks/mod.rs), the runtime.rs composition -root gate comment, and the audit.rs module doc. L3 (robustness): strip_test_module matched #[cfg(test)]\nmod tests specifically and only the first occurrence. Generalize the anchor to #[cfg(test)]\nmod (any module name) so a refactor that renames the test module or adds a second #[cfg(test)] mod block is still fully stripped, preventing false positives in the FORBIDDEN_INSTALLED_PRIMITIVES architecture scan. M1 (test coverage): the build_reborn_runtime third-party wiring test already landed in tests/runtime.rs (9866793). Add the reviewer-requested TODO preserving the removed test's Cancelled-outcome coverage gap: the stub local-dev gateway cancels the turn before any capability dispatches, so the test exercises discovery + projection + tenant-threading at build/start but not end-to-end hook enforcement. NOTE: third-party discovery is intentionally tolerant (skips unparseable manifests), so this test catches compile-time field/arg regressions and build-path failures but not a silent manifest-read drop; documented for the reviewer. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(hooks): correct in-memory backend warning text; delegate test ctor (nearai#3938) Address serrrfirat review (2026-06-03): - Low: the in-memory backend warning claimed the LRU cap is shared across tenants, but the Reborn composition constructs a fresh InMemoryPredicateStateBackend per tenant. Rewrite the warn! text and doc comment so the real limitation (process-local replay dedup for multi-host deployments) is accurate, and note the backend is per-tenant in this composition. - Nit: PredicateEvaluator::with_backend (test-only) and with_state_backend had identical bodies; delegate with_backend to with_state_backend so they stay in lockstep. The Medium finding (hooks_config assertion in build_runtime_input caller test) was already addressed in 218a1de. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 17, 2026
…earai#4559) * docs: trace commons agent onboarding design spec Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: address spec review findings (trust anchoring, key staging, consumption atomicity, replay validation) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: spec review round 2 nits (server-anchored tenant wording, pending-key cleanup) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: implementation plan for trace commons agent onboarding Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: address plan review findings (scope threading refactor, dispatch model, dev-deps, LazyLock hazard) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: plan review round 2 fixes (literal dep versions, context constructor threading depth) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: incorporate server-agent coordination feedback (optional community/profile/leaderboard URLs) From TraceCommons/trace-commons#136-nearai#141 comments: onboard response gains optional browser-surface navigation hints, sanitized client-side (HTTPS or dropped), never part of issuer trust anchoring. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): onboarding wire types matching trace-commons-server contract Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): invite URL parsing with origin trust anchoring Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): device keypair lifecycle with pending staging and self-signed workload JWTs Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): auth_mode and device_key_id policy fields with legacy-compatible defaults Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): onboard() orchestration with trust anchoring and retry-safe key staging Wire invite parsing, device key staging, onboard POST, issuer origin trust anchoring, ingest_url HTTPS enforcement, keypair promotion, and policy write into onboard_at_dir(). Refactors invite.rs to extract pub(crate) is_https_or_loopback, origin_of, and host_only helpers shared with mod.rs (one source of truth for origin/bracket handling). Adds axum mock-issuer tests covering the happy path, mismatch rejection, terminal vs transient error key retention, insecure ingest URL, loopback ingest allowance, community URL sanitisation, and retry key reuse. Partial-failure lockout fix (spec §2.2): promote() no longer deletes the pending file. The flow now writes the tenant key file, then the policy, and only discards the pending file after BOTH durably succeed. If the policy write fails the pending key survives, so a retry reloads the same key (server idempotency returns the original registration) and harmlessly overwrites the tenant file — no permanent lockout from a consumed invite with a regenerated keypair. Regression test simulates a policy-write failure (policy.json pre-created as a non-empty dir so the atomic rename fails), asserts Err(Persist) with the pending key intact, then asserts a retry succeeds reusing the same device_key_id. Response validation (defense-in-depth): reject schema_version != the v1 response constant as MalformedResponse, and cross-check the response device_key_id against the locally derived id (we never trust the response value for policy; a disagreement is now treated as a tamper signal and rejected). Both covered by tests. The onboard response body is read with the 64 KB cap enforced per-chunk during streaming (mirroring read_bounded_trace_upload_claim_response) rather than buffering the whole body first, so a hostile server cannot force a large allocation. Also fixes a pre-existing test-isolation defect surfaced by the added load: the remote-request timeout test configured a 50ms timeout via the process-global IRONCLAW_TRACE_REMOTE_REQUEST_TIMEOUT_MS env var. set_var is process-global, so under parallel execution the 50ms value leaked into other tests' trace HTTP clients, producing spurious `operation timed out` failures against fast local mocks. Replace the env mutation with a task-scoped TEST_REMOTE_REQUEST_TIMEOUT_OVERRIDE task-local (visible only within the awaiting test's own task tree, zero production change; documents the spawn caveat), and decouple the timing assertion from a tight wall-clock race so it no longer flakes when reqwest's timer is delayed under an oversubscribed runtime. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): device-key self-signed workload JWT branch in upload-claim refresh Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(engine): trace_commons onboard and status first-party tools with agent guidance Add two model-visible first-party capabilities to the Reborn engine: - builtin.trace_commons.onboard: drives operator-invite enrollment flow with explicit per-conversation consent gate (confirmed=true required before any network call); maps OnboardOutcome/OnboardError to clean agent-readable JSON - builtin.trace_commons.status: read-only enrollment state inspector Wires ironclaw_reborn_traces into ironclaw_host_runtime, creates schema files (schemas/builtin/trace-commons-{onboard,status}.{input,output}.v1.json) and prompt doc files (prompts/builtin/trace-commons-{onboard,status}.md) at the manifest-derived paths. Includes 11 unit tests covering input parsing, consent refusal, success/error value formatting, and status formatting. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: add Task 11 — credits visibility (console display + agent-queryable balance) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(engine): e2e trace commons onboarding through capability dispatch Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(traces): document agent onboarding flow in trace-commons internal doc Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: correct Task 11 console scope (credit endpoint already exists; frontend = coordinate with designer) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): trace_commons.credits agent-queryable balance tool Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(gateway): minimal Trace Commons credits card in settings Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(traces): store upload-claim endpoint in policy; preserve primary onboard error; block metadata/link-local/multicast issuers Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): route agent onboarding HTTP through host network-egress policy (nearai#4560) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * build: update Cargo.lock for trace-commons onboarding dev-deps Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(traces): drop orphaned schema/prompt files (main resolves builtin schemas inline; prompt_doc_ref dropped) Post-merge cleanup: main's first_party_tools now resolves builtin input schemas via the inline schemas.rs match (trace_commons arms added during the merge) and sets prompt_doc_ref: None for all builtins, so the physical trace-commons-*.json schema files and trace-commons-*.md prompt docs are no longer referenced. The onboard consent contract remains in the capability description and is enforced in dispatch_onboard. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Fix Trace Commons invite hash contract * fix(traces): grant trace_commons capabilities in local-dev policy The three builtin.trace_commons.* capabilities were declared in the first-party package but had no [[grants]] entries in local_dev_capability_policy.toml, so local-dev runs (repl/serve) filtered them out of the model-visible tool surface entirely. The provider-level authority_effects ceiling had external_write, but the per-capability grants were never added. onboard gets the local_dev_wildcard egress profile (invite origins are operator-chosen; private/metadata IP ranges stay blocked by the shared enforcer). status/credits are read-only. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(traces): add Reborn e2e coverage for trace_commons first-party tools Closes the coverage gate failure: builtin.trace_commons.{onboard,status, credits} were declared in the first-party package but missing from REBORN_FIRST_PARTY_E2E_COVERED_CAPABILITIES, failing reborn_builtin_first_party_capability_e2e_coverage_is_complete on both the Reborn root tests and all-features CI jobs. Adds a trace_commons host-runtime harness (network policy populated so the onboard Network-effect obligation passes) and a parity test driving all three capabilities through the scripted model loop: onboard with confirmed=false exercises the deterministic consent gate with no network, status and credits return the unenrolled/zero-credit defaults. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(traces): community profile second opt-in (token mint + profile set) After device-key enrollment, public leaderboard attribution is a second, separate opt-in: IronClaw mints a short-lived profile token from the claim issuer with consent_scopes=[public_attribution] and empty allowed_uses (such a claim cannot submit traces), then either prints it for the web profile page or performs the profile update itself. The browser cannot sign device-key requests, so the token must be minted by IronClaw — previously this step was impossible and agent guidance invented flows. - ConsentScope::PublicAttribution mirrors the server protocol enum; default_allowed_uses_for_scope returns empty for it. - mint_profile_attribution_token_for_scope / set_community_profile_for_scope / withdraw_community_profile_for_scope reuse the hardened issuer HTTP path (allowlist validation, pinned DNS, no redirects, bounded reads, token never in errors). PUT/DELETE /v1/community/profile per the server contract; handle (3-32 ASCII alnum/-/_) and bio (<=280 bytes) validated client-side. - CLI: ironclaw-reborn traces profile token|set|withdraw. - Onboard tool next_steps now describes the profile second opt-in so agent guidance stops inventing browser login flows. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(traces): autonomous turn-end trace capture in the Reborn runtime The Reborn binary could onboard, report status/credits, and manage profiles, but never captured or submitted traces — the autonomous pipeline existed only in the v1 agent loop. This wires it into the Reborn runtime composition: - TraceCaptureTurnEventSink subscribes best-effort to the turn lifecycle bus (the existing turn_event_sink injection seam). On Completed/Failed events with an explicit owner it spawns a detached task that reads the owner's standing policy (one file read for non-enrolled users), loads the recent thread history (last 24 messages, 5 turns — v1 parity), adapts user/assistant text rows into the neutral ConversationMessage shape, redacts + scores locally, and queues + immediately flushes eligible envelopes. All failures are debug!-logged and never touch the turn lifecycle path. - A periodic flush worker (300s, 25/scope — v1 parity) retries queued envelopes for the runtime owner plus every scope observed since boot, with CancellationToken shutdown alongside the other workers. - TraceClientAutonomousCaptureRequest gains outcome_override so the lifecycle event's terminal status (authoritative in Reborn, where transcripts carry no structured outcome payload) marks failed turns as TaskSuccess::Failure; v1 passes None (no behavior change). - Tool-result rows and credit-notice delivery are documented follow-ups (refs-only records; no composition-level outbound channel surface). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(traces): end-to-end auto-capture through send_user_message Proves the full Reborn auto-submission chain with a real runtime: a completed turn for an enrolled owner scope lands a redacted envelope in that scope's submission queue with no manual trace command — turn completion -> lifecycle bus -> capture sink -> thread-history read -> redact/score -> eligibility -> queue (+ local-failing immediate flush leaves the entry for the retry worker). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Expose Trace Commons profile token tool * Expose Trace Commons profile set tool * Allow Trace Commons profile setup from agent * feat(webui-v2): Trace Commons credits card in WebChat v2 settings Adds GET /api/webchat/v2/traces/credit and a read-only Trace Commons settings tab to the v2 SPA, giving webui-v2-beta parity with the v1 console's credits card. - Route follows the descriptor system end to end: bearer-auth required, NoBody, 120/60 per-caller read rate limit; descriptor-driven body/rate-limit enforcement applies automatically. - RebornServicesApi::trace_credits derives the trace scope exclusively from the authenticated caller's user id (never from query/body) and reads contributor-local state via ironclaw_reborn_traces (policy + trace_credit_report), soft-falling back to an unenrolled zero-state on missing/unreadable local state, mirroring builtin.trace_commons.credits. - SPA: Trace Commons subtab (enrollment, pending/final credit, delayed ledger delta, submission counts, last submission/sync, recent credit explanations) with the server-authoritative framing and a not-enrolled empty state pointing at agent onboarding. - Tests: descriptor contract row, handler oneshot, and three composed- router serve tests (200 zero-state, 401 without bearer, enrolled policy reporting with per-test scope isolation). - Drive-by: cfg-gate openai_user_id in webui_serve.rs to clear a pre-existing unused-variable warning under default features. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Exempt Trace Commons profile setup from local-dev gate * Route Trace Commons profile writes to ingest * review(4559): address serrrfirat feedback - Drop stray working-note markdown files from the repo root (they rode in via an early origin/main merge and are not this PR's documentation). - trace_commons_dispatch_e2e: setup_base_dir is now a OnceLock that every test calls first — the previous 'single-threaded during init' claim was wrong under tokio's multi-threaded test runtime, and two of three tests skipped the setup entirely. - settings.js: extract shared appendDisplayGroup + declarative row defs; loadTraceCommonsCredits drops from ~120 lines of manual DOM to a rows array; also removes a double-escape (textContent + escapeHtml) on explanation lines. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(webui-v2): add traceCommons i18n keys to all locales The credits card added the traceCommons.* key set to en.js only; the i18n consistency test (all_locales_share_the_en_key_set) requires every locale to carry the same key set. Adds translated entries to ar, de, es, fr, hi, ja, ko, pt-BR, uk, and zh-CN. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(reborn): fail loud with source context on malformed local-dev master key The local-dev secret store resolver read the cached key file (and the SECRETS_MASTER_KEY env fallback) and passed the material straight into SecretsCrypto::new several layers deep. A corrupt or low-entropy key (e.g. a 64-char all-zeros value, which passes the length floor but has one distinct byte) surfaced only as the opaque "Invalid master key", with no pointer to the file the operator must fix. - Add ironclaw_secrets::validate_master_key_material as the single source of truth for master-key rules; SecretsCrypto::new delegates to it. - resolve_local_dev_secret_master_key now validates at the source (cached file vs SECRETS_MASTER_KEY env) and returns a RebornBuildError::InvalidConfig naming the offending path/env var and the actual constraint, before any crypto is constructed. - A malformed env value is now rejected before being persisted to the cached key file (no more poisoned-cache state). Tests: malformed-file path-context rejection, malformed-env source-context rejection, valid cached file accepted. Refs nearai#4741 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(webui-v2): add Trace Commons credits card to chat sidebar Surface trace contribution credits at a glance in the chat sidebar, above the conversation list. Previously credits were only visible under Settings -> Trace Commons. - New SidebarTraceCredits component reuses the existing useTraceCredits hook (/api/webchat/v2/traces/credit) — no new endpoint. Renders only when enrolled; loading/error/not-enrolled render nothing to keep the sidebar clean. Shows final credit and accepted/submitted counts and clicks through to Settings -> Trace Commons for the full ledger. - useTraceCredits now refetches (60s interval + on window focus) so the card and the Settings tab reflect newly-accepted submissions live. - Add one compact i18n key (traceCommons.cardAccepted) across all 11 locales; reuse existing keys for the rest. - Source-shape regression test in assets.rs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reborn-traces): reconstruct tool calls in turn-end trace capture The Reborn capture adapter dropped every tool-result row, so captured trace envelopes were text-only. That left the two highest-value scoring levers — replayability (0.20) and tool coverage (0.15) — permanently at zero, so even agentic tool-using turns scored as plain chat and stayed below the 0.35 submission gate. Nothing ever submitted. conversation_messages_from_records now reconstructs a `tool_calls` message from each run of ToolResultReference rows that carry `tool_result_provider_call` replay metadata, collapsing consecutive rows into one message positioned between the user message and the assistant response (the shape capture_turns_from_conversation_messages' per-turn lookahead consumes). Tool names always flow through so the value scorecard sees required_tools/replayable; raw tool payloads stay consent-gated downstream by include_tool_payloads. Rows without provider metadata remain dropped. TDD: - adapter unit tests: single tool call -> tool_calls message; consecutive calls collapse into one; ref without provider metadata still dropped. - integration guard: a captured tool-using turn's queued envelope carries replay.required_tools + replayable=true. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reborn-traces): read capture history from context window, not display projection Tool-call reconstruction (previous commit) had no data to work with: the capture history source read SessionThreadService::list_thread_history, whose product-display projection (history_message) hard-nulls tool_result_provider_call. So even though tool calls persist with full provider metadata, the adapter received None on every tool row, dropped them, and produced a text-only envelope that scored below the 0.35 submission gate. Nothing ever submitted. SessionThreadHistorySource now reads load_context_window (the model-context/replay view, which preserves tool_result_provider_call) and maps ContextMessage -> ThreadMessageRecord via context_window_to_records. This is the semantically correct source for trace capture anyway: the replay transcript, not the display transcript. TDD: a caller-level test (per .claude/rules/testing.md "test through the caller") drives SessionThreadHistorySource against a real InMemorySessionThreadService with an appended tool result, asserting the returned tool row keeps provider_call. Failed on list_thread_history (None), passes on load_context_window. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reborn-traces): auto-submit traces with PII risk below High Previously any non-Low residual PII risk was blocked from auto-submission two ways: the manual-approval eligibility gate held everything != Low, and the value scorecard halved the score (privacy_gate Medium 0.5) and subtracted a 0.60-weighted penalty. A minimal tool trace scores ~0.36 at Low (barely over the 0.35 gate), so any Medium penalty collapsed it to 0 — nothing below High could ever submit. Treat below-High residual risk as clean for auto-submission (the deterministic redactor has already scrubbed detected PII): - trace_autonomous_eligibility manual-approval gate now holds only High (== High, was != Low). - privacy_gate: Low|Medium => 1.0 (was Medium 0.5); High => 0.0. - privacy_risk_score: Low|Medium => 0.0 (was Medium 0.5); High => 1.0. High remains fully blocked: privacy_gate zeros its score and the gate holds it for manual review. The 0.35 submission gate leaves no headroom for a partial Medium discount on a minimal trace, so below-High is clean rather than partially penalized. TDD: medium_pii_tool_trace_auto_submits_while_high_is_held asserts a Medium-risk tool trace clears 0.35 and auto-submits while High is held. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(reborn-traces): design for Trace Commons held-trace review Held traces are currently dropped on the autonomous capture path with no visibility or authorize path. This plan reuses the existing hold-sidecar machinery (TraceQueueHold / .held.json / read_trace_queue_holds_for_scope / ManualReview) and adds: retain held traces, surface a held count+list on the /traces/credit response, a card/tab UI, and a promote-as-is authorize endpoint. Four independently-shippable TDD slices. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reborn-traces): retain manual-review held traces instead of dropping (slice 1) Autonomous turn-end capture dropped every held trace (logged at debug, envelope discarded), so PII-gated traces were unrecoverable and invisible. Slice 1 of the held-review feature retains manual-review holds: - TraceQueueEligibility::Hold now carries a typed TraceQueueHoldKind (ManualReview for the High residual-PII gate; PolicyGate for score / tool-allowlist / submission-class gates), replacing reason-string classification at the flush call site. - TraceClientAutonomousCaptureOutcome::Held carries the built envelope and its kind so callers can persist it. - New queue_trace_envelope_as_held_for_scope: queues the envelope plus a ManualReview .held.json sidecar under one scope lock; the flush worker already skips held sidecars, so it is retained but not submitted. - capture_turn_trace retains ManualReview holds and still drops PolicyGate holds (low-value traces never pollute the review surface). TDD: held-retain function (RED on missing sidecar -> GREEN), eligibility kind classification, and caller-level capture tests (an AWS-key message forces High PII -> retained ManualReview hold; a sub-threshold trace is dropped, not retained). Refs docs/plans/2026-06-10-trace-commons-held-review.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(webui-v2): surface manual-review held count + list on /traces/credit (slice 2) Held traces retained by slice 1 were invisible to the UI. Slice 2 surfaces them on the existing trace-credits response so one fetch powers the whole card/tab. - ironclaw_reborn_traces: manual_review_holds_for_scope() returns only ManualReview holds (excludes PolicyGate value-gates and transient RetryableSubmissionFailure retry holds), via an extracted retain_manual_review_holds filter. - RebornTraceCreditsResponse gains manual_review_hold_count + holds[] ({ submission_id, reason }). Sanitized: submission id and the already privacy-safe hold reason only, never raw trace content. TDD: retain_manual_review_holds filter unit test (excludes policy/retry), disk-level manual_review_holds_for_scope test, and the facade zero-state test asserts the new fields default empty. webui_v2 handler contract tests (42) still pass with the propagated fields. Refs docs/plans/2026-06-10-trace-commons-held-review.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(webui-v2): show held-for-review traces on card + Settings tab (slice 3) Surface the manual-review held count/list from slice 2 in the UI. Both render only when there are holds, so the common (nothing-held) state is unchanged. - Sidebar card: "{count} held for review" line when manual_review_hold_count > 0. - Settings -> Trace Commons tab: a "Held for review" section listing each held trace's sanitized reason + submission id from holds[]. - No hook/api change: fetchTraceCredits already returns the raw response, so credits.holds / credits.manual_review_hold_count are available. - Three i18n keys (cardHeld, heldTitle, heldDescription) across all 11 locales. The per-trace Authorize action ships with its endpoint in slice 4 (so the UI never offers a button that 404s). Source-shape assertions extended. Refs docs/plans/2026-06-10-trace-commons-held-review.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(webui-v2): authorize held traces for submission (slice 4) Complete the held-review feature with a promote-as-is authorize action across the stack. ironclaw_reborn_traces: - TraceContributionEnvelope gains `manual_review_authorized`; an authorized envelope submits past every gate in trace_autonomous_eligibility (the flush re-evaluates eligibility each pass, so removing the hold sidecar alone is not enough to promote). - authorize_manual_review_hold_for_scope: stamps the envelope (durable consent record) BEFORE removing the .held.json sidecar, so a crash between the two leaves the trace held (fail closed). Only ManualReview holds are authorizable; unknown submissions return Ok(false), not an error. ironclaw_product_workflow: - RebornServicesApi::authorize_trace_hold derives scope from the authenticated caller (the path submission id is never cross-scope authority), validates the id, and returns RebornTraceHoldAuthorizeResponse. ironclaw_webui_v2: - POST /api/webchat/v2/traces/holds/{submission_id}/authorize — NoBody, mutation rate limit, bearer auth. Descriptor + handler + router + contract table (now 46 routes). Frontend: - authorizeTraceHold api, an authorize mutation in useTraceCredits that invalidates the credits query on success, and a per-hold Authorize button on the Settings tab. `authorize`/`authorizing` i18n in all 11 locales. TDD: authorize promotes a High-PII held envelope past all gates; facade zero-state; webui_v2 descriptor/handler contracts; composition serve (47); source-shape assertions. clippy/fmt clean across crates. Refs docs/plans/2026-06-10-trace-commons-held-review.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): loopback dev claim exception + profile_set consent gate Address the two codex P2 findings from review: - Preserve loopback claim uploads after onboarding: the loopback-HTTP dev invite form stores a loopback claim/ingest endpoint in the policy, but the claim/ingest validators required https and rejected loopback hosts, so a successful loopback onboarding could never mint a claim or submit credits. The validators and the pinned DNS resolution now honor the same literal-loopback exception as invite parsing (shared is_loopback_host predicate); for loopback hosts the pinned resolution additionally requires all resolved addresses to be loopback. Non-loopback http, internal hostnames, and private ranges stay rejected, and the issuer allowlist still applies. - Require explicit confirmation before community profile updates: trace_commons.profile_set now has the same hard confirmed=true input gate as onboarding — it short-circuits with consent_required before the enrollment check and any network write, since the capability is approval-gate-exempt in local-dev policy. Schema and manifest document the field. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(merge): thread attachments field through trace-capture record construction main added ThreadMessageRecord.attachments (Vec<AttachmentRef>); the trace-capture reconstruction path and its two test helpers construct records and must set it. The capture path reconstructs records from a context window for redaction/scoring and carries no attachment refs of its own, so Vec::new() is correct. * fix(traces): adapt v1 autonomous capture to new Held variant shape The merge brought in slice 1 of the held-trace-review feature, which changed TraceClientAutonomousCaptureOutcome::Held from { submission_id, reason } to { kind, reason, envelope } so manual-review holds can be retained instead of dropped. The v1 autonomous-capture path in thread_ops.rs still matched the old shape, breaking the `--no-default-features --features libsql` build (and default build). Adapt the v1 path to the new shape and give it the same retain-or-drop parity as the Reborn capture path (ironclaw_reborn_composition::trace_capture): ManualReview holds are retained via queue_held_envelope_for_scope (the on-disk held queue is shared, so a v1-captured hold surfaces in the v2 review UI); policy/value gates are dropped as before, just logged. Behavior mirrors the tested Reborn path (send_user_message_auto_queues_trace_for_enrolled_scope); the v1 autonomous-capture path is a detached tokio::spawn with no unit-testable seam, so no focused regression test is added. [skip-regression-check] * fix(traces): set manual_review_authorized in reborn-cli test envelope fixture The merge brought in the held-trace-review manual_review_authorized field on TraceContributionEnvelope. The reborn-cli trace_queue test fixture constructs the envelope directly and missed the field, breaking `cargo clippy --all-features --tests` and `Tests (all-features)` (the fixture is test-only, so the libsql binary build did not surface it). Fresh queued envelopes are not yet authorized, so false is correct. [skip-regression-check] * test(traces): pass confirmed=true in profile_set parity step The trace_commons first-party-tools parity test invoked profile_set without confirmed=true and asserted the NotEnrolled enrollment-gate result. Commit 6bc776d added the public-attribution consent gate to dispatch_profile_set, which now short-circuits to consent_required before the enrollment check when confirmed is unset — so the test's NotEnrolled assertion failed (the gate output carries no error_code). Pass confirmed=true so the call clears the consent gate and reaches the enrollment check, deterministically returning NotEnrolled with no network (the scope never onboarded). Matches the unit-test pattern established for the other profile_set tests in the same change. [skip-regression-check] * fix(traces): onboarding-security + contribution correctness (coderabbit batch 1) Addresses 6 coderabbit findings in ironclaw_reborn_traces: - device_key.rs: re-assert 0o700 on pre-existing key dirs (not just on create), so broader perms on an existing device_keys/ or pending/ can't leave invite/tenant hashes enumerable. - device_key.rs: fail closed on load when on-disk public_key/device_key_id don't match the loaded private key (tampered/partial files no longer load an inconsistent identity that only fails later at remote auth). - invite.rs: scope the staged pending-key filename by invite ORIGIN, not just code, so two issuers reusing one invite code can't share a device key (invite_hash stays code-only as the server allowlist subject). - onboarding/mod.rs: reject ingest_url values with embedded userinfo before persisting, so a malicious onboarding response can't smuggle credentials into policy.json + outbound requests. - contribution.rs: preserve mount path prefixes when deriving the community-profile endpoint (mirrors trace_submission_status_endpoint); a prefixed deployment no longer 404s on profile PUT/DELETE. - contribution.rs: fail closed in trace_autonomous_eligibility on envelopes with no allowed-uses (public_attribution-only) instead of relying on the remote to bounce them. Updated two retry tests that encoded the cross-issuer key-sharing bug now fixed: they retried against a second mock on a different port; a new spawn_flaky_mock_issuer keeps the retry on the same origin so it exercises genuine same-issuer pending-key reuse. Added regression tests for each fix. * fix(trace-commons): address coderabbit review findings on nearai#4559 - index.html: add type="button" to the Trace Commons settings subtab to prevent accidental form submission. - settings.js + i18n/en.js: route the Trace Commons credits copy through I18n.t(...) and register the matching locale keys (matches the existing surface pattern; en-only like settings.traceCommons, fallback covers rest). - factory.rs: drive the malformed SECRETS_MASTER_KEY env case through the real caller resolve_local_dev_secret_master_key (via an env-parameterized inner) and assert the rejected key is never persisted to the cached file. - trace_commons_dispatch_e2e.rs: give each test a distinct user/extension scope so onboarding state can no longer bleed across tests. - local_dev_capability_policy.toml: exempt builtin.trace_commons.onboard from the REPL approval gate (it has its own confirmed=true consent gate, mirroring profile_set). - docs: fix the onboard prompt-file reference, match the held-trace JSON shape to RebornTraceHold (submission_id + reason only), and resolve the wire-protocol ownership split (types live locally in onboarding/protocol.rs, no shared trace-commons-protocol crate). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(traces): tenant-scoping + token leak + read-failure + unbounded scopes (coderabbit batch 2) Addresses the coupled backend findings: - Tenant-scope Trace Commons local state across the Reborn paths: new trace_scope_key(tenant, user) helper keys policy / device-key / credit / profile / capture state by tenant+user, so the same user id in two tenants no longer shares state. Applied in host_runtime trace_commons dispatchers, product_workflow credits/hold, and composition trace-capture (v1 stays user-only — legacy single-tenant). Updated the affected runtime/sink tests and added a non-owner attribution assertion. - Do not return the raw profile token from the model-visible profile_token capability: persist it to a 0600 <scope>/profile_token.jwt and return the file path + instructions instead, keeping the bearer credential off the LLM transcript. - Stop masking genuine local-state read failures as zero/not-enrolled: the status capability and the WebUI credits path now propagate a read/parse failure (NotFound is already softened inside read_*_for_scope) so an enrolled user with a corrupt policy file is not told they have nothing. - Bound ObservedTraceScopes: the periodic flush worker now prunes drained scopes (new trace_scope_has_pending_queue) after each tick, so the set is bounded by actual pending backlog instead of growing one entry per caller ever seen. Note: a v1 caller-level test for the ManualReview hold-retention path is not included — v1 ingress blocks secrets outright and the outbound leak detector redacts them, so the High-residual-PII condition that produces a ManualReview hold cannot be reproduced through process_user_input. The retention logic is identical to and covered by the Reborn-side capture_retains_manual_review_hold_for_high_pii_trace. * test(traces): enroll under tenant-scoped key in webui_v2_serve credits test trace_credits_reports_enrolled_for_caller_with_enabled_policy wrote the policy under the bare user id, but the credits route now keys local state by trace_scope_key(tenant, user). Enroll (and clean up) under the composite TENANT/user scope so the route sees the enrollment. * fix(factory): fail closed on explicit-but-unusable SECRETS_MASTER_KEY An explicitly-set-but-unusable local-dev master key silently fell through to generating + persisting a fresh key, leaving local-dev secrets encrypted under an unintended master key the operator never chose: - resolve_local_dev_secret_master_key used std::env::var(...).ok(), which drops VarError::NotUnicode -> treated as absent. Now only NotPresent is absent; a non-Unicode value returns InvalidConfig. - resolve_local_dev_secret_master_key_with_env collapsed a set-but-empty (or whitespace-only) value to None via .filter(). Now a set-but-empty value returns InvalidConfig instead of generating a key. Added resolve_local_dev_secret_master_key_rejects_set_but_empty_env_without_persisting asserting empty/whitespace env values fail closed and persist nothing. (coderabbit follow-up on nearai#3794) * fix(factory): reject empty SECRETS_MASTER_KEY before the cached-file read Follow-up to the prior fix: the empty-env rejection lived in the env branch, which only runs when no cached key file exists. On a rebuild where .reborn-local-dev-secrets-master-key already exists, the cached key was returned first, so an explicitly-set-but-empty SECRETS_MASTER_KEY was still silently ignored. Hoist the empty/whitespace rejection (and env normalization) above the cached-file read so it fails closed regardless of cached state. Added resolve_local_dev_secret_master_key_rejects_empty_env_even_with_cached_file asserting the empty env is rejected and the cached key is left unchanged. * fix(traces): address 14:54 coderabbit re-review (tenant-seed, IO errors, effects, test) Four outside-diff findings from the re-review: - runtime.rs: seed ObservedTraceScopes with the runtime owner's trace_scope_key(tenant, owner) composite, not the bare owner id, so startup pending-queue discovery matches how capture keys state; the enrolled-scope test cleanup now removes the composite scope dir too. - runtime.rs: the trace-queue polling test helper no longer swallows read_dir errors via unwrap_or_default() — only NotFound is the expected pre-capture fallback; any other IO error panics instead of masking as 'no queued traces'. - trace_commons.rs manifests + local_dev grants: onboard (device-key material) and profile_token (0600 token file) now declare Read/WriteFilesystem effects, and the local-dev grants allow them, so the effect model accurately models the local secret-material writes. - local_dev_authorization test: added local_dev_trace_commons_onboard_skips_approval_gate (the onboard exemption was the actual fix; the profile_set-only test would pass even if the onboard TOML exemption were dropped). * fix(factory): validate non-empty SECRETS_MASTER_KEY before the cached-file read Follow-up: the prior fix rejected an *empty* env value before the cached read but still validated a non-empty *malformed* value only after it. So a valid cache + SECRETS_MASTER_KEY=0000... silently ignored the explicit bad secret config on rebuilds. Move validate_resolved_master_key into the up-front env normalization so any explicit-but-unusable env key (empty OR malformed) fails closed regardless of cached state. Added resolve_local_dev_secret_master_key_rejects_malformed_env_even_with_cached_file. * fix(traces): address 15:41 coderabbit re-review (credits read-failure + 2 test guards) - trace_commons.rs dispatch_credits: stop masking genuine records read/parse failures as 'no records' (NotFound is already softened inside read_local_trace_records_for_scope); report RecordsReadFailed, mirroring dispatch_status. - runtime.rs trace-queue polling helper: fail loud on per-ENTRY read_dir IO errors too (map + unwrap_or_else panic) instead of filter_map(e.ok()), so a broken entry can't be silently dropped while claiming the queue holds one. - local_dev_authorization approval-gate test: assert the effects DO require approval without the exemption (local_dev_effects_require_approval), so the test can't pass via a non-gating default policy if the TOML exemption were dropped. * fix(traces): address Henri review — backend findings (atomic token, error mapping, validation, egress test) - persist_profile_token now writes atomically (unique 0600 temp + fsync + rename) so a reader never observes a half-written or overwritten bearer credential under overlapping mints (Henri perf/security Medium). - dispatch_onboard error mapping: OnboardError::DeviceKey is reported as a distinct DeviceKeyError (re-run onboarding) instead of being collapsed into PersistError's check-disk-and-permissions guidance (Henri bugs Medium). - parse_profile_set_input enforces the manifest's declared schema at parse time: handle 3-32 ASCII letters/digits/-/_, bio <= 280 bytes (Henri conventions Medium). Added schema-limit test. - Added dispatch_onboard_confirmed_without_host_egress_is_network_denied covering the NetworkDenied host-egress-miswiring branch (Henri tests Medium). * fix(traces): address Henri review — frontend findings (enrolled empty-state + polling) - v1 credits: TraceCreditResponse now carries `enrolled` (read from the standing policy), and settings.js keys the opt-in empty state on `!data.enrolled` instead of `!submissions_total` — an enrolled user with zero submissions now sees their zero-credit view, not the not-enrolled prompt (Henri bugs Medium). - useTraceCredits: each fetch rebuilds the full server-side credit view, so the aggressive 60s poll made an open tab steady O(history) work. Relaxed to a 5-min interval + staleTime + no background polling, keeping a focus refetch for liveness; mutation invalidation still updates promptly. Added a TODO to incrementalize the server-side view (Henri perf Medium). * perf(traces): memoize server-side credit view by on-disk input signature Bounds the trace-credits polling cost to O(new submissions) instead of O(total history). New scoped_credit_view(scope) caches the computed credit report + manual-review holds keyed by a cheap change signature (submissions file mtime+len, plus a hash of the held-trace sidecars). On the steady-state polling case (unchanged history) a request is a couple of stat()s + a clone rather than reading/parsing the full submissions file and re-aggregating. On any change the signature differs and it recomputes once. Cache is bounded (4096 scopes, cleared on overflow). Wired through the polled WebUI path (local_trace_credits_for_user) and the model-visible credits capability (dispatch_credits). Added scoped_credit_view_reflects_record_changes_via_signature covering the cache-hit path and signature-based invalidation on record changes. Completes the TODO from the Henri perf-review follow-up (#5). * fix(traces): gate profile_set behind runtime approval (Henri #1 High) profile_set publishes a public community profile (an external write to a public surface). Its `confirmed=true` input is model-controlled, so a prompt-injected or confused model could supply it. Make the runtime approval gate the primary, user-controlled consent control: - Drop `builtin.trace_commons.profile_set` from the local-dev approval-gate exemption list (keep `onboard`, which runs its own in-turn confirmed=true consent before the network POST). - Set profile_set's manifest default_permission to Ask (was Allow). - Split the local-dev authorization test into `local_dev_trace_commons_profile_set_requires_approval_gate` (asserts Decision::RequireApproval) and `local_dev_trace_commons_onboard_skips_approval_gate` (asserts Decision::Allow), via a shared `trace_commons_authorize_decision` helper that first asserts the effects would gate without an exemption. Also fix a pre-existing trace_commons harness gap: onboard + profile_token gained a WriteFilesystem effect (device-key persistence) but the `trace_commons_tools` harness allow-set was never updated, so those capabilities were filtered out of the model-visible surface and the parity/visibility tests failed with driver_unavailable. Grant WriteFilesystem in the harness allow-set. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(traces): extract onboarding test harness to sibling file (Henri nearai#8) The onboarding module's ~840-line `#[cfg(test)] mod tests` block (mock issuer harness, retry/idempotency coverage, URL-validation tests) made `onboarding/mod.rs` a 1319-line file dominated by test scaffolding. Move the module body into `onboarding/tests.rs` declared `#[cfg(test)] mod tests;`, leaving mod.rs focused on production logic (now 480 lines). No test behavior changes; `use super::*;` still resolves to the onboarding module. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(webui-v2): update embedded-asset assertion for incrementalized credits poll The Henri #5 polling fix changed useTraceCredits.js from refetchInterval 60_000 to 300_000 (plus refetchIntervalInBackground: false and staleTime: 60_000), but the embedded-asset test in assets.rs still asserted the old 60_000 value and failed in CI. Update the assertion to lock the new infrequent-poll + paused-while-hidden + focus-refetch shape. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): address CodeRabbit review + stale capability-policy test CodeRabbit findings on the gating/refactor commits: - Major: format_profile_token returned the absolute host path of the token file (token_file) on the model-visible surface, which violates the "never expose absolute paths" guideline. Replace with an opaque token_delivery marker; the token is still persisted 0600 for out-of-band retrieval by a bearer-auth UI/CLI. Update the message + test accordingly. - Major (fail-loud): profile_token_error_value and profile_set_error_value collapsed "could not read policy" into NotEnrolled, sending enrolled users back through onboarding on unreadable/corrupt state. Split into a distinct PolicyReadFailed result in both formatters (matches dispatch_status). - Minor: stale comment claiming profile_set is approval-gate-exempt (it is now PermissionMode::Ask and NOT exempt) — corrected. - Minor: inaccurate harness comments (profile_token writes profile_token.jwt not device-key material; yolo auto-approves all Trace Commons Ask-gated tools, not just onboard) — corrected. Also fix bundled_local_dev_capability_policy_parses, which still asserted the pre-gating policy shape: profile_set as exempt (now onboard exempt / profile_set NOT exempt), onboard's grant missing the read/write filesystem effects, and profile_token/profile_set sharing one effect-set assertion even though profile_token now carries WriteFilesystem and profile_set does not. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * style(traces): collapse single-line use block after Path import removal rustfmt collapses `use std::{panic, path::PathBuf, sync::Arc}` to one line once Path was dropped; the prior commit skipped re-running fmt after that edit, reddening the Formatting CI check. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): consent-gate profile_token + drop fixed-origin profile URL (CodeRabbit) Two Major CodeRabbit security findings on the profile tools: - profile_token minted and persisted a bearer credential with no in-turn consent gate. PermissionMode::Ask can be auto-approved under local-yolo, so a model call could mint a credential without explicit per-conversation consent. Add a hard confirmed=true gate (schema + parse + consent_required short-circuit) before minting, mirroring dispatch_onboard / dispatch_profile_set. - format_profile_token and profile_set_success_value hardcoded https://tracecommons.ai/profile. The token is scoped to the user's ENROLLED issuer (which may be self-hosted or loopback), so steering the user to paste a bearer profile-management token at a fixed origin could leak it to the wrong host. Drop the fixed profile_url; route through the enrolled profile flow / local UI/CLI out of band. Tests: new dispatch_profile_token_without_confirmed_returns_consent_required_no_mint; existing without-enrollment test now passes confirmed=true; profile_set success test asserts no fixed origin; parity step mints with confirmed=true. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): route agent-invoked profile writes through host egress (CodeRabbit #3) profile_token (upload-claim mint) and profile_set (community-profile PUT/DELETE) previously made network writes via the ironclaw_reborn_traces crate-local reqwest client, bypassing the host RuntimeHttpEgress pipeline (private-IP filtering, redaction, byte accounting) that onboard already uses. Add a `ContributionHttpSink` port (mirroring `OnboardingHttpSink`): when a sink is injected, the mint POST and the profile PUT/DELETE run through host egress; when `None`, the existing hardened crate-local client is used unchanged. host_runtime supplies `HostEgressContributionSink` (wraps RuntimeHttpEgress, sanitizes errors via stable_runtime_reason, never leaks URL/token), and dispatch_profile_token / dispatch_profile_set fail closed with NetworkDenied if egress is absent (after the enrollment pre-check, so a not-enrolled user still gets NotEnrolled guidance). The background trace-upload / status-sync worker and the CLI keep the crate-local client (pass `None`): that lane is a durable, model-input-free internal task that sends only already-redacted envelopes to the operator-enrolled endpoint and does its own SSRF/private-IP validation, so host egress adds complexity without security benefit. Justification recorded in a comment on `trace_remote_http_client`. New public surface: ContributionHttpSink/Request/Response/Error/Method, mint_profile_attribution_token_for_scope_via_sink, set_community_profile_for_scope_via_sink. Existing public fns keep their signatures (None path) so CLI/worker/tests are unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jun 25, 2026
…ols (nearai#5061) * feat(reborn): skill-learning turn-end seam + extraction prompt Add the post-completion seam for learning reusable skills from successful runs, composed additively alongside trace capture (no behavior change to existing paths): - SkillLearningTurnEventSink: on a successful turn completion, reads the run transcript (load_context_window, preserving tool calls) and gates substantive runs (>=3 tool actions, >=5 messages) as skill-extraction candidates. Modeled on trace_capture.rs; detached, debug!-only. - CompositeTurnEventSink: fans the single turn_event_sink slot out to both trace capture and skill learning. - assets/prompts/skill_extraction.md: one-shot transcript -> SKILL.md prompt for the next increment (the distillation LLM call). - docs/plans: design + implementation log. Distillation, staging-for-approval, and the scoped skill write land in follow-up increments. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(skill-learning): distillation logic crate (transcript -> SKILL.md) New leaf crate ironclaw_skill_learning owns the pure skill-learning logic, kept out of the composition root (per architecture guardrails) and reusable by both the autonomous sink and a future explicit CLI command: - distill_skill(transcript, &dyn SkillInferencePort) -> DistillOutcome: runs the extraction prompt through an abstracted inference port, then validates the output with ironclaw_skills::parse_skill_md (the SAME parser the install path uses) so a distilled skill is guaranteed installable. - parse_distillation: tolerates SKIP declines and accidental code-fence wraps; rejects chatty/invalid output. Inference is abstracted behind SkillInferencePort so the crate has no LLM/runtime/filesystem dependency. - Moves the extraction prompt here (co-located with the parser contract it must satisfy). 6 unit tests, clippy clean. Composition wiring (inference adapter over the runtime's non-run inference port + scoped write) lands next. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(skill-learning): wire distillation into the turn-end sink On a successful, substantive run, the skill-learning sink now actually distills a SKILL.md (instead of just logging a candidate): - SkillLearningInferenceAdapter bridges a strong-model LlmProvider to the logic crate's SkillInferencePort, passing the learning model as a per-request override (NEAR AI honours it) — so distillation runs against a STRONGER model than the run's, without touching the run's model gateway. - build_skill_learning_provider builds that provider from the run's resolved NEAR config with only the model overridden (IRONCLAW_SKILL_LEARNING_MODEL), reusing existing credentials. No churn to build_llm_gateway / the gateway return tuple. - The sink formats the run transcript (tool names included) and calls distill_skill, logging the distilled skill / skip / error. The scoped write + stage-for-approval land in the next increment. - Skill learning is gated on root-llm-provider (it needs an LLM) and is active only when the learning model is configured; otherwise only trace capture runs. CompositeTurnEventSink fans the single turn_event_sink slot to both. Verified: check (default + root-llm-provider) 0 warnings; test + clippy (root-llm-provider,test-support,libsql) green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(skill-learning): install distilled skills (scoped write + safety scan) The skill-learning sink now persists the distilled skill so it appears in Settings->Skills and loads into the next run (per the user's "scan + visible" choice; the pre-approval gate is the next increment): - SkillWriter seam (composition trait): the sink depends on a small write abstraction; PortSkillWriter implements it over the runtime's existing RebornLocalSkillManagementPort (install_for_scope, falling back to update_for_scope on re-learn). Tests use a stub writer (no filesystem). - Scope is derived from the EVENT: ResourceScope::local_default(owner, ...) with tenant_id overridden to the run's tenant, so the write lands where the WebUI lists it and the next run reads it (NOT the `default` tenant). - Distilled content is injection-scanned (ironclaw_safety:: validate_trusted_trigger_prompt with a Sanitizer, mirroring the WebUI facade) before install — it becomes trusted prompt text in the next run. - Sink wiring now also requires local_runtime (the skill port lives there); reuses local_runtime.skill_management rather than building a new port. Verified: check (default + root-llm-provider) 0 warnings; test + clippy (root-llm-provider,test-support,libsql) green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(skill-learning): live "learned a skill" bubble on WebChat v2 When a skill is distilled + installed, the sink now emits a live notification to the run's thread stream, rendered by the EXISTING WebChat v2 chat bubble (reuses the SkillActivation projection — zero new wire variants): - LiveProjectionPublisher::publish_skill_learned: publishes a SkillActivation live item from raw pieces (owner, turn scope, run_id, name, feedback), the post-run analogue of the in-run SkillActivationObserver (which only fires at prompt-build for skill selection). Gated on root-llm-provider. - SkillLearnedNotifier seam (same testable pattern as SkillWriter): LiveSkillLearnedNotifier wraps the publisher; the sink emits the bubble after a successful install. Tests use a stub notifier. - runtime wiring clones the live projection publisher before the milestone-sink builder consumes it, and passes a notifier into the sink. The learned skill already appears in the existing Settings->Skills page (installed live in the prior increment); this adds the in-chat moment. Pre-approval gate (decision #3) is the next increment. Verified: check (default + root-llm-provider) 0 warnings; test + clippy (root-llm-provider,test-support,libsql) green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(skill-learning): refresh implementation log; rename refinement (drop GEPA) - Phase 2 renamed to "Skill Refinement (eval-driven reflective improvement)"; removed the "GEPA-lite" name (DSPy/Hermes term) per review. - Implementation log updated to reflect increments 2 (logic crate), 2b (sink wiring + the SystemInferencePort rejection), 3 (scoped install + scan), and 4 (live learned-skill bubble), plus the per-increment verification gate. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(skill-learning): e2e fixes from a live ironclaw-reborn run Validated the whole loop end-to-end against a running `ironclaw-reborn serve` with a NEAR AI `openai/gpt-5.5` learning model: a completed multi-tool run was distilled into a real SKILL.md (with pitfalls captured from the transcript), injection-scanned, installed under the correct (tenant=reborn-cli, user) scope, and shown in Settings->Skills. Three real bugs the run surfaced: - Drop the temperature override: reasoning models (gpt-5.x) reject any non-default temperature with HTTP 400 ("temperature does not support 0.2"). - Bump the distillation output ceiling to 16384: a reasoning learning model spends tokens on reasoning before emitting the SKILL.md, so a 4096 cap would truncate it. - Lower the eligibility gate to >=2 tool actions / >=3 messages: an efficient agent can complete a skill-worthy multi-step task in two tool calls (e.g. `shell` mkdir + batch write). The gate is only a cheap pre-filter; the learning model's own SKIP judgement is the real quality gate. - Loosen the extraction prompt: distill any multi-step tool procedure (capture the general repeatable procedure); only skip purely conversational runs. Also removes the temporary info-level diagnostics added while debugging. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(skill-learning): auto-activate learned skills on criteria match Local-dev composition hard-coded the skill selection mode to `ExplicitOnly`, so a learned skill only activated when the user typed `$name`/`/name`. That left the learn loop half-open: skills were distilled and installed but never reused unless named explicitly. Switch local-dev to `ExplicitAndCriteria` (the upstream default) so a learned skill auto-activates when a later request matches its keywords/patterns, closing the learn→reuse loop. Explicit mentions still force-activate; criteria selection is additive and bounded by `max_active_skills` / `max_context_tokens`. The selector-config unit test now locks `ExplicitAndCriteria` so a revert to explicit-only trips a clear failure. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(skill-learning): durable learned-skill feedback + dedup consolidation Two gaps in the learn loop, both surfaced while dogfooding against a live ironclaw-reborn run: 1. No visible "learned a skill" feedback. The post-run sink published a live `SkillActivation` projection bubble, but that is ephemeral — only delivered to a stream connected at publish time, ~seconds after the run when distillation finishes. Add a DURABLE path: after install, append a finalized assistant note to the run's thread, so the feedback renders from `get_timeline` and survives a reload even when no live stream was open. The spawned extraction body is lifted into `ExtractionJob::run` so the durable announce is testable end-to-end through its caller (the spawn is otherwise fire-and-forget). Two regression tests also lock that the live `SkillActivation` bubble drains to the WebUI projection stream (fresh and resume-from-advanced-cursor paths). 2. Near-duplicate skills accreted. The distiller names the same kind of task slightly differently each run, so the user's skill list filled with siblings (file-create-read-count-summary, file-character-count-roundtrip, create-read-count-file-characters …) that never get reused together. Before installing, `PortSkillWriter` now lists existing learned skills and, when one covers the same ground (Jaccard over the combined name/keyword/tag token sets ≥ 0.45), refines it in place under its existing name instead of installing a second one. Only `User`-source skills are merge targets; system/registry skills are never touched. `update_skill` requires the document name to match the target, so the merged content is retargeted first. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(skill-learning): self-evolving skill refinement on recurring tasks Builds on near-duplicate consolidation: when a learned task recurs and the freshly distilled candidate matches an existing learned skill, the existing skill is now *refined* in place rather than overwritten — the self-evolution step. The learning model folds the candidate's new evidence into the existing SKILL.md (converged steps, the UNION of real gotchas, a bumped version), so a skill gets strictly better each time its task comes around. - `ironclaw_skill_learning::refine_skill` + `parse_refinement` + `RefineOutcome`, driven by `prompts/skill_refinement.md`. Pure domain logic, validated by the install-path parser; tolerates a `KEEP` decline (existing already subsumes the candidate) and a code-fence wrap, same as distillation. - Composition `SkillRefiner`/`LlmSkillRefiner` seam: maps the model outcome to a `MergeAction` — `Replace` (refined, retargeted to the existing name, and injection-scanned), `KeepExisting` (leave the existing skill untouched), or `Overwrite` (fall back to plain consolidation when refinement is unavailable or the model output is unusable). The refined document is retargeted defensively (never trust the model to preserve the name) and re-scanned before install. - `PortSkillWriter` reads the existing skill and consults the refiner on the merge path; wired in `runtime.rs` from the same learning inference adapter. Unit-tested end to end through the refiner (replace+bump, model-rename retarget, keep, unparseable→overwrite) and in the logic crate (parse/keep/reject). The prompt's merge quality is verified live against the NEAR AI learning model. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(skill-evolution): log increments 5-8 (durable feedback, auto-consume, dedup, refinement) Records tonight's work and the one known gap: the live SkillActivation bubble is published but not delivered in the running server (empirically confirmed), its mechanism passes deterministic tests, and it could not be instrumented live without the NEAR AI key — so a durable timeline note is the reliable fix shipped instead. Carries forward the remaining work: pinning the live-SSE gap, an eval-driven refinement loop, the pre-approval gate, CLI commands, and a one-off consolidation of the siblings already on disk. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(skill-learning): live-validation fixes for durable feedback + refinement Found by re-running the loop end to end against a live ironclaw-reborn with the NEAR AI learning model (the in-memory fakes missed both): 1. Durable note never persisted. The durable store dedups assistant drafts by `turn_run_id` and returns the existing one, so `announce_learned_skill` reusing the run's id handed back the run's already-finalized reply and the finalize failed `MessageNotDraft` ("message … is not an assistant draft"). Use a distinct `skill-learned:{run_id}` id so the note is its own message. The regression test now seeds the run's finalized reply first (reproducing the collision the fresh-thread test missed) and asserts the note is a separate, finalized message. 2. Re-learning the SAME skill name overwrote the refined version instead of refining it. The distiller derives the name from the task, so it often repeats; the old path skipped the similarity check for the same name and fell to a plain install→update-on-conflict, resetting an evolved v2 back to a fresh v1. `find_merge_target` now routes BOTH an exact-name re-learn and a renamed sibling through refinement, so the version climbs consistently (verified live: create-read-count-file-characters v1 -> v2, "refined existing learned skill", skill count held at 3, durable note rendered, zero finalize errors). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(skill-evolution): record live end-to-end validation + the two fixes it found Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(skills): per-skill auto-activation flag honored by the selector Foundation for user-facing skill activation control. Adds a manifest `auto_activate` flag (frontmatter, defaults true so existing skills are unaffected) and has the activation selector honor it: a skill with `auto_activate: false` is excluded from criteria (keyword/regex) selection but stays available for an explicit `$name` / `/name` mention. State lives in the skill's own SKILL.md — no new storage layer. - `SkillManifest.auto_activate` (`#[serde(default = "default_auto_activate")]`). - `set_skill_auto_activate(content, enabled)`: line-edits the frontmatter flag, preserving the rest of the document byte-for-byte so a toggle does not reformat the skill (re-parses cleanly with the same name). Unit-tested (default-true, insert-then-replace). - `select_skill_activations` builds a criteria-candidate set filtered by `auto_activate`; explicit mentions still resolve against the full set. Existing skill construction sites updated for the new field; full workspace + reborn binary build green (266 crate tests pass). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(skills): API + DTO to toggle a skill's auto-activation Wires the per-skill auto-activation flag end to end on the backend so the WebChat v2 UI can flip it: - `POST /api/webchat/v2/skills/{name}/auto-activate` ({ enabled }) — reads the skill, line-edits the frontmatter flag via `set_skill_auto_activate`, re-scans it for injection (parity with install/update), and persists. Added as a default method on `SkillsProductFacade` / `RebornServicesApi` (fail-closed unavailable) with the real implementation in the composition facade, plus the handler, descriptor, route, and exports. Descriptor contract test updated. - `SkillSummary.auto_activate` + `RebornSkillInfo.auto_activate` so the skills list reports each skill's current state for the UI toggle (defaults true). Full backend chain builds (reborn binary green); ironclaw_skills, extension ports, webui_v2, and product_workflow test suites pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(webui): per-skill auto-activation toggle in Settings → Skills Adds an "Auto-activate: On/Off" switch to each manageable skill card. Off makes the skill explicit-only (`/name`); on restores keyword/criteria auto-activation. Wires `setSkillAutoActivate` through settings-api → useSkills mutation (invalidates the skills query) → SkillsTab handler → SkillGroup → SkillCard, reading the `auto_activate` field the v2 skills DTO now reports. Mirrors the existing skill install/update/remove mutation pattern. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(skills): global "auto-activate learned skills" master switch (live) Add a global toggle that disables default auto-activation while keeping explicit /name invocation. ON (default) selects ExplicitAndCriteria; OFF selects ExplicitOnly. It takes effect on the next turn with no restart, via one process-global Arc<AtomicBool> shared by reference between the activation selector (reads it every turn in select_skill_activations) and the WebUI skills facade (writes it). Not persisted by design — resets to ON on restart. Vertical: - factory.rs: RebornLocalRuntimeServices.skill_auto_activate_learned (default true), one instance shared with the selector and the facade. - activation.rs/skills.rs: thread the flag into SelectableSkillContextSource; gate the criteria branch on it. Explicit mentions always activate. - webui.rs: LocalSkillsProductFacade holds Option<Arc<AtomicBool>>; set_auto_activate_learned stores into it, list_skills surfaces it. When no flag-reading selector is wired (production assembly) the facade gets None and the toggle fails closed (503) instead of writing to an orphan flag — fixes a review finding where the production toggle silently no-oped and read back true. - reborn_services.rs/types.rs: facade + API trait method, delegation, and RebornSkillListResponse.auto_activate_learned DTO field (serde default true). - webui_v2: POST /api/webchat/v2/skills/auto-activate-learned route, handler, descriptor + contract row. - frontend: setAutoActivateLearned API, useSkills mutation, Settings → Skills LearnedAutoActivateCard master switch. Tests (regression): - global_auto_activate_flag_gates_criteria_and_honors_live_toggle: drives the real selector with a live flag flip (off → empty, flip on → activates). - set_auto_activate_learned_flips_shared_flag_and_surfaces_in_list. - set_auto_activate_learned_fails_closed_when_no_selector_is_wired. - set_auto_activate_learned_forwards_enabled_flag_to_facade (through the caller). - descriptor contract row. Live-validated against ironclaw-reborn serve: GET skills auto_activate_learned True → toggle OFF → False → toggle ON → True. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(skill-evolution): record increment 9 — global auto-activate master switch Document the live global toggle (shared Arc<AtomicBool>, ExplicitAndCriteria ⇄ ExplicitOnly, not persisted), the production orphan-flag review finding and its fail-closed fix, the regression tests, and the live end-to-end validation. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(skill-learning): scope extraction eligibility to the completed run The post-turn ExtractionJob loads the recent THREAD window (no run filter) and the eligibility gate counted tool-result messages across that whole window. A trivial follow-up turn after a tool-heavy task could re-pass the gate on the previous run's stale tool results and re-distill it — wasted inference plus a stale-transcript refine that can regress an evolved skill. Count tool actions only for the completed run, read from the history projection (which keeps message kind + turn_run_id and only nulls the tool metadata the transcript needs). The full window is still used as the multi-turn distillation context, which is intentional. The producer writes turn_run_id = run_id.to_string(), matching self.run_id. Localized to skill_learning.rs — no change to the shared ContextMessage / agent-loop model-context path. Regression test: eligibility_counts_tool_actions_for_the_completed_run_only (trivial follow-up under a fresh run id over stale prior-run tool results does not distill; a run with its own tool actions reaches distillation). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(webui): make the skill auto-activation switch read as the global control it is The master switch gates the entire criteria-selection pass, so it affects every skill (learned, user-authored, and bundled), not only learned ones — but the Settings card said "Auto-activate learned skills". Rename the user-facing card to "Default skill auto-activation" with global wording (frontend strings only; behavior and the wire field are unchanged). Also give the card a light-red background and a black status line when disabled, as a persistent "default is off" cue. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(skill-evolution): increment 10 — review-driven hardening Record the three review findings and their disposition: run-scoped extraction eligibility (fixed), the global master-switch relabel (fixed), and the deferred learned-skill prompt-injection approval gate with its residual risk and the reason the obvious low-risk mitigations don't apply (auto_activate=false is filtered out of criteria selection; trust attenuation needs a dedicated learned-skill source/dir first). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Fix skill manifest test fixture * Avoid bundled skill collision in runtime test * Avoid review keyword in filesystem skill test * Allow auto-activated skills in runtime asset test * fix(skill-learning): guard the two data-loss paths in learned-skill writes merge() now keeps the existing accumulated skill (KeepExisting) on a refiner error, an unparseable response, or a rejected injection scan of the merged doc, instead of overwriting it with the raw single-run candidate — a transient model hiccup must not discard a skill that accreted gotchas over many runs. Overwrite is reserved for the genuine no-existing-content case. install_or_update now matches SkillManagementErrorKind::Conflict specifically and fails loud on any other install error (filesystem/validation/resource), instead of treating every install failure as a name conflict and overwriting a live skill. Addresses review #1/#2 (data-loss/overwrite paths). Self-learning stays off by default (sink wired only with IRONCLAW_SKILL_LEARNING_MODEL + nearai), so these paths are unreachable in a default deployment; the broader hold-for-review / approval hardening lands in the stacked follow-up PR. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Address review feedback on skill activation --------- Co-authored-by: krishna <krishna@krishnadeMacBook-Air.local> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Robert Yan <mstr.raphael@gmail.com>
JZKK720
pushed a commit
that referenced
this pull request
Jul 18, 2026
…rai#4841 + nearai#5389/nearai#5390/nearai#5403/nearai#5613) (nearai#5692) * reborn: add failure explanations and retryable failed runs * test(loop_support): set inline_messages on the contract-test LoopModelRequest nearai#4841 added the `inline_messages` field (serde default) to LoopModelRequest but missed this one construction in the thread_loop_support_contract integration test, breaking that test target's compile. Production builds default it to Vec::new(); match that. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(reborn): resolve main-merge CI breaks on nearai#4841 (retry_turn stub + finer failure categories) The main→nearai#4841 merge surfaced two semantic conflicts the auto-merge missed: - Clippy: main added `TurnCoordinator::retry_turn`; the StaticTurnCoordinator test stub in openai_compat_serve/tests.rs didn't implement it. Add the stub (returns Unavailable, matching its other methods). - Test ironclaw_reborn: main's chaos tests (nearai#5296) assert the coarse failure categories "driver_unavailable"/"model_error", but nearai#4841 refined production to finer, accurate categories — a full checkpoint-state disk now yields "host_stage_unavailable_checkpoint" and an offline model provider yields "model_unavailable" (both deliberate named categories with dedicated failure_summary messages). Update the two stale assertions to match nearai#4841's intended categorization. Verified: the two turn_runner_worker_full_reborn_fails_* tests pass; clippy ironclaw_reborn_composition --all-features clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Persist retry busy idempotency records * test(reborn): expect specific model unavailable failure * fix(reborn): preserve same-run checkpoint refs * fix(ci): update reborn compile drift * fix(ci): update composition test drift * feat(reborn): make model-fixable capability failures recoverable (batch 1) Turns recoverable→bork mis-mappings into model-visible tool errors so the agent self-corrects instead of the run dying. On top of nearai#4841. - agent_loop keystone: capability_error_class re-buckets Dispatcher / InvalidOutput / Unknown(_) / non-exhaustive default from Permanent (Abort) to OperationFailed (ToolErrorResult), aligning with the host_runtime disposition layer (which never intends a capability failure to abort). Cancelled / Permanent stay terminal. This makes "model called a nonexistent tool" (UnknownCapability/UnknownProvider -> InvalidOutput) recoverable. - outbound_delivery: outbound_delivery_outcome routes recoverable RebornServicesErrorCode to Ok(Failed/Denied) (only Internal -> Err); fixed the safe_summary that interpolated the model-supplied target_id (Invariant 2); expired/not-yet-approved approval-lease arms -> Ok(Denied) instead of terminal. - host_runtime: malformed model-supplied SandboxProcessPlan -> recoverable Failed{InvalidInput} outcome (defense-in-depth; the live gate in loop_support's host_runtime_input_for_capability is fixed in a follow-up). Lib tests green: agent_loop 354, host_runtime 297, loop_support 337, reborn 259; outbound_delivery 26. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(loop_support): malformed sandbox plan is recoverable, not run-ending Completes the sandbox-plan fix. The live terminal gate is loop_support's host_runtime_input_for_capability: a malformed/invalid model-supplied SandboxProcessPlan returned AgentLoopHostError::InvalidInvocation, which capability_host_error maps to terminal HostUnavailable{Capability} (run dies). Now the invoke path downgrades that InvalidInvocation to a model-visible Ok(CapabilityOutcome::Failed{InvalidInput}) so the agent can correct the plan and the run continues. The helper only emits InvalidInvocation for the sandbox-plan parse/validation case; its host-internal serialization failure keeps its Internal Err. Updated both locked tests to assert the recoverable outcome. loop_support lib: 337 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(llm): provider error fidelity for accurate recover/explain (batch 2) Maps provider failures to the right LlmError variant so model-call errors are explained accurately and context overflow recovers via context-shrink instead of borking. - rig_adapter (OpenAI/Anthropic/Ollama/Tinfoil/openai_compatible): map_rig_error now detects auth failures (401/403/invalid key) -> AuthFailed (non-retryable, non-breaker-tripping) instead of generic RequestFailed, so a bad key surfaces as a credentials problem rather than wasted retries + opaque run-bork. - Codex (openai_codex_provider + codex_chatgpt): a stream ending without response.completed is now a retryable InvalidResponse/EmptyResponse instead of a silent successful Stop; codex_chatgpt now maps SSE error/response.failed events; both detect 413/context-overflow -> ContextLengthExceeded. - github_copilot + anthropic_oauth: detect 413 (and 400+context body) -> ContextLengthExceeded so context-shrink recovery fires (401/429/5xx untouched). ironclaw_llm lib: 915 passed, 0 failed. clippy + fmt clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(host_runtime): unknown method/capability is immediate model-visible error (batch 3) Sweep finding: a method/capability the model named that does not exist (RuntimeDispatchErrorKind::MethodMissing / UndeclaredCapability) mapped to RuntimeFailureKind::Backend -> RetrySameCall, so it burned the retry budget before becoming model-visible. Retrying never resolves a nonexistent target. Now maps to InvalidInput -> ModelVisibleToolError: the model gets an immediate "no such method/capability" tool error and self-corrects. Updated the pinning table entries. host_runtime lib: 297 passed. Batch-3 sweep conclusion: after the keystone + batches 1-2, no remaining recoverable->bork CORRECTNESS defects exist (tool backends fully clean; no hard-Err bypass on model-fixable conditions; nothing mapped to terminal Cancelled/Permanent). This was the last shape-#3 quality nit worth fixing now. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(reborn): FailureLane classifier + two-bucket enforcement test Item #2 foundation, built ON TOP of nearai#4841's failure-surfacing machinery (reuses category + FailureExplanationProvider + retryable rather than a parallel RunFailureReason taxonomy). - FailureLane enum (Retriable | Explainable | Security), wire-stable snake_case. - failure_lane(category, retryable): retryable -> Retriable, else Explainable. Security is reserved for the ingress safety/leak refusal path (minimal security-stop policy) and is never produced at the run boundary; the match on category is the seam for a future mid-run safety-abort category. - ALL_RUN_FAILURE_CATEGORIES: canonical list of every category the run boundary can produce. - ENFORCEMENT TEST (every_failure_category_is_explainable_and_classified): locks the two-bucket invariant — every failure category resolves to a SPECIFIC user explanation (never the generic fallback) AND a definite lane. A new category that forgets its sentence, or regresses to the generic fallback, fails here. Plus canonical_list_covers_loop_failure_kinds guards against list drift. reborn_composition failure_lane: 5 passed. clippy clean (the one pre-existing needless_return in local_runtime_profile.rs is unrelated). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(reborn): retry-disposition policy (hybrid retry core) Operationalizes the hybrid retry decision on top of the FailureLane classifier. Pure decision function; the auto-redrive scheduler is its consumer. - RetryDisposition { Auto | UserInitiated | NoRetry }, wire-stable snake_case. - retry_disposition(category, retryable): no checkpoint -> NoRetry; transient host/lease/store/provider/tool faults -> Auto (silent re-drive from checkpoint, bounded by the scheduler); model/provider/config/model-fixable faults -> UserInitiated (retry affordance; a silent re-drive would just re-fail). Conservative Auto allowlist (anything not clearly transient -> UserInitiated). - RetryDisposition::failure_lane() ties it back to FailureLane; a test asserts the two layers agree for every category in ALL_RUN_FAILURE_CATEGORIES. reborn_composition retry_disposition: 5 passed. clippy + fmt clean. Follow-up: the scheduler wiring that calls retry_disposition() to auto-requeue (the behavior-flipping "Auto" half) — this is its tested decision core. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(reborn): carry secret-scrubbed raw cause to the model Reborn over-sanitized capability/host failures: a real cause like `missing input_schema_ref at /system/extensions/.../list_calendars.input.v1.json` was collapsed to the generic "host runtime rejected capability request" because there was no model-visible field to carry the raw cause and the summary validator rejected any string containing `/`. Policy shift: redact secret VALUES only; let paths, codes, schema refs, and raw error text reach the model so it can retry or explain. Foundation + Tier-1 vertical: - AgentLoopHostError gains an optional model-visible `detail: Option<String>` channel (+ `with_detail`). - CapabilityFailureDetail gains a free-text `Diagnostic { text }` variant. - ToolObservationDetail::GenericFailure gains a bounded, leniently-validated `detail` (allows `/ { } [ ] < >`, rejects NUL/control + caps length) — the channel that already reaches the model and bypasses the strict summary validator. - Relax ONLY the false-positive word bans in validate_loop_safe_summary and validate_tool_result_safe_summary (drop "provider error", "stack trace", "tool input", "traceback", "host path", "raw runtime", "invalid api key"); keep the delimiter ban, control-char ban, length cap, and credential markers. - Tier-1 producers stop dropping the cause: raw_agent_loop_host_error threads the value-scrubbed raw_detail into AgentLoopHostError.detail; the runtime model-visible failure path carries a value-scrubbed Diagnostic when the strict summary validator drops the reason; capability_helpers forwards the diagnostic into the model-visible observation. - Boxed ProviderArgumentError.error to keep result_large_err quiet after the AgentLoopHostError/CapabilityFailureDetail size growth. Tests cover the anchor (path string reaches the model-visible detail), secret value redaction, the relaxed/retained summary markers, and legacy GenericFailure JSON round-trip. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(reborn): MCP per-cause error tokens + explainer detail plumbing (tiers 2a, 3.2) - ironclaw_mcp: replace the flat "response_error"/"request_denied" literals with per-cause diagnostic tokens (mcp_http_status_<code>, mcp_jsonrpc_error code=..., mcp_parse_failed, ...), bounded + control-char-stripped, no public signature change. The model now learns the real HTTP status / JSON-RPC code. - reborn_composition: FailureExplanationInput gains a `detail` field rendered into the failure-explanation prompt (secret-scrubbed via sanitize_model_visible_text). Wired end-to-end in the projection; sourced once TurnLifecycleEvent carries detail (upstream chain in a follow-up commit). mcp --lib 18 passed; reborn_composition --lib failure_explanation tests pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(reborn): thread secret-scrubbed failure detail to the explainer (tiers 2b, 3.1) Complete the model-error-detail chain so the failure explainer (and the model) receive the real cause of a model/provider/driver fault instead of only a sanitized category. Only secret VALUES are withheld (scrubbed via the existing value-level redactors); the descriptive cause now flows end-to-end. Carrier `detail: Option<String>` (serde default + skip_serializing_if, so pre-detail persisted rows rehydrate as None) added and threaded through: - ironclaw_loop_support: HostManagedModelError.detail + with_detail; threaded in model_gateway_error into AgentLoopHostError.detail. - ironclaw_agent_loop: AgentLoopExecutorError::HostUnavailableWithDiagnostics gains detail; model-stage construction carries error.detail. - ironclaw_turns: AgentLoopDriverError::Failed.detail; TurnLifecycleEvent.detail (Failed events only, via failure_detail_for_event in the runner/memory path). - ironclaw_reborn: map_provider_error puts the scrubbed provider reason into HostManagedModelError.detail; planned_driver carries HostUnavailable detail into AgentLoopDriverError::Failed; turn_runner/turn_run_executor carry it into the failure record. - ironclaw_reborn_composition: detail_for_turn_event sources from event.detail, feeding the FailureExplanationInput.detail already rendered in the explainer prompt. Construction-site churn: detail added to TurnLifecycleEvent / AgentLoopDriverError test fixtures and the event_projections pending-gate test support. Verified per crate (--lib): turns 355, loop_support 342, agent_loop 261, reborn 175, event_projections 26 — all green; composition --lib 1042 passed (1 pre-existing live_progress_stream failure, unrelated). turns integration contracts compile; clippy clean across the chain. Pre-existing base-branch breakage in loop_support thread_loop_support_contract (inline_messages) is unrelated to this change. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(turns): allow secret-scrubbed model-visible detail on Failed events The detail channel (TurnLifecycleEvent.detail) intentionally carries a secret-scrubbed description of the real failure cause to the model/explainer; update the guardrail so the spec matches the behavior. Only secret values are withheld; raw unscrubbed backend strings still stay behind host adapters. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): carry detail on AgentLoopDriverError::Failed in integration tests The tier-2b/3.1 detail field on AgentLoopDriverError::Failed broke construction and pattern sites in reborn's integration test targets (concurrent_workers, loop_driver_host) that the original --lib gate never compiled. Constructions get detail: None; the driver_host_error helper carries error.detail; exhaustive match patterns bind detail: _. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): update integration contracts for per-cause MCP tokens + detail channel Two integration test targets asserted pre-refinement model-visible strings that the --lib gate never compiled (test-through-the-caller gap): - mcp_adapter_contract: 5 assertions expected the flat "response_error"/ "request_denied" tokens; update them to the per-cause tokens the Tier 2a change now emits (mcp_invalid_protocol_version, mcp_jsonrpc_id_mismatch, mcp_invalid_session_id, mcp_http_status_500, mcp_denied_credential_source). - llm_gateway: the offline-provider test asserted the error Debug leaked NO provider detail at all. Tier 2b deliberately surfaces the secret-scrubbed non-secret cause on the detail channel. Rewrite (and rename) the test to the current policy: assert the non-secret reason ("connection refused", endpoint URL) reaches the model via `detail`, while the credential token (sk-provider-secret) is scrubbed from both `detail` and the full Debug. This strengthens the secret-scrubbing guard. mcp full suite green; llm_gateway full suite green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(host_runtime): missing first-party handler is InvalidInput, not Backend first_party_missing_handler_fails_closed_without_side_effect_handler asserted the dispatch failure kind was Backend, but nearai#5389 deliberately reclassified an UndeclaredCapability/MethodMissing dispatch failure (a capability the model named that has no registered handler) to InvalidInput — a model-fixable, model-visible tool error that must not burn the retry budget on a call that can never resolve by retrying (see the From<DispatchFailureKind> mapping in production.rs). The test still fails closed (Failed outcome) and still carries "dispatch failed: UndeclaredCapability"; only the kind assertion is updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): LoopFailureKind fault matrix — executor layer + exhaustive category table Phase 1 of the per-error coverage harness (docs/plans/2026-07-03-loop-failure-matrix.md): - turns: all_failure_kinds category table now exhaustive (13/13 — adds the previously-missing CheckpointUnavailable + CompactionUnavailable) with a same-crate exhaustive-match guard so a new variant breaks compilation. - agent_loop: new table-driven executor failure matrix (executor/tests/failure_matrix.rs) driving every executor-reachable LoopFailureKind at its real origin via MockHost seams, asserting per row: P1 (reason_kind + sanitized category/safe_summary), P3 (no fabricated final assistant reply), and explanation_message_refs presence per the explainable set. New fail_transcript_with test knob on MockHost + DriverMockHost (test code only). - Four divergences found and documented (doc §5a), asserted as actual behavior, none silently fixed: Approval+SkipAndContinue completes (gate enforcement gap), NoProgressDetected missing its explanation attach, single Denied recovers-and-completes (no-borking working as designed), TranscriptWriteFailed/CheckpointRejected legacy-only enum origins. Validated (bounded): cargo test -p ironclaw_turns all_failure_kinds; cargo test -p ironclaw_agent_loop failure_matrix; check/clippy -D warnings/ fmt on ironclaw_agent_loop — all green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): LoopFailureKind fault matrix — binary/driver layer rows Phase 2 of the per-error coverage harness (docs/plans/2026-07-03-loop-failure-matrix.md): - planned_driver: resume with missing checkpoint payload asserts LoopFailureKind::CheckpointUnavailable + "checkpoint_unavailable"; in-flight model Cancelled (no cooperative cancel signal) asserts map_executor_error yields "interrupted_unexpectedly". - e2e: binary-level divergence lock — the same in-flight Cancelled run projects "driver_failed" at the runner boundary (category overwritten; doc §5a.5, candidate follow-up to preserve the driver-mapped category). - e2e: non-model P4 row — capability-stage invocation failure is retryable ("host_stage_unavailable_capability", checkpoint preserved, no fabricated reply) and retry_run resumes to completion, so P4 is no longer proven only through the model stage. New scripted capability-invocation-error mode in the test harness (test support only). Validated (bounded): targeted planned_driver tests + full reborn_failure_retry_resume_e2e (15 passed) + clippy -D warnings + fmt. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): reconcile fault matrix to the recoverability stack (nearai#5389/nearai#5390) This matrix PR sits on top of the recoverability stack (main←nearai#4841←nearai#5389←nearai#5390←nearai#5403), so it validates the FIXED + classified system rather than nearai#4841's base behavior: - Add executor rows for nearai#5389's model-fixable capability failures that are now RECOVERABLE (InvalidInput / InvalidOutput / PolicyDenied): where the base branch terminated the run, the stack now surfaces a model-visible tool error and the loop completes. Rows assert the recovered outcome. - Add FailureLane / RetryDisposition alignment (binary/e2e layer, since ironclaw_agent_loop can't depend on ironclaw_reborn_composition): each real failure path's (category, retryable) is asserted to map to the expected nearai#5390 FailureLane bucket + RetryDisposition — proving the real paths feed the classifier correctly. Complements nearai#5390's classifier unit tests (which test the functions directly) rather than duplicating them. - Doc: §6 relationship to nearai#5390; §5a marks divergences RESOLVED by the stack (capability recoverability) vs still-open (Approval SkipAndContinue completes; NoProgressDetected lacks explanation; Transcript/Checkpoint are planned-executor host errors; interrupted_unexpectedly projected as driver_failed at the runner boundary). The cherry-picked base matrix assertions still passed unchanged on the stack — they assert actual behavior, so the additive fixes did not break them; this commit adds the stack-specific coverage on top. Validated (bounded, on the stack): cargo test ironclaw_turns all_failure_kinds; cargo test ironclaw_agent_loop failure_matrix; cargo test --test reborn_failure_retry_resume_e2e (19 passed); clippy -D warnings; fmt --check. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): address gemini failure matrix comments (nearai#5613) * style: cargo fmt on batch-1 recoverable-error changes Reproduces the stack's skipped fmt commit (c19db45) against the restacked base. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tests): restack reconciliation — retry_run stub + superseded IssueCode import - webui_v2_router_smoke's MinimalWebuiServices gained the retry_run rejecting stub the trait now requires (base nearai#4841 break: nearai#5633's smoke fake predates nearai#4841's retry_run addition; every other RebornServicesApi fake already has it). - drop CapabilityInputIssueCode from the ironclaw_turns re-export and ironclaw_loop_support import: the restacked base's CapabilityInputIssue carries DispatchInputIssueCode instead, and the stack's name was import-only on this branch. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(restack): outbound set-target routes through outbound_delivery_outcome; MCP tests use Option error info - set-target handler: replace the superseded nearai#5445 NotFound special-case + outbound_delivery_host_error (deleted by the recoverability batch) with the outbound_delivery_outcome disposition the stack pins in unit tests; matches the list handler. - parse_mcp_response framing tests (base-side, written against the old `error: bool`): assert against the stack's richer `Option<JsonRpcErrorInfo>` — same intent, error presence still pinned. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(restack): turn_stream_auth fake event carries the new detail field Base-side projection test predates the stack's TurnLifecycleEvent.detail addition; None matches the auth-gate fixture's intent. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): align failure_category_demasked pin to the fidelity taxonomy TraceLlm exhaustion (gateway cannot serve the call) maps to ModelErrorClass::Unavailable -> "model_unavailable" under the batch-2 provider-error fidelity mapping; the scenario's "model_error" pin predated it. The scenario's intent — the de-masked TRUE category survives, never the "driver_protocol_violation" sentinel — is unchanged and still asserted exactly. Capability-surface direction (not a security boundary): both categories are classified, retriable lanes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(agent_loop): run fault-matrix rows on 16MiB threads The restacked executor's future (failure explanations + digests + the recovery re-entry) outgrows the 2MiB default test-thread stack in debug on the in-run recovery rows. Production loop threads run 8MiB stacks (ironclaw_reborn_cli serve runtime); mirror the repo's big-stack test-thread pattern (traces/tests.rs, process_port.rs) per row. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(host_runtime): egress contract pins the per-cause MCP denial token The SecretStoreLease-over-production-egress denial now surfaces mcp_denied_credential_source (McpRequestDeniedCause::DeniedCredentialSource) instead of the flat request_denied. Deny-before-transport is unchanged and still asserted (zero recorded requests); ironclaw_mcp's own adapter contract pins the same token. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(review): strip model-visible failure detail from public run-state + align feature-gated tests Security (IronLoop HIGH): RebornGetRunStateResponse forwarded SanitizedFailure.detail — free-form, model-visible backend cause text, scrubbed only for secret VALUES — straight to the browser. Add SanitizedFailure::public_projection() (keeps category, drops detail) and project the public WebUI shape through it. Regression tests: the strip helper (status.rs) and the caller (get_run_state contract test now programs a detail-bearing failure and asserts the DTO omits it). Inherent to the stack, not the reconciliation (original top-of-stack forwarded it raw too). Feature-gated test alignments (only run under --all-features/libsql, so the default-feature local sweep missed them; CI crate buckets caught them): - factory web-access: missing first-party handler is InvalidInput, not Backend (nearai#5389 reclassify; capability still fails closed, only disposition changed). - outbound local_dev (x2): set-target routes through outbound_delivery_outcome (matches original top-of-stack), so the missing-target summary is the fixed "invalid outbound delivery request"; error_kind stays recoverable InvalidInput. - ironclaw_mcp parse_mcp_response_rejects_empty: per-cause tokens replace the flat "response_error" (mcp_parse_failed / mcp_no_payload). Also: cargo fmt (capability_port import reflow from the restack edit). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(review): drop untrusted MCP server error message from model-visible reason + align tool_call category Security (IronLoop HIGH): parse_json_rpc_error_info copied the untrusted MCP server JSON-RPC error.message into the McpClientError::Client reason (model-visible, 'stable sanitized reason') with only length-bounding — not redaction. MCP servers can echo request args, paths, provider diagnostics, or credential-shaped values. Remove message end-to-end (JsonRpcErrorInfo field, parse, cause variant, render); keep only the standardized protocol code=<n>, which is the safe diagnostic. No redaction util is reachable from this leaf crate, and the stable-reason surface should carry stable tokens, not free text. Regression: the former 'reason carries message' test now asserts the message does NOT leak. Feature-gated integration test (ran only under --features libsql): tests/integration/tool_call.rs disabled-spawn-subagent-called-anyway now asserts 'model_unavailable' (InvalidOutput -> Unavailable fidelity category), not the stale 'model_error'. The security property (disabled capability never dispatched) is unchanged and still asserted. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(reborn): cancel-path provider error is model_context_overflow, not model_error fail_model() -> ErrLlm -> LlmError::ContextLengthExceeded, which the batch-2 provider fidelity mapping now categorizes as the accurate model_context_overflow (was the generic model_error). Both cancel tests still pin the load-bearing behavior: reaches Failed after bounded context-shrink recovery (no retry-forever), and the per-thread busy lock releases on Failed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(merge): implement retry_turn on TurnCoordinator doubles added by main Merging origin/main brought two new TurnCoordinator test doubles (UnusedTurnCoordinator in src/runtime.rs, SpyTurnCoordinator in tests/runtime.rs) that predate this stack's retry_turn addition to the TurnCoordinator trait. Add the impls (unimplemented!/delegate, matching each double's existing method style) so the merged tree compiles. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(deny): ignore RUSTSEC-2026-0204 (crossbeam Debug-fmt invalid deref) Newly published advisory (after this branch and main), transitive via crossbeam-epoch. The affected path is the `fmt::Pointer`/`Debug` impl for `Atomic`/`Shared` when the pointer is already invalid — a formatting path we do not exercise. Ignore with justification per the existing advisories convention; remove when the fixed crossbeam-utils release propagates. Verified `cargo deny check advisories` = ok locally. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(reborn): preserve scrubbed failure detail into TurnRunExecutorError (IronLoop) The driver-failed Err path in execute_claimed_run converted the computed SanitizedFailure back to TurnRunExecutorError::new(category), dropping the scrubbed model-visible detail. The scheduler records error.failure(), so production driver failures persisted only the category and TurnLifecycleEvent.detail stayed empty — the failure explainer got the fallback summary instead of the real provider/model cause. Add TurnRunExecutorError::from_failure(SanitizedFailure) (the struct already holds a full SanitizedFailure) and use it at the call site so detail survives across the host-runtime boundary. Caller regression test: driver Failed{detail: Some(..)} -> execute_claimed_run -> err.failure().detail() is preserved. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(reborn): JSON-frame untrusted failure detail in the explainer prompt (IronLoop) detail is untrusted provider/tool/runtime error text (e.g. MCP server or provider bodies). sanitize_model_visible_text redacts credential tokens but keeps newlines/instructions, so appending it raw let a crafted error inject extra prompt fields or directives (a fake fallback_summary:, an 'ignore previous instructions') into the failure explainer — whose output becomes the public failure_summary. That is a prompt-injection path into user-visible messaging, widened by the detail-preservation fix. Frame detail as data: JSON-string-escape it so newlines/quotes are escaped and it stays a single quoted 'detail: "..."' value. failure_category and fallback_summary are host-authored (category-derived) and unchanged. Regression test: a detail embedding newline+fallback_summary+directive is neutralized (exactly one real fallback_summary line, no directive line, detail present as an escaped quoted value). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jul 18, 2026
… inspection (nearai#5280) * docs: spec for Trace Commons instance enrollment, profiles, and trace inspection Cross-repo design (ironclaw + trace-commons-server) for three coexisting capabilities: instance-wide enrollment, per-user contributor accounts via login-links, and submitted-trace inspection. Introduces a trace-credential resolver so the existing user-invite model and the new instance-wide model both function on one instance, with personal-invite enrollment taking precedence. Server change is additive (optional per-user subject through claim issuance + login-link + account resolution). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: Slice 0 plan — trace-commons-server per-user subject TDD plan for the one server change the whole effort depends on: accept an optional opaque subject in the upload-claim request and derive a per-user, tenant-namespaced principal at device-key issuance. Submission attribution, login-link account resolution, and trace readback all become per-user automatically from the shared bearer principal; absent subject reproduces today's behavior. Targets trace-commons-server (contributor-account-slice1). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: IronClaw plans for Trace Commons slices 1-4 Slice 1: trace-credential resolver (personal-invite wins, instance fallback with per-user subject) + admin-gated instance enrollment. Slice 2: per-user subject plumbing through upload-claim request + submission. Slice 3: trace_commons.account_login_link first-party capability (profiles). Slice 4: per-user submitted-trace inspection across reborn_traces → product_workflow facade → webui_v2 handler → frontend. Each plan is bite-sized TDD against verbatim-extracted current code. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): trace-credential resolver (personal invite wins, instance fallback w/ subject) * refactor(traces): single dir-parameterized policy-path site (remove resolver duplication) Extract `trace_contribution_dir_for_scope_at`, `trace_policy_path_at`, `read_trace_policy_for_scope_at`, and `write_trace_policy_for_scope_at` as the canonical base-dir-parameterized path helpers. All public functions (`trace_contribution_dir_for_scope`, `read_trace_policy_for_scope`, `write_trace_policy_for_scope`) now delegate to the `_at` variants with `ironclaw_base_dir()` — signatures unchanged. The inline `read_policy` closure in `resolve_trace_credentials_at` that re-implemented path layout is deleted; it now calls `read_trace_policy_for_scope_at` directly. The test `write_policy_at` helper's bespoke path construction is replaced with a call to `write_trace_policy_for_scope_at`. The now-dead `trace_policy_path` function is removed. Path layout is encoded in exactly one place. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): instance-level enrollment write path (scope None) * test(traces): make instance-enrollment test hermetic (tempdir, no global base) Rework `instance_onboard_writes_instance_level_policy` to operate entirely under a `tempfile::tempdir()`: - Compute instance_dir as base.path().join("trace_contributions") (scope=None layout, no users/<hash> segment) rather than calling the global LazyLock. - Call `onboard_at_dir_with_sink` directly against the tempdir so the test never touches the real ~/.ironclaw tree. - Assert policy.json by reading and deserializing it from the tempdir. - Remove all manual std::fs::remove_* cleanup lines; tempdir drops automatically. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(admin): AdminScope::enroll_instance_trace_commons (admin-gated instance enrollment) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): carry optional per-user subject in upload-claim request Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): thread resolver subject into submission claim context * test(traces): claim request carries per-user subject end-to-end * feat(traces): mint_account_login_link_via_sink (POST /v1/account/login-links) Add `mint_account_login_link_via_sink` to ironclaw_reborn_traces: - `TraceUploadClaimContext::for_account(subject)` constructor for account-management call contexts (no trace/submission ids, no consent scopes). - `AccountLoginLink { account_id, url }` return type. - `account_login_links_url(policy)` helper that derives the login-links URL from the upload-claim issuer URL (strip /v1/trace-upload-claim, append /v1/account/login-links). - `mint_account_login_link_inner(base_dir, ...)` private dir-parameterised core: resolves credentials, selects correct scope_dir for DeviceKey auth (instance enrollment → instance scope dir; personal → user scope dir), mints bearer, POSTs subject, parses response. - `mint_account_login_link_via_sink(tenant_id, user_id, sink)` public entry point wrapping the inner function with the real base dir. Tests (hermetic, tempdir-isolated): - `mint_account_login_link_posts_subject_and_returns_url`: verifies the posted subject equals `local_pseudonymous_contributor_id(trace_scope_key(...))` for instance-enrolled users via an axum mock serving both the upload-claim issuer and the login-links endpoint. - `mint_account_login_link_errors_when_not_enrolled`: verifies error path. - `ReqwestContributionSink` test helper added to the test module. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(traces): error instead of silent misroute in account_login_links_url Replace the unwrap_or_else fallback (which silently used the full issuer URL as a base when the /v1/trace-upload-claim suffix was absent) with an explicit anyhow error. Add two unit tests: one asserting an Err on a wrong-suffix URL, one asserting the correct .../v1/account/login-links URL on a valid issuer. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(host_runtime): add consent-gated trace_commons.account_login_link capability Mints a Trace Commons browser login URL via host network egress, mirroring dispatch_profile_token. Includes consent gate, enrollment pre-check, HostEgressContributionSink routing, and two e2e tests. Also fixes a sanitizer bug: validate_runtime_request was rejecting authorization headers on all requests, including RuntimeKind::FirstParty. FirstParty requests are host-internal and trusted to carry bearer tokens; the sensitive-header and manual-credentials guards now only apply to untrusted plugin runtimes (WASM/MCP/Script). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(host_runtime): route trace bearer via credential injection; restore FirstParty sensitive-header guard Commit 9e25d99 blanket-exempted all RuntimeKind::FirstParty requests from the egress sensitive-header and manual-credentials guards so the host-minted Trace Commons bearer could pass. builtin.http is also FirstParty but forwards model-supplied headers, so this let the model smuggle Authorization/Cookie/ x-api-key headers (or user:pass@ URLs) to allowlisted hosts. Revert the sanitize.rs exemption (guards now apply to ALL runtimes again) and deliver the trace bearer through the staged credential-injection path instead: the HostEgressContributionSink stages the minted token one-shot via RuntimeSecretMaterialStager and declares a StagedObligation Authorization-header injection, mirroring the SlackProtocolHttpEgress pattern. The stager is now exposed to first-party handlers via InvocationServices. Covers the profile_token, profile_set/community-profile, and account_login_link bearer paths. Regression tests: FirstParty + raw authorization header -> denied; FirstParty + user:pass@ URL -> denied. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): fetch_account_traces_via_sink (GET /v1/account/traces, per-user) - Add ContributionHttpMethod::Get variant; update all exhaustive match sites in ironclaw_reborn_traces and HostEgressContributionSink in ironclaw_host_runtime. - Extract account_api_base_url() shared helper; account_login_links_url and new account_traces_url both delegate to it (DRY). - Add AccountTraceItem (Debug, Clone, Serialize, Deserialize; serde defaults). - Add fetch_account_traces_via_sink / fetch_account_traces_inner mirroring mint_account_login_link pattern: unenrolled -> Ok(vec![]), non-2xx -> Ok(vec![]), transport error -> Err. - Tests: hermetic axum mock (GET /v1/account/traces), unenrolled empty-list, URL shape with/without limit. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(reborn): trace_account_traces facade method + wire types Adds RebornAccountTrace / RebornAccountTracesResponse wire types and a trace_account_traces default method on RebornServicesApi, mirroring the trace_credits egress pattern (crate-local hardened reqwest, no host-egress sink). Also adds fetch_account_traces (direct path) to ironclaw_reborn_traces::contribution so the facade can fetch server traces without coupling to RuntimeHttpEgress. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(reborn): GET /api/webchat/v2/traces/account handler + contract test * feat(reborn-ui): render submitted Trace Commons traces in settings Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(traces): document flush-gate limitation, hermetic account-traces contract test, annotate sink scaffold Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * style(traces): cargo fmt across Trace Commons slice changes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): resolver-aware flush gate (instance-enrolled users can contribute) The autonomous trace-flush gate read only the per-scope (personal-invite) policy and aborted when it was disabled, so instance-enrolled users (whose enrollment lives at scope None) could never contribute traces — and the per-user scope_dir would also fail to load the instance device key. Introduce a single EffectiveFlushTarget resolver (resolve_effective_flush_target, mirroring resolve_trace_credentials but keyed on the already-composed scope string) that returns the policy, device-key dir, and per-user subject in one policy-read/path pass. The flush gate now proceeds for instance-only enrollment, loads the device key from the instance (None) dir, and attributes uploads via the per-user pseudonymous subject. The redundant subject_for_scope helper (which re-read the same policies with silent .ok() error swallowing) is removed and its logic folded into the new helper with proper error propagation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): include per-user subject in upload-claim cache key Under instance enrollment every user shares the same instance device-key dir (scope None), so the upload-claim cache key — which keyed on scope_dir but not subject — collided across users. A bearer minted for one subject could be served from cache to another, mis-attributing traces / leaking across users. Add a hashed subject component to the DeviceKey cache key and a regression test proving two subjects sharing a scope_dir get distinct keys (and a no-subject personal-invite context stays distinct from both). Found by Codex review of PR nearai#5280. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): address CodeRabbit review on PR nearai#5280 - account_login_link manifest: declare ReadFilesystem effect (it reads local enrollment/policy/device-key state before egress), matching profile_token. (CR #2) - account-traces fetch: always send a bounded, clamped limit ([1, 500], default 200) so None never triggers an unbounded server history fetch. (CR #3) - direct fetch path: bound the response body with a hard byte ceiling (256 KiB) via a chunked bounded reader, instead of buffering unbounded. (CR #5) - account-traces fetch (both sink + direct): stop swallowing every non-2xx as an empty list — 404 = legitimate empty (no account yet), all other non-2xx surface as Err so the WebUI renders a sanitized unavailable state. Add regression tests (500 -> err, 404 -> empty). (CR #6) - trace-commons-tab.js: render missing final_credit as "—" not "0.00"; surface useAccountTraces() query errors instead of collapsing them to "no traces". (CR nearai#7, nearai#8) - handlers contract test: capture the forwarded caller in the trace_account_traces stub and assert the route threads the authenticated user id (test-through-the-caller). (CR nearai#9) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ci): cover trace_commons.account_login_link + backfill trace i18n keys PR nearai#5280 added the builtin.trace_commons.account_login_link capability and a submitted-traces UI section, but left three guardrail/parity tests un-updated, turning CI red: - ironclaw_host_runtime builtin_first_party_package_declares_expected_capabilities: register account_login_link in the expected id list and its Ask-permission arm. - reborn_builtin_first_party_capability_e2e_coverage_is_complete: add genuine e2e coverage by exercising account_login_link in the existing trace_commons parity test (confirmed=true on a not-enrolled scope returns a deterministic NotEnrolled, no network), grant it in the harness allow-set, and add it to the model-visible surface test and the covered-capability list. - ironclaw_webui_v2_static all_locales_share_the_en_key_set: backfill the six new traceCommons.* submitted-traces keys into all ten non-en locales. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): address CodeRabbit review — withhold login URL, type errors, wire i18n - Security (host_runtime): dispatch_account_login_link returned the one-time login `url` (a code-bearing account-access credential) on the model-visible surface, persisting it into the LLM transcript and any downstream logging. Follow the profile_token pattern: persist the URL to a 0600 private file (atomic temp+rename) and return an opaque `link_delivery` marker instead. E2e test now asserts the URL/code never appears in the result and is delivered out-of-band to the private file. - Typed error (product_workflow): account_traces_for_user flattened backend errors into String before the WebUI boundary. Introduce AccountTracesError (thiserror) that names the failing operation and preserves the full cause chain ({:#}); the boundary keeps returning a sanitized, diagnosable 500. Also document that fetch_account_traces(None) is already server-bounded (default 200, clamp 500, 256 KiB response cap) — the "unbounded fetch" concern was resolved by prior hardening. - i18n (webui_v2_static): the traceStatus and traceReceivedAt keys backfilled for locale parity were unused by the consumer. Wire traceStatus as the status badge's accessible title/aria-label and render traceReceivedAt as the timestamp label, so all six submitted-traces keys are now consumed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): address CodeRabbit re-review — async persist + typed identifiers - Blocking I/O (host_runtime): persist_account_login_link does mkdir/write/ fsync/rename with std::fs on the async dispatch path. Wrap the persist call in tokio::task::spawn_blocking so it never stalls a Tokio worker (coding guideline: all I/O async). Atomic temp+rename behavior is unchanged; a join failure maps to the same sanitized "could not write" result. - Typed identifiers (product_workflow): account_traces_for_user took bare &str tenant/user; the caller already holds TenantId/UserId newtypes. Take &TenantId/&UserId and only cross to &str at the ironclaw_reborn_traces boundary. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): instance-aware enrollment across trace_commons dispatch + UI Instance-only-enrolled users (admin-provisioned instance policy, no personal invite) were falsely rejected across the Trace Commons surface: the dispatch gates and the profile mints read only the personal per-scope policy, and the submitted-traces UI was gated behind the personal-credits branch. Addresses CodeRabbit re-review (#3, #4, #5) on PR nearai#5280. reborn_traces: - Add instance-aware entry points mint_profile_attribution_token_for_user_via_sink and set_community_profile_for_user_via_sink that resolve enrollment via resolve_trace_credentials (personal OR instance) and build the claim context with the instance scope_dir + per-user pseudonymous subject, mirroring mint_account_login_link_inner. Refactor the token mint to share a context-based core. New tests assert the per-user subject reaches the issuer. host_runtime (trace_commons dispatch): - Route the enrollment gates in dispatch_status, dispatch_profile_token, dispatch_profile_set, and dispatch_account_login_link through resolve_trace_credentials so instance-only contributors pass. status now reports the resolved (instance or personal) policy. profile_token/profile_set call the new instance-aware mints. - #4: preserve the stage_secret_material_once failure cause (log it) instead of discarding it with map_err(|_|); wire message stays sanitized. webui_v2_static (#3): - Lift the submitted-traces section out of the credits/empty-state branch so instance-enrolled users with no personal credits still see their traces and any tracesQuery errors. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(traces): isolated dispatch-layer e2e for instance-only enrollment Extract the trace_commons dispatch e2e helpers into a shared tests/support/trace_commons_dispatch.rs module (base-dir setup, mock issuer, runtime/dispatch helpers, find_persisted_login_link, test_jwt_eddsa) so a second test binary can reuse them. Add trace_commons_instance_dispatch_e2e.rs — a SEPARATE binary (fresh process = private IRONCLAW_BASE_DIR) that provisions the process-global instance policy (scope None) without bleeding into the personal-invite suite. It pins the CodeRabbit #5 fix at the layer it manifests: an instance-only-enrolled user (no personal invite) passes dispatch_status and dispatch_account_login_link and mints under the shared instance device key with a per-user pseudonymous subject (asserted via the subject on the login-links POST). No production changes; trace_commons_dispatch_e2e.rs behavior is unchanged (5 tests still pass) — only its helpers moved to the shared module. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): sanitize bearer-staging log + typed IDs on mint entry points Addresses CodeRabbit overnight review on PR nearai#5280. - Security (#1): the trace-bearer staging error was debug-logged via `?error`, which can leak secret-store/backend detail on the credential path. The host_runtime logging guideline forbids backend error detail here — log only the safe fact of failure; the wire message stays sanitized. (Supersedes the earlier "preserve cause" change specifically on this bearer-material path.) - Typed identities (#2): the three agent-facing Trace Commons mint entry points (mint_account_login_link_via_sink, mint_profile_attribution_token_for_user_via_sink, set_community_profile_for_user_via_sink) now take &TenantId/&UserId instead of adjacent &str, so callers can't transpose tenant/user and misattribute a contributor. Identity stays typed to the public boundary and is stringified only when handing off to the dir-parameterised `_inner` cores / resolver (the storage edge). Adds ironclaw_host_api as a reborn_traces dependency (no cycle: host_api does not depend on reborn_traces). Dispatch callers pass the typed scope ids directly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): sanitize persist-path logs, preserve handle-validation cause Two follow-up CodeRabbit findings on PR nearai#5280: - Security (Major): dispatch_account_login_link's spawn_blocking persist arms logged %error / %join_error at debug. Filesystem errors (mkdir/write/fsync/ rename) can carry raw host paths, which the host_runtime guideline forbids in logs. Drop the interpolation; log only the generic fact, keep the message sanitized — same treatment as the bearer-staging path. - Maintainability (Minor): SecretHandle::new(TRACE_COMMONS_BEARER_HANDLE) used map_err(|_| ...), discarding the cause (non-exemptible per the guideline). The handle name is a compile-time constant, so its validation error carries no secret/path — bind and log it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): typed login-link errors, per-request bearer handle, doc accuracy Addresses the third CodeRabbit review round on PR nearai#5280. - Security (Major): the trace-bearer staging used a constant SecretHandle (TRACE_COMMONS_BEARER_HANDLE). The injection store is a HashMap keyed by (scope, capability, handle) with overwrite-on-insert, so two concurrent same-scope Trace Commons egresses could race and stage/consume the wrong bearer. Suffix the handle with a per-request uuid so every staged bearer key is distinct. Localized to the shared HostEgressContributionSink, so all trace_commons flows benefit. - Correctness (Major): account_login_link_error_value classified failures by substring-matching upstream error wording, coupling the public error_code contract to phrasing. Introduce a typed AccountLoginLinkError (thiserror) in reborn_traces; mint_account_login_link_via_sink returns it, producing the specific variant at each failure site. The host maps variants -> error_code with no substring checks. NotEnrolled (the only tested code) is preserved; the two bearer-derived codes collapse into EnrollmentIncomplete (both meant "re-run onboarding"), and persist failures get a distinct LocalStateWrite. - Docs (Minor): the persist_account_login_link comments promised 0600 across platforms though only Unix enforces it. Softened to "private local file (0600 on Unix)". Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(traces): type profile_token/profile_set error mappers (systemic) Follow-up to the account_login_link typed-error change: convert the remaining substring-based error mappers so all four trace_commons dispatch flows derive the public error_code contract from typed variants instead of matching upstream error wording. (onboard was already typed via OnboardError.) - reborn_traces: add ProfileAttributionError (shared by the profile_token and profile_set token mints) and CommunityProfileError (profile_set wrapper adding InvalidProfile). mint_profile_attribution_token_for_user_via_sink and set_community_profile_for_user_via_sink now return these; each failure site produces the specific variant (NotEnrolled / PolicyRead / EnrollmentIncomplete / Backend / LocalStateWrite, plus InvalidProfile for profile_set). - host_runtime: profile_token_error_value / profile_set_error_value now match on the typed variants — no error.contains(...) anywhere in the file. NotEnrolled and InvalidProfile (the tested codes) are preserved; the issuer/device/refused substrings collapse into EnrollmentIncomplete, consistent with the account_login_link mapping. Also sanitized the profile_token persist-failure log (host-path leak class), matching the login-link path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): split enrollment precondition from backend in token mints CodeRabbit re-review: collapsing every mint_profile_attribution_token_with_context (and bearer_token) failure into EnrollmentIncomplete mislabels transient transport/status/serde failures as "re-run onboarding". Split the local precondition (upload-claim issuer URL configured) from post-resolution failures: a missing issuer URL maps to EnrollmentIncomplete via an explicit upload_claim_issuer_missing() check (typed, no substring), while the claim mint / bearer fetch / PUT failures now map to Backend. Applied consistently across profile_token, profile_set, and account_login_link so the error_code contract reflects the real failure class. URL-derivation preconditions (ingest/login-links URL) stay EnrollmentIncomplete. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): check login-link URL precondition before minting bearer Fail-closed ordering: the local account_login_links_url derivation ran after bearer_token, so a malformed/absent login-links URL would mint a device-key bearer and hit the issuer before failing. Move that local precondition ahead of all secret/egress work so incomplete enrollment fails closed with no side effects. (profile_token/profile_set already order local preconditions first.) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): make upload-claim cache key match issuer payload exactly The subject cache-key component trimmed/collapsed context.subject, but the DeviceKey issuer request sends it unchanged — so None, Some(""), and whitespace variants could share a cache key while minting different payloads, letting one user's claim be served from cache to another (cross-user trace mis-attribution). This is nearai#5280's per-user-subject cache-key path. Hash the exact optional bytes the request sends (DeviceKey → subject, WorkloadTokenEnv → None) with a None/Some discriminator. Extend the cache-key test with the Some("")-vs-None and whitespace-variant collision cases. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): check response-size cap before growing the buffer Both bounded response readers (upload-claim and account-traces) enforced the hard byte ceiling only after extend_from_slice, so a single oversized chunk could push the buffer past the advertised limit before the error returned. Compute bytes.len() + chunk.len() (checked_add) and validate before appending. Pre-existing pattern (from nearai#4559), fixed here per review. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): fail loud when trace policy cannot be statted read_trace_policy_for_scope_at used Path::exists(), which maps stat/permission errors to false — silently treating an unreadable policy as missing and default-disabled, flipping enrollment/flush behavior. Use try_exists() and propagate the stat error with context; only a confirmed non-existent path returns the not-enrolled default. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): capture traces for instance-only enrolled users Codex P1: capture_turn_trace gated on the per-user scope policy (read_trace_policy_for_scope(Some(scope)) + policy.enabled), so an instance-only enrolled user — whose per-user policy is absent/disabled — had every turn dropped before an envelope was queued, leaving the instance-aware flush gate nothing to submit. The headline instance-enrollment feature never captured for exactly the users it targets. Gate capture on the effective enrollment instead, mirroring the flush gate: add resolve_effective_capture_policy (personal-invite policy if enabled, else the admin-provisioned instance policy at scope None, else None) and prepare the envelope under that governing policy. Add a resolver test covering the personal / instance-only / neither cases. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Remove accidentally committed frontend node_modules, restore .gitignore The merge commit f34dfa7 dropped crates/ironclaw_webui_v2_static/frontend/.gitignore and swept 1065 node_modules files into the index. Untrack them and restore the ignore. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address PR review feedback: egress hardening, effect declaration, instance status sync - fetch_account_traces_direct now uses a pinned-DNS, private-IP-filtered HTTP client (shared pinned_trace_commons_http_client) instead of an unrestricted reqwest lookup, closing the DNS-rebinding window between claim validation and the bearer-authenticated account-traces GET. - account_login_link capability manifest declares EffectKind::WriteFilesystem for the local delivery-file write. - Queue-flush status sync (and the public sync entry point) now run off the resolved effective flush target (policy, device-key dir, per-user subject) instead of re-reading the per-scope policy, so instance-enrolled users get final credit status after submission; subject is threaded into the status-sync claim context. - Each login-link mint persists to a unique account_login_link.<uuid>.url file so concurrent mints cannot clobber each other; stale link files are pruned best-effort after one hour. - resolve_trace_credentials takes typed &TenantId/&UserId at the public boundary; call sites drop their .as_str() conversions. - Login-link/account-traces requests honor the policy-configured issuer timeout; the sink-path traces fetch uses ACCOUNT_TRACES_MAX_RESPONSE_BYTES. - Removed the AdminScope::enroll_instance_trace_commons wrapper from the v1 monolith (crate-side entry point is onboard_instance_with_sink; noted in the slice1 plan). - Tests: direct account-traces path covered for 500/404; new regression test pins instance-target status sync (subject + instance device-key dir). - Plan docs: server login-link contract callout, no developer-local paths, resolver errors propagate, 404-only zero-state, scope_dir threading. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Pin DNS resolution on the background trace submit/status/revoke lane The background lane (queue flush submission, status sync, revocation) previously relied only on enrollment-time endpoint validation (validate_trace_commons_ingest_url); the per-request client did a fresh unrestricted DNS lookup. Replace trace_remote_http_client with pinned_trace_remote_http_client: per-request host resolution through resolve_trace_upload_claim_issuer_host (private/internal IPs rejected, literal-loopback local-dev exception) pinned via resolve_to_addrs, so an endpoint host that passed validation at enrollment cannot later rebind to an internal address and receive bearer-authenticated requests. Timeout behavior (env/test task-local override) is unchanged. Regression test: pinned_trace_remote_client_rejects_private_endpoint_hosts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address CodeRabbit follow-up: sanitize status log, sync plan snippets - trace commons status dispatcher no longer formats the resolver error into the log (it can embed the policy file's host path); logs the safe fact only, matching the sibling dispatchers. - slice4 plan: AccountTraceItem snippet derives Deserialize (matches shipped code, which parses the response). - slice3 plan: login-link parsing snippet fails loud on missing account_id/url instead of unwrap_or_default (matches shipped code). - slice1 plan: the AdminScope wrapper task is marked SUPERSEDED up front so the plan no longer gives conflicting guidance. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address round-2 review: opt-out precedence, salted subjects, UI branch tests - Explicit per-user opt-out (scoped policy present with enabled=false, as written by 'traces opt-out') now blocks the instance-enrollment fallback in resolve_trace_credentials and resolve_effective_flush_target (and thus capture) — only a never-configured scope falls through to the instance policy. Regression test covers all three resolution surfaces. - Instance-enrollment subjects are now salted: a per-instance random salt (persisted 0600 at the instance trace dir, create_new race-safe) feeds sha256(salt:scope), so the server or ledger holders cannot dictionary-match guessable tenant/user ids against an unsalted scope hash. Unsalted local_pseudonymous_contributor_id remains for local state keying/log refs. - contribution.rs carries the architecture-rule file-size justification referencing decomposition tracking issue nearai#4088; state_scope field docs now say which state it does (and does not) locate. - Submitted-traces UI: extracted the pure tracesSectionMode decision (error wins over list; list needs enrolled + non-empty) and covered it plus the row formatters in trace-commons-tab.test.mjs. - Docs: slice4 plan points at crates/ironclaw_webui_v2 (static crate was folded in), slice3 signature snippet matches the typed contract, and the webui_v2 CLAUDE.md route table gains the three trace routes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Update crossbeam-epoch 0.9.18 -> 0.9.20 for RUSTSEC-2026-0204 Lockfile-only patch bump of a transitive dep (via termimad/crossbeam) to clear the new advisory failing cargo-deny; verified locally with cargo deny check advisories. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Route login-link/account-traces claim mint through the caller's sink The sink-based entry points (mint_account_login_link_via_sink, fetch_account_traces_via_sink) used the sink for the final POST/GET but minted the upload-claim bearer via DefaultTraceUploadCredentialProvider, whose issuer request takes the direct reqwest path — so an agent-invoked account_login_link performed a network call outside RuntimeHttpEgress. New trace_upload_bearer_token_via threads Option<sink> into the claim mint (cache behavior unchanged; the default provider passes None), and both sink paths pass Some(sink), matching the profile-token/profile-set flows. Tests now use a RecordingSink to pin the invariant that both the claim mint and the follow-up request route through the sink. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
JZKK720
pushed a commit
that referenced
this pull request
Jul 18, 2026
…earai#6230) The scheduled `Reborn Playwright` suite (nightly, not PR-gated) had 7 deterministic failures in the extensions shard, from three UI changes that merged without e2e coverage running on their PRs: 1. Catalog-error banner (production bug). nearai#6088 keyed the "Extension catalog unavailable" (danger) vs "Some extension data is unavailable" (warning) banner to the *tab* rather than the *failure cause*, so a catalog (registry) failure on the channels/mcp tab showed the partial warning text. Fix: `CatalogErrorBanner` now selects text+tone from which query failed (catalog vs enrichment); blocking-vs-inline placement still follows the tab. Locked by the extensions-page unit test (asserts the page passes `isCatalogError` per cause) and the `catalog_failure_shows_retry` e2e scenario. 2. Remove confirmation (stale test). nearai#6084 replaced the native `window.confirm` with the shared React `ConfirmDialog`, but the 5 extension remove/reinstall e2e tests still waited on a native dialog event that never fires. Replaced `_capture_next_confirm` with `_resolve_remove_confirm`, which drives the modal (`confirm-dialog-confirm`/`-cancel` testids). 3. Telegram configure (stale test). nearai#6159 retired the per-user Telegram bot-token secret form in favor of the WebGeneratedCode pairing panel (its own commit note flags the "stale per-user bot-token telegram tests"). Replaced the obsolete token-entry test with one asserting the configure modal hosts the pairing panel and no longer offers a bot-token form. Only #1 changes production code; #2/#3 align stale tests to shipped, intended behavior. Verified locally: the 7 tests, the full 43-test extensions file, all 812 frontend unit tests, tsc, and the conventions lint all pass. [skip-regression-check] #2/#3 are test-only alignments; #1's regression is the pre-existing catalog_failure e2e scenario plus the extended unit test, both included here. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Auto-generated by release CI. Updates SHA256 checksums and version-pinned artifact URLs in registry manifests to match the released WASM artifacts. Only extensions whose version changed since the last release are included.